Inside four LLM guards: what the weights say about safeguards for agentic systems
I audited four candidate guards for an agentic assistant layer by layer. They all decide in one token and late in the network, and the two dedicated guards miss prompt injection hidden in documents. A small generalist catches almost everything if it writes its justification before the verdict.
AI safetyLLM safeguardsPrompt injectionInterpretabilityAgentic AI