Writing /
Inside four LLM guards: what the weights say about safeguards for agentic systems
I audited four candidate guards for an agentic assistant layer by layer. They all decide in one token and late in the network, and the two dedicated guards miss prompt injection hidden in documents. A small generalist catches almost everything if it writes its justification before the verdict.
An agent that reads email, CVs and invoices needs a first barrier: a model that looks at every input and decides whether to block it. I audited four candidates for that job on the same machine and the same 46 cases, and then looked inside them, layer by layer, to see how each one reaches its verdict.
The candidates: two dedicated guards, Shieldstral 3B and Qwen3Guard-Gen 8B, and two general models used as guards through a written policy, gpt-oss 20B and Gemma 4 E2B. The test set mixes prompts, documents with hidden instructions and images, in Spanish and English, for a corporate assistant with tools.
Parameter counts mislead
gpt-oss is the largest on disk and one of the smallest in compute: 91 % of its weights are experts, and each token uses 4 of 32. Qwen3Guard, the “8B”, computes all 8.2 B every time.
The verdict is always one token
Each model answers “block or not?” differently. Shieldstral reads two logits and never generates. Qwen3Guard writes Safety: Unsafe, and the third token decides. gpt-oss reasons in an analysis channel and then writes JSON. Gemma writes JSON, with or without thinking first.
Underneath, all four reduce to one token whose probability you can read. That turns a fixed “strict / loose” switch into a threshold you set for your own risk tolerance.
The decision forms late, and can be undone
In all four, the verdict appears in the last third of the network. What differs is what happens next. Shieldstral reaches 1.00 and its last block pulls it back to 0.71. Gemma, without reasoning, senses the risk (0.30) and its final layers discard it; with its own reasoning in front, it holds at 1.00 from layer 27 onwards.
What each one catches
The gaps cluster in two places. Documents with indirect prompt injection slip past Qwen3Guard (none of five in its default mode, one in strict mode) and past Shieldstral in two of five: both learned to judge whether text is harmful, not whether it tries to instruct an AI. Neither technical report evaluates indirect injection. Images are invisible to the text-only guards.
Justify before you decide
The first comparison mixed three things: reasoning, forced JSON and field order. Separating them shows the cause. Removing the JSON schema changes nothing. Putting a rationale field before violation lifts F1 from 0.74 to 0.91-0.93 at 0.22 s per input. Full reasoning adds little more (0.94-0.96) at about seven times the latency, and varies between runs.
The runtime is part of the model
Two surprises came from outside the weights. Ollama’s GGUF of Shieldstral shrinks images to about 896 px against 1540 px in the reference, so small text in screenshots loses resolution. And the standard transformers loader silently sets Gemma’s per-layer output scales to 1.0; without correcting it, any internal analysis measures a different model. Auditing a model means auditing how it runs.
What I would deploy
- Shieldstral as a fast first filter (0.04 s per text, no false positives here), not as an injection defence.
- A second, policy-driven check for documents and tool outputs, where injection lives. A small generalist that writes its justification first is cheap and strong.
- Scores, not labels. Read the decision token and set the threshold on your own data.
- A separate path for images and audio, or an explicit decision not to trust them.
Method and limits
Static inventory of every tensor in the GGUF files Ollama serves; reading of technical reports, model cards and reference code; activation probes (logit lens, gradient × input, attention, MoE routes, attention sinks) on the same quantised weights, validated against Ollama’s own outputs; and a controlled experiment with three replicas per variant. With 46 cases, one case moves F1 by 2-4 points: use these results to see patterns, not to rank models by decimals. The system, its policy and the case texts stay private; the method and the numbers are public.
The companion tutorial, Inside an LLM guard, reproduces the core techniques on small models that run on a laptop CPU.