Tutorials / T4
Inside an LLM guard: the verdict token, the logit lens and the injection blind spot
Open a safety guard on a laptop CPU: read its verdict as a probability, watch the decision form layer by layer, and see why a harmful-content guard lets prompt injection hidden in documents straight through.
Why this tutorial
In Inside four LLM guards I audited four safety guards on a GPU with 46 cases. This notebook takes the core techniques down to models that run on a laptop, so you can apply them to your own guard: 0.6 B-parameter versions, 20 synthetic cases, CPU only.
What you do in the notebook
- Find where the policy lives. Qwen3Guard carries its policy inside the chat template; you render it and count its tokens.
- Read the verdict as a probability. The decision is one token (
Safe,UnsafeorControversial). Reading its logits gives a continuous score and a threshold you set yourself. - Run a logit lens, with the rule that matters: only trust layers where the label tokens hold at least 1 % of the probability.
- Test the injection blind spot with emails, CVs and invoices that hide instructions for the AI.
- Prompt a small generalist as a guard and compare writing the verdict first against writing a rationale first.
What it shows
| Finding | Audit (8 B / E2B, GPU) | This notebook (0.6 B, CPU) |
|---|---|---|
| Verdict is one readable token | yes | yes |
| Verdict forms in the last third | layers 24–28 of 36 | layers 19–22 of 28 |
| Injected documents, max P(Unsafe) | 0.014 | 0.0098 |
| Generalist, rationale first vs verdict first | F1 0.74 → 0.91–0.93 | F1 0.74 → 0.50 (worse) |
Three of the four findings carry over. The fourth does not, and the notebook says why: a rationale only helps when the model can reason about the document while writing it. At 0.6 B, the rationales paraphrase the user’s task (“a request to summarise a supplier email…”) and commit the model to a benign reading. The lesson that transfers is the method: measure both formats on your own cases instead of assuming either one.
There is also a weak signal worth knowing: injected documents carry more Controversial probability than their benign twins (median 0.17 against 0.006). The guard notices something is odd without calling it unsafe. Five cases are far too few to build a detector on it.
Takeaways
- A harmful-content guard is not an injection defence. Documents and tool outputs need their own check.
- Read the decision token, not the label, and choose your threshold on your own data.
- In a logit lens, early “confident” values are renormalised noise. Filter by probability mass.
- Prompt format can help or hurt. Test it at the model size you will deploy.
References: Qwen3Guard Technical Report, arXiv 2510.14276; nostalgebraist, interpreting GPT: the logit lens, 2020; Greshake et al., Not what you’ve signed up for, 2023; OWASP Top 10 for LLM Applications, LLM01.