Working papers · speech, health & trustworthy AIVigo, 2026

Writing

Essays and analyses on speech, health and the safety of generative AI. Each one links to the tutorial or papers behind it.

Inside four LLM guards: what the weights say about safeguards for agentic systems

I audited four candidate guards for an agentic assistant layer by layer. They all decide in one token and late in the network, and the two dedicated guards miss prompt injection hidden in documents. A small generalist catches almost everything if it writes its justification before the verdict.

AI safetyLLM safeguardsPrompt injectionInterpretabilityAgentic AI

How LLM guards decide: a tour inside four safety models

Where the policy lives, which layers read it, what the router does and why the last block can overturn a verdict: a tour through the internals of four LLM guards, measured on the weights they actually run.

AI safetyLLM safeguardsInterpretabilityTransformersMixture of experts