Story · arXiv
Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection (arXiv)
paper · Story page
SkillGuard treats the moment untrusted skill output enters an agent's context as contamination and cuts the agent's future capabilities so deployer-defined forbidden states become unreachable. It's a harness-level reference monitor with no extra model call, evaluated on four AgentDojo suites with Gemini 2.5 Flash and Llama3.3-70B.
In plain words
- Researchers created SkillGuard to limit what an artificial intelligence system can do after it reads untrusted information.
- Attackers can place harmful instructions inside outside information, which may influence later actions with greater permissions.
- SkillGuard tracks how each available tool can affect the system, then blocks routes toward actions the deployer forbids.
- It applies these limits immediately without asking another artificial intelligence system to review the decision.
- This could help deployers contain attacks without adding another artificial intelligence check.
Appeared in
- Anthropic's deliberately misaligned model, Fable 5.1, and a fix for reward hacking
Sep 02, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.