Story · arXiv
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents (arXiv)
paper · Story page
Fifteen detectors, Prompt Guard 2 among them, were scored on tool outputs replayed from AgentDojo and tau-bench and on the BIPIA benchmark. Rankings barely survive the move: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate.
In plain words
- Software that spots attempts to trick artificial intelligence often missed them when researchers changed the test.
- The researchers tested instructions hidden inside information that the artificial intelligence would receive while doing tasks.
- The best performer in one test caught only 2% of attacks in another, while wrongly flagging 1% of harmless material.
- Teams choosing safety software cannot assume that a high test score means it will catch attacks in their own tasks.
Appeared in
- Branch steering breaks the Dual-LLM guarantee for computer-use agents
Oct 06, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.