Story · Semantic Scholar
Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication (Semantic Scholar)
paper · Story page
Four defenses and an undefended control ran on the AgentDojo banking benchmark across GPT-5.4, GPT-5.4-mini, and Claude Sonnet 4.6, with two independent replications (Tool Filter was tested only on the OpenAI models). On GPT-5.4-mini two defenses were associated with lower attack rates and lower benign utility, and none of the four paired comparisons survived Holm correction.
In plain words
- Researchers tested ways to protect artificial intelligence (AI) assistants from malicious instructions hidden in material they read.
- The attacks try to make an assistant follow instructions in outside content instead of its assigned task.
- For one tested system, some protections were linked to fewer successful attacks but worse performance on ordinary tasks.
- For people choosing protections, the apparent security improvements remained uncertain after researchers adjusted their calculations for testing several options.
Appeared in
- Prefix caching changes agent runs, and cross-family reviewers beat self-review
Sep 08, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.