Story · arXiv
Recognition Without Enforcement: Configuration-Dependent Failures in LLM Agent Instruction Arbitration and External Control (arXiv)
paper · Story page
Agents can identify forged authority when asked, yet certain prompt-model pairings still execute the conflicting tool call. Average execution under novel spoofing attacks is 1.21%; the gap between recognizing and refusing is configuration-dependent, not an immutable property of model weights.
In plain words
- Researchers found that some action-taking artificial intelligence systems recognized fake authority but still followed its conflicting command.
- They tested different system-and-instruction pairings to see when the conflicting action occurred.
- Restrictive rules and varied wording prevented failures in the same systems, while permissive combinations caused repeatable failures.
- Organizations cannot rely on recognition alone because surrounding instructions and policies determine whether suspicious commands are actually blocked.
Appeared in
- Realistic prompts drop coding-agent scores, and tool filtering beats prompt rules
Sep 01, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.