Story · arXiv
Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents (arXiv)
paper · Story page
A preprint on indirect prompt injections that stay dormant in retrieved content until an attacker-chosen trigger fires, measured against nine production agents.
In plain words
- Researchers found a way to trick artificial intelligence assistants into following hidden instructions later.
- Attackers hide instructions in material an assistant reads, with orders to wait until a chosen condition is met.
- Successful attacks make the assistant use its tools to carry out the hidden instructions.
- Across nine working assistants, the attacks succeeded in 43% to 83% of trials, compared with at most 3% for direct commands.
- People using these assistants can face unwanted computer actions even when the software rejects an attacker's direct commands.
Appeared in
- Dormant prompt injections land on nine production agents where direct orders fail
Sep 23, 2026 · lead story
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.