Story · arXiv
OverThink: Slowdown Attacks on Reasoning LLMs (arXiv)
paper · Story page
Decoy reasoning problems injected into external context make a model burn far more reasoning tokens while still answering correctly: 13x on FreshQA, 46x on SQuAD, 12x on MuSR. Tested injection points for coding agents include skills, README files and code, and this is a revision of an older paper.
In plain words
- Researchers found a way to make artificial intelligence assistants do much more unnecessary work while still answering correctly.
- Attackers hide distracting puzzles in material the assistant reads while working on a task.
- The assistant spends extra effort solving those puzzles, increasing the computing work needed for its answer.
- For people running these assistants, correct answers can still come with inflated computing costs.
Appeared in
- Approve one operation, run another: the binding failure in shipped agent products
Sep 22, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.