Story · arXiv
How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making (arXiv)
paper · Story page

Across nine models, agent task success follows a geometric law set by one per-step reliability parameter, which rises with model scale but saturates well below 1. On the tool-use task, every model tested fell from near-perfect to near zero within sixteen dependent steps, which the authors argue is why benchmark optimism doesn't survive production horizons.
In plain words
- Researchers found that artificial intelligence systems become much less reliable as jobs require more dependent steps.
- Each step has a chance of failure, so mistakes accumulate as a task grows longer.
- They tested nine systems across four task types, five lengths, and three ways of handling prior information.
- Companies cannot assume strong results on short tests will carry over to long production workflows.
Appeared in
- GPT-6 Astra's score hinges on its harness, and agents rot geometrically with each step
Sep 04, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
