Story · arXiv

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making (arXiv)

paper · Story page

A staircase of sixteen steps. Many small tokens crowd the lowest step, fewer stand on each higher step, and the top steps are empty. A thin curve above traces the decline.

Across nine models, agent task success follows a geometric law set by one per-step reliability parameter, which rises with model scale but saturates well below 1. On the tool-use task, every model tested fell from near-perfect to near zero within sixteen dependent steps, which the authors argue is why benchmark optimism doesn't survive production horizons.

In plain words

  • Researchers found that artificial intelligence systems become much less reliable as jobs require more dependent steps.
  • Each step has a chance of failure, so mistakes accumulate as a task grows longer.
  • They tested nine systems across four task types, five lengths, and three ways of handling prior information.
  • Companies cannot assume strong results on short tests will carry over to long production workflows.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.