Story · arXiv

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis (arXiv)

paper · Story page

Five easels in a row each hold a copy of the previous easel's drawing, losing detail at every step until the last holds only a fragment; a small robot at the end looks back down the row.

LongDS builds 68 long-horizon data-analysis tasks from real Kaggle notebooks, spanning 2,225 turns. The best of five models reaches 48.45%, accuracy falls nearly 47 points from early to late turns, and extra steps don't necessarily help, suggesting the bottleneck is keeping analytical state correct.

In plain words

  • Researchers tested five artificial intelligence systems on 68 data-analysis tasks that unfolded across many rounds.
  • Each new request could depend on earlier work, requiring the system to preserve, revise, restore, or combine previous analyses.
  • The best system averaged 48.45% accuracy, while performance fell nearly 47 points from early to late rounds.
  • The results suggest extra attempts do not reliably help when the system loses track of the current analysis.
  • People using artificial intelligence for lengthy analysis may need stronger ways to preserve earlier work and decisions.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.