Story · arXiv
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis (arXiv)
paper · Story page

LongDS builds 68 long-horizon data-analysis tasks from real Kaggle notebooks, spanning 2,225 turns. The best of five models reaches 48.45%, accuracy falls nearly 47 points from early to late turns, and extra steps don't necessarily help, suggesting the bottleneck is keeping analytical state correct.
In plain words
- Researchers tested five artificial intelligence systems on 68 data-analysis tasks that unfolded across many rounds.
- Each new request could depend on earlier work, requiring the system to preserve, revise, restore, or combine previous analyses.
- The best system averaged 48.45% accuracy, while performance fell nearly 47 points from early to late rounds.
- The results suggest extra attempts do not reliably help when the system loses track of the current analysis.
- People using artificial intelligence for lengthy analysis may need stronger ways to preserve earlier work and decisions.
Appeared in
- Realistic prompts drop coding-agent scores, and tool filtering beats prompt rules
Sep 01, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
