Story · arXiv

ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time (arXiv)

paper · Story page

ArrivalBench re-runs the pipeline an agent leaves behind against late, duplicated, out-of-order and retried records, then requires the final table to match a full recomputation of the log. Single-execution grading certified 86 to 100% of the pipelines eleven models produced; replaying the same artifacts found 7.0 to 79.2% of those silently wrong.

In plain words

  • Researchers found that programs written by artificial intelligence could pass a test yet mishandle records arriving over time.
  • They reran the programs with records arriving late, arriving more than once, or arriving in the wrong order.
  • They checked the final results against a fresh calculation using the complete record of what had happened.
  • For teams using these programs, a successful trial can hide mistakes that appear later without the program crashing.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.