Story · arXiv
ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time (arXiv)
paper · Story page
ArrivalBench re-runs the pipeline an agent leaves behind against late, duplicated, out-of-order and retried records, then requires the final table to match a full recomputation of the log. Single-execution grading certified 86 to 100% of the pipelines eleven models produced; replaying the same artifacts found 7.0 to 79.2% of those silently wrong.
In plain words
- Researchers found that programs written by artificial intelligence could pass a test yet mishandle records arriving over time.
- They reran the programs with records arriving late, arriving more than once, or arriving in the wrong order.
- They checked the final results against a fresh calculation using the complete record of what had happened.
- For teams using these programs, a successful trial can hide mistakes that appear later without the program crashing.
Appeared in
- Branch steering breaks the Dual-LLM guarantee for computer-use agents
Oct 06, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.