Story · arXiv
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus (arXiv)
paper · Story page
Only 78 of the 125 all-fail tasks in a frozen Terminal-Bench 3 and Frontier-Bench 0.1 record survive as certified-unsolved candidates. The rest involve broken oracles, infrastructure failures, verifier bypasses, or tasks with insufficient evidence of solvability.
In plain words
- Researchers found that many failed artificial intelligence tests had problems with the tests themselves.
- They checked whether the tasks could be solved and whether the systems judging answers worked properly.
- Only 78 of 125 tasks with no successful solution appeared to be fair tests that the assistants still could not pass.
- People comparing artificial intelligence systems could mistake a broken test for evidence that an assistant lacks ability.
Appeared in
- Forged control tokens blank an agent's reasoning while the tool call still fires
Sep 25, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.