Story · arXiv
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead (arXiv)
paper · Story page
Across 254 submissions, read without running a model, the two leading Verified entries each resolve 396 of 500 instances and exact paired McNemar tests separate none of the 29 adjacent pairs in the top thirty. Within-model scaffold ranges reach 29.8 points against that group's 8.8-point spread.
In plain words
- Researchers checked whether published test scores could reliably rank the leading automated coding tools.
- They compared different versions of the test, checking which programming tasks each tool completed successfully.
- The leading tools often succeeded and failed on the same tasks.
- For people choosing tools, the results could not reliably distinguish neighboring entries in one version's top 30.
Appeared in
- Every audited chat tokenizer lets prompt text forge control tokens
Sep 17, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.