Story · arXiv
What Does a Harness Buy? Tokens, Mostly (arXiv)
paper · Story page
Five models through Claude Code, mini-SWE-agent and OpenCode on SWE-bench Verified, with reruns to calibrate how far a score drifts when nothing changes. The heaviest and lightest harness land within five points on 447 tasks with two models, and on a 45-task hard subset swapping the harness flips as many tasks as rerunning it.
In plain words
- Researchers compared different software setups for artificial intelligence coding assistants.
- They kept the artificial intelligence unchanged while switching the surrounding instructions and tools.
- The most elaborate and simplest setups performed similarly on the larger test, while OpenCode performed worse.
- For developers, an improved score may just reflect the variation researchers saw when repeating the same test.
Appeared in
- Claude Code's Bash tool changed 12% of calls carrying code or escapes
Oct 07, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.