Story · arXiv
What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents (arXiv)
arXiv, 25 pp · Story page

The authors replay frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent from one Qwen3-8B warm start and score each checkpoint on a sealed SWE-bench Verified oracle. Across 24,000 evaluations the evaluation harness moved mean solve rate from 2.14% to 9.27%, a factor of 4.3, while pooling rewards across harnesses moved the held-out result by 0.25 points with an interval spanning zero.
In plain words
- An artificial intelligence (AI) coding assistant showed no clear benefit from mixing training feedback when tested in an unfamiliar software setup.
- Each setup is software that runs the assistant while it works on coding tasks.
- The researchers compared training that combined success scores across setups with training that kept scores separate for each setup.
- For developers, changing the software running the assistant affected success much more than mixing feedback in these tests.
Appeared in
- Prefix caching changes agent runs, and cross-family reviewers beat self-review
Sep 08, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
