Story · arXiv

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents (arXiv)

arXiv, 25 pp · Story page

Five identical small robots each standing in a differently shaped doorway, with a bar of a different height above each doorway.

The authors replay frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent from one Qwen3-8B warm start and score each checkpoint on a sealed SWE-bench Verified oracle. Across 24,000 evaluations the evaluation harness moved mean solve rate from 2.14% to 9.27%, a factor of 4.3, while pooling rewards across harnesses moved the held-out result by 0.25 points with an interval spanning zero.

In plain words

  • An artificial intelligence (AI) coding assistant showed no clear benefit from mixing training feedback when tested in an unfamiliar software setup.
  • Each setup is software that runs the assistant while it works on coding tasks.
  • The researchers compared training that combined success scores across setups with training that kept scores separate for each setup.
  • For developers, changing the software running the assistant affected success much more than mixing feedback in these tests.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.