Story · arXiv

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses (arXiv)

paper · Story page

Eight generated harnesses ran against nine byte-identical copies of one baseline, three executions each, on 386 MATH-500 tasks. They lost persistently on 100 tasks and won persistently on one; the frozen selector gained 0.00 percentage points.

In plain words

  • Researchers tested whether changing the software guiding an artificial intelligence system improved its answers to math problems.
  • They compared several newly generated programs with repeated runs of the same starting program.
  • Generated programs consistently did worse on 100 problems and better on only one, with that win depending on how answers were read.
  • For developers, choosing among the generated programs before each problem produced no improvement in this test.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.