Story · arXiv
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses (arXiv)
paper · Story page
Eight generated harnesses ran against nine byte-identical copies of one baseline, three executions each, on 386 MATH-500 tasks. They lost persistently on 100 tasks and won persistently on one; the frozen selector gained 0.00 percentage points.
In plain words
- Researchers tested whether changing the software guiding an artificial intelligence system improved its answers to math problems.
- They compared several newly generated programs with repeated runs of the same starting program.
- Generated programs consistently did worse on 100 problems and better on only one, with that win depending on how answers were read.
- For developers, choosing among the generated programs before each problem produced no improvement in this test.
Appeared in
- A forged chat-template marker loses most of its authority as ordinary subwords
Oct 01, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.