Story · arXiv (via papers.cool)
Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite (arXiv (via papers.cool))
arXiv, 25 pp · Story page
Paired same-model runs on a private, contamination-controlled suite found no clear average advantage for vendor-native harnesses over deepagents, with Opus 4.8 or GPT-5.5. The Opus average hides opposite strata on a split chosen after seeing the data.
In plain words
- Researchers found no clear overall advantage for coding assistants using tools from their own supplier.
- They tested the same artificial intelligence on the same tasks while changing its tools and instructions.
- For the tested Opus coding assistant, the supplier's tools did better on contests but worse on existing software projects.
- Researchers chose that comparison after seeing the results, so it needs confirmation in another study.
- People choosing coding assistants cannot assume that the same company's tools will help them solve more tasks.
Appeared in
- A shell beats a typed tool catalog, and agents barely report their work
Sep 16, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.