Story · arXiv
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments (arXiv)
paper · Story page
StarHarness treats the agent harness as the optimization target: prompts, tool interfaces, skills, MCP providers, subagent structure and loop settings evolve while model weights stay fixed. On three enterprise benchmarks it reports 20-35 percentage points over the default after 4-12 accepted changes, with gains holding on held-out tasks and transferring across GPT and Qwen models.
In plain words
- Researchers improved workplace task performance by changing the surrounding instructions and tools while leaving the underlying artificial intelligence unchanged.
- They grouped tasks by the default setup's mistakes, tested proposed changes separately, and reserved other tasks for final checks.
- Changes covered instructions, tool designs, reusable procedures, supporting services, helper organization, and rules for repeating steps.
- After 4-12 accepted changes, performance on three workplace test suites improved by 20-35 percentage points over the default setup.
- Businesses may reuse these improvements across unfamiliar tasks and different artificial intelligence systems without changing the systems themselves.
Appeared in
- A website summary hijacks Claude Code Auto Mode, and handoffs turn must into maybe
Aug 28, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.