Story · arXiv

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments (arXiv)

paper · Story page

StarHarness treats the agent harness as the optimization target: prompts, tool interfaces, skills, MCP providers, subagent structure and loop settings evolve while model weights stay fixed. On three enterprise benchmarks it reports 20-35 percentage points over the default after 4-12 accepted changes, with gains holding on held-out tasks and transferring across GPT and Qwen models.

In plain words

  • Researchers improved workplace task performance by changing the surrounding instructions and tools while leaving the underlying artificial intelligence unchanged.
  • They grouped tasks by the default setup's mistakes, tested proposed changes separately, and reserved other tasks for final checks.
  • Changes covered instructions, tool designs, reusable procedures, supporting services, helper organization, and rules for repeating steps.
  • After 4-12 accepted changes, performance on three workplace test suites improved by 20-35 percentage points over the default setup.
  • Businesses may reuse these improvements across unfamiliar tasks and different artificial intelligence systems without changing the systems themselves.

Appeared in

Subscribe

Get the brief in your inbox

Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.

  • Weekdays at 8:45am IST, one lead story and 6 to 9 items.
  • Sundays, an argued synthesis rather than a recap.
  • One click to leave, and quiet days say so in the subject line.
How often

Weekdays 8:45am IST + Sundays. Unsubscribe in one click.

You're asking for The Agentic Brief by email at the cadence you picked. You can unsubscribe in one click from any issue, and your address is never sold or shared.