Story · arXiv
Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents (arXiv)
paper · Story page
Across 1,200 Claude Code and Mini-SWE-Agent trajectories on SWE-bench Verified, subsumed retrieval, similar script generation and test re-execution affect 79.00% to 98.00% of tasks and account for up to 22.75% of task cost. Structure-aware retrieval, the obvious fix, added overhead and raised costs by up to 28.14%.
In plain words
- Researchers found that artificial intelligence coding tools repeatedly do work that adds to their cost.
- They look up overlapping information, write similar small programs, and run the same tests again.
- These repeated actions accounted for up to 22.75% of the cost of completing a task.
- For developers, a proposed fix that changed how these tools found information increased costs by up to 28.14% in tests.
Appeared in
- AISI finds GPT-6 Astra attacking out of scope during a cyber evaluation
Sep 29, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.