Story · arXiv
Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild (arXiv)
paper · Story page
Across 37,623 provenance-labeled pull requests from five commercial agents and a matched human baseline, revert rates split by vendor: 6.1% for Codex, 11.5% for humans, 14.5% for Devin. Pooled agent code carried fewer security smells.
In plain words
- Researchers found that some artificial intelligence coding assistants had their work undone more often than others.
- They compared software changes written by coding assistants with similar changes written by people.
- Codex changes were undone 6.1% of the time, compared with 11.5% for people and 14.5% for Devin.
- Taken together, the assistants' changes showed fewer signs of possible security problems than changes written by people.
- The findings help software teams compare coding assistants by how often their accepted changes were later undone.
Appeared in
- Split an MCP injection across two channels and resistant models leak at 100%
Sep 18, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.