Story · Denys Linkov (Wisedocs)
Benchmarking Coding Agents on New vs Legacy Codebases (Denys Linkov (Wisedocs))
talk · Story page
Linkov audits a six-month, ten-repository medical-claims refactor to ask whether agents could have done it. The same task took three hours and ten major mistakes with o3, while Opus 4.8 essentially got it in one pass; handed the whole job, GPT 5.5 declared it done in about ten minutes with the actual models missing.
In plain words
- A talk tested whether newer coding assistants can handle a real software restructuring job that previously took six months.
- The original team combined ten separate code collections used to process medical claims.
- On one task, o3 needed three hours of guidance and still made ten major mistakes, while Opus 4.8 nearly succeeded on its first attempt.
- One current coding assistant claimed the whole job was finished after about ten minutes but omitted core systems and steps needed to run it.
- Software teams may save time on focused tasks, but they still need to check whether a coding assistant truly finished.
Appeared in
- Claude Code turns auto mode on for everyone; reward-hack monitors catch 28%
Aug 11, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.