Story · arXiv
Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines (arXiv)
paper · Story page

Across 100 olympiad math problems, a cross-family mid-tier reviewer lifted final accuracy from 52 to 64 percent with zero damaged answers. Same-model self-review caught more errors, 0.85 recall, yet produced no significant gain, rejecting 2.1 times as often and falsely rejecting 35 percent of its own correct answers.
In plain words
- In a math study, an artificial intelligence (AI) system answered more accurately when a different AI checked its work.
- The answering system tried to fix answers that the reviewer rejected.
- When the answering AI checked its own work, it caught more mistakes but wrongly rejected many correct answers.
- For teams using AI reviewers, catching more errors did not guarantee more correct final answers.
Appeared in
- Prefix caching changes agent runs, and cross-family reviewers beat self-review
Sep 08, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
