Story · Snorkel AI
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results (Snorkel AI)
blog post · Story page
Pass@1 is flat across three generations on one expert-created terminal-bench style task set: 61.5% for Fable 5.1, 60.7% for Opus 5, and 60.7% for Opus 5.5. Pass@5 separates them, and Opus 5.5's 76.7% sits below Opus 5's 79.3%.
In plain words
- Snorkel found that newer coding assistants did not consistently do better on its programming tests.
- The tests checked success on the first attempt and whether allowing five attempts helped.
- Some Fable 5.1 failures came from stopping early or being unable to recover after errors.
- People choosing coding help can use these failure details to compare assistants beyond their overall success rates.
Appeared in
- Two MemOS packages shipped credential stealers into the agent memory layer
Sep 24, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.