Story · Hamel Husain
Claude’s new auto eval tool (Hamel Husain)
blog post · Story page

Husain and Isaac Flath livestreamed Claude Code's new build_eval and hill-climb commands over traces from an apartment leasing assistant. His sharpest complaint is ordering: it asked them to pick a failure before they'd read any conversations.
In plain words
- Anthropic added tools to Claude Code that help developers test and improve applications built with artificial intelligence.
- The tools create tests, check how answers are judged, and help improve the application's results.
- Hamel Husain and Isaac Flath were asked to choose a problem before reading the apartment leasing assistant's conversations.
- Husain argues that developers should read those conversations first to decide which problems are worth testing.
Appeared in
- A forged chat-template marker loses most of its authority as ordinary subwords
Oct 01, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
