Story · arXiv
An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop (arXiv)
paper · Story page

A research loop built against drift: immutable experiment cards so a falsified hypothesis can't be retconned, subagents locked to mechanical roles, and a preference oracle that alone makes subjective calls. The oracle changed research direction, not the best score.
In plain words
- Researchers built an artificial intelligence research system that keeps testing its original ideas instead of quietly changing them.
- Every experiment used a permanent card linking its prediction to its result, so failed ideas could not be rewritten.
- Supporting workers performed fixed mechanical jobs, while one preference tool alone made judgments based on the user's research tastes.
- With or without the preference tool, the system rejected roughly three quarters of its own ideas.
- Researchers could guide which questions get studied without changing how results are judged.
Appeared in
- In LangChain's benchmark, only 7% of agent turns needed a frontier model
Aug 12, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
