Story · AlphaSignal
UCLA Finds AI Reward Hack Monitors Collapse to 28% on Real Cheating (AlphaSignal)
newsletter item · Story page
Reward-hacking monitors trained on synthetic examples fall to 28% on real model cheating; the synthetic data doesn't reflect how models actually exploit RL rewards. If a monitor is your safety layer, it may be watching for the wrong thing.
Appeared in
- Auto mode becomes the Claude Code default while agent monitoring keeps failing its tests
Aug 11, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.