Story · arXiv
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks (arXiv)
paper · Story page
Planted optional shortcuts that inflate public test scores
Appeared in
- Anthropic's deliberately misaligned model, Fable 5.1, and a fix for reward hacking
Sep 02, 2026 · quick links
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.