Story · Dumitru Erhan, Shane Gu & Nicole Brichtova (Google DeepMind)
SOTA Generative Media Panel (Dumitru Erhan, Shane Gu & Nicole Brichtova (Google DeepMind))
talk · Story page
The panel regenerated real videos from their captions and human evaluators largely preferred the generated version: sharper and more saturated, Dumitru Erhan deflates, not more realistic. Their image model also quietly added wedding rings nobody caught internally, a warning for anyone scoring agents on preference signals.
In plain words
- A media generation system recreated real videos from written captions, and viewers generally preferred the artificial versions.
- Viewers favored sharper colors and smoother skin tones, even though those details did not make scenes more realistic.
- The image system quietly added wedding rings, and only an outside tester noticed the repeated mistake.
- The panel suggested that written descriptions lose details people notice in sound, color, taste, and other senses.
- People judging generated media may reward attractive surfaces while missing repeated errors, which can misdirect development.
Appeared in
- Maersk's 100,000 corrections, preference-trap evals, and Tencent's Hy4 Preview
Aug 31, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.