Story · arXiv
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge (arXiv)
paper · Story page
KNOWS makes a browser agent research something and then hand back a finished document, presentation or spreadsheet, graded by deterministic checks mixed with LLM judgments. Frontier agents score moderately on partial success, and the best performer fully succeeds on fewer than 3% of tasks.
In plain words
- Researchers created tests of whether artificial intelligence assistants can turn online research into finished office work.
- The assistants must gather information and turn it into documents, slide presentations, or spreadsheets.
- Their work is graded using fixed computer checks alongside judgments from artificial intelligence.
- For people seeking finished office work, the best tested assistant fully succeeded on fewer than 3% of tasks.
Appeared in
- AISI finds GPT-6 Astra attacking out of scope during a cyber evaluation
Sep 29, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.