Story · Jo Kristian Bergum (Hornet.dev)
The unreasonable effectiveness of BM25 for agentic search (Jo Kristian Bergum (Hornet.dev))
talk · Story page
The lexical function didn't change, Bergum argues, the user did: a model knows entities, dates and product identifiers, so it writes far longer queries than a person would and fires a dozen in a row. He points at a benchmark where accuracy is high with answer-bearing documents in context and drops once the model has to fetch them.
In plain words
- Bergum argues that an old method of matching search words to documents works well for artificial intelligence.
- The software can search repeatedly using more names, dates, and product details than a person would usually type.
- In a test, it answered more accurately when given documents containing the answers than when it had to find them.
- For people building automated research tools, finding the right documents can be an obstacle to accurate answers.
Appeared in
- Every audited chat tokenizer lets prompt text forge control tokens
Sep 17, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.