Story · Stefania Druga (Sakana.ai)
Memory Harnesses for Long-Running Research Agents (Stefania Druga (Sakana.ai))
talk · Story page

Druga held the model fixed and varied only the recall policy. When everything fits in context, memory adds cost and nothing else; on long-horizon tasks a ranked decisions ledger beat vector RAG and gated recall, and even oracle memory didn't reach the ceiling.
In plain words
- Stefania Druga found that extra memory helped artificial intelligence research systems only when needed information no longer fit in their current conversation.
- When every paper already fit, memory kept accuracy unchanged while increasing cost.
- On longer work, a ranked list of past decisions outperformed systems with no memory and systems that searched stored text.
- Even receiving the correct memory did not ensure the system used that information correctly.
- This matters for people building long-running research tools, because carefully choosing recalled information can improve results and reduce cost.
Appeared in
- Mind viruses spread between LLM agents, and a one-line warning nearly stops them
Aug 13, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
