Story · Louis-François Bouchard, Omar Solano and Samridhi Vaid (Towards AI)
Context Engineering in 2026 (Louis-François Bouchard, Omar Solano and Samridhi Vaid (Towards AI))
talk · Story page

Across 11 context presets on an open-source AI tutor, sending the full history beat every compaction technique on recall, cost and latency, because 97% of tokens came from cache and rewriting the context invalidates it. Full history recovered specific details 95% of the time against 32% after summarizing.
In plain words
- A test of 11 setups found that keeping complete conversations beat every tested shortening method on memory, cost, and response speed.
- Previously processed conversation text can be reused, making repeated text much cheaper to handle.
- Rewriting conversations as summaries prevents that reuse, so shortening must be very large before it saves money.
- Keeping full conversations recovered specific details 95% of the time, compared with 32% after summarizing.
- People building conversational artificial intelligence may save money and preserve details by keeping history, unless local memory limits prevent it.
Appeared in
- Deno's Claw Patrol treats agents as untrusted software, and full history beats compaction
Aug 18, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.
