Story · arXiv
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving (arXiv)
paper · Story page
An eighty-episode agentic tool-use workload measures how prefix caching changes trajectories under fixed seeds, across two engines and four weight formats.
In plain words
- Researchers found that reusing earlier calculations changed the steps an artificial intelligence (AI) assistant took.
- When requests start with the same instructions, the software can reuse earlier calculations instead of repeating them.
- Differences became more common when the assistant stored its internal numbers with less detail.
- For people repeating these tests, turning off reuse produced identical results each time within every tested setup.
Appeared in
- Prefix caching changes agent runs, and cross-family reviewers beat self-review
Sep 08, 2026 · lead story
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.