Story · arXiv (via papers.cool)
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs (arXiv (via papers.cool))
paper · Story page
An audit of 256 deployed chat tokenizers finds all of them forgeable, and proposes a tokenizer-level fix that leaves the token stream unchanged on attack-free data.
In plain words
- Researchers proposed a fix for a flaw that lets people fake who is speaking in a chatbot conversation.
- Chatbots use special labels to identify speakers, but people could create those labels by typing particular text.
- The fix keeps these labels separate from anything a person can type.
- Chatbot builders could block this trick while preserving ordinary message text exactly.
Appeared in
- Every audited chat tokenizer lets prompt text forge control tokens
Sep 17, 2026 · lead story
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.