Story · r/LocalLLaMA
Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38x faster decisions and +8.7 points accuracy, for under 2 GB extra memory (r/LocalLLaMA)
Reddit thread · Story page
A 0.8B model picks between options you define and returns a calibrated probability for each, now with nine LoRA adapters of about 40 MB apiece for recurring agent decisions like tool choice and prompt-injection detection. The speed and accuracy numbers are the author's own, on the author's hardware.
In plain words
- A developer released extra training for Jeff, a small artificial intelligence system that chooses between options.
- Jeff picks from choices you provide and estimates how likely each choice is to be right.
- Separate training packages teach Jeff jobs such as choosing tools or checking whether answers agree with their sources.
- The developer reports faster, more accurate decisions on their own computers when Jeff makes choices before a larger system.
Appeared in
- Six ways an agent harness can run something other than what you approved
Oct 02, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.