Story · r/LocalLLaMA
The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks (r/LocalLLaMA)
Reddit thread · Story page
One user reports running Qwen 3.8 27B for about 21 days on one RTX 3090 to build a CUDA inference engine, on roughly 12 human messages, getting working kernels that never beat llama.cpp. Compaction ate about 83 hours, and the same card had to host the agents and run the engine under test.
In plain words
- One user reports letting an artificial intelligence system write software for about 21 days.
- Its task was to build software that could run artificial intelligence more efficiently on the same computer.
- Written rules told it when to ask the user for help.
- The system had to stop running so its new software could be tested on the same hardware.
- For this user, days of largely independent work produced working software that was no faster than an existing alternative.
Appeared in
- Giving an agent the decoder's confidence never beat a plain deterministic gate
Sep 21, 2026 · in the sections
Subscribe
Get the brief in your inbox
Pick daily, weekly, or both. Nothing is gated either way: every issue is on the site and in the feeds.
- Weekdays at 8:45am IST, one lead story and 6 to 9 items.
- Sundays, an argued synthesis rather than a recap.
- One click to leave, and quiet days say so in the subject line.