LoRA fine-tuning of Qwen3.8-Flash-Next on 40GB of VRAM
10/09/2026 — 10/11, 02:16·1 sources (all)·1 reports (all)
Key takeaways
- A developer updated a LoRA over GGUF approach on Reddit's LocalLLaMA board that trains Qwen3.8-Flash-Next (125B-A6B plus a 51B engram) withi…
A developer updated a LoRA over GGUF approach on Reddit's LocalLLaMA board that trains Qwen3.8-Flash-Next (125B-A6B plus a 51B engram) within 40 GiB of VRAM and no CPU offload. Keeping the engram on disk does not slow training. On a Strix Halo machine, sharded training at 2048 context runs at 9.5 s/it, roughly 200 tokens/s.
Latest turnA developer on Reddit's LocalLLaMA forum updated a LoRA over GGUF recipe that trains Qwen3.8-Flash-Next (125B-A6B plus 51B engram) in 40 GiB of VRAM with no CPU offloading, keeping engram on disk without slowing training. On Strix Halo it runs context chunk size 2048 at 9.5 s/it, roughly 200 tokens/s.
Related AI tools
- Heat index
- 32
- All sources
- 1
- All reports
- 1
- First seen
- 2 days ago
Reports on this story headlines open the original
A developer on Reddit's LocalLLaMA forum updated a LoRA over GGUF recipe that trains Qwen3.8-Flash-Next (125B-A6B plus 51B engram) in 40 GiB of VRAM with no CPU offloading, keeping engram on disk without slowing training. On Strix Halo it runs context chunk size 2048 at 9.5 s/it, roughly 200 tokens/s.
Reddit · LocalLLaMAAI score 55
Other stories people are talking about
- Heat index 42Qwen3.8-27B local test hits 159 tok/s via speculative decoding1 sources
- Heat index 83User writes CUDA megakernel for Qwen3.8-27B with Claude Opus 5.5, up to 1.9x faster1 sources
- Heat index 94Qwen open-sources Qwen-Image-2.1-Turbo, an 8-step image model2 sources
- Heat index 59Developer releases Basalt, a Blackwell-tuned inference engine1 sources
- Heat index 28Atomic Agent Desktop goes open source, runs Qwen and Gemma locally1 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
