Qwen3.8-27B local test hits 159 tok/s via speculative decoding
10/09/2026 — 10/11, 00:18·1 sources (all)·1 reports (all)
Key takeaways
- A developer tested Qwen3.8-27B locally, using Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8 to boost decoding speed from 14 to…
Background & analysis
On October 9, 2026, benchmark data on running Qwen3.8-27B locally surfaced on Reddit, and a developer later published a write-up on Juejin breaking down the setup. According to that account, the tester combined Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8 to push decoding speed from 14 tok/s to 159 tok/s.
The reported figures are 159.4 tok/s on a Sapphire Radeon AI PRO R9700 and 64.4 tok/s on the Strix Halo mobile platform. For comparison, the article cites a Mac Studio with M3 Ultra running Qwen3.8-27B Q4_K_M through Ollama, where the weights take up about 17 GB and sustained generation sits at roughly 14 tok/s.
The write-up attributes the speedup to two layers of optimization: draft tree verification and fusion of the gate/up GEMM operators. It also mentions comparing hardware and software configurations alongside llm-bench quality results, though the summary does not include the specific outcomes of those comparisons. The public information currently stops at this benchmark data and the breakdown of the optimization path, with no further developments reported in the available material.
AI-generated from 1 reports · updated 2 hours ago
Latest turnA developer benchmarked Qwen3.8-27B locally and raised decoding from 14 tok/s to 159 tok/s using Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8. Benchmarks shared on Reddit show 159.4 tok/s on macOS with a Sapphire Radeon AI PRO R9700 and 64.4 tok/s on Strix Halo. The write-up breaks down the draft-tree verification and gate/up GEMM operator fusion behind the gains.
Related AI tools
- Heat index
- 43
- All sources
- 1
- All reports
- 1
- First seen
- 1 day ago
Reports on this story headlines open the original
A developer benchmarked Qwen3.8-27B locally and raised decoding from 14 tok/s to 159 tok/s using Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8. Benchmarks shared on Reddit show 159.4 tok/s on macOS with a Sapphire Radeon AI PRO R9700 and 64.4 tok/s on Strix Halo. The write-up breaks down the draft-tree verification and gate/up GEMM operator fusion behind the gains.
掘金AI score 55
Other stories people are talking about
- Heat index 86NewUser writes CUDA megakernel for Qwen3.8-27B with Claude Opus 5.5, up to 1.9x faster1 sources
- Heat index 33LoRA fine-tuning of Qwen3.8-Flash-Next on 40GB of VRAM1 sources
- Heat index 28Atomic Agent Desktop goes open source, runs Qwen and Gemma locally1 sources
- Heat index 97Qwen open-sources Qwen-Image-2.1-Turbo, an 8-step image model2 sources
- Heat index 61Developer releases Basalt, a Blackwell-tuned inference engine1 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
