HaiAI123

Curated Global AI Tools Directory

Industry update
Flat

Qwen3.8-27B local test hits 159 tok/s via speculative decoding

10/09/2026 — 10/11, 00:18·1 sources (all)·1 reports (all)

Key takeaways

  • A developer tested Qwen3.8-27B locally, using Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8 to boost decoding speed from 14 to…

Background & analysis

On October 9, 2026, benchmark data on running Qwen3.8-27B locally surfaced on Reddit, and a developer later published a write-up on Juejin breaking down the setup. According to that account, the tester combined Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8 to push decoding speed from 14 tok/s to 159 tok/s.

The reported figures are 159.4 tok/s on a Sapphire Radeon AI PRO R9700 and 64.4 tok/s on the Strix Halo mobile platform. For comparison, the article cites a Mac Studio with M3 Ultra running Qwen3.8-27B Q4_K_M through Ollama, where the weights take up about 17 GB and sustained generation sits at roughly 14 tok/s.

The write-up attributes the speedup to two layers of optimization: draft tree verification and fusion of the gate/up GEMM operators. It also mentions comparing hardware and software configurations alongside llm-bench quality results, though the summary does not include the specific outcomes of those comparisons. The public information currently stops at this benchmark data and the breakdown of the optimization path, with no further developments reported in the available material.

AI-generated from 1 reports · updated 2 hours ago

Latest turnA developer benchmarked Qwen3.8-27B locally and raised decoding from 14 tok/s to 159 tok/s using Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8. Benchmarks shared on Reddit show 159.4 tok/s on macOS with a Sapphire Radeon AI PRO R9700 and 64.4 tok/s on Strix Halo. The write-up breaks down the draft-tree verification and gate/up GEMM operator fusion behind the gains.

Related AI tools

24-hour heatHeat index 43 · peak 84 · 23h ago
24 hours agonow

Reports on this story headlines open the original

Oct 9
  1. A developer benchmarked Qwen3.8-27B locally and raised decoding from 14 tok/s to 159 tok/s using Q8 DFlash2 speculative decoding with LemonSeed Engine 0.5.8. Benchmarks shared on Reddit show 159.4 tok/s on macOS with a Sapphire Radeon AI PRO R9700 and 64.4 tok/s on Strix Halo. The write-up breaks down the draft-tree verification and gate/up GEMM operator fusion behind the gains.

    掘金AI score 55

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to AI industry updates →