Strata runs ~120B Qwen 3.8 Flash Next on a 2060 laptop at 10 tok/s
10/03/2026 — 10/03, 23:12·1 sources·1 reports
Story overview
On October 3, 2026, a post appeared on Reddit's LocalLLaMA forum describing a laptop with a 2060 GPU, 32 GB of RAM and 6 GB of VRAM running Qwen 3.8 Flash Next, a model of roughly 120B parameters, quantized to q2_0, on the Strata engine. According to the author, decoding came in at about 10 tok/s at a 50k context. The setup used a 4 bit KV cache and no vision.
The author said he had expected prompt processing of around 30 token/s and decode of at most 3 tok/s, given that the full Qwen 3.8 Next has around 120B parameters (not counting Engrams) and the machine has just 32 GB of RAM. He had assumed the run would crawl to a halt because of SSD swapping. The measured 10 token/s at 50k context struck him as remarkable for such an old device paired with such a large model, and he called the engine the real deal. So far the matter rests with this single post; no further developments have been reported.
AI-generated from 1 reports · updated 2 hours ago
Latest turnA user ran Qwen 3.8 Flash Next — a roughly 120B-parameter model — on a laptop with 32 GB RAM and 6 GB VRAM using the Strata engine. With a 4-bit KV cache and no vision, decode reached about 10 tok/s at 50k context, well above the 3 tok/s the user had expected.
- Heat index
- 91
- Sources
- 1
- Reports
- 1
- First seen
- 4 hours ago
Reports on this story headlines open the original
A user ran Qwen 3.8 Flash Next — a roughly 120B-parameter model — on a laptop with 32 GB RAM and 6 GB VRAM using the Strata engine. With a 4-bit KV cache and no vision, decode reached about 10 tok/s at 50k context, well above the 3 tok/s the user had expected.
Other stories people are talking about
- 552RisingApple tightens macOS Full Disk Access over AI agent risks9 sources
- 392NVIDIA launches 64GB DGX Spark desktop AI computer at $4,9998 sources
- 376RisingMeta open-sources Muse Gadgets firmware and SDK6 sources
- 289Claude Opus 5.5 and GPT-6 Sol: comparing cost per task4 sources
- 264Anthropic Reportedly Targets Nov 9 IPO Listing Before Thanksgiving5 sources
- 249NewSurgeAleph Alpha releases Kolibri-1, a 78B MoE model with 1M context3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
