Qwen3.8-Flash-Next 177B Hits 11-15 tok/s on a Single RTX 5070
10/03/2026 — 10/03, 15:12·1 sources·1 reports
Story overview
On October 3, 2026, a post on the LocalLLaMA subreddit described a local inference setup for Qwen3.8-Flash-Next 177B. According to the author, he had been working on a llama.cpp-based expert streaming approach on Windows, running the model at UD-IQ3_XXS quantization on a single RTX 5070 with 12GB of VRAM, alongside 32GB of DDR4-2400 memory, a Ryzen 5 5600GT, and a PCIe Gen3 link. The headline figure for the setup is a range of 11-15 tok/s on that one GPU.
The post reports several numbers behind that range. The benchmark came in at about 11.5 tok/s, which the author says is up from roughly 7 tok/s on the setup he inherited. In ordinary conversation he saw 14-15 tok/s. On a long coding prompt, the model generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.
The post does not list context length, batch size, or other generation parameters, and the point of comparison is the author's own inherited setup rather than any published baseline. So far this is the only report on the configuration: the figures come from the poster's own account, and the material contains no third-party reproduction or independent verification.
AI-generated from 1 reports · updated 1 hour ago
Latest turnA llama.cpp-based expert streaming setup runs Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on a single RTX 5070 12GB under Windows at about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. Everyday chat reaches 14–15 tok/s, and one long coding prompt produced 4,892 tokens at 10.15 tok/s while generating a working single-file Snake game. The rest of the machine is 32GB DDR4-2400, a Ryzen 5 5600GT and PCIe Gen3.

Reports on this story headlines open the original
A llama.cpp-based expert streaming setup runs Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on a single RTX 5070 12GB under Windows at about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. Everyday chat reaches 14–15 tok/s, and one long coding prompt produced 4,892 tokens at 10.15 tok/s while generating a working single-file Snake game. The rest of the machine is 32GB DDR4-2400, a Ryzen 5 5600GT and PCIe Gen3.
Reddit · LocalLLaMAAI score 75
Other stories people are talking about
- 508NVIDIA launches 64GB DGX Spark desktop AI computer at $4,9998 sources
- 494Apple tightens macOS Full Disk Access over AI agent risks7 sources
- 367Meta open-sources Muse Gadgets firmware and SDK5 sources
- 246SurgeMicrosoft AI releases MAI-Transcribe-2-Streaming real-time transcription model4 sources
- 215SurgeAmazon Weighs Moving $8 Billion of Nvidia Chips into a Financing Vehicle3 sources
- 208Anthropic launches Claude Frontier Academy with $100M3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
