HaiAI123

Curated Global AI Tools Directory

NewHot story
96
heat index
New

Qwen3.8-Flash-Next 177B Hits 11-15 tok/s on a Single RTX 5070

10/03/2026 — 10/03, 15:12·1 sources·1 reports

Story overview

On October 3, 2026, a post on the LocalLLaMA subreddit described a local inference setup for Qwen3.8-Flash-Next 177B. According to the author, he had been working on a llama.cpp-based expert streaming approach on Windows, running the model at UD-IQ3_XXS quantization on a single RTX 5070 with 12GB of VRAM, alongside 32GB of DDR4-2400 memory, a Ryzen 5 5600GT, and a PCIe Gen3 link. The headline figure for the setup is a range of 11-15 tok/s on that one GPU.

The post reports several numbers behind that range. The benchmark came in at about 11.5 tok/s, which the author says is up from roughly 7 tok/s on the setup he inherited. In ordinary conversation he saw 14-15 tok/s. On a long coding prompt, the model generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.

The post does not list context length, batch size, or other generation parameters, and the point of comparison is the author's own inherited setup rather than any published baseline. So far this is the only report on the configuration: the figures come from the poster's own account, and the material contains no third-party reproduction or independent verification.

AI-generated from 1 reports · updated 1 hour ago

Latest turnA llama.cpp-based expert streaming setup runs Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on a single RTX 5070 12GB under Windows at about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. Everyday chat reaches 14–15 tok/s, and one long coding prompt produced 4,892 tokens at 10.15 tok/s while generating a working single-file Snake game. The rest of the machine is 32GB DDR4-2400, a Ryzen 5 5600GT and PCIe Gen3.

Reports on this story headlines open the original

Today
  1. A llama.cpp-based expert streaming setup runs Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on a single RTX 5070 12GB under Windows at about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. Everyday chat reaches 14–15 tok/s, and one long coding prompt produced 4,892 tokens at 10.15 tok/s while generating a working single-file Snake game. The rest of the machine is 32GB DDR4-2400, a Ryzen 5 5600GT and PCIe Gen3.

    Reddit · LocalLLaMAAI score 75

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to the hot board →