ninfer vs llama.cpp: Qwen3.8 27B Runs 2.3x Faster
10/03/2026 — 10/03, 18:49·1 sources·1 reports
Story overview
On October 3, 2026, a user on Reddit's LocalLLaMA forum posted a hands-on comparison of ninfer and llama.cpp, both running Qwen3.8 27B on an RTX 5090. The test used a single oneshot run with thinking enabled at xhigh, a 120k context window, default sampling settings, and the same prompt for both backends — described by the poster as covering physics and spin with the full rule set.
According to the post summary, ninfer decoded at roughly 147 tokens/s, while llama.cpp with Q4_K_M and MTP reached about 141 tokens/s. Wall-clock time came in at 7 minutes 40 seconds for ninfer versus 8 minutes 13 seconds for llama.cpp. The poster also said that switching to ninfer's NVFP4 build produced a larger output.
One point of disagreement is worth flagging: the headline says ninfer is 2.3x faster than llama.cpp, but the decode figures in the summary differ by only about 4%. The excerpt shown is also incomplete — it cuts off in the middle of a table row that lists ninfer's precision for the non-NVFP4 build alongside MTP. That leaves the underlying numbers only partly verifiable from the material available.
So far this remains a single user-generated benchmark post. No other reports, follow-up tests, or official statements from either project have appeared.
AI-generated from 1 reports · updated 2 hours ago
Latest turnA user benchmarked ninfer against llama.cpp running Qwen3.8 27B on an RTX 5090 with 120k context and thinking enabled: ninfer decoded at roughly 147 tokens/s versus 141 tokens/s for llama.cpp with Q4_K_M and MTP, finishing the same prompt in 7m40s versus 8m13s. The NVFP4 ninfer build produced a longer output, the author notes.
- Heat index
- 92
- Sources
- 1
- Reports
- 1
- First seen
- 4 hours ago
Reports on this story headlines open the original
A user benchmarked ninfer against llama.cpp running Qwen3.8 27B on an RTX 5090 with 120k context and thinking enabled: ninfer decoded at roughly 147 tokens/s versus 141 tokens/s for llama.cpp with Q4_K_M and MTP, finishing the same prompt in 7m40s versus 8m13s. The NVFP4 ninfer build produced a longer output, the author notes.
Reddit · LocalLLaMAAI score 82
Other stories people are talking about
- 529RisingApple tightens macOS Full Disk Access over AI agent risks8 sources
- 453NVIDIA launches 64GB DGX Spark desktop AI computer at $4,9998 sources
- 334SurgeClaude Opus 5.5 and GPT-6 Sol: comparing cost per task4 sources
- 327Meta open-sources Muse Gadgets firmware and SDK5 sources
- 219Microsoft AI releases MAI-Transcribe-2-Streaming real-time transcription model4 sources
- 192Amazon Weighs Moving $8 Billion of Nvidia Chips into a Financing Vehicle3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
