Developer releases Basalt, a Blackwell-tuned inference engine
10/10/2026 — 10/10, 23:19·1 sources (all)·1 reports (all)
Key takeaways
- A developer has released Basalt, an inference engine heavily tuned from a Strata fork that currently supports only Qwen3.8 Flash-Next on Bla…
Background & analysis
On October 10, 2026, a developer announced Basalt, an inference engine, in the LocalLLaMA section of Reddit. According to the post, Basalt is built on a heavily tuned fork of Strata and is positioned as a specialized engine that supports only Qwen3.8 Flash-Next and the Blackwell architecture, including dual-GPU setups such as a 5090 paired with a 5060 Ti. The developer's own machine is a 5090 plus a 5060 Ti, running on a 9950X with 32GB of DDR5 RAM.
On performance, the developer claims Basalt can reach up to 2.6x the throughput of the original Strata on the same weights. The reported measurements include 665 tok/s for structured output, 354 tok/s for prose output, and a prefill figure of 7,317 at 64k context, using IQ3_XXS quantization at 400 W.
So far the engine has only been announced by its author. The report does not mention a release schedule, how the project will be distributed, or any third-party verification, and no other source has responded to the speed figures.
AI-generated from 1 reports · updated 3 hours ago
Latest turnA developer released Basalt, a heavily tuned fork of Strata that supports only Qwen3.8 Flash-Next and the Blackwell architecture, including dual-GPU setups like a 5090 plus 5060 Ti. The author claims up to 2.6x the throughput of stock Strata on the same weights, reporting 665 tok/s structured and 354 tok/s prose, 7,317 prefill at 64k context, IQ3_XXS and 400 W.
Related AI tools
- Heat index
- 61
- All sources
- 1
- All reports
- 1
- First seen
- 18 hours ago
Reports on this story headlines open the original
A developer released Basalt, a heavily tuned fork of Strata that supports only Qwen3.8 Flash-Next and the Blackwell architecture, including dual-GPU setups like a 5090 plus 5060 Ti. The author claims up to 2.6x the throughput of stock Strata on the same weights, reporting 665 tok/s structured and 354 tok/s prose, 7,317 prefill at 64k context, IQ3_XXS and 400 W.
Reddit · LocalLLaMAAI score 53
Other stories people are talking about
- Heat index 33LoRA fine-tuning of Qwen3.8-Flash-Next on 40GB of VRAM1 sources
- Heat index 86NewUser writes CUDA megakernel for Qwen3.8-27B with Claude Opus 5.5, up to 1.9x faster1 sources
- Heat index 97Qwen open-sources Qwen-Image-2.1-Turbo, an 8-step image model2 sources
- Heat index 43Qwen3.8-27B local test hits 159 tok/s via speculative decoding1 sources
- Heat index 28Atomic Agent Desktop goes open source, runs Qwen and Gemma locally1 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
