Kyojin runs two 300B-class MoE models on one 128 GB mini PC
10/03/2026 — 10/03, 23:12·1 sources·1 reports
Story overview
On October 3, 2026, a post on Reddit's LocalLLaMA forum introduced an inference engine called Kyojin. According to the post, Kyojin is built on top of ExLlamaV3 and targets Strix Halo (gfx1151, ROCm). The goal is to pack two 300B-class MoE models so that each one fits on a single 128 GB mini PC.
The post describes this as Kyojin's first release, with all measurements taken on a Ryzen AI Max+ 395. The two models are GLM-5.3-Flash and MiMo-V2.6-Flash-MOPD, occupying 99.7 GB and 105 GB respectively. GLM-5.3-Flash reaches about 580 tok/s prefill at 3.5K context and 546 tok/s at 64K, with decode between 26 and 30 tok/s (MTP). MiMo-V2.6-Flash prefill is reported at about 650 tok/s at 4K context, while its decode figures vary by workload: 32 tok/s for prose, 35 tok/s for chat, and 44 tok/s for code with speculative decoding, alongside a 29 tok/s figure labeled plain. The headline framing states that MiMo-V2.6-Flash decodes at up to 44 tok/s.
So far this is the only report on the release, and no further sources or follow-up measurements are available to check the numbers against.
AI-generated from 1 reports · updated 2 hours ago
Latest turnA new engine called Kyojin, built on ExLlamaV3, runs two 300B-class MoE models on a single 128 GB Strix Halo mini PC. On a Ryzen AI Max+ 395, GLM-5.3-Flash reaches about 580 tok/s prefill at 3.5K context, while MiMo-V2.6-Flash hits up to 44 tok/s decode.
- Heat index
- 92
- Sources
- 1
- Reports
- 1
- First seen
- 3 hours ago
Reports on this story headlines open the original
A new engine called Kyojin, built on ExLlamaV3, runs two 300B-class MoE models on a single 128 GB Strix Halo mini PC. On a Ryzen AI Max+ 395, GLM-5.3-Flash reaches about 580 tok/s prefill at 3.5K context, while MiMo-V2.6-Flash hits up to 44 tok/s decode.
Reddit · LocalLLaMAAI score 85
Other stories people are talking about
- 552RisingApple tightens macOS Full Disk Access over AI agent risks9 sources
- 392NVIDIA launches 64GB DGX Spark desktop AI computer at $4,9998 sources
- 376RisingMeta open-sources Muse Gadgets firmware and SDK6 sources
- 289Claude Opus 5.5 and GPT-6 Sol: comparing cost per task4 sources
- 264Anthropic Reportedly Targets Nov 9 IPO Listing Before Thanksgiving5 sources
- 249NewSurgeAleph Alpha releases Kolibri-1, a 78B MoE model with 1M context3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
