llama.cpp fuses shared experts into MMVQ for MoE speedups
10/03/2026 — 10/03, 03:02·1 sources·1 reports
Story overview
On October 3, 2026, a post on Reddit's r/LocalLLaMA surfaced a new pull request in llama.cpp, PR #29184, which folds shared experts into MMVQ to speed up inference for MoE models. The submission came from user u/jacek2023, who noted in the accompanying description that the gain is not universal: the speedup applies only to some MoE architectures. Qwen 35B A3B was given as an example of an architecture that benefits.
That is the full extent of what the report states. The PR number, the change it makes, the condition attached to it, and the single architecture named as an example are the concrete details available. No benchmark figures were included, so there is no stated measure of how much faster inference becomes under the change. The author also did not enumerate the other MoE architectures that might be affected, leaving the boundary of the speedup undefined beyond the Qwen 35B A3B example.
Nor does the material say where the pull request stands in the review process. There is no indication of whether PR #29184 has been merged, whether it is still under discussion, or whether it has been scheduled for any upcoming release of llama.cpp. The thread links to the pull request page and its comments, but the summary available here does not describe what those comments contain or whether reviewers raised any objections.
As it stands, the item is a single community post pointing at an open code contribution, with the scope of the claimed speedup limited by the author to selected MoE architectures and illustrated with one example. Any further detail about performance numbers, broader applicability, or the merge status would have to come from reporting beyond what is currently available. For now, the change remains a proposal discussed in the LocalLLaMA community rather than a confirmed addition to the project.
AI-generated from 1 reports · updated 2 hours ago
Latest turnA pull request (PR #29184) in llama.cpp fuses shared experts into MMVQ, speeding up inference for some MoE architectures. The gains are concentrated on specific MoE models such as Qwen 35B A3B, and not every MoE architecture benefits.

- Heat index
- 96
- Sources
- 1
- Reports
- 1
- First seen
- 2 hours ago
Reports on this story headlines open the original
A pull request (PR #29184) in llama.cpp fuses shared experts into MMVQ, speeding up inference for some MoE architectures. The gains are concentrated on specific MoE models such as Qwen 35B A3B, and not every MoE architecture benefits.
Other stories people are talking about
- 610SurgeNVIDIA launches 64GB DGX Spark desktop AI computer at $4,9997 sources
- 480NewApple tightens macOS Full Disk Access over AI agent risks5 sources
- 433RisingGoogle releases Gemini 4 Argon as next-gen frontier AI model12 sources
- 255SurgeSuno launches voice generation feature for narration and music3 sources
- 219SurgeMicrosoft AI releases MAI-Transcribe-2-Streaming real-time transcription model3 sources
- 186NewAnthropic launches Claude Frontier Academy with $100M2 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
