HaiAI123

Curated Global AI Tools Directory

NewHot story
96
heat index
New

llama.cpp fuses shared experts into MMVQ for MoE speedups

10/03/2026 — 10/03, 03:02·1 sources·1 reports

Story overview

On October 3, 2026, a post on Reddit's r/LocalLLaMA surfaced a new pull request in llama.cpp, PR #29184, which folds shared experts into MMVQ to speed up inference for MoE models. The submission came from user u/jacek2023, who noted in the accompanying description that the gain is not universal: the speedup applies only to some MoE architectures. Qwen 35B A3B was given as an example of an architecture that benefits.

That is the full extent of what the report states. The PR number, the change it makes, the condition attached to it, and the single architecture named as an example are the concrete details available. No benchmark figures were included, so there is no stated measure of how much faster inference becomes under the change. The author also did not enumerate the other MoE architectures that might be affected, leaving the boundary of the speedup undefined beyond the Qwen 35B A3B example.

Nor does the material say where the pull request stands in the review process. There is no indication of whether PR #29184 has been merged, whether it is still under discussion, or whether it has been scheduled for any upcoming release of llama.cpp. The thread links to the pull request page and its comments, but the summary available here does not describe what those comments contain or whether reviewers raised any objections.

As it stands, the item is a single community post pointing at an open code contribution, with the scope of the claimed speedup limited by the author to selected MoE architectures and illustrated with one example. Any further detail about performance numbers, broader applicability, or the merge status would have to come from reporting beyond what is currently available. For now, the change remains a proposal discussed in the LocalLLaMA community rather than a confirmed addition to the project.

AI-generated from 1 reports · updated 2 hours ago

Latest turnA pull request (PR #29184) in llama.cpp fuses shared experts into MMVQ, speeding up inference for some MoE architectures. The gains are concentrated on specific MoE models such as Qwen 35B A3B, and not every MoE architecture benefits.

Related tools
24-hour heatpeak 99 · 2h ago
24 hours agonow

Reports on this story headlines open the original

Today
  1. A pull request (PR #29184) in llama.cpp fuses shared experts into MMVQ, speeding up inference for some MoE architectures. The gains are concentrated on specific MoE models such as Qwen 35B A3B, and not every MoE architecture benefits.

    Reddit · LocalLLaMAAI score 70

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to the hot board →