Xiaomi releases MiMo-V2.6-Pro and Flash MoE models with Agentic RL report
10/10/2026 — 10/10, 18:20·1 sources (all)·1 reports (all)·In the 2026-10-11 briefing
Key takeaways
- Xiaomi's MiMo team released the MiMo-V2.6-Pro and MiMo-V2.6-Flash MoE models alongside a technical report. The report details how it scales…
Background & analysis
On October 8, the MiMo team at Xiaomi released a technical report titled "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement," describing a large-scale Agentic RL training process. The approach scales the training batch, task environments and harness, along with the Grader Compute used to judge agent behavior, in order to explore a path toward continuous model self-improvement.
The release covers two models, MiMo-V2.6-Pro and MiMo-V2.6-Flash. Pro has 1.02T total parameters and activates 42B parameters per inference; Flash has 310B total parameters and activates 15B. Both use a sparse Mixture-of-Experts (MoE) architecture and mix local sliding window attention (SWA) with global attention (GA) in the text backbone, supporting training on context lengths in the million-token range.
According to the report, MiMo-V2.6 splits reinforcement learning compute into three parts: Rollout, Grading and Training, each scaled independently. In practice, every RL step first samples 1568 prompts, and each prompt generates 16 rollouts at once, producing roughly 25,000 trajectories per step and 2.7 billion to 3.7 billion tokens in total. The RL post-training compute cost is about $2.6 million.
In an October 11 report, InfoQ Chinese described the series as the first to systematically scale Agentic RL simultaneously across the three dimensions of Rollout, environment and Grader, and noted that the million-token context is used to support long-horizon agent behavior modeling. The report also links to the technical report at https://arxiv.org/pdf/2610.11959. Public information currently stops at the technical report and the two model releases; no later versions or deployment plans were mentioned.
AI-generated from 1 reports · updated 7 hours ago
Latest turnXiaomi's MiMo team released the MiMo-V2.6-Pro and MiMo-V2.6-Flash models along with a technical report describing an Agentic RL pipeline that scales rollouts, environments, and grader compute together. Pro has 1.02T total parameters with 42B activated; one RL step produces about 25,000 trajectories, or 2.7–3.7 billion tokens, and its RL post-training compute cost roughly $2.6 million.
- Heat index
- 100
- All sources
- 1
- All reports
- 1
- First seen
- 7 hours ago
Reports on this story headlines open the original
Xiaomi's MiMo team released the MiMo-V2.6-Pro and MiMo-V2.6-Flash models along with a technical report describing an Agentic RL pipeline that scales rollouts, environments, and grader compute together. Pro has 1.02T total parameters with 42B activated; one RL step produces about 25,000 trajectories, or 2.7–3.7 billion tokens, and its RL post-training compute cost roughly $2.6 million.
InfoQ 中文AI score 62
Other stories people are talking about
- Heat index 610OpenAI Defends Firing of Three Safety Researchers, Denies Safety Concerns15 sources
- Heat index 493OpenAI in Early Talks for $30B Raise at $1.4T Valuation14 sources
- Heat index 447TypeSafe AI raises $870M led by a16z at $7.5B valuation8 sources
- Heat index 314Microsoft launches Decision-1 for structured decisions5 sources
- Heat index 262Anthropic Tightens Claude Usage Rules on Elections and Weapons7 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
