HaiAI123

Curated Global AI Tools Directory

Industry update
100
heat index
New

Xiaomi releases MiMo-V2.6-Pro and Flash MoE models with Agentic RL report

10/10/2026 — 10/10, 18:20·1 sources (all)·1 reports (all)·In the 2026-10-11 briefing

Key takeaways

  • Xiaomi's MiMo team released the MiMo-V2.6-Pro and MiMo-V2.6-Flash MoE models alongside a technical report. The report details how it scales…

Background & analysis

On October 8, the MiMo team at Xiaomi released a technical report titled "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement," describing a large-scale Agentic RL training process. The approach scales the training batch, task environments and harness, along with the Grader Compute used to judge agent behavior, in order to explore a path toward continuous model self-improvement.

The release covers two models, MiMo-V2.6-Pro and MiMo-V2.6-Flash. Pro has 1.02T total parameters and activates 42B parameters per inference; Flash has 310B total parameters and activates 15B. Both use a sparse Mixture-of-Experts (MoE) architecture and mix local sliding window attention (SWA) with global attention (GA) in the text backbone, supporting training on context lengths in the million-token range.

According to the report, MiMo-V2.6 splits reinforcement learning compute into three parts: Rollout, Grading and Training, each scaled independently. In practice, every RL step first samples 1568 prompts, and each prompt generates 16 rollouts at once, producing roughly 25,000 trajectories per step and 2.7 billion to 3.7 billion tokens in total. The RL post-training compute cost is about $2.6 million.

In an October 11 report, InfoQ Chinese described the series as the first to systematically scale Agentic RL simultaneously across the three dimensions of Rollout, environment and Grader, and noted that the million-token context is used to support long-horizon agent behavior modeling. The report also links to the technical report at https://arxiv.org/pdf/2610.11959. Public information currently stops at the technical report and the two model releases; no later versions or deployment plans were mentioned.

AI-generated from 1 reports · updated 7 hours ago

Latest turnXiaomi's MiMo team released the MiMo-V2.6-Pro and MiMo-V2.6-Flash models along with a technical report describing an Agentic RL pipeline that scales rollouts, environments, and grader compute together. Pro has 1.02T total parameters with 42B activated; one RL step produces about 25,000 trajectories, or 2.7–3.7 billion tokens, and its RL post-training compute cost roughly $2.6 million.

24-hour heatHeat index 100
Just charted — trend is building up
24 hours agonow

Reports on this story headlines open the original

Today
  1. Xiaomi's MiMo team released the MiMo-V2.6-Pro and MiMo-V2.6-Flash models along with a technical report describing an Agentic RL pipeline that scales rollouts, environments, and grader compute together. Pro has 1.02T total parameters with 42B activated; one RL step produces about 25,000 trajectories, or 2.7–3.7 billion tokens, and its RL post-training compute cost roughly $2.6 million.

    InfoQ 中文AI score 62

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to AI industry updates →