Today, we are officially releasing and open-sourcing the Xiaomi MiMo-V2.6 series. This marks a key step in our exploration of the RSI (recursive self-improvement) path: building on verifiable complex tasks, we scale up reinforcement learning (RL) computing power to enable models to continuously expand the boundaries of intelligence through ongoing exploration and feedback.
Where the path is flat and close, travelers are many; where it is rugged and distant, few reach the end. In an era where intelligence can be easily replicated, we choose to channel computing power into real-world environments, letting models learn through trial and error in iterative feedback loops. This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we've taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.
The MiMo-V2.6 series comprises two native fully multimodal models, namely Pro and Flash. Benefiting from the expanded RL computing power, MiMo-V2.6-Pro scores 46 points in the Artificial Analysis Intelligence Index (AA Composite Intelligence Index), surpassing Kimi K3 and Qwen3.8 Max to become the most powerful open-source model available; however, there is still a gap when compared with the strongest closed-source models Claude Fable 5.1 and GPT-6 Astra.
The MiMo-V2.6 series adopts the same API pricing as the V2.5 series. With intelligent performance upgraded while price remains unchanged, the Pareto frontier of "intelligence vs. cost" has thus been pushed outward once again. The MiMo-V2.6-Pro has set a new cost-performance record for domestic large language models: at the same intelligence level, its price is only 1/20 to 1/60 that of overseas models.
Scale RL on a large scale and fully open-source it
During the RL training phase, MiMo-V2.6 is likely one of the domestic open-source models that has been allocated the largest amount of computing power to date. After large-scale, multi-task reinforcement learning training, MiMo-V2.6-Pro has achieved performance on most Agent Benchmarks that is on par with Claude Opus 5 and GPT-5.6 Sol, while MiMo-V2.6-Flash has comprehensively outperformed MiMo-V2.5-Pro.
Throughout the entire process, we overcame fundamental research and engineering challenges in RL training, and documented the official experimental journey via live sharing. In less than 6 days, MiMo-V2.6-Flash and MiMo-V2.6-Pro completed 30 steps each with a cumulative total of approximately 750,000 trajectories, at training costs of around $850,000 and $2.62 million respectively. The average pass rate of training tasks has been relatively improved by 25% and 12% respectively, and the out-of-sample long-range software engineering evaluation benchmark DeepSWE v1.1 has been improved by approximately 17 points (from 48.8 to 65.7) and approximately 14 points (from 58.4 to 72.6) respectively.
This training mainly expands RL computing power from three dimensions:
- Larger Batch Size and Higher Throughput: By combining a large batch size and a fully asynchronous architecture, each update uses 1,568 samples, supports training with 1M context length, and the number of tokens per training step reaches 3.5–3.7B.
- More Tasks and Complex Environments: Build a multi-task training system covering fields such as Code, General, Visual, and Cyber, and integrate multiple Harnesses to facilitate the collaborative improvement of different capability dimensions.
- Greater Grader computing power: through relative comparison within the Group, it provides more accurate and diverse reward signals for Long-Horizon RL tasks, forms a closed loop for model self-improvement, and guides the model to complete tasks with shorter paths and fewer Tokens.
As the training scale expands, we freeze the MoE Router to suppress expert load drift, and establish a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators, so as to improve training stability and reward reliability.
We have open-sourced the aforementioned technical achievements and supporting resources, including the complete technical report, training environment and RL code, to help more researchers reproduce and verify relevant results.
From Vibe Coding to Vibe World
MiMo-V2.6 integrates 3D spatial reasoning, multimodal perception and computer user operation (CUA) capabilities, further expanding the boundaries of what programming can achieve; it can extend natural language-driven programming tasks into the "Vibe World" oriented towards interactive world construction—including 3D open-world games, Blender 3D modeling, embodied intelligence with a Franka Panda arm, and Computer Use Agents.
Advance cutting-edge scientific research
Without undergoing specialized reinforcement learning tailored for scientific research tasks, MiMo-V2.6 has already demonstrated application potential across multiple research domains—from materials design (MOF candidates for PFAS adsorption) to formal mathematical proof (Lean 4 formalization of Li and Yorke's Period Three Implies Chaos).
Fully open source
We have fully open-sourced the weights and technical report of the MiMo-V2.6-Pro and Flash models, simultaneously released the MiMo-V2.6-Distill-Qwen-9B along with supporting reinforcement learning (RL) research resources:
- 7k+ high-quality RL task environments covering software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development
- End-to-end RL training framework built on verl, uni-agent and mini-swe-agent
- Lightweight and composable mini-harnesses
Open Source Link: https://huggingface.co/collections/XiaomiMiMo/mimo-v26
Note: When calling the API, please use the all-lowercase model names mimo-v2.6-pro, mimo-v2.6-flash, and mimo-v2.6-pro-ultraspeed.
Update Time: September 22, 2026