---
title: "MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement"
slug: mimo-v2-6-scaling-up-reinforcement-learning-for-self-improvement
url: https://listedarticles.com/articles/mimo-v2-6-scaling-up-reinforcement-learning-for-self-improvement
canonical_url: https://mimo.mi.com/docs/en-US/news/latest/v2-6
content_type: announcement
language: en
published_at: 2026-09-22T00:00:00.000Z
updated_at: 2026-09-22T03:12:27.455Z
author: "Xiaomi MiMo"
author_url: https://mimo.mi.com
authored_by: human
publisher: "Xiaomi MiMo"
publisher_url: https://mimo.mi.com
topics: ["AI", "LLMs", "Open Source", "Research", "Machine Learning"]
license: all-rights-reserved
word_count: 839
reading_minutes: 4
citation: "Xiaomi MiMo, Xiaomi MiMo. \"MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement.\" 22 Sept 2026. https://mimo.mi.com/docs/en-US/news/latest/v2-6 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement

> Xiaomi open-sources MiMo-V2.6 Pro and Flash after large-scale live RL (~$3.5M, 750k trajectories), claiming top open-weight AA Index scores, agent parity with frontier models, and 7k+ RL environments.

Today, we are officially releasing and open-sourcing the Xiaomi MiMo-V2.6 series. This marks a key step in our exploration of the RSI (recursive self-improvement) path: building on verifiable complex tasks, we scale up reinforcement learning (RL) computing power to enable models to continuously expand the boundaries of intelligence through ongoing exploration and feedback.

Where the path is flat and close, travelers are many; where it is rugged and distant, few reach the end. In an era where intelligence can be easily replicated, we choose to channel computing power into real-world environments, letting models learn through trial and error in iterative feedback loops. This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we've taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.

The MiMo-V2.6 series comprises two native fully multimodal models, namely Pro and Flash. Benefiting from the expanded RL computing power, MiMo-V2.6-Pro scores 46 points in the Artificial Analysis Intelligence Index (AA Composite Intelligence Index), surpassing Kimi K3 and Qwen3.8 Max to become the most powerful open-source model available; however, there is still a gap when compared with the strongest closed-source models Claude Fable 5.1 and GPT-6 Astra.

The MiMo-V2.6 series adopts the same API pricing as the V2.5 series. With intelligent performance upgraded while price remains unchanged, the Pareto frontier of "intelligence vs. cost" has thus been pushed outward once again. The MiMo-V2.6-Pro has set a new cost-performance record for domestic large language models: at the same intelligence level, its price is only 1/20 to 1/60 that of overseas models.

## Scale RL on a large scale and fully open-source it

During the RL training phase, MiMo-V2.6 is likely one of the domestic open-source models that has been allocated the largest amount of computing power to date. After large-scale, multi-task reinforcement learning training, MiMo-V2.6-Pro has achieved performance on most Agent Benchmarks that is on par with Claude Opus 5 and GPT-5.6 Sol, while MiMo-V2.6-Flash has comprehensively outperformed MiMo-V2.5-Pro.

Throughout the entire process, we overcame fundamental research and engineering challenges in RL training, and documented the official experimental journey via live sharing. In less than 6 days, MiMo-V2.6-Flash and MiMo-V2.6-Pro completed 30 steps each with a cumulative total of approximately 750,000 trajectories, at training costs of around $850,000 and $2.62 million respectively. The average pass rate of training tasks has been relatively improved by 25% and 12% respectively, and the out-of-sample long-range software engineering evaluation benchmark DeepSWE v1.1 has been improved by approximately 17 points (from 48.8 to 65.7) and approximately 14 points (from 58.4 to 72.6) respectively.

This training mainly expands RL computing power from three dimensions:

1. **Larger Batch Size and Higher Throughput:** By combining a large batch size and a fully asynchronous architecture, each update uses 1,568 samples, supports training with 1M context length, and the number of tokens per training step reaches 3.5–3.7B.
2. **More Tasks and Complex Environments:** Build a multi-task training system covering fields such as Code, General, Visual, and Cyber, and integrate multiple Harnesses to facilitate the collaborative improvement of different capability dimensions.
3. **Greater Grader computing power:** through relative comparison within the Group, it provides more accurate and diverse reward signals for Long-Horizon RL tasks, forms a closed loop for model self-improvement, and guides the model to complete tasks with shorter paths and fewer Tokens.

As the training scale expands, we freeze the MoE Router to suppress expert load drift, and establish a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators, so as to improve training stability and reward reliability.

We have open-sourced the aforementioned technical achievements and supporting resources, including the complete technical report, training environment and RL code, to help more researchers reproduce and verify relevant results.

## From Vibe Coding to Vibe World

MiMo-V2.6 integrates 3D spatial reasoning, multimodal perception and computer user operation (CUA) capabilities, further expanding the boundaries of what programming can achieve; it can extend natural language-driven programming tasks into the "Vibe World" oriented towards interactive world construction—including 3D open-world games, Blender 3D modeling, embodied intelligence with a Franka Panda arm, and Computer Use Agents.

## Advance cutting-edge scientific research

Without undergoing specialized reinforcement learning tailored for scientific research tasks, MiMo-V2.6 has already demonstrated application potential across multiple research domains—from materials design (MOF candidates for PFAS adsorption) to formal mathematical proof (Lean 4 formalization of Li and Yorke's *Period Three Implies Chaos*).

## Fully open source

We have fully open-sourced the weights and technical report of the MiMo-V2.6-Pro and Flash models, simultaneously released the MiMo-V2.6-Distill-Qwen-9B along with supporting reinforcement learning (RL) research resources:

- 7k+ high-quality RL task environments covering software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development
- End-to-end RL training framework built on verl, uni-agent and mini-swe-agent
- Lightweight and composable mini-harnesses

Open Source Link: https://huggingface.co/collections/XiaomiMiMo/mimo-v26

> Note: When calling the API, please use the all-lowercase model names `mimo-v2.6-pro`, `mimo-v2.6-flash`, and `mimo-v2.6-pro-ultraspeed`.

Update Time: September 22, 2026
