{"article":{"slug":"mimo-v2-6-scaling-up-reinforcement-learning-for-self-improvement","title":"MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement","subtitle":null,"summary":"Xiaomi open-sources MiMo-V2.6 Pro and Flash after large-scale live RL (~$3.5M, 750k trajectories), claiming top open-weight AA Index scores, agent parity with frontier models, and 7k+ RL environments.","content_type":"announcement","language":"en","canonical_url":"https://mimo.mi.com/docs/en-US/news/latest/v2-6","author":{"name":"Xiaomi MiMo","url":"https://mimo.mi.com","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Xiaomi MiMo","url":"https://mimo.mi.com","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":839,"reading_minutes":4,"published_at":"2026-09-22T00:00:00.000Z","added_at":"2026-09-22T03:12:27.455Z","updated_at":"2026-09-22T03:12:27.455Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/mimo-v2-6-scaling-up-reinforcement-learning-for-self-improvement","markdown_url":"https://listedarticles.com/articles/mimo-v2-6-scaling-up-reinforcement-learning-for-self-improvement.md","example":false,"citation":"Xiaomi MiMo, Xiaomi MiMo. \"MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement.\" 22 Sept 2026. https://mimo.mi.com/docs/en-US/news/latest/v2-6 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://mimo.mi.com/docs/en-US/news/latest/v2-6"},"body_markdown":"Today, we are officially releasing and open-sourcing the Xiaomi MiMo-V2.6 series. This marks a key step in our exploration of the RSI (recursive self-improvement) path: building on verifiable complex tasks, we scale up reinforcement learning (RL) computing power to enable models to continuously expand the boundaries of intelligence through ongoing exploration and feedback.\n\nWhere the path is flat and close, travelers are many; where it is rugged and distant, few reach the end. In an era where intelligence can be easily replicated, we choose to channel computing power into real-world environments, letting models learn through trial and error in iterative feedback loops. This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we've taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.\n\nThe MiMo-V2.6 series comprises two native fully multimodal models, namely Pro and Flash. Benefiting from the expanded RL computing power, MiMo-V2.6-Pro scores 46 points in the Artificial Analysis Intelligence Index (AA Composite Intelligence Index), surpassing Kimi K3 and Qwen3.8 Max to become the most powerful open-source model available; however, there is still a gap when compared with the strongest closed-source models Claude Fable 5.1 and GPT-6 Astra.\n\nThe MiMo-V2.6 series adopts the same API pricing as the V2.5 series. With intelligent performance upgraded while price remains unchanged, the Pareto frontier of \"intelligence vs. cost\" has thus been pushed outward once again. The MiMo-V2.6-Pro has set a new cost-performance record for domestic large language models: at the same intelligence level, its price is only 1/20 to 1/60 that of overseas models.\n\n## Scale RL on a large scale and fully open-source it\n\nDuring the RL training phase, MiMo-V2.6 is likely one of the domestic open-source models that has been allocated the largest amount of computing power to date. After large-scale, multi-task reinforcement learning training, MiMo-V2.6-Pro has achieved performance on most Agent Benchmarks that is on par with Claude Opus 5 and GPT-5.6 Sol, while MiMo-V2.6-Flash has comprehensively outperformed MiMo-V2.5-Pro.\n\nThroughout the entire process, we overcame fundamental research and engineering challenges in RL training, and documented the official experimental journey via live sharing. In less than 6 days, MiMo-V2.6-Flash and MiMo-V2.6-Pro completed 30 steps each with a cumulative total of approximately 750,000 trajectories, at training costs of around $850,000 and $2.62 million respectively. The average pass rate of training tasks has been relatively improved by 25% and 12% respectively, and the out-of-sample long-range software engineering evaluation benchmark DeepSWE v1.1 has been improved by approximately 17 points (from 48.8 to 65.7) and approximately 14 points (from 58.4 to 72.6) respectively.\n\nThis training mainly expands RL computing power from three dimensions:\n\n1. **Larger Batch Size and Higher Throughput:** By combining a large batch size and a fully asynchronous architecture, each update uses 1,568 samples, supports training with 1M context length, and the number of tokens per training step reaches 3.5–3.7B.\n2. **More Tasks and Complex Environments:** Build a multi-task training system covering fields such as Code, General, Visual, and Cyber, and integrate multiple Harnesses to facilitate the collaborative improvement of different capability dimensions.\n3. **Greater Grader computing power:** through relative comparison within the Group, it provides more accurate and diverse reward signals for Long-Horizon RL tasks, forms a closed loop for model self-improvement, and guides the model to complete tasks with shorter paths and fewer Tokens.\n\nAs the training scale expands, we freeze the MoE Router to suppress expert load drift, and establish a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators, so as to improve training stability and reward reliability.\n\nWe have open-sourced the aforementioned technical achievements and supporting resources, including the complete technical report, training environment and RL code, to help more researchers reproduce and verify relevant results.\n\n## From Vibe Coding to Vibe World\n\nMiMo-V2.6 integrates 3D spatial reasoning, multimodal perception and computer user operation (CUA) capabilities, further expanding the boundaries of what programming can achieve; it can extend natural language-driven programming tasks into the \"Vibe World\" oriented towards interactive world construction—including 3D open-world games, Blender 3D modeling, embodied intelligence with a Franka Panda arm, and Computer Use Agents.\n\n## Advance cutting-edge scientific research\n\nWithout undergoing specialized reinforcement learning tailored for scientific research tasks, MiMo-V2.6 has already demonstrated application potential across multiple research domains—from materials design (MOF candidates for PFAS adsorption) to formal mathematical proof (Lean 4 formalization of Li and Yorke's *Period Three Implies Chaos*).\n\n## Fully open source\n\nWe have fully open-sourced the weights and technical report of the MiMo-V2.6-Pro and Flash models, simultaneously released the MiMo-V2.6-Distill-Qwen-9B along with supporting reinforcement learning (RL) research resources:\n\n- 7k+ high-quality RL task environments covering software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development\n- End-to-end RL training framework built on verl, uni-agent and mini-swe-agent\n- Lightweight and composable mini-harnesses\n\nOpen Source Link: https://huggingface.co/collections/XiaomiMiMo/mimo-v26\n\n> Note: When calling the API, please use the all-lowercase model names `mimo-v2.6-pro`, `mimo-v2.6-flash`, and `mimo-v2.6-pro-ultraspeed`.\n\nUpdate Time: September 22, 2026\n","body_html":"<p>Today, we are officially releasing and open-sourcing the Xiaomi MiMo-V2.6 series. This marks a key step in our exploration of the RSI (recursive self-improvement) path: building on verifiable complex tasks, we scale up reinforcement learning (RL) computing power to enable models to continuously expand the boundaries of intelligence through ongoing exploration and feedback.</p>\n<p>Where the path is flat and close, travelers are many; where it is rugged and distant, few reach the end. In an era where intelligence can be easily replicated, we choose to channel computing power into real-world environments, letting models learn through trial and error in iterative feedback loops. This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we&#39;ve taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.</p>\n<p>The MiMo-V2.6 series comprises two native fully multimodal models, namely Pro and Flash. Benefiting from the expanded RL computing power, MiMo-V2.6-Pro scores 46 points in the Artificial Analysis Intelligence Index (AA Composite Intelligence Index), surpassing Kimi K3 and Qwen3.8 Max to become the most powerful open-source model available; however, there is still a gap when compared with the strongest closed-source models Claude Fable 5.1 and GPT-6 Astra.</p>\n<p>The MiMo-V2.6 series adopts the same API pricing as the V2.5 series. With intelligent performance upgraded while price remains unchanged, the Pareto frontier of &quot;intelligence vs. cost&quot; has thus been pushed outward once again. The MiMo-V2.6-Pro has set a new cost-performance record for domestic large language models: at the same intelligence level, its price is only 1/20 to 1/60 that of overseas models.</p>\n<h2 id=\"scale-rl-on-a-large-scale-and-fully-open-source-it\">Scale RL on a large scale and fully open-source it</h2>\n<p>During the RL training phase, MiMo-V2.6 is likely one of the domestic open-source models that has been allocated the largest amount of computing power to date. After large-scale, multi-task reinforcement learning training, MiMo-V2.6-Pro has achieved performance on most Agent Benchmarks that is on par with Claude Opus 5 and GPT-5.6 Sol, while MiMo-V2.6-Flash has comprehensively outperformed MiMo-V2.5-Pro.</p>\n<p>Throughout the entire process, we overcame fundamental research and engineering challenges in RL training, and documented the official experimental journey via live sharing. In less than 6 days, MiMo-V2.6-Flash and MiMo-V2.6-Pro completed 30 steps each with a cumulative total of approximately 750,000 trajectories, at training costs of around $850,000 and $2.62 million respectively. The average pass rate of training tasks has been relatively improved by 25% and 12% respectively, and the out-of-sample long-range software engineering evaluation benchmark DeepSWE v1.1 has been improved by approximately 17 points (from 48.8 to 65.7) and approximately 14 points (from 58.4 to 72.6) respectively.</p>\n<p>This training mainly expands RL computing power from three dimensions:</p>\n<ol><li><strong>Larger Batch Size and Higher Throughput:</strong> By combining a large batch size and a fully asynchronous architecture, each update uses 1,568 samples, supports training with 1M context length, and the number of tokens per training step reaches 3.5–3.7B.</li><li><strong>More Tasks and Complex Environments:</strong> Build a multi-task training system covering fields such as Code, General, Visual, and Cyber, and integrate multiple Harnesses to facilitate the collaborative improvement of different capability dimensions.</li><li><strong>Greater Grader computing power:</strong> through relative comparison within the Group, it provides more accurate and diverse reward signals for Long-Horizon RL tasks, forms a closed loop for model self-improvement, and guides the model to complete tasks with shorter paths and fewer Tokens.</li></ol>\n<p>As the training scale expands, we freeze the MoE Router to suppress expert load drift, and establish a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators, so as to improve training stability and reward reliability.</p>\n<p>We have open-sourced the aforementioned technical achievements and supporting resources, including the complete technical report, training environment and RL code, to help more researchers reproduce and verify relevant results.</p>\n<h2 id=\"from-vibe-coding-to-vibe-world\">From Vibe Coding to Vibe World</h2>\n<p>MiMo-V2.6 integrates 3D spatial reasoning, multimodal perception and computer user operation (CUA) capabilities, further expanding the boundaries of what programming can achieve; it can extend natural language-driven programming tasks into the &quot;Vibe World&quot; oriented towards interactive world construction—including 3D open-world games, Blender 3D modeling, embodied intelligence with a Franka Panda arm, and Computer Use Agents.</p>\n<h2 id=\"advance-cutting-edge-scientific-research\">Advance cutting-edge scientific research</h2>\n<p>Without undergoing specialized reinforcement learning tailored for scientific research tasks, MiMo-V2.6 has already demonstrated application potential across multiple research domains—from materials design (MOF candidates for PFAS adsorption) to formal mathematical proof (Lean 4 formalization of Li and Yorke&#39;s <em>Period Three Implies Chaos</em>).</p>\n<h2 id=\"fully-open-source\">Fully open source</h2>\n<p>We have fully open-sourced the weights and technical report of the MiMo-V2.6-Pro and Flash models, simultaneously released the MiMo-V2.6-Distill-Qwen-9B along with supporting reinforcement learning (RL) research resources:</p>\n<ul><li>7k+ high-quality RL task environments covering software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development</li><li>End-to-end RL training framework built on verl, uni-agent and mini-swe-agent</li><li>Lightweight and composable mini-harnesses</li></ul>\n<p>Open Source Link: <a href=\"https://huggingface.co/collections/XiaomiMiMo/mimo-v26\" rel=\"nofollow ugc noopener\">https://huggingface.co/collections/XiaomiMiMo/mimo-v26</a></p>\n<blockquote><p>Note: When calling the API, please use the all-lowercase model names <code>mimo-v2.6-pro</code>, <code>mimo-v2.6-flash</code>, and <code>mimo-v2.6-pro-ultraspeed</code>.</p></blockquote>\n<p>Update Time: September 22, 2026</p>","headings":[{"level":2,"text":"Scale RL on a large scale and fully open-source it","id":"scale-rl-on-a-large-scale-and-fully-open-source-it"},{"level":2,"text":"From Vibe Coding to Vibe World","id":"from-vibe-coding-to-vibe-world"},{"level":2,"text":"Advance cutting-edge scientific research","id":"advance-cutting-edge-scientific-research"},{"level":2,"text":"Fully open source","id":"fully-open-source"}]}}