Today, we are officially releasing and open-sourcing the Xiaomi MiMo-V2.6 series. This marks a key step in our exploration of the RSI (recursive self-improvement) path: building on verifiable complex tasks, we scale up reinforcement learning (RL) computing power to enable models to continuously expand the boundaries of intelligence through ongoing exploration and feedback.

Where the path is flat and close, travelers are many; where it is rugged and distant, few reach the end. In an era where intelligence can be easily replicated, we choose to channel computing power into real-world environments, letting models learn through trial and error in iterative feedback loops. This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we've taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.