Xiaomi officially released and open-sourced the native multimodal model MiMo-V2.6 series (Pro/Flash) on September 22, scaling up reinforcement learning (RL) computing power based on verifiable complex tasks and exploring recursive self-improvement (RSI) paths.

QQ20260922-092637.jpg

MiMo-V2.6-Pro scored 46 points in the Artificial Analysis comprehensive intelligence index, surpassing Kimi K3 and Qwen3.8Max to become the strongest open-source model currently, with most Agent benchmarks comparable to Claude Opus5 and GPT-5.6Sol, but still lagging behind the strongest closed-source models.

Both models completed Live RL training in less than 6 days, with costs of approximately $850,000 and $2.62 million, totaling about 750,000 trajectories. The software engineering benchmark DeepSWE v1.1 improved by about 17 points and 14 points respectively. The models added 3D spatial reasoning and computer operation (CUA) capabilities, enabling tasks such as 3D modeling and robotic arm control. MiMo-V2.6 continues to use the V2.5 API pricing, with prices at 1/20 to 1/60 of overseas models for the same level of intelligence. The MiMo Desktop client was launched simultaneously, with UltraSpeed mode output speed increasing up to 20 times.

In addition, Xiaomi simultaneously open-sourced model weights, technical reports, 7k+ RL task environments, and end-to-end training frameworks, making the path and cost of large-scale RL transparent, providing a reproducible public foundation for Agentic RL research.

Open-source link:

https://huggingface.co/collections/XiaomiMiMo/mimo-v26