On July 16, Xiaomi officially released the Xiaomi-Robotics-1, a embodied foundation model designed for real mobile operation tasks. The model was pre-trained on 100,000 hours of real-world data and completed training by combining cross-body data, marking a systematic step forward for Xiaomi in advancing embodied intelligence models along the "Scaling Law" (scale law) path.

Traditional robot strategy models are often limited by hardware dependencies and scarce data scales. To break through this bottleneck, the Xiaomi team introduced 100,000 hours of real-world trajectories collected through the UMI (Universal Manipulation Interface) device during the pre-training phase, covering multiple scenarios such as home, commercial, and industrial environments, and combined it with an efficient visual language model to complete full-scale automatic annotation within two weeks.

In the post-training phase, the team used approximately 10,000 hours of cross-body data for body and instruction alignment, enabling Xiaomi-Robotics-1 to have "out-of-the-box" multi-type mobile operation capabilities. Experiments show that as the amount of training data and model size (offering three versions: 2B, 5B, and 10B) increases, the model demonstrates clear scaling growth trends in action prediction accuracy and success rate of unseen scenario tasks, and has set new SOTA (state-of-the-art) records on multiple public simulation benchmarks such as RoboCasa365 and RoboDojo.
This release of Xiaomi-Robotics-1 not only demonstrates Xiaomi's R&D strength in the field of physical AI but also successfully verifies a scalable embodied intelligence training path: "large-scale pre-training - cross-body post-training - fine-tuning with a small amount of data," providing a highly referenceable paradigm for robots to transition from lab demonstrations to complex and realistic physical worlds.




