An open-source robot foundation model combines large-scale pre-training, real-world robot data, and natural-language instruction learning to improve robot manipulation and adaptation across diverse environments.
Xiaomi has introduced Xiaomi-Robotics-1, an open-source vision-language-action foundation model designed to improve robotic manipulation through large-scale pre-training and real-robot alignment. Trained on more than 100,000 hours of real-world manipulation trajectories collected across over 1,700 scenarios, the model is intended to perform a wide range of mobile manipulation tasks while adapting efficiently to new applications with minimal additional training.
This model uses a two-step training approach. In the pre-training step, it acquires generic action generation skills using embodiment-free manipulation data, while in the post-training phase, the skills learned by the model are aligned with the embodiment and natural language instructions. According to Xiaomi, an increase in the amount of training data and the size of the model improves real-world robotic performance and there are no signs of performance saturation when scaling up the model.
After the post-training step, the foundation model can carry out manipulation tasks in unseen environments and new tasks with just a few hours of demonstration data. According to Xiaomi, there have been successful demonstrations of manipulation activities like packing, sorting, and manipulating objects in households while maintaining state-of-the-art performance on several robotics simulators used for generalisation benchmarking.
Additionally, a scalable pipeline for automatic data labelling is developed to label manipulation trajectories with language descriptions of scene changes, making it possible to efficiently learn from large datasets. The project also makes the technical report, source code, and model resources available to encourage further research in embodied AI.
By using large datasets, scalable training techniques, and open-source resources, the project shows how to develop better robot foundation models by reducing barriers to research in embodied AI.















































































