‹ 返回 2026-07-21

Xiaomi-Robotics-1:利用超过10万小时的真实世界轨迹数据来扩展视觉-语言-动作模型

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

▲ 47 💬 1 2026-07-21

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou

摘要

我们介绍了Xiaomi-Robotics-1这一基础性的视觉语言-动作模型。该模型能够:(1) 按照各种语言指令在未知环境中执行多种移动操作任务;以及(2) 在很少的微调数据下就能有效适应新的任务需求。我们提出了一种两阶段训练方法,包括预训练和后训练阶段。在预训练过程中,我们通过使用超过10万小时的真实场景操作数据来训练模型,从而赋予其广泛且可通用的动作生成能力。更重要的是,我们开发了一种可扩展的自标记流程,利用自然语言来描述场景状态变化,从而为动作学习提供丰富且准确的输入信息。在后训练阶段,我们试图让这些能力与人类常用的机器人指令相匹配。大量实验表明,随着数据量和模型规模的扩大,Xiaomi-Robotics-1的性能会得到显著提升。这种性能提升在后训练阶段同样适用——更强的预训练模型能够在未知环境中实现更好的机器人性能。此外,Xiaomi-Robotics-1还可以作为强大的机器人基础模型,通过高效的微调方式来处理复杂的、需要精细操作的任务。在多个模拟测试中,Xiaomi-Robotics-1的表现优于现有最先进的方法。特别是在RoboCasa365测试中,其成功率为57.6%,超过了之前的最佳水平46.6%。在RoboDojo测试中,其平均得分达到了20.07,远远高于之前的最先进水平13.07。代码和模型检查点也将被发布。项目页面:https://robotics.xiaomi.com/xiaomi-robotics-1.html

English Abstract

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html