Xiaomi-Robotics-1:利用超过10万小时的真实世界轨迹数据来扩展视觉-语言-动作模型
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
摘要
我们介绍了Xiaomi-Robotics-1这一基础性的视觉语言-动作模型。该模型能够:(1) 按照各种语言指令在未知环境中执行多种移动操作任务;以及(2) 在很少的微调数据下就能有效适应新的任务需求。我们提出了一种两阶段训练方法,包括预训练和后训练阶段。在预训练过程中,我们通过使用超过10万小时的真实场景操作数据来训练模型,从而赋予其广泛且可通用的动作生成能力。更重要的是,我们开发了一种可扩展的自标记流程,利用自然语言来描述场景状态变化,从而为动作学习提供丰富且准确的输入信息。在后训练阶段,我们试图让这些能力与人类常用的机器人指令相匹配。大量实验表明,随着数据量和模型规模的扩大,Xiaomi-Robotics-1的性能会得到显著提升。这种性能提升在后训练阶段同样适用——更强的预训练模型能够在未知环境中实现更好的机器人性能。此外,Xiaomi-Robotics-1还可以作为强大的机器人基础模型,通过高效的微调方式来处理复杂的、需要精细操作的任务。在多个模拟测试中,Xiaomi-Robotics-1的表现优于现有最先进的方法。特别是在RoboCasa365测试中,其成功率为57.6%,超过了之前的最佳水平46.6%。在RoboDojo测试中,其平均得分达到了20.07,远远高于之前的最先进水平13.07。代码和模型检查点也将被发布。项目页面:https://robotics.xiaomi.com/xiaomi-robotics-1.html
English Abstract
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html