‹ 返回 2026-06-18

ACE-Ego-0:将以自我为中心的人类数据与机器人数据整合,用于VLA预训练

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

▲ 40 💬 2 2026-06-18

Hao Li, Ganlong Zhao, Yufei Liu, Haotian Hou, Guoquan Ye, Tongyan Fang, Chunxiao Liu, Siyuan Huang, Jianbo Liu, Xiaogang Wang, Hongsheng Li

摘要

视觉-语言-动作模型能够充分利用大规模且多样化的具体化数据,不过,收集机器人运动轨迹的数据需要耗费大量资源且效率低下。最近的研究表明,大规模以人类为中心的视频数据可以在预训练阶段提供有用的监督信息。然而,由于动作空间、身体结构、时间动态以及监督质量方面的差异,对人和机器人数据进行联合训练仍然具有挑战性。我们提出了ACE-EGO-0这一统一的视觉-语言-动作预训练框架,该框架能够结合不同来源的数据资源。为了从以人类为中心的视频中提取大规模的监督信息,我们构建了一个可扩展的从视频到动作的转换流程,将原始人类视频转换为机器人格式的动作轨迹。为了使这些标签与机器人演示结果保持一致,ACE-EGO-0采用了一种基于相机空间动作、形态条件化以及时间对齐的动作分块处理的统一动作表示方式。为了有效利用来自以人类为中心的视频的噪声性伪动作监督信息,我们设计了一个考虑可靠性的训练目标,同时引入了人类辅助损失机制,从而将监督信息集中在可靠的信号上。我们在4.53千小时的机器人和仿真数据基础上,以及1.48千小时的带有伪动作标签的以人类为中心的视频数据上实现了ACE-EGO-0的实验。实验结果表明,在考虑可靠性因素的情况下,结合大规模的人类监督信息能够有效提升预训练和监督微调的效果。ACE-EGO-0在RoboCasa GR1桌面平台和RoboTwin 2.0平台上取得了领先的性能表现,同时也在真实世界中的双手动操作任务中展现了出色的性能。

English Abstract

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on camera-space actions, morphology conditioning, and time-aligned action chunking. To robustly leverage noisy pseudo-action supervision from egocentric human videos, we formulate a reliability-aware training objective with a human auxiliary loss that concentrates supervision on reliable signals. We instantiate ACE-EGO-0 on 4.53K hours of robot and simulation data, together with 1.48K hours of pseudo-action-labeled egocentric human data. Experiments show that incorporating large-scale human supervision under reliability-aware weighting consistently improves both unified joint pretraining and supervised fine-tuning. ACE-EGO-0 achieves state-of-the-art performance on RoboCasa GR1 TableTop and RoboTwin 2.0, while demonstrating strong transfer to real-world bimanual manipulation.