‹ 返回 2026-06-21

HumanScale:以自我为中心的人体视频在实体预训练中能够优于真实机器人数据

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

▲ 5 💬 2 2026-06-21

Juncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai, Kewei Zhang, Ye Huang, Bo Liang, Shukai Gong, Jiankai Tu, Xiaotian Tang, Jiaxin Li, Kaiqi Chen, Duomin Wang, Yuqi Wang, Bingyi Kang, Eric Huang, Zhiyang Dou, Zhen Dong, Enze Xie, Wojciech Matusik, Tat-Seng Chua, Daquan Zhou

摘要

具身基础模型有望像大型语言模型一样从数据规模扩展中受益,但面临更为严格的数据限制问题。由于具备精确的动作监督机制以及与真实机器人的一致性,远程操控的机器人轨迹仍被作为预训练的数据来源,但其可扩展性受到高收集成本、数据获取难度以及行为和环境多样性不足等因素的限制。这些限制使得以人类视角拍摄的视频成为一种更可扩展、成本更低且更具多样性的预训练数据来源。不过,与远程操控的机器人数据相比,这种方法的有效性尚未得到充分研究。为了解答这个问题,我们在固定的训练后评估流程下,对以人类视角拍摄的视频和远程操控的机器人轨迹作为具身基础模型的预训练数据来源进行了系统比较研究。令人惊讶的是,我们发现通过精心设计的过滤和标注流程处理以人类视角拍摄的数据,不仅可以作为预训练的替代方案,还能带来更好的性能表现。在相同的预训练数据量下,使用以人类视角拍摄的数据进行训练的模型在真实机器人动作预测方面的验证损失降低了24%,而在真实机器人任务执行中的成功率则分别提高了52.5%和90%。这一发现证明了一种可扩展的具身基础模型训练方法:先以人类视角拍摄的视频进行预训练,以学习多样化的世界表示;然后利用少量经过标注的机器人数据来调整模型,使其与动作空间更加一致。我们希望这项研究能够促进对以人类视角拍摄的数据的进一步探索,并为昂贵的机器人数据收集前的数据质量评估提供指导。

English Abstract

Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and validation protocols. Surprisingly, we find that egocentric data, when processed through a carefully designed filtering and labeling pipeline, is not merely a viable substitute for model pretraining but can lead to superior performance. With the same amount of pretraining data, models pretrained on egocentric data achieve a 24% lower validation loss on real-robot action prediction, as well as 52.5% and 90% higher success rates on in-distribution and out-of-distribution real-robot task execution, respectively. This finding verifies a scalable paradigm for embodied foundation models: pretrain on egocentric human video to learn diverse world representations, then adapt with a small amount of labeled real-robot data for action-space alignment. We hope this study encourages broader exploration of egocentric data and offers guidance for data quality assessment before costly robot data collection.