‹ 返回 2026-06-30

翻译作为一种连接手段:将人类的操作技能转移给机器人

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

▲ 32 💬 1 2026-06-30

Sijin Chen, Kaixuan Jiang, Haixin Shi, Yanhui Wang, Weiheng Zhong, Haosheng Li, Bo Jiang, Yuxiao Liu, Xihui Liu

摘要

我们研究能否从人类的行为中学习新的操作技能,从而让具有平行夹持器的生物手动机器人能够掌握这些技能。人类行为数据价格低廉、数量丰富且种类多样,因此成为扩展机器人学习能力的有效资源。不过,将人类的技能转移到机器人身上仍然很困难:大多数相关研究都把人类视为另一种生物手动6DoF执行器,此时手部姿态的估计值存在噪声,而人类手指的接触模式与平行夹持器的接触模式也有本质区别。我们认为,从人类数据中学习包含旋转行为的动作信号并不理想,因此我们提出了一种过渡性动作表示方式——即初始头部-相机框架内的手腕相对位移,这是一种人类和机器人都共有的动作空间。为了应对不同执行器中某些动作成分的缺失问题,我们构建了一个类似π_0的视觉-语言-动作模型,该模型使用交错排列的动作标记以及注意力掩码机制。在一系列新的生物手动操作任务中,这种过渡性动作表示方式能够比噪声较大的6DoF人类动作更有效地将人类的操作知识传递给机器人,而且其效果还随着人类数据的数量增加而提升。

English Abstract

We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, abundant, and diverse, making it one of the most promising resources for scaling up robot learning. Yet transferring skills from humans to robots remains hard: most prior work treats humans as just another bi-manual 6DoF embodiment, where hand-pose estimates are noisy and the contact patterns of human fingers differ fundamentally from those of a parallel gripper. We argue that learning rotation-inclusive action signals from human data is therefore sub-optimal, and instead propose a bridging action representation: the relative wrist translation within the initial head-camera frame, an action space shared by humans and robots. To handle the potential absence of certain action components in different embodiments, we build a π_0-like vision-language-action model with interleaved action tokens and attention masking. On a suite of novel bi-manual manipulation tasks, our bridging action transfers human manipulation knowledge to robots far more effectively than noisy 6DoF human actions and scales with the amount of human data.