EgoPhys:从以自我为中心的视频中学习可泛化的可变形物体物理模型
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
摘要
人类通过日常互动自然而然地理解物体的物理特性。然而,要准确预测诸如弹性材料、织物等复杂可变形物体的动态行为,仍然是一个巨大的挑战,这对计算机视觉和机器人技术来说都是一个严峻的考验。我们提出了EgoPhys框架——该框架能够利用泛化的先验知识,从以观察者为中心的RGB视频中构建出可变形物体的数字孪生模型。EgoPhys克服了现有方法的局限性,使得从以观察者为中心的视频中生成可变形物体的数字孪生成为可能。通过将每个物体的逆物理方程压缩成紧凑的代码集,EgoPhys能够预测那些尚未被观测到的物体的弹簧刚度场,而无需进行针对每个弹簧的优化计算。通过使用来自各种以观察者为中心的交互场景中的泛化先验信息,EgoPhys在物体重建、未来预测以及零样本泛化能力方面均优于传统方法。为了支持训练和评估过程,我们整理了一个包含多种可变形物体、场景及操作方式的以观察者为中心的交互数据集。我们在真实的xArm6机器人上实现了EgoPhys框架,结果表明:从单个以观察者为中心的视频中初始化出的数字孪生模型可以作为内部世界表示,从而帮助规划可变形物体的运动路径。这表明,以观察者为中心的RGB观测数据是一种可行的解决方案,可以应用于从现实世界到仿真环境的迁移任务中。
English Abstract
Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformable dynamics, such as elastic materials and fabrics, remains a major challenge for computer vision and robotics. We present EgoPhys, a framework that constructs deformable physical digital twins from egocentric RGB-only video using generalizable priors. EgoPhys overcomes the limitations of existing methods to enable controllable deformable digital twin generation from egocentric videos by distilling per-object inverse-physics solutions into a compact codebook, enabling prediction of dense spring stiffness fields for unseen objects without per-spring test-time optimization. Trained with generalizable priors from diverse egocentric interactions, EgoPhys outperforms baselines in reconstruction, future prediction, and zero-shot generalization. To support training and evaluation, we curate an egocentric interaction dataset covering diverse deformable objects, scenes, and manipulation styles. We deploy EgoPhys on a real xArm6 robot, demonstrating that a digital twin initialized from a single egocentric human play video can serve as an internal world representation to aid in deformable-object planning, highlighting egocentric RGB observations as a scalable path toward real-to-sim pipelines.