MolmoMotion:通过语言教学进行3D中点轨迹预测
MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction
摘要
运动预测是视觉智能的核心:智能体必须预测物体未来的运动情况,才能规划行动、理解物理相互作用,并生成更真实的未来场景。我们认为,以世界坐标表示的3D点能够提供一种通用的表示方式——这种表示方式与视角无关、易于处理、体积较小,且可直接用于后续任务。我们正式定义了基于目标的3D点运动预测任务:给定较短的视觉历史信息、目标物体上的若干3D查询点以及目标的描述信息,模型就能预测每个点的未来3D运动轨迹。我们还提出了一种完整的解决方案来研究这一任务:(1) MolmoMotion-1M是一个包含1.16百万个无约束视频的3D点运动数据集,这些数据都带有动作描述和物体相关的信息;(2) PointMotionBench是一个经过人类验证的基准数据集,涵盖了111种物体类别和61种运动类型;(3) MolmoMotion则是一种通用的运动预测模型,它既支持自回归坐标预测,也支持基于流匹配的方法生成运动轨迹。MolmoMotion能够准确预测各种不同指令下的运动模式,其性能显著优于PointMotionBench上现有的运动预测模型。最后,我们表明所学习的3D运动先验知识可以很好地应用于后续任务中:它可以提高机器人操作任务的训练效率与泛化能力,而其预测的轨迹能够为生成式模型提供有效的运动指导,从而生成更真实的物体运动效果。
English Abstract
Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize realistic futures. We argue that 3D points in world coordinates provide a general representation that is class-agnostic, view-stable, compact, and directly useful for downstream tasks. We formalize the task of goal-conditioned 3D point motion forecasting: given a short visual history, a set of 3D query points on an object of interest, and a language description of the intended goal, the model predicts the future 3D trajectory of each point. We introduce a full stack to study this task at scale: (1) MolmoMotion-1M is a large corpus of action-described, object-grounded 3D point trajectories annotated from 1.16M unconstrained videos; (2) PointMotionBench is a human-verified benchmark spanning 111 object categories and 61 motion types; and (3) MolmoMotion is a general motion forecasting model that supports both autoregressive coordinate prediction and flow-matching-based trajectory generation. MolmoMotion accurately predicts diverse motion patterns with different language instructions, and significantly outperforms existing motion prediction baselines on PointMotionBench. Finally, we show that the learned 3D motion prior transfers well to downstream applications: it improves training efficiency and generalization for robot manipulation, and its predicted trajectories provide effective motion guidance for generative models to synthesize videos with more realistic object motion.