‹ 返回 2026-06-30

PhysisForcing:用于机器人操作的物理强化世界模拟器

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

▲ 38 2026-06-30

Peiwen Zhang, Yufan Deng, Shangkun Sun, Juncheng Ma, Duomin Wang, Jonas Du, Zilin Pan, Ye Huang, Hao Liang, Songyan Huang, Ruihua Zhang, Enze Xie, Ming-Yu Liu, Daquan Zhou

摘要

视频生成模型已成为实现物理世界模拟的有效方法。不过,无论是通用领域的视频生成器,还是针对特定机器人优化的模型,都可能出现不可信的物理现象,比如不连续的运动轨迹以及不一致的机器人与物体之间的互动关系,这些现象限制了它们作为世界模拟器的可靠性。通过大量实验发现,这种物理上的不稳定性主要源于两个因素:移动物体的变形,以及相互作用对象之间在时空维度上的不合理关联,尤其是在接触过程中。基于这一发现,我们提出了PhysisForcing这一可扩展的训练框架。该框架通过联合优化像素级和语义级特征,将监督重点集中在具有物理意义的区域上,从而增强模型的物理一致性。该框架包含两种损失函数:一种用于监督DiT特征,利用参考点轨迹来实施监督;另一种则用于确保DiT特征与从冻结视频理解编码器中提取的区域间关系保持一致。在R-Bench、PAI-Bench和EZS-Bench上的实验表明,PhysisForcing能够显著提升视频生成的性能,使得R-Bench上的Wan2.2-I2V-A14B和Cosmos3-Nano基础模型分别提升了22.3%和9.2%(相比传统的微调方法分别提升了7.1%和3.7%),而Cosmos3-Nano版本则取得了最优秀的成绩。除了视频生成功能外,作为WorldArena动作规划协议下的世界模型,PhysisForcing还能将闭环成功率从16.0%提升至24.0%,同时还能提高下游策略的成功率,这表明物理上一致的视频模型能够产生更有效的机器人操作表示。

English Abstract

Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models can still produce physically implausible manipulations, including discontinuous motion trajectories and inconsistent robot-object interactions, which limits their reliability as world simulators. Through extensive experiments, we find that such physical instability mainly arises from two factors: deformation of moving objects and implausible spatio-temporal correlations among interacting entities, particularly during contact. Building on this observation, we propose PhysisForcing, a scalable training framework that strengthens physical consistency by focusing supervision on physics-informative regions through joint optimization of pixel-level and semantic-level features. The framework consists of a pixel-level trajectory alignment loss, which supervises DiT features using reference point trajectories, and a semantic-level relational alignment loss, which aligns DiT features with inter-region relations extracted from a frozen video understanding encoder. Extensive experiments on R-Bench, PAI-Bench, and EZS-Bench show that PhysisForcing consistently improves embodied video generation over strong baselines, improving the Wan2.2-I2V-A14B and Cosmos3-Nano base models on R-Bench by 22.3\% and 9.2\% (7.1\% and 3.7\% over vanilla finetuning), with the Cosmos3-Nano variant attaining the best overall score. Beyond generation, as a world model under the WorldArena action-planner protocol it raises the closed-loop success rate from 16.0\% to 24.0\% and further improves downstream policy success, indicating that physically aligned video models yield stronger representations for robotic manipulation.