‹ 返回 2026-06-24

世界行动模式:一项调查

World Action Models: A Survey

▲ 37 💬 1 2026-06-24

Qiuhong Shen, Shihua Zhang, Yue Liao, Qi Li, Zhenxiong Tan, Shizun Wang, Shuicheng Yan, Xinchao Wang

摘要

世界行动模型(WAMs)是一种具身预测-行动模型,能够将对未来的预测转化为可执行的行动。最新的WAMs则利用大型视频生成模型进行改造;而另一种类型则依靠语言或视觉-语言核心机制来实现视频生成功能,而不需要专门的视频生成模块。这种快速发展使得各种模型之间的界限变得模糊——比如广义世界模型、视频生成模型、以行动为目标的视频世界模型、视觉-语言-行动策略以及WAMs之间。本综述旨在为这一领域提供一个统一的框架。首先,它明确了这些不同模型之间的区别;然后,通过两种互补的视角来整理现有研究内容。第一种视角关注每种模型在生成过程中需要具备哪些能力,包括渲染出的未来场景、潜在的未来状态以及无需视频生成功能的行动推理能力。第二种视角则根据预测基础、核心机制、行动关联以及应用方式来分类各模型。这种分类方式有助于我们统一讨论交互性、因果关系、持久性、物理合理性以及泛化能力等问题。此外,还有数据处理、评估方法以及当前面临的挑战等相关内容。在这些方面中,一种一致的设计模式逐渐显现出来:WAMs并非简单的带有行动功能的视频生成器,而是那些需要在表示丰富度与计算资源、内存使用、延迟时间以及动作标签成本之间做出权衡的预测-行动方法。该领域正在朝着那种能够生成较少未来信息的同时又能保留必要控制能力的方向发展。该综述的官方网站可在https://world-action-models.github.io/查看。

English Abstract

World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large video generation models, and a parallel line relies on language or vision-language backbones without a video-generation core. This rapid expansion has blurred the boundary among broad world models, video generation models, action-grounded video world models, Vision-Language-Action policies, and WAMs. This survey gives the field a common account. It first clarifies these boundaries, then organizes existing works through two complementary views. The first view asks what each method is required to generate, spanning rendered futures, latent futures, and video-generation-free action reasoning. The second view decomposes each method by predictive substrate, backbone, action coupling, and deployment regime. This anatomy supports a unified discussion of interactability, causality, persistence, physical plausibility, and generalization, followed by data, evaluation, and open challenges. Across these axes, a consistent design pattern emerges: WAMs are not simply video generators with action heads, but predictive-action methods whose design choices trade representational richness against compute, memory, latency, and action-label cost. The field is moving toward methods that generate less of the future while preserving what control requires. The survey homepage is available at https://world-action-models.github.io/.