‹ 返回 2026-06-21

当前的世界模型缺乏一个稳定的核心状态

Current World Models Lack a Persistent State Core

▲ 9 💬 1 2026-06-21

Jinpeng Lu, Dexu Zhu, Haoyuan Shi, Linghan Cai, Guo Tang, Yinda Chen, Jie Cao, Duyu Tang, Yi Zhang, Yong Dai, Xiaozhu Ju

摘要

世界模型越来越被视为实现人工通用智能的关键步骤。然而,对物理世界的建模并非只是按需生成令人信服的画面而已:它还需要一个随时间不断变化的内部世界状态,这种状态与观察行为无关。如此一来,无论相机是否在观察,物体都能持续存在,事件也能按其逻辑发展下去——就像月亮在无人注视时仍然保持其轨道一样。这一要求是现有评估标准的不足之处:这些标准只关注表面特性,如真实性、运动效果以及相机的控制能力,却从未考虑生成的世界在未被观察后是否仍会继续变化。我们提出了WRBench这一首个系统性的评估基准,它将相机运动视为对可观察性的干预手段,并将评估过程转化为以人类标准为依据的评估体系,即判断相机是否执行了所需的互动操作、场景在可见状态下是否保持连续且可识别,以及返回的目标是否与所引发的事件保持一致。在来自23种模型的9,600个视频中,我们发现了一个重要结论:当前系统将观察到的世界视为一个跟踪场景,当目标被忽略后,它仍然保持原来的状态,而不是在目标被忽视后继续推进事件的发展。由于这种缺陷在不同控制模式、不同模型版本以及不同规模下都会出现,因此,稳定的世界状态演化并不能由更清晰的图像、更严格的控制、更丰富的几何先验或更多的参数数量来实现。因此,我们认为,物理状态的稳定性以及在不同视角下的世界线的一致性应成为世界模型设计的首要目标,这样,世界模型才能准确反映世界的实际发展情况,而非仅仅描述下一帧的画面内容。

English Abstract

World models are increasingly regarded as a decisive step toward artificial general intelligence, yet modeling the physical world demands more than rendering convincing frames on demand: it requires an internal world state that keeps evolving over time, decoupled from observation, so that objects endure and events run to their conclusions whether or not a camera is watching, much as the moon holds to its orbit when no one is looking. This requirement is a blind spot of existing benchmarks, which reward surface properties such as fidelity, motion, and camera controllability while never asking whether a generated world keeps evolving once it is unobserved. We introduce WRBench, the first systematic diagnostic benchmark that treats camera motion as an intervention on observability and resolves evaluation into a human-calibrated chain that asks whether the camera executes the requested interaction, whether the scene stays continuous and identifiable while in view, and whether a returning target remains consistent with the event that was set in motion. Across 9{,}600 videos from 23 models spanning four control paradigms, one finding proves stubborn: current systems maintain the observed world as a tracking shot, resuming a returning target in the state at which it was abandoned rather than advancing the event while it went unseen. Because this failure recurs across control paradigms, model families, and increments of scale, robust world-state evolution does not follow from cleaner imagery, tighter control, richer geometric priors, or sheer parameter count We therefore argue that the stability of the physical state kernel and the consistency of worldlines under viewpoint intervention should become first-class objectives of world-model design, so that a world model captures how the world will unfold rather than how the next frame appears.