MemoBench:在动态变化的环境中对全球模型进行基准测试
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
摘要
视频生成模型旨在模拟动态环境,目前已有多种基准测试用于评估不同帧之间的记忆一致性。不过,大多数测试仅针对目标物体仍然处于视野中的情况来评估一致性;而那些要求物体脱离视野的测试则适用于那些在遮挡期间没有任何变化的静态场景。为了填补这一空白,我们提出了MemoBench这一诊断性基准测试——该测试基于“物体消失与重新出现”的模式,即目标物体经历某种物理过程后从视野中消失,然后在重新出现时必须以更新后的状态被正确恢复。我们收集了来自合成场景和真实场景的360度真实视频片段,并设计了一套评估体系,该体系结合了自动化指标与基于VQA的评估方法,涵盖四个关键评估维度。对八种最先进的模型的评估结果显示了关于“物体消失与重新出现”模式下记忆一致性的重要见解以及存在的挑战。
English Abstract
Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few that force objects out of view evaluate static scenes where nothing changes during occlusion. To bridge this gap, we introduce MemoBench, a diagnostic benchmark built around the disappear-and-reappear paradigm in dynamically changing environments: a target object undergoes a physical process, disappears from view, and must be correctly recovered in its updated state upon reappearance. We curate 360 ground-truth clips spanning synthetic and real-world scenes, and design an evaluation suite combining automated metrics with VQA-based assessment across four diagnostic pillars. Evaluation of eight state-of-the-art models reveals key insights and open challenges regarding memory consistency under the disappear-and-reappear paradigm.