‹ 返回 2026-06-25

机器人操作的世界价值模型

World Value Models for Robotic Manipulation

▲ 1 2026-06-25

Zhihao Wang, Jianxiong Li, Yu Cui, Yuan Gao, Xianyuan Zhan, Junzhi Yu, Xiao Ma

摘要

通用价值模型在从大规模、质量不一的数据中实现机器人策略学习方面发挥着关键作用。从数学角度来看,准确的价值估计需要深入的时间理解能力,即模型既要利用历史背景来确认当前认知,又要能够预测未来结果。然而,目前大多数机器人价值模型都是基于视觉-语言模型构建的,这些模型主要是基于静态或时间上稀疏的视觉数据进行预训练的,因此缺乏进行价值估计所需的时间建模能力。与视觉-语言模型不同,世界模型天然具备时间建模和未来规划的能力,因此非常适合用于构建可泛化的价值函数。基于这一认识,我们将世界模型与价值估计相结合,提出了一种新的通用机器人价值模型——世界价值模型(World Value Model,WVM)。该模型能够准确反映任务进展,从而评估数据的质量。在标准测试集上,WVM取得了最先进的价值-顺序相关性指标成绩。除了包含专家数据的标准评估体系外,我们还引入了次优价值基准测试,该测试包含800条次优轨迹,这些轨迹都带有高保真度、人类标注的帧注释。我们的评估结果显示,WVM在次优价值基准测试上仍然保持最先进的性能,证明了其在处理专家数据和次优数据方面的鲁棒性。当用于策略学习时,WVM能够提升各种策略提取方法的操作性能,无论是在模拟环境还是现实世界中都能提供可靠的指导,从而帮助从混合质量的数据中学习到有效的策略。

English Abstract

Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. Mathematically, accurate value estimation demands deep temporal understanding, requiring models to both ground the current belief using historical context and plan over future outcomes. However, most existing robotic value models are built on Vision-Language Model (VLM) backbones that are pretrained primarily on static or temporally sparse visual observations, lacking the requisite temporal modeling capabilities for value estimation. Unlike VLMs, world models naturally excel at temporal modeling and future planning, making them ideal foundations for learning generalizable value functions. Driven by this insight, we marry world models with value estimation to construct a new generalist robotic value model, World Value Model (WVM), that offers accurate task progressions to assess data quality. On standard benchmarks, WVM delivers state-of-the-art (SOTA) Value-Order Correlation (VOC) results. Complementing standard evaluation suites that contains only expert data, we further introduce Suboptimal-Value-Bench, a multi-embodiment benchmark consisting of 800 suboptimal trajectories with high-fidelity, human-labeled frame annotations. Our evaluations show that WVM maintains its SOTA performance on Suboptimal-Value-Bench, establishing its robustness in handling both expert and suboptimal data. When deployed for policy learning, WVM improves manipulation performance across various policy extraction approaches in both simulated and real-world deployment, providing robust guidance for learning from mixed-quality data.