‹ 返回 2026-06-26

EBench:通用移动操作策略的要素诊断

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

▲ 13 💬 1 2026-06-26

Ning Gao, Jinliang Zheng, Xing Gao, Haoxiang Ma, Hanqing Wang, Yukai Wang, Jiantong Chen, Zanxin Chen, Shujie Zhang, Mingda Jia, Xuekun Jiang, Zihou Zhu, Xinyu Li, Shuai Wang, Hao Li, Wenzhe Cai, Yuqiang Yang, Xudong Xu, Zhaoyang Lyu, Yao Mu, Tai Wang, Jiangmiao Pang, Jia Zeng, Weinan Zhang, Chunhua Shen

摘要

我们提出了EBench这一仿真基准测试工具,它能够超越单一的成功率指标,来评估各种通用的移动操作策略。EBench包含26种不同且具有挑战性的操作任务,这些任务依据5个能力维度和4个泛化维度进行标注。我们对当前最先进的通用操作模型进行了评估,包括π_0、π_{0.5}、XVLA和InternVLA-A1等模型。结果显示,那些成功率接近理想的模型在能力表现上存在显著差异:π_{0.5}的测试成功率最高,且训练与测试阶段的性能保持情况也最好;而InternVLA-A1在移动操作方面表现优异,但在需要灵活操作的任务中则表现不佳;XVLA则在一系列独立的技能方面具有优势。除了能力分析之外,EBench还从4个不同的角度来分析模型的泛化能力,从而揭示不同分布变化因素对模型表现的影响。这些结果有助于我们了解各个模型的优缺点。我们希望这一基准测试工具能够提供丰富的诊断信息,从而帮助人们不断改进通用操作模型。

English Abstract

We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar. EBench comprises 26 diverse and challenging manipulation tasks annotated along 5 capability dimensions and 4 generalization dimensions. We evaluate state-of-the-art generalist manipulation models including π_0, π_{0.5}, XVLA, and InternVLA-A1, and reveal that models with near success rates exhibit strikingly different capability profiles: π_{0.5} achieves the highest test success rate and the best train--test retention, whereas InternVLA-A1 dominates mobile manipulation but collapses on dexterous tasks, and XVLA exhibits strengths on a disjoint set of atomic skills compared to other policies. Beyond capability profiling, EBench analyzes the generalization ability from 4 representative perspectives, identifying the impact of different distribution shift factors. The results reveal strengths and weaknesses of models behind an overall score. We hope this benchmark offers a broad set of diagnostic signals to guide iteration on generalist manipulation models.