EBench:通用移动操作策略的要素诊断
EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
摘要
我们提出了EBench这一仿真基准测试工具,它能够超越单一的成功率指标,来评估各种通用的移动操作策略。EBench包含26种不同且具有挑战性的操作任务,这些任务依据5个能力维度和4个泛化维度进行标注。我们对当前最先进的通用操作模型进行了评估,包括π_0、π_{0.5}、XVLA和InternVLA-A1等模型。结果显示,那些成功率接近理想的模型在能力表现上存在显著差异:π_{0.5}的测试成功率最高,且训练与测试阶段的性能保持情况也最好;而InternVLA-A1在移动操作方面表现优异,但在需要灵活操作的任务中则表现不佳;XVLA则在一系列独立的技能方面具有优势。除了能力分析之外,EBench还从4个不同的角度来分析模型的泛化能力,从而揭示不同分布变化因素对模型表现的影响。这些结果有助于我们了解各个模型的优缺点。我们希望这一基准测试工具能够提供丰富的诊断信息,从而帮助人们不断改进通用操作模型。
English Abstract
We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar. EBench comprises 26 diverse and challenging manipulation tasks annotated along 5 capability dimensions and 4 generalization dimensions. We evaluate state-of-the-art generalist manipulation models including π_0, π_{0.5}, XVLA, and InternVLA-A1, and reveal that models with near success rates exhibit strikingly different capability profiles: π_{0.5} achieves the highest test success rate and the best train--test retention, whereas InternVLA-A1 dominates mobile manipulation but collapses on dexterous tasks, and XVLA exhibits strengths on a disjoint set of atomic skills compared to other policies. Beyond capability profiling, EBench analyzes the generalization ability from 4 representative perspectives, identifying the impact of different distribution shift factors. The results reveal strengths and weaknesses of models behind an overall score. We hope this benchmark offers a broad set of diagnostic signals to guide iteration on generalist manipulation models.