‹ 返回 2026-06-21

超越静态排行榜:用于评估LLM智能体的预测有效性

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

▲ 23 💬 1 2026-06-21

Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin, Yusheng Li, Tianjun Feng, Chun-Yi Tsai, Yihan Sun, Wei Alexander Xin, Akshat Bhandari, Tanisha Rathod, Aaron Fan, Sanskruti Vijay Shejwal, Tomas Pasiecznik, Sagar Chethan Kumar, Tanmay Agarwal, Rohith Kanathur, Sam Colman, Amaan Sheikh, Dev Bahl, Ann Li, Krish Veera, Alimurtaza Mustafa Merchant, Shambhawi Baswaraj Bhure, Sajal Kumar Goyla, Chengrui Li, Kirthana Natarajan, Rui Li, Thomas Ajai, Rujing Li, Vivek G. Iyer, Sanjaii Vijayakumar, Yitong Bai, Ayal Yakobe, Darief Maes, Yassine Jebbouri, Tianyang Xu, Thai Quoc On, Vera Mazeeva, Winston Li, Yuval Shemla, Yeshitha Bhuvanesh, Rushin Bhatt, Siddharth Chethan Gowda, Alisha Vinod, Caroline Cahill, Shriya Aishani Rachakonda, Yunfeng Chen, Aryaman Agrawal, Aman Upganlawar, Mao Le Jonathan Ang, Yubin Sally Go, Madhav Rajkondawar, Yang-Jung Chen, Trisha Maturi, Ananya Kapoor, Andrew Li, Shrey Arora, Mana Abbaszadeh, Shen Li, Charles Xu, Byeolah Kwon

摘要

代理评估基准正在快速发展,但没有任何一个基准能够涵盖部署过程中涉及的四到五个关键维度。本文汇总了迄今为止基于MCP的工业代理评估基准中规模最大的深度研究:共有14项并行研究,涉及新的资产类别(包括多模态视觉功能)、不同的调度方式、检索策略、推理模式、基础设施优化以及评估方法等方面。将这些研究与之前的七个代理评估基准结合起来后,我们认为,基于汇总分数的排名方式无法准确反映实际部署中的代理评估情况。由汇总分数得出的排名并不适用于非分布环境下的场景;最近的公开与隐藏竞争分析结果也直接证明了这种排名的不稳定性。我们提出了基于预测有效性作为排名依据的体系,即样本内和样本外排名之间的相关性,而不是仅依赖样本内的平均值。此外,我们还提出了一种包含12个等级的评估体系,该体系能够揭示与部署相关的维度——HELM及其后续版本所存在的缺陷。这一排名方式通过三个可验证的非分布环境标准来实施,这些标准都有明确的阈值;现有证据部分支持这一方法,但证据数量不足,无法得到确认。最后,我们提出了一种预先注册的试点设计,以及对未来新一代代理评估基准的期望。

English Abstract

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.