超越静态排行榜:用于评估LLM智能体的预测有效性
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
摘要
代理评估基准正在快速发展,但没有任何一个基准能够涵盖部署过程中涉及的四到五个关键维度。本文汇总了迄今为止基于MCP的工业代理评估基准中规模最大的深度研究:共有14项并行研究,涉及新的资产类别(包括多模态视觉功能)、不同的调度方式、检索策略、推理模式、基础设施优化以及评估方法等方面。将这些研究与之前的七个代理评估基准结合起来后,我们认为,基于汇总分数的排名方式无法准确反映实际部署中的代理评估情况。由汇总分数得出的排名并不适用于非分布环境下的场景;最近的公开与隐藏竞争分析结果也直接证明了这种排名的不稳定性。我们提出了基于预测有效性作为排名依据的体系,即样本内和样本外排名之间的相关性,而不是仅依赖样本内的平均值。此外,我们还提出了一种包含12个等级的评估体系,该体系能够揭示与部署相关的维度——HELM及其后续版本所存在的缺陷。这一排名方式通过三个可验证的非分布环境标准来实施,这些标准都有明确的阈值;现有证据部分支持这一方法,但证据数量不足,无法得到确认。最后,我们提出了一种预先注册的试点设计,以及对未来新一代代理评估基准的期望。
English Abstract
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.