挑战之旅:重新评估非熟悉环境中的代理人能力
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
摘要
随着代理系统不断发展并在现实场景中得到广泛应用,对这些系统的性能进行准确评估的需求日益增加。然而,目前的评估标准通常基于那些任务相对简单的常用应用,只关注有限的几种能力,而忽略了其他重要方面。这导致现代代理系统在性能上已经达到了极限,无法揭示其局限性。为此,我们提出了GauntletBench这一基于网络的评估工具,用于评估代理在具有挑战性的场景中的泛化能力。该工具重点关注三种尚未被充分探索的能力:时间感知、图形理解以及3D推理能力,同时涵盖五种较少被使用的专业应用:视频编辑、工作流构建、3D建模、飞行分析以及电路设计。每个应用都包含20个以视觉处理为核心的任务,总计100个任务。我们的评估工具包含了一个模块化框架,包括适用于开源和闭源代理框架的环境、一个可控的网页应用、结构合理的任务集,以及具有多种指标的自动化评估引擎。与普遍预期相反,我们的实验结果表明,当前的代理系统距离达到人类水平还相差甚远。即使是最先进的代理系统,在GauntletBench上的成功率也仅为19.1%,这凸显了这些被忽视的能力及泛化能力的局限性。相比之下,非专业的人类标注者在这些具有挑战性但可行的任务上能够取得超过80%的成功率,这反映出当前代理能力与复杂现实场景所需的能力之间存在的巨大差距。
English Abstract
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 20 vision-intensive tasks (100 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 19.1% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.