‹ 返回 2026-06-24

EnterpriseClawBench:基于真实工作场景的代理对比分析

EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions

▲ 56 💬 1 2026-06-24

Jincheng Zhong, Weizhi Wang, Che Jiang, Kai Tian, Zhenzhao Yuan, Junlin Yang, Dianqiao Lei, Kaiyan Zhang

摘要

企业代理越来越多地在工作环境中运行:它们会读取各种不同的文件、调用各种工具,并生成各种业务成果。我们推出了EnterpriseClawBench这一企业代理评估标准,该标准基于真实的代理交互场景而构建。从大量工作场景数据中,EnterpriseClawBench能够生成852个可复现的任务,每个任务都伴随着相应的参数、重新编写的提示词、角色类别、技能子类型、固定规则以及语义标准。由于这些场景中包含企业内部的特定内容,因此我们不会公开这些评估数据;我们的贡献在于其构建和评估流程。在EnterpriseClawBench中,最佳配置仅为0.663分(使用GPT-5.5模型时的得分)。这些结果表明,对企业代理的评估应该考虑多种因素,如模型组合、成果质量、成本、运行时间以及技能传递效果,而不是将性能简化为一个单一分数。代码链接:https://github.com/FrontisAI/EnterpriseClawBench

English Abstract

Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions. Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics. Because the sessions contain internal enterprise content, we do not release the benchmark data; instead, our reusable contribution is the construction and evaluation protocol. On EnterpriseClawBench, the best configuration reaches only 0.663 (Codex with GPT-5.5). These results show that enterprise agent evaluation must report harness--model combinations, artifact delivery, visual quality, cost, runtime, and skill-transfer behavior, rather than collapsing performance into a single score. Code: https://github.com/FrontisAI/EnterpriseClawBench