‹ 返回 2026-06-24

物理科学深度研究:多智能体框架与全面基准测试

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark

▲ 12 💬 1 2026-06-24

Yigeng Jiang, Tengchao Yang, Taoyong Cui, Jiaxing Wan, Yuan Wang, Weida Wang, Zhiyu Liu, Chuyi Peng, Binzhao Luo, Maoli Gao, Huaihai Huang, Yuqianer Zeng, Ziyang Zheng, Dongchen Huang, Chao Chen, Zichao Liu, Weiping Shen, Shuchen Pu, Siyu Zhou, Runmin Ma, Yusong Hu, Fei Chao, Bo Zhang, Xiawu Zheng, Zifu Wang, Lei Bai, Yunqi Cai, Shufei Zhang

摘要

深度研究智能体是一种基于大型语言模型构建的系统,旨在实现自主、多步骤的科学推理功能。它们具有巨大的潜力,可以加速物理科学领域的研究进程。不过,目前对于这类系统在相关领域的实际性能仍缺乏全面且深入的评价。为了填补这一空白,我们推出了PhySciBench——一个与物理科学研究高度相关的评估工具。该工具包含200个由专家筛选出的问题,涵盖了物理学和化学两个领域,且涉及六个不同的任务类别,这些任务类别反映了现实世界中的科学工作流程。对最新技术的模型和智能体系统的评估结果显示,它们的性能有限;即便是最强大的基准模型Gemini Deep Research,其准确率也仅为33.5%。通过对失败案例的分析,我们发现有三个常见的缺陷:扩展推理过程中的脆弱性、跨步骤的知识传递不足,以及缺乏以物理原理为基础的自我验证机制。基于这些发现,我们开发了DelveAgent——一种模块化多智能体框架,它具备自适应规划机制、双粒度记忆系统以及以物理原理为基础的反思机制。在四个科学评估测试中,DelveAgent的准确率提高了7.5个百分点,而推理成本则降低了约三分之一。这些结果表明,PhySciBench作为评估物理科学领域人工智能系统的关键工具具有重要意义,同时也说明专门的架构设计能够有效提升自主科学研究的可靠性。

English Abstract

Deep research agents are Large Language Model (LLM)-based systems designed for autonomous, multi-step scientific reasoning, and they hold immense potential for accelerating research in the physical sciences. However, comprehensive and in-depth evaluations of their capabilities within this domain remain lacking. To address this gap, we introduce PhySciBench, a benchmark highly relevant to physical science research, comprising 200 expert-curated questions, balanced between physics and chemistry, across six task categories that reflect real-world scientific workflows. Evaluations of state-of-the-art models and agent systems on PhySciBench reveal limited performance; even the strongest baseline, Gemini Deep Research, achieves an accuracy of only 33.5%. Analysis of failure cases identifies three recurrent deficiencies: fragility in extended reasoning chains, limited knowledge transfer across steps, and a lack of physics-grounded self-verification. Motivated by these findings, we develop DelveAgent, a modular multi-agent framework equipped with an adaptive planning loop, dual-granularity memory, and a hierarchical physics-grounded reflection mechanism. Across four scientific benchmarks, DelveAgent improves accuracy by up to 7.5 percentage points while reducing inference costs to approximately one-third of the strongest baseline. These results establish the significance of PhySciBench as a critical benchmark for evaluating AI systems in the physical sciences and demonstrate that architectural specialization can effectively enhance the reliability of autonomous scientific research.