物理科学深度研究:多智能体框架与全面基准测试
Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark
摘要
深度研究智能体是一种基于大型语言模型构建的系统,旨在实现自主、多步骤的科学推理功能。它们具有巨大的潜力,可以加速物理科学领域的研究进程。不过,目前对于这类系统在相关领域的实际性能仍缺乏全面且深入的评价。为了填补这一空白,我们推出了PhySciBench——一个与物理科学研究高度相关的评估工具。该工具包含200个由专家筛选出的问题,涵盖了物理学和化学两个领域,且涉及六个不同的任务类别,这些任务类别反映了现实世界中的科学工作流程。对最新技术的模型和智能体系统的评估结果显示,它们的性能有限;即便是最强大的基准模型Gemini Deep Research,其准确率也仅为33.5%。通过对失败案例的分析,我们发现有三个常见的缺陷:扩展推理过程中的脆弱性、跨步骤的知识传递不足,以及缺乏以物理原理为基础的自我验证机制。基于这些发现,我们开发了DelveAgent——一种模块化多智能体框架,它具备自适应规划机制、双粒度记忆系统以及以物理原理为基础的反思机制。在四个科学评估测试中,DelveAgent的准确率提高了7.5个百分点,而推理成本则降低了约三分之一。这些结果表明,PhySciBench作为评估物理科学领域人工智能系统的关键工具具有重要意义,同时也说明专门的架构设计能够有效提升自主科学研究的可靠性。
English Abstract
Deep research agents are Large Language Model (LLM)-based systems designed for autonomous, multi-step scientific reasoning, and they hold immense potential for accelerating research in the physical sciences. However, comprehensive and in-depth evaluations of their capabilities within this domain remain lacking. To address this gap, we introduce PhySciBench, a benchmark highly relevant to physical science research, comprising 200 expert-curated questions, balanced between physics and chemistry, across six task categories that reflect real-world scientific workflows. Evaluations of state-of-the-art models and agent systems on PhySciBench reveal limited performance; even the strongest baseline, Gemini Deep Research, achieves an accuracy of only 33.5%. Analysis of failure cases identifies three recurrent deficiencies: fragility in extended reasoning chains, limited knowledge transfer across steps, and a lack of physics-grounded self-verification. Motivated by these findings, we develop DelveAgent, a modular multi-agent framework equipped with an adaptive planning loop, dual-granularity memory, and a hierarchical physics-grounded reflection mechanism. Across four scientific benchmarks, DelveAgent improves accuracy by up to 7.5 percentage points while reducing inference costs to approximately one-third of the strongest baseline. These results establish the significance of PhySciBench as a critical benchmark for evaluating AI systems in the physical sciences and demonstrate that architectural specialization can effectively enhance the reliability of autonomous scientific research.