‹ 返回 2026-06-30

利用 Google Paper Assistant 工具实现科学论文的自动化评审

Towards Automating Scientific Review with Google's Paper Assistant Tool

▲ 3 2026-06-30

Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, Vincent Cohen-Addad

摘要

人工智能正在推动科学发现领域的革命性变革,使得从假设提出到数学定理验证等过程都得以加速。然而,这种快速发展也带来了系统性挑战:传统的人工同行评审方式无法应对人工智能辅助下的科学研究所带来的巨大压力。为了解决这一矛盾,我们必须利用人工智能来加速验证和评审过程本身。为了围绕这一转变展开讨论,我们提出了一种分类体系,将人工智能与人类的合作分为四个不同的阶段,并探讨了每个阶段所涉及的各种权衡问题。 作为迈向这一未来的一步,我们推出了Paper Assistant Tool(PAT)这一智能AI工具,它专门用于深度科学领域的评审与验证工作。PAT能够处理完整的科学论文,并提供全面的评估结果,包括对理论结果的检验、实验的验证、改进建议以及潜在缺陷的识别。通过运用推理扩展技术,PAT能够发现比单一模型更深层的问题,其在SPOT基准测试中的数学错误识别准确率提升了34%。在STOC和ICML这两个重要的计算机科学会议上,PAT被用作作者提交前使用的工具,结果显示它能够识别出关键错误并提出有效的改进建议。通过及早发现错误,PAT减轻了审稿人的认知负担,同时又能保留他们对评审过程的控制权。

English Abstract

Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements, and identifying potential flaws. By utilizing inference scaling techniques, PAT is able to identify deeper issues than a single model call alone, achieving a 34% improvement over zero-shot recall on mathematical errors in the SPOT benchmark. Pilot deployments of PAT as a pre-submission tool for authors at two major Computer Science conferences -- STOC and ICML -- demonstrate its ability to identify critical errors and suggest substantive improvements to research papers. By catching errors early, PAT eases the cognitive burden placed on referees, while preserving their control over the outcomes of the review process.