通过研究平台将人工智能科学家的研究整合、综合与验证工作外包出去
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
摘要
人工智能系统能够越来越自动完成科学研究流程,但那些将先前的证据、提出的想法、实验以及最终结论联系起来的推理过程,往往仍然隐藏在模型推理之中。在这里,我们介绍了Xcientist这一研究工具——它将研究的整合与实验验证过程转化为可检查、受契约约束的过程。Xcientist将文献证据、想法描述、实施计划、测试记录以及修复痕迹都作为持久的研究成果保存下来,这样所生成的机制就能保持其依据,被执行、测试并修正,而不会失去其证据基础。我们认为,自动化研究中的一种失败模式就是“主张的漂移”,即那些可以运行的成果不再能支持最初所声称的机制。在无需训练的记忆系统、图结构交通预测以及多尺度物理信息神经网络中,Xcientist能够保留从问题提出到机制设计、验证以及有限修改的完整过程。这些结果表明,人工智能科学家不应仅根据他们的最终成果来评价,还应看他们的整合与验证过程是否仍然具有可追溯性、可检查性以及科学上的可问责性。
English Abstract
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, idea states, implementation plans, ablation records and repair traces as persistent research artifacts, so that generated mechanisms can be grounded, executed, tested and revised without losing their evidential basis. We identify claim drift as a failure mode of automated research, where runnable artifacts no longer support the mechanism originally claimed. Across training-free memory systems, graph-structured traffic forecasting and multi-scale physics-informed neural networks, Xcientist preserves traceable trajectories from problem formulation to mechanism design, validation and bounded revision. These results suggest that AI scientists should be evaluated not only by their final artifacts, but by whether their synthesis and validation processes remain attributable, inspectable and scientifically accountable.