‹ 返回 2026-06-26

OpenBioRQ:尚未解决的生物医学研究问题与药物相关

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

▲ 1 💬 1 2026-06-26

Minbyul Jeong

摘要

有效的引用看起来就像证据一样——但链接能够被正确解析,并不意味着被引用的论文确实支持该主张。我发现当前的人工智能模型很少伪造引用(超过99%的链接都是有效的),不过大约有15.9%的链接指向了错误的论文。现有的评估标准没有考虑到这种错误情况:当某个问题有固定的答案时,模型可以依据那个答案来识别正确的来源,而无需独立验证该来源是否真的支持该主张。我引入了\openbiorq{}这个基准测试工具,它涵盖了12个领域中的12,553个尚未解决的生物医学研究问题,将开放性问题视为一种“忠实性与回避性”的测试指标。据我所知,这是第一个将人工智能模型与那些没有固定答案的未解决问题相结合的生物医学基准测试工具。这种开放性是通过真实的后续证据来验证的,而不是通过模型的参数知识来确定的。难度则是基于实证数据来确定的:我以那些三个开放权重参考模型都无法回答的问题作为衡量标准,而不是依靠主观上的难度标签。在最难的那些问题上,来自同一系列的模型只能解决约17%的问题,而三个独立的先进模型则能解决29%-60%的问题。因此,这个基准测试非常具有挑战性,不会让人感到饱和(即使是最优秀的模型也仍有约33%-40%的问题无法解决),而且能够区分不同能力的模型。除了难度之外,我还发现人工智能模型在最困难的问题上会陷入崩溃状态,即不再使用它们的工具。对于最容易出现这种情况的模型来说,完全阻止工具的使用几乎不会改变其得分——也就是说,工具在最需要的时候就不再起作用了。每个问题的检查清单使得不同评审者之间的评分一致性从Spearman的0.35提升到了0.82。

English Abstract

A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over 99% resolve), yet roughly 15.9% link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source supports the claim. I introduce \openbiorq{}, a retrieval-grounded agentic benchmark of 12{,}553 unsolved biomedical research questions across 12 domains that treats open questions as a faithfulness-and-abstention probe. To my knowledge, this is the first biomedical benchmark to combine an agentic setting -- where the model must issue multiple tool calls -- with unsolved questions that have no answer key. Openness is verified against real follow-up evidence rather than a model's parametric knowledge. Difficulty is empirical: I anchor it on questions that three open-weight reference models fail to answer, rather than on subjective hardness labels. On this hardest subset, held-out models from the same lineage as the difficulty anchors solve only ~17%, while three independent frontier agents (Gemini-3-Pro, Opus-4.7, GPT-5.5) span a wide 29-60% range. The benchmark is thus hard, non-saturating (the best agent still leaves ~33-40\% unsolved), and discriminating across capability tiers. Beyond difficulty, I observe agentic collapse on the hardest questions, where agents stop using their tools. For the most collapse-prone model, blocking tool access entirely barely changes its score -- so tools stop paying off exactly where they are needed most. A frozen per-question checklist raises inter-judge agreement from Spearman 0.35 to 0.82.