‹ 返回 2026-06-21

LegalHalluLens:基于类型化幻觉检测与校准的多智能体辩论技术,用于实现可信的法律人工智能

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

▲ 3 💬 2 2026-06-21

Lalit Yadav, Akshaj Gurugubelli

摘要

在法律工作流程中使用的AI系统出现错误的概率约为52%。不过,这一平均值掩盖了错误发生的具体位置和方向,使得合规人员无法获得有效的信息来确保系统的可靠使用。我们提出了LegalHalluLens这一审计框架,该框架包含三个组成部分:针对四种与法律相关的索赔类型(数字、时间、义务/权利、事实)的类型化错误特征描述,这些特征基于CUAD数据(Hendrycks等人,2021年);一个风险方向指数,能够将遗漏与虚构之间的偏差转化为可比较的单一指标;以及一个根据各种参数和方向进行校准的辩论处理流程。通过对510份合同中的249,252条条款进行分析,我们发现义务/数字与时间索赔之间存在约38-40%的差异,而汇总报告则忽略了这一差异。我们还发现,两个错误率均为52%的系统可能具有相反的风险方向指数。该辩论处理流程可将虚假检测减少45%,且每个类别的改进情况与商业API相比,其核心参数仅为4B个活跃参数。类型化特征和风险方向指数能够揭示那些汇总指标所隐藏的故障模式。此外,这些诊断结果还可以作为多智能体辩论处理流程的校准输入,其中针对特定故障模式的怀疑性挑战和不对称机制能比一般化的辩论方式更有效。该框架有助于实现对方向的感知式采购、责任分配以及法律AI系统的设计,使其能够在实际环境中得到应用。

English Abstract

AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment. We present LegalHalluLens, an auditing framework with three components: typed hallucination profiles across four legally-motivated claim categories (numeric, temporal, obligation/entitlement, factual) over CUAD (Hendrycks et al., 2021); a Risk Direction Index (RDI) that reduces omission-versus-invention bias to a single deployment-comparable scalar; and a typed debate pipeline calibrated to both magnitudes and directions. Across 510 contracts and 249,252 clause-level instances we measure a within-model gap of approximately 38-40 pp between obligation/numeric and temporal claims that aggregate reporting hides, and show that two systems with matched 52% rates can carry opposite RDIs. The debate pipeline reduces fabricated detections by 45% with per-category gains tracking the diagnosis, matching commercial APIs with a substantially smaller backbone (4B active parameters). Typed profiles and RDI surface failure modes that aggregate metrics hide; we further show these diagnostics serve as calibration inputs for multi-agent debate pipelines, where Skeptic challenges and asymmetric gates targeted at measured failure modes outperform generically-tuned debate. The framework supports direction-aware procurement, accountability, and agent design for legal AI deployed in the wild.