在噪声感知下的选择性控制:模块化网络中被综合指标隐藏的治理失败
Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks
摘要
内容审核系统虽然能够在所有准确性指标上达到较高水平,但仍然有可能对某些用户造成实际损害——尤其是那些处于不同社区、且与其他用户缺乏联系的用户。我们在基于代理的模型中验证了这一点:在一个由N=240个学习代理组成的社区结构中,每个代理都会发布无害、有益或危险的内容;而监管机构则负责移除或惩罚那些被误判为有害的内容。随着噪声的增加,整体有效性几乎没有变化(单向方差分析,p=0.96);从总体来看,似乎没有任何问题。然而,真正的损害却集中在这些“桥梁用户”身上——他们的有益内容被错误地屏蔽,而危险内容则被错误地保留下来。在误报频繁的情况下,这种治理成本会远远超过执行成本。总体准确率掩盖了谁受到了损害,而易于监控的指标就是用户的连接数(度),而度几乎可以完美地反映某个用户的“桥梁特性”(r=0.96)。
English Abstract
A content-moderation system can score well on every standard accuracy metric and still cause real harm, if its mistakes fall on the few users who connect otherwise separate communities. We show this in an agent-based model where N=240 learning agents on a community-structured network each post harmless, productive, or dangerous content, and a regulator removes or penalizes whatever a noisy classifier flags. Overall usefulness barely moves as the noise changes (one-way ANOVA, p=0.96): by aggregate measures, nothing looks wrong. The damage instead concentrates on these bridge users, whose useful posts are wrongly suppressed and whose dangerous posts are wrongly spared. A governance loss (L_gov) that prices these two mistakes separately from the cost of enforcement more than doubles under false-positive-heavy noise. Aggregate accuracy hides who is harmed, and the cheap quantity to audit is how many connections a user has (degree), a near-perfect proxy for the betweenness that defines a bridge (r=0.96).