‹ 返回 2026-06-25

文本到图像模型是归纳式“火鸡”吗?一种用于因果推理的反事实基准测试

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

▲ 5 💬 1 2026-06-25

Jiayi Lei, Yuandong Pu, Xingyu Han, Rongpeng Zhu, Jing Xu, Jinyao Wang, Zijian Zhou, Bin Fu, Yuewen Cao, Yihao Liu, Yongsheng Li

摘要

文本到图像生成模型在根据自然语言提示生成视觉上逼真的图像方面取得了显著进展。不过,其成功究竟源于对因果关系的真正理解,还是因为对视觉与文本之间关系的复杂匹配能力,目前仍不清楚。受Russell的归纳主义理论启发,我们提出了“反事实世界”这一评估标准,用于检验文本到图像模型是否能够在与现实情况相悖的规则下生成图像。该标准将每种场景分为三个层次:基于常规世界知识的真实场景生成、带有明确视觉指令的反事实场景生成,以及需要从改变后的规则中进行因果推理的反事实场景生成。我们使用基于视觉语言模型的评估工具CF-Eval来评估开源和闭源文本到图像模型。此外,我们还引入了两个指标:优先性抵抗率,用于衡量模型克服现实中的固有先验的能力;推理保持率,用于评估模型能否在没有明确视觉提示的情况下继续进行反事实场景生成。实验表明,所有模型在从真实场景到反事实场景转换时都表现出明显的性能下降。进一步分析表明,这种失败是因为当前的文本到图像模型将世界知识和视觉特征视为紧密关联的模式。因此,它们在训练数据中频繁出现的视觉特征使得它们在面对反事实场景时只能依赖熟悉的常识性先验。

English Abstract

Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their success reflects genuine causal understanding or sophisticated pattern matching over visual-textual correlations. Inspired by Russell's inductivist turkey, we introduce Counterfactual-World (CF-World), a counterfactual benchmark designed to investigate whether text-to-image models can generate images under rules that systematically contradict real-world priors. CF-World organizes each scenario into three progressive levels: factual generation under ordinary world knowledge, explicit counterfactual generation with direct visual instructions, and implicit counterfactual generation requiring causal deduction from altered rules. We evaluate both open-source and closed-source T2I models using a Vision Language Model (VLM)-based evaluator (CF-Eval). Furthermore, we introduce two metrics: Prior Resistance Rate (PRR), which measures a model's ability to overcome entrenched real-world priors, and Reasoning Retention Rate (RRR), which assesses whether models can maintain reasoning-dependent counterfactual generation without explicit visual cues. Experiments show that all models exhibit sharp degradation from factual to counterfactual settings. Further analyses suggest that these failures arise because current T2I models encode world knowledge and visual appearances as tightly coupled patterns. Consequently, their heavy reliance on frequent visual co-occurrences within the training data forces them to default to familiar commonsense priors when tasked with rendering counterfactual worlds.