ChLogic:评估中文表达中逻辑推理的稳健性
ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions
摘要
大型语言模型在标准化的逻辑推理测试中的表现越来越出色,不过这种能力是否能在超越英语的情况下依然保持,目前尚不清楚。我们提出了ChLogic这一英语-中文对齐的测试基准,该基准旨在评估模型在相同的潜在逻辑结构被用英语和不同的中文表达方式呈现时,其逻辑推理能力的稳定性如何。该测试基准基于正式的逻辑模板构建,包含三个数据集:(i)通用对齐集,由来自九个逻辑模板家族的60个通用命题组成;(ii)困难对齐集,由40个困难问题组成;(iii)仅中文集,涵盖15种与语言相关的现象类型。每个对齐项都包含一个英语参考表达以及五个对应的中文表达方式。在Qwen3、Ministral和GLM模型上的实验表明,英语-中文之间的性能差距仍然存在。从标准中文翻译成英文通常能提升在通用对齐集上的性能,但在困难对齐集上则效果不一;其中,Qwen3-32B和GLM-5.1在翻译后的表现反而更差。这些结果表明,中文的表达方式、翻译过程中的误差以及模型特有的行为共同影响了多语言逻辑推理的能力。总体而言,ChLogic为评估多语言推理能力的鲁棒性提供了一个有效的测试工具。
English Abstract
Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogic, an English--Chinese aligned benchmark that tests whether models preserve logical reasoning performance when the same latent logical structure is expressed in English and diverse Chinese surface realizations. Built from formal logical templates, the benchmark contains three data sets: (i) the General aligned set, derived from 60 General Propositions across nine template families; (ii) the Difficult aligned set, derived from 40 Difficult Problems; and (iii) the Chinese-only set, covering 15 language-specific phenomenon types. Each aligned item pairs one English reference expression with five Chinese realizations. Experiments on Qwen3, Ministral, and GLM models reveal a persistent English--Chinese performance gap. Back-translation from standard Chinese into English often improves performance on the General aligned set, but produces mixed effects on the Difficult aligned set, where Qwen3-32B and GLM-5.1 perform worse after translation. These results indicate that Chinese surface realization, translation artifacts, and model-specific behavior jointly affect multilingual logical reasoning. Overall, ChLogic provides a useful stress test for the robustness of multilingual reasoning.