Answer Compliance and Cross-Lingual Chain-of-Thought Gaps
Yandan Zheng ⋅ Anh Tuan Luu
Abstract
**Motivation.** Multilingual chain-of-thought (CoT) evaluation [1] must separate panel-level ranking changes from model-level shifts in cross-lingual accuracy gaps. Kendall's $\tau$ correlates models' English and target-language accuracy rankings. Under a strict output protocol, $\Delta\tau=\tau_{\mathrm{CoT}}-\tau_{\mathrm{direct}}$ spans $0.256$–$0.400$ across the three commonsense settings. Yet leave-two-model-out panels on 100 Cross-lingual Choice of Plausible Alternatives items (XCOPA100) span $-0.307$ to $0.643$. This sensitivity motivates a paired audit of each model and setting. **Method.** We evaluate ten open-weight instruction-tuned models from Gemma, DeepSeek, Ministral, Llama, Aya, Command R, Qwen, Yi, and Phi, spanning 4 to 15.7 billion parameters. The six settings are English-Chinese XCOPA100 and XCOPA500 for two-choice causal commonsense, English-Chinese XStoryCloze for story completion, and Multilingual Grade School Math (MGSM) for free-response arithmetic with English paired with Chinese, Japanese, and French [2–4]. XCOPA100 is nested in XCOPA500. The three MGSM pairs share each model's English outputs. **PairLift** defines the final-answer gap shift for model $M$, English (EN), and target language X as $\delta_M=[\mathrm{acc}_M^{\mathrm{EN,CoT}}-\mathrm{acc}_M^{\mathrm{X,CoT}}]-[\mathrm{acc}_M^{\mathrm{EN,direct}}-\mathrm{acc}_M^{\mathrm{X,direct}}]$. $\delta_M>0$ means that CoT increases the English-minus-X accuracy gap. Before generation, we fixed the strict protocol. Its evaluation-language prompts require a final line of exactly *Answer: A*, *Answer: B*, or *Answer: <number>*. CoT also requests reasoning in that language. Both modes use one template, a 768-token limit, temperature $0.7$, seed 42, and four additional seeds. The strict extractor reads only the final *Answer:* field. Missing or invalid fields count as incorrect in primary end-to-end accuracy. We report schema compliance, extractability, and truncation separately. A secondary diagnostic recomputes $\delta_M$ on items with valid answers in all four arms. Because validity is determined after generation, this diagnostic estimates performance on a post-treatment, cell-specific complete-case subset. Both analyses use a two-way seed-by-item bootstrap. A flag requires $|\delta_M|\geq0.10$, a 95\% confidence interval (CI) excluding zero, and a sign-flip test that remains significant after Benjamini–Hochberg (BH) false-discovery-rate correction at $0.05$ across all 60 cells. **Results.** The primary grid has 28 positive and 32 negative gap shifts. The rule selects 44 cells, with 20 positive and 24 negative effects. Median $|\delta_M|$ is $0.455$, with a range of $0.121$ to $0.818$. Across 522,200 outputs from both modes and five seeds, strict extractability falls from $82.6\%$ under direct answering to $70.1\%$ under CoT in English, while it rises from $80.9\%$ to $86.6\%$ in target languages. Only $0.08\%$ of English CoT outputs reach the token limit. Cell-level extractability interactions are bidirectional, with 23 positive and 37 negative shifts, so aggregate rates conceal individual reversals. The 44 selected cells retain a median $27.2\%$ of items in the four-arm-valid subset, with a range of $2.6\%$ to $83.8\%$. Their median secondary $|\delta_M|$ is $0.057$. The secondary rule retains 11 primary-direction effects across seven models, split into six positive and five negative effects. Ten occur on MGSM, with four in Chinese, four in Japanese, and two in French. The remaining effect is DeepSeek-V2-Lite on XCOPA500. On XCOPA100, Qwen2.5-14B accuracy rises from $0.444$ to $0.964$ in English and from $0.930$ to $0.940$ in Chinese, giving primary $\delta_M=0.510$. English direct accuracy and extractability are both $0.444$. Only 42 to 46 items per seed are valid in all four arms, where secondary $\delta_M=-0.023$ with 95\% CI $[-0.073,0]$. Two greedy XCOPA100 spot checks give $\delta_M=0.450$ for Qwen2.5-14B and $-0.810$ for Phi-4. **Implication.** Large end-to-end shifts coincide with answer-schema compliance changes and contract sharply in the four-arm-valid diagnostic. The 11 retained effects split six positive and five negative, giving no common English-advantage direction in this panel. Multilingual CoT reports that omit arm-by-language extraction diagnostics can mix final-answer accuracy with compliance. The scope covers models up to 15.7 billion parameters and one prompt template. **References.** [1] Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. *NeurIPS*, 2022. [2] Ponti et al. XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning. *EMNLP*, 2020. [3] Lin et al. Few-Shot Learning with Multilingual Generative Language Models. *EMNLP*, 2022. [4] Shi et al. Language Models Are Multilingual Chain-of-Thought Reasoners. *ICLR*, 2023.
Chat is not available.
Successful Page Load