Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了多语言大模型在保持语义不变的结构变化下的鲁棒性问题,通过构建Hindi和Malayalam两种语言的数据集并采用重排序及主动被动语态转换方法进行实验。
📝 Abstract
Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the structural sensitivity of multilingual LLMs using two linguistically grounded perturbation settings in Hindi and Malayalam: constrained constituent reordering and active-passive voice transformation. We introduce a benchmark dataset IndicReStruct, with two variants, GSM8K-Reordered and GSM8K-Voice, constructed from GSM8K while preserving semantic meaning. Across six state-of-the-art LLMs and multiple prompting strategies, we observe consistent and significant degradation in mathematical reasoning performance under structurally perturbed inputs. To further understand these failures, we perform qualitative error analysis and mechanistic interpretability experiments using residual-stream activation patching. Our analyses show that reasoning failures frequently arise from disruptions in entity-quantity alignment and that intermediate transformer layers contribute most strongly toward reasoning restoration. Overall, our findings suggest that current multilingual LLMs remain highly sensitive to surface syntactic realization and lack robust compositional invariance under structurally different but semantically equivalent inputs.
Problem

Research questions and friction points this paper is trying to address.

structural sensitivity
multilingual LLMs
semantics-preserving perturbations
robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

structural sensitivity
semantics-preserving perturbations
multilingual LLMs
entity-quantity alignment
transformer layers