Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
This study addresses the instability of large language models (LLMs) when presented with inputs that are semantically equivalent but vary in lexical or syntactic form—a vulnerability that undermines their reliability in evaluation settings. For the first time, the authors integrate linguistic principles to construct perturbation datasets based on semantic equivalence through lexical substitutions (synonym replacement) and syntactic transformations (dependency structure alterations). They systematically evaluate the robustness of 23 prominent LLMs across the MMLU, SQuAD, and AMEGA benchmarks, complemented by statistical significance testing. The findings reveal that lexical perturbations consistently and significantly degrade model performance, while syntactic perturbations yield variable effects. Notably, model scale shows no consistent correlation with robustness, suggesting that current LLMs rely excessively on surface-level patterns rather than deep semantic understanding—thereby challenging the prevailing paradigm of assessing model capabilities solely through raw benchmark scores.