Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过眼动追踪数据比较了不同粤语适应程度的语言模型,探讨了粤语特定训练是否能更好预测粤语阅读。结果表明更广泛的粤语训练提高了预测准确性。
📝 Abstract
Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is still open for Cantonese, where recent NLP evaluations reported mixed benefits from Cantonese-specific training relative to Mandarin-oriented or general-purpose models. Using naturalistic Cantonese eye-tracking data, we compare two within-family adaptation contrasts: CKIP GPT-2 Tiny versus its lightly Cantonese-adapted JED351 derivative, and Qwen2.5-7B versus CantoneseLLM-7B, which underwent substantially more extensive Cantonese continued pretraining and instruction tuning. From each model, we derive lexical surprisal, POS surprisal, entropy before the target, and entropy reduction. Lexical surprisal and the joint four-metric model consistently favor CantoneseLLM-7B, followed by Qwen2.5-7B, CKIP, and JED351, whereas entropy reduction favors CKIP. These results suggest that more extensive Cantonese-specific training can be associated with stronger predictive fit, while model rankings also depend on the information-theoretic measure being evaluated.
Problem

Research questions and friction points this paper is trying to address.

Cantonese
language models
reading prediction
psycholinguistic alignment
information-theoretic measures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cantonese-specific training
information-theoretic measures
predictive fit
language-variety adaptation
💼 Related Jobs
No related jobs found.
Z
Ziqi Zhang
Department of Language Science and Technology, The Hong Kong Polytechnic University
Emmanuele Chersoni
Emmanuele Chersoni
Hong Kong Polytechnic University
Computational Linguistics
M
Mohammad Momenian
Department of Linguistics, The University of Hong Kong