From Phonemes to Meaning: Evaluating Large Language Models on Tamil
Low-resource, morphologically rich languages like Tamil lack native linguistic evaluation benchmarks, hindering reliable assessment of large language models (LLMs). Method: We introduce ILAKKANAM—the first linguistically grounded, culturally authentic Tamil evaluation benchmark—constructed from real Sri Lankan K–12 examination items. It covers five linguistic dimensions (morphology, syntax, semantics, pragmatics, and factual knowledge) via 820 expert-annotated, native-language questions organized within a grade-based difficulty framework to avoid cultural and linguistic distortions from English translation. Contribution/Results: Our systematic evaluation of leading closed- and open-weight LLMs reveals that Gemini 2.5 achieves highest accuracy; open-source models consistently underperform. Accuracy declines markedly with increasing grade level (i.e., rising linguistic complexity), and improvements in linguistic competence show no strong correlation with language identification capability. ILAKKANAM establishes a reproducible, culturally grounded paradigm for evaluating LLMs in low-resource languages.