Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of machine translation for low-resource languages like Tamil, which suffer from scarce parallel corpora, significant domain divergence, and morphological complexity. We systematically evaluate multilingual Transformer models—including NLLB and mBART—as well as the TamilLaMA large language model across multiple English–Tamil bilingual datasets. Innovatively integrating few-shot in-context prompting with attention alignment visualization, our approach is the first to apply these techniques to English–Tamil bidirectional translation. Performance is quantified using BLEU and chrF metrics. Results demonstrate that translation quality is highly sensitive to data quality and domain alignment; few-shot prompting yields structurally coherent Tamil translations; and attention visualization substantially enhances model interpretability.
📝 Abstract
The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and morphological complexity. This research presents the comprehensive evaluation of the performance of several multilingual translation models on English-Tamil and Tamil-English translations across multiple datasets: NTREX, EnTamV2, WikiMatrix and PMIndia. This study evaluates supervised NMT systems, NLLB and mBART, using both the BLEU and chrF metric, and examines how these systems perform on data of different quality levels and domains. This performs an attention-based analysis to increase model interpretability by visualising the alignments of tokens in an English source text and their Tamil translations and vice-versa to provide insight into how they make translations. This study also demonstrates that using in-context prompting can provide an excellent way to perform a few-shot translation of English to Tamil and Tamil-English using a Tamil capable TamilLaMA model, and compare this to supervised approaches qualitatively. These findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.
Problem

Research questions and friction points this paper is trying to address.

low-resource languages
machine translation
Tamil
morphological complexity
domain variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

low-resource machine translation
attention-based interpretability
in-context prompting
few-shot translation
Tamil language processing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sriharshaa S
School of Computing, SASTRA Deemed to be University
S
Sangeetha Sivanesan
Department of Computer Applications, National Institute of Technology