Automatic Essay Scoring and Feedback Generation in Basque Language Learning
This study addresses automated essay scoring (AES) and pedagogical feedback generation for Basque—a low-resource language. We introduce the first publicly available, expert-annotated CEFR C1-level Basque essay dataset (3,200 essays), annotated across multidimensional quality criteria including accuracy, lexical richness, and coherence, alongside expert feedback and representative error examples. Methodologically, we propose a novel feedback quality evaluation framework integrating automatic consistency assessment with expert validation, and develop interpretable, teaching-oriented AES and feedback generation models via supervised fine-tuning of RoBERTa-EusCrawl and Latxa 8B/70B. Results show that fine-tuned Latxa significantly outperforms GPT-5 and Claude Sonnet 4.5 in scoring consistency and feedback utility, while detecting a broader spectrum of linguistic errors. This work establishes a high-quality benchmark dataset, reproducible methodology, and open-source tools for low-resource NLP research.