Institution profile

HiTZ Center

Academic institution
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Automatic Essay Scoring and Feedback Generation in Basque Language Learning

Dec 09, 2025

This study addresses automated essay scoring (AES) and pedagogical feedback generation for Basque—a low-resource language. We introduce the first publicly available, expert-annotated CEFR C1-level Basque essay dataset (3,200 essays), annotated across multidimensional quality criteria including accuracy, lexical richness, and coherence, alongside expert feedback and representative error examples. Methodologically, we propose a novel feedback quality evaluation framework integrating automatic consistency assessment with expert validation, and develop interpretable, teaching-oriented AES and feedback generation models via supervised fine-tuning of RoBERTa-EusCrawl and Latxa 8B/70B. Results show that fine-tuned Latxa significantly outperforms GPT-5 and Claude Sonnet 4.5 in scoring consistency and feedback utility, while detecting a broader spectrum of linguistic errors. This work establishes a high-quality benchmark dataset, reproducible methodology, and open-source tools for low-resource NLP research.

0 citationsRead paper

BERnaT: Basque Encoders for Representing Natural Textual Diversity

Dec 03, 2025

Language models often exhibit representational bias and reduced robustness due to overreliance on high-quality, standardized corpora, neglecting dialectal, historical, and informal linguistic variants. To address this, we propose *full-spectrum language modeling*, instantiated with Basque as a case study. We construct a multilingual, multi-source pretraining dataset integrating standardized texts, social media content, and historical corpora, and train three Transformer-based encoder configurations. Our key contribution is a novel hierarchical evaluation framework that—uniquely—partitions NLU tasks into standard and diverse subsets, enabling systematic quantification of models’ capacity to capture linguistic variation. Experiments demonstrate that models trained on mixed-domain data achieve significant performance gains on diverse linguistic contexts while maintaining competitive accuracy on standard benchmarks. This validates both the efficacy and necessity of explicitly modeling linguistic diversity in language model pretraining.

0 citationsRead paper

Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding

Jul 21, 2025

This paper systematically evaluates the genuine metaphor comprehension capabilities of large language models (LLMs), investigating whether their performance on metaphor interpretation tasks stems from deep semantic understanding or reliance on superficial cues—such as lexical overlap, sentence length, or syntactic patterns—in natural language inference (NLI) and question answering (QA). Method: Leveraging multiple public metaphor datasets, the authors design diverse prompting strategies and controlled ablation experiments to quantify correlations between model performance and surface-level linguistic features. Contribution/Results: Results demonstrate that LLMs’ metaphor interpretation is predominantly driven by surface features, in-context learning, and pretraining-induced linguistic priors—not by robust semantic reasoning. Consequently, standard evaluation protocols yield overly optimistic estimates of metaphor understanding. The paper proposes a more rigorous, bias-mitigated evaluation paradigm for metaphor comprehension and open-sources all data, code, and analysis frameworks to enable reproducible research and establish a new benchmark for trustworthy metaphor modeling.

0 citationsRead paper
Recent publications

Latest Papers

Automatic Essay Scoring and Feedback Generation in Basque Language Learning

Dec 09, 2025

This study addresses automated essay scoring (AES) and pedagogical feedback generation for Basque—a low-resource language. We introduce the first publicly available, expert-annotated CEFR C1-level Basque essay dataset (3,200 essays), annotated across multidimensional quality criteria including accuracy, lexical richness, and coherence, alongside expert feedback and representative error examples. Methodologically, we propose a novel feedback quality evaluation framework integrating automatic consistency assessment with expert validation, and develop interpretable, teaching-oriented AES and feedback generation models via supervised fine-tuning of RoBERTa-EusCrawl and Latxa 8B/70B. Results show that fine-tuned Latxa significantly outperforms GPT-5 and Claude Sonnet 4.5 in scoring consistency and feedback utility, while detecting a broader spectrum of linguistic errors. This work establishes a high-quality benchmark dataset, reproducible methodology, and open-source tools for low-resource NLP research.

0 citationsRead paper

BERnaT: Basque Encoders for Representing Natural Textual Diversity

Dec 03, 2025

Language models often exhibit representational bias and reduced robustness due to overreliance on high-quality, standardized corpora, neglecting dialectal, historical, and informal linguistic variants. To address this, we propose *full-spectrum language modeling*, instantiated with Basque as a case study. We construct a multilingual, multi-source pretraining dataset integrating standardized texts, social media content, and historical corpora, and train three Transformer-based encoder configurations. Our key contribution is a novel hierarchical evaluation framework that—uniquely—partitions NLU tasks into standard and diverse subsets, enabling systematic quantification of models’ capacity to capture linguistic variation. Experiments demonstrate that models trained on mixed-domain data achieve significant performance gains on diverse linguistic contexts while maintaining competitive accuracy on standard benchmarks. This validates both the efficacy and necessity of explicitly modeling linguistic diversity in language model pretraining.

0 citationsRead paper

Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding

Jul 21, 2025

This paper systematically evaluates the genuine metaphor comprehension capabilities of large language models (LLMs), investigating whether their performance on metaphor interpretation tasks stems from deep semantic understanding or reliance on superficial cues—such as lexical overlap, sentence length, or syntactic patterns—in natural language inference (NLI) and question answering (QA). Method: Leveraging multiple public metaphor datasets, the authors design diverse prompting strategies and controlled ablation experiments to quantify correlations between model performance and surface-level linguistic features. Contribution/Results: Results demonstrate that LLMs’ metaphor interpretation is predominantly driven by surface features, in-context learning, and pretraining-induced linguistic priors—not by robust semantic reasoning. Consequently, standard evaluation protocols yield overly optimistic estimates of metaphor understanding. The paper proposes a more rigorous, bias-mitigated evaluation paradigm for metaphor comprehension and open-sources all data, code, and analysis frameworks to enable reproducible research and establish a new benchmark for trustworthy metaphor modeling.

0 citationsRead paper