🤖 AI Summary
This study investigates the capacity of large language models (LLMs) to detect, reproduce, and correct authentic speech errors produced by native Spanish speakers, thereby exposing fundamental limitations in their emulation of human linguistic cognition. Methodologically, it integrates interdisciplinary approaches: (1) constructing a curated, annotated corpus of 500+ naturally occurring Spanish errors; (2) conducting high-temporal-resolution EEG experiments to characterize real-time neural processing dynamics; and (3) systematically evaluating state-of-the-art LLMs—including GPT and Gemini—on error detection, generative generalization, and fault-tolerant reasoning. This work achieves the first tri-level, cross-domain modeling of linguistic errors across cognitive representation, neurophysiological response, and AI behavioral output. Results advance theoretical understanding of Spanish linguistic competence and variation, while providing an empirical foundation and methodological framework for developing cognitively aligned, robust, and error-resilient NLP systems.
📝 Abstract
Linguistic errors are not merely deviations from normative grammar; they offer a unique window into the cognitive architecture of language and expose the current limitations of artificial systems that seek to replicate them. This project proposes an interdisciplinary study of linguistic errors produced by native Spanish speakers, with the aim of analyzing how current large language models (LLM) interpret, reproduce, or correct them. The research integrates three core perspectives: theoretical linguistics, to classify and understand the nature of the errors; neurolinguistics, to contextualize them within real-time language processing in the brain; and natural language processing (NLP), to evaluate their interpretation against linguistic errors. A purpose-built corpus of authentic errors of native Spanish (+500) will serve as the foundation for empirical analysis. These errors will be tested against AI models such as GPT or Gemini to assess their interpretative accuracy and their ability to generalize patterns of human linguistic behavior. The project contributes not only to the understanding of Spanish as a native language but also to the development of NLP systems that are more cognitively informed and capable of engaging with the imperfect, variable, and often ambiguous nature of real human language.