🤖 AI Summary
This study investigates whether large language models implicitly infer Colombian nationality, socioeconomic status, or associated stereotypes through linguistic cues, even when such attributes are not explicitly mentioned. We introduce the first application of natural language autoencoders (NLAs) to bias detection in Latin American Spanish varieties, analyzing residual stream activations from layer 20 of the Qwen2.5-7B-Instruct model using multilingual prompt pairs and hierarchical positional probing. Our findings reveal that the model encodes nationality- and stereotype-related information about Colombia within its internal representations prior to output generation, particularly when inputs contain explicit or implicit Colombian cues. This work bridges activation-level interpretability with fairness evaluation for underrepresented linguistic groups, establishing a novel paradigm for detecting implicit bias in multilingual language models.
📝 Abstract
Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts. We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt. Our dataset contains 30 prompts arranged as 15 matched Spanish-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls. We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype representations appear before they are verbalized in the model output. This work connects activation-level interpretability with bias evaluation for underrepresented Spanish varieties.