Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
This study investigates whether large language models implicitly infer Colombian nationality, socioeconomic status, or associated stereotypes through linguistic cues, even when such attributes are not explicitly mentioned. We introduce the first application of natural language autoencoders (NLAs) to bias detection in Latin American Spanish varieties, analyzing residual stream activations from layer 20 of the Qwen2.5-7B-Instruct model using multilingual prompt pairs and hierarchical positional probing. Our findings reveal that the model encodes nationality- and stereotype-related information about Colombia within its internal representations prior to output generation, particularly when inputs contain explicit or implicit Colombian cues. This work bridges activation-level interpretability with fairness evaluation for underrepresented linguistic groups, establishing a novel paradigm for detecting implicit bias in multilingual language models.