🤖 AI Summary
This study addresses the critical gap in bias research—its predominant focus on gender and race while neglecting religion—by systematically investigating the implicit representation of religious identity in large language models (LLMs) and its associations with violence and geographic concepts. Using mechanistic interpretability, we apply sparse autoencoders (SAEs) integrated with the Neuronpedia API to analyze latent-layer feature activations across five mainstream LLMs. Results reveal strong conceptual coherence for religious categories in model representations; notably, Islam exhibits significantly stronger activation overlap with violence-related features than other religions, while religion–geography associations partially mirror real-world demographic distributions. This work uncovers the co-embedding of cultural stereotypes and empirical statistical patterns within LLM representations, providing the first neuron-level empirical framework for understanding the structural origins of religious bias in foundation models.
📝 Abstract
Despite growing research on bias in large language models (LLMs), most work has focused on gender and race, with little attention to religious identity. This paper explores how religion is internally represented in LLMs and how it intersects with concepts of violence and geography. Using mechanistic interpretability and Sparse Autoencoders (SAEs) via the Neuronpedia API, we analyze latent feature activations across five models. We measure overlap between religion- and violence-related prompts and probe semantic patterns in activation contexts. While all five religions show comparable internal cohesion, Islam is more frequently linked to features associated with violent language. In contrast, geographic associations largely reflect real-world religious demographics, revealing how models embed both factual distributions and cultural stereotypes. These findings highlight the value of structural analysis in auditing not just outputs but also internal representations that shape model behavior.