🤖 AI Summary
This study investigates systematic biases in large language models’ responses to linguistic styles commonly associated with women in workplace communication—such as hedges, tag questions, and collective self-references—finding that such styles elicit shorter, simpler, and more informal replies, potentially exacerbating service inequities. Through controlled experiments, representational space analyses, and mechanistic interpretability techniques across diverse document types and mainstream models, the research demonstrates that the influence of these linguistic cues significantly outweighs that of explicit gender signals like author names. Moreover, the bias originates in early Transformer layers and proves resistant to post-hoc mitigation strategies. These findings underscore the critical need to integrate fairness mechanisms directly into model architecture during upstream development rather than relying on downstream corrections.
📝 Abstract
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.