Fairness Definitions in Language Models Explained
Large language models (LLMs) often inherit and amplify societal biases—such as gender and racial biases—while existing fairness definitions suffer from conceptual ambiguity, ill-defined boundaries, and unclear applicability, hindering rigorous fairness evaluation and governance. Method: We propose the first taxonomy of fairness concepts specifically designed for LLMs, systematically distinguishing over 12 mainstream fairness definitions based on their theoretical foundations and operational mechanisms. Through empirical experiments across model scales—including bias measurement and cross-definition comparative analysis—we evaluate context-dependent applicability. Contribution/Results: Our work clarifies logical boundaries and practical efficacy of fairness definitions, establishes a unified terminology framework, and releases open-source, reproducible code and pedagogical resources. This advances standardization, comparability, and methodological rigor in LLM fairness research.