MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
This work addresses the limitations of current ethical evaluations of large language models, which predominantly rely on moral foundations theory while overlooking critical dimensions such as social values and personality traits that shape human moral judgment. To bridge this gap, the authors introduce MOSAIC, the first large-scale, multidimensional ethical benchmark that integrates nine standardized scales from moral philosophy, psychology, and social theory, along with four contextualized game-theoretic tasks, to form an extensible and ready-to-use evaluation framework. Experiments across three mainstream models demonstrate that moral foundations alone are insufficient for comprehensively characterizing AI ethical behavior. The project further contributes an open-source dataset and a Python evaluation library to advance more holistic and human-aligned assessments of AI ethics.