Toward a Theory of Value in AI Alignment

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the prevalent lack of a clear definition of β€œhuman values” in current AI alignment research, which often reduces values to preferences or binary choices while overlooking their cultural embeddedness and pluralism. Through systematic analysis of 94 relevant publications, the work identifies implicit theoretical assumptions about values in the field and examines how values are operationalized in model training, evaluation, and automated scoring. Drawing on philosophical and social scientific perspectives, the authors employ content annotation and textual analysis to trace the technical pathways through which values are encoded. The findings reveal that most approaches conflate preferences with values and rely heavily on synthetic data, potentially suppressing value diversity. The paper offers a critical theoretical reflection on AI alignment and calls for the development of value frameworks grounded in deeper philosophical insight and greater cultural sensitivity.
πŸ“ Abstract
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on preferences as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignments philosophical commitments explicit, we seek to bring great specificity and under explored perspectives in the debate on whether and how AI can address human values.
Problem

Research questions and friction points this paper is trying to address.

AI alignment
human values
value theory
foundation models
AI safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI alignment
human values
preference reduction
synthetic data
value theory