Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of bias from implicit developer decisions and ineffective vote aggregation in moral AI preference alignment. It presents the first systematic empirical analysis of how feature specification, sample selection, and question framing influence alignment outcomes. Through multi-scenario experiments, ideological analysis, and moral foundations modeling, the research demonstrates that ethical features lack cross-domain transferability, political orientation significantly shapes preferences, and question wording can shift ideological gaps by up to one scale point. These findings reveal the underlying mechanisms by which implicit decisions construct preference outcomes, underscoring the necessity for comprehensive auditing and disclosure throughout the alignment pipeline to ensure fairness and transparency in moral AI systems.
📝 Abstract
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones. We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral AI elicitation. Across two phases (N = 809) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.
Problem

Research questions and friction points this paper is trying to address.

Participatory Moral AI
Moral Preference Elicitation
Developer Bias
AI Alignment
Value Sensitive Design
Innovation

Methods, ideas, or system contributions that make the work stand out.

Moral AI Elicitation
Developer Bias
Framing Effect
Ideological Divergence
Alignment Audit