🤖 AI Summary
This study addresses the absence of a unified moral benchmark in AI alignment and the divergent public expectations regarding the ethical standards applicable to AI systems, their designers, and human agents. Through moral psychology experiments employing a trolley-problem paradigm with manipulated agents—human, robot, and engineer—it provides the first empirical evidence that explicitly highlighting the human origins of AI design significantly shifts public moral judgments toward deontological reasoning. Although moral evaluations of robot and human actors do not differ significantly, AI actions traceable to human designers are subjected to stricter deontological constraints. Building on these findings, the paper introduces the “alignment target problem,” underscoring the complexity of moral judgment in AI alignment and the critical role of design attribution in shaping normative expectations.
📝 Abstract
The quest to align machine behavior with human values raises fundamental questions about the moral frameworks that should govern AI decision-making. Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Research into agent-type value forks has challenged this assumption by showing that people do not always hold AI systems to the same moral standards as humans. Yet this challenge is subject to two further questions: whether people evaluate AI behavior differently when its human origins are made visible, and whether people hold the humans who program AI systems to different moral standards than either the humans or the machines under evaluation. An experimental study on 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and company engineers programming a repair robot. We find no significant variation in the moral standards applied to the repairman and the robot. However, moral judgments shifted substantially when robot actions were described as the product of human design. Participants exhibited markedly more deontological reasoning when evaluating the robot programmed by engineers or the engineers programming it, suggesting that making human design visible activates heightened moral constraints. These findings provide evidence that people apply meaningfully different moral standards to AI systems, to humans acting in the same situation, and to the humans who design them. We call this divergence the alignment target problem. Whether these plural normative standards can be reconciled into a coherent framework for AI governance in high-stakes domains remains an open question.