Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过将AI对齐问题重新表述为凸影响空间上的线性优化,利用福利经济学工具解决多用户偏好聚合难题,提出并验证了新的对齐协议。
📝 Abstract
When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\unicode{x2013}$reinforcement learning from human feedback$\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.
Problem

Research questions and friction points this paper is trying to address.

Social Choice
Algorithmic Impact
Alignment Problem
Welfare Consequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

linear optimization
convex impact space
welfare economics
mechanism design
alignment protocols
🔎 Similar Papers
2024-06-06AAAI Conference on Artificial IntelligenceCitations: 1