AI Alignment through a Game-theoretic Lens: A Survey

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过博弈论视角审视AI对齐问题,针对偏好多样性、对齐优先级和时间动态性三大挑战,综述了现有方法并探讨了构建鲁棒、自适应及可验证AI系统所面临的难题。
📝 Abstract
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.
Problem

Research questions and friction points this paper is trying to address.

AI Alignment
Game-theoretic Lens
Preference Diversity
Dynamic Interactions
Complex Human Values
Innovation

Methods, ideas, or system contributions that make the work stand out.

Game-theoretic Analysis
Preference Diversity
Temporal Dynamics