Mechanism Design for Alignment and Control

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对AI代理的偏好与能力未知的问题,通过设计一种单边模仿结构机制来激励其诚实与服从,并应用于多个具体场景。
📝 Abstract
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.
Problem

Research questions and friction points this paper is trying to address.

mechanism design
AI agents
alignment
capabilities
honesty
Innovation

Methods, ideas, or system contributions that make the work stand out.

mechanism design
AI agents
revelation principle
nested cyclical monotonicity
higher-order beliefs
🔎 Similar Papers
No similar papers found.