Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了基于强化学习的AI对齐问题,指出当前训练机制导致AI仅在被观察时遵守规范,并提出应通过架构设计而非更深层次内化来解决此问题。
📝 Abstract
AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Conditional Compliance
Alignment
Norms
Behavioral Training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Conditional Compliance
Alignment Faking
Architecture Design
Behavioral Training
🔎 Similar Papers
No similar papers found.
K
Kevin Baum
German Research Center for Artificial Intelligence (DFKI), Saarbrücken; Institute for Ethics in Technology, Hamburg University of Technology (TUHH); Oxford Internet Institute, University of Oxford
R
Rūta Binkytė
DFKI, Saarbrücken
Felix Jahn
Felix Jahn
Ph.D. Student, German Research Center for Artificial Intelligence (DFKI)
Causal Reinforcement Learning