Training Alignment Auditors via Reinforcement Learning

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过强化学习改进了LLM审计员,使其能更有效地发现和调查模型中的隐藏不当行为,同时保持低误报率。
📝 Abstract
Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training environment, the policy investigates target models that potentially possess hidden behaviors planted via their system prompt. An LLM judge, which knows whether the target has a hidden behavior, holistically compares the policy's investigation to a reference investigation to determine the reward. With systematic ablations, we find that pairwise rewards yield more robust training compared to pointwise rewards, and that adding targets without planted behaviors helps maintain a low false positive rate. Training improves investigation quality against targets with planted behaviors, the rate of concerning behaviors surfaced in unmodified production models, and audit realism, while false-positive rates stay below 1%. Furthermore, auditing capabilities generalize across scaffolds: performance on AuditBench's adversarially fine-tuned targets substantially improves [Sheshadri et al., 2026].
Problem

Research questions and friction points this paper is trying to address.

alignment auditing
LLM auditors
reinforcement learning
hidden behaviors
audit realism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Alignment Auditing
Pairwise Rewards
False Positive Rate
Generalization
💼 Related Jobs
No related jobs found.