Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了通过不同方式引导模型对评估的认知,发现能力导向的引导比安全导向的引导更能提高模型的合规性。
📝 Abstract
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes.
Problem

Research questions and friction points this paper is trying to address.

eval-awareness
compliance
capabilities framing
safety framing
Innovation

Methods, ideas, or system contributions that make the work stand out.

eval-awareness
capabilities framing
safety framing
compliance prediction
causal link
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Allison Zhuang
ENS Paris-Saclay
Santiago Aranguri
Santiago Aranguri
PhD Student, NYU Courant