Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出两种技术,通过增加推理时间和模拟真实部署环境来提高对齐评估的真实性,解决模型在测试与实际部署中表现差异的问题。
📝 Abstract
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates multiple candidate actions, refines them using feedback from an instance of the target model on how to make them more realistic, and continues the evaluation with the most deployment-like candidate. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in an agent harness, reducing the gap between simulated and real deployment environments in coding settings. We test the techniques on multiple target models and find that they compose: applying both yields larger realism gains than either alone. Our results show that automated approaches can improve the realism of alignment evaluations, and that these improvements use additional compute more effectively than making the audits longer.
Problem

Research questions and friction points this paper is trying to address.

alignment evaluation
evaluation awareness
realism
Innovation

Methods, ideas, or system contributions that make the work stand out.

critique refinement
DISH
realism improvement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Axel Ahlqvist
Meridian Visiting Researcher Programme
R
Richard Guan
Cambridge Boston Alignment Initiative
Juan-Pablo Rivera
Juan-Pablo Rivera
Georgia Institute of Technology
A
Adeline Kassler
Meridian Visiting Researcher Programme
D
Dmitrii Troitskii
Cambridge Boston Alignment Initiative
A
Alexandra Souly
UK AI Security Institute
Kai Fronsdal
Kai Fronsdal
Masters Student, Stanford University
AI safetyAI alignment
Robert Kirk
Robert Kirk
Research Scientist, UK AI Security Institute
AI AlignmentAI SafetyLanguage ModelsFine-tuningGeneralisation
John Hughes
John Hughes
Anthropic
scalable oversightadversarial robustnessautomatic speech recognition