TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对大型推理模型中间推理痕迹中的不安全内容检测问题,提出了TRACE基准测试,涵盖从提示到最终响应的全过程,并提供了证据注释。
📝 Abstract
Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchmarks for unsafe content detection focus primarily on prompts and final responses, leaving reasoning traces largely unexamined. Moreover, these benchmarks typically provide only binary safety labels, without evidence annotations that justify the judgments. To address these limitations, we introduce TRACE, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses. TRACE includes prompts in two languages spanning nine risk categories and ten attack strategies. For each prompt, four LRMs generate reasoning traces and final responses, and we annotate the safety of each component and extract supporting evidence from the corresponding source text. Evaluating 18 guardrail models on TRACE reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models struggle to accurately extract supporting evidence. These findings highlight the need for guardrail models that can reliably detect and precisely localize unsafe content across the LRM inference pipeline.
Problem

Research questions and friction points this paper is trying to address.

Large Reasoning Models
unsafe content
reasoning traces
safety evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence-grounded
safety evaluation
reasoning traces
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhenyu Wu
King Abdullah University of Science and Technology
S
Siyuan Chen
King Abdullah University of Science and Technology
Changchun Yang
Changchun Yang
KAUST
medical image analysisimaging science
J
Jiaqi Dong
King Abdullah University of Science and Technology
M
Min Zhou
King Abdullah University of Science and Technology
Ali Almadan
Ali Almadan
Aramco
T
Talal Hammad
Aramco
F
Faisal Wahbo
Aramco
A
Aminullah Tora
Aramco
M
Mona Alshahrani
Aramco
X
Xin Gao
King Abdullah University of Science and Technology