Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current safety evaluations of large language models predominantly rely on single canonical prompts, which inadequately capture model safety under surface-form variations that preserve semantic equivalence. To address this limitation, this work proposes a multi-form joint evaluation framework that generates semantically consistent but linguistically diverse prompts via machine back-translation and code-switching. The framework integrates human-anchored judgments (using Claude, κ=0.86), vendor-neutral safety criteria, and statistical tests (McNemar’s test and bootstrap) to systematically quantify the impact of surface form on safety assessments. Experiments reveal that canonical prompts significantly underestimate risk: 13% of unsafe behaviors manifest only in non-canonical forms. Across five mainstream models, multi-form evaluation yields unsafe rates 3.3–12.9 percentage points higher than the worst-performing single form—differences that are all statistically significant, with Gemini 2.5 Pro showing the greatest sensitivity. Notably, just 3–4 prompt forms suffice to uncover 85% of unsafe behaviors.
📝 Abstract
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much of any variation is decoding/judge noise rather than signal? We instantiate this in safety, a high-stakes setting with no gold label to average toward. To avoid prior confounds, we pre-author the reformulations (refusal-free, mostly non-LLM: machine back-translation and a Matrix-Language-Frame code-switch generator) so an identical surface form reaches every model, score all responses with one human-anchored, vendor-neutral judge (Claude, kappa = 0.86 vs. human on unsafe compliance, stable across languages, cross-checked by GPT-4o), and verify intent preservation. On 370 seeds x 5 surface forms x 5 models, no single transformation is uniformly most dangerous (6 of 20 per-transformation McNemar tests survive correction, most protective). Yet evaluating only the canonical prompt underestimates unsafe compliance: the union of unsafe outcomes across forms exceeds even the worst single form by 3.3-12.9 pp, with bootstrap 95% CIs excluding zero for all five models, and 5-13% of seeds safe on canonical are unsafe under some reformulation -- above a zero stochasticity floor (canonical resampled five times at temperature 0 gives 0/370 new exposures). The size of this gap is model-dependent (largest on Gemini 2.5 Pro). One form recovers only ~53% of a model's observed unsafe surface and about three reach 85% -- a redundancy characterization of this form set, not of a defined population. A benign control (XSTest) suggests the instability is bidirectional, though the benign and harmful pools are not item-matched. We release the dataset, code, and per-response labels.
Problem

Research questions and friction points this paper is trying to address.

surface-form sensitivity
LLM safety
benchmarking
prompt reformulation
unsafe compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

surface-form sensitivity
safety evaluation
prompt reformulation
LLM robustness
benchmarking bias
💼 Related Jobs
No related jobs found.
Y
Yongxi Zhou
Northeastern University, Massachusetts, USA
J
Junwei Yao
Northeastern University, Massachusetts, USA
Yuanzhe Liu
Yuanzhe Liu
CS Ph.D. Student at Rensselaer Polytechnic Institute
multi agentcode optimizationcontrollable music generation
Z
Zihan Dong
Georgia Institute of Technology, Georgia, USA
W
Wenbo Ye
University of Southern California, California, USA
J
Jiaxi Wen
Northeastern University, Massachusetts, USA
L
Lai Yun Choi
Northeastern University, Massachusetts, USA