Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

πŸ“… 2026-08-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the vulnerability of large language models to surface-form perturbations in safety alignment, which often leads to harmful requests being erroneously approved while harmless ones are incorrectly rejected. To mitigate this issue, the authors propose WIFA, an automatic intent-grouping data augmentation method that requires no external teacher models or human annotations. WIFA constructs pairs of structurally similar but semantically opposite samples to disentangle surface form from underlying intent. Combined with a two-stage fine-tuning strategy (WIFA-Boost) and Anchor-group Consistency-based Rejection Training (A-GCRT), the approach significantly enhances the model’s ability to discern true user intent. Experiments on Qwen demonstrate that WIFA-Boost effectively improves rejection of adversarially reformulated harmful queries, while A-GCRT reduces over-rejection on OR-Bench from 25.7% to 17.4%, outperforming existing baselines.
πŸ“ Abstract
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels. We use WIFA as a common data layer for two complementary fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and Anchored Group-Consistent Refusal Training (A-GCRT), which regularizes refusal/compliance decision scores across same-intent wrappers and anchors harmful and benign groups on opposite sides of a margin. In the Qwen setting, WIFA-Boost reaches the strongest transformed-harmful refusal, while A-GCRT reduces OR-Bench over-refusal from 25.7\% for the base model to 17.4\%; reproduced baselines do not match these operating points. Llama results and ablations over data structure, two-stage order, and A-GCRT components support this intent-group interpretation without claiming universal below-base over-refusal.
Problem

Research questions and friction points this paper is trying to address.

harmful refusal
surface-form shortcuts
wrapper-based prompts
over-refusal
intent-group supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wrapper-Based Intent-Form Augmentation
Intent-Group Supervision
Refusal Consistency
Safety Tuning
Over-Refusal Reduction
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
P
Ping Wu
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Key Laboratory of Safe AI and Superalignment
H
Haibo Tong
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Key Laboratory of Safe AI and Superalignment
F
Feifei Zhao
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Key Laboratory of Safe AI and Superalignment
Han Shen
Han Shen
Research Engineer, Ant Group; Ph.D., Rensselaer Polytechnic Institute
OptimizationReinforcement LearningAlignment
Y
Yu Shi
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
Y
Yilin Zhao
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Key Laboratory of Safe AI and Superalignment
S
Sicheng Shen
BrainCog Lab, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Key Laboratory of Safe AI and Superalignment
Guobin Shen
Guobin Shen
Institute of Automation, Chinese Academy of Sciences
bio-inspired neural networksspiking neural networksmachine learningcognitive science
Yun Luo
Yun Luo
Shanghai AI Lab
natural language processinggraph neural network
Yi Zeng
Yi Zeng
Institute of Automation, Chinese Academy of Sciences
Brain-inspired AIAI SafetyAI Ethics and Governance