Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

๐Ÿ“… 2026-08-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of detecting harmful dialogue, which continuously evolves through type shifting and lexical evasion. To tackle this, the authors propose the BRACE framework, which formalizes the invariant ordered reasoning chain (ORC)โ€”comprising four differentiable stages: topic โ†’ indicator โ†’ severity โ†’ typeโ€”and integrates it as a structured regularization term into the model. Through mechanisms including intermediate supervision, prototype feature enhancement, and path disentanglement, BRACE effectively mitigates type confusion under semantic ambiguity. Evaluated across five categories of harmful content spanning four domains, the framework achieves macro F1 scores of 0.934 and 0.949 using RoBERTa-wwm-ext and a Qwen3-1.7B LoRA decoder, respectively, demonstrating significant improvements in dynamic harmful content detection.
๐Ÿ“ Abstract
Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type characteristics, which can help us capture the key information in the frequently changing lexical expressions. We propose BRACE, which encodes the ORC as four differentiable stages (Topic -> Indicator -> Severity -> Type) with intermediate supervision, serving as a structured regularizer blended with direct heads, and supported by prototype-based feature augmentation and feature path disentanglement. The evaluation results show that, across 4 domains and 5 harm categories, BRACE achieves harm-type macro F1 of 0.934 (RoBERTa-wwm-ext, 3-seed mean), with decoder backbones (Qwen3-1.7B LoRA) reaching 0.949. Ablation studies show that all components contribute to BRACE, and the structural decomposition of ORC enables BRACE to distinguish harmful types with semantic ambiguity. Disclaimer: This paper may contain content that is disturbing to some readers.
Problem

Research questions and friction points this paper is trying to address.

harmful chat dialogue
type-shifting
lexical evasion
harm detection
semantic ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ordered Reasoning Chain
harmful dialogue detection
structured regularization
prototype-based augmentation
feature disentanglement
๐Ÿ”Ž Similar Papers
No similar papers found.
H
Haojie Yu
State Key Laboratory of Complex System Modeling and Simulation Technology, Beijing, China; Science and Technology on Integrated Information System Laboratory, Institute of Software, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences
Ziyou Jiang
Ziyou Jiang
Institute of Software Chinese Academy of Sciences
software engineering
Junjie Wang
Junjie Wang
Institute of Software, Chinese Academy of Sciences
Software Engineering
Mingyang Li
Mingyang Li
Associate Professor, Industrial and Management Systems Engineering, The University of South Florida
data sciencereliability and qualitysystem informaticscomplex systems modeling and optimizationcomputational intelligence
Y
Yuekai Huang
State Key Laboratory of Complex System Modeling and Simulation Technology, Beijing, China; Science and Technology on Integrated Information System Laboratory, Institute of Software, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences
J
Jie Huang
State Key Laboratory of Complex System Modeling and Simulation Technology, Beijing, China; Science and Technology on Integrated Information System Laboratory, Institute of Software, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences
Qing Wang
Qing Wang
Institute of Software Chinese Academy of Sciences
Software engineering