VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of verifiable reasoning benchmarks and deterministic evaluation methods for Chinese abusive speech moderation by constructing a specialized evaluation benchmark. It innovatively proposes field-anchored chain-of-thought reasoning with structured output protocols and designs deterministic metrics that eliminate reliance on LLM judges, thereby enabling auditable verification of reasoning processes. The research reveals that strong label performance often masks deficiencies in record completeness. Furthermore, it provides a reproducible evaluation framework supporting zero-shot prompting and taxonomy-guided classification. Collectively, this work fills a critical gap in structured reasoning verification within the domain of content moderation, offering robust methodologies for ensuring both accuracy and transparency in automated detection systems.
📝 Abstract
The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.
Problem

Research questions and friction points this paper is trying to address.

Chinese abusive speech moderation
verifiable structured reasoning
benchmark
deterministic verification
chain-of-thought rationales
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifiable Structured Reasoning
Field-anchored Chain-of-Thought
Deterministic Evaluation Protocol
Chinese Abusive Speech Moderation
Auditable Benchmark
🔎 Similar Papers
No similar papers found.
M
Mingyu Yuan
MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics
S
Shengtao Wen
MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics
Lingbing Guo
Lingbing Guo
Tianjin University
Machine learningArtificial Intelligence
Zhen Bi
Zhen Bi
Zhejiang University, Huzhou University
Knowledge GraphLanguage ModelOn-device LLM
X
Xiang Chen
MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics