Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建C-SafeQA基准,针对中文内容安全评估问题,利用多模型裁决和专家审计方法评估大型语言模型响应的安全性。
📝 Abstract
Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by four full-model LLM deployments, yielding 37,660 query-response records labeled safe, unsafe, or disputed. Reference labels are generated through agreement-aware multi-model adjudication and blind audits of stratified subsets by three safety experts. C-SafeQA supports both evaluation of target-model safety and auditing of seven automated safety judges against shared reference labels. Unsafe-response rates range from 0.93% to 3.35% on base queries and from 11.68% to 30.05% on adversarial queries. On the adversarial subset, judges show substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate, and no judge dominates all metrics. Both acrostic transformations reduce unsafe recall for all seven judges, revealing mechanism-specific evaluator weaknesses. Dataset records, metadata, verification code, and judge scripts are publicly released to support recomputation, while benchmark construction, target-response generation, and private adjudication remain outside the release boundary.
Problem

Research questions and friction points this paper is trying to address.

Safety Evaluation
Chinese Content
Adversarial Queries
Policy Violation
Automated Safety Judges
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chinese Safety QA
Adversarial Queries
Policy-Grounded Benchmark
Multi-Model Adjudication
Safety Judges
💼 Related Jobs
No related jobs found.
R
Rui Yang
Anhui SparkShield Intelligent Technology
S
Shuang Huang
Anhui SparkShield Intelligent Technology
Junhua Liu
Junhua Liu
University of Southern California
Multimedia SystemsVR/AR/XRAI/ML Systems
Z
Ziqi Zhao
Anhui SparkShield Intelligent Technology
Q
Qingzhong Yan
Anhui SparkShield Intelligent Technology
Y
Yuhang Sun
Anhui SparkShield Intelligent Technology
C
Cong Liu
iFLYTEK
G
Guoping Hu
iFLYTEK
R
Rui Mei
Anhui SparkShield Intelligent Technology, Peking University
Jing Shao
Jing Shao
Research Scientist, Shanghai AI Laboratory/Shanghai Jiao Tong University
Computer VisionMulti-Modal Large Language Model