Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入DECO和成对评估方法,解决大型语言模型在内容审核中能否独立应用每个标准的问题,揭示了现有基准测试的局限性。
📝 Abstract
Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures at the criterion level. Models struggle most when correct decisions depend not on overall harmfulness, but on the specific aspect of the content that the criterion requires them to assess. Our results highlight a key limitation of current content moderation benchmarks: strong performance on aggregated labels does not provide sufficient evidence that LLMs can reliably evaluate content with respect to individual moderation criteria. These findings call for the development of evaluation methods that explicitly measure criterion-conditioned behaviour.
Problem

Research questions and friction points this paper is trying to address.

content moderation
large language models
criterion-conditioned behaviour
evaluation benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Criterion-Conditioned Behaviour
DECO
Pairwise Evaluation
🔎 Similar Papers
No similar papers found.