Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了基础模型通过指令驱动和示例驱动两种方法在内容审核中的有效性,结果表明这两种方法都能显著提高在线内容的审核性能。
📝 Abstract
The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.
Problem

Research questions and friction points this paper is trying to address.

Content Moderation
Foundation Models
Policy Operationalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

instruction-driven approach
example-driven approach
foundation models
content moderation
ModerationBench
🔎 Similar Papers
No similar papers found.