Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究评估了现有AI安全基准对小型语言模型的有效性和可靠性,发现这些基准在评估小型语言模型时存在局限性,特别是在处理复杂提示和模型架构时。
📝 Abstract
Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-scale assessment of the effectiveness and robustness of these automated pipelines by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which assigns a score of 0, 1, or 0.5 to harmful, safe, or ambiguous/irrelevant responses, respectively. Across the benchmarks, ambiguous judgments dominate and correlate with prompt complexity and model architecture, indicating that {\em LLM-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment}. In general, the ambiguity rate increases with lexical density, output perplexity, and output length and decreases with lexical sophistication, self-coherence, and reply-prompt similarity. This reveals a capability-safety confound that mixes model capability with apparent safety. Since ambiguity is prevalent, aggregate mean-score leaderboards are mathematically brittle: model rankings change significantly under reasonable ambiguity treatments, even when the underlying outputs remain unchanged.
Problem

Research questions and friction points this paper is trying to address.

Small Language Models
Safety Benchmarks
Ambiguity
Capability-Safety Confound
Innovation

Methods, ideas, or system contributions that make the work stand out.

small language models
safety benchmarks
ambiguity in evaluation
capability-safety confound
benchmark robustness
💼 Related Jobs
No related jobs found.