Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过改变计算条件测试AI模型在负责任的AI基准测试中的结论稳定性,使用不同批次、量化和基准缩减方法评估模型的准确性、偏见等性能。
📝 Abstract
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations. Rather than treating preserved aggregate accuracy as sufficient, we compare accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy against a full-benchmark BF16 baseline. Larger batching keeps accuracy within 0.35 percentage points of baseline and produces comparatively small subgroup changes, while reducing energy in five of six model--dataset settings. INT8 largely preserves quality but uses 1.79--4.26$\times$ baseline energy. INT4 causes larger, model- and context-dependent changes. Reduced benchmarks provide the most consistent savings, but very small subsets are substantially more sensitive to which items are retained. Efficient evaluation should therefore be treated as a measurement intervention whose validity must be checked across the conclusions the benchmark is intended to support. Our project website is https://vectorinstitute.github.io/sustainable-rai-evaluation/ and the code is available at https://github.com/VectorInstitute/sustainable-rai-evaluation.
Problem

Research questions and friction points this paper is trying to address.

Efficient evaluation
Benchmarking
Model behavior
Conclusion robustness
Responsible-AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stress-Testing
Efficient Responsible-AI Evaluation
Benchmarking
Model Behavior
Energy Consumption
🔎 Similar Papers