The Benchmark Trap: Structures of Power and Injustice in AI Evaluations

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the exacerbation of power concentration and structural inequities by AI benchmarks through an innovative integration of oppression theory into evaluative critique. Synthesizing social theory, critical technology studies, and institutional discourse analysis, it reconstructs the sociopolitical dimensions of benchmarking. The research elucidates how benchmarks, as sociotechnical artifacts, entrench power asymmetries and constrain research trajectories, thereby revealing the systemic harms perpetuated by "benchmark culture." Transcending purely technical evaluation paradigms, this work warns that such practices undermine both cognitive robustness and socially beneficial progress in AI. Ultimately, it provides essential theoretical foundations for establishing equitable AI governance frameworks.
📝 Abstract
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
Problem

Research questions and friction points this paper is trying to address.

AI benchmarks
structural injustice
power structures
oppression
benchmark trap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark Trap
Structural Injustice
Socio-technical Artefacts
AI Evaluation
Power Structures
🔎 Similar Papers