Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
It remains unclear whether existing commonsense reasoning benchmarks effectively predict the performance of large language models on real-world downstream tasks. This study systematically evaluates 23 models across a range of commonsense benchmarks—including their revised versions—and diverse downstream tasks, offering the first comprehensive quantification of these benchmarks’ predictive validity. Through model ranking comparisons, resource-controlled correlation analyses, and leave-one-model-family-out cross-validation—augmented by non-commonsense control tasks to enhance rigor—the work finds that revised benchmarks do not substantially improve predictive power. Commonsense benchmarks exhibit stable cross-model-family predictive validity only for a limited subset of specific downstream tasks, with overall performance highly task-dependent, thereby challenging their utility as general-purpose indicators of model capability.
📝 Abstract
Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.
Problem

Research questions and friction points this paper is trying to address.

commonsense benchmarks
predictive validity
large language models
downstream tasks
criterion validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

commonsense reasoning
benchmark validity
downstream task prediction
large language models
criterion validity
🔎 Similar Papers
No similar papers found.