🤖 AI Summary
This study addresses the distortion in reliability reporting common in psychological research due to selective computation. For the first time, it applies a unified marginal reliability estimator across 889 psychometric datasets from the Item Response Warehouse, enabling standardized reliability assessment of cognitive tests, clinical screening instruments, and personality and attitude scales. Findings reveal that, depending on the reliability definition used, between 30% and 52% of datasets fall below the conventional .80 reliability threshold, with observed variability primarily attributable to genuine differences rather than estimation noise. The work underscores the substantial impact of reliability definition on substantive conclusions and advocates for more transparent and consistent reporting standards in psychometric practice.
📝 Abstract
Nearly every quantitative study in psychology reports a reliability coefficient, so the field knows a great deal about the reliability that authors choose to publish. It knows much less about the reliability of the data psychology actually produces, because published coefficients pass through decisions about what to compute and what to report. We therefore measure reliability directly, applying the same estimators under the same rules to 889 datasets from the Item Response Warehouse, a public collection of item-response data that spans cognitive tests, clinical screeners, personality inventories, and attitude scales. Three findings emerge. Low reliability is common: even under a lenient definition of reliability, 30% of datasets fall below the conventional .80 threshold. The variation across datasets is real rather than statistical, since estimation noise accounts for only about one percent of it. Finally, the answer depends on the definition itself: under a strict definition the share below .80 rises to 52%, a difference large enough to change what one concludes about the field. We conclude that a reliability report should say which definition it uses, attach a measure of uncertainty, and give a strict coefficient alongside a lenient one.