🤖 AI Summary
This study systematically investigates the prevalence and psychometric consequences of deviations from the normality assumption of latent trait distributions in item response theory (IRT). Analyzing 504 real-world datasets, the authors employed flexible nonparametric and semiparametric methods to estimate trait distributions and compared these against conventional normality assumptions, evaluating impacts on reliability, item parameters, predicted responses, and individual scores. The large-scale empirical analysis reveals—for the first time—that more than half of the datasets exhibit cumulative distribution discrepancies exceeding 10 percentage points (with roughly one-fifth surpassing 20 points), frequently manifesting as skewed, heavy-tailed, flat, or multimodal shapes, with substantial variation across domains. These deviations exert particularly pronounced effects when full models are refitted. The findings underscore the necessity of routinely reporting distributional sensitivity analyses in IRT applications.
📝 Abstract
Item response models usually assume a normal trait distribution, yet little is known about how often fitted distributions in real studies differ substantially from normality or which reported results are most affected. We analyzed 504 itemresponse data sets from 273 studies in the Item Response Warehouse, fitting each with the normal assumption and with a flexible distribution estimated from the responses, and we compared reliability, item estimates, predicted test responses, and person scores between the two calibrations. In more than half of the data sets the two fitted distributions differed by at least 10 percentage points of cumulative probability at some point on the trait scale, with a median maximum difference of about 11 points; about one-third reached 15 points, and nearly one-fifth reached 20 points. The estimated shapes included skewness, heavy tails, flat regions, and occasional multimodality. Differences were larger in several attitudinal, affective, and behavioral domains, although these contrasts weakened when data sets were compared within item-model families. Reliability usually changed little when item estimates were held fixed, whereas refitting the full model produced larger changes in some data sets, and item estimates, predicted test responses, and person scores did not change in parallel. Alternative flexible methods generally identified the same data sets as most unusual but disagreed somewhat about magnitude. Applied analyses should state the distributional assumption, examine a flexible alternative, and report sensitivity separately for each result used in interpretation or decision making.