Score
Applying psychometric models to estimate latent traits from item responses, evaluate item discrimination and bias, and produce reliable comparisons and per-task confidence measures for tests, essays, or questionnaire instruments.
This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
Current psychological measurement item validation for large language models (LLMs) lacks efficient construct validity assessment methods. Method: This paper proposes a virtual validation framework grounded in mediation modeling: LLMs generate trait–response mediators—such as cognitive biases and social desirability tendencies—that reflect individual differences and drive simulated respondents’ diverse response behaviors, thereby evaluating items’ robustness in measuring target constructs (Big Five, Schwartz Values, VIA Strengths). Contribution/Results: This work is the first systematic investigation of LLMs’ potential for psychometric validity validation without requiring large-scale human-annotated data. Experiments demonstrate that LLMs reliably generate theoretically grounded mediators and accurately reproduce expected response patterns across all three major theoretical frameworks. The framework successfully supports item selection and validity evaluation, achieving substantial reductions in validation cost while maintaining methodological rigor.
This study investigates the applicability and ecological validity of human psychometric instruments—such as gender/racial bias and moral judgment scales—when adapted for evaluating large language models (LLMs). Employing a comprehensive assessment framework comprising multi-round item design, prompt variation testing, convergent validity analysis, and behavioral alignment with downstream tasks, we systematically evaluated the reliability and validity of 12 widely used psychological tests across LLMs. Results indicate moderate internal consistency (Cronbach’s α ≈ 0.65–0.78) but critically low ecological validity: model psychometric scores exhibit weak or even negative correlations with actual discriminatory outputs and fairness-related decision-making in realistic scenarios. The core contribution is the first proposal and empirical validation of an ecological validity evaluation paradigm specifically tailored for LLMs, demonstrating fundamental limitations in directly transplanting human-centered scales. These findings provide critical empirical grounding for theoretical reconceptualization and methodological innovation in AI-oriented psychological assessment.
To address limitations in scalability of psychometric tools, high barriers to integrating novel methods, and insufficient support for reproducibility and pedagogy, this study introduces ShinyItemAnalysis (SIA), a modular extensibility framework built in R and Shiny. SIA pioneers a module mechanism supporting embedded datasets, object-oriented design, and compiled code—enabling researchers to develop, package, and share interactive psychometric methods (e.g., IRT, EFA, DIF detection) as standalone R packages while seamlessly leveraging SIA’s integrated data processing, visualization, and core analytical capabilities. The framework has been instantiated with multiple open-source example modules. These advances substantially enhance method accessibility, computational reproducibility, and instructional utility, thereby fostering an open, interactive psychometric tool ecosystem across psychology, education, and the social sciences.
This study addresses the distortion in reliability reporting common in psychological research due to selective computation. For the first time, it applies a unified marginal reliability estimator across 889 psychometric datasets from the Item Response Warehouse, enabling standardized reliability assessment of cognitive tests, clinical screening instruments, and personality and attitude scales. Findings reveal that, depending on the reliability definition used, between 30% and 52% of datasets fall below the conventional .80 reliability threshold, with observed variability primarily attributable to genuine differences rather than estimation noise. The work underscores the substantial impact of reliability definition on substantive conclusions and advocates for more transparent and consistent reporting standards in psychometric practice.
This study addresses the potential pitfalls of directly applying item response theory (IRT)—originally designed for human assessment—to the evaluation of artificial intelligence systems, where mismatched data-generating mechanisms may compromise inference validity. It presents the first systematic evaluation of IRT’s applicability to large language model benchmarks, examining the feasibility, scalability, and reliability of four estimation approaches—marginal maximum likelihood, Markov chain Monte Carlo (MCMC), variational inference, and neural pseudo-twin estimators—across 18,000 simulated conditions. The findings reveal that classical methods are computationally infeasible at scale, while scalable alternatives introduce bias when the number of models is small or their ability distribution deviates from normality. The work quantifies, for the first time, the failure boundaries of IRT in AI evaluation and establishes required sample sizes and diagnostic criteria for its reliable application.
This study presents the first systematic evaluation of whether large language models (LLMs) can capture item discriminability—the core psychometric property that enables reading comprehension items to differentiate between students of varying ability levels. Under a zero-shot setting, 42 LLMs were assessed using two complementary approaches: direct prediction of item discriminability and classical test theory (CTT) calibration based on model-synthesized responses. Results indicate that although LLM outputs contain non-random signals, their discriminability patterns show limited alignment with human-calibrated benchmarks: the best Spearman correlation for direct prediction reaches only 0.152, improving modestly to 0.241 with the CTT approach—still far from reliably replicating the discrimination structure inherent in human responses. This work thus reveals significant limitations of current LLMs in psychometric modeling.
Current practices of directly applying human psychometric instruments to large language models (LLMs) to construct “psychological profiles” suffer from fundamental biases that may mislead research on their usability, safety, and agentic behavior. This study employs a psychometric framework to administer multiple personality and risk-preference scales to 56 instruction-tuned LLMs alongside large human samples, integrating self-report questionnaires, behavioral tasks, and variance decomposition into a multimodal assessment system. Findings reveal that 81–90% of inter-model differences stem from directional response biases rather than genuine traits; this bias diminishes with increasing model capability but persists nonetheless. Scale reliability is almost entirely predicted by a newly proposed metric—“response orthogonality.” These results demonstrate that LLM “psychological profiles” can be artificially manipulated through item selection, exposing critical limitations in prevailing evaluation paradigms.
This study systematically investigates the prevalence and psychometric consequences of deviations from the normality assumption of latent trait distributions in item response theory (IRT). Analyzing 504 real-world datasets, the authors employed flexible nonparametric and semiparametric methods to estimate trait distributions and compared these against conventional normality assumptions, evaluating impacts on reliability, item parameters, predicted responses, and individual scores. The large-scale empirical analysis reveals—for the first time—that more than half of the datasets exhibit cumulative distribution discrepancies exceeding 10 percentage points (with roughly one-fifth surpassing 20 points), frequently manifesting as skewed, heavy-tailed, flat, or multimodal shapes, with substantial variation across domains. These deviations exert particularly pronounced effects when full models are refitted. The findings underscore the necessity of routinely reporting distributional sensitivity analyses in IRT applications.