Bayesian Consensus Calibration of Continuously Evolving IRT Item Banks
本文提出共识校准方法,通过分治策略独立校准各时段并重构后验分布,解决了大规模、稀疏且频繁更新的IRT题库的高效校准问题。
本文提出共识校准方法,通过分治策略独立校准各时段并重构后验分布,解决了大规模、稀疏且频繁更新的IRT题库的高效校准问题。
This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.
Traditional long assessments exhibit low statistical efficiency in estimating population-level latent trait means. This work proposes a closed-form approximation method within the item response theory (IRT) framework, wherein item information functions are approximated via Gaussian convolution. By integrating marginal maximum likelihood estimation with a stochastic item sampling mechanism, the approach explicitly models the trade-off between test length and sample size and incorporates a variance inflation factor accounting for item pool mismatch. The method enables unbiased inference of population means even with extremely short tests—including single-item assessments. Validation through exact benchmarks, Monte Carlo simulations, and resampling experiments on NAEP mathematics data demonstrates that estimates remain approximately unbiased even when based on a single item, substantially enhancing measurement efficiency.
This study addresses the challenge in AI-based assessment systems where dynamic item banks contain newly introduced items with uncertain parameters, rendering traditional static test assembly inadequate for simultaneously ensuring measurement precision, adherence to content blueprint constraints, item bank sustainability, and efficient calibration of new items. To overcome this, the authors formulate linear test assembly as a constrained sequential decision-making problem and propose a Test-level Stochastic Constraint Hybrid (SCH) framework. SCH uniquely incorporates items with uncertain parameters directly into the assembly process, extending multi-armed bandit methodology from the item level to the test level by using Fisher information as the reward signal. Experimental results demonstrate that SCH significantly improves measurement accuracy, accelerates new-item calibration, and maintains item bank health while satisfying content constraints, outperforming six existing methods.
This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.
本文提出共识校准方法,通过分治策略独立校准各时段并重构后验分布,解决了大规模、稀疏且频繁更新的IRT题库的高效校准问题。
This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.
Traditional long assessments exhibit low statistical efficiency in estimating population-level latent trait means. This work proposes a closed-form approximation method within the item response theory (IRT) framework, wherein item information functions are approximated via Gaussian convolution. By integrating marginal maximum likelihood estimation with a stochastic item sampling mechanism, the approach explicitly models the trade-off between test length and sample size and incorporates a variance inflation factor accounting for item pool mismatch. The method enables unbiased inference of population means even with extremely short tests—including single-item assessments. Validation through exact benchmarks, Monte Carlo simulations, and resampling experiments on NAEP mathematics data demonstrates that estimates remain approximately unbiased even when based on a single item, substantially enhancing measurement efficiency.
This study addresses the challenge in AI-based assessment systems where dynamic item banks contain newly introduced items with uncertain parameters, rendering traditional static test assembly inadequate for simultaneously ensuring measurement precision, adherence to content blueprint constraints, item bank sustainability, and efficient calibration of new items. To overcome this, the authors formulate linear test assembly as a constrained sequential decision-making problem and propose a Test-level Stochastic Constraint Hybrid (SCH) framework. SCH uniquely incorporates items with uncertain parameters directly into the assembly process, extending multi-armed bandit methodology from the item level to the test level by using Fisher information as the reward signal. Experimental results demonstrate that SCH significantly improves measurement accuracy, accelerates new-item calibration, and maintains item bank health while satisfying content constraints, outperforming six existing methods.
This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.