Institution profile

Duolingo

Industry researchnorthamerica · us
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Analytically Corrected Bayesian Modularization for Local Item Calibration

Aug 11, 2026

This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.

0 citationsRead paper

Group-Level Inference with Micro-Assessments

Aug 03, 2026

Traditional long assessments exhibit low statistical efficiency in estimating population-level latent trait means. This work proposes a closed-form approximation method within the item response theory (IRT) framework, wherein item information functions are approximated via Gaussian convolution. By integrating marginal maximum likelihood estimation with a stochastic item sampling mechanism, the approach explicitly models the trade-off between test length and sample size and incorporates a variance inflation factor accounting for item pool mismatch. The method enables unbiased inference of population means even with extremely short tests—including single-item assessments. Validation through exact benchmarks, Monte Carlo simulations, and resampling experiments on NAEP mathematics data demonstrates that estimates remain approximately unbiased even when based on a single item, substantially enhancing measurement efficiency.

0 citationsRead paper

Stochastic Constrained Test Assembly for AI-Enabled Assessment Systems

Jul 10, 2026

This study addresses the challenge in AI-based assessment systems where dynamic item banks contain newly introduced items with uncertain parameters, rendering traditional static test assembly inadequate for simultaneously ensuring measurement precision, adherence to content blueprint constraints, item bank sustainability, and efficient calibration of new items. To overcome this, the authors formulate linear test assembly as a constrained sequential decision-making problem and propose a Test-level Stochastic Constraint Hybrid (SCH) framework. SCH uniquely incorporates items with uncertain parameters directly into the assembly process, extending multi-armed bandit methodology from the item level to the test level by using Fisher information as the reward signal. Experimental results demonstrate that SCH significantly improves measurement accuracy, accelerates new-item calibration, and maintains item bank health while satisfying content constraints, outperforming six existing methods.

0 citationsRead paper

Learning Item Embeddings and Hyperparameters for IRT Calibration via Monte Carlo EM

Jul 07, 2026

This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.

0 citationsRead paper
Recent publications

Latest Papers

Analytically Corrected Bayesian Modularization for Local Item Calibration

Aug 11, 2026

This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.

0 citationsRead paper

Group-Level Inference with Micro-Assessments

Aug 03, 2026

Traditional long assessments exhibit low statistical efficiency in estimating population-level latent trait means. This work proposes a closed-form approximation method within the item response theory (IRT) framework, wherein item information functions are approximated via Gaussian convolution. By integrating marginal maximum likelihood estimation with a stochastic item sampling mechanism, the approach explicitly models the trade-off between test length and sample size and incorporates a variance inflation factor accounting for item pool mismatch. The method enables unbiased inference of population means even with extremely short tests—including single-item assessments. Validation through exact benchmarks, Monte Carlo simulations, and resampling experiments on NAEP mathematics data demonstrates that estimates remain approximately unbiased even when based on a single item, substantially enhancing measurement efficiency.

0 citationsRead paper

Stochastic Constrained Test Assembly for AI-Enabled Assessment Systems

Jul 10, 2026

This study addresses the challenge in AI-based assessment systems where dynamic item banks contain newly introduced items with uncertain parameters, rendering traditional static test assembly inadequate for simultaneously ensuring measurement precision, adherence to content blueprint constraints, item bank sustainability, and efficient calibration of new items. To overcome this, the authors formulate linear test assembly as a constrained sequential decision-making problem and propose a Test-level Stochastic Constraint Hybrid (SCH) framework. SCH uniquely incorporates items with uncertain parameters directly into the assembly process, extending multi-armed bandit methodology from the item level to the test level by using Fisher information as the reward signal. Experimental results demonstrate that SCH significantly improves measurement accuracy, accelerates new-item calibration, and maintains item bank health while satisfying content constraints, outperforming six existing methods.

0 citationsRead paper

Learning Item Embeddings and Hyperparameters for IRT Calibration via Monte Carlo EM

Jul 07, 2026

This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.

0 citationsRead paper