Institution profile

RTI International

Academic institutionnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Design-Based Supervised Learning with Noisy Human Labels

Jul 16, 2026

This study addresses the bias in classifier training and statistical inference caused by noisy human-reviewed labels. To mitigate this issue, the authors propose Partially Adjudicated Design-based Supervised Learning (PA-DSL), a novel framework that integrates partial expert adjudication into design-based supervised learning. By combining probability sampling audits, label noise correction, and design-weighted estimation, PA-DSL leverages recoverable signals from noisy labels while ensuring unbiased estimation. The method is applicable to various downstream tasks where audit and adjudication probabilities are known. Experiments on synthetic data and semi-synthetic Wikipedia Detox datasets demonstrate that, compared to approaches using only adjudicated labels, PA-DSL reduces root mean squared error by 10%–17% while maintaining nominal coverage.

0 citationsRead paper

Enhancing Computational Efficiency in NetLogo: Best Practices for Running Large-Scale Agent-Based Models on AWS and Cloud Infrastructures

Feb 16, 2026

This study addresses the significant computational overhead, performance instability, and high costs commonly encountered when running large-scale agent-based models (ABMs) in NetLogo. To tackle these challenges, the authors propose the first cloud deployment optimization framework specifically designed for large-scale NetLogo ABMs, which systematically integrates memory management, JVM parameter tuning, BehaviorSpace execution strategies, and AWS instance selection. Through experiments on the canonical wolf-sheep predation model, the framework quantitatively evaluates the impact of different cloud instances on performance and cost. The results demonstrate a 32% reduction in computational expenses while substantially improving runtime stability and efficiency.

0 citationsRead paper

Privacy Amplification for Synthetic data using Range Restriction

Feb 04, 2026

This study addresses the challenge of balancing privacy preservation and data utility in synthetic data generation. The authors propose a range-restricted privacy mechanism that formalizes data owners’ prior knowledge about sensitive value domains into two probabilistic adjustment strategies, applying protection only to truly sensitive subsets. By integrating range constraints and belief modeling within a risk-weighted pseudo-posterior framework, the method achieves localized amplification of differential privacy guarantees. Experimental results demonstrate that, compared to conventional pseudo-posterior mechanisms under asymptotic differential privacy, the proposed approach significantly enhances privacy protection in sensitive regions while effectively preserving the overall utility of the generated synthetic data.

0 citationsRead paper

Developing synthetic microdata through machine learning for firm-level business surveys

Dec 05, 2025

Traditional anonymization of enterprise-level commercial survey data faces significant re-identification risks and struggles to balance confidentiality with analytical utility. To address this, we propose a generative machine learning–based method for synthesizing microdata. Our approach integrates multidimensional distribution matching across geographic and industrial dimensions, employs domain-specific quality metrics, and enforces statistical moment constraints to ensure high-fidelity synthetic data that closely replicates key statistical properties and economic inference outcomes of the original dataset. We successfully generated a synthetic dataset for the 2007 Business Owner Survey and fully reproduced an empirical study published in *Small Business Economics*, thereby validating the method’s statistical validity and reproducibility. This work bridges a critical technical gap in secure commercial survey data dissemination and provides a scalable, regulatory-compliant framework for sharing sensitive microdata.

0 citationsRead paper

Uncertainty Quantification for Multi-level Models Using the Survey-Weighted Pseudo-Posterior

Oct 10, 2025

To address inadequate calibration of both local (e.g., group-level random effects) and global parameters in multilevel Bayesian models under complex survey designs, this paper proposes an improved survey-weighted pseudo-posterior framework with an automated post-processing pipeline—enabling, for the first time, consistent uncertainty calibration for both parameter types within hierarchical structures. The method integrates mixed-effects modeling, weighted likelihood construction, and posterior reweighting calibration, ensuring asymptotic consistency theoretically and seamless compatibility with mainstream Bayesian software computationally. Simulation studies and empirical analysis using the National Survey on Drug Use and Health (NSDUH) demonstrate that the proposed approach substantially improves estimation accuracy and reliability of uncertainty quantification, particularly reducing bias in random-effects inference. The corresponding algorithms are publicly available and integrated into the R package `csSampling`, offering a scalable, reproducible solution for hierarchical Bayesian analysis under complex sampling.

0 citationsRead paper
Recent publications

Latest Papers

Design-Based Supervised Learning with Noisy Human Labels

Jul 16, 2026

This study addresses the bias in classifier training and statistical inference caused by noisy human-reviewed labels. To mitigate this issue, the authors propose Partially Adjudicated Design-based Supervised Learning (PA-DSL), a novel framework that integrates partial expert adjudication into design-based supervised learning. By combining probability sampling audits, label noise correction, and design-weighted estimation, PA-DSL leverages recoverable signals from noisy labels while ensuring unbiased estimation. The method is applicable to various downstream tasks where audit and adjudication probabilities are known. Experiments on synthetic data and semi-synthetic Wikipedia Detox datasets demonstrate that, compared to approaches using only adjudicated labels, PA-DSL reduces root mean squared error by 10%–17% while maintaining nominal coverage.

0 citationsRead paper

Enhancing Computational Efficiency in NetLogo: Best Practices for Running Large-Scale Agent-Based Models on AWS and Cloud Infrastructures

Feb 16, 2026

This study addresses the significant computational overhead, performance instability, and high costs commonly encountered when running large-scale agent-based models (ABMs) in NetLogo. To tackle these challenges, the authors propose the first cloud deployment optimization framework specifically designed for large-scale NetLogo ABMs, which systematically integrates memory management, JVM parameter tuning, BehaviorSpace execution strategies, and AWS instance selection. Through experiments on the canonical wolf-sheep predation model, the framework quantitatively evaluates the impact of different cloud instances on performance and cost. The results demonstrate a 32% reduction in computational expenses while substantially improving runtime stability and efficiency.

0 citationsRead paper

Privacy Amplification for Synthetic data using Range Restriction

Feb 04, 2026

This study addresses the challenge of balancing privacy preservation and data utility in synthetic data generation. The authors propose a range-restricted privacy mechanism that formalizes data owners’ prior knowledge about sensitive value domains into two probabilistic adjustment strategies, applying protection only to truly sensitive subsets. By integrating range constraints and belief modeling within a risk-weighted pseudo-posterior framework, the method achieves localized amplification of differential privacy guarantees. Experimental results demonstrate that, compared to conventional pseudo-posterior mechanisms under asymptotic differential privacy, the proposed approach significantly enhances privacy protection in sensitive regions while effectively preserving the overall utility of the generated synthetic data.

0 citationsRead paper

Developing synthetic microdata through machine learning for firm-level business surveys

Dec 05, 2025

Traditional anonymization of enterprise-level commercial survey data faces significant re-identification risks and struggles to balance confidentiality with analytical utility. To address this, we propose a generative machine learning–based method for synthesizing microdata. Our approach integrates multidimensional distribution matching across geographic and industrial dimensions, employs domain-specific quality metrics, and enforces statistical moment constraints to ensure high-fidelity synthetic data that closely replicates key statistical properties and economic inference outcomes of the original dataset. We successfully generated a synthetic dataset for the 2007 Business Owner Survey and fully reproduced an empirical study published in *Small Business Economics*, thereby validating the method’s statistical validity and reproducibility. This work bridges a critical technical gap in secure commercial survey data dissemination and provides a scalable, regulatory-compliant framework for sharing sensitive microdata.

0 citationsRead paper

Uncertainty Quantification for Multi-level Models Using the Survey-Weighted Pseudo-Posterior

Oct 10, 2025

To address inadequate calibration of both local (e.g., group-level random effects) and global parameters in multilevel Bayesian models under complex survey designs, this paper proposes an improved survey-weighted pseudo-posterior framework with an automated post-processing pipeline—enabling, for the first time, consistent uncertainty calibration for both parameter types within hierarchical structures. The method integrates mixed-effects modeling, weighted likelihood construction, and posterior reweighting calibration, ensuring asymptotic consistency theoretically and seamless compatibility with mainstream Bayesian software computationally. Simulation studies and empirical analysis using the National Survey on Drug Use and Health (NSDUH) demonstrate that the proposed approach substantially improves estimation accuracy and reliability of uncertainty quantification, particularly reducing bias in random-effects inference. The corresponding algorithms are publicly available and integrated into the R package `csSampling`, offering a scalable, reproducible solution for hierarchical Bayesian analysis under complex sampling.

0 citationsRead paper