Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach

📅 2025-09-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
In high-dimensional covariate settings, conventional experimental stratification designs suffer from low efficiency, while traditional variable selection and weighting methods struggle to simultaneously ensure covariate balance and interpretability. Method: We propose a generative stratification framework that—novelly—integrates large language models (LLMs) into the pre-experimental design stage. Leveraging generative modeling, it automatically fuses heterogeneous covariate information to construct semantically informed strata, without requiring manual specification of variable importance or functional forms. Contribution/Results: Theoretically and empirically, our method reduces the variance of treatment effect estimation by 10–50% relative to simple randomization; further gains accrue when combined with classical stratification. Crucially, it extends the application paradigm of LLMs to causal inference design—offering a scalable, interpretable solution for high-dimensional experimental design.

Technology Category

Application Category

📝 Abstract
Pre-experiment stratification, or blocking, is a well-established technique for designing more efficient experiments and increasing the precision of the experimental estimates. However, when researchers have access to many covariates at the experiment design stage, they often face challenges in effectively selecting or weighting covariates when creating their strata. This paper proposes a Generative Stratification procedure that leverages Large Language Models (LLMs) to synthesize high-dimensional covariate data to improve experimental design. We demonstrate the value of this approach by applying it to a set of experiments and find that our method would have reduced the variance of the treatment effect estimate by 10%-50% compared to simple randomization in our empirical applications. When combined with other standard stratification methods, it can be used to further improve the efficiency. Our results demonstrate that LLM-based simulation is a practical and easy-to-implement way to improve experimental design in covariate-rich settings.
Problem

Research questions and friction points this paper is trying to address.

Improves experimental design by leveraging LLMs for covariate synthesis
Reduces treatment effect estimate variance through generative stratification
Addresses high-dimensional covariate challenges in pre-experiment blocking
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLMs synthesize high-dimensional covariate data
Generative stratification reduces treatment effect variance
Combines with standard methods for efficiency improvement
🔎 Similar Papers
No similar papers found.
G
George Gui
Columbia Business School
S
Seungwoo Kim
Columbia Business School