Design-Based Inference under Deep Domain Stratification: Language of Instruction and Private-Institution Choice in India's NSS 71st Round

πŸ“… 2026-08-15
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of multidimensional subgroup inference in large-scale surveys and the absence of auditable frameworks by proposing an auditable design-based inference framework grounded in nested contribution ledgers. Leveraging first-order linearization and double-sample replication variance estimation, we construct granularity-stability profiles to evaluate subgroup reliability in stratified multistage surveys. Theoretically, we establish the unbiasedness and asymptotic efficiency of the proposed estimators. Empirically, the method’s efficacy is validated using data from an Indian education survey, demonstrating a reproducible workflow for fine-domain inference. Collectively, this work provides rigorous statistical safeguards and a foundational audit basis for analyzing complex survey data, ensuring both transparency and reliability in high-dimensional subgroup estimation.
πŸ“ Abstract
Large household surveys support precise national estimates but can become statistically fragile after repeated disaggregation by geography, sector, sex, age, and outcome category. This paper develops an auditable design-based framework for deciding how far such disaggregation can be taken in a stratified multistage survey. The framework is built around a nested contribution ledger that reconstructs each domain total through the first stage probability proportional to size expansion, the certainty-plus-random hamlet-group selection, and the second-stage household expansion. Nonlinear domain parameters are expressed as ratios of these totals and analyzed by first-order linearization. The two independent National Sample Survey subsamples then provide a natural replication variance estimator. A granularity-stability profile combines the resulting relative standard error with replicate support and concentration diagnostics, so that a detailed estimate is accompanied by evidence about whether the design can sustain it. Finite population unbiasedness of the total estimator, asymptotic validity of the ratio linearization, and unbiasedness of the two-subsample variance estimator for linearized totals are established. The method is illustrated with the 71st-round Social Consumption: Education survey, focusing on home language versus medium of instruction and reported reasons for preferring private educational institutions in India and Himachal Pradesh. The application preserves the substantive analysis in the original project while replacing ad hoc calculation with a reproducible inferential workflow.
Problem

Research questions and friction points this paper is trying to address.

Design-Based Inference
Domain Stratification
Disaggregation
Statistical Fragility
Granularity-Stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Design-Based Inference
Deep Domain Stratification
Granularity-Stability Profile
Nested Contribution Ledger
Replication Variance Estimator
πŸ’Ό Related Jobs
No related jobs found.