🤖 AI Summary
Traditional survey instruments struggle to accurately capture the latent cognitive content embedded in occupational tasks, limiting the measurement of key variables such as exposure to artificial intelligence (AI). This study proposes leveraging large language models (LLMs) as a tool for latent variable measurement, using Claude Haiku 4.5 to generate cognitive dimension scores from 18,796 task descriptions in O*NET and constructing an augmented human capital index (AHC_o) to quantify occupation-level AI exposure. The research provides the first systematic validation of LLMs’ capacity to reliably measure latent constructs in labor economics, distinguishing between AI’s augmentation and substitution effects. Results show that AHC_o exhibits high correlation with existing AI exposure metrics (r ≤ 0.85), strong cross-model scoring reliability (Pearson r = 0.76), and a 25% reduction in measurement error using ORIV estimation relative to OLS, demonstrating significantly improved measurement precision.
📝 Abstract
This paper establishes the theoretical and practical foundations for using Large Language Models (LLMs) as measurement instruments for latent economic variables -- specifically variables that describe the cognitive content of occupational tasks at a level of granularity not achievable with existing survey instruments. I formalize four conditions under which LLM-generated scores constitute valid instruments: semantic exogeneity, construct relevance, monotonicity, and model invariance. I then apply this framework to the Augmented Human Capital Index (AHC_o), constructed from 18,796 O*NET task statements scored by Claude Haiku 4.5, and validated against six existing AI exposure indices. The index shows strong convergent validity (r = 0.85 with Eloundou GPT-gamma, r = 0.79 with Felten AIOE) and discriminant validity.
Principal component analysis confirms that AI-related occupational measures span two distinct dimensions -- augmentation and substitution. Inter-rater reliability across two LLM models (n = 3,666 paired scores) yields Pearson r = 0.76 and Krippendorff's alpha = 0.71. Prompt sensitivity analysis across four alternative framings shows that task-level rankings are robust. Obviously Related Instrumental Variables (ORIV) estimation recovers coefficients 25% larger than OLS, consistent with classical measurement error attenuation. The methodology generalizes beyond labor economics to any domain where semantic content must be quantified at scale.