🤖 AI Summary
This study addresses the longstanding gap in long-term per capita GDP estimates for hundreds of regions across Europe and North America over the past 700 years. Method: It innovatively employs large-scale biographical data—encoding birthplace, occupation, and social status—as proxy variables to train a supervised machine learning regression model. The approach integrates high-dimensional text feature engineering with cross-regional extrapolation techniques. Contribution/Results: The model achieves an out-of-sample R² of 0.90 and generates high-accuracy, regionally granular, multi-century per capita GDP series. Relative to existing datasets, it quadruples the volume of historical GDP estimates. Validated against multiple economic proxies—including urbanization rates, average stature, and subjective well-being—the estimates robustly replicate well-documented macrohistorical patterns, such as the North–South European economic reversal and the growth-enhancing role of Atlantic port cities. This work substantially expands both the empirical foundation and methodological toolkit for long-run macroeconomic analysis.
📝 Abstract
Significance The scarcity of historical GDP per capita data limits our ability to explore questions of long-term economic development. Here, we introduce a machine learning method using detailed data on famous biographies to estimate the historical GDP per capita of hundreds of regions in Europe and North America. Our model generates accurate out-of-sample estimates (R2 = 90%) that quadruple the availability of historical GDP per capita data and correlate positively with proxies of economic output such as urbanization, body height, well-being, and church building activity. We use these estimates to reproduce the reversal of fortunes experienced by southern and northern Europe and the historical role played by Atlantic ports. These findings show that machine learning can effectively augment the historical availability of economic data.