Augmenting the availability of historical GDP per capita estimates through machine learning

📅 2024-09-16
🏛️ Proceedings of the National Academy of Sciences of the United States of America
📈 Citations: 2
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the longstanding gap in long-term per capita GDP estimates for hundreds of regions across Europe and North America over the past 700 years. Method: It innovatively employs large-scale biographical data—encoding birthplace, occupation, and social status—as proxy variables to train a supervised machine learning regression model. The approach integrates high-dimensional text feature engineering with cross-regional extrapolation techniques. Contribution/Results: The model achieves an out-of-sample R² of 0.90 and generates high-accuracy, regionally granular, multi-century per capita GDP series. Relative to existing datasets, it quadruples the volume of historical GDP estimates. Validated against multiple economic proxies—including urbanization rates, average stature, and subjective well-being—the estimates robustly replicate well-documented macrohistorical patterns, such as the North–South European economic reversal and the growth-enhancing role of Atlantic port cities. This work substantially expands both the empirical foundation and methodological toolkit for long-run macroeconomic analysis.

Technology Category

Application Category

📝 Abstract
Significance The scarcity of historical GDP per capita data limits our ability to explore questions of long-term economic development. Here, we introduce a machine learning method using detailed data on famous biographies to estimate the historical GDP per capita of hundreds of regions in Europe and North America. Our model generates accurate out-of-sample estimates (R2 = 90%) that quadruple the availability of historical GDP per capita data and correlate positively with proxies of economic output such as urbanization, body height, well-being, and church building activity. We use these estimates to reproduce the reversal of fortunes experienced by southern and northern Europe and the historical role played by Atlantic ports. These findings show that machine learning can effectively augment the historical availability of economic data.
Problem

Research questions and friction points this paper is trying to address.

Estimating historical GDP per capita using biographical data
Validating estimates with economic output proxies
Tracking regional economic shifts over 700 years
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine learning estimates GDP from biographies
Elastic net regression for feature selection
External validation with economic proxies
P
Philipp Koch
EcoAustria – Institute for Economic Research, Vienna, Austria; Center for Collective Learning, ANITI, IRIT, Université de Toulouse, Toulouse, France
V
Viktor Stojkoski
Faculty of Economics, University Ss. Cyril and Methodius, Skopje, North Macedonia; Center for Collective Learning, ANITI, IRIT, Université de Toulouse, Toulouse, France
C
César A Hidalgo
Toulouse School of Economics, Université de Toulouse, Toulouse, France; Center for Collective Learning, CIAS, Corvinus University, Budapest, Hungary; Center for Collective Learning, ANITI, IRIT, Université de Toulouse, Toulouse, France