Clustering country-level all-cause mortality data: a review

📅 2025-12-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of methodological syntheses on clustering applications for national all-cause mortality data. Leveraging the Human Mortality Database (HMD), it systematically applies k-means, hierarchical, and functional clustering to advance mortality forecasting, convergence analysis, and inequality assessment. Results confirm persistent east–west mortality divergence across Europe and demonstrate that cluster-based modeling substantially improves forecast accuracy over single-country models. The study’s primary contributions are threefold: (1) the first integrated review of data sources, clustering methodologies, and empirical findings in this domain; (2) identification of two critical research gaps—insufficient validation of clustering quality and underrepresentation of low-income countries in mortality datasets; and (3) provision of methodological foundations and practical guidance for cross-national mortality comparison and policy-relevant country grouping. (149 words)

Technology Category

Application Category

📝 Abstract
Mortality data are relevant to demography, public health, and actuarial science. Whilst clustering is increasingly used to explore patterns in such data, no study has reviewed its application to country-level all-cause mortality. This review therefore summarises recent work and addresses key questions: why clustering is used, which mortality data are analysed, which methods are most common, and what main findings emerge. To address these questions, we examine studies applying clustering to country-level all-cause mortality, focusing on mortality indices, data sources, and methodological choices, and we replicate some approaches using Human Mortality Database (HMD) data. Our analysis reveals that clustering is mainly motivated by forecasting and by studying convergence and inequality. Most studies use HMD data from developed countries and rely on k-means, hierarchical, or functional clustering. Main findings include a persistent East-West European division across applications, with clustering generally improving forecast accuracy over single-country models. Overall, this review highlights the methodological range in the literature, summarises clustering results, and identifies gaps, such as the limited evaluation of clustering quality and the underuse of data from countries outside the high-income world.
Problem

Research questions and friction points this paper is trying to address.

Reviews clustering applications to country-level all-cause mortality data.
Examines motivations, data sources, and methods used in mortality clustering.
Identifies research gaps like limited quality evaluation and regional data underuse.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clustering used for forecasting and inequality analysis
Common methods include k-means and hierarchical clustering
Replication with Human Mortality Database data validates approaches
🔎 Similar Papers
No similar papers found.
P
Pedro Menezes de Araujo
School of Mathematics and Statistics, University College Dublin
Isobel Claire Gormley
Isobel Claire Gormley
Professor, School of Mathematics and Statistics, University College Dublin
Computational StatisticsBayesian StatisticsStatistical MethodologyApplied Statistics
T
Thomas Brendan Murphy
School of Mathematics and Statistics, University College Dublin