Cardiovascular Disease Prediction using Machine Learning: A Comparative Analysis

📅 2025-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses cardiovascular disease (CVD) risk prediction using a large-scale clinical dataset of 68,119 individuals. It systematically evaluates the impact of multiple risk factors—including age, blood pressure, cholesterol levels, smoking, and alcohol consumption—via statistical tests (t-tests, chi-square tests, and ANOVA) to identify significant associations. Notably, an unexpected negative correlation between smoking and alcohol use was detected, suggesting potential data bias and underscoring the need for careful preprocessing. Among several benchmark models, CatBoost achieved superior probabilistic calibration and discrimination: accuracy of 73.4%, Brier score of 0.1824, and expected calibration error (ECE) of only 0.0064—substantially outperforming logistic regression and other baselines. The work demonstrates CatBoost’s efficacy and reliability for CVD risk prediction while advancing interpretability and clinical trustworthiness through a hybrid statistical–machine learning framework that integrates rigorous hypothesis testing with high-performance modeling.

Technology Category

Application Category

📝 Abstract
Cardiovascular diseases (CVDs) are a main cause of mortality globally, accounting for 31% of all deaths. This study involves a cardiovascular disease (CVD) dataset comprising 68,119 records to explore the influence of numerical (age, height, weight, blood pressure, BMI) and categorical gender, cholesterol, glucose, smoking, alcohol, activity) factors on CVD occurrence. We have performed statistical analyses, including t-tests, Chi-square tests, and ANOVA, to identify strong associations between CVD and elderly people, hypertension, higher weight, and abnormal cholesterol levels, while physical activity (a protective factor). A logistic regression model highlights age, blood pressure, and cholesterol as primary risk factors, with unexpected negative associations for smoking and alcohol, suggesting potential data issues. Model performance comparisons reveal CatBoost as the top performer with an accuracy of 0.734 and an ECE of 0.0064 and excels in probabilistic prediction (Brier score = 0.1824). Data challenges, including outliers and skewed distributions, indicate a need for improved preprocessing to enhance predictive reliability.
Problem

Research questions and friction points this paper is trying to address.

Predicting cardiovascular disease risk using machine learning models
Identifying key factors like age, blood pressure, and cholesterol
Addressing data challenges to improve prediction accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine learning for CVD prediction
CatBoost model excels in accuracy
Statistical analysis identifies key risk factors
🔎 Similar Papers
No similar papers found.
R
Risshab Srinivas Ramesh
Department of Computer Science and Engineering Ramaiah Institute of Technology Bangalore, India
R
Roshani T S Udupa
Department of Computer Science and Engineering Ramaiah Institute of Technology Bangalore, India
M
Monisha J
Department of Computer Science and Engineering Ramaiah Institute of Technology Bangalore, India
K
Kushi K K S
Department of Computer Science and Engineering Ramaiah Institute of Technology Bangalore, India