Cardiovascular disease (CVD) represents one of the leading causes of global death and hence the imperative of accurate and early risk prediction. This study compares the relative effectiveness of a set of ensemble machine learning classifiers, such as Random Forest, Gradient Boosting, Extra Trees, XGBoost, LightGBM, and a Voting Ensemble, to predict individuals into a high, intermediate and low risk groups regarding CVD. CAIR-CVD-2025 dataset went through rigorous preprocessing measures that included categorical encoding, imputation of missing data, feature standardization as well as class balance calibration using Synthetic Minority Over-Sampling Technique. Different metrics and tools have been used, e.g., macro-averaged precision, recall, F1 score, ROC-AUC, confusion matrices, overall accuracy, and computational run time. The obtained results support the applicability of machine learning, and, in particular, ensemble methodologies, to develop the field of CVD early risk prediction and make preventive healthcare interventions.

AI-Driven Ensemble Approaches for Early Prediction of Cardiovascular Disease Risk / Hasnain, S.I., Faris, M., Pepe, C., Ali, M.F., Zanoli, S.M.. - (2026), pp. 190-195. (27th International Carpathian Control Conference, ICCC 2026 La Contessa Castle Hotel, hun 2026) [10.1109/ICCC71363.2026.11593356].

AI-Driven Ensemble Approaches for Early Prediction of Cardiovascular Disease Risk

Pepe, C.;Ali, M. F.;Zanoli, S. M.
2026-01-01

Abstract

Cardiovascular disease (CVD) represents one of the leading causes of global death and hence the imperative of accurate and early risk prediction. This study compares the relative effectiveness of a set of ensemble machine learning classifiers, such as Random Forest, Gradient Boosting, Extra Trees, XGBoost, LightGBM, and a Voting Ensemble, to predict individuals into a high, intermediate and low risk groups regarding CVD. CAIR-CVD-2025 dataset went through rigorous preprocessing measures that included categorical encoding, imputation of missing data, feature standardization as well as class balance calibration using Synthetic Minority Over-Sampling Technique. Different metrics and tools have been used, e.g., macro-averaged precision, recall, F1 score, ROC-AUC, confusion matrices, overall accuracy, and computational run time. The obtained results support the applicability of machine learning, and, in particular, ensemble methodologies, to develop the field of CVD early risk prediction and make preventive healthcare interventions.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/362038
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact