Cardiovascular disease (CVD) represents one of the leading causes of global death and hence the imperative of accurate and early risk prediction. This study compares the relative effectiveness of a set of ensemble machine learning classifiers, such as Random Forest, Gradient Boosting, Extra Trees, XGBoost, LightGBM, and a Voting Ensemble, to predict individuals into a high, intermediate and low risk groups regarding CVD. CAIR-CVD-2025 dataset went through rigorous preprocessing measures that included categorical encoding, imputation of missing data, feature standardization as well as class balance calibration using Synthetic Minority Over-Sampling Technique. Different metrics and tools have been used, e.g., macro-averaged precision, recall, F1 score, ROC-AUC, confusion matrices, overall accuracy, and computational run time. The obtained results support the applicability of machine learning, and, in particular, ensemble methodologies, to develop the field of CVD early risk prediction and make preventive healthcare interventions.
AI-Driven Ensemble Approaches for Early Prediction of Cardiovascular Disease Risk / Hasnain, S.I., Faris, M., Pepe, C., Ali, M.F., Zanoli, S.M.. - (2026), pp. 190-195. (27th International Carpathian Control Conference, ICCC 2026 La Contessa Castle Hotel, hun 2026) [10.1109/ICCC71363.2026.11593356].
AI-Driven Ensemble Approaches for Early Prediction of Cardiovascular Disease Risk
Pepe, C.;Ali, M. F.;Zanoli, S. M.
2026-01-01
Abstract
Cardiovascular disease (CVD) represents one of the leading causes of global death and hence the imperative of accurate and early risk prediction. This study compares the relative effectiveness of a set of ensemble machine learning classifiers, such as Random Forest, Gradient Boosting, Extra Trees, XGBoost, LightGBM, and a Voting Ensemble, to predict individuals into a high, intermediate and low risk groups regarding CVD. CAIR-CVD-2025 dataset went through rigorous preprocessing measures that included categorical encoding, imputation of missing data, feature standardization as well as class balance calibration using Synthetic Minority Over-Sampling Technique. Different metrics and tools have been used, e.g., macro-averaged precision, recall, F1 score, ROC-AUC, confusion matrices, overall accuracy, and computational run time. The obtained results support the applicability of machine learning, and, in particular, ensemble methodologies, to develop the field of CVD early risk prediction and make preventive healthcare interventions.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


