Early detection of renal pathologies (cysts, calculi, neoplasms) is critical to allow timely treatment of patients. For the renal evaluation, Computed Tomography (CT) is still the method of preference, because it provides better spatial resolution and a consistent diagnostic accuracy. However, the interpretation of CT scans by clinicians is still tedious, time consuming, and vulnerable to inter-observer error. In order to overcome issues that arise from manual analyses of CT scan, the current research study presents a hybrid Deep Learning framework that combines MobileNetV2 and a Vision Transformer (ViT) to achieve automated multi-classification of renal lesions based on CT scans. MobileNetV2 is a lightweight convolutional backbone able to extract discriminative local spatial features in a depth-separable convolution and maintain computational efficiency in this architecture. The extracted features are subsequently processed by a ViT, which models longrange contextual relationships using multi-head self-attention. The model combines convolutional inductive bias with global relationship modeling to operate effectively on a publicly available Kidney lesions Kaggle dataset. The pre-processing and data augmentation were done on the four-categories of the datasets and training and testing were performed by a tailored hybrid model. The empirical results show high diagnostic accuracy of 96.78 %, a macro-precision of 94.96 %, a macro-recall of 96.74 %, and an overall macro-F1 of 95.76 %. The Area Under the ROC Curve (AUC) is 0.997. These findings demonstrate that the proposed hybrid framework can effectively classify renal CT images in clinical settings, offering both high accuracy and cost efficiency.

Hybrid MobileNetV2-Vision Transformer Framework for Automated Multi-Class Renal Lesion Classification in CT Imaging / Hasnain, S.I., Khowaja, R., Faris, M., Pepe, C., Ali, M.F., Zanoli, S.M.. - (2026), pp. 295-300. (34th Mediterranean Conference on Control and Automation, MED 2026 ita 2026) [10.1109/MED70602.2026.11598046].

Hybrid MobileNetV2-Vision Transformer Framework for Automated Multi-Class Renal Lesion Classification in CT Imaging

Pepe C.;Ali M. F.;Zanoli S. M.
2026-01-01

Abstract

Early detection of renal pathologies (cysts, calculi, neoplasms) is critical to allow timely treatment of patients. For the renal evaluation, Computed Tomography (CT) is still the method of preference, because it provides better spatial resolution and a consistent diagnostic accuracy. However, the interpretation of CT scans by clinicians is still tedious, time consuming, and vulnerable to inter-observer error. In order to overcome issues that arise from manual analyses of CT scan, the current research study presents a hybrid Deep Learning framework that combines MobileNetV2 and a Vision Transformer (ViT) to achieve automated multi-classification of renal lesions based on CT scans. MobileNetV2 is a lightweight convolutional backbone able to extract discriminative local spatial features in a depth-separable convolution and maintain computational efficiency in this architecture. The extracted features are subsequently processed by a ViT, which models longrange contextual relationships using multi-head self-attention. The model combines convolutional inductive bias with global relationship modeling to operate effectively on a publicly available Kidney lesions Kaggle dataset. The pre-processing and data augmentation were done on the four-categories of the datasets and training and testing were performed by a tailored hybrid model. The empirical results show high diagnostic accuracy of 96.78 %, a macro-precision of 94.96 %, a macro-recall of 96.74 %, and an overall macro-F1 of 95.76 %. The Area Under the ROC Curve (AUC) is 0.997. These findings demonstrate that the proposed hybrid framework can effectively classify renal CT images in clinical settings, offering both high accuracy and cost efficiency.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/362032
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact