Due to their outstanding performance, Vision Transformers (ViTs) are becoming one of the most widely used architectures for computer vision tasks. However, the interpretation of the output returned by a ViT is still a challenging problem, and yet its solution is crucial for users to gain confidence in these new architectures. In this setting, some methods to address this issue rely only on attention scores; they cannot offer any causal guarantee because the highlighted patches do not necessarily drive the prediction. Other methods use mask-based perturbations; they are computationally expensive because they must repeatedly test many mask combinations of the image and rerun the model to measure confidence changes. To address these limitations, we present SImilarity-based GRAphs for Vision Transformer Explainability (SIGRATE), a hybrid framework that combines attention cues with graph-based mask generation to leverage the strengths of both approaches and overcome their weaknesses. For each attention layer of a ViT, SIGRATE extracts the embeddings of the image patches and constructs the corresponding similarity graph; then, it generates a set of binary masks starting from specific patches and performing walks in the similarity graph. Finally, it merges the masks of all attention layers into a heatmap using the coverage bias formula. We have tested SIGRATE on two ViT models (i.e., ViT-Base and DeiT-Base) and on two datasets, i.e., a subset of the ImageNet validation set and BloodMNIST. Our experiments show that SIGRATE provides promising performance in terms of Insertion, Deletion, and Pointing Game. Finally, we have performed a qualitative analysis demonstrating SIGRATE’s capabilities, as well as a hyperparameter analysis and an ablation study showing the impact of design choices on SIGRATE’s performance and efficiency.

Generating Masks from Similarity-based Graphs for Vision Transformer Explainability / Marchetti, M., Traini, D., Ursino, D., Virgili, L.. - In: NEURAL COMPUTING & APPLICATIONS. - ISSN 0941-0643. - 38:(2026). [10.1007/s00521-026-12392-6]

Generating Masks from Similarity-based Graphs for Vision Transformer Explainability

M. Marchetti
Co-primo
;
D. Traini
Co-primo
;
D. Ursino
Co-primo
;
L. Virgili
Ultimo
2026-01-01

Abstract

Due to their outstanding performance, Vision Transformers (ViTs) are becoming one of the most widely used architectures for computer vision tasks. However, the interpretation of the output returned by a ViT is still a challenging problem, and yet its solution is crucial for users to gain confidence in these new architectures. In this setting, some methods to address this issue rely only on attention scores; they cannot offer any causal guarantee because the highlighted patches do not necessarily drive the prediction. Other methods use mask-based perturbations; they are computationally expensive because they must repeatedly test many mask combinations of the image and rerun the model to measure confidence changes. To address these limitations, we present SImilarity-based GRAphs for Vision Transformer Explainability (SIGRATE), a hybrid framework that combines attention cues with graph-based mask generation to leverage the strengths of both approaches and overcome their weaknesses. For each attention layer of a ViT, SIGRATE extracts the embeddings of the image patches and constructs the corresponding similarity graph; then, it generates a set of binary masks starting from specific patches and performing walks in the similarity graph. Finally, it merges the masks of all attention layers into a heatmap using the coverage bias formula. We have tested SIGRATE on two ViT models (i.e., ViT-Base and DeiT-Base) and on two datasets, i.e., a subset of the ImageNet validation set and BloodMNIST. Our experiments show that SIGRATE provides promising performance in terms of Insertion, Deletion, and Pointing Game. Finally, we have performed a qualitative analysis demonstrating SIGRATE’s capabilities, as well as a hyperparameter analysis and an ablation study showing the impact of design choices on SIGRATE’s performance and efficiency.
2026
Computer vision; Patch embedding; Similarity graph; Vision transformer; Visual explainability
File in questo prodotto:
File Dimensione Formato  
Marchetti_Generating-masks-from-similarity_2026.pdf

accesso aperto

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza d'uso: Creative commons
Dimensione 5.93 MB
Formato Adobe PDF
5.93 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/360212
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact