Vision Transformers (ViTs) have achieved remarkable performance across various computer vision tasks, yet their high computational cost remains a significant limitation in many real-world scenarios. As noted in prior studies, decreasing the number of tokens processed by the attention layers of a ViT directly reduces the required operations. Building on this idea, and drawing inspiration from signal processing, we reinterpret the token embeddings of a ViT layer as a signal, which allows us to apply the Discrete Wavelet Transform (DWT) to separate low- and high-frequency components. Guided by this insight, we present Token REduction via WAvelet decomposition (TREWA), a token-pruning strategy built upon DWT. For each image in a batch, TREWA selects a pruning level by comparing that image’s attention entropy with the batch one. It then applies the DWT to the token embeddings and forwards only the low-frequency coefficients (i.e., those capturing the image’s main semantic structure) to the next attention layer, discarding 50–75% of the tokens. We evaluate TREWA on four benchmark datasets in both pre-trained and training from scratch settings, comparing it against state-of-the-art pruning methods. Our results show a superior trade-off between accuracy and computational efficiency, validating the effectiveness of our frequency-domain token pruning strategy for accelerating ViTs.

Token Reduction in Vision Transformers via Discrete Wavelet Decomposition / Buratti, C., Marchetti, M., Parlapiano, F., Traini, D., Ursino, D., Virgili, L.. - (2026), pp. 583-597. (28th International Conference on Pattern Recognition, ICPR 2026 Lyon 17 - 22 August 2026) [10.1007/978-3-032-31654-7_39].

Token Reduction in Vision Transformers via Discrete Wavelet Decomposition

C. Buratti
Primo
;
M. Marchetti;F. Parlapiano;D. Traini
;
D. Ursino;L. Virgili
Ultimo
2026-01-01

Abstract

Vision Transformers (ViTs) have achieved remarkable performance across various computer vision tasks, yet their high computational cost remains a significant limitation in many real-world scenarios. As noted in prior studies, decreasing the number of tokens processed by the attention layers of a ViT directly reduces the required operations. Building on this idea, and drawing inspiration from signal processing, we reinterpret the token embeddings of a ViT layer as a signal, which allows us to apply the Discrete Wavelet Transform (DWT) to separate low- and high-frequency components. Guided by this insight, we present Token REduction via WAvelet decomposition (TREWA), a token-pruning strategy built upon DWT. For each image in a batch, TREWA selects a pruning level by comparing that image’s attention entropy with the batch one. It then applies the DWT to the token embeddings and forwards only the low-frequency coefficients (i.e., those capturing the image’s main semantic structure) to the next attention layer, discarding 50–75% of the tokens. We evaluate TREWA on four benchmark datasets in both pre-trained and training from scratch settings, comparing it against state-of-the-art pruning methods. Our results show a superior trade-off between accuracy and computational efficiency, validating the effectiveness of our frequency-domain token pruning strategy for accelerating ViTs.
2026
9783032316530
File in questo prodotto:
File Dimensione Formato  
Buratti_Token-Reduction-Vision-Transformers_2026.pdf

Solo gestori archivio

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza d'uso: Tutti i diritti riservati
Dimensione 388.41 kB
Formato Adobe PDF
388.41 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/354992
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact