In this work we present LDAI, a mixed generative-predictive DNN framework aimed at repairing long-duration corrupted audio fragments in musical sequences. The proposed pipeline incorporates a convolutional encoder-decoder structure for waveform mapping and a low-level conditional WGAN for latent inpainting. Leveraging some recent insights of generative network modeling, we managed to reduce the effort required by the task, by transferring the inpainting process to a low dimensional level, thus ensuring robust and fast training convergence, and richer musical note generation. When dealing with gaps up to 1024 ms, we observed better performance than the baseline method of our previous study, namely bin2bin-MIDI, which began to struggle in repairing gaps exceeding 750 ms. The results, evaluated through perceptual metrics and estimated subjective scores, indicated improvements as follows: +0.89 points in PESQ, +3.1% in STOI, +4.1% in PLCMOS, and +0.87% in DNSMOS. Additionally, the Fréchet Audio Distance, aimed at assessing generative tasks, exhibited a remarkable 40% decrease.

Low-Dimensional Audio Inpainting with Mixed Generative-Predictive Model / Aironi, C., Gabrielli, L., Squartini, S.. - 459:(2026), pp. 413-423. [10.1007/978-981-95-4072-3_35]

Low-Dimensional Audio Inpainting with Mixed Generative-Predictive Model

Aironi C.
Primo
;
Gabrielli L.;Squartini S.
Ultimo
2026-01-01

Abstract

In this work we present LDAI, a mixed generative-predictive DNN framework aimed at repairing long-duration corrupted audio fragments in musical sequences. The proposed pipeline incorporates a convolutional encoder-decoder structure for waveform mapping and a low-level conditional WGAN for latent inpainting. Leveraging some recent insights of generative network modeling, we managed to reduce the effort required by the task, by transferring the inpainting process to a low dimensional level, thus ensuring robust and fast training convergence, and richer musical note generation. When dealing with gaps up to 1024 ms, we observed better performance than the baseline method of our previous study, namely bin2bin-MIDI, which began to struggle in repairing gaps exceeding 750 ms. The results, evaluated through perceptual metrics and estimated subjective scores, indicated improvements as follows: +0.89 points in PESQ, +3.1% in STOI, +4.1% in PLCMOS, and +0.87% in DNSMOS. Additionally, the Fréchet Audio Distance, aimed at assessing generative tasks, exhibited a remarkable 40% decrease.
2026
Neural Networks: Overview of Current Theories and Applications
9789819540716
9789819540723
File in questo prodotto:
File Dimensione Formato  
LDAI_paper_25_final.pdf

embargo fino al 01/04/2027

Tipologia: Documento in post-print (versione successiva alla peer review e accettata per la pubblicazione)
Licenza d'uso: Licenza specifica dell'editore
Dimensione 1.14 MB
Formato Adobe PDF
1.14 MB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/361553
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact