Dysarthria severity assessment traditionally relies on time-intensive perceptual evaluation by Speech-Language Pathologists. This work presents a parameter-efficient approach for automatic severity classification on the HeyJay! dataset. A frozen wav2vec 2.0 backbone is paired with a lightweight MLP-based decoder incorporating learnable layer weighting, temporal attention pooling, and a compact embedding network. Under speaker-independent five-fold cross-validation, the system achieves 64.6% utterance accuracy, 61.2% weighted F1, and 80.2% speaker-level accuracy via majority voting. Learned weights reveal that upper Transformer layers (17--21) are most informative for severity discrimination. Error analysis shows that 80.3% of errors involve adjacent classes, with 62.1% concentrated at the Moderate--Severe boundary, matching the 62.5% inter-rater disagreement observed among expert SLPs at the same threshold. These results establish a first baseline on this dataset.
Dysarthria Severity Classification on the HeyJay! Dataset: A Parameter-Efficient Approach Using Self-Supervised Speech Representations / Lillini, D., Thebaud, T., Migliorelli, L., Dehak, N., Squartini, S., Velazquez, L.M.. - (2026), pp. 256-262. (Odyssey 2026 Lisbon, Portugal 23-26, June 2026) [10.21437/odyssey.2026-38].
Dysarthria Severity Classification on the HeyJay! Dataset: A Parameter-Efficient Approach Using Self-Supervised Speech Representations
Lillini, Davide
;Migliorelli, Lucia;Squartini, Stefano;
2026-01-01
Abstract
Dysarthria severity assessment traditionally relies on time-intensive perceptual evaluation by Speech-Language Pathologists. This work presents a parameter-efficient approach for automatic severity classification on the HeyJay! dataset. A frozen wav2vec 2.0 backbone is paired with a lightweight MLP-based decoder incorporating learnable layer weighting, temporal attention pooling, and a compact embedding network. Under speaker-independent five-fold cross-validation, the system achieves 64.6% utterance accuracy, 61.2% weighted F1, and 80.2% speaker-level accuracy via majority voting. Learned weights reveal that upper Transformer layers (17--21) are most informative for severity discrimination. Error analysis shows that 80.3% of errors involve adjacent classes, with 62.1% concentrated at the Moderate--Severe boundary, matching the 62.5% inter-rater disagreement observed among expert SLPs at the same threshold. These results establish a first baseline on this dataset.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


