Speaker Identification with Short Sequences of Speech Frames

Biagetti, Giorgio; Crippa, Paolo; Curzi, Alessandro; Orcioni, Simone; Turchetti, Claudio

doi:10.5220/0005191701780185

In biometric person identification systems, speaker identification plays a crucial role as the voice is the more natural signal to produce and the simplest to acquire. Mel frequency cepstral coefficients (MFCCs) have been widely adopted for decades in speech processing to capture the speech-specific characteristics with a reduced dimensionality. However, although their ability to de-correlate the vocal source and the vocal tract filter make them suitable for speech recognition, they show up some drawbacks in speaker recognition. This paper presents an experimental evaluation showing that reducing the dimension of features by using the discrete Karhunen-Loève transform (DKLT), guarantees better performance with respect to conventional MFCC features. In particular with short sequences of speech frames, that is with utterance duration of less than 1 s, the performance of truncated DKLT representation are always better than MFCC.

Speaker Identification with Short Sequences of Speech Frames / Biagetti, G., Crippa, P., Curzi, A., Orcioni, S., Turchetti, C.. - 2:(2015), pp. 178-185. (4th International Conference on Pattern Recognition Applications and Methods (ICPRAM 2015) Lisbona, Portogallo 10 - 12 Gennaio 2015) [10.5220/0005191701780185].