Pertinent Prosodic Features for Speaker Identification by Voice

Authors:
Halim Sayoud;Siham Ouamour
Affiliations:
USTHB University, Algeria;USTHB University, Algeria
Venue:
International Journal of Mobile Computing and Multimedia Communications
Year:
2010

Citing 9
Cited 0

Speaker identification and verification using Gaussian mixture speaker models

Speech Communication
Neural networks for discrimination and modelization of speakers

Speech Communication
Second-order statistical measures for text-independent speaker identification

Speech Communication
GSM speech coding and speaker recognition

ICASSP '00 Proceedings of the Acoustics, Speech, and Signal Processing, 2000. on IEEE International Conference - Volume 02
Behavior of a Bayesian adaptation method for incremental enrollment in speaker verification

ICASSP '00 Proceedings of the Acoustics, Speech, and Signal Processing, 2000. on IEEE International Conference - Volume 02
Automatic speech recognition and speech variability: A review

Speech Communication
Statistical modeling of heterogeneous features for speech processing tasks

Statistical modeling of heterogeneous features for speech processing tasks
Modeling Prosodic Features With Joint Factor Analysis for Speaker Verification

IEEE Transactions on Audio, Speech, and Language Processing
A Study of Interspeaker Variability in Speaker Verification

IEEE Transactions on Audio, Speech, and Language Processing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Most existing systems of speaker recognition use "state of the art" acoustic features. However, many times one can only recognize a speaker by his or her prosodic features, especially by the accent. For this reason, the authors investigate some pertinent prosodic features that can be associated with other classic acoustic features, in order to improve the recognition accuracy. The authors have developed a new prosodic model using a modified LVQ Learning Vector Quantization algorithm, which is called MLVQ Modified LVQ. This model is composed of three reduced prosodic features: the mean of the pitch, original duration, and low-frequency energy. Since these features are heterogeneous, a new optimized metric has been proposed that is called Optimized Distance for Heterogeneous Features ODHEF. Tests of speaker identification are done on Arabic corpus because the NIST evaluations showed that speaker verification scores depend on the spoken language and that some of the worst scores were got for the Arabic language. Experimental results show good performances of the new prosodic approach.