Employing second-order circular suprasegmental hidden Markov models to enhance speaker identification performance in shouted talking environments

Authors:
Ismail Shahin
Affiliations:
Electrical and Computer Engineering Department, University of Sharjah, Sharjah, United Arab Emirates
Venue:
EURASIP Journal on Audio, Speech, and Music Processing
Year:
2010

Citing 7
Cited 0

Speaker-dependent-feature extraction, recognition and processing techniques

Speech Communication - Special issue on speaker characterization in speech terminology
Fundamentals of speech recognition

Fundamentals of speech recognition
Hidden Markov Models for Speech Recognition

Hidden Markov Models for Speech Recognition
Emotions, speech and the ASR framework

Speech Communication - Special issue on speech and emotion
Improving speaker identification performance under the shouted talking condition using the second-order hidden Markov models

EURASIP Journal on Applied Signal Processing
Speaker identification in the shouted environment using Suprasegmental Hidden Markov Models

Signal Processing
Modulation spectral features for robust far-field speaker identification

IEEE Transactions on Audio, Speech, and Language Processing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Speaker identification performance is almost perfect in neutral talking environments. However, the performance is deteriorated significantly in shouted talking environments. This work is devoted to proposing, implementing, and evaluating new models called Second-Order Circular Suprasegmental Hidden Markov Models (CSPHMM2s) to alleviate the deteriorated performance in the shouted talking environments. These proposed models possess the characteristics of both Circular Suprasegmental Hidden Markov Models (CSPHMMs) and Second-Order Suprasegmental Hidden Markov Models (SPHMM2s). The results of this work show that CSPHMM2s outperform each of First-Order Left-to-Right Suprasegmental Hidden Markov Models (LTRSPHMM1s), Second-Order Left-to-Right Suprasegmental Hidden Markov Models (LTRSPHMM2s), and First-Order Circular Suprasegmental Hidden Markov Models (CSPHMM1s) in the shouted talking environments. In such talking environments and using our collected speech database, average speaker identification performance based on LTRSPHMM1s, LTRSPHMM2s, CSPHMM1s, and CSPHMM2s is 74.6%, 78.4%, 78.7%, and 83.4%, respectively. Speaker identification performance obtained based on CSPHMM2s is close to that obtained based on subjective assessment by human listeners.