HMM-based reconstruction of unreliable spectrographic data for noise robust speech recognition

Authors:
Bengt J. Borgström;Abeer Alwan
Affiliations:
Department of Electrical Engineering, University of California, Los Angeles, CA;Department of Electrical Engineering, University of California, Los Angeles, CA
Venue:
IEEE Transactions on Audio, Speech, and Language Processing
Year:
2010

Citing 11
Cited 0

Fundamentals of digital image processing

Fundamentals of digital image processing
Fundamentals of speech recognition

Fundamentals of speech recognition
Cepstral domain segmental feature vector normalization for noise robust speech recognition

Speech Communication - Special issue on robust speech recognition
Robust automatic speech recognition with missing and unreliable acoustic data

Speech Communication
Clustering Algorithms

Clustering Algorithms
Acoustical and Environmental Robustness in Automatic Speech Recognition

Acoustical and Environmental Robustness in Automatic Speech Recognition
Missing Data Techniques for Robust Speech Recognition

ICASSP '97 Proceedings of the 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP '97)-Volume 2 - Volume 2
Reconstruction of incomplete spectrograms for robust speech recognition

Reconstruction of incomplete spectrograms for robust speech recognition
Digital Speech Transmission: Enhancement, Coding And Error Concealment

Digital Speech Transmission: Enhancement, Coding And Error Concealment
Improved A Posteriori Speech Presence Probability Estimation Based on a Likelihood Ratio With Fixed Priors

IEEE Transactions on Audio, Speech, and Language Processing
Hidden Markov model-based packet loss concealment for voice over IP

IEEE Transactions on Audio, Speech, and Language Processing

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper presents a framework for efficient HMM-based estimation of unreliable spectrographic speech data. It discusses the role of hidden Markov models (HMMs) during minimum mean-square error (MMSE) spectral reconstruction. We develop novel HMM-based reconstruction algorithms which exploit intra-channel (across-time) correlation and/or inter-channel (across-frequency) correlation. For the sake of computational efficiency, this paper utilizes approximations to HMM-based decoding methods by developing models constructed from lower resolution quantizers. State configurations for lower resolution models are obtained through a tree-structured mapping of quantizer centroids, and model parameters are adapted accordingly. HMM downsampling avoids expensive retraining of models, and eliminates unnecessary memory requirements. Explicit general formulae are presented for the adaptation of steady-state and transitional statistics. Adaptation of observation statistics are derived from stochastic models of noise spectral magnitude estimation accuracies. The proposed estimation methods are applied in combination with oracle masks, which provide an upper performance bound, as well as masks derived from speech presence probability, which represent a more realistic scenario. Both methods are shown to boost noise robust recognition accuracies significantly relative to the Mel-frequency cepstral coefficient (MFCC) baseline system. Furthermore, HMM downsampling greatly reduces the complexity of the HMM-based reconstruction method while negligibly affecting results.