New signal decomposition method based speech enhancement

Authors:
C. Tantibundhit;J. R. Boston;C. C. Li;J. D. Durrant;S. Shaiman;K. Kovacyk;A. El-Jaroudi
Affiliations:
Department of Electrical Engineering and Computer Engineering, Thammasat University, Rangsit Campus, Pathumthani 12120, Thailand;Department of Electrical and Computer Engineering, University of Pittsburgh, Pittsburgh, PA 15261, USA;Department of Electrical and Computer Engineering, University of Pittsburgh, Pittsburgh, PA 15261, USA;Department of Communication Science and Disorders, University of Pittsburgh, Pittsburgh, PA 15261, USA;Department of Communication Science and Disorders, University of Pittsburgh, Pittsburgh, PA 15261, USA;Department of Communication Science and Disorders, University of Pittsburgh, Pittsburgh, PA 15261, USA;Department of Electrical and Computer Engineering, University of Pittsburgh, Pittsburgh, PA 15261, USA
Venue:
Signal Processing
Year:
2007

Citing 7
Cited 3

Ten lectures on wavelets

Ten lectures on wavelets
The effect of cue-enhancement on the intelligibility of nonsense word and sentence materials presented in noise

Speech Communication
Hybrid representations for audiophonic signal encoding

Signal Processing - Image and Video Coding beyond Standards
A Greedy EM Algorithm for Gaussian Mixture Learning

Neural Processing Letters
Speech decomposition and enhancement

Speech decomposition and enhancement
Speech enhancement using transient speech components

Speech enhancement using transient speech components
Wavelet-based statistical signal processing using hidden Markovmodels

IEEE Transactions on Signal Processing

Denoising and recognition using hidden Markov models with observation distributions modeled by hidden Markov trees

Pattern Recognition
Minimum classification error learning for sequential data in the wavelet domain

Pattern Recognition
Joint time-frequency segmentation algorithm for transient speech decomposition and speech enhancement

IEEE Transactions on Audio, Speech, and Language Processing

Quantified Score

Hi-index	0.08

Visualization

Abstract

The auditory system, like the visual system, may be sensitive to abrupt stimulus changes, and the transient component in speech may be particularly critical to speech perception. If this component can be identified and selectively amplified, improved speech perception in background noise may be possible. This paper describes an algorithm to decompose speech into tonal, transient, and residual components. The modified discrete cosine transform (MDCT) was used to capture the tonal component and the wavelet transform was used to capture transient features. A hidden Markov chain (HMC) model and a hidden Markov tree (HMT) model were applied to capture statistical dependencies between the MDCT coefficients and between the wavelet coefficients, respectively. The transient component identified by the wavelet transform was selectively amplified and recombined with the original speech to generate modified speech, with energy adjusted to equal the energy of the original speech. The intelligibility of the original and modified speech was evaluated in eleven human subjects using the modified rhyme protocol. Word recognition rate results show that the modified speech can improve speech intelligibility at low SNR levels (8% at -15dB, 14% at -20dB, and 18% at -25dB) and has minimal effect on intelligibility at higher SNR levels.