A temporal frequency warped (TFW) 2D psychoacoustic filter for robust speech recognition system

Authors:
Peng Dai;Ing Yann Soon
Affiliations:
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798, Singapore;School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798, Singapore
Venue:
Speech Communication
Year:
2012

Citing 5
Cited 1

Should recognizers have ears?

Speech Communication - Special issue on robust speech recognition
Speech and Audio Signal Processing: Processing and Perception of Speech and Music

Speech and Audio Signal Processing: Processing and Perception of Speech and Music
Noise robust voice activity detection based on periodic to aperiodic component ratio

Speech Communication
2D psychoacoustic filtering for robust speech recognition

ICICS'09 Proceedings of the 7th international conference on Information, communications and signal processing
A temporal warped 2D psychoacoustic modeling for robust speech recognition system

Speech Communication

An improved model of masking effects for robust speech recognition system

Speech Communication

Quantified Score

Hi-index	0.00

Visualization

Abstract

In this paper, a novel hybrid feature extraction algorithm is proposed, which implements forward masking, lateral inhibition, and temporal integration with a simple 2D psychoacoustic filter. The proposed algorithm consists of two key parts, the 2D psychoacoustic filter and cepstral mean variance normalization (CMVN). Mathematical derivation is provided to show the correctness of the 2D psychoacoustic filter based on the characteristic functions of masking effects. The effectiveness of the proposed algorithm is tested on the AURORA2 database. Extensive comparison is made against lateral inhibition (LI), forward masking (FM), CMVN, RASTA filter, the ETSI standard advanced front-end feature extraction algorithm (AFE), and the temporal warped 2D psychoacoustic filter. Experimental results show significant improvements from the proposed algorithm, a relative improvement of nearly 46.78% over the baseline mel-frequency cepstral coefficients (MFCC) system in noisy conditions.