Feature vs. Model Based Vocal Tract Length Normalization for a Speech Recognition-Based Interactive Toy

Authors:
Chun Keung Chau;Chak Shun Lai;Bertram Emil Shi
Affiliations:
-;-;-
Venue:
AMT '01 Proceedings of the 6th International Computer Science Conference on Active Media Technology
Year:
2001

Citing 5
Cited 0

Fundamentals of speech recognition

Fundamentals of speech recognition
Fast implementation methods for Viterbi-based word-spotting

ICASSP '96 Proceedings of the Acoustics, Speech, and Signal Processing, 1996. on Conference Proceedings., 1996 IEEE International Conference - Volume 01
Speaker normalization on conversational telephone speech

ICASSP '96 Proceedings of the Acoustics, Speech, and Signal Processing, 1996. on Conference Proceedings., 1996 IEEE International Conference - Volume 01
A parametric approach to vocal tract length normalization

ICASSP '96 Proceedings of the Acoustics, Speech, and Signal Processing, 1996. on Conference Proceedings., 1996 IEEE International Conference - Volume 01
A study of speech recognition for children and the elderly

ICASSP '96 Proceedings of the Acoustics, Speech, and Signal Processing, 1996. on Conference Proceedings., 1996 IEEE International Conference - Volume 01

Quantified Score

Hi-index	0.00

Visualization

Abstract

We describe an architecture for speech recognition based interactive toys and discuss the strategies we have adopted to deal with the requirements for the speech recognizer imposed by this application. In particular, we focus on the fact that speech recognizers used in interactive toys must deal with users whose age ranges from children to adults. The large variations in vocal tract length between children and adults can significantly degrade the performance of speech recognizers. We compare two approaches to vocal tract length normalization: feature-based VTLN and model-based VTLN. We describe why intuitively, one might expect that due to the coarser frequency information used by the model-based approach, that feature-based VTLN would outperform model-based VTLN. However, our results indicate that there is very little difference in performance between the two schemes.