Developments in continuous speech dictation using the 1995 ARPA NAB news task

Authors:
J. L. Gauvain;L. Lamel;G. Adda;D. Matrouf
Affiliations:
Lab. d'Inf. pour la Mecanique et les Sci. de l'Ingenieur, CNRS, Orsay, France;-;-;-
Venue:
ICASSP '96 Proceedings of the Acoustics, Speech, and Signal Processing, 1996. on Conference Proceedings., 1996 IEEE International Conference - Volume 01
Year:
1996

Citing 0
Cited 4

Korean large vocabulary continuous speech recognition with morpheme-based recognition units

Speech Communication
Speech Recognition Issues for Dutch Spoken Document Retrieval

TSD '01 Proceedings of the 4th International Conference on Text, Speech and Dialogue
Robust Romanian language automatic speech recognizer based on multistyle training

WSEAS Transactions on Computer Research
Classroom lecture recognition

PROPOR'06 Proceedings of the 7th international conference on Computational Processing of the Portuguese Language

Quantified Score

Hi-index	0.00

Visualization

Abstract

We report on the LIMSI recognizer evaluated in the ARPA 1995 North American Business (NAB) news benchmark test. In contrast to previous evaluations, the new Hub 3 test aims at improving basic SI, CSR performance on unlimited-vocabulary read speech recorded under more varied acoustical conditions (background environmental noise and unknown microphones). The LIMSI recognizer is an HMM-based system with a Gaussian mixture. Decoding is carried out in multiple forward acoustic passes, where more refined acoustic and language models are used in successive passes and information is transmitted via word graphs. In order to deal with the varied acoustic conditions, channel compensation is performed iteratively, refining the noise estimates before the first three decoding passes. The final decoding pass is carried out with speaker-adapted models obtained via unsupervised adaptation using the MLLR method. On the Sennheiser microphone (average SNR 29 dB) a word error of 9.1% was obtained, which can be compared to 17.5% on the secondary microphone data (average SNR 15 dB) using the same recognition system.