Coping with out-of-vocabulary words: Open versus huge vocabulary asr

Authors:
Matteo Gerosa;Marcello Federico
Affiliations:
FBK-irst - Fondazione Bruno Kessler, Via Sommarive 18, 38100 Povo (TN), Italy;FBK-irst - Fondazione Bruno Kessler, Via Sommarive 18, 38100 Povo (TN), Italy
Venue:
ICASSP '09 Proceedings of the 2009 IEEE International Conference on Acoustics, Speech and Signal Processing
Year:
2009

Citing 0
Cited 3

Comparing SMT methods for automatic generation of pronunciation variants

IceTAL'10 Proceedings of the 7th international conference on Advances in natural language processing
Automatic generation of a pronunciation dictionary with rich variation coverage using SMT methods

CICLing'11 Proceedings of the 12th international conference on Computational linguistics and intelligent text processing - Volume Part II
Web-based tools and methods for rapid pronunciation dictionary creation

Speech Communication

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper investigates methods for coping with out-of-vocabulary words in a large vocabulary speech recognition task, namely the automatic transcription of Italian broadcast news. Two alternative ways for augmenting a 64K(thousand)-word recognition vocabulary and language model are compared: introducing extra words with their phonetic transcription up to 1.2M (million) words, or extending the language model with so-called graphones, i.e. subword units made of phone-character sequences. Graphones and phonetic transcriptions of words are automatically generated by adapting an off-the-shelf statistical machine translation toolkit. We found that the word-based and graphone-based extensions allow both for better recognition performance, with the former performing significantly better than the latter. In addition, the word-based extension approach shows interesting potential even under conditions of little supervision. In fact, by training the grapheme to phoneme translation system with only 2K manually verified transcriptions, the final word error rate increases by just 3% relative, with respect to starting from a lexicon of 64K words.