Diphone subspace mixture trajectory models for HMM Complementation

  • Authors:
  • K. Reinhard;M. Niranjan

  • Affiliations:
  • Ericsson Eurolab Deutschland GmbH, Research Mobile Communications, Neumeyerstrasse 50, D-90411 Nuremberg, Germany and Engineering Department, Cambridge University, Trumpington Street, Cambridge CB ...;Department of Computer Science, Sheffield University, Portobello Street, Sheffield, S1 4DP, UK and Engineering Department, Cambridge University, Trumpington Street, Cambridge CB2 1PZ, UK

  • Venue:
  • Speech Communication
  • Year:
  • 2002

Quantified Score

Hi-index 0.00

Visualization

Abstract

This paper describes an extension of the previously reported attempt of capturing segmental transition information for speech recognition tasks [Speech Communication 27 (1) (1999) 19]. Representations in the subspace with multiple projected trajectories are discussed, employing EM-based methods to find optimal anchor points. Experimental work is carried out to illustrate that useful discriminant information is preserved in the subspace trajectories. These experiments include the development of "matched filters" to spot particular diphones in continuous speech, and the inclusion of diphone-based discriminant information into a phone-based HMM recognition framework to rerank multiple hypotheses. The difficulties in constructing the models due to the limited coverage of a sufficient amount of tokens within the phone balanced TIMIT database are discussed. The influence of the restricted diphone coverage on the rescoring results is reported. Improvements in phone recognition accuracy have been obtained on a speaker-by-speaker basis. Obtained improvements over baseline HMMs augmented with first-order derivatives suggest the importance of explicitly modelled between-phone information.