Speaker clustering for speech recognition using vocal tract parameters

Authors:
Masaki Naito;Li Deng;Yoshinori Sagisaka
Affiliations:
ATR Interpreting Telecommunications Research Laboratory, 2-2 Hikaridai, Seika-cho, Soraku-gun, Kyoto 619-0288, Japan;Department of Electrical and Computer Engineering, University of Waterloo, Waterloo, Ont., Canada;ATR Interpreting Telecommunications Research Laboratory, 2-2 Hikaridai, Seika-cho, Soraku-gun, Kyoto 619-0288, Japan
Venue:
Speech Communication
Year:
2002

Citing 0
Cited 2

Selecting Representative Speakers for a Speech Database on the Basis of Heterogeneous Similarity Criteria

Speaker Classification II
On enabling techniques for personal audio content management

MIR '08 Proceedings of the 1st ACM international conference on Multimedia information retrieval

Quantified Score

Hi-index	0.00

Visualization

Abstract

We propose speaker clustering methods for speech recognition based on vocal tract (VT) size related articulatory parameters associated with individual speakers. Two parameters characterizing gross VT dimensions are first derived from the formant frequencies of two vowels and are then used to cluster speakers. The resulting speaker clusters are significantly different from speaker clusters obtained by conventional acoustic criteria. Then phoneme recognition experiments are carried out by using speaker-clustered HMMs (SC-HMMs) trained for each cluster. The proposed method requires a small amount of speech data for speaker clustering and for selecting the most suitable SC-HMM for a target speaker, but gives higher recognition rates than conventional speaker clustering methods based on acoustic criteria.