Generation of robust phonetic set and decision tree for Mandarin using chi-square testing

Authors:
Yeou-Jiunn Chen;Chung-Hsien Wu;Yu-Hsien Chiu;Hsiang-Chuan Liao
Affiliations:
Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan, ROC;Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan, ROC;Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan, ROC;Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan, Taiwan, ROC
Venue:
Speech Communication
Year:
2002

Citing 5
Cited 3

Principles of multivariate analysis: a user's perspective

Principles of multivariate analysis: a user's perspective
Fundamentals of speech recognition

Fundamentals of speech recognition
Decision tree state tying based on penalized Bayesian information criterion

ICASSP '99 Proceedings of the Acoustics, Speech, and Signal Processing, 1999. on 1999 IEEE International Conference - Volume 01
Refining tree-based state clustering by means of formal concept analysis, balanced decision trees and automatically generated model-sets

ICASSP '99 Proceedings of the Acoustics, Speech, and Signal Processing, 1999. on 1999 IEEE International Conference - Volume 02
Irrelevant variability normalization in learning HMM state tying from data based on phonetic decision-tree

ICASSP '99 Proceedings of the Acoustics, Speech, and Signal Processing, 1999. on 1999 IEEE International Conference - Volume 02

Generation of Phonetic Units for Mixed-Language Speech Recognition Based on Acoustic and Contextual Analysis

IEEE Transactions on Computers
Stochastic vector mapping-based feature enhancement using prior-models and model adaptation for noisy speech recognition

Speech Communication
Proposing an interactive speaking improvement system for EFL learners

Expert Systems with Applications: An International Journal

Quantified Score

Hi-index	0.01

Visualization

Abstract

A phonetic representation of a language is used to describe the corresponding pronunciation and synthesize the acoustic model of any vocabulary. A phonetic representation with smaller phonetic units such as SAMPA-C for Mandarin Chinese and decision trees for parameter sharing are broadly applied to deal with the problem of large numbers of recognition units. However, the confusable phonetic representation in SAMPA-C generally degrades the recognition performance. In this paper, a statistical method based on chi-square testing is used to investigate the phonetic unit characteristics that are confusing and develop a more reliable phonetic set, named modified SAMPA-C. A corresponding question set for the modified SAMPA-C and a two-level splitting criterion are also proposed to effectively and efficiently construct the decision trees. Experiments using continuous Mandarin telephone speech recognition were conducted. Experimental results show that an encouraging improvement in recognition performance can be obtained. The proposed approaches represent a good compromise between the demands of accurate acoustic modeling and the limitations imposed by insufficient training data.