Some aspects of ASR transcription based unsupervised speaker adaptation for HMM speech synthesis

Authors:
Bálint Tóth;Tibor Fegyó;Géza Németh
Affiliations:
Department of Telecommunications and Media Informatics, Budapest University of Technology and Economics;Department of Telecommunications and Media Informatics, Budapest University of Technology and Economics;Department of Telecommunications and Media Informatics, Budapest University of Technology and Economics
Venue:
TSD'10 Proceedings of the 13th international conference on Text, speech and dialogue
Year:
2010

Citing 4
Cited 1

Speech spectrum conversion based on speaker interpolation and multi-functional representation with weighting by radial basis function networks

Speech Communication - Special issue: voice conversion: state of the art and perspectives
Speech Synthesis with Various Emotional Expressions and Speaking Styles by Style Interpolation and Morphing

IEICE - Transactions on Information and Systems
Hybrid Voice Conversion of Unit Selection and Generation Using Prosody Dependent HMM

IEICE - Transactions on Information and Systems
Adaptation of pitch and spectrum for HMM-based speech synthesis using MLLR

ICASSP '01 Proceedings of the Acoustics, Speech, and Signal Processing, 200. on IEEE International Conference - Volume 02

Improvements of Hungarian hidden Markov model-based text-to-speech synthesis

Acta Cybernetica

Quantified Score

Hi-index	0.00

Visualization

Abstract

Statistical parametric synthesis offers numerous techniques to create new voices. Speaker adaptation is one of the most exciting ones. However, it still requires high quality audio data with low signal to noise ration and precise labeling. This paper presents an automatic speech recognition based unsupervised adaptation method for Hidden Markov Model (HMM) speech synthesis and its quality evaluation. The adaptation technique automatically controls the number of phone mismatches. The evaluation involves eight different HMM voices, including supervised and unsupervised speaker adaptation. The effects of segmentation and linguistic labeling errors in adaptation data are also investigated. The results show that unsupervised adaptation can contribute to speeding up the creation of new HMM voices with comparable quality to supervised adaptation.