Prosody modelling for TTS systems using statistical methods

Authors:
Zdeněk Chaloupka;Petr Horák
Affiliations:
Institute of Photonics and Electronics, Academy of Sciences of the Czech Republic, Czech Republic;Institute of Photonics and Electronics, Academy of Sciences of the Czech Republic, Czech Republic
Venue:
COST'11 Proceedings of the 2011 international conference on Cognitive Behavioural Systems
Year:
2011

Citing 1
Cited 0

TectoMT: highly modular MT system with tectogrammatics used as transfer layer

StatMT '08 Proceedings of the Third Workshop on Statistical Machine Translation

Quantified Score

Hi-index	0.00

Visualization

Abstract

The main drawback of older methods of prosody modelling is the monotony of the output, which is perceived as uncomfortable by the users, especially when listening to longer passages. The present paper proposes a prosodic generator designed to increase the variability of synthesized speech in reading devices for the blind. The method used is based on text segmentation into several prosodic patterns by means of vector quantisation and the subsequent training of corresponding HMMs (Hidden Markov Models) on F0 parameters. The path through the model's states is then used to generate sentence prosody. We also tried to utilize morphological information in order to increase prosody naturalness. The evaluation of the quality of the proposed prosodic generators was carried out by means of listening tests.