HMM-Based Speech Synthesis Utilizing Glottal Inverse Filtering

Authors:
T. Raitio;A. Suni;J. Yamagishi;H. Pulakka;J. Nurminen;M. Vainio;P. Alku
Affiliations:
Dept. of Signal Process. & Acoust., Aalto Univ., Helsinki, Finland;-;-;-;-;-;-
Venue:
IEEE Transactions on Audio, Speech, and Language Processing
Year:
2011

Citing 0
Cited 7

Mixed source model and its adapted vocal tract filter estimate for voice transformation and synthesis

Speech Communication
Evaluation of glottal closure instant detection in a range of voice qualities

Speech Communication
Animated Lombard speech: Motion capture, facial animation and visual intelligibility of speech produced in adverse conditions

Computer Speech and Language
Synthesis and perception of breathy, normal, and Lombard speech in the presence of noise

Computer Speech and Language
Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction

IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)
Phonetic feature extraction for context-sensitive glottal source processing

Speech Communication
Pitch-Scaled Spectrum Based Excitation Model for HMM-based Speech Synthesis

Journal of Signal Processing Systems

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper describes an hidden Markov model (HMM)-based speech synthesizer that utilizes glottal inverse filtering for generating natural sounding synthetic speech. In the proposed method, speech is first decomposed into the glottal source signal and the model of the vocal tract filter through glottal inverse filtering, and thus parametrized into excitation and spectral features. The source and filter features are modeled individually in the framework of HMM and generated in the synthesis stage according to the text input. The glottal excitation is synthesized through interpolating and concatenating natural glottal flow pulses, and the excitation signal is further modified according to the spectrum of the desired voice source characteristics. Speech is synthesized by filtering the reconstructed source signal with the vocal tract filter. Experiments show that the proposed system is capable of generating natural sounding speech, and the quality is clearly better compared to two HMM-based speech synthesis systems based on widely used vocoder techniques.