Speech recognition using energy, MFCCs and rho parameters to classify syllables in the spanish language

Authors:
Sergio Suárez Guerra;José Luis Oropeza Rodríguez;Edgardo Manuel Felipe Riveron;Jesús Figueroa Nazuno
Affiliations:
Computing Research Center, National Polytechnic Institute, Mexico;Computing Research Center, National Polytechnic Institute, Mexico;Computing Research Center, National Polytechnic Institute, Mexico;Computing Research Center, National Polytechnic Institute, Mexico
Venue:
MICAI'06 Proceedings of the 5th Mexican international conference on Artificial Intelligence
Year:
2006

Citing 3
Cited 0

Fundamentals of speech recognition

Fundamentals of speech recognition
Integrating Syllable Boundary Information Into Speech Recognition

ICASSP '97 Proceedings of the 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP '97)-Volume 2 - Volume 2
Incorporating information from syllable-length time scales into automatic speech recognition

Incorporating information from syllable-length time scales into automatic speech recognition

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper presents an approach for the automatic speech re-cognition using syllabic units. Its segmentation is based on using the Short-Term Total Energy Function (STTEF) and the Energy Function of the High Frequency (ERO parameter) higher than 3,5 KHz of the speech signal. Training for the classification of the syllables is based on ten related Spanish language rules for syllable splitting. Recognition is based on a Continuous Density Hidden Markov Models and the bigram model language. The approach was tested using two voice corpus of natural speech, one constructed for researching in our laboratory (experimental) and the other one, the corpus Latino40 commonly used in speech researches. The use of ERO and MFCCs parameter increases speech recognition by 5.5% when compared with recognition using STTEF in discontinuous speech and improved more than 2% in continuous speech with three states. When the number of states is incremented to five, the recognition rate is improved proportionally to 98% for the discontinuous speech and to 81% for the continuous one.