Statistical multimodal integration for audio-visual speech processing

Authors:
S. Nakamura
Affiliations:
ATR Spoken Language Translation Res. Labs., Kyoto
Venue:
IEEE Transactions on Neural Networks
Year:
2002

Citing 0
Cited 4

Educational violin transcription by fusing multimedia streams

Proceedings of the international workshop on Educational multimedia and multimedia education
Temporal filtering of visual speech for audio-visual speech recognition in acoustically and visually challenging environments

Proceedings of the 9th international conference on Multimodal interfaces
Some experiments in audio-visual speech processing

NOLISP'07 Proceedings of the 2007 international conference on Advances in nonlinear speech processing
Review Article: Multimodal interaction: A review

Pattern Recognition Letters

Quantified Score

Hi-index	0.00

Visualization

Abstract

Sensory information is indispensable for living things. It is also important for living things to integrate multiple types of senses to understand their surroundings. In human communications, human beings must further integrate the multimodal senses of audition and vision to understand intention. In this paper, we describe speech related modalities since speech is the most important media to transmit human intention. To date, there have been a lot of studies concerning technologies in speech communications, but performance levels still have room for improvement. For instance, although speech recognition has achieved remarkable progress, the speech recognition performance still seriously degrades in acoustically adverse environments. On the other hand, perceptual research has proved the existence of the complementary integration of audio speech and visual face movements in human perception mechanisms. Such research has stimulated attempts to apply visual face information to speech recognition and synthesis. This paper introduces works on audio-visual speech recognition, speech to lip movement mapping for audio-visual speech synthesis, and audio-visual speech translation.