Combining modality specific deep neural networks for emotion recognition in video

Authors:
Samira Ebrahimi Kahou;Christopher Pal;Xavier Bouthillier;Pierre Froumenty;Çaglar Gülçehre;Roland Memisevic;Pascal Vincent;Aaron Courville;Yoshua Bengio;Raul Chandias Ferrari;Mehdi Mirza;Sébastien Jean;Pierre-Luc Carrier;Yann Dauphin;Nicolas Boulanger-Lewandowski;Abhishek Aggarwal;Jeremie Zumer;Pascal Lamblin;Jean-Philippe Raymond;Guillaume Desjardins;Razvan Pascanu;David Warde-Farley;Atousa Torabi;Arjun Sharma;Emmanuel Bengio;Kishore Reddy Konda;Zhenzhou Wu
Affiliations:
Ecole Polytechnique, Montreal, Canada;Ecole Polytechnique, Montreal, Canada;University of Montreal, Montreal, Canada;Ecole Polytechnique, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;Univeristy of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;University of Montreal, Montreal, Canada;Goethe Universität Frankfurt, Frankfurt, Canada;McGill University, Montreal, Canada
Venue:
Proceedings of the 15th ACM on International conference on multimodal interaction
Year:
2013

Citing 7
Cited 0

Gabor-Based Kernel Partial-Least-Squares Discrimination Features for Face Recognition

Informatica
LIBSVM: A library for support vector machines

ACM Transactions on Intelligent Systems and Technology (TIST)
Random search for hyper-parameter optimization

The Journal of Machine Learning Research
Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis

CVPR '11 Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition
Face detection, pose estimation, and landmark localization in the wild

CVPR '12 Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Collecting Large, Richly Annotated Facial-Expression Databases from Movies

IEEE MultiMedia
Emotion recognition in the wild challenge 2013

Proceedings of the 15th ACM on International conference on multimodal interaction

Quantified Score

Hi-index	0.00

Visualization

Abstract

In this paper we present the techniques used for the University of Montréal's team submissions to the 2013 Emotion Recognition in the Wild Challenge. The challenge is to classify the emotions expressed by the primary human subject in short video clips extracted from feature length movies. This involves the analysis of video clips of acted scenes lasting approximately one-two seconds, including the audio track which may contain human voices as well as background music. Our approach combines multiple deep neural networks for different data modalities, including: (1) a deep convolutional neural network for the analysis of facial expressions within video frames; (2) a deep belief net to capture audio information; (3) a deep autoencoder to model the spatio-temporal information produced by the human actions depicted within the entire scene; and (4) a shallow network architecture focused on extracted features of the mouth of the primary human subject in the scene. We discuss each of these techniques, their performance characteristics and different strategies to aggregate their predictions. Our best single model was a convolutional neural network trained to predict emotions from static frames using two large data sets, the Toronto Face Database and our own set of faces images harvested from Google image search, followed by a per frame aggregation strategy that used the challenge training data. This yielded a test set accuracy of 35.58%. Using our best strategy for aggregating our top performing models into a single predictor we were able to produce an accuracy of 41.03% on the challenge test set. These compare favorably to the challenge baseline test set accuracy of 27.56%.