Spatiotemporal Localization and Categorization of Human Actions in Unsegmented Image Sequences

Authors:
A. Oikonomopoulos;I. Patras;M. Pantic
Affiliations:
Dept. of Comput., Imperial Coll. London, London, UK;-;-
Venue:
IEEE Transactions on Image Processing
Year:
2011

Citing 0
Cited 5

Editors Choice Article: Structured learning of local features for human action classification and localization

Image and Vision Computing
Higher rank Support Tensor Machines for visual recognition

Pattern Recognition
Human behavior analysis in video surveillance: A Social Signal Processing perspective

Neurocomputing
An on-line, real-time learning method for detecting anomalies in videos using spatio-temporal compositions

Computer Vision and Image Understanding
Editor's Choice Article: Human activity recognition in videos using a single example

Image and Vision Computing

Quantified Score

Hi-index	0.01

Visualization

Abstract

In this paper we address the problem of localization and recognition of human activities in unsegmented image sequences. The main contribution of the proposed method is the use of an implicit representation of the spatiotemporal shape of the activity which relies on the spatiotemporal localization of characteristic ensembles of feature descriptors. Evidence for the spatiotemporal localization of the activity is accumulated in a probabilistic spatiotemporal voting scheme. The local nature of the proposed voting framework allows us to deal with multiple activities taking place in the same scene, as well as with activities in the presence of clutter and occlusion. We use boosting in order to select characteristic ensembles per class. This leads to a set of class specific codebooks where each codeword is an ensemble of features. During training, we store the spatial positions of the codeword ensembles with respect to a set of reference points, as well as their temporal positions with respect to the start and end of the action instance. During testing, each activated codeword ensemble casts votes concerning the spatiotemporal position and extend of the action, using the information that was stored during training. Mean Shift mode estimation in the voting space provides the most probable hypotheses concerning the localization of the subjects at each frame, as well as the extend of the activities depicted in the image sequences. We present classification and localization results for a number of publicly available datasets, and for a number of sequences where there is a significant amount of clutter and occlusion.