Detecting people using mutually consistent poselet activations
European Conference on Computer Vision (ECCV), Springer, LNCS, Sept. 2010
Abstract: Bourdev and Malik (ICCV 09) introduced a new notion of parts, poselets, constructed to be tightly clustered both in the configuration
space of keypoints, as well as in the appearance space of image patches. In this paper we develop a new algorithm for detecting people
using poselets. Unlike that work which used 3D annotations of keypoints, we use only 2D annotations which are much easier for naive human annotators. The main algorithmic contribution is in how we use the pattern of poselet activations. Individual poselet activations are noisy, but considering the spatial context of each can provide vital disambiguating information, just as object detection can be improved by considering the detection scores of nearby objects in the scene. This can be done by training a two-layer feed-forward network with weights set using a max margin technique. The refined poselet activations are then clustered into mutually consistent hypotheses where consistency is based on empirically determined spatial keypoint distributions. Finally, bounding boxes are predicted for each person hypothesis and shape masks are aligned to edges in the image to provide a segmentation. To the best of our knowledge, the resulting system is the current best performer on the task of people detection and segmentation with an average precision of 47.8% and 40.5% respectively on PASCAL VOC 2009.
DownloadsImages and movies
BibTex reference
@InProceedings{Bro10d, author = "L. Bourdev and S. Maji and T. Brox and J. Malik", title = "Detecting people using mutually consistent poselet activations", booktitle = "European Conference on Computer Vision (ECCV)", series = "Lecture Notes in Computer Science", month = "Sept.", year = "2010", publisher = "Springer", url = "http://lmbweb.informatik.uni-freiburg.de/Publications/2010/Bro10d" }