EDBT 2026 Demo / reviewers in the wild / expert
Mykhaylo Andriluka
dblp:54/4495
· DBLP profile ↗
37ranked-venue papers
9as first author
4since 2021 · last 2024
0000-0002-3109-042XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 8 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
32 papers |
Face, body and person analysis · 29% Video understanding and tracking · 23% 3D vision · 21% | |
| Computer graphics and multimedia
2 papers |
Computer animation and physical simulation · 100% |
Topics — the 30 heaviest of 59, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
human pose estimation |
2.7 | 12 | 2022 | Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular Video · CVPR 2022 PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018 MARCOnI - ConvNet-Based MARker-Less Motion Capture in Outdoor and Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking |
0.8 | 4 | 2017 | Multiple People Tracking by Lifted Multicut and Person Re-identification · CVPR 2017 ArtTrack: Articulated Multi-Person Tracking in the Wild · CVPR 2017 Detection and Tracking of Occluded People · Int. J. Comput. Vis. 2014 |
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation |
0.8 | 1 | 2024 | Learned Neural Physics Simulation for Articulated 3D Human Pose Reconstruction · ECCV (84) 2024 |
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation |
0.7 | 3 | 2016 | DeeperCut: A Deeper, Stronger, and Faster Multi-person Pose Estimation Model · ECCV (6) 2016 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation · CVPR 2016 3D Pictorial Structures for Multiple Human Pose Estimation · CVPR 2014 |
Machine learning › Optimization for machine learning
learned optimizer |
0.7 | 1 | 2023 | Transformer-Based Learned Optimization · CVPR 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.7 | 1 | 2023 | Transformer-Based Learned Optimization · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection |
0.7 | 5 | 2016 | End-to-End People Detection in Crowded Scenes · CVPR 2016 Learning People Detectors for Tracking in Crowded Scenes · ICCV 2013 Learning people detection models from few training samples · CVPR 2011 |
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation |
0.6 | 5 | 2017 | ArtTrack: Articulated Multi-Person Tracking in the Wild · CVPR 2017 Strong Appearance and Expressive Spatial Models for Human Pose Estimation · ICCV 2013 People-tracking-by-detection and people-detection-by-tracking · CVPR 2008 |
Computer vision › 3D vision › 3d motion analysis
human motion reconstruction |
0.6 | 1 | 2022 | Differentiable Dynamics for Articulated 3d Human Motion Reconstruction · CVPR 2022 |
Computer vision › 3D vision › 3d motion analysis
physics-based motion reconstruction |
0.6 | 1 | 2022 | Differentiable Dynamics for Articulated 3d Human Motion Reconstruction · CVPR 2022 |
Computer vision › 3D vision › pose estimation › model-based pose estimation
physics-based pose estimation |
0.6 | 1 | 2022 | Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular Video · CVPR 2022 |
Robotics › Motion planning and robot control
trajectory optimization |
0.6 | 1 | 2022 | Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular Video · CVPR 2022 |
Computer animation and physical simulation
differentiable simulation |
0.6 | 1 | 2022 | Differentiable Dynamics for Articulated 3d Human Motion Reconstruction · CVPR 2022 |
Computer vision › Video understanding and tracking
activity recognition |
0.5 | 3 | 2016 | Recognizing Fine-Grained and Composite Activities Using Hand-Centric Features and Script Data · Int. J. Comput. Vis. 2016 Script Data for Attribute-Based Recognition of Composite Activities · ECCV (1) 2012 A database for fine grained activity detection of cooking activities · CVPR 2012 |
Computer vision › 3D vision › motion capture › human motion capture
markerless motion capture |
0.5 | 2 | 2017 | MARCOnI - ConvNet-Based MARker-Less Motion Capture in Outdoor and Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2017 Efficient ConvNet-based marker-less motion capture in general scenes with a low number of cameras · CVPR 2015 |
Computer vision › 3D vision › pose estimation
multi-view pose estimation |
0.5 | 2 | 2017 | MARCOnI - ConvNet-Based MARker-Less Motion Capture in Outdoor and Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2017 Efficient ConvNet-based marker-less motion capture in general scenes with a low number of cameras · CVPR 2015 |
Computer vision › Image recognition and object detection
image annotation |
0.4 | 1 | 2020 | Panoptic Image Annotation with a Collaborative Assistant · ACM Multimedia 2020 |
Computer vision › Segmentation and scene understanding
panoptic segmentation |
0.4 | 1 | 2020 | Panoptic Image Annotation with a Collaborative Assistant · ACM Multimedia 2020 |
Computer vision › Face, body and person analysis › human pose estimation
pictorial structures |
0.4 | 3 | 2013 | Poselet Conditioned Pictorial Structures · CVPR 2013 Discriminative Appearance Models for Pictorial Structures · Int. J. Comput. Vis. 2012 Pictorial structures revisited: People detection and articulated pose estimation · CVPR 2009 |
Computer vision › Video understanding and tracking › activity recognition
complex activity recognition |
0.4 | 2 | 2016 | Recognizing Fine-Grained and Composite Activities Using Hand-Centric Features and Script Data · Int. J. Comput. Vis. 2016 Script Data for Attribute-Based Recognition of Composite Activities · ECCV (1) 2012 |
Computer vision › Video understanding and tracking
action recognition |
0.3 | 1 | 2018 | Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos · Int. J. Comput. Vis. 2018 |
Computer vision › Video understanding and tracking › object tracking
articulated object tracking |
0.3 | 1 | 2018 | PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018 |
Computer vision › Segmentation and scene understanding
interactive segmentation |
0.3 | 1 | 2018 | Fluid Annotation: A Human-Machine Collaboration Interface for Full Image Annotation · ACM Multimedia 2018 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person pose tracking |
0.3 | 1 | 2018 | PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018 |
Computer vision › 3D vision
3d human pose estimation |
0.3 | 2 | 2014 | 3D Pictorial Structures for Multiple Human Pose Estimation · CVPR 2014 Monocular 3D pose estimation and tracking by detection · CVPR 2010 |
Computer vision › 3D vision
3d reconstruction |
0.3 | 1 | 2017 | MARCOnI - ConvNet-Based MARker-Less Motion Capture in Outdoor and Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Computer vision › Video understanding and tracking › motion tracking
articulated motion tracking |
0.3 | 1 | 2017 | MARCOnI - ConvNet-Based MARker-Less Motion Capture in Outdoor and Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Computer vision › Video understanding and tracking › multi-object tracking
lifted multicut |
0.3 | 1 | 2017 | Multiple People Tracking by Lifted Multicut and Person Re-identification · CVPR 2017 |
Machine learning › Generative modeling
synthetic training data |
0.3 | 2 | 2012 | Articulated people detection and pose estimation: Reshaping the future · CVPR 2012 Learning people detection models from few training samples · CVPR 2011 |
Computer vision › Image recognition and object detection › object detection › multi-object detection
crowded scene detection |
0.2 | 1 | 2016 | End-to-End People Detection in Crowded Scenes · CVPR 2016 |
Methods — techniques the papers use, named apart from their topics
trajectory optimization · 1.7neural physics simulation · 1.5differentiable physics simulation · 1.1benchmark dataset · 1.0transformer · 0.7BFGS · 0.7physics engine · 0.6convolutional neural network · 0.6generative model-based tracking · 0.5discriminative joint detection · 0.5panoptic segmentation · 0.4collaborative assistant · 0.4neural network · 0.3machine learning · 0.3evaluation server · 0.3minimum cost subgraph multicut · 0.2long-range re-identification · 0.2activity taxonomy · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learned Neural Physics Simulation for Articulated 3D Human Pose Reconstruction
Mykhaylo Andriluka, Baruch Tabanpour, C. Daniel Freeman, Cristian Sminchisescu |
ECCV (84) | 1 |
| 2023 | Transformer-Based Learned OptimizationabstractWe propose a new approach to learned optimization where we represent the computation of an optimizer's update step using a neural network. The parameters of the optimizer are then learned by training on a set of optimization tasks with the objective to perform minimization efficiently. Our innovation is a new neural network architecture, Optimus, for the learned optimizer inspired by the classic BFGS algorithm. As in BFGS, we estimate a preconditioning matrix as a sum of rank-one updates but use a Transformerbased neural network to predict these updates jointly with the step length and direction. In contrast to several recent learned optimization-based approaches [24, 27], our formulation allows for conditioning across the dimensions of the parameter space of the target problem while remaining applicable to optimization tasks of variable dimensionality without retraining. We demonstrate the advantages of our approach on a benchmark composed of objective functions traditionally used for the evaluation of optimization algorithms, as well as on the real world-task of physics-based visual reconstruction of articulated 3d human motion. Erik Gärtner, Luke Metz, Mykhaylo Andriluka, C. Daniel Freeman, Cristian Sminchisescu |
CVPR | 3 |
| 2022 | Differentiable Dynamics for Articulated 3d Human Motion ReconstructionabstractWe introduce DiffPhy, a differentiable physics-based model for articulated 3d human motion reconstruction from video. Applications of physics-based reasoning in human motion analysis have so far been limited, both by the complexity of constructing adequate physical models of articulated human motion, and by the formidable challenges of performing stable and efficient inference with physics in the loop. We jointly address such modeling and inference challenges by proposing an approach that combines a physically plausible body representation with anatomical joint limits, a differentiable physics simulator, and optimization techniques that ensure good performance and robustness to suboptimal local optima. In contrast to several recent methods [39], [42], [55], our approach readily supports full-body contact including interactions with objects in the scene. Most importantly, our model connects end-to-end with images, thus supporting direct gradient-based physics optimization by means of image-based loss functions. We validate the model by demonstrating that it can accurately reconstruct physically plausible 3d human motion from monocular video, both on public benchmarks with available 3d ground-truth, and on videos from the internet. Erik Gärtner, Mykhaylo Andriluka, Erwin Coumans, Cristian Sminchisescu |
CVPR | 2 |
| 2022 | Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular VideoabstractWe focus on the task of estimating a physically plausi-ble articulated human motion from monocular video. Ex-isting approaches that do not consider physics often pro-duce temporally inconsistent output with motion artifacts, while state-of-the-art physics-based approaches have either been shown to work only in controlled laboratory conditions or consider simplified body-ground contact limited to feet. This paper explores how these shortcomings can be addressed by directly incorporating a fully-featured physics engine into the pose estimation process. Given an uncon-trolled, real-world scene as input, our approach estimates the ground-plane location and the dimensions of the physi-cal body model. It then recovers the physical motion by per-forming trajectory optimization. The advantage of our for-mulation is that it readily generalizes to a variety of scenes that might have diverse ground properties and supports any form of self-contact and contact between the articu-lated body and scene geometry. We show that our approach achieves competitive results with respect to existing physics-based methods on the Human3.6M benchmark [13], while being directly applicable without re-training to more complex dynamic motions from the AIST benchmark [36] and to uncontrolled internet videos. Erik Gärtner, Mykhaylo Andriluka, Cristian Sminchisescu |
CVPR | 2 |
| 2020 | Panoptic Image Annotation with a Collaborative AssistantabstractThis paper aims to reduce the time to annotate images for panoptic segmentation, which requires annotating segmentation masks and class labels for all object instances and stuff regions. We formulate our approach as a collaborative process between an annotator and an automated assistant who take turns to jointly annotate an image using a predefined pool of segments. Actions performed by the annotator serve as a strong contextual signal. The assistant intelligently reacts to this signal by annotating other parts of the image on its own, which reduces the amount of work required by the annotator. We perform thorough experiments on the COCO panoptic dataset, both in simulation and with human annotators. These demonstrate that our approach is significantly faster than the recent machine-assisted interface of [Andriluka 18 ACMMM], and $2.4\times$ to $5\times$ faster than manual polygon drawing. Finally, we show on ADE20k that our method can be used to efficiently annotate new datasets, bootstrapping from a very small amount of annotated data. Jasper R. R. Uijlings, Mykhaylo Andriluka, Vittorio Ferrari |
ACM Multimedia | 2 |
| 2018 | PoseTrack: A Benchmark for Human Pose Estimation and TrackingabstractExisting systems for video-based pose estimation and tracking struggle to perform well on realistic videos with multiple people and often fail to output body-pose trajectories consistent over time. To address this shortcoming this paper introduces PoseTrack which is a new large-scale benchmark for video-based human pose estimation and articulated tracking. Our new benchmark encompasses three tasks focusing on i) single-frame multi-person pose estimation, ii) multi-person pose estimation in videos, and iii) multi-person articulated tracking. To establish the benchmark, we collect, annotate and release a new dataset that features videos with multiple people labeled with person tracks and articulated pose. A public centralized evaluation server is provided to allow the research community to evaluate on a held-out test set. Furthermore, we conduct an extensive experimental study on recent approaches to articulated pose tracking and provide analysis of the strengths and weaknesses of the state of the art. We envision that the proposed benchmark will stimulate productive research both by providing a large and representative training dataset as well as providing a platform to objectively evaluate and compare the proposed methods. The benchmark is freely accessible at https://posetrack.net/. Mykhaylo Andriluka, Umar Iqbal 0001, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, Bernt Schiele |
CVPR | 1 |
| 2018 | Fluid Annotation: A Human-Machine Collaboration Interface for Full Image AnnotationabstractWe introduce Fluid Annotation, an intuitive human-machine collaboration interface for annotating the class label and outline of every object and background region in an image. Fluid annotation is based on three principles:(I) Strong Machine-Learning aid. We start from the output of a strong neural network model, which the annotator can edit by correcting the labels of existing regions, adding new regions to cover missing objects, and removing incorrect regions.The edit operations are also assisted by the model.(II) Full image annotation in a single pass. As opposed to performing a series of small annotation tasks in isolation [51,68], we propose a unified interface for full image annotation in a single pass.(III) Empower the annotator.We empower the annotator to choose what to annotate and in which order. This enables concentrating on what the ma-chine does not already know, i.e. putting human effort only on the errors it made. This helps using the annotation budget effectively. Mykhaylo Andriluka, Jasper R. R. Uijlings, Vittorio Ferrari |
ACM Multimedia | 1 |
| 2018 | Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
Serena Yeung-Levy, Olga Russakovsky, Ning Jin 0002, Mykhaylo Andriluka, Greg Mori, Li Fei-Fei 0001 |
Int. J. Comput. Vis. | 4 |
| 2017 | ArtTrack: Articulated Multi-Person Tracking in the WildabstractIn this paper we propose an approach for articulated tracking of multiple people in unconstrained videos. Our starting point is a model that resembles existing architectures for single-frame pose estimation but is substantially faster. We achieve this in two ways: (1) by simplifying and sparsifying the body-part relationship graph and leveraging recent methods for faster inference, and (2) by offloading a substantial share of computation onto a feed-forward convolutional architecture that is able to detect and associate body joints of the same person even in clutter. We use this model to generate proposals for body joint locations and formulate articulated tracking as spatio-temporal grouping of such proposals. This allows to jointly solve the association problem for all people in the scene by propagating evidence from strong detections through time and enforcing constraints that each proposal can be assigned to one person only. We report results on a public MPII Human Pose benchmark and on a new MPII Video Pose dataset of image sequences with multiple people. We demonstrate that our model achieves state-of-the-art results while using only a fraction of time and is able to leverage temporal information to improve state-of-the-art for crowded scenes. Eldar Insafutdinov, Mykhaylo Andriluka, Leonid Pishchulin, Siyu Tang 0001, Evgeny Levinkov, Bjoern Andres, Bernt Schiele |
CVPR | 2 |
| 2017 | Multiple People Tracking by Lifted Multicut and Person Re-identificationabstractTracking multiple persons in a monocular video of a crowded scene is a challenging task. Humans can master it even if they loose track of a person locally by re-identifying the same person based on their appearance. Care must be taken across long distances, as similar-looking persons need not be identical. In this work, we propose a novel graph-based formulation that links and clusters person hypotheses over time by solving an instance of a minimum cost lifted multicut problem. Our model generalizes previous works by introducing a mechanism for adding long-range attractive connections between nodes in the graph without modifying the original set of feasible solutions. This allows us to reward tracks that assign detections of similar appearance to the same person in a way that does not introduce implausible solutions. To effectively match hypotheses over longer temporal gaps we develop new deep architectures for re-identification of people. They combine holistic representations extracted with deep networks and body pose layout obtained with a state-of-the-art pose estimation model. We demonstrate the effectiveness of our formulation by reporting a new state-of-the-art for the MOT16 benchmark. The code and pre-trained models are publicly available. Siyu Tang 0001, Mykhaylo Andriluka, Bjoern Andres, Bernt Schiele |
CVPR | 2 |
| 2017 | MARCOnI - ConvNet-Based MARker-Less Motion Capture in Outdoor and Indoor ScenesabstractMarker-less motion capture has seen great progress, but most state-of-the-art approaches fail to reliably track articulated human body motion with a very low number of cameras, let alone when applied in outdoor scenes with general background. In this paper, we propose a method for accurate marker-less capture of articulated skeleton motion of several subjects in general scenes, indoors and outdoors, even from input filmed with as few as two cameras. The new algorithm combines the strengths of a discriminative image-based joint detection method with a model-based generative motion tracking algorithm through an unified pose optimization energy. The discriminative part-based pose detection method is implemented using Convolutional Networks (ConvNet) and estimates unary potentials for each joint of a kinematic skeleton model. These unary potentials serve as the basis of a probabilistic extraction of pose constraints for tracking by using weighted sampling from a pose posterior that is guided by the model. In the final energy, we combine these constraints with an appearance-based model-to-image similarity term. Poses can be computed very efficiently using iterative local optimization, since joint detection with a trained ConvNet is fast, and since our formulation yields a combined pose estimation energy with analytic derivatives. In combination, this enables to track full articulated joint angles at state-of-the-art accuracy and temporal stability with a very low number of cameras. Our method is efficient and lends itself to implementation on parallel computing hardware, such as GPUs. We test our method extensively and show its advantages over related work on many indoor and outdoor data sets captured by ourselves, as well as data sets made available to the community by other research labs. The availability of good evaluation data sets is paramount for scientific progress, and many existing test data sets focus on controlled indoor settings, do not feature much variety in the scenes, and often lack a large corpus of data with ground truth annotation. We therefore further contribute with a new extensive test data set called MPI-MARCOnI for indoor and outdoor marker-less motion capture that features 12 scenes of varying complexity and varying camera count, and that features ground truth reference data from different modalities, ranging from manual joint annotations to marker-based motion capture results. Our new method is tested on these data, and the data set will be made available to the community. Ahmed Elhayek, Edilson de Aguiar, Arjun Jain, Jonathan Tompson, Leonid Pishchulin, Mykhaylo Andriluka, Christoph Bregler, Bernt Schiele, Christian Theobalt |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2016 | DeepCut: Joint Subset Partition and Labeling for Multi Person Pose EstimationabstractThis paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts between people in close proximity of each other. This joint formulation is in contrast to previous strategies, that address the problem by first detecting people and subsequently estimating their body pose. We propose a partitioning and labeling formulation of a set of body-part hypotheses generated with CNN-based part detectors. Our formulation, an instance of an integer linear program, implicitly performs non-maximum suppression on the set of part candidates and groups them to form configurations of body parts respecting geometric and appearance constraints. Experiments on four different datasets demonstrate state-of-the-art results for both single person and multi person pose estimation. Leonid Pishchulin, Eldar Insafutdinov, Siyu Tang 0001, Bjoern Andres, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
CVPR | 5 |
| 2016 | End-to-End People Detection in Crowded ScenesabstractCurrent people detectors operate either by scanning an image in a sliding window fashion or by classifying a discrete set of proposals. We propose a model that is based on decoding an image into a set of people detections. Our system takes an image as input and directly outputs a set of distinct detection hypotheses. Because we generate predictions jointly, common post-processing steps such as nonmaximum suppression are unnecessary. We use a recurrent LSTM layer for sequence generation and train our model end-to-end with a new loss function that operates on sets of detections. We demonstrate the effectiveness of our approach on the challenging task of detecting people in crowded scenes1. Russell Stewart, Mykhaylo Andriluka, Andrew Y. Ng |
CVPR | 2 |
| 2016 | DeeperCut: A Deeper, Stronger, and Faster Multi-person Pose Estimation Model
Eldar Insafutdinov, Leonid Pishchulin, Bjoern Andres, Mykhaylo Andriluka, Bernt Schiele |
ECCV (6) | 4 |
| 2016 | Recognizing Fine-Grained and Composite Activities Using Hand-Centric Features and Script Data
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, Bernt Schiele |
Int. J. Comput. Vis. | 5 |
| 2016 | 3D Pictorial Structures Revisited: Multiple Human Pose EstimationabstractWe address the problem of 3D pose estimation of multiple humans from multiple views. The transition from single to multiple human pose estimation and from the 2D to 3D space is challenging due to a much larger state space, occlusions and across-view ambiguities when not knowing the identity of the humans in advance. To address these problems, we first create a reduced state space by triangulation of corresponding pairs of body parts obtained by part detectors for each camera view. In order to resolve ambiguities of wrong and mixed parts of multiple humans after triangulation and also those coming from false positive detections, we introduce a 3D pictorial structures (3DPS) model. Our model builds on multi-view unary potentials, while a prior model is integrated into pairwise and ternary potential functions. To balance the potentials' influence, the model parameters are learnt using a Structured SVM (SSVM). The model is generic and applicable to both single and multiple human pose estimation. To evaluate our model on single and multiple human pose estimation, we rely on four different datasets. We first analyse the contribution of the potentials and then compare our results with related work where we demonstrate superior performance. Vasileios Belagiannis, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele, Nassir Navab, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Emotion recognition from embedded bodily expressions and speech during dyadic interactionsabstractPrevious work on emotion recognition from bodily expressions focused on analysing such expressions in isolation, of individuals or in controlled settings, from a single camera view, or required intrusive motion tracking equipment. We study the problem of emotion recognition from bodily expressions and speech during dyadic (person-person) interactions in a real kitchen instrumented with ambient cameras and microphones. We specifically focus on bodily expressions that are embedded in regular interactions and background activities and recorded without human augmentation to increase naturalness of the expressions. We present a human-validated dataset that contains 224 high-resolution, multi-view video clips and audio recordings of emotionally charged interactions between eight couples of actors. The dataset is fully annotated with categorical labels for four basic emotions (anger, happiness, sadness, and surprise) and continuous labels for valence, activation, power, and anticipation provided by five annotators for each actor. We evaluate vision and audio-based emotion recognition using dense trajectories and a standard audio pipeline and provide insights into the importance of different body parts and audio features for emotion recognition. Philipp Müller 0001, Sikandar Amin, Mykhaylo Andriluka, Andreas Bulling |
ACII | 4 |
| 2015 | Efficient ConvNet-based marker-less motion capture in general scenes with a low number of camerasabstractWe present a novel method for accurate marker-less capture of articulated skeleton motion of several subjects in general scenes, indoors and outdoors, even from input filmed with as few as two cameras. Our approach unites a discriminative image-based joint detection method with a model-based generative motion tracking algorithm through a combined pose optimization energy. The discriminative part-based pose detection method, implemented using Convolutional Networks (ConvNet), estimates unary potentials for each joint of a kinematic skeleton model. These unary potentials are used to probabilistically extract pose constraints for tracking by using weighted sampling from a pose posterior guided by the model. In the final energy, these constraints are combined with an appearance-based model-to-image similarity term. Poses can be computed very efficiently using iterative local optimization, as ConvNet detection is fast, and our formulation yields a combined pose estimation energy with analytic derivatives. In combination, this enables to track full articulated joint angles at state-of-the-art accuracy and temporal stability with a very low number of cameras. Ahmed Elhayek, Edilson de Aguiar, Arjun Jain, Jonathan Tompson, Leonid Pishchulin, Mykhaylo Andriluka, Christoph Bregler, Bernt Schiele, Christian Theobalt |
CVPR | 6 |
| 2015 | Subgraph decomposition for multi-target trackingabstractTracking multiple targets in a video, based on a finite set of detection hypotheses, is a persistent problem in computer vision. A common strategy for tracking is to first select hypotheses spatially and then to link these over time while maintaining disjoint path constraints [14, 15, 24]. In crowded scenes multiple hypotheses will often be similar to each other making selection of optimal links an unnecessary hard optimization problem due to the sequential treatment of space and time. Embracing this observation, we propose to link and cluster plausible detections jointly across space and time. Specifically, we state multi-target tracking as a Minimum Cost Subgraph Multicut Problem. Evidence about pairs of detection hypotheses is incorporated whether the detections are in the same frame, neighboring frames or distant frames. This facilitates long-range re-identification and within-frame clustering. Results for published benchmark sequences demonstrate the superiority of this approach. Siyu Tang 0001, Bjoern Andres, Mykhaylo Andriluka, Bernt Schiele |
CVPR | 3 |
| 2014 | 2D Human Pose Estimation: New Benchmark and State of the Art AnalysisabstractHuman pose estimation has made significant progress during the last years. However current datasets are limited in their coverage of the overall pose estimation challenges. Still these serve as the common sources to evaluate, train and compare different models on. In this paper we introduce a novel benchmark "MPII Human Pose" that makes a significant advance in terms of diversity and difficulty, a contribution that we feel is required for future developments in human body models. This comprehensive dataset was collected using an established taxonomy of over 800 human activities [1]. The collected images cover a wider variety of human activities than previous datasets including various recreational, occupational and householding activities, and capture people from a wider range of viewpoints. We provide a rich set of labels including positions of body joints, full 3D torso and head orientation, occlusion labels for joints and body parts, and activity labels. For each image we provide adjacent video frames to facilitate the use of motion information. Given these rich annotations we perform a detailed analysis of leading human pose estimation approaches and gaining insights for the success and failures of these methods. Mykhaylo Andriluka, Leonid Pishchulin, Peter V. Gehler, Bernt Schiele |
CVPR | 1 |
| 2014 | 3D Pictorial Structures for Multiple Human Pose EstimationabstractIn this work, we address the problem of 3D pose estimation of multiple humans from multiple views. This is a more challenging problem than single human 3D pose estimation due to the much larger state space, partial occlusions as well as across view ambiguities when not knowing the identity of the humans in advance. To address these problems, we first create a reduced state space by triangulation of corresponding body joints obtained from part detectors in pairs of camera views. In order to resolve the ambiguities of wrong and mixed body parts of multiple humans after triangulation and also those coming from false positive body part detections, we introduce a novel 3D pictorial structures (3DPS) model. Our model infers 3D human body configurations from our reduced state space. The 3DPS model is generic and applicable to both single and multiple human pose estimation. In order to compare to the state-of-the art, we first evaluate our method on single human 3D pose estimation on HumanEva-I [22] and KTH Multiview Football Dataset II [8] datasets. Then, we introduce and evaluate our method on two datasets for multiple human 3D pose estimation. In order to compare to the state-of-the art, we first evaluate our method on single human 3D pose estimation on HumanEva-I [22] and KTH Multiview Football Dataset II [8] datasets. Then, we introduce and evaluate our method on two datasets for multiple human 3D pose estimation. Vasileios Belagiannis, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele, Nassir Navab, Slobodan Ilic |
CVPR | 3 |
| 2014 | Detection and Tracking of Occluded People
Siyu Tang 0001, Mykhaylo Andriluka, Bernt Schiele |
Int. J. Comput. Vis. | 2 |
| 2013 | Multi-view Pictorial Structures for 3D Human Pose EstimationabstractPictorial structure models are the de facto standard for 2D human pose estimation. Numerous refinements and improvements have been proposed such as discriminatively trained body part detectors, flexible body models, and local and global mixtures. While these techniques allow to achieve state-of-the-art performance for 2D pose estimation, they have not yet been extended to enable pose estimation in 3D. This paper thus proposes a multi-view pictorial structures model that builds on recent advances in 2D pose estimation and incorporates evidence across multiple viewpoints to allow for robust 3D pose estimation. We evaluate our multi-view pictorial structures approach on the HumanEva-I and MPII Cooking dataset. In comparison to related work for 3D pose estimation our approach achieves similar or better results while operating on single-frames only and not relying on activity specific motion models or tracking. Notably, our approach outperforms state-of-the-art for activities with more complex motions. Sikandar Amin, Mykhaylo Andriluka, Marcus Rohrbach, Bernt Schiele |
BMVC | 2 |
| 2013 | Poselet Conditioned Pictorial StructuresabstractIn this paper we consider the challenging problem of articulated human pose estimation in still images. We observe that despite high variability of the body articulations, human motions and activities often simultaneously constrain the positions of multiple body parts. Modelling such higher order part dependencies seemingly comes at a cost of more expensive inference, which resulted in their limited use in state-of-the-art methods. In this paper we propose a model that incorporates higher order part dependencies while remaining efficient. We achieve this by defining a conditional model in which all body parts are connected a-priori, but which becomes a tractable tree-structured pictorial structures model once the image observations are available. In order to derive a set of conditioning variables we rely on the poselet-based features that have been shown to be effective for people detection but have so far found limited application for articulated human pose estimation. We demonstrate the effectiveness of our approach on three publicly available pose estimation benchmarks improving or being on-par with state of the art in each case. Leonid Pishchulin, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
CVPR | 2 |
| 2013 | Strong Appearance and Expressive Spatial Models for Human Pose EstimationabstractTypical approaches to articulated pose estimation combine spatial modelling of the human body with appearance modelling of body parts. This paper aims to push the state-of-the-art in articulated pose estimation in two ways. First we explore various types of appearance representations aiming to substantially improve the body part hypotheses. And second, we draw on and combine several recently proposed powerful ideas such as more flexible spatial models as well as image-conditioned spatial models. In a series of experiments we draw several important conclusions: (1) we show that the proposed appearance representations are complementary, (2) we demonstrate that even a basic tree-structure spatial human body model achieves state-of-the-art performance when augmented with the proper appearance representation, and (3) we show that the combination of the best performing appearance model with a flexible image-conditioned spatial model achieves the best result, significantly improving over the state of the art, on the ``Leeds Sports Poses'' and ``Parse'' benchmarks. Leonid Pishchulin, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
ICCV | 2 |
| 2013 | Learning People Detectors for Tracking in Crowded ScenesabstractPeople tracking in crowded real-world scenes is challenging due to frequent and long-term occlusions. Recent tracking methods obtain the image evidence from object (people) detectors, but typically use off-the-shelf detectors and treat them as black box components. In this paper we argue that for best performance one should explicitly train people detectors on failure cases of the overall tracker instead. To that end, we first propose a novel joint people detector that combines a state-of-the-art single person detector with a detector for pairs of people, which explicitly exploits common patterns of person-person occlusions across multiple viewpoints that are a frequent failure case for tracking in crowded scenes. To explicitly address remaining failure modes of the tracker we explore two methods. First, we analyze typical failures of trackers and train a detector explicitly on these cases. And second, we train the detector with the people tracker in the loop, focusing on the most common tracker failures. We show that our joint multi-person detector significantly improves both detection accuracy as well as tracker performance, improving the state-of-the-art on standard benchmarks. Siyu Tang 0001, Mykhaylo Andriluka, Anton Milan, Konrad Schindler, Stefan Roth 0001, Bernt Schiele |
ICCV | 2 |
| 2012 | Detection and Tracking of Occluded PeopleabstractWe consider the problem of detection and tracking of multiple people in crowded street scenes. State-of-the-art methods perform well in scenes with relatively few people, but are severely challenged by scenes with many subjects that partially occlude each other. This limitation is due to the fact that current people detectors fail when persons are strongly occluded. We observe that typical occlusions are due to overlaps between people and propose a people detector tailored to various occlusion levels. Instead of treating partial occlusions as distractions, we leverage the fact that person/person occlusions result in very characteristic appearance patterns that can help to improve detection results. We demonstrate the performance of our occlusion-aware person detector on a new dataset of people with controlled but severe levels of occlusion and on two challenging publicly available benchmarks outperforming single person detectors in each case. Siyu Tang 0001, Mykhaylo Andriluka, Bernt Schiele |
BMVC | 2 |
| 2012 | Articulated people detection and pose estimation: Reshaping the futureabstractState-of-the-art methods for human detection and pose estimation require many training samples for best performance. While large, manually collected datasets exist, the captured variations w.r.t. appearance, shape and pose are often uncontrolled thus limiting the overall performance. In order to overcome this limitation we propose a new technique to extend an existing training set that allows to explicitly control pose and shape variations. For this we build on recent advances in computer graphics to generate samples with realistic appearance and background while modifying body shape and pose. We validate the effectiveness of our approach on the task of articulated human detection and articulated pose estimation. We report close to state of the art results on the popular Image Parsing [25] human pose estimation benchmark and demonstrate superior performance for articulated human detection. In addition we define a new challenge of combined articulated human detection and pose estimation in real-world scenes. Leonid Pishchulin, Arjun Jain, Mykhaylo Andriluka, Thorsten Thormählen, Bernt Schiele |
CVPR | 3 |
| 2012 | A database for fine grained activity detection of cooking activitiesabstractWhile activity recognition is a current focus of research the challenging problem of fine-grained activity recognition is largely overlooked. We thus propose a novel database of 65 cooking activities, continuously recorded in a realistic setting. Activities are distinguished by fine-grained body motions that have low inter-class variability and high intra-class variability due to diverse subjects and ingredients. We benchmark two approaches on our dataset, one based on articulated pose tracks and the second using holistic video features. While the holistic approach outperforms the pose-based approach, our evaluation suggests that fine-grained activities are more difficult to detect and the body model can help in those cases. Providing high-resolution videos as well as an intermediate pose representation we hope to foster research in fine-grained activity recognition. Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele |
CVPR | 3 |
| 2012 | Script Data for Attribute-Based Recognition of Composite Activities
Marcus Rohrbach, Michaela Regneri, Mykhaylo Andriluka, Sikandar Amin, Manfred Pinkal, Bernt Schiele |
ECCV (1) | 3 |
| 2012 | Discriminative Appearance Models for Pictorial Structures
Mykhaylo Andriluka, Stefan Roth 0001, Bernt Schiele |
Int. J. Comput. Vis. | 1 |
| 2011 | Learning people detection models from few training samplesabstractPeople detection is an important task for a wide range of applications in computer vision. State-of-the-art methods learn appearance based models requiring tedious collection and annotation of large data corpora. Also, obtaining data sets representing all relevant variations with sufficient accuracy for the intended application domain at hand is often a non-trivial task. Therefore this paper investigates how 3D shape models from computer graphics can be leveraged to ease training data generation. In particular we employ a rendering-based reshaping method in order to generate thousands of synthetic training samples from only a few persons and views. We evaluate our data generation method for two different people detection models. Our experiments on a challenging multi-view dataset indicate that the data from as few as eleven persons suffices to achieve good performance. When we additionally combine our synthetic training samples with real data we even outperform existing state-of-the-art methods. Leonid Pishchulin, Arjun Jain, Christian Wojek, Mykhaylo Andriluka, Thorsten Thormählen, Bernt Schiele |
CVPR | 4 |
| 2010 | Monocular 3D pose estimation and tracking by detectionabstractAutomatic recovery of 3D human pose from monocular image sequences is a challenging and important research topic with numerous applications. Although current methods are able to recover 3D pose for a single person in controlled environments, they are severely challenged by real-world scenarios, such as crowded street scenes. To address this problem, we propose a three-stage process building on a number of recent advances. The first stage obtains an initial estimate of the 2D articulation and viewpoint of the person from single frames. The second stage allows early data association across frames based on tracking-by-detection. These two stages successfully accumulate the available 2D image evidence into robust estimates of 2D limb positions over short image sequences (= tracklets). The third and final stage uses those tracklet-based estimates as robust image observations to reliably recover 3D pose. We demonstrate state-of-the-art performance on the HumanEva II benchmark, and also show the applicability of our approach to articulated 3D tracking in realistic street conditions. Mykhaylo Andriluka, Stefan Roth 0001, Bernt Schiele |
CVPR | 1 |
| 2010 | Vision based victim detection from unmanned aerial vehiclesabstractFinding injured humans is one of the primary goals of any search and rescue operation. The aim of this paper is to address the task of automatically finding people lying on the ground in images taken from the on-board camera of an unmanned aerial vehicle (UAV). In this paper we evaluate various state-of-the-art visual people detection methods in the context of vision based victim detection from an UAV. The top performing approaches in this comparison are those that rely on flexible part-based representations and discriminatively trained part detectors. We discuss their strengths and weaknesses and demonstrate that by combining multiple models we can increase the reliability of the system. We also demonstrate that the detection performance can be substantially improved by integrating the height and pitch information provided by on-board sensors. Jointly these improvements allow us to significantly boost the detection performance over the current de-facto standard, which provides a substantial step towards making autonomous victim detection for UAVs practical. Mykhaylo Andriluka, Paul Schnitzspan, Stefan Kohlbrecher, Karen Petersen, Oskar von Stryk, Stefan Roth 0001, Bernt Schiele |
IROS | 1 |
| 2010 | A Semantic World Model for Urban Search and Rescue Based on Heterogeneous Sensors
Paul Schnitzspan, Stefan Kohlbrecher, Karen Petersen, Mykhaylo Andriluka, Oliver Schwahn, Uwe Klingauf, Stefan Roth 0001, Bernt Schiele, Oskar von Stryk |
RoboCup | 5 |
| 2009 | Pictorial structures revisited: People detection and articulated pose estimationabstractNon-rigid object detection and articulated pose estimation are two related and challenging problems in computer vision. Numerous models have been proposed over the years and often address different special cases, such as pedestrian detection or upper body pose estimation in TV footage. This paper shows that such specialization may not be necessary, and proposes a generic approach based on the pictorial structures framework. We show that the right selection of components for both appearance and spatial modeling is crucial for general applicability and overall performance of the model. The appearance of body parts is modeled using densely sampled shape context descriptors and discriminatively trained AdaBoost classifiers. Furthermore, we interpret the normalized margin of each classifier as likelihood in a generative model. Non-Gaussian relationships between parts are represented as Gaussians in the coordinate system of the joint between parts. The marginal posterior of each part is inferred using belief propagation. We demonstrate that such a model is equally suitable for both detection and pose estimation tasks, outperforming the state of the art on three recently proposed datasets. Mykhaylo Andriluka, Stefan Roth 0001, Bernt Schiele |
CVPR | 1 |
| 2008 | People-tracking-by-detection and people-detection-by-trackingabstractBoth detection and tracking people are challenging problems, especially in complex real world scenes that commonly involve multiple people, complicated occlusions, and cluttered or even moving backgrounds. People detectors have been shown to be able to locate pedestrians even in complex street scenes, but false positives have remained frequent. The identification of particular individuals has remained challenging as well. Tracking methods are able to find a particular individual in image sequences, but are severely challenged by real-world scenarios such as crowded street scenes. In this paper, we combine the advantages of both detection and tracking in a single framework. The approximate articulation of each person is detected in every frame based on local features that model the appearance of individual body parts. Prior knowledge on possible articulations and temporal coherency within a walking cycle are modeled using a hierarchical Gaussian process latent variable model (hGPLVM). We show how the combination of these results improves hypotheses for position and articulation of each person in several subsequent frames. We present experimental results that demonstrate how this allows to detect and track multiple people in cluttered scenes with reoccurring occlusions. Mykhaylo Andriluka, Stefan Roth 0001, Bernt Schiele |
CVPR | 1 |