VLDB 2026 Research / reviewers in the wild / expert
Peter V. Gehler
dblp:78/1502 · also Peter Vincent Gehler
· DBLP profile ↗
48ranked-venue papers
6as first author
9since 2021 · last 2023
0000-0002-5812-4052ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
38 papers |
Trustworthy machine learning · 16% 3D vision · 14% Face, body and person analysis · 12% | |
| Computer graphics and multimedia
11 papers |
Computational photography and imaging · 55% Visual content generation and editing · 20% Image and video processing · 16% |
Topics — the 30 heaviest of 93, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
human pose estimation |
1.7 | 8 | 2017 | Unite the People: Closing the Loop Between 3D and 2D Human Representations · CVPR 2017 Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image · ECCV (5) 2016 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation · CVPR 2016 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
1.1 | 2 | 2022 | Assaying Out-Of-Distribution Generalization in Transfer Learning · NeurIPS 2022 The Role of Pretrained Representations for the OOD Generalization of RL Agents · ICLR 2022 |
Computational photography and imaging
intrinsic image decomposition |
0.6 | 3 | 2017 | Reflectance Adaptive Filtering Improves Intrinsic Image Estimation · CVPR 2017 Intrinsic Video · ECCV (2) 2014 Recovering Intrinsic Images with a Global Sparsity Prior on Reflectance · NIPS 2011 |
Machine learning › Time series and sequential data
anomaly detection |
0.6 | 1 | 2022 | Towards Total Recall in Industrial Anomaly Detection · CVPR 2022 |
Natural language and speech › Language models and text generation
compositional generalization |
0.6 | 1 | 2022 | Visual Representation Learning Does Not Generalize Strongly Within the Same Domain · ICLR 2022 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution |
0.6 | 1 | 2022 | You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory Prediction · ICLR 2022 |
Machine learning › Time series and sequential data › anomaly detection
industrial anomaly detection |
0.6 | 1 | 2022 | Towards Total Recall in Industrial Anomaly Detection · CVPR 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.6 | 1 | 2022 | You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory Prediction · ICLR 2022 |
Machine learning › Representation and self-supervised learning › pre-training
pre-trained representations |
0.6 | 1 | 2022 | The Role of Pretrained Representations for the OOD Generalization of RL Agents · ICLR 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | Assaying Out-Of-Distribution Generalization in Transfer Learning · NeurIPS 2022 |
Robotics › Autonomous driving
trajectory prediction |
0.6 | 1 | 2022 | You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory Prediction · ICLR 2022 |
Computer vision › Segmentation and scene understanding › video segmentation
video semantic segmentation |
0.6 | 2 | 2017 | Semantic Video CNNs Through Representation Warping · ICCV 2017 Video Propagation Networks · CVPR 2017 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.6 | 1 | 2022 | Visual Representation Learning Does Not Generalize Strongly Within the Same Domain · ICLR 2022 |
Computer vision › 3D vision
human mesh recovery |
0.5 | 2 | 2017 | Unite the People: Closing the Loop Between 3D and 2D Human Representations · CVPR 2017 Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image · ECCV (5) 2016 |
Computational photography and imaging
color constancy |
0.5 | 2 | 2020 | Providing a Single Ground-Truth for Illuminant Estimation for the ColorChecker Dataset · IEEE Trans. Pattern Anal. Mach. Intell. 2020 Bayesian color constancy revisited · CVPR 2008 |
Computational photography and imaging › color constancy
illuminant estimation |
0.5 | 2 | 2020 | Providing a Single Ground-Truth for Illuminant Estimation for the ColorChecker Dataset · IEEE Trans. Pattern Anal. Mach. Intell. 2020 Bayesian color constancy revisited · CVPR 2008 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.5 | 1 | 2021 | Dynamic Inference with Neural Interpreters · NeurIPS 2021 |
Natural language and speech › Language models and text generation › knowledge editing
negative flip reduction |
0.5 | 1 | 2021 | Backward-Compatible Prediction Updates: A Probabilistic Approach · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness
prediction consistency |
0.5 | 1 | 2021 | Backward-Compatible Prediction Updates: A Probabilistic Approach · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning
systematic generalization |
0.5 | 1 | 2021 | Dynamic Inference with Neural Interpreters · NeurIPS 2021 |
Computer vision › Vision and language
video captioning |
0.5 | 1 | 2021 | CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations · ICCV 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.5 | 3 | 2017 | Semantic Video CNNs Through Representation Warping · ICCV 2017 On Parameter Learning in CRF-Based Approaches to Object Class Image Segmentation · ECCV (6) 2010 A Generative Model of People in Clothing · ICCV 2017 |
Computer vision › 3D vision › stereo vision
event-based stereo |
0.4 | 1 | 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo · ICCV 2019 |
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction |
0.4 | 1 | 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo · ICCV 2019 |
Computer vision › 3D vision
3d object detection |
0.4 | 2 | 2015 | Multi-View and 3D Deformable Part Models · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Teaching 3D geometry to deformable part models · CVPR 2012 |
Computer vision › 3D vision › object pose estimation
viewpoint estimation |
0.4 | 2 | 2015 | Multi-View and 3D Deformable Part Models · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Teaching 3D geometry to deformable part models · CVPR 2012 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.3 | 2 | 2017 | Superpixel Convolutional Networks Using Bilateral Inceptions · ECCV (1) 2016 Reflectance Adaptive Filtering Improves Intrinsic Image Estimation · CVPR 2017 |
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation |
0.3 | 1 | 2018 | Deep Directional Statistics: Pose Estimation with Uncertainty Quantification · ECCV (9) 2018 |
Computer vision › Segmentation and scene understanding › scene parsing
facade parsing |
0.3 | 1 | 2018 | Efficient 2D and 3D Facade Segmentation Using Auto-Context · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.3 | 1 | 2018 | Deep Directional Statistics: Pose Estimation with Uncertainty Quantification · ECCV (9) 2018 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 1.4reinforcement learning · 0.6pre-trained representations · 0.6outlier detection · 0.6memory bank of nominal patch features · 0.6imagenet embeddings · 0.6fine-tuning · 0.6feature attribution · 0.6calibration error · 0.6adversarial attack · 0.6negative sampling · 0.5contrastive learning · 0.5temporal bilateral network · 0.3spatial refinement network · 0.3online propagation · 0.3loss function design · 0.3joint bilateral filtering · 0.3generative model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TeST: Test-time Self-Training under Distribution ShiftabstractDespite their recent success, deep neural networks continue to perform poorly when they encounter distribution shifts at test time. Many recently proposed approaches try to counter this by aligning the model to the new distribution prior to inference. With no labels available this requires unsupervised objectives to adapt the model on the observed test data. In this paper, we propose Test-Time Self-Training (TeST): a technique that takes as input a model trained on some source data and a novel data distribution at test time, and learns invariant and robust representations using a student-teacher framework. We find that models adapted using TeST significantly improve over baseline test-time adaptation algorithms. TeST achieves competitive performance to modern domain adaptation algorithms [4], [43], while having access to 5-10x less data at time of adaption. We thoroughly evaluate a variety of baselines on two tasks: object detection and image segmentation and find that models adapted with TeST. We find that TeST sets the new state-of-the art for test-time domain adaptation algorithms. Samarth Sinha, Peter V. Gehler, Francesco Locatello, Bernt Schiele |
WACV | 2 |
| 2022 | Towards Total Recall in Industrial Anomaly DetectionabstractBeing able to spot defective parts is a critical component in large-scale industrial manufacturing. A particular challenge that we address in this work is the cold-start problem: fit a model using nominal (non-defective) example images only. While handcrafted solutions per class are possible, the goal is to build systems that work well simultaneously on many different tasks automatically. The best peforming approaches combine embeddings from ImageNet models with an outlier detection model. In this paper, we extend on this line of work and propose PatchCore, which uses a maximally representative memory bank of nominal patch-features. PatchCore offers competitive inference times while achieving state-of-the-art performance for both detection and localization. On the challenging, widely used MVTec AD benchmark PatchCore achieves an image-level anomaly detection AUROC score of up to 99.6%, more than halving the error compared to the next best competitor. We further report competitive results on two additional datasets and also find competitive results in the few samples regime. Code: github.com/amazon-research/patchcore-inspection. Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf, Thomas Brox, Peter V. Gehler |
CVPR | 6 |
| 2022 | You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory Prediction
Osama Makansi, Julius von Kügelgen, Francesco Locatello, Peter V. Gehler, Dominik Janzing, Thomas Brox, Bernhard Schölkopf |
ICLR | 4 |
| 2022 | Visual Representation Learning Does Not Generalize Strongly Within the Same Domain
Lukas Schott, Julius von Kügelgen, Frederik Träuble, Peter V. Gehler, Chris Russell 0001, Matthias Bethge, Bernhard Schölkopf, Francesco Locatello, Wieland Brendel |
ICLR | 4 |
| 2022 | The Role of Pretrained Representations for the OOD Generalization of RL Agents
Frederik Träuble, Andrea Dittadi, Manuel Wüthrich, Felix Widmaier, Peter V. Gehler, Ole Winther, Francesco Locatello, Olivier Bachem, Bernhard Schölkopf, Stefan Bauer |
ICLR | 5 |
| 2022 | Assaying Out-Of-Distribution Generalization in Transfer LearningabstractSince out-of-distribution generalization is a generally ill-posed problem, various proxy targets (e.g., calibration, adversarial robustness, algorithmic corruptions, invariance across shifts) were studied across different research programs resulting in different recommendations. While sharing the same aspirational goal, these approaches have never been tested under the same experimental conditions on real data. In this paper, we take a unified view of previous work, highlighting message discrepancies that we address empirically, and providing recommendations on how to measure the robustness of a model and how to improve it. To this end, we collect 172 publicly available dataset pairs for training and out-of-distribution evaluation of accuracy, calibration error, adversarial attacks, environment invariance, and synthetic corruptions. We fine-tune over 31k networks, from nine different architectures in the many- and few-shot setting. Our findings confirm that in- and out-of-distribution accuracies tend to increase jointly, but show that their relation is largely dataset-dependent, and in general more nuanced and more complex than posited by previous, smaller scale studies. Florian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel, Max Horn, Dominik Zietlow, David Kernert, Chris Russell 0001, Thomas Brox, Bernt Schiele, Bernhard Schölkopf, Francesco Locatello |
NeurIPS | 3 |
| 2021 | CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video RepresentationsabstractContrastive learning allows us to flexibly define powerful losses by contrasting positive pairs from sets of negative samples. Recently, the principle has also been used to learn cross-modal embeddings for video and text, yet without exploiting its full potential. In particular, previous losses do not take the intra-modality similarities into account, which leads to inefficient embeddings, as the same content is mapped to multiple points in the embedding space. With CrossCLR, we present a contrastive loss that fixes this issue. Moreover, we define sets of highly related samples in terms of their input embeddings and exclude them from the negative samples to avoid issues with false negatives. We show that these principles consistently improve the quality of the learned embeddings. The joint embeddings learned with CrossCLR extend the state of the art in video-text retrieval on Youcook2 and LSMDC datasets and in video captioning on Youcook2 dataset by a large margin. We also demonstrate the generality of the concept by learning improved joint embeddings for other pairs of modalities. Mohammadreza Zolfaghari, Peter V. Gehler, Thomas Brox |
ICCV | 3 |
| 2021 | Dynamic Inference with Neural InterpretersabstractModern neural network architectures can leverage large amounts of data to generalize well within the training distribution. However, they are less capable of systematic generalization to data drawn from unseen but related distributions, a feat that is hypothesized to require compositional reasoning and reuse of knowledge. In this work, we present Neural Interpreters, an architecture that factorizes inference in a self-attention network as a system of modules, which we call functions. Inputs to the model are routed through a sequence of functions in a way that is end-to-end learned. The proposed architecture can flexibly compose computation along width and depth, and lends itself well to capacity extension after training. To demonstrate the versatility of Neural Interpreters, we evaluate it in two distinct settings: image classification and visual abstract reasoning on Raven Progressive Matrices. In the former, we show that Neural Interpreters perform on par with the vision transformer using fewer parameters, while being transferrable to a new task in a sample efficient manner. In the latter, we find that Neural Interpreters are competitive with respect to the state-of-the-art in terms of systematic generalization. Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter V. Gehler, Yoshua Bengio, Francesco Locatello, Bernhard Schölkopf |
NeurIPS | 4 |
| 2021 | Backward-Compatible Prediction Updates: A Probabilistic ApproachabstractWhen machine learning systems meet real world applications, accuracy is only one of several requirements. In this paper, we assay a complementary perspective originating from the increasing availability of pre-trained and regularly improving state-of-the-art models. While new improved models develop at a fast pace, downstream tasks vary more slowly or stay constant. Assume that we have a large unlabelled data set for which we want to maintain accurate predictions. Whenever a new and presumably better ML models becomes available, we encounter two problems: (i) given a limited budget, which data points should be re-evaluated using the new model?; and (ii) if the new predictions differ from the current ones, should we update? Problem (i) is about compute cost, which matters for very large data sets and models. Problem (ii) is about maintaining consistency of the predictions, which can be highly relevant for downstream applications; our demand is to avoid negative flips, i.e., changing correct to incorrect predictions. In this paper, we formalize the Prediction Update Problem and present an efficient probabilistic approach as answer to the above questions. In extensive experiments on standard classification benchmark data sets, we show that our method outperforms alternative strategies along key metrics for backward-compatible prediction updates. Frederik Träuble, Julius von Kügelgen, Matthäus Kleindessner, Francesco Locatello, Bernhard Schölkopf, Peter V. Gehler |
NeurIPS | 6 |
| 2020 | Providing a Single Ground-Truth for Illuminant Estimation for the ColorChecker DatasetabstractThe ColorChecker dataset is one of the most widely used image sets for evaluating and ranking illuminant estimation algorithms. However, this single set of images has at least 3 different sets of ground-truth (i.e., correct answers) associated with it. In the literature it is often asserted that one algorithm is better than another when the algorithms in question have been tuned and tested with the different ground-truths. In this short correspondence we present some of the background as to why the 3 existing ground-truths are different and go on to make a new single and recommended set of correct answers. Experiments reinforce the importance of this work in that we show that the total ordering of a set of algorithms may be reversed depending on whether we use the new or legacy ground-truth data. Ghalia Hemrit, Graham D. Finlayson, Arjan Gijsenij, Peter V. Gehler, Simone Bianco 0001, Mark S. Drew, Brian V. Funt, Lilong Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep StereoabstractToday, a frame-based camera is the sensor of choice for machine vision applications. However, these cameras, originally developed for acquisition of static images rather than for sensing of dynamic uncontrolled visual environments, suffer from high power consumption, data rate, latency and low dynamic range. An event-based image sensor addresses these drawbacks by mimicking a biological retina. Instead of measuring the intensity of every pixel in a fixed time-interval, it reports events of significant pixel intensity changes. Every such event is represented by its position, sign of change, and timestamp, accurate to the microsecond. Asynchronous event sequences require special handling, since traditional algorithms work only with synchronous, spatially gridded data. To address this problem we introduce a new module for event sequence embedding, for use in difference applications. The module builds a representation of an event sequence by firstly aggregating information locally across time, using a novel fully-connected layer for an irregularly sampled continuous domain, and then across discrete spatial domain. Based on this module, we design a deep learning-based stereo method for event-based cameras. The proposed method is the first learning-based stereo method for an event-based camera and the only method that produces dense results. We show that large performance increases on the Multi Vehicle Stereo Event Camera Dataset (MVSEC), which became the standard set for benchmarking of event-based stereo methods. Stepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler |
ICCV | 4 |
| 2018 | Neural Body Fitting: Unifying Deep Learning and Model Based Human Pose and Shape EstimationabstractDirect prediction of 3D body pose and shape remains a challenge even for highly parameterized deep learning models. Mapping from the 2D image space to the prediction space is difficult: perspective ambiguities make the loss function noisy and training data is scarce. In this paper, we propose a novel approach (Neural Body Fitting (NBF)). It integrates a statistical body model within a CNN, leveraging reliable bottom-up semantic body part segmentation and robust top-down body model constraints. NBF is fully differentiable and can be trained using 2D and 3D annotations. In detailed experiments, we analyze how the components of our model affect performance, especially the use ofpart segmentations as an explicit intermediate representation, and present a robust, efficiently trainable framework for 3D human pose estimation from 2D images with competitive results on standard benchmarks. Code will be made available at http://github.com/mohomran/ neural_body_fitting. Mohamed Omran, Christoph Lassner, Gerard Pons-Moll, Peter V. Gehler, Bernt Schiele |
3DV | 4 |
| 2018 | Deep Directional Statistics: Pose Estimation with Uncertainty Quantification
Sergey Prokudin, Peter V. Gehler, Sebastian Nowozin |
ECCV (9) | 2 |
| 2018 | Efficient 2D and 3D Facade Segmentation Using Auto-ContextabstractThis paper introduces a fast and efficient segmentation technique for 2D images and 3D point clouds of building facades. Facades of buildings are highly structured and consequently most methods that have been proposed for this problem aim to make use of this strong prior information. Contrary to most prior work, we are describing a system that is almost domain independent and consists of standard segmentation methods. We train a sequence of boosted decision trees using auto-context features. This is learned using stacked generalization. We find that this technique performs better, or comparable with all previous published methods and present empirical results on all available 2D and 3D facade benchmark datasets. The proposed method is simple to implement, easy to extend, and very efficient at test-time inference. Raghudeep Gadde, Varun Jampani, Renaud Marlet, Peter V. Gehler |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Video Propagation NetworksabstractWe propose a technique that propagates information forward through video data. The method is conceptually simple and can be applied to tasks that require the propagation of structured information, such as semantic labels, based on video content. We propose a Video Propagation Network that processes video frames in an adaptive manner. The model is applied online: it propagates information forward without the need to access future frames. In particular we combine two components, a temporal bilateral network for dense and video adaptive filtering, followed by a spatial network to refine features and increased flexibility. We present experiments on video object segmentation and semantic video segmentation and show increased performance comparing to the best previous task-specific methods, while having favorable runtime. Additionally we demonstrate our approach on an example regression task of color propagation in a grayscale video. Varun Jampani, Raghudeep Gadde, Peter V. Gehler |
CVPR | 3 |
| 2017 | Unite the People: Closing the Loop Between 3D and 2D Human Representationsabstract3D models provide a common ground for different representations of human bodies. In turn, robust 2D estimation has proven to be a powerful tool to obtain 3D fits in-the-wild. However, depending on the level of detail, it can be hard to impossible to acquire labeled data for training 2D estimators on large scale. We propose a hybrid approach to this problem: with an extended version of the recently introduced SMPLify method, we obtain high quality 3D body model fits for multiple human pose datasets. Human annotators solely sort good and bad fits. This procedure leads to an initial dataset, UP-3D, with rich annotations. With a comprehensive set of experiments, we show how this data can be used to train discriminative models that produce results with an unprecedented level of detail: our models predict 31 segments and 91 landmark locations on the body. Using the 91 landmark pose estimator, we present state-of-the art results for 3D human pose and shape estimation using an order of magnitude less training data and without assumptions about gender or pose in the fitting procedure. We show that UP-3D can be enhanced with these improved fits to grow in quantity and quality, which makes the system deployable on large scale. The data, code and models are available for research purposes. Christoph Lassner, Javier Romero 0002, Martin Kiefel, Federica Bogo, Michael J. Black, Peter V. Gehler |
CVPR | 6 |
| 2017 | Reflectance Adaptive Filtering Improves Intrinsic Image EstimationabstractSeparating an image into reflectance and shading layers poses a challenge for learning approaches because no large corpus of precise and realistic ground truth decompositions exists. The Intrinsic Images in the Wild (IIW) dataset provides a sparse set of relative human reflectance judgments, which serves as a standard benchmark for intrinsic images. A number of methods use IIW to learn statistical dependencies between the images and their reflectance layer. Although learning plays an important role for high performance, we show that a standard signal processing technique achieves performance on par with current state-of-the-art. We propose a loss function for CNN learning of dense reflectance predictions. Our results show a simple pixel-wise decision, without any context or prior knowledge, is sufficient to provide a strong baseline on IIW. This sets a competitive baseline which only two other approaches surpass. We then develop a joint bilateral filtering method that implements strong prior knowledge about reflectance constancy. This filtering operation can be applied to any intrinsic image algorithm and we improve several previous results achieving a new state-of-the-art on IIW. Our findings suggest that the effect of learning-based approaches may have been over-estimated so far. Explicit prior knowledge is still at least as important to obtain high performance in intrinsic image decompositions. Thomas Nestmeyer, Peter V. Gehler |
CVPR | 2 |
| 2017 | Semantic Video CNNs Through Representation WarpingabstractIn this work, we propose a technique to convert CNN models for semantic segmentation of static images into CNNs for video data. We describe a warping method that can be used to augment existing architectures with very lit- tle extra computational cost. This module is called Net- Warp and we demonstrate its use for a range of network architectures. The main design principle is to use opti- cal flow of adjacent frames for warping internal network representations across time. A key insight of this work is that fast optical flow methods can be combined with many different CNN architectures for improved performance and end-to-end training. Experiments validate that the proposed approach incurs only little extra computational cost, while improving performance, when video streams are available. We achieve new state-of-the-art results on the CamVid and Cityscapes benchmark datasets and show consistent improvements over different baseline networks. Our code and models are available at http://segmentation.is.tue.mpg.de. Raghudeep Gadde, Varun Jampani, Peter V. Gehler |
ICCV | 3 |
| 2017 | A Generative Model of People in ClothingabstractWe present the first image-based generative model of people in clothing for the full body. We sidestep the commonly used complex graphics rendering pipeline and the need for high-quality 3D scans of dressed people. Instead, we learn generative models from a large image database. The main challenge is to cope with the high variance in human pose, shape and appearance. For this reason, pure image-based approaches have not been considered so far. We show that this challenge can be overcome by splitting the generating process in two parts. First, we learn to generate a semantic segmentation of the body and clothing. Second, we learn a conditional model on the resulting segments that creates realistic images. The full model is differentiable and can be conditioned on pose, shape or color. The result are samples of people in different clothing items and styles. The proposed model can generate entirely new people with realistic clothing. In several experiments we present encouraging results that suggest an entirely data-driven approach to people generation is possible. Christoph Lassner, Gerard Pons-Moll, Peter V. Gehler |
ICCV | 3 |
| 2016 | Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs and Bilateral Neural NetworksabstractBilateral filters have wide spread use due to their edge-preserving properties. The common use case is to manually choose a parametric filter type, usually a Gaussian filter. In this paper, we will generalize the parametrization and in particular derive a gradient descent algorithm so the filter parameters can be learned from data. This derivation allows to learn high dimensional linear filters that operate in sparsely populated feature spaces. We build on the permutohedral lattice construction for efficient filtering. The ability to learn more general forms of high-dimensional filters can be used in several diverse applications. First, we demonstrate the use in applications where single filter applications are desired for runtime reasons. Further, we show how this algorithm can be used to learn the pairwise potentials in densely connected conditional random fields and apply these to different image segmentation tasks. Finally, we introduce layers of bilateral filters in CNNs and propose bilateral neural networks for the use of highdimensional sparse data. This view provides new ways to encode model structure into network architectures. A diverse set of experiments empirically validates the usage of general forms of filters. Varun Jampani, Martin Kiefel, Peter V. Gehler |
CVPR | 3 |
| 2016 | DeepCut: Joint Subset Partition and Labeling for Multi Person Pose EstimationabstractThis paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts between people in close proximity of each other. This joint formulation is in contrast to previous strategies, that address the problem by first detecting people and subsequently estimating their body pose. We propose a partitioning and labeling formulation of a set of body-part hypotheses generated with CNN-based part detectors. Our formulation, an instance of an integer linear program, implicitly performs non-maximum suppression on the set of part candidates and groups them to form configurations of body parts respecting geometric and appearance constraints. Experiments on four different datasets demonstrate state-of-the-art results for both single person and multi person pose estimation. Leonid Pishchulin, Eldar Insafutdinov, Siyu Tang 0001, Bjoern Andres, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
CVPR | 6 |
| 2016 | Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter V. Gehler, Javier Romero 0002, Michael J. Black |
ECCV (5) | 4 |
| 2016 | Superpixel Convolutional Networks Using Bilateral Inceptions
Raghudeep Gadde, Varun Jampani, Martin Kiefel, Daniel Kappler, Peter V. Gehler |
ECCV (1) | 5 |
| 2016 | Barrista: Caffe Well-ServedabstractThe caffe framework is one of the leading deep learning toolboxes in the machine learning and computer vision community. While it offers efficiency and configurability, it falls short of a full interface to Python. With increasingly involved procedures for training deep networks and reaching depths of hundreds of layers, creating configuration files and keeping them consistent becomes an error prone process. We introduce the barrista framework, offering full, pythonic control over caffe. It separates responsibilities and offers code to solve frequently occurring tasks for pre-processing, training and model inspection. It is compatible to all caffe versions since mid 2015 and can import and export .prototxt files. Examples are included, e.g., a deep residual network implemented in only 172 lines (for arbitrary depths), comparing to 2320 lines in the official implementation for the equivalent model. Christoph Lassner, Daniel Kappler, Martin Kiefel, Peter V. Gehler |
ACM Multimedia | 4 |
| 2015 | Efficient Facade Segmentation Using Auto-contextabstractIn this paper we propose a system for the problem of facade segmentation. Building facades are highly structured images and consequently most methods that have been proposed for this problem, aim to make use of this strong prior information. We are describing a system that is almost domain independent and consists of standard segmentation methods. A sequence of boosted decision trees is stacked using auto-context features and learned using the stacked generalization technique. We find that this, albeit standard, technique performs better, or equals, all previous published empirical results on all available facade benchmark datasets. The proposed method is simple to implement, easy to extend, and very efficient at test time inference. Varun Jampani, Raghudeep Gadde, Peter V. Gehler |
WACV | 3 |
| 2015 | The informed sampler: A discriminative approach to Bayesian inference in generative computer vision models
Varun Jampani, Sebastian Nowozin, Matthew Loper, Peter V. Gehler |
Comput. Vis. Image Underst. | 4 |
| 2015 | Multi-View and 3D Deformable Part ModelsabstractAs objects are inherently 3D, they have been modeled in 3D in the early days of computer vision. Due to the ambiguities arising from mapping 2D features to 3D models, 3D object representations have been neglected and 2D feature-based models are the predominant paradigm in object detection nowadays. While such models have achieved outstanding bounding box detection performance, they come with limited expressiveness, as they are clearly limited in their capability of reasoning about 3D shape or viewpoints. In this work, we bring the worlds of 3D and 2D object representations closer, by building an object detector which leverages the expressive power of 3D object representations while at the same time can be robustly matched to image evidence. To that end, we gradually extend the successful deformable part model [1] to include viewpoint information and part-level 3D geometry information, resulting in several different models with different level of expressiveness. We end up with a 3D object model, consisting of multiple object parts represented in 3D and a continuous appearance model. We experimentally verify that our models, while providing richer object hypotheses than the 2D object models, provide consistently better joint object localization and viewpoint estimation than the state-of-the-art multi-view and 3D object detectors on various benchmarks (KITTI [2] , 3D object classes [3] , Pascal3D+ [4] , Pascal VOC 2007 [5] , EPFL multi-view cars[6] ). Bojan Pepik, Michael Stark 0003, Peter V. Gehler, Bernt Schiele |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | 2D Human Pose Estimation: New Benchmark and State of the Art AnalysisabstractHuman pose estimation has made significant progress during the last years. However current datasets are limited in their coverage of the overall pose estimation challenges. Still these serve as the common sources to evaluate, train and compare different models on. In this paper we introduce a novel benchmark "MPII Human Pose" that makes a significant advance in terms of diversity and difficulty, a contribution that we feel is required for future developments in human body models. This comprehensive dataset was collected using an established taxonomy of over 800 human activities [1]. The collected images cover a wider variety of human activities than previous datasets including various recreational, occupational and householding activities, and capture people from a wider range of viewpoints. We provide a rich set of labels including positions of body joints, full 3D torso and head orientation, occlusion labels for joints and body parts, and activity labels. For each image we provide adjacent video frames to facilitate the use of motion information. Given these rich annotations we perform a detailed analysis of leading human pose estimation approaches and gaining insights for the success and failures of these methods. Mykhaylo Andriluka, Leonid Pishchulin, Peter V. Gehler, Bernt Schiele |
CVPR | 3 |
| 2014 | Efficient Nonlinear Markov Models for Human MotionabstractDynamic Bayesian networks such as Hidden Markov Models (HMMs) are successfully used as probabilistic models for human motion. The use of hidden variables makes them expressive models, but inference is only approximate and requires procedures such as particle filters or Markov chain Monte Carlo methods. In this work we propose to instead use simple Markov models that only model observed quantities. We retain a highly expressive dynamic model by using interactions that are nonlinear and non-parametric. A presentation of our approach in terms of latent variables shows logarithmic growth for the computation of exact log-likelihoods in the number of latent states. We validate our model on human motion capture data and demonstrate state-of-the-art performance on action recognition and motion completion tasks. Andreas M. Lehrmann, Peter V. Gehler, Sebastian Nowozin |
CVPR | 2 |
| 2014 | Human Pose Estimation with Fields of Parts
Martin Kiefel, Peter V. Gehler |
ECCV (5) | 2 |
| 2014 | Intrinsic Video
Naejin Kong, Peter V. Gehler, Michael J. Black |
ECCV (2) | 2 |
| 2014 | Branch&Rank for Efficient Object DetectionabstractRanking hypothesis sets is a powerful concept for efficient object detection. In this work, we propose a branch&rank scheme that detects objects with often less than 100 ranking operations. This efficiency enables the use of strong and also costly classifiers like non-linear SVMs with RBF- $$\chi ^2$$ kernels. We thereby relieve an inherent limitation of branch&bound methods as bounds are often not tight enough to be effective in practice. Our approach features three key components: a ranking function that operates on sets of hypotheses and a grouping of these into different tasks. Detection efficiency results from adaptively sub-dividing the object search space into decreasingly smaller sets. This is inherited from branch&bound, while the ranking function supersedes a tight bound which is often unavailable (except for rather limited function classes). The grouping makes the system effective: it separates image classification from object recognition, yet combines them in a single formulation, phrased as a structured SVM problem. A novel aspect of branch&rank is that a better ranking function is expected to decrease the number of classifier calls during detection. We use the VOC’07 dataset to demonstrate the algorithmic properties of branch&rank. Alain D. Lehmann, Peter V. Gehler, Luc Van Gool |
Int. J. Comput. Vis. | 2 |
| 2013 | Occlusion Patterns for Object Class DetectionabstractDespite the success of recent object class recognition systems, the long-standing problem of partial occlusion remains a major challenge, and a principled solution is yet to be found. In this paper we leave the beaten path of methods that treat occlusion as just another source of noise - instead, we include the occluder itself into the modelling, by mining distinctive, reoccurring occlusion patterns from annotated training data. These patterns are then used as training data for dedicated detectors of varying sophistication. In particular, we evaluate and compare models that range from standard object class detectors to hierarchical, part-based representations of occluder/occludee pairs. In an extensive evaluation we derive insights that can aid further developments in tackling the occlusion challenge. Bojan Pepik, Michael Stark 0003, Peter V. Gehler, Bernt Schiele |
CVPR | 3 |
| 2013 | Poselet Conditioned Pictorial StructuresabstractIn this paper we consider the challenging problem of articulated human pose estimation in still images. We observe that despite high variability of the body articulations, human motions and activities often simultaneously constrain the positions of multiple body parts. Modelling such higher order part dependencies seemingly comes at a cost of more expensive inference, which resulted in their limited use in state-of-the-art methods. In this paper we propose a model that incorporates higher order part dependencies while remaining efficient. We achieve this by defining a conditional model in which all body parts are connected a-priori, but which becomes a tractable tree-structured pictorial structures model once the image observations are available. In order to derive a set of conditioning variables we rely on the poselet-based features that have been shown to be effective for people detection but have so far found limited application for articulated human pose estimation. We demonstrate the effectiveness of our approach on three publicly available pose estimation benchmarks improving or being on-par with state of the art in each case. Leonid Pishchulin, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
CVPR | 3 |
| 2013 | A Non-parametric Bayesian Network Prior of Human PoseabstractHaving a sensible prior of human pose is a vital ingredient for many computer vision applications, including tracking and pose estimation. While the application of global non-parametric approaches and parametric models has led to some success, finding the right balance in terms of flexibility and tractability, as well as estimating model parameters from data has turned out to be challenging. In this work, we introduce a sparse Bayesian network model of human pose that is non-parametric with respect to the estimation of both its graph structure and its local distributions. We describe an efficient sampling scheme for our model and show its tractability for the computation of exact log-likelihoods. We empirically validate our approach on the Human 3.6M dataset and demonstrate superior performance to global models and parametric networks. We further illustrate our model's ability to represent and compose poses not present in the training set (compositionality) and describe a speed-accuracy trade-off that allows real-time scoring of poses. Andreas M. Lehrmann, Peter V. Gehler, Sebastian Nowozin |
ICCV | 2 |
| 2013 | Strong Appearance and Expressive Spatial Models for Human Pose EstimationabstractTypical approaches to articulated pose estimation combine spatial modelling of the human body with appearance modelling of body parts. This paper aims to push the state-of-the-art in articulated pose estimation in two ways. First we explore various types of appearance representations aiming to substantially improve the body part hypotheses. And second, we draw on and combine several recently proposed powerful ideas such as more flexible spatial models as well as image-conditioned spatial models. In a series of experiments we draw several important conclusions: (1) we show that the proposed appearance representations are complementary, (2) we demonstrate that even a basic tree-structure spatial human body model achieves state-of-the-art performance when augmented with the proper appearance representation, and (3) we show that the combination of the best performing appearance model with a flexible image-conditioned spatial model achieves the best result, significantly improving over the state of the art, on the ``Leeds Sports Poses'' and ``Parse'' benchmarks. Leonid Pishchulin, Mykhaylo Andriluka, Peter V. Gehler, Bernt Schiele |
ICCV | 3 |
| 2012 | Teaching 3D geometry to deformable part modelsabstractCurrent object class recognition systems typically target 2D bounding box localization, encouraged by benchmark data sets, such as Pascal VOC. While this seems suitable for the detection of individual objects, higher-level applications such as 3D scene understanding or 3D object tracking would benefit from more fine-grained object hypotheses incorporating 3D geometric information, such as viewpoints or the locations of individual parts. In this paper, we help narrowing the representational gap between the ideal input of a scene understanding system and object class detector output, by designing a detector particularly tailored towards 3D geometric reasoning. In particular, we extend the successful discriminatively trained deformable part models to include both estimates of viewpoint and 3D parts that are consistent across viewpoints. We experimentally verify that adding 3D geometric information comes at minimal performance loss w.r.t. 2D bounding box localization, but outperforms prior work in 3D viewpoint estimation and ultra-wide baseline matching. Bojan Pepik, Michael Stark 0003, Peter V. Gehler, Bernt Schiele |
CVPR | 3 |
| 2012 | 3D2PM - 3D Deformable Part Models
Bojan Pepik, Peter V. Gehler, Michael Stark 0003, Bernt Schiele |
ECCV (6) | 2 |
| 2011 | Branch&Rank: Non-Linear Object DetectionabstractBranch&rank is an object detection scheme that overcomes the inherent limitation of branch&bound: this method works with arbitrary (classifier) functions whereas tight bounds exist only for simple functions. Objects are usually detected with less than 100 classifier evaluation, which paves the way for using strong (and thus costly) classifiers: We utilize non-linear SVMs with RBF-χ2 kernels without a cascade-like approximation. Our approach features three key components: a ranking function that operates on sets of hypotheses and a grouping of these into different tasks. Detection efficiency results from adaptively sub-dividing the object search space into decreasingly smaller sets. This is inherited from branch&bound, while the ranking function supersedes a tight bound which is often unavailable (except for too simple function classes). The grouping makes the system effective: it separates image classification from object recognition, yet combines them in a single, structured SVM formulation. A novel aspect of branch&rank is that a better ranking function is expected to decrease the number of classifier calls during detection. We demonstrate the algorithmic properties using the VOC'07 dataset. © 2011. The copyright of this document resides with its authors. Alain D. Lehmann, Peter V. Gehler, Luc Van Gool |
BMVC | 2 |
| 2011 | Learning Output Kernels with Block Coordinate Descent
Francesco Dinuzzo, Cheng Soon Ong, Peter V. Gehler, Gianluigi Pillonetto |
ICML | 3 |
| 2011 | Recovering Intrinsic Images with a Global Sparsity Prior on ReflectanceabstractWe address the challenging task of decoupling material properties from lighting properties given a single image. In the last two decades virtually all works have concentrated on exploiting edge information to address this problem. We take a different route by introducing a new prior on reflectance, that models reflectance values as being drawn from a sparse set of basis colors. This results in a Random Field model with global, latent variables (basis colors) and pixel-accurate output reflectance values. We show that without edge information high-quality results can be achieved, that are on par with methods exploiting this source of information. Finally, we present competitive results by integrating an additional edge model. We believe that our approach is a solid starting point for future development in this domain. Peter V. Gehler, Carsten Rother, Martin Kiefel, Lumin Zhang, Bernhard Schölkopf |
NIPS | 1 |
| 2010 | Scene Carving: Scene Consistent Image Retargeting
Alex Mansfield, Peter V. Gehler, Luc Van Gool, Carsten Rother |
ECCV (1) | 2 |
| 2010 | On Parameter Learning in CRF-Based Approaches to Object Class Image Segmentation
Sebastian Nowozin, Peter V. Gehler, Christoph H. Lampert |
ECCV (6) | 2 |
| 2009 | Let the kernel figure it out; Principled learning of pre-processing for kernel classifiersabstractMost modern computer vision systems for high-level tasks, such as image classification, object recognition and segmentation, are based on learning algorithms that are able to separate discriminative information from noise. In practice, however, the typical system consists of a long pipeline of pre-processing steps, such as extraction of different kinds of features, various kinds of normalizations, feature selection, and quantization into aggregated representations such as histograms. Along this pipeline, there are many parameters to set and choices to make, and their effect on the overall system performance is a-priori unclear. In this work, we shorten the pipeline in a principled way. We move pre-processing steps into the learning system by means of kernel parameters, letting the learning algorithm decide upon suitable parameter values. Learning to optimize the pre-processing choices becomes learning the kernel parameters. We realize this paradigm by extending the recent Multiple Kernel Learning formulation from the finite case of having a fixed number of kernels which can be combined to the general infinite case where each possible parameter setting induces an associated kernel. We evaluate the new paradigm extensively on image classification and object classification tasks. We show that it is possible to learn optimal discriminative codebooks and optimal spatial pyramid schemes, consistently outperforming all previous state-of-the-art approaches. Peter V. Gehler, Sebastian Nowozin |
CVPR | 1 |
| 2009 | On feature combination for multiclass object classificationabstractA key ingredient in the design of visual object classification systems is the identification of relevant class specific aspects while being robust to intra-class variations. While this is a necessity in order to generalize beyond a given set of training images, it is also a very difficult problem due to the high variability of visual appearance within each class. In the last years substantial performance gains on challenging benchmark datasets have been reported in the literature. This progress can be attributed to two developments: the design of highly discriminative and robust image features and the combination of multiple complementary features based on different aspects such as shape, color or texture. In this paper we study several models that aim at learning the correct weighting of different features from training data. These include multiple kernel learning as well as simple baseline methods. Furthermore we derive ensemble methods inspired by Boosting which are easily extendable to several multiclass setting. All methods are thoroughly evaluated on object classification datasets using a multitude of feature descriptors. The key results are that even very simple baseline methods, that are orders of magnitude faster than learning techniques are highly competitive with multiple kernel learning. Furthermore the Boosting type methods are found to produce consistently better results in all experiments. We provide insight of when combination methods can be expected to work and how the benefit of complementary features can be exploited most efficiently. Peter V. Gehler, Sebastian Nowozin |
ICCV | 1 |
| 2008 | Bayesian color constancy revisitedabstractComputational color constancy is the task of estimating the true reflectances of visible surfaces in an image. In this paper we follow a line of research that assumes uniform illumination of a scene, and that the principal step in estimating reflectances is the estimation of the scene illuminant. We review recent approaches to illuminant estimation, firstly those based on formulae for normalisation of the reflectance distribution in an image - so-called grey-world algorithms, and those based on a Bayesian formulation of image formation. In evaluating these previous approaches we introduce a new tool in the form of a database of 568 high-quality, indoor and outdoor images, accurately labelled with illuminant, and preserved in their raw form, free of correction or normalisation. This has enabled us to establish several properties experimentally. Firstly automatic selection of grey-world algorithms according to image properties is not nearly so effective as has been thought. Secondly, it is shown that Bayesian illuminant estimation is significantly improved by the improved accuracy of priors for illuminant and reflectance that are obtained from the new dataset. Peter V. Gehler, Carsten Rother, Andrew Blake 0001, Tom Minka, Toby Sharp |
CVPR | 1 |
| 2006 | The rate adapting poisson model for information retrieval and object recognitionabstractProbabilistic modelling of text data in the bag-of-words representation has been dominated by directed graphical models such as pLSI, LDA, NMF, and discrete PCA. Recently, state of the art performance on visual object recognition has also been reported using variants of these models. We introduce an alternative undirected graphical model suitable for modelling count data. This "Rate Adapting Poisson" (RAP) model is shown to generate superior dimensionally reduced representations for subsequent retrieval or classification. Models are trained using contrastive divergence while inference of latent topical representations is efficiently achieved through a simple matrix multiplication. Peter V. Gehler, Alex Holub, Max Welling |
ICML | 1 |
| 2005 | Products of Edge-pertsabstractImages represent an important and abundant source of data. Understanding their statistical structure has important applications such as image compression and restoration. In this paper we propose a particular kind of probabilistic model, dubbed the "products of edge-perts model" to describe the structure of wavelet transformed images. We develop a practical denoising algorithm based on a single edge-pert and show state-ofthe-art denoising performance on benchmark images. Peter V. Gehler, Max Welling |
NIPS | 1 |