Paul Wohlhart

dblp:23/10068 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0002-3669-2809ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-authorSystems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Motion planning and robot control · 24% Transfer learning and domain adaptation · 17% Face, body and person analysis · 11%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.932020
Generalized Feedback Loop for Joint Hand-Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Efficiently Creating 3D Training Data for Fine Hand Pose Estimation · CVPR 2016
Training a Feedback Loop for Hand Pose Estimation · ICCV 2015
Robotics › Motion planning and robot control › robot learning › robot policy learning
generalist robot policy
0.812024
Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration · ICRA 2024
Robotics › Motion planning and robot control
robot learning
0.812024
Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration · ICRA 2024
Robotics › Motion planning and robot control › robot learning
robot policy learning
0.812024
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches · ICLR 2024
Machine learning › Reinforcement learning › generalization in reinforcement learning
task generalization
0.812024
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches · ICLR 2024
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.722019
Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019
Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018
Robotics › Robot manipulation
grasping
0.722019
Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019
Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018
Computer vision › 3D vision
pose estimation
0.722020
Generalized Feedback Loop for Joint Hand-Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Learning descriptors for object recognition and 3D pose estimation · CVPR 2015
Computer vision › Image recognition and object detection
object detection
0.532014
Accurate Object Detection with Joint Classification-Regression Random Forests · CVPR 2014
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013
Alternating Decision Forests · CVPR 2013
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.532014
Accurate Object Detection with Joint Classification-Regression Random Forests · CVPR 2014
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013
Alternating Decision Forests · CVPR 2013
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.522019
Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019
Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018
Computer vision › 3D vision › object pose estimation
hand-object pose estimation
0.412020
Generalized Feedback Loop for Joint Hand-Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Robotics › Robot manipulation
learning from demonstration
0.412020
Watch, Try, Learn: Meta-Learning from Demonstrations and Rewards · ICLR 2020
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412020
Watch, Try, Learn: Meta-Learning from Demonstrations and Rewards · ICLR 2020
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization
0.412019
Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
pixel-level domain adaptation
0.312018
Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018
Machine learning › Representation and self-supervised learning › representation learning › metric learning
mahalanobis distance metric learning
0.322013
Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers · ICCV 2013
Large scale metric learning from equivalence constraints · CVPR 2012
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.322013
Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers · ICCV 2013
Large scale metric learning from equivalence constraints · CVPR 2012
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.322016
Learning descriptors for object recognition and 3D pose estimation · CVPR 2015
Efficiently Creating 3D Training Data for Fine Hand Pose Estimation · CVPR 2016
Computer vision › 3D vision › pose estimation › 3d hand pose estimation
depth-based hand pose estimation
0.212016
Efficiently Creating 3D Training Data for Fine Hand Pose Estimation · CVPR 2016
Robotics › Motion planning and robot control
trajectory representation
0.212024
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches · ICLR 2024
Machine learning › Deep learning architectures and training
feedback loop
0.212015
Training a Feedback Loop for Hand Pose Estimation · ICCV 2015
Computer vision › Image recognition and object detection
object recognition
0.212015
Learning descriptors for object recognition and 3D pose estimation · CVPR 2015
Computer vision › Image recognition and object detection › object detection
bounding box regression
0.212014
Accurate Object Detection with Joint Classification-Regression Random Forests · CVPR 2014
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.212013
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013
Computer vision › Face, body and person analysis
head pose estimation
0.212013
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013
Machine learning › Representation and self-supervised learning
prototype learning
0.212013
Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers · ICCV 2013
Computer vision › Face, body and person analysis
person re-identification
0.112012
Large scale metric learning from equivalence constraints · CVPR 2012
Computer vision › 3D vision
depth image analysis
0.112020
Generalized Feedback Loop for Joint Hand-Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Machine learning › Deep learning architectures and training › data-centric deep learning
training data generation
0.112016
Efficiently Creating 3D Training Data for Fine Hand Pose Estimation · CVPR 2016

Methods — techniques the papers use, named apart from their topics

imitation learning · 1.2convolutional neural network · 0.9reinforcement learning · 0.8transformer policy · 0.8large-scale pretraining · 0.8image generation · 0.8deep network · 0.7meta-learning · 0.4feedback loop · 0.4image-to-image translation · 0.4
YearPublicationVenuePosition
2024 RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
abstract
Generalization remains one of the most important desiderata for robust robot learning systems. While recently proposed approaches show promise in generalization to novel objects, semantic concepts, or visual distribution shifts, generalization to new tasks remains challenging. For example, a language-conditioned policy trained on pick-and-place tasks will not be able to generalize to a folding task, even if the arm trajectory of folding is similar to pick-and-place. Our key insight is that this kind of generalization becomes feasible if we represent the task through rough trajectory sketches. We propose a policy conditioning method using such rough trajectory sketches, which we call RT-Trajectory, that is practical, easy to specify, and allows the policy to effectively perform new tasks that would otherwise be challenging to perform. We find that trajectory sketches strike a balance between being detailed enough to express low-level motion-centric guidance while being coarse enough to allow the learned policy to interpret the trajectory sketch in the context of situational visual observations. In addition, we show how trajectory sketches can provide a useful interface to communicate with robotic policies -- they can be specified through simple human inputs like drawings or videos, or through automated methods such as modern image-generating or waypoint-generating methods. We evaluate RT-Trajectory at scale on a variety of real-world robotic tasks, and find that RT-Trajectory is able to perform a wider range of tasks compared to language-conditioned and goal-conditioned policies, when provided the same training data.
Jiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu 0006, Montse Gonzalez Arenas, Kanishka Rao, Wenhao Yu 0003, Chuyuan Fu, Keerthana Gopalakrishnan, Priya Sundaresan, Peng Xu 0010, Hao Su 0001, Karol Hausman, Chelsea Finn, Ted Xiao
ICLR3
2024 Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration
abstract
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io.
Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin
ICRA180
2020 Watch, Try, Learn: Meta-Learning from Demonstrations and Rewards
Allan Zhou, Eric Jang, Daniel Kappler, Mohi Khansari, Paul Wohlhart, Mrinal Kalakrishnan, Sergey Levine, Chelsea Finn
ICLR6
2020 Generalized Feedback Loop for Joint Hand-Object Pose Estimation
abstract
We propose an approach to estimating the 3D pose of a hand, possibly handling an object, given a depth image. We show that we can correct the mistakes made by a Convolutional Neural Network trained to predict an estimate of the 3D pose by using a feedback loop. The components of this feedback loop are also Deep Networks, optimized using training data. This approach can be generalized to a hand interacting with an object. Therefore, we jointly estimate the 3D pose of the hand and the 3D pose of the object. Our approach performs en-par with state-of-the-art methods for 3D hand pose estimation, and outperforms state-of-the-art methods for joint hand-object pose estimation when using depth images only. Also, our approach is efficient as our implementation runs in real-time on a single GPU.
Markus Oberweger, Paul Wohlhart, Vincent Lepetit
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks
abstract
Real world data, especially in the domain of robotics, is notoriously costly to collect. One way to circumvent this can be to leverage the power of simulation to produce large amounts of labelled data. However, training models on simulated images does not readily transfer to real-world ones. Using domain adaptation methods to cross this "reality gap" requires a large amount of unlabelled real-world data, whilst domain randomization alone can waste modeling power. In this paper, we present Randomized-to-Canonical Adaptation Networks (RCANs), a novel approach to crossing the visual reality gap that uses no real-world data. Our method learns to translate randomized rendered images into their equivalent non-randomized, canonical versions. This in turn allows for real images to also be translated into canonical sim images. We demonstrate the effectiveness of this sim-to-real approach by training a vision-based closed-loop grasping reinforcement learning agent in simulation, and then transferring it to the real world to attain 70% zero-shot grasp success on unseen objects, a result that almost doubles the success of learning the same task directly on domain randomization alone. Additionally, by joint finetuning in the real-world with only 5,000 real-world grasps, our method achieves 91%, attaining comparable performance to a state-of-the-art system trained with 580,000 real-world grasps, resulting in a reduction of real-world data by more than 99%.
Stephen James, Paul Wohlhart, Mrinal Kalakrishnan, Dmitry Kalashnikov, Alex Irpan, Julian Ibarz, Sergey Levine, Raia Hadsell, Konstantinos Bousmalis
CVPR2
2018 Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping
abstract
Instrumenting and collecting annotated visual grasping datasets to train modern machine learning algorithms can be extremely time-consuming and expensive. An appealing alternative is to use off-the-shelf simulators to render synthetic data for which ground-truth annotations are generated automatically. Unfortunately, models trained purely on simulated data often fail to generalize to the real world. We study how randomized simulated environments and domain adaptation methods can be extended to train a grasping system to grasp novel objects from raw monocular RGB images. We extensively evaluate our approaches with a total of more than 25,000 physical test grasps, studying a range of simulation conditions and domain adaptation methods, including a novel extension of pixel-level domain adaptation that we term the GraspGAN. We show that, by using synthetic data and domain adaptation, we are able to reduce the number of real-world samples needed to achieve a given level of performance by up to 50 times, using only randomly generated simulated objects. We also show that by using only unlabeled real-world data and our GraspGAN methodology, we obtain real-world grasping performance without any real-world labels that is similar to that achieved with 939,777 labeled real-world samples.
Konstantinos Bousmalis, Alex Irpan, Paul Wohlhart, Matthew Kelcey, Mrinal Kalakrishnan, Laura Downs, Julian Ibarz, Peter Pastor, Kurt Konolige, Sergey Levine, Vincent Vanhoucke
ICRA3
2016 Efficiently Creating 3D Training Data for Fine Hand Pose Estimation
abstract
While many recent hand pose estimation methods critically rely on a training set of labelled frames, the creation of such a dataset is a challenging task that has been overlooked so far. As a result, existing datasets are limited to a few sequences and individuals, with limited accuracy, and this prevents these methods from delivering their full potential. We propose a semi-automated method for efficiently and accurately labeling each frame of a hand depth video with the corresponding 3D locations of the joints: The user is asked to provide only an estimate of the 2D reprojections of the visible joints in some reference frames, which are automatically selected to minimize the labeling work by efficiently optimizing a sub-modular loss function. We then exploit spatial, temporal, and appearance constraints to retrieve the full 3D poses of the hand over the complete sequence. We show that this data can be used to train a recent state-of-the-art hand pose estimation method, leading to increased accuracy.
Markus Oberweger, Gernot Riegler, Paul Wohlhart, Vincent Lepetit
CVPR3
2015 Learning descriptors for object recognition and 3D pose estimation
abstract
Detecting poorly textured objects and estimating their 3D pose reliably is still a very challenging problem. We introduce a simple but powerful approach to computing descriptors for object views that efficiently capture both the object identity and 3D pose. By contrast with previous manifold-based approaches, we can rely on the Euclidean distance to evaluate the similarity between descriptors, and therefore use scalable Nearest Neighbor search methods to efficiently handle a large number of objects under a large range of poses. To achieve this, we train a Convolutional Neural Network to compute these descriptors by enforcing simple similarity and dissimilarity constraints between the descriptors. We show that our constraints nicely untangle the images from different objects and different views into clusters that are not only well-separated but also structured as the corresponding sets of poses: The Euclidean distance between descriptors is large when the descriptors are from different objects, and directly related to the distance between the poses when the descriptors are from the same object. These important properties allow us to outperform state-of-the-art object views representations on challenging RGB and RGB-D data.
Paul Wohlhart, Vincent Lepetit
CVPR1
2015 Training a Feedback Loop for Hand Pose Estimation
abstract
We propose an entirely data-driven approach to estimating the 3D pose of a hand given a depth image. We show that we can correct the mistakes made by a Convolutional Neural Network trained to predict an estimate of the 3D pose by using a feedback loop. The components of this feedback loop are also Deep Networks, optimized using training data. They remove the need for fitting a 3D model to the input data, which requires both a carefully designed fitting function and algorithm. We show that our approach outperforms state-of-the-art methods, and is efficient as our implementation runs at over 400 fps on a single GPU.
Markus Oberweger, Paul Wohlhart, Vincent Lepetit
ICCV2
2015 You Should Use Regression to Detect Cells
Philipp Kainz, Martin Urschler, Samuel Schulter, Paul Wohlhart, Vincent Lepetit
MICCAI (3)4
2014 Accurate Object Detection with Joint Classification-Regression Random Forests
abstract
In this paper, we present a novel object detection approach that is capable of regressing the aspect ratio of objects. This results in accurately predicted bounding boxes having high overlap with the ground truth. In contrast to most recent works, we employ a Random Forest for learning a template-based model but exploit the nature of this learning algorithm to predict arbitrary output spaces. In this way, we can simultaneously predict the object probability of a window in a sliding window approach as well as regress its aspect ratio with a single model. Furthermore, we also exploit the additional information of the aspect ratio during the training of the Joint Classification-Regression Random Forest, resulting in better detection models. Our experiments demonstrate several benefits: (i) Our approach gives competitive results on standard detection benchmarks. (ii) The additional aspect ratio regression delivers more accurate bounding boxes than standard object detection approaches in terms of overlap with ground truth, especially when tightening the evaluation criterion. (iii) The detector itself becomes better by only including the aspect ratio information during training.
Samuel Schulter, Christian Leistner, Paul Wohlhart, Peter M. Roth, Horst Bischof
CVPR3
2013 Alternating Decision Forests
abstract
This paper introduces a novel classification method termed Alternating Decision Forests (ADFs), which formulates the training of Random Forests explicitly as a global loss minimization problem. During training, the losses are minimized via keeping an adaptive weight distribution over the training samples, similar to Boosting methods. In order to keep the method as flexible and general as possible, we adopt the principle of employing gradient descent in function space, which allows to minimize arbitrary losses. Contrary to Boosted Trees, in our method the loss minimization is an inherent part of the tree growing process, thus allowing to keep the benefits of common Random Forests, such as, parallel processing. We derive the new classifier and give a discussion and evaluation on standard machine learning data sets. Furthermore, we show how ADFs can be easily integrated into an object detection application. Compared to both, standard Random Forests and Boosted Trees, ADFs give better performance in our experiments, while yielding more compact models in terms of tree depth.
Samuel Schulter, Paul Wohlhart, Christian Leistner, Amir Saffari, Peter M. Roth, Horst Bischof
CVPR2
2013 Optimizing 1-Nearest Prototype Classifiers
abstract
The development of complex, powerful classifiers and their constant improvement have contributed much to the progress in many fields of computer vision. However, the trend towards large scale datasets revived the interest in simpler classifiers to reduce runtime. Simple nearest neighbor classifiers have several beneficial properties, such as low complexity and inherent multi-class handling, however, they have a runtime linear in the size of the database. Recent related work represents data samples by assigning them to a set of prototypes that partition the input feature space and afterwards applies linear classifiers on top of this representation to approximate decision boundaries locally linear. In this paper, we go a step beyond these approaches and purely focus on 1-nearest prototype classification, where we propose a novel algorithm for deriving optimal prototypes in a discriminative manner from the training samples. Our method is implicitly multi-class capable, parameter free, avoids noise over fitting and, since during testing only comparisons to the derived prototypes are required, highly efficient. Experiments demonstrate that we are able to outperform related locally linear methods, while even getting close to the results of more complex classifiers.
Paul Wohlhart, Martin Köstinger, Michael Donoser, Peter M. Roth, Horst Bischof
CVPR1
2013 Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers
abstract
In this paper, we raise important issues concerning the evaluation complexity of existing Mahalanobis metric learning methods. The complexity scales linearly with the size of the dataset. This is especially cumbersome on large scale or for real-time applications with limited time budget. To alleviate this problem we propose to represent the dataset by a fixed number of discriminative prototypes. In particular, we introduce a new method that jointly chooses the positioning of prototypes and also optimizes the Mahalanobis distance metric with respect to these. We show that choosing the positioning of the prototypes and learning the metric in parallel leads to a drastically reduced evaluation effort while maintaining the discriminative essence of the original dataset. Moreover, for most problems our method performing k-nearest prototype (k-NP) classification on the condensed dataset leads to even better generalization compared to k-NN classification using all data. Results on a variety of challenging benchmarks demonstrate the power of our method. These include standard machine learning datasets as well as the challenging Public Figures Face Database. On the competitive machine learning benchmarks we are comparable to the state-of-the-art while being more efficient. On the face benchmark we clearly outperform the state-of-the-art in Mahalanobis metric learning with drastically reduced evaluation effort.
Martin Köstinger, Paul Wohlhart, Peter M. Roth, Horst Bischof
ICCV2
2013 Alternating Regression Forests for Object Detection and Pose Estimation
abstract
We present Alternating Regression Forests (ARFs), a novel regression algorithm that learns a Random Forest by optimizing a global loss function over all trees. This interrelates the information of single trees during the training phase and results in more accurate predictions. ARFs can minimize any differentiable regression loss without sacrificing the appealing properties of Random Forests, like low computational complexity during both, training and testing. Inspired by recent developments for classification [19], we derive a new algorithm capable of dealing with different regression loss functions, discuss its properties and investigate the relations to other methods like Boosted Trees. We evaluate ARFs on standard machine learning benchmarks, where we observe better generalization power compared to both standard Random Forests and Boosted Trees. Moreover, we apply the proposed regressor to two computer vision applications: object detection and head pose estimation from depth images. ARFs outperform the Random Forest baselines in both tasks, illustrating the importance of optimizing a common loss function for all trees.
Samuel Schulter, Christian Leistner, Paul Wohlhart, Peter M. Roth, Horst Bischof
ICCV3
2012 Detecting Partially Occluded Objects with an Implicit Shape Model Random Field
Paul Wohlhart, Michael Donoser, Peter M. Roth, Horst Bischof
ACCV (1)1
2012 Discriminative Hough Forests for Object Detection
abstract
Object detection models based on the Implicit Shape Model (ISM) [3] use small, local parts that vote for object centers in images. Since these parts vote completely independently from each other, this often leads to false-positive detections due to random constellations of parts. Thus, we introduce a verification step, which considers the activations of all voting elements that contribute to a detection. The levels of activation of each voting element of the ISM form a new description vector for an object hypothesis, which can be examined in order to discriminate between correct and incorrect detections. In particular, we observe the levels of activation of the voting elements in Hough Forests [2], which can be seen as a variant of ISM. In Hough Forests, the voting elements are all the positive training patches used to train the Forest. Each patch of the input image is classified by all decision trees in the Hough Forest. Whenever an input patch falls into the same leaf node as a patch from training, a certain amount of weight is added to the detection hypothesis at the relative position of the object center, which was recorded when cropping out the training patch. The total amount of weight one voting element (offset vector) adds to a detection hypothesis (the total activation) can be calculated by summing over all input patches and trees in the forest. Stacking the activations of all elements gives an activation vector for a hypothesis. We learn classifiers to discriminate correct and wrong part constellations based on these activation vectors and thus assign a better confidence to each detection. We use linear models as well as a histogram intersection kernel SVM. In the linear classifier, one weight is learned for each voting element. We additionally show how to use these weights, not only as a post processing step, but directly in the voting process. This has two advantages: First, it circumvents the explicit calculation of the activation vector for later reclassification, which is computationally more demanding. Second, the non-maxima suppression is performed on cleaner Hough maps, which allows for reducing the size of the suppression neighborhood and thus increases the recall at high levels of precision.
Paul Wohlhart, Samuel Schulter, Martin Köstinger, Peter M. Roth, Horst Bischof
BMVC1
2012 Large scale metric learning from equivalence constraints
abstract
In this paper, we raise important issues on scalability and the required degree of supervision of existing Mahalanobis metric learning methods. Often rather tedious optimization procedures are applied that become computationally intractable on a large scale. Further, if one considers the constantly growing amount of data it is often infeasible to specify fully supervised labels for all data points. Instead, it is easier to specify labels in form of equivalence constraints. We introduce a simple though effective strategy to learn a distance metric from equivalence constraints, based on a statistical inference perspective. In contrast to existing methods we do not rely on complex optimization problems requiring computationally expensive iterations. Hence, our method is orders of magnitudes faster than comparable methods. Results on a variety of challenging benchmarks with rather diverse nature demonstrate the power of our method. These include faces in unconstrained environments, matching before unseen object instances and person re-identification across spatially disjoint cameras. In the latter two benchmarks we clearly outperform the state-of-the-art.
Martin Köstinger, Martin Hirzer, Paul Wohlhart, Peter M. Roth, Horst Bischof
CVPR3
2011 Learning to recognize faces from videos and weakly related information cues
abstract
Videos are often associated with additional information that could be valuable for interpretation of their content. This especially applies for the recognition of faces within video streams, where often cues such as transcripts and subtitles are available. However, this data is not completely reliable and might be ambiguously labeled. To overcome these limitations, we take advantage of semi-supervised (SSL) and multiple instance learning (MIL) and propose a new semi-supervised multiple instance learning (SSMIL) algorithm. Thus, during training we can weaken the prerequisite of knowing the label for each instance and can integrate unlabeled data, given only probabilistic information in form of priors. The benefits of the approach are demonstrated for face recognition in videos on a publicly available benchmark dataset. In fact, we show exploring new information sources can considerably improve the classification results.
Martin Köstinger, Paul Wohlhart, Peter M. Roth, Horst Bischof
AVSS2
2010 Automatic Detection and Reading of Dangerous Goods Plates
abstract
In this paper, we present an efficient solution for automatic detection and reading of dangerous goods plates on trucks and trains. According to the ADR agreement dangerous goods transports are marked with an orange plate covering the hazard class and the identification number for the hazardous substances. Since under real-world conditions high resolution images (often at low quality) have to be processed an efficient and robust system is required. In particular, we propose a multi-stage system consisting of an acquisition step, a saliency region detector (to reduce the run-time), a plate detector, and a robust recognition step based on an Optical Character Recognition (OCR). To demonstrate the system, we show qualitative and quantitative localization/recognition results on two challenging data sets. In fact, building on proven robust and efficient methods, we show excellent detection and classification results under hard environmental conditions at low run-time.
Peter M. Roth, Martin Köstinger, Paul Wohlhart, Horst Bischof, Josef A. Birchbauer
AVSS3