Peter M. Roth

dblp:18/4857 · DBLP profile ↗
← Back
53ranked-venue papers
3as first author
5since 2021 · last 2022
0000-0001-9566-1298ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 36 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
3D vision · 26% Image recognition and object detection · 14% Face, body and person analysis · 14%
Computer graphics and multimedia
3 papers
Virtual and augmented reality · 54% Rendering · 46%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
1.252022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Accurate Object Detection with Joint Classification-Regression Random Forests · CVPR 2014
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
1.132020
Geometric Correspondence Fields: Learned Differentiable Rendering for 3D Pose Refinement in the Wild · ECCV (16) 2020
GP2C: Geometric Projection Parameter Consensus for Joint 3D Pose and Focal Length Estimation in the Wild · ICCV 2019
3D Pose Estimation and 3D Model Retrieval for Objects in the Wild · CVPR 2018
Virtual and augmented reality
mixed reality
0.822021
Neural Cameras: Learning Camera Characteristics for Coherent Mixed Reality Rendering · ISMAR 2021
Learning Lightprobes for Mixed Reality Illumination · ISMAR 2017
Computer vision › 3D vision
3d object detection
0.612022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Machine learning › Trustworthy machine learning
interpretability
0.612022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection
0.612022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map
0.612022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.532014
Accurate Object Detection with Joint Classification-Regression Random Forests · CVPR 2014
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013
Alternating Decision Forests · CVPR 2013
Rendering › physically based rendering › wave optics rendering
coherent rendering
0.512021
Neural Cameras: Learning Camera Characteristics for Coherent Mixed Reality Rendering · ISMAR 2021
Computer vision › 3D vision › pose estimation
pose refinement
0.412020
Geometric Correspondence Fields: Learned Differentiable Rendering for 3D Pose Refinement in the Wild · ECCV (16) 2020
Rendering
differentiable rendering
0.412020
Geometric Correspondence Fields: Learned Differentiable Rendering for 3D Pose Refinement in the Wild · ECCV (16) 2020
Medical and health informatics › medical imaging
medical image analysis
0.412019
Biomedical image augmentation using Augmentor · Bioinform. 2019
Computer vision › Video understanding and tracking
multi-object tracking
0.422014
Occlusion Geodesics for Online Multi-object Tracking · CVPR 2014
Robust Real-Time Tracking of Multiple Objects by Volumetric Mass Densities · CVPR 2013
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.432013
Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers · ICCV 2013
Large scale metric learning from equivalence constraints · CVPR 2012
Relaxed Pairwise Learned Metric for Person Re-identification · ECCV (6) 2012
Computer vision › 3D vision › 3d shape analysis
3d shape retrieval
0.312018
3D Pose Estimation and 3D Model Retrieval for Objects in the Wild · CVPR 2018
Computer vision › 3D vision › object pose estimation › 6d object pose estimation
category-level object pose estimation
0.312018
3D Pose Estimation and 3D Model Retrieval for Objects in the Wild · CVPR 2018
Computer vision › Video understanding and tracking › multi-object tracking
tracking-by-detection
0.322014
Occlusion Geodesics for Online Multi-object Tracking · CVPR 2014
Hough-based tracking of non-rigid objects · ICCV 2011
Machine learning › Representation and self-supervised learning › representation learning › metric learning
mahalanobis distance metric learning
0.322013
Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers · ICCV 2013
Large scale metric learning from equivalence constraints · CVPR 2012
Computer vision › 3D vision › visual localization
geo-localization
0.312017
Learning to Align Semantic Segmentation and 2.5D Maps for Geolocalization · CVPR 2017
Robotics › Robot navigation and mapping
localization
0.312017
Learning to Align Semantic Segmentation and 2.5D Maps for Geolocalization · CVPR 2017
Computer vision › Face, body and person analysis
person re-identification
0.322012
Relaxed Pairwise Learned Metric for Person Re-identification · ECCV (6) 2012
Large scale metric learning from equivalence constraints · CVPR 2012
Computer vision › Segmentation and scene understanding
semantic segmentation
0.312017
Learning to Align Semantic Segmentation and 2.5D Maps for Geolocalization · CVPR 2017
Virtual and augmented reality › tracking and registration
photometric registration
0.312017
Learning Lightprobes for Mixed Reality Illumination · ISMAR 2017
Computer vision › Image recognition and object detection › object detection
bounding box regression
0.212014
Accurate Object Detection with Joint Classification-Regression Random Forests · CVPR 2014
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.212014
Occlusion Geodesics for Online Multi-object Tracking · CVPR 2014
Computer vision › Video understanding and tracking › multi-object tracking
online multi-object tracking
0.212014
Occlusion Geodesics for Online Multi-object Tracking · CVPR 2014
Computer vision › 3D vision › point cloud processing
LiDAR point cloud
0.212022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Robotics › Autonomous driving
perception
0.212022
OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data · CVPR 2022
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.212013
Robust Real-Time Tracking of Multiple Objects by Volumetric Mass Densities · CVPR 2013
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.212013
Alternating Regression Forests for Object Detection and Pose Estimation · ICCV 2013

Methods — techniques the papers use, named apart from their topics

geometric correspondence fields · 0.9differentiable rendering · 0.9deep learning · 0.8convolutional neural network · 0.6subsampling · 0.6perturbation-based attribution · 0.6neural network · 0.5image database · 0.5stochastic augmentation · 0.4reprojection error minimization · 0.4geometric optimization · 0.4boosting · 0.3metric learning · 0.3synthetic training data · 0.3
YearPublicationVenuePosition
2022 OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR Data
abstract
While 3D object detection in LiDAR point clouds is well-established in academia and industry, the explainability of these models is a largely unexplored field. In this paper, we propose a method to generate attribution maps for the detected objects in order to better understand the behavior of such models. These maps indicate the importance of each 3D point in predicting the specific objects. Our method works with black-box models: We do not require any prior knowledge of the architecture nor access to the model's internals, like parameters, activations or gradients. Our efficient perturbation-based approach empirically estimates the importance of each point by testing the model with randomly generated subsets of the input point cloud. Our sub-sampling strategy takes into account the special characteristics of LiDAR data, such as the depth-dependent point density. We show a detailed evaluation of the attribution maps and demonstrate that they are interpretable and highly informative. Furthermore, we compare the attribution maps of recent 3D object detection architectures to provide insights into their decision-making processes.
David Schinagl, Georg Krispel, Horst Possegger, Peter M. Roth, Horst Bischof
CVPR4
2022 Counting Everything in Remote Sensing - the Need for Benchmarks
abstract
Object counting and object density estimation is a basic technology serving many applications. For that reason, a multitude of benchmarks are existing in the computer vision community allowing to evaluate and compare different solutions for object counting. However, in remote sensing, such benchmarks are not available, which hinders fair evaluations between different methods and paradigms. This work should lead to discussions on how to optimally design and set up object counting benchmarks in remote sensing for multiple object categories and also for various satellite data, including multi-resolution, multispectral, and multi-modal data from optical and SAR sensors.
Roland Perko, Sead Mustafic, Alexander Almer, Peter M. Roth
IGARSS4
2021 Protocol Design Issues for Object Density Estimation and Counting in Remote Sensing
abstract
Generic object density estimation and counting plays an important role in various remote sensing applications. Therefore, this work discusses available Earth observation based datasets for both tasks. Critical aspects of the current evaluation protocol are identified and analyzed, which yields novel design issues for an optimal future protocol. Moreover, the parameters for an ideal benchmark dataset are defined which allows appropriate evaluation and ranking of methodologies.
Roland Perko, Alexander Almer, Mario Theuermann, Manfred Klopschitz, Thomas Schnabel, Peter M. Roth
IGARSS6
2021 Neural Cameras: Learning Camera Characteristics for Coherent Mixed Reality Rendering
abstract
Coherent rendering is important for generating plausible Mixed Reality presentations of virtual objects within a user’s real-world environment. Besides photo-realistic rendering and correct lighting, visual coherence requires simulating the imaging system that is used to capture the real environment. While existing approaches either focus on a specific camera or a specific component of the imaging system, we introduce Neural Cameras, the first approach that jointly simulates all major components of an arbitrary modern camera using neural networks. Our system allows for adding new cameras to the framework by learning the visual properties from a database of images that has been captured using the physical camera. We present qualitative and quantitative results and discuss future direction for research that emerge from using Neural Cameras.
David Mandl, Peter M. Roth, Tobias Langlotz, Christoph Ebner, Shohei Mori, Stefanie Zollmann, Peter Mohr, Denis Kalkofen
ISMAR2
2021 Performing arithmetic using a neural network trained on images of digit permutation pairs
abstract
Abstract In this paper, a neural network is trained to perform simple arithmetic using images of concatenated handwritten digit pairs. A convolutional neural network was trained with images consisting of two side-by-side handwritten digits, where the image’s label is the summation of the two digits contained in the combined image. Crucially, the network was tested on permutation pairs that were not present during training in an effort to see if the network could learn the task of addition, as opposed to simply mapping images to labels. A dataset was generated for all possible permutation pairs of length 2 for the digits 0–9 using MNIST as a basis for the images, with one thousand samples generated for each permutation pair. For testing the network, samples generated from previously unseen permutation pairs were fed into the trained network, and its predictions measured. Results were encouraging, with the network achieving an accuracy of over 90% on some permutation train/test splits. This suggests that the network learned at first digit recognition, and subsequently the further task of addition based on the two recognised digits. As far as the authors are aware, no previous work has concentrated on learning a mathematical operation in this way. This paper is an attempt to demonstrate that a network can learn more than a direct mapping from image to label, but is learning to analyse two separate regions of an image and combining what was recognised to produce the final output label.
Marcus D. Bloice, Peter M. Roth, Andreas Holzinger
J. Intell. Inf. Syst.2
2020 Geometric Correspondence Fields: Learned Differentiable Rendering for 3D Pose Refinement in the Wild
Alexander Grabner, Yaming Wang, Peizhao Zhang, Peihong Guo, Tong Xiao 0003, Peter Vajda, Peter M. Roth, Vincent Lepetit
ECCV (16)7
2020 Performing Arithmetic Using a Neural Network Trained on Digit Permutation Pairs
Marcus D. Bloice, Peter M. Roth, Andreas Holzinger
ISMIS2
2020 L*ReLU: Piece-wise Linear Activation Functions for Deep Fine-grained Visual Categorization
abstract
Deep neural networks paved the way for significant improvements in image visual categorization during the last years. However, even though the tasks are highly varying, differing in complexity and difficulty, existing solutions mostly build on the same architectural decisions. This also applies to the selection of activation functions (AFs), where most approaches build on Rectified Linear Units (Re- LUs). In this paper, however, we show that the choice of a proper AF has a significant impact on the classification accuracy, in particular, if fine, subtle details are of relevance. Therefore, we propose to model the degree of absence and the degree presence of features via the AF by using piecewise linear functions, which we refer to as L*ReLU. In this way, we can ensure the required properties, while still inheriting the benefits in terms of computational efficiency from ReLUs. We demonstrate our approach for the task of Fine-grained Visual Categorization (FGVC), running experiments on seven different benchmark datasets. The results do not only demonstrate superior results but also that for different tasks, having different characteristics, different AFs are selected.
Mina Basirat, Peter M. Roth
WACV2
2020 Smart Hypothesis Generation for Efficient and Robust Room Layout Estimation
abstract
We propose a novel method to efficiently estimate the spatial layout of a room from a single monocular RGB image. As existing approaches based on low-level feature extraction, followed by a vanishing point estimation are very slow and often unreliable in realistic scenarios, we build on semantic segmentation of the input image. To obtain better segmentations, we introduce a robust, accurate and very efficient hypothesize-and-test scheme. The key idea is to use three segmentation hypotheses, each based on a different number of visible walls. For each hypothesis, we predict the image locations of the room corners and select the hypothesis for which the layout estimated from the room corners is consistent with the segmentation. We demonstrate the efficiency and robustness of our method on three challenging benchmark datasets, where we significantly outperform the state-of-the-art.
Martin Hirzer, Peter M. Roth, Vincent Lepetit
WACV2
2020 ALCN: Adaptive Local Contrast Normalization
Mahdi Rad, Peter M. Roth, Vincent Lepetit
Comput. Vis. Image Underst.2
2019 Location Field Descriptors: Single Image 3D Model Retrieval in the Wild
abstract
We present Location Field Descriptors, a novel approach for single image 3D model retrieval in the wild. In contrast to previous methods that directly map 3D models and RGB images to an embedding space, we establish a common low-level representation in the form of location fields from which we compute pose invariant 3D shape descriptors. Location fields encode correspondences between 2D pixels and 3D surface coordinates and, thus, explicitly capture 3D shape and 3D pose information without appearance variations which are irrelevant for the task. This early fusion of 3D models and RGB images results in three main advantages: First, the bottleneck location field prediction acts as a regularizer during training. Second, major parts of the system benefit from training on a virtually infinite amount of synthetic data. Finally, the predicted location fields are visually interpretable and unblackbox the system. We evaluate our proposed approach on three challenging real-world datasets (Pix3D, Comp, and Stanford) with different object categories and significantly outperform the state-of-the-art by up to 20% absolute in multiple 3D retrieval metrics.
Alexander Grabner, Peter M. Roth, Vincent Lepetit
3DV2
2019 GP2C: Geometric Projection Parameter Consensus for Joint 3D Pose and Focal Length Estimation in the Wild
abstract
We present a joint 3D pose and focal length estimation approach for object categories in the wild. In contrast to previous methods that predict 3D poses independently of the focal length or assume a constant focal length, we explicitly estimate and integrate the focal length into the 3D pose estimation. For this purpose, we combine deep learning techniques and geometric algorithms in a two-stage approach: First, we estimate an initial focal length and establish 2D-3D correspondences from a single RGB image using a deep network. Second, we recover 3D poses and refine the focal length by minimizing the reprojection error of the predicted correspondences. In this way, we exploit the geometric prior given by the focal length for 3D pose estimation. This results in two advantages: First, we achieve significantly improved 3D translation and 3D pose accuracy compared to existing methods. Second, our approach finds a geometric consensus between the individual projection parameters, which is required for precise 2D-3D alignment. We evaluate our proposed approach on three challenging real-world datasets (Pix3D, Comp, and Stanford) with different object categories and significantly outperform the state-of-the-art by up to 20% absolute in multiple different metrics.
Alexander Grabner, Peter M. Roth, Vincent Lepetit
ICCV2
2019 Multiple View Geometry in Remote Sensing: An Empirical Study Based on Pléiades Satellite Images
abstract
In contrast to the fields of computer vision and photogrammetry, multiple view geometry has not been extensively exploited in the remote sensing domain so far. Therefore, an empirical study is conducted based on multi view Pléiades data that depicts a scene from multiple orbits and multiple incidence angles. First, an accuracy analysis of the 2D and 3D geo-location performance is elaborated showing that ground control points can be modelled with a root mean square residual error below 30 cm in East, North, and height. Second, digital surface models are reconstructed from all possible stereo pairs and are additionally fused in the multiple view geometry sense. It is shown that employing more data increases the accuracy of the digital surface model while reducing the amount of the non-reconstructed regions.
Roland Perko, Mathias Schardt, Livia Piermattei, Stefan Auer, Peter M. Roth
IGARSS5
2019 Biomedical image augmentation using Augmentor
abstract
MOTIVATION: Image augmentation is a frequently used technique in computer vision and has been seeing increased interest since the popularity of deep learning. Its usefulness is becoming more and more recognized due to deep neural networks requiring larger amounts of data to train, and because in certain fields, such as biomedical imaging, large amounts of labelled data are difficult to come by or expensive to produce. In biomedical imaging, features specific to this domain need to be addressed. RESULTS: Here we present the Augmentor software package for image augmentation. It provides a stochastic, pipeline-based approach to image augmentation with a number of features that are relevant to biomedical imaging, such as z-stack augmentation and randomized elastic distortions. The software has been designed to be highly extensible meaning an operation that might be specific to a highly specialized task can easily be added to the library, even at runtime. Although it has been designed as a general software library, it has features that are particularly relevant to biomedical imaging and the techniques required for this domain. AVAILABILITY AND IMPLEMENTATION: Augmentor is a Python package made available under the terms of the MIT licence. Source code can be found on GitHub under https://github.com/mdbloice/Augmentor and installation is via the pip package manager (A Julia version of the package, developed in parallel by Christof Stocker, is also available under https://github.com/Evizero/Augmentor.jl).
Marcus D. Bloice, Peter M. Roth, Andreas Holzinger
Bioinform.2
2018 3D Pose Estimation and 3D Model Retrieval for Objects in the Wild
abstract
We propose a scalable, efficient and accurate approach to retrieve 3D models for objects in the wild. Our contribution is twofold. We first present a 3D pose estimation approach for object categories which significantly outperforms the state-of-the-art on Pascal3D+. Second, we use the estimated pose as a prior to retrieve 3D models which accurately represent the geometry of objects in RGB images. For this purpose, we render depth images from 3D models under our predicted pose and match learned image descriptors of RGB images against those of rendered depth images using a CNN-based multi-view metric learning approach. In this way, we are the first to report quantitative results for 3D model retrieval on Pascal3D+, where our method chooses the same models as human annotators for 50% of the validation images on average. In addition, we show that our method, which was trained purely on Pascal3D+, retrieves rich and accurate 3D models from ShapeNet given RGB images of objects in the wild.
Alexander Grabner, Peter M. Roth, Vincent Lepetit
CVPR2
2017 Accurate Camera Registration in Urban Environments Using High-Level Feature Matching
Anil Armagan, Martin Hirzer, Peter M. Roth, Vincent Lepetit
BMVC3
2017 Efficient 3D Tracking in Urban Environments with Semantic Segmentation
Martin Hirzer, Peter M. Roth, Vincent Lepetit
BMVC2
2017 Adaptive Local Contrast Normalization for Robust Object Detection and Pose Estimation
Mahdi Rad, Vincent Lepetit, Peter M. Roth
BMVC3
2017 Learning to Align Semantic Segmentation and 2.5D Maps for Geolocalization
abstract
We present an efficient method for geolocalization in urban environments starting from a coarse estimate of the location provided by a GPS and using a simple untextured 2.5D model of the surrounding buildings. Our key contribution is a novel efficient and robust method to optimize the pose: We train a Deep Network to predict the best direction to improve a pose estimate, given a semantic segmentation of the input image and a rendering of the buildings from this estimate. We then iteratively apply this CNN until converging to a good pose. This approach avoids the use of reference images of the surroundings, which are difficult to acquire and match, while 2.5D models are broadly available. We can therefore apply it to places unseen during training.
Anil Armagan, Martin Hirzer, Peter M. Roth, Vincent Lepetit
CVPR3
2017 Learning Lightprobes for Mixed Reality Illumination
abstract
This paper presents the first photometric registration pipeline for Mixed Reality based on high quality illumination estimation using convolutional neural networks (CNNs). For easy adaptation and deployment of the system, we train the CNNs using purely synthetic images and apply them to real image data. To keep the pipeline accurate and efficient, we propose to fuse the light estimation results from multiple CNN instances and show an approach for caching estimates over time. For optimal performance, we furthermore explore multiple strategies for the CNN training. Experimental results show that the proposed method yields highly accurate estimates for photo-realistic augmentations.
David Mandl, Kwang Moo Yi, Peter Mohr, Peter M. Roth, Pascal Fua, Vincent Lepetit, Dieter Schmalstieg, Denis Kalkofen
ISMAR4
2014 Occlusion Geodesics for Online Multi-object Tracking
abstract
Robust multi-object tracking-by-detection requires the correct assignment of noisy detection results to object trajectories. We address this problem by proposing an online approach based on the observation that object detectors primarily fail if objects are significantly occluded. In contrast to most existing work, we only rely on geometric information to efficiently overcome detection failures. In particular, we exploit the spatio-temporal evolution of occlusion regions, detector reliability, and target motion prediction to robustly handle missed detections. In combination with a conservative association scheme for visible objects, this allows for real-time tracking of multiple objects from a single static camera, even in complex scenarios. Our evaluations on publicly available multi-object tracking benchmark datasets demonstrate favorable performance compared to the state-of-the-art in online and offline multi-object tracking.
Horst Possegger, Thomas Mauthner, Peter M. Roth, Horst Bischof
CVPR3
2014 Accurate Object Detection with Joint Classification-Regression Random Forests
abstract
In this paper, we present a novel object detection approach that is capable of regressing the aspect ratio of objects. This results in accurately predicted bounding boxes having high overlap with the ground truth. In contrast to most recent works, we employ a Random Forest for learning a template-based model but exploit the nature of this learning algorithm to predict arbitrary output spaces. In this way, we can simultaneously predict the object probability of a window in a sliding window approach as well as regress its aspect ratio with a single model. Furthermore, we also exploit the additional information of the aspect ratio during the training of the Joint Classification-Regression Random Forest, resulting in better detection models. Our experiments demonstrate several benefits: (i) Our approach gives competitive results on standard detection benchmarks. (ii) The additional aspect ratio regression delivers more accurate bounding boxes than standard object detection approaches in terms of overlap with ground truth, especially when tightening the evaluation criterion. (iii) The detector itself becomes better by only including the aspect ratio information during training.
Samuel Schulter, Christian Leistner, Paul Wohlhart, Peter M. Roth, Horst Bischof
CVPR4
2013 Unsupervised Object Discovery and Segmentation in Videos
abstract
Unsupervised object discovery is the task of finding recurring objects over an unsorted set of images without any human supervision, which becomes more and more important as the amount of visual data grows exponentially. Existing approaches typically build on still images and rely on different prior knowledge to yield accurate results. In contrast, we propose a novel video-based approach, allowing also for exploiting motion information, which is a strong and physically valid indicator for foreground objects, thus, tremendously easing the task. In particular, we show how to integrate motion information in parallel with appearance cues into a common conditional random field formulation to automatically discover object categories from videos. In the experiments, we show that our system can successfully extract, group, and segment most foreground objects and is also able to discover stationary objects in the given videos. Furthermore, we demonstrate that the unsupervised learned appearance models also yield reasonable results for object detection on still images.
Samuel Schulter, Christian Leistner, Peter M. Roth, Horst Bischof
BMVC3
2013 Robust Real-Time Tracking of Multiple Objects by Volumetric Mass Densities
abstract
Combining foreground images from multiple views by projecting them onto a common ground-plane has been recently applied within many multi-object tracking approaches. These planar projections introduce severe artifacts and constrain most approaches to objects moving on a common 2D ground-plane. To overcome these limitations, we introduce the concept of an occupancy volume - exploiting the full geometry and the objects' center of mass - and develop an efficient algorithm for 3D object tracking. Individual objects are tracked using the local mass density scores within a particle filter based approach, constrained by a Voronoi partitioning between nearby trackers. Our method benefits from the geometric knowledge given by the occupancy volume to robustly extract features and train classifiers on-demand, when volumetric information becomes unreliable. We evaluate our approach on several challenging real-world scenarios including the public APIDIS dataset. Experimental evaluations demonstrate significant improvements compared to state-of-the-art methods, while achieving real-time performance.
Horst Possegger, Sabine Sternig, Thomas Mauthner, Peter M. Roth, Horst Bischof
CVPR4
2013 Alternating Decision Forests
abstract
This paper introduces a novel classification method termed Alternating Decision Forests (ADFs), which formulates the training of Random Forests explicitly as a global loss minimization problem. During training, the losses are minimized via keeping an adaptive weight distribution over the training samples, similar to Boosting methods. In order to keep the method as flexible and general as possible, we adopt the principle of employing gradient descent in function space, which allows to minimize arbitrary losses. Contrary to Boosted Trees, in our method the loss minimization is an inherent part of the tree growing process, thus allowing to keep the benefits of common Random Forests, such as, parallel processing. We derive the new classifier and give a discussion and evaluation on standard machine learning data sets. Furthermore, we show how ADFs can be easily integrated into an object detection application. Compared to both, standard Random Forests and Boosted Trees, ADFs give better performance in our experiments, while yielding more compact models in terms of tree depth.
Samuel Schulter, Paul Wohlhart, Christian Leistner, Amir Saffari, Peter M. Roth, Horst Bischof
CVPR5
2013 Optimizing 1-Nearest Prototype Classifiers
abstract
The development of complex, powerful classifiers and their constant improvement have contributed much to the progress in many fields of computer vision. However, the trend towards large scale datasets revived the interest in simpler classifiers to reduce runtime. Simple nearest neighbor classifiers have several beneficial properties, such as low complexity and inherent multi-class handling, however, they have a runtime linear in the size of the database. Recent related work represents data samples by assigning them to a set of prototypes that partition the input feature space and afterwards applies linear classifiers on top of this representation to approximate decision boundaries locally linear. In this paper, we go a step beyond these approaches and purely focus on 1-nearest prototype classification, where we propose a novel algorithm for deriving optimal prototypes in a discriminative manner from the training samples. Our method is implicitly multi-class capable, parameter free, avoids noise over fitting and, since during testing only comparisons to the derived prototypes are required, highly efficient. Experiments demonstrate that we are able to outperform related locally linear methods, while even getting close to the results of more complex classifiers.
Paul Wohlhart, Martin Köstinger, Michael Donoser, Peter M. Roth, Horst Bischof
CVPR4
2013 Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor Classifiers
abstract
In this paper, we raise important issues concerning the evaluation complexity of existing Mahalanobis metric learning methods. The complexity scales linearly with the size of the dataset. This is especially cumbersome on large scale or for real-time applications with limited time budget. To alleviate this problem we propose to represent the dataset by a fixed number of discriminative prototypes. In particular, we introduce a new method that jointly chooses the positioning of prototypes and also optimizes the Mahalanobis distance metric with respect to these. We show that choosing the positioning of the prototypes and learning the metric in parallel leads to a drastically reduced evaluation effort while maintaining the discriminative essence of the original dataset. Moreover, for most problems our method performing k-nearest prototype (k-NP) classification on the condensed dataset leads to even better generalization compared to k-NN classification using all data. Results on a variety of challenging benchmarks demonstrate the power of our method. These include standard machine learning datasets as well as the challenging Public Figures Face Database. On the competitive machine learning benchmarks we are comparable to the state-of-the-art while being more efficient. On the face benchmark we clearly outperform the state-of-the-art in Mahalanobis metric learning with drastically reduced evaluation effort.
Martin Köstinger, Paul Wohlhart, Peter M. Roth, Horst Bischof
ICCV3
2013 Alternating Regression Forests for Object Detection and Pose Estimation
abstract
We present Alternating Regression Forests (ARFs), a novel regression algorithm that learns a Random Forest by optimizing a global loss function over all trees. This interrelates the information of single trees during the training phase and results in more accurate predictions. ARFs can minimize any differentiable regression loss without sacrificing the appealing properties of Random Forests, like low computational complexity during both, training and testing. Inspired by recent developments for classification [19], we derive a new algorithm capable of dealing with different regression loss functions, discuss its properties and investigate the relations to other methods like Boosted Trees. We evaluate ARFs on standard machine learning benchmarks, where we observe better generalization power compared to both standard Random Forests and Boosted Trees. Moreover, we apply the proposed regressor to two computer vision applications: object detection and head pose estimation from depth images. ARFs outperform the Random Forest baselines in both tasks, illustrating the importance of optimizing a common loss function for all trees.
Samuel Schulter, Christian Leistner, Paul Wohlhart, Peter M. Roth, Horst Bischof
ICCV4
2013 Hough-based tracking of non-rigid objects
Martin Godec, Peter M. Roth, Horst Bischof
Comput. Vis. Image Underst.2
2013 Segmentation-based tracking by support fusion
Markus Heber, Martin Godec, Matthias Rüther, Peter M. Roth, Horst Bischof
Comput. Vis. Image Underst.4
2012 Detecting Partially Occluded Objects with an Implicit Shape Model Random Field
Paul Wohlhart, Michael Donoser, Peter M. Roth, Horst Bischof
ACCV (1)3
2012 Person Re-identification by Efficient Impostor-Based Metric Learning
abstract
Recognizing persons over a system of disjunct cameras is a hard task for human operators and even harder for automated systems. In particular, realistic setups show difficulties such as different camera angles or different camera properties. Additionally, also the appearance of exactly the same person can change dramatically due to different views (e.g., frontal/back) of carried objects. In this paper, we mainly address the first problem by learning the transition from one camera to the other. This is realized by learning a Mahalanobis metric using pairs of labeled samples from different cameras. Building on the ideas of Large Margin Nearest Neighbor classification, we obtain a more efficient solution which additionally provides much better generalization properties. To demonstrate these benefits, we run experiments on three different publicly available datasets, showing state-of-the-art or even better results, however, on much lower computational efforts. This is in particular interesting since we use quite simple color and texture features, whereas other approaches build on rather complex image descriptions!
Martin Hirzer, Peter M. Roth, Horst Bischof
AVSS2
2012 Discriminative Hough Forests for Object Detection
abstract
Object detection models based on the Implicit Shape Model (ISM) [3] use small, local parts that vote for object centers in images. Since these parts vote completely independently from each other, this often leads to false-positive detections due to random constellations of parts. Thus, we introduce a verification step, which considers the activations of all voting elements that contribute to a detection. The levels of activation of each voting element of the ISM form a new description vector for an object hypothesis, which can be examined in order to discriminate between correct and incorrect detections. In particular, we observe the levels of activation of the voting elements in Hough Forests [2], which can be seen as a variant of ISM. In Hough Forests, the voting elements are all the positive training patches used to train the Forest. Each patch of the input image is classified by all decision trees in the Hough Forest. Whenever an input patch falls into the same leaf node as a patch from training, a certain amount of weight is added to the detection hypothesis at the relative position of the object center, which was recorded when cropping out the training patch. The total amount of weight one voting element (offset vector) adds to a detection hypothesis (the total activation) can be calculated by summing over all input patches and trees in the forest. Stacking the activations of all elements gives an activation vector for a hypothesis. We learn classifiers to discriminate correct and wrong part constellations based on these activation vectors and thus assign a better confidence to each detection. We use linear models as well as a histogram intersection kernel SVM. In the linear classifier, one weight is learned for each voting element. We additionally show how to use these weights, not only as a post processing step, but directly in the voting process. This has two advantages: First, it circumvents the explicit calculation of the activation vector for later reclassification, which is computationally more demanding. Second, the non-maxima suppression is performed on cleaner Hough maps, which allows for reducing the size of the suppression neighborhood and thus increases the recall at high levels of precision.
Paul Wohlhart, Samuel Schulter, Martin Köstinger, Peter M. Roth, Horst Bischof
BMVC4
2012 Large scale metric learning from equivalence constraints
abstract
In this paper, we raise important issues on scalability and the required degree of supervision of existing Mahalanobis metric learning methods. Often rather tedious optimization procedures are applied that become computationally intractable on a large scale. Further, if one considers the constantly growing amount of data it is often infeasible to specify fully supervised labels for all data points. Instead, it is easier to specify labels in form of equivalence constraints. We introduce a simple though effective strategy to learn a distance metric from equivalence constraints, based on a statistical inference perspective. In contrast to existing methods we do not rely on complex optimization problems requiring computationally expensive iterations. Hence, our method is orders of magnitudes faster than comparable methods. Results on a variety of challenging benchmarks with rather diverse nature demonstrate the power of our method. These include faces in unconstrained environments, matching before unseen object instances and person re-identification across spatially disjoint cameras. In the latter two benchmarks we clearly outperform the state-of-the-art.
Martin Köstinger, Martin Hirzer, Paul Wohlhart, Peter M. Roth, Horst Bischof
CVPR4
2012 Relaxed Pairwise Learned Metric for Person Re-identification
Martin Hirzer, Peter M. Roth, Martin Köstinger, Horst Bischof
ECCV (6)2
2012 Hough Regions for Joining Instance Localization and Segmentation
Hayko Riemenschneider, Sabine Sternig, Michael Donoser, Peter M. Roth, Horst Bischof
ECCV (3)4
2012 Dense appearance modeling and efficient learning of camera transitions for person re-identification
abstract
One central task in many visual surveillance scenarios is person re-identification, i.e., recognizing an individual person across a network of spatially disjoint cameras. Most successful recognition approaches are either based on direct modeling of the human appearance or on machine learning. In this work, we aim at taking advantage of both directions of research. On the one hand side, we compute a descriptive appearance representation encoding the vertical color structure of pedestrians. To improve the classification results, we additionally estimate the transition between two cameras using a pair-wisely estimated metric. In particular, we introduce 4D spatial color histograms and adopt Large Margin Nearest Neighbor (LMNN) metric learning. The approach is demonstrated for two publicly available datasets, showing competitive results, however, on lower computational costs.
Martin Hirzer, Csaba Beleznai, Martin Köstinger, Peter M. Roth, Horst Bischof
ICIP4
2012 On-line inverse multiple instance boosting for classifier grids
abstract
Classifier grids have shown to be a considerable choice for object detection from static cameras. By applying a single classifier per image location the classifier's complexity can be reduced and more specific and thus more accurate classifiers can be estimated. In addition, by using an on-line learner a highly adaptive but stable detection system can be obtained. Even though long-term stability has been demonstrated such systems still suffer from short-term drifting if an object is not moving over a long period of time. The goal of this work is to overcome this problem and thus to increase the recall while preserving the accuracy. In particular, we adapt ideas from multiple instance learning (MIL) for on-line boosting. In contrast to standard MIL approaches, which assume an ambiguity on the positive samples, we apply this concept to the negative samples: inverse multiple instance learning. By introducing temporal bags consisting of background images operating on different time scales, we can ensure that each bag contains at least one sample having a negative label, providing the theoretical requirements. The experimental results demonstrate superior classification results in presence of non-moving objects.
Sabine Sternig, Peter M. Roth, Horst Bischof
Pattern Recognit. Lett.2
2011 AVSS 2011 demo session: OUTLIER - online learning and visualization of unusual events
abstract
Summary form only given. We introduce to the surveillance community the VIRAT Video Dataset[1], which is a new large-scale surveillance video dataset designed to assess the performance of event recognition algorithms in realistic scenes1.
Josef A. Birchbauer, Samuel Schulter, René Schuster, Georg Poier, Peter Schallauer, Peter M. Roth, Horst Bischof
AVSS7
2011 Learning to recognize faces from videos and weakly related information cues
abstract
Videos are often associated with additional information that could be valuable for interpretation of their content. This especially applies for the recognition of faces within video streams, where often cues such as transcripts and subtitles are available. However, this data is not completely reliable and might be ambiguously labeled. To overcome these limitations, we take advantage of semi-supervised (SSL) and multiple instance learning (MIL) and propose a new semi-supervised multiple instance learning (SSMIL) algorithm. Thus, during training we can weaken the prerequisite of knowing the label for each instance and can integrate unlabeled data, given only probabilistic information in form of priors. The benefits of the approach are demonstrated for face recognition in videos on a publicly available benchmark dataset. In fact, we show exploring new information sources can considerably improve the classification results.
Martin Köstinger, Paul Wohlhart, Peter M. Roth, Horst Bischof
AVSS3
2011 Next-generation 3D visualization for visual surveillance
abstract
Existing visual surveillance systems typically require that human operators observe video streams from different cameras, which becomes infeasible if the number of observed cameras is ever increasing. In this paper, we present a new surveillance system that combines automatic video analysis (i.e., single person tracking and crowd analysis) and interactive visualization. Our novel visualization takes advantage of a high resolution display and given 3D information to focus the operator's attention to interesting/ critical areas of the observed area. This is realized by embedding the results of automatic scene analysis techniques into the visualization. By providing different visualization modes, the user can easily switch between the different modes and can select the mode which provides most information. The system is demonstrated for a real setup on a university campus.
Peter M. Roth, Volker Settgast, Peter Widhalm, Marcel Lancelle, Josef A. Birchbauer, Norbert Brändle, Sven Havemann, Horst Bischof
AVSS1
2011 On-line Hough Forests
abstract
Hough forests have emerged as a powerful and versatile method, which achieves state-of-the-art results on various computer vision applications, ranging from object detection over pose estimation to action recognition. The original method operates in offline mode, assuming to have access to the entire training set at once. This limits its applicability in domains where data arrives sequentially or when large amounts of data have to be exploited. In these cases, on-line approaches naturally would be beneficial. To this end, we propose an on-line extension of Hough forests, which is based on the principle of letting the trees evolve on-line while the data arrives sequentially, for both classification and regression. We further propose a modified version of off-line Hough forests, which only needs a small subset of the training data for optimization. In the experiments, we show that using these formulations, the classification results of classic Hough forests could be reached or even outperformed, while being orders of magnitudes faster. Furthermore, our method allows for tracking arbitrary objects without requiring any prior knowledge. We present state-of-the-art tracking results on publicly available data sets. © 2011. The copyright of this document resides with its authors.
Samuel Schulter, Christian Leistner, Peter M. Roth, Horst Bischof, Luc Van Gool
BMVC3
2011 Hough-based tracking of non-rigid objects
abstract
Online learning has shown to be successful in tracking of previously unknown objects. However, most approaches are limited to a bounding-box representation with fixed aspect ratio. Thus, they provide a less accurate fore- ground/background separation and cannot handle highly non-rigid and articulated objects. This, in turn, increases the amount of noise introduced during online self-training. In this paper, we present a novel tracking-by-detection approach to overcome this limitation based on the generalized Hough-transform. We extend the idea of Hough Forests to the online domain and couple the voting- based detection and back-projection with a rough segmentation based on GrabCut. This significantly reduces the amount of noisy training samples during online learning and thus effectively prevents the tracker from drifting. In the experiments, we demonstrate that our method successfully tracks a variety of previously unknown objects even under heavy non-rigid transformations, partial occlusions, scale changes and rotations. Moreover, we compare our tracker to state-of-the-art methods (both bounding-box- based as well as part-based) and show robust and accurate tracking results on various challenging sequences.
Martin Godec, Peter M. Roth, Horst Bischof
ICCV2
2010 Temporal Feature Weighting for Prototype-Based Action Recognition
Thomas Mauthner, Peter M. Roth, Horst Bischof
ACCV (2)2
2010 Automatic Detection and Reading of Dangerous Goods Plates
abstract
In this paper, we present an efficient solution for automatic detection and reading of dangerous goods plates on trucks and trains. According to the ADR agreement dangerous goods transports are marked with an orange plate covering the hazard class and the identification number for the hazardous substances. Since under real-world conditions high resolution images (often at low quality) have to be processed an efficient and robust system is required. In particular, we propose a multi-stage system consisting of an acquisition step, a saliency region detector (to reduce the run-time), a plate detector, and a robust recognition step based on an Optical Character Recognition (OCR). To demonstrate the system, we show qualitative and quantitative localization/recognition results on two challenging data sets. In fact, building on proven robust and efficient methods, we show excellent detection and classification results under hard environmental conditions at low run-time.
Peter M. Roth, Martin Köstinger, Paul Wohlhart, Horst Bischof, Josef A. Birchbauer
AVSS1
2010 Learning of Scene-Specific Object Detectors by Classifier Co-Grids
abstract
Recently, classifier grids have shown to be a considerable alternative to sliding window approaches for object detection from static cameras. The main drawback of such methods is that they are biased by the initial model. In fact, the classifiers can be adapted to changing environmental conditions but due to conservative updates no new object-specific information is acquired. Thus, the goal of this work is to increase the recall of scene-specific classifiers while preserving their accuracy and speed. In particular, we introduce a co-training strategy for classifier grids using a robust on-line learner. Thus, the robustness is preserved while the recall can be increased. The co-training strategy robustly provides negative as well as positive updates. In addition, the number of negative updates can be drastically reduced, which additionally speeds up the system. In the experimental results these benefits are demonstrated on different publicly available surveillance benchmark data sets.
Sabine Sternig, Peter M. Roth, Horst Bischof
AVSS2
2010 Inverse Multiple Instance Learning for Classifier Grids
abstract
Recently, classifier grids have shown to be a considerable alternative for object detection from static cameras. However, one drawback of such approaches is drifting if an object is not moving over a long period of time. Thus, the goal of this work is to increase the recall of such classifiers while preserving their accuracy and speed. In particular, this is realized by adapting ideas from Multiple Instance Learning within a boosting framework. Since the set of positive samples is well defined, we apply this concept to the negative samples extracted from the scene: Inverse Multiple Instance Learning. By introducing temporal bags, we can ensure that each bag contains at least one sample having a negative label, providing the required stability. The experimental results demonstrate that using the proposed approach state-of-the-art detection results can by obtained, however, showing superior classification results in presence of non-moving objects.
Sabine Sternig, Peter M. Roth, Horst Bischof
ICPR2
2009 Semantic Classification in Aerial Imagery by Integrating Appearance and Height Information
Stefan Kluckner, Thomas Mauthner, Peter M. Roth, Horst Bischof
ACCV (2)3
2009 Semantic Image Classification using Consistent Regions and Individual Context
abstract
This paper proposes an efficient approach for semantic image classification by integrating additional contextual constraints such as class co-occurrences into a randomized forest classification framework. The randomized forest classifier performs an initial yet local classification on the pixel level by using powerful covariance matrix based descriptors as feature representation. Furthermore, we exploit multiple unsupervised image partitions to provide a reliable spatial region support and to capture the real object boundaries. An information theoretic driven approach detects consistently classified regions and generates a representative segmentation incorporating the classification result on the pixel level. Moreover, we use a conditional random field formulation to obtain a final labeling including context information individually generated for each test image. To illustrate state-of-the-art performance, we run experiments on the two versions of the MSRC [21] dataset with 9 and 21 object classes and on the PASCAL VOC2007 [5] image collection.
Stefan Kluckner, Thomas Mauthner, Peter M. Roth, Horst Bischof
BMVC3
2009 Classifier grids for robust adaptive object detection
abstract
In this paper we present an adaptive but robust object detector for static cameras by introducing classifier grids. Instead of using a sliding window for object detection we propose to train a separate classifier for each image location, obtaining a very specific object detector with a low false alarm rate. For each classifier corresponding to a grid element we estimate two generative representations in parallel, one describing the object's class and one describing the background. These are combined in order to obtain a discriminative model. To enable to adapt to changing environments these classifiers are learned on-line (i.e., boosting). Continuously learning (24 hours a day, 7 days a week) requires a stable system. In our method this is ensured by a fixed object representation while updating only the representation of the background. We demonstrate the stability in a long-term experiment by running the system for a whole week, which shows a stable performance over time. In addition, we compare the proposed approach to state-of-the-art methods in the field of person and car detection. In both cases we obtain competitive results.
Peter M. Roth, Sabine Sternig, Helmut Grabner, Horst Bischof
CVPR1
2009 A level set framework using a new incremental, robust Active Shape Model for object segmentation and tracking
Michael Fussenegger, Peter M. Roth, Horst Bischof, Rachid Deriche, Axel Pinz
Image Vis. Comput.2
2007 Incremental LDA Learning by Combining Reconstructive and Discriminative Approaches
abstract
Incremental subspace methods have proven to enable efficient training if large amounts of training data have to be processed or if not all data is available in advance. In this paper we focus on incremental LDA learning which provides good classification results while it assures a compact data representation. In contrast to existing incremental LDA methods we additionally consider reconstructive information when incrementally building the LDA subspace. Hence, we get a more flexible representation that is capable to adapt to new data. Moreover, this allows to add new instances to existing classes as well as to add new classes. The experimental results show that the proposed approach outperforms other incremental LDA methods even approaching classification results obtained by batch learning. 1
Martina Uray, Danijel Skocaj, Peter M. Roth, Horst Bischof, Ales Leonardis
BMVC3
2007 Eigenboosting: Combining Discriminative and Generative Information
abstract
A major shortcoming of discriminative recognition and detection methods is their noise sensitivity, both during training and recognition. This may lead to very sensitive and brittle recognition systems focusing on irrelevant information. This paper proposes a method that selects generative and discriminative features. In particular, we boost classical Haar-like features and use the same features to approximate a generative model (i.e., eigenimages). A modified error function for boosting ensures that only features are selected that show a good discrimination and reconstruction. This allows a robust feature selection using boosting. Thus, we can handle problems where discriminant classifiers fail while still retaining the discriminative power. Our experiments show that we can significantly improve the recognition performance when learning from noisy data. Moreover, the feature type used allows efficient recognition and reconstruction.
Helmut Grabner, Peter M. Roth, Horst Bischof
CVPR2