EDBT 2026 Demo / reviewers in the wild / expert
Horst Bischof
dblp:69/3793
· DBLP profile ↗
290ranked-venue papers
15as first author
28since 2021 · last 2026
0000-0002-9096-6671ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 226 · 12 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 224 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 1 since 2021Systems, architecture and hardware · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spectral Basis Learning for Expressive Graph Neural Networks in Link PredictionabstractGraph Neural Networks (GNNs) excel in handling graph-structured data but often underperform in link prediction tasks compared to classical methods, mainly due to the limitations of the commonly used message-passing principle. Notably, their ability to distinguish non-isomorphic graphs is limited by the 1-dimensional Weisfeiler-Lehman test (WL). Our study presents a novel method to enhance the expressivity of GNNs by embedding induced subgraphs into the eigenbasis of the graph Laplacian. We introduce a Learnable Lanczos algorithm with Linear Constraints (LLwLC), proposing two novel subgraph extraction strategies: encoding vertex-deleted subgraphs and applying Neumann eigenvalue constraints. For the former, we demonstrate the ability to distinguish graphs that are indistinguishable by 2-WL, while maintaining efficiency. The latter focuses on link representations enabling differentiation between k-regular graphs and node automorphism, a vital aspect for link prediction tasks. Our approach results in a lightweight architecture, reducing the need for extensive training datasets. Empirically, our method improves performance in challenging link prediction tasks across benchmark datasets, establishing its practical utility and supporting our theoretical findings. Notably, LLwLC achieves 20x and 10x speedups by requiring only 5% and 10% of the data from the PubMed and OGBL-Vessel datasets, while comparing to the state-of-the-art. Niloofar Azizi, Nils M. Kriege, Nicholas J. A. Harvey, Horst Bischof |
AAAI | 4 |
| 2024 | Vision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection
Christian Fruhwirth-Reisinger, Wei Lin 0019, Dusan Malic, Horst Bischof, Horst Possegger |
BMVC | 4 |
| 2024 | Into the Fog: Evaluating Robustness of Multiple Object Tracking
Nadezda Kirillova, Muhammad Jehanzeb Mirza, Horst Bischof, Horst Possegger |
BMVC | 3 |
| 2024 | MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly DetectionabstractWe propose a novel approach to video anomaly detection: we treat feature vectors extracted from videos as re-alizations of a random variable with a fixed distribution and model this distribution with a neural network. This lets us estimate the likelihood of test videos and detect video anomalies by thresholding the likelihood estimates. We train our video anomaly detector using a modification of de-noising score matching, a method that injects training data with noise to facilitate modeling its distribution. To elim-inate hyperparameter selection, we model the distribution of noisy video features across a range of noise levels and introduce a regularizer that tends to align the models for different levels of noise. At test time, we combine anomaly indications at multiple noise scales with a Gaussian mix-ture model. Running our video anomaly detector induces minimal delays as inference requires merely extracting the features and forward-propagating them through a shallow neural network and a Gaussian mixture model. Our ex-periments on five popular video anomaly detection bench-marks demonstrate state-of-the-art performance, both in the object-centric and in the frame-centric setup. Jakub Micorek, Horst Possegger, Dominik Narnhofer, Horst Bischof, Mateusz Kozinski |
CVPR | 4 |
| 2024 | Occlusion Handling in 3D Human Pose Estimation with Perturbed Positional Encoding
Niloofar Azizi, Mohsen Fayyaz, Horst Bischof |
ECCV (13) | 3 |
| 2024 | Robust Localization of Key Fob Using Channel Impulse Response of Ultra Wide Band Sensors for Keyless Entry SystemsabstractUsing neural networks for localization of key fob within and surrounding a car as a security feature for keyless entry is fast emerging. In this paper we study: 1) the performance of pre-computed features of neural networks based UWB (ultra wide band) localization classification forming the baseline of our experiments. 2) Investigate the inherent robustness of various neural networks; therefore, we include the study of robustness of the adversarial examples without any adversarial training in this work. 3) Propose a multi-head self-supervised neural network architecture which outperforms the baseline neural networks without any adversarial training. The model’s performance improved by 67% at certain ranges of adversarial magnitude for fast gradient sign method and 37% each for basic iterative method and projected gradient descent method. Abhiram Kolli, Filippo Casamassima, Horst Possegger, Horst Bischof |
ICASSP | 4 |
| 2024 | Action-By-Detection: Efficient Forklift Action Detection for Autonomous Mobile Robots in WarehousesabstractUnderstanding actions of other agents increases the efficiency of autonomous mobile robots (AMRs) since they encompass intention and indicate future movements. We propose a new method that allows us to infer vehicle actions using a shallow image-based classification model. The actions are classified via bird’s-eye view scene crops, where we project the detections of a 3D object detection model onto a context map. We learn map context information and aggregate temporal sequence information without requiring object tracking. This results in a highly efficient classification model that can easily be deployed on embedded AMR hardware. To evaluate our approach, we create new large-scale synthetic datasets showing warehouse traffic based on real vehicle models and geometry. Alexander Prutsch, Horst Possegger, Horst Bischof |
ICRA | 3 |
| 2024 | Efficient Motion Prediction: A Lightweight & Accurate Trajectory Prediction Model With Fast Training and Inference SpeedabstractFor efficient and safe autonomous driving, it is essential that autonomous vehicles can predict the motion of other traffic agents. While highly accurate, current motion prediction models often impose significant challenges in terms of training resource requirements and deployment on embedded hardware. We propose a new efficient motion prediction model, which achieves highly competitive benchmark results while training only a few hours on a single GPU. Due to our lightweight architectural choices and the focus on reducing the required training resources, our model can easily be applied to custom datasets. Furthermore, its low inference latency makes it particularly suitable for deployment in autonomous applications with limited computing resources. Alexander Prutsch, Horst Bischof, Horst Possegger |
IROS | 2 |
| 2024 | MAELi: Masked Autoencoder for Large-Scale LiDAR Point CloudsabstractThe sensing process of large-scale LiDAR point clouds inevitably causes large blind spots, i.e. regions not visible to the sensor. We demonstrate how these inherent sampling properties can be effectively utilized for self-supervised representation learning by designing a highly effective pretraining framework that considerably reduces the need for tedious 3D annotations to train state-of-the-art object detectors. Our Masked AutoEncoder for LiDAR point clouds (MAELi) intuitively leverages the sparsity of LiDAR point clouds in both the encoder and decoder during reconstruction. This results in more expressive and useful initialization, which can be directly applied to downstream perception tasks, such as 3D object detection or semantic segmentation for autonomous driving. In a novel reconstruction approach, MAELi distinguishes between empty and occluded space and employs a new masking strategy that targets the LiDAR’s inherent spherical projection. Thereby, without any ground truth whatsoever and trained on single frames only, MAELi obtains an understanding of the underlying 3D scene geometry and semantics. To demonstrate the potential of MAELi, we pre-train backbones in an end-to-end manner and show the effectiveness of our unsupervised pre-trained weights on the tasks of 3D object detection and semantic segmentation. Georg Krispel, David Schinagl, Christian Fruhwirth-Reisinger, Horst Possegger, Horst Bischof |
WACV | 5 |
| 2024 | ATS: Adaptive Temperature Scaling for Enhancing Out-of-Distribution Detection MethodsabstractOut-of-distribution (OOD) detection is essential to ensure the reliability and robustness of machine learning models in real-world applications. Post-hoc OOD detection methods have gained significant attention due to the fact that they offer the advantage of not requiring additional re-training, which could degrade model performance and increase training time. However, most existing post-hoc methods rely only on the encoder output (features), logits, or the softmax probability, meaning they have no access to information that might be lost in the feature extraction process. In this work, we address this limitation by introducing Adaptive Temperature Scaling (ATS), a novel approach that dynamically calculates a temperature value based on activations of the intermediate layers. Fusing this sample-specific adjustment with class-dependent logits, our ATS captures additional statistical information before they are lost in the feature extraction process, leading to a more robust and powerful OOD detection method. We conduct extensive experiments to demonstrate the efficacy of our approach. Notably, our method can be seamlessly combined with SOTA post-hoc OOD detection methods that rely on the logits, thereby enhancing their performance and improving their robustness. Gerhard Krumpl, Henning Avenhaus, Horst Possegger, Horst Bischof |
WACV | 4 |
| 2023 | A Comprehensive Crossroad Camera Dataset to Improve Traffic Safety of Mobility Aid Users
Ludwig Mohr, Nadezda Kirillova, Horst Possegger, Horst Bischof |
BMVC | 4 |
| 2023 | Video Test-Time Adaptation for Action RecognitionabstractAlthough action recognition systems can achieve top performance when evaluated on in-distribution test points, they are vulnerable to unanticipated distribution shifts in test data. However, test-time adaptation of video action recognition models against common distribution shifts has so far not been demonstrated. We propose to address this problem with an approach tailored to spatio-temporal models that is capable of adaptation on a single video sample at a step. It consists in a feature distribution alignment technique that aligns online estimates of test set statistics towards the training statistics. We further enforce prediction consistency over temporally augmented views of the same test video sample. Evaluations on three benchmark action recognition datasets show that our proposed technique is architecture-agnostic and able to significantly boost the performance on both, the state of the art convolutional architecture TANet and the Video Swin Transformer. Our proposed method demonstrates a substantial performance gain over existing test-time adaptation approaches in both evaluations of a single distribution shift and the challenging case of random distribution shifts. Code will be available at https://github.com/wlin-at/ViTTA. Wei Lin 0019, Muhammad Jehanzeb Mirza, Mateusz Kozinski, Horst Possegger, Hilde Kuehne, Horst Bischof |
CVPR | 6 |
| 2023 | ActMAD: Activation Matching to Align Distributions for Test-Time-TrainingabstractTest-Time-Training (TTT) is an approach to cope with out-of-distribution (OOD) data by adapting a trained model to distribution shifts occurring at test-time. We propose to perform this adaptation via Activation Matching (ActMAD): We analyze activations of the model and align activation statistics of the OOD test data to those of the training data. In contrast to existing methods, which model the distribution of entire channels in the ultimate layer of the feature extractor, we model the distribution of each feature in multiple layers across the network. This results in a more fine-grained supervision and makes ActMAD attain state of the art performance on CIFAR-100C and Imagenet-C. ActMAD is also architecture-and task-agnostic, which lets us go beyond image classification, and score 15.4% improvement over previous approaches when evaluating a KITTI-trained object detector on KITTI-Fog. Our experiments highlight that ActMAD can be applied to online adaptation in realistic scenarios, requiring little data to attain its full performance. Muhammad Jehanzeb Mirza, Pol Jané-Soneira, Wei Lin 0019, Mateusz Kozinski, Horst Possegger, Horst Bischof |
CVPR | 6 |
| 2023 | MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeabstractLarge scale Vision Language (VL) models have shown tremendous success in aligning representations between visual and text modalities. This enables remarkable progress in zero-shot recognition, image generation & editing, and many other exciting tasks. However, VL models tend to over-represent objects while paying much less attention to verbs, and require additional tuning on video data for best zero-shot action recognition performance. While previous work relied on large-scale, fully-annotated data, in this work we propose an unsupervised approach. We adapt a VL model for zero-shot and few-shot action recognition using a collection of unlabeled videos and an unpaired action dictionary. Based on that, we leverage Large Language Models and VL models to build a text bag for each unlabeled video via matching, text expansion and captioning. We use those bags in a Multiple Instance Learning setup to adapt an image-text backbone to video data. Although finetuned on unlabeled video data, our resulting models demonstrate high transferability to numerous unseen zero-shot downstream tasks, improving the base VL model performance by up to 14%, and even comparing favorably to fully-supervised baselines in both zero-shot and few-shot video recognition transfer. The code is released at https://github.com/wlin-at/MAXI. Wei Lin 0019, Leonid Karlinsky, Nina Shvetsova, Horst Possegger, Mateusz Kozinski, Rameswar Panda, Rogério Feris, Hilde Kuehne, Horst Bischof |
ICCV | 9 |
| 2023 | MATE: Masked Autoencoders are Online 3D Test-Time LearnersabstractOur MATE is the first Test-Time-Training (TTT) method designed for 3D data, which makes deep networks trained for point cloud classification robust to distribution shifts occurring in test data. Like existing TTT methods from the 2D image domain, MATE also leverages test data for adaptation. Its test-time objective is that of a Masked Autoencoder: a large portion of each test point cloud is removed before it is fed to the network, tasked with reconstructing the full point cloud. Once the network is updated, it is used to classify the point cloud. We test MATE on several 3D object classification datasets and show that it significantly improves robustness of deep networks to several types of corruptions commonly occurring in 3D point clouds. We show that MATE is very efficient in terms of the fraction of points it needs for the adaptation. It can effectively adapt given as few as 5% of tokens of each test sample, making it extremely lightweight. Our experiments show that MATE also achieves competitive performance by adapting sparsely on the test data, which further reduces its computational overhead, making it ideal for real-time applications. Muhammad Jehanzeb Mirza, Inkyu Shin, Wei Lin 0019, Andreas Schriebl, Kunyang Sun, Jaesung Choe, Mateusz Kozinski, Horst Possegger, In-So Kweon, Kuk-Jin Yoon, Horst Bischof |
ICCV | 11 |
| 2023 | GACE: Geometry Aware Confidence Enhancement for Black-box 3D Object Detectors on LiDAR-DataabstractWidely-used LiDAR-based 3D object detectors often neglect fundamental geometric information readily available from the object proposals in their confidence estimation. This is mostly due to architectural design choices, which were often adopted from the 2D image domain, where geometric context is rarely available. In 3D, however, considering the object properties and its surroundings in a holistic way is important to distinguish between true and false positive detections, e.g. occluded pedestrians in a group. To address this, we present GACE, an intuitive and highly efficient method to improve the confidence estimation of a given black-box 3D object detector. We aggregate geometric cues of detections and their spatial relationships, which enables us to properly assess their plausibility and consequently, improve the confidence estimation. This leads to consistent performance gains over a variety of state-of-the-art detectors. Across all evaluated detectors, GACE proves to be especially beneficial for the vulnerable road user classes, i.e. pedestrians and cyclists. David Schinagl, Georg Krispel, Christian Fruhwirth-Reisinger, Horst Possegger, Horst Bischof |
ICCV | 5 |
| 2023 | Sit Back and Relax: Learning to Drive Incrementally in All Weather ConditionsabstractIn autonomous driving scenarios, current object detection models show strong performance when tested in clear weather. However, their performance deteriorates significantly when tested in degrading weather conditions. In addition, even when adapted to perform robustly in a sequence of different weather conditions, they are often unable to perform well in all of them and suffer from catastrophic forgetting. To efficiently mitigate forgetting, we propose Domain-Incremental Learning through Activation Matching (DILAM), which employs unsupervised feature alignment to adapt only the affine parameters of a clear weather pre-trained network to different weather conditions. We propose to store these affine parameters as a memory bank for each weather condition and plug-in their weather-specific parameters during driving (i.e. test time) when the respective weather conditions are encountered. Our memory bank is extremely lightweight, since affine parameters account for less than 2% of a typical object detector. Furthermore, contrary to previous domain-incremental learning approaches, we do not require the weather label when testing and propose to automatically infer the weather condition by a majority voting linear classifier. Stefan Leitner, Muhammad Jehanzeb Mirza, Wei Lin 0019, Jakub Micorek, Marc Masana, Mateusz Kozinski, Horst Possegger, Horst Bischof |
IV | 8 |
| 2023 | LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsabstractRecently, large-scale pre-trained Vision and Language (VL) models have set a new state-of-the-art (SOTA) in zero-shot visual classification enabling open-vocabulary recognition of potentially unlimited set of categories defined as simple language prompts. However, despite these great advances, the performance of these zero-shot classifiers still falls short of the results of dedicated (closed category set) classifiers trained with supervised fine-tuning. In this paper we show, for the first time, how to reduce this gap without any labels and without any paired VL data, using an unlabeled image collection and a set of texts auto-generated using a Large Language Model (LLM) describing the categories of interest and effectively substituting labeled visual instances of those categories. Using our label-free approach, we are able to attain significant performance improvements over the zero-shot performance of the base VL model and other contemporary methods and baselines on a wide variety of datasets, demonstrating absolute improvement of up to $11.7\%$ ($3.8\%$ on average) in the label-free setting. Moreover, despite our approach being label-free, we observe $1.3\%$ average gains over leading few-shot prompting baselines that do use 5-shot supervision. Muhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin 0019, Horst Possegger, Mateusz Kozinski, Rogério Feris, Horst Bischof |
NeurIPS | 7 |
| 2023 | SAILOR: Scaling Anchors via Insights into Latent Object RepresentationabstractLiDAR 3D object detection models are inevitably biased towards their training dataset. The detector clearly exhibits this bias when employed on a target dataset, particularly towards object sizes. However, object sizes vary heavily between domains due to, for instance, different labeling policies or geographical locations. State-of-the-art unsupervised domain adaptation approaches outsource methods to overcome the object size bias. Mainstream size adaptation approaches exploit target domain statistics, contradicting the original unsupervised assumption. Our novel unsupervised anchor calibration method addresses this limitation. Given a model trained on the source data, we estimate the optimal target anchors in a completely unsupervised manner. The main idea stems from an intuitive observation: by varying the anchor sizes for the target domain, we inevitably introduce noise or even remove valuable object cues. The latent object representation, perturbed by the anchor size, is closest to the learned source features only under the optimal target anchors. We leverage this observation for anchor size optimization. Our experimental results show that, without any retraining, we achieve competitive results even compared to state-of-the-art weakly-supervised size adaptation approaches. In addition, our anchor calibration can be combined with such existing methods, making them completely unsupervised. Dusan Malic, Christian Fruhwirth-Reisinger, Horst Possegger, Horst Bischof |
WACV | 4 |
| 2022 | A Visual Surveillance System to Observe Realistic Road User Behavior for Improved Pedestrian and Cyclist Safety at CrossroadsabstractPedestrians and cyclists suffer the most serious injuries in traffic accidents. Existing Pedestrian Protection Systems and Road Safety Systems rely on an ideal model of pedestrian behavior and do not consider that people tend to take shortcuts, appear at unexpected places or can be distracted on the road, for example, by using a smartphone or wearing headphones. Collecting and analyzing realistic road user behavior is a crucial component to improve pedestrian and cyclist safety. However, such real-world data is still missing. To address this, we propose a visual surveillance system with two perpendicular partially overlapping fields of view, combined with a fully automated deep learning-based pipeline to process and collect video observations, detect and extract road user trajectories in real-world coordinates and estimate human attributes, such as age, gender, smartphone usage, etc. We demonstrate our prototype by deploying it in two locations in a European city. Nadezda Kirillova, Horst Possegger, Horst Bischof |
AVSS | 3 |
| 2022 | The Norm Must Go On: Dynamic Unsupervised Domain Adaptation by NormalizationabstractDomain adaptation is crucial to adapt a learned model to new scenarios, such as domain shifts or changing data distributions. Current approaches usually require a large amount of labeled or unlabeled data from the shifted domain. This can be a hurdle in fields which require continuous dynamic adaptation or suffer from scarcity of data, e.g. autonomous driving in challenging weather conditions. To address this problem of continuous adaptation to distribution shifts, we propose Dynamic Unsupervised Adaptation (DUA). By continuously adapting the statistics of the batch normalization layers we modify the feature representations of the model. We show that by sequentially adapting a model with only a fraction of unlabeled data, a strong performance gain can be achieved. With even less than 1% of unlabeled data from the target domain, DUA already achieves competitive results to strong baselines. In addition, the computational overhead is minimal in contrast to previous approaches. Our approach is simple, yet effective and can be applied to any architecture which uses batch normalization as one of its components. We show the utility of DUA by evaluating it on a variety of domain adaptation datasets and tasks including object recognition, digit recognition and object detection. Muhammad Jehanzeb Mirza, Jakub Micorek, Horst Possegger, Horst Bischof |
CVPR | 4 |
| 2022 | OccAM's Laser: Occlusion-based Attribution Maps for 3D Object Detectors on LiDAR DataabstractWhile 3D object detection in LiDAR point clouds is well-established in academia and industry, the explainability of these models is a largely unexplored field. In this paper, we propose a method to generate attribution maps for the detected objects in order to better understand the behavior of such models. These maps indicate the importance of each 3D point in predicting the specific objects. Our method works with black-box models: We do not require any prior knowledge of the architecture nor access to the model's internals, like parameters, activations or gradients. Our efficient perturbation-based approach empirically estimates the importance of each point by testing the model with randomly generated subsets of the input point cloud. Our sub-sampling strategy takes into account the special characteristics of LiDAR data, such as the depth-dependent point density. We show a detailed evaluation of the attribution maps and demonstrate that they are interpretable and highly informative. Furthermore, we compare the attribution maps of recent 3D object detection architectures to provide insights into their decision-making processes. David Schinagl, Georg Krispel, Horst Possegger, Peter M. Roth, Horst Bischof |
CVPR | 5 |
| 2022 | 3D Human Pose Estimation Using Möbius Graph Convolutional Networks
Niloofar Azizi, Horst Possegger, Emanuele Rodolà, Horst Bischof |
ECCV (1) | 4 |
| 2022 | CycDA: Unsupervised Cycle Domain Adaptation to Learn from Image to Video
Wei Lin 0019, Anna Kukleva, Kunyang Sun, Horst Possegger, Hilde Kuehne, Horst Bischof |
ECCV (3) | 6 |
| 2022 | OnlyCaps-Net, a Capsule only Based Neural Network for 2D and 3D Semantic Segmentation
Savinien Bonheur, Franz Thaler, Michael Pienn, Horst Olschewski, Horst Bischof, Martin Urschler |
MICCAI (5) | 5 |
| 2021 | FAST3D: Flow-Aware Self-Training for 3D Object Detectors
Christian Fruhwirth-Reisinger, Michael Opitz, Horst Possegger, Horst Bischof |
BMVC | 4 |
| 2021 | DRT: Detection Refinement for Multiple Object Tracking
Bisheng Wang, Christian Fruhwirth-Reisinger, Horst Possegger, Horst Bischof, Guo Cao |
BMVC | 4 |
| 2021 | Semi-Supervised Learning Of Monocular 3D Hand Pose Estimation From Multi-View ImagesabstractMost modern hand pose estimation methods rely on Convolutional Neural Networks (CNNs), which typically require a large training dataset to perform well. Exploiting unlabeled data provides a way to reduce the required amount of annotated data. We propose to take advantage of a geometry-aware representation of the human hand, which we learn from multiview images without annotations. The objective for learning this representation is simply based on learning to predict a different view. Our results show that using this objective yields clearly superior pose estimation results compared to directly mapping an input image to the 3Djoint locations of the hand if the amount of 3D annotations is limited. We further show the effect of the objective for either case, using the objective for pre-learning as well as to simultaneously learn to predict novel views and to estimate the 3D pose of the hand. Georg Poier, Horst Possegger, Horst Bischof |
ICIP | 4 |
| 2020 | FuseSeg: LiDAR Point Cloud Segmentation Fusing Multi-Modal DataabstractWe introduce a simple yet effective fusion method of LiDAR and RGB data to segment LiDAR point clouds. Utilizing the dense native range representation of a LiDAR sensor and the setup calibration, we establish point correspondences between the two input modalities. Subsequently, we are able to warp and fuse the features from one domain into the other. Therefore, we can jointly exploit information from both data sources within one single network. To show the merit of our method, we extend SqueezeSeg, a point cloud segmentation network, with an RGB feature branch and fuse it into the original structure. Our extension called FuseSeg leads to an improvement of up to 18% IoU on the KITTI benchmark. In addition to the improved accuracy, we also achieve real-time performance at 50 fps, five times as fast as the recording speed of the KITTI LiDAR data. Georg Krispel, Michael Opitz, Georg Waltner, Horst Possegger, Horst Bischof |
WACV | 5 |
| 2020 | Deep Metric Learning with BIER: Boosting Independent Embeddings RobustlyabstractLearning similarity functions between image pairs with deep neural networks yields highly correlated activations of embeddings. In this work, we show how to improve the robustness of such embeddings by exploiting the independence within ensembles. To this end, we divide the last embedding layer of a deep network into an embedding ensemble and formulate the task of training this ensemble as an online gradient boosting problem. Each learner receives a reweighted training sample from the previous learners. Further, we propose two loss functions which increase the diversity in our ensemble. These loss functions can be applied either for weight initialization or during training. Together, our contributions leverage large embedding sizes more effectively by significantly reducing correlation of the embedding and consequently increase retrieval accuracy of the embedding. Our method works with any differentiable loss function and does not introduce any additional parameters during test time. We evaluate our metric learning method on image retrieval tasks and show that it improves over state-of-the-art methods on the CUB-200-2011, Cars-196, Stanford Online Products, In-Shop Clothes Retrieval and VehicleID datasets. Therefore, our findings suggest that by dividing deep networks at the end into several smaller and diverse networks, we can significantly reduce overfitting. Michael Opitz, Georg Waltner, Horst Possegger, Horst Bischof |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | MURAUER: Mapping Unlabeled Real Data for Label AUstERityabstractData labeling for learning 3D hand pose estimation models is a huge effort. Readily available, accurately labeled synthetic data has the potential to reduce the effort. However, to successfully exploit synthetic data, current state-of-the-art methods still require a large amount of labeled real data. In this work, we remove this requirement by learning to map from the features of real data to the features of synthetic data mainly using a large amount of synthetic and unlabeled real data. We exploit unlabeled data using two auxiliary objectives, which enforce that (i) the mapped representation is pose specific and (ii) at the same time, the distributions of real and synthetic data are aligned. While pose specifity is enforced by a self-supervisory signal requiring that the representation is predictive for the appearance from different views, distributions are aligned by an adversarial term. In this way, we can significantly improve the results of the baseline system, which does not use unlabeled data and outperform many recent approaches already with about 1% of the labeled real data. This presents a step towards faster deployment of learning based hand pose estimation, making it accessible for a larger range of applications. Georg Poier, Michael Opitz, David Schinagl, Horst Bischof |
WACV | 4 |
| 2019 | HiBsteR: Hierarchical Boosted Deep Metric Learning for Image RetrievalabstractWhen the number of categories is growing into thousands, large-scale image retrieval becomes an increasingly hard task. Retrieval accuracy can be improved by learning distance metric methods that separate categories in a transformed embedding space. Unlike most methods that utilize a single embedding to learn a distance metric, we build on the idea of boosted metric learning, where an embedding is split into a boosted ensemble of embeddings. While in general metric learning is directly applied on fine labels to learn embeddings, we take this one step further and incorporate hierarchical label information into the boosting framework and show how to properly adapt loss functions for this purpose. We show that by introducing several sub-embeddings which focus on specific hierarchical classes, the retrieval accuracy can be improved compared to standard flat label embeddings. The proposed method is especially suitable for exploiting hierarchical datasets or when additional labels can be retrieved without much effort. Our approach improves R@1 over state-of-the-art methods on the biggest available retrieval dataset (Stanford Online Products) and sets new reference baselines for hierarchical metric learning on several other datasets (CUB-200-2011, VegFru, FruitVeg-81). We show that the clustering quality in terms of NMI score is superior to previous works. Georg Waltner, Michael Opitz, Horst Possegger, Horst Bischof |
WACV | 4 |
| 2019 | Integrating spatial configuration into heatmap regression based CNNs for landmark localizationabstractIn many medical image analysis applications, only a limited amount of training data is available due to the costs of image acquisition and the large manual annotation effort required from experts. Training recent state-of-the-art machine learning methods like convolutional neural networks (CNNs) from small datasets is a challenging task. In this work on anatomical landmark localization, we propose a CNN architecture that learns to split the localization task into two simpler sub-problems, reducing the overall need for large training datasets. Our fully convolutional SpatialConfiguration-Net (SCN) learns this simplification due to multiplying the heatmap predictions of its two components and by training the network in an end-to-end manner. Thus, the SCN dedicates one component to locally accurate but ambiguous candidate predictions, while the other component improves robustness to ambiguities by incorporating the spatial configuration of landmarks. In our extensive experimental evaluation, we show that the proposed SCN outperforms related methods in terms of landmark localization error on a variety of size-limited 2D and 3D landmark localization datasets, i.e., hand radiographs, lateral cephalograms, hand MRIs, and spine CTs. Christian Payer, Darko Stern, Horst Bischof, Martin Urschler |
Medical Image Anal. | 3 |
| 2019 | Segmenting and tracking cell instances with cosine embeddings and recurrent hourglass networksabstractDifferently to semantic segmentation, instance segmentation assigns unique labels to each individual instance of the same object class. In this work, we propose a novel recurrent fully convolutional network architecture for tracking such instance segmentations over time, which is highly relevant, e.g., in biomedical applications involving cell growth and migration. Our network architecture incorporates convolutional gated recurrent units (ConvGRU) into a stacked hourglass network to utilize temporal information, e.g., from microscopy videos. Moreover, we train our network with a novel embedding loss based on cosine similarities, such that the network predicts unique embeddings for every instance throughout videos, even in the presence of dynamic structural changes due to mitosis of cells. To create the final tracked instance segmentations, the pixel-wise embeddings are clustered among subsequent video frames by using the mean shift algorithm. After showing the performance of the instance segmentation on a static in-house dataset of muscle fibers from H&E-stained microscopy images, we also evaluate our proposed recurrent stacked hourglass network regarding instance segmentation and tracking performance on six datasets from the ISBI celltracking challenge, where it delivers state-of-the-art results. Christian Payer, Darko Stern, Marlies Feiner, Horst Bischof, Martin Urschler |
Medical Image Anal. | 4 |
| 2019 | Guest Editorial Special Issue on Discriminative Learning for Model Optimization and Statistical InferenceabstractModel optimization and statistical inference have played a central role in various applications of computational intelligence, data analytics, and computer vision. Traditional approaches are usually based on model-centric learning. That is, even after model training, it is still required to design proper algorithms and to specify hand-crafted parameters for optimization and inference. Recently, discriminative learning has demonstrated its power for process-centric learning. Taking domain expertise and problem structure into account, problem-specific deep architectures can be formed by unfolding the model inference as an iterative process, and the parameters of the optimization process can then be learned from training data. These solutions are closely related with bilevel optimization, partial differential equation (PDE), as well as meta learning, and can provide new insights into the studies of versatile statistical and optimization models, such as sparse representation, structured regression, and conditional random fields. Moreover, generic deep network architectures are often referred to as “black-box” methods, while discriminative process-centric learning can provide a new perspective for the understanding and development of generic deep architectures. To sum up, connecting discriminative learning with model optimization and inference is not only helpful in analyzing convergence and generalization of deep architectures but also offers new perspectives for understanding and developing generic deep learning models. Wangmeng Zuo, Xi Peng 0001, Ling Shao 0001, Danil V. Prokhorov, Horst Bischof |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | An Intent-Based Automated Traffic Light for PedestriansabstractWe propose a fully automated, vision-based traffic light for pedestrians. Traditional industrial solutions only report people standing in a constrained waiting zone near the crosswalk. However, reporting only people below the traffic light does not allow for efficient traffic scheduling. For example, some pedestrians do not want to cross the street and walk past the traffic light, or just wait for another person to arrive. In contrast, our system leverages intent prediction to estimate which pedestrians are actually going to cross the road by analyzing both short-term and long-term trajectory cues. In this way, we can decrease the waiting times and pave the road for optimal and adaptive traffic light scheduling. We conduct a long-term evaluation in a European capital that proves the applicability and reliability of our system and demonstrates that it is not only able to replace existing push-button solutions but also yields additional information that can be used to further optimize traffic light scheduling. Christian Ertler, Horst Possegger, Michael Opitz, Horst Bischof |
AVSS | 4 |
| 2018 | Learning Pose Specific Representations by Predicting Different ViewsabstractThe labeled data required to learn pose estimation for articulated objects is difficult to provide in the desired quantity, realism, density, and accuracy. To address this issue, we develop a method to learn representations, which are very specific for articulated poses, without the need for labeled training data. We exploit the observation that the object pose of a known object is predictive for the appearance in any known view. That is, given only the pose and shape parameters of a hand, the hand's appearance from any viewpoint can be approximated. To exploit this observation, we train a model that - given input from one view - estimates a latent representation, which is trained to be predictive for the appearance of the object when captured from another viewpoint. Thus, the only necessary supervision is the second view. The training process of this model reveals an implicit pose representation in the latent space. Importantly, at test time the pose representation can be inferred using only a single view. In qualitative and quantitative experiments we show that the learned representations capture detailed pose information. Moreover, when training the proposed method jointly with labeled and unlabeled data, it consistently surpasses the performance of its fully supervised counterpart, while reducing the amount of needed labeled samples by at least one order of magnitude. Georg Poier, David Schinagl, Horst Bischof |
CVPR | 3 |
| 2018 | Semantically Aware Urban 3D Reconstruction with Plane-Based Regularization
Thomas Holzmann, Michael Maurer, Friedrich Fraundorfer, Horst Bischof |
ECCV (14) | 4 |
| 2018 | Instance Segmentation and Tracking with Cosine Embeddings and Recurrent Hourglass Networks
Christian Payer, Darko Stern, Thomas Neff, Horst Bischof, Martin Urschler |
MICCAI (2) | 4 |
| 2018 | Guest Editorial Introduction to the Special Issue on Large Scale and Nonlinear Similarity Learning for Intelligent Video AnalysisabstractLearning similarity and distance measures has become increasingly important for the analysis, matching, retrieval, recognition, and categorization of video and multimedia data. With the ubiquitous use of digital imaging devices, mobile terminals and social networks, there are massive volumes of heterogeneous and homogeneous video and multimedia data from multiple sources, views, and domains, e.g., news media websites, microblog, mobile phone, social networking, etc. Similarity and distance-based constraints can also be extended and incorporated to boost classification and relationship learning. Moreover, the spatio-temporal coherence among video data can also be utilized for self-supervised learning of similarity and distance metrics. This trend has brought several challenging issues for developing similarity and metric learning methods for large scale and weakly annotated data, where outliers and incorrectly annotated data are inevitable. Recently, scalability has been investigated to cope with lightweight and large scale metric learning, while nonlinear similarity models have shown their great potentials in learning invariant representation and nonlinear measures of video and multimedia data. Wangmeng Zuo, Liang Lin 0004, Alan L. Yuille, Horst Bischof, Lei Zhang 0006, Fatih Porikli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Spatiotemporal Saliency Estimation by Spectral Foreground DetectionabstractWe present a novel approach for spatiotemporal saliency detection by optimizing a unified criterion of color contrast, motion contrast, appearance, and background cues. To this end, we first abstract the video by temporal superpixels. Second, we propose a novel graph structure exploiting the saliency cues to assign the edge weights. The salient segments are then extracted by applying a spectral foreground detection method, quantum cuts, on this graph. We evaluate our approach on several public datasets for video saliency and activity localization to demonstrate the favorable performance of the proposed video quantum cuts compared to the state of the art. Çaglar Aytekin, Horst Possegger, Thomas Mauthner, Serkan Kiranyaz, Horst Bischof, Moncef Gabbouj |
IEEE Trans. Multim. | 5 |
| 2017 | OctNetFusion: Learning Depth Fusion from DataabstractIn this paper, we present a learning based approach to depth fusion, i.e., dense 3D reconstruction from multiple depth images. The most common approach to depth fusion is based on averaging truncated signed distance functions, which was originally proposed by Curless and Levoy in 1996. While this method is simple and provides great results, it is not able to reconstruct (partially) occluded surfaces and requires a large number frames to filter out sensor noise and outliers. Motivated by the availability of large 3D model repositories and recent advances in deep learning, we present a novel 3D CNN architecture that learns to predict an implicit surface representation from the input depth maps. Our learning based method significantly outperforms the traditional volumetric fusion approach in terms of noise reduction and outlier suppression. By learning the structure of real world 3D objects and scenes, our approach is further able to reconstruct occluded regions and to fill in gaps in the reconstruction. We demonstrate that our learning based approach outperforms both vanilla TSDF fusion as well as TV-L1 fusion on the task of volumetric fusion. Further, we demonstrate state-of-the-art 3D shape completion results. Gernot Riegler, Ali O. Ulusoy, Horst Bischof, Andreas Geiger 0001 |
3DV | 3 |
| 2017 | Primitive-based Surface Regularization for Urban 3D Reconstruction
Thomas Holzmann, Martin R. Oswald, Marc Pollefeys, Friedrich Fraundorfer, Horst Bischof |
BMVC | 5 |
| 2017 | Scalable Surface Reconstruction from Point Clouds with Extreme Scale and Density DiversityabstractIn this paper we present a scalable approach for robustly computing a 3D surface mesh from multi-scale multi-view stereo point clouds that can handle extreme jumps of point density (in our experiments three orders of magnitude). The backbone of our approach is a combination of octree data partitioning, local Delaunay tetrahedralization and graph cut optimization. Graph cut optimization is used twice, once to extract surface hypotheses from local Delaunay tetrahedralizations and once to merge overlapping surface hypotheses even when the local tetrahedralizations do not share the same topology. This formulation allows us to obtain a constant memory consumption per sub-problem while at the same time retaining the density independent interpolation properties of the Delaunay-based optimization. On multiple public datasets, we demonstrate that our approach is highly competitive with the state-of-the-art in terms of accuracy, completeness and outlier resilience. Further, we demonstrate the multi-scale potential of our approach by processing a newly recorded dataset with 2 billion points and a point density variation of more than four orders of magnitude - requiring less than 9GB of RAM per process. Christian Mostegel, Rudolf Prettenthaler, Friedrich Fraundorfer, Horst Bischof |
CVPR | 4 |
| 2017 | BIER - Boosting Independent Embeddings RobustlyabstractLearning similarity functions between image pairs with deep neural networks yields highly correlated activations of large embeddings. In this work, we show how to improve the robustness of embeddings by exploiting independence in ensembles. We divide the last embedding layer of a deep network into an embedding ensemble and formulate training this ensemble as an online gradient boosting problem. Each learner receives a reweighted training sample from the previous learners. This leverages large embedding sizes more effectively by significantly reducing correlation of the embedding and consequently increases retrieval accuracy of the embedding. Our method does not introduce any additional parameters and works with any differentiable loss function. We evaluate our metric learning method on image retrieval tasks and show that it improves over state-ofthe- art methods on the CUB-200-2011, Cars-196, Stanford Online Products, In-Shop Clothes Retrieval and VehicleID datasets by a significant margin. Michael Opitz, Georg Waltner, Horst Possegger, Horst Bischof |
ICCV | 4 |
| 2017 | Efficient 3D scene abstraction using line segments
Manuel Hofer, Michael Maurer, Horst Bischof |
Comput. Vis. Image Underst. | 3 |
| 2017 | Face inpainting based on high-level facial attributes
Mahdi Jampour, Chen Li 0031, Lap-Fai Yu, Kun Zhou 0001, Stephen Lin 0001, Horst Bischof |
Comput. Vis. Image Underst. | 6 |
| 2017 | Evaluations on multi-scale camera networks for precise and geo-accurate reconstructions from aerial and terrestrial images with user guidance
Markus Rumpler, Alexander Tscharf, Christian Mostegel, Shreyansh Daftry, Christof Hoppe, Rudolf Prettenthaler, Friedrich Fraundorfer, Gerhard Mayer, Horst Bischof |
Comput. Vis. Image Underst. | 9 |
| 2017 | Pose-specific non-linear mappings in feature space towards multiview facial expression recognition
Mahdi Jampour, Vincent Lepetit, Thomas Mauthner, Horst Bischof |
Image Vis. Comput. | 4 |
| 2016 | Regularized 3D Modeling from Noisy Building ReconstructionsabstractIn this paper, we present a method for regularizing noisy 3D reconstructions, which is especially well suited for scenes containing planar structures like buildings. At horizontal structures, the input model is divided into slices and for each slice, an inside/outside labeling is computed. With the outlines of each slice labeling, we create an irregularly shaped volumetric cell decomposition of the whole scene. Then, an optimized inside/outside labeling of these cells is computed by solving an energy minimization problem. For the cell labeling optimization we introduce a novel smoothness term, where lines in the images are used to improve the regularization result. We show that our approach can take arbitrary dense meshed point clouds as input and delivers well regularized building models, which can be textured afterwards. Thomas Holzmann, Friedrich Fraundorfer, Horst Bischof |
3DV | 3 |
| 2016 | Efficient Model Averaging for Deep Neural Networks
Michael Opitz, Horst Possegger, Horst Bischof |
ACCV (2) | 3 |
| 2016 | A Deep Primal-Dual Network for Guided Depth Super-Resolution
Gernot Riegler, David Ferstl, Matthias Rüther, Horst Bischof |
BMVC | 4 |
| 2016 | Using Self-Contradiction to Learn Confidence Measures in Stereo VisionabstractLearned confidence measures gain increasing importance for outlier removal and quality improvement in stereo vision. However, acquiring the necessary training data is typically a tedious and time consuming task that involves manual interaction, active sensing devices and/or synthetic scenes. To overcome this problem, we propose a new, flexible, and scalable way for generating training data that only requires a set of stereo images as input. The key idea of our approach is to use different view points for reasoning about contradictions and consistencies between multiple depth maps generated with the same stereo algorithm. This enables us to generate a huge amount of training data in a fully automated manner. Among other experiments, we demonstrate the potential of our approach by boosting the performance of three learned confidence measures on the KITTI2012 dataset by simply training them on a vast amount of automatically generated training data rather than a limited amount of laser ground truth data. Christian Mostegel, Markus Rumpler, Friedrich Fraundorfer, Horst Bischof |
CVPR | 4 |
| 2016 | Grid Loss: Detecting Occluded Faces
Michael Opitz, Georg Waltner, Georg Poier, Horst Possegger, Horst Bischof |
ECCV (3) | 5 |
| 2016 | ATGV-Net: Accurate Depth Super-Resolution
Gernot Riegler, Matthias Rüther, Horst Bischof |
ECCV (3) | 3 |
| 2016 | Regressing Heatmaps for Multiple Landmark Localization Using CNNs
Christian Payer, Darko Stern, Horst Bischof, Martin Urschler |
MICCAI (2) | 3 |
| 2016 | Special Issue on Visual Tracking
Xue Mei, Tianzhu Zhang 0001, Huchuan Lu, Ming-Hsuan Yang 0001, Kyoung Mu Lee, Horst Bischof |
Comput. Vis. Image Underst. | 6 |
| 2015 | Learning Depth Calibration of Time-of-Flight CamerasabstractWe present a novel method for an automatic calibration of modern consumer Timeof-Flight (ToF) cameras. Usually, these sensors come equipped with an integrated color camera. Albeit they deliver acquisitions at high frame rates they usually suffer from incorrect calibration and low accuracy due to multiple error sources. Using information from both cameras together with a simple planar target, we will show how to accurately calibrate both color and depth camera, and tackle most error sources inherent to ToF technology in a unified calibration framework. Automatic feature detection minimizes user interaction during calibration. We utilize a Random Regression Forest to optimize the manufacturer supplied depth measurements. We show the improvements to commonly used depth calibration methods in a qualitative and quantitative evaluation on multiple scenes acquired by an accurate reference system for the application of dense 3D reconstruction. David Ferstl, Christian Reinbacher, Gernot Riegler, Matthias Rüther, Horst Bischof |
BMVC | 5 |
| 2015 | Hybrid One-Shot 3D Hand Pose Estimation by Exploiting UncertaintiesabstractModel-based approaches to 3D hand tracking have been shown to perform well in a wide range of scenarios. However, they require initialisation and cannot recover easily from tracking failures that occur due to fast hand motions. Data-driven approaches, on the other hand, can quickly deliver a solution, but the results often suffer from lower accuracy or missing anatomical validity compared to those obtained from model-based approaches. In this work we propose a hybrid approach for hand pose estimation from a single depth image. First, a learned regressor is employed to deliver multiple initial hypotheses for the 3D position of each hand joint. Subsequently, the kinematic parameters of a 3D hand model are found by deliberately exploiting the inherent uncertainty of the inferred joint proposals. This way, the method provides anatomically valid and accurate solutions without requiring manual initialisation or suffering from track losses. Quantitative results on several standard datasets demonstrate that the proposed method outperforms state-of-the-art representatives of the model-based, data-driven and hybrid paradigms. Georg Poier, Konstantinos Roditakis, Samuel Schulter, Damien Michel, Horst Bischof, Antonis A. Argyros |
BMVC | 5 |
| 2015 | Depth Restoration via Joint Training of a Global Regression Model and CNNsabstract[1] Bredies, Kunisch and Pock. Total Generalized Variation, SIAM Journal on Imaging Sciences, 3(3):492-526, 2012 [2] Ferstl, Reinbacher, Ranftl, Ruther and Bischof. Image Guided Depth Upsampling using Anisotropic Total Generalized Variation, ICCV, 2013 [3} Kunisch and Pock. A Bilevel Optimization Approach for Parameter Learning in Variational Models. SIAM Journal on Imaging Sciences, 6(2):938-983, 2013 [4] Martull, Peris and Fukui. Realistic CG Stereo Image Dataset with Ground Truth Disparity Maps. ICPRW, 2012 [5] Ranftl and Pock. A Deep Variational Model for Image Segmentation, GCPR, 2014 References Gernot Riegler, René Ranftl, Matthias Rüther, Thomas Pock, Horst Bischof |
BMVC | 5 |
| 2015 | Encoding based saliency detection for videos and imagesabstractWe present a novel video saliency detection method to support human activity recognition and weakly supervised training of activity detection algorithms. Recent research has emphasized the need for analyzing salient information in videos to minimize dataset bias or to supervise weakly labeled training of activity detectors. In contrast to previous methods we do not rely on training information given by either eye-gaze or annotation data, but propose a fully unsupervised algorithm to find salient regions within videos. In general, we enforce the Gestalt principle of figure-ground segregation for both appearance and motion cues. We introduce an encoding approach that allows for efficient computation of saliency by approximating joint feature distributions. We evaluate our approach on several datasets, including challenging scenarios with cluttered background and camera motion, as well as salient object detection in images. Overall, we demonstrate favorable performance compared to state-of-the-art methods in estimating both ground-truth eye-gaze and activity annotations. Thomas Mauthner, Horst Possegger, Georg Waltner, Horst Bischof |
CVPR | 4 |
| 2015 | In defense of color-based model-free trackingabstractIn this paper, we address the problem of model-free online object tracking based on color representations. According to the findings of recent benchmark evaluations, such trackers often tend to drift towards regions which exhibit a similar appearance compared to the object of interest. To overcome this limitation, we propose an efficient discriminative object model which allows us to identify potentially distracting regions in advance. Furthermore, we exploit this knowledge to adapt the object representation beforehand so that distractors are suppressed and the risk of drifting is significantly reduced. We evaluate our approach on recent online tracking benchmark datasets demonstrating state-of-the-art results. In particular, our approach performs favorably both in terms of accuracy and robustness compared to recent tracking algorithms. Moreover, the proposed approach allows for an efficient implementation to enable online object tracking in real-time. Horst Possegger, Thomas Mauthner, Horst Bischof |
CVPR | 3 |
| 2015 | Event-driven stereo matching for real-time 3D panoramic visionabstractThis paper presents a stereo matching approach for a novel multi-perspective panoramic stereo vision system, making use of asynchronous and non-simultaneous stereo imaging towards real-time 3D 360° vision. The method is designed for events representing the scenes visual contrast as a sparse visual code allowing the stereo reconstruction of high resolution panoramic views. We propose a novel cost measure for the stereo matching, which makes use of a similarity measure based on event distributions. Thus, the robustness to variations in event occurrences was increased. An evaluation of the proposed stereo method is presented using distance estimation of panoramic stereo views and ground truth data. Furthermore, our approach is compared to standard stereo methods applied on event-data. Results show that we obtain 3D reconstructions of 1024 × 3600 round views and outperform depth reconstruction accuracy of state-of-the-art methods on event data. Stephan Schraml, Ahmed Nabil Belbachir, Horst Bischof |
CVPR | 3 |
| 2015 | Fast and accurate image upscaling with super-resolution forestsabstractThe aim of single image super-resolution is to reconstruct a high-resolution image from a single low-resolution input. Although the task is ill-posed it can be seen as finding a non-linear mapping from a low to high-dimensional space. Recent methods that rely on both neighborhood embedding and sparse-coding have led to tremendous quality improvements. Yet, many of the previous approaches are hard to apply in practice because they are either too slow or demand tedious parameter tweaks. In this paper, we propose to directly map from low to high-resolution patches using random forests. We show the close relation of previous work on single image super-resolution to locally linear regression and demonstrate how random forests nicely fit into this framework. During training the trees, we optimize a novel and effective regularized objective that not only operates on the output space but also on the input space, which especially suits the regression task. During inference, our method comprises the same well-known computational efficiency that has made random forests popular for many computer vision problems. In the experimental part, we demonstrate on standard benchmarks for single image super-resolution that our approach yields highly accurate state-of-the-art results, while being fast in both training and evaluation. Samuel Schulter, Christian Leistner, Horst Bischof |
CVPR | 3 |
| 2015 | Variational Depth Superresolution Using Example-Based Edge RepresentationsabstractIn this paper we propose a novel method for depth image superresolution which combines recent advances in example based upsampling with variational superresolution based on a known blur kernel. Most traditional depth superresolution approaches try to use additional high resolution intensity images as guidance for superresolution. In our method we learn a dictionary of edge priors from an external database of high and low resolution examples. In a novel variational sparse coding approach this dictionary is used to infer strong edge priors. Additionally to the traditional sparse coding constraints the difference in the overlap of neighboring edge patches is minimized in our optimization. These edge priors are used in a novel variational superresolution as anisotropic guidance of the higher order regularization. Both the sparse coding and the variational superresolution of the depth are solved based on a primal-dual formulation. In an exhaustive numerical and visual evaluation we show that our method clearly outperforms existing approaches on multiple real and synthetic datasets. David Ferstl, Matthias Rüther, Horst Bischof |
ICCV | 3 |
| 2015 | Conditioned Regression Models for Non-blind Single Image Super-ResolutionabstractSingle image super-resolution is an important task in the field of computer vision and finds many practical applications. Current state-of-the-art methods typically rely on machine learning algorithms to infer a mapping from low-to high-resolution images. These methods use a single fixed blur kernel during training and, consequently, assume the exact same kernel underlying the image formation process for all test images. However, this setting is not realistic for practical applications, because the blur is typically different for each test image. In this paper, we loosen this restrictive constraint and propose conditioned regression models (including convolutional neural networks and random forests) that can effectively exploit the additional kernel information during both, training and inference. This allows for training a single model, while previous methods need to be re-trained for every blur kernel individually to achieve good results, which we demonstrate in our evaluations. We also empirically show that the proposed conditioned regression models (i) can effectively handle scenarios where the blur kernel is different for each image and (ii) outperform related approaches trained for only a single kernel. Gernot Riegler, Samuel Schulter, Matthias Rüther, Horst Bischof |
ICCV | 4 |
| 2015 | Building with drones: Accurate 3D facade reconstruction using MAVsabstractAutomatic reconstruction of 3D models from images using multi-view Structure-from-Motion methods has been one of the most fruitful outcomes of computer vision. These advances combined with the growing popularity of Micro Aerial Vehicles as an autonomous imaging platform, have made 3D vision tools ubiquitous for large number of Architecture, Engineering and Construction applications among audiences, mostly unskilled in computer vision. However, to obtain high-resolution and accurate reconstructions from a large-scale object using SfM, there are many critical constraints on the quality of image data, which often become sources of inaccuracy as the current 3D reconstruction pipelines do not facilitate the users to determine the fidelity of input data during the image acquisition. In this paper, we present and advocate a closed-loop interactive approach that performs incremental reconstruction in real-time and gives users an online feedback about the quality parameters like Ground Sampling Distance (GSD), image redundancy, etc on a surface mesh. We also propose a novel multi-scale camera network design to prevent scene drift caused by incremental map building, and release the first multi-scale image sequence dataset as a benchmark. Further, we evaluate our system on real outdoor scenes, and show that our interactive pipeline combined with a multi-scale camera network approach provides compelling accuracy in multi-view reconstruction tasks when compared against the state-of-the-art methods. Shreyansh Daftry, Christof Hoppe, Horst Bischof |
ICRA | 3 |
| 2014 | aTGV-SF: Dense Variational Scene Flow through Projective Warping and Higher Order RegularizationabstractIn this paper we present a novel method to accurately estimate the dense 3D motion field, known as scene flow, from depth and intensity acquisitions. The method is formulated as a convex energy optimization, where the motion warping of each scene point is estimated through a projection and back-projection directly in 3D space. We utilize higher order regularization which is weighted and directed according to the input data by an anisotropic diffusion tensor. Our formulation enables the calculation of a dense flow field which does not penalize smooth and non-rigid movements while aligning motion boundaries with strong depth boundaries. An efficient parallelization of the numerical algorithm leads to runtimes in the order of 1s and therefore enables the method to be used in a variety of applications. We show that this novel scene flow calculation outperforms existing approaches in terms of speed and accuracy. Furthermore, we demonstrate applications such as camera pose estimation and depth image super resolution, which are enabled by the high accuracy of the proposed method. We show these applications using modern depth sensors such as Microsoft Kinect or the PMD Nano Time-of-Flight sensor. David Ferstl, Christian Reinbacher, Gernot Riegler, Matthias Rüther, Horst Bischof |
3DV | 5 |
| 2014 | Improving Sparse 3D Models for Man-Made Environments Using Line-Based 3D ReconstructionabstractTraditional Structure-from-Motion (SfM) approaches work well for richly textured scenes with a high number of distinctive feature points. Since man-made environments often contain texture less objects, the resulting point cloud suffers from a low density in corresponding scene parts. The missing 3D information heavily affects all kinds of subsequent post-processing tasks (e.g. Meshing), and significantly decreases the visual appearance of the resulting 3D model. We propose a novel 3D reconstruction approach, which uses the output of conventional SfM pipelines to generate additional complementary 3D information, by exploiting line segments. We use appearance-less epipolar guided line matching to create a potentially large set of 3D line hypotheses, which are then verified using a global graph clustering procedure. We show that our proposed method outperforms the current state-of-the-art in terms of runtime and accuracy, as well as visual appearance of the resulting reconstructions. Manuel Hofer, Michael Maurer, Horst Bischof |
3DV | 3 |
| 2014 | CP-Census: A Novel Model for Dense Variational Scene Flow from RGB-D Data
David Ferstl, Gernot Riegler, Matthias Rüther, Horst Bischof |
BMVC | 4 |
| 2014 | Semi-Global 3D Line Modeling for Incremental Structure-from-Motion
Manuel Hofer, Michael Donoser, Horst Bischof |
BMVC | 3 |
| 2014 | Hough Networks for Head Pose Estimation and Facial Feature Localization
Gernot Riegler, David Ferstl, Matthias Rüther, Horst Bischof |
BMVC | 4 |
| 2014 | Occlusion Geodesics for Online Multi-object TrackingabstractRobust multi-object tracking-by-detection requires the correct assignment of noisy detection results to object trajectories. We address this problem by proposing an online approach based on the observation that object detectors primarily fail if objects are significantly occluded. In contrast to most existing work, we only rely on geometric information to efficiently overcome detection failures. In particular, we exploit the spatio-temporal evolution of occlusion regions, detector reliability, and target motion prediction to robustly handle missed detections. In combination with a conservative association scheme for visible objects, this allows for real-time tracking of multiple objects from a single static camera, even in complex scenarios. Our evaluations on publicly available multi-object tracking benchmark datasets demonstrate favorable performance compared to the state-of-the-art in online and offline multi-object tracking. Horst Possegger, Thomas Mauthner, Peter M. Roth, Horst Bischof |
CVPR | 4 |
| 2014 | Accurate Object Detection with Joint Classification-Regression Random ForestsabstractIn this paper, we present a novel object detection approach that is capable of regressing the aspect ratio of objects. This results in accurately predicted bounding boxes having high overlap with the ground truth. In contrast to most recent works, we employ a Random Forest for learning a template-based model but exploit the nature of this learning algorithm to predict arbitrary output spaces. In this way, we can simultaneously predict the object probability of a window in a sliding window approach as well as regress its aspect ratio with a single model. Furthermore, we also exploit the additional information of the aspect ratio during the training of the Joint Classification-Regression Random Forest, resulting in better detection models. Our experiments demonstrate several benefits: (i) Our approach gives competitive results on standard detection benchmarks. (ii) The additional aspect ratio regression delivers more accurate bounding boxes than standard object detection approaches in terms of overlap with ground truth, especially when tightening the evaluation criterion. (iii) The detector itself becomes better by only including the aspect ratio information during training. Samuel Schulter, Christian Leistner, Paul Wohlhart, Peter M. Roth, Horst Bischof |
CVPR | 5 |
| 2014 | Active monocular localization: Towards autonomous monocular exploration for multirotor MAVsabstractThe main contribution of this paper is to bridge the gap between passive monocular SLAM and autonomous robotic systems. While passive monocular SLAM strives to reconstruct the scene and determine the current camera pose for any given camera motion, not every camera motion is equally suited for these tasks. In this work we propose methods to evaluate the quality of camera motions with respect to the generation of new useful map points and localization maintenance. In our experiments, we demonstrate the effectiveness of our measures using a low-cost quadrocopter. The proposed system only requires a single passive camera as exteroceptive sensor. Due to its explorative nature, the system achieves autonomous way-point navigation in challenging, unknown, GPS-denied environments. Christian Mostegel, Andreas Wendel, Horst Bischof |
ICRA | 3 |
| 2014 | Towards Automatic Bone Age Estimation from MRI: Localization of 3D Anatomical Landmarks
Thomas Ebner, Darko Stern, Rene Donner, Horst Bischof, Martin Urschler |
MICCAI (2) | 4 |
| 2014 | Fully Automatic Bone Age Estimation from Left Hand MR Images
Darko Stern, Thomas Ebner, Horst Bischof, Sabine Grassegger, Thomas Ehammer, Martin Urschler |
MICCAI (2) | 3 |
| 2014 | Structured Labels in Random Forests for Semantic Labelling and Object DetectionabstractEnsembles of randomized decision trees, known as Random Forests, have become a valuable machine learning tool for addressing many computer vision problems. Despite their popularity, few works have tried to exploit contextual and structural information in random forests in order to improve their performance. In this paper, we propose a simple and effective way to integrate contextual information in random forests, which is typically reflected in the structured output space of complex problems like semantic image labelling. Our paper has several contributions: We show how random forests can be augmented with structured label information and be used to deliver structured low-level predictions. The learning task is carried out by employing a novel split function evaluation criterion that exploits the joint distribution observed in the structured label space. This allows the forest to learn typical label transitions between object classes and avoid locally implausible label configurations. We provide two approaches for integrating the structured output predictions obtained at a local level from the forest into a concise, global, semantic labelling. We integrate our new ideas also in the Hough-forest framework with the view of exploiting contextual information at the classification level to improve the performance on the task of object detection. Finally, we provide experimental evidence for the effectiveness of our approach on different tasks: Semantic image labelling on the challenging MSRCv2 and CamVid databases, reconstruction of occluded handwritten Chinese characters on the Kaist database and pedestrian detection on the TU Darmstadt databases. Peter Kontschieder, Samuel Rota Bulò, Marcello Pelillo, Horst Bischof |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Augmented Reality for Construction Site Monitoring and DocumentationabstractAugmented reality (AR) allows for an on-site presentation of information that is registered to the physical environment. Applications from civil engineering, which require users to process complex information, are among those which can benefit particularly highly from such a presentation. In this paper, we will describe how to use AR to support monitoring and documentation of construction site progress. For these tasks, the responsible staff usually requires fast and comprehensible access to progress information to enable comparison to the as-built status as well as to as-planned data. Instead of tediously searching and mapping related information to the actual construction site environment, our AR system allows for the access of information right where it is needed. This is achieved by superimposing progress as well as as-planned information onto the user's view of the physical environment. For this purpose, we present an approach that uses aerial 3-D reconstruction to automatically capture progress information and a mobile AR client for on-site visualization. Within this paper, we will describe in greater detail how to capture 3-D, how to register the AR system within the physical outdoor environment, how to visualize progress information in a comprehensible way in an AR overlay, and how to interact with this kind of information. By implementing such an AR system, we are able to provide an overview about the possibilities and future applications of AR in the construction industry. Stefanie Zollmann, Christof Hoppe, Stefan Kluckner, Christian Poglitsch, Horst Bischof, Gerhard Reitmayr |
Proc. IEEE | 5 |
| 2013 | Flexible and User-Centric Camera Calibration using Planar Fiducial MarkersabstractThe benefit of accurate camera calibration for recovering 3D structure from images is a well-studied topic. Recently 3D vision tools for end-user applications have become popular among large audiences, mostly unskilled in computer vision. This motivates the need for a flexible and user-centric camera calibration method which drastically releases the critical requirements on the calibration target and ensures that low-quality or faulty images provided by end users do not degrade the overall calibration and in effect the resulting 3D model. In this paper we present and advocate an approach to camera cal-ibration using fiducial markers, aiming at the accuracy of target calibration techniques without the requirement for a precise calibration pattern, to ease the calibration effort for the end-user. An extensive set of experiments with real images is presented which demonstrates improvements in the estimation of the parameters of the camera model as well as accuracy in the multi-view stereo reconstruction of large scale scenes. Pixel re-projection errors and ground truth errors obtained by our method are significantly lower compared to popular calibration routines, even though paper-printable and easy-to-use targets are employed. 1 Shreyansh Daftry, Michael Maurer, Andreas Wendel, Horst Bischof |
BMVC | 4 |
| 2013 | Incremental Line-based 3D Reconstruction using Geometric ConstraintsabstractGenerating accurate 3D models for man-made environments can be a challenging task due to the presence of texture-less objects or wiry structures. Since traditional point-based 3D reconstruction approaches may fail to integrate these structures into the resulting point cloud, a different feature representation is necessary. We present a novel approach which uses point features for camera estimation and additional line segments for 3D reconstruction. To avoid appearance-based line matching, we use purely geometric constraints for hypothesis generation and verification. Therefore, the proposed method is able to reconstruct both wiry structures as well as solid objects. The algorithm is designed to generate incremental results using online Structure-from-Motion and linebased 3D modelling in parallel. We show that the proposed method outperforms previous descriptor-less line matching approaches in terms of run-time while delivering accurate Manuel Hofer, Andreas Wendel, Horst Bischof |
BMVC | 3 |
| 2013 | Incremental Surface Extraction from Sparse Structure-from-Motion Point CloudsabstractIn this paper we propose a new method to incrementally extract a surface from a consecutively growing Structure-from-Motion (SfM) point cloud in real-time. Our method is based on a Delaunay triangulation (DT) on the 3D points. The core idea is to robustly label all tetrahedra into freeand occupied space using a random field formulation and to extract the surface as the interface between differently labeled tetrahedra. For this reason, we propose a new energy function that achieves the same accuracy as state-of-the-art methods but reduces the computational effort significantly. Furthermore, our new formulation allows us to extract the surface in an incremental manner, i. e. whenever the point cloud is updated we adapt our energy function. Instead of minimizing the updated energy with a standard graph cut, we employ the dynamic graph cut of Kohli et al. [1] which enables efficient minimization of a series of similar random fields by re-using the previous solution. In such a way we are able to extract the surface from an increasingly growing point cloud nearly independent of the overall scene size. Energy Function for Surface Extraction Our method formulates surface extraction as a binary labeling problem, with the goal of assigning each tetrahedron either a free or occupied label. For this reason, we model the probabilities that a tetrahedron is free- or occupied space analyzing the set of rays that connect all 3D points to image features. Following the idea of the truncated signed distance function (TSDF), which is known from voxel-based surface reconstructions, a tetrahedron in front of a 3D point X has a high probability to be free space, whereas a tetrahedron behind X is presumably occupied space. We further assume that it is very unlikely that neighboring tetrahedra obtain different labels, except for pairs of tetrahedra that have a ray through the face connecting both. Such a labeling problem can be elegantly formulated as a pairwise random field and since our priors are submodular, we can efficiently find a global optimal labeling solution e. g. using graph cuts. In contrast to existing methods like [2], our energy depends only on the visibility information that is directly connected to the four 3D points that span the tetrahedraVi. Hence a modification of the tetrahedral structure by inserting new points has only limited effect on the energy function. This property enables us to easily adopt the energy function to a modified tetrahedral structure. Incremental Surface Extraction To enable efficient incremental surface reconstruction, our method has to consecutively integrate new scene information (3D points as well as visibility information) in the energy function and to minimize the modified energy efficiently. Integrating new visibility information, i. e. adding rays for newly available 3D points, affects only those terms of the energy function that relate Christof Hoppe, Manfred Klopschitz, Michael Donoser, Horst Bischof |
BMVC | 4 |
| 2013 | Unsupervised Object Discovery and Segmentation in VideosabstractUnsupervised object discovery is the task of finding recurring objects over an unsorted set of images without any human supervision, which becomes more and more important as the amount of visual data grows exponentially. Existing approaches typically build on still images and rely on different prior knowledge to yield accurate results. In contrast, we propose a novel video-based approach, allowing also for exploiting motion information, which is a strong and physically valid indicator for foreground objects, thus, tremendously easing the task. In particular, we show how to integrate motion information in parallel with appearance cues into a common conditional random field formulation to automatically discover object categories from videos. In the experiments, we show that our system can successfully extract, group, and segment most foreground objects and is also able to discover stationary objects in the given videos. Furthermore, we demonstrate that the unsupervised learned appearance models also yield reasonable results for object detection on still images. Samuel Schulter, Christian Leistner, Peter M. Roth, Horst Bischof |
BMVC | 4 |
| 2013 | Diffusion Processes for Retrieval RevisitedabstractIn this paper we revisit diffusion processes on affinity graphs for capturing the intrinsic manifold structure defined by pair wise affinity matrices. Such diffusion processes have already proved the ability to significantly improve subsequent applications like retrieval. We give a thorough overview of the state-of-the-art in this field and discuss obvious similarities and differences. Based on our observations, we are then able to derive a generic framework for diffusion processes in the scope of retrieval applications, where the related work represents specific instances of our generic formulation. We evaluate our framework on several retrieval tasks and are able to derive algorithms that e.\, g.~achieve a 100\% bulls eye score on the popular MPEG7 shape retrieval data set. Michael Donoser, Horst Bischof |
CVPR | 2 |
| 2013 | Robust Real-Time Tracking of Multiple Objects by Volumetric Mass DensitiesabstractCombining foreground images from multiple views by projecting them onto a common ground-plane has been recently applied within many multi-object tracking approaches. These planar projections introduce severe artifacts and constrain most approaches to objects moving on a common 2D ground-plane. To overcome these limitations, we introduce the concept of an occupancy volume - exploiting the full geometry and the objects' center of mass - and develop an efficient algorithm for 3D object tracking. Individual objects are tracked using the local mass density scores within a particle filter based approach, constrained by a Voronoi partitioning between nearby trackers. Our method benefits from the geometric knowledge given by the occupancy volume to robustly extract features and train classifiers on-demand, when volumetric information becomes unreliable. We evaluate our approach on several challenging real-world scenarios including the public APIDIS dataset. Experimental evaluations demonstrate significant improvements compared to state-of-the-art methods, while achieving real-time performance. Horst Possegger, Sabine Sternig, Thomas Mauthner, Peter M. Roth, Horst Bischof |
CVPR | 5 |
| 2013 | Alternating Decision ForestsabstractThis paper introduces a novel classification method termed Alternating Decision Forests (ADFs), which formulates the training of Random Forests explicitly as a global loss minimization problem. During training, the losses are minimized via keeping an adaptive weight distribution over the training samples, similar to Boosting methods. In order to keep the method as flexible and general as possible, we adopt the principle of employing gradient descent in function space, which allows to minimize arbitrary losses. Contrary to Boosted Trees, in our method the loss minimization is an inherent part of the tree growing process, thus allowing to keep the benefits of common Random Forests, such as, parallel processing. We derive the new classifier and give a discussion and evaluation on standard machine learning data sets. Furthermore, we show how ADFs can be easily integrated into an object detection application. Compared to both, standard Random Forests and Boosted Trees, ADFs give better performance in our experiments, while yielding more compact models in terms of tree depth. Samuel Schulter, Paul Wohlhart, Christian Leistner, Amir Saffari, Peter M. Roth, Horst Bischof |
CVPR | 6 |
| 2013 | Optimizing 1-Nearest Prototype ClassifiersabstractThe development of complex, powerful classifiers and their constant improvement have contributed much to the progress in many fields of computer vision. However, the trend towards large scale datasets revived the interest in simpler classifiers to reduce runtime. Simple nearest neighbor classifiers have several beneficial properties, such as low complexity and inherent multi-class handling, however, they have a runtime linear in the size of the database. Recent related work represents data samples by assigning them to a set of prototypes that partition the input feature space and afterwards applies linear classifiers on top of this representation to approximate decision boundaries locally linear. In this paper, we go a step beyond these approaches and purely focus on 1-nearest prototype classification, where we propose a novel algorithm for deriving optimal prototypes in a discriminative manner from the training samples. Our method is implicitly multi-class capable, parameter free, avoids noise over fitting and, since during testing only comparisons to the derived prototypes are required, highly efficient. Experiments demonstrate that we are able to outperform related locally linear methods, while even getting close to the results of more complex classifiers. Paul Wohlhart, Martin Köstinger, Michael Donoser, Peter M. Roth, Horst Bischof |
CVPR | 5 |
| 2013 | Multi-modality depth map fusion using primal-dual optimizationabstractWe present a novel fusion method that combines complementary 3D and 2D imaging techniques. Consider a Time-of-Flight sensor that acquires a dense depth map on a wide depth range but with a comparably small resolution. Complementary, a stereo sensor generates a disparity map in high resolution but with occlusions and outliers. In our method, we fuse depth data, and optionally also intensity data using a primal-dual optimization, with an energy functional that is designed to compensate for missing parts, filter strong outliers and reduce the acquisition noise. The numerical algorithm is efficiently implemented on a GPU to achieve a processing speed of 10 to 15 frames per second. Experiments on synthetic, real and benchmark datasets show that the results are superior compared to each sensor alone and to competing optimization techniques. In a practical example, we are able to fuse a Kinect triangulation sensor and a small size Time-of-Flight camera to create a gaming sensor with superior resolution, acquisition range and accuracy. David Ferstl, René Ranftl, Matthias Rüther, Horst Bischof |
ICCP | 4 |
| 2013 | Image Guided Depth Upsampling Using Anisotropic Total Generalized VariationabstractIn this work we present a novel method for the challenging problem of depth image up sampling. Modern depth cameras such as Kinect or Time-of-Flight cameras deliver dense, high quality depth measurements but are limited in their lateral resolution. To overcome this limitation we formulate a convex optimization problem using higher order regularization for depth image up sampling. In this optimization an an isotropic diffusion tensor, calculated from a high resolution intensity image, is used to guide the up sampling. We derive a numerical algorithm based on a primal-dual formulation that is efficiently parallelized and runs at multiple frames per second. We show that this novel up sampling clearly outperforms state of the art approaches in terms of speed and accuracy on the widely used Middlebury 2007 datasets. Furthermore, we introduce novel datasets with highly accurate ground truth, which, for the first time, enable to benchmark depth up sampling methods using real sensor data. David Ferstl, Christian Reinbacher, René Ranftl, Matthias Rüther, Horst Bischof |
ICCV | 5 |
| 2013 | Joint Learning of Discriminative Prototypes and Large Margin Nearest Neighbor ClassifiersabstractIn this paper, we raise important issues concerning the evaluation complexity of existing Mahalanobis metric learning methods. The complexity scales linearly with the size of the dataset. This is especially cumbersome on large scale or for real-time applications with limited time budget. To alleviate this problem we propose to represent the dataset by a fixed number of discriminative prototypes. In particular, we introduce a new method that jointly chooses the positioning of prototypes and also optimizes the Mahalanobis distance metric with respect to these. We show that choosing the positioning of the prototypes and learning the metric in parallel leads to a drastically reduced evaluation effort while maintaining the discriminative essence of the original dataset. Moreover, for most problems our method performing k-nearest prototype (k-NP) classification on the condensed dataset leads to even better generalization compared to k-NN classification using all data. Results on a variety of challenging benchmarks demonstrate the power of our method. These include standard machine learning datasets as well as the challenging Public Figures Face Database. On the competitive machine learning benchmarks we are comparable to the state-of-the-art while being more efficient. On the face benchmark we clearly outperform the state-of-the-art in Mahalanobis metric learning with drastically reduced evaluation effort. Martin Köstinger, Paul Wohlhart, Peter M. Roth, Horst Bischof |
ICCV | 4 |
| 2013 | Alternating Regression Forests for Object Detection and Pose EstimationabstractWe present Alternating Regression Forests (ARFs), a novel regression algorithm that learns a Random Forest by optimizing a global loss function over all trees. This interrelates the information of single trees during the training phase and results in more accurate predictions. ARFs can minimize any differentiable regression loss without sacrificing the appealing properties of Random Forests, like low computational complexity during both, training and testing. Inspired by recent developments for classification [19], we derive a new algorithm capable of dealing with different regression loss functions, discuss its properties and investigate the relations to other methods like Boosted Trees. We evaluate ARFs on standard machine learning benchmarks, where we observe better generalization power compared to both standard Random Forests and Boosted Trees. Moreover, we apply the proposed regressor to two computer vision applications: object detection and head pose estimation from depth images. ARFs outperform the Random Forest baselines in both tasks, illustrating the importance of optimizing a common loss function for all trees. Samuel Schulter, Christian Leistner, Paul Wohlhart, Peter M. Roth, Horst Bischof |
ICCV | 5 |
| 2013 | Hough-based tracking of non-rigid objects
Martin Godec, Peter M. Roth, Horst Bischof |
Comput. Vis. Image Underst. | 3 |
| 2013 | Segmentation-based tracking by support fusion
Markus Heber, Martin Godec, Matthias Rüther, Peter M. Roth, Horst Bischof |
Comput. Vis. Image Underst. | 5 |
| 2013 | Global localization of 3D anatomical structures by pre-filtered Hough Forests and discrete optimizationabstractThe accurate localization of anatomical landmarks is a challenging task, often solved by domain specific approaches. We propose a method for the automatic localization of landmarks in complex, repetitive anatomical structures. The key idea is to combine three steps: (1) a classifier for pre-filtering anatomical landmark positions that (2) are refined through a Hough regression model, together with (3) a parts-based model of the global landmark topology to select the final landmark positions. During training landmarks are annotated in a set of example volumes. A classifier learns local landmark appearance, and Hough regressors are trained to aggregate neighborhood information to a precise landmark coordinate position. A non-parametric geometric model encodes the spatial relationships between the landmarks and derives a topology which connects mutually predictive landmarks. During the global search we classify all voxels in the query volume, and perform regression-based agglomeration of landmark probabilities to highly accurate and specific candidate points at potential landmark locations. We encode the candidates' weights together with the conformity of the connecting edges to the learnt geometric model in a Markov Random Field (MRF). By solving the corresponding discrete optimization problem, the most probable location for each model landmark is found in the query volume. We show that this approach is able to consistently localize the model landmarks despite the complex and repetitive character of the anatomical structures on three challenging data sets (hand radiographs, hand CTs, and whole body CTs), with a median localization error of 0.80 mm, 1.19 mm and 2.71 mm, respectively. Rene Donner, Bjoern Menze, Horst Bischof, Georg Langs |
Medical Image Anal. | 3 |
| 2012 | Detecting Partially Occluded Objects with an Implicit Shape Model Random Field
Paul Wohlhart, Michael Donoser, Peter M. Roth, Horst Bischof |
ACCV (1) | 4 |
| 2012 | Person Re-identification by Efficient Impostor-Based Metric LearningabstractRecognizing persons over a system of disjunct cameras is a hard task for human operators and even harder for automated systems. In particular, realistic setups show difficulties such as different camera angles or different camera properties. Additionally, also the appearance of exactly the same person can change dramatically due to different views (e.g., frontal/back) of carried objects. In this paper, we mainly address the first problem by learning the transition from one camera to the other. This is realized by learning a Mahalanobis metric using pairs of labeled samples from different cameras. Building on the ideas of Large Margin Nearest Neighbor classification, we obtain a more efficient solution which additionally provides much better generalization properties. To demonstrate these benefits, we run experiments on three different publicly available datasets, showing state-of-the-art or even better results, however, on much lower computational efforts. This is in particular interesting since we use quite simple color and texture features, whereas other approaches build on rather complex image descriptions! Martin Hirzer, Peter M. Roth, Horst Bischof |
AVSS | 3 |
| 2012 | Learning Edge-Specific Kernel Functions For Pairwise Graph Matching
Michael Donoser, Martin Urschler, Horst Bischof |
BMVC | 3 |
| 2012 | Online Feedback for Structure-from-Motion Image AcquisitionabstractThe quality and completeness of 3D models obtained by Structure-fromMotion (SfM) heavily depend on the image acquisition process. If the user gets feedback about the reconstruction quality already during the acquisition, he can optimize this process. The goal of this paper is to support a user during image acquisition by giving online feedback of the current reconstruction quality. We propose an online SfM method that integrates wide-baseline still-images in an online fashion into a consistent reconstruction and we derive a surface model given the SfM point cloud. To guide the user to scene parts that are captured not very well, we colour the mesh according to redundancy and resolution information. In the experiments, we show that our approach makes the final SfM result predictable already during image acquisition. The method is suited for large-scale reconstructions as obtained by flying micro aerial vehicles as well as on small indoor environments. We propose a method that supports a user in the acquisition process in two ways: (a) sparse online SfM with accuracy close to offline methods and (b) surface extraction and quality visualization. The workflow of our method is shown in Figure 1. Christof Hoppe, Manfred Klopschitz, Markus Rumpler, Andreas Wendel, Stefan Kluckner, Horst Bischof, Gerhard Reitmayr |
BMVC | 6 |
| 2012 | On Cross-Spectral Stereo Matching using Dense Gradient FeaturesabstractWe address the problem of scene depth recovery within cross-spectral stereo imagery (each image sensed over a differing spectral range).We compare several robust matching techniques which are able to capture local similarities between the structure of crossspectral images and a range of stereo optimisation techniques for the computation of valid depth estimates in this case.Specifically we deal with the recovery of dense depth information from thermal (far infrared spectrum) and optical (visible spectrum) image pairs where large differences in the characteristics of image pairs make this task significantly more challenging than the common stereo case.We show that the use of dense gradient features, based on Histograms of Oriented Gradient (HOG) descriptors, for pixel matching in combination with a strong match optimisation approach can produce largely valid, yet coarse, dense depth estimates suitable for object localisation or environment navigation.The proposed solution is compared and shown to work favourably against prior approaches based on using Mutual Information (MI) or Local Self-Similarity (LSS) descriptors. Peter Pinggera, Toby P. Breckon, Horst Bischof |
BMVC | 3 |
| 2012 | Discriminative Hough Forests for Object DetectionabstractObject detection models based on the Implicit Shape Model (ISM) [3] use small, local parts that vote for object centers in images. Since these parts vote completely independently from each other, this often leads to false-positive detections due to random constellations of parts. Thus, we introduce a verification step, which considers the activations of all voting elements that contribute to a detection. The levels of activation of each voting element of the ISM form a new description vector for an object hypothesis, which can be examined in order to discriminate between correct and incorrect detections. In particular, we observe the levels of activation of the voting elements in Hough Forests [2], which can be seen as a variant of ISM. In Hough Forests, the voting elements are all the positive training patches used to train the Forest. Each patch of the input image is classified by all decision trees in the Hough Forest. Whenever an input patch falls into the same leaf node as a patch from training, a certain amount of weight is added to the detection hypothesis at the relative position of the object center, which was recorded when cropping out the training patch. The total amount of weight one voting element (offset vector) adds to a detection hypothesis (the total activation) can be calculated by summing over all input patches and trees in the forest. Stacking the activations of all elements gives an activation vector for a hypothesis. We learn classifiers to discriminate correct and wrong part constellations based on these activation vectors and thus assign a better confidence to each detection. We use linear models as well as a histogram intersection kernel SVM. In the linear classifier, one weight is learned for each voting element. We additionally show how to use these weights, not only as a post processing step, but directly in the voting process. This has two advantages: First, it circumvents the explicit calculation of the activation vector for later reclassification, which is computationally more demanding. Second, the non-maxima suppression is performed on cleaner Hough maps, which allows for reducing the size of the suppression neighborhood and thus increases the recall at high levels of precision. Paul Wohlhart, Samuel Schulter, Martin Köstinger, Peter M. Roth, Horst Bischof |
BMVC | 5 |
| 2012 | Structured Local Predictors for image labellingabstractIn this paper we introduce Structured Local Predictors (SLP) — A new formulation that considers the image labelling problem from a structured learning point of view. SLP are locally operating models, which provide a per-pixel labelling by exploiting contextual relations, learned from complex interactions between labels and a customizable intermediate representation of the image data. Our first key contribution is to handle flexible configurations of pairwise interactions between image pixels while allowing them to be made arbitrarily dependent on the image data. Moreover, we pose the parameter learning process as a convex, structured-learning problem, which can be efficiently solved in a globally optimal way due to the introduction of a continuous, structured output space. Finally, we provide an interface to our model by means of a quantization space, allowing to define task-specific intermediate representations for the input data. In our experiments we demonstrate the broad applicability of our model for tasks like inpainting and semantic labelling. Samuel Rota Bulò, Peter Kontschieder, Marcello Pelillo, Horst Bischof |
CVPR | 4 |
| 2012 | Large scale metric learning from equivalence constraintsabstractIn this paper, we raise important issues on scalability and the required degree of supervision of existing Mahalanobis metric learning methods. Often rather tedious optimization procedures are applied that become computationally intractable on a large scale. Further, if one considers the constantly growing amount of data it is often infeasible to specify fully supervised labels for all data points. Instead, it is easier to specify labels in form of equivalence constraints. We introduce a simple though effective strategy to learn a distance metric from equivalence constraints, based on a statistical inference perspective. In contrast to existing methods we do not rely on complex optimization problems requiring computationally expensive iterations. Hence, our method is orders of magnitudes faster than comparable methods. Results on a variety of challenging benchmarks with rather diverse nature demonstrate the power of our method. These include faces in unconstrained environments, matching before unseen object instances and person re-identification across spatially disjoint cameras. In the latter two benchmarks we clearly outperform the state-of-the-art. Martin Köstinger, Martin Hirzer, Paul Wohlhart, Peter M. Roth, Horst Bischof |
CVPR | 5 |
| 2012 | Irregular lattices for complex shape grammar facade parsingabstractHigh-quality urban reconstruction requires more than multi-view reconstruction and local optimization. The structure of facades depends on the general layout, which has to be optimized globally. Shape grammars are an established method to express hierarchical spatial relationships, and are therefore suited as representing constraints for semantic facade interpretation. Usually inference uses numerical approximations, or hard-coded grammar schemes. Existing methods inspired by classical grammar parsing are not applicable on real-world images due to their prohibitively high complexity. This work provides feasible generic facade reconstruction by combining low-level classifiers with mid-level object detectors to infer an irregular lattice. The irregular lattice preserves the logical structure of the facade while reducing the search space to a manageable size. We introduce a novel method for handling symmetry and repetition within the generic grammar. We show competitive results on two datasets, namely the Paris 2010 and the Graz 50. The former includes only Hausmannian, while the latter includes Classicism, Biedermeier, Historicism, Art Nouveau and post-modern architectural styles. Hayko Riemenschneider, Ulrich Krispel, Wolfgang Thaller, Michael Donoser, Sven Havemann, Dieter W. Fellner, Horst Bischof |
CVPR | 7 |
| 2012 | Joint motion estimation and segmentation of complex scenes with label costs and occlusion modelingabstractWe propose a unified variational formulation for joint motion estimation and segmentation with explicit occlusion handling. This is done by a multi-label representation of the flow field, where each label corresponds to a parametric representation of the motion. We use a convex formulation of the multi-label Potts model with label costs and show that the asymmetric map-uniqueness criterion can be integrated into our formulation by means of convex constraints. Explicit occlusion handling eliminates errors otherwise created by the regularization. As occlusions can occur only at object boundaries, a large number of objects may be required. By using a fast primal-dual algorithm we are able to handle several hundred motion segments. Results are shown on several classical motion segmentation and optical flow examples. Markus Unger, Manuel Werlberger, Thomas Pock, Horst Bischof |
CVPR | 4 |
| 2012 | Dense reconstruction on-the-flyabstractWe present a novel system that is capable of generating live dense volumetric reconstructions based on input from a micro aerial vehicle. The distributed reconstruction pipeline is based on state-of-the-art approaches to visual SLAM and variational depth map fusion, and is designed to exploit the individual capabilities of the system components. Results are visualized in real-time on a tablet interface, which gives the user the opportunity to interact. We demonstrate the performance of our approach by capturing several indoor and outdoor scenes on-the-fly and by evaluating our results with respect to a ground-truth model. Andreas Wendel, Michael Maurer, Gottfried Munda, Thomas Pock, Horst Bischof |
CVPR | 5 |
| 2012 | Relaxed Pairwise Learned Metric for Person Re-identification
Martin Hirzer, Peter M. Roth, Martin Köstinger, Horst Bischof |
ECCV (6) | 4 |
| 2012 | Hough Regions for Joining Instance Localization and Segmentation
Hayko Riemenschneider, Sabine Sternig, Michael Donoser, Peter M. Roth, Horst Bischof |
ECCV (3) | 5 |
| 2012 | Simultaneous Shape and Pose Adaption of Articulated Models Using Linear Optimization
Matthias Straka, Stefan Hauswiesner, Matthias Rüther, Horst Bischof |
ECCV (1) | 4 |
| 2012 | Depth coded shape from focusabstractWe present a novel shape from focus method for high- speed shape reconstruction in optical microscopy. While the traditional shape from focus approach heavily depends on presence of surface texture, and requires a considerable amount of measurement time, our method is able to perform reconstruction from only two images. Our method relies the rapid projection of a binary pattern sequence, while object is continuously moved through the camera focus range and a single image is continuously exposed. Deconvolution of the integral image allows a direct decoding of binary pattern and its associated depth. Experiments a synthetic dataset and on real scenes show that a depth map can be reconstructed at only 3% of memory costs and fraction of the computational effort compared with traditional shape from focus. Martin Lenz, David Ferstl, Matthias Rüther, Horst Bischof |
ICCP | 4 |
| 2012 | Dense appearance modeling and efficient learning of camera transitions for person re-identificationabstractOne central task in many visual surveillance scenarios is person re-identification, i.e., recognizing an individual person across a network of spatially disjoint cameras. Most successful recognition approaches are either based on direct modeling of the human appearance or on machine learning. In this work, we aim at taking advantage of both directions of research. On the one hand side, we compute a descriptive appearance representation encoding the vertical color structure of pedestrians. To improve the classification results, we additionally estimate the transition between two cameras using a pair-wisely estimated metric. In particular, we introduce 4D spatial color histograms and adopt Large Margin Nearest Neighbor (LMNN) metric learning. The approach is demonstrated for two publicly available datasets, showing competitive results, however, on lower computational costs. Martin Hirzer, Csaba Beleznai, Martin Köstinger, Peter M. Roth, Horst Bischof |
ICIP | 5 |
| 2012 | RoμNect: Hand mounted depth sensing using a commodity gaming sensor
Christian Reinbacher, Matthias Rüther, Horst Bischof |
ICPR | 3 |
| 2012 | Geo-referenced 3D reconstruction: Fusing public geographic data and aerial imageryabstractWe present an image-based 3D reconstruction pipeline for acquiring geo-referenced semi-dense 3D models. Multiple overlapping images captured from a micro aerial vehicle platform provide a highly redundant source for multi-view reconstructions. Publicly available geo-spatial information sources are used to obtain an approximation to a digital surface model (DSM). Models obtained by the semi-dense reconstruction are automatically aligned to the DSM to allow the integration of highly detailed models into the original DSM and to provide geographic context. Michael Maurer, Markus Rumpler, Andreas Wendel, Christof Hoppe, Arnold Irschara, Horst Bischof |
ICRA | 6 |
| 2012 | Interactive 4D overview and detail visualization in augmented realityabstractIn this paper we present an approach for visualizing time-oriented data of dynamic scenes in an on-site AR view. Visualizations of time-oriented data have special challenges compared to the visualization of arbitrary virtual objects. Usually, the 4D data occludes a large part of the real scene. Additionally, the data sets from different points in time may occlude each other. Thus, it is important to design adequate visualization techniques that provide a comprehensible visualization. In this paper we introduce a visualization concept that uses overview and detail techniques to present 4D data in different detail levels. These levels provide at first an overview of the 4D scene, at second information about the 4D change of a single object and at third detailed information about object appearance and geometry for specific points in time. Combining the three levels of detail with interactive transitions such as magic lenses or distorted viewing techniques enables the user to understand the relationship between them. Finally we show how to apply this concept for construction site documentation and monitoring. Stefanie Zollmann, Denis Kalkofen, Christof Hoppe, Stefan Kluckner, Horst Bischof, Gerhard Reitmayr |
ISMAR | 5 |
| 2012 | Pushing the limits of stereo using variational stereo estimationabstractWe examine high accuracy stereo estimation for binocular sequences that where obtained from a mobile platform. The ultimate goal is to improve the range of stereo systems without altering the setup. Based on a well-known variational optical flow model, we introduce a novel stereo model that features a second-order regularization, which both allows sub-pixel accurate solutions and piecewise planar disparity maps. The model incorporates a robust fidelity term to account for adverse illumination conditions that frequently arise in real-world scenes. Using several sequences that were taken from a mobile platform we show the robustness and accuracy of the proposed model. René Ranftl, Stefan K. Gehrig, Thomas Pock, Horst Bischof |
Intelligent Vehicles Symposium | 4 |
| 2012 | Context-Sensitive Decision Forests for Object DetectionabstractIn this paper we introduce Context-Sensitive Decision Forests - A new perspective to exploit contextual information in the popular decision forest framework for the object detection problem. They are tree-structured classifiers with the ability to access intermediate prediction (here: classification and regression) information during training and inference time. This intermediate prediction is available to each sample, which allows us to develop context-based decision criteria, used for refining the prediction process. In addition, we introduce a novel split criterion which in combination with a priority based way of constructing the trees, allows more accurate regression mode selection and hence improves the current context information. In our experiments, we demonstrate improved results for the task of pedestrian detection on the challenging TUD data set when compared to state-of-the-art methods. Peter Kontschieder, Samuel Rota Bulò, Antonio Criminisi, Pushmeet Kohli, Marcello Pelillo, Horst Bischof |
NIPS | 6 |
| 2012 | Large-scale, dense city reconstruction from user-contributed photos
Arnold Irschara, Christopher Zach, Manfred Klopschitz, Horst Bischof |
Comput. Vis. Image Underst. | 4 |
| 2012 | Evolutionary Hough Games for coherent object detection
Peter Kontschieder, Samuel Rota Bulò, Michael Donoser, Marcello Pelillo, Horst Bischof |
Comput. Vis. Image Underst. | 5 |
| 2012 | Fast variational multi-view segmentation through backprojection of spatial constraints
Christian Reinbacher, Matthias Rüther, Horst Bischof |
Image Vis. Comput. | 3 |
| 2012 | On-line inverse multiple instance boosting for classifier gridsabstractClassifier grids have shown to be a considerable choice for object detection from static cameras. By applying a single classifier per image location the classifier's complexity can be reduced and more specific and thus more accurate classifiers can be estimated. In addition, by using an on-line learner a highly adaptive but stable detection system can be obtained. Even though long-term stability has been demonstrated such systems still suffer from short-term drifting if an object is not moving over a long period of time. The goal of this work is to overcome this problem and thus to increase the recall while preserving the accuracy. In particular, we adapt ideas from multiple instance learning (MIL) for on-line boosting. In contrast to standard MIL approaches, which assume an ambiguity on the positive samples, we apply this concept to the negative samples: inverse multiple instance learning. By introducing temporal bags consisting of background images operating on different time scales, we can ensure that each bag contains at least one sample having a negative label, providing the theoretical requirements. The experimental results demonstrate superior classification results in presence of non-moving objects. Sabine Sternig, Peter M. Roth, Horst Bischof |
Pattern Recognit. Lett. | 3 |
| 2011 | AVSS 2011 demo session: OUTLIER - online learning and visualization of unusual eventsabstractSummary form only given. We introduce to the surveillance community the VIRAT Video Dataset[1], which is a new large-scale surveillance video dataset designed to assess the performance of event recognition algorithms in realistic scenes1. Josef A. Birchbauer, Samuel Schulter, René Schuster, Georg Poier, Peter Schallauer, Peter M. Roth, Horst Bischof |
AVSS | 8 |
| 2011 | AVSS 2011 demo session: Construction site monitoring from highly-overlapping MAV imagesabstractSummary form only given. We report on a disruption in organizational dynamics arising from the introduction of model-driven development tools in General Motors. The introduction altered the balance of collaboration deeply, and the organization is still negotiating with its aftermath. Our report illustrates one consequence of tool adoption in groups, and that these consequences should be understood to facilitate technical change. Stefan Kluckner, Josef A. Birchbauer, Claudia Windisch, Christof Hoppe, Arnold Irschara, Andreas Wendel, Stefanie Zollmann, Gerhard Reitmayr, Horst Bischof |
AVSS | 9 |
| 2011 | Learning to recognize faces from videos and weakly related information cuesabstractVideos are often associated with additional information that could be valuable for interpretation of their content. This especially applies for the recognition of faces within video streams, where often cues such as transcripts and subtitles are available. However, this data is not completely reliable and might be ambiguously labeled. To overcome these limitations, we take advantage of semi-supervised (SSL) and multiple instance learning (MIL) and propose a new semi-supervised multiple instance learning (SSMIL) algorithm. Thus, during training we can weaken the prerequisite of knowing the label for each instance and can integrate unlabeled data, given only probabilistic information in form of priors. The benefits of the approach are demonstrated for face recognition in videos on a publicly available benchmark dataset. In fact, we show exploring new information sources can considerably improve the classification results. Martin Köstinger, Paul Wohlhart, Peter M. Roth, Horst Bischof |
AVSS | 4 |
| 2011 | Next-generation 3D visualization for visual surveillanceabstractExisting visual surveillance systems typically require that human operators observe video streams from different cameras, which becomes infeasible if the number of observed cameras is ever increasing. In this paper, we present a new surveillance system that combines automatic video analysis (i.e., single person tracking and crowd analysis) and interactive visualization. Our novel visualization takes advantage of a high resolution display and given 3D information to focus the operator's attention to interesting/ critical areas of the observed area. This is realized by embedding the results of automatic scene analysis techniques into the visualization. By providing different visualization modes, the user can easily switch between the different modes and can select the mode which provides most information. The system is demonstrated for a real setup on a university campus. Peter M. Roth, Volker Settgast, Peter Widhalm, Marcel Lancelle, Josef A. Birchbauer, Norbert Brändle, Sven Havemann, Horst Bischof |
AVSS | 8 |
| 2011 | Semantic Image Labelling as a Label Puzzle GameabstractIn this work we introduce a novel solution to the semantic image labelling problem, i.e. the task of assigning semantic object class labels to individual pixels in a test image.Conventional methods are typically relying on random fields for modelling interactions between neighboring pixels and obtaining smooth labelling results using unary and pairwise cost functions.Instead, we consider the labelling problem as a puzzle game, where the final labelling is obtained by assembling discriminatively learned candidate sets of label puzzle pieces, each representing a topological and semantically plausible label configuration.The puzzle game is set up by means of a modified random forest classifier, designed to learn the local, topological label-structure and hence the local context associated to the training data.To solve the puzzle game we propose an iterative optimization technique that maximizes an agreement function by alternatingly seeking for the best label puzzle piece per pixel and the resulting semantic labelling per image.We provide both, theoretical properties of our puzzle solver algorithm as well as experimental results on the challenging MSRC and CamVid databases.In a direct comparison with a conditional random field we obtain superior results, indicating the practicability of our proposed method. Peter Kontschieder, Samuel Rota Bulò, Michael Donoser, Marcello Pelillo, Horst Bischof |
BMVC | 5 |
| 2011 | Discriminative Learning of Contour Fragments for Object DetectionabstractThe goal of this work is to discriminatively learn contour fragment descriptors for the task of object detection. Unlike previous methods that incorporate learning techniques only for object model generation or for verification after detection, we present a holistic object detection system using solely shape as underlying cue. In the learning phase, we interrelate local shape descriptions (fragments) of the object contour with the corresponding spatial location of the object centroid. We introduce a novel shape fragment descriptor that abstracts spatially connected edge points into a matrix consisting of angular relations between the points. Our proposed descriptor fulfills important properties like distinctiveness, robustness and insensitivity to clutter. During detection, we hypothesize object locations in a generalized Hough voting scheme. The back-projected votes from the fragments allow to approximately delineate the object contour. We evaluate our method e.g. on the well-known ETHZ shape data base, where we achieve an average detection score of 87:5% at 1:0 FPPI only from Hough voting, outperforming the highest scoring Hough voting approaches by almost 8%. Peter Kontschieder, Hayko Riemenschneider, Michael Donoser, Horst Bischof |
BMVC | 4 |
| 2011 | GPSlam: Marrying Sparse Geometric and Dense Probabilistic Visual MappingabstractWe propose a novel, hybrid SLAM system to construct a dense occupancy grid map based on sparse visual features and dense depth information. While previous approaches deemed the occupancy grid usable only in 2D mapping, and in combination with a probabilistic approach, we show that geometric SLAM can produce consistent, robust and dense occupancy information, and maintain it even during erroneous exploration and loop closure. We require only a single hypothesis of the occupancy map and employ a weighted inverse mapping scheme to align it to sparse geometric information. We propose a novel map-update criterion to prevent inconsistencies, and a robust measure to discriminate exploration from localization. Katrin Pirker, Matthias Rüther, Gerald Schweighofer, Horst Bischof |
BMVC | 4 |
| 2011 | On-line Hough ForestsabstractHough forests have emerged as a powerful and versatile method, which achieves state-of-the-art results on various computer vision applications, ranging from object detection over pose estimation to action recognition. The original method operates in offline mode, assuming to have access to the entire training set at once. This limits its applicability in domains where data arrives sequentially or when large amounts of data have to be exploited. In these cases, on-line approaches naturally would be beneficial. To this end, we propose an on-line extension of Hough forests, which is based on the principle of letting the trees evolve on-line while the data arrives sequentially, for both classification and regression. We further propose a modified version of off-line Hough forests, which only needs a small subset of the training data for optimization. In the experiments, we show that using these formulations, the classification results of classic Hough forests could be reached or even outperformed, while being orders of magnitudes faster. Furthermore, our method allows for tracking arbitrary objects without requiring any prior knowledge. We present state-of-the-art tracking results on publicly available data sets. © 2011. The copyright of this document resides with its authors. Samuel Schulter, Christian Leistner, Peter M. Roth, Horst Bischof, Luc Van Gool |
BMVC | 4 |
| 2011 | Skeletal Graph Based Human Pose Estimation in Real-TimeabstractWe propose a new method to quickly and robustly estimate the 3D pose of the human skeleton from volumetric body scans without the need for visual markers. The core principle of our algorithm is to apply a fast center-line extraction to 3D voxel data and robustly fit a skeleton model to the resulting graph. Our algorithm allows for automatic, single-frame initialization and tracking of the human pose while being fast enough for real-time applications at up to 30 frames per second. We provide an extensive qualitative and quantitative evaluation of our method on real and synthetic datasets which demonstrates the stability of our algorithm even when applied to long motion sequences. Matthias Straka, Stefan Hauswiesner, Matthias Rüther, Horst Bischof |
BMVC | 4 |
| 2011 | Supervised local subspace learning for continuous head pose estimationabstractHead pose estimation from images has recently attracted much attention in computer vision due to its diverse applications in face recognition, driver monitoring and human computer interaction. Most successful approaches to head pose estimation formulate the problem as a nonlinear regression between image features and continuous 3D angles (i.e. yaw, pitch and roll). However, regression-like methods suffer from three main drawbacks: (1) They typically lack generalization and overfit when trained using a few samples. (2) They fail to get reliable estimates over some regions of the output space (angles) when the training set is not uniformly sampled. For instance, if the training data contains under-sampled areas for some angles. (3) They are not robust to image noise or occlusion. To address these problems, this paper presents Supervised Local Subspace Learning (SL2), a method that learns a local linear model from a sparse and non-uniformly sampled training set. SL2learns a mixture of local tangent spaces that is robust to under-sampled regions, and due to its regularization properties it is also robust to over-fitting. Moreover, because SL2is a generative model, it can deal with image noise. Experimental results on the CMU Multi-PIE and BU-3DFE database show the effectiveness of our approach in terms of accuracy and computational complexity. Dong Huang 0007, Markus Storer, Fernando De la Torre, Horst Bischof |
CVPR | 4 |
| 2011 | Improving classifiers with unlabeled weakly-related videosabstractCurrent state-of-the-art object classification systems are trained using large amounts of hand-labeled images. In this paper, we present an approach that shows how to use unlabeled video sequences, comprising weakly-related object categories towards the target class, to learn better classifiers for tracking and detection. The underlying idea is to exploit the space-time consistency of moving objects to learn classifiers that are robust to local transformations. In particular, we use dense optical flow to find moving objects in videos in order to train part-based random forests that are insensitive to natural transformations. Our method, which is called Video Forests, can be used in two settings: first, labeled training data can be regularized to force the trained classifier to generalize better towards small local transformations. Second, as part of a tracking-by-detection approach, it can be used to train a general codebook solely on pair-wise data that can then be applied to tracking of instances of a priori unknown object categories. In the experimental part, we show on benchmark datasets for both tracking and detection that incorporating unlabeled videos into the learning of visual classifiers leads to improved results. Christian Leistner, Martin Godec, Samuel Schulter, Amir Saffari, Manuel Werlberger, Horst Bischof |
CVPR | 6 |
| 2011 | Hough-based tracking of non-rigid objectsabstractOnline learning has shown to be successful in tracking of previously unknown objects. However, most approaches are limited to a bounding-box representation with fixed aspect ratio. Thus, they provide a less accurate fore- ground/background separation and cannot handle highly non-rigid and articulated objects. This, in turn, increases the amount of noise introduced during online self-training. In this paper, we present a novel tracking-by-detection approach to overcome this limitation based on the generalized Hough-transform. We extend the idea of Hough Forests to the online domain and couple the voting- based detection and back-projection with a rough segmentation based on GrabCut. This significantly reduces the amount of noisy training samples during online learning and thus effectively prevents the tracker from drifting. In the experiments, we demonstrate that our method successfully tracks a variety of previously unknown objects even under heavy non-rigid transformations, partial occlusions, scale changes and rotations. Moreover, we compare our tracker to state-of-the-art methods (both bounding-box- based as well as part-based) and show robust and accurate tracking results on various challenging sequences. Martin Godec, Peter M. Roth, Horst Bischof |
ICCV | 3 |
| 2011 | Structured class-labels in random forests for semantic image labellingabstractIn this paper we propose a simple and effective way to integrate structural information in random forests for semantic image labelling. By structural information we refer to the inherently available, topological distribution of object classes in a given image. Different object class labels will not be randomly distributed over an image but usually form coherently labelled regions. In this work we provide a way to incorporate this topological information in the popular random forest framework for performing low-level, unary classification. Our paper has several contributions: First, we show how random forests can be augmented with structured label information. In the second part, we introduce a novel data splitting function that exploits the joint distributions observed in the structured label space for learning typical label transitions between object classes. Finally, we provide two possibilities for integrating the structured output predictions into concise, semantic labellings. In our experiments on the challenging MSRC and CamVid databases, we compare our method to standard random forest and conditional random field classification results. Peter Kontschieder, Samuel Rota Bulò, Horst Bischof, Marcello Pelillo |
ICCV | 3 |
| 2011 | Natural landmark-based monocular localization for MAVsabstractHighly accurate localization of a micro aerial vehicle (MAV) with respect to a scene is important for a wide range of applications, in particular surveillance and inspection. Most existing approaches to visual localization focus on indoor environments, while such tasks require outdoor navigation. Within this work, we introduce a novel algorithm for monocular visual localization for MAVs based on the concept of virtual views in 3D space. Under the assumption that significant parts of the scene do not alter their geometry and serve as natural landmarks, the accuracy of our visual approach outperforms consumer grade GPS systems. In an experimental setup we compare our approach to a state-of-the-art visual SLAM algorithm and evaluate the performance by geometric validation from an observer's view. As our method directly allows global registration, it is neither prone to drift nor bias. This makes it well suited for long-term autonomous navigation. Andreas Wendel, Arnold Irschara, Horst Bischof |
ICRA | 3 |
| 2011 | CD SLAM - Continuous localization and mapping in a dynamic worldabstractWhen performing large-scale perpetual localization and mapping one faces problems like memory consumption or repetitive and dynamic scene elements requiring robust data association. We propose a visual SLAM method which handles short- and long-term scene dynamics in large environments using a single camera only. Through visibility-dependent map filtering and efficient keyframe organization we reach a considerable performance gain only through incorporation of a slightly more complex map representation. Experiments on a large, mixed indoor/outdoor dataset over a time period of two weeks demonstrate the scalability and robustness of our approach. Katrin Pirker, Matthias Rüther, Horst Bischof |
IROS | 3 |
| 2011 | Robust planar target tracking and pose estimation from a single concavityabstractIn this paper we introduce a novel real-time method to track weakly textured planar objects and to simultaneously estimate their 3D pose. The basic idea is to adapt the classic tracking-by-detection approach, which seeks for the object to be tracked independently in each frame, for tracking non-textured objects. In order to robustly estimate the 3D pose of such objects in each frame, we have to tackle three demanding problems. First, we need to find a stable representation of the object which is discriminable against the background and highly repetitive. Second, we have to robustly relocate this representation in every frame, also during considerable viewpoint changes. Finally, we have to estimate the pose from a single, closed object contour. Of course, all demands shall be accommodated at low computational costs and in real-time. To attack the above mentioned problems, we propose to exploit the properties of Maximally Stable Extremal Regions (MSERs) for detecting the required contours in an efficient manner and to apply random ferns as efficient and robust classifier for tracking. To estimate the 3D pose, we construct a perspectively invariant frame on the closed contour which is intrinsically provided by the extracted MSER. In our experiments we obtain robust tracking results with accurate poses on various challenging image sequences at a single requirement: One MSER used for tracking has to have at least one concavity that sufficiently deviates from its convex hull. Michael Donoser, Peter Kontschieder, Horst Bischof |
ISMAR | 3 |
| 2011 | Neural Process Reconstruction from Sparse User Scribbles
Mike Roberts 0001, Won-Ki Jeong, Amelio Vázquez Reina, Markus Unger, Horst Bischof, Jeff Lichtman, Hanspeter Pfister |
MICCAI (1) | 5 |
| 2011 | Special issue on Optimization for vision, graphics and medical imaging: Theory and applications
Nikos Komodakis, Georg Langs, Horst Bischof, Nikos Paragios |
Comput. Vis. Image Underst. | 3 |
| 2010 | Temporal Feature Weighting for Prototype-Based Action Recognition
Thomas Mauthner, Peter M. Roth, Horst Bischof |
ACCV (2) | 3 |
| 2010 | Interactive Multi-label Segmentation
Jakob Santner, Thomas Pock, Horst Bischof |
ACCV (1) | 3 |
| 2010 | Audio-Visual Co-Training for Vehicle ClassificationabstractIn this paper, we introduce a fully autonomous vehicle classification system that continuously learns from largeamounts of unlabeled data. For that purpose, we proposea novel on-line co-training method based on visual and acoustic information. Our system does not need complicated microphone arrays or video calibration and automatically adapts to specific traffic scenes. These specialized detectors are more accurate and more compact than general classifiers, which allows for light-weight usage in low-cost and portable embedded systems. Hence, we implemented our system on an off-the-shelf embedded platform. In the experimental part, we show that the proposed method is able to cover the desired task and outperforms single-cue systems. Furthermore, our co-training framework minimizes the labeling effort without degrading the overall system performance. Martin Godec, Christian Leistner, Horst Bischof, Andreas Starzacher, Bernhard Rinner |
AVSS | 3 |
| 2010 | Automatic Detection and Reading of Dangerous Goods PlatesabstractIn this paper, we present an efficient solution for automatic detection and reading of dangerous goods plates on trucks and trains. According to the ADR agreement dangerous goods transports are marked with an orange plate covering the hazard class and the identification number for the hazardous substances. Since under real-world conditions high resolution images (often at low quality) have to be processed an efficient and robust system is required. In particular, we propose a multi-stage system consisting of an acquisition step, a saliency region detector (to reduce the run-time), a plate detector, and a robust recognition step based on an Optical Character Recognition (OCR). To demonstrate the system, we show qualitative and quantitative localization/recognition results on two challenging data sets. In fact, building on proven robust and efficient methods, we show excellent detection and classification results under hard environmental conditions at low run-time. Peter M. Roth, Martin Köstinger, Paul Wohlhart, Horst Bischof, Josef A. Birchbauer |
AVSS | 4 |
| 2010 | Learning of Scene-Specific Object Detectors by Classifier Co-GridsabstractRecently, classifier grids have shown to be a considerable alternative to sliding window approaches for object detection from static cameras. The main drawback of such methods is that they are biased by the initial model. In fact, the classifiers can be adapted to changing environmental conditions but due to conservative updates no new object-specific information is acquired. Thus, the goal of this work is to increase the recall of scene-specific classifiers while preserving their accuracy and speed. In particular, we introduce a co-training strategy for classifier grids using a robust on-line learner. Thus, the robustness is preserved while the recall can be increased. The co-training strategy robustly provides negative as well as positive updates. In addition, the number of negative updates can be drastically reduced, which additionally speeds up the system. In the experimental results these benefits are demonstrated on different publicly available surveillance benchmark data sets. Sabine Sternig, Peter M. Roth, Horst Bischof |
AVSS | 3 |
| 2010 | Linked edges as stable region boundariesabstractMany of the recently popular shape based category recognition methods require stable, connected and labeled edges as input. This paper introduces a novel method to find the most stable region boundaries in grayscale images for this purpose. In contrast to common edge detection algorithms as Canny, which only analyze local discontinuities in image brightness, our method integrates mid-level information by analyzing regions that support the local gradient magnitudes. We use a component tree where every node contains a single connected region obtained from thresholding the gradient magnitude image. Edges in the tree are defined by an inclusion relationship between nested regions in different levels of the tree. Region boundaries which are similar in shape (i. e. have a low chamfer distance) across several levels of the tree are included in the final result. Since the component tree can be calculated in quasi-linear time and chamfer matching between nodes in the component tree is reduced to analysis of the distance transformation, results are obtained in an efficient manner. The proposed detection algorithm labels all identified edges during calculation, thus avoiding the cumbersome post-processing of connecting and labeling edge responses. We evaluate our method on two reference data sets and demonstrate improved performance for shape prototype based localization of objects in images. Michael Donoser, Hayko Riemenschneider, Horst Bischof |
CVPR | 3 |
| 2010 | Variational segmentation of elongated volumetric structuresabstractWe present an interactive approach for segmenting thin volumetric structures. The proposed segmentation model is based on an anisotropic weighted Total Variation energy with a global volumetric constraint and is minimized using an efficient numerical approach and a convex relaxation. The algorithm is globally optimal w.r.t. the relaxed problem for any volumetric constraint. The binary solution of the relaxed problem equals the globally optimal solution of the original problem. Implemented on today's user-programmable graphics cards, it allows real-time user interaction. The method is applied to and evaluated on the task of articular cartilage segmentation of human knee joints and segmentation of tubular structures like liver vessels and airway trees. Christian Reinbacher, Thomas Pock, Christian Bauer 0001, Horst Bischof |
CVPR | 4 |
| 2010 | Online multi-class LPBoostabstractOnline boosting is one of the most successful online learning algorithms in computer vision. While many challenging online learning problems are inherently multi-class, online boosting and its variants are only able to solve binary tasks. In this paper, we present Online Multi-Class LPBoost (OMCLP) which is directly applicable to multi-class problems. From a theoretical point of view, our algorithm tries to maximize the multi-class soft-margin of the samples. In order to solve the LP problem in online settings, we perform an efficient variant of online convex programming, which is based on primal-dual gradient descent-ascent update strategies. We conduct an extensive set of experiments over machine learning benchmark datasets, as well as, on Caltech 101 category recognition dataset. We show that our method is able to outperform other online multi-class methods. We also apply our method to tracking where, we present an intuitive way to convert the binary tracking by detection problem to a multi-class problem where background patterns which are similar to the target class, become virtual classes. Applying our novel model, we outperform or achieve the state-of-the-art results on benchmark tracking videos. Amir Saffari, Martin Godec, Thomas Pock, Christian Leistner, Horst Bischof |
CVPR | 5 |
| 2010 | PROST: Parallel robust online simple trackingabstractTracking-by-detection is increasingly popular in order to tackle the visual tracking problem. Existing adaptive methods suffer from the drifting problem, since they rely on self-updates of an on-line learning method. In contrast to previous work that tackled this problem by employing semi-supervised or multiple-instance learning, we show that augmenting an on-line learning method with complementary tracking approaches can lead to more stable results. In particular, we use a simple template model as a non-adaptive and thus stable component, a novel optical-flow-based mean-shift tracker as highly adaptive element and an on-line random forest as moderately adaptive appearance-based learner. We combine these three trackers in a cascade. All of our components run on GPUs or similar multi-core systems, which allows for real-time performance. We show the superiority of our system over current state-of-the-art tracking methods in several experiments on publicly available data. Jakob Santner, Christian Leistner, Amir Saffari, Thomas Pock, Horst Bischof |
CVPR | 5 |
| 2010 | Motion estimation with non-local total variation regularizationabstractState-of-the-art motion estimation algorithms suffer from three major problems: Poorly textured regions, occlusions and small scale image structures. Based on the Gestalt principles of grouping we propose to incorporate a low level image segmentation process in order to tackle these problems. Our new motion estimation algorithm is based on non-local total variation regularization which allows us to integrate the low level image segmentation process in a unified variational framework. Numerical results on the Middlebury optical flow benchmark data set demonstrate that we can cope with the aforementioned problems. Manuel Werlberger, Thomas Pock, Horst Bischof |
CVPR | 3 |
| 2010 | On-line semi-supervised multiple-instance boostingabstractA recent dominating trend in tracking called tracking-by-detection uses on-line classifiers in order to redetect objects over succeeding frames. Although these methods usually deliver excellent results and run in real-time they also tend to drift in case of wrong updates during the self-learning process. Recent approaches tackled this problem by formulating tracking-by-detection as either one-shot semi-supervised learning or multiple instance learning. Semi-supervised learning allows for incorporating priors and is more robust in case of occlusions while multiple-instance learning resolves the uncertainties where to take positive updates during tracking. In this work, we propose an on-line semi-supervised learning algorithm which is able to combine both of these approaches into a coherent framework. This leads to more robust results than applying both approaches separately. Additionally, we introduce a combined loss that simultaneously uses labeled and unlabeled samples, which makes our tracker more adaptive compared to previous on-line semi-supervised methods. Experimentally, we demonstrate that by using our semi-supervised multiple-instance approach and utilizing robust learning methods, we are able to outperform state-of-the-art methods on various benchmark tracking videos. Bernhard Zeisl, Christian Leistner, Amir Saffari, Horst Bischof |
CVPR | 4 |
| 2010 | MIForests: Multiple-Instance Learning with Randomized Trees
Christian Leistner, Amir Saffari, Horst Bischof |
ECCV (6) | 3 |
| 2010 | Using Partial Edge Contour Matches for Efficient Object Category Localization
Hayko Riemenschneider, Michael Donoser, Horst Bischof |
ECCV (5) | 3 |
| 2010 | Robust Multi-View Boosting with Priors
Amir Saffari, Christian Leistner, Martin Godec, Horst Bischof |
ECCV (3) | 4 |
| 2010 | An omnidirectional Time-of-Flight camera and its application to indoor SLAMabstractPhotonic mixer devices (PMDs) are able to create reliable depth maps of indoor environments. Yet, their application in mobile robotics, especially in simultaneous localization and mapping (SLAM) applications, is hampered by the limited field of view. Enhancing the field of view by optical devices is not trivial, because the active light source and the sensor rays need to be redirected in a defined manner. In this work we propose an omnidirectional PMD sensor which is well suited for indoor SLAM and easy to calibrate. Using a single sensor and multiple planar mirrors, we are able to reliably navigate in indoor environments to create geometrically consistent maps, even on optically difficult surfaces. Katrin Pirker, Matthias Rüther, Horst Bischof, Gerald Schweighofer, Heinz Mayer |
ICARCV | 3 |
| 2010 | The narcissistic robot: Robot calibration using a mirrorabstractWe present a novel method for calibration of a robotic manipulator. The robot kinematic chain and its tool are observed by a hand mounted camera through a mirror. We show the possibility of enabling hand-eye, hand-tool, and kinematic robot calibration without incorporating accurate external references, except the mirror. Using this particularly simple setup, hand-eye calibration becomes independent of the kinematic chain and parameter observability constraints in kinematic calibration become more relaxed, which makes pose planning for robot calibration more convenient. Matthias Rüther, Martin Lenz, Horst Bischof |
ICARCV | 3 |
| 2010 | Object Tracking by Structure Tensor AnalysisabstractCovariance matrices have recently been a popular choice for versatile tasks like recognition and tracking due to their powerful properties as local descriptor and their low computational demands. This paper outlines similarities of covariance matrices to the well-known structure tensor. We show that the generalized version of the structure tensor is a powerful descriptor and that it can be calculated in constant time by exploiting the properties of integral images. To measure the similarities between several structure tensors, we describe an approximation scheme which allows comparison in a Euclidean space. Such an approach is also much more efficient than the common, computationally demanding Riemannian Manifold distances. Experimental evaluation proves the applicability for the task of object tracking demonstrating improved performance compared to covariance tracking. Michael Donoser, Stefan Kluckner, Horst Bischof |
ICPR | 3 |
| 2010 | Shape Prototype Signatures for Action RecognitionabstractRecognizing human actions in video sequences is frequently based on analyzing the shape of the human silhouette as the main feature. In this paper we introduce a method for recognizing different actions by comparing signatures of similarities to pre-defined shape prototypes. In training, we build a vocabulary of shape prototypes by clustering a training set of human silhouettes and calculate prototype similarity signatures for all training videos. During testing a prototype signature is calculated for the test video and is aligned to each training signature by dynamic time warping. A simple voting scheme over the similarities to the training videos provides action classification results and temporal alignments to the training videos. Experimental evaluation on a reference data set demonstrates that state-of-the-art results are achieved. Michael Donoser, Hayko Riemenschneider, Horst Bischof |
ICPR | 3 |
| 2010 | Shape Guided Maximally Stable Extremal Region (MSER) TrackingabstractMaximally Stable Extremal Regions (MSERs) are one of the most prominent interest region detectors in computer vision due to their powerful properties and low computational demands. In general MSERs are detected in single images, but given image sequences as input, the repeatability of MSER detection can be improved by exploiting correspondences between subsequent frames by feature based analysis. Such an approach fails during fast movements, in heavily cluttered scenes and in images containing several similar sized regions because of the simple feature based analysis. In this paper we propose an extension of MSER tracking by considering shape similarity as strong cue for defining the frame-to-frame correspondences. Efficient calculation of shape similarity scores ensures that real-time capability is maintained. Experimental evaluation demonstrates improved repeatability and an application for tracking weakly textured, planar objects. Michael Donoser, Hayko Riemenschneider, Horst Bischof |
ICPR | 3 |
| 2010 | On-Line Random Naive Bayes for TrackingabstractRandomized learning methods (i.e., Forests or Ferns) have shown excellent capabilities for various computer vision applications. However, it was shown that the tree structure in Forests can be replaced by even simpler structures, e.g., Random Naive Bayes classifiers, yielding similar performance. The goal of this paper is to benefit from these findings to develop an efficient on-line learner. Based on the principals of on-line Random Forests, we adapt the Random Naive Bayes classifier to the on-line domain. For that purpose, we propose to use on-line histograms as weak learners, which yield much better performance than simple decision stumps. Experimentally we show, that the approach is applicable to incremental learning on machine learning datasets. Additionally, we propose to use an IIR filtering-like forgetting function for the weak learners to enable adaptivity and evaluate our classifier on the task of tracking by detection. Martin Godec, Christian Leistner, Amir Saffari, Horst Bischof |
ICPR | 4 |
| 2010 | Detecting Paper Fibre Cross Sections in Microtomy ImagesabstractThe goal of this work is the fully-automated detection of cellulose fibre cross sections in microtomy images. A lack of significant appearance information makes edges the only reliable cue for detection. We present a novel and highly discriminative edge fragment descriptor that represents angular relations between fragment points. We train a Random Forest with a plurality of these descriptors including their respective center votes. In such a way, the Random Forest exploits the knowledge about the object centroid for detection using a generalized Hough voting scheme. In the experiments we found that our method is able to robustly detect fibre cross sections in microtomy images and can therefore serve as initialization for successive fibre segmentation or tracking algorithms. Peter Kontschieder, Michael Donoser, Johannes Kritzinger, Wolfgang Bauer, Horst Bischof |
ICPR | 5 |
| 2010 | Pose Estimation of Known Objects by Efficient Silhouette MatchingabstractPose estimation is essential for automated handling of objects. In many computer vision applications only the object silhouettes can be acquired reliably, because untextured or slightly transparent objects do not allow for other features. We propose a pose estimation method for known objects, based on hierarchical silhouette matching and unsupervised clustering. The search hierarchy is created by an unsupervised clustering scheme, which makes the method less sensitive to parametrization, and still exploits spatial neighborhood for efficient hierarchy generation. Our evaluation shows a decrease in matching time of 80% compared to an exhaustive matching and scalability to large models. Christian Reinbacher, Matthias Rüther, Horst Bischof |
ICPR | 3 |
| 2010 | Novel Multi View Structure Estimation Based on Barycentric CoordinatesabstractTraditionally, multi-view stereo algorithms estimate three-dimensional structure from corresponding points by linear triangulation or bundle-adjustment. This introduces systematic errors in case of inaccurate camera calibration and partial occlusion. The errors are not negligible in applications requiring high accuracy like micro-metrology or quality inspection. We show how accuracy of structure estimation can be significantly increased by using a barycentric coordinate representation for central perspective projection. Experiments show a reduction of geometric error by 50% compared with bundle adjustment. The error remains almost constantly low, even under partial occlusion. Matthias Rüther, Horst Bischof |
ICPR | 2 |
| 2010 | Inverse Multiple Instance Learning for Classifier GridsabstractRecently, classifier grids have shown to be a considerable alternative for object detection from static cameras. However, one drawback of such approaches is drifting if an object is not moving over a long period of time. Thus, the goal of this work is to increase the recall of such classifiers while preserving their accuracy and speed. In particular, this is realized by adapting ideas from Multiple Instance Learning within a boosting framework. Since the set of positive samples is well defined, we apply this concept to the negative samples extracted from the scene: Inverse Multiple Instance Learning. By introducing temporal bags, we can ensure that each bag contains at least one sample having a negative label, providing the required stability. The experimental results demonstrate that using the proposed approach state-of-the-art detection results can by obtained, however, showing superior classification results in presence of non-moving objects. Sabine Sternig, Peter M. Roth, Horst Bischof |
ICPR | 3 |
| 2010 | Intensity-Based Congealing for Unsupervised Joint Image AlignmentabstractWe present an approach for unsupervised alignment of an ensemble of images called congealing. Our algorithm is based on image registration using the mutual information measure as a cost function. The cost function is optimized by a standard gradient descent method in a multiresolution scheme. As opposed to other congealing methods, which use the SSD measure, the mutual information measure is better suited as a similarity measure for registering images since no prior assumptions on the relation of intensities between images are required. We present alignment results on the MNIST handwritten digit database and on facial images obtained from the CVL database. Markus Storer, Martin Urschler, Horst Bischof |
ICPR | 3 |
| 2010 | Generalized sparse MRF appearance models
Rene Donner, Georg Langs, Branislav Micusík, Horst Bischof |
Image Vis. Comput. | 4 |
| 2010 | Segmentation of interwoven 3d tubular tree structures utilizing shape priors and graph cuts
Christian Bauer 0001, Thomas Pock, Erich Sorantin, Horst Bischof, Reinhard Beichel |
Medical Image Anal. | 4 |
| 2010 | Localization and Trajectory Reconstruction in Surveillance Cameras with Nonoverlapping ViewsabstractThis paper proposes a method that localizes two surveillance cameras and simultaneously reconstructs object trajectories in 3D space. The method is an extension of the Direct Reference Plane method, which formulates the localization and the reconstruction as a system of linear equations that is globally solvable by Singular Value Decomposition. The method's assumptions are static synchronized cameras, smooth trajectories, known camera internal parameters, and the rotation between the cameras in a world coordinate system. The paper describes the method in the context of self-calibrating cameras, where the internal parameters and the rotation can be jointly obtained assuming a man-made scene with orthogonal structures. Experiments with synthetic and real--image data show that the method can recover the camera centers with an error less than half a meter even in the presence of a 4 meter gap between the fields of view. Roman P. Pflugfelder, Horst Bischof |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Modelling fingerprint ridge orientation using Legendre polynomials
Surinder Ram, Horst Bischof, Josef A. Birchbauer |
Pattern Recognit. | 2 |
| 2010 | Context information from search engines for document recognition
Michael Donoser, Silke Wagner, Horst Bischof |
Pattern Recognit. Lett. | 3 |
| 2010 | Global Solutions of Variational Models with Convex RegularizationabstractWe propose an algorithmic framework for computing global solutions of variational models with convex regularity terms that permit quite arbitrary data terms. While the minimization of variational problems with convex data and regularity terms is straightforward (using, for example, gradient descent), this is no longer trivial for functionals with nonconvex data terms. Using the theoretical framework of calibrations, the original variational problem can be written as the maximum flux of a particular vector field going through the boundary of the subgraph of the unknown function. Upon relaxation this formulation turns the problem into a convex problem, although in a higher dimension. In order to solve this problem, we propose a fast primal-dual algorithm which significantly outperforms existing algorithms. In experimental results we show the application of our method to outlier filtering of range images and disparity estimation in stereo images using a variety of convex regularity terms. Thomas Pock, Daniel Cremers, Horst Bischof, Antonin Chambolle |
SIAM J. Imaging Sci. | 3 |
| 2009 | Efficient Partial Shape Matching of Outer Contours
Michael Donoser, Hayko Riemenschneider, Horst Bischof |
ACCV (1) | 3 |
| 2009 | Semantic Classification in Aerial Imagery by Integrating Appearance and Height Information
Stefan Kluckner, Thomas Mauthner, Peter M. Roth, Horst Bischof |
ACCV (2) | 4 |
| 2009 | Beyond Pairwise Shape Similarity Analysis
Peter Kontschieder, Michael Donoser, Horst Bischof |
ACCV (3) | 3 |
| 2009 | Semantic Image Classification using Consistent Regions and Individual ContextabstractThis paper proposes an efficient approach for semantic image classification by integrating additional contextual constraints such as class co-occurrences into a randomized forest classification framework. The randomized forest classifier performs an initial yet local classification on the pixel level by using powerful covariance matrix based descriptors as feature representation. Furthermore, we exploit multiple unsupervised image partitions to provide a reliable spatial region support and to capture the real object boundaries. An information theoretic driven approach detects consistently classified regions and generates a representative segmentation incorporating the classification result on the pixel level. Moreover, we use a conditional random field formulation to obtain a final labeling including context information individually generated for each test image. To illustrate state-of-the-art performance, we run experiments on the two versions of the MSRC [21] dataset with 9 and 21 object classes and on the PASCAL VOC2007 [5] image collection. Stefan Kluckner, Thomas Mauthner, Peter M. Roth, Horst Bischof |
BMVC | 4 |
| 2009 | Bag of Optical Flow Volumes for Image Sequence RecognitionabstractThis paper introduces a novel 3D interest point detector and feature representation for describing image sequences. The approach considers image sequences as spatio-temporal volumes and detects Maximally Stable Volumes (MSVs) in efficiently calcu-lated optical flow fields. This provides a set of binary optical flow volumes highlighting the dominant motions in the sequences. 3D interest points are sampled on the surface of the volumes which balance well between density and informativeness. The binary opti-cal flow volumes are used as feature representation in a 3D shape context descriptor. A standard bag-of-words approach then allows building discriminant optical flow volume signatures for predicting class labels of previously unseen image sequences by machine learning algorithms. We evaluate the proposed method for the task of action recognition on the well-known Weizmann dataset, and show that we outperform recently proposed state-of-the-art 3D interest point detection and description methods. 1 Hayko Riemenschneider, Michael Donoser, Horst Bischof |
BMVC | 3 |
| 2009 | Interactive Texture Segmentation using Random Forests and Total VariationabstractCommon methods for interactive texture segmentation rely on probability maps based on low dimensional features such as e.g. intensity or color, that are usually modeled using basic learning algorithms such as histograms or Gaussian Mixture Models. The use of low level features allows for fast generation of these hypotheses but limits applicability to a small class of images. We address this problem by learning complex descriptors with Random Forests and exploiting their inherent parallelism in a GPU implementation. The segmentation itself is based on a convex energy functional that uses weighted Total Variation regularization and a point-wise data term allowing for continuous foreground/background membership hypotheses. Its globally optimal solution is obtained by a fast primal-dual algorithm providing a reasonable convergence criterion. As a result, we present a versatile interactive texture segmentation framework. We show experiments with natural, artificial and medical data and demonstrate superior results compared to two recent approaches. Jakob Santner, Markus Unger, Thomas Pock, Christian Leistner, Amir Saffari, Horst Bischof |
BMVC | 6 |
| 2009 | Anisotropic Huber-L1 Optical FlowabstractTV regularization is an L1 penalization of the flow gradient magnitudes, and due to the tendency of the L1 norm to favor sparse solutions (i.e. lots of ‘zeros’), the fill-in effect caused by the regularizer leads to piecewise constant solutions in weakly textured areas. This effect, known as ‘staircasing’ in a 1D setting, can be reduced significantly by using a quadratic penalization for small gradient magnitudes while sticking to linear penalization for larger magnitudes to maintain the discontinuity preserving properties known from TV. A comparison of isotropic TV and isotropic Huber regularity is shown in Fig. 1 by means of rendering the disparities u1 of the Dimetrodon dataset. The color coded flow (cf. Fig. 1(a)) is superimposed as texture. Based on the two observations that motion discontinuities often occur along object boundaries and that in turn object boundaries often coincide Manuel Werlberger, Werner Trobin, Thomas Pock, Andreas Wedel, Daniel Cremers, Horst Bischof |
BMVC | 6 |
| 2009 | Fast human detection in crowded scenes by contour integration and local shape estimationabstractThe complexity of human detection increases significantly with a growing density of humans populating a scene. This paper presents a Bayesian detection framework using shape and motion cues to obtain a maximum a posteriori (MAP) solution for human configurations consisting of many, possibly occluded pedestrians viewed by a stationary camera. The paper contains two novel contributions for the human detection task: 1. computationally efficient detection based on shape templates using contour integration by means of integral images which are built by oriented string scans; (2) a non-parametric approach using an approximated version of the shape context descriptor which generates informative object parts and infers the presence of humans despite occlusions. The outputs of the two detectors are used to generate a spatial configuration of hypothesized human body locations. The configuration is iteratively optimized while taking into account the depth ordering and occlusion status of the hypotheses. The method achieves fast computation times even in complex scenarios with a high density of people. Its validity is demonstrated on a substantial amount of image data using the CAVIAR and our own datasets. Evaluation results and comparison with state of the art are presented. Csaba Beleznai, Horst Bischof |
CVPR | 2 |
| 2009 | From structure-from-motion point clouds to fast location recognitionabstractEfficient view registration with respect to a given 3D reconstruction has many applications like inside-out tracking in indoor and outdoor environments, and geo-locating images from large photo collections. We present a fast location recognition technique based on structure from motion point clouds. Vocabulary tree-based indexing of features directly returns relevant fragments of 3D models instead of documents from the images database. Additionally, we propose a compressed 3D scene representation which improves recognition rates while simultaneously reducing the computation time and the memory consumption. The design of our method is based on algorithms that efficiently utilize modern graphics processing units to deliver real-time performance for view registration. We demonstrate the approach by matching hand-held outdoor videos to known 3D urban models, and by registering images from online photo collections to the corresponding landmarks. Arnold Irschara, Christopher Zach, Jan-Michael Frahm, Horst Bischof |
CVPR | 4 |
| 2009 | A convex relaxation approach for computing minimal partitionsabstractIn this work we propose a convex relaxation approach for computing minimal partitions. Our approach is based on rewriting the minimal partition problem (also known as Potts model) in terms of a primal dual Total Variation functional. We show that the Potts prior can be incorporated by means of convex constraints on the dual variables. For minimization we propose an efficient primal dual projected gradient algorithm which also allows a fast implementation on parallel hardware. Although our approach does not guarantee to find global minimizers of the Potts model we can give a tight bound on the energy between the computed solution and the true minimizer. Furthermore we show that our relaxation approach dominates recently proposed relaxations. As a consequence, our approach allows to compute solutions closer to the true minimizer. For many practical problems we even find the global minimizer. We demonstrate the excellent performance of our approach on several multi-label image segmentation and stereo problems. Thomas Pock, Antonin Chambolle, Daniel Cremers, Horst Bischof |
CVPR | 4 |
| 2009 | Classifier grids for robust adaptive object detectionabstractIn this paper we present an adaptive but robust object detector for static cameras by introducing classifier grids. Instead of using a sliding window for object detection we propose to train a separate classifier for each image location, obtaining a very specific object detector with a low false alarm rate. For each classifier corresponding to a grid element we estimate two generative representations in parallel, one describing the object's class and one describing the background. These are combined in order to obtain a discriminative model. To enable to adapt to changing environments these classifiers are learned on-line (i.e., boosting). Continuously learning (24 hours a day, 7 days a week) requires a stable system. In our method this is ensured by a fixed object representation while updating only the representation of the background. We demonstrate the stability in a long-term experiment by running the system for a whole week, which shows a stable performance over time. In addition, we compare the proposed approach to state-of-the-art methods in the field of person and car detection. In both cases we obtain competitive results. Peter M. Roth, Sabine Sternig, Helmut Grabner, Horst Bischof |
CVPR | 4 |
| 2009 | Regularized multi-class semi-supervised boostingabstractMany semi-supervised learning algorithms only deal with binary classification. Their extension to the multi-class problem is usually obtained by repeatedly solving a set of binary problems. Additionally, many of these methods do not scale very well with respect to a large number of unlabeled samples, which limits their applications to large-scale problems with many classes and unlabeled samples. In this paper, we directly address the multi-class semi-supervised learning problem by an efficient boosting method. In particular, we introduce a new multi-class margin-maximizing loss function for the unlabeled data and use the generalized expectation regularization for incorporating cluster priors into the model. Our approach enables efficient usage of very large data sets. The performance and efficiency of our method is demonstrated on both standard machine learning data sets as well as on challenging object categorization tasks. Amir Saffari, Christian Leistner, Horst Bischof |
CVPR | 3 |
| 2009 | Saliency driven total variation segmentationabstractThis paper introduces an unsupervised color segmentation method. The underlying idea is to segment the input image several times, each time focussing on a different salient part of the image and to subsequently merge all obtained results into one composite segmentation. We identify salient parts of the image by applying affinity propagation clustering to efficiently calculated local color and texture models. Each salient region then serves as an independent initialization for a figure/ground segmentation. Segmentation is done by minimizing a convex energy functional based on weighted total variation leading to a global optimal solution. Each salient region provides an accurate figure/ ground segmentation highlighting different parts of the image. These highly redundant results are combined into one composite segmentation by analyzing local segmentation certainty. Our formulation is quite general, and other salient region detection algorithms in combination with any semi-supervised figure/ground segmentation approach can be used. We demonstrate the high quality of our method on the well-known Berkeley segmentation database. Furthermore we show that our method can be used to provide good spatial support for recognition frameworks. Michael Donoser, Martin Urschler, Martin Hirzer, Horst Bischof |
ICCV | 4 |
| 2009 | Semi-Supervised Random ForestsabstractRandom Forests (RFs) have become commonplace in many computer vision applications. Their popularity is mainly driven by their high computational efficiency during both training and evaluation while still being able to achieve state-of-the-art accuracy. This work extends the usage of Random Forests to Semi-Supervised Learning (SSL) problems. We show that traditional decision trees are optimizing multi-class margin maximizing loss functions. From this intuition, we develop a novel multi-class margin definition for the unlabeled data, and an iterative deterministic annealing-style training algorithm maximizing both the multi-class margin of labeled and unlabeled samples. In particular, this allows us to use the predicted labels of the unlabeled data as additional optimization variables. Furthermore, we propose a control mechanism based on the out-of-bag error, which prevents the algorithm from degradation if the unlabeled data is not useful for the task. Our experiments demonstrate state-of-the-art semi-supervised learning performance in typical machine learning problems and constant improvement using unlabeled data for the Caltech-101 object categorization task. Christian Leistner, Amir Saffari, Jakob Santner, Horst Bischof |
ICCV | 4 |
| 2009 | An algorithm for minimizing the Mumford-Shah functionalabstractIn this work we revisit the Mumford-Shah functional, one of the most studied variational approaches to image segmentation. The contribution of this paper is to propose an algorithm which allows to minimize a convex relaxation of the Mumford-Shah functional obtained by functional lifting. The algorithm is an efficient primal-dual projection algorithm for which we prove convergence. In contrast to existing algorithms for minimizing the full Mumford-Shah this is the first one which is based on a convex relaxation. As a consequence the computed solutions are independent of the initialization. Experimental results confirm that the proposed algorithm determines smooth approximations while preserving discontinuities of the underlying signal. Thomas Pock, Daniel Cremers, Horst Bischof, Antonin Chambolle |
ICCV | 3 |
| 2009 | Structure- and motion-adaptive regularization for high accuracy optic flowabstractThe accurate estimation of motion in image sequences is of central importance to numerous computer vision applications. Most competitive algorithms compute flow fields by minimizing an energy made of a data and a regularity term. To date, the best performing methods rely on rather simple purely geometric regularizes favoring smooth motion. In this paper, we revisit regularization and show that appropriate adaptive regularization substantially improves the accuracy of estimated motion fields. In particular, we systematically evaluate regularizes which adoptively favor rigid body motion (if supported by the image data) and motion field discontinuities that coincide with discontinuities of the image structure. The proposed algorithm relies on sequential convex optimization, is real-time capable and outperforms all previously published algorithms by more than one average rank on the Middlebury optic flow benchmark. Andreas Wedel, Daniel Cremers, Thomas Pock, Horst Bischof |
ICCV | 4 |
| 2009 | Multiple target detection and tracking with guaranteed framerates on mobile phonesabstractIn this paper we present a novel method for real-time pose estimation and tracking on low-end devices such as mobile phones. The presented system can track multiple known targets in real-time and simultaneously detect new targets for tracking. We present a method to automatically and dynamically balance the quality of detection and tracking to adapt to a variable time budget and ensure a constant frame rate. Results from real data of a mobile phone Augmented Reality system demonstrate the efficiency and robustness of the described approach. The system can track 6 planar targets on a mobile phone simultaneously at framerates of 23 fps. Daniel Wagner 0003, Dieter Schmalstieg, Horst Bischof |
ISMAR | 3 |
| 2009 | Weakly Supervised Group-Wise Model Learning Based on Discrete Optimization
Rene Donner, Horst Wildenauer, Horst Bischof, Georg Langs |
MICCAI (1) | 3 |
| 2009 | Editorial Special Issue ECCV 2006
Horst Bischof, Ales Leonardis |
Int. J. Comput. Vis. | 1 |
| 2009 | A level set framework using a new incremental, robust Active Shape Model for object segmentation and tracking
Michael Fussenegger, Peter M. Roth, Horst Bischof, Rachid Deriche, Axel Pinz |
Image Vis. Comput. | 3 |
| 2009 | Comparison and Evaluation of Methods for Liver Segmentation From CT DatasetsabstractThis paper presents a comparison study between 10 automatic and six interactive methods for liver segmentation from contrast-enhanced CT images. It is based on results from the "MICCAI 2007 Grand Challenge" workshop, where 16 teams evaluated their algorithms on a common database. A collection of 20 clinical images with reference segmentations was provided to train and tune algorithms in advance. Participants were also allowed to use additional proprietary training data for that purpose. All teams then had to apply their methods to 10 test datasets and submit the obtained results. Employed algorithms include statistical shape models, atlas registration, level-sets, graph-cuts and rule-based systems. All results were compared to reference segmentations five error measures that highlight different aspects of segmentation accuracy. All measures were combined according to a specific scoring system relating the obtained values to human expert variability. In general, interactive methods reached higher average scores than automatic approaches and featured a better consistency of segmentation quality. However, the best automatic methods (mainly based on statistical shape models with some additional free deformation) could compete well on the majority of test images. The study provides an insight in performance of different segmentation approaches under real-world conditions and highlights achievements and limitations of current image analysis techniques. Tobias Heimann, Bram van Ginneken, Martin Styner, Yulia Arzhaeva, Volker Aurich, Christian Bauer 0001, Andreas Beck 0001, Christoph Becker 0002, Reinhard Beichel, György Bekes, Fernando Bello, Gerd Karl Binnig, Horst Bischof, Alexander Bornik, Peter Cashman, Ying Chi, Andrés Cordova, Benoit M. Dawant, Márta Fidrich, Jacob D. Furst, Daisuke Furukawa, Lars Grenacher, Joachim Hornegger, Dagmar Kainmüller, Richard Kitney, Hidefumi Kobatake, Hans Lamecker, Thomas Lange, Brian Lennon, Rui Li 0012, Senhu Li, Hans-Peter Meinzer, Gábor Németh, Daniela Raicu, Anne-Mareike Rau, Eva M. van Rikxoort, Mikaël Rousson, László Ruskó, Kinda Anna Saddi, Günter Schmidt 0001, Dieter Seghers, Akinobu Shimizu, Pieter Slagmolen, Erich Sorantin, Grzegorz Soza, Ruchaneewan Susomboon, Jonathan M. Waite, Andreas Wimmer, Ivo Wolf |
IEEE Trans. Medical Imaging | 13 |
| 2009 | Automatic Quantification of Joint Space Narrowing and Erosions in Rheumatoid ArthritisabstractRheumatoid arthritis (RA) is a chronic disease that affects and potentially destroys the joints of the appendicular skeleton. The precise and reproducible quantification of the progression of joint space narrowing and the erosive bone destructions caused by RA is crucial during treatment and in imaging biomarkers in clinical trials. Current manual scoring methods exhibit high interreader variability, even after intensive training, and thus, impede the efficient monitoring of the disease. We propose a fully automatic quantitative assessment of the radiographic changes that result from RA, to increase the accuracy, reproducibility, and speed of image interpretation. Initial joint location estimates are obtained by local linear mappings based on texture features. Bone contours are delineated by active shape models comprised of statistical models of bone shape and local texture. These models are refined by snakes which increase the accuracy and allow for a fitting of pathological deviations from the training population. The method then measures joint space widths and detects erosions on the bone contour. Joint space widths are measured with a coefficient of variation of 2%-7% for repeated measurements and erosion detection exhibits an area under the receiver operating characteristic (ROC) curve of 0.89. Model landmarks serve as a reference system along the contour. These landmarks enable the definition of joint regions and more specific follow-up monitoring. The automatic quantification allows for a remote analysis, relevant for multicenter clinical trials, and reduces the workload of clinical experts since parts of the process can be managed by nonexpert personnel. Georg Langs, Philipp Peloschek, Horst Bischof, Franz Kainberger |
IEEE Trans. Medical Imaging | 3 |
| 2008 | Fast Non-Rigid Object Boundary TrackingabstractThis paper introduces a method which provides robust tracking results and accurately segmented object boundaries in short computation time. The first step of the algorithm is to apply a novel edge detector on efficiently calculated color probability maps in an object-specific Fisher color space. The proposed edge detector exploits context information by finding the maximally stable boundaries of connected regions in threshold results outperforming purely local edge detectors. Finally, based on the estimated edge maps a probabilistic particle filtering framework hypothesizes rigid transformations for initializing an active contour model to provide accurate object segmentations in each frame. Experimental evaluations show that robust tracking results with accurate segmentations are obtained on challenging data sets. 1 Michael Donoser, Horst Bischof |
BMVC | 2 |
| 2008 | TVSeg - Interactive Total Variation Based Image SegmentationabstractInteractive object extraction is an important part in any image editing software. We present a two step segmentation algorithm that first obtains a binary segmentation and then applies matting on the border regions to obtain a smooth alpha channel. The proposed segmentation algorithm is based on the minimization of the Geodesic Active Contour energy. A fast Total Variation minimization algorithm is used to find the globally optimal solution. We show how user interaction can be incorporated and outline an efficient way to exploit color information. A novel matting approach, based on energy minimization, is presented. Experimental evaluations are discussed, and the algorithm is compared to state of the art object extraction algorithms. The GPU based binaries are available online. Markus Unger, Thomas Pock, Werner Trobin, Daniel Cremers, Horst Bischof |
BMVC | 5 |
| 2008 | Semi-supervised boosting using visual similarity learningabstractThe required amount of labeled training data for object detection and classification is a major drawback of current methods. Combining labeled and unlabeled data via semi-supervised learning holds the promise to ease the tedious and time consuming labeling effort. This paper presents a novel semi-supervised learning method which combines the power of learned similarity functions and classifiers. The approach capable of exploiting both labeled and unlabeled data is formulated in a boosting framework. One classifier (the learned similarity) serves as a prior which is steadily improved via training a second classifier on labeled and unlabeled samples. We demonstrate the approach on challenging computer vision applications. First, we show how we can train a classifier using only a few labeled samples and many unlabeled data. Second, we improve (specialize) a state-of-the-art detector by using labeled and unlabeled data. Christian Leistner, Helmut Grabner, Horst Bischof |
CVPR | 3 |
| 2008 | What can missing correspondences tell us about 3D structure and motion?abstractPractically all existing approaches to structure and motion computation use only positive image correspondences to verify the camera pose hypotheses. Incorrect epipolar geometries are solely detected by identifying outliers among the found correspondences. Ambiguous patterns in the images are often incorrectly handled by these standard methods. In this work we propose two approaches to overcome such problems. First, we apply non-monotone reasoning on view triplets using a Bayesian formulation. In contrast to two-view epipolar geometry, image triplets allow the prediction of features in the third image. Absence of these features (i.e. missing correspondences) enables additional inference about the view triplet. Furthermore, we integrate these view triplet handling into an incremental procedure for structure and motion computation. Thus, our approach is able to refine the maintained 3D structure when additional image data is provided. Christopher Zach, Arnold Irschara, Horst Bischof |
CVPR | 3 |
| 2008 | Semi-supervised On-Line Boosting for Robust Tracking
Helmut Grabner, Christian Leistner, Horst Bischof |
ECCV (1) | 3 |
| 2008 | A Convex Formulation of Continuous Multi-label Problems
Thomas Pock, Thomas Schoenemann, Gottfried Munda, Horst Bischof, Daniel Cremers |
ECCV (3) | 4 |
| 2008 | SERBoost: Semi-supervised Boosting with Expectation Regularization
Amir Saffari, Helmut Grabner, Horst Bischof |
ECCV (3) | 3 |
| 2008 | Continuous Energy Minimization Via Repeated Binary Fusion
Werner Trobin, Thomas Pock, Daniel Cremers, Horst Bischof |
ECCV (4) | 4 |
| 2008 | Fusion of Feature- and Area-Based Information for Urban Buildings Modeling from Aerial Imagery
Lukas Zebedin, Joachim Bauer, Konrad F. Karner, Horst Bischof |
ECCV (4) | 4 |
| 2008 | An active boosting-based learning framework for real-time hand detectionabstractHuman hand detection problem has important applications in sign language and human machine interfaces. In this work, we present a novel approach for learning a vision-based hand detection system. The main contribution is a robust on-line boosting-based framework for real-time detection of a hand in unconstrained environments. The use of efficient representative features allows fast computation while dealing with vast changing of hand appearances and background. Interactive on-line training allows efficiently train and improve the detector. Moreover, we propose a strategy to efficiently improve the performance meanwhile reduce hand labeling effort. Besides, if necessary, we use a verification process to prevent “drifting” of classifier over time. The proposed method is practically favorable as it meets the requirements of real-time performance, accuracy and robustness. It works well with reasonable amount of training samples and is computational efficient. Experiments for detection of hands in challenging data sets show the outperform of our approach. Thi Thuy Nguyen, Nguyen Dang Binh, Horst Bischof |
FG | 3 |
| 2008 | Classifier fusion for robust ICAO compliant face analysisabstractBiometrics is a huge and very fast growing domain of methods for uniquely recognizing humans based on one or more intrinsic physical or behavioral traits with applications in many different areas, e.g., surveillance, person verification and identification. The International Civil Aviation Organization (ICAO) provides a number of specifications to prepare automated recognition from travel document photos. The goal of these specifications is to increase security in civil aviation on the basis of standardized biometric data. Due to this international standard, there is a high demand for automatically checking face images to assist civil service employees in decision-making. In this work, we present a face normalization and analysis system implementing several parts of the ICAO specification. Our key contribution of this analysis is the fusion of different established classifiers to boost performance of the overall system. Our results show the superior checking quality on facial images due to utilizing classifier fusion compared to a single classifier decision. Markus Storer, Martin Urschler, Horst Bischof, Josef A. Birchbauer |
FG | 3 |
| 2008 | Using covariance matrices for unsupervised texture segmentationabstractIn this paper we propose an efficient unsupervised texture segmentation method. We introduce a texture extension of a state-of-the-art color segmentation algorithm. We show how to use covariance matrices of low level features for texture description. These features are efficiently calculated using integral images. Furthermore, a multi-scale extension allows to provide accurate texture segmentation results. An experimental evaluation on a synthetic texture database and images of the Berkeley image database demonstrate the improved performance of the algorithm. Michael Donoser, Horst Bischof |
ICPR | 2 |
| 2008 | Real time appearance based hand trackingabstractThis paper introduces a real-time method for tracking hands through image sequences. Our method combines efficiently calculated color likelihood maps with a modified version of the maximally stable extremal region (MSER)-tracker. The proposed algorithm allows to robustly track hands through image sequences and additionally provides accurate hand segmentations per frame. Experimental evaluation proves the high accuracy of the segmentation results and a first application for human-computer interaction is presented. Michael Donoser, Horst Bischof |
ICPR | 2 |
| 2008 | Using web search engines to improve text recognitionabstractIn this paper we introduce a framework for automated text recognition from images. We first describe a simple but efficient text detection and recognition method based on analysis of maximally stable extremal regions (MSERs) and simple template matching which allows to provide initial character recognition results. The main emphasis of the paper is on introducing a novel method for exploiting contextual information to improve the obtained recognition results. We propose to analyze the results of Web search engine queries on two levels of detail, which both allow to significantly improve the overall text recognition performance. The experimental evaluations on reference data sets prove that even based on a low quality single character recognition method the proposed Web search engine extension enables reasonable text recognition results. Michael Donoser, Horst Bischof, Silke Wagner |
ICPR | 2 |
| 2008 | A probabilistic approach for tracking fibersabstractThis paper describes a combination of an automated image acquisition method and a probabilistic tracking method for analysis of the 3D microstructure of a sheet of paper. A prototype which combines microtomy and light microscopy enables efficient and fully automated digitization of paper samples in high resolution. A particle filter based tracking method then allows to segment individual fibers from the obtained 3D data sets. The capability of accessing the properties of individual fibers enables analysis of e. g. 3D fiber mass distribution, 3D fiber orientation or fiber morphology. Experiments show that the method provides results consistent with the knowledge of paper experts. Michael Donoser, Thomas Mauthner, Horst Bischof, Johannes Kritzinger |
ICPR | 3 |
| 2008 | Training sequential on-line boosting classifier for visual trackingabstractOn-line boosting allows to adapt a trained classifier to changing environmental conditions or to use sequentially available training data. Yet, two important problems in the on-line boosting training remain unsolved: (i) classifier evaluation speed optimization and, (ii) automatic classifier complexity estimation. In this paper we show how the on-line boosting can be combined with Waldpsilas sequential decision theory to solve both of the problems. The properties of the proposed on-line WaldBoost algorithm are demonstrated on a visual tracking problem. The complexity of the classifier is changing dynamically depending on the difficulty of the problem. On average, a speedup of a factor of 5-10 is achieved compared to the non-sequential on-line boosting. Helmut Grabner, Jan Sochman, Horst Bischof, Jiri Matas |
ICPR | 3 |
| 2008 | Robust tracking of spatial related componentsabstractThis paper introduces a hierarchical approach for multi-component tracking, where the object-to-be-tracked is modeled as a group of spatial related parts. We propose to use a robust particle filtering framework for tracking the individual components and outline how the spatial coherency between the parts can be efficiently integrated by analyzing a two-level hierarchy of particle filters. Including spatial information allows to handle common tracking problems like occlusions, clutter or blur. Furthermore, the dynamic calculation of particle set uncertainties allows a dynamic adaption of stiffness values for the spatial model to e. g. force occluded parts to stay in spatial relation. The experimental section proves the robustness of the proposed tracker on challenging sequences of the VIVID-PETS database. Thomas Mauthner, Michael Donoser, Horst Bischof |
ICPR | 3 |
| 2008 | Tracking across non-overlapping views via geometryabstractTracking across non-overlapping camera views is still an unsolved problem. Appearance is a popular cue that does not work when the views are considerably different. This paper proposes tracking in the 3-D space without using appearance. We propose to use the geometry between the two cameras. Linear inhomogeneous triangulation is expanded by a Gaussian random walk model. Tracking is then triangulation followed by a re-projection given the assumption that two persons cannot occupy the same space at the same time. A first experiment with a single person shows the success of this new tracking approach even with a 2m gap between the fields of view of the cameras. Roman P. Pflugfelder, Horst Bischof |
ICPR | 2 |
| 2008 | Online object recognition by MSER trajectoriesabstractThis work presents a robust online learning and recognition system. The basic idea is to exploit information from tracking an object during the recognition and/or learning stage to obtain increased robustness and better recognition results. Object tracking by means of an extended MSER tracker is utilized to detect local features and construct their trajectories. Compact object representations are formed by summarizing the trajectories. All steps are performed online including the MSER detection, tracking, summarization, SIFT description as well as learning and recognition based on a vocabulary tree. The proposed method is evaluated on realistic video sequences which prove the increased performance for robust online recognition. Hayko Riemenschneider, Michael Donoser, Horst Bischof |
ICPR | 3 |
| 2007 | Detecting, Tracking and Recognizing License Plates
Michael Donoser, Clemens Arth, Horst Bischof |
ACCV (2) | 3 |
| 2007 | Flea, Do You Remember Me?
Michael Grabner, Helmut Grabner, Joachim Pehserl, Petra Korica-Pehserl, Horst Bischof |
ACCV (1) | 5 |
| 2007 | People tracking across two distant self-calibrated camerasabstractPeople tracking is of fundamental importance in multi-camera surveillance systems. In recent years, many approaches for multi-camera tracking have been discussed. Most methods use either various image features or the geometric relation between the cameras or both as a cue. It is a desire to know the geometry for distant cameras, because geometry is not influenced by, for example, drastic changes in object appearance or in scene illumination. However, the determination of the camera geometry is cumbersome. The paper tries to solve this problem and contributes in two different ways. On the one hand, an approach is presented that calibrates two distant cameras automatically. We continue previous work and focus especially on the calibration of the extrinsic parameters. Point correspondences are used for this task which are acquired by detecting points on top of people's heads. On the other hand, qualitative experimental results with the PETS 2006 benchmark data show that the self-calibration is accurate enough for a solely geometric tracking of people across distant cameras. Reliable features for a matching are hardly available in such cases. Roman P. Pflugfelder, Horst Bischof |
AVSS | 2 |
| 2007 | Sparse MRF Appearance Models for Fast Anatomical Structure LocalisationabstractImage segmentation methods like active shape models, active appearance models or snakes require an initialisation that guarantees a considerable overlap with the object to be segmented. In this paper we present an approach that localises anatomical structures in a global manner by means of Markov Random Fields (MRF). It does not need initialisation, but finds the most plausible match of the query structure in the image. It provides for precise, reliable and fast detection of the structure and can serve as initialisation for more detailed segmentation steps. Sparse MRF Appearance Models (SAMs) encode a priori information about the geometric configurations of interest points, local features at these points and local features along the edges of adjacent points. This information is used to formulate a Markov Random Field and the mapping of the modeled object (e.g. a sequence of vertebrae) to the query image interest points is performed by the MAX-SUM algorithm. The local image information is captured by novel symmetry-based interest points and local descriptors derived from Gradient Vector Flow. Experimental results are reported for two data-sets showing the applicability to complex medical data. 1 Rene Donner, Branislav Micusík, Georg Langs, Horst Bischof |
BMVC | 4 |
| 2007 | Incremental LDA Learning by Combining Reconstructive and Discriminative ApproachesabstractIncremental subspace methods have proven to enable efficient training if large amounts of training data have to be processed or if not all data is available in advance. In this paper we focus on incremental LDA learning which provides good classification results while it assures a compact data representation. In contrast to existing incremental LDA methods we additionally consider reconstructive information when incrementally building the LDA subspace. Hence, we get a more flexible representation that is capable to adapt to new data. Moreover, this allows to add new instances to existing classes as well as to add new classes. The experimental results show that the proposed approach outperforms other incremental LDA methods even approaching classification results obtained by batch learning. 1 Martina Uray, Danijel Skocaj, Peter M. Roth, Horst Bischof, Ales Leonardis |
BMVC | 4 |
| 2007 | Binary Co-occurrences of Weak DescriptorsabstractThis paper demonstrates that a reliable and efficient object recognition system based only on binary joint occurrences of quantized descriptors can be built. Specifically, we show that a high recognition performance can be obtained even with very weak (non discriminative) descriptors. The binary joint occurrence representation despite being high dimensional is very sparse and therefore efficient. In order to obtain reliable joint occurrences we present a fast hierarchical quantization algorithm. We illustrate our results using different descriptors (PCA-SIFT, Spin images, SIFT) on a challenging, specific object recognition task and consider the scaling behavior of the method. 1 Horst Bischof |
BMVC | 2 |
| 2007 | Real-Time License Plate Recognition on an Embedded DSP-PlatformabstractIn this paper we present a full-featured license plate detection and recognition system. The system is implemented on an embedded DSP platform and processes a video stream in real-time. It consists of a detection and a character recognition module. The detector is based on the AdaBoost approach presented by Viola and Jones. Detected license plates are segmented into individual characters by using a region-based approach. Character classification is performed with support vector classification. In order to speed up the detection process on the embedded device, a Kalman tracker is integrated into the system. The search area of the detector is limited to locations where the next location of a license plate is predicted. Furthermore, classification results of subsequent frames are combined to improve the class accuracy. The major advantages of our system are its real-time capability and that it does not require any additional sensor input (e.g. from infrared sensors) except a video stream. We evaluate our system on a large number of vehicles and license plates using bad quality video and show that the low resolution can be partly compensated by combining classification results of subsequent frames. Clemens Arth, Florian Limberger, Horst Bischof |
CVPR | 3 |
| 2007 | Robust Local Features and their Application in Self-Calibration and Object Recognition on Embedded SystemsabstractIn recent years many powerful computer vision algorithms have been invented, making automatic or semiautomatic solutions to many popular vision tasks, such as visual object recognition or camera calibration, possible. On the other hand embedded vision platforms and solutions such as smart cameras have successfully emerged, however, only offering limited computational and memory resources. The first contribution of this paper is the investigation of a set of robust local feature detectors and descriptors for application on embedded systems. We briefly describe the methods involved, i.e. the DoG (difference of Gaussian) and MSER (maximally stable extremal regions) detector as well as the PCA-SIFT descriptor, and discuss their suitability for smart systems and their qualification for given tasks. The second contribution of this work is the experimental evaluation of these methods on two challenging tasks, namely fully embedded object recognition on a moderate size database and on the task of robust camera calibration. Our approach is fortified by encouraging results we present at length. Clemens Arth, Christian Leistner, Horst Bischof |
CVPR | 3 |
| 2007 | ROI-SEG: Unsupervised Color Segmentation by Combining Differently Focused Sub ResultsabstractThis paper presents a novel unsupervised color segmentation scheme named ROI-SEG, which is based on the main idea of combining a set of different sub-segmentation results. We propose an efficient algorithm to compute sub-segmentations by an integral image approach for calculating Bhattacharyya distances and a modified version of the maximally stable extremal region (MSER) detector. The sub-segmentation algorithm gets a region-of-interest (ROI) as input and detects connected regions having similar color appearance as the ROI. We further introduce a method to identify ROIs representing the predominant color and texture regions of an image. Passing each of the identified ROIs to the sub-segmentation algorithm provides a set of different segmentations, which are then combined by analyzing a local quality criterion. The entire approach is fully unsupervised and does not need a priori information about the image scene. The method is compared to state-of-the-art algorithms on the Berkeley image database, where it shows competitive results at reduced computational costs. Michael Donoser, Horst Bischof |
CVPR | 2 |
| 2007 | Learning Features for TrackingabstractWe treat tracking as a matching problem of detected key-points between successive frames. The novelty of this paper is to learn classifier-based keypoint descriptions allowing to incorporate background information. Contrary to existing approaches, we are able to start tracking of the object from scratch requiring no off-line training phase before tracking. The tracker is initialized by a region of interest in the first frame. Afterwards an on-line boosting technique is used for learning descriptions of detected keypoints lying within the region of interest. New frames provide new samples for updating the classifiers which increases their stability. A simple mechanism incorporates temporal information for selecting stable features. In order to ensure correct updates a verification step based on estimating homographies using RANSAC is performed. The approach can be used for real-time applications since on-line updating and evaluating classifiers can be done efficiently. Michael Grabner, Helmut Grabner, Horst Bischof |
CVPR | 3 |
| 2007 | Eigenboosting: Combining Discriminative and Generative InformationabstractA major shortcoming of discriminative recognition and detection methods is their noise sensitivity, both during training and recognition. This may lead to very sensitive and brittle recognition systems focusing on irrelevant information. This paper proposes a method that selects generative and discriminative features. In particular, we boost classical Haar-like features and use the same features to approximate a generative model (i.e., eigenimages). A modified error function for boosting ensures that only features are selected that show a good discrimination and reconstruction. This allows a robust feature selection using boosting. Thus, we can handle problems where discriminant classifiers fail while still retaining the discriminative power. Our experiments show that we can significantly improve the recognition performance when learning from noisy data. Moreover, the feature type used allows efficient recognition and reconstruction. Helmut Grabner, Peter M. Roth, Horst Bischof |
CVPR | 3 |
| 2007 | Mumford-Shah Meets Stereo: Integration of Weak Depth HypothesesabstractRecent results on stereo indicate that an accurate segmentation is crucial for obtaining faithful depth maps. Variational methods have successfully been applied to both image segmentation and computational stereo. In this paper we propose a combination in a unified framework. In particular, we use a Mumford-Shah-like functional to compute a piecewise smooth depth map of a stereo pair. Our approach has two novel features: First, the regularization term of the functional combines edge information obtained from the color segmentation with flow-driven depth discontinuities emerging during the optimization procedure. Second, we propose a robust data term which adoptively selects the best matches obtained from different weak stereo algorithms. We integrate these features in a theoretically consistent framework. The final depth map is the minimizer of the energy functional, which can be solved by the associated functional derivatives. The underlying numerical scheme allows an efficient implementation on modern graphics hardware. We illustrate the performance of our algorithm using the Middlebury database as well as on real imagery. Thomas Pock, Christopher Zach, Horst Bischof |
CVPR | 3 |
| 2007 | Recognizing cars in aerial imagery to improve orthophotosabstractThe automatic creation of 3D models of urban spaces has become a very active field of research. This has been inspired by recent applications in the location-awareness on the Internet, as demonstrated in maps.live.com and similar websites. The level of automation in creating 3D city models has increased considerably, and has benefited from an increase in the redundancy of the source imagery, namely digital aerial photography. In this paper we argue that the next big step forward is to replace photographic texture by an interpretation of what the texture describes, and to achieve this fully automatically. One calls the result "semantic knowledge". For example we want to know that a certain part of the image is a car, a person, a building, a tree, a shrub, a window, a door, instead of just a collection of 3D points or triangles with a superimposed photographic texture. We investigate object recognition methods to make this next big step. We demonstrate an early result of using the on-line variant of a Boosting algorithm to indeed detect cars in aerial digital imagery to a satisfactory and useful level of completeness. And we show that we can use this semantic knowledge to produce improved orthophotos. We expect that also the 3D models will be improved by the knowledge of cars. Franz Leberl, Horst Bischof, Helmut Grabner, Stefan Kluckner |
GIS | 2 |
| 2007 | Towards Wiki-based Dense City ModelingabstractThis work reports on the advances and on the current status of a terrestrial city modeling approach, which uses images contributed by end-users as input. Hence, the Wiki principle well known from textual knowledge databases is transferred to the goal of incrementally building a virtual representation of the occupied habitat. In order to achieve this objective, many state-of-the-art computer vision methods must be applied and modified according to this task. We describe the utilized 3D vision methods and show initial results obtained from the current image database acquired by in-house participants. Arnold Irschara, Christopher Zach, Horst Bischof |
ICCV | 3 |
| 2007 | A 3D Teacher for Car Detection in Aerial ImagesabstractThis paper demonstrates how to reduce the hand labeling effort considerably by 3D information in an object detection task. In particular, we demonstrate how an efficient car detector for aerial images with minimal hand labeling effort can be build. We use an on-line boosting algorithm to incrementally improve the detection results. Initially, we train the classifier with a single positive (car) example, randomly drawn from a fixed number of given samples. When applying this detector to an image we obtain many false positive detections. We use information from a stereo matcher to detect some of these false positives (e.g. detected cars on a facade) and feed back this information to the classifier as negative updates. This improves the detector considerably, thus reducing the number of false positives. We show that we obtain similar results to hand labeling by iteratively applying this strategy. The performance of our algorithm is demonstrated on digital aerial images of urban environments. Stefan Kluckner, Georg Pacher, Helmut Grabner, Horst Bischof, Joachim Bauer |
ICCV | 4 |
| 2007 | A Globally Optimal Algorithm for Robust TV-L1 Range Image IntegrationabstractRobust integration of range images is an important task for building high-quality 3D models. Since range images, and in particular range maps from stereo vision, may have a substantial amount of outliers, any integration approach aiming at high-quality models needs an increased level of robustness. Additionally, a certain level of regularization is required to obtain smooth surfaces. Computational efficiency and global convergence are further preferable properties. The contribution of this paper is a unified framework to solve all these issues. Our method is based on minimizing an energy functional consisting of a total variation (TV) regularization force and an L1 data fidelity term. We present a novel and efficient numerical scheme, which combines the duality principle for the TV term with a point-wise optimization step. We demonstrate the superior performance of our algorithm on the well-known Middlebury multi-view database and additionally on real-world multi-view images. Christopher Zach, Thomas Pock, Horst Bischof |
ICCV | 3 |
| 2007 | Dual-Layer Visual Vocabulary Tree Hypotheses for Object RecognitionabstractThis paper introduces an efficient method to substantially increase the recognition performance of a vocabulary tree based recognition system. We propose to enhance the hypothesis obtained by a standard inverse object voting algorithm with reliable descriptor co-occurrences. The algorithm operates on different layers of a standard k-means tree benefiting from the advantages of different levels of information abstraction. The visual vocabulary tree shows good results when a large number of distinctive descriptors form a large visual vocabulary. Co-occurrences perform well even on a coarse object representation with a small number of visual words. An arbitration strategy with minimal computational effort combines the specific strengths of the particular representations. We demonstrate the achieved performance boost and robustness to occlusions in a challenging object recognition task. Sandra Ober, Clemens Arth, Horst Bischof |
ICIP (6) | 4 |
| 2007 | Object Localization Based on Markov Random Fields and Symmetry Interest Points
Rene Donner, Branislav Micusík, Georg Langs, Lech Szumilas, Philipp Peloschek, Klaus Friedrich 0004, Horst Bischof |
MICCAI (2) | 7 |
| 2007 | Robust Autonomous Model Learning from 2D and 3D Data Sets
Georg Langs, Rene Donner, Philipp Peloschek, Horst Bischof |
MICCAI (1) | 4 |
| 2007 | A Duality Based Algorithm for TV- L 1-Optical-Flow Image Registration
Thomas Pock, Martin Urschler, Christopher Zach, Reinhard Beichel, Horst Bischof |
MICCAI (2) | 5 |
| 2007 | Algorithmic Differentiation: Application to Variational Problems in Computer VisionabstractMany vision problems can be formulated as minimization of appropriate energy functionals. These energy functionals are usually minimized, based on the calculus of variations (Euler-Lagrange equation). Once the Euler-Lagrange equation has been determined, it needs to be discretized in order to implement it on a digital computer. This is not a trivial task and, is moreover, error-prone. In this paper, we propose a flexible alternative. We discretize the energy functional and, subsequently, apply the mathematical concept of algorithmic differentiation to directly derive algorithms that implement the energy functional's derivatives. This approach has several advantages: First, the computed derivatives are exact with respect to the implementation of the energy functional. Second, it is basically straightforward to compute second-order derivatives and, thus, the Hessian matrix of the energy functional. Third, algorithmic differentiation is a process which can be automated. We demonstrate this novel approach on three representative vision problems (namely, denoising, segmentation, and stereo) and show that state-of-the-art results are obtained with little effort. Thomas Pock, Michael Pock, Horst Bischof |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Multiple appearance models
Georg Langs, Philipp Peloschek, Rene Donner, Horst Bischof |
Pattern Recognit. | 4 |
| 2007 | Weighted and robust learning of subspace representations
Danijel Skocaj, Ales Leonardis, Horst Bischof |
Pattern Recognit. | 3 |
| 2006 | Fast Approximated SIFT
Michael Grabner, Helmut Grabner, Horst Bischof |
ACCV (1) | 3 |
| 2006 | Real-Time Tracking via On-line BoostingabstractVery recently tracking was approached using classification techniques such as support vector machines. The object to be tracked is discriminated by a classifier from the background. In a similar spirit we propose a novel on-line AdaBoost feature selection algorithm for tracking. The distinct advantage of our method is its capability of on-line training. This allows to adapt the classifier while tracking the object. Therefore appearance changes of the object (e.g. out of plane rotations, illumination changes) are handled quite naturally. Moreover, depending on the background the algorithm selects the most discriminating features for tracking resulting in stable tracking results. By using fast computable features (e.g. Haar-like wavelets, orientation histograms, local binary patterns) the algorithm runs in real-time. We demonstrate the performance of the algorithm on several (publically available) video sequences. 1 Helmut Grabner, Michael Grabner, Horst Bischof |
BMVC | 3 |
| 2006 | Efficient Maximally Stable Extremal Region (MSER) TrackingabstractThis paper introduces a tracking method for the well known local MSER (Maximally Stable Extremal Region) detector. The component tree is used as an efficient data structure, which allows the calculation of MSERs in quasi-linear time. It is demonstrated that the tree is able to manage the required data for tracking. We show that by means of MSER tracking the computational time for the detection of single MSERs can be improved by a factor of 4 to 10. Using a weighted feature vector for data association improves the tracking stability. Furthermore, the component tree enables backward tracking which further improves the robustness. The novel MSER tracking algorithm is evaluated on a variety of scenes. In addition, we demonstrate three different applications, tracking of license plates, faces and fibers in paper, showing in all three scenarios improved speed and stability. Michael Donoser, Horst Bischof |
CVPR (1) | 2 |
| 2006 | On-line Boosting and VisionabstractBoosting has become very popular in computer vision, showing impressive performance in detection and recognition tasks. Mainly off-line training methods have been used, which implies that all training data has to be a priori given; training and usage of the classifier are separate steps. Training the classifier on-line and incrementally as new data becomes available has several advantages and opens new areas of application for boosting in computer vision. In this paper we propose a novel on-line AdaBoost feature selection method. In conjunction with efficient feature extraction methods the method is real time capable. We demonstrate the multifariousness of the method on such diverse tasks as learning complex background models, visual tracking and object detection. All approaches benefit significantly by the on-line training. Helmut Grabner, Horst Bischof |
CVPR (1) | 2 |
| 2006 | Color Blob Segmentation by MSER AnalysisabstractThis paper presents an efficient color blob segmentation concept, which combines an ordering relationship based on analyzing Bhattacharyya distances with a modified version of the maximally stable extremal region (MSER) detector. After definition of the region-of-interest by a one-time user input, connected regions are detected within the input image with low computational effort. Single image and video sequence analysis results are presented, which prove the applicability of the concept. Additionally, the possible extension to 3D segmentation is shown by an application that analyzes the 3D microstructure of a sheet of paper. Michael Donoser, Horst Bischof, Mario Wiltsche |
ICIP | 2 |
| 2006 | Automatic Point Landmark Matching for Regularizing Nonlinear Intensity Registration: Application to Thoracic CT Images
Martin Urschler, Christopher Zach, Hendrik Ditt, Horst Bischof |
MICCAI (2) | 4 |
| 2006 | Piecewise planar scene reconstruction from sparse correspondences
Friedrich Fraundorfer, Konrad Schindler, Horst Bischof |
Image Vis. Comput. | 3 |
| 2006 | Fast Active Appearance Model Search Using Canonical Correlation AnalysisabstractA fast AAM search algorithm based on canonical correlation analysis (CCA-AAM) is introduced. It efficiently models the dependency between texture residuals and model parameters during search. Experiments show that CCA-AAMs, while requiring similar implementation effort, consistently outperform standard search with regard to convergence speed by a factor of four. Rene Donner, Michael Reiter, Georg Langs, Philipp Peloschek, Horst Bischof |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2005 | A Clique of Active Appearance Models by Minimum Description LengthabstractAutonomous model building is a crucial trend in model based methods like AAMs. This paper introduces an approach that deals with non-linearities by detecting distinct sub-parts in the data. Sub-models each representing an individual sub-part are derived from a minimum description length criterion. Thereby the resulting clique of models is more compact and obtains a better generalization behavior than a single model. The proposed AAM clique generation deals with non-linearities in the data in a generic information theoretic manner reducing the necessity of user interaction during training. Georg Langs, Philipp Peloschek, Rene Donner, Horst Bischof |
BMVC | 4 |
| 2005 | Optimal Sub-Shape Models by Minimum Description LengthabstractActive shape models are powerful and widely used tool to interpret complex image data. By building models of shape variation they enable search algorithms to use a priori knowledge in an efficient and gainful way. However, due to the linearity of PCA, non-linearities like rotations or independently moving sub-parts in the data can deteriorate the resulting model considerably. Although non-linear extensions of active shape models have been proposed and application specific solutions have been used, they still need a certain amount of user interaction during model building. In this paper the task of building/choosing optimal models is tackled in a more generic information theoretic fashion. In particular, we propose an algorithm based on the minimum description length principle to find an optimal subdivision of the data into sub-parts, each adequate for linear modeling. This results in an overall more compact model configuration. Which in turn leads to a better model in terms of modes of variations. The proposed method is evaluated on synthetic data, medical images and hand contours. Georg Langs, Philipp Peloschek, Horst Bischof |
CVPR (2) | 3 |
| 2005 | Efficient representation of in-plane rotation within a PCA framework
Martin Sengel, Horst Bischof |
Image Vis. Comput. | 2 |
| 2005 | Robust active appearance models and their application to medical image analysisabstractActive appearance models (AAMs) have been successfully used for a variety of segmentation tasks in medical image analysis. However, gross disturbances of objects can occur in routine clinical setting caused by pathological changes or medical interventions. This poses a problem for AAM-based segmentation, since the method is inherently not robust. In this paper, a novel robust AAM (RAAM) matching algorithm is presented. Compared to previous approaches, no assumptions are made regarding the kind of gray-value disturbance and/or the expected magnitude of residuals during matching. The method consists of two main stages. First, initial residuals are analyzed by means of a mean-shift-based mode detection step. Second, an objective function is utilized for the selection of a mode combination not representing the gross outliers. We demonstrate the robustness of the method in a variety of examples with different noise conditions. The RAAM performance is quantitatively demonstrated in two substantially different applications, diaphragm segmentation and rheumatoid arthritis assessment. In all cases, the robust method shows an excellent behavior, with the new method tolerating up to 50% object area covered by gross gray-level disturbances. Reinhard Beichel, Horst Bischof, Franz Leberl, Milan Sonka |
IEEE Trans. Medical Imaging | 2 |
| 2004 | Rapid Object Recognition from Discriminative Regions of Interest
Gerald Fritz, Christin Seifert, Lucas Paletta, Horst Bischof |
AAAI | 4 |
| 2004 | Camera Calibration from a Single Night Sky Image
Andreas Klaus, Joachim Bauer, Konrad F. Karner, Pierre Elbischger, Roland Perko, Horst Bischof |
CVPR (1) | 6 |
| 2004 | Learning to Focus Attention on Discriminative Regions for Object Detection
Gerald Fritz, Christin Seifert, Lucas Paletta, Horst Bischof |
ECAI | 4 |
| 2004 | Human detection in groups using a fast mean shift procedureabstractDetecting individual humans within groups becomes a non-trivial task when performing automatic visual surveillance in crowded scenes. This paper proposes a novel way to detect individual humans directly from the difference image using a fast variant of the mean shift mode seeking procedure and verifying the hypothesized configuration by a model-based approach. The method runs in real-time. Promising result are demonstrated for challenging image sequences. Csaba Beleznai, Bernhard Frühstück, Horst Bischof |
ICIP | 3 |
| 2004 | Illumination Insensitive Robot Self-Localization Using Panoramic Eigenspaces
Gerald Steinbauer-Wagner, Horst Bischof |
RoboCup | 2 |
| 2004 | Illumination insensitive recognition using eigenspaces
Horst Bischof, Horst Wildenauer, Ales Leonardis |
Comput. Vis. Image Underst. | 1 |
| 2004 | Automatic analysis of collagen fiber orientation in the outermost layer of human arteries
Pierre Elbischger, Horst Bischof, Peter Regitnig, Gerhard A. Holzapfel |
Pattern Anal. Appl. | 2 |
| 2003 | Shape-based detection of humans for video surveillance applicationsabstractIn this paper we describe a surveillance system that is not only able to detect blobs and track them but also determines if a blob is a person. The given blob is segmented into sub-regions. A person model is fit to these regions such that a likelihood measure is maximized. The likelihood measure depends on the number of identified body parts, their length, location, and aspect ratio. The method is translation, rotation, and scale invariant and computationally efficient. The results obtained for test video sequences are very encouraging. Herbert Ramoser, Csaba Beleznai, Thomas Schlögl, Horst Bischof |
ICIP (3) | 5 |
| 2003 | Robust DNA microarray image analysis
Norbert Brändle, Horst Bischof, Hilmar Lapp |
Mach. Vis. Appl. | 2 |
| 2003 | Kernel and subspace methods for computer vision
Ales Leonardis, Horst Bischof |
Pattern Recognit. | 2 |
| 2003 | Appearance models based on kernel canonical correlation analysis
Thomas Melzer, Michael Reiter, Horst Bischof |
Pattern Recognit. | 3 |
| 2002 | A Robust PCA Algorithm for Building Representations from Panoramic Images
Danijel Skocaj, Horst Bischof, Ales Leonardis |
ECCV (4) | 2 |
| 2002 | Fast object recognition and pose determinationabstractAddresses the problem of fast object recognition and pose determination of segmented objects. It combines the well-studied parametric eigenspace method with statistical moments of image signatures resulting in a computationally and memory efficient algorithm. The approach is suited for time or memory critical applications, e.g. in embedded systems. A variety of experiments on a set of 1620 images compare the recognition and pose estimation performance to the standard eigenspace technique. The results show that despite the reduced memory and speed requirements the recognition rate is identical to the standard method; only under heavy noise conditions is the pose estimation accuracy slightly lower. Martin Sengel, Vassili Kravtchenko-Berejnoi, Horst Bischof |
ICIP (3) | 4 |
| 2002 | Multiple eigenspaces
Ales Leonardis, Horst Bischof, Jasna Maver |
Pattern Recognit. | 2 |
| 2001 | Nonlinear Feature Extraction Using Generalized Canonical Correlation Analysis
Thomas Melzer, Michael Reiter, Horst Bischof |
ICANN | 3 |
| 2001 | Illumination Insensitive Eigenspaces
Horst Bischof, Horst Wildenauer, Ales Leonardis |
ICCV | 1 |
| 2001 | Memory efficient fingerprint verificationabstractFingerprint recognition and verification are often based on local fingerprint features, usually ridge endings or terminations, also called minutiae. By exploiting the structural uniqueness of the image region around a minutia, the fingerprint recognition performance can be significantly enhanced. However, for most fingerprint images the number of minutia image regions (MIRs) becomes dramatically large, which imposes - especially for embedded systems - an enormous memory requirement. Therefore, we are investigating different algorithms for compression of minutia regions. The requirement for these algorithms is to achieve a high compression rate (about 20) with minimum loss in the matching performance of minutia image region matching. We investigate the matching performance for compression algorithms based on the principal component and the wavelet transformation. The matching results are presented in form of normalized ROC curves and interpreted in terms of compression rates and the MIR dimension. Csaba Beleznai, Herbert Ramoser, B. Wachmann, Josef A. Birchbauer, Horst Bischof, Walter G. Kropatsch |
ICIP (2) | 5 |
| 2001 | View-based object representations using RBF networks
Horst Bischof, Ales Leonardis |
Image Vis. Comput. | 1 |
| 2000 | Robust Spot Fitting for Genetic Spot Array ImagesabstractAddresses the problem of reliably fitting parametric and semi-parametric models to high density spot array images obtained in gene expression experiments. The goal is to measure the amount of genetic material at specific spot locations. Many spots can be modelled accurately by a Gaussian shape. In order to deal with highly overlapping spots the authors use robust M-estimators. When the parametric method fails, they use a novel, robust semi-parametric method which can handle spots of different shapes accurately. They present the results for real data and compare the complexity of the two methods. Horng-Yang Chen, Norbert Brändle, Horst Bischof, Hilmar Lapp |
ICIP | 3 |
| 2000 | Multiple Eigenspaces by MDLabstractWe propose an approach to constructing multiple eigenspaces from a set of training images based on the minimum description length (MDL) principle. The main idea is to systematically build a redundant set of eigenspaces, which are treated as hypotheses that are then subject to a selection procedure. The selection procedure, based on the MDL principle, selects the final resulting set of eigenspaces as an optimal representation of the training set. We have tested the proposed method on a number of standard image sets, and the significance of the approach with respect to the recognition rate has been clearly demonstrated. Ales Leonardis, Horst Bischof |
ICPR | 2 |
| 2000 | Fuzzy C-Means in an MDL-FrameworkabstractIn this paper we present a minimum description length (MDL) framework for fuzzy clustering algorithms. This framework enables us to find an optimal number of cluster centers. We applied our approach to the fuzzy c-means algorithm for which we designed a computationally efficient procedure. We report the results of our approach on a 2D clustering problem and on RGB color image segmentation. Alexander Selb, Horst Bischof, Ales Leonardis |
ICPR | 2 |
| 2000 | Content Based Image Retrieval Using Interest Points and Texture FeaturesabstractContent based image retrieval is the task of searching images from a database, which are visually similar to a given example image. We present methods for content based image retrieval based on texture similarity using interest points and Gabor features. Interest point detectors are used in computer vision to detect image points with special properties, which can be geometric (corners) or non-geometric (contrast etc.). Gabor functions and Gabor filters are regarded as excellent tools for feature extraction and texture segmentation. The article combines these methods and generates a textural description of images. Special emphasis is devoted to distance measures on texture descriptions. Experimental results of a query system are given. Christian Wolf 0001, Walter G. Kropatsch, Horst Bischof, Jean-Michel Jolion |
ICPR | 3 |
| 2000 | Robust Parametric and Semi-Parametric Spot Fitting for Spot Array Images
Norbert Brändle, Horng-Yang Chen, Horst Bischof, Hilmar Lapp |
ISMB | 3 |
| 2000 | Recognizing Objects by Their Appearance Using Eigenimages
Horst Bischof, Ales Leonardis |
SOFSEM | 1 |
| 2000 | Robust Recognition Using Eigenimages
Ales Leonardis, Horst Bischof |
Comput. Vis. Image Underst. | 2 |
| 1999 | Automatic Grid Fitting for Genetic Spot Array Images Containing Guide Spots
Norbert Brändle, Hilmar Lapp, Horst Bischof |
CAIP | 3 |
| 1999 | MDL Principle for Robust Vector Quantisation
Horst Bischof, Ales Leonardis, Alexander Selb |
Pattern Anal. Appl. | 1 |
| 1998 | Robust Recognition of Scaled Eigenimages through a Hierarchical ApproachabstractRecently, we have proposed a new approach to estimation of the coefficients of eigenimages, which is robust against occlusion, varying background, and other types of non-Gaussian noise. In this paper we show that our method for estimating the coefficients can be applied to convolved and subsampled images yielding the same value of the coefficients. This enables an efficient multiresolution approach, where the values of the coefficients can directly be propagated through the scales. This property is used to extend our robust method to the problem of scaled images. We performed extensive experimental evaluations to confirm our theoretical results. Horst Bischof, Ales Leonardis |
CVPR | 1 |
| 1998 | MDL-based design of vector quantizersabstractWe develop a framework for vector quantization networks based on the minimum description length (MDL) principle. This MDL framework is used to derive conditions for the removal of superfluous units from the network. We design a computationally efficient algorithm for finding the optimal number of reference vectors as well as their positions. We illustrate our approach on 2D clustering problems and present applications on image coding. Horst Bischof, Ales Leonardis |
ICPR | 1 |
| 1998 | A robust subspace classifierabstractIn this paper we study the problem of missing features and the issues of robustness of subspace classification methods. We propose a new robust method for subspace classification which can cope with missing features and/or outliers. The main idea of our method is to use a robust projection of the patterns onto a subspace. We demonstrate our approach on cervicomotography data and compare our results to the results obtained by using various decision tree algorithms. Horst Bischof, Ales Leonardis, Florian Pezzei |
ICPR | 1 |
| 1998 | An efficient MDL-based construction of RBF networks
Ales Leonardis, Horst Bischof |
Neural Networks | 2 |
| 1998 | Finding optimal neural networks for land use classificationabstractThe authors present a fully automatic and computationally efficient algorithm based on the minimum description length principle (MDL) for optimizing multilayer perceptron (MLP) classifiers. They demonstrate their method on the problem of multispectral Landsat image classification. They compare their results with a hand-designed MLP and a Gaussian maximum likelihood classifier, in which their method produces better classification accuracy with a smaller number of hidden units. Horst Bischof, Ales Leonardis |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 1997 | Computational Complexity Reduction in Eigenspace Approaches
Ales Leonardis, Horst Bischof |
CAIP | 2 |
| 1997 | Adaptive combination of PCA and VQ networksabstractIn this paper we consider the principal component analysis (PCA) and vector quantization (VQ) neural networks for image compression. We present a method where the PCA and VQ steps are adaptively combined. A learning algorithm for this combined network is derived. We demonstrate that this approach can improve the results of the successive application of the individually optimal methods. Andreas Weingessel, Horst Bischof, Kurt Hornik, Friedrich Leisch |
IEEE Trans. Neural Networks | 2 |
| 1996 | Dealing with occlusions in the eigenspace approachabstractThe basic limitations of the current appearance-based matching methods using eigenimages are non-robust estimation of coefficients and inability to cope with problems related to occlusions and segmentation. In this paper we present a new approach which successfully solves these problems. The major novelty of our approach lies in the way how the coefficients of the eigenimages are determined. Instead of computing the coefficients by a projection of the data onto the eigenimages, we extract them by a hypothesize-and-test paradigm using subsets of image points. Competing hypotheses are then subject to a selection procedure based on the Minimum Description Length principle. The approach enables us not only to reject outliers and to deal with occlusions but also to simultaneously use multiple classes of eigenimages. Ales Leonardis, Horst Bischof |
CVPR | 2 |
| 1996 | Complexity optimization of adaptive RBF networksabstractWe propose an extension of RBF networks which includes a mechanism for optimizing the complexity of the network. The approach involves two procedures: adaptation (training) and selection. The first procedure adaptively changes the locations and the width of the centers of the basis functions and trains the linear weights. The selection procedure performs the elimination of some of the basis functions using an objective function. By iteratively combining these two procedures we achieve a controlled way of training and modifying RBF networks, which balances accuracy, learning time, and complexity of the resulting network. Tamás Leonardis, Horst Bischof |
ICPR | 2 |
| 1996 | Hierarchies of autoassociatorsabstractThe principal component pyramid is a hierarchical neural network which can successfully be employed in image compression and feature extraction of images. Previously, the construction of the network from the corresponding pyramid was done on a case by case basis. In this paper we automate this process by giving formulas describing the size of the network and the number of weight constraints in the net. Andreas Weingessel, Horst Bischof, Kurt Hornik |
ICPR | 2 |
| 1996 | Voronoi Pyramids Controlled by Hopfield Neural Networks
Etienne Bertin 0002, Horst Bischof, Pascal Bertolino |
Comput. Vis. Image Underst. | 2 |
| 1994 | Voronoi pyramids and Hopfield networksabstractPresents an algorithm for image segmentation with irregular pyramids. Instead of starting with the original pixel grid, the authors first apply an adaptive Voronoi tessellation to the image. For irregular pyramid construction the authors present a Hopfield neural network which controls the decimation process. The validity of the authors' approach is demonstrated by several examples in image segmentation. Horst Bischof, Etienne Bertin 0002, Pascal Bertolino |
ICPR (3) | 1 |
| 1994 | Fuzzy curve pyramidabstractThis paper describes an extension of the binary curve pyramid to curves with strengths. In particular we propose fuzzy relations to represent curve strength. We show how fuzzy relations can be processed in a pyramidal framework. The advantage gained by this method is that the properties of the binary curve pyramid are preserved, and that we gain some additional properties, which can be used in the pyramid construction phase. Horst Bischof, Walter G. Kropatsch |
ICPR (1) | 1 |
| 1994 | PCA-Pyramids for Image CompressionabstractThis paper presents a new method for image compression by neural networks. First, we show that we can use neural networks in a py(cid:173) ramidal framework, yielding the so-called PCA pyramids. Then we present an image compression method based on the PCA pyramid, which is similar to the Laplace pyramid and wavelet transform. Some experimental results with real images are reported. Finally, we present a method to combine the quantization step with the learning of the PCA pyramid. Horst Bischof, Kurt Hornik |
NIPS | 1 |
| 1992 | Neural Network "Surgery": Transplantation of Hidden Units
Axel Pinz, Horst Bischof |
ECAI | 2 |
| 1992 | Visualization methods for neural networksabstractThe interpretation of neural network behavior is of particular interest in neural network research. Visualization methods provide the necessary means to simultaneously analyze the huge amount of information hidden in the network. The authors propose a framework for visualization methods suited for feed forward neural networks. The basic idea is to use the spatial information available outside the network to arrange the data to be visualized (weights, activations of units) in the spatial domain of the display. Several examples which illustrate the proposed framework are presented.> Horst Bischof, Axel Pinz, Walter G. Kropatsch |
ICPR (2) | 1 |
| 1992 | Combining pyramidal and fractal image codingabstractCompared with the richness in detail of images generated by iterated function systems, the memory requirement of them is very low. Displaying these images is very time consuming, and currently there is no method published to derive an iterated function system directly from a given image. This is known as the inverse problem. The authors introduce discrete transformations and direct computation of a discrete attractor with deterministic noniterative algorithms. This results in an essential saving of time. Furthermore they present a solution of the inverse problem in the one dimensional discrete space.> Walter G. Kropatsch, Michael A. Neuhauser, Irene J. Leitgeb, Horst Bischof |
ICPR (3) | 4 |
| 1992 | Multispectral classification of Landsat-images using neural networksabstractThe authors report the application of three-layer back-propagation networks for classification of Landsat TM data on a pixel-by-pixel basis. The results are compared to Gaussian maximum likelihood classification. First, it is shown that the neural network is able to perform better than the maximum likelihood classifier. Secondly, in an extension of the basic network architecture it is shown that textural information can be integrated into the neural network classifier without the explicit definition of a texture measure. Finally, the use of neural networks for postclassification smoothing is examined.> Horst Bischof, Werner Schneider, Axel Pinz |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 1990 | Constructing a neural network for the interpretation of the species of trees in aerial photographsabstractA neural network with a three-layer feedforward architecture was used to interpret the species of trees in aerial photographs. Weight-visualization (WV) diagrams were developed to interpret the weights and the behavior of the hidden units easily. Several networks that were trained on different parameter settings were combined to construct a better-performing network using the WV diagrams as a tool.> Axel Pinz, Horst Bischof |
ICPR (1) | 2 |