VLDB 2026 Research / reviewers in the wild / expert
Martin Humenberger
dblp:53/2870
· DBLP profile ↗
21ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0003-0600-9164ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Self-Supervised Learning of Neural Implicit Feature Fields for Camera Pose RefinementabstractVisual localization techniques rely upon some underlying scene representation to localize against. These representations can be explicit such as 3D SFM map or implicit, such as a neural network that learns to encode the scene. The former requires sparse feature extractors and matchers to build the scene representation. The latter might lack geometric grounding not capturing the 3D structure of the scene well enough. This paper proposes to jointly learn the scene representation along with a 3D dense feature field and a 2D feature extractor whose outputs are embedded in the same metric space. Through a contrastive framework we align this volumetric field with the image-based extractor and regularize the latter with a ranking loss from learned surface information. We learn the underlying geometry of the scene with an implicit field through volumetric rendering and design our feature field to leverage intermediate geometric information encoded in the implicit field. The resulting features are discriminative and robust to viewpoint change while maintaining rich encoded information. Visual localization is then achieved by aligning the image-based features and the rendered volumetric features. We show the effectiveness of our approach on real-world scenes, demonstrating that our approach outperforms prior and concurrent work on leveraging implicit scene representations for localization. Maxime Pietrantoni, Gabriela Csurka, Martin Humenberger, Torsten Sattler |
3DV | 3 |
| 2023 | SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual LocalizationabstractInspired by properties of semantic segmentation, in this paper we investigate how to leverage robust image segmentation in the context of privacy-preserving visual localization. We propose a new localization framework, SegLoc, that leverages image segmentation to create robust, compact, and privacy-preserving scene representations, i.e., 3D maps. We build upon the correspondence-supervised, fine-grained segmentation approach from [42], making it more robust by learning a set of cluster labels with discriminative clustering, additional consistency regularization terms and we jointly learn a global image representation along with a dense local representation. In our localization pipeline, the former will be used for retrieving the most similar images, the latter to refine the retrieved poses by minimizing the label inconsistency between the 3D points of the map and their projection onto the query image. In various experiments, we show that our proposed representation allows to achieve (close-to) state-of-the-art pose estimation results while only using a compact 3D map that does not contain enough information about the original images for an attacker to reconstruct personal information. Maxime Pietrantoni, Martin Humenberger, Torsten Sattler, Gabriela Csurka |
CVPR | 2 |
| 2023 | 4DHumanOutfit: A multi-subject 4D dataset of human motion sequences in varying outfits exhibiting large displacements
Matthieu Armando, Laurence Boissieux, Edmond Boyer, Jean-Sébastien Franco, Martin Humenberger, Christophe Legras, Vincent Leroy 0003, Mathieu Marsot, Julien Pansiot, Sergi Pujades, Rim Rekik, Grégory Rogez, Anilkumar Swamy, Stefanie Wuhrer |
Comput. Vis. Image Underst. | 5 |
| 2022 | Investigating the Role of Image Retrieval for Visual Localization
Martin Humenberger, Yohann Cabon, Noé Pion, Philippe Weinzaepfel, Nicolas Guérin, Torsten Sattler, Gabriela Csurka |
Int. J. Comput. Vis. | 1 |
| 2021 | Large-Scale Localization Datasets in Crowded Indoor SpacesabstractEstimating the precise location of a camera using visual localization enables interesting applications such as augmented reality or robot navigation. This is particularly useful in indoor environments where other localization technologies, such as GNSS, fail. Indoor spaces impose interesting challenges on visual localization algorithms: occlusions due to people, textureless surfaces, large viewpoint changes, low light, repetitive textures, etc. Existing indoor datasets are either comparably small or do only cover a subset of the mentioned challenges. In this paper, we introduce 5 new indoor datasets for visual localization in challenging real-world environments. They were captured in a large shopping mall and a large metro station in Seoul, South Korea, using a dedicated mapping platform consisting of 10 cameras and 2 laser scanners. In order to obtain accurate ground truth camera poses, we developed a robust LiDAR SLAM which provides initial poses that are then refined using a novel structure-from-motion based optimization. We present a benchmark of modern visual localization algorithms on these challenging datasets showing superior performance of structure-based methods using robust image features. The datasets are available at: https://naverlabs.com/datasets Soohyun Ryu, Suyong Yeon, Yonghan Lee 0001, Deokhwa Kim, Cheolho Han, Yohann Cabon, Philippe Weinzaepfel, Nicolas Guérin, Gabriela Csurka, Martin Humenberger |
CVPR | 11 |
| 2021 | On the Limits of Pseudo Ground Truth in Visual Camera Re-localisationabstractBenchmark datasets that measure camera pose accuracy have driven progress in visual re-localisation research. To obtain poses for thousands of images, it is common to use a reference algorithm to generate pseudo ground truth. Popular choices include Structure-from-Motion (SfM) and Simultaneous-Localisation-and-Mapping (SLAM) using additional sensors like depth cameras if available. Re-localisation benchmarks thus measure how well each method replicates the results of the reference algorithm. This begs the question whether the choice of the reference algorithm favours a certain family of re-localisation methods. This paper analyzes two widely used re-localisation datasets and shows that evaluation outcomes indeed vary with the choice of the reference algorithm. We thus question common beliefs in the re-localisation literature, namely that learning-based scene coordinate regression outperforms classical feature-based methods, and that RGB-D-based methods outperform RGB-based methods. We argue that any claims on ranking re-localisation methods should take the type of the reference algorithm, and the similarity of the methods to the reference algorithm, into account. Eric Brachmann, Martin Humenberger, Carsten Rother, Torsten Sattler |
ICCV | 2 |
| 2021 | Robust Automatic Monocular Vehicle Speed Estimation for Traffic SurveillanceabstractEven though CCTV cameras are widely deployed for traffic surveillance and have therefore the potential of becoming cheap automated sensors for traffic speed analysis, their large-scale usage toward this goal has not been reported yet. A key difficulty lies in fact in the camera calibration phase. Existing state-of-the-art methods perform the calibration using image processing or keypoint detection techniques that require high-quality video streams, yet typical CCTV footage is low-resolution and noisy. As a result, these methods largely fail in real-world conditions. In contrast, we propose two novel calibration techniques whose only inputs come from an off-the-shelf object detector. Both methods consider multiple detections jointly, leveraging the fact that cars have similar and well-known 3D shapes with normalized dimensions. The first one is based on minimizing an energy function corresponding to a 3D reprojection error, the second one instead learns from synthetic training data to predict the scene geometry directly. Noticing the lack of speed estimation benchmarks faithfully reflecting the actual quality of surveillance cameras, we introduce a novel dataset collected from public CCTV streams. Experimental results conducted on three diverse benchmarks demonstrate excellent speed estimation accuracy that could enable the wide use of CCTV cameras for traffic analysis, even in challenging conditions where state-of-the-art methods completely fail. Additional information can be found on our project web page: https://rebrand.ly/nle-cctv Jérôme Revaud, Martin Humenberger |
ICCV | 2 |
| 2020 | Benchmarking Image Retrieval for Visual LocalizationabstractVisual localization, i.e., camera pose estimation in a known scene, is a core component of technologies such as autonomous driving and augmented reality. State-of-the-art localization approaches often rely on image retrieval techniques for one of two tasks: (1) provide an approximate pose estimate or (2) determine which parts of the scene are potentially visible in a given query image. It is common practice to use state-of-the-art image retrieval algorithms for these tasks. These algorithms are often trained for the goal of retrieving the same landmark under a large range of viewpoint changes. However, robustness to viewpoint changes is not necessarily desirable in the context of visual localization. This paper focuses on understanding the role of image retrieval for multiple visual localization tasks. We introduce a benchmark setup and compare state-of-the-art retrieval representations on multiple datasets. We show that retrieval performance on classical landmark retrieval/recognition tasks correlates only for some but not all tasks to localization performance. This indicates a need for retrieval approaches specifically designed for localization tasks. Our benchmark and evaluation protocols are available at https://github.com/naver/kapture-localization. Noé Pion, Martin Humenberger, Gabriela Csurka, Yohann Cabon, Torsten Sattler |
3DV | 2 |
| 2020 | Estimating Low-Rank Region Likelihood MapsabstractLow-rank regions capture geometrically meaningful structures in an image which encompass typical local features such as edges, corners and all kinds of regular, symmetric, often repetitive patterns, that are commonly found in man-made environment. While such patterns are challenging current state-of-the-art feature correspondence methods, the recovered homography of a low-rank texture readily provides 3D structure with respect to a 3D plane, without any prior knowledge of the visual information on that plane. However, the automatic and efficient detection of the broad class of low-rank regions is unsolved. Herein, we propose a novel self-supervised low-rank region detection deep network that predicts a low-rank likelihood map from an image. The evaluation of our method on real-world datasets shows not only that it reliably predicts low-rank regions in the image similarly to our baseline method, but thanks to the data augmentations used in the training phase it generalizes well to difficult cases (e.g. day/night lighting, low contrast, underexposure) where the baseline prediction fails. Gabriela Csurka, Zoltan Kato, Andor Juhasz, Martin Humenberger |
CVPR | 4 |
| 2019 | Visual Localization by Learning Objects-Of-Interest Dense Match RegressionabstractWe introduce a novel CNN-based approach for visual localization from a single RGB image that relies on densely matching a set of Objects-of-Interest (OOIs). In this paper, we focus on planar objects which are highly descriptive in an environment, such as paintings in museums or logos and storefronts in malls or airports. For each OOI, we define a reference image for which 3D world coordinates are available. Given a query image, our CNN model detects the OOIs, segments them and finds a dense set of 2D-2D matches between each detected OOI and its corresponding reference image. Given these 2D-2D matches, together with the 3D world coordinates of each reference image, we obtain a set of 2D-3D matches from which solving a Perspective-n-Point problem gives a pose estimate. We show that 2D-3D matches for reference images, as well as OOI annotations can be obtained for all training images from a single instance annotation per OOI by leveraging Structure-from-Motion reconstruction. We introduce a novel synthetic dataset, VirtualGallery, which targets challenges such as varying lighting conditions and different occlusion levels. Our results show that our method achieves high precision and is robust to these challenges. We also experiment using the Baidu localization dataset captured in a shopping mall. Our approach is the first deep regression-based method to scale to such a larger environment. Philippe Weinzaepfel, Gabriela Csurka, Yohann Cabon, Martin Humenberger |
CVPR | 4 |
| 2019 | R2D2: Reliable and Repeatable Detector and DescriptorabstractInterest point detection and local feature description are fundamental steps in many computer vision applications. Classical approaches are based on a detect-then-describe paradigm where separate handcrafted methods are used to first identify repeatable keypoints and then represent them with a local descriptor. Neural networks trained with metric learning losses have recently caught up with these techniques, focusing on learning repeatable saliency maps for keypoint detection or learning descriptors at the detected keypoint locations. In this work, we argue that repeatable regions are not necessarily discriminative and can therefore lead to select suboptimal keypoints. Furthermore, we claim that descriptors should be learned only in regions for which matching can be performed with high confidence. We thus propose to jointly learn keypoint detection and description together with a predictor of the local descriptor discriminativeness. This allows to avoid ambiguous areas, thus leading to reliable keypoint detection and description. Our detection-and-description approach simultaneously outputs sparse, repeatable and reliable keypoints that outperforms state-of-the-art detectors and descriptors on the HPatches dataset and on the recent Aachen Day-Night localization benchmark. Jérôme Revaud, César Roberto de Souza, Martin Humenberger, Philippe Weinzaepfel |
NeurIPS | 3 |
| 2017 | Analyzing Computer Vision Data - The Good, the Bad and the UglyabstractIn recent years, a great number of datasets were published to train and evaluate computer vision (CV) algorithms. These valuable contributions helped to push CV solutions to a level where they can be used for safety-relevant applications, such as autonomous driving. However, major questions concerning quality and usefulness of test data for CV evaluation are still unanswered. Researchers and engineers try to cover all test cases by using as much test data as possible. In this paper, we propose a different solution for this challenge. We introduce a method for dataset analysis which builds upon an improved version of the CV-HAZOP checklist, a list of potential hazards within the CV domain. Picking stereo vision as an example, we provide an extensive survey of 28 datasets covering the last two decades. We create a tailored checklist and apply it to the datasets Middlebury, KITTI, Sintel, Freiburg, and HCI to present a thorough characterization and quantitative comparison. We confirm the usability of our checklist for identification of challenging stereo situations by applying nine state-of-the-art stereo matching algorithms on the analyzed datasets, showing that hazard frames correlate with difficult frames. We show that challenging datasets still allow a meaningful algorithm evaluation even for small subsets. Finally, we provide a list of missing test cases that are still not covered by current datasets as inspiration for researchers who want to participate in future dataset creation. Oliver Zendel 0001, Katrin Honauer, Markus Murschitz, Martin Humenberger, Gustavo Fernández Domínguez |
CVPR | 4 |
| 2017 | How Good Is My Test Data? Introducing Safety Analysis for Computer VisionabstractGood test data is crucial for driving new developments in computer vision (CV), but two questions remain unanswered: which situations should be covered by the test data, and how much testing is enough to reach a conclusion? In this paper we propose a new answer to these questions using a standard procedure devised by the safety community to validate complex systems: the hazard and operability analysis (HAZOP). It is designed to systematically identify possible causes of system failure or performance loss. We introduce a generic CV model that creates the basis for the hazard analysis and—for the first time—apply an extensive HAZOP to the CV domain. The result is a publicly available checklist with more than 900 identified individual hazards. This checklist can be utilized to evaluate existing test datasets by quantifying the covered hazards. We evaluate our approach by first analyzing and annotating the popular stereo vision test datasets Middlebury and KITTI. Second, we demonstrate a clearly negative influence of the hazards in the checklist on the performance of six popular stereo matching algorithms. The presented approach is a useful tool to evaluate and improve test datasets and creates a common basis for future dataset designs. Oliver Zendel 0001, Markus Murschitz, Martin Humenberger, Wolfgang Herzner |
Int. J. Comput. Vis. | 3 |
| 2016 | Guided Matching Based on Statistical Optical Flow for Fast and Robust Correspondence Analysis
Josef Maier, Martin Humenberger, Markus Murschitz, Oliver Zendel 0001, Markus Vincze |
ECCV (7) | 2 |
| 2015 | CV-HAZOP: Introducing Test Data Validation for Computer VisionabstractTest data plays an important role in computer vision (CV) but is plagued by two questions: Which situations should be covered by the test data and have we tested enough to reach a conclusion? In this paper we propose a new solution answering these questions using a standard procedure devised by the safety community to validate complex systems: The Hazard and Operability Analysis (HAZOP). It is designed to systematically search and identify difficult, performance-decreasing situations and aspects. We introduce a generic CV model that creates the basis for the hazard analysis and, for the first time, apply an extensive HAZOP to the CV domain. The result is a publicly available checklist with more than 900 identified individual hazards. This checklist can be used to evaluate existing test datasets by quantifying the amount of covered hazards. We evaluate our approach by first analyzing and annotating the popular stereo vision test datasets Middlebury and KITTI. Second, we compare the performance of six popular stereo matching algorithms at the identified hazards from our checklist with their average performance and show, as expected, a clear negative influence of the hazards. The presented approach is a useful tool to evaluate and improve test datasets and creates a common basis for future dataset designs. Oliver Zendel 0001, Markus Murschitz, Martin Humenberger, Wolfgang Herzner |
ICCV | 3 |
| 2015 | Fast PatchMatch stereo matching using cross-scale cost fusion for automotive applicationsabstractDue to recent developments of low-cost image sensors and high-performance embedded processing hardware, future cars and automotive systems will increasingly use binocular stereo vision for environmental perception. However, research and development in stereo vision is still ongoing since there are many challenges unsolved. In this paper, we propose a fast and accurate stereo matching algorithm, designed for automotive applications. It convincingly handles real-world scenes containing complex, textureless, and slanted surfaces. To achieve that, we propose an improved PatchMatch stereo algorithm that combines a census-based cost function with Semi-Global Matching optimization integrated in a cross-scale fusion processing scheme. To further accelerate the algorithm, we propose a novel enhancement approach for PatchMatch-based approximation which allows us to skip the random search or at least significantly reduce the number of iterations. Our method is ranked in the upper third of the KITTI benchmark and among the top performers in terms of processing time. Ji-Ho Cho, Martin Humenberger |
Intelligent Vehicles Symposium | 2 |
| 2012 | Optimization of a Neural Network for Computer Vision Based Fall Detection with Fixed-Point Arithmetic
Christoph Sulzbachner, Martin Humenberger, Ágoston Srp, Ferenc Vajda |
ICONIP (4) | 2 |
| 2012 | CARE: A dynamic stereo vision sensor system for fall detectionabstractThis paper presents a recently developed dynamic stereo vision sensor system and its application for fall detection towards safety for elderly at home. The system consists of (1) two optical detector chips with 304×240 event-driven pixels which are only sensitive to relative light intensity changes, (2) an FPGA for interfacing the detectors, early data processing, and stereo matching for depth map reconstruction, (3) a digital signal processor for interpreting the sensor data in real-time for fall recognition, and (4) a wireless communication module for instantly alerting caring institutions. This system was designed for incident detection in private homes of elderly to foster safety and security. The two main advantages of the system, compared to existing wearable systems are from the application's point of view: (a) the stationary installation has a better acceptance for independent living comparing to permanent wearing devices, and (b) the privacy of the system is systematically ensured since the vision detector does not produce real images such as classic video sensors. The system can actually process about 300 kevents per second. It was evaluated using 500 fall cases acquired with a stuntman. More than 90% positive detections were reported. We will show a live demonstration during ISCAS2012 of the sensor system and its capabilities. Ahmed Nabil Belbachir, Martin Litzenberger, Stephan Schraml, Michael Hofstätter, Michael D. Bauer, Peter Schön, Martin Humenberger, Christoph Sulzbachner, Tommi Lunden, M. Merne |
ISCAS | 7 |
| 2010 | A fast stereo matching algorithm suitable for embedded real-time systems
Martin Humenberger, Christian Zinner, Wilfried Kubinger, Markus Vincze |
Comput. Vis. Image Underst. | 1 |
| 2007 | Hardware implementation of an SAD based stereo vision algorithmabstractThis paper presents the hardware implementation of a stereo vision core algorithm, that runs in real-time and is targeted at automotive applications. The algorithm is based on the sum of absolute differences (SAD) and computes the disparity map using 320 times 240 input images with a maximum disparity of 100 pixels. The hardware operates at a frequency of 65 MHz and achieves a frame rate of 425 fps by calculating the data highly parallel and pipelined. Thus an implemented and basically optimized software solution, running on an Intel Pentium 4 with 3 GHz clock frequency is 166 times outperformed. Kristian Ambrosch, Wilfried Kubinger, Martin Humenberger, Andreas Steininger |
CVPR | 3 |
| 2007 | Boosting the performance of embedded vision systems using a DSP/FPGA co-processor systemabstractSensor systems for robotics and autonomous systems usually require small-sized and power-aware dedicated solutions for their realization. Therefore, an embedded system is the first choice - but the drawback is weaker computing power compared to state-of-the-art PC-based systems. This paper describes our approach on "boosting" the computing power of embedded vision systems. Our system consists of a platform based on a digital signal processor (DSP) enhanced by an additional field programmable gate array (FPGA) used as a co-processor. In our novel approach called resource optimized co-processing, both the DSP and the FPGA are driven in parallel for the execution of crucial parts of the vision algorithms. Thereby, through efficient usage of system resources a significant increase of the system performance is possible. The approach and the achievable profit in computing power is outlined in the paper. Based on an example case - the realization of a robot soccer embedded vision sensor - the usefulness and the powerfulness of our approach is demonstrated. In this demo case, by applying resource optimized co-processing, the most crucial and computing-intensive function was executed twice as fast as before. Thus, we were able to fulfill the stringent real-time requirements of the vision system. Franz Rinnerthaler, Wilfried Kubinger, Josef Langer, Martin Humenberger, Stefan Borbely |
SMC | 4 |