Hannes Kruppa

dblp:79/1290 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0009-2616-3469ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Traffic Sign Localization and Orientation Classification for Automated Map Updating
abstract
High Definition (HD) maps, containing detailed road information, are essential for autonomous driving and many geo-related tasks. Recent developments in computer vision make it possible to automate the labor-intensive HD map maintenance work, such as localizing traffic signs within a road network. However, updating traffic signs to HD maps is non-trivial, as it not only requires precise geo-location but also requires confirming whether a sign belongs to a specific road. In our work, we develop an end-to-end automated traffic sign update system, termed AutoTS, which is capable of using an image sequence collected during vehicle operation to extract the geo-location of a traffic sign and determine whether it belongs to the road driven on, from its orientation. In AutoTS, we design a noise and sparsity adaptive localization module, which can filter noisy location points and derive a geo-location from sparse location points. To identify the orientation of traffic signs, we devise a position-aware orientation classification module, which uses the ROI feature and the position-aware SIFT feature to explore the orientation characteristic and understand the road context. To facilitate the evaluation of the proposed method, we construct a traffic sign localization and orientation classification benchmark, KITTI-TS. Our AutoTS achieves an MAE of 2.38 meters in traffic sign localization, while the accuracy in orientation classification reaches 88.89%.
Xianjing Han, Wenmiao Hu, Xuemeng Song, Hannes Kruppa, See-Kiong Ng, Roger Zimmermann
IEEE Trans. Circuits Syst. Video Technol.4
2025 GAN-Assisted Road Segmentation from Satellite Imagery
abstract
Geo-information extraction from satellite imagery has become crucial to carry out large-scale ground surveys in a short amount of time. With the increasing number of commercial satellites launched into orbit in recent years, high-resolution RGB color remote sensing imagery has attracted a lot of attention. However, because of the high cost of image acquisition and even more complicated annotation procedures, there are limited high-resolution satellite datasets available. Compared to close-range imagery datasets, existing satellite datasets have a much lower number of images and cover only a few scenarios (cities, background environments, etc.). They may not be sufficient for training robust learning models that fit all environmental conditions or be representative enough for training regional models that optimize for local scenarios. Instead of collecting and annotating more data, using synthetic images could be another solution to boost the performance of a model. This article proposes a GAN-assisted training scheme for road segmentation from high-resolution RGB color satellite images, which includes three critical components: (a) synthetic training sample generation, (b) synthetic training sample selection, and (c) assisted training strategy. Apart from the GeoPalette and cSinGAN image generators introduced in our prior work, this article explains in detail how to generate new training pairs using OpenStreetMap (OSM) and introduces a new set of evaluation metrics for selecting synthetic training pairs from a pool of generated samples. We conduct extensive quantitative and qualitative experiments to compare different image generators and training strategies. Our experiments on the downstream road segmentation task show that (1) our proposed metrics are more aligned with the trained model performance compared to commonly used GAN evaluation metrics such as the Fréchet inception distance (FID); and (2) by using synthetic data with the best training strategy, the model performance, mean Intersection over Union (mean IoU), is improved from 60.92% to 64.44%, when 1,000 real training pairs are available for learning, which reaches a similar level of performance as a model that is standard-trained with 4,000 real images (64.59%), i.e., enabling a 4-fold reduction in real dataset size.
Wenmiao Hu, Yifang Yin, Ying Kiat Tan, An Tran, Hannes Kruppa, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.5
2023 PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search
abstract
Satellite-based street-view information extraction by cross-view matching refers to a task that extracts the location and orientation information of a given street-view image query by using one or multiple geo-referenced satellite images. Recent work has initiated a new research direction to find accurate information within a local area covered by one satellite image centered at a location prior (e.g., from GPS). It can be used as a standalone solution or complementary step following a large-scale search with multiple satellite candidates. However, these existing works require an accurate initial orientation (angle) prior (e.g., from IMU) and/or do not efficiently search through all possible poses. To allow efficient search and to give accurate prediction regardless of the existence or the accuracy of the angle prior, we present PetalView extractors with multi-scale search. The PetalView extractors give semantically meaningful features that are equivalent across two drastically different views, and the multi-scale search strategy efficiently inspects the satellite image from coarse to fine granularity to provide sub-meter and sub-degree precision extraction. Moreover, when an angle prior is given, we propose a learnable prior angle mixer to utilize this information. Our method obtains the best performance on the VIGOR dataset and successfully improves the performance on KITTI dataset test~1 set with the recall within 1 meter (r@1m) for location estimation to 68.88% and recall within 1 degree (r@1d) 21.10% when no angle prior is available, and with angle prior achieves stable estimations at r@1m and r@1d above 70% and 21%, up to a 40-degree noise level.
Wenmiao Hu, Yichen Zhang 0002, Yuxuan Liang 0002, Xianjing Han, Yifang Yin, Hannes Kruppa, See-Kiong Ng, Roger Zimmermann
ACM Multimedia6
2022 Beyond Geo-localization: Fine-grained Orientation of Street-view Images by Cross-view Matching with Satellite Imagery
abstract
Street-view imagery provides us with novel experiences to explore different places remotely. Carefully calibrated street-view images (e.g., Google Street View) can be used for different downstream tasks, e.g., navigation, map features extraction. As personal high-quality cameras have become much more affordable and portable, an enormous amount of crowdsourced street-view images are uploaded to the internet, but commonly with missing or noisy sensor information. To prepare this hidden treasure for "ready-to-use" status, determining missing location information and camera orientation angles are two equally important tasks. Recent methods have achieved high performance on geo-localization of street-view images by cross-view matching with a pool of geo-referenced satellite imagery. However, most of the existing works focus more on geo-localization than estimating the image orientation. In this work, we re-state the importance of finding fine-grained orientation for street-view images, formally define the problem and provide a set of evaluation metrics to assess the quality of the orientation estimation. We propose two methods to improve the granularity of the orientation estimation, achieving 82.4% and 72.3% accuracy for images with estimated angle errors below 2 degrees for CVUSA and CVACT datasets, corresponding to 34.9% and 28.2% absolute improvement compared to previous works. Integrating fine-grained orientation estimation in training also improves the performance on geo-localization, giving top 1 recall 95.5%/85.5% and 86.8%/80.4% for orientation known/unknown tests on the two datasets.
Wenmiao Hu, Yichen Zhang 0002, Yuxuan Liang 0002, Yifang Yin, Andrei Georgescu, An Tran, Hannes Kruppa, See-Kiong Ng, Roger Zimmermann
ACM Multimedia7
2022 A Context-enriched Satellite Imagery Dataset and an Approach for Parking Lot Detection
abstract
Automatic detection of geoinformation from satellite images has been a fundamental yet challenging problem, which aims to reduce the manual effort of human annotators in maintaining an up-to-date digital map. There are currently several high-resolution satellite imagery datasets that are publicly available. However, the associated ground-truth annotations are limited to road, building, and land use, while the annotations of other geographic objects or attributes are mostly not available. To bridge the gap, we present Grab-Pklot, the first high-resolution and context-enriched satellite imagery dataset for parking lot detection. Our dataset consists of 1344 satellite images with the ground-truth annotations of carparks in Singapore. Motivated by the observation that carparks are mostly co-appear with other geographic objects, we associate each satellite image in our dataset with the surrounding contextual information of road and building, given in the format of multi-channel images. As a side contribution, we present a fusion-based segmentation approach to demonstrate that the parking lot detection accuracy can be improved by modeling the correlations between parking lots and other geographic objects. Experiments on our dataset provide baseline results as well as new insights into the challenges and opportunities in parking lot detection from satellite images.
Yifang Yin, Wenmiao Hu, An Tran, Hannes Kruppa, Roger Zimmermann, See-Kiong Ng
WACV4
2021 GeoPalette: Road Segmentation with Limited Satellite Imagery
abstract
In recent years, Geo-information extraction from high-resolution satellite imagery has attracted a lot of attention. However, because of the high cost of image acquisition and annotation, there are limited datasets available. Compared to close-range imagery datasets, existing satellite datasets have a much lower number of images and cover only a few scenarios (cities, background environments, etc.). They may not be sufficient for training robust learning models that fit all environmental conditions or be representative enough for training regional models that optimize for local scenarios. In this study, we propose GeoPalette, a Generative Adversarial Network (GAN) based tool to generate additional synthetic training samples for boosting model performance when the training dataset is limited. Our experiments on road segmentation show that using additional synthetic data can improves the model performance mean Intersection over Union (mIoU) from 60.92% to 64.44%, when 1,000 real training pairs are available for learning, which reaches a similar level of performance as a model is standard-trained on 4,000 real pairs (64.59%), i.e., a 4-fold reduction in real dataset size.
Wenmiao Hu, Yifang Yin, Ying Kiat Tan, An Tran, Hannes Kruppa, Roger Zimmermann
SIGSPATIAL/GIS5
2021 TEST-GCN: Topologically Enhanced Spatial-Temporal Graph Convolutional Networks for Traffic Forecasting
abstract
Accurate traffic forecasting is a fundamental challenge of location-based systems. Recent works were able to achieve state-of-the-art results by incorporating Graph Convolutional Networks (GCN) to capture spatial dependencies in the data. However, these works rely on a fixed latent feature representation of the underlying graph structure, failing to exploit the rich spatial information offered by the road network. In this paper, we propose the Topologically Enhanced Spatial-Temporal Graph Convolutional Network (TEST-GCN), a novel graph convolution model for road traffic speed forecasting based on floating car data, aiming to better capture the spatial dependencies in the data by fully exploiting the characteristics of the road network. We introduce the node and edge embedding layers, using topological attributes to iteratively improve the latent feature representation of the road network. We show that our model effectively captures both spatial and temporal dependencies in the data, consistently outperforming state-of-the-art methods in road traffic speed prediction, achieving approximately 50 % reduction in model size and 33% improvement in empirical computational times.
Muhammad Afif Ali, Suriya Venkatesan, Victor C. Liang, Hannes Kruppa
ICDM4
2003 Using Local Context To Improve Face Detection
abstract
Most face detection algorithms locate faces by classifying the content of a detection window iterating over all positions and scales of the input image. Recent developments have accelerated this process up to real-time performance at high levels of accuracy. However, even the best of today's computational systems are far from being able to compete with the detection capabilities of the human visual system. Psychophysical experiments have shown the importance of local context in the face detection process. In this paper we investigate the role of local context for face detection algorithms. In experiments on two large data sets we find that using local context can significantly increase the number of correct detections, particularly in low resolution cases, uncommon poses or individual appearances as well as occlusions.
Hannes Kruppa, Bernt Schiele
BMVC1
2001 Hierarchical Combination of Object Models using Mutual Information
abstract
Combining different and complementary object models promises to increase the robustness and generality of today's computer vision algorithms. This paper introduces a new method for combining different object models by determining a configuration of the models which maximizes their mutual information. The combination scheme consequently creates a unified hypothesis from multiple object models "on the fly" without prior training. To validate the effectiveness of the proposed method, the approach is applied to the detection of faces combining the output of three different models. 1
Hannes Kruppa, Bernt Schiele
BMVC1