Hanno Gottschalk

dblp:154/1735 · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
22since 2021 · last 2026
0000-0003-2167-2028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VideoGAN-Based Trajectory Proposal for Automated Vehicles
Annajoyce Mariani, Kira Maag, Hanno Gottschalk
ICAART (2)3
2025 On the Influence of Shape, Texture and Color for Learning Semantic Segmentation
abstract
Recent research has investigated the shape and texture biases of pre-trained deep neural networks (DNNs) in image classification. Those works test how much a trained DNN relies on specific image cues like texture. The present study shifts the focus to understanding the cue influence during training, analyzing what DNNs can learn from shape, texture, and color cues in absence of the others; investigating their individual and combined influence on the learning success. We analyze these cue influences at multiple levels by decomposing datasets into cue-specific versions. Addressing semantic segmentation, we learn the given task from these reduced cue datasets, creating cue experts. Early fusion of cues is performed by constructing appropriate datasets. This is complemented by a late fusion of experts which allows us to study cue influence location-dependent on pixel level. Experiments on Cityscapes, PASCAL Context, and a synthetic CARLA dataset show that while no single cue dominates, the shape + color expert predominantly improves the prediction of small objects and border pixels. The cue performance order is consistent for the tested convolutional and transformer architecture, indicating similar cue extraction capabilities, although pre-trained transformers are said to be more biased towards shape than convolutional neural networks.
Annika Mütze, Natalie Grabowsky, Edgar Heinert, Matthias Rottmann, Hanno Gottschalk
ECAI5
2025 Robust Evolutionary Multi-Objective Neural Architecture Search for Reinforcement Learning (EMNAS-RL)
abstract
This paper introduces Evolutionary Multi-Objective Neural Architecture Search (EMNAS) for the first time to optimize neural network architectures in large-scale Reinforcement Learning (RL) for Autonomous Driving (AD).EMNAS uses genetic algorithms to automate network design, tailored to enhance rewards and reduce model size without compromising performance.Additionally, parallelization techniques are employed to accelerate the search, and teacher-student methodologies are implemented to ensure scalable optimization.Experimental results demonstrate that tailored EMNAS outperforms manually designed models, achieving higher rewards with fewer parameters.
Nihal Acharya Adde, Alexandra Gianzina, Hanno Gottschalk, Andreas Ebert
ESANN3
2025 Adaptive Neural Networks for Intelligent Data-Driven Development
abstract
Advances in machine learning methods for computer vision tasks have led to their consideration for safety-critical applications like autonomous driving. However, effectively integrating these methods into the automotive development lifecycle remains challenging. Since the performance of machine learning algorithms relies heavily on the training data provided, the data and model development lifecycle play a key role in successfully integrating these components into the product development lifecycle. Existing models frequently encounter difficulties recognizing or adapting to novel instances not present in the original training dataset. This poses a significant risk for reliable deployment in dynamic environments. To address this challenge, we propose an adaptive neural network architecture and an iterative development framework that enables users to efficiently incorporate previously unknown objects into the current perception system. Our approach builds on continuous learning, emphasizing the necessity of dynamic updates to reflect real-world deployment conditions. Specifically, we introduce a pipeline with three key components: (1) a scalable network extension strategy to integrate new classes while preserving existing performance, (2) a dynamic OoD detection component that requires no additional retraining for newly added classes, and (3) a retrieval-based data augmentation process tailored for safety-critical deployments. The integration of these components establishes a pragmatic and adaptive pipeline for the continuous evolution of perception systems in the context of autonomous driving.
Youssef Shoeb, Azarm Nowzad, Hanno Gottschalk
IV3
2025 FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
abstract
While Multimodal Large Language Models (MLLMs) offer strong perception and reasoning capabilities for image-text input, Visual Question Answering (VQA) focusing on small image details still remains a challenge. Although visual cropping techniques seem promising, recent approaches have several limitations: the need for task-specific fine-tuning, low efficiency due to uninformed exhaustive search, or incompatibility with efficient attention implementations. We address these shortcomings by proposing a training-free visual cropping method, dubbed FOCUS, that leverages MLLM-internal representations to guide the search for the most relevant image region. This is accomplished in four steps: first, we identify the target object(s) in the VQA prompt; second, we compute an object relevance map using the key-value (KV) cache; third, we propose and rank relevant image regions based on the map; and finally, we perform the fine-grained VQA task using the top-ranked region. As a result of this informed search strategy, FOCUS achieves strong performance across four fine-grained VQA datasets and three types of MLLMs. It outperforms three popular visual cropping methods in both accuracy and efficiency, and matches the best-performing baseline, ZoomEye, while requiring 3 – 6.5 × less compute.
Liangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger, Thorsten Bagdonat, Hanno Gottschalk, Leo Schwinn
NeurIPS6
2025 ECD: Efficient Contrastive Decoding with Probabilistic Hallucination Detection
Laura Fieback, Nishilkumar Balar, Jakob Spiegelberg, Hanno Gottschalk
ECML/PKDD (7)4
2024 Strong but Simple: A Baseline for Domain Generalized Dense Perception by CLIP-Based Transfer Learning
Christoph Hümmer, Manuel Schwonberg, Liangwei Zhou, Hu Cao, Alois C. Knoll, Hanno Gottschalk
ACCV (10)6
2024 ResBuilder: Automated Learning of Depth with Residual Structures
Julian Burghoff, Matthias Rottmann, Jill von Conta, Sebastian Schoenen, Andreas Witte, Hanno Gottschalk
ICANN (1)6
2024 Have We Ever Encountered This Before? Retrieving Out-of-Distribution Road Obstacles from Driving Scenes
abstract
In the life cycle of highly automated systems operating in an open and dynamic environment, the ability to adjust to emerging challenges is crucial. For systems integrating data-driven AI-based components, rapid responses to deployment issues require fast access to related data for testing and reconfiguration. In the context of automated driving, this especially applies to road obstacles not included in the training data, commonly referred to as out-of-distribution (OoD) road obstacles. Given the availability of large uncurated driving scene recordings, a pragmatic approach is to query a database to retrieve similar scenarios featuring the same safety concerns due to OoD road obstacles. In this work, we extend beyond identifying OoD road obstacles in video streams and offer a comprehensive approach to extract sequences of OoD road obstacles using text queries, thereby proposing a way of curating a collection of OoD data for subsequent analysis. Our proposed method leverages the recent advances in OoD segmentation and multi-modal foundation models to identify and efficiently extract safety-relevant scenes from unlabeled videos. We present a first approach for the novel task of text-based OoD object retrieval, which addresses the question "Have we ever encountered this before?".
Youssef Shoeb, Robin Chan 0001, Gesina Schwalbe, Azarm Nowzad, Fatma Güney, Hanno Gottschalk
WACV6
2023 Risk Stratification of Malignant Melanoma Using Neural Networks
Julian Burghoff, Leonhard Ackermann, Younes Salahdine, Veronika Bram, Katharina Wunderlich, Julius Balkenhol, Thomas Dirschka, Hanno Gottschalk
ICANN (4)8
2023 Who Breaks Early, Looses: Goal Oriented Training of Deep Neural Networks Based on Port Hamiltonian Dynamics
Julian Burghoff, Marc Heinrich Monells, Hanno Gottschalk
ICANN (10)3
2023 LU-Net: Invertible Neural Networks Based on Matrix Factorization
abstract
LU-Net is a simple and fast architecture for invertible neural networks (INN) that is based on the factorization of quadratic weight matrices A=LU, where L is a lower triangular matrix with ones on the diagonal and U an upper triangular matrix. Instead of learning a fully occupied matrix A, we learn L and U separately. If combined with an invertible activation function, such layers can easily be inverted whenever the diagonal entries of U are different from zero. Also, the computation of the determinant of the Jacobian matrix of such layers is cheap. Consequently, the LU architecture allows for cheap computation of the likelihood via the change of variables formula and can be trained according to the maximum likelihood principle. In our numerical experiments, we test the LU-Net architecture as generative model on several academic datasets. We also provide a detailed comparison with conventional invertible neural networks in terms of performance, training as well as run time.
Robin Chan 0001, Sarina Penquitt, Hanno Gottschalk
IJCNN3
2023 Augmentation-based Domain Generalization for Semantic Segmentation
abstract
Unsupervised Domain Adaptation (UDA) and domain generalization (DG) are two research areas that aim to tackle the lack of generalization of Deep Neural Networks (DNNs) towards unseen domains. While UDA methods have access to unlabeled target images, domain generalization does not involve any target data and only learns generalized features from a source domain. Image-style randomization or augmentation is a popular approach to improve network generalization without access to the target domain. Complex methods are often proposed that disregard the potential of simple image augmentations for out-of-domain generalization. For this reason, we systematically study the in- and out-of-domain generalization capabilities of simple, rule-based image augmentations like blur, noise, color jitter and many more. Based on a full factorial design of experiment design we provide a systematic statistical evaluation of augmentations and their interactions. Our analysis provides both, expected and unexpected, outcomes. Expected, because our experiments confirm the common scientific standard that combination of multiple different augmentations outperforms single augmentations. Unexpected, because combined augmentations perform competitive to state-of-the-art domain generalization approaches, while being significantly simpler and without training overhead. On the challenging synthetic-to-real domain shift between Synthia and Cityscapes we reach 39.5% mIoU compared to 40.9% mIoU of the best previous work. When additionally employing the recent vision transformer architecture DAFormer we outperform these benchmarks with a performance of 44.2% mIoU.
Manuel Schwonberg, Fadoua El Bouazati, Nico M. Schmidt, Hanno Gottschalk
IV4
2023 Gradient-Based Quantification of Epistemic Uncertainty for Deep Object Detectors
abstract
The majority of uncertainty quantification methods for deep object detectors are based on the network output, such as sampling strategies like Monte-Carlo dropout or deep ensembles with straight-forward transfers to object detection. Here, we study gradient-based uncertainty features for object detection. We show that they contain information orthogonal to that of common, output-based uncertainty approximation methods. Meta classification and meta regression are used to produce confidence estimates using gradient features and other methods which are applicable to numerous object detection architectures. Our results show that gradient uncertainty itself performs on par with state-of-the-art methods across different detectors and datasets. We find that combined meta classifiers outperform standalone models. This suggests that sampling strategies may be supplemented by gradient-based uncertainty to obtain improved confidences, contributing to the probabilistic reliability of object detectors in down-stream applications.
Tobias Riedlinger, Matthias Rottmann, Marius Schubert, Hanno Gottschalk
WACV4
2022 Two Video Data Sets for Tracking and Retrieval of Out of Distribution Objects
Kira Maag, Robin Chan 0001, Svenja Uhlemeyer, Kamil Kowol, Hanno Gottschalk
ACCV (5)5
2022 A-Eye: Driving with the Eyes of AI for Corner Case Generation
Kamil Kowol, Stefan Bracke, Hanno Gottschalk
CHIRA3
2022 Towards unsupervised open world semantic segmentation
abstract
For the semantic segmentation of images, state-of-the-art deep neural networks (DNNs) achieve high segmentation accuracy if that task is restricted to a closed set of classes. However, as of now DNNs have limited ability to operate in an open world, where they are tasked to identify pixels belonging to unknown objects and eventually to learn novel classes, incrementally. Humans have the capability to say: “I don’t know what that is, but I’ve already seen something like that”. Therefore, it is desirable to perform such an incremental learning task in an unsupervised fashion. We introduce a method where unknown objects are clustered based on visual similarity. Those clusters are utilized to define new classes and serve as training data for unsupervised incremental learning. More precisely, the connected components of a predicted semantic segmentation are assessed by a segmentation quality estimate. Connected components with a low estimated prediction quality are candidates for a subsequent clustering. Additionally, the component-wise quality assessment allows for obtaining predicted segmentation masks for the image regions potentially containing unknown objects. The respective pixels of such masks are pseudo-labeled and afterwards used for re-training the DNN, i.e., without the use of ground truth generated by humans. In our experiments we demonstrate that, without access to ground truth and even with few data, a DNN’s class space can be extended by a novel class, achieving considerable segmentation accuracy.
Svenja Uhlemeyer, Matthias Rottmann, Hanno Gottschalk
UAI3
2021 YOdar: Uncertainty-based Sensor Fusion for Vehicle Detection with Camera and Radar Sensors
abstract
In this work, we present an uncertainty-based method for sensor fusion with camera and radar data. The outputs of two neural networks, one processing camera and the other one radar data, are combined in an uncertainty aware manner. To this end, we gather the outputs and corresponding meta information for both networks. For each predicted object, the gathered information is post-processed by a gradient boosting method to produce a joint prediction of both networks. In our experiments we combine the YOLOv3 object detection network with a customized $1D$ radar segmentation network and evaluate our method on the nuScenes dataset. In particular we focus on night scenes, where the capability of object detection networks based on camera data is potentially handicapped. Our experiments show, that this approach of uncertainty aware fusion, which is also of very modular nature, significantly gains performance compared to single sensor baselines and is in range of specifically tailored deep learning based fusion approaches.
Kamil Kowol, Matthias Rottmann, Stefan Bracke, Hanno Gottschalk
ICAART (2)4
2021 Entropy Maximization and Meta Classification for Out-of-Distribution Detection in Semantic Segmentation
abstract
Deep neural networks (DNNs) for the semantic segmentation of images are usually trained to operate on a predefined closed set of object classes. This is in contrast to the "open world" setting where DNNs are envisioned to be deployed to. From a functional safety point of view, the ability to detect so-called "out-of-distribution" (OoD) samples, i.e., objects outside of a DNN’s semantic space, is crucial for many applications such as automated driving. A natural baseline approach to OoD detection is to threshold on the pixel-wise softmax entropy. We present a two-step procedure that significantly improves that approach. Firstly, we utilize samples from the COCO dataset as OoD proxy and introduce a second training objective to maximize the softmax entropy on these samples. Starting from pretrained semantic segmentation networks we re-train a number of DNNs on different in-distribution datasets and consistently observe improved OoD detection performance when evaluating on completely disjoint OoD datasets. Secondly, we perform a transparent post-processing step to discard false positive OoD samples by so-called "meta classification." To this end, we apply linear models to a set of hand-crafted metrics derived from the DNN’s softmax probabilities. In our experiments we consistently observe a clear additional gain in OoD detection performance, cutting down the number of detection errors by 52% when comparing the best baseline with our results. We achieve this improvement sacrificing only marginally in original segmentation performance. Therefore, our method contributes to safer DNNs with more reliable overall system performance.
Robin Chan 0001, Matthias Rottmann, Hanno Gottschalk
ICCV3
2021 MetaBox+: A New Region based Active Learning Method for Semantic Segmentation using Priority Maps
abstract
We present a novel region based active learning method for semantic image segmentation, called MetaBox+. For acquisition, we train a meta regression model to estimate the segment-wise Intersection over Union (IoU) of each predicted segment of unlabeled images. This can be understood as an estimation of segment-wise prediction quality. Queried regions are supposed to minimize to competing targets, i.e., low predicted IoU values / segmentation quality and low estimated annotation costs. For estimating the latter we propose a simple but practical method for annotation cost estimation. We compare our method to entropy based methods, where we consider the entropy as uncertainty of the prediction. The comparison and analysis of the results provide insights into annotation costs as well as robustness and variance of the methods. Numerical experiments conducted with two different networks on the Cityscapes dataset clearly demonstrate a reduction of annotation effort compared to random acquisition. Noteworthily, we achieve 95%of the mean Intersection over Union (mIoU), using MetaBox+ compared to when training with the full dataset, with only 10.47% / 32.01% annotation effort for the two networks, respectively.
Pascal Colling, Lutz Roese-Koerner, Hanno Gottschalk, Matthias Rottmann
ICPRAM3
2021 False Positive Detection and Prediction Quality Estimation for LiDAR Point Cloud Segmentation
abstract
We present a novel post-processing tool for semantic segmentation of LiDAR point cloud data, called LidarMetaSeg, which estimates the prediction quality segmentwise. For this purpose we compute dispersion measures based on network probability outputs as well as feature measures based on point cloud input features and aggregate them on segment level. These aggregated measures are used to train a meta classification model to predict whether a predicted segment is a false positive or not and a meta regression model to predict the segmentwise intersection over union. Both models can then be applied to semantic segmentation inferences without knowing the ground truth. In our experiments we use different LiDAR segmentation models and datasets and analyze the power of our method. We show that our results outperform other standard approaches.
Pascal Colling, Matthias Rottmann, Lutz Roese-Koerner, Hanno Gottschalk
ICTAI4
2021 Improving Video Instance Segmentation by Light-weight Temporal Uncertainty Estimates
abstract
Instance segmentation with neural networks is an essential task in environment perception. In many works, it has been observed that neural networks can predict false positive instances with high confidence values and true positives with low ones. Thus, it is important to accurately model the uncertainties of neural networks in order to prevent safety issues and foster interpretability. In applications such as automated driving, the reliability of neural networks is of highest interest. In this paper, we present a time-dynamic approach to model uncertainties of instance segmentation networks and apply this to the detection of false positives as well as the estimation of prediction quality. The availability of image sequences in online applications allows for tracking instances over multiple frames. Based on an instances history of shape and uncertainty information, we construct temporal instance-wise aggregated metrics. The latter are used as input to post-processing models that estimate the prediction quality in terms of instance-wise intersection over union. The proposed method only requires a readily trained neural network (that may operate on single frames) and video sequence input. In our experiments, we further demonstrate the use of the proposed method by replacing the traditional score value from object detection and thereby improving the overall performance of the instance segmentation network.
Kira Maag, Matthias Rottmann, Serin Varghese, Fabian Hüger, Peter Schlicht, Hanno Gottschalk
IJCNN6
2020 Detection of False Positive and False Negative Samples in Semantic Segmentation
abstract
In recent years, deep learning methods have outperformed other methods in image recognition. This has fostered imagination of potential application of deep learning technology including safety relevant applications like the interpretation of medical images or autonomous driving. The passage from assistance of a human decision maker to ever more automated systems however increases the need to properly handle the failure modes of deep learning modules. In this contribution, we review a set of techniques for the self-monitoring of machine-learning algorithms based on uncertainty quantification. In particular, we apply this to the task of semantic segmentation, where the machine learning algorithm decomposes an image according to semantic categories. We discuss false positive and false negative error modes at instance-level and review techniques for the detection of such errors that have been recently proposed by the authors. We also give an outlook on future research directions.
Matthias Rottmann, Kira Maag, Robin Chan 0001, Fabian Hüger, Peter Schlicht, Hanno Gottschalk
DATE6
2020 Time-Dynamic Estimates of the Reliability of Deep Semantic Segmentation Networks
abstract
In the semantic segmentation of street scenes with neural networks, the reliability of predictions is of highest interest. The assessment of neural networks by means of uncertainties is a common ansatz to prevent safety issues. As in applications like automated driving, video streams of images are available, we present a time-dynamic approach to investigating uncertainties and assessing the prediction quality of neural networks. We track segments over time and gather aggregated metrics per segment, thus obtaining time series of metrics from which we assess prediction quality. This is done by either classifying between intersection over union equal to 0 and greater than 0 or predicting the intersection over union directly. We study different models for these two tasks and analyze the influence of the time series length on the predictive power of our metrics.
Kira Maag, Matthias Rottmann, Hanno Gottschalk
ICTAI3
2020 Controlled False Negative Reduction of Minority Classes in Semantic Segmentation
abstract
In semantic segmentation datasets, classes of high importance are oftentimes underrepresented, e.g., humans in street scenes. Neural networks are usually trained to reduce the overall number of errors, attaching identical loss to errors of all kinds. However, this is not necessarily aligned with human intuition. For instance, an overlooked pedestrian seems more severe than an incorrectly detected one. One possible remedy is to deploy different decision rules by introducing class priors that assign more weight to underrepresented classes. While reducing the false negatives of the underrepresented class, at the same time this leads to a considerable increase of false positive indications. In this work, we combine decision rules with methods for false positive detection. Therefore, we fuse false negative detection with uncertainty based false positive meta classification. We present the efficiency of our method for the semantic segmentation of street scenes on the Cityscapes dataset based on predicted instances of the "human" class. In the latter we employ an advanced false positive detection method using uncertainty measures aggregated over instances. We, thereby, achieve improved trade-offs between false negative and false positive samples of the underrepresented classes.
Robin Chan 0001, Matthias Rottmann, Fabian Hüger, Peter Schlicht, Hanno Gottschalk
IJCNN5
2020 Prediction Error Meta Classification in Semantic Segmentation: Detection via Aggregated Dispersion Measures of Softmax Probabilities
abstract
We present a method that "meta" classifies whether segments predicted by a semantic segmentation neural network intersect with the ground truth. For this purpose, we employ measures of dispersion for predicted pixel-wise class probability distributions, like classification entropy, that yield heat maps of the input scene's size. We aggregate these dispersion measures segment-wise and derive metrics that are well correlated with the segment-wise IoU of prediction and ground truth. This procedure yields an almost plug and play post-processing tool to rate the prediction quality of semantic segmentation networks on segment level. This is especially relevant for monitoring neural networks in online applications like automated driving or medical imaging where reliability is of utmost importance. In our tests, we use publicly available state-of-the-art networks trained on the Cityscapes dataset and the BraTS2017 dataset and analyze the predictive power of different metrics as well as different sets of metrics. To this end, we compute logistic LASSO regression fits for the task of classifying IoU = 0 vs. IoU > 0 per segment and obtain AUROC values of up to 91.55%. We complement these tests with linear regression fits to predict the segment-wise IoU and obtain prediction standard deviations of down to 0.130 as well as R2values of up to 84.15%. We show that these results clearly outperform standard approaches.
Matthias Rottmann, Pascal Colling, Thomas-Paul Hack, Robin Chan 0001, Fabian Hüger, Peter Schlicht, Hanno Gottschalk
IJCNN7
2018 Deep Bayesian Active Semi-Supervised Learning
abstract
In many applications the process of generating label information is expensive and time consuming. We present a new method that combines active and semi-supervised deep learning to achieve high generalization performance from a deep convolutional neural network with as few known labels as possible. In a setting where a small amount of labeled data as well as a large amount of unlabeled data is available, our method first learns the labeled data set. This initialization is followed by an expectation maximization algorithm, where further training reduces classification entropy on the unlabeled data by targeting a low entropy fit which is consistent with the labeled data. In addition the algorithm asks at a specified frequency an oracle for labels of data with entropy above a certain entropy quantile. Using this active learning component we obtain an agile labeling process that achieves high accuracy, but requires only a small amount of known labels. For the MNIST dataset we report an error rate of 2.06% using only 300 labels and 1.06% for 1,000 labels. These results are obtained without employing any special network architecture or data augmentation.
Matthias Rottmann, Karsten Kahl, Hanno Gottschalk
ICMLA3