Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Volker Fischer 0003

dblp:84/4102-3 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-5437-4030ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Trustworthy machine learning · 33% Deep learning architectures and training · 20% 3D vision · 8%

Topics — the 30 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
shortcut learning
1.122022
Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Computer vision › Vision and language › vision-language model
contrastive vision-language model
0.912025
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models · ICLR 2025
Machine learning › Representation and self-supervised learning › multimodal representation learning › cross-modal representation learning
modality gap
0.912025
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models · ICLR 2025
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.812024
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.722019
Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019
Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.622017
On Detecting Adversarial Perturbations · ICLR (Poster) 2017
Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017
Machine learning › Transfer learning and domain adaptation
domain generalization
0.612022
Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022
Machine learning › Learning theory
generalization
0.512021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.512021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Trustworthy machine learning › robustness
robustness to corruption
0.512021
Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021
Computer vision › Image recognition and object detection
shape bias
0.512021
Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.512021
Fostering Generalization in Single-View 3D Reconstruction by Learning a Hierarchy of Local and Global Shape Priors · CVPR 2021
Machine learning › Deep learning architectures and training
equivariant neural network
0.412020
SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks · NeurIPS 2020
Computer vision › 3D vision › point cloud analysis
point cloud learning
0.412020
SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks · NeurIPS 2020
Machine learning › Trustworthy machine learning › language model interpretability
context attribution
0.412019
Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019
Machine learning › Trustworthy machine learning
interpretability
0.412019
Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map
0.412019
Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.312018
The streaming rollout of deep networks - towards fully model-parallel execution · NeurIPS 2018
Machine learning › Deep learning architectures and training
recurrent neural network
0.312018
The streaming rollout of deep networks - towards fully model-parallel execution · NeurIPS 2018
Machine learning › Trustworthy machine learning › robustness › adversarial examples
universal adversarial perturbation
0.312017
Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017
Natural language and speech › Language models and text generation
in-context learning
0.212024
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.212015
Inverse reinforcement learning of behavioral models for online-adapting navigation strategies · ICRA 2015
Robotics › Motion planning and robot control
robot learning
0.212015
Inverse reinforcement learning of behavioral models for online-adapting navigation strategies · ICRA 2015
Robotics › Robot navigation and mapping
social navigation
0.212015
Inverse reinforcement learning of behavioral models for online-adapting navigation strategies · ICRA 2015
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.212022
Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022
Machine learning › Deep learning architectures and training
convolutional neural network
0.112021
Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021
Computer vision › Image recognition and object detection
image classification
0.112021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Computer vision › 3D vision › geometric prior
shape prior learning
0.112021
Fostering Generalization in Single-View 3D Reconstruction by Learning a Hierarchy of Local and Global Shape Priors · CVPR 2021

Methods — techniques the papers use, named apart from their topics

training dynamics analysis · 0.8synthetic task design · 0.8shape bias analysis · 0.5metrics · 0.5hierarchical shape priors · 0.5diagnostic benchmark · 0.5depth map learning · 0.5data augmentation · 0.5self-attention · 0.4equivariance · 0.4
YearPublicationVenuePosition
2025 Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
abstract
Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, like zero-shot object recognition, they perform surprisingly poor on other tasks, like attribute recognition. Previous work has attributed these challenges to the modality gap, a separation of image and text in the shared representation space, and to a bias towards objects over other factors, such as attributes. In this analysis paper, we investigate both phenomena thoroughly. We evaluated off-the-shelf VLMs and while the gap's influence on performance is typically overshadowed by other factors, we find indications that closing the gap indeed leads to improvements. Moreover, we find that, contrary to intuition, only few embedding dimensions drive the gap and that the embedding spaces are differently organized. To allow for a clean study of object bias, we introduce a definition and a corresponding measure of it. Equipped with this tool, we find that object bias does not lead to worse performance on other concepts, such as attributes per se. However, why do both phenomena, modality gap and object bias, emerge in the first place? To answer this fundamental question and uncover some of the inner workings of contrastive VLMs, we conducted experiments that allowed us to control the amount of shared information between the modalities. These experiments revealed that the driving factor behind both the modality gap and the object bias, is an information imbalance between images and captions, and unveiled an intriguing connection between the modality gap and entropy of the logits.
Simon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer 0003, Thomas Brox
ICLR4
2024 Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
abstract
In this work, we study rapid improvements of the training loss in transformers when being confronted with multi-step decision tasks. We found that transformers struggle to learn the intermediate task and both training and validation loss saturate for hundreds of epochs. When transformers finally learn the intermediate task, they do this rapidly and unexpectedly. We call these abrupt improvements Eureka-moments, since the transformer appears to suddenly learn a previously incomprehensible concept. We designed synthetic tasks to study the problem in detail, but the leaps in performance can be observed also for language modeling and in-context learning (ICL). We suspect that these abrupt transitions are caused by the multi-step nature of these tasks. Indeed, we find connections and show that ways to improve on the synthetic multi-step tasks can be used to improve the training of language modeling and ICL. Using the synthetic data we trace the problem back to the Softmax function in the self-attention block of transformers and show ways to alleviate the problem. These fixes reduce the required number of training steps, lead to higher likelihood to learn the intermediate task, to higher final accuracy and training becomes more robust to hyper-parameters.
David T. Hoffmann, Simon Schrodi, Jelena Bratulic, Nadine Behrmann, Volker Fischer 0003, Thomas Brox
ICML5
2022 Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain
Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Claudia Blaiotta, Mauricio Munoz, Volker Fischer 0003
ECCV (25)5
2021 Fostering Generalization in Single-View 3D Reconstruction by Learning a Hierarchy of Local and Global Shape Priors
abstract
Single-view 3D object reconstruction has seen much progress, yet methods still struggle generalizing to novel shapes unseen during training. Common approaches pre- dominantly rely on learned global shape priors and, hence, disregard detailed local observations. In this work, we address this issue by learning a hierarchy of priors at different levels of locality from ground truth input depth maps. We argue that exploiting local priors allows our method to efficiently use input observations, thus improving generalization in visible areas of novel shapes. At the same time, the combination of local and global priors enables meaningful hallucination of unobserved parts resulting in consistent 3D shapes. We show that the hierarchical approach generalizes much better than the global approach. It generalizes not only between different instances of a class but also across classes and to unseen arrangements of objects.
Jan Bechtold, Maxim Tatarchenko, Volker Fischer 0003, Thomas Brox
CVPR3
2021 DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities
abstract
Common deep neural networks (DNNs) for image classification have been shown to rely on shortcut opportunities (SO) in the form of predictive and easy-to-represent visual factors. This is known as shortcut learning and leads to impaired generalization. In this work, we show that common DNNs also suffer from shortcut learning when predicting only basic visual object factors of variation (FoV) such as shape, color, or texture. We argue that besides shortcut opportunities, generalization opportunities (GO) are also an inherent part of real-world vision data and arise from partial independence between predicted classes and FoVs. We also argue that it is necessary for DNNs to exploit GO to overcome shortcut learning. Our core contribution is to introduce the Diagnostic Vision Benchmark suite DiagViB-6, which includes datasets and metrics to study a network’s shortcut vulnerability and generalization capability for six independent FoV. In particular, DiagViB-6 allows controlling the type and degree of SO and GO in a dataset. We benchmark a wide range of popular vision architectures and show that they can exploit GO only to a limited extent.
Elias Eulig, Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Kilian Rambach, William Beluch, Xiahan Shi, Volker Fischer 0003
ICCV7
2021 Does enhanced shape bias improve neural network robustness to common corruptions?
Chaithanya Kumar Mummadi, Ranjitha Subramaniam, Robin Hutmacher, Julien Vitay, Volker Fischer 0003, Jan Hendrik Metzen
ICLR5
2020 SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks
abstract
We introduce the SE(3)-Transformer, a variant of the self-attention module for 3D point-clouds, which is equivariant under continuous 3D roto-translations. Equivariance is important to ensure stable and predictable performance in the presence of nuisance transformations of the data input. A positive corollary of equivariance is increased weight-tying within the model. The SE(3)-Transformer leverages the benefits of self-attention to operate on large point clouds with varying number of points, while guaranteeing SE(3)-equivariance for robustness. We evaluate our model on a toy N-body particle simulation dataset, showcasing the robustness of the predictions under rotations of the input. We further achieve competitive performance on two real-world datasets, ScanObjectNN and QM9. In all cases, our model outperforms a strong, non-equivariant attention baseline and an equivariant model without attention.
Fabian Fuchs, Daniel E. Worrall, Volker Fischer 0003, Max Welling
NeurIPS3
2019 Grid Saliency for Context Explanations of Semantic Segmentation
abstract
Recently, there has been a growing interest in developing saliency methods that provide visual explanations of network predictions. Still, the usability of existing methods is limited to image classification models. To overcome this limitation, we extend the existing approaches to generate grid saliencies, which provide spatially coherent visual explanations for (pixel-level) dense prediction networks. As the proposed grid saliency allows to spatially disentangle the object and its context, we specifically explore its potential to produce context explanations for semantic segmentation networks, discovering which context most influences the class predictions inside a target object area. We investigate the effectiveness of grid saliency on a synthetic dataset with an artificially induced bias between objects and their context as well as on the real-world Cityscapes dataset using state-of-the-art segmentation networks. Our results show that grid saliency can be successfully used to provide easily interpretable context explanations and, moreover, can be employed for detecting and localizing contextual biases present in the data.
Lukas Hoyer, Mauricio Munoz, Prateek Katiyar, Anna Khoreva, Volker Fischer 0003
NeurIPS5
2018 Functionally Modular and Interpretable Temporal Filtering for Robust Segmentation
Volker Fischer 0003, Michael Herman, Sven Behnke
BMVC2
2018 Hierarchical Recurrent Filtering for Fully Convolutional DenseNets
Volker Fischer 0003, Michael Herman, Sven Behnke
ESANN2
2018 The streaming rollout of deep networks - towards fully model-parallel execution
abstract
Deep neural networks, and in particular recurrent networks, are promising candidates to control autonomous agents that interact in real-time with the physical world. However, this requires a seamless integration of temporal features into the network’s architecture. For the training of and inference with recurrent neural networks, they are usually rolled out over time, and different rollouts exist. Conventionally during inference, the layers of a network are computed in a sequential manner resulting in sparse temporal integration of information and long response times. In this study, we present a theoretical framework to describe rollouts, the level of model-parallelization they induce, and demonstrate differences in solving specific tasks. We prove that certain rollouts, also for networks with only skip and no recurrent connections, enable earlier and more frequent responses, and show empirically that these early responses have better performance. The streaming rollout maximizes these properties and enables a fully parallel execution of the network reducing runtime on massively parallel devices. Finally, we provide an open-source toolbox to design, train, evaluate, and interact with streaming rollouts.
Volker Fischer 0003, Jan Köhler, Thomas Pfeil
NeurIPS1
2017 Learning Semantic Prediction using Pretrained Deep Feedforward Networks
Volker Fischer 0003, Michael Herman, Sven Behnke
ESANN2
2017 Universal Adversarial Perturbations Against Semantic Image Segmentation
abstract
While deep learning is remarkably successful on perceptual tasks, it was also shown to be vulnerable to adversarial perturbations of the input. These perturbations denote noise added to the input that was generated specifically to fool the system while being quasi-imperceptible for humans. More severely, there even exist universal perturbations that are input-agnostic but fool the network on the majority of inputs. While recent work has focused on image classification, this work proposes attacks against semantic image segmentation: we present an approach for generating (universal) adversarial perturbations that make the network yield a desired target segmentation as output. We show empirically that there exist barely perceptible universal noise patterns which result in nearly the same predicted segmentation for arbitrary inputs. Furthermore, we also show the existence of universal noise which removes a target class (e.g., all pedestrians) from the segmentation while leaving the segmentation mostly unchanged otherwise.
Jan Hendrik Metzen, Chaithanya Kumar Mummadi, Thomas Brox, Volker Fischer 0003
ICCV4
2017 On Detecting Adversarial Perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer 0003, Bastian Bischoff
ICLR (Poster)3
2016 Multispectral Pedestrian Detection using Deep Fusion Convolutional Neural Networks
Volker Fischer 0003, Michael Herman, Sven Behnke
ESANN2
2015 Inverse reinforcement learning of behavioral models for online-adapting navigation strategies
abstract
To increase the acceptance of autonomous systems in populated environments, it is indispensable to teach them social behavior. We would expect a social robot, which plans its motions among humans, to consider both the social acceptability of its behavior as well as task constraints, such as time limits. These requirements are often contradictory and therefore resulting in a trade-off. For example, a robot has to decide whether it is more important to quickly achieve its goal or to comply with social conventions, such as the proximity to humans, i.e., the robot has to react adaptively to task-specific priorities. In this paper, we present a method for priority-adaptive navigation of mobile autonomous systems, which optimizes the social acceptability of the behavior while meeting task constraints. We learn acceptability-dependent behavioral models from human demonstrations by using maximum entropy (MaxEnt) inverse reinforcement learning (IRL). These models are generative and describe the learned stochastic behavior. We choose the optimum behavioral model by maximizing the social acceptability under constraints on expected time-limits and reliabilities. This approach is evaluated in the context of driving behaviors based on the highway scenario of Levine et al. [1].
Michael Herman, Volker Fischer 0003, Tobias Gindele, Wolfram Burgard
ICRA2