VLDB 2026 Research / reviewers in the wild / expert
Volker Fischer 0003
dblp:84/4102-3
· DBLP profile ↗
16ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-5437-4030ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Trustworthy machine learning · 33% Deep learning architectures and training · 20% 3D vision · 8% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
shortcut learning |
1.1 | 2 | 2022 | Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022 DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021 |
Computer vision › Vision and language › vision-language model
contrastive vision-language model |
0.9 | 1 | 2025 | Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models · ICLR 2025 |
Machine learning › Representation and self-supervised learning › multimodal representation learning › cross-modal representation learning
modality gap |
0.9 | 1 | 2025 | Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.8 | 1 | 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.7 | 2 | 2019 | Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019 Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.6 | 2 | 2017 | On Detecting Adversarial Perturbations · ICLR (Poster) 2017 Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.6 | 1 | 2022 | Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022 |
Machine learning › Learning theory
generalization |
0.5 | 1 | 2021 | DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.5 | 1 | 2021 | DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021 |
Machine learning › Trustworthy machine learning › robustness
robustness to corruption |
0.5 | 1 | 2021 | Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021 |
Computer vision › Image recognition and object detection
shape bias |
0.5 | 1 | 2021 | Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.5 | 1 | 2021 | Fostering Generalization in Single-View 3D Reconstruction by Learning a Hierarchy of Local and Global Shape Priors · CVPR 2021 |
Machine learning › Deep learning architectures and training
equivariant neural network |
0.4 | 1 | 2020 | SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks · NeurIPS 2020 |
Computer vision › 3D vision › point cloud analysis
point cloud learning |
0.4 | 1 | 2020 | SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › language model interpretability
context attribution |
0.4 | 1 | 2019 | Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2019 | Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map |
0.4 | 1 | 2019 | Grid Saliency for Context Explanations of Semantic Segmentation · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › distributed training
model parallelism |
0.3 | 1 | 2018 | The streaming rollout of deep networks - towards fully model-parallel execution · NeurIPS 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | The streaming rollout of deep networks - towards fully model-parallel execution · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › robustness › adversarial examples
universal adversarial perturbation |
0.3 | 1 | 2017 | Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017 |
Natural language and speech › Language models and text generation
in-context learning |
0.2 | 1 | 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.2 | 1 | 2015 | Inverse reinforcement learning of behavioral models for online-adapting navigation strategies · ICRA 2015 |
Robotics › Motion planning and robot control
robot learning |
0.2 | 1 | 2015 | Inverse reinforcement learning of behavioral models for online-adapting navigation strategies · ICRA 2015 |
Robotics › Robot navigation and mapping
social navigation |
0.2 | 1 | 2015 | Inverse reinforcement learning of behavioral models for online-adapting navigation strategies · ICRA 2015 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.2 | 1 | 2022 | Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2021 | Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2021 | DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021 |
Computer vision › 3D vision › geometric prior
shape prior learning |
0.1 | 1 | 2021 | Fostering Generalization in Single-View 3D Reconstruction by Learning a Hierarchy of Local and Global Shape Priors · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
training dynamics analysis · 0.8synthetic task design · 0.8shape bias analysis · 0.5metrics · 0.5hierarchical shape priors · 0.5diagnostic benchmark · 0.5depth map learning · 0.5data augmentation · 0.5self-attention · 0.4equivariance · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language ModelsabstractContrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, like zero-shot object recognition, they perform surprisingly poor on other tasks, like attribute recognition. Previous work has attributed these challenges to the modality gap, a separation of image and text in the shared representation space, and to a bias towards objects over other factors, such as attributes. In this analysis paper, we investigate both phenomena thoroughly. We evaluated off-the-shelf VLMs and while the gap's influence on performance is typically overshadowed by other factors, we find indications that closing the gap indeed leads to improvements. Moreover, we find that, contrary to intuition, only few embedding dimensions drive the gap and that the embedding spaces are differently organized. To allow for a clean study of object bias, we introduce a definition and a corresponding measure of it. Equipped with this tool, we find that object bias does not lead to worse performance on other concepts, such as attributes per se. However, why do both phenomena, modality gap and object bias, emerge in the first place? To answer this fundamental question and uncover some of the inner workings of contrastive VLMs, we conducted experiments that allowed us to control the amount of shared information between the modalities. These experiments revealed that the driving factor behind both the modality gap and the object bias, is an information imbalance between images and captions, and unveiled an intriguing connection between the modality gap and entropy of the logits. Simon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer 0003, Thomas Brox |
ICLR | 4 |
| 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization ProblemsabstractIn this work, we study rapid improvements of the training loss in transformers when being confronted with multi-step decision tasks. We found that transformers struggle to learn the intermediate task and both training and validation loss saturate for hundreds of epochs. When transformers finally learn the intermediate task, they do this rapidly and unexpectedly. We call these abrupt improvements Eureka-moments, since the transformer appears to suddenly learn a previously incomprehensible concept. We designed synthetic tasks to study the problem in detail, but the leaps in performance can be observed also for language modeling and in-context learning (ICL). We suspect that these abrupt transitions are caused by the multi-step nature of these tasks. Indeed, we find connections and show that ways to improve on the synthetic multi-step tasks can be used to improve the training of language modeling and ICL. Using the synthetic data we trace the problem back to the Softmax function in the self-attention block of transformers and show ways to alleviate the problem. These fixes reduce the required number of training steps, lead to higher likelihood to learn the intermediate task, to higher final accuracy and training becomes more robust to hyper-parameters. David T. Hoffmann, Simon Schrodi, Jelena Bratulic, Nadine Behrmann, Volker Fischer 0003, Thomas Brox |
ICML | 5 |
| 2022 | Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain
Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Claudia Blaiotta, Mauricio Munoz, Volker Fischer 0003 |
ECCV (25) | 5 |
| 2021 | Fostering Generalization in Single-View 3D Reconstruction by Learning a Hierarchy of Local and Global Shape PriorsabstractSingle-view 3D object reconstruction has seen much progress, yet methods still struggle generalizing to novel shapes unseen during training. Common approaches pre- dominantly rely on learned global shape priors and, hence, disregard detailed local observations. In this work, we address this issue by learning a hierarchy of priors at different levels of locality from ground truth input depth maps. We argue that exploiting local priors allows our method to efficiently use input observations, thus improving generalization in visible areas of novel shapes. At the same time, the combination of local and global priors enables meaningful hallucination of unobserved parts resulting in consistent 3D shapes. We show that the hierarchical approach generalizes much better than the global approach. It generalizes not only between different instances of a class but also across classes and to unseen arrangements of objects. Jan Bechtold, Maxim Tatarchenko, Volker Fischer 0003, Thomas Brox |
CVPR | 3 |
| 2021 | DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization OpportunitiesabstractCommon deep neural networks (DNNs) for image classification have been shown to rely on shortcut opportunities (SO) in the form of predictive and easy-to-represent visual factors. This is known as shortcut learning and leads to impaired generalization. In this work, we show that common DNNs also suffer from shortcut learning when predicting only basic visual object factors of variation (FoV) such as shape, color, or texture. We argue that besides shortcut opportunities, generalization opportunities (GO) are also an inherent part of real-world vision data and arise from partial independence between predicted classes and FoVs. We also argue that it is necessary for DNNs to exploit GO to overcome shortcut learning. Our core contribution is to introduce the Diagnostic Vision Benchmark suite DiagViB-6, which includes datasets and metrics to study a network’s shortcut vulnerability and generalization capability for six independent FoV. In particular, DiagViB-6 allows controlling the type and degree of SO and GO in a dataset. We benchmark a wide range of popular vision architectures and show that they can exploit GO only to a limited extent. Elias Eulig, Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Kilian Rambach, William Beluch, Xiahan Shi, Volker Fischer 0003 |
ICCV | 7 |
| 2021 | Does enhanced shape bias improve neural network robustness to common corruptions?
Chaithanya Kumar Mummadi, Ranjitha Subramaniam, Robin Hutmacher, Julien Vitay, Volker Fischer 0003, Jan Hendrik Metzen |
ICLR | 5 |
| 2020 | SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksabstractWe introduce the SE(3)-Transformer, a variant of the self-attention module for 3D point-clouds, which is equivariant under continuous 3D roto-translations. Equivariance is important to ensure stable and predictable performance in the presence of nuisance transformations of the data input. A positive corollary of equivariance is increased weight-tying within the model. The SE(3)-Transformer leverages the benefits of self-attention to operate on large point clouds with varying number of points, while guaranteeing SE(3)-equivariance for robustness. We evaluate our model on a toy N-body particle simulation dataset, showcasing the robustness of the predictions under rotations of the input. We further achieve competitive performance on two real-world datasets, ScanObjectNN and QM9. In all cases, our model outperforms a strong, non-equivariant attention baseline and an equivariant model without attention. Fabian Fuchs, Daniel E. Worrall, Volker Fischer 0003, Max Welling |
NeurIPS | 3 |
| 2019 | Grid Saliency for Context Explanations of Semantic SegmentationabstractRecently, there has been a growing interest in developing saliency methods that provide visual explanations of network predictions. Still, the usability of existing methods is limited to image classification models. To overcome this limitation, we extend the existing approaches to generate grid saliencies, which provide spatially coherent visual explanations for (pixel-level) dense prediction networks. As the proposed grid saliency allows to spatially disentangle the object and its context, we specifically explore its potential to produce context explanations for semantic segmentation networks, discovering which context most influences the class predictions inside a target object area. We investigate the effectiveness of grid saliency on a synthetic dataset with an artificially induced bias between objects and their context as well as on the real-world Cityscapes dataset using state-of-the-art segmentation networks. Our results show that grid saliency can be successfully used to provide easily interpretable context explanations and, moreover, can be employed for detecting and localizing contextual biases present in the data. Lukas Hoyer, Mauricio Munoz, Prateek Katiyar, Anna Khoreva, Volker Fischer 0003 |
NeurIPS | 5 |
| 2018 | Functionally Modular and Interpretable Temporal Filtering for Robust Segmentation
Volker Fischer 0003, Michael Herman, Sven Behnke |
BMVC | 2 |
| 2018 | Hierarchical Recurrent Filtering for Fully Convolutional DenseNets
Volker Fischer 0003, Michael Herman, Sven Behnke |
ESANN | 2 |
| 2018 | The streaming rollout of deep networks - towards fully model-parallel executionabstractDeep neural networks, and in particular recurrent networks, are promising candidates to control autonomous agents that interact in real-time with the physical world. However, this requires a seamless integration of temporal features into the network’s architecture. For the training of and inference with recurrent neural networks, they are usually rolled out over time, and different rollouts exist. Conventionally during inference, the layers of a network are computed in a sequential manner resulting in sparse temporal integration of information and long response times. In this study, we present a theoretical framework to describe rollouts, the level of model-parallelization they induce, and demonstrate differences in solving specific tasks. We prove that certain rollouts, also for networks with only skip and no recurrent connections, enable earlier and more frequent responses, and show empirically that these early responses have better performance. The streaming rollout maximizes these properties and enables a fully parallel execution of the network reducing runtime on massively parallel devices. Finally, we provide an open-source toolbox to design, train, evaluate, and interact with streaming rollouts. Volker Fischer 0003, Jan Köhler, Thomas Pfeil |
NeurIPS | 1 |
| 2017 | Learning Semantic Prediction using Pretrained Deep Feedforward Networks
Volker Fischer 0003, Michael Herman, Sven Behnke |
ESANN | 2 |
| 2017 | Universal Adversarial Perturbations Against Semantic Image SegmentationabstractWhile deep learning is remarkably successful on perceptual tasks, it was also shown to be vulnerable to adversarial perturbations of the input. These perturbations denote noise added to the input that was generated specifically to fool the system while being quasi-imperceptible for humans. More severely, there even exist universal perturbations that are input-agnostic but fool the network on the majority of inputs. While recent work has focused on image classification, this work proposes attacks against semantic image segmentation: we present an approach for generating (universal) adversarial perturbations that make the network yield a desired target segmentation as output. We show empirically that there exist barely perceptible universal noise patterns which result in nearly the same predicted segmentation for arbitrary inputs. Furthermore, we also show the existence of universal noise which removes a target class (e.g., all pedestrians) from the segmentation while leaving the segmentation mostly unchanged otherwise. Jan Hendrik Metzen, Chaithanya Kumar Mummadi, Thomas Brox, Volker Fischer 0003 |
ICCV | 4 |
| 2017 | On Detecting Adversarial Perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer 0003, Bastian Bischoff |
ICLR (Poster) | 3 |
| 2016 | Multispectral Pedestrian Detection using Deep Fusion Convolutional Neural Networks
Volker Fischer 0003, Michael Herman, Sven Behnke |
ESANN | 2 |
| 2015 | Inverse reinforcement learning of behavioral models for online-adapting navigation strategiesabstractTo increase the acceptance of autonomous systems in populated environments, it is indispensable to teach them social behavior. We would expect a social robot, which plans its motions among humans, to consider both the social acceptability of its behavior as well as task constraints, such as time limits. These requirements are often contradictory and therefore resulting in a trade-off. For example, a robot has to decide whether it is more important to quickly achieve its goal or to comply with social conventions, such as the proximity to humans, i.e., the robot has to react adaptively to task-specific priorities. In this paper, we present a method for priority-adaptive navigation of mobile autonomous systems, which optimizes the social acceptability of the behavior while meeting task constraints. We learn acceptability-dependent behavioral models from human demonstrations by using maximum entropy (MaxEnt) inverse reinforcement learning (IRL). These models are generative and describe the learned stochastic behavior. We choose the optimum behavioral model by maximizing the social acceptability under constraints on expected time-limits and reliabilities. This approach is evaluated in the context of driving behaviors based on the highway scenario of Levine et al. [1]. Michael Herman, Volker Fischer 0003, Tobias Gindele, Wolfram Burgard |
ICRA | 2 |