EDBT 2026 Demo / reviewers in the wild / expert
David T. Hoffmann
dblp:246/5191
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 30% Language models and text generation · 19% Representation and self-supervised learning · 17% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
in-context learning |
1.2 | 2 | 2026 | Unlocking In-Context Learning for Natural Datasets Across Modalities · Int. J. Comput. Vis. 2026 Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024 |
Natural language and speech › Language models and text generation › in-context learning
emergent in-context learning |
1.0 | 1 | 2026 | Unlocking In-Context Learning for Natural Datasets Across Modalities · Int. J. Comput. Vis. 2026 |
Computer vision › Vision and language › vision-language model
contrastive vision-language model |
0.9 | 1 | 2025 | Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models · ICLR 2025 |
Machine learning › Representation and self-supervised learning › multimodal representation learning › cross-modal representation learning
modality gap |
0.9 | 1 | 2025 | Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models · ICLR 2025 |
Computer vision › 3D vision
scene flow estimation |
0.9 | 1 | 2025 | Floxels: Fast Unsupervised Voxel Based Scene Flow Estimation · CVPR 2025 |
Computer vision › 3D vision › scene flow estimation
self-supervised scene flow |
0.9 | 1 | 2025 | Floxels: Fast Unsupervised Voxel Based Scene Flow Estimation · CVPR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.8 | 1 | 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems · ICML 2024 |
Computer vision › 3D vision
human body modeling |
0.6 | 2 | 2021 | AGORA: Avatars in Geography Optimized for Regression Analysis · CVPR 2021 Learning Multi-human Optical Flow · Int. J. Comput. Vis. 2020 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives · AAAI 2022 |
Machine learning › Representation and self-supervised learning › contrastive learning › contrastive loss
InfoNCE |
0.6 | 1 | 2022 | Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives · AAAI 2022 |
Computer vision › 3D vision
3d human pose estimation |
0.5 | 1 | 2021 | AGORA: Avatars in Geography Optimized for Regression Analysis · CVPR 2021 |
Computer vision › Face, body and person analysis
human pose estimation |
0.5 | 1 | 2021 | AGORA: Avatars in Geography Optimized for Regression Analysis · CVPR 2021 |
Computer vision › Video understanding and tracking
motion analysis |
0.4 | 1 | 2020 | Learning Multi-human Optical Flow · Int. J. Comput. Vis. 2020 |
Computer vision › 3D vision › motion estimation
optical flow |
0.4 | 1 | 2020 | Learning Multi-human Optical Flow · Int. J. Comput. Vis. 2020 |
Machine learning › Deep learning architectures and training
training dynamics |
0.3 | 1 | 2026 | Unlocking In-Context Learning for Natural Datasets Across Modalities · Int. J. Comput. Vis. 2026 |
Computer vision › 3D vision › 3d scene modeling › scene representation
voxel-based scene representation |
0.3 | 1 | 2025 | Floxels: Fast Unsupervised Voxel Based Scene Flow Estimation · CVPR 2025 |
Machine learning › Learning theory
classification |
0.2 | 1 | 2022 | Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives · AAAI 2022 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.2 | 1 | 2022 | Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
autoregressive model · 1.0voxel grid model · 0.9test-time optimization · 0.9multi-frame loss · 0.9training dynamics analysis · 0.8synthetic task design · 0.8contrastive learning · 0.6InfoNCE · 0.6fine-tuning · 0.5body model fitting · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlocking In-Context Learning for Natural Datasets Across ModalitiesabstractAbstract Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updating the model’s weights. While ICL offers fast adaptation across natural language tasks and domains, its emergence is less straightforward for modalities beyond text. In this work, we systematically uncover properties present in LLMs that support the emergence of ICL for autoregressive models and various modalities by promoting the learning of the mechanisms needed for ICL. We identify exact token repetitions in the training data sequences as an important factor for ICL. Such repetitions further improve stability and reduce transiency in ICL performance. We analyse in detail the training dynamics of such data sequences and explain how token repetitions enhance the ICL learning mechanisms. Moreover, we emphasise the importance of the training task difficulty for the emergence of ICL. Finally, by applying our novel insights on ICL emergence, we unlock ICL capabilities across various visual datasets used for few-shot classification, and confirm the generalisability of our insights to much harder real-world examples of large-scale object classification, and a more challenging EEG classification task. Code is available at https://github.com/jelenab98/unlocking_icl Jelena Bratulic, Sudhanshu Mittal, David T. Hoffmann, Samuel Böhm, Robin Schirrmeister, Tonio Ball, Christian Rupprecht 0001, Thomas Brox |
Int. J. Comput. Vis. | 3 |
| 2025 | Floxels: Fast Unsupervised Voxel Based Scene Flow EstimationabstractScene flow estimation is a foundational task for many robotic applications, including robust dynamic object detection, automatic labeling, and sensor synchronization. Two types of approaches to the problem have evolved: 1) Supervised and 2) optimization-based methods. Supervised methods are fast during inference and achieve high-quality results, however, they are limited by the need for large amounts of labeled training data and are susceptible to domain gaps. In contrast, unsupervised test-time optimization methods do not face the problem of domain gaps but usually suffer from substantial runtime, exhibit artifacts, or fail to converge to the right solution. In this work, we mitigate several limitations of existing optimization-based methods. To this end, we 1) introduce a simple voxel grid-based model that improves over the standard MLP-based formulation in multiple dimensions and 2) introduce a new multi-frame loss formulation. 3) We combine both contributions in our new method, termed Floxels. On the Argoverse 2 benchmark, Floxels is surpassed only by EulerFlow among unsupervised methods while achieving comparable performance at a fraction of the computational cost. Floxels achieves a massive speedup of more than ~60 – 140× over Euler-Flow, reducing the runtime from a day to 10× minutes per sequence. Over the faster but low-quality baseline, NSFP, Floxels achieves a speedup of ∼14×. David T. Hoffmann, Syed Haseeb Raza, Hanqiu Jiang, Denis Tananaev, Steffen Klingenhoefer, Martin Meinke |
CVPR | 1 |
| 2025 | Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language ModelsabstractContrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, like zero-shot object recognition, they perform surprisingly poor on other tasks, like attribute recognition. Previous work has attributed these challenges to the modality gap, a separation of image and text in the shared representation space, and to a bias towards objects over other factors, such as attributes. In this analysis paper, we investigate both phenomena thoroughly. We evaluated off-the-shelf VLMs and while the gap's influence on performance is typically overshadowed by other factors, we find indications that closing the gap indeed leads to improvements. Moreover, we find that, contrary to intuition, only few embedding dimensions drive the gap and that the embedding spaces are differently organized. To allow for a clean study of object bias, we introduce a definition and a corresponding measure of it. Equipped with this tool, we find that object bias does not lead to worse performance on other concepts, such as attributes per se. However, why do both phenomena, modality gap and object bias, emerge in the first place? To answer this fundamental question and uncover some of the inner workings of contrastive VLMs, we conducted experiments that allowed us to control the amount of shared information between the modalities. These experiments revealed that the driving factor behind both the modality gap and the object bias, is an information imbalance between images and captions, and unveiled an intriguing connection between the modality gap and entropy of the logits. Simon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer 0003, Thomas Brox |
ICLR | 2 |
| 2024 | Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization ProblemsabstractIn this work, we study rapid improvements of the training loss in transformers when being confronted with multi-step decision tasks. We found that transformers struggle to learn the intermediate task and both training and validation loss saturate for hundreds of epochs. When transformers finally learn the intermediate task, they do this rapidly and unexpectedly. We call these abrupt improvements Eureka-moments, since the transformer appears to suddenly learn a previously incomprehensible concept. We designed synthetic tasks to study the problem in detail, but the leaps in performance can be observed also for language modeling and in-context learning (ICL). We suspect that these abrupt transitions are caused by the multi-step nature of these tasks. Indeed, we find connections and show that ways to improve on the synthetic multi-step tasks can be used to improve the training of language modeling and ICL. Using the synthetic data we trace the problem back to the Softmax function in the self-attention block of transformers and show ways to alleviate the problem. These fixes reduce the required number of training steps, lead to higher likelihood to learn the intermediate task, to higher final accuracy and training becomes more robust to hyper-parameters. David T. Hoffmann, Simon Schrodi, Jelena Bratulic, Nadine Behrmann, Volker Fischer 0003, Thomas Brox |
ICML | 1 |
| 2022 | Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked PositivesabstractThis paper introduces Ranking Info Noise Contrastive Estimation (RINCE), a new member in the family of InfoNCE losses that preserves a ranked ordering of positive samples. In contrast to the standard InfoNCE loss, which requires a strict binary separation of the training pairs into similar and dissimilar samples, RINCE can exploit information about a similarity ranking for learning a corresponding embedding space. We show that the proposed loss function learns favorable embeddings compared to the standard InfoNCE whenever at least noisy ranking information can be obtained or when the definition of positives and negatives is blurry. We demonstrate this for a supervised classification task with additional superclass labels and noisy similarity scores. Furthermore, we show that RINCE can also be applied to unsupervised training with experiments on unsupervised representation learning from videos. In particular, the embedding yields higher classification accuracy, retrieval rates and performs better on out-of-distribution detection than the standard InfoNCE loss. David T. Hoffmann, Nadine Behrmann, Juergen Gall, Thomas Brox, Mehdi Noroozi |
AAAI | 1 |
| 2021 | AGORA: Avatars in Geography Optimized for Regression AnalysisabstractWhile the accuracy of 3D human pose estimation from images has steadily improved on benchmark datasets, the best methods still fail in many real-world scenarios. This suggests that there is a domain gap between current datasets and common scenes containing people. To obtain ground-truth 3D pose, current datasets limit the complexity of clothing, environmental conditions, number of subjects, and occlusion. Moreover, current datasets evaluate sparse 3D joint locations corresponding to the major joints of the body, ignoring the hand pose and the face shape. To evaluate the current state-of-the-art methods on more challenging images, and to drive the field to address new problems, we introduce AGORA, a synthetic dataset with high realism and highly accurate ground truth. Here we use 4240 commercially-available, high-quality, textured human scans in diverse poses and natural clothing; this includes 257 scans of children. We create reference 3D poses and body shapes by fitting the SMPL-X body model (with face and hands) to the 3D scans, taking into account clothing. We create around 14K training and 3K test images by rendering between 5 and 15 people per image using either image-based lighting or rendered 3D environments, taking care to make the images physically plausible and photoreal. In total, AGORA consists of 173K individual person crops. We evaluate existing state-of-the-art methods for 3D human pose estimation on this dataset. and find that most methods perform poorly on images of children. Hence, we extend the SMPL-X model to better capture the shape of children. Additionally, we fine-tune methods on AGORA and show improved performance on both AGORA and 3DPW, confirming the realism of the dataset. We provide all the registered 3D reference training data, rendered images, and a web-based evaluation site at https://agora.is.tue.mpg.de/. Priyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann, Shashank Tripathi, Michael J. Black |
CVPR | 4 |
| 2020 | Learning Multi-human Optical FlowabstractAbstract The optical flow of humans is well known to be useful for the analysis of human action. Recent optical flow methods focus on training deep networks to approach the problem. However, the training data used by them does not cover the domain of human motion. Therefore, we develop a dataset of multi-human optical flow and train optical flow networks on this dataset. We use a 3D model of the human body and motion capture data to synthesize realistic flow fields in both single- and multi-person images. We then train optical flow networks to estimate human flow fields from pairs of images. We demonstrate that our trained networks are more accurate than a wide range of top methods on held-out test data and that they can generalize well to real image sequences. The code, trained models and the dataset are available for research. Anurag Ranjan, David T. Hoffmann, Dimitrios Tzionas, Siyu Tang 0001, Javier Romero 0002, Michael J. Black |
Int. J. Comput. Vis. | 2 |