VLDB 2026 Research / reviewers in the wild / expert
Jan-Nico Zaech
dblp:217/2208 · also Jan-Nico Zäch
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-2566-0841ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsabstractDifferentiable simulators represent an environment’s dynamics as a differentiable function. Within robotics and autonomous driving, this property is used in Analytic Policy Gradients (APG), which relies on backpropagating through the dynamics to train accurate policies for diverse tasks. Here we show that differentiable simulation also has an important role in world modeling, where it can impart predictive, prescriptive, and counterfactual capabilities to an agent. Specifically, we design three novel task setups in which the differentiable dynamics are combined within an end-to-end computation graph not with a policy, but a state predictor. This allows us to learn relative odometry, optimal planners, and optimal inverse states. We collectively call these predictors Analytic World Models (AWMs) and demonstrate how differentiable simulation enables their efficient, end-to-end learning. In autonomous driving scenarios, they have broad applicability and can augment an agent’s decision-making beyond reactive control. Asen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza 0001, Luc Van Gool |
AAAI | 3 |
| 2026 | Autonomous Vehicle Path Planning by Searching with Differentiable SimulationabstractPlanning allows an agent to safely refine its actions before executing them in the real world. In autonomous driving, this is crucial to avoid collisions and navigate in complex, dense traffic scenarios. One way to plan is to search for the best action sequence. However, this is challenging when all necessary components – policy, next-state predictor, and critic – have to be learned. Here we propose Differentiable Simulation for Search (DSS), a framework that leverages the differentiable simulator Waymax as both a next state predictor and a critic. It relies on the simulator’s hardcoded dynamics, making state predictions highly accurate, while utilizing the simulator’s differentiability to effectively search across action sequences. Our DSS agent optimizes its actions using gradient descent over imagined future trajectories. We show experimentally that DSS – the combination of planning gradients and stochastic search – significantly improves tracking and path planning accuracy compared to sequence prediction, imitation learning, model-free RL, and other planning methods. Asen Nachkov, Jan-Nico Zaech, Danda Pani Paudel, Xi Wang 0021, Luc Van Gool |
AAAI | 2 |
| 2025 | Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
Anna-Maria Halacheva, Yang Miao 0004, Jan-Nico Zaech, Xi Wang 0021, Luc Van Gool, Danda Pani Paudel |
ICCV | 3 |
| 2025 | ReVLA: Reverting Visual Domain Limitation of Robotic Foundation ModelsabstractRecent progress in large language models and access to large-scale robotic datasets has sparked a paradigm shift in robotics models transforming them into generalists able to adapt to various tasks, scenes, and robot modalities. A large step for the community are open Vision Language Action models which showcase strong performance in a wide variety of tasks. In this work, we study the visual generalization capabilities of three existing robotic foundation models, and propose a corresponding evaluation framework. Our study shows that the existing models do not exhibit robustness to visual out-of-domain scenarios. This is potentially caused by limited variations in the training data and/or catastrophic forgetting, leading to domain limitations in the vision foundation models. We further explore OpenVLA, which uses two pre-trained vision foundation models and is, therefore, expected to generalize to out-of-domain experiments. However, we showcase catastrophic forgetting by DINO-v2 in OpenVLA through its failure to fulfill the task of depth regression. To overcome the aforementioned issue of visual catastrophic forgetting, we propose a gradual backbone reversal approach founded on model merging. This enables OpenVLA - which requires the adaptation of the visual backbones during initial training - to regain its visual generalization ability. Regaining this capability enables our ReVLA model to improve over OpenVLA by a factor of 77% and 66% for grasping and lifting in visual OOD tasks. Comprehensive evaluations, episode rollouts and model weights are available on the ReVLA Page Sombit Dey, Jan-Nico Zaech, Nikolay Nikolov, Luc Van Gool, Danda Pani Paudel |
ICRA | 2 |
| 2025 | LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationabstractWe propose LangHOPS, the first Multimodal Large Language Model (MLLM)-based framework for open-vocabulary object–part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from open-vocabulary candidate categories. Unlike prior approaches that rely on heuristic or learnable visual grouping, our approach grounds object–part hierarchies in language space. It integrates the MLLM into the object-part parsing pipeline to leverage rich knowledge and reasoning capabilities, and link multi-granularity concepts within the hierarchies. We evaluate LangHOPS across multiple challenging scenarios, including in-domain and cross-dataset object-part instance segmentation, and zero-shot semantic segmentation. LangHOPS achieves state-of-the-art results, surpassing previous methods by 5.5% Average Precision(AP) (in-domain) and 4.8% (cross-dataset) on the PartImageNet dataset and by 2.5% mIOU on unseen object parts in ADE20K (zero-shot). Ablation studies further validate the effectiveness of the language-grounded hierarchy and MLLM-driven part query refinement strategy. Yang Miao 0004, Jan-Nico Zaech, Xi Wang 0021, Fabien Despinoy, Danda Pani Paudel, Luc Van Gool |
NeurIPS | 2 |
| 2024 | Probabilistic Sampling of Balanced K-Means using Adiabatic Quantum ComputingabstractAdiabatic quantum computing (AQC) is a promising approach for discrete and often NP-hard optimization prob-lems. Current AQCs allow to implement problems of re-search interest, which has sparked the development of quan-tum representations for many computer vision tasks. De-spite requiring multiple measurements from the noisy AQC, current approaches only utilize the best measurement, dis-carding information contained in the remaining ones. In this work, we explore the potential of using this information for probabilistic balanced k-means clustering. Instead of discarding non-optimal solutions, we propose to use them to compute calibrated posterior probabilities with little ad-ditional compute cost. This allows us to identify ambiguous solutions and data points, which we demonstrate on a D-Wave AQC on synthetic tasks and real visual data. Jan-Nico Zaech, Martin Danelljan, Tolga Birdal, Luc Van Gool |
CVPR | 1 |
| 2024 | One-Shot Sparse Neural Architecture Search for Resource-Constrained DevicesabstractEmploying high-performance neural network models is challenging for resource-constrained devices as the models require strong computing power and large memory space. One-shot neural architecture search (NAS) helps to find a suitable neural architecture more efficiently by using a single one-shot network that shares weights with sub-networks. However, the existing one-shot NAS methods pay little attention to the sparsity achievable from the one-shot network. Therefore, there still may be room to reduce the resource requirement of the network. This work presents a new one-shot sparse NAS method, which tries to find an optimal sparsity for the network using soft channel masking during the architecture search. The experimental results show that the proposed method can find a more sparse architecture with little accuracy drop. Shenghui Song 0003, Jan-Nico Zaech, Seonyeong Heo |
RTCSA | 2 |
| 2024 | Optimizing Long-Term Robot Tracking with Multi-Platform Sensor FusionabstractMonitoring a fleet of robots requires stable long-term tracking with re-identification, which is yet an unsolved challenge in many scenarios. One application of this is the analysis of autonomous robotic soccer games at RoboCup. Tracking in these games requires handling of identically looking players, strong occlusions, and non-professional video recordings, but also offers state information estimated by the robots. In order to make effective use of the information coming from the robot sensors, we propose a robust tracking and identification pipeline. It fuses external non-calibrated camera data with the robots’ internal states using quadratic optimization for tracklet matching. The approach is validated using game recordings from previous RoboCup World Cup tournaments. Giuliano Albanese, Arka Mitra, Jan-Nico Zaech, Yupeng Zhao, Ajad Chhatkuli, Luc Van Gool |
WACV | 3 |
| 2022 | Adiabatic Quantum Computing for Multi Object TrackingabstractMulti-Object Tracking (MOT) is most often approached in the tracking-by-detection paradigm, where object detections are associated through time. The association step naturally leads to discrete optimization problems. As these optimization problems are often NP-hard, they can only be solved exactly for small instances on current hardware. Adiabatic quantum computing (AQC) offers a solution for this, as it has the potential to provide a considerable speedup on a range of NP-hard optimization problems in the near future. However, current MOT formulations are unsuitable for quantum computing due to their scaling properties. In this work, we therefore propose the first MOT formulation designed to be solved with AQC. We employ an Ising model that represents the quantum mechanical system implemented on the AQC. We show that our approach is competitive compared with state-of-the-art optimization-based approaches, even when using of-the-shelf integer programming solvers. Finally, we demonstrate that our MOT problem is already solvable on the current generation of real quantum computers for small examples, and analyze the properties of the measured solutions. Jan-Nico Zaech, Alexander Liniger, Martin Danelljan, Dengxin Dai, Luc Van Gool |
CVPR | 1 |
| 2022 | Unsupervised Robust Domain Adaptation without Source DataabstractWe study the problem of robust domain adaptation in the context of unavailable target labels and source data. The considered robustness is against adversarial perturbations. This paper aims at answering the question of finding the right strategy to make the target model robust and accurate in the setting of unsupervised domain adaptation without source data. The major findings of this paper are: (i) robust source models can be transferred robustly to the target; (ii) robust domain adaptation can greatly benefit from nonrobust pseudo-labels and the pair-wise contrastive loss. The proposed method of using non-robust pseudo-labels performs surprisingly well on both clean and adversarial samples, for the task of image classification. We show a consistent performance improvement of over 10% in accuracy against the tested baselines on four benchmark datasets. Our source code will be made publicly available. Peshal Agarwal, Danda Pani Paudel, Jan-Nico Zaech, Luc Van Gool |
WACV | 3 |
| 2021 | Decoder Fusion RNN: Context and Interaction Aware Decoders for Trajectory PredictionabstractForecasting the future behavior of all traffic agents in the vicinity is a key task to achieve safe and reliable autonomous driving systems. It is a challenging problem as agents adjust their behavior depending on their intentions, the others’ actions, and the road layout. In this paper, we propose Decoder Fusion RNN (DF-RNN), a recurrent, attention-based approach for motion forecasting. Our network is composed of a recurrent behavior encoder, an inter-agent multi-headed attention module, and a context-aware decoder. We design a map encoder that embeds polyline segments, combines them to create a graph structure, and merges their relevant parts with the agents’ embeddings. We fuse the encoded map information with further inter-agent interactions only inside the decoder and propose to use explicit training as a method to effectively utilize the information available. We demonstrate the efficacy of our method by testing it on the Argoverse motion forecasting dataset and show its state-of-the-art performance on the public benchmark. Edoardo Mello Rella, Jan-Nico Zaech, Alexander Liniger, Luc Van Gool |
IROS | 2 |
| 2020 | Action Sequence Predictions of Vehicles in Urban Environments using Map and Social ContextabstractThis work studies the problem of predicting the sequence of future actions for surrounding vehicles in real-world driving scenarios. To this aim, we make three main contributions. The first contribution is an automatic method to convert the trajectories recorded in real-world driving scenarios to action sequences with the help of HD maps. The method enables automatic dataset creation for this task from large-scale driving data. Our second contribution lies in applying the method to the well-known traffic agent tracking and prediction dataset Argoverse, resulting in 228,000 action sequences. Additionally, 2,245 action sequences were manually annotated for testing. The third contribution is to propose a novel action sequence prediction method by integrating past positions and velocities of the traffic agents, map information and social context into a single end-to-end trainable neural network. Our experiments prove the merit of the data creation method and the value of the created dataset - prediction performance improves consistently with the size of the dataset and shows that our action prediction method outperforms comparing models. Jan-Nico Zaech, Dengxin Dai, Alexander Liniger, Luc Van Gool |
IROS | 1 |
| 2019 | Learning to Avoid Poor Images: Towards Task-aware C-arm Cone-beam CT Trajectories
Jan-Nico Zaech, Cong Gao 0003, Bastian Bier, Russell H. Taylor, Andreas K. Maier, Nassir Navab, Mathias Unberath |
MICCAI (5) | 1 |
| 2018 | X-ray-transform Invariant Anatomical Landmark Detection for Pelvic Trauma Surgery
Bastian Bier, Mathias Unberath, Jan-Nico Zaech, Javad Fotouhi, Mehran Armand, Greg Osgood, Nassir Navab, Andreas K. Maier |
MICCAI (4) | 3 |
| 2018 | DeepDRR - A Catalyst for Machine Learning in Fluoroscopy-Guided Procedures
Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee, Bastian Bier, Javad Fotouhi, Mehran Armand, Nassir Navab |
MICCAI (4) | 2 |