EDBT 2026 Demo / reviewers in the wild / expert
Zuria Bauer
dblp:228/1780
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-8447-2344ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 39% Robot navigation and mapping · 23% Generative modeling · 19% | |
| Computer graphics and multimedia
2 papers |
Rendering · 85% Virtual and augmented reality · 15% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot navigation and mapping › visual odometry
self-supervised visual odometry |
1.0 | 1 | 2026 | Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame · Int. J. Comput. Vis. 2026 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.0 | 1 | 2026 | Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame · Int. J. Comput. Vis. 2026 |
Robotics › Robot navigation and mapping
visual odometry |
1.0 | 1 | 2026 | Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame · Int. J. Comput. Vis. 2026 |
Computer vision › 3D vision
3d object detection |
0.9 | 1 | 2025 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection · ICCV 2025 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | Video Perception Models for 3D Scene Synthesis · NeurIPS 2025 |
Computer vision › 3D vision › 3d generation
3d scene generation |
0.9 | 1 | 2025 | Video Perception Models for 3D Scene Synthesis · NeurIPS 2025 |
Computer vision › 3D vision › 3d reconstruction › learning-based 3d reconstruction
feed-forward reconstruction |
0.9 | 1 | 2025 | Video Perception Models for 3D Scene Synthesis · NeurIPS 2025 |
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection |
0.9 | 1 | 2025 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection · ICCV 2025 |
Computer vision › Image recognition and object detection
object detection |
0.9 | 1 | 2025 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection · ICCV 2025 |
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-set object detection |
0.9 | 1 | 2025 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection · ICCV 2025 |
Machine learning › Generative modeling › scene generation
scene layout generation |
0.9 | 1 | 2025 | Video Perception Models for 3D Scene Synthesis · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
video diffusion model |
0.9 | 1 | 2025 | Video Perception Models for 3D Scene Synthesis · NeurIPS 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | Video Perception Models for 3D Scene Synthesis · NeurIPS 2025 |
Computer vision › 3D vision
visual localization |
0.9 | 1 | 2025 | CroCoDL: Cross-device Collaborative Dataset for Localization · CVPR 2025 |
Rendering
image-based rendering |
0.8 | 1 | 2024 | MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References · ECCV (9) 2024 |
Rendering
novel view synthesis |
0.8 | 1 | 2024 | MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References · ECCV (9) 2024 |
Robotics › Robot navigation and mapping
localization |
0.3 | 1 | 2026 | Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame · Int. J. Comput. Vis. 2026 |
Computer vision › 3D vision
3d scene understanding |
0.3 | 1 | 2025 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
pose estimation · 1.7image retrieval · 1.7feature extraction · 1.7uncertainty propagation · 1.0transformer · 1.0CNN · 1.0large language model · 0.9geometry prior conditioning · 0.9canonical image space · 0.93d bounding box head · 0.9image matching · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-FrameabstractAbstract Visual odometry (VO) is fundamental to autonomous navigation, robotics, and augmented reality. While self-supervised learning has eliminated the need for expensive ground-truth labels in monocular VO, dynamic objects and occlusions that violate the static scene assumption lead to erroneous pose estimates. Existing uncertainty-based methods filter unreliable regions but rely solely on single-frame information, neglecting temporal consistency across consecutive frames. We present Combined Projected Uncertainty (CoProU), a principled probabilistic formulation that propagates and fuses uncertainties across temporal frames. Our key insight is that robust uncertainty estimation requires combining target frame uncertainty with projected uncertainty from reference frames, enabling effective identification of dynamic regions and temporal inconsistencies. We demonstrate CoProU’s versatility through two complementary frameworks. CoProU-VO-2F employs a decoupled architecture with CNN-based pose encoder and vision transformer-based depth encoder for two-frame visual odometry. CoProU-VO-MF extends our approach to multi-frame scenarios using a unified transformer architecture with coupled encoders that produce shared representations for ego-motion and geometry estimation. This demonstrates that CoProU, though originally formulated for frame pairs, generalizes naturally to multi-frame settings through pairwise application. Comprehensive experiments validate our contributions. CoProU-VO-2F achieves substantial improvements over state-of-the-art two-frame methods, reducing ATE by up to 63% on KITTI and 33% on nuScenes. CoProU-VO-MF achieves 45% lower average ATE across KITTI, nuScenes, and Waymo compared to the large-scale pretrained VGGT baseline. Extensive ablation studies confirm the effectiveness of temporal uncertainty propagation and CoProU’s adaptability across different architectural paradigms. Please check out our Project Page . Jingchao Xie, Oussema Dhaouadi, Johannes Meier, Zuria Bauer, Marc Pollefeys, Daniel Cremers |
Int. J. Comput. Vis. | 5 |
| 2025 | CroCoDL: Cross-device Collaborative Dataset for LocalizationabstractAccurate localization plays a pivotal role in the autonomy of systems operating in unfamiliar environments, particularly when interaction with humans is expected. High-accuracy visual localization systems encompass various components, such as image retrievers, feature extractors, matchers, reconstruction and pose estimation methods. This complexity translates to the necessity of robust evaluation settings and pipelines. However, existing datasets and benchmarks primarily focus on single-agent scenarios, overlooking the critical issue of cross-device localization. Different agents with different sensors will show their own specific strengths and weaknesses, and the data they have available varies substantially. This work addresses this gap by enhancing an existing augmented reality visual localization benchmark with data from legged robots, and evaluating human-robot, cross-device mapping and localization. Our contributions extend beyond device diversity and include high environment variability, spanning ten distinct locations ranging from disaster sites to art exhibitions. Each scene in our dataset features recordings from robot agents, hand-held and head-mounted devices, and high-accuracy ground truth LiDAR scanners, resulting in a comprehensive multi-agent dataset and benchmark. This work represents a significant advancement in the field of visual localization benchmarking, with key in-sights into the performance of cross-device localization methods across diverse settings. Hermann Blum, Alessandro Mercurio, Joshua O'Reilly, Tim Engelbracht, Mihai Dusmanu, Marc Pollefeys, Zuria Bauer |
CVPR | 7 |
| 2025 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionabstractMonocular 3D object detection is valuable for various applications such as robotics and AR/VR. Existing methods are confined to closed-set settings, where the training and testing sets consist of the same scenes and/or object categories. However, real-world applications often introduce new environments and novel object categories, posing a challenge to these methods. In this paper, we address monocular 3D object detection in an open-set setting and introduce the first end-to-end 3D Monocular Open-set Object Detector (3D-MOOD). We propose to lift the open-set 2D detection into 3D space through our designed 3D bounding box head, enabling end-to-end joint training for both 2D and 3D tasks to yield better overall performance. We condition the object queries with geometry prior and overcome the generalization for 3D estimation across diverse scenes. To further improve performance, we design the canonical image space for more efficient cross-dataset training. We evaluate 3D-MOOD on both closed-set settings (Omni3D) and open-set settings (Omni3D to Argoverse 2, ScanNet), and achieve new state-of-the-art results. Code and models are available at royyang0714.github.io/3D-MOOD. Yung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Siyuan Li 0008, Rui Huang 0012, Yuqian Fu, Marc Pollefeys, Hermann Blum, Zuria Bauer |
ICCV | 9 |
| 2025 | Video Perception Models for 3D Scene SynthesisabstractAutomating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or strong visual priors from image generation models. However, current LLMs exhibit limited 3D spatial reasoning, undermining the realism and global coherence of synthesized scenes, while image-generation-based methods often constrain viewpoint control and introduce multi-view inconsistencies. In this work, we present Video Perception models for 3D Scene synthesis (VIPScene), a novel framework that exploits the encoded commonsense knowledge of the 3D physical world in video generation models to ensure coherent scene layouts and consistent object placements across views. VIPScene accepts both text and image prompts and seamlessly integrates video generation, feedforward 3D reconstruction, and open-vocabulary perception models to semantically and geometrically analyze each object in a scene. This enables flexible scene synthesis with high realism and structural consistency. For a more sufficient evaluation on coherence and plausibility, we further introduce First-Person View Score (FPVScore), utilizing a continuous first-person perspective to capitalize on the reasoning ability of multimodal large language models. Extensive experiments show that VIPScene significantly outperforms existing methods and generalizes well across diverse scenarios. Rui Huang 0012, Guangyao Zhai, Zuria Bauer, Marc Pollefeys, Federico Tombari, Leonidas J. Guibas, Gao Huang 0001, Francis Engelmann |
NeurIPS | 3 |
| 2024 | MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References
Lukas Bösiger, Mihai Dusmanu, Marc Pollefeys, Zuria Bauer |
ECCV (9) | 4 |
| 2021 | NVS-MonoDepth: Improving Monocular Depth Prediction with Novel View SynthesisabstractBuilding upon the recent progress in novel view synthesis, we propose its application to improve monocular depth estimation. In particular, we propose a novel training method split in three main steps. First, the prediction results of a monocular depth network are warped to an additional view point. Second, we apply an additional image synthesis network, which corrects and improves the quality of the warped RGB image. The output of this network is required to look as similar as possible to the ground-truth view by minimizing the pixel-wise RGB reconstruction error. Third, we reapply the same monocular depth estimation onto the synthesized second view point and ensure that the depth predictions are consistent with the associated ground truth depth. Experimental results prove that our method achieves state-of-the-art or comparable performance on the KITTI and NYU-Depth-v2 datasets with a lightweight and simple vanilla U-Net architecture. Zuria Bauer, Zuoyue Li, Sergio Orts, Miguel Cazorla, Marc Pollefeys, Martin R. Oswald |
3DV | 1 |
| 2020 | Enhancing perception for the visually impaired with deep learning techniques and low-cost wearable sensorsabstractAs estimated by the World Health Organization, there are millions of people who lives with some form of vision impairment . As a consequence, some of them present mobility problems in outdoor environments . With the aim of helping them, we propose in this work a system which is capable of delivering the position of potential obstacles in outdoor scenarios. Our approach is based on non-intrusive wearable devices and focuses also on being low-cost. First, a depth map of the scene is estimated from a color image, which provides 3D information of the environment. Then, an urban object detector is in charge of detecting the semantics of the objects in the scene. Finally, the three-dimensional and semantic data is summarized in a simpler representation of the potential obstacles the users have in front of them. This information is transmitted to the user through spoken or haptic feedback. Our system is able to run at about 3.8 fps and achieved a 87.99% mean accuracy in obstacle presence detection. Finally, we deployed our system in a pilot test which involved an actual person with vision impairment, who validated the effectiveness of our proposal for improving its navigation capabilities in outdoors. Zuria Bauer, Alejandro Dominguez, Edmanuel Cruz, Francisco Gomez-Donoso, Sergio Orts, Miguel Cazorla |
Pattern Recognit. Lett. | 1 |
| 2020 | COMBAHO: A deep learning system for integrating brain injury patients in societyabstractIn the last years, the care of dependent people, either by disease, accident, disability, or age, is one of the current priority research topics in developed countries. Moreover, such care is intended to be at patients home, in order to minimize the cost of therapies. Patients rehabilitation will be fulfilled when their integration in society is achieved, either in the family or in a work environment. To address this challenge, we propose the development and evaluation of an assistant for people with acquired brain injury or dependents. This assistant is twofold: in the patient’s home is based on the design and use of an intelligent environment with abilities to monitor and active learning, combined with an autonomous social robot for interactive assistance and stimulation. On the other hand, it is complemented with an outdoor assistant, to help patients under disorientation or complex situations. This involves the integration of several existing technologies and provides solutions to a variety of technological challenges. Deep leaning-based techniques are proposed as core technology to solve these problems. José García Rodríguez 0001, Francisco Gomez-Donoso, Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, Miguel Cazorla, Sergio Orts, Zuria Bauer, John Alejandro Castro-Vargas, Félix Escalona, David Ivorra-Piqueres, Pablo Martinez-Gonzalez, Eugenio Aguirre, Miguel García-Silvente, Marcelo García-Pérez, José María Cañas, Francisco Martín 0001, Jonatan Gines Clavero, Francisco Rivas-Montero |
Pattern Recognit. Lett. | 7 |
| 2018 | Finding the Place: How to Train and Use Convolutional Neural Networks for a Dynamically Learning RobotabstractFor a robot, the ability to adapt his knowledge automatically and customize its behavior is a key feature. Furthermore, a robot should be able to carry out its tasks at a long-term basis, performing it seamlessly in presence of changes in their surroundings. To do that, it is essential that the robot dynamically learn from their environment, but to perform a fully retraining of a deep learning architecture when the model needs new knowledge is a highly time consuming task. This work focus on exploring several strategies to include new data to an already learned model, applied to the semantic localization problem focusing in the accuracy of the final model and their training time. Exhaustive experimentation is carried out and each result is discussed consequently. Edmanuel Cruz, José Carlos Rangel, Francisco Gomez-Donoso, Zuria Bauer, Miguel Cazorla, José García Rodríguez 0001 |
IJCNN | 4 |