Benjamin Ummenhofer

dblp:86/10064 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0002-7467-4273ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
3D vision · 58% Segmentation and scene understanding · 16% Robot navigation and mapping · 12%
Computer graphics and multimedia
3 papers
Image and video processing · 41% Rendering · 38% Geometric modeling and processing · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.242021
Adaptive Surface Reconstruction with Multiscale Convolutional Kernels · ICCV 2021
Global, Dense Multiscale Reconstruction for a Billion Points · Int. J. Comput. Vis. 2017
Global, Dense Multiscale Reconstruction for a Billion Points · ICCV 2015
Rendering
physically based rendering
0.912025
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors · NeurIPS 2025
Image and video processing
super-resolution
0.912025
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors · NeurIPS 2025
Image and video processing › super-resolution
texture super-resolution
0.912025
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors · NeurIPS 2025
Robotics › Robot navigation and mapping
SLAM
0.822020
DeepTAM: Deep Tracking and Mapping with Convolutional Neural Networks · Int. J. Comput. Vis. 2020
DeepTAM: Deep Tracking and Mapping · ECCV (16) 2018
Rendering
neural radiance fields
0.812024
Mesh2NeRF: Direct Mesh Supervision for Neural Radiance Field Representation and Generation · ECCV (9) 2024
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.722021
Adaptive Surface Reconstruction with Multiscale Convolutional Kernels · ICCV 2021
Global, Dense Multiscale Reconstruction for a Billion Points · ICCV 2015
Computer vision › Segmentation and scene understanding
3d semantic segmentation
0.612022
Segment-Fusion: Hierarchical Context Fusion for Robust 3D Semantic Segmentation · CVPR 2022
Computer vision › Segmentation and scene understanding
instance segmentation
0.612022
Segment-Fusion: Hierarchical Context Fusion for Robust 3D Semantic Segmentation · CVPR 2022
Machine learning › Deep learning architectures and training
physics-informed neural network
0.612022
Guaranteed Conservation of Momentum for Learning Particle-based Fluid Dynamics · NeurIPS 2022
Computational science and engineering › computational physics
physics simulation
0.612022
Guaranteed Conservation of Momentum for Learning Particle-based Fluid Dynamics · NeurIPS 2022
Computer animation and physical simulation
fluid simulation
0.412020
Lagrangian Fluid Simulation with Continuous Convolutions · ICLR 2020
Machine learning › Transfer learning and domain adaptation
domain generalization
0.412019
CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth · CVPR 2019
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.412019
CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth · CVPR 2019
Computer vision › 3D vision
depth estimation
0.312017
DeMoN: Depth and Motion Network for Learning Monocular Stereo · CVPR 2017
Computer vision › 3D vision › stereo vision
monocular stereo
0.312017
DeMoN: Depth and Motion Network for Learning Monocular Stereo · CVPR 2017
Computer vision › 3D vision
point cloud processing
0.312017
Global, Dense Multiscale Reconstruction for a Billion Points · Int. J. Comput. Vis. 2017
Computer vision › 3D vision
structure from motion
0.312017
DeMoN: Depth and Motion Network for Learning Monocular Stereo · CVPR 2017
Geometric modeling and processing
mesh processing
0.312025
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors · NeurIPS 2025
Geometric modeling and processing › shape representation
mesh representation
0.212024
Mesh2NeRF: Direct Mesh Supervision for Neural Radiance Field Representation and Generation · ECCV (9) 2024
Computer vision › 3D vision › structure from motion
large-scale reconstruction
0.212015
Global, Dense Multiscale Reconstruction for a Billion Points · ICCV 2015
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.212013
Point-Based 3D Reconstruction of Thin Objects · ICCV 2013
Computer vision › 3D vision › 3d reconstruction › object reconstruction
thin object reconstruction
0.212013
Point-Based 3D Reconstruction of Thin Objects · ICCV 2013
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
continuous convolution
0.112020
Lagrangian Fluid Simulation with Continuous Convolutions · ICLR 2020
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.112020
DeepTAM: Deep Tracking and Mapping with Convolutional Neural Networks · Int. J. Comput. Vis. 2020

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.6temporal coherence training · 1.1resampling · 1.1antisymmetric continuous convolution · 1.1zero-shot super-resolution · 0.9image priors · 0.9differentiable rendering · 0.9neural radiance field · 0.8mesh supervision · 0.8hierarchical networks · 0.6hierarchical network · 0.6graph segmentation · 0.6connected component labeling · 0.6attention-based hierarchical fusion · 0.6octree · 0.5multi-scale kernel · 0.5neural network · 0.4continuous convolution · 0.4
YearPublicationVenuePosition
2025 PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors
abstract
We present PBR-SR, a novel method for physically based rendering (PBR) texture super resolution (SR). It outputs high-resolution, high-quality PBR textures from low-resolution (LR) PBR input in a zero-shot manner. PBR-SR leverages an off-the-shelf super-resolution model trained on natural images, and iteratively minimizes the deviations between super-resolution priors and differentiable renderings. These enhancements are then back-projected into the PBR map space in a differentiable manner to produce refined, high-resolution textures. To mitigate the effects of view inconsistency and lighting sensitivity inherent to view-based super-resolution, our approach incorporates 2D prior constraints across multi-view renderings, enabling iterative refinement of shared upscaled textures. In parallel, we incorporate identity constraints directly in the PBR texture domain to ensure the upscaled textures remain faithful to the LR input. PBR-SR operates without any additional training or data requirements, relying entirely on pretrained image priors. We demonstrate that our approach produces high-fidelity PBR textures for both artist-designed and AI-generated meshes, outperforming both direct SR models application and prior texture optimization methods. Our results show high-quality outputs in both PBR and rendering evaluations, supporting advanced applications such as relighting.
Yujin Chen, Yinyu Nie, Benjamin Ummenhofer, Reiner Birkl, Michael Paulitsch, Matthias Nießner
NeurIPS3
2025 HoloScene: Simulation-Ready Interactive 3D Worlds from a Single Video
abstract
Digitizing the physical world into accurate simulation‑ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and scene-understanding methods commonly fall short in one or more critical aspects, such as geometry completeness, object interactivity, physical plausibility, photorealistic rendering, or realistic physical properties for reliable dynamic simulation. To address these limitations, we introduce HoloScene, a novel interactive 3D reconstruction framework that simultaneously achieves these requirements. HoloScene leverages a comprehensive interactive scene-graph representation, encoding object geometry, appearance, and physical properties alongside hierarchical and inter-object relationships. Reconstruction is formulated as an energy-based optimization problem, integrating observational data, physical constraints, and generative priors into a unified, coherent objective. Optimization is efficiently performed via a hybrid approach combining sampling-based exploration with gradient-based refinement. The resulting digital twins exhibit complete and precise geometry, physical stability, and realistic rendering from novel viewpoints. Evaluations conducted on multiple benchmark datasets demonstrate superior performance, while practical use-cases in interactive gaming and real-time digital-twin manipulation illustrate HoloScene's broad applicability and effectiveness.
Hongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet, Katelyn Gao, Michael Paulitsch, Benjamin Ummenhofer, Shenlong Wang
NeurIPS7
2024 Objects With Lighting: A Real-World Dataset for Evaluating Reconstruction and Rendering for Object Relighting
abstract
Reconstructing an object from photos and placing it virtually in a new environment goes beyond the standard novel view synthesis task as the appearance of the object has to not only adapt to the novel viewpoint but also to the new lighting conditions and yet evaluations of inverse rendering methods rely on novel view synthesis data or simplistic synthetic datasets for quantitative analysis. This work presents a real-world dataset for measuring the reconstruction and rendering of objects for relighting. To this end, we capture the environment lighting and ground truth images of the same objects in multiple environments allowing to reconstruct the objects from images taken in one environment and quantify the quality of the rendered views for the unseen lighting environments. Further, we introduce a simple baseline composed of off-the-shelf methods and test several state-of-the-art methods on the relighting task and show that novel view synthesis is not a reliable proxy to measure performance. Code and dataset are available at https://github.com/isl-org/objects-with-lighting.
Benjamin Ummenhofer, Sanskar Agrawal, Rene Sepúlveda, Yixing Lao, Tianhang Cheng, Stephan R. Richter, Shenlong Wang, Germán Ros 0001
3DV1
2024 Mesh2NeRF: Direct Mesh Supervision for Neural Radiance Field Representation and Generation
Yujin Chen, Yinyu Nie, Benjamin Ummenhofer, Reiner Birkl, Michael Paulitsch, Matthias Müller 0011, Matthias Nießner
ECCV (9)3
2022 Segment-Fusion: Hierarchical Context Fusion for Robust 3D Semantic Segmentation
abstract
3D semantic segmentation is a fundamental building block for several scene understanding applications such as autonomous driving, robotics and AR/VR. Several state-of-the-art semantic segmentation models suffer from the part-misclassification problem, wherein parts of the same object are labelled incorrectly. Previous methods have utilized hierarchical, iterative methods to fuse semantic and instance information, but they lack learnability in context fusion, and are computationally complex and heuristic driven. This paper presents Segment-Fusion, a novel attention-based method for hierarchical fusion of semantic and instance in-formation to address the part misclassifications. The presented method includes a graph segmentation algorithmfor grouping points into segments that pools point-wise features into segment-wise features, a learnable attention-based net-work to fuse these segments based on their semantic and instance features, and followed by a simple yet effective connected component labelling algorithm to convert seg-ment features to instance labels. Segment-Fusion can be flexibly employed with any network architecture for semantic/instance segmentation. It improves the qualitative and quantitative performance of several semantic segmentation backbones by upto 5% on the ScanNet and S3DIS datasets.
Anirud Thyagharajan, Benjamin Ummenhofer, Prashant Laddha, Om Ji Omer, Sreenivas Subramoney
CVPR2
2022 Guaranteed Conservation of Momentum for Learning Particle-based Fluid Dynamics
abstract
We present a novel method for guaranteeing linear momentum in learned physics simulations. Unlike existing methods, we enforce conservation of momentum with a hard constraint, which we realize via antisymmetrical continuous convolutional layers. We combine these strict constraints with a hierarchical network architecture, a carefully constructed resampling scheme, and a training approach for temporal coherence. In combination, the proposed method allows us to increase the physical accuracy of the learned simulator substantially. In addition, the induced physical bias leads to significantly better generalization performance and makes our method more reliable in unseen test cases. We evaluate our method on a range of different, challenging fluid scenarios. Among others, we demonstrate that our approach generalizes to new scenarios with up to one million particles. Our results show that the proposed algorithm can learn complex dynamics while outperforming existing approaches in generalization and training performance. An implementation of our approach is available at https://github.com/tum-pbs/DMCF.
Lukas Prantl, Benjamin Ummenhofer, Vladlen Koltun, Nils Thürey
NeurIPS2
2021 Adaptive Surface Reconstruction with Multiscale Convolutional Kernels
abstract
We propose generalized convolutional kernels for 3D reconstruction with ConvNets from point clouds. Our method uses multiscale convolutional kernels that can be applied to adaptive grids as generated with octrees. In addition to standard kernels in which each element has a distinct spatial location relative to the center, our elements have a distinct relative location as well as a relative scale level. Making our kernels span multiple resolutions allows us to apply ConvNets to adaptive grids for large problem sizes where the input data is sparse but the entire domain needs to be processed. Our ConvNet architecture can predict the signed and unsigned distance fields for large data sets with millions of input points and is faster and more accurate than classic energy minimization or recent learning approaches. We demonstrate this in a zero-shot setting where we only train on synthetic data and evaluate on the Tanks and Temples dataset of real-world large-scale 3D scenes.
Benjamin Ummenhofer, Vladlen Koltun
ICCV1
2020 Lagrangian Fluid Simulation with Continuous Convolutions
Benjamin Ummenhofer, Lukas Prantl, Nils Thürey, Vladlen Koltun
ICLR1
2020 DeepTAM: Deep Tracking and Mapping with Convolutional Neural Networks
Huizhong Zhou, Benjamin Ummenhofer, Thomas Brox
Int. J. Comput. Vis.2
2019 CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth
abstract
Single-view depth estimation suffers from the problem that a network trained on images from one camera does not generalize to images taken with a different camera model. Thus, changing the camera model requires collecting an entirely new training dataset. In this work, we propose a new type of convolution that can take the camera parameters into account, thus allowing neural networks to learn calibration-aware patterns. Experiments confirm that this improves the generalization capabilities of depth prediction networks considerably, and clearly outperforms the state of the art when the train and test images are acquired with different cameras.
José M. Fácil, Benjamin Ummenhofer, Huizhong Zhou, Luis Montesano, Thomas Brox, Javier Civera 0001
CVPR2
2018 DeepTAM: Deep Tracking and Mapping
Huizhong Zhou, Benjamin Ummenhofer, Thomas Brox
ECCV (16)2
2017 DeMoN: Depth and Motion Network for Learning Monocular Stereo
abstract
In this paper we formulate structure from motion as a learning problem. We train a convolutional network end-to-end to compute depth and camera motion from successive, unconstrained image pairs. The architecture is composed of multiple stacked encoder-decoder networks, the core part being an iterative network that is able to improve its own predictions. The network estimates not only depth and motion, but additionally surface normals, optical flow between the images and confidence of the matching. A crucial component of the approach is a training loss based on spatial relative differences. Compared to traditional two-frame structure from motion methods, results are more accurate and more robust. In contrast to the popular depth-from-single-image networks, DeMoN learns the concept of matching and, thus, better generalizes to structures not seen during training.
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, Thomas Brox
CVPR1
2017 Global, Dense Multiscale Reconstruction for a Billion Points
Benjamin Ummenhofer, Thomas Brox
Int. J. Comput. Vis.1
2015 Global, Dense Multiscale Reconstruction for a Billion Points
abstract
We present a variational approach for surface reconstruction from a set of oriented points with scale information. We focus particularly on scenarios with non-uniform point densities due to images taken from different distances. In contrast to previous methods, we integrate the scale information in the objective and globally optimize the signed distance function of the surface on a balanced octree grid. We use a finite element discretization on the dual structure of the octree minimizing the number of variables. The tetrahedral mesh is generated efficiently from the dual structure, and also memory efficiency is optimized, such that robust data terms can be used even on very large scenes. The surface normals are explicitly optimized and used for surface extraction to improve the reconstruction at edges and corners.
Benjamin Ummenhofer, Thomas Brox
ICCV1
2013 Point-Based 3D Reconstruction of Thin Objects
abstract
3D reconstruction deals with the problem of finding the shape of an object from a set of images. Thin objects that have virtually no volume pose a special challenge for reconstruction with respect to shape representation and fusion of depth information. In this paper we present a dense point-based reconstruction method that can deal with this special class of objects. We seek to jointly optimize a set of depth maps by treating each pixel as a point in space. Points are pulled towards a common surface by pair wise forces in an iterative scheme. The method also handles the problem of opposed surfaces by means of penalty forces. Efficient optimization is achieved by grouping points to super pixels and a spatial hashing approach for fast neighborhood queries. We show that the approach is on a par with state-of-the-art methods for standard multi view stereo settings and gives superior results for thin objects.
Benjamin Ummenhofer, Thomas Brox
ICCV1