Nan Yang 0007

dblp:51/1629-7 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-1497-9630ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 56% Robot navigation and mapping · 34% Autonomous driving · 10%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
visual odometry
1.632025
4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions · Int. J. Comput. Vis. 2025
D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry · CVPR 2020
Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry · ECCV (8) 2018
Robotics › Autonomous driving
perception
1.322025
4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions · Int. J. Comput. Vis. 2025
DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation · ICRA 2020
Computer vision › 3D vision
3d scene reconstruction
1.222023
Behind the Scenes: Density Fields for Single View Reconstruction · CVPR 2023
MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera · CVPR 2021
Computer vision › 3D vision
depth estimation
1.032021
MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera · CVPR 2021
D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry · CVPR 2020
Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry · ECCV (8) 2018
Robotics › Robot navigation and mapping › localization
long-term localization
0.912025
4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions · Int. J. Comput. Vis. 2025
Robotics › Robot navigation and mapping
place recognition
0.912025
4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions · Int. J. Comput. Vis. 2025
Computer vision › 3D vision
visual localization
0.912025
4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions · Int. J. Comput. Vis. 2025
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.912025
4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions · Int. J. Comput. Vis. 2025
Computer vision › 3D vision
novel view synthesis
0.712023
Behind the Scenes: Density Fields for Single View Reconstruction · CVPR 2023
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.712023
Behind the Scenes: Density Fields for Single View Reconstruction · CVPR 2023
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.512021
MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera · CVPR 2021
Computer vision › 3D vision › 3d reconstruction › single-view 3d reconstruction
monocular dense reconstruction
0.512021
MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera · CVPR 2021
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.512021
MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera · CVPR 2021
Computer vision › 3D vision
3d reconstruction
0.412020
D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry · CVPR 2020
Robotics › Robot navigation and mapping › visual odometry
monocular visual odometry
0.412020
D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry · CVPR 2020
Computer vision › 3D vision
object pose estimation
0.412020
DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation · ICRA 2020
Computer vision › 3D vision
pose estimation
0.412020
D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry · CVPR 2020
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised monocular depth estimation
0.412020
D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry · CVPR 2020

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.1volume rendering · 0.7neural radiance field · 0.7semi-supervised learning · 0.5multi-view stereo · 0.5moving object mask prediction · 0.5cost volume · 0.5photometric alignment · 0.4nonlinear optimization · 0.4direct visual odometry · 0.4
YearPublicationVenuePosition
2025 4Seasons: Benchmarking Visual SLAM and Long-Term Localization for Autonomous Driving in Challenging Conditions
abstract
Abstract In this paper, we present a novel visual SLAM and long-term localization benchmark for autonomous driving in challenging conditions based on the large-scale 4Seasons dataset. The proposed benchmark provides drastic appearance variations caused by seasonal changes and diverse weather and illumination conditions. While significant progress has been made in advancing visual SLAM on small-scale datasets with similar conditions, there is still a lack of unified benchmarks representative of real-world scenarios for autonomous driving. We introduce a new unified benchmark for jointly evaluating visual odometry, global place recognition, and map-based visual localization performance which is crucial to successfully enable autonomous driving in any condition. The data has been collected for more than one year, resulting in more than 300 km of recordings in nine different environments ranging from a multi-level parking garage to urban (including tunnels) to countryside and highway. We provide globally consistent reference poses with up to centimeter-level accuracy obtained from the fusion of direct stereo-inertial odometry with RTK GNSS. We evaluate the performance of several state-of-the-art visual odometry and visual localization baseline approaches on the benchmark and analyze their properties. The experimental results provide new insights into current approaches and show promising potential for future research. Our benchmark and evaluation protocols will be available at https://go.vision.in.tum.de/4seasons .
Patrick Wenzel, Nan Yang 0007, Rui Wang 0037, Niclas Zeller, Daniel Cremers
Int. J. Comput. Vis.2
2024 FIRe: Fast Inverse Rendering using Directional and Signed Distance Functions
abstract
Neural 3D implicit representations learn priors that are useful for diverse applications, such as single- or multiple-view 3D reconstruction. A major downside of existing approaches while rendering an image is that they require evaluating the network multiple times per camera ray so that the high computational time forms a bottleneck for downstream applications. We address this problem by introducing a novel neural scene representation that we call the directional distance function (DDF). To this end, we learn a signed distance function (SDF) along with our DDF model to represent a class of shapes. Specifically, our DDF is defined on the unit sphere and predicts the distance to the surface along any given direction. Therefore, our DDF allows rendering images with just a single network evaluation per camera ray. Based on our DDF, we present a novel fast algorithm (FIRe) to reconstruct 3D shapes given a posed depth map. We evaluate our proposed method on 3D reconstruction from single-view depth images, where we empirically show that our algorithm reconstructs 3D shapes more accurately and it is more than 15 times faster (per iteration) than competing methods.
Tarun Yenamandra, Ayush Tewari, Nan Yang 0007, Florian Bernard 0001, Christian Theobalt, Daniel Cremers
WACV3
2023 Behind the Scenes: Density Fields for Single View Reconstruction
abstract
Inferring a meaningful geometric scene representation from a single image is a fundamental problem in computer vision. Approaches based on traditional depth map prediction can only reason about areas that are visible in the image. Currently, neural radiance fields (NeRFs) can capture true 3D including color, but are too complex to be generated from a single image. As an alternative, we propose to predict an implicit density field from a single image. It maps every location in the frustum of the image to volumetric density. By directly sampling color from the available views instead of storing color in the density field, our scene representation becomes significantly less complex compared to NeRFs, and a neural network can predict it in a single forward pass. The network is trained through self-supervision from only video data. Our formulation allows volume rendering to perform both depth prediction and novel view synthesis. Through experiments, we show that our method is able to predict meaningful geometry for regions that are occluded in the input image. Additionally, we demonstrate the potential of our approach on three datasets for depth prediction and novel-view synthesis.
Felix Wimbauer, Nan Yang 0007, Christian Rupprecht 0001, Daniel Cremers
CVPR2
2021 MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera
abstract
In this paper, we propose MonoRec, a semi-supervised monocular dense reconstruction architecture that predicts depth maps from a single moving camera in dynamic environments. MonoRec is based on a multi-view stereo setting which encodes the information of multiple consecutive images in a cost volume. To deal with dynamic objects in the scene, we introduce a MaskModule that predicts moving object masks by leveraging the photometric inconsistencies encoded in the cost volumes. Unlike other multi-view stereo methods, MonoRec is able to reconstruct both static and moving objects by leveraging the predicted masks. Furthermore, we present a novel multi-stage training scheme with a semi-supervised loss formulation that does not require LiDAR depth values. We carefully evaluate MonoRec on the KITTI dataset and show that it achieves state-of-theart performance compared to both multi-view and singleview methods. With the model trained on KITTI, we furthermore demonstrate that MonoRec is able to generalize well to both the Oxford RobotCar dataset and the more challenging TUM-Mono dataset recorded by a handheld camera. Code and related materials are available at https://vision.in.tum.de/research/monorec.
Felix Wimbauer, Nan Yang 0007, Lukas von Stumberg, Niclas Zeller, Daniel Cremers
CVPR2
2020 LM-Reloc: Levenberg-Marquardt Based Direct Visual Relocalization
abstract
We present LM-Reloc-a novel approach for visual relocalization based on direct image alignment. In contrast to prior works that tackle the problem with a feature-based formulation, the proposed method does not rely on feature matching and RANSAC. Hence, the method can utilize not only corners but any region of the image with gradients. In particular, we propose a loss formulation inspired by the classical Levenberg-Marquardt algorithm to train LM-Net. The learned features significantly improve the robustness of direct image alignment, especially for relocalization across different conditions. To further improve the robustness of LM-Net against large image baselines, we propose a pose estimation network, CorrPoseNet, which regresses the relative pose to bootstrap the direct image alignment. Evaluations on the CARLA and Oxford RobotCar relocalization tracking benchmark show that our approach delivers more accurate results than previous state-of-the-art methods while being comparable in terms of robustness.
Lukas von Stumberg, Patrick Wenzel, Nan Yang 0007, Daniel Cremers
3DV3
2020 D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
abstract
We propose D3VO as a novel framework for monocular visual odometry that exploits deep networks on three levels -- deep depth, pose and uncertainty estimation. We first propose a novel self-supervised monocular depth estimation network trained on stereo videos without any external supervision. In particular, it aligns the training image pairs into similar lighting condition with predictive brightness transformation parameters. Besides, we model the photometric uncertainties of pixels on the input images, which improves the depth estimation accuracy and provides a learned weighting function for the photometric residuals in direct (feature-less) visual odometry. Evaluation results show that the proposed network outperforms state-of-the-art self-supervised depth estimation networks. D3VO tightly incorporates the predicted depth, pose and uncertainty into a direct visual odometry method to boost both the front-end tracking as well as the back-end non-linear optimization. We evaluate D3VO in terms of monocular visual odometry on both the KITTI odometry benchmark and the EuRoC MAV dataset. The results show that D3VO outperforms state-of-the-art traditional monocular VO methods by a large margin. It also achieves comparable results to state-of-the-art stereo/LiDAR odometry on KITTI and to the state-of-the-art visual-inertial odometry on EuRoC MAV, while using only a single camera.
Nan Yang 0007, Lukas von Stumberg, Rui Wang 0037, Daniel Cremers
CVPR1
2020 DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation
abstract
Scene understanding from images is a challenging problem encountered in autonomous driving. On the object level, while 2D methods have gradually evolved from computing simple bounding boxes to delivering finer grained results like instance segmentations, the 3D family is still dominated by estimating 3D bounding boxes. In this paper, we propose a novel approach to jointly infer the 3D rigid-body poses and shapes of vehicles from a stereo image pair using shape priors. Unlike previous works that geometrically align shapes to point clouds from dense stereo reconstruction, our approach works directly on images by combining a photometric and a silhouette alignment term in the energy function. An adaptive sparse point selection scheme is proposed to efficiently measure the consistency with both terms. In experiments, we show superior performance of our method on 3D pose and shape estimation over the previous geometric approach and demonstrate that our method can also be applied as a refinement step and significantly boost the performances of several state-of-the-art deep learning based 3D object detectors. All related materials and demonstration videos are available at the project page https://vision.in.tum.de/research/vslam/direct-shape.
Rui Wang 0037, Nan Yang 0007, Jörg Stückler, Daniel Cremers
ICRA2
2018 Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry
Nan Yang 0007, Rui Wang 0037, Jörg Stückler, Daniel Cremers
ECCV (8)1