VLDB 2026 Research / reviewers in the wild / expert
Huangying Zhan
dblp:210/9949
· DBLP profile ↗
15ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-2899-8314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
3D vision · 70% Robot navigation and mapping · 22% Reinforcement learning · 8% |
Topics — the 28 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
depth estimation |
3.4 | 7 | 2024 | SC-DepthV3: Robust Self-Supervised Monocular Depth Estimation for Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Auto-Rectify Network for Unsupervised Indoor Depth Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Unsupervised Scale-Consistent Depth Learning from Video · Int. J. Comput. Vis. 2021 |
Computer vision › 3D vision
3d reconstruction |
2.5 | 3 | 2025 | PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance Fields · ICRA 2025 ActiveGAMER: Active GAussian Mapping through Efficient Rendering · CVPR 2025 NARUTO: Neural Active Reconstruction from Uncertain Target Observations · CVPR 2024 |
Computer vision › 3D vision › depth estimation
self-supervised depth estimation |
1.3 | 3 | 2022 | Auto-Rectify Network for Unsupervised Indoor Depth Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video · NeurIPS 2019 Self-supervised Learning for Single View Depth and Surface Normal Estimation · ICRA 2019 |
Robotics › Robot navigation and mapping › active perception
active mapping |
0.9 | 1 | 2025 | ActiveGAMER: Active GAussian Mapping through Efficient Rendering · CVPR 2025 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | ActiveGAMER: Active GAussian Mapping through Efficient Rendering · CVPR 2025 |
Robotics › Robot navigation and mapping › view planning
next-best-view planning |
0.9 | 1 | 2025 | ActiveGAMER: Active GAussian Mapping through Efficient Rendering · CVPR 2025 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.8 | 2 | 2020 | Visual Odometry Revisited: What Should Be Learnt? · ICRA 2020 Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction · CVPR 2018 |
Robotics › Robot navigation and mapping
visual odometry |
0.8 | 2 | 2020 | Visual Odometry Revisited: What Should Be Learnt? · ICRA 2020 Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction · CVPR 2018 |
Robotics › Robot navigation and mapping › active perception
active reconstruction |
0.8 | 1 | 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target Observations · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
neural surface reconstruction |
0.8 | 1 | 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target Observations · CVPR 2024 |
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised monocular depth estimation |
0.8 | 1 | 2024 | SC-DepthV3: Robust Self-Supervised Monocular Depth Estimation for Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Reinforcement learning › exploration
uncertainty-guided exploration |
0.8 | 1 | 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target Observations · CVPR 2024 |
Computer vision › 3D vision › depth estimation › scene depth estimation
indoor depth estimation |
0.6 | 1 | 2022 | Auto-Rectify Network for Unsupervised Indoor Depth Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › 3D vision › depth estimation
unsupervised depth learning |
0.5 | 1 | 2021 | Unsupervised Scale-Consistent Depth Learning from Video · Int. J. Comput. Vis. 2021 |
Robotics › Robot navigation and mapping › visual odometry
monocular visual odometry |
0.4 | 1 | 2020 | Visual Odometry Revisited: What Should Be Learnt? · ICRA 2020 |
Computer vision › 3D vision › motion estimation
optical flow |
0.4 | 1 | 2020 | Visual Odometry Revisited: What Should Be Learnt? · ICRA 2020 |
Computer vision › 3D vision
surface normal estimation |
0.4 | 1 | 2019 | Self-supervised Learning for Single View Depth and Surface Normal Estimation · ICRA 2019 |
Robotics › Robot navigation and mapping
SLAM |
0.4 | 2 | 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target Observations · CVPR 2024 Visual Odometry Revisited: What Should Be Learnt? · ICRA 2020 |
Computer vision › 3D vision › 3d shape representation
deformation field |
0.3 | 1 | 2018 | Efficient Dense Point Cloud Object Reconstruction Using Deformation Vector Fields · ECCV (12) 2018 |
Computer vision › 3D vision › 3d reconstruction
point cloud reconstruction |
0.3 | 1 | 2018 | Efficient Dense Point Cloud Object Reconstruction Using Deformation Vector Fields · ECCV (12) 2018 |
Computer vision › 3D vision
pose estimation |
0.3 | 1 | 2018 | Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction · CVPR 2018 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.3 | 1 | 2025 | ActiveGAMER: Active GAussian Mapping through Efficient Rendering · CVPR 2025 |
Computer vision › 3D vision
neural radiance field |
0.3 | 1 | 2025 | PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance Fields · ICRA 2025 |
Computer vision › 3D vision
novel view synthesis |
0.3 | 1 | 2025 | ActiveGAMER: Active GAussian Mapping through Efficient Rendering · CVPR 2025 |
Computer vision › 3D vision › 3d scene modeling
scene representation |
0.3 | 1 | 2025 | PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance Fields · ICRA 2025 |
Robotics › Robot navigation and mapping › SLAM
neural SLAM |
0.2 | 1 | 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target Observations · CVPR 2024 |
Robotics › Robot navigation and mapping › state estimation
egomotion |
0.2 | 1 | 2022 | Auto-Rectify Network for Unsupervised Indoor Depth Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › 3D vision › motion estimation
ego-motion estimation |
0.1 | 1 | 2019 | Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 1.4self-supervised learning · 1.1plane fitting · 0.9neural radiance field · 0.9memory bank · 0.9keyframe selection · 0.9information gain · 0.93d gaussian splatting · 0.9multi-resolution hashgrid · 0.8active ray sampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Voxel Grid Optimization for High-Fidelity RGB-D Supervised Surface Reconstruction
Xiangyu Xu 0004, Qingan Yan, Changjiang Cai, Huangying Zhan, Pan Ji, Junsong Yuan 0001, Yi Xu 0002 |
CGI (2) | 4 |
| 2025 | ActiveGAMER: Active GAussian Mapping through Efficient RenderingabstractWe introduce ActiveGAMER, an active mapping system that utilizes 3D Gaussian Splatting (3DGS) to achieve high-quality scene mapping and efficient exploration. Unlike recent NeRF-based methods, which are computationally demanding and limit mapping performance, our approach leverages the efficient rendering capabilities of 3DGS to enable effective and efficient exploration in complex environments. The core of our system is a rendering-based information gain module that identifies the most informative viewpoints for next-best-view planning, enhancing both geometric and photometric reconstruction accuracy. ActiveGAMER also integrates a carefully balanced framework, combining coarse-to-fine exploration, post-refinement, and a global-local keyframe selection strategy to maximize reconstruction completeness and fidelity. Our system autonomously explores and reconstructs environments with state-of-the-art geometric and photometric accuracy and completeness, significantly surpassing existing approaches in both aspects. Extensive evaluations on benchmark datasets such as Replica and MP3D highlight ActiveGAMER’s effectiveness in active mapping tasks. Huangying Zhan, Xiangyu Xu 0004, Qingan Yan, Changjiang Cai, Yi Xu 0002 |
CVPR | 2 |
| 2025 | PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance FieldsabstractIdentifying spatially complete planar primitives from visual data is a crucial task in computer vision. Prior methods are largely restricted to either 2D segment recovery or simplifying 3D structures, even with extensive plane annotations. We present PlanarNeRF, a novel framework capable of detecting dense 3D planes through online learning. Drawing upon the neural field representation, PlanarNeRF brings three major contributions. First, it enhances 3D plane detection with concurrent appearance and geometry knowledge. Second, a lightweight plane fitting module is used to estimate plane parameters. Third, a novel global memory bank structure with an update mechanism is introduced, ensuring consistent cross-frame correspondence. The flexible architecture of PlanarNeRF allows it to function in both 2D-supervised and self-supervised solutions, in each of which it can effectively learn from sparse training signals, significantly improving training efficiency. Through extensive experiments, we demonstrate the effectiveness of PlanarNeRF in various real-world scenarios and remarkable improvement in 3D plane detection over existing works. Zheng Chen 0016, Qingan Yan, Huangying Zhan, Changjiang Cai, Xiangyu Xu 0004, Yuzhong Huang, Ziyue Feng, Yi Xu 0002, Lantao Liu |
ICRA | 3 |
| 2025 | Understanding while Exploring: Semantics-driven Active MappingabstractEffective robotic autonomy in unknown environments demands proactive exploration and precise understanding of both geometry and semantics. In this paper, we propose ActiveSGM, an active semantic mapping framework designed to predict the informativeness of potential observations before execution. Built upon a 3D Gaussian Splatting (3DGS) mapping backbone, our approach employs semantic and geometric uncertainty quantification, coupled with a sparse semantic representation, to guide exploration. By enabling robots to strategically select the most beneficial viewpoints, ActiveSGM efficiently enhances mapping completeness, accuracy, and robustness to noisy semantic data, ultimately supporting more adaptive scene exploration. Our experiments on the Replica and Matterport3D datasets highlight the effectiveness of ActiveSGM in active semantic mapping tasks. Huangying Zhan, Hairong Yin, Yi Xu 0002, Philippos Mordohai |
NeurIPS | 2 |
| 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target ObservationsabstractWe present NARUTO, a neural active reconstruction system that combines a hybrid neural representation with uncertainty learning, enabling high-fidelity surface reconstruction. Our approach leverages a multi-resolution hashgrid as the mapping backbone, chosen for its exceptional convergence speed and capacity to capture high-frequency local features. The centerpiece of our work is the incorporation of an uncertainty learning module that dynamically quantifies reconstruction uncertainty while actively reconstructing the environment. By harnessing learned uncertainty, we propose a novel uncertainty aggregation strategy for goal searching and efficient path planning. Our system autonomously explores by targeting uncertain observations and reconstructs environments with remarkable completeness and fidelity. We also demonstrate the utility of this uncertainty-aware approach by enhancing SOTA neural SLAM systems through an active ray sampling strategy. Extensive evaluations of NARUTO in various environments, using an indoor scene simulator, confirm its superior performance and state-of-the-art status in active reconstruction, as evidenced by its impressive results on benchmark datasets like Replica and MP3D. Project page: oppo-usresearch.github.io/NARUTO-website/ Ziyue Feng, Huangying Zhan, Zheng Chen 0016, Qingan Yan, Xiangyu Xu 0004, Changjiang Cai, Qilun Zhu, Yi Xu 0002 |
CVPR | 2 |
| 2024 | SC-DepthV3: Robust Self-Supervised Monocular Depth Estimation for Dynamic ScenesabstractSelf-supervised monocular depth estimation has shown impressive results in static scenes. It relies on the multi-view consistency assumption for training networks, however, that is violated in dynamic object regions and occlusions. Consequently, existing methods show poor accuracy in dynamic scenes, and the estimated depth map is blurred at object boundaries because they are usually occluded in other training views. In this paper, we propose SC-DepthV3 for addressing the challenges. Specifically, we introduce an external pretrained monocular depth estimation model for generating single-image depth prior, namely pseudo-depth, based on which we propose novel losses to boost self-supervised training. As a result, our model can predict sharp and accurate depth maps, even when training from monocular videos of highly dynamic scenes. We demonstrate the significantly superior performance of our method over previous methods on six challenging datasets, and we provide detailed ablation studies for the proposed terms. Libo Sun 0002, Jiawang Bian, Huangying Zhan, Wei Yin 0006, Ian D. Reid 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Auto-Rectify Network for Unsupervised Indoor Depth EstimationabstractSingle-View depth estimation using the CNNs trained from unlabelled videos has shown significant promise. However, excellent results have mostly been obtained in street-scene driving scenarios, and such methods often fail in other settings, particularly indoor videos taken by handheld devices. In this work, we establish that the complex ego-motions exhibited in handheld settings are a critical obstacle for learning depth. Our fundamental analysis suggests that the rotation behaves as noise during training, as opposed to the translation (baseline) which provides supervision signals. To address the challenge, we propose a data pre-processing method that rectifies training images by removing their relative rotations for effective learning. The significantly improved performance validates our motivation. Towards end-to-end learning without requiring pre-processing, we propose an Auto-Rectify Network with novel loss functions, which can automatically learn to rectify images during training. Consequently, our results outperform the previous unsupervised SOTA method by a large margin on the challenging NYUv2 dataset. We also demonstrate the generalization of our trained model in ScanNet and Make3D, and the universality of our proposed learning method on 7-Scenes and KITTI datasets. Jiawang Bian, Huangying Zhan, Naiyan Wang, Tat-Jun Chin, Chunhua Shen, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | NVSS: High-quality Novel View Selfie SynthesisabstractWe present a novel method to synthesize novel view selfies from a mobile phone captured video. This is challenging due to the inconsistent geometry that is caused by the person’s unavoidable movement. Recent methods reconstruct the whole deformable scene implicitly with a deformation field. We argue that they are inefficient and hard to fit diverse real-world videos. In contrast, we use an explicit reconstruction for generalization and efficiency, where we separately track, reconstruct, and synthesize the foreground and background to overcome the geometry inconsistency. Several novel and effective modules are proposed for better performance and visual results. We demonstrate the advantage of the proposed method against the existing alternatives in a collection of our captured selfie videos with the support of quantitative and qualitative results. Jiawang Bian, Huangying Zhan, Ian D. Reid 0001 |
3DV | 2 |
| 2021 | Unsupervised Scale-Consistent Depth Learning from Video
Jiawang Bian, Huangying Zhan, Naiyan Wang, Le Zhang 0001, Chunhua Shen, Ming-Ming Cheng, Ian D. Reid 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Visual Odometry Revisited: What Should Be Learnt?abstractIn this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance are based on geometry and have to be carefully designed for different application scenarios. Moreover, most monocular systems suffer from scale-drift issue. Some recent deep learning works learn VO in an end-to-end manner but the performance of these deep systems is still not comparable to geometry-based methods. In this work, we revisit the basics of VO and explore the right way for integrating deep learning with epipolar geometry and Perspective-n-Point (PnP) method. Specifically, we train two convolutional neural networks (CNNs) for estimating single-view depths and two-view optical flows as intermediate outputs. With the deep predictions, we design a simple but robust frame-to-frame VO algorithm (DF-VO) which outperforms pure deep learning-based and geometry-based methods. More importantly, our system does not suffer from the scale-drift issue being aided by a scale consistent single-view depth CNN. Extensive experiments on KITTI dataset shows the robustness of our system and a detailed ablation study shows the effect of different factors in our system. Code is available at here: DF-VO. Huangying Zhan, Chamara Saroj Weerasekera, Jiawang Bian, Ian D. Reid 0001 |
ICRA | 1 |
| 2019 | Self-supervised Learning for Single View Depth and Surface Normal EstimationabstractIn this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single image. In contrast to most existing frameworks which represent outdoor scenes as fronto-parallel planes at piece-wise smooth depth, we propose to predict depth with surface orientation while assuming that natural scenes have piece-wise smooth normals. We show that a simple depth-normal consistency as a soft-constraint on the predictions is sufficient and effective for training both these networks simultaneously. The trained normal network provides state-of-the-art predictions while the depth network, relying on much realistic smooth normal assumption, outperforms the traditional self-supervised depth prediction network by a large margin on the KITTI benchmark. Huangying Zhan, Chamara Saroj Weerasekera, Ravi Garg, Ian D. Reid 0001 |
ICRA | 1 |
| 2019 | Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular VideoabstractRecent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static scene assumption in geometric image reconstruction. More significantly, due to lack of proper constraints, networks output scale-inconsistent results over different samples, i.e., the ego-motion network cannot provide full camera trajectories over a long video sequence because of the per-frame scale ambiguity. This paper tackles these challenges by proposing a geometry consistency loss for scale-consistent predictions and an induced self-discovered mask for handling moving objects and occlusions. Since we do not leverage multi-task learning like recent works, our framework is much simpler and more efficient. Comprehensive evaluation results demonstrate that our depth estimator achieves the state-of-the-art performance on the KITTI dataset. Moreover, we show that our ego-motion network is able to predict a globally scale-consistent camera trajectory for long video sequences, and the resulting visual odometry accuracy is competitive with the recent model that is trained using stereo videos. To the best of our knowledge, this is the first work to show that deep networks trained using unlabelled monocular videos can predict globally scale-consistent camera trajectories over a long video sequence. Jiawang Bian, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, Ian D. Reid 0001 |
NeurIPS | 4 |
| 2018 | Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature ReconstructionabstractDespite learning based methods showing promising results in single view depth estimation and visual odometry, most existing approaches treat the tasks in a supervised manner. Recent approaches to single view depth estimation explore the possibility of learning without full supervision via minimizing photometric error. In this paper, we explore the use of stereo sequences for learning depth and visual odometry. The use of stereo sequences enables the use of both spatial (between left-right pairs) and temporal (forward backward) photometric warp error, and constrains the scene depth and camera motion to be in a common, real-world scale. At test time our framework is able to estimate single view depth and two-view odometry from a monocular sequence. We also show how we can improve on a standard photometric warp loss by considering a warp of deep features. We show through extensive experiments that: (i) jointly training for single view depth and visual odometry improves depth prediction because of the additional constraint imposed on depths and achieves competitive results for visual odometry; (ii) deep feature-based warping loss improves upon simple photometric warp loss for both single view depth estimation and visual odometry. Our method outperforms existing learning based methods on the KITTI driving dataset in both tasks. The source code is available at https://github.com/Huangying-Zhan/Depth-VO-Feat. Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera, Kejie Li, Ian D. Reid 0001 |
CVPR | 1 |
| 2018 | Efficient Dense Point Cloud Object Reconstruction Using Deformation Vector Fields
Kejie Li, Trung Pham, Huangying Zhan, Ian D. Reid 0001 |
ECCV (12) | 3 |
| 2017 | Deep learning for 2D scan matching and loop closureabstractAlthough 2D LiDAR based Simultaneous Localization and Mapping (SLAM) is a relatively mature topic nowadays, the loop closure problem remains challenging due to the lack of distinctive features in 2D LiDAR range scans. Existing research can be roughly divided into correlation based approaches e.g. scan-to-submap matching and feature based methods e.g. bag-of-words (BoW). In this paper, we solve loop closure detection and relative pose transformation using 2D LiDAR within an end-to-end Deep Learning framework. The algorithm is verified with simulation data and on an Unmanned Aerial Vehicle (UAV) flying in indoor environment. The loop detection ConvNet alone achieves an accuracy of 98.2% in loop closure detection. With a verification step using the scan matching ConvNet, the false positive rate drops to around 0.001%. The proposed approach processes 6000 pairs of raw LiDAR scans per second on a Nvidia GTX1080 GPU. Huangying Zhan, Ben M. Chen, Ian D. Reid 0001, Gim Hee Lee |
IROS | 2 |