VLDB 2026 Research / reviewers in the wild / expert
Hengkai Guo
dblp:188/3108
· DBLP profile ↗
14ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Systems, architecture and hardware · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video Depth Anything: Consistent Depth Estimation for Super-Long VideosabstractDepth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation models or introducing priors from optical flow and camera poses. Nonetheless, these methods are only applicable to short videos (< 10 seconds) and require a trade-off between quality and computational efficiency. We propose Video Depth Anything for high-quality, consistent depth estimation in super-long videos (over several minutes) without sacrificing efficiency. We base our model on Depth Anything V2 and replace its head with an efficient spatial-temporal head. We design a straightforward yet effective temporal consistency loss by constraining the temporal depth gradient, eliminating the need for additional geometric priors. The model is trained on a joint dataset of video depth and unlabeled images, similar to Depth Anything V2. Moreover, a novel key-frame-based strategy is developed for long video inference. Experiments show that our model can be applied to arbitrarily long videos without compromising quality, consistency, or generalization ability. Comprehensive evaluations on multiple video benchmarks demonstrate that our approach sets a new state-of-the-art in zero-shot video depth estimation. We offer models of different scales to support a range of scenarios, with our smallest model capable of real-time performance at 30 FPS. Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Jiashi Feng, Bingyi Kang |
CVPR | 2 |
| 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Imageabstract3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data. In this work, we introduce a novel framework dubbed ZeroPlane, a Transformer-based model targeting zero-shot 3D plane detection and reconstruction from a single image, over diverse domains and environments. To enable data-driven models across multiple domains, we have curated a large-scale planar benchmark, comprising over 14 datasets and 560,000 high-resolution, dense planar annotations for diverse indoor and outdoor scenes. To address the challenge of achieving desirable planar geometry on multi-dataset training, we propose to disentangle the representation of plane normal and offset, and employ an exemplar-guided, classification-then-regression paradigm to learn plane and offset respectively. Additionally, we employ advanced backbones as image encoder, and present an effective pixel-geometry-enhanced plane embedding module to further facilitate planar reconstruction. Extensive experiments across multiple zero-shot evaluation datasets have demonstrated that our approach significantly outperforms previous methods on both reconstruction accuracy and generalizability, especially over in-the-wild data. Our code and data are available at: https://github.com/jcliu0428/ZeroPlane. Rui Yu 0002, Sili Chen, Sharon X. Huang, Hengkai Guo |
CVPR | 5 |
| 2024 | MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane ReconstructionabstractThis paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and establishes a plane reconstruction pipeline based on monocular geometric cues, resulting in accurate, robust and scalable 3D plane detection and reconstruction in the wild. Specifically, we first leverage large-scale pre-trained neural networks to obtain the depth and surface normals from a single image. These monocular geometric cues are then incorporated into a proximity-guided RANSAC framework to sequentially fit each plane instance. We exploit effective 3D point proximity and model such proximity via a graph within RANSAC to guide the plane fitting from noisy monocular depths, followed by image-level multi-plane joint optimization to improve the consistency among all plane instances. We further design a simple but effective pipeline to extend this single-view solution to sparse-view 3D plane reconstruction. Extensive experiments on a list of datasets demonstrate our superior zero-shot generalizability over baselines, achieving state-of-the-art plane reconstruction performance in a transferring setting. Our code is available at https://github.com/thuzhaowang/MonoPlane. Wang Zhao 0001, Yishu Li, Sili Chen, Sharon X. Huang, Yong-Jin Liu 0001, Hengkai Guo |
IROS | 8 |
| 2022 | ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild
Wang Zhao 0001, Shaohui Liu, Hengkai Guo, Wenping Wang 0001, Yong-Jin Liu 0001 |
ECCV (32) | 3 |
| 2022 | Iterative Feature Matching for Self-Supervised Indoor Depth EstimationabstractIn this paper, we propose an iterative feature matching framework for self-supervised depth estimation in indoor scenes. Conventional methods usually leverage the structure-from-motion supervision to help the photometric optimization escape from the local minima, which have complex ego-motion and large regions with non-texture or repeated-texture. However, the supervision is limited as the reconstruction is usually sparse. To address this, we propose an iterative feature matching framework called IFMNet to jointly learn depths and search for correspondences. With the predicted depths from the previous iteration, we present an online optimized grid searching algorithm to find more accurate correspondences. Given these new correspondences, we compute the triangulated depths and improve the depth network with adaptive bin-wise online hard example mining. Experimental results on the NYU Depth V2 and SceneNet datasets verify the effectiveness of our approach. Yi Wei 0003, Hengkai Guo, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Confidence-based Iterative Solver of Depths and Surface Normals for Deep Multi-view StereoabstractIn this paper, we introduce a deep multi-view stereo (MVS) system that jointly predicts depths, surface normals and per-view confidence maps. The key to our approach is a novel solver that iteratively solves for per-view depth map and normal map by optimizing an energy potential based on the locally planar assumption. Specifically, the algorithm updates depth map by propagating from neigh-boring pixels with slanted planes, and updates normal map with local probabilistic plane fitting. Both two steps are monitored by a customized confidence map. This solver is not only effective as a post-processing tool for plane-based depth refinement and completion, but also differentiable such that it can be efficiently integrated into deep learning pipelines. Our multi-view stereo system employs multiple optimization steps of the solver over the initial prediction of depths and surface normals. The whole system can be trained end-to-end, decoupling the challenging problem of matching pixels within poorly textured regions from the cost-volume based neural network. Experimental results on ScanNet and RGB-D Scenes V2 demonstrate state-of-the-art performance of the proposed deep MVS system on multi-view depth estimation, with our proposed solver consistently improving the depth quality over both conventional and deep learning based MVS pipelines. Code is available at https://github.com/thuzhaowang/idn-solver. Wang Zhao 0001, Shaohui Liu, Yi Wei 0003, Hengkai Guo, Yong-Jin Liu 0001 |
ICCV | 4 |
| 2020 | GPO: Global Plane Optimization for Fast and Accurate Monocular SLAM InitializationabstractInitialization is essential to monocular Simultaneous Localization and Mapping (SLAM) problems. This paper focuses on a novel initialization method for monocular SLAM based on planar features. The algorithm starts by homography estimation in a sliding window. It then proceeds to a global plane optimization (GPO) to obtain camera poses and the plane normal. 3D points can be recovered using planar constraints without triangulation. The proposed method fully exploits the plane information from multiple frames and avoids the ambiguities in homography decomposition. We validate our algorithm on the collected chessboard dataset against baseline implementations and present extensive analysis. Experimental results show that our method outperforms the ne-tuned baselines in both accuracy and real-time. Sicong Du, Hengkai Guo, Yilun Lin 0002, Xiangbing Meng, Linfu Wen, Fei-Yue Wang 0001 |
ICRA | 2 |
| 2020 | Geometric Pretraining for Monocular Depth EstimationabstractImageNet-pretrained networks have been widely used in transfer learning for monocular depth estimation. These pretrained networks are trained with classification losses for which only semantic information is exploited while spatial information is ignored. However, both semantic and spatial information is important for per-pixel depth estimation. In this paper, we design a novel self-supervised geometric pretraining task that is tailored for monocular depth estimation using uncalibrated videos. The designed task decouples the structure information from input videos by a simple yet effective conditional autoencoder-decoder structure. Using almost unlimited videos from the internet, networks are pretrained to capture a variety of structures of the scene and can be easily transferred to depth estimation tasks using calibrated images. Extensive experiments are used to demonstrate that the proposed geometric-pretrained networks perform better than ImageNet-pretrained networks in terms of accuracy, few-shot learning and generalization ability. Using existing learning methods, geometric-transferred networks achieve new state-of-the-art results by a large margin. The pretrained networks will be open source soon1. Hengkai Guo, Linfu Wen, Shaojie Shen |
ICRA | 3 |
| 2020 | Pose guided structured region ensemble network for cascaded hand pose estimation
Xinghao Chen 0001, Guijin Wang, Hengkai Guo, Cairong Zhang |
Neurocomputing | 3 |
| 2018 | Region ensemble network: Towards good practices for deep 3D hand pose estimation
Guijin Wang, Xinghao Chen 0001, Hengkai Guo, Cairong Zhang |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Motion feature augmented recurrent neural network for skeleton-based dynamic hand gesture recognitionabstractDynamic hand gesture recognition has attracted increasing interests because of its importance for human computer interaction. In this paper, we propose a new motion feature augmented recurrent neural network for skeleton-based dynamic hand gesture recognition. Finger motion features are extracted to describe finger movements and global motion features are utilized to represent the global movement of hand skeleton. These motion features are then fed into a bidirectional recurrent neural network (RNN) along with the skeleton sequence, which can augment the motion features for RNN and improve the classification performance. Experiments demonstrate that our proposed method is effective and outperforms start-of-the-art methods. Xinghao Chen 0001, Hengkai Guo, Guijin Wang, Li Zhang 0023 |
ICIP | 2 |
| 2017 | Region ensemble network: Improving convolutional network for hand pose estimationabstractHand pose estimation from monocular depth images is an important and challenging problem for human-computer interaction. Recently deep convolutional networks (ConvNet) with sophisticated design have been employed to address it, but the improvement over traditional methods is not so apparent. To promote the performance of directly 3D coordinate regression, we propose a tree-structured Region Ensemble Network (REN), which partitions the convolution outputs into regions and integrates the results from multiple regressors on each regions. Compared with multi-model ensemble, our model is completely end-to-end training. The experimental results demonstrate that our approach achieves the best performance among state-of-the-arts on two public datasets. Hengkai Guo, Guijin Wang, Xinghao Chen 0001, Cairong Zhang, Fei Qiao, Huazhong Yang |
ICIP | 1 |
| 2017 | Two-stream binocular network: Accurate near field finger detection based on binocular imagesabstractFingertip detection plays an important role in human computer interaction. Previous works transform binocular images into depth images. Then depth-based hand pose estimation methods are used to predict 3D positions of fingertips. Different from previous works, we propose a new framework, named Two-Stream Binocular Network (TSBnet) to detect fingertips from binocular images directly. TSBnet first shares convolutional layers for low level features of right and left images. Then it extracts high level features in two-stream convolutional networks separately. Further, we add a new layer: binocular distance measurement layer to improve performance of our model. To verify our scheme, we build a binocular hand image dataset, containing about 117k pairs of images in training set and 10k pairs of images in test set. Our methods achieve an average error of 10.9mm on our test set, outperforming previous work by 5.9mm (relatively 35.1%). Guijin Wang, Cairong Zhang, Hengkai Guo, Xinghao Chen 0001, Huazhong Yang |
VCIP | 4 |
| 2016 | Accurate fingertip detection from binocular mask imagesabstractAccurate fingertip detection is important for hand-based human computer interaction. Different from prior methods based on depth image, in this paper we propose a novel scheme for accurate fingertip detection from binocular mask images without explicitly computing the depth map. To demonstrate our proposed scheme, we build a new hand dataset containing synthetic and real binocular images. A deep convolutional neural network (CNN) is utilized as a baseline method to demonstrate the proposed scheme. The mask images are extracted from binocular images and fed into the CNN to predict the 3D positions of fingertips and palm center. Experiments show that this method achieves the mean error of 8.30mm on synthetic data and 4.64mm on real data, which is promosing for accurate 3D interaction. The proposed scheme runs at 130fps on a CPU and is promising for real-time applications. Xinghao Chen 0001, Guijin Wang, Hengkai Guo |
VCIP | 3 |