VLDB 2026 Research / reviewers in the wild / expert
Haiping Wang 0004
dblp:68/7989-4
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-8370-4585ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 83% Vision and language · 7% Generative modeling · 6% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
point cloud registration |
2.7 | 4 | 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators · ICLR 2024 RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local Rotations · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History Reweighting · CVPR 2023 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
1.0 | 1 | 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting · AAAI 2026 |
Computer vision › 3D vision
3d scene understanding |
1.0 | 1 | 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting · AAAI 2026 |
Computer vision › 3D vision
neural rendering |
1.0 | 1 | 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting · AAAI 2026 |
Computer vision › 3D vision › 3d scene understanding
open-vocabulary 3d scene understanding |
1.0 | 1 | 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting · AAAI 2026 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene Reconstruction · ICCV 2025 |
Computer vision › 3D vision › 3d scene understanding
3d visual grounding |
0.9 | 1 | 2025 | CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMs · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMs · ICLR 2025 |
Computer vision › 3D vision
novel view synthesis |
0.9 | 1 | 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene Reconstruction · ICCV 2025 |
Computer vision › 3D vision › 3d scene reconstruction
single-view scene reconstruction |
0.9 | 1 | 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene Reconstruction · ICCV 2025 |
Computer vision › 3D vision › image registration › multimodal registration
image-to-point cloud registration |
0.8 | 1 | 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators · ICLR 2024 |
Computer vision › 3D vision
feature matching |
0.7 | 1 | 2023 | RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local Rotations · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision › point cloud registration
multi-view registration |
0.7 | 1 | 2023 | Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History Reweighting · CVPR 2023 |
Robotics › Robot navigation and mapping › SLAM › graph optimization
pose graph optimization |
0.7 | 1 | 2023 | Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History Reweighting · CVPR 2023 |
Computer vision › 3D vision › local feature descriptor
3d local descriptors |
0.6 | 1 | 2022 | You Only Hypothesize Once: Point Cloud Registration with Rotation-equivariant Descriptors · ACM Multimedia 2022 |
Computer vision › 3D vision › local feature descriptor
rotation equivariant descriptor |
0.6 | 1 | 2022 | You Only Hypothesize Once: Point Cloud Registration with Rotation-equivariant Descriptors · ACM Multimedia 2022 |
Machine learning › Generative modeling
diffusion model |
0.5 | 2 | 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene Reconstruction · ICCV 2025 FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators · ICLR 2024 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › image restoration
diffusion-based inpainting |
0.3 | 1 | 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene Reconstruction · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features |
0.2 | 1 | 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators · ICLR 2024 |
Geometric modeling and processing
correspondence estimation |
0.2 | 1 | 2022 | You Only Hypothesize Once: Point Cloud Registration with Rotation-equivariant Descriptors · ACM Multimedia 2022 |
Methods — techniques the papers use, named apart from their topics
RANSAC · 1.8knowledge distillation · 1.0gaussian splatting · 1.0SAM · 1.0multiview consistency sampling · 0.9multi-view reconstruction · 0.9large language model · 0.9diffusion-based RGB-D inpainting · 0.9monocular depth estimation · 0.8diffusion feature extraction · 0.8group equivariant feature learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splattingabstract3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary queries for renderings on arbitrary viewpoints. The main challenge of distilling 2D features for 3D fields lies in the multiview inconsistency of extracted 2D features, which provides unstable supervision for the 3D feature field. GAGS addresses this challenge with two novel strategies. First, GAGS associates the prompt point density of SAM with the camera distances to scene objects, which significantly improves the multiview consistency of segmentation results. Second, GAGS further decodes a granularity factor to guide the distillation process and this granularity factor can be learned in a unsupervised manner to only select the multiview consistent 2D features in the distillation process. Experimental results on two datasets show that GAGS improves visual grounding accuracy by an average of 10.9% and semantic segmentation accuracy by an average of 7.0%, with an inference speed 2× faster than baseline methods. Yuning Peng, Haiping Wang 0004, Yuan Liu 0025, Chenglu Wen, Zhen Dong 0005, Bisheng Yang |
AAAI | 2 |
| 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionabstractIn this paper, we propose VistaDream a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most existing methods only concentrate on building the consistency between the input image and the generated images while losing the consistency between the generated images. VistaDream addresses this problem by a two-stage pipeline. In the first stage, VistaDream begins with building a global coarse 3D scaffold by zooming out a little step with inpainted boundaries and an estimated depth map. Then, on this global scaffold, we use iterative diffusion-based RGB-D inpainting to generate novel-view images to inpaint the holes of the scaffold. In the second stage, we further enhance the consistency between the generated novel-view images by a novel training-free Multiview Consistency Sampling (MCS) that introduces multi-view consistency constraints in the reverse sampling process of diffusion models. Experimental results demonstrate that without training or fine-tuning existing diffusion models, VistaDream achieves consistent and high-quality novel view synthesis using just single-view images and outperforms baseline methods by a large margin. The code, videos, and interactive demos are available at https://vistadream-project-page.github.io/. Haiping Wang 0004, Yuan Liu 0025, Ziwei Liu 0002, Wenping Wang 0001, Zhen Dong 0005, Bisheng Yang |
ICCV | 1 |
| 2025 | CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMsabstractIn this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point cloud remains an open problem. Previous 3D visual grounding system mainly concentrates on localizing an object in an image or a small-scale point cloud, which is not accurate and efficient enough to scale up to a city-scale point cloud. We address this problem with a multi-modality LLM which consists of two stages, a coarse localization and a fine-grained matching. Given the text descriptions, the coarse localization stage locates possible regions on a projected 2D map of the point cloud while the fine-grained matching stage accurately determines the most matched object in these possible regions. We conduct experiments on the CityRefer dataset and a new synthetic dataset annotated by us, both of which demonstrate our method can produce accurate 3D visual grounding on a city-scale 3D point cloud. Haiping Wang 0004, Yuan Liu 0025, Zhiyang Dou, Yuexin Ma, Sibei Yang, Wenping Wang 0001, Zhen Dong 0005, Bisheng Yang |
ICLR | 2 |
| 2025 | WHU-Synthetic: A Synthetic Perception Dataset for 3-D Multitask Model ResearchabstractEnd-to-end models capable of handling multiple subtasks in parallel have become a new trend, thereby presenting significant challenges and opportunities for the integration of multiple tasks within the domain of 3-D vision. The limitations of 3-D data acquisition conditions have not only restricted the exploration of many innovative research problems but have also caused existing 3-D datasets to predominantly focus on single tasks. This has resulted in a lack of systematic approaches and theoretical frameworks for 3-D multitask learning, with most efforts merely serving as auxiliary support to the primary task. In this article, we introduce WHU-Synthetic, a large-scale 3-D synthetic perception dataset designed for multitask learning, from the initial data augmentation (upsampling and depth completion), through scene understanding (segmentation), to macrolevel tasks (place recognition and 3-D reconstruction). Collected in the same environmental domain, we ensure inherent alignment across subtasks to construct multitask models without separate training methods. In addition, we implement several novel settings, making it possible to realize certain ideas that are difficult to achieve in real-world scenarios. This supports more adaptive and robust multitask perception tasks, such as sampling on city-level models, providing point clouds with different densities, and simulating temporal changes. Using our dataset, we conduct several experiments to investigate mutual benefits between subtasks, revealing new observations, challenges, and opportunities for future research. The dataset is accessible at:https://github.com/WHU-USI3DV/WHU-Synthetic. Chen Long, Conglang Zhang, Boheng Li, Haiping Wang 0004, Zhe Chen 0028, Zhen Dong 0005 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsabstractMatching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a $3.0\times$ higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts. The code and additional results are available at \url{https://whu-usi3dv.github.io/FreeReg/}. Haiping Wang 0004, Yuan Liu 0025, Bing Wang 0013, Yujing Sun 0001, Zhen Dong 0005, Wenping Wang 0001, Bisheng Yang |
ICLR | 1 |
| 2024 | A Novel Method for Registration of MLS and Stereo Reconstructed Point CloudsabstractCross-source point cloud registration is a prerequisite for effectively leveraging the complementary information of multiple 3D sensors. However, existing point cloud registration methods have primarily focused on the registration of mono-source point clouds and typically fail to register cross-source data with varying noise patterns and capture characteristics. In this paper, we present a new algorithm for cross-source point cloud registration between MLS point clouds and stereo-reconstructed point clouds. Our method has two key designs. Firstly, we design a novel descriptor with in-plane rotation-equivariance by leveraging the accessible gravity prior, yielding strong descriptiveness, better robustness, and improved efficiency. Secondly, based on the noise pattern of stereo-reconstructed point clouds, a novel disparity-weighted correspondence scoring strategy is proposed to strengthen the registration accuracy. In comparison to existing registration baselines, our method achieves a 32.6% higher Registration Recall on cross-source datasets of KITTI and KITTI-360 and a 23.1% higher Registration Recall on mono-source datasets of KITTI. Notably, our method also outperforms RANSAC-based methods in terms of computational efficiency with a 10× ~ 70× speedup. The source code and datasets have been available at https://github.com/WHU-USI3DV/MSReg. Haiping Wang 0004, Zhen Dong 0005, Yuan Liu 0025, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History ReweightingabstractIn this paper, we present a new method for the multi-view registration of point cloud. Previous multiview registration methods rely on exhaustive pairwise registration to construct a densely-connected pose graph and apply Iteratively Reweighted Least Square (IRLS) on the pose graph to compute the scan poses. However, constructing a densely-connected graph is time-consuming and contains lots of outlier edges, which makes the subsequent IRLS struggle to find correct poses. To address the above problems, we first propose to use a neural network to estimate the overlap between scan pairs, which enables us to construct a sparse but reliable pose graph. Then, we design a novel history reweighting function in the IRLS scheme, which has strong robustness to outlier edges on the graph. In comparison with existing multiview registration methods, our method achieves 11% higher registration recall on the 3DMatch dataset and ~ 13% lower registration errors on the ScanNet dataset while reducing ~ 70% required pairwise registrations. Comprehensive ablation studies are conducted to demonstrate the effectiveness of our designs. The source code is available at https://github.com/WHU-USI3DV/SGHR. Haiping Wang 0004, Yuan Liu 0025, Zhen Dong 0005, Yulan Guo, Yu-Shen Liu, Wenping Wang 0001, Bisheng Yang |
CVPR | 1 |
| 2023 | RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local RotationsabstractWe present RoReg, a novel point cloud registration framework that fully exploits oriented descriptors and estimated local rotations in the whole registration pipeline. Previous methods mainly focus on extracting rotation-invariant descriptors for registration but unanimously neglect the orientations of descriptors. In this paper, we show that the oriented descriptors and the estimated local rotations are very useful in the whole registration pipeline, including feature description, feature detection, feature matching, and transformation estimation. Consequently, we design a novel oriented descriptor RoReg-Desc and apply RoReg-Desc to estimate the local rotations. Such estimated local rotations enable us to develop a rotation-guided detector, a rotation coherence matcher, and a one-shot-estimation RANSAC, all of which greatly improve the registration performance. Extensive experiments demonstrate that RoReg achieves state-of-the-art performance on the widely-used 3DMatch and 3DLoMatch datasets, and also generalizes well to the outdoor ETH dataset. In particular, we also provide in-depth analysis on each component of RoReg, validating the improvements brought by oriented descriptors and the estimated local rotations. Source code and supplementary material are available at https://github.com/HpWang-whu/RoReg. Haiping Wang 0004, Yuan Liu 0025, Qingyong Hu, Bing Wang 0013, Zhen Dong 0005, Yulan Guo, Wenping Wang 0001, Bisheng Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | You Only Hypothesize Once: Point Cloud Registration with Rotation-equivariant DescriptorsabstractIn this paper, we propose a novel local descriptor-based framework, called You Only Hypothesize Once (YOHO), for the registration of two unaligned point clouds. In contrast to most existing local descriptors which rely on a fragile local reference frame to gain rotation invariance, the proposed descriptor achieves the rotation invariance by recent technologies of group equivariant feature learning, which brings more robustness to point density and noise. Meanwhile, the descriptor in YOHO also has a rotation-equivariant part, which enables us to estimate the registration from just one correspondence hypothesis. Such property reduces the searching space for feasible transformations, thus greatly improving both the accuracy and the efficiency of YOHO. Extensive experiments show that YOHO achieves superior performances with much fewer needed RANSAC iterations on four widely-used datasets, the 3DMatch/3DLoMatch datasets, the ETH dataset and the WHU-TLS dataset. More details are shown in our project page: https://hpwang-whu.github.io/YOHO/. Haiping Wang 0004, Yuan Liu 0025, Zhen Dong 0005, Wenping Wang 0001 |
ACM Multimedia | 1 |