VLDB 2026 Research / reviewers in the wild / expert
Zhicheng Wang 0022
dblp:78/1664-22
· DBLP profile ↗
17ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HOLO: Holistic Lightweight Optimization for Scene Understanding with Auto-Annotation and Multimodal LearningabstractVision-language models (VLMs) have achieved remarkable success in various domains. However, their application to 3D scene understanding remains largely underexplored. Existing 3D VLMs predominantly focus on object-level tasks and often emphasize instance-centric representations within a scene, lacking holistic scene-level descriptions. In this work, we propose an automated annotation framework that leverages multi-view images to partition 3D scenes into localized point cloud sub-regions, which are then enriched with precise semantic information-all without any manual intervention. We processed ScanNet v2 and ScanNet++ to construct SceneCap, a large-scale dataset designed for scene-level description. To demonstrate the benefits of our framework for scene understanding, we introduce Natural Interactive Universal Multimodal Observer, NIUMO-LLM, a lightweight yet high-performing model training on SceneCap. We further demonstrate that NIUMO-LLM achieves state-of-the-art (SOTA) performance on both scene description benchmarks and object-level tasks, requiring only 12 hours of training on a single NVIDIA A800 GPU. This design significantly reduces computational demands, lowering the barrier for MLLM-related research. Xiaoyun Hu, Xiaohan Yan, Nan Wang 0041, Zhicheng Wang 0022 |
WACV | 5 |
| 2025 | Semantic-Guided Gaussian Splatting with Deferred RenderingabstractRevolutionizing novel view synthesis, 3D Gaussian Splatting has unlocked new horizons in 3D visual representation. Despite the efficiency and impressive rendering capabilities of GS, the accurate inverse rendering of reflectant non-Lambertian surfaces remains a significant challenge, particularly in the context of diverse reflective materials settings, leading to inconsistent renderings and undermining the technology’s potential in applications ranging from digital asserts production to virtual reality. We propose Semantic-Guided Gaussian Splatting (SGGS), which aims to address this challenge by leveraging the capabilities of semantic features derived from cutting-edge 2D foundation models, revolutionizing material properties optimization for Gaussians. By integrating this high-level understanding, we enhance the model’s resilience against reflective surfaces and significantly improve multi-view consistency, which is a crucial step towards seamless immersive experiences. Our experiments systematically demonstrate that SGGS outperforms previous methods in terms of both rendering quality and geometry. Nan Wang 0041, Xiaohan Yan, Zhicheng Wang 0022 |
ICASSP | 4 |
| 2025 | PUGS: Zero-Shot Physical Understanding with Gaussian SplattingabstractCurrent robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts various physical properties in a zero-shot manner. We propose two techniques during the reconstruction phase: a geometryaware regularization loss function to improve the shape quality and a region-aware feature contrastive loss function to promote region affinity. Two other new techniques are designed during inference: a feature-based property propagation module and a volume integration module tailored for the Gaussian representation. Our framework is named as zero-shot physical understanding with Gaussian splatting, or PUGS. PUGS achieves new state-of-the-art results on the standard benchmark of ABO-500 mass prediction. We provide extensive quantitative ablations and qualitative visualization to demonstrate the mechanism of our designs. We show the proposed methodology can help address challenging real-world grasping tasks. Our codes, data, and models are available at https://github.com/EverNorif/PUGS Yinghao Shuai, Yuantao Chen, Zijian Jiang, Nan Wang 0041, Jv Zheng, Jianzhu Ma, Meng Yang 0035, Zhicheng Wang 0022, Wenbo Ding 0001, Hao Zhao 0002 |
ICRA | 10 |
| 2025 | RE0: Recognize Everything with 3D Zero-Shot Instance SegmentationabstractRecognizing objects in the 3D world is a significant challenge for robotics. Due to the lack of high-quality 3D data, directly training a general-purpose segmentation model in 3D is almost infeasible. Meanwhile, vision foundation models (VFM) have revolutionized the 2D computer vision field with outstanding performance, making the use of VFM to assist 3D perception a promising direction. However, most existing VFM-assisted methods do not effectively address the 2D-3D inconsistency problem or adequately provide corresponding semantic information for 3D instance objects. To address these two issues, this paper introduces a novel framework for 3D zero-shot instance segmentation called RE0. For the given 3D point clouds and multi-view RGB-D images with poses, we leverage the 3D geometric information, projection relationships, and CLIP semantic features. Specifically, we utilize CropFormer to extract mask information from multi-view posed images, combined with projection relationships to assign point-level labels to each point in the point cloud, and achieve instance-level consistency through inter-frame information interaction. Then, we employ projection relationships again to assign CLIP semantic features to the point cloud and achieve aggregation of small-scale point clouds. Notably, RE0 does not require any additional training and can be implemented by supporting only one inference of CropFormer and one inference of CLIP. Experiments on ScanNet200 and ScanNet++ show that our method achieves higher quality segmentation than the previous zero-shot methods. Our codes and demos are available at https://recognizeeverything.github.io/, with only one RTX 3090 GPU required. Xiaohan Yan, Zijian Jiang, Yinghao Shuai, Nan Wang 0041, Wenbo Ji, Jinyu He, Zhicheng Wang 0022 |
ICRA | 10 |
| 2024 | AttenPoint: Exploring Point Cloud Segmentation Through Attention-Based Modules
Xiaohan Yan, Nan Wang 0041, Zhicheng Wang 0022 |
PRCV (6) | 5 |
| 2022 | Digging Into Normal Incorporated Stereo MatchingabstractDespite the remarkable progress facilitated by learning-based stereo matching algorithms, disparity estimation in low-texture, occluded, and bordered regions still remain bottlenecks that limit the performance. To tackle these challenges, geometric guidance like plane information is necessary as it provides intuitive guidance about disparity consistency and affinity similarity. In this paper, we propose a normal incorporated joint learning that framework consisting of two specific modules named non-local disparity propagation(NDP) and affinity-aware residual learning(ARL). The estimated normal map is first utilized for calculating a non-local affinity matrix as well as a non-local offset to perform spatial propagation at the disparity level. To enhance geometric consistency, especially in low-texture regions, the estimated normal map is then leveraged to calculate a local affinity matrix which provides the residual learning with information about where the correction should refer and thus improve the residual learning efficiency. Extensive experiments on several public datasets including Scene Flow, KITTI 2015, and Middlebury 2014 validate the effectiveness of our proposed method. By the time we finished this work, our approach ranked 1st for stereo matching across foreground pixels on the KITTI 2015 dataset and 3rd on the Scene Flow dataset among all the published works. Zihua Liu, Songyan Zhang, Zhicheng Wang 0022, Masatoshi Okutomi |
ACM Multimedia | 3 |
| 2021 | EDNet: Efficient Disparity Estimation With Cost Volume Combination and Attention-Based Spatial ResidualabstractExisting state-of-the-art disparity estimation works mostly leverage the 4D concatenation volume and construct a very deep 3D convolution neural network (CNN) for disparity regression, which is inefficient due to the high memory consumption and slow inference speed. In this paper, we propose a network named EDNet for efficient disparity estimation. Firstly, we construct a combined volume which incorporates contextual information from the squeezed concatenation volume and feature similarity measurement from the correlation volume. The combined volume can be next aggregated by 2D convolutions which are faster and require less memory than 3D convolutions. Secondly, we propose an attention-based spatial residual module to generate attention-aware residual features. The attention mechanism is applied to provide intuitive spatial evidence about inaccurate regions with the help of error maps at multiple scales and thus improve the residual learning efficiency. Extensive experiments on the Scene Flow and KITTI datasets show that EDNet outperforms the previous 3D CNN based works and achieves state-of-the-art performance with significantly faster speed and less memory consumption. Songyan Zhang, Zhicheng Wang 0022, Qiang Wang 0022, Jinshuo Zhang, Xiaowen Chu 0001 |
CVPR | 2 |
| 2016 | Context-Aware Gaussian Fields for Non-rigid Point Set RegistrationabstractPoint set registration (PSR) is a fundamental problem in computer vision and pattern recognition, and it has been successfully applied to many applications. Although widely used, existing PSR methods cannot align point sets robustly under degradations, such as deformation, noise, occlusion, outlier, rotation, and multi-view changes. This paper proposes context-aware Gaussian fields (CA-LapGF) for nonrigid PSR subject to global rigid and local non-rigid geometric constraints, where a laplacian regularized term is added to preserve the intrinsic geometry of the transformed set. CA-LapGF uses a robust objective function and the quasi-Newton algorithm to estimate the likely correspondences, and the non-rigid transformation parameters between two point sets iteratively. The CA-LapGF can estimate non-rigid transformations, which are mapped to reproducing kernel Hilbert spaces, accurately and robustly in the presence of degradations. Experimental results on synthetic and real images reveal that how CA-LapGF outperforms state-of-the-art algorithms for non-rigid PSR. Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002, Qiangqiang Zhou |
CVPR | 2 |
| 2016 | Salient object detection using biogeography-based optimization to combine features
Zhicheng Wang 0022, Xiaobei Wu |
Appl. Intell. | 1 |
| 2016 | Learning coherent vector fields for robust point matching under manifold regularization
Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002, Yingchun Ren |
Neurocomputing | 2 |
| 2016 | Removing mismatches for retinal image registration via multi-attribute-driven regularized mixture model
Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002, Qiangqiang Zhou |
Inf. Sci. | 2 |
| 2016 | An automatic panoramic image mosaic method based on graph model
Zhicheng Wang 0022, Yufei Chen 0002, Zewei Zhu |
Multim. Tools Appl. | 1 |
| 2015 | Contour-Based Plant Leaf Image Segmentation Using Visual Saliency
Qiangqiang Zhou, Zhicheng Wang 0022, Yufei Chen 0002 |
ICIG (2) | 2 |
| 2015 | Fuzzy Correspondences and Kernel Density Estimation for Contaminated Point Set RegistrationabstractPoint set registration problem is challenging to solve in the presence of outliers. In this paper, we proposed a registration method based on fuzzy correspondences and kernel density estimation. The main idea of our method is that the moving point set consists of inliers represented using a mixture of Gaussian, and outliers represented via an additional uniform distribution, then we use the fuzzy correspondences to estimate the Gaussian elements in the mixture model. There are four parts of the paper: we formulate the contaminated point set registration problem as a mixture model according to the well known Gaussian mixture model (GMM) based method firstly. Secondly, Gaussian elements are estimated by fuzzy correspondences to increase the registration accuracy efficiently. Thirdly, the optimal transformation between two contaminated point sets is expressed by representation theorem, and solved by EM algorithm iteratively. Finally, we compare our proposed method with several state-of-the-art methods, and the results show that our method gets better performances than the other methods in most tested scenarios. Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002 |
SMC | 2 |
| 2015 | A robust non-rigid point set registration method based on asymmetric gaussian representation
Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002 |
Comput. Vis. Image Underst. | 2 |
| 2014 | Robust Point Matching Using Mixture of Asymmetric Gaussians for Nonrigid Transformation
Gang Wang 0008, Zhicheng Wang 0022, Qiangqiang Zhou |
ACCV (4) | 2 |
| 2007 | On Computation Complexity of the Concurrently Enabled Transition Set Problem
Zhicheng Wang 0022, Shumei Wang |
TAMC | 3 |