Zhoutao Wang

dblp:212/5839 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0002-5798-2407ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PrimitiveGroup: A Fast Primitive Segmentation Framework on Industrial Point Clouds
abstract
Efficiency is crucial for primitive segmentation in industrial applications. Previous state-of-the-art (SOTA) methods suffer from low efficiency due to reliance on time-consuming feature clustering. To circumvent the need for slowly grouping points in a high-dimensional space, we propose a fast framework named PrimitiveGroup (PG), which makes full use of spatial relations and various feature consistency to group points efficiently in the 3-D space. Moreover, to improve the accuracy of point grouping in the 3-D space, we also introduce an adaptive long-range offset prediction module which expands the neighborhood perception range and adaptively focuses on those neighborhoods exhibiting higher semantic and instance correlation. A hybrid consistency aggregation that considers not only spatial distance and semantic constraints but also other geometric consistency of each point is proposed to decompose mixed points belonging to primitives with overlapping or adjacent centroids. Experimental results on the ABCParts and the ANSI datasets show that PG not only achieves competitive performance compared to recent SOTA methods but also operates as the fastest deep learning method in primitive segmentation, 22 times faster than the existing fastest method. Meanwhile, PG achieves promising robustness on noisy point clouds and industrial real scans.
Anyi Huang, Zhoutao Wang, Zikuan Li, Mingqiang Wei, Jun Wang 0039
IEEE Trans. Ind. Informatics2
2025 DPPNet: A Depth Pixel-Wise Potential-Aware Network for RGB-D Salient Object Detection
abstract
Depth cues are essential for visual perception tasks like Salient Object Detection (SOD). Due to varying depth reliability across scenes, some researchers propose evaluating the overall quality of the depth maps and discarding the less reliable ones to avoid contamination. However, these methods often fail to fully utilize valuable information in depth maps, leading to sub-optimal performance particularly when depth quality is unreliable. Since low-quality depth maps still contain useful information that potentially improves model performance, we propose a Depth Pixel-wise Potential-aware Network to leverage these depth cues effectively. This network includes two novel components designed: 1) A learning strategy for explicitly modeling the confidence of each depth pixel to assist the model in locating valid information in the depth map. 2) A cross-modal adaptive multiple fusion module that fuses features from both RGB and depth modalities. It aims to mitigate the contamination effect of unreliable depth maps and fully exploit the benefits of multiple fusion strategies. Experimental results show that on four publicly available datasets, our method outperforms 17 mainstream methods on various evaluation metrics.
Junbin Yuan, Zhoutao Wang, Qingzhen Xu, Bharadwaj Veeravalli, Xulei Yang
IEEE Trans. Multim.3
2024 FlyCore: Fast Low-Frequency Coarse Registration of Large-Scale Outdoor LiDAR Point Clouds
abstract
Fast and accurate registration of outdoor LiDAR point clouds poses a considerable challenge for their large-scale (e.g., 300 K points) and intricate (e.g., noise and outliers) distributions. In this article, we present a fast low-frequency coarse registration method for large-scale outdoor LiDAR point clouds, dubbed FlyCore. Different from existing methods, FlyCore is very fast for practical applications and bridges current refinement registration methods smoothly for their accuracy improvements. Specifically, we first construct spherical feature spaces for a pair of point clouds based on their keypoints and saliency uncertainties independently. Then, we perform harmonic decomposition on these spherical feature spaces, utilizing the low-frequency components of spherical harmonics (SHs) to implement point cloud registration. FlyCore demonstrates less sensitivity to noise and outliers compared to feature-based registration techniques. Also, FlyCore achieves exceptionally low time complexity by eliminating the need for feature matching and iterative procedures, ensuring fine alignment with only a few iterations. Experimental validations, utilizing two extensive LiDAR datasets featuring urban and natural scenarios, confirm the effectiveness and accuracy improvement of existing fine registration methods facilitated by our FlyCore.
Zikuan Li, Kaijun Zhang, Zhoutao Wang, Sibo Wu, Xiao-Ping Zhang 0002, Mingqiang Wei, Jun Wang 0039
IEEE Trans. Geosci. Remote. Sens.3
2022 MODNet: Multi-offset Point Cloud Denoising Network Customized for Multi-scale Patches
abstract
Abstract The intricacy of 3D surfaces often results cutting‐edge point cloud denoising (PCD) models in surface degradation including remnant noise, wrongly‐removed geometric details. Although using multi‐scale patches to encode the geometry of a point has become the common wisdom in PCD, we find that simple aggregation of extracted multi‐scale features can not adaptively utilize the appropriate scale information according to the geometric information around noisy points. It leads to surface degradation, especially for points close to edges and points on complex curved surfaces. We raise an intriguing question – if employing multi‐scale geometric perception information to guide the network to utilize multi‐scale information, can eliminate the severe surface degradation problem? To answer it, we propose a Multi‐offset Denoising Network (MODNet) customized for multi‐scale patches. First, we extract the low‐level feature of three scales patches by patch feature encoders. Second, a multi‐scale perception module is designed to embed multi‐scale geometric information for each scale feature and regress multi‐scale weights to guide a multi‐offset denoising displacement. Third, a multi‐offset decoder regresses three scale offsets, which are guided by the multi‐scale weights to predict the final displacement by weighting them adaptively. Experiments demonstrate that our method achieves new state‐of‐the‐art performance on both synthetic and real‐scanned datasets. Our code is publicly available at https://github.com/hay-001/MODNet .
Anyi Huang, Qian Xie 0001, Zhoutao Wang, Dening Lu, Mingqiang Wei, Jun Wang 0039
Comput. Graph. Forum3
2022 Multi-feature Fusion VoteNet for 3D Object Detection
abstract
In this article, we propose a Multi-feature Fusion VoteNet (MFFVoteNet) framework for improving the 3D object detection performance in cluttered and heavily occluded scenes. Our method takes the point cloud and the synchronized RGB image as inputs to provide object detection results in 3D space. Our detection architecture is built on VoteNet with three key designs. First, we augment the VoteNet input with point color information to enhance the difference of various instances in a scene. Next, we integrate an image feature module into the VoteNet to provide a strong object class signal that can facilitate deterministic detections in occlusion. Moreover, we propose a Projection Non-Maximum Suppression (PNMS) method in 3D object detection to eliminate redundant proposals and hence provide more accurate positioning of 3D objects. We evaluate the proposed MFFVoteNet on two challenging 3D object detection datasets, i.e., ScanNetv2 and SUN RGB-D. Extensive experiments show that our framework can effectively improve the performance of 3D object detection.
Zhoutao Wang, Qian Xie 0001, Mingqiang Wei, Kun Long, Jun Wang 0039
ACM Trans. Multim. Comput. Commun. Appl.1
2021 MLVSNet: Multi-level Voting Siamese Network for 3D Visual Tracking
abstract
Benefiting from the excellent performance of Siamese-based trackers, huge progress on 2D visual tracking has been achieved. However, 3D visual tracking is still under-explored. Inspired by the idea of Hough voting in 3D object detection, in this paper, we propose a Multi-level Voting Siamese Network (MLVSNet) for 3D visual tracking from outdoor point cloud sequences. To deal with sparsity in outdoor 3D point clouds, we propose to perform Hough voting on multi-level features to get more vote centers and retain more useful information, instead of voting only on the fi-nal level feature as in previous methods. We also design an efficient and lightweight Target-Guided Attention (TGA) module to transfer the target information and highlight the target points in the search area. Moreover, we propose a Vote-cluster Feature Enhancement (VFE) module to exploit the relationships between different vote clusters. Extensive experiments on the 3D tracking benchmark of KITTI dataset demonstrate that our MLVSNet outperforms state-of-the-art methods with significant margins. Code will be available at https://github.com/CodeWZT/MLVSNet.
Zhoutao Wang, Qian Xie 0001, Yukun Lai, Jing Wu 0004, Kun Long, Jun Wang 0039
ICCV1
2021 VENet: Voting Enhancement Network for 3D Object Detection
abstract
Hough voting, as has been demonstrated in VoteNet, is effective for 3D object detection, where voting is a key step. In this paper, we propose a novel VoteNet-based 3D detector with vote enhancement to improve the detection accuracy in cluttered indoor scenes. It addresses the limitations of current voting schemes, i.e., votes from neighboring objects and background have significant negative impacts. Before voting, we replace the classic MLP with the proposed Attentive MLP (AMLP) in the backbone network to get better feature description of seed points. During voting, we design a new vote attraction loss (VALoss) to enforce vote centers to locate closely and compactly to the corresponding object centers. After voting, we then devise a vote weighting module to integrate the foreground/background prediction into the vote aggregation process to enhance the capability of the original VoteNet to handle noise from background voting. The three proposed strategies all contribute to more effective voting and improved performance, resulting in a novel 3D object detector, termed VENet. Experiments show that our method outperforms state-of-the-art methods on benchmark datasets. Ablation studies demonstrate the effectiveness of the proposed components.
Qian Xie 0001, Yukun Lai, Jing Wu 0004, Zhoutao Wang, Dening Lu, Mingqiang Wei, Jun Wang 0039
ICCV4
2021 Vote-Based 3D Object Detection with Context Modeling and SOB-3DNMS
Qian Xie 0001, Yukun Lai, Jing Wu 0004, Zhoutao Wang, Kai Xu 0004, Jun Wang 0039
Int. J. Comput. Vis.4
2020 MLCVNet: Multi-Level Context VoteNet for 3D Object Detection
abstract
In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects individually, without giving any consideration on contextual information between these objects. Comparatively, we propose Multi-Level Context VoteNet (MLCVNet) to recognize 3D objects correlatively, building on the state-of-the-art VoteNet. We introduce three context modules into the voting and classifying stages of VoteNet to encode contextual information at different levels. Specifically, a Patch-to-Patch Context (PPC) module is employed to capture contextual information between the point patches, before voting for their corresponding object centroid points. Subsequently, an Object-to-Object Context (OOC) module is incorporated before the proposal and classification stage, to capture the contextual information between object candidates. Finally, a Global Scene Context (GSC) module is designed to learn the global scene context. We demonstrate these by capturing contextual information at patch, object and scene levels. Our method is an effective way to promote detection accuracy, achieving new state-of-the-art detection performance on challenging 3D object detection datasets, i.e., SUN RGBD and ScanNet. We also release our code at https://github.com/NUAAXQ/MLCVNet.
Qian Xie 0001, Yukun Lai, Jing Wu 0004, Zhoutao Wang, Kai Xu 0004, Jun Wang 0039
CVPR4
2019 A novel edge-oriented framework for saliency detection enhancement
Qingzhen Xu, Fengyun Wang, Yongyi Gong, Zhoutao Wang, Qi Li 0001
Image Vis. Comput.4
2018 Thermal comfort research on human CT data modeling
Qingzhen Xu, Zhoutao Wang, Fengyun Wang
Multim. Tools Appl.2
2017 An Edge-oriented Framework for Saliency Detection
abstract
Confusing visual appearance and scattered small-scale patterns commonly exist in natural images, which forms a challenge for prior saliency detection methods. Inspired by the sensitivity to edge information of Human Visual Systems, we propose a universal edge-oriented framework to improve the performance of existing salient detection methods. Firstly, edge probability map is extracted from images and utilized to get edge-based over segmentation. Secondly, merging segments by a hierarchical model to generate edge regions. Finally, the proposed framework turns saliency detection to assign a saliency value to each edge region. Experimental results demonstrate the effectiveness of our framework.
Qingzhen Xu, Fengyun Wang, Yongyi Gong, Zhoutao Wang
BIBE4