EDBT 2026 Demo / reviewers in the wild / expert
Hanyun Wang
dblp:131/2115
· DBLP profile ↗
40ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0002-8320-4230ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | P3D: Plug-and-play prompt-driven framework for RGB-thermal semantic segmentationabstract• A plug-and-play prompt-driven framework for RGB-thermal image semantic segmentation. • LoRA-based fine-tuning strategy for SAM series model integration. • A model-agnostic encoder to generate statistical distributed prompts for training. The semantic segmentation of RGB-thermal images is critical for applications with low-light conditions. Existing works primarily focus on feature fusion strategies and model design to enhance performance. While Visual Foundation Models (VFMs) have been introduced in previous studies to improve generalization and segmentation accuracy, they suffer from poor compatibility with other models thus requiring full model retraining. Additionally, the domain gap and modality gap between VFM pre-training datasets and RGB-thermal semantic segmentation datasets pose significant challenges to VFM adaptation for downstream tasks. To address these issues, in this paper a plug-and-play prompt driven framework P 3 D is proposed. Unlike existing VFM-based methods that require complete retraining for each specific architecture, P 3 D is designed with a model-agnostic training strategy that enables one-time training and seamless integration with various existing methods without requiring retraining. First, a dual-branch LoRA (Low-Rank Adaptation) fine-tuned (DBLF) image encoder for the RGB and thermal image branches is proposed to narrow the domain gap and modality gap when incorporating SAM series models into our task. Second, a unified prompt generation and representation (UPGR) encoder is proposed. It generates diverse prompts using semantic labels during the training stage, ensuring the generated prompts are model-agnostic and compatible with existing methods. Finally, a cross-modality spatial-channel attention (CM-SCA) decoder is developed to fuse the embeddings from two-modality images and prompts for the final prediction. Extensive experiments are conducted on three popular benchmarks. Results demonstrate that P 3 D not only improves the performance of existing models but also outperforms current state-of-the-art (SOTA) methods leveraging < 1% trainable parameters. More importantly, by simply plugging P 3 D into existing methods, we consistently achieve significant performance improvements without retraining these base models, demonstrating the practical value of our plug-and-play design. Yongqi Sun, Chenguang Dai, Hanyun Wang, Longguang Wang, Wenke Li, Anzhu Yu |
Pattern Recognit. | 3 |
| 2026 | Efficient Occupancy Prediction Guided Point Cloud Geometry CompressionabstractEfficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signal-to- noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based models struggle with limited receptive fields to capture long-range dependencies, while Transformer-built architectures always neglect fine-grained details due to their reliance on global self-attention. This paper presents a Transformer-efficient occupancy prediction Network, termed TopNet, to overcome these challenges by developing several novel components designed to enhance both global context modeling and local structure preservation: Locally-enhanced Context Encoding (LeCE) for improving local structural awareness and enhancing the translation-invariance of the octree nodes, Adaptive-Length Sliding Window Attention (AL-SWA) for capturing both global and local dependencies while adaptively adjusting attention weights based on the input window length, Spatial-Gated-enhanced Channel Mixer (SG-CM) for efficient feature aggregation from ancestors and siblings, and Latent-guided Node Occupancy Predictor (LNOP) for improving prediction accuracy of spatially adjacent octree nodes in local context. Comprehensive experiments across three large-scale outdoor sparse LiDAR datasets, including SemanticKITTI, nuScenes, and LiDAR-CS, as well as two indoor dense human body datasets, including 8iVFB and MVUB, and one indoor dense scenario dataset, ScanNet, demonstrate that our TopNet achieves state-of-the-art compression performance with fewer parameters. Yifan Zhang 0030, Ting Liu 0017, Xinpu Liu, Ke Xu 0013, Jianwei Wan, Yulan Guo, Hanyun Wang |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | OctGLP-Net: Learning Octree-Structured Context Entropy Model With Global-Local Perception for Point Cloud Geometry CompressionabstractThe insufficient exploitation of spatial correlations among octree node context and feature interactions across spatial and channel dimensions limits the reconstruction performance of current point cloud geometry compression (PCGC) models. To solve these issues, this paper presents an octree-structured context entropy model OctGLP-Net with global-local perception, which mainly consists of a Local Spatial Perception (LocSP) module, a Scaled-cosine Attention based Spatial Interaction (SASI) module, and a Locally-enhanced Feed-forward Spatial and Channel Interaction (LFSCI) module. First, we propose the LocSP to extract local context information from octree nodes. Then, we introduce the SASI to fully exploit the spatial correlations among nodes. Finally, to effectively interact local features across spatial and channel dimensions, we devise a LFSCI network by employing depth-wise and point-wise convolutions to realize fine-grained local feature extraction from octree nodes. Experimental results on both sparse LiDAR and dense object benchmark datasets demonstrate that our method achieves state-of-the-art lossy/lossless compression performance. Our method obtains higher reconstruction quality (D1/D2 PSNR) and smaller chamfer distance (CD) at similar bits per point (BPP) on the SemanticKITTI, nuScenes, and LiDAR-CS datasets, and lower bitrate on the Owlii, 8iVFB and MVUB datasets. Significantly, the proposed OctGLP-Net also exhibits strong generalization abilities when applied to unseen nuScenes and LiDAR-CS datasets. In addition, downstream object detection task on the nuScenes dataset with different compression precisions further demonstrate the superiority and robustness of our method. Ke Xu 0013, Xinpu Liu, Jianwei Wan, Yulan Guo, Hanyun Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry CompressionabstractEfficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signalto-noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based models struggle with limited receptive fields to capture long-range dependencies, while Transformer-built architectures always neglect fine-grained details due to their reliance on global selfattention. In this paper, we propose a Transformer-efficient occupancy prediction Network, termed TopNet, to overcome these challenges by developing several novel components: Locally-enhanced Context Encoding (LeCE) for enhancing the translation-invariance of the octree nodes, Adaptive-Length Sliding Window Attention (ALSWA) for capturing both global and local dependencies while adaptively adjusting attention weights based on the input window length, Spatial-Gated-enhanced Channel Mixer (SG-CM) for efficient feature aggregation from ancestors and siblings, and Latent-guided Node Occupancy Predictor (LNOP) for improving prediction accuracy of spatially adjacent octree nodes. Comprehensive experiments across both indoor and outdoor point cloud datasets demonstrate that our TopNet achieves state-ofthe-art performance with fewer parameters, further advancing the reduction-efficiency boundaries of PCGC. The code is available at https://github.com/xinjiewang1995/TopNet. Yifan Zhang 0030, Ting Liu 0017, Xinpu Liu, Ke Xu 0013, Jianwei Wan, Yulan Guo, Hanyun Wang |
CVPR | 8 |
| 2025 | ASRL: Adaptive Sparse Representation Learning for LiDAR Point Cloud Geometry Compression
Ke Xu 0013, Bin Deng 0002, Yulan Guo, Hanyun Wang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Cross-Domain Incremental Feature Learning for ALS Point Cloud Semantic Segmentation With Few SamplesabstractFeature learning of airborne laser scanning (ALS) point clouds is challenged by both the limited annotated samples and imbalanced class distribution. An intuitive way involves pretraining on a well-annotated source dataset and fine-tuning on a limited target dataset. However, cross-domain challenges such as heterogeneous point cloud density, varying terrain features, and inconsistent object categories complicate transfer learning for 3-D land cover classification. In this article, we address these issues by separating the cross-domain ALS point cloud semantic segmentation into two subsequent subtasks, i.e., the cross-domain transfer learning subtask and the intradomain class-incremental learning subtask, and we use a well-annotated photogrammetric point cloud dataset as the source dataset. To mitigate domain discrepancies, the first subtask employs domain adversarial training to learn from base categories that are shared between source and target datasets. Then, the second subtask incrementally learns new categories that are specific within the target dataset using an incremental feature-semantic distillation module and a semantic adversarial learning module while retaining base category knowledge. Experimental results evaluated on three ALS point cloud datasets (ISPRS, DALES, and H3D) with different semantics show state-of-the-art cross-domain performance with few labeled samples. Compared with few-shot learning methods, our method shows promising generalization ability particularly on domain-specific categories, greatly alleviating the dependence on ALS point cloud annotations. Mofan Dai, Shuai Xing, Qing Xu 0005, Jiechen Pan, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Ms-DANet: Multiscale Difference-Aware Network for 3-D Point Cloud Change DetectionabstractWith the rapid advancements of 3-D acquisition technology, 3-D change detection has gained lots of attentions recently. Existing deep learning-based point cloud change detection methods usually adopt a common encoder-decoder structure to learn pointwise features. However, these feature learning backbones are not specifically designed for change detection task, and ignore the local structure discrepancies during feature learning. To address these issues, this article proposes a multiscale difference-aware network (Ms-DANet) for 3-D point cloud change detection. First, we propose a difference-guided multiscale feature learning (DG-MsFL) module to enhance the feature differences between bi-temporal point clouds at multiple scales during feature encoding, and use these differences to guide the network focusing more on the local structures with large discrepancies. Next, we introduce a multiscale difference feature fusion (Ms-DFF) module to fuse the multiscale feature differences to learn more discriminative features during feature decoding. Finally, we treat the point cloud change detection task as a semantic classification problem, and propose a multiscale loss (Ms-Loss) function to promote the network training. We conduct experiments on the real-world street-level point cloud change detection dataset SLPCCD and the simulated airborne urban point cloud change detection dataset URB3DCD. The experimental results show that Ms-DANet obtains a significant improvement on both the real-world and simulated point cloud change detection datasets, demonstrating its effectiveness and robustness across various sensors and data modalities. Jinhao Lu, Chenguang Dai, Zhenchao Zhang 0001, Xuanguang Liu, Ruqin Zhou, Song Ji, Haiyan Guan, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | Satellite Video Object Detection Based on Enhanced 3DTV Regularization and Gaussian PriorabstractSatellite videos have played important roles in many applications in recent years due to the advantages of continuous providing high temporal resolution remote sensing images. Although much progress has been achieved for moving object detection (MOD) in satellite videos, the low-rank characteristics of background and the intensity variations of moving objects across frames have not been fully exploited. In this article, we propose an efficient method for MOD in satellite videos, which models the background with enhanced 3-D total variation (E-3DTV) regularization and the moving objects with Gaussian prior. Specifically, considering that the gradient maps on the spatial and temporal dimensions exhibit different physical meanings, we model the background with different Laplacian sparsity priors for the gradient maps along the spatial and temporal dimensions for 3DTV regularization. Different from current methods, which model moving objects with sparsity characteristics in each frame alone, we utilize Gaussian prior to model intensity changing characteristics of moving objects across frames. After integrating background model and moving object model into low-rank sparse matrix factorization framework, the alternating direction method of multipliers (ADMM) is adopted to iteratively optimize the parameters of background and moving object models. We conduct experiments on VISO and SkySat datasets, and the results demonstrate that our method achieves superior MOD performance with high computational efficiency compared to state-of-the-art methods. Wei An 0003, Ting Liu 0017, Yang Sun 0006, Zaiping Lin, Yulan Guo, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | GCFI-Net: Global-Local Cross-Spatial-Channel Feature Interaction Network for Point Cloud Geometry CompressionabstractEfficiently compressing large-scale point cloud data under limited bandwidth and computing resource conditions has become a critical issue to be addressed in mobile computing platforms. Although the octree structure can efficiently represent large-scale and complex point clouds, existing octree-based Point Cloud Geometry Compression (PCGC) approaches typically focus on exploiting either spatial or channel features individually, neglecting the interaction across spatial-channel dimensions. In addition, current approaches are also limited to small-scale point clouds due to reliance on global Transformer or local convolutional neural network (CNN). To solve these issues, we introduce GCFI-Net, a global-local cross-spatial-channel feature interaction network for predicting the occupancy probability distribution of each octree node in this paper. In the GCFI-Net, we propose a Multiscale Convolutional Fusion-based Spatial Interaction (MCFSI) module to capture global context and model spatial interactions, and a Global-Local Cross-Channel Interaction (GLCCI) module with dual pathways to integrate global and local cross-channel information. Additionally, we propose a Multiscale-enhanced Spatial and Channel Interaction (MSCI) module to aggregate features from ancestor and sibling nodes, which further enhances the octree node representation ability. Extensive experiments on large-scale sparse LiDAR and dense human body point clouds demonstrate that the proposed GCFI-Net achieves superior compression performance with fewer parameters compared to state-of-the-art PCGC methods. Yifan Zhang 0030, Xinpu Liu, Ke Xu 0013, Jianwei Wan, Yulan Guo, Hanyun Wang |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | DuInNet: Dual-Modality Feature Interaction for Point Cloud CompletionabstractTo further promote the development of multimodal point cloud completion, we contribute a large-scale multimodal point cloud completion benchmark ModelNet-MPC with richer shape categories and more diverse test data, which contains nearly 400,000 pairs of high-quality point clouds and rendered images of 40 categories. Besides the fully supervised point cloud completion task, two additional tasks including denoising completion and zero-shot learning completion are proposed in ModelNet-MPC, to simulate real-world scenarios and verify the robustness to noise and the transfer ability across categories of current methods. Meanwhile, considering that existing multimodal completion pipelines usually adopt a unidirectional fusion mechanism and ignore the shape prior contained in the image modality, we propose a Dual-Modality Feature Interaction Network (DuInNet) in this paper. DuInNet iteratively interacts features between point clouds and images to learn both geometric and texture characteristics of shapes with the dual feature interactor. To adapt to specific tasks such as fully supervised, denoising, and zero-shot learning point cloud completions, an adaptive point generator is proposed to generate complete point clouds in blocks with different weights for these two modalities. Extensive experiments on the ShapeNet-ViPC and ModelNet-MPC benchmarks demonstrate that DuInNet exhibits superiority, robustness and transfer ability in all completion tasks over state-of-the-art methods. The code and dataset will be available athttps://github.com/xinpuliu/DuInNet. Xinpu Liu, Baolin Hou, Hanyun Wang, Ke Xu 0013, Jianwei Wan, Yulan Guo |
IEEE Trans. Multim. | 3 |
| 2025 | Supervised Contrastive Learning for Indoor Point Cloud OversegmentationabstractPoint cloud oversegmentation method can obtain a series of superpoints by grouping points that are semantically and geometrically consistent. The generated superpoints can be treated as the basic processing units in various downstream tasks to improve task performance and processing efficiency. However, due to the high semantic and geometric complexity of point cloud scenes, obtaining high-quality superpoints is still challenging. Aiming to generate high-quality indoor superpoints, we propose an end-to-end supervised contrastive learning framework SCL-OverSeg for indoor point cloud oversegmentation. Firstly, to solve the challenge of balancing the importance of geometric similarity and spatial proximity constraint between points and superpoints in indoor scenes, we integrate the geometric similarity and spatial proximity constraint into the supervision signal by generating the superpoint ground truth. To solve the challenge of superpoints crossing objects, we propose to utilize instance labels rather than semantic labels to generate the ideal superpoint ground truth as the object-level supervision signal. Secondly, to construct the distinguishable embedding space facilitating to the assignments of points to superpoints, we propose point-superpoint contrastive learning to compel the network to project each point to be closer to the reasonable superpoint in embedding space. Besides, with the instance labels, to improve the superpoint performance on object boundaries, we propose the object boundary contrastive learning to enhance the feature distinguishability between tough points across the object boundaries. Extensive experiments demonstrate that SCL-OverSeg can effectively improve indoor oversegmentation performance, especially on object boundaries. The relevant codes will be available onhttps://github.com/sssssyf/SCL-OverSeg. Yifan Sun 0008, Chenguang Dai, Wenke Li, Song Ji, Anzhu Yu, Yiping Chen 0002, Hanyun Wang |
IEEE Trans. Multim. | 8 |
| 2024 | ICPR 2024 Competition on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Han Wang 0049, Furui Chen, Silei Liu, Xiaomin Huang, Shining Wang, Ying Li 0017, Peng Wang 0015, Shiyong Peng, Xiaokai Bi, Renbin Zou, Wenjing Deng, Zhen Cui 0001 |
ICPR (34) | 8 |
| 2024 | GRLoR: A Unified Global Retrieval and Local Reranking Framework for 3-D Place RecognitionabstractThree-dimensional place recognition aims to search point cloud in a large database that matches the query. It is an essential task in remote sensing applications, such as smart city management and disaster monitoring. The existing methods commonly leverage global descriptors to perform point cloud retrieval for place recognition. However, these methods rely on spatial aggregation to obtain global descriptors, which are neither discriminative nor general. In this letter, we propose a unified global retrieval and local reranking (namely, GRLoR) framework for 3-D place recognition. Specifically, we first utilize a self-attention mechanism to capture the channel dependencies of local features and design a spatial-fusion pooling (SFP) approach to obtain a discriminative global descriptor for retrieval. We then construct a feature correlation module for local reranking, which uses a cross-attention mechanism to determine whether the point cloud pair matches correctly by predicting the similarity of local regions. Experiments conducted on several public benchmarks validate the superiority performance of our method. For instance, it outperforms the strongest model by an average of about 1% on the public datasets in terms of AR@1. Wenshuo Liu, Sheng Ao, Ye Zhang 0037, Hanyun Wang, Yulan Guo |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | TBSCD-Net: A Siamese Multitask Network Integrating Transformers and Boundary Regularization for Semantic Change Detection From VHR Satellite ImagesabstractSemantic change detection (SCD) from very high-resolution images involves two key challenges: (1) the global features of bitemporal images tend to be extracted insufficiently, leading to imprecise land cover semantic classification results, and (2) the detected changed objects exhibit ambiguous boundaries, resulting in low geometric accuracy. To address these two issues, we propose an SCD method called TBSCD-Net based on a multi-task learning framework to simultaneously identify different types of semantic changes and regularize change boundaries. Firstly, we construct a hybrid encoder combining transformer and convolutional neural network (TCEncoder) to enhance the extraction of global context information. A bitemporal semantic linkage module (Bi-SLM) is embedded into the TCEncoder to enhance the semantic correlations between bitemporal images. Secondly, we introduce a boundary-region joint extractor based on Laplacian operators (LOBRE) to regularize the changed objects. We evaluated the effectiveness of the proposed method using the SECOND dataset and a Fuzhou GF-2 SCD dataset (FZ-SCD) and compared it with seven existing methods. The proposed method performed better than the other evaluated methods as it achieved 24.42% Sek and 20.18% GTC on the SECOND dataset and 23.10% Sek and 23.15% GTC on the FZ-SCD dataset. The results of ablation studies on the FZ-SCD dataset also verified the effectiveness of the developed modules for SCD. Xuanguang Liu, Chenguang Dai, Zhenchao Zhang 0001, Mengmeng Li 0002, Hanyun Wang, Hongliang Ji, Yujie Li 0010 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Multiprototype Relational Network for Few-Shot ALS Point Cloud Semantic Segmentation by Transferring Knowledge From Photogrammetric Point CloudsabstractExisting airborne laser scanning (ALS) point cloud semantic segmentation approaches are limited by their overreliances on sufficient point-wise annotations that further confine their generalization ability to new scenes. To overcome these problems, a novel three-stage multi-prototype relational network (Thr-MPRNet) is proposed for few-shot ALS point cloud semantic segmentation by transferring knowledge from well-annotated photogrammetric point clouds. In MPRNet, a 3D few-shot learning structure containing a feature learner and a relation learner is built to learn meta-knowledge from multiple point-wise tasks, and a multi-prototype generator is designed to represent the semantic distribution of point clouds that can dynamically adapt to large-scale scenarios. Then, to transfer knowledge across different domains, MPRNet is trained in a unified framework with three task-based learning stages. Prior knowledge is first meta-learned from the source photogrammetric point clouds and then transferred to novel target datasets with a few labeled ALS point clouds. Finally, the MPRNet can be flexibly generalized to the unlabeled target ALS point clouds without further retraining from scratch. In the experiments, the SensatUrban dataset is used as the source photogrammetric point clouds, and two ALS point cloud datasets (ISPRS and DALES) are used to evaluate the few-shot semantic segmentation ability of the proposed method. The experiments demonstrate that Thr-MPRNet obtains promising generalization performance on different target datasets. More importantly, it outperforms supervised networks with 10% labeled samples. In summary, the proposed method achieves state-of-the-art cross-domain semantic segmentation performance and greatly alleviates the dependence on ALS point cloud annotations. Mofan Dai, Shuai Xing, Qing Xu 0005, Jiechen Pan, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Deep Semantic Graph Matching for Large-Scale Outdoor Point Cloud RegistrationabstractCurrent point cloud registration methods are mainly based on local geometric information and usually ignore the semantic information contained in the scenes. In this paper, we treat the point cloud registration problem as a semantic instance matching and registration task, and propose a deep semantic graph matching method (DeepSGM) for large-scale outdoor point cloud registration. Firstly, the semantic categorical labels of 3D points are obtained using a semantic segmentation network. The adjacent points with the same category labels are then clustered together using the Euclidean clustering algorithm to obtain the semantic instances, which are represented by three kinds of attributes including spatial location information, semantic categorical information, and global geometric shape information. Secondly, the semantic adjacency graph is constructed based on the spatial adjacency relations of semantic instances. To fully explore the topological structures between semantic instances in the same scene and across different scenes, the spatial distribution features and the semantic categorical features are learned with graph convolutional networks, and the global geometric shape features are learned with a PointNet-like network. These three kinds of features are further enhanced with the self-attention and cross-attention mechanisms. Thirdly, the semantic instance matching is formulated as an optimal transport problem, and solved through an optimal matching layer. Finally, the geometric transformation matrix between two point clouds is first estimated by the SVD algorithm and then refined by the ICP algorithm. Experimental results conducted on the KITTI Odometry dataset demonstrate that the proposed method improves the registration performance and outperforms various state-of-the-art methods. Shaocong Liu, Tao Wang 0075, Yan Zhang 0159, Ruqin Zhou, Li Li 0100, Chenguang Dai, Longguang Wang, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2023 | BUFFER: Balancing Accuracy, Efficiency, and Generalizability in Point Cloud RegistrationabstractAn ideal point cloud registration framework should have superior accuracy, acceptable efficiency, and strong generalizability: However, this is highly challenging since existing registration techniques are either not accurate enough, far from efficient, or generalized poorly. It remains an open question that how to achieve a satisfying balance between this three key elements. In this paper, we propose BUFFER, a point cloud registration method for balancing accuracy, efficiency, and generalizability. The key to our approach is to take advantage of both point-wise and patch-wise techniques, while overcoming the inherent drawbacks simultaneously. Different from a simple combination of existing methods, each component of our network has been carefully crafted to tackle specific issues. Specifically, a Point-wise Learner is first introduced to enhance computational efficiency by predicting keypoints and improving the representation capacity of features by estimating point orientations, a Patch-wise Embedder which leverages a lightweight local feature learner is then deployed to extract efficient and general patch features. Additionally, an Inliers Generator which combines simple neural layers and general features is presented to search inlier correspondences. Extensive experiments on real-world scenarios demonstrate that our method achieves the best of both worlds in accuracy, efficiency, and generalization. In particular, our method not only reaches the highest success rate on unseen domains, but also is almost 30 times faster than the strong baselines specializing in generalization. Code is available at https://github.com/aosheng1996/BUFFER. Sheng Ao, Qingyong Hu, Hanyun Wang, Kai Xu 0004, Yulan Guo |
CVPR | 3 |
| 2023 | Sparse Representation based Deep Residual Geometry Compression Network for Large-scale Point CloudsabstractThe increasing applications of 3D point clouds require efficient compression techniques to achieve high-quality and low-delay services. However, the computational efficiency and rate-distortion performance for large-scale dense point clouds are still challenging, and the phenomenon of reconstruction ability degradation also exists when the network is deep. To solve these challenges, we propose a novel fully end-to-end point cloud compression model based on sparse convolution. Specifically, we adopt a long-range-residual aided architecture to avoid the reconstruction degradation and high computational complexity of deep networks. Further, we propose a multi-scale geometry compression module to construct an end-to-end network that avoids the accumulation of reconstruction distortion during decoding. Experiments on the large-scale Moving Picture Experts Group (MPEG) PCC benchmarks show that our model outperforms the latest Video-based Point Cloud Compression (V-PCC) scheme in terms of lossy geometry compression by 50.4% in D1 BD-rate and 50.8% in D2 BD-rate, while maintaining affordable processing speed and memory consumption. Pengpeng Yu, Dian Zuo, Yueer Huang, Ruishan Huang, Hanyun Wang, Yulan Guo, Fan Liang 0001 |
ICME | 5 |
| 2023 | OctPCGC-Net: Learning Octree-Structured Context Entropy Model for Point Cloud Geometry Compression
Hanyun Wang, Ke Xu 0013, Jianwei Wan, Yulan Guo |
PRCV (2) | 2 |
| 2023 | PMNet: A Point-to-Mesh Network for 3-D Semantic Instance ReconstructionabstractSemantic instance reconstruction attracts increasing attention in several areas such as mobile mapping, scene reconstruction, and robot navigation. Although much progresses have been made in recent years, the reconstruction performance is highly sensitive to occlusions and noises. To address these issues, we incorporate point cloud completion into a novel semantic instance reconstruction network PMNet, which consists of a 3-D object detection module, a point cloud completion module, and a mesh generation module. Based on the candidate instance proposals and their proposal features obtained in the object detection module, a point encoder layer is proposed to learn the local geometric features from the point cloud belonging to the detected instances, and a feature transformation layer is utilized to align the proposal features with the local geometric features. These two types of features are then fused and fed into the point cloud decoder to predict the complete point cloud of each instance. The mesh is finally reconstructed for each instance by the mesh generation module. Quantitative and qualitative experiments conducted on the ScanNetv2 dataset demonstrate that the proposed PMNet achieves the best reconstruction performance on real-world point clouds. Junhui Wan, Zhiheng Fu, Minglin Chen, Peng Zhang 0079, Hanyun Wang, Yulan Guo |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Scene Overlap Prediction for LiDAR-Based Place RecognitionabstractRecently, LiDAR (Light Detection and Ranging)-based place recognition, has been widely concerned because of its robustness to light conditions, seasonal changes, and viewpoint variations. Unlike most of existing methods which represent the whole point cloud scenes with global descriptors, we treat the LiDAR-based place recognition problem as a scene overlap prediction task and propose an end-to-end overlap prediction network, which consists of a feature learning backbone, a feature enhancement module, and an overlap prediction module. Based on the prediction result for each point, the overlapping ratios between two point clouds are computed and used to predict whether these two point clouds are at the same place. In addition, to promote the computational efficiency and reduce the model complexity, a lightweight feature learning backbone is also adopted. The experiments conducted on the KITTI Odometry dataset demonstrate that the proposed method achieves superior performance compared with state-of-the-art methods. The lightweight method also obtains 2x inference speed with little performance degradation compared with the vanilla method. Yingjian Zhang, Chenguang Dai, Ruqin Zhou, Zhenchao Zhang 0001, Hongliang Ji, Huixin Fan, Hanyun Wang |
IEEE Geosci. Remote. Sens. Lett. | 8 |
| 2023 | AnchorPoint: Query Design for Transformer-Based 3D Object Detection and TrackingabstractWith the success of Transformers in natural language processing, object detection with Transformers (DETR) has attracted widespread attentions. In previous Transformer-based 2D detectors, the object queries are a set of learning embeddings. However, it is very hard to apply these detectors to the 3D domain due to the lack of explicit physical meanings and position priors of learned object queries. In this paper, we introduce the concept of anchors and propose a novel query design based on anchor points. In our query design, we use the foreground points as the anchor points and encode these anchor points as the object queries. Consequently, each object query has an explicit physical meaning and only focus on its nearby object. Additionally, we also propose an instance-aware sampling strategy to select a small set of representation foreground points from the scene point cloud. Extensive experiments on several large-scale 3D object detection datasets demonstrate that the proposed AnchorPoint detector achieves promising accuracy and efficiency. In particularly, AnchorPoint achieves an average precision (AP) of 83.21 at 61 frame-per-second (FPS) on the moderate level of the KITTI-DET Car subset. Moreover, we model each object as its corresponding anchor point, and extend the AnchorPoint model to 3D multi-object tracking by adding an extra tracking head. We show that our method achieves comparable performance to existing state-of-the-art methods on the KITTI-MOT dataset. Hao Liu 0061, Yanni Ma, Hanyun Wang, Yulan Guo |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | 3DAC: Learning Attribute Compression for Point CloudsabstractWe study the problem of attribute compression for large-scale unstructured 3D point clouds. Through an in-depth exploration of the relationships between different encoding steps and different attribute channels, we introduce a deep compression network, termed 3DAC, to explicitly compress the attributes of 3D point clouds and reduce storage usage in this paper. Specifically, the point cloud attributes such as color and reflectance are firstly converted to transform coefficients. We then propose a deep entropy model to model the probabilities of these coefficients by considering information hidden in attribute transforms and previous encoded attributes. Finally, the estimated probabilities are used to further compress these transform coefficients to a final attributes bitstream. Extensive experiments conducted on both indoor and outdoor large-scale open point cloud datasets, including ScanNet and SemanticKITTI, demonstrated the superior compression rates and reconstruction quality of the proposed method. Guangchi Fang, Qingyong Hu, Hanyun Wang, Yiling Xu, Yulan Guo |
CVPR | 3 |
| 2022 | The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and ResultsabstractIn this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis. Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003 |
ICPR | 7 |
| 2022 | MaskNet++: Inlier/outlier identification for two point clouds
Ruqin Zhou, Hanyun Wang, Xixing Li, Yulan Guo, Chenguang Dai, Wanshou Jiang |
Comput. Graph. | 2 |
| 2022 | Deep learning for 3D visionabstractWith the rapid development of 3D imaging sensors, such as depth cameras and laser scanning systems, 3D data has become increasingly accessible. Meanwhile, the boost of various deep learning algorithms, such as convolutional neural networks and transformers, further increases the usability of 3D vision systems. Driven by these factors, 3D vision has become an emerging and core component for numerous applications, such as autonomous driving, augmented reality, virtual reality and robotics. Although remarkable progress has been achieved in this area during the last few years, there are still several challenges that need to be addressed, such as the noisy, sparse, and irregular nature of point clouds, the high cost to label 3D data and the necessity to integrate geometry-based and learning-based techniques. Besides, 3D data produced by different 3D imaging sensors (e.g. structured light, stereo, LiDAR and time-of-flight) can be highly different. It is, therefore, necessary to investigate general algorithms that can mitigate the domain gap between different types of 3D data. This special issue aims to collect and present the latest research development in learning-based 3D vision theories and their applications and to inspire future research in this area. In total, there are eight papers accepted for publication in this special issue through careful peer reviews and revisions. These accepted papers are broadly categorised into three topics, and the summary of each topic is given below. TOPIC A—OPTICAL FLOW AND DEPTH ESTIMATION Han et al., in their paper ‘DEMVSNet: Denoising and Depth Inference for Unstructured Multi-View Stereo on Noised Images’, proposed a DEMVSNet to simultaneously address the depth estimation and image denoising problems for unstructured multi-view stereo. The multi-scales feature maps for each image are wrapped to construct cost volumes containing both the depth and RGB information through differentiable homography and Gaussian probability mapping. The cost volume regularisation module is then adopted to predict the probability of depth and RGB. To avoid overfitting in multi-task learning, the gradient normalisation algorithm is utilised to dynamically fine-tune the weights between the depth prediction task and the denoising task. To evaluate the performance of proposed DEMVSNet, a noisy Technical University of Denmark dataset is generated by adding Gaussian-Poisson noise to each image, and the experimental results demonstrate the superiority of DEMVSNet on both the denoising and multi-view stereo reconstruction tasks. Lin et al., in their paper ‘EAGAN: Event-Based Attention Generative Adversarial Networks for Optical Flow and Depth Estimation’, proposed an event-based attention generative adversarial network named EAGAN to simultaneously deal with optical flow and depth estimation based on monocular event camera. The generator of EAGAN is similar to U-net except that a transformer structure is introduced between the encoder and decoder. The position-coding features learnt from the transformer is added to features learnt from the encoding layer, which helps to capture the correlation between sequence information. The discriminator of EAGAN is based on a fully convolutional network and aims to distinguish whether the depth image or the optical flow image is generated by the generator. Experimental results conducted on the multi-vehicle stereo event camera dataset demonstrate the effectiveness of EAGAN on both the depth and optical flow estimation tasks. TOPIC B—POSE ESTIMATION Gao et al., in their paper ‘Efficient 6D Object Pose Estimation based on Attentive Multi-Scale Contextual Information’, proposed an end-to-end 6D pose estimation network to utilise multi-scale contextual features learnt from two heterogeneous data. First, interesting objects are detected from an RGB-D image using an existing semantic segmentation method. Then, pixel-wise geometric and colour features are learnt from 3D point clouds and 2D images respectively. Next, three pixelwise feature attention mechanism modules are utilised to exploit the inter-channel relationship of multimodal features. Finally, multi-scale features are extracted at three different scales and 6D pose is estimated through a dense regression module. Experimental results conducted on the LineMOD and YCB-Video datasets demonstrate that the proposed method achieves state-of-the-art performances in terms of average point distance and average closest point distance. Liu et al., in their paper ‘Auto Calibration of Multi-Camera System for Human Pose Estimation’, proposed an iterative joint estimation of intrinsic and extrinsic parameters for a multi-camera system. Specifically, keypoints are detected with high confidence to estimate the essential matrix between two cameras, and the valid extrinsic parameters are estimated by assuming that the intrinsic parameters are known a priori. Then, the reconstructed 3D human body coordinates are projected into the pixel coordinate system, and the intrinsic parameters are estimated by minimising the projection errors. The experimental results show that the proposed method achieves better performance than commonly used calibration tools. TOPIC C—POINT CLOUD PROCESSING AND UNDERSTANDING Liu et al., in their paper ‘Point Cloud Completion by Dynamic Transformer with Adaptive Neighbourhood Feature Fusion’, utilised the adaptive neighbourhood feature extraction (ANE) module and genetic hierarchical point generation (GHG) module to accomplish the point cloud completion task. The ANE module selects k nearest points both in the spatial and feature spaces adaptively according to different target shapes. The GHG module generates finer point clouds hierarchically according to the local shape characteristics, and the shape information of current points is transferred to the next stage through a dynamic transformer structure. The experimental results conducted on the Point Completion Network and Completion3D datasets demonstrate the superiority of the proposed method. Wang et al., in their paper ‘PCCN-RE: Point Cloud Colourisation Network Based on Relevance Embedding’, proposed a highly authentic point cloud colorisation network based on conditional generative adversarial (cGANs) networks. The generator network predicts the colours from the coordinates of each point, while the discriminator utilises the coordinates and the generated colours to determine the reality of input colourised point clouds. Three key components are contained in the generator. Specifically, the relevance embedding structure captures the most related local information, the weighted pooling structure aggregates the local features based on the correlation values of the covariance matrix, and the enhanced spatial transform network keeps the point clouds invariant to the geometric transformations based on weighted pooling and maximal pooling. The experimental results show that the proposed method achieves the highest Peak Signal to Noise Ratio and Structural Similarity Index on the ShapeNetCore dataset. Fang et al., in their paper ‘Sparse Point-Voxel Aggregation Network for Efficient Point Cloud Semantic Segmentation’, proposed a sparse point-voxel aggregation network to overcome high computational costs in the point cloud semantic segmentation task. In the encoding layer, the local context features are learnt through a sparse convolutional network performed on the voxelised point cloud, and the individual point features are learnt through multi-layer perceptron (MLP)-based network performed on the original point cloud. In the decoding layer, these two kinds of features are aggregated at different encoding layers through simple MLP layers. The experimental results show that the proposed method achieves state-of-the-art performance on the SemanticKITTI and S3DIS datasets. Wang et al., in their paper ‘Scale Robust Point Matching-Net: End-to-End Scale Point Matching Using Lie Group’, proposed an end-to-end scale point cloud matching network named SRPM-Net based on Lie Group. The extracted pointwise features are composed of point absolute coordinates, relative coordinates and point pair features of neighbouring points, and the local context features are aggregated through an attentive pooling layer. The matching matrix is computed via the exponential map of Lie group, which represents the feature similarity of points in two point clouds. The final transformation estimation problem is transferred as estimating the coefficients of the Lie algebra optimisation problem and is optimised through an iterative linear optimisation approach. The experimental results show that SRPM-Net achieves the best performance on the ModelNet40 and Stanford 3D scanning datasets. SUMMARY/CONCLUSION The papers published in this Special Issue show that traditional topics, such as optical flow and depth estimation, pose estimation, and point cloud processing have developed very fast in recent years. In addition, many topics have emerged in deep learning-based 3D vision, such as multi-task joint learning and multimodality intelligence. Future research in this field is expected to boost the theoretical development and potential applications of 3D vision. Yulan Guo, Hanyun Wang, Ronald Clark, Stefano Berretti, Mohammed Bennamoun |
IET Comput. Vis. | 2 |
| 2022 | Domain Adaptation for Object Classification in Point Clouds via Asymmetrical Siamese and Conditional Adversarial NetworkabstractNowadays, researchers have developed various deep neural networks for processing point clouds effectively. Due to the enormous parameters in deep learning-based models, a lot of manual efforts have to be invested into annotating sufficient training samples. To mitigate such manual efforts of annotating samples for a new scanning device, this letter focuses on proposing a new neural network to achieve domain adaptation in 3D object classification. Specifically, to minimize the data discrepancy of intra-class objects in different domains, an Asymmetrical Siamese module is designed to align the intra-class features. To preserve the discriminative information for distinguishing inter-class objects in different domains, a Conditional Adversarial module is leveraged to consider the classification information conveyed from the classifier. To verify the effectiveness of the proposed method on object classification in heterogeneous point clouds, evaluations are conducted on three point cloud datasets, which are collected in different scenarios by different laser scanning devices. Furthermore, the comparative experiments also demonstrate the superior performance of the proposed method on the classification accuracy. Huan Luo 0001, Lingkai Li, Lina Fang, Hanyun Wang, Cheng Wang 0003, Wenzhong Guo, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Continuous Mapping Convolution for Large-Scale Point Clouds Semantic SegmentationabstractIn this letter, we introduce MappingConvSeg, a continuous convolution network for semantic segmentation of large-scale point clouds. In particular, a conceptually simple, end-to-end learnable, and continuous convolution operator is proposed for learning spatial correlation of unstructured 3-D point clouds. For each local point set, the unstructured point features are first mapped onto a series of learned kernel points based on the spatial relationship, and the continuous convolution is then applied to capture specific local geometrical patterns. Taking the proposed mapping convolution operation as the building block, a hierarchical network is then built for large-scale point cloud semantic segmentation. Experimental results conducted on two public benchmarks, including Toronto-3D and Stanford large-scale 3-D Indoor Spaces (S3DIS) dataset, demonstrate the superiority of the proposed method. Kunping Yan, Qingyong Hu, Hanyun Wang, Li Li 0100, Song Ji |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Contour deformation network for instance segmentation
Kefeng Lv, Hanyun Wang, Huaigang Jiang, Chenguang Dai |
Pattern Recognit. Lett. | 4 |
| 2022 | RoadCapsFPN: Capsule Feature Pyramid Network for Road Extraction From VHR Optical Remote Sensing ImageryabstractRoad detection plays an important role in a wide range of applications. However, due to size variations, spectral diversities, occlusions, and complex scenarios, it is still challenging to accurately extract roads from very-high resolution (VHR) optical remote sensing images. This paper proposes a capsule feature pyramid network for extracting road networks from VHR optical images, termed as RoadCapsFPN. By designing a capsule feature pyramid network, the RoadCapsFPN extracts and integrates multiscale capsule features to recover a high-resolution and semantically strong road feature representation. Next, we also design a contextual feature module, including dense atrous convolution (DAC) and residual multi-kernel pooling (RMP) units, to further exploit rich contextual properties of the roads at a high-resolution perspective. Benefitting from the multiscale feature abstraction and context augmentation, our RoadCapsFPN shows impressing results in processing variedly-sized and diversely-spectral roads in complex environments. Two testing datasets, Google and Massichusate Roads Datasets, are used for evaluating the proposed RoadCapsFPN via four testing indicators -precision,recall, intersection-over-union (IoU), and$F_{1}$-score. Comparative studies also confirm the superior performance of the RoadCapsFPN in accurately extracting road networks. Haiyan Guan, Yongtao Yu, Dilong Li, Hanyun Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Deep Learning for 3D Point Clouds: A SurveyabstractPoint cloud learning has lately attracted increasing attention due to its wide applications in many areas, such as computer vision, autonomous driving, and robotics. As a dominating technique in AI, deep learning has been successfully used to solve various 2D vision problems. However, deep learning on point clouds is still in its infancy due to the unique challenges faced by the processing of point clouds with deep neural networks. Recently, deep learning on point clouds has become even thriving, with numerous methods being proposed to address different problems in this area. To stimulate future research, this paper presents a comprehensive review of recent progress in deep learning methods for point clouds. It covers three major tasks, including 3D shape classification, 3D object detection and tracking, and 3D point cloud segmentation. It also presents comparative results on several publicly available datasets, together with insightful observations and inspiring future research directions. Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu 0061, Li Liu 0002, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Adv-Depth: Self-Supervised Monocular Depth Estimation With an Adversarial LossabstractLoss function plays a key role in self-supervised monocular depth estimation methods. Current reprojection loss functions are hand-designed and mainly focus on local patch similarity but overlook the global distribution differences between a synthetic image and a target image. In this paper, we leverage global distribution differences by introducing an adversarial loss into the training stage of self-supervised depth estimation. Specifically, we formulate this task as a novel view synthesis problem. We use a depth estimation module and a pose estimation module to form a generator, and then design a discriminator to learn the global distribution differences between real and synthetic images. With the learned global distribution differences, the adversarial loss can be back-propagated to the depth estimation module to improve its performance. Experiments on the KITTI dataset have demonstrated the effectiveness of the adversarial loss. The adversarial loss is further combined with the reprojection loss to achieve the state-of-the-art performance on the KITTI dataset. Kunhong Li 0001, Zhiheng Fu, Hanyun Wang, Zonghao Chen, Yulan Guo |
IEEE Signal Process. Lett. | 3 |
| 2016 | Exploiting location information to detect light pole in mobile LiDAR point cloudsabstractWith rapid development of light detection and ranging (LiDAR) technologies, three dimensional point clouds increasingly become a new approach to sense the world. In our previous work, light poles were detected from mobile LiDAR point clouds without using their locations. In this paper, we improve our previous work by considering location information between two neighboring light poles to reduce false alarm. In the proposed method, the potential light poles are first detected by the extended Hough Forest Framework. Then, a gaussian distribution is exploited to model the distance between two light poles by using locations of those detected light poles. Finally, inaccurately detected light poles are removed by considering the distance between two adjacent objects. We evaluate our proposed method on mobile LiDAR point clouds acquired by RIEGL VMX-450 system. On the basis of the experimental test instances, we demonstrate improved accuracy on light pole detection. Huan Luo 0001, Cheng Wang 0003, Hanyun Wang, Ziyi Chen 0001, Dawei Zai, Shanxin Zhang, Jonathan Li 0001 |
IGARSS | 3 |
| 2016 | Vehicle Detection in High-Resolution Aerial Images Based on Fast Sparse Representation Classification and Multiorder FeatureabstractThis paper presents an algorithm for vehicle detection in high-resolution aerial images through a fast sparse representation classification method and a multiorder feature descriptor that contains information of texture, color, and high-order context. To speed up computation of sparse representation, a set of small dictionaries, instead of a large dictionary containing all training items, is used for classification. To extract the context information of a patch, we proposed a high-order context information extraction method based on the proposed fast sparse representation classification method. To effectively extract the color information, the RGB color space is transformed into color name space. Then, the color name information is embedded into the grids of histogram of oriented gradient feature to represent the low-order feature of vehicles. By combining low- and high-order features together, a multiorder feature is used to describe vehicles. We also proposed a sample selection strategy based on our fast sparse representation classification method to construct a complete training subset. Finally, a set of dictionaries, which are trained by the multiorder features of the selected training subset, is used to detect vehicles based on superpixel segmentation results of aerial images. Experimental results illustrate the satisfactory performance of our algorithm. Ziyi Chen 0001, Cheng Wang 0003, Huan Luo 0001, Hanyun Wang, Yiping Chen 0002, Chenglu Wen, Yongtao Yu, Liujuan Cao, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2016 | Patch-Based Semantic Labeling of Road Scene Using Colorized Mobile LiDAR Point CloudsabstractSemantic labeling of road scenes using colorized mobile LiDAR point clouds is of great significance in a variety of applications, particularly intelligent transportation systems. However, many challenges, such as incompleteness of objects caused by occlusion, overlapping between neighboring objects, interclass local similarities, and computational burden brought by a huge number of points, make it an ongoing open research area. In this paper, we propose a novel patch-based framework for labeling road scenes of colorized mobile LiDAR point clouds. In the proposed framework, first, three-dimensional (3-D) patches extracted from point clouds are used to construct a 3-D patch-based match graph structure (3D-PMG), which transfers category labels from labeled to unlabeled point cloud road scenes efficiently. Then, to rectify the transferring errors caused by local patch similarities in different categories, contextual information among 3-D patches is exploited by combining 3D-PMG with Markov random fields. In the experiments, the proposed framework is validated on colorized mobile LiDAR point clouds acquired by the RIEGL VMX-450 mobile LiDAR system. Comparative experiments show the superior performance of the proposed framework for accurate semantic labeling of road scenes. Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Zhipeng Cai 0003, Ziyi Chen 0001, Hanyun Wang, Yongtao Yu, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2016 | Spatial-Related Traffic Sign Inspection for Inventory Purposes Using Mobile Laser Scanning DataabstractThis paper presents a spatial-related traffic sign inspection process for sign type, position, and placement using mobile laser scanning (MLS) data acquired by a RIEGL VMX-450 system and presents its potential for traffic sign inventory applications. First, the paper describes an algorithm for traffic sign detection in complicated road scenes based on the retroreflectivity properties of traffic signs in MLS point clouds. Then, a point cloud-to-image registration process is proposed to project the traffic sign point clouds onto a 2-D image plane. Third, based on the extracted traffic sign points, we propose a traffic sign position and placement inspection process by creating geospatial relations between the traffic signs and road environment. For further inventory applications, we acquire several spatial-related inventory measurements. Finally, a traffic sign recognition process is conducted to assign sign type. With the acquired sign type, position, and placement data, a spatial-associated sign network is built. Experimental results indicate satisfactory performance of the proposed detection, recognition, position, and placement inspection algorithms. The experimental results also prove the potential of MLS data for automatic traffic sign inventory applications. Chenglu Wen, Jonathan Li 0001, Huan Luo 0001, Yongtao Yu, Zhipeng Cai 0003, Hanyun Wang, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2015 | Road Boundaries Detection Based on Local Normal Saliency From Mobile Laser Scanning DataabstractThe accurate extraction of roads is a prerequisite for the automatic extraction of other road features. This letter describes a method for detecting road boundaries from mobile laser scanning (MLS) point clouds in an urban environment. The key idea of our method is directly constructing a saliency map on 3-D unorganized point clouds to extract road boundaries. The method consists of four major steps, i.e., road partition with the assistance of the vehicle trajectory, salient map construction and salient points extraction, curb detection and curb lowest points extraction, and road boundaries fitting. The performance of the proposed method is evaluated on the point clouds of an urban scene collected by a RIEGL VMX-450 MLS system. The completeness, correctness, and quality of the extracted road boundaries are 95.41%, 99.35%, and 94.81%, respectively. Experimental results demonstrate that our method is feasible for detecting road boundaries in MLS point clouds. Hanyun Wang, Huan Luo 0001, Chenglu Wen, Jun Cheng 0002, Peng Li 0064, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Iterative Tensor Voting for Pavement Crack Extraction Using Mobile Laser Scanning DataabstractThe assessment of pavement cracks is one of the essential tasks for road maintenance. This paper presents a novel framework, called ITVCrack, for automated crack extraction based on iterative tensor voting (ITV), from high-density point clouds collected by a mobile laser scanning system. The proposed ITVCrack comprises the following: 1) the preprocessing involving the separation of road points from nonroad points using vehicle trajectory data; 2) the generation of the georeferenced feature (GRF) image from the road points; and 3) the ITV-based crack extraction from the noisy GRF image, followed by an accurate delineation of the curvilinear cracks. Qualitatively, the method is applicable for pavement cracks with low contrast, low signal-to-noise ratio, and bad continuity. Besides the application to GRF images, the proposed framework demonstrates much better crack extraction performance when quantitatively compared to existing methods on synthetic data and pavement images. Haiyan Guan, Jonathan Li 0001, Yongtao Yu, Michael A. Chapman, Hanyun Wang, Cheng Wang 0003, Ruifang Zhai |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2014 | Object Detection in Terrestrial Laser Scanning Point Clouds Based on Hough ForestabstractThis letter presents a novel rotation-invariant method for object detection from terrestrial 3-D laser scanning point clouds acquired in complex urban environments. We utilize the Implicit Shape Model to describe object categories, and extend the Hough Forest framework for object detection in 3-D point clouds. A 3-D local patch is described by structure and reflectance features and then mapped to the probabilistic vote about the possible location of the object center. Objects are detected at the peak points in the 3-D Hough voting space. To deal with the arbitrary azimuths of objects in real world, circular voting strategy is introduced by rotating the offset vector. To deal with the interference of adjacent objects, distance weighted voting is proposed. Large-scale real-world point cloud data collected by terrestrial mobile laser scanning systems are used to evaluate the performance. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 3-D object detection methods. Hanyun Wang, Cheng Wang 0003, Huan Luo 0001, Peng Li 0064, Ming Cheng 0002, Chenglu Wen, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | Learn Multiple-Kernel SVMs for Domain Adaptation in Hyperspectral DataabstractThis letter presents a novel semisupervised method for addressing a domain adaptation problem in the classification of hyperspectral data. To overcome the influence of distribution bias between the source and target domains, we introduce the domain transfer multiple-kernel learning to simultaneously minimize the maximum mean discrepancy criterion and the structural risk functional of support vector machines. Then, the pairwise binary classifiers are merged as the multiclass classifier for solving the classification problem in hyperspectral data. Both bias and nonbias sampling strategies are introduced to evaluate the robustness of the proposed method against the spectral distribution bias. The results obtained from real data sets show that the proposed method can achieve higher classification accuracy even with cross-domain distribution bias and provide robust solutions with different labeled and unlabeled data sizes. Cheng Wang 0003, Hanyun Wang, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |