Ming Cheng 0002

dblp:82/104-2 · DBLP profile ↗
← Back
56ranked-venue papers
3as first author
38since 2021 · last 2026
0000-0001-6480-6482ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 21 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 11 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range Conditions
abstract
Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and cross-distance scenarios under real-world conditions. To address this gap, we present LRGait, the first LiDAR-Camera multimodal benchmark designed for robust long-range gait recognition across diverse outdoor distances and environments. We further propose EMGaitNet, an end-to-end framework tailored for long-range multimodal gait recognition. To bridge the modality gap between RGB images and point clouds, we introduce a semantic-guided fusion pipeline. A CLIP-based Semantic Mining (SeMi) module first extracts human body-part-aware semantic cues, which are then employed to align 2D and 3D features via a Semantic-Guided Alignment (SGA) module within a unified embedding space. A Symmetric Cross-Attention Fusion (SCAF) module hierarchically integrates visual contours and 3D geometric features, and a Spatio-Temporal (ST) module captures global gait dynamics. Extensive experiments on various gait datasets validate the effectiveness of our method.
Zhiyang Lu, Tianren Wu, Changwang Zhang, Ming Cheng 0002
AAAI7
2026 Physically-Based LiDAR Smoke Simulation for Robust 3D Object Detection
abstract
3D object detection in adverse weather is crucial for autonomous driving, especially in smoke where LiDAR data becomes sparse and noisy. Due to the lack of real smoke data, this paper introduces a physics-based simulation framework to generate realistic LiDAR point clouds of smoke and augment large-scale driving datasets. First, we present a 3D fluid dynamics-based smoke simulation framework in Unity, which models the realistic spatial diffusion and temporal evolution of smoke particles. Coupled with a physically accurate LiDAR perception module, our system captures complex light interactions—such as beam attenuation, scattering, and multi-path effects—to generate high-fidelity, physically consistent smoke point clouds. Second, we propose a range image-based data fusion strategy that seamlessly integrates the simulated smoke point clouds into large-scale real-world LiDAR datasets (e.g., Waymo). This approach accurately emulates LiDAR scanning characteristics and naturally incorporates occlusion effects, enabling realistic smoke integration without compromising spatial consistency. To validate our approach, we collect a real-world LiDAR smoke dataset (LiSmoke) and conduct extensive experiments using state-of-the-art 3D detectors. Results demonstrate that models trained with our augmented synthetic data achieve significant improvements in smoke-affected scenarios, while maintaining competitive performance in clear-weather conditions. Our work provides a cost-effective solution for enhancing perception robustness in safety-critical environments.
Shijun Zheng, Weiquan Liu, Ming Cheng 0002, Cheng Wang 0003
AAAI6
2026 Language as a Bridge: Semantic-Guided Cross-Modal Gait Recognition via Text Prototype and Feature Decoupling
abstract
Gait recognition aims to identify individuals based on walking patterns in a long-range, contactless manner. While camera-based methods have advanced significantly, their performance deteriorates under poor lighting conditions. LiDAR offers a promising alternative by capturing accurate 3D gait information regardless of illumination. However, effectively integrating heterogeneous data from diverse sensors, such as LiDAR and cameras, remains a key challenge for cross-modal gait recognition. Existing approaches often minimize modality discrepancy directly, which can lead to class collapse and damage to inter-class discriminability. To overcome these limitations, we propose a Semantic-Guided Cross-modal Gait recognition framework, SG-CrossGait, that introduces text features as the prototype space to bridge camera and LiDAR modalities. We design structured Gait Description Factors (GDF) and leverage multimodal large language models (MLLMs) for automatic factor annotation and text generation, enriching existing datasets with textual descriptions, yielding SUSTech1K-Text and FreeGait-Text. A CLIP-based pipeline aligns multi-grained representations from both modalities to the text prototype space. We further propose the Dual-stream Cross-attention Fusion (DCF) module for fine-grained feature integration and the Semantic-Guided Feature Decoupling (SGFD) module to disentangle shared and modality-specific features. A Multi-task Training (MT) scheme incorporating Gait Attribute Recognition (GAR) further enhances intra-class compactness. Extensive experiments validate the effectiveness of our approach. On SUSTech1K-Text, our method achieves 61% accuracy in LiDAR-to-Camera recognition, outperforming the state-of-the-art method by 8.3%. We also release the Gait-Text benchmark to promote future research at the intersection of gait analysis and vision-language learning. Code and datasets are available at: https://github.com/O-VIGIA/SCCG.git.
Zhiyang Lu, Wankang Zeng, Ming Cheng 0002, Cheng Wang 0003
IEEE Trans. Inf. Forensics Secur.3
2026 UniCrossGait: Unified Cross-Modal Gait Recognition Based on Knowledge Distillation
abstract
Gait recognition provides a crucial biometric identification technique for intelligent security since it is long-range and nonintrusive. Previous methods have generally focused on gait recognition within a single modality, but with the constant advancement of multimedia sensors, there is a growing requirement for the development of multimodal and cross-modal gait recognition. The various modalities are aligned directly in a hard manner by current cross-modal approaches, which could result in class collapse caused by imbalanced learning and make learning discriminative representations difficult for the weaker modality. To overcome these limitations, we propose a unified cross-modal gait recognition framework, UniCrossGait. Specifically, we employ the knowledge distillation paradigm to transfer the knowledge of a fused multimodal teacher network, Exist Methods UniCrossGait(Ours) Figure 1: Illustration of comparison between exist methods and which serves as a shared intermediate representation, to unimodal student networks, which enables them to learn modality-invariant representations. To perform knowledge distillation effectively, we design the Direction-level Feature Imitation loss and Intra&Inter sample Correlation loss. Furthermore, we construct the Warm up operator and Distillation Balancing operator to prevent the network from falling into suboptimal solutions due to pure imitation and imbalanced distillation. Our method achieves state of-the-art (SOTA) results on cross-modal gait recognition datasets, and extensive ablation experiments demonstrate the effectiveness of the proposed paradigm and modules. Our code is available at https://github.com/O-VIGIA/UniCrossGait.
Zhiyang Lu, Ming Cheng 0002, Cheng Wang 0003
IEEE Trans. Multim.3
2025 SSRFlow: Semantic-Aware Fusion with Spatial Temporal Re-Embedding for Real-World Scene Flow
abstract
Scene flow, which provides the 3D motion field of the first frame from two consecutive point clouds, is vital for dynamic scene perception. However, contemporary scene flow methods face three major challenges. Firstly, they only consider the context of individual point clouds before flow embedding, leading to embedded points struggling to perceive the consistent semantic relationship of another frame. To address this issue, we propose a novel approach called Dual Cross Attentive (DCA) for the latent fusion and alignment between two frames based on semantic contexts. This is then integrated into Global Fusion Flow Embedding (GF) to initialize flow embedding based on global correlations in both contextual and Euclidean spaces. Secondly, deformations exist in non-rigid objects after the warping layer, which distorts the spatiotemporal relation between the consecutive frames. For a more precise estimation of residual flow at next-level, the Spatial Temporal Re-embedding (STR) module is devised to update the point sequence features at current-level. Lastly, poor generalization is often observed due to the significant domain gap between synthetic and LiDAR-scanned datasets. We leverage novel domain adaptive losses to effectively bridge the gap of motion inference from synthetic to real-world. Experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance across various datasets, with particularly outstanding results in real-world LiDAR-scanned situations.
Zhiyang Lu, Qinghan Chen, Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003
3DV5
2025 AENet: An Asymmetric Encoder Network with Feature Selection for Change Detection
abstract
Change detection (CD) plays a crucial role in applications such as disaster monitoring, urban planning, and natural resource management. Recent advances in satellite imaging technology have led to the availability of bi-temporal images with very high resolution, making binary CD increasingly popular. Although deep learning methods have outperformed traditional approaches in CD tasks, most existing models rely on symmetric dual-branch encoders with shared parameters. However, these architectures may find it increasingly complex to further effectively capture essential CD features from bi-temporal images. To address this limitation, we propose an innovative CD information encoding structure called the Asymmetric Encoder Network (AENet). AENet comprises two asymmetric encoders and a CD information fusion module. Unlike traditional symmetric architectures, our approach enables different CD information to be selected from each other during the fusion process, capturing key change information that enhances the sensitivity of the network. Additionally, we introduce a feature extraction module for multi-scale spatial aggregation applied to this structure. Experimental results demonstrate that AENet achieves higher F1 scores on three publicly available datasets while maintaining moderate computational complexity and fewer parameters.
Zhentao Xiong, Ming Cheng 0002, Youliang Chu, Cheng Wang 0003
IJCNN2
2025 MOJO: MOtion Pattern Learning and JOint-Based Fine-Grained Mining for Person Re-Identification Based on 4D LiDAR Point Clouds
Zhiyang Lu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003
IEEE Trans. Inf. Forensics Secur.3
2024 DiffLoc: Diffusion Model for Outdoor LiDAR Localization
abstract
Absolute pose regression (APR) estimates global pose in an end-to-end manner, achieving impressive results in learn-based LiDAR localization. However, compared to the top-performing methods reliant on 3D-3D correspondence matching, APR's accuracy still has room for improvement. We recognize APR's lack of robust features learning and iterative denoising process leads to suboptimal results. In this paper, we propose DiffLoc, a novel framework that formulates LiDAR localization as a conditional generation of poses. First, we propose to utilize the foundation model and static-object-aware pool to learn robust features. Second, we incorporate the iterative denoising process into APR via a diffusion model conditioned on the learned geometrically robust features. In addition, due to the unique nature of diffusion models, we propose to adapt our models to two additional applications: (1) using multiple inferences to evaluate pose uncertainty, and (2) seamlessly introducing geometric constraints on denoising steps to improve prediction accuracy. Extensive experiments conducted on the Oxford Radar RobotCar and NCLT datasets demonstrate that DiffLoc outperforms better than the state-of-the-art methods. Especially on the NCLT dataset, we achieve 35% and 34.7% improvement on position and orientation accuracy, respectively. Our code is released at https://github.com/liw95/DiffLoc.
Wen Li 0005, Yuyang Yang, Shangshu Yu, Guosheng Hu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003
CVPR6
2024 Density-guided Translator Boosts Synthetic-to-Real Unsupervised Domain Adaptive Segmentation of 3D Point Clouds
abstract
3D synthetic-to-real unsupervised domain adaptive seg-mentation is crucial to annotating new domains. Self-training is a competitive approach for this task, but its performance is limited by different sensor sampling patterns (i.e., variations in point density) and incomplete training strate-gies. In this work, we propose a density-guided translator (DGT), which translates point density between domains, and integrates it into a two-stage self-training pipeline named DGT-ST. First, in contrast to existing works that simulta-neously conduct data generation and feature/output align-ment within unstable adversarial training, we employ the non-learnable DGT to bridge the domain gap at the in-put level. Second, to provide a well-initialized model for self-training, we propose a category-level adversarial net-work in stage one that utilizes the prototype to prevent neg-ative transfer. Finally, by leveraging the designs above, a domain-mixed self-training method with source-aware consistency loss is proposed in stage two to narrow the domain gap further. Experiments on two synthetic-to-real segmentation tasks (SynLiDAR → semanticKITTI and SynL- iDAR → semanticPOSS) demonstrate that DGT-ST outper-forms state-of-the-art methods, achieving 9.4% and 4.3% mIoU improvements, respectively. Code is available at https://github.com/yuan-zm/DGT-ST.
Zhimin Yuan, Wankang Zeng, Yanfei Su, Weiquan Liu, Ming Cheng 0002, Yulan Guo, Cheng Wang 0003
CVPR5
2024 Bridging LiDAR Gaps: A Multi-LiDARs Domain Adaptation Dataset for 3D Semantic Segmentation
Shaoyang Chen, Bochun Yang, Yan Xia 0003, Ming Cheng 0002, Cheng Wang 0003
IJCAI4
2024 An Entropy-Based Pseudo-Label Mixup Method for Source-Free Domain Adaptation
Qinghan Chen, Zhiyang Lu, Ming Cheng 0002
PRCV (2)3
2024 Domain adaptive remote sensing image semantic segmentation with prototype guidance
Wankang Zeng, Ming Cheng 0002, Zhimin Yuan, Youming Wu, Weiquan Liu, Cheng Wang 0003
Neurocomputing2
2024 Multilevel Interactive Enhanced Network for Infrared Small-Target Detection
abstract
Infrared small target detection (IRSTD) aims to identify small and faint targets amidst cluttered background in infrared images, which is vital for applications like maritime surveillance. Traditional methods struggle due to low signal-to-noise ratio (SNR) and contrast. However, recent CNN-based approaches show promise, leveraging deep learning’s strong modeling capabilities. In this letter, we propose a multilevel interactive enhanced network (MIE-Net). In MIE-Net, we use multiple backbones that have progressively decreasing numbers of blocks. Features transfer and information interaction are carried out between different backbones. We designed an attention mechanism-based feature filter (AFF) to reduce background noise interference by filtering the low-level features with high-level features. Furthermore, we proposed a global information enhancement module (GIEM), through which features are enhanced as they are delivered, while further mitigating the problem of small target loss. Experiments on public datasets validate the effectiveness of our method. MIE-Net outperforms the current state-of-the-art (SOTA) methods by approximately 6% in terms of the intersection over union (IoU). There was also about a 2% increase in average area under the curve (AUC).
Youliang Chu, Ming Cheng 0002, Zhiyang Lu, Zhentao Xiong, Cheng Wang 0003
IEEE Geosci. Remote. Sens. Lett.2
2024 GSDDet: Ground Sample Distance-Guided Object Detection for Remote Sensing Images
abstract
Object detection for remote sensing images (ODRSI) is an important task in computer vision. Effective algorithms inspired by oriented object detection have been proposed recently. However, a major challenge still remains. Different categories of objects may be similar under different scales, causing cross-scale confusion. Different from natural images, remote sensing images have a consistent scale within the same image, which is generally referred to as Ground Sample Distance (GSD). In this paper, we show, that GSD can be utilized to address the cross-scale confusion problem, and effectively boost the performance of ODRSI. Specifically, we propose GSDDet, which embeds the deep features that represent GSD constraints to decrease the cross-scale confusion between different object categories. In GSDDet, a deep GSD classification network is first designed to extract the GSD deep features from remote sensing images. Then, the GSD deep feature is coupled with an attention framework to detect multiple categories of objects. Due to the simplicity of our framework, GSDDet can be applied to improve both one-stage and two-stage methods. Experiments demonstrate that GSDDet outperforms state-of-the-art methods on challenging benchmarks, including DOTA-v1.0, DOTA-v1.5, and HRSC2016. The source code will be released upon publication.
Yunuo Yang, Cheng Wang 0003, Zhipeng Cai 0003, Pinqing Song, Guanjie Huang, Ming Cheng 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 SR-Adv: Salient Region Adversarial Attacks on 3D Point Clouds for Autonomous Driving
abstract
Autonomous driving safety based on LiDAR perception is increasingly becoming a hot spot. Specifically, 3D adversarial examples always make the prediction results of deep neural network models unpredictable, which poses a major security risk to autonomous driving systems. However, the vulnerability of 3D neural network models to adversarial examples is less explored. At present, existing adversarial attack methods often obtain 3D adversarial examples by perturbing the entire point cloud, which requires a large number of perturbed points, i.e. requires a large perturbation budget. In this paper, we propose a Salient Region Adversarial attack method (SR-Adv) to generate adversarial point clouds, by perturbing fewer regions and fewer points. To our knowledge, we are the first to propose region-based attacks for 3D point clouds. First, the proposed SR-Adv employs game theory to extract salient regions of point clouds. This mechanism assigns a value to each region to measure its importance to the 3D neural network model prediction results and realizes the vulnerability analysis of the 3D model. Second, we propose a novel optimization-based gradient attack algorithm to achieve adversarial attacks on salient regions. We evaluate the proposed SR-Adv attack method on the synthetic datasets ModelNet40 and ShapeNetPart as well as the real-world dataset KITTI and NuScenes. Experimental results show that the proposed SR-Adv achieves a state-of-the-art attack success rate and better imperceptibility by perturbing fewer points on 3D point clouds.
Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Ping Zhong 0001, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.7
2023 GMA3D: Local-Global Attention Learning to Estimate Occluded Motions of Scene Flow
Zhiyang Lu, Ming Cheng 0002
PRCV (2)2
2023 Adaptive local adversarial attacks on 3D point clouds
Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003
Pattern Recognit.6
2023 Multistage Scene-Level Constraints for Large-Scale Point Cloud Weakly Supervised Semantic Segmentation
abstract
Compared to fully supervised 3D large-scale point cloud segmentation methods, which necessitate extensive manual point-wise annotations, weakly supervised segmentation has emerged as a popular approach for significantly reducing labeling costs while maintaining effectiveness. However, the existing methods have exhibited inferior segmentation performance and unsatisfactory generalization capabilities in some scenarios with unique structures (e.g., building facades). In this paper, we propose an effective and generalized weakly supervised semantic segmentation framework, called multi-stage scene-level constraints (MSC), to solve the above problem. To address the issue regarding inadequate labeled data, we use pseudo-labels for unlabeled data and propose an uncertainty-guided adaptive reweighting strategy to reduce the negative impact of erroneous pseudo-labeled data on the model learning process. To address the class imbalance issue, we employ multi-stage scene-level constraints (i.e., encoder, decoder, and classifier stages) to treat each class equally and improve perception ability of the model for each class. Evaluations conducted on multiple large-scale point cloud datasets collected in different scenarios, including building facades, indoor scenes, outdoor scenes, and UAV scenes, show that our MSC achieves a large gain over the existing weakly supervised methods and even surpasses some fully supervised methods.
Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 Spatial Adaptive Fusion Consistency Contrastive Constraint: Weakly Supervised Building Facade Point Cloud Semantic Segmentation
abstract
Semantic segmentation of building facade point clouds has diverse applications. The development of semantic segmentation methods is inextricably linked to datasets. The available building facade datasets suffer from a lack of abundant semantic categories and data completeness. To compensate for these shortcomings, we propose a new building facade dataset characterized by various categories and relatively complete 3-D building facades. In addition, most existing methods focus on fully supervised learning, which relies on manually labeling large-scale point cloud data and results in high time and labor costs. In this article, we propose an effective weakly supervised building facade segmentation approach, called spatial adaptive fusion consistency contrastive constraint (SAF-C3), to solve the above problem. We first design a multirandom point cloud augmentor as an auxiliary supervision branch to enhance the learning ability of the original network branch. Then, we present a spatial adaptive fusion (SAF) module to extract discriminative features for building facade point clouds. Finally, we propose a spatial consistency contrastive constraint to explore the contrastive property in feature space and to ensure the predictive consistency among the augmentation and original branches. The proposed method achieves a significant performance improvement against the state-of-the-art methods on two building facade point cloud datasets through extensive experiments. In particular, the performance of SAF-C3 with 1% labels significantly surpasses the baseline network with 100% labels.
Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Zhihong Zhang 0001, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 Novel Convolutions for Semantic Segmentation of Remote Sensing Images
abstract
The networks are required to be capable of learning low-level features well when applied to remote sensing image semantic segmentation task. To capture accurate and abundant low-level semantic information, the early feature extractor layer is crucial to the whole network, because all the subsequent features are inferred from that base. To address the low-level feature extraction issue and overcome the shortcomings of traditional convolution like too many parameters or limited receptive field, some novel convolution units have been proposed in the literature. In this paper, we propose two elaborately designed and portable yet effective convolution units, i.e., directional convolution (DC) and large field convolution (LFC), combined as the extractor of low-level semantic features. DC is designed to extract directional features from specific directions, and LFC can achieve a large receptive field with few parameters. Experiment results on two public datasets provide evidence that our convolution units can help deep learning networks improve performance stably and comprehensively compared to the baseline networks.
Ruijie Xiao, Chuan Zhong, Wankang Zeng, Ming Cheng 0002, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.4
2023 Prototype-Guided Multitask Adversarial Network for Cross-Domain LiDAR Point Clouds Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) segmentation aims to leverage labeled source data to make accurate predictions on unlabeled target data. The key is to make the segmentation network learn domain-invariant representations. In this work, we propose a prototype-guided multitask adversarial network (PMAN) to achieve this. First, we propose an intensity-aware segmentation network (IAS-Net) that leverages the private intensity information of target data to substantially facilitate feature learning of the target domain. Second, the category-level cross-domain feature alignment strategy is introduced to flee the side effects of global feature alignment. It employs the prototype (class centroid) and includes two essential operations: 1) build an auxiliary nonparametric classifier to evaluate the semantic alignment degree of each point based on the prediction consistency between the main and auxiliary classifiers and 2) introduce two class-conditional point-to-prototype learning objectives for better alignment. One is to explicitly perform category-level feature alignment in a progressive manner, and the other aims to shape the source feature representation to be discriminative. Extensive experiments reveal that our PMAN outperforms state-of-the-art results on two benchmark datasets.
Zhimin Yuan, Ming Cheng 0002, Wankang Zeng, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 STCLoc: Deep LiDAR Localization With Spatio-Temporal Constraints
abstract
LiDAR localization is of great importance to autonomous vehicles and robotics. Absolute pose regression, directly estimating the mapping from a scene to a 6-DoF pose, has achieved impressive results in learning-based localization. Different from traditional map-based methods, it does not need a pre-built 3D map during inference. However, current regression networks typically suffer from scene ambiguities, especially in challenging traffic environments, leading to large wrong predictions (e.g., outliers) and limited applications. To address this problem, a novel LiDAR localization framework with spatio-temporal constraints is proposed, termed STCLoc, to reduce scene ambiguities and achieve more accurate localization. First, we propose to regularize regression in the spatial dimension with a novel classification task to reduce outliers. Specifically, the classification task categorizes the point cloud in terms of position and orientation and then couples it with the regression task to conduct multi-task learning. Second, to learn discriminative features to reduce scene ambiguities, we propose using attention-based feature aggregation to capture the correlation in LiDAR sequences. We conduct extensive experiments on two benchmark datasets, where the localization takes 97ms on each dataset. Results show that our model outperforms state-of-the-art methods by 43.33%/36.76% (position/orientation) on the Oxford Radar RobotCar dataset, verifying the effectiveness of our method. The source code is available on the project website athttps://github.com/PSYZ1234/STCLoc.
Shangshu Yu, Cheng Wang 0003, Yitai Lin, Chenglu Wen, Ming Cheng 0002, Guosheng Hu
IEEE Trans. Intell. Transp. Syst.5
2023 Category-Level Adversaries for Outdoor LiDAR Point Clouds Cross-Domain Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) is a low-cost way to deal with the lack of annotations in a new domain. For outdoor point clouds in urban transportation scenes, the mismatch of sampling patterns and the transferability difference between classes make cross-domain segmentation extremely difficult. To overcome these challenges, we propose a category-level adversarial framework. Firstly, we propose a multi-scale domain conditioned block that facilitates to extract the critical low-level domain-dependent knowledge and reduce the domain gap caused by distinct LiDAR sampling patterns. Secondly, we make full use of multiple representation forms (i.e., point-based sets and voxel-based cells) and utilize the prediction consistency between the two forms to measure how well each point is semantically aligned. The model then focuses on the poorly-aligned points without affecting the well-aligned points. Experimental results on three autonomous driving point cloud datasets show that the proposed method outperforms existing methods by a large margin, especially on the low-beam to high-beam cross-domain segmentation task.
Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.3
2022 FedRME: Federated Road Markings Extraction from Mobile LiDAR Point Clouds
abstract
Road markings extraction (RME) from 3D point clouds acquired by mobile LiDAR systems has been widely used for road safety and autonomous driving. However, due to the increasing awareness of personal data protection and national informat ion security regulations, most autonomous driving companies are not willing to share their private point clouds data with the community. Therefore, such restriction of centralized training might inevitably inhibit the effectiveness of RME procedure. Federated learning (FL) is a distributed machine learning architecture that could address the aforementioned privacy-accuracy dilemma to collaboratively learn a global RME model from multiple clients without sharing raw data. In this paper, we propose a novel FedRME, a federated road markings extraction system to collaboratively learn a global RME model with multiple privacy-preserved local models from 3D mobile LiDAR point clouds. FedRME adopt the classical FedAvg model to construct a generalizable global feature embedding model without accessing local data. Moreover, to tackle data heterogeneity problem that local models vary in point clouds volumes and categories, we design a dynamic weighting mechanism to optimize the cooperative training effectiveness before server aggregation. Experimental results on three real-world mobil e LiDAR point clouds datasets with federated learning settings demonstrate that FedRME not only achieves superior performance but also reduces computation by up to 25%.The source code is available at https://github.com/WwZzz/easyFL#FedRME.
Xiaoliang Fan, Haibing Jin, Xiaotian Sun 0005, Ming Cheng 0002, Cheng Wang 0003
CSCWD5
2022 Multi-Graph Fusion Networks for Urban Region Embedding
abstract
Learning the embeddings for urban regions from human mobility data can reveal the functionality of regions, and then enables the correlated but distinct tasks such as crime prediction. Human mobility data contains rich but abundant information, which yields to the comprehensive region embeddings for cross domain tasks. In this paper, we propose multi-graph fusion networks (MGFN) to enable the cross domain prediction tasks. First, we integrate the graphs with spatio-temporal similarity as mobility patterns through a mobility graph fusion module. Then, in the mobility pattern joint learning module, we design the multi-level cross-attention mechanism to learn the comprehensive embeddings from multiple mobility patterns based on intra-pattern and inter-pattern messages. Finally, we conduct extensive experiments on real-world urban datasets. Experimental results demonstrate that the proposed MGFN outperforms the state-of-the-art methods by up to 12.35% improvement. https://github.com/wushangbin/MGFN
Shangbin Wu, Xiaoliang Fan, Shirui Pan, Chuanpan Zheng, Ming Cheng 0002, Cheng Wang 0003
IJCAI7
2022 2D3D-MVPNet: Learning cross-domain feature descriptors for 2D-3D matching based on multi-view projections of point clouds
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xiaoliang Fan, Yangbin Lin, Xuesheng Bian, Shangbin Wu, Ming Cheng 0002, Jonathan Li 0001
Appl. Intell.8
2022 Planar Primitive Group-Based Point Cloud Registration for Autonomous Vehicle Localization in Underground Parking Lots
abstract
We present a registration strategy based on a planar primitive group for indoor environments between a large-scale point cloud and a small-scale point cloud, providing a localization solution that is fully independent of prior information about the initial positions of the two point cloud coordinate systems. The algorithm first divides the point cloud into planes by region growing and then refines the plane boundaries by local$k$-means clustering. The planes are grouped to obtain planar primitive groups that are used to infer potential matching regions for an effective coarse registration and then fine-tuning. The algorithm is applied for autonomous vehicle localization in underground garages. Evaluation on four datasets shows that the algorithm can provide decimeter-level localization accuracy in seconds.
Lili Lin, Ming Cheng 0002, Chenglu Wen, Cheng Wang 0003
IEEE Geosci. Remote. Sens. Lett.3
2022 Local Fusion Attention Network for Semantic Segmentation of Building Facade Point Clouds
abstract
Automatic building facade point cloud semantic segmentation is an important step in 3-D urban building reconstruction. How to correctly segment the components (e.g., windows, walls, and columns) from the building facade is still a challenging task. According to the characteristics of building facade point clouds, we introduce local fusion attention network (LFA-Net), an efficient neural network that learns LFA features from building facade point clouds, for better capturing the local neighborhood structure information of each point. The core of LFA-Net is the LFA module, which consists of three neural units: local graph attention (LGA), local aggregation attention (LAA), and fusion attention (FA). The LFA-Net is the standard encoder-decoder architecture. Experiments demonstrate that our LFA-Net outperforms the state-of-the-art methods on the large-scale building facade point cloud dataset.
Yanfei Su, Weiquan Liu, Ming Cheng 0002, Zhimin Yuan, Cheng Wang 0003
IEEE Geosci. Remote. Sens. Lett.3
2022 Learning scale awareness in keypoint extraction and description
Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Yifan Peng 0001, Chenglu Wen, Ming Cheng 0002
Pattern Recognit.7
2022 DLA-Net: Learning dual local attention features for semantic segmentation of large-scale building facade point clouds
Yanfei Su, Weiquan Liu, Zhimin Yuan, Ming Cheng 0002, Zhihong Zhang 0001, Xuelun Shen, Cheng Wang 0003
Pattern Recognit.4
2022 LiDAR-based localization using universal encoding and memory-aware regression
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003
Pattern Recognit.4
2022 Corrigendum to "LiDAR-based localization using universal encoding and memory-aware regression" Pattern Recognition Volume 128 (2022) 108685
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003
Pattern Recognit.4
2022 Dense Point Cloud Completion Based on Generative Adversarial Network
abstract
Point cloud completion aims to reconstruct complete point clouds from partial point clouds, which is widely used in various fields such as autonomous driving and robotics. Most existing methods are sparse point cloud completion, where the number of point clouds after completion is relatively small and the details are insufficient. This article proposes a novel end-to-end generative adversarial network-based dense point cloud completion architecture (DPCG-Net). We design two generative adversarial network (GAN)-based modules that translate point cloud completion into mapping between global feature distributions obtained by encoding partial point clouds and ground truth, respectively. The first designed generator module proposes skip connections to fully connected layer-based network for regenerating global feature and changing the global feature distribution derived from the encoder module to approximate the ground truth global feature distribution. The second proposed discriminator module divides high-dimensional global feature vectors into several smaller batches for judgment to guarantee the similarity between the regenerated global feature and the ground truth. We perform quantitative and qualitative experiments on the ShapeNet and KITTI datasets. Experiments on ShapeNet demonstrate that our model outperforms other models in cases where the lack of a large proportion of point clouds results in a large loss of spatial structure, especially when 80% of point clouds are missing. Moreover, KITTI experiments reveal that it is also valid for realistic situations. In addition, application in classification shows that the classification accuracy of point clouds completed with DPCG-Net is as high as 86.5% under the condition of 80% missing point clouds.
Ming Cheng 0002, Guoyan Li, Yiping Chen 0002, Jun Chen 0005, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 DFAN: Dual-Branch Feature Alignment Network for Domain Adaptation on Point Clouds
abstract
Unsupervised domain adaptation (UDA) significantly reduces the gap between the source domain and the target domain in machine learning and computer vision tasks. Most UDA approaches are applied to images and videos, and only a few methods implement domain adaptation on 3-D computer vision problems. The existing UDA approaches operating on point clouds try to extract domain-invariant features in different domains for feature alignment. However, higher commonality brings less diversity and results a loss of detailed information. In this article, we propose a novel dual-branch feature alignment network (DFAN) architecture for domain adaptation on point cloud visual tasks to better exploit the respective characteristics of local and global features. Our approach specializes in the extraction and alignment of global and local features with different strategies in each branch to complement each other. We also introduce a hierarchical alignment strategy for local feature alignment and a distribution alignment strategy for global feature alignment. Experiments on the PointDA-10 and PointSegDA datasets show that our approach achieves state-of-the-art performance on the UDA of point cloud classification and segmentation tasks. The ablation study demonstrates the effectiveness of the dual-branch design and the feature alignment strategies.
Liangwei Shi, Zhimin Yuan, Ming Cheng 0002, Yiping Chen 0002, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.3
2022 GCN-Based Pavement Crack Detection Using Mobile LiDAR Point Clouds
abstract
Mobile Laser Scanning (MLS) system can provide high-density and accurate 3D point clouds that enable rapid pavement crack detection for road maintenance tasks. Supervised learning-based algorithms have been proved pretty effective for handling such a large amount of inhomogeneous and unstructured point clouds. However, these algorithms often rely on a lot of annotated data, which is labor-intensive and time-consuming. This paper presents a semi-supervised point-level approach to overcome this challenge. We propose a graph-widen module to construct a reasonable graph structure for point clouds, increasing the detection performance of graph convolutional networks (GCN). The constructed graph characterizes the local features from a small amount of annotated data, avoiding information loss and dramatically reduces the dependence on annotated data. The MLS point clouds acquired by a commercial RIEGL VMX-450 system are used in this study. The experimental results demonstrate that our method outperforms the state-of-the-art point-level methods in terms of recall, F1 score, and efficiency while achieving comparable accuracy.
Huifang Feng 0002, Wen Li 0005, Yiping Chen 0002, Sarah Narges Fatholahi, Ming Cheng 0002, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.6
2021 Predicting the spread of COVID-19 in China with human mobility data
abstract
The coronavirus disease 2019 (COVID-19) break-out in late December 2019 has spread rapidly worldwide. Existing studies have shown that there is a significant correlation between large-scale human movements and the spread of the epidemic. However, there is a lack of quantification of these correlations, and it is still challenging to predict the spread of the epidemic at early stage. In this paper, we address this issue by conducting a statistical analysis on the spatio-temporal relationship between human mobility and the epidemic spread. Specifically, we proposed an improved SEIR model to adapt to the COVID-19 epidemic, so that we can predict the spread of the epidemic at the early stage using human mobility data and the early confirmed cases. We evaluated our model in various provinces and cities in China, and the results are superior to various baselines, verifying the effectiveness of the method.
Shangbin Wu, Xiaoliang Fan, Longbiao Chen, Ming Cheng 0002, Cheng Wang 0003
SIGSPATIAL/GIS4
2021 Learning Cross-Domain Descriptors for 2D-3D Matching with Hard Triplet Loss and Spatial Transformer Network
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Yanfei Su, Xiuhong Lin, Zhimin Yuan, Ming Cheng 0002
ICIG (3)9
2021 Y-Net: Learning Domain Robust Feature Representation for ground camera image and large-scale image-based point cloud registration
Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Baiqi Lai, Xuelun Shen, Ming Cheng 0002, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001
Inf. Sci.7
2019 LO-Net: Deep Real-Time Lidar Odometry
abstract
We present a novel deep convolutional network pipeline, LO-Net, for real-time lidar odometry estimation. Unlike most existing lidar odometry (LO) estimations that go through individually designed feature selection, feature matching, and pose estimation pipeline, LO-Net can be trained in an end-to-end manner. With a new mask-weighted geometric constraint loss, LO-Net can effectively learn feature representation for LO estimation, and can implicitly exploit the sequential dependencies and dynamics in the data. We also design a scan-to-map module, which uses the geometric and semantic information learned in LO-Net, to improve the estimation accuracy. Experiments on benchmark datasets demonstrate that LO-Net outperforms existing learning based approaches and has similar accuracy with the state-of-the-art geometry-based approach, LOAM.
Qing Li 0032, Shaoyang Chen, Cheng Wang 0003, Xin Li 0003, Chenglu Wen, Ming Cheng 0002, Jonathan Li 0001
CVPR6
2019 RF-Net: An End-To-End Image Matching Network Based on Receptive Field
abstract
This paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable matching framework is desirable and challenging. The very recent approach, LF-Net, successfully embeds the entire feature extraction pipeline into a jointly trainable pipeline, and produces the state-of-the-art matching results. This paper introduces two modifications to the structure of LF-Net. First, we propose to construct receptive feature maps, which lead to more effective keypoint detection. Second, we introduce a general loss function term, neighbor mask, to facilitate training patch selection. This results in improved stability in descriptor training. We trained RF-Net on the open dataset HPatches, and compared it with other methods on multiple benchmark datasets. Experiments show that RF-Net outperforms existing state-of-the-art methods.
Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Zenglei Yu, Jonathan Li 0001, Chenglu Wen, Ming Cheng 0002
CVPR7
2019 NormalNet: A voxel-based CNN for 3D object classification and retrieval
Cheng Wang 0003, Ming Cheng 0002, Ferdous Sohel, Mohammed Bennamoun, Jonathan Li 0001
Neurocomputing2
2018 A Modified Framework for Ship Detection from Compact Polarization SAR Image
abstract
In recent years, the compact polarimetric SAR (CP SAR) imaging mode has received much attention because of its advantages in swath width compared to the quad-polarization mode and in information of scattering targets compared to the linear dual-polarization mode, respectively. For object detection (e.g., ship, sea ice, and drilling platform) in maritime monitoring which has been widely discussed, it is still a challenge to eliminate the false alarms caused by ocean clutter. Accordingly, a modified feature-based framework for ship detection using CP SAR data was proposed in this paper. In particular, a Guard-filter was used after feature extraction to reduce the effect of ocean clutter. Several simulated CP SAR imageries based on Gaofen-3 quad-polarization SAR image were selected for further investigations on ship detection. A SVM classifier was included in the framework to do a pixel-wise classification. The result shows that the proposed method has a better performance compared to traditional methods (e.g., constant false alarm rate and polarimetric whitening filter) and the feature-based method without filtering. Overall false alarm rates are decreased by 35% with the proposed method.
Qiancong Fan, Feng Chen 0022, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS3
2018 3-D Road Boundary Extraction From Mobile Laser Scanning Data via Supervoxels and Graph Cuts
abstract
Effective extraction of road boundaries plays a significant role in intelligent transportation applications, including autonomous driving, vehicle navigation, and mapping. This paper presents a new method to automatically extract 3-D road boundaries from mobile laser scanning (MLS) data. The proposed method includes two main stages: supervoxel generation and 3-D road boundary extraction. Supervoxels are generated by selecting smooth points as seeds and assigning points into facets centered on these seeds using several attributes (e.g., geometric, intensity, and spatial distance). 3-D road boundaries are then extracted using the α-shape algorithm and the graph cuts-based energy minimization algorithm. The proposed method was tested on two data sets acquired by a RIEGL VMX-450 MLS system. Experimental results show that road boundaries can be robustly extracted with an average completeness over 95%, an average correctness over 98%, and an average quality over 94% on two data sets. The effectiveness and superiority of the proposed method over the state-of-the-art methods is demonstrated.
Dawei Zai, Jonathan Li 0001, Yulan Guo, Ming Cheng 0002, Yangbin Lin, Huan Luo 0001, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.4
2017 Tree Classification in Complex Forest Point Clouds Based on Deep Learning
abstract
Recently, the classification of tree species using 3-D point clouds has drawn wide attention in surveys and forestry investigations. This letter proposes a new voxel-based deep learning method to classify tree species in 3-D point clouds collected from complex forest scenes. The proposed method includes three steps: 1) individual tree extraction based on the density of the point clouds; 2) low-level feature representation through voxel-based rasterization; and 3) classification of tree species by a deep learning model. Two data sets of 3-D forest point clouds acquired by terrestrial laser scanning systems are used to evaluate the proposed method. The method achieves an average classification accuracy of 93.1% and 95.6% on the two data sets. Furthermore, in comparative experiments, the proposed method exhibits performance superior to that of the other 3-D tree species classification methods.
Xinhuai Zou, Ming Cheng 0002, Cheng Wang 0003, Yan Xia 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2017 Traffic Sign Occlusion Detection Using Mobile Laser Scanning Point Clouds
abstract
For survey and maintenance of traffic signs, this paper presents a novel traffic sign occlusion detection method using 3-D point clouds and trajectory data acquired by a mobile laser scanning system. To produce a maintenance guide, our method aims to obtain the degree of occlusion by analyzing the spatial relationship between traffic signs, surroundings, and drivers on the road. First, a detection method considering both reflectance and geometric features is developed to capture traffic signs. Next, to simulate the driver's view, a trajectory-based method is proposed to determine driver's observation location and the corresponding observed traffic sign. Finally, to determine whether a traffic sign is in occlusion, a hidden point removal algorithm is adopted and carried out. Furthermore, we develop two indices to evaluate the degree of occlusion. The proposed method is tested using two point cloud data sets collected by an RIEGL VMX-450 system along a 23.68-km-long urban road. The obtained results illustrate the feasibility of the proposed occlusion detection method.
Pengdi Huang, Ming Cheng 0002, Yiping Chen 0002, Huan Luo 0001, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.2
2016 Road network extraction via deep learning and line integral convolution
abstract
In this paper, we propose a learning-based road network extraction scheme from high resolution satellite. First, the convolutional neural network (CNN), which is able to capture large context of local structures, are applied to predict the probability of a pixel belonging to road regions, and assign labels to each pixel to describe whether it is road. Then, a line integral convolution based algorithm is developed to smooth the rough map to connect small gaps. Finally, by combining with some common image processing operators, road centerlines are able to be acquired. Attribute to the learning capacity of CNN, and the line integral convolution based connection scheme, the proposed road extraction method is able to provide high quality results comparing to current state-of-art road extraction methods.
Peikang Li, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002, Lun Luo
IGARSS5
2016 Superpixel-based coastline extraction in SAR images with speckle noise removal
abstract
Coastline extraction in Synthetic aperture radar (SAR) images is a fundamental and challenging task due to the speckle noise. In this paper, we propose a new method for automatic coastline extraction in SAR images. In our method, we combine K-means and speckle noise removal methods together to increase the dissimilarity between sea and land. To enhance the robustness to speckle noise, and preserve the targets boundaries, we treat superpixels as basic regions instead of pixels in traditional pixel-based methods. Finally, an adaptive threshold is applied to classify these regions into sea or land. Based on the classifications, a canny detector is employed to detect the coastline. We evaluate our proposed method on SAR images and the improved coastline extraction method superpixel-based is verified on remote sensing images with RGB channels. The experimental results demonstrate its superior performance on coastline extraction.
Xiaofang Liu, Hong Jia, Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002
IGARSS6
2016 Robust vehicle detection by combining deep features with exemplar classification
Liujuan Cao, Qilin Jiang, Ming Cheng 0002, Cheng Wang 0003
Neurocomputing3
2016 Local quality assessment of point clouds for indoor mobile mapping
Fangfang Huang, Chenglu Wen, Huan Luo 0001, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
Neurocomputing4
2015 Big Data Analytics and Visualization with Spatio-Temporal Correlations for Traffic Accidents
Xiaoliang Fan, Baoqin He, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002, Huaqiang Huang, Xiao Liu 0004
ICA3PP (2)5
2015 Turning mobile laser scanning points into 2D/3D on-road object models: Current status
abstract
Traditional road surveying methods rely largely on in-situ measurements, which are time consuming and labor intensive. Recent Mobile Laser Scanning (MLS) techniques enable collection of road data at a normal driving speed. However, extracting required information from collected MLS data remains a challenging task. This paper focuses on examining the current status of automated on-road object extraction techniques from 3D MLS points over the last five years. Several kinds of on-road objects are included in this paper: curbs and road surfaces, road markings, pavement cracks, as well as manhole and sewer well covers. We evaluate the extraction techniques according to their method design, degree of automation, precision, and computational efficiency. Given the large volume of MLS data, to date most MLS object extraction techniques are aiming to improve their precision and efficiency.
Zongliang Zhang, Ming Cheng 0002, Xinqu Chen, Menglan Zhou, Jonathan Li 0001, Hongshan Nie
IGARSS2
2015 Single-image super-resolution in RGB space via group sparse representation
abstract
Super‐resolution (SR) is the problem of generating a high‐resolution (HR) image from one or more low‐resolution (LR) images. This study presents a new approach to single‐image super‐resolution based on group sparse representation. Two dictionaries are constructed corresponding to the LR and HR image patches, respectively. The sparse coefficients of an input LR image patch in terms of the LR dictionary are used to recover the HR patch from the HR dictionary. When constructing the dictionaries, the three colour channels in a training image patch are considered a group composed of three atoms. The whole group is selected simultaneously when representing an image patch so that the correlations between the colour channels can be retained. A dictionary training method is also designed in which the two dictionaries are trained jointly to ensure that the corresponding LR and HR patches have the same sparse coefficients. Experimental results demonstrate the effectiveness of the proposed method and its robustness to noise.
Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
IET Image Process.1
2015 Combinative hypergraph learning for semi-supervised image classification
Binghui Wei, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
Neurocomputing2
2014 Earthwork volumes estimation in asphalt pavement reconstruction using a mobile laser scanning systerm
abstract
This paper presents a novel method for estimating earthwork volumes in asphalt pavement reconstruction using a mobile laser scanning (MLS) system. First, based on the static targets, this method registers two point cloud datasets into the same coordinate system, which respectively are acquired in the reconstructing road before and after asphalting. Next, road surface points are detected from each point cloud using a curb-based method, and further divided into a set of blocks. Afterwards, the blocks are perpendicularly partitioned into grids, where two surface features are extracted using the RANSAC. Finally, the volume of each grid is calculated according to these two surface features. The proposed algorithm has been tested on two sets of point clouds acquired by a RIEGL VMX-450 MLS system in the reconstructing road before and after asphalting. The results demonstrate the accuracy and efficiency of the proposed algorithm in estimating earthwork volumes.
Fukai Jia, Jonathan Li 0001, Cheng Wang 0003, Yongtao Yu, Ming Cheng 0002, Dawei Zai
IGARSS5
2014 Sparse Representation Based Pansharpening Using Trained Dictionary
abstract
Sparse representation has been used to fuse high-resolution panchromatic (HRP) and low-resolution multispectral (LRM) images. However, the approach faces the difficulty that the dictionary is generated from the high-resolution multispectral (HRM) images, which are unknown. In this letter, a two-step method is proposed to train the dictionary from the HRP and LRM images. In the first step, coarse HRM images are obtained by additive wavelet fusion method. The initial dictionary is composed of randomly sampled patches from the coarse HRM images. In the second step, a linear constraint K-SVD method is designed to train the dictionary to improve its representation ability. Experimental results using QuickBird and IKONOS data indicate that the trained dictionary yields comparable fusion products with raw-patch-dictionary sampled from HRM images.
Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2014 Object Detection in Terrestrial Laser Scanning Point Clouds Based on Hough Forest
abstract
This letter presents a novel rotation-invariant method for object detection from terrestrial 3-D laser scanning point clouds acquired in complex urban environments. We utilize the Implicit Shape Model to describe object categories, and extend the Hough Forest framework for object detection in 3-D point clouds. A 3-D local patch is described by structure and reflectance features and then mapped to the probabilistic vote about the possible location of the object center. Objects are detected at the peak points in the 3-D Hough voting space. To deal with the arbitrary azimuths of objects in real world, circular voting strategy is introduced by rotating the offset vector. To deal with the interference of adjacent objects, distance weighted voting is proposed. Large-scale real-world point cloud data collected by terrestrial mobile laser scanning systems are used to evaluate the performance. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 3-D object detection methods.
Hanyun Wang, Cheng Wang 0003, Huan Luo 0001, Peng Li 0064, Ming Cheng 0002, Chenglu Wen, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5