Yan Xia 0003

dblp:17/6518-3 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0001-6684-9814ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 TRASE: Tracking-Free 4D Segmentation and Editing
abstract
Understanding dynamic 3D scenes is crucial for extended reality (XR) and autonomous driving. Incorporating semantic information into 3D reconstruction enables holistic scene representations, unlocking immersive and interactive applications. To this end, we introduce TRASE, a novel tracking-free 4D segmentation method for dynamic scene understanding. TRASE learns a 4D segmentation feature field in a weakly-supervised manner, leveraging a soft-mined contrastive learning objective guided by SAM masks. The resulting feature space is semantically coherent and well-separated, and final object-level segmentation is obtained via unsupervised clustering. This enables fast editing, such as object removal, composition, and style transfer, by directly manipulating the scene's Gaussians. We evaluate TRASE on five dynamic benchmarks, demonstrating state-of-the-art segmentation performance from unseen viewpoints and its effectiveness across various interactive editing tasks. Our project page is available at: https://yunjinli.github.io/project-sadg/
Yun-Jin Li, Mariia Gladkova, Yan Xia 0003, Daniel Cremers
3DV3
2026 AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) achieve impressive performance once optimized on massive datasets. Such datasets often contain sensitive or copyrighted content, raising significant data privacy concerns. Regulatory frameworks mandating the 'right to be forgotten' drive the need for machine unlearning. This technique allows for the removal of target data without resource-consuming retraining. However, while well-studied for text, visual concept unlearning in MLLMs remains underexplored. A primary challenge is precisely removing a target visual concept without disrupting model performance on related entities. To address this, we introduce AUVIC, a novel visual concept unlearning framework for MLLMs. AUVIC applies adversarial perturbations to enable precise forgetting. This approach effectively isolates the target concept while avoiding unintended effects on similar entities. To evaluate our method, we construct VCUBench. It is the first benchmark designed to assess visual concept unlearning in group contexts. Experimental results demonstrate that AUVIC achieves state-of-the-art target forgetting rates while incurs minimal performance degradation on non-target concepts.
Jinhe Bi, Yan Xia 0003, Jindong Gu, Volker Tresp
AAAI5
2026 TARGO and TARGO-Net: Benchmarking Target-Driven Object Grasping Under Occlusions
abstract
Predicting 6-DoF grasp poses from a single RGB-D frame has recently achieved impressive accuracy, yet performance collapses when the target object is heavily occluded by clutter. In this paper, we establish the first benchmark dataset for TARget-driven Grasping under Occlusions, named TARGO, and our model that remains robust under occlusion, TARGO-Net. Our main contributions are: 1) We first recognize the visual occlusion challenge in 6-DoF grasping using single RGB-D images, and found that even the current SOTA models suffer under high occlusion. 2) We propose TARGO dataset, which can be used to train and test 6-DoF grasp models under different visual occlusion severities, and evaluate model robustness in real-world scenarios. 3) We further devise TARGO-Net, a transformer-based grasping model involving a target completion module and target-scene cross-attention, that performs most robustly across all visual occlusion levels. 4) We discover that other than visual occlusion, number of occluders and target minimum dimension also contribute to grasp success. The dataset and codes are publicly available at https://targo-benchmark.github.io .
Yan Xia 0003, Ziyuan Qin 0002, Guanqi Zhan, Kaichen Zhou, Hao Dong 0003, Daniel Cremers
Int. J. Comput. Vis.1
2025 VXP: Voxel-Cross-Pixel Large-Scale Camera-LiDAR Place Recognition
abstract
Cross-modal place recognition methods are flexible GPS-alternatives under varying environment conditions and sensor setups. However, this task is non-trivial since extracting consistent and robust global descriptors from different modalities is challenging. To tackle this issue, we propose Voxel-Cross-Pixel (VXP), a novel camera-to-LiDAR place recognition framework that enforces local similarities in a self-supervised manner and effectively brings global context from images and LiDAR scans into a shared feature space. Specifically, VXP is trained in three stages: first, we deploy a visual transformer to compactly represent input images. Secondly, we establish local correspondences between image-based and point cloud-based feature spaces using our novel geometric alignment module. We then aggregate local similarities into an expressive shared latent space. Extensive experiments on the three benchmarks (Oxford RobotCar, ViViD++ and KITTI) demonstrate that our method surpasses the state-of-the-art cross-modal retrieval by a large margin. Our evaluations show that the proposed method is accurate, efficient and light-weight. Our project page is available at: https://yunjinli.github.io/projects-vxp/.
Yun-Jin Li, Mariia Gladkova, Yan Xia 0003, Rui Wang 0037, Daniel Cremers
3DV3
2025 L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object Detection
abstract
LiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, it faces challenges of significant differences in terms of data quality and the degree of degradation in adverse weather. To address these issues, we introduce L4DR, a weather-robust 3D object detection method that effectively achieves LiDAR and 4D Radar fusion. Our L4DR proposes Multi-Modal Encoding (MME) and Foreground-Aware Denoising (FAD) modules to reconcile sensor gaps, which is the first exploration of the complementarity of early fusion between LiDAR and 4D radar. Additionally, we design an Inter-Modal and IntraModal ({IM}2) parallel feature extraction backbone coupled with a Multi-Scale Gated Fusion (MSGF) module to counteract the varying degrees of sensor degradation under adverse weather conditions. Experimental evaluation on a VoD dataset with simulated fog proves that L4DR is more adaptable to changing weather conditions. It delivers a significant performance increase under different fog levels, improving the 3D mAP by up to 20.0% over the traditional LiDAR-only approach. Moreover, the results on the K-Radar dataset validate the consistent performance improvement of L4DR in realworld adverse weather conditions.
Xun Huang 0003, Ziyu Xu 0002, Qiming Xia, Yan Xia 0003, Jonathan Li 0001, Kyle Gao, Chenglu Wen, Cheng Wang 0003
AAAI6
2025 SparseAlign: a Fully Sparse Framework for Cooperative Object Detection
abstract
Cooperative perception can increase the view field and decrease the occlusion of an ego vehicle, hence improving the perception performance and safety of autonomous driving. Despite the success of previous works on cooperative object detection, they mostly operate on dense Bird’s Eye View (BEV) feature maps, which are computationally demanding and can hardly be extended to long-range detection problems. More efficient fully sparse frameworks are rarely explored. In this work, we design a fully sparse framework, SparseAlign, with three key features: an enhanced sparse 3D backbone, a query-based temporal context learning module, and a robust detection head specially tailored for sparse features. Extensive experimental results on both OPV2V and DAIR-V2X datasets show that our framework, despite its sparsity, outperforms the state of the art with less communication bandwidth requirements. In addition, experiments on the OPV2Vt and DAIR-V2Xt datasets for time-aligned cooperative object detection also show a significant performance gain compared to the baseline works.
Yunshuang Yuan, Yan Xia 0003, Daniel Cremers, Monika Sester
CVPR2
2025 Localizing Events in Videos with Multimodal Queries
abstract
Localizing events in videos based on semantic queries is a pivotal task in video understanding research and user-oriented applications like video search. Yet, current research predominantly relies on natural language queries (NLQs), overlooking the potential of using multimodal queries (MQs) that incorporate images to flexibly represent semantic queries, particularly when it is difficult to express non-verbal or unfamiliar concepts in words. To bridge this gap, we introduce ICQ, a new benchmark designed for localizing events in videos with MQs, alongside an evaluation dataset ICQ-Highlight. To adapt and reevaluate existing video localization models for this new task, we propose 3 Multimodal Query Adaptation methods and a novel Surrogate Fine-Tuning strategy, serving as strong baseline methods. ICQ systematically benchmarks 12 state-of-the-art backbone models, spanning from specialized video localization models to Video Large Language Models. Our extensive experiments highlight the high potential of using MQs in real-world applications. We believe this is a first step toward video event localization with MQs1.
Gengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia 0003, Daniel Cremers, Philip Torr 0001, Volker Tresp, Jindong Gu
CVPR4
2025 CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Driving
abstract
28031
Rui Song 0007, Chenwei Liang, Yan Xia 0003, Walter Zimmer, Hu Cao, Holger Caesar, Andreas Festag, Alois C. Knoll
ICCV3
2025 Trafficloc: Localizing Traffic Surveillance Cameras in 3D Scenes
abstract
We tackle the problem of localizing traffic cameras within a 3D reference map and propose a novel image-to-point cloud registration (I2P) method, TrafficLoc, in a coarse-tofine matching fashion. To overcome the lack of large-scale real-world intersection datasets, we first introduce Carla Intersection, a new simulated dataset with 75 urban and rural intersections in Carla. We find that current I2P methods struggle with cross-modal matching under large viewpoint differences, especially at traffic intersections. TrafficLoc thus employs a novel Geometry-guided Attention Loss (GAL) to focus only on the corresponding geometric regions under different viewpoints during 2D-3D feature fusion. To address feature inconsistency in paired image patch-point groups, we further propose Inter-intra Contrastive Learning (ICL) to enhance separating 2D patch/3D group features within each intra-modality and introduce Dense Training Alignment (DTA) with soft-argmax for improving position regression. Extensive experiments show our TrafficLoc greatly improves the performance over the SOTA I2P methods (up to 86%) on Carla Intersection and generalizes well to real-world data. TrafficLoc also achieves new SOTA performance on KITTI and NuScenes datasets, demonstrating the superiority across both in-vehicle and traffic cameras. Our project page is publicly available at https://tum-luk.github.io/projects/trafficloc/.
Yan Xia 0003, Yunxiang Lu, Rui Song 0003, Oussema Dhaouadi, João F. Henriques, Daniel Cremers
ICCV1
2025 MonoCT: Overcoming Monocular 3D Detection Domain Shift with Consistent Teacher Models
abstract
We tackle the problem of monocular 3D object detection across different sensors, environments, and camera setups. In this paper, we introduce a novel unsupervised domain adaptation approach, MonoCT, that generates highly accurate pseudo labels for self-supervision. Inspired by our observation that accurate depth estimation is critical to mitigating domain shifts, MonoCT introduces a novel Generalized Depth Enhancement (GDE) module with an ensemble concept to improve depth estimation accuracy. Moreover, we introduce a novel Pseudo Label Scoring (PLS) module by exploring inner-model consistency measurement and a Diversity Maximization (DM) strategy to further generate high-quality pseudo labels for self-training. Extensive experiments on six benchmarks show that MonoCT outperforms existing SOTA domain adaptation methods by large margins (~21% minimum for AP Mod.) and generalizes well to car, traffic camera and drone views.
Johannes Meier, Louis Inchingolo, Oussema Dhaouadi, Yan Xia 0003, Jacques Kaiser, Daniel Cremers
ICRA4
2025 Feature-aligned Fisheye Object Detection Network for Autonomous Driving
abstract
Fisheye cameras, renowned for their panoramic field of view (FOV) of 360°, are crucial for surround-view perception in autonomous driving. However, research on object perception in fisheye images lags behind that of standard images. To address this gap, we propose a feature-aligned fisheye object detection network specifically tailored for autonomous driving. Current fisheye perception algorithms often overlook the misalignment issues that typically arise in object detectors. To tackle these challenges in the feature pyramid network (FPN), we introduce a feature-aligned pyramid module (FaPM), which learns pixel transformation offsets to contextually align feature maps. Additionally, we present a location-aligned detection head (LaDH) to align the spatial distribution of classification and regression localization. Integrating these modules into a detection framework results in a novel feature-aligned fisheye object detector. Our method undergoes extensive evaluation on the WoodScape dataset, achieving a mean average precision (mAP) of 32.2%, surpassing the performance of existing methods.
Hu Cao, Dongyi Sun, Rui Song 0007, Yan Xia 0003, Alois C. Knoll
IROS4
2025 L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing Imagery
abstract
We tackle the challenge of LiDAR-based place recognition, which traditionally depends on costly and time-consuming prior 3D maps. To overcome this, we first construct LiRSI-XA dataset, which encompasses approximately $110,000$ remote sensing submaps and $13,000$ LiDAR point cloud submaps captured in urban scenes, and propose a novel method, L2RSI, for cross-view LiDAR place recognition using high-resolution Remote Sensing Imagery. This approach enables large-scale localization capabilities at a reduced cost by leveraging readily available overhead images as map proxies. L2RSI addresses the dual challenges of cross-view and cross-modal place recognition by learning feature alignment between point cloud submaps and remote sensing submaps in the semantic domain. Additionally, we introduce a novel probability propagation method based on particle estimation to refine position predictions, effectively leveraging temporal and spatial information. This approach enables large-scale retrieval and cross-scene generalization without fine-tuning. Extensive experiments on LiRSI-XA demonstrate that, within a $100km^2$ retrieval range, L2RSI accurately localizes $83.27\%$ of point cloud submaps within a $30m$ radius for top-$1$ retrieved location. Our project page is publicly available at https://shizw695.github.io/L2RSI/.
Ziwei Shi, Yan Xia 0003, Cheng Wang 0003
NeurIPS4
2025 ZAHA: Introducing the Level of Facade Generalization and the Large-Scale Point Cloud Facade Semantic Segmentation Benchmark Dataset
abstract
Facade semantic segmentation is a long-standing challenge in photogrammetry and computer vision. Although the last decades have witnessed the influx of facade segmentation methods, there is a lack of comprehensive facade classes and data covering the architectural variability. In ZAHA11Project page: https://github.com/OloOcki/zaha, we introduce Level of Facade Generalization (LoFG), novel hierarchical facade classes designed based on international urban modeling standards, ensuring compatibility with real-world challenging classes and uniform methods' comparison. Realizing the LoFG, we present to date the largest semantic 3D facade segmentation dataset, providing 601 million annotated points at five and 15 classes of LoFG2 and LoFG3, respectively. More-over, we analyze the performance of baseline semantic segmentation methods on our introduced LoFG classes and data, complementing it with a discussion on the unresolved challenges for facade segmentation. We firmly believe that ZAHA shall facilitate further development of 3D facade semantic segmentation methods, enabling robust segmentation indispensable in creating urban digital twins.
Olaf Wysocki, Thomas Froech, Yan Xia 0003, Magdalena Wysocki, Ludwig Hoegner, Daniel Cremers, Christoph Holst
WACV4
2024 Text2Loc: 3D Point Cloud Localization from Natural Language
abstract
We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to 2 × over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at https://yan-xia.github.io/projects/text2loc/.
Yan Xia 0003, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers
CVPR1
2024 Embracing Events and Frames with Hierarchical Feature Refinement Network for Object Detection
Hu Cao, Zehua Zhang 0009, Yan Xia 0003, Jiahao Xia 0001, Guang Chen 0001, Alois C. Knoll
ECCV (84)3
2024 Boosting 3D Single Object Tracking with 2D Matching Distillation and 3D Pre-training
Qiangqiang Wu, Yan Xia 0003, Jia Wan 0001, Antoni B. Chan
ECCV (12)2
2024 Bridging LiDAR Gaps: A Multi-LiDARs Domain Adaptation Dataset for 3D Semantic Segmentation
Shaoyang Chen, Bochun Yang, Yan Xia 0003, Ming Cheng 0002, Cheng Wang 0003
IJCAI3
2024 OPOCA: One Point One Class Annotation for LiDAR Point Cloud Semantic Segmentation
abstract
This paper tackles the problem of requiring a large amount of data annotation in LiDAR point cloud semantic segmentation (PCSS) task by proposing OPOCA, a weakly supervised network that only annotates one point per class in a single LiDAR scan. To compensate for the supervisory losses due to extremely few annotated labels, a large number of pseudo labels is first generated using a Pseudo Label Spreading Module (PLSM), whereas the potential ambiguity and inaccuracy is further addressed by a carefully-designed Spread Distance Loss (SDL) and a Range Image Auxiliary Module (RIAM). Moreover, we propose an iterative Self Training Module (STM) to increase the high-quality pseudo labels for the next round of training. Extensive experiments on various benchmark datasets (SemanticKITTI, Waymo Open dataset, and SemanticPOSS) demonstrate the rationality of each module and the superior performance of the proposed network over the current baseline with below 0.1 ‰ labels.
Pufan Zou, Yan Xia 0003, Chenglu Wen, Cheng Wang 0003, Guoqing Zhou 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 CASSPR: Cross Attention Single Scan Place Recognition
abstract
Place recognition based on point clouds (LiDAR) is an important component for autonomous robots or self-driving vehicles. Current SOTA performance is achieved on accumulated LiDAR submaps using either point-based or voxel-based structures. While voxel-based approaches nicely integrate spatial context across multiple scales, they do not exhibit the local precision of point-based methods. As a result, existing methods struggle with fine-grained matching of subtle geometric features in sparse single-shot Li-DAR scans. To overcome these limitations, we propose CASSPR as a method to fuse point-based and voxel-based approaches using cross attention transformers. CASSPR leverages a sparse voxel branch for extracting and aggregating information at lower resolution and a point-wise branch for obtaining fine-grained local information. CASSPR uses queries from one branch to try to match structures in the other branch, ensuring that both extract self-contained descriptors of the point cloud (rather than one branch dominating), but using both to inform the out-put global descriptor of the point cloud. Extensive experiments show that CASSPR surpasses the state-of-the-art by a large margin on several datasets (Oxford RobotCar, TUM, USyd). For instance, it achieves AR@1 of 85.6% on the TUM dataset, surpassing the strongest prior model by ~15%. Our code is publicly available.1
Yan Xia 0003, Mariia Gladkova, Rui Wang 0037, Qianyun Li, Uwe Stilla, João F. Henriques, Daniel Cremers
ICCV1
2023 SegTrans: Semantic Segmentation With Transfer Learning for MLS Point Clouds
abstract
3D point cloud semantic segmentation plays an essential role in fine-grained scene understanding from photogrammetry to autonomous driving. Although recent efforts have been made to push the 3D semantic segmentation forward, many solutions cannot generalize well to new data with different sensor configurations. For example, when transferring the segmentation model learned from terrestrial laser scanning (TLS) data to mobile laser scanning (MLS) data, the performance drops dramatically. Besides, rich-labeled data is usually required. However, labeling point cloud data is time-consuming and label-intensive in practice. In light of this, we propose SegTrans, an unsupervised domain adaption method for the point cloud semantic segmentation task, which largely improves the generalization performance from one labeled dataset (source domain) to another unlabeled dataset (target domain). Specifically, we first introduce a data selection module (DSM) to tackle the discrepancy between different datasets at the data level. Then an adversarial learning module (ALM) with an adversarial loss is iteratively implemented to align the domain-specific feature in both the source and target domains, which only consists of two fully connected layers. Experiments show the overall accuracy of the proposed method achieves 88% OA on the TUM City Campus dataset (MLS dataset) when trained on the Semantic3D dataset (TLS dataset).
Shuo Shen 0003, Yan Xia 0003, Andreas Eich, Yusheng Xu, Bisheng Yang, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.2
2023 A Lightweight and Detector-Free 3D Single Object Tracker on Point Clouds
abstract
Recent works on 3D single object tracking treat the task as a target-specific 3D detection task, where an off-the-shelf 3D detector is commonly employed for the tracking. However, it is non-trivial to perform accurate target-specific detection since the point cloud of objects in raw LiDAR scans is usually sparse and incomplete. In this paper, we address this issue by explicitly leveraging temporal motion cues and propose DMT, a Detector-free Motion-prediction-based 3D Tracking network that completely removes the usage of complicated 3D detectors and is lighter, faster, and more accurate than previous trackers. Specifically, the motion prediction module is first introduced to estimate a potential target center of the current frame in a point-cloud-free manner. Then, an explicit voting module is proposed to directly regress the 3D box from the estimated target center. Extensive experiments on KITTI and NuScenes datasets demonstrate that our DMT can still achieve better performance ($\sim $10% improvement over the NuScenes dataset) and a faster tracking speed (i.e., 72 FPS) than state-of-the-art approaches without applying any complicated 3D detectors. Our code is released athttps://github.com/jimmy-dq/DMT.
Yan Xia 0003, Qiangqiang Wu, Wei Li 0111, Antoni B. Chan, Uwe Stilla
IEEE Trans. Intell. Transp. Syst.1
2021 SOE-Net: A Self-Attention and Orientation Encoding Network for Point Cloud Based Place Recognition
abstract
We tackle the problem of place recognition from point cloud data and introduce a self-attention and orientation encoding network (SOE-Net) that fully explores the relationship between points and incorporates long-range context into point-wise local descriptors. Local information of each point from eight orientations is captured in a PointOE module, whereas long-range feature dependencies among local descriptors are captured with a self-attention unit. Moreover, we propose a novel loss function called Hard Positive Hard Negative quadruplet loss (HPHN quadruplet), that achieves better performance than the commonly used metric learning loss. Experiments on various benchmark datasets demonstrate superior performance of the proposed network over the current state-of-the-art approaches. Our code is released publicly at https://github.com/Yan-Xia/SOE-Net.
Yan Xia 0003, Yusheng Xu, Shuang Li 0008, Rui Wang 0037, Juan Du 0012, Daniel Cremers, Uwe Stilla
CVPR1
2021 ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion
abstract
We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net. Specifically, the Siamese auto-encoder neural network is adopted to map the partial and complete input point cloud into a shared latent space, which can capture detailed shape prior. Then we design an iterative refinement unit to generate complete shapes with fine-grained details by integrating prior information. Experiments are conducted on the PCN dataset and the Completion3D benchmark, demonstrating the state-of-the-art performance of the proposed ASFM-Net. Our method achieves the 1st place in the leaderboard of Completion3D and outperforms existing methods with a large margin, about 12%. The codes and trained models are released publicly at https://github.com/Yan-Xia/ASFM-Net.
Yaqi Xia, Yan Xia 0003, Wei Li 0111, Rui Song 0003, Kailang Cao, Uwe Stilla
ACM Multimedia2
2020 Toward Efficient 3-D Colored Mapping in GPS-/GNSS-Denied Environments
abstract
Efficient 3-D mapping provides useful and detailed 3-D data for many applications. In this letter, we present a multisensor calibration and mapping method, to provide highly efficient and relatively accurate colored mapping for GPS-/global navigation satellite system-denied environments. The sensor data include 3-D laser scanning point clouds and camera images. A simultaneous localization and mapping (SLAM)-assisted calibration method is first proposed for multiple multibeam light detection and ranging (LiDAR) and multiple camera calibration. An improved SLAM method with loop closure is proposed for 3-D mapping. With the proposed calibration and mapping methods, centimeter-level colored point clouds can be obtained efficiently. The proposed method was tested with both backpacked and car-mounted systems on indoor and outdoor scenes. Experimental results show the effectiveness and efficiency of the proposed calibration and mapping methods.
Chenglu Wen, Yudi Dai, Yan Xia 0003, Yuhan Lian, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2017 Tree Classification in Complex Forest Point Clouds Based on Deep Learning
abstract
Recently, the classification of tree species using 3-D point clouds has drawn wide attention in surveys and forestry investigations. This letter proposes a new voxel-based deep learning method to classify tree species in 3-D point clouds collected from complex forest scenes. The proposed method includes three steps: 1) individual tree extraction based on the density of the point clouds; 2) low-level feature representation through voxel-based rasterization; and 3) classification of tree species by a deep learning model. Two data sets of 3-D forest point clouds acquired by terrestrial laser scanning systems are used to evaluate the proposed method. The method achieves an average classification accuracy of 93.1% and 95.6% on the two data sets. Furthermore, in comparative experiments, the proposed method exhibits performance superior to that of the other 3-D tree species classification methods.
Xinhuai Zou, Ming Cheng 0002, Cheng Wang 0003, Yan Xia 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4