VLDB 2026 Research / reviewers in the wild / expert
Chenglu Wen
dblp:140/4398
· DBLP profile ↗
108ranked-venue papers
7as first author
75since 2021 · last 2026
0000-0002-6189-1236ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 49 since 2021Applied, interdisciplinary, general and emerging computing · 45 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 39 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors ReasoningabstractUnsupervised 3D object detection leverages heuristic algorithms to discover potential objects, offering a promising route to reduce annotation costs in autonomous driving. Existing approaches mainly generate pseudo labels and refine them through self-training iterations. However, these pseudo-labels are often incorrect at the beginning of training, resulting in misleading the optimization process. Moreover, effectively filtering and refining them remains a critical challenge. In this paper, we propose $\textbf{OWL}$ for unsupervised 3D object detection by occupancy guided warm-up and large-model priors reasoning. OWL first employs an Occupancy Guided Warm-up (OGW) strategy to initialize the backbone weight with spatial perception capabilities, mitigating the interference of incorrect pseudo-labels on network convergence. Furthermore, OWL introduces an Instance-Cued Reasoning (ICR) module that leverages the prior knowledge of large models to assess pseudo-label quality, enabling precise filtering and refinement. Finally, we design a WAS (Weight-adapted Self-training) strategy to dynamically re-weight pseudo-labels, improving the performance through self-training. Extensive experiments on Waymo Open Dataset (WOD) and KITTI demonstrate that OWL outperforms state-of-the-art unsupervised methods by over 15.0\% mAP, revealing the effectiveness of our method. Xusheng Guo, Wanfa Zhang, Shijia Zhao, Qiming Xia, Xiaolong Xie, Chenglu Wen |
AAAI | 8 |
| 2026 | V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR LocalizationabstractMulti-agents rely on accurate poses to share and align observations, enabling a collaborative perception of the environment. However, traditional GNSS-based localization often fails in GNSS-denied environments, making consistent feature alignment difficult in collaboration. To tackle this challenge, we propose a robust GNSS-free collaborative perception framework based on LiDAR localization. Specifically, we propose a lightweight Pose Generator with Confidence (PGC) to estimate compact pose and confidence representations. To alleviate the effects of localization errors, we further develop the Pose-Aware Spatio-Temporal Alignment Transformer (PASTAT), which performs confidence-aware spatial alignment while capturing essential temporal context. Additionally, we present a new simulation dataset, V2VLoc, which can be adapted for both LiDAR localization and collaborative detection tasks. V2VLoc comprises three subsets: Town1Loc, Town4Loc, and V2VDet. Town1Loc and Town4Loc offer multi-traversal sequences for training in localization tasks, whereas V2VDet is specifically intended for the collaborative detection task. Extensive experiments conducted on the V2VLoc dataset demonstrate that our approach achieves state-of-the-art performance under GNSS-denied conditions. We further conduct extended experiments on the real-world V2V4Real dataset to validate the effectiveness and generalizability of PASTAT. Wenkai Lin, Qiming Xia, Wen Li 0005, Xun Huang 0003, Chenglu Wen |
AAAI | 5 |
| 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splattingabstract3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary queries for renderings on arbitrary viewpoints. The main challenge of distilling 2D features for 3D fields lies in the multiview inconsistency of extracted 2D features, which provides unstable supervision for the 3D feature field. GAGS addresses this challenge with two novel strategies. First, GAGS associates the prompt point density of SAM with the camera distances to scene objects, which significantly improves the multiview consistency of segmentation results. Second, GAGS further decodes a granularity factor to guide the distillation process and this granularity factor can be learned in a unsupervised manner to only select the multiview consistent 2D features in the distillation process. Experimental results on two datasets show that GAGS improves visual grounding accuracy by an average of 10.9% and semantic segmentation accuracy by an average of 7.0%, with an inference speed 2× faster than baseline methods. Yuning Peng, Haiping Wang 0004, Yuan Liu 0025, Chenglu Wen, Zhen Dong 0005, Bisheng Yang |
AAAI | 4 |
| 2026 | WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object Detectionabstract3D object detection is critical for autonomous driving, yet it remains fundamentally challenging to simultaneously maximize computational efficiency and capture long-range spatial dependencies.We observed that Mamba-based models, with their linear state-space design, capture long-range dependencies at lower cost, offering a promising balance between efficiency and accuracy.However, existing methods rely on axis-aligned scanning within a fixed window, inevitably discarding spatial information. To address this problem, we propose WinMamba, a novel Mamba-based 3D feature-encoding backbone composed of stacked WinMamba blocks. To enhance the backbone with robust multi-scale representation, the WinMamba block incorporates a window-scale-adaptive module that compensates voxel features across varying resolutions during sampling. Meanwhile, to obtain rich contextual cues within the linear state space, we equip the WinMamba layer with a learnable positional encoding and a window-shift strategy.Extensive experiments on the KITTI and Waymo datasets demonstrate that WinMamba significantly outperforms the baseline. Ablation studies further validate the individual contributions of the WSF and AWF modules in improving detection accuracy. The code will be made publicly available. Longhui Zheng, Qiming Xia, Xiaolu Chen, Zhaoliang Liu, Chenglu Wen |
AAAI | 5 |
| 2026 | DOtA++: Unsupervisely and Collaboratively Detect Objects From Multi-Agent Observations With Multi-Modal Prior ConstraintsabstractEnhancing perception performance via multi-agent collaboration has gained increasing attention in the field of autonomous driving. However, as the number of agents grows, the manual annotation required for training collaborative detectors increases significantly. To tackle this problem, we introduce an unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, named DOtA, without using labels from external. DOtA first generates preliminary labels by an initial detector, which is trained by internally shared information of collaborative agents. DOtA then optimizes these preliminary labels by utilizing the physical rule constraints derived from the surrounding area of the object. Building on DOtA, we further propose DOtA++, an enhanced version that improves performance by leveraging composite prior constraints. Beyond physical rule constraints, DOtA++ further uses image data as an auxiliary modality to introduce multi-agent observation consistency constraints, boosting object classification, while also incorporating point cloud geometric distribution constraints to improve structural description. Extensive experiments on widely-used benchmarks demonstrate that DOtA and DOtA++ effectively perceive potential objects in the scene without manual annotations. In particular, DOtA++ shows 10.7% mAP improvement over traditional unsupervised methods on V2X-R dataset. Qiming Xia, Longhui Zheng, Shijia Zhao, Xun Huang 0003, Chenglu Wen, Cheng Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | SSRFlow: Semantic-Aware Fusion with Spatial Temporal Re-Embedding for Real-World Scene FlowabstractScene flow, which provides the 3D motion field of the first frame from two consecutive point clouds, is vital for dynamic scene perception. However, contemporary scene flow methods face three major challenges. Firstly, they only consider the context of individual point clouds before flow embedding, leading to embedded points struggling to perceive the consistent semantic relationship of another frame. To address this issue, we propose a novel approach called Dual Cross Attentive (DCA) for the latent fusion and alignment between two frames based on semantic contexts. This is then integrated into Global Fusion Flow Embedding (GF) to initialize flow embedding based on global correlations in both contextual and Euclidean spaces. Secondly, deformations exist in non-rigid objects after the warping layer, which distorts the spatiotemporal relation between the consecutive frames. For a more precise estimation of residual flow at next-level, the Spatial Temporal Re-embedding (STR) module is devised to update the point sequence features at current-level. Lastly, poor generalization is often observed due to the significant domain gap between synthetic and LiDAR-scanned datasets. We leverage novel domain adaptive losses to effectively bridge the gap of motion inference from synthetic to real-world. Experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance across various datasets, with particularly outstanding results in real-world LiDAR-scanned situations. Zhiyang Lu, Qinghan Chen, Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
3DV | 4 |
| 2025 | L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object DetectionabstractLiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, it faces challenges of significant differences in terms of data quality and the degree of degradation in adverse weather. To address these issues, we introduce L4DR, a weather-robust 3D object detection method that effectively achieves LiDAR and 4D Radar fusion. Our L4DR proposes Multi-Modal Encoding (MME) and Foreground-Aware Denoising (FAD) modules to reconcile sensor gaps, which is the first exploration of the complementarity of early fusion between LiDAR and 4D radar. Additionally, we design an Inter-Modal and IntraModal ({IM}2) parallel feature extraction backbone coupled with a Multi-Scale Gated Fusion (MSGF) module to counteract the varying degrees of sensor degradation under adverse weather conditions. Experimental evaluation on a VoD dataset with simulated fog proves that L4DR is more adaptable to changing weather conditions. It delivers a significant performance increase under different fog levels, improving the 3D mAP by up to 20.0% over the traditional LiDAR-only approach. Moreover, the results on the K-Radar dataset validate the consistent performance improvement of L4DR in realworld adverse weather conditions. Xun Huang 0003, Ziyu Xu 0002, Qiming Xia, Yan Xia 0003, Jonathan Li 0001, Kyle Gao, Chenglu Wen, Cheng Wang 0003 |
AAAI | 9 |
| 2025 | Seg2Box: 3D Object Detection by Point-Wise Semantics SupervisionabstractLIDAR-based 3D object detection and semantic segmentation are critical tasks in 3D scene understanding. Traditional detection and segmentation methods supervise their models through bounding box labels and semantic mask labels. However, these two independent labels inherently contain significant redundancy. This paper aims to eliminate the redundancy by supervising 3D object detection using only semantic labels. However, the challenge arises due to the incomplete geometry structure and boundary ambiguity of point cloud instances, leading to inaccurate pseudo-labels and poor detection results. To address these challenges, we propose a novel method, named Seg2Box. We first introduce a Multi-Frame Multi-Scale Clustering (MFMS-C) module, which leverages the spatio-temporal consistency of point clouds to generate accurate box-level pseudo-labels. Additionally, the Semantic-Guiding Iterative-Mining Self-Training (SGIM-ST) module is proposed to enhance the performance by progressively refining the pseudo-labels and mining the instances without generating pseudo-labels. Experiments on the Waymo Open Dataset and nuScenes Dataset show that our method significantly outperforms other competitive methods by 23.7% and 10.3% in mAP, respectively. The results demonstrate the great label-efficient potential and advancement of our method. Maoji Zheng, Ziyu Xu 0002, Qiming Xia, Chenglu Wen, Cheng Wang 0003 |
AAAI | 5 |
| 2025 | AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label CorrectionabstractRecently, Visual Foundation Models (VFMs) have shown a remarkable generalization performance in 3D perception tasks. However, their effectiveness in large-scale outdoor datasets remains constrained by the scarcity of accurate supervision signals, the extensive noise caused by variable outdoor conditions, and the abundance of unknown objects. In this work, we propose a novel label-free learning method, Adaptive Label Correction (AdaCo), for 3D semantic segmentation. AdaCo first introduces the Cross-modal Label Generation Module (CLGM), providing cross-modal supervision with the formidable interpretive capabilities of the VFMs. Subsequently, AdaCo incorporates the Adaptive Noise Corrector (ANC), updating and adjusting the noisy samples within this supervision iteratively during training. Moreover, we develop an Adaptive Robust Loss (ARL) function to modulate each sample's sensitivity to noisy supervision, preventing potential underfitting issues associated with robust loss. Our proposed AdaCo can effectively mitigate the performance limitations of label-free learning networks in 3D semantic segmentation tasks. Extensive experiments on two outdoor benchmark datasets highlight the superior performance of our method. Pufan Zou, Shijia Zhao, Qiming Xia, Chenglu Wen, Cheng Wang 0003 |
AAAI | 5 |
| 2025 | LightLoc: Learning Outdoor LiDAR Localization at Light SpeedabstractScene coordinate regression achieves impressive results in outdoor LiDAR localization but requires days of training. Since training needs to be repeated for each new scene, long training times make these impractical for applications requiring time-sensitive system upgrades, such as autonomous driving, drones, robotics, etc. We identify large coverage areas and vast amounts of data in large-scale outdoor scenes as key challenges that limit fast training. In this paper, we propose LightLoc, the first method capable of efficiently learning localization in a new scene at light speed. Beyond freezing the scene-agnostic feature backbone and training only the scene-specific prediction heads, we introduce two novel techniques to address these challenges. First, we introduce sample classification guidance to assist regression learning, reducing ambiguity from similar samples and improving training efficiency. Second, we propose redundant sample downsampling to remove well-learned frames during training, reducing training time without compromising accuracy. In addition, the fast training and confidence estimation characteristics of sample classification enable its integration into SLAM, effectively eliminating error accumulation. Extensive experiments on large-scale outdoor datasets demonstrate that LightLoc achieves state-of-the-art performance with just 1 hour of training—50× faster than existing methods. Our Code is available at https://github.com/liw95/LightLoc. Wen Li 0005, Shangshu Yu, Dunqiang Liu, Chenglu Wen, Cheng Wang 0003 |
CVPR | 7 |
| 2025 | Boosting Adversarial Transferability through Augmentation in Hypothesis SpaceabstractAdversarial examples can mislead deep neural networks with subtle perturbations, causing them to make incorrect predictions. Notably, adversarial examples crafted for one model can also deceive other models, a phenomenon known as the transferability of adversarial examples. To improve transferability, existing studies have designed increasingly complex mechanisms, but the improvements achieved remain relatively limited and are often difficult to adapt to other modalities, further restricting the scalability of these methods. In this work, we observe a mirroring relationship between model generalization and adversarial example transferability. Motivated by this observation, we propose an augmentation-based attack, called OPS (OperatorPerturbation-based Stochastic optimization), which constructs a stochastic optimization problem by input transformation operators and random perturbations, and solves this problem to generate adversarial examples with better transferability. Extensive experiments on both images and 3D point clouds demonstrate that OPS significantly outperforms existing state-of-the-art methods in terms of both performance and cost, showcasing the universality and superiority of our approach. The code is available at https://github.com/the-full/OPS. Weiquan Liu, Qingshan Xu 0001, Shijun Zheng, Shujun Huang, Chenglu Wen, Cheng Wang 0003 |
CVPR | 8 |
| 2025 | DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based RefinementabstractLiDAR odometry is a critical module in autonomous driving systems, responsible for accurate localization by estimating the relative pose transformation between consecutive point cloud frames. However, existing studies frequently encounter challenges with unreliable pose estimation, due to the lack of in-depth understanding of scenario and the presence of noise interference. To address this challenge, we propose DiffLO, a semantic-aware LiDAR odometry network with diffusion-based refinement. To mitigate the impact of challenging cases such as dynamic, repetitive patterns, and low textures, we introduce a semantic distillation method that integrates semantic information into the odometry task. This allows the network to gain a semantic understanding of the scene, enabling it to focus more on the objects that are beneficial for pose estimation. Additionally, to enhance the robustness, we propose a diffusion-based refinement method. This method uses pose-related features as conditional constraints for generative diversity, iteratively refining the pose estimation to achieve greater accuracy. Comparative experiments on the KITTI odometry dataset demonstrate that the proposed method achieves state-of-the-art performance among existing learning-based approaches. Furthermore, the proposed DiffLO method outperforms the classic A-LOAM on most evaluation sequences. The code will be released at https://github.com/hytree7/difflo. Yongshu Huang, Minghang Zhu, Sheng Ao, Chenglu Wen, Cheng Wang 0003 |
CVPR | 5 |
| 2025 | V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object DetectionabstractCurrent Vehicle-to-Everything (V2X) systems have significantly enhanced 3D object detection using LiDAR and camera data. However, they face performance degradation in adverse weather. Weather-robust 4D radar, with Doppler velocity and additional geometric information, offers a promising solution to this challenge. To this end, we present V2X-R, the first simulated V2X dataset incorporating LiDAR, camera, and 4D radar modalities. V2X-R contains 12,079 scenarios with 37,727 frames of LiDAR and 4D radar point clouds, 150,908 images, and 170,859 annotated 3D vehicle bounding boxes. Subsequently, we propose a novel cooperative LiDAR-4D radar fusion pipeline for 3D object detection and implement it with multiple fusion strategies. To achieve weather-robust detection, we additionally propose a Multi-modal Denoising Diffusion (MDD) module in our fusion pipeline. MDD utilizes weather-robust 4D radar feature as a condition to guide the diffusion model in denoising noisy LiDAR features. Experiments show that our LiDAR-4D radar fusion pipeline demonstrates superior performance in the V2X-R dataset. Over and above this, our MDD module further improved the foggy/snowy performance of the basic fusion model by up to 5.73%/6.70% and barely disrupting normal performance. The dataset and code will be publicly available at: https://github.com/ylwhxht/V2X-R. Xun Huang 0003, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li 0003, Cheng Wang 0003, Chenglu Wen |
CVPR | 8 |
| 2025 | Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual LabelsabstractUnsupervised 3D object detection serves as an important solution for offline 3D object annotation. However, due to the data sparsity and limited views, the clustering-based label fitting in unsupervised object detection often generates low-quality pseudo-labels. Multi-agent collaborative dataset, which involves the sharing of complementary observations among agents, holds the potential to break through this bottleneck. In this paper, we introduce a novel unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, termed DOtA, without using labels from external. DOtA first uses the internally shared ego-pose and ego-shape of collaborative agents to initialize the detector, leveraging the generalization performance of neural networks to infer preliminary labels. Subsequently, DOtA uses the complementary observations between agents to perform multi-scale encoding on preliminary labels, then decodes high-quality and low-quality labels. These labels are further used as prompts to guide a correct feature learning process, thereby enhancing the performance of the unsupervised object detection task. Extensive experiments on the V2V4Real and OPV2V datasets show that our DOtA outperforms state-of-the-art unsupervised 3D object detection methods. Additionally, we also validate the effectiveness of the DOtA labels under various collaborative perception frameworks. The code is available at https://github.com/xmuqimingxia/DOtA. Qiming Xia, Wenkai Lin, Haoen Xiang, Xun Huang 0003, Siheng Chen, Zhen Dong 0005, Cheng Wang 0003, Chenglu Wen |
CVPR | 8 |
| 2025 | ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World CoordinateabstractHuman Motion Recovery (HMR) research mainly focuses on ground-based motions such as running. The study on capturing climbing motion, an off-ground motion, is sparse. This is partly due to the limited availability of climbing motion datasets, especially large-scale and challenging 3D labeled datasets. To address the insufficiency of climbing motion datasets, we collect AscendMotion, a large-scale well-annotated, and challenging climbing motion dataset. It consists of 412k RGB, LiDAR frames, and IMU measurements, including the challenging climbing motions of 22 skilled climbing coaches across 12 different rock walls. Capturing the climbing motions is challenging as it requires precise recovery of not only the complex pose but also the global position of climbers. Although multiple global HMR methods have been proposed, they cannot faithfully capture climbing motions. To address the limitations of HMR methods for climbing, we propose Climbing-Cap, a motion recovery method that reconstructs continuous 3D human climbing motion in a global coordinate system. One key insight is to use the RGB and LiDAR modalities to separately reconstruct motions in camera coordinates and global coordinates and to optimize them jointly. We demonstrate the quality of the AscendMotion dataset and present promising results from ClimbingCap. The AscendMotion dataset and source code release publicly at http://www.lidarhumanmotion.net/climbingcap/ Xincheng Lin, Yuhua Luo, Shuqi Fan, Yudi Dai, Qixin Zhong, Lincai Zhong, Yuexin Ma, Lan Xu 0003, Chenglu Wen, Cheng Wang 0003 |
CVPR | 10 |
| 2025 | SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic PromptsabstractRecently, sparsely-supervised 3D object detection has gained great attention, achieving performance close to fully-supervised 3D detectors while requiring only a few annotated instances. Nevertheless, these methods suffer challenges when accurate labels are extremely absent. In this paper, we propose a boosting strategy, termed SP3D, explicitly utilizing the cross-modal semantic prompts generated from Large Multimodal Models (LMMs) to boost the 3D detector with robust feature discrimination capability under sparse annotation settings. Specifically, we first develop a Confident Points Semantic Transfer (CPST) module that generates accurate cross-modal semantic prompts through boundary-constrained center cluster selection. Based on these accurate semantic prompts, which we treat as seed points, we introduce a Dynamic Cluster Pseudo-label Generation (DCPG) module to yield pseudo-supervision signals from the geometry shape of multi-scale neighbor points. Additionally, we design a Distribution Shape score (DS score) that chooses high-quality supervision signals for the initial training of the 3D detector. Experiments on the KITTI dataset and Waymo Open Dataset (WOD) have validated that SP3D can enhance the performance of sparsely supervised detectors by a large margin under meager labeling conditions. Moreover, we verified SP3D in the zero-shot setting, where its performance exceeded that of the state-of-the-art methods. The code is available at https://github.com/xmuqimingxia/SP3D. Shijia Zhao, Qiming Xia, Xusheng Guo, Pufan Zou, Maoji Zheng, Chenglu Wen, Cheng Wang 0003 |
CVPR | 7 |
| 2025 | Pretend Benign: A Stealthy Adversarial Attack by Exploiting Vulnerabilities in Cooperative Perception
Dongyu Pan, Qiming Xia, Cheng Wang 0003, Chenglu Wen |
ICCV | 7 |
| 2025 | Motal: Unsupervised 3D Object Detection by Modality and Task-Specific Knowledge Transfer
Xusheng Guo, Xin Li 0003, Cheng Wang 0003, Chenglu Wen |
ICCV | 7 |
| 2025 | DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement LearningabstractDiffusion models have been widely adopted in image and language generation and are now being applied to reinforcement learning. However, the application of diffusion models in offline cooperative Multi-Agent Reinforcement Learning (MARL) remains limited. Although existing studies explore this direction, they suffer from scalability or poor cooperation issues due to the lack of design principles for diffusion-based MARL. The Individual-Global-Max (IGM) principle is a popular design principle for cooperative MARL. By satisfying this principle, MARL algorithms achieve remarkable performance with good scalability. In this work, we extend the IGM principle to the Individual-Global-identically-Distributed (IGD) principle. This principle stipulates that the generated outcome of a multi-agent diffusion model should be identically distributed as the collective outcomes from multiple individual-agent diffusion models. We propose DoF, a diffusion factorization framework for Offline MARL. It uses noise factorization function to factorize a centralized diffusion model into multiple diffusion models. We theoretically show that the noise factorization functions satisfy the IGD principle. Furthermore, DoF uses data factorization function to model the complex relationship among data generated by multiple diffusion models. Through extensive experiments, we demonstrate the effectiveness of DoF. The source code is available at [https://github.com/xmu-rl-3dv/DoF](https://github.com/xmu-rl-3dv/DoF). Ziwei Deng, Chenxing Lin, Yongquan Fu, Weiquan Liu, Chenglu Wen, Cheng Wang 0003 |
ICLR | 7 |
| 2025 | Unsupervised 3D Object Detection by Commonsense ClueabstractTraditional 3D object detectors, whether fully-, semi-, or weakly-supervised, rely heavily on extensive human annotations. In contrast, this paper introduces an unsupervised 3D object detector that automatically discerns object patterns without such annotations. To achieve this, we propose a Commonsense Prototype-based Detector (CPD) for unsupervised 3D object detection. CPD first constructs Commonsense Prototypes (CProto) to represent the geometric center and size of objects. It then generates high-quality pseudo-labels and guides detector convergence using size and geometry priors from CProto. Building on CPD, we further introduce CPD++, an enhanced version that improves performance by leveraging motion cues. CPD++ learns localization from stationary objects and recognition from moving objects, facilitating the mutual transfer of localization and recognition knowledge between these two object types. Both CPD and CPD++ outperform existing state-of-the-art unsupervised 3D detectors. Furthermore, when trained on Waymo Open Dataset (WOD) and tested on KITTI, CPD++ achieves 89.25% 3D Average Precision (AP) on the moderate car class at a 0.5 IoU threshold, reaching 95.3% of the performance attained by fully supervised counterparts. These results underscore the significant advancements brought by our method. Shijia Zhao, Xun Huang 0003, Qiming Xia, Chenglu Wen, Li Jiang 0009, Xin Li 0003, Cheng Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | QPCR: Weakly Supervised Semantic Segmentation of Large-Scale Point Cloud via Gradual Query Points Component ReasoningabstractSemantic segmentation of large-scale point clouds under weak supervision is challenging due to the limited annotations. Current methods typically rely on the implicit use of annotation information to supervise the network, thereby constraining the capacity to characterize features at annotation points. In this article, we propose a query points component reasoning (QPCR) framework, which enhances the utilization of annotation information by introducing them into the middle layer, thereby refining the features used for semantic segmentation. First, we propose the semantic annotations component code reasoning (SACC Reasoning) module, which uses annotations of query points to supervise the semantic information of gradual query points. The semantic annotations component code (SACC) Reasoning module can guide the learning of feature representations of semantic annotations at multiple scales. Second, we create the semantic primitive component code reasoning (SPCC Reasoning) module, which addresses the issue of supervising the network using existing annotation point information that may suppress active features. The semantic primitive component code (SPCC) Reasoning module infers semantic primitive labels of query points using the decoding layer and uses them to guide the corresponding coding layer in learning semantic primitives. Finally, our QPCR method achieves state-of-the-art (SOTA) results on public benchmark datasets. Specifically, even with only 1% of annotated points, our QPCR method achieves the highest mIoU accuracy, i.e., 65.4% on S3DIS Area 5, 56.4% on SensatUrban and 55.3% mIoU on SemanticKITTI, respectively. Lixin Zhan, Jie Jiang 0017, Tianjian Zhou, Chenglu Wen, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | MOJO: MOtion Pattern Learning and JOint-Based Fine-Grained Mining for Person Re-Identification Based on 4D LiDAR Point Clouds
Zhiyang Lu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Fade3D: Fast and Deployable 3D Object Detection for Autonomous Drivingabstract3D object detection is an essential scene perception capability for autonomous vehicles. In intelligent transportation systems, autonomous vehicles require minimal inference latency to sense their surroundings in real-time. However, advanced 3D detection methods often suffer from high inference latency. This limits the real-time deployment of 3D detection models in the real world. To address this problem, this paper proposes a fast and deployable 3D object detection method from the LiDAR point cloud for autonomous driving, namedFade3D. Firstly, we propose a Lightweight Input Encoder (LIE) to extract the most critical features from point clouds. Then, we develop a Spatial Feature Enhancement BEV backbone (SFENet) that efficiently encodes geometry features into compact representations. Additionally, we design an IoU-aware Loss Re-weighting (ILR) that enhances performance by shifting more attention to hard samples. Leveraging LIE and SFENet, our approach is independent of point cloud density and number, achieving significant speed advantages in processing large-scale point clouds and being deployment-friendly. Extensive experiments on KITTI and Waymo Open Dataset (WOD) datasets comparing various baseline detectors demonstrate its universality and superiority. Specifically, our method demonstrates impressive real-time inference capabilities, achieving 51.5 Hz on an RTX3090 GPU and 12.4 Hz on a Jetson Orin embedded development board. Code will be available at https://github.com/wayyeah/Fade3D Qiming Xia, Zhen Dong 0005, Ruofei Zhong, Cheng Wang 0003, Chenglu Wen |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Sunshine to Rainstorm: Cross-Weather Knowledge Distillation for Robust 3D Object DetectionabstractLiDAR-based 3D object detection models inevitably struggle under rainy conditions due to the degraded and noisy scanning signals. Previous research has attempted to address this by simulating the noise from rain to improve the robustness of detection models. However, significant disparities exist between simulated and actual rain-impacted data points. In this work, we propose a novel rain simulation method, termed DRET, that unifies Dynamics and Rainy Environment Theory to provide a cost-effective means of expanding the available realistic rain data for 3D detection training. Furthermore, we present a Sunny-to-Rainy Knowledge Distillation (SRKD) approach to enhance 3D detection under rainy conditions. Extensive experiments on the Waymo-Open-Dataset show that, when combined with the state-of-the-art DSVT model and other classical 3D detectors, our proposed framework demonstrates significant detection accuracy improvements, without losing efficiency. Remarkably, our framework also improves detection capabilities under sunny conditions, therefore offering a robust solution for 3D detection regardless of whether the weather is rainy or sunny. Xun Huang 0003, Xin Li 0003, Xiaoliang Fan, Chenglu Wen, Cheng Wang 0003 |
AAAI | 5 |
| 2024 | SPEAL: Skeletal Prior Embedded Attention Learning for Cross-Source Point Cloud RegistrationabstractPoint cloud registration, a fundamental task in 3D computer vision, has remained largely unexplored in cross-source point clouds and unstructured scenes. The primary challenges arise from noise, outliers, and variations in scale and density. However, neglected geometric natures of point clouds restricts the performance of current methods. In this paper, we propose a novel method termed SPEAL to leverage skeletal representations for effective learning of intrinsic topologies of point clouds, facilitating robust capture of geometric intricacy. Specifically, we design the Skeleton Extraction Module to extract skeleton points and skeletal features in an unsupervised manner, which is inherently robust to noise and density variances. Then, we propose the Skeleton-Aware GeoTransformer to encode high-level skeleton-aware features. It explicitly captures the topological natures and inter-point-cloud skeletal correlations with the noise-robust and density-invariant skeletal representations. Next, we introduce the Correspondence Dual-Sampler to facilitate correspondences by augmenting the correspondence set with skeletal correspondences. Furthermore, we construct a challenging novel cross-source point cloud dataset named KITTI CrossSource for benchmarking cross-source point cloud registration methods. Extensive quantitative and qualitative experiments are conducted to demonstrate our approach’s superiority and robustness on both cross-source and same-source datasets. To the best of our knowledge, our approach is the first to facilitate point cloud registration with skeletal geometric priors. Kezheng Xiong, Maoji Zheng, Qingshan Xu 0001, Chenglu Wen, Cheng Wang 0003 |
AAAI | 4 |
| 2024 | DiffLoc: Diffusion Model for Outdoor LiDAR LocalizationabstractAbsolute pose regression (APR) estimates global pose in an end-to-end manner, achieving impressive results in learn-based LiDAR localization. However, compared to the top-performing methods reliant on 3D-3D correspondence matching, APR's accuracy still has room for improvement. We recognize APR's lack of robust features learning and iterative denoising process leads to suboptimal results. In this paper, we propose DiffLoc, a novel framework that formulates LiDAR localization as a conditional generation of poses. First, we propose to utilize the foundation model and static-object-aware pool to learn robust features. Second, we incorporate the iterative denoising process into APR via a diffusion model conditioned on the learned geometrically robust features. In addition, due to the unique nature of diffusion models, we propose to adapt our models to two additional applications: (1) using multiple inferences to evaluate pose uncertainty, and (2) seamlessly introducing geometric constraints on denoising steps to improve prediction accuracy. Extensive experiments conducted on the Oxford Radar RobotCar and NCLT datasets demonstrate that DiffLoc outperforms better than the state-of-the-art methods. Especially on the NCLT dataset, we achieve 35% and 34.7% improvement on position and orientation accuracy, respectively. Our code is released at https://github.com/liw95/DiffLoc. Wen Li 0005, Yuyang Yang, Shangshu Yu, Guosheng Hu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
CVPR | 5 |
| 2024 | Commonsense Prototype for Outdoor Unsupervised 3D Object DetectionabstractThe prevalent approaches of unsupervised 3D object de-tection follow cluster-based pseudo-label generation and iterative self-training processes. However, the challenge arises due to the sparsity of LiDAR scans, which leads to pseudo-labels with erroneous size and position, resulting in subpar detection performance. To tackle this problem, this paper introduces a Commonsense Prototype-based Detector, termed CPD, for unsupervised 3D object de-tection. CPD first constructs Commonsense Prototype (CProto) characterized by high-quality bounding box and dense points, based on commonsense intuition. Subse-quently, CPD refines the low-quality pseudo-labels by lever-aging the size prior from CProto. Furthermore, CPD en-hances the detection accuracy of sparsely scanned objects by the geometric knowledge from CProto. CPD outper-forms state-of-the-art unsupervised 3D detectors on Waymo Open Dataset (WOD), PandaSet, and KITTI datasets by a large margin. Besides, by training CPD on WOD and testing on KITTI, CPD attains 90.85% and 81.01% 3D Aver-age Precision on easy and moderate car classes, respectively. These achievements position CPD in close prox-imity to fully supervised detectors, highlighting the sig-nificance of our method. The code will be available at https://github.com/hailanyi/CPD. Shijia Zhao, Xun Huang 0003, Chenglu Wen, Xin Li 0003, Cheng Wang 0003 |
CVPR | 4 |
| 2024 | HINTED: Hard Instance Enhanced Detector with Mixed-Density Feature Fusion for Sparsely-Supervised 3D Object DetectionabstractCurrent sparsely-supervised object detection methods largely depend on high threshold settings to derive high-quality pseudo labels from detector predictions. However, hard instances within point clouds frequently display incomplete structures, causing decreased confidence scores in their assigned pseudo-labels. Previous methods inevitably result in inadequate positive supervision for these instances. To address this problem, we propose a novel Hard INsTance Enhanced Detector (HINTED), for sparsely-supervised 3D object detection. Firstly, we design a self-boosting teacher (SBT) model to generate more potential pseudo-labels, enhancing the effectiveness of information transfer. Then, we introduce a mixed-density student (MDS) model to concentrate on hard instances during the training phase, thereby improving detection accuracy. Our extensive experiments on the KITTI dataset validate our method's superior performance. Compared with leading sparsely-supervised methods, HINTED significantly improves the detection performance on hard instances, no-tably outperforming fully-supervised methods in detecting challenging categories like cyclists. HINTED also significantly outperforms the state-of-the-art semi-supervised method on challenging categories. The code is available at https://github.com/xmuqimingxia/HINTED. Qiming Xia, Shijia Zhao, Leyuan Xing, Xun Huang 0003, Jinhao Deng, Xin Li 0003, Chenglu Wen, Cheng Wang 0003 |
CVPR | 9 |
| 2024 | RELI11D: A Comprehensive Multimodal Human Motion Dataset and MethodabstractComprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB, LiDAR, or IMU data. However, solely using these modalities or a combination of them may not be adequate for HPE, particularly for complex and fast movements. For holistic human motion understanding, we present RELI11D, a high-quality multimodal human motion dataset involves LiDAR, IMU system, RGB camera, and Event camera. It records the motions of 10 actors performing 5 sports in 7 scenes, including 3.32 hours of synchronized LiDAR point clouds, IMU measurement data, RGB videos and Event steams. Through extensive experiments, we demonstrate that the RELI 11 D presents considerable challenges and opportunities as it contains many rapid and complex motions that require precise location. To address the challenge of integrating different modalities, we propose LEIR, a multimodal baseline that effectively utilizes LiDAR Point Cloud, Event stream, and RGB through our cross-attention fusion strategy. We show that LEIR exhibits promising results for rapid motions and daily motions and that utilizing the characteristics of multiple modalities can indeed improve HPE performance. Both the dataset and source code release publicly in http://www.lidarhumanmotion.net/reli11d/, fostering collaboration and enabling further exploration in this field. Shuqiang Cai, Shuqi Fan, Xincheng Lin, Yudi Dai, Chenglu Wen, Lan Xu 0003, Yuexin Ma, Cheng Wang 0003 |
CVPR | 8 |
| 2024 | LiSA: LiDAR Localization with Semantic AwarenessabstractLiDAR localization is a fundamental task in robotics and computer vision, which estimates the pose of a Li- DAR point cloud within a global map. Scene Coordinate Regression (SCR) has demonstrated state-of-the-art performance in this task. In SCR, a scene is represented as a neural network, which outputs the world coordinates for each point in the input point cloud. However, SCR treats all points equally during localization, ignoring the fact that not all objects are beneficial for localization. For exam-ple, dynamic objects and repeating structures often negatively impact SCR. To address this problem, we introduce LiSA, the first method that incorporates semantic aware-ness into SCR to boost the localization robustness and accuracy. To avoid extra computation or network parame-ters during inference, we distill the knowledge from a seg-mentation model to the original SCR network. Experi-ments show the superior performance of LiSA on standard LiDAR localization benchmarks compared to state-of-the- art methods. Applying knowledge distillation not only pre-serves high efficiency but also achieves higher localization accuracy than introducing extra semantic segmentation modules. We also analyze the benefit of semantic in-formation for LiDAR localization. Our code is released at https://github.com/Ybchun/LiSA. Bochun Yang, Zijun Li 0006, Wen Li 0005, Zhipeng Cai 0003, Chenglu Wen, Matthias Müller 0011, Cheng Wang 0003 |
CVPR | 5 |
| 2024 | CMD: A Cross Mechanism Domain Adaptation Dataset for 3D Object Detection
Jinhao Deng, Xun Huang 0003, Qiming Xia, Xin Li 0003, Wei Li 0111, Chenglu Wen, Cheng Wang 0003 |
ECCV (57) | 9 |
| 2024 | FedPFT: Federated Proxy Fine-Tuning of Foundation Models
Zhaopeng Peng, Xiaoliang Fan, Zheng Wang 0076, Shirui Pan, Chenglu Wen, Ruisheng Zhang, Cheng Wang 0003 |
IJCAI | 6 |
| 2024 | FedSAC: Dynamic Submodel Allocation for Collaborative Fairness in Federated LearningabstractCollaborative fairness stands as an essential element in federated learning to encourage client participation by equitably distributing rewards based on individual contributions. Existing methods primarily focus on adjusting gradient allocations among clients to achieve collaborative fairness. However, they frequently overlook crucial factors such as maintaining consistency across local models and catering to the diverse requirements of high-contributing clients. This oversight inevitably decreases both fairness and model accuracy in practice. To address these issues, we propose FedSAC, a novel Federated learning framework with dynamic Submodel Allocation for Collaborative fairness, backed by a theoretical convergence guarantee. First, we present the concept of "bounded collaborative fairness (BCF)", which ensures fairness by tailoring rewards to individual clients based on their contributions. Second, to implement the BCF, we design a submodel allocation module with a theoretical guarantee of fairness. This module incentivizes high-contributing clients with high-performance submodels containing a diverse range of crucial neurons, thereby preserving consistency across local models. Third, we further develop a dynamic aggregation module to adaptively aggregate submodels, ensuring the equitable treatment of low-frequency neurons and consequently enhancing overall model accuracy. Extensive experiments conducted on three public benchmarks demonstrate that FedSAC outperforms all baseline methods in both fairness and model accuracy. We see this work as a significant step towards incentivizing broader client participation in federated learning. The source code is available at https://github.com/wangzihuixmu/FedSAC. Zheng Wang 0076, Lingjuan Lyu, Zhaopeng Peng, Chenglu Wen, Rongshan Yu, Cheng Wang 0003, Xiaoliang Fan |
KDD | 6 |
| 2024 | HmPEAR: A Dataset for Human Pose Estimation and Action RecognitionabstractWe introduce HmPEAR, a novel dataset crafted for advancing research in 3D Human Pose Estimation (3D HPE) and Human Action Recognition (HAR), with a primary focus on outdoor environments. This dataset offers a synchronized collection of imagery, LiDAR point clouds, 3D human poses, and action categories. In total, the dataset encompasses over 300,000 frames collected from 10 distinct scenes and 25 diverse subjects. Among these, 250,000 frames of data contain 3D human pose annotations captured using an advanced motion capture system and further optimized for accuracy. Furthermore, the dataset annotates 40 types of daily human actions, resulting in over 6,000 action clips. Through extensive experimentation, we have demonstrated the quality of HmPEAR and highlighted the challenges it presents to current methodologies. Additionally, we propose baselines leveraging sequential images and point clouds for 3D HPE and HAR, which underscore the mutual reinforcement between them, highlighting the potential for cross-task synergies. The dataset is available at http://www.lidarhumanmotion.net/hmpear. Yitai Lin, Zhijie Wei, Wanfa Zhang, Xiping Lin, Yudi Dai, Chenglu Wen, Lan Xu 0003, Cheng Wang 0003 |
ACM Multimedia | 6 |
| 2024 | Mining and Transferring Feature-Geometry Coherence for Unsupervised Point Cloud RegistrationabstractPoint cloud registration, a fundamental task in 3D vision, has achieved remarkable success with learning-based methods in outdoor environments. Unsupervised outdoor point cloud registration methods have recently emerged to circumvent the need for costly pose annotations. However, they fail to establish reliable optimization objectives for unsupervised training, either relying on overly strong geometric assumptions, or suffering from poor-quality pseudo-labels due to inadequate integration of low-level geometric and high-level contextual information. We have observed that in the feature space, latent new inlier correspondences tend to cluster
around respective positive anchors that summarize features of existing inliers. Motivated by this observation, we propose a novel unsupervised registration method termed INTEGER to incorporate high-level contextual information for reliable pseudo-label mining. Specifically, we propose the Feature-Geometry Coherence Mining module to dynamically adapt the teacher for each mini-batch of data during training and discover reliable pseudo-labels by considering both high-level feature representations and low-level geometric cues. Furthermore, we propose Anchor-Based Contrastive Learning to facilitate contrastive learning with anchors for a robust feature space. Lastly, we introduce a Mixed-Density Student to learn density-invariant features, addressing challenges related to density variation and low overlap in the outdoor scenario. Extensive experiments on KITTI and nuScenes datasets demonstrate that our INTEGER achieves competitive performance in terms of accuracy and generalizability. Kezheng Xiong, Haoen Xiang, Qingshan Xu 0001, Chenglu Wen, Jonathan Jun Li, Cheng Wang 0003 |
NeurIPS | 4 |
| 2024 | Federated Graph Learning for Cross-Domain RecommendationabstractCross-domain recommendation (CDR) offers a promising solution to the data sparsity problem by enabling knowledge transfer across source and target domains. However, many recent CDR models overlook crucial issues such as privacy as well as the risk of negative transfer (which negatively impact model performance), especially in multi-domain settings. To address these challenges, we propose FedGCDR, a novel federated graph learning framework that securely and effectively leverages positive knowledge from multiple source domains. First, we design a positive knowledge transfer module that ensures privacy during inter-domain knowledge transmission. This module employs differential privacy-based knowledge extraction combined with a feature mapping mechanism, transforming source domain embeddings from federated graph attention networks into reliable domain knowledge. Second, we design a knowledge activation module to filter out potential harmful or conflicting knowledge from source domains, addressing the issues of negative transfer. This module enhances target domain training by expanding the graph of the target domain to generate reliable domain attentions and fine-tunes the target model for improved negative knowledge filtering and more accurate predictions. We conduct extensive experiments on 16 popular domains of the Amazon dataset, demonstrating that FedGCDR significantly outperforms state-of-the-art methods. Zhaopeng Peng, Jianzhong Qi 0001, Chaochao Chen 0001, Weike Pan, Chenglu Wen, Cheng Wang 0003, Xiaoliang Fan |
NeurIPS | 7 |
| 2024 | HiSC4D: Human-Centered Interaction and 4D Scene Capture in Large-Scale Space Using Wearable IMUs and LiDARabstractWe introduce HiSC4D, a novel Human-centered interaction and 4D Scene Capture method, aimed at accurately and efficiently creating a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, rich human-human interactions, and human-environment interactions. By utilizing body-mounted IMUs and a head-mounted LiDAR, HiSC4D can capture egocentric human motions in unconstrained space without the need for external devices and pre-built maps. This affords great flexibility and accessibility for human-centered interaction and 4D scene capturing in various environments. Taking into account that IMUs can capture human spatially unrestricted poses but are prone to drifting for long-period using, and while LiDAR is stable for global localization but rough for local positions and orientations, HiSC4D employs a joint optimization method, harmonizing all sensors and utilizing environment cues, yielding promising results for long-term capture in large scenes. To promote research of egocentric human interaction in large scenes and facilitate downstream tasks, we also present a dataset, containing 8 sequences in 4 large scenes (200 to 5,000 [Formula: see text]), providing 36 k frames of accurate 4D human motions with SMPL annotations and dynamic scenes, 31k frames of cropped human point clouds, and scene mesh of the environment. A variety of scenarios, such as the basketball gym and commercial street, alongside challenging human motions, such as daily greeting, one-on-one basketball playing, and tour guiding, demonstrate the effectiveness and the generalization ability of HiSC4D. The dataset and code will be publicly available for research purposes. Yudi Dai, Xiping Lin, Chenglu Wen, Lan Xu 0003, Yuexin Ma, Cheng Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | LiDARCapV2: 3D human pose estimation with human-object interaction from LiDAR point clouds
Qihong Mao, Chenglu Wen, Lan Xu 0003, Cheng Wang 0003 |
Pattern Recognit. | 4 |
| 2024 | SCSQ-Net: A Shared Kernel Point Convolution Semantic Query Network for Weakly Supervised Classification of Multispectral LiDAR Point Clouds
Haiyan Guan, Yongtao Yu, Yufu Zang, Chenglu Wen |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | OPOCA: One Point One Class Annotation for LiDAR Point Cloud Semantic SegmentationabstractThis paper tackles the problem of requiring a large amount of data annotation in LiDAR point cloud semantic segmentation (PCSS) task by proposing OPOCA, a weakly supervised network that only annotates one point per class in a single LiDAR scan. To compensate for the supervisory losses due to extremely few annotated labels, a large number of pseudo labels is first generated using a Pseudo Label Spreading Module (PLSM), whereas the potential ambiguity and inaccuracy is further addressed by a carefully-designed Spread Distance Loss (SDL) and a Range Image Auxiliary Module (RIAM). Moreover, we propose an iterative Self Training Module (STM) to increase the high-quality pseudo labels for the next round of training. Extensive experiments on various benchmark datasets (SemanticKITTI, Waymo Open dataset, and SemanticPOSS) demonstrate the rationality of each module and the superior performance of the proposed network over the current baseline with below 0.1 ‰ labels. Pufan Zou, Yan Xia 0003, Chenglu Wen, Cheng Wang 0003, Guoqing Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | WHU-Railway3D: A Diverse Dataset and Benchmark for Railway Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation (PCSS) shows great potential in generating accurate 3D semantic maps for digital twin railways. Deep learning-based methods have seen substantial advancements, driven by numerous PCSS datasets. Nevertheless, existing datasets tend to neglect railway scenes, with limitations in scale, categories, and scene diversity. This motivated us to establish WHU-Railway3D, a diverse PCSS dataset specifically designed for railway scenes. WHU-Railway3D is categorized into urban, rural, and plateau railways based on scene complexity and semantic class distribution. The dataset spans approximately 30 km with 4.6 billion points labeled into 11 classes, such as rails, masts, overhead lines, and fences. In addition to 3D coordinates, WHU-Railway3D provides rich attribute information such as reflected intensity, scanning angle, and number of returns. Cutting-edge methods are extensively evaluated on the dataset, followed by in-depth analysis. Lastly, key challenges and potential future work are identified to stimulate further innovative research. The dataset is accessible athttps://github.com/WHU-USI3DV/WHU-Railway3D. Yuzhou Zhou, Bing Wang 0013, Jianping Li 0004, Zhen Dong 0005, Chenglu Wen, Zhiliang Ma, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | NIDALoc: Neurobiologically Inspired Deep LiDAR LocalizationabstractAbsolute pose regression has shown great potential in LiDAR localization, which learns to regress 6-DoF LiDAR poses through deep networks. However, recent regression methods suffer from scene ambiguities in challenging scenarios, leading to inaccurate and unstable localization. Inspired by neurobiological localization mechanisms, i.e., the firing mechanism of place cells, head-direction cells, and grid cells in mammalian brains, we propose a novel LiDAR localization framework called NIDALoc to achieve more robust and accurate results. First, we propose a Hebbian memory module, motivated by place cells, to preserve historical information, which helps refine local view features to reduce scene ambiguities. Specifically, the memory module stores scene information and then recalls it when revisiting an old place. Second, we propose a novel pose constrained framework, consisting of an orientation classification task and a grid center regression task, to regularize orientation and position estimation, respectively. The framework based on head-direction cells and grid cells constrains the absolute pose regression to reduce wrong predictions. Extensive experiments on two outdoor datasets demonstrate the effectiveness of NIDALoc, which outperforms state-of-the-art localization methods, especially in large-scale challenging scenes. The source code is available on the project website at https://github.com/PSYZ1234/NIDALoc. Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Chenglu Wen, Yunuo Yang, Bailu Si, Guosheng Hu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | SR-Adv: Salient Region Adversarial Attacks on 3D Point Clouds for Autonomous DrivingabstractAutonomous driving safety based on LiDAR perception is increasingly becoming a hot spot. Specifically, 3D adversarial examples always make the prediction results of deep neural network models unpredictable, which poses a major security risk to autonomous driving systems. However, the vulnerability of 3D neural network models to adversarial examples is less explored. At present, existing adversarial attack methods often obtain 3D adversarial examples by perturbing the entire point cloud, which requires a large number of perturbed points, i.e. requires a large perturbation budget. In this paper, we propose a Salient Region Adversarial attack method (SR-Adv) to generate adversarial point clouds, by perturbing fewer regions and fewer points. To our knowledge, we are the first to propose region-based attacks for 3D point clouds. First, the proposed SR-Adv employs game theory to extract salient regions of point clouds. This mechanism assigns a value to each region to measure its importance to the 3D neural network model prediction results and realizes the vulnerability analysis of the 3D model. Second, we propose a novel optimization-based gradient attack algorithm to achieve adversarial attacks on salient regions. We evaluate the proposed SR-Adv attack method on the synthetic datasets ModelNet40 and ShapeNetPart as well as the real-world dataset KITTI and NuScenes. Experimental results show that the proposed SR-Adv achieves a state-of-the-art attack success rate and better imperceptibility by perturbing fewer points on 3D point clouds. Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Ping Zhong 0001, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Transformation-Equivariant 3D Object Detection for Autonomous Drivingabstract3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augmentation are required for robust detection. Recent equivariant networks explicitly model the transformation variations by applying shared networks on multiple transformed point clouds, showing great potential in object geometry modeling. However, it is difficult to apply such networks to 3D object detection in autonomous driving due to its large computation cost and slow reasoning speed. In this work, we present TED, an efficient Transformation-Equivariant 3D Detector to overcome the computation cost and speed issues. TED first applies a sparse convolution backbone to extract multi-channel transformation-equivariant voxel features; and then aligns and aggregates these equivariant features into lightweight and compact representations for high-performance 3D object detection. On the highly competitive KITTI 3D car detection leaderboard, TED ranked 1st among all submissions with competitive efficiency. Code is available at https://github.com/hailanyi/TED. Chenglu Wen, Wei Li 0111, Xin Li 0003, Ruigang Yang, Cheng Wang 0003 |
AAAI | 2 |
| 2023 | SLOPER4D: A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban EnvironmentsabstractWe present SLOPER4D, a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild. Employing a head-mounted device integrated with a LiDAR and camera, we record 12 human subjects' activities over 10 diverse urban scenes from an egocentric view. Frame-wise annotations for 2D key points, 3D pose parameters, and global translations are provided, together with reconstructed scene point clouds. To obtain accurate 3D ground truth in such large dynamic scenes, we propose a joint optimization method to fit local SMPL meshes to the scene and fine-tune the camera calibration during dynamic motions frame by frame, resulting in plausible and scene-natural 3D human poses. Even-tually, SLOPER4D consists of 15 sequences of human motions, each of which has a trajectory length of more than 200 meters (up to 1,300 meters) and covers an area of more than 200 m2(up to 30,000 m2), including more than 100k LiDAR frames, 300k video frames, and 500k IMU-based motion frames. With SLOPER4D, we provide a detailed and thorough analysis of two critical tasks, including camera-based 3D HPE and LiDAR-based 3D HPE in urban environments, and benchmark a new task, GHPE. The in-depth analysis demonstrates SLOPER4D poses significant challenges to existing methods and produces great research opportunities. The dataset and code are released at http://www.lidarhumanmotion.net/sloper4d/. Yudi Dai, Yitai Lin, Xiping Lin, Chenglu Wen, Lan Xu 0003, Hongwei Yi, Yuexin Ma, Cheng Wang 0003 |
CVPR | 4 |
| 2023 | SGLoc: Scene Geometry Encoding for Outdoor LiDAR LocalizationabstractLiDAR-based absolute pose regression estimates the global pose through a deep network in an end-to-end manner, achieving impressive results in learning-based localization. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding the scene geometry and the unsatisfactory quality of the data. In this work, we propose a novel LiDAR localization frame-work, SGLoc, which decouples the pose estimation to point cloud correspondence regression and pose estimation via this correspondence. This decoupling effectively encodes the scene geometry because the decoupled correspondence regression step greatly preserves the scene geometry, leading to significant performance improvement. Apart from this decoupling, we also design a tri-scale spatial feature aggregation module and inter-geometric consistency constraint loss to effectively capture scene geometry. Moreover, we empirically find that the ground truth might be noisy due to GPS/INS measuring errors, greatly reducing the pose estimation performance. Thus, we propose a pose quality evaluation and enhancement method to measure and correct the ground truth pose. Extensive experiments on the Oxford Radar RobotCar and NCLT datasets demonstrate the effectiveness of SGLoc, which outperforms state-of-the-art regression-based localization methods by 68.5% and 67.6% on position accuracy, respectively. Wen Li 0005, Shangshu Yu, Cheng Wang 0003, Guosheng Hu, Chenglu Wen |
CVPR | 6 |
| 2023 | Virtual Sparse Convolution for Multimodal 3D Object DetectionabstractRecently, virtuall pseudo-point-based 3D object detection that seamlessly fuses RGB images and LiDAR data by depth completion has gained great attention. However, virtual points generated from an image are very dense, introducing a huge amount of redundant computation during detection. Meanwhile, noises brought by inaccurate depth completion significantly degrade detection precision. This paper proposes a fast yet effective backbone, termed Vir-ConvNet, based on a new operator VirConv (Virtual Sparse Convolution), for virtual-point-based 3D object detection. VirConv consists of two key designs: (1) StVD (Stochastic Voxel Discard) and (2) NRConv (Noise-Resistant Sub-manifold Convolution). StVD alleviates the computation problem by discarding large amounts of nearby redundant voxels. NRConv tackles the noise problem by encoding voxel features in both 2D image and 3D LiDAR space. By integrating VirConv, we first develop an efficient pipeline VirConv-L based on an early fusion design. Then, we build a high-precision pipeline Vir Conv-T based on a transformed refinement scheme. Finally, we develop a semi-supervised pipeline VirConv-S based on a pseudo-label framework. On the KITTI car 3D detection test leader-board, our VirConv-L achieves 85% AP with a fast running speed of 56ms. Our VirConv-T and VirConv-S attains a high-precision of 86.3% and 87.2% AP, and currently rank 2ndand 1st11On the date of CVPR deadline, i. e., Nov.11, 2022, respectively. The code is available at https://github.com/hailanyi/VirConv. Chenglu Wen, Shaoshuai Shi, Xin Li 0003, Cheng Wang 0003 |
CVPR | 2 |
| 2023 | CIMI4D: A Large Multimodal Climbing Motion Dataset under Human-scene InteractionsabstractMotion capture is a long-standing research problem. Although it has been studied for decades, the majority of research focus on ground-based movements such as walking, sitting, dancing, etc. Off- grounded actions such as climbing are largely overlooked. As an important type of action in sports and firefighting field, the climbing movements is challenging to capture because of its complex back poses, intricate human-scene interactions, and difficult global localization. The research community does not have an indepth understanding of the climbing action due to the lack of specific datasets. To address this limitation, we collect CIMI4D, a large rock Climbing Motion dataset from 12 persons climbing 13 different climbing walls. The dataset consists of around 180,000 frames of pose inertial measurements, LiDAR point clouds, RGB videos, high-precision static point cloud scenes, and reconstructed scene meshes. Moreover, we frame-wise annotate touch rock holds to facilitate a detailed exploration of human-scene interaction. The core of this dataset is a blending optimization process, which corrects for the pose as it drifts and is affected by the magnetic conditions. To evaluate the merit of CIMI4D, we perform four tasks which include human pose estimations (with/without scene constraints), pose prediction, and pose generation. The experimental results demonstrate that CIMI4D presents great challenges to existing methods and enables extensive research opportunities. We share the dataset with the research community in http://www.lidarhumanmotion.net/cimi4d/. Yudi Dai, Chenglu Wen, Lan Xu 0003, Yuexin Ma, Cheng Wang 0003 |
CVPR | 5 |
| 2023 | CoIn: Contrastive Instance Feature Mining for Outdoor 3D Object Detection with Very Limited AnnotationsabstractRecently, 3D object detection with sparse annotations has received great attention. However, current detectors usually perform poorly under very limited annotations. To address this problem, we propose a novel Contrastive Instance feature mining method, named CoIn. To better identify indistinguishable features learned through limited supervision, we design a Multi-Class contrastive learning module (MCcont) to enhance feature discrimination. Meanwhile, we propose a feature-level pseudo-label mining framework consisting of an instance feature mining module (InF-Mining) and a Labeled-to-Pseudo contrastive learning module (LPcont). These two modules exploit latent instances in feature space to supervise the training of detectors with limited annotations. Extensive experiments with KITTI dataset, Waymo open dataset, and nuScenes dataset show that under limited annotations, our method greatly improves the performance of baseline detectors: CenterPoint, Voxel-RCNN, and CasA. Combining CoIn with an iterative training strategy, we propose a CoIn++ pipeline, which requires only 2% annotations in the KITTI dataset to achieve performance comparable to the fully supervised methods. The code is available at https://github.com/xmuqimingxia/CoIn. Qiming Xia, Jinhao Deng, Chenglu Wen, Shaoshuai Shi, Xin Li 0003, Cheng Wang 0003 |
ICCV | 3 |
| 2023 | Semantic Segmentation of Point Cloud With Novel Neural Radiation Field ConvolutionabstractPoint cloud semantic segmentation predicts the semantic class of each point, which can help AI machines perceive the real 3-D world. Recently, precision for large-scale point cloud is limited by complex scenarios, data occlusion, and massive data, which remains. This letter proposes neural radiance field convolution (NeRFConv) for large-scale point cloud semantic analysis. First, to conquer some existing methods represent the point cloud local space as the relative position without considering the rotation property of the point cloud. A new spatial direction representation called neural radiance field 7-D is constructed to denote the local point cloud space, which is unaffected by the rotation of the coordinate axes. Then, to be robust against the varying density of the point cloud, an adaptive convolution method is proposed to aggregate the local point cloud features and obtain more expressive local features. Finally, the proposed method achieves state-of-the-art performance on the test dataset. In particular, mIoU of 58.1% and 58.3% were achieved the SensatUrban and SementicKITTI datasets, respectively. Wei Li 0151, Lixin Zhan, Weidong Min, Chenglu Wen |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | DeepSIR: Deep semantic iterative registration for LiDAR point clouds
Qing Li 0032, Cheng Wang 0003, Chenglu Wen, Xin Li 0003 |
Pattern Recognit. | 3 |
| 2023 | Adaptive local adversarial attacks on 3D point clouds
Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
Pattern Recognit. | 5 |
| 2023 | SemanticFlow: Semantic Segmentation of Sequential LiDAR Point Clouds From Sparse Frame AnnotationsabstractSequential point clouds acquired by light detection and ranging (LiDAR) technology provide accurate spatial information for environmental sensing. However, semantic segmentation of point cloud sequences relies on many manual point-wise annotations, which are error-prone and expensive. Existing mainstream weakly supervised methods tackle this by reducing the percentage of labeled points, but they are mostly designed for static indoor scenes and are hard to apply practically. From the viewpoint of realistic annotation procedures and the nature of point cloud sequences, this paper proposes a novel semantic segmentation method, SemanticFlow, for LiDAR point cloud sequences using sparse frames with annotations. The proposed method achieves competitive performance compared with fully supervised methods. Specifically, we designed a bidirectional cross-frame pseudo label propagation module that uses scene flow to learn the correlation and propagate pseudo labels across neighboring frames. In addition, a label refinement mechanism is proposed to select reliable pseudo labels for learning. Extensive experiments on SemanticKITTI, SemanticPOSS, and Synthia 4D datasets demonstrate that our sparse frame annotation method is compatible with some fully supervised counterparts. Junhao Zhao, Chenglu Wen, Bo Yang 0027, Yulan Guo, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | STCLoc: Deep LiDAR Localization With Spatio-Temporal ConstraintsabstractLiDAR localization is of great importance to autonomous vehicles and robotics. Absolute pose regression, directly estimating the mapping from a scene to a 6-DoF pose, has achieved impressive results in learning-based localization. Different from traditional map-based methods, it does not need a pre-built 3D map during inference. However, current regression networks typically suffer from scene ambiguities, especially in challenging traffic environments, leading to large wrong predictions (e.g., outliers) and limited applications. To address this problem, a novel LiDAR localization framework with spatio-temporal constraints is proposed, termed STCLoc, to reduce scene ambiguities and achieve more accurate localization. First, we propose to regularize regression in the spatial dimension with a novel classification task to reduce outliers. Specifically, the classification task categorizes the point cloud in terms of position and orientation and then couples it with the regression task to conduct multi-task learning. Second, to learn discriminative features to reduce scene ambiguities, we propose using attention-based feature aggregation to capture the correlation in LiDAR sequences. We conduct extensive experiments on two benchmark datasets, where the localization takes 97ms on each dataset. Results show that our model outperforms state-of-the-art methods by 43.33%/36.76% (position/orientation) on the Oxford Radar RobotCar dataset, verifying the effectiveness of our method. The source code is available on the project website athttps://github.com/PSYZ1234/STCLoc. Shangshu Yu, Cheng Wang 0003, Yitai Lin, Chenglu Wen, Ming Cheng 0002, Guosheng Hu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Category-Level Adversaries for Outdoor LiDAR Point Clouds Cross-Domain Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is a low-cost way to deal with the lack of annotations in a new domain. For outdoor point clouds in urban transportation scenes, the mismatch of sampling patterns and the transferability difference between classes make cross-domain segmentation extremely difficult. To overcome these challenges, we propose a category-level adversarial framework. Firstly, we propose a multi-scale domain conditioned block that facilitates to extract the critical low-level domain-dependent knowledge and reduce the domain gap caused by distinct LiDAR sampling patterns. Secondly, we make full use of multiple representation forms (i.e., point-based sets and voxel-based cells) and utilize the prediction consistency between the two forms to measure how well each point is semantically aligned. The model then focuses on the poorly-aligned points without affecting the well-aligned points. Experimental results on three autonomous driving point cloud datasets show that the proposed method outperforms existing methods by a large margin, especially on the low-beam to high-beam cross-domain segmentation task. Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | GoComfort: Comfortable Navigation for Autonomous Vehicles Leveraging High-Precision Road Damage CrowdsensingabstractRecent years have witnessed rapid advances in autonomous driving technologies. Autonomous vehicles are more likely to be accepted if they drive comfortably to avoid potholes, bumps, and other road damage conditions, especially when there are elderly and disabled passengers on board. Traditionally, sensing road damage conditions is either labor-intensive by field investigation and reporting, or inaccurate by surveillance cameras and driving recorders due to limited perspective. In this paper, we propose GoComfort, a crowdsensing-based framework to provide low-cost and fine-grained comfortable navigation for autonomous vehicles with road damage identification leveraging high-precision road sensing data. First, we propose to exploit city-wide autonomous vehicle fleets as crowdsensing participants, and employ an edge-cloud-hybrid computing paradigm to efficiently collect high-precision road damage-related data, including 3D LiDAR point clouds and street view images. Second, we design an accurate road damage identification model fusing spatial structures of point clouds and texture features of street view images, and use an active learning-based method to address the sparse labels issue. Finally, we devise two comfortable navigation scenarios, i.e., fine-grained road damage avoidance and coarse-grained city-wide navigation, and propose a hierarchical road damage assessment diagram for comfortable route planning. Experiments using real-world road sensing data in Xiamen, China show that our approach identifies road damage conditions with an accuracy of 87.5%, and achieves a user acceptance rate of 93.8% in riding comfort evaluation, outperforming the state-of-the-art baselines. Longbiao Chen, Xin He 0030, Xiantao Zhao, Yunyi Huang, Binbin Zhou 0005, Wei Chen 0001, Yongchuan Li, Chenglu Wen, Cheng Wang 0003 |
IEEE Trans. Mob. Comput. | 9 |
| 2022 | HSC4D: Human-centered 4D Scene Capture in Large-scale Indoor-outdoor Space Using Wearable IMUs and LiDARabstractWe propose Human-centered 4D Scene Capture (HSC4D) to accurately and efficiently create a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, and rich interactions between humans and environments. Using only body-mounted IMUs and LiDAR, HSC4D is space-free without any external devices' constraints and map-free without pre-built maps. Considering that IMUs can capture human poses but always drift for long-period use, while LiDAR is stable for global localization but rough for local positions and orientations, HSC4D makes both sensors complement each other by a joint optimization and achieves promising results for long-term capture. Relationships between humans and environments are also explored to make their interaction more realistic. To facilitate many down-stream tasks, like AR, VR, robots, autonomous driving, etc., we propose a dataset containing three large scenes (1k-5k m2) with accurate dynamic human motions and locations. Diverse scenarios (climbing gym, multi-story building, slope, etc.) and challenging human activities (exercising, walking up/down stairs, climbing, etc.) demonstrate the effectiveness and the generalization ability of HSC4D. The dataset and code is available at lidarhumanmotion.net/hsc4d. Yudi Dai, Yitai Lin, Chenglu Wen, Lan Xu 0003, Jingyi Yu 0001, Yuexin Ma, Cheng Wang 0003 |
CVPR | 3 |
| 2022 | LiDARCap: Long-range Markerless 3D Human Motion Capture with LiDAR Point CloudsabstractExisting motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captured by LiDAR at a much longer range to overcome this limitation. Our dataset also includes the ground truth human motions acquired by the IMU system and the synchronous RGB images. We further present a strong base-line method, LiDARCap, for LiDAR point cloud human motion capture. Specifically, we first utilize$PointNet++$to encode features of points and then employ the inverse kinematics solver and SMPL optimizer to regress the pose through aggregating the temporally encoded features hierarchically. Quantitative and qualitative experiments show that our method outperforms the techniques based only on RGB images. Ablation experiments demonstrate that our dataset is challenging and worthy of further research. Finally, the experiments on the KITTI Dataset and the Waymo Open Dataset show that our method can be generalized to different LiDAR sensor settings. Jialian Li, Chenglu Wen, Yuexin Ma, Lan Xu 0003, Jingyi Yu 0001, Cheng Wang 0003 |
CVPR | 5 |
| 2022 | A LiDAR-Based 3D Indoor Mapping Framework with Mismatch Detection and OptimizationabstractIn this paper, we propose a novel LiDAR-based mapping framework for geometry-featureless scenarios, combining learning-based mismatch detection and intensity-assisted registration optimization. The mismatch detection method hybridizes point cloud features and temporal features of pose to detect the mismatch position, and the optimization method uses the intensity of captured road markers by LiDAR to optimize mismatch pose. Experiments on underground garages data demonstrate that the proposed method performs better localization and mapping accuracy than the LOAM. Weiquan Liu, Chenglu Wen, Yongfei Shi, Xiaocheng Yan, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 3 |
| 2022 | Vehicle Completion in Traffic Scene Using 3D LiDAR Point Cloud DataabstractModern autonomous vehicles perceive their surroundings with the help of 3D computer vision technologies (e.g., object detection, vehicle recognition). The 3D point cloud obtained by vehicle-mounted mobile LiDAR is the primary data for 3D visual tasks in the driving system. Due to partial observations, a complete point cloud of the surrounding vehicle cannot be obtained. In this paper, we proposed Point Voxel Completion Network (PVCNet), an end-to-end learning-based model for 3D vehicle completion. Unlike existing point cloud completion methods, PVCNet only predicts the missing parts of the input and preserves the original spatial structure of the point cloud. PVCNet consists of two branches, the voxel branch and the point cloud branch. The voxel branch extracts the local features of the input and predicts the position of the missing point cloud in voxel form. The point cloud branch extracts the global features of the input point cloud. Combined with features from two branches, PVCNet can generate the missing points. The proposed PVCNet is validated on our hybrid dataset and KITTI dataset and performed competitive results in the task of urban vehicle point cloud completion. Chongrong Wu, Yitai Lin, Chenglu Wen, Yongfei Shi, Cheng Wang 0003 |
IGARSS | 4 |
| 2022 | BLNet: Bidirectional learning network for point cloudsabstractThe key challenge in processing point clouds lies in the inherent lack of ordering and irregularity of the 3D points. By relying on perpoint multi-layer perceptions (MLPs), most existing point-based approaches only address the first issue yet ignore the second one. Directly convolving kernels with irregular points will result in loss of shape information. This paper introduces a novel point-based bidirectional learning network (BLNet) to analyze irregular 3D points. BLNet optimizes the learning of 3D points through two iterative operations: feature-guided point shifting and feature learning from shifted points, so as to minimise intra-class variances, leading to a more regular distribution. On the other hand, explicitly modeling point positions leads to a new feature encoding with increased structure-awareness. Then, an attention pooling unit selectively combines important features. This bidirectional learning alternately regularizes the point cloud and learns its geometric features, with these two procedures iteratively promoting each other for more effective feature learning. Experiments show that BLNet is able to learn deep point features robustly and efficiently, and outperforms the prior state-of-the-art on multiple challenging tasks. Wenkai Han, Chenglu Wen, Cheng Wang 0003, Xin Li 0003 |
Comput. Vis. Media | 3 |
| 2022 | Planar Primitive Group-Based Point Cloud Registration for Autonomous Vehicle Localization in Underground Parking LotsabstractWe present a registration strategy based on a planar primitive group for indoor environments between a large-scale point cloud and a small-scale point cloud, providing a localization solution that is fully independent of prior information about the initial positions of the two point cloud coordinate systems. The algorithm first divides the point cloud into planes by region growing and then refines the plane boundaries by local$k$-means clustering. The planes are grouped to obtain planar primitive groups that are used to infer potential matching regions for an effective coarse registration and then fine-tuning. The algorithm is applied for autonomous vehicle localization in underground garages. Evaluation on four datasets shows that the algorithm can provide decimeter-level localization accuracy in seconds. Lili Lin, Ming Cheng 0002, Chenglu Wen, Cheng Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Learning scale awareness in keypoint extraction and description
Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Yifan Peng 0001, Chenglu Wen, Ming Cheng 0002 |
Pattern Recognit. | 6 |
| 2022 | LiDAR-based localization using universal encoding and memory-aware regression
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 3 |
| 2022 | Corrigendum to "LiDAR-based localization using universal encoding and memory-aware regression" Pattern Recognition Volume 128 (2022) 108685
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 3 |
| 2022 | Indoor 3D Human Trajectory Reconstruction Using Surveillance Camera Videos and Point Cloudsabstract3D human trajectory reconstruction in an indoor scene is critical in various applications, such as indoor navigation and human activity recognition. This task is challenging due to occlusion and clutters of indoor scenes, flexible human body joints, and severe lack of relevant datasets. Although several methods have been proposed to reconstruct a 3D human trajectory, they either can recover only 2D positions or require human initiative cooperation. In this paper, we propose a novel framework for 3D human trajectory reconstruction in an indoor scene using monocular surveillance videos and static point clouds without any initiative cooperation. The proposed framework consists of three modules: 3D pose estimation, depth regression, and trajectory reconstruction. We first estimate 3D pose from videos. Especially, we reconstruct a half-body 3D pose to deal with the occlusion problem. Then, we propose a depth regression approach to iteratively regress the depth of a 3D pose. Unlike data-driven approaches, our depth regression approach does not require training data and can be integrated into any 3D pose model. Finally, we exploit the geometric constraints from the point cloud to optimize the 3D trajectory. We evaluated the 3D pose estimation and depth regression modules on the H3.6M datasets. Due to the lack of evaluation datasets, we also built a trajectory dataset to evaluate the trajectory reconstruction performance. Empirical evaluation shows that our framework achieves accurate trajectory reconstruction results on real-world videos. Yudi Dai, Chenglu Wen, Yulan Guo, Longbiao Chen, Cheng Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | CasA: A Cascade Attention Network for 3-D Object Detection From LiDAR Point Cloudsabstract3D object detection from LiDAR point clouds has gained great attention in recent years due to its wide applications in smart cities and autonomous driving. Cascade framework shows its advancement in 2D object detection but is less investigated in 3D space. Conventional cascade structures use multipleseparatesub-networks to sequentially refine region proposals. Such methods, however, have limited ability to measure proposal quality in all stages, and hard to achieve a desirable performance improvement in 3D space. This paper proposes a new cascade framework, termed CasA, for 3D object detection from LiDAR point clouds. CasA consists of a Region Proposal Network (RPN) and a Cascade Refinement Network (CRN). In CRN, we designed a new Cascade Attention Module that uses multiple sub-networks and attention modules to aggregate the object features from different stages and progressively refine region proposals. CasA can be integrated into various two-stage 3D detectors and improve their performance. Extensive experiments on KITTI and Waymo datasets with various baseline detectors demonstrate the universality and superiority of our CasA. In particular, based on one variant of Voxel-RCNN, we achieve state-of-the-art results on the KITTI dataset. On the KITTI online 3D object detection leaderboard, we achieve a high detection performance of 83.06%, 47.09%, and 73.47% Average Precision (AP) in the moderate Car, Pedestrian, and Cyclist classes, respectively. Code is available at https://github.com/hailanyi/CasA. Jinhao Deng, Chenglu Wen, Xin Li 0003, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | 3D Multi-Object Tracking in Point Clouds Based on Prediction Confidence-Guided Data AssociationabstractThis paper proposes a new 3D multi-object tracker to more robustly track objects that are temporarily missed by detectors. Our tracker can better leverage object features for 3D Multi-Object Tracking (MOT) in point clouds. The proposed tracker is based on a novel data association scheme guided by prediction confidence, and it consists of two key parts. First, we design a new predictor that employs a constant acceleration (CA) motion model to estimate future positions, and outputs a prediction confidence to guide data association through increased awareness of detection quality. Second, we introduce a new aggregated pairwise cost to exploit features of objects in point clouds for faster and more accurate data association. The proposed cost consists of geometry, appearance and motion components. Specifically, we formulate the geometry cost using resolutions (lengths, widths and heights), centroids, and orientations of 3D bounding boxes (BBs), the appearance cost using appearance features from the deep learning-based detector backbone network, and the motion cost by associating different motion vectors. Extensive multi-object tracking experiments on the KITTI tracking benchmark demonstrated that our method outperforms, by a large margin, the state-of-the-art methods in both tracking accuracy and speed. Wenkai Han, Chenglu Wen, Xin Li 0003, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Metric Learning for 2D Image Patch and 3D Point Cloud Volume MatchingabstractSimilarity measure of cross-domain descriptors (2D descriptors and 3D descriptors) between 2D image patches and 3D point cloud volumes provides stable retrieval performance and establishes the spatial relationship between 2D and 3D space, which plays the potential applications in geospatial space, such as 2D and 3D interaction of remote sensing, Augmented Reality (AR) and robot navigation. However, the mature handcrafted descriptors of 2D image patches and 3D point cloud volumes are extremely different, resulting in the huge challenge for 2D image patch and 3D point cloud volume matching. In this paper, we propose a novel network which combines both unified descriptor training and descriptor comparison function training for 2D image patch and 3D point cloud volume matching. First, two feature extraction networks are applied for jointly learning the local descriptors for 2D image patches and 3D point cloud volumes, respectively. Second, a fully connected network is introduced to compute the similarity between 2D descriptors and 3D descriptors. Motivated by the successful indicator system on evaluating 2D patch feature representation, we use the false positive rate at 95% recall (FPR95) and precision based on cross-domain descriptors as the measured metric. The experimental results show that our proposed network achieve state-of-the-art performance in the matching of 2D image patches and 3D point cloud volumes. Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Xiuhong Lin, Chenglu Wen, Jonathan Li 0001 |
IGARSS | 7 |
| 2021 | Federated Learning with Fair AveragingabstractFairness has emerged as a critical problem in federated learning (FL). In this work, we identify a cause of unfairness in FL -- conflicting gradients with large differences in the magnitudes. To address this issue, we propose the federated fair averaging (FedFV) algorithm to mitigate potential conflicts among clients before averaging their gradients. We first use the cosine similarity to detect gradient conflicts, and then iteratively eliminate such conflicts by modifying both the direction and the magnitude of the gradients. We further show the theoretical foundation of FedFV to mitigate the issue conflicting gradients and converge to Pareto stationary solutions. Extensive experiments on a suite of federated datasets confirm that FedFV compares favorably against state-of-the-art methods in terms of fairness, accuracy and efficiency. The source code is available at https://github.com/WwZzz/easyFL. Zheng Wang 0076, Xiaoliang Fan, Jianzhong Qi 0001, Chenglu Wen, Cheng Wang 0003, Rongshan Yu |
IJCAI | 4 |
| 2021 | Tracklet Proposal Network for Multi-Object Tracking on Point CloudsabstractThis paper proposes the first tracklet proposal network, named PC-TCNN, for Multi-Object Tracking (MOT) on point clouds. Our pipeline first generates tracklet proposals, then refines these tracklets and associates them to generate long trajectories. Specifically, object proposal generation and motion regression are first performed on a point cloud sequence to generate tracklet candidates. Then, spatial-temporal features of each tracklet are exploited and their consistency is used to refine the tracklet proposal. Finally, the refined tracklets across multiple frames are associated to perform MOT on the point cloud sequence. The PC-TCNN significantly improves the MOT performance by introducing the tracklet proposal design. On the KITTI tracking benchmark, it attains an MOTA of 91.75%, outperforming all submitted results on the online leaderboard. Qing Li 0032, Chenglu Wen, Xin Li 0003, Xiaoliang Fan, Cheng Wang 0003 |
IJCAI | 3 |
| 2021 | Cooperative indoor 3D mapping and modeling using LiDAR data
Chenglu Wen, Jinbin Tan, Fashuai Li, Chongrong Wu, Yitai Lin, Cheng Wang 0003 |
Inf. Sci. | 1 |
| 2021 | FeatFlow: Learning geometric features for 3D motion estimation
Qing Li 0032, Cheng Wang 0003, Xin Li 0003, Chenglu Wen |
Pattern Recognit. | 4 |
| 2021 | Mapping and Semantic Modeling of Underground Parking Lots Using a Backpack LiDAR SystemabstractPresented in this paper is a novel method for the mapping and semantic modeling of an underground parking lot using 3D point clouds collected by a low-cost Backpack Laser Scanning (BLS) or LiDAR system. Our method consists of two parts: a Simultaneous Localization and Mapping (SLAM) algorithm based on Sparse Point Clouds (SPC) and a semantic modeling algorithm based on a modified PointNet model. The main contributions of this paper are as follows: (1) a probability frontend framework for the alignment of point clouds using the local point cloud surface variance as the weight of registration, which modifies registration failure caused by the lack of features in sparse point clouds, (2) a robust submap-based strategy for loop closure detection and back-end optimization under sparse point clouds, and (3) a modified PointNet model for classifying the point clouds of underground parking lots into four categories: ceiling, floor, wall, others. Experimental results show that our SPC-SLAM algorithm achieves centimeter-level accuracy (0.09% trajectory error rate) after closed loop processing in a Global Navigation Satellite System (GNSS)-denied underground parking lot, and precision of 84.8% in semantic segmentation. Jonathan Li 0001, Chenglu Wen, Cheng Wang 0003, John S. Zelek |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Special Issue on 3D Sensing in Intelligent TransportationabstractHigh-Accuracy and high-efficiency 3-D sensing and associated data processing techniques are urgently needed for today’s roadway inventory, infrastructure health monitoring, autonomous driving, connected vehicles, urban modeling, and smart cities. 3D geospatial data acquired by digital photogrammetry or laser scanning or LiDAR systems have become one of the most critical data sources to support the above-mentioned applications. While progress has been made to applying 3D sensory data to those applications related to intelligent transportation systems (ITS), such as road network extraction, platform localization, obstacle avoidance, high-definition map generation, and transportation infrastructure inventory, many essential questions remain regarding the processing and understanding such massive 3D datasets in ITS-related applications. The authors have selected four articles for review in this Special issue. A summary of these articles is outlined below. Chenglu Wen, Ayman Habib 0001, Jonathan Li 0001, Charles K. Toth, Cheng Wang 0003, Hongchao Fan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Point2Node: Correlation Learning of Dynamic-Node for Point Cloud Feature ModelingabstractFully exploring correlation among points in point clouds is essential for their feature modeling. This paper presents a novel end-to-end graph model, named Point2Node, to represent a given point cloud. Point2Node can dynamically explore correlation among all graph nodes from different levels, and adaptively aggregate the learned features. Specifically, first, to fully explore the spatial correlation among points for enhanced feature description, in a high-dimensional node graph, we dynamically integrate the node's correlation with self, local, and non-local nodes. Second, to more effectively integrate learned features, we design a data-aware gate mechanism to self-adaptively aggregate features at the channel level. Extensive experiments on various point cloud benchmarks demonstrate that our method outperforms the state-of-the-art. Wenkai Han, Chenglu Wen, Cheng Wang 0003, Xin Li 0003, Qing Li 0032 |
AAAI | 2 |
| 2020 | Pairwise registration of TLS point clouds by deep multi-scale local features
Wei Li 0151, Cheng Wang 0003, Chenglu Wen, Congren Lin, Jonathan Li 0001 |
Neurocomputing | 3 |
| 2020 | Toward Efficient 3-D Colored Mapping in GPS-/GNSS-Denied EnvironmentsabstractEfficient 3-D mapping provides useful and detailed 3-D data for many applications. In this letter, we present a multisensor calibration and mapping method, to provide highly efficient and relatively accurate colored mapping for GPS-/global navigation satellite system-denied environments. The sensor data include 3-D laser scanning point clouds and camera images. A simultaneous localization and mapping (SLAM)-assisted calibration method is first proposed for multiple multibeam light detection and ranging (LiDAR) and multiple camera calibration. An improved SLAM method with loop closure is proposed for 3-D mapping. With the proposed calibration and mapping methods, centimeter-level colored point clouds can be obtained efficiently. The proposed method was tested with both backpacked and car-mounted systems on indoor and outdoor scenes. Experimental results show the effectiveness and efficiency of the proposed calibration and mapping methods. Chenglu Wen, Yudi Dai, Yan Xia 0003, Yuhan Lian, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | DeepSTD: Mining Spatio-Temporal Disturbances of Multiple Context Factors for Citywide Traffic Flow PredictionabstractDeep learning techniques have been widely applied to traffic flow prediction, considering underlying routine patterns, and multiple context factors (e.g., time and weather). However, the complex spatio-temporal dependencies between inherent traffic patterns and multiple disturbances have not been fully addressed. In this paper, we propose a two-phase end-to-end deep learning framework, namely DeepSTD to uncover the spatio-temporal disturbances (STD) to predict the citywide traffic flow. In the STD Modeling phase, we propose an STD modeling method to model both the different regional disturbances caused by various region functions and the spatio-temporal propagating effects. In the Prediction phase, we eliminate the STD from the historical traffic flow to enhance the leaning of inherent traffic patterns and combine the STD at the prediction time interval to consider the future disturbances. The experimental results on two real-world datasets demonstrate that DeepSTD outperforms the state-of-the-art methods. Chuanpan Zheng, Xiaoliang Fan, Chenglu Wen, Longbiao Chen, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Urban 3D modeling with mobile laser scanning: a reviewabstractMobile laser scanning (MLS) systems mainly comprise laser scanners and mobile mapping platforms. Typical MLS systems are able to acquire three-dimensional point clouds with 1-10 centimeter point spacing at a normal driving or walking speed in the street or indoor environments. The MLS' advantages of efficiency and stability make it a quite practical tool for three-dimensional urban modeling. This paper reviews the latest advances in 3D modeling of the LiDAR-based mobile mapping system (MMS) point cloud, including LiDAR Simultaneous Localization and Mapping (SLAM), point cloud registration, feature extraction, object extraction, semantic segmentation, and deep learning processing. Then typical urban modeling applications based on MMS are also discussed. Cheng Wang 0003, Chenglu Wen, Yudi Dai, Shangshu Yu, Minghao Liu 0007 |
Virtual Real. Intell. Hardw. | 2 |
| 2019 | LO-Net: Deep Real-Time Lidar OdometryabstractWe present a novel deep convolutional network pipeline, LO-Net, for real-time lidar odometry estimation. Unlike most existing lidar odometry (LO) estimations that go through individually designed feature selection, feature matching, and pose estimation pipeline, LO-Net can be trained in an end-to-end manner. With a new mask-weighted geometric constraint loss, LO-Net can effectively learn feature representation for LO estimation, and can implicitly exploit the sequential dependencies and dynamics in the data. We also design a scan-to-map module, which uses the geometric and semantic information learned in LO-Net, to improve the estimation accuracy. Experiments on benchmark datasets demonstrate that LO-Net outperforms existing learning based approaches and has similar accuracy with the state-of-the-art geometry-based approach, LOAM. Qing Li 0032, Shaoyang Chen, Cheng Wang 0003, Xin Li 0003, Chenglu Wen, Ming Cheng 0002, Jonathan Li 0001 |
CVPR | 5 |
| 2019 | RF-Net: An End-To-End Image Matching Network Based on Receptive FieldabstractThis paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable matching framework is desirable and challenging. The very recent approach, LF-Net, successfully embeds the entire feature extraction pipeline into a jointly trainable pipeline, and produces the state-of-the-art matching results. This paper introduces two modifications to the structure of LF-Net. First, we propose to construct receptive feature maps, which lead to more effective keypoint detection. Second, we introduce a general loss function term, neighbor mask, to facilitate training patch selection. This results in improved stability in descriptor training. We trained RF-Net on the open dataset HPatches, and compared it with other methods on multiple benchmark datasets. Experiments show that RF-Net outperforms existing state-of-the-art methods. Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Zenglei Yu, Jonathan Li 0001, Chenglu Wen, Ming Cheng 0002 |
CVPR | 6 |
| 2019 | Partial 3D Object Retrieval and Completeness Evaluation for Urban Street Sceneabstract3D objects detected from real-world data are usually incomplete in different degrees. Objects with different degrees of incompleteness should be treated and processed separately. This paper proposes a framework for partial 3D object retrieval and completeness evaluation in an urban street scene based on mobile laser scanning (MLS) point cloud data. The framework consists of three parts. A deep learning method is first used to detect objects from 3D point cloud data. Then, for each detected object, the most similar object in the reference dataset, which contains complete objects, is obtained by a partial 3D shape retrieval method. Last, a completeness evaluation of the detected object is conducted by calculating the completeness index that reflects the integrity of the detected object, and a missing part prediction is given to guide further completion. The proposed framework is validated on the public dataset KITTI and our own point cloud dataset. The experiment includes 3D detection, the partial 3D shape retrieval, and the completeness evaluation. Results show the good performance of the object detection and partial shape retrieval, also a reasonable evaluation of objects completeness. Chenglu Wen, Xiaotian Sun 0005, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2019 | Non-Reference Quality Evaluation for Indoor 3d Point CloudsabstractThis paper proposes a novel approach for indoor point clouds quality evaluation, which works well without reference point clouds. In this paper, we mainly evaluate indoor point clouds quality in two aspects: the smoothness of the walls and the degree of occlusion of the walls and the floor. Our approach involves three steps. Firstly, with the S3DIS dataset, we use a deep learning method to train a detector to label walls and floor from indoor scenes. Next, we calculate the normal vector of the wall and the normal vector of each point on the wall. The degree of smoothness of the wall is judged according to the angle between the normal vector of the wall and the normal vectors of the points. Finally, according to the cause of occlusion (objects on the floor or in front of the wall), the occlusion degree of the wall and the floor is obtained. The experimental results demonstrate that the proposed method is suitable for non-reference indoor point cloud quality evaluation. Yuhan Lian, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2019 | Joint 2-D-3-D Traffic Sign Landmark Data Set for Geo-Localization Using Mobile Laser Scanning DataabstractThis paper presents a framework to build a joint 2-D-3-D traffic sign landmark data set for geo-localization using mobile laser scanning (MLS) data. The MLS data include 3-D point clouds and corresponding multi-view images. First, an integrated method, based on a deep learning network and the retro-reflective properties of traffic signs, is developed to accurately extract traffic signs from MLS point clouds. Next, the semantic and spatial properties of the traffic signs (type, location, position, and geometric characteristics) are obtained. Then, a joint 2-D-3-D traffic sign landmark data set is built, and a semantic-spatial organization graph is used to organize the traffic sign data set. Last, based on the traffic sign landmark data set, a geo-localization method for a driving car is proposed to estimate the driving trajectory. It can be used for auxiliary positioning of autonomous vehicles. Experimental results demonstrate the reliability of our proposed method for traffic sign detection and the potential of building 2-D-3-D traffic sign landmark data set for driving trajectory estimation from MLS data. Changbin You, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001, Ayman Habib 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Dection and Health Analysis of Individual Tree in Urban Environment with Multi-Sensor PlatformabstractWith the technology enhanced, 3D mobile light detection and ranging (LiDAR) can produce more accurate 3D information for the objects. Meanwhile, hyperspectral remote sensing has more number of wavelengths and provides a higher resolution spectrum of objects. This paper proposes a multi-sensor platform to provide these two data for health detection at the individual tree level in urban environments. We firstly locate and segment the suspected tree objects by ground removal and Euclidean distance clustering. Then we take use of spectrum to remove non-tree objects, e.g., buildings, light poles. After that, we use LiDAR data to compute the geometric parameters of each tree and hyperspectral data to analyze its health situation. Yunhe Feng, Chenglu Wen, Pengdi Huang, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2018 | Discriminative Learning of Point Cloud Feature Descriptors Based on Siamese NetworkabstractIt is challenging to direct extract the feature descriptors of the object in the point cloud, although deep learning has been widely used with the classification and detection in the point cloud, those methods hidden feature presentation in the network. Since the point cloud scanned by the Laser Scanner usually have different point density, unordered and even the different occlusion, which go beyond the reach of handcrafted descriptors, e.g. FPH, FPFH, VFH, ROPS. In this paper, we aim to direct extract the feature descriptors of the point cloud object through the raw point cloud. Inspired by the recent success of the Siamese networks[6], PointNet[7] and PointNet++[8], we propose a novel network to direct extract the feature descriptors of the whole point cloud object. We train our network with the Euclidean distance as the loss function which reflects feature descriptors similarity. The experiment object datasets were acquired by Mobile Laser Scanning (MLS) system which contains 6 categories. Experiment result shows that our network has a robust generalization, which can well direct extract the feature descriptors of the whole point cloud object. Xuelun Shen, Cheng Wang 0003, Chenglu Wen, Weiquan Liu, Xiaotian Sun 0005, Jonathan Li 0001 |
IGARSS | 3 |
| 2018 | H-Net: Neural Network for Cross-domain Image Patch MatchingabstractDescribing the same scene with different imaging style or rendering image from its 3D model gives us different domain images. Different domain images tend to have a gap and different local appearances, which raise the main challenge on the cross-domain image patch matching. In this paper, we propose to incorporate AutoEncoder into the Siamese network, named as H-Net, of which the structural shape resembles the letter H. The H-Net achieves state-of-the-art performance on the cross-domain image patch matching. Furthermore, we improved H-Net to H-Net++. The H-Net++ extracts invariant feature descriptors in cross-domain image patches and achieves state-of-the-art performance by feature retrieval in Euclidean space. As there is no benchmark dataset including cross-domain images, we made a cross-domain image dataset which consists of camera images, rendering images from UAV 3D model, and images generated by CycleGAN algorithm. Experiments show that the proposed H-Net and H-Net++ outperform the existing algorithms. Our code and cross-domain image dataset are available at https://github.com/Xylon-Sean/H-Net. Weiquan Liu, Xuelun Shen, Cheng Wang 0003, Zhihong Zhang 0001, Chenglu Wen, Jonathan Li 0001 |
IJCAI | 5 |
| 2018 | Line Structure-Based Indoor and Outdoor Integration Using Backpacked and TLS Point Cloud DataabstractThis letter presents a line structure-based method for integration of centimeter-level indoor backpacked scanning point clouds and millimeter-level outdoor terrestrial laser scanning point clouds. Using 3-D lines for registration, instead of matching points directly, can improve the robustness of the method and adapt to multisource point cloud data of different qualities. Considering the limited overlapping between indoor and outdoor scenes, line structures are extracted from overlapped wall areas that may be included in interior and exterior data. Here, a patch-based method labels a point cloud into wall, ceiling, floor categories, as well as assigning the candidate overlapping walls. Then, lines structures are extracted from the wall plane point cloud. Potential door and window line structures are detected and refined for point cloud registration. Last, an iterative closest point-based method is used to fine tune the registration results. Our results show that the proposed method effectively integrates a promising map of indoor and outdoor scenes. Chenglu Wen, Xiaotian Sun 0005, Shiwei Hou, Jinbin Tan, Yudi Dai, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Semantic Labeling of Mobile LiDAR Point Clouds via Active Learning and Higher Order MRFabstractUsing mobile Light Detection and Ranging point clouds to accomplish road scene labeling tasks shows promise for a variety of applications. Most existing methods for semantic labeling of point clouds require a huge number of fully supervised point cloud scenes, where each point needs to be manually annotated with a specific category. Manually annotating each point in point cloud scenes is labor intensive and hinders practical usage of those methods. To alleviate such a huge burden of manual annotation, in this paper, we introduce an active learning method that avoids annotating the whole point cloud scenes by iteratively annotating a small portion of unlabeled supervoxels and creating a minimal manually annotated training set. In order to avoid the biased sampling existing in traditional active learning methods, a neighbor-consistency prior is exploited to select the potentially misclassified samples into the training set to improve the accuracy of the statistical model. Furthermore, lots of methods only consider short-range contextual information to conduct semantic labeling tasks, but ignore the long-range contexts among local variables. In this paper, we use a higher order Markov random field model to take into account more contexts for refining the labeling results, despite of lacking fully supervised scenes. Evaluations on three data sets show that our proposed framework achieves a high accuracy in labeling point clouds although only a small portion of labels is provided. Moreover, comparative experiments demonstrate that our proposed framework is superior to traditional sampling methods and exhibits comparable performance to those fully supervised models. Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Ziyi Chen 0001, Dawei Zai, Yongtao Yu, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Quality evaluation of point cloud model for interior structure of a common buildingabstractThis paper presents a standardized quality criteria to evaluate the 3D point cloud model of the indoor building which is based on point cloud's data accuracy, the prior characteristics of the building and the coincidence errors of the point cloud model. Our assessment framework involves three steps: the point cloud data acquisition, model generation and quality evaluation. In model generation progress, incapacity of scanning the whole building information one time since its multi-storied spatial structure, the building model need to be registered and merged. In evaluation step, taking into account the need of mapping, indoor location and navigation, the establishment of an interior spatial building model requires accurate measurement. Therefore, we adopt data noise analysis to give a judgement. Then, since the geometric characteristics of the building model are varying, the geometric analysis is proposed to evaluate acquisition errors and registration error. Comparative experiments demonstrate our method give integrate, realistic and reliable quality framework for the indoor building point cloud model. Yiping Chen 0002, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001, JinYong Chen |
IGARSS | 3 |
| 2017 | Rapid traffic sign damage inspection in natural scenes using mobile laser scanning dataabstractThis paper proposes a novel approach for traffic sign detection and rapid damage inspection in natural scenes based on mobile laser scanning (MLS) data, including images and point clouds. The inspection results assist traffic management departments to take immediate measures to update and maintain traffic signs after natural disasters leading to many damaged traffic signs. Our approach involves four steps: Firstly, we use a deep learning network, Fast regions with convolutional neural network (Fast R-CNN), to train a traffic sign detector in an open benchmark, where the images are more variable and have a higher resolution. Then, traffic signs in images are detected by using the trained detector. Next, the area of the traffic sign, based on the sign area in the image, is roughly detected in MLS point clouds. Then, an accurate traffic sign is detected. Finally, some placement parameters of the traffic sign are measured for damage inspection and further inventory. Our proposed approach is validated on a set of point-clouds acquired by a RIEGL VMX-450 MLS system. Experimental results demonstrate that the rapidity and reliability of our proposed approach in traffic sign detection and damage inspection are robust. Changbin You, Chenglu Wen, Huan Luo 0001, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2017 | Rapid Localization and Extraction of Street Light Poles in Mobile LiDAR Point Clouds: A Supervoxel-Based ApproachabstractThis paper presents a supervoxel-based approach for automated localization and extraction of street light poles in point clouds acquired by a mobile LiDAR system. The method consists of five steps: preprocessing, localization, segmentation, feature extraction, and classification. First, the raw point clouds are divided into segments along the trajectory, the ground points are removed, and the remaining points are segmented into supervoxels. Then, a robust localization method is proposed to accurately identify the pole-like objects. Next, a localization-guided segmentation method is proposed to obtain pole-like objects. Subsequently, the pole features are classified using the support vector machine and random forests. The proposed approach was evaluated on three datasets with 1,055 street light poles and 701 million points. Experimental results show that our localization method achieved an average recall value of 98.8%. A comparative study proved that our method is more robust and efficient than other existing methods for localization and extraction of street light poles. Chenglu Wen, Yulan Guo, Yongtao Yu, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Automated segmentation of LiDAR point clouds for building rooftop extractionabstractLIDAR (Light Detection and Ranging) is an especially effective tool for acquiring geo-referenced point clouds of urban site. Accurate extraction of elevated features such as building rooftops is vitally important in various applications. However, it is still challenging to determine an accurate rooftop contour from the irregularly distributed LIDAR point clouds. In this paper an efficient LIDAR segmentation method is presented in order to achieve automated rooftop extraction. First, we apply a voxel-based upward growing algorithm that filters out the ground points from the raw point cloud scenes. Second, we employ a Euclidean based clustering method on non-ground points by making use of nearest neighbors. Then we introduce RANSAC (RANdom SAmple Consensus) technique to estimate primitive planes for fitting rooftop facets. Finally, we use concave hull and L0regularization to determine the rooftop contour. Accurate experimental results demonstrate the validity of our segmentation method for rooftop extraction. Cheng Wang 0003, Jonathan Li 0001, Zongliang Zhang, Dawei Zai, Pengdi Huang, Chenglu Wen |
IGARSS | 7 |
| 2016 | Local quality assessment of point clouds for indoor mobile mapping
Fangfang Huang, Chenglu Wen, Huan Luo 0001, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001 |
Neurocomputing | 2 |
| 2016 | An Indoor Backpack System for 2-D and 3-D Mapping of Building InteriorsabstractThis letter presents a backpack mapping system for creating indoor 2-D and 3-D maps of building interiors. For many applications, indoor mobile mapping provides a 3-D structure via an indoor map. Because there are significant roll and pitch motions of the indoor mobile mapping system, the need arises for a moving mobile system with 6 degrees of freedom (DOFs) (x, y, and z positions and roll, yaw, and pitch angles). First, we present a 6-DOF pose estimation algorithm by fusing 2-D laser scanner data with inertial sensor data using an extended Kalman filter-based method. The estimated 6-DOF pose is used as the initialized transformation for consecutive map alignment in 3-D map building. The 6-DOF pose gives a full 3-D estimation of the system pose and is used to accelerate the map alignment process and also align the two maps directly when there are few or no overlapping areas between the maps. Our results show that the proposed system effectively builds a consistent 2-D grid map and a 3-D point cloud map of an indoor environment. Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Vehicle Detection in High-Resolution Aerial Images via Sparse Representation and SuperpixelsabstractThis paper presents a study of vehicle detection from high-resolution aerial images. In this paper, a superpixel segmentation method designed for aerial images is proposed to control the segmentation with a low breakage rate. To make the training and detection more efficient, we extract meaningful patches based on the centers of the segmented superpixels. After the segmentation, through a training sample selection iteration strategy that is based on the sparse representation, we obtain a complete and small training subset from the original entire training set. With the selected training subset, we obtain a dictionary with high discrimination ability for vehicle detection. During training and detection, the grids of histogram of oriented gradient descriptor are used for feature extraction. To further improve the training and detection efficiency, a method is proposed for the defined main direction estimation of each patch. By rotating each patch to its main direction, we give the patches consistent directions. Comprehensive analyses and comparisons on two data sets illustrate the satisfactory performance of the proposed algorithm. Ziyi Chen 0001, Cheng Wang 0003, Chenglu Wen, Xiuhua Teng, Yiping Chen 0002, Haiyan Guan, Huan Luo 0001, Liujuan Cao, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Vehicle Detection in High-Resolution Aerial Images Based on Fast Sparse Representation Classification and Multiorder FeatureabstractThis paper presents an algorithm for vehicle detection in high-resolution aerial images through a fast sparse representation classification method and a multiorder feature descriptor that contains information of texture, color, and high-order context. To speed up computation of sparse representation, a set of small dictionaries, instead of a large dictionary containing all training items, is used for classification. To extract the context information of a patch, we proposed a high-order context information extraction method based on the proposed fast sparse representation classification method. To effectively extract the color information, the RGB color space is transformed into color name space. Then, the color name information is embedded into the grids of histogram of oriented gradient feature to represent the low-order feature of vehicles. By combining low- and high-order features together, a multiorder feature is used to describe vehicles. We also proposed a sample selection strategy based on our fast sparse representation classification method to construct a complete training subset. Finally, a set of dictionaries, which are trained by the multiorder features of the selected training subset, is used to detect vehicles based on superpixel segmentation results of aerial images. Experimental results illustrate the satisfactory performance of our algorithm. Ziyi Chen 0001, Cheng Wang 0003, Huan Luo 0001, Hanyun Wang, Yiping Chen 0002, Chenglu Wen, Yongtao Yu, Liujuan Cao, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2016 | Patch-Based Semantic Labeling of Road Scene Using Colorized Mobile LiDAR Point CloudsabstractSemantic labeling of road scenes using colorized mobile LiDAR point clouds is of great significance in a variety of applications, particularly intelligent transportation systems. However, many challenges, such as incompleteness of objects caused by occlusion, overlapping between neighboring objects, interclass local similarities, and computational burden brought by a huge number of points, make it an ongoing open research area. In this paper, we propose a novel patch-based framework for labeling road scenes of colorized mobile LiDAR point clouds. In the proposed framework, first, three-dimensional (3-D) patches extracted from point clouds are used to construct a 3-D patch-based match graph structure (3D-PMG), which transfers category labels from labeled to unlabeled point cloud road scenes efficiently. Then, to rectify the transferring errors caused by local patch similarities in different categories, contextual information among 3-D patches is exploited by combining 3D-PMG with Markov random fields. In the experiments, the proposed framework is validated on colorized mobile LiDAR point clouds acquired by the RIEGL VMX-450 mobile LiDAR system. Comparative experiments show the superior performance of the proposed framework for accurate semantic labeling of road scenes. Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Zhipeng Cai 0003, Ziyi Chen 0001, Hanyun Wang, Yongtao Yu, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Spatial-Related Traffic Sign Inspection for Inventory Purposes Using Mobile Laser Scanning DataabstractThis paper presents a spatial-related traffic sign inspection process for sign type, position, and placement using mobile laser scanning (MLS) data acquired by a RIEGL VMX-450 system and presents its potential for traffic sign inventory applications. First, the paper describes an algorithm for traffic sign detection in complicated road scenes based on the retroreflectivity properties of traffic signs in MLS point clouds. Then, a point cloud-to-image registration process is proposed to project the traffic sign point clouds onto a 2-D image plane. Third, based on the extracted traffic sign points, we propose a traffic sign position and placement inspection process by creating geospatial relations between the traffic signs and road environment. For further inventory applications, we acquire several spatial-related inventory measurements. Finally, a traffic sign recognition process is conducted to assign sign type. With the acquired sign type, position, and placement data, a spatial-associated sign network is built. Experimental results indicate satisfactory performance of the proposed detection, recognition, position, and placement inspection algorithms. The experimental results also prove the potential of MLS data for automatic traffic sign inventory applications. Chenglu Wen, Jonathan Li 0001, Huan Luo 0001, Yongtao Yu, Zhipeng Cai 0003, Hanyun Wang, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Bag of Contextual-Visual Words for Road Scene Object Detection From Mobile Laser Scanning DataabstractThis paper proposes a novel algorithm for detecting road scene objects (e.g., light poles, traffic signposts, and cars) from 3-D mobile-laser-scanning point cloud data for transportation-related applications. To describe local abstract features of point cloud objects, a contextual visual vocabulary is generated by integrating spatial contextual information of feature regions. Objects of interest are detected based on the similarity measures of the bag of contextual-visual words between the query object and the segmented semantic objects. Quantitative evaluations on two selected data sets show that the proposed algorithm achieves an average recall, precision, quality, and F-score of 0.949, 0.970, 0.922, and 0.959, respectively, in detecting light poles, traffic signposts, and cars. Comparative studies demonstrate the superior performance of the proposed algorithm over other existing methods. Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003, Chenglu Wen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2015 | Using mobile LiDAR point clouds for traffic sign detection and sign visibility estimationabstractThis paper presents a novel method for traffic sign detection and visibility evaluation from mobile Light Detection and Ranging (LiDAR) point clouds and the corresponding images. Our algorithm involves two steps. Firstly, a detection algorithm based on high retro-reflectivity of the traffic sign from the MLS point clouds is designed for sign detection in complicated road scenes. To solve the spatial features of traffic signs, we also create geo-referenced relations between traffic signs and roads according to the normal of ground. Secondly, we propose a visibility estimation method to evaluate the visibility level of the traffic sign based on a combination of visual appearance and spatial-related features. The proposed algorithm is validated on a set of transportation-related point-clouds acquired by a RIEGL VMX-450 LiDAR system. The experiment results demonstrate that the efficiency and reliability of the proposed algorithm in detection traffic signs are robust, and also prove the potential of using mobile LiDAR data for traffic sign visibility evaluation. Chenglu Wen, Huan Luo 0001, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2015 | Occluded Boundary Detection for Small-Footprint Groundborne LIDAR Point Cloud Guided by Last EchoabstractOccluded boundary detection in a 3-D point cloud is an indispensable preprocessing step for many applications, such as point cloud completion. Meanwhile, existing methods do not have the ability of distinguishing occluded and complete surface borders, such as the border of a sign board. To solve this problem, this letter presents an occluded boundary detection method for small-footprint LIDAR point clouds. The main novelty of this letter is using the last-echo information for occluded boundary detection. Seed boundary (SB) points are subsequently detected using this last-echo information. Finally, the SB points are grown into neighboring points using an occluded boundary growth algorithm. To the best of our knowledge, this method is the first method that uses the last-echo information to detect occluded boundaries. Experimental results with comparisons indicate that the proposed method can accurately and efficiently detect an occluded boundary without contamination from a complete surface border. These advantages allow the proposed method to benefit further applications such as point cloud completion, as demonstrated in the application section. Zhipeng Cai 0003, Cheng Wang 0003, Chenglu Wen, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Road Boundaries Detection Based on Local Normal Saliency From Mobile Laser Scanning DataabstractThe accurate extraction of roads is a prerequisite for the automatic extraction of other road features. This letter describes a method for detecting road boundaries from mobile laser scanning (MLS) point clouds in an urban environment. The key idea of our method is directly constructing a saliency map on 3-D unorganized point clouds to extract road boundaries. The method consists of four major steps, i.e., road partition with the assistance of the vehicle trajectory, salient map construction and salient points extraction, curb detection and curb lowest points extraction, and road boundaries fitting. The performance of the proposed method is evaluated on the point clouds of an urban scene collected by a RIEGL VMX-450 MLS system. The completeness, correctness, and quality of the extracted road boundaries are 95.41%, 99.35%, and 94.81%, respectively. Experimental results demonstrate that our method is feasible for detecting road boundaries in MLS point clouds. Hanyun Wang, Huan Luo 0001, Chenglu Wen, Jun Cheng 0002, Peng Li 0064, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2014 | Automatic extraction of power lines from mobile laser scanning dataabstractThis paper presents a stepwise algorithm for extracting power-lines from mobile laser scanning (MLS) data. This algorithm first extracts non-road points from MLS data by estimating road ranges with regard to scanning mechanism and applying elevation-difference and slope criteria to the road ranges scan-line by scan-line. Then, three filters, in terms of height, spatial density, and size-and-shape, are proposed to extract power-line points in the identified non-road points, followed by Hough transform and Euclidean distance clustering. Finally, a 3D power line is modelled as a horizontal line in X-Y plane and a vertical catenary curve defined by a hyperbolic cosine function in X-Z plane. The proposed algorithm has been tested on a sample of point clouds acquired by a RIEGL VMX-450 MLS system. The results demonstrate the applicability of the proposed algorithm in extracting power transmission lines. Haiyan Guan, Jonathan Li 0001, Yongjun Zhou, Yongtao Yu, Cheng Wang 0003, Chenglu Wen |
IGARSS | 6 |
| 2014 | Object Detection in Terrestrial Laser Scanning Point Clouds Based on Hough ForestabstractThis letter presents a novel rotation-invariant method for object detection from terrestrial 3-D laser scanning point clouds acquired in complex urban environments. We utilize the Implicit Shape Model to describe object categories, and extend the Hough Forest framework for object detection in 3-D point clouds. A 3-D local patch is described by structure and reflectance features and then mapped to the probabilistic vote about the possible location of the object center. Objects are detected at the peak points in the 3-D Hough voting space. To deal with the arbitrary azimuths of objects in real world, circular voting strategy is introduced by rotating the offset vector. To deal with the interference of adjacent objects, distance weighted voting is proposed. Large-scale real-world point cloud data collected by terrestrial mobile laser scanning systems are used to evaluate the performance. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 3-D object detection methods. Hanyun Wang, Cheng Wang 0003, Huan Luo 0001, Peng Li 0064, Ming Cheng 0002, Chenglu Wen, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2014 | Three-Dimensional Indoor Mobile Mapping With Fusion of Two-Dimensional Laser Scanner and RGB-D Camera DataabstractThree-dimensional mobile mapping in indoor environment, mostly global navigation satellite system-denied space, is to consecutively align the frames to build a global 3-D map of an indoor environment. One of the major difficulties of the current solutions is the failure at the insufficient overlapping between the frames, which is the reality of a lack of correspondences between the frames. To overcome this problem, a 3-D indoor mobile mapping system that integrates a 2-D laser scanner, and an RGB-Depth camera is presented in this letter. In this system, a fusion-iterative closest point (ICP) method, which combines the 2-D mobile platform pose from a Rao-Blackwellized particle filter estimation, an ICP, and a generalized-ICP method, is proposed for the consecutive frame alignment. Fusion-ICP achieves effective frame alignment, particularly in solving the insufficient overlapping frame alignment problem. Comparative experiments were conducted to evaluate the mapping system. The experimental results demonstrate the effectiveness and efficiency of our system for 3-D indoor mobile mapping. Chenglu Wen, Qingyuan Zhu, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | Hyperspectral image classification using Primal Laplacian SVM in preconditioned conjugate gradient solutionabstractWith the introduction of manifold assumption, Laplacian Support Vector Machine (LapSVM) has advantages over the traditional SVM classifiers. However the dual solution of LapSVM is still a major barrier on the further application of LapSVM. Primal optimization is a promising solution to this problem. In this paper, we introduce a novel primal Laplacian Support Vector Machine with Precondition Conjugate Gradient method (PCG) to the problem of hyperspectral images classification which is one type of primal optimization solution. To prove the effectiveness of the proposed method, we apply it into the hyperspectral image data set Indian Pine. The experiment results show higher accuracy and better generalization ability than dual strategy. Xiaoli Ma, Cheng Wang 0003, Chenglu Wen, Jonathan Li 0001 |
IGARSS | 4 |