Qiming Xia

dblp:276/3620 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-6440-5694ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning
abstract
Unsupervised 3D object detection leverages heuristic algorithms to discover potential objects, offering a promising route to reduce annotation costs in autonomous driving. Existing approaches mainly generate pseudo labels and refine them through self-training iterations. However, these pseudo-labels are often incorrect at the beginning of training, resulting in misleading the optimization process. Moreover, effectively filtering and refining them remains a critical challenge. In this paper, we propose $\textbf{OWL}$ for unsupervised 3D object detection by occupancy guided warm-up and large-model priors reasoning. OWL first employs an Occupancy Guided Warm-up (OGW) strategy to initialize the backbone weight with spatial perception capabilities, mitigating the interference of incorrect pseudo-labels on network convergence. Furthermore, OWL introduces an Instance-Cued Reasoning (ICR) module that leverages the prior knowledge of large models to assess pseudo-label quality, enabling precise filtering and refinement. Finally, we design a WAS (Weight-adapted Self-training) strategy to dynamically re-weight pseudo-labels, improving the performance through self-training. Extensive experiments on Waymo Open Dataset (WOD) and KITTI demonstrate that OWL outperforms state-of-the-art unsupervised methods by over 15.0\% mAP, revealing the effectiveness of our method.
Xusheng Guo, Wanfa Zhang, Shijia Zhao, Qiming Xia, Xiaolong Xie, Chenglu Wen
AAAI4
2026 V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR Localization
abstract
Multi-agents rely on accurate poses to share and align observations, enabling a collaborative perception of the environment. However, traditional GNSS-based localization often fails in GNSS-denied environments, making consistent feature alignment difficult in collaboration. To tackle this challenge, we propose a robust GNSS-free collaborative perception framework based on LiDAR localization. Specifically, we propose a lightweight Pose Generator with Confidence (PGC) to estimate compact pose and confidence representations. To alleviate the effects of localization errors, we further develop the Pose-Aware Spatio-Temporal Alignment Transformer (PASTAT), which performs confidence-aware spatial alignment while capturing essential temporal context. Additionally, we present a new simulation dataset, V2VLoc, which can be adapted for both LiDAR localization and collaborative detection tasks. V2VLoc comprises three subsets: Town1Loc, Town4Loc, and V2VDet. Town1Loc and Town4Loc offer multi-traversal sequences for training in localization tasks, whereas V2VDet is specifically intended for the collaborative detection task. Extensive experiments conducted on the V2VLoc dataset demonstrate that our approach achieves state-of-the-art performance under GNSS-denied conditions. We further conduct extended experiments on the real-world V2V4Real dataset to validate the effectiveness and generalizability of PASTAT.
Wenkai Lin, Qiming Xia, Wen Li 0005, Xun Huang 0003, Chenglu Wen
AAAI2
2026 WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object Detection
abstract
3D object detection is critical for autonomous driving, yet it remains fundamentally challenging to simultaneously maximize computational efficiency and capture long-range spatial dependencies.We observed that Mamba-based models, with their linear state-space design, capture long-range dependencies at lower cost, offering a promising balance between efficiency and accuracy.However, existing methods rely on axis-aligned scanning within a fixed window, inevitably discarding spatial information. To address this problem, we propose WinMamba, a novel Mamba-based 3D feature-encoding backbone composed of stacked WinMamba blocks. To enhance the backbone with robust multi-scale representation, the WinMamba block incorporates a window-scale-adaptive module that compensates voxel features across varying resolutions during sampling. Meanwhile, to obtain rich contextual cues within the linear state space, we equip the WinMamba layer with a learnable positional encoding and a window-shift strategy.Extensive experiments on the KITTI and Waymo datasets demonstrate that WinMamba significantly outperforms the baseline. Ablation studies further validate the individual contributions of the WSF and AWF modules in improving detection accuracy. The code will be made publicly available.
Longhui Zheng, Qiming Xia, Xiaolu Chen, Zhaoliang Liu, Chenglu Wen
AAAI2
2026 DOtA++: Unsupervisely and Collaboratively Detect Objects From Multi-Agent Observations With Multi-Modal Prior Constraints
abstract
Enhancing perception performance via multi-agent collaboration has gained increasing attention in the field of autonomous driving. However, as the number of agents grows, the manual annotation required for training collaborative detectors increases significantly. To tackle this problem, we introduce an unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, named DOtA, without using labels from external. DOtA first generates preliminary labels by an initial detector, which is trained by internally shared information of collaborative agents. DOtA then optimizes these preliminary labels by utilizing the physical rule constraints derived from the surrounding area of the object. Building on DOtA, we further propose DOtA++, an enhanced version that improves performance by leveraging composite prior constraints. Beyond physical rule constraints, DOtA++ further uses image data as an auxiliary modality to introduce multi-agent observation consistency constraints, boosting object classification, while also incorporating point cloud geometric distribution constraints to improve structural description. Extensive experiments on widely-used benchmarks demonstrate that DOtA and DOtA++ effectively perceive potential objects in the scene without manual annotations. In particular, DOtA++ shows 10.7% mAP improvement over traditional unsupervised methods on V2X-R dataset.
Qiming Xia, Longhui Zheng, Shijia Zhao, Xun Huang 0003, Chenglu Wen, Cheng Wang 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 DGFusion: Dual-Guided Fusion for Robust Multi-Modal 3D Object Detection
abstract
As a critical task in autonomous driving perception systems, 3D object detection is used to identify and track key objects, such as vehicles and pedestrians. However, detecting distant, small, or occluded objects (hard instances) remains a challenge, which directly compromises the safety of autonomous driving systems. We observe that existing multi-modal 3D object detection methods often follow a single-guided paradigm, failing to account for the differences in information density of hard instances between modalities. In this work, we propose DGFusion, based on the Dual-guided paradigm, which fully inherits the advantages of the Point-guide-Image paradigm and integrates the Image-guide-Point paradigm to address the limitations of the single paradigms. The core of DGFusion, the Difficulty-aware Instance Pair Matcher (DIPM), performs instance-level feature matching based on difficulty to generate easy and hard instance pairs, while the Dual-guided Modules exploit the advantages of both pair types to enable effective multi-modal feature fusion. Experimental results demonstrate that our DGFusion outperforms the baseline methods, with respective improvements of +1.0% mAP, +0.8% NDS, and +1.3% average recall on nuScenes. Extensive experiments demonstrate consistent robustness gains for hard instance detection across ego-distance, size, visibility, and small-scale training scenarios.
Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lei Yang 0060, Ziying Song
IEEE Trans. Circuits Syst. Video Technol.5
2025 L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object Detection
abstract
LiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, it faces challenges of significant differences in terms of data quality and the degree of degradation in adverse weather. To address these issues, we introduce L4DR, a weather-robust 3D object detection method that effectively achieves LiDAR and 4D Radar fusion. Our L4DR proposes Multi-Modal Encoding (MME) and Foreground-Aware Denoising (FAD) modules to reconcile sensor gaps, which is the first exploration of the complementarity of early fusion between LiDAR and 4D radar. Additionally, we design an Inter-Modal and IntraModal ({IM}2) parallel feature extraction backbone coupled with a Multi-Scale Gated Fusion (MSGF) module to counteract the varying degrees of sensor degradation under adverse weather conditions. Experimental evaluation on a VoD dataset with simulated fog proves that L4DR is more adaptable to changing weather conditions. It delivers a significant performance increase under different fog levels, improving the 3D mAP by up to 20.0% over the traditional LiDAR-only approach. Moreover, the results on the K-Radar dataset validate the consistent performance improvement of L4DR in realworld adverse weather conditions.
Xun Huang 0003, Ziyu Xu 0002, Qiming Xia, Yan Xia 0003, Jonathan Li 0001, Kyle Gao, Chenglu Wen, Cheng Wang 0003
AAAI5
2025 Seg2Box: 3D Object Detection by Point-Wise Semantics Supervision
abstract
LIDAR-based 3D object detection and semantic segmentation are critical tasks in 3D scene understanding. Traditional detection and segmentation methods supervise their models through bounding box labels and semantic mask labels. However, these two independent labels inherently contain significant redundancy. This paper aims to eliminate the redundancy by supervising 3D object detection using only semantic labels. However, the challenge arises due to the incomplete geometry structure and boundary ambiguity of point cloud instances, leading to inaccurate pseudo-labels and poor detection results. To address these challenges, we propose a novel method, named Seg2Box. We first introduce a Multi-Frame Multi-Scale Clustering (MFMS-C) module, which leverages the spatio-temporal consistency of point clouds to generate accurate box-level pseudo-labels. Additionally, the Semantic-Guiding Iterative-Mining Self-Training (SGIM-ST) module is proposed to enhance the performance by progressively refining the pseudo-labels and mining the instances without generating pseudo-labels. Experiments on the Waymo Open Dataset and nuScenes Dataset show that our method significantly outperforms other competitive methods by 23.7% and 10.3% in mAP, respectively. The results demonstrate the great label-efficient potential and advancement of our method.
Maoji Zheng, Ziyu Xu 0002, Qiming Xia, Chenglu Wen, Cheng Wang 0003
AAAI3
2025 AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label Correction
abstract
Recently, Visual Foundation Models (VFMs) have shown a remarkable generalization performance in 3D perception tasks. However, their effectiveness in large-scale outdoor datasets remains constrained by the scarcity of accurate supervision signals, the extensive noise caused by variable outdoor conditions, and the abundance of unknown objects. In this work, we propose a novel label-free learning method, Adaptive Label Correction (AdaCo), for 3D semantic segmentation. AdaCo first introduces the Cross-modal Label Generation Module (CLGM), providing cross-modal supervision with the formidable interpretive capabilities of the VFMs. Subsequently, AdaCo incorporates the Adaptive Noise Corrector (ANC), updating and adjusting the noisy samples within this supervision iteratively during training. Moreover, we develop an Adaptive Robust Loss (ARL) function to modulate each sample's sensitivity to noisy supervision, preventing potential underfitting issues associated with robust loss. Our proposed AdaCo can effectively mitigate the performance limitations of label-free learning networks in 3D semantic segmentation tasks. Extensive experiments on two outdoor benchmark datasets highlight the superior performance of our method.
Pufan Zou, Shijia Zhao, Qiming Xia, Chenglu Wen, Cheng Wang 0003
AAAI4
2025 V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object Detection
abstract
Current Vehicle-to-Everything (V2X) systems have significantly enhanced 3D object detection using LiDAR and camera data. However, they face performance degradation in adverse weather. Weather-robust 4D radar, with Doppler velocity and additional geometric information, offers a promising solution to this challenge. To this end, we present V2X-R, the first simulated V2X dataset incorporating LiDAR, camera, and 4D radar modalities. V2X-R contains 12,079 scenarios with 37,727 frames of LiDAR and 4D radar point clouds, 150,908 images, and 170,859 annotated 3D vehicle bounding boxes. Subsequently, we propose a novel cooperative LiDAR-4D radar fusion pipeline for 3D object detection and implement it with multiple fusion strategies. To achieve weather-robust detection, we additionally propose a Multi-modal Denoising Diffusion (MDD) module in our fusion pipeline. MDD utilizes weather-robust 4D radar feature as a condition to guide the diffusion model in denoising noisy LiDAR features. Experiments show that our LiDAR-4D radar fusion pipeline demonstrates superior performance in the V2X-R dataset. Over and above this, our MDD module further improved the foggy/snowy performance of the basic fusion model by up to 5.73%/6.70% and barely disrupting normal performance. The dataset and code will be publicly available at: https://github.com/ylwhxht/V2X-R.
Xun Huang 0003, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li 0003, Cheng Wang 0003, Chenglu Wen
CVPR3
2025 Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual Labels
abstract
Unsupervised 3D object detection serves as an important solution for offline 3D object annotation. However, due to the data sparsity and limited views, the clustering-based label fitting in unsupervised object detection often generates low-quality pseudo-labels. Multi-agent collaborative dataset, which involves the sharing of complementary observations among agents, holds the potential to break through this bottleneck. In this paper, we introduce a novel unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, termed DOtA, without using labels from external. DOtA first uses the internally shared ego-pose and ego-shape of collaborative agents to initialize the detector, leveraging the generalization performance of neural networks to infer preliminary labels. Subsequently, DOtA uses the complementary observations between agents to perform multi-scale encoding on preliminary labels, then decodes high-quality and low-quality labels. These labels are further used as prompts to guide a correct feature learning process, thereby enhancing the performance of the unsupervised object detection task. Extensive experiments on the V2V4Real and OPV2V datasets show that our DOtA outperforms state-of-the-art unsupervised 3D object detection methods. Additionally, we also validate the effectiveness of the DOtA labels under various collaborative perception frameworks. The code is available at https://github.com/xmuqimingxia/DOtA.
Qiming Xia, Wenkai Lin, Haoen Xiang, Xun Huang 0003, Siheng Chen, Zhen Dong 0005, Cheng Wang 0003, Chenglu Wen
CVPR1
2025 SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts
abstract
Recently, sparsely-supervised 3D object detection has gained great attention, achieving performance close to fully-supervised 3D detectors while requiring only a few annotated instances. Nevertheless, these methods suffer challenges when accurate labels are extremely absent. In this paper, we propose a boosting strategy, termed SP3D, explicitly utilizing the cross-modal semantic prompts generated from Large Multimodal Models (LMMs) to boost the 3D detector with robust feature discrimination capability under sparse annotation settings. Specifically, we first develop a Confident Points Semantic Transfer (CPST) module that generates accurate cross-modal semantic prompts through boundary-constrained center cluster selection. Based on these accurate semantic prompts, which we treat as seed points, we introduce a Dynamic Cluster Pseudo-label Generation (DCPG) module to yield pseudo-supervision signals from the geometry shape of multi-scale neighbor points. Additionally, we design a Distribution Shape score (DS score) that chooses high-quality supervision signals for the initial training of the 3D detector. Experiments on the KITTI dataset and Waymo Open Dataset (WOD) have validated that SP3D can enhance the performance of sparsely supervised detectors by a large margin under meager labeling conditions. Moreover, we verified SP3D in the zero-shot setting, where its performance exceeded that of the state-of-the-art methods. The code is available at https://github.com/xmuqimingxia/SP3D.
Shijia Zhao, Qiming Xia, Xusheng Guo, Pufan Zou, Maoji Zheng, Chenglu Wen, Cheng Wang 0003
CVPR2
2025 Pretend Benign: A Stealthy Adversarial Attack by Exploiting Vulnerabilities in Cooperative Perception
Dongyu Pan, Qiming Xia, Cheng Wang 0003, Chenglu Wen
ICCV3
2025 Unsupervised 3D Object Detection by Commonsense Clue
abstract
Traditional 3D object detectors, whether fully-, semi-, or weakly-supervised, rely heavily on extensive human annotations. In contrast, this paper introduces an unsupervised 3D object detector that automatically discerns object patterns without such annotations. To achieve this, we propose a Commonsense Prototype-based Detector (CPD) for unsupervised 3D object detection. CPD first constructs Commonsense Prototypes (CProto) to represent the geometric center and size of objects. It then generates high-quality pseudo-labels and guides detector convergence using size and geometry priors from CProto. Building on CPD, we further introduce CPD++, an enhanced version that improves performance by leveraging motion cues. CPD++ learns localization from stationary objects and recognition from moving objects, facilitating the mutual transfer of localization and recognition knowledge between these two object types. Both CPD and CPD++ outperform existing state-of-the-art unsupervised 3D detectors. Furthermore, when trained on Waymo Open Dataset (WOD) and tested on KITTI, CPD++ achieves 89.25% 3D Average Precision (AP) on the moderate car class at a 0.5 IoU threshold, reaching 95.3% of the performance attained by fully supervised counterparts. These results underscore the significant advancements brought by our method.
Shijia Zhao, Xun Huang 0003, Qiming Xia, Chenglu Wen, Li Jiang 0009, Xin Li 0003, Cheng Wang 0003
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Fade3D: Fast and Deployable 3D Object Detection for Autonomous Driving
abstract
3D object detection is an essential scene perception capability for autonomous vehicles. In intelligent transportation systems, autonomous vehicles require minimal inference latency to sense their surroundings in real-time. However, advanced 3D detection methods often suffer from high inference latency. This limits the real-time deployment of 3D detection models in the real world. To address this problem, this paper proposes a fast and deployable 3D object detection method from the LiDAR point cloud for autonomous driving, namedFade3D. Firstly, we propose a Lightweight Input Encoder (LIE) to extract the most critical features from point clouds. Then, we develop a Spatial Feature Enhancement BEV backbone (SFENet) that efficiently encodes geometry features into compact representations. Additionally, we design an IoU-aware Loss Re-weighting (ILR) that enhances performance by shifting more attention to hard samples. Leveraging LIE and SFENet, our approach is independent of point cloud density and number, achieving significant speed advantages in processing large-scale point clouds and being deployment-friendly. Extensive experiments on KITTI and Waymo Open Dataset (WOD) datasets comparing various baseline detectors demonstrate its universality and superiority. Specifically, our method demonstrates impressive real-time inference capabilities, achieving 51.5 Hz on an RTX3090 GPU and 12.4 Hz on a Jetson Orin embedded development board. Code will be available at https://github.com/wayyeah/Fade3D
Qiming Xia, Zhen Dong 0005, Ruofei Zhong, Cheng Wang 0003, Chenglu Wen
IEEE Trans. Intell. Transp. Syst.2
2024 HINTED: Hard Instance Enhanced Detector with Mixed-Density Feature Fusion for Sparsely-Supervised 3D Object Detection
abstract
Current sparsely-supervised object detection methods largely depend on high threshold settings to derive high-quality pseudo labels from detector predictions. However, hard instances within point clouds frequently display incomplete structures, causing decreased confidence scores in their assigned pseudo-labels. Previous methods inevitably result in inadequate positive supervision for these instances. To address this problem, we propose a novel Hard INsTance Enhanced Detector (HINTED), for sparsely-supervised 3D object detection. Firstly, we design a self-boosting teacher (SBT) model to generate more potential pseudo-labels, enhancing the effectiveness of information transfer. Then, we introduce a mixed-density student (MDS) model to concentrate on hard instances during the training phase, thereby improving detection accuracy. Our extensive experiments on the KITTI dataset validate our method's superior performance. Compared with leading sparsely-supervised methods, HINTED significantly improves the detection performance on hard instances, no-tably outperforming fully-supervised methods in detecting challenging categories like cyclists. HINTED also significantly outperforms the state-of-the-art semi-supervised method on challenging categories. The code is available at https://github.com/xmuqimingxia/HINTED.
Qiming Xia, Shijia Zhao, Leyuan Xing, Xun Huang 0003, Jinhao Deng, Xin Li 0003, Chenglu Wen, Cheng Wang 0003
CVPR1
2024 CMD: A Cross Mechanism Domain Adaptation Dataset for 3D Object Detection
Jinhao Deng, Xun Huang 0003, Qiming Xia, Xin Li 0003, Wei Li 0111, Chenglu Wen, Cheng Wang 0003
ECCV (57)5
2024 SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object Detection
abstract
LiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12%$\text {AP}_{\text {3-D}}$on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released inhttps://github.com/liulin813/SparseDet.git.
Ziying Song, Qiming Xia, Feiyang Jia, Caiyan Jia, Lei Yang 0060, Hongyu Pan
IEEE Trans. Geosci. Remote. Sens.3
2023 CoIn: Contrastive Instance Feature Mining for Outdoor 3D Object Detection with Very Limited Annotations
abstract
Recently, 3D object detection with sparse annotations has received great attention. However, current detectors usually perform poorly under very limited annotations. To address this problem, we propose a novel Contrastive Instance feature mining method, named CoIn. To better identify indistinguishable features learned through limited supervision, we design a Multi-Class contrastive learning module (MCcont) to enhance feature discrimination. Meanwhile, we propose a feature-level pseudo-label mining framework consisting of an instance feature mining module (InF-Mining) and a Labeled-to-Pseudo contrastive learning module (LPcont). These two modules exploit latent instances in feature space to supervise the training of detectors with limited annotations. Extensive experiments with KITTI dataset, Waymo open dataset, and nuScenes dataset show that under limited annotations, our method greatly improves the performance of baseline detectors: CenterPoint, Voxel-RCNN, and CasA. Combining CoIn with an iterative training strategy, we propose a CoIn++ pipeline, which requires only 2% annotations in the KITTI dataset to achieve performance comparable to the fully supervised methods. The code is available at https://github.com/xmuqimingxia/CoIn.
Qiming Xia, Jinhao Deng, Chenglu Wen, Shaoshuai Shi, Xin Li 0003, Cheng Wang 0003
ICCV1
2023 3-D HANet: A Flexible 3-D Heatmap Auxiliary Network for Object Detection
abstract
3-D object detection is a vital part of outdoor scene perception. Learning the complete size and accurate positioning of objects from an incomplete point cloud spatial structure is essential to 3-D object detection. We propose a novel flexible 3-D heatmap auxiliary network (3-D HANet) for object detection. To obtain complete structure and location information from an incomplete point cloud structure, we propose a 3-D heatmap to reflect object information. Also, we design a plug-and-play auxiliary network based on 3-D heatmap, which improves the accuracy of the entire detection network without extra computation in the inference stage. We validate the 3-D HANet on the basis of three classic 3-D object detection networks: PointPillars, sparsely embedded convolutional detection (SECOND), and structure aware single-stage 3-D object detection from point cloud (SASSD). Experimental results show that our auxiliary network augments the feature extraction ability of the backbone network, which is manifested in that the predicted boxes and the ground-truth boxes are more suitable in size and more aligned in direction. Furthermore, we conducted verification experiments on the state-of-the-art (SOTA) detector, CasA, and made a further improvement on the official ranking of the KITTI dataset.
Qiming Xia, Yidong Chen 0006, Guo-Rong Cai, Guikun Chen, Daoshun Xie, Jinhe Su, Zongyue Wang
IEEE Trans. Geosci. Remote. Sens.1
2022 SSA3D: Semantic Segmentation Assisted One-Stage Three-Dimensional Vehicle Object Detection
abstract
One-stage 3D object detection using mobile light detection and ranging (LiDAR) has developed rapidly in recent years. Specifically, one-stage methods have attracted attention because of their high efficiency and light weight compared with two-stage methods. Inspired by this, we present the semantic segmentation assisted one-stage three-dimensional vehicle object detection (SSA3D), a network for the rapid detection of objects that keeps the advantages of the semantic segmentation module in the two-stage methods without increasing redundant computational load. First, we modified the sampling of the farthest point to improve the quality of the sampling points. This helps to reduce sampling outlier points and bad points that are difficult to perceive in the spatial structure information surrounding the point. Second, a neighbor attention group module is devoted to selectively add extra weight to neighbor points because of the different importance of neighbor points for the corresponding sampling point. Correctly increasing the weight is helpful to obtain richer spatial structure information. Finally, a delicate box generation module is included as a voted center point layer based on the generalized Hoff vote method and an anchor-free regression. We used the feature aggregation module as the backbone and the feature propagation module as the auxiliary network to achieve efficiency. At the same time, the auxiliary network retains the ability to extract point-wise features from the state-of-the-art semantic segmentation network. In experiments, we evaluated and tested the SSA3D on a common KITTI dataset and achieved improved performance in the class of car accuracy.
Shangfeng Huang, Guo-Rong Cai, Zongyue Wang, Qiming Xia, Ruisheng Wang 0001
IEEE Trans. Intell. Transp. Syst.4
2020 Multi-layer Pointpillars: Multi-layer Feature Abstraction for Object Detection from Point Cloud
Shangfeng Huang, Qiming Xia, Yanhao Lin, Haiyan Lian, Zongyue Wang, Guo-Rong Cai, Jinhe Su
PRCV (1)2