Li Wang 0092

dblp:58/6810-92 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-9325-2391ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Global relationship awareness 3-dimensional object detection using 4-dimensional radar
Pianzhang Duan, Li Wang 0092, Ziying Song, Ying Li 0036, Wei Fan 0011, Bin Xu 0003
Eng. Appl. Artif. Intell.2
2026 V 2 -Fusion: Virtual voxel enhanced 4D radar-image feature fusion for 3D object detection
Li Wang 0092, Xinyu Zhang 0001, Yuxuan Fan, Tao Xie 0010, Lei Yang 0060, Bin Xu 0003
Expert Syst. Appl.1
2026 CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection
abstract
The sparse cross-modality detector offers more advantages than its counterpart, the Bird’s-Eye-View (BEV) detector, particularly in terms of adaptability for downstream tasks and computational cost savings. However, existing sparse detectors overlook the quality of token representation, leaving it with a sub-optimal foreground quality and limited performance. In this paper, we identify that the geometric structure preserved and the class distribution are the key to improving the performance of the sparse detector, and propose a Sparse Selector (SS). The core module of SS is Ray-Aware Supervision (RAS), which preserves rich geometric information during the training stage, and Class-Balanced Supervision, which adaptively reweights the salience of class semantics, ensuring that tokens associated with small objects are retained during token sampling. Thereby, outperforming other sparse multi-modal detectors in the representation of tokens. Additionally, we design Ray Positional Encoding (Ray PE) to address the distribution differences between the LiDAR modality and the image. Finally, we integrate the aforementioned module into an end-to-end sparse multi-modality detector, dubbed CrossRay3D. Experiments show that, on the challenging nuScenes benchmark, CrossRay3D achieves state-of-the-art performance with 72.4% mAP and 74.7% NDS, while running$1.84\times $faster than other leading methods. Moreover, CrossRay3D demonstrates strong robustness even in scenarios where LiDAR or camera data are partially or entirely missing. The code is available onhttps://github.com/xuehaipiaoxiang/CrossRay3D
Huiming Yang, Wenzhuo Liu, Yicheng Qiao, Lei Yang 0060, Xianzhu Zeng, Li Wang 0092, Zhiwei Li 0011, Zijian Zeng 0001, Zhiying Jiang, Huaping Liu 0001, Kunfeng Wang
IEEE Trans. Intell. Transp. Syst.6
2025 V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception
abstract
Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In recent years, a series of cooperative perception datasets have emerged; however, these datasets primarily focus on cameras and LiDAR, neglecting 4D Radar—a sensor used in single-vehicle autonomous driving to provide robust perception in adverse weather conditions. In this paper, to bridge the gap created by the absence of 4D Radar datasets in cooperative perception, we present V2X-Radar, the first large-scale, real-world multi-modal dataset featuring 4D Radar. V2X-Radar dataset is collected using a connected vehicle platform and an intelligent roadside unit equipped with 4D Radar, LiDAR, and multi-view cameras. The collected data encompasses sunny and rainy weather conditions, spanning daytime, dusk, and nighttime, as well as various typical challenging scenarios. The dataset consists of 20K LiDAR frames, 40K camera images, and 20K 4D Radar data, including 350K annotated boxes across five categories. To support various research domains, we have established V2X-Radar-C for cooperative perception, V2X-Radar-I for roadside perception, and V2X-Radar-V for single-vehicle perception. Furthermore, we provide comprehensive benchmarks across these three sub-datasets.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Jiaqi Ma 0003, Zhiying Song, Ziying Song, Li Wang 0092, Yang Shen 0005, Chen Lv 0001
NeurIPS9
2025 BEVHeight++: Toward Robust Visual Centric 3D Object Detection
abstract
While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric detection methods perform poorly on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight++, to address this issue. In essence, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. By incorporating both height and depth encoding techniques, we achieve a more accurate and robust projection from 2D to BEV spaces. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. In terms of the ego-vehicle scenario, BEVHeight++ surpasses depth-only methods with increases of +2.8% NDS and +1.7% mAP on the nuScenes test set, and even higher gains of +9.3% NDS and +8.8% mAP on the nuScenes-C benchmark with object-level distortion. Consistent and substantial performance improvements are achieved across the KITTI, KITTI-360, and Waymo datasets as well.
Lei Yang 0060, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Yi Huang 0038, Xinyu Zhang 0001, Kaicheng Yu
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 FMRT: Learning Accurate Feature Matching With Reconciliatory Transformer
abstract
Local Feature Matching, a pivotal component of numerous computer vision tasks (e.g., structure from motion and visual localization), has been effectively addressed by Transformer-based methods. Nevertheless, these methods solely incorporate long-range context information among keypoints with a fixed receptive field, which constrains the network from appropriately reconciling the importance of features with diverse receptive fields to realize complete image perception, hence limiting feature matching accuracy. In addition, these methods employ a conventional handcrafted encoding approach to incorporate positional information of keypoints into visual descriptors, which limits the capability of networks to extract effective positional encoding message. In this study, we propose FMRT, a novel detector-free method that reconciles local features with diverse receptive fields adaptively and utilizes parallel networks to realize reliable positional encoding. Specifically, FMRT proposes a dedicated reconciliatory transformer (RecFormer) that contains a global perception attention layer to identify visual descriptors with different receptive fields and integrate global context information under various scales, a perception weight layer to measure the importance of various receptive fields adaptively, and a local perception feed-forward network to extract deep aggregated multi-scale local feature representation. Moreover, we introduce a novel axis-wise position encoder (AWPE) that views positional encoding as two keypoints encoding tasks along the row and column dimensions, decouples the x- and y-coordinates of keypoints into two independent 1D vectors, and designs two parallel network branches to explicitly encodes geometric correlations among keypoints, hence realizing reliable positional encoding. Extensive experiments indicate that FMRT yields impressive performance on multiple tasks, including relative pose estimation, visual localization, homography estimation, and image matching. Besides, we integrate FMRT into a localization framework and conduct a visual localization experiment in a real scene, which further demonstrate the superiority of FMRT. Note to Practitioners—This paper presents a novel approach to enhancing the performance of local feature matching in computer vision tasks. Traditional methods often rely on fixed receptive fields for integrating context among keypoints, which can limit the perception of the complete image and, consequently, the precision of feature matching. Our work introduces a Reconciliatory Transformer that not only addresses these limitations by effectively reconciling the importance of features across varying receptive fields but also improves the integration of positional information into visual descriptors. The techniques developed here can be adapted to a wide range of systems, e.g., image matching for computer vision and visual localization for autonomous driving, offering practitioners a tool to significantly improve the fidelity of feature matching, which is foundational for accurate interaction with the surrounding environment.
Li Wang 0092, Xinyu Zhang 0001, Tao Xie 0010, Lei Yang 0060, Wenhao Yu 0006, Yang Shen 0005, Bin Xu 0003, Jun Li 0082
IEEE Trans Autom. Sci. Eng.1
2025 Steering Angle-Guided Multimodal Fusion Lane Detection for Autonomous Driving
abstract
Lane detection is a critical part of autonomous driving technology. When difficult situations are encountered (i.e., adverse light, severe occlusion), the lane detection task is still challenging. However, previous methods strongly depend on the extracted image features and ignore other features. It is necessary to consider the information from other modalities to assist the model for lane detection, especially in the task of curved lane detection. In this paper, considering that the vehicle steering angle is closely related to the visual feature of lane lines, we propose a novel model named Image-Angle Fusion Network (IAFNet) to solve the lane detection problem by fusing vehicle steering angle features with image features. To make the steering angle features better match the image features, we use the tensor outer product to extend the dimensionality of the steering angle information. A lightweight Image-Angle cross-attention module (LIA-CAM) is proposed to learn the implicit relationship between steering angles and visual features of lane lines, aimed at improving the performance of our model in difficult situations. To guide the network to retain the correct steering angle information, we introduced regression prediction loss of steering angle. Besides, we also released a new dataset based on the Udacity dataset: ImageAngle-Udacity (IA-Udacity) dataset. Extensive experiments on the IA-Udacity dataset show that our method outperforms the current state-of-the-art methods showing both higher efficiency and accuracy. Code and data are available onhttps://github.com/gongyan1/LIA-CAM.
Xinyu Zhang 0001, Jianli Lu, Xinmin Jiang, Hao Liu 0114, Zhiwei Li 0011, Li Wang 0092, Qingshan Yang, Xingang Wu
IEEE Trans. Intell. Transp. Syst.8
2025 UMD-Net: A Unified Multi-Task Assistive Driving Network Based on Multimodal Fusion
abstract
In recent years, researchers have focused on identifying tasks related to driver state, traffic environment, and others to enhance the safety of autonomous driving assistance systems. However, current research on these tasks is conducted independently, neglecting the interconnections between the driver, traffic environment, and vehicle. In this paper, we propose a Unified Multi-task Assistive Driving Network Based on Multimodal Fusion (UMD-Net), the first unified model capable of recognizing four tasks simultaneously by utilizing multimodal data: driver behavior recognition, driver emotion recognition, traffic context recognition, and vehicle behavior recognition. In order to better enhance the synergistic effects between multiple tasks, we designed the position-sensitive multi-directional attention feature extraction subnetwork and recursive dynamic feature fusion module. The former captures the key features of multi-view images by different directions of attention mechanism to improve the generalization of the model across multiple tasks. The latter dynamically adjusts the fusion weight according to the multimodal features to enhance the representation ability of important features in multi-task learning. Our model was evaluated on the public dataset AIDE, achieving the best performance across all four tasks and a high accuracy of 95.31% in the traffic context recognition task, demonstrating the superiority of our approach. The code is available on https://github.com/Wenzhuo-Liu/UMD-Net.
Wenzhuo Liu, Yicheng Qiao, Zhiwei Li 0011, Wenshuo Wang 0001, Wei Zhang 0012, Jiayin Zhu, Yanhuan Jiang, Li Wang 0092, Hong Wang 0014, Huaping Liu 0001, Kunfeng Wang
IEEE Trans. Intell. Transp. Syst.8
2025 SGV3D: Toward Scenario Generalization for Vision-Based Roadside 3D Object Detection
abstract
Roadside perception can significantly enhance the safety of autonomous vehicles by extending their perceptual capabilities beyond the visual range and addressing occluded regions. However, current state-of-the-art vision-based roadside detection methods exhibit high accuracy on labeled scenes but perform poorly on new scenes. This limitation arises because roadside cameras remain stationary after installation and can only gather data from a single scene, leading the algorithm to overfit these roadside backgrounds and camera positions. To tackle this issue, we propose an innovativeScenarioGeneralization Framework forVision-based Roadside3DObject Detection, calledSGV3D. Specifically, we utilize a Background-suppressed Module (BSM) to reduce background overfitting in vision-centric pipelines by diminishing background features during the 2D to bird’s-eye-view projection. Furthermore, by introducing the Semi-supervised Data Generation Pipeline (SSDG) that employs unlabeled images from new scenes, we generate diverse foreground instances with varying camera poses, mitigating the risk of overfitting to specific camera positions. Experiments conducted on two large-scale roadside benchmarks demonstrate that SGV3D, with only a minimal increase in latency, effectively improves the scenario generalization capabilities of vision-based roadside 3D object detectors. The code is available here (https://github.com/yanglei18/SGV3D).
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Zhiwei Li 0011, Yang Shen 0005, Chen Lv 0001, Hong Wang 0014
IEEE Trans. Intell. Transp. Syst.4
2024 GraphBEV: Towards Robust BEV Feature Alignment for Multi-modal 3D Object Detection
Ziying Song, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092
ECCV (26)8
2024 RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
Ziying Song, Guoxing Zhang, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092
IJCAI8
2024 Robustness-Aware 3D Object Detection in Autonomous Driving: A Review and Outlook
abstract
In the realm of modern autonomous driving, the perception system is indispensable for accurately assessing the state of the surrounding environment, thereby enabling informed prediction and planning. The key step to this system is related to 3D object detection that utilizes vehicle-mounted sensors such as LiDAR and cameras to identify the size, the category, and the location of nearby objects. Despite the surge in 3D object detection methods aimed at enhancing detection precision and efficiency, there is a gap in the literature that systematically examines their resilience against environmental variations, noise, and weather changes. This study emphasizes the importance of robustness, alongside accuracy and latency, in evaluating perception systems under practical scenarios. Our work presents an extensive survey of camera-only, LiDAR-only, and multi-modal 3D object detection algorithms, thoroughly evaluating their trade-off between accuracy, latency, and robustness, particularly on datasets like KITTI-C and nuScenes-C to ensure fair comparisons. Among these, multi-modal 3D detection approaches exhibit superior robustness, and a novel taxonomy is introduced to reorganize the literature for enhanced clarity. This survey aims to offer a more practical perspective on the current capabilities and the constraints of 3D object detection algorithms in real-world applications, thus steering future research towards robustness-centric advancements.
Ziying Song, Feiyang Jia, Yadan Luo, Caiyan Jia, Lei Yang 0060, Li Wang 0092
IEEE Trans. Intell. Transp. Syst.8
2024 MonoGAE: Roadside Monocular 3D Object Detection With Ground-Aware Embeddings
abstract
Although the majority of recent autonomous driving systems concentrate on developing perception methods based on ego-vehicle sensors, there is an overlooked alternative approach that involves leveraging intelligent roadside cameras to help extend the ego-vehicle perception ability beyond the visual range. We discover that most existing monocular 3D object detectors rely on the ego-vehicle prior assumption that the optical axis of the camera is parallel to the ground. However, the roadside camera is installed on a pole with a pitched angle, which makes the existing methods not optimal for roadside scenes. In this paper, we introduce a novel framework for Roadside Monocular 3D object detection with ground-aware embeddings, named MonoGAE. Specifically, the ground plane is a stable and strong prior knowledge due to the fixed installation of cameras in roadside scenarios. In order to reduce the domain gap between the ground geometry information and high-dimensional image features, we employ a supervised training paradigm with a ground plane to predict high-dimensional ground-aware embeddings. These embeddings are subsequently integrated with image features through cross-attention mechanisms. Furthermore, to improve the detector’s robustness to the divergences in cameras’ installation poses, we replace the ground plane depth map with a novel pixel-level refined ground plane equation map. Our approach demonstrates a substantial performance advantage over all previous monocular 3D object detectors on widely recognized 3D detection benchmarks for roadside cameras. The code and pre-trained models will be released soon.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Yi Huang 0038, Hong Wang 0014
IEEE Trans. Intell. Transp. Syst.6
2024 Auto-Points: Automatic Learning for Point Cloud Analysis With Neural Architecture Search
abstract
Pure point-based neural networks have recently shown tremendous promise for point cloud tasks, including 3D object classification, 3D object part segmentation, 3D semantic segmentation, and 3D object detection. Nevertheless, it is a laborious process to construct a network for each task due to the artificial parameters and hyperparameters involved, e.g., the depths and widths of the network and the number of sampled points at each stage. In this work, we propose Auto-Points, a novel one-shot search framework that automatically seeks the optimal architecture configuration for point cloud tasks. Technically, we introduce a set abstraction mixer (SAM) layer that is capable of scaling up flexibly along the depth and width of the network. Each SAM layer consists of numerous child candidates, which simplifies architecture search and enables us to discover the optimum design for each point cloud task pursuant to resource constraint from an enormous search space. To fully optimize the child candidates, we develop a weight-entwinement neural architecture search (NAS) technique that entwines the weights of different candidates in the same layer during supernet training such that all candidates can be extremely optimized. Benefiting from the proposed techniques, the trained supernet allows the searched subnets to be exceptionally well-optimized without further retraining or finetuning. In particular, the searched models deliver superior performances on multiple extensively employed benchmarks, 93.9% overall accuracy (OA) on ModelNet40, 89.1% OA on ScanObjectNN, 87.1% instance average IoU on ShapeNetPart, 69.1% mIoU on S3DIS, 70.4% [email protected] on ScanNet V2, and 64.4% [email protected] on SUN RGB-D.
Li Wang 0092, Tao Xie 0010, Xinyu Zhang 0001, Linqi Yang, Yilong Ren, Haiyang Yu 0002, Jun Li 0082, Huaping Liu 0001
IEEE Trans. Multim.1
2024 FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object Detection
abstract
In this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance.
Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082
IEEE Trans. Multim.2
2023 BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection
abstract
While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric bird's eye view detection methods have inferior performances on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight, to address this issue. In essence, instead of predicting the pixel-wise depth, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. The code is available at https://github.com/ADLab-AutoDrive/BEVHeight.
Lei Yang 0060, Kaicheng Yu, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Xinyu Zhang 0001
CVPR6
2023 CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive Network
abstract
We present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specifically, we leverage residual MLP (Res-MLP) block for effective feature extraction and scale it gracefully along the depth and width of the network to meet the demands of different tasks. Based on the block, we propose a novel nested layer-wise processing policy, which identifies the optimal architecture for each task while provides partial sharing parameters and partial non-sharing parameters inside each layer of the block. Such policy tackles the inherent challenges of multi-task learning on point cloud, e.g., diverse model topologies resulting from task skew and conflicting gradients induced by heterogeneous dataset domains. Finally, we propose a sign-based gradient surgery to promote the training of CO-Net, thereby emphasizing the usage of task-shared parameters and guaranteeing that each task can be thoroughly optimized. Experimental results reveal that models optimized by CO-Net jointly for all point cloud tasks maintain much fewer computation cost and overall storage cost yet outpace prior methods by a significant margin. We also demonstrate that CO-Net allows incremental learning and prevents catastrophic amnesia when adapting to a new point cloud task.
Tao Xie 0010, Ke Wang 0028, Siyi Lu, Jie Xu 0066, Li Wang 0092, Lijun Zhao 0003, Xinyu Zhang 0001, Ruifeng Li 0001
ICCV8
2023 SAT-GCN: Self-attention graph convolutional network-based 3D object detection for autonomous driving
Li Wang 0092, Ziying Song, Xinyu Zhang 0001, Jun Li 0082, Huaping Liu 0001
Knowl. Based Syst.1
2023 Lite-FPN for keypoint-based monocular 3D object detection
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu
Knowl. Based Syst.4
2023 Mix-Teaching: A Simple, Unified and Effective Semi-Supervised Learning Framework for Monocular 3D Object Detection
abstract
Semi-supervised learning (SSL) has promising potential for improving model performance using both labelled and unlabelled data. Since recovering 3D information from 2D images is an ill-posed problem, the current state-of-the-art methods of monocular 3D object detection (Mono3D) have relatively low precision and recall, making semi-supervised learning for Mono3D tasks challenging and understudied. In this work, we propose a unified and effective semi-supervised learning framework called Mix-Teaching that can be applied to most monocular 3D object detectors. Based on the idea of decomposition and recombination, unlabelled samples are firstly decomposed into collections of image patches with high-quality predictions and collections of background images containing no objects. The student model is then trained on the mixed images containing dense instances with high-quality pseudo-labels generated by the recombination operation. In addition, we propose an uncertainty-based filter to distinguish high-quality pseudo-labels from noisy predictions during the decomposition process. As results in KITTI and nuScenes benchmarks, Mix-Teaching consistently improves MonoFlex and GUPNet by significant margins under various labeling ratios. Our method achieves around +6.34%$AP_{3D}$improvement against the GUPNet on the validation set when using only 10% labelled data. Using the full training set and the additional 38K raw images from KITTI, it can further improve the MonoFlex by +4.65% absolute improvement on$AP_{3D}$for car detection, reaching 18.54%$AP_{3D}$, which ranks the 1st place among all monocular based methods on the KITTI test leaderboard.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu, Huaping Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 CAMO-MOT: Combined Appearance-Motion Optimization for 3D Multi-Object Tracking With Camera-LiDAR Fusion
abstract
3D Multi-object tracking (MOT) ensures consistency during continuous dynamic detection, conducive to subsequent motion planning and navigation tasks in autonomous driving. However, camera-based methods suffer in the case of occlusions and it can be challenging to track the irregular motion of objects for LiDAR-based methods accurately. Some fusion methods work well but do not consider the untrustworthy issue of appearance features under occlusion. At the same time, the false detection problem also significantly affects tracking. As such, we propose a novel camera-LiDAR fusion 3D MOT framework based on Combined Appearance-Motion Optimization (CAMO-MOT), which uses both camera and LiDAR data and significantly reduces tracking failures caused by occlusion and false detection. For occlusion problems, we are the first to propose an occlusion head to select the best object appearance features multiple times effectively, reducing the influence of occlusions. To decrease the impact of false detection in tracking, we design a motion cost matrix based on confidence scores which improve the positioning and object prediction accuracy in 3D space. As existing multi-object tracking methods always evaluate each category separately and do not consider the mismatch between objects of different categories, we also propose to build a multi-category cost to implement multi-object tracking in multi-category scenes. A series of validation experiments are conducted on the KITTI and nuScenes tracking benchmarks. Our proposed method achieves state-of-the-art performance with 79.99% HOTA and the lowest identity switches (IDS) value (23 for Car and 137 for Pedestrian) among all multi-modal MOT methods on the KITTI test dataset. And our method achieves state-of-the-art performance among all algorithms on the nuScenes test dataset with 75.3% AMOTA.
Li Wang 0092, Xinyu Zhang 0001, Wenyuan Qin, Jinghan Gao, Lei Yang 0060, Zhiwei Li 0011, Jun Li 0082, Hong Wang 0014, Huaping Liu 0001
IEEE Trans. Intell. Transp. Syst.1
2022 InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection
abstract
Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggested. In recent years, the fusion of LiDAR and Radar has gained ever-increasing attention, especially 4D Radar, which can adapt to bad weather conditions due to its penetrability. Although features have been fused from multiple sensing modalities, most methods cannot learn interactions from different modalities, which does not make for their best use. Inspired by the self-attention mechanism, we present InterFusion, an interaction-based fusion framework, to fuse 16-line LiDAR with 4D Radar. It aggregates features from two modalities and identifies cross-modal relations between Radar and LiDAR features. In experimental evaluations on the Astyx HiRes 2019 dataset, our method outperformed the baseline by 4.20% mAP in 3D and 10.76% BEV mAP for the car class at the moderate level.
Li Wang 0092, Xinyu Zhang 0001, Baowei Xv, Jinzhao Zhang, Haibing Ren, Pingping Lu, Jun Li 0082, Huaping Liu 0001
IROS1
2022 Fast Detection of Multi-Direction Remote Sensing Ship Object Based on Scale Space Pyramid
abstract
Ships in remote sensing images are usually arranged in arbitrary direction, small in size, and densely arranged. As a result, existing object detection algorithms cannot detect ships quickly and accurately. In order to solve the above problems, a lightweight object detection network for fast detection of ships is proposed. The network is composed of backbone network, four-scale fusion network and rotation branch. First, a lightweight network unit S-LeanNet is designed and used to build a low-computing and accurate backbone network. Then, a four-scale feature fusion module is designed to generate a four-scale feature pyramid, which contains more features such as ship shape and texture, and at the same time is conducive to the detection of small ships. Finally, a novel rotation branch module is designed, using balance L1 loss function and R-NMS for post-processing, to realize the precise positioning and regression of the rotating bounding box in one step. Experimental results show that the detection precision of our method in the DOT A remote sensing data set is compared with the latest SCRDet detection method, the precision is increased by 1.1%, and the operating speed is increased by 8 times, which can meet the fast detection requirements of ships.
Ziying Song, Li Wang 0092, Caiyan Jia, Jiangfeng Bi, Haiyue Wei, Yongchao Xia, Lijun Zhao 0003
MSN2
2018 Feature-Based and Convolutional Neural Network Fusion Method for Visual Relocalization
abstract
Relocalization is one of the necessary modules for mobile robots in long-term autonomous movement in an environment. Currently, visual relocalization algorithms mainly include feature-based methods and CNN-based (Convolutional Neural Network) methods. Feature-based methods can achieve high localization accuracy in feature-rich scenes, but the error is quite large or it even fails in cases with motion blur, texture-less scene and changing view angle. CNN-based methods usually have better robustness but poor localization accuracy. For this reason, a visual relocalization algorithm that combines the advantages of the two methods is proposed in this paper. The BoVW (Bag of Visual Words) model is used to search for the most similar image in the training dataset. PnP (Perspective n Points) and RANSAC (Random Sample Consensus) are employed to estimate an initial pose. Then the number of inliers is utilized as a criterion whether the feature-based method or the CNN-based method is to be leveraged. Compared with a previous CNN-based method, PoseNet, the average position error is reduced by 45.6% and the average orientation error is reduced by 67.4% on Microsoft's 7-Scenes datasets, which verifies the effectiveness of the proposed algorithm.
Li Wang 0092, Ruifeng Li 0001, Seah Hock Soon, Chee Kwang Quah, Lijun Zhao 0003
ICARCV1
2018 A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization
abstract
In this paper, we propose a deep learning approach to tackle the automatic summarization tasks by incorporating topic information into the convolutional sequence-to-sequence (ConvS2S) model and using self-critical sequence training (SCST) for optimization. Through jointly attending to topics and word-level alignment, our approach can improve coherence, diversity, and informativeness of generated summaries via a biased probability generation mechanism. On the other hand, reinforcement training, like SCST, directly optimizes the proposed model with respect to the non-differentiable metric ROUGE, which also avoids the exposure bias during inference. We carry out the experimental evaluation with state-of-the-art methods over the Gigaword, DUC-2004, and LCSTS datasets. The empirical results demonstrate the superiority of our proposed method in the abstractive summarization.
Li Wang 0092, Junlin Yao, Yunzhe Tao, Wei Liu 0005, Qiang Du 0001
IJCAI1
2015 Unsupervised feature selection based on spectral regression from manifold learning for facial expression recognition
abstract
In this study, an unsupervised feature selection method is proposed for facial feature recognition (FER) in the absence of class labels. The contribution is the descriptive feature components selector spectral regression representative coefficient scores based on graph manifold learning from high‐dimensional feature space. The spectral regression analysis and L1‐regularised least square are then used to compute the importance of features in the original space, so that less representative features with lower coefficient scores will be removed without prior distribution assumption. To verify the performance of the authors’ method, some classifiers are used to classify facial expressions on three benchmark facial expression databases. The recognition results indicate the availability and effectiveness of the proposed method for FER.
Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001
IET Comput. Vis.1