EDBT 2026 Demo / reviewers in the wild / expert
Zhiwei Li 0011
dblp:47/3951-11
· DBLP profile ↗
15ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0001-7071-199XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Structure-Equivalent Prior: Unifying Temporal Dynamics and 3D Evolution in 4D Latent SpaceabstractRecent advances in deep learning-based 3D representation have achieved remarkable success, particularly in modeling static high-fidelity geometries. However, the extension of these techniques to dynamic 3D scenes introduces a critical challenge of effectively representing spatio-temporal dependencies, i.e., jointly modeling detailed spatial structures within frames and temporal dynamics across frames. To address this challenge, this paper proposes that the temporal evolution observed in dynamic 3D scenes is fundamentally attributable to the deformation of underlying spatial structures. To capture this relationship, we introduce a unified continuous 4D latent space representation incorporating a structure-equivalence prior, named SEP-4D. The core of SEP-4D is an efficient 4D tensor decomposition-fusion approach. This method fuses decomposed learnable 2D feature planes via a plane-wise spatio-temporal fusion mechanism of planar distributions, explicitly enforcing the principle that temporal evolution originates from geometric deformations of the 3D structure. To mitigate the associated computational demands, we sample the 3D probability volumes generated by VAE-based fusion into a spatio-temporally consistent 4D latent representation. The efficacy of our approach is validated through experiments on the fundamental task of 4D occupancy reconstruction. Extensive results demonstrate that, by leveraging the inherent equivalence of temporal dynamics and structural deformation, our method achieves high-quality reconstruction across various sequence lengths. Notably, for 4-frame scenes, we attain an impressive 91.68% mIoU, significantly outperforming state-of-the-art baselines on standard benchmarks. Jingyuan Gao, Tianyu Shen, Ruosen Hao, Te Guo 0003, Zhiwei Li 0011, Kunfeng Wang |
AAAI | 5 |
| 2026 | CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D DetectionabstractThe sparse cross-modality detector offers more advantages than its counterpart, the Bird’s-Eye-View (BEV) detector, particularly in terms of adaptability for downstream tasks and computational cost savings. However, existing sparse detectors overlook the quality of token representation, leaving it with a sub-optimal foreground quality and limited performance. In this paper, we identify that the geometric structure preserved and the class distribution are the key to improving the performance of the sparse detector, and propose a Sparse Selector (SS). The core module of SS is Ray-Aware Supervision (RAS), which preserves rich geometric information during the training stage, and Class-Balanced Supervision, which adaptively reweights the salience of class semantics, ensuring that tokens associated with small objects are retained during token sampling. Thereby, outperforming other sparse multi-modal detectors in the representation of tokens. Additionally, we design Ray Positional Encoding (Ray PE) to address the distribution differences between the LiDAR modality and the image. Finally, we integrate the aforementioned module into an end-to-end sparse multi-modality detector, dubbed CrossRay3D. Experiments show that, on the challenging nuScenes benchmark, CrossRay3D achieves state-of-the-art performance with 72.4% mAP and 74.7% NDS, while running$1.84\times $faster than other leading methods. Moreover, CrossRay3D demonstrates strong robustness even in scenarios where LiDAR or camera data are partially or entirely missing. The code is available onhttps://github.com/xuehaipiaoxiang/CrossRay3D Huiming Yang, Wenzhuo Liu, Yicheng Qiao, Lei Yang 0060, Xianzhu Zeng, Li Wang 0092, Zhiwei Li 0011, Zijian Zeng 0001, Zhiying Jiang, Huaping Liu 0001, Kunfeng Wang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving PerceptionabstractAdvanced driver assistance systems require a comprehensive understanding of the driver’s mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper proposes MMTL-UniAD, a unified multi-modal multitask learning framework that simultaneously recognizes driver behavior (e.g., looking around, talking), driver emotion (e.g., anxiety, happiness), vehicle behavior (e.g., parking, turning), and traffic context (e.g., traffic jam, traffic smooth). A key challenge is avoiding negative transfer between tasks, which can impair learning performance. To address this, we introduce two key components into the framework: one is the multi-axis region attention network to extract global context-sensitive features, and the other is the dual-branch multimodal embedding to learn multi-modal embeddings from both task-shared and task-specific features. The former uses a multi-attention mechanism to extract task-relevant features, mitigating negative transfer caused by task-unrelated features. The latter employs a dual-branch structure to adaptively adjust task-shared and task-specific parameters, enhancing cross-task knowledge transfer while reducing task conflicts. We assess MMTL-UniAD on the AIDE dataset, using a series of ablation studies, and show that it outperforms state-of-the-art methods across all four tasks. The code is available on https://github.com/Wenzhuo-Liu/MMTL-UniAD. Wenzhuo Liu, Wenshuo Wang 0001, Yicheng Qiao, Qiannan Guo, Jiayin Zhu, Zilong Chen, Huiming Yang, Zhiwei Li 0011, Tiao Tan, Huaping Liu 0001 |
CVPR | 9 |
| 2025 | TEM3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive DrivingabstractMulti-task learning (MTL) can advance assistive driving by exploring inter-task correlations through shared representations. However, existing methods face two critical limitations: single-modality constraints limiting comprehensive scene understanding and inefficient architectures impeding real-time deployment. This paper proposes TEM3-Learning (Time-Efficient Multimodal Multi-task Learning), a novel framework that jointly optimizes driver emotion recognition, driver behavior recognition, traffic context recognition, and vehicle behavior recognition through a two-stage architecture. The first component, the mamba-based multi-view temporal-spatial feature extraction subnetwork (MTS-Mamba), introduces a forward-backward temporal scanning mechanism and global-local spatial attention to efficiently extract low-cost temporal-spatial features from multi-view sequential images. The second component, the MTL-based gated multimodal feature integrator (MGMI), employs task-specific multi-gating modules to adaptively highlight the most relevant modality features for each task, effectively alleviating the negative transfer problem in MTL. Evaluation on the AIDE dataset, our proposed model achieves state-of-the-art accuracy across all four tasks, maintaining a lightweight architecture with fewer than 6 million parameters and delivering an impressive 142.32 FPS inference speed. Rigorous ablation studies further validate the effectiveness of the proposed framework and the independent contributions of each module. The code is available on https://github.com/Wenzhuo-Liu/TEM3-Learning. Wenzhuo Liu, Yicheng Qiao, Qiannan Guo, Zilong Chen, Meihua Zhou, Zhiwei Li 0011, Huaping Liu 0001, Wenshuo Wang 0001 |
IROS | 9 |
| 2025 | SAMOccNet:Refined SAM-based surrounding semantic occupancy perception for autonomous driving
Qifan Tan, Wenzhuo Liu, Han Bi, Lei Yang 0060, Yicheng Qiao, Zhuo Zhao, Yanhuan Jiang, Qiannan Guo, Huaping Liu 0001, Zhiwei Li 0011 |
Neurocomputing | 11 |
| 2025 | Steering Angle-Guided Multimodal Fusion Lane Detection for Autonomous DrivingabstractLane detection is a critical part of autonomous driving technology. When difficult situations are encountered (i.e., adverse light, severe occlusion), the lane detection task is still challenging. However, previous methods strongly depend on the extracted image features and ignore other features. It is necessary to consider the information from other modalities to assist the model for lane detection, especially in the task of curved lane detection. In this paper, considering that the vehicle steering angle is closely related to the visual feature of lane lines, we propose a novel model named Image-Angle Fusion Network (IAFNet) to solve the lane detection problem by fusing vehicle steering angle features with image features. To make the steering angle features better match the image features, we use the tensor outer product to extend the dimensionality of the steering angle information. A lightweight Image-Angle cross-attention module (LIA-CAM) is proposed to learn the implicit relationship between steering angles and visual features of lane lines, aimed at improving the performance of our model in difficult situations. To guide the network to retain the correct steering angle information, we introduced regression prediction loss of steering angle. Besides, we also released a new dataset based on the Udacity dataset: ImageAngle-Udacity (IA-Udacity) dataset. Extensive experiments on the IA-Udacity dataset show that our method outperforms the current state-of-the-art methods showing both higher efficiency and accuracy. Code and data are available onhttps://github.com/gongyan1/LIA-CAM. Xinyu Zhang 0001, Jianli Lu, Xinmin Jiang, Hao Liu 0114, Zhiwei Li 0011, Li Wang 0092, Qingshan Yang, Xingang Wu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | MIPD: A Multi-Sensory Interactive Perception Dataset for Embodied Intelligent DrivingabstractDuring the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional sensory information in order to facilitate interaction with the environment. However, the current multi-modal fusion sensing schemes often neglect these additional sensory inputs, hindering the realization of fully autonomous driving. This paper considers multi-sensory information and proposes a multi-modal interactive perception dataset named MIPD, enabling expanding the current autonomous driving algorithm framework, for supporting the research on embodied intelligent driving. In addition to the conventional camera, lidar, and 4D radar data, our dataset incorporates multiple sensor inputs including sound, light intensity, vibration intensity and vehicle speed to enrich the dataset comprehensiveness. Comprising 126 consecutive sequences, many exceeding twenty seconds, MIPD features over 8,500 meticulously synchronized and annotated frames. Moreover, it encompasses many challenging scenarios, covering various road and lighting conditions. The dataset has undergone thorough experimental validation, producing valuable insights for the exploration of next-generation autonomous driving frameworks. Data, development kit and more details will be available athttps://github.com/BUCT-IUSRC/Dataset__MIPD Zhiwei Li 0011, Tingzhen Zhang, Meihua Zhou, Dandan Tang, Wenzhuo Liu, Qiaoning Yang, Tianyu Shen, Kunfeng Wang, Huaping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | UMD-Net: A Unified Multi-Task Assistive Driving Network Based on Multimodal FusionabstractIn recent years, researchers have focused on identifying tasks related to driver state, traffic environment, and others to enhance the safety of autonomous driving assistance systems. However, current research on these tasks is conducted independently, neglecting the interconnections between the driver, traffic environment, and vehicle. In this paper, we propose a Unified Multi-task Assistive Driving Network Based on Multimodal Fusion (UMD-Net), the first unified model capable of recognizing four tasks simultaneously by utilizing multimodal data: driver behavior recognition, driver emotion recognition, traffic context recognition, and vehicle behavior recognition. In order to better enhance the synergistic effects between multiple tasks, we designed the position-sensitive multi-directional attention feature extraction subnetwork and recursive dynamic feature fusion module. The former captures the key features of multi-view images by different directions of attention mechanism to improve the generalization of the model across multiple tasks. The latter dynamically adjusts the fusion weight according to the multimodal features to enhance the representation ability of important features in multi-task learning. Our model was evaluated on the public dataset AIDE, achieving the best performance across all four tasks and a high accuracy of 95.31% in the traffic context recognition task, demonstrating the superiority of our approach. The code is available on https://github.com/Wenzhuo-Liu/UMD-Net. Wenzhuo Liu, Yicheng Qiao, Zhiwei Li 0011, Wenshuo Wang 0001, Wei Zhang 0012, Jiayin Zhu, Yanhuan Jiang, Li Wang 0092, Hong Wang 0014, Huaping Liu 0001, Kunfeng Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | SGV3D: Toward Scenario Generalization for Vision-Based Roadside 3D Object DetectionabstractRoadside perception can significantly enhance the safety of autonomous vehicles by extending their perceptual capabilities beyond the visual range and addressing occluded regions. However, current state-of-the-art vision-based roadside detection methods exhibit high accuracy on labeled scenes but perform poorly on new scenes. This limitation arises because roadside cameras remain stationary after installation and can only gather data from a single scene, leading the algorithm to overfit these roadside backgrounds and camera positions. To tackle this issue, we propose an innovativeScenarioGeneralization Framework forVision-based Roadside3DObject Detection, calledSGV3D. Specifically, we utilize a Background-suppressed Module (BSM) to reduce background overfitting in vision-centric pipelines by diminishing background features during the 2D to bird’s-eye-view projection. Furthermore, by introducing the Semi-supervised Data Generation Pipeline (SSDG) that employs unlabeled images from new scenes, we generate diverse foreground instances with varying camera poses, mitigating the risk of overfitting to specific camera positions. Experiments conducted on two large-scale roadside benchmarks demonstrate that SGV3D, with only a minimal increase in latency, effectively improves the scenario generalization capabilities of vision-based roadside 3D object detectors. The code is available here (https://github.com/yanglei18/SGV3D). Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Zhiwei Li 0011, Yang Shen 0005, Chen Lv 0001, Hong Wang 0014 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | SIFDriveNet: Speed and Image Fusion for Driving Behavior Classification NetworkabstractDriving behavior classification is an important direction in the field of social transportation systems and advanced driving assistance system (ADAS), which has attracted more and more attention in recent years. An accurate driving behavior classification algorithm plays a great role in traffic safety, energy saving, and other fields. In this article, we propose a novel vehicle speed and image fusion for driving behavior classification network (SIFDriveNet), which classifies driver behaviors into normal driving, aggressive driving, and drowsy driving. Our method has the following key advantages. First, in the research of driving behavior classification, we are the first to introduce a 2-D image with rich roadside information and convert speeds into a 2-D spectrogram expressing the time–frequency characteristics of speeds through short-time Fourier transform (STFT) while unifying the data space of image information and speed information. Second, we propose a tensor fusion method based on weight decomposition to fully fuse the vectors of the two modalities. This method maps the tensor outer product results to the low-dimensional space through weight decomposition and has a low computational cost while maintaining the fusion effect of the tensor outer product. In addition, we evaluated our model on the public UAH-DriveSet and compared it with the most advanced model. Experimental results show that our model has a better performance, and F1-score is 97.9% on all roads. Especially on the secondary road, our F1-score is 99.4%. Also, our model has strong generalization, and we have reached 99.3% F1 in distracted driving multimodal dataset. In addition, the inference speed reaches 411 FPS, enabling real-time needs. The code is available onhttps://github.com/alu222/SIFDriveNet. Jianli Lu, Wenzhuo Liu, Zhiwei Li 0011, Xinmin Jiang, Xin Gao 0028, Xingang Wu |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | FMDNet: Feature-Attention-Embedding-Based Multimodal-Fusion Driving-Behavior-Classification NetworkabstractDriving behavior classification is a critical component of social transportation systems and advanced driver assistance systems, and it has gained increasing attention in recent years. Accurate classification algorithms for driving behavior play a significant role in enhancing traffic safety, energy conservation, and related fields. In this article, we propose a novel driving behavior classification network named feature-attention-embedding-based multimodal-fusion driving-behavior-classification network (FMDNet). FMDNet incorporates eight types of data, including acceleration along the x-axis, y-axis, z-axis, roll angle, pitch angle, yaw angle, roadside image, and vehicle speed, to classify driving behavior. To effectively fuse features extracted from different modalities, taking into account their varying importance, we introduce the feature attention embedding-based fusion module (FAEF) as our fusion strategy. This fusion strategy enhances the network's capability to capture meaningful features by incorporating two feature attention embedding units that delve deeper into the interplay between different modes. Furthermore, we provide further validation of the effectiveness of our approach through extensive ablation experiments to investigate and analyze the impact of various modal data on the classification of driving behavior. Our proposed FMDNet achieves state-of-the-art performance on the public UAH-DriveSet dataset, demonstrating its effectiveness with an impressive F1-score of 99.0%. Additionally, the robustness of our model is confirmed on distracted dataset, achieving a remarkable F1-score of 99.7%. The model's outstanding performance on both the UAH-DriveSet dataset and the distracted-dataset highlights its capabilities and potential for real-world applications.https://github.com/Wenzhuo-Liu/FMDNet Wenzhuo Liu, Jianli Lu, Junbin Liao, Yicheng Qiao, Guoying Zhang, Jiayin Zhu, Bozhang Xu, Zhiwei Li 0011 |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2024 | Informative Data Selection With Uncertainty for Multimodal Object DetectionabstractNoise has always been nonnegligible trouble in object detection by creating confusion in model reasoning, thereby reducing the informativeness of the data. It can lead to inaccurate recognition due to the shift in the observed pattern, that requires a robust generalization of the models. To implement a general vision model, we need to develop deep learning models that can adaptively select valid information from multimodal data. This is mainly based on two reasons. Multimodal learning can break through the inherent defects of single-modal data, and adaptive information selection can reduce chaos in multimodal data. To tackle this problem, we propose a universal uncertainty-aware multimodal fusion model. It adopts a multipipeline loosely coupled architecture to combine the features and results from point clouds and images. To quantify the correlation in multimodal information, we model the uncertainty, as the inverse of data information, in different modalities and embed it in the bounding box generation. In this way, our model reduces the randomness in fusion and generates reliable output. Moreover, we conducted a completed investigation on the KITTI 2-D object detection dataset and its derived dirty data. Our fusion model is proven to resist severe noise interference like Gaussian, motion blur, and frost, with only slight degradation. The experiment results demonstrate the benefits of our adaptive fusion. Our analysis on the robustness of multimodal fusion will provide further insights for future research. Xinyu Zhang 0001, Zhiwei Li 0011, Zhenhong Zou, Xin Gao 0028, Yijin Xiong, Dafeng Jin, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | CAMO-MOT: Combined Appearance-Motion Optimization for 3D Multi-Object Tracking With Camera-LiDAR Fusionabstract3D Multi-object tracking (MOT) ensures consistency during continuous dynamic detection, conducive to subsequent motion planning and navigation tasks in autonomous driving. However, camera-based methods suffer in the case of occlusions and it can be challenging to track the irregular motion of objects for LiDAR-based methods accurately. Some fusion methods work well but do not consider the untrustworthy issue of appearance features under occlusion. At the same time, the false detection problem also significantly affects tracking. As such, we propose a novel camera-LiDAR fusion 3D MOT framework based on Combined Appearance-Motion Optimization (CAMO-MOT), which uses both camera and LiDAR data and significantly reduces tracking failures caused by occlusion and false detection. For occlusion problems, we are the first to propose an occlusion head to select the best object appearance features multiple times effectively, reducing the influence of occlusions. To decrease the impact of false detection in tracking, we design a motion cost matrix based on confidence scores which improve the positioning and object prediction accuracy in 3D space. As existing multi-object tracking methods always evaluate each category separately and do not consider the mismatch between objects of different categories, we also propose to build a multi-category cost to implement multi-object tracking in multi-category scenes. A series of validation experiments are conducted on the KITTI and nuScenes tracking benchmarks. Our proposed method achieves state-of-the-art performance with 79.99% HOTA and the lowest identity switches (IDS) value (23 for Car and 137 for Pedestrian) among all multi-modal MOT methods on the KITTI test dataset. And our method achieves state-of-the-art performance among all algorithms on the nuScenes test dataset with 75.3% AMOTA. Li Wang 0092, Xinyu Zhang 0001, Wenyuan Qin, Jinghan Gao, Lei Yang 0060, Zhiwei Li 0011, Jun Li 0082, Hong Wang 0014, Huaping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | IPS300+: a Challenging multi-modal data sets for Intersection Perception SystemabstractDue to high complexity and occlusion, insufficient perception in the crowded urban intersection can be a serious safety risk for both human drivers and autonomous algorithms, whereas CVIS (Cooperative Vehicle Infrastructure System) is a proposed solution for full-participants perception under this scenario. However, the research on roadside multi-modal perception is still in its infancy, and there is no open-source data sets for such scene. Accordingly, this paper fills the gap. Through an IPS (Intersection Perception System) installed at the diagonal of the intersection, this paper proposes a high-quality multi-modal data sets for the intersection perception task. The center of the experimental intersection covers an area of 3000m2, and the extended distance reaches 300m, which is typical for CVIS. The first batch of open-source data includes 14198 frames, and each frame has an average of 319.84 labels, which is 9.6 times larger than the most crowded data sets (H3D data sets in 2019) by now. Our data sets is available at: http://www.openmpd.com/column/IPS300. Huanan Wang, Xinyu Zhang 0001, Zhiwei Li 0011, Jun Li 0082, Zhu Lei, Haibing Ren |
ICRA | 3 |
| 2021 | Channel Attention in LiDAR-camera Fusion for Lane Line Segmentation
Xinyu Zhang 0001, Zhiwei Li 0011, Xin Gao 0028, Dafeng Jin, Jun Li 0082 |
Pattern Recognit. | 2 |