Zhanwen Liu

dblp:170/5800 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-8823-0833ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Learning to plan efficient robot paths via advantage-shaped reward in dynamic environment
Yisheng An, Yukun Xiao, Yaxin Wei, Zhanwen Liu, Niannian Shi, Faqin Jia
Eng. Appl. Artif. Intell.5
2026 KD-DiffSeg: knowledge distillation guided LiDAR-camera diffusion framework for 3D semantic segmentation
Wenfeng Leng, Chuan Hu 0003, Yakang Wang, Zhanwen Liu, Xi Zhang 0016
Expert Syst. Appl.5
2026 SCSMamba: Spatial-channel sparse Mamba for event-based object detection
Zhanwen Liu, Shangyu Xie, Wenyue Liu, Xiangmo Zhao
Expert Syst. Appl.2
2026 DCGM: Synergizing Differential Cross-Modal Fusion and Graph-Mamba for Referring Multi-Object Tracking
Zhibiao Xue, Zhanwen Liu, Chunmian Lin, Daniel Jian Sun
Knowl. Based Syst.2
2026 Controlled consistent diffusion network for long-tail trajectory generation
Xiangmo Zhao, Zhanwen Liu, Rongjie Yu, Naikan Ding
Knowl. Based Syst.3
2026 Not all regions are equal: Spatially adaptive representation learning for efficient visual object tracking
abstract
The sparsity of information in natural images poses great challenges for trackers to strike a balance between accuracy and efficiency. Existing methods commonly process all regions in the template and search area equally without considering their difference. As a result, considerable redundant computation is involved and limited inference efficiency is achieved. To remedy this, in this paper, we argue that not all regions are equal during the representation learning for visual object tracking. Specifically, we develop a sparse mask Transformer (SMTransformer) that is able to achieve spatially adaptive representation learning. Particularly, a deformable patch embedding module is constructed to adapt the receptive field to focus on the object in the template. In addition, sparse mask module is developed to dynamically identify regions with low object existence probabilities, thereby reducing the search region progressively for higher computational efficiency. With these two modules, our SMTransformer can significantly reduce the redundant computation for superior efficiency while maintaining high performance. Extensive experiments are conducted on a wide range of benchmark datasets and the results demonstrate the state-of-the-art performance of the proposed SMTransformer against previous methods in terms of both accuracy and efficiency.
Hongke Xu, Zhanwen Liu, Longguang Wang
Neural Networks4
2025 Towards Real-World Event-Guided Motion Deblurring
Zhanwen Liu, Yang Wang 0015, Shangyu Xie, Huanna Song
WISA2
2025 SMamba: Sparse Mamba for Event-based Object Detection
abstract
Transformer-based methods have achieved remarkable performance in event-based object detection, owing to the global modeling ability. However, they neglect the influence of non-event and noisy regions and process them uniformly, leading to high computational overhead. To mitigate computation cost, some researchers propose window attention based sparsification strategies to discard unimportant regions, which sacrifices the global modeling ability and results in suboptimal performance. To achieve better trade-off between accuracy and efficiency, we propose Sparse Mamba (SMamba), which performs adaptive sparsification to reduce computational effort while maintaining global modeling capability. Specifically, a Spatio-Temporal Continuity Assessment module is proposed to measure the information content of tokens and discard uninformative ones by leveraging the spatiotemporal distribution differences between activity and noise events. Based on the assessment results, an Information-Prioritized Local Scan strategy is designed to shorten the scan distance between high-information tokens, facilitating interactions among them in the spatial dimension. Furthermore, to extend the global interaction from 2D space to 3D representations, a Global Channel Interaction module is proposed to aggregate channel information from a global spatial perspective. Results on three datasets (Gen1, 1Mpx, and eTram) demonstrate that our model outperforms other methods in both performance and efficiency.
Yang Wang 0015, Zhanwen Liu, Meng Li 0017, Yisheng An, Xiangmo Zhao
AAAI3
2025 Multi-class Agent Trajectory Prediction with Selective State Spaces for autonomous driving
Zhanwen Liu
Eng. Appl. Artif. Intell.2
2025 Modified You Only Look Once Network Model for Enhanced Traffic Scene Detection Performance for Small Targets
abstract
ABSTRACT In order to address the challenge of small target recognition in traffic scenes, we propose a model based on you only look once version 8X (Yolov8X) network model, which has been combined with receptive fields block (RFB) and multidimensional collaborative attention (MCA). First, the model employs the RFB to extract reliable and distinctive features, thereby enhancing the precision of small target identification. Furthermore, the MCA structure is introduced to simulate multidimensional attention through three parallel branches, thereby enhancing the feature expression ability of the model. This fragment describes a compression transformation and an excitation transformation that captures the differentiated feature representation of the command. These transformations facilitate the network's ability to locate and predict the location of small objects more accurately. Utilizing these transformations enhances the expressiveness and diversity of features, thereby improving the detection performance of small objects. Furthermore, data augmentation and hyperparameter optimization techniques are employed to enhance the model's generalisability. The validation results on the Argoverse 1.1 autonomous driving dataset demonstrate that the enhanced network model outperforms the prevailing detectors, achieving an F1 score of 78.6, an average precision of 55.1, and an average recall of 72.4. The algorithm's excellent performance for small target detection was demonstrated through visual analysis, proving its high application value and potential for promotion in fields such as autonomous driving.
Shuai Ren 0001, Ke Wang 0058, Zhanwen Liu
IET Image Process.6
2025 Scenario-Based Accelerated Testing for SOTIF in Autonomous Driving: A Review
abstract
The development of intelligent driving systems has drawn significant attention to enhancing the safety of autonomous vehicles and their intended functionality. Despite this, current accelerated testing approaches remain inadequate in assessing system reliability, as they fail to simulate scenarios involving collisions between vehicles and pedestrians and identify unknown risks. To address these limitations, scenario-based testing methods have been proposed, which seek to identify critical scenarios with a high frequency of exposure to safety risks. A comprehensive review of these methods is thus of paramount significance. In this article, we provide a timely and systematic literature review of existing accelerated testing for autonomous vehicles. We propose a taxonomy of these methods, discuss each subfield, and highlight open problems and future directions. Our objective is to provide a clear and concise overview of the state of the art in this field and to offer insights into the effectiveness of scenario-based testing approaches. By doing so, we aim to facilitate the identification of critical scenarios and the assessment of risk exposure frequencies, which are essential for enhancing the safety and reliability of autonomous vehicles.
Lei Tang 0002, Zhanwen Liu, Yunji Liang, Yuanyuan Niu, Wei Zhu 0004, Zongtao Duan
IEEE Internet Things J.3
2025 Multi-Modal Fusion Based on Depth Adaptive Mechanism for 3D Object Detection
abstract
Lidars and cameras are critical sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, accurate and robust fusion methods are still under exploration due to non-homogenous representations. In this paper, we find that the complementary roles of point clouds and images vary with depth. An important reason is that the point cloud appearance changes significantly with increasing distance from the Lidar, while the image's edge, color, and texture information are not sensitive to depth. To address this, we propose a fusion module based on the Depth Attention Mechanism (DAM), which mainly consists of two operations: gated feature generation and point cloud division. The former adaptively learns the importance of bimodal features without additional annotations, while the latter divides point clouds to achieve differential fusion of multi-modal features at different depths. This fusion module can enhance the representation ability of original features for different point sets and provide more comprehensive features by using the dual splicing strategy of concatenation and index connection. Additionally, considering point density as a feature and its negative correlation with depth, we build an Adaptive Threshold Generation Network (ATGN) to generate the depth threshold by extracting density information, which can divide point clouds more reasonably. Experiments on the KITTI dataset demonstrate the effectiveness and competitiveness of our proposed models.
Zhanwen Liu, Juanru Cheng, Jin Fan 0004, Yang Wang 0015, Xiangmo Zhao
IEEE Trans. Multim.1
2024 MSTF: Multiscale Transformer for Incomplete Trajectory Prediction
abstract
Motion forecasting plays a pivotal role in autonomous driving systems, enabling vehicles to execute collision warnings and rational local-path planning based on predictions of the surrounding vehicles. However, prevalent methods often assume complete observed trajectories, neglecting the potential impact of missing values induced by object occlusion, scope limitation, and sensor failures. Such oversights inevitably compromise the accuracy of trajectory predictions. To tackle this challenge, we propose an end-to-end framework, termed Multi-scale Transformer (MSTF), meticulously crafted for incomplete trajectory prediction. MSTF integrates a Multiscale Attention Head (MAH) and an Information Increment-based Pattern Adaptive (IIPA) module. Specifically, the MAH component concurrently captures multiscale motion representation of trajectory sequence from various temporal granularities, utilizing a multi-head attention mechanism. This approach facilitates the modeling of global dependencies in motion across different scales, thereby mitigating the adverse effects of missing values. Additionally, the IIPA module adaptively extracts continuity representation of motion across time steps by analyzing missing patterns in the data. The continuity representation delineates motion trend at a higher level, guiding MSTF to generate predictions consistent with motion continuity. We evaluate our proposed MSTF model using two large-scale real-world datasets. Experimental results demonstrate that MSTF surpasses state-of-the-art (SOTA) models in the task of incomplete trajectory prediction, showcasing its efficacy in addressing the challenges posed by missing values in motion forecasting for autonomous driving systems.
Zhanwen Liu, Yang Wang 0015, Jiaqi Ma 0003, Xiangmo Zhao
IV1
2024 Intention-convolution and hybrid-attention network for vehicle trajectory prediction
Zhanwen Liu, Yang Wang 0015, Xiangmo Zhao
Expert Syst. Appl.2
2024 Enhancing Traffic Object Detection in Variable Illumination With RGB-Event Fusion
abstract
Traffic object detection under variable illumination is challenging due to the information loss caused by the limited dynamic range of conventional frame-based cameras. To address this issue, we introduce bio-inspired event cameras and propose a novel Structure-aware Fusion Network (SFNet) that extracts sharp and complete object structures from the event stream to compensate for the lost information in images through cross-modality fusion, enabling the network to obtain illumination-robust representations for traffic object detection. Specifically, to mitigate the sparsity or blurriness issues arising from diverse motion states of traffic objects in fixed-interval event sampling methods, we propose the Reliable Structure Generation Network (RSGNet) to generate Speed Invariant Frames (SIF), ensuring the integrity and sharpness of object structures. Next, we design a novel Adaptive Feature Complement Module (AFCM) which guides the adaptive fusion of two modality features to compensate for the information loss in the images by perceiving the global lightness distribution of the images, thereby generating illumination-robust representations. Finally, considering the lack of large-scale and high-quality annotations in the existing event-based object detection datasets, we build a DSEC-Det dataset, which consists of 53 sequences with 63,931 images and more than 208,000 labels for 8 classes. Extensive experimental results demonstrate that our proposed SFNet can overcome the perceptual boundaries of conventional cameras and outperform the frame-based methods, e.g., YOLOX by 7.9% in mAP50 and 3.8% in mAP50:95. Our code and dataset will be available athttps://github.com/YN-Yang/SFNet.
Zhanwen Liu, Yang Wang 0015, Xiangmo Zhao, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Regional attention network with data-driven modal representation for multimodal trajectory prediction
Zhanwen Liu, Xiangmo Zhao
Expert Syst. Appl.2
2023 YOLO-F: YOLO for Flame Detection
abstract
Flame detection is of great significance in a fire prevention system. YOLOv4 has poor real-time performance on flame detection caused by the complex structure and high parameter size. To address this problem, a novel flame detection framework, YOLO for flame (YOLO-F), is proposed in this paper. The backbone of YOLOv4 is simplified from the original 53 convolutional layers to 34 convolutional layers to reduce the number of parameters by simplifying the structure of the CSPBlock. Based on the FPN, an effective and light-weight feature pyramid architecture, namely FPNs-SE, is then proposed and the neck part of YOLOv4 is replaced by FPNs-SE to enhance the feature extraction ability of different scales. In addition, the CIoU loss in the YOLOv4 ignores the similarity measure of the area between the predicted bounding box and the ground-truth bounding box. An effective loss named ACIoU is proposed in this paper in order to handle the above issue and further improve the detection accuracy. The proposed methods are tested on FLAME dataset and network crawled dataset, respectively. The mAP, recall, and precision of YOLO-F are higher by 2.01%, 4.0%, 2.0% on average than those of YOLOv4. With input size of [Formula: see text] and on a single GTX 1660, the operating speed of our method can reach 24.53[Formula: see text]fps, which is improved by 38.04% compared with YOLOv4. The experimental results show that our method is more robust to the small flame and flame-like objects and can achieve the best balance of detection speed and accuracy. The code is made available at https://github.com/Windxy/YOLO-F .
Yuan Xu 0012, Yuanxin Xing, Zhanwen Liu
Int. J. Pattern Recognit. Artif. Intell.4
2019 Restoration algorithm for noisy complex illumination
abstract
Although promising results have been achieved in the restoration of complex illumination images with the Retinex algorithm, there are still some drawbacks in the processing of Retinex. Considering the noise characteristics of complex illumination images, in this study, we propose a novel restoration algorithm for noisy complex illumination, which combines guided adaptive multi‐scale Retinex (GAMSR) and improvement BayesShrink threshold filtering (IBTF) based on double‐density dual‐tree complex wavelet transform (DDDTCWT) domain. Extensive restoration experiments are conducted on three typical types images and the same image with different noises. On the basis of a series of evaluation indexes, we compare our method to those of state‐of‐the‐art algorithms. The results show that (i) SSIM of the proposed IBTF is superior to traditional Bayes threshold method by 15% as the standard variance is 100. (ii) PSNR of the proposed GAMSR enhances 15% to traditional MSR. (iii) The clarity of final results for restoration speeds up three times than that of original images, and the information entropy is improved slightly too. Therefore, the proposed method can effectively enhance the details, edges and textures of the image under complex illumination and noises.
Zhanwen Liu, Tao Gao 0001, Fanjie Kong, Ziheng Jiao, Aodong Yang, Bo Liu 0006
IET Comput. Vis.1
2019 Combination of modified U-Net and domain adaptation for road detection
abstract
Road detection is one of the crucial tasks for scene understanding in autonomous driving. Recently, methods based on deep learning had rapidly grown and addressed this task excellently, because they can extract more abundant features. In this study, the authors consider the visual road detection problem as a classification for each pixel of the given image, which is road or non‐road. There is complex illumination encounter in traffic applications, so that the detection model has poor adaptability. They address this problem by proposing a deep network architecture, which combines the network U‐Net‐prior and domain adaptation model (DAM). U‐Net‐prior is a modified segmentation network which integrates location prior and shape prior into U‐Net. DAM is a model for reducing the gap between training images and test images, which is optimised in adversarial learning to make the features extracted from different datasets close to each other. They validate the effectiveness of each component of the algorithm, and compare the overall architecture with other state‐of‐the‐art methods, and the results show that the architecture achieves top accuracies with the shortest run time in monocular‐vision‐based methods, simultaneously, compared with the methods based on other sensors, the architecture also achieves a competitive result.
Xiangmo Zhao, Zhanwen Liu
IET Image Process.5
2018 Infrared and Visible Image Fusion Based on Compressive Sensing and OSS-ICA-Bases
abstract
Aimed at the problems that most existing fusion methods tolerate one or more drawback such as noise, blur and key information loss, a novel and valid fusion algorithm is proposed to efficiently extract the object information in infrared image and preserve abundant background information in visible image. Firstly, non-subsampled shearlet transform (NSST) is employed to decompose the visible and infrared images into high frequency subbands and low frequency subbands. Secondly, a fusion rule based on compressed sensing (CS) was put into high frequency subbands and a fusion rule based on online same scene independent component analysis bases (OSS-ICA-bases) was input into low frequency subbands. Finally the fusion image was reconstructed by an inverse NSST on these merged coefficients. Because the OSS-ICA-bases could suppress the noise and fuses the complementary information well, CS enables the high frequency subbands to be accurately reconstructed from fewer sparse fused coefficients, NSST can obtain the asymptotic optimal representation and has the better sparse representation ability, the proposed algorithm can obtain a better result. Experiments also show that our approach can achieve better performance than other methods in terms of subjective visual effect and objective assessment.
Zhanwen Liu
ICIP1
2017 Image feature representation with orthogonal symmetric local weber graph structure
Tao Gao 0001, Xiangmo Zhao, Ting Chen 0003, Zhanwen Liu
Neurocomputing4
2017 Illumination-insensitive image representation via synergistic weighted center-surround receptive field model and weber law
Tao Gao 0001, Xiangmo Zhao, Ting Chen 0003, Zhanwen Liu
Pattern Recognit.4
2016 Example-based super-resolution via social images
Yi Tang 0003, Hong Chen 0004, Zhanwen Liu, Biqin Song, Qi Wang 0009
Neurocomputing3