Shuang Li 0016

dblp:43/6294-16 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-4215-2195ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 End-Point Drive and Reverse Enhanced Decoding- Based Traffic Participants Trajectory Prediction Under Bird's Eye View
abstract
Trajectory prediction under Bird’s Eye View (BEV) refers to predicting the movement intention of agents based on the historical observation trajectories, which is of great significance for autonomous driving, driving safety and social navigation. The traffic prediction trajectory is multimodality with multiple reasonable trajectories and various prediction time, which often suffer complex agents interaction, cumulative errors and a large number of agents. To overcome these problems, we explore the BEV-based traffic participants trajectory prediction problem and propose the novel End-point Drive and Reverse Enhanced Decoding Network (EDRED-TPNet), based on the mechanisms of end-point driving and reverse enhanced decoding. Firstly, the Encoder based on Dynamic Spatio-temporal Graph and Multimodality Coding Fusion (ST-MC-Encoder) are constructed to effectively represent complex traffic scenario with changeable agents, encode social interactions with historical trajectories, social interaction, future trajectory and multimodality. Secondly, the End-Point Drive Module is proposed to predict the end point before predicting the complete trajectory, thus providing more accurate trajectory prediction; Lastly, to further improve the long-term prediction performance, the Reverse Enhanced Decoder (RE-Decoder) is proposed to fuse forward and reverse hidden state vectors to obtain diverse trajectories that conform to physical and social acceptability rules. We build the first AAV-captured Trajectory Prediction Dataset (UTP-Dataset) for traffic participants trajectory prediction. Experimental results show that the proposed methods can fulfill the multi-target trajectory prediction task in complex traffic scenarios and achieve high performance.
Chunsheng Liu 0001, Jincan Xie, Faliang Chang, Shuang Li 0016, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.5
2025 UAV-Based Vehicle Re-Identification via Counterfactual Attention Learning and Hard-Sensitive Binomial Joint Promotion
abstract
Autonomous Aerial Vehicles (AAVs) based Vehicle Re-Identification (ReID) brings high flexibility to the ReID system, and also brings challenges of complicated shooting views and special occlusions. In this study, we propose a novel framework for UAV-based vehicle ReID, calledCounterfactual Attention and Joint Promotion based ReID(CAJP-ReID), to deal with fine-grained feature extraction and occlusions. Based on the mechanism of using randomly generated counterfactual attention intervention to train, theCounterfactual Attention Learning Network(CAL-Net) is proposed to learn the fine-grained features of vehicle images, for distinguishing similar vehicles in different ID categories. In order to enhance the diversity of the dataset and the robustness of the network to occluded images, theCounterfactual Attention Enhancement Module(CAE-Module) is proposed based on counterfactual attention mechanism, by the data-enriching mechanism of cropping and erasing. TheHard-sensitive Binomial Joint Promotion Loss(HBJP-Loss) is proposed to comprehensively consider the relative distance and absolute distance of positive and negative samples of vehicle images, and further improves the accuracy of vehicle re-identification. Experiments on two public datasets show that the proposed method achieves State-of-the-Art performance.
Chunsheng Liu 0001, Baoqi Xue, Shuang Li 0016, Faliang Chang, Nanjun Li, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.3
2024 Lightweight-Shaped Object Grasping Detection Network Based on Feature Fusion
abstract
Robotic grasping techniques for regular targets with known shapes are now well established. However, unknown shaped objects have complex features such as texture, shape, and appearance, which leads to inaccurate recognition and localization of shaped objects during grasp detection. To improve the generalization ability of the grasping detection network for unfamiliar shaped objects, we propose a lightweight shaped object grasping detection network (LSOGD) based on feature fusion, which solves the problem that the network repeatedly extracts features from images and ultimately improves the accuracy of model detection by combining different features. The effectiveness of LSOGD is confirmed by performance evaluation on the Cornell dataset and Jacquard dataset, where the detection accuracy reaches 97.9% and 96.7% for unknown objects, respectively. In addition, due to the small proportion of shaped objects in the current publicly available dataset, we added a portion of industrial-shaped pieces based on the selection of some shaped objects in the Cornell grasping dataset to build a shaped object dataset named X-Cornell on which the accuracy of our proposed model for grasping and detecting the unknown shaped objects is 94.6%. Finally, an actual robot grasping experiment was conducted using a Realsense d435i camera and a Kinova robotic arm, and the success rate of grasping shaped objects was 94%.
Peng Zhang 0105, Yupei Xing, Shuang Li 0016, Dongri Shan
Int. J. Pattern Recognit. Artif. Intell.3
2024 Transfer and supplement AdaBoost for extracting region proposals of CNN in transfer-learning application
Shuang Li 0016, Chunsheng Liu 0001
Multim. Tools Appl.1
2024 Drone-captured vehicle re-identification via perspective mask segmentation and hard sample learning
Chunsheng Liu 0001, Baoqi Xue, Shuang Li 0016, Faliang Chang
Multim. Tools Appl.3
2024 Traffic Scenario Understanding and Video Captioning via Guidance Attention Captioning Network
abstract
Describing a traffic scenario from the driver’s perspective is a challenging process for Advanced Driving Assistance System (ADAS), involving different sub-tasks of detection, tracking, segmentation, etc. Previous methods mainly focus on independent sub-tasks and have difficulties to comprehensively describe the incidents. In this study, this problem is novelly treated as a video captioning task, and a Guidance Attention Captioning Network (GAC-Network) structure is proposed for describing the incidents in a concise single sentence. In GAC-Network, an Attention based Encoder-Decoder Net (AED-Net) is built as the main network; with the temporal spatial attention mechanisms, the AED-Net make it possible to effectively reject the unimportant traffic behaviors and redundant backgrounds. Considering various driving scenarios, the Spatio-Temporal Layer Normalization is used to improve the generalization ability. To generate captions for incidents in driving, the novel Guidance Module is proposed to boost the encoder-decoder model to generate words in a caption, which have better relationship to the past and future words. Because there is no public dataset for captioning of driving scenarios, the Traffic Video Captioning (TVC) dataset is released for the video captioning task in driving scenarios. Experimental results show that the proposed methods can fulfill the captioning task for complex driving scenarios, and achieve higher performance than the methods for comparison, including at least 2.5%, 1.8%, 3.6%, and 13.1% better results on BLEU_1, METEOR, ROUGE_L and CIDEr, respectively.
Chunsheng Liu 0001, Faliang Chang, Shuang Li 0016, Penghui Hao, Yansha Lu, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.4
2022 Adaptive Short-Temporal Induced Aware Fusion Network for Predicting Attention Regions Like a Driver
abstract
Driver attention prediction can solve the problem of ‘Where should the driver pay attention?’, Most previous methods are designed to predict regional attention with redundant regions. Furthermore, popular spatial-temporal feature extraction networks such as ConvLSTM and 3D-CNN are difficult to achieve real-time. To overcome these difficulties, we propose an Adaptive Short-temporal Induced Aware Fusion Network (ASIAF-Net) for region-level and object-level driver attention prediction.1In ASIAF-Net, we design anAttention Related Spatial Feature Encoder(AF-Encoder) and anInduced Aware Fusion Network(IAF-Net) as the main network; with anAssociation Analysis Cell(AAC), the AF-Encoder makes it possible to effectively capture the relationship information of different objects. Considering most vital visual cues from moving objects, we propose aSelf-adaptive Short-temporal Feature Extraction Module(SSFE-Module) to obtain inter-frame motion features. In IAF-Net, aMulti-scale Driver Attention Region Prediction Branchis designed to predict the regional attention, and anObject Saliency Estimation Branchis proposed to fuse the perception results and the regional attention map to estimate the object-level attention. Experiments show that the proposed ASIAF-Net can predict driver’s attention on regions and objects more robustly and precisely than state-of-the-art methods on three datasets, and that it achieves real-time on our ADAS platform.
Chunsheng Liu 0001, Faliang Chang, Shuang Li 0016, Hui Liu 0040
IEEE Trans. Intell. Transp. Syst.4
2022 Temporal Shift and Spatial Attention-Based Two-Stream Network for Traffic Risk Assessment
abstract
On-board vision based traffic risk assessment is a challenging task for intelligent driving systems, which has some special challenges including spatial-temporal feature extraction, different judgements of risks, real-time requirement, lacking data, etc. To overcome these difficulties, we propose a novelTemporal Shift and Spatial Attention based Two-stream Network(TSSAT-Net) for on-board vision based traffic risk assessment. Firstly, we build new judgement measures that integrate actual driving experience and scenario complexity, and release an on-board vision based traffic risk assessment dataset. Secondly, a novelweighted Temporal Shift Module(weighted-TSM) based two-stream network is proposed; unlike previous methods that rely on complex and time-consuming 3D CNN or LSTM calculation, the proposed two-stream network can effectively extract spatial-temporal features using a weighted temporal shift mechanism with just 2D CNN calculation requirement. Thirdly, a spatial and channel attention mechanism is proposed to make the TSSAT-Net more focus on the features closely related to traffic risks, avoiding redundant information in complex traffic scenarios. Experiments based on the released comprehensive dataset show that our method achieves the state-of-the-art classification accuracy in real-time.
Chunsheng Liu 0001, Zijian Li 0013, Faliang Chang, Shuang Li 0016, Jincan Xie
IEEE Trans. Intell. Transp. Syst.4
2022 Posture Calibration Based Cross-View & Hard-Sensitive Metric Learning for UAV-Based Vehicle Re-Identification
abstract
Machine vision based vehicle re-identification (ReID) plays an important role in some Intelligent Transportation Systems (ITS). Yet, most previous methods mainly focus on fixed surveillance cameras instead of Unmanned Aerial Vehicle (UAV). With high flexibility, the UAV-based vehicle ReID problem has some special challenges including complicated shooting angles, low discrimination of top-down features, and large variance in vehicle scales, etc. To overcome these challenges, we propose a novel structure for UAV-based vehicle ReID without license plates. Firstly, a triple-head segmentation net is proposed for segmenting UAV-captured vehicles under different heights and directions. Secondly, a posture calibration model is designed to uniform the vehicle postures based on the segmentation results, with the purpose of reducing the influence of different postures. Thirdly, the novel Cross-View & Hard-Sensitive Metric Learning (CHSML) method is proposed to train a ReID network with cross-view training constraint and hard sensitive principles; the mechanism of CHSML takes the cross-view samples of same ID as a training unit to learn the potential visual relationship in cross-view and builds a hard sensitive weight matrix to make learning more focus on hard samples, which improves the low ReID accuracy brought by cross-view or hard samples. Moreover, to facilitate the research of UAV-based vehicle ReID, a large-scale UAV-based vehicle ReID dataset called VeRi-UAV is released with 17 516 vehicles of 453 IDs. The experiments show that the proposed structure gains a better performance compared with the representative methods in the UAV-based vehicle ReID task.
Chunsheng Liu 0001, Ye Song, Faliang Chang, Shuang Li 0016, Ruimin Ke, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.4
2021 Adaptive multi-level feature fusion and attention-based network for arbitrary-oriented object detection in remote sensing imagery
Luchang Chen, Chunsheng Liu 0001, Faliang Chang, Shuang Li 0016, Zhaoying Nie
Neurocomputing4
2021 Bi-Directional Dense Traffic Counting Based on Spatio-Temporal Counting Feature and Counting-LSTM Network
abstract
Machine vision based vehicle counting and traffic flow estimation are challenging problems especially for dense traffic scenarios. Previousline of interest(LOI) counting methods rarely focus on dense scenarios and their performance largely relies on the accuracy of tracking. Avoiding the use of complex tracking methods, an LOI counting framework is proposed to address the bi-directional LOI counting problem in dense scenarios. There are three main contributions. Firstly, instead of treating the LOI vehicle counting problem as a combination of detecting and tracking of individual vehicles, the bi-directional traffic flow is taken as a whole and a novelspatio-temporal counting feature(STCF) is proposed for extracting bi-directional traffic flow features in dense traffic scenarios. Secondly, without relying on a multi-target tracking process for tracking and counting each vehicle, a counting network is proposed, called thecounting Long Short-Term Memory(cLSTM) network, to do analysis of the bi-directional STCF features and vehicle counting in successive video frames. Lastly, an estimation model is designed for estimating traffic flow parameters including speed, volume and density. Experiments performed on the UA-DETRAC dataset and the captured videos show that the proposed vehicle counting method outperforms the tested representative LOI counting methods in both accuracy and speed, and that the proposed framework can efficiently estimate traffic flow parameters including speed, volume and density in real time.
Shuang Li 0016, Faliang Chang, Chunsheng Liu 0001
IEEE Trans. Intell. Transp. Syst.1