Dachuan Li

dblp:123/5741 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-7267-9951ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 SD-MoE: Scenario-Driven MoE Forecasting for Intelligent Elastic Scaling in Cloud Clusters
Xianzhao Guo, Weipeng Cao, Minxian Xu, Dachuan Li, Chuanfei Xu, Zhong Ming 0001
CCGrid4
2025 UA-PnP: Uncertainty-Aware End-to-End Bird's Eye View Visual Perception and Prediction for Autonomous Driving
abstract
Robust and accurate perception and prediction of the driving scenarios are crucial for autonomous driving vehicles (ADV). State-of-the-art ADV frameworks have evolved from conventional modular design to an end-to-end (E2E) pipeline that enables joint feature learning and optimization. However, the evaluation of uncertainties in the intermediate features propagated between perception and prediction units is missing in current E2E pipelines. Consequently, adverse and extreme environment factors may incur highly untrustworthy features that ultimately result in degraded perception and prediction. In this work, we propose a novel uncertainty-aware E2E visual perception and prediction framework that utilized Bird's Eye View (BEV) representations. A feature distribution estimation network is introduced to explicitly quantify the uncertainties in the intermediate BEV features extracted from the images. To better exploit temporal information and generate more robust features for scene prediction, an uncertainty-aware transformer is designed to utilize the guidance of the quantified feature uncertainty via the attention mechanism. In addition, an evidential decoder generates accurate future instance segmentations along with the associated uncertainties. Comprehensive experiments conducted on real-world dataset validate the superiority of our proposed framework over conventional pipelines. Codes are available at: https://github.com/Huang121381/UAPnP.
Zijian Huang 0014, Dachuan Li, Qi Hao 0003
ICRA2
2025 PA-TCP: Interpretable End-to-End Autonomous Driving Through Parallel Adaptive Attention Mechanism and State Representation
abstract
A safe and interpretable end-to-end autonomous driving system is essential for real-world applications. However, existing methods struggle with incomplete feature understanding, the black box problem, and poor interpretability, making it hard to adapt to complex environments and be accepted by users. In this study, we propose an end-to-end autonomous driving framework, PA-TCP, which enhances safety and interpretability through a hybrid attention mechanism and efficient state representation. Specifically, we introduce a parallel-weighted compound attention module that dynamically captures and prioritizes critical environmental features for vehicle driving. This module leverages a parallel architecture to simultaneously combine spatial and channel attention mechanisms through learned adaptive weights, enabling fine-grained feature selection and robust scene understanding in challenging scenarios. Next, we integrate vehicle dynamics, navigation commands, and contextual information through a linear-based Squeeze-and-Excitation attention framework, which systematically identifies and emphasizes the most task-relevant features while achieving a balance between representation capability and computational overhead. Extensive experiments on the CARLA simulation platform demonstrate the superiority of our approach over the baseline method TCP, including a 25.08% increase in driving score, a 16.2% increase in route completion, and a 6.8% increase in infraction score. We also demonstrate its effectiveness regarding generalization capabilities.
Dongzhuo Wang, Yang Li 0093, Weisi Chen, Yao Mu 0001, Dachuan Li
IV6
2025 Risk-Aware Reinforcement Learning for Non-Conservative Motion Planning in Uncertain Autonomous Driving Environments
abstract
Reinforcement learning (RL) offers a powerful paradigm for adaptive motion planning in complex driving environments. However, applying RL to autonomous driving remains challenging due to uncertainty from partial observability and the stochastic, multimodal behaviors of traffic participants. This paper presents a novel risk-aware RL framework for non-conservative motion planning under uncertainty. By integrating Partially Observable Markov Decision Processes (POMDP) with a deep RL-based policy optimization scheme, the proposed approach explicitly models aleatoric uncertainty via a Gaussian Mixture Bayesian Belief Updater and a time-varying risk field. Additionally, an Adaptive Context-aware Attention (ACA) module is employed to prioritize critical targets for enhanced interaction modeling dynamically. Extensive experiments on the CARLA simulator show that the framework generalizes well across diverse traffic conditions, improving average reward by 65.74% and 64.02% in low-speed dense and high-speed sparse scenarios. It remains robust in challenging situations such as overtaking and sudden lane changes in the PeMS dataset. Furthermore, distributed deployment tests confirm a real-time performance of 10 Hz on a hardware-in-the-loop platform, demonstrating the feasibility of practical deployment.
Chuan Hu 0003, Dongang Liu, Dachuan Li, Jinxiang Wang 0002, Xiaolin Tang
IEEE Trans. Intell. Transp. Syst.4
2025 Decision Making of Automated Vehicles in Mixed Environment Based on Bayesian Sequential Games
abstract
Automated Vehicles (AVs) will coexist with Human-Driven Vehicles (HDVs) for a long time. AVs must navigate safely among HDVs while maintaining smooth traffic flow. To facilitate this, the decision making system of AVs must accurately assess HDV intentions while accounting for inherent uncertainties. Current HDV intention prediction models often misclassify these intentions, leading to unsafe navigation decisions. This study introduces a three-stage Bayesian sequential game-based decision making architecture designed for AV operation. In the first stage, the AV utilizes a temporal neural network to classify vehicle intentions. In the second stage, a sequential game is solved to determine optimal actions by predicting future HDV states. The final stage, serving as a validation stage, identifies and corrects misclassifications from the first stage by predicting HDV future positions, incorporating models that account for potential deviations from the ground truth. Simulation results indicate a 93.5±0.5% accuracy in initial intention predictions, facilitating swift and effective decision making. The validation stage further enhances safety by promptly correcting errors, ensuring reliable navigation for AVs in HDV environments.
Harikrishnan Vijayakumar, Dezong Zhao, Jianglin Lan, David Flynn, Dachuan Li, Quan Zhou 0006, Yuanjian Zhang 0001
IEEE Trans. Intell. Transp. Syst.6
2024 Depth-Aware Multi-Modal Fusion for Generalized Zero-Shot Learning
abstract
Realizing Generalized Zero-Shot Learning (GZSL) based on large models is emerging as a prevailing trend. However, most existing methods merely regard large models as black boxes, solely leveraging the features output by the final layer while disregarding potential performance enhancements from other layers. Indeed, numerous researchers have visually depicted variations in the features learned across different layers of neural networks. Motivated by this observation, we propose a Vision Transformer (ViT)-based GZSL method named Depth-Aware Multi-Modal ViT (DAM2ViT), which exploits multi-level features of ViT. DAM2ViT incorporates a multi-modal interaction block to align semantic information of categories across multiple layers, thereby augmenting the model's capacity to learn associations between visual and semantic spaces. Extensive experiments conducted on three benchmark datasets (i.e., CUB, SUN, AWA2) have showcased that DAM2ViT achieves competitive results compared to state-of-the-art methods.
Weipeng Cao, Xuyang Yao, Zhiwu Xu 0001, Yinghui Pan, Yixuan Sun, Dachuan Li, Bohua Qiu, Muheng Wei
INDIN6
2024 ICSGD-Momentum: SGD Momentum Based on Inter-gradient Collision
abstract
Deep neural networks (DNNs) are widely used in fields like computer vision and natural language processing. A key component of DNN training is the optimizer. SGD-Momentum is popular in many DNN methodologies, such as ResNet and DenseNet, due to its simplicity and effectiveness. However, its slow convergence rate limits its use. To overcome this, we introduce inter-gradient collision into SGD-Momentum, inspired by the elastic collision model in physics. This new method, called ICSGD-Momentum, aims to improve convergence. We provide theoretical proof of convergence and establish a regret bound for ICSGD-Momentum. Experiments on benchmarks including function optimization, CIFAR-100, ImageNet, Penn Treebank, COCO, and YCB-Video show that ICSGD-Momentum accelerates training and enhances the generalization performance of DNNs compared to optimizers like SGD-Momentum, Adam, Radam, Adabound, and AdaBelief.
Weidong Zou, Weipeng Cao, Yuanqing Xia, Bineng Zhong 0001, Dachuan Li
INDIN5
2023 Sustainable and Transferable Traffic Sign Recognition for Intelligent Transportation Systems
abstract
Traffic Sign Recognition (TSR) is an essential component of Intelligent Transportation Systems (ITS) and intelligent vehicles. TSR systems based on deep learning have grown in popularity in recent years. However, since these models belong to the closed-world-oriented learning paradigm, they are only capable of accurately identifying traffic signs that are easy to collect and cannot adapt to the real world. Furthermore, the sample utilization of these methods is insufficient, the resource consumption of model training may become unbearable as the data scale grows. To address this problem, we propose a novel “knowledge + data” co- driven solution (i.e., Joint Semantic Representation algorithm, JSR) for TSR. JSR creates a hybrid feature representation by extracting general and principal visual features from traffic sign images. It also realizes the model’s reasoning ability to zero-shot TSR based on prior knowledge of traffic sign design standards. The effectiveness of JSR is demonstrated by experiments on four benchmark datasets and two self-built TSR datasets.
Weipeng Cao, Yuhao Wu 0001, Chinmay Chakraborty, Dachuan Li, Liang Zhao 0004, Soumya K. Ghosh 0001
IEEE Trans. Intell. Transp. Syst.4
2022 Runtime Safety Assurance for Learning-enabled Control of Autonomous Driving Vehicles
abstract
Providing safety guarantees for Autonomous Vehicle (AV) systems with machine-learning based controllers remains a challenging issue. In this work, we propose Simplex-Drive, a framework that can achieve runtime safety assurance for machine-learning enabled controllers of AVs. The proposed Simplex-Drive consists of an unverified Deep Reinforcement Learning (DRL)-based advanced controller (AC) that achieves desirable performance in complex scenarios, a Velocity-Obstacle (VO) based baseline safe controller (BC) with provably safety guarantees, and a verified mode management unit that monitors the operation status and switches the control authority between AC and BC based on safety-related conditions. We provide a formal correctness proof of Simplex-Drive and conduct a lane-changing case study in dense traffic scenarios. The simulation experiment results demonstrate that Simplex-Drive can always ensure the operation safety without sacrificing control performance, even if the DRL policy may lead to deviations from the safe status.
Shengduo Chen, Yaowei Sun, Dachuan Li, Qiang Wang 0020, Qi Hao 0003, Joseph Sifakis
ICRA3
2022 BLSHF: Broad Learning System with Hybrid Features
Weipeng Cao, Dachuan Li, Meikang Qiu
KSEM (2)2
2021 Unit-Modulus Wireless Federated Learning Via Penalty Alternating Minimization
abstract
Wireless federated learning (FL) is an emerging machine learning paradigm that trains a global parametric model from distributed datasets via wireless communications. This paper proposes a unit-modulus wireless FL (UMWFL) framework, which simultaneously uploads local model parameters and computes global model parameters via optimized phase shifting. The proposed framework avoids sophisticated baseband signal processing, leading to both low communication delays and implementation costs. A training loss bound is derived and a penalty alternating minimization (PAM) algorithm is proposed to minimize the nonconvex nonsmooth loss bound. Experimental results in the Car Learning to Act (CARLA) platform show that the proposed UMWFL framework with PAM algorithm achieves smaller training losses and testing errors than those of the benchmark scheme.
Shuai Wang 0004, Dachuan Li, Rui Wang 0007, Qi Hao 0003, Yik-Chung Wu, Derrick Wing Kwan Ng
GLOBECOM2
2020 SUSTech POINTS: A Portable 3D Point Cloud Interactive Annotation Platform System
abstract
The major challenges of developing 3D point cloud annotation systems for autonomous driving datasets include convenient user-data interfaces, efficient operations on geometric data units, and scalable annotation tools. This paper presents a Portable pOint-cloud Interactive aNnotation plaTform System (i.e. SUSTech POINTS), which contains a set of user-friendly interfaces and efficient annotation tools to help achieve high-quality data annotations with high efficiency. The novelty of this work is threefold: (1) developing a set of visualization modules for fast annotation error localization and convenient annotator-data interactions; (2) developing a set of interactive tools for annotators labeling 3D point clouds and 2D images in high speed; (3) developing an annotation transfer method to label the same objects in different data frames. The developed POINTS system is tested with public datasets such as KITTI and a private dataset (SUSTech SCAPES). The experimental results show that the developed platform can help improve the annotation accuracy and efficiency compared with using other open-source annotation platforms.
E. Li, Dachuan Li, Xiangbin Wu, Qi Hao 0003
IV4
2020 Safe and efficient collision avoidance control for autonomous vehicles
abstract
We study a novel principle for safe and efficient collision avoidance that adopts a mathematically elegant and general framework making as much as possible abstraction of the controlled vehicle’s dynamics and of its environment. Vehicle dynamics is characterized by pre-computed functions for accelerating and braking to a given speed. Environment is modeled by a function of time giving the free distance ahead of the controlled vehicle under the assumption that the obstacles are either fixed or are moving in the same direction. The main result is a control policy enforcing the vehicle’s speed so as to avoid collision and efficiently use the free distance ahead, provided some initial safety condition holds.The studied principle is applied to the design of a synchronous controller. We show that the controller is safe by construction. Furthermore, we show that the efficiency strictly increases for decreasing granularity of discretization. We present the implementation and experimental evaluations in the Carla autonomous driving simulator and investigate various performance issues.
Qiang Wang 0020, Dachuan Li, Joseph Sifakis
MEMOCODE2
2018 Head pose estimation with neural networks from surveillant images
abstract
Estimating head pose of pedestrians is a crucial task in autonomous driving system. It plays a significant role in many research fields, such as pedestrian intention judgment and human-vehicle interaction, etc. While most of the current studies focus on driver’s-view images, we reckon that surveillant images are also worthy of attention since more global information can be obtained from them than driver’s-view images. In this paper, we propose a method for head pose estimation from surveillant images. This approach consists of two stages, head detection and pose estimation. Since the head of pedestrian takes up a very small number of pixels in a surveillant image, a two-step strategy is used to improve the performance in head detection. Firstly, we train a model to extract body region from the source image. Secondly, a head detector is trained to locate head position from the extracted body regions. We use YOLOv3 as our detection network for both body and head detection. For head pose estimation, we treat it as classification task of 10 categories. We use ResNet-50 as the backbone of the classifier, of which the input is the result of head detection. A serial of experiments demonstrate the good performance of our proposed method.
Yichao Cai 0001, Xiao Zhou 0006, Dachuan Li, Yifei Ming, Xingang Mou
ICMV3
2013 Invariant Observer Design of a RGB-D Aided Inertial System for MAV in GPS-Denied Environments
abstract
This paper presents an non-linear observer framework that uses low-cost inertial measurement units (IMU) and RGB-D sensor to provide position, attitude and velocity estimates for micro aerial vehicles (MAV) in GPS-denied environments. The data fusion of inertial measurements and RGB-D motion estimates is performed through the observer, which is based on the invariant observer approach for symmetry possessing systems. The gains of the invariant observer are computed using an adaption of the invariant Extended Kalman Filter (IEKF). In addition, a robust RGB-D odometry algorithm is proposed to estimate the relative motions using successive images captured by the RGB-D sensor, which is used as an aiding measurement for accurate state estimation. The proposed approach guarantees a simplified form of estimation error dynamics, as well as simplified calculation of observer gains. The resulting observer is implemented on a quad rotor MAV and successfully validated through indoor flight tests. Experimental results demonstrate that the proposed observer is effective in estimating the motion states of the MAV in GPS-denied environments.
Dachuan Li, Qing Li 0010, Nong Cheng, Jingyan Song, Qin-fan Wu, Liangwen Tang
SMC1