Dezong Zhao

dblp:96/8136 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
36since 2021 · last 2026
0000-0002-9848-372XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 24 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Closed-loop feedback optimization for autonomous vehicles using deep reinforcement learning
Sifan Wu 0004, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao, Daxin Tian
Expert Syst. Appl.4
2026 Co-Design of Bandwidth-Awareness Scheduling Protocol and Vehicular Platoon Control Subject to DoS Attacks Over VANETs
abstract
This paper deals with the dynamic event-triggered platoon control problem for automated vehicles subject to denial-of-service attacks over the resource-constrained vehicular adhoc networks. A unified framework is established where the longitudinal dynamics of linearized third-order vehicular platoon, the constant time headway strategy as well as the resource-constrained communication are simultaneously considered. For the resource-efficient purpose, a dynamic event-triggered mechanism is developed where the triggering thresholds are adaptively adjusted based on a bandwidth-awareness parameter α as well as the state variables to trade a balance between communication efficiency and control performance. Furthermore, a set of Bernoulli random variables are introduced to describe the nature of denial-of-service attacks. The main objective of this paper is to co-design a dynamic event-triggered mechanism and a platoon controller to ensure exponential mean-square stability andH∞ performance under resource-limited networks and denial-of-service attacks. With the aid of the established Lyapunov function, the sufficient condition of the desired platoon controller is designed, and then the corresponding parameters are derived by recurring to a set of linear matrix inequalities. Finally, a 6 vehicles platoon is employed to validate effectiveness of the developed co-design scheme.
Yinbo Gu, Dezong Zhao, Zhiquan Liu 0001
IEEE Trans Autom. Sci. Eng.3
2026 Enhancing Data Efficiency With a Trustworthy Counterfactual Generative Model
abstract
Leveraging limited data to synthesize an additional training set is essential for robotic vision, particularly in dynamic environments where collecting large datasets is impractical. Traditional robotic vision systems rely on extensive training data for object recognition and scene understanding but struggle to generalize to real-world variations, such as lighting conditions, occlusions, and sensor noise. This article proposes causal diffuse variational autoencoder (causal DiffuseVAE), a novel method integrating causal inference with high-fidelity image synthesis to generate counterfactual images. By combining the disentanglement properties of variational autoencoders (VAEs) with the generative capabilities of diffusion models, causal DiffuseVAE produces realistic, interpretable simulations of variations, such as shadows and occlusions. This combination enables data-efficient generative modeling by learning from small subsets and synthesizing missing or unseen samples. In addition, causal inference ensures that generated data follow real-world dependencies, making it robust and interpretable for deployment in unpredictable environments. Four baseline approaches are evaluated across six different datasets, demonstrating that causal DiffuseVAE consistently outperforms the four baseline approaches.
Zhaoan Ye, Dezong Zhao, Li Zhang 0013, Xidong Yan, Qinglin Bi, David Flynn
IEEE Trans. Ind. Informatics2
2026 Evaluating Scenario-Based Decision-Making for Interactive Autonomous Driving Using Rational Criteria: A Survey
abstract
Autonomous vehicles (AVs) promise substantial gains in safety, reliability, and decarbonization, yet safe and efficient interaction in dynamic, heterogeneous traffic remains a key barrier to large-scale deployment. Deep reinforcement learning (DRL) has emerged as a data-driven approach for learning adaptive decision policies that handle complex, unpredictable environments better than rule-based methods. However, different scenarios impose distinct requirements, necessitating scenario-specific algorithms. This survey systematically reviews DRL for four typical scenarios (highways, on-ramp merging, roundabouts, and unsignalized intersections), summarizes road features and recent advances, and evaluates methods using five criteria: driving safety, driving efficiency, training efficiency, unselfishness, and interpretability (DDTUI). Each DDTUI criterion is analyzed with respect to the reviewed algorithms. In addition, a dedicated scenario-centric learning transferability analysis is introduced that systematically evaluates whether each reviewed method demonstrates scene-specific learning improvements and assesses how effectively their designs transfer across the four scenarios. Finally, the challenges for future DRL-based decision-making algorithms are summarized.
Zhen Tian 0002, Dezong Zhao, David Flynn, Shuja Ansari, Chongfeng Wei
IEEE Trans. Intell. Transp. Syst.3
2026 NavDrive: Safety-Enhanced End-to-End Autonomous Driving With Navigation-Guided Diffusion Policy
abstract
Automated vehicles (AVs) are transforming urban transportation systems, as end-to-end autonomous driving models show great promise in enhancing traffic safety and operational efficiency. Despite these advances, their performance in highly interactive driving scenarios remains limited due to insufficient decision-making diversity and the absence of explicit safety guarantees. To address these challenges, we propose NavDrive, a safety-enhanced end-to-end autonomous driving framework that formulates planning as a multi-modal generative process.Specifically, NavDrive integrates navigation-based guidance into a diffusion policy. To focus on decision-critical information, a Decision-Aware Channel Fusion (DCF) module adaptively emphasizes regions involving key interactions between the ego vehicle and surrounding agents. Furthermore, a safety-aware generative planner refines trajectory samples toward feasible regions via the Target-Prior Diffusion Transformer (TDiT), which explicitly embeds physical constraints to ensure safe and human-aligned driving behaviors. Extensive experiments on the NAVSIM and nuScenes benchmarks demonstrate that NavDrive consistently outperforms existing baselines, delivering substantial gains in planning quality, safety, and robustness under complex and adverse conditions. The details will be available athttps://github.com/zgchongbo/NavDrive
Daxin Tian, Jianshan Zhou, Xuting Duan, Dezong Zhao, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.5
2025 SEMPose: A single end-to-end network for multi-object pose estimation
Shibei Xue, Dezong Zhao
Neurocomputing4
2025 An MPC-Based Distributed Bidirectional Control Strategy for Virtual Coupling With Unreliable Train-to-Train Communications
abstract
Virtual coupling (VC) is perceived to be promising in raising rail traffic capacity. In a train-to-train (T2T) based VC system, a communication network that ensures high quality of service (QoS) plays a critical role in enhancing both the coupling efficiency and the safety of the train platoon. However, unreliable communication environments characterized by issues such as time delays, packet loss, and network attacks present significant security risks to virtually coupled train sets (VCTS). How to cope with the impact caused by unstable communication and realize safe and stable VCTS formation are an important challenge for the VC system. In this paper, we propose a model predictive control (MPC) based distributed bidirectional control (DBC) strategy to tackle these challenges. We propose a control framework that integrates MPC with linear feedback-feedforward control to achieve real-time optimal control of the VC system, utilizing a bidirectional communication topology. To stabilize the VCTS, we derive local and string stability conditions to be satisfied by the controller parameters under asymmetric time-lagged unreliable networks, and utilize them as real-time constraints for the MPC controller. Furthermore, an analysis of the scalability of the proposed strategy has been conducted to improve its adaptability. Simulation results demonstrate that the proposed MPC-based DBC strategy significantly reduces the VCTS formation time and the maximum fluctuation of VCTS by 28.57% to 41.86%, and 28.84% to 52.10%, respectively, across various unreliable communication scenarios.
Daxin Tian, Jianshan Zhou, Xuting Duan, Jie Zhang 0125, Zhengguo Sheng, Dezong Zhao, Dongpu Cao
IEEE Internet Things J.7
2025 Cluster search optimisation of deep neural networks for audio emotion classification
abstract
Automated patient monitoring solutions greatly benefit from audio emotion classification, although the considerable variance in individual expression and interpretation of emotions poses a challenge. Current approaches often employ standard Audio Spectrogram Transformer (AST) and deep learning models such as Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN)-based networks. However, their performance can be enhanced by integrating neural architecture search techniques using swarm optimisation algorithms. In this research, we explore AST with hyperparameter optimisation for speech emotion recognition. Three deep learning architectures with optimisable τ b -block structures and variable filter numbers, i.e. 1DCNN, bidirectional LSTM (BiLSTM) and CNN-BiLSTM, are also proposed, enabling the optimisation of network depth and width. A novel Cluster Search Optimisation (CSO) algorithm is introduced. It incorporates Cluster Centroid Search, a Cluster Distance Improvement metric and reinforcement learning to dispatch different search actions based on clustering convergence and Q -learning strategies, respectively. A novel Noise Tempered K-means (NTKM) clustering model is also proposed with the integration of Gaussian-based noise insertion and cluster compactness-separation measurement, to further fine-tune the cluster centriods obtained using OPTICS clustering. CSO is used for hyperparameter and architecture search for AST and aforementioned deep networks. Attention mechanisms are also integrated with CSO-optimised networks to further enhance feature learning. We evaluate the resulting models against those devised by other optimisation algorithms across the EMO-DB, SAVEE, and TESS datasets. The empirical results demonstrate that CSO-optimised AST and CNN-BiLSTM with attention mechanisms outperform other architectures and yield favourable comparison results against those from existing state-of-the-art audio emotion classification methods. • Evolving transformer and deep networks are devised for audio emotion recognition. • A Cluster Search Optimisation algorithm is proposed to adapt hyperparameters. • It incorporates Noise Tempered K-means clustering and Cluster Distance Improvement. • The Q-learning algorithm is used to optimise search behaviours. • Our study indicates CSO-optimised deep networks’ effectiveness across datasets.
Sam Slade, Li Zhang 0013, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Arjun Panesar, Philip Fei Wu, Rong Gao 0001
Knowl. Based Syst.6
2025 Enhancing Negotiation Policies via Spatio-Temporal Directed Graphs for Autonomous Interaction
abstract
Mixed traffic environments, comprising autonomous vehicles (AVs) and human-driven vehicles (HDVs), present substantial challenges for developing negotiation policies. These policies are essential for enabling AVs to make adaptive decisions and achieve harmonious interactions with HDVs in dynamic and complex scenarios. Despite their potential in autonomous driving, reinforcement learning-based decision-making schemes remain insufficient in addressing interaction-critical situations. To overcome this limitation, this study proposes a graph-enhanced negotiation-aware policy optimization (GNPO) framework, which embeds interaction awareness into all key components, including representation, understanding, and response. To support interaction representation, the spatio-temporal directed graph (STDG) captures the dynamic and asymmetric nature of interactions. It integrates multiple directed graph topologies, each encoding distinct social priors, to collectively depict realistic interactive behaviors. Additionally, a multimodal feature extractor fuses environmental perception, interaction cues, and task-related information, enabling comprehensive state understanding. Building on this representation, GNPO incorporates an interaction-critical reward function and an entropy-regularized adaptive policy optimization scheme, both designed to promote harmonious and context-aware behaviors in dense traffic scenarios. The proposed framework is validated in both interaction-intensive simulated environments and real-data-based digital twin scenarios, with its effectiveness quantitatively demonstrated by substantial improvements in success rate, safety, and efficiency. Qualitative evaluations further show that GNPO is capable of generating human-like cooperative and competitive behaviors.
Chuyu Ma, Dezong Zhao, Guochen Liu
IEEE Trans Autom. Sci. Eng.3
2025 DMP: Difference-Guided Motion Prediction for Vision-Centric Autonomous Driving
abstract
Vision-centric motion prediction concentrates on accurately determining the instance mask and its future trajectory from surround-view cameras, which manifests inherent merits such as holistic perspective and fully-differentiable spirit. Nonetheless, it is still impeded by sparse bird’s-eye view (BEV) representation and unfavorable temporal context across frames, resulting in a sub-optimal solution to decision-making and vehicle navigation. In this work, we propose a novelDifference-guideMotionPrediction for vision-centric autonomous driving, that is DMP, where it integrates BEV map refinement with spatial-temporal relation modeling in a hierarchical manner. Specifically, a bidirectional view projection strategy is introduced for the complementary BEV feature generation via depth-consistency correction. To promote spatiotemporal context aggregation, we design a difference-guided motion approach by offset approximation to align motion-aware cues between adjacent frames, and a dual-stream pyramid module is further developed for historical information fusion and future instance segmentation during specific durations. Extensive experiments on the large-scale nuScenes dataset demonstrate that it outperforms the baselines by a remarkable margin and delivers competitive motion prediction across diverse scenarios and range settings, suggesting its effectiveness and superiority. The details will be available athttps://github.com/pupu-chenyanyan/DMP-VAD.
Chunmian Lin, Xuting Duan, Jianshan Zhou, Kan Guo, Dezong Zhao, Dongpu Cao, Daxin Tian
IEEE Trans. Intell. Transp. Syst.6
2025 Dynamic Game-Theoretical Decision-Making Framework for Vehicle-Pedestrian Interaction With Human Bounded Rationality
abstract
Human-involved interactive environments pose significant challenges for autonomous vehicle decision-making processes due to the complexity and uncertainty of human behavior. It is crucial to develop an explainable and trustworthy decision-making system for autonomous vehicles interacting with pedestrians. Previous studies often used traditional game theory to describe interactions for its interpretability. However, it assumes complete human rationality and unlimited reasoning abilities, which is unrealistic. To solve this limitation and improve model accuracy, this paper proposes a novel framework that integrates the partially observable markov decision process with behavioral game theory to dynamically model AV-pedestrian interactions at the unsignalized intersection. Both the AV and the pedestrian are modeled as dynamic-belief-induced quantal cognitive hierarchy (DB-QCH) models, considering human reasoning limitations and bounded rationality in the decision-making process. In addition, a dynamic belief updating mechanism allows the AV to update its understanding of the opponent’s rationality degree in real-time based on observed behaviors and adapt its strategies accordingly. The analysis results indicate that our models effectively simulate vehicle-pedestrian interactions and our proposed AV decision-making approach performs well in safety, efficiency, and smoothness. It captures key patterns of the driving behavior operated by real human drivers in virtual reality(VR) experiments and even achieves more comfortable navigation compared to our previous VR experimental data.
Meiting Dang, Dezong Zhao, Yafei Wang 0001, Chongfeng Wei
IEEE Trans. Intell. Transp. Syst.2
2025 Game Theory-Based Harmonious Decision-Making for Autonomous Bus Lane Change
abstract
Blended traffic, comprising autonomous buses (ABs) and human-driven vehicles (HDVs), is becoming increasingly common, yet the lane change decision-making for ABs remains challenging due to complex interactions with heterogeneous HDVs. To address the challenge above, this paper proposes a game theory-based harmonious decision-making (GTHD) algorithm considering nuance of driving styles of HDVs, achieving human-like performance in interactions with HDVs. Technically, a game theoretic model of the GTHD uses predictions of the opposing vehicle’s motion and the information from preplanned trajectories. Besides, a prior estimation for driving styles is obtained utilizing clustering of historical data, and refined in real time through Bayesian estimation. Then, the driving style estimation is utilized to modify the game theoretic model. The modified model provides a closer depiction of the opponent’s preferences, meanwhile adjusts self-preferences to adapt to the opponent. The efficacy of GTHD is validated using a hardware and human in loop simulator and datasets in MLC scenarios. It is shown that the GTHD achieves human-like performance with 91.50%-98.50% accuracy compared with human bus driver under different conditions, better than several lane change models based on data driven methods. The code is open source and available athttps://github.com/guofan999/GTHD.
Kaichen Jiang, Dezong Zhao, Jinbo Hao, Caimei Wang, Hui Xie 0004
IEEE Trans. Intell. Transp. Syst.5
2025 Decision Making of Automated Vehicles in Mixed Environment Based on Bayesian Sequential Games
abstract
Automated Vehicles (AVs) will coexist with Human-Driven Vehicles (HDVs) for a long time. AVs must navigate safely among HDVs while maintaining smooth traffic flow. To facilitate this, the decision making system of AVs must accurately assess HDV intentions while accounting for inherent uncertainties. Current HDV intention prediction models often misclassify these intentions, leading to unsafe navigation decisions. This study introduces a three-stage Bayesian sequential game-based decision making architecture designed for AV operation. In the first stage, the AV utilizes a temporal neural network to classify vehicle intentions. In the second stage, a sequential game is solved to determine optimal actions by predicting future HDV states. The final stage, serving as a validation stage, identifies and corrects misclassifications from the first stage by predicting HDV future positions, incorporating models that account for potential deviations from the ground truth. Simulation results indicate a 93.5±0.5% accuracy in initial intention predictions, facilitating swift and effective decision making. The validation stage further enhances safety by promptly correcting errors, ensuring reliable navigation for AVs in HDV environments.
Harikrishnan Vijayakumar, Dezong Zhao, Jianglin Lan, David Flynn, Dachuan Li, Quan Zhou 0006, Yuanjian Zhang 0001
IEEE Trans. Intell. Transp. Syst.2
2025 Efficient and Energy-Saving Cooperative Motion Planning for Multiple Connected and Autonomous Vehicles at Unsignalized Intersections
abstract
Unsignalized intersections represent a typical societal road scenario, where severe spatio-temporal conflicts occur in the central area. The indiscriminate competition for spatio-temporal resources by vehicles at intersections leads to low traffic efficiency and frequent accidents. This paper proposes a universal and unified multi-vehicle cooperative motion planning framework for intersections, coupling optimization scheduling with vehicle motion control tasks, with the aim of achieving more rational resource allocation and enhanced efficiency. Specifically, the proposed optimal control conflict-based search (OPC-CBS) algorithm constructs a conflict search tree for spatio-temporal conflict detection, and further formulates a multi-objective optimization control problem based on conflict objects. This effectively transforms the large-scale global optimization problem into a small-scale multi-stage optimization problem, achieving a balance between optimality and computational efficiency. The algorithm efficiently establishes the passing order and motion trajectories of connected and automated vehicles (CAVs) in continuous spatio-temporal domains. Simulation experiments demonstrate that the proposed algorithm can comprehensively address multiple objectives such as vehicle kinematic constraints, environmental constraints, and performance constraints in complex and dynamic scenarios. It achieves approximately an 89.06% improvement in solution efficiency while reducing energy consumption by around 17.11%.
Qi Wang 0186, Daxin Tian, Xuting Duan, Guochang Qi, Jianshan Zhou, Dezong Zhao
IEEE Trans. Intell. Transp. Syst.6
2025 Efficient Robust Model Predictive Control for Behaviorally Stable Vehicle Platoons
abstract
With increasing emphasis on vehicular automation and traffic efficiency, the management and coordination of platoon-based systems have become important. This research introduces a unique control framework based on a behavioral stability strategy, designed to enhance the cohesion of vehicle platoons and improve their ability to resist disturbances. Our approach integrates a vehicle scheduling system with a real-time platoon control mechanism to enhance the behavioral stability, robustness, and safety of the platoon. Given the heterogeneous nature of vehicles, we propose an optimal platoon formation model. This model strategically determines the number of platoons, arranges the sequence of vehicles within each platoon, and selects optimal cruising speeds to maximize platoon cohesion. To further enhance system robustness, a centralized robust model predictive controller is deployed for each platoon, ensuring stability against stochastic perturbations in vehicle dynamics and guaranteeing platoon safety. Finally, we conduct a simulation study involving multiple platoons with 20 heterogeneous vehicles to validate the effectiveness of the multi-layer optimization model.
Peiyu Zhang 0001, Daxin Tian, Jianshan Zhou, Xuting Duan, Zhengguo Sheng, Dezong Zhao, Dongpu Cao, Luzheng Bi
IEEE Trans. Intell. Transp. Syst.6
2025 CUDA-X: Unsupervised Domain-Adaptive Vehicle-to-Everything Collaboration via Knowledge Transfer and Alignment
abstract
Recently emerged vehicle-to-everything (V2X) perception has revealed great potential to overcome the limitation of single-vehicle intelligence aided by vigorous interaction among on-road agents, while prior endeavors are practically developed on parameter-specific simulation or configuration-dynamic real-world setting, overlooking the transferability across various scenarios. In this article, we propose unsupervised domain-adaptive vehicle-to-everything collaboration framework dubbed CUDA-X, which is built on top of a de facto collective model with key-point information exchange and instance adaptation. Specifically, collaborative knowledge transfer (CKT) is responsible for domain-agnostic feature reconstruction from nearby car or infrastructure by spatial-channel pooling operation in an elementwise manner. To promote the candidate alignment, a brand-new bin-based location correction (BLC) provides an auxiliary supervision for cross-dataset box refinement via residual coordinate encoding (RCE), and category-aware pooling alignment (CPA) is further designed for pulling the category-specific instance closer between source and target samples. We benchmark CUDA-X against the counterparts on four prevalent cooperative perception datasets, i.e., OPV2V, V2X-Sim, V2V4Real, and DAIR-V2X: it establishes the new state-of-the-art vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) performances regardless of simulation or reality. We expect that this appealing attempt would provide an in-depth insight into domain generalization in the context of multiagent perception, and the code is publicly available soon.
Daxin Tian, Chunmian Lin, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao
IEEE Trans. Neural Networks Learn. Syst.6
2024 Developing a new integrated advanced driver assistance system in a connected vehicle environment
Siyuan Gong, Dezong Zhao, N. N. Sze, Mohammed A. Quddus 0001, Helai Huang
Expert Syst. Appl.3
2024 Video Deepfake classification using particle swarm optimization-based evolving ensemble models
abstract
The recent breakthrough of deep learning based generative models has led to the escalated generation of photo-realistic synthetic videos with significant visual quality. Automated reliable detection of such forged videos requires the extraction of fine-grained discriminative spatial-temporal cues. To tackle such challenges, we propose weighted and evolving ensemble models comprising 3D Convolutional Neural Networks (CNNs) and CNN-Recurrent Neural Networks (RNNs) with Particle Swarm Optimization (PSO) based network topology and hyper-parameter optimization for video authenticity classification. A new PSO algorithm is proposed, which embeds Muller's method and fixed-point iteration based leader enhancement, reinforcement learning-based optimal search action selection, a petal spiral simulated search mechanism, and cross-breed elite signal generation based on adaptive geometric surfaces. The PSO variant optimizes the RNN topologies in CNN-RNN, as well as key learning configurations of 3D CNNs, with the attempt to extract effective discriminative spatial-temporal cues. Both weighted and evolving ensemble strategies are used for ensemble formulation with aforementioned optimized networks as base classifiers. In particular, the proposed PSO algorithm is used to identify optimal subsets of optimized base networks for dynamic ensemble generation to balance between ensemble complexity and performance. Evaluated using several well-known synthetic video datasets, our approach outperforms existing studies and various ensemble models devised by other search methods with statistical significance for video authenticity classification. The proposed PSO model also illustrates statistical superiority over a number of search methods for solving optimization problems pertaining to a variety of artificial landscapes with diverse geometrical layouts.
Li Zhang 0013, Dezong Zhao, Chee Peng Lim, Houshyar Asadi, Haoqian Huang, Yonghong Yu, Rong Gao 0001
Knowl. Based Syst.2
2024 Distributed Robust Model Predictive Control for Virtual Coupling Under Structural and External Uncertainty
abstract
Virtual coupling is expected to primarily improve the capacity of a railway system. Virtual coupled systems are affected by multi-source disturbances due to the complex operating environment. However, existing research only partially considers the effects of structural or external disturbances, which limits the stability and robustness of the virtually coupled train set (VCTS). In this paper, we aim to tackle the challenges arising from both structural and external disturbances in virtual coupling. We specifically propose a distributed robust model predictive control (DRMPC) solution based on a linearized model by joining linear feedback and feedforward control into a model predictive control (MPC) framework with a discrete Kalman filter (DKF). We also theoretically derive and prove a set of sufficient conditions for both local and string stabilities under structural uncertainty. The stability conditions are incorporated into the constraint space of the distributed MPC framework in order to guarantee system stability in the presence of structural and external uncertainties. The simulation results validate that our proposed control method can stabilize train platooning under both structural and external disturbances. Our control method particularly reduces the spacing and velocity tracking errors by approximately 97.55% and 99.97% on average, respectively, as compared to several baselines.
Daxin Tian, Jianshan Zhou, Xuting Duan, Zhengguo Sheng, Dezong Zhao, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.6
2024 Adaptive Dual-Channel Event-Triggered Fuzzy Control for Autonomous Underwater Vehicles With Multiple Obstacles Environment
abstract
This article investigates the formation control of autonomous underwater vehicles (AUVs) suffering from unknown sea loads, unmoulded structure, limited communication and multiple static and moving obstacles. Given the challenge, a novel adaptive dual-channel event-triggered control scheme is proposed for formation tracking and obstacles avoidance. To economize the communication resources, the dual-channel event-triggered mechanism is designed in the sensor-to-controller and controller-to-actuator channels respectively. By adopting the approximation of fuzzy systems in the form of one-parameter integrated learning, the uncertainties consisted of the unmoulded structure and unknown sea loads are compressed together to be compensated online, which ensures a lower computational cost. To solve the multiple obstacles, the modified artificial potential field approach is employed, and the derived repulsive potential field can ensure that the multi-AUV formation can avoid obstacles smoothly regardless of static or moving obstacles. It is showed by the Lyapunov stability theorem that the tracking errors are guaranteed to be semi-globally uniformly ultimately bounded. Finally, three simulation examples illustrate the effectiveness and superiority of the proposed scheme.
Shang Liu 0003, Dezong Zhao, Houbing Song
IEEE Trans. Intell. Transp. Syst.3
2024 V2VFormer++: Multi-Modal Vehicle-to-Vehicle Cooperative Perception via Global-Local Transformer
abstract
Multi-vehicle cooperative perception has recently emerged for facilitating long-range and large-scale perception ability of connected automated vehicles (CAVs). Nonetheless, enormous efforts formulate collaborative perception as LiDAR-only 3D detection paradigm, neglecting the significance and complementary of dense image. In this work, we construct the first multi-modal vehicle-to-vehicle cooperative perception framework dubbed as V2VFormer++, where individual camera-LiDAR representation is incorporated with dynamic channel fusion (DCF) at bird’s-eye-view (BEV) space and ego-centric BEV maps from adjacent vehicles are aggregated by global-local transformer module. Specifically, channel-token mixer (CTM) with MLP design is developed to capture global response among neighboring CAVs, and position-aware fusion (PAF) further investigate the spatial correlation between each ego-networked map in a local perspective. In this manner, we could strategically determine which CAVs are desirable for collaboration and how to aggregate the foremost information from them. Quantitative and qualitative experiments are conducted on both publicly-available OPV2V and V2X-Sim 2.0 benchmarks, and our proposed V2VFormer++ reports the state-of-the-art cooperative perception performance, demonstrating its effectiveness and advancement. Moreover, ablation study and visualization analysis further suggest the strong robustness against diverse disturbances from real-world scenarios.
Daxin Tian, Chunmian Lin, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.6
2024 A Spatial-State-Based Omni-Directional Collision Warning System for Intelligent Vehicles
abstract
Collision warning systems (CWSs) have been recognized as effective tools in preventing vehicle collisions. Existing systems mainly provide safety warnings based on single-directional approaches, such as rear-end, lateral, and forward collision warnings. Such systems cannot provide omni-directorial enhancements on driver’s perception. Meanwhile, due to the unclear and overlapped activation areas of above single-directional CWSs, multiple kinds of warnings may be triggered mistakenly for a collision. The multi-triggering may confuse drivers about the position of dangerous targets. To this end, this paper develops a spatial-state-based omni-directional collision warning system (S-OCWS), aiming to help drivers identify the specific danger by providing the unique warning. First, the operational domains of rear-end, lateral, and forward collisions are theoretically distinguished. This distinction is attained by a geometric approach with a rigorous mathematical derivation, based on the spatial states and the relative motion states of itself and the target vehicle in real time. Then, a theoretical omni-directional collision warning model is established using time-to-collision (TTC) to clarify activation conditions for different collision warnings. Finally, the effectiveness of the S-OCWS is validated in field tests. Results indicate that the S-OCWS can help drivers quickly and properly respond to the warnings without compromising their control over lateral offsets. In particular, the probability of drivers giving proper responses to FCW doubles when the S-OCWS is on, compared to when the system is off. In addition, the S-OCWS shortens the responses time of nonprofessional drivers, and therefore enhances their safety in driving.
Siyuan Gong, Dezong Zhao, N. N. Sze, Mohammed A. Quddus 0001, Helai Huang
IEEE Trans. Intell. Transp. Syst.3
2024 Neural Inference Search for Multiloss Segmentation Models
abstract
Semantic segmentation is vital for many emerging surveillance applications, but current models cannot be relied upon to meet the required tolerance, particularly in complex tasks that involve multiple classes and varied environments. To improve performance, we propose a novel algorithm, neural inference search (NIS), for hyperparameter optimization pertaining to established deep learning segmentation models in conjunction with a new multiloss function. It incorporates three novel search behaviors, i.e., Maximized Standard Deviation Velocity Prediction, Local Best Velocity Prediction, and n -dimensional Whirlpool Search. The first two behaviors are exploratory, leveraging long short-term memory (LSTM)-convolutional neural network (CNN)-based velocity predictions, while the third employs n -dimensional matrix rotation for local exploitation. A scheduling mechanism is also introduced in NIS to manage the contributions of these three novel search behaviors in stages. NIS optimizes learning and multiloss parameters simultaneously. Compared with state-of-the-art segmentation methods and those optimized with other well-known search algorithms, NIS-optimized models show significant improvements across multiple performance metrics on five segmentation datasets. NIS also reliably yields better solutions as compared with a variety of search methods for solving numerical benchmark functions.
Sam Slade, Li Zhang 0013, Haoqian Huang, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Hanhe Lin, Rong Gao 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 Structurally Optimized Neural Fuzzy Modeling for Model Predictive Control
abstract
This article investigates the local linear model tree (LOLIMOT), a typical neural fuzzy model, in the multiple-input–multiple-output model predictive control (MPC). In the conventional LOLIMOT, the structural parameters including centers and variances of its Gaussian kernels are set based on equally dividing the input data space. In this article, after the structural parameters are initially obtained from the input space partition, they are optimized by the gradient descent search, from which the space partitions are further adjusted. This makes it better for the model structure to fit the input data statistics, leading to improved modeling performance with a small model size. The MPC based on the proposed structurally optimized LOLIMOT is then implemented and verified with both numerical and diesel engine plants. Validation results show that the proposed MPC has significantly a better controlling performance than the MPC based on the conventional LOLIMOT, making it an attractive solution in practice.
Xiaoyan Hu 0009, Yu Gong 0001, Dezong Zhao, Wen Gu
IEEE Trans. Ind. Informatics3
2023 Data-Driven Robust Predictive Control for Mixed Vehicle Platoons Using Noisy Measurement
abstract
This paper investigates cooperative adaptive cruise control (CACC) for mixed platoons consisting of both human-driven vehicles (HVs) and automated vehicles (AVs). This research is critical because the penetration rate of AVs in the transportation system will remain unsaturated for a long time. Uncertainties and randomness are prevalent in human driving behaviours and highly affect the platoon safety and stability, which need to be considered in the CACC design. A further challenge is the difficulty to know the exact models of the HVs and the exact powertrain parameters of both AVs and HVs. To address these challenges, this paper proposes a data-driven model predictive control (MPC) that does not need the exact models of HVs or powertrain parameters. The MPC design adopts the technique of data-driven reachability to predict the future trajectory of the mixed platoon within a given horizon based on noisy vehicle measurements. Compared to the classic adaptive cruise control (ACC) and existing data-driven adaptive dynamic programming (ADP), the proposed MPC ensures satisfaction of constraints such as acceleration limit and safe inter-vehicular gap. With this salient feature, the proposed MPC has provably guarantee in establishing a safe and robustly stable mixed platoon despite of the velocity changes of the leading vehicle. The efficacy and advantage of the proposed MPC are verified through comparison with the classic ACC and data-driven ADP methods on both small and large mixed platoons.
Jianglin Lan, Dezong Zhao, Daxin Tian
IEEE Trans. Intell. Transp. Syst.2
2023 DA-RDD: Toward Domain Adaptive Road Damage Detection Across Different Countries
abstract
Recent advances on road damage detection relies on a large amount of labeled data, whilst collecting pavement image is labor-intensive and time-consuming. Unsupervised Domain Adaptation (UDA) provides a promising solution to adapt a source domain to the target domain, however, cross-domain crack detection is still an open problem. In this paper, we propose domain adaptive road damage detection termed as DA-RDD, by incorporating image-level with instance-level feature alignment for domain-invariant representation learning in an adversarial manner. Specifically, importance weighting is introduced to evaluate the intermediate samples for image-level alignment between domains, and we aggregate RoI-wise feature with multi-scale contextual information to recover the crack details for progressive domain alignment at instance level. Additionally, a large-scale road damage dataset (based on Road Damage Dataset 2020 (RDD2020)) named as RDD2021 is constructed with$100k$synthetic labeled distress images. Extensive experimental results on damage detection across different countries demonstrate the universality and superiority of DA-RDD, and empirical studies on RDD2021 further claim its effectiveness and advancement. To our best knowledge, it is the first time to investigate domain adaptative pavement crack detection, and we expect the contributions in this work would facilitate the development of generalized road damage detection in the future.
Chunmian Lin, Daxin Tian, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.5
2023 Probabilistic Approach for Road-Users Detection
abstract
Object detection in autonomous driving applications implies the detection and tracking of semantic objects that are commonly native to urban driving environments, as pedestrians and vehicles. One of the major challenges in state-of-the-art deep-learning based object detection are false positives which occur with overconfident scores. This is highly undesirable in autonomous driving and other critical robotic-perception domains because of safety concerns. This paper proposes an approach to alleviate the problem of overconfident predictions by introducing a novel probabilistic layer to deep object detection networks in testing. The suggested approach avoids the traditional Sigmoid or Softmax prediction layer which often produces overconfident predictions. It is demonstrated that the proposed technique reduces overconfidence in the false positives without degrading the performance on the true positives. The approach is validated on the 2D-KITTI objection detection through the YOLOV4 and SECOND (Lidar-based detector). The proposed approach enables interpretable probabilistic predictions without the requirement of re-training the network and therefore is very practical.
Gledson Melotti, Weihao Lu 0003, Pedro Conde, Dezong Zhao, Alireza Asvadi, Nuno Gonçalves 0001, Cristiano Premebida
IEEE Trans. Intell. Transp. Syst.4
2023 A Cooperation-Aware Lane Change Method for Automated Vehicles
abstract
Lane change for automated vehicles (AVs) is an important but challenging task in complex dynamic traffic environments. Due to difficulties in guaranteeing safety as well as a high efficiency, AVs are inclined to choose relatively conservative strategies for lane change. To avoid the conservatism, this paper presents a cooperation-aware lane change method utilizing interactions between vehicles. We first propose an interactive trajectory prediction method to explore possible cooperations between an AV and the others. Further, an evaluation on safety, efficiency and comfort is designed to make a decision on lane change. Thereafter, we propose a motion planning algorithm based on model predictive control (MPC), which incorporates AV’s decision and surrounding vehicles’ interactive behaviors into constraints so as to avoid collisions during lane change. Quantitative testing results show that compared with the methods without an interactive prediction, our method enhances driving efficiencies of the AV and other vehicles by 14.8% and 2.6%, respectively, which indicates that a proper utilization of vehicle interactions can effectively reduce the conservatism of the AV and promote the cooperation between the AV and others.
Zihao Sheng, Shibei Xue, Dezong Zhao, Min Jiang 0009, Dewei Li 0001
IEEE Trans. Intell. Transp. Syst.4
2023 3D-DFM: Anchor-Free Multimodal 3-D Object Detection With Dynamic Fusion Module for Autonomous Driving
abstract
Recent advances in cross-modal 3D object detection rely heavily on anchor-based methods, and however, intractable anchor parameter tuning and computationally expensive postprocessing severely impede an embedded system application, such as autonomous driving. In this work, we develop an anchor-free architecture for efficient camera-light detection and ranging (LiDAR) 3D object detection. To highlight the effect of foreground information from different modalities, we propose a dynamic fusion module (DFM) to adaptively interact images with point features via learnable filters. In addition, the 3D distance intersection-over-union (3D-DIoU) loss is explicitly formulated as a supervision signal for 3D-oriented box regression and optimization. We integrate these components into an end-to-end multimodal 3D detector termed 3D-DFM. Comprehensive experimental results on the widely used KITTI dataset demonstrate the superiority and universality of 3D-DFM architecture, with competitive detection accuracy and real-time inference speed. To the best of our knowledge, this is the first work that incorporates an anchor-free pipeline with multimodal 3D object detection.
Chunmian Lin, Daxin Tian, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao
IEEE Trans. Neural Networks Learn. Syst.5
2022 CL3D: Camera-LiDAR 3D Object Detection With Point Feature Enhancement and Point-Guided Fusion
abstract
Camera-LiDAR 3D object detection has been extensively investigated due to its significance for many real-world applications. However, there are still of great challenges to address the intrinsic data difference and perform accurate feature fusion among two modalities. To these ends, we propose a two-stream architecture termed as CL3D, that integrates with point enhancement module, point-guided fusion module with IoU-aware head for cross-modal 3D object detection. Specifically, pseudo LiDAR is firstly generated from RGB image, and point enhancement module (PEM) is then designed to enhance the raw LiDAR with pseudo point. Moreover, point-guided fusion module (PFM) is developed to find image-point correspondence at different resolutions, and incorporate semantic with geometric features in a point-wise manner. We also investigate the inconsistency between localization confidence and classification score in 3D detection, and introduce IoU-aware prediction head (IoU Head) for accurate box regression. Comprehensive experiments are conducted on publicly available KITTI dataset, and CL3D reports the outstanding detection performance compared to both single- and multi-modal 3D detectors, demonstrating its effectiveness and competitiveness.
Chunmian Lin, Daxin Tian, Xuting Duan, Jianshan Zhou, Dezong Zhao, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.5
2022 SA-YOLOv3: An Efficient and Accurate Object Detector Using Self-Attention Mechanism for Autonomous Driving
abstract
Object detection is becoming increasingly significant for autonomous-driving system. However, poor accuracy or low inference performance limits current object detectors in applying to autonomous driving. In this work, a fast and accurate object detector termed as SA-YOLOv3, is proposed by introducing dilated convolution and self-attention module (SAM) into the architecture of YOLOv3. Furthermore, loss function based on GIoU and focal loss is reconstructed to further optimize detection performance. With an input size of$512\times 512$, our proposed SA-YOLOv3 improves YOLOv3 by 2.58 mAP and 2.63 mAP on KITTI and BDD100K benchmarks, with real-time inference (more than 40 FPS). When compared with other state-of-the-art detectors, it reports better trade-off in terms of detection accuracy and speed, indicating the suitability for autonomous-driving application. To our best knowledge, it is the first method that incorporates YOLOv3 with attention mechanism, and we expect this work would guide for autonomous-driving research in the future.
Daxin Tian, Chunmian Lin, Jianshan Zhou, Xuting Duan, Yue Cao 0002, Dezong Zhao, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.6
2022 Robust Min-Max Model Predictive Vehicle Platooning With Causal Disturbance Feedback
abstract
Platoon-based vehicular cyber-physical systems have gained increasing attention due to their potentials in improving traffic efficiency, capacity, and saving energy. However, external uncertain disturbances arising from mismatched model errors, sensor noises, communication delays and unknown environments can impose a great challenge on the constrained control of vehicle platooning. In this paper, we propose a closed-loop min-max model predictive control (MPC) with causal disturbance feedback for vehicle platooning. Specifically, we first develop a compact form of a centralized vehicle platooning model subject to external disturbances, which also incorporates the lower-level vehicle dynamics. We then formulate the uncertain optimal control of the vehicle platoon as a worst-case constrained optimization problem and derive its robust counterpart by semidefinite relaxation. Thus, we design a causal disturbance feedback structure with the robust counterpart, which leads to a closed-loop min-max MPC platoon control solution. Even though the min-max MPC follows a centralized paradigm, its robust counterpart can keep the convexity and enable the efficient and practical implementation of current convex optimization techniques. We also derive a linear matrix inequality (LMI) condition for guaranteeing the recursive feasibility and input-to-state practical stability (ISpS) of the platoon system. Finally, simulation results are provided to verify the effectiveness and advantage of the proposed MPC in terms of constraint satisfaction, platoon stability and robustness against different external disturbances.
Jianshan Zhou, Daxin Tian, Zhengguo Sheng, Xuting Duan, Guixian Qu, Dezong Zhao, Dongpu Cao, Xuemin Shen
IEEE Trans. Intell. Transp. Syst.6
2021 Semantic Feature Mining for 3D Object Classification and Segmentation
abstract
Deep learning on 3D point clouds has drawn much attention, due to its large variety of applications in intelligent perception for automated and robotic systems. Unlike structured 2D images, it is challenging to extract features and implement convolutional networks over these unordered points. Although a number of previous works achieved high accuracies for point cloud recognition, they tend to process local point information in such a way that semantic information is not fully encoded. In this paper, we propose a deep neural network for 3D point cloud processing that utilizes effective feature aggregation methods emphasizing both generalizability and relevance. In particular, our method uses fixed-radius grouping for pooling layers and spherical kernel convolution for semantics mining. To address the issue of gradient degradation and memory consumption of a deep network, a parallel feature feed-forward mechanism and bottleneck layers are implemented to reduce the number of parameters. Experiments show that our algorithm achieves state-of-the-art results and competitive accuracy in both classification and part segmentation while maintaining an efficient architecture.
Weihao Lu 0003, Dezong Zhao, Cristiano Premebida, Wen-Hua Chen 0001, Daxin Tian
ICRA2
2021 Joint Optimization of Resource Scheduling and Mobility for UAV-Assisted Vehicle Platoons
abstract
In the era of the Internet of Everything, autonomous driving has put forward a higher ambition for data transmission capabilities. This paper studies joint scheduling of computation and communication resources in the collaborative networking of unmanned aerial vehicles (UAV s) and platooning vehicles in mobile edge computing (MEC) framework to maximize the energy efficiency. Considering the movement characteristics of vehicles, we integrate mobility, communication, computation, and energy consumption to establish a collective optimization problem. Since this multivariate coupled model is non-convex, we further propose a joint optimization method (JOM) algorithm based on the convex approximation theory, particularly quadratic programming. Experimental results verify that this algorithm converges quickly within a dozen iterations and proves to be superior to several other benchmark schemes.
Yang Liu 0291, Jianshan Zhou, Daxin Tian, Zhengguo Sheng, Xuting Duan, Guixian Qu, Dezong Zhao
VTC Fall7
2021 A Dynamic Model Averaging for the Discovery of Time-Varying Weather-Cycling Patterns
abstract
It has been well recognized that weather variations significantly impact cycling experiences of users. However, the weather-cycling dynamic relationship over time is not well studied in the literature. In this paper, in order to bridge this gap, we propose a Dynamic Model Averaging and Dynamic Model Selection (DMA and DMS) to reveal the characteristics of time-varying responses and the associated influencing factors for young people's shared bike trips. Without loss of generality, dynamic models with unknown observational variances are also proposed. We take New York City as an instance and analyze the drifts of patterns of New York CitiBike trips under six weather factors from various aspects. The results suggest that the bike trips' responses to some weather factors fluctuate dynamically while others maintain at a relatively stable level. It is concluded that a few main influencing factors are adequate to represent the travel patterns. It is observed that dynamic models, with the strength of alleviating multicollinearity, present better forecast performance than classic models. This work can facilitate the decision makers and managers to oversee and optimise travel experience of users in real time.
Guanying Jiang, Xiaobo Qu 0002, Dezong Zhao
IEEE Trans. Intell. Transp. Syst.4
2021 Knowledge Implementation and Transfer With an Adaptive Learning Network for Real-Time Power Management of the Plug-in Hybrid Vehicle
abstract
Essential decision-making tasks such as power management in future vehicles will benefit from the development of artificial intelligence technology for safe and energy-efficient operations. To develop the technique of using neural network and deep learning in energy management of the plug-in hybrid vehicle and evaluate its advantage, this article proposes a new adaptive learning network that incorporates a deep deterministic policy gradient (DDPG) network with an adaptive neuro-fuzzy inference system (ANFIS) network. First, the ANFIS network is built using a new global K-fold fuzzy learning (GKFL) method for real-time implementation of the offline dynamic programming result. Then, the DDPG network is developed to regulate the input of the ANFIS network with the real-world reinforcement signal. The ANFIS and DDPG networks are integrated to maximize the control utility (CU), which is a function of the vehicle's energy efficiency and the battery state-of-charge. Experimental studies are conducted to testify the performance and robustness of the DDPG-ANFIS network. It has shown that the studied vehicle with the DDPG-ANFIS network achieves 8% higher CU than using the MATLAB ANFIS toolbox on the studied vehicle. In five simulated real-world driving conditions, the DDPG-ANFIS network increased the maximum mean CU value by 138% over the ANFIS-only network and 5% over the DDPG-only network.
Quan Zhou 0006, Dezong Zhao, Bin Shuai, Huw Williams, Hongming Xu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 A Game-Based Computation Offloading Method in Vehicular Multiaccess Edge Computing Networks
abstract
Multiaccess edge computing (MEC) is a new paradigm to meet the requirements for low latency and high reliability of applications in vehicular networking. More computation-intensive and delay-sensitive applications can be realized through computation offloading of vehicles in vehicular MEC networks. However, the resources of a MEC server are not unlimited. Vehicles need to determine their task offloading strategies in real time under a dynamic-network environment to achieve optimal performance. In this article, we propose a multiuser noncooperative computation offloading game to adjust the offloading probability of each vehicle in vehicular MEC networks and design the payoff function considering the distance between the vehicle and MEC access point, application and communication model, and multivehicle competition for MEC resources. Moreover, we construct a distributed best response algorithm based on the computation offloading game model to maximize the utility of each vehicle and demonstrate that the strategy in this algorithm can converge to a unique and stable equilibrium under certain conditions. Furthermore, we conduct a series of experiments and comparisons with other offloading methods to analyze the effectiveness and performance of the proposed algorithms. The fast convergence and the improved performance of this algorithm are verified by numerical results.
Ping Lang, Daxin Tian, Jianshan Zhou, Xuting Duan, Yue Cao 0002, Dezong Zhao
IEEE Internet Things J.7
2020 A data-driven operational integrated driving behavioral model on highways
Hengcong Guo, Dezong Zhao
Neural Comput. Appl.4
2016 Quality assessment metric of stereo images considering cyclopean integration and visual saliency
Yafang Wang, Baihua Li, Wen Lu 0004, Qinggang Meng, Zhihan Lyu, Dezong Zhao, Zhiqun Gao
Inf. Sci.7