EDBT 2026 Demo / reviewers in the wild / expert
Lu Dong 0002
dblp:40/2448-2
· DBLP profile ↗
43ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0001-6737-1381ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 7 first-author · 19 since 2021Computer networks · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robustness-enhanced cooperative adaptive cruise control for multi-task scenarios via generalised joint multi-agent reinforcement learning
Lu Dong 0002, Min Hua, Quan Zhou 0006, Changyin Sun 0001 |
Neurocomputing | 1 |
| 2026 | Enabling diverse styles coverless image steganography with two-stage latent transformation and diffusion model
Jiangtao Guo 0001, Buwei Tian, Junyong Jiang, Lu Dong 0002 |
J. Inf. Secur. Appl. | 4 |
| 2026 | Gradient Perturbation Guidance for Boosting Sparse Adversarial Attack TransferabilityabstractSparse adversarial attacks perturb only a few pixels to achieve an attack, making them harder to detect and more dangerous. Recently, generative sparse attacks decouple the generation of sparse adversarial examples (AEs) into dense perturbations and sparse masks. By modeling the data distribution from clean examples to sparse AEs, generative sparse attacks mitigate the poor transferability that arises from over-reliance on gradients. These methods put effort into deriving optimal sparse masks on the generated perturbation. However, the quality of perturbation generation has always been overlooked, which limits the transferability of sparse AEs. To explore the influence of perturbation quality, we conduct empirical analyses of sparse gradient-based perturbations. The results show that directly applying sparsity to gradient-based perturbations disrupts their holistic adversarial information, leading to degraded attack performance. Therefore, it is critical to extract key adversarial knowledge from gradient-based perturbations while preserving their overall integrity to guide sparse adversarial attacks. Motivated by this observation, we propose to extract essential adversarial information from gradient-based AEs to guide the generator to produce higher-quality dense perturbations and stronger transferable sparse AEs. Specifically, we introduce the Gradient Perturbation Guidance (GPG) sparse adversarial attack, which integrates gradient adversarial feature guidance and gradient perturbation guidance regularization. The former guides the generator to capture gradient-based adversarial features during encoding, while the latter refines adversarial knowledge from gradient-based perturbations during decoding. Extensive experiments on ImageNet-1K show that our GPG significantly boosts transferability compared to state-of-the-art methods under consistent sparsity constraints. Our code is available at Github. Chengze Jiang, Minjing Dong, Jie Gui, Lu Dong 0002, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Breaking the Bottleneck of Heterogeneity in Federated Learning Through GBG: Grouping, Block Training, and Global DistillationabstractFederated learning (FL) collaboratively trains a global model across multiple clients without sharing local data, effectively utilizing data while preserving privacy. However, real-world data are often non-independently and identically distributed ( non-IID) and heterogeneous across clients, making standard FL less suitable for practical deployment. To tackle these challenges, we propose a novel framework named GBG which includes grouping, block training and global distillation. First, it introduces a novel grouping mechanism that clusters heterogeneous clients into different groups. Clients in these groups perform different training tasks. Second, it employs Symmetric Balanced Incomplete Block Design (SBIBD) to construct intra-group blocks and establishes a new training paradigm called block training. Finally, by incorporating mutual learning, the GBG framework enables the effective development of both personalized block-level models and a global model. Moreover, we demonstrate the convergence of block training in combination with existing works. Additional experiments further demonstrate that the GBG framework achieves favorable results in both testing accuracy and training error. Zehu Zhang, Lu Dong 0002, Jiangtao Guo 0001, Junyong Jiang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | SBA: A Swift and Stealthy Backdoor Attack Framework for Federated LearningabstractFederated Learning (FL) enables collaborative model training while protecting the privacy of individual participants. However, this decentralized framework is vulnerable to certain risks, particularly backdoor attacks. In such attacks, malicious actors embed triggers into the global model, causing incorrect classifications for inputs with specific features. Traditional backdoor attacks typically use uniform triggers, which makes them more likely to be detected. Furthermore, These methods tend to overlook the fact that poisoned samples may cause abnormal confidence levels in non-target classes. These methods also require an extended attack window to achieve the desired success rate, which reduces their efficiency and increases the chances of detection. This paper introduces a novel backdoor attack method called Swift Backdoor Attack (SBA), designed to overcome these challenges. SBA improves both the stealth and success rate of backdoor attacks by generating adaptive triggers tailored to specific inputs. It also incorporates Non-Target Class Suppression (NTCS) loss and an entropy gain mechanism to enhance the attack’s effectiveness. Our approach suppresses unnecessary activations in the neural network when encountering poisoned samples, making the attack harder to be detected. Experimental results demonstrate that this method significantly increases the success rate of backdoor attacks within a shorter attack window, while evading existing detection mechanisms. This provides a more efficient and covert strategy for executing backdoor attacks in FL. Our code is available at https://github.com/bigcheeses/SBA. Junhan Wang, Zhangming Wu, Zhuoyue Wang, Lu Dong 0002 |
ICASSP | 4 |
| 2025 | MP-CSAS: A Privacy-Preserving Speed Advisory Framework for Mixed Traffic Environment Based on Consortium BlockchainabstractGlobal climate change has emerged as a pressing global challenge, underscoring the imperative for governments and urban traffic management authorities to prioritize carbon emission reduction in the transportation sector. In mixed traffic environments, where internal combustion engine vehicles and electric vehicles coexist, the disparity in carbon emissions between these vehicle types poses a significant challenge for the formulation of effective transportation coordination policies. Consensus-based speed advisory systems (CSAS) have been extensively employed to enhance fleet energy efficiency and mitigate emissions. This paper develops a novel vehicle speed advisory framework for mixed traffic environments, termed the MP-CSAS, where M stands for Mixed Traffic and P for Privacy-preserving, which leverages blockchain technology, privacy-preserving mechanisms, and secure car-following strategies. By incorporating a coordination factor, the framework enables policymakers to dynamically optimize carbon emission reduction strategies while safeguarding vehicle user data privacy and ensuring operational safety. Simulation results demonstrate that the MP-CSAS framework effectively minimizes fleet carbon emissions while preserving data confidentiality and ensuring system security. This study contributes a forward-looking decision-making paradigm for road infrastructure providers and policymakers, equipping them with scientifically grounded and adaptive strategies to achieve sustainable and safe transportation objectives. Lu Dong 0002, Weichao Zhuang, Guodong Yin, Boli Chen |
IEEE Internet Things J. | 1 |
| 2025 | Joint Optimization of Computation Offloading and Caching for Satellite Edge Computing Based on GRU-SAC AlgorithmabstractConsidering the satellites’ wide coverage and independence from geographical constraints, satellite edge computing (SEC) has demonstrated broad application prospects. In this paper, a joint offloading and caching framework is proposed, addressing issues such as redundant data transmission, heterogeneous resources, and the high-speed movement of satellites in SEC. In the framework, latency and system energy consumption are reduced by dynamically caching reusable data and formulating an appropriate offloading strategy. Considering the complexity of the problem, we propose a GRU-SAC algorithm that trains multiple distinct deep neural networks (DNNs) to output offloading and caching actions. It combines maximum entropy with policy gradient to enhance strategy stability and avoid local optima. Furthermore, the algorithm predicts the future request probabilities of tasks as part of the state space, aiding in strategy training and adapting to dynamically changing environments. Simulation results substantiate the effectiveness of our proposed method in reducing latency and system energy consumption. In the simulation environment comprising five users and two satellites, the latency decreased from nearly 132 to around 129, and the energy consumption decreased from nearly 22 to around 18. Lu Dong 0002, Aoting Xu, Jian Liu 0006, Yubin Jia |
IEEE Internet Things J. | 1 |
| 2025 | Optimal DoS Attack Energy Allocation in Cyber-Physical Systems Based on Deep Reinforcement LearningabstractThe security of wireless cyber-physical systems (CPSs) faces significant challenges from channel congestion attacks, which disrupt data transmission and threaten system stability. Existing research on attack energy allocation primarily focuses on single-sensor, single-channel scenarios or assumes discrete energy distribution, limiting the development of optimal strategies. To address this, we model the remote state estimation problem in wireless CPSs as a Markov Decision Process (MDP) and propose an action mapping mechanism. Based on these, we propose two reinforcement learning algorithms to address the continuous attack energy allocation problem in multi-sensor, multi-channel environments. The first algorithm, Genetic Algorithm-Driven Model-based Action Mapping Algorithm (GA-MB-AM), is a model-based reinforcement learning approach that uses a genetic algorithm for network initialization. The second algorithm is Soft Actor-Critic-based Action Mapping Algorithm (SAC-AM), which improves performance in dynamic environments, such as systems with energy harvesting. Simulations demonstrate the superiority of our approach over traditional methods. Junyong Jiang, Buwei Tian, Yubin Jia, Lu Dong 0002 |
IEEE Internet Things J. | 5 |
| 2025 | Fed-SecTP: A Federated-Learning-Based Framework for Secure Vehicle Trajectory Prediction Using Surrounding Vehicle DataabstractAccurate vehicle trajectory prediction process depends on seamless data sharing within the Internet of Vehicles. However, such interconnected data exchange introduces significant security risks. Specifically, network attacks can compromise data integrity, thereby degrading prediction accuracy. Concurrently, the need to protect sensitive vehicle data, such as driving trajectories and user account information, results in data silos that hinder the free flow of information essential for effective prediction. Existing studies have largely addressed either privacy preservation or attack mitigation in isolation, lacking a unified solution that simultaneously tackles both challenges. To address this gap, we propose Fed-SecTP, an integrated dual-module secure federated learning framework. The first module employs a Temporal Convolutional Network (TCN) with multi-head attention to detect and filter network attacks in real-time. The second module combines TCN with a Bidirectional Long Short-Term Memory (Bi-LSTM) network for trajectory prediction and leverages FedProx for federated learning, thereby enabling privacy-preserving model training without sharing raw data. Experimental results demonstrate that Fed-SecTP achieves high prediction accuracy and robustness even when up to 50% of the data is compromised by attacks, while ensuring secure data processing. This framework offers a reliable and comprehensive solution for autonomous vehicle trajectory prediction. Hao Sun 0029, Lu Dong 0002, Boli Chen, Weichao Zhuang, Guodong Yin |
IEEE Internet Things J. | 5 |
| 2025 | RIS-Assisted Energy-Efficient UAV Data Collection Method Based on Deep Reinforcement Learning
Lu Dong 0002, Mengjiao Lu, Yongbao Wu |
Mob. Networks Appl. | 1 |
| 2025 | Memory-Event-Based Distributed T-S Fuzzy Security Control for a Class of Cyber-Physical Systems Under Replay AttackabstractIn this paper, a distributed T-S fuzzy security control strategy based on memory events is proposed for a class of multi-input-multi-output (MIMO) cyber-physical systems (CPSs) that subjected to replay attacks. This strategy can also address common uncertainties in practical systems, including communication latency, unmodeled dynamics, and external unknown disturbances. It is worth noting that the replay attack model in the paper is highly generalized. And the attack location, target, and frequency are all uncertain. Therefore, a suitable distributed memory event-based strategy (DMEBS) is designed. It can dynamically adjust the usage of historical data to optimize the release of sampled data, thereby determining when to update the control laws of each subsystem, ensuring system performance while greatly saving communication resources. In addition, the stability of the system is demonstrated by establishing a suitable Lyapunov-Krasovskii (L-K) functional, ensuring the elimination of the Zeno phenomenon. Finally, the effectiveness of the proposed method is validated through simulations conducted on two commonly encountered practical systems. Yi Shui, Lu Dong 0002, Ya Zhang 0001, Changyin Sun 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | Finite-Time Vehicle Platoon Control for CPVS Under DoS AttacksabstractThis paper addresses the problem of vehicle platoon control for third-order nonlinear cyber physical vehicle systems (CPVSs) within finite-time under denial-of-service (DoS) attacks. Unlike existing approaches that assume known systematic matrix of the leader vehicle, this study proposes a data-driven learning algorithm to learn unknown systematic matrix of the leader vehicle. Additionally, a finite-time distributed observer is introduced, thereby enabling follower vehicles to achieve finite-time state observation of the leader vehicle under DoS attacks. Moreover, a novel low-pass filter chain is designed to construct a new variable with high-order derivatives. Utilizing the new variable, a finite-time resilient decentralized controller is formulated, incorporating fuzzy adaptive methods and backstepping techniques to achieve finite-time vehicle platoon control under DoS attacks. Finally, simulation experiments validate the effectiveness of the proposed method. Sha Fan, Dan Zhang 0001, Lu Dong 0002, Chao Deng 0008 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Reinforcement Learning with Safe Action Generation for Autonomous RacingabstractReinforcement learning (RL) has been successfully applied to tackle many complex decision-making tasks. However, a critical problem is the unsafe exploration when deploying classical model-free RL methods in safety-critical systems. In this paper, we focus on safe RL for autonomous racing, which is formulated as a multi-objective optimization problem (MOP). We propose a novel mechanism, named Safe Action Generation (SAG), to address the above challenge. The mechanism mainly consists of two modules: a risk prediction module and an action generation module. The risk prediction module monitors the state of the vehicle in real-time, and once the state is evaluated as risky, the action generation module is activated to prevent unsafe behavior by restricting the exploration of the RL agent. Extensive experiments in The Open Racing Car Simulator (TORCS) demonstrate that our approach can complete one lap at an average speed of 108 km/h, Compared to the baseline without the SAG mechanism, our approach reduces the instances of driving off the track by 97.7%. By using the proposed framework, the learned policy can improve vehicle safety and generalization performance. Yuanda Wang, Lu Dong 0002 |
CEC | 3 |
| 2024 | Hierarchical reinforcement learning for kinematic control tasks with parameterized action spaces
Jingyu Cao, Lu Dong 0002, Changyin Sun 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Correction: Hierarchical reinforcement learning for kinematic control tasks with parameterized action spaces
Jingyu Cao, Lu Dong 0002, Changyin Sun 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Hierarchical multi-agent reinforcement learning for cooperative tasks with sparse rewards in continuous domain
Jingyu Cao, Lu Dong 0002, Yuanda Wang, Changyin Sun 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Semi-Supervised Feature Distillation and Unsupervised Domain Adversarial Distillation for Underwater Image EnhancementabstractAt present, deep learning has demonstrated outstanding performance in the area of underwater image enhancement. However, these approaches often demand substantial computational resources and extended training time. Knowledge distillation is a widely used technique for model compression, and nowadays it has delivered outstanding results across various fields. However, it has not been utilized in the field of underwater image enhancement. To tackle the aforementioned issues, this paper introduces a knowledge distillation technique for underwater image enhancement for the first time. It is a semi-supervised self-inter feature distillation and unsupervised self-domain adversarial distillation approach. It specifically includes adaptive local self-feature distillation technique, information lossless multi-scale inter-feature distillation technique, and self-domain adversarial distillation approach in LAB-RGB space. Self-feature distillation enhances the performance of the student network by correcting other lossy feature maps with the maximum effective feature map. Inter-feature distillation enables the student network to maximize the potential information learned from the teacher network. Furthermore, an information loss-free pooling approach is suggested to achieve multi-scale loss-free information extraction. Self-domain adversarial distillation boosts the performance of student networks through unsupervised adaptive enhancement in LAB space and unsupervised domain adversarial distillation in RGB space. Finally, a self-inter alternate knowledge distillation training measure is proposed, aiming to maximize the respective benefits of self-inter knowledge distillation. Through extensive comparative experiments, it can be found that student networks with dissimilar structures trained using the knowledge distillation technique designed in this paper achieve outstanding underwater image enhancement results. Nianzu Qiao, Changyin Sun 0001, Lu Dong 0002, Quanbo Ge |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Intermittent Fixed-Time Fuzzy Consensus of Nonlinear Multiagent Systems With Unknown Control Directions and Event-Based CommunicationabstractIn this article, a new event-based aperiodic intermittent fixed-time consensus control strategy is developed for multiagent systems (MASs) with nonlinear uncertainties and unknown control directions. Concretely, the intermittent control strategy is considered to construct the intermittent event-based control (IEBC) algorithm, which leads to substantial savings in communication resources. In addition, an enhanced triggering algorithm is further designed to eliminate continuous monitoring of neighbors' and its own states. Consequently, different from the existing event-based control strategy, the IEBC algorithms proposed herein can further reduce communication energy consumption. To handle the nonlinear uncertainties in MASs, fuzzy logic systems are utilized, enhancing the algorithm to tackle more general problems. Considering the problem of unknown control directions, a controller employing Nussbaum-type functions is formulated. Moreover, the fixed-time consensus control algorithm is incorporated, and the estimation of the convergence time is not reliant on the initial states. Finally, the feasibility of the proposed algorithm is demonstrated through a numerical example. Jian Liu 0006, Jinglong Shi, Lu Dong 0002, Changyin Sun 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | Switching-Event-Based Interval Type-2 Fuzzy Control for a Class of Uncertain Nonlinear SystemsabstractIn this article, a switching-event-based interval type-2 variable universe fuzzy tracking control strategy is proposed for a class of uncertain nonlinear systems. Remarkably, the nonlinearities and the large unknown uncertainties (including parametric and structural) can be allowed. Therefore, a novel switching-event-based mechanism is proposed. It not only determines when to update the control law, but also when to update the design parameters. At the same time, generous computing resources are saved. In addition, the interval type-2 fuzzy control technology is introduced to resist unknown uncertainties by identifying the controlled model online. Moreover, the Lyapunov function is designed to prove that the Zeno phenomenon does not occur and the closed-loop system is asymptotically stable. Finally, two practical simulations are given to demonstrate the effectiveness of the proposed method. Yi Shui, Lu Dong 0002, Ya Zhang 0001, Changyin Sun 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Switching-Event-Based Interval Type-2 T-S Variable Direction Fuzzy Control for Time-Delay Systems With Unknown Control DirectionsabstractIn this paper, a switching-event-based interval type2 (IT2) T-S variable direction fuzzy tracking control strategy is proposed for time-varying delay systems with unknown control directions. To deal with the time-varying delay problem, the T-S fuzzy logic system (TSFLS) is used to approximate the unknown nonlinear functions. A novel logic-based switching mechanism is proposed to handle the problem of unknown control directions. At the same time, in order to ensure the stability and tracking performance of the system, an auxiliary controller is designed. The proposed controller not only ensures the tracking performance well, but also reduces the communication burden between the controller and the actuator. In addition, through designing appropriate Lyapunov-Krasoviskii (L-K) functional for tracking error, the system is proved to be asymptotically stable and the Zeno phenomenon is excluded. Wherein the main parameters discussed are tracking error and time interval for eventtriggering. Finally, the effectiveness of the proposed method is verified by using both a mathematical model and an actual physical model. Yi Shui, Lu Dong 0002, Ya Zhang 0001, Changyin Sun 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Dynamic Event-Based Hierarchical Fuzzy Prescribed Performance Control for Underactuated Systems With Uncertain Dead ZoneabstractIn this article, an event-based hierarchical fuzzy prescribed performance control strategy is proposed for a class of underactuated systems with input dead zones. It is worth noting that the slope of the input dead zone is uncertain (time-varying/fuzzy). Therefore, a suitable hierarchical fuzzy logic system (HFLS) is designed to compensate for uncertain dead zones while significantly reducing the number of fuzzy rules. In addition, a dynamic event-based mechanism (DEBM) is proposed, which not only determines when to update the control law of the upper-level fuzzy system, but also when to update the parameters of the lower-level fuzzy controller. Moreover, this strategy can better tolerate interference while achieving the specified transient and steady-state performance of the system, and greatly save computing/communication resources. Furthermore, a Lyapunov function is designed to prove the stability of the system and eliminate the Zeno phenomenon. Finally, simulations are conducted using a quadcopter under two uncertain dead zone conditions to verify the effectiveness of the method. Yi Shui, Lu Dong 0002, Ya Zhang 0001, Changyin Sun 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Functionality-Verification Attack Framework Based on Reinforcement Learning Against Static Malware DetectorsabstractCurrent adversarial attacks are capable of achieving effective evasion against machine learning-based static malware detectors. However, these methods have problems such as long example generation times and lack of functionality validation. To address these issues, we propose an enhanced adversarial example generation framework based on reinforcement learning. This framework improves the example generation efficiency by redesigning the state space and action space employed by the agents. Furthermore, we incorporate the functionality of adversarial example validation for the first time as a component of the example generation process within the framework, significantly enhancing the efficiency of verification. Multiple popular detectors are chosen as victim models to assess the effectiveness of the attack framework. The vulnerabilities of these detectors are elucidated through explanations of the detectors and the analysis of attack results. Finally, a policy distillation approach based on transfer learning is employed to enhance the generalizability of the framework. By learning expert knowledge from agents trained against different detectors, the framework could launch effective attacks against various detectors. The effectiveness of the proposed framework is verified through experiment results. Buwei Tian, Junyong Jiang, Zichen He, Lu Dong 0002, Changyin Sun 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | A Local-and-Global Attention Reinforcement Learning Algorithm for Multiagent Cooperative NavigationabstractThe cooperative navigation algorithm is the crucial technology for multirobot systems to accomplish autonomous collaborative operations, and it is still a challenge for researchers. In this work, we propose a new multiagent reinforcement learning algorithm called multiagent local-and-global attention actor-critic (MLGA2C) for multiagent cooperative navigation. Inspired by the attention mechanism, we design the local-and-global attention module to dynamically extract and encode critical environmental features. Meanwhile, based on the centralized training and decentralized execution (CTDE) paradigm, we extend a new actor-critic method to handle feature encoding and make navigation decisions. We also evaluate the proposed algorithm in two cooperative navigation scenarios: static target navigation and dynamic pedestrian target tracking. The multiple experimental results show that our algorithm performs well in cooperative navigation tasks with increasing agents. Chunwei Song, Zichen He, Lu Dong 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-scale and Self-mutual Feature Distillation
Nianzu Qiao, Jia Sun 0004, Lu Dong 0002 |
ICIC (2) | 3 |
| 2023 | Robust Navigation with Cross-Modal Fusion and Knowledge TransferabstractRecently, learning-based approaches show promising results in navigation tasks. However, the poor generalization capability and the simulation-reality gap prevent a wide range of applications. We consider the problem of improving the generalization of mobile robots and achieving sim-to-real transfer for navigation skills. To that end, we propose a cross-modal fusion method and a knowledge transfer framework for better generalization. This is realized by a teacher-student distillation architecture. The teacher learns a discriminative representation and the near-perfect policy in an ideal environment. By imitating the behavior and representation of the teacher, the student is able to align the features from noisy multi-modal input and reduce the influence of variations on navigation policy. We evaluate our method in simulated and real-world environments. Experiments show that our method outperforms the baselines by a large margin and achieves robust navigation performance with varying working conditions. Wenzhe Cai, Guangran Cheng, Lingyue Kong, Lu Dong 0002, Changyin Sun 0001 |
ICRA | 4 |
| 2023 | Credit assignment in heterogeneous multi-agent reinforcement learning for fully cooperative tasks
Wenzhang Liu, Yuanda Wang, Lu Dong 0002, Changyin Sun 0001 |
Appl. Intell. | 4 |
| 2023 | Multi-objective deep reinforcement learning for crowd-aware robot navigation with dynamic human preference
Guangran Cheng, Yuanda Wang, Lu Dong 0002, Wenzhe Cai, Changyin Sun 0001 |
Neural Comput. Appl. | 3 |
| 2023 | Multiagent Soft Actor-Critic Based Hybrid Motion Planner for Mobile RobotsabstractIn this article, a novel hybrid multirobot motion planner that can be applied under no explicit communication and local observable conditions is presented. The planner is model-free and can realize the end-to-end mapping of multirobot state and observation information to final smooth and continuous trajectories. The planner is a front-end and back-end separated architecture. The design of the front-end collaborative waypoints searching module is based on the multiagent soft actor-critic (MASAC) algorithm under the centralized training with decentralized execution (CTDE) diagram. The design of the back-end trajectory optimization module is based on the minimal snap method with safety zone constraints. This module can output the final dynamic-feasible and executable trajectories. Finally, multigroup experimental results verify the effectiveness of the proposed motion planner. Zichen He, Lu Dong 0002, Chunwei Song, Changyin Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Hybrid Reinforcement Learning for Optimal Control of Non-Linear Switching SystemabstractBased on the reinforcement learning mechanism, a data-based scheme is proposed to address the optimal control problem of discrete-time non-linear switching systems. In contrast to conventional systems, in the switching systems, the control signal consists of the active mode (discrete) and the control inputs (continuous). First, the Hamilton-Jacobi-Bellman equation of the hybrid action space is derived, and a two-stage value iteration method is proposed to learn the optimal solution. In addition, a neural network structure is designed by decomposing the Q-function into the value function and the normalized advantage value function, which is quadratic with respect to the continuous control of subsystems. In this way, the Q-function and the continuous policy can be simultaneously updated at each iteration step so that the training of hybrid policies is simplified to a one-step manner. Moreover, the convergence analysis of the proposed algorithm with consideration of approximation error is provided. Finally, the algorithm is applied evaluated on three different simulation examples. Compared to the related work, the results demonstrate the potential of our method. Xiaofeng Li 0014, Lu Dong 0002, Lei Xue 0003, Changyin Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Critic Learning-Based Control for Robotic Manipulators With Prescribed ConstraintsabstractIn this article, the optimal control problem for robotic manipulators (RMs) with prescribed constraints is addressed. Considering the environmental conditions and requirements of practical applications, prescribed constraints are imposed on the system states to guarantee the control performance and normal operation of the robotic system. Accordingly, an error transformation function is adopted to cope with the prescribed constraints and generate an equivalent unconstrained error for the convenience of the intelligent control design. In order to improve the learning ability and optimize the control performance, critic learning (CL) is introduced to the control design of the constrained RM based on the transformed equivalent unconstrained system. In addition, the stability analysis is given to illustrate the feasibility of the proposed CL-based control. Finally, simulations are conducted on a two-degree-of-freedom (DOF)-constrained RM to further validate the effectiveness of the proposed controller. Yuncheng Ouyang, Lu Dong 0002, Changyin Sun 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Asynchronous Multithreading Reinforcement-Learning-Based Path Planning and Tracking for Unmanned Underwater VehicleabstractThe underwater unmanned vehicle (UUV) is widely used in various marine operations, in which path planning and trajectory tracking are the critical technologies to achieve autonomous motion planning. Unlike previous research methods, this article proposes the asynchronous multithreading proximal policy optimization-based path planning (AMPPO-PP) and trajectory tracking (AMPPO-TT) algorithms and applies these two methods to different task scenarios of UUVs. Taking advantage of the AMPPO, the expensive online computational procedure is converted to an offline training process. The proposed algorithms enable the UUV to learn autonomous planning, tracking, and emergency obstacle avoiding. Besides, the algorithm architecture of the AMPPO-PP and the AMPPO-TT is described in detail. By refining the reward in each timestep and utilizing the reward-shaping trick, the reward sparsity is avoided. The goal-distance heuristic reward function is used to make the UUV explore more directionally. Various simulation environments are developed from simple to complex, along with multiple comparative experiments to verify the effectiveness of the proposed algorithms. Zichen He, Lu Dong 0002, Changyin Sun 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Neural Network-Based Finite-Time Distributed Formation-Containment Control of Two-Layer Quadrotor UAVsabstractIn this article, quadrotor unmanned aerial vehicles (QUAVs) are organized as a two-layer structure, where the first layer is leader QUAVs and the second layer is follower QUAVs. In this structure, only the leader QUAVs can receive the desired tracking information of position and attitude. Although the followers cannot obtain the given tracking information directly, they can obtain the corresponding information from leaders and other followers through the communication network based on the graph theory. In terms of this case, a distributed formation-containment (FC) control method is proposed to handle the related flight problems. We aim to develop a formation control for the leader QUAVs and a containment control for the follower QUAVs with the graph theory. Furthermore, a neural network (NN) technique is utilized to cope with the uncertainty of each QUAV. In order to guarantee good flight performance when tracking, the finite-time stability theorem is introduced into the control design to make each QUAV achieve satisfactory tacking performance in finite time. Finally, numerical simulations are conducted in the platform of two-layer 16 QUAVs to validate the feasibility and effectiveness of the proposed NN-based finite-time FC control. Yuncheng Ouyang, Lei Xue 0003, Lu Dong 0002, Changyin Sun 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Solver-Critic: A Reinforcement Learning Method for Discrete-Time-Constrained-Input SystemsabstractIn this article, a solver-critic (SC) architecture is developed for optimal control problems of discrete-time (DT)-constrained-input systems. The proposed design consists of three parts: 1) a critic network; 2) an action solver; and 3) a target network. The critic network first approximates the action-value function using the sum-of-squares (SOS) polynomial. Then, the action solver adopts the SOS programming to obtain control inputs within the constraint set. The target network introduces the soft update mechanism into policy evaluation to stabilize the learning process. By using the proposed architecture, the constrained-input control problem can be solved without adding the nonquadratic functionals into the reward function. In this article, the theoretical analysis of the convergence property is presented. Besides, the effects of both different initial Q -functions and different discount factors are investigated. It is proven that the learned policy converges to the optimal solution of the Hamilton-Jacobi-Bellman equation. Four numerical examples are provided to validate the theoretical analysis and also demonstrate the effectiveness of our approach. Lu Dong 0002, Changyin Sun 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Reinforcement Learning With Task Decomposition for Cooperative Multiagent SystemsabstractIn this article, we study cooperative multiagent systems (MASs) with multiple tasks by using reinforcement learning (RL)-based algorithms. The target for a single-agent RL system is represented by its scalar reward signals. However, for an MAS with multiple cooperative tasks, the holistic reward signal consists of multiple parts to represent the tasks, which makes the problem complicated. Existing multiagent RL algorithms search distributed policies with holistic reward signals directly, making it difficult to obtain an optimal policy for each task. This article provides efficient learning-based algorithms such that each agent can learn a joint optimal policy to accomplish these multiple tasks cooperatively with other agents. The main idea of the algorithms is to decompose the holistic reward signal for each agent into multiple parts according to the subtasks, and then the proposed algorithms learn multiple value functions with the decomposed reward signals and update the policy with the sum of distributed value functions. In addition, this article presents a theoretical analysis of the proposed approach. Finally, the simulation results for both discrete decision-making and continuous control problems have demonstrated the effectiveness of the proposed algorithms. Changyin Sun 0001, Wenzhang Liu, Lu Dong 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Event-triggered receding horizon control via actor-critic design
Lu Dong 0002, Changyin Sun 0001 |
Sci. China Inf. Sci. | 1 |
| 2020 | Neural network based tracking control for an elastic joint robot with input constraint via actor-critic design
Yuncheng Ouyang, Lu Dong 0002, Yanling Wei 0001, Changyin Sun 0001 |
Neurocomputing | 2 |
| 2020 | Cooperative control for multi-player pursuit-evasion games with reinforcement learning
Yuanda Wang, Lu Dong 0002, Changyin Sun 0001 |
Neurocomputing | 2 |
| 2019 | Functional Nonlinear Model Predictive Control Based on Adaptive Dynamic ProgrammingabstractThis paper presents a functional model predictive control (MPC) approach based on an adaptive dynamic programming (ADP) algorithm with the abilities of handling control constraints and disturbances for the optimal control of nonlinear discrete-time systems. In the proposed ADP-based nonlinear MPC (NMPC) structure, a neural-network-based identification is established first to reconstruct the unknown system dynamics. Then, the actor-critic scheme is adopted with a critic network to estimate the index performance function and an action network to approximate the optimal control input. Meanwhile, as the MPC strategy can effectively determine the current control by solving a finite horizon open-loop optimal control problem, in the proposed algorithm, the infinite horizon is decomposed into a series of finite horizons to obtain the optimal control. In each finite horizon, the finite ADP algorithm solves the optimal control problem subject to the terminal constraint, the control constraint, and the disturbance. The uniform ultimate boundedness of the closed-loop system is verified by the Lyapunov approach. Finally, the ADP-based NMPC is conducted on two different cases and the simulation results demonstrate the quick response and strong robustness of the proposed method. Lu Dong 0002, Jun Yan 0007, Haibo He, Changyin Sun 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Robust optimal control for time-delay systems with dynamic uncertainties via ADPabstractThis paper considers a robust optimal control design for a class of nonlinear discrete-time systems with unknown time-varying delays and dynamic uncertainties. An iterative control strategy based on adaptive dynamic programming (ADP) has been proposed. Neural networks are applied to realize the state prediction, the control input estimation and the performance index function approximation. The estimated control input and performance index function are updated iteratively. Furthermore, it has been proven that the approximated performance index function can converge to the optimal solution of the Hamilton-Jacobia-Bellman (HJB) equation. Finally, the proposed algorithm has been conducted to a numerical simulation. The simulation results demonstrate the effectiveness of the new design. Lu Dong 0002, Jun Li 0033, Wankou Yang, Changyin Sun 0001 |
IJCNN | 1 |
| 2017 | Adaptive Event-Triggered Control Based on Heuristic Dynamic Programming for Nonlinear Discrete-Time SystemsabstractThis paper presents the design of a novel adaptive event-triggered control method based on the heuristic dynamic programming (HDP) technique for nonlinear discrete-time systems with unknown system dynamics. In the proposed method, the control law is only updated when the event-triggered condition is violated. Compared with the periodic updates in the traditional adaptive dynamic programming (ADP) control, the proposed method can reduce the computation and transmission cost. An actor-critic framework is used to learn the optimal event-triggered control law and the value function. Furthermore, a model network is designed to estimate the system state vector. The main contribution of this paper is to design a new trigger threshold for discrete-time systems. A detailed Lyapunov stability analysis shows that our proposed event-triggered controller can asymptotically stabilize the discrete-time systems. Finally, we test our method on two different discrete-time systems, and the simulation results are included. Lu Dong 0002, Xiangnan Zhong, Changyin Sun 0001, Haibo He |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Event-Triggered Adaptive Dynamic Programming for Continuous-Time Systems With Control ConstraintsabstractIn this paper, an event-triggered near optimal control structure is developed for nonlinear continuous-time systems with control constraints. Due to the saturating actuators, a nonquadratic cost function is introduced and the Hamilton-Jacobi-Bellman (HJB) equation for constrained nonlinear continuous-time systems is formulated. In order to solve the HJB equation, an actor-critic framework is presented. The critic network is used to approximate the cost function and the action network is used to estimate the optimal control law. In addition, in the proposed method, the control signal is transmitted in an aperiodic manner to reduce the computational and the transmission cost. Both the networks are only updated at the trigger instants decided by the event-triggered condition. Detailed Lyapunov analysis is provided to guarantee that the closed-loop event-triggered system is ultimately bounded. Three case studies are used to demonstrate the effectiveness of the proposed method. Lu Dong 0002, Xiangnan Zhong, Changyin Sun 0001, Haibo He |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Dual heuristic dynamic programming based event-triggered control for nonlinear continuous-time systemsabstractA novel event-triggered approach for a class of nonlinear continuous-time system is proposed in this paper to reduce the computation cost of the dual heuristic dynamic programming (DHP) algorithm. Two neural networks are included in our design. A critic network is used to estimate the partial derivatives of the cost function with respect to its inputs, and an action network is used to approximate the optimal control law. Instead of periodical sampling in the traditional DHP approach, under the event-triggered mechanism, both of the neural networks are only updated at the jump instants, and kept constant during the inter-event time. With the designed trigger threshold, the proposed DHP-based event-triggered approach can save computation time significantly while obtaining competitive control performance when comparing with those of the traditional DHP approach. Two simulation tests are presented to verify the theoretical results. Lu Dong 0002, Changyin Sun 0001, Haibo He |
IJCNN | 1 |
| 2015 | Predictive event-triggered control based on heuristic dynamic programming for nonlinear continuous-time systemsabstractIn this paper, a novel predictive event-triggered control method based on heuristic dynamic programming (HDP) algorithm is developed for nonlinear continuous-time systems. A model network is used to estimate the system state vector, so that the event-triggered instant is available to predict one step ahead of time. Furthermore, an actor-critic structure is used to approximate the optimal event-triggered control law and performance index function. Although event-triggered adaptive dynamic programming (ADP) has been investigated in the community before, to our best knowledge, this is the first study of using a “predictive” approach through a model network to design the event-triggered ADP. This is the key contribution of this work. Compared to the existing event-triggered ADP methods, our simulations demonstrate that the predictive event-triggered approach can achieve improved control performance and lower computational cost in comparison with the existing methods. Lu Dong 0002, Xiangnan Zhong, Changyin Sun 0001, Haibo He |
IJCNN | 1 |