Hao Xu 0002

dblp:43/6008-2 · DBLP profile ↗
← Back
44ranked-venue papers
10as first author
17since 2021 · last 2025
0000-0003-4130-7925ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 9 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Extended mean field game theoretical optimal distributed control for large scale multi-agent systems: An efficiency-complexity tradeoff
Shawon Dey, Hao Xu 0002
Inf. Sci.2
2025 Safe adaptive control for multi-agent systems under uncertain dynamic environments: a robust barrier-certified learning approach
Shawon Dey, Hao Xu 0002
Neural Comput. Appl.2
2025 Online Receding Horizon Safe-Critical Control for Complex Nonlinear System Under Uncertainty: A Biological Brain-Inspired Dual Learning Approach
abstract
This research introduces a real-time, computationally efficient receding horizon control (RHC) framework with adaptive safety mechanisms under uncertainty. Achieving real-time optimal performance and safety with uncertain nonlinear systems and environments is a challenging task due to the heavy computational complexity involved in learning optimal solutions and managing system uncertainties. To address this, a novel RHC-based safe critical mechanism is designed, enhancing the classical RHC by integrating performance with a real-time safety framework that adapts to environmental uncertainties. Particularly, a human brain-inspired dual-learning algorithm is developed with a slow learning phase that gradually learns the system and environmental uncertainties and further achieves the optimal RHC solution by employing a situation-aware physics-informed neural network (SA-PINN). A fast-learning mechanism is then developed to maintain system safety with quick decision-making capabilities, utilizing a fast-learned adaptive control barrier (FLA-CBF) function that incorporates an adaptive bound relaxation. This approach enables the slow-learning phase to gradually learn the optimal solutions and adapt to uncertainties, thus enhancing the fast-learning phase, which supports real-time control execution with guaranteed safety. In addition, Lyapunov stability analysis and a series of experimental results have been provided to showcase the efficacy of the developed algorithm.
Shawon Dey, Hao Xu 0002
IEEE Trans Autom. Sci. Eng.2
2025 Distributed Adaptive Flocking Control for Large-Scale Multiagent Systems
abstract
This article presents a novel distributed flocking control method for large-scale multiagent systems (LS-MASs) operating in uncertain environments. When dealing with a massive number of flocking agents in uncertain environments, existing flocking methods encounter the problem of communication complexity and "Curse of dimensionality" caused by the exponential growth of agent interactions while solving PDE-based optimal flocking control for large-scale systems. The mean field game (MFG) method addresses this issue by transforming interactions among all agents into the interaction of each individual agent with average effects represented by a probability density function (pdf) of other agents. However, relying solely on a pdf term to consider other agents' states can result in inefficient flocking performance due to the absence of a proficient coordination mechanism encompassing all agents involved in flocking. To overcome these difficulties and achieve the desired flocking performance for LS-MASs, the agents are decomposed into a finite number of subgroups. Each subgroup comprises a leader and followers, and a hybrid game theory is developed to manage both inter- and intragroup interactions. The method incorporates a cooperative game that links leaders from different groups to formulate distributed flocking control, a Stackelberg game that teams up leaders and followers within the same group to extend collective flocking behavior, and an MFG for followers to address the challenges of LS-MASs. Furthermore, to achieve distributed adaptive flocking using the hybrid game structure, we propose a hierarchical actor-critic-mass-based reinforcement learning technique. This approach incorporates a multiactor-critic method for leaders and an actor-critic-mass algorithm for followers, enabling adaptive flocking control in a distributed manner for large-scale agents. Finally, numerical simulation including comparison study and Lyapunov analysis demonstrates the effectiveness of the developed method.
Shawon Dey, Hao Xu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 An Imbalanced Mean-Field Game Theoretical Large-Scale Multiagent Optimization With Constraints
abstract
A novel optimization algorithm has been developed for distributed large-scale multiagent systems (LS-MASs), specifically focusing on achieving a terminal density constraint. While the recent advancement in mean field game (MFG) offers a feasible distributed solution to address the “Curse of Dimensionality” problem, it compromises the optimality of large-scale homogeneous agents and lacks the capability to achieve arbitrary fixed terminal probability density function (PDF) constraint, especially when deviating from the normal distribution. To tackle this issue, a novel approach called the imbalanced mean-field game (Imb-MFG) theory has been designed alongside an adaptive PDF decomposition method and distributed reinforcement learning (RL) that can effectively obtain optimal solutions in LS-MAS even with fixed terminal density constraints in a distributed manner. In particular, a method based on the induction theory has been developed for estimating the parameter of the final PDF constraint, enabling the decomposition of MFG-PDF into multiple imbalanced normal distributions. Subsequently, the Imb-MFG theory is developed by integrating the decomposed multigroup MFG agents with a K-means clustering algorithm with constraint. The developed Imb-MFG approach decomposes a single PDF into multiple imbalanced normal distributions, which are combined to achieve arbitrary terminal PDF constraints. To achieve the solution of the Imb-MFG theory, a multiactor-critic-mass (M-ACM) algorithm is developed. This algorithm is developed to concurrently learn the solution for coupled Fokker-Planck–Kolmogorov (FPK) and Hamilton-Jacobi–Bellman equations. The algorithm’s convergence is ensured through the Lyapunov analysis. The effectiveness of this algorithm is validated through a simulation study.
Shawon Dey, Hao Xu 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Data-Driven Learning Based Secrecy Rate Optimization for RIS-Assisted Secure Wireless Networks: A Multi-Armed Bandit Approach
abstract
In this paper, the secrecy rate optimization problem for Reconfigurable Intelligent Surface (RIS) assisted secure wireless networks is investigated. However, to fully stimulate the potential of RIS, a novel algorithm is needed to optimize the security of RIS-assisted wireless networks. Specifically, a secrecy optimization for multiple RIS-assisted wireless networks with uncertain environments as well as eavesdropping is given first. Then, to solve the formulated optimal secrecy problem, the Multi-Armed Bandit (MAB) has been adopted and further integrated with a deep reinforcement learning algorithm along with Deep Deterministic Policy Gradients (DDPG) technique. This developed algorithm can effectively enhance the secrecy of RIS-assisted wireless networks, i.e. minimizing secrecy outage probability, etc., not only through selecting appropriate RIS, but also by optimizing transmit power and phase shift matrix in selected RIS. Eventually, the effectiveness of the developed algorithm has been demonstrated through a series of numerical simulations.
Hongshan Mao, Hao Xu 0002
CCNC2
2024 Decentralized Multi-agent Reinforcement Learning for Large-scale Mobile Wireless Sensor Network Control Using Mean Field Games
abstract
In this paper, the real-time optimal transmission power control problem is investigated for large-scale mobile wireless sensor networks (MWSN). Controlling large-scale MWSN has two novel challenges, i.e., 1) increasing navigation complexity due to a large number of mobile wireless sensors, and 2) limited energy prohibits peer-to-peer communication between large-scale mobile sensors. To overcome these challenges, the novel mean field game theory is adopted and integrated along with the emerging decentralized reinforcement learning technique. Specifically, the optimal transmission control problem and the optimal navigation problem are formulated as mean field games with two objectives. Then, a novel Actor-Critic-Mass multi-agent reinforcement learning algorithm is developed to learn the decentralized optimal transmission power control and motion control. To learn the decentralized optimal navigation and transmission power control policies, the coupled Hamiltonian-Jacobian-Bellman (HJB) and Fokker-Planck-Kolmogorov (FPK) equations are derived in mean field game formulation. The learned decentralized policies can be guaranteed to converge close to the optimal value, i.e., the Nash Equilibrium, even with large-scale MWSNs in uncertain environments. Finally, the numerical simulations have been provided to demonstrate the effectiveness of the proposed design.
Zejian Zhou, Lijun Qian, Hao Xu 0002
ICCCN3
2024 Robust Barrier-Certified Safe Learning-based Adaptive Control for Multi-Agent Systems in Presence of Uncertain Environments
abstract
This paper develops a decentralized safe learningbased adaptive control for multi-agent systems operating in uncertain environments. Due to the fact that the safe set of the local agent depends on the other agents’ states, the uncertainty of these external systems leads to an uncertain safe set. As a result, the safe control design of the local agent system in a multi-agent setting becomes intractable. To address this challenge, a neural network (NN) based adaptive observer is developed to estimate the state of the unknown external agents. Based on the state estimation of external agents, an adaptive interplay control barrier function (AI-CBF) is formulated. The AI-CBF is designed by considering both the local agent’s state and the NN-based estimated states of external agents. Notably, the limitation of forward invariance for the approximated safe set without guaranteeing the same for the actual safe set is acknowledged in AI-CBF design. The AI-CBF incorporates the bounds on state estimation errors of external agents to guarantee the strict safety requirements of the local agent while learning external agent dynamics. Then, a control framework is formulated using a quadratic programming (QP) method that integrates the safety and stability of the system.
Shawon Dey, Hao Xu 0002
IJCNN2
2024 Reverse-graph enhanced graph neural networks for session-based recommendation
Hao Xu 0002, Bo Yang 0011, Xiangkun Liu
Expert Syst. Appl.1
2024 Multi-level category-aware graph neural network for session-based recommendation
Bo Yang 0011, Hao Xu 0002, Wang Hu 0001
Expert Syst. Appl.3
2023 Joint Optimal Placement and Dynamic Resource Allocation for multi-UAV Enhanced Reconfigurable Intelligent Surface Assisted Wireless Network
abstract
In this paper, the optimal placement and dynamic resource allocation problem has been investigated for multi-UAV enhanced reconfigurable intelligent surface (RIS) assisted wireless network with uncertain time-varying wireless channels. This paper aims to stimulate the potential of RIS by adding mobility to RIS through unmanned aerial vehicles (UAV). A novel UAV optimal placement and dynamic resource allocation technique needs to be developed jointly. A novel online rein-forcement learning based optimal resource allocation algorithm has been designed. Firstly, a deep Q-learning based K-means clustering algorithm is utilized to optimize the deployment of the multi-UAV. Then, an online actor-critic reinforcement learning algorithm is developed to learn the optimal transmit power control as well as mobile RIS phase shift control policy. Compared with conventional learning algorithms, the developed algorithm can learn the optimal resource allocation and multi-UAV placement for mobile RIS-assisted wireless networks in real-time even with uncertain and time-varying wireless channels. Eventually, numerical simulations are provided to demonstrate the effectiveness of developed schemes.
Yuzhu Zhang, Lijun Qian, Hao Xu 0002
CCNC3
2022 Decentralized optimal large scale multi-player pursuit-evasion strategies: A mean field game approach with reinforcement learning
Zejian Zhou, Hao Xu 0002
Neurocomputing2
2022 Category-aware Multi-relation Heterogeneous Graph Neural Networks for session-based recommendation
Hao Xu 0002, Bo Yang 0011, Xiangkun Liu, Wenqi Fan, Qing Li 0001
Knowl. Based Syst.1
2022 A Novel Mean-Field-Game-Type Optimal Control for Very Large-Scale Multiagent Systems
abstract
In this article, a decentralized adaptive optimal controller based on the emerging mean-field game (MFG) and self-organizing neural networks (NNs) has been developed for multiagent systems (MASs) with a large population and uncertain dynamics. This design can effectively break the "curse of dimensionality" and reduce the computational complexity by appropriately integrating emerging MFG theory with self-organizing NNs-based reinforcement learning techniques. First, the decentralized optimal control for massive MASs has been formulated into an MFG. To unfold the MFG, the coupled Hamilton-Jacobian-Bellman (HJB) equation and Fokker-Planck-Kolmogorov (FPK) equation needed to be solved simultaneously, which is challenging in real time. Therefore, a novel actor-critic-mass (ACM) structure has been developed along with self-organizing NNs subsequently. In the developed ACM structure, each agent has three NNs, including: 1) mass NN learning the mass MAS's overall behavior via online estimating the solution of the FPK equation; 2) critic NN obtaining the optimal cost function through learning the HJB equation solution along with time; and 3) actor NN estimating the decentralized optimal control by using the critic and mass NNs along with the optimal control theory. To reduce the NNs' computational complexity, a self-organizing NN has been adopted and integrated into a developed ACM structure that can adjust the NNs' architecture based on the NNs' learning performance and the computation cost. Finally, numerical simulation has been provided to demonstrate the effectiveness of the developed schemes.
Zejian Zhou, Hao Xu 0002
IEEE Trans. Cybern.2
2022 Large-Scale Multiagent System Tracking Control Using Mean Field Games
abstract
This article studies the tracking control problem with a large-scale group of agents. Unlike traditional control techniques used in multiagent systems (MASs), a new type of intelligent design is needed to handle the intractable "Curse of Dimensionality" caused by the extremely large number of agents. To address this challenge, the mean field game (MFG) theory has been embedded into reinforcement learning to advance intelligent tracking control with large-scale MAS. Specifically, MFG-based control can calculate the optimal strategy based on one unified fix-dimension probability density function (pdf) instead of high-dimensional large-scale MAS information collected from individual agents. Moreover, the approximate dynamic programming technique is adopted to generate a new type of MFG-based algorithm. Each agent has three neural networks (NNs) to approximate the solution of the mean field type control. In addition to the algorithm development, the performance of the NNs is also analyzed using the Lyapunov method. Finally, the linear and nonlinear tracking control simulations are given to evaluate the algorithm's performance.
Zejian Zhou, Hao Xu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2021 Reinforcement Learning-based Decentralized Optimal Control for Large-Scale Multi-agent System by Using Neural Networks and Discrete-time Mean Field Games
abstract
This paper proposed the Actor-critic-mass (ACM) algorithm to solve the discrete-time mean field type tracking control problem without discretization errors. The traditional large-scale multi-agent control problem suffers from computational explosion and communication difficulty due to the team's nearly infinite agent number. To deal with those challenges, the Mean Field Games (MFG) have been introduced in the multi-agent control problem to disconnect the computation and communication complexity with the agent number. However, traditional mean field type control problems mainly studied the continuous- times systems limited by the concerns of losing the solutions' existence or uniqueness caused by discretization. In this paper, the Feynman-Kac formula is utilized to reformulate the continuous MFG equation system into a backward discretized system. Specifically, the Hamiltonian-Jacobi-Bellman (HJB) equation and backward Kolmogorov equation in the form of a backward stochastic differential equation (BSDE) are introduced. Meanwhile, three neural networks (NN), i.e., the actor, critic, and mass NN, are developed to approximate the MFG equation system's discretized solution and the optimal control. Moreover, the stability and convergence of the NNs and the closed-loop system are provided via the Lyapunov stability analysis. Finally, a series of numerical simulations have shown the new discrete algorithm's performance.
Zejian Zhou, Yuzhu Zhang, Hao Xu 0002
IJCNN3
2021 Decentralized Adaptive Optimal Tracking Control for Massive Autonomous Vehicle Systems With Heterogeneous Dynamics: A Stackelberg Game
abstract
In this article, a decentralized optimal tracking control problem has been studied for a large-scale autonomous vehicle system with heterogeneous system dynamics. Due to the ultralarge number of agents, the notorious “curse of dimension” problem as well as the unrealistic assumption of the existence of reliable very large-scale communication links in uncertain environments have challenged the traditional multiagent system (MAS) algorithms for decades. The emerging mean-field game (MFG) theory has recently been widely adopted to generate a decentralized control method that deals with those challenges by encoding the large scale MASs’ information into a novel time-varying probability density functions (PDF) which can be obtained locally. However, the traditional MFG methods assume all agents are homogeneous, which is unrealistic in practical industrial applications, e.g., Internet of Things (IoTs), and so on. Therefore, a novel mean-field Stackelberg game (MFSG) is formulated based on the Stackelberg game, where all the agents have been classified as two different categories where one major leader’s decision dominates the other minor agents. Moreover, a hierarchical structure that treats all minor agents as a mean-field group is developed to tackle the assumption of homogeneous agents. Then, the actor-actor–critic–critic-mass ($A^{2}C^{2}M$) algorithm with five neural networks is designed to learn the optimal policies by solving the MFSG. The Lyapunov theory is utilized to prove the convergence of$A^{2}C^{2}M$neural networks and the closed-loop system’s stability. Finally, a series of numerical simulations are conducted to demonstrate the effectiveness of the developed method.
Zejian Zhou, Hao Xu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2020 Distributed Consensus Control of Multiple UAVs in a Constrained Environment
abstract
In this paper, we investigate the consensus problem of multiple unmanned aerial vehicles (UAVs) in the presence of environmental constraints under a general communication topology containing a directed spanning tree. First, based on a position transformation function, we propose a novel dynamic reference position and yaw angle for each UAV to cope with both the asymmetric topology and the constraints. Then, the backstepping-like design methodology is presented to derive a local tracking controller for each UAV such that its position and yaw angle can converge to the reference ones. The proposed protocol is distributed in the sense that, the input update of each UAV dynamically relies only on local state information from its neighborhood set and the constraints, and it does not require any additional centralized information. It is demonstrated that under the proposed protocol, all UAVs reach consensus without violation of the environmental constraints. Finally, simulation and experimental results are provided to demonstrate the performance of the protocol.
Gang Wang 0024, Na Zhao 0008, Yunfeng Ji, Yantao Shen 0001, Hao Xu 0002, Peng Li 0019
ICRA6
2020 Zero-Sum Game Based Cooperative Control for Onboard Pulsed Power Load Accommodation
abstract
In this paper, the noncooperative control problem of onboard pulsed power load is formulated as a two-player zero-sum game. One player is the optimal controller which is designed to optimize a predefined cost function. The other player is the disturbance that represents the overall damping effect of the system, including that of unmodeled system dynamics. Neuro-dynamic programming based control design is developed to solve the nonlinear optimal control problem under disturbance. The neural network based control algorithm can achieve the near-optimal control without the acknowledgment of system dynamics. In addition, the control design can relax the requirements for initial admissible control conditions and predetermined control references. The effectiveness of the proposed algorithm is demonstrated through the real-time simulations on both simplified and detailed shipboard power system models.
Jiajun Duan, Hao Xu 0002, Wenxin Liu 0001, Jian-Chun Peng, Hui Jiang 0006
IEEE Trans. Ind. Informatics2
2019 MSNET-Blockchain: A New Framework for Securing Mobile Satellite Communication Network
abstract
In this paper, the security problem for mobile satellite communication networks (MSNET) has been investigated. With the rapidly growth of communication needs, mobile satellite systems represent a significant solution to provide high-quality communication services to mobile users in under-populated regions, in emergency areas, on planes, trains and ships. However, lacking an effective framework to secure mobile satellite communication networks seriously limited the practicality of satellite services. Therefore, a new security framework have been developed in this paper to address the security challenges in mobile satellite communication network. Firstly, the mobile satellite communication networks have been formulated as delay-tolerance network (DTN). Then, the blockchain technique has been adopted and used in two aspects, i.e. 1) integrating with DTN structure to secure the data communication, 2) combing with the practical satellite constellation management algorithm to defend the unexpected cyber attacks physically. Through integrating emerging blockchain techniques with both communication and physical aspects, the developed framework cannot only effectively detect the cyber attacks, but also better defend the mobile satellite communication networks through communication and satellite management aspects. Eventually, the numerical simulation and experimental tests have been provided to demonstrate the effectiveness of developed MSNET-Blockchain framework.
Ming Feng, Hao Xu 0002
SECON2
2019 A biologically-inspired distributed fault tolerant flocking control for multi-agent system in presence of uncertain dynamics and unknown disturbance
Mohammad Jafari 0001, Hao Xu 0002
Eng. Appl. Artif. Intell.2
2018 The Deformable Quad-Rotor Enabled and Wasp-Pedal-Carrying Inspired Aerial Gripper
abstract
The paper presents the development of a novel deformable quad-rotor enabled aerial gripper. The mechanism of our deformable quad-rotor is based on simultaneous expansion or contraction of the quad-rotor body, which is generated by controlling a rigid elements based morphing structure (REMS). Such deformation results in a highly deformable quad-rotor that can not only perform morphological adaptation in response to environmental changes and obstacles, but also improve the flight performance by contracting to facilitate the agility/maneuverability or by expanding to enhance the stability. Meanwhile, inspired by the wasp grasping behavior, such controllable expansion and contraction from the REMS ingeniously enable a new function of aerial gripper. In this paper, we start to detail the mechanism and design of the REMS based deformable quad-rotor, then present the quad-rotor deformation enabled aerial gripper design, its dynamics modeling, the grasping function and analysis. The simulation was conducted in order to graphically show the elicited aerodynamic flow situation during expansion or contraction of the quad-rotor with and without carrying payload. Experiments were further implemented to validate the grasping function of the gripper and the flight performance of the quad-rotor. Finally, two case studies on the new aerial gripper were performed. All results demonstrate the excellent performance of the deformable quad-rotor enabled aerial gripper, that is, it has the advantages of both flight maneuverability and grasping capability during performing tasks.
Na Zhao 0008, Yudong Luo, Hongbin Deng, Yantao Shen 0001, Hao Xu 0002
IROS5
2018 Distributed Control of Inverter-Interfaced Microgrids With Bounded Transient Line Currents
abstract
Distributed generators (DGs) in a microgrid are tightly coupled through power lines, whose dynamics should not be ignored. If not properly handled, large transient line currents may trigger false protection even under normal operating conditions. Droop-based control adjustments also unnecessarily increase frequency and voltage oscillations. Targeting at these problems, this paper presents a distributed control solution for inverter-interfaced microgrids. The objective of primary control is to realize the desired regulations of bus voltages and frequency as well as suppression of transient line currents. The objective of secondary control is to maintain fair load sharing. At secondary control level, a consensus algorithm is introduced to calculate the references for phase angles of bus voltages based on fair load sharing and dc power flow. At primary control level, a feedback linearization based control algorithm with dynamic control bounds is designed for voltage regulation and transient line current suppression. In addition to a common reference frame, the subsystem controllers only require measurements of local and neighboring subsystems. The effectiveness of the proposed control solution is demonstrated through simulations based on both simplified and detailed models.
Jiajun Duan, Cheng Wang 0012, Hao Xu 0002, Wenxin Liu 0001, Jian-Chun Peng, Hui Jiang 0006
IEEE Trans. Ind. Informatics3
2018 Power Converters Based Advanced Experimental Platform for Integrated Study of Power and Controls
abstract
With increasing interest in smart grid and renewable energy, significant investments have been allocated to promote related studies. Since there is a wide spectrum of topics to study, it is necessary to have an advanced experimental platform that can accommodate both system- and component-level studies, both hardware and algorithm designs, and both teaching and research. Unfortunately, such experimental platform is not commercially available. In this paper, an advanced power electronics based experimental platform is introduced. The system is consisted of one OPAL-RT real-time simulator, two one-bus microgrid testbeds, and two modular multilevel converters. The subsystems can form a multiple-bus microgrid testbed if connected through emulated power lines. The system can provide real-time simulation, controller hardware-in-the-loop (HIL) simulation, and power HIL simulation for power systems study. Various converter topologies can be configured with the modular converters for power electronics study. Both real-time simulator and DSP control boards can be used to implement advanced control algorithms. The designs and experiences shared in the paper will benefit many researchers that are in need of such system and promote power and energy related studies.
Wenxin Liu 0001, Jang-Mok Kim, Cheng Wang 0012, Won-Sang Im, Hao Xu 0002
IEEE Trans. Ind. Informatics6
2017 Approximate Optimal Control of Affine Nonlinear Continuous-Time Systems Using Event-Sampled Neurodynamic Programming
abstract
This paper presents an approximate optimal control of nonlinear continuous-time systems in affine form by using the adaptive dynamic programming (ADP) with event-sampled state and input vectors. The knowledge of the system dynamics is relaxed by using a neural network (NN) identifier with event-sampled inputs. The value function, which becomes an approximate solution to the Hamilton-Jacobi-Bellman equation, is generated by using event-sampled NN approximator. Subsequently, the NN identifier and the approximated value function are utilized to obtain the optimal control policy. Both the identifier and value function approximator weights are tuned only at the event-sampled instants leading to an aperiodic update scheme. A novel adaptive event sampling condition is designed to determine the sampling instants, such that the approximation accuracy and the stability are maintained. A positive lower bound on the minimum inter-sample time is guaranteed to avoid accumulation point, and the dependence of inter-sample time upon the NN weight estimates is analyzed. A local ultimate boundedness of the resulting nonlinear impulsive dynamical closed-loop system is shown. Finally, a numerical example is utilized to evaluate the performance of the near-optimal design. The net result is the design of an event-sampled ADP-based controller for nonlinear continuous-time systems.
Avimanyu Sahoo, Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.2
2016 Adaptive Neural Network-Based Event-Triggered Control of Single-Input Single-Output Nonlinear Discrete-Time Systems
abstract
This paper presents a novel adaptive neural network (NN) control of single-input and single-output uncertain nonlinear discrete-time systems under event sampled NN inputs. In this control scheme, the feedback signals are transmitted, and the NN weights are tuned in an aperiodic manner at the event sampled instants. After reviewing the NN approximation property with event sampled inputs, an adaptive state estimator (SE), consisting of linearly parameterized NNs, is utilized to approximate the unknown system dynamics in an event sampled context. The SE is viewed as a model and its approximated dynamics and the state vector, during any two events, are utilized for the event-triggered controller design. An adaptive event-trigger condition is derived by using both the estimated NN weights and a dead-zone operator to determine the event sampling instants. This condition both facilitates the NN approximation and reduces the transmission of feedback signals. The ultimate boundedness of both the NN weight estimation error and the system state vector is demonstrated through the Lyapunov approach. As expected, during an initial online learning phase, events are observed more frequently. Over time with the convergence of the NN weights, the inter-event times increase, thereby lowering the number of triggered events. These claims are illustrated through the simulation results.
Avimanyu Sahoo, Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.2
2016 Neural Network-Based Event-Triggered State Feedback Control of Nonlinear Continuous-Time Systems
abstract
This paper presents a novel approximation-based event-triggered control of multi-input multi-output uncertain nonlinear continuous-time systems in affine form. The controller is approximated using a linearly parameterized neural network (NN) in the context of event-based sampling. After revisiting the NN approximation property in the context of event-based sampling, an event-triggered condition is proposed using the Lyapunov technique to reduce the network resource utilization and to generate the required number of events for the NN approximation. In addition, a novel weight update law for aperiodic tuning of the NN weights at triggered instants is proposed to relax the knowledge of complete system dynamics and to reduce the computation when compared with the traditional NN-based control. Nonetheless, a nonzero positive lower bound for the inter-event times is guaranteed to avoid the accumulation of events or Zeno behavior. For analyzing the stability, the event-triggered system is modeled as a nonlinear impulsive dynamical system and the Lyapunov technique is used to show local ultimate boundedness of all signals. Furthermore, in order to overcome the unnecessary triggered events when the system states are inside the ultimate bound, a dead-zone operator is used to reset the event-trigger errors to zero. Finally, the analytical design is substantiated with numerical results.
Avimanyu Sahoo, Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.2
2016 Near Optimal Event-Triggered Control of Nonlinear Discrete-Time Systems Using Neurodynamic Programming
abstract
This paper presents an event-triggered near optimal control of uncertain nonlinear discrete-time systems. Event-driven neurodynamic programming (NDP) is utilized to design the control policy. A neural network (NN)-based identifier, with event-based state and input vectors, is utilized to learn the system dynamics. An actor-critic framework is used to learn the cost function and the optimal control input. The NN weights of the identifier, the critic, and the actor NNs are tuned aperiodically once every triggered instant. An adaptive event-trigger condition to decide the trigger instants is derived. Thus, a suitable number of events are generated to ensure a desired accuracy of approximation. A near optimal performance is achieved without using value and/or policy iterations. A detailed analysis of nontrivial inter-event times with an explicit formula to show the reduction in computation is also derived. The Lyapunov technique is used in conjunction with the event-trigger condition to guarantee the ultimate boundedness of the closed-loop system. The simulation results are included to verify the performance of the controller. The net result is the development of event-driven NDP.
Avimanyu Sahoo, Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.2
2015 Neural Network-Based Finite Horizon Stochastic Optimal Control Design for Nonlinear Networked Control Systems
abstract
The stochastic optimal control of nonlinear networked control systems (NNCSs) using neuro-dynamic programming (NDP) over a finite time horizon is a challenging problem due to terminal constraints, system uncertainties, and unknown network imperfections, such as network-induced delays and packet losses. Since the traditional iteration or time-based infinite horizon NDP schemes are unsuitable for NNCS with terminal constraints, a novel time-based NDP scheme is developed to solve finite horizon optimal control of NNCS by mitigating the above-mentioned challenges. First, an online neural network (NN) identifier is introduced to approximate the control coefficient matrix that is subsequently utilized in conjunction with the critic and actor NNs to determine a time-based stochastic optimal control input over finite horizon in a forward-in-time and online manner. Eventually, Lyapunov theory is used to show that all closed-loop signals and NN weights are uniformly ultimately bounded with ultimate bounds being a function of initial conditions and final time. Moreover, the approximated control input converges close to optimal value within finite time. The simulation results are included to show the effectiveness of the proposed scheme.
Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.1
2015 Finite-Horizon Near-Optimal Output Feedback Neural Network Control of Quantized Nonlinear Discrete-Time Systems With Input Constraint
abstract
The output feedback-based near-optimal regulation of uncertain and quantized nonlinear discrete-time systems in affine form with control constraint over finite horizon is addressed in this paper. First, the effect of input constraint is handled using a nonquadratic cost functional. Next, a neural network (NN)-based Luenberger observer is proposed to reconstruct both the system states and the control coefficient matrix so that a separate identifier is not needed. Then, approximate dynamic programming-based actor-critic framework is utilized to approximate the time-varying solution of the Hamilton-Jacobi-Bellman using NNs with constant weights and time-dependent activation functions. A new error term is defined and incorporated in the NN update law so that the terminal constraint error is also minimized over time. Finally, a novel dynamic quantizer for the control inputs with adaptive step size is designed to eliminate the quantization error overtime, thus overcoming the drawback of the traditional uniform quantizer. The proposed scheme functions in a forward-in-time manner without offline training phase. Lyapunov analysis is used to investigate the stability. Simulation results are given to show the effectiveness and feasibility of the proposed method.
Hao Xu 0002, Qiming Zhao, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.1
2015 Neural Network-Based Finite-Horizon Optimal Control of Uncertain Affine Nonlinear Discrete-Time Systems
abstract
In this paper, the finite-horizon optimal control design for nonlinear discrete-time systems in affine form is presented. In contrast with the traditional approximate dynamic programming methodology, which requires at least partial knowledge of the system dynamics, in this paper, the complete system dynamics are relaxed utilizing a neural network (NN)-based identifier to learn the control coefficient matrix. The identifier is then used together with the actor-critic-based scheme to learn the time-varying solution, referred to as the value function, of the Hamilton-Jacobi-Bellman (HJB) equation in an online and forward-in-time manner. Since the solution of HJB is time-varying, NNs with constant weights and time-varying activation functions are considered. To properly satisfy the terminal constraint, an additional error term is incorporated in the novel update law such that the terminal constraint error is also minimized over time. Policy and/or value iterations are not needed and the NN weights are updated once a sampling instant. The uniform ultimate boundedness of the closed-loop system is verified by standard Lyapunov stability theory under nonautonomous analysis. Numerical examples are provided to illustrate the effectiveness of the proposed method.
Qiming Zhao, Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.2
2014 Neural network-based adaptive optimal consensus control of leaderless networked mobile robots
abstract
A novel neural network (NN)-based optimal adaptive consensus control scheme is introduced in this paper for networked mobile robots in the presence of unknown robot dynamics. Throughout the paper, two NNs are used. The unknown formation dynamics of each robot is identified by using the first NN. The second NN is utilized to approximate a novel value function derived in this paper as a function of augmented error vector, which is comprised of the regulation and consensus-based formation errors of each robot. A novel near optimal controller is developed by using approximated value function and identified formation dynamics. The Lyapunov stability theorem is employed to derive the NN weight tuning laws and demonstrate the consensus achievement of the overall formation. The simulation results are depicted to show performance of our theoretical claims.
Haci Mehmet Guzey, Hao Xu 0002, Sarangapani Jagannathan
ADPRL2
2014 Event-based optimal regulator design for nonlinear networked control systems
abstract
This paper presents a novel stochastic event-based near optimal control strategy to regulate a networked control system (NCS) represented as an uncertain nonlinear continuous time system. An online stochastic actor-critic neural network (NN) based approach is utilized to achieve the near optimal regulation in the presence of network constraints, such as, network induced time-varying delays and random packet losses under event-based transmission of the feedback signals. The transformed nonlinear NCS in discrete-time after the incorporation the delays and packet losses is utilized for the actor-critic NN based controller design. To relax the knowledge of the control coefficient matrix, a NN based identifier is used. Event sampled state vector is utilized as NN inputs and their respective weights are updated non-periodically at the occurrence of events. Further, an event-trigger condition is designed by using the Lyapunov technique to ensure ultimate boundedness of all the closed-loop signals and save network resources and computation. Moreover, policy and value iterations are not utilized for the stochastic optimal regulator design. Finally, the analytical design is verified by using a numerical example by carrying out Monte-Carlo simulations.
Avimanyu Sahoo, Hao Xu 0002, Sarangapani Jagannathan
ADPRL2
2014 Model-free Q-learning over finite horizon for uncertain linear continuous-time systems
abstract
In this paper, a novel optimal control over finite horizon has been introduced for linear continuous-time systems by using adaptive dynamic programming (ADP). First, a new time-varying Q-function parameterization and its estimator are introduced. Subsequently, Q-function estimator is tuned online by using both Bellman equation in integral form and terminal cost. Eventually, near optimal control gain is obtained by using the Q-function estimator. All the closed-loop signals are shown to be bounded by using Lyapunov stability analysis where bounds are functions of initial conditions and final time while the estimated control signal converges close to the optimal value. The simulation results illustrate the effectiveness of the proposed scheme.
Hao Xu 0002, Sarangapani Jagannathan
ADPRL1
2014 Near optimal event-based control of nonlinear discrete time systems in affine form with measured input and output data
abstract
In this paper, an event-based near optimal control of uncertain nonlinear discrete time systems is presented by using input-output data and approximate dynamic programming (ADP). The nonlinear system dynamics in affine form are transformed into an input-output form. Then, three neural networks (NN) with event sampled input-output vector are used, namely, the identifier NN to relax the knowledge of the system dynamics, a critic NN to approximate the value function which is the solution to the Hamilton-Jacobi Bellman (HJB) equation, and an actor NN to approximate the optimal control policy, in an online manner without utilizing value or policy iterations. In addition, the NN weights of all the three NNs are tuned only at event-triggered instants leading to a novel non-periodic update rule to reduce computation when compared to traditional NN based scheme. Further, an event-trigger condition to decide the trigger instants is derived. Finally, the Lyapunov technique is used in conjunction with the event-trigger condition to guarantee the uniform ultimate boundedness (UUB) of the closed-loop system. The analytical design is substantiated with numerical results via simulation.
Avimanyu Sahoo, Hao Xu 0002, Sarangapani Jagannathan
IJCNN2
2014 Finite horizon stochastic optimal control of nonlinear two-player zero-sum games under communication constraint
abstract
In this paper, the finite horizon stochastic optimal control of nonlinear two-player zero-sum games, referred to as Nonlinear Networked Control Systems (NNCS) two-player zero-sum game, between control and disturbance input players in the presence of unknown system dynamics and a communication network with delays and packet losses is addressed by using neuro dynamic programming (NDP). The overall objective being to find the optimal control input while maximizing the disturbance attenuation. First, a novel online neural network (NN) identifier is introduced to estimate the unknown control and disturbance coefficient matrices which are needed in the generation of optimal control input. Then, the critic and two actor NNs have been introduced to learn the time-varying solution to the Hamilton-Jacobi-Isaacs (HJI) equation and determine the stochastic optimal control and disturbance policies in a forward-in-time manner. Eventually, with the proposed novel NN weight update laws, Lyapunov theory is utilized to demonstrate that all closed-loop signals and NN weights are uniformly ultimately bounded (ÜUB) during the finite horizon with ultimate bounds being a function of initial conditions and final time. Further, the approximated control input and disturbance signals tend close to the saddle-point equilibrium within finite-time. Simulation results are included.
Hao Xu 0002, Sarangapani Jagannathan
IJCNN1
2013 Finite horizon stochastic optimal control of uncertain linear networked control system
abstract
In this paper, finite horizon stochastic optimal control issue has been studied for linear networked control system (LNCS) in the presence of network imperfections such as network-induced delays and packet losses by using adaptive dynamic programming (ADP) approach. Due to an uncertainty in system dynamics resulting from network imperfections, the stochastic optimal control design uses a novel adaptive estimator (AE) to solve the optimal regulation of uncertain LNCS in a forward-in-time manner in contrast with backward-in-time Riccati equation-based optimal control with known system dynamics. Tuning law for unknown parameters of AE has been derived. Lyapunov theory is used to show that all the signals are uniformly ultimately bounded (UUB) with ultimate bounds being a function of initial values and final time. In addition, the estimated control input converges to optimal control input within finite horizon. Simulation results are included to show the effectiveness of the proposed scheme.
Hao Xu 0002, Sarangapani Jagannathan
ADPRL1
2013 Finite-horizon optimal control design for uncertain linear discrete-time systems
abstract
In this paper, the finite-horizon optimal adaptive control design for linear discrete-time systems with unknown system dynamics by using adaptive dynamic programming (ADP) is presented. In the presence of full state feedback, the terminal state constraint is incorporated in solving the optimal feedback control via the Bellman equation. The optimal regulation of the uncertain linear system is solved in a forward-in-time and online manner without using value and/or policy iterations. Due to the nature of finite horizon, the stability of the closed-loop system is involved but verified by using Lyapunov theory. The effectiveness of the proposed method is verified by simulation results.
Qiming Zhao, Hao Xu 0002, Sarangapani Jagannathan
ADPRL2
2013 Solutions to finite horizon cost problems using actor-critic reinforcement learning
abstract
Actor-critic reinforcement learning algorithms have shown to be a successful tool in learning the optimal control for a range of (repetitive) tasks on systems with (partially) unknown dynamics, which may or may not be nonlinear. Most of the reinforcement learning literature published up to this point only deals with modeling the task at hand as a Markov decision process with an infinite horizon cost function. In practice, however, it is sometimes desired to have a solution for the case where the cost function is defined over a finite horizon, which means that the optimal control problem will be time-varying and thus harder to solve. This paper adapts two previously introduced actor-critic algorithms from the infinite horizon setting to the finite horizon setting and applies them to learning a task on a nonlinear system, without needing any assumptions or knowledge about the system dynamics, using radial basis function networks. Simulations on a typical nonlinear motion control problem are carried out, showing that actor-critic algorithms are capable of solving the difficult problem of time-varying optimal control. Moreover, the benefit of using a model learning technique is shown.
Ivo Grondman, Hao Xu 0002, Sarangapani Jagannathan, Robert Babuska
IJCNN2
2013 Neural network based finite horizon stochastic optimal controller design for nonlinear networked control systems
abstract
Existing neuro-dynamic programming (NDP) techniques are not applicable for optimizing real-time NNCS with terminal constraints during the finite horizon. Therefore, a novel time-based NDP scheme is developed in this paper to solve finite horizon optimal control of NNCS. First, an online neural network (NN) identifier is introduced to approximate the control coefficient matrix. Then, the critic and action NNs are utilized to determine time-based finite horizon stochastic optimal control for NNCS in a forward-in-time manner. By incorporating novel NN weight update laws, Lyapunov theory is used to show that all closed-loop signals and NN weights are uniformly ultimately bounded (UUB) with ultimate bounds being a function of initial conditions and final time. Moreover, the approximated control input converges close to target value within finite time. Simulation results are included to show the effectiveness of the proposed scheme.
Hao Xu 0002, Sarangapani Jagannathan
IJCNN1
2013 Finite-horizon neural network-based optimal control design for affine nonlinear continuous-time systems
abstract
In this paper, the finite-horizon optimal control design for affine nonlinear continuous-time systems in the presence of known system dynamics is presented. A neural network (NN) is utilized to learn the time-varying solution of the Hamilton-Jacobi-Bellman (HJB) equation in an online and forward in time manner. To handle the time varying nature of the value function, the NN with constant weights and time-varying activation function is considered. The update law for tuning the NN weights is derived based on normalized gradient descent approach. To satisfy the terminal constraint and ensure stability, additional terms, one corresponding to the terminal constraint, and the other to stabilize the nonlinear system are added to the novel updating law. A uniformly ultimately boundedness of the non-autonomous closed-loop system is verified by using standard Lyapunov theory. The effectiveness of the proposed method is verified by simulation results.
Qiming Zhao, Hao Xu 0002, Travis Dierks, Sarangapani Jagannathan
IJCNN2
2013 Stochastic Optimal Controller Design for Uncertain Nonlinear Networked Control System via Neuro Dynamic Programming
abstract
The stochastic optimal controller design for the nonlinear networked control system (NNCS) with uncertain system dynamics is a challenging problem due to the presence of both system nonlinearities and communication network imperfections, such as random delays and packet losses, which are not unknown a priori. In the recent literature, neuro dynamic programming (NDP) techniques, based on value and policy iterations, have been widely reported to solve the optimal control of general affine nonlinear systems. However, for realtime control, value and policy iterations-based methodology are not suitable and time-based NDP techniques are preferred. In addition, output feedback-based controller designs are preferred for implementation. Therefore, in this paper, a novel NNCS representation incorporating the system uncertainties and network imperfections is introduced first by using input and output measurements for facilitating output feedback. Then, an online neural network (NN) identifier is introduced to estimate the control coefficient matrix, which is subsequently utilized for the controller design. Subsequently, the critic and action NNs are employed along with the NN identifier to determine the forward-in-time, time-based stochastic optimal control of NNCS without using value and policy iterations. Here, the value function and control inputs are updated once a sampling instant. By using novel NN weight update laws, Lyapunov theory is used to show that all the closed-loop signals and NN weights are uniformly ultimately bounded in the mean while the approximated control input converges close to its target value with time. Simulation results are included to show the effectiveness of the proposed scheme.
Hao Xu 0002, Sarangapani Jagannathan
IEEE Trans. Neural Networks Learn. Syst.1
2012 Adaptive dynamic programming-based state quantized networked control system without value and/or policy iterations
abstract
In this paper, the Bellman equation is used to solve the stochastic optimal control of unknown linear discrete-time system with communication imperfections including random delays, packet losses and quantization. A dynamic quantizer for the sensor measurements is proposed which essentially provides system states to the controller. To eliminate the effect of the quantization error, the dynamics of the quantization error bound and an update law for tuning its range are derived. Subsequently, by using adaptive dynamic programming technique, the infinite horizon optimal regulation of the uncertain NCS is solved in a forward-in-time manner without using value and/or policy iterations by using Q-function and reinforcement learning. The asymptotic stability of the closed-loop system is verified by standard Lyapunov stability theory. Finally, the effectiveness of the proposed method is verified by simulation results.
Qiming Zhao, Hao Xu 0002, Sarangapani Jagannathan
IJCNN2
2012 A cross layer approach to the novel distributed scheduling protocol and event-triggered controller design for Cyber Physical Systems
abstract
In Cyber Physical Systems (CPS), multiple real-time dynamic systems are connected to their corresponding controllers through a shared communication network in contrast with a dedicated line. For such CPS, the existing scheduling schemes for instance Centralized or Distributed Scheduling schemes are unsuitable since the behavior of real-time dynamic systems is ignored during the network protocol design. Therefore, in this paper, a novel distributed scheduling protocol design via cross-layer approach is proposed to optimize the performance of CPS by maximizing the utility function which is generated by using the information collected from both the application and network layers. Subsequently, a novel adaptive model based optimal event-triggered control scheme is derived for each real-time dynamic system with unknown system dynamics in the application layer. Compared with traditional scheduling algorithms, the proposed distributed scheduling scheme via cross-layer approach can not only allocates the network resources efficiently but also improves the performance of the overall real-time dynamic system.
Hao Xu 0002, Sarangapani Jagannathan
LCN1