Xiaodong Xu 0002

dblp:43/2085-2 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0003-4795-9967ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Neural network with reduced-order GSINDy-Koopman operators for modeling and robust control of nonlinear processes
Zihang Xu, Yu Xiao 0006, Yuan Yuan 0017, Xiaodong Xu 0002, Biao Luo 0001
Neurocomputing4
2026 Reinforcement Learning-Based MPC for Output Regulation of Uncertain Constrained Linear System
abstract
This paper presents a reinforcement learning-based model predictive control (RL-MPC) output regulation approach for uncertain constrained linear systems. The proposed approach addresses key challenges in trajectory tracking under unknown parameters, state-input constraints, and external disturbances. By transforming the original system dynamics and constraint sets via regulator equations, the output regulation problem (ORP) is reformulated as a constrained stabilization for the tracking error dynamics. An MPC scheme is then constructed based on the uncertain transformed model, providing constraint-satisfying control policies applied to the actual system. Meanwhile, RL employs the parameterized MPC as a function approximator, recursively updating both regulator and MPC parameters online to enable adaptive policy optimization via continuous state feedback. A safety validation mechanism is integrated into the parameter learning to ensure recursive feasibility and closed-loop stability. The effectiveness of the proposed scheme is verified by a speed regulation simulation of a permanent magnet synchronous motor (PMSM) system.
Yu Xiao 0006, Yuan Yuan 0017, Biao Huang 0001, Xiaodong Xu 0002
IEEE Trans Autom. Sci. Eng.5
2026 Time-Varying HJBE-Based Adaptive Safe Critic Control Design for Stochastic Asymmetric Constrained Multiagent Systems
abstract
In this article, we investigate the problem of adaptive safe critic control design for stochastic multiagent systems (MASs) subject to asymmetric state and input constraints. To systematically address asymmetric state constraints, a unified transformation function (UTF) is proposed to convert the constrained consensus control problem into the stability analysis of an unconstrained error system. In addition, a nonquadratic cost function is incorporated to address input limitations effectively. Building upon these developments, a time-varying Hamilton-Jacobi-Bellman equation (HJBE) is formulated by integrating the Bellman optimality principle with Itô's lemma, thereby accommodating stochastic disturbances and enhancing controller robustness. To improve data utilization and eliminate reliance on explicit drift dynamics, an integral reinforcement learning (IRL) algorithm is developed within this framework. Furthermore, a time-varying single-critic network is designed to approximate the solution to the HJBE and generate optimal control policies, thereby considerably reducing computational complexity. To further enhance learning efficiency and relax the persistent excitation (PE) condition, the experience replay (ER) technique is incorporated into the update process of the critic weight. Finally, two simulation examples are provided to verify the feasibility and effectiveness of the proposed approach.
Yuhao Zhou 0001, Biao Luo 0001, Xiaodong Xu 0002, Yalin Wang 0003, Weihua Gui 0001
IEEE Trans. Cybern.3
2026 IRL-Based Optimal Consensus Control of MASs With Predefined Time Convergence Performance
abstract
This study addresses the adaptive optimal consensus control problem for nonlinear multiagent systems (MASs). To enhance the convergence speed of the consensus error, a predefined time performance technique is integrated into the optimal control framework. Unlike the conventional Hamilton–Jacobi–Bellman equation (HJBE), a time-varying HJBE is formulated to solve the optimal control problem for MASs. To further improve efficiency, an integral reinforcement learning (IRL) algorithm is developed, which eliminates the need for precise system dynamics during controller design. In addition, a single critic network is employed to simultaneously evaluate system performance and execute control actions, effectively reducing computational complexity. The experience replay technique is incorporated into the update law for the critic network weights, thus alleviating the requirement for persistent excitation. Finally, a simulation example is presented to validate the feasibility and effectiveness of the proposed method.
Yuhao Zhou 0001, Biao Luo 0001, Xiaodong Xu 0002
IEEE Trans. Syst. Man Cybern. Syst.4
2025 Adaptive neural boundary control for multi-agent manipulators system with uncertainties through cooperative disturbance observers network
Zhibo Zhao, Yuan Yuan 0017, Xiaodong Xu 0002, Biao Luo 0001, Tingwen Huang
Eng. Appl. Artif. Intell.3
2025 Auto-tuning strategy for deep Koopman robust model predictive control design based on advanced Metaheuristics
Zijia Meng, Yu Xiao 0006, Xiaodong Xu 0002, Biao Luo 0001, Yuncheng Du, Weihua Gui 0001
Neurocomputing3
2025 Neural operator-based composite learning adaptive backstepping control for linear 2×2 hyperbolic PDE systems
Yu Xiao 0006, Xiaodong Xu 0002, Biao Luo 0001, Weihua Gui 0001, Tingwen Huang
Neurocomputing3
2025 Offline reward shaping with scaling human preference feedback for deep reinforcement learning
Biao Luo 0001, Xiaodong Xu 0002, Tingwen Huang
Neural Networks3
2025 Reinforcement Learning-Based 3D Trajectory Tracking Control of Hypersonic Gliding Vehicles With Time-Varying Uncertainties
abstract
In this paper, a robust three-dimensional trajectory tracking control scheme based on reinforcement learning is proposed for the glide phase of a hypersonic gliding vehicle (HGV) with time-varying uncertainties. First, the non-affine nonlinear full-state kinematics and dynamics model of the HGV glide phase is constructed. Then, without linearizing the system, the desired multiplanar reference trajectories for HGVs are planned based on the pseudo-spectral theory under the input constraints, initial conditions, and terminal conditions. Subsequently, the full-state error system is generated by subtracting the reference system state from the actual state of the HGV system with time-varying uncertainty. For the full-state HGV error system with time-varying uncertainty and input constraints, we design a reinforcement learning-based optimal control scheme for its nominal system and establish the equivalence between this optimal control and the robust control of the original HGV error system. A single-evaluation network structure is used in the concrete implementation to reduce the computational cost. A rigorous theory is given to demonstrate the uniform ultimate boundedness of the closed-loop system and the weight error. Finally, we perform simulation traces for reference trajectories with different optimization performances to verify the effectiveness of the proposed method. Note to Practitioners—There are various constraints and uncertainties in the glide phase of HGVs, which is the hinge connecting the initial descent phase and the terminal management phase. How to design robust trajectory tracking controllers for the glide phase of HGVs with complex environments and large span of flight parameters is of great significance to aerial guidance practitioners. In this paper, an RL-based three-dimensional trajectory robust tracking guidance method is proposed for the HGV glide phase system, which can resist time-varying uncertainties and satisfy flight constraints. The uniform ultimate boundedness of the closed-loop system is proved using the Lyapunov method. The proposed tracking algorithm is effective for reference trajectories with different performance indexes.
Biao Luo 0001, Xiaodong Xu 0002
IEEE Trans Autom. Sci. Eng.4
2025 Composite Learning Based Adaptive Control of Linear 2 × 2 Hyperbolic PDE Systems
abstract
This article considers the adaptive stability control of a class of linear hyperbolic PDE systems. The PDE model is subject to constant but in-domain and boundary unknown parameters. A novel adaptive controller is developed by leveraging the swapping design technique and composite parameter learning law. With swapping design, several linear and static combinations, including carefully designed filters, unknown parameters, and error terms, are constructed to express the system states. From the static combinations, a composite learning based forgetting-factor least squares law is introduced to guarantee exponential parameter convergence without the persistent excitation (PE). Although inaccurate parameter estimation in the adaptive backstepping control results in asymptotic stability of the system, accurate parameter estimation ensures the exponential convergence of closed-loop system and concomitantly improves the transient performance. Finally, a comparative numerical simulation is performed to validate the effectiveness and advantage of the developed adaptive control scheme.
Yu Xiao 0006, Yun Feng 0001, Biao Luo 0001, Han-Xiong Li, Xiaodong Xu 0002
IEEE Trans. Cybern.5
2025 Multistep Q-Learning-Based Optimal Consensus Control of Linear Discrete-Time Multiagent Systems
abstract
This article considers the optimal consensus control for the multiagent systems problem. By developing the multiagent multistep Q-learning (MaMsQL), the methodology achieves enhanced efficiency while addressing the issue of the complex interaction dynamics between agents, environmental uncertainty, thus ultimately meeting demand of balancing exploration and exploitation. First, associated with the performance index, the Q-function is established to prove that all optimal Q-functions form a Nash equilibrium outcome, thereby the consensus problem is converted to finding the optimal Q-functions. Then, the MaMsQL method is developed with theoretical proof of its convergence. Finally, the method is implemented through a specially designed Actor-Critic network. By virtue of the comparison with multiagent single step Q-learning, the effectiveness and superiority of this method are verified through simulation examples.
Jialin Xiao, Biao Luo 0001, Xiaodong Xu 0002, Chunhua Yang 0001, Weihua Gui 0001
IEEE Trans. Cybern.3
2025 Consensus of Nonlinear Uncertain Delayed Multiagent Systems Modeled by PDEs via Adaptive Boundary Control
abstract
Under the influence of nonlinearity, time-varying delay, and uncertainty, the consensus problem is concerned in this study for multiagent systems modeled by partial differential equations, which means that both the time and space variables are included in the dynamic behavior of each agent. First, with a directed graph, an adaptive boundary controller is developed under boundary measurements, which can effectively reduce the control cost with dynamic control gains and a few actuators and sensors installed at the boundary of the spatial domain. Then, through the designed adaptive boundary controller, the linear matrix inequality (LMI)-based consensus conditions are obtained to ensure the exponential stability of the consensus error systems derived by utilizing the inequality techniques and Lyapunov direct approach. Lastly, two numerical examples demonstrate the effectiveness of the presented adaptive boundary control protocols.
Xu Zhang 0051, Biao Luo 0001, Zipeng Wang 0001, Xiaodong Xu 0002, Chunhua Yang 0001
IEEE Trans. Cybern.4
2025 Adaptive Neural Consensus Observer Networks Design for a Class of Semilinear Parabolic PDE Systems
abstract
This article concerns the investigation on the consensus problem for the joint state-uncertainty estimation of a class of parabolic partial differential equation (PDE) systems with parametric and nonparametric uncertainties. We propose a two-layer network consisting of informed and uninformed boundary observers where novel adaptation laws are developed for the identification of uncertainties. Particularly, all observer agents in the network transmit their information with each other across the entire network. The proposed adaptation laws include a penalty term of the mismatch between the parameter estimates generated by the other observer agents. Moreover, for the nonparametric uncertainties, radial basis function (RBF) neural networks are employed for the universal approximation of unknown nonlinear functions. Given the persistently exciting condition, it is shown that the proposed network of adaptive observers can achieve exponential joint state-uncertainty estimation in the presence of parametric uncertainties and ultimate bounded estimation in the presence of nonparametric uncertainties based on the Lyapunov stability theory. The effects of the proposed consensus method are demonstrated through a typical reaction-diffusion system example, which implies convincing numerical findings.
Mingxing Cai, Yuan Yuan 0017, Biao Luo 0001, Fanbiao Li, Xiaodong Xu 0002, Chunhua Yang 0001, Weihua Gui 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Adaptive Boundary Control for Synchronization of Reaction-Diffusion Neural Networks With Random Time-Varying Delay
abstract
This article addresses the synchronization problem of reaction-diffusion neural networks (RDNNs) with random time-varying delay (RTVD) via boundary control (BC) (including adaptive BC and BC with constant-valued gain) under distributed measurements or boundary measurements. First, a novel BC strategy with constant-valued gain is designed, which considers three cases of the measurements, that is, distributed measurements, boundary measurements, and both coexist. Subsequently, an adaptive BC scheme under boundary measurements is proposed, where the control gain is regulated effectively. Next, based on the inequality techniques and Lyapunov direct approach, the delay-dependent synchronization conditions are gained and some linear matrix inequalities (LMIs) based theorems are given. Then, the BC design for the delayed RDNNs is transformed into an LMI feasibility problem. Finally, the developed BC approaches are validated by the simulation results.
Xu Zhang 0051, Biao Luo 0001, Zipeng Wang 0001, Xiaodong Xu 0002, Chunhua Yang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Concurrent learning adaptive boundary observer design for linear coupled hyperbolic partial differential equation systems
Linbin Teng, Yuan Yuan 0017, Xiaodong Xu 0002, Chunhua Yang 0001, Biao Luo 0001, Stevan Dubljevic, Tingwen Huang
Knowl. Based Syst.3
2024 Neural operators for robust output regulation of hyperbolic PDEs
Yu Xiao 0006, Yuan Yuan 0017, Biao Luo 0001, Xiaodong Xu 0002
Neural Networks4
2024 RL-Based Adaptive Optimal Bipartite Consensus Control for Nonlinear Heterogeneous MASs via Event-Triggered State Feedback
abstract
This article investigates a leader-following bipartite consensus issue for uncertain nonlinear heterogeneous multiagent systems (MASs). Initially, within the framework of optimal control theory, we employ the reinforcement learning (RL) algorithm to derive an approximate solution to the Hamilton-Jacobi-Bellman equation (HJBE). Specifically, the neural networks (NNs) are utilized to construct the Actor-Critic structure with the aim of implementing control behavior and evaluating system performance, respectively. An additional network is employed to address nonlinear uncertainties existing in the system. Furthermore, we design a static threshold event-triggered mechanism (ETM) to achieve the event-triggered state feedback-based control strategy. By utilizing this event-triggered state information, we reconstruct the approximate optimal controller and update laws of neural network weights, effectively reducing the communication burden while ensuring that all signals of the MASs remain bounded. Finally, two simulation examples are carried out to demonstrate the feasibility of the proposed method.
Yuhao Zhou 0001, Biao Luo 0001, Xin Wang 0028, Xiaodong Xu 0002, Lin Xiao 0002
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Boundary Optimal Control for Parabolic Distributed Parameter Systems With Value Iteration
abstract
A reinforcement learning-based boundary optimal control algorithm for parabolic distributed parameter systems is developed in this article. First, a spatial Riccati-like equation and an integral optimal controller are derived in infinite-time horizon based on the principle of the variational method, which avoids the complex semigroups and operator theories. Using state data along the system trajectory, a value iteration algorithm via the Bellman optimality principle is proposed to obtain the solution of the spatial Riccati-like equation and the optimal control law. The convergence of the value iteration algorithm is proved. Subsequently, an approximation scheme based on weighted residuals is developed to implement the value iteration algorithm, where radial basis functions are chosen as the basic functions to approximate the solution of the spatial Riccati-like equation. Simulations on the diffusion-reaction process demonstrate the effectiveness of the developed method.
Biao Luo 0001, Xiaodong Xu 0002, Chunhua Yang 0001
IEEE Trans. Cybern.3
2024 Adaptive Neural Tracking Control of a Class of Hyperbolic PDE With Uncertain Actuator Dynamics
abstract
This article investigates the adaptive neural tracking control problem for a class of hyperbolic PDE with boundary actuator dynamics described by a set of nonlinear ordinary differential equations (ODEs). Particularly, the control input appears in the ODE subsystem with unknown nonlinearities requiring to be estimated and compensated, which makes the control task rather difficult. It is the first time to consider tracking control of such a class of systems, rendering our contributions essentially different from the existing literature that merely focus on the stabilization problem. By formulating a virtual exosystem to generate a reference trajectory, we propose a novel design of the adaptive geometric controller for the considered system where neural networks (NNs) are employed to approximately estimate nonlinearities, and finite and infinite-dimensional backstepping techniques are leveraged. Moreover, rigorously theoretical proofs based on the Lyapunov theory are provided to analyze the stability of the closed-loop system. Finally, we illustrate the results through two numerical simulations.
Yu Xiao 0006, Yuan Yuan 0017, Chunhua Yang 0001, Biao Luo 0001, Xiaodong Xu 0002, Stevan Dubljevic
IEEE Trans. Cybern.5
2024 Decentralized Multiagent Reinforcement Learning Based State-of-Charge Balancing Strategy for Distributed Energy Storage System
abstract
State-of-charge (SoC) balancing in distributed energy storage systems (DESS) is crucial but challenging. Traditional deep reinforcement learning approaches struggle with real-world multiagent cooperation for SoC balance in these decentralized systems. To address these significant hurdles, this article pioneers an innovative fully-decentralized multiagent reinforcement learning (FDMARL) strategy, specifically tailored for DESS. First, the SoC balancing problem is formulated into a finite decentralized Markov decision process with action constraints derived from the microgrid context. Then, the average consensus algorithm is introduced for expanding the agent's observation and obtaining global information through a communication network. To improve multiagent system cooperation, a novel demand balance algorithm is proposed to refine agent actions for precise demand distribution. By the above modules, the FDMARL reveals outstanding performance in a fully-decentralized system without any expert experience or modeling. Finally, numerous simulations were carried out on pymgrid, (i.e., a Python-based open-source microgrid platform), which shows that: The FDMARL exceeds traditional reinforcement learning in decentralized cooperation, effectively manages large-scale systems with random states and demand/PV series, and maintains robustness with ESU failure or integration.
Zheng Xiong, Biao Luo 0001, Bing-Chuan Wang, Xiaodong Xu 0002, Xiaodong Liu 0011, Tingwen Huang
IEEE Trans. Ind. Informatics4
2024 Concurrent Learning Robust Adaptive Fault Tolerant Boundary Regulation of Hyperbolic Distributed Parameter Systems
abstract
This article develops a robust adaptive boundary output regulation approach for a class of complex anticollocated hyperbolic partial differential equations subjected to multiplicative unknown faults in both the boundary sensor and actuator. The regulator design is based on the internal model principle, which amounts to stabilize a coupled cascade system, which consists of a finite-dimensional internal model driven by a hyperbolic distributed parameter system (DPS). To this end, a systematic sliding mode equipped with a backstepping approach is developed such that the robust state feedback control can be realized. Moreover, since the available information is a faulty boundary measurement at the right side point, state estimation is required. However, due to the presence of boundary unknown faults, we need to solve an issue of joint fault-state estimation. Restrictive persistent excitation conditions are usually required to guarantee the exact estimation of faults but are unrealistic in practice. To this end, a novel concurrent learning (CL) adaptive observer is proposed so that exponential convergence is obtained. It is the first time that the spirit of CL is introduced to the field of DPSs. Consequently, the observer-based adaptive boundary fault tolerant control scheme is developed, and rigorous theoretical analysis is given such that the exponential output regulation can be achieved. Finally, the effectiveness of the proposed methodology is demonstrated via comparative simulations.
Yuan Yuan 0017, Xiaodong Xu 0002, Chunhua Yang 0001, Biao Luo 0001, Stevan Dubljevic
IEEE Trans. Neural Networks Learn. Syst.2
2024 Adaptive Exponential Fault Estimation for 1-D Linear Parabolic PDEs With Process Uncertainties
abstract
The problem of fault estimation is addressed for one-dimensional (1-D) linear boundary control and boundary observation (BCBO) parabolic partial differential equations (PDEs) with a faulty boundary measurement. The considered plant is subjected to simultaneous unknown multiplicative faults entering the boundary input and boundary measurement. Difficulties arise due to the coupling between the sensor fault parameter and unknown boundary state appearing in the measurement. With the only boundary input and faulty boundary measurement, it is rather challenging to estimate the accurate values of faults and state simultaneously. Therefore, most existing results only consider correct and healthy measurement for PDE systems. To this end, novel adaptation laws and an adaptive observer are designed in this work to provide exponential convergent joint fault-state estimation, where we design and leverage a set of novel filters. It is first time that unknown multiplicative fault parameter in the measurement can be estimated accurately in the PDE systems.
Yuan Yuan 0017, Xiaodong Xu 0002, Chunhua Yang 0001, Tingwen Huang, Stevan Dubljevic
IEEE Trans. Syst. Man Cybern. Syst.2
2023 LJIR: Learning Joint-Action Intrinsic Reward in cooperative multi-agent reinforcement learning
Biao Luo 0001, Tianmeng Hu, Xiaodong Xu 0002
Neural Networks4
2023 Adaptive Fuzzy Boundary Observer Design for Uncertain Linear Coupled Hyperbolic Partial Differential Equation Systems
abstract
Joint uncertainties and state estimation of a class of linear coupled hyperbolic partial differential equation systems in the presence of unstructured and structured uncertainties are studied in this paper. For unstructured uncertainties which are completely unknown, by employing Takagi-Sugeno fuzzy logic system to approximate the unstructured uncertainties, a novel adaptive fuzzy boundary observer is developed to estimate both unknown system states as well as unknown weights in the fuzzy logic system, and the estimation errors are ultimately bounded. Therein, in the design of the proposed observer, a set of swapping filters and infinite dimensional backstepping technique are combined. On the other hand, for structured uncertainties that can be described in a concrete parameterized form, the proposed method can easily achieve the exact estimation of weights and states to their true values. The rigorous proof is provided to show that the ultimately bounded estimation errors for the case of unstructured uncertainties and the exponential convergent estimation errors for the case of structured uncertainties can be realized. Finally, three illustrative simulations are carried out to show the feasibility and effectiveness of the developed methods in this paper.
Linbin Teng, Yuan Yuan 0017, Biao Luo 0001, Chunhua Yang 0001, Stevan Dubljevic, Tingwen Huang, Xiaodong Xu 0002
IEEE Trans. Fuzzy Syst.7