VLDB 2026 Research / reviewers in the wild / expert
Qinglai Wei
dblp:44/4380
· DBLP profile ↗
164ranked-venue papers
62as first author
70since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 119 · 49 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 6 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 18 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward more generalizable time-series outlier detection in large-volume manufacturing via contextual convolution
Qinglai Wei, Yanbin Du, Dongdong Ni |
Neurocomputing | 2 |
| 2026 | Optimal parallel control for nonlinear multi-agent systems with state constraints
Qinglai Wei |
Neurocomputing | 2 |
| 2026 | Physics-informed spiking neural networks for continuous-time dynamic systems
Qinglai Wei, Qizhi Yang, Liyuan Han, Tielin Zhang |
Neurocomputing | 1 |
| 2026 | Attack Detection and Active Attack Defense for Cyber-Physical Systems via Zonotopic Observer and Reachability AnalysisabstractThis article concentrates on the attack detection and active attack defense strategies for discrete-time linear cyber-physical systems (CPSs) with unknown but bounded (UBB) disturbance and noise in the presence of both actuator and sensor attacks. First, a novel zonotopic observer is constructed to estimate the set-valued state and actuator attack by introducing augmentation techniques. To mitigate the effects of uncertainty and enhance estimation accuracy, the $H_{\infty }$ technique is introduced to construct the observer. Unlike most existing works, the constructed observer simultaneously estimates the system state and actuator attacks. Then, by combining the designed observer with reachability analysis, a set-valued abnormal detector and a residual-based abnormal detector are designed to detect actuator and sensor attacks, respectively. In addition, by incorporating the obtained state reachable sets and the $H_{\infty }$ technique, an active attack defense mechanism is designed to mitigate the impact of attacks on system performance. The proposed defense strategy does not introduce any performance loss in the absence of attacks. Finally, the superiority of the developed method is demonstrated by its application to a numerical simulation and an autonomous aircraft system. Zhihua Guo 0001, Qinglai Wei, Xudong Zhao 0001, Bohui Wang, Ben Niu 0003, Hao Liu 0012 |
IEEE Trans. Cybern. | 2 |
| 2025 | Guest Editorial: Emerging Trends in Safety-Critical Issues for Intelligent Automation SystemsabstractIntelligent Automation Systems, such as automated storage and retrieval systems, self-driving vehicles, various types of autonomous robots, and autonomous workshop plants, act independently of direct human supervision. Their impact on society and human life will be significant. These automation systems are safety-critical, complex, and powerful with a higher-level functionality. Safety and reliability to perform their tasks safely and minimize failures is one of the key challenges and becomes costly and difficult to achieve. The development of novel safety and reliability technologies dealing with theoretical aspects and for intelligent automation systems has become a hot spot in recent years. With the novel safety and reliability technologies, the intelligent automation systems can detect system failures, identify operation risks, predict unknown safety hazards and vulnerabilities, and avoid situations that pose risks to humans, property, or the automation systems themselves. Recent developments confirm that there are still areas of research to be explored within the safety-critical approaches. This Special Issue on Emerging Trends in Safety-critical Issues for Intelligent Automation Systems of IEEE Transactions on Automation Science and Engineering (TASE) aims to present recent advances in theories, methods, and applications that address safety-critical challenges in autonomous intelligent systems. The objective is to compile state-of-the-art research that contributes to the development of resilient and trustworthy automation solutions in safety-sensitive scenarios. After a thorough and rigorous peer-review process, 27 high-quality articles were selected from numerous submissions worldwide. These articles closely align with the Special Issue’s scope and can be categorized into seven key topics. - Risk assessment and model-based safety and cybersecurity analysis [A1], [A2], [A3], [A4], [A5]. - Intelligent fault detection and fault-tolerant control [A6], [A7], [A8], [A9], [A10]. - Reliability and traceability of decision-making for intelligent automation systems [A11], [A12], [A13]. - Conflict detection and resolution in intelligent automation systems [A14], [A15]. - Safety- and security-related issues including functional safety and system security [A16], [A17], [A18]. - Design, development, validation, and applications of intelligent automation systems such as UAVs, UGVs, and UUVs [A19], [A20], [A21], [A22], [A23]. - Human-robot collaboration, risk assessment of intelligent automation [A24], [A25], [A26], [A27]. We believe this collection will provide a valuable reference for academia and industry alike, and inspire further research in this rapidly evolving field. Chao Huang 0006, Qinglai Wei, Huaguang Zhang, Andrey V. Savkin, Mohammed Chadli, Hailong Huang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Parallel Control With Event-Based Adaptive Critic Implementation for Robust Optimal Tracking of Uncertain Nonlinear SystemsabstractThis paper investigates event-based robust optimal parallel tracking control for a class of uncertain nonlinear systems via adaptive dynamic programming (ADP). Analysis reveals that optimal control of the nominal system with sufficient feedback gain leads to robust optimal control. Then, optimal control is implemented online employing a critic neural network (NN), and event-triggered mechanisms (ETMs) are explored to update the weights intermittently. Among them, a unique dynamic event-triggered mechanism (DETM), which releases data in the light of an auxiliary term designed for stability verification, merits significant focus, and the comparison emphasizes its potential for better handling practical control challenges. Finally, experimental findings highlight the feasibility of the proposed robust control method while validating the characteristics and superiority of the ETMs. Note to Practitioners—As a significant issue in practical applications, event-based robust optimal control motivates this study. Distinguished from the previous efforts, this study develops the practical control difficulty into virtual space via parallel control and proportionally increases the nominal system’s optimal control rule for robust optimal control. The relaxation of both the system’s prior knowledge and the assumption related to the input dynamics allows the proposed method to be more compatible with practical systems. Moreover, ETMs are adopted to better balance control performance and resource occupation, with segmented DETM providing a fresh insight to effectively address practical control issues, i.e., releasing more data to prevent the system from being damaged by the disturbance while invoking less data to improve resource efficiency when the system is stable. Finally, the experimental comparisons validate the theoretical results. Shanshan Jiao 0001, Qinglai Wei, Jun Xiao 0005 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | A Multi-Observer Based Optimal Control Method for Nonlinear Systems Under Sensor AttacksabstractA multi-observer based optimal control method is presented for nonlinear continuous-time systems under sensor attacks. On the basis of the ideas of sensor redundancy and multi-observers, a neural network (NN) based multi-observer method is introduced to find the attack-free set of sensors in order to deal with the influences of sensor attacks. Then, the adaptive dynamic programming (ADP) method with critic NN approximation is developed by utilizing the correct state estimate. The properties of the multi-observer based ADP method are analyzed by utilizing the Lyapunov theory. Numerical analysis is conducted to show the efficiency of the presented method. Note to Practitioners—In practical engineering, the communication networks may provide new access points for cyberattacks which try to degrade the system control performance because of the physical constraints. These cyberattacks can affect the quality of the system data transmission, which brings great challenges to the effective optimal control of systems. Aiming at the above problems, this paper designs a multi-observer based optimal control method for nonlinear continuous-time systems under sensor attacks. A neural network (NN) based multi-observer method is provided to find the attack-free set of sensors. Then, the adaptive dynamic programming (ADP) method is presented to realize the optimal control of systems based on the state estimate. Numerical analysis is given to show the correctness of the presented method. Hongyang Li 0002, Qinglai Wei |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Learning-Based Parallel Control for Unknown Nonaffine Nonzero-Sum GamesabstractIn this paper, a novel nonzero-sum game (NSG) method is developed for completely unknown nonaffine nonlinear discrete-time (DT) systems, which is referred to as model-free NSG (MNSG). First, novel dynamic control laws are developed for NSGs using parallel control, namely introducing controls into feedback. Subsequently, an augmentedN-player NSG is formulated according to the originalN-player NSG to derive the dynamic control laws. Furthermore, we show that the control stabilities of the original and augmentedN-player NSGs are equivalent. In the meantime, we prove that optimal control of the augmentedN-player NSG is equivalent to near-optimal control of the originalN-player NSG, and the Nash equilibrium of the originalN-player NSG can be achieved. Then, a model-free learning scheme is developed to obtain the solution of the augmentedN-player NSG using online policy iteration, and neither using a model network to predict unknown dynamics nor off-policy reinforcement learning (RL) is needed in the scheme. Lastly, numerical analysis, including the NSG of a DT system with unknown control-nonaffine dynamics and coupled controls, confirms the correctness of our MNSG method. The associated code is available at: https://github.com/lujingweihh/Adaptive-dynamic-programming-algorithms/tree/main/model_free_nonzero_sum_games_discrete_time. Jingwei Lu, Qinglai Wei, Lefei Li |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Dynamic Decentralized Event-Triggered Tracking Control of Continuous Stirred Tank Reactor SystemsabstractWe develop a dynamic decentralized event-triggered tracking control strategy for cascade interconnected continuous stirred tank reactor (CSTR) systems subject to asymmetric input limits. Initially, we construct auxiliary augmented subsystems related to the cascade interconnected CSTR systems. Then, by introducing modified nonquadratic cost functions for the auxiliary augmented subsystems, we convert the decentralized constrained tracking control problem into an array of unconstrained optimal regulation problems. After that, with the construction of dynamic event-triggering mechanisms, we propose the event-triggered Hamilton-Jacobi-Bellman equations (ET-HJBEs) associated with the transformed optimal regulation problems. To approximately solve the ET-HJBEs, we design critic neural networks (CNNs) in the adaptive dynamic programming framework with the CNNs’ weights being updated through the gradient descent method. Furthermore, we use Lyapunov method to prove that the CNNs’ weight estimation errors and the tracking error are stable in the sense of uniform ultimate boundedness. Finally, simulation results of the three-reactor cascade interconnected CSTR systems are provided to validate the present dynamic decentralized event-triggered tracking control scheme. Xiong Yang 0001, Jianling Meng, Qinglai Wei |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Parallel Control With Adaptive Critic-Actor Learning Implementation for State and Input Time-Delayed Nonlinear Continuous-Time SystemsabstractThis study seeks to develop a constructive approach that settles the optimal control issue for nonlinear systems with known time delays. The feedback system, which depends on the state and control input, is built to identify the actual control rule utilizing the backstepping integral technique. Optimal control of the augmented system established on parallel control delivers a solution for nonlinear time-delayed systems. At the cost of the modified gain condition, the value function is specified in terms of the state and input delays, transforming the optimal control issue into a minimax task. Then, the critic-actor framework is employed to reconstruct the cost function and control rule while maintaining the persistently exciting (PE) condition so that the online optimal control algorithm is investigated. In addition, the Lyapunov proof discusses the system's stability. Ultimately, the remarkable properties become visible through experimental findings. Shanshan Jiao 0001, Qinglai Wei, Fei-Yue Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2025 | Reinforcement Learning for H∞ Optimal Control of Unknown Continuous-Time Linear SystemsabstractDesigning the optimal control for the practical systems is challenging due to the unknown system dynamics and unavoidable external disturbances. In this article, the $H_{\infty } $ optimal control problem is investigated for continuous-time linear systems with unknown dynamics. The existing reinforcement learning-based $H_{\infty } $ optimal control methods require persistence of excitation (PE) condition or data storage mechanism to guarantee the convergence of the algorithms. However, PE condition is hard to be monitored online and data storage mechanism requires to store huge amounts of past system data. In order to solve these problems, the initial excitation-based reinforcement learning algorithms are presented to learn the optimal control policy under an online-verifiable initial excitation condition. The properties of the initial excitation-based reinforcement learning algorithms are analyzed, which show that the presented algorithms converge to the optimum under the initial excitation condition. Numerical analysis is provided which demonstrates the correctness of the presented algorithms. Hongyang Li 0002, Qinglai Wei, Xiangmin Tan |
IEEE Trans. Cybern. | 2 |
| 2025 | Event-/Self-Triggered Adaptive Optimal Consensus Control for Nonlinear Multiagent System With Unknown Dynamics and DisturbancesabstractIn this article, the optimal consensus tracking control for nonlinear multiagent systems (MASs) with unknown dynamics and disturbances is investigated via adaptive dynamic programming (ADP) technology. Taking into account the disturbance as control inputs, the optimal control problem for the nonlinear MASs is reformulated as a multiplayer zero-sum differential game. In addition, a single network ADP structure is constructed to approach the optimal consensus control policies. Subsequently, an event triggering mechanism is implemented to reduce the workload of the controller and conserve computing and communication resources. Since then, in order to further streamline the intricacies of controller design, this work is extended to self-triggered cases to alleviate the need for hardware devices to continuously monitor signals. By using the Lyapunov method, the stability of the nonlinear MASs and the uniform ultimate boundedness (UUB) of the weight estimation error of the critic neural network (NN) is proved. Finally, the simulation results for an MAS consisting of a single-link robot validate the effectiveness of the proposed control method. Qinglai Wei, Hao Jiang 0037 |
IEEE Trans. Cybern. | 1 |
| 2025 | Secure Containment Control for Multi-UAV Systems by Fixed-Time Convergent Reinforcement LearningabstractThis article concerns the secure containment control problem for multiple autonomous aerial vehicles. The cyber attacker can manipulate control commands, resulting in containment failure in the position loop. Within a zero-sum graphical game framework, secure containment controllers and malicious attackers are regarded as game players, and the attack-defense process is recast as a min-max optimization problem. Acquiring optimal distributed secure control policies requires solving the game-related Hamilton-Jacobi-Isaacs (HJI) equations. Based on the critic-only neural network (NN) structure, the reinforcement learning (RL) method is employed in solving coupled HJI equations. The fixed-time convergence technique is introduced to improve the convergence rate of RL, and the experience replay mechanism is utilized to relax the persistence of excitation condition. The associated NN convergence and closed-loop stability are analyzed. In the attitude loop, the optimal feedback control law is obtained by solving Hamilton-Jacobi-Bellman equations using the fixed-time convergent RL method. The simulation example and the quadrotor experiment are given to show the effectiveness of the proposed scheme. Feisheng Yang, Zhenyu Gong, Qinglai Wei, Yifei Lei |
IEEE Trans. Cybern. | 3 |
| 2025 | Multi-Viewpoint and Multi-Evaluation With Felicitous Inductive Bias Boost Machine Abstract Reasoning AbilityabstractGreat efforts have been made to investigate AI's ability in abstract reasoning, along with the proposal of various versions of RAVEN's progressive matrices (RPM) as benchmarks. Previous studies suggest that, even after extensive training, neural networks may still struggle to make decisive decisions regarding RPM problems without sophisticated designs or additional semantic information in the form of meta-data. Through comprehensive experiments, we demonstrate that neural networks endowed with appropriate inductive biases, either intentionally designed or fortuitously matched, can efficiently solve RPM problems without the need for extra meta-data augmentation. Our work also reveals the importance of employing a multi-viewpoint with multi-evaluation approach as a key learning strategy for successful reasoning. Nevertheless, we acknowledge the unique role of metadata by demonstrating that a pre-training model supervised by meta-data leads to an RPM solver with improved performance. Codes are available in: https://github.com/QinglaiWeiCASIA/RavenSolver. Qinglai Wei, Diancheng Chen, Beiming Yuan |
IEEE Trans. Image Process. | 1 |
| 2025 | Synergetic Learning Neuro-Control for Unknown Affine Nonlinear Systems With Asymptotic Stability GuaranteesabstractFor completely unknown affine nonlinear systems, in this article, a synergetic learning algorithm (SLA) is developed to learn an optimal control. Unlike the conventional Hamilton-Jacobi-Bellman equation (HJBE) with system dynamics, a model-free HJBE (MF-HJBE) is deduced by means of off-policy reinforcement learning (RL). Specifically, the equivalence between HJBE and MF-HJBE is first bridged from the perspective of the uniqueness of the solution of the HJBE. Furthermore, it is proven that once the solution of MF-HJBE exists, its corresponding control input renders the system asymptotically stable and optimizes the cost function. To solve the MF-HJBE, the two agents composing the synergetic learning (SL) system, the critic agent and the actor agent, can evolve in real-time using only the system state data. By building an experience reply (ER)-based learning rule, it is proven that when the critic agent evolves toward the optimal cost function, the actor agent not only evolves toward the optimal control, but also guarantees the asymptotic stability of the system. Finally, simulations of the F16 aircraft system and the Van der Pol oscillator are conducted and the results support the feasibility of the developed SLA. Liao Zhu, Qinglai Wei, Ping Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Parallel Control for Nonzero-Sum Games With Completely Unknown Nonlinear Dynamics via Reinforcement LearningabstractThis article utilizes parallel control to investigate the problem of continuous-time (CT) nonzero-sum games (NZSGs) for completely unknown nonlinear systems via reinforcement learning (RL), and a parallel control-based NZSG (PNZSG) method is developed without reconstructing unknown dynamics or employing off-policy integral RL (IRL). First, novel dynamic control policies (DCPs) are developed for NZSGs by introducing controls into feedback, and an augmented system with augmented performance indices is constructed to derive the DCPs. Then, we theoretically analyze the effect of the DCPs on the control stability and performance indices, and the optimality of PNZSG is proven to be equivalent to the optimality of the original NZSGs. Subsequently, an IRL technique is employed to achieve the developed PNZSG method, and we show that no prior knowledge of the dynamics of NZSGs is needed to deploy the developed PNZSG method because of the augmented system and performance indices. Finally, numerical examples, including cooperative adaptive cruise control (CACC) of a vehicular platoon, demonstrate the correctness of the developed PNZSG method. The associated code is available at:https://github.com/lujingweihh/Adaptive-dynamic-programming-algorithms/tree/main/model_free_nonzero_sum_games. Jingwei Lu, Qinglai Wei, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Data-Driven Optimal Control for Continuous- Time Linear Nonzero-Sum GamesabstractIn this paper, the data-driven optimal control problem is studied for continuous-time linear nonzero-sum games. Two kinds of reinforcement learning algorithms, i.e., reinforcement learning algorithm with data-storage based least-square method and reinforcement learning algorithm with filter based least-square method, are presented to obtain the Nash equilibrium solution. The properties of the presented rein-forcement learning algorithms are analyzed. Simulation results show the efficiency of the presented reinforcement learning algorithms. Hongyang Li 0002, Qinglai Wei, Ruizhuo Song |
ICARCV | 2 |
| 2024 | A novel energy management method for multiple residential energy systems with energy exchange
Hongyang Li 0002, Qinglai Wei |
Neurocomputing | 2 |
| 2024 | Nearly optimal stabilization of unknown continuous-time nonlinear systems: A new parallel control approach
Jingwei Lu, Xingxia Wang, Qinglai Wei, Fei-Yue Wang 0001 |
Neurocomputing | 3 |
| 2024 | Class-incremental learning with Balanced Embedding Discrimination Maximization
Qinglai Wei, Weiqin Zhang |
Neural Networks | 1 |
| 2024 | Data-Driven Optimal Output Cluster Synchronization Control of Heterogeneous Multi-Agent SystemsabstractThis paper presents a novel data-driven optimal output cluster synchronization control method for heterogeneous multi-agent systems with disturbances based on adaptive dynamic programming. Traditional cluster synchronization control methods require the information of system matrices, which limit the application scope in the reality. In order to solve this problem, a novel data-driven optimal control method is presented, where the input-state data is utilized without the information of system matrices of state equations, to realize the output cluster synchronization of multi-agent systems. The major contributions are displayed as follows: 1) a novel data-driven optimal output cluster synchronization control method is presented which requires the input-state data of multi-agent systems; 2) a novel distributed adaptive observer is designed which can avoid the effects of negative edge weights between the clusters; 3) the output cluster synchronization control problem is transformed into the output regulation problem, and adaptive dynamic programming method is presented which can realize the disturbance rejection. First, the output cluster synchronization control problem is formulated. Next, a novel optimal output cluster synchronization control method is presented based on distributed adaptive observer and adaptive dynamic programming. Numerical experiment shows the good performance of the presented method.Note to Practitioners—Most of the existing cluster synchronization control methods require the information of system matrices. However, it is hard to obtain the accurate system models in the reality, which limits the application scope of the existing methods. On the other hand, the disturbances in practice make the effective cluster synchronization control of multi-agent systems very challenging. Aiming at the above problems, this paper designs novel data-driven optimal control laws for heterogeneous multi-agent systems with disturbances to realize the output cluster synchronization. A novel distributed adaptive observer is designed to estimate the states and system matrices of leaders. Then, the adaptive dynamic programming method is presented to obtain the optimal control law of each agent based on the estimated information. Comparative experiment is provided to show the good performance of the presented method. Hongyang Li 0002, Qinglai Wei |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Online Adaptive Dynamic Programming for Optimal Self-Learning Control of VTOL Aircraft Systems With DisturbancesabstractIn this paper, a novel online adaptive dynamic programming control algorithm is developed to solve optimal control problems for Vertical Take Off and Landing (VTOL) aircraft systems. Considering the input disturbances of the VTOL systems and adding perturbation interference term to the value function, a robust policy iteration-based ADP method will be introduced to obtain the optimal controller with an approximate optimal control approach. The convergence property is developed to ensure the value function converges to a finite neighborhood of the optimal one. The uniform ultimate boundedness (UUB) of the iterative errors will be strictly proved by Lyapunov stability theory. Finally, we will give mathematical simulation results with the developed method. Compared with sliding mode control (SMC) and linear quadratic regulator (LQR) methods, the simulation results illustrate better overall performance of the novel method. Note to Practitioners—As an effective tool to solve the optimal control problem of complex nonlinear systems, adaptive dynamic programming (ADP) can effectively overcome the computational problems caused by “curse of dimensionality”. VTOL aircraft system is a typical strongly coupled, underactuate, non-minimum phase system, and the sudden random disturbances will make the control task more challenging. Aiming at the particularity of VTOL system, a new online policy iteration-based ADP algorithm is developed to obtain optimal control with an approximate optimal control strategy in this paper. Decoupling and disturbances can be effectively avoided. convergence will be analyzed to make sure the value function converge to a finite neighborhood of the optimal value function. The UUB property of the system will be proved by the Lyapunov approach. Qinglai Wei, Zesheng Yang, Huaizhong Su, Lijian Wang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | A Novel Parallel Control Method for Optimal Consensus of Nonlinear Multiagent SystemsabstractThis work concentrates on the initial introduction of parallel control to investigate an optimal consensus control strategy for continuous-time nonlinear multiagent systems (MASs) via adaptive dynamic programming (ADP). First, the control input is integrated into the feedback system for parallel control, facilitating an augmented system's optimal consensus control with an appropriate augmented performance index function to be established, which is identical to the original system's suboptimal control with a conventional performance index. Second, the feasibility of the proposed control scheme is evaluated based on the policy iteration algorithm, and the convergence of the algorithm is demonstrated. Then, an online learning algorithm becomes available to implement the ADP-based optimal parallel consensus control protocol without prior knowledge of the system. The Lyapunov approach is employed to indicate that the signals are convergent. Ultimately, the experimental data support the theoretical results. Shanshan Jiao 0001, Qinglai Wei, Fei-Yue Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | Observer-Based Optimal Backstepping Security Control for Nonlinear Systems Using Reinforcement Learning StrategyabstractThis article considers an observer-based optimal backstepping security control for nonlinear systems using reinforcement learning (RL) strategy. The main challenge faced is the design of optimal contoller under the deception attacks. Therefore, this article introduces an improved security RL algorithm based on neural network technology under the design framework of critic-actor to resist attacks and optimize the entire system. Second, compared with some existing results, how to relax the general assumption about deception attack is also a difficult research topic. In this article, an unusual observer that uses the attacked system output is designed to estimate the real unavailable states caused by deception attacks, so that the impact of deception attacks is eliminated and the output feedback control is also achieved. By selecting the virtual controllers and the real controller as corresponding optimized controllers within the framework of the RL algorithm, the control strategy can ensure that all signals in the closed-loop system are semi-globally ultimately bounded. Finally, two simulation experiments will be run to demonstrate the effectiveness of the strategy. Qinglai Wei, Xiangmin Tan, Jun Xiao 0005, Qi Dong 0005 |
IEEE Trans. Cybern. | 1 |
| 2024 | Learning and Controlling Multiscale Dynamics in Spiking Neural Networks Using Recursive Least Square ModificationsabstractInvasive brain-computer interfaces (BCIs) have the capability to simultaneously record discrete signals across multiple scales, but how to effectively process and analyze these potentially related signals remains an open challenge. This article introduces an innovative approach that merges modern control theory with spiking neural networks (SNNs) to bridge the gap among multiscale discrete information. Specifically, the macroscopic point-to-point trajectory is formulated as an optimal control problem with fixed terminal time and state, and it is iteratively solved using the direct dynamic programming (DDP) algorithm. Additionally, SNN is utilized to simulate microscale neural activities in the premotor cortex, employing the product of the weighted adjacency matrix and the mesoscale firing rate to approximate the macroscopic trajectory. The error between actual macroscale behavior and the preceding approximation is then used to update the weighted adjacency matrix through the recursive least square (RLS) method. Analysis and simulation of various tasks, including low-dimensional point-to-point tasks, high-dimensional complex Lorenz systems, and center-out-and-back tasks, verify the feasibility and interpretability of our method in processing multiscale signals ranging from spiking neurons to motion trajectory through the integration of SNN and control theory. Qinglai Wei, Liyuan Han, Tielin Zhang |
IEEE Trans. Cybern. | 1 |
| 2024 | Robust Optimal Parallel Tracking Control Based on Adaptive Dynamic ProgrammingabstractThis article focuses on a novel robust optimal parallel tracking control method for continuous-time (CT) nonlinear systems subject to uncertainties. First, the designed virtual controller facilitates the transformation of the original nonlinear system into an affine system with an augmented state vector, which promotes the introduction of the optimal parallel tracking control problem. Then, this article generates fresh insight into counteracting the effects of uncertainty by developing a novel parallel control system that invokes the formulated virtual control law and an auxiliary variable obtained from the relationship between the solutions of the optimal control problems for the uncertain system and the nominal one. Next, critic neural networks (NNs) approximate the Hamilton-Jacobi-Bellman (HJB) equations' solution to implement the proposed robust optimal control method via adaptive dynamic programming (ADP). Finally, simulation experiments demonstrate the proposed method's remarkable effectiveness. Qinglai Wei, Shanshan Jiao 0001, Fei-Yue Wang 0001, Qi Dong 0005 |
IEEE Trans. Cybern. | 1 |
| 2024 | Isoperimetric Constraint Inference for Discrete-Time Nonlinear Systems Based on Inverse Optimal ControlabstractIn this article, the problem of inferring unknown isoperimetric constraints is considered given optimal state and control trajectories that solve the optimal control problem with isoperimetric constraints. By exploiting Pontryagin's principle, the recovery equations for unknown isoperimetric constraints are established. Under verifiable dimensionality condition and matrix rank condition, the proposed method is guaranteed to infer the unknown isoperimetric constraints exactly. Furthermore, the proposed method is extended to multiple trajectory setting. Finally, the effectiveness of the proposed method is illustrated by two simulation examples with various settings. Qinglai Wei, Tao Li 0058, Jie Zhang 0116, Hongyang Li 0002, Xin Wang 0137, Jun Xiao 0005 |
IEEE Trans. Cybern. | 1 |
| 2024 | Adaptive Dynamic Programming for Robust Event-Driven Tracking Control of Nonlinear Systems With Asymmetric Input ConstraintsabstractThis article considers the robust dynamic event-driven tracking control problem of nonlinear systems having mismatched disturbances and asymmetric input constraints. Initially, to tackle the asymmetric constraints, a novel nonquadratic value function is constructed for the original system. This makes the asymmetrically constrained tracking control problem transformed into an unconstrained optimal regulation problem. Then, a dynamic event-driven mechanism is proposed. Meanwhile, the event-driven Hamilton-Jacobi-Bellman equation (ED-HJBE) is developed for the optimal regulation problem in order to acquire the optimal control with distinctly decreased computational burden. To solve the ED-HJBE, a single critic neural network (CNN) is designed in the adaptive dynamic programming framework. Meanwhile, the gradient descent method is employed to update the CNN's weights. After that, both the weight estimation error and the tracking error are proved to be uniformly ultimately bounded via Lyapunov's direct method. Finally, simulations of the spring-mass-damper system and the pendulum plant are separately utilized to validate the established theoretical claims. Xiong Yang 0001, Qinglai Wei |
IEEE Trans. Cybern. | 2 |
| 2024 | Optimal Spin Polarization Control for the Spin-Exchange Relaxation-Free System Using Adaptive Dynamic ProgrammingabstractThis work is the first to solve the 3-D spin polarization control (3DSPC) problem of atomic ensembles, which controls the spin polarization to achieve arbitrary states with the cooperation of multiphysics fields. First, a novel adaptive dynamic programming (ADP) structure is proposed based on the developed multicritic multiaction neural network (MCMANN) structure with nonquadratic performance functions, as a way to solve the multiplayer nonzero-sum game (MP-NZSG) problem in 3DSPC under the constraints of asymmetric saturation inputs. Then, we utilize the MCMANNs to implement the multicritic multiaction ADP (MCMA-ADP) algorithm, whose convergence is proven by the compression mapping principle. Finally, the MCMA-ADP is deployed in the spin-exchange relaxation-free (SERF) system to provide a set of control laws in 3DSPC that fully exploits the multiphysics fields to achieve arbitrary spin polarization states. Numerical simulations support the theoretical results. Zhuo Wang 0003, Sixun Liu, Tao Li 0058, Feng Li 0066, Bodong Qin, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Constrained-Cost Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear SystemsabstractFor discrete-time nonlinear systems, this research is concerned with optimal control problems (OCPs) with constrained cost, and a novel value iteration with constrained cost (VICC) method is developed to solve the optimal control law with the constrained cost functions. The VICC method is initialized through a value function constructed by a feasible control law. It is proven that the iterative value function is nonincreasing and converges to the solution of the Bellman equation with constrained cost. The feasibility of the iterative control law is proven. The method to find the initial feasible control law is given. Implementation using neural networks (NNs) is introduced, and the convergence is proven by considering the approximation error. Finally, the property of the present VICC method is shown by two simulation examples. Qinglai Wei, Tao Li 0058 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | VGN: Value Decomposition With Graph Attention Networks for Multiagent Reinforcement LearningabstractAlthough value decomposition networks and the follow on value-based studies factorizes the joint reward function to individual reward functions for a kind of cooperative multiagent reinforcement problem, in which each agent has its local observation and shares a joint reward signal, most of the previous efforts, however, ignored the graphical information between agents. In this article, a new value decomposition with graph attention network (VGN) method is developed to solve the value functions by introducing the dynamical relationships between agents. It is pointed out that the decomposition factor of an agent in our approach can be influenced by the reward signals of all the related agents and two graphical neural network-based algorithms (VGN-Linear and VGN-Nonlinear) are designed to solve the value functions of each agent. It can be proved theoretically that the present methods satisfy the factorizable condition in the centralized training process. The performance of the present methods is evaluated on the StarCraft Multiagent Challenge (SMAC) benchmark. Experiment results show that our method outperforms the state-of-the-art value-based multiagent reinforcement algorithms, especially when the tasks are with very hard level and challenging for existing methods. Qinglai Wei, Yugu Li, Jie Zhang 0116, Fei-Yue Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | A Self-Attention-Based Deep Reinforcement Learning Approach for AGV Dispatching SystemsabstractThe automated guided vehicle (AGV) dispatching problem is to develop a rule to assign transportation tasks to certain vehicles. This article proposes a new deep reinforcement learning approach with a self-attention mechanism to dynamically dispatch the tasks to AGV. The AGV dispatching system is modeled as a less complicated Markov decision process (MDP) using vehicle-initiated rules to dispatch a workcenter to an idle AGV. In order to deal with the highly dynamical environment, the self-attention mechanism is introduced to calculate the importance of different information. The invalid action masking technique is performed to alleviate false actions. A multimodal structure is employed to mix the features of various sources. Comparative experiments are performed to show the effectiveness of the proposed method. The properties of the learned policies are also investigated under different environment settings. It is discovered that the policies explore and learn the properties of different systems, and also smooth the traffic congestion. Under certain environment settings, the policy converges to a heuristic rule that assigns the idle AGV to the workcenter with the shortest queue length, which shows the adaptiveness of the proposed method. Qinglai Wei, Yutian Yan, Jie Zhang 0116, Jun Xiao 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Dynamic Event-Sampled Control of Interconnected Nonlinear Systems Using Reinforcement LearningabstractWe develop a decentralized dynamic event-based control strategy for nonlinear systems subject to matched interconnections. To begin with, we introduce a dynamic event-based sampling mechanism, which relies on the system's states and the variables generated by time-based differential equations. Then, we prove that the decentralized event-based controller for the whole system is composed of all the optimal event-based control policies of nominal subsystems. To derive these optimal event-based control policies, we design a critic-only architecture to solve the related event-based Hamilton-Jacobi-Bellman equations in the reinforcement learning framework. The implementation of such an architecture uses only critic neural networks (NNs) with their weight vectors being updated through the gradient descent method together with concurrent learning. After that, we demonstrate that the asymptotic stability of closed-loop nominal subsystems and the uniformly ultimate boundedness stability of critic NNs' weight estimation errors are guaranteed by using Lyapunov's approach. Finally, we provide simulations of a matched nonlinear-interconnected plant to validate the present theoretical claims. Xiong Yang 0001, Mengmeng Xu 0004, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multistep Look-Ahead Policy Iteration for Optimal Control of Discrete-Time Nonlinear Systems With Isoperimetric ConstraintsabstractIn this article, a novel multistep look-ahead policy iteration with isoperimetric constraints (MLPIIC) method is developed to solve infinite horizon optimal control problems (OCPs) with isoperimetric constraints for discrete-time nonlinear systems. In order to overcome the difficulty that Bellman’s principle of optimality does not hold directly in OCPs with isoperimetric constraints, a method to approximate OCPs with isoperimetric constraints by OCPs with new constraints is developed. For the MLPIIC method initialized with an admissible control law, the convergence and optimality of the iterative value function and the feasibility of the iterative control law are proven. Utilizing the function approximator, the implementation of the MLPIIC method is described. Finally, simulation results are provided. Tao Li 0058, Qinglai Wei, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Initial Excitation-Based Optimal Control for Continuous-Time Linear Nonzero-Sum GamesabstractIn this article, the initial excitation-based optimal control methods are presented for continuous-time linear nonzero-sum games. The traditional reinforcement learning-based optimal control methods for continuous-time linear nonzero-sum games require the persistent excitation condition or data storage to guarantee the convergence of the algorithms. To relax the above conditions, the initial excitation-based policy iteration and value iteration algorithms are presented to obtain the Nash equilibrium solution under an online-verifiable initial excitation condition. The properties of the initial excitation-based policy iteration and value iteration algorithms are analyzed. Simulation examples are provided to show the efficiency of the presented methods. Hongyang Li 0002, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Online Off-Policy Reinforcement Learning for Optimal Control of Unknown Nonlinear Systems Using Neural NetworksabstractIn this article, a real-time online off-policy reinforcement learning (RL) method is developed for the optimal control problem of unknown continuous-time nonlinear systems. First, by applying the temporal difference technique to the iterative procedure of off-policy RL, the iterative value function and the iterative policy input can be learned in real-time online. It is proven that the fitting error of neural network (NN) weights is exponentially convergent in each iteration. Second, a model-free Hamilton–Jacobi–Bellman equation (MF-HJBE) is deduced by taking the limit of the iterative procedure of off-policy RL. In this manner, it not only eliminates system dynamics in the classical HJBE, but also vanishes the iteration index. By applying temporal difference to the MF-HJBE, a real-time online tuning rule is designed to learn the optimal value function and the optimal policy input. It is proven that the fitting error of NN weights caused by the real-time online tuning rule is exponentially convergent. Note that the two online tuning rules, the iterative one and the real-time one, use only current and previous state data extracted from system trajectories. Meanwhile, it is proven using the Lyapunov’s direct method that the system solution is uniformly ultimately bounded. Finally, simulation results demonstrate the validity of the proffered method. Liao Zhu, Qinglai Wei, Ping Guo 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | Event-triggered near-optimal tracking control based on adaptive dynamic programming for discrete-time systems
Joonhyup Lee, Qinglai Wei, Anting Zhang |
Neurocomputing | 3 |
| 2023 | Raven solver: From perception to reasoning
Qinglai Wei, Diancheng Chen, Beiming Yuan, Peijun Ye 0001 |
Inf. Sci. | 1 |
| 2023 | Synergetic learning for unknown nonlinear H∞ control using neural networks
Liao Zhu, Ping Guo 0002, Qinglai Wei |
Neural Networks | 3 |
| 2023 | Event-Triggered Near-Optimal Control for Unknown Discrete-Time Nonlinear Systems Using Parallel ControlabstractThis article uses parallel control to investigate the problem of event-triggered near-optimal control (ETNOC) for unknown discrete-time (DT) nonlinear systems. First, to achieve parallel control, an augmented nonlinear system (ANS) with an augmented performance index (API) is proposed to introduce the control input into the feedback system. The control stability relationship between the ANS and the original system is analyzed, and it is shown that, by choosing a proper API, optimal control of the ANS with the API can be seen as near-optimal control of the original system with the original performance index (OPI). Second, based on parallel control, a novel event-triggered scheme is proposed, and then a novel ETNOC method is developed using the time-triggered optimal value function of the ANS with the API. The control stability is proved, and an upper bound, which is related to the design parameter, is provided for the actual performance index in advance. Then, to implement the developed ETNOC method for unknown DT nonlinear systems, a novel online learning algorithm is developed without reconstructing unknown systems, and neural network (NN) and adaptive dynamic programming (ADP) techniques are employed in the developed algorithm. The convergence of the signals in the closed-loop system (CLS) is shown using the Lyapunov approach, and the assumption of boundedness of input dynamics is not required. Finally, two simulations justify the theoretical conjectures. Jingwei Lu, Qinglai Wei, Tianmin Zhou, Fei-Yue Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2023 | A Novel Parallel Control Method for Continuous-Time Linear Output Regulation With DisturbancesabstractIn this article, a novel linear parallel control method is developed for output regulation problems with disturbances. The traditional feedback regulators are passive regulation methods. In order to solve this problem, the parallel controllers are presented, where the time variation of the control is constructed instead of the control value itself, to stabilize the output of the system. The main contributions of the developed method include two aspects: 1) a novel parallel regulator structure is presented, which can provide greater flexibility comparing with traditional methods and 2) the necessary and sufficient conditions for the existence of parallel regulators are analyzed. First, the structure of the linear parallel regulators is provided. Next, considering the situations that the full system information is obtained and the information of error is obtained, respectively, the properties of the parallel regulators are analyzed, and the regulator designs for the two situations are provided. Finally, numerical examples are provided to verify the correctness of the present method. Qinglai Wei, Hongyang Li 0002, Fei-Yue Wang 0001 |
IEEE Trans. Cybern. | 1 |
| 2023 | A Data-Driven Iterative Learning Approach for Optimizing the Train Control StrategyabstractThe energy-efficient train control (EETC) problem is investigated in this article. And a soft actor-critic (SAC)-based method is proposed to optimize the train driving strategy. First, EETC problem is converted to the inverse problem, i.e., minimizing the trip time of the journey with constant energy consumption. Based on the conversion, the EETC problem is reformulated as a finite Markov decision process, which can be solved by deep reinforcement learning algorithms. Second, an optimization method based on the SAC method is designed to calculate the optimal driving strategy of the train with introducing the reservoir sampling method. Finally, some case studies are conducted to verify the effectiveness and performance of the proposed method. Simulation results demonstrate that a good energy-saving performance can be achieved. In single interval, the SAC-based method can reduce about 1.65% of the energy consumption compared with numerical method. And the energy consumption reduction can be extended to be 6.49% when the proposed approach is applied in multiple intervals. Shuai Su, Qingyang Zhu, Junqing Liu, Tao Tang 0004, Qinglai Wei, Yuan Cao 0002 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | A Novel Data-Based Fault-Tolerant Control Method for Multicontroller Linear Systems via Distributed Policy IterationabstractThis article presents a novel data-based fault-tolerant control method for multicontroller linear systems via distributed policy iteration. The traditional fault-tolerant control methods based on policy iteration may cause a huge-computational burden under the situation of high-dimension control laws. In order to solve this problem, a novel distributed policy iteration method is presented, where only one iterative control law is updated in each iteration, to realize the fault-tolerant control of multicontroller linear systems. The main contributions can be highlighted as follows: 1) a novel data-based distributed policy iteration method is presented to reduce the computational burden; 2) the fault-tolerant control method is presented via designing fault compensators; and 3) the developed data-based method only requires the partial system information. First, the model-based fault-tolerant control via distributed policy iteration and fault compensation is provided. Based on the model-based method, a data-based fault-tolerant control method is presented. Finally, numerical experiments are given to show the performance of the presented method. Qinglai Wei, Hongyang Li 0002, Tao Li 0058, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Continuous-Time Stochastic Policy Iteration of Adaptive Dynamic ProgrammingabstractIn this article, we study the optimal control problem of continuous-time (CT) time-invariant nonlinear systems with stochastic nonlinear disturbances. A new stochastic adaptive dynamic programming (ADP) method is developed to solve the Hamilton–Jacobi–Bellman equation (HJBE). Under the conditional expectation, the value function and the control law are successively approximated simultaneously. The asymptotic stability of the closed-loop stochastic system in probability is analyzed by the stochastic Lyapunov direct method, and the convergence of the developed ADP method is given. Finally, four simulations illustrate the effectiveness of the developed method. Qinglai Wei, Tianmin Zhou, Jingwei Lu, Yu Liu 0078, Shuai Su, Jun Xiao 0005 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Adaptive Dynamic Programming for Nonlinear-Constrained H∞ ControlabstractThis article considers the$H_{\infty }$control problem of nonlinear systems having unavailable dynamics and asymmetric saturating actuators. Initially, such an$H_{\infty }$control problem is converted into the zero-sum game with a nonquadratic cost function being introduced. Then, in order to solve the Hamilton–Jacobi–Isaacs equation arising in the zero-sum game, a simultaneous policy iteration (SPI) algorithm is developed under the adaptive dynamic programming framework. Meanwhile, it is proved that the convergence of the SPI algorithm in essence amounts to the convergence of the sequential PI algorithm. To implement the SPI algorithm, the critic, the actor, and the perturbation neural networks (NNs) are, respectively, constructed to estimate the cost function, the control policy, and the perturbation. The three NNs’ weights are simultaneously determined by using the least-squares method together with the Monte Carlo integration technique. A remarkable characteristic of such an SPI algorithm is that arbitrary control policies and perturbations are applicable in the learning process. This makes system’s information be able to be replaced by the data collected along system’s trajectories in advance. More importantly, the persistence of the excitation condition is not required. Finally, simulations of two nonlinear examples are given to validate the present SPI algorithm. Xiong Yang 0001, Mengmeng Xu 0004, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Approximate Dynamic Programming for Event-Driven H∞ Constrained ControlabstractWe study the dynamic event-driven H∞ constrained control problem through approximate dynamic programming (ADP). Differing from the existing literature considering systems with either symmetric constraints or asymmetric constraints, we consider the two different constraints simultaneously. Initially, by constructing a generalized nonquadratic value function, we transform the H∞ constrained control problem into an unconstrained two-player zero-sum game. Then, we present an event-driven Hamilton–Jacobi–Isaacs equation (ED-HJIE) corresponding to the zero-sum game for lowering down the computational load. To solve the ED-HJIE, we propose a dynamic triggering mechanism together with a sole critic neural network (CNN) being built under the ADP framework. The CNN’s weights are tuned via the gradient descent approach. After that, we prove uniform ultimate boundedness of the closed-loop system and the CNN’s weight estimation error via Lyapunov’s method. Finally, we separately use an F16 aircraft plant and an inverted pendulum system to validate the present theoretical claims. Xiong Yang 0001, Mengmeng Xu 0004, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Data-driven adaptive-critic optimal output regulation towards water level control of boiler-turbine systems
Qinglai Wei, Xin Wang 0137, Yu Liu 0078, Gang Xiong 0001 |
Expert Syst. Appl. | 1 |
| 2022 | Monte Carlo-based reinforcement learning control for unmanned aerial vehicle systems
Qinglai Wei, Zesheng Yang, Huaizhong Su, Lijian Wang |
Neurocomputing | 1 |
| 2022 | A New Neuro-Optimal Nonlinear Tracking Control Method via Integral Reinforcement Learning with Applications to Nuclear Systems
Weifeng Zhong, Mengxuan Wang, Qinglai Wei, Jingwei Lu |
Neurocomputing | 3 |
| 2022 | Event-triggered optimal control for discrete-time multi-player non-zero-sum games using parallel control
Jingwei Lu, Qinglai Wei, Tianmin Zhou, Fei-Yue Wang 0001 |
Inf. Sci. | 2 |
| 2022 | Optimal synchronization control for multi-agent systems with input saturation: a nonzero-sum gameabstractThis paper presents a novel optimal synchronization control method for multi-agent systems with input saturation. The multi-agent game theory is introduced to transform the optimal synchronization control problem into a multi-agent nonzero-sum game. Then, the Nash equilibrium can be achieved by solving the coupled Hamilton—Jacobi—Bellman (HJB) equations with nonquadratic input energy terms. A novel off-policy reinforcement learning method is presented to obtain the Nash equilibrium solution without the system models, and the critic neural networks (NNs) and actor NNs are introduced to implement the presented method. Theoretical analysis is provided, which shows that the iterative control laws converge to the Nash equilibrium. Simulation results show the good performance of the presented method. Hongyang Li 0002, Qinglai Wei |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | Institutionalized and systematized gaming for multi-agent systems
Fei-Yue Wang 0001, Qi Dong 0005, Qinglai Wei |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | Parallel cognition: hybrid intelligence for human-machine interaction and managementabstractAs an interdisciplinary research approach, traditional cognitive science adopts mainly the experiment, induction, modeling, and validation paradigm. Such models are sometimes not applicable in cyber-physical-social-systems (CPSSs), where the large number of human users involves severe heterogeneity and dynamics. To reduce the decision-making conflicts between people and machines in human-centered systems, we propose a new research paradigm called parallel cognition that uses the system of intelligent techniques to investigate cognitive activities and functionals in three stages: descriptive cognition based on artificial cognitive systems (ACSs), predictive cognition with computational deliberation experiments, and prescriptive cognition via parallel behavioral prescription. To make iteration of these stages constantly on-line, a hybrid learning method based on both a psychological model and user behavioral data is further proposed to adaptively learn an individual’s cognitive knowledge. Preliminary experiments on two representative scenarios, urban travel behavioral prescription and cognitive visual reasoning, indicate that our parallel cognition learning is effective and feasible for human behavioral prescription, and can thus facilitate human-machine cooperation in both complex engineering and social systems. Peijun Ye 0001, Xiao Wang 0002, Wenbo Zheng 0001, Qinglai Wei, Fei-Yue Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | Event-Triggered Near-Optimal Control of Discrete-Time Constrained Nonlinear Systems With Application to a Boiler-Turbine SystemabstractThis article presents a novel event-triggered near-optimal control (ETNOC) method for discrete-time (DT) constrained nonlinear systems. First, the tracking error system is constructed to convert the tracking control problem to the regulation problem. By introducing the tracking error system, the asymmetric control constraints design for the original constrained system can be converted to the symmetric control constraints design for the tracking error system. Second, a novel triggering condition is developed using the time-triggered optimal value function and control law. It is proven that the closed-loop system (CLS) is asymptotically stable under the developed ETNOC method, and there exists a predetermined upper bound for the real performance index. Then, to implement the developed ETNOC method, a parallel control approach with neural networks (NNs) and adaptive dynamic programming techniques is proposed to predict the next state of the system and obtain the optimal value function and control law. The stability analysis of the CLS is provided in the consideration of the estimation errors of the NN weights and state. Finally, the effectiveness of the developed ETNOC method is validated by an application to a boiler-turbine system. Qinglai Wei, Jingwei Lu, Tianmin Zhou, Xiang Cheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Spiking Adaptive Dynamic Programming Based on Poisson Process for Discrete-Time Nonlinear SystemsabstractIn this article, a new iterative spiking adaptive dynamic programming (SADP) method based on the Poisson process is developed to solve optimal impulsive control problems. For a fixed time interval, combining the Poisson process and the maximum likelihood estimation (MLE), the three-tuple of state, spiking interval, and probability of Poisson distribution can be computed, and then, the iterative value functions and iterative control laws can be obtained. A property analysis method is developed to show that the value functions converge to optimal performance index function as the iterative index increases from zero to infinity. Finally, two simulation examples are given to verify the effectiveness of the developed algorithm. Qinglai Wei, Liyuan Han, Tielin Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Model-Free Adaptive Optimal Control for Unknown Nonlinear Multiplayer Nonzero-Sum GameabstractIn this article, an online adaptive optimal control algorithm based on adaptive dynamic programming is developed to solve the multiplayer nonzero-sum game (MP-NZSG) for discrete-time unknown nonlinear systems. First, a model-free coupled globalized dual-heuristic dynamic programming (GDHP) structure is designed to solve the MP-NZSG problem, in which there is no model network or identifier. Second, in order to relax the requirement of systems dynamics, an online adaptive learning algorithm is developed to solve the Hamilton-Jacobi equation using the system states of two adjacent time steps. Third, a series of critic networks and action networks are used to approximate value functions and optimal policies for all players. All the neural network (NN) weights are updated online based on real-time system states. Fourth, the uniformly ultimate boundedness analysis of the NN approximation errors is proved based on the Lyapunov approach. Finally, simulation results are given to demonstrate the effectiveness of the developed scheme. Qinglai Wei, Liao Zhu, Ruizhuo Song, Pinjia Zhang, Derong Liu 0001, Jun Xiao 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Decentralized Event-Driven Constrained Control Using Adaptive Critic DesignsabstractWe study the decentralized event-driven control problem of nonlinear dynamical systems with mismatched interconnections and asymmetric input constraints. To begin with, by introducing a discounted cost function for each auxiliary subsystem, we transform the decentralized event-driven constrained control problem into a group of nonlinear$H_{2}$-constrained optimal control problems. Then, we develop the event-driven Hamilton–Jacobi–Bellman equations (ED-HJBEs), which arise in the nonlinear$H_{2}$-constrained optimal control problems. Meanwhile, we demonstrate that all the solutions of the ED-HJBEs together keep the overall system stable in the sense of uniform ultimate boundedness (UUB). To solve the ED-HJBEs, we build a critic-only architecture under the framework of adaptive critic designs. The architecture only employs critic neural networks and updates their weight vectors via the gradient descent method. After that, based on the Lyapunov approach, we prove that the UUB stability of all signals in the closed-loop auxiliary subsystems is assured. Finally, simulations of an illustrated nonlinear interconnected plant are provided to validate the present designs. Xiong Yang 0001, Yuanheng Zhu, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Event-Triggered Optimal Parallel Tracking Control for Discrete-Time Nonlinear SystemsabstractA novel event-triggered optimal tracking control (ETOTC) method is developed for discrete-time nonlinear systems in this study. For the time-invariant desired trajectory, we prove that the tracking error is asymptotically stable, and an upper bound of the real performance index can be predetermined by a design parameter. For the time-varying desired trajectory, the developed triggering condition reduces communication costs by relaxing the restriction of the asymptotic stability of the closed-loop system, and we prove that the tracking error is uniformly ultimately bounded (UUB). The developed ETOTC method entails obtaining the next state of the real system. Therefore, a parallel control approach is proposed to predict the next state by constructing a parallel system for the real system. Neural networks (NNs) and adaptive dynamic programming (ADP) techniques are utilized in the parallel control approach. Moreover, the stability analysis of the closed-loop system is shown, and the tracking error and NN weight estimation errors are proved to be UUB using the Lyapunov approach. Finally, we validate the developed ETOTC method through two simulations. Jingwei Lu, Qinglai Wei, Tianmin Zhou, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Discrete-Time Self-Learning Parallel ControlabstractIn this article, a new self-learning parallel control method, which is based on adaptive dynamic programming (ADP) technique, is developed for solving the optimal control problem of discrete- time time-varying nonlinear systems. It aims to obtain an approximate optimal control law sequence and simultaneously guarantees the convergence of the value function. Establishing the time-varying artificial system by neural networks in a certain time-horizon, a control-sequence-improvement ADP algorithm is developed to obtain the control law sequence. For the first time, the criteria of the parallel execution are presented, such that the value function is proven to converge to a finite neighborhood of the optimal performance index function. Finally, numerical results and analysis are presented to demonstrate the effectiveness of the parallel control method. Qinglai Wei, Jingwei Lu, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | Topology Prediction and Structural Controllability Analysis of Complex Networks Without Connection InformationabstractIn this article, we consider complex networks without connection information. The absence of global structure induces a great obstacle in the structural controllability analysis of these networks. Thus, a topology predicting method based on the connection probability matrix is proposed to provide a global structure for structural controllability analysis in this article. Furthermore, the modified principles of predicting global network topologies are established to acquire a more accurate global connection relationship. Eventually, the drive node set of these networks is determined by predicting global topologies. The accuracy of the proposed topology predicting method is verified by numerical simulations in the context of artificial networks and real networks. The results reveal that the global topology and structural controllability of complex networks with large scale and high edge density could be accurately predicted by utilizing the proposed method. Dongsheng Yang 0001, Qinglai Wei, Huaguang Zhang, Ting Li 0021 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Adaptive Critics for Decentralized Stabilization of Constrained-Input Nonlinear Interconnected SystemsabstractThis article considers the decentralized stabilization problem of continuous-time nonlinear systems subject to unmatched interconnections and asymmetric input constraints. Initially, with the nonquadratic value functions being introduced to constrained auxiliary subsystems, the decentralized stabilization problem is converted into an array of nonlinear optimal control problems. It is proved that all solutions of these nonlinear optimal control problems together assure asymptotic stability of the entire system. Then, in the framework of adaptive critics, the critic-only architecture is built to solve the Hamilton–Jacobi–Bellman equations associated with these solutions. The critic-only architecture is implemented via critic neural networks (NNs) with their weight vectors being tuned through an improved gradient descent method. A remarkable feature of the present gradient descent approach is that it simultaneously utilizes previously stored and instantaneous state data, which makes the persistence of excitation conditions relaxed. After that, with Lyapunov’s techniques being employed, asymptotic stability of the closed-loop auxiliary subsystems and uniform ultimate boundedness of the critic NNs’ weight estimation errors are demonstrated. Finally, simulations of an unmatched interconnected nonlinear plant are provided to validate the present decentralized control method. Xiong Yang 0001, Yingjiang Zhou, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Policy Iteration Algorithm for Constrained Cost Optimal Control of Discrete-Time Nonlinear SystemabstractIn this paper, optimal control problems with constraints on summation of auxiliary utility function are called constrained cost optimal control problems and a constrained cost policy iteration adaptive dynamic programming (ADP) algorithm is developed to solve constrained cost optimal control problems for discrete-time nonlinear systems. A convergence analysis is developed to guarantee that the iterative value functions nonin-creasingly convergent to the approximate optimal value function. It is also proven that any of the iterative control policy is feasible and can stabilize the nonlinear systems. Finally, a simulation example is given to illustrate the performance of the developed constrained cost policy iteration algorithm. Tao Li 0058, Qinglai Wei, Hongyang Li 0002, Ruizhuo Song |
IJCNN | 2 |
| 2021 | Optimal Tracking Control of the Boiler-turbine System Based on Adaptive Dynamic ProgrammingabstractTo guarantee the efficient performance of the power plant, an adaptive tacking controller for the nonlinear boiler-turbine system based on offline policy iteration adaptive dynamic prorgamming (ADP) method is proposed in this paper. The optimal tracking controller is obtained through offline learning, which can maintain the characteristics of load changes in drum boiler-turbine type power plants. To implement the proposed method, neural networks (NNs) are used to construct the cost function and approximate optimal solution is achived. Then convergence of the method is analyzed. Simulation studies on the typical boiler-turbine system demonstrate that the proposed control strategy can achieve a satisfactory performance during a short period. Aiguo Gao, Qinglai Wei |
IJCNN | 3 |
| 2021 | A partial policy iteration ADP algorithm for nonlinear neuro-optimal control with discounted total reward
Mingming Liang, Qinglai Wei |
Neurocomputing | 2 |
| 2021 | Discrete-Time Non-Zero-Sum Games With Completely Unknown DynamicsabstractIn this article, off-policy reinforcement learning (RL) algorithm is established to solve the discrete-time N -player nonzero-sum (NZS) games with completely unknown dynamics. The N -coupled generalized algebraic Riccati equations (GARE) are derived, and then policy iteration (PI) algorithm is used to obtain the N -tuple of iterative control and iterative value function. As the system dynamics is necessary in PI algorithm, off-policy RL method is developed for discrete-time N -player NZS games. The off-policy N -coupled Hamilton-Jacobi (HJ) equation is derived based on quadratic value functions. According to the Kronecker product, the N -coupled HJ equation is decomposed into unknown parameter part and the system operation data part, which makes the N -coupled HJ equation solved independent of system dynamics. The least square is used to calculate the iterative value function and N -tuple of iterative control. The existence of Nash equilibrium is proved. The result of the proposed method for discrete-time unknown dynamics NZS games is indicated by the simulation examples. Ruizhuo Song, Qinglai Wei, Huaguang Zhang, Frank L. Lewis |
IEEE Trans. Cybern. | 2 |
| 2021 | Continuous-Time Distributed Policy Iteration for Multicontroller Nonlinear SystemsabstractIn this article, a novel distributed policy iteration algorithm is established for infinite horizon optimal control problems of continuous-time nonlinear systems. In each iteration of the developed distributed policy iteration algorithm, only one controller's control law is updated and the other controllers' control laws remain unchanged. The main contribution of the present algorithm is to improve the iterative control law one by one, instead of updating all the control laws in each iteration of the traditional policy iteration algorithms, which effectively releases the computational burden in each iteration. The properties of distributed policy iteration algorithm for continuous-time nonlinear systems are analyzed. The admissibility of the present methods has also been analyzed. Monotonicity, convergence, and optimality have been discussed, which show that the iterative value function is nonincreasingly convergent to the solution of the Hamilton-Jacobi-Bellman equation. Finally, numerical simulations are conducted to illustrate the effectiveness of the proposed method. Qinglai Wei, Hongyang Li 0002, Xiong Yang 0001, Haibo He |
IEEE Trans. Cybern. | 1 |
| 2021 | Generalized Actor-Critic Learning Optimal Control in Smart Home Energy ManagementabstractThis article is concerned with a new generalized actor-critic learning (GACL) optimal control method. It aims at the optimal energy control and management for smart home systems, which is expected to minimize the consumption cost for home users. In the present GACL optimal control method, it is the first time that three iteration processes, which are global iteration, local iteration, and interior iteration, respectively, are established to obtain the optimal energy control law. The main contribution of the developed method is to establish a common iteration structure for both value and policy iterations in adaptive dynamic programming based on a control law sequence in each iteration for periodic time-varying systems, instead of a single control law, and simultaneously accelerates the convergence rate. The monotonicity, convergence, and optimality of the iterative value function for the GACL optimal control method are proven. Finally, numerical results and comparisons are displayed to show the superiority of the developed method. Qinglai Wei, Zehua Liao |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Adaptive Critic Designs for Optimal Event-Driven Control of a CSTR SystemabstractThis article presents an optimal event-driven control scheme for a continuous stirred tank reactor (CSTR) system. The CSTR system differs from most of studied plants in that its equilibrium point is nonzero. In order to obtain the optimal event-driven control without coordinate transformations, we first introduce a discounted cost for such a system. Then, under the framework of adaptive critic designs, we use a single critic network to solve the Hamilton-Jacobi-Bellman equation related to the discounted-cost optimal event-driven control problem. To update the critic network weights, we employ the gradient descent method together with the experience replay technique. An advantage of the experience replay technique is that it brings about an easy-checked persistency of excitation-like condition. The stability analysis of all the signals in the closed-loop system is conducted via the Lyapunov approach. Finally, we validate the present optimal event-driven control strategy through simulations of the CSTR system. Xiong Yang 0001, Qinglai Wei |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Adaptive Critic Learning for Constrained Optimal Event-Triggered Control With Discounted CostabstractThis article studies an optimal event-triggered control (ETC) problem of nonlinear continuous-time systems subject to asymmetric control constraints. The present nonlinear plant differs from many studied systems in that its equilibrium point is nonzero. First, we introduce a discounted cost for such a system in order to obtain the optimal ETC without making coordinate transformations. Then, we present an event-triggered Hamilton-Jacobi-Bellman equation (ET-HJBE) arising in the discounted-cost constrained optimal ETC problem. After that, we propose an event-triggering condition guaranteeing a positive lower bound for the minimal intersample time. To solve the ET-HJBE, we construct a critic network under the framework of adaptive critic learning. The critic network weight vector is tuned through a modified gradient descent method, which simultaneously uses historical and instantaneous state data. By employing the Lyapunov method, we prove that the uniform ultimate boundedness of all signals in the closed-loop system is guaranteed. Finally, we provide simulations of a pendulum system and an oscillator system to validate the obtained optimal ETC strategy. Xiong Yang 0001, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Adaptive Dynamic Programming for Control: A Survey and Recent AdvancesabstractThis article reviews the recent development of adaptive dynamic programming (ADP) with applications in control. First, its applications in optimal regulation are introduced, and some skilled and efficient algorithms are presented. Next, the use of ADP to solve game problems, mainly nonzero-sum game problems, is elaborated. It is followed by applications in large-scale systems. Note that although the functions presented in this article are based on continuous-time systems, various applications of ADP in discrete-time systems are also analyzed. Moreover, in each section, not only some existing techniques are discussed, but also possible directions for future work are pointed out. Finally, some overall prospects for the future are given, followed by conclusions of this article. Through a comprehensive and complete investigation of its applications in many existing fields, this article fully demonstrates that the ADP intelligent control method is promising in today's artificial intelligence era. Furthermore, it also plays a significant role in promoting economic and social development. Derong Liu 0001, Shan Xue 0004, Bo Zhao 0015, Biao Luo 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | Deep learning neural networks: Methods, systems, and applications
Qinglai Wei, Nikola K. Kasabov, Marios M. Polycarpou, Zhigang Zeng |
Neurocomputing | 1 |
| 2020 | Nash Q-learning based equilibrium transfer for integrated energy management game with We-Energy
Lingxiao Yang, Qiuye Sun, Dazhong Ma, Qinglai Wei |
Neurocomputing | 4 |
| 2020 | Event-triggered adaptive dynamic programming for discrete-time multi-player games
Qinglai Wei, Derong Liu 0001 |
Inf. Sci. | 2 |
| 2020 | Continuous-Time Time-Varying Policy IterationabstractA novel policy iteration algorithm, called the continuous-time time-varying (CTTV) policy iteration algorithm, is presented in this paper to obtain the optimal control laws for infinite horizon CTTV nonlinear systems. The adaptive dynamic programming (ADP) technique is utilized to obtain the iterative control laws for the optimization of the performance index function. The properties of the CTTV policy iteration algorithm are analyzed. Monotonicity, convergence, and optimality of the iterative value function have been analyzed, and the iterative value function can be proven to monotonically converge to the optimal solution of the Hamilton-Jacobi-Bellman (HJB) equation. Furthermore, the iterative control law is guaranteed to be admissible to stabilize the nonlinear systems. In the implementation of the presented CTTV policy algorithm, the approximate iterative control laws and iterative value function are obtained by neural networks. Finally, the numerical results are given to verify the effectiveness of the presented method. Qinglai Wei, Zehua Liao, Zhanyu Yang, Benkai Li, Derong Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Discrete-Time Impulsive Adaptive Dynamic ProgrammingabstractIn this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal impulsive control problems for infinite horizon discrete-time nonlinear systems. Considering the constraint of the impulsive interval, in each iteration, the iterative impulsive value function under each possible impulsive interval is obtained, and then the iterative value function and iterative control law are achieved. A new convergence analysis method is developed which proves an iterative value function to converge to the optimum as the iteration index increases to infinity. The properties of the iterative control law are analyzed, and the detailed implementation of the optimal impulsive control law is presented. Finally, two simulation examples with comparisons are given to show the effectiveness of the developed method. Qinglai Wei, Ruizhuo Song, Zehua Liao, Benkai Li, Frank L. Lewis |
IEEE Trans. Cybern. | 1 |
| 2020 | Optimal Elevator Group Control via Deep Asynchronous Actor-Critic LearningabstractIn this article, a new deep reinforcement learning (RL) method, called asynchronous advantage actor-critic (A3C) method, is developed to solve the optimal control problem of elevator group control systems (EGCSs). The main contribution of this article is that the optimal control law of EGCSs is designed via a new deep RL method, such that the elevator system sends passengers to the desired destination floors as soon as possible. Deep convolutional and recurrent neural networks, which can update themselves during applications, are designed to dispatch elevators. Then, the structure of the A3C method is developed, and the training phase for the learning optimal law is discussed. Finally, simulation results illustrate that the developed method effectively reduces the average waiting time in a complex building environment. Comparisons with traditional algorithms further verify the effectiveness of the developed method. Qinglai Wei, Yu Liu 0078, Marios M. Polycarpou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Editorial Special Issue on Adaptive Dynamic Programming and Reinforcement LearningabstractThe past decade has witnessed a surge in research activities related to adaptive dynamic programming (ADP) and reinforcement learning (RL), particularly for control applications. Several books [item 1)–5) in the Appendix] and survey papers [item 6)–10) in the Appendix] have been published on the subject. Both ADP and RL provide approximate solutions to dynamic programming problems. In a 1995 article by Bartoet al.[item 11) in the Appendix], they introduced the so-called “adaptive real-time dynamic programming,” which was specifically to apply ADP for real-time control. Later, in 2002, Murrayet al.[item 12) in the Appendix] developed an ADP algorithm for optimal control of continuous-time affine nonlinear systems. On the other hand, the most famous algorithms in RL are the temporal difference algorithm [item 13) in the Appendix] and the Q-learning algorithm [item 14) and 15) in the Appendix]. Derong Liu 0001, Frank L. Lewis, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Distributed Adaptive Dynamic Programming Algorithm for Office Energy Control with Multiple Batteries
Chao Li 0024, Bo Zhao 0015, Qinglai Wei, Derong Liu 0001 |
IJCNN | 4 |
| 2019 | Event-Triggered Adaptive Control for Discrete-Time Zero-Sum GamesabstractIn this paper, an event-triggered adaptive dynamic programming (ADP) method is developed for the discrete-time nonlinear two-player zero-sum games. First, an event-triggered ADP algorithm is presented to solve the Hamilton-Jacobi-Isaacs (HJI) equation. Then, a novel double event-triggered scheme is designed, the control inputs and the disturbance inputs will be updated only when the triggering conditions are satisfied. Therefore, the computational burden and the communication cost can be reduced. The algorithm is implemented by two neural networks, and the stability of the two-player system is proved. Finally, an example is employed to illustrate the effectiveness of the developed method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 2 |
| 2019 | A Solution of Two-Person Zero Sum Differential Games with Incomplete State Information
Kanghao Du, Ruizhuo Song, Qinglai Wei, Bo Zhao 0015 |
ISNN (1) | 3 |
| 2019 | An echo state network based approach to room classification of office buildings
Bo Zhao 0015, Chao Li 0024, Qinglai Wei, Derong Liu 0001 |
Neurocomputing | 4 |
| 2019 | Editorial: Booming of Neural Networks and Learning SystemsabstractAs you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community. Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He |
IEEE Trans. Neural Networks Learn. Syst. | 16 |
| 2018 | Adaptive Critic Designs of Optimal Control for Ice Storage Air Conditioning Systems
Zehua Liao, Qinglai Wei |
ICONIP (7) | 2 |
| 2018 | Local Tracking Control for Unknown Interconnected Systems via Neuro-Dynamic Programming
Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Ding Wang 0001, Yancai Xu, Qinglai Wei |
ICONIP (7) | 6 |
| 2018 | Reinforcement learning for robust adaptive control of partially unknown nonlinear systems subject to unmatched uncertainties
Xiong Yang 0001, Haibo He, Qinglai Wei, Biao Luo 0001 |
Inf. Sci. | 3 |
| 2018 | Adaptive Dynamic Programming for Discrete-Time Zero-Sum GamesabstractIn this paper, a novel adaptive dynamic programming (ADP) algorithm, called "iterative zero-sum ADP algorithm," is developed to solve infinite-horizon discrete-time two-player zero-sum games of nonlinear systems. The present iterative zero-sum ADP algorithm permits arbitrary positive semidefinite functions to initialize the upper and lower iterations. A novel convergence analysis is developed to guarantee the upper and lower iterative value functions to converge to the upper and lower optimums, respectively. When the saddle-point equilibrium exists, it is emphasized that both the upper and lower iterative value functions are proved to converge to the optimal solution of the zero-sum game, where the existence criteria of the saddle-point equilibrium are not required. If the saddle-point equilibrium does not exist, the upper and lower optimal performance index functions are obtained, respectively, where the upper and lower performance index functions are proved to be not equivalent. Finally, simulation results and comparisons are shown to illustrate the performance of the present method. Qinglai Wei, Derong Liu 0001, Qiao Lin 0003, Ruizhuo Song |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Discrete-Time Stable Generalized Self-Learning Optimal Control With Approximation ErrorsabstractIn this paper, a generalized policy iteration (GPI) algorithm with approximation errors is developed for solving infinite horizon optimal control problems for nonlinear systems. The developed stable GPI algorithm provides a general structure of discrete-time iterative adaptive dynamic programming algorithms, by which most of the discrete-time reinforcement learning algorithms can be described using the GPI structure. It is for the first time that approximation errors are explicitly considered in the GPI algorithm. The properties of the stable GPI algorithm with approximation errors are analyzed. The admissibility of the approximate iterative control law can be guaranteed if the approximation errors satisfy the admissibility criteria. The convergence of the developed algorithm is established, which shows that the iterative value function is convergent to a finite neighborhood of the optimal performance index function, if the approximate errors satisfy the convergence criterion. Finally, numerical examples and comparisons are presented. Qinglai Wei, Benkai Li, Ruizhuo Song |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Discrete-Time Local Value Iteration Adaptive Dynamic Programming: Convergence AnalysisabstractIn this paper, convergence properties are established for the newly developed discrete-time local value iteration adaptive dynamic programming (ADP) algorithm. The present local iterative ADP algorithm permits an arbitrary positive semidefinite function to initialize the algorithm. Employing a state-dependent learning rate function, for the first time, the iterative value function and iterative control law can be updated in a subset of the state space instead of the whole state space, which effectively relaxes the computational burden. A new analysis method for the convergence property is developed to prove that the iterative value functions will converge to the optimum under some mild constraints. Monotonicity of the local value iteration ADP algorithm is presented, which shows that under some special conditions of the initial value function and the learning rate function, the iterative value function can monotonically converge to the optimum. Finally, three simulation examples and comparisons are given to illustrate the performance of the developed algorithm. Qinglai Wei, Frank L. Lewis, Derong Liu 0001, Ruizhuo Song, Hanquan Lin |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2017 | Batch Process Fault Monitoring Based on LPGD-kNN and Its Applications in Semiconductor Industry
Ting Li 0021, Dongsheng Yang 0001, Qinglai Wei, Huaguang Zhang |
ICONIP (1) | 3 |
| 2017 | Partially-Directed-Topology-Based Consensus Control for Linear Multi-agent Systems
Chunping Shi, Qinglai Wei, Derong Liu 0001 |
ICONIP (6) | 2 |
| 2017 | An Event-Triggered Heuristic Dynamic Programming Algorithm for Discrete-Time Nonlinear Systems
Qinglai Wei, Derong Liu 0001 |
ICONIP (1) | 2 |
| 2017 | A Generalized Policy Iteration Adaptive Dynamic Programming Algorithm for Optimal Control of Discrete-Time Nonlinear Systems with Actuator Saturation
Qiao Lin 0003, Qinglai Wei, Bo Zhao 0015 |
ISNN (2) | 2 |
| 2017 | Local Policy Iteration Adaptive Dynamic Programming for Discrete-Time Nonlinear Systems
Qinglai Wei, Yancai Xu, Qiao Lin 0003, Derong Liu 0001, Ruizhuo Song |
ISNN (2) | 1 |
| 2017 | Neural-network-based synchronous iteration learning method for multi-player zero-sum games
Ruizhuo Song, Qinglai Wei, Biao Song |
Neurocomputing | 2 |
| 2017 | Off-policy neuro-optimal control for unknown complex-valued nonlinear systems based on policy iteration
Ruizhuo Song, Qinglai Wei, Wendong Xiao |
Neural Comput. Appl. | 2 |
| 2017 | Optimization of electricity consumption in office buildings based on adaptive dynamic programming
Qinglai Wei, Derong Liu 0001 |
Soft Comput. | 2 |
| 2017 | Discrete-Time Optimal Control via Local Policy Iteration Adaptive Dynamic ProgrammingabstractIn this paper, a discrete-time optimal control scheme is developed via a novel local policy iteration adaptive dynamic programming algorithm. In the discrete-time local policy iteration algorithm, the iterative value function and iterative control law can be updated in a subset of the state space, where the computational burden is relaxed compared with the traditional policy iteration algorithm. Convergence properties of the local policy iteration algorithm are presented to show that the iterative value function is monotonically nonincreasing and converges to the optimum under some mild conditions. The admissibility of the iterative control law is proven, which shows that the control system can be stabilized under any of the iterative control laws, even if the iterative control law is updated in a subset of the state space. Finally, two simulation examples are given to illustrate the performance of the developed method. Qinglai Wei, Derong Liu 0001, Qiao Lin 0003, Ruizhuo Song |
IEEE Trans. Cybern. | 1 |
| 2017 | Discrete-Time Deterministic Q-Learning: A Novel Convergence AnalysisabstractIn this paper, a novel discrete-time deterministic Q -learning algorithm is developed. In each iteration of the developed Q -learning algorithm, the iterative Q function is updated for all the state and control spaces, instead of updating for a single state and a single control in traditional Q -learning algorithm. A new convergence criterion is established to guarantee that the iterative Q function converges to the optimum, where the convergence criterion of the learning rates for traditional Q -learning algorithms is simplified. During the convergence analysis, the upper and lower bounds of the iterative Q function are analyzed to obtain the convergence criterion, instead of analyzing the iterative Q function itself. For convenience of analysis, the convergence properties for undiscounted case of the deterministic Q -learning algorithm are first developed. Then, considering the discounted factor, the convergence criterion for the discounted case is established. Neural networks are used to approximate the iterative Q function and compute the iterative control law, respectively, for facilitating the implementation of the deterministic Q -learning algorithm. Finally, simulation results and comparisons are given to illustrate the performance of the developed algorithm. Qinglai Wei, Frank L. Lewis, Qiuye Sun, Ruizhuo Song |
IEEE Trans. Cybern. | 1 |
| 2017 | Off-Policy Integral Reinforcement Learning Method to Solve Nonlinear Continuous-Time Multiplayer Nonzero-Sum GamesabstractThis paper establishes an off-policy integral reinforcement learning (IRL) method to solve nonlinear continuous-time (CT) nonzero-sum (NZS) games with unknown system dynamics. The IRL algorithm is presented to obtain the iterative control and off-policy learning is used to allow the dynamics to be completely unknown. Off-policy IRL is designed to do policy evaluation and policy improvement in the policy iteration algorithm. Critic and action networks are used to obtain the performance index and control for each player. The gradient descent algorithm makes the update of critic and action weights simultaneously. The convergence analysis of the weights is given. The asymptotic stability of the closed-loop system and the existence of Nash equilibrium are proved. The simulation study demonstrates the effectiveness of the developed method for nonlinear CT NZS games with unknown system dynamics. Ruizhuo Song, Frank L. Lewis, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Discrete-Time Local Value Iteration Adaptive Dynamic Programming: Admissibility and Termination AnalysisabstractIn this paper, a novel local value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon optimal control problems for discrete-time nonlinear systems. The focuses of this paper are to study admissibility properties and the termination criteria of discrete-time local value iteration ADP algorithms. In the discrete-time local value iteration ADP algorithm, the iterative value functions and the iterative control laws are both updated in a given subset of the state space in each iteration, instead of the whole state space. For the first time, admissibility properties of iterative control laws are analyzed for the local value iteration ADP algorithm. New termination criteria are established, which terminate the iterative local ADP algorithm with an admissible approximate optimal control law. Finally, simulation results are given to illustrate the performance of the developed algorithm.In this paper, a novel local value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon optimal control problems for discrete-time nonlinear systems. The focuses of this paper are to study admissibility properties and the termination criteria of discrete-time local value iteration ADP algorithms. In the discrete-time local value iteration ADP algorithm, the iterative value functions and the iterative control laws are both updated in a given subset of the state space in each iteration, instead of the whole state space. For the first time, admissibility properties of iterative control laws are analyzed for the local value iteration ADP algorithm. New termination criteria are established, which terminate the iterative local ADP algorithm with an admissible approximate optimal control law. Finally, simulation results are given to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Qiao Lin 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Optimal Constrained Neuro-Dynamic Programming Based Self-learning Battery Management in Microgrids
Qinglai Wei, Derong Liu 0001 |
ICONIP (3) | 1 |
| 2016 | An adaptive dynamic programming based method for optimization of electricity consumption in office buildingsabstractIn this paper, an adaptive dynamic programming (ADP) based method is developed to optimize electricity consumption of rooms in office buildings through optimal battery management. Rooms in office buildings are generally divided into office room, computer room, storage room, meeting room, etc., each of which has different characteristics of electricity consumption, as divided into electricity consumption from sockets, lights and air-conditioners in this paper. The developed method based on ADP is elaborated, and different optimization strategies of electricity consumption in different categories of rooms are proposed in accordance with the developed method. Finally, a detailed case study on an office building is given to show the practical effect of the developed method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 2 |
| 2016 | Optimal self-learning control scheme for discrete-time nonlinear systems using local value iterationabstractIn this paper, an optimal self-learning control scheme for discrete-time nonlinear systems is developed using a new local value iteration based adaptive dynamic programming (ADP) algorithm. The developed local value iteration algorithm permits an arbitrary positive semi-definite function to initialize the algorithm. In the developed local value iteration algorithm, the iterative value function and iterative control law are updated by a subset of the state space. A new analysis method of the convergence property is presented to show that the iterative value functions will converge to the optimum. The convergence criterion for the local value iteration algorithm is presented. A simulation example is given to demonstrate the validity of the present optimal control scheme. Qinglai Wei, Derong Liu 0001 |
IJCNN | 1 |
| 2016 | Discrete-Time Two-Player Zero-Sum Games for Nonlinear Systems Using Iterative Adaptive Dynamic Programming
Qinglai Wei, Derong Liu 0001 |
ISNN | 1 |
| 2016 | Energy consumption prediction of office buildings based on echo state networks
Derong Liu 0001, Qinglai Wei |
Neurocomputing | 3 |
| 2016 | Guaranteed cost neural tracking control for a class of uncertain nonlinear systems using adaptive dynamic programming
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei, Ding Wang 0001 |
Neurocomputing | 3 |
| 2016 | ADP-based optimal sensor scheduling for target tracking in energy harvesting wireless sensor networks
Ruizhuo Song, Qinglai Wei, Wendong Xiao |
Neural Comput. Appl. | 2 |
| 2016 | Neuro-optimal tracking control for a class of discrete-time nonlinear systems via generalized value iteration adaptive dynamic programming approach
Qinglai Wei, Derong Liu 0001, Yancai Xu |
Soft Comput. | 1 |
| 2016 | Off-Policy Actor-Critic Structure for Optimal Control of Unknown Systems With DisturbancesabstractAn optimal control method is developed for unknown continuous-time systems with unknown disturbances in this paper. The integral reinforcement learning (IRL) algorithm is presented to obtain the iterative control. Off-policy learning is used to allow the dynamics to be completely unknown. Neural networks are used to construct critic and action networks. It is shown that if there are unknown disturbances, off-policy IRL may not converge or may be biased. For reducing the influence of unknown disturbances, a disturbances compensation controller is added. It is proven that the weight errors are uniformly ultimately bounded based on Lyapunov techniques. Convergence of the Hamiltonian function is also proven. The simulation study demonstrates the effectiveness of the proposed optimal control method for unknown systems with disturbances. Ruizhuo Song, Frank L. Lewis, Qinglai Wei, Huaguang Zhang |
IEEE Trans. Cybern. | 3 |
| 2016 | Value Iteration Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear SystemsabstractIn this paper, a value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon undiscounted optimal control problems for discrete-time nonlinear systems. The present value iteration ADP algorithm permits an arbitrary positive semi-definite function to initialize the algorithm. A novel convergence analysis is developed to guarantee that the iterative value function converges to the optimal performance index function. Initialized by different initial functions, it is proven that the iterative value function will be monotonically nonincreasing, monotonically nondecreasing, or nonmonotonic and will converge to the optimum. In this paper, for the first time, the admissibility properties of the iterative control laws are developed for value iteration algorithms. It is emphasized that new termination criteria are established to guarantee the effectiveness of the iterative control laws. Neural networks are used to approximate the iterative value function and compute the iterative control law, respectively, for facilitating the implementation of the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the present method. Qinglai Wei, Derong Liu 0001, Hanquan Lin |
IEEE Trans. Cybern. | 1 |
| 2016 | Data-Driven Zero-Sum Neuro-Optimal Control for a Class of Continuous-Time Unknown Nonlinear Systems With Disturbance Using ADPabstractThis paper is concerned with a new data-driven zero-sum neuro-optimal control problem for continuous-time unknown nonlinear systems with disturbance. According to the input-output data of the nonlinear system, an effective recurrent neural network is introduced to reconstruct the dynamics of the nonlinear system. Considering the system disturbance as a control input, a two-player zero-sum optimal control problem is established. Adaptive dynamic programming (ADP) is developed to obtain the optimal control under the worst case of the disturbance. Three single-layer neural networks, including one critic and two action networks, are employed to approximate the performance index function, the optimal control law, and the disturbance, respectively, for facilitating the implementation of the ADP method. Convergence properties of the ADP method are developed to show that the system state will converge to a finite neighborhood of the equilibrium. The weight matrices of the critic and the two action networks are also convergent to finite neighborhoods of their optimal ones. Finally, the simulation results will show the effectiveness of the developed data-driven ADP methods. Qinglai Wei, Ruizhuo Song |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Neural-Network-Based Distributed Adaptive Robust Control for a Class of Nonlinear Multiagent Systems With Time Delays and External NoisesabstractA class of nonlinear multiagent systems with time delays and external noises is investigated, and a distributed adaptive robust control protocol is developed. It is the first time for a class of multiagent systems to take both time delays and external noises into consideration. By virtue of Lyapunov-Krasovskii functional and Young's inequality, the effects of time delay can be eliminated. Then, to exclude external noises, a robustifying term is introduced to eliminate the negative effects of these noises. Moreover, neural networks are utilized to learn the unknown nonlinear terms to adapt to the complex external environment. Finally, a numerical simulation is conducted to validate the effectiveness of our distributed control protocol. Hongwen Ma, Zhuo Wang 0003, Ding Wang 0001, Derong Liu 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2015 | Robust Tracking Control of Uncertain Nonlinear Systems Using Adaptive Dynamic Programming
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei |
ICONIP (3) | 3 |
| 2015 | A New Discrete-Time Iterative Adaptive Dynamic Programming Algorithm Based on Q-LearningabstractIn this paper, a novel Q -learning based policy iteration adaptive dynamic programming (ADP) algorithm is developed to solve the optimal control problems for discrete-time nonlinear systems. The idea is to use a policy iteration ADP technique to construct the iterative control law which stabilizes the system and simultaneously minimizes the iterative Q function. Convergence property is analyzed to show that the iterative Q function is monotonically non-increasing and converges to the solution of the optimality equation. Finally, simulation results are presented to show the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001 |
ISNN | 1 |
| 2015 | A novel policy iteration based deterministic Q-learning for discrete-time nonlinear systems
Qinglai Wei, Derong Liu 0001 |
Sci. China Inf. Sci. | 1 |
| 2015 | Nearly finite-horizon optimal control for a class of nonaffine time-delay nonlinear systems based on adaptive dynamic programming
Ruizhuo Song, Qinglai Wei, Qiuye Sun |
Neurocomputing | 2 |
| 2015 | Neural-network-based adaptive optimal tracking control scheme for discrete-time nonlinear systems with approximation errors
Qinglai Wei, Derong Liu 0001 |
Neurocomputing | 1 |
| 2015 | Nonlinear neuro-optimal tracking control via stable iterative Q-learning algorithm
Qinglai Wei, Ruizhuo Song, Qiuye Sun |
Neurocomputing | 1 |
| 2015 | Optimal distributed synchronization control for continuous-time heterogeneous multi-agent differential graphical games
Qinglai Wei, Derong Liu 0001, Frank L. Lewis |
Inf. Sci. | 1 |
| 2015 | Reinforcement-Learning-Based Robust Controller Design for Continuous-Time Uncertain Nonlinear Systems Subject to Input ConstraintsabstractThe design of stabilizing controller for uncertain nonlinear systems with control constraints is a challenging problem. The constrained-input coupled with the inability to identify accurately the uncertainties motivates the design of stabilizing controller based on reinforcement-learning (RL) methods. In this paper, a novel RL-based robust adaptive control algorithm is developed for a class of continuous-time uncertain nonlinear systems subject to input constraints. The robust control problem is converted to the constrained optimal control problem with appropriately selecting value functions for the nominal system. Distinct from typical action-critic dual networks employed in RL, only one critic neural network (NN) is constructed to derive the approximate optimal control. Meanwhile, unlike initial stabilizing control often indispensable in RL, there is no special requirement imposed on the initial control. By utilizing Lyapunov's direct method, the closed-loop optimal control system and the estimated weights of the critic NN are proved to be uniformly ultimately bounded. In addition, the derived approximate optimal control is verified to guarantee the uncertain nonlinear system to be stable in the sense of uniform ultimate boundedness. Two simulation examples are provided to illustrate the effectiveness and applicability of the present approach. Derong Liu 0001, Xiong Yang 0001, Ding Wang 0001, Qinglai Wei |
IEEE Trans. Cybern. | 4 |
| 2015 | Multiple Actor-Critic Structures for Continuous-Time Optimal Control Using Input-Output DataabstractIn industrial process control, there may be multiple performance objectives, depending on salient features of the input-output data. Aiming at this situation, this paper proposes multiple actor-critic structures to obtain the optimal control via input-output data for unknown nonlinear systems. The shunting inhibitory artificial neural network (SIANN) is used to classify the input-output data into one of several categories. Different performance measure functions may be defined for disparate categories. The approximate dynamic programming algorithm, which contains model module, critic network, and action network, is used to establish the optimal control in each category. A recurrent neural network (RNN) model is used to reconstruct the unknown system dynamics using input-output data. NNs are used to approximate the critic and action networks, respectively. It is proven that the model error and the closed unknown system are uniformly ultimately bounded. Simulation results demonstrate the performance of the proposed optimal control scheme for the unknown nonlinear system. Ruizhuo Song, Frank L. Lewis, Qinglai Wei, Huaguang Zhang, Zhong-Ping Jiang, Daniel S. Levine 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Infinite Horizon Self-Learning Optimal Control of Nonaffine Discrete-Time Nonlinear SystemsabstractIn this paper, a novel iterative adaptive dynamic programming (ADP)-based infinite horizon self-learning optimal control algorithm, called generalized policy iteration algorithm, is developed for nonaffine discrete-time (DT) nonlinear systems. Generalized policy iteration algorithm is a general idea of interacting policy and value iteration algorithms of ADP. The developed generalized policy iteration algorithm permits an arbitrary positive semidefinite function to initialize the algorithm, where two iteration indices are used for policy improvement and policy evaluation, respectively. It is the first time that the convergence, admissibility, and optimality properties of the generalized policy iteration algorithm for DT nonlinear systems are analyzed. Neural networks are used to implement the developed algorithm. Finally, numerical examples are presented to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Xiong Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Generalized Policy Iteration Adaptive Dynamic Programming for Discrete-Time Nonlinear SystemsabstractThis paper is concerned with a novel generalized policy iteration algorithm for solving optimal control problems for discrete-time nonlinear systems. The idea is to use an iterative adaptive dynamic programming algorithm to obtain iterative control laws which make the iterative value functions converge to the optimum. Initialized by an admissible control law, it is shown that the iterative value functions are monotonically nonincreasing and converge to the optimal solution of Hamilton-Jacobi-Bellman equation, under the assumption that a perfect function approximation is employed. The admissibility property is analyzed, which shows that any of the iterative control laws can stabilize the nonlinear system. Neural networks are utilized to implement the generalized policy iteration algorithm, by approximating the iterative value function and computing the iterative control law, respectively, to achieve approximate optimal control. Finally, numerical examples are presented to verify the effectiveness of the present generalized policy iteration algorithm. Derong Liu 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2014 | Optimal self-learning battery control in smart residential grids by iterative Q-learning algorithmabstractIn this paper, a novel dual iterative Q-learning algorithm is developed to solve the optimal battery management and control problems in smart residential environments. The main idea is to use adaptive dynamic programming (ADP) technique to obtain the optimal battery management and control scheme iteratively for residential energy systems. In the developed dual iterative Q-learning algorithm, two iterations, including external and internal iterations, are introduced, where internal iteration minimizes the total cost of power loads in each period and the external iteration makes the iterative Q function converge to the optimum. For the first time, the convergence property of iterative Q-learning method is proven to guarantee the convergence property of the iterative Q function. Finally, numerical results are given to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Yu Liu 0078, Qiang Guan |
ADPRL | 1 |
| 2014 | Discrete-Time Nonlinear Generalized Policy Iteration for Optimal Control Using Neural Networks
Qinglai Wei, Derong Liu 0001, Xiong Yang 0001 |
ICONIP (1) | 1 |
| 2014 | Neural-network-based optimal control for a class of complex-valued nonlinear systems with input saturationabstractThis paper proposes an optimal control scheme based on adaptive dynamic programming (ADP) algorithm for complex-valued systems with input saturation. The equivalence transformation is used to obtain the real dynamic system. Then the performance index function is defined. Based on the transformed system, an ADP optimal control method is established. The update methods for critic network neural network and action network are given. It is proved that the closed-loop system is uniformly ultimately bounded based on Lyapunov approach. Finally, the simulation study was given to show the effectiveness of the proposed optimal control scheme. Ruizhuo Song, Qinglai Wei, Zenglian Zhang, Biao Song |
IJCNN | 2 |
| 2014 | Near-optimal online control of uncertain nonlinear continuous-time systems based on concurrent learningabstractThis paper presents a novel observer-critic architecture for solving the near-optimal control problem of uncertain nonlinear continuous-time systems. Two neural networks (NNs) are employed in the architecture: an observer NN is constructed to get the knowledge of uncertain system dynamics and a critic NN is utilized to derive the optimal control. The observer NN and the critic NN are tuned simultaneously. By using the recorded and instantaneous data together, the optimal control can be derived without the persistence of excitation condition. Meanwhile, the closed-loop system is guaranteed to be stable in the sense of uniform ultimate boundedness. No initial stabilizing control is required in the developed algorithm. An illustrated example is provided to demonstrate the effectiveness of the present approach. Xiong Yang 0001, Derong Liu 0001, Qinglai Wei |
IJCNN | 3 |
| 2014 | Reinforcement-Learning-Based Controller Design for Nonaffine Nonlinear Systems
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei |
ISNN | 3 |
| 2014 | Special issue on International Symposium on Neural Networks
Zhigang Zeng, Amir Hussain 0001, Qinglai Wei |
Neurocomputing | 3 |
| 2014 | Data-based analysis of discrete-time linear systems in noisy environment: Controllability and observability
Derong Liu 0001, Qinglai Wei |
Inf. Sci. | 3 |
| 2014 | Stable iterative adaptive dynamic programming algorithm with approximation errors for discrete-time nonlinear systems
Qinglai Wei, Derong Liu 0001 |
Neural Comput. Appl. | 1 |
| 2014 | Discrete-time online learning control for a class of unknown nonaffine nonlinear systems using reinforcement learning
Xiong Yang 0001, Derong Liu 0001, Ding Wang 0001, Qinglai Wei |
Neural Networks | 4 |
| 2014 | Neural-network-based approach to finite-time optimal control for a class of unknown nonlinear systems
Ruizhuo Song, Wendong Xiao, Qinglai Wei, Changyin Sun 0001 |
Soft Comput. | 3 |
| 2014 | Adaptive Dynamic Programming for Optimal Tracking Control of Unknown Nonlinear Systems With Application to Coal GasificationabstractIn this paper, we establish a new data-based iterative optimal learning control scheme for discrete-time nonlinear systems using iterative adaptive dynamic programming (ADP) approach and apply the developed control scheme to solve a coal gasification optimal tracking control problem. According to the system data, neural networks (NNs) are used to construct the dynamics of coal gasification process, coal quality and reference control, respectively, where the mathematical model of the system is unnecessary. The approximation errors from neural network construction of the disturbance and the controls are both considered. Via system transformation, the optimal tracking control problem with approximation errors and disturbances is effectively transformed into a two-person zero-sum optimal control problem. A new iterative ADP algorithm is then developed to obtain the optimal control laws for the transformed system. Convergence property is developed to guarantee that the performance index function converges to a finite neighborhood of the optimal performance index function, and the convergence criterion is also obtained. Finally, numerical results are given to illustrate the performance of the present method. Qinglai Wei, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2014 | A Novel Iterative $\theta $-Adaptive Dynamic Programming for Discrete-Time Nonlinear SystemsabstractThis paper is concerned with a new iterative θ-adaptive dynamic programming (ADP) technique to solve optimal control problems of infinite horizon discrete-time nonlinear systems. The idea is to use an iterative ADP algorithm to obtain the iterative control law which optimizes the iterative performance index function. In the present iterative θ-ADP algorithm, the condition of initial admissible control in policy iteration algorithm is avoided. It is proved that all the iterative controls obtained in the iterative θ-ADP algorithm can stabilize the nonlinear system which means that the iterative θ-ADP algorithm is feasible for implementations both online and offline. Convergence analysis of the performance index function is presented to guarantee that the iterative performance index function will converge to the optimum monotonically. Neural networks are used to approximate the performance index function and compute the optimal control policy, respectively, for facilitating the implementation of the iterative θ-ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the established method. Qinglai Wei, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2014 | Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic ProgrammingabstractIn this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal control problems for infinite horizon discrete-time nonlinear systems with finite approximation errors. First, a new generalized value iteration algorithm of ADP is developed to make the iterative performance index function converge to the solution of the Hamilton-Jacobi-Bellman equation. The generalized value iteration algorithm permits an arbitrary positive semi-definite function to initialize it, which overcomes the disadvantage of traditional value iteration algorithms. When the iterative control law and iterative performance index function in each iteration cannot accurately be obtained, for the first time a new "design method of the convergence criteria" for the finite-approximation-error-based generalized value iteration algorithm is established. A suitable approximation error can be designed adaptively to make the iterative performance index function converge to a finite neighborhood of the optimal performance index function. Neural networks are used to implement the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the developed method. Qinglai Wei, Fei-Yue Wang 0001, Derong Liu 0001, Xiong Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Policy Iteration Adaptive Dynamic Programming Algorithm for Discrete-Time Nonlinear SystemsabstractThis paper is concerned with a new discrete-time policy iteration adaptive dynamic programming (ADP) method for solving the infinite horizon optimal control problem of nonlinear systems. The idea is to use an iterative ADP technique to obtain the iterative control law, which optimizes the iterative performance index function. The main contribution of this paper is to analyze the convergence and stability properties of policy iteration method for discrete-time nonlinear systems for the first time. It shows that the iterative performance index function is nonincreasingly convergent to the optimal solution of the Hamilton-Jacobi-Bellman equation. It is also proven that any of the iterative control laws can stabilize the nonlinear systems. Neural networks are used to approximate the performance index function and compute the optimal control law, respectively, for facilitating the implementation of the iterative ADP algorithm, where the convergence of the weight matrices is analyzed. Finally, the numerical results and analysis are presented to illustrate the performance of the developed method. Derong Liu 0001, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Convergence analysis of continuous-time systems based on feedforward neural networksabstractIn this paper, we construct a feedforward neural network (NN) based system containing two NNs. The convergence of the NN based system is analyzed in detail. For setting up the NN based system, an NN observer is first designed to estimate the system states. Then, based on the observed states, a feedforward neuro-control system is constructed by using adaptive dynamic programming (ADP). In this design, two NN structures are used: a three-layer feedforward NN to constitute the observer which can be applied to the systems with high degrees of nonlinearity and without a priori knowledge about system dynamics, and a critic NN to approximate the value function. Moreover, the weight update laws for the critic NN are generated using a gradient-descent method based on a modified temporal difference error, which is independent of the system dynamics. Finally, uniform ultimate boundedness (UUB) of the NN based system is proved. Yuzhu Huang, Derong Liu 0001, Qinglai Wei |
ISCAS | 3 |
| 2013 | Neural Network H ∞ Tracking Control of Nonlinear Systems Using GHJI Method
Derong Liu 0001, Yuzhu Huang, Qinglai Wei |
ISNN (2) | 3 |
| 2013 | Optimal Tracking Control Scheme for Discrete-Time Nonlinear Systems with Approximation Errors
Qinglai Wei, Derong Liu 0001 |
ISNN (2) | 1 |
| 2013 | The neural paradigm for complex systems: new algorithms and applications
Stefano Squartini, Jinhu Lü 0001, Qinglai Wei |
Neural Comput. Appl. | 3 |
| 2013 | Dual iterative adaptive dynamic programming for a class of discrete-time nonlinear systems with time-delays
Qinglai Wei, Ding Wang 0001, Dehua Zhang |
Neural Comput. Appl. | 1 |
| 2013 | Multi-objective optimal control for a class of nonlinear time-delay systems via adaptive dynamic programming
Ruizhuo Song, Wendong Xiao, Qinglai Wei |
Soft Comput. | 3 |
| 2013 | Finite-Approximation-Error-Based Optimal Control Approach for Discrete-Time Nonlinear SystemsabstractIn this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal control problems for infinite-horizon discrete-time nonlinear systems with finite approximation errors. The idea is to use an iterative ADP algorithm to obtain the iterative control law that makes the iterative performance index function reach the optimum. When the iterative control law and the iterative performance index function in each iteration cannot be accurately obtained, the convergence conditions of the iterative ADP algorithm are obtained. When convergence conditions are satisfied, it is shown that the iterative performance index functions can converge to a finite neighborhood of the greatest lower bound of all performance index functions under some mild assumptions. Neural networks are used to approximate the performance index function and compute the optimal control policy, respectively, for facilitating the implementation of the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the present method. Derong Liu 0001, Qinglai Wei |
IEEE Trans. Cybern. | 2 |
| 2013 | Optimal Home Energy Management Under Dynamic Electrical and Thermal ConstraintsabstractThe optimization of energy consumption, with consequent costs reduction, is one of the main challenges in present and future smart grids. Of course, this has to occur keeping the living comfort for the end-user unchanged. In this work, an approach based on the mixed-integer linear programming paradigm, which is able to provide an optimal solution in terms of tasks power consumption and management of renewable resources, is developed. The proposed algorithm yields an optimal task scheduling under dynamic electrical constraints, while simultaneously ensuring the thermal comfort according to the user needs. On purpose, a suitable thermal model based on heat-pump usage has been considered in the framework. Some computer simulations using real data have been performed, and obtained results confirm the efficiency and robustness of the algorithm, also in terms of achievable cost savings. Francesco De Angelis 0002, Matteo Boaro, Danilo Fuselli, Stefano Squartini, Francesco Piazza, Qinglai Wei |
IEEE Trans. Ind. Informatics | 6 |
| 2012 | Generalized Hamilton-Jacobi-Isaacs Formulation-Based Neural Network H ∞ Control for Constrained Input Nonlinear Systems
Yuzhu Huang, Derong Liu 0001, Qinglai Wei |
ICONIP (1) | 3 |
| 2012 | Nearly Optimal Control for Nonlinear Systems with Dead-Zone Control Input Based on the Iterative ADP Approach
Dehua Zhang, Derong Liu 0001, Qinglai Wei |
ICONIP (1) | 3 |
| 2012 | Adaptive dynamic programming with stable value iteration algorithm for discrete-time nonlinear systemsabstractIn this paper, a new stable value iteration adaptive dynamic programming (ADP) algorithm, named “θ-ADP” algorithm, is proposed for solving the optimal control problems of infinite horizon discrete-time nonlinear systems. By introducing a parameter θ in the iterative ADP algorithm, it is proved that any of iterative control obtained in the proposed algorithm can stabilize the nonlinear system which overcomes the disadvantage of traditional value iteration algorithms. Neural networks are used to approximate the performance index function and compute the optimal control policy, respectively, for facilitating the implementation of the iterative θ-ADP algorithm. Finally, a simulation example is given to illustrate the performance of the proposed method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 1 |
| 2012 | Optimal Task and Energy Scheduling in Dynamic Residential Scenarios
Francesco De Angelis 0002, Matteo Boaro, Danilo Fuselli, Stefano Squartini, Francesco Piazza, Qinglai Wei, Ding Wang 0001 |
ISNN (1) | 6 |
| 2012 | Optimal Battery Management with ADHDP in Smart Home Environments
Danilo Fuselli, Francesco De Angelis 0002, Matteo Boaro, Derong Liu 0001, Qinglai Wei, Stefano Squartini, Francesco Piazza |
ISNN (2) | 5 |
| 2012 | Temperature Control in Water-Gas Shift Reaction with Adaptive Dynamic Programming
Yuzhu Huang, Derong Liu 0001, Qinglai Wei |
ISNN (2) | 3 |
| 2012 | Self-learning Control Schemes for Two-Person Zero-Sum Differential Games of Continuous-Time Nonlinear Systems with Saturating Controllers
Qinglai Wei, Derong Liu 0001 |
ISNN (2) | 1 |
| 2012 | Finite-horizon neuro-optimal tracking control for a class of discrete-time nonlinear systems using adaptive dynamic programming approach
Ding Wang 0001, Derong Liu 0001, Qinglai Wei |
Neurocomputing | 3 |
| 2012 | An iterative ϵ-optimal control scheme for a class of discrete-time nonlinear systems with unfixed initial state
Qinglai Wei, Derong Liu 0001 |
Neural Networks | 1 |
| 2012 | Neural-Network-Based Optimal Control for a Class of Unknown Discrete-Time Nonlinear Systems Using Globalized Dual Heuristic ProgrammingabstractIn this paper, a neuro-optimal control scheme for a class of unknown discrete-time nonlinear systems with discount factor in the cost function is developed. The iterative adaptive dynamic programming algorithm using globalized dual heuristic programming technique is introduced to obtain the optimal controller with convergence analysis in terms of cost function and control law. In order to carry out the iterative algorithm, a neural network is constructed first to identify the unknown controlled system. Then, based on the learned system model, two other neural networks are employed as parametric structures to facilitate the implementation of the iterative algorithm, which aims at approximating at each iteration the cost function and its derivatives and the control law, respectively. Finally, a simulation example is provided to verify the effectiveness of the proposed optimal control approach. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao, Qinglai Wei |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2011 | Optimal control for discrete-time nonlinear systems with unfixed initial state using adaptive dynamic programmingabstractA new ε-optimal control algorithm based on the adaptive dynamic programming (ADP) is proposed to solve the finite horizon optimal control problem for a class of discrete-time nonlinear systems with unfixed initial state. The proposed algorithm makes the performance index function converges iteratively to the greatest lower bound of all performance indices within an error bound according to ε with finite time. The number of optimal control steps can also be obtained by the proposed ADP approach for the situation when the initial state of the system is unfixed. A simulation example is given to show the performance of the present method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 1 |
| 2011 | Finite Horizon Optimal Tracking Control for a Class of Discrete-Time Nonlinear Systems
Qinglai Wei, Ding Wang 0001, Derong Liu 0001 |
ISNN (2) | 1 |
| 2011 | Optimal Tracking Control for a Class of Nonlinear Discrete-Time Systems With Time Delays Based on Heuristic Dynamic ProgrammingabstractIn this paper, a novel heuristic dynamic programming (HDP) iteration algorithm is proposed to solve the optimal tracking control problem for a class of nonlinear discrete-time systems with time delays. The novel algorithm contains state updating, control policy iteration, and performance index iteration. To get the optimal states, the states are also updated. Furthermore, the "backward iteration" is applied to state updating. Two neural networks are used to approximate the performance index function and compute the optimal control policy for facilitating the implementation of HDP iteration algorithm. At last, we present two examples to demonstrate the effectiveness of the proposed HDP iteration algorithm. Huaguang Zhang, Ruizhuo Song, Qinglai Wei, Tieyan Zhang |
IEEE Trans. Neural Networks | 3 |
| 2010 | Adaptive dynamic programming for a class of discrete-time non-affine nonlinear systems with time-delaysabstractIn this paper, an optimal control scheme for a class of non-affine nonlinear systems with time-delays in state and control variables is developed using a new iterative adaptive dynamic programming (ADP) algorithm. By introducing delay matrix functions, the explicit expression of optimal control is solved using the dynamic programming theory and the optimal control can iteratively be solved using the present technique. Convergence analysis is presented to show the performance index function to reach the optimum by the present method. Neural networks are used to approximate the performance index function, compute the optimal control policy, solve delay matrix functions and model the nonlinear system, respectively, for facilitating the implementation of the iterative ADP algorithm. A simulation example is given to demonstrate the validity of the present optimal control scheme. Derong Liu 0001, Qinglai Wei |
IJCNN | 2 |
| 2010 | Optimal control laws for time-delay systems with saturating actuators based on heuristic dynamic programming
Ruizhuo Song, Huaguang Zhang, Qinglai Wei |
Neurocomputing | 4 |
| 2009 | Model-free multiobjective approximate dynamic programming for discrete-time nonlinear systems with general performance index functions
Qinglai Wei, Huaguang Zhang |
Neurocomputing | 1 |
| 2008 | Adaptive Dynamic Programming for a Class of Nonlinear Control Systems with General Separable Performance Index
Qinglai Wei, Derong Liu 0001, Huaguang Zhang |
ISNN (2) | 1 |
| 2008 | A Novel Infinite-Time Optimal Tracking Control Scheme for a Class of Discrete-Time Nonlinear Systems via the Greedy HDP Iteration AlgorithmabstractIn this paper, we aim to solve the infinite-time optimal tracking control problem for a class of discrete-time nonlinear systems using the greedy heuristic dynamic programming (HDP) iteration algorithm. A new type of performance index is defined because the existing performance indexes are very difficult in solving this kind of tracking problem, if not impossible. Via system transformation, the optimal tracking problem is transformed into an optimal regulation problem, and then, the greedy HDP iteration algorithm is introduced to deal with the regulation problem with rigorous convergence analysis. Three neural networks are used to approximate the performance index, compute the optimal control policy, and model the nonlinear system for facilitating the implementation of the greedy HDP iteration algorithm. An example is given to demonstrate the validity of the proposed optimal tracking control scheme. Huaguang Zhang, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | On-Line Learning Control for Discrete Nonlinear Systems Via an Improved ADDHP Method
Huaguang Zhang, Qinglai Wei, Derong Liu 0001 |
ISNN (1) | 2 |