EDBT 2026 Demo / reviewers in the wild / expert
Frank L. Lewis
dblp:82/4281
· DBLP profile ↗
210ranked-venue papers
7as first author
81since 2021 · last 2026
0000-0003-4074-1615ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 136 · 4 first-author · 45 since 2021Human-computer interaction and ubiquitous computing · 37 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 15 since 2021Systems, architecture and hardware · 15 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prescribed-rate target tracking for time-delayed systems using output measurements
Ci Chen 0002, Frank L. Lewis, Kan Xie 0002, Shengli Xie 0001 |
Neural Networks | 3 |
| 2026 | Dropout Resilience and Model-Free H∞ Off-Policy Learning for Unknown Target Tracking
Wenzhao Liu, Xumin Huang, Yu Wang 0050, Frank L. Lewis, Ci Chen 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Adaptive Nash Equilibria Seeking in Directed Networks With Application to Energy AllocationabstractThis paper focuses on the Nash equilibrium problem for non-cooperative games over general directed graphs with application to energy allocation. Unlike previous Nash equilibrium-seeking works that typically required strongly connected or undirected graphs, we now relax the network to contain a spanning tree. To overcome the inaccessibility caused by the general graph, we introduce an information collector who gathers player profiles and acts as the leader. Under this framework, each player preserves the state observation from leader through the spanning-tree network, and uses gradient descent to adjust variables for converging to an equilibrium point. We also present an enhanced version that invests the solo coupling gain with adaptation through distributed consensus error. The proposed schemes ensure the globally asymptotic stability of the Nash equilibrium-seeking process through Lyapunov analysis. Finally, numerical studies, as well as hardware-in-the-loop experiments for energy allocation on the RT-Lab platform, are provided to demonstrate the effectiveness of the proposed protocols. Zhiyang Zheng, Yu Wang 0050, Zhaoyu Xiang, Frank L. Lewis, Shengli Xie 0001, Ci Chen 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Inverse Reinforcement Learning for Disturbed Networked Nonlinear Systems With Data DropoutsabstractThis article develops inverse reinforcement learning (IRL) control algorithms for nonlinear networked control systems (NCSs) to mimic trajectories of a target system governed by an unknown optimal cost function, despite the presence of random data dropouts and external disturbances. Data dropouts occur during: 1) reception of target trajectory data by the controller; 2) reception of state feedback data by the controller; and 3) reception of control input data by the actuator. By organically integrating $H_{\infty } $ control to account for disturbances and dropout-induced uncertainty, a model-based IRL algorithm is first developed. Building on this, a neural-network-based data-driven IRL algorithm is developed to infer the cost function and optimal control policy using available data while reducing dependence on system models. The proposed methods enable effective trajectory imitation under partial model knowledge, data dropouts, and disturbances, as demonstrated through simulation studies. Wenqian Xue, Jialu Fan, Frank L. Lewis, Bosen Lian |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2026 | Complexity Dynamics and Fuzzy Optimal Prescribed-Time Control of a Large Network of Bidirectionally Coupled FO PMSG
Shaohua Luo, Ya Zhang 0001, Ye Cao 0001, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2026 | Initially Excited Asynchronous Reinforcement Learning Control With Monotonicity and StabilityabstractThis article addresses the existing reinforcement learning (RL) control issues of continuous-time Markov jump systems (MJSs), including their synchronous iteration structure using nonlatest updates, nonmonotonic convergence, and requirement of initial admissible control and all-time persistent excitation (PE) condition. We propose advanced model-based and model-free RL algorithms that: 1) have asynchronous decoupled Lyapunov iteration equations to approximate the optimal control solutions using the latest updates; 2) determine the initial admissible control policy (IACP) and initial value function matrix without relying on engineering experiences; and 3) relax PE condition with a milder initial excitation (IE) condition. Rigorous theoretical analyses are provided to establish the monotonic convergence to the optimal control solution and the closed-loop stability at each iteration. Finally, the simulation and comparison results verify the proposed algorithms. Wenqian Xue, Frank L. Lewis, Bosen Lian |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | An Optimal Synchronization Control Method of PLL Utilizing Adaptive Dynamic Programming to Synchronize Inverter-Based Resources With Unbalanced, Low-Inertia, and Very Weak GridsabstractWhen it comes to integrating inverter-based resources (IBRs) into modern grids with varying characteristics like unbalanced systems, low-inertia networks, or very weak grids, synthesizing the synchronization control method (SCM) of the IBR’s phase-locked loop can be a challenging task. This paper provides a unique solution to enhance the three-phase IBR’s SCM using the adaptive dynamic programming (ADP) method based on reinforcement learning. By making the SCM more intelligent and self-learning, IBRs can be easily integrated into diverse grids. To this end, this article investigates the synchronization process’s detailed dynamics, including all incorporating disturbances and parameters required for the first step in designing the ADP method. Afterward, this research synthesizes an optimal controller using an ADP method. It is a data-driven and practically sound approach to the problem under investigation. The new methodology is based on the adaptive optimal control employing measurement feedback to control the output regulation problem of uncertain synchronization process dynamics via the internal model principle. The proposed SCM design deploys an ADP learning methodology to tackle uncertain parameters and unknown disturbance signals to synchronize IBRs during transients, thereby enhancing IBRs’ synchronization in challenging conditions of modern power systems with unbalanced, low-inertia, and very weak grids. For comparison purposes, this paper applies a robust controller based on the well-established$\mu$synthesis approach (benefiting from the well-known$D\text{-}K$iteration process). Comparative simulations are performed; experiments are conducted to reveal the effectiveness and practicality of the ADP-based optimal SCM proposed in this paper.Note to Practitioners—As different nations strive to combat global warming and accelerate decarbonization, power and energy systems are undergoing a significant shift. Inverter-based resources are being used as an essential component to achieve these goals. However, studies have revealed that designing synchronization control methods of the inverter-based resources’ phase-locked loop in unbalanced, low-inertia, and very weak grids is challenging due to the need for accurate dynamic models and other factors. This study revisits the synchronization process’s detailed dynamics. It also proposes a novel adaptive dynamic programming strategy using intelligent self-learning approaches to the synchronization control method associated with inverter-based resources. This method utilizes an optimal control to synthesize the adaptive dynamic programming control strategy for the inverter-based resources’ synchronization process. Besides, it employs measurement feedback to control the output regulation problem of uncertain dynamics of inverter-based resources’ synchronization process via the internal model principle. As a result, this paper makes this process data-driven. It utilizes a learning methodology using adaptive dynamic programming to address uncertain parameters and unknown disturbance signals associated with the dynamics derived and formulated for the problem under investigation. Thus, the proposed method applies to controlling inverter-based resources’ synchronization process even in cases with slow parameter variations caused by different factors. It can compensate for all functional disturbance signals affecting the dynamics of the systems. In fact, unlike traditional methods that need an exact dynamic model of the inverter-based resources’ synchronization process to design and tune the controller to achieve a proper transient response, the proposed control system trains itself and does so. This study’s simulations and experiments reveal that the above points give the proposed approach a competitive edge over the existing methodologies. Masoud Davari, Weinan Gao, Amir Aghazadeh, Frede Blaabjerg, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Asymptotic Leader-Following Consensus of Heterogeneous Multi-Agent Systems With Unknown and Time-Varying Control GainsabstractThis paper investigates the consensus tracking control problem for a class of nonlinear multi-input multi-output (MIMO) heterogeneous multi-agent systems (HMASs), where the dimension of the dynamics of each agent is allowed to be different from each other. In addition, the control gain matrices (CGMs) with unknown time-varying coefficients and actuator faults with unpredictable jumps are also involved in the considered MIMO HMASs, all of which would cause damage for the system performance. First, a novel distributed auxiliary filter by feat of time-varying technology is introduced, which allows the zero-error estimation for the desired trajectory to be achieved and the asymptotic consensus tracking result to be further realized. Then, a cooperative adaptive control solution is proposed to ensure the asymptotical consensus tracking control result, in spite of the inherent unknown time-varying coefficients, unpredictable jumps caused by the unknown actuator faults, unknown disturbances and uncertain system parameters, distinguishing itself from those existing cooperative control works for HMASs where only ultimately uniformly bounded (UUB) result is derived. This is achieved mainly by the introduction of a series of Nussbaum functions and the employment of the adaptive estimation techniques. The effectiveness of the proposed control algorithm is confirmed by the simulation conducted on a group of HMASs involving unmanned aerial vehicles (UAVs) and autonomous surface vessels (ASVs).Note to Practitioners—In large-scale complex communication networks, uncertain HMASs with different structures and functions are capable of exchanging information and collaborating with each other to accomplish more complex and diverse tasks. Simultaneously, the probability of actuator faults within the HMASs increases dramatically, and the fault of a single agent may evolve into the failure of the whole system. Thus, the safety and reliability of HMASs are extremely important. Additionally, the control direction (the symbol of control gain) caused by both CGMs with coupling property of MIMO systems and actuator faults with unpredictable jumps, may not be guaranteed to be known in practical application, such as ship autopilot systems or uncalibrated visual servoing. This will have a considerable influence on the system’s control performance. On account of the threat of nonlinear uncertainties, and unknown control direction yield CGMs and actuator faults to HMASs, an adaptive fault-tolerant control solution based on an effective distributed time-varying auxiliary filter, is developed for MIMO HMASs to guarantee the asymptotical consensus tracking control result. Further, the proposed control scheme has been illustrated to be feasible through simulation experiment conducted on a group of HMASs consisting of UAVs and ASVs. Dahui Luo, Yujuan Wang 0001, Zeqiang Li, Yongduan Song 0001, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Model-Free Inverse H-Infinity Control for Imitation LearningabstractThis paper proposes a data-driven model-free inverse reinforcement learning (IRL) algorithm tailored for solving an inverse$H_{\infty } $control problem. In the problem, both an expert and a learner engage in$H_{\infty } $control to reject disturbances and the learner’s objective is to imitate the expert’s behavior by reconstructing the expert’s performance function through IRL techniques. Introducing zero-sum game principles, we first formulate a model-based single-loop IRL policy iteration algorithm that includes three key steps: updating the policy, action, and performance function using a new correction formula and the standard inverse optimal control principles. Building upon the model-based approach, we propose a model-free single-loop off-policy IRL algorithm that eliminates the need for initial stabilizing policies and prior knowledge of the dynamics of expert and learner. Also, we provide rigorous proof of convergence, stability, and Nash optimality to guarantee the effectiveness and reliability of the proposed algorithms. Furthermore, we showcase the efficiency of our algorithm through simulations and experiments, highlighting its advantages compared to the existing methods.Note to Practitioners—Generally, the cost function for optimal tracking or imitation control is manually defined, which is a challenging task and may result in large tracking errors and slow tracking. In such cases, IRL is a powerful tool for reconstructing proper cost functions. Real-world systems, as demonstrated in practical cases, are frequently exposed to external disturbances and come with unknown models. Employing$H_{\infty } $control is an effective strategy to handle disturbances. However, applying model-free IRL to solve the inverse problem of$H_{\infty } $control for imitation remains an underexplored domain. This paper explores model-free inverse$H_{\infty } $control for imitating expert behaviors, specifically addressing the time-consuming nature of the existing IRL studies that employ a two-loop iteration structure. We propose an efficient single-loop IRL algorithm with a new framework to do this. It is data-driven and model-free, eliminating the need to find an initial stabilizing control policy, which is typically challenging. Additionally, it ensures convergence, stability, and optimality with provable guarantees. Wenqian Xue, Bosen Lian, Yusuf Kartal, Jialu Fan, Tianyou Chai, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Inverse Reinforcement Learning for Discrete-Time Systems With Data DropoutsabstractThis article proposes inverse reinforcement learning (IRL) algorithms for tracking control of linear networked control systems under random state dropouts during wireless transmission. The controlled system aims to track the optimal trajectory of a target system, despite the cost function governing the target's behaviors being unknown. The problem is complicated by random state dropouts occurring in two crucial scenarios: 1) the reception of the target's state and 2) feedback of the controlled system's states. Our approach enables the controlled system to infer the target's cost function and optimal control policy, thereby facilitating effective tracking. Specifically, we develop a model-based IRL algorithm that integrates the Smith predictor for state estimation. Then, we advance a state-dropout-aware inverse Q-learning algorithm that uses solely accessible system data, eliminating the need for system models. The theoretical validity of the proposed algorithms is rigorously established, and their practical effectiveness is validated through numerical simulations. Jialu Fan, Wenqian Xue, Bosen Lian, Yunfang Cui, Frank L. Lewis |
IEEE Trans. Cybern. | 6 |
| 2025 | Dynamic Analysis and Neural-Adaptive Prescribed-Time Control of the FO Memristive Magnetic-Field Electromechanical TransducerabstractThis article is concerned with dynamic analysis and neural-adaptive prescribed-time control of the magnetic-field electromechanical transducer incorporating a memristor. First, a fractional-order (FO) mathematical model is developed, which comprehensively characterizes fractional properties of various dielectrics and establishes the relationship between magnetic flux and electric charge. The dynamical analysis explores internal evolution and complexity performance concerning a single factor or double factors among the FO, system parameter, and memristor configuration by the Bifurcation diagram, sample entropy, and complexity from multiple perspectives. Subsequently, a neural-adaptive prescribed-time control scheme is proposed to transform detrimental chaotic oscillations into orderly motions, achieve the pregiven tracking precision and accommodating both actuator fault and system uncertainty. The controller design consists of three key steps: 1) a deferred constraint function is imposed on the tracking error starting from anywhere to get assignable tracking precision within a specified time, ensuring collision avoidance; 2) a type-2 fuzzy wavelet neural network (FWNN) is utilized effectively to handle parameter perturbations and system uncertainties; and 3) a second-order FO tracking differentiator (TD) is utilized to address the "explosion of complexity" of traditional backstepping under actuator fault model. It is shown that the proposed scheme is able to ensure the boundness of all signals of the closed-loop system. Finally, extensive simulation experiments are conducted to validate the effectiveness and robustness of the rendered scheme. Shaohua Luo, Yongduan Song 0001, Ya Zhang 0001, Hassen M. Ouakad, Frank L. Lewis |
IEEE Trans. Cybern. | 5 |
| 2025 | Perception-Based Feedback Tracking via Scalarized Pixel Data of Visual Frames
Haokun Guo, Murad Abu-Khalaf, Zhiyang Zheng, Junli Gao, Frank L. Lewis, Ci Chen 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Inverse Value Iteration and Q-Learning: Algorithms, Stability, and RobustnessabstractThis article proposes a data-driven model-free inverse Q-learning algorithm for continuous-time linear quadratic regulators (LQRs). Using an agent's trajectories of states and optimal control inputs, the algorithm reconstructs its cost function that captures the same trajectories. This article first poses a model-based inverse value iteration scheme using the agent's system dynamics. Then, an online model-free inverse Q-learning algorithm is developed to recover the agent's cost function only using the demonstrated trajectories. It is more efficient than the existing inverse reinforcement learning (RL) algorithms as it avoids the repetitive RL in inner loops. The proposed algorithms do not need initial stabilizing control policies and solve for unbiased solutions. The proposed algorithm's asymptotic stability, convergence, and robustness are guaranteed. Theoretical analysis and simulation examples show the effectiveness and advantages of the proposed algorithms. Bosen Lian, Wenqian Xue, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Neuroadaptive Control With Enhanced Stability and ReliabilityabstractThe performance of neural network (NN)-driven control systems hinges on the reliability and functionality of the NN unit in the controller. Maintaining the compact set condition for NN training signals (inputs) during operation is crucial for preserving the NN's universal learning and approximation capabilities, yet this requirement is often overlooked in existing studies. This article introduces a constraint transformation-based design method that ensures excitation signals always originate from a fixed region, regardless of initial conditions. By meeting the compactness condition required by the universal approximation theorem, this approach safeguards the functionality of the NN-driven control unit. Additionally, a decaying damping rate is employed to enable the tracking error to asymptotically converge to zero, rather than being ultimately uniformly bounded (UUB). To further ensure robust operation even if the NN underperforms due to an insufficient number of neurons or violation of the compact set condition, a new control strategy is developed based on the worst case behavior of NNs. This "fail-secure" mechanism significantly enhances the reliability of the NN-based control scheme. The effectiveness and benefits of the proposed method are confirmed through numerical simulations, demonstrating its potential to substantially improve the robustness and performance of NN-driven control systems. Kaili Xiang, Ruotong Ming, Siyu Chen 0020, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Output Feedback Secure Control for Fully Quantized Nonlinear Systems Under Irregular Input/Output DoS AttacksabstractThis article investigates the stability of fully quantized nonlinear systems under irregular denial-of-service (IDoS) attacks with unpredictable targets and frequencies. Unlike most existing works that focus only on DoS attacks in fixed or single-type channels, this work considers a more general attack mode, where attacks can occur in the input, output, or both simultaneously. During an attack, the signal in the affected channel is interrupted, rendering it inaccessible. Even if the system is safe, only quantized output signal is available. To address this, a novel state estimator is designed that switches different observer gains depending on the attacked channel. On this basis, we develop an output-feedback adaptive control algorithm that utilizes the estimated signals and incorporates a dynamic filtering technique. The control strategy effectively resists IDoS attacks, ensuring that all closed-loop signals remain semiglobally uniformly ultimately bounded (SUUB), and the stability error converges to a small residual set near the origin. Numerical simulation further confirmed the advantages and effectiveness of the proposed scheme. Yongduan Song 0001, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Output Regulation Based on Zero-Sum Game for Discrete-Time System Driven by Exogenous SignalabstractThis article proposes a novelQ-learning algorithm that relies solely on input-output data to address the output regulation control problem of complex discrete-time systems affected by exogenous signals. Unlike traditional methods, this algorithm does not require detailed system information, state knowledge, or data about external systems or exogenous signals. Additionally, the control strategy does not depend on state information, but on input-output data processed by a set of filters. We provide upper and lower bounds on the discount factor, eliminating the need to solve the Riccati equation. These bounds ensure that the value function remains finite, and we prove the stability of the system when using control inputs derived from the value function with the given discount factor. Furthermore, theQ-learning algorithm, when applied with input data containing probing noise, is shown to yieldQ-function estimates that are independent of the probing noise. Finally, a simulation involving a grid-connected inverter is presented, demonstrating the effectiveness of the proposed algorithm in a practical setting. Ruizhuo Song, Gaofu Yang, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Novel Event-Triggered Control for Time-Varying Leader-Follower MASs on Directed GraphsabstractThe article studies the fully distributed leader–follower and adaptive event-triggered problem with the guarantee of positive minimum interevent times (MIET) for time-varying MAS on directed graphs. First, a novel fully distributed adaptive event-triggered scheme that includes a time-varying matrix gain and two dynamic gains is designed, and the requirement for global topology information can be removed. A novel dynamic triggering mechanism is then put forward for each follower, where an auxiliary dynamic parameter is introduced into the triggering function to guarantee the existence of positive MIETs. Meanwhile, it is worth mentioning that a direct measurement parameter using only the sampling information is leveraged to avoid the use of continuous communication with neighbors. Finally, simulation results are presented to confirm the performance of the proposed method. Lina Xia, Qing Li 0015, Ruizhuo Song, Xiang-Gui Guo, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Error-based adaptive optimal tracking control of nonlinear discrete-time systems
Jinliang Ding, Frank L. Lewis, Tianyou Chai |
Sci. China Inf. Sci. | 3 |
| 2024 | Online Policy Iteration Algorithms for Linear Continuous-Time H-Infinity Regulation With Completely Unknown DynamicsabstractThis paper proposes two online policy-iteration (PI) algorithms for solving linear continuous-time$H_\infty$regulation problems with unknown dynamics. Our results are completely learning-orientated in the sense that prior model knowledge of initial stabilizing control policies arising from solving the Game Algebraic Riccati Equation (GARE) associated with the$H_\infty$regulation problem is now removed, which thereby resolves a long-standing challenge in the existing PI works to achieve model-free learning. To this end, two offline PI algorithms, consisting of the single-looped and the double-looped, are first proposed by nesting a homotopy-based initialization to solve a series of Lyapunov equations associated with the GARE. Then, two online PI algorithms are further proposed by utilizing the system data to avoid the model requirement for online solving the GARE. The single-looped PI algorithm has the feature of simultaneously learning control and disturbance policies, while a double-looped PI updates control policies before carrying out a series of learning disturbance policies. These two online PI algorithms work in a model-free manner and do not require prior knowledge of the system matrices over the whole learning period including the control policy initialization and can lead to the desired control policy with satisfactory system performance. We demonstrate the effectiveness of the proposed learning algorithms with an example of a power systemNote to Practitioners—Solving the$H_\infty$regulation problem for linear continuous-time systems can be achieved by finding the Nash equilibrium of the two-player zero-sum game. However, it is a challenge for control practitioners to design the$H_\infty$regulation controller with completely unknown dynamics due to the fact that it is nontrivial to obtain precise prior knowledge of models/dynamics for many engineering systems. The current methods usually utilized system data to solve the Nash equilibrium solution by offline or online iterative computation, but most of them still needed prior knowledge of the system dynamics for policy seeking such as stabilizing/admissible control or disturbance policies in the initialization. To address such a challenge, this paper develops two homotopy-based online PI algorithms that solve the$H_\infty$regulation problem in a fully model-free manner. It is shown that the developed algorithms can find the Nash equilibrium solution by online measuring the system data, and overcomes the difficulty of finding an initial stabilizing control policy with unknown system dynamics. The validity of the algorithms is illustrated through a simulation study. Ci Chen 0002, Kan Xie 0002, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Low Complexity Distributed Synchronization of Uncertain Nonlinear Multi-Agent Systems With Global Funnel PerformanceabstractThis paper investigates the synchronization control problem for a family of high-order nonlinear multi-agent systems with mismatched and nonparametric uncertainties. Under the directed interconnection topology, a distributed robust control protocol with low complexity is proposed such that the desired performance indexes, such as the fast convergence speed and the small steady-state error, can be pre-specified irrespective of initial conditions, distinguishing itself from the existing related literatures where the prescribed performance is subject to certain initial conditions. Further, no knowledge regarding the bounds of uncertainties and no approximation operators are involved in the control laws. It is shown that all the closed-loop signals are ensured to be globally uniformly bounded. A numerical example is provided to demonstrate the effectiveness of this approach. Zeqiang Li, Yujuan Wang 0001, Yongduan Song 0001, Frank L. Lewis |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Cooperative Control of Multiagent Systems: A Quantization Feedback-Based Event-Triggered ApproachabstractThis article addresses the synchronization tracking problem for high-order uncertain nonlinear multiagent systems via intermittent feedback under a directed graph. By resorting to a novel storer-based triggering transmission strategy in the state channels, we propose an event-triggered neuroadaptive control method with quantitative state feedback that exhibits several salient features: 1) avoiding continuous control updates by making the parameter estimations updated intermittently at the trigger instants; 2) resulting in lower-frequency triggering transmissions by using one event detector to monitor the triggering condition such that each agent only needs to broadcast information at its own trigger times; and 3) saving communication and computation resources by designing the intermittent updating of neural network weights using a dual-phase technique during the triggering period. Besides, it is shown that the proposed scheme is capable of steering the tracking/disagreement errors into an adjustable neighborhood close to the origin, and the existence of a strictly positive dwell time is proved to circumvent Zeno behavior. Both theoretical analysis and numerical simulation authenticate and validate the efficiency of the proposed protocols. Hongwei Cao, Xiucai Huang, Yongduan Song 0001, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2024 | Data-Efficient Reinforcement Learning for Complex Nonlinear SystemsabstractThis article proposes a data-efficient model-free reinforcement learning (RL) algorithm using Koopman operators for complex nonlinear systems. A high-dimensional data-driven optimal control of the nonlinear system is developed by lifting it into the linear system model. We use a data-driven model-based RL framework to derive an off-policy Bellman equation. Building upon this equation, we deduce the data-efficient RL algorithm, which does not need a Koopman-built linear system model. This algorithm preserves dynamic information while reducing the required data for optimal control learning. Numerical and theoretical analyses of the Koopman eigenfunctions for dataset truncation are discussed in the proposed model-free data-efficient RL algorithm. We validate our framework on the excitation control of the power system. Vrushabh S. Donge, Bosen Lian, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 3 |
| 2024 | Performance-Based Distributed Control of Multiagent Systems: A Dual Phase ApproachabstractIn this article, we investigate the distributed tracking control problem for networked uncertain nonlinear strict-feedback systems with unknown time-varying gains under a directed interaction topology. A dual phase performance-guaranteed approach is established. In the first phase, a fully distributed robust filter is constructed for each agent to estimate the desired trajectory with prescribed performance such that the control directions of all agents are allowed to be nonidentical. In the second phase, by establishing a novel lemma regarding Nussbaum function, a new adaptive control protocol is developed for each agent based on backstepping technique, which not only steers the output to track the corresponding estimated signal asymptotically with arbitrarily prescribed transient response but also extends the application scope of the proposed control scheme largely since the unknown control gains are allowed to be time-varying and even state-dependent. In such a way, the underlying problem is tackled with the output tracking error converging into an arbitrarily preassigned residual set exhibiting an arbitrarily predefined convergence rate. Besides, all the internal signals are ensured to be semi-globally ultimately uniformly bounded (SGUUB). Finally, two examples are provided to illustrate the effectiveness of the co-designed scheme. Zeqiang Li, Yujuan Wang 0001, Yongduan Song 0001, Xiucai Huang, Frank L. Lewis |
IEEE Trans. Cybern. | 5 |
| 2024 | Reinforcement Learning for Synchronization of Heterogeneous Multiagent Systems by Improved Q-FunctionsabstractThis article dedicates to investigating a methodology for enhancing adaptability to environmental changes of reinforcement learning (RL) techniques with data efficiency, by which a joint control protocol is learned using only data for multiagent systems (MASs). Thus, all followers are able to synchronize themselves with the leader and minimize their individual performance. To this end, an optimal synchronization problem of heterogeneous MASs is first formulated, and then an arbitration RL mechanism is developed for well addressing key challenges faced by the current RL techniques, that is, insufficient data and environmental changes. In the developed mechanism, an improved Q-function with an arbitration factor is designed for accommodating the fact that control protocols tend to be made by historic experiences and instinctive decision-making, such that the degree of control over agents' behaviors can be adaptively allocated by on-policy and off-policy RL techniques for the optimal multiagent synchronization problem. Finally, an arbitration RL algorithm with critic-only neural networks is proposed, and theoretical analysis and proofs of synchronization and performance optimality are provided. Simulation results verify the effectiveness of the proposed method. Jinna Li, Weiran Cheng, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 5 |
| 2024 | Inverse Q-Learning Using Input-Output DataabstractThis article addresses the problem of learning the objective function of linear discrete-time systems that use static output-feedback (OPFB) control by designing inverse reinforcement learning (RL) algorithms. Most of the existing inverse RL methods require the availability of states and state-feedback control from the expert or demonstrated system. In contrast, this article considers inverse RL in a more general case where the demonstrated system uses static OPFB control with only input-output measurements available. We first develop a model-based inverse RL algorithm to reconstruct an input-output objective function of a demonstrated discrete-time system using its system dynamics and the OPFB gain. This objective function infers the demonstrations and OPFB gain of the demonstrated system. Then, an input-output Q -function is built for the inverse RL problem upon the state reconstruction technique. Given demonstrated inputs and outputs, a data-driven inverse Q -learning algorithm reconstructs the objective function without the knowledge of the demonstrated system dynamics or the OPFB gain. This algorithm yields unbiased solutions even though exploration noises exist. Convergence properties and the nonunique solution nature of the proposed algorithms are studied. Numerical simulation examples verify the effectiveness of the proposed methods. Bosen Lian, Wenqian Xue, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 3 |
| 2024 | Inverse Reinforcement Learning for Trajectory Imitation Using Static Output Feedback ControlabstractThis article studies the trajectory imitation control problem of linear systems suffering external disturbances and develops a data-driven static output feedback (OPFB) control-based inverse reinforcement learning (RL) approach. An Expert-Learner structure is considered where the learner aims to imitate expert's trajectory. Using only measured expert's and learner's own input and output data, the learner computes the policy of the expert by reconstructing its unknown value function weights and thus, imitates its optimally operating trajectory. Three static OPFB inverse RL algorithms are proposed. The first algorithm is a model-based scheme and serves as basis. The second algorithm is a data-driven method using input-state data. The third algorithm is a data-driven method using only input-output data. The stability, convergence, optimality, and robustness are well analyzed. Finally, simulation experiments are conducted to verify the proposed algorithms. Wenqian Xue, Bosen Lian, Jialu Fan, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 5 |
| 2024 | Cooperative Finitely Excited Learning for Dynamical GamesabstractIn this article, we propose a way to enhance the learning framework for zero-sum games with dynamics evolving in continuous time. In contrast to the conventional centralized actor-critic learning, a novel cooperative finitely excited learning approach is developed to combine the online recorded data with instantaneous data for efficiency. By using an experience replay technique for each agent and distributed interaction amongst agents, we are able to replace the classical persistent excitation condition with an easy-to-check cooperative excitation condition. This approach also guarantees the consensus of the distributed actor-critic learning on the solution to the Hamilton-Jacobi-Isaacs (HJI) equation. It is shown that both the closed-loop stability of the equilibrium point and convergence to the Nash equilibrium can be guaranteed. Simulation results demonstrate the efficacy of this approach compared to previous methods. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2024 | Policy Iteration-Based Learning Design for Linear Continuous-Time Systems Under Initial Stabilizing OPFB PolicyabstractPolicy iteration (PI), an iterative method in reinforcement learning, has the merit of interactions with a little-known environment to learn a decision law through policy evaluation and improvement. However, the existing PI-based results for output-feedback (OPFB) continuous-time systems relied heavily on an initial stabilizing full state-feedback (FSFB) policy. It thus raises the question of violating the OPFB principle. This article addresses such a question and establishes the PI under an initial stabilizing OPFB policy. We prove that an off-policy Bellman equation can transform any OPFB policy into an FSFB policy. Based on this transformation property, we revise the traditional PI by appending an additional iteration, which turns out to be efficient in approximating the optimal control under the initial OPFB policy. We show the effectiveness of the proposed learning methods through theoretical analysis and a case study. Ci Chen 0002, Frank L. Lewis, Shengli Xie 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Dynamical Analysis, Circuit Design, and Fuzzy Prescribed Performance Backstepping Control of the FO Weakly Coupled MEMS ResonatorsabstractThis article investigates the dynamical analysis, circuit design, and fuzzy prescribed performance backstepping control of fractional-order (FO) weakly coupled micro-electro-mechanical system (MEMS) resonators with an event-triggered input. The math model of such MEMS resonators coupled by bridge-type coupling beam is constructed based on the Lagrange motion equation and Caputo definition. The dynamical analysis reveals evolution rules and a tendency of system dynamics involved periodic/multiperiodic state and chaotic oscillation for the coupling stiffness, alternating voltage and FO. The designed FO analog circuit and digital circuit based on the field-programmable gate array further validate the mentioned dynamics and are convenient for latter engineering development. To realize stabilization control purposes like chaos suppression, accelerated convergence, high accuracy tracking, performance constraint, and communication resource saving, a fuzzy prescribed performance backstepping controller, in which theβ-cut type-2 fuzzy logic system is used to deal with unknown system function, a transfer function with prescribed performance function is proposed to guarantee the constraint boundary of the tracking error, an accelerated tracking differentiator with a speed function is established to increase convergence rate and solve the “complexity explosion,” an event trigger mechanism is provided to ease the communication burden and a continuous frequency distributed model is utilized to reflect the essence of infinite dimension of such FO system, is constructed in the framework of a backstepping. The proposed scheme not only guarantees that all system signals in closed-loop system are bounded, but also achieves the mentioned stabilization control purposes. Finally, abundant simulation results validate the feasibility of the presented scheme. Shaohua Luo, Yongduan Song 0001, Frank L. Lewis, Guangwei Deng |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | Compensator-Based Self-Learning: Optimal Operational Control for Two-Time-Scale Systems With Input ConstraintsabstractThe practical industrial operation systems are not ideally immune to the effect of unmodeled dynamics and the industrial processes generally are operated at multitime-scales, which cause troubles for optimizing the industrial operation. The novelty of this article is that a self-learning composite compensation control method is developed for two-time-scale optimal operation systems, with well dealing with unmodeled dynamics, unknown operation process and input constraints. First, the two-time scales system is decomposed into fast and slow subsystems based on singular perturbation theory. Then, the critic-only reinforcement learning technique and H$\infty$control are employed for designing the composite controller. Finally, the efficacy is verified by an industrial mixed separation thickening process and a numerical example. Jinna Li, Frank L. Lewis, Meng Zheng 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Integrated Intelligent Guidance and Motion Control of USVs With Anticipatory Collision Avoidance Decision-MakingabstractIn crowded waters, multiple vessel encounter situations increase the collision risks (CRs) of unmanned surface vehicles (USVs) and hence the frequent collision avoidance (COLAV) maneuvers of USVs increase their actual sailing distances. This paper innovatively proposes a risk-prediction-based deep reinforcement learning (RPDRL) approach for the integrated intelligent guidance and motion control of USVs with anticipatory COLAV decision-making. The data sizes of detected vessels’ motion states are different due to the uncertainties in the number of vessels detected by the navigation systems of a USV. To address this problem, these data are, for the first time, converted into the corresponding same-sized raster data as the states in the RPDRL approach. A new CR assessment model of the USV collisions with all the detected vessels is built to calculate the rewards in the RPDRL approach. Furthermore, actor and critic deep convolutional neural networks are created to make the anticipatory COLAV decisions which are the engine command and rudder command. Simulations and simulation comparison results on a USV demonstrate that the USV sails along a shorter path with a lower CR under the anticipatory COLAV decisions from our proposed RPDRL approach compared with a velocity obstacle method and a model predictive control method, and hence the economy and safety of USVs’ autonomous navigation are enhanced. Yihan Tao, Jialu Du, Frank L. Lewis |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Model-Free Q-Learning for the Tracking Problem of Linear Discrete-Time SystemsabstractIn this article, a model-free Q-learning algorithm is proposed to solve the tracking problem of linear discrete-time systems with completely unknown system dynamics. To eliminate tracking errors, a performance index of the Q-learning approach is formulated, which can transform the tracking problem into a regulation one. Compared with the existing adaptive dynamic programming (ADP) methods and Q-learning approaches, the proposed performance index adds a product term composed of a gain matrix and the reference tracking trajectory to the control input quadratic form. In addition, without requiring any prior knowledge of the dynamics of the original controlled system and command generator, the control policy obtained by the proposed approach can be deduced by an iterative technique relying on the online information of the system state, the control input, and the reference tracking trajectory. In each iteration of the proposed method, the desired control input can be updated by the iterative criteria derived from a precondition of the controlled system and the reference tracking trajectory, which ensures that the obtained control policy can eliminate tracking errors in theory. Moreover, to effectively use less data to obtain the optimal control policy, the off-policy approach is introduced into the proposed algorithm. Finally, the effectiveness of the proposed algorithm is verified by a numerical simulation. Jinliang Ding, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Data-Driven Inverse Reinforcement Learning Control for Linear Multiplayer GamesabstractThis article proposes a data-driven inverse reinforcement learning (RL) control algorithm for nonzero-sum multiplayer games in linear continuous-time differential dynamical systems. The inverse RL problem in the games is solved by a learner reconstructing the unknown expert players' cost functions from demonstrated expert's optimal state and control input trajectories. The learner, thus, obtains the same control feedback gains and trajectories as the expert, only using data along system trajectories without knowing system dynamics. This article first proposes a model-based inverse RL policy iteration framework that has: 1) policy evaluation step for reconstructing cost matrices using Lyapunov functions; 2) state-reward weight improvement step using inverse optimal control (IOC); and 3) policy improvement step using optimal control. Based on the model-based policy iteration algorithm, this article further develops an online data-driven off-policy inverse RL algorithm without knowing any knowledge of system dynamics or expert control gains. Rigorous convergence and stability analysis of the algorithms are provided. It shows that the off-policy inverse RL algorithm guarantees unbiased solutions while probing noises are added to satisfy the persistence of excitation (PE) condition. Finally, two different simulation examples validate the effectiveness of the proposed algorithms. Bosen Lian, Vrushabh S. Donge, Frank L. Lewis, Tianyou Chai, Ali Davoudi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Distributed Minmax Strategy for Multiplayer Games: Stability, Robustness, and AlgorithmsabstractThis article studies a distributed minmax strategy for multiplayer games and develops reinforcement learning (RL) algorithms to solve it. The proposed minmax strategy is distributed, in the sense that it finds each player’s optimal control policy without knowing all the other players’ policies. Each player obtains its distributed control policy by solving a distributed algebraic Riccati equation in a multiplayer noncooperative game. This policy is found against the worst policies of all the other players. We guarantee the existence of distributed minmax solutions and study their$\mathcal {L}_{2}$and asymptotic stabilities. Under mild conditions, the resulting minmax control policies are shown to improve robust gain and phase margins of multiplayer systems compared to the standard linear–quadratic regulator controller. Distributed minmax solutions are found using both model-based policy iteration and data-driven off-policy RL algorithms. Simulation examples verify the proposed formulation and its computational efficiency over the nondistributed Nash solutions. Bosen Lian, Vrushabh S. Donge, Wenqian Xue, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Nearly Optimal Control for Mixed Zero-Sum Game Based on Off-Policy Integral Reinforcement LearningabstractIn this article, we solve a class of mixed zero-sum game with unknown dynamic information of nonlinear system. A policy iterative algorithm that adopts integral reinforcement learning (IRL), which does not depend on system information, is proposed to obtain the optimal control of competitor and collaborators. An adaptive update law that combines critic-actor structure with experience replay is proposed. The actor function not only approximates optimal control of every player but also estimates auxiliary control, which does not participate in the actual control process and only exists in theory. The parameters of the actor-critic structure are simultaneously updated. Then, it is proven that the parameter errors of the polynomial approximation are uniformly ultimately bounded. Finally, the effectiveness of the proposed algorithm is verified by two given simulations. Ruizhuo Song, Gaofu Yang, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Data-Based Optimal Synchronization of Heterogeneous Multiagent Systems in Graphical Games via Reinforcement LearningabstractThis article studies the optimal synchronization of linear heterogeneous multiagent systems (MASs) with partial unknown knowledge of the system dynamics. The object is to realize system synchronization as well as minimize the performance index of each agent. A framework of heterogeneous multiagent graphical games is formulated first. In the graphical games, it is proved that the optimal control policy relying on the solution of the Hamilton-Jacobian-Bellmen (HJB) equation is not only in Nash equilibrium, but also the best response to fixed control policies of its neighbors. To solve the optimal control policy and the minimum value of the performance index, a model-based policy iteration (PI) algorithm is proposed. Then, according to the model-based algorithm, a data-based off-policy integral reinforcement learning (IRL) algorithm is put forward to handle the partially unknown system dynamics. Furthermore, a single-critic neural network (NN) structure is used to implement the data-based algorithm. Based on the data collected by the behavior policy of the data-based off-policy algorithm, the gradient descent method is used to train NNs to approach the ideal weights. In addition, it is proved that all the proposed algorithms are convergent, and the weight-tuning law of the single-critic NNs can promote optimal synchronization. Finally, a numerical example is proposed to show the effectiveness of the theoretical analysis. Chunping Xiong, Qian Ma 0001, Jian Guo 0007, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Global Tracking Control With Guaranteed Performance for Nonlinear Systems Driven by Delayed InputsabstractThis article addresses the global prescribed performance tracking control problem for uncertain nonlinear systems with delayed input and unmatched nonlinearity. By introducing a new error transformation consisting of a normalized function and a time-varying scaling function, together with an auxiliary function, we establish a robust state-feedback control scheme capable of handling the unknown and time-varying input delays without imposing any constraints on the initial condition between tracking/virtual error and performance function, leading to a global control solution to the challenging problem. Besides, the unknown delayed input is handled skillfully by the union of Lyapunov–Krasovskii functional and proof by contradiction at the final step. Furthermore, the designed controller is simple in structure and inexpensive in calculation since it does not involve any approximation tools/estimation techniques and avoids iterative computation in backstepping-like methods. It is shown that all internal signals are globally ultimately uniformly bounded, and the tracking/virtual error converges to a preassigned arbitrarily small region at a certain predefined rate. Both theoretical analysis and numerical example confirm the validity of the developed method. Yongduan Song 0001, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Asymptotic Output Tracking With Malfunctioning Actuators and Twisted/Biased FeedbackabstractIt is highly desirable yet challenging to maintain stable system operation in the presence of unexpected malfunctioning sensors and actuators arising from internal component faults and/or external malicious attacks. In this note, we investigate the reliable control problem for a class of nonlinear systems with mismatched modeling uncertainties and unknown control gain matrix as well as abnormal actuating and sensoring units. Such malfunctioning sensors and actuators not only bring about additional modeling uncertainties but also literally pollute the original control influence gain matrix, making the underlying problem further complicated. By using the backstepping-like design procedure, we present an adaptive control solution with relaxed controllability conditions, capable of achieving asymptotic stabilization under severely twisted feedback information due to malicious attacks/sensor failures. Besides a primary actuator, we also propose a strategy to use additional actuators as the backup ones to enhance system survivability, where the actuator replacement automatically and seamlessly takes place from the primary actuator to a backup one, once a severe failure at the primary actuator is detected. Numerical simulation also confirms the effectiveness and benefits of the proposed method. Yongduan Song 0001, Changyun Wen, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Robust Optimal Output Regulation for Nonlinear Systems With Unknown ParametersabstractIn this article, a robust optimal output regulation framework for nonlinear systems with unknown parameters is proposed, in which a suboptimal feedforward–feedback controller is designed. Specifically, a novel internal model principle-based feedforward controller is designed to cope with the adverse effects of unknown parameters. Considering the system performance and control cost, an optimal-feedback controller is then designed via the adaptive dynamic programming method. It is proved that under the proposed suboptimal controller, all the signals of the closed-loop systems remain bounded and the tracking error is arbitrarily small. Furthermore, the predefined performance index is minimized. Finally, two simulation examples are given to verify the effectiveness of the proposed framework. Qian Ma 0001, Frank L. Lewis, Shengyuan Xu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | On the Uniformness of Full-State Error Prescribed Performance for Strict-Feedback SystemsabstractMost existing results on full-state error prescribed performance control for multiple-input multiple-output (MIMO) strict-feedback nonlinear systems typically impose demanding constraining conditions on the initial full-state errors, rendering the performance boundary nonuniform with respective to initial conditions, and consequently tedious offline computations for initial error (especially the initial virtual error) constraint verification is inevitable, which is highly undesirable or even impractical for the design and implementation of the corresponding controls. In this article, we present a novel adaptive control solution that allows the performance uniformness (with respect to the initial condition) and the transient behavior (with respect to error overshoot) to be addressed simultaneously under a unified framework. The key design steps and features include: 1) by constructing a nonlinear transformation based on the time-varying scaling function, the developed performance boundary is uniform to any initial condition; 2) the demanding condition on the initial values of full-state errors in the existing prescribed performance works is removed, allowing the designer more freedom to select design parameters and rendering the solution more user-friendly and less demanding in design and implementation; and 3) by making use of the minimum eigenvalues of the resultant diagonal matrix and imposing a critical negative feedback term in the control design, the stability of closed-loop system is ensured by the developed uniform control strategy. The effectiveness of the proposed approach is verified by simulations. Lianhua Li, Kai Zhao 0004, Yongduan Song 0001, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Structural Analysis of the Stochastic Influence Model for Identifiability and Reduced-Order EstimationabstractThe influence model (IM) is a reduced-order stochastic network model that captures the spatiotemporal dynamics in a network of interactive Markov chains. Identifiability and reduced-order estimation of the IM from observation data are crucial for IM applications. Despite the tractability of IM analysis with its reduced-order representation, the identifiability and estimation of IM are challenging due to the tight coupling of both network and local level interactions. The limited identifiability studies in the literature only apply to homogeneous IMs and existing methods for IM estimation incur high-computational cost. In this article, we solve the identifiability problem by providing succinct if-and-only-if conditions for both the homogeneous and heterogeneous IMs. This is obtained through a structural analysis that establishes a novel connection between the high-order and low-order representations of the IMs. The identifiability analysis further leads to reduced-order parameter estimation algorithms of the homogeneous and heterogeneous IMs with reduced computation. Yan Wan 0001, Chenyuan He, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Optimal operational self-learning control for multi-time scale industrial processes with signal compensations
Jinna Li, Frank L. Lewis |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Consensus of Nonlinear Multiagent Systems With Uncertainties Using Reinforcement Learning Based Sliding Mode ControlabstractThis paper investigates distributed control protocols design for uncertain nonlinear multi-agent systems with the goal of achieving the optimal consensus. The critical challenges encountered when designing the optimal distributed control protocols are mainly caused by the internal coupling of agents, uncertainty and nonlinear dynamics. Communication delay among agents makes overcoming these challenges even more difficult. To this end, a novel sliding mode control design method is developed based on the sliding mode control principle and the reinforcement learning technique. The remarkable highlights of the developed method in this paper include the design of distributed sliding mode controllers and the integrated framework of sliding mode control and reinforcement learning, which bring the outcome of successfully learning the composite distributed control protocols for multi-agent systems. Thus, all agents can successfully eliminate the negative impacts brought by system uncertainties and communication delay among agents, and finally follow the leader with a nearly optimal approach. The reachability of sliding mode surfaces and the optimal consensus are rigorously proven and analyzed. Finally, simulation results illustrate the effectiveness of the developed method. Jinna Li, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Time-Varying Formation of Heterogeneous Multiagent Systems via Reinforcement Learning Subject to Switching TopologiesabstractThis paper investigates the optimal formation control of a heterogeneous multiagent system consisting of multiple quadrotors and ground vehicles via reinforcement learning to achieve the time-varying formation under switching topologies. A distributed observer is firstly constructed to generate references using local information for each vehicle to form time-varying formation and the convergence of the observer under switching topologies is proven. Then, reinforcement learning methods are provided for the heterogeneous vehicle group to realize the optimal tracking control without information of vehicle dynamical model. Simulation tests are given to confirm the effectiveness of the proposed method. Deyuan Liu, Hao Liu 0004, Jinhu Lü 0001, Frank L. Lewis |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Static Output-Feedback H∞ Control Design Procedures for Continuous-Time Systems With Different Levels of Model Knowledgeabstractcontrol design criterion of a continuous-time linear system. The goal is to obtain a static output-feedback controller while the design criterion is formulated with an exponential term, divergent or convergent, depending on the designer's choice. Two offline policy-iteration algorithms are presented first, which form the foundations for a family of online off-policy designs. These algorithms cover all different cases of partial or complete model knowledge and provide the designer with a collection of design alternatives. It is shown that such a design for partial model knowledge can reduce the number of unknown matrices to be solved online. In particular, if the disturbance input matrix of the model is given, off-policy learning can be done with no disturbance excitation. This alternative is useful in situations where a measurable disturbance is not available in the learning phase. The utility of these design procedures is demonstrated for the case of an optimal lane tracking controller of an automated car. Shai A. Arogeti, Frank L. Lewis |
IEEE Trans. Cybern. | 2 |
| 2023 | Anomaly Detection and Correction of Optimizing Autonomous Systems With Inverse Reinforcement LearningabstractThis article considers autonomous systems whose behaviors seek to optimize an objective function. This goes beyond standard applications of condition-based maintenance, which seeks to detect faults or failures in nonoptimizing systems. Normal agents optimize a known accepted objective function, whereas abnormal or misbehaving agents may optimize a renegade objective that does not conform to the accepted one. We provide a unified framework for anomaly detection and correction in optimizing autonomous systems described by differential equations using inverse reinforcement learning (RL). We first define several types of anomalies and false alarms, including noise anomaly, objective function anomaly, intention (control gain) anomaly, abnormal behaviors, noise-anomaly false alarms, and objective false alarms. We then propose model-free inverse RL algorithms to reconstruct the objective functions and intentions for given system behaviors. The inverse RL procedure for anomaly detection and correction has the training phase, detection phase, and correction phase. First, inverse RL in the training phase infers the objective function and intention of the normal behavior system using offline stored data. Second, in the detection phase, inverse RL infers the objective function and intention for online observed test system behaviors using online observation data. They are then compared with that of the nominal system to identify anomalies. Third, correction is executed for the anomalous system to learn the normal objective and intention. Simulations and experiments on a quadrotor unmanned aerial vehicle (UAV) verify the proposed methods. Bosen Lian, Yusuf Kartal, Frank L. Lewis, Dariusz G. Mikulski, Gregory R. Hudas, Yan Wan 0001, Ali Davoudi |
IEEE Trans. Cybern. | 3 |
| 2023 | Neuroadaptive Optimal Fixed-Time Synchronization and its Circuit Realization for Unidirectionally Coupled FO Self-Sustained Electromechanical Seismograph SystemsabstractThis article investigates the neuroadaptive optimal fixed-time synchronization and its circuit realization along with dynamical analysis for unidirectionally coupled fractional-order (FO) self-sustained electromechanical seismograph systems under subharmonic and superharmonic oscillations. The synchronization model of the coupled FO seismograph system is established based on drive and response seismic detectors. The dynamical analysis reveals this coupled system generating transient chaos and homoclinic/heteroclinic oscillations. The test results of the constructed equivalent analog circuit further testify its complex nonlinear dynamics. Then, a neuroadaptive optimal fixed-time synchronization controller integrated with the FO hyperbolic tangent tracking differentiator (HTTD), interval type-2 fuzzy neural network (IT2FNN) with transformation, and prescribed performance function (PPF) together with the constraint condition is developed in the backstepping recursive design. Furthermore, it is proved that all signals of this closed-loop system are bounded, and the tracking errors fall into a trap of the prescribed constraint along with the minimized cost function. Extensive studies confirm the effectiveness of the proposed scheme. Shaohua Luo, Yongduan Song 0001, Frank L. Lewis, Roberto Garrappa |
IEEE Trans. Cybern. | 3 |
| 2023 | Distributed Adaptive Nash Equilibrium Solution for Differential Graphical GamesabstractThis article investigates differential graphical games for linear multiagent systems with a leader on fixed communication graphs. The objective is to make each agent synchronize to the leader and, meanwhile, optimize a performance index, which depends on the control policies of its own and its neighbors. To this end, a distributed adaptive Nash equilibrium solution is proposed for the differential graphical games. This solution, in contrast to the existing ones, is not only Nash but also fully distributed in the sense that each agent only uses local information of its own and its immediate neighbors without using any global information of the communication graph. Moreover, the asymptotic stability and global Nash equilibrium properties are analyzed for the proposed distributed adaptive Nash equilibrium solution. As an illustrative example, the differential graphical game solution is applied to the microgrid secondary control problem to achieve fully distributed voltage synchronization with optimized performance. Yang-Yang Qian, Mushuang Liu, Yan Wan 0001, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 4 |
| 2023 | Dynamic Analysis and Fuzzy Fixed-Time Optimal Synchronization Control of Unidirectionally Coupled FO Permanent Magnet Synchronous Generator SystemabstractThis article focuses on dynamic analysis and the fuzzy fixed-time optimal synchronization control problem of unidirectionally coupled fractional-order (FO) permanent magnet synchronous generator (PMSG) system. The synchronization model between FO master and slave PMSGs with capacitive and resistive couplings is built. The dynamic analysis fully reveals its abundant dynamical behaviors including chaotic oscillations and gives stability/instability boundaries with the designed numerical method. In controller design, the hierarchical type-2 fuzzy neural network (HT2FNN) with a transformation is designed to approximate unknown functions, the fixed-time command filter matched up with the compensating signal is proposed to achieve precise estimate and fast convergence, and a fixed-time preconfigured performance function integrated with a smooth and invertible function is built to realize fixed-time convergence and performance constraint. Then a fuzzy fixed-time optimal synchronization control scheme fusing with the HT2FNN, filter, performance function and optimal control is developed under the FO backstepping theory. The stability analysis proves that all signals of the closed-loop system are bounded along with the cost function being minimized. Finally, numerical simulation results verify the feasibility and advantages of our scheme. Shaohua Luo, Yongduan Song 0001, Frank L. Lewis, Roberto Garrappa, Shaobo Li 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Distributed 3-D Time-Varying Formation Control of Underactuated AUVs With Communication Delays Based on Data-Driven State PredictorabstractCommunication delays are a crucial issue in autonomous underwater vehicle (AUV) formation control. To solve this issue, this article for the first time develops an active communication delay compensation (ACDC) mechanism with an innovatively developed data-driven state predictor (DDSP) of AUVs. The DDSP of each AUV can online estimate the current motion states of its neighbors solely depending on its received delayed motion state information of its neighbors. Incorporating the ACDC, prescribed performance control method, and neural networks into the backstepping approach, this article proposes a distributed 3-D time-varying formation prescribed performance control strategy, where a novel auxiliary dynamic system is created to alleviate the adverse effect of input saturations. The proposed formation control strategy ensures AUVs to maintain the desired 3-D time-varying formation pattern, while achieving the asymptotic stability with respect to formation errors and satisfying the performance constraints simultaneously. Simulations are performed to validate our proposed formation control strategy. Jialu Du, Jian Li 0074, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Inverse Reinforcement Learning for Adversarial Apprentice GamesabstractThis article proposes new inverse reinforcement learning (RL) algorithms to solve our defined Adversarial Apprentice Games for nonlinear learner and expert systems. The games are solved by extracting the unknown cost function of an expert by a learner using demonstrated expert's behaviors. We first develop a model-based inverse RL algorithm that consists of two learning stages: an optimal control learning and a second learning based on inverse optimal control. This algorithm also clarifies the relationships between inverse RL and inverse optimal control. Then, we propose a new model-free integral inverse RL algorithm to reconstruct the unknown expert cost function. The model-free algorithm only needs online demonstration of the expert and learner's trajectory data without knowing system dynamics of either the learner or the expert. These two algorithms are further implemented using neural networks (NNs). In Adversarial Apprentice Games, the learner and the expert are allowed to suffer from different adversarial attacks in the learning process. A two-player zero-sum game is formulated for each of these two agents and is solved as a subproblem for the learner in inverse RL. Furthermore, it is shown that the cost functions that the learner learns to mimic the expert's behavior are stabilizing and not unique. Finally, simulations and comparisons show the effectiveness and the superiority of the proposed algorithms. Bosen Lian, Wenqian Xue, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Local Learning Enabled Iterative Linear Quadratic Regulator for Constrained Trajectory PlanningabstractTrajectory planning is one of the indispensable and critical components in robotics and autonomous systems. As an efficient indirect method to deal with the nonlinear system dynamics in trajectory planning tasks over the unconstrained state and control space, the iterative linear quadratic regulator (iLQR) has demonstrated noteworthy outcomes. In this article, a local-learning-enabled constrained iLQR algorithm is herein presented for trajectory planning based on hybrid dynamic optimization and machine learning. Rather importantly, this algorithm attains the key advantage of circumventing the requirement of system identification, and the trajectory planning task is achieved with a simultaneous refinement of the optimal policy and the neural network system in an iterative framework. The neural network can be designed to represent the local system model with a simple architecture, and thus it leads to a sample-efficient training pipeline. In addition, in this learning paradigm, the constraints of the general form that are typically encountered in trajectory planning tasks are preserved. Several illustrative examples on trajectory planning are scheduled as part of the test itinerary to demonstrate the effectiveness and significance of this work. Jun Ma 0008, Zilong Cheng, Ziyu Lin, Frank L. Lewis, Tong Heng Lee |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Inverse Reinforcement Q-Learning Through Expert Imitation for Discrete-Time SystemsabstractIn inverse reinforcement learning (RL), there are two agents. An expert target agent has a performance cost function and exhibits control and state behaviors to a learner. The learner agent does not know the expert's performance cost function but seeks to reconstruct it by observing the expert's behaviors and tries to imitate these behaviors optimally by its own response. In this article, we formulate an imitation problem where the optimal performance intent of a discrete-time (DT) expert target agent is unknown to a DT Learner agent. Using only the observed expert's behavior trajectory, the learner seeks to determine a cost function that yields the same optimal feedback gain as the expert's, and thus, imitates the optimal response of the expert. We develop an inverse RL approach with a new scheme to solve the behavior imitation problem. The approach consists of a cost function update based on an extension of RL policy iteration and inverse optimal control, and a control policy update based on optimal control. Then, under this scheme, we develop an inverse reinforcement Q-learning algorithm, which is an extension of RL Q-learning. This algorithm does not require any knowledge of agent dynamics. Proofs of stability, convergence, and optimality are given. A key property about the nonunique solution is also shown. Finally, simulation experiments are presented to show the effectiveness of the new approach. Wenqian Xue, Bosen Lian, Jialu Fan, Patrik Kolaric, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Data-Driven H∞ Optimal Output Feedback Control for Linear Discrete-Time Systems Based on Off-Policy Q-Learningabstractstatic OPFB control problem of linear discrete-time (DT) systems. The primary contribution of the proposed algorithms lies in a newly developed OPFB control algorithm form for completely unknown systems. Under the premise of satisfying disturbance attenuation conditions, the conditions for the existence of the optimal OPFB solution are given. The convergence of the proposed Q -learning methods, and the difference and equivalence of two algorithms are rigorously proven. Moreover, considering the effects brought by probing noise for the persistence of excitation (PE), the proposed off-policy Q -learning method has the advantage of being immune to probing noise and avoiding biasedness of solution. Simulation results are presented to verify the effectiveness of the proposed approaches. Li Zhang 0151, Jialu Fan, Wenqian Xue, Victor G. Lopez, Jinna Li, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Online Optimal Event-Triggered H∞ Control for Nonlinear Systems With Constrained State and InputabstractTaking safety and performance into consideration, the state and control input of the actual engineering system are often constrained. For this kind of problem, this article puts forward an online dual event-triggered (ET) adaptive dynamic programming (ADP) optimal control algorithm for a class of nonlinear systems with constrained state and input. First, the original system is transformed into another system through the barrier function, after that, a suitable value function with a nonquadratic utility function is designed to obtain the optimal control pair. In addition, on the premise of the asymptotic stability of the system, the trigger condition is devised, and the intersampling time analysis is proved that the algorithm can avoid the Zeno phenomenon. What is more, the critic, action, and disturbance neural networks (NNs) are trained to approximate value function and control sequences, subsequently, the approximation error is proved to be uniformly ultimately boundedness (UUB). Finally, two comparative experiments based on the robot arm model are simulated to verify that the algorithm can make control policies update only when the system has the requirement and keep satisfactory control effect, which can effectively decrease the number of data transfers and reduce the calculation burden. Ruizhuo Song, Lu Liu 0002, Lina Xia, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Prescribed-Time Control and Its Latest DevelopmentsabstractPrescribed-time (PT) control for nonlinear systems, originated from Song et al., has gained increasing attention among the control community. The salient feature of PT control lies in its ability to achieve system stability within a finite settling time user-assignable in advance irrespective of initial conditions. It is such a unique feature that has enticed many follow-up studies on this technically important area, motivating numerous research advancements. In this article, we provide a comprehensive survey on the recent developments in PT control. Through a concise introduction to the concept of PT control, and a unique taxonomy covering: 1) from robust PT control to adaptive PT control; 2) from PT control for single-input–single-output (SISO) systems to multi-input–multioutput (MIMO) systems; and 3) from PT control for an isolated system to multiagent systems, we present an accessible review of this interesting topic. We highlight key techniques, and fundamental assumptions adopted in various developments as well as some new design ideas. We also discuss several possible future research directions toward PT control. Yongduan Song 0001, Hefu Ye, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | H∞-Based Minimal Energy Adaptive Control With Preset Convergence RateabstractThis work studies the${H}_{\infty }$-based minimal energy control with a preset convergence rate (PCR) problem for a class of disturbed linear time-invariant continuous-time systems with matched external disturbance. This problem aims to design an optimal controller so that the energy of the control input satisfies a predetermined requirement. Moreover, the closed-loop system asymptotic stability with PCR is ensured simultaneously. To deal with this problem, a modified game algebraic Riccati equation (MGARE) is proposed, which is different from the game algebraic Riccati equation in the traditional${H}_{\infty } $control problem due to the state cost being lost. Therefore, a unique positive-definite solution of the MGARE is theoretically analyzed with its existing conditions. In addition, based on this formulation, a novel approach is proposed to solve the actuator magnitude saturation problem with the system dynamics being exactly known. To relax the requirement of the knowledge of system dynamics, a model-free policy iteration approach is proposed to compute the solution of this problem. Finally, the effectiveness of the proposed approaches is verified through two simulation examples. Yi Jiang 0007, Kai Zhang 0004, Jin Wu 0002, Chengxi Zhang, Wenqian Xue, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 7 |
| 2022 | Robustness Analysis of Distributed Kalman Filter for Estimation in Sensor NetworksabstractMotivated by the guaranteed stability margins of linear quadratic regulators (LQRs) and standard Kalman filter (KF) in the frequency domain, this article extends these results to the distributed Kalman-consensus filter (DKCF) for distributed estimation in sensor networks. In particular, we study the robustness margins of DKCF in two cases, one of which is based on the direct target observation while the other uses estimates from neighbor sensors in the network. The loop transfer functions of the two cases are established, and gain and phase margin robustness results are derived for both. The robustness margins of DKCF are improved compared to the single-agent KF. Furthermore, as communication topology varies in sensor networks, graph overall coupling strengths change. We also analyze the correlation between overall coupling strengths and the robustness margins of DKCF. Bosen Lian, Frank L. Lewis, Gary A. Hewer, Katia Estabridis, Tianyou Chai |
IEEE Trans. Cybern. | 2 |
| 2022 | Distributed Kalman Consensus Filter for Estimation With Moving TargetsabstractConsensus-based distributed Kalman filters for estimation with targets have attracted considerable attention. Most of the existing Kalman filters use the average consensus approach, which tends to have a low convergence speed. They also rarely consider the impacts of limited sensing range and target mobility on the information flow topology. In this article, we address these issues by designing a novel distributed Kalman consensus filter (DKCF) with an information-weighted consensus structure for random mobile target estimation in continuous time. A new moving target information-flow topology for the measurement of targets is developed based on the sensors' sensing ranges, targets' random mobility, and local information-weighted neighbors. Novel necessary and sufficient conditions about the convergence of the proposed DKCF are developed. Under these conditions, the estimates of all sensors converge to the consensus values. Simulation and comparative studies show the effectiveness and the superiority of this new DKCF. Bosen Lian, Yan Wan 0001, Ya Zhang 0001, Mushuang Liu, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Cybern. | 5 |
| 2022 | Robust Inverse Q-Learning for Continuous-Time Linear Systems in Adversarial EnvironmentsabstractThis article proposes robust inverse Q -learning algorithms for a learner to mimic an expert's states and control inputs in the imitation learning problem. These two agents have different adversarial disturbances. To do the imitation, the learner must reconstruct the unknown expert cost function. The learner only observes the expert's control inputs and uses inverse Q -learning algorithms to reconstruct the unknown expert cost function. The inverse Q -learning algorithms are robust in that they are independent of the system model and allow for the different cost function parameters and disturbances between two agents. We first propose an offline inverse Q -learning algorithm which consists of two iterative learning loops: 1) an inner Q -learning iteration loop and 2) an outer iteration loop based on inverse optimal control. Then, based on this offline algorithm, we further develop an online inverse Q -learning algorithm such that the learner mimics the expert behaviors online with the real-time observation of the expert control inputs. This online computational method has four functional approximators: a critic approximator, two actor approximators, and a state-reward neural network (NN). It simultaneously approximates the parameters of Q -function and the learner state reward online. Convergence and stability proofs are rigorously studied to guarantee the algorithm performance. Bosen Lian, Wenqian Xue, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Cybern. | 3 |
| 2022 | Event-Driven Off-Policy Reinforcement Learning for Control of Interconnected SystemsabstractIn this article, we introduce a novel approximate optimal decentralized control scheme for uncertain input-affine nonlinear-interconnected systems. In the proposed scheme, we design a controller and an event-triggering mechanism (ETM) at each subsystem to optimize a local performance index and reduce redundant control updates, respectively. To this end, we formulate a noncooperative dynamic game at every subsystem in which we collectively model the interconnection inputs and the event-triggering error as adversarial players that deteriorate the subsystem performance and model the control policy as the performance optimizer, competing against these adversarial players. To obtain a solution to this game, one has to solve the associated Hamilton-Jacobi-Isaac (HJI) equation, which does not have a closed-form solution even when the subsystem dynamics are accurately known. In this context, we introduce an event-driven off-policy integral reinforcement learning (OIRL) approach to learn an approximate solution to this HJI equation using artificial neural networks (NNs). We then use this NN approximated solution to design the control policy and event-triggering threshold at each subsystem. In the learning framework, we guarantee the Zeno-free behavior of the ETMs at each subsystem using the exploration policies. Finally, we derive sufficient conditions to guarantee uniform ultimate bounded regulation of the controlled system states and demonstrate the efficacy of the proposed framework with numerical examples. Vignesh Narayanan, Hamidreza Modares, Sarangapani Jagannathan, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2022 | Inverse Reinforcement Learning in Tracking Control Based on Inverse Optimal ControlabstractThis article provides a novel inverse reinforcement learning (RL) algorithm that learns an unknown performance objective function for tracking control. The algorithm combines three steps: 1) an optimal control update; 2) a gradient descent correction step; and 3) an inverse optimal control (IOC) update. The new algorithm clarifies the relation between inverse RL and IOC. It is shown that the reward weight of an unknown performance objective that generates a target control policy may not be unique. We characterize the set of all weights that generate the same target control policy. We develop a model-based algorithm and, further, two model-free algorithms for systems with unknown model information. Finally, simulation experiments are presented to show the effectiveness of the proposed algorithms. Wenqian Xue, Patrik Kolaric, Jialu Fan, Bosen Lian, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 6 |
| 2022 | Data-Driven Optimal Formation Control for Quadrotor Team With Unknown DynamicsabstractIn this article, the data-driven optimal formation control problem is addressed for a heterogeneous quadrotor team with a virtual leader. Each quadrotor is considered as a highly nonlinear system with six degrees of freedom and the accurate dynamic information of the quadrotor is difficult to obtain in practical applications. An optimal cascade formation controller, including a position controller and an attitude controller, is proposed to track a virtual leader and form a predesigned formation. By using the reinforcement learning (RL) approach, the optimal formation controller is learned from the quadrotor system data without any knowledge of dynamic information of the quadrotors. Simulation results of a heterogeneous multiquadrotor system in a formation flight are given to show the effectiveness of the proposed controllers. Wanbing Zhao, Hao Liu 0004, Frank L. Lewis |
IEEE Trans. Cybern. | 3 |
| 2022 | A Three-Level Game-Theoretic Decision-Making Framework for Autonomous VehiclesabstractIn this paper, a three-level decision-making framework is developed to generate safe and effective decisions for autonomous vehicles (AVs). A key component in this decision framework is a normal-form game to capture the interactions between the ego vehicle and its surrounding vehicles. The payoffs in the normal-form game are designed to capture both safety reward and the reward gained by obeying (or the price paid by violating) “soft” traffic rules, e.g., first-come-first-go. This game formulation enables the ego to 1) make appropriate decisions considering the payoffs and possible actions of its surrounding vehicles, and 2) take intelligent actions in emergencies that may sacrifice some soft traffic rules to ensure safety. Moreover, we introduce parameters in the payoff matrix to tune the ego vehicle’s behavior, e.g., aggressiveness level. A neural network is developed to learn the tuning parameters via supervised learning. In addition, to enable the ego to respond timely to different surrounding vehicles’ driving styles, driving style characterization is incorporated into the payoff design for the normal-form game. Simulation studies are conducted to demonstrate the performance of the developed algorithms in two-vehicle intersection-crossing and lane-changing scenarios. Mushuang Liu, Yan Wan 0001, Frank L. Lewis, Subramanya Nageshrao, Dimitar P. Filev |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Semi-Definite Relaxation-Based ADMM for Cooperative Planning and Control of Connected Autonomous VehiclesabstractThis paper investigates the cooperative planning and control problem for multiple connected autonomous vehicles (CAVs) in different scenarios. In the existing literature, most of the methods suffer from significant problems in computational efficiency. Furthermore, as the optimization problem is nonlinear and nonconvex, it typically poses great difficulty in determining the optimal solution. To address this issue, this work proposes a novel and completely parallel computation framework by leveraging the alternating direction method of multipliers (ADMM). The nonlinear and nonconvex optimization problem in the autonomous driving problem can be divided into two manageable sub-problems; and the resulting sub-problems can be solved by using effective optimization methods in a parallel framework. Here, the differential dynamic programming (DDP) algorithm is capable of addressing the nonlinearity of the system dynamics rather effectively; and the nonconvex coupling constraints with small dimensions can be resolved by invoking the notion of semi-definite relaxation (SDR), which can also be solved in a very short time. Due to the parallel computation and efficient relaxation of nonconvex constraints, our proposed approach effectively realizes real-time implementation; and thus extra assurance of driving safety is provided. In addition, two transportation scenarios for multiple CAVs are used to illustrate the effectiveness and efficiency of the proposed method. Zilong Cheng, Jun Ma 0008, Sunan Huang 0001, Frank L. Lewis, Tong Heng Lee |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Adaptive Interleaved Reinforcement Learning: Robust Stability of Affine Nonlinear Systems With Unknown UncertaintyabstractThis article investigates adaptive robust controller design for discrete-time (DT) affine nonlinear systems using an adaptive dynamic programming. A novel adaptive interleaved reinforcement learning algorithm is developed for finding a robust controller of DT affine nonlinear systems subject to matched or unmatched uncertainties. To this end, the robust control problem is converted into the optimal control problem for nominal systems by selecting an appropriate utility function. The performance evaluation and control policy update combined with neural networks approximation are alternately implemented at each time step for solving a simplified Hamilton-Jacobi-Bellman (HJB) equation such that the uniformly ultimately bounded (UUB) stability of DT affine nonlinear systems can be guaranteed, allowing for all realization of unknown bounded uncertainties. The rigorously theoretical proofs of convergence of the proposed interleaved RL algorithm and UUB stability of uncertain systems are provided. Simulation results are given to verify the effectiveness of the proposed method. Jinna Li, Jinliang Ding, Tianyou Chai, Frank L. Lewis, Sarangapani Jagannathan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Optimal Synchronization of Unidirectionally Coupled FO Chaotic Electromechanical Devices With the Hierarchical Neural NetworkabstractThis article solves the problem of optimal synchronization, which is important but challenging for coupled fractional-order (FO) chaotic electromechanical devices composed of mechanical and electrical oscillators and electromagnetic filed by using a hierarchical neural network structure. The synchronization model of the FO electromechanical devices with capacitive and resistive couplings is built, and the phase diagrams reveal that the dynamic properties are closely related to sets of physical parameters, coupling coefficients, and FOs. To force the slave system to move from its original orbits to the orbits of the master system, an optimal synchronization policy, which includes an adaptive neural feedforward policy and an optimal neural feedback policy, is proposed. The feedforward controller is developed in the framework of FO backstepping integrated with the hierarchical neural network to estimate unknown functions of dynamic system in which the mentioned network has the formula transformation and hierarchical form to reduce the numbers of weights and membership functions. Also, an adaptive dynamic programming (ADP) policy is proposed to address the zero-sum differential game issue in the optimal neural feedback controller in which the hierarchical neural network is designed to yield solutions of the constrained Hamilton-Jacobi-Isaacs (HJI) equation online. The presented scheme not only ensures uniform ultimate boundedness of closed-loop coupled FO chaotic electromechanical devices and realizes optimal synchronization but also achieves a minimum value of cost function. Simulation results further show the validity of the presented scheme. Shaohua Luo, Frank L. Lewis, Yongduan Song 0001, Hassen M. Ouakad |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Directed Graph Clustering Algorithms, Topology, and Weak LinksabstractIn this article, a general approach for directed graph clustering and two new density-based clustering objectives are presented. First, using an equivalence between the clustering objective functions and a trace maximization expression, the directed graph clustering objectives are converted into the corresponding weighted kernel$k$-means problems. Then, a nonspectral algorithm, which covers both the direction and weight information of the directed graphs, is thus proposed. Next, with Rayleigh’s quotient, the upper and lower bounds of clustering objectives are obtained. After that, we introduce a new definition of weak links to characterize the effectiveness of clustering. Finally, illustrative examples are given to demonstrate effectiveness of the results. This article provides a glance at the potential connection between density-based and pattern-based clustering. Compared with other approaches for directed graph clustering, the method proposed in this article naturally avoids the loss of the nonsymmetric edge data because there is no need for any additional symmetrization. Xiao Zhang 0007, Bosen Lian, Frank L. Lewis, Yan Wan 0001, Daizhan Cheng |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | An Optimal Primary Frequency Control Based on Adaptive Dynamic Programming for Islanded Modernized MicrogridsabstractIn many pilot research and development (R&D) microgrid projects, engine-based generators are employed in their power systems, either generating electrical energy or being mixed with the heat and power technology. One of the critical tasks of such engine-based generation units is the frequency regulation in the islanded mode of modernized microgrid (MMG) operation; MMGs are microgrids equipped with advanced controls to address more emerging scenarios in smart grids. For having a stable and reliable MMG, we need to synthesize an optimal, robust, primary frequency controller for the islanded mode of MMG of the future. This task is challenging because of unknown mechanical parameters, occurrence of uncertain disturbances, uncertainty of loads, operating point variations, and the appearance of engine delays, and hence nonminimum phase dynamics. This article presents an innovative primary frequency control for the engine generators regulating the frequency of an islanded MMG in the context of smart grids. The proposed approach is based on an adaptive optimal output-feedback control algorithm using adaptive dynamic programming (ADP). The convergence of algorithms, along with the stability analysis of the closed-loop system, is also shown in this article. Finally, as experimental validation, hardware-in-the-loop (HIL) test results are provided in order to examine the effectiveness of the proposed methodology practically.Note to Practitioners—This article was motivated by the problem of primary frequency controls in modernized microgrids (MMGs) using engine generators, which are still one of the prime sources of regulating frequency in pilot research and development (R&D) microgrid projects. Although MMGs will be integral parts of the smart grid of the future, their primary controls in the islanded mode are not advanced enough and not considering existing theoretical challenges scientifically. Existing approaches to regulate frequency using industrially accepted methods are highly model-based and not optimal. Besides, they are not considering the nonminimum phase dynamics. These dynamics are mainly associated with the engine delays—an inherent issue of mechanical parts—for islanded microgrids. This article suggests a new adaptive optimal output-feedback control approach based on the adaptive dynamic programming (ADP) to the abovementioned problem under consideration. By using the proposed methodology, MMGs can deal with the issues mentioned earlier, which are challenging. The proposed approach is optimally rejecting uncertain disturbances (considering the load uncertainty and operating point variations) and reducing the impacts of nonminimum phase dynamics caused by the engine delay. Based on our currently available hardware-in-the-loop (HIL) device’s capability of modeling power systems’ components in real time, our HIL-based experiments demonstrate that this approach is feasible. Masoud Davari, Weinan Gao, Zhong-Ping Jiang, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | Robust Trajectory Tracking in Satellite Time-Varying Formation FlyingabstractThe robust time-varying formation control problem for a group of satellites is addressed. By the static state feedback control strategy and the disturbance estimation theory, a formation flying controller is proposed for the satellite group to form desired time-varying formation patterns and trajectories, and achieve the satellite attitude consensus. The dynamics of each satellite is subject to nonlinearities, parametric perturbations, and external disturbances. Robustness analysis shows that the trajectory and attitude tracking errors of the global closed-loop control system can converge into a given neighborhood of the origin in a finite time. The numerical simulation results validate the effectiveness and advantages of the proposed formation flying controller. Hao Liu 0004, Frank L. Lewis |
IEEE Trans. Cybern. | 3 |
| 2021 | Discrete-Time Non-Zero-Sum Games With Completely Unknown DynamicsabstractIn this article, off-policy reinforcement learning (RL) algorithm is established to solve the discrete-time N -player nonzero-sum (NZS) games with completely unknown dynamics. The N -coupled generalized algebraic Riccati equations (GARE) are derived, and then policy iteration (PI) algorithm is used to obtain the N -tuple of iterative control and iterative value function. As the system dynamics is necessary in PI algorithm, off-policy RL method is developed for discrete-time N -player NZS games. The off-policy N -coupled Hamilton-Jacobi (HJ) equation is derived based on quadratic value functions. According to the Kronecker product, the N -coupled HJ equation is decomposed into unknown parameter part and the system operation data part, which makes the N -coupled HJ equation solved independent of system dynamics. The least square is used to calculate the iterative value function and N -tuple of iterative control. The existence of Nash equilibrium is proved. The result of the proposed method for discrete-time unknown dynamics NZS games is indicated by the simulation examples. Ruizhuo Song, Qinglai Wei, Huaguang Zhang, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2021 | Accelerated Adaptive Fuzzy Optimal Control of Three Coupled Fractional-Order Chaotic Electromechanical TransducersabstractIn this article, we investigate the issue of the accelerated adaptive fuzzy optimal control of three coupled fractional-order chaotic electromechanical transducers. A small network where every transducer has the nearest-neighbor coupling configuration is used to form the coupled fractional-order chaotic electromechanical transducers. The mathematical model of the coupled electromechanical transducers with nearest-neighbors is established and the dynamical analysis reveals that its behaviors are very sensitive to external excitation and fractional order. In the controller design, the recurrent nonsingleton type-2 sequential fuzzy neural network (RNT2SFNN) with the transformation is designed to estimate unknown functions of dynamics system in the feedforward fuzzy controller, and it is constructed to approximate the critic value and actor control functions by using policy iteration (PI) in the optimal feedback controller. Meanwhile, the speed functions are employed to achieve accelerated convergence within a pregiven finite time and a tracking differentiator is used to solve the explosion of terms associated with traditional backstepping. The whole control strategy consists of a feedforward controller integrating with the RNT2SFNN, tracking differentiator, and speed function in the framework of the backstepping control and a feedback controller fusing with the RNT2SFNN and PI under an actor/critic structure to solve the Hamilton-Jacobi-Bellman equation. The proposed scheme not only guarantees the boundness of all signals and realizes the chaos suppression, synchronization, and accelerated convergence, but also minimizes the cost function. Simulations demonstrate and validate the effectiveness of the proposed scheme. Shaohua Luo, Frank L. Lewis, Yongduan Song 0001, Hassen M. Ouakad |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | Statistical Properties and Airspace Capacity for Unmanned Aerial Vehicle Networks Subject to Sense-and-Avoid Safety ProtocolsabstractRandom mobility models (RMMs) capture the random mobility patterns of mobile agents, and have been widely used as the modeling framework for the evaluation and design of mobile networks. All existing RMMs in the literature assume independent movements of mobile agents, which does not hold for unmanned aircraft systems (UASs). In particular, UASs must maintain a safe separation distance to avoid collision. In this paper, we propose a new modeling framework of random mobility models equipped with physical sense-and-avoid protocols to capture the flexible, variable, and uncertain movement patterns of UASs subject to separation safety constraints. For the random direction (RD) RMM equipped with a commonly used sense-and-avoid (S&A) protocol, named sense-and-stop (S&S), we provide its statistical properties including stationary location distribution and stationary inter-vehicle distance distribution, using the Markov analysis. This study provides knowledge on the impact of S&A protocols to critical UAS networking statistics. In addition, we define collision probabilities and airspace capacity concepts for UASs based on the inter-vehicle distance distribution, and derive their closed-form expressions. This analytical framework mathematically bridges local autonomy with global airspace capacity, and allows the impact analysis of local autonomy configurations for effective UAS airspace capacity management. Mushuang Liu, Yan Wan 0001, Frank L. Lewis, Ella M. Atkins, Dapeng Oliver Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Adaptive NN Distributed Control for Time-Varying Networks of Nonlinear Agents With Antagonistic InteractionsabstractThis article proposes an adaptive neural network (NN) distributed control algorithm for a group of high-order nonlinear agents with nonidentical unknown control directions (UCDs) under signed time-varying topologies. An important lemma on the convergence property is first established for agents with antagonistic time-varying interactions, and then by using Nussbaum-type functions, a new class of NN distributed control algorithms is proposed. If the signed time-varying topologies are cut-balanced and uniformly in time structurally balanced, then convergence is achieved for a group of nonlinear agents. Moreover, the proposed algorithms are adopted to achieve the bipartite consensus of high-order nonlinear agents with nonidentical UCDs under signed graphs, which are uniformly quasi-strongly δ -connected. Finally, simulation examples are given to illustrate the effectiveness of the NN distributed control algorithms. Qingling Wang, Haris E. Psillakis, Changyin Sun 0001, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Off-Policy Reinforcement Learning for Tracking in Continuous-Time Systems on Two Time ScalesabstractThis article applies a singular perturbation theory to solve an optimal linear quadratic tracker problem for a continuous-time two-time-scale process. Previously, singular perturbation was applied for system regulation. It is shown that the two-time-scale tracking problem can be separated into a linear-quadratic tracker (LQT) problem for the slow system and a linear-quadratic regulator (LQR) problem for the fast system. We prove that the solutions to these two reduced-order control problems can approximate the LQT solution of the original control problem. The reduced-order slow LQT and fast LQR control problems are solved by off-policy integral reinforcement learning (IRL) using only measured data from the system. To test the effectiveness of the proposed method, we use an industrial thickening process as a simulation example and compare our method to a method with the known system model and a method without time-scale separation. Wenqian Xue, Jialu Fan, Victor G. Lopez, Yi Jiang 0007, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Robust Formation Control for Cooperative Underactuated Quadrotors via Reinforcement LearningabstractIn this article, the model-free robust formation control problem is addressed for cooperative underactuated quadrotors involving unknown nonlinear dynamics and disturbances. Based on the hierarchical control scheme and the reinforcement learning theory, a robust controller is proposed without knowledge of each quadrotor dynamics, consisting of a distributed observer to estimate the position state of the leader, a position controller to achieve the desired formation, and an attitude controller to control the rotational motion. Simulation results on the multiquadrotor system confirm the effectiveness of the proposed model-free robust formation control method. Wanbing Zhao, Hao Liu 0004, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Nonlinear Stochastic Estimators on the Special Euclidean Group SE(3) Using Uncertain IMU and Vision MeasurementsabstractTwo novel robust nonlinear stochastic full pose (i.e., attitude and position) estimators on the Special Euclidean Group$\mathbb {SE}(3)$are proposed using the available uncertain measurements. The resulting estimators utilize the basic structure of the deterministic pose estimators adopting it to the stochastic sense. The proposed estimators for six degrees of freedom (DOF) pose estimations consider the group velocity vectors to be contaminated with constant bias and Gaussian random noise, unlike nonlinear deterministic pose estimators which disregard the noise component in the estimator derivations. The proposed estimators ensure that the closed-loop error signals are semi-globally uniformly ultimately bounded in mean square. The efficiency and robustness of the proposed estimators are demonstrated by the numerical results which test the estimators against high levels of noise and bias associated with the group velocity and body-frame measurements and large initialization error. Hashim A. Hashim, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | On the Identifiability of the Influence Model for Stochastic Spatiotemporal Spread ProcessesabstractThe influence model is a discrete-time stochastic model that succinctly captures the interactions of a network of interacting Markov chains. The model produces a reduced-order representation of stochastic networks, and can be used to describe and tractably analyze probabilistic spatiotemporal spread dynamics, and hence has found broad usage in network applications, such as social networks, traffic management, and failure cascades in power systems. This article provides sufficient and necessary conditions for the identifiability of the influence model, and also develops estimators for model structures through exploiting the model's special properties. In addition, we analyze conditions for the identifiability of the partially observed influence model (POIM), for which not all of the sites can be measured. We develop an expectation-maximization (EM) algorithm-based estimator for POIMs. Chenyuan He, Yan Wan 0001, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Robust Time-Varying Formation Control for Tail-Sitters in Flight Mode TransitionsabstractThis paper mainly addresses the formation control problem for a group of tail-sitters in transition flight between forward and vertical flight. A robust formation control method is proposed to achieve the aggressive time-varying formation subject to nonlinear dynamics and uncertainties. For each tail-sitter, the proposed control method results in a composite controller that includes a trajectory tracking controller and an attitude controller to achieve the translational and rotational motion control, respectively. It is proven that tracking errors of the proposed global closed-loop system can converge to a given neighborhood around the origin in a finite time. Finally, the simulation studies for multiple tail-sitters to accomplish the time-varying formation in transition flight are presented to show the effectiveness of the proposed control strategy. Deyuan Liu, Hao Liu 0004, Frank L. Lewis, Kimon P. Valavanis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Robust Distributed Formation Controller Design for a Group of Unmanned Underwater VehiclesabstractThe formation of unmanned underwater vehicles (UUVs) has wide potential for applications in various marine activities. This paper studies the robust formation protocol design problem for multiple UUVs, whose dynamics are subject to nonlinearity, parametric uncertainties, and external disturbances. A robust distributed formation control scheme is proposed, which yields a control structure involving a position loop and an attitude loop to govern the translational motion and rotational motion, respectively. Theoretical analysis is given to show the robustness properties of the global closed-loop control system. Simulation results are provided to validate the effectiveness of the proposed formation control method. Hao Liu 0004, Yanhu Wang, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Global Social Cost Minimization With Possibly Nonconvex Objective Functions: An Extremum Seeking-Based ApproachabstractA social cost minimization problem is addressed in this article. In the considered problem, a network of agents work collaboratively to minimize the social cost function, which is defined as the sum of the agents’ local objective functions. The engaged agents are supposed to be equipped with an undirected and connected communication graph. Different from most of the existing works, the social cost function in the considered problem is allowed to be nonconvex and possibly admits local extrema. To avoid local extrema and achieve the global minimization of the social cost function, an extremum-seeking-based approach is proposed by introducing a dynamic average consensus protocol to the sinusoidal-dither-signal-based extremum seeking scheme. The dynamic average consensus protocol is leveraged in the proposed extremum-seeker for information sharing and the sinusoidal probing signal is utilized for information extraction. For the avoidance of local extrema, the amplitude of the dither signal is designed to be adaptive. Through Lyapunov stability analysis, it is shown that the proposed method enables the decision variable to converge to a neighborhood of the global minimum point if the conditions on the network connectivity, the existence of unique global minimum and achievability of the global minimum are satisfied. The theoretical result is verified via simulating a numerical example. Maojiao Ye, Guanghui Wen, Shengyuan Xu 0001, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Local Policy Optimization for Trajectory-Centric Reinforcement LearningabstractThe goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact that global policy optimization for non-linear systems could be a very challenging problem both algorithmically and numerically. However, a lot of robotic manipulation tasks are trajectory-centric, and thus do not require a global model or policy. Due to inaccuracies in the learned model estimates, an open-loop trajectory optimization process mostly results in very poor performance when used on the real system. Motivated by these problems, we try to formulate the problem of trajectory optimization and local policy synthesis as a single optimization problem. It is then solved simultaneously as an instance of nonlinear programming. We provide some results for analysis as well as achieved performance of the proposed technique under some simplifying assumptions. Patrik Kolaric, Devesh K. Jha, Arvind U. Raghunathan, Frank L. Lewis, Mouhacine Benosman, Diego Romeres, Daniel Nikovski |
ICRA | 4 |
| 2020 | Heterogeneous formation control of multiple UAVs with limited-input leader via reinforcement learning
Hao Liu 0004, Qingyao Meng, Fachun Peng, Frank L. Lewis |
Neurocomputing | 4 |
| 2020 | Robust optimal control for a class of nonlinear systems with unknown disturbances based on disturbance observer and policy iteration
Ruizhuo Song, Frank L. Lewis |
Neurocomputing | 2 |
| 2020 | Optimal Output Regulation of Linear Discrete-Time Systems With Unknown Dynamics Using Reinforcement LearningabstractThis paper presents a model-free optimal approach based on reinforcement learning for solving the output regulation problem for discrete-time systems under disturbances. This problem is first broken down into two optimization problems: 1) a constrained static optimization problem is established to find the solution to the output regulator equations (i.e., the feedforward control input) and 2) a dynamic optimization problem is established to find the optimal feedback control input. Solving these optimization problems requires the knowledge of the system dynamics. To obviate this requirement, a model-free off-policy algorithm is presented to find the solution to the dynamic optimization problem using only measured data. Then, based on the solution to the dynamic optimization problem, a model-free approach is provided for the static optimization problem. It is shown that the proposed algorithm is insensitive to the probing noise added to the control input for satisfying the persistence of excitation condition. Simulation results are provided to verify the effectiveness of the proposed approach. Yi Jiang 0007, Bahare Kiumarsi-Khomartash, Jialu Fan, Tianyou Chai, Jinna Li, Frank L. Lewis |
IEEE Trans. Cybern. | 6 |
| 2020 | Nonzero-Sum Game Reinforcement Learning for Performance Optimization in Large-Scale Industrial ProcessesabstractThis article presents a novel technique to achieve plant-wide performance optimization for large-scale unknown industrial processes by integrating the reinforcement learning method with the multiagent game theory. A main advantage of this technique is that plant-wide optimal performance is achieved by a distributed approach where multiple agents solve simplified local nonzero-sum optimization problems so that a global Nash equilibrium is reached. To this end, first, the plant-wide performance optimization problem is reformulated by decomposition into local optimization subproblems for each production index in a multiagent framework. Then, the nonzero-sum graphical game theory is utilized to compute the operational indices for each unit process with the purpose of reaching the global Nash equilibrium, resulting in production indices following their prescribed target values. The stability and the global Nash equilibrium of this multiagent graphical game solution are rigorously proved. The reinforcement learning methods are then developed for each agent to solve the nonzero-sum graphical game problem using data measurements available in the system in real time. The plant dynamics do not have to be known. Finally, the emulation results are given to show the effectiveness of the proposed automated decision algorithm by using measured data from a large mineral processing plant in Gansu Province, China. Jinna Li, Jinliang Ding, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2020 | Robust Formation Control for Multiple Quadrotors With Nonlinearities and DisturbancesabstractIn this paper, the robust formation control problem is investigated for a group of quadrotors. Each quadrotor dynamics exhibits the features of underactuation, high nonlinearities and couplings, and disturbances in both the translational and rotational motions. A distributed robust controller is developed, which consists of a position controller to govern the translational motion for the desired formation and an attitude controller to control the rotational motion of each quadrotor. Theoretical analysis and simulation studies of a formation of multiple uncertain quadrotors are presented to validate the effectiveness of the proposed formation control scheme. Hao Liu 0004, Teng Ma 0005, Frank L. Lewis, Yan Wan 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | Resilient and Robust Synchronization of Multiagent Systems Under Attacks on Sensors and ActuatorsabstractResilient and robust distributed control protocols for multiagent systems under attacks on sensors and actuators are designed. A distributed H∞control protocol is designed to attenuate the disturbance or attack effects. However, the H∞controller is too conservative in the presence of attacks. Therefore, it is augmented with a distributed adaptive compensator to mitigate the adverse effects of attacks. The proposed controller can make the synchronization error arbitrarily small in the presence of faulty attacks, and satisfy global L2-gain performance in the presence of malicious attacks or disturbances. A significant advantage of the proposed method is that it requires no restriction on the number of agents or agents' neighbors under attacks on sensors and/or actuators, and it recovers even compromised agents under attacks on actuators. Simulation examples verify the effectiveness of the proposed method. Hamidreza Modares, Bahare Kiumarsi-Khomartash, Frank L. Lewis, Frank T. Ferrese, Ali Davoudi |
IEEE Trans. Cybern. | 3 |
| 2020 | Lagrange Stability and Finite-Time Stabilization of Fuzzy Memristive Neural Networks With Hybrid Time-Varying DelaysabstractThis paper focuses on Lagrange exponential stability and finite-time stabilization of Takagi-Sugeno (T-S) fuzzy memristive neural networks with discrete and distributed time-varying delays (DFMNNs). By resorting to theories of differential inclusions and the comparison strategy, an algebraic condition is developed to confirm Lagrange exponential stability of the underlying DFMNNs in Filippov's sense, and the exponentially attractive set is estimated. When external input is not considered, global exponential stability of DFMNNs is derived directly, which includes some existing ones as special cases. Furthermore, finite-time stabilization of the addressed DFMNNs is analyzed by exploiting inequality techniques and the comparison approach via designing a nonlinear state feedback controller. The boundedness assumption of activation functions is removed herein. Finally, two simulations are presented to demonstrate the validness of the outcomes, and an application is performed in pseudorandom number generation. Yin Sheng, Frank L. Lewis, Zhigang Zeng, Tingwen Huang |
IEEE Trans. Cybern. | 2 |
| 2020 | Discrete-Time Impulsive Adaptive Dynamic ProgrammingabstractIn this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal impulsive control problems for infinite horizon discrete-time nonlinear systems. Considering the constraint of the impulsive interval, in each iteration, the iterative impulsive value function under each possible impulsive interval is obtained, and then the iterative value function and iterative control law are achieved. A new convergence analysis method is developed which proves an iterative value function to converge to the optimum as the iteration index increases to infinity. The properties of the iterative control law are analyzed, and the detailed implementation of the optimal impulsive control law is presented. Finally, two simulation examples with comparisons are given to show the effectiveness of the developed method. Qinglai Wei, Ruizhuo Song, Zehua Liao, Benkai Li, Frank L. Lewis |
IEEE Trans. Cybern. | 5 |
| 2020 | Robust Fault-Tolerant Formation Control for Tail-Sitters in Aggressive Flight Mode TransitionsabstractIn this paper, the fault-tolerant time-varying formation control problem for a group of tail-sitters with multiple actuator faults and uncertainties is studied. A robust distributed fault-tolerant formation control strategy is developed to achieve aggressive time-varying formation flying in flight mode transitions. For each tail-sitter, the designed controller can be divided into an inner attitude controller and an outer position controller to govern the rotational and translational motions, respectively. The information of the actuator faults does not need to be identified online and the tracking errors of the global closed-loop control system can converge into a given neighborhood of the origin in a finite time. Simulation results are presented to show the effectiveness of the proposed control strategy. Deyuan Liu, Hao Liu 0004, Frank L. Lewis, Yan Wan 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | New Methods for Optimal Operational Control of Industrial Processes Using Reinforcement Learning on Two Time ScalesabstractCurrent challenges in industrial processes control include achieving optimum operation for systems with two-time-scale dynamics and unknown models. This paper presents, for the first time, the integration of singular perturbation theory and reinforcement learning to solve this problem. To this end, an optimal operational control (OOC) problem with two time scales is formulated to reach the desired operational indices. Then, a singularly perturbed dynamics for two-time-scale industrial operational processes is developed by introducing a perturbed scale, resulting in the separation of the original system dynamics. Thus, the original optimization problem is decomposed into a reduced slow subproblem and a boundary fast subproblem. The fact that the sum of the separate solutions of these subproblems is approximately equal to the solution of the OOC problem is proven. Then, two Q-learning algorithms are proposed to obtain a composite feedback control. Finally, an industrial thickener example is employed to show the effectiveness of the proposed method. Wenqian Xue, Jialu Fan, Victor G. Lopez, Jinna Li, Yi Jiang 0007, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 7 |
| 2020 | Adaptive Optimal Control for Stochastic Multiplayer Differential Games Using On-Policy and Off-Policy Reinforcement LearningabstractControl-theoretic differential games have been used to solve optimal control problems in multiplayer systems. Most existing studies on differential games either assume deterministic dynamics or dynamics corrupted with additive noise. In realistic environments, multidimensional environmental uncertainties often modulate system dynamics in a more complicated fashion. In this article, we study stochastic multiplayer differential games, where the players' dynamics are modulated by randomly time-varying parameters. We first formulate two differential games for systems of general uncertain linear dynamics, including the two-player zero-sum and multiplayer nonzero-sum games. We then show that optimal control policies, which constitute the Nash equilibrium solutions, can be derived from the corresponding Hamiltonian functions. Stability is proven using the Lyapunov type of analysis. In order to solve the stochastic differential games online, we integrate reinforcement learning (RL) and an effective uncertainty sampling method called the multivariate probabilistic collocation method (MPCM). Two learning algorithms, including the on-policy integral RL (IRL) and off-policy IRL, are designed for the formulated games, respectively. We show that the proposed learning algorithms can effectively find the Nash equilibrium solutions for the stochastic multiplayer differential games. Mushuang Liu, Yan Wan 0001, Frank L. Lewis, Victor G. Lopez |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | H∞ Static Output-Feedback Control Design for Discrete-Time Systems Using Reinforcement LearningabstractThis paper provides necessary and sufficient conditions for the existence of the static output-feedback (OPFB) solution to the H∞control problem for linear discrete-time systems. It is shown that the solution of the static OPFB H∞control is a Nash equilibrium point. Furthermore, a Q-learning algorithm is developed to find the H∞OPFB solution online using data measured along the system trajectories and without knowing the system matrices. This is achieved by solving a game algebraic Riccati equation online and using the measured data. A simulation example shows the effectiveness of the proposed method. Amir Parviz Valadbeigi, Ali Khaki-Sedigh, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Model-Free Online Neuroadaptive Controller With Intent Estimation for Physical Human-Robot InteractionabstractWith the rise of collaborative robots, the need for safe, reliable, and efficient physical human-robot interaction (pHRI) has grown. High-performance pHRI requires robust and stable controllers suitable for multiple degrees of freedom (DoF) and highly nonlinear robots. In this article, we describe a cascade-loop pHRI controller, which relies on human force and pose measurements and can adapt to varying robot dynamics online. It can also adapt to different users and simplifies the interaction by making the robot behave according to a prescribed dynamic model. In our controller formulation, two neural networks (NNs) in the “outer-loop” predict human motion intent and estimate a reference trajectory for the robot that the “inner-loop” controller follows. The inner-loop imposes a prescribed error dynamics (PED) with the help of a model-free neuroadaptive controller (NAC), which uses a NN to feedback linearize the robot dynamics. Lyapunov stability analysis gives weight tuning laws that guarantee that the error signals are bounded and the desired reference trajectory is achieved. Our control scheme was implemented on a Personal Robot 2 robot and validated through an exploratory experimental study in point-to-point collaborative motion. Results indicate fast convergence of our controller, and the resulting tracking error, motion jerk, and human control effort are comparable with other methods that require prior training, knowledge, and calibration. Sven Cremer, Sumit K. Das, Indika Wijayasinghe, Dan O. Popa, Frank L. Lewis |
IEEE Trans. Robotics | 5 |
| 2020 | Model-Free Optimal Output Regulation for Linear Discrete-Time Lossy Networked Control SystemsabstractIn this article, a new model-free approach is proposed to solve the output regulation problem for networked control systems, where the system state can be lost in the feedback process. The goal of the output regulation is to design a control law that can make the system achieve asymptotic stability of the tracking error while maintaining the stability of the closed-loop system. The solvability of the output regulation problem depends on the solvability of a set of matrix equations called the regulator equations. First, a restructured dynamic system is established by using the Smith predictor; then, an off-policy algorithm based on reinforcement learning is developed to calculate the feedback gain using only the measured data when dropout occurs. Based on the solution to the feedback gain, a model-free solution is provided for solving the forward gain using the regulator equations. The simulation results demonstrate the effectiveness of the proposed approach for discrete-time networked systems with unknown dynamics and dropout. Jialu Fan, Yi Jiang 0007, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | Editorial Special Issue on Adaptive Dynamic Programming and Reinforcement LearningabstractThe past decade has witnessed a surge in research activities related to adaptive dynamic programming (ADP) and reinforcement learning (RL), particularly for control applications. Several books [item 1)–5) in the Appendix] and survey papers [item 6)–10) in the Appendix] have been published on the subject. Both ADP and RL provide approximate solutions to dynamic programming problems. In a 1995 article by Bartoet al.[item 11) in the Appendix], they introduced the so-called “adaptive real-time dynamic programming,” which was specifically to apply ADP for real-time control. Later, in 2002, Murrayet al.[item 12) in the Appendix] developed an ADP algorithm for optimal control of continuous-time affine nonlinear systems. On the other hand, the most famous algorithms in RL are the temporal difference algorithm [item 13) in the Appendix] and the Q-learning algorithm [item 14) and 15) in the Appendix]. Derong Liu 0001, Frank L. Lewis, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Robust Optimal Control for Disturbed Nonlinear Zero-Sum Differential Games Based on Single NN and Least SquaresabstractThis paper establishes an approximate optimal critic learning algorithm based on single neural network (NN) policy iteration (PI) aiming at solving for continuous-time (CT) 2-player zero-sum games (ZSGs). In fact, we have to face the problem that the errors will disturb the dynamics and in turn identifying dynamics will generate errors. In order to prevent the effect of errors, in this paper, a single NN-based online PI algorithm is developed for the CT system, which is disturbed nonlinear ZSG. With plenty of online data, the Hamilton-Jacobi-Isaacs equation can be solved without complete dynamics. Then by the least-squares method, we can obtain the NN weights. Moreover, in the process of dealing with the undisturbed system, we find the way that obtains NN weights in this paper is equal to the way that obtains the optimal solution by the Gauss-Newton method. Based on the convergence of the Gauss-Newton method, we can efficiently obtain the optimal controller for the undisturbed system by utilizing online data. After getting the controller of the undisturbed system, it is time to take disturbance into consideration, so that we design a robust control pair to overcome the disturbance. In order to demonstrate the effectiveness of this algorithm, we design a set of simulations. The results verify that we can solve the disturbed nonlinear ZSG by this algorithm. Ruizhuo Song, Junsong Li, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Learning and Uncertainty-Exploited Directional Antenna Control for Robust Aerial NetworkingabstractAerial communication using directional antennas (ACDA) is a promising solution to enable long-distance and broad-band unmanned aerial vehicle (UAV)-to-UAV communication. The automatic alignment of directional antennas allows transmission energy to focus in certain direction and hence significantly extends communication range and rejects interference. In this paper, we develop reinforcement learning (RL)-based on-line directional antennas control solutions for the ACDA system. The novel stochastic optimal control algorithm integrates RL, an effective uncertainty evaluation method called multivariate probabilistic collocation method (MPCM), and unscented Kalman Filter (UKF) for the nonlinear random switching dynamics. Simulation studies are conducted to illustrate and validate the proposed solutions. Mushuang Liu, Yan Wan 0001, Songwei Li 0003, Frank L. Lewis |
VTC Fall | 4 |
| 2019 | Optimal control using adaptive resonance theory and Q-learning
Bahare Kiumarsi-Khomartash, Bakur AlQaudi, Hamidreza Modares, Frank L. Lewis, Daniel S. Levine 0001 |
Neurocomputing | 4 |
| 2019 | Adaptive Compensation for Nonlinear Time-Varying Multiagent Systems With Actuator Failures and Unknown Control DirectionsabstractThis paper investigates a problem of designing an adaptive asymptotic cooperative control scheme for nonlinear time-varying multiagent systems, which can simultaneously tolerate unknown actuator failures and unknown control directions. To address such the problem, we propose a conditional inequality, which allows multiple piecewise Nussbaum functions to acquire the control robustness. Benefiting from this robustness, a part of failure uncertainties and system errors are compensated for, while the remaining parts are handled by adaptive control technique. Moreover, structural properties of the proposed adaptive laws are utilized so that Barbalat's lemma is applicable to make all the followers asymptotically converge to the leader based on the neighborhood information. Kan Xie 0002, Ci Chen 0002, Frank L. Lewis, Shengli Xie 0001 |
IEEE Trans. Cybern. | 3 |
| 2019 | Stability and Stabilization of Takagi-Sugeno Fuzzy Systems With Hybrid Time-Varying DelaysabstractThis paper investigates the stability and stabilization of Takagi-Sugeno (T-S) fuzzy systems with discrete and distributed time-varying delays. First, pth moment global exponential stability (p≥1) of the addressed delayed fuzzy systems is considered by virtue of the comparison approach and inequality techniques. The developed algebraic criteria include some existing outcomes as special cases. Second, global exponential stabilization of the underlying delayed fuzzy systems is performed under a fuzzy state feedback controller. Third, considering that only a few studies have been concerned with finite-time stabilization of T-S fuzzy systems, by employing the comparison strategy and a nonlinear controller, finite-time stabilization of the nominated delayed fuzzy systems is presented. The result obtained herein establishes a general theoretical framework to analyze the finite-time behavior of delayed T-S fuzzy systems. Finally, simulation examples are conducted to illustrate the validity of the results. Yin Sheng, Frank L. Lewis, Zhigang Zeng, Tingwen Huang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2019 | Operational Control of Mineral Grinding Processes Using Adaptive Dynamic Programming and Reference GovernorabstractOperation performance of mineral grinding processes is measured by the grinding product particle size and the circulating load, as two of the most crucial operational indices that measure the product quality and operation efficiency, respectively. In this paper, a data-driven method is proposed for the operational control design of mineral grinding processes with input constraints. A reference governor is introduced to take into account the input constraints and the infeasible setpoint issue. The reference governor generates feasible setpoints that keep control inputs within allowed regions. The lookup table embedded in the reference governor mapping steady-state outputs to inputs provides feasible setpoints for output regulation and baseline for inputs. An ad hoc optimization guarantees that the input constraints are not violated, with the priority of regulating the grinding product particle size if regulation of both indices is not feasible. Since the dynamic model of the controlled plant is complicated because of the strongly nonlinear and intricately coupled nature of ball mills and hydrocyclones, a novel policy iteration algorithm is proposed for optimal regulator design without system modeling. Simulation results comparing performances of a mineral grinding process with and without the reference governor show the effectiveness of the proposed method. Xinglong Lu, Bahare Kiumarsi-Khomartash, Tianyou Chai, Yi Jiang 0007, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Multistability of Delayed Hybrid Impulsive Neural Networks With Application to Associative MemoriesabstractThe important topic of multistability of continuous-and discrete-time neural network (NN) models has been investigated rather extensively. Concerning the design of associative memories, multistability of delayed hybrid NNs is studied in this paper with an emphasis on the impulse effects. Arising from the spiking phenomenon in biological networks, impulsive NNs provide an efficient model for synaptic interconnections among neurons. Using state-space decomposition, the coexistence of multiple equilibria of hybrid impulsive NNs is analyzed. Multistability criteria are then established regrading delayed hybrid impulsive neurodynamics, for which both the impulse effects on the convergence rate and the basins of attraction of the equilibria are discussed. Illustrative examples are given to verify the theoretical results and demonstrate an application to the design of associative memories. It is shown by an experimental example that delayed hybrid impulsive NNs have the advantages of high storage capacity and high fault tolerance when used for associative memories. Bin Hu 0008, Zhi-Hong Guan, Guanrong Chen, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Off-Policy Interleaved $Q$ -Learning: Optimal Control for Affine Nonlinear Discrete-Time SystemsabstractIn this paper, a novel off-policy interleaved Q-learning algorithm is presented for solving optimal control problem of affine nonlinear discrete-time (DT) systems, using only the measured data along the system trajectories. Affine nonlinear feature of systems, unknown dynamics, and off-policy learning approach pose tremendous challenges on approximating optimal controllers. To this end, on-policy Q-learning method for optimal control of affine nonlinear DT systems is reviewed first, and its convergence is rigorously proven. The bias of solution to Q-function-based Bellman equation caused by adding probing noises to systems for satisfying persistent excitation is also analyzed when using on-policy Q-learning approach. Then, a behavior control policy is introduced followed by proposing an off-policy Q-learning algorithm. Meanwhile, the convergence of algorithm and no bias of solution to optimal control problem when adding probing noise to systems are investigated. Third, three neural networks run by the interleaved Q-learning approach in the actor-critic framework. Thus, a novel off-policy interleaved Q-learning algorithm is derived, and its convergence is proven. Simulation results are given to verify the effectiveness of the proposed method. Jinna Li, Tianyou Chai, Frank L. Lewis, Zhengtao Ding, Yi Jiang 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Exponential Stabilization of Fuzzy Memristive Neural Networks With Hybrid Unbounded Time-Varying DelaysabstractThis paper is concerned with exponential stabilization for a class of Takagi-Sugeno fuzzy memristive neural networks (FMNNs) with unbounded discrete and distributed time-varying delays. Under the framework of Filippov solutions, algebraic criteria are established to guarantee exponential stabilization of the addressed FMNNs with hybrid unbounded time delays via designing a fuzzy state feedback controller by exploiting inequality techniques, calculus theorems, and theories of fuzzy sets. The obtained results in this paper enhance and generalize some existing ones. Meanwhile, a general theoretical framework is proposed to investigate the dynamical behaviors of various neural networks with mixed infinite time delays. Finally, two simulation examples are performed to illustrate the validity of the derived outcomes. Yin Sheng, Frank L. Lewis, Zhigang Zeng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Model-Free Reinforcement Learning for Fully Cooperative Multi-Agent Graphical GamesabstractIn this paper, the optimal coordinated control problem for the homogeneous multi-agent graphical games with completely unknown dynamics is investigated. The off-policy reinforcement learning is proposed to approach the solution of the Hamilton-Jacobi equation under the framework of centralized training and decentralized execution. The actor-critic structure is adopted to learn the optimal control policies. Note that the critic network is centralized using the information from all the agents, and the parameter sharing scheme is adopted for the single actor network during the training process. For the execution process, the centralized critic network is not required, and only the trained actor network is used for each agent to obtain the control input based on its individual observation. For the implementation purpose, the neural network approximators with the actor-critic structure are constructed to approach the optimal centralized value function and the optimal policies for the multiagent graphical games. Finally, a simulation example is provided to demonstrate the effectiveness of the proposed algorithm. Dongbin Zhao, Frank L. Lewis |
IJCNN | 3 |
| 2018 | Actor-Critic Off-Policy Learning for Optimal Control of Multiple-Model Discrete-Time SystemsabstractIn this paper, motivated by human neurocognitive experiments, a model-free off-policy reinforcement learning algorithm is developed to solve the optimal tracking control of multiple-model linear discrete-time systems. First, an adaptive self-organizing map neural network is used to determine the system behavior from measured data and to assign a responsibility signal to each of system possible behaviors. A new model is added if a sudden change of system behavior is detected from the measured data and the behavior has not been previously detected. A value function is represented by partially weighted value functions. Then, the off-policy iteration algorithm is generalized to multiple-model learning to find a solution without any knowledge about the system dynamics or reference trajectory dynamics. The off-policy approach helps to increase data efficiency and speed of tuning since a stream of experiences obtained from executing a behavior policy is reused to update several value functions corresponding to different learning policies sequentially. Two numerical examples serve as a demonstration of the off-policy algorithm performance. Jan Skach, Bahare Kiumarsi-Khomartash, Frank L. Lewis, Ondrej Straka |
IEEE Trans. Cybern. | 3 |
| 2018 | Optimal Robust Output Containment of Unknown Heterogeneous Multiagent System Using Off-Policy Reinforcement LearningabstractThis paper investigates optimal robust output containment problem of general linear heterogeneous multiagent systems (MAS) with completely unknown dynamics. A model-based algorithm using offline policy iteration (PI) is first developed, where the -copy internal model principle is utilized to address the system parameter variations. This offline PI algorithm requires the nominal model of each agent, which may not be available in most real-world applications. To address this issue, a discounted performance function is introduced to express the optimal robust output containment problem as an optimal output-feedback design problem with bounded -gain. To solve this problem online in real time, a Bellman equation is first developed to evaluate a certain control policy and find the updated control policies, simultaneously, using only the state/output information measured online. Then, using this Bellman equation, a model-free off-policy integral reinforcement learning algorithm is proposed to solve the optimal robust output containment problem of heterogeneous MAS, in real time, without requiring any knowledge of the system dynamics. Simulation results are provided to verify the effectiveness of the proposed method. Shan Zuo, Yongduan Song 0001, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 3 |
| 2018 | Data-Driven Flotation Industrial Process Operational Optimal Control Based on Reinforcement LearningabstractThis paper studies the operational optimal control problem for the industrial flotation process, a key component in the mineral processing concentrator line. A new model-free data-driven method is developed here for real-time solution of this problem. A novel formulation is given for the optimal selection of the process control inputs that guarantees optimal tracking of the operational indices while maintaining the inputs within specified bounds. Proper tracking of prescribed operational indices, namely concentrate grade and tail grade, is essential in the proper economic operation of the flotation process. The difficulty in establishing an accurate mathematic model is overcome, and optimal controls are learned online in real time, using a novel form of reinforcement learning we call interleaved learning for online computation of the operational optimal control solution. Simulation experiments are provided to verify the effectiveness of the proposed interleaved learning method and to show that it performs significantly better than standard policy iteration and value iteration. Yi Jiang 0007, Jialu Fan, Tianyou Chai, Jinna Li, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | Tracking Control for Linear Discrete-Time Networked Control Systems With Unknown Dynamics and DropoutabstractThis paper develops a new method for solving the optimal control tracking problem for networked control systems (NCSs), where network-induced dropout can occur and the system dynamics are unknown. First, a novel dropout Smith predictor is designed to predict the current state based on historical data measurements over the communication network. Then, it is shown that the quadratic form of the performance index is preserved even with dropout, and the optimal tracker solution with dropout is given based on a novel dropout generalized algebraic Riccati equation. New algorithms for off-line policy iteration (PI), online PI, and Q-learning PI are presented for NCS with dropout. The Q-learning algorithm adaptively learns the optimal control online using data measured over the communication network based on reinforcement learning, including dropout, without requiring any knowledge of the system dynamics. Simulation results are provided to show that the proposed approaches give proper optimal tracking performance for the NCS with unknown dynamics and dropout. Yi Jiang 0007, Jialu Fan, Tianyou Chai, Frank L. Lewis, Jinna Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Optimal and Autonomous Control Using Reinforcement Learning: A SurveyabstractThis paper reviews the current state of the art on reinforcement learning (RL)-based feedback control solutions to optimal regulation and tracking of single and multiagent systems. Existing RL solutions to both optimal and control problems, as well as graphical games, will be reviewed. RL methods learn the solution to optimal control and game problems online and using measured data along the system trajectories. We discuss Q-learning and the integral RL algorithm as core algorithms for discrete-time (DT) and continuous-time (CT) systems, respectively. Moreover, we discuss a new direction of off-policy RL for both CT and DT systems. Finally, we review several applications. Bahare Kiumarsi-Khomartash, Kyriakos G. Vamvoudakis, Hamidreza Modares, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Adaptive Asymptotic Neural Network Control of Nonlinear Systems With Unknown Actuator QuantizationabstractIn this paper, we propose an adaptive neural-network-based asymptotic control algorithm for a class of nonlinear systems subject to unknown actuator quantization. To this end, we exploit the sector property of the quantization nonlinearity and transform actuator quantization control problem into analyzing its upper bounds, which are then handled by a dynamic loop gain function-based approach. In our adaptive control scheme, there is only one parameter required to be estimated online for updating weights of neural networks. Within the framework of Lyapunov theory, it is shown that the proposed algorithm ensures that all the signals in the closed-loop system are ultimately bounded. Moreover, an asymptotic tracking error is obtained by means of introducing Barbalat's lemma to the proposed adaptive law. Kan Xie 0002, Ci Chen 0002, Frank L. Lewis, Shengli Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Special Issue on Deep Reinforcement Learning and Adaptive Dynamic ProgrammingabstractThe sixteen papers in this special section focus on deep reinforcement learning and adaptive dynamic programming (deep RL/ADP). Deep RL is able to output control signal directly based on input images, which incorporates both the advantages of the perception of deep learning (DL) and the decision making of RL or adaptive dynamic programming (ADP). This mechanism makes the artificial intelligence much closer to human thinking modes. Deep RL/ADP has achieved remarkable success in terms of theory and applications since it was proposed. Successful applications cover video games, Go, robotics, smart driving, healthcare, and so on. However, it is still an open problem to perform the theoretical analysis on deep RL/ADP, e.g., the convergence, stability, and optimality analyses. The learning efficiency needs to be improved by proposing new algorithms or combined with other methods. More practical demonstrations are encouraged to be presented. Therefore, the aim of this special issue is to call for the most advanced research and state-of-the-art works in the field of deep RL/ADP. Dongbin Zhao, Derong Liu 0001, Frank L. Lewis, José C. Príncipe, Stefano Squartini |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Neuro-Adaptive Distributed Control With Prescribed Performance for the Synchronization of Unknown Nonlinear Networked SystemsabstractThis paper proposes a neuro-adaptive distributive cooperative tracking control with prescribed performance function (PPF) for highly nonlinear multiagent systems. PPF allows error tracking from a predefined large set to be trapped into a predefined small set. The key idea is to transform the constrained system into unconstrained one through transformation of the output error. Agents' dynamics are assumed to be completely unknown, and the controller is developed for strongly connected structured network. The proposed controller allows all agents to follow the trajectory of the leader node, while satisfying necessary dynamic requirements. The proposed approach guarantees uniform ultimate boundedness of the transformed error and the adaptive neural network weights. Simulations include two examples to validate the robustness and smoothness of the proposed controller against highly nonlinear heterogeneous networked system with time varying uncertain parameters and external disturbances. Sami El-Ferik, Hashim A. Hashim, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | Discrete-Time Local Value Iteration Adaptive Dynamic Programming: Convergence AnalysisabstractIn this paper, convergence properties are established for the newly developed discrete-time local value iteration adaptive dynamic programming (ADP) algorithm. The present local iterative ADP algorithm permits an arbitrary positive semidefinite function to initialize the algorithm. Employing a state-dependent learning rate function, for the first time, the iterative value function and iterative control law can be updated in a subset of the state space instead of the whole state space, which effectively relaxes the computational burden. A new analysis method for the convergence property is developed to prove that the iterative value functions will converge to the optimum under some mild constraints. Monotonicity of the local value iteration ADP algorithm is presented, which shows that under some special conditions of the initial value function and the learning rate function, the iterative value function can monotonically converge to the optimum. Finally, three simulation examples and comparisons are given to illustrate the performance of the developed algorithm. Qinglai Wei, Frank L. Lewis, Derong Liu 0001, Ruizhuo Song, Hanquan Lin |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2017 | Distributed Fault-Tolerant Control of Networked Uncertain Euler-Lagrange Systems Under Actuator FaultsabstractThis paper investigates the distributed fault-tolerant control problem of networked Euler-Lagrange systems with actuator and communication link faults. An adaptive fault-tolerant cooperative control scheme is proposed to achieve the coordinated tracking control of networked uncertain Lagrange systems on a general directed communication topology, which contains a spanning tree with the root node being the active target system. The proposed algorithm is capable of compensating for the actuator bias fault, the partial loss of effectiveness actuation fault, the communication link fault, the model uncertainty, and the external disturbance simultaneously. The control scheme does not use any fault detection and isolation mechanism to detect, separate, and identify the actuator faults online, which largely reduces the online computation and expedites the responsiveness of the controller. To validate the effectiveness of the proposed method, a test-bed of multiple robot-arm cooperative control system is developed for real-time verification. Experiments on the networked robot-arms are conduced and the results confirm the benefits and the effectiveness of the proposed distributed fault-tolerant control algorithms. Gang Chen 0014, Yongduan Song 0001, Frank L. Lewis |
IEEE Trans. Cybern. | 3 |
| 2017 | Off-Policy Reinforcement Learning: Optimal Operational Control for Two-Time-Scale Industrial ProcessesabstractIndustrial flow lines are composed of unit processes operating on a fast time scale and performance measurements known as operational indices measured at a slower time scale. This paper presents a model-free optimal solution to a class of two time-scale industrial processes using off-policy reinforcement learning (RL). First, the lower-layer unit process control loop with a fast sampling period and the upper-layer operational index dynamics at a slow time scale are modeled. Second, a general optimal operational control problem is formulated to optimally prescribe the set-points for the unit industrial process. Then, a zero-sum game off-policy RL algorithm is developed to find the optimal set-points by using data measured in real-time. Finally, a simulation experiment is employed for an industrial flotation process to show the effectiveness of the proposed method. Jinna Li, Bahare Kiumarsi-Khomartash, Tianyou Chai, Frank L. Lewis, Jialu Fan |
IEEE Trans. Cybern. | 4 |
| 2017 | Policy Gradient Adaptive Dynamic Programming for Data-Based Optimal ControlabstractThe model-free optimal control problem of general discrete-time nonlinear systems is considered in this paper, and a data-based policy gradient adaptive dynamic programming (PGADP) algorithm is developed to design an adaptive optimal controller method. By using offline and online data rather than the mathematical system model, the PGADP algorithm improves control policy with a gradient descent scheme. The convergence of the PGADP algorithm is proved by demonstrating that the constructed Q -function sequence converges to the optimal Q -function. Based on the PGADP algorithm, the adaptive control method is developed with an actor-critic structure and the method of weighted residuals. Its convergence properties are analyzed, where the approximate Q -function converges to its optimum. Computer simulation results demonstrate the effectiveness of the PGADP-based adaptive control method. Biao Luo 0001, Derong Liu 0001, Huai-Ning Wu, Ding Wang 0001, Frank L. Lewis |
IEEE Trans. Cybern. | 5 |
| 2017 | Discrete-Time Deterministic Q-Learning: A Novel Convergence AnalysisabstractIn this paper, a novel discrete-time deterministic Q -learning algorithm is developed. In each iteration of the developed Q -learning algorithm, the iterative Q function is updated for all the state and control spaces, instead of updating for a single state and a single control in traditional Q -learning algorithm. A new convergence criterion is established to guarantee that the iterative Q function converges to the optimum, where the convergence criterion of the learning rates for traditional Q -learning algorithms is simplified. During the convergence analysis, the upper and lower bounds of the iterative Q function are analyzed to obtain the convergence criterion, instead of analyzing the iterative Q function itself. For convenience of analysis, the convergence properties for undiscounted case of the deterministic Q -learning algorithm are first developed. Then, considering the discounted factor, the convergence criterion for the discounted case is established. Neural networks are used to approximate the iterative Q function and compute the iterative control law, respectively, for facilitating the implementation of the deterministic Q -learning algorithm. Finally, simulation results and comparisons are given to illustrate the performance of the developed algorithm. Qinglai Wei, Frank L. Lewis, Qiuye Sun, Ruizhuo Song |
IEEE Trans. Cybern. | 2 |
| 2017 | Output Containment Control of Linear Heterogeneous Multi-Agent Systems Using Internal Model PrincipleabstractThis paper studies the output containment control of linear heterogeneous multi-agent systems, where the system dynamics and even the state dimensions can generally be different. Since the states can have different dimensions, standard results from state containment control do not apply. Therefore, the control objective is to guarantee the convergence of the output of each follower to the dynamic convex hull spanned by the outputs of leaders. This can be achieved by making certain output containment errors go to zero asymptotically. Based on this formulation, two different control protocols, namely, full-state feedback and static output-feedback, are designed based on internal model principles. Sufficient local conditions for the existence of the proposed control protocols are developed in terms of stabilizing the local followers' dynamics and satisfying a certain H∞ criterion. Unified design procedures to solve the proposed two control protocols are presented by formulation and solution of certain local state-feedback and static output-feedback problems, respectively. Numerical simulations are given to validate the proposed control protocols. Shan Zuo, Yongduan Song 0001, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 3 |
| 2017 | Off-Policy Reinforcement Learning for Synchronization in Multiagent Graphical GamesabstractThis paper develops an off-policy reinforcement learning (RL) algorithm to solve optimal synchronization of multiagent systems. This is accomplished by using the framework of graphical games. In contrast to traditional control protocols, which require complete knowledge of agent dynamics, the proposed off-policy RL algorithm is a model-free approach, in that it solves the optimal synchronization problem without knowing any knowledge of the agent dynamics. A prescribed control policy, called behavior policy, is applied to each agent to generate and collect data for learning. An off-policy Bellman equation is derived for each agent to learn the value function for the policy under evaluation, called target policy, and find an improved policy, simultaneously. Actor and critic neural networks along with least-square approach are employed to approximate target control policies and value functions using the data generated by applying prescribed behavior policies. Finally, an off-policy RL algorithm is presented that is implemented in real time and gives the approximate optimal control policy for each agent using only measured data. It is shown that the optimal distributed policies found by the proposed algorithm satisfy the global Nash equilibrium and synchronize all agents to the leader. Simulation results illustrate the effectiveness of the proposed method. Jinna Li, Hamidreza Modares, Tianyou Chai, Frank L. Lewis, Lihua Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Guest Editorial Special Issue on New Developments in Neural Network Structures for Signal Processing, Autonomous Decision, and Adaptive ControlabstractThere has been continuously increasing interest in applying neural networks (NNs) to identification and adaptive control of practical systems that are characterized by nonlinearity, uncertainty, communication constraints, and complexity. The past few years have witnessed a variety of new developments in NN-based approaches for behavior learning, information processing, autonomous decision, and system control. Biologically inspired NN structures can significantly enhance the capabilities of information processing, control, and computational performance. New discoveries in neurocognitive psychology, sociology, and elsewhere reveal new neurological learning structures with more powerful capabilities in complex problem solving and fast decision in dynamic environments. The goal of the special issue is to consolidate recent new developments in NN structures for signal processing, autonomous decision, and adaptive control with application to complex systems. It includes contributions from a wide range of research aspects relevant to the topic, ranging from neural computing, adaptive control, cooperative control, autonomous decision systems, mathematical and computational models, neuropsychology decision and control, algorithms and simulation, to applications and/or case studies. This issue contains 24 papers and the contents of which are summarized below. Yongduan Song 0001, Frank L. Lewis, Marios M. Polycarpou, Danil V. Prokhorov, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Off-Policy Integral Reinforcement Learning Method to Solve Nonlinear Continuous-Time Multiplayer Nonzero-Sum GamesabstractThis paper establishes an off-policy integral reinforcement learning (IRL) method to solve nonlinear continuous-time (CT) nonzero-sum (NZS) games with unknown system dynamics. The IRL algorithm is presented to obtain the iterative control and off-policy learning is used to allow the dynamics to be completely unknown. Off-policy IRL is designed to do policy evaluation and policy improvement in the policy iteration algorithm. Critic and action networks are used to obtain the performance index and control for each player. The gradient descent algorithm makes the update of critic and action weights simultaneously. The convergence analysis of the weights is given. The asymptotic stability of the closed-loop system and the existence of Nash equilibrium are proved. The simulation study demonstrates the effectiveness of the developed method for nonlinear CT NZS games with unknown system dynamics. Ruizhuo Song, Frank L. Lewis, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | An H∞ performance allocation approach to distributed output regulation of linear heterogeneous multi-agent systemsabstractThis paper is concerned with cooperative output regulation of heterogeneous multi-agent systems. Agents are allowed to be heterogeneous general linear time-invariant systems and the communication graph is not restricted to the acyclic type. New distributed state-feedback controllers with extra scalar parameters are constructed. A global sufficient solvability condition is first derived and simplified stability conditions for closed-loop poles in a specified region are then obtained in terms of an H∞-type performance of local sub-systems coupled through a matrix associated with the graph. Linear matrix inequality conditions are further presented for allocating the H∞-type performance levels and designing controllers for each sub-system. A numerical example is presented for illustrating the advantages of the proposed design method. Both continuous- and discrete-time multi-agent systems are investigated in a unified framework. Xianwei Li 0001, Yeng Chai Soh, Lihua Xie 0001, Frank L. Lewis |
ICARCV | 4 |
| 2016 | Optimal output synchronization of nonlinear multi-agent systems using approximate dynamic programmingabstractOptimal output synchronization of multi-agent leader-follower systems is considered. The agents are assumed heterogeneous so that the dynamics may be non-identical. An optimal control protocol is designed for each agent based on the leader state and the agent local state. A distributed observer is designed to provide the leader state for each agent. A model-free approximate dynamic programming algorithm is then developed to solve the optimal output synchronization problem online in real time. No knowledge of the agents' dynamics is required. The proposed approach does not require explicitly solving of the output regulator equations, though it implicitly solves them by imposing optimality. A simulation example verifies the suitability of the proposed approach. Hamidreza Modares, Frank L. Lewis, Ali Davoudi |
IJCNN | 2 |
| 2016 | Synchronization for an array of neural networks with hybrid coupling by a novel pinning control strategy
Dawei Gong, Frank L. Lewis |
Neural Networks | 2 |
| 2016 | Cooperative Output Regulation of Singular Heterogeneous Multiagent SystemsabstractThis paper investigates the cooperative output regulation problem of singular heterogeneous multiagent systems. General distributed observers are proposed for every agent obtaining the estimated state of the exosystem. The feedforward control technique and reduced-order approach are used to design distributed singular output feedback controllers and distributed normal output feedback controllers. The proposed cooperative dynamic controller is dependent on the plant parameters and the interaction topologies. A simulation example is provided to demonstrate the effectiveness of the proposed design method. Qian Ma 0001, Shengyuan Xu 0001, Frank L. Lewis, Baoyong Zhang |
IEEE Trans. Cybern. | 3 |
| 2016 | Optimal Output-Feedback Control of Unknown Continuous-Time Linear Systems Using Off-policy Reinforcement LearningabstractA model-free off-policy reinforcement learning algorithm is developed to learn the optimal output-feedback (OPFB) solution for linear continuous-time systems. The proposed algorithm has the important feature of being applicable to the design of optimal OPFB controllers for both regulation and tracking problems. To provide a unified framework for both optimal regulation and tracking, a discounted performance function is employed and a discounted algebraic Riccati equation (ARE) is derived which gives the solution to the problem. Conditions on the existence of a solution to the discounted ARE are provided and an upper bound for the discount factor is found to assure the stability of the optimal control solution. To develop an optimal OPFB controller, it is first shown that the system state can be constructed using some limited observations on the system output over a period of the history of the system. A Bellman equation is then developed to evaluate a control policy and find an improved policy simultaneously using only some limited observations on the system output. Then, using this Bellman equation, a model-free Off-policy RL-based OPFB controller is developed without requiring the knowledge of the system state or the system dynamics. It is shown that the proposed OPFB method is more powerful than the static OPFB as it is equivalent to a state-feedback control policy. The proposed method is successfully used to solve a regulation and a tracking problem. Hamidreza Modares, Frank L. Lewis, Zhong-Ping Jiang |
IEEE Trans. Cybern. | 2 |
| 2016 | Optimized Assistive Human-Robot Interaction Using Reinforcement LearningabstractAn intelligent human-robot interaction (HRI) system with adjustable robot behavior is presented. The proposed HRI system assists the human operator to perform a given task with minimum workload demands and optimizes the overall human-robot system performance. Motivated by human factor studies, the presented control structure consists of two control loops. First, a robot-specific neuro-adaptive controller is designed in the inner loop to make the unknown nonlinear robot behave like a prescribed robot impedance model as perceived by a human operator. In contrast to existing neural network and adaptive impedance-based control methods, no information of the task performance or the prescribed robot impedance model parameters is required in the inner loop. Then, a task-specific outer-loop controller is designed to find the optimal parameters of the prescribed robot impedance model to adjust the robot's dynamics to the operator skills and minimize the tracking error. The outer loop includes the human operator, the robot, and the task performance details. The problem of finding the optimal parameters of the prescribed robot impedance model is transformed into a linear quadratic regulator (LQR) problem which minimizes the human effort and optimizes the closed-loop behavior of the HRI system for a given task. To obviate the requirement of the knowledge of the human model, integral reinforcement learning is used to solve the given LQR problem. Simulation results on an x - y table and a robot arm, and experimental implementation results on a PR2 robot confirm the suitability of the proposed method. Hamidreza Modares, Isura Ranatunga, Frank L. Lewis, Dan O. Popa |
IEEE Trans. Cybern. | 3 |
| 2016 | Off-Policy Actor-Critic Structure for Optimal Control of Unknown Systems With DisturbancesabstractAn optimal control method is developed for unknown continuous-time systems with unknown disturbances in this paper. The integral reinforcement learning (IRL) algorithm is presented to obtain the iterative control. Off-policy learning is used to allow the dynamics to be completely unknown. Neural networks are used to construct critic and action networks. It is shown that if there are unknown disturbances, off-policy IRL may not converge or may be biased. For reducing the influence of unknown disturbances, a disturbances compensation controller is added. It is proven that the weight errors are uniformly ultimately bounded based on Lyapunov techniques. Convergence of the Hamiltonian function is also proven. The simulation study demonstrates the effectiveness of the proposed optimal control method for unknown systems with disturbances. Ruizhuo Song, Frank L. Lewis, Qinglai Wei, Huaguang Zhang |
IEEE Trans. Cybern. | 2 |
| 2016 | Data-Based Multiobjective Plant-Wide Performance Optimization of Industrial Processes Under Dynamic EnvironmentsabstractThis paper provides a method for automatically selecting optimal operational indices for unit processes in an industrial plant using measured data and without knowing dynamical models of the unit process. A dynamic multiobjective optimization problem is defined to find operational indices that lead to plant-wide production indices close to their target values. A case-based reasoning (CBR) technique is also employed, which uses the stored experience of a human expert to determine appropriate operational indices for given target production indices. The solutions of the optimization problem and CBR technique are combined to form baseline operational indices. The dynamic models of the production indices, however, are time varying and affected by disturbances and online corrections of these baseline operational indices are required. To this end, reinforcement learning (RL) is used to provide a data-driven optimization technique to compensate for disturbances and model approximation errors and variations. The data-driven RL approach is used in two different time scales. The samples of the predicted production indices are used at a fast sampling rate, i.e., at each sample time, and the samples of actual production indices are used at a slower sampling rate, i.e., after each operational run, to correct the baseline operational indices. The effectiveness of this automated decision procedure has been demonstrated by successful implementation of the proposed approach on a large mineral processing plant in Gansu Province, China. Jinliang Ding, Hamidreza Modares, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 4 |
| 2016 | Distributed Fault-Tolerant Control of Virtually and Physically Interconnected Systems With Application to High-Speed Trains Under Traction/Braking FailuresabstractThis paper investigates the tracking control problem of dynamic systems consisting of physically connected subsystems with virtual connections through local communication, where unknown unidentical nonlinearities, time-varying yet undetectable actuation faults, and varying actuation authorities are involved. The local communication nature and the physical uncertain interactions among the subsystems, together with the unpredictable actuation failures and control authority variation, make the underlying problem nontrivial, calling for a control solution that is not only decentralized (distributed) but also adaptive and fault tolerant. In this paper, with the aid of the concepts of generalized parameter estimation error and virtual regrouping, a distributed and fault-tolerant control design approach is presented by using local (neighboring) information exchange only. This method is applied to develop tracking and braking control schemes for high-speed trains subject to traction and braking failures. The proposed distributed control is capable of simultaneously coping with the physical interactions among the subsystems, compensating the uncertain control gains, and accommodating the undetectable actuation faults, as authenticated and verified by theoretical analysis and numerical simulations. Yujuan Wang 0001, Yongduan Song 0001, Hui Gao 0003, Frank L. Lewis |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2015 | Intent aware adaptive admittance control for physical Human-Robot InteractionabstractEffective physical Human-Robot Interaction (pHRI) needs to account for variable human dynamics and also predict human intent. Recently, there has been a lot of progress in adaptive impedance and admittance control for human-robot interaction. Not as many contributions have been reported on online adaptation schemes that can accommodate users with varying physical strength and skill level during interaction with a robot. The goal of this paper is to present and evaluate a novel adaptive admittance controller that can incorporate human intent, nominal task models, as well as variations in the robot dynamics. An outer-loop controller is developed using an ARMA model which is tuned using an adaptive inverse control technique. An inner-loop neuroadaptive controller linearizes the robot dynamics. Working in conjunction and online, this two-loop technique offers an elegant way to decouple the pHRI problem. Experimental results are presented comparing the performance of different types of admittance controllers. The results show that efficient online adaptation of the robot admittance model for different human subjects can be achieved. Specifically, the adaptive admittance controller reduces jerk which results in a smooth human-robot interaction. Isura Ranatunga, Sven Cremer, Dan O. Popa, Frank L. Lewis |
ICRA | 4 |
| 2015 | A neural network model of decisions on the Asian Disease ProblemabstractTversky and Kahneman [1] found that human decisions can be inconsistent across descriptions of the options. An example is the Asian Disease Problem, whereby preferences between two public health programs are different when options are framed in terms of deaths versus lives saved. Several variants of the Asian Disease paradigm and an analogous problem were run [2, 3]: the results showed that the strength of the framing effect depended on whether one of options explicitly contained the possibility of no lives lost or saved. This result was explained by fuzzy trace theory whereby decisions are based not on details of the options but on their gist (underlying meaning). We simulate these results using a neural network model that blends fuzzy trace theory with adaptive resonance theory and contains analogs of orbitofrontal cortex, amygdala, anterior cingulate, and striatum. Bakur AlQaudi, Daniel S. Levine 0001, Frank L. Lewis |
IJCNN | 3 |
| 2015 | Optimal control of nonlinear discrete time-varying systems using a new neural network approximation structure
Bahare Kiumarsi-Khomartash, Frank L. Lewis, Daniel S. Levine 0001 |
Neurocomputing | 2 |
| 2015 | Optimal distributed synchronization control for continuous-time heterogeneous multi-agent differential graphical games
Qinglai Wei, Derong Liu 0001, Frank L. Lewis |
Inf. Sci. | 3 |
| 2015 | Optimal Tracking Control of Unknown Discrete-Time Linear Systems Using Input-Output Measured DataabstractIn this paper, an output-feedback solution to the infinite-horizon linear quadratic tracking (LQT) problem for unknown discrete-time systems is proposed. An augmented system composed of the system dynamics and the reference trajectory dynamics is constructed. The state of the augmented system is constructed from a limited number of measurements of the past input, output, and reference trajectory in the history of the augmented system. A novel Bellman equation is developed that evaluates the value function related to a fixed policy by using only the input, output, and reference trajectory data from the augmented system. By using approximate dynamic programming, a class of reinforcement learning methods, the LQT problem is solved online without requiring knowledge of the augmented system dynamics only by measuring the input, output, and reference trajectory from the augmented system. We develop both policy iteration (PI) and value iteration (VI) algorithms that converge to an optimal controller that require only measuring the input, output, and reference trajectory data. The convergence of the proposed PI and VI algorithms is shown. A simulation example is used to verify the effectiveness of the proposed control scheme. Bahare Kiumarsi-Khomartash, Frank L. Lewis, Mohammad-Bagher Naghibi-Sistani, Ali Karimpour |
IEEE Trans. Cybern. | 2 |
| 2015 | Continuous-Time Q-Learning for Infinite-Horizon Discounted Cost Linear Quadratic Regulator ProblemsabstractThis paper presents a method of Q-learning to solve the discounted linear quadratic regulator (LQR) problem for continuous-time (CT) continuous-state systems. Most available methods in the existing literature for CT systems to solve the LQR problem generally need partial or complete knowledge of the system dynamics. Q-learning is effective for unknown dynamical systems, but has generally been well understood only for discrete-time systems. The contribution of this paper is to present a Q-learning methodology for CT systems which solves the LQR problem without having any knowledge of the system dynamics. A natural and rigorous justified parameterization of the Q-function is given in terms of the state, the control input, and its derivatives. This parameterization allows the implementation of an online Q-learning algorithm for CT systems. The simulation results supporting the theoretical development are also presented. Muthukumar Palanisamy, Hamidreza Modares, Frank L. Lewis, Muhammad Aurangzeb |
IEEE Trans. Cybern. | 3 |
| 2015 | Actor-Critic-Based Optimal Tracking for Partially Unknown Nonlinear Discrete-Time SystemsabstractThis paper presents a partially model-free adaptive optimal control solution to the deterministic nonlinear discrete-time (DT) tracking control problem in the presence of input constraints. The tracking error dynamics and reference trajectory dynamics are first combined to form an augmented system. Then, a new discounted performance function based on the augmented system is presented for the optimal nonlinear tracking problem. In contrast to the standard solution, which finds the feedforward and feedback terms of the control input separately, the minimization of the proposed discounted performance function gives both feedback and feedforward parts of the control input simultaneously. This enables us to encode the input constraints into the optimization problem using a nonquadratic performance function. The DT tracking Bellman equation and tracking Hamilton-Jacobi-Bellman (HJB) are derived. An actor-critic-based reinforcement learning algorithm is used to learn the solution to the tracking HJB equation online without requiring knowledge of the system drift dynamics. That is, two neural networks (NNs), namely, actor NN and critic NN, are tuned online and simultaneously to generate the optimal bounded control policy. A simulation example is given to show the effectiveness of the proposed method. Bahare Kiumarsi-Khomartash, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | H∞ Tracking Control of Completely Unknown Continuous-Time Systems via Off-Policy Reinforcement LearningabstractThis paper deals with the design of an H ∞ tracking controller for nonlinear continuous-time systems with completely unknown dynamics. A general bounded L2 -gain tracking problem with a discounted performance function is introduced for the H ∞ tracking. A tracking Hamilton-Jacobi-Isaac (HJI) equation is then developed that gives a Nash equilibrium solution to the associated min-max optimization problem. A rigorous analysis of bounded L2 -gain and stability of the control solution obtained by solving the tracking HJI equation is provided. An upper-bound is found for the discount factor to assure local asymptotic stability of the tracking error dynamics. An off-policy reinforcement learning algorithm is used to learn the solution to the tracking HJI equation online without requiring any knowledge of the system dynamics. Convergence of the proposed algorithm to the solution to the tracking HJI equation is shown. Simulation examples are provided to verify the effectiveness of the proposed method. Hamidreza Modares, Frank L. Lewis, Zhong-Ping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Multiple Actor-Critic Structures for Continuous-Time Optimal Control Using Input-Output DataabstractIn industrial process control, there may be multiple performance objectives, depending on salient features of the input-output data. Aiming at this situation, this paper proposes multiple actor-critic structures to obtain the optimal control via input-output data for unknown nonlinear systems. The shunting inhibitory artificial neural network (SIANN) is used to classify the input-output data into one of several categories. Different performance measure functions may be defined for disparate categories. The approximate dynamic programming algorithm, which contains model module, critic network, and action network, is used to establish the optimal control in each category. A recurrent neural network (RNN) model is used to reconstruct the unknown system dynamics using input-output data. NNs are used to approximate the critic and action networks, respectively. It is proven that the model error and the closed unknown system are uniformly ultimately bounded. Simulation results demonstrate the performance of the proposed optimal control scheme for the unknown nonlinear system. Ruizhuo Song, Frank L. Lewis, Qinglai Wei, Huaguang Zhang, Zhong-Ping Jiang, Daniel S. Levine 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Team-oriented adaptive droop control for autonomous AC microgridsabstractThis paper proposes a distributed control strategy for voltage and reactive power regulation in ac Microgrids. First, the control module introduces a voltage regulator that maintains the average voltage of the system on the rated value, keeping all bus voltages within an acceptable range. Dynamic consensus protocol is used to estimate the average voltage across the Microgrid. This estimation is further utilized by the voltage regulator to elevate/lower the voltage-reactive power (Q-E) droop characteristic, compensating the drop caused by the droop mechanism. The second module, the reactive power regulator, dynamically fine-tunes the Q-E coefficients to handle the proportional reactive power sharing. Accordingly, locally supplied reactive power of any source is compared with neighbor sources and the local droop coefficient is adjusted to mitigate and, ultimately, eliminate the load mismatch. The proposed controllers are fully distributed; i.e., each source requires information exchange with only a few other sources, those in direct contact through the communication infrastructure. A Microgrid test bench is used to verify the proposed control methodology, where different test scenarios such as load change, link failure, and inverter outage are carried out. Qobad Shafiee, Vahidreza Nasirian, Josep M. Guerrero, Frank L. Lewis, Ali Davoudi |
IECON | 4 |
| 2014 | A Multiobjective Distributed Control Framework for Islanded AC MicrogridsabstractThis paper proposes a distributed two-layer control structure for ac microgrids. Inverter-based distributed generators (DGs) can operate either as voltage-controlled voltage source inverters (VCVSI) or current-controlled voltage source inverters (CCVSI). VCVSIs provide the voltage and frequency support, whereas CCVSIs regulate the generated active and reactive powers. The proposed control structure has two main layers. The first layer deals with the voltage and frequency control of VCVSIs. The second layer regulates the active and reactive powers of CCVSIs. These controllers are implemented through two communication networks with one-way communication links and are fully distributed; each DG only requires its own information and the information of its neighbors on the communication network graph. The proposed control framework is verified on a microgrid test system and IEEE 34 test feeder. Ali Bidram, Ali Davoudi, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 3 |
| 2014 | A Distributed Auction-Based Algorithm for the Nonconvex Economic Dispatch ProblemabstractThis paper presents a distributed algorithm based on auction techniques and consensus protocols to solve the nonconvex economic dispatch problem. The optimization problem of the nonconvex economic dispatch includes several constraints such as valve-point loading effect, multiple fuel option, and prohibited operating zones. Each generating unit locally evaluates quantities used as bids in the auction mechanism. These units send their bids to their neighbors in a communication graph that supports the power system and which provides the required information flow. A consensus procedure is used to share the bids among the network agents and resolves the auction. As a result, the power distribution of generating units is updated and the generation cost is minimized. The effectiveness of this approach is demonstrated by simulations on standard test systems. Giulio Binetti, Ali Davoudi, David Naso, Biagio Turchiano, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 5 |
| 2013 | Approximate dynamic programming solutions of multi-agent graphical games using actor-critic network structuresabstractThis paper studies a new class of multi-agent discrete-time dynamical graphical games, where interactions between agents are restricted by a communication graph structure. The paper brings together discrete Hamiltonian mechanics, optimal control theory, cooperative control, game theory, reinforcement learning, and neural network structures to solve the multi-agent dynamical graphical games. Graphical game Bellman equations are derived and shown to be equivalent to certain graphical game Hamilton Jacobi Bellman equations developed herein. Reinforcement Learning techniques are used to solve these dynamical graphical games. Heuristic Dynamic Programming and Dual Heuristic Programming, are extended to solve the graphical games using only neighborhood information. Online adaptive learning structure is implemented using actor-critic networks to solve these graphical games. Mohammed I. Abouheaf, Frank L. Lewis |
IJCNN | 2 |
| 2013 | Guest Editorial Advances in Theories and Industrial Applications of Networked Control SystemsabstractThe articles in this special section focus on advancements in theories and industrial applications of networked control systems in the industrial informatics industry. Lixian Zhang 0001, Huijun Gao, Frank L. Lewis, Okyay Kaynak |
IEEE Trans. Ind. Informatics | 3 |
| 2013 | Adaptive Optimal Control of Unknown Constrained-Input Systems Using Policy Iteration and Neural NetworksabstractThis paper presents an online policy iteration (PI) algorithm to learn the continuous-time optimal control solution for unknown constrained-input systems. The proposed PI algorithm is implemented on an actor-critic structure where two neural networks (NNs) are tuned online and simultaneously to generate the optimal bounded control policy. The requirement of complete knowledge of the system dynamics is obviated by employing a novel NN identifier in conjunction with the actor and critic NNs. It is shown how the identifier weights estimation error affects the convergence of the critic NN. A novel learning rule is developed to guarantee that the identifier weights converge to small neighborhoods of their ideal values exponentially fast. To provide an easy-to-check persistence of excitation condition, the experience replay technique is used. That is, recorded past experiences are used simultaneously with current data for the adaptation of the identifier weights. Stability of the whole system consisting of the actor, critic, system state, and system identifier is guaranteed while all three networks undergo adaptation. Convergence to a near-optimal control law is also shown. The effectiveness of the proposed method is illustrated with a simulation example. Hamidreza Modares, Frank L. Lewis, Mohammad-Bagher Naghibi-Sistani |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Managing Complex Mechatronics R&D: A Systems Design ApproachabstractTo compress research and development (R&D) cycle times of high-tech mechatronic products with conformance performance metrics, managing R&D projects to allow engineers from electrical, mechanical, and manufacturing disciplines receive real-time design feedback and assessment are essential. In this paper, we propose a systems design procedure to integrate mechanical design, structure prototyping, and servo evaluation through careful comprehension of the servo-mechanical-prototype production cycle commonly employed in mechatronic industries. Our approach focuses on the Modal Parametric Identification of key feedback parameters for fast exchange of design specifications and information. This enables efficient conduct of product design evaluations, and supports schedule compression of the R&D project life cycle in the highly competitive consumer electronics industry. Using the commercial hard disk drive as a case example, we demonstrate how our approach allow inter-disciplinary specifications to be communicated among engineers from different backgrounds to speed up the R&D process for the next generation of intelligent manufacturing. This provides the management of technology team with powerful decision-making tools for project strategy formulation, and improvements in project outcome are potentially massive because of the low costs of change. Chee Khiang Pang, Tsan Sheng Ng, Frank L. Lewis, Tong Heng Lee |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2011 | An approximate Dynamic Programming based controller for an underactuated 6DoF quadrotorabstractThis paper discusses how the principles of Adaptive Dynamic Programming (ADP) can be applied to the control of a quadrotor helicopter platform flying in an uncontrolled environment and subjected to various disturbances and model uncertainties. ADP is based on reinforcement learning using an actor-critic structure. Due to the complexity of the quadrotor system, the learning process has to use as much information as possible about the system and the environment. Various methods to improve the learning speed and efficiency are presented. Neural networks with local activation functions are used as function approximators because the state-space can not be explored efficiently due to its size and the limited time available. The complex dynamics is controlled by a single critic and by multiple actors thus avoiding the curse of dimensionality. After a number of iterations, the overall actor-critic structure stores information (knowledge) about the system dynamics and the optimal controller that can accomplish the explicit or implicit goal specified in the cost function. Petru Emanuel Stingu, Frank L. Lewis |
ADPRL | 2 |
| 2011 | Online adaptive learning of optimal control solutions using integral reinforcement learningabstractIn this paper we introduce an online algorithm that uses integral reinforcement knowledge for learning the continuous-time optimal control solution for nonlinear systems with infinite horizon costs and partial knowledge of the system dynamics. This algorithm is a data based approach to the solution of the Hamilton-Jacobi-Bellman equation and it does not require explicit knowledge on the system's drift dynamics. The adaptive algorithm is based on policy iteration, and it is implemented on an actor/critic structure. Both actor and critic neural networks are adapted simultaneously a persistence of excitation condition is required to guarantee convergence of the critic to the actual optimal value function. Novel tuning algorithms are given for both critic and actor networks, with extra terms in the actor tuning law being required to guarantee closed-loop dynamical stability. The convergence to the optimal controller is proven, and stability of the system is also guaranteed. Simulation examples support the theoretical result. Kyriakos G. Vamvoudakis, Draguna L. Vrabie, Frank L. Lewis |
ADPRL | 3 |
| 2011 | Decentralized task sequencing and multiple mission control for heterogeneous robotic networksabstractIn this paper a novel decentralized approach for task sequencing within a multiple missions control framework is presented. The main contribution of this work concerns the decentralization of a control framework for multiple mission execution in order to enhance the robustness of the system, and the application of the latter to a heterogeneous robotic network. The proposed approach is based on the Matrix-based Discrete Event Framework (MDEF). This formalism is adapted to networks of heterogeneous robots, i.e., robots with different capabilities, and to the decentralized control of mission execution using a consensus-based approach which guarantees the agreement among robots on executed actions and their consequences. Donato Di Paola, Andrea Gasparri, David Naso, Giovanni Ulivi, Frank L. Lewis |
ICRA | 5 |
| 2011 | Dominant Feature Identification for Industrial Fault Detection and Isolation Applications
Junhong Zhou, Chee Khiang Pang, Frank L. Lewis, Zhao-Wei Zhong |
Expert Syst. Appl. | 3 |
| 2011 | Guest Editorial Data-Based Control, Modeling, and OptimizationabstractThe 21 papers in this special section focus on data-based control, modeling, and optimization. Tianyou Chai, Zhongsheng Hou, Frank L. Lewis, Amir Hussain 0001, Dongbin Zhao |
IEEE Trans. Neural Networks | 3 |
| 2011 | Distributed Adaptive Tracking Control for Synchronization of Unknown Networked Lagrangian SystemsabstractThis paper investigates the cooperative tracking control problem for a group of Lagrangian vehicle systems with directed communication graph topology. All the vehicles can have different dynamics. A design method for a distributed adaptive protocol is given which guarantees that all the networked systems synchronize to the motion of a target system. The dynamics of the networked systems, as well as the target system, are all assumed unknown. A neural network (NN) is used at each node to approximate the distributed dynamics. The resulting protocol consists of a simple decentralized proportional-plus-derivative term and a nonlinear term with distributed adaptive tuning laws at each node. The case with nonconstant NN approximation error is considered. There, a robust term is added to suppress the external disturbances and the approximation errors of the NNs. Simulation examples are included to demonstrate the effectiveness of the proposed algorithms. Gang Chen 0014, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Reinforcement Learning for Partially Observable Dynamic Processes: Adaptive Dynamic Programming Using Measured Output DataabstractApproximate dynamic programming (ADP) is a class of reinforcement learning methods that have shown their importance in a variety of applications, including feedback control of dynamical systems. ADP generally requires full information about the system internal states, which is usually not available in practical situations. In this paper, we show how to implement ADP methods using only measured input/output data from the system. Linear dynamical systems with deterministic behavior are considered herein, which are systems of great interest in the control system community. In control system theory, these types of methods are referred to as output feedback (OPFB). The stochastic equivalent of the systems dealt with in this paper is a class of partially observable Markov decision processes. We develop both policy iteration and value iteration algorithms that converge to an optimal controller that requires only OPFB. It is shown that, similar to Q -learning, the new methods have the important advantage that knowledge of the system dynamics is not needed for the implementation of these learning algorithms or for the OPFB control. Only the order of the system, as well as an upper bound on its "observability index," must be known. The learned OPFB controller is in the form of a polynomial autoregressive moving-average controller that has equivalent performance with the optimal state variable feedback gain. Frank L. Lewis, Kyriakos G. Vamvoudakis |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | A Cost Function Based Single Network Adaptive Critic architecture for optimal control synthesis for a class of nonlinear systemsabstractApproximate dynamic programming implemented with an Adaptive Critic (AC) neural network (NN) structure has evolved as a powerful alternative technique that eliminates the need for excessive computations and storage requirements in solving optimal control problems. A typical AC structure consists of two interacting NNs. In this paper, a new architecture, called the “Cost Function Based Single Network Adaptive Critic (J-SNAC)” is presented. This approach is applicable to a wide class of nonlinear systems where the optimal control (stationary) equation can be explicitly expressed in terms of the state and cost variables. Selection of this terminology is guided by the fact that it eliminates the use of one NN (namely the action network) that is part of a typical dual network AC. In order to demonstrate the benefits and the control synthesis technique in using the J-SNAC, two problems have been solved with the AC and the J-SNAC approaches. Results are presented that show savings of about 50% of the computational costs by J-SNAC while having the same accuracy levels of the dual network structure in solving for optimal control. Convergence of the J-SNAC iterations is discussed as well as the reduction of the iterative process to the familiar algebraic Ricatti equation in the case of linear systems. S. N. Balakrishnan, Frank L. Lewis |
IJCNN | 3 |
| 2010 | Adaptive Dynamic Programming algorithm for finding online the equilibrium solution of the two-player zero-sum differential gameabstractThis paper will present an Approximate/Adaptive Dynamic Programming (ADP) algorithm for determining online the Nash equilibrium solution for the two-player zero-sum differential game with linear dynamics and infinite horizon quadratic cost. The algorithm is built around an iterative method that has been developed in the control engineering community for solving the continuous-time game algebraic Riccati equation (CT-GARE) that is underlying the game problem. We here show how the ADP techniques will enhance the capabilities of the offline method allowing an online solution without the requirement of complete knowledge of the system dynamics. While working in the framework of control applications we will be referring to the two players as controller and disturbance. Both players are competing in real time and the equilibrium solution policies will be determined based on online measured data from the system. The two players are not learning concurrently. The algorithm is built on interplay between a learning phase, performed by the controller that is learning in order to optimize its behavior, and a policy update step, performed by the disturbance that is increasing its detrimental effect. The update of the disturbance policy will give way for further improvement for the, no longer optimal, controller policy. The control policy will be learned online using a continuous-time heuristic dynamic programming procedure. The feasibility of the ADP scheme is demonstrated in simulation on a power system. The goal is to determine the best control policy that will face in an optimal manner the highest load disturbance. Draguna L. Vrabie, Frank L. Lewis |
IJCNN | 2 |
| 2009 | Online policy iteration based algorithms to solve the continuous-time infinite horizon optimal control problemabstractIn this paper we discuss two online algorithms based on policy iterations for learning the continuous-time (CT) optimal control solution when nonlinear systems with infinite horizon quadratic cost are considered. For the first time we present an online adaptive algorithm implemented on an actor/critic structure which involves synchronous continuous-time adaptation of both actor and critic neural networks. This is a version of generalized policy iteration for CT systems. The convergence to the optimal controller based on the novel algorithm is proven while stability of the system is guaranteed. The characteristics and requirements of the new online learning algorithm are discussed in relation with the regular online policy iteration algorithm for CT systems which we have previously developed. The latter solves the optimal control problem by performing sequential updates on the actor and critic networks, i.e. while one is learning the other one is held constant. In contrast, the new algorithm relies on simultaneous adaptation of both actor and critic networks. To support the new theoretical result a simulation example is then considered. Kyriakos G. Vamvoudakis, Draguna L. Vrabie, Frank L. Lewis |
ADPRL | 3 |
| 2009 | Algorithm and stability of ATC receding horizon controlabstractReceding horizon control (RHC), also known as model predictive control (MPC), is a suboptimal control scheme that solves a finite horizon open-loop optimal control problem in an infinite horizon context and yields a measured state feedback control law. A lot of efforts have been made to study the closed-loop stability, leading to various stability conditions involving constraints on either the terminal state, or the terminal cost, or the horizon size, or their different combinations. In this paper, we propose a modified RHC scheme, called adaptive terminal cost RHC (ATC-RHC). The control law generated by ATC-RHC algorithm converges to the solution of the infinite horizon optimal control problem. Moreover, it ensures the closed-loop system to be uniformly ultimately exponentially stable without imposing any constraints on the terminal state, the horizon size, or the terminal cost. Finally we show that when the horizon size is one, the underlying problems of ATC-RHC and heuristic dynamic programming (HDP) are the same. Thus, ATC-RHC can be implemented using HDP techniques without knowing the system matrix A. Hongwei Zhang 0005, Jie Huang 0001, Frank L. Lewis |
ADPRL | 3 |
| 2009 | Online actor critic algorithm to solve the continuous-time infinite horizon optimal control problemabstractIn this paper we discuss an online algorithm based on policy iteration for learning the continuous-time (CT) optimal control solution with infinite horizon cost for nonlinear systems with known dynamics. We present an online adaptive algorithm implemented as an actor/critic structure which involves simultaneous continuous-time adaptation of both actor and critic neural networks. We call this dasiasynchronouspsila policy iteration. A persistence of excitation condition is shown to guarantee convergence of the critic to the actual optimal value function. Novel tuning algorithms are given for both critic and actor networks, with extra terms in the actor tuning law being required to guarantee closed-loop dynamical stability. The convergence to the optimal controller is proven, and stability of the system is also guaranteed. Simulation examples show the effectiveness of the new algorithm. Kyriakos G. Vamvoudakis, Frank L. Lewis |
IJCNN | 2 |
| 2009 | Generalized Policy Iteration for continuous-time systemsabstractIn this paper we present a unified point of view over the approximate dynamic programming (ADP) algorithms which have been developed in the last years for continuous-time (CT) systems. We introduce here, in a continuous-time formulation, the generalized policy iteration (GPI), and show that in effect it represents a spectrum of algorithms which has at one end the exact policy iteration (PI) algorithm and at the other the value iteration (VI) algorithm. At the middle part of the spectrum we formulate for the first time the optimistic policy iteration (OPI) algorithm for CT systems. We introduce the GPI starting from a new formulation for the PI algorithm which involves an iterative process to solve for the value function at the policy evaluation step. The GPI algorithm is implemented on an actor/critic structure. The results allow implementation of a family of adaptive controllers which converge online to the solution of the optimal control problem, without knowing or identifying the internal dynamics of the system. Simulation results are provided to verify the convergence to the optimal control solution. Draguna L. Vrabie, Frank L. Lewis |
IJCNN | 2 |
| 2009 | Neural network approach to continuous-time direct adaptive optimal control for partially unknown nonlinear systems
Draguna L. Vrabie, Frank L. Lewis |
Neural Networks | 2 |
| 2009 | Intelligent Diagnosis and Prognosis of Tool Wear Using Dominant Feature IdentificationabstractIdentification and prediction of a lifetime of industrial cutting tools using minimal sensors is crucial to reduce production costs and downtime in engineering systems. In this paper, we provide a formal decision software tool to extract the dominant features enabling tool wear prediction. This decision tool is based on a formal mathematical approach that selects dominant features using the singular value decomposition of real-time measurements from the sensors of an industrial cutting tool. Selection of dominant features is important, as retaining only essential features allows reduced signal processing or even reduction in the number of required sensors, which cuts costs. It is shown that the proposed method of dominant feature selection is optimal in the sense that it minimizes the least-squares estimation error. The identified dominant features are used with the recursive least squares (RLS) algorithm to identify parameters in forecasting the time series of cutting tool wear. Experimental results on an industrial high-speed milling machine show the effectiveness in predicting the tool wear using only the dominant features. Junhong Zhou, Chee Khiang Pang, Frank L. Lewis, Zhao-Wei Zhong |
IEEE Trans. Ind. Informatics | 3 |
| 2008 | Using blending control to suppress multi-frequency disturbancesabstractIn this paper, rejecting multi-frequency narrow-band disturbances is formulated as a control blending problem according to each disturbance characteristic. Each disturbance rejection is accomplished by using H2optimal control method. Based on all the H2optimal controls, the blending technique is employed to yield a single controller which is capable to achieve the rejection for all disturbances. Rejections for two and three disturbances through the control design for the VCM actuator in a hard disk drive are taken as application examples in the current paper. Simulation and experimental results show that the ultimate controller results in a simultaneous attenuation to disturbances with frequencies higher or lower than the closed-loop system bandwidth. Moreover, the method turns out to be able to lift phase and thus prevent phase margin loss when it is used to deal with disturbances near bandwidth. Chunling Du, Lihua Xie 0001, Frank L. Lewis |
ICARCV | 3 |
| 2008 | Real time controller design to solve the "pull-in" instability of MEMS actuatorabstractThe purpose of this paper is to improve the performance of MEMS parallel plate actuators and RF switches systems. Electrostatic micro actuators are normally driven by static open-loop voltage control schemes. A major problem in this control strategy is that at a distance of two-thirds of the zero-bias capacitive gap, the actuator position becomes unstable and collapses. This phenomenon is known as “snap-through” or “pull-in”. In this paper a new closed-loop feedback control scheme using switching technique will be proposed to solve the problem above to provide the stable and controllable range to full gap and controlling movable plate through whole gap to achieve desired switching on-off time. The other issue to be considered is that the control strategy must be applicable for real time implementation. Mohammad Hossein Nikpanah, Youyi Wang, Frank L. Lewis, Ai Qun Liu |
ICARCV | 3 |
| 2008 | Neural Network-Based Adaptive Optimal Controller - A Continuous-Time Formulation
Draguna L. Vrabie, Frank L. Lewis, Daniel S. Levine 0001 |
ICIC (3) | 2 |
| 2008 | Integrated Supervisory and Operational Control of a Warehouse With a Matrix-Based ApproachabstractThis paper considers a matrix-based discrete event control approach for a warehouse. The control system is organized in two modules: a dynamic model and a controller. The model provides a complete description of the discrete event dynamics of the warehouse, and is used as a means to track the stock-keeping units, and identify and inhibit control actions that violate system's constraints. The controller has several functions. At the supervisory level, it is in charge of inhibiting operations that may lead to deadlocks, commanding the actual start of the task, and the release of the resources once a task is completed. At the operational level, it is in charge of performing decisions regarding the order in which allowable tasks waiting for service should be performed. All the modules are implemented using the same matrix-based formalism, and thus integrated with each other. The main advantages of the approach are the inherent modularity (the matrix-based control is obtained by assembling individual atomic components), and the integration between the various modules, which permits a better overall resource utilization. Simulation examples describing an actual industrial warehouse are finally provided to emphasize the main advantages of the proposed approach. Vincenzo Giordano, Jing Bing Zhang, David Naso, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2008 | Neurodynamic Programmingand Zero-Sum Games for Constrained Control SystemsabstractIn this paper, neural networks are used along with two-player policy iterations to solve for the feedback strategies of a continuous-time zero-sum game that appears in L2-gain optimal control, suboptimal Hinfincontrol, of nonlinear systems affine in input with the control policy having saturation constraints. The result is a closed-form representation, on a prescribed compact set chosen a priori, of the feedback strategies and the value function that solves the associated Hamilton-Jacobi-Isaacs (HJI) equation. The closed-loop stability, L2-gain disturbance attenuation of the neural network saturated control feedback strategy, and uniform convergence results are proven. Finally, this approach is applied to the rotational/translational actuator (RTAC) nonlinear benchmark problem under actuator saturation, offering guaranteed stability and disturbance attenuation. Murad Abu-Khalaf, Frank L. Lewis, Jie Huang 0001 |
IEEE Trans. Neural Networks | 2 |
| 2008 | Discrete-Time Nonlinear HJB Solution Using Approximate Dynamic Programming: Convergence ProofabstractConvergence of the value-iteration-based heuristic dynamic programming (HDP) algorithm is proven in the case of general nonlinear systems. That is, it is shown that HDP converges to the optimal control and the optimal value function that solves the Hamilton-Jacobi-Bellman equation appearing in infinite-horizon discrete-time (DT) nonlinear optimal control. It is assumed that, at each iteration, the value and action update equations can be exactly solved. The following two standard neural networks (NN) are used: a critic NN is used to approximate the value function, whereas an action network is used to approximate the optimal control policy. It is stressed that this approach allows the implementation of HDP without knowing the internal dynamics of the system. The exact solution assumption holds for some classes of nonlinear systems and, specifically, in the specific case of the DT linear quadratic regulator (LQR), where the action is linear and the value quadratic in the states and NNs have zero approximation error. It is stressed that, for the LQR, HDP may be implemented without knowing the system A matrix by using two NNs. This fact is not generally appreciated in the folklore of HDP for the DT LQR, where only one critic NN is generally used. Asma Al-Tamimi, Frank L. Lewis, Murad Abu-Khalaf |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Issues on Stability of ADP Feedback Controllers for Dynamical SystemsabstractThis paper traces the development of neural-network (NN)-based feedback controllers that are derived from the principle of adaptive/approximate dynamic programming (ADP) and discusses their closed-loop stability. Different versions of NN structures in the literature, which embed mathematical mappings related to solutions of the ADP-formulated problems called "adaptive critics" or "action-critic" networks, are discussed. Distinction between the two classes of ADP applications is pointed out. Furthermore, papers in "model-free" development and model-based neurocontrollers are reviewed in terms of their contributions to stability issues. Recent literature suggests that work in ADP-based feedback controllers with assured stability is growing in diverse forms. S. N. Balakrishnan, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2008 | Guest Editorial: Special Issue on Adaptive Dynamic Programming and Reinforcement Learning in Feedback ControlabstractThe 18 papers in this special issue focus on adaptive dynamic programming and reinforcement learning in feedback control. Frank L. Lewis, Derong Liu 0001, George G. Lendaris |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | Model-free Approximate Dynamic Programming Schemes for Linear SystemsabstractIn this paper, we present online model-free adaptive critic (AC) schemes based on approximate dynamic programming (ADP) to solve optimal control problems in both discrete-time and continuous-time domains for linear systems with unknown dynamics. In the discrete-time case, it is shown that the proposed ADP algorithm is in fact solving the underlying Generalized Algebraic Riccati Equation (GARE) of the corresponding optimal control problem or zero-sum game. In the continuous-time domain, an ADP scheme is introduced to solve for the underlying ARE of the optimal control problem. It is shown that this continuous-time ADP scheme is in fact a Quasi-Newton method to solve the ARE. In both time domains, the adaptive critic algorithms are easy to initialize since initial policies are not required to be stabilizing. It is also shown, on a power system control example, that both discrete-time and continuous-time approaches to ADP converge to the same continuous time optimal control solution provided that the utility function is appropriately chosen. Asma Al-Tamimi, Draguna L. Vrabie, Murad Abu-Khalaf, Frank L. Lewis |
IJCNN | 4 |
| 2007 | Fixed-Final-Time-Constrained Optimal Control of Nonlinear Systems Using Neural Network HJB ApproachabstractIn this paper, fixed-final time-constrained optimal control laws using neural networks (NNS) to solve Hamilton-Jacobi-Bellman (HJB) equations for general affine in the constrained nonlinear systems are proposed. An NN is used to approximate the time-varying cost function using the method of least squares on a predefined region. The result is an NN nearly -constrained feedback controller that has time-varying coefficients found by a priori offline tuning. Convergence results are shown. The results of this paper are demonstrated in two examples, including a nonholonomic system. Frank L. Lewis, Murad Abu-Khalaf |
IEEE Trans. Neural Networks | 2 |
| 2007 | Guest Editorial Special Issue on Neural Networks for Feedback Control SystemsabstractThe twenty-two papers in this special issue are devoted to neural networks for feedback control systems. Covers some of the following topics: reinforcement learning; applications; neurocontrol systems; discrete time systems; and network architectures and training methods. Frank L. Lewis, Jie Huang 0001, Thomas Parisini, Danil V. Prokhorov, Donald C. Wunsch II |
IEEE Trans. Neural Networks | 1 |
| 2007 | Energy-efficient wireless sensor network design and implementation for condition-based maintenanceabstractA new application architecture is designed for continuous, real-time, distributed wireless sensor networks. We develop a wireless sensor network for machinery condition-based maintenance (CBM) in small machinery spaces using commercially available products. We develop a hardware platform, networking architecture, and medium access communication protocol. We implement a single-hop sensor network to facilitate real-time monitoring and extensive data processing for machine monitoring. A new radio battery consumption model is presented and the battery consumption equation is used to select the most suitable topology and design an energy efficient communication protocol for wireless sensor networks. A new streamlined matrix formulation is developed that allows the base station to compute the best periodic sleep times for all the nodes in the network. We combine scheduling and contention to design a hybrid MAC protocol, which achieves 100% collision avoidance by using our modified RTS-CTS contention mechanism known as UC-TDMA protocol. A LabVIEW graphical user interface is described that allows for signal processing, including FFT, various moments, and kurtosis. A wireless CBM sensor network implementation on a heating and air conditioning plant is presented as a case study. Ankit Tiwari, Prasanna Ballal, Frank L. Lewis |
ACM Trans. Sens. Networks | 3 |
| 2007 | Adaptive Critic Designs for Discrete-Time Zero-Sum Games With Application to Hinfty ControlabstractIn this correspondence, adaptive critic approximate dynamic programming designs are derived to solve the discrete-time zero-sum game in which the state and action spaces are continuous. This results in a forward-in-time reinforcement learning algorithm that converges to the Nash equilibrium of the corresponding zero-sum game. The results in this correspondence can be thought of as a way to solve the Riccati equation of the well-known discrete-time H(infinity) optimal control problem forward in time. Two schemes are presented, namely: 1) a heuristic dynamic programming and 2) a dual-heuristic dynamic programming, to solve for the value function and the costate of the game, respectively. An H(infinity) autopilot design for an F-16 aircraft is presented to illustrate the results. Asma Al-Tamimi, Murad Abu-Khalaf, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2006 | Accuracy Based Adaptive Sampling and Multi-Sensor Scheduling for Collaborative Target TrackingabstractTracking is an essential capability in many wireless sensor network (WSN) applications. Due to the probabilistic nature of the target movement and the limited detection region and miss detection of sensor, this existing single sensor sensing scheme may result in tracking failure when a scheduled sensor fails to detect the target. We present an adaptive multi-sensor scheduling algorithm for collaborative target tracking in WSNs to improve the tracking reliability and power efficient. This proposed scheme determines the sampling intervals based on the predicted tracking accuracy, and selects a number of sensors to form a temporary tasking group for the next time step, based on a specified detection probability. One of the tasking group members is selected to be the leader node, which acts as a local temporary fusion center Jianyong Lin, Frank L. Lewis, Wendong Xiao, Lihua Xie 0001 |
ICARCV | 2 |
| 2006 | Adaptive Sampling using Non-linear EKF with Mobile Robotic Wireless Sensor NodesabstractThe use of robotics in distributed monitoring applications requires mobile wireless sensors that are deployed efficiently. Efficiency can be defined in multiple ways, such as in terms of the amount of energy expenditure, communication bandwidth or information content. A very important aspect of mobile sensor deployment includes sampling algorithms at location most likely to yield useful information about a field variable of interest. In this paper, we use inexpensive mobile robot nodes built in our lab (ARRI-Bots) as wireless sensor deployment agents, and we use them to demonstrate information efficient algorithms (e.g., "adaptive sampling"). Each mobile robot node is characterized by sensor measurement noise in addition to localization uncertainty. We use the extended Kalman filter (EKF) to derive quantitative information measures for sampling locations most likely to yield optimal information about the sampled field distribution. We present simulation and experimental results using this approach Dan O. Popa, Muhammad Faizan Mysorewala, Frank L. Lewis |
ICARCV | 3 |
| 2006 | EKF-based Adaptive Sampling with Mobile Robotic Sensor NodesabstractThe use of robotics in environmental monitoring applications requires distributed sensor systems optimized for effective estimation of relevant models subject to energy and environmental constraints. The mobile robot nodes are agents facilitating the repositioning of sensors in order to estimate a field distribution. This field distribution could be, for instance, water salinity in a lake, or air pollution over an industrial area. Each mobile robot node is characterized by sensor measurement noise in addition to localization uncertainty. This paper addresses an important problem for the robotic deployment of sensor networks, namely adaptive sampling (AS) by selection and repositioning of nodes in order to optimally estimate the parameters of distributed variable field models. The AS problem is posed as a sensor fusion problem within the extended Kalman filter (EKF) framework. We present simulation and experimental results of 2D deployment scenarios using low-cost mobile sensor robots developed in our lab Dan O. Popa, Muhammad Faizan Mysorewala, Frank L. Lewis |
IROS | 3 |
| 2006 | Data-Logging and Supervisory Control in Wireless Sensor NetworksabstractWireless sensor networks (WSN) are increasingly used in a multitude of applications such as environmental and structural health monitoring, and condition-based maintenance. Even though the sensors collect a vast amount of data, only a tiny fraction of this data may be useful. This paper introduces and implements a data-logging & supervisory control architecture to manage the information gathered by the WSN and make decisions based on this information. We present an application module which would be able to effectively manipulate the sensor data and to support a discrete event controller (DEC). The DEC is responsible for generating rule-based tasks in order to address information-centric issues such as data-logging, alarm & event reporting and security. A combined data-logging and supervisory control framework (DSC) is proposed to address data pre-processing challenges such as acquiring and recording signals, online analysis, offline analysis, report generation, and data sharing Aditya N. Das, Frank L. Lewis, Dan O. Popa |
SNPD | 2 |
| 2006 | Deployment Algorithms and In-Door Experimental Vehicles for Studying Mobile Wireless Sensor NetworksabstractWireless communication has been traditionally used in robotics to transmit sensory and telemetry information between a robot and a base station. Because research in mobile robotics has typically focused on navigation, mapping and sensor fusion, network oriented problems such as communication bandwidth optimization, coverage and fault tolerance are not usually considered in this context. The motivation behind this research is formulating and solving combined robot navigation issues (such as obstacle avoidance, environment mapping and coverage) with sensor network issues (such as congestion control, routing and node energy minimization). In this paper we present several types of algorithms for mobile wireless sensor nodes (MWSN) as well as experimental results with a fleet of mobile robots and sensors in our lab. The algorithms include adaptive sampling (AS) for distributed field estimation, potential fields (PF) for communication bandwidth optimization, and a discrete event controller (DEC) for mission planning Muhammad Faizan Mysorewala, Dan O. Popa, Vincenzo Giordano, Frank L. Lewis |
SNPD | 4 |
| 2006 | Supervisory control of mobile sensor networks: math formulation, simulation, and implementationabstractThis paper uses a novel discrete-event controller (DEC) for the coordination of cooperating heterogeneous wireless sensor networks (WSNs) containing both unattended ground sensors (UGSs) and mobile sensor robots. The DEC sequences the most suitable tasks for each agent and assigns sensor resources according to the current perception of the environment. A matrix formulation makes this DEC particularly useful for WSN, where missions change and sensor agents may be added or may fail. WSN have peculiarities that complicate their supervisory control. Therefore, this paper introduces several new tools for DEC design and operation, including methods for generating the required supervisory matrices based on mission planning, methods for modifying the matrices in the event of failed nodes, or nodes entering the network, and a novel dynamic priority assignment weighting approach for selecting the most appropriate and useful sensors for a given mission task. The resulting DEC represents a complete dynamical description of the WSN system, which allows a fast programming of deployable WSN, a computer simulation analysis, and an efficient implementation. The DEC is actually implemented on an experimental wireless-sensor-network prototyping system. Both simulation and experimental results are presented to show the effectiveness and versatility of the developed control architecture. Vincenzo Giordano, Prasanna Ballal, Frank L. Lewis, Biagio Turchiano, Jing Bing Zhang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2005 | Neural networks for feedback control of robots and dynamical systemsabstractSummary form only given. Over the past years we have developed a family of feedback controllers that can confront these systems using neural networks as the basic control block structure. The learning abilities of neural networks considered as intelligent systems allow these controllers to learn online and improve their performance through tuning of the weights. We present a catalog of neural network controllers designed based on feedback linearization, backstepping, singular perturbations, and dynamic inversion techniques. These neural network controllers are all tuned online in real time based on the system errors. Then, we present some recent results on H-infinity feedback control for constrained input nonlinear systems. The constraints on the input to the system are encoded via a quasi-norm that allows nonquadratic supply rates along with dissipativity theory to formulate the robust output feedback control problem using Hamilton-Jacobi-Isaac (HJI) equations. An iterative solution technique based on a game theoretic interpretation is presented. To provide a computationally tractable controller design method, the solution is approximated at each iteration with a neural network. The result is a closed loop control based on a neural net that has been tuned a priori offline. Frank L. Lewis |
IJCNN | 1 |
| 2004 | Wireless sensor network for machine condition based maintenanceabstractA new application architecture is designed for continuous, real-time, distributed wireless sensor networks. We develop a wireless sensor network for machinery condition-based maintenance (CBM) using commercially available products, including a hardware platform, networking architecture, and medium access communication protocol. We implement a single-hop sensor network to facilitate real-time monitoring and extensive data processing for machine monitoring. A LabVIEW graphical user interface is described that allows for signal processing, including FFT, various moments, and kurtosis. A wireless CBM sensor network implementation on a heating and air conditioning plant is presented as a case study. Ankit Tiwari, Frank L. Lewis, Shuzhi Sam Ge |
ICARCV | 2 |
| 2003 | A controller-observer scheme for a robotic cellabstractThis paper presents a controller-observer scheme for Discrete Event Systems (DES) and its implementation into a robotic cell. In this approach the observer provides the current state of a system having some internal states that are not measurable from the outputs. The controller confines the system behavior into a specified behavior. This scheme is complemented with a decision-making module that determines the best transition to fire. All the entities of this scheme and the system are modeled using interpreted Petri nets (IPN). Raul Campos-Rodriguez, Ernesto López-Mellado, Antonio Ramírez-Treviño, José Mireles, Frank L. Lewis |
SMC | 5 |
| 2003 | Two-time scale fuzzy logic controller of flexible link robot arm
Frank L. Lewis |
Fuzzy Sets Syst. | 2 |
| 2003 | Neural-network predictive control for nonlinear dynamic systems with time-delayabstractA new recurrent neural-network predictive feedback control structure for a class of uncertain nonlinear dynamic time-delay systems in canonical form is developed and analyzed. The dynamic system has constant input and feedback time delays due to a communications channel. The proposed control structure consists of a linearized subsystem local to the controlled plant and a remote predictive controller located at the master command station. In the local linearized subsystem, a recurrent neural network with on-line weight tuning algorithm is employed to approximate the dynamics of the time-delay-free nonlinear plant. No linearity in the unknown parameters is required. No preliminary off-line weight learning is needed. The remote controller is a modified Smith predictor that provides prediction and maintains the desired tracking performance; an extra robustifying term is needed to guarantee stability. Rigorous stability proofs are given using Lyapunov analysis. The result is an adaptive neural net compensation scheme for unknown nonlinear systems with time delays. A simulation example is provided to demonstrate the effectiveness of the proposed control strategy. Jin-Quan Huang, Frank L. Lewis |
IEEE Trans. Neural Networks | 2 |
| 2002 | Neural-network approximation of piecewise continuous functions: application to friction compensationabstractOne of the most important properties of neural nets (NNs) for control purposes is the universal approximation property. Unfortunately,, this property is generally proven for continuous functions. In most real industrial control systems there are nonsmooth functions (e.g., piecewise continuous) for which approximation results in the literature are sparse. Examples include friction, deadzone, backlash, and so on. It is found that attempts to approximate piecewise continuous functions using smooth activation functions require many NN nodes and many training iterations, and still do not yield very good results. Therefore, a novel neural-network structure is given for approximation of piecewise continuous functions of the sort that appear in friction, deadzone, backlash, and other motion control actuator nonlinearities. The novel NN consists of neurons having standard sigmoid activation functions, plus some additional neurons having a special class of nonsmooth activation functions termed "jump approximation basis function." Two types of nonsmooth jump approximation basis functions are determined- a polynomial-like basis and a sigmoid-like basis. This modified NN with additional neurons having "jump approximation" activation functions can approximate any piecewise continuous function with discontinuities at a finite number of known points. Applications of the new NN structure are made to rigid-link robotic systems with friction nonlinearities. Friction is a nonlinear effect that can limit the performance of industrial control systems; it occurs in all mechanical systems and therefore is unavoidable in control systems. It can cause tracking errors, limit cycles, and other undesirable effects. Often, inexact friction compensation is used with standard adaptive techniques that require models that are linear in the unknown parameters. It is shown here how a certain class of augmented NN, capable of approximating piecewise continuous functions, can be used for friction compensation. Rastko R. Selmic, Frank L. Lewis |
IEEE Trans. Neural Networks | 2 |
| 2000 | Backlash Compensation in Discrete Time Nonlinear Systems Using Dynamic Inversion by Neural NetworksabstractA dynamics inversion compensation scheme is designed for control of nonlinear discrete-time systems with input backlash. The compensator uses the backstepping technique with neural networks (NN) for inverting the backlash nonlinearity in the feedforward path. The technique provides a general procedure for using NN to determine the dynamics pre-inverse of an invertible discrete time dynamical system. A discrete-time tuning algorithm is given for the NN weights so that the backlash compensation scheme becomes adaptive, guaranteeing bounded tracking and backlash errors, and also bounded parameter estimates. A rigorous proof of stability and performance is given and a simulation example verifies the performance. Unlike standard discrete-time adaptive control techniques, no certainty equivalence assumption is needed. Javier Campos, Frank L. Lewis, Rastko R. Selmic |
ICRA | 2 |
| 2000 | Robust backstepping control of a class of nonlinear systems using fuzzy logic
Sarangapani Jagannathan, Frank L. Lewis |
Inf. Sci. | 2 |
| 2000 | Robust backstepping control of induction motors using neural networksabstractIn this paper, we present a new robust control technique for induction motors using neural networks (NNs). The method is systematic and robust to parameter variations. Motivated by the well-known backstepping design technique, we first treat certain signals in the system as fictitious control inputs to a simpler subsystem. A two-layer NN is used in this stage to design the fictitious controller. Then we apply a second two-layer NN to robustly realize the fictitious NN signals designed in the previous step. A new tuning scheme is proposed which can guarantee the boundedness of tracking error and weight updates. A main advantage of our method is that we do not require regression matrices, so that no preliminary dynamical analysis is needed. Another salient feature of our NN approach is that the off-line learning phase is not needed. Full state feedback is needed for implementation. Load torque and rotor resistance can be unknown but bounded. Chiman Kwan, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2000 | Optimal design of CMAC neural-network controller for robot manipulatorsabstractThis paper is concerned with the application of quadratic optimization for motion control to feedback control of robotic systems using cerebellar model arithmetic computer (CMAC) neural networks. Explicit solutions to the Hamilton-Jacobi-Bellman (H-J-B) equation for optimal control of robotic systems are found by solving an algebraic Riccati equation. It is shown how the CMAC can cope with nonlinearities through optimization with no preliminary off-line learning phase required. The adaptive-learning algorithm is derived from Lyapunov stability analysis, so that both system-tracking stability and error convergence can be guaranteed in the closed-loop system. The filtered-tracking error or critic gain and the Lyapunov function for the nonlinear analysis are derived from the user input in terms of a specified quadratic-performance index. Simulation results from a two-link robot manipulator show the satisfactory performance of the proposed control schemes even in the presence of large modeling uncertainties and external disturbances. Young Ho Kim, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2000 | Robust backstepping control of nonlinear systems using neural networksabstractA controller is proposed for the robust backstepping control of a class of general nonlinear systems using neural networks (NNs). A tuning scheme is proposed which can guarantee the boundedness of tracking error and weight updates. Compared with adaptive backstepping control schemes, we do not require the unknown parameters to be linear parametrizable. No regression matrices are needed, so no preliminary dynamical analysis is needed. One salient feature of our NN approach is that there is no need for the off-line learning phase. Three nonlinear systems, including a one-link robot, an induction motor, and a rigid-link flexible-joint robot, were used to demonstrate the effectiveness of the proposed scheme. Chiman Kwan, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2000 | Computational complexity of determining resource loops in re-entrant flow linesabstractThis paper presents a comparison study of the computational complexity of the general job shop protocol and the more structured flow line protocol in a flexible manufacturing system. It is shown that the representative problem of finding resource invariants is NP-complete in the case of the job shop, while in the flow line case it admits a closed form solution. The importance of correctly selecting part flow and job routing protocols in flexible manufacturing systems to reduce complexity is thereby conclusively demonstrated. Frank L. Lewis, Bill G. Horne, Chaouki T. Abdallah |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 1999 | Deadzone compensation in discrete time using adaptive fuzzy logicabstractA fuzzy logic (FL) compensator is designed for control of nonlinear discrete-time systems with input deadzone. The classification property of FL systems makes them a natural candidate for the rejection of errors induced by the deadzone, which has regions in which it behaves differently. A discrete-time tuning algorithm is given for the FL parameters so that the deadzone compensation scheme becomes adaptive, guaranteeing bounded tracking errors and parameter estimates. A rigorous proof of stability and performance is given and a simulation example verifies performance. Unlike standard discrete-time adaptive control techniques, no certainty equivalence assumption is needed. Javier Campos, Frank L. Lewis |
IEEE Trans. Fuzzy Syst. | 2 |
| 1999 | Neural network output feedback control of robot manipulatorsabstractA robust neural network output feedback scheme is developed for the motion control of robot manipulators without measuring joint velocities. A neural network observer is presented to estimate the joint velocities. It is shown that all the signals in a closed-loop system composed of a robot, an observer, and a controller is uniformly ultimately bounded. This amounts to a separation principle for the design of nonlinear dynamic trackers for robotic systems. The neural network weights in both the observer and the controller are tuned online, with no off-line learning phase required. No exact knowledge of the robot dynamics is required so that the neural network controller is model-free and so applicable to a class of nonlinear systems which have a similar structure to robot manipulators. Simulation results on 2-link robot manipulator are reported to show the performance of the proposed output feedback control scheme. Young Ho Kim, Frank L. Lewis |
IEEE Trans. Robotics Autom. | 2 |
| 1999 | Hybrid control for a class of underactuated mechanical systemsabstractThis paper considers a stabilizing hybrid scheme to control a class of underactuated mechanical systems. The hybrid controller consists of a collection of state feedback controllers plus a discrete-event supervisor. When the continuous-state hits a switching boundary, a new controller is applied to the plant. Lyapunov theory is used to determine the switching boundaries and to guarantee the stability of the closed-loop hybrid system. This approach is applied to the well-known swing up and balancing control problem of the inverted pendulum. Rafael Fierro, Frank L. Lewis, J. Andy Lowe |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 1998 | A Fuzzy System Compensator for BacklashabstractWe design a fuzzy system to compensate the delays due to backlash nonlinearity. The fuzzy compensator is constructed from some common-sense rules. We consider the general case of the unknown backlash parameter and develop an adaptation algorithm to estimate the backlash parameter online. We prove that under certain conditions the fuzzy compensator with the adaptation algorithm guarantees that the backlash output converges to the desired trajectory. Simulation and hardware implementation results show that the fuzzy compensator is robust to the estimation errors in the backlash parameters. Application to an industrial CNC machine tool is described. Tim K. T. Woo, Li-Xin Wang, Frank L. Lewis, Zexiang Li 0001 |
ICRA | 3 |
| 1998 | Control of a nonholonomic mobile robot using neural networksabstractA control structure that makes possible the integration of a kinematic controller and a neural network (NN) computed-torque controller for nonholonomic mobile robots is presented. A combined kinematic/torque control law is developed using backstepping and stability is guaranteed by Lyapunov theory. This control algorithm can be applied to the three basic nonholonomic navigation problems: tracking a reference trajectory, path following, and stabilization about a desired posture. Moreover, the NN controller proposed in this work can deal with unmodeled bounded disturbances and/or unstructured unmodeled dynamics in the vehicle. On-line NN weight tuning algorithms do no require off-line learning yet guarantee small tracking errors and bounded control signals are utilized. Rafael Fierro, Frank L. Lewis |
IEEE Trans. Neural Networks | 2 |
| 1998 | Robust neural-network control of rigid-link electrically driven robotsabstractA robust neural-network (NN) controller is proposed for the motion control of rigid-link electrically driven (RLED) robots. Two-layer NN's are used to approximate two very complicated nonlinear functions. The main advantage of our approach is that the NN weights are tuned on-line, with no off-line learning phase required. Most importantly, we can guarantee the uniformly ultimately bounded (UUB) stability of tracking errors and NN weights. When compared with standard adaptive robot controllers, we do not require lengthy and tedious preliminary analysis to determine a regression matrix. The controller can be regarded as a universal reusable controller because the same controller can be applied to any type of RLED robots without any modifications. Chiman Kwan, Frank L. Lewis, Darren M. Dawson |
IEEE Trans. Neural Networks | 2 |
| 1997 | Deadzone compensation in motion control systems using adaptive fuzzy logic controlabstractA deadzone compensator is designed for industrial positioning systems using a fuzzy logic (FL) controller. The classification property of FL systems makes them a natural candidate for the rejection of errors induced by the deadzone, which has regions in which it behaves differently. A tuning algorithm is given for the FL parameters, so that the deadzone compensation scheme becomes adaptive, guaranteeing small tracking errors and bounded parameter estimates. The adaptive FL deadzone compensator is implemented on an actual industrial CNC machine tool to show its efficacy. Tim K. T. Woo, Frank L. Lewis, Li-Xin Wang, Zexiang Li 0001 |
ICRA | 2 |
| 1997 | A framework for hybrid control designabstractThis paper presents a hybrid system framework which considers simultaneously the control and decision-making issues. This reconfigurable framework can accommodate a wide range of situations, from aircraft control systems to mobile manipulators. A continuous-state plant is supervised by a discrete-event system which is based on a theory of linked finite state machines. The composite system is viewed as an iterative process where a task is carried out by changing the structure of the continuous-state plant. An algorithm for a hybrid control design is provided and illustrated through a mobile manipulator example. Rafael Fierro, Frank L. Lewis |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 1996 | Adaptive-fuzzy logic control of robot manipulatorsabstractThis paper presents a methodology for the design of fuzzy logic controllers (FLC) that guarantees prescribed performance for a general robotic system. It is shown that the proposed online learning algorithm learns the stabilizing membership functions (MFs) online from initial MFs that are selected using simple design criteria. From a practical standpoint, the controller structure leads to efficient implementation and fills a void that existed in the lack of repeatable design methodologies for FLC implementation. Sesh Commuri, Frank L. Lewis |
ICRA | 2 |
| 1996 | Discrete-time adaptive fuzzy logic control of robotic systemsabstractThis paper demonstrates tracking control of a class of feedback linearizable unknown nonlinear dynamical systems, such as a robotic systems, using a discrete-time fuzzy logic controller (FLC). Designing a discrete-time FLC is significant because almost all FLC's are implemented on digital computers. A repeatable design algorithm and a stability proof are examined for an adaptive fuzzy logic controller that uses fuzzy basis functions based on the fuzzy system, unlike most standard adaptive control approaches which use basis vectors depending on the unknown plant. An /spl epsiv/-modification sort of approach to adapt the fuzzy system parameters is examined. Using this adaptive fuzzy logic controller, uniform ultimate boundedness of the closed-loop signals is presented and that the controller achieves tracking. In fact, the fuzzy system designed is a model-free universal fuzzy controller that works for any system in the given class of systems. Sarangapani Jagannathan, Frank L. Lewis |
ICRA | 2 |
| 1996 | Output feedback control of rigid robots using dynamic neural networksabstractA robust neural network (NN) output feedback scheme is proposed for the motion control of rigid robots. A dynamic NN observer is presented to estimate the joint speeds. The stability of a closed-loop system composed of a robot, a NN observer, and a NN controller is proven. The NN weights in both the observer and the controller are tuned online, with no off-line learning phase required. Most importantly, we can guarantee the boundness of the estimated velocities, the position tracking errors, and the NN weights. Also no exact knowledge of the robot dynamics is required so that the NN controller is model-free and so applicable to any type of rigid robot. When compared with adaptive-type controllers, we do not require persistent excitation conditions, linearity in the unknown system parameters, or the tedious computation of a regression matrix. Thus the new NN approach represents an improvement over adaptive techniques. Young Ho Kim, Frank L. Lewis |
ICRA | 2 |
| 1996 | Multilayer discrete-time neural-net controller with guaranteed performanceabstractA family of novel multilayer discrete-time neural-net (NN) controllers is presented for the control of a class of multi-input multi-output (MIMO) dynamical systems. The neural net controller includes modified delta rule weight tuning and exhibits a learning while-functioning-features. The structure of the NN controller is derived using a filtered error/passivity approach. Linearity in the parameters is not required and certainty equivalence is not used. This overcomes several limitations of standard adaptive control. The notion of persistency of excitation (PE) for multilayer NN is defined and explored. New online improved tuning algorithms for discrete-time systems are derived, which are similar to sigma or epsilon-modification for the case of continuous-time systems, that include a modification to the learning rate parameter plus a correction term. These algorithms guarantee tracking as well as bounded NN weights in nonideal situations so that PE is not needed. An extension of these novel weight tuning updates to NN with an arbitrary number of hidden layers is discussed. The notions of discrete-time passive NN, dissipative NN, and robust NN are introduced. The NN makes the closed-loop system passive. Sarangapani Jagannathan, Frank L. Lewis |
IEEE Trans. Neural Networks | 2 |
| 1996 | Multilayer neural-net robot controller with guaranteed tracking performanceabstractA multilayer neural-net (NN) controller for a general serial-link rigid robot arm is developed. The structure of the NN controller is derived using a filtered error/passivity approach. No off-line learning phase is needed for the proposed NN controller and the weights are easily initialized. The nonlinear nature of the NN, plus NN functional reconstruction inaccuracies and robot disturbances, mean that the standard delta rule using backpropagation tuning does not suffice for closed-loop dynamic control. Novel online weight tuning algorithms, including correction terms to the delta rule plus an added robust signal, guarantee bounded tracking errors as well as bounded NN weights. Specific bounds are determined, and the tracking error bound can be made arbitrarily small by increasing a certain feedback gain. The correction terms involve a second-order forward-propagated wave in the backpropagation network. New NN properties including the notions of a passive NN, a dissipative NN, and a robust NN are introduced. Frank L. Lewis, Aydin Yesildirek, Kai Liu 0020 |
IEEE Trans. Neural Networks | 1 |
| 1995 | Neural net robot controller with guaranteed tracking performanceabstractA neural net (NN) controller for a general serial-link robot arm is developed. The NN has two layers so that linearity in the parameters holds, but the "net functional reconstruction error" and robot disturbance input are taken as nonzero. The structure of the NN controller is derived using a filtered error/passivity approach, leading to new NN passivity properties. Online weight tuning algorithms including a correction term to backpropagation, plus an added robustifying signal, guarantee tracking as well as bounded NN weights. The NN controller structure has an outer tracking loop so that the NN weights are conveniently initialized at zero, with learning occurring online in real-time. It is shown that standard backpropagation, when used for real-time closed-loop control, can yield unbounded NN weights if (1) the net cannot exactly reconstruct a certain required control function or (2) there are bounded unknown disturbances in the robot dynamics. The role of persistency of excitation is explored. Frank L. Lewis, Kai Liu 0020, Aydin Yesildirek |
IEEE Trans. Neural Networks | 1 |
| 1990 | Decentralized continuous robust controller for mobile robotsabstractA composite system model for the wheeled mobile robot is proposed. The overall system can be thought of as a composite of two subsystems; one is the vehicle with m degrees of freedom, and the other is the robot arm with n degrees of freedom. The interconnections between them are unknown interactive forces. To overcome the effects of the unknown interactive forces and some mismatches between the estimated parameter values and the actual ones, a decentralized continuous robust controller is proposed. Conditions are derived for stability when the interconnections are not known. The theoretical analysis and computer simulation results show that the performance of the continuous robust controller is much better than that of the nonrobust controller. An interesting conclusion is that one can choose arbitrarily large weighting gains to decrease the trajectory errors without increasing the magnitude of the physical control torques.> Kai Liu 0020, Frank L. Lewis |
ICRA | 2 |