VLDB 2026 Research / reviewers in the wild / expert
Kyriakos G. Vamvoudakis
dblp:01/8083
· DBLP profile ↗
28ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0003-1978-4848ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Safe Physics-Informed Machine Learning for Optimal Predefined-Time Stabilization: A Lyapunov-Based ApproachabstractIn this article, we introduce the notion of safe predefined-time stability and address an optimal safe predefined-time stabilization problem. In particular, safe predefined-time stability characterizes parameter-dependent nonlinear dynamical systems whose trajectories starting in a given set of admissible states remain in the set of admissible states for all time and converge to an equilibrium point in a predefined time. Furthermore, we provide a Lyapunov theorem establishing sufficient conditions for safe predefined-time stability. We address the optimal safe predefined-time stabilization problem by synthesizing feedback controllers that guarantee closed-loop system safe predefined-time stability while optimizing a given performance measure. Specifically, safe predefined-time stability of the closed-loop system is guaranteed via a Lyapunov function satisfying a differential inequality while simultaneously serving as a solution to the steady-state Hamilton-Jacobi-Bellman (HJB) equation ensuring optimality. Given that the HJB equation is generally difficult to solve, we develop a physics-informed machine learning-based algorithm for learning the safely predefined-time stabilizing solution to the steady-state HJB equation. Several simulation results are provided to demonstrate the efficacy of the proposed approach. Nick-Marios T. Kokolakis, Zhen Zhang 0029, Shanqing Liu, Kyriakos G. Vamvoudakis, Jérôme Darbon, George Em Karniadakis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Guest Editorial Special Issue on Learning From Imperfect Data for Industrial AutomationabstractWith the rapid development of advanced sensing, communication, and the industrial Internet of Things, it has become much easier to obtain, transmit, and, store a massive amount of real-world data. However, imperfect data is inevitable in real-world systems, such as the existence of outliers, contaminated, incomplete, inaccurate, and even missing information in the data. This phenomenon is called data imperfection, which usually makes traditional datadriven modeling and automation methods either unfeasible or ending at undesired inaccuracies. This has been a wellknown challenge to data-driven methods when applied to real-world systems, such as process industry, manufacturing, energy networks, and transportation systems. Ping Zhou 0003, Xuewu Dai, Kyriakos G. Vamvoudakis, Jan Faigl, Hong Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Cooperative Finitely Excited Learning for Dynamical GamesabstractIn this article, we propose a way to enhance the learning framework for zero-sum games with dynamics evolving in continuous time. In contrast to the conventional centralized actor-critic learning, a novel cooperative finitely excited learning approach is developed to combine the online recorded data with instantaneous data for efficiency. By using an experience replay technique for each agent and distributed interaction amongst agents, we are able to replace the classical persistent excitation condition with an easy-to-check cooperative excitation condition. This approach also guarantees the consensus of the distributed actor-critic learning on the solution to the Hamilton-Jacobi-Isaacs (HJI) equation. It is shown that both the closed-loop stability of the equilibrium point and convergence to the Nash equilibrium can be guaranteed. Simulation results demonstrate the efficacy of this approach compared to previous methods. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Frank L. Lewis |
IEEE Trans. Cybern. | 3 |
| 2024 | UWB Ranging and IMU Data Fusion: Overview and Nonlinear Stochastic Filter for Inertial NavigationabstractThis paper proposes a nonlinear stochastic complementary filter design for inertial navigation that takes advantage of a fusion of Ultra-wideband (UWB) and Inertial Measurement Unit (IMU) technology ensuring semi-global uniform ultimate boundedness (SGUUB) of the closed loop error signals in mean square. The proposed filter estimates the vehicle’s orientation, position, linear velocity, and noise covariance. The filter is designed to mimic the nonlinear navigation motion kinematics and is posed on a matrix Lie Group, the extended form of the Special Euclidean Group$\mathbb {SE}_{2}\left ({3}\right)$. The Lie Group based structure of the proposed filter provides unique and global representation avoiding singularity (a common shortcoming of Euler angles) as well as non-uniqueness (a common limitation of unit-quaternion). Unlike Kalman-type filters, the proposed filter successfully addresses IMU measurement noise considering unknown upper-bounded covariance. Although the navigation estimator is proposed in a continuous form, the discrete version is also presented. Moreover, the unit-quaternion implementation has been provided in the Appendix. Experimental validation performed using a publicly available real-world six-degrees-of-freedom (6 DoF) flight dataset obtained from an unmanned Micro Aerial Vehicle (MAV) illustrating the robustness of the proposed navigation technique. Hashim A. Hashim, Abdelrahman E. E. Eltoukhy, Kyriakos G. Vamvoudakis |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Guest Editorial Special Issue on Reinforcement Learning-Based Control: Data-Efficient and Resilient MethodsabstractAs an important branch of machine learning, reinforcement learning (RL) has proved its efficiency in many emerging applications in science and engineering. A remarkable advantage of RL is that it enables agents to maximize their cumulative rewards through online exploration and interactions with unknown (or partially unknown) and uncertain environments, which is regarded as a variant of data-driven adaptive optimal control methods. However, the successful implementation of RL-based control systems usually relies on a good quantity of online data due to its data-driven nature. Therefore, it is imperative to develop data-efficient RL methods for control systems to reduce the required number of interactions with the external environment. Moreover, network-aware issues, such as cyberattacks, dropout packet and communication latency, and actuator and sensor faults, are challenging conundrums that threaten the safety, security, stability, and reliability of network control systems. Consequently, it is significant to develop safe and resilient RL mechanisms. Weinan Gao, Na Li 0002, Kyriakos G. Vamvoudakis, F. Richard Yu, Zhong-Ping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Adaptive Neural Network Stochastic-Filter-Based Controller for Attitude Tracking With Disturbance RejectionabstractThis article proposes a real-time neural network (NN) stochastic filter-based controller on the Lie group of the special orthogonal group [Formula: see text] as a novel approach to the attitude tracking problem. The introduced solution consists of two parts: a filter and a controller. First, an adaptive NN-based stochastic filter is proposed, which estimates attitude components and dynamics using measurements supplied by onboard sensors directly. The filter design accounts for measurement uncertainties inherent to the attitude dynamics, namely, unknown bias and noise corrupting angular velocity measurements. The closed-loop signals of the proposed NN-based stochastic filter have been shown to be semiglobally uniformly ultimately bounded (SGUUB). Second, a novel control law on [Formula: see text] coupled with the proposed estimator is presented. The control law addresses unknown disturbances. In addition, the closed-loop signals of the proposed filter-based controller have been shown to be SGUUB. The proposed approach offers robust tracking performance by supplying the required control signal given data extracted from low-cost inertial measurement units. While the filter-based controller is presented in continuous form, the discrete implementation is also presented. In addition, the unit-quaternion form of the proposed approach is given. The effectiveness and robustness of the proposed filter-based controller are demonstrated using its discrete form and considering low sampling rate, high initialization error, high level of measurement uncertainties, and unknown disturbances. Hashim A. Hashim, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Safety-Aware Pursuit-Evasion Games in Unknown Environments Using Gaussian Processes and Finite-Time Convergent Reinforcement LearningabstractThis article develops a safe pursuit-evasion game for enabling finite-time capture, optimal performance as well as adaptation to an unknown cluttered environment. The pursuit-evasion game is formulated as a zero-sum differential game wherein the pursuer seeks to minimize its relative distance to the target while the evader attempts to maximize it. A critic-only reinforcement learning (RL)-based algorithm is then proposed for learning online and in finite time the pursuit-evasion policies and thus enabling finite-time capture of the evader. Safety is ensured by means of barrier functions associated with the obstacles, which are integrated into the running cost. Using Gaussian processes (GPs), a learning-based mechanism is devised for safely learning the unknown environment. Simulation results illustrate the efficacy of the proposed approach. Nick-Marios T. Kokolakis, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Online and Robust Intermittent Motion Planning in Dynamic and Changing EnvironmentsabstractIn this article, we propose RRT- , an online and intermittent kinodynamic motion planning framework for dynamic environments with unknown robot dynamics and unknown disturbances. We leverage RRT for global path planning and rapid replanning to produce waypoints as a sequence of boundary-value problems (BVPs). For each BVP, we formulate a finite-horizon, continuous-time zero-sum game, where the control input is the minimizer, and the worst case disturbance is the maximizer. We propose a robust intermittent Q-learning controller for waypoint navigation with completely unknown system dynamics, external disturbances, and intermittent control updates. We execute a relaxed persistence of excitation technique to guarantee that the Q-learning controller converges to the optimal controller. We provide rigorous Lyapunov-based proofs to guarantee the closed-loop stability of the equilibrium point. The effectiveness of the proposed RRT- is illustrated with Monte Carlo numerical experiments in numerous dynamic and changing environments. George P. Kontoudis, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Decentralized and Privacy-Preserving Learning of Approximate Stackelberg Solutions in Energy Trading Games With Demand Response AggregatorsabstractIn the pathway to 2030 electricity generation decarbonization and 2050 net-zero economies, scalable integration of distributed load can support environmental goals and also help alleviate smart grid operational issues through its electricity market participation. In this work, a novel Stackelberg game theoretic framework is proposed for trading the energy bidirectionally between the demand-response (DR) aggregator and the prosumers (distributed load). This formulation allows for flexible energy arbitrage and additional monetary rewards while ensuring that the prosumers’ desired daily energy demand is met. Then, a scalable (linear with the number of prosumers and the number of learning samples), the decentralized privacy-preserving algorithm is proposed to find approximate equilibria with online sampling and learning of the prosumers’ cumulative best response, which finds applications beyond this energy game. Moreover, cost bounds are provided on the quality of the approximate equilibrium solution. Finally, the real data from the California day-ahead market and the UC Davis campus building energy demands are utilized to demonstrate the efficacy of the proposed framework and the algorithm. Stella Kampezidou, Justin K. Romberg, Kyriakos G. Vamvoudakis, Dimitri N. Mavris |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Nonlinear Deterministic Observer for Inertial Navigation Using Ultra-Wideband and IMU Sensor FusionabstractNavigation in Global Positioning Systems (GPS)-denied environments requires robust estimators reliant on fusion of inertial sensors able to estimate rigid-body's orientation, position, and linear velocity. Ultra-wideband (UWB) and Inertial Measurement Unit (IMU) represent low-cost measurement technology that can be utilized for successful Inertial Navigation. This paper presents a nonlinear deterministic navigation observer in a continuous form that directly employs UWB and IMU measurements. The estimator is developed on the extended Special Euclidean Group$\mathbb{SE}_{2}$(3) and ensures exponential convergence of the closed loop error signals starting from almost any initial condition. The discrete version of the proposed observer is tested using a publicly available real-world dataset of a drone flight. Hashim A. Hashim, Abdelrahman E. E. Eltoukhy, Kyriakos G. Vamvoudakis, Mohammed I. Abouheaf |
IROS | 3 |
| 2023 | Recursive Reasoning With Reduced Complexity and Intermittency for Nonequilibrium Learning in Stochastic GamesabstractIn this article, we propose a computationally and communicationally efficient approach for decision-making in nonequilibrium stochastic games. In particular, due to the inherent complexity of computing Nash equilibria, as well as the innate tendency of agents to choose nonequilibrium strategies, we construct two models of bounded rationality based on recursive reasoning. In the first model, named level-$k$thinking, each agent assumes that everyone else has a cognitive level immediately lower than theirs and—given such an assumption—chooses their policy to be a best response to them. In the second model, named cognitive hierarchy, each agent conjectures that the rest of the agents have a cognitive level that is lower than theirs, but follows a distribution instead of being deterministic. To explicitly compute the boundedly rational policies, a level-recursive algorithm and a level-paralleled algorithm are constructed, where the latter one can have an overall reduced computational complexity. To further reduce the complexity in the communication layer, modifications of the proposed nonequilibrium strategies are presented, which do not require the action of a boundedly rational agent to be updated at each step of the stochastic game. Simulations are performed that demonstrate our results. Filippos Fotiadis, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Intermittent Learning Through Operant Conditioning for Cyber-Physical SystemsabstractThis article presents a novel scheme, namely, an intermittent learning scheme based on Skinner's operant conditioning techniques that approximates the optimal policy while decreasing the usage of the communication buses transferring information. While traditional reinforcement learning schemes continuously evaluate and subsequently improve, every action taken by a specific learning agent based on received reinforcement signals, this form of continuous transmission of reinforcement signals and policy improvement signals can cause overutilization of the system's inherently limited resources. Moreover, the highly complex nature of the operating environment for cyber-physical systems (CPSs) creates a gap for malicious individuals to corrupt the signal transmissions between various components. The proposed schemes will increase uncertainty in the learning rate and the extinction rate of the acquired behavior of the learning agents. In this article, we investigate the use of fixed/variable interval and fixed/variable ratio schedules in CPSs along with their rate of success and loss in their optimal behavior incurred during intermittent learning. Simulation results show the efficacy of the proposed approach. Prachi Pratyusha Sahoo, Aris Kanellopoulos, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Hamiltonian-Driven Adaptive Dynamic Programming With Approximation ErrorsabstractIn this article, we consider an iterative adaptive dynamic programming (ADP) algorithm within the Hamiltonian-driven framework to solve the Hamilton-Jacobi-Bellman (HJB) equation for the infinite-horizon optimal control problem in continuous time for nonlinear systems. First, a novel function, "min-Hamiltonian," is defined to capture the fundamental properties of the classical Hamiltonian. It is shown that both the HJB equation and the policy iteration (PI) algorithm can be formulated in terms of the min-Hamiltonian within the Hamiltonian-driven framework. Moreover, we develop an iterative ADP algorithm that takes into consideration the approximation errors during the policy evaluation step. We then derive a sufficient condition on the iterative value gradient to guarantee closed-loop stability of the equilibrium point as well as convergence to the optimal value. A model-free extension based on an off-policy reinforcement learning (RL) technique is also provided. Finally, numerical results illustrate the efficacy of the proposed framework. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Wei He 0001, Cheng-Zhong Xu 0001, Donald C. Wunsch II |
IEEE Trans. Cybern. | 3 |
| 2021 | A Secure Control Learning Framework for Cyber-Physical Systems Under Sensor and Actuator AttacksabstractIn this article, we develop a learning-based secure control framework for cyber-physical systems in the presence of sensor and actuator attacks. Specifically, we use a bank of observer-based estimators to detect the attacks while introducing a threat-detection level function. Under nominal conditions, the system operates with a nominal-feedback controller with the developed attack monitoring process checking the reliance of the measurements. If there exists an attacker injecting attack signals to a subset of the sensors and/or actuators, then the attack mitigation process is triggered and a two-player, zero-sum differential game is formulated with the defender being the minimizer and the attacker being the maximizer. Next, we solve the underlying joint state estimation and attack mitigation problem and learn the secure control policy using a reinforcement-learning-based algorithm. Finally, two illustrative numerical examples are provided to show the efficacy of the proposed framework. Yuanqiang Zhou, Kyriakos G. Vamvoudakis, Wassim M. Haddad 0001, Zhong-Ping Jiang |
IEEE Trans. Cybern. | 2 |
| 2021 | Guest Editorial: Industrial Artificial Intelligence for Smart ManufacturingabstractThis Special Section presents the latest developments on intelligent modeling, neural networks, deep learning, and adaptive estimation, and their applications in industrial applications. Through a rigorous peer-review process, eleven articles have been accepted, which are summarized below. Tao Yang 0003, Jinliang Ding, Kyriakos G. Vamvoudakis, S. Joe Qin |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Guest Editorial Introduction to the Special Issue on Unmanned Aircraft System Traffic ManagementabstractAdvances of unmanned aircraft system (UAS) technology have spurred a rapid investment of commercial UAS use in broad public domains, such as cargo transport, agriculture support, emergency response, on-demand communication, and infrastructure health monitoring. Urban unmanned aerial transportation that can transport passengers over short distances is also on the way. With the forthcoming dense operations of UAS particularly over urban regions, ensuring airspace safety becomes an urgent issue. Yan Wan 0001, Ella M. Atkins, Dengfeng Sun, Kyriakos G. Vamvoudakis, Konstadinos G. Goulias |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Safe Approximate Dynamic Programming via Kernelized Lipschitz EstimationabstractWe develop a method for obtaining safe initial policies for reinforcement learning via approximate dynamic programming (ADP) techniques for uncertain systems evolving with discrete-time dynamics. We employ the kernelized Lipschitz estimation to learn multiplier matrices that are used in semidefinite programming frameworks for computing admissible initial control policies with provably high probability. Such admissible controllers enable safe initialization and constraint enforcement while providing exponential stability of the equilibrium of the closed-loop system. Ankush Chakrabarty, Devesh K. Jha, Gregery T. Buzzard, Yebin Wang, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Hamiltonian-Driven Hybrid Adaptive Dynamic ProgrammingabstractThis article presents a model-based hybrid adaptive dynamic programming (ADP) framework consisting of continuous feedback-based policy evaluation and policy improvement steps as well as an intermittent policy implementation procedure. This results in an intermittent ADP with a quantifiable performance and guaranteed closed-loop stability of the equilibrium point. To investigate the effect of aperiodic sampling on the communication bandwidth and the control performance of the intermittent ADP algorithms, we use a Hamiltonian-driven unified framework. With such a framework, it is shown that there is a tradeoff between the communication burden and the control performance. We finally show that the developed policies exhibit Zeno-free behaviors. Simulation examples show the efficiency of the proposed framework along with quantifiable comparisons of the policies with different intermittent information. Yongliang Yang 0001, Kyriakos G. Vamvoudakis, Hamidreza Modares, Yixin Yin, Donald C. Wunsch II |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Dynamic Intermittent Feedback Design for $H_{\infty}$ Containment Control on a Directed GraphabstractThis article develops a novel distributed intermittent control framework with the ultimate goal of reducing the communication burden in containment control of multiagent systems communicating via a directed graph. Agents are assumed to be under disturbance and communicate on a directed graph. Both static and dynamic intermittent protocols are proposed. Intermittent H∞containment control design is considered to attenuate the effect of the disturbance and the game algebraic Riccati equation (GARE) is employed to design the coupling and feedback gains for both static and dynamic intermittent feedback. A novel scheme is then used to unify continuous, static, and dynamic intermittent containment protocols. Finally, simulation results verify the efficacy of the proposed approach. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Yixin Yin, Donald C. Wunsch II |
IEEE Trans. Cybern. | 3 |
| 2020 | Safe Intermittent Reinforcement Learning With Static and Dynamic Event GeneratorsabstractIn this article, we present an intermittent framework for safe reinforcement learning (RL) algorithms. First, we develop a barrier function-based system transformation to impose state constraints while converting the original problem to an unconstrained optimization problem. Second, based on optimal derived policies, two types of intermittent feedback RL algorithms are presented, namely, a static and a dynamic one. We finally leverage an actor/critic structure to solve the problem online while guaranteeing optimality, stability, and safety. Simulation results show the efficacy of the proposed approach. Yongliang Yang 0001, Kyriakos G. Vamvoudakis, Hamidreza Modares, Yixin Yin, Donald C. Wunsch II |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Kinodynamic Motion Planning With Continuous-Time Q-Learning: An Online, Model-Free, and Safe Navigation FrameworkabstractThis paper presents an online kinodynamic motion planning algorithmic framework using asymptotically optimal rapidly-exploring random tree (RRT*) and continuous-time Q-learning, which we term as RRT-Q*. We formulate a model-free Q-based advantage function and we utilize integral reinforcement learning to develop tuning laws for the online approximation of the optimal cost and the optimal policy of continuous-time linear systems. Moreover, we provide rigorous Lyapunov-based proofs for the stability of the equilibrium point, which results in asymptotic convergence properties. A terminal state evaluation procedure is introduced to facilitate the online implementation. We propose a static obstacle augmentation and a local replanning framework, which are based on topological connectedness, to locally recompute the robot's path and ensure collision-free navigation. We perform simulations and a qualitative comparison to evaluate the efficacy of the proposed methodology. George P. Kontoudis, Kyriakos G. Vamvoudakis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Optimal and Autonomous Control Using Reinforcement Learning: A SurveyabstractThis paper reviews the current state of the art on reinforcement learning (RL)-based feedback control solutions to optimal regulation and tracking of single and multiagent systems. Existing RL solutions to both optimal and control problems, as well as graphical games, will be reviewed. RL methods learn the solution to optimal control and game problems online and using measured data along the system trajectories. We discuss Q-learning and the integral RL algorithm as core algorithms for discrete-time (DT) and continuous-time (CT) systems, respectively. Moreover, we discuss a new direction of off-policy RL for both CT and DT systems. Finally, we review several applications. Bahare Kiumarsi-Khomartash, Kyriakos G. Vamvoudakis, Hamidreza Modares, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Game-theoretic tracking control for actuator attack attenuation in cyber-physical systemsabstractIn this paper a game-theoretic adaptive learning algorithm based on an action-dependent value function (Q-function) is proposed to solve the optimal tracking control problem with adversarial inputs and completely unknown system and reference dynamics. In order to convert the tracking problem to a regulation problem we augment the system and the reference states and pick appropriately the user-defined matrices. An integral reinforcement learning approach is used to estimate the parameters of the Q-function while also guaranteeing closed-loop stability, trajectory tracking and convergence of the policies to a saddle point. A simulation result of a Quanser helicopter shows the efficacy of the proposed approach. Kyriakos G. Vamvoudakis |
IJCNN | 1 |
| 2016 | Asymptotically Stable Adaptive-Optimal Control Algorithm With Saturating Actuators and Relaxed Persistence of ExcitationabstractThis paper proposes a control algorithm based on adaptive dynamic programming to solve the infinite-horizon optimal control problem for known deterministic nonlinear systems with saturating actuators and nonquadratic cost functionals. The algorithm is based on an actor/critic framework, where a critic neural network (NN) is used to learn the optimal cost, and an actor NN is used to learn the optimal control policy. The adaptive control nature of the algorithm requires a persistence of excitation condition to be a priori validated, but this can be relaxed using previously stored data concurrently with current data in the update of the critic NN. A robustifying control term is added to the controller to eliminate the effect of residual errors, leading to the asymptotically stability of the closed-loop system. Simulation results show the effectiveness of the proposed approach for a controlled Van der Pol oscillator and also for a power system plant. Kyriakos G. Vamvoudakis, Marcio Fantini Miranda, João Pedro Hespanha |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Online adaptive learning of optimal control solutions using integral reinforcement learningabstractIn this paper we introduce an online algorithm that uses integral reinforcement knowledge for learning the continuous-time optimal control solution for nonlinear systems with infinite horizon costs and partial knowledge of the system dynamics. This algorithm is a data based approach to the solution of the Hamilton-Jacobi-Bellman equation and it does not require explicit knowledge on the system's drift dynamics. The adaptive algorithm is based on policy iteration, and it is implemented on an actor/critic structure. Both actor and critic neural networks are adapted simultaneously a persistence of excitation condition is required to guarantee convergence of the critic to the actual optimal value function. Novel tuning algorithms are given for both critic and actor networks, with extra terms in the actor tuning law being required to guarantee closed-loop dynamical stability. The convergence to the optimal controller is proven, and stability of the system is also guaranteed. Simulation examples support the theoretical result. Kyriakos G. Vamvoudakis, Draguna L. Vrabie, Frank L. Lewis |
ADPRL | 1 |
| 2011 | Reinforcement Learning for Partially Observable Dynamic Processes: Adaptive Dynamic Programming Using Measured Output DataabstractApproximate dynamic programming (ADP) is a class of reinforcement learning methods that have shown their importance in a variety of applications, including feedback control of dynamical systems. ADP generally requires full information about the system internal states, which is usually not available in practical situations. In this paper, we show how to implement ADP methods using only measured input/output data from the system. Linear dynamical systems with deterministic behavior are considered herein, which are systems of great interest in the control system community. In control system theory, these types of methods are referred to as output feedback (OPFB). The stochastic equivalent of the systems dealt with in this paper is a class of partially observable Markov decision processes. We develop both policy iteration and value iteration algorithms that converge to an optimal controller that requires only OPFB. It is shown that, similar to Q -learning, the new methods have the important advantage that knowledge of the system dynamics is not needed for the implementation of these learning algorithms or for the OPFB control. Only the order of the system, as well as an upper bound on its "observability index," must be known. The learned OPFB controller is in the form of a polynomial autoregressive moving-average controller that has equivalent performance with the optimal state variable feedback gain. Frank L. Lewis, Kyriakos G. Vamvoudakis |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2009 | Online policy iteration based algorithms to solve the continuous-time infinite horizon optimal control problemabstractIn this paper we discuss two online algorithms based on policy iterations for learning the continuous-time (CT) optimal control solution when nonlinear systems with infinite horizon quadratic cost are considered. For the first time we present an online adaptive algorithm implemented on an actor/critic structure which involves synchronous continuous-time adaptation of both actor and critic neural networks. This is a version of generalized policy iteration for CT systems. The convergence to the optimal controller based on the novel algorithm is proven while stability of the system is guaranteed. The characteristics and requirements of the new online learning algorithm are discussed in relation with the regular online policy iteration algorithm for CT systems which we have previously developed. The latter solves the optimal control problem by performing sequential updates on the actor and critic networks, i.e. while one is learning the other one is held constant. In contrast, the new algorithm relies on simultaneous adaptation of both actor and critic networks. To support the new theoretical result a simulation example is then considered. Kyriakos G. Vamvoudakis, Draguna L. Vrabie, Frank L. Lewis |
ADPRL | 1 |
| 2009 | Online actor critic algorithm to solve the continuous-time infinite horizon optimal control problemabstractIn this paper we discuss an online algorithm based on policy iteration for learning the continuous-time (CT) optimal control solution with infinite horizon cost for nonlinear systems with known dynamics. We present an online adaptive algorithm implemented as an actor/critic structure which involves simultaneous continuous-time adaptation of both actor and critic neural networks. We call this dasiasynchronouspsila policy iteration. A persistence of excitation condition is shown to guarantee convergence of the critic to the actual optimal value function. Novel tuning algorithms are given for both critic and actor networks, with extra terms in the actor tuning law being required to guarantee closed-loop dynamical stability. The convergence to the optimal controller is proven, and stability of the system is also guaranteed. Simulation examples show the effectiveness of the new algorithm. Kyriakos G. Vamvoudakis, Frank L. Lewis |
IJCNN | 1 |