VLDB 2026 Research / reviewers in the wild / expert
Hamidreza Modares
dblp:30/8043
· DBLP profile ↗
37ranked-venue papers
9as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 9 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Confidence Data-Driven Safe Tracking Control DesignabstractThis article presents a high-confidence data-driven safe tracking control design for stochastic linear discrete-time systems. The high-confidence safe reference tracking for an ellipsoidal safe set is first formalized using the concept of probabilistic set-based $\lambda $ -contractivity. A data-driven controller, composed of feedback and feedforward elements, is then designed to enforce the $\lambda $ -contractivity of the safe set. The feedback control gain is learned by 1) providing a data-driven representation of the closed-loop system, which contains a decision variable that affects the control gain and 2) optimizing the decision variable to ensure the $\lambda $ -contractivity. This feedback term can be learned using a data set that is not even rich enough to identify the full system model. A feedforward gain learning algorithm and a data-driven reference governor are provided to satisfy the required conditions on equilibrium terms. It is shown that under certain conditions on the equilibrium terms, the learned tracking controller guarantees the system's safety and stability with high probability. The reference governor dynamically manipulates the desired reference signal based on the data quality to prevent any breach of safety constraints in a probabilistic manner. It is shown that the output of the reference governor eventually converges to the desired goal states if inside the safe set and high-quality data is available. Therefore, the tracking controller guarantees convergence of the system output to its desired goal while ensuring safety with a high probability. The simulation results on a drone hovering and a test system, comparing the results with the existing literature, confirm that the presented high-confidence data-driven safe tracking control outperforms certainty-equivalent safe control methods. Nariman Niknejad, Ramin Esmzad, Hamidreza Modares |
IEEE Trans. Cybern. | 3 |
| 2025 | Online Learning of Noisy Functions via a Data-Regularized Gradient-Descent ApproachabstractOnline first-order algorithms for function identification and regression with noisy data often rely on replacing actual gradients with their constructed noisy estimates. Stochastic gradient descent (SGD) and its mini-batch variant rely on a single sample or a set of randomly selected samples, respectively, to estimate the noisy gradient. SGD algorithms with a constant learning rate for functions that are not strongly convex converge sublinearly. In this article, we show that the strong convexity requirement is satisfied for time-varying regressors if a strong data richness condition, that is, the persistence of excitation (PE) condition, is satisfied on the collected data. To improve the convergence under easy-to-verify data requirements, an online data-regularized concurrent learning-based SGD (CL-based SGD) with a fixed learning rate is then presented for function approximation with noisy data. First, instead of randomly selecting a mini-batch of data, as the mini-batch SGD does, a fixed-size memory of past experiences is repeatedly used in the update law along with the current streaming data. Then, a data selection strategy is used to provide probabilistic convergence guarantees with a highly improved convergence rate (i.e., linear instead of sublinear) to a narrow bound. We finally leverage the Lyapunov theory to provide probabilistic guarantees that assure convergence of the parameters to a probabilistic ultimate bound exponentially fast, provided that a rank condition on the stored data is satisfied. It is shown that the ultimate bound and the exponential convergence to a bounded error region with high probability depend on the condition number of the recorded data matrix. This analysis shows how the quality of the memory data affects the ultimate bound and can reduce the effects of the noise variance on the error bounds. Simulation examples verify the effectiveness of the presented learning approach. Farzaneh Tatari, Ramin Esmzad, Hamidreza Modares |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Discrete-Time Nonlinear System Identification: A Fixed-Time Concurrent Learning ApproachabstractThis article develops a fixed-time identifier for modeling unknown discrete-time nonlinear systems without requiring the standard persistence of excitation (PE) condition. A data-driven update law based on a modified gradient descent (GD) update law is presented to learn the system parameters, which relies on the concurrent learning approach under which the recorded past data and the current data are employed concurrently. Fixed-time convergence guarantees are provided for the modified GD update law under the condition that the recorded data fulfills a rank condition, which is less restrictive than the standard PE condition. To guarantee fixed-time convergence, fixed-time Lyapunov analysis is leveraged. Compared to typical GD-based update laws, two main advantages of the presented approach are: 1) the modified GD update law guarantees fixed-time convergence instead of asymptotic convergence and 2) the convergence guarantee is provided under an easy-to-check rank condition rather than the standard PE condition, which is hard or even impossible to check online. Simulation results are provided to verify the obtained results. Farzaneh Tatari, Nariman Niknejad, Hamidreza Modares |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Fixed-Time Stable Gradient Flows for Optimal Adaptive Control of Continuous-Time Nonlinear SystemsabstractThis paper introduces an inclusive class of fixed‐time stable continuous‐time gradient flows (GFs). This class of GFs is then leveraged to learn optimal control solutions for nonlinear systems in fixed time. It is shown that the presented GF guarantees convergence within a fixed time from any initial condition to the exact minimum of functions that satisfy the Polyak–Łojasiewicz (PL) inequality. The presented fixed‐time GF is then utilized to design fixed‐time optimal adaptive control algorithms. To this end, a fixed‐time reinforcement learning (RL) algorithm is developed on the basis of a single network adaptive critic (SNAC) to learn the solution to an infinite‐horizon optimal control problem in a fixed‐time convergent, online, adaptive, and forward‐in‐time manner. It is shown that the PL inequality in the presented RL algorithm amounts to a mild inequality condition on a few collected samples. This condition is much weaker than the standard persistence of excitation (PE) and finite duration PE that relies on a rank condition of a dataset. This is crucial for learning‐enabled control systems as control systems can commit to learning an optimal controller from the beginning, in sharp contrast to existing results that rely on the PE and rank condition, and can only commit to learning after rich data samples are collected. Simulation results are provided to validate the performance and efficacy of the presented fixed‐time RL algorithm. Mahdi Niroomand, Reihaneh Kardehi Moghaddam, Hamidreza Modares, Mohammad-Bagher Naghibi-Sistani |
Int. J. Intell. Syst. | 3 |
| 2024 | Cooperative Finitely Excited Learning for Dynamical GamesabstractIn this article, we propose a way to enhance the learning framework for zero-sum games with dynamics evolving in continuous time. In contrast to the conventional centralized actor-critic learning, a novel cooperative finitely excited learning approach is developed to combine the online recorded data with instantaneous data for efficiency. By using an experience replay technique for each agent and distributed interaction amongst agents, we are able to replace the classical persistent excitation condition with an easy-to-check cooperative excitation condition. This approach also guarantees the consensus of the distributed actor-critic learning on the solution to the Hamilton-Jacobi-Isaacs (HJI) equation. It is shown that both the closed-loop stability of the equilibrium point and convergence to the Nash equilibrium can be guaranteed. Simulation results demonstrate the efficacy of this approach compared to previous methods. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Frank L. Lewis |
IEEE Trans. Cybern. | 2 |
| 2024 | Safe Reinforcement Learning via a Model-Free Safety CertifierabstractThis article presents a data-driven safe reinforcement learning (RL) algorithm for discrete-time nonlinear systems. A data-driven safety certifier is designed to intervene with the actions of the RL agent to ensure both safety and stability of its actions. This is in sharp contrast to existing model-based safety certifiers that can result in convergence to an undesired equilibrium point or conservative interventions that jeopardize the performance of the RL agent. To this end, the proposed method directly learns a robust safety certifier while completely bypassing the identification of the system model. The nonlinear system is modeled using linear parameter varying (LPV) systems with polytopic disturbances. To prevent the requirement for learning an explicit model of the LPV system, data-based λ -contractivity conditions are first provided for the closed-loop system to enforce robust invariance of a prespecified polyhedral safe set and the system's asymptotic stability. These conditions are then leveraged to directly learn a robust data-based gain-scheduling controller by solving a convex program. A significant advantage of the proposed direct safe learning over model-based certifiers is that it completely resolves conflicts between safety and stability requirements while assuring convergence to the desired equilibrium point. Data-based safety certification conditions are then provided using Minkowski functions. They are then used to seemingly integrate the learned backup safe gain-scheduling controller with the RL controller. Finally, we provide a simulation example to verify the effectiveness of the proposed approach. Amir Modares, Nasser Sadati, Babak Esmaeili 0004, Farnaz Adib Yaghmaie, Hamidreza Modares |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Finite-time Koopman Identifier: A Unified Batch-online Learning Framework for Joint Learning of Koopman Structure and ParametersabstractIn this paper, a unified batch-online learning approach is introduced to learn a linear representation of nonlinear system dynamics using the Koopman operator. The presented system modeling approach leverages a novel incremental Koopman-based update law that regains a mini-collection of samples stored in a memory to minimize not only the instantaneous Koopman operator’s identification errors but also the identification errors for the collection of retrieved samples. Discontinuous modifications of gradient flows are presented for the online update law to assure finite-time convergence under easy-to-verify conditions defined on the batch of data. Therefore, this unified online-batch framework allows joint sample- and time-domain analysis to converge the Koopman operator’s parameters. More specifically, it is shown that if the collected mini-batch of samples guarantees a rank condition, then finite-time guarantee in the time domain can be certified, and the settling time depends on the quality of collected samples being reused in the update law. Moreover, the efficiency of the proposed Koopman-based update law is further analyzed by showing that the identification regret in continuous time grows sub-linearly with time. Furthermore, to avoid learning corrupted dynamics due to the selection of an inappropriate set of Koopman observables, a higher-layer meta-learner employs a discrete Bayesian optimization algorithm to obtain the best library of observable functions for the operator. Since finite-time convergence of the Koopman model for each set of observables is guaranteed under a rank condition on stored data, the fitness of each set of observables can be obtained based on the identification error on the stored samples in the proposed framework and even without implementing any controller based on the learned system. Finally, to confirm the effectiveness of the proposed scheme, two simulation examples are presented. Majid Mazouchi, Subramanya Nageshrao, Hamidreza Modares |
J. Mach. Learn. Res. | 3 |
| 2023 | Fixed-Time System Identification Using Concurrent LearningabstractThis article presents a fixed-time (FxT) system identifier for continuous-time nonlinear systems. A novel adaptive update law with discontinuous gradient flows of the identification errors is presented, which leverages concurrent learning (CL) to guarantee the learning of uncertain nonlinear dynamics in a fixed time, as opposed to asymptotic or exponential time. More specifically, the CL approach retrieves a batch of samples stored in a memory, and the update law simultaneously minimizes the identification error for the current stream of samples and past memory samples. Rigorous analyses are provided based on FxT Lyapunov stability to certify FxT convergence to the stable equilibria of the gradient descent flow of the system identification error under easy-to-verify rank conditions. The performance of the proposed method in comparison with the existing methods is illustrated in the simulation results. Farzaneh Tatari, Majid Mazouchi, Hamidreza Modares |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Model-Free λ-Policy Iteration for Discrete-Time Linear Quadratic RegulationabstractThis article presents a model-free λ -policy iteration ( λ -PI) for the discrete-time linear quadratic regulation (LQR) problem. To solve the algebraic Riccati equation arising from solving the LQR in an iterative manner, we define two novel matrix operators, named the weighted Bellman operator and the composite Bellman operator. Then, the λ -PI algorithm is first designed as a recursion with the weighted Bellman operator, and its equivalent formulation as a fixed-point iteration with the composite Bellman operator is shown. The contraction and monotonic properties of the composite Bellman operator guarantee the convergence of the λ -PI algorithm. In contrast to the PI algorithm, the λ -PI does not require an admissible initial policy, and the convergence rate outperforms the value iteration (VI) algorithm. Model-free extension of the λ -PI algorithm is developed using the off-policy reinforcement learning technique. It is also shown that the off-policy variants of the λ -PI algorithm are robust against the probing noise. Finally, simulation examples are conducted to validate the efficacy of the λ -PI algorithm. Yongliang Yang 0001, Bahare Kiumarsi-Khomartash, Hamidreza Modares, Cheng-Zhong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Secure Event-Triggered Distributed Kalman Filters for State Estimation Over Wireless Sensor NetworksabstractIn this article, we analyze the adverse effects of cyber–physical attacks as well as mitigate their impacts on the event-triggered distributed Kalman filter (DKF). We first show that although event-triggered mechanisms are highly desirable, the attacker can leverage the event-triggered mechanism to cause nontriggering misbehavior, which significantly harms the network connectivity and its collective observability. We also show that an attacker can mislead the event-triggered mechanism to achieve continuous-triggering misbehavior, which not only drains the communication resources but also harms the network’s performance. An information-theoretic approach is presented next to detect attacks on both sensors and communication channels. In contrast to the existing results, the restrictive Gaussian assumption on the attack signal’s probability distribution is not required. To mitigate attacks, a meta-Bayesian approach is presented that incorporates the outcome of the attack detection mechanism to perform second-order inference. The proposed second-order inference forms confidence and trust values about the truthfulness or legitimacy of sensors’ own estimates and those of their neighbors, respectively. Each sensor communicates its confidence to its neighbors. Sensors then incorporate the confidence they receive from their neighbors and the trust they formed about their neighbors into their posterior update laws to successfully discard corrupted information. Finally, the simulation result validates the effectiveness of the presented resilient event-triggered DKF. Aquib Mustafa, Majid Mazouchi, Hamidreza Modares |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Event-Driven Off-Policy Reinforcement Learning for Control of Interconnected SystemsabstractIn this article, we introduce a novel approximate optimal decentralized control scheme for uncertain input-affine nonlinear-interconnected systems. In the proposed scheme, we design a controller and an event-triggering mechanism (ETM) at each subsystem to optimize a local performance index and reduce redundant control updates, respectively. To this end, we formulate a noncooperative dynamic game at every subsystem in which we collectively model the interconnection inputs and the event-triggering error as adversarial players that deteriorate the subsystem performance and model the control policy as the performance optimizer, competing against these adversarial players. To obtain a solution to this game, one has to solve the associated Hamilton-Jacobi-Isaac (HJI) equation, which does not have a closed-form solution even when the subsystem dynamics are accurately known. In this context, we introduce an event-driven off-policy integral reinforcement learning (OIRL) approach to learn an approximate solution to this HJI equation using artificial neural networks (NNs). We then use this NN approximated solution to design the control policy and event-triggering threshold at each subsystem. In the learning framework, we guarantee the Zeno-free behavior of the ETMs at each subsystem using the exploration policies. Finally, we derive sufficient conditions to guarantee uniform ultimate bounded regulation of the controlled system states and demonstrate the efficacy of the proposed framework with numerical examples. Vignesh Narayanan, Hamidreza Modares, Sarangapani Jagannathan, Frank L. Lewis |
IEEE Trans. Cybern. | 2 |
| 2022 | Hamiltonian-Driven Adaptive Dynamic Programming With Approximation ErrorsabstractIn this article, we consider an iterative adaptive dynamic programming (ADP) algorithm within the Hamiltonian-driven framework to solve the Hamilton-Jacobi-Bellman (HJB) equation for the infinite-horizon optimal control problem in continuous time for nonlinear systems. First, a novel function, "min-Hamiltonian," is defined to capture the fundamental properties of the classical Hamiltonian. It is shown that both the HJB equation and the policy iteration (PI) algorithm can be formulated in terms of the min-Hamiltonian within the Hamiltonian-driven framework. Moreover, we develop an iterative ADP algorithm that takes into consideration the approximation errors during the policy evaluation step. We then derive a sufficient condition on the iterative value gradient to guarantee closed-loop stability of the equilibrium point as well as convergence to the optimal value. A model-free extension based on an off-policy reinforcement learning (RL) technique is also provided. Finally, numerical results illustrate the efficacy of the proposed framework. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Wei He 0001, Cheng-Zhong Xu 0001, Donald C. Wunsch II |
IEEE Trans. Cybern. | 2 |
| 2022 | Robust Actor-Critic Learning for Continuous-Time Nonlinear Systems With Unmodeled DynamicsabstractThis article considers the robust optimal control problem for a class of nonlinear systems in the presence of unmodeled dynamics. An adaptive optimal controller is designed using the online actor–critic learning and is robustified against unmodeled dynamics. To deal with unmodeled dynamics, an auxiliary signal with the system state as its input signal is designed to capture the input-to-state stability. In addition to the critic network for value function approximation, a novel robustifying term is developed and introduced into the actor network to ensure robustness during the learning process. It is shown that both the actor and the critic weights learning converge to their optimal values while guaranteeing the boundedness of all the signals in the closed loop. Simulation examples are conducted to verify the efficacy of the presented scheme. Yongliang Yang 0001, Weinan Gao, Hamidreza Modares, Cheng-Zhong Xu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2022 | Distributed Consensus Control of Vehicular Platooning Under Delay, Packet Dropout and Noise: Relative State and Relative Input-Output Control StrategiesabstractIn this article, we study the distributed control problem of vehicular platooning by taking into account network imperfections, namely the input delay and the packet dropout. The delay occurs within the inter network of each vehicle from its controller to its actuators and the packet dropout occurs within the intra network of vehicles when information is exchanged between vehicles. Relative state-feedback and relative input-output feedback control strategies are adopted for the vehicular platooning under network imperfection. To this end, and to avoid the requirement of relative state information, the relative state measurements are constructed for platoon systems based on a distributed observer via only relative input-output measurements and in the presence of delay, packet dropout and noise. Conditions for synchronization are provided in terms of the unique positive-definite solutions to some algebraic Riccati equations. Simulations validate the theoretical results. Arezou Elahi, Alireza Alfi, Hamidreza Modares |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Data-Driven Dynamic Multiobjective Optimal Control: An Aspiration-Satisfying Reinforcement Learning Approach
Majid Mazouchi, Yongliang Yang 0001, Hamidreza Modares |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | H∞ Consensus of Homogeneous Vehicular Platooning Systems With Packet Dropout and Communication DelayabstractConsensus of vehicular platooning systems under external disturbances and network imperfections, i.e., communication delay and random packet dropout, is considered in this article. Third-order dynamics are considered for vehicles to take into account the rate of change of acceleration, i.e., the jerk, for which its control can provide comfort to the passengers. Designing consensus protocols for platoon of vehicles with third-order dynamics that can deal with both disturbance and network imperfection is not straightforward and requires new developments. Using the Lyapunov–Krasovskii functional, sufficient conditions are provided to assure that: 1) the consensus error dynamics are asymptotically mean-square consensus stable and 2) a given disturbance attenuation level is achieved in the presence of both network imperfections and disturbances. Simulation results validate the efficiency of the presented approach in dealing with disturbances and network imperfections. Arezou Elahi, Alireza Alfi, Hamidreza Modares |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Heterogeneous formation control of multiple rotorcrafts with unknown dynamics by reinforcement learning
Hao Liu 0004, Fachun Peng, Hamidreza Modares, Bahare Kiumarsi-Khomartash |
Inf. Sci. | 3 |
| 2021 | Hamiltonian-Driven Hybrid Adaptive Dynamic ProgrammingabstractThis article presents a model-based hybrid adaptive dynamic programming (ADP) framework consisting of continuous feedback-based policy evaluation and policy improvement steps as well as an intermittent policy implementation procedure. This results in an intermittent ADP with a quantifiable performance and guaranteed closed-loop stability of the equilibrium point. To investigate the effect of aperiodic sampling on the communication bandwidth and the control performance of the intermittent ADP algorithms, we use a Hamiltonian-driven unified framework. With such a framework, it is shown that there is a tradeoff between the communication burden and the control performance. We finally show that the developed policies exhibit Zeno-free behaviors. Simulation examples show the efficiency of the proposed framework along with quantifiable comparisons of the policies with different intermittent information. Yongliang Yang 0001, Kyriakos G. Vamvoudakis, Hamidreza Modares, Yixin Yin, Donald C. Wunsch II |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Resilient and Robust Synchronization of Multiagent Systems Under Attacks on Sensors and ActuatorsabstractResilient and robust distributed control protocols for multiagent systems under attacks on sensors and actuators are designed. A distributed H∞control protocol is designed to attenuate the disturbance or attack effects. However, the H∞controller is too conservative in the presence of attacks. Therefore, it is augmented with a distributed adaptive compensator to mitigate the adverse effects of attacks. The proposed controller can make the synchronization error arbitrarily small in the presence of faulty attacks, and satisfy global L2-gain performance in the presence of malicious attacks or disturbances. A significant advantage of the proposed method is that it requires no restriction on the number of agents or agents' neighbors under attacks on sensors and/or actuators, and it recovers even compromised agents under attacks on actuators. Simulation examples verify the effectiveness of the proposed method. Hamidreza Modares, Bahare Kiumarsi-Khomartash, Frank L. Lewis, Frank T. Ferrese, Ali Davoudi |
IEEE Trans. Cybern. | 1 |
| 2020 | Dynamic Intermittent Feedback Design for $H_{\infty}$ Containment Control on a Directed GraphabstractThis article develops a novel distributed intermittent control framework with the ultimate goal of reducing the communication burden in containment control of multiagent systems communicating via a directed graph. Agents are assumed to be under disturbance and communicate on a directed graph. Both static and dynamic intermittent protocols are proposed. Intermittent H∞containment control design is considered to attenuate the effect of the disturbance and the game algebraic Riccati equation (GARE) is employed to design the coupling and feedback gains for both static and dynamic intermittent feedback. A novel scheme is then used to unify continuous, static, and dynamic intermittent containment protocols. Finally, simulation results verify the efficacy of the proposed approach. Yongliang Yang 0001, Hamidreza Modares, Kyriakos G. Vamvoudakis, Yixin Yin, Donald C. Wunsch II |
IEEE Trans. Cybern. | 2 |
| 2020 | Safe Intermittent Reinforcement Learning With Static and Dynamic Event GeneratorsabstractIn this article, we present an intermittent framework for safe reinforcement learning (RL) algorithms. First, we develop a barrier function-based system transformation to impose state constraints while converting the original problem to an unconstrained optimization problem. Second, based on optimal derived policies, two types of intermittent feedback RL algorithms are presented, namely, a static and a dynamic one. We finally leverage an actor/critic structure to solve the problem online while guaranteeing optimality, stability, and safety. Simulation results show the efficacy of the proposed approach. Yongliang Yang 0001, Kyriakos G. Vamvoudakis, Hamidreza Modares, Yixin Yin, Donald C. Wunsch II |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Optimal control using adaptive resonance theory and Q-learning
Bahare Kiumarsi-Khomartash, Bakur AlQaudi, Hamidreza Modares, Frank L. Lewis, Daniel S. Levine 0001 |
Neurocomputing | 3 |
| 2019 | Resilient Autonomous Control of Distributed Multiagent Systems in Contested EnvironmentsabstractAn autonomous and resilient controller is proposed for leader-follower multiagent systems under uncertainties and cyber-physical attacks. The leader is assumed nonautonomous with a nonzero control input, which allows changing the team behavior or mission in response to the environmental changes. A resilient learning-based control protocol is presented to find optimal solutions to the synchronization problem in the presence of attacks and system dynamic uncertainties. An observer-based distributed H∞controller is first designed to prevent propagating the effects of attacks on sensors and actuators throughout the network, as well as to attenuate the effect of these attacks on the compromised agent itself. Nonhomogeneous game algebraic Riccati equations are derived to solve the H∞optimal synchronization problem and off-policy reinforcement learning (RL) is utilized to learn their solution without requiring any knowledge of the agent's dynamics. A trust-confidence-based distributed control protocol is then proposed to mitigate attacks that hijack the entire node and attacks on communication links. A confidence value is defined for each agent based solely on its local evidence. The proposed resilient RL algorithm employs the confidence value of each agent to indicate the trustworthiness of its own information and broadcast it to its neighbors to put weights on the data they receive from it during and after learning. If the confidence value of an agent is low, it employs a trust mechanism to identify compromised agents and remove the data it receives from them from the learning process. The simulation results are provided to show the effectiveness of the proposed approach. Rohollah Moghadam, Hamidreza Modares |
IEEE Trans. Cybern. | 2 |
| 2018 | Optimal and Autonomous Control Using Reinforcement Learning: A SurveyabstractThis paper reviews the current state of the art on reinforcement learning (RL)-based feedback control solutions to optimal regulation and tracking of single and multiagent systems. Existing RL solutions to both optimal and control problems, as well as graphical games, will be reviewed. RL methods learn the solution to optimal control and game problems online and using measured data along the system trajectories. We discuss Q-learning and the integral RL algorithm as core algorithms for discrete-time (DT) and continuous-time (CT) systems, respectively. Moreover, we discuss a new direction of off-policy RL for both CT and DT systems. Finally, we review several applications. Bahare Kiumarsi-Khomartash, Kyriakos G. Vamvoudakis, Hamidreza Modares, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Leader-Follower Output Synchronization of Linear Heterogeneous Systems With Active Leader Using Reinforcement LearningabstractThis paper develops optimal control protocols for the distributed output synchronization problem of leader-follower multiagent systems with an active leader. Agents are assumed to be heterogeneous with different dynamics and dimensions. The desired trajectory is assumed to be preplanned and is generated by the leader. Other follower agents autonomously synchronize to the leader by interacting with each other using a communication network. The leader is assumed to be active in the sense that it has a nonzero control input so that it can act independently and update its control to keep the followers away from possible danger. A distributed observer is first designed to estimate the leader's state and generate the reference signal for each follower. Then, the output synchronization of leader-follower systems with an active leader is formulated as a distributed optimal tracking problem, and inhomogeneous algebraic Riccati equations (AREs) are derived to solve it. The resulting distributed optimal control protocols not only minimize the steady-state error but also optimize the transient response of the agents. An off-policy reinforcement learning algorithm is developed to solve the inhomogeneous AREs online in real time and without requiring any knowledge of the agents' dynamics. Finally, two simulation examples are conducted to illustrate the effectiveness of the proposed algorithm. Yongliang Yang 0001, Hamidreza Modares, Donald C. Wunsch II, Yixin Yin |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Off-Policy Reinforcement Learning for Synchronization in Multiagent Graphical GamesabstractThis paper develops an off-policy reinforcement learning (RL) algorithm to solve optimal synchronization of multiagent systems. This is accomplished by using the framework of graphical games. In contrast to traditional control protocols, which require complete knowledge of agent dynamics, the proposed off-policy RL algorithm is a model-free approach, in that it solves the optimal synchronization problem without knowing any knowledge of the agent dynamics. A prescribed control policy, called behavior policy, is applied to each agent to generate and collect data for learning. An off-policy Bellman equation is derived for each agent to learn the value function for the policy under evaluation, called target policy, and find an improved policy, simultaneously. Actor and critic neural networks along with least-square approach are employed to approximate target control policies and value functions using the data generated by applying prescribed behavior policies. Finally, an off-policy RL algorithm is presented that is implemented in real time and gives the approximate optimal control policy for each agent using only measured data. It is shown that the optimal distributed policies found by the proposed algorithm satisfy the global Nash equilibrium and synchronize all agents to the leader. Simulation results illustrate the effectiveness of the proposed method. Jinna Li, Hamidreza Modares, Tianyou Chai, Frank L. Lewis, Lihua Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Optimal output synchronization of nonlinear multi-agent systems using approximate dynamic programmingabstractOptimal output synchronization of multi-agent leader-follower systems is considered. The agents are assumed heterogeneous so that the dynamics may be non-identical. An optimal control protocol is designed for each agent based on the leader state and the agent local state. A distributed observer is designed to provide the leader state for each agent. A model-free approximate dynamic programming algorithm is then developed to solve the optimal output synchronization problem online in real time. No knowledge of the agents' dynamics is required. The proposed approach does not require explicitly solving of the output regulator equations, though it implicitly solves them by imposing optimality. A simulation example verifies the suitability of the proposed approach. Hamidreza Modares, Frank L. Lewis, Ali Davoudi |
IJCNN | 1 |
| 2016 | Optimal Output-Feedback Control of Unknown Continuous-Time Linear Systems Using Off-policy Reinforcement LearningabstractA model-free off-policy reinforcement learning algorithm is developed to learn the optimal output-feedback (OPFB) solution for linear continuous-time systems. The proposed algorithm has the important feature of being applicable to the design of optimal OPFB controllers for both regulation and tracking problems. To provide a unified framework for both optimal regulation and tracking, a discounted performance function is employed and a discounted algebraic Riccati equation (ARE) is derived which gives the solution to the problem. Conditions on the existence of a solution to the discounted ARE are provided and an upper bound for the discount factor is found to assure the stability of the optimal control solution. To develop an optimal OPFB controller, it is first shown that the system state can be constructed using some limited observations on the system output over a period of the history of the system. A Bellman equation is then developed to evaluate a control policy and find an improved policy simultaneously using only some limited observations on the system output. Then, using this Bellman equation, a model-free Off-policy RL-based OPFB controller is developed without requiring the knowledge of the system state or the system dynamics. It is shown that the proposed OPFB method is more powerful than the static OPFB as it is equivalent to a state-feedback control policy. The proposed method is successfully used to solve a regulation and a tracking problem. Hamidreza Modares, Frank L. Lewis, Zhong-Ping Jiang |
IEEE Trans. Cybern. | 1 |
| 2016 | Optimized Assistive Human-Robot Interaction Using Reinforcement LearningabstractAn intelligent human-robot interaction (HRI) system with adjustable robot behavior is presented. The proposed HRI system assists the human operator to perform a given task with minimum workload demands and optimizes the overall human-robot system performance. Motivated by human factor studies, the presented control structure consists of two control loops. First, a robot-specific neuro-adaptive controller is designed in the inner loop to make the unknown nonlinear robot behave like a prescribed robot impedance model as perceived by a human operator. In contrast to existing neural network and adaptive impedance-based control methods, no information of the task performance or the prescribed robot impedance model parameters is required in the inner loop. Then, a task-specific outer-loop controller is designed to find the optimal parameters of the prescribed robot impedance model to adjust the robot's dynamics to the operator skills and minimize the tracking error. The outer loop includes the human operator, the robot, and the task performance details. The problem of finding the optimal parameters of the prescribed robot impedance model is transformed into a linear quadratic regulator (LQR) problem which minimizes the human effort and optimizes the closed-loop behavior of the HRI system for a given task. To obviate the requirement of the knowledge of the human model, integral reinforcement learning is used to solve the given LQR problem. Simulation results on an x - y table and a robot arm, and experimental implementation results on a PR2 robot confirm the suitability of the proposed method. Hamidreza Modares, Isura Ranatunga, Frank L. Lewis, Dan O. Popa |
IEEE Trans. Cybern. | 1 |
| 2016 | Data-Based Multiobjective Plant-Wide Performance Optimization of Industrial Processes Under Dynamic EnvironmentsabstractThis paper provides a method for automatically selecting optimal operational indices for unit processes in an industrial plant using measured data and without knowing dynamical models of the unit process. A dynamic multiobjective optimization problem is defined to find operational indices that lead to plant-wide production indices close to their target values. A case-based reasoning (CBR) technique is also employed, which uses the stored experience of a human expert to determine appropriate operational indices for given target production indices. The solutions of the optimization problem and CBR technique are combined to form baseline operational indices. The dynamic models of the production indices, however, are time varying and affected by disturbances and online corrections of these baseline operational indices are required. To this end, reinforcement learning (RL) is used to provide a data-driven optimization technique to compensate for disturbances and model approximation errors and variations. The data-driven RL approach is used in two different time scales. The samples of the predicted production indices are used at a fast sampling rate, i.e., at each sample time, and the samples of actual production indices are used at a slower sampling rate, i.e., after each operational run, to correct the baseline operational indices. The effectiveness of this automated decision procedure has been demonstrated by successful implementation of the proposed approach on a large mineral processing plant in Gansu Province, China. Jinliang Ding, Hamidreza Modares, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Ind. Informatics | 2 |
| 2015 | Continuous-Time Q-Learning for Infinite-Horizon Discounted Cost Linear Quadratic Regulator ProblemsabstractThis paper presents a method of Q-learning to solve the discounted linear quadratic regulator (LQR) problem for continuous-time (CT) continuous-state systems. Most available methods in the existing literature for CT systems to solve the LQR problem generally need partial or complete knowledge of the system dynamics. Q-learning is effective for unknown dynamical systems, but has generally been well understood only for discrete-time systems. The contribution of this paper is to present a Q-learning methodology for CT systems which solves the LQR problem without having any knowledge of the system dynamics. A natural and rigorous justified parameterization of the Q-function is given in terms of the state, the control input, and its derivatives. This parameterization allows the implementation of an online Q-learning algorithm for CT systems. The simulation results supporting the theoretical development are also presented. Muthukumar Palanisamy, Hamidreza Modares, Frank L. Lewis, Muhammad Aurangzeb |
IEEE Trans. Cybern. | 2 |
| 2015 | H∞ Tracking Control of Completely Unknown Continuous-Time Systems via Off-Policy Reinforcement LearningabstractThis paper deals with the design of an H ∞ tracking controller for nonlinear continuous-time systems with completely unknown dynamics. A general bounded L2 -gain tracking problem with a discounted performance function is introduced for the H ∞ tracking. A tracking Hamilton-Jacobi-Isaac (HJI) equation is then developed that gives a Nash equilibrium solution to the associated min-max optimization problem. A rigorous analysis of bounded L2 -gain and stability of the control solution obtained by solving the tracking HJI equation is provided. An upper-bound is found for the discount factor to assure local asymptotic stability of the tracking error dynamics. An off-policy reinforcement learning algorithm is used to learn the solution to the tracking HJI equation online without requiring any knowledge of the system dynamics. Convergence of the proposed algorithm to the solution to the tracking HJI equation is shown. Simulation examples are provided to verify the effectiveness of the proposed method. Hamidreza Modares, Frank L. Lewis, Zhong-Ping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Adaptive Optimal Control of Unknown Constrained-Input Systems Using Policy Iteration and Neural NetworksabstractThis paper presents an online policy iteration (PI) algorithm to learn the continuous-time optimal control solution for unknown constrained-input systems. The proposed PI algorithm is implemented on an actor-critic structure where two neural networks (NNs) are tuned online and simultaneously to generate the optimal bounded control policy. The requirement of complete knowledge of the system dynamics is obviated by employing a novel NN identifier in conjunction with the actor and critic NNs. It is shown how the identifier weights estimation error affects the convergence of the critic NN. A novel learning rule is developed to guarantee that the identifier weights converge to small neighborhoods of their ideal values exponentially fast. To provide an easy-to-check persistence of excitation condition, the experience replay technique is used. That is, recorded past experiences are used simultaneously with current data for the adaptation of the identifier weights. Stability of the whole system consisting of the actor, critic, system state, and system identifier is guaranteed while all three networks undergo adaptation. Convergence to a near-optimal control law is also shown. The effectiveness of the proposed method is illustrated with a simulation example. Hamidreza Modares, Frank L. Lewis, Mohammad-Bagher Naghibi-Sistani |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | A general insight into the effect of neuron structure on classification
Hadi Sadoghi Yazdi, Alireza Rowhanimanesh, Hamidreza Modares |
Knowl. Inf. Syst. | 3 |
| 2011 | Solving nonlinear optimal control problems using a hybrid IPSO-SQP algorithm
Hamidreza Modares, Mohammad-Bagher Naghibi-Sistani |
Eng. Appl. Artif. Intell. | 1 |
| 2010 | Parameter estimation of bilinear systems based on an adaptive particle swarm optimization
Hamidreza Modares, Alireza Alfi, Mohammad-Bagher Naghibi-Sistani |
Eng. Appl. Artif. Intell. | 1 |
| 2010 | Parameter identification of chaotic dynamic systems through an improved particle swarm optimization
Hamidreza Modares, Alireza Alfi, Mohammad Mehdi Fateh |
Expert Syst. Appl. | 1 |