EDBT 2026 Demo / reviewers in the wild / expert
Bosen Lian
dblp:245/5752
· DBLP profile ↗
25ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-3275-9551ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 10 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-Driven Inverse Reinforcement Learning for Markov Multiplayer Tidal Turbine SystemsabstractThis paper designs data-driven inverse reinforcement learning (IRL) for Markov jump tidal turbine control systems, where the stochastic tidal flow is modeled as a finite-state Markov process, and the pitch and generator speed controllers are treated as players in a nonzero-sum differential game. By exploiting demonstrated optimal state and control input trajectories, the cost function weights can be reconstructed without requiring knowledge of the system dynamics. First, a model-based IRL algorithm with three components is developed: (i) policy evaluation based on a modified Lyapunov function, (ii) policy improvement using the optimal control, and (iii) weight updates via inverse optimal control. An equivalent data-driven IRL algorithm is designed, which incorporates transition rate identification into an integral off-policy method that uncovers cost and control parameters without prior knowledge of system dynamics. Rigorous convergence analysis and simulation studies validate the proposed approach. Bosen Lian |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Inverse Reinforcement Learning for Disturbed Networked Nonlinear Systems With Data DropoutsabstractThis article develops inverse reinforcement learning (IRL) control algorithms for nonlinear networked control systems (NCSs) to mimic trajectories of a target system governed by an unknown optimal cost function, despite the presence of random data dropouts and external disturbances. Data dropouts occur during: 1) reception of target trajectory data by the controller; 2) reception of state feedback data by the controller; and 3) reception of control input data by the actuator. By organically integrating $H_{\infty } $ control to account for disturbances and dropout-induced uncertainty, a model-based IRL algorithm is first developed. Building on this, a neural-network-based data-driven IRL algorithm is developed to infer the cost function and optimal control policy using available data while reducing dependence on system models. The proposed methods enable effective trajectory imitation under partial model knowledge, data dropouts, and disturbances, as demonstrated through simulation studies. Wenqian Xue, Jialu Fan, Frank L. Lewis, Bosen Lian |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | Distributed FilterNet Reinforcement Learning for Achieving Output Consensus in Heterogeneous Multiplayer Multiagent SystemsabstractWe study the leader-follower consensus problem in multiagent systems with heterogeneous agent dynamics and multiple internal players per agent, each with distinct and interaffected objectives. Formulated as a multiplayer differential game per agent, the goal is to achieve output consensus among all agents while ensuring Nash equilibrium controls across each agent's internal players. To address this challenge, we introduce a distributed control framework that integrates both feedforward (regulator-based) and feedback (game-theoretic Riccati-based) components. We further design a FilterNet reinforcement learning (RL) architecture that solves the control solutions while eliminating the need for large-scale distributed data storage. Organized into four layers, FilterNet handles admissible policy identification, online initialization, asynchronous updates for Nash policy convergence, and real-time regulator solutions. This design reduces data requirements, ensures initial excitation, and accelerates convergence. Theoretical guarantees establish conditions for solvability and convergence. Numerical simulations and comparisons with existing methods confirm the effectiveness and superiority of the proposed approach. Bosen Lian, Changyun Wen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | Initially Excited Asynchronous Reinforcement Learning Control With Monotonicity and StabilityabstractThis article addresses the existing reinforcement learning (RL) control issues of continuous-time Markov jump systems (MJSs), including their synchronous iteration structure using nonlatest updates, nonmonotonic convergence, and requirement of initial admissible control and all-time persistent excitation (PE) condition. We propose advanced model-based and model-free RL algorithms that: 1) have asynchronous decoupled Lyapunov iteration equations to approximate the optimal control solutions using the latest updates; 2) determine the initial admissible control policy (IACP) and initial value function matrix without relying on engineering experiences; and 3) relax PE condition with a milder initial excitation (IE) condition. Rigorous theoretical analyses are provided to establish the monotonic convergence to the optimal control solution and the closed-loop stability at each iteration. Finally, the simulation and comparison results verify the proposed algorithms. Wenqian Xue, Frank L. Lewis, Bosen Lian |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Robust Node Localization With Fault Detection in Anisotropic NetworksabstractIn Internet of Things (IoT) applications, node localization in wireless sensor networks (WSNs) is critical for ensuring service quality. However, the presence of faulty nodes can significantly degrade localization accuracy. This paper proposes the Hop Distance Credibility Z-Score DV-Hop (HDCZ-DV-Hop) algorithm, which enhances traditional hopsize metrics by integrating hop distance credibility with Z-Score analysis for more accurate fault node detection. An adaptive Z-Score threshold mechanism is introduced to accommodate various network topologies, particularly anisotropic networks with irregular shapes. Simulation results, conducted with a network border length L of 100m and optimal adaptive Z-Score thresholds, show fault node detection rates of 97%, 96%, and 77%, and RMSE values of 4.43m, 7.21m, and 9.04m for Square, C-shaped, and S-shaped topologies, respectively. Compared to other fault node identification methods, such as the Jenks natural breaks and Pauta Criterion, the HDCZ-DV-Hop algorithm significantly reduces RMSE across different network shapes. Ruizhuo Song, Shi Xing, Bosen Lian |
IEEE Internet Things J. | 3 |
| 2025 | Data-Driven Weighted H∞ Control of Persistent Dwell Time Switched Systems With Optimal Disturbance Attenuation GuaranteedabstractPersistent dwell-time switched systems (PDTSSs) are being increasingly employed for modeling the dynamics of systems with non-uniform time-dependent switching rules. The existing attempts to design control methods for PDTSSs necessitate a prior knowledge of system dynamics. In this article, we propose a data-driven reinforcement learning (RL) method to solve the weighted${\mathcal {H}}_{\infty }$control problem for PDTSSs with completely unknown system dynamics. The proposed method aims to calculate the weighted${\mathcal {H}}_{\infty }$control gain for PDTSSs with optimal disturbance attenuation guaranteed by solving a set of linear matrix inequalities related to system data, as opposed to solving it with the exact model information. To ensure the stability of the overall switched system, we establish the criterion for exponential mean-square stability of the closed-loop PDTSSs, as well as analyze the convergence of the proposed data-driven RL algorithm. Finally, the efficacy of the designed algorithms is illustrated via an electric circuit model. Note to Practitioners—First, at a practical level, we consider a more general PDT switching rule that can effectively model the phenomenon of simultaneous fast and slow switching in real systems, which finds wide application in various engineering domains, including capacitor monolithic filters, robot manipulators, and power grids. Then, we aim to address the challenges associated with acquiring accurate system model information and ensuring stable operation with optimal disturbance attenuation performance. On the one hand, the existing$ {\mathcal {H}}_{\infty }$control methods of PDTSSs only guarantee system stability with a prescribed level of disturbance attenuation, and achieving mean-square stability while maintaining an optimal level of disturbance attenuation remains a challenge. On the other hand, existing$\mathcal {H} _{\infty }$control methods incorporate precise system dynamics that are difficult or costly to acquire in practical systems. To tackle the aforementioned challenges, we propose a data-driven weighted$\mathcal {H} _{\infty }$control approach for PDTSSs with unknown system dynamics. This method ensures that the closed-loop PDTSSs exhibit exponential mean-square stability while achieving an optimal disturbance attenuation level. Furthermore, by leveraging RL algorithms, our controller design process only relies on system data rather than precise system dynamics. Bosen Lian |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Model-Free Inverse H-Infinity Control for Imitation LearningabstractThis paper proposes a data-driven model-free inverse reinforcement learning (IRL) algorithm tailored for solving an inverse$H_{\infty } $control problem. In the problem, both an expert and a learner engage in$H_{\infty } $control to reject disturbances and the learner’s objective is to imitate the expert’s behavior by reconstructing the expert’s performance function through IRL techniques. Introducing zero-sum game principles, we first formulate a model-based single-loop IRL policy iteration algorithm that includes three key steps: updating the policy, action, and performance function using a new correction formula and the standard inverse optimal control principles. Building upon the model-based approach, we propose a model-free single-loop off-policy IRL algorithm that eliminates the need for initial stabilizing policies and prior knowledge of the dynamics of expert and learner. Also, we provide rigorous proof of convergence, stability, and Nash optimality to guarantee the effectiveness and reliability of the proposed algorithms. Furthermore, we showcase the efficiency of our algorithm through simulations and experiments, highlighting its advantages compared to the existing methods.Note to Practitioners—Generally, the cost function for optimal tracking or imitation control is manually defined, which is a challenging task and may result in large tracking errors and slow tracking. In such cases, IRL is a powerful tool for reconstructing proper cost functions. Real-world systems, as demonstrated in practical cases, are frequently exposed to external disturbances and come with unknown models. Employing$H_{\infty } $control is an effective strategy to handle disturbances. However, applying model-free IRL to solve the inverse problem of$H_{\infty } $control for imitation remains an underexplored domain. This paper explores model-free inverse$H_{\infty } $control for imitating expert behaviors, specifically addressing the time-consuming nature of the existing IRL studies that employ a two-loop iteration structure. We propose an efficient single-loop IRL algorithm with a new framework to do this. It is data-driven and model-free, eliminating the need to find an initial stabilizing control policy, which is typically challenging. Additionally, it ensures convergence, stability, and optimality with provable guarantees. Wenqian Xue, Bosen Lian, Yusuf Kartal, Jialu Fan, Tianyou Chai, Frank L. Lewis |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Distributed Global Nash Containment Control and Learning of Multiplayer Multiagent SystemsabstractThis paper studies multiplayer multiagent differential graphical games, where each agent represents a linear dynamic system with multiple heterogeneous control inputs, referred to as players. The objective of the games is to control all followers to achieve containment within the convex hull of the leaders. Meanwhile, each player minimizes its cost function, which is also jointly affected by the other players and neighboring states in the same and neighboring agents. We design global Nash equilibrium (NE) (called Nash) control policies in the games, enabling all players to play the optimal control policies. The global Nash control solution is ensured to be distributed as it is computed using only local agents’ states. The solvability of the games, the asymptotic stability of the local error system, and the distributed global NE property are guaranteed. In addition, a data-driven integral reinforcement learning (RL) algorithm ensures the online computation of the distributed Nash control without requiring explicit knowledge of the system model matrices, by utilizing online measurements of system trajectories. The algorithm solves optimal control and inverse optimal control (IOC) as subproblems. A simulation example validates the algorithms. Bosen Lian, Wenqian Xue, Dariusz G. Mikulski, Gregory R. Hudas |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Inverse Reinforcement Learning for Discrete-Time Systems With Data DropoutsabstractThis article proposes inverse reinforcement learning (IRL) algorithms for tracking control of linear networked control systems under random state dropouts during wireless transmission. The controlled system aims to track the optimal trajectory of a target system, despite the cost function governing the target's behaviors being unknown. The problem is complicated by random state dropouts occurring in two crucial scenarios: 1) the reception of the target's state and 2) feedback of the controlled system's states. Our approach enables the controlled system to infer the target's cost function and optimal control policy, thereby facilitating effective tracking. Specifically, we develop a model-based IRL algorithm that integrates the Smith predictor for state estimation. Then, we advance a state-dropout-aware inverse Q-learning algorithm that uses solely accessible system data, eliminating the need for system models. The theoretical validity of the proposed algorithms is rigorously established, and their practical effectiveness is validated through numerical simulations. Jialu Fan, Wenqian Xue, Bosen Lian, Yunfang Cui, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2025 | Inverse Value Iteration and Q-Learning: Algorithms, Stability, and RobustnessabstractThis article proposes a data-driven model-free inverse Q-learning algorithm for continuous-time linear quadratic regulators (LQRs). Using an agent's trajectories of states and optimal control inputs, the algorithm reconstructs its cost function that captures the same trajectories. This article first poses a model-based inverse value iteration scheme using the agent's system dynamics. Then, an online model-free inverse Q-learning algorithm is developed to recover the agent's cost function only using the demonstrated trajectories. It is more efficient than the existing inverse reinforcement learning (RL) algorithms as it avoids the repetitive RL in inner loops. The proposed algorithms do not need initial stabilizing control policies and solve for unbiased solutions. The proposed algorithm's asymptotic stability, convergence, and robustness are guaranteed. Theoretical analysis and simulation examples show the effectiveness and advantages of the proposed algorithms. Bosen Lian, Wenqian Xue, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Reinforcement Learning-Based H∞ Control of 2-D Markov Jump Roesser Systems With Optimal Disturbance AttenuationabstractThis article investigates model-free reinforcement learning (RL)-based ${\mathcal {H}}_{\infty }$ control problem for discrete-time 2-D Markov jump Roesser systems (2-D MJRSs) with optimal disturbance attenuation level. This is compared to existing studies on ${\mathcal {H}}_{\infty }$ control of 2-D MJRSs with optimal disturbance attenuation levels that are off-line and use full system dynamics. We design a comprehensive model-free RL algorithm to solve optimal ${\mathcal {H}}_{\infty }$ control policy, optimize disturbance attenuation level, and search for the initial stabilizing control policy, via online horizontal and vertical data along 2-D MJRSs trajectories. The optimal disturbance attenuation level is obtained by solving a set of linear matrix inequalities based on online measurement data. The initial stabilizing control policy is obtained via a data-driven parallel value iteration (VI) algorithm. Besides, we further certify the performance including the convergence of the RL algorithm and the asymptotic mean-square stability of the closed-loop systems. Finally, simulation results and comparisons demonstrate the effectiveness of the proposed algorithms. Bosen Lian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Data-Efficient Reinforcement Learning for Complex Nonlinear SystemsabstractThis article proposes a data-efficient model-free reinforcement learning (RL) algorithm using Koopman operators for complex nonlinear systems. A high-dimensional data-driven optimal control of the nonlinear system is developed by lifting it into the linear system model. We use a data-driven model-based RL framework to derive an off-policy Bellman equation. Building upon this equation, we deduce the data-efficient RL algorithm, which does not need a Koopman-built linear system model. This algorithm preserves dynamic information while reducing the required data for optimal control learning. Numerical and theoretical analyses of the Koopman eigenfunctions for dataset truncation are discussed in the proposed model-free data-efficient RL algorithm. We validate our framework on the excitation control of the power system. Vrushabh S. Donge, Bosen Lian, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 2 |
| 2024 | Inverse Q-Learning Using Input-Output DataabstractThis article addresses the problem of learning the objective function of linear discrete-time systems that use static output-feedback (OPFB) control by designing inverse reinforcement learning (RL) algorithms. Most of the existing inverse RL methods require the availability of states and state-feedback control from the expert or demonstrated system. In contrast, this article considers inverse RL in a more general case where the demonstrated system uses static OPFB control with only input-output measurements available. We first develop a model-based inverse RL algorithm to reconstruct an input-output objective function of a demonstrated discrete-time system using its system dynamics and the OPFB gain. This objective function infers the demonstrations and OPFB gain of the demonstrated system. Then, an input-output Q -function is built for the inverse RL problem upon the state reconstruction technique. Given demonstrated inputs and outputs, a data-driven inverse Q -learning algorithm reconstructs the objective function without the knowledge of the demonstrated system dynamics or the OPFB gain. This algorithm yields unbiased solutions even though exploration noises exist. Convergence properties and the nonunique solution nature of the proposed algorithms are studied. Numerical simulation examples verify the effectiveness of the proposed methods. Bosen Lian, Wenqian Xue, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Cybern. | 1 |
| 2024 | Inverse Reinforcement Learning for Trajectory Imitation Using Static Output Feedback ControlabstractThis article studies the trajectory imitation control problem of linear systems suffering external disturbances and develops a data-driven static output feedback (OPFB) control-based inverse reinforcement learning (RL) approach. An Expert-Learner structure is considered where the learner aims to imitate expert's trajectory. Using only measured expert's and learner's own input and output data, the learner computes the policy of the expert by reconstructing its unknown value function weights and thus, imitates its optimally operating trajectory. Three static OPFB inverse RL algorithms are proposed. The first algorithm is a model-based scheme and serves as basis. The second algorithm is a data-driven method using input-state data. The third algorithm is a data-driven method using only input-output data. The stability, convergence, optimality, and robustness are well analyzed. Finally, simulation experiments are conducted to verify the proposed algorithms. Wenqian Xue, Bosen Lian, Jialu Fan, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 2 |
| 2024 | Data-Driven Inverse Reinforcement Learning Control for Linear Multiplayer GamesabstractThis article proposes a data-driven inverse reinforcement learning (RL) control algorithm for nonzero-sum multiplayer games in linear continuous-time differential dynamical systems. The inverse RL problem in the games is solved by a learner reconstructing the unknown expert players' cost functions from demonstrated expert's optimal state and control input trajectories. The learner, thus, obtains the same control feedback gains and trajectories as the expert, only using data along system trajectories without knowing system dynamics. This article first proposes a model-based inverse RL policy iteration framework that has: 1) policy evaluation step for reconstructing cost matrices using Lyapunov functions; 2) state-reward weight improvement step using inverse optimal control (IOC); and 3) policy improvement step using optimal control. Based on the model-based policy iteration algorithm, this article further develops an online data-driven off-policy inverse RL algorithm without knowing any knowledge of system dynamics or expert control gains. Rigorous convergence and stability analysis of the algorithms are provided. It shows that the off-policy inverse RL algorithm guarantees unbiased solutions while probing noises are added to satisfy the persistence of excitation (PE) condition. Finally, two different simulation examples validate the effectiveness of the proposed algorithms. Bosen Lian, Vrushabh S. Donge, Frank L. Lewis, Tianyou Chai, Ali Davoudi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Distributed Minmax Strategy for Multiplayer Games: Stability, Robustness, and AlgorithmsabstractThis article studies a distributed minmax strategy for multiplayer games and develops reinforcement learning (RL) algorithms to solve it. The proposed minmax strategy is distributed, in the sense that it finds each player’s optimal control policy without knowing all the other players’ policies. Each player obtains its distributed control policy by solving a distributed algebraic Riccati equation in a multiplayer noncooperative game. This policy is found against the worst policies of all the other players. We guarantee the existence of distributed minmax solutions and study their$\mathcal {L}_{2}$and asymptotic stabilities. Under mild conditions, the resulting minmax control policies are shown to improve robust gain and phase margins of multiplayer systems compared to the standard linear–quadratic regulator controller. Distributed minmax solutions are found using both model-based policy iteration and data-driven off-policy RL algorithms. Simulation examples verify the proposed formulation and its computational efficiency over the nondistributed Nash solutions. Bosen Lian, Vrushabh S. Donge, Wenqian Xue, Frank L. Lewis, Ali Davoudi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Anomaly Detection and Correction of Optimizing Autonomous Systems With Inverse Reinforcement LearningabstractThis article considers autonomous systems whose behaviors seek to optimize an objective function. This goes beyond standard applications of condition-based maintenance, which seeks to detect faults or failures in nonoptimizing systems. Normal agents optimize a known accepted objective function, whereas abnormal or misbehaving agents may optimize a renegade objective that does not conform to the accepted one. We provide a unified framework for anomaly detection and correction in optimizing autonomous systems described by differential equations using inverse reinforcement learning (RL). We first define several types of anomalies and false alarms, including noise anomaly, objective function anomaly, intention (control gain) anomaly, abnormal behaviors, noise-anomaly false alarms, and objective false alarms. We then propose model-free inverse RL algorithms to reconstruct the objective functions and intentions for given system behaviors. The inverse RL procedure for anomaly detection and correction has the training phase, detection phase, and correction phase. First, inverse RL in the training phase infers the objective function and intention of the normal behavior system using offline stored data. Second, in the detection phase, inverse RL infers the objective function and intention for online observed test system behaviors using online observation data. They are then compared with that of the nominal system to identify anomalies. Third, correction is executed for the anomalous system to learn the normal objective and intention. Simulations and experiments on a quadrotor unmanned aerial vehicle (UAV) verify the proposed methods. Bosen Lian, Yusuf Kartal, Frank L. Lewis, Dariusz G. Mikulski, Gregory R. Hudas, Yan Wan 0001, Ali Davoudi |
IEEE Trans. Cybern. | 1 |
| 2023 | Inverse Reinforcement Learning for Adversarial Apprentice GamesabstractThis article proposes new inverse reinforcement learning (RL) algorithms to solve our defined Adversarial Apprentice Games for nonlinear learner and expert systems. The games are solved by extracting the unknown cost function of an expert by a learner using demonstrated expert's behaviors. We first develop a model-based inverse RL algorithm that consists of two learning stages: an optimal control learning and a second learning based on inverse optimal control. This algorithm also clarifies the relationships between inverse RL and inverse optimal control. Then, we propose a new model-free integral inverse RL algorithm to reconstruct the unknown expert cost function. The model-free algorithm only needs online demonstration of the expert and learner's trajectory data without knowing system dynamics of either the learner or the expert. These two algorithms are further implemented using neural networks (NNs). In Adversarial Apprentice Games, the learner and the expert are allowed to suffer from different adversarial attacks in the learning process. A two-player zero-sum game is formulated for each of these two agents and is solved as a subproblem for the learner in inverse RL. Furthermore, it is shown that the cost functions that the learner learns to mimic the expert's behavior are stabilizing and not unique. Finally, simulations and comparisons show the effectiveness and the superiority of the proposed algorithms. Bosen Lian, Wenqian Xue, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Inverse Reinforcement Q-Learning Through Expert Imitation for Discrete-Time SystemsabstractIn inverse reinforcement learning (RL), there are two agents. An expert target agent has a performance cost function and exhibits control and state behaviors to a learner. The learner agent does not know the expert's performance cost function but seeks to reconstruct it by observing the expert's behaviors and tries to imitate these behaviors optimally by its own response. In this article, we formulate an imitation problem where the optimal performance intent of a discrete-time (DT) expert target agent is unknown to a DT Learner agent. Using only the observed expert's behavior trajectory, the learner seeks to determine a cost function that yields the same optimal feedback gain as the expert's, and thus, imitates the optimal response of the expert. We develop an inverse RL approach with a new scheme to solve the behavior imitation problem. The approach consists of a cost function update based on an extension of RL policy iteration and inverse optimal control, and a control policy update based on optimal control. Then, under this scheme, we develop an inverse reinforcement Q-learning algorithm, which is an extension of RL Q-learning. This algorithm does not require any knowledge of agent dynamics. Proofs of stability, convergence, and optimality are given. A key property about the nonunique solution is also shown. Finally, simulation experiments are presented to show the effectiveness of the new approach. Wenqian Xue, Bosen Lian, Jialu Fan, Patrik Kolaric, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Robustness Analysis of Distributed Kalman Filter for Estimation in Sensor NetworksabstractMotivated by the guaranteed stability margins of linear quadratic regulators (LQRs) and standard Kalman filter (KF) in the frequency domain, this article extends these results to the distributed Kalman-consensus filter (DKCF) for distributed estimation in sensor networks. In particular, we study the robustness margins of DKCF in two cases, one of which is based on the direct target observation while the other uses estimates from neighbor sensors in the network. The loop transfer functions of the two cases are established, and gain and phase margin robustness results are derived for both. The robustness margins of DKCF are improved compared to the single-agent KF. Furthermore, as communication topology varies in sensor networks, graph overall coupling strengths change. We also analyze the correlation between overall coupling strengths and the robustness margins of DKCF. Bosen Lian, Frank L. Lewis, Gary A. Hewer, Katia Estabridis, Tianyou Chai |
IEEE Trans. Cybern. | 1 |
| 2022 | Distributed Kalman Consensus Filter for Estimation With Moving TargetsabstractConsensus-based distributed Kalman filters for estimation with targets have attracted considerable attention. Most of the existing Kalman filters use the average consensus approach, which tends to have a low convergence speed. They also rarely consider the impacts of limited sensing range and target mobility on the information flow topology. In this article, we address these issues by designing a novel distributed Kalman consensus filter (DKCF) with an information-weighted consensus structure for random mobile target estimation in continuous time. A new moving target information-flow topology for the measurement of targets is developed based on the sensors' sensing ranges, targets' random mobility, and local information-weighted neighbors. Novel necessary and sufficient conditions about the convergence of the proposed DKCF are developed. Under these conditions, the estimates of all sensors converge to the consensus values. Simulation and comparative studies show the effectiveness and the superiority of this new DKCF. Bosen Lian, Yan Wan 0001, Ya Zhang 0001, Mushuang Liu, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Cybern. | 1 |
| 2022 | Robust Inverse Q-Learning for Continuous-Time Linear Systems in Adversarial EnvironmentsabstractThis article proposes robust inverse Q -learning algorithms for a learner to mimic an expert's states and control inputs in the imitation learning problem. These two agents have different adversarial disturbances. To do the imitation, the learner must reconstruct the unknown expert cost function. The learner only observes the expert's control inputs and uses inverse Q -learning algorithms to reconstruct the unknown expert cost function. The inverse Q -learning algorithms are robust in that they are independent of the system model and allow for the different cost function parameters and disturbances between two agents. We first propose an offline inverse Q -learning algorithm which consists of two iterative learning loops: 1) an inner Q -learning iteration loop and 2) an outer iteration loop based on inverse optimal control. Then, based on this offline algorithm, we further develop an online inverse Q -learning algorithm such that the learner mimics the expert behaviors online with the real-time observation of the expert control inputs. This online computational method has four functional approximators: a critic approximator, two actor approximators, and a state-reward neural network (NN). It simultaneously approximates the parameters of Q -function and the learner state reward online. Convergence and stability proofs are rigorously studied to guarantee the algorithm performance. Bosen Lian, Wenqian Xue, Frank L. Lewis, Tianyou Chai |
IEEE Trans. Cybern. | 1 |
| 2022 | Inverse Reinforcement Learning in Tracking Control Based on Inverse Optimal ControlabstractThis article provides a novel inverse reinforcement learning (RL) algorithm that learns an unknown performance objective function for tracking control. The algorithm combines three steps: 1) an optimal control update; 2) a gradient descent correction step; and 3) an inverse optimal control (IOC) update. The new algorithm clarifies the relation between inverse RL and IOC. It is shown that the reward weight of an unknown performance objective that generates a target control policy may not be unique. We characterize the set of all weights that generate the same target control policy. We develop a model-based algorithm and, further, two model-free algorithms for systems with unknown model information. Finally, simulation experiments are presented to show the effectiveness of the proposed algorithms. Wenqian Xue, Patrik Kolaric, Jialu Fan, Bosen Lian, Tianyou Chai, Frank L. Lewis |
IEEE Trans. Cybern. | 4 |
| 2022 | Directed Graph Clustering Algorithms, Topology, and Weak LinksabstractIn this article, a general approach for directed graph clustering and two new density-based clustering objectives are presented. First, using an equivalence between the clustering objective functions and a trace maximization expression, the directed graph clustering objectives are converted into the corresponding weighted kernel$k$-means problems. Then, a nonspectral algorithm, which covers both the direction and weight information of the directed graphs, is thus proposed. Next, with Rayleigh’s quotient, the upper and lower bounds of clustering objectives are obtained. After that, we introduce a new definition of weak links to characterize the effectiveness of clustering. Finally, illustrative examples are given to demonstrate effectiveness of the results. This article provides a glance at the potential connection between density-based and pattern-based clustering. Compared with other approaches for directed graph clustering, the method proposed in this article naturally avoids the loss of the nonsymmetric edge data because there is no need for any additional symmetrization. Xiao Zhang 0007, Bosen Lian, Frank L. Lewis, Yan Wan 0001, Daizhan Cheng |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Integrated Sliding Mode Control and Neural Networks Based Packet Disordering Prediction for Nonlinear Networked Control SystemsabstractIn this paper, we propose a new scheme based on neural networks for predicting the packet disordering and sliding mode control (SMC) to stabilize the nonlinear networked control systems (NCSs). It is assumed that the packet disordering is unknown in the NCSs. The stochastic configuration networks (SCNs), which randomly assign the input weights and biases and analytically evaluate the output weights, are designed to solve the problem of unknown packet disordering. A new SMC scheme is developed by integrating the SCNs algorithm to learn and control the system in advance. Specifically, a novel measurement of packet disordering is constructed for the quantization of the packet disordering. In addition, the newest signal principle leads to the existence of stochastic parameters, thereby resulting in a Markovian jumping system. The effectiveness of the proposed approach is verified by some simulation results. Bosen Lian, Qingling Zhang 0001, Jinna Li |
IEEE Trans. Neural Networks Learn. Syst. | 1 |