Xiangnan Zhong

dblp:139/5954 · DBLP profile ↗
← Back
37ranked-venue papers
16as first author
11since 2021 · last 2026
0000-0002-8367-0215ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 14 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Decoupling Shared and Personalized Knowledge: A Dual-Branch Federated Learning Framework for Multi-Domain with Non-IID Data
abstract
Federated learning (FL) enables collaborative model training without centralizing data. In multi-domain scenarios with non-identically and independently distributed (non-IID) data, prediction performance is often hindered by catastrophic forgetting of specialized local knowledge and negative transfer from conflicting client updates. To address these challenges, we propose a personalized FL framework with dual-branch (pFedDB) structure and a two-phase training protocol. The dual-branch architecture separates the model into a shared branch for cross-client aggregation and a private branch that remains on each local client. The private branch is never overwritten by server updates, which prevents the catastrophic forgetting of domain-specific knowledge. This structure also significantly reduces communication overhead per round as only the shared branch is transmitted. To mitigate negative transfer, our two-phase protocol first establishes a personalized knowledge anchor by training a single-branch expert model on each client’s local data. In the second phase, the locally trained model is cloned to initialize private and shared branches. Only the shared branch is aggregated in federated training. This process enables the shared branch to learn a general representation that complements the established local expertise. This design consistently improves the performance of every client over its single-domain baseline, overcoming the challenge of negative transfer among clients. Experiments on our new Chest-X-Ray-4 suite and three public benchmarks show that the proposed pFedDB method obtains 30% saving in communication overhead per round and competitive or better accuracy performance than recent FL methods.
Yiran Pang, Zhen Ni, Xiangnan Zhong
AAAI3
2025 A New Federated Learning Approach for Imbalanced Medical Image Datasets
abstract
Medical image classification plays a vital role in disease diagnosis, but often suffers from class imbalance and privacy concerns, particularly in rare disease categories. Such an imbalance can bias machine learning models toward majority classes, reducing the diagnostic accuracy for minority conditions across geographically diverse populations. To address these challenges, we propose integrating the synthetic minority oversampling technique (SMOTE) into the federated learning (FL) framework for 2D medical image analysis, enabling privacy-preserving and balanced learning. The SMOTE technique enhances neural networks’ generalization by generating informative synthetic samples, outperforming raw duplication, and reducing overfitting. In our approach, by operating on features obtained from deep models like VGG19, we avoid pixel-level interpolation artifacts and retain critical semantic information. The SMOTE technique is applied to the extracted features at the local client level, where it generates new samples for the minority class by interpolating between existing samples to achieve the local data balance. The proposed approach allows each client to balance its local data without compromising privacy, promoting fair learning across all classes in disease classification tasks. We evaluate our approach on two publicly available chest X-ray datasets for binary and multi-class classifications. Our method achieves 83% accuracy and 97% recall in binary classification and 50% accuracy in multiclass settings, highlighting its potential to contribute to diagnostic fairness and robustness in healthcare applications.
Aniruddha Tiwari, Zhen Ni, Xiangnan Zhong, Ahmed Imteaj
ICMLA3
2025 A fast federated reinforcement learning approach with phased weight-adjustment technique
Yiran Pang, Zhen Ni, Xiangnan Zhong
Neurocomputing3
2025 Intelligent Control in Asymmetric Decision-Making: An Event-Triggered RL Approach for Mismatched Uncertainties
abstract
Artificial intelligence (AI)-based multiplayer systems have attracted increasing attention across diverse fields. While most research focuses on simultaneous-move multiplayer games to achieve Nash equilibrium, there are complex applications that involve hierarchical decision-making, where certain players act before others. This power asymmetry increases the complexity of strategic interactions, especially in the presence of mismatched uncertainties that can compromise data reliability and decision-making. To this end, this article develops a novel event-triggered reinforcement learning (RL) approach for hierarchical multiplayer systems with mismatched uncertainties. Specifically, by establishing an auxiliary augment system and designing appropriate cost functions for the high-level leader and low-level followers, we reformulate the hierarchical robust control problem as an optimization task within the Stackelberg–Nash game framework. Furthermore, an event-triggered scheme is designed to reduce the computational overhead and a neural-RL-based method is developed to automatically learn the event-triggered control policies for hierarchical players. Theoretical analyses are conducted to 1) demonstrate the stability preservation of the designed robust-optimal transformation; 2) verify the achievement of Stackelberg–Nash equilibrium under the developed event-triggered policies; and 3) guarantee the boundedness of the impulsive closed-loop system. Finally, the simulation studies validate the effectiveness of the developed method.
Xiangnan Zhong, Zhen Ni
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Federated Learning for Crowd Counting in Smart Surveillance Systems
abstract
Crowd counting in smart surveillance systems plays a crucial role in Internet of Things (IoT) and smart cities, and can affect various aspects, such as public safety, crowd management, and urban planning. Using surveillance data to centrally train a crowd counting model raises significant privacy concerns. Traditional methods try to alleviate the concern by reducing the focus on individuals, but the concern still needs to be thoroughly resolved. In this work, we develop a horizontal federated learning (HFL) framework to train the crowd counting models which can preserve privacy simultaneously. This framework enables the smart surveillance system to learn from model aggregation without accessing the private data stored on local devices. Therefore, it eliminates the need for video data transmission, reduces communication costs, and avoids raw data leakage. Due to the lack of federated learning (FL) crowd counting data sets, we design four non-independent and identically distributed (non-IID) partitioning strategies, including feature-skew, quantity-skew, scene-skew, and time-skew, to simulate real-world FL scenarios. In addition, we present an efficient fully convolutional network (e-FCN) for each client to demonstrate the practical applicability of the proposed framework. The e-FCN adopts an encoder-decoder architecture with fewer parameters, making it communication-friendly and easier to train. This design can achieve competitive performance compared to more complex models in surveillance crowd counting in literature. Finally, we evaluate the proposed HFL framework with e-FCN under our skew strategies on multiple real-world data sets, including crowd surveillance, ShanghaiTech PartB, WorldExpo’10, FDST, CityUHK-X, UCSD, and MALL. Extensive experiments allow us to present our developed Federated Crowd Counting benchmark as a reference for future research and provide guidance for FL algorithm selection in smart surveillance system deployment.
Yiran Pang, Zhen Ni, Xiangnan Zhong
IEEE Internet Things J.3
2024 Kernelized Deep Learning for Matrix Factorization Recommendation System Using Explicit and Implicit Information
abstract
In the current matrix factorization recommendation approaches, the item and the user latent factor vectors are with the same dimension. Thus, the linear dot product is used as the interactive function between the user and the item to predict the ratings. However, the relationship between real users and items is not entirely linear and the existing recommendation model of matrix factorization faces the challenge of data sparsity. To this end, we propose a kernelized deep neural network recommendation model in this article. First, we encode the explicit user-item rating matrix in the form of column vectors and project them to higher dimensions to facilitate the simulation of nonlinear user-item interaction for enhancing the connection between users and items. Second, the algorithm of association rules is used to mine the implicit relation between users and items, rather than simple feature extraction of users or items, for improving the recommendation performance when the datasets are sparse. Third, through the autoencoder and kernelized network processing, the implicit data are connected with the explicit data by the multilayer perceptron network for iterative training instead of doing simple linear weighted summation. Finally, the predicted rating is output through the hidden layer. Extensive experiments were conducted on four public datasets in comparison with several existing well-known methods. The experimental results indicated that our proposed method has obtained improved performance in data sparsity and prediction accuracy.
Xiaoyao Zheng, Zhen Ni, Xiangnan Zhong, Yonglong Luo
IEEE Trans. Neural Networks Learn. Syst.3
2023 An Improved Trust-Region Method for Off-Policy Deep Reinforcement Learning
abstract
Reinforcement learning (RL) is a powerful tool for training agents to interact with complex environments. In particular, trust-region methods are widely used for policy optimization in model-free RL. However, these methods suffer from high sample complexity due to their on-policy nature, which requires interactions with the environment for each update. To address this issue, off-policy trust-region methods have been proposed, but they have shown limited success in highdimensional continuous control problems compared to other off-policy DRL methods. To improve the performance and sample efficiency of trust-region policy optimization, we propose an off-policy trust-region RL algorithm. Our algorithm is based on a theoretical result on a closed-form solution to trust-region policy optimization and is effective in optimizing complex nonlinear policies. We demonstrate the superiority of our algorithm over prior trust-region DRL methods and show that it achieves excellent performance on a range of continuous control tasks in the Multi-Joint dynamics with Contact (MuJoCo) environment, comparable to state-of-the-art off-policy algorithms.
Hepeng Li, Xiangnan Zhong, Haibo He
IJCNN2
2022 Multi-Virtual-Agent Reinforcement Learning for a Stochastic Predator-Prey Grid Environment
abstract
Generalization problem of reinforcement learning is crucial especially for dynamic environments. Conventional reinforcement learning methods solve the problems with some ideal assumptions and are difficult to be applied in dynamic environments directly. In this paper, we propose a new multi-virtual- agent reinforcement learning (MVARL) approach for a predator-prey grid game. The designed method can find the optimal solution even when the predator moves. Specifically, we design virtual agents to interact with simulated changing environments in parallel instead of using actual agents. Moreover, a global agent learns information from these virtual agents and interacts with the actual environment at the same time. This method can not only effectively improve the generalization performance of reinforcement learning in dynamic environments, but also reduce the overall computational cost. Two simulation studies are considered in this paper to validate the effectiveness of the designed method. We also compare the results with the conventional reinforcement learning methods. The results indicate that our proposed method can improve the robustness of reinforcement learning method and contribute to the generalization to certain extent.
Yanbin Lin, Zhen Ni, Xiangnan Zhong
IJCNN3
2022 Semicentralized Deep Deterministic Policy Gradient in Cooperative StarCraft Games
abstract
In this article, we propose a novel semicentralized deep deterministic policy gradient (SCDDPG) algorithm for cooperative multiagent games. Specifically, we design a two-level actor-critic structure to help the agents with interactions and cooperation in the StarCraft combat. The local actor-critic structure is established for each kind of agents with partially observable information received from the environment. Then, the global actor-critic structure is built to provide the local design an overall view of the combat based on the limited centralized information, such as the health value. These two structures work together to generate the optimal control action for each agent and to achieve better cooperation in the games. Comparing with the fully centralized methods, this design can reduce the communication burden by only sending limited information to the global level during the learning process. Furthermore, the reward functions are also designed for both local and global structures based on the agents' attributes to further improve the learning performance in the stochastic environment. The developed method has been demonstrated on several scenarios in a real-time strategy game, i.e., StarCraft. The simulation results show that the agents can effectively cooperate with their teammates and defeat the enemies in various StarCraft scenarios.
Xiangnan Zhong
IEEE Trans. Neural Networks Learn. Syst.2
2021 A Reinforcement Learning-Based Control Approach for Unknown Nonlinear Systems with Persistent Adversarial Inputs
abstract
This paper develops an intelligent control method based on reinforcement learning techniques for unknown nonlinear continuous-time systems in an adversarial environment. The developed method can automatically learn the optimal control input for the system and also predict the worst case adversarial input that one adversary can bring into. Besides, we assume that the agent can only observe partial information of the environment during the learning process. Therefore, a neural network-based observer is developed to adaptively reconstruct the hidden states and dynamics. Then, theoretical analysis is provided to show the stability of the developed intelligent control and the accuracy of the established observer. This method has been applied on a torsional pendulum system and the results demonstrate the effectiveness of the designed approach.
Xiangnan Zhong, Haibo He
IJCNN1
2021 Approximate Dynamic Programming for Nonlinear-Constrained Optimizations
abstract
In this paper, we study the constrained optimization problem of a class of uncertain nonlinear interconnected systems. First, we prove that the solution of the constrained optimization problem can be obtained through solving an array of optimal control problems of constrained auxiliary subsystems. Then, under the framework of approximate dynamic programming, we present a simultaneous policy iteration (SPI) algorithm to solve the Hamilton-Jacobi-Bellman equations corresponding to the constrained auxiliary subsystems. By building an equivalence relationship, we demonstrate the convergence of the SPI algorithm. Meanwhile, we implement the SPI algorithm via an actor-critic structure, where actor networks are used to approximate optimal control policies and critic networks are applied to estimate optimal value functions. By using the least squares method and the Monte Carlo integration technique together, we are able to determine the weight vectors of actor and critic networks. Finally, we validate the developed control method through the simulation of a nonlinear interconnected plant.
Xiong Yang 0001, Haibo He, Xiangnan Zhong
IEEE Trans. Cybern.3
2020 Event-triggered Multi-agent Optimal Regulation Using Adaptive Dynamic Programming
abstract
This paper develops an event-triggered multi-agent control method based on adaptive dynamic programming (ADP) techniques. Different from the traditional ADP-based multi-agent control with fixed sampling period, our method designs an adaptive controller only based on the efficiently reduced samples. The sampling instants are decided by an adaptive triggering condition to guarantee the stability of the event-triggered learning process. The theoretical analysis of the proposed method is also provided in this paper. It is proved that the designed event-triggered ADP controller can make all the agents synchronize to the leader's dynamics with reduced sampled data, and also reach Nash equilibrium at the same time. Therefore, the proposed method can save the computational resources in the learning process. Finally, the simulation results verify the theoretical analysis and also demonstrate the performance of the developed method.
Xiangnan Zhong, Haibo He
IJCNN1
2020 Event-triggered ADP control of a class of non-affine continuous-time nonlinear systems using output information
Yang Yang 0052, Chuang Xu, Dong Yue 0001, Xiangnan Zhong, Xuefeng Si
Neurocomputing4
2020 GrHDP Solution for Optimal Consensus Control of Multiagent Discrete-Time Systems
abstract
This paper develops a new online learning consensus control scheme for multiagent discrete-time systems by goal representation heuristic dynamic programming (GrHDP) techniques. The agents in the whole system are interacted with each other through a communication graph structure. Therefore, each agent can only receive the information from itself and its neighbors. Our goal is to design the GrHDP method to achieve consensus control which makes all the agents track the desired dynamics and simultaneously makes the performance indices reach Nash equilibrium. The new local internal reinforcement signals and local performance indices are provided for each agent and the corresponding distributed control laws are designed. Then, GrHDP algorithm is developed to solve the multiagent consensus control problem with the proof of convergence. It is shown that the designed local internal reinforcement signals are bounded signals and the local performance indices can monotonically converge to their optimal values. Moreover, the desired distributed control laws can also achieve optimal. Two simulation studies, including one with four agents and another with ten agents, are applied to validate the theoretical analysis and also demonstrate the effectiveness of the proposed method.
Xiangnan Zhong, Haibo He
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Prioritizing Useful Experience Replay for Heuristic Dynamic Programming-Based Learning Systems
abstract
The adaptive dynamic programming controller usually needs a long training period because the data usage efficiency is relatively low by discarding the samples once used. Prioritized experience replay (ER) promotes important experiences and is more efficient in learning the control process. This paper proposes integrating an efficient learning capability of prioritized ER design into heuristic dynamic programming (HDP). First, a one time-step backward state-action pair is used to design the ER tuple and, thus, avoids the model network. Second, a systematic approach is proposed to integrate the ER into both critic and action networks of HDP controller design. The proposed approach is tested for two case studies: a cart-pole balancing task and a triple-link pendulum balancing task. For fair comparison, we set the same initial weight parameters and initial starting states for both traditional HDP and the proposed approach under the same simulation environment. The proposed approach improves the required average number of trials to succeed by 60.56% for cart-pole, and 56.89% for triple-link balancing tasks, in comparison with the traditional HDP approach. Also, we have added results of ER-based HDP for comparison. Moreover, theoretical convergence analysis is presented to guarantee the stability of the proposed control design.
Zhen Ni, Naresh Malla, Xiangnan Zhong
IEEE Trans. Cybern.3
2019 Event-Triggered Globalized Dual Heuristic Programming and Its Application to Networked Control Systems
abstract
Networked control systems (NCSs) provide many benefits, such as higher control accuracy and better robustness with the successively increasing computational complexity and communication burden. This results in the traditional adaptive dynamic programming control method having difficulty meeting the real-time requirements of industrial systems. In this paper, a novel event-triggered globalized dual heuristic programming method is proposed to reduce the required samples while guaranteeing the stability of the system. In the proposed method, the NCSs can communicate and update the control law only when the designed event-triggered condition is violated. Furthermore, the Elman neural network, which is a dynamic feedback network with a memory function is implemented to reconstruct the state variables as an approximator, and it depends only on the input and output data. To obtain fewer event-triggered times, two optimization methods, i.e., the unscented Kalman filter and the multiobjective quantum particle swarm optimization, are used to optimize the initial weights of the networks and the positive constant in the event-triggered condition, respectively. The simulation results on industrial system of aluminum electrolysis production are included to verify the performance of the controller.
Xiangnan Zhong, Wei Zhou 0017, Haibo He
IEEE Trans. Ind. Informatics3
2018 Data-Driven Reinforcement Learning Design for Multi-agent Systems with Unknown Disturbances
abstract
In this paper, we develop a new data-driven reinforcement learning method to solve the multi-agent consensus control problem with unknown disturbances. Due to the existence of disturbances, the transmitted information between each pair of the agents becomes unreliable, which makes that the data- driven reinforcement learning based control becomes difficult. Therefore, we solve the problem by developing an appropriate performance index for each agent to convert the robust consensus problem to an auxiliary optimal control problem. The equivalence of the transformation is proved to show that the solution of the auxiliary optimal control system can asymptotically stabilize the original robust system and synchronize all the agents at the same time. Neural network techniques are applied to implement the proposed method. Finally, the simulation results demonstrate the theoretical analysis and verify the effectiveness of the proposed method.
Xiangnan Zhong, Zhen Ni
IJCNN1
2018 Model-Free Adaptive Control for Unknown Nonlinear Zero-Sum Differential Game
abstract
In this paper, we present a new model-free globalized dual heuristic dynamic programming (GDHP) approach for the discrete-time nonlinear zero-sum game problems. First, the online learning algorithm is proposed based on the GDHP method to solve the Hamilton-Jacobi-Isaacs equation associated with optimal regulation control problem. By setting backward one step of the definition of performance index, the requirement of system dynamics, or an identifier is relaxed in the proposed method. Then, three neural networks are established to approximate the optimal saddle point feedback control law, the disturbance law, and the performance index, respectively. The explicit updating rules for these three neural networks are provided based on the data generated during the online learning along the system trajectories. The stability analysis in terms of the neural network approximation errors is discussed based on the Lyapunov approach. Finally, two simulation examples are provided to show the effectiveness of the proposed method.
Xiangnan Zhong, Haibo He, Ding Wang 0001, Zhen Ni
IEEE Trans. Cybern.1
2017 Towards enabling deep learning techniques for adaptive dynamic programming
abstract
Human-level control through deep learning and deep reinforcement learning have revealed the unique and powerful potentials through a very complex Go game. The AlphaGo, developed by Google DeepMind, has beat the top Go game player early this year. The scientific and technological advancement behind the success of AlphaGo attracted researchers from multiple areas, including machine learning, artificial intelligence, computational intelligence and so on. Adaptive dynamic programming (ADP) methods have the similar fundamental principle with reinforcement learning, and show strong performance for continuous time and continuous state systems. Deep learning techniques are also possible to be integrated for ADP designs. In this paper, we discuss the key techniques and components in deep reinforcement learning and then present the successful applications for computer games and maze navigation. Future opportunities for deep learning enabled ADP will be discussed at the end.
Zhen Ni, Naresh Malla, Xiangnan Zhong
IJCNN3
2017 An Event-Triggered ADP Control Approach for Continuous-Time System With Unknown Internal States
abstract
This paper proposes a novel event-triggered adaptive dynamic programming (ADP) control method for nonlinear continuous-time system with unknown internal states. Comparing with the traditional ADP design with a fixed sample period, the event-triggered method samples the state and updates the controller only when it is necessary. Therefore, the computation cost and transmission load are reduced. Usually, the event-triggered method is based on the system entire state which is either infeasible or very difficult to obtain in practice applications. This paper integrates a neural-network-based observer to recover the system internal states from the measurable feedback. Both the proposed observer and the controller are aperiodically updated according to the designed triggering condition. Neural network techniques are applied to estimate the performance index and help calculate the control action. The stability analysis of the proposed method is also demonstrated by Lyapunov construct for both the continuous and jump dynamics. The simulation results verify the theoretical analysis and justify the efficiency of the proposed method.
Xiangnan Zhong, Haibo He
IEEE Trans. Cybern.1
2017 Gr-GDHP: A New Architecture for Globalized Dual Heuristic Dynamic Programming
abstract
Goal representation globalized dual heuristic dynamic programming (Gr-GDHP) method is proposed in this paper. A goal neural network is integrated into the traditional GDHP method providing an internal reinforcement signal and its derivatives to help the control and learning process. From the proposed architecture, it is shown that the obtained internal reinforcement signal and its derivatives can be able to adjust themselves online over time rather than a fixed or predefined function in literature. Furthermore, the obtained derivatives can directly contribute to the objective function of the critic network, whose learning process is thus simplified. Numerical simulation studies are applied to show the performance of the proposed Gr-GDHP method and compare the results with other existing adaptive dynamic programming designs. We also investigate this method on a ball-and-beam balancing system. The statistical simulation results are presented for both the Gr-GDHP and the GDHP methods to demonstrate the improved learning and controlling performance.
Xiangnan Zhong, Zhen Ni, Haibo He
IEEE Trans. Cybern.1
2017 Q-Learning-Based Vulnerability Analysis of Smart Grid Against Sequential Topology Attacks
abstract
Recent studies on sequential attack schemes revealed new smart grid vulnerability that can be exploited by attacks on the network topology. Traditional power systems contingency analysis needs to be expanded to handle the complex risk of cyber-physical attacks. To analyze the transmission grid vulnerability under sequential topology attacks, this paper proposes a Q-learning-based approach to identify critical attack sequences with consideration of physical system behaviors. A realistic power flow cascading outage model is used to simulate the system behavior, where attacker can use the Q-learning to improve the damage of sequential topology attack toward system failures with the least attack efforts. Case studies based on three IEEE test systems have demonstrated the learning ability and effectiveness of Q-learning-based vulnerability analysis.
Jun Yan 0007, Haibo He, Xiangnan Zhong, Yufei Tang
IEEE Trans. Inf. Forensics Secur.3
2017 Adaptive Event-Triggered Control Based on Heuristic Dynamic Programming for Nonlinear Discrete-Time Systems
abstract
This paper presents the design of a novel adaptive event-triggered control method based on the heuristic dynamic programming (HDP) technique for nonlinear discrete-time systems with unknown system dynamics. In the proposed method, the control law is only updated when the event-triggered condition is violated. Compared with the periodic updates in the traditional adaptive dynamic programming (ADP) control, the proposed method can reduce the computation and transmission cost. An actor-critic framework is used to learn the optimal event-triggered control law and the value function. Furthermore, a model network is designed to estimate the system state vector. The main contribution of this paper is to design a new trigger threshold for discrete-time systems. A detailed Lyapunov stability analysis shows that our proposed event-triggered controller can asymptotically stabilize the discrete-time systems. Finally, we test our method on two different discrete-time systems, and the simulation results are included.
Lu Dong 0002, Xiangnan Zhong, Changyin Sun 0001, Haibo He
IEEE Trans. Neural Networks Learn. Syst.2
2017 Event-Triggered Adaptive Dynamic Programming for Continuous-Time Systems With Control Constraints
abstract
In this paper, an event-triggered near optimal control structure is developed for nonlinear continuous-time systems with control constraints. Due to the saturating actuators, a nonquadratic cost function is introduced and the Hamilton-Jacobi-Bellman (HJB) equation for constrained nonlinear continuous-time systems is formulated. In order to solve the HJB equation, an actor-critic framework is presented. The critic network is used to approximate the cost function and the action network is used to estimate the optimal control law. In addition, in the proposed method, the control signal is transmitted in an aperiodic manner to reduce the computational and the transmission cost. Both the networks are only updated at the trigger instants decided by the event-triggered condition. Detailed Lyapunov analysis is provided to guarantee that the closed-loop event-triggered system is ultimately bounded. Three case studies are used to demonstrate the effectiveness of the proposed method.
Lu Dong 0002, Xiangnan Zhong, Changyin Sun 0001, Haibo He
IEEE Trans. Neural Networks Learn. Syst.2
2016 Comparative studies of power grid security with network connectivity and power flow information using unsupervised learning
abstract
The modern electric power grid has become highly integrated in order to increase reliability of power transmission from the generating units to end consumers. This integrated nature and its upgrade toward an intelligent smart grid make the power grid vulnerable when facing cyber or physical attacks as well as intentional attacks. Therefore, determining the most vulnerable components (e.g., buses or generators) is critically important for power grid defense. In this paper, a new definition of load is proposed by taking power flow into consideration in comparison with the load definition based on degree or network connectivity. Unsupervised learning techniques (e.g., K-means algorithm and self-organizing map (SOM)) are introduced to cluster the nodes (i.e., buses) in IEEE-39 bus and IEEE-57 bus benchmarks. Then most vulnerable node in each cluster is determined based on their load information to form initial victim set. We use percentage of failure (PoF) to compare the performance of clustering based approach and traditional load based approach during cascading failure process. With the simulation results, the unsupervised learning (clustering based) approaches are more efficient in finding the most vulnerable nodes and our proposed definition of load is relatively useful in studying power grid security.
Shiva Poudel, Zhen Ni, Xiangnan Zhong, Haibo He
IJCNN3
2016 Convergence analysis of GrDHP-based optimal control for discrete-time nonlinear system
abstract
Adaptive dynamic programming (ADP) has been investigated for its new architectures, algorithms and applications for years. Recently, the goal representation (Gr) design has been demonstrated with promising results to improve ADP control performance from certain perspectives. This paper is focused on the theoretical analysis of the goal representation dual heuristic dynamic programming (GrDHP). Starting from the general formulation of the GrDHP design, we provide the iterative algorithm for this method. The corresponding convergence analysis is showed in terms of the internal reinforcement signal, the performance index, and their derivatives. Our analysis assumes that the system is controllable and stabilizable. Then, neural-network-based implementation of this method is presented. Simulation study validates the theoretical analysis of this paper and also shows the effectiveness of the GrDHP method.
Xiangnan Zhong, Zhen Ni, Haibo He
IJCNN1
2016 Fuzzy-Based Goal Representation Adaptive Dynamic Programming
abstract
In this paper, a novel nonlinear learning controller called fuzzy-based goal representation adaptive dynamic programming (Fuzzy-GrADP) is proposed. In the proposed GrADP method, a goal representation network is introduced to generate an adaptive internal reinforcement signal to the critic network to help the controller provide a general mapping between the input and output actions. Moreover, in the proposed architecture, the action network in the GrADP is improved by using the fuzzy hyperbolic model, which combines the merits of the fuzzy model and the neural network model. Based on the back-propagation technique, the parameters in the membership functions and the fuzzy rules are all undergo training and online adapting. The proposed controller is tested on two numerical benchmarks, and the simulation results show that the proposed controller outperforms the original adaptive dynamic fuzzy controller and the pure neural network-based GrADP controller. In addition, the proposed controller is further applied on a large multimachine power system for static var compensator damping control, where simulation results demonstrate the effectiveness of the proposed approach on real applications. Furthermore, in order to demonstrate the theoretical guarantee of the proposed method, Lyapunov stability analysis to support the proposed Fuzzy-GrADP approach has also been carried out.
Yufei Tang, Haibo He, Zhen Ni, Xiangnan Zhong, Dongbin Zhao, Xin Xu 0001
IEEE Trans. Fuzzy Syst.4
2016 A Theoretical Foundation of Goal Representation Heuristic Dynamic Programming
abstract
Goal representation heuristic dynamic programming (GrHDP) control design has been developed in recent years. The control performance of this design has been demonstrated in several case studies, and also showed applicable to industrial-scale complex control problems. In this paper, we develop the theoretical analysis for the GrHDP design under certain conditions. It has been shown that the internal reinforcement signal is a bounded signal and the performance index can converge to its optimal value monotonically. The existence of the admissible control is also proved. Although the GrHDP control method has been investigated in many areas before, to the best of our knowledge, this is the first study of presenting the theoretical foundation of the internal reinforcement signal and how such an internal reinforcement signal can provide effective information to improve the control performance. Numerous simulation studies are used to validate the theoretical analysis and also demonstrate the effectiveness of the GrHDP design.
Xiangnan Zhong, Zhen Ni, Haibo He
IEEE Trans. Neural Networks Learn. Syst.1
2015 Predictive event-triggered control based on heuristic dynamic programming for nonlinear continuous-time systems
abstract
In this paper, a novel predictive event-triggered control method based on heuristic dynamic programming (HDP) algorithm is developed for nonlinear continuous-time systems. A model network is used to estimate the system state vector, so that the event-triggered instant is available to predict one step ahead of time. Furthermore, an actor-critic structure is used to approximate the optimal event-triggered control law and performance index function. Although event-triggered adaptive dynamic programming (ADP) has been investigated in the community before, to our best knowledge, this is the first study of using a “predictive” approach through a model network to design the event-triggered ADP. This is the key contribution of this work. Compared to the existing event-triggered ADP methods, our simulations demonstrate that the predictive event-triggered approach can achieve improved control performance and lower computational cost in comparison with the existing methods.
Lu Dong 0002, Xiangnan Zhong, Changyin Sun 0001, Haibo He
IJCNN2
2015 A boundedness theoretical analysis for GrADP design: A case study on maze navigation
abstract
A new theoretical analysis towards the goal representation adaptive dynamic programming (GrADP) design proposed in [1], [2] is investigated in this paper. Unlike the proofs of convergence for adaptive dynamic programming (ADP) in literature, here we provide a new insight for the error bound between the estimated value function and the expected value function. Then we employ the critic network in GrADP approach to approximate the Q value function, and use the action network to provide the control policy. The goal network is adopted to provide the internal reinforcement signal for the critic network over time. Finally, we illustrate that the estimated Q value function is close to the expected value function in an arbitrary small bound on the maze navigation example.
Zhen Ni, Xiangnan Zhong, Haibo He
IJCNN2
2015 Event-triggered adaptive dynamic programming for continuous-time nonlinear system using measured input-output data
abstract
In this paper, we propose a novel event-triggered adaptive dynamic programming (ADP) method using only the input-output data. Event-triggered method is widely used for its computational efficiency capacity. Comparing with the traditional method which updates the controller periodically, the event-triggered method only updates the controller when it is necessary and therefore the computation is reduced. Generally, the triggered condition is based on the system current and sampled states. In this paper, we consider a neural-network-based observer to recover the system dynamics using the measured input-output data. The triggered instants are calculated according to the recovered state. Stability analysis of the proposed approach is presented. We verify our proposed method through a robot-arm example.
Xiangnan Zhong, Zhen Ni, Haibo He
IJCNN1
2015 A neural network based online learning and control approach for Markov jump systems
Xiangnan Zhong, Haibo He, Huaguang Zhang, Zhanshan Wang 0001
Neurocomputing1
2015 Model-Free Dual Heuristic Dynamic Programming
abstract
Model-based dual heuristic dynamic programming (MB-DHP) is a popular approach in approximating optimal solutions in control problems. Yet, it usually requires offline training for the model network, and thus resulting in extra computational cost. In this brief, we propose a model-free DHP (MF-DHP) design based on finite-difference technique. In particular, we adopt multilayer perceptron with one hidden layer for both the action and the critic networks design, and use delayed objective functions to train both the action and the critic networks online over time. We test both the MF-DHP and MB-DHP approaches with a discrete time example and a continuous time example under the same parameter settings. Our simulation results demonstrate that the MF-DHP approach can obtain a control performance competitive with that of the traditional MB-DHP approach while requiring less computational resources.
Zhen Ni, Haibo He, Xiangnan Zhong, Danil V. Prokhorov
IEEE Trans. Neural Networks Learn. Syst.3
2014 Data-driven partially observable dynamic processes using adaptive dynamic programming
abstract
Adaptive dynamic programming (ADP) has been widely recognized as one of the “core methodologies” to achieve optimal control for intelligent systems in Markov decision process (MDP). Generally, ADP control design requires all the information of the system dynamics. However, in many practical situations, the measured input and output data can only represent part of the system states. This means the complete information of the system cannot be available in many real-world cases, which narrows the range of application of the ADP design. In this paper, we propose a data-driven ADP method to stabilize the system with partially observable dynamics based on neural network techniques. A state network is integrated into the typical actor-critic architecture to provide an estimated state from the measured input/output sequences. The theoretical analysis and the stability discussion of this data-driven ADP method are also provided. Two examples are studied to verify our proposed method.
Xiangnan Zhong, Zhen Ni, Yufei Tang, Haibo He
ADPRL1
2014 Event-triggered reinforcement learning approach for unknown nonlinear continuous-time system
abstract
This paper provides an adaptive event-triggered method using adaptive dynamic programming (ADP) for the nonlinear continuous-time system. Comparing to the traditional method with fixed sampling period, the event-triggered method samples the state only when an event is triggered and therefore the computational cost is reduced. We demonstrate the theoretical analysis on the stability of the event-triggered method, and integrate it with the ADP approach. The system dynamics are assumed unknown. The corresponding ADP algorithm is given and the neural network techniques are applied to implement this method. The simulation results verify the theoretical analysis and justify the efficiency of the proposed event-triggered technique using the ADP approach.
Xiangnan Zhong, Zhen Ni, Haibo He, Xin Xu 0001, Dongbin Zhao
IJCNN1
2014 Optimal Control for Unknown Discrete-Time Nonlinear Markov Jump Systems Using Adaptive Dynamic Programming
abstract
In this paper, we develop and analyze an optimal control method for a class of discrete-time nonlinear Markov jump systems (MJSs) with unknown system dynamics. Specifically, an identifier is established for the unknown systems to approximate system states, and an optimal control approach for nonlinear MJSs is developed to solve the Hamilton-Jacobi-Bellman equation based on the adaptive dynamic programming technique. We also develop detailed stability analysis of the control approach, including the convergence of the performance index function for nonlinear MJSs and the existence of the corresponding admissible control. Neural network techniques are used to approximate the proposed performance index function and the control law. To demonstrate the effectiveness of our approach, three simulation studies, one linear case, one nonlinear case, and one single link robot arm case, are used to validate the performance of the proposed optimal control method.
Xiangnan Zhong, Haibo He, Huaguang Zhang, Zhanshan Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 Robust controller design of continuous-time nonlinear system using neural network
abstract
In this paper, we propose an optimal control method based on the solution of Hamilton-Jacobi-Bellman (HJB) equation for the continuous-time nonlinear system with bounded unknown perturbation. The robust control system is converted into the corresponding optimal control system with appropriate performance index and the equivalence of the transformation is proved, i.e., the solution of the optimal control problem can globally asymptotically stabilize the robust control system. Adaptive dynamic programming (ADP) based approach is presented to iteratively approximate the optimal performance index and obtain the optimal control policy. A neural network with adaptive weights is applied to implement this approach. An example is given to illustrate the proposed method.
Xiangnan Zhong, Haibo He, Danil V. Prokhorov
IJCNN1