Zhen Ni

dblp:55/9614 · DBLP profile ↗
← Back
64ranked-venue papers
17as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 13 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decoupling Shared and Personalized Knowledge: A Dual-Branch Federated Learning Framework for Multi-Domain with Non-IID Data
abstract
Federated learning (FL) enables collaborative model training without centralizing data. In multi-domain scenarios with non-identically and independently distributed (non-IID) data, prediction performance is often hindered by catastrophic forgetting of specialized local knowledge and negative transfer from conflicting client updates. To address these challenges, we propose a personalized FL framework with dual-branch (pFedDB) structure and a two-phase training protocol. The dual-branch architecture separates the model into a shared branch for cross-client aggregation and a private branch that remains on each local client. The private branch is never overwritten by server updates, which prevents the catastrophic forgetting of domain-specific knowledge. This structure also significantly reduces communication overhead per round as only the shared branch is transmitted. To mitigate negative transfer, our two-phase protocol first establishes a personalized knowledge anchor by training a single-branch expert model on each client’s local data. In the second phase, the locally trained model is cloned to initialize private and shared branches. Only the shared branch is aggregated in federated training. This process enables the shared branch to learn a general representation that complements the established local expertise. This design consistently improves the performance of every client over its single-domain baseline, overcoming the challenge of negative transfer among clients. Experiments on our new Chest-X-Ray-4 suite and three public benchmarks show that the proposed pFedDB method obtains 30% saving in communication overhead per round and competitive or better accuracy performance than recent FL methods.
Yiran Pang, Zhen Ni, Xiangnan Zhong
AAAI2
2025 A New Federated Learning Approach for Imbalanced Medical Image Datasets
abstract
Medical image classification plays a vital role in disease diagnosis, but often suffers from class imbalance and privacy concerns, particularly in rare disease categories. Such an imbalance can bias machine learning models toward majority classes, reducing the diagnostic accuracy for minority conditions across geographically diverse populations. To address these challenges, we propose integrating the synthetic minority oversampling technique (SMOTE) into the federated learning (FL) framework for 2D medical image analysis, enabling privacy-preserving and balanced learning. The SMOTE technique enhances neural networks’ generalization by generating informative synthetic samples, outperforming raw duplication, and reducing overfitting. In our approach, by operating on features obtained from deep models like VGG19, we avoid pixel-level interpolation artifacts and retain critical semantic information. The SMOTE technique is applied to the extracted features at the local client level, where it generates new samples for the minority class by interpolating between existing samples to achieve the local data balance. The proposed approach allows each client to balance its local data without compromising privacy, promoting fair learning across all classes in disease classification tasks. We evaluate our approach on two publicly available chest X-ray datasets for binary and multi-class classifications. Our method achieves 83% accuracy and 97% recall in binary classification and 50% accuracy in multiclass settings, highlighting its potential to contribute to diagnostic fairness and robustness in healthcare applications.
Aniruddha Tiwari, Zhen Ni, Xiangnan Zhong, Ahmed Imteaj
ICMLA2
2025 A Visual Feature-Based Inverse Reinforcement Learning Approach for Reward Reconstruction from Image Observations
abstract
Inverse reinforcement learning (IRL) approaches are useful for applications without a defined reward during the Markov decision process, as they reconstruct the reward function and build the policy from expert demonstrations. Many IRL methods adopt the feature space with both state and action variables. However, a vast amount of demonstrations in the format of video or image may not include an action space. To this end, we propose a visual feature-based inverse reinforcement learning approach (VF-IRL) targeting image observations. Specifically, we reformulate the reward recovering via learning from the image observations only. We use a convolutional neural network (CNN) technique to extract key visual features from images, eliminating the need for predefined features or model-based information. By calculating the loss function based on the difference between expert and learner feature expectations, the gradient computation is independent of the number of learner’s trajectories. Thus, the proposed method reduces the computational cost, supports continuous-time scenarios, and further enhances the scalability. Experimental study on the CarRacing example demonstrates the proposed method’s effectiveness in recovering the reward function under diverse scenarios and outperforming state-of-the-art imitation learning and inverse reinforcement learning approaches.
Yanbin Lin, Zhen Ni
IJCNN2
2025 Multimodal information-driven intelligent damage detection framework for full components of highway prefabricated bridges
Xinxiang Xu, Jiawang Zhan, Zhihang Wang, Zhen Ni
Adv. Eng. Informatics4
2025 A fast federated reinforcement learning approach with phased weight-adjustment technique
Yiran Pang, Zhen Ni, Xiangnan Zhong
Neurocomputing2
2025 A Robust Multi-Virtual-Agent Inverse Reinforcement Learning Approach With Data Aggregation for Perturbed Environments
abstract
Learning control in environments with uncertainties and perturbations remains a challenging issue in the field of artificial intelligence. Though conventional imitation learning (IL) and inverse reinforcement learning (IRL) methods have made some progress in handling perturbations, the repeatability and resilience are somehow limited. To alleviate this issue, we propose a multi-virtual-agent IRL (MVIRL) method to produce stable policies. Specifically, we design multiple virtual agents interacting with pertinent environments. The proposed MVIRL method can recover a resilient reward function from multiple demonstration sources. This recovered reward function provides adequate information and comprehensive coverage of perturbations by considering the upper and lower bounds. Moreover, using maximum discrimination for the worst case and applying data aggregation, the proposed method requires fewer demonstrations than existing methods and improves the ability to handle uncertainties. Case studies with gravity and noise interruptions are considered to validate the effectiveness of the proposed method. The proposed MVIRL method obtains better performance than comparable IL and IRL methods in terms of average return (Avg Return) and standard deviation (SD) metrics, and it is more robust to the level of uncertainties.
Yanbin Lin, Zhen Ni
IEEE Trans. Neural Networks Learn. Syst.2
2025 Intelligent Control in Asymmetric Decision-Making: An Event-Triggered RL Approach for Mismatched Uncertainties
abstract
Artificial intelligence (AI)-based multiplayer systems have attracted increasing attention across diverse fields. While most research focuses on simultaneous-move multiplayer games to achieve Nash equilibrium, there are complex applications that involve hierarchical decision-making, where certain players act before others. This power asymmetry increases the complexity of strategic interactions, especially in the presence of mismatched uncertainties that can compromise data reliability and decision-making. To this end, this article develops a novel event-triggered reinforcement learning (RL) approach for hierarchical multiplayer systems with mismatched uncertainties. Specifically, by establishing an auxiliary augment system and designing appropriate cost functions for the high-level leader and low-level followers, we reformulate the hierarchical robust control problem as an optimization task within the Stackelberg–Nash game framework. Furthermore, an event-triggered scheme is designed to reduce the computational overhead and a neural-RL-based method is developed to automatically learn the event-triggered control policies for hierarchical players. Theoretical analyses are conducted to 1) demonstrate the stability preservation of the designed robust-optimal transformation; 2) verify the achievement of Stackelberg–Nash equilibrium under the developed event-triggered policies; and 3) guarantee the boundedness of the impulsive closed-loop system. Finally, the simulation studies validate the effectiveness of the developed method.
Xiangnan Zhong, Zhen Ni
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Federated Learning for Crowd Counting in Smart Surveillance Systems
abstract
Crowd counting in smart surveillance systems plays a crucial role in Internet of Things (IoT) and smart cities, and can affect various aspects, such as public safety, crowd management, and urban planning. Using surveillance data to centrally train a crowd counting model raises significant privacy concerns. Traditional methods try to alleviate the concern by reducing the focus on individuals, but the concern still needs to be thoroughly resolved. In this work, we develop a horizontal federated learning (HFL) framework to train the crowd counting models which can preserve privacy simultaneously. This framework enables the smart surveillance system to learn from model aggregation without accessing the private data stored on local devices. Therefore, it eliminates the need for video data transmission, reduces communication costs, and avoids raw data leakage. Due to the lack of federated learning (FL) crowd counting data sets, we design four non-independent and identically distributed (non-IID) partitioning strategies, including feature-skew, quantity-skew, scene-skew, and time-skew, to simulate real-world FL scenarios. In addition, we present an efficient fully convolutional network (e-FCN) for each client to demonstrate the practical applicability of the proposed framework. The e-FCN adopts an encoder-decoder architecture with fewer parameters, making it communication-friendly and easier to train. This design can achieve competitive performance compared to more complex models in surveillance crowd counting in literature. Finally, we evaluate the proposed HFL framework with e-FCN under our skew strategies on multiple real-world data sets, including crowd surveillance, ShanghaiTech PartB, WorldExpo’10, FDST, CityUHK-X, UCSD, and MALL. Extensive experiments allow us to present our developed Federated Crowd Counting benchmark as a reference for future research and provide guidance for FL algorithm selection in smart surveillance system deployment.
Yiran Pang, Zhen Ni, Xiangnan Zhong
IEEE Internet Things J.2
2024 Safe Reinforcement Learning and Adaptive Optimal Control With Applications to Obstacle Avoidance Problem
abstract
This paper presents a novel composite obstacle avoidance control method to generate safe motion trajectories for autonomous systems in an adaptive manner. First, system safety is described using forward invariance, and the barrier function is encoded into the cost function such that the obstacle avoidance problem can be characterized by an infinite-horizon optimal control problem. Next, a safe reinforcement learning framework is proposed by combining model-based policy iteration and state-following-based approximation. Upon real-time data and extrapolated experience data, this learning design is implemented through the actor-critic structure, in which critic networks are tuned by gradient-descent adaption and actor networks produce adaptive control policies via gradient projection. Then, system stability and weight convergence are theoretically analyzed using Lyapunov method. Finally, the proposed learning-based controller is demonstrated on a two-dimensional single integrator system and a nonlinear unicycle kinematic system. Simulation results reveal that the system or agent can smoothly reach the target point while keeping a safe distance from each obstacle; at the same time, other three avoidance control methods are used to provide side-by-side comparisons and to verify some claimed advantages of the present method.Note to Practitioners—This paper is motivated by the obstacle avoidance problem of real-time navigation of an agent to the target point, which applies to practical autonomous systems such as vehicles and robots. Pre-generative methods and reactive methods have been widely employed to generate safe motion trajectories in the obstacle environment. However, these methods cannot strike a good balance between safety and optimality. In this paper, the obstacle avoidance problem is formulated in the sense of optimal control, and a safe reinforcement learning method is designed to generate safe motion trajectories. This method combines the advantages of model-based policy iteration and state-following-based approximation, in which the former ensures regional optimality while the latter ensures local safety. Based on the proposed adaptive tuning laws, engineers are able to design learning-based avoidance controllers in the environment with static obstacles. In future research, we will address the dynamic avoidance problem against moving obstacles.
Ke Wang 0037, Chaoxu Mu, Zhen Ni, Derong Liu 0001
IEEE Trans Autom. Sci. Eng.3
2024 Kernelized Deep Learning for Matrix Factorization Recommendation System Using Explicit and Implicit Information
abstract
In the current matrix factorization recommendation approaches, the item and the user latent factor vectors are with the same dimension. Thus, the linear dot product is used as the interactive function between the user and the item to predict the ratings. However, the relationship between real users and items is not entirely linear and the existing recommendation model of matrix factorization faces the challenge of data sparsity. To this end, we propose a kernelized deep neural network recommendation model in this article. First, we encode the explicit user-item rating matrix in the form of column vectors and project them to higher dimensions to facilitate the simulation of nonlinear user-item interaction for enhancing the connection between users and items. Second, the algorithm of association rules is used to mine the implicit relation between users and items, rather than simple feature extraction of users or items, for improving the recommendation performance when the datasets are sparse. Third, through the autoencoder and kernelized network processing, the implicit data are connected with the explicit data by the multilayer perceptron network for iterative training instead of doing simple linear weighted summation. Finally, the predicted rating is output through the hidden layer. Extensive experiments were conducted on four public datasets in comparison with several existing well-known methods. The experimental results indicated that our proposed method has obtained improved performance in data sparsity and prediction accuracy.
Xiaoyao Zheng, Zhen Ni, Xiangnan Zhong, Yonglong Luo
IEEE Trans. Neural Networks Learn. Syst.2
2022 An Efficient Distributed Reinforcement Learning for Enhanced Multi-Microgrid Management
abstract
Economic dispatch in a multi-microgrid (MMG) system involves an increasing number of states from distributed energy resources (DERs) compared to a single microgrid. In these cases, traditional reinforcement learning (RL) approaches may become computationally expensive or less effective in finding the least-cost solution. This paper presents a novel RL approach that employs local learning agents to interact with individual microgrid environments in a distributed manner and a global agent to search for actions to minimize system cost at the MMG system level. The proposed distributed RL framework is more efficient in learning the dispatch policy compared to conventional approaches. Case studies are performed on a 3-microgrid system with different types of DERs. Results substantiate the effectiveness of the proposed approach in comparison with conventional methods in terms of operation costs, computation time, and peak-to-average ratio.
Avijit Das, Zhen Ni, Di Wu 0021
IJCNN2
2022 Multi-Virtual-Agent Reinforcement Learning for a Stochastic Predator-Prey Grid Environment
abstract
Generalization problem of reinforcement learning is crucial especially for dynamic environments. Conventional reinforcement learning methods solve the problems with some ideal assumptions and are difficult to be applied in dynamic environments directly. In this paper, we propose a new multi-virtual- agent reinforcement learning (MVARL) approach for a predator-prey grid game. The designed method can find the optimal solution even when the predator moves. Specifically, we design virtual agents to interact with simulated changing environments in parallel instead of using actual agents. Moreover, a global agent learns information from these virtual agents and interacts with the actual environment at the same time. This method can not only effectively improve the generalization performance of reinforcement learning in dynamic environments, but also reduce the overall computational cost. Two simulation studies are considered in this paper to validate the effectiveness of the designed method. We also compare the results with the conventional reinforcement learning methods. The results indicate that our proposed method can improve the robustness of reinforcement learning method and contribute to the generalization to certain extent.
Yanbin Lin, Zhen Ni, Xiangnan Zhong
IJCNN2
2022 Hierarchical attention and feature projection for click-through rate prediction
Chengliang Zhong, Shouxiang Fan, Xiaodong Mu, Zhen Ni
Appl. Intell.5
2022 An approach of method-level bug localization
abstract
Abstract Bug localization is an important field in software engineering research. The traditional bug localization approaches based on information retrieval separate words through lexical analysis. In this way, the comments of the source code are ignored or treated as plain text, which will lose some semantic information. In this paper, MBL_SHL, an automatic Method‐level Bug Localization approach, which utilises code Summarization, Historical fixed bugs and code Length, is presented. Based on the code summarization technology, this approach first supplements the comment for uncommented code, and then calculates the Word2vec vector and Term Frequency–Inverse Document Frequency vector for the bug report, methods and comments, respectively. After that the authors calculate separately the similarity between the bug report and each method, the bug report and each comment. The code length information and historical fix information are also considered as a weight and a part of the score, respectively, to calculate the final score of each method. Finally, the scores are sorted to determine the list of methods that may need to be modified when fixing the software bugs. We built a method‐granular bug localization dataset, which contains five open‐source projects. The experimental results show that the proposed approach significantly outperforms the existing approaches on the method level.
Zhen Ni, Lili Bo, Bin Li 0006, Tianhao Chen, Xiaobing Sun 0001, Xiaoxue Wu 0001
IET Softw.1
2022 Adaptive Learning and Sampled-Control for Nonlinear Game Systems Using Dynamic Event-Triggering Strategy
abstract
Static event-triggering-based control problems have been investigated when implementing adaptive dynamic programming algorithms. The related triggering rules are only current state-dependent without considering previous values. This motivates our improvements. This article aims to provide an explicit formulation for dynamic event-triggering that guarantees asymptotic stability of the event-sampled nonzero-sum differential game system and desirable approximation of critic neural networks. This article first deduces the static triggering rule by processing the coupling terms of Hamilton-Jacobi equations, and then, Zeno-free behavior is realized by devising an exponential term. Subsequently, a novel dynamic-triggering rule is devised into the adaptive learning stage by defining a dynamic variable, which is mathematically characterized by a first-order filter. Moreover, mathematical proofs illustrate the system stability and the weight convergence. Theoretical analysis reveals the characteristics of dynamic rule and its relations with the static rules. Finally, a numerical example is presented to substantiate the established claims. The comparative simulation results confirm that both static and dynamic strategies can reduce the communication that arises in the control loops, while the latter undertakes less communication burden due to fewer triggered events.
Chaoxu Mu, Ke Wang 0037, Zhen Ni
IEEE Trans. Neural Networks Learn. Syst.3
2020 One-Shot Unsupervised Domain Adaptation for Object Detection
abstract
The existing unsupervised domain adaptation (UDA) methods require not only labeled source samples but also a large number of unlabeled target samples for domain adaptation. Collecting these target samples is generally time-consuming, which hinders the rapid deployment of these UDA methods in new domains. Besides, most of these UDA methods are developed for image classification. In this paper, we address a new problem called one-shot unsupervised domain adaptation for object detection, where only one unlabeled target sample is available. To the best of our knowledge, this is the first time this problem is investigated. To solve this problem, a one-shot feature alignment (OSFA) algorithm is proposed to align the low-level features of the source domain and the target domain. Specifically, the domain shift is reduced by aligning the average activation of the feature maps in the lower layer of CNN. The proposed OSFA is evaluated under two scenarios: adapting from clear weather to foggy weather; adapting from synthetic images to real-world images. Experimental results show that the proposed OSFA can significantly improve the object detection performance in target domain compared to the baseline model without domain adaptation.
Zhiqiang Wan, Lusi Li, Hepeng Li, Haibo He, Zhen Ni
IJCNN5
2020 Analyzing bug fix for automatic bug cause classification
abstract
During the bug fixing process, developers usually need to analyze the source code to induce the bug cause, which is useful for bug understanding and localization. The bug fixes of historical bugs usually reflects the bug causes when fixing them. This paper aims at exploiting the corresponding relationship between bug causes and bug fixes to automatically classify bugs into their cause categories. First, we define the code-related bug classification criterion from the perspective of the cause of bugs. Then, we propose a new model to exploit the knowledge in the bug fix by constructing fix trees from the diff source code at Abstract Syntax Tree (AST) level, and representing each fix tree based on the encoding method of Tree-based Convolutional Neural Network (TBCNN). Finally, the corresponding relationship between bug causes and bug fixes is analyzed by automatically classifying bugs into their cause categories. We collected 2000 real-world bugs from two open source projects Mozilla and Radare2 to evaluate our approach. The experimental results show the existence of observational correlation between the bug fix and the cause of the historical bugs, and the proposed fix tree can effectively express the characteristics of the historical bugs for bug cause classification.
Zhen Ni, Bin Li 0006, Xiaobing Sun 0001, Tianhao Chen, Ben Tang, Xinchen Shi
J. Syst. Softw.1
2020 Robot-Assisted Pedestrian Regulation Based on Deep Reinforcement Learning
abstract
Pedestrian regulation can prevent crowd accidents and improve crowd safety in densely populated areas. Recent studies use mobile robots to regulate pedestrian flows for desired collective motion through the effect of passive human-robot interaction (HRI). This paper formulates a robot motion planning problem for the optimization of two merging pedestrian flows moving through a bottleneck exit. To address the challenge of feature representation of complex human motion dynamics under the effect of HRI, we propose using a deep neural network to model the mapping from the image input of pedestrian environments to the output of robot motion decisions. The robot motion planner is trained end-to-end using a deep reinforcement learning algorithm, which avoids hand-crafted feature detection and extraction, thus improving the learning capability for complex dynamic problems. Our proposed approach is validated in simulated experiments, and its performance is evaluated. The results demonstrate that the robot is able to find optimal motion decisions that maximize the pedestrian outflow in different flow conditions, and the pedestrian-accumulated outflow increases significantly compared to cases without robot regulation and with random robot motion.
Zhiqiang Wan, Chao Jiang 0001, Muhammad Fahad 0003, Zhen Ni, Yi Guo 0004, Haibo He
IEEE Trans. Cybern.4
2020 Cooperative Differential Game-Based Optimal Control and Its Application to Power Systems
abstract
Differential games have been extensively applied to optimal control problems. Nash equilibrium captures the tradeoff among players' policies when every player independently tries to minimize a predefined index. When considering potential cooperation, Pareto equilibrium plays an important role in cooperative differential games. This article studies the cooperative control of multiplayer systems on the quadratic infinite horizon. First, by defining a joint cost function using a parameter set, a cooperative differential game is reformulated as a general optimal control problem, where all players form a grand coalition. Then, the joint cost function is approximated by a critic neural network, and for the first time, a novel adaptive dynamic programming algorithm with two learning stages is proposed to determine the parameter selection and then obtain Pareto optimal solutions. A numerical example demonstrates that this algorithm can achieve optimal policies and Pareto frontier. As for its application, the cooperative control of a two-area interconnected power system is investigated, where the primary frequency control and secondary frequency control are regarded as two players. Simulation results indicate that the proposed scheme can obtain binding cooperation agreements, such that cooperative control scheme can get better overall performance compared to Nash control method and another three control methods.
Chaoxu Mu, Ke Wang 0037, Zhen Ni, Changyin Sun 0001
IEEE Trans. Ind. Informatics3
2020 Low-cohesion differential privacy protection for industrial Internet
Jun Hou 0002, Qianmu Li, Shicheng Cui, Shunmei Meng, Zhen Ni
J. Supercomput.6
2020 A Learning-Based Solution for an Adversarial Repeated Game in Cyber-Physical Power Systems
abstract
Due to the rapidly expanding complexity of the cyber-physical power systems, the probability of a system malfunctioning and failing is increasing. Most of the existing works combining smart grid (SG) security and game theory fail to replicate the adversarial events in the simulated environment close to the real-life events. In this article, a repeated game is formulated to mimic the real-life interactions between the adversaries of the modern electric power system. The optimal action strategies for different environment settings are analyzed. The advantage of the repeated game is that the players can generate actions independent of the previous actions' history. The solution of the game is designed based on the reinforcement learning algorithm, which ensures the desired outcome in favor of the players. The outcome in favor of a player means achieving higher mixed strategy payoff compared to the other player. Different from the existing game-theoretic approaches, both the attacker and the defender participate actively in the game and learn the sequence of actions applying to the power transmission lines. In this game, we consider several factors (e.g., attack and defense costs, allocated budgets, and the players' strengths) that could affect the outcome of the game. These considerations make the game close to real-life events. To evaluate the game outcome, both players' utilities are compared, and they reflect how much power is lost due to the attacks and how much power is saved due to the defenses. The players' favorable outcome is achieved for different attack and defense strengths (probabilities). The IEEE 39 bus system is used here as the test benchmark. Learned attack and defense strategies are applied in a simulated power system environment (PowerWorld) to illustrate the postattack effects on the system.
Shuva Paul, Zhen Ni, Chaoxu Mu
IEEE Trans. Neural Networks Learn. Syst.2
2019 A Composite Extended Nearest Neighbor Model for Day-Ahead Load Forecasting
abstract
Day-ahead load forecasting is an important task for the reduction of electricity waste and efficient management of a smart grid. The electricity load profile data reveals the correlation of electricity load demand with weather condition, day type (working day or holiday), time of the day and season of the year. Thus the load forecasting problem has a high degree of complexity with consideration of those variables as input. To solve the problem of day-ahead short-term load forecasting (STLF), the proposed solution first classifies load profile data into different classes. To this end, a recent developed classification approach called extended nearest neighbor (ENN) algorithm is adopted. Then, a composite ENN model is proposed for day-ahead load forecasting. The composite ENN model consists of three individual ENN models which are combined together by tuned weight factors for predicting final forecasting output. Unlike other statistical and computational intelligence approaches, the composite ENN model predicts electricity load demand from the maximum gain of intra-class coherence. By exploiting intraclass coherence from the generalized class-wise statistic of all available training samples, the composite ENN algorithm is able to learn from global distribution and therefore improve the accuracy of load forecasting. The proposed method is validated on two case study: (i) Australian National Energy Market Data and (ii) Brookings, South Dakota, USA Data. For case study 1, mean absolute percent error (MAPE) of composite ENN based load forecasting is decreased by 44.68% compared to composite kNN based load forecasting and mean absolute error (MAE) is decreased by 45.52%. Similarly for case study 2, the decrease of MAPE and MAE values are 27.72% and 31.65% respectively.
Md. Rashedul Haq, Zhen Ni
IJCNN2
2019 Advancing Network Function Virtualization Platforms with Programmable NICs
abstract
Network Function Virtualization seeks to run high performance middleboxes in a flexible, more configurable software environment. Even with advances such as kernel bypass and zero-copy IO, middlebox platforms still struggle to meet stringent throughput and latency requirements. To achieve line rates as network bandwidths rise, these platforms often must make tradeoffs such as inefficiently dedicating more CPU cores or weakening security and isolation properties. In this paper we explore how advances in programmable “smart NICs” can be leveraged by software middlebox platforms to improve performance, resource efficiency, and security. Our evaluation shows several use cases for smart NICs, which improve performance significantly while reducing resource consumption and providing strong isolation.
Zhen Ni, Guyue Liu, Dennis Afanasev, Timothy Wood 0001, Jinho Hwang
LANMAN1
2019 Prioritizing Useful Experience Replay for Heuristic Dynamic Programming-Based Learning Systems
abstract
The adaptive dynamic programming controller usually needs a long training period because the data usage efficiency is relatively low by discarding the samples once used. Prioritized experience replay (ER) promotes important experiences and is more efficient in learning the control process. This paper proposes integrating an efficient learning capability of prioritized ER design into heuristic dynamic programming (HDP). First, a one time-step backward state-action pair is used to design the ER tuple and, thus, avoids the model network. Second, a systematic approach is proposed to integrate the ER into both critic and action networks of HDP controller design. The proposed approach is tested for two case studies: a cart-pole balancing task and a triple-link pendulum balancing task. For fair comparison, we set the same initial weight parameters and initial starting states for both traditional HDP and the proposed approach under the same simulation environment. The proposed approach improves the required average number of trials to succeed by 60.56% for cart-pole, and 56.89% for triple-link balancing tasks, in comparison with the traditional HDP approach. Also, we have added results of ER-based HDP for comparison. Moreover, theoretical convergence analysis is presented to guarantee the stability of the proposed control design.
Zhen Ni, Naresh Malla, Xiangnan Zhong
IEEE Trans. Cybern.1
2019 Editorial: Booming of Neural Networks and Learning Systems
abstract
As you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community.
Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He
IEEE Trans. Neural Networks Learn. Syst.23
2019 A Multistage Game in Smart Grid Security: A Reinforcement Learning Solution
abstract
Existing smart grid security research investigates different attack techniques and cascading failures from the attackers' viewpoints, while the defenders' or the operators' protection strategies are somehow neglected. Game theoretic methods are applied for the attacker-defender games in the smart grid security area. Yet, most of the existing works only use the one-shot game and do not consider the dynamic process of the electric power grid. In this paper, we propose a new solution for a multistage game (also called a dynamic game) between the attacker and the defender based on reinforcement learning to identify the optimal attack sequences given certain objectives (e.g., transmission line outages or generation loss). Different from a one-shot game, the attacker here learns a sequence of attack actions applying for the transmission lines and the defender protects a set of selected lines. After each time step, the cascading failure will be measured, and the line outage (and/or generation loss) will be used as the feedback for the attacker to generate the next action. The performance is evaluated on W&W 6-bus and IEEE 39-bus systems. A comparison between a multistage attack and a one-shot attack is conducted to show the significance of the multistage attack. Furthermore, different protection strategies are evaluated in simulation, which shows that the proposed reinforcement learning solution can identify optimal attack sequences under several attack objectives. It also indicates that attacker's learned information helps the defender to enhance the security of the system.
Zhen Ni, Shuva Paul
IEEE Trans. Neural Networks Learn. Syst.1
2019 Learning Human-Robot Interaction for Robot-Assisted Pedestrian Flow Optimization
abstract
Due to the fast-is-slower phenomenon in emergency escape, it is desirable to regulate pedestrian flows at the exit or a bottleneck. We propose a new robot-assisted pedestrian regulation and study passive human-robot interaction (HRI). A learning-based motion control approach is presented for a robot to efficiently interact with pedestrians for desirable collective motion. We first formulate the problem into an optimal control framework using the pedestrian dynamics description based on existing social force models with embedded HRI forces. To solve the defined optimal control problem, we propose an adaptive dynamic programming (ADP) approach to provide adjustable motion parameters of the robot to efficiently interact with pedestrians so that the regulated pedestrian flow tracks a desired velocity. The ADP control process only uses observed flow information rather than the models of pedestrians, and the ADP method provides feedback control with online learning and control capability. Simulation results demonstrate that the proposed approach can regulate pedestrian flows to desirable speeds by online learning.
Chao Jiang 0001, Zhen Ni, Yi Guo 0004, Haibo He
IEEE Trans. Syst. Man Cybern. Syst.2
2018 A Study of Linear Programming and Reinforcement Learning for One-Shot Game in Smart Grid Security
abstract
Smart grid attacks can be applied on a single component or multiple components. The corresponding defense strategies are totally different. In this paper, we investigate the solutions (e.g., linear programming and reinforcement learning) for one-shot game between the attacker and defender in smart power systems. We designed one-shot game with multi-line- switching attack and solved it using linear programming. We also designed the game with single-line-switching attack and solved it using reinforcement learning. The pay-off and utility/reward of the game is calculated based on the generation loss due to initiated attack by the attacker. Defender's defense action is considered while evaluating the pay-off from attacker's and defender's action. The linear programming based solution gives the probability of choosing best attack actions against different defense actions. The reinforcement learning based solution gives the optimal action to take under selected defense action. The proposed game is demonstrated on 6 bus system and IEEE 30 bus system and optimal solutions are analyzed.
Shuva Paul, Zhen Ni
IJCNN2
2018 Data-Driven Reinforcement Learning Design for Multi-agent Systems with Unknown Disturbances
abstract
In this paper, we develop a new data-driven reinforcement learning method to solve the multi-agent consensus control problem with unknown disturbances. Due to the existence of disturbances, the transmitted information between each pair of the agents becomes unreliable, which makes that the data- driven reinforcement learning based control becomes difficult. Therefore, we solve the problem by developing an appropriate performance index for each agent to convert the robust consensus problem to an auxiliary optimal control problem. The equivalence of the transformation is proved to show that the solution of the auxiliary optimal control system can asymptotically stabilize the original robust system and synchronize all the agents at the same time. Neural network techniques are applied to implement the proposed method. Finally, the simulation results demonstrate the theoretical analysis and verify the effectiveness of the proposed method.
Xiangnan Zhong, Zhen Ni
IJCNN2
2018 Measuring Performance and Isolation Tradeoffs for NFV
abstract
Network function virtualization (NFV) allows network services, such as firewalls and routing, to be deployed into a virtual environment and run on commodity hardware. Recently service providers and developers can deploy their network functions (NF) prototypes on a shared infrastructure, and all the NFs are being controlled by the manager of the platform. NFV platforms run these NFs together, and share the system resources to optimize the utilization. This means that limited resources such as CPU cores or memory have to be shared. Several recent NFV systems run network services with one shared memory region, so that they can achieve high performance with zero-copy I/O. This resource sharing brings security problems since it allows malicious NFs to easily modify data from other NFs. To enhance the security of NFV, we are designing a platform to provide stronger memory isolation between different NFs. Our approach is based on the architecture developed for our OpenNetVM platform, which supports lightweight NFs, flexible management, but assumes a single shared memory pool for all NFs.
Zhen Ni, Timothy Wood 0001
LANMAN1
2018 CRIMES: Using Evidence to Secure the Cloud
abstract
Cloud applications are appealing targets to attackers, yet current cloud infrastructures have few ways of helping defend their customers from attacks. However, the use of virtual machines, and the economy of scale found in cloud platforms, provides an opportunity to offer strong security guarantees to tenants at low cost to the cloud provider. We present CRIMES, an evidence based, modular security framework for cloud platforms that uses speculative execution coupled with memory introspection tools to detect malicious behavior in real time. By buffering VM outputs (i.e., outgoing network packets and disk writes) until a scan has been completed, CRIMES gives strong guarantees about the amount of damage an attack can do, while minimizing overheads. When an attack is detected, CRIMES rolls back to a recent checkpoint and performs automated forensic analysis to help pinpoint the source of an attack. Our evaluation demonstrates that CRIMES incurs less overhead compared to memory protection tools such as AddressSanitizer, while offering valuable forensic analysis for buffer overflow attacks and malware detection across multiple applications and the OS.
Sundaresan Rajasekaran, Harpreet Singh Chawla, Zhen Ni, Emery D. Berger, Timothy Wood 0001
Middleware3
2018 Model-Free Adaptive Control for Unknown Nonlinear Zero-Sum Differential Game
abstract
In this paper, we present a new model-free globalized dual heuristic dynamic programming (GDHP) approach for the discrete-time nonlinear zero-sum game problems. First, the online learning algorithm is proposed based on the GDHP method to solve the Hamilton-Jacobi-Isaacs equation associated with optimal regulation control problem. By setting backward one step of the definition of performance index, the requirement of system dynamics, or an identifier is relaxed in the proposed method. Then, three neural networks are established to approximate the optimal saddle point feedback control law, the disturbance law, and the performance index, respectively. The explicit updating rules for these three neural networks are provided based on the data generated during the online learning along the system trajectories. The stability analysis in terms of the neural network approximation errors is discussed based on the Lyapunov approach. Finally, two simulation examples are provided to show the effectiveness of the proposed method.
Xiangnan Zhong, Haibo He, Ding Wang 0001, Zhen Ni
IEEE Trans. Cybern.4
2017 Towards enabling deep learning techniques for adaptive dynamic programming
abstract
Human-level control through deep learning and deep reinforcement learning have revealed the unique and powerful potentials through a very complex Go game. The AlphaGo, developed by Google DeepMind, has beat the top Go game player early this year. The scientific and technological advancement behind the success of AlphaGo attracted researchers from multiple areas, including machine learning, artificial intelligence, computational intelligence and so on. Adaptive dynamic programming (ADP) methods have the similar fundamental principle with reinforcement learning, and show strong performance for continuous time and continuous state systems. Deep learning techniques are also possible to be integrated for ADP designs. In this paper, we discuss the key techniques and components in deep reinforcement learning and then present the successful applications for computer games and maze navigation. Future opportunities for deep learning enabled ADP will be discussed at the end.
Zhen Ni, Naresh Malla, Xiangnan Zhong
IJCNN1
2017 VRer: Context-Based Venue Recommendation using embedded space ranking SVM in location-based social network
Bin Xia 0003, Zhen Ni, Tao Li 0001, Qianmu Li, Qifeng Zhou
Expert Syst. Appl.2
2017 A new history experience replay design for model-free adaptive dynamic programming
Naresh Malla, Zhen Ni
Neurocomputing2
2017 Data-Driven Tracking Control With Adaptive Dynamic Programming for a Class of Continuous-Time Nonlinear Systems
abstract
A data-driven adaptive tracking control approach is proposed for a class of continuous-time nonlinear systems using a recent developed goal representation heuristic dynamic programming (GrHDP) architecture. The major focus of this paper is on designing a multivariable tracking scheme, including the filter-based action network (FAN) architecture, and the stability analysis in continuous-time fashion. In this design, the FAN is used to observe the system function, and then generates the corresponding control action together with the reference signals. The goal network will provide an internal reward signal adaptively based on the current system states and the control action. This internal reward signal is assigned as the input for the critic network, which approximates the cost function over time. We demonstrate its improved tracking performance in comparison with the existing heuristic dynamic programming (HDP) approach under the same parameter and environment settings. The simulation results of the multivariable tracking control on two examples have been presented to show that the proposed scheme can achieve better control in terms of learning speed and overall performance.
Chaoxu Mu, Zhen Ni, Changyin Sun 0001, Haibo He
IEEE Trans. Cybern.2
2017 Gr-GDHP: A New Architecture for Globalized Dual Heuristic Dynamic Programming
abstract
Goal representation globalized dual heuristic dynamic programming (Gr-GDHP) method is proposed in this paper. A goal neural network is integrated into the traditional GDHP method providing an internal reinforcement signal and its derivatives to help the control and learning process. From the proposed architecture, it is shown that the obtained internal reinforcement signal and its derivatives can be able to adjust themselves online over time rather than a fixed or predefined function in literature. Furthermore, the obtained derivatives can directly contribute to the objective function of the critic network, whose learning process is thus simplified. Numerical simulation studies are applied to show the performance of the proposed Gr-GDHP method and compare the results with other existing adaptive dynamic programming designs. We also investigate this method on a ball-and-beam balancing system. The statistical simulation results are presented for both the Gr-GDHP and the GDHP methods to demonstrate the improved learning and controlling performance.
Xiangnan Zhong, Zhen Ni, Haibo He
IEEE Trans. Cybern.2
2017 Air-Breathing Hypersonic Vehicle Tracking Control Based on Adaptive Dynamic Programming
abstract
In this paper, we propose a data-driven supplementary control approach with adaptive learning capability for air-breathing hypersonic vehicle tracking control based on action-dependent heuristic dynamic programming (ADHDP). The control action is generated by the combination of sliding mode control (SMC) and the ADHDP controller to track the desired velocity and the desired altitude. In particular, the ADHDP controller observes the differences between the actual velocity/altitude and the desired velocity/altitude, and then provides a supplementary control action accordingly. The ADHDP controller does not rely on the accurate mathematical model function and is data driven. Meanwhile, it is capable to adjust its parameters online over time under various working conditions, which is very suitable for hypersonic vehicle system with parameter uncertainties and disturbances. We verify the adaptive supplementary control approach versus the traditional SMC in the cruising flight, and provide three simulation studies to illustrate the improved performance with the proposed approach.
Chaoxu Mu, Zhen Ni, Changyin Sun 0001, Haibo He
IEEE Trans. Neural Networks Learn. Syst.2
2016 Supplementary control for virtual synchronous machine based on adaptive dynamic programming
abstract
Virtual synchronous machine (VSM) can be used in micro-grids (MGs) for supporting virtual inertia to mitigate frequency fluctuations in the system. The performance of VSM depends on the performance of the current controlled voltage source inverter. Generally used proportional-integral (PI) controller has limited transient performance in current control. Adaptive dynamic programming (ADP) controller has been investigated, designed and tested to improve transient stability in the electric power system. In this paper, ADP controller has been proposed as a supplementary controller for fine tuning conventional PI current controller thereby improving the performance of VSM. The grid connected inverter system was simulated in MATLAB/Simulink to analyze transient stability problems. ADP as a supplementary controller not only reduced the overshoot of d-axis current at sudden load changes but also increased the speed of the current control. Furthermore, the system was tested for single phase line to ground fault and performed well based on the proposed ADP supplementary control approach.
Naresh Malla, Dipesh Shrestha, Zhen Ni, Reinaldo Tonkoski
CEC3
2016 Comparative studies of power grid security with network connectivity and power flow information using unsupervised learning
abstract
The modern electric power grid has become highly integrated in order to increase reliability of power transmission from the generating units to end consumers. This integrated nature and its upgrade toward an intelligent smart grid make the power grid vulnerable when facing cyber or physical attacks as well as intentional attacks. Therefore, determining the most vulnerable components (e.g., buses or generators) is critically important for power grid defense. In this paper, a new definition of load is proposed by taking power flow into consideration in comparison with the load definition based on degree or network connectivity. Unsupervised learning techniques (e.g., K-means algorithm and self-organizing map (SOM)) are introduced to cluster the nodes (i.e., buses) in IEEE-39 bus and IEEE-57 bus benchmarks. Then most vulnerable node in each cluster is determined based on their load information to form initial victim set. We use percentage of failure (PoF) to compare the performance of clustering based approach and traditional load based approach during cascading failure process. With the simulation results, the unsupervised learning (clustering based) approaches are more efficient in finding the most vulnerable nodes and our proposed definition of load is relatively useful in studying power grid security.
Shiva Poudel, Zhen Ni, Xiangnan Zhong, Haibo He
IJCNN2
2016 Convergence analysis of GrDHP-based optimal control for discrete-time nonlinear system
abstract
Adaptive dynamic programming (ADP) has been investigated for its new architectures, algorithms and applications for years. Recently, the goal representation (Gr) design has been demonstrated with promising results to improve ADP control performance from certain perspectives. This paper is focused on the theoretical analysis of the goal representation dual heuristic dynamic programming (GrDHP). Starting from the general formulation of the GrDHP design, we provide the iterative algorithm for this method. The corresponding convergence analysis is showed in terms of the internal reinforcement signal, the performance index, and their derivatives. Our analysis assumes that the system is controllable and stabilizable. Then, neural-network-based implementation of this method is presented. Simulation study validates the theoretical analysis of this paper and also shows the effectiveness of the GrDHP method.
Xiangnan Zhong, Zhen Ni, Haibo He
IJCNN2
2016 Robot-assisted pedestrian regulation in an exit corridor
abstract
Due to the faster-is-slower phenomenon in emergency escape, it is desirable to regulate pedestrian flow at the exit or a bottleneck. Modification of pedestrian facilities was previously studied to increase the efficiency and safety by the transportation community. We propose a robot-assisted pedestrian regulation scheme and study passive human-robot interaction (HRI), where the robot acts as a dynamic obstacle that interacts with pedestrians. Such a robot-assisted solution replaces expensive infrastructure modification with real-time reconfigurability. In the paper, we first formulate a robot-assisted flow optimization problem based on the social force models of pedestrian dynamics with embedded HRI forces. We then present an online learning algorithm based on adaptive dynamic programming (ADP) to generate motion control so that the robot can replan and adapt its motion to real-time pedestrian flows. The ADP control process uses observed flow information only but not the models of pedestrians, and provides feedback control with online learning and control capability. Simulation results demonstrate efficiency of the proposed method.
Chao Jiang 0001, Zhen Ni, Yi Guo 0004, Haibo He
IROS2
2016 Fuzzy-Based Goal Representation Adaptive Dynamic Programming
abstract
In this paper, a novel nonlinear learning controller called fuzzy-based goal representation adaptive dynamic programming (Fuzzy-GrADP) is proposed. In the proposed GrADP method, a goal representation network is introduced to generate an adaptive internal reinforcement signal to the critic network to help the controller provide a general mapping between the input and output actions. Moreover, in the proposed architecture, the action network in the GrADP is improved by using the fuzzy hyperbolic model, which combines the merits of the fuzzy model and the neural network model. Based on the back-propagation technique, the parameters in the membership functions and the fuzzy rules are all undergo training and online adapting. The proposed controller is tested on two numerical benchmarks, and the simulation results show that the proposed controller outperforms the original adaptive dynamic fuzzy controller and the pure neural network-based GrADP controller. In addition, the proposed controller is further applied on a large multimachine power system for static var compensator damping control, where simulation results demonstrate the effectiveness of the proposed approach on real applications. Furthermore, in order to demonstrate the theoretical guarantee of the proposed method, Lyapunov stability analysis to support the proposed Fuzzy-GrADP approach has also been carried out.
Yufei Tang, Haibo He, Zhen Ni, Xiangnan Zhong, Dongbin Zhao, Xin Xu 0001
IEEE Trans. Fuzzy Syst.3
2016 Adaptive Modulation for DFIG and STATCOM With High-Voltage Direct Current Transmission
abstract
This paper develops an adaptive modulation approach for power system control based on the approximate/adaptive dynamic programming method, namely, the goal representation heuristic dynamic programming (GrHDP). In particular, we focus on the fault recovery problem of a doubly fed induction generator (DFIG)-based wind farm and a static synchronous compensator (STATCOM) with high-voltage direct current (HVDC) transmission. In this design, the online GrHDP-based controller provides three adaptive supplementary control signals to the DFIG controller, STATCOM controller, and HVDC rectifier controller, respectively. The mechanism is to observe the system states and their derivatives and then provides supplementary control to the plant according to the utility function. With the GrHDP design, the controller can adaptively develop an internal goal representation signal according to the observed power system states, therefore, to achieve more effective learning and modulating. Our control approach is validated on a wind power integrated benchmark system with two areas connected by HVDC transmission lines. Compared with the classical direct HDP and proportional integral control, our GrHDP approach demonstrates the improved transient stability under system faults. Moreover, experiments under different system operating conditions with signal transmission delays are also carried out to further verify the effectiveness and robustness of the proposed approach.
Yufei Tang, Haibo He, Zhen Ni, Jinyu Wen, Tingwen Huang
IEEE Trans. Neural Networks Learn. Syst.3
2016 A Theoretical Foundation of Goal Representation Heuristic Dynamic Programming
abstract
Goal representation heuristic dynamic programming (GrHDP) control design has been developed in recent years. The control performance of this design has been demonstrated in several case studies, and also showed applicable to industrial-scale complex control problems. In this paper, we develop the theoretical analysis for the GrHDP design under certain conditions. It has been shown that the internal reinforcement signal is a bounded signal and the performance index can converge to its optimal value monotonically. The existence of the admissible control is also proved. Although the GrHDP control method has been investigated in many areas before, to the best of our knowledge, this is the first study of presenting the theoretical foundation of the internal reinforcement signal and how such an internal reinforcement signal can provide effective information to improve the control performance. Numerous simulation studies are used to validate the theoretical analysis and also demonstrate the effectiveness of the GrHDP design.
Xiangnan Zhong, Zhen Ni, Haibo He
IEEE Trans. Neural Networks Learn. Syst.2
2015 A comparative study between motivated learning and reinforcement learning
abstract
This paper analyzes advanced reinforcement learning techniques and compares some of them to motivated learning. Motivated learning is briefly discussed indicating its relation to reinforcement learning. A black box scenario for comparative analysis of learning efficiency in autonomous agents is developed and described. This is used to analyze selected algorithms. Reported results demonstrate that in the selected category of problems, motivated learning outperformed all reinforcement learning algorithms we compared with.
James T. Graham, Janusz A. Starzyk, Zhen Ni, Haibo He, Teck-Hou Teng, Ah-Hwee Tan
IJCNN3
2015 A boundedness theoretical analysis for GrADP design: A case study on maze navigation
abstract
A new theoretical analysis towards the goal representation adaptive dynamic programming (GrADP) design proposed in [1], [2] is investigated in this paper. Unlike the proofs of convergence for adaptive dynamic programming (ADP) in literature, here we provide a new insight for the error bound between the estimated value function and the expected value function. Then we employ the critic network in GrADP approach to approximate the Q value function, and use the action network to provide the control policy. The goal network is adopted to provide the internal reinforcement signal for the critic network over time. Finally, we illustrate that the estimated Q value function is close to the expected value function in an arbitrary small bound on the maze navigation example.
Zhen Ni, Xiangnan Zhong, Haibo He
IJCNN1
2015 Event-triggered adaptive dynamic programming for continuous-time nonlinear system using measured input-output data
abstract
In this paper, we propose a novel event-triggered adaptive dynamic programming (ADP) method using only the input-output data. Event-triggered method is widely used for its computational efficiency capacity. Comparing with the traditional method which updates the controller periodically, the event-triggered method only updates the controller when it is necessary and therefore the computation is reduced. Generally, the triggered condition is based on the system current and sampled states. In this paper, we consider a neural-network-based observer to recover the system dynamics using the measured input-output data. The triggered instants are calculated according to the recovered state. Stability analysis of the proposed approach is presented. We verify our proposed method through a robot-arm example.
Xiangnan Zhong, Zhen Ni, Haibo He
IJCNN2
2015 Data-driven heuristic dynamic programming with virtual reality
Dezhong Zheng, Haibo He, Zhen Ni
Neurocomputing4
2015 Model-Free Dual Heuristic Dynamic Programming
abstract
Model-based dual heuristic dynamic programming (MB-DHP) is a popular approach in approximating optimal solutions in control problems. Yet, it usually requires offline training for the model network, and thus resulting in extra computational cost. In this brief, we propose a model-free DHP (MF-DHP) design based on finite-difference technique. In particular, we adopt multilayer perceptron with one hidden layer for both the action and the critic networks design, and use delayed objective functions to train both the action and the critic networks online over time. We test both the MF-DHP and MB-DHP approaches with a discrete time example and a continuous time example under the same parameter settings. Our simulation results demonstrate that the MF-DHP approach can obtain a control performance competitive with that of the traditional MB-DHP approach while requiring less computational resources.
Zhen Ni, Haibo He, Xiangnan Zhong, Danil V. Prokhorov
IEEE Trans. Neural Networks Learn. Syst.1
2015 GrDHP: A General Utility Function Representation for Dual Heuristic Dynamic Programming
abstract
A general utility function representation is proposed to provide the required derivable and adjustable utility function for the dual heuristic dynamic programming (DHP) design. Goal representation DHP (GrDHP) is presented with a goal network being on top of the traditional DHP design. This goal network provides a general mapping between the system states and the derivatives of the utility function. With this proposed architecture, we can obtain the required derivatives of the utility function directly from the goal network. In addition, instead of a fixed predefined utility function in literature, we conduct an online learning process for the goal network so that the derivatives of the utility function can be adaptively tuned over time. We provide the control performance of both the proposed GrDHP and the traditional DHP approaches under the same environment and parameter settings. The statistical simulation results and the snapshot of the system variables are presented to demonstrate the improved learning and controlling performance. We also apply both approaches to a power system example to further demonstrate the control capabilities of the GrDHP approach.
Zhen Ni, Haibo He, Dongbin Zhao, Xin Xu 0001, Danil V. Prokhorov
IEEE Trans. Neural Networks Learn. Syst.1
2014 Data-driven partially observable dynamic processes using adaptive dynamic programming
abstract
Adaptive dynamic programming (ADP) has been widely recognized as one of the “core methodologies” to achieve optimal control for intelligent systems in Markov decision process (MDP). Generally, ADP control design requires all the information of the system dynamics. However, in many practical situations, the measured input and output data can only represent part of the system states. This means the complete information of the system cannot be available in many real-world cases, which narrows the range of application of the ADP design. In this paper, we propose a data-driven ADP method to stabilize the system with partially observable dynamics based on neural network techniques. A state network is integrated into the typical actor-critic architecture to provide an estimated state from the measured input/output sequences. The theoretical analysis and the stability discussion of this data-driven ADP method are also provided. Two examples are studied to verify our proposed method.
Xiangnan Zhong, Zhen Ni, Yufei Tang, Haibo He
ADPRL2
2014 Experimental studies on indoor sign recognition and classification
abstract
Previous works on outdoor traffic sign recognition and classification have been demonstrated useful to the driver assistant system and the possibility to the autonomous vehicles. This motivates our research on the assistance for visual impairment or visual disabled pedestrians in the indoor environment. In this paper, we build an indoor sign database and investigate the recognition and classification for the indoor sign problem. We adopt the classical techniques on extracting the features, including the principle component analysis (PCA), dense scale invariant feature transform (DSIFT), histogram of oriented gradients (HOG), and conduct the state-of-art classification techniques, such as the neural network (NN), support vector machine (SVM) and k-nearest neighbors (KNN). We provide the experimental results on this newly built database and also discuss the insight for the possibility of indoor navigation for the blind or visual-disabled people.
Zhen Ni, Si-Yao Fu, Bo Tang 0011, Haibo He, Xinming Huang 0001
CIDM1
2014 Event-triggered reinforcement learning approach for unknown nonlinear continuous-time system
abstract
This paper provides an adaptive event-triggered method using adaptive dynamic programming (ADP) for the nonlinear continuous-time system. Comparing to the traditional method with fixed sampling period, the event-triggered method samples the state only when an event is triggered and therefore the computational cost is reduced. We demonstrate the theoretical analysis on the stability of the event-triggered method, and integrate it with the ADP approach. The system dynamics are assumed unknown. The corresponding ADP algorithm is given and the neural network techniques are applied to implement this method. The simulation results verify the theoretical analysis and justify the efficiency of the proposed event-triggered technique using the ADP approach.
Xiangnan Zhong, Zhen Ni, Haibo He, Xin Xu 0001, Dongbin Zhao
IJCNN2
2014 Reactive power control of grid-connected wind farm based on adaptive dynamic programming
Yufei Tang, Haibo He, Zhen Ni, Jinyu Wen, Xianchao Sui
Neurocomputing3
2013 Real-time tracking on adaptive critic design with uniformly ultimately bounded condition
abstract
In this paper, we proposed a new nonlinear tracking controller based on heuristic dynamic programming (HDP) with the tracking filter. Specifically, we integrate a goal network into the regular HDP design and provide the critic network with detailed internal reward signal to help the value function approximation. The architecture is explicitly explained with the tracking filter, goal network, critic network and action network, respectively. We provide the stability analysis of our proposed controller with Lyapunov approach. It is shown that the filtered tracking errors and the weights estimation errors in neural networks are all uniformly ultimately bounded (UUB) under certain conditions. Finally, we compare our proposed approach with regular HDP approach in virtual reality (VR)/Simulink environment to justify the improved control performance.
Zhen Ni, Haibo He, Dongbin Zhao, Xin Xu 0001
ADPRL1
2013 Heuristic dynamic programming with internal goal representation
Zhen Ni, Haibo He
Soft Comput.1
2013 Adaptive Learning in Tracking Control Based on the Dual Critic Network Design
abstract
In this paper, we present a new adaptive dynamic programming approach by integrating a reference network that provides an internal goal representation to help the systems learning and optimization. Specifically, we build the reference network on top of the critic network to form a dual critic network design that contains the detailed internal goal representation to help approximate the value function. This internal goal signal, working as the reinforcement signal for the critic network in our design, is adaptively generated by the reference network and can also be adjusted automatically. In this way, we provide an alternative choice rather than crafting the reinforcement signal manually from prior knowledge. In this paper, we adopt the online action-dependent heuristic dynamic programming (ADHDP) design and provide the detailed design of the dual critic network structure. Detailed Lyapunov stability analysis for our proposed approach is presented to support the proposed structure from a theoretical point of view. Furthermore, we also develop a virtual reality platform to demonstrate the real-time simulation of our approach under different disturbance situations. The overall adaptive learning performance has been tested on two tracking control benchmarks with a tracking filter. For comparative studies, we also present the tracking performance with the typical ADHDP, and the simulation results justify the improved performance with our approach.
Zhen Ni, Haibo He, Jinyu Wen
IEEE Trans. Neural Networks Learn. Syst.1
2013 Goal Representation Heuristic Dynamic Programming on Maze Navigation
abstract
Goal representation heuristic dynamic programming (GrHDP) is proposed in this paper to demonstrate online learning in the Markov decision process. In addition to the (external) reinforcement signal in literature, we develop an adaptively internal goal/reward representation for the agent with the proposed goal network. Specifically, we keep the actor-critic design in heuristic dynamic programming (HDP) and include a goal network to represent the internal goal signal, to further help the value function approximation. We evaluate our proposed GrHDP algorithm on two 2-D maze navigation problems, and later on one 3-D maze navigation problem. Compared to the traditional HDP approach, the learning performance of the agent is improved with our proposed GrHDP approach. In addition, we also include the learning performance with two other reinforcement learning algorithms, namely Sarsa(λ) and Q-learning, on the same benchmarks for comparison. Furthermore, in order to demonstrate the theoretical guarantee of our proposed method, we provide the characteristics analysis toward the convergence of weights in neural networks in our GrHDP approach.
Zhen Ni, Haibo He, Jinyu Wen, Xin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2012 Reinforcement learning control based on multi-goal representation using hierarchical heuristic dynamic programming
abstract
We are interested in developing a multi-goal generator to provide detailed goal representations that help to improve the performance of the adaptive critic design (ACD). In this paper we propose a hierarchical structure of goal generator networks to cascade external reinforcement into more informative internal goal representations in the ACD. This is in contrast with previous designs in which the external reward signal is assigned to the critic network directly. The ACD control system performance is evaluated on the ball-and-beam balancing benchmark under noise-free and various noisy conditions. Simulation results in the form of a comparative study demonstrate effectiveness of our approach.
Zhen Ni, Haibo He, Dongbin Zhao, Danil V. Prokhorov
IJCNN1
2012 A three-network architecture for on-line learning and optimization based on adaptive dynamic programming
Haibo He, Zhen Ni
Neurocomputing2
2011 Adaptive dynamic programming with balanced weights seeking strategy
abstract
In this paper we propose to integrate the recursive Levenberg-Marquardt method into the adaptive dynamic programming (ADP) design for improved learning and adaptive control performance. Our key motivation is to consider a balanced weight updating strategy with the consideration of both robustness and convergence during the online learning process. Specifically, a modified recursive Levenberg-Marquardt (LM) method is integrated into both the action network and critic network of the ADP design, and a detailed learning algorithm is proposed to implement this approach. We test the performance of our approach based on the triple link inverted pendulum, a popular benchmark in the community, to demonstrate online learning and control strategy. Experimental results and comparative study under different noise conditions demonstrate the effectiveness of this approach.
Haibo He, Zhen Ni
ADPRL3
2011 An online actor-critic learning approach with Levenberg-Marquardt algorithm
abstract
This paper focuses on the efficiency improvement of online actor-critic design base on the Levenberg-Marquardt (LM) algorithm rather than traditional chain rule. Over the decades, several generations of adaptive/approximate dynamic programming (ADP) structures have been proposed in the community and demonstrated many successfully applications. Neural network with backpropagation has been one of the most important approaches to tune the parameters in such ADP designs. In this paper, we aim to study the integration of Levenberg-Marquardt method into the regular actor-critic design to improve weights updating and learning for a quadratic convergence under certain condition. Specifically, for the critic network design, we adopt the LM method targeting improved learning performance, while for the action network, we use the neural network with backpropagation to provide an appropriate control action. A detailed learning algorithm is presented, followed by benchmark tests of pendulum swing up and balance and cart-pole balance tasks. Various simulation results and comparative study demonstrated the effectiveness of this approach.
Zhen Ni, Haibo He, Danil V. Prokhorov
IJCNN1
2011 An Adaptive Dynamic Programming Approach for Closely-Coupled MIMO System Control
Haibo He, Zhen Ni
ISNN (3)4