EDBT 2026 Demo / reviewers in the wild / expert
Derong Liu 0001
dblp:l/DerongLiu
· DBLP profile ↗
314ranked-venue papers
47as first author
74since 2021 · last 2026
0000-0003-3715-4778ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 231 · 30 first-author · 55 since 2021Human-computer interaction and ubiquitous computing · 38 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-authorSystems, architecture and hardware · 11 · 2 first-author · 4 since 2021Computer networks · 5 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compensating Distribution Drifts in Continual Learning with Pre-trained Vision TransformersabstractRecent advances have shown that sequential fine-tuning (SeqFT) of pre-trained vision transformers (ViTs), followed by classifier refinement using approximate distributions of class features, can be an effective strategy for class-incremental learning (CIL). However, this approach is susceptible to distribution drift, caused by the sequential optimization of shared backbone parameters. This results in a mismatch between the distributions of the previously learned classes and that of the updated model, ultimately degrading the effectiveness of classifier performance over time. To address this issue, we introduce a latent space transition operator and propose Sequential Learning with Drift Compensation (SLDC). SLDC aims to align feature distributions across tasks to mitigate the impact of drift. First, we present a linear variant of SLDC, which learns a linear operator by solving a regularized least-squares problem that maps features before and after fine-tuning. Next, we extend this with a weakly nonlinear SLDC variant, which assumes that the ideal transition operator lies between purely linear and fully nonlinear transformations. This is implemented using learnable, weakly nonlinear mappings that balance flexibility and generalization. To further reduce representation drift, we apply knowledge distillation (KD) in both algorithmic variants. Extensive experiments on standard CIL benchmarks demonstrate that SLDC significantly improves the performance of SeqFT. Notably, by combining KD to address representation drift with SLDC to compensate distribution drift, SeqFT achieves performance comparable to joint training across all evaluated datasets. Xuan Rao, Simian Xu, Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Cesare Alippi |
AAAI | 5 |
| 2026 | Computation-aware Transformer-based encoding for efficient latent spatial neural architecture search
Jiamin Xiao, Bo Zhao 0015, Derong Liu 0001, Yonghua Wang 0001, Jiacai Huang |
Neurocomputing | 3 |
| 2026 | Control barrier function-based self-learning robust control of safety-critical nonlinear systems
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001 |
Neurocomputing | 3 |
| 2026 | Federated Learning Adaptive Dynamic Programming for Massive Multiagent Mean-Field Games-Based Optimal ConsensusabstractMassive multiagent systems typically involve a very large number of interactions and conflicts of interest among agents, which presents a significant challenge for achieving stable and efficient adaptive optimal consensus control in real-time. To fill this gap, this article develops a novel federated learning adaptive dynamic programming (FL-ADP) control scheme to solve the massive multiagent mean-field games (MFGs)-based optimal consensus problem. First, the complex interactions of each individual agent with all other agents can be approximated by an average or collective influence in the context of MFGs. Then, a novel undiscounted performance index function involving the mean-field coupling term, the tracking errors and their derivatives is proposed to circumvent the potential impact of the improper discount factor selection and achieve better control performance. By designing the critic-mass neural network structure, the coupled Hamilton-Jacobi-Bellman and Fokker-Planck-Kolmogorov equations are solved to derive the approximate optimal control policy and quantify the probability density function of the collective behavior simultaneously. Additionally, to comply with the required convergence condition of the MFG, a novel event-triggered federated learning mechanism is formulated, which achieves a balance between communication resource consumption and the guarantee of algorithm convergence. On the basis of Lyapunov's direct method, the tracking errors and the weight estimation errors of all agents are guaranteed to be uniformly ultimately bounded. Simulation results of massive multi-uncrewed aerial vehicle systems affirm the rationality and effectiveness of the proposed method. Mingduo Lin, Guoling Yuan, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2026 | Dynamic Event-Triggered Control for Human-Machine Cooperative Systems Based on Dynamic Authority AllocationabstractThis article addresses the challenging problem of constrained optimal control for human–machine systems subject to external disturbances and the bounded rationality of the human operator. To this end, a novel game-theoretic framework is proposed. Unlike monolithic game formulations, the framework uniquely disaggregates the control problem by transforming it into a multifaceted game via logarithmic barrier functions (BFs): it models human–machine cooperation as a positive-sum game oriented toward shared objectives, and disturbance rejection as a zero-sum game tailored for robustness enhancement. To capture the nonideal human decision-making, we integrate the level-$k$reasoning framework to model the operator’s bounded cognitive dynamics. The corresponding coupled Hamilton–Jacobi–Isaacs (HJI) equations for this human–machine game are derived, and critically, a rigorous proof of global asymptotic stability (GAS) for the transformed system is provided, establishing a solid theoretical foundation. For online implementation without requiring prior knowledge of the system dynamics, we develop a resource-efficient learning architecture based on the adaptive dynamic programming (ADP) and a novel dynamic event-triggered mechanism (DETM). A key feature of this architecture is a fuzzy logic-based module for dynamic authority allocation, which adaptively adjusts control sharing in real time. Rigorous analysis demonstrates that all signals in the closed-loop system are uniformly ultimately bounded and that Zeno behavior is precluded. Simulation results are presented to validate the effectiveness and superiority of the proposed control strategy. Dehua Zhang, Linlin Liang, Chunbin Qin, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2025 | Distributed Fault-Tolerant Consensus Control Based on Zero-Sum Differential Games for Nonlinear Multi-agent Systems
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001 |
ISNN | 4 |
| 2025 | Dynamic event-triggering adaptive dynamic programming for robust stabilization of partially unknown nonlinear systems
Yishen Hong, Shan Xue 0004, Derong Liu 0001, Yonghua Wang 0001 |
Neurocomputing | 3 |
| 2025 | Improved cost function-based fault tolerant control for nonlinear systems with simultaneous faults
Chujian Zeng, Bo Zhao 0015, Derong Liu 0001 |
Neurocomputing | 3 |
| 2025 | Event-triggered robust hierarchical control for uncertain multiplayer Stackelberg games via adaptive dynamic programming
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Marios M. Polycarpou, Shiguo Peng, Shunchao Zhang |
Neurocomputing | 3 |
| 2025 | On robust learning of memory attractors with noisy deep associative memory networks
Xuan Rao, Bo Zhao 0015, Derong Liu 0001 |
Neural Networks | 3 |
| 2025 | Event-triggered control for input-constrained nonzero-sum games through particle swarm optimized neural networks
Qiuye Wu, Bo Zhao 0015, Derong Liu 0001 |
Neural Networks | 3 |
| 2025 | Dynamic Self-Triggered Intelligent Path Tracking Control for Autonomous Agricultural Vehicles via Reinforcement LearningabstractThis paper investigates the path tracking control of unmanned agricultural vehicles with disturbances under a dynamic self-triggered mechanism via reinforcement learning (RL). To begin with, a path tracking offset system is constructed based on the kinematic model of the unmanned agricultural vehicle, which transforms the path tracking control problem into an optimal control problem. Subsequently, a novel dynamic self-triggered second-order integral sliding mode control policy is developed to mitigate the impact of disturbances and to derive a nominal path tracking offset model. Afterward, to further alleviate the computing and communication burdens, a novel dynamic self-triggered mechanism is proposed for the optimal control policy. It can predict the next update time based on current information, thus avoiding the need for continuous monitoring of the triggering condition. Furthermore, a single critic network architecture is constructed to obtain an approximate path tracking control policy, and it is proven by Lyapunov stability theory that this policy ensures unmanned agricultural vehicles can maintain the predefined working path in the presence of disturbances. Finally, the effectiveness of the proposed path tracking control method is demonstrated by simulation experiments. Yongwei Zhang 0002, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Event-Triggered Impulsive Controller Design of Continuous Nonlinear Systems Using Liquid-Updating ADPabstractThis paper designs and optimizes the event-triggered impulsive controller (ETIC) of continuous-time nonlinear systems. A generalized-event-driven system model (GEM) is designed to characterize the impulsive dynamics over the impulsive actions. Using the GEM, we construct the ETIC which is further optimized by the proposed event-triggered impulsive adaptive dynamic programming (ETIADP) method. By utilizing a new value updating technique, the liquid-updating ETIADP (LADP) is presented such that the computing devices with low memory capacities can be used to carry out the optimization scheme. By analyzing the admissibility, convergence and error bound properties of ETIADP and LADP, it is proved that the optimal impulsive performance index function and ETIC can be successfully obtained. An experimental study is given to validate the effectiveness of the proposed approaches. Mingming Liang, Yonghua Wang 0001, Derong Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Dynamic Event-Triggered Control for Hierarchical Differential GamesabstractThis paper proposes a novel dynamic event-triggered control method for a class of completely unknown nonaffine hierarchical differential games, incorporating asymmetric boundaries in both system states and control strategies. To tackle this problem, dynamic feedback and mapping functions are first introduced to construct an unconstrained affine augmented system. Then, integral reinforcement learning techniques are used to derive the Hamilton-Jacobi equation without the original system dynamics. Furthermore, dynamic event-triggered control is employed to alleviate the network transmission burden. During the algorithm implementation, critic neural networks are designed for each agent. Analysis results show that the states and weights are ultimately uniformly bounded. Finally, simulation results using the torsional pendulum system and RLC circuit system validate the effectiveness of the present method. Shan Xue 0004, Biao Luo 0001, Weidong Zhang 0004, Derong Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Relaxed Optimal Control With Self-Learning Horizon for Discrete-Time Stochastic DynamicsabstractThe innovation of optimal learning control methods is profoundly propelled due to the improvement of the learning ability. In this article, we investigate the synthesis of initialization and acceleration for optimal learning control algorithms. This approach contrasts with traditional methods that concentrate solely on either the improvement of initialization or acceleration. Specifically, we establish a novel relaxed policy iteration (PI) algorithm with self-learning horizon for stochastic optimal control. Notably, by suitably utilizing self-learning horizon, we can directly evaluate inadmissible policies to reduce the initialization burden. Meanwhile, the inadmissible policy can be rapidly optimized with few learning iterations. Then, several critical conclusions of relaxed optimal control are established by discussing algorithm convergence and system stability. Furthermore, to provide the convincing application potentials, a class of unconventional problems is effectively solved by the relaxed PI algorithm, including the dynamics with external noises and nonzero equilibrium. Finally, we present a series of nonlinear benchmarks with practical applications to comprehensively evaluate the performance of relaxed PI. The experimental results obtained from these diverse benchmarks uniformly highlight the effectiveness of self-learning horizon mechanism. Ding Wang 0001, Jiangyu Wang, Ao Liu 0012, Derong Liu 0001, Junfei Qiao 0001 |
IEEE Trans. Cybern. | 4 |
| 2025 | Integral Reinforcement Learning-Based Dynamic Event-Triggered Nonzero-Sum Games of USVsabstractIn this article, an integral reinforcement learning (IRL) method is developed for dynamic event-triggered nonzero-sum (NZS) games to achieve the Nash equilibrium of unmanned surface vehicles (USVs) with state and input constraints. Initially, a mapping function is designed to map the state and control of the USV into a safe environment. Subsequently, IRL-based coupled Hamilton-Jacobi equations, which avoid dependence on system dynamics, are derived to solve the Nash equilibrium. To conserve computational resources and reduce network transmission burdens, a static event-triggered control is initially designed, followed by the development of a more flexible dynamic form. Finally, a critic neural network is designed for each player to approximate its value function and control policy. Rigorous proofs are provided for the uniform ultimate boundedness of the state and the weight estimation errors. The effectiveness of the present method is demonstrated through simulation experiments. Shan Xue 0004, Weidong Zhang 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2025 | Optimal Learning Output Tracking Control: A Model-Free Policy Optimization Method With Convergence AnalysisabstractOptimal learning output tracking control (OLOTC) in a model-free manner has received increasing attention in both the intelligent control and the reinforcement learning (RL) communities. Although the model-free tracking control has been achieved via off-policy learning and Q-learning, another popular RL idea of direct policy learning, with its easy-to-implement feature, is still rarely considered. To fill this gap, this article aims to develop a novel model-free policy optimization (PO) algorithm to achieve the OLOTC for unknown linear discrete-time (DT) systems. The iterative control policy is parameterized to directly improve the discounted value function of the augmented system via the gradient-based method. To implement this algorithm in a model-free manner, a model-free two-point policy gradient (PG) algorithm is designed to approximate the gradient of discounted value function by virtue of the sampled states and the reference trajectories. The global convergence of model-free PO algorithm to the optimal value function is demonstrated with the sufficient quantity of samples and proper conditions. Finally, numerical simulation results are provided to validate the effectiveness of the present method. Mingduo Lin, Bo Zhao 0015, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | FX-DARTS: Designing Topology-Unconstrained Architectures With Differentiable Architecture Search and Entropy-BasedSuper-Network ShrinkingabstractStrong priors are imposed on the search space of differentiable architecture search (DARTS), such that cells of the same type share the same topological structure and each intermediate node retains two operators from distinct nodes. While these priors reduce optimization difficulties and improve the applicability of searched architectures, they hinder the subsequent development of automated machine learning (auto-ML) and prevent the optimization algorithm from exploring more powerful neural networks through improved architectural flexibility. This article aims to reduce these prior constraints by eliminating restrictions on cell topology and modifying the discretization mechanism for super-networks. Specifically, the flexible DARTS (FX-DARTS) method, which leverages an entropy-based super-network shrinking (ESS) framework, is presented to address the challenges arising from the elimination of prior constraints. Notably, FX-DARTS enables the derivation of neural architectures without strict prior rules while maintaining the stability in the enlarged search space. Experimental results on image classification benchmarks demonstrate that FX-DARTS is capable of exploring a set of neural architectures with competitive trade-offs between performance and computational complexity within a single search procedure. Xuan Rao, Bo Zhao 0015, Derong Liu 0001, Cesare Alippi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Parallel Multistep Evaluation With Efficient Data Utilization for Safe Neural Critic Control and Its Application to Orbital Maneuver SystemsabstractData-driven methods have significantly advanced optimal learning control, but some approaches overlook systematic considerations of data utilization, including safety, efficiency, and error accumulation. To address the neglects in safe neural critic control, this article introduces a parallel multistep evaluation mechanism that combines data from the system interaction with data generated by data-driven models. Based on this evaluation mechanism, we propose a novel parallel multistep Q-learning algorithm that enhances data utilization efficiency and mitigates the error accumulation. Furthermore, we formulate a novel control barrier function (CBF) to ensure safety during learning and control processes, which is capable of dealing with asymmetric constraints and adjusting the constraint strength. In addition, the analysis reveals that multistep information introduced by data-driven models influences the learning performance of actor-critic neural networks (NNs). Finally, parallel multistep Q-learning, which makes use of data in aspects of safety, efficiency, and error bounds, is validated within an orbital maneuver system. Jiangyu Wang, Ding Wang 0001, Derong Liu 0001, Junfei Qiao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | A Hybrid Adaptive Dynamic Programming for Optimal Tracking Control of USVsabstractThis article presents an efficient method for solving the optimal tracking control policy of unmanned surface vehicles (USVs) using a hybrid adaptive dynamic programming (ADP) approach. This approach integrates data-driven integral reinforcement learning (IRL) and dynamic event-driven (DED) mechanisms into the solution of the control policy of the established augmented system while obtaining both the feedforward and feedback components of the tracking controller. For the USV model and the reference trajectory, an augmented system is established, and the tracking Hamilton-Jacobi-Bellman (HJB) equation is derived based on IRL, aiming to fully utilize system data information and reduce model dependency. For the solution of the tracking HJB equation, the DED-based controller update rule is used to further reduce the burden of network transmission. In implementing the ADP method, the DED experience replay-based weight update rule is utilized to recycle data resources. Experiments show that compared with the static event-driven (SED) approach, the DED approach reduces the sample size by 78% and increases the average interval by about four times. Shan Xue 0004, Weidong Zhang 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Self-Triggered Approximate Optimal Neuro-Control for Nonlinear Systems Through Adaptive Dynamic ProgrammingabstractIn this article, a novel self-triggered approximate optimal neuro-control scheme is presented for nonlinear systems by utilizing adaptive dynamic programming (ADP). According to the Bellman principle of optimality, the cost function of the general nonlinear system is approximated by building a critic neural network with a nested updating weight vector. Thus, the Hamilton-Jacobi-Bellman equation is solved to indirectly obtain the approximate optimal neuro-control input. In order to reduce the computation, the communication bandwidth, and the energy consumption, an appropriate self-triggering condition is designed as an alternative way to predict the updating time instants of the approximate optimal neuro-control policy. On the basis of Lyapunov's direct method, the stability of the closed-loop nonlinear system is analyzed and guaranteed to be uniformly ultimately bounded. Simulation results of two practical systems illustrate the present ADP-based self-triggered approximate optimal neuro-control scheme to be reasonable and effective. Bo Zhao 0015, Shunchao Zhang, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Distributed Optimal Containment Control of Wheeled Mobile Robots via Adaptive Dynamic ProgrammingabstractIn this article, the distributed optimal containment (DOC) control of wheeled mobile robots (WMRs) is investigated via adaptive dynamic programming. To begin with, a novel performance index function which contains containment errors and their derivatives is designed for each following WMR without requiring the discount factor, which simplifies the controller design process and enhances the practicality of the control method. Subsequently, the DOC control of WMRs is formulated as a differential graphical game whose Nash equilibrium can be formed by using the optimal responses of all following WMRs. Moreover, a critic-only structure is built to obtain an approximate DOC control law, which provides a solution for the coupled Hamilton–Jacobi–Bellman equation of each following WMR. Stability analysis demonstrates that the containment error of each following WMR is uniformly ultimately bounded. Finally, a group of WMRs are utilized to verify the effectiveness of the present DOC control scheme. Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Massive Multi-agent Mean-Field Game Using Online Federated Adaptive Critic-Density Learning
Mingduo Lin, Guoling Yuan, Bo Zhao 0015, Derong Liu 0001 |
ICONIP (4) | 4 |
| 2024 | ENAO: Evolutionary Neural Architecture Optimization in the Approximate Continuous Latent Space of a Deep Generative ModelabstractNeural architecture search (NAS) has emerged as a transformative approach for automating the design of neural networks, demonstrating exceptional performance across a variety of tasks. Numerous NAS methods aim to optimize neural architectures within discrete or continuous search spaces, but each method possesses its own inherent limitations. Additionally, the search efficiency is notably impeded by suboptimal encoding methods, presenting an ongoing challenge. In response to these obstacles, this paper introduces a novel approach, evolutionary neural architecture optimization (ENAO), which optimizes architectures in an approximate continuous search space. ENAO begins with training a deep generative model to embed discrete architectures into a condensed latent space, leveraging unsupervised representation learning. Subsequently, evolutionary algorithm is employed to refine neural architectures within this approximate continuous latent space. Empirical comparisons against several NAS benchmarks underscore the effectiveness of the ENAO method. Thanks to its foundation in deep unsupervised representation learning, ENAO demonstrates a distinguished ability to identify high-quality architectures with fewer evaluations and achieve state-of-the-art result in NAS-Bench-201 dataset. Overall, the ENAO method is a promising approach for optimizing neural network architectures in an approximate continuous search space with evolutionary algorithms and may be a useful tool for researchers and practitioners in the field of NAS. Xuan Rao, Shaojie Liu, Bo Zhao 0015, Derong Liu 0001 |
IJCNN | 5 |
| 2024 | Temporal Normalization Flow for Probabilistic Time Series ForecastingabstractTime series data has the characteristics of strong randomness, complex data structures, and high non-stationarity. Time series probabilistic forecasting is of great significance for quantifying the uncertainty in time series. This study presents a novel probabilistic forecasting approach that is distribution-free and relies on the normalization flow technology. The developed method utilizes normalization flow to transform complex target distributions into simpler ones, which is convenient for probabilistic modeling of the target distributions. Additionally, the method we have developed employs both convolution and attention mechanisms to identify temporal patterns across both short and extended timeframes within the time series. Comprehensive empirical assessments across various real-world time series datasets confirm that the proposed method surpasses standard models in predictive accuracy. Jiarui Ye, Bo Zhao 0015, Derong Liu 0001 |
INDIN | 3 |
| 2024 | Uncertainty-based Continual Learning for Neural Networks with Low-rank Variance MatricesabstractBayesian inference has provided the continual learning (CL) with an elegant framework where past experiences and new knowledge are consolidated into the posterior constantly. Typical approaches rely on Bayesian neural networks whose parameters are updated by variational inference, namely, maximizing the evidence lower bound of log-likelihood. In this paper, we discuss the effects of local reparameterization on the optimization of such networks in the context of CL. The empirical results show that it does not only increase the inference speed of neural networks, but also enhance the CL performance in some scenarios. Additionally, motivated by the observation that variance matrices have low-rank structures, we propose the d-tied variational continual learning (d-tied-VCL) to improve the parameter efficiency of variational continual learning (VCL). Experiments on random classification, per-muted MNIST, and split CIFAR100 show that even VCL with rank-1 variance matrices achieves competitive performance. Xuan Rao, Bo Zhao 0015, Derong Liu 0001 |
SMC | 3 |
| 2024 | Stable approximate Q-learning under discounted cost for data-based adaptive tracking control
Zhantao Liang, Mingming Ha, Derong Liu 0001, Yonghua Wang 0001 |
Neurocomputing | 3 |
| 2024 | Dynamic compensator-based near-optimal control for unknown nonaffine systems via integral reinforcement learning
Jinquan Lin, Bo Zhao 0015, Derong Liu 0001, Yonghua Wang 0001 |
Neurocomputing | 3 |
| 2024 | Semi-supervised accuracy predictor-based multi-objective neural architecture search
Songyi Xiao, Bo Zhao 0015, Derong Liu 0001 |
Neurocomputing | 3 |
| 2024 | Safe Reinforcement Learning and Adaptive Optimal Control With Applications to Obstacle Avoidance ProblemabstractThis paper presents a novel composite obstacle avoidance control method to generate safe motion trajectories for autonomous systems in an adaptive manner. First, system safety is described using forward invariance, and the barrier function is encoded into the cost function such that the obstacle avoidance problem can be characterized by an infinite-horizon optimal control problem. Next, a safe reinforcement learning framework is proposed by combining model-based policy iteration and state-following-based approximation. Upon real-time data and extrapolated experience data, this learning design is implemented through the actor-critic structure, in which critic networks are tuned by gradient-descent adaption and actor networks produce adaptive control policies via gradient projection. Then, system stability and weight convergence are theoretically analyzed using Lyapunov method. Finally, the proposed learning-based controller is demonstrated on a two-dimensional single integrator system and a nonlinear unicycle kinematic system. Simulation results reveal that the system or agent can smoothly reach the target point while keeping a safe distance from each obstacle; at the same time, other three avoidance control methods are used to provide side-by-side comparisons and to verify some claimed advantages of the present method.Note to Practitioners—This paper is motivated by the obstacle avoidance problem of real-time navigation of an agent to the target point, which applies to practical autonomous systems such as vehicles and robots. Pre-generative methods and reactive methods have been widely employed to generate safe motion trajectories in the obstacle environment. However, these methods cannot strike a good balance between safety and optimality. In this paper, the obstacle avoidance problem is formulated in the sense of optimal control, and a safe reinforcement learning method is designed to generate safe motion trajectories. This method combines the advantages of model-based policy iteration and state-following-based approximation, in which the former ensures regional optimality while the latter ensures local safety. Based on the proposed adaptive tuning laws, engineers are able to design learning-based avoidance controllers in the environment with static obstacles. In future research, we will address the dynamic avoidance problem against moving obstacles. Ke Wang 0037, Chaoxu Mu, Zhen Ni, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | A Novel Online Adaptive Dynamic Programming Algorithm With Adjustable Convergence RateabstractThis article develops a novel online adaptive dynamic programming algorithm with adjustable convergence rate to address the optimal control problem of nonlinear systems. Relaxation factors are introduced to tune the convergence rate of value function sequence online. A novel update law based on recursive least squares is developed to adjust the weight of critic neural network at the sampling instant. The uniform ultimate boundedness of the neural network estimation error and the closed-loop system state are analyzed by utilizing the Lyapunov technique. Finally, the effectiveness of the present algorithm is demonstrated by executing three simulation examples. Yonghua Wang 0001, Zheliang Zhang, Yongwei Zhang 0002, Mingming Liang, Derong Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Novel Discounted Adaptive Critic Control Designs With Accelerated Learning FormulationabstractInspired by the successive relaxation method, a novel discounted iterative adaptive dynamic programming framework is developed, in which the iterative value function sequence possesses an adjustable convergence rate. The different convergence properties of the value function sequence and the stability of the closed-loop systems under the new discounted value iteration (VI) are investigated. Based on the properties of the given VI scheme, an accelerated learning algorithm with convergence guarantee is presented. Moreover, the implementations of the new VI scheme and its accelerated learning design are elaborated, which involve value function approximation and policy improvement. A nonlinear fourth-order ball-and-beam balancing plant is used to verify the performance of the developed approaches. Compared with the traditional VI, the present discounted iterative adaptive critic designs greatly accelerate the convergence rate of the value function and reduce the computational cost simultaneously. Mingming Ha, Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Event-Triggered Robust Adaptive Dynamic Programming for Multiplayer Stackelberg-Nash Games of Uncertain Nonlinear SystemsabstractIn this article, an event-triggered robust adaptive dynamic programming (ETRADP) algorithm is developed to solve a class of multiplayer Stackelberg-Nash games (MSNGs) for uncertain nonlinear continuous-time systems. Considering the different roles of players in the MSNG, the hierarchical decision-making process is described as the designed value functions for the leader and all followers, which assist to transform the robust control problem of the uncertain nonlinear system into an optimal regulation problem of the nominal system. Then, an online policy iteration algorithm is formulated to solve the derived coupled Hamilton-Jacobi equation. Meanwhile, an event-triggered mechanism is designed to alleviate computational and communication burdens. Moreover, critic neural networks (NNs) are constructed to obtain the event-triggered approximate optimal control polices for all players, which constitute the Stackelberg-Nash equilibrium of the MSNG. By using Lyapunov's direct method, the stability of the closed-loop uncertain nonlinear system is guaranteed under the ETRADP-based control scheme in the sense of uniform ultimate boundedness. Finally, a numerical simulation is provided to demonstrate the effectiveness of the present ETRADP-based control scheme. Mingduo Lin, Bo Zhao 0015, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Rapid Adaptation for Active Pantograph Control in High-Speed Railway via Deep Meta Reinforcement LearningabstractActive pantograph control is the most promising technique for reducing contact force (CF) fluctuation and improving the train's current collection quality. Existing solutions, however, suffer from two significant limitations: 1) they are incapable of dealing with the various pantograph types, catenary line operating conditions, changing operating speeds, and contingencies well and 2) it is challenging to implement in practical systems due to the lack of rapid adaptability to a new pantograph-catenary system (PCS) operating conditions and environmental disturbances. In this work, we alleviate these problems by developing a revolutionary context-based deep meta-reinforcement learning (CB-DMRL) algorithm. The proposed CB-DMRL algorithm combines Bayesian optimization (BO) with deep reinforcement learning (DRL), allowing the general agent to adapt to new tasks quickly and efficiently. We evaluated the CB-DMRL algorithm's performance on a proven PCS model. The experimental results demonstrate that meta-training DRL policies with latent space swiftly adapt to new operating conditions and unknown perturbations. The meta-agent adapts quickly after two iterations with a high reward, which require only ten spans, approximately equal to 0.5 km of PCS interaction data. Compared with state-of-the-art DRL algorithms and traditional solutions, the proposed method can promptly traverse scenario changes and reduce CF fluctuations, resulting in an excellent performance. Hui Wang 0063, Zhigang Liu 0001, Zhiwei Han, Yanbo Wu, Derong Liu 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | Adaptive Tracking Control for Underactuated Double Pendulum Overhead Cranes With Variable Cable LengthabstractAlthough the literature on control of overhead crane systems is extensive and relatively mature, there is still a need to develop strategies that can simultaneously handle factors such as the double pendulum effect, variable cable length, input saturation, input dead zones, and external disturbances. This article is concerned with adaptive tracking control for underactuated overhead cranes in the presence of the above-mentioned challenging effects. The proposed controller is composed of the following two components. First, a tracking signal vector that effectively reduces system swing magnitudes is constructed to improve the transient performance and guarantee smooth operation of the system. Second, an adaptive law is designed to estimate and compensate for the overall effects of the friction, the external disturbances, and certain nonlinearities. The system stability has been proved rigorously via the Lyapunov method and Barbalat's lemma. Extensions to the cases with input saturation and dead zones have also been discussed. Extensive numerical simulations have been conducted to verify the performance and robustness of the proposed controller, in comparison to some existing methods. Fuxing Yao, Ai-Guo Wu 0001, Mehdi Golestani, Derong Liu 0001, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Cybern. | 4 |
| 2024 | Explainable Intelligent Fault Diagnosis for Nonlinear Dynamic Systems: From Unsupervised to Supervised LearningabstractThe increased complexity and intelligence of automation systems require the development of intelligent fault diagnosis (IFD) methodologies. By relying on the concept of a suspected space, this study develops explainable data-driven IFD approaches for nonlinear dynamic systems. More specifically, we parameterize nonlinear systems through a generalized kernel representation for system modeling and the associated fault diagnosis. An important result obtained is a unified form of kernel representations, applicable to both unsupervised and supervised learning. More importantly, through a rigorous theoretical analysis, we discover the existence of a bridge (i.e., a bijective mapping) between some supervised and unsupervised learning-based entities. Notably, the designed IFD approaches achieve the same performance with the use of this bridge. In order to have a better understanding of the results obtained, both unsupervised and supervised neural networks are chosen as the learning tools to identify the generalized kernel representations and design the IFD schemes; an invertible neural network is then employed to build the bridge between them. This article is a perspective article, whose contribution lies in proposing and formalizing the fundamental concepts for explainable intelligent learning methods, contributing to system modeling and data-driven IFD designs for nonlinear dynamic systems. Hongtian Chen, Zhigang Liu 0001, Cesare Alippi, Biao Huang 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Neuro-Optimal Event-Triggered Impulsive Control for Stochastic Systems via ADPabstractThis article presents a novel neural-network-based optimal event-triggered impulsive control method. First, a novel general-event-based impulsive transition matrix (GITM) is constructed to represent the probability distribution evolving characteristics regarding all system states across the impulsive actions, rather than the prefixed timing sequence. On the foundation of this GITM, the event-triggered impulsive adaptive dynamic programming (ETIADP) algorithm and its high-efficiency version (HEIADP) are developed to deal with the optimization problems for stochastic systems with event-triggered impulsive controls. It is shown that the obtained controller design scheme can reduce the computational and communication burden caused by updating the controller periodically. By analyzing the admissibility, monotonicity, and optimality properties of ETIADP and HEIADP, we further establish the approximation error bound of the neural networks to address the connection between the ideal and neural-network-based realizations of the present methods. It is proven that the iterative value functions of both the ETIADP and HEIADP algorithms fall in a small neighborhood of the optimum as the iteration index increases to infinity. By adopting a novel task synchronization mechanism, the proposed HEIADP algorithm fully utilizes the computing resources of multiprocessor systems (MPSs), while significantly reducing the memory requirement compared to traditional ADP approaches. Finally, we carry out a numerical study to show that the proposed methods can fulfill the desired goals. Mingming Liang, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Guest Editorial: Special Issue on Explainable Representation Learning-Based Intelligent Inspection and Maintenance of Complex SystemsabstractOver the past decade, representation learning has received particular attention in the intelligent inspection and maintenance of complex systems thanks to its overwhelming advantages in discovering and mining hidden knowledge representations. The room for in-depth investigations of representation learning-related topics remains open, especially explainable approaches for intelligent inspection and maintenance of complex systems. The primary objective of this special issue, entitled “Explainable Representation Learning-based Intelligent Inspection and Maintenance of Complex Systems,” of IEEE Transactions on Neural Networks and Learning Systems is to provide the related latest achievements made by researchers and practitioners on the one hand and to identify critical issues and challenges for future investigation on the other hand. Zhigang Liu 0001, Cesare Alippi, Hongtian Chen, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Synchronization of Delayed Memristor-Based Neural Networks via Pinning Control With Local InformationabstractIn this article, a novel pinning control method, only requiring information from partial nodes, is developed to synchronize drive-response memristor-based neural networks (MNNs) with time delay. An improved mathematical model of MNNs is established to describe the dynamic behaviors of MNNs accurately. In the existing literature, pinning controllers for synchronization of drive-response systems were designed based on information of all nodes, but in some specific situations, the control gains may be very large and challenging to realize in practice. To overcome this problem, a novel pinning control policy is developed to achieve synchronization of delayed MNNs, which depends only on local information of MNNs, for reducing communication and calculation burdens. Furthermore, sufficient conditions for synchronization of delayed MNNs are provided. Finally, numerical simulation and comparative experiments are conducted to verify the effectiveness and superiority of the proposed pinning control method. Zhanyu Yang, Bo Zhao 0015, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Distributed Fault Tolerant Consensus Control of Nonlinear Multiagent Systems via Adaptive Dynamic ProgrammingabstractThis article develops a distributed fault-tolerant consensus control (DFTCC) approach for multiagent systems by using adaptive dynamic programming. By establishing a local fault observer, the potential actuator faults of each agent are estimated. Subsequently, the DFTCC problem is transformed into an optimal consensus control problem by designing a novel local value function for each agent which contains the estimated fault, the consensus errors, and the control laws of the local agent and its neighbors. In order to solve the coupled Hamilton-Jacobi-Bellman equation of each agent, a critic-only structure is established to obtain the approximate local optimal consensus control law of each agent. Moreover, by using Lyapunov's direct method, it is proven that the approximate local optimal consensus control law guarantees the uniform ultimate boundedness of the consensus error of all agents, which means that all following agents with potential actuator faults synchronize to the leader. Finally, two simulation examples are provided to validate the effectiveness of the present DFTCC scheme. Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Shunchao Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Liquid-Updating Impulsive Adaptive Dynamic Programming for Continuous Nonlinear SystemsabstractThis article focuses on designing the optimal impulsive controller (IMC) of continuous-time nonlinear systems. A general-event-based dynamics (GED) is constructed to describe the state transition characteristics across the impulsive actions for the continuous-time nonlinear systems. Then, based on the GED, a new IMC design scheme and an impulsive adaptive dynamic programming (IADP) algorithm are developed, which possess strong generality and feasibility. Next, by introducing a novel policy-improving mechanism the liquid-updating IADP (LIADP) algorithm is established, which is more flexible to fit into the memory-limited computing devices, thus improving the flexibility and realization efficiency of the ADP-based approaches. The developed methods are proved to converge to the optimal impulsive performance index function and obtain the optimal impulsive controllers for the continuous-time nonlinear systems. Finally, a numerical study is provided to verify the effectiveness of the present IADP and LIADP algorithms. Mingming Liang, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Evolution-Guided Adaptive Dynamic Programming for Nonlinear Optimal ControlabstractIn this article, an evolution-guided adaptive dynamic programming (EGADP) algorithm is developed to address the optimal regulation problems for the nonlinear systems. In the traditional adaptive dynamic programming algorithms, policy improvement is typically reliant on the gradient information, according to the first order necessity condition. However, these methods encounter limitations when calculating the gradient information becomes infeasible or system dynamics is not differentiable. In response to this challenge, the evolutionary computation is harnessed by EGADP to search for a superior policy during policy improvement. Therefore, compared with the traditional methods, scenarios that gradient information is unavailable can effectively be handled by EGADP. Additionally, the convergence of the algorithm is proven to enhance the rigorousness of the developed method. Finally, the three simulation experiments with realistic physical backgrounds are conducted to comprehensively demonstrate the effectiveness of the established method from different perspectives. Ding Wang 0001, Haiming Huang, Derong Liu 0001, Junfei Qiao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Event-Triggered Decentralized Integral Sliding Mode Control for Input-Constrained Nonlinear Large-Scale Systems With Actuator FailuresabstractIn this article, an event-triggered decentralized integral sliding mode control (ETDISMC) method is investigated for a class of input-constrained nonlinear large-scale systems with actuator failures based on adaptive dynamic programming (ADP). An integral sliding mode control method is developed to maintain the subsystem trajectories on the sliding mode surface, eliminate the effect of actuator failures, and obtain the sliding mode dynamics (SMDs). Then, the control problem is transformed into an optimal control (OC) problem for the nominal form of the SMDs by constructing a modified local value function. To obtain the event-triggered OC law, a critic-only structure is applied to approximate the local optimal value function of the nominal subsystem for solving the event-triggered Hamilton–Jacobi–Bellman equation. An event-triggered ADP control method is developed to decrease the updating frequency of the OC law and to reduce the computational burden. In addition, an experience replay-based weight updating policy is presented to relax the persistence of excitation condition. Furthermore, we prove that the developed method can guarantee the closed-loop system to be asymptotically stable by using Lyapunov’s direct method. Finally, a numerical example and a practical system are employed for simulation to demonstrate the effectiveness of the proposed ETDISMC scheme. Shunchao Zhang, Bo Zhao 0015, Derong Liu 0001, Yongwei Zhang 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Event-Triggered Constrained H∞ Control Using Concurrent Learning and ADP
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Dongsheng Guo 0001 |
ICONIP (8) | 3 |
| 2023 | Fault tolerant control for a class of nonlinear systems with multiple faults using neuro-dynamic programming
Chujian Zeng, Bo Zhao 0015, Derong Liu 0001 |
Neurocomputing | 3 |
| 2023 | Adaptive dynamic programming-based hierarchical decision-making of non-affine systems
Danyu Lin, Shan Xue 0004, Derong Liu 0001, Mingming Liang, Yonghua Wang 0001 |
Neural Networks | 3 |
| 2023 | Event-triggered adaptive dynamic programming for decentralized tracking control of input constrained unknown nonlinear interconnected systems
Qiuye Wu, Bo Zhao 0015, Derong Liu 0001, Marios M. Polycarpou |
Neural Networks | 3 |
| 2023 | Policy gradient adaptive dynamic programming for nonlinear discrete-time zero-sum games with unknown dynamics
Mingduo Lin, Bo Zhao 0015, Derong Liu 0001 |
Soft Comput. | 3 |
| 2023 | Deep Learning-Based Trajectory Planning and Control for Autonomous Ground Vehicle Parking ManeuverabstractIn this paper, a novel integrated real-time trajectory planning and tracking control framework capable of dealing with autonomous ground vehicle (AGV) parking maneuver problems is presented. In the motion planning component, a newly-proposed idea of utilizing deep neural networks (DNNs) for approximating optimal parking trajectories is further extended by taking advantages of a recurrent network structure. The main aim is to fully exploit the inherent relationships between different vehicle states in the training process. Furthermore, two transfer learning strategies are applied such that the developed motion planner can be adapted to suit various AGVs. In order to follow the planned maneuver trajectory, an adaptive learning tracking control algorithm is designed and served as the motion controller. By adapting the network parameters, the stability of the proposed control scheme, along with the convergence of tracking errors, can be theoretically guaranteed. In order to validate the effectiveness and emphasize key features of our proposal, a number of experimental studies and comparative analysis were executed. The obtained results reveal that the proposed strategy can enable the AGV to fulfill the parking mission with enhanced motion planning and control performance.Note to Practitioners—This article was motivated by the problem of optimal automatic parking planning and tracking control for autonomous ground vehicles (AGVs) maneuvering in a restricted environment (e.g., constrained parking regions). A number of challenges may arise when dealing with this problem (e.g., the model uncertainties involved in the vehicle dynamics, system variable limits, and the presence of external disturbances). Existing approaches to address such a problem usually exploit the merit of optimization-based planning/control techniques such as model predictive control and dynamic programming in order for an optimal solution. However, two practical issues may require further considerations: 1). The nonlinear (re)optimization process tends to consume a large amount of computing power and it might not be affordable in real-time; 2). Existing motion planning and control algorithms might not be easily adapted to suit various types of AGVs. To overcome the aforementioned issues, we present an idea of utilizing the recurrent deep neural network (RDNN) for planning optimal parking maneuver trajectories and an adaptive learning NN-based (ALNN) control scheme for robust trajectory tracking. In addition, by introducing two transfer learning strategies, the proposed RDNN motion planner can be adapted to suit different AGVs. In our follow-up research, we will explore the possibility of extending the developed methodology for large-scale AGV parking systems collaboratively operating in a more complex cluttered environment. Runqi Chai, Derong Liu 0001, Antonios Tsourdos, Yuanqing Xia, Senchun Chai |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | Adaptive Dynamic Programming-Based Event-Triggered Robust Control for Multiplayer Nonzero-Sum Games With Unknown DynamicsabstractIn this article, the event-triggered robust control of unknown multiplayer nonlinear systems with constrained inputs and uncertainties is investigated by using adaptive dynamic programming. To relax the requirement of system dynamics, a neural network-based identifier is constructed by using the system input-output data. Subsequently, by designing a nonquadratic value function, which contains the bounded functions, the system states, and the control inputs of all players, the event-triggered robust stabilization problem is converted into an event-triggered constrained optimal control problem. To obtain the approximate solution of the event-triggered Hamilton-Jacobi (HJ) equation, a critic network for each player is established with a novel weight updating law to relax the persistence of excitation condition based on the experience replay technique. Furthermore, according to the Lyapunov stability theorem, the present event-triggered robust optimal control ensures the multiplayer system to be uniformly ultimately bounded. Finally, two simulation examples are employed to show the effectiveness of the present method. Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Shunchao Zhang |
IEEE Trans. Cybern. | 3 |
| 2023 | An Efficient Impulsive Adaptive Dynamic Programming Algorithm for Stochastic SystemsabstractIn this study, a novel general impulsive transition matrix is defined, which can reveal the transition dynamics and probability distribution evolution patterns for all system states between two impulsive "events," instead of two regular time indexes. Based on this general matrix, the policy iteration-based impulsive adaptive dynamic programming (IADP) algorithm along with its variant, which is a more efficient IADP (EIADP) algorithm, are developed in order to solve the optimal impulsive control problems of discrete stochastic systems. Through analyzing the monotonicity, stability, and convergency properties of the obtained iterative value functions and control laws, it is proved that the IADP and EIADP algorithms both converge to the optimal impulsive performance index function. By dividing the whole impulsive policy into smaller pieces, the proposed EIADP algorithm updates the iterative policies in a "piece-by-piece" manner according to the actual hardware constraints. This feature of the EIADP method enables these ADP-based algorithms to be fully optimized to run on all "sizes" of computing devices including the ones with low memory spaces. A simulation experiment is conducted to validate the effectiveness of the present methods. Mingming Liang, Yonghua Wang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | A Novel Value Iteration Scheme With Adjustable Convergence RateabstractIn this article, a novel value iteration scheme is developed with convergence and stability discussions. A relaxation factor is introduced to adjust the convergence rate of the value function sequence. The convergence conditions with respect to the relaxation factor are given. The stability of the closed-loop system using the control policies generated by the present VI algorithm is investigated. Moreover, an integrated VI approach is developed to accelerate and guarantee the convergence by combining the advantages of the present and traditional value iterations. Also, a relaxation function is designed to adaptively make the developed value iteration scheme possess fast convergence property. Finally, the theoretical results and the effectiveness of the present algorithm are validated by numerical examples. Mingming Ha, Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Adaptive Dynamic Programming-Based Cooperative Motion/Force Control for Modular Reconfigurable Manipulators: A Joint Task Assignment ApproachabstractThis article develops a cooperative motion/force control (CMFC) scheme based on adaptive dynamic programming (ADP) for modular reconfigurable manipulators (MRMs) with the joint task assignment approach. By separating terms depending on local variables only, the dynamic model of the entire MRM system can be regarded as a set of joint modules interconnected by coupling torque. In addition, the Jacobian matrix, which reflects the interaction force of the MRM end-effector, can be mapped into each joint. Using this approach, both the motion and force tasks on the end-effector of the entire MRM system can be assigned to each joint module cooperatively. Then, by substituting the actual states of coupled joint modules with their desired ones, the norm-boundedness assumption on the interconnection of joint module can be relaxed. By using the measured input-output data of each joint module, a neural network (NN)-based robust decentralized observer, which guarantees the observation error to be asymptotically stable is established. An improved local value function is constructed for each joint module to reflect the interconnection. Then, the local Hamilton-Jacobi-Bellman equation is solved by constructing a local critic NN with a nested learning structure. Hereafter, the ADP-based CMFC is obtained by the assistance of force feedback compensation. Based on the Lyapunov stability analysis, the closed-loop MRM system is guaranteed to be uniformly ultimately bounded under the present ADP-based CMFC scheme. The simulation on a two-degree of freedom MRM system demonstrates the effectiveness of the present control approach. Bo Zhao 0015, Yongwei Zhang 0002, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Event-Triggered Local Control for Nonlinear Interconnected Systems Through Particle Swarm Optimization-Based Adaptive Dynamic ProgrammingabstractThis article investigates local control problems for nonlinear interconnected systems by using adaptive dynamic programming (ADP) with particle swarm optimization (PSO). Through constructing a proper local value function, a local critic neural network, whose weight vector is tuned via the PSO algorithm, is employed to solve the local Hamilton–Jacobi–Bellman equation. By introducing the event-triggering mechanism, the sampling time instants of each interconnected subsystem are determined by establishing a proper event-triggering condition. Then, the ADP-based event-triggered local control policy can be derived indirectly to ensure the closed-loop nonlinear interconnected system to be asymptotically stable through the Lyapunov stability analysis. The positive lower bound on the minimal intersampling instant for each interconnected subsystem is provided to exclude the Zeno behavior. Simulation results of a practical system and a numerical example demonstrate the effectiveness of the present event-triggered local control scheme. Bo Zhao 0015, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Neural network-based event-triggered integral reinforcement learning for constrained H∞ tracking control with experience replay
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004 |
Neurocomputing | 3 |
| 2022 | DMPP: Differentiable multi-pruner and predictor for neural network pruning
Bo Zhao 0015, Derong Liu 0001 |
Neural Networks | 3 |
| 2022 | Event-triggered integral reinforcement learning for nonzero-sum games with asymmetric input saturation
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004 |
Neural Networks | 3 |
| 2022 | Offline and Online Adaptive Critic Control Designs With Stability Guarantee Through Value IterationabstractThis article is concerned with the stability of the closed-loop system using various control policies generated by value iteration. Some stability properties involving admissibility criteria, the attraction domain, and so forth, are investigated. An offline integrated value iteration (VI) scheme with a stability guarantee is developed by combining the advantages of VI and policy iteration, which is convenient to obtain admissible control policies. Also, based on the concept of attraction domain, an online adaptive dynamic programming algorithm using immature control policies is developed. Remarkably, it is ensured that the state trajectory under the online algorithm converges to the origin. Particularly, for linear systems, the online ADP algorithm with a general scheme possesses more enhanced stability property. The theoretical results reveal that the stability of the linear system can be guaranteed even if the control policy sequence includes finite unstable elements. The numerical results verify the effectiveness of the present algorithms. Mingming Ha, Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Event-Triggered ADP for Tracking Control of Partially Unknown Constrained Uncertain SystemsabstractAn event-triggered adaptive dynamic programming (ADP) algorithm is developed in this article to solve the tracking control problem for partially unknown constrained uncertain systems. First, an augmented system is constructed, and the solution of the optimal tracking control problem of the uncertain system is transformed into an optimal regulation of the nominal augmented system with a discounted value function. The integral reinforcement learning is employed to avoid the requirement of augmented drift dynamics. Second, the event-triggered ADP is adopted for its implementation, where the learning of neural network weights not only relaxes the initial admissible control but also executes only when the predefined execution rule is violated. Third, the tracking error and the weight estimation error prove to be uniformly ultimately bounded, and the existence of a lower bound for the interexecution times is analyzed. Finally, simulation results demonstrate the effectiveness of the present event-triggered ADP method. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004 |
IEEE Trans. Cybern. | 3 |
| 2022 | Leader-Following Mean-Square Consensus of Stochastic Multiagent Systems With ROUs and RONs via Distributed Event-Triggered Impulsive ControlabstractBased on the distributed event-triggered impulsive mechanism, the leader-following mean-square consensus of stochastic multiagent systems with randomly occurring uncertainties and randomly occurring nonlinearities is investigated for the first time in this article. In order to make better use of the limited communication resources, we proposed some novel communication rules among agents and corresponding control protocol. Moreover, some new triggering functions are designed for different types of agents, which cannot only ensure that the Zeno behavior can be excluded but also make the upper bound of impulsive interval in the total time sequence satisfy a newly proposed constraint condition. When the expected value of the triggering function of the i th agent is non-negative within an event time interval, the impulsive control will be triggered. If the system achieves the consensus, the triggering events of all agents will not occur after some time. The original system is transformed into the delay system by using the input delay approach. Based on the Lyapunov stability theory, several sufficient delay-independent criteria for mean-square consensus are derived by a class of Halanay impulsive differential inequalities. Finally, the effectiveness of theoretical results is illustrated by numerical simulation examples. Shiguo Peng, Derong Liu 0001, Yonghua Wang 0001, Tao Chen 0039 |
IEEE Trans. Cybern. | 3 |
| 2022 | Model-Free Adaptive Optimal Control for Unknown Nonlinear Multiplayer Nonzero-Sum GameabstractIn this article, an online adaptive optimal control algorithm based on adaptive dynamic programming is developed to solve the multiplayer nonzero-sum game (MP-NZSG) for discrete-time unknown nonlinear systems. First, a model-free coupled globalized dual-heuristic dynamic programming (GDHP) structure is designed to solve the MP-NZSG problem, in which there is no model network or identifier. Second, in order to relax the requirement of systems dynamics, an online adaptive learning algorithm is developed to solve the Hamilton-Jacobi equation using the system states of two adjacent time steps. Third, a series of critic networks and action networks are used to approximate value functions and optimal policies for all players. All the neural network (NN) weights are updated online based on real-time system states. Fourth, the uniformly ultimate boundedness analysis of the NN approximation errors is proved based on the Lyapunov approach. Finally, simulation results are given to demonstrate the effectiveness of the developed scheme. Qinglai Wei, Liao Zhu, Ruizhuo Song, Pinjia Zhang, Derong Liu 0001, Jun Xiao 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Policy Gradient Adaptive Critic Designs for Model-Free Optimal Tracking Control With Experience ReplayabstractA model-free optimal tracking controller is designed for discrete-time nonlinear systems through policy gradient adaptive critic designs (PGACDs) with experience replay (ER). By using system transformation, optimal tracking control problems are converted into optimal regulation problems. An off-policy PGACD algorithm is developed to minimize the iterative$Q$-function and improve the tracking control performance. The proposed method is realized based on the critic network and the actor network (AN), which are applied to approximate the iterative$Q$-function and the iterative control policy, respectively. Then, the policy gradient technique is introduced to derive a novel weight updating law of the AN explicitly by using measured system data only. The convergence of the iteration is established through theoretical analysis, and the uniform ultimate boundedness is demonstrated for the closed-loop system under the PGACD-based controller by using Lyapunov’s direct method. To guarantee the stability and increase the data usage efficiency of the learning process, an ER-based learning framework is designed to improve the realizability of the proposed method. Finally, simulation results of two examples are provided to demonstrate the performance of the off-policy PGACD algorithm. Mingduo Lin, Bo Zhao 0015, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Constrained Event-Triggered H∞ Control Based on Adaptive Dynamic Programming With Concurrent LearningabstractIn this article, an event-triggered$H_{\infty }$control method is proposed based on adaptive dynamic programming (ADP) with concurrent learning for unknown continuous-time nonlinear systems with control constraints. First, a system identification technique based on neural networks (NNs) is adopted to identify completely unknown systems. Second, a critic NN is employed to approximate the value function. A novel weight updating rule is developed based on the event-triggered control law and time-triggered disturbance law, which reduces controller execution times and guarantees the stability of the system. Subsequently, concurrent learning is applied to the weight updating rule to relax the demand for the traditional persistence of excitation condition that is difficult to implement online. Finally, the comparison between the time-triggered method and event-triggered method in simulation demonstrates the effectiveness of the developed constrained event-triggered ADP method. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Yin Yang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Event-Triggered Control of Discrete-Time Zero-Sum Games via Deterministic Policy Gradient Adaptive Dynamic ProgrammingabstractIn order to address zero-sum game problems for discrete-time (DT) nonlinear systems, this article develops a novel event-triggered control (ETC) approach based on the deterministic policy gradient (PG) adaptive dynamic programming (ADP) algorithm. By adopting the input and output data, the proposed ETC method updates the control law and the disturbance law with a gradient descent algorithm. Compared with the conventional PG ADP-based control scheme, the present controller is updated aperiodically to reduce the computational and communication burden. Then, the actor-critic-disturbance framework is adopted to obtain the optimal control law and the worst disturbance law, which guarantee the input-to-state stability of the closed-loop system. Moreover, a novel neural network weight updating law which guarantees the uniform ultimate boundedness of weight estimation errors is provided based on the experience replay technique. Finally, the validity of the present method is verified by simulation of two DT nonlinear systems. Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001, Shunchao Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Neural-network-based discounted optimal control via an integrated value iteration with accuracy guarantee
Mingming Ha, Ding Wang 0001, Derong Liu 0001 |
Neural Networks | 3 |
| 2021 | Observer-based event-triggered control for zero-sum games of input constrained multi-player nonlinear systems
Shunchao Zhang, Bo Zhao 0015, Derong Liu 0001, Yongwei Zhang 0002 |
Neural Networks | 3 |
| 2021 | Particle swarm optimized neural networks based local tracking control scheme of unknown nonlinear interconnected systems
Bo Zhao 0015, Fangchao Luo, Haowei Lin, Derong Liu 0001 |
Neural Networks | 4 |
| 2021 | Event-triggered adaptive dynamic programming for multi-player zero-sum games with unknown dynamics
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001 |
Soft Comput. | 3 |
| 2021 | Periodic Event-Triggered Suboptimal Control With Sampling Period and Performance AnalysisabstractIn this paper, the periodic event-triggered suboptimal control (PETSOC) method is developed for continuous-time linear systems. Different from event-triggered control, where the triggering condition is monitored continuously, the developed PETSOC method only verifies the triggering condition periodically at sampling instants, which further reduces computational resources. First, the control gain of the PETSOC is designed based on the algebraic Riccati equation. Subsequently, the periodic event-triggering condition is proposed for the suboptimal control method, which is only verified at sampling instants periodically. The sampling period is determined and analyzed based on the continuous form of the triggering condition. Moreover, the stability and the performance upper bound of the closed-loop system with the PETSOC are proved. Finally, the effectiveness of the developed PETSOC is validated through simulation on an unstable batch reactor. Biao Luo 0001, Tingwen Huang, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Policy Iteration Q-Learning for Data-Based Two-Player Zero-Sum Game of Linear Discrete-Time SystemsabstractIn this article, the data-based two-player zero-sum game problem is considered for linear discrete-time systems. This problem theoretically depends on solving the discrete-time game algebraic Riccati equation (DTGARE), while it requires complete system dynamics. To avoid solving the DTGARE, the Q -function is introduced and a data-based policy iteration Q -learning (PIQL) algorithm is developed to learn the optimal Q -function by using data collected from the real system. Writing the Q -function in a quadratic form, it is proved that the PIQL algorithm is equivalent to the Newton iteration method in the Banach space by using the Fréchet derivative. Then, the convergence of the PIQL algorithm can be guaranteed by Kantorovich's theorem. For the realization of the PIQL algorithm, the off-policy learning scheme is proposed using real data rather than the system model. Finally, the efficiency of the developed data-based PIQL method is validated through simulation studies. Biao Luo 0001, Yin Yang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Sliding-Mode Surface-Based Approximate Optimal Control for Uncertain Nonlinear Systems With Asymptotically Stable Critic StructureabstractThis article develops a novel sliding-mode surface (SMS)-based approximate optimal control scheme for a large class of nonlinear systems affected by unknown mismatched perturbations. The observer-based perturbation estimation procedure is employed to establish the online updated value function. The solution to the Hamilton-Jacobi-Bellman equation is approximated by an SMS-based critic neural network whose weights error dynamics is designed to be asymptotically stable by nested update laws. The sliding-mode control strategy is combined with the approximate optimal control design procedure to obtain a faster control action. The stability is proved based on the Lyapunov's direct method. The simulation results show the effectiveness of the developed control scheme. Bo Zhao 0015, Derong Liu 0001, Cesare Alippi |
IEEE Trans. Cybern. | 2 |
| 2021 | Event-Triggered Adaptive Dynamic Programming for Unmatched Uncertain Nonlinear Continuous-Time SystemsabstractIn this article, an event-triggered adaptive dynamic programming (ADP) method is proposed to solve the robust control problem of unmatched uncertain systems. First, the robust control problem with unmatched uncertainties is transformed into the optimal control design for an auxiliary system. Subsequently, to reduce controller executions and save computational and communication resources, an event-triggering mechanism is introduced. By using a critic neural network (NN) to approximate the value function, novel concurrent learning is developed to learn NN weights, which avoids the requirement of an initial admissible control and the persistence of excitation condition. Moreover, it is proven that the developed event-triggered ADP controller guarantees the robustness of the uncertain system and the uniform ultimate boundedness of the NN weight estimation error. Finally, by using the F-16 aircraft and the inverted pendulum with unmatched uncertainties as examples, the simulation results show the effectiveness of the developed event-triggered ADP method. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Robust Exponential Synchronization for Memristor Neural Networks With Nonidentical Characteristics by Pinning ControlabstractIn this paper, robust exponential synchronization of memristor-based neural networks (MNNs) with nonidentical characteristics is investigated. Coefficient mismatch, time-varying delay mismatch, and activation function mismatch are considered between the drive and the response MNNs. Pinning control strategy is developed to realize robust exponential synchronization and the stability criteria is established by using the Lyapunov function method and differential inclusion theory. Furthermore, the stable region of controller parameters is computed to guarantee that the synchronization errors enter a predetermined error bound within given settling time. Finally, the effectiveness of the proposed methods is verified by the numerical simulations. The methods presented in this paper offer novel schemes for robust exponential synchronization of nonidentical MNNs. Yueheng Li, Biao Luo 0001, Derong Liu 0001, Yin Yang 0001, Zhanyu Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Adaptive Dynamic Programming for Control: A Survey and Recent AdvancesabstractThis article reviews the recent development of adaptive dynamic programming (ADP) with applications in control. First, its applications in optimal regulation are introduced, and some skilled and efficient algorithms are presented. Next, the use of ADP to solve game problems, mainly nonzero-sum game problems, is elaborated. It is followed by applications in large-scale systems. Note that although the functions presented in this article are based on continuous-time systems, various applications of ADP in discrete-time systems are also analyzed. Moreover, in each section, not only some existing techniques are discussed, but also possible directions for future work are pointed out. Finally, some overall prospects for the future are given, followed by conclusions of this article. Through a comprehensive and complete investigation of its applications in many existing fields, this article fully demonstrates that the ADP intelligent control method is promising in today's artificial intelligence era. Furthermore, it also plays a significant role in promoting economic and social development. Derong Liu 0001, Shan Xue 0004, Bo Zhao 0015, Biao Luo 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Adaptive synchronization of memristor-based neural networks with discontinuous activations
Yueheng Li, Biao Luo 0001, Derong Liu 0001, Zhanyu Yang, Yunli Zhu |
Neurocomputing | 3 |
| 2020 | Adaptive dynamic programming based event-triggered control for unknown continuous-time nonlinear systems with input constraints
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
Neurocomputing | 3 |
| 2020 | Deterministic policy gradient adaptive dynamic programming for model-free optimal control
Yongwei Zhang 0002, Bo Zhao 0015, Derong Liu 0001 |
Neurocomputing | 3 |
| 2020 | Event-triggered constrained control with DHP implementation for nonaffine discrete-time systems
Mingming Ha, Ding Wang 0001, Derong Liu 0001 |
Inf. Sci. | 3 |
| 2020 | Event-triggered adaptive dynamic programming for discrete-time multi-player games
Qinglai Wei, Derong Liu 0001 |
Inf. Sci. | 3 |
| 2020 | Improved value iteration for neural-network-based stochastic optimal control design
Mingming Liang, Ding Wang 0001, Derong Liu 0001 |
Neural Networks | 3 |
| 2020 | Integral reinforcement learning based event-triggered control with input saturation
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001 |
Neural Networks | 3 |
| 2020 | Continuous-Time Time-Varying Policy IterationabstractA novel policy iteration algorithm, called the continuous-time time-varying (CTTV) policy iteration algorithm, is presented in this paper to obtain the optimal control laws for infinite horizon CTTV nonlinear systems. The adaptive dynamic programming (ADP) technique is utilized to obtain the iterative control laws for the optimization of the performance index function. The properties of the CTTV policy iteration algorithm are analyzed. Monotonicity, convergence, and optimality of the iterative value function have been analyzed, and the iterative value function can be proven to monotonically converge to the optimal solution of the Hamilton-Jacobi-Bellman (HJB) equation. Furthermore, the iterative control law is guaranteed to be admissible to stabilize the nonlinear systems. In the implementation of the presented CTTV policy algorithm, the approximate iterative control laws and iterative value function are obtained by neural networks. Finally, the numerical results are given to verify the effectiveness of the presented method. Qinglai Wei, Zehua Liao, Zhanyu Yang, Benkai Li, Derong Liu 0001 |
IEEE Trans. Cybern. | 5 |
| 2020 | Event-Triggered Optimal Control With Performance Guarantees Using Adaptive Dynamic ProgrammingabstractThis paper studies the problem of event-triggered optimal control (ETOC) for continuous-time nonlinear systems and proposes a novel event-triggering condition that enables designing ETOC methods directly based on the solution of the Hamilton-Jacobi-Bellman (HJB) equation. We provide formal performance guarantees by proving a predetermined upper bound. Moreover, we also prove the existence of a lower bound for interexecution time. For implementation purposes, an adaptive dynamic programming (ADP) method is developed to realize the ETOC using a critic neural network (NN) to approximate the value function of the HJB equation. Subsequently, we prove that semiglobal uniform ultimate boundedness can be guaranteed for states and NN weight errors with the ADP-based ETOC. Simulation results demonstrate the effectiveness of the developed ADP-based ETOC method. Biao Luo 0001, Yin Yang 0001, Derong Liu 0001, Huai-Ning Wu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Relaxed Stability Criteria for Neural Networks With Time-Varying Delay Using Extended Secondary Delay Partitioning and Equivalent Reciprocal Convex Combination TechniquesabstractThis article investigates global asymptotic stability for neural networks (NNs) with time-varying delay, which is differentiable and uniformly bounded, and the delay derivative exists and is upper-bounded. First, we propose the extended secondary delay partitioning technique to construct the novel Lyapunov-Krasovskii functional, where both single-integral and double-integral state variables are considered, while the single-integral ones are only solved by the traditional secondary delay partitioning. Second, a novel free-weight matrix equality (FWME) is presented to resolve the reciprocal convex combination problem equivalently and directly without Schur complement, which eliminates the need of positive definite matrices, and is less conservative and restrictive compared with various improved reciprocal convex inequalities. Furthermore, by the present extended secondary delay partitioning, equivalent reciprocal convex combination technique, and Bessel-Legendre inequality, two different relaxed sufficient conditions ensuring global asymptotic stability for NNs are obtained, for time-varying delays, respectively, with unknown and known lower bounds of the delay derivative. Finally, two examples are given to illustrate the superiority and effectiveness of the presented method. Shenquan Wang, Wenchengyu Ji, Yulian Jiang, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Reinforcement Learning-Based Optimal Stabilization for Unknown Nonlinear Systems Subject to Inputs With Uncertain ConstraintsabstractThis article presents a novel reinforcement learning strategy that addresses an optimal stabilizing problem for unknown nonlinear systems subject to uncertain input constraints. The control algorithm is composed of two parts, i.e., online learning optimal control for the nominal system and feedforward neural networks (NNs) compensation for handling uncertain input constraints, which are considered as the saturation nonlinearities. Integrating the input-output data and recurrent NN, a Luenberger observer is established to approximate the unknown system dynamics. For nominal systems without input constraints, the online learning optimal control policy is derived by solving Hamilton-Jacobi-Bellman equation via a critic NN alone. By transforming the uncertain input constraints to saturation nonlinearities, the uncertain input constraints can be compensated by employing a feedforward NN compensator. The convergence of the closed-loop system is guaranteed to be uniformly ultimately bounded by using the Lyapunov stability analysis. Finally, the effectiveness of the developed stabilization scheme is illustrated by simulation studies. Bo Zhao 0015, Derong Liu 0001, Chaomin Luo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Event-Triggered Adaptive Critic Control Design for Discrete-Time Constrained Nonlinear SystemsabstractIn this paper, through event-triggered approach, the constrained near-optimal control problem for a class of nonlinear discrete-time systems is investigated and solved by heuristic dynamic programming (HDP) technique. The proposed method can reduce the amount of computation remarkably without deteriorating the system stability. In order to overcome the control constraints and reduce the computational burden, a nonquadratic performance index is introduced. Then, stability analysis of the event-triggered system with control constraints and an event-triggered constrained controller design algorithm are given. Three neural networks are used in the HDP scheme, which are designed to identify the unknown nonlinear system, approximate value function, and control law, respectively. In the model neural network, an effective method is developed to initialize its weights. Finally, two examples are included to demonstrate the present method. Mingming Ha, Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Model-Free H∞ Optimal Tracking Control of Constrained Nonlinear Systems via an Iterative Adaptive Learning AlgorithmabstractIn this paper, an H∞optimal tracking controller for completely unknown discrete-time nonlinear systems with control constraints is obtained by using an iterative adaptive learning algorithm. An augmented system is established by integrating the tracking error system and the reference trajectory. As an identifier of the unknown systems, a neural network (NN) is introduced with asymptotic stability of the estimation error. An action-disturbance-critic NN structure is proposed to implement the iterative dual heuristic programming algorithm with convergence guarantee of the costate function and the control policy. Simulation results and comparisons are provided to illustrate the superior performance of the designed optimal tracking controller. Jiaxu Hou, Ding Wang 0001, Derong Liu 0001, Yun Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Neuro-Optimal Control for Discrete Stochastic Processes via a Novel Policy Iteration AlgorithmabstractIn this paper, a novel policy iteration adaptive dynamic programming (ADP) algorithm is presented which is called “local policy iteration ADP algorithm” to obtain the optimal control for discrete stochastic processes. In the proposed local policy iteration ADP algorithm, the iterative decision rules are updated in a local space of the whole state space. Hence, we can significantly reduce the computational burden for the CPU in comparison with the conventional policy iteration algorithm. By analyzing the convergence properties of the proposed algorithm, it is shown that the iterative value functions are monotonically nonincreasing. Besides, the iterative value functions can converge to the optimum in a local policy space. In addition, this local policy space will be described in detail for the first time. Under a few weak constraints, it is also shown that the iterative value function will converge to the optimal performance index function of the global policy space. Finally, a simulation example is presented to validate the effectiveness of the developed method. Mingming Liang, Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Editorial Special Issue on Adaptive Dynamic Programming and Reinforcement LearningabstractThe past decade has witnessed a surge in research activities related to adaptive dynamic programming (ADP) and reinforcement learning (RL), particularly for control applications. Several books [item 1)–5) in the Appendix] and survey papers [item 6)–10) in the Appendix] have been published on the subject. Both ADP and RL provide approximate solutions to dynamic programming problems. In a 1995 article by Bartoet al.[item 11) in the Appendix], they introduced the so-called “adaptive real-time dynamic programming,” which was specifically to apply ADP for real-time control. Later, in 2002, Murrayet al.[item 12) in the Appendix] developed an ADP algorithm for optimal control of continuous-time affine nonlinear systems. On the other hand, the most famous algorithms in RL are the temporal difference algorithm [item 13) in the Appendix] and the Q-learning algorithm [item 14) and 15) in the Appendix]. Derong Liu 0001, Frank L. Lewis, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Event-Triggered Adaptive Dynamic Programming for Zero-Sum Game of Partially Unknown Continuous-Time Nonlinear SystemsabstractIn this paper, the zero-sum game problem is considered for partially unknown continuous-time nonlinear systems, and an event-triggered adaptive dynamic programming (ADP) method is developed to solve the problem. First, an identifier neural network (NN) and a critic NN are applied to approximate the drift system dynamics and the optimal value function, respectively. Subsequently, an event-triggered approach is developed based on ADP, which samples the states and updates the weights of NNs at the same time when the event-triggering condition is violated, such that the computational complexity is reduced. It is proved that the states and the error of NN weights are uniformly ultimately bounded. Finally, the effectiveness of the developed ADP-based event-triggered method is verified through simulation studies. Shan Xue 0004, Biao Luo 0001, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Adaptive Synchronization of Delayed Memristive Neural Networks With Unknown ParametersabstractIn this paper, the drive-response synchronization of the delayed memristive neural networks (MNNs) with unknown parameters is studied. With the realization of practical memristors, more and more researchers start to investigate MNNs, and their synchronization problem has became a hot topic. However, the majority of the existing works are based on the strict condition that the weights of MNNs are known and determined. When the parameters are unknown, the obtained results may be inapplicable. Thus, it is worthwhile to investigate the synchronization problem of the delayed MNNs with unknown parameters. Due to the parameter uncertainties of MNNs, a novel response system and an adaptive control method are proposed under different assumptions. The update laws for weights in the response system and the gains of adaptive controllers are developed to synchronize the proposed response system with the delayed MNNs. Furthermore, the proposed methods can be applied to various cases and the corresponding stability theories are established. Finally, the numerical simulations are conducted to verify the effectiveness of the developed methods. Zhanyu Yang, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Local Near-Optimal Control for Interconnected Systems with Time-Varying Delays
Qiuye Wu, Haowei Lin, Bo Zhao 0015, Derong Liu 0001 |
ICONIP (2) | 4 |
| 2019 | Pinning Control for Synchronization of Drive-Response Memristive Neural Networks with Nonidentical ParametersabstractIn this paper, the asymptotic synchronization for drive-response memristive neural networks(MNNs) with nonidentical parameters is investigated. Parameter inconformity is ubiquitous between drive and response systems due to environmental or internal influence. However, the majority of previous results were based on the well-matched MNNs. Thus, it is meaningful to study the synchronization problem of MNNs with nonidentical parameters. First, coefficient mismatches are dealt within the framework of set-valued maps and differential inclusions. Furthermore, in order to reduce the control cost, a pinning control strategy is adopted to drive two nonidentical MNNs to achieve asymptotic synchronization. And the sufficient stability conditions are given based on Lyapunov functional method. Finally, the effectiveness of proposed pinning controller is verified by a numerical example. Yueheng Li, Biao Luo 0001, Derong Liu 0001, Zhanyu Yang |
IJCNN | 3 |
| 2019 | Distributed Adaptive Dynamic Programming Algorithm for Office Energy Control with Multiple Batteries
Chao Li 0024, Bo Zhao 0015, Qinglai Wei, Derong Liu 0001 |
IJCNN | 5 |
| 2019 | Event-Triggered Adaptive Control for Discrete-Time Zero-Sum GamesabstractIn this paper, an event-triggered adaptive dynamic programming (ADP) method is developed for the discrete-time nonlinear two-player zero-sum games. First, an event-triggered ADP algorithm is presented to solve the Hamilton-Jacobi-Isaacs (HJI) equation. Then, a novel double event-triggered scheme is designed, the control inputs and the disturbance inputs will be updated only when the triggering conditions are satisfied. Therefore, the computational burden and the communication cost can be reduced. The algorithm is implemented by two neural networks, and the stability of the two-player system is proved. Finally, an example is employed to illustrate the effectiveness of the developed method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 3 |
| 2019 | An echo state network based approach to room classification of office buildings
Bo Zhao 0015, Chao Li 0024, Qinglai Wei, Derong Liu 0001 |
Neurocomputing | 5 |
| 2019 | Output Tracking Control Based on Adaptive Dynamic Programming With Multistep Policy EvaluationabstractIn this paper, the optimal output tracking control problem of discrete-time nonlinear systems is considered. First, the augmented system is derived and the tracking control problem is converted to the regulation problem with a discounted performance index, which relies on the solution of the Bellman equation. It is known that policy iteration and value iteration are two classical algorithms for solving the Bellman equation. Through analysis of the two algorithms, it is found that policy iteration converges fast while requires an initial admissible control policy, and value iteration avoids the requirement of an initial admissible control policy but converges slowly. To achieve the tradeoff between policy iteration and value iteration, the multistep heuristic dynamic programming (MsHDP) is proposed by using multistep policy evaluation scheme. The convergence of MsHDP algorithm is proved by demonstrating that it converges to the solution of the Bellman equation. Subsequently, neural network-based actor-critic structure is developed to implement the MsHDP algorithm. The effectiveness and advantages of the developed MsHDP method are validated through comparative simulation studies. Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Jiangjiang Liu 0004 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Event-Triggered Optimal Neuro-Controller Design With Reinforcement Learning for Unknown Nonlinear SystemsabstractThis paper develops an optimal control scheme for continuous-time unknown nonlinear systems using the event-triggering mechanism. Different from designing controllers using the time-triggering mechanism, the event-triggered controller is updated only when the system state deviates more than a certain threshold from a prescribed value. To obtain the event-triggered optimal controller, we develop an identifier-critic architecture under the framework of reinforcement learning. The identifier network, composed of a feedforward neural network (FNN), aims to derive the knowledge of unknown system dynamics, and the critic network, constituted of an FNN, intends to derive the event-triggered optimal controller. The identifier network is tuned via the combination of a standard back-propagation algorithm and an e-modification method, and the critic network is updated using a modification of the gradient descent method. By introducing an additional stability term to update the critic network, the initial admissible control is no longer required. Meanwhile, by using historical and instantaneous state data together, the persistence of excitation condition is relaxed. A stability analysis of the closed-loop system is provided based on the Lyapunov method. The effectiveness of the proposed designs is illustrated through simulations of a nonlinear example and a single link robot arm system. Xiong Yang 0001, Haibo He, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | Event-Triggered Adaptive Dynamic Programming for Continuous-Time Nonlinear Two-Player Zero-Sum Game
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
ICONIP (7) | 3 |
| 2018 | Local Tracking Control for Unknown Interconnected Systems via Neuro-Dynamic Programming
Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Ding Wang 0001, Yancai Xu, Qinglai Wei |
ICONIP (7) | 2 |
| 2018 | Robust synchronization of memristive neural networks with strong mismatch characteristics via pinning control
Yueheng Li, Biao Luo 0001, Derong Liu 0001, Zhanyu Yang |
Neurocomputing | 3 |
| 2018 | Neural robust stabilization via event-triggering mechanism and adaptive learning technique
Ding Wang 0001, Derong Liu 0001 |
Neural Networks | 2 |
| 2018 | Neural network robust tracking control with adaptive critic framework for uncertain nonlinear systems
Ding Wang 0001, Derong Liu 0001, Yun Zhang 0001, Hongyi Li 0001 |
Neural Networks | 2 |
| 2018 | Distributed algorithm for dissensus of a class of networked multiagent systems using output information
Hongwen Ma, Derong Liu 0001, Ding Wang 0001, Xiong Yang 0001, Hongliang Li 0002 |
Soft Comput. | 2 |
| 2018 | Adaptive Q-Learning for Data-Based Optimal Output Regulation With Experience ReplayabstractIn this paper, the data-based optimal output regulation problem of discrete-time systems is investigated. An off-policy adaptive -learning (QL) method is developed by using real system data without requiring the knowledge of system dynamics and the mathematical model of utility function. By introducing the -function, an off-policy adaptive QL algorithm is developed to learn the optimal -function. An adaptive parameter in the policy evaluation is used to achieve tradeoff between the current and future -functions. The convergence of adaptive QL algorithm is proved and the influence of the adaptive parameter is analyzed. To realize the adaptive QL algorithm with real system data, the actor-critic neural network (NN) structure is developed. The least-squares scheme and the batch gradient descent method are developed to update the critic and actor NN weights, respectively. The experience replay technique is employed in the learning process, which leads to simple and convenient implementation of the adaptive QL method. Finally, the effectiveness of the developed adaptive QL method is verified through numerical simulations. Biao Luo 0001, Yin Yang 0001, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2018 | Intelligent Optimal Control With Critic Learning for a Nonlinear Overhead Crane SystemabstractIn this paper, for achieving the discounted optimal feedback stabilization of a nonlinear overhead crane system, we establish an intelligent control strategy to obtain the solution of the corresponding Hamilton-Jacobi-Bellman equation. Specifically, neural networks are employed to serve as a necessary component to the control system, which exhibits strong online learning ability. A novel updating rule compared to the traditional adaptive critic algorithms is developed, which eliminates the requirement of the initial stabilizing controller and brings in unique advantages to the adaptive critic control design. Stability analysis of the closed-loop system based on the well-known Lyapunov approach and experimental simulation considering the nonlinear overhead dynamics with different case studies are performed to verify the effectiveness of the present control method both in theory and applications. Ding Wang 0001, Haibo He, Derong Liu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Manifold Regularized Reinforcement LearningabstractThis paper introduces a novel manifold regularized reinforcement learning scheme for continuous Markov decision processes. Smooth feature representations for value function approximation can be automatically learned using the unsupervised manifold regularization method. The learned features are data-driven, and can be adapted to the geometry of the state space. Furthermore, the scheme provides a direct basis representation extension for novel samples during policy learning and control. The performance of the proposed scheme is evaluated on two benchmark control tasks, i.e., the inverted pendulum and the energy storage problem. Simulation results illustrate the concepts of the proposed scheme and show that it can obtain excellent performance. Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Adaptive Constrained Optimal Control Design for Data-Based Nonlinear Discrete-Time Systems With Critic-Only StructureabstractReinforcement learning has proved to be a powerful tool to solve optimal control problems over the past few years. However, the data-based constrained optimal control problem of nonaffine nonlinear discrete-time systems has rarely been studied yet. To solve this problem, an adaptive optimal control approach is developed by using the value iteration-based Q-learning (VIQL) with the critic-only structure. Most of the existing constrained control methods require the use of a certain performance index and only suit for linear or affine nonlinear systems, which is unreasonable in practice. To overcome this problem, the system transformation is first introduced with the general performance index. Then, the constrained optimal control problem is converted to an unconstrained optimal control problem. By introducing the action-state value function, i.e., Q-function, the VIQL algorithm is proposed to learn the optimal Q-function of the data-based unconstrained optimal control problem. The convergence results of the VIQL algorithm are established with an easy-to-realize initial condition . To implement the VIQL algorithm, the critic-only structure is developed, where only one neural network is required to approximate the Q-function. The converged Q-function obtained from the critic-only VIQL method is employed to design the adaptive constrained optimal controller based on the gradient descent scheme. Finally, the effectiveness of the developed adaptive control method is tested on three examples with computer simulation. Biao Luo 0001, Derong Liu 0001, Huai-Ning Wu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Learning and Guaranteed Cost Control With Event-Based Adaptive Critic ImplementationabstractThis paper focuses on the event-triggered guaranteed cost control design of nonlinear systems via a self-learning technique. In brief, an event-based guaranteed cost control strategy of nonlinear systems subjects to matched uncertainties is developed, thereby balancing the performance of guaranteed cost and the actuality of limited communication resource. The original control design is transformed into an optimal control problem with an event-based mechanism, where the relationship of guaranteed cost performance compared to the time-based formulation is discussed. A critic neural network is constructed for implementing the event-based optimal control design with stability guarantee. Simulation experiments are carried out to verify the theoretical results in detail. Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Neural Network Learning and Robust Stabilization of Nonlinear Systems With Dynamic UncertaintiesabstractDue to the existence of dynamical uncertainties, it is important to pay attention to the robustness of nonlinear control systems, especially when designing adaptive critic control strategies. In this paper, based on the neural network learning component, the robust stabilization scheme of nonlinear systems with general uncertainties is developed. Through system transformation and employing adaptive critic technique, the approximate optimal controller of the nominal plant can be applied to accomplish robust stabilization for the original uncertain dynamics. The neural network weight vector is very convenient to initialize by virtue of the improved critic learning formulation. Under the action of the approximate optimal control law, the stability issues for the closed-loop form of nominal and uncertain plants are analyzed, respectively. Simulation illustrations via a typical nonlinear system and a practical power system are included to verify the control performance. Ding Wang 0001, Derong Liu 0001, Chaoxu Mu, Yun Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | On Mixed Data and Event Driven Design for Adaptive-Critic-Based Nonlinear H∞ ControlabstractIn this paper, based on the adaptive critic learning technique, the control for a class of unknown nonlinear dynamic systems is investigated by adopting a mixed data and event driven design approach. The nonlinear control problem is formulated as a two-player zero-sum differential game and the adaptive critic method is employed to cope with the data-based optimization. The novelty lies in that the data driven learning identifier is combined with the event driven design formulation, in order to develop the adaptive critic controller, thereby accomplishing the nonlinear control. The event driven optimal control law and the time driven worst case disturbance law are approximated by constructing and tuning a critic neural network. Applying the event driven feedback control, the closed-loop system is built with stability analysis. Simulation studies are conducted to verify the theoretical results and illustrate the control performance. It is significant to observe that the present research provides a new avenue of integrating data-based control and event-triggering mechanism into establishing advanced adaptive critic systems. Ding Wang 0001, Chaoxu Mu, Derong Liu 0001, Hongwen Ma |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Adaptive Dynamic Programming for Discrete-Time Zero-Sum GamesabstractIn this paper, a novel adaptive dynamic programming (ADP) algorithm, called "iterative zero-sum ADP algorithm," is developed to solve infinite-horizon discrete-time two-player zero-sum games of nonlinear systems. The present iterative zero-sum ADP algorithm permits arbitrary positive semidefinite functions to initialize the upper and lower iterations. A novel convergence analysis is developed to guarantee the upper and lower iterative value functions to converge to the upper and lower optimums, respectively. When the saddle-point equilibrium exists, it is emphasized that both the upper and lower iterative value functions are proved to converge to the optimal solution of the zero-sum game, where the existence criteria of the saddle-point equilibrium are not required. If the saddle-point equilibrium does not exist, the upper and lower optimal performance index functions are obtained, respectively, where the upper and lower performance index functions are proved to be not equivalent. Finally, simulation results and comparisons are shown to illustrate the performance of the present method. Qinglai Wei, Derong Liu 0001, Qiao Lin 0003, Ruizhuo Song |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Special Issue on Deep Reinforcement Learning and Adaptive Dynamic ProgrammingabstractThe sixteen papers in this special section focus on deep reinforcement learning and adaptive dynamic programming (deep RL/ADP). Deep RL is able to output control signal directly based on input images, which incorporates both the advantages of the perception of deep learning (DL) and the decision making of RL or adaptive dynamic programming (ADP). This mechanism makes the artificial intelligence much closer to human thinking modes. Deep RL/ADP has achieved remarkable success in terms of theory and applications since it was proposed. Successful applications cover video games, Go, robotics, smart driving, healthcare, and so on. However, it is still an open problem to perform the theoretical analysis on deep RL/ADP, e.g., the convergence, stability, and optimality analyses. The learning efficiency needs to be improved by proposing new algorithms or combined with other methods. More practical demonstrations are encouraged to be presented. Therefore, the aim of this special issue is to call for the most advanced research and state-of-the-art works in the field of deep RL/ADP. Dongbin Zhao, Derong Liu 0001, Frank L. Lewis, José C. Príncipe, Stefano Squartini |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Data-Based Optimal Control for Weakly Coupled Nonlinear Systems Using Policy IterationabstractIn this paper, a data-based online learning algorithm is established to solve the optimal control problem for weakly coupled continuous-time nonlinear systems with completely unknown dynamics. Using the weak coupling theory, we reformulate the original problem into three reduced-order optimal control problems. We establish an online model-free integral policy iteration algorithm to solve the decoupled optimal control problems without system dynamics. To implement the data-based online learning algorithm, the actor-critic technique based on neural networks and the least squares method are used. Two simulation examples are given to verify the effectiveness of the developed algorithm. Chao Li 0024, Derong Liu 0001, Ding Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | Discrete-Time Local Value Iteration Adaptive Dynamic Programming: Convergence AnalysisabstractIn this paper, convergence properties are established for the newly developed discrete-time local value iteration adaptive dynamic programming (ADP) algorithm. The present local iterative ADP algorithm permits an arbitrary positive semidefinite function to initialize the algorithm. Employing a state-dependent learning rate function, for the first time, the iterative value function and iterative control law can be updated in a subset of the state space instead of the whole state space, which effectively relaxes the computational burden. A new analysis method for the convergence property is developed to prove that the iterative value functions will converge to the optimum under some mild constraints. Monotonicity of the local value iteration ADP algorithm is presented, which shows that under some special conditions of the initial value function and the learning rate function, the iterative value function can monotonically converge to the optimum. Finally, three simulation examples and comparisons are given to illustrate the performance of the developed algorithm. Qinglai Wei, Frank L. Lewis, Derong Liu 0001, Ruizhuo Song, Hanquan Lin |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | Decentralized Control for Large-Scale Nonlinear Systems With Unknown Mismatched Interconnections via Policy IterationabstractIn this paper, the decentralized control problem is solved based on a policy iteration algorithm for large-scale nonlinear systems with unknown mismatched interconnections. The unknown interconnection is approximated by a neural network with local states of isolated subsystem and substituted reference states of coupled subsystems. Then, an adaptive estimation term is utilized to construct the improved local performance index function that reflects the substitution error. Hereafter, the closed-loop large-scale nonlinear system is guaranteed to be ultimately uniformly bounded by the implementation of a set of developed decentralized optimal control policies. Two simulation examples are given to verify the effectiveness of the presented scheme. The significant contribution of this scheme lies in that it removes the common assumptions on satisfying matching condition and upper boundedness of interconnections, when designing the decentralized optimal control for large-scale nonlinear systems. Bo Zhao 0015, Ding Wang 0001, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | Partially-Directed-Topology-Based Consensus Control for Linear Multi-agent Systems
Chunping Shi, Qinglai Wei, Derong Liu 0001 |
ICONIP (6) | 3 |
| 2017 | An Event-Triggered Heuristic Dynamic Programming Algorithm for Discrete-Time Nonlinear Systems
Qinglai Wei, Derong Liu 0001 |
ICONIP (1) | 3 |
| 2017 | Synchronization of Memristor-Based Time-Delayed Neural Networks via Pinning Control
Zhanyu Yang, Biao Luo 0001, Derong Liu 0001 |
ICONIP (3) | 3 |
| 2017 | Neuro-control of Nonlinear Systems with Unknown Input Constraints
Bo Zhao 0015, Derong Liu 0001 |
ICONIP (1) | 3 |
| 2017 | Local Policy Iteration Adaptive Dynamic Programming for Discrete-Time Nonlinear Systems
Qinglai Wei, Yancai Xu, Qiao Lin 0003, Derong Liu 0001, Ruizhuo Song |
ISNN (2) | 4 |
| 2017 | Fault detection and control co-design for discrete-time delayed fuzzy networked control systems subject to quantization and multiple packet dropouts
Shenquan Wang, Yulian Jiang, Derong Liu 0001 |
Fuzzy Sets Syst. | 4 |
| 2017 | Bounded robust control design for uncertain nonlinear systems using single-network adaptive dynamic programming
Yuzhu Huang, Ding Wang 0001, Derong Liu 0001 |
Neurocomputing | 3 |
| 2017 | Self-tuned local feedback gain based decentralized fault tolerant control for a class of large-scale nonlinear systems
Bo Zhao 0015, Derong Liu 0001 |
Neurocomputing | 3 |
| 2017 | Multi-step heuristic dynamic programming for optimal control of nonlinear discrete-time systems
Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Xiong Yang 0001, Hongwen Ma |
Inf. Sci. | 2 |
| 2017 | Observer based adaptive dynamic programming for fault tolerant control of a class of nonlinear systems
Bo Zhao 0015, Derong Liu 0001 |
Inf. Sci. | 2 |
| 2017 | Pinning synchronization of memristor-based neural networks with time-varying delays
Zhanyu Yang, Biao Luo 0001, Derong Liu 0001, Yueheng Li |
Neural Networks | 3 |
| 2017 | Optimization of electricity consumption in office buildings based on adaptive dynamic programming
Qinglai Wei, Derong Liu 0001 |
Soft Comput. | 3 |
| 2017 | Policy Gradient Adaptive Dynamic Programming for Data-Based Optimal ControlabstractThe model-free optimal control problem of general discrete-time nonlinear systems is considered in this paper, and a data-based policy gradient adaptive dynamic programming (PGADP) algorithm is developed to design an adaptive optimal controller method. By using offline and online data rather than the mathematical system model, the PGADP algorithm improves control policy with a gradient descent scheme. The convergence of the PGADP algorithm is proved by demonstrating that the constructed Q -function sequence converges to the optimal Q -function. Based on the PGADP algorithm, the adaptive control method is developed with an actor-critic structure and the method of weighted residuals. Its convergence properties are analyzed, where the approximate Q -function converges to its optimum. Computer simulation results demonstrate the effectiveness of the PGADP-based adaptive control method. Biao Luo 0001, Derong Liu 0001, Huai-Ning Wu, Ding Wang 0001, Frank L. Lewis |
IEEE Trans. Cybern. | 2 |
| 2017 | Improving the Critic Learning for Event-Based Nonlinear H∞ Control Designabstractcontrol problem is regarded as a two-player zero-sum game and the adaptive critic mechanism is used to achieve the minimax optimization under event-based environment. Then, based on an improved updating rule, the event-based optimal control law and the time-based worst-case disturbance law are obtained approximately by training a single critic neural network. The initial stabilizing control is no longer required during the implementation process of the new algorithm. Next, the closed-loop system is formulated as an impulsive model and its stability issue is handled by incorporating the improved learning criterion. The infamous Zeno behavior of the present event-based design is also avoided through theoretical analysis on the lower bound of the minimal intersample time. Finally, the applications to an aircraft dynamics and a robot arm plant are carried out to verify the efficient performance of the present novel design method. Ding Wang 0001, Haibo He, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Adaptive Critic Nonlinear Robust Control: A SurveyabstractAdaptive dynamic programming (ADP) and reinforcement learning are quite relevant to each other when performing intelligent optimization. They are both regarded as promising methods involving important components of evaluation and improvement, at the background of information technology, such as artificial intelligence, big data, and deep learning. Although great progresses have been achieved and surveyed when addressing nonlinear optimal control problems, the research on robustness of ADP-based control strategies under uncertain environment has not been fully summarized. Hence, this survey reviews the recent main results of adaptive-critic-based robust control design of continuous-time nonlinear systems. The ADP-based nonlinear optimal regulation is reviewed, followed by robust stabilization of nonlinear systems with matched uncertainties, guaranteed cost control design of unmatched plants, and decentralized stabilization of interconnected systems. Additionally, further comprehensive discussions are presented, including event-based robust control design, improvement of the critic learning rule, nonlinear H∞control design, and several notes on future perspectives. By applying the ADP-based optimal and robust control methods to a practical power system and an overhead crane plant, two typical examples are provided to verify the effectiveness of theoretical results. Overall, this survey is beneficial to promote the development of adaptive critic control methods with robustness guarantee and the construction of higher level intelligent systems. Ding Wang 0001, Haibo He, Derong Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Discrete-Time Optimal Control via Local Policy Iteration Adaptive Dynamic ProgrammingabstractIn this paper, a discrete-time optimal control scheme is developed via a novel local policy iteration adaptive dynamic programming algorithm. In the discrete-time local policy iteration algorithm, the iterative value function and iterative control law can be updated in a subset of the state space, where the computational burden is relaxed compared with the traditional policy iteration algorithm. Convergence properties of the local policy iteration algorithm are presented to show that the iterative value function is monotonically nonincreasing and converges to the optimum under some mild conditions. The admissibility of the iterative control law is proven, which shows that the control system can be stabilized under any of the iterative control laws, even if the iterative control law is updated in a subset of the state space. Finally, two simulation examples are given to illustrate the performance of the developed method. Qinglai Wei, Derong Liu 0001, Qiao Lin 0003, Ruizhuo Song |
IEEE Trans. Cybern. | 2 |
| 2017 | Discrete-Time Local Value Iteration Adaptive Dynamic Programming: Admissibility and Termination AnalysisabstractIn this paper, a novel local value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon optimal control problems for discrete-time nonlinear systems. The focuses of this paper are to study admissibility properties and the termination criteria of discrete-time local value iteration ADP algorithms. In the discrete-time local value iteration ADP algorithm, the iterative value functions and the iterative control laws are both updated in a given subset of the state space in each iteration, instead of the whole state space. For the first time, admissibility properties of iterative control laws are analyzed for the local value iteration ADP algorithm. New termination criteria are established, which terminate the iterative local ADP algorithm with an admissible approximate optimal control law. Finally, simulation results are given to illustrate the performance of the developed algorithm.In this paper, a novel local value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon optimal control problems for discrete-time nonlinear systems. The focuses of this paper are to study admissibility properties and the termination criteria of discrete-time local value iteration ADP algorithms. In the discrete-time local value iteration ADP algorithm, the iterative value functions and the iterative control laws are both updated in a given subset of the state space in each iteration, instead of the whole state space. For the first time, admissibility properties of iterative control laws are analyzed for the local value iteration ADP algorithm. New termination criteria are established, which terminate the iterative local ADP algorithm with an admissible approximate optimal control law. Finally, simulation results are given to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Qiao Lin 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Event-Driven Adaptive Robust Control of Nonlinear Systems With Uncertainties Through NDP StrategyabstractIn this paper, we construct an event-driven adaptive robust control approach for continuous-time uncertain nonlinear systems through a neural dynamic programming (NDP) strategy. Through system transformation and theoretical analysis, the robustness of the original uncertain system can be achieved by designing an event-driven optimal controller with respect to the nominal system under a suitable triggering condition. In addition, it is also observed that the event-driven controller has a certain degree of gain margin. Then, the NDP technique is employed to perform the main controller design task, followed by the uniform ultimate boundedness stability proof with the feedback action of the event-driven adaptive control law. The comparative effect of the present control strategy is also illustrated via two simulation examples. The established method provides a new avenue of combining adaptive dynamic programming-based self-learning control, event-triggered adaptive control, and robust control, to investigate the nonlinear adaptive robust feedback design under uncertain environment. Ding Wang 0001, Chaoxu Mu, Haibo He, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | Event-Based Constrained Robust Control of Affine Systems Incorporating an Adaptive Critic MechanismabstractThis paper focuses on establishing an event-based constrained robust control strategy for a class of continuous-time affine nonlinear systems by incorporating the adaptive critic mechanism (ACM). The main objective is to integrate the event-based framework, the constrained optimal control method, and the neural network learning ability, thereby achieving the nonlinear robust state feedback of input-constrained nonlinear systems under event-based environment. Through theoretical analysis, it is shown that the nonlinear robust control law subject to input limitations can be obtained by designing an event-based constrained optimal controller with respect to the nominal system. Then, the ACM is adopted to facilitate the constrained optimal control implementation, where a critic neural network is constructed to serve as the learning approximator. The system stability issue is proved by employing the Lyapunov theory and the constrained robust control performance is illustrated through simulation experiments of several dynamical plants. Ding Wang 0001, Chaoxu Mu, Xiong Yang 0001, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | Error Bound Analysis of Q-Function for Discounted Optimal Control Problems With Policy IterationabstractIn this paper, we present error bound analysis of the Q-function for the action-dependent adaptive dynamic programming for solving discounted optimal control problems of unknown discrete-time nonlinear systems. The convergence of Q-functions derived by a policy iteration algorithm under ideal conditions is given. Considering the approximated errors of the Q-function and control policy in the policy evaluation step and policy improvement step, we establish error bounds of approximate Q-functions in each iteration. With the given boundedness conditions, the approximate Q-function will converge to a finite neighborhood of the optimal Q-function. To implement the presented algorithm, two three-layer neural networks are employed to approximate the Q-function and the control policy, respectively. Finally, a simulation example is utilized to verify the validity of the presented algorithm. Ding Wang 0001, Hongliang Li 0002, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Improvement of reliable H∞ control for discrete-time fuzzy systems with time-varying delaysabstractThis paper is concerned with the reliable H∞control problem for discrete-time Takagi-Sugeno (T-S) fuzzy systems with time-varying delays and stochastic actuator faults based on a novel summation inequality. A discrete-time homogeneous Markov chain is used to represent the stochastic behavior of actuator faults. By employing a fuzzy basis-dependent Lyapunov functional and the novel summation inequality, an improved sufficient condition is established to ensure that the resultant closed-loop system is stochastically stable with an H∞performance index. Meanwhile, the solvability condition for the reliable H∞control is also established, by which the reliable H∞fuzzy controller can be solved from linear matrix inequalities. A numerical example is provided to demonstrate the effectiveness of the present approach. Shenquan Wang, Yulian Jiang, Derong Liu 0001 |
FUZZ-IEEE | 4 |
| 2016 | Data-Based Optimal Tracking Control of Nonaffine Nonlinear Discrete-Time Systems
Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Chao Li 0024 |
ICONIP (4) | 2 |
| 2016 | Neural Dynamic Programming for Event-Based Nonlinear Adaptive Robust Stabilization
Ding Wang 0001, Hongwen Ma, Derong Liu 0001, Huidong Wang |
ICONIP (1) | 3 |
| 2016 | Optimal Constrained Neuro-Dynamic Programming Based Self-learning Battery Management in Microgrids
Qinglai Wei, Derong Liu 0001 |
ICONIP (3) | 2 |
| 2016 | Decentralized Stabilization for Nonlinear Systems with Unknown Mismatched Interconnections
Bo Zhao 0015, Ding Wang 0001, Derong Liu 0001 |
ICONIP (3) | 4 |
| 2016 | Neural-network-based robust optimal control of uncertain nonlinear systems using model-free policy iteration algorithmabstractIn this paper, we establish a robust optimal control law for a class of continuous-time uncertain nonlinear systems by using a neural-network-based model-free policy iteration approach. The robust control law of the original uncertain nonlinear system is derived by adding a feedback gain to the optimal control law of the nominal system. It is proven that this robust control law can achieve optimality under a specified cost function. Then, the neural-network-based model-free policy iteration algorithm is developed to solve the Hamilton-Jacobi-Bellman equation corresponding to the nominal system without system dynamics. The actor-critic technique and the least squares implementation method are used to obtain the optimal control policy of the nominal system. A numerical simulation is given to verify the applicability of the present robust optimal control scheme. Chao Li 0024, Ding Wang 0001, Derong Liu 0001 |
IJCNN | 3 |
| 2016 | Distributed control of second-order nonlinear time-delayed multiagent systems with disturbance using neural networksabstractIn this paper, a class of second-order nonlinear time-delayed multiagent systems with disturbance is investigated. In order to improve the adaptivity, neural networks are used to learn the unknown dynamics. Then, by utilizing Lyapunov-Krasovskii functional, time delays can be eliminated. Moreover, a robustifying term is introduced to constrain external disturbance. With divide-and-conquer idea, the distributed controller is divided into five different parts to make the multiagent systems reach consensus. To circumvent the singularity induced by the time-delay elimination part, a σ-function is developed. Finally, the simulation results demonstrate the validity of the distributed controller. Hongwen Ma, Derong Liu 0001, Ding Wang 0001 |
IJCNN | 2 |
| 2016 | An adaptive dynamic programming based method for optimization of electricity consumption in office buildingsabstractIn this paper, an adaptive dynamic programming (ADP) based method is developed to optimize electricity consumption of rooms in office buildings through optimal battery management. Rooms in office buildings are generally divided into office room, computer room, storage room, meeting room, etc., each of which has different characteristics of electricity consumption, as divided into electricity consumption from sockets, lights and air-conditioners in this paper. The developed method based on ADP is elaborated, and different optimization strategies of electricity consumption in different categories of rooms are proposed in accordance with the developed method. Finally, a detailed case study on an office building is given to show the practical effect of the developed method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 3 |
| 2016 | Optimal self-learning control scheme for discrete-time nonlinear systems using local value iterationabstractIn this paper, an optimal self-learning control scheme for discrete-time nonlinear systems is developed using a new local value iteration based adaptive dynamic programming (ADP) algorithm. The developed local value iteration algorithm permits an arbitrary positive semi-definite function to initialize the algorithm. In the developed local value iteration algorithm, the iterative value function and iterative control law are updated by a subset of the state space. A new analysis method of the convergence property is presented to show that the iterative value functions will converge to the optimum. The convergence criterion for the local value iteration algorithm is presented. A simulation example is given to demonstrate the validity of the present optimal control scheme. Qinglai Wei, Derong Liu 0001 |
IJCNN | 2 |
| 2016 | Adaptive dynamic programming based fault compensation control for nonlinear systems with actuator failuresabstractThis paper develops a novel fault compensation control scheme based on adaptive dynamic programming for nonlinear systems with actuator failures. The control scheme consists of a policy iteration algorithm and a fault compensation. For fault-free dynamic models, the Hamilton-Jacobi-Bellman equation is solved by policy iteration algorithm via constructing a critic neural network, and then the approximate optimal control policy can be derived directly. On the other hand, the online fault compensation is achieved without the fault detection and isolation mechanism by reconstructing the actuator failure. The closed-loop system is guaranteed to be asymptotically stable based on Lyapunov stability theorem. Two simulation examples are given to demonstrate the effectiveness of the present fault compensation control scheme. Bo Zhao 0015, Derong Liu 0001 |
IJCNN | 2 |
| 2016 | Discrete-Time Two-Player Zero-Sum Games for Nonlinear Systems Using Iterative Adaptive Dynamic Programming
Qinglai Wei, Derong Liu 0001 |
ISNN | 2 |
| 2016 | Energy consumption prediction of office buildings based on echo state networks
Derong Liu 0001, Qinglai Wei |
Neurocomputing | 2 |
| 2016 | Decentralized guaranteed cost control of interconnected systems with uncertainties: A learning-based optimal control strategy
Ding Wang 0001, Derong Liu 0001, Chaoxu Mu, Hongwen Ma |
Neurocomputing | 2 |
| 2016 | Distributed control algorithm for bipartite consensus of the nonlinear time-delayed multi-agent systems with neural networks
Ding Wang 0001, Hongwen Ma, Derong Liu 0001 |
Neurocomputing | 3 |
| 2016 | Event-based input-constrained nonlinear H∞ state feedback with adaptive critic and neural implementation
Ding Wang 0001, Chaoxu Mu, Derong Liu 0001 |
Neurocomputing | 4 |
| 2016 | Data-driven controller design for general MIMO nonlinear systems via virtual reference feedback tuning and neural networks
Derong Liu 0001, Ding Wang 0001, Hongwen Ma |
Neurocomputing | 2 |
| 2016 | Guaranteed cost neural tracking control for a class of uncertain nonlinear systems using adaptive dynamic programming
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei, Ding Wang 0001 |
Neurocomputing | 2 |
| 2016 | Data-based robust optimal control of continuous-time affine nonlinear systems with matched uncertainties
Ding Wang 0001, Chao Li 0024, Derong Liu 0001, Chaoxu Mu |
Inf. Sci. | 3 |
| 2016 | Data-based robust adaptive control for a class of unknown nonlinear constrained-input systems via integral reinforcement learning
Xiong Yang 0001, Derong Liu 0001, Biao Luo 0001, Chao Li 0024 |
Inf. Sci. | 2 |
| 2016 | Online approximate solution of HJI equation for unknown constrained-input nonlinear continuous-time systems
Xiong Yang 0001, Derong Liu 0001, Hongwen Ma, Yancai Xu |
Inf. Sci. | 2 |
| 2016 | A neural-network-based online optimal control approach for nonlinear robust decentralized stabilization
Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Hongwen Ma, Chao Li 0024 |
Soft Comput. | 2 |
| 2016 | Neuro-optimal tracking control for a class of discrete-time nonlinear systems via generalized value iteration adaptive dynamic programming approach
Qinglai Wei, Derong Liu 0001, Yancai Xu |
Soft Comput. | 2 |
| 2016 | Value Iteration Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear SystemsabstractIn this paper, a value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon undiscounted optimal control problems for discrete-time nonlinear systems. The present value iteration ADP algorithm permits an arbitrary positive semi-definite function to initialize the algorithm. A novel convergence analysis is developed to guarantee that the iterative value function converges to the optimal performance index function. Initialized by different initial functions, it is proven that the iterative value function will be monotonically nonincreasing, monotonically nondecreasing, or nonmonotonic and will converge to the optimum. In this paper, for the first time, the admissibility properties of the iterative control laws are developed for value iteration algorithms. It is emphasized that new termination criteria are established to guarantee the effectiveness of the iterative control laws. Neural networks are used to approximate the iterative value function and compute the iterative control law, respectively, for facilitating the implementation of the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the present method. Qinglai Wei, Derong Liu 0001, Hanquan Lin |
IEEE Trans. Cybern. | 2 |
| 2016 | Model-Free Optimal Tracking Control via Critic-Only Q-LearningabstractModel-free control is an important and promising topic in control fields, which has attracted extensive attention in the past few years. In this paper, we aim to solve the model-free optimal tracking control problem of nonaffine nonlinear discrete-time systems. A critic-only Q-learning (CoQL) method is developed, which learns the optimal tracking control from real system data, and thus avoids solving the tracking Hamilton-Jacobi-Bellman equation. First, the Q-learning algorithm is proposed based on the augmented system, and its convergence is established. Using only one neural network for approximating the Q-function, the CoQL method is developed to implement the Q-learning algorithm. Furthermore, the convergence of the CoQL method is proved with the consideration of neural network approximation error. With the convergent Q-function obtained from the CoQL method, the adaptive optimal tracking control is designed based on the gradient descent scheme. Finally, the effectiveness of the developed CoQL method is demonstrated through simulation studies. The developed CoQL method learns with off-policy data and implements with a critic-only structure, thus it is easy to realize and overcome the inadequate exploration problem. Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Ding Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Neural-Network-Based Distributed Adaptive Robust Control for a Class of Nonlinear Multiagent Systems With Time Delays and External NoisesabstractA class of nonlinear multiagent systems with time delays and external noises is investigated, and a distributed adaptive robust control protocol is developed. It is the first time for a class of multiagent systems to take both time delays and external noises into consideration. By virtue of Lyapunov-Krasovskii functional and Young's inequality, the effects of time delay can be eliminated. Then, to exclude external noises, a robustifying term is introduced to eliminate the negative effects of these noises. Moreover, neural networks are utilized to learn the unknown nonlinear terms to adapt to the complex external environment. Finally, a numerical simulation is conducted to validate the effectiveness of our distributed control protocol. Hongwen Ma, Zhuo Wang 0003, Ding Wang 0001, Derong Liu 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | An Approximate Optimal Control Approach for Robust Stabilization of a Class of Discrete-Time Nonlinear Systems With UncertaintiesabstractIn this correspondence paper, the robust stabilization of a class of discrete-time nonlinear systems with uncertainties is investigated by using an approximate optimal control approach. The robust control problem is transformed into an optimal control problem under some proper restrictions on the bound of the uncertainties. For the purpose of dealing with the transformed optimal control, the discrete-time generalized Hamilton-Jacobi-Bellman equation is introduced and then solved using the successive approximation method with neural network implementation. In addition, a numerical simulation is included to illustrate the effectiveness of the robust control strategy. Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Biao Luo 0001, Hongwen Ma |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2016 | Data-Based Adaptive Critic Designs for Nonlinear Robust Optimal Control With Uncertain DynamicsabstractIn this paper, the infinite-horizon robust optimal control problem for a class of continuous-time uncertain nonlinear systems is investigated by using data-based adaptive critic designs. The neural network identification scheme is combined with the traditional adaptive critic technique, in order to design the nonlinear robust optimal control under uncertain environment. First, the robust optimal controller of the original uncertain system with a specified cost function is established by adding a feedback gain to the optimal controller of the nominal system. Then, a neural network identifier is employed to reconstruct the unknown dynamics of the nominal system with stability analysis. Hence, the data-based adaptive critic designs can be developed to solve the Hamilton-Jacobi-Bellman equation corresponding to the transformed optimal control problem. The uniform ultimate boundedness of the closed-loop system is also proved by using the Lyapunov approach. Finally, two simulation examples are presented to illustrate the effectiveness of the developed control strategy. Ding Wang 0001, Derong Liu 0001, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2015 | Distributed Control for Nonlinear Time-Delayed Multi-Agent Systems with Connectivity Preservation Using Neural Networks
Hongwen Ma, Derong Liu 0001, Ding Wang 0001 |
ICONIP (3) | 2 |
| 2015 | Robust Tracking Control of Uncertain Nonlinear Systems Using Adaptive Dynamic Programming
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei |
ICONIP (3) | 2 |
| 2015 | Approximate policy iteration with unsupervised feature learning based on manifold regularizationabstractIn this paper, we develop a novel approximate policy iteration reinforcement learning algorithm with unsupervised feature learning based on manifold regularization. The proposed algorithm can automatically learn data-driven smooth basis representations for value function approximation, which can preserve the intrinsic geometry of the state space of Markov decision processes. Moreover, it can provide a direct basis extension for new samples in both policy learning and policy control processes. We evaluate the effectiveness and efficiency of the proposed algorithm on the inverted pendulum task. Simulation results show that this algorithm can learn smooth basis representations and excellent control policies. Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001 |
IJCNN | 2 |
| 2015 | Data-driven virtual reference controller design for high-order nonlinear systems via neural networkabstractThis paper is concerned with data-driven methods for virtual reference controller design of high-order nonlinear systems via neural network. Virtual reference feedback tuning (VRFT) is a one-shot direct data-based method to design controller of linear or nonlinear systems. In this paper, we recall the model reference control problem of high-order nonlinear systems and design a new objective function of VRFT. In ideal conditions, the two problems are demonstrated to have the same solution. For the first time, we prove that the value of the optimization problem for model reference control is bounded by that of the objective function of VRFT. A three-layer neural network is employed as a general approximator of the designed controller and two simulations are given to verify the validity of our method. Derong Liu 0001, Ding Wang 0001 |
IJCNN | 2 |
| 2015 | H ∞ Control Synthesis for Linear Parabolic PDE Systems with Model-Free Policy IterationabstractThe H ∞ control problem is considered for linear parabolic partial differential equation (PDE) systems with completely unknown system dynamics. We propose a model-free policy iteration (PI) method for learning the H ∞ control policy by using measured system data without system model information. First, a finite-dimensional system of ordinary differential equation (ODE) is derived, which accurately describes the dominant dynamics of the parabolic PDE system. Based on the finite-dimensional ODE model, the H ∞ control problem is reformulated, which is theoretically equivalent to solving an algebraic Riccati equation (ARE). To solve the ARE without system model information, we propose a least-square based model-free PI approach by using real system data. Finally, the simulation results demonstrate the effectiveness of the developed model-free PI method. Biao Luo 0001, Derong Liu 0001, Xiong Yang 0001, Hongwen Ma |
ISNN | 2 |
| 2015 | A New Discrete-Time Iterative Adaptive Dynamic Programming Algorithm Based on Q-LearningabstractIn this paper, a novel Q -learning based policy iteration adaptive dynamic programming (ADP) algorithm is developed to solve the optimal control problems for discrete-time nonlinear systems. The idea is to use a policy iteration ADP technique to construct the iterative control law which stabilizes the system and simultaneously minimizes the iterative Q function. Convergence property is analyzed to show that the iterative Q function is monotonically non-increasing and converges to the solution of the optimality equation. Finally, simulation results are presented to show the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001 |
ISNN | 2 |
| 2015 | A novel policy iteration based deterministic Q-learning for discrete-time nonlinear systems
Qinglai Wei, Derong Liu 0001 |
Sci. China Inf. Sci. | 2 |
| 2015 | Neural-network-based decentralized control of continuous-time nonlinear interconnected systems with unknown dynamics
Derong Liu 0001, Chao Li 0024, Hongliang Li 0002, Ding Wang 0001, Hongwen Ma |
Neurocomputing | 1 |
| 2015 | Centralized and decentralized event-triggered control for group consensus with fixed topology in continuous time
Hongwen Ma, Derong Liu 0001, Ding Wang 0001, Fuxiao Tan, Chao Li 0024 |
Neurocomputing | 2 |
| 2015 | Computational Energy Management in Smart Grids
Stefano Squartini, Derong Liu 0001, Francesco Piazza, Dongbin Zhao, Haibo He |
Neurocomputing | 2 |
| 2015 | Reliable observer-based H∞ control for discrete-time fuzzy systems with time-varying delays and stochastic actuator faults via scaled small gain theorem
Shenquan Wang, Yulian Jiang, Derong Liu 0001 |
Neurocomputing | 4 |
| 2015 | Neural-network-based adaptive optimal tracking control scheme for discrete-time nonlinear systems with approximation errors
Qinglai Wei, Derong Liu 0001 |
Neurocomputing | 2 |
| 2015 | Convergence analysis and application of fuzzy-HDP for nonlinear discrete-time HJB systems
Yuanheng Zhu, Dongbin Zhao, Derong Liu 0001 |
Neurocomputing | 3 |
| 2015 | Optimal distributed synchronization control for continuous-time heterogeneous multi-agent differential graphical games
Qinglai Wei, Derong Liu 0001, Frank L. Lewis |
Inf. Sci. | 2 |
| 2015 | Reinforcement learning solution for HJB equation arising in constrained optimal control problem
Biao Luo 0001, Huai-Ning Wu, Tingwen Huang, Derong Liu 0001 |
Neural Networks | 4 |
| 2015 | Reinforcement-Learning-Based Robust Controller Design for Continuous-Time Uncertain Nonlinear Systems Subject to Input ConstraintsabstractThe design of stabilizing controller for uncertain nonlinear systems with control constraints is a challenging problem. The constrained-input coupled with the inability to identify accurately the uncertainties motivates the design of stabilizing controller based on reinforcement-learning (RL) methods. In this paper, a novel RL-based robust adaptive control algorithm is developed for a class of continuous-time uncertain nonlinear systems subject to input constraints. The robust control problem is converted to the constrained optimal control problem with appropriately selecting value functions for the nominal system. Distinct from typical action-critic dual networks employed in RL, only one critic neural network (NN) is constructed to derive the approximate optimal control. Meanwhile, unlike initial stabilizing control often indispensable in RL, there is no special requirement imposed on the initial control. By utilizing Lyapunov's direct method, the closed-loop optimal control system and the estimated weights of the critic NN are proved to be uniformly ultimately bounded. In addition, the derived approximate optimal control is verified to guarantee the uncertain nonlinear system to be stable in the sense of uniform ultimate boundedness. Two simulation examples are provided to illustrate the effectiveness and applicability of the present approach. Derong Liu 0001, Xiong Yang 0001, Ding Wang 0001, Qinglai Wei |
IEEE Trans. Cybern. | 1 |
| 2015 | TNNLS Call for Reviewers and Special IssuesabstractIt is my pleasure to report that we have a couple of very successful on-going special issues, including a Special Issue on neurodynamic systems for optimization and applications and a Special Issue on neural networks and learning systems applications in smart grid. They received 101 and 79 submissions, respectively. You can visithttp://www.tnnls.orgfor the call for papers of other Special Issues. If you and your colleagues would like to propose a Special Issue for TNNLS, please follow the instructions posted athttp://cis.ieee.org/tnnlsto submit your proposals. Special Issue proposals are always welcome on timely and hot topics related to neural networks and learning systems. Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Farewell Editorial: Smooth Transition of IEEE TNNLSabstractTime flies! It has been six years since I took over as the Editor-in-Chief of the IEEE TRANSACTIONS ON NEURAL NETWORKS in January 2010. This issue not only marks the last issue of 2015 but is also my last one as the Editor-in-Chief. Serving as the Editor-in-Chief was truly an exciting and good learning experience for me. In the past six years, our TRANSACTIONS has grown significantly.Most importantly, we implemented a title change to the IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, and our annual submission numbers have been doubled to reach 1200 submissions now. It has been a remarkable six years for me to work with amazing authors, reviewers, associate editors, and the IEEE staff. Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Error Bounds of Adaptive Dynamic Programming Algorithms for Solving Undiscounted Optimal Control ProblemsabstractIn this paper, we establish error bounds of adaptive dynamic programming algorithms for solving undiscounted infinite-horizon optimal control problems of discrete-time deterministic nonlinear systems. We consider approximation errors in the update equations of both value function and control policy. We utilize a new assumption instead of the contraction assumption in discounted optimal control problems. We establish the error bounds for approximate value iteration based on a new error condition. Furthermore, we also establish the error bounds for approximate policy iteration and approximate optimistic policy iteration algorithms. It is shown that the iterative approximate value function can converge to a finite neighborhood of the optimal value function under some conditions. To implement the developed algorithms, critic and action neural networks are used to approximate the value function and control policy, respectively. Finally, a simulation example is given to demonstrate the effectiveness of the developed algorithms. Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Infinite Horizon Self-Learning Optimal Control of Nonaffine Discrete-Time Nonlinear SystemsabstractIn this paper, a novel iterative adaptive dynamic programming (ADP)-based infinite horizon self-learning optimal control algorithm, called generalized policy iteration algorithm, is developed for nonaffine discrete-time (DT) nonlinear systems. Generalized policy iteration algorithm is a general idea of interacting policy and value iteration algorithms of ADP. The developed generalized policy iteration algorithm permits an arbitrary positive semidefinite function to initialize the algorithm, where two iteration indices are used for policy improvement and policy evaluation, respectively. It is the first time that the convergence, admissibility, and optimality properties of the generalized policy iteration algorithm for DT nonlinear systems are analyzed. Neural networks are used to implement the developed algorithm. Finally, numerical examples are presented to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Xiong Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Generalized Policy Iteration Adaptive Dynamic Programming for Discrete-Time Nonlinear SystemsabstractThis paper is concerned with a novel generalized policy iteration algorithm for solving optimal control problems for discrete-time nonlinear systems. The idea is to use an iterative adaptive dynamic programming algorithm to obtain iterative control laws which make the iterative value functions converge to the optimum. Initialized by an admissible control law, it is shown that the iterative value functions are monotonically nonincreasing and converge to the optimal solution of Hamilton-Jacobi-Bellman equation, under the assumption that a perfect function approximation is employed. The admissibility property is analyzed, which shows that any of the iterative control laws can stabilize the nonlinear system. Neural networks are utilized to implement the generalized policy iteration algorithm, by approximating the iterative value function and computing the iterative control law, respectively, to achieve approximate optimal control. Finally, numerical examples are presented to verify the effectiveness of the present generalized policy iteration algorithm. Derong Liu 0001, Qinglai Wei |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2014 | Optimal self-learning battery control in smart residential grids by iterative Q-learning algorithmabstractIn this paper, a novel dual iterative Q-learning algorithm is developed to solve the optimal battery management and control problems in smart residential environments. The main idea is to use adaptive dynamic programming (ADP) technique to obtain the optimal battery management and control scheme iteratively for residential energy systems. In the developed dual iterative Q-learning algorithm, two iterations, including external and internal iterations, are introduced, where internal iteration minimizes the total cost of power loads in each period and the external iteration makes the iterative Q function converge to the optimum. For the first time, the convergence property of iterative Q-learning method is proven to guarantee the convergence property of the iterative Q function. Finally, numerical results are given to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Yu Liu 0078, Qiang Guan |
ADPRL | 2 |
| 2014 | Discrete-Time Nonlinear Generalized Policy Iteration for Optimal Control Using Neural Networks
Qinglai Wei, Derong Liu 0001, Xiong Yang 0001 |
ICONIP (1) | 2 |
| 2014 | Data-driven iterative adaptive dynamic programming algorithm for approximate optimal control of unknown nonlinear systemsabstractIn this paper, we develop a data-driven iterative adaptive dynamic programming algorithm to learn offline the approximate optimal control of unknown discrete-time nonlinear systems. We do not use a model network to identify the unknown system, but utilize the available offline data to learn the approximate optimal control directly. First, the data-driven iterative adaptive dynamic programming algorithm is presented with a convergence analysis. Then, the error bounds for this algorithm are provided considering the approximation errors of function approximation structures. To implement the developed algorithm, two neural networks are used to approximate the state-action value function and the control policy. Finally, two simulation examples are given to demonstrate the effectiveness of the developed algorithm. Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001, Chao Li 0024 |
IJCNN | 2 |
| 2014 | Unmanned aerial vehicles (UAV) heading optimal tracking control using online kernel-based HDP algorithmabstractUAV can work in places that are dangerous, or not easy to reach for humans. However, due to active control and operating difficulties, it is still a challenge to develop fully autonomous flight in complex environments. This paper applies a novel heuristic dynamic programming for the UAV heading optimal tracking controller design, using kernel-based heuristic dynamic programming (KHDP). Kernel-based HDP is developed by integrating kernel methods and approximately linear dependence (ALD) analysis with the critic learning of HDP algorithm. Compared with conventional HDP where neural networks are widely used and their features were manually designed, the proposed algorithm can obtain better generalization capability and learning efficiency through applying the sparse kernel machine into the critic learning process of HDP algorithm. Simulation and experimental results of UAV heading optimal tracking control problems demonstrate the effectiveness of the proposed kernel-based HDP algorithm. Fuxiao Tan, Derong Liu 0001, Xin-Ping Guan, Bin Luo 0001 |
IJCNN | 2 |
| 2014 | Near-optimal online control of uncertain nonlinear continuous-time systems based on concurrent learningabstractThis paper presents a novel observer-critic architecture for solving the near-optimal control problem of uncertain nonlinear continuous-time systems. Two neural networks (NNs) are employed in the architecture: an observer NN is constructed to get the knowledge of uncertain system dynamics and a critic NN is utilized to derive the optimal control. The observer NN and the critic NN are tuned simultaneously. By using the recorded and instantaneous data together, the optimal control can be derived without the persistence of excitation condition. Meanwhile, the closed-loop system is guaranteed to be stable in the sense of uniform ultimate boundedness. No initial stabilizing control is required in the developed algorithm. An illustrated example is provided to demonstrate the effectiveness of the present approach. Xiong Yang 0001, Derong Liu 0001, Qinglai Wei |
IJCNN | 2 |
| 2014 | Reinforcement-Learning-Based Controller Design for Nonaffine Nonlinear Systems
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei |
ISNN | 2 |
| 2014 | Neural-network-based optimal tracking control scheme for a class of unknown discrete-time nonlinear systems using iterative ADP algorithm
Yuzhu Huang, Derong Liu 0001 |
Neurocomputing | 2 |
| 2014 | Dual Heuristic dynamic Programming for nonlinear discrete-time uncertain systems with state delay
Bin Wang 0034, Dongbin Zhao, Cesare Alippi, Derong Liu 0001 |
Neurocomputing | 4 |
| 2014 | Data-based analysis of discrete-time linear systems in noisy environment: Controllability and observability
Derong Liu 0001, Qinglai Wei |
Inf. Sci. | 1 |
| 2014 | Neural-network-based robust optimal control design for a class of uncertain nonlinear systems via adaptive dynamic programming
Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Hongwen Ma |
Inf. Sci. | 2 |
| 2014 | Stable iterative adaptive dynamic programming algorithm with approximation errors for discrete-time nonlinear systems
Qinglai Wei, Derong Liu 0001 |
Neural Comput. Appl. | 2 |
| 2014 | Discrete-time online learning control for a class of unknown nonaffine nonlinear systems using reinforcement learning
Xiong Yang 0001, Derong Liu 0001, Ding Wang 0001, Qinglai Wei |
Neural Networks | 2 |
| 2014 | Approximate optimal solution of the DTHJB equation for a class of nonlinear affine systems with unknown dead-zone constraints
Dehua Zhang, Derong Liu 0001, Ding Wang 0001 |
Soft Comput. | 2 |
| 2014 | Integral Reinforcement Learning for Linear Continuous-Time Zero-Sum Games With Completely Unknown DynamicsabstractIn this paper, we develop an integral reinforcement learning algorithm based on policy iteration to learn online the Nash equilibrium solution for a two-player zero-sum differential game with completely unknown linear continuous-time dynamics. This algorithm is a fully model-free method solving the game algebraic Riccati equation forward in time. The developed algorithm updates value function, control and disturbance policies simultaneously. The convergence of the algorithm is demonstrated to be equivalent to Newton's method. To implement this algorithm, one critic network and two action networks are used to approximate the game value function, control and disturbance policies, respectively, and the least squares method is used to estimate the unknown parameters. The effectiveness of the developed scheme is demonstrated in the simulation by designing an H∞state feedback controller for a power system. Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2014 | Policy Iteration Algorithm for Online Design of Robust Control for a Class of Continuous-Time Nonlinear SystemsabstractIn this paper, a novel strategy is established to design the robust controller for a class of continuous-time nonlinear systems with uncertainties based on the online policy iteration algorithm. The robust control problem is transformed into the optimal control problem by properly choosing a cost function that reflects the uncertainties, regulation, and control. An online policy iteration algorithm is presented to solve the Hamilton-Jacobi-Bellman (HJB) equation by constructing a critic neural network. The approximate expression of the optimal control policy can be derived directly. The closed-loop system is proved to possess the uniform ultimate boundedness. The equivalence of the neural-network-based HJB solution of the optimal control problem and the solution of the robust control problem is established as well. Two simulation examples are provided to verify the effectiveness of the present robust control scheme. Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2014 | Adaptive Dynamic Programming for Optimal Tracking Control of Unknown Nonlinear Systems With Application to Coal GasificationabstractIn this paper, we establish a new data-based iterative optimal learning control scheme for discrete-time nonlinear systems using iterative adaptive dynamic programming (ADP) approach and apply the developed control scheme to solve a coal gasification optimal tracking control problem. According to the system data, neural networks (NNs) are used to construct the dynamics of coal gasification process, coal quality and reference control, respectively, where the mathematical model of the system is unnecessary. The approximation errors from neural network construction of the disturbance and the controls are both considered. Via system transformation, the optimal tracking control problem with approximation errors and disturbances is effectively transformed into a two-person zero-sum optimal control problem. A new iterative ADP algorithm is then developed to obtain the optimal control laws for the transformed system. Convergence property is developed to guarantee that the performance index function converges to a finite neighborhood of the optimal performance index function, and the convergence criterion is also obtained. Finally, numerical results are given to illustrate the performance of the present method. Qinglai Wei, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2014 | A Novel Iterative $\theta $-Adaptive Dynamic Programming for Discrete-Time Nonlinear SystemsabstractThis paper is concerned with a new iterative θ-adaptive dynamic programming (ADP) technique to solve optimal control problems of infinite horizon discrete-time nonlinear systems. The idea is to use an iterative ADP algorithm to obtain the iterative control law which optimizes the iterative performance index function. In the present iterative θ-ADP algorithm, the condition of initial admissible control in policy iteration algorithm is avoided. It is proved that all the iterative controls obtained in the iterative θ-ADP algorithm can stabilize the nonlinear system which means that the iterative θ-ADP algorithm is feasible for implementations both online and offline. Convergence analysis of the performance index function is presented to guarantee that the iterative performance index function will converge to the optimum monotonically. Neural networks are used to approximate the performance index function and compute the optimal control policy, respectively, for facilitating the implementation of the iterative θ-ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the established method. Qinglai Wei, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2014 | Neural-Network-Based Online HJB Solution for Optimal Robust Guaranteed Cost Control of Continuous-Time Uncertain Nonlinear SystemsabstractIn this paper, the infinite horizon optimal robust guaranteed cost control of continuous-time uncertain nonlinear systems is investigated using neural-network-based online solution of Hamilton-Jacobi-Bellman (HJB) equation. By establishing an appropriate bounded function and defining a modified cost function, the optimal robust guaranteed cost control problem is transformed into an optimal control problem. It can be observed that the optimal cost function of the nominal system is nothing but the optimal guaranteed cost of the original uncertain system. A critic neural network is constructed to facilitate the solution of the modified HJB equation corresponding to the nominal system. More importantly, an additional stabilizing term is introduced for helping to verify the stability, which reinforces the updating process of the weight vector and reduces the requirement of an initial stabilizing control. The uniform ultimate boundedness of the closed-loop system is analyzed by using the Lyapunov approach as well. Two simulation examples are provided to verify the effectiveness of the present control approach. Derong Liu 0001, Ding Wang 0001, Fei-Yue Wang 0001, Hongliang Li 0002, Xiong Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic ProgrammingabstractIn this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal control problems for infinite horizon discrete-time nonlinear systems with finite approximation errors. First, a new generalized value iteration algorithm of ADP is developed to make the iterative performance index function converge to the solution of the Hamilton-Jacobi-Bellman equation. The generalized value iteration algorithm permits an arbitrary positive semi-definite function to initialize it, which overcomes the disadvantage of traditional value iteration algorithms. When the iterative control law and iterative performance index function in each iteration cannot accurately be obtained, for the first time a new "design method of the convergence criteria" for the finite-approximation-error-based generalized value iteration algorithm is established. A suitable approximation error can be designed adaptively to make the iterative performance index function converge to a finite neighborhood of the optimal performance index function. Neural networks are used to implement the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the developed method. Qinglai Wei, Fei-Yue Wang 0001, Derong Liu 0001, Xiong Yang 0001 |
IEEE Trans. Cybern. | 3 |
| 2014 | Editorial TNNLS Special Issues and Editorial Board ChangesabstractPresents a listing of changes to the Editorial Board for this publication. Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Policy Iteration Adaptive Dynamic Programming Algorithm for Discrete-Time Nonlinear SystemsabstractThis paper is concerned with a new discrete-time policy iteration adaptive dynamic programming (ADP) method for solving the infinite horizon optimal control problem of nonlinear systems. The idea is to use an iterative ADP technique to obtain the iterative control law, which optimizes the iterative performance index function. The main contribution of this paper is to analyze the convergence and stability properties of policy iteration method for discrete-time nonlinear systems for the first time. It shows that the iterative performance index function is nonincreasingly convergent to the optimal solution of the Hamilton-Jacobi-Bellman equation. It is also proven that any of the iterative control laws can stabilize the nonlinear systems. Neural networks are used to approximate the performance index function and compute the optimal control law, respectively, for facilitating the implementation of the iterative ADP algorithm, where the convergence of the weight matrices is analyzed. Finally, the numerical results and analysis are presented to illustrate the performance of the developed method. Derong Liu 0001, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Decentralized Stabilization for a Class of Continuous-Time Nonlinear Interconnected Systems Using Online Learning Optimal Control ApproachabstractIn this paper, using a neural-network-based online learning optimal control approach, a novel decentralized control strategy is developed to stabilize a class of continuous-time nonlinear interconnected large-scale systems. First, optimal controllers of the isolated subsystems are designed with cost functions reflecting the bounds of interconnections. Then, it is proven that the decentralized control strategy of the overall system can be established by adding appropriate feedback gains to the optimal control policies of the isolated subsystems. Next, an online policy iteration algorithm is presented to solve the Hamilton-Jacobi-Bellman equations related to the optimal control problem. Through constructing a set of critic neural networks, the cost functions can be obtained approximately, followed by the control policies. Furthermore, the dynamics of the estimation errors of the critic networks are verified to be uniformly and ultimately bounded. Finally, a simulation example is provided to illustrate the effectiveness of the present decentralized control scheme. Derong Liu 0001, Ding Wang 0001, Hongliang Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | A Comprehensive Review of Stability Analysis of Continuous-Time Recurrent Neural NetworksabstractStability problems of continuous-time recurrent neural networks have been extensively studied, and many papers have been published in the literature. The purpose of this paper is to provide a comprehensive review of the research on stability of continuous-time recurrent neural networks, including Hopfield neural networks, Cohen-Grossberg neural networks, and related models. Since time delay is inevitable in practice, stability results of recurrent neural networks with different classes of time delays are reviewed in detail. For the case of delay-dependent stability, the results on how to deal with the constant/variable delay in recurrent neural networks are summarized. The relationship among stability results in different forms, such as algebraic inequality forms, M-matrix forms, linear matrix inequality forms, and Lyapunov diagonal stability forms, is discussed and compared. Some necessary and sufficient stability conditions for recurrent neural networks without time delays are also discussed. Concluding remarks and future directions of stability analysis of recurrent neural networks are given. Huaguang Zhang, Zhanshan Wang 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Detecting and Reacting to Changes in Sensing Units: The Active Classifier CaseabstractThe ability to detect concept drift, i.e., a structural change in the acquired datastream, and react accordingly is a major achievement for intelligent sensing units. This ability allows the unit, for actively tuning the application, to maintain high performance, changing online the operational strategy, detecting and isolating possible occurring faults to name a few tasks. In the paper, we consider a just-in-time strategy for adaptation; the sensing unit reacts exactly when needed, i.e., when concept drift is detected. Change detection tests (CDTs), designed to inspect structural changes in industrial and environmental data, are coupled here with adaptive k-nearest neighbor and support vector machine classifiers, and suitably retrained when the change is detected. Computational complexity and memory requirements of the CDT and the classifier, due to precious limited resources in embedded sensing, are taken into account in the application design. We show that a hierarchical CDT coupled with an adaptive resource-aware classifier is a suitable tool for processing and classifying sequential streams of data. Cesare Alippi, Derong Liu 0001, Dongbin Zhao, Li Bu |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2014 | Online Synchronous Approximate Optimal Learning Algorithm for Multi-Player Non-Zero-Sum Games With Unknown DynamicsabstractIn this paper, we develop an online synchronous approximate optimal learning algorithm based on policy iteration to solve a multiplayer nonzero-sum game without the requirement of exact knowledge of dynamical systems. First, we prove that the online policy iteration algorithm for the nonzero-sum game is mathematically equivalent to the quasi-Newton's iteration in a Banach space. Then, a model neural network is established to identify the unknown continuous-time nonlinear system using input-output data. For each player, a critic neural network and an action neural network are used to approximate its value function and control policy, respectively. Our algorithm only needs to tune the weights of critic neural networks, so there will be less computational complexity during the learning process. All the neural network weights are updated online in real-time, continuously and synchronously. Furthermore, the uniform ultimate bounded stability of the closed-loop system is proved based on Lyapunov approach. Finally, two simulation examples are given to demonstrate the effectiveness of the developed scheme. Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Integral Policy Iteration for Zero-Sum Games with Completely Unknown Nonlinear Dynamics
Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001 |
ICONIP (1) | 2 |
| 2013 | Observer-Based Adaptive Output Feedback Control for Nonaffine Nonlinear Discrete-Time Systems Using Reinforcement Learning
Xiong Yang 0001, Derong Liu 0001, Ding Wang 0001 |
ICONIP (1) | 2 |
| 2013 | Convergence analysis of continuous-time systems based on feedforward neural networksabstractIn this paper, we construct a feedforward neural network (NN) based system containing two NNs. The convergence of the NN based system is analyzed in detail. For setting up the NN based system, an NN observer is first designed to estimate the system states. Then, based on the observed states, a feedforward neuro-control system is constructed by using adaptive dynamic programming (ADP). In this design, two NN structures are used: a three-layer feedforward NN to constitute the observer which can be applied to the systems with high degrees of nonlinearity and without a priori knowledge about system dynamics, and a critic NN to approximate the value function. Moreover, the weight update laws for the critic NN are generated using a gradient-descent method based on a modified temporal difference error, which is independent of the system dynamics. Finally, uniform ultimate boundedness (UUB) of the NN based system is proved. Yuzhu Huang, Derong Liu 0001, Qinglai Wei |
ISCAS | 2 |
| 2013 | Neural Network H ∞ Tracking Control of Nonlinear Systems Using GHJI Method
Derong Liu 0001, Yuzhu Huang, Qinglai Wei |
ISNN (2) | 1 |
| 2013 | Optimal Tracking Control Scheme for Discrete-Time Nonlinear Systems with Approximation Errors
Qinglai Wei, Derong Liu 0001 |
ISNN (2) | 2 |
| 2013 | Neural-network-based zero-sum game for discrete-time nonlinear systems via iterative adaptive dynamic programming algorithm
Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001 |
Neurocomputing | 1 |
| 2013 | Neuro-optimal control for a class of unknown nonlinear dynamic systems using SN-DHP technique
Ding Wang 0001, Derong Liu 0001 |
Neurocomputing | 2 |
| 2013 | An iterative adaptive dynamic programming algorithm for optimal control of unknown discrete-time nonlinear systems with constrained inputs
Derong Liu 0001, Ding Wang 0001, Xiong Yang 0001 |
Inf. Sci. | 1 |
| 2013 | Data-based stability analysis of a class of nonlinear discrete-time systems
Zhuo Wang 0003, Derong Liu 0001 |
Inf. Sci. | 2 |
| 2013 | A self-learning scheme for residential energy system control and management
Derong Liu 0001 |
Neural Comput. Appl. | 2 |
| 2013 | Adaptive optimal control for a class of continuous-time affine nonlinear systems with unknown internal dynamics
Derong Liu 0001, Xiong Yang 0001, Hongliang Li 0002 |
Neural Comput. Appl. | 1 |
| 2013 | A neural-network-based iterative GDHP approach for solving a class of nonlinear optimal control problems with control constraints
Ding Wang 0001, Derong Liu 0001, Dongbin Zhao, Yuzhu Huang, Dehua Zhang |
Neural Comput. Appl. | 2 |
| 2013 | Special issue on intelligent control and information processing
Dongbin Zhao, Cesare Alippi, Derong Liu 0001, Huaguang Zhang |
Soft Comput. | 3 |
| 2013 | A supervised Actor-Critic approach for adaptive cruise control
Dongbin Zhao, Bin Wang 0034, Derong Liu 0001 |
Soft Comput. | 3 |
| 2013 | Finite-Approximation-Error-Based Optimal Control Approach for Discrete-Time Nonlinear SystemsabstractIn this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal control problems for infinite-horizon discrete-time nonlinear systems with finite approximation errors. The idea is to use an iterative ADP algorithm to obtain the iterative control law that makes the iterative performance index function reach the optimum. When the iterative control law and the iterative performance index function in each iteration cannot be accurately obtained, the convergence conditions of the iterative ADP algorithm are obtained. When convergence conditions are satisfied, it is shown that the iterative performance index functions can converge to a finite neighborhood of the greatest lower bound of all performance index functions under some mild assumptions. Neural networks are used to approximate the performance index function and compute the optimal control policy, respectively, for facilitating the implementation of the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the present method. Derong Liu 0001, Qinglai Wei |
IEEE Trans. Cybern. | 1 |
| 2013 | A Data-Based State Feedback Control Method for a Class of Nonlinear SystemsabstractIn this paper, a data-based state feedback control method is developed for a class of nonlinear systems. It is a real-time control method, which requires little prior knowledge about the system dynamics, and does not need to know or to build the mathematical model of the system. We apply a fast sampling technique to sample the state signal, which contains useful information of the system. The zero-order hold (ZOH) and the control switch are also used to obtain system information. The feedback gain matrix is calculated and adjusted according to these sampled data. Theoretical analysis on the convergence and simulation results demonstrate the feasibility of this data-based control method. Zhuo Wang 0003, Derong Liu 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2013 | Editorial A Successful Change From TNN to TNNLS and a Very Successful YearabstractThis issue marks the first anniversary issue of IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS after it changed its name from IEEE TRANSACTIONS ON NEURAL NETWORKS. I am happy to report that we had a great year! The number of new submissions in a year exceeded 1,000 for the first time in the history of TNN/TNNLS. IEEE TNN had a very successful development for 22 years from 1990 to 2011, and we have good reasons to believe that IEEE TNNLS will have many more years of successful growth. Derong Liu 0001, Charles W. Anderson, Ahmad Taher Azar, Giorgio Battistelli, Eduardo Bayro-Corrochano, Cristiano Cervellera, David A. Elizondo, Maurizio Filippone, Giorgio Gnecco, Tingwen Huang, Weifeng Liu 0016, Wenlian Lu, Ana Madureira, Igor Skrjanc, Thomas Villmann, Q. M. Jonathan Wu, Shengli Xie 0001, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Generalized Hamilton-Jacobi-Isaacs Formulation-Based Neural Network H ∞ Control for Constrained Input Nonlinear Systems
Yuzhu Huang, Derong Liu 0001, Qinglai Wei |
ICONIP (1) | 2 |
| 2012 | Nearly Optimal Control for Nonlinear Systems with Dead-Zone Control Input Based on the Iterative ADP Approach
Dehua Zhang, Derong Liu 0001, Qinglai Wei |
ICONIP (1) | 2 |
| 2012 | H∞ control of unknown discrete-time nonlinear systems with control constraints using adaptive dynamic programmingabstractIn this paper, we solve the H∞robust optimal control problem for discrete-time nonlinear systems with control saturation constraints using the iterative adaptive dynamic programming algorithm. First, a heuristic dynamic programming algorithm is derived to solve the Hamilton-Jacobi-Isaacs equation associated with the H∞control problem, and a convergence analysis is provided. Then, a dual heuristic dynamic programming algorithm with nonquadratic performance functional is developed to overcome the control saturation constraints. Finally, to facilitate the implementation of the algorithm, four neural networks are used to approximate the unknown nonlinear system, the control policy, the disturbance policy, and the value function. Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001 |
IJCNN | 1 |
| 2012 | Adaptive dynamic programming with stable value iteration algorithm for discrete-time nonlinear systemsabstractIn this paper, a new stable value iteration adaptive dynamic programming (ADP) algorithm, named “θ-ADP” algorithm, is proposed for solving the optimal control problems of infinite horizon discrete-time nonlinear systems. By introducing a parameter θ in the iterative ADP algorithm, it is proved that any of iterative control obtained in the proposed algorithm can stabilize the nonlinear system which overcomes the disadvantage of traditional value iteration algorithms. Neural networks are used to approximate the performance index function and compute the optimal control policy, respectively, for facilitating the implementation of the iterative θ-ADP algorithm. Finally, a simulation example is given to illustrate the performance of the proposed method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 2 |
| 2012 | Optimal Battery Management with ADHDP in Smart Home Environments
Danilo Fuselli, Francesco De Angelis 0002, Matteo Boaro, Derong Liu 0001, Qinglai Wei, Stefano Squartini, Francesco Piazza |
ISNN (2) | 4 |
| 2012 | Temperature Control in Water-Gas Shift Reaction with Adaptive Dynamic Programming
Yuzhu Huang, Derong Liu 0001, Qinglai Wei |
ISNN (2) | 2 |
| 2012 | Self-learning Control Schemes for Two-Person Zero-Sum Differential Games of Continuous-Time Nonlinear Systems with Saturating Controllers
Qinglai Wei, Derong Liu 0001 |
ISNN (2) | 2 |
| 2012 | Optimal Cooperative Spectrum Sensing in Cognitive Radio with Taguchi MethodabstractSpectrum sensing is an essential topic in cognitive radio networks to detect primary users. Cooperative spectrum sensing with linear fusion scheme is studied in multiband channel systems. The problem turns out to be a nonconvex optimization problem in general with thresholds of the energy detectors and the weights of linear fusion scheme as parameters. In this paper, we use the Taguchi method to estimate the gradient of the aggregate throughput function, determine thresholds of the energy detector and linear combination weights of the linear fusion rule. One of the advantages of our approach is to optimize the thresholds and the linear weights simultaneously. In addition, we modify the input parameters of experiments to satisfy the constraints before we calculate the cost and analyze the trends, and it improves the computation and search efficiency. Simulation results show that the proposed method can be used in all classes of CR. Our approach has a very good performance with short computational time and it is also insensitive to initial values of parameters and gives a relatively stable result. Derong Liu 0001 |
VTC Fall | 2 |
| 2012 | Finite-horizon neuro-optimal tracking control for a class of discrete-time nonlinear systems using adaptive dynamic programming approach
Ding Wang 0001, Derong Liu 0001, Qinglai Wei |
Neurocomputing | 2 |
| 2012 | An iterative ϵ-optimal control scheme for a class of discrete-time nonlinear systems with unfixed initial state
Qinglai Wei, Derong Liu 0001 |
Neural Networks | 2 |
| 2012 | Neural-Network-Based Optimal Control for a Class of Unknown Discrete-Time Nonlinear Systems Using Globalized Dual Heuristic ProgrammingabstractIn this paper, a neuro-optimal control scheme for a class of unknown discrete-time nonlinear systems with discount factor in the cost function is developed. The iterative adaptive dynamic programming algorithm using globalized dual heuristic programming technique is introduced to obtain the optimal controller with convergence analysis in terms of cost function and control law. In order to carry out the iterative algorithm, a neural network is constructed first to identify the unknown controlled system. Then, based on the learned system model, two other neural networks are employed as parametric structures to facilitate the implementation of the iterative algorithm, which aims at approximating at each iteration the cost function and its derivatives and the control law, respectively. Finally, a simulation example is provided to verify the effectiveness of the proposed optimal control approach. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao, Qinglai Wei |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2012 | Neural Networks and Learning Systems Come TogetherabstractThis issue marks the beginning of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS). By adding "Learning Systems" to the title, we now state explicitly the scope of the Transactions to include neural networks as well as related learning systems. This issue marks a new era in the history of our Transactions. The Transactions is now ready to face the challenges of the next 10-20 years. With the evolution of the fields of neural networks in particular and computational intelligence in general, the IEEE Transactions on Neural Networks and Learning Systems will continue to grow and to succeed in this ever-changing world. Also included are a few comments about the review process of TNN manuscripts and the introduction of 14 new TNNLS Associate Editors. Short biographies are included for the new Associate Editors. Bart Baesens, Pantelis Bouboulis, Sergio Cruces, Carlotta Domeniconi, Shiro Ikeda, Xuelong Li 0001, Patricia Melin, Vadrevu Sree Hari Rao, Björn W. Schuller, Huajin Tang, Cong Wang 0033, Jian Yang 0003, Derong Zhao, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 15 |
| 2011 | Adaptive dynamic programming for optimal control of unknown nonlinear discrete-time systemsabstractAn intelligent optimal control scheme for unknown nonlinear discrete-time systems with discount factor in the cost function is proposed in this paper. An iterative adaptive dynamic programming (ADP) algorithm via globalized dual heuristic programming (GDHP) technique is developed to obtain the optimal controller with convergence analysis. Three neural networks are used as parametric structures to facilitate the implementation of the iterative algorithm, which will approximate at each iteration the cost function, the optimal control law, and the unknown nonlinear system, respectively. Two simulation examples are provided to verify the effectiveness of the presented optimal control approach. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao |
ADPRL | 1 |
| 2011 | Residential energy system control and management using adaptive dynamic programmingabstractIn this paper, we apply adaptive dynamic programming to the residential energy system control and management, with an emphasis on home battery use connected to power grids. The proposed scheme is built upon a self-learning architecture with only a single critic module instead of the action-critic dual module architecture. The novelty of the present scheme is its ability to improve the performance as it learns and gains more experience in real-time operations under uncertain changes of the environment. Simulation results demonstrate that the proposed scheme can achieve the minimum electricity cost for residential customers. Derong Liu 0001 |
IJCNN | 2 |
| 2011 | Neural-network-based optimal control for a class of nonlinear cdiscrete-time systems with control constraints using the citerative GDHP algorithmabstractIn this paper, a neural-network-based optimal control scheme for a class of nonlinear discrete-time systems with control constraints is proposed. The iterative adaptive dynamic programming (ADP) algorithm via globalized dual heuristic programming (GDHP) technique is developed to design the optimal controller with convergence proof. Three neural networks are used to facilitate the implementation of the iterative algorithm, which will approximate at each iteration the cost function, the optimal control law, and the controlled nonlinear discrete-time system, respectively. A simulation study is carried out to demonstrate the effectiveness of the present approach in dealing with the nonlinear constrained optimal control problem. Derong Liu 0001, Ding Wang 0001, Dongbin Zhao |
IJCNN | 1 |
| 2011 | Optimal control for discrete-time nonlinear systems with unfixed initial state using adaptive dynamic programmingabstractA new ε-optimal control algorithm based on the adaptive dynamic programming (ADP) is proposed to solve the finite horizon optimal control problem for a class of discrete-time nonlinear systems with unfixed initial state. The proposed algorithm makes the performance index function converges iteratively to the greatest lower bound of all performance indices within an error bound according to ε with finite time. The number of optimal control steps can also be obtained by the proposed ADP approach for the situation when the initial state of the system is unfixed. A simulation example is given to show the performance of the present method. Qinglai Wei, Derong Liu 0001 |
IJCNN | 2 |
| 2011 | Optimal Control for a Class of Unknown Nonlinear Systems via the Iterative GDHP Algorithm
Ding Wang 0001, Derong Liu 0001 |
ISNN (2) | 2 |
| 2011 | Finite Horizon Optimal Tracking Control for a Class of Discrete-Time Nonlinear Systems
Qinglai Wei, Ding Wang 0001, Derong Liu 0001 |
ISNN (2) | 3 |
| 2011 | Nonlinear Torque and Air-to-Fuel Ratio Control of SPARK Ignition Engines Using Neuro-Sliding Mode TechniquesabstractThis paper presents a new approach for the calibration and control of spark ignition engines using a combination of neural networks and sliding mode control technique. Two parallel neural networks are utilized to realize a neuro-sliding mode control (NSLMC) for self-learning control of automotive engines. The equivalent control and the corrective control terms are the outputs of the neural networks. Instead of using error backpropagation algorithm, the network weights of equivalent control are updated using the Levenberg-Marquardt algorithm. Moreover, a new approach is utilized to update the gain of corrective control. Both modifications of the NSLMC are aimed at improving the transient performance and speed of convergence. Using the data from a test vehicle with a V8 engine, we built neural network models for the engine torque (TRQ) and the air-to-fuel ratio (AFR) dynamics and developed NSLMC controllers to achieve tracking control. The goal of TRQ control and AFR control is to track the commanded values under various operating conditions. From simulation studies, the feasibility and efficiency of the approach are illustrated. For both control problems, excellent tracking performance has been achieved. Hossein Javaherian, Derong Liu 0001 |
Int. J. Neural Syst. | 3 |
| 2011 | An Introduction to Parallel Control and Management for High-Speed Railway SystemsabstractThis paper introduces a framework of parallel control and management for high-speed railway systems (HRSs). First, based on multiagent modeling, an artificial HRS that is consistent with realistic operations of the actual HRS is constructed. Then, different kinds of computational experiments are performed on the artificial HRS, followed by analysis and synthesis with a case. Finally, through an interactive and parallel operation between the actual and artificial HRSs, a set of practical control and management strategies can be achieved for the actual HRS. With the primary objective of ensuring reliability and safety of HRSs, this study could enhance the quality of services and the integrated transportability with other existing modes of transportation systems to provide appropriate recommendations and strategies for forming an overall effective comprehensive transportation system. Tao Tang 0004, Hairong Dong 0001, Ding Wen, Derong Liu 0001, Shigen Gao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2011 | Editorial: The Blossoming of the IEEE Transactions on Neural Networks
Derong Liu 0001 |
IEEE Trans. Neural Networks | 1 |
| 2011 | Adaptive Dynamic Programming for Finite-Horizon Optimal Control of Discrete-Time Nonlinear Systems With varepsilon-Error BoundabstractIn this paper, we study the finite-horizon optimal control problem for discrete-time nonlinear systems using the adaptive dynamic programming (ADP) approach. The idea is to use an iterative ADP algorithm to obtain the optimal control law which makes the performance index function close to the greatest lower bound of all performance indices within an ε-error bound. The optimal number of control steps can also be obtained by the proposed ADP algorithms. A convergence analysis of the proposed ADP algorithms in terms of performance index function and control policy is made. In order to facilitate the implementation of the iterative ADP algorithms, neural networks are used for approximating the performance index function, computing the optimal control policy, and modeling the nonlinear system. Finally, two simulation examples are employed to illustrate the applicability of the proposed method. Fei-Yue Wang 0001, N. Jin, Derong Liu 0001, Q. Wei |
IEEE Trans. Neural Networks | 3 |
| 2011 | Data-Based Controllability and Observability Analysis of Linear Discrete-Time SystemsabstractIn this brief, we develop data-based methods for analyzing the controllability and observability of linear discrete-time systems which have unknown system parameters. These data-based methods will only use measured data to construct the controllability matrix as well as the observability matrix, in order to verify the corresponding properties. The advantages of our methods are threefold. First, they can directly verify system properties based on measured data without knowing system parameters. Second, our calculation precision is higher than traditional approaches, which need to identify the unknown parameters. Third, our methods have lower computational complexities when constructing the controllability and observability matrices. Zhuo Wang 0003, Derong Liu 0001 |
IEEE Trans. Neural Networks | 2 |
| 2010 | Adaptive dynamic programming for a class of discrete-time non-affine nonlinear systems with time-delaysabstractIn this paper, an optimal control scheme for a class of non-affine nonlinear systems with time-delays in state and control variables is developed using a new iterative adaptive dynamic programming (ADP) algorithm. By introducing delay matrix functions, the explicit expression of optimal control is solved using the dynamic programming theory and the optimal control can iteratively be solved using the present technique. Convergence analysis is presented to show the performance index function to reach the optimum by the present method. Neural networks are used to approximate the performance index function, compute the optimal control policy, solve delay matrix functions and model the nonlinear system, respectively, for facilitating the implementation of the iterative ADP algorithm. A simulation example is given to demonstrate the validity of the present optimal control scheme. Derong Liu 0001, Qinglai Wei |
IJCNN | 1 |
| 2010 | Editorial: the IEEE transactions on neural networks 2010 and beyond
Derong Liu 0001 |
IEEE Trans. Neural Networks | 1 |
| 2009 | Constrained optimal control of affine nonlinear discrete-time systems using GHJB methodabstractThe infinite-horizon optimal control problem of nonlinear discrete-time systems with actuator saturation is considered in this paper. In order to deal with actuator saturation, a novel nonquadratic functional is introduced, and the constrained Generalized Hamilton-Jacobi-Bellman (GHJB) equation and Hamilton-Jacobi-Bellman (HJB) equation for nonlinear discrete-time systems are derived in terms of non-quadratic functionals. The optimal saturated controller is obtained by a novel iterative algorithm based on the constrained GHJB equation, and a convergence proof is presented, where a neural network is used to approximate the value function. Finally, a nearly optimal saturated controller is obtained. The effectiveness of this algorithm is demonstrated by a numerical example. Lili Cui, Huaguang Zhang, Derong Liu 0001, Yongsu Kim |
ADPRL | 3 |
| 2009 | Neural-network-based reinforcement learning controller for nonlinear systems with non-symmetric dead-zone inputsabstractA novel adaptive-critic-based NN controller using reinforcement learning is developed for a class of nonlinear systems with non-symmetric dead-zone inputs. The adaptive critic NN controller uses two NNs: the critic NN is used to approximate the strategic utility function, and the output of action NN is used to approximate the unknown nonlinear function and to minimize the strategic utility function. The tuning of the NNs is performed online without an explicit offline learning phase. The uniformly ultimate boundedness of the close-loop tracking error is derived by using using the Lyapunov method. Finally, a numerical example is included to show the effectiveness of the theoretical results. Huaguang Zhang, Derong Liu 0001, Yongsu Kim |
ADPRL | 3 |
| 2009 | A novel control scheme for a class of nonlinear systems with time delays based on fuzzy hyperbolic modelabstractThis paper concerns the problem of mixed H2/Hinfincontrol of a class of nonlinear continuous-time systems with time delays, which can be represented by a delayed fuzzy hyperbolic model (DFHM). The main advantage of using DFHM over T-S fuzzy model is that no premise structure identification is needed and no completeness design of premise variables space is needed. In addition, a DFHM is not only a kind of valid global description but also a kind of nonlinear model essentially. Based on delay-independent Lyapunov functional approach, some sufficient conditions for the existence of such a DFHM-based mixed H2/Hinfincontroller are provided, which are given in terms of the feasibility of a set of linear matrix inequalities (LMIs). Simulation results show the validity of the proposed method. Jun Yang 0008, Huaguang Zhang, Derong Liu 0001 |
FUZZ-IEEE | 3 |
| 2009 | Adaptive dynamic programming for discrete-time systems with infinite horizon and ELEMENT OF -error bound in the performance costabstractIn this paper, we present our work on infinite horizon adaptive dynamic programming problem, which is referred to as isin-adaptive dynamic programming, for discrete-time systems with discount factor 0isin*, which is determined from an isin-optimal cost Visin*, is obtained to approximate the optimal controller. The isin-optimal controller muisin*can always control the state to approach the equilibrium state, while the performance cost is close to the biggest lower bound of all performance costs within an error according to isin. An algorithm for finding the isin-optimal controller is developed and numerical experiments are given to illustrate the performance of the algorithm. Derong Liu 0001 |
IJCNN | 1 |
| 2009 | LMI Based Global Asymptotic Stability Criterion for Recurrent Neural Networks with Infinite Distributed Delays
Zhanshan Wang 0001, Huaguang Zhang, Derong Liu 0001, Jian Feng 0001 |
ISNN (1) | 3 |
| 2009 | A Biologically Inspired Adaptive Nonlinear Control Strategy for Applications to Powertrain ControlabstractIn this paper, an adaptive nonlinear control strategy derived from a biological control system is developed and its applications to the automotive engine are presented. The biological adaptive nonlinear control strategy inspired by the functions of baroreceptor reflex is realized by a parallel controller. The controller consists of a linear controller and a nonlinear controller that interact via a reciprocal lateral inhibitory mechanism. The linear controller design is based on a PID controller, while the nonlinear controller is constructed from neural networks which are updated online. In order to provide superior control performances, the controller must be robust to external unknown disturbances, un-modeled dynamics and plant uncertainties and also be able to perform well under a wide range of operating conditions. In the linear operating region, the linear controller takes control. If the controlled process is far away from the linear regime or is disturbed by the noise, the output of linear controller may be inappropriate, and therefore the nonlinear controller is activated to compensate for the inadequacy of the linear controller in a dynamic environment and in the presence of distances and process parameter variations. These situations can be addressed by adjusting the amount of lateral inhibition and learning the characteristics of the controlled system such that desirable controller outputs are produced in any particular operating region. The novelty of the biological adaptive nonlinear control strategy is that each controller modulates the other controller via the reciprocal lateral inhibitory connections. The good transient performance, computational efficiency, real-time adaptability and superior learning ability are illustrated through extensive numerical simulations for engine torque management driven by the biological adaptive nonlinear control strategy. Hossein Javaherian, Derong Liu 0001 |
SMC | 3 |
| 2009 | Generalized Receding Horizon Control of Fuzzy Systems Based on Numerical Optimization AlgorithmabstractThe optimal control of fuzzy systems with constraints is still an open problem. Our focus concerns the optimal control problem of fuzzy systems derived from receding horizon control (RHC) schemes. We consider methods to numerically compute the value function for general fuzzy systems. The numerical method that is developed using the finite difference with sigmoidal transformation is a stable and convergent algorithm for the Hamilton-Jacobi-Bellman (HJB) equation. An optimization procedure is developed to increase the calculation accuracy with less computation time. A parallel-processing method is employed in the optimization procedure. The optimization results are applied to the controller design of general fuzzy dynamic systems. Employing the principle of conventional RHC schemes, RHC-form controllers are designed for some classes of fuzzy dynamic systems. The basic ideas are as follows. First, the value function is calculated by numerical methods. Then, the value function is used as controller-design parameters to redesign RHC controllers for fuzzy systems, which is motivated by the inverse Lyapunov function design method. It is proven that the closed-loop system is asymptotically stable. An engineering implementation of the controller redesign scheme is discussed. Meanwhile, the parallel-processing framework that can improve the closed-loop performance is also introduced. Chonghui Song, Jinchun Ye, Derong Liu 0001, Qi Kang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2009 | Neural-Network-Based Near-Optimal Control for a Class of Discrete-Time Affine Nonlinear Systems With Control ConstraintsabstractIn this paper, the near-optimal control problem for a class of nonlinear discrete-time systems with control constraints is solved by iterative adaptive dynamic programming algorithm. First, a novel nonquadratic performance functional is introduced to overcome the control constraints, and then an iterative adaptive dynamic programming algorithm is developed to solve the optimal feedback control problem of the original constrained system with convergence analysis. In the present control scheme, there are three neural networks used as parametric structures for facilitating the implementation of the iterative algorithm. Two examples are given to demonstrate the convergence and feasibility of the proposed optimal control scheme. Huaguang Zhang, Derong Liu 0001 |
IEEE Trans. Neural Networks | 3 |
| 2008 | e-Adaptive Dynamic Programming for discrete-time systemsabstractDynamic programming for discrete-time systems is difficult due to the ldquocurse of dimensionalityrdquo: one has to find a series of control actions that must be taken in sequence. This sequence will lead to the optimal performance cost, but the total cost of those actions will be unknown until the end of that sequence. In this paper, we present our work on dynamic programming for discrete-time system, which is referred as epsiv-adaptive dynamic programming. A single controller, epsiv-optimal controller u*(epsiv)is determined from an epsiv-optimal cost J*(epsiv)is given to approximate the optimal controller. The epsiv-optimal controller u*(epsiv)can always control the state to approach to the equilibrium state, while the performance cost is close to the biggest lower bound of all performance costs within an error according to epsiv. Derong Liu 0001 |
IJCNN | 1 |
| 2008 | ADHDP for the pH Value Control in the Clarifying Process of Sugar Cane Juice
Xiaofeng Lin 0002, Shengyong Lei, Chunning Song, Shaojian Song, Derong Liu 0001 |
ISNN (1) | 5 |
| 2008 | Adaptive Dynamic Programming for a Class of Nonlinear Control Systems with General Separable Performance Index
Qinglai Wei, Derong Liu 0001, Huaguang Zhang |
ISNN (2) | 2 |
| 2008 | Neural networks: Algorithms and applications
Derong Liu 0001, Huaguang Zhang, Sanqing Hu |
Neurocomputing | 1 |
| 2008 | Neurodynamic programming: a case study of the traveling salesman problem
Jia Ma, Tao Yang 0011, Zeng-Guang Hou, Min Tan 0001, Derong Liu 0001 |
Neural Comput. Appl. | 5 |
| 2008 | Wavelet Basis Function Neural Networks for Sequential LearningabstractIn this letter, we develop the wavelet basis function neural networks (WBFNNs). It is analogous to radial basis function neural networks (RBFNNs) and to wavelet neural networks (WNNs). In WBFNNs, both the scaling function and the wavelet function of a multiresolution approximation (MRA) are adopted as the basis for approximating functions. A sequential learning algorithm for WBFNNs is presented and compared to the sequential learning algorithm of RBFNNs. Experimental results show that WBFNNs have better generalization property and require shorter training time than RBFNNs. Derong Liu 0001 |
IEEE Trans. Neural Networks | 2 |
| 2008 | A Neural Network Method for Detection of Obstructive Sleep Apnea and Narcolepsy Based on Pupil Size and EEGabstractElectroencephalogram (EEG) is able to indicate states of mental activity ranging from concentrated cognitive efforts to sleepiness. Such mental activity can be reflected by EEG energy. In particular, intrusion of EEG theta wave activity into the beta activity of active wakefulness has been interpreted as ensuing sleepiness. Pupil behavior can also provide information regarding alertness. This paper develops an innovative signal classification method that is capable of differentiating subjects with sleep disorders which cause excessive daytime sleepiness (EDS) from normal control subjects who do not have a sleep disorder based on EEG and pupil size. Subjects with sleep disorders include persons with untreated obstructive sleep apnea (OSA) and narcolepsy. The Yoss pupil staging rule is used to scale levels of wakefulness and at the same time theta energy ratios are calculated from the same 2-s sliding windows by Fourier or wavelet transforms. Then, an artificial neural network (NN) of modified adaptive resonance theory (ART2) is utilized to identify the two groups within a combined group of subjects including those with OSA and healthy controls. This grouping from the NN is then compared with the actual diagnostic classification of subjects as OSA or controls and is found to be 91% accurate in differentiating between the two groups. The same algorithm results in 90% correct differentiation between narcoleptic and control subjects. Derong Liu 0001, Zhongyu Pang, Stephen R. Lloyd |
IEEE Trans. Neural Networks | 1 |
| 2008 | Global Asymptotic Stability of Recurrent Neural Networks With Multiple Time-Varying DelaysabstractIn this paper, several sufficient conditions are established for the global asymptotic stability of recurrent neural networks with multiple time-varying delays. The Lyapunov-Krasovskii stability theory for functional differential equations and the linear matrix inequality (LMI) approach are employed in our investigation. The results are shown to be generalizations of some previously published results and are less conservative than existing results. The present results are also applied to recurrent neural networks with constant time delays. Huaguang Zhang, Zhanshan Wang 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks | 3 |
| 2008 | Robust Stability Analysis for Interval Cohen-Grossberg Neural Networks With Unknown Time-Varying DelaysabstractIn this paper, robust stability problems for interval Cohen-Grossberg neural networks with unknown time-varying delays are investigated. Using linear matrix inequality, M -matrix theory, and Halanay inequality techniques, new sufficient conditions independent of time-varying delays are derived to guarantee the uniqueness and the global robust stability of the equilibrium point of interval Cohen-Grossberg neural networks with time-varying delays. All these results have no restriction on the rate of change of the time-varying delays. Compared to some existing results, these new criteria are less conservative and are more convenient to check. Two numerical examples are used to show the effectiveness of the present results. Huaguang Zhang, Zhanshan Wang 0001, Derong Liu 0001 |
IEEE Trans. Neural Networks | 3 |
| 2008 | Guest Editorial: Special Issue on Adaptive Dynamic Programming and Reinforcement Learning in Feedback ControlabstractThe 18 papers in this special issue focus on adaptive dynamic programming and reinforcement learning in feedback control. Frank L. Lewis, Derong Liu 0001, George G. Lendaris |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Adaptive Critic Learning Techniques for Engine Torque and Air-Fuel Ratio ControlabstractA new approach for engine calibration and control is proposed. In this paper, we present our research results on the implementation of adaptive critic designs for self-learning control of automotive engines. A class of adaptive critic designs that can be classified as (model-free) action-dependent heuristic dynamic programming is used in this research project. The goals of the present learning control design for automotive engines include improved performance, reduced emissions, and maintained optimum performance under various operating conditions. Using the data from a test vehicle with a V8 engine, we developed a neural network model of the engine and neural network controllers based on the idea of approximate dynamic programming to achieve optimal control. We have developed and simulated self-learning neural network controllers for both engine torque (TRQ) and exhaust air-fuel ratio (AFR) control. The goal of TRQ control and AFR control is to track the commanded values. For both control problems, excellent neural network controller transient performance has been achieved. Derong Liu 0001, Hossein Javaherian, Olesia Kovalenko |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | Delay-Dependent Guaranteed Cost Control for Uncertain Stochastic Fuzzy Systems With Multiple Time DelaysabstractThis paper studies the guaranteed cost control problem for a class of uncertain stochastic nonlinear systems with multiple time delays represented by the Takagi-Sugeno fuzzy model with uncertain parameters. By constructing a new stochastic Lyapunov-Krasovskii functional, sufficient conditions for delay-dependent guaranteed cost control are obtained which do not require system transformation or relaxation matrices. Conditions for the existence of an optimal guaranteed cost controller are presented in the linear matrix inequality format. Simulation examples are provided to demonstrate the effectiveness of the proposed approach in this paper. Huaguang Zhang, Yingchun Wang 0003, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2007 | Wavelet Basis Function Neural NetworksabstractIn this paper, a new kind of neural networks for sequential learning is proposed, which are called wavelet basis function neural networks (WBFNNs). They are analogous to radial basis function neural networks (RBFNNs) and to wavelet neural networks (WNNs). In WBFNNs, both the scaling function and the wavelet function of a multiresolution approximation (MRA) are adopted as the basis for approximating functions. A sequential learning algorithm for WBFNNs is presented and compared to the sequential learning algorithm for RBFNNs. Experimental results show that WBFNNs has better generalization property and require shorter training time than RBFNNs. Derong Liu 0001, Zhongyu Pang |
IJCNN | 2 |
| 2007 | Temperature Control in Precalcinator with Dual Heuristic Dynamic ProgrammingabstractIn cement manufacturing, pre-calcining process (PCP) contains an added precalcinator between the preheater and the kiln. Since raw meal decomposition rate in the kiln is over 90%, optimal control is necessary for keeping the temperature stable to ensure normal decomposition. Due to the diversity of raw meal composition and the consistent thermal transformation, temperature control in the precalcinator is a complex nonlinear control problem. Many factors may influence it, and there are couplings between various factors and uncertainties. It is difficult to solve the dynamic stability control problem for the temperature in precalcinator. This paper addresses temperature control in the precalcinator using dual heuristic dynamic programming (DHP). The approach is mainly composed of three parts: model network, action network and critic network. All three parts use neural networks (NN) to simulate the temperature trend. The self-learning and adaptive abilities were analyzed in the MATLAB environment and simulation results demonstrate the advantage of adaptation and optimization solutions compared to conventional control. Xiaofeng Lin 0002, Derong Liu 0001 |
IJCNN | 3 |
| 2007 | Neural Network Strategy for Sampling of Particle Filters on the Tracking ProblemabstractSequential Monte Carlo methods, namely particle filters, are popular statistic techniques for sampling sequentially from a complex probability distribution. Sampling is a key step for particle filters and has vital effects on simulation results. Since degeneracy of particles in samples sometimes is very severe, there exist only a few particles with significant weights. Thus the sample diversity is reduced significantly so that only a few particles are used to represent the corresponding probability distribution. Therefore, resampling has to be used very often during the whole procedure. This paper addresses a new method which can avoid the phenomenon of particle degeneracy. A backpropagation neural network is used to adjust low weight particles in order to increase their weights and particles with high weights may be split into two small ones if needed. Our simulation results on a typical tracking problem show that not only the phenomenon of particle degeneracy is effectively avoided but also tracking results are much better than those of the traditional particle filter. Zhongyu Pang, Derong Liu 0001, Zhuo Wang 0003 |
IJCNN | 2 |
| 2007 | Global Asymptotic Stability of Recurrent Neural Networks with Time Varying DelaysabstractIn this paper, two sufficient conditions are established for the global asymptotic stability of recurrent neural networks with multiple time varying delays. The Lyapunov-Krasovskii stability theory for functional differential equations and the linear matrix inequality approach are employed in our investigation. Our results are shown to be generalizations of some previously published results and are less conservative than existing results. The present results are also applicable to recurrent neural networks with constant time delays. Huanxin Guan, Huaguang Zhang, Zhanshan Wang 0001, Derong Liu 0001 |
ISCAS | 4 |
| 2007 | A Brief Overview of the Complex Biological and Engineering NetworksabstractOver the last few decades, complex networks have been intensively studied throughout many fields of science, especially in biological and engineering sciences. This paper briefly reviews the main advances in the complex biological and engineering networks, aiming to bridge the gap between the complex biological and engineering networks. Biologists pay more attention to the mechanisms and local dynamics of individuals, however, engineers are more interesting in the global dynamical behaviors. It is the time for the biologists and engineers to work together for better understanding the complex networks. Jinhu Lü 0001, Derong Liu 0001 |
ISCAS | 2 |
| 2007 | Classification of Obstructive Sleep Apnea by Neural Networks
Zhongyu Pang, Derong Liu 0001, Stephen R. Lloyd |
ISNN (2) | 2 |
| 2007 | On-Line Learning Control for Discrete Nonlinear Systems Via an Improved ADDHP Method
Huaguang Zhang, Qinglai Wei, Derong Liu 0001 |
ISNN (1) | 3 |
| 2007 | Fuzzy H∞ Filter Design for a Class of Nonlinear Discrete-Time Systems With Multiple Time DelaysabstractThis paper studies the fuzzy Hinfinfilter design problem for signal estimation of nonlinear discrete-time systems with multiple time delays and unknown bounded disturbances. First, the Takagi-Sugeno (T-S) fuzzy model is used to represent the state-space model of nonlinear discrete-time systems with time delays. Next, we design a stable fuzzy Hinfinfilter based on the T-S fuzzy model, which guarantees asymptotic stability and a prescribed Hinfinindex for the filtering error system, irrespective of the time delays and uncertain disturbances. A sufficient condition for the existence of such a filter is established by using the linear matrix inequality (LMI) approach. The proposed LMI problem can be efficiently solved with global convergence guarantee using convex optimization techniques such as the interior point algorithm. Simulation examples are provided to illustrate the design procedure of the present method. Huaguang Zhang, Shuxian Lun, Derong Liu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2007 | Guest Editorial: Networking, Sensing, and Control for Networked Control Systems: Architectures, Algorithms, and ApplicationsabstractThe eight papers in this special issue focus on networking, sensing, and control for networked control systems (NCS). The papers, which are briefly summarized, examine the architectures, algorithms and applications of NCS. Fei-Yue Wang 0001, Derong Liu 0001, Simon X. Yang, Li Li 0013 |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2006 | Decentralized H-infinity Control of Fuzzy Large-scale SystemsabstractThis paper investigates decentralized Hinfincontroller design for the fuzzy large-scale systems, which are composed of a number of T-S fuzzy subsystems with interconnections. First, according to the Lyapunov direct method and the decentralized control theory of large-scale systems, a sufficient condition in the form of LMIs which guarantees the existence of state feedback Hinfincontrol for fuzzy large-scale systems is proposed. Second, if states are not all available, decentralized fuzzy observers are proposed to estimate the states of each subsystem for decentralized control. Consequently, the fuzzy observer-based state feedback decentralized fuzzy controllers are obtained to solve Hinfincontrol problem. Finally, a numerical example is given to demonstrate the fuzzy observer-based controller design procedure and its effectiveness. Huaguang Zhang, Derong Liu 0001 |
FUZZ-IEEE | 3 |
| 2006 | Automotive Engine Torque and Air-Fuel Ratio Control Using Dual Heuristic Dynamic ProgrammingabstractIn this paper, we present our work on self-learning engine control using "derivative" adaptive critics. In the derivative version of adaptive critic designs, the derivative of the cost function of dynamic programming with respect to the state of the process is estimated. As the goal of dynamic programming is to maximize the overall cost function, direct estimation of the cost function derivatives will naturally lead to better algorithm performance. In our previous studies of the regular version of adaptive critic techniques, a model of the process for controller training was not necessary. However, for the implementation of the "derivative" version of adaptive critic designs, a process model in the form of neural network is required. In this way, the derivative information can be propagated back to the controller in order to update the controller parameters. The objectives of the present learning control design for automotive engines are to improve performance, reduce emissions and maintain optimum performance under various operating conditions. Using data from the test vehicle equipped with a 5.3L V8 engine, we built a neural network model of the engine. The model is then used in the development of self-learning neural network controllers based on the idea of approximate dynamic programming to achieve optimal control for both engine torque and exhaust air-fuel ratio control. The goals of the engine torque and air-fuel ratio control are to track the commanded torque and to regulate the air-fuel ratio at specified set-points, respectively. Hossein Javaherian, Derong Liu 0001, Olesia Kovalenko |
IJCNN | 2 |
| 2006 | Sequential blind extraction of instantaneous mixtures with arbitrary rankabstractIn this paper, we present new extractability conditions for blind source extraction of linear instantaneous mixtures. Two general conditions for source extraction of arbitrarily mixed nonzero sources are presented. A sufficient condition is provided to guarantee that sequential extraction can be continued. We also show an important property for inseparable mixtures; that is, any two extracted signals involving the same sources are proportional to each other. For sub-Gaussian or sup-Gaussian source signals with only mutual independence, cost functions based on fourth-order cumulants are introduced to sequentially extract all separable single sources and all inseparable mixtures. By minimizing the cost functions, gradient-based methods are developed. Our algorithms are guaranteed to converge. Finally, simulation results show the operation characteristics and the effectiveness of our methods Sanqing Hu, Derong Liu 0001, Jun Wang 0002 |
ISCAS | 2 |
| 2006 | Global exponential stability of generalized neural networks with time-varying delaysabstractIn this paper, we essentially drop the requirement of Lipschitz condition on the activation functions. Only using physical parameters of neural networks, we propose some new criteria concerning global exponential stability of generalized neural networks with time-varying delays. Since these new criteria do not require the activation functions to be differentiate, bounded or monotone nondecreasing and the connection weight matrices to be symmetric, they are mild and more general than previously known criteria Gang Wang 0026, Huaguang Zhang, Derong Liu 0001 |
ISCAS | 3 |
| 2006 | A New Fuzzy Identification Method Based on Adaptive Critic Designs
Huaguang Zhang, Derong Liu 0001 |
ISNN (1) | 3 |
| 2006 | Motif discoveries in unaligned molecular sequences using self-organizing neural networksabstractIn this paper, we study the problem of motif discoveries in unaligned DNA and protein sequences. The problem of motif identification in DNA and protein sequences has been studied for many years in the literature. Major hurdles at this point include computational complexity and reliability of the search algorithms. We propose a self-organizing neural network structure for solving the problem of motif identification in DNA and protein sequences. Our network contains several layers, with each layer performing classifications at different levels. The top layer divides the input space into a small number of regions and the bottom layer classifies all input patterns into motifs and nonmotif patterns. Depending on the number of input patterns to be classified, several layers between the top layer and the bottom layer are needed to perform intermediate classifications. We maintain a low computational complexity through the use of the layered structure so that each pattern's classification is performed with respect to a small subspace of the whole input space. Our self-organizing neural network will grow as needed (e.g., when more motif patterns are classified). It will give the same amount of attention to each input pattern and will not omit any potential motif patterns. Finally, simulation results show that our algorithm outperforms existing algorithms in certain aspects. In particular, simulation results show that our algorithm can identify motifs with more mutations than existing algorithms. Our algorithm works well for long DNA sequences as well. Derong Liu 0001, Xiaoxu Xiong, Bhaskar DasGupta, Huaguang Zhang |
IEEE Trans. Neural Networks | 1 |
| 2005 | A self-organizing neural network approach for the identification of motifs with insertions and deletions in protein sequencesabstractCurrent popular algorithms of motif identification in protein sequences face two difficulties, large computation and insertions and deletions of letters. In this paper, we provide a new strategy that solves this problem in a more efficient and effective way. We build a self-organizing neural network with multiple levels of subnetworks to classify subsequences obtained from the protein sequences. We maintain a low computational complexity through the use of this multi-level structure so that the classification of each subsequence is performed with respect to a small subspace of the whole input space. The new definition of pairwise distance between motif patterns provided in this paper can deal with more insertions/deletions allowed in a motif than other algorithms. In the simulation result, our algorithm significantly outperforms existing algorithms in both accuracy and reliability aspects. Xiaoxu Xiong, Derong Liu 0001, Huaguang Zhang |
IJCNN | 2 |
| 2005 | Robust Stability Analysis of a Class of Hopfield Neural Networks with Multiple Delays
Huaguang Zhang, Ce Ji, Derong Liu 0001 |
ISNN (1) | 3 |
| 2005 | Exponential Stability Analysis of Neural Networks with Multiple Time Delays
Huaguang Zhang, Zhanshan Wang 0001, Derong Liu 0001 |
ISNN (1) | 3 |
| 2005 | On the global output convergence of a class of recurrent neural networks with time-varying inputs
Sanqing Hu, Derong Liu 0001 |
Neural Networks | 2 |
| 2005 | Identification of motifs with insertions and deletions in protein sequences using self-organizing neural networks
Derong Liu 0001, Xiaoxu Xiong, Zeng-Guang Hou, Bhaskar DasGupta |
Neural Networks | 1 |
| 2005 | A self-learning call admission control scheme for CDMA cellular networksabstractIn the present paper, a call admission control scheme that can learn from the network environment and user behavior is developed for code division multiple access (CDMA) cellular networks that handle both voice and data services. The idea is built upon a novel learning control architecture with only a single module instead of two or three modules in adaptive critic designs (ACDs). The use of adaptive critic approach for call admission control in wireless cellular networks is new. The call admission controller can perform learning in real-time as well as in offline environments and the controller improves its performance as it gains more experience. Another important contribution in the present work is the choice of utility function for the present self-learning control approach which makes the present learning process much more efficient than existing learning control methods. The performance of our algorithm will be shown through computer simulation and compared with existing algorithms. Derong Liu 0001, Huaguang Zhang |
IEEE Trans. Neural Networks | 1 |
| 2005 | Multiuser detection using the Taguchi method for DS-CDMA systemsabstractWe study multiuser detection for direct-sequence code-division multiple-access systems in a multipath environment. Systems with unknown channel information are considered and the well-known result for maximum likelihood multiuser detector is directly used in our work. Due to the high computational cost of the maximum likelihood detector, most existing works have investigated simplified, linearized, and/or suboptimal solutions that have less computational requirements. In our approach, we use the Taguchi method that involves the use of orthogonal arrays in estimating the gradient of the likelihood function. The Taguchi method has been widely used in experimental designs for problems with multiple parameters where the optimization of a cost function is required. In this work, we choose the likelihood function as the cost function in the Taguchi method. The use of the Taguchi method for multiuser detection is a novel idea, and it leads to efficient algorithms that can find a satisfactory solution by maximizing the likelihood function in a small number of iterations. One of the advantages of the present Taguchi method is that it is blind since no channel estimation is required to detect the transmitted data, which is not the case in many existing methods. Simulation results show that the Taguchi multiuser detector significantly outperforms the conventional receivers, is insensitive to initial values of parameters, and has performance close to that of minimum mean square error detectors and decorrelating detectors. In addition, our algorithm is suitable for parallel implementations. Derong Liu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2004 | Applications of fuzzy decision-making in pipeline leak localizationabstractMonitoring oil transporting pipelines is an important task for economical and safe operation, loss prevention, and environmental protection from crude oil emission. The leak detection of oil pipelines, therefore, plays a key role in the overall integrity, monitoring a pipeline system. This paper proposes a fuzzy decision-making approach to oil pipeline leak localization. The two main methods, pressure gradient localization and negative pressure wave localization, are combined with fuzzy logical decision-making to form a novel fault diagnosis scheme. The combination scheme can improve the precision of localization. An application example, a 14 km long oil pipeline leak detection and localization is illustrated, and the effectiveness of the proposed approach is demonstrated using the practical results. Jian Feng 0001, Huaguang Zhang, Derong Liu 0001 |
FUZZ-IEEE | 3 |
| 2004 | Roughness of fuzzy sets based on two new operatorsabstractIn this paper, two new operators are introduced for the rough set theory. Using them, two inequalities well-known in the rough set theory can now be modified to become equalities. With this change, no information may be lost in the new expressions. Hence, many properties in the rough set theory can be improved and in particular, the union, the intersection, and the complement operations can be redefined based on the two equalities. Finally, the roughness properties of fuzzy sets are analyzed using the new operations. Hongli Liang, Huaguang Zhang, Derong Liu 0001 |
FUZZ-IEEE | 3 |
| 2004 | A generalized fuzzy hyperbolic modeling and control schemeabstractA new generalized fuzzy hyperbolic model (GFHM) is proposed, which is proved to be a universal approximator. GFHM can be used as identifier for nonlinear dynamic systems and the back-propagation training algorithm is given. The feature of GFHM is that as the number of input variables or (and) fuzzy subsets increases, the number of the unknown parameters of GFHM would increase linearly. Finally, the adaptive fuzzy control scheme is presented, which can guarantee that the closed-loop system is globally asymptotically stable. The simulation results show the applicability of the modeling scheme and the effectiveness of the proposed adaptive control scheme. Huaguang Zhang, Derong Liu 0001 |
FUZZ-IEEE | 3 |
| 2004 | A new condition for the global robust exponential periodicity of interval neural networks with delaysabstractWe study the robust exponential periodicity of a class of interval-delayed neural networks. A new condition ensuring the existence, uniqueness and global robust exponential stability of the periodic solution of interval-delayed neural networks with periodic input is established. Changyin Sun 0001, Derong Liu 0001, Chun-Bo Feng |
IJCNN | 2 |
| 2004 | Two new operators in rough set theory with applications to fuzzy sets
Huaguang Zhang, Hongli Liang, Derong Liu 0001 |
Inf. Sci. | 3 |
| 2004 | Call Admission Policies Based on Calculated Power Control Setpoints in SIR-Based Power-Controlled DS-CDMA Cellular Networks
Derong Liu 0001, Sanqing Hu |
Wirel. Networks | 1 |
| 2003 | A self-learning adaptive critic approach for call admission control in wireless cellular networksabstractWe apply in the present paper adaptive critic designs to the call admission control problem in CDMA cellular network. A novel learning control architecture of adaptive critic designs is used in our study. The call admission controller performs learning in real-time as well as in off-line environments and the controller improves its performance through continuous learning. Derong Liu 0001 |
ICC | 1 |
| 2003 | New results on exponential periodicity of delayed neural networksabstractIn this paper, exponential periodicity of neural networks with delays is investigated. Without assuming the boundedness and differentiability of the activation functions, some new sufficient conditions ensuring existence and uniqueness of periodic solution for a general class of neural networks with delays are obtained. This work gives some improvements to previous ones. Changyin Sun 0001, Derong Liu 0001, Chun-Bo Feng |
IJCNN | 2 |
| 2002 | Call admission control algorithms for DS-CDMA cellular networks supporting multimedia servicesabstractWe develop call admission control algorithms for DS-CDMA cellular networks that support multimedia services. When a new call (or a handoff call) arrives at a base station requesting admission, our algorithms calculate the desired power levels for the new call and all existing calls. These calculations are based on the interference received at the base station and the desired quality of service target for each call. If the desired power levels to be received at the base station for some calls are larger than the maximum allowable power limits, the admission request is rejected. Otherwise, the admission request is granted. When higher priority is desired for handoff calls, we allow different thresholds for new calls and handoff calls. Our algorithms are presented in a basic form and in an adaptive form. Derong Liu 0001 |
ICME (1) | 1 |
| 2002 | N-bit parity neural networks: new solutions based on linear programming
Derong Liu 0001, Myron E. Hohil, Stanley H. Smith |
Neurocomputing | 1 |
| 2001 | A heuristic approach for measurement-based admission control with variable-size windowabstractA measurement-based admission control algorithm using a variable-size window is developed. The size of the measurement window is varied according to the estimate of the variance of the traffic in the measurement window. We use a small measurement window when the traffic is smooth (small variations) and we use large measurement window when the traffic is bursty (large variations). Our simulation studies provide results that complement two existing works (see Breslau, L. et al., 2000; Lee, S.K. and Song, J.S., 1998). On the one hand, we conclude that when the desired utilization rate is low, measurement-based admission control algorithms using a variable-size measurement window can improve the performance over algorithms using a fixed-size window. On the other hand, when the desired utilization rate is high, measurement-based admission control algorithms using a variable-size window and using a fixed-size window provide nearly identical performance results. Derong Liu 0001 |
GLOBECOM | 1 |
| 2001 | Nested auto-regressive processes for MPEG-encoded video traffic modelingabstractThis paper presents a new traffic model for MPEG-encoded video sequences. The hybrid gamma/Pareto distribution is used for all three types of frames in MPEG-encoded video sequences, and the present model takes scene changes into account. The autocorrelation structure is modeled using two second-order auto-regressive (AR) processes nested with each other. One AR process is used to generate the mean frame size of the scenes to model the long-range dependence, and another AR process is used to generate the fluctuations within the scenes to model the short range dependence. The parameters of the AR processes are estimated from measurements of empirical video sequences. Simulation results show that the present model captures the autocorrelation structure in the empirical traces at both small and large lags. The MPEG traffic model presented in this paper is used to predict the queueing performance of single and multiplexed MPEG video sequences at an asynchronous transfer mode multiplexer. Comparison study shows that the present model provides accurate prediction for quality of service measures, such as cell-loss ratio under different traffic loads and various buffer sizes. Derong Liu 0001, Endre I. Sára |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2000 | A new traffic model for MPEG encoded videos in ATM networksabstractThis paper presents a new traffic model for MPEG encoded video sequences. Two second-order autoregressive (AR) processes are used to model the autocorrelation structure. One AR process is used to generate the mean frame size of the scenes to model the long range dependence and another AR process is used to generate the fluctuations within the scenes to model the short range dependence. The scene length distribution is fitted using a geometric distribution. The first AR process is therefore "stretched" unevenly according to the geometric distribution to generate the mean frame size sequence. The two AR processes are not simply superposed; instead, they are nested with each other. The parameters of the AR processes are estimated from measurements of empirical video sequences. Simulation results show that the present model captures the autocorrelation structure in the empirical traces for both small and large lags. The MPEG traffic model presented in this paper is used to predict the queueing performance of single and multiplexed MPEG video sequences at an asynchronous transfer mode multiplexer. Comparison study shows that the present model provides accurate prediction for quality of service measures such as cell loss ratio under different traffic loads and various buffer sizes. Derong Liu 0001, Endre I. Sára |
ICCCN | 1 |
| 2000 | Synthesis approach for bidirectional associative memories based on the perceptron training algorithm
Ismail Salih, Stanley H. Smith, Derong Liu 0001 |
Neurocomputing | 3 |
| 2000 | Neural network-based model reference adaptive control systemabstractIn this paper, an approach to model reference adaptive control based on neural networks is proposed and analyzed for a class of first-order continuous-time nonlinear dynamical systems. The controller structure can employ either a radial basis function network or a feedforward neural network to compensate adaptively the nonlinearities in the plant. A stable controller-parameter adjustment mechanism, which is determined using the Lyapunov theory, is constructed using a sigma-modification-type updating law. The evaluation of control error in terms of the neural network learning error is performed. That is, the control error converges asymptotically to a neighborhood of zero, whose size is evaluated and depends on the approximation error of the neural network. In the design and analysis of neural network-based control systems, it is important to take into account the neural network learning error and its influence on the control error of the plant. Simulation results showing the feasibility and performance of the proposed approach are given. Hector D. Patiño, Derong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1999 | An adaptive critic approach for self-learning stock tradingabstractThe paper describes a stock trading system with self learning capability using adaptive critic designs. The same approach can be formulated for other financial applications such as trading of bonds, options, futures, commodities, and the like. Derong Liu 0001, Yan Kong, Edward G. Luxford |
CIFEr | 1 |
| 1999 | Solving the N-bit parity problem using neural networks
Myron E. Hohil, Derong Liu 0001, Stanley H. Smith |
Neural Networks | 2 |
| 1997 | A new synthesis approach for feedback neural networks based on the perceptron training algorithmabstractIn this paper, a new synthesis approach is developed for associative memories based on the perceptron training algorithm. The design (synthesis) problem of feedback neural networks for associative memories is formulated as a set of linear inequalities such that the use of perceptron training is evident. The perceptron training in the synthesis algorithms is guaranteed to converge for the design of neural networks without any constraints on the connection matrix. For neural networks with constraints on the diagonal elements of the connection matrix, results concerning the properties of such networks and concerning the existence of such a network design are established. For neural networks with sparsity and/or symmetry constraints on the connection matrix, design algorithms are presented. Applications of the present synthesis approach to the design of associative memories realized by means of other feedback neural network models are studied. To demonstrate the applicability of the present results and to compare the present synthesis approach with existing design methods, specific examples are considered. Derong Liu 0001, Zanjun Lu |
IEEE Trans. Neural Networks | 1 |
| 1996 | Robustness analysis and design of a class of neural networks with sparse interconnecting structureabstractWe first conduct an analysis of the robustness properties of a class of neural networks with applications to associative memories. Specifically, for a network with nominal parameters which stores a set of desired bipolar memories, we establish sufficient conditions under which the same set of bipolar memories is also stored in the network with perturbed parameters. This result enables us to establish a synthesis procedure for neural networks whose stored memories are invariant under perturbations. Our synthesis procedure is capable of generating artificial neural networks with prespecified sparsity constraints (on the interconnecting structure) and with nonsymmetric and symmetric interconnection matrices. To demonstrate the applicability of the present results, we consider several specific examples. Derong Liu 0001, Anthony N. Michel |
Neurocomputing | 1 |
| 1993 | Asympototic stability of two-dimensional digital filters with overflow nonlinearities
Derong Liu 0001, Anthony N. Michel |
ISCAS | 1 |
| 1993 | Analysis and synthesis of a class of neural networks with sparse interconnections
Derong Liu 0001, Anthony N. Michel |
ISCAS | 1 |