Ding Wang 0001

dblp:99/4292-1 · DBLP profile ↗
← Back
214ranked-venue papers
70as first author
121since 2021 · last 2026
0000-0002-7149-5712ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 139 · 42 first-author · 71 since 2021Applied, interdisciplinary, general and emerging computing · 33 · 13 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 29 · 12 first-author · 16 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval
abstract
Retrieval-Augmented Generation (RAG) plays a crucial role in grounding Large Language Models by leveraging external knowledge, whereas the effectiveness is often compromised by the retrieval of contextually flawed or incomplete information. To address this, knowledge graph-based RAG methods have evolved towards hierarchical structures, organizing knowledge into multi-level summaries. However, these approaches still suffer from two critical, unaddressed challenges: high-level conceptual summaries exist as disconnected ``semantic islands'', lacking the explicit relations needed for cross-community reasoning; and the retrieval process itself remains structurally unaware, often degenerating into an inefficient flat search that fails to exploit the graph's rich topology. To overcome these limitations, we introduce LeanRAG, a framework that features a deeply collaborative design combining knowledge aggregation and retrieval strategies. LeanRAG first employs a novel semantic aggregation algorithm that forms entity clusters and constructs new explicit relations among aggregation-level summaries, creating a fully navigable semantic network. Then, a bottom-up, structure-guided retrieval strategy anchors queries to the most relevant fine-grained entities and then systematically traverses the graph's semantic pathways to gather concise yet contextually comprehensive evidence sets. The LeanRAG can mitigate the substantial overhead associated with path retrieval on graphs and minimize redundant information retrieval. Extensive experiments on four challenging QA benchmarks with different domains demonstrate that LeanRAG significantly outperforms existing methods in response quality while reducing 46% retrieval redundancy.
Yaoze Zhang, Pinlong Cai, Guohang Yan, Song Mao, Ding Wang 0001, Botian Shi
AAAI7
2026 ReBrain: Brain MRI Reconstruction from Sparse CT Slice via Retrieval-Augmented Diffusion
abstract
Magnetic Resonance Imaging (MRI) plays a crucial role in brain disease diagnosis, but it is not always feasible for certain patients due to physical or clinical constraints. Recent studies attempt to synthesize MRI from Computed Tomography (CT) scans; however, low-dose protocols often result in highly sparse CT volumes with poor throughplane resolution, making accurate reconstruction of the full brain MRI volume particularly challenging. To address this, we propose ReBrain, a retrieval-augmented diffusion framework for brain MRI reconstruction. Given any 3D CT scan with limited slices, we first employ a Brownian Bridge Diffusion Model (BBDM) to synthesize MRI slices along the 2D dimension. Simultaneously, we retrieve structurally and pathologically similar CT slices from a comprehensive prior database via a fine-tuned retrieval model. These retrieved slices are used as references, incorporated through a ControlNet branch to guide the generation of intermediate MRI slices and ensure structural continuity. We further account for rare retrieval failures when the database lacks suitable references and apply spherical linear interpolation to provide supplementary guidance. Extensive experiments on SynthRAD2023 and BraTS demonstrate that ReBrain achieves state-of-the-art performance in cross-modal reconstruction under sparse conditions.
Weihua Cheng, Yujin Kang, Yirong Chen, Ding Wang 0001, Guosun Zeng
WACV6
2026 Online Q-learning design augmented with fuzzy-adaptive parameter adjustment for wastewater treatment processes
Ding Wang 0001
Neurocomputing2
2026 Data-driven fast optimal regulation with adjustable convergence for asymmetric constrained systems
Ding Wang 0001, Junfei Qiao 0001
Neurocomputing2
2026 Adaptive evolutionary inverse reinforcement learning for large-scale interconnected systems
Ding Wang 0001, Jiangyu Wang, Junfei Qiao 0001
Neurocomputing2
2026 Neural-network-based robust critic learning control with advanced value iteration for continuous-time dynamical systems
Ao Liu 0012, Ding Wang 0001
Neural Networks2
2026 Adaptive critic designs for event-based multi-agent systems with asymmetric constraints
Wenting Yan, Ding Wang 0001, Xinrui Ma, Junfei Qiao 0001
Neural Networks2
2026 Evolution-Guided Q-Learning for Optimal Regulation of Unknown Continuous-Time Systems
abstract
In practical applications, it is challenging to obtain accurate system models, which limits the applicability of model-based control methods. To address this issue, an evolution-guided iterative Q-learning (EIQL) approach is developed in this paper to solve the optimal regulation problem of continuous-time (CT) systems. Incorporating the data-driven mechanism, the offline data-set is utilized for learning, eliminating dependence on the exact system model. Within the Q-learning structure, an actor-critic framework is employed to facilitate policy improvement and Q-function updating. Specifically, an improved particle swarm optimization (PSO) algorithm is developed to mitigate gradient vanishing and overcome the premature convergence issue of standard PSO. Additionally, theoretical analyses are conducted to establish monotonicity and convergence of the designed Q-function. Finally, the constructed EIQL method is validated on real-world physical systems, demonstrating its effectiveness and advantages over conventional approaches.
Ding Wang 0001, Qinna Hu, Zeqiang Yuan, Ao Liu 0012, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.1
2026 Incremental Q-Learning for Data-Driven Adaptive Critic Control of Wastewater Treatment Processes With State Constraints
abstract
With the widespread adoption of massive multiple-input multiple-output (MIMO) technology, itThe wastewater treatment process (WWTP) can effectively convert wastewater into reusable water, with dissolved oxygen (DO) concentration playing a crucial role in treatment efficiency. To achieve online control of DO concentration, a data-driven incrementalQ-learning (IQL) method with utility reshaping is proposed. The core idea is to integrate the incremental information of control inputs, the norm value of tracking errors, and a control barrier function into the utility function. First, the IQL method designs an online controller by utilizing tracking errors and the incremental information of control inputs, which reduces the approximation pressure on neural networks caused by excessively large control inputs. Second, the IQL leverages the norm value of tracking errors to influence the magnitude of the incremental control inputs, thereby enhancing the adaptive capability and online stability. Third, an adjustable control barrier function is employed to ensure that the tracking errors do not exceed the predefined constraints, thus meeting the specified tracking accuracy requirements. Fourth, a stability criterion related to learning rates is established to demonstrate the convergence of neural network weights. Finally, the superior control performance of the IQL method is validated via simulation results under two distinct weather conditions.
Ding Wang 0001, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.2
2026 Adaptive Tracking Control for Nonlinear Systems Under False Data Injection Attacks via Intermittent State Triggering
abstract
This article investigates the adaptive tracking control strategy for a class of nonlinear systems subjected to false data injection (FDI) attacks, incorporating an improved event-triggered mechanism. A significant breakthrough of this study lies in the challenge that, after FDI, not all states of the system can be utilized for stability design, thereby making it more complicated to achieve tracking control. This article eliminates the restrictive assumption, required in some existing results, that the attack signal at the first step must be known. Instead, we propose to estimate the tracking error directly. This approach not only facilitates the tracking control of nonlinear systems but also enhances the generalizability and practical applicability of the solution. To conserve system resources, an improved event-triggered condition is proposed that utilizes the triggered attacked-output. Consequently, the controllers and adaptive laws are implemented using the sampled states rather than continuous real-time states, thereby minimizing unnecessary computations and communications. By constructing Lyapunov functions, the proposed control strategy ensures that all signals in the closed-loop system are globally bounded. Finally, the simulation results are displayed to validate the effectiveness of the proposed control strategy.
Ben Niu 0003, Xudong Zhao 0001, Ding Wang 0001
IEEE Trans. Cybern.4
2026 Fixed-Time Adaptive Control for Uncertain High-Order Nonlinear CPSs Against Dual-Channel Attacks
abstract
This article proposes an adaptive fixed-time control strategy for uncertain high-order nonlinear cyber-physical system subject to deception attacks on both the sensor-to-controller (S-C) and controller-to-actuator (C-A) communication channels. First, to mitigate the impact of dual-channel attacks, mathematical tools are employed to decouple the high-order term induced by C-A channel attacks, and a robust controller is designed based on the compromised state information. Second, novel Lyapunov functions and an adaptive mechanism are constructed to transform nonlinear uncertainties into a linearly parameterized form with unknown parameters. This transformation effectively compensates for unknown control coefficients and mitigates the severe nonlinear growth caused by high-order dynamics under attacks. Third, a control strategy independent of the system's powers is developed, which relaxes the requirement for prior knowledge of system powers and avoids singularities. Finally, the effectiveness and feasibility of the proposed strategy are confirmed via simulation results.
Haiqing Huang, Ben Niu 0003, Xudong Zhao 0001, Ding Wang 0001
IEEE Trans. Cybern.6
2026 Hybrid Event-Triggered Tracking Control With Critic Learning for Nonlinear Networked Systems
abstract
In this article, a novel hybrid event-triggered (ET) control framework is constructed based on the adaptive critic technique, aiming to address the optimal tracking issue of discrete-time nonlinear networked control systems. First, an augmented plant is created by combining the system state with the reference trajectory, transforming the optimal tracking control design into the optimal regulation problem of the reconstructed nonlinear error system. Subsequently, to conserve communication network resources and ensure the stability of the error system, a hybrid ET mechanism is developed to determine a constant interval for event silence. This approach not only alleviates the limited network bandwidth but also eliminates the need for continuous evaluation of triggering conditions, as seen in traditional event-based methods. Regarding algorithm implementation, the model, critic, and action networks are established to execute the online adaptive critic algorithm, which allows the tracking control policy to be adjusted in real-time to reach the optimal level. Finally, an experimental plant with nonlinear characteristics is presented to illustrate the overall performance of the proposed online tracking control method with the hybrid ET mechanism.
Ding Wang 0001, Lingzhi Hu, Dongbin Zhao
IEEE Trans. Cybern.1
2026 Event-Triggered Safe Critic Learning Control via Swarm Intelligence Optimization
abstract
This article develops an event-triggered safe critic learning control (ESCLC) algorithm for nonlinear systems subject to asymmetric state constraints by integrating a safe critic learning control (SCLC) framework with an event-triggering mechanism. The SCLC algorithm innovatively incorporates control barrier functions into the safe value function design, addressing the challenge of deriving optimal control policies that guarantee system safety. Convergence of the SCLC algorithm is rigorously established within the value iteration framework, along with a criterion for assessing the admissibility of control policies. To enhance the application value of the algorithm in resource-constrained scenarios, an event-triggering mechanism is incorporated into the SCLC framework, yielding the ESCLC algorithm. The resulting closed-loop system under the ESCLC algorithm is proved to be asymptotically stable, and an upper bound on the actual value function is derived to ensure bounded performance degradation. In addition, a policy improvement method based on particle swarm optimization is designed that eliminates dependence on the system control matrix. Finally, the effectiveness of the ESCLC algorithm is verified through simulation experiments on a torsion pendulum system and a ball-and-beam system.
Ding Wang 0001, Xin Li 0055, Wenjing Li 0004, Junfei Qiao 0001
IEEE Trans. Cybern.1
2026 Accelerated Intelligent Critic Tracking Predictive Control With Data Experience Replay for Unknown Nonlinear Systems
abstract
In this article, the accelerated intelligent critic tracking predictive control with data experience replay (AICTPC-DER) framework is constructed to address the trajectory tracking problem of the nonlinear systems with unknown dynamics. The receding optimization mechanism of model predictive control and the intelligent critic scheme are deeply integrated to realize real-time optimization of online policies. First, the time-series data of the unknown system is collected to establish a deep neural network as the prediction model. Afterward, in order to improve the efficiency of solving optimization problems online, the accelerated critic architecture with experience replay via collecting tracking error data is established based on the conventional adaptive critic designs. Simultaneously, the theoretical properties of the AICTPC-DER algorithm are comprehensively analyzed. Finally, a large number of simulation results verify the effectiveness and progressiveness of the AICTPC-DER algorithm in solving tracking problems, among which the advantages of the accelerated factor and the DEP mechanism are apparent. From the comparative experiments, it can be seen that the developed algorithm exhibits superior control performance.
Ding Wang 0001, Peng Xin, Ao Liu 0012, Junfei Qiao 0001
IEEE Trans. Cybern.1
2026 Data-Driven Adaptive Critic Designs for Hybrid Lifelong Learning in Wastewater Treatment Processes
abstract
Wastewater treatment yields significant societal benefits in resource recycling, economic development, and public health. Dissolved oxygen (DO) concentration during the wastewater treatment process serves as a critical indicator for assessing effluent quality. Therefore, maintaining DO within an appropriate range is essential. This study proposes a data-driven tracking controller based on an action-dependent heuristic dynamic programming (ADHDP) approach incorporating lifelong learning (LL) to achieve precise DO concentration tracking. First, the LL-ADHDP controller, acting as an auxiliary controller, is integrated with a prior-knowledge-based controller to achieve model-free tracking control. The online LL-ADHDP approach enhances the approximation accuracy of both the critic and action networks. Second, integrating the LL mechanism into these networks mitigates catastrophic forgetting and improves overall robustness. Third, the method is applied to the benchmark simulation model no. 1. Experimental results demonstrate its superior tracking performance. Finally, simulations with diverse reference trajectories confirm the good dynamic performance and effectively reduce the tracking error of the proposed LL-ADHDP method.
Zhaoyu Ji, Xiang Liu 0020, Ding Wang 0001, Menghua Li, Junfei Qiao 0001
IEEE Trans. Ind. Informatics3
2026 Multiagent Adaptive Critic Control With Expert Knowledge for Wastewater Treatment Plants
abstract
In this article, a multiagent adaptive critic control algorithm is developed for the multivariable control problem of wastewater treatment plants. The wastewater treatment plant is regarded as a type of decentralized interconnected system in this algorithm. To reduce the design difficulty of control strategies, multiple agents are created, and each agent is only responsible for optimizing the control strategy of a single subsystem. In addition, based on a well-designed Q-function, the agent considers the impact on other subsystems while optimizing its own control strategy. Therefore, the control algorithm exhibits good performance in both the implementation difficulty and the control accuracy. To ensure the initial performance of the control algorithm, the prior control strategy with expert knowledge is integrated into the control algorithm. Considering the disturbance factors in wastewater treatment plants, the direct control strategy is modified into an incremental control strategy, which improves the anti-interference performance of the control algorithm. The stability of the proposed algorithm is proved by constructing Lyapunov functions. The superiority of the control algorithm is verified through quantitative comparison results with other control algorithms.
Ding Wang 0001, Xin Li 0055, Junfei Qiao 0001
IEEE Trans. Ind. Informatics1
2026 Incremental Critic Learning Control With Policy Transfer for Wastewater Treatment Processes
Ding Wang 0001, Xin Li 0055, Ao Liu 0012, Junfei Qiao 0001
IEEE Trans. Ind. Informatics1
2026 Adaptive Output-Feedback Fault-Tolerant Control for Distributed Optimization of Nonlinear MASs Under Event-Triggered Communication
Yonglin Yu, Xudong Zhao 0001, Ben Niu 0003, Ding Wang 0001, Xinjun Wang 0001
IEEE Trans. Reliab.4
2026 Static and Dynamic Event-Triggered Control for Disturbed Interconnected Systems Through Adaptive Critic Designs
abstract
This article develops decentralized static and dynamic event-triggered control (ETC) strategies for continuous-time (CT) nonlinear systems featuring unmatched external disturbances and matched interconnections. Both static and dynamic ETC schemes can be used separately to alleviate the computational load and communication resources. The proposed static and dynamic event-triggering conditions differ from existing approaches in terms of condition design and stability proof. The static event-triggered mechanism (ETM) relies solely on the system state, whereas the dynamic ETM integrates the system state with an internal variable derived from a differential equation. The decentralized event-triggered (ET) controller is established for the large-scale system, which is comprised of many ETC policies for each individual subsystem. Based on the critic-only architecture, the optimal policies are obtained within the adaptive critic framework by solving the relevant ET Hamilton–Jacobi–Isaacs (HJI) equation. Specifically, the weight vectors of critic neural networks (NNs) are updated using a novel gradient descent strategy with momentum. Finally, the effectiveness of the developed method is validated through two simulation examples.
Ding Wang 0001, Zeqiang Yuan, Ao Liu 0012, Qinna Hu, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2025 Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
abstract
Multimodal reasoning in Large Language Models (LLMs) struggles with incomplete knowledge and hallucination artifacts, challenges that textual Knowledge Graphs (KGs) only partially mitigate due to their modality isolation. While Multimodal Knowledge Graphs (MMKGs) promise enhanced cross-modal understanding, their practical construction is impeded by semantic narrowness of manual text annotations and inherent noise in visual-semantic entity linkages. In this paper, we propose Vision-align-to-Language integrated Knowledge Graph (VaLiK), a novel approach for constructing MMKGs that enhances LLMs reasoning through cross-modal information supplementation. Specifically, we cascade pre-trained Vision-Language Models (VLMs) to align image features with text, transforming them into descriptions that encapsulate image-specific information. Furthermore, we developed a cross-modal similarity verification mechanism to quantify semantic consistency, effectively filtering out noise introduced during feature alignment. Even without manually annotated image captions, the refined descriptions alone suffice to construct the MMKG. Compared to conventional MMKGs construction paradigms, our approach achieves substantial storage efficiency gains while maintaining direct entity-to-image linkage capability. Experimental results on multimodal reasoning tasks demonstrate that LLMs augmented with VaLiK outperform previous state-of-the-art models. Our code is published at https://github.com/Wings-Of-Disaster/VaLiK.
Siyuan Meng, Yanting Gao, Song Mao, Pinlong Cai, Guohang Yan, Yirong Chen, Zilin Bian, Ding Wang 0001, Botian Shi
ICCV9
2025 HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
abstract
While Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge, conventional single-agent RAG remains fundamentally limited in resolving complex queries demanding coordinated reasoning across heterogeneous data ecosystems. We present HM-RAG, a novel Hierarchical Multi-agent Multimodal RAG framework that pioneers collaborative intelligence for dynamic knowledge synthesis across structured, unstructured, and graph-based data. The framework is composed of a three-tiered architecture with specialized agents: a Decomposition Agent that dissects complex queries into contextually coherent sub-tasks via semantic-aware query rewriting and schema-guided context augmentation; Multi-source Retrieval Agents that carry out parallel, modality-specific retrieval using plug-and-play modules designed for vector, graph, and web-based databases; and a Decision Agent that uses consistency voting to integrate multi-source answers and resolve discrepancies in retrieval results through Expert Model Refinement. This architecture attains comprehensive query understanding by combining textual, graph-relational, and web-derived evidence, resulting in a remarkable 12.95% improvement in answer accuracy and a 3.56% boost in question classification accuracy over baseline RAG systems on the ScienceQA and CrisisMMD benchmarks. Notably, HM-RAG establishes state-of-the-art results in zero-shot settings on both datasets. Its modular architecture ensures seamless integration of new data modalities while maintaining strict data governance, marking a significant advancement in addressing the critical challenges of multimodal reasoning and knowledge synthesis in RAG systems.
Ruoyu Yao, Siyuan Meng, Ding Wang 0001, Jun Ma 0008
ACM Multimedia6
2025 Adjustable behavior-guided adaptive dynamic programming for neural learning control
Guohan Tang, Ding Wang 0001, Ao Liu 0012, Junfei Qiao 0001
Neurocomputing2
2025 Temporal difference learning with multi-step returns for intelligent optimal control of dynamic systems
Peng Xin, Ding Wang 0001, Ao Liu 0012, Junfei Qiao 0001
Neurocomputing2
2025 Enhancing offline reinforcement learning for wastewater treatment via transition filter and prioritized approximation loss
abstract
Wastewater treatment plays a crucial role in urban society, requiring efficient control strategies to optimize its performance. In this paper, we propose an enhanced offline reinforcement learning (RL) approach for wastewater treatment. Our algorithm improves the learning process. It uses a transition filter to sort out low-performance transitions and employs prioritized approximation loss to achieve prioritized experience replay with uniformly sampled loss. Additionally, the variational autoencoder is introduced to address the problem of distribution shift in offline RL. The proposed approach is evaluated on a nonlinear system and wastewater treatment simulation platform, demonstrating its effectiveness in achieving optimal control. The contributions of this paper include the development of an improved offline RL algorithm for wastewater treatment and the integration of transition filtering and prioritized approximation loss. Evaluation results demonstrate that the proposed algorithm achieves lower tracking error and cost.
Ruyue Yang, Ding Wang 0001, Menghua Li, Chengyu Cui, Junfei Qiao 0001
Neurocomputing2
2025 Evolution-guided Q-learning for tracking control of unknown dynamic systems
Zeqiang Yuan, Ding Wang 0001, Jiangyu Wang, Junfei Qiao 0001
Neurocomputing2
2025 A Zonotopic Secure Estimation Framework for Cyber-Physical Systems Under Dynamic Event-Triggered Mechanism
abstract
This paper proposes a zonotopic secure estimation framework for discrete-time cyber-physical systems (CPSs) subject to unknown-but-bounded (UBB) disturbances and false data injection (FDI) attacks. A decentralized dynamic event-triggered mechanism (DETM) is developed, allowing each sensor to adapt its triggering threshold using local output deviations, thereby reducing communication while preserving estimation accuracy. Subsequently, a zonotopic interval observer is developed to estimate the system state under DETM. The observer propagates zonotopic error bounds and is designed via LMIs to ensure stability and l1 performance. Furthermore, a zonotope-based attack reconstruction approach is formulated. The attack signal is conservatively enclosed in a residual-based zonotope, and a threshold test is used to isolate attacked channels. Finally, simulation results confirm that the method reduces communication significantly while maintaining reliable estimation, validating its use in resource-constrained CPSs.
Jianing Hu, Zhihua Guo 0001, Ben Niu 0003, Ding Wang 0001, Xinjun Wang 0001, Hao Liu 0012
IEEE Internet Things J.5
2025 Fault Detection and Isolation for Multiagent Systems Under Event-Triggered Communication: A Zonotopic Joint Estimation Method
abstract
This paper investigates the fault detection and isolation problem of a class of leader-follower multi-agent systems subject to unknown-but-bounded disturbances under an event-triggered communication mechanism (ETCM). Firstly, an adaptive ETCM is proposed to avoid the continuous acquisition of outputs from neighboring agents, thereby reducing the communication burden. Secondly, a joint state and fault observer is designed, and H∞ analysis is utilized to improve the robustness of the observer against external disturbances and event-triggered errors. Next, zonotopic reachability analysis is utilized to derive the estimated intervals of the system state and fault signal, based on which a fault detection and isolation method is developed to determine both when the system is faulty and which agent is at fault. Finally, a fighter formation model is employed to verify the effectiveness of the proposed approach.
Mingliang Tian, Zhihua Guo 0001, Ben Niu 0003, Xinjun Wang 0001, Ding Wang 0001, Huanqing Wang 0001
IEEE Internet Things J.6
2025 Self-triggered neural tracking control for discrete-time nonlinear systems via adaptive critic learning
Lingzhi Hu, Ding Wang 0001, Gongming Wang, Junfei Qiao 0001
Neural Networks2
2025 Neural-network-based accelerated safe Q-learning for optimal control of discrete-time nonlinear systems with state constraints
Ding Wang 0001, Junfei Qiao 0001
Neural Networks2
2025 An Improved Trajectory Tracking Mechanism With Adaptive Critic for Event-Based Multiplayer Zero-Sum Games
abstract
In this paper, based on the adaptive critic control method, an improved event-based trajectory tracking mechanism of continuous-time (CT) nonlinear multiplayer zero-sum games (MZSGs) is established. It is worthy of note that previous papers studying the trajectory tracking issue of nonlinear CT MZSGs only apply to the case where the reference trajectory eventually converges to zero. Consequently, this paper develops an improved mechanism to overcome this weakness. Later, an event-triggered framework is brought in to reduce the amount of computation and improve control efficiency. In this process, an innovative triggering condition is provided. At the same time, the infamous Zeno behavior is ruled out through theoretical analysis. Furthermore, the event-based near-optimal controls and event-based near-worst disturbances for tracking error dynamics are gained by building and adjusting a single critic neural network. Immediately after, by utilizing the Lyapunov method, the uniform ultimate boundedness stability of the tracking error and the weight estimation error is ensured. Lastly, an example containing two case studies is offered to validate the validity of the established mechanism. Note to Practitioners—Complex industrial processes often involve multiple control inputs and may also be affected by multiple disturbances at the same time, which can be referred to as a MZSG. Since many industrial processes can be viewed as a tracking question of nonlinear systems and the event-triggered mechanism can decrease the computational cost, the tracking problem for event-based nonlinear MZSGs is studied in this paper, which is significant for control practitioners. Moreover, the Hamilton-Jacobi-Isaacs equation is often difficult to solve when dealing with the game problem. Hence, an adaptive critic technique is presented to acquire the near-optimal controls and the near-worst disturbances, which replaces the traditional actor-critic framework and thus simplifies the theoretical analysis. Note that this paper proposes an innovative triggering condition to relax the restriction on the choice of disturbance rejection level. Compared to previous works dealing with the tracking problem of nonlinear MZSGs, the method presented in this paper makes the choice of the reference trajectory more flexible and thus enhances the applicability in general industrial processes. Finally, stability analysis and simulation results are given. Note that for different practical situations, practitioners can adjust the related parameters to achieve the tracking control of MZSGs and minimize the computational cost.
Menghua Li, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.2
2025 Adaptive Prescribed-Time Consensus Tracking Control Scheme of Nonlinear Multi-Agent Systems Under Deception Attacks
abstract
This article investigates the adaptive prescribed-time consensus tracking control problem for nonlinear multi-agent systems (MASs), where the states of systems are unmeasured and the actuators suffer from the deception attacks. Firstly, a novel coordinate transformation technology is developed by introducing a time-varying constraint function, such that the prescribed-time tracking control problem of nonlinear MASs is converted into the constraint problem of the error variables. Then, a new attack compensator is proposed to address the unknown time-varying attack gains caused by the actuator deception attacks. Further, the state observers are designed to estimate the unavailable state variables and fuzzy-logic systems (FLSs) are employed to handle the unknown functions that exist within the systems. In addition, the attack compensator-based controller ensures the boundedness of all signals, while the error variables converge to the predefined region in a specified time. The upper bound of the whole tracking errors in the mean square sense can be decreased by selecting the appropriate design parameters. At last, the simulation example illustrates the availability of the developed control method. Note to Practitioners—In the industry, consensus tracking control of nonlinear MASs exists in many different systems, such as mobile robot networks, intelligent transportation management, surveillance and monitoring. Since the above systems operate in a network environment, the security problems of the systems cannot be ignored. Hence, considering the unmeasured states, the unknown functions, and the unknown time-varying attack gains existing simultaneously in the studied systems, it is a challenging and meaningful task to achieve the desired security control objectives. On the other hand, based on a time-varying constraint function, this article presents an adaptive prescribed-time consensus tracking control scheme for the nonlinear MASs under the deception attacks. It provides a viable strategy for industrial applications.
Ben Niu 0003, Yahui Gao, Guangju Zhang, Xudong Zhao 0001, Huanqing Wang 0001, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.6
2025 Adaptive Fuzzy Resilient Fixed-Time Bipartite Consensus Tracking Control for Nonlinear MASs Under Sensor Deception Attacks
abstract
This paper studies the adaptive fuzzy resilient fixed-time bipartite consensus tracking control problem for a class of nonlinear multi-agent systems (MASs) under sensor deception attacks. Firstly, in order to reduce the impact of unknown sensor deception attacks on the nonlinear MASs, a novel coordinate transformation technique is proposed, which is composed of the states after being attacked. Then, in the case of unbalanced directed topological graph, a partition algorithm (PA) is utilized to implement the bipartite consensus tracking control, which is more widely applicable than the previous control strategies that only apply to balanced directed topological graph. Moreover, the fixed-time control strategy is extended to nonlinear MASs under sensor deception attacks, and the singularity problem that exists in fixed-time control is successfully avoided by employing a novel switching function. The developed distributed adaptive resilient fixed-time control strategy ensures that all the signals in the closed-loop system are bounded and the bipartite consensus tracking control is achieved in fixed time. Finally, the designed control strategy’s validity is demonstrated by means of a simulation experiment.Note to Practitioners—Currently, there are many practical application scenarios for nonlinear MASs, such as intelligent transportation, unmanned aerial vehicle cluster formation, etc. This paper investigates the adaptive fuzzy resilient fixed-time bipartite consensus tracking control problem for a class of nonlinear MASs under sensor deception attacks. In practice, these two situations are common: 1) The topology graph describing the communication relationship of nonlinear MASs is unbalanced 2) The nonlinear MASs is subjected to external malicious cyber attacks. Therefore, a novel coordinate transformation technique is proposed to reduce the impact of sensor deception attacks on the nonlinear MASs, and a partition algorithm is employed to implement bipartite consensus tracking control based on the unbalanced communication topology graph. At the same time, the nonsingular fixed-time control strategy can significantly improve the convergence of the studied nonlinear MASs. Furthermore, the nonlinear nonstrict-feedback model and backstepping design method used in this paper are general and practical.
Ben Niu 0003, Zihao Shang, Guangju Zhang, Huanqing Wang 0001, Xudong Zhao 0001, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.7
2025 Discounted Stable Adaptive Critic Design for Zero-Sum Games With Application Verifications
abstract
In this paper, an adaptive critic design with performance guarantee is established based on the discounted value iteration algorithm to settle with the optimal regulation problem for discrete-time zero-sum games. Value iteration is implemented to obtain the approximate optimal solutions to the Hamilton-Jacobi-Isaacs equation for nonlinear systems and the game algebraic Riccati equation for linear systems. Then, we focus on system stability affected by the introduction of the discount factor and the admissibility of the policy pairs in the value iteration process. The appropriate selection range of the discount factor and the criteria for ensuring system stability are established to assist in obtaining the stabilized optimal policy pair, which not only makes the cost function converge to the optimal value, but also guarantees the asymptotic stability of the closed-loop system. Finally, practical examples for the power system and the ball-beam system are conducted to demonstrate the effectiveness of the presented method. Note to Practitioners—Since there exist a multitude of dynamic systems with uncertainty and interference, the zero-sum game problems are ubiquitous, especially when dealing with dynamic systems featuring antagonistic properties. As an important research direction in the field of optimal control, zero-sum games usually involve designing policy pairs that can optimize the system performance in the presence of adversarial disturbances. Due to the excellent adaptability, value iteration in adaptive dynamic programming is employed to deal with this kind of issues. In addition to focusing on the optimality of policies, the system stability during the control process is equally significance, where the stability is the premise of all operations. Therefore, we are dedicated to providing guidance on the optimal regulation of discrete-time zero-sum games with performance guarantee, which contributes to obtain the stable optimal policy pair. Theoretical analysis of the stability is provided and the asymptotic stability of the system is ensured, which improves the performance of the designed controller. Furthermore, simulation experiments for practical applications are conducted, which verify the feasibility and effectiveness of the proposed control design.
Ding Wang 0001, Menghua Li, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.2
2025 Prescribed Performance Adaptive Containment Control for Full-State Constrained Nonlinear Multiagent Systems: A Disturbance Observer-Based Design Strategy
abstract
This paper focuses on the prescribed performance adaptive containment control problem for a class of nonlinear nonstrict-feedback multiagent systems (MASs) with unknown disturbances and full-state constraints. First, the radial basis function neural networks (RBF NNs) technology is employed to approximate the unknown nonlinear functions in the system, and the problem of “explosion of complexity” caused by repeated derivation of virtual controls is solved by using the dynamic surface control (DSC) technology. Then, the nonlinear disturbance observers are designed to estimate the external disturbance, and the barrier Lyapunov functions (BLFs) and the prescribed performance function (PPF) are combined to achieve the control objective of prescribed performance without violating the full-state constraints. The theoretical result shows that all signals in the closed-loop system are semiglobally uniformly ultimately bounded (SGUUB), and the local neighborhood containment errors can converge to the specified boundary. Finally, two simulation examples show the effectiveness of the proposed method.Note to Practitioners—The containment control problem is a hot topic in the field of control, which plays an important role in practical engineering. Especially for this problem of nonlinear MASs, the mathematical models are difficult to be obtained accurately. This paper investigates the prescribed performance adaptive containment control problem for the nonlinear nonstrict-feedback MASs, whose model can be extended to more complex engineering applications, such as unmanned aerial vehicle formations and intelligent traffic management. It is worth noting that external disturbances and state constraint problems often exist in practical applications. Therefore, the disturance observers are designed to compensate for the system disturbances, which can eliminate the impacts of disturbances on the systems. By introducing BLFs, it is ensured that all states of the system are constrained within the specified regions. To sum up, the paper proposes a prescribed performance adaptive containment control strategy, which contributes to the development of containment control for MASs in practical applications.
Jihang Sui, Ben Niu 0003, Xudong Zhao 0001, Ding Wang 0001, Bocheng Yan
IEEE Trans Autom. Sci. Eng.5
2025 Intelligent Critic Design With Policy Transfer for Wastewater Treatment Processes
abstract
Wastewater treatment plays a meaningful role in environmental protection, water recycling and public health, and its optimal control offers significant economic and social value. However, the wastewater treatment plant is a large-scale nonlinear system, and its own operation is affected by a series of uncontrollable factors such as the flow rate and composition of the incoming wastewater. This makes it extremely challenging for traditional control methods to realize precise control. To address these issues, this paper proposes a novel knowledge-guided policy optimization control method by integrating policy transfer and adaptive dynamic programming. First, an adaptive critic policy transfer framework based on the source domain knowledge selection method is introduced to construct knowledge extraction and data mining structures to enhance the learning efficiency of the target domain agent. Second, a source domain knowledge selection module is proposed to adaptively adjust the knowledge extracted by the target domain agent. Third, an adaptive termination module is designed to determine when the immature policy in the source domain should terminate during the target domain learning process. Finally, the theoretical stability of the method in this paper is demonstrated by designing a reasonable Lyapunov function. The system performance of the wastewater treatment plant under dry, rainy, and stormy weather conditions is evaluated, ultimately demonstrating the superior performance and adaptability of the proposed method.
Ding Wang 0001, Ning Gao 0008, Xin Li 0055, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.1
2025 Safe Optimal Tracking Control via Multi-Step Critic Learning for Unknown Nonlinear Systems
abstract
For unknown nonlinear systems, a safe optimal tracking control algorithm is developed based on multi-step critic learning. By integrating the control barrier function into the critic learning framework, the algorithm ensures that the tracking error is kept within a safe region and converges to zero with the minimal cost. Utilizing historical operational data of the system, a model network is established to identify the unknown system dynamics. By constructing feedforward control, the tracking problem of the original system is transformed into a regulation problem of the error system. To enhance the convergence speed of the algorithm, a critic learning framework with multi-step policy evaluation is designed. Furthermore, a criterion is developed to determine the admissibility of the safe tracking control policy at each iteration step. The convergence of the proposed algorithm is also proved. Finally, the simulation results demonstrate effectiveness of the algorithm and the validity of the theoretical results.
Ding Wang 0001, Xin Li 0055, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.1
2025 Online Digital Twin Adaptive Critic Design With Long Short-Term Memory for Wastewater Treatment Plants
abstract
The wastewater treatment system is a complex unknown system with nonlinear and uncertain characteristics. It is necessary to control the concentration of the dissolved oxygen and the nitrate nitrogen at the set value in the wastewater treatment process. However, traditional control methods are difficult to meet the accuracy requirements for the required concentration value. At the same time, the digital twin (DT) technology has been widely applied to practical industrial systems like the wastewater treatment system in recent years. Based on the above backgrounds, an online DT adaptive critic design (DTACD) is developed by combining the long short-term memory (LSTM) neural network with the action-critic structure. First, a set of historical data is collected, which is used to build a digital model using LSTM. Then, the digital model is used to guide the real wastewater treatment system to control the nitrate nitrogen concentration and the dissolved oxygen concentration at set values. Finally, we select the Benchmark Simulation Model No.1 as the model for the simulation experiment. Compared with other methods, DTACD shows a better performance.
Ding Wang 0001, Hongyu Ma, Honggui Han, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.1
2025 Evolution-Guided Q-Learning With Dual Swarm Intelligence for Model-Free Optimal Control
abstract
In this article, a novel accelerated evolution-guided Q-learning (EGQL) algorithm is introduced to address optimal control problems for unknown nonlinear systems. A novel adaptive evolutionary algorithm is introduced here to enhance problem-solving strategies in two key aspects, which is a concept referred to dual swarm intelligence. First, the novel evolutionary algorithm is incorporated into policy improvement, replacing the gradient descent approach in traditional Q-learning. The method eliminates the need for gradient information in the approximate Q-function, while also enabling more precise policy solutions. Second, the commonly used polynomial model is replaced with a neural network, which is further optimized by the evolutionary algorithm to approximate the critic network. The accuracy of the approximate Q-function is significantly enhanced, improving the precision of value function update. Furthermore, the monotonicity and convergence properties of the algorithm are thoroughly analyzed. Finally, the effectiveness of the EGQL method is validated through two simulation experiments, demonstrating its superiority over traditional Q-learning.
Ding Wang 0001, Zeqiang Yuan, Guohan Tang, Jiangyu Wang, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.1
2025 Optimizing Wastewater Treatment Control Using Policy-Constrained Offline Reinforcement Learning and Lyapunov Density Model
abstract
Wastewater treatment plants (WWTPs) constitute a crucial component of contemporary societal infrastructure, playing a critical role in purifying and recycling wastewater from both municipal and industrial sources. The effectiveness and reliability of control strategies in these processes are paramount. This paper presents an offline reinforcement learning (RL) controller enhanced with the Lyapunov density model (LDM) to solve the optimal tracking control problem in wastewater treatment processes. The offline RL algorithm employs off-policy learning to derive the optimal control policy from the offline dataset without risky online trials. By incorporating the density estimation of the offline dataset, the LDM estimates the probability that the system adheres to the distribution of the dataset under the learned policy. The introduction of LDM helps constrain the learned policy, thereby enhancing its reliability when encountering unknown states. Additionally, a weighting scheme, based on the coefficient of variation, is applied to determine the relative weights of the LDM and the critic in the update of the control policy. Simulations conducted on a nonlinear system and the wastewater treatment process demonstrate that the proposed method learns a more effective policy from the offline dataset, leading to reduced implementation costs.
Ruyue Yang, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.2
2025 Event-Triggered Adaptive Finite-Time Control for a Robotic Manipulator System With Global Prescribed Performance and Asymptotic Tracking
abstract
This article studies the dynamic event-triggered adaptive finite-time tracking control issue for a robotic manipulator (RM) system with disturbances. First, a new global prescribed performance function (PPF) is designed based on a scaling function such that the tracking error evolves within the constrained bounds and the restriction related to the initial conditions is removed. Then, the finite-time command filter (FTCF) is used to avoid the direct derivations of virtual controllers and the singularity issue of the conventional backstepping technique. Moreover, the filtering errors caused by the FTCF are removed by the designed error compensation mechanism. A novel dynamic event-triggered mechanism (DETM) using the dynamic auxiliary variable is designed to save communication resources. The proposed control scheme can guarantee that all signals of the RM are globally bounded within a finite time, and the tracking error can asymptotically reach zero. Finally, a simulation example and several comparative simulations show the validity of the proposed scheme.
Jihang Sui, Ben Niu 0003, Yongsheng Ou, Xudong Zhao 0001, Ding Wang 0001
IEEE Trans. Cybern.5
2025 Relaxed Optimal Control With Self-Learning Horizon for Discrete-Time Stochastic Dynamics
abstract
The innovation of optimal learning control methods is profoundly propelled due to the improvement of the learning ability. In this article, we investigate the synthesis of initialization and acceleration for optimal learning control algorithms. This approach contrasts with traditional methods that concentrate solely on either the improvement of initialization or acceleration. Specifically, we establish a novel relaxed policy iteration (PI) algorithm with self-learning horizon for stochastic optimal control. Notably, by suitably utilizing self-learning horizon, we can directly evaluate inadmissible policies to reduce the initialization burden. Meanwhile, the inadmissible policy can be rapidly optimized with few learning iterations. Then, several critical conclusions of relaxed optimal control are established by discussing algorithm convergence and system stability. Furthermore, to provide the convincing application potentials, a class of unconventional problems is effectively solved by the relaxed PI algorithm, including the dynamics with external noises and nonzero equilibrium. Finally, we present a series of nonlinear benchmarks with practical applications to comprehensively evaluate the performance of relaxed PI. The experimental results obtained from these diverse benchmarks uniformly highlight the effectiveness of self-learning horizon mechanism.
Ding Wang 0001, Jiangyu Wang, Ao Liu 0012, Derong Liu 0001, Junfei Qiao 0001
IEEE Trans. Cybern.1
2025 Composite-Observer-Based Adaptive Consensus Tracking Control for Nonlinear MASs With Unknown Control Directions Against Deception Attacks
abstract
This article primarily studies the adaptive output-feedback consensus tracking control issue for nonlinear multiagent systems (MASs) with unknown control directions against deception attacks. First, a composite observer combining the state observer and the disturbance observer is developed to concurrently estimate the states of confronting deception attacks and unmeasurable disturbances. Moreover, to resolve the unknown gains resulting from deception attacks, the adaptive attack compensator is proposed. Furthermore, in view of the logarithm Lyapunov function in the final step of the design process and the intelligent approximation technique, a new composite-observer-based adaptive consensus tracking control strategy is constructed. The suggested control strategy ensures the boundedness of all the closed-loop signals while also achieving synchronous tracking of the leader's output by the followers. Last but not least, the effectiveness of the suggested control strategy is validated through two simulation examples.
Luyao Wen, Ben Niu 0003, Xudong Zhao 0001, Guangdeng Zong, Ding Wang 0001, Wencheng Wang 0002, Yuqiang Jiang
IEEE Trans. Cybern.5
2025 Accelerated Value Iteration-Based Safe Q-Learning for Data-Driven Optimal Tracking Control
abstract
In this article, an accelerated value iteration-based safe Q-learning (SQL) algorithm is developed to design the tracking controller for unknown nonlinear systems. First, an augmented Q-function, consisting of a quadratic utility function and an adjustable positive-definite control barrier function (CBF), is devised to ensure both the optimality and safety of the tracking controller. The quadratic utility function, associated with optimality, guarantees that the tracking controller can eliminate the ultimate tracking error, regardless of the reference trajectory. The adjustable positive-definite CBF, pertaining to safety, ensures that the tracking error converges faster toward zero while remaining within the safe set at all times. Second, an accelerated iterative learning mechanism, comprising policy evaluation (PE) and policy improvement (PI), is employed to discover the safe optimal tracking control policy. Integrating the difference between two iterative Q-functions into the current PE process can expedite the convergence rate of the SQL algorithm. A policy optimization technique based on Nesterov Momentum method is utilized to accelerate the PI process of the SQL algorithm. When faced with a large amount of offline data, the two-stage accelerated learning effectively reduces computational pressure. Furthermore, convergence of the Q-function sequence and safety of the optimal tracking policy are theoretically analyzed. Finally, by using neural networks and the action-critic structure, two simulation examples are performed to verify the availability of accelerated SQL methods.
Ding Wang 0001, Shijie Song 0001, Junfei Qiao 0001
IEEE Trans. Cybern.2
2025 Adaptive Fuzzy Consensus Tracking Using Monotone Tube Boundary Functions for Nonlinear MASs Under FDI Attacks: A Single-Parameter Integration Approach
abstract
This article principally presents the adaptive fuzzy dynamic event-triggered output-feedback consensus tracking control problem for constrained nonlinear multiagent systems encountering false data injection attacks, using monotone tube boundary functions (MTBFs). First, the composite observer is designed to concurrently estimate the unpredictable system states and disturbances. Then, a novel parameter integration approach is constructed via the fuzzy approximation technique to reduce the computational burden, where the uncertainty terms (state errors, composite errors, fuzzy weights, and injection attacks) can be converted into a linear parameterized form with merely one unknown scalar. In addition, a set of MTBFs is designed to prevent the excessive overshoot or jitter of system errors in the transient-state performance. Furthermore, an adaptive dynamic event-triggered control strategy is constructed on the basis of the auxiliary dynamic variable. All the closed-loop signals are shown to be bounded, and the followers can achieve the synchronous tracking of the leader's output. In the end, the simulation results demonstrate the efficiency of the proposed control strategy.
Ben Niu 0003, Luyao Wen, Xudong Zhao 0001, Guangdeng Zong, Ding Wang 0001, Baoyi Zhang
IEEE Trans. Fuzzy Syst.5
2025 Adaptive Fuzzy Nonsingular Fixed-Time Bipartite Consensus Tracking Using Adding Power Integration Technique for Stochastic Nonlinear Constrained MASs
abstract
In this article, the adaptive fuzzy fixed-time bipartite consensus tracking control problem is studied for stochastic nonlinear multi-agent systems with unknown control gains and time-varying output constraints. First, in order to address the difficulties arising from the unknown control gains, the Nussbaum technique is employed. In the meantime, the$tan$-type nonlinear mapping function is introduced, which guarantees the predefined output constraints are not violated. Then, different from the previous control strategies in which they only focused on the balanced directed topology, the classification optimization algorithm is presented to accomplish the bipartite consensus tracking control according to the structurally unbalanced directed topology. Besides, by combining the adaptive backstepping technique with the adding power integration methodology, the nonsingular fixed-time control strategy is proposed. The proposed adaptive fuzzy fixed-time control strategy ensures that the bipartite consensus tracking errors converge to a region near zero in fixed time and all the signals in the closed-loop system are bounded in probability. Last, the effectiveness of the presented control scheme is demonstrated with a simulation example.
Ben Niu 0003, Zihao Shang, Ding Wang 0001, Huanqing Wang 0001, Wencheng Wang 0002
IEEE Trans. Fuzzy Syst.5
2025 Model-Free Neuro-Fuzzy Q-Learning Control With Swarm Intelligence
abstract
In this paper, a novel neuro-fuzzy-based evolution-guided Q-learning (EGQL) algorithm is established for solving the optimal control problem of unknown nonlinear systems. To enhance the accuracy for approximating the Q-function, the adaptive neuro-fuzzy inference system (ANFIS) is leveraged, which offers superior precision compared to traditional polynomial approximations commonly used in adaptive dynamic programming (ADP). Despite its advantages, the ANFIS-based approximation faces challenges in obtaining the derivative of the Q-function with respect to the control input. To address this limitation, evolutionary algorithms are integrated into EGQL, eliminating the need for gradient information by directly minimizing the Q-function values to derive optimal control strategies. This integration enables precise and robust exploration of the solution space, resulting in accurate and reliable control policies. Furthermore, convergence and monotonic improvement are ensured by the EGQL algorithm, making it suitable for uncertain and nonlinear environments. The effectiveness and superiority of the ANFIS-based EGQL algorithm are validated through simulation results. The developed algorithm achieves a 3.1% reduction in total cost compared to the traditional approach, demonstrating superior control performance.
Ding Wang 0001, Zeqiang Yuan, Ao Liu 0012, Junfei Qiao 0001
IEEE Trans. Fuzzy Syst.1
2025 Integrated Online Q-Learning Design for Wastewater Treatment Processes
abstract
The efficiency and economy of the nonlinear optimal control process in wastewater treatment plants are two crucial indicators, corresponding to achieving the control objective faster and reducing the preset cost function. To accomplish this, an integrated online Q-learning (IOQL) algorithm, driven by a prior policy and an exploration policy, is proposed for nonlinear discrete-time systems characterized by nonaffine features and unknown structures. First, a prior policy based on historical or artificial experience is designed to reduce the training time of the controller and provide a more stable learning environment. By introducing a weighting factor, the impact of the prior policy on the overall learning process can be adjusted. Second, an exploration policy is trained online through new experiences collected from the real environment. By leveraging two policies, we can swiftly and smoothly adjust the critic network for approximating the cost function and the action network for approximating the exploration policy, which can gradually enhance the control outcomes. Third, a stability condition with reasonable bounds is presented for the IOQL design. Finally, experimental and comparative results pertaining to a wastewater treatment plant, specifically evaluating learning speed and cost consumption, clearly demonstrate the significant advantages and superiority of the IOQL algorithm.
Ding Wang 0001, Junfei Qiao 0001
IEEE Trans. Ind. Informatics2
2025 Online Self-Triggered Transmission Control With Critic Learning for Discrete Nonlinear Systems
abstract
In this article, a novel online self-triggered transmission control (STTC) framework is constructed based on the critic learning technique, which aims at tackling the optimal regulation issue of discrete-time nonlinear systems. On the premise of ensuring the system stability, a self-sampling function is designed only related to the sampling state, so that the next triggering moment can be determined. This not only effectively reduces the computational burden, but also avoids continuous judgment for the triggering condition similar to traditional event-based methods. Furthermore, the developed control method can be found to possess excellent triggering performance through theoretical analysis. Then, the model, critic, and action networks are established to execute the online critic learning algorithm, which make the control policy is adjusted in real-time to the optimal level. Finally, an experimental plant with nonlinear characteristics is given to illustrate the overall performance of the proposed online STTC method.
Lingzhi Hu, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Parallel Multistep Evaluation With Efficient Data Utilization for Safe Neural Critic Control and Its Application to Orbital Maneuver Systems
abstract
Data-driven methods have significantly advanced optimal learning control, but some approaches overlook systematic considerations of data utilization, including safety, efficiency, and error accumulation. To address the neglects in safe neural critic control, this article introduces a parallel multistep evaluation mechanism that combines data from the system interaction with data generated by data-driven models. Based on this evaluation mechanism, we propose a novel parallel multistep Q-learning algorithm that enhances data utilization efficiency and mitigates the error accumulation. Furthermore, we formulate a novel control barrier function (CBF) to ensure safety during learning and control processes, which is capable of dealing with asymmetric constraints and adjusting the constraint strength. In addition, the analysis reveals that multistep information introduced by data-driven models influences the learning performance of actor-critic neural networks (NNs). Finally, parallel multistep Q-learning, which makes use of data in aspects of safety, efficiency, and error bounds, is validated within an orbital maneuver system.
Jiangyu Wang, Ding Wang 0001, Derong Liu 0001, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Reinforcement Learning for Robust Dynamic Event-Driven Constrained Control
abstract
We consider a robust dynamic event-driven control (EDC) problem of nonlinear systems having both unmatched perturbations and unknown styles of constraints. Specifically, the constraints imposed on the nonlinear systems' input could be symmetric or asymmetric. Initially, to tackle such constraints, we construct a novel nonquadratic cost function for the constrained auxiliary system. Then, we propose a dynamic event-triggering mechanism relied on the time-based variable and the system states simultaneously for cutting down the computational load. Meanwhile, we show that the robust dynamic EDC of original nonlinear-constrained systems could be acquired by solving the event-driven optimal control problem of the constrained auxiliary system. After that, we develop the corresponding event-driven Hamilton-Jacobi-Bellman equation, and then solve it through a unique critic neural network (CNN) in the reinforcement learning framework. To relax the persistence of excitation condition in tuning CNN's weights, we incorporate experience replay into the gradient descent method. With the aid of Lyapunov's approach, we prove that the closed-loop auxiliary system and the weight estimation error are uniformly ultimately bounded stable. Finally, two examples, including a nonlinear plant and the pendulum system, are utilized to validate the theoretical claims.
Xiong Yang 0001, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Attack Detection and Reconstruction for CPSs Based on Unknown Input Observer and Reachability Analysis
Chaojiang Liang, Ben Niu 0003, Zhihua Guo 0001, Xinjun Wang 0001, Hao Liu 0012, Ding Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2025 Nonperiodic and Periodic Event-Triggered Online H∞ Control for Constrained Nonlinear Systems
abstract
This article presents two new event-triggered control (ETC) schemes based on the online critic learning technique, which aims at tackling the optimal regulation problem of discrete-time constrained nonlinear systems with the disturbance input. First, a novel stability criterion condition is designed to obtain an initial admissible policy pair by using an offline iterative method under the time-triggered control framework. Then, starting from the stability of the constrained system, a nonperiodic ETC method and a periodic ETC method are developed by adopting an online learning algorithm. In addition, four kinds of neural networks are constructed for the implementation of the event-based online$H_{\infty }$optimal control strategy. Finally, two experimental examples with physical backgrounds are provided to illustrate the effectiveness and superiority of the developed schemes.
Ding Wang 0001, Lingzhi Hu, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2025 Finite-Time Adaptive Bipartite Consensus Tracking Control for Constrained Nonlinear MASs: A Locally Optimal Bipartition Strategy
abstract
This article introduces a finite-time adaptive bipartite consensus tracking control algorithm leveraging an improved nonlinear mapping (NM), capable of implementing the prescribed asymmetric full-state constraint requirements for nonlinear heterogeneous multiagent systems (MASs) under actuator deception attacks. First, a locally optimal bipartition strategy is proposed in this article to enable bipartite consensus tracking control within unbalanced directed graph. Unlike most previous full-state-constrained controls employing the barrier Lyapunov function (BLF), the proposed algorithm is derived from building a state-dependent NM function that can achieve one-to-one mapping for the original function and constructing a new coordinate transformation, leading to the solution of full-state-constrained performance control without the feasibility conditions on the control signals. By means of the proposed practical finite-time stability criterion, the novel design aims at constructing finite-time controllers that can ensure the existence and nonsingularity of the finite-time term in each control signal, which are oblivious in the existing results on finite time control. What’s more, the designed controller even under actuator deception attack achieves that:1)the bipartite consensus tracking errors converge to a small area containing zero in a finite time;2)the prescribed asymmetric full-state constraints are obeyed. At last, the availability of the proposed design solution is clarified by virtue of single-link robot systems.
Ben Niu 0003, Xudong Zhao 0001, Ding Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2025 Adaptive Prescribed Finite-Time Bipartite Consensus Control for Nonaffine Nonlinear MASs Under Structurally Unbalanced Topology
abstract
This article investigates an adaptive bipartite consensus tracking control algorithm for a class of heterogeneous nonaffine nonlinear multiagent systems (MASs) with prescribed finite-time tracking performance under an unbalanced communication topology. In the case of an unbalanced digraph, a novel locally optimal bipartition strategy is proposed to transform the unbalanced communication topology into a structurally balanced one, thereby enabling the implementation of bipartite consensus tracking control. To achieve the expected tracking performance, the design philosophy focuses on developing a prescribed finite-time performance function (PFTPF), capable of preassigning the convergence time and accuracy precisely beforehand. The explored adaptive control algorithm can ensure that the whole signals concerning the closed-loop MASs remain bounded while the bipartite consensus errors converge to a predetermined range around zero within the prescribed finite time. Ultimately, the simulation results on robotic systems prove the availability of the developed design solution.
Yongduan Song 0001, Xudong Zhao 0001, Huanqing Wang 0001, Ding Wang 0001, Ben Niu 0003
IEEE Trans. Syst. Man Cybern. Syst.5
2025 Intelligent Critic Learning for Data-Driven Output Tracking Control With Lightweight Parallelization
abstract
This article investigates critical challenges in optimal output tracking control, including residual tracking errors, potential instability caused by discount factors, and premature convergence due to inefficient termination criteria. We tackle these issues by developing a data-driven parallelQ-learning algorithm. Specifically, a utility function directly linked to system states is proposed to avoid the instability and error amplification issues in traditional discounted approaches. In addition, the algorithm uses dual lightweight controllers that use convergence properties to enhance learning efficiency. Based on dual controllers, a novel termination criterion is introduced to prevent premature convergence during the training process. Numerical simulations demonstrate that the proposed method eliminates tracking errors, accelerates convergence compared with traditional algorithms, and ensures stable convergence across diverse system dynamics.
Jiangyu Wang, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2025 Iterative Q-Learning Design for Zero-Sum Games With Evolving Policies
abstract
This article aims to achieve data-based online evolving control for zero-sum games with unknown dynamics. First of all, the value-iteration-basedQ-learning framework is established. Relevant properties of the iterativeQ-learning framework are analyzed, including the convergence and monotonicity. Then, the stability property is investigated and the online data is employed for off-policy learning. More importantly, two effective algorithms are designed to achieve online evolving control. In one algorithm, the monotonically nondecreasingQ-learning sequence requires the admissible criterion to guarantee the stability with the simpleQ-function initialization. In another algorithm, the monotonically nonincreasingQ-function sequence can ensure the stability without the admissible criterion, but it requires an elaborate initialQ-function. In the end, by including two examples of real physical backgrounds, the excellent performance of online evolving control is exhibited with the given algorithms.
Ding Wang 0001, Yuan Wang 0056, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2025 Adaptive Secure Bipartite Consensus Tracking Control for Nonlinear Multiagent Systems Under FDI Attacks With Predefined Accuracy
abstract
This article mainly considers the adaptive secure bipartite consensus tracking control (BCTC) problem for nonlinear multiagent systems (MASs) under false data injection (FDI) attacks with predefined accuracy. Since FDI attacks produce unknown attack gains, which increases the difficulty of the controller design, an adaptive secure control strategy is given based on the essential property of Nussbaum functions. By improving the traditional coordinate transformation in the current literatures that can only achieve unilateral consensus control, a backstepping-based control algorithm is put forward attaining bilateral consensus control. In addition, the appropriate Lyapunov functions are generated by a class of non-negative functions to construct the adaptive secure bipartite consensus controllers, which not only makes certain that the bilateral errors ultimately converge to a predefined interval, but also guarantees that all the closed-loop signals within the investigated system are bounded. Conclusively, a practical example is provided to validate the effectiveness of the proposed control strategy.
Luyao Wen, Ben Niu 0003, Ding Wang 0001, Yuqiang Jiang, Huanqing Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2025 UKF-Based Multistep Heuristic Dynamic Programming for Optimal Event-Triggering Control of Nonlinear Systems With Asymmetric Input Constraints
abstract
In this article, an unscented Kalman filter (UKF)-based multistep heuristic dynamic programming (MsHDP) optimal control algorithm is developed for nonlinear discrete-time (DT) systems with uncertainty and asymmetric input constraints. The Hamilton–Jacobi–Bellman (HJB) equation is solved by the UKF-based MsHDP algorithm, which has the advantages of faster convergence speed and handling unknown disturbances in the system. The convergence of the developed algorithm is proved under certain conditions, and the system stability is guaranteed. To reduce the communication needs, a dynamic event-triggering mechanism is designed. Then, an event-based EC structure is proposed to implement the UKF-based MsHDP algorithm, where the UKF is used to estimate the future state of uncertain systems and the critic neural network (NN) is used to approximate cost function. Finally, simulation results are provided to verify the effectiveness of the developed algorithm.
Kun Zhang 0005, Ning Liu 0025, Xiangpeng Xie 0001, Ding Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Data on the Move: Traffic-Oriented Data Trading Platform Powered by AI Agent with Common Sense
abstract
In the digital era, data has become a pivotal asset, advancing technologies such as autonomous driving. Despite this, data trading faces challenges like the absence of robust pricing methods and the lack of trustworthy trading mechanisms. To address these challenges, we introduce a traffic-oriented data trading platform named Data on The Move (DTM), integrating traffic simulation, data trading, and Artificial Intelligent (AI) agents. The DTM platform supports evident-based data value evaluation and AI-based trading mechanisms. Leveraging the common sense capabilities of Large Language Models (LLMs) to assess traffic state and data value, DTM can determine reasonable traffic data pricing through multi-round interaction and simulations. Moreover, DTM provides a pricing method validation by simulating traffic systems, multi-agent interactions, and the heterogeneity and irrational behaviors of individuals in the trading market. Within the DTM platform, entities such as connected vehicles and traffic light controllers could engage in information collecting, data pricing, trading, and decision-making. Simulation results demonstrate that our proposed AI agent-based pricing approach enhances data trading by offering rational prices, as evidenced by the observed improvement in traffic efficiency. This underscores the effectiveness and practical value of DTM, offering new perspectives for the evolution of data markets and smart cities. To the best of our knowledge, this is the first study employing LLMs in data pricing and a pioneering data trading practice in the field of intelligent vehicles and smart cities.
Yi Yu 0012, Shengyue Yao, Yexuan Fu, Jingru Yu, Ding Wang 0001, Xuhong Wang, Cen Chen 0001, Yilun Lin 0002
IV6
2024 Reinforcement learning control with n-step information for wastewater treatment systems
Xin Li 0055, Ding Wang 0001, Junfei Qiao 0001
Eng. Appl. Artif. Intell.2
2024 Multilayer adaptive critic design with digital twin for data-driven optimal tracking control and industrial applications
Ding Wang 0001, Hongyu Ma, Junfei Qiao 0001
Eng. Appl. Artif. Intell.1
2024 Adaptive critic design with weight allocation for intelligent learning control of wastewater treatment plants
Ding Wang 0001, Hongyu Ma, Ning Gao 0008, Junfei Qiao 0001
Eng. Appl. Artif. Intell.1
2024 Supplementary heuristic dynamic programming for wastewater treatment process control
Ding Wang 0001, Xin Li 0055, Peng Xin, Ao Liu 0012, Junfei Qiao 0001
Expert Syst. Appl.1
2024 Model-free intelligent critic design with error analysis for neural tracking control
Ning Gao 0008, Ding Wang 0001, Lingzhi Hu
Neurocomputing2
2024 Novel generalized policy iteration for efficient evolving control of nonlinear systems
Haiming Huang, Ding Wang 0001
Neurocomputing2
2024 Evolution-guided value iteration for optimal tracking control
Haiming Huang, Ding Wang 0001, Qinna Hu
Neurocomputing2
2024 Adjustable iterative Q-learning for advanced neural tracking control with stability guarantee
Yuan Wang 0056, Ding Wang 0001, Ao Liu 0012, Junfei Qiao 0001
Neurocomputing2
2024 Advanced optimal tracking integrating a neural critic technique for asymmetric constrained zero-sum games
Menghua Li, Ding Wang 0001, Junfei Qiao 0001
Neural Networks2
2024 Neural Q-learning for discrete-time nonlinear zero-sum games with adjustable convergence rate
Yuan Wang 0056, Ding Wang 0001, Junfei Qiao 0001
Neural Networks2
2024 Neural critic learning with accelerated value iteration for nonlinear model predictive control
Peng Xin, Ding Wang 0001, Ao Liu 0012, Junfei Qiao 0001
Neural Networks2
2024 Adaptive Event-Triggered Control for Non-Strict Feedback Nonlinear CPSs With Time Delays Against Deception Attacks and Actuator Faults
abstract
An adaptive event-triggered control strategy for non-strict feedback nonlinear cyber-physical systems with time delays against deception attacks and actuator faults is presented. The most prominent difficulty lies in after the system signals are damaged by malicious deception attacks, all the exact states in the system are unavailable. Another design difficulty is the common existence of unknown nonlinearities and time-varying time delays, which makes it difficult to get the desired controller. Thus, to stabilize the studied cyber-physical systems, an unusual coordinate transformation is presented, where the attack gains and the problem of unavailability of states are considered simultaneously. Furthermore, the effect of the malicious deception attacks on the studied system is tackled by using Nussbaum technology and designing the Lyapunov functions with the attack gains. The unknown nonlinear functions, the mismatch problem for the control input caused by the nonlinearities of time-varying time delays are processed by neural networks technology and Lyapunov-Krasovskii functions. As a result, a controller based on event-triggered mechanism is constructed to reduce the waste of communication resources. The presented control strategy can ensure the signals in the closed-loop system are bounded. Finally, the Matlab simulation experiments are showed to verify the effectiveness of the developed strategy.Note to Practitioners—This paper studies the adaptive event-triggered control strategy for non-strict feedback nonlinear cyber-physical systems with time delay under deception attacks and actuator faults. Cyber-physical systems are widely used in power system, aviation system and other practical systems. However, due to the openness of network communication channels, attackers are prone to transmitting incorrect data to controllers and actuators, resulting in unusable system states, which is a challenging problem to ensure system stability by only using the compromised system states. In addition, the safety control problem of cyber-physical systems with time delays and actuator faults is investigated simultaneously, which poses significant challenges to the design of an desired controller. Therefore, by constructing Lyapunov functions and designing adaptive laws, not only the system under deception attacks is stabilized, but also the control strategy studied is made more practical.
Wen-Di Chen, Ben Niu 0003, Huanqing Wang 0001, Haitao Li 0001, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.5
2024 Offline Data-Driven Adaptive Critic Design With Variational Inference for Wastewater Treatment Process Control
abstract
Wastewater treatment is indispensable to the functioning of urban society, and its optimal control has enormous social benefits. However, precise modelling of the unstable and complex treatment process is challenging yet crucial to the adaptive dynamic programming method. In this article, an adaptive critic algorithm with variational inference is designed to address the optimal control problem of nonlinear discrete-time systems, along with the convergence analysis. Based on the recorded system trajectory, the variational autoencoder is utilized to approximate the behavior policy of the offline dataset without system modelling and online interaction. Through policy iteration learning, the actor-critic structure can amend the policy generated by the variational autoencoder to achieve the optimal control objective. Simulations on a nonlinear system and the wastewater treatment process have verified that the proposed approach outperformed the behavior policy. Driven by the wastewater treatment process data derived from the incremental proportional-integral-derivative controller, the proposed approach can produce an optimal control policy of less tracking error and cost.Note to Practitioners—When dealing with an unknown system with complex dynamics, it is more feasible to improve the acceptable performance of the existing control policy based on the system’s trajectory than to obtain an excelling policy. Motivated by batch reinforcement learning, learning from offline data can avoid the online interaction between the system and the adaptive dynamic programming algorithm, which could lead to exploratory errors during online learning. Specifically, using a model-free adaptive dynamic programming algorithm, the parameters of the controller are instantly updated based on the experience replay buffer sampled from the online trajectory data. However, online exploration determines the update, and there is no guarantee that the system will converge every time. As a specific type of adaptive dynamic programming algorithm, adaptive critic design uses a critic network to approximate the expected future cost and an actor network to generate a control input that minimizes the expected future cost. In this article, using the converged trajectory as the offline dataset, a revised variational autoencoder is used to approximate the behavior policy of the offline dataset. As a generative model, the variational autoencoder considers a random variable that adheres to a prior distribution while producing outputs. Through offline learning, the actor network can amend the approximated policy based on the evaluation from the critic network while being constrained within the limited variation of the generative model. Finally, the objective of the optimal control task can be achieved by following the designated cost design. However, a dataset containing disturbances could impede offline learning, which needs to be addressed.
Junfei Qiao 0001, Ruyue Yang, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.3
2024 Adaptive Finite-Time Bipartite Consensus Tracking Control for Heterogeneous Nonlinear MASs With Time-Varying Output Constraints
abstract
In this paper, an adaptive finite-time bipartite consensus tracking control strategy is presented for a class of heterogeneous nonlinear nonstrict-feedback multi-agent systems (MASs) with output constraints. Firstly, to deal with the time-varying output constraints problem, an improved tan-type nonlinear mapping (NM) function is presented for the first time. And based on the improved NM function, a novel tracking error is constructed to design controller for each agent, which guarantees the bipartite consensus tracking is achieved while constraints requirement is not violated. Then, a state observer is designed to estimate the unmeasurable states of each agent. Moreover, in the case of unbalanced directed topological graph, a partition algorithm (PA) is employed to implement bipartite consensus tracking control. The developed distributed adaptive finite-time control strategy ensures that all the signals in the closed-loop system are bounded and the bipartite consensus tracking control is achieved in finite time. Finally, the validity of the designed control strategy is demonstrated by a simulation experiment.Note to Practitioners—At present, nonlinear MASs are widely used in practice, such as robots formation control, vehicular platoon systems control, etc. This paper investigated the adaptive finite-time bipartite consensus tracking control problem for a class of heterogeneous nonlinear nonstrict-feedback MASs with output constraints. In the scenarios of practical application, these two situations are common: 1) The communication topology graph of nonlinear MASs is unbalanced. 2) The output of each agent is constrained. Therefore, this paper presents an improved tan-type NM method to deal with the time-varying output constraints problem, and a partition algorithm is employed to implement bipartite consensus tracking control based on the unbalanced communication topology graph. Meanwhile, the nonsingular finite-time control strategy effectively improves the convergence of the studied nonlinear MASs. In addition, the system model and backstepping technology used in this paper are general and practical.
Zihao Shang, Yuqiang Jiang, Ben Niu 0003, Xudong Zhao 0001, Ding Wang 0001, Bin Li 0005
IEEE Trans Autom. Sci. Eng.5
2024 Adaptive Event-Triggered Consensus Tracking Control Schemes for Uncertain Constrained Nonlinear Multi-Agent Systems
abstract
The majority of the results on constrained nonlinear multi-agent systems (MASs) control focused on output or state constraints without considering the saving of communication resources. In this paper, for a class of uncertain nonlinear MASs, we first present a new adaptive bounded consensus tracking control scheme in which the asymmetric and full-state constraints are jointly synthesized with a switching threshold event-triggered strategy, such that the communication resources are effectively utilized. The key to accomplishing the asymmetric and full-state constraints is that a kind of improved$tan$-type barrier Lyapunov functions are constructed. The controller constructed for each agent by the switching threshold event-triggered strategy guarantees that the asymmetric and full-state constraints are not violated and the output of each agent can track the leader’s trajectory with an adjustable bounded tracking error. Furthermore, to achieve the asymptotic consensus tracking control, we give another kind of novel$tan$-type barrier Lyapunov functions to design the desired controller for each agent. A simulation example of five single-link robots is proposed to illustrate the effectiveness of our control scheme.Note to Practitioners—In this paper, the adaptive bounded consensus tracking control problem is considered for nonlinear MASs subject to full-state constraints, whose models are capable of describing a multitude of critical applications, including the formation of unmanned vehicles and robots. The research on the tracking control problem will be rather complicated yet challenging if the asymmetric and full-state constraints are taken into account in a complex environment. Additionally, the incorporation of an event-triggered mechanism aids in the reduction of communication resource usage, enhancing the ease of implementation and improving the user-friendliness of the proposed control scheme.
Ben Niu 0003, Jiaming Zhang 0003, Huanqing Wang 0001, Yuqiang Jiang, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.6
2024 Adaptive Critic Tracking Design for Data-Based Nonaffine Predictive Control
abstract
In recent years, model predictive control (MPC) is widely utilized to address the tracking problem of the practical industrial processes. In this paper, in terms of the advantages of adaptive dynamic programming (ADP), the adaptive critic trajectory tracking predictive control (ACTTPC) framework is designed to tackle tracking predictive control problems for unknown nonaffine systems. First, the unknown system dynamics are approximated by the established model network. Meanwhile, the feedforward steady control is considered to assist with accomplishing the tracking mission. Further, in each prediction horizon, the adaptive critic learning method is utilized to solve the open-loop optimization problem satisfying some conditions. Afterwards, the Lyapunov stability of the augmented error system is fully proved, and the convergence of the ACTTPC algorithm is analyzed in detail. Finally, a nonaffine system and a torsional pendulum plant are applied to validate the effectiveness of the presented approach in solving the tracking problems.Note to Practitioners—Many industrial processes are nonlinear nonaffine systems, which causes a great challenge to solve Hamilton-Jacobi-Bellman (HJB) equations for nonlinear MPC (NMPC). Therefore, it is quite valuable to solve the NMPC problem by using the advantages of ADP in addressing nonlinear HJB equations. In this paper, the ACTTPC algorithm is designed to guide the trajectory tracking predictive control for unknown system dynamics in the practical industrial processes. Generally speaking, the mathematical model of complex industrial systems is difficultly established. Hence, the model network is built via selecting a batch of data and it is seen as the prediction model. The introduced feedforward steady control can not only assist realizing trajectory tracking but also maintain stable tracking effect. Meanwhile, the feedback predictive control input is solved via the ACTTPC algorithm. The simulation experiments are conducted to prove the effectiveness of the presented algorithm. Moreover, the pseudo-code and the relevant experimental parameters are given. For different systems and reference trajectories, the practitioners can realize tracking tasks via modulating the related parameters.
Ding Wang 0001, Peng Xin, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.1
2024 Event-Triggered-Based Consensus Neural Network Tracking Control for Nonlinear Pure-Feedback Multiagent Systems With Delayed Full-State Constraints
abstract
This paper investigates the consensusability for a class of nonlinear pure-feedback multiagent systems (MASs) with delayed full-state constraints. First, by using a novel state-shifting transformation (SST), the constrained states are encapsulated into some barrier functions such that the new converted state variables and their dynamic model are generated. Then, the essential relationship between the post-transformation tracking errors and the pre-transformation tracking errors are further explained. By employing the radial basis function neural networks (RBF NNs) to deal with the unknown nonlinearities for each agent and introducing a relative threshold method to overcome the waste problem of system transmission resources, an event-triggered control protocol is developed for the considered system. The proposed protocol has its own advantages: 1) This is the first work to consider the pure-feedback MASs with delayed full-state constraints. By transforming the constrained states into unconstrained variables, the states of the MASs can be converged into ideal ranges within a predefined time$T$; 2) Under this control protocol, only one adaptive law is necessary to be updated online in each follower and the use efficiency of system resources is improved by using reasonable event triggering mechanism. Finally, the effectiveness of the suggested consensus control protocol is demonstrated by the simulation results.Note to Practitioners—Delayed constraints often encountered in practical applications, which represent a type of constraints that may be violated initially but required to be satisfied sometime after system operation, so it is very necessary and important to converge states into ideal ranges within a predefined time when delayed constraints required. What’s more, considering the requirement of high control accuracy in practical application and the current Barrier Lyapunov Function (BLF) methods need to restrict the virtual control signals within some given regions, such restrictive conditions make the design and implementation of the traditional controllers particularly difficult. In this article, the control protocol for a class of nonlinear pure-feedback MASs with delayed full-state constraints and event-triggered communication are constructed to promote the development of asymptotic tracking control methods, and a novel state-shifting transformation is used to ensure the system states could be converged into the ideal constrained intervals in a predefined time.
Xiao-An Wang, Guangju Zhang, Ben Niu 0003, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.4
2024 Novel Discounted Adaptive Critic Control Designs With Accelerated Learning Formulation
abstract
Inspired by the successive relaxation method, a novel discounted iterative adaptive dynamic programming framework is developed, in which the iterative value function sequence possesses an adjustable convergence rate. The different convergence properties of the value function sequence and the stability of the closed-loop systems under the new discounted value iteration (VI) are investigated. Based on the properties of the given VI scheme, an accelerated learning algorithm with convergence guarantee is presented. Moreover, the implementations of the new VI scheme and its accelerated learning design are elaborated, which involve value function approximation and policy improvement. A nonlinear fourth-order ball-and-beam balancing plant is used to verify the performance of the developed approaches. Compared with the traditional VI, the present discounted iterative adaptive critic designs greatly accelerate the convergence rate of the value function and reduce the computational cost simultaneously.
Mingming Ha, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Cybern.2
2024 Intelligent-Critic-Based Tracking Control of Discrete-Time Input-Affine Systems and Approximation Error Analysis With Application Verification
abstract
In recent years, the application of function approximators, such as neural networks and polynomials, has ushered in a new stage of development in solving optimal control problems. However, considering the existence of approximation errors, the stability of the controlled system cannot be guaranteed. Therefore, in view of the prevalence of approximation errors, we investigate optimal tracking control problems for discrete-time systems. First, a novel value function is introduced into the intelligent critic framework. Second, an implicit method is utilized to demonstrate the boundedness of the iterative value functions with approximation errors. An explicit method is applied to prove the stability of the system with approximation errors. Furthermore, an evolving policy is designed to iteratively tackle the optimal tracking control problem and demonstrate the stability of the system. Finally, the effectiveness of the developed method is verified through numerical as well as practical examples.
Ding Wang 0001, Ning Gao 0008, Mingming Ha, Junfei Qiao 0001
IEEE Trans. Cybern.1
2024 Adaptive Fuzzy Practical Predefined-Time Bipartite Consensus Tracking Control for Heterogeneous Nonlinear MASs With Actuator Faults
abstract
This paper focuses on the adaptive fuzzy practical predefined-time (PT) bipartite consensus tracking control (BCTC) problem for heterogeneous nonlinear multi-agent systems (HNMASs) with actuator faults. First, the fuzzy logic systems (FLSs) are used to approximate the unknown nonlinear functions. Then, the partial loss of effectiveness and bias fault of actuator are considered simultaneously in the HNMASs, which is effectively handled by using adaptive compensation technology. In addition, the developed bipartite consensus tracking control protocol based on a practical predefined-time strategy not only ensures the fast convergence of the studied systems, but also predetermines the convergence time not relied on the initial conditions. The theoretical result shows that all signals of the closed-loop system are semiglobally uniformly predefined-time bounded (SGUPTB), and the BCTC performance is guaranteed within the predefined time. Finally, the simulation example based on the agents of different orders shows the validity of the obtained results.
Ben Niu 0003, Jihang Sui, Xudong Zhao 0001, Ding Wang 0001, Xinliang Zhao
IEEE Trans. Fuzzy Syst.4
2024 Action-Dependent Heuristic Dynamic Programming With Experience Replay for Wastewater Treatment Processes
abstract
The wastewater treatment process (WWTP) is beneficial for maintaining sufficient water resources and recycling wastewater. A crucial link of WWTP is to ensure that the dissolved oxygen (DO) concentration is continuously maintained at the predetermined value, which can actually be considered as a tracking problem. In this article, an experience replay-based action-dependent heuristic dynamic programming (ER-ADHDP) method is developed to design the model-free tracking controller to accomplish the tracking goal of the DO concentration. First, the online ER-ADHDP controller is regarded as a supplementary controller to conduct the model-free tracking control alongside a stabilizing controller with a priori knowledge. The online ER-ADHDP method can adaptively adjust weight parameters of critic and action networks, thereby continuously ameliorating the tracking result over time. Second, the ER technique is integrated into the critic and action networks to promote the data utilization efficiency and accelerate the learning process. Third, a rational stability result is provided to theoretically ensure the usefulness of the ER-ADHDP tracking design. Finally, simulation experiments including different reference trajectories are conducted to show the superb tracking performance and excellent adaptability of the proposed ER-ADHDP method.
Junfei Qiao 0001, Ding Wang 0001, Menghua Li
IEEE Trans. Ind. Informatics3
2024 Adaptive Critic Control Design With Knowledge Transfer for Wastewater Treatment Applications
abstract
The wastewater treatment process (WWTP) is of great significance to environmental protection. To improve the efficiency of the WWTP, it is crucial to ensure that the dissolved oxygen (DO) concentration tracks the set value efficiently. Due to the nonlinear and time-varying dynamics of the WWTP, traditional control methods cannot accurately control the DO concentration. To overcome these challenges, this article proposes an online transferred heuristic dynamic programming (TrHDP) control design by combining transfer learning with adaptive critic design. First, we use the historical sample data to construct a mathematical model of the WWTP and learn the prior knowledge from the model. Then, the online control process of the DO concentration is guided by utilizing the prior knowledge. In order to avoid negative transfer and save computing resources, we design a novel decay function with the truncation mechanism. In addition, we prove the stability of the TrHDP control scheme by constructing a Lyapunov function. Finally, the performance of the TrHDP scheme is verified by the Benchmark Simulation Model No. 1. Compared with other methods, the TrHDP method possesses higher control accuracy for the DO concentration and overcomes the disadvantage of low learning efficiency of general online methods.
Ding Wang 0001, Xin Li 0055, Junfei Qiao 0001
IEEE Trans. Ind. Informatics1
2024 Asymmetric Constrained Optimal Tracking Control With Critic Learning of Nonlinear Multiplayer Zero-Sum Games
abstract
By utilizing a neural-network-based adaptive critic mechanism, the optimal tracking control problem is investigated for nonlinear continuous-time (CT) multiplayer zero-sum games (ZSGs) with asymmetric constraints. Initially, we build an augmented system with the tracking error system and the reference system. Moreover, a novel nonquadratic function is introduced to address asymmetric constraints. Then, we derive the tracking Hamilton-Jacobi-Isaacs (HJI) equation of the constrained nonlinear multiplayer ZSG. However, it is extremely hard to get the analytical solution to the HJI equation. Hence, an adaptive critic mechanism based on neural networks is established to estimate the optimal cost function, so as to obtain the near-optimal control policy set and the near worst disturbance policy set. In the process of neural critic learning, we only utilize one critic neural network and develop a new weight updating rule. After that, by using the Lyapunov approach, the uniform ultimate boundedness stability of the tracking error in the augmented system and the weight estimation error of the critic network is verified. Finally, two simulation examples are provided to demonstrate the efficacy of the established mechanism.
Junfei Qiao 0001, Menghua Li, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Advanced Optimal Tracking Control With Stability Guarantee via Novel Value Learning Formulation
abstract
In this article, to solve the optimal tracking control problem (OTCP) for discrete-time (DT) nonlinear systems, general value iteration (GVI) scheme and online value iteration (VI) algorithms with novel value function are discussed. First, the disadvantage of the traditional value function for the OTCP is presented and the novel value function is introduced. Second, we analyze the monotonicity and convergence of GVI and establish the admissibility condition of GVI to evaluate the admissibility of the current iterative control. Note that a novel approach is introduced to analyze the admissibility. Third, based on the attraction domain, improved control policies with online VI can be obtained by judging the location of the current tracking error and reference point. Finally, the stability of the online VI-based control system is guaranteed. Besides, we provide two simulation examples to show the performance of the proposed methods.
Ding Wang 0001, Mingming Ha, Menghua Li, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Convergence and Stability of Optimal Regulation via Generalized N-Step Value Gradient Learning
abstract
In this article, the generalized N -step value gradient learning (GNSVGL) algorithm, which takes a long-term prediction parameter λ into account, is developed for infinite horizon discounted near-optimal control of discrete-time nonlinear systems. The proposed GNSVGL algorithm can accelerate the learning process of adaptive dynamic programming (ADP) and has a better performance by learning from more than one future reward. Compared with the traditional N -step value gradient learning (NSVGL) algorithm with zero initial functions, the proposed GNSVGL algorithm is initialized with positive definite functions. Considering different initial cost functions, the convergence analysis of the value-iteration-based algorithm is provided. The stability condition for the iterative control policy is established to determine the value of the iteration index, under which the control law can make the system asymptotically stable. Under such a condition, if the system is asymptotically stable at the current iteration, then the iterative control laws after this step are guaranteed to be stabilizing. Two critic neural networks and one action network are constructed to approximate the one-return costate function, the λ -return costate function, and the control law, respectively. It is emphasized that one-return and λ -return critic networks are combined to train the action neural network. Finally, via conducting simulation studies and comparisons, the superiority of the developed algorithm is confirmed.
Ding Wang 0001, Mingming Ha, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Cooperative ETM-Based Adaptive Neural Network Tracking Control for Nonlinear Pure-Feedback MASs: A Special-Shaped Laplacian Matrix Method
abstract
This article solves the cooperative adaptive tracking control problem for nonlinear pure-feedback multi-agent systems (MASs). Compared with the previous achievements of adaptive control of pure-feedback MASs, the partial derivative of the nonaffine function may not exist by using decoupling technology. In the controller design framework based on the backstepping technique, the additional state variables are processed using the special properties of the radial basis function neural networks (RBF NNs). A special-shaped Laplacian matrix is proposed to unify the leader gain form in the tracking error design process (the coefficient in the second term of tracking error). Furthermore, an event trigger mechanism (ETM) is introduced to save resources. The constructed controller under the ETM can not only stabilize the system states but also make the tracking error reach a small accuracy. Finally, the simulation results demonstrated the feasibility of the proposed method.
Qiangqiang Zhu, Ben Niu 0003, Ding Wang 0001, Shengtao Li
IEEE Trans. Neural Networks Learn. Syst.3
2024 Decentralized Event-Triggered Asymmetric Constrained Control Through Adaptive Critic Designs for Nonlinear Interconnected Systems
abstract
In this article, a decentralized event-triggered control mechanism is established to solve the interconnected issue of continuous-time nonlinear systems with asymmetric input constraints and matched interconnections based on the adaptive critic technology. First, by inserting the discount factor, a novel nonquadratic cost function is constructed for the constrained subsystem with nonzero equilibrium point. Meanwhile, the decentralized event-triggered control issue is transformed into a set of optimal control issues. Then, the execution of nominal subsystem is based on the event-triggered mechanism (ETM) with an event-triggering condition which increases the algorithm efficiency. Moreover, we derive the associated event-triggered Hamilton–Jacobi–Bellman (HJB) equation which arising in the discounted-cost optimal event-triggered control issues of nominal subsystems. In the implementation, an adaptive critic framework is employed to approximate the optimal cost function. Later, the experience replay (ER) approach is introduced into a novel weight tuning mechanism, which converts the traditional persistence of excitation (PE) condition into an easy-checked rank condition. Theoretically, the stability of the system and the exclusion of Zeno behavior are demonstrated. Finally, one representative example is simulated to validate the efficacy of the constructed framework.
Ding Wang 0001, Menghua Li, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Adjustable Iterative Q-Learning Schemes for Model-Free Optimal Tracking Control
abstract
This article puts emphasis on the deterministic value-iteration-based$Q$-learning (VIQL) algorithm with adjustable convergence speed, followed by the application verification on trajectory tracking for completely unknown nonaffine systems. It is worth emphasizing that, under the effect of learning rates, the convergence speed can be adjusted and the new convergence criterion of the VIQL framework is investigated. The merit of the adjustable VIQL scheme is that it can quicken the learning speed and decrease the number of iterations, thereby reducing the computation burden. To carry out the model-free VIQL algorithm, the offline data of system states and reference trajectories are collected to provide the reference control, the tracking error, and the tracking control, which promotes the parameter updating of the adjustable VIQL algorithm via the off-policy learning scheme. By this updating operation, the convergent optimal tracking policy can guarantee that arbitrary initial state tracks the desired trajectory and can completely obviate the terminal tracking error. Finally, numerical simulations are conducted to indicate the validity of the designed tracking control algorithm.
Junfei Qiao 0001, Ding Wang 0001, Mingming Ha
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Evolution-Guided Adaptive Dynamic Programming for Nonlinear Optimal Control
abstract
In this article, an evolution-guided adaptive dynamic programming (EGADP) algorithm is developed to address the optimal regulation problems for the nonlinear systems. In the traditional adaptive dynamic programming algorithms, policy improvement is typically reliant on the gradient information, according to the first order necessity condition. However, these methods encounter limitations when calculating the gradient information becomes infeasible or system dynamics is not differentiable. In response to this challenge, the evolutionary computation is harnessed by EGADP to search for a superior policy during policy improvement. Therefore, compared with the traditional methods, scenarios that gradient information is unavailable can effectively be handled by EGADP. Additionally, the convergence of the algorithm is proven to enhance the rigorousness of the developed method. Finally, the three simulation experiments with realistic physical backgrounds are conducted to comprehensively demonstrate the effectiveness of the established method from different perspectives.
Ding Wang 0001, Haiming Huang, Derong Liu 0001, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Novel Parallel Formulation for Iterative Reinforcement Learning Control
abstract
Parallelization is widely employed to improve the exploration ability of controllers. However, it is rare to provide a lightweight scheme for reducing homogeneous policies with theoretical guarantees. This article is concerned with a novel parallel scheme for solving optimal control problems. In brief, we design a novel global indicator that inherits the theoretical guarantees of a class of iterative reinforcement learning algorithms. By generating a tentative function, the global indicator can guide and communicate with parallel controllers to accelerate the learning process. Using two typical exploration policies, the novel parallel scheme can rapidly compress the neighborhood of the optimal cost function. Besides, two parallel algorithms based on value iteration and Q-learning are established to improve the data efficiency through different extensions. Finally, two benchmark problems are presented to demonstrate the learning effectiveness of the novel parallel scheme.
Ding Wang 0001, Jiangyu Wang, Lingzhi Hu, Liguo Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Intelligent Optimal Control of Constrained Nonlinear Systems via Receding-Horizon Heuristic Dynamic Programming
abstract
For addressing the approximate optimal control problem of nonlinear affine systems with the terminal state constraint and asymmetric control constraints, the constrained receding-horizon heuristic dynamic programming (RH-HDP) algorithm is established in this article. In consideration of the RH mechanism of model predictive control (MPC), the approximate optimal control problem based on the HDP algorithm is transformed into a battery of subproblems. Then, the terminal state constraint related with the current prediction horizon is considered such that the terminal state is forced into the neighborhood of the system equilibrium point. In addition, the asymmetric control constraints are introduced to release the pressure of actuator saturation, so that the control input is well confined within the given constraint range. Meanwhile, relevant results of the stability proof are also displayed based on the Lyapunov approach. Finally, the constrained RH-HDP algorithm has been applied in two kinds of systems to verify its effectiveness. Comparative experiments with the traditional HDP algorithm have been carried out to verify the superiority of the present algorithm.
Ding Wang 0001, Peng Xin, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2023 Data-driven tracking control design with reinforcement learning involving a wastewater treatment application
Ding Wang 0001, Xin Li 0055, Lingzhi Hu, Junfei Qiao 0001
Eng. Appl. Artif. Intell.1
2023 Event-based online learning control design with eligibility trace for discrete-time unknown nonlinear systems
Ding Wang 0001, Jiangyu Wang, Lingzhi Hu
Eng. Appl. Artif. Intell.1
2023 Event-triggered constrained neural critic control of nonlinear continuous-time multiplayer nonzero-sum games
Menghua Li, Ding Wang 0001, Junfei Qiao 0001
Inf. Sci.2
2023 Discounted linear Q-learning control with novel tracking cost and its stability
Ding Wang 0001, Mingming Ha
Inf. Sci.1
2023 Dichotomy value iteration with parallel learning design towards discrete-time zero-sum games
Jiangyu Wang, Ding Wang 0001, Xin Li 0055, Junfei Qiao 0001
Neural Networks2
2023 Off-Policy Model-Free Learning for Multi-Player Non-Zero-Sum Games With Constrained Inputs
abstract
In this paper, multi-player non-zero-sum games with control constraints are studied by utilizing a novel model-free approach based on adaptive dynamic programming framework. First, the model-based policy iteration (PI) method is provided, which requires the system dynamics, and the convergence is demonstrated. Then, aiming to eliminate the need for the system dynamics, a model-free iterative method is obtained by using the off-policy integral reinforcement learning (IRL) scheme based on the PI approach. Moreover, the system data is collected in order to construct the model-free approach. Besides, we analyze the convergence of the off-policy IRL approach by proving the equivalence between the model-free iterative approach and the model-based iterative approach. Remarkably, in the implementation of the scheme, the control policy and cost function are approximated by utilizing the actor-critic networks. The least square algorithm is utilized to learn the actor-critic networks weights depended on the collected data sets. Finally, two cases are provided to demonstrate the effectiveness of the established framework.
Ding Wang 0001, Junfei Qiao 0001, Menghua Li
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Switching Event-Triggered Adaptive Resilient Dynamic Surface Control for Stochastic Nonlinear CPSs With Unknown Deception Attacks
abstract
This work concentrates on the adaptive resilient dynamic surface controller design problem for uncertain nonlinear lower triangular stochastic cyber-physical systems (CPSs) subject to unknown deception attacks based on a switching threshold event-triggered mechanism. The adverse effect of deception attacks on the stochastic CPSs is that the exact system state variables become unavailable. Furthermore, it should be emphasized that the coexistence of unknown nonlinearities, stochastic perturbations, and unknown sensor and actuator attacks makes it a very difficult and challenging event to implement the control design. To get the desired controller, radial basis function (RBF) neural networks (NNs) are introduced so that the design obstacle caused from the unknown nonlinearities can be easily solved. On this basis, in order to save resources and effectively transmit, the event-triggered control scheme based on a switching threshold strategy is further considered. In the backstepping design process, the dynamic surface control (DSC) technique is presented to deal with the issue of "explosion of complexity." By skillfully designing a new coordinate transformation and the attack compensators, the problem of unknown deception attacks is successfully handled. Under our proposed control scheme, all the closed-loop signals are bounded in probability and the stabilization errors converge to an adjustable neighborhood of the origin in probability. Finally, the simulation results on the double chemical reactor show the validity of the proposed design scheme.
Ben Niu 0003, Huanqing Wang 0001, Ding Wang 0001, Xudong Zhao 0001
IEEE Trans. Cybern.5
2023 Evolving and Incremental Value Iteration Schemes for Nonlinear Discrete-Time Zero-Sum Games
abstract
In this article, evolving and incremental value iteration (VI) frameworks are constructed to address the discrete-time zero-sum game problem. First, the evolving scheme means that the closed-loop system is regulated by using the evolving policy pair. During the control stage, we are committed to establishing the stability criterion in order to guarantee the availability of evolving policy pairs. Second, a novel incremental VI algorithm, which takes the historical information of the iterative process into account, is developed to solve the regulation and tracking problems for the nonlinear zero-sum game. Via introducing different incremental factors, it is highlighted that we can adjust the convergence rate of the iterative cost function sequence. Finally, two simulation examples, including linear and nonlinear systems, are conducted to demonstrate the performance and the validity of the proposed evolving and incremental VI schemes.
Ding Wang 0001, Mingming Ha, Junfei Qiao 0001
IEEE Trans. Cybern.2
2023 Time-/Event-Triggered Adaptive Neural Asymptotic Tracking Control of Nonlinear Interconnected Systems With Unmodeled Dynamics and Prescribed Performance
abstract
This article proposes two adaptive asymptotic tracking control schemes for a class of interconnected systems with unmodeled dynamics and prescribed performance. By applying an inherent property of radial basis function (RBF) neural networks (NNs), the design difficulties aroused from the unknown interactions among subsystems and unmodeled dynamics are overcome. Then, in order to ensure that the tracking errors can be suppressed in the specified range, the constrained control problem is transformed into the stabilization problem by using an auxiliary function. Based on the adaptive backstepping method, a time-triggered controller is constructed. It is proven that under the framework of Barbalat's lemma, all the variables in the closed-loop system are bounded and the tracking errors are further ensured to converge to zero asymptotically. Furthermore, the event-triggered strategy with a variable threshold is adopted to make more precise control such that the better system performance can be obtained, which reduces the system communication burden under the condition of limited communication resources. Finally, an illustrative example is provided to demonstrate the effectiveness of the proposed control scheme.
Ben Niu 0003, Jiaming Zhang 0003, Ding Wang 0001, Zhenhua Wang 0004
IEEE Trans. Neural Networks Learn. Syst.4
2023 A Novel Value Iteration Scheme With Adjustable Convergence Rate
abstract
In this article, a novel value iteration scheme is developed with convergence and stability discussions. A relaxation factor is introduced to adjust the convergence rate of the value function sequence. The convergence conditions with respect to the relaxation factor are given. The stability of the closed-loop system using the control policies generated by the present VI algorithm is investigated. Moreover, an integrated VI approach is developed to accelerate and guarantee the convergence by combining the advantages of the present and traditional value iterations. Also, a relaxation function is designed to adaptively make the developed value iteration scheme possess fast convergence property. Finally, the theoretical results and the effectiveness of the present algorithm are validated by numerical examples.
Mingming Ha, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Neuro-Optimal Trajectory Tracking With Value Iteration of Discrete-Time Nonlinear Dynamics
abstract
In this article, a novel neuro-optimal tracking control approach is developed toward discrete-time nonlinear systems. By constructing a new augmented plant, the optimal trajectory tracking design is transformed into an optimal regulation problem. For discrete-time nonlinear dynamics, the steady control input corresponding to the reference trajectory is given. Then, the value-iteration-based tracking control algorithm is provided and the convergence of the value function sequence is established. Therein, the approximation error between the iterative value function and the optimal cost is estimated. The uniformly ultimately bounded stability of the closed-loop system is also discussed in detail. Moreover, the iterative heuristic dynamic programming (HDP) algorithm is implemented by involving the critic and action components, where some new updating rules of the action network are provided. Finally, two examples are used to demonstrate the optimality of the present controller as well as the effectiveness of the proposed method.
Ding Wang 0001, Mingming Ha, Long Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Adaptive Critic for Event-Triggered Unknown Nonlinear Optimal Tracking Design With Wastewater Treatment Applications
abstract
In this article, an event-based near-optimal tracking control algorithm is developed for a class of nonaffine systems. First, in order to gain the tracking control strategy, the costate function is established through the iterative dual heuristic dynamic programming (DHP) algorithm. Then, the event-based control method is employed to improve the utilization efficiency of resources and ensure that the closed-loop system has an excellent control performance. Meanwhile, the input-to-state stability (ISS) is proven for the event-based tracking plant. In addition, three kinds of neural networks are used in the event-based DHP algorithm, which aims to identify the nonaffine nonlinear system, estimate the costate function, and approximate the tracking control law. Finally, a numerical experimental simulation is conducted to verify the effectiveness of the proposed scheme. Moreover, in order to further validate the feasibility, the algorithm is applied to the wastewater treatment plant to effectively control the concentrations of dissolved oxygen and nitrate nitrogen.
Ding Wang 0001, Lingzhi Hu, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 System Stability of Learning-Based Linear Optimal Control With General Discounted Value Iteration
abstract
For discounted optimal regulation design, the stability of the controlled system is affected by the discount factor. If an inappropriate discount factor is employed, the optimal control policy might be unstabilizing. Therefore, in this article, the effect of the discount factor on the stabilization of control strategies is discussed. We develop the system stability criterion and the selection rules of the discount factor with respect to the linear quadratic regulator problem under the general discounted value iteration algorithm. Based on the monotonicity of the value function sequence, the method to judge the stability of the controlled system is established during the iteration process. In addition, once some stability conditions are satisfied at a certain iteration step, all control policies after this iteration step are stabilizing. Furthermore, combined with the undiscounted optimal control problem, the practical rule of how to select an appropriate discount factor is constructed. Finally, several simulation examples with physical backgrounds are conducted to demonstrate the present theoretical results.
Ding Wang 0001, Mingming Ha, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Stability and Admissibility Analysis for Zero-Sum Games Under General Value Iteration Formulation
abstract
In this article, the general value iteration (GVI) algorithm for discrete-time zero-sum games is investigated. The theoretical analysis focuses on stability properties of the systems and also the admissibility properties of the iterative policy pair. A new criterion is established to determine the admissibility of the current policy pair. Besides, based on the admissibility criterion, the improved GVI algorithm toward zero-sum games is developed to guarantee that all iterative policy pairs are admissible if the current policy pair satisfies the criterion. On the basis of the attraction domain, we demonstrate that the state trajectory will stay in the region using the fixed or the evolving policy pair if the initial state belongs to the domain. It is emphasized that the evolving policy pair can stabilize the controlled system. These theoretical results are applied to linear and nonlinear systems via offline and online critic control design.
Ding Wang 0001, Mingming Ha, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Adaptive Neural Control of Nonlinear Nonstrict Feedback Systems With Full-State Constraints: A Novel Nonlinear Mapping Method
abstract
In this work, a neural-networks (NNs)-based adaptive asymptotic tracking control scheme is presented for a class of uncertain nonstrict feedback nonlinear systems with time-varying full-state constraints. First, we construct a novel exponentially decaying nonlinear mapping to map the constrained system states to new system states without constraints. Instead of the traditional barrier Lyapunov function methods, the feasible conditions which require the virtual control signals satisfying the constraint requirements are removed. By employing the Nussbaum design method to eliminate the effect of unknown control gains, the general assumption about the signs of the unknown control gains is relaxed. Then, the nonstrict feedback form of the system can be pulled back to the strict feedback form through the basic properties of radial basis function NNs. Simultaneously, the intermediate control signals and the desired controller are constructed by the backstepping process and the Nussbaum design method. The designed controller can ensure that all signals in the whole closed-loop system are bounded without the violation of the constraints and hold the asymptotic tracking performance. In the end, a practical example about a brush dc motor driving a one-link robot manipulator is given to illustrate the effectiveness of the proposed design scheme.
Jiaming Zhang 0003, Ben Niu 0003, Ding Wang 0001, Huanqing Wang 0001, Peiyong Duan, Guangdeng Zong
IEEE Trans. Neural Networks Learn. Syst.3
2023 Dual Event-Triggered Constrained Control Through Adaptive Critic for Discrete-Time Zero-Sum Games
abstract
In this article, through adaptive critic, a dual event-triggered (DET) constrained control scheme is established for discrete-time nonlinear zero-sum games. The neural networks are trained from the dual heuristic dynamic programming technique to obtain the approximate optimal policy pair. Two corresponding independent triggering conditions are constructed for the control input and the disturbance to improve the utilization efficiency and ensure the independence between them. In addition, in order to overcome the challenge caused by the actuator saturation, we constrain the control input to a bounded range. Meanwhile, the asymptotically stability is proved for the DET control system. Finally, experimental simulations are conducted to verify the effectiveness of the proposed algorithm.
Ding Wang 0001, Lingzhi Hu, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2023 Novel Discounted Optimal Tracking Design Under Offline and Online Formulations for Asymmetric Constrained Systems
abstract
In this thesis, we construct improved value iteration (VI) and online VI structures, in a bid to tackle the optimal tracking control problem for discrete-time nonlinear systems. Note that asymmetric control restraints and the discount factor are considered. First, related properties are discussed for novel VI, involving the monotonicity of the iterative cost function sequence and the admissibility of the iterative tracking control policy. Second, the stability condition for the discount factor is provided to ensure the stability of all iterative tracking control policies, which are created by stabilizing VI. Third, by combining novel VI and stabilizing VI, an improved VI algorithm is developed, where iterative cost function sequences are monotonically nondecreasing and nonincreasing during novel VI and stabilizing VI stages, respectively. Fourth, with the appropriate discount factor, an online VI algorithm is proposed by integrating the attraction domain with improved VI. Also, under the online VI structure, the asymptotic stability proof is performed for the tracking error system. Finally, regarding theoretical contributions are illustrated by a simulation case.
Ding Wang 0001, Peng Xin, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2022 Neural critic learning for tracking control design of constrained nonlinear multi-person zero-sum games
Menghua Li, Ding Wang 0001, Junfei Qiao 0001
Neurocomputing2
2022 Novel optimal trajectory tracking for nonlinear affine systems with an advanced critic learning structure
Ding Wang 0001, Huiling Zhao
Neural Networks1
2022 Offline and Online Adaptive Critic Control Designs With Stability Guarantee Through Value Iteration
abstract
This article is concerned with the stability of the closed-loop system using various control policies generated by value iteration. Some stability properties involving admissibility criteria, the attraction domain, and so forth, are investigated. An offline integrated value iteration (VI) scheme with a stability guarantee is developed by combining the advantages of VI and policy iteration, which is convenient to obtain admissible control policies. Also, based on the concept of attraction domain, an online adaptive dynamic programming algorithm using immature control policies is developed. Remarkably, it is ensured that the state trajectory under the online algorithm converges to the origin. Particularly, for linear systems, the online ADP algorithm with a general scheme possesses more enhanced stability property. The theoretical results reveal that the stability of the linear system can be guaranteed even if the control policy sequence includes finite unstable elements. The numerical results verify the effectiveness of the present algorithms.
Mingming Ha, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Cybern.2
2022 Self-Learning Robust Control Synthesis and Trajectory Tracking of Uncertain Dynamics
abstract
In this article, we investigate the self-learning robust control synthesis and tracking design of general uncertain dynamical systems. Based on the adaptive critic learning, the robust stabilization method is developed with the help of conducting problem transformation. In addition, by considering the optimal control solution with a discounted cost function, the established method is extended to address the robust trajectory tracking design problem. The Lyapunov stability analysis is also conducted for proving the robustness of the related control plants. Finally, the simulation verification with the three case studies is provided in terms of robust stabilization and trajectory tracking, respectively.
Ding Wang 0001, Long Cheng 0001, Jun Yan 0007
IEEE Trans. Cybern.1
2022 An Approximate Neuro-Optimal Solution of Discounted Guaranteed Cost Control Design
abstract
The adaptive optimal feedback stabilization is investigated in this article for discounted guaranteed cost control of uncertain nonlinear dynamical systems. Via theoretical analysis, the guaranteed cost control problem involving a discounted utility is transformed to the design of a discounted optimal control policy for the nominal plant. The size of the neighborhood with respect to uniformly ultimately bounded stability is discussed. Then, for deriving the approximate optimal solution of the modified Hamilton-Jacobi-Bellman equation, an improved self-learning algorithm under the framework of adaptive critic designs is established. It facilitates the neuro-optimal control implementation without an additional requirement of the initial admissible condition. The simulation verification toward several dynamics is provided, involving the F16 aircraft plant, in order to illustrate the effectiveness of the discounted guaranteed cost control method.
Ding Wang 0001, Junfei Qiao 0001, Long Cheng 0001
IEEE Trans. Cybern.1
2022 Dynamic Transfer Reference Point-Oriented MOEA/D Involving Local Objective-Space Knowledge
abstract
The decomposition-based evolutionary algorithm (MOEA/D) has attained excellent performance in solving optimization problems involving multiple conflicting objectives. However, the Pareto-optimal front (POF) of many multiobjective optimization problems (MOPs) has irregular properties, which weakens the performance of MOEA/D. To address this issue, we devise a dynamic transfer reference point-oriented MOEA/D with local objective-space knowledge (DTR-MOEA/D). The design principle is based on three original and rigorous mechanisms. First, the individuals are projected onto a line segment (two-objective case) or a 3-D plane (three-objective case) after being normalized in the objective space. The line segment or the plane is divided into three different regions: 1) the central region; 2) the middle region; and 3) the edge region. Second, a dynamic transfer criterion of the reference point is developed based on the population density relationships in different regions. Third, a strategy of population diversity enhancement guided by local objective-space knowledge is adopted to improve the diversity of the population. Finally, the experimental results conducted on 16 benchmark MOPs and eight modified MOPs with irregular POF shapes verify that the proposed DTR-MOEA/D has attained a strong competitiveness compared with other representative algorithms.
Yingbo Xie, Shengxiang Yang, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans. Evol. Comput.3
2022 Policy Gradient Adaptive Critic Design With Dynamic Prioritized Experience Replay for Wastewater Treatment Process Control
abstract
With the industrialization of modern society, the pollution of water resources becomes more and more serious. Although purifying urban sewage through the wastewater treatment plants eases the burden of fragile ecosystems, the nonlinearities and uncertainties of biochemical reactions are difficult to address. In this article, a dynamic prioritized policy gradient adaptive dynamic programming (ADP) method is developed to solve the optimal control problem of nonaffine nonlinear discrete-time systems, along with convergence analysis of the algorithm. To the best of our knowledge, it is indispensable to conduct system modeling during the previous ADP research on wastewater treatment process control. By introducing the dynamic prioritized replay buffer and neural networks, the proposed ADP controller can track the setpoints of the wastewater treatment plant and alleviate the effects of disturbance without system modeling. The test results verify that the devised control method outperforms the proportional-integral-derivative strategy with less oscillation when unknown interference occurred.
Ruyue Yang, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans. Ind. Informatics2
2022 Adaptive Optimal Control for Unknown Constrained Nonlinear Systems With a Novel Quasi-Model Network
abstract
A policy-iteration-based algorithm is presented in this article for optimal control of unknown continuous-time nonlinear systems subject to bounded inputs by utilizing the adaptive dynamic programming (ADP). Three neural networks (NNs), called critic network, actor network, and quasi-model network, are utilized in the proposed algorithm to give approximations of the control law, the cost function, and the function constituted by partial derivatives of value functions with respect to states and unknown input gain dynamics, respectively. At each iteration, based on the least sum of squares method, the parameters of critic and quasi-model networks will be tuned simultaneously, which eliminates the necessity of separately learning the system model in advance. Then, the control law is improved by satisfying the necessary optimality condition. Then, the proposed algorithm's optimality and convergence properties are exhibited. Finally, the simulation results demonstrate the availability of the proposed algorithm.
Xiumei Han, Xudong Zhao 0001, Hamid Reza Karimi, Ding Wang 0001, Guangdeng Zong
IEEE Trans. Neural Networks Learn. Syst.4
2022 Time-/Event-Triggered Adaptive Neural Asymptotic Tracking Control for Nonlinear Systems With Full-State Constraints and Application to a Single-Link Robot
abstract
This study proposes the time-/event-triggered adaptive neural control strategies for the asymptotic tracking problem of a class of uncertain nonlinear systems with full-state constraints. First, we design a time-triggered strategy. The effect caused by the residuals of the estimation via radial basis function (RBF) neural networks (NNs), and the reasonable upper bounds on the first derivative of the reference signal and the derivative of each virtual control, can be eliminated by designing appropriate adaptive laws and utilizing the basic properties of RBF NNs. Moreover, the construction of the barrier Lyapunov functions (BLFs) in this work ensures the compliance of the full-state constraints and also holds the asymptotic output tracking performance. Then, based on the time-triggered strategy, we further design a relative threshold event-triggered strategy. The proposed event-triggered adaptive neural controller can solve the main control objective of this work, that is: 1) the full-state constraint requirements of the system are not violated and 2) the output signal asymptotically tracks the reference signal. Compared with the traditional method, the event-triggered strategy can improve the utilization of communication channels and resources and has greater practical significance. Finally, an example of single-link robot under the proposed two strategies illustrates the validity of the constructed controllers.
Jiaming Zhang 0003, Ben Niu 0003, Ding Wang 0001, Huanqing Wang 0001, Ping Zhao 0002, Guangdeng Zong
IEEE Trans. Neural Networks Learn. Syst.3
2021 Adaptive-critic-based hybrid intelligent optimal tracking for a class of nonlinear discrete-time systems
Ding Wang 0001, Mingming Ha, Lingzhi Hu
Eng. Appl. Artif. Intell.1
2021 Adaptive neural tracking control of high-order nonlinear systems with quantized input
Huanqing Wang 0001, Ding Wang 0001, Ben Niu 0003, Ming Chen 0020
Neurocomputing3
2021 A novel decomposition-based multiobjective evolutionary algorithm using improved multiple adaptive dynamic selection strategies
Yingbo Xie, Junfei Qiao 0001, Ding Wang 0001
Inf. Sci.3
2021 Neural-network-based discounted optimal control via an integrated value iteration with accuracy guarantee
Mingming Ha, Ding Wang 0001, Derong Liu 0001
Neural Networks2
2021 Neural optimal tracking control of constrained nonaffine systems with a wastewater treatment application
Ding Wang 0001, Mingming Ha
Neural Networks1
2020 Asymptotically stable critic designs for approximate optimal stabilization of nonlinear systems subject to mismatched external disturbances
Bo Zhao 0015, Ding Wang 0001
Neurocomputing3
2020 Event-triggered constrained control with DHP implementation for nonaffine discrete-time systems
Mingming Ha, Ding Wang 0001, Derong Liu 0001
Inf. Sci.2
2020 Improved value iteration for neural-network-based stochastic optimal control design
Mingming Liang, Ding Wang 0001, Derong Liu 0001
Neural Networks2
2020 Intelligent Critic Control With Robustness Guarantee of Disturbed Nonlinear Plants
abstract
In this paper, the author focuses on establishing an intelligent critic control framework with robustness guarantee for disturbed nonlinear systems. Combining the neural network learning ability with adaptive critic designs, a general structure of intelligent critic control is developed to address the robustness problems, which broadens the application scope of adaptive dynamic programming and the related learning control methods. First, the problem transformation is conducted for changing the robust stabilization problem into optimal control design, where a special discounted cost function is well defined. Then, a recurrent neural network is constructed to learn the unknown nominal plant with stability proof. Moreover, the critic network implementation is presented with the help of the obtained neural identifier and the adaptive learning architecture. In addition, extension discussions and several simulation examples are provided to display the robustness verification results of the intelligent critic strategy.
Ding Wang 0001
IEEE Trans. Cybern.1
2020 Robust Policy Learning Control of Nonlinear Plants With Case Studies for a Power System Application
abstract
In view of the prevalence of dynamic uncertainties, we study the robust policy learning control of nonlinear plants in this paper. The auxiliary system and policy learning techniques are integrated to accomplish robust stabilization of mismatched nonlinear systems. First, the uncertain dynamics is handled by proper transformation, so as to construct an optimal regulation problem with respect to an augmented auxiliary system. Then, the integral policy iteration algorithm is employed for optimal control design without requiring system dynamics. The equivalence results involved in problem transformation and algorithm improvement are analyzed. After that, the actor-critic structure is adopted with least squares implementation for approximate calculation. Finally, the experimental simulation with an application to a power system is provided, which demonstrates the validity of the adaptive robust control strategy. The present policy learning algorithm does not rely on whole information of system dynamics and the established robust control technique is applicable for nonlinear plants subjected to mismatched uncertainties.
Ding Wang 0001
IEEE Trans. Ind. Informatics1
2020 Adaptive Neural Output-Feedback Controller Design of Switched Nonlower Triangular Nonlinear Systems With Time Delays
abstract
In this article, we study the issue of adaptive neural output-feedback controller design for a class of uncertain switched time-delay nonlinear systems with nonlower triangular structure. The prominent contribution of this article is that the delay-dependent stability criterion of nonswitched nonlinear systems is successfully extended to that of switched nonlower triangular nonlinear systems. The design algorithm is listed as follows. First, a switched state observer is designed such that the error dynamic system can be generated. Second, neural networks, adaptive backstepping technique, and variable separation method are, respectively, applied to construct a common controller for all subsystems, in which the Lyapunov-Krasovskii functionals are deliberately constructed such that the average dwell-time scheme can be employed to guarantee the stability and performance of the closed-loop system, despite the existence of time delays. Third, the stability analysis process confirms in detail that all the variables of the closed-loop system are semiglobally uniformly ultimately bounded. Finally, simulation study is given to show the validity of the proposed control approach.
Ben Niu 0003, Ding Wang 0001, Ming Liu 0014, Xinmin Song, Huanqing Wang 0001, Peiyong Duan
IEEE Trans. Neural Networks Learn. Syst.2
2020 Event-Triggered Adaptive Critic Control Design for Discrete-Time Constrained Nonlinear Systems
abstract
In this paper, through event-triggered approach, the constrained near-optimal control problem for a class of nonlinear discrete-time systems is investigated and solved by heuristic dynamic programming (HDP) technique. The proposed method can reduce the amount of computation remarkably without deteriorating the system stability. In order to overcome the control constraints and reduce the computational burden, a nonquadratic performance index is introduced. Then, stability analysis of the event-triggered system with control constraints and an event-triggered constrained controller design algorithm are given. Three neural networks are used in the HDP scheme, which are designed to identify the unknown nonlinear system, approximate value function, and control law, respectively. In the model neural network, an effective method is developed to initialize its weights. Finally, two examples are included to demonstrate the present method.
Mingming Ha, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Model-Free H∞ Optimal Tracking Control of Constrained Nonlinear Systems via an Iterative Adaptive Learning Algorithm
abstract
In this paper, an H∞optimal tracking controller for completely unknown discrete-time nonlinear systems with control constraints is obtained by using an iterative adaptive learning algorithm. An augmented system is established by integrating the tracking error system and the reference trajectory. As an identifier of the unknown systems, a neural network (NN) is introduced with asymptotic stability of the estimation error. An action-disturbance-critic NN structure is proposed to implement the iterative dual heuristic programming algorithm with convergence guarantee of the costate function and the control policy. Simulation results and comparisons are provided to illustrate the superior performance of the designed optimal tracking controller.
Jiaxu Hou, Ding Wang 0001, Derong Liu 0001, Yun Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Neuro-Optimal Control for Discrete Stochastic Processes via a Novel Policy Iteration Algorithm
abstract
In this paper, a novel policy iteration adaptive dynamic programming (ADP) algorithm is presented which is called “local policy iteration ADP algorithm” to obtain the optimal control for discrete stochastic processes. In the proposed local policy iteration ADP algorithm, the iterative decision rules are updated in a local space of the whole state space. Hence, we can significantly reduce the computational burden for the CPU in comparison with the conventional policy iteration algorithm. By analyzing the convergence properties of the proposed algorithm, it is shown that the iterative value functions are monotonically nonincreasing. Besides, the iterative value functions can converge to the optimum in a local policy space. In addition, this local policy space will be described in detail for the first time. Under a few weak constraints, it is also shown that the iterative value function will converge to the optimal performance index function of the global policy space. Finally, a simulation example is presented to validate the effectiveness of the developed method.
Mingming Liang, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Approximate neural optimal control with reinforcement learning for a torsional pendulum device
Ding Wang 0001, Junfei Qiao 0001
Neural Networks1
2019 Adaptive Neural State-Feedback Tracking Control of Stochastic Nonlinear Switched Systems: An Average Dwell-Time Method
abstract
In this paper, the problem of adaptive neural state-feedback tracking control is considered for a class of stochastic nonstrict-feedback nonlinear switched systems with completely unknown nonlinearities. In the design procedure, the universal approximation capability of radial basis function neural networks is used for identifying the unknown compounded nonlinear functions, and a variable separation technique is employed to overcome the design difficulty caused by the nonstrict-feedback structure. The most outstanding novelty of this paper is that individual Lyapunov function of each subsystem is constructed by flexibly adopting the upper and lower bounds of the control gain functions of each subsystem. Furthermore, by combining the average dwell-time scheme and the adaptive backstepping design, a valid adaptive neural state-feedback controller design algorithm is presented such that all the signals of the switched closed-loop system are in probability semiglobally uniformly ultimately bounded, and the tracking error eventually converges to a small neighborhood of the origin in probability. Finally, the availability of the developed control scheme is verified by two simulation examples.
Ben Niu 0003, Ding Wang 0001, Naif D. Alotaibi, Fuad E. Alsaadi
IEEE Trans. Neural Networks Learn. Syst.2
2019 A Novel Neural-Network-Based Adaptive Control Scheme for Output-Constrained Stochastic Switched Nonlinear Systems
abstract
In this paper, a novel neural-network (NN)-based adaptive tracking controller design method is presented for the single-input/single-output nonlinear stochastic switched systems in lower triangular structures with an output constraint. First, a well-designed nonlinear mapping is introduced to transform the switched stochastic system to a new system without constraints, which implies the controller design of the transformed system is equivalent to that of the stochastic switched system. Then radial basis function NNs are applied to model the unknown nonlinearities and the adaptive backstepping technique is employed to construct two classes of adaptive neural controllers under different adaptive laws. It is proved that both controllers can assure all the signals in the closed-loop remain bounded in probability, and the tracking error finally converges to a neighborhood of the origin without violating the constraint. Furthermore, the use of the nonlinear mapping to deal with the asymmetric output constraint is also studied as a generalization result. Two illustrative examples with numerical data and simulation results are given to show the validity and performance of the proposed control schemes.
Ben Niu 0003, Ding Wang 0001, Xue-Jun Xie, Naif D. Alotaibi, Fuad E. Alsaadi
IEEE Trans. Syst. Man Cybern. Syst.2
2018 Local Tracking Control for Unknown Interconnected Systems via Neuro-Dynamic Programming
Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Ding Wang 0001, Yancai Xu, Qinglai Wei
ICONIP (7)4
2018 Connectivity preserved nonlinear time-delayed multiagent systems using neural networks and event-based mechanism
Hongwen Ma, Ding Wang 0001
Neural Comput. Appl.2
2018 Neural robust stabilization via event-triggering mechanism and adaptive learning technique
Ding Wang 0001, Derong Liu 0001
Neural Networks1
2018 Neural network robust tracking control with adaptive critic framework for uncertain nonlinear systems
Ding Wang 0001, Derong Liu 0001, Yun Zhang 0001, Hongyi Li 0001
Neural Networks1
2018 Distributed algorithm for dissensus of a class of networked multiagent systems using output information
Hongwen Ma, Derong Liu 0001, Ding Wang 0001, Xiong Yang 0001, Hongliang Li 0002
Soft Comput.3
2018 Decentralized adaptive optimal stabilization of nonlinear systems with matched interconnections
Chaoxu Mu, Changyin Sun 0001, Ding Wang 0001, Aiguo Song, Chengshan Qian
Soft Comput.3
2018 Data-Driven Finite-Horizon Approximate Optimal Control for Discrete-Time Nonlinear Systems Using Iterative HDP Approach
abstract
This paper presents a data-based finite-horizon optimal control approach for discrete-time nonlinear affine systems. The iterative adaptive dynamic programming (ADP) is used to approximately solve Hamilton-Jacobi-Bellman equation by minimizing the cost function in finite time. The idea is implemented with the heuristic dynamic programming (HDP) involved the model network, which makes the iterative control at the first step can be obtained without the system function, meanwhile the action network is used to obtain the approximate optimal control law and the critic network is utilized for approximating the optimal cost function. The convergence of the iterative ADP algorithm and the stability of the weight estimation errors based on the HDP structure are intensively analyzed. Finally, two simulation examples are provided to demonstrate the theoretical results and show the performance of the proposed method.
Chaoxu Mu, Ding Wang 0001, Haibo He
IEEE Trans. Cybern.2
2018 Model-Free Adaptive Control for Unknown Nonlinear Zero-Sum Differential Game
abstract
In this paper, we present a new model-free globalized dual heuristic dynamic programming (GDHP) approach for the discrete-time nonlinear zero-sum game problems. First, the online learning algorithm is proposed based on the GDHP method to solve the Hamilton-Jacobi-Isaacs equation associated with optimal regulation control problem. By setting backward one step of the definition of performance index, the requirement of system dynamics, or an identifier is relaxed in the proposed method. Then, three neural networks are established to approximate the optimal saddle point feedback control law, the disturbance law, and the performance index, respectively. The explicit updating rules for these three neural networks are provided based on the data generated during the online learning along the system trajectories. The stability analysis in terms of the neural network approximation errors is discussed based on the Lyapunov approach. Finally, two simulation examples are provided to show the effectiveness of the proposed method.
Xiangnan Zhong, Haibo He, Ding Wang 0001, Zhen Ni
IEEE Trans. Cybern.3
2018 Intelligent Optimal Control With Critic Learning for a Nonlinear Overhead Crane System
abstract
In this paper, for achieving the discounted optimal feedback stabilization of a nonlinear overhead crane system, we establish an intelligent control strategy to obtain the solution of the corresponding Hamilton-Jacobi-Bellman equation. Specifically, neural networks are employed to serve as a necessary component to the control system, which exhibits strong online learning ability. A novel updating rule compared to the traditional adaptive critic algorithms is developed, which eliminates the requirement of the initial stabilizing controller and brings in unique advantages to the adaptive critic control design. Stability analysis of the closed-loop system based on the well-known Lyapunov approach and experimental simulation considering the nonlinear overhead dynamics with different case studies are performed to verify the effectiveness of the present control method both in theory and applications.
Ding Wang 0001, Haibo He, Derong Liu 0001
IEEE Trans. Ind. Informatics1
2018 Manifold Regularized Reinforcement Learning
abstract
This paper introduces a novel manifold regularized reinforcement learning scheme for continuous Markov decision processes. Smooth feature representations for value function approximation can be automatically learned using the unsupervised manifold regularization method. The learned features are data-driven, and can be adapted to the geometry of the state space. Furthermore, the scheme provides a direct basis representation extension for novel samples during policy learning and control. The performance of the proposed scheme is evaluated on two benchmark control tasks, i.e., the inverted pendulum and the energy storage problem. Simulation results illustrate the concepts of the proposed scheme and show that it can obtain excellent performance.
Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Learning and Guaranteed Cost Control With Event-Based Adaptive Critic Implementation
abstract
This paper focuses on the event-triggered guaranteed cost control design of nonlinear systems via a self-learning technique. In brief, an event-based guaranteed cost control strategy of nonlinear systems subjects to matched uncertainties is developed, thereby balancing the performance of guaranteed cost and the actuality of limited communication resource. The original control design is transformed into an optimal control problem with an event-based mechanism, where the relationship of guaranteed cost performance compared to the time-based formulation is discussed. A critic neural network is constructed for implementing the event-based optimal control design with stability guarantee. Simulation experiments are carried out to verify the theoretical results in detail.
Ding Wang 0001, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Adaptive Neural Output-Feedback Control for a Class of Nonlower Triangular Nonlinear Systems With Unmodeled Dynamics
abstract
This paper presents the development of an adaptive neural controller for a class of nonlinear systems with unmodeled dynamics and immeasurable states. An observer is designed to estimate system states. The structure consistency of virtual control signals and the variable partition technique are combined to overcome the difficulties appearing in a nonlower triangular form. An adaptive neural output-feedback controller is developed based on the backstepping technique and the universal approximation property of the radial basis function (RBF) neural networks. By using the Lyapunov stability analysis, the semiglobally and uniformly ultimate boundedness of all signals within the closed-loop system is guaranteed. The simulation results show that the controlled system converges quickly, and all the signals are bounded. This paper is novel at least in the two aspects: 1) an output-feedback control strategy is developed for a class of nonlower triangular nonlinear systems with unmodeled dynamics and 2) the nonlinear disturbances and their bounds are the functions of all states, which is in a more general form than existing results.
Huanqing Wang 0001, Peter Xiaoping Liu, Shuai Li 0002, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Neural Network Learning and Robust Stabilization of Nonlinear Systems With Dynamic Uncertainties
abstract
Due to the existence of dynamical uncertainties, it is important to pay attention to the robustness of nonlinear control systems, especially when designing adaptive critic control strategies. In this paper, based on the neural network learning component, the robust stabilization scheme of nonlinear systems with general uncertainties is developed. Through system transformation and employing adaptive critic technique, the approximate optimal controller of the nominal plant can be applied to accomplish robust stabilization for the original uncertain dynamics. The neural network weight vector is very convenient to initialize by virtue of the improved critic learning formulation. Under the action of the approximate optimal control law, the stability issues for the closed-loop form of nominal and uncertain plants are analyzed, respectively. Simulation illustrations via a typical nonlinear system and a practical power system are included to verify the control performance.
Ding Wang 0001, Derong Liu 0001, Chaoxu Mu, Yun Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 On Mixed Data and Event Driven Design for Adaptive-Critic-Based Nonlinear H∞ Control
abstract
In this paper, based on the adaptive critic learning technique, the control for a class of unknown nonlinear dynamic systems is investigated by adopting a mixed data and event driven design approach. The nonlinear control problem is formulated as a two-player zero-sum differential game and the adaptive critic method is employed to cope with the data-based optimization. The novelty lies in that the data driven learning identifier is combined with the event driven design formulation, in order to develop the adaptive critic controller, thereby accomplishing the nonlinear control. The event driven optimal control law and the time driven worst case disturbance law are approximated by constructing and tuning a critic neural network. Applying the event driven feedback control, the closed-loop system is built with stability analysis. Simulation studies are conducted to verify the theoretical results and illustrate the control performance. It is significant to observe that the present research provides a new avenue of integrating data-based control and event-triggering mechanism into establishing advanced adaptive critic systems.
Ding Wang 0001, Chaoxu Mu, Derong Liu 0001, Hongwen Ma
IEEE Trans. Neural Networks Learn. Syst.1
2018 Event-Based Robust Control for Uncertain Nonlinear Systems Using Adaptive Dynamic Programming
abstract
In this paper, the robust control problem for a class of continuous-time nonlinear system with unmatched uncertainties is investigated using an event-based control method. First, the robust control problem is transformed into a corresponding optimal control problem with an augmented control and an appropriate cost function. Under the event-based mechanism, we prove that the solution of the optimal control problem can asymptotically stabilize the uncertain system with an adaptive triggering condition. That is, the designed event-based controller is robust to the original uncertain system. Note that the event-based controller is updated only when the triggering condition is satisfied, which can save the communication resources between the plant and the controller. Then, a single network adaptive dynamic programming structure with experience replay technique is constructed to approach the optimal control policies. The stability of the closed-loop system with the event-based control policy and the augmented control policy is analyzed using the Lyapunov approach. Furthermore, we prove that the minimal intersample time is bounded by a nonzero positive constant, which excludes Zeno behavior during the learning process. Finally, two simulation examples are provided to demonstrate the effectiveness of the proposed control scheme.
Dongbin Zhao, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Data-Based Optimal Control for Weakly Coupled Nonlinear Systems Using Policy Iteration
abstract
In this paper, a data-based online learning algorithm is established to solve the optimal control problem for weakly coupled continuous-time nonlinear systems with completely unknown dynamics. Using the weak coupling theory, we reformulate the original problem into three reduced-order optimal control problems. We establish an online model-free integral policy iteration algorithm to solve the decoupled optimal control problems without system dynamics. To implement the data-based online learning algorithm, the actor-critic technique based on neural networks and the least squares method are used. Two simulation examples are given to verify the effectiveness of the developed algorithm.
Chao Li 0024, Derong Liu 0001, Ding Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2018 Decentralized Control for Large-Scale Nonlinear Systems With Unknown Mismatched Interconnections via Policy Iteration
abstract
In this paper, the decentralized control problem is solved based on a policy iteration algorithm for large-scale nonlinear systems with unknown mismatched interconnections. The unknown interconnection is approximated by a neural network with local states of isolated subsystem and substituted reference states of coupled subsystems. Then, an adaptive estimation term is utilized to construct the improved local performance index function that reflects the substitution error. Hereafter, the closed-loop large-scale nonlinear system is guaranteed to be ultimately uniformly bounded by the implementation of a set of developed decentralized optimal control policies. Two simulation examples are given to verify the effectiveness of the presented scheme. The significant contribution of this scheme lies in that it removes the common assumptions on satisfying matching condition and upper boundedness of interconnections, when designing the decentralized optimal control for large-scale nonlinear systems.
Bo Zhao 0015, Ding Wang 0001, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Robust Control of Uncertain Nonlinear Systems Based on Adaptive Dynamic Programming
Jing Na, Jun Zhao 0015, Guanbin Gao, Ding Wang 0001
ICONIP (3)4
2017 Adaptation-Oriented Near-Optimal Control and Robust Synthesis of an Overhead Crane System
Ding Wang 0001
ICONIP (6)1
2017 Neural adaptive control of microgrid frequency regulation with wind power
abstract
Due to the uncertainty of power load demand and the stochastic power generation from renewable energy, frequency fluctuation becomes a major concern of power system, especially for a microgrid. In this paper, an improved proportional-integral (PI) controller based on the neural adaptive control method is proposed to deal with the load frequency control (LFC) problem in a microgrid with wind power. The designed neural adaptive control auxiliary controller is used to provide the adaptive supplement control signal to PI controller in a real-time manner. Simulation studies on a benchmark microgrid system are carried out between the proposed compound controller and traditional PI controller. The simulation results demonstrate the proposed method has a superior performance for stabilizing the frequency over the traditional PI control under disturbances from the load change and wind power.
Weiqiang Liu 0008, Chaoxu Mu, Ding Wang 0001, Chao Ren 0003
IECON3
2017 Developing nonlinear adaptive optimal regulators through an improved neural learning mechanism
Ding Wang 0001, Chaoxu Mu
Sci. China Inf. Sci.1
2017 Bounded robust control design for uncertain nonlinear systems using single-network adaptive dynamic programming
Yuzhu Huang, Ding Wang 0001, Derong Liu 0001
Neurocomputing2
2017 Adaptive tracking control for a class of continuous-time uncertain nonlinear systems using the approximate solution of HJB equation
Chaoxu Mu, Changyin Sun 0001, Ding Wang 0001, Aiguo Song
Neurocomputing3
2017 Neural-network-based adaptive guaranteed cost control of nonlinear dynamical systems with matched uncertainties
Chaoxu Mu, Ding Wang 0001
Neurocomputing2
2017 A novel neural optimal control framework with nonlinear dynamics: Closed-loop stability and simulation verification
Ding Wang 0001, Chaoxu Mu
Neurocomputing1
2017 Policy Gradient Adaptive Dynamic Programming for Data-Based Optimal Control
abstract
The model-free optimal control problem of general discrete-time nonlinear systems is considered in this paper, and a data-based policy gradient adaptive dynamic programming (PGADP) algorithm is developed to design an adaptive optimal controller method. By using offline and online data rather than the mathematical system model, the PGADP algorithm improves control policy with a gradient descent scheme. The convergence of the PGADP algorithm is proved by demonstrating that the constructed Q -function sequence converges to the optimal Q -function. Based on the PGADP algorithm, the adaptive control method is developed with an actor-critic structure and the method of weighted residuals. Its convergence properties are analyzed, where the approximate Q -function converges to its optimum. Computer simulation results demonstrate the effectiveness of the PGADP-based adaptive control method.
Biao Luo 0001, Derong Liu 0001, Huai-Ning Wu, Ding Wang 0001, Frank L. Lewis
IEEE Trans. Cybern.4
2017 Improving the Critic Learning for Event-Based Nonlinear H∞ Control Design
abstract
control problem is regarded as a two-player zero-sum game and the adaptive critic mechanism is used to achieve the minimax optimization under event-based environment. Then, based on an improved updating rule, the event-based optimal control law and the time-based worst-case disturbance law are obtained approximately by training a single critic neural network. The initial stabilizing control is no longer required during the implementation process of the new algorithm. Next, the closed-loop system is formulated as an impulsive model and its stability issue is handled by incorporating the improved learning criterion. The infamous Zeno behavior of the present event-based design is also avoided through theoretical analysis on the lower bound of the minimal intersample time. Finally, the applications to an aircraft dynamics and a robot arm plant are carried out to verify the efficient performance of the present novel design method.
Ding Wang 0001, Haibo He, Derong Liu 0001
IEEE Trans. Cybern.1
2017 Adaptive Critic Nonlinear Robust Control: A Survey
abstract
Adaptive dynamic programming (ADP) and reinforcement learning are quite relevant to each other when performing intelligent optimization. They are both regarded as promising methods involving important components of evaluation and improvement, at the background of information technology, such as artificial intelligence, big data, and deep learning. Although great progresses have been achieved and surveyed when addressing nonlinear optimal control problems, the research on robustness of ADP-based control strategies under uncertain environment has not been fully summarized. Hence, this survey reviews the recent main results of adaptive-critic-based robust control design of continuous-time nonlinear systems. The ADP-based nonlinear optimal regulation is reviewed, followed by robust stabilization of nonlinear systems with matched uncertainties, guaranteed cost control design of unmatched plants, and decentralized stabilization of interconnected systems. Additionally, further comprehensive discussions are presented, including event-based robust control design, improvement of the critic learning rule, nonlinear H∞control design, and several notes on future perspectives. By applying the ADP-based optimal and robust control methods to a practical power system and an overhead crane plant, two typical examples are provided to verify the effectiveness of theoretical results. Overall, this survey is beneficial to promote the development of adaptive critic control methods with robustness guarantee and the construction of higher level intelligent systems.
Ding Wang 0001, Haibo He, Derong Liu 0001
IEEE Trans. Cybern.1
2017 Event-Driven Adaptive Robust Control of Nonlinear Systems With Uncertainties Through NDP Strategy
abstract
In this paper, we construct an event-driven adaptive robust control approach for continuous-time uncertain nonlinear systems through a neural dynamic programming (NDP) strategy. Through system transformation and theoretical analysis, the robustness of the original uncertain system can be achieved by designing an event-driven optimal controller with respect to the nominal system under a suitable triggering condition. In addition, it is also observed that the event-driven controller has a certain degree of gain margin. Then, the NDP technique is employed to perform the main controller design task, followed by the uniform ultimate boundedness stability proof with the feedback action of the event-driven adaptive control law. The comparative effect of the present control strategy is also illustrated via two simulation examples. The established method provides a new avenue of combining adaptive dynamic programming-based self-learning control, event-triggered adaptive control, and robust control, to investigate the nonlinear adaptive robust feedback design under uncertain environment.
Ding Wang 0001, Chaoxu Mu, Haibo He, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Event-Based Constrained Robust Control of Affine Systems Incorporating an Adaptive Critic Mechanism
abstract
This paper focuses on establishing an event-based constrained robust control strategy for a class of continuous-time affine nonlinear systems by incorporating the adaptive critic mechanism (ACM). The main objective is to integrate the event-based framework, the constrained optimal control method, and the neural network learning ability, thereby achieving the nonlinear robust state feedback of input-constrained nonlinear systems under event-based environment. Through theoretical analysis, it is shown that the nonlinear robust control law subject to input limitations can be obtained by designing an event-based constrained optimal controller with respect to the nominal system. Then, the ACM is adopted to facilitate the constrained optimal control implementation, where a critic neural network is constructed to serve as the learning approximator. The system stability issue is proved by employing the Lyapunov theory and the constrained robust control performance is illustrated through simulation experiments of several dynamical plants.
Ding Wang 0001, Chaoxu Mu, Xiong Yang 0001, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Error Bound Analysis of Q-Function for Discounted Optimal Control Problems With Policy Iteration
abstract
In this paper, we present error bound analysis of the Q-function for the action-dependent adaptive dynamic programming for solving discounted optimal control problems of unknown discrete-time nonlinear systems. The convergence of Q-functions derived by a policy iteration algorithm under ideal conditions is given. Considering the approximated errors of the Q-function and control policy in the policy evaluation step and policy improvement step, we establish error bounds of approximate Q-functions in each iteration. With the given boundedness conditions, the approximate Q-function will converge to a finite neighborhood of the optimal Q-function. To implement the presented algorithm, two three-layer neural networks are employed to approximate the Q-function and the control policy, respectively. Finally, a simulation example is utilized to verify the validity of the presented algorithm.
Ding Wang 0001, Hongliang Li 0002, Derong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2016 Neural Dynamic Programming for Event-Based Nonlinear Adaptive Robust Stabilization
Ding Wang 0001, Hongwen Ma, Derong Liu 0001, Huidong Wang
ICONIP (1)1
2016 Decentralized Stabilization for Nonlinear Systems with Unknown Mismatched Interconnections
Bo Zhao 0015, Ding Wang 0001, Derong Liu 0001
ICONIP (3)2
2016 Neural-network-based robust optimal control of uncertain nonlinear systems using model-free policy iteration algorithm
abstract
In this paper, we establish a robust optimal control law for a class of continuous-time uncertain nonlinear systems by using a neural-network-based model-free policy iteration approach. The robust control law of the original uncertain nonlinear system is derived by adding a feedback gain to the optimal control law of the nominal system. It is proven that this robust control law can achieve optimality under a specified cost function. Then, the neural-network-based model-free policy iteration algorithm is developed to solve the Hamilton-Jacobi-Bellman equation corresponding to the nominal system without system dynamics. The actor-critic technique and the least squares implementation method are used to obtain the optimal control policy of the nominal system. A numerical simulation is given to verify the applicability of the present robust optimal control scheme.
Chao Li 0024, Ding Wang 0001, Derong Liu 0001
IJCNN2
2016 Distributed control of second-order nonlinear time-delayed multiagent systems with disturbance using neural networks
abstract
In this paper, a class of second-order nonlinear time-delayed multiagent systems with disturbance is investigated. In order to improve the adaptivity, neural networks are used to learn the unknown dynamics. Then, by utilizing Lyapunov-Krasovskii functional, time delays can be eliminated. Moreover, a robustifying term is introduced to constrain external disturbance. With divide-and-conquer idea, the distributed controller is divided into five different parts to make the multiagent systems reach consensus. To circumvent the singularity induced by the time-delay elimination part, a σ-function is developed. Finally, the simulation results demonstrate the validity of the distributed controller.
Hongwen Ma, Derong Liu 0001, Ding Wang 0001
IJCNN3
2016 Decentralized guaranteed cost control of interconnected systems with uncertainties: A learning-based optimal control strategy
Ding Wang 0001, Derong Liu 0001, Chaoxu Mu, Hongwen Ma
Neurocomputing1
2016 Distributed control algorithm for bipartite consensus of the nonlinear time-delayed multi-agent systems with neural networks
Ding Wang 0001, Hongwen Ma, Derong Liu 0001
Neurocomputing1
2016 Event-based input-constrained nonlinear H∞ state feedback with adaptive critic and neural implementation
Ding Wang 0001, Chaoxu Mu, Derong Liu 0001
Neurocomputing1
2016 Data-driven controller design for general MIMO nonlinear systems via virtual reference feedback tuning and neural networks
Derong Liu 0001, Ding Wang 0001, Hongwen Ma
Neurocomputing3
2016 Guaranteed cost neural tracking control for a class of uncertain nonlinear systems using adaptive dynamic programming
Xiong Yang 0001, Derong Liu 0001, Qinglai Wei, Ding Wang 0001
Neurocomputing4
2016 Data-based robust optimal control of continuous-time affine nonlinear systems with matched uncertainties
Ding Wang 0001, Chao Li 0024, Derong Liu 0001, Chaoxu Mu
Inf. Sci.1
2016 A neural-network-based online optimal control approach for nonlinear robust decentralized stabilization
Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Hongwen Ma, Chao Li 0024
Soft Comput.1
2016 Experience Replay for Optimal Control of Nonzero-Sum Game Systems With Unknown Dynamics
abstract
In this paper, an approximate online equilibrium solution is developed for an N -player nonzero-sum (NZS) game systems with completely unknown dynamics. First, a model identifier based on a three-layer neural network (NN) is established to reconstruct the unknown NZS games systems. Moreover, the identifier weight vector is updated based on experience replay technique which can relax the traditional persistence of excitation condition to a simplified condition on recorded data. Then, the single-network adaptive dynamic programming (ADP) with experience replay algorithm is proposed for each player to solve the coupled nonlinear Hamilton- (HJ) equations, where only the critic NN weight vectors are required to tune for each player. The feedback Nash equilibrium is provided by the solution of the coupled HJ equations. Based on the experience replay technique, a novel critic NN weights tuning law is proposed to guarantee the stability of the closed-loop system and the convergence of the value functions. Furthermore, a Lyapunov-based stability analysis shows that the uniform ultimate boundedness of the closed-loop system is achieved. Finally, two simulation examples are given to verify the effectiveness of the proposed control scheme.
Dongbin Zhao, Ding Wang 0001, Yuanheng Zhu
IEEE Trans. Cybern.3
2016 Editorial IEEE Transactions on Neural Networks and Learning Systems 2016 and Beyond
abstract
“Happy New Year!” At the beginning of 2016, I would like to take this opportunity to wish everyone a very happy, healthy, and prosperous new year! It is my great honor and privilege to serve as the Editor-in-Chief (EiC) of the IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS (TNNLS), and I am excited to write this Editorial to start a new journey with you all.
Haibo He, Nitesh V. Chawla, Yoonsuck Choe, Andries P. Engelbrecht, Jaya deva, Lyle N. Long, Ali A. Minai, Feiping Nie 0001, Umut Ozertem, Barak A. Pearlmutter, Ling Shao 0001, Jennie Si, Jochen J. Steil, Brijesh K. Verma, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.16
2016 Model-Free Optimal Tracking Control via Critic-Only Q-Learning
abstract
Model-free control is an important and promising topic in control fields, which has attracted extensive attention in the past few years. In this paper, we aim to solve the model-free optimal tracking control problem of nonaffine nonlinear discrete-time systems. A critic-only Q-learning (CoQL) method is developed, which learns the optimal tracking control from real system data, and thus avoids solving the tracking Hamilton-Jacobi-Bellman equation. First, the Q-learning algorithm is proposed based on the augmented system, and its convergence is established. Using only one neural network for approximating the Q-function, the CoQL method is developed to implement the Q-learning algorithm. Furthermore, the convergence of the CoQL method is proved with the consideration of neural network approximation error. With the convergent Q-function obtained from the CoQL method, the adaptive optimal tracking control is designed based on the gradient descent scheme. Finally, the effectiveness of the developed CoQL method is demonstrated through simulation studies. The developed CoQL method learns with off-policy data and implements with a critic-only structure, thus it is easy to realize and overcome the inadequate exploration problem.
Biao Luo 0001, Derong Liu 0001, Tingwen Huang, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2016 Neural-Network-Based Distributed Adaptive Robust Control for a Class of Nonlinear Multiagent Systems With Time Delays and External Noises
abstract
A class of nonlinear multiagent systems with time delays and external noises is investigated, and a distributed adaptive robust control protocol is developed. It is the first time for a class of multiagent systems to take both time delays and external noises into consideration. By virtue of Lyapunov-Krasovskii functional and Young's inequality, the effects of time delay can be eliminated. Then, to exclude external noises, a robustifying term is introduced to eliminate the negative effects of these noises. Moreover, neural networks are utilized to learn the unknown nonlinear terms to adapt to the complex external environment. Finally, a numerical simulation is conducted to validate the effectiveness of our distributed control protocol.
Hongwen Ma, Zhuo Wang 0003, Ding Wang 0001, Derong Liu 0001, Qinglai Wei
IEEE Trans. Syst. Man Cybern. Syst.3
2016 An Approximate Optimal Control Approach for Robust Stabilization of a Class of Discrete-Time Nonlinear Systems With Uncertainties
abstract
In this correspondence paper, the robust stabilization of a class of discrete-time nonlinear systems with uncertainties is investigated by using an approximate optimal control approach. The robust control problem is transformed into an optimal control problem under some proper restrictions on the bound of the uncertainties. For the purpose of dealing with the transformed optimal control, the discrete-time generalized Hamilton-Jacobi-Bellman equation is introduced and then solved using the successive approximation method with neural network implementation. In addition, a numerical simulation is included to illustrate the effectiveness of the robust control strategy.
Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Biao Luo 0001, Hongwen Ma
IEEE Trans. Syst. Man Cybern. Syst.1
2016 Data-Based Adaptive Critic Designs for Nonlinear Robust Optimal Control With Uncertain Dynamics
abstract
In this paper, the infinite-horizon robust optimal control problem for a class of continuous-time uncertain nonlinear systems is investigated by using data-based adaptive critic designs. The neural network identification scheme is combined with the traditional adaptive critic technique, in order to design the nonlinear robust optimal control under uncertain environment. First, the robust optimal controller of the original uncertain system with a specified cost function is established by adding a feedback gain to the optimal controller of the nominal system. Then, a neural network identifier is employed to reconstruct the unknown dynamics of the nominal system with stability analysis. Hence, the data-based adaptive critic designs can be developed to solve the Hamilton-Jacobi-Bellman equation corresponding to the transformed optimal control problem. The uniform ultimate boundedness of the closed-loop system is also proved by using the Lyapunov approach. Finally, two simulation examples are presented to illustrate the effectiveness of the developed control strategy.
Ding Wang 0001, Derong Liu 0001, Dongbin Zhao
IEEE Trans. Syst. Man Cybern. Syst.1
2015 Distributed Control for Nonlinear Time-Delayed Multi-Agent Systems with Connectivity Preservation Using Neural Networks
Hongwen Ma, Derong Liu 0001, Ding Wang 0001
ICONIP (3)3
2015 Approximate policy iteration with unsupervised feature learning based on manifold regularization
abstract
In this paper, we develop a novel approximate policy iteration reinforcement learning algorithm with unsupervised feature learning based on manifold regularization. The proposed algorithm can automatically learn data-driven smooth basis representations for value function approximation, which can preserve the intrinsic geometry of the state space of Markov decision processes. Moreover, it can provide a direct basis extension for new samples in both policy learning and policy control processes. We evaluate the effectiveness and efficiency of the proposed algorithm on the inverted pendulum task. Simulation results show that this algorithm can learn smooth basis representations and excellent control policies.
Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001
IJCNN3
2015 Data-driven virtual reference controller design for high-order nonlinear systems via neural network
abstract
This paper is concerned with data-driven methods for virtual reference controller design of high-order nonlinear systems via neural network. Virtual reference feedback tuning (VRFT) is a one-shot direct data-based method to design controller of linear or nonlinear systems. In this paper, we recall the model reference control problem of high-order nonlinear systems and design a new objective function of VRFT. In ideal conditions, the two problems are demonstrated to have the same solution. For the first time, we prove that the value of the optimization problem for model reference control is bounded by that of the objective function of VRFT. A three-layer neural network is employed as a general approximator of the designed controller and two simulations are given to verify the validity of our method.
Derong Liu 0001, Ding Wang 0001
IJCNN3
2015 Neural-network-based decentralized control of continuous-time nonlinear interconnected systems with unknown dynamics
Derong Liu 0001, Chao Li 0024, Hongliang Li 0002, Ding Wang 0001, Hongwen Ma
Neurocomputing4
2015 Centralized and decentralized event-triggered control for group consensus with fixed topology in continuous time
Hongwen Ma, Derong Liu 0001, Ding Wang 0001, Fuxiao Tan, Chao Li 0024
Neurocomputing3
2015 Model-Free Optimal Control for Affine Nonlinear Systems With Convergence Analysis
abstract
In this paper, a self-learning control scheme is proposed for the infinite horizon optimal control of affine nonlinear systems based on the action dependent heuristic dynamic programming algorithm. The policy iteration technique is introduced to derive the optimal control policy with feasibility and convergence analysis. It shows that the “greedy” control action for each state is uniquely existent, the learned control policy after each policy iteration is admissible, and the optimal control policy is able to be obtained. Two three-layer perceptron neural networks are employed to implement the scheme. The critic network is trained by a novel rule to conform to the Bellman equation, and the action network is trained to yield a better control policy. Both training processes alternate until the optimal control policy is achieved. Two simulation examples are provided to validate the effectiveness of the approach. Note to Practitioners - The objective of designing optimal controllers without mathematical models is sought by control practitioners, whereas existing approaches usually derive optimal controllers by accessing the mathematical models or identified models. This paper proposes a new approach which derives optimal controllers by numerical iteration method without accessing any knowledge of the mathematical models. It gives evaluation for every state-action pair in the whole state-action space through the collected data of the underlying system, and then selects the action with the best evaluation for each state. What is required initial admissible control policy. Theorems show that optimal controllers can be acquired and simulation studies verify effectiveness. Further research will extend this approach to online self-learning optimal control approach, thus it can adapt the variation of underlying systems.
Dongbin Zhao, Zhongpu Xia, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.3
2015 Reinforcement-Learning-Based Robust Controller Design for Continuous-Time Uncertain Nonlinear Systems Subject to Input Constraints
abstract
The design of stabilizing controller for uncertain nonlinear systems with control constraints is a challenging problem. The constrained-input coupled with the inability to identify accurately the uncertainties motivates the design of stabilizing controller based on reinforcement-learning (RL) methods. In this paper, a novel RL-based robust adaptive control algorithm is developed for a class of continuous-time uncertain nonlinear systems subject to input constraints. The robust control problem is converted to the constrained optimal control problem with appropriately selecting value functions for the nominal system. Distinct from typical action-critic dual networks employed in RL, only one critic neural network (NN) is constructed to derive the approximate optimal control. Meanwhile, unlike initial stabilizing control often indispensable in RL, there is no special requirement imposed on the initial control. By utilizing Lyapunov's direct method, the closed-loop optimal control system and the estimated weights of the critic NN are proved to be uniformly ultimately bounded. In addition, the derived approximate optimal control is verified to guarantee the uncertain nonlinear system to be stable in the sense of uniform ultimate boundedness. Two simulation examples are provided to illustrate the effectiveness and applicability of the present approach.
Derong Liu 0001, Xiong Yang 0001, Ding Wang 0001, Qinglai Wei
IEEE Trans. Cybern.3
2015 Error Bounds of Adaptive Dynamic Programming Algorithms for Solving Undiscounted Optimal Control Problems
abstract
In this paper, we establish error bounds of adaptive dynamic programming algorithms for solving undiscounted infinite-horizon optimal control problems of discrete-time deterministic nonlinear systems. We consider approximation errors in the update equations of both value function and control policy. We utilize a new assumption instead of the contraction assumption in discounted optimal control problems. We establish the error bounds for approximate value iteration based on a new error condition. Furthermore, we also establish the error bounds for approximate policy iteration and approximate optimistic policy iteration algorithms. It is shown that the iterative approximate value function can converge to a finite neighborhood of the optimal value function under some conditions. To implement the developed algorithms, critic and action neural networks are used to approximate the value function and control policy, respectively. Finally, a simulation example is given to demonstrate the effectiveness of the developed algorithms.
Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2014 Data-driven iterative adaptive dynamic programming algorithm for approximate optimal control of unknown nonlinear systems
abstract
In this paper, we develop a data-driven iterative adaptive dynamic programming algorithm to learn offline the approximate optimal control of unknown discrete-time nonlinear systems. We do not use a model network to identify the unknown system, but utilize the available offline data to learn the approximate optimal control directly. First, the data-driven iterative adaptive dynamic programming algorithm is presented with a convergence analysis. Then, the error bounds for this algorithm are provided considering the approximation errors of function approximation structures. To implement the developed algorithm, two neural networks are used to approximate the state-action value function and the control policy. Finally, two simulation examples are given to demonstrate the effectiveness of the developed algorithm.
Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001, Chao Li 0024
IJCNN3
2014 Full-range adaptive cruise control based on supervised adaptive dynamic programming
Dongbin Zhao, Zhaohui Hu, Zhongpu Xia, Cesare Alippi, Yuanheng Zhu, Ding Wang 0001
Neurocomputing6
2014 Neural-network-based robust optimal control design for a class of uncertain nonlinear systems via adaptive dynamic programming
Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002, Hongwen Ma
Inf. Sci.1
2014 Discrete-time online learning control for a class of unknown nonaffine nonlinear systems using reinforcement learning
Xiong Yang 0001, Derong Liu 0001, Ding Wang 0001, Qinglai Wei
Neural Networks3
2014 Approximate optimal solution of the DTHJB equation for a class of nonlinear affine systems with unknown dead-zone constraints
Dehua Zhang, Derong Liu 0001, Ding Wang 0001
Soft Comput.3
2014 Integral Reinforcement Learning for Linear Continuous-Time Zero-Sum Games With Completely Unknown Dynamics
abstract
In this paper, we develop an integral reinforcement learning algorithm based on policy iteration to learn online the Nash equilibrium solution for a two-player zero-sum differential game with completely unknown linear continuous-time dynamics. This algorithm is a fully model-free method solving the game algebraic Riccati equation forward in time. The developed algorithm updates value function, control and disturbance policies simultaneously. The convergence of the algorithm is demonstrated to be equivalent to Newton's method. To implement this algorithm, one critic network and two action networks are used to approximate the game value function, control and disturbance policies, respectively, and the least squares method is used to estimate the unknown parameters. The effectiveness of the developed scheme is demonstrated in the simulation by designing an H∞state feedback controller for a power system.
Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001
IEEE Trans Autom. Sci. Eng.3
2014 Policy Iteration Algorithm for Online Design of Robust Control for a Class of Continuous-Time Nonlinear Systems
abstract
In this paper, a novel strategy is established to design the robust controller for a class of continuous-time nonlinear systems with uncertainties based on the online policy iteration algorithm. The robust control problem is transformed into the optimal control problem by properly choosing a cost function that reflects the uncertainties, regulation, and control. An online policy iteration algorithm is presented to solve the Hamilton-Jacobi-Bellman (HJB) equation by constructing a critic neural network. The approximate expression of the optimal control policy can be derived directly. The closed-loop system is proved to possess the uniform ultimate boundedness. The equivalence of the neural-network-based HJB solution of the optimal control problem and the solution of the robust control problem is established as well. Two simulation examples are provided to verify the effectiveness of the present robust control scheme.
Ding Wang 0001, Derong Liu 0001, Hongliang Li 0002
IEEE Trans Autom. Sci. Eng.1
2014 Neural-Network-Based Online HJB Solution for Optimal Robust Guaranteed Cost Control of Continuous-Time Uncertain Nonlinear Systems
abstract
In this paper, the infinite horizon optimal robust guaranteed cost control of continuous-time uncertain nonlinear systems is investigated using neural-network-based online solution of Hamilton-Jacobi-Bellman (HJB) equation. By establishing an appropriate bounded function and defining a modified cost function, the optimal robust guaranteed cost control problem is transformed into an optimal control problem. It can be observed that the optimal cost function of the nominal system is nothing but the optimal guaranteed cost of the original uncertain system. A critic neural network is constructed to facilitate the solution of the modified HJB equation corresponding to the nominal system. More importantly, an additional stabilizing term is introduced for helping to verify the stability, which reinforces the updating process of the weight vector and reduces the requirement of an initial stabilizing control. The uniform ultimate boundedness of the closed-loop system is analyzed by using the Lyapunov approach as well. Two simulation examples are provided to verify the effectiveness of the present control approach.
Derong Liu 0001, Ding Wang 0001, Fei-Yue Wang 0001, Hongliang Li 0002, Xiong Yang 0001
IEEE Trans. Cybern.2
2014 Decentralized Stabilization for a Class of Continuous-Time Nonlinear Interconnected Systems Using Online Learning Optimal Control Approach
abstract
In this paper, using a neural-network-based online learning optimal control approach, a novel decentralized control strategy is developed to stabilize a class of continuous-time nonlinear interconnected large-scale systems. First, optimal controllers of the isolated subsystems are designed with cost functions reflecting the bounds of interconnections. Then, it is proven that the decentralized control strategy of the overall system can be established by adding appropriate feedback gains to the optimal control policies of the isolated subsystems. Next, an online policy iteration algorithm is presented to solve the Hamilton-Jacobi-Bellman equations related to the optimal control problem. Through constructing a set of critic neural networks, the cost functions can be obtained approximately, followed by the control policies. Furthermore, the dynamics of the estimation errors of the critic networks are verified to be uniformly and ultimately bounded. Finally, a simulation example is provided to illustrate the effectiveness of the present decentralized control scheme.
Derong Liu 0001, Ding Wang 0001, Hongliang Li 0002
IEEE Trans. Neural Networks Learn. Syst.2
2014 Online Synchronous Approximate Optimal Learning Algorithm for Multi-Player Non-Zero-Sum Games With Unknown Dynamics
abstract
In this paper, we develop an online synchronous approximate optimal learning algorithm based on policy iteration to solve a multiplayer nonzero-sum game without the requirement of exact knowledge of dynamical systems. First, we prove that the online policy iteration algorithm for the nonzero-sum game is mathematically equivalent to the quasi-Newton's iteration in a Banach space. Then, a model neural network is established to identify the unknown continuous-time nonlinear system using input-output data. For each player, a critic neural network and an action neural network are used to approximate its value function and control policy, respectively. Our algorithm only needs to tune the weights of critic neural networks, so there will be less computational complexity during the learning process. All the neural network weights are updated online in real-time, continuously and synchronously. Furthermore, the uniform ultimate bounded stability of the closed-loop system is proved based on Lyapunov approach. Finally, two simulation examples are given to demonstrate the effectiveness of the developed scheme.
Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2013 Integral Policy Iteration for Zero-Sum Games with Completely Unknown Nonlinear Dynamics
Hongliang Li 0002, Derong Liu 0001, Ding Wang 0001
ICONIP (1)3
2013 Observer-Based Adaptive Output Feedback Control for Nonaffine Nonlinear Discrete-Time Systems Using Reinforcement Learning
Xiong Yang 0001, Derong Liu 0001, Ding Wang 0001
ICONIP (1)3
2013 Neural-network-based zero-sum game for discrete-time nonlinear systems via iterative adaptive dynamic programming algorithm
Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001
Neurocomputing3
2013 Neuro-optimal control for a class of unknown nonlinear dynamic systems using SN-DHP technique
Ding Wang 0001, Derong Liu 0001
Neurocomputing1
2013 An iterative adaptive dynamic programming algorithm for optimal control of unknown discrete-time nonlinear systems with constrained inputs
Derong Liu 0001, Ding Wang 0001, Xiong Yang 0001
Inf. Sci.2
2013 A neural-network-based iterative GDHP approach for solving a class of nonlinear optimal control problems with control constraints
Ding Wang 0001, Derong Liu 0001, Dongbin Zhao, Yuzhu Huang, Dehua Zhang
Neural Comput. Appl.1
2013 Dual iterative adaptive dynamic programming for a class of discrete-time nonlinear systems with time-delays
Qinglai Wei, Ding Wang 0001, Dehua Zhang
Neural Comput. Appl.2
2012 H∞ control of unknown discrete-time nonlinear systems with control constraints using adaptive dynamic programming
abstract
In this paper, we solve the H∞robust optimal control problem for discrete-time nonlinear systems with control saturation constraints using the iterative adaptive dynamic programming algorithm. First, a heuristic dynamic programming algorithm is derived to solve the Hamilton-Jacobi-Isaacs equation associated with the H∞control problem, and a convergence analysis is provided. Then, a dual heuristic dynamic programming algorithm with nonquadratic performance functional is developed to overcome the control saturation constraints. Finally, to facilitate the implementation of the algorithm, four neural networks are used to approximate the unknown nonlinear system, the control policy, the disturbance policy, and the value function.
Derong Liu 0001, Hongliang Li 0002, Ding Wang 0001
IJCNN3
2012 Optimal Task and Energy Scheduling in Dynamic Residential Scenarios
Francesco De Angelis 0002, Matteo Boaro, Danilo Fuselli, Stefano Squartini, Francesco Piazza, Qinglai Wei, Ding Wang 0001
ISNN (1)7
2012 Finite-horizon neuro-optimal tracking control for a class of discrete-time nonlinear systems using adaptive dynamic programming approach
Ding Wang 0001, Derong Liu 0001, Qinglai Wei
Neurocomputing1
2012 Neural-Network-Based Optimal Control for a Class of Unknown Discrete-Time Nonlinear Systems Using Globalized Dual Heuristic Programming
abstract
In this paper, a neuro-optimal control scheme for a class of unknown discrete-time nonlinear systems with discount factor in the cost function is developed. The iterative adaptive dynamic programming algorithm using globalized dual heuristic programming technique is introduced to obtain the optimal controller with convergence analysis in terms of cost function and control law. In order to carry out the iterative algorithm, a neural network is constructed first to identify the unknown controlled system. Then, based on the learned system model, two other neural networks are employed as parametric structures to facilitate the implementation of the iterative algorithm, which aims at approximating at each iteration the cost function and its derivatives and the control law, respectively. Finally, a simulation example is provided to verify the effectiveness of the proposed optimal control approach.
Derong Liu 0001, Ding Wang 0001, Dongbin Zhao, Qinglai Wei
IEEE Trans Autom. Sci. Eng.2
2011 Adaptive dynamic programming for optimal control of unknown nonlinear discrete-time systems
abstract
An intelligent optimal control scheme for unknown nonlinear discrete-time systems with discount factor in the cost function is proposed in this paper. An iterative adaptive dynamic programming (ADP) algorithm via globalized dual heuristic programming (GDHP) technique is developed to obtain the optimal controller with convergence analysis. Three neural networks are used as parametric structures to facilitate the implementation of the iterative algorithm, which will approximate at each iteration the cost function, the optimal control law, and the unknown nonlinear system, respectively. Two simulation examples are provided to verify the effectiveness of the presented optimal control approach.
Derong Liu 0001, Ding Wang 0001, Dongbin Zhao
ADPRL2
2011 Neural-network-based optimal control for a class of nonlinear cdiscrete-time systems with control constraints using the citerative GDHP algorithm
abstract
In this paper, a neural-network-based optimal control scheme for a class of nonlinear discrete-time systems with control constraints is proposed. The iterative adaptive dynamic programming (ADP) algorithm via globalized dual heuristic programming (GDHP) technique is developed to design the optimal controller with convergence proof. Three neural networks are used to facilitate the implementation of the iterative algorithm, which will approximate at each iteration the cost function, the optimal control law, and the controlled nonlinear discrete-time system, respectively. A simulation study is carried out to demonstrate the effectiveness of the present approach in dealing with the nonlinear constrained optimal control problem.
Derong Liu 0001, Ding Wang 0001, Dongbin Zhao
IJCNN2
2011 Optimal Control for a Class of Unknown Nonlinear Systems via the Iterative GDHP Algorithm
Ding Wang 0001, Derong Liu 0001
ISNN (2)1
2011 Finite Horizon Optimal Tracking Control for a Class of Discrete-Time Nonlinear Systems
Qinglai Wei, Ding Wang 0001, Derong Liu 0001
ISNN (2)2