EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Zhao 0001
dblp:02/8134-1
· DBLP profile ↗
23ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0002-1182-4502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 16 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hybrid Resilient and Fault-Tolerant Control of Wind Turbines via Actor-Critic Reinforcement Learning With Prescribed Performance
Jingjie Xie, Hongyang Dong, Xiaowei Zhao 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Adaptive Resilient Output Feedback Control of Power Buffers in DC MicrogridsabstractDirect current microgrids (DCmGs) are attractive due to their high efficiency and ease of deployment, yet decentralized controllers remain vulnerable to false-data-injection (FDI) attacks and abrupt load changes. This article develops a resilient decentralized control architecture that achieves performance guarantees under bounded FDI attacks using only local measurements. The key idea is a decentralized high-gain observer that reconstructs each power buffer input impedance and stored energy while explicitly accounting for network coupling. On this basis, a smooth nonlinear feedback term provides attack compensation without chattering. In contrast to existing schemes, our design comes with explicit linear matrix inequality conditions that, first, certify asymptotic stability in the attack-free case and, second, ensure uniform ultimate boundedness of the closed loop under FDI, thereby yielding transparent tuning ranges for the controller and observer gains. Hardware-in-the-loop experiments, in which the DCmG is subjected to staged FDI attack windows and load steps, are conducted to demonstrate the proposed design with robust and practically tunable resilience for DCmGs. Yongliang Yang 0001, Zhenzhuo Shan, Guilong Liu, Qiaohui He, Xiaowei Zhao 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | Multiagent Inductive Policy OptimizationabstractPolicy optimization methods are promising to tackle high-complexity reinforcement learning (RL) tasks with multiple agents. In this article, we derive a general trust region for policy optimization methods by considering the effect of subpolicy combinations among agents in multiagent environments. Based on this trust region, we propose an inductive objective to train the policy function, which can ensure agents learn monotonically improving policies. Furthermore, we observe that the policy always updates very weakly before falling into a local optimum. To address this, we introduce a cost regarding policy distance in the inductive objective to strengthen the motivation of agents to explore new policies. This approach strikes a balance during training, where the policy update step size remains within the constraints of the trust region, preventing excessive updates while avoiding getting stuck in local optima. Simulations on wind farm (WF) control tasks and two multiagent benchmarks demonstrate the high performance of the proposed multiagent inductive policy optimization (MAIPO) method. Xiaowei Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Wind Farm Control via Offline Reinforcement Learning With Adversarial TrainingabstractIn a wind farm, wakes produced by upstream wind turbines significantly diminish the wind capture of downstream ones, resulting in reduced power generation for the entire farm. Reinforcement learning (RL) control can alleviate the wake effect by enhancing coordination among turbines in arrays. However, training such a collaborative control policy through online RL necessitates millions of interactions with a computational fluid dynamics-based simulator, making it computationally expensive. Wind farm operators possess extensive datasets from past operations, which can tackle the sample generation difficulty. Nevertheless, online RL struggles with the direct use of these offline datasets due to covariate shift and biased value estimation. In this paper, we introduce an offline RL method named Multi-Agent Offline Behavior Imitation (MAOBI) to cope with this challenge. First, by employing f-GANs (Generative Adversarial Networks), MAOBI estimates the divergence between the learning policy and the behavior policy based on samples generated by them. By minimizing this divergence, MAOBI can mimic the behavior policy hidden in the offline datasets. The method then identifies state-action pairs that yield high returns, further improving the control policy. Results demonstrate that MAOBI-trained control policies achieve performance comparable to state-of-the-art online RL methods when deployed in a high-fidelity wind farm simulator. Note to Practitioners—This study addresses a pressing demand in the wind industry: developing high-performance control systems for wind farms to maximize power generation. Although online RL is a popular approach for training collaborative control policies, its reliance on constructing wind farm environments consumes substantial computational resources, such as a supercomputer, to generate training samples. To overcome this limitation, we propose an offline RL algorithm. This algorithm leverages historical operational data from wind farms to search for optimal control policies, eliminating the need for a virtual environment. Our method offers advantages such as being model-free, communication-free, and capable of real-time response. Xiaowei Zhao 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Finite-Horizon Optimal Control for Nonlinear Multi-Input Systems With Online Adaptive Integral Reinforcement LearningabstractIn this paper, a novel adaptive integral reinforcement learning (AIRL) is utilized to online handle the finite-horizon optimum control policies of the partially unknown multi-input nonlinear system. Firstly, the concept of Nash equilibrium is introduced to make the multiple cost functions reach the saddle point. Then, dual neural networks (NNs) are applied to approach the performance index functions based on the integral reinforcement signal. Simultaneously, two novel learning algorithms are proposed to update the NN weights, in which the convergence of weights is proved. Then, the optimal strategies can be obtained by using the obtained weights. The designed controllers based on the data-driven AIRL scheme can avoid the internal state of the system and the derivatives of NN activations in the weight learning process. Finally, the stability of the controlled system is analyzed. An F-16 aircraft model and another nonlinear system are utilized to prove the validity and rationality of the algorithm.Note to Practitioners—There exist many multi-input systems in practical engineering, which includes multi-engine driven F-16 aircraft, large radar servo system and large artillery systems, etc. The finite-horizon optimal control of these systems is crucial for the better performance of system state. However, an accurate engineering model is difficult to obtain, and the finite-horizon cannot generally be achieved. To address these issues, this paper proposes the finite-horizon optimal control for these systems based on a novel adaptive integral reinforcement learning (AIRL). The AIRL can realize an optimal performance for these multi-input systems in finite-time without internal system dynamics, which is a good development for the multi-input system in practical engineering. Yongfeng Lv, Jun Zhao 0015, Xiaowei Zhao 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Wind Turbine Fault-Tolerant Control via Incremental Model-Based Reinforcement LearningabstractA reinforcement learning (RL) based fault-tolerant control strategy is developed in this paper for wind turbine torque & pitch control under actuator & sensor faults subject to unknown system models. An incremental model-based heuristic dynamic programming (IHDP) approach, along with a critic-actor structure, is designed to enable fault-tolerance capability and achieve optimal control. Particularly, an incremental model is embedded in the critic-actor structure to quickly learn the potential system changes, such as faults, in real-time. Different from the current IHDP methods that need the intensive evaluation of the state and input matrices, only the input matrix of the incremental model is dynamically evaluated and updated by an online recursive least square estimation procedure in our proposed method. Such a design significantly enhances the online model evaluation efficiency and control performance, especially under faulty conditions. In addition, a value function and a target critic network are incorporated into the main critic-actor structure to improve our method’s learning effectiveness. Case studies for wind turbines under various working conditions are conducted based on the fatigue, aerodynamics, structures, and turbulence (FAST) simulator to demonstrate the proposed method’s solid fault-tolerance capability and adaptability.Note to Practitioners—This work achieves high-performance wind turbine control under unknown actuator & sensor faults. Such a task is still an open problem due to the complexity of turbine dynamics and potential uncertainties in practical situations. A novel data-driven and model-free control strategy based on reinforcement learning is proposed to handle these issues. The designed method can quickly capture the potential changes in the system and adjust its control policy in real-time, rendering strong adaptability and fault-tolerant abilities. It provides data-driven innovations for complex operational tasks of wind turbines and demonstrates the feasibility of applying reinforcement learning to handle fault-tolerant control problems. The proposed method has a generic structure and has the potential to be implemented in other renewable energy systems. Jingjie Xie, Hongyang Dong, Xiaowei Zhao 0001, Shuyue Lin |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | ADP-Based Optimal Control for Discrete-Time Systems With Safe Constraints and DisturbancesabstractIn this paper, a novel adaptive dynamic programming (ADP)-based optimal control method is developed for discrete-time systems subject to constraints and disturbances. Particularly, a safe policy iteration scheme is designed to handle state and input constraints, including both hard and soft constraints, by converting the original policy improvement strategy into a constrained optimization problem with a prescribed state cost function. After that, an actor-critic-disturbance framework is introduced to address the constrained optimal control problem. The robust safety against disturbances is treated as a two-player zero-sum game, where the actor and disturbance neural networks are used to approximate the optimal control input and the disturbance policy, respectively. The convergence property of the proposed algorithm is analyzed, and the multi-step version of the proposed ADP scheme is derived based on this property. Simulation results are demonstrated and discussed to validate the effectiveness and performance of the proposed method.Note to Practitioners—Addressing constraints in optimal control problems is essential for guaranteeing the safe operation of controlled systems. However, conventional ADP algorithms struggle to simultaneously manage state and control input constraints during the search for the optimal solution. In real-world applications, another critical and common issue is the presence of external disturbances, where disturbances that cause the control object to deviate from the safe region must be constrained while seeking an optimal control policy. Bearing these factors in mind, this study presents a novel ADP scheme for solving optimal control problems of discrete-time systems, taking into account state and control constraints as well as the impact of disturbances. Moreover, the convergence analysis of the proposed SADP scheme is provided, offering a powerful theoretical foundation for guaranteeing the safety and feasibility of the controlled system during operation. Jun Ye 0007, Hongyang Dong, Yougang Bian, Hongmao Qin, Xiaowei Zhao 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | A Transformer-Based Motion Deblurring Network for UAV ImagesabstractWhen performing surveying and mapping missions using drones, motion blur is an unavoidable issue induced by several factors such as vibration, turbulence and wind during operation. Such blurring can significantly degrade the image quality, adversely affecting the accuracy and reliability of downstream applications. In this paper, to effectively eliminate the motion blur in the UAV-captured images, we propose the NAFormer based on the well-established Nonlinear Activation Free Network (NAFNet), which introduces Transformer-based blocks to further enhance its motion-deblurring ability to UAV images. The experiments based on the UAVid dataset demonstrate the effectiveness of the proposed framework. Xiaowei Zhao 0001 |
IGARSS | 2 |
| 2024 | Self-Detection Fine-Tuning: A Framework for Enhancing The Performance of Learning-Based Deblurring Models on UAV DataabstractAffected by unstable disturbances, UAV-based visual data inevitably contains partially blurred segments, consequently reducing the overall effectiveness of the data. Additionally, considering the pronounced regional characteristics of UAV visual data, existing deep learning-based deblurring algorithms face challenges due to the absence of local UAV datasets. To bridge this gap, this paper proposes the Self-Detection Fine-Tuning (SDFT) framework. SDFT only takes a pre-trained deblurring neural network and a blend of clear and blurred UAV data as inputs. This framework filters input data through pre-trained neural networks, obtaining datasets suitable for neural network fine-tuning, thus eliminating the need for any additional datasets or algorithms. Experimental results validate the effectiveness of the SDFT framework in handling UAV-acquired visual data. Weitao Yue, Xiaowei Zhao 0001 |
IGARSS | 2 |
| 2024 | Power Regulation and Load Mitigation of Floating Wind Turbines via Reinforcement LearningabstractFloating offshore wind turbines (FOWTs) are often subjected to heavy structural loads due to challenging operating conditions, which can negatively impact power generation and lead to structural fatigue. This paper proposes a novel reinforcement learning (RL)-based control scheme to address this issue. It combines individual pitch control (IPC) and collective pitch control (CPC) to balance two key objectives: load reduction and power regulation. Specifically, a novel incremental model-based dual heuristic programming (IDHP) strategy is developed as the IPC solution to reduce structural loads. It integrates the online-learned FOWT dynamics into the dual heuristic programming process, making the entire control scheme data-driven and free from dependence on analytical models. Furthermore, the proposed method differs from existing IDHP methods in that only partial system dynamics need to be learned, resulting in a simplified design structure and improved training efficiency. Tests using a high-fidelity FOWT simulator demonstrate the effectiveness of the proposed method.Note to Practitioners—This work achieves power regulation and load reduction simultaneously for FOWTs to guarantee the reliability of wind turbine operations. Such a task is still an open problem because existing FOWT controllers commonly rely on accurate turbine models and lack adaptability to potential uncertainties and errors in practical situations. A new data-driven, model-free control strategy based on the RL technique is developed to address these issues. Our method has the ability to capture potential changes in system dynamics by updating a so-called incremental model via online measurements. Unlike current advances in this direction that need to approximate the whole system dynamics, the proposed control algorithm only needs to update partial system information for the incremental model. This naturally simplifies the design structure and enhances learning effectiveness while providing adaptability and robustness against uncertainties and errors. The proposed control strategy can also be extended and implemented in other systems, such as autonomous systems and other renewable energy systems. Jingjie Xie, Hongyang Dong, Xiaowei Zhao 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Adaptive Fuzzy Practical Bipartite Synchronization for Multiagent Systems With Intermittent Feedback Under Multiple Unknown Control DirectionsabstractIn this article, we propose an adaptive fuzzy control design for the distributed competitive control problem of multiagent systems (MASs) with multiple unknown control directions. The bipartite synchronization control is investigated by using the fuzzy backstepping control framework and fuzzy logic systems. To broaden the application field for the distributed protocol design, we consider practical bipartite synchronization for a group of MASs consisting of followers subject to heterogeneous unknown control directions. To address these multiple unknown control directions, a novel Nussbaum-type function is developed. Moreover, to reduce the communication bandwidth, this article proposes two threshold strategies for event-triggered control to avoid any unnecessary sampling while taking flexibility into consideration, further improving the efficiency and feasibility of the developed bipartite protocol design. The experimental results indicate that the proposed control method can effectively realize bipartite synchronization of MASs with multiple unknown control directions. Guilong Liu, Yongliang Yang 0001, Xiaowei Zhao 0001, Choon Ki Ahn |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | LSwinSR: UAV Imagery Super-Resolution Based on Linear Swin TransformerabstractSuper-resolution, which aims to reconstruct high-resolution (HR) images from low-resolution (LR) images, has drawn considerable attention and has been intensively studied in computer vision and remote sensing communities. Super-resolution technology is especially beneficial for unmanned aerial vehicles (UAVs), as the number and resolution of images captured by UAVs are highly limited by physical constraints such as flight altitude and load capacity. In the wake of the successful application of deep learning methods in the super-resolution task, in recent years, a series of super-resolution algorithms have been developed. In this article, for the super-resolution of UAV images, a novel network based on the state-of-the-art Swin Transformer is proposed with better efficiency and competitive accuracy. Meanwhile, as one of the essential applications of the UAV is land cover and land use monitoring, simple image quality assessments such as the peak-signal-to-noise ratio (PSNR) and the structural similarity index measure (SSIM) are not enough to comprehensively measure the performance of an algorithm. Therefore, we further investigate the effectiveness of super-resolution methods using the accuracy of semantic segmentation. The code is available athttps://github.com/lironui/GeoSR. Rui Li 0036, Xiaowei Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Reinforcement Learning-Based Multiobjective Control of Grid-Connected Wind FarmsabstractThis article aims to design a multiobjective controller with a primary focus on mitigating the wake interference in wind farms to increase their long-term power generation compared with the method without coordination in tackling the wake effect. What is more, this farm-level controller endeavors to keep turbines online in the event of a voltage fault and offer ancillary services to the grid while minimizing the farm power losses. Model-free reinforcement learning (RL) is promising in coping with such challenges since it can train a multiobjective control policy by interacting with the environment. In this article, we first construct a grid-connected wind farm simulator as the environment to collect training samples for model-free RL. Then, we erect three wind farm control objectives: maximizing the power generation of the wind farm by overcoming the wake effect, facilitating the fault-ride-through capability of wind turbines, and enhancing the frequency stability of the grid. Finally, an appropriate reward function is designed to induce RL agents to achieve these objectives. Results show that the proposed method can obviously increase the power generation compared with the method ignoring the wake effect. Furthermore, the method showcases it is capable of controlling turbines to ride through voltage faults of the grid and assisting the grid to return to its nominal frequency. Xiaowei Zhao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Inverse-model-based iterative learning control for unknown MIMO nonlinear system with neural networkabstractThis paper provides an inverse-model-based iterative learning control (ILC) for the unknown multi-input multi-output (MIMO) nonlinear system with neural network (NN), where a novel gradient adaptive law is used to update the NN weights both hidden and output layers such a faster convergence can be achieved. First, a three-layer NN structure is introduced to observe the MIMO nonlinear system with input–output data, and a new gradient algorithm is proposed to update the unknown parameters of both hidden and output layers. Then, the input dynamic can be obtained with the NN observer, and the inversion-model-based control is designed. Moreover, the ideal inversion control can be obtained based on the reference signal, and the inverse ILC is designed. The stability of the NN observer and the convergence of the inverse-model-based control are analyzed. Finally, a SCARA manipulator MIMO model is simulated to illustrate the correctness of the proposed methods. Yongfeng Lv, Xuemei Ren, Jianyan Tian, Xiaowei Zhao 0001 |
Neurocomputing | 4 |
| 2023 | Data-Based Optimal Microgrid Management for Energy Trading With Integral Q-Learning SchemeabstractThis article proposes an integral${Q}$-learning scheme to study the optimum management strategy of the battery for energy trading in the microgrid. The obtained optimal strategies are managed to minimize the cost of the microgrid, and simultaneously guarantee the better performance of the battery such that service life of battery is extended. First, the microgrid model is constructed, where renewable energy and profiled loads are considered. To satisfy the microgrid demand and the environmental energy consumption, an integral${Q}$-learning scheme will be developed for the input policy of the microgrid energy system such that the model dynamics are avoided. Moreover, the minimized cost and the property of the battery are considered in the performance function. Rather than the traditional${Q}$-learning iterative scheme, this article proposes a self-learning architecture for the integral-${Q}$learning scheme with two neural networks. The first network can learn the optimal${Q}$-function. Another network is used to learn the optimal action such that the${Q}$-function is minimized and load demand can be satisfied. The network weights are updated with the gradient method and the stability analysis is presented. Finally, the experiment data is used to verify the proposed integral${Q}$-learning scheme. Yongfeng Lv, Zhaolong Wu, Xiaowei Zhao 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Reinforcement Learning-Based Wind Farm Control: Toward Large Farm Applications via Automatic Grouping and Transfer LearningabstractThe high system complexity and strong wake effects bring significant challenges to wind farm operations. Conventional wind farm control methods may lead to degraded power generation efficiency. A reinforcement learning (RL)-based approach is proposed in this paper to handle these issues, which can increase the long-term farm-level power generation subject to strong wake effects while without requiring analytical wind farm models. The proposed method is significantly distinct from existing RL-based wind farm control approaches, whose computational complexities usually increase heavily with the increase of total turbine numbers. In contrast, our method can greatly reduce training loads and enhance learning efficiency via two novel designs: (1) automatic grouping and (2) multi-agent-based transfer learning (MATL). Automatic Grouping can divide a large wind farm into small turbine groups by analyzing the aerodynamic interactions between turbines and utilizing some key principles from the graph theory. It enables the separated conduction of RL algorithms on small turbine groups, avoiding the complex training process and high computational costs of applying RL on the entire farm. Based on Automatic Grouping, MATL can further reduce the computational complexity by allowing agents (i.e. wind turbines) to inherit control policies under potential group changes. Case studies with a dynamical simulator show that the proposed method achieves clear power generation increases than the benchmark. It also dramatically reduces computational costs compared with typical RL-based wind farm control methods, paving the way for the application of RL in general wind farms. Hongyang Dong, Xiaowei Zhao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | A Multiagent Reinforcement Learning Approach for Wind Farm Frequency ControlabstractAs wind turbines (WTs) become more prevalent, there is an increasing interest in actively controlling their power output to participate in the frequency regulation for the power grid. Conventional frequency regulation controllers use fixed gains, making it difficult for the WT to adjust its kinetic energy uptake to its operating conditions and to collaborate effectively with other WTs in the wind farm. In addition, the design of conventional frequency controllers does not consider their impacts on the mechanical structure. To address these issues, in this article, we model the cooperative frequency control problem for all the WTs in a wind farm as a decentralized partially observable Markov decision process and use a multiagent deep reinforcement learning algorithm to solve it. We also develop a grid-connected wind farm simulation model based on MATLAB/Simulink and OpenFAST, which can reflect the detailed interactions between the electrical and mechanical components of WTs. Simulation results show that the proposed strategy is effective in reducing frequency drops and has less impact on mechanical structure deflections compared with traditional methods. Yanchang Liang, Xiaowei Zhao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Multi-H∞ Controls for Unknown Input-Interference Nonlinear System With Reinforcement LearningabstractThis article studies the multi- [Formula: see text] controls for the input-interference nonlinear systems via adaptive dynamic programming (ADP) method, which allows for multiple inputs to have the individual selfish component of the strategy to resist weighted interference. In this line, the ADP scheme is used to learn the Nash-optimization solutions of the input-interference nonlinear system such that multiple [Formula: see text] performance indices can reach the defined Nash equilibrium. First, the input-interference nonlinear system is given and the Nash equilibrium is defined. An adaptive neural network (NN) observer is introduced to identify the input-interference nonlinear dynamics. Then, the critic NNs are used to learn the multiple [Formula: see text] performance indices. A novel adaptive law is designed to update the critic NN weights by minimizing the Hamiltonian-Jacobi-Isaacs (HJI) equation, which can be used to directly calculate the multi- [Formula: see text] controls effectively by using input-output data such that the actor structure is avoided. Moreover, the control system stability and updated parameter convergence are proved. Finally, two numerical examples are simulated to verify the proposed ADP scheme for the input-interference nonlinear system. Yongfeng Lv, Jing Na, Xiaowei Zhao 0001, Yingbo Huang, Xuemei Ren |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Wind-Farm Power Tracking Via Preview-Based Robust Reinforcement LearningabstractThis article aims to address the wind-farm power tracking problem, which requires the farm's total power generation to track time-varying power references and, therefore, allows the wind farm to participate in ancillary services such as frequency regulation. A novel preview-based robust deep reinforcement learning (PR-DRL) method is proposed to handle such tasks which are subject to uncertain environmental conditions and strong aerodynamic interactions among wind turbines. To our knowledge, this is for the first time that a data-driven model-free solution is developed for wind-farm power tracking. Particularly, reference signals are treated as preview information and embedded in the system as specially designed augmented states. The control problem is then transformed into a zero-sum game to quantify the influence of unknown wind conditions and future reference signals. Built upon the$H_\infty$control theory, the proposed PR-DRL method can successfully approximate the resulting zero-sum game's solution and achieve wind-farm power tracking. Time-series measurements and long short-term memory networks are employed in our DRL structure to handle the non-Markovian property induced by the time-delayed feature of aerodynamic interactions. Tests based on a dynamic wind-farm simulator demonstrate the effectiveness of the proposed PR-DRL wind-farm control strategy. Hongyang Dong, Xiaowei Zhao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Network Identification Using μ-PMU and Smart Meter MeasurementsabstractThe network identification plays a very prominent role for the network operator to accomplish the various objectives such as state-estimation, monitoring, control, planning, and real-time analytics. The network structure varies from time-to-time and its details are often not available with the network operator. To address this issue, in this article, an alternating direction method of multipliers (ADMM) based framework is presented herein to identify the network topology and line parameters using smart meter and microphasor measurement unit (μ) measurements. The presented algorithm is divided into two sections 1) approximate parameter evaluation through regression, to extract the partial topology information and 2) complete network topology identification through the ADMM framework. This algorithm accomplishes the objectives of identifying the network configuration, branch parameters (e.g., conductance and susceptance), and change in branch parameters. Simulation results demonstrate the effectiveness of the presented algorithm on the benchmarked IEEE 13-bus and IEEE 123-bus feeders under various operating scenarios. Furthermore, the presented framework illustrates excellent network identification even with the presence of the stochastic nature of renewable power generation. The presented algorithm exhibits an excellent performance even with the consideration of noise in both measurements. In addition, the comparative performance is carried out on the benchmarked unbalanced IEEE 13-bus and balanced IEEE 33-bus feeders to highlight the efficacy of the presented framework over the state-of-art framework. Priyank Shah, Xiaowei Zhao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Wind Farm Power Generation Control Via Double-Network-Based Deep Reinforcement LearningabstractA model-free deep reinforcement learning (DRL) method is proposed in this article to maximize the total power generation of wind farms through the combination of induction control and yaw control. Specifically, a novel double-network (DN)-based DRL approach is designed to generate control policies for thrust coefficients and yaw angles simultaneously and separately. Two sets of critic-actor networks are constructed to this end. They are linked by a central power-related reward, providing a coordinated control structure while inheriting the critic-actor mechanism's advantages. Compared with conventional DRL methods, the proposed DN-based DRL strategy can adapt to the distinctive and incompatible features of different control inputs, guaranteeing a reliable training process and ensuring superior performance. Also, the prioritized experience replay strategy is utilized to improve the training efficiency of deep neural networks. Simulation tests based on a dynamic wind farm simulator show that the proposed method can significantly increase the power generation for wind farms with different layouts. Jingjie Xie, Hongyang Dong, Xiaowei Zhao 0001, Aris Karcanias |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Optimal Tracking Control for Uncertain Nonlinear Systems With Prescribed Performance via Critic-Only ADPabstractThis article addresses the tracking control problem for a class of nonlinear systems described by Euler–Lagrange equations with uncertain system parameters. The proposed control scheme is capable of guaranteeing prescribed performance from two aspects: 1) a special parameter estimator with prescribed-performance properties is embedded in the control scheme. The estimator not only ensures the exponential convergence of the estimation errors under relaxed excitation conditions but also can restrict all estimates to predetermined bounds during the whole estimation process and 2) the proposed controller can strictly guarantee the user-defined performance specifications on tracking errors, including convergence rate, maximum overshoot, and residual set. More importantly, it has the optimizing ability for the tradeoff between performance and control cost. A state transformation method is employed to transform the constrained optimal tracking control problem to an unconstrained stationary optimal problem. Then, a critic-only adaptive dynamic programming algorithm is designed to approximate the solution of the Hamilton–Jacobi–Bellman equation and the corresponding optimal control policy. Uniformly ultimately bounded stability is guaranteed via a Lyapunov-based stability analysis. Finally, numerical simulation results demonstrate the effectiveness of the proposed control scheme. Hongyang Dong, Xiaowei Zhao 0001, Biao Luo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Reinforcement Learning-Based Structural Control of Floating Wind TurbinesabstractThe structural control of floating wind turbines using active tuned mass damper is investigated in this article. To our knowledge, this is for the first time that reinforcement learning-based control approach is employed to this type of application. Specifically, an adaptive dynamic programming (ADP) algorithm is used to derive the optimal control law based on the nonlinear structural dynamics, and the large-scale machine learning platform Tensorflow is employed for the design and implementation of the neural network (NN) structure. Three fully connected NNs, i.e., a plant network, a critic network, and an action network, are included in the proposed NN structure. Their training requires the gradient information flowing through the whole network, which is tackled by automatic differentiation, a popular technique for deriving the gradients of complex networks automatically. While to our knowledge, the network structures in the existing literature are rather simple and the training of the hidden layer is usually ignored. This allows their gradients to be derived analytically, which is infeasible with complex network structures. Thus, automatic differentiation greatly improves the employed ADP algorithm’s ability in solving complex problems. The simulation results of structural control of floating wind turbines show that ADP controller performs very well in both normal and extreme conditions, with the standard deviation of the platform pitch displacement being reduced by around 40%. A clear advantage of ADP controllers over the$H_{\infty }$controller is observed, especially in extreme conditions. Moreover, our design considers the tradeoff between the control performance and power consumption. Jincheng Zhang 0004, Xiaowei Zhao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |