Zhong-Ping Jiang

dblp:41/245 · DBLP profile ↗
← Back
55ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-4868-9359ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Computer networks · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Physically structured and robust distributed recurrent neural networks for modeling hydropower generator cooling systems
Xianning Li, Zhun Yin, Wenbo Jia, Hong Wang 0001, Zhong-Ping Jiang
Neurocomputing5
2026 Model-Free Output Regulation of Networked Systems Under Unknown Hybrid Attacks
abstract
This article considers the output regulation problem for an unknown discrete-time system subject to the random combination of denial-of-service, replay, and deception attacks on both sensor-controller and controller-actuator channels. We propose a learning-based receding-horizon control with historical output signals. It offers two advantages over state and output feedback regulators in the sense that it requires neither exact knowledge of system dynamics nor a direct measurement of external disturbance on one hand, and on the other hand, it can counteract the adverse impact of hybrid attacks on the executive capability of the actuator, regardless of the seriously tampered data on the sensor-controller channel. To overcome technical difficulties from hybrid attacks on both channels, we generalize the Markov-parameter-based time-series control method to generate a data packet containing the current and future control inputs, which are further compromised on the controller-actuator channel. Thus, a recovery procedure is additionally designed to solve the model-free output regulation problem by distinguishing the undamaged predicted inputs based on the proposed hybrid attack detection procedure.
Xiran Cui, Zhengguang Wu, Yi Dong 0001, Zhong-Ping Jiang
IEEE Trans. Cybern.4
2026 Corrections to "A Fully Data-Driven Value Iteration for Stochastic LQR: Convergence, Robustness, and Stability"
abstract
In the above article, an error in [1, Th. 2].
Leilei Cui 0002, Zhong-Ping Jiang, Petter N. Kolm, Grégoire G. Macqueron
IEEE Trans. Neural Networks Learn. Syst.2
2026 DeeP-TE: Data-Enabled Predictive Traffic Engineering
abstract
Routing configurations of a network should constantly adapt to traffic variations to achieve good network performance. Adaptive routing faces two main challenges: 1) how to accurately measure/estimate time-varying traffic matrices? 2) how to control the network and application performance degradation caused by frequent route changes? In this paper, we develop a novel data-enabled predictive traffic engineering (DeeP-TE) algorithm that minimizes the network congestion by gracefully adapting routing configurations over time. Our control algorithm can generate routing updates directly from the historical routing data and the corresponding link rate data, without direct traffic matrix measurement or estimation. Numerical experiments on real network topologies with real traffic matrices demonstrate that the proposed DeeP-TE routing adaptation algorithm can achieve close-to-optimal control effectiveness with significantly lower routing variations than the baseline methods.
Zhun Yin, Lifan Mei, Yong Liu 0013, Zhong-Ping Jiang
IEEE Trans. Netw.5
2025 Experimental Results in Cyber-Physical Transportation Systems: A Case Study in Cybersecurity
abstract
This paper presents experimental results from a learning-based control framework for cyber-physical transportation systems. Building on theoretical guarantees that establish an upper bound on denial-of-service (DoS) attack durations to maintain closed-loop stability, we deploy a resilient learning-based lane-changing control algorithm on a remote-controlled (RC) autonomous vehicle equipped with GPS, IMU, and camera sensors, interfaced with an Nvidia Jetson AGX Xavier board. The algorithm leverages real-time sensor data to make suboptimal yet robust lane-change decisions while enduring intermittent DoS attacks that disrupt communication. Our experiments confirm the resilience of this learning-based approach, demonstrating safe and efficient maneuvers under adversarial conditions in obstacle-rich driving scenarios. By highlighting these experimental findings, this work underscores the importance of cybersecurity in next-generation vehicle control algorithms for autonomous transportation applications.
Won Yong Ha, Kaan Özbay, Zhong-Ping Jiang
IV4
2025 Physics-informed Machine Learning with Heuristic Feedback Control Layer for Autonomous Vehicle Control
abstract
This paper proposes a novel physics-informed ma-chine learning framework for motion planning and control of autonomous vehicles. By integrating longitudinal and lat-eral control, a nonlinear control problem is formulated using Model Predictive Control (MPC). To address computational challenges, a self-supervised framework, Recurrent Predictive Control (RPC), is introduced, leveraging differentiable neural networks and recurrent neural networks to train a neural network controller. Additionally, a heuristic feedback control layer is designed to reduce steady-state errors in the closed-loop tracking. Through numerical simulations and co-simulations using Simulink and CarSim, five neural network controllers are compared with an MPC controller in a lane-changing sce-nario. The proposed RPC framework improves computational efficiency by 95% compared to MPC, enhances generalization performance compared to Approximate MPC, and reduces performance loss by 17% compared to Differentiable Predictive Control. The heuristic feedback control layer further reduces steady-state errors and improves convergence speed during training.
Xianning Li, Yebin Wang, Kaan Özbay, Zhong-Ping Jiang
IV4
2025 Robust Lyapunov Optimization for LEO Satellite Networks Routing Control
abstract
Low Earth Orbit (LEO) satellite networks are emerging as crucial components of space-air-ground integrated networks (SAGINs), extending beyond terrestrial capabilities to provide global data transmission services for the Internet of Things (IoT) and mobile devices. The proliferation of connected devices has led to increased data volumes and highly variable, bursty traffic patterns, thus posing significant challenges for network stability and necessitating effective routing control mechanisms. Traditional Lyapunov optimization methods have been fundamental in network optimization, offering stability guarantees under the assumption that traffic flows remain strictly within the network's capacity region. However, this assumption is often violated in LEO satellite networks due to their dynamic and bursty nature, thereby rendering conventional approaches inadequate for ensuring stability. To address this challenge, we propose a robust Lyapunov optimization framework tailored for LEO satellite networks. Our method relaxes the strict requirements of traditional Lyapunov optimization by allowing the network to tolerate finite violations of the capacity region while still ensuring overall system stability. This approach demonstrates that, for a stabilizable network system, it is not necessary for traffic to remain within the capacity region at every time slot. We validate the effectiveness of the proposed robust Lyapunov optimization through extensive simulations under various traffic conditions and LEO satellite network configurations. The results confirm that LEO satellite networks can maintain stability despite finite violations of the capacity region, ensuring reliable performance amid dynamic and bursty traffic demands.
Zhemin Huang 0002, Zhong-Ping Jiang, Zhu Han 0001, Yong Liu 0013
IEEE Trans. Mob. Comput.2
2024 Guest Editorial Special Issue on Reinforcement Learning-Based Control: Data-Efficient and Resilient Methods
abstract
As an important branch of machine learning, reinforcement learning (RL) has proved its efficiency in many emerging applications in science and engineering. A remarkable advantage of RL is that it enables agents to maximize their cumulative rewards through online exploration and interactions with unknown (or partially unknown) and uncertain environments, which is regarded as a variant of data-driven adaptive optimal control methods. However, the successful implementation of RL-based control systems usually relies on a good quantity of online data due to its data-driven nature. Therefore, it is imperative to develop data-efficient RL methods for control systems to reduce the required number of interactions with the external environment. Moreover, network-aware issues, such as cyberattacks, dropout packet and communication latency, and actuator and sensor faults, are challenging conundrums that threaten the safety, security, stability, and reliability of network control systems. Consequently, it is significant to develop safe and resilient RL mechanisms.
Weinan Gao, Na Li 0002, Kyriakos G. Vamvoudakis, F. Richard Yu, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.5
2024 Reducing Urban Traffic Congestion Using Deep Learning and Model Predictive Control
abstract
This article proposes a deep learning (DL)-based control algorithm-DL velocity-based model predictive control (VMPC)-for reducing traffic congestion with slowly time-varying traffic signal controls. This control algorithm consists of system identification using DL and traffic signal control using VMPC. For the training process of DL, we established a modeling error entropy loss as the criteria inspired by the theory of stochastic distribution control (SDC) originated by the fourth author. Simulation results show that the proposed algorithm can reduce traffic congestion with a slowly varying traffic signal control input. Results of an ablation study demonstrate that this algorithm compares favorably to other model-based controllers in terms of prediction error, signal varying speed, and control effectiveness.
Zhun Yin, Tong Liu 0026, Chieh Ross Wang, Hong Wang 0001, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.5
2023 Data-driven cooperative output regulation of multi-agent systems under distributed denial of service attacks
Weinan Gao, Zhong-Ping Jiang
Sci. China Inf. Sci.2
2022 Learning-Based Adaptive Optimal Control for Connected Vehicles in Mixed Traffic: Robustness to Driver Reaction Time
abstract
Through vehicle-to-vehicle (V2V) communication, both human-driven and autonomous vehicles can actively exchange data, such as velocities and bumper-to-bumper distances. Employing the shared data, control laws with improved performance can be designed for connected and autonomous vehicles (CAVs). In this article, taking into account human-vehicle interaction and heterogeneous driver behavior, an adaptive optimal control design method is proposed for a platoon mixed with multiple preceding human-driven vehicles and one CAV at the tail. It is shown that by using reinforcement learning and adaptive dynamic programming techniques, a near-optimal controller can be learned from real-time data for the CAV with V2V communications, but without the precise knowledge of the accurate car-following parameters of any driver in the platoon. The proposed method allows the CAV controller to adapt to different platoon dynamics caused by the unknown and heterogeneous driver-dependent parameters. To improve the safety performance during the learning process, our off-policy learning algorithm can leverage both the historical data and the data collected in real time, which leads to considerably reduced learning time duration. The effectiveness and efficiency of our proposed method is demonstrated by rigorous proofs and microscopic traffic simulations.
Mengzhe Huang, Zhong-Ping Jiang, Kaan Özbay
IEEE Trans. Cybern.2
2022 Reinforcement Learning and Adaptive Optimal Control for Continuous-Time Nonlinear Systems: A Value Iteration Approach
abstract
This article studies the adaptive optimal control problem for continuous-time nonlinear systems described by differential equations. A key strategy is to exploit the value iteration (VI) method proposed initially by Bellman in 1957 as a fundamental tool to solve dynamic programming problems. However, previous VI methods are all exclusively devoted to the Markov decision processes and discrete-time dynamical systems. In this article, we aim to fill up the gap by developing a new continuous-time VI method that will be applied to address the adaptive or nonadaptive optimal control problems for continuous-time systems described by differential equations. Like the traditional VI, the continuous-time VI algorithm retains the nice feature that there is no need to assume the knowledge of an initial admissible control policy. As a direct application of the proposed VI method, a new class of adaptive optimal controllers is obtained for nonlinear systems with totally unknown dynamics. A learning-based control algorithm is proposed to show how to learn robust optimal controllers directly from real-time data. Finally, two examples are given to illustrate the efficacy of the proposed methodology.
Tao Bian, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.2
2022 Reinforcement Learning-Based Cooperative Optimal Output Regulation via Distributed Adaptive Internal Model
abstract
In this article, a data-driven distributed control method is proposed to solve the cooperative optimal output regulation problem of leader-follower multiagent systems. Different from traditional studies on cooperative output regulation, a distributed adaptive internal model is originally developed, which includes a distributed internal model and a distributed observer to estimate the leader's dynamics. Without relying on the dynamics of multiagent systems, we have proposed two reinforcement learning algorithms, policy iteration and value iteration, to learn the optimal controller through online input and state data, and estimated values of the leader's state. By combining these methods, we have established a basis for connecting data-distributed control methods with adaptive dynamic programming approaches in general since these are the theoretical foundation from which they are built.
Weinan Gao, Mohammed Mynuddin, Donald C. Wunsch II, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.4
2021 Robust Reinforcement Learning: A Case Study in Linear Quadratic Regulation
abstract
This paper studies the robustness of reinforcement learning algorithms to errors in the learning process. Specifically, we revisit the benchmark problem of discrete-time linear quadratic regulation (LQR) and study the long-standing open question: Under what conditions is the policy iteration method robustly stable from a dynamical systems perspective? Using advanced stability results in control theory, it is shown that policy iteration for LQR is inherently robust to small errors in the learning process and enjoys small-disturbance input-to-state stability: whenever the error in each iteration is bounded and small, the solutions of the policy iteration algorithm are also bounded, and, moreover, enter and stay in a small neighborhood of the optimal LQR solution. As an application, a novel off-policy optimistic least-squares policy iteration for the LQR problem is proposed, when the system dynamics are subjected to additive stochastic disturbances. The proposed new results in robust reinforcement learning are validated by a numerical example.
Bo Pang 0002, Zhong-Ping Jiang
AAAI2
2021 Balance Control of a Novel Wheel-legged Robot: Design and Experiments
abstract
This paper presents a balance control technique for a novel wheel-legged robot. We first derive a dynamic model of the robot and then apply a linear feedback controller based on output regulation and linear quadratic regulator (LQR) methods to maintain the standing of the robot on the ground without moving backward and forward mightily. To take into account nonlinearities of the model and obtain a large domain of stability, a nonlinear controller based on the interconnection and damping assignment - passivity-based control (IDA-PBC) method is exploited to control the robot in more general scenarios. Physical experiments are performed with various control tasks. Experimental results demonstrate that the proposed linear output regulator can maintain the standing of the robot, while the proposed nonlinear controller can balance the robot under an initial starting angle far away from the equilibrium point, or under a changing robot height.
Shuai Wang 0007, Leilei Cui 0002, Jingfan Zhang, Jie Lai, Yu Zheng 0001, Zhengyou Zhang, Zhong-Ping Jiang
ICRA9
2021 Detection and localization of biased load attacks in smart grids via interval observer
Xiaoyuan Luo, Zhong-Ping Jiang, Xin-Ping Guan
Inf. Sci.4
2021 An Optimal Primary Frequency Control Based on Adaptive Dynamic Programming for Islanded Modernized Microgrids
abstract
In many pilot research and development (R&D) microgrid projects, engine-based generators are employed in their power systems, either generating electrical energy or being mixed with the heat and power technology. One of the critical tasks of such engine-based generation units is the frequency regulation in the islanded mode of modernized microgrid (MMG) operation; MMGs are microgrids equipped with advanced controls to address more emerging scenarios in smart grids. For having a stable and reliable MMG, we need to synthesize an optimal, robust, primary frequency controller for the islanded mode of MMG of the future. This task is challenging because of unknown mechanical parameters, occurrence of uncertain disturbances, uncertainty of loads, operating point variations, and the appearance of engine delays, and hence nonminimum phase dynamics. This article presents an innovative primary frequency control for the engine generators regulating the frequency of an islanded MMG in the context of smart grids. The proposed approach is based on an adaptive optimal output-feedback control algorithm using adaptive dynamic programming (ADP). The convergence of algorithms, along with the stability analysis of the closed-loop system, is also shown in this article. Finally, as experimental validation, hardware-in-the-loop (HIL) test results are provided in order to examine the effectiveness of the proposed methodology practically.Note to Practitioners—This article was motivated by the problem of primary frequency controls in modernized microgrids (MMGs) using engine generators, which are still one of the prime sources of regulating frequency in pilot research and development (R&D) microgrid projects. Although MMGs will be integral parts of the smart grid of the future, their primary controls in the islanded mode are not advanced enough and not considering existing theoretical challenges scientifically. Existing approaches to regulate frequency using industrially accepted methods are highly model-based and not optimal. Besides, they are not considering the nonminimum phase dynamics. These dynamics are mainly associated with the engine delays—an inherent issue of mechanical parts—for islanded microgrids. This article suggests a new adaptive optimal output-feedback control approach based on the adaptive dynamic programming (ADP) to the abovementioned problem under consideration. By using the proposed methodology, MMGs can deal with the issues mentioned earlier, which are challenging. The proposed approach is optimally rejecting uncertain disturbances (considering the load uncertainty and operating point variations) and reducing the impacts of nonminimum phase dynamics caused by the engine delay. Based on our currently available hardware-in-the-loop (HIL) device’s capability of modeling power systems’ components in real time, our HIL-based experiments demonstrate that this approach is feasible.
Masoud Davari, Weinan Gao, Zhong-Ping Jiang, Frank L. Lewis
IEEE Trans Autom. Sci. Eng.3
2021 Cooperative Formation Control Under Switching Topology: An Experimental Case Study in Multirotors
abstract
This article presents a novel design algorithm for the cooperative formation control of multirotors with directed and switching topology. A key strategy is to transform the formation control problem into an output agreement problem for which a class of cooperative controllers with successive loops is developed to achieve output agreement. For practical implementation, velocities and accelerations of the controlled multirotors are restricted to within desired ranges by introducing appropriate saturations to the loops. It is proved that each controlled multirotor admits an invariant set property, and the formation control objective can be achieved if a mild joint connectivity condition is satisfied by the switching topology. Along the way, this article also proves a result of independent interest in the output agreement problem subject to both velocity and control input constraints with switching topology. Numerical simulations and physical experiments are employed to verify the effectiveness of the proposed design.
Zhanxiu Wang, Tengfei Liu 0004, Zhong-Ping Jiang
IEEE Trans. Cybern.3
2021 A Secure Control Learning Framework for Cyber-Physical Systems Under Sensor and Actuator Attacks
abstract
In this article, we develop a learning-based secure control framework for cyber-physical systems in the presence of sensor and actuator attacks. Specifically, we use a bank of observer-based estimators to detect the attacks while introducing a threat-detection level function. Under nominal conditions, the system operates with a nominal-feedback controller with the developed attack monitoring process checking the reliance of the measurements. If there exists an attacker injecting attack signals to a subset of the sensors and/or actuators, then the attack mitigation process is triggered and a two-player, zero-sum differential game is formulated with the defender being the minimizer and the attacker being the maximizer. Next, we solve the underlying joint state estimation and attack mitigation problem and learn the secure control policy using a reinforcement-learning-based algorithm. Finally, two illustrative numerical examples are provided to show the efficacy of the proposed framework.
Yuanqiang Zhou, Kyriakos G. Vamvoudakis, Wassim M. Haddad 0001, Zhong-Ping Jiang
IEEE Trans. Cybern.4
2021 Distributed Event-Triggered Formation Control of Multiagent Systems via Complex-Valued Laplacian
abstract
Event-triggered formation control of multiagent systems under an undirected communication graph is investigated using complex-valued Laplacian. Both continuous-time and discrete-time models are considered. The dynamics of each agent is described by complex-valued differential or difference equations. For each agent, only the discrete-time information of its neighbors is used in the design of formation controllers and event triggers. Triggering time instants for any agent are determined by certain events that depend on the states of its neighboring agents. Continuous updating of controllers and continuous communication among neighboring agents are avoided. The obtained results show that formation can reach specific but arbitrary formation shape. Furthermore, it is shown that the closed-loop system does not exhibit the Zeno phenomenon for the continuous-time dynamics case or the Zeno-like behavior for the discrete-time dynamics case. Finally, numerical simulations for both the continuous-time and the discrete-time dynamics cases are presented to illustrate the effectiveness of the proposed distributed event-triggered control methods.
Wei Zhu 0004, Wenji Cao, Zhong-Ping Jiang
IEEE Trans. Cybern.3
2021 Event-Triggered Adaptive Optimal Control With Output Feedback: An Adaptive Dynamic Programming Approach
abstract
This article presents an event-triggered output-feedback adaptive optimal control method for continuous-time linear systems. First, it is shown that the unmeasurable states can be reconstructed by using the measured input and output data. An event-based feedback strategy is then proposed to reduce the number of controller updates and save communication resources. The discrete-time algebraic Riccati equation is iteratively solved through event-triggered adaptive dynamic programming based on both policy iteration (PI) and value iteration (VI) methods. The convergence of the proposed algorithm and the closed-loop stability is carried out by using the Lyapunov techniques. Two numerical examples are employed to verify the effectiveness of the design methodology.
Fuyu Zhao, Weinan Gao, Zhong-Ping Jiang, Tengfei Liu 0004
IEEE Trans. Neural Networks Learn. Syst.3
2020 Nonlinear Balance Control of an Unmanned Bicycle: Design and Experiments
abstract
In this paper, nonlinear control techniques are exploited to balance an unmanned bicycle with enlarged stability domain. We consider two cases. For the first case when the autonomous bicycle is balanced by the flywheel, the steering angle is set to zero, and the torque of the flywheel is used as the control input. The controller is designed based on the Interconnection and Damping Assignment Passivity Based Control (IDA-PBC) method. For the second case when the bicycle is balanced by the handlebar, the bicycle's velocity is high, and the flywheel is turned off. The angular velocity of the handlebar is used as the control input and the balance controller is designed based on feedback linearization. In these cases, the global stability of the closed-loop unmanned bicycle is theoretically proved based on Lyapunov theory. The experiments are conducted to validate the efficacy of the proposed nonlinear balance controllers.
Leilei Cui 0002, Shuai Wang 0007, Jie Lai, Xiangyu Chen 0001, Zhengyou Zhang, Zhong-Ping Jiang
IROS7
2020 Gain Scheduled Controller Design for Balancing an Autonomous Bicycle
abstract
In this paper, the gain scheduling technique is applied to design a balance controller for an autonomous bicycle with an inertia wheel. Previously, two different balance controllers are needed depending on whether the bicycle is stationary or dynamic. The switch between the two different controllers may cause the instability of the autonomous bicycle. Our proposed gain scheduled controller can balance the autonomous bicycle in both stationary and dynamic cases. A physical system is built and experiments are carried out to demonstrate the effectiveness of the gain scheduled controller.
Shuai Wang 0007, Leilei Cui 0002, Jie Lai, Xiangyu Chen 0001, Yu Zheng 0001, Zhengyou Zhang, Zhong-Ping Jiang
IROS8
2020 Detection and Isolation of False Data Injection Attacks in Smart Grid via Unknown Input Interval Observer
abstract
This article investigates the detection and isolation of false data injection (FDI) attacks in a smart grid based on the unknown input (UI) interval observer. Recent studies have shown that the FDI attacks can bypass the traditional bad data detection methods by using the vulnerability of state estimation. For this reason, the emergency of FDI attacks brings enormous risk to the security of smart grid. To solve this crucial problem, an UI interval observer-based detection and the isolation scheme against FDI attacks are proposed. We first design the UI interval observers to obtain interval state estimation accurately, based on the constructed physical dynamics grid model. Through the capabilities of the designed UI interval observers, the accurate interval estimation state can be decoupled from unknown disturbances. Based on the characteristics of the interval residuals, a UI interval observer-based global detection algorithm was proposed. Particularly, the interval residual-based detection criteria can address the limitation of the precomputed threshold in traditional bad data detection methods. On this basis, we further consider the detection and isolation of FDI attacks under structure vulnerability. Namely, there exist undetectable FDI attacks in the grid system. Taking the attack undetectability problem into account, a logic judgment matrix-based local detection and isolation algorithm against FDI attacks are developed. Based on the combinations of observable sensor cases, local control centers can further detect and isolate the attack set under structure vulnerability. Finally, the effectiveness of the developed detection and isolation algorithms against FDI attacks is demonstrated on the IEEE 8-bus and IEEE 118-bus smart grid system, respectively.
Xiaoyuan Luo, Zhong-Ping Jiang, Xin-Ping Guan
IEEE Internet Things J.4
2020 Model-Free Robust Optimal Feedback Mechanisms of Biological Motor Control
abstract
Sensorimotor tasks that humans perform are often affected by different sources of uncertainty. Nevertheless, the central nervous system (CNS) can gracefully coordinate our movements. Most learning frameworks rely on the internal model principle, which requires a precise internal representation in the CNS to predict the outcomes of our motor commands. However, learning a perfect internal model in a complex environment over a short period of time is a nontrivial problem. Indeed, achieving proficient motor skills may require years of training for some difficult tasks. Internal models alone may not be adequate to explain the motor adaptation behavior during the early phase of learning. Recent studies investigating the active regulation of motor variability, the presence of suboptimal inference, and model-free learning have challenged some of the traditional viewpoints on the sensorimotor learning mechanism. As a result, it may be necessary to develop a computational framework that can account for these new phenomena. Here, we develop a novel theory of motor learning, based on model-free adaptive optimal control, which can bypass some of the difficulties in existing theories. This new theory is based on our recently developed adaptive dynamic programming (ADP) and robust ADP (RADP) methods and is especially useful for accounting for motor learning behavior when an internal model is inaccurate or unavailable. Our preliminary computational results are in line with experimental observations reported in the literature and can account for some phenomena that are inexplicable using existing models.
Tao Bian, Daniel M. Wolpert, Zhong-Ping Jiang
Neural Comput.3
2020 Active Defense-Based Resilient Sliding Mode Control Under Denial-of-Service Attacks
abstract
This paper investigates the problem of the resilient control for cyber-physical systems (CPSs) in the presence of malicious sensor denial-of-service (DoS) attacks, which result in the loss of state information. The concepts of DoS frequency and DoS duration are introduced to describe the DoS attacks. According to the attack situation, that is, whether the attack is successfully implemented or not, the original physical system is rewritten as a switched version. A resilient sliding mode control scheme is designed to guarantee that the physical process is exponentially stable, which is a foundation of the main results. Then, a zero-sum game is employed to establish an effective mixed defense mechanism. Furthermore, a defense-based resilient sliding mode control scheme is proposed and the desired control performance is achieved. Compared with the existing results, the differences mainly lie in two aspects, that is, one where a switched model is obtained, based on which the average dwell-time like approach is utilized to derive the resilient control scheme, and the other where the zero-sum game in employed to make the attacks satisfy the concepts of DoS frequency and DoS duration. Finally, simulation results are given to demonstrate the effectiveness of the proposed resilient control approach.
Chengwei Wu 0001, Ligang Wu 0001, Jianxing Liu, Zhong-Ping Jiang
IEEE Trans. Inf. Forensics Secur.4
2019 Data-Driven Shared Steering Control of Semi-Autonomous Vehicles
abstract
This paper presents a cooperative/shared framework of the driver and his/her semi-autonomous vehicle in order to achieve desired steering performance. In particular, a copilot controller and the driver together operate and control the vehicle. Exploiting the classical small-gain theory, our proposed shared steering controller is developed independent of the unmeasurable internal states of the human driver, and only relies on his/her steering torque. Furthermore, by adopting data-driven adaptive dynamic programming and an iterative learning scheme, the shared steering controller is studied from the measurable data of the driver and the vehicle. Meanwhile, the accurate knowledge of the driver and the vehicle dynamics is unnecessary, which settles the problem of their potential parametric variations/uncertainties in practice. The effectiveness of the proposed method is validated by rigorous analysis and demonstrated by numerical simulations.
Mengzhe Huang, Weinan Gao, Yebin Wang, Zhong-Ping Jiang
IEEE Trans. Hum. Mach. Syst.4
2019 Distributed Model Predictive Consensus With Self-Triggered Mechanism in General Linear Multiagent Systems
abstract
This paper investigates the consensus problem of general linear discrete-time multiagent systems by using distributed model predictive control (DMPC) with self-triggered mechanism. First, a novel DMPC-based consensus algorithm is proposed, where each agent only needs to obtain its neighbors' predicted state sequences once at each time step. We prove that the resultant DMPC optimization problem is feasible, and the proposed algorithm guarantees the dynamic consensus of agents. Then, to further reduce the communication cost and the energy consumption of control updates, a self-triggered DMPC-based consensus algorithm is proposed with the control input and the triggering interval jointly optimized. Numerical examples including the benchmark problem with platooning vehicles are provided to verify the effectiveness and advantages of the proposed algorithms.
Jingyuan Zhan, Zhong-Ping Jiang, Yebin Wang, Xiang Li 0010
IEEE Trans. Ind. Informatics2
2019 Reinforcement-Learning-Based Cooperative Adaptive Cruise Control of Buses in the Lincoln Tunnel Corridor with Time-Varying Topology
abstract
The exclusive bus lane (XBL) is one of the most popular bus transit systems in the U.S. The Lincoln Tunnel utilizes an XBL through the tunnel in the AM peak period. This paper proposes a novel data-driven cooperative adaptive cruise control (CACC) algorithm that aims to minimize a cost function for connected and autonomous buses along the XBL. Different from existing model-based CACC algorithms, the proposed approach employs the idea of reinforcement learning, which does not rely on accurate knowledge of bus dynamics. Considering a time-varying topology, where each autonomous vehicle can only receive information from preceding vehicles that are within its communication range, a distributed controller is learned real-time by online headway, velocity, and acceleration data collected from the system trajectories. The convergence of the proposed algorithm and the stability of the closed-loop system are rigorously analyzed. The effectiveness of the proposed approach is demonstrated using a well-calibrated Paramics microscopic traffic simulation model of the XBL corridor. The simulation results show that the travel time in the autonomous version of the XBL are close to the present day travel time even when the bus volume is increased by 30%.
Weinan Gao, Jingqin Gao, Kaan Özbay, Zhong-Ping Jiang
IEEE Trans. Intell. Transp. Syst.4
2019 Adaptive Optimal Output Regulation of Time-Delay Systems via Measurement Feedback
abstract
This brief proposes a novel solution to problems related to the measurement feedback adaptive optimal output regulation of discrete-time linear systems with input time-delay. Based on reinforcement learning and adaptive dynamic programming, an approximate optimal control policy is obtained via recursive numerical algorithms using online information. Convergence proofs for the proposed algorithms are given. Notably, the exact knowledge of the plant and the exosystem is not needed. The learned control policy is only a function of retrospective input and measurement output data. Theoretical analysis and an application to a grid-connected inverter show that the proposed methodologies serve as effective tools for solving adaptive and optimal output regulation problems.
Weinan Gao, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.2
2018 An Incremental Multi-view Active Learning Algorithm for PolSAR Data Classification
abstract
The fast and accurate classification of polarimetric synthetic aperture radar (PolSAR) data in dynamically changing environments is an important and challenging task. In this paper, we propose an Incremental Multi-view Passive-Aggressive Active learning algorithm, named IMPAA, for PolSAR data classification. This algorithm can deal with online two-view multi-class categorization problem by exploiting the relationship between the polarimetric-color and texture feature sets of PolSAR data. In addition, the IMPAA algorithm can handle the dynamic large-scale datasets where not only the amount of data but also the number of classes gradually increases. Moreover, this algorithm only queries the class labels of some informative incoming samples to update the classifier based on the disagreement of different views' predictors and a randomized rule. Experiments on real PolSAR data demonstrate that the proposed method can use a smaller fraction of queried labels to achieve low online classification errors compared with previously known methods.
Xiangli Nie, Yongkang Luo 0001, Hong Qiao, Bo Zhang 0006, Zhong-Ping Jiang
ICPR5
2018 Learning-Based Adaptive Optimal Tracking Control of Strict-Feedback Nonlinear Systems
abstract
This paper proposes a novel data-driven control approach to address the problem of adaptive optimal tracking for a class of nonlinear systems taking the strict-feedback form. Adaptive dynamic programming (ADP) and nonlinear output regulation theories are integrated for the first time to compute an adaptive near-optimal tracker without any a priori knowledge of the system dynamics. Fundamentally different from adaptive optimal stabilization problems, the solution to a Hamilton-Jacobi-Bellman (HJB) equation, not necessarily a positive definite function, cannot be approximated through the existing iterative methods. This paper proposes a novel policy iteration technique for solving positive semidefinite HJB equations with rigorous convergence analysis. A two-phase data-driven learning method is developed and implemented online by ADP. The efficacy of the proposed adaptive optimal tracking control methodology is demonstrated via a Van der Pol oscillator with time-varying exogenous signals.
Weinan Gao, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.2
2017 Data-Driven Nonlinear Adaptive Optimal Control of Connected Vehicles
Weinan Gao, Zhong-Ping Jiang
ICONIP (6)2
2017 Adaptive Dynamic Programming for Human Postural Balance Control
Eric Mauro, Tao Bian, Zhong-Ping Jiang
ICONIP (4)3
2017 Data-Driven Adaptive Optimal Control of Connected Vehicles
abstract
In this paper, a data-driven non-model-based approach is proposed for the adaptive optimal control of a class of connected vehicles that is composed of n human-driven vehicles only transmitting motional data and an autonomous vehicle in the tail receiving the broadcasted data from preceding vehicles by wireless vehicle-to-vehicle (V2V) communication devices. Considering the cases of range-limited V2V communication and input saturation, several optimal control problems are formulated to minimize the errors of distance and velocity and to optimize the fuel usage. By employing an adaptive dynamic programming technique, the optimal controllers are obtained without relying on the knowledge of system dynamics. The effectiveness of the proposed approaches is demonstrated via the online learning control of the connected vehicles in Paramics' traffic microsimulation.
Weinan Gao, Zhong-Ping Jiang, Kaan Özbay
IEEE Trans. Intell. Transp. Syst.2
2016 Editor's Note
Zhong-Ping Jiang, Jie Huang 0001
Sci. China Inf. Sci.2
2016 Further results on quantized stabilization of nonlinear cascaded systems with dynamic uncertainties
Tengfei Liu 0004, Zhong-Ping Jiang
Sci. China Inf. Sci.2
2016 A junction-by-junction feedback-based strategy with convergence analysis for dynamic traffic assignment
Tengfei Liu 0004, Zhong-Ping Jiang
Sci. China Inf. Sci.3
2016 Optimal Output-Feedback Control of Unknown Continuous-Time Linear Systems Using Off-policy Reinforcement Learning
abstract
A model-free off-policy reinforcement learning algorithm is developed to learn the optimal output-feedback (OPFB) solution for linear continuous-time systems. The proposed algorithm has the important feature of being applicable to the design of optimal OPFB controllers for both regulation and tracking problems. To provide a unified framework for both optimal regulation and tracking, a discounted performance function is employed and a discounted algebraic Riccati equation (ARE) is derived which gives the solution to the problem. Conditions on the existence of a solution to the discounted ARE are provided and an upper bound for the discount factor is found to assure the stability of the optimal control solution. To develop an optimal OPFB controller, it is first shown that the system state can be constructed using some limited observations on the system output over a period of the history of the system. A Bellman equation is then developed to evaluate a control policy and find an improved policy simultaneously using only some limited observations on the system output. Then, using this Bellman equation, a model-free Off-policy RL-based OPFB controller is developed without requiring the knowledge of the system state or the system dynamics. It is shown that the proposed OPFB method is more powerful than the static OPFB as it is equivalent to a state-feedback control policy. The proposed method is successfully used to solve a regulation and a tracking problem.
Hamidreza Modares, Frank L. Lewis, Zhong-Ping Jiang
IEEE Trans. Cybern.3
2015 Optimal Codesign of Nonlinear Control Systems Based on a Modified Policy Iteration Method
abstract
This brief studies the optimal codesign of nonlinear control systems: simultaneous design of physical plants and related optimal control policies. Nonlinearity of the optimal codesign problem could come from either a nonquadratic cost function or the plant. After formulating the optimal codesign into a nonconvex optimization problem, an iterative scheme is proposed in this brief by adding an additional step of system-equivalence-based policy improvement to the conventional policy iteration. We have proved rigorously that the closed-loop system performance can be improved after each step of the proposed policy iteration scheme, and the convergence to a suboptimal solution is guaranteed. It is also shown that under certain conditions, this additional policy improvement step can be conducted by solving a quadratic programming problem. The linear version of the proposed methodology is addressed in the context of linear quadratic regulator. Finally, the effectiveness of the proposed methodology is illustrated through the optimal codesign of a load-positioning system.
Yu Jiang 0003, Yebin Wang, Scott A. Bortoff, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.4
2015 H∞ Tracking Control of Completely Unknown Continuous-Time Systems via Off-Policy Reinforcement Learning
abstract
This paper deals with the design of an H ∞ tracking controller for nonlinear continuous-time systems with completely unknown dynamics. A general bounded L2 -gain tracking problem with a discounted performance function is introduced for the H ∞ tracking. A tracking Hamilton-Jacobi-Isaac (HJI) equation is then developed that gives a Nash equilibrium solution to the associated min-max optimization problem. A rigorous analysis of bounded L2 -gain and stability of the control solution obtained by solving the tracking HJI equation is provided. An upper-bound is found for the discount factor to assure local asymptotic stability of the tracking error dynamics. An off-policy reinforcement learning algorithm is used to learn the solution to the tracking HJI equation online without requiring any knowledge of the system dynamics. Convergence of the proposed algorithm to the solution to the tracking HJI equation is shown. Simulation examples are provided to verify the effectiveness of the proposed method.
Hamidreza Modares, Frank L. Lewis, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.3
2015 Multiple Actor-Critic Structures for Continuous-Time Optimal Control Using Input-Output Data
abstract
In industrial process control, there may be multiple performance objectives, depending on salient features of the input-output data. Aiming at this situation, this paper proposes multiple actor-critic structures to obtain the optimal control via input-output data for unknown nonlinear systems. The shunting inhibitory artificial neural network (SIANN) is used to classify the input-output data into one of several categories. Different performance measure functions may be defined for disparate categories. The approximate dynamic programming algorithm, which contains model module, critic network, and action network, is used to establish the optimal control in each category. A recurrent neural network (RNN) model is used to reconstruct the unknown system dynamics using input-output data. NNs are used to approximate the critic and action networks, respectively. It is proven that the model error and the closed unknown system are uniformly ultimately bounded. Simulation results demonstrate the performance of the proposed optimal control scheme for the unknown nonlinear system.
Ruizhuo Song, Frank L. Lewis, Qinglai Wei, Huaguang Zhang, Zhong-Ping Jiang, Daniel S. Levine 0001
IEEE Trans. Neural Networks Learn. Syst.5
2013 Movement Duration, Fitts's Law, and an Infinite-Horizon Optimal Feedback Control Model for Biological Motor Systems
abstract
Optimization models explain many aspects of biological goal-directed movements. However, most such models use a finite-horizon formulation, which requires a prefixed movement duration to define a cost function and solve the optimization problem. To predict movement duration, these models have to be run multiple times with different prefixed durations until an appropriate duration is found by trial and error. The constrained minimum time model directly predicts movement duration; however, it does not consider sensory feedback and is thus applicable only to open-loop movements. To address these problems, we analyzed and simulated an infinite-horizon optimal feedback control model, with linear plants, that contains both control-dependent and control-independent noise and optimizes steady-state accuracy and energetic costs per unit time. The model applies the steady-state estimator and controller continuously to guide an effector to, and keep it at, target position. As such, it integrates movement control and posture maintenance without artificially dividing them with a precise, prefixed time boundary. Movement pace is determined by the model parameters, and the duration is an emergent property with trial-to-trial variability. By considering the mean duration, we derived both the log and power forms of Fitts's law as different approximations of the model. Moreover, the model reproduces typically observed velocity profiles and occasional transient overshoots. For unbiased sensory feedback, the effector reaches the target without bias, in contrast to finite-horizon models that systematically undershoot target when energetic cost is considered. Finally, the model does not involve backward and forward sweeps in time, its stability is easily checked, and the same solution applies to movements of different initial conditions and distances. We argue that biological systems could use steady-state solutions as default control mechanisms and might seek additional optimization of transient costs when justified or demanded by task or context.
Ning Qian, Yu Jiang 0003, Zhong-Ping Jiang, Pietro Mazzoni
Neural Comput.3
2013 Robust Adaptive Dynamic Programming With an Application to Power Systems
abstract
This brief presents a novel framework of robust adaptive dynamic programming (robust-ADP) aimed at computing globally stabilizing and suboptimal control policies in the presence of dynamic uncertainties. A key strategy is to integrate ADP theory with techniques in modern nonlinear control with a unique objective of filling up a gap in the past literature of ADP without taking into account dynamic uncertainties. Neither the system dynamics nor the system order are required to be precisely known. As an illustrative example, the computational algorithm is applied to the controller design of a two-machine power system.
Yu Jiang 0003, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.2
2011 Approximate Dynamic Programming for Optimal Stationary Control With Control-Dependent Noise
abstract
This brief studies the stochastic optimal control problem via reinforcement learning and approximate/adaptive dynamic programming (ADP). A policy iteration algorithm is derived in the presence of both additive and multiplicative noise using Itô calculus. The expectation of the approximated cost matrix is guaranteed to converge to the solution of some algebraic Riccati equation that gives rise to the optimal cost value. Moreover, the covariance of the approximated cost matrix can be reduced by increasing the length of time interval between two consecutive iterations. Finally, a numerical example is given to illustrate the efficiency of the proposed ADP methodology.
Yu Jiang 0003, Zhong-Ping Jiang
IEEE Trans. Neural Networks2
2009 Stability and control of nonlinear systems described by retarded functional equations: a review of recent results
Iasson Karafyllis, Zhong-Ping Jiang
Sci. China Ser. F Inf. Sci.2
2009 Velocity-Scheduling Control for a Unicycle Mobile Robot: Theory and Experiments
abstract
Improvement over classical dynamic feedback linearization for a unicycle mobile robots is proposed. Compared to classical extension, the technique uses a higher-dimensional state extension, which allows rejecting a constant disturbance on the robot rotational axis. The proposed dynamic extension acts as a velocity scheduler that specifies, at each time instant, the ideal translational velocity that the robot should have. By using a higher-order extension, both the magnitude and the orientation of the velocity vector can be generated, which introduces robustness in the control scheme. Stability for both asymptotic convergence to a point and trajectory tracking is proven. The theoretical results are illustrated first in simulation, and then experimentally on the autonomous mobile robotFouzyIII.
Davide Buccieri, Damien Perritaz, Philippe Müllhaupt, Zhong-Ping Jiang, Dominique Bonvin
IEEE Trans. Robotics4
2008 Distributed power control and random access for spectrum sharing with QoS constraint
Bo Yang 0006, Yanyan Shen, Gang Feng 0001, Chengnian Long, Zhong-Ping Jiang, Xin-Ping Guan
Comput. Commun.5
2007 Nonlinear Bistable Stochastic Resonance Filters for Image Processing
abstract
Nonlinear bistable double-well stochastic resonance systems have been successfully used for one-dimensional signal processing, based on the concept of parameter-tuning stochastic resonance. This paper investigates the applications of parameter-tuning stochastic resonance in image processing. First, a two-dimensional stochastic resonance system is introduced as a nonlinear filter for image processing. The equation satisfied by the dynamic probability density function of the images processed by this stochastic resonance filter and its solutions are then discussed. Finally, this nonlinear filter is used to process a black-white image corrupted by additive white Gaussian noise to reveal the possibility to extend the concept of parameter-tuning stochastic resonance to two-dimensional cases. This provides an innovative approach for image processing.
Bohou Xu, Zhong-Ping Jiang, Xingxing Wu, Daniel W. Repperger
ICASSP (1)2
2006 Modeling and performance analysis of ad hoc broadcasting schemes
Hao Zhang 0012, Zhong-Ping Jiang
Perform. Evaluation2
2004 Recent developments in decentralized nonlinear control
abstract
In this paper, we review some of the recent results in decentralized control for large-scale nonlinear systems. Problems of decentralized stabilization, adaptive tracking, disturbance attenuation with internal stability are discussed, with a special emphasis on output feedback control. A practical example of double inverted pendulums is given to demonstrate the novelty of our approaches.
Zhong-Ping Jiang
ICARCV1
2004 On output feedback stabilization of uncertain chained systems
abstract
This paper deals with chained form systems with strongly nonlinear disturbances and drift terms. The objective is to design robust nonlinear output feedback laws such that the closed-loop systems are globally exponentially stable. The systematic strategy combines the input-state-scaling technique with the so-called backstepping procedure.
Zairong Xi, Gang Feng 0001, Zhong-Ping Jiang, Daizhan Cheng
ICARCV3
2004 Analysis of two ad hoc broadcasting protocols
abstract
Broadcasting techniques are employed in route discovery procedure in on-demand routing protocols such as DSR and AODV. The traditional broadcasting method 'flooding' is not suitable for large heavy-loaded ad hoc networks. Numerous protocols have been proposed to save wireless network capacity and avoid packets collisions. Ns-2 simulations have been used widely to compare the performance of these protocols in various situations. Sometimes it is difficult to measure the efficiency of the protocols based only on the simulation results because usually they cannot cover all the possible situations especially for random networks and random schemes. With rigorous theoretic analysis, we may achieve statistic results which can be used to choose proper parameters (such as threshold values in some schemes) for optimal performance. In this paper, we analyze two popular ad hoc broadcasting protocols in grid networks. The results have shown the relation between protocol efficiency and selection of the parameters.
Hao Zhang 0012, Zhong-Ping Jiang
WCNC2
2004 A global output-feedback controller for simultaneous tracking and stabilization of unicycle-type mobile robots
abstract
We present a time-varying global output-feedback controller that solves both tracking and stabilization for unicycle-type mobile robots simultaneously at the torque level. The controller synthesis is based on a coordinate transformation, Lyapunov's direct method, and backstepping technique. Simulations demonstrate the result.
Khac Duc Do 0001, Zhong-Ping Jiang, Jie Pan 0001
IEEE Trans. Robotics2
2000 Stable neural controller design for unknown nonlinear systems using backstepping
abstract
Despite the vast development of neural controllers in the literature, their stability properties are usually addressed inadequately. With most neural control schemes, the choices of neural-network structure, initial weights, and training speed are often nonsystematic, due to the lack of understanding of the stability behavior of the closed-loop system. In this paper, we propose, from an adaptive control perspective, a neural controller for a class of unknown, minimum phase, feedback linearizable nonlinear system with known relative degree. The control scheme is based on the backstepping design technique in conjunction with a linearly parameterized neural-network structure. The resulting controller, however, moves the complex mechanics involved in a typical backstepping design from offline to online. With appropriate choice of the network size and neural basis functions, the same controller can be trained online to control different nonlinear plants with the same relative degree, with semiglobal stability as shown by simple Lyapunov analysis. Meanwhile, the controller also preserves some of the performance properties of the standard backstepping controllers. Simulation results are shown to demonstrate these properties and to compare the neural controller with a standard backstepping controller.
Youping Zhang, Pei-Yaun Peng, Zhong-Ping Jiang
IEEE Trans. Neural Networks Learn. Syst.3