EDBT 2026 Demo / reviewers in the wild / expert
Ke Wang 0037
dblp:181/2613-37
· DBLP profile ↗
23ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-8306-1663ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TNLight: A triple-network enhanced double Q-learning method for hierarchical and cooperative traffic signal control
Chaoxu Mu, Dayu Hou, Ke Wang 0037, Song Zhu, Ge Guo 0001 |
Inf. Sci. | 3 |
| 2026 | Six-DoF Coupled Dynamics Modeling and Intelligent Vibration Suppression of CMG With the Flexible Vibration IsolatorabstractThe control moment gyroscope (CMG) is widely used for high-precision attitude control in aerospace applications due to its excellent torque output and energy efficiency. However, the micro-vibrations generated during CMG operation can propagate through flexible vibration isolation platforms, which may affect the performance of sensitive spacecraft payloads. To address this critical issue, this article proposes a systematic algorithm that combines coupled dynamic modeling with intelligent control. First, a six-degree-of-freedom (six-DoF) dynamic model accounting for the flexibility of the vibration isolator is developed, which combines the flywheel rotor, gimbal system, and the vibration isolator with a flexible platform to characterize their coupled dynamics. Based on this model, an active-passive isolation control algorithm is designed, which combines feedforward disturbance compensation with a policy iteration algorithm based on adaptive dynamic programming (ADP). The feedforward component enhanced by neural networks effectively approximates and cancels out the interference caused by CMG in real time, while ADP optimization ensures adaptive and optimal vibration suppression. The simulation results demonstrate the performance of the proposed algorithm, confirming a significant reduction in vibration transmission and a marked improvement in isolation efficiency. Jiankun Yang, Ke Wang 0037, Chaoxu Mu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Integrated Task and Motion Planner Using Hierarchical Reinforcement Learning for Multi-Robot CollaborationabstractThe task planning and motion planning problems in multi task and multi robot collaboration are usually considered as independent parts to be solved. However, in long-horizon, multi-step collaborative scenarios, this independent solution approach is difficult to effectively handle the coupling problem between task planning and motion planning, such as the inability to simultaneously consider robot motion collisions during task allocation. To address these limitations, we develop a unified hierarchical reinforcement learning framework that enables agents to learn effective policies in multi task and multi robot collaborative motion planning, supplemented by two techniques: 1) using a shared graph attention network and distributed police networks to assign target tasks to each robot at a high level, and 2) training decentralized motion planning policies for each robot at a low level to control the workspace state and target end actuator posture. The high-level and low-level policies are trained in parallel and stabilized through expert dataset guidance. To verify the effectiveness of the method, we conduct experiments on a general object placement platform and an engine hydraulic column assembly platform. The results indicate that this method can more efficiently complete multi task and multi robot collaborative work. Chaoxu Mu, Ke Wang 0037, Lei Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Uncertainty-Aware Model-Based Multi-Agent Deep Reinforcement Learning for Robust Active Voltage ControlabstractThe large-scale injection of new energy systems into active distribution networks (ADNs) has caused voltage violations, challenging the stable operation of power grids. Recently, deep reinforcement learning (DRL) has emerged with great advantages in replacing traditional optimization methods for voltage regulation in ADNs. However, existing DRL studies face sampling inefficiency and overlook the robustness issue resulting from the uncertainties brought by renewable energy systems in ADNs, which seriously affects the application of DRL in the real world. In this paper, a novel uncertainty-aware model-based multi-agent deep reinforcement learning (MADRL) framework is proposed for robust active voltage control (AVC). First, a probabilistic ensemble of neural networks with different initializations is designed for uncertainties in environment model learning, and a hybrid data augmentation method is proposed to improve the learning efficiency and final performance of MADRL. Then, a multi-agent distributional soft actor-critic (MADSAC) framework is developed for robust voltage regulation by tackling various uncertainties in the ADN environment. Simulations are performed on the IEEE 33-bus distribution network and IEEE 141-bus distribution network to validate that the proposed model-based MADSAC algorithm can significantly improve sampling efficiency, robustness and performance in AVC. Zhaoyang Liu 0005, Ke Wang 0037, Chenyi Si, Chaoxu Mu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Cooperative Control of Heterogeneous Connected Vehicle Platoons: A Dynamic Event-Triggered Reinforcement Learning Approach
Ke Wang 0037, Yichun Liu, Chaoxu Mu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | VLIN-RL: A Unified Vision-Language Interpreter and Reinforcement Learning Motion Planner Framework for Robot Dynamic TasksabstractRecently, with the development of Large Language Models (LLMs), Embodied AI represented by Vision-Language-Action Models (VLAs) has played a significant role in realizing the natural language interaction between humans and robots. Current VLA models can process and understand visual information and language instructions, while guiding robots to complete interactive tasks with the environment based on human language instructions. However, when tackling with the real-time and dynamic tasks, VLAs have poor robustness and real-time planning and adjustment ability against changes in target objects, instructions, and environments. To handle these limitations, we propose VLIN-RL, a unified framework that consists of the Vision-Language Interpreter (VLIN) that owns excellent vision language information understanding and advanced task planning abilities and reinforcement learning (RL)-based motion planner with enhanced flexibility and broader applicability. If the environmental state changes during task execution, the RL planning module in VLIN-RL will directly make dynamic adjustments at the subtask level based on visual feedback to achieve the task goals, without the need for time-consuming reprocessing by VLIN. Experiments demonstrate that our model can complete multi-robot manipulation tasks more efficiently and stably. Finally, our work is verified by the pick-grasp tasks and real manipulators experiments. The test video is available at https://github.com/jzwsoulferryman/VLIN-RL.git. Zewu Jiang, Ke Wang 0037, Chenyi Si |
IROS | 3 |
| 2025 | Online value iteration-driven intelligent control for CMG gimbal servo system under multi-source disturbances
Chaoxu Mu, Jiankun Yang, Ke Wang 0037 |
Expert Syst. Appl. | 3 |
| 2025 | Safety-Critical Path Planning for Obstacle Avoidance Based on Reinforcement Learning and Control Barrier FunctionsabstractThis article presents a safety-critical control framework for navigation in complex environments with numerous obstacles. An online robust path planning scheme is developed by integrating reinforcement learning (RL) with control barrier functions (CBFs). First, a disturbance observer is designed to estimate the unknown disturbance along with a derived upper bound of the estimation error. Then, a nominal controller is designed using RL, where a critic neural network (NN) structure is established by using state-following (StaF) kernel function. Additionally, by employing a state extrapolation technique, the learning process leverages both real-time and simulated experience data. To ensure safety, obstacle avoidance is formulated as a forward invariance problem of safe sets defined by CBFs. Subsequently, the CBF-based safety-critical constraints are integrated into a quadratic programming (QP) framework to modify the nominal controller. Furthermore, these CBFs are incorporated into a composite CBF using smooth approximation, enabling efficient constraint consolidation. Then, an explicit safe control policy is proposed that guarantees collision-free path planning. Finally, the effectiveness of the proposed scheme is demonstrated through numerical simulations, and comparative results show the advantages over the existing methods in motion trajectory. Ke Wang 0037, Chaoxu Mu, Tie Qiu 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Resilience-Based Output Formation-Containment Control of Nonlinear MASs Against DOS AttacksabstractThis paper addresses the problem of distributed formation-containment tracking control for uncertain multi-agent systems (MASs) with completely unknown system nonlinearities, denial-of-service (DOS) attacks, and switching communication topologies. To enhance the system robustness, neural networks (NNs) are utilized to identify the unknown nonlinear terms. Additionally, a novel distributed observer is designed to reconstruct the external unmeasured attack dynamics. To handle the switching topologies of MASs, a mechanism using piecewise continuous functions is proposed to counteract the unexpected controller actions during switching time instants. Furthermore, the clever design of barrier Lyapunov functions aids in achieving the required predefined performance. A fractional power nonlinear filter is introduced to tackle the problem of computational complexity. By applying the local neighborhood states information and the lyapunov stability theory, the presented control method evaluates the stability of the MASs and shows that the designed controller not only enables the system output to track a formation trajectory in the presence of external attacks but also converges the consensus errors into a predefined set. Finally, simulation results are provided to validate the effectiveness of the proposed control strategy. Chaoxu Mu, Ke Wang 0037, Song Zhu, Zhijia Zhao 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Fuzzy Dynamic Event-Triggered Control for Constrained Stochastic Game SystemsabstractIn this article we focus on solving the optimal control problem for stochastic game systems while considering multiple constraints. Initially, applying the asymmetric time-varying mapping functions, the constrained nonzero-sum stochastic differential games are transformed into the ones in unconstrained forms. Moreover, the actor-critic architecture is designed to achieve the Nash equilibrium, the critic fuzzy logic systems (FLSs) and actor FLSs are utilized to approximate the optimal value functions and the control policies, respectively. Subsequently, a dynamic event-triggered mechanism is devised and an additional dynamic variable is defined to characterize the past triggering information, which provides a larger inter-event time. Mathematical analysis reveals the stability of closed-loop system, the weight convergence and the characteristics of dynamic event-triggered mechanism. Finally, simulation results prove the effectiveness of the proposed method. Chaoxu Mu, Chenyi Si, Ke Wang 0037, Qing Wang 0010, Jinpeng Yu 0001, Jinshan Bian |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | Data-Model Hybrid-Driven Safe Reinforcement Learning for Adaptive Avoidance Control Against Unsafe Moving ZonesabstractWith the gradual application of reinforcement learning (RL), safety has emerged as a paramount concern. This article presents a novel data-model hybrid-driven safe RL (SRL) scheme to address the challenge of avoidance control in the operation domain containing multiple moving unsafe zones. First, the avoidance problem is transformed into the optimal control problem of an augmented system by encoding a barrier function (BF) term into the cost function. Then, using the idea of integral RL (IRL), an adaptive learning algorithm is proposed for generating safe control policies, in which the actor-critic neural network (NN) structure is established with the aid of state-following (StaF) kernel function. The policy iteration process is executed by this structure; specifically, the critic network undergoes gradient-descent adaptation, while the actor network employs gradient projection updating. Particularly, via a state extrapolation technique, both real-time experience and simulated experience are utilized in the learning process. Next, closed-loop stability and weight convergence are theoretically substantiated. Finally, the effectiveness of the proposed scheme is demonstrated on a single integrator system, a nonlinear numerical system, and a unicycle kinematic system; besides, its advantages over the existing control methods are illustrated by comparisons. Ke Wang 0037, Chaoxu Mu, Anguo Zhang, Changyin Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Dynamic Event-Triggered Model-Free Reinforcement Learning for Cooperative Control of Multiagent SystemsabstractIn this article, a novel model-free dynamic event-triggered adaptive learning control scheme is developed for continuous-time linear multiagent systems. This control scheme is different from model-based control scheme in the sense that prior knowledge of the system's model is not required. To further reduce transmission data, an event-triggered control policy based on static event-triggered mechanism (SETM) and dynamic event-triggered mechanism (DETM) is proposed. Compared to SETM, DETM may significantly produce larger average event intervals and maintain control performance. In addition, based on off-policy integral reinforcement learning, an adaptive iteration method is proposed with convergence proof. Numerical tests on both linear and nonlinear multiagent systems are conducted to demonstrate that the proposed scheme can guarantee learning performance and larger triggering intervals. Finally, the learning control scheme is tested on the multiarea power system, which can illustrate the reliability and practicality of this method. Specifically, the load frequency control problem of the multiarea power system is studied using three control schemes, revealing that DETM can achieve a better frequency response at the lowest information transmission rate and ensure the overall quality and reliability of the power system. Ke Wang 0037, Zhuo Tang, Chaoxu Mu |
IEEE Trans. Reliab. | 1 |
| 2024 | Q-learning based tracking control with novel finite-horizon performance index
Ke Wang 0037, Chaoxu Mu, Haoxian Shi |
Inf. Sci. | 2 |
| 2024 | Safe Reinforcement Learning and Adaptive Optimal Control With Applications to Obstacle Avoidance ProblemabstractThis paper presents a novel composite obstacle avoidance control method to generate safe motion trajectories for autonomous systems in an adaptive manner. First, system safety is described using forward invariance, and the barrier function is encoded into the cost function such that the obstacle avoidance problem can be characterized by an infinite-horizon optimal control problem. Next, a safe reinforcement learning framework is proposed by combining model-based policy iteration and state-following-based approximation. Upon real-time data and extrapolated experience data, this learning design is implemented through the actor-critic structure, in which critic networks are tuned by gradient-descent adaption and actor networks produce adaptive control policies via gradient projection. Then, system stability and weight convergence are theoretically analyzed using Lyapunov method. Finally, the proposed learning-based controller is demonstrated on a two-dimensional single integrator system and a nonlinear unicycle kinematic system. Simulation results reveal that the system or agent can smoothly reach the target point while keeping a safe distance from each obstacle; at the same time, other three avoidance control methods are used to provide side-by-side comparisons and to verify some claimed advantages of the present method.Note to Practitioners—This paper is motivated by the obstacle avoidance problem of real-time navigation of an agent to the target point, which applies to practical autonomous systems such as vehicles and robots. Pre-generative methods and reactive methods have been widely employed to generate safe motion trajectories in the obstacle environment. However, these methods cannot strike a good balance between safety and optimality. In this paper, the obstacle avoidance problem is formulated in the sense of optimal control, and a safe reinforcement learning method is designed to generate safe motion trajectories. This method combines the advantages of model-based policy iteration and state-following-based approximation, in which the former ensures regional optimality while the latter ensures local safety. Based on the proposed adaptive tuning laws, engineers are able to design learning-based avoidance controllers in the environment with static obstacles. In future research, we will address the dynamic avoidance problem against moving obstacles. Ke Wang 0037, Chaoxu Mu, Zhen Ni, Derong Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Fuzzy-Based Optimal Control for Stochastic Nonlinear Systems With Constrained Inputs via Dynamic Event-TriggeringabstractA dynamic event-based optimal learning scheme is provided in the paper for nonlinear systems subject to constrained inputs and stochastic disturbances. An actor-critic structure is constructed to learn the stochastic optimal solution, which includes critic fuzzy logic system (FLS) and actor FLS. The experience replay technique and gradient-descent adaption method are used to periodically tune the critic FLS, which can approximate the optimal cost function. Based on the static event-triggered control mechanism (ETCM), a dynamic ETCM is designed, which can incorporate past triggering information and generate a longer inter-event time. The actor FLS is updated at aperiodic jumping points, which can approximate the optimal control policy. The combination of dynamic ETCM and learning structure ensures the stochastic stability of closed-loop system. The efficiency of the controller is illustrated on a numerical example and a manipulator system. Chenyi Si, Chaoxu Mu, Ke Wang 0037, Song Zhu, Jinpeng Yu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Learning-Based Control With Decentralized Dynamic Event-Triggering for Vehicle SystemsabstractThe optimal control of multi-input system can be described by a multiplayer nonzero-sum differential game. This article theoretically presents an event-based adaptive learning scheme to approximate the Nash equilibrium, and practically addresses the cruise control problem for Caltech vehicle systems. This design is deployed in two aspects. On one hand, the reinforcement learning is implemented through critic neural network architecture and recalling stored experience data. On the other hand, in view of that each player’s preference is different, the decentralized triggering manner is considered to reduce communication. Based on the continuous state, the local sampled state is defined for each player, and a static triggering mechanism is formulated first. The decentralized dynamic triggering is then promoted by designing an auxiliary variable whose dynamics are constructed using static triggering information. Next, the proposed learning scheme is examined on a four-player numerical system. Finally, the learning-based controller is tested on a single-vehicle system under different tracking commands, and then, it is extended to multivehicle systems to realize cooperative optimization by introducing a novel game-in-game structure. Ke Wang 0037, Chaoxu Mu |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Data-based decentralized learning scheme for nonlinear systems with mismatched interconnections
Chaoxu Mu, Jiangwen Peng, Ke Wang 0037 |
Neurocomputing | 4 |
| 2022 | Dynamic Event-Triggering Neural Learning Control for Partially Unknown Nonlinear SystemsabstractThis article presents an event-sampled integral reinforcement learning algorithm for partially unknown nonlinear systems using a novel dynamic event-triggering strategy. This is a novel attempt to introduce the dynamic triggering into the adaptive learning process. The core of this algorithm is the policy iteration technique, which is implemented by two neural networks. A critic network is periodically tuned using the integral reinforcement signal, and an actor network adopts the event-based communication to update the control policy only at triggering instants. For overcoming the deficiency of static triggering, a dynamic triggering rule is proposed to determine the occurrence of events, in which an internal dynamic variable characterized by a first-order filter is defined. Theoretical results indicate that the impulsive system driven by events is asymptotically stable, the network weight is convergent, and the Zeno behavior is successfully avoided. Finally, three examples are provided to demonstrate that the proposed dynamic triggering algorithm can reduce samples and transmissions even more, with guaranteed learning performance. Chaoxu Mu, Ke Wang 0037, Tie Qiu 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Adaptive Learning and Sampled-Control for Nonlinear Game Systems Using Dynamic Event-Triggering StrategyabstractStatic event-triggering-based control problems have been investigated when implementing adaptive dynamic programming algorithms. The related triggering rules are only current state-dependent without considering previous values. This motivates our improvements. This article aims to provide an explicit formulation for dynamic event-triggering that guarantees asymptotic stability of the event-sampled nonzero-sum differential game system and desirable approximation of critic neural networks. This article first deduces the static triggering rule by processing the coupling terms of Hamilton-Jacobi equations, and then, Zeno-free behavior is realized by devising an exponential term. Subsequently, a novel dynamic-triggering rule is devised into the adaptive learning stage by defining a dynamic variable, which is mathematically characterized by a first-order filter. Moreover, mathematical proofs illustrate the system stability and the weight convergence. Theoretical analysis reveals the characteristics of dynamic rule and its relations with the static rules. Finally, a numerical example is presented to substantiate the established claims. The comparative simulation results confirm that both static and dynamic strategies can reduce the communication that arises in the control loops, while the latter undertakes less communication burden due to fewer triggered events. Chaoxu Mu, Ke Wang 0037, Zhen Ni |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Learning Control Supported by Dynamic Event Communication Applying to Industrial SystemsabstractFor the practical control system, the controller is normally implemented on a digital platform with a time-triggered scheme. This scheme maybe produces redundant control and resources wasting, and hence, an event-triggered scheme is gradually favored. In this article, the robust learning control scheme is proposed aiming at a class of disturbed control systems, in which the system information is processed by a novel dynamic event communication. First, the robust optimal control problem with external disturbances is redescribed as a zero-sum differential game, and with integral reinforcement learning, a model-independent weight tuning law is devised for a critic neural network. Then, in order to further reduce the computational burden, an additional dynamic variable is put forward to incorporate the past triggering information. The application of a single-link joint arm system demonstrates that the proposed scheme can guarantee learning performance and robust control effect, along with larger triggering intervals. Finally, the load frequency control problem of single-area power system is studied. On one hand, the comparative results of five control schemes reveal that the dynamic event scheme can achieve the better frequency response at the lowest information transmission rate. On the other hand, the advantages of the proposed method are illustrated by comparing with other three event-triggered schemes. Chaoxu Mu, Ke Wang 0037, Changyin Sun 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Policy-Iteration-Based Learning for Nonlinear Player Game Systems With Constrained InputsabstractThis article investigates the optimal control problem for nonlinear nonzero-sum differential game in the environment of no initial admissible policies while considering the control constraint. An adaptive learning algorithm is thus developed based on policy iteration technique to approximately obtain the Nash equilibrium using real-time data. A two-player continuous-time system is used to present this approximate mechanism, which is implemented as a critic-actor architecture for every player. The constraint is incorporated into this optimization by introducing the nonquadratic value function, and the associated constrained Hamilton-Jacobi equation is derived. The critic neural network (NN) and actor NN are utilized to learn the value function and the optimal control policy, respectively, in the light of novel weight tuning laws. In order to tackle the stability during the learning phase, two stable operators are designed for two actors. The proposed algorithm is proved to be convergent as a Newton's iteration, and the stability of this closed-loop system is also ensured by Lyapunov analysis. Finally, two simulation examples demonstrate the effectiveness of the proposed learning scheme by considering different constraint scenes. Chaoxu Mu, Ke Wang 0037, Changyin Sun 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Hierarchical optimal control for input-affine nonlinear systems through the formulation of Stackelberg game
Chaoxu Mu, Ke Wang 0037, Dongbin Zhao |
Inf. Sci. | 2 |
| 2020 | Cooperative Differential Game-Based Optimal Control and Its Application to Power SystemsabstractDifferential games have been extensively applied to optimal control problems. Nash equilibrium captures the tradeoff among players' policies when every player independently tries to minimize a predefined index. When considering potential cooperation, Pareto equilibrium plays an important role in cooperative differential games. This article studies the cooperative control of multiplayer systems on the quadratic infinite horizon. First, by defining a joint cost function using a parameter set, a cooperative differential game is reformulated as a general optimal control problem, where all players form a grand coalition. Then, the joint cost function is approximated by a critic neural network, and for the first time, a novel adaptive dynamic programming algorithm with two learning stages is proposed to determine the parameter selection and then obtain Pareto optimal solutions. A numerical example demonstrates that this algorithm can achieve optimal policies and Pareto frontier. As for its application, the cooperative control of a two-area interconnected power system is investigated, where the primary frequency control and secondary frequency control are regarded as two players. Simulation results indicate that the proposed scheme can obtain binding cooperation agreements, such that cooperative control scheme can get better overall performance compared to Nash control method and another three control methods. Chaoxu Mu, Ke Wang 0037, Zhen Ni, Changyin Sun 0001 |
IEEE Trans. Ind. Informatics | 2 |