Qingrui Zhang

dblp:35/8076 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep Reinforcement Learning-Driven Parameter Tuning for Adaptive Control Systems in Hypersonic Flight Vehicle
abstract
Hypersonic flight vehicle faces critical challenges of control from highly nonlinear and time-varying uncertainties, which impose stringent requirements for real-time parameter adaptation under safety constraints. This paper proposes a reinforcement learning-based adaptive tracking control algorithm to address these issues. The crucial contributions of our design, as opposed to the state-of-the-art approaches, lie in three aspects: (a) a hybrid design of model-based control and reinforcement learning to alleviate the safety, stability and generalization issues of learning-based methods specifically for the demanding hypersonic flight environment; (b) the establishment of a reinforcement learning-based optimization framework that dynamically adjusts control parameters in a real-time optimal fashion to improve the tracking performance under dynamic uncertainties and flight regime transitions, which is substantially different from most conventional methods with constant parameters; (c) the theoretical analysis of both the closed-loop stability of the adaptive control and the convergence performance of the learning algorithm, which distinguishes our design from most existing reinforcement learning-based methods that have no stability or convergence guarantee and is particularly critical for safety-critical hypersonic flight vehicle applications. Numerical simulations show that the proposed method achieves a reduction in the integral of tracking error of 8.31% under model perturbations and 34.3% under changing reference trajectories, compared to the baseline method, while maintaining comparable control energy consumption.
Maolong Lv, Qingrui Zhang, Zehong Dong, Zongyu Zuo
IEEE Trans Autom. Sci. Eng.2
2025 A Recursive Total Least Squares Solution for Bearing-Only Target Motion Analysis and Circumnavigation
abstract
Bearing-only Target Motion Analysis (TMA) is a promising technique for passive tracking in various applications as a bearing angle is easy to measure. Despite its advantages, bearing-only TMA is challenging due to the nonlinearity of the bearing measurement model and the lack of range information, which impairs observability and estimator convergence. This paper addresses these issues by proposing a Recursive Total Least Squares (RTLS) method for online target localization and tracking using mobile observers. The RTLS approach, inspired by previous results on Total Least Squares (TLS), mitigates biases in position estimation and improves computational efficiency compared to pseudo-linear Kalman filter (PLKF) methods. Additionally, we propose a circumnavigation controller to enhance system observability and estimator convergence by guiding the mobile observer in orbit around the target. Extensive simulations and experiments are performed to demonstrate the effectiveness and robustness of the proposed method. The proposed algorithm is also compared with the state-of-the-art approaches, which confirms its superior performance in terms of both accuracy and stability.
Xueming Liu, Zhoujingzi Qiu, Tianjiang Hu, Qingrui Zhang
IROS5
2025 HAC-LOCO: Learning Hierarchical Active Compliance Control for Quadruped Locomotion under Continuous External Disturbances
abstract
Despite recent remarkable achievements in quadruped control, it remains challenging to ensure robust and compliant locomotion in the presence of unforeseen external disturbances. Existing methods prioritize locomotion robustness over compliance, often leading to stiff, high-frequency motions, and energy inefficiency. This paper, therefore, presents a two-stage hierarchical learning framework that can learn to take active reactions to external force disturbances based on force estimation. In the first stage, a velocity-tracking policy is trained alongside an auto-encoder to distill historical proprioceptive features. A neural network-based estimator is learned through supervised learning, which estimates body velocity and external forces based on proprioceptive measurements. In the second stage, a compliance action module, inspired by impedance control, is learned based on the pre-trained encoder and policy. This module is employed to actively adjust velocity commands in response to external forces based on real-time force estimates. With the compliance action module, a quadruped robot can robustly handle minor disturbances while appropriately yielding to significant forces, thus striking a balance between robustness and compliance. Simulations and real-world experiments have demonstrated that our method has superior performance in terms of robustness, energy efficiency, and safety. Experiment comparison shows that our method outperforms the state-of-the-art RL-based locomotion controllers. Ablation studies are given to show the critical roles of the compliance action module.
Qingrui Zhang
IROS4
2025 EASpace: Enhanced Action Space for Policy Transfer
abstract
Formulating expert policies as macro actions promises to alleviate the long-horizon issue via structured exploration and efficient credit assignment. However, traditional option-based multipolicy transfer methods suffer from inefficient exploration of macro action's length and insufficient exploitation of useful long-duration macro actions. In this article, a novel algorithm named enhanced action space (EASpace) is proposed, which formulates macro actions in an alternative form to accelerate the learning process using multiple available suboptimal expert policies. Specifically, EASpace formulates each expert policy into multiple macro actions with different execution times. All the macro actions are then integrated into the primitive action space directly. An intrinsic reward, which is proportional to the execution time of macro actions, is introduced to encourage the exploitation of useful macro actions. The corresponding learning rule that is similar to intraoption Q-learning is employed to improve the data efficiency. Theoretical analysis is presented to show the convergence of the proposed learning rule. The efficiency of EASpace is illustrated by a grid-based game and a multiagent pursuit problem. The proposed algorithm is also implemented in physical systems to validate its effectiveness.
Qingrui Zhang, Bo Zhu 0005, Tianjiang Hu
IEEE Trans. Neural Networks Learn. Syst.2
2024 GRF-based Predictive Flocking Control with Dynamic Pattern Formation
abstract
It is promising but challenging to design flocking control for a robot swarm to autonomously follow changing patterns or shapes in a optimal distributed manner. The optimal flocking control with dynamic pattern formation is, therefore, investigated in this paper. A predictive flocking control algorithm is proposed based on a Gibbs random field (GRF), where bio-inspired potential energies are used to charaterize "robot-robot" and "robot-environment" interactions. Specialized performance-related energies, e.g., motion smoothness, are introduced in the proposed design to improve the flocking behaviors. The optimal control is obtained by maximizing a posterior distribution of a GRF. A region-based shape control is accomplished for pattern formation in light of a mean shift technique. The proposed algorithm is evaluated via the comparison with two state-of-the-art flocking control methods in an environment with obstacles. Both numerical simulations and real-world experiments are conducted to demonstrate the efficiency of the proposed design.
Chenghao Yu, Dengyu Zhang, Qingrui Zhang
ICRA3
2024 PA-LOCO: Learning Perturbation-Adaptive Locomotion for Quadruped Robots
abstract
Locomotion control is still a challenging task for quadruped robots traversing diverse terrains amidst unforeseen disturbances. Recently, privileged learning has been employed to learn reliable and robust quadrupedal locomotion over various terrains based on a teacher-student architecture. However, its one-encoder structure is not adequate in addressing external force perturbations. The student policy would experience inevitable performance degradation due to the feature embedding discrepancy between the feature encoder of the teacher policy and the one of the student policy. Hence, this paper presents a privileged learning framework with multiple feature encoders and a residual policy network for robust and reliable quadruped locomotion subject to various external perturbations. The multi-encoder structure can decouple latent features from different privileged information, ultimately leading to enhanced performance of the learned policy in terms of robustness, stability, and reliability. The efficiency of the proposed feature encoding module is analyzed in depth using extensive simulation data. The introduction of the residual policy network helps mitigate the performance degradation experienced by the student policy that attempts to clone the behaviors of a teacher policy. The proposed framework is evaluated on a Unitree GO1 robot, showcasing its performance enhancement over the state-of-the-art privileged learning algorithm through extensive experiments conducted on diverse terrains. Ablation studies are conducted to illustrate the efficiency of the residual policy network.
Zhiyuan Xiao, Qingrui Zhang
IROS4
2024 A Game-Theoretic Incentive Mechanism for Multi-Distributor Multi-Agent Federated Learning
abstract
In a multi-distributor multi-agent federated learning architecture, it is desired to maximize the overall system benefits by optimizing the association relationship between the base station and users. To address this issue, a novel incentive mechanism based on Stackelberg game is proposed. Firstly, each mobile device determines its own strategy based on the utility function. Then, each base station adjusts its own strategy based on the optimal solution of the mobile devive to maximize its own utiltiy function. By relaxing the binary variables that represent the correlation relationship into continuous variables, the 0–1 optimization problem for base stations and mobile devices is transformed into a linear problem and solved. The simulation results show that the proposed algorithm is better than random algorithm in terms of training loss and testing accuracy.
Mingkai Zhu, Qingrui Zhang
WCNC4
2023 Reinforced Potential Field for Multi-Robot Motion Planning in Cluttered Environments
abstract
Motion planning is challenging for multiple robots in cluttered environments without communication, especially in view of real-time efficiency, motion safety, distributed computation, and trajectory optimality, etc. In this paper, a reinforced potential field method is developed for distributed multi-robot motion planning, which is a synthesized design of reinforcement learning and artificial potential fields. An observation embedding with a self-attention mechanism is presented to model the robot-robot and robot-environment interactions. A soft wall-following rule is developed to improve the trajectory smoothness. Our method belongs to reactive planning, but environment properties are implicitly encoded. The total amount of robots in our method can be scaled up to any number. The performance improvement over a vanilla APF and RL method has been demonstrated via numerical simulations. Experiments are also performed using quadrotors to further illustrate the competence of our method.
Dengyu Zhang, Bo Zhu 0005, Qingrui Zhang
IROS5
2022 Multi-robot Cooperative Pursuit via Potential Field-Enhanced Reinforcement Learning
abstract
It is of great challenge, though promising, to coordinate collective robots for hunting an evader in a decentralized manner purely in light of local observations. In this paper, this challenge is addressed by a novel hybrid cooperative pursuit algorithm that combines reinforcement learning with the artificial potential field method. In the proposed algorithm, decentralized deep reinforcement learning is employed to learn cooperative pursuit policies that are adaptive to dynamic environments. The artificial potential field method is integrated into the learning process as predefined rules to improve the data efficiency and generalization ability. It is shown by numerical simulations that the proposed hybrid design outperforms the pursuit policies either learned from vanilla reinforcement learning or designed by the potential field method. Furthermore, experiments are conducted by transferring the learned pursuit policies into real-world mobile robots. Experimental results demonstrate the feasibility and potential of the proposed algorithm in learning multiple cooperative pursuit strategies.
Qingrui Zhang, Tianjiang Hu
ICRA3
2022 Model-Reference Reinforcement Learning for Collision-Free Tracking Control of Autonomous Surface Vehicles
abstract
This paper presents a novel model-reference reinforcement learning algorithm for the intelligent tracking control of uncertain autonomous surface vehicles with collision avoidance. The proposed control algorithm combines a conventional control method with reinforcement learning to enhance control accuracy and intelligence. In the proposed control design, a nominal system is considered for the design of a baseline tracking controller using a conventional control approach. The nominal system also defines the desired behaviour of uncertain autonomous surface vehicles in an obstacle-free environment. Thanks to reinforcement learning, the overall tracking controller is capable of compensating for model uncertainties and achieving collision avoidance at the same time in environments with obstacles. In comparison to traditional deep reinforcement learning methods, our proposed learning-based control can provide stability guarantees and better sample efficiency. We demonstrate the performance of the new algorithm using an example of autonomous surface vehicles.
Qingrui Zhang, Wei Pan 0004, Vasso Reppa
IEEE Trans. Intell. Transp. Syst.1
2021 Reinforcement Learning Compensated Extended Kalman Filter for Attitude Estimation
abstract
Inertial measurement units are widely used in different fields to estimate the attitude. Many algorithms have been proposed to improve estimation performance. However, most of them still suffer from 1) inaccurate initial estimation, 2) inaccurate initial filter gain, and 3) non-Gaussian process and/or measurement noise. This paper will leverage reinforcement learning to compensate for the classical extended Kalman filter estimation, i.e., to learn the filter gain from the sensor measurements. We also analyse the convergence of the estimate error. The effectiveness of the proposed algorithm is validated on both simulated data and real data.
Yujie Tang 0004, Liang Hu 0002, Qingrui Zhang, Wei Pan 0004
IROS3
2020 Lyapunov-Based Reinforcement Learning for Decentralized Multi-agent Control
Qingrui Zhang, Hao Dong 0003, Wei Pan 0004
DAI1
2013 Robust stability analysis of Markov jump standard genetic regulatory networks with mixed time delays and uncertainties
Yanzheng Zhu, Qingrui Zhang, Zuolong Wei, Lixian Zhang 0001
Neurocomputing2