Man Li 0002

dblp:62/3171-2 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-9776-1628ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Online safe tracking control with barrier-like functions: Coordinating dynamic output performance and obstacle avoidance
Ambreen Basheer, Man Li 0002, Weiming Fu, Jiahu Qin
Neurocomputing2
2024 Ensuring Safety in LLM-Driven Robotics: A Cross-Layer Sequence Supervision Mechanism
abstract
Integrating Large Language Models (LLMs) into robotics significantly enhances autonomous task planning. However, ensuring that multi-step task plans (action sequence) generated by LLMs comply with pre-defined safety constraints during planning and execution remains a challenge, limiting their adaptability in complex environments. To address this issue, a mechanism that can monitor and adjust the plan generated by the LLM-driven task planner and guide the motion planner to avoid potential risks during action execution is required. Therefore, this paper proposes a cross-layer sequence supervision mechanism. Specifically, we employ linear temporal logic syntax to express safety constraints and convert them into a set of nondeterministic Büchi automatons to build a cross-layer safety supervisor. For the task planning layer, the safety supervisor provides a closed-loop correction mechanism that can identify violations in the task plan in real time and guide LLM-driven planners to correct this plan to ensure compliance. For the motion planning layer, the safety supervisor introduces virtual "obstacle" information into the task plan to form the task plan tuple. Based on this plan tuple, the motion planner can proactively prevent unsafe behaviors during action execution. Extensive experimentation demonstrates significant improvements in safety with this cross-layer supervision mechanism, highlighting its potential to enhance LLM-driven robotic technology. Experiment details can be found in https://youtu.be/BDdSSEP6HJw.
Qingchen Liu, Jiahu Qin, Man Li 0002
IROS4
2024 Game-Based Approximate Optimal Motion Planning for Safe Human-Swarm Interaction
abstract
Safety as a fundamental requirement for human-swarm interaction has attracted a lot of attention in recent years. Most existing approaches solve a constrained optimization problem at each time step, which has a high real-time requirement. To deal with this challenge, this article formulates the safe human-swarm interaction problem as a Stackerberg-Nash game, in which the optimization is performed over the entire time domain. The leader robot is supposed to be in a dominant position, interacting directly with the human operator to realize trajectory tracking and responsible for guiding the swarm to avoid obstacles. The follower robots always take their best responses to leader's behavior with the purpose of achieving the desired formation. Following the bottom-up principle, we first design the best-response controllers, that is, Nash equilibrium strategies, for the followers. Then, a Lyapunov-like control barrier function-based safety controller and a learning-based formation tracking controller for the leader are designed to realize safe and robust cooperation. We show that the designed controllers can make the robotic swarms move in a desired geometric formation following the human command and modify their motion trajectories autonomously when the human command is unsafe. The effectiveness of the proposed approach is verified through simulation and experiments. The experiment results further show that safety can still be guaranteed even when there exists a dynamic obstacle.
Man Li 0002, Jiahu Qin, Jiacheng Li 0005, Qingchen Liu, Yang Shi 0001, Yu Kang 0001
IEEE Trans. Cybern.1
2024 Differential Game-Based Control for Nonlinear Human-Robot Interaction System With Unknown Desired Trajectory
abstract
Differential game is an effective technique to describe the negotiation between the humans and robots, which is widely used to realize the trajectory tracking tasks in the human-robot interaction (HRI). However, most existing works consider the control-affine HRI systems and assume the desired trajectory is available to both the human and the robot, which limit the scope of applications. To overcome these difficulties, this work focuses on the nonaffine HRI system and supposes that the desired trajectory is not available to the robot. A novel differential game framework encoding the desired trajectory estimator is proposed, where the desired trajectory is estimated via the Gaussian process regression (GPR) technique. To address the challenge arising from the nonlinearity of the HRI system, we equivalently transform the original problem into the one in a differentially flat space, and seek the equilibrium strategies for the transformed problem substitutionally. We further prove that the trajectory tracking error satisfies a probabilistic bound, whose confidence interval tightens as the decrease of noise variance during the interaction. Comparative simulation results show that our method outperforms the learning-based method in terms of robustness, parameters setting, and time consumption. Experiment results further show that the tracking error under the proposed human-robot cooperative algorithm is reduced by 55% compared to the human direct control.
Kang Tong, Man Li 0002, Jiahu Qin, Qichao Ma 0001, Jie Zhang 0110, Qingchen Liu
IEEE Trans. Cybern.2
2022 Multiplayer Stackelberg-Nash Game for Nonlinear System via Value Iteration-Based Integral Reinforcement Learning
abstract
In this article, we study a multiplayer Stackelberg-Nash game (SNG) pertaining to a nonlinear dynamical system, including one leader and multiple followers. At the higher level, the leader makes its decision preferentially with consideration of the reaction functions of all followers, while, at the lower level, each of the followers reacts optimally to the leader's strategy simultaneously by playing a Nash game. First, the optimal strategies for the leader and the followers are derived from down to the top, and these strategies are further shown to constitute the Stackelberg-Nash equilibrium points. Subsequently, to overcome the difficulty in calculating the equilibrium points analytically, we develop a novel two-level value iteration-based integral reinforcement learning (VI-IRL) algorithm that relies only upon partial information of system dynamics. We establish that the proposed method converges asymptotically to the equilibrium strategies under the weak coupling conditions. Moreover, we introduce effective termination criteria to guarantee the admissibility of the policy (strategy) profile obtained from a finite number of iterations of the proposed algorithm. In the implementation of our scheme, we employ neural networks (NNs) to approximate the value functions and invoke the least-squares methods to update the involved weights. Finally, the effectiveness of the developed algorithm is verified by two simulation examples.
Man Li 0002, Jiahu Qin, Nikolaos M. Freris, Daniel W. C. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2022 Bio-Inspired Dynamic Collective Choice in Large-Population Systems: A Robust Mean-Field Game Perspective
abstract
Inspired by the collective decision making in biological systems, such as honeybee swarm searching for a new colony, we study a dynamic collective choice problem for large-population systems with the purpose of realizing certain advantageous features observed in biology. This problem focuses on the situation where a large number of heterogeneous agents subject to adversarial disturbances move from initial positions toward one of the destinations in a finite time while trying to remain close to the average trajectory of all agents. To overcome the complexity of this problem resulting from the large population and the heterogeneity of agents, and also to enforce some specific choices by individuals, we formulate the problem under consideration as a robust mean-field game with non-convex and non-smooth cost functions. Through Nash equivalence principle, we first deal with a single-player$H_{\infty }$tracking problem by taking the population behavior as a fixed trajectory, and then establish a mean-field system to estimate the population behavior. Optimal control strategies and worst disturbances, independent of the population size, are designed, which give a way to realize the collective decision-making behavior emerged in biological systems. We further prove that the designed strategies constitute$\epsilon _{N}$-Nash equilibrium, where$\epsilon _{N}$goes toward zero as the number of agents increases to infinity. The effectiveness of the proposed results are illustrated through two simulation examples.
Man Li 0002, Jiahu Qin, Yaonan Wang 0001, Yu Kang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Model-Free Reinforcement Learning for Fully Cooperative Consensus Problem of Nonlinear Multiagent Systems
abstract
This article presents an off-policy model-free algorithm based on reinforcement learning (RL) to optimize the fully cooperative (FC) consensus problem of nonlinear continuous-time multiagent systems (MASs). First, the optimal FC consensus problem is transformed into solving the coupled Hamilton-Jacobian-Bellman (HJB) equation. Then, we propose a policy iteration (PI)-based algorithm, which is further proved to be effective to solve the coupled HJB equation. To implement this scheme in a model-free way, a model-free Bellman equation is derived to find the optimal value function and the optimal control policy for each agent. Then, based on the least-squares approach, the tuning law for actor and critic weights is derived by employing actor and critic neural networks into the model-free Bellman equation to approximate the target policies and the value function. Finally, we propose an off-policy model-free integral RL (IRL) algorithm, which can be used to optimize the FC consensus problem of the whole system in real time by using measured data. The effectiveness of this proposed algorithm is verified by the simulation results.
Man Li 0002
IEEE Trans. Neural Networks Learn. Syst.2
2021 Semiglobal Cluster Consensus for Heterogeneous Systems With Input Saturation
abstract
In this article, the semiglobal cluster consensus problem is investigated for heterogeneous generic linear systems with input saturation. A general case in a leaderless framework is studied first, and then in order to broaden the scope of application, we consider a special case in which the leader nodes are pinned intermittently. To tackle the above problems, we propose a linear control scheme by using the low-gain feedback technique under the assumptions that each node is asymptotically null controllable and the underlying topology of each cluster (the extended cluster under the intermittent pinning control) has a directed spanning tree. The Lyapunov-based method and the low-gain feedback technique are developed for convergence analysis. It is shown that for both cases, the convergence rate is explicitly specified, which depends on the low-gain parameter and system matrices. Finally, two numerical examples are provided to verify the effectiveness of the theoretical findings.
Man Li 0002, Changyin Sun 0001
IEEE Trans. Cybern.2
2021 Hierarchical Optimal Synchronization for Linear Systems via Reinforcement Learning: A Stackelberg-Nash Game Perspective
abstract
Considering the fact that in the real world, a certain agent may have some sort of advantage to act before others, a novel hierarchical optimal synchronization problem for linear systems, composed of one major agent and multiple minor agents, is formulated and studied in this article from a Stackelberg-Nash game perspective. The major agent herein makes its decision prior to others, and then, all the minor agents determine their actions simultaneously. To seek the optimal controllers, the Hamilton-Jacobi-Bellman (HJB) equations in coupled forms are established, whose solutions are further proven to be stable and constitute the Stackelberg-Nash equilibrium. Due to the introduction of the asymmetric roles for agents, the established HJB equations are more strongly coupled and more difficult to solve than that given in most existing works. Therefore, we propose a new reinforcement learning (RL) algorithm, i.e., a two-level value iteration (VI) algorithm, which does not rely on complete system matrices. Furthermore, the proposed algorithm is shown to be convergent, and the converged values are exactly the optimal ones. To implement this VI algorithm, neural networks (NNs) are employed to approximate the value functions, and the gradient descent method is used to update the weights of NNs. Finally, an illustrative example is provided to verify the effectiveness of the proposed algorithm.
Man Li 0002, Jiahu Qin, Qichao Ma 0001, Wei Xing Zheng 0001, Yu Kang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Distributed time-varying group formation control for generic linear systems with observer-based protocols
Man Li 0002, Qichao Ma 0001, Chongjian Zhou, Jiahu Qin, Yu Kang 0001
Neurocomputing1
2019 Optimal Synchronization Control of Multiagent Systems With Input Saturation via Off-Policy Reinforcement Learning
abstract
In this paper, we aim to investigate the optimal synchronization problem for a group of generic linear systems with input saturation. To seek the optimal controller, Hamilton-Jacobi-Bellman (HJB) equations involving nonquadratic input energy terms in coupled forms are established. The solutions to these coupled HJB equations are further proven to be optimal and the induced controllers constitute interactive Nash equilibrium. Due to the difficulty to analytically solve HJB equations, especially in coupled forms, and the possible lack of model information of the systems, we apply the data-based off-policy reinforcement learning algorithm to learn the optimal control policies. A byproduct of this off-policy algorithm is shown that it is insensitive to probing noise that is exerted to the system to maintain persistence of excitation condition. In order to implement this off-policy algorithm, we employ actor and critic neural networks to approximate the controllers and the cost functions. Furthermore, the estimated control policies obtained by this presented implementation are proven to converge to the optimal ones under certain conditions. Finally, an illustrative example is provided to verify the effectiveness of the proposed algorithm.
Jiahu Qin, Man Li 0002, Yang Shi 0001, Qichao Ma 0001, Wei Xing Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.2