Yunkai Lv

dblp:250/2030 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-5212-8629ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Safe Multi-Agent Reinforcement Learning via Distributional Safety Critic and Maximum Entropy Optimization
abstract
Deploying multi-agent reinforcement learning (MARL) in safety-critical systems faces significant challenges due to insufficient agent exploration and inadequate safety constraint guarantees. Current approaches are constrained by two fundamental limitations: inefficient exploration leading to suboptimal policies, and expected-cost-based constraint frameworks failing to ensure full-process safety. To address these challenges, this paper proposes a novel safety-aware maximum entropy MARL framework using Conditional Value-at-Risk (CVaR) as a joint safety metric, which quantifies constraint satisfaction under worst-case scenarios for multi-agent systems. Moreover, we develop the Worst-Case Multi-Agent Soft Actor-Critic (WCMASAC) algorithm, incorporating sequential update mechanisms and maximum entropy optimization for heterogeneous agents, enhanced with distributed safety critics. Theoretically, we establish the monotonic improvement property, guaranteed constraint satisfaction, and convergence to a generalized Nash equilibrium for WCMASAC. Extensive experiments on Safety-Gymnasium based benchmarks demonstrate that WCMASAC outperforms state-of-the-art baselines in both task reward acquisition and safety constraint violation reduction, while exhibiting superior exploration efficiency and risk-aware control capabilities.
Lingyue Zhang, Kaitian Chen, Yunkai Lv, Huaicheng Yan 0001
AAAI5
2026 Decentralized Pursuit of an Evader With Probabilistic Collision-Free for Differential Drive Robots
abstract
This article addresses the pursuit-evasion problem among differential drive robots in an obstacle environment with perception uncertainty. To calculate probabilistic collision-free trajectories during the pursuit process, this article introduces the chance-constraint pursuit Voronoi cell (CCPVC), which consists of separation hyperplanes between robots and separation hyperplanes between robots and obstacles. The optimization problems are formulated to compute the separation hyperplanes, and the solution methods are provided. By incorporating two buffer terms, CCPVC exhibits favorable probabilistic collision avoidance properties. Furthermore, a nearest point finding algorithm specifically designed for pursuit scenarios, along with a distributed pursuit control policy tailored for differential drive robots are proposed based on CCPVC. Rigorous proofs for the probabilistic collision avoidance guarantees of CCPVC and the control law during the pursuit process are provided, respectively. Finally, the effectiveness of the proposed methods is validated through simulations and experiments.
Kai Rao, Huaicheng Yan 0001, Yunkai Lv, Youmin Zhang 0001
IEEE Trans. Cybern.3
2026 Safety-Certified Distributed Formation for Uncertain Multirobot Systems With Minimum Restrictive Connectivity Maintenance
abstract
This article addresses the distributed formation control problem for safety-certified uncertain multirobot systems with minimum restrictive connectivity maintenance (MRCM). Compared to this problem with deterministic models, it presents inherently more challenging requirements for safety and stability analysis. An integrated quadratic programming controller incorporating input-to-state safety control barrier functions (ISSf-CBFs) is developed to simultaneously enforce safety certification and MRCM. This controller builds upon a nominal design leveraging input-to-state stability and ISSf lemmas. The derived ISSf-CBFs provide formal guarantees for both safety certification and MRCM under uncertainties. The key to obtaining this controller lies in constructing ISSf-CBFs that enable safety certification and MRCM, from which the corresponding constraints are formally derived. Unlike existing methods, the strategy maximizes formation performance by executing formation tasks at full capability whenever possible, and only modifying the nominal controller when safety or MRCM constraints are violated. Finally, simulation and experiment validate the effectiveness of the proposed controller.
Yunkai Lv, Zhuping Wang, Hao Zhang 0008
IEEE Trans. Ind. Informatics2
2026 Path-Guided Cooperative Circumnavigation of Autonomous Vehicles With Collision Avoidance: An Input-to-State Safety-Certified Solution
abstract
This paper investigates path-guided cooperative circumnavigation control for autonomous vehicles with collision avoidance. Compared to the trajectory-guided control frameworks, which require vehicles to track time-dependent trajectories, the proposed method eliminates temporal constraints but introduces new design challenges. Coordination error variables based on separation angles and path information are defined, and a temporally unconstrained guidance law incorporating linear and angular velocities is proposed. The developed guidance law is inspired by chase-and-wait strategies, achieving structural simplicity. To address collision avoidance involving both inter-vehicle interactions and external obstacles for autonomous vehicles under model uncertainties and environmental disturbances, an input-to-state safe control barrier function (ISSf-CBF) is employed. The constructed ISSf-CBF systematically maps state-space safety constraints to control input constraints, thereby ensuring collision-free operations in dynamic environments. The key to achieving robust collision-free path-guided circumnavigation lies in synthesizing an input-to-state safety-certified controller by defining ISSf-CBF constraints. Simulation results verify the validity and effectiveness of the theoretical results.
Zhuping Wang, Hao Zhang 0008, Yunkai Lv
IEEE Trans. Intell. Transp. Syst.4
2025 Distributed Pursuit of an Evader with Adaptive Robust Path Control Under State Measurement Uncertainty
abstract
This paper presents a distributed pursuit frame-work for environments with obstacles considering state measurement uncertainty. Our framework consists of two primary components: the computation of safe pursuit regions based on Voronoi cell (VC) and the solution of an adaptive robust path controller based on Control Barrier Function (CBF). Initially, the chance constrained obstacle-aware Voronoi cell (CCOVC) for each pursuer is constructed by calculating separation hyperplane and buffer terms. Subsequently, we formulate chance CBF and chance Control Lyapunov Function (CLF) constraints, using convex approximation to determine their upper bounds. We then find the adaptive robust path controller by solving a Quadratically Constrained Quadratic Program (QCQP). The advantage of this framework lies in its capability to adaptively compute the path controller and ensure robust collision avoidance among pursuers and with obstacles. Simulation and experimental results demonstrate the effectiveness and robustness of the proposed framework.
Kai Rao, Huaicheng Yan 0001, Penghui Yang 0002, Yunkai Lv
ICRA5
2024 Local-Bearing-Based Prescribed-Time Distributed Localization of Multiagent Systems With Noisy Measurement
abstract
This work investigates stability and localizability of local-bearing-based multiagent systems without common orientation in the presence of measurement noise, which are more general but also more challenging to deal with than global-bearing-based multiagent systems under ideal environment. Based on local-bearing unbiased estimator constructed from the historical information and a newly designed time-varying gain, a robust prescribed-time orientation estimation algorithm is proposed to ensure that the local reference frame of the follower agent is aligned with the global one. The local bearing information is more easily obtained than global one. Therefore, the new orientation estimation result is expected to be more widely applicable. The robust orientation estimation algorithm is then applied to the problem of localization estimation, and a prescribed-time distributed localization estimation algorithm is developed. The distinctive advantage of this work is that only the local bearing information is used, the fast and controllable localization estimation is achieved. The global convergence is derived based on the cascade system. Some simulation and experiment results are provided to prove the effectiveness of the proposed estimation algorithms.
Yunkai Lv, Hao Zhang 0008, Zhuping Wang, Huaicheng Yan 0001
IEEE Trans. Ind. Informatics1
2024 Distributed Localization for Multi-Agent Systems With Random Noise Based on Iterative Learning
abstract
This article is concerned with the real-time localization problem for the dynamic multi-agent systems with measurement and communication noises under directed graphs. The barycentric coordinates are introduced to describe the relative position between agents. A novel robust distributed localization estimation algorithm based on iterative learning is proposed. The relative-distance unbiased estimator constructed from the historical iterative information is used to suppress the measurement noise. The designed stochastic approximation method with two iterative-varying gains is used to inhibit the communication noise. Under the zero-mean and independent distributed conditions on the measurement and communication noises, the asymptotic convergence of the proposed methods is derived. The numerical simulation and the QBot-2e robot experiment are conducted to test and verify the effectiveness and the practicability of the proposed methods.
Yunkai Lv, Hao Zhang 0008, Zhuping Wang, Huaicheng Yan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Distributed sensor network localization based on local bearing measurement
Hao Zhang 0008, Yunkai Lv, Zhuping Wang, Zhian Zhan, Huaicheng Yan 0001
Sci. China Inf. Sci.2
2023 Distributed Localization Estimation for Dynamic Multiagent Systems
abstract
This article investigates the real-time localization problem of dynamic multiagent systems with repetitive operation characteristics under directed graph. A distributed localization estimation algorithm based on iterative learning is proposed. The barycentric coordinates calculated based on the relative distance are used to estimate the real coordinates of the agent. Different from the traditional estimation methods along the time axis, the proposed method utilizes the information of iteration axis simultaneously. In this method, the current estimation coordinates are updated by using the estimation coordinates of the same sampling time in previous iteration, the estimation accuracy is improved, and the velocity constraint is removed. Additionally, the real-time localization problem of dynamic multiagent systems under arbitrary deployment is concerned. An improved distributed localization estimation algorithm with signed coefficients based on iterative learning is proposed. Meanwhile, the results are also extended to the localization estimation of multiagent systems with arbitrary deployment in 3-D space. By introducing Richardson iteration and infinite norm, the global asymptotic convergence of the proposed methods is guaranteed. Finally, numerical simulations and the Qbot-2e robot experiment are provided to show the effectiveness and validity of the obtained results.
Yunkai Lv, Hao Zhang 0008, Zhuping Wang, Shun-Feng Su
IEEE Trans. Ind. Informatics1