VLDB 2026 Research / reviewers in the wild / expert
Peng Yi 0001
dblp:98/1202-1
· DBLP profile ↗
19ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-2494-1505ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dynamic Aggregative Game-Based Distributed Approach for Defensive Spacecraft Formations
Peng Yi 0001, Jinlong Lei, Dechao Ran, Lu Cao 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | A Decentralized Designed Distributed Observer for Linear Interconnected SystemsabstractThis article addresses the problem of distributed state estimation (DSE) for discrete-time interconnected systems, where the observed system is composed of subsystems interconnected through state-to-state and state-to-output couplings. Inspired by the leader-follower consensus method, we propose a distributed observer that enables each subsystem to estimate the entire state of the interconnected system. Under certain structural assumptions, we derive necessary and sufficient conditions for the stability of the estimation error dynamics. We further present a decentralized design of the proposed observer, where the operation and construction of the observer can be completed by each subsystem using its locally available information, including the system's basic configuration, local measurements, and data exchanged with neighboring subsystems. In addition, we demonstrate that our distributed estimation framework can be applied to solve the distributed estimation problem for linear time-invariant (LTI) systems with fixed composition by employing an observability decomposition method. Finally, we illustrate the effectiveness of our scheme by applying it to vehicle platooning. Shuaiting Huang, Lingying Huang, Peng Yi 0001, Hong Chen 0003, Guodong Shi, Junfeng Wu 0001 |
IEEE Trans. Cybern. | 3 |
| 2026 | An Automated Reinforcement Learning Reward Design Framework With Large Language Model for Cooperative Platoon CoordinationabstractReinforcement Learning (RL) has demonstrated excellent decision-making potential in platoon coordination problems. However, due to the variability of coordination goals, the complexity of the decision problem, and the time-consumption of trial-and-error in manual design, finding a well performance reward function to guide RL training to solve complex platoon coordination problems remains challenging. In this paper, we formally define the Platoon Coordination Reward Design Problem (PCRDP), extending the RL-based cooperative platoon coordination problem to incorporate automated reward function generation. To address PCRDP, we propose a Large Language Model (LLM)-based Platoon coordination Reward Design (PCRD) framework, which systematically automates reward function discovery through LLM-driven initialization and iterative optimization. In this method, LLM first initializes reward functions based on environment code and task requirements with an Analysis and Initial Reward (AIR) module, and then iteratively optimizes them based on training feedback with an evolutionary module. The AIR module guides LLM to deepen their understanding of code and tasks through a chain of thought, effectively mitigating hallucination risks in code generation. The evolutionary module fine-tunes and reconstructs the reward function, achieving a balance between exploration diversity and convergence stability for training. To validate our approach, we establish six challenging coordination scenarios with varying complexity levels within the Yangtze River Delta transportation network simulation. Comparative experimental results demonstrate that RL agents utilizing PCRD-generated reward functions consistently outperform human-engineered reward functions, achieving an average of 10% higher performance metrics in all scenarios. Dixiao Wei, Peng Yi 0001, Jinlong Lei, Yiguang Hong, Hairong Dong 0001, Yuchuan Du |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | ADMM-iCLQG: A Fast Solver of Constrained Dynamic Games for Planning Multi-Vehicle Feedback TrajectoryabstractTrajectory planning for autonomous driving requires safe, robust and efficient algorithms in complex, dynamic environments where multiple vehicles interact. Constrained Dynamic Games (CDG) provide a unifying framework to model such multi-agent interactions, with solutions corresponding to Generalized Nash Equilibria (GNE). Existing solving methods of CDG can be divided into the open-loop strategies, which is computational-efficient but lack robustness to disturbances, and the feedback strategies, which offer adaptability but often struggle to guarantee real-time performance when handling hard constraints. To address these limitations, we introduce ADMM-iCLQG, a solver of CDG tailored for planning a feedback trajectory in real-time. We iteratively approximate the nonlinear CDG with a linearly Constrained Linear Quadratic Game (CLQG). By leveraging the Alternating Direction Method of Multipliers (ADMM), the solver efficiently manages constraints while computing feedback strategies that converge to GNE. We demonstrate the algorithm effectiveness through extensive simulations, showing its superiority in computational efficiency and reliability. Furthermore, physical experiments confirm the approach's potential of real-time implementation, achieving a planning frequency exceeding 40 Hz for up to six vehicles. Hui Min, Peng Yi 0001, Jinlong Lei |
IV | 2 |
| 2025 | Multiview Landmark-Assisted UAV Swarm 6-DoF Pose Estimation Using Resonant Beam and VIOabstractAs an essential aerial platform in Internet of Things (IoT) applications, UAV swarms require high-precision attitude estimation in GPS-limited and dynamic environments, which supports higher-level IoT functions such as smart logistics, disaster response, and environmental monitoring. However, most odometry-based pose estimation methods in dynamic scenarios without GPS encounter issues with cumulative errors over time. In this paper, we propose a method to reduce these cumulative errors by constraining the absolute positioning of Visual-Inertial Odometry (VIO) using relative poses obtained from multi-view landmark and Resonant-Beam (RBeam) sensors between UAVs. For synchronous moments, we estimate the 6 Degree-of-Freedom (DoF) relative pose using the Angle of Arrival and Time of Flight data from the RBeam, combined with nonlinear optimization. For asynchronous moments, a two-stage visual estimation method is introduced, combining multi-camera epipolar geometry for rotation recovery and depth reconstruction optimization for translation recovery, enabling the estimation of asynchronous relative poses. Finally, we design a global objective function based on a sliding window and factor graph, integrating RBeam, multi-view landmark, and VIO for absolute 6-DoF pose optimization of the UAV swarm. Simulation results demonstrate that optimizing with the addition of RBeam synchronous relative poses improves overall positioning accuracy by 32.94% compared to pure VIO. Incorporating both RBeam synchronous and visual asynchronous relative poses further enhances overall positioning accuracy by 37.65%. Additionally, the UAV’s attitude benefits from the rotational constraints provided by the RBeam, achieving over 30% improvement in the pitch and yaw directions. Mengyuan Xu, Wen Fang 0001, Qingwen Liu 0001, Peng Yi 0001, Yiguang Hong |
IEEE Internet Things J. | 6 |
| 2025 | Emergence of cooperation promoted by higher-order strategy updatesabstractCooperation is fundamental to human societies, and the interaction structure among individuals profoundly shapes its emergence and evolution. In real-world scenarios, cooperation prevails in multi-group (higher-order) populations, beyond just dyadic behaviors. Despite recent studies on group dilemmas in higher-order networks, the exploration of cooperation driven by higher-order strategy updates remains limited due to the intricacy and indivisibility of group-wise interactions. Here we investigate four categories of higher-order mechanisms for strategy updates in public goods games and establish their mathematical conditions for the emergence of cooperation. Such conditions uncover the impact of both higher-order strategy updates and network properties on evolutionary outcomes, notably highlighting the enhancement of cooperation by overlaps between groups. Interestingly, we discover that the group-mutual comparison update - selecting a high-fitness group and then imitating a random individual within this group - can prominently promote cooperation. Our analyses further unveil that, compared to pairwise interactions, higher-order strategy updates generally improve cooperation in most higher-order networks. These findings underscore the pivotal role of higher-order strategy updates in fostering collective cooperation in complex social systems. Dini Wang, Peng Yi 0001, Yiguang Hong, Jie Chen 0003 |
PLoS Comput. Biol. | 2 |
| 2025 | Multiple Pursuers Versus One Evader Reach-Avoid Differential Games With Asymmetric ObservationsabstractThis paper considers a multiple-pursuer-one-evader reach-avoid (RA) differential game in a two-dimensional (2D) plane that is split by a straight line into a goal region and a play region. The evader aims to enter the goal region from the play region without being captured, while the pursuers try to intercept the evader in the play region. In terms of information structures, we assume that the evader has perfect information about the game, while the pursuers have uncertain observations of the evader’s location and speed. First, a robust estimation of the safe region (RESR) is constructed by computing the safe region boundary beliefs with respect to these uncertainties. Based on the RESR, the winning condition for the one-on-one RA game is analyzed, and an optimal interception strategy is derived. A proximity strategy is also proposed when the winning condition is not satisfied. Then, the concept of effective coalition is studied and its properties are proposed to enhance the robustness of the strategy. Additionally, a receding horizon cooperation algorithm with a strategy switching mechanism is designed for the pursuers to cooperatively capture the evader under distance-related observation uncertainties. Finally, simulations and experiments are conducted to validate the effectiveness of the proposed strategies. Note to Practitioners—This paper is motivated by practical considerations of the reach-avoid game, in which players have different capacities for observation. In most related works, pursuers and evaders are assumed to have perfect information about the game. Optimal strategies are derived from safe regions (SR) determined by perfect information. Our objective is to challenge this assumption and study the robust pursuit strategy from the perspective of the pursuers under observational uncertainty. We approach observational uncertainty in two ways. On the one hand, to obtain the SR of the evader, we construct the robust estimation of the safe region (RESR) and identify its$3\sigma $principle boundary with theoretical guarantees. On the other hand, to mitigate the impact of observational uncertainty on pursuit strategies, we propose the concept of effective coalition, which can compress the RESR of the evader and capture the evader farther if acting individually. By combining these contributions, we design a receding horizon strategy for the pursuit coalition that is tailored for the autonomous multi-agent system in robotics, defense, and surveillance. At present, this method cannot be used directly in the multi-pursuer versus multi-evader game scenario. Future studies will investigate the scalability of this method. Hongwei Fang, Peng Yi 0001, Di Deng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Multi-Agent Deep Reinforcement Learning for Distributed and Autonomous Platoon Coordination via Speed-Regulation Over Large-Scale Transportation NetworksabstractTruck platooning technology enables a group of trucks to travel closely together, with which the platoon can save fuel, improve traffic flow efficiency, and improve safety. In this paper, we consider the platoon coordination problem in a large-scale transportation network to promote cooperation among trucks and optimize overall efficiency. Involving the regulation of both speed and departure times at hubs, we formulate the coordination problem as a complicated dynamic stochastic integer programming under network and information constraints. To obtain an autonomous, distributed, and robust platoon coordination policy, we formulate the problem into a model of the Decentralized-Partial Observable Markov Decision Process. Then, we propose a Multi-Agent Deep Reinforcement Learning framework named Trcuk Attention-QMIX (TA-QMIX) to train an efficient online decision policy. TA-QMIX utilizes the attention mechanism to enhance the representation of truck fuel gains and delay times and provides explicit truck cooperation information during the training process, promoting trucks’ willingness to cooperate. The training framework adopts centralized training with decentralized execution, thus training a policy for trucks to make decisions online using only nearby information. Hence, the policy can be autonomously executed on a large-scale network. Finally, we perform comparison experiments and ablation experiments in the transportation network of the Yangtze River Delta region in China to verify the effectiveness of the proposed framework. In a repeated comparative experiment with 5,000 trucks, our method averages 19. 17% fuel savings with an average delay of only 9.57 minutes per truck and a decision time of 0.001 seconds. Dixiao Wei, Peng Yi 0001, Jinlong Lei, Xingyi Zhu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | A Local Information Aggregation-Based Multiagent Reinforcement Learning for Robot Swarm Dynamic Task AllocationabstractIn this article, we explore how to optimize task allocation for robot swarms in dynamic environments, emphasizing the necessity of formulating robust, flexible, and scalable strategies for robot cooperation. We introduce a novel framework using a decentralized partially observable Markov decision process (Dec-POMDP), specifically designed for distributed robot swarm networks. At the core of our methodology is the local information aggregation multiagent deep deterministic policy gradient (LIA-MADDPG) algorithm, which merges centralized training with distributed execution. During the centralized training phase, a local information aggregation (LIA) module is meticulously designed to gather critical data from neighboring robots, enhancing decision-making efficiency. In the distributed execution phase, a strategy improvement method is proposed to dynamically adjust task allocation based on changing and partially observable environmental conditions. Our empirical evaluations show that the LIA module can be seamlessly integrated into various centralized training and decentralized execution (CTDE)-based multiagent reinforcement learning (MARL) methods, significantly enhancing their performance. Additionally, by comparing LIA-MADDPG with six conventional reinforcement learning algorithms and a heuristic algorithm, we demonstrate its superior scalability, rapid adaptation to environmental changes, and ability to maintain both stability and convergence speed. These results underscore LIA-MADDPG's outstanding performance and its potential to significantly improve dynamic task allocation in robot swarms through enhanced local collaboration and adaptive strategy execution. Yang Lv 0007, Jinlong Lei, Peng Yi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Motion Planning at Intersections with Safe Differential Games based on Control Barrier FunctionabstractMotion planning at intersections is a challenging problem in autonomous driving due to the complicated interactions. The existing pipeline of "planning after predicting" is too conservative, can reduce traffic efficiency. Using game theory to model the non-cooperative coupling relationships between multiple vehicles can resolve the above problems, but such methods cannot guarantee safety without collision. This paper presents motion planning for autonomous driving with safe differential games based on Control Barrier Function (CBF), and also provides a safety-critical generalized Nash equilibrium seeking algorithm. We handle the hard CBF constraints through augmented Lagrangian multiplier method. Motivated by iterative Linear-Quadratic Game (iLQG) algorithm, we use the Taylor expansion method to approximate the model into an Linear-Quadratic (LQ) structure, and then incrementally solve this problem with an iterative feedback LQ game algorithm. Through Carla simulation and hardware testing, our results indicate that the algorithm can find a balance between safety and efficiency while maintaining real-time implementation performance. Peng Yi 0001, Qingwen Liu 0001, Yiguang Hong |
IV | 1 |
| 2024 | Distributed Optimization With Projection-Free Dynamics: A Frank-Wolfe PerspectiveabstractWe consider solving distributed constrained optimization in this article. To avoid projection operations due to constraints in the scenario with large-scale variable dimensions, we propose distributed projection-free dynamics by employing the Frank-Wolfe method, also known as the conditional gradient. Technically, we find a feasible descent direction by solving an alternative linear suboptimization. To make the approach available over multiagent networks with weight-balanced digraphs, we design dynamics to simultaneously achieve both the consensus of local decision variables and the global gradient tracking of auxiliary variables. Then, we present the rigorous convergence analysis of the continuous-time dynamical systems. Also, we derive its discrete-time scheme with an accordingly proved convergence rate of O(1/k) . Furthermore, to clarify the advantage of our proposed distributed projection-free dynamics, we make detailed discussions and comparisons with both existing distributed projection-based dynamics and other distributed Frank-Wolfe algorithms. Guanpu Chen, Peng Yi 0001, Yiguang Hong, Jie Chen 0003 |
IEEE Trans. Cybern. | 2 |
| 2024 | Distributed Nash Equilibrium Seeking Dynamics With Discrete CommunicationabstractIn this brief, we aim to provide a distributed Nash equilibrium seeking algorithm in continuous time with discrete communications. A group of agents are considered playing a continuous-kernel noncooperative game over a network. The agents need to seek the Nash equilibrium when each player cannot get the overall action profiles in real time, rather are only able to get information from its networked neighbors. Meanwhile, a continuous-time dynamics is discussed for the players to update their variables, but the communications over the network are only assumed to allow at discrete-time instants, since continuous-time communications are prohibitive and cumbersome in practice. First, the periodic communication is considered at a fixed interval, and the solvability of Nash equilibrium seeking is shown with discrete communications. Then, an event-trigger communication scheme is proposed to further reduce the communication rounds. Nevertheless, the event-trigger communication scheme requires each player continuously monitoring its local states. To alleviate the monitoring burden, a periodic event detection mechanism is further developed. The exponential convergence of the dynamics with the three discrete communication schemes is proven. Finally, the comparative simulation studies are designed to illustrate the algorithm performance with different communication schemes and parameter settings. Rui Yu 0001, Yutao Tang, Peng Yi 0001, Li Li 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Distributed Time-Varying Convex Optimization With Dynamic QuantizationabstractIn this work, we design a distributed algorithm for time-varying convex optimization over networks with quantized communications. Each agent has its local time-varying objective function, while the agents need to cooperatively track the optimal solution trajectories of global time-varying functions. The distributed algorithm is motivated by the alternating direction method of multipliers, but the agents can only share quantization information through an undirected graph. To reduce the tracking error due to information loss in quantization, we apply the dynamic quantization scheme with a decaying scaling function. The tracking error is explicitly characterized with respect to the limit of the decaying scaling function in quantization. Furthermore, we are able to show that the algorithm could asymptotically track the optimal solution when time-varying functions converge, even with quantization information loss. Finally, the theoretical results are validated via numerical simulation. Ziqin Chen, Peng Yi 0001, Li Li 0008, Yiguang Hong |
IEEE Trans. Cybern. | 2 |
| 2020 | Online Convex Optimization Over Erdos-Renyi Random NetworksabstractThe work studies how node-to-node communications over an Erd\H{o}s-R\'enyi random network influence distributed online convex optimization, which is vital in solving large-scale machine learning in antagonistic or changing environments. At per step, each node (computing unit) makes a local decision, experiences a loss evaluated with a convex function, and communicates the decision with other nodes over a network. The node-to-node communications are described by the Erd\H{o}s-R\'enyi rule, where independently each link takes place with a probability $p$ over a prescribed connected graph. The objective is to minimize the system-wide loss accumulated over a finite time horizon. We consider standard distributed gradient descents with full gradients, one-point bandits and two-points bandits for convex and strongly convex losses, respectively. We establish how the regret bounds scale with respect to time horizon $T$, network size $N$, decision dimension $d$, and an algebraic network connectivity. The regret bounds scaling with respect to $T$ match those obtained by state-of-the-art algorithms and fundamental limits in the corresponding centralized online optimization problems, e.g., $\mathcal{O}(\sqrt{T}) $ and $\mathcal{O}(\ln(T)) $ regrets are established for convex and strongly convex losses with full gradient feedback and two-points information, respectively. For classical Erd\H{o}s-R\'enyi networks over all-to-all possible node communications, the regret scalings with respect to the probability $p$ are analytically established, based on which the tradeoff between the communication overhead and computation accuracy is clearly demonstrated. Numerical studies have validated the theoretical findings. Jinlong Lei, Peng Yi 0001, Yiguang Hong, Jie Chen 0003, Guodong Shi |
NeurIPS | 2 |
| 2020 | Synthesis of recurrent neural dynamics for monotone inclusion with application to Bayesian inference
Peng Yi 0001, ShiNung Ching |
Neural Networks | 1 |
| 2020 | Asynchronous Distributed Algorithms for Seeking Generalized Nash Equilibria Under Full and Partial-Decision InformationabstractThis paper investigates asynchronous algorithms for distributedly seeking generalized Nash equilibria with delayed information in multiagent networks. In the game model, a shared affine constraint couples all players' local decisions. Each player is assumed to only access its private objective function, private feasible set, and a local block matrix of the affine constraint. We first give an algorithm for the case when each agent is able to fully access all other players' decisions. By using auxiliary variables related to communication links and the edge Laplacian matrix, each player can carry on its iteration asynchronously with only private data and possibly delayed information from its neighbors. Then, we consider the case when agents cannot know all other players' decisions, called a partial-decision information case. We introduce a local estimation of the overall agents' decisions and incorporate consensus dynamics on these local estimations. The two algorithms do not need any centralized clock coordination, fully exploit the local computation resource, and remove the idle time due to waiting for the "slowest" agent. Both algorithms are developed by preconditioned forward-backward operator splitting, and their convergence is shown by relating them to asynchronous fixed-point iterations, under proper assumptions and fixed and nondiminishing step-size choices. Numerical studies verify the algorithms' convergence and efficiency. Peng Yi 0001, Lacra Pavel |
IEEE Trans. Cybern. | 1 |
| 2019 | Multiple Timescale Online Learning Rules for Information Maximization with Energetic ConstraintsabstractA key aspect of the neural coding problem is understanding how representations of afferent stimuli are built through the dynamics of learning and adaptation within neural networks. The infomax paradigm is built on the premise that such learning attempts to maximize the mutual information between input stimuli and neural activities. In this letter, we tackle the problem of such information-based neural coding with an eye toward two conceptual hurdles. Specifically, we examine and then show how this form of coding can be achieved with online input processing. Our framework thus obviates the biological incompatibility of optimization methods that rely on global network awareness and batch processing of sensory signals. Central to our result is the use of variational bounds as a surrogate objective function, an established technique that has not previously been shown to yield online policies. We obtain learning dynamics for both linear-continuous and discrete spiking neural encoding models under the umbrella of linear gaussian decoders. This result is enabled by approximating certain information quantities in terms of neuronal activity via pairwise feedback mechanisms. Furthermore, we tackle the problem of how such learning dynamics can be realized with strict energetic constraints. We show that endowing networks with auxiliary variables that evolve on a slower timescale can allow for the realization of saddle-point optimization within the neural dynamics, leading to neural codes with favorable properties in terms of both information and energy. Peng Yi 0001, ShiNung Ching |
Neural Comput. | 1 |
| 2018 | Privacy Preservation in Distributed Subgradient Optimization AlgorithmsabstractIn this paper, some privacy-preserving features for distributed subgradient optimization algorithms are considered. Most of the existing distributed algorithms focus mainly on the algorithm design and convergence analysis, but not the protection of agents' privacy. Privacy is becoming an increasingly important issue in applications involving sensitive information. In this paper, we first show that the distributed subgradient synchronous homogeneous-stepsize algorithm is not privacy preserving in the sense that the malicious agent can asymptotically discover other agents' subgradients by transmitting untrue estimates to its neighbors. Then a distributed subgradient asynchronous heterogeneous-stepsize projection algorithm is proposed and accordingly its convergence and optimality is established. In contrast to the synchronous homogeneous-stepsize algorithm, in the new algorithm agents make their optimization updates asynchronously with heterogeneous stepsizes. The introduced two mechanisms of projection operation and asynchronous heterogeneous-stepsize optimization can guarantee that agents' privacy can be effectively protected. Youcheng Lou, Lean Yu, Shou-Yang Wang, Peng Yi 0001 |
IEEE Trans. Cybern. | 4 |
| 2017 | Distributed Optimization Design of Continuous-Time Multiagent Systems With Unknown-Frequency DisturbancesabstractIn this paper, a distributed optimization problem is studied for continuous-time multiagent systems with unknown-frequency disturbances. A distributed gradient-based control is proposed for the agents to achieve the optimal consensus with estimating unknown frequencies and rejecting the bounded disturbance in the semi-global sense. Based on convex optimization analysis and adaptive internal model approach, the exact optimization solution can be obtained for the multiagent system disturbed by exogenous disturbances with uncertain parameters. Xinghu Wang, Yiguang Hong, Peng Yi 0001, Haibo Ji, Yu Kang 0001 |
IEEE Trans. Cybern. | 3 |