Xin Yu 0009

dblp:54/1184-9 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-3354-8625ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structural entropy guided hierarchical symmetric multi-agent reinforcement learning
Yongkai Tian, Xin Yu 0009, Yirong Qi, Li Wang 0170, Pu Feng, Wenjun Wu 0001, Rongye Shi, Jie Luo 0004
Expert Syst. Appl.2
2026 Embedded mean field reinforcement learning for perimeter-defense game
Li Wang 0170, Xin Yu 0009, Xuxin Lv, Gangzheng Ai, Wenjun Wu 0001
Expert Syst. Appl.2
2025 Symmetry-Guided Multi-Agent Inverse Reinforcement Learning
abstract
In robotic systems, the performance of reinforcement learning depends on the rationality of predefined reward functions. However, manually designed reward functions often lead to policy failures due to inaccuracies. Inverse Reinforcement Learning (IRL) addresses this problem by inferring implicit reward functions from expert demonstrations. Nevertheless, existing methods rely heavily on large amounts of expert demonstrations to accurately recover the reward function. The high cost of collecting expert demonstrations in robotic applications, particularly in multi-robot systems, severely hinders the practical deployment of IRL. Consequently, improving sample efficiency has emerged as a critical challenge in multi-agent inverse reinforcement learning (MIRL). Inspired by the symmetry inherent in multi-agent systems, this work theoretically demonstrates that leveraging symmetry enables the recovery of more accurate reward functions. Building upon this insight, we propose a universal framework that integrates symmetry into existing multi-agent adversarial IRL algorithms, thereby significantly enhancing sample efficiency. Experimental results from multiple challenging tasks have demonstrated the effectiveness of this framework. Further validation in physical multi-robot systems has shown the practicality of our method.
Yongkai Tian, Yirong Qi, Xin Yu 0009, Wenjun Wu 0001, Jie Luo 0004
IROS3
2025 CLGA: A Collaborative LLM Framework for Dynamic Goal Assignment in Multi-Robot Systems
abstract
Goal assignment is a critical challenge in multi-robot systems. The emergence of large language models (LLMs) has enabled the use of natural language commands for tackling goal assignment problems. However, applying LLMs directly to these tasks presents two limitations: 1) limited accuracy and 2) excessive decision delays due to their autoregressive nature, hindering adaptability to unexpected changes. To address these issues, inspired by dual-process theory, we propose a framework called Collaborative LLMs for dynamic Goal Assignment (CLGA). Specifically, we leverage LLMs for pre-planning tasks and invoke an external solver to generate an initial goal assignment solution, ensuring solution accuracy. During execution, small-scale models enable real-time adjustments to respond to dynamic environmental changes. This approach integrates the strengths of slow, precise pre-planning and fast, adaptive online adjustments, allowing agents to efficiently handle real-world challenges. Additionally, we introduce a benchmark dataset for NLP-based goal assignment to advance research in this domain. Simulation and real-world experiments demonstrate that CLGA significantly enhances task execution efficiency and flexibility in multi-robot systems. The prompt, experimental videos, and datasets associated with this work are available at https://sites.google.com/view/project-clga/.
Xin Yu 0009, Yandong Wang 0002, Rongye Shi, Gangzheng Ai, Zhiqiang Pu, Wenjun Wu 0001
IROS1
2025 Empirical Study on Robustness and Resilience in Cooperative Multi-Agent Reinforcement Learning
abstract
In cooperative Multi-Agent Reinforcement Learning (MARL), it is a common practice to tune hyperparameters in ideal simulated environments to maximize cooperative performance. However, policies tuned for cooperation often fail to maintain robustness and resilience under real-world uncertainties. Building trustworthy MARL systems requires a deep understanding of \emph{robustness}, which ensures stability under uncertainties, and \emph{resilience}, the ability to recover from disruptions—a concept extensively studied in control systems but largely overlooked in MARL. In this paper, we present a large-scale empirical study comprising over 82,620 experiments to evaluate cooperation, robustness, and resilience in MARL across 4 real-world environments, 13 uncertainty types, and 15 hyperparameters. Our key findings are: (1) Under mild uncertainty, optimizing cooperation improves robustness and resilience, but this link weakens as perturbations intensify. Robustness and resilience also varies by algorithm and uncertainty type. (2) Robustness and resilience do not generalize across uncertainty modalities or agent scopes: policies robust to action noise for all agents may fail under observation noise on a single agent. (3) Hyperparameter tuning is critical for trustworthy MARL: surprisingly, standard practices like parameter sharing, GAE, and PopArt can hurt robustness, while early stopping, high critic learning rates, and Leaky ReLU consistently help. By optimizing hyperparameters only, we observe substantial improvement in cooperation, robustness and resilience across all MARL backbones, with the phenomenon also generalizing to robust MARL methods across these backbones.
Zihao Mao, Zonglei Jing, Zhuohang bian, Jun Guo 0009, Li Wang 0170, Zhuoran Han, Ruixiao Xu, Xin Yu 0009, Chengdong Ma, Yuqing Ma, Bo An 0001, Yaodong Yang 0001, Weifeng Lv, Xianglong Liu 0001
NeurIPS10
2025 Attacking cooperative multi-agent reinforcement learning by adversarial minority influence
Jun Guo 0009, Jingqiao Xiu, Yuwei Zheng, Pu Feng, Xin Yu 0009, Jiakai Wang, Aishan Liu, Yaodong Yang 0001, Bo An 0001, Wenjun Wu 0001, Xianglong Liu 0001
Neural Networks6
2025 Lyapunov-Informed Multi-Agent Reinforcement Learning for Multi-Robot Cooperation Tasks
abstract
Multi-Agent Reinforcement Learning (MARL) has shown great potential in solving complex tasks. Despite great success, low training efficiency remains a pervasive and long-standing challenge in MARL. To tackle this issue, it is promising to leverage prior knowledge or environmental properties to inform and improve the MARL. We notice that many multi-agent tasks specify certain goal states where special rewards are granted, guiding agents to achieve the goal. Inspired by the theory of Lyapunov stability, an intuitive optimal policy to the tasks should be able to asymptotically converge to the goal states from any initial, making the goal states stable equilibria. Focusing on this type of tasks, we introduce the concept of Lyapunov Markov game (LMG), a new subclass of the cooperative Markov game, featuring a set of goal states and goal-oriented reward function. We then provide a theoretical bound on scaled value distance as a necessary condition to obtain a stable suboptimal policy in LMG. Motivated by this insight, we further propose the Lyapunov-informed MARL, which leverages a newly-designed Lyapunov-informed reward. Theoretical work is conducted to show that the Lyapunov-informed MARL enjoys a broadened bound, facilitating the training process to find a stable suboptimal policy more easily and then converge to an optimal policy more efficiently. Extensive experiments and real-world multi-robot implementations are conducted to show the superior performance of the proposed approach over advanced baseline models.
Pu Feng, Rongye Shi, Size Wang, Qizhen Wu, Xin Yu 0009, Wenjun Wu 0001
IEEE Trans Autom. Sci. Eng.5
2025 Symmetry-Informed MARL: A Decentralized and Cooperative UAV Swarm Control Approach for Communication Coverage
abstract
Uncrewed aerial vehicle-mounted base stations (UAV-MBSs) provide flexible wireless connectivity, extending communication coverage in underserved areas. Recently, multi-agent reinforcement learning (MARL) has shown great potential for cooperative UAV swarm control to support efficient communication coverage in dynamic and complex environments. However, existing MARL-based methods often suffer from low sample efficiency due to its trial-and-error training characteristics, limiting its ability to control large UAV swarms with continuous state-action space and partial observation. We notice that UAV swarm systems in communication coverage tasks exhibit a spatial symmetry property, e.g., a rotation in the spatial observation of a UAV results in a same rotation in its optimal action. Exploiting this property, we formulate the task as a symmetric decentralized partially observable Markov decision process and introduce symmetry-informed MARL, featuring a novel network called the symmetry-informed graph neural network (SiGNN) to serve as the policy/value networks. SiGNN leverages the inherent symmetry in multi-UAV systems by embedding the symmetry into the network structure, thereby enhancing the training efficiency to handle large swarms with continuous control. Theoretical analysis shows that the SiGNN strictly preserves symmetry properties, which guarantees the effectiveness of the approach. Experiments in simulation were conducted to handle communication coverage using up to 20 UAVs with continuous control. Experimental results demonstrate that SiGNN-based MARL outperforms advanced baselines, verifying its superior sample efficiency, scalability and robustness.
Rongye Shi, Xin Yu 0009, Yandong Wang 0002, Yongkai Tian, Zhenyu Liu 0003, Wenjun Wu 0001, Xiao-Ping Zhang 0002, Manuela M. Veloso
IEEE Trans. Mob. Comput.2
2024 Leveraging Partial Symmetry for Multi-Agent Reinforcement Learning
abstract
Incorporating symmetry as an inductive bias into multi-agent reinforcement learning (MARL) has led to improvements in generalization, data efficiency, and physical consistency. While prior research has succeeded in using perfect symmetry prior, the realm of partial symmetry in the multi-agent domain remains unexplored. To fill in this gap, we introduce the partially symmetric Markov game, a new subclass of the Markov game. We then theoretically show that the performance error introduced by utilizing symmetry in MARL is bounded, implying that the symmetry prior can still be useful in MARL even in partial symmetry situations. Motivated by this insight, we propose the Partial Symmetry Exploitation (PSE) framework that is able to adaptively incorporate symmetry prior in MARL under different symmetry-breaking conditions. Specifically, by adaptively adjusting the exploitation of symmetry, our framework is able to achieve superior sample efficiency and overall performance of MARL algorithms. Extensive experiments are conducted to demonstrate the superior performance of the proposed framework over baselines. Finally, we implement the proposed framework in real-world multi-robot testbed to show its superiority.
Xin Yu 0009, Rongye Shi, Pu Feng, Yongkai Tian, Shuhao Liao, Wenjun Wu 0001
AAAI1
2024 Exploiting Hierarchical Symmetry in Multi-Agent Reinforcement Learning
abstract
Achieving high sample efficiency is a critical research area in reinforcement learning. This becomes extremely difficult in multi-agent reinforcement learning (MARL), as the capacity of the joint state and action space grows exponentially with the number of agents. The reliance of MARL solely on exploration and trial-and-error, without incorporating prior knowledge, exacerbates the issue of low sample efficiency. Currently, introducing symmetry into MARL is an effective approach to address this issue. Yet the concept of hierarchical symmetry, which maintains symmetry across different levels of a multi-agent system (MAS), has not been explored in existing methods. This paper focuses on multi-agent cooperative tasks and proposes a method incorporating hierarchical symmetry, termed the Hierarchical Equivariant Policy Network (HEPN) which is O(n)-equivariant. Specifically, HEPN utilizes clustering to perform hierarchical information extraction in MAS, and employs graph neural networks to model agent interactions. We conducted extensive experiments across various multi-agent tasks. The results indicate that our method achieves faster convergence speeds and higher convergence rewards compared to baseline algorithms. Additionally, we have deployed our algorithm in a physical multi-robot system, confirming its effectiveness in real-world environments. Supplementary materials are available at https://yongkai-tian.github.io/HEPN/.
Yongkai Tian, Xin Yu 0009, Yirong Qi, Li Wang 0170, Pu Feng, Wenjun Wu 0001, Rongye Shi, Jie Luo 0004
ECAI2
2024 GraphRARE: Reinforcement Learning Enhanced Graph Neural Network with Relative Entropy
abstract
Graph neural networks (GNNs) have shown ad-vantages in graph-based analysis tasks. However, most existing methods have the homogeneity assumption and show poor performance on heterophilic graphs, where the linked nodes have dissimilar features and different class labels, and the semantically related nodes might be multi-hop away. To address this limitation, this paper presents GraphRARE, a general framework built upon node relative entropy and deep reinforcement learning, to strengthen the expressive capability of GNNs. An innovative node relative entropy, which considers node features and structural similarity, is used to measure mutual information between node pairs. In addition, to avoid the sub-optimal solutions caused by mixing useful information and noises of remote nodes, a deep reinforcement learning-based algorithm is developed to optimize the graph topology. This algorithm selects informative nodes and discards noisy nodes based on the defined node relative en-tropy. Extensive experiments are conducted on seven real-world datasets. The experimental results demonstrate the superiority of GraphRARE in node classification and its capability to optimize the original graph topology.
Tianhao Peng 0002, Wenjun Wu 0001, Haitao Yuan 0002, Zhifeng Bao, Zhao Pengrui, Xin Yu 0009, Xuetao Lin, Yu Liang 0003, Yanjun Pu
ICDE6
2024 Byzantine Robust Cooperative Multi-Agent Reinforcement Learning as a Bayesian Game
abstract
In this study, we explore the robustness of cooperative multi-agent reinforcement learning (c-MARL) against Byzantine failures, where any agent can enact arbitrary, worst-case actions due to malfunction or adversarial attack. To address the uncertainty that any agent can be adversarial, we propose a Bayesian Adversarial Robust Dec-POMDP (BARDec-POMDP) framework, which views Byzantine adversaries as nature-dictated types, represented by a separate transition. This allows agents to learn policies grounded on their posterior beliefs about the type of other agents, fostering collaboration with identified allies and minimizing vulnerability to adversarial manipulation. We define the optimal solution to the BARDec-POMDP as an ex interim robust Markov perfect Bayesian equilibrium, which we proof to exist and the corresponding policy weakly dominates previous approaches as time goes to infinity. To realize this equilibrium, we put forward a two-timescale actor-critic algorithm with almost sure convergence under specific conditions. Experiments on matrix game, Level-based Foraging and StarCraft II indicate that, our method successfully acquires intricate micromanagement skills and adaptively aligns with allies under worst-case perturbations, showing resilience against non-oblivious adversaries, random allies, observation-based attacks, and transfer-based attacks.
Jun Guo 0009, Jingqiao Xiu, Ruixiao Xu, Xin Yu 0009, Jiakai Wang, Aishan Liu, Yaodong Yang 0001, Xianglong Liu 0001
ICLR5
2024 AdaptAUG: Adaptive Data Augmentation Framework for Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning has emerged as a promising approach for the control of multi-robot systems. Nevertheless, the low sample efficiency of MARL poses a significant obstacle to its broader application in robotics. While data augmentation appears to be a straightforward solution for improving sample efficiency, it usually incurs training instability, making the sample efficiency worse. Moreover, manually choosing suitable augmentations for a variety of tasks is a tedious and time-consuming process. To mitigate these challenges, our research theoretically analyzes the implications of data augmentation on MARL algorithms. Guided by these insights, we present AdaptAUG, an adaptive framework designed to selectively identify beneficial data augmentations, thereby achieving superior sample efficiency and overall performance in multi-robot tasks. Extensive experiments in both simulated and real-world multi-robot scenarios validate the effectiveness of our proposed framework.
Xin Yu 0009, Yongkai Tian, Li Wang 0170, Pu Feng, Wenjun Wu 0001, Rongye Shi
ICRA1
2024 Hierarchical Consensus-Based Multi-Agent Reinforcement Learning for Multi-Robot Cooperation Tasks
abstract
In multi-agent reinforcement learning (MARL), the Centralized Training with Decentralized Execution (CTDE) framework is pivotal but struggles due to a gap: global state guidance in training versus reliance on local observations in execution, lacking global signals. Inspired by human societal consensus mechanisms, we introduce the Hierarchical Consensus-based Multi-Agent Reinforcement Learning (HC-MARL) framework to address this limitation. HC-MARL employs contrastive learning to foster a global consensus among agents, enabling cooperative behavior without direct communication. This approach enables agents to form a global consensus from local observations, using it as an additional piece of information to guide collaborative actions during execution. To cater to the dynamic requirements of various tasks, consensus is divided into multiple layers, encompassing both short-term and long-term considerations. Short-term observations prompt the creation of an immediate, low-layer consensus, while long-term observations contribute to the formation of a strategic, high-layer consensus. This process is further refined through an adaptive attention mechanism that dynamically adjusts the influence of each consensus layer. This mechanism optimizes the balance between immediate reactions and strategic planning, tailoring it to the specific demands of the task at hand. Extensive experiments and real-world applications in multi-robot systems showcase our framework’s superior performance, marking significant advancements over baselines.
Pu Feng, Junkang Liang, Size Wang, Xin Yu 0009, Xin Ji, Rongye Shi, Wenjun Wu 0001
IROS4
2023 ESP: Exploiting Symmetry Prior for Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has achieved promising results in recent years. However, most existing reinforcement learning methods require a large amount of data for model training. In addition, data-efficient reinforcement learning requires the construction of strong inductive biases, which are ignored in the current MARL approaches. Inspired by the symmetry phenomenon in multi-agent systems, this paper proposes a framework for exploiting prior knowledge by integrating data augmentation and a well-designed consistency loss into the existing MARL methods. In addition, the proposed framework is model-agnostic and can be applied to most of the current MARL algorithms. Experimental tests on multiple challenging tasks demonstrate the effectiveness of the proposed framework. Moreover, the proposed framework is applied to a physical multi-robot testbed to show its superiority.
Xin Yu 0009, Rongye Shi, Pu Feng, Yongkai Tian, Jie Luo 0004, Wenjun Wu 0001
ECAI1
2021 Swarm Inverse Reinforcement Learning for Biological Systems
abstract
Complex global behavior can emerge from local interactions in biological systems. Many models have been introduced to describe the interaction rules of biological individuals. Nonetheless, most research efforts cannot capture the inner cognitive and sequential decision process of individual animals in their swarms. In this paper, we formulate this problem as homogeneous Markov game and focus on identifying the potential reward function of individual animals so as to understand their collective behaviors. We propose an inverse reinforcement learning method PS-AIRL specifically for biological systems, where the parameter sharing paradigm is combined with a deep inverse reinforcement learning. Theoretical analysis and experimental evaluation show that PS-AIRL can learn the policy and the reward function from collective behavior demonstrations. Moreover, our methods can be applied to a wide range of biological behavioral studies.
Xin Yu 0009, Wenjun Wu 0001, Pu Feng, Yongkai Tian
BIBM1