Pu Feng

dblp:295/8419 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-6219-1741ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structural entropy guided hierarchical symmetric multi-agent reinforcement learning
Yongkai Tian, Xin Yu 0009, Yirong Qi, Li Wang 0170, Pu Feng, Wenjun Wu 0001, Rongye Shi, Jie Luo 0004
Expert Syst. Appl.5
2026 Offline constrained policy optimization with safe anchoring
Diyuan Hou, Longyang Huang, Pu Feng, Wenjun Wu 0001
Neural Networks3
2025 Neural Algorithmic Reasoners informed Large Language Model for Multi-Agent Path Finding
abstract
The development and application of large language models (LLM) have demonstrated that foundational models can be utilized to solve a wide array of tasks. However, their performance in multi-agent path finding (MAPF) tasks has been less than satisfactory, with only a few studies exploring this area. MAPF is a complex problem requiring both planning and multi-agent coordination. To improve the performance of LLM in MAPF tasks, we propose a novel framework, LLM-NAR, which leverages neural algorithmic reasoners (NAR) to inform LLM for MAPF. LLM-NAR consists of three key components: an LLM for MAPF, a pre-trained graph neural network-based NAR, and a cross-attention mechanism. This is the first work to propose using a neural algorithmic reasoner to integrate GNNs with the map information for MAPF, thereby guiding LLM to achieve superior performance. LLM-NAR can be easily adapted to various LLM models. Both simulation and real-world experiments demonstrate that our method significantly outperforms existing LLM-based approaches in solving MAPF problems.
Pu Feng, Size Wang, Yuhong Cao, Junkang Liang, Rongye Shi, Wenjun Wu 0001
IJCNN1
2025 Attacking cooperative multi-agent reinforcement learning by adversarial minority influence
Jun Guo 0009, Jingqiao Xiu, Yuwei Zheng, Pu Feng, Xin Yu 0009, Jiakai Wang, Aishan Liu, Yaodong Yang 0001, Bo An 0001, Wenjun Wu 0001, Xianglong Liu 0001
Neural Networks5
2025 Lyapunov-Informed Multi-Agent Reinforcement Learning for Multi-Robot Cooperation Tasks
abstract
Multi-Agent Reinforcement Learning (MARL) has shown great potential in solving complex tasks. Despite great success, low training efficiency remains a pervasive and long-standing challenge in MARL. To tackle this issue, it is promising to leverage prior knowledge or environmental properties to inform and improve the MARL. We notice that many multi-agent tasks specify certain goal states where special rewards are granted, guiding agents to achieve the goal. Inspired by the theory of Lyapunov stability, an intuitive optimal policy to the tasks should be able to asymptotically converge to the goal states from any initial, making the goal states stable equilibria. Focusing on this type of tasks, we introduce the concept of Lyapunov Markov game (LMG), a new subclass of the cooperative Markov game, featuring a set of goal states and goal-oriented reward function. We then provide a theoretical bound on scaled value distance as a necessary condition to obtain a stable suboptimal policy in LMG. Motivated by this insight, we further propose the Lyapunov-informed MARL, which leverages a newly-designed Lyapunov-informed reward. Theoretical work is conducted to show that the Lyapunov-informed MARL enjoys a broadened bound, facilitating the training process to find a stable suboptimal policy more easily and then converge to an optimal policy more efficiently. Extensive experiments and real-world multi-robot implementations are conducted to show the superior performance of the proposed approach over advanced baseline models.
Pu Feng, Rongye Shi, Size Wang, Qizhen Wu, Xin Yu 0009, Wenjun Wu 0001
IEEE Trans Autom. Sci. Eng.1
2025 Robust Multi-Agent Reinforcement Learning by Mutual Information Regularization
abstract
In cooperative multi-agent reinforcement learning (MARL), ensuring robustness against cooperative agents making unpredictable or worst-case adversarial actions is crucial for real-world deployment. In multi-agent settings, each agent may be perturbed or unperturbed, leading to an exponential increase in potential threat scenarios as the number of agents grows. Existing robust MARL methods either enumerate, or approximate all possible threat scenarios, leading to intense computation and insufficient robustness. In contrast, humans develop robust behaviors by maintaining a general level of caution rather than preparing for every possible threat. Inspired by human decision making, we frame robust MARL as a control-as-inference problem, and optimize worst-case robustness across all threat scenarios implicitly optimized through off-policy evaluation. Specifically, we introduce mutual information regularization as robust regularization (MIR3), which maximizes a lower bound on robustness during routine training, serving as a kind of caution for MARL without adversarial inputs. Further insights show that MIR3 acts as an information bottleneck, preventing agents from over-reacting to others and aligning policies with robust action priors. In the presence of worst-case adversaries, our MIR3 significantly surpasses baseline methods in robustness and training efficiency, and maintaining cooperative performance in StarCraft II, quadrotor swarm control, and robot swarm control. When deploying the robot swarm control algorithm in the real world, our method also outperforms the best baseline by 14.29% in reward. See code and demo videos at https://github.com/DIG-Beihang/MIR3.
Ruixiao Xu, Jingqiao Xiu, Yuwei Zheng, Pu Feng, Yuqing Ma, Bo An 0001, Yaodong Yang 0001, Xianglong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Leveraging Partial Symmetry for Multi-Agent Reinforcement Learning
abstract
Incorporating symmetry as an inductive bias into multi-agent reinforcement learning (MARL) has led to improvements in generalization, data efficiency, and physical consistency. While prior research has succeeded in using perfect symmetry prior, the realm of partial symmetry in the multi-agent domain remains unexplored. To fill in this gap, we introduce the partially symmetric Markov game, a new subclass of the Markov game. We then theoretically show that the performance error introduced by utilizing symmetry in MARL is bounded, implying that the symmetry prior can still be useful in MARL even in partial symmetry situations. Motivated by this insight, we propose the Partial Symmetry Exploitation (PSE) framework that is able to adaptively incorporate symmetry prior in MARL under different symmetry-breaking conditions. Specifically, by adaptively adjusting the exploitation of symmetry, our framework is able to achieve superior sample efficiency and overall performance of MARL algorithms. Extensive experiments are conducted to demonstrate the superior performance of the proposed framework over baselines. Finally, we implement the proposed framework in real-world multi-robot testbed to show its superiority.
Xin Yu 0009, Rongye Shi, Pu Feng, Yongkai Tian, Shuhao Liao, Wenjun Wu 0001
AAAI3
2024 Exploiting Hierarchical Symmetry in Multi-Agent Reinforcement Learning
abstract
Achieving high sample efficiency is a critical research area in reinforcement learning. This becomes extremely difficult in multi-agent reinforcement learning (MARL), as the capacity of the joint state and action space grows exponentially with the number of agents. The reliance of MARL solely on exploration and trial-and-error, without incorporating prior knowledge, exacerbates the issue of low sample efficiency. Currently, introducing symmetry into MARL is an effective approach to address this issue. Yet the concept of hierarchical symmetry, which maintains symmetry across different levels of a multi-agent system (MAS), has not been explored in existing methods. This paper focuses on multi-agent cooperative tasks and proposes a method incorporating hierarchical symmetry, termed the Hierarchical Equivariant Policy Network (HEPN) which is O(n)-equivariant. Specifically, HEPN utilizes clustering to perform hierarchical information extraction in MAS, and employs graph neural networks to model agent interactions. We conducted extensive experiments across various multi-agent tasks. The results indicate that our method achieves faster convergence speeds and higher convergence rewards compared to baseline algorithms. Additionally, we have deployed our algorithm in a physical multi-robot system, confirming its effectiveness in real-world environments. Supplementary materials are available at https://yongkai-tian.github.io/HEPN/.
Yongkai Tian, Xin Yu 0009, Yirong Qi, Li Wang 0170, Pu Feng, Wenjun Wu 0001, Rongye Shi, Jie Luo 0004
ECAI5
2024 AdaptAUG: Adaptive Data Augmentation Framework for Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning has emerged as a promising approach for the control of multi-robot systems. Nevertheless, the low sample efficiency of MARL poses a significant obstacle to its broader application in robotics. While data augmentation appears to be a straightforward solution for improving sample efficiency, it usually incurs training instability, making the sample efficiency worse. Moreover, manually choosing suitable augmentations for a variety of tasks is a tedious and time-consuming process. To mitigate these challenges, our research theoretically analyzes the implications of data augmentation on MARL algorithms. Guided by these insights, we present AdaptAUG, an adaptive framework designed to selectively identify beneficial data augmentations, thereby achieving superior sample efficiency and overall performance in multi-robot tasks. Extensive experiments in both simulated and real-world multi-robot scenarios validate the effectiveness of our proposed framework.
Xin Yu 0009, Yongkai Tian, Li Wang 0170, Pu Feng, Wenjun Wu 0001, Rongye Shi
ICRA4
2024 Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation
Yanni Xue, Haojie Hao, Jiakai Wang, Qiang Sheng 0001, Renshuai Tao, Pu Feng, Xianglong Liu 0001
IJCAI7
2024 Hierarchical Consensus-Based Multi-Agent Reinforcement Learning for Multi-Robot Cooperation Tasks
abstract
In multi-agent reinforcement learning (MARL), the Centralized Training with Decentralized Execution (CTDE) framework is pivotal but struggles due to a gap: global state guidance in training versus reliance on local observations in execution, lacking global signals. Inspired by human societal consensus mechanisms, we introduce the Hierarchical Consensus-based Multi-Agent Reinforcement Learning (HC-MARL) framework to address this limitation. HC-MARL employs contrastive learning to foster a global consensus among agents, enabling cooperative behavior without direct communication. This approach enables agents to form a global consensus from local observations, using it as an additional piece of information to guide collaborative actions during execution. To cater to the dynamic requirements of various tasks, consensus is divided into multiple layers, encompassing both short-term and long-term considerations. Short-term observations prompt the creation of an immediate, low-layer consensus, while long-term observations contribute to the formation of a strategic, high-layer consensus. This process is further refined through an adaptive attention mechanism that dynamically adjusts the influence of each consensus layer. This mechanism optimizes the balance between immediate reactions and strategic planning, tailoring it to the specific demands of the task at hand. Extensive experiments and real-world applications in multi-robot systems showcase our framework’s superior performance, marking significant advancements over baselines.
Pu Feng, Junkang Liang, Size Wang, Xin Yu 0009, Xin Ji, Rongye Shi, Wenjun Wu 0001
IROS1
2023 Towards Benchmarking and Assessing Visual Naturalness of Physical World Adversarial Attacks
abstract
Physical world adversarial attack is a highly practical and threatening attack, which fools real world deep learning systems by generating conspicuous and maliciously crafted real world artifacts. In physical world attacks, evaluating naturalness is highly emphasized since human can easily detect and remove unnatural attacks. However, current studies evaluate naturalness in a case-by-case fashion, which suffers from errors, bias and inconsistencies. In this paper, we take the first step to benchmark and assess visual naturalness of physical world attacks, taking autonomous driving scenario as the first attempt. First, to benchmark attack naturalness, we contribute the first Physical Attack Naturalness (PAN) dataset with human rating and gaze. PAN verifies several insights for the first time: naturalness is (disparately) affected by contextual features (i.e., environmental and semantic variations) and correlates with behavioral feature (i.e., gaze signal). Second, to automatically assess attack naturalness that aligns with human ratings, we further introduce Dual Prior Alignment (DPA) network, which aims to embed human knowledge into model reasoning process. Specifically, DPA imitates human reasoning in naturalness assessment by rating prior alignment and mimics human gaze behavior by attentive prior alignment. We hope our work fosters researches to improve and automatically assess naturalness of physical world attacks. Our code and dataset can be found at https://github.com/zhangsn-19/PAN.
Gujun Chen, Pu Feng, Jiakai Wang, Aishan Liu, Xin Yi 0001, Xianglong Liu 0001
CVPR5
2023 ESP: Exploiting Symmetry Prior for Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has achieved promising results in recent years. However, most existing reinforcement learning methods require a large amount of data for model training. In addition, data-efficient reinforcement learning requires the construction of strong inductive biases, which are ignored in the current MARL approaches. Inspired by the symmetry phenomenon in multi-agent systems, this paper proposes a framework for exploiting prior knowledge by integrating data augmentation and a well-designed consistency loss into the existing MARL methods. In addition, the proposed framework is model-agnostic and can be applied to most of the current MARL algorithms. Experimental tests on multiple challenging tasks demonstrate the effectiveness of the proposed framework. Moreover, the proposed framework is applied to a physical multi-robot testbed to show its superiority.
Xin Yu 0009, Rongye Shi, Pu Feng, Yongkai Tian, Jie Luo 0004, Wenjun Wu 0001
ECAI3
2021 Swarm Inverse Reinforcement Learning for Biological Systems
abstract
Complex global behavior can emerge from local interactions in biological systems. Many models have been introduced to describe the interaction rules of biological individuals. Nonetheless, most research efforts cannot capture the inner cognitive and sequential decision process of individual animals in their swarms. In this paper, we formulate this problem as homogeneous Markov game and focus on identifying the potential reward function of individual animals so as to understand their collective behaviors. We propose an inverse reinforcement learning method PS-AIRL specifically for biological systems, where the parameter sharing paradigm is combined with a deep inverse reinforcement learning. Theoretical analysis and experimental evaluation show that PS-AIRL can learn the policy and the reward function from collective behavior demonstrations. Moreover, our methods can be applied to a wide range of biological behavioral studies.
Xin Yu 0009, Wenjun Wu 0001, Pu Feng, Yongkai Tian
BIBM3
2021 AR-TV and AR-Diànshì: Cultural Differences in Users' Preferences for Augmented Reality Television
abstract
As Augmented Reality television gains momentum, it is important to understand whether cultural differences among viewers favor different expectations and preferences for immersion in such new television environments. A previous study documented the preferences of 172 participants from various European countries for twenty application scenarios for ARTV, such as virtual objects coming out of the TV screen into the room. In this work, we conduct an empirical generalization of this previous study to understand potential cultural differences in users’ preferences for and expectations of ARTV. To this end, we report insights from data collected from a sample of 147 participants from China, which we compare against the preferences expressed by the participants from Europe from the original study. Our findings reveal similarities, but also differences in terms of expectations of ARTV across the two cultural groups. We draw implications for future research on culturally-aware augmentations of the television watching experience.
Irina Popovici, Radu-Daniel Vatavu, Pu Feng, Wenjun Wu 0001
IMX3