VLDB 2026 Research / reviewers in the wild / expert
Pei Xu 0003
dblp:125/0711-3
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-1766-5634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | No-Regret Strategy Solving in Imperfect-Information Games via Pre-Trained EmbeddingabstractHigh-quality information set abstraction remains a core challenge in solving large-scale imperfect-information extensive-form games (IIEFGs)--such as no-limit Texas Hold’em--where the finite nature of spatial resources hinders solving strategies for the full game. State-of-the-art AI methods rely on pre-trained discrete clustering for abstraction, yet their hard classification irreversibly discards critical information: specifically, the quantifiable subtle differences between information sets--vital for strategy solving--thus compromising the quality of such solving. Inspired by the word embedding paradigm in natural language processing, this paper proposes the Embedding CFR algorithm, a novel approach for solving strategies in IIEFGs within an embedding space. The algorithm pre-trains and embeds the features of individual information sets into an interconnected low-dimensional continuous space, where the resulting vectors more precisely capture both the distinctions and connections between information sets. Embedding CFR introduces a strategy-solving process driven by regret accumulation and strategy updates in this embedding space, with supporting theoretical analysis verifying its ability to reduce cumulative regret. Experiments on poker show that with the same spatial overhead, Embedding CFR achieves significantly faster exploitability convergence compared to cluster-based abstraction algorithms, confirming its effectiveness. Furthermore, to our knowledge, it is the first algorithm in poker AI that pre-trains information set abstractions via low-dimensional embedding for strategy solving. Yanchang Fu, Shengda Liu, Pei Xu 0003, Kaiqi Huang |
AAAI | 3 |
| 2026 | RefRea: Reference-Guided Reasoning with Meta-Cognition for Accurate Language Model AgentsabstractIn recent years, with the rapid development of large language models (LLMs), LLM-based agents have achieved remarkable progress across a wide range of tasks. However, reasoning inconsistencies in LLMs still significantly limit the performance of agents in complex decision-making scenarios. Cognitive science research suggests that individuals can benefit from observing others' explicit thinking processes to improve their strategy-making. Inspired by this mechanism, we propose Reference-guided Reasoning with meta-cognition (RefRea), a novel approach that enhances decision-making by introducing a reference language model to guide and calibrate the reasoning model's actions. RefRea enhances reasoning accuracy and stability by integrating a reference model and a meta-cognition module. The reference model relies solely on validated meta-cognition for consistent guidance, while the reasoning model interacts with the environment using both validated and exploratory meta-cognition. Guidance is provided by comparing the action similarity between the reference and reasoning models. This process is supported by the meta-cognition module, which generates summary knowledge by reflecting on action history and environmental feedback, leading to more adaptive and reliable behavior. We evaluate our algorithm in the text-based reasoning environment ScienceWorld. Experimental results demonstrate that RefRea outperforms state-of-the-art methods. Comprehensive ablation studies further highlight the effectiveness of both the reference model and the meta-cognition module. Yuxiang Mai, Qiyue Yin, Wancheng Ni, Xiaogang Ouyang, Pei Xu 0003, Kaiqi Huang |
AAAI | 6 |
| 2025 | Uncertainty-Aware Opponent Modeling for Deep Reinforcement Learning
Likun Yang, Pei Xu 0003, Shiyue Cao, Xiaotang Chen, Kaiqi Huang |
AAMAS | 2 |
| 2025 | Constructive Conflict-Driven Multi-Agent Reinforcement Learning for Strategic DiversityabstractIn recent years, diversity has emerged as a useful mechanism to enhance the efficiency of multi-agent reinforcement learning (MARL). However, existing methods predominantly focus on designing policies based on individual agent characteristics, often neglecting the interplay and mutual influence among agents during policy formation. To address this gap, we propose Competitive Diversity through Constructive Conflict (CoDiCon), a novel approach that incorporates competitive incentives into cooperative scenarios to encourage policy exchange and foster strategic diversity among agents. Drawing inspiration from sociological research, which highlights the benefits of moderate competition and constructive conflict in group decision-making, we design an intrinsic reward mechanism using ranking features to introduce competitive motivations. A centralized intrinsic reward module generates and distributes varying reward values to agents, ensuring an effective balance between competition and cooperation. By optimizing the parameterized centralized reward module to maximize environmental rewards, we reformulate the constrained bilevel optimization problem to align with the original task objectives. We evaluate our algorithm against state-of-the-art methods in the SMAC and GRF environments. Experimental results demonstrate that CoDiCon achieves superior performance, with competitive intrinsic rewards effectively promoting diverse and adaptive strategies among cooperative agents. Yuxiang Mai, Qiyue Yin, Wancheng Ni, Pei Xu 0003, Kaiqi Huang |
IJCAI | 4 |
| 2025 | Learning differentiable categorical regions with Gumbel-Softmax for person re-identification
Wenjie Yang 0005, Pei Xu 0003 |
Neurocomputing | 2 |
| 2025 | Learning Individual Potential-Based Rewards in Multiagent Reinforcement LearningabstractA great challenge for applying multiagent reinforcement learning (MARL) in the field of game artificial intelligence (AI) is to enable agents to learn diversified policies to handle different game-specific problems, while receiving only a shared team reward. At present, a common approach is reward shaping, which focuses on designing rewards for agents to guide cooperation. However, most of the existing methods require prior knowledge on the environment for reward design or alter the optimal policies after imposing extra rewards. Besides, previous MARL methods that rely on manually designed rewards can hardly generalize across different game environments. To this end, we propose a new MARL method that learns individual potential-based rewards for agents. Specifically, we learn a parameterized potential function for each agent to generate individual rewards in the discounted temporal difference form. The whole update procedure is modeled as the bilevel optimization problem, where the lower level is to optimize policies with potential-based rewards, and the upper level is to optimize parameterized potential functions toward maximizing the environment return. We theoretically prove that the individual potential-based rewards can guarantee policy invariance for agents, so that the optimization objective is consistent with the original MARL problem. We evaluate our method with a number of existing state-of-the-art MARL methods on predator–prey andStarCraft IIgame environments. Empirical results show that our proposed method significantly outperforms baseline methods and achieves better game AI that enjoys high performance and generalization. Pei Xu 0003, Junge Zhang |
IEEE Trans. Games | 2 |
| 2025 | Exploration via Embracing Diversity in Reinforcement Learning for Sparse-Reward Procedurally-Generated TasksabstractA key challenge in reinforcement learning is how to guide agents to efficiently explore sparse reward environments. In order to overcome this challenge, the state-of-the-art methods introduce additional intrinsic rewards based on state-related information, such as the novelty of states. Unfortunately, these methods frequently fail in procedurally-generated tasks, where a different environment is generated in each episode so that the agent is not likely to visit the same state more than once. Recently, some exploration methods designed specifically for procedurally-generated tasks have been proposed. However, they still only consider state-related information, which leads to relatively inefficient exploration. In this work, we propose a novel exploration method, which utilizes cross-episode policy-related information and intraepisode state-related information to jointly encourage exploration in procedurally-generated tasks. In term of policy-related information, we first use an imitator-based unbalanced policy diversity to measure the difference between the agent’s current policy and the agent’s previous policies, and then encourage the agent to maximize this difference. In term of state-related information, we encourage the agent to maximize the state diversity within an episode, thereby visiting as many different states as possible in an episode. We show that our method significantly improves sample efficiency over state-of-the-art methods on three challenging benchmarks, including MiniGrid, MiniWorld, and the sparse-reward version of Procgen. Pei Xu 0003, Hao Chen 0103, Wenjie Yang 0005, Kaiqi Huang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | ADMN: Agent-Driven Modular Network for Dynamic Parameter Sharing in Cooperative Multi-Agent Reinforcement Learning
Yang Yu 0056, Qiyue Yin, Junge Zhang, Pei Xu 0003, Kaiqi Huang |
IJCAI | 4 |
| 2024 | Population-Based Diverse Exploration for Sparse-Reward Multi-Agent Tasks
Pei Xu 0003, Junge Zhang, Kaiqi Huang |
IJCAI | 1 |
| 2023 | Subspace-Aware Exploration for Sparse-Reward Multi-Agent TasksabstractExploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. One possible solution to this issue is to exploit inherent task structures for an acceleration of exploration. In this paper, we present a novel exploration approach, which encodes a special structural prior on the reward function into exploration, for sparse-reward multi-agent tasks. Specifically, a novel entropic exploration objective which encodes the structural prior is proposed to accelerate the discovery of rewards. By maximizing the lower bound of this objective, we then propose an algorithm with moderate computational cost, which can be applied to practical tasks. Under the sparse-reward setting, we show that the proposed algorithm significantly outperforms the state-of-the-art algorithms in the multiple-particle environment, the Google Research Football and StarCraft II micromanagement tasks. To the best of our knowledge, on some hard tasks (such as 27m_vs_30m}) which have relatively larger number of agents and need non-trivial strategies to defeat enemies, our method is the first to learn winning strategies under the sparse-reward setting. Pei Xu 0003, Junge Zhang, Qiyue Yin, Chao Yu 0004, Yaodong Yang 0001, Kaiqi Huang |
AAAI | 1 |
| 2023 | Exploration via Joint Policy Diversity for Sparse-Reward Multi-Agent TasksabstractExploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. Previous works argue that complex dynamics between agents and the huge exploration space in MARL scenarios amplify the vulnerability of classical count-based exploration methods when combined with agents parameterized by neural networks, resulting in inefficient exploration. In this paper, we show that introducing constrained joint policy diversity into a classical count-based method can significantly improve exploration when agents are parameterized by neural networks. Specifically, we propose a joint policy diversity to measure the difference between current joint policy and previous joint policies, and then use a filtering-based exploration constraint to further refine the joint policy diversity. Under the sparse-reward setting, we show that the proposed method significantly outperforms the state-of-the-art methods in the multiple-particle environment, the Google Research Football, and StarCraft II micromanagement tasks. To the best of our knowledge, on the hard 3s_vs_5z task which needs non-trivial strategies to defeat enemies, our method is the first to learn winning strategies without domain knowledge under the sparse-reward setting. Pei Xu 0003, Junge Zhang, Kaiqi Huang |
IJCAI | 1 |
| 2022 | Deep Reinforcement Learning With Part-Aware Exploration Bonus in Video GamesabstractReinforcement learning algorithms rely on carefully engineering environment rewards that are extrinsic to agents. However, environments with dense rewards are rare, motivating the need for developing reward functions that are intrinsic to agents. Curiosity is a type of successful intrinsic reward function, which uses the prediction error as an reward signal. In prior work, the prediction problem used to generate intrinsic rewards is optimized in the pixel space rather than a learnable feature space to avoid randomness caused by feature changes. However, these methods ignore small but important elements of the states that are often associated with locations of the character, which makes it impossible to generate accurate internal rewards for efficient exploration. In this article, we first demonstrate the effectiveness of introducing prior learned features for existing prediction-based exploration methods. Then, an attention map mechanism is designed to discretize learned features, thereby updating the learned feature and meanwhile reducing the impact of randomness on intrinsic rewards caused by the learning process of features. We verify our method on some video games from the standard reinforcement learning Atari benchmark, achieving clear improvements over random network distillation, which is one of the most advanced exploration methods, in almost all Atari games. Pei Xu 0003, Qiyue Yin, Junge Zhang, Kaiqi Huang |
IEEE Trans. Games | 1 |
| 2018 | Densely Connected Single-Shot DetectorabstractOne-stage object detection approach which utilizes multi-scale feature maps to predict objects is currently the best real-time detector. However, in this approach, the high-resolution feature maps which are responsible for detecting small objects are harder to learn a proper abstraction of objects than the low-resolution feature maps. The problem is that these feature maps have to transform sufficient low-level information to the next layer while learning high-level abstraction. In this paper, we develop a transformation module which adopts the dense structure to simplify the learning problem of high-resolution feature maps. In addition, we utilize the inception module to enrich the representation power of high-resolution feature maps. Extensive experiments on most object detection datasets clearly demonstrate the effectiveness of our method. In particular, on PASCAL VOC 2007/2012, our method outperforms all the existing one-stage methods. Our model based on the VGG-16 network also achieves competitive result on MS COCO. Pei Xu 0003, Xin Zhao 0012, Kaiqi Huang |
ICPR | 1 |