VLDB 2026 Research / reviewers in the wild / expert
Xiaolin Ai
dblp:211/9844
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0001-7943-8336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement LearningabstractGuiding the policy of multi-agent reinforcement learning to align with human common sense is a difficult problem, largely due to the complexity of modeling common sense as a reward, especially in complex and long-horizon multi-agent tasks. Recent works have shown the effectiveness of reward shaping, such as potential-based rewards, to enhance policy alignment. The existing works, however, primarily rely on experts to design rule-based rewards, which are often labor-intensive and lack a high-level semantic understanding of common sense. To solve this problem, we propose a hierarchical vision-based reward shaping method. At the bottom layer, a visual-language model (VLM) serves as a generic potential function, guiding the policy to align with human common sense through its intrinsic semantic understanding. To help the policy adapts to uncertainty and changes in long-horizon tasks, the top layer features an adaptive skill selection module based on a visual large language model (vLLM). The module uses instructions, video replays, and training records to dynamically select suitable potential function from a pre-designed pool. Besides, our method is theoretically proven to preserve the optimal policy. Extensive experiments conducted in the Google Research Football environment demonstrate that our method not only achieves a higher win rate but also effectively aligns the policy with human common sense. Shijie Wang 0006, Zhiqiang Pu, Siyao Zhao, Xiaolin Ai |
AAAI | 5 |
| 2025 | OMoE: Diversifying Mixture of Low-Rank Adaptation by Orthogonal FinetuningabstractBuilding mixture-of-experts (MoE) architecture for low-rank adaptation (LoRA) is emerging as a potential direction in parameter-efficient fine-tuning (PEFT) for its modular design and remarkable performance. However, simply stacking the number of experts cannot guarantee significant improvement. In this work, we first conduct qualitative analysis to indicate that experts collapse to similar representations in vanilla MoE, limiting the capacity of modular design and computational efficiency. Ulteriorly, our analysis reveals that the performance of previous MoE variants may be limited by a lack of diversity among experts. Motivated by these findings, we propose Orthogonal Mixture-of-Experts (OMoE), a resource-efficient MoE variant that trains experts in an orthogonal manner to promote diversity. In OMoE, a Gram-Schmidt process is leveraged to enforce that the experts’ representations lie within the Stiefel manifold. By applying orthogonal constraints directly to the architecture, OMoE keeps the learning objective unchanged, without compromising optimality. Our method is simple and alleviates memory bottlenecks, as it incurs minimal experts compared to vanilla MoE models. Experiments on diverse benchmarks demonstrate that OMoE can consistently achieve stable and efficient performance improvement when compared with the state-of-the-art methods while significantly reducing the number of required experts. Jinyuan Feng, Zhiqiang Pu, Xiaolin Ai, Huimu Wang |
ECAI | 5 |
| 2025 | A Policy Resonance Approach to Solve the Problem of Responsibility Diffusion in Multiagent Reinforcement LearningabstractState-of-the-art (SOTA) multiagent reinforcement algorithms distinguish themselves in many ways from their single-agent equivalences. However, most of them still totally inherit the single-agent exploration-exploitation strategy. Naively inheriting this strategy from single-agent algorithms causes potential collaboration failures, in which the agents blindly follow mainstream behaviors and reject taking minority responsibility. We name this problem the responsibility diffusion (RD) as it shares similarities with the same-name social psychology effect. In this work, we start by theoretically analyzing the cause of this RD problem, which can be traced back to the exploration-exploitation dilemma of multiagent systems (especially large-scale multiagent systems). We address this RD problem by proposing a policy resonance (PR) approach which modifies the collaborative exploration strategy of agents by refactoring the joint agent policy while keeping individual policies approximately invariant. Next, we show that SOTA algorithms can equip this approach to promote the collaborative performance of agents in complex cooperative tasks. Experiments are performed in multiple test benchmark tasks to illustrate the effectiveness of this approach. Qingxu Fu, Tenghai Qiu, Jianqiang Yi, Zhiqiang Pu, Xiaolin Ai, Wanmai Yuan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Co-Opetition Network-Based Group Decision-Making Under Incomplete InformationabstractThe integration of cooperation and competition strategies in game theory emphasizes the systematic nature of strategy spaces and group interactions, forming the basis for achieving win-win scenarios. This is particularly crucial under coalition constraints and incomplete information. Addressing these challenges, this article introduces a comprehensive mathematical method for policy formation using co-opetition topological networks. This method enables autonomous decision-making and game equilibrium in group decision scenarios, considering individual preferences amidst constraints like incomplete information and alliance limitations. Leveraging the complementary entropy theorem on superiority, inferiority, and fuzzy measures, we propose a cognitive model for information interaction and attribute fusion. Utilizing the ordered weighted averaging operator and average tree solutions aids in identifying optimal alliance structures. We subsequently discuss evaluating missing information to complete the topological network. Updating the cognitive model and value function, we develop a Gaussian oscillation heuristic algorithm to explore alliance and component strategy spaces. Simulation results are provided and analyzed to illustrate the performance and effectiveness of our approach. Xiwen Ma, Zhihuan Hu, Kairong Duan, Xiaolin Ai, Wei Xie 0009, Jingsong Yang, Weidong Zhang 0004 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Neural Formation A*: A Knowledge-Data Hybrid-Driven Path Planning Algorithm for Multi-agent Formation Cooperation
Qi'ang Cai, Xiaolin Ai, Zhiqiang Pu, Jianqiang Yi, Feng Lv |
ICANN (4) | 2 |
| 2024 | Heterogeneous Observation Aggregation Network for Multi-agent Reinforcement LearningabstractLearning effective policies is challenging for a multi-agent system in partially observable environments, where agents need to extract relevant features from local observations. Most approaches in multi-agent reinforcement learning (MARL) are limited to feature extraction for homogenous agents. They struggle to deal with local observations in heterogeneous multi-agent scenarios, where agents have different observation spaces and are necessitated to process semantically varied information. To address this issue, we analyze the observational heterogeneity of multi-agent systems, and propose a heterogeneous-graph-based approach for feature extraction in MARL. We model agent observations as heterogeneous graphs, and design a heterogeneous observation aggregation network (HOA-Net) for processing these graph-based observations. HOA-Net is specifically designed to address various forms of observational heterogeneity. It employs class-specific weighting networks and computes across-class attentions for observed entities, effectively reducing the number of learnable parameters. The proposed method is evaluated on SMAC and an Unreal-Engine-based heterogeneous multi-agent testbed. Experimental results demonstrate that our method significantly outperforms other baselines in effectively aggregating an agent’s observation, and finally enhancing the performance of heterogeneous multi-agent systems. Xiaolin Ai, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi |
IJCNN | 2 |
| 2024 | Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningabstractReinforcement learning (RL) has emerged as a pivotal technique for fine-tuning large language models (LLMs) on specific tasks. However, prevailing RL fine-tuning methods predominantly rely on PPO and its variants. Though these algorithms are effective in general RL settings, they often exhibit suboptimal performance and vulnerability to distribution collapse when applied to the fine-tuning of LLMs. In this paper, we propose CORY, extending the RL fine-tuning of LLMs to a sequential cooperative multi-agent reinforcement learning framework, to leverage the inherent coevolution and emergent capabilities of multi-agent systems. In CORY, the LLM to be fine-tuned is initially duplicated into two autonomous agents: a pioneer and an observer. The pioneer generates responses based on queries, while the observer generates responses using both the queries and the pioneer’s responses. The two agents are trained together. During training, the agents exchange roles periodically, fostering cooperation and coevolution between them. Experiments evaluate CORY's performance by fine-tuning GPT-2 and Llama-2 under subjective and objective reward functions on the IMDB Review and GSM8K datasets, respectively. Results show that CORY outperforms PPO in terms of policy optimality, resistance to distribution collapse, and training robustness, thereby underscoring its potential as a superior methodology for refining LLMs in real-world applications. Zhiqiang Pu, Boyin Liu, Xiaolin Ai, Yanyan Liang 0001, Min Chen 0038 |
NeurIPS | 5 |
| 2024 | An Affine-Based Maneuver Control Method for Multi-Agent Cooperative Transportation System Over Switching FormationsabstractThis paper addresses the maneuver control prob-lem for the multi-agent cooperative transportation systems (MACTSs) with double-integrator dynamics over switching formations. The switching formations consist of a pair of operations: one is the switching directed graphs and the other is the corresponding formation configurations. In real cooperative transportation scenarios, varying interaction re-lationships necessitate distinct system configurations. But it is challenging for researchers to design control law with the evolutions of not only the communication topology but also the system configuration. Drawing inspiration from advancements on switching topologies, we utilize the characteristics of the stress matrix to build up a new LMI inequality to design a novel class of distributed affine-based controller. The global convergence will be achieved as long as the feedback gain matrix and the switching signal satisfy three specific conditions. And we give the corresponding algorithm to calculate the required control parameters. A Lyapunov function is constructed to demonstrate the system's stability and simulation example is provided in detail at the end of this paper to validate our method's efficacy. Xiaolin Ai, Zhiqiang Pu, Feng Lv |
SMC | 2 |
| 2023 | Learning Superior Cooperative Policy in Adversarial Multi-Team Reinforcement LearningabstractMulti-agent Reinforcement Learning (MARL) has become a powerful tool for addressing multi-agent challenges. Existing studies have explored numerous models to use MARL to solve single-team cooperation (competition) problems and adversarial problems with opponents controlled by static knowledge-based policies. However, most studies in the literature often ignore adversarial multi-team problems involving dynamically evolving opponents. We investigate adversarial multi-team problems where all participating teams use MARL learners to learn policies against each other. Two objectives are achieved in this study. Firstly, we design an adversarial team-versus-team learning framework to generate cooperative multi-agent policies to compete against opponents without preprogrammed opponent partners or any supervision. Secondly, we explore the key factors to achieve win-rate superiority during dynamic competitions. Then we put forward a novel FeedBack MARL (FBMARL) algorithm that takes advantage of feedback loops to adjust optimizer hyper-parameters based on real-time game statistics. Finally, the effectiveness of our FBMARL model is tested in a benchmark environment named Multi-Team Decentralized Collective Assault (MT-DCA). The results demonstrate that our feedback MARL model can achieve superior performance over baseline competitor MARL learners in 2-team and 3-team dynamic competitions. Qingxu Fu, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi, Xiaolin Ai, Wanmai Yuan |
IJCNN | 5 |
| 2023 | Learning Cooperative Policies with Graph Networks in Distributed Swarm SystemsabstractDeriving efficient cooperative policies in uncertain dynamic environments poses huge challenges for a distributed swarm system due to the limited capability of the agents and the complex dynamics of the environment. In this paper, a novel distributed method based on deep reinforcement learning using observation-level and communication-level graph networks is proposed to learn cooperative policies for the distributed swarm system. Specifically, a relational directed graph attention neural network is designed to model observation-level graphs composed of heterogeneous relational graphs among each agent and each type of entities (e.g., obstacles, other teammates, opponents), for extracting different relational representations. Moreover, a relevant directed graph attention network is presented to cut off the ineffective communication among irrelevant agents, and model a relevant communication topology between each agent and relevant homogeneous neighbor agents as an communication-level graph, for promoting efficient inter-agent interactions. Furthermore, a distributed actor-critic algorithm with full parameter sharing is implemented to learn cooperative swarm policies by using distributed critics, which avoids the curse of dimensionality under a centralized critic. Various simulation results validate the effectiveness and generalization of the proposed method, and demonstrate that the proposed method outperforms existing state-of-the-art methods on coverage and pursuit tasks. Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Xiaolin Ai, Wanmai Yuan |
IJCNN | 5 |
| 2023 | A Deep Reinforcement Learning Approach Combined With Model-Based Paradigms for Multiagent Formation Control With Collision AvoidanceabstractGenerating collision-free formation control strategy for multiagent systems faces huge challenges in collaborative navigation tasks, especially in a highly dynamic and uncertain environment. Two typical methodologies for solving this problem are the conventional model-based paradigm and the data-driven paradigm, particularly the widely used deep reinforcement learning (DRL) method. However, both the model-based and data-driven paradigms encounter inherent drawbacks. In this paper, we present two novel general schemes that combine these two paradigms together in an online mode. Specifically, the two paradigms are combined in a parallel and a serial structure in these two schemes, respectively. In the parallel scheme, the outputs of the model-based and DRL-based controllers are lumped together. In the serial scheme, the output of the model-based controller is fed as an input of the DRL-based controller. The interpretation of the two combined schemes is suggested from a control-oriented perspective, where the parallel DRL controller is viewed as a complementary uncertainty compensator and the serial DRL controller is taken as an inverse dynamics estimator. Finally, comprehensive simulations are conducted to demonstrate the superiority of the proposed schemes, and the effectiveness is further verified by deploying our schemes to a physical experiment platform based on a set of three-wheeled omnidirectional robots. Zhiqiang Pu, Xiaolin Ai, Tenghai Qiu, Jianqiang Yi |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |