VLDB 2026 Research / reviewers in the wild / expert
Xueguang Lyu
dblp:232/1694
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 54% Language models and text generation · 23% Multi-agent systems · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.6 | 2 | 2026 | LLM Collaboration with Multi-Agent Reinforcement Learning · AAAI 2026 A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement Learning · AAAI 2022 |
Natural language and speech › Language models and text generation › LLM agents
LLM collaboration |
1.0 | 1 | 2026 | LLM Collaboration with Multi-Agent Reinforcement Learning · AAAI 2026 |
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent LLM coordination |
1.0 | 1 | 2026 | LLM Collaboration with Multi-Agent Reinforcement Learning · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution |
0.6 | 1 | 2022 | A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement Learning · AAAI 2022 |
Machine learning › Reinforcement learning
actor-critic methods |
0.2 | 1 | 2022 | A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement Learning · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
multi-turn reinforcement learning · 1.0group relative policy optimization · 1.0policy gradient analysis · 0.6bias-variance analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM Collaboration with Multi-Agent Reinforcement LearningabstractA large amount of work has been done in Multi-Agent Systems (MAS) for modeling and solving problems with multiple interacting agents. However, most LLMs are pretrained independently and not specifically optimized for coordination. Existing LLM fine-tuning frameworks rely on individual rewards, which require complex reward designs for each agent to encourage collaboration. To address these challenges, we model LLM collaboration as a cooperative Multi-Agent Reinforcement Learning (MARL) problem. We develop a multi-agent, multi-turn algorithm, Multi-Agent Group Relative Policy Optimization (MAGRPO), to solve it, building on current RL approaches for LLMs as well as MARL techniques. Our experiments on LLM writing and coding collaboration demonstrate that fine-tuning MAS with MAGRPO enables agents to generate high-quality responses efficiently through effective cooperation. Our approach opens the door to using MARL methods for LLM collaboration and highlights the associated challenges. Xueguang Lyu, Christopher Amato |
AAAI | 3 |
| 2023 | On Centralized Critics in Multi-Agent Reinforcement LearningabstractCentralized Training for Decentralized Execution, where agents are trained offline in a centralized fashion and execute online in a decentralized manner, has become a popular approach in Multi-Agent Reinforcement Learning (MARL). In particular, it has become popular to develop actor-critic methods that train decentralized actors with a centralized critic where the centralized critic is allowed access global information of the entire system, including the true system state. Such centralized critics are possible given offline information and are not used for online execution. While these methods perform well in a number of domains and have become a de facto standard in MARL, using a centralized critic in this context has yet to be sufficiently analyzed theoretically or empirically. In this paper, we therefore formally analyze centralized and decentralized critic approaches, and analyze the effect of using state-based critics in partially observable environments. We derive theories contrary to the common intuition: critic centralization is not strictly beneficial, and using state values can be harmful. We further prove that, in particular, state-based critics can introduce unexpected bias and variance compared to history-based critics. Finally, we demonstrate how the theory applies in practice by comparing different forms of critics on a wide range of common multi-agent benchmarks. The experiments show practical issues such as the difficulty of representation learning with partial observability, which highlights why the theoretical problems are often overlooked in the literature. Xueguang Lyu, Andrea Baisero, Brett Daley, Christopher Amato |
J. Artif. Intell. Res. | 1 |
| 2022 | A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement LearningabstractCentralized Training for Decentralized Execution, where training is done in a centralized offline fashion, has become a popular solution paradigm in Multi-Agent Reinforcement Learning. Many such methods take the form of actor-critic with state-based critics, since centralized training allows access to the true system state, which can be useful during training despite not being available at execution time. State-based critics have become a common empirical choice, albeit one which has had limited theoretical justification or analysis. In this paper, we show that state-based critics can introduce bias in the policy gradient estimates, potentially undermining the asymptotic guarantees of the algorithm. We also show that, even if the state-based critics do not introduce any bias, they can still result in a larger gradient variance, contrary to the common intuition. Finally, we show the effects of the theories in practice by comparing different forms of centralized critics on a wide range of common benchmarks, and detail how various environmental properties are related to the effectiveness of different types of critics. Xueguang Lyu, Andrea Baisero, Christopher Amato |
AAAI | 1 |