VLDB 2026 Research / reviewers in the wild / expert
Wenshuai Zhao
dblp:246/5109
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 83% Learning paradigms · 7% Representation and self-supervised learning · 5% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
3.3 | 4 | 2025 | Learning Progress Driven Multi-Agent Curriculum · ICML 2025 AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025 Optimistic Multi-Agent Policy Gradient · ICML 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
1.6 | 2 | 2025 | Learning Progress Driven Multi-Agent Curriculum · ICML 2025 Backpropagation Through Agents · AAAI 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
policy factorization |
1.6 | 2 | 2025 | AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025 Backpropagation Through Agents · AAAI 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution |
0.9 | 1 | 2025 | AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › equilibrium learning
correlated equilibrium |
0.9 | 1 | 2025 | AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025 |
Machine learning › Learning paradigms
curriculum learning |
0.9 | 1 | 2025 | Learning Progress Driven Multi-Agent Curriculum · ICML 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent policy gradient |
0.8 | 1 | 2024 | Optimistic Multi-Agent Policy Gradient · ICML 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › cooperative multi-agent reinforcement learning
relative overgeneralization |
0.8 | 1 | 2024 | Optimistic Multi-Agent Policy Gradient · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model |
0.7 | 1 | 2023 | Simplified Temporal Consistency Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | Simplified Temporal Consistency Reinforcement Learning · ICML 2023 |
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency |
0.7 | 1 | 2023 | Simplified Temporal Consistency Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › sparse reward reinforcement learning
sparse reward tasks |
0.3 | 1 | 2025 | Learning Progress Driven Multi-Agent Curriculum · ICML 2025 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.2 | 1 | 2023 | Simplified Temporal Consistency Reinforcement Learning · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
individual-global-consistency · 0.9curriculum learning · 0.9correlated equilibrium · 0.9TD-error learning progress · 0.9proximal policy optimization · 0.8policy gradient · 0.8optimistic update · 0.8advantage clipping · 0.8planning · 0.7latent temporal consistency · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AgentMixer: Multi-Agent Correlated Policy FactorizationabstractIn multi-agent reinforcement learning, centralized training with decentralized execution (CTDE) methods typically assumes that agents make decisions based on their local observations independently, which may not lead to a correlated joint policy with coordination. Coordination can be explicitly encouraged during training and individual policies can be trained to imitate the correlated joint policy. However, this may lead to an asymmetric learning failure due to the observation mismatch between the joint and individual policies. Inspired by the concept of correlated equilibrium, we introduce a strategy modification called AgentMixer that allows agents to correlate their policies. AgentMixer combines individual partially observable policies into a joint fully observable policy non-linearly. To enable decentralized execution, we introduce Individual-Global-Consistency to guarantee mode consistency during joint training of the centralized and decentralized policies and prove that AgentMixer converges to an ϵ-approximate Correlated Equilibrium. In the Multi-Agent MuJoCo, SMAC-v2, Matrix Game, and Predator-Prey benchmarks, AgentMixer outperforms or matches state-of-the-art methods. Wenshuai Zhao, Joni Pajarinen |
AAAI | 2 |
| 2025 | Learning Progress Driven Multi-Agent CurriculumabstractThe number of agents can be an effective curriculum variable for controlling the difficulty of multi-agent reinforcement learning (MARL) tasks. Existing work typically uses manually defined curricula such as linear schemes. We identify two potential flaws while applying existing reward-based automatic curriculum learning methods in MARL: (1) The expected episode return used to measure task difficulty has high variance; (2) Credit assignment difficulty can be exacerbated in tasks where increasing the number of agents yields higher returns which is common in many MARL tasks. To address these issues, we propose to control the curriculum by using a TD-error based learning progress measure and by letting the curriculum proceed from an initial context distribution to the final task specific one. Since our approach maintains a distribution over the number of agents and measures learning progress rather than absolute performance, which often increases with the number of agents, we alleviate problem (2). Moreover, the learning progress measure naturally alleviates problem (1) by aggregating returns. In three challenging sparse-reward MARL benchmarks, our approach outperforms state-of-the-art baselines. Wenshuai Zhao, Joni Pajarinen |
ICML | 1 |
| 2024 | Backpropagation Through AgentsabstractA fundamental challenge in multi-agent reinforcement learning (MARL) is to learn the joint policy in an extremely large search space, which grows exponentially with the number of agents. Moreover, fully decentralized policy factorization significantly restricts the search space, which may lead to sub-optimal policies. In contrast, the auto-regressive joint policy can represent a much richer class of joint policies by factorizing the joint policy into the product of a series of conditional individual policies. While such factorization introduces the action dependency among agents explicitly in sequential execution, it does not take full advantage of the dependency during learning. In particular, the subsequent agents do not give the preceding agents feedback about their decisions. In this paper, we propose a new framework Back-Propagation Through Agents (BPTA) that directly accounts for both agents' own policy updates and the learning of their dependent counterparts. This is achieved by propagating the feedback through action chains. With the proposed framework, our Bidirectional Proximal Policy Optimisation (BPPO) outperforms the state-of-the-art methods. Extensive experiments on matrix games, StarCraftII v2, Multi-agent MuJoCo, and Google Research Football demonstrate the effectiveness of the proposed method. Wenshuai Zhao, Joni Pajarinen |
AAAI | 2 |
| 2024 | Optimistic Multi-Agent Policy GradientabstractRelative overgeneralization (RO) occurs in cooperative multi-agent learning tasks when agents converge towards a suboptimal joint policy due to overfitting to suboptimal behaviors of other agents. No methods have been proposed for addressing RO in multi-agent policy gradient (MAPG) methods although these methods produce state-of-the-art results. To address this gap, we propose a general, yet simple, framework to enable optimistic updates in MAPG methods that alleviate the RO problem. Our approach involves clipping the advantage to eliminate negative values, thereby facilitating optimistic updates in MAPG. The optimism prevents individual agents from quickly converging to a local optimum. Additionally, we provide a formal analysis to show that the proposed method retains optimality at a fixed point. In extensive evaluations on a diverse set of tasks including the Multi-agent MuJoCo and Overcooked benchmarks, our method outperforms strong baselines on 13 out of 19 tested tasks and matches the performance on the rest. Wenshuai Zhao, Yi Zhao 0014, Juho Kannala, Joni Pajarinen |
ICML | 1 |
| 2023 | Simplified Temporal Consistency Reinforcement LearningabstractReinforcement learning (RL) is able to solve complex sequential decision-making tasks but is currently limited by sample efficiency and required computation. To improve sample efficiency, recent work focuses on model-based RL which interleaves model learning with planning. Recent methods further utilize policy learning, value estimation, and, self-supervised learning as auxiliary objectives. In this paper we show that, surprisingly, a simple representation learning approach relying only on a latent dynamics model trained by latent temporal consistency is sufficient for high-performance RL. This applies when using pure planning with a dynamics model conditioned on the representation, but, also when utilizing the representation as policy and value function features in model-free RL. In experiments, our approach learns an accurate dynamics model to solve challenging high-dimensional locomotion tasks with online planners while being 4.1$\times$ faster to train compared to ensemble-based methods. With model-free RL without planning, especially on high-dimensional tasks, such as the Deepmind Control Suite Humanoid and Dog tasks, our approach outperforms model-free methods by a large margin and matches model-based methods’ sample efficiency while training 2.4$\times$ faster. Yi Zhao 0014, Wenshuai Zhao, Rinu Boney, Juho Kannala, Joni Pajarinen |
ICML | 2 |