Wenshuai Zhao

dblp:246/5109 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 83% Learning paradigms · 7% Representation and self-supervised learning · 5%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
3.342025
Learning Progress Driven Multi-Agent Curriculum · ICML 2025
AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025
Optimistic Multi-Agent Policy Gradient · ICML 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment
1.622025
Learning Progress Driven Multi-Agent Curriculum · ICML 2025
Backpropagation Through Agents · AAAI 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
policy factorization
1.622025
AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025
Backpropagation Through Agents · AAAI 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.912025
AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › equilibrium learning
correlated equilibrium
0.912025
AgentMixer: Multi-Agent Correlated Policy Factorization · AAAI 2025
Machine learning › Learning paradigms
curriculum learning
0.912025
Learning Progress Driven Multi-Agent Curriculum · ICML 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent policy gradient
0.812024
Optimistic Multi-Agent Policy Gradient · ICML 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning › cooperative multi-agent reinforcement learning
relative overgeneralization
0.812024
Optimistic Multi-Agent Policy Gradient · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model
0.712023
Simplified Temporal Consistency Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
model-based reinforcement learning
0.712023
Simplified Temporal Consistency Reinforcement Learning · ICML 2023
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency
0.712023
Simplified Temporal Consistency Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › sparse reward reinforcement learning
sparse reward tasks
0.312025
Learning Progress Driven Multi-Agent Curriculum · ICML 2025
Machine learning › Reinforcement learning
model-free reinforcement learning
0.212023
Simplified Temporal Consistency Reinforcement Learning · ICML 2023

Methods — techniques the papers use, named apart from their topics

individual-global-consistency · 0.9curriculum learning · 0.9correlated equilibrium · 0.9TD-error learning progress · 0.9proximal policy optimization · 0.8policy gradient · 0.8optimistic update · 0.8advantage clipping · 0.8planning · 0.7latent temporal consistency · 0.7
YearPublicationVenuePosition
2025 AgentMixer: Multi-Agent Correlated Policy Factorization
abstract
In multi-agent reinforcement learning, centralized training with decentralized execution (CTDE) methods typically assumes that agents make decisions based on their local observations independently, which may not lead to a correlated joint policy with coordination. Coordination can be explicitly encouraged during training and individual policies can be trained to imitate the correlated joint policy. However, this may lead to an asymmetric learning failure due to the observation mismatch between the joint and individual policies. Inspired by the concept of correlated equilibrium, we introduce a strategy modification called AgentMixer that allows agents to correlate their policies. AgentMixer combines individual partially observable policies into a joint fully observable policy non-linearly. To enable decentralized execution, we introduce Individual-Global-Consistency to guarantee mode consistency during joint training of the centralized and decentralized policies and prove that AgentMixer converges to an ϵ-approximate Correlated Equilibrium. In the Multi-Agent MuJoCo, SMAC-v2, Matrix Game, and Predator-Prey benchmarks, AgentMixer outperforms or matches state-of-the-art methods.
Wenshuai Zhao, Joni Pajarinen
AAAI2
2025 Learning Progress Driven Multi-Agent Curriculum
abstract
The number of agents can be an effective curriculum variable for controlling the difficulty of multi-agent reinforcement learning (MARL) tasks. Existing work typically uses manually defined curricula such as linear schemes. We identify two potential flaws while applying existing reward-based automatic curriculum learning methods in MARL: (1) The expected episode return used to measure task difficulty has high variance; (2) Credit assignment difficulty can be exacerbated in tasks where increasing the number of agents yields higher returns which is common in many MARL tasks. To address these issues, we propose to control the curriculum by using a TD-error based learning progress measure and by letting the curriculum proceed from an initial context distribution to the final task specific one. Since our approach maintains a distribution over the number of agents and measures learning progress rather than absolute performance, which often increases with the number of agents, we alleviate problem (2). Moreover, the learning progress measure naturally alleviates problem (1) by aggregating returns. In three challenging sparse-reward MARL benchmarks, our approach outperforms state-of-the-art baselines.
Wenshuai Zhao, Joni Pajarinen
ICML1
2024 Backpropagation Through Agents
abstract
A fundamental challenge in multi-agent reinforcement learning (MARL) is to learn the joint policy in an extremely large search space, which grows exponentially with the number of agents. Moreover, fully decentralized policy factorization significantly restricts the search space, which may lead to sub-optimal policies. In contrast, the auto-regressive joint policy can represent a much richer class of joint policies by factorizing the joint policy into the product of a series of conditional individual policies. While such factorization introduces the action dependency among agents explicitly in sequential execution, it does not take full advantage of the dependency during learning. In particular, the subsequent agents do not give the preceding agents feedback about their decisions. In this paper, we propose a new framework Back-Propagation Through Agents (BPTA) that directly accounts for both agents' own policy updates and the learning of their dependent counterparts. This is achieved by propagating the feedback through action chains. With the proposed framework, our Bidirectional Proximal Policy Optimisation (BPPO) outperforms the state-of-the-art methods. Extensive experiments on matrix games, StarCraftII v2, Multi-agent MuJoCo, and Google Research Football demonstrate the effectiveness of the proposed method.
Wenshuai Zhao, Joni Pajarinen
AAAI2
2024 Optimistic Multi-Agent Policy Gradient
abstract
Relative overgeneralization (RO) occurs in cooperative multi-agent learning tasks when agents converge towards a suboptimal joint policy due to overfitting to suboptimal behaviors of other agents. No methods have been proposed for addressing RO in multi-agent policy gradient (MAPG) methods although these methods produce state-of-the-art results. To address this gap, we propose a general, yet simple, framework to enable optimistic updates in MAPG methods that alleviate the RO problem. Our approach involves clipping the advantage to eliminate negative values, thereby facilitating optimistic updates in MAPG. The optimism prevents individual agents from quickly converging to a local optimum. Additionally, we provide a formal analysis to show that the proposed method retains optimality at a fixed point. In extensive evaluations on a diverse set of tasks including the Multi-agent MuJoCo and Overcooked benchmarks, our method outperforms strong baselines on 13 out of 19 tested tasks and matches the performance on the rest.
Wenshuai Zhao, Yi Zhao 0014, Juho Kannala, Joni Pajarinen
ICML1
2023 Simplified Temporal Consistency Reinforcement Learning
abstract
Reinforcement learning (RL) is able to solve complex sequential decision-making tasks but is currently limited by sample efficiency and required computation. To improve sample efficiency, recent work focuses on model-based RL which interleaves model learning with planning. Recent methods further utilize policy learning, value estimation, and, self-supervised learning as auxiliary objectives. In this paper we show that, surprisingly, a simple representation learning approach relying only on a latent dynamics model trained by latent temporal consistency is sufficient for high-performance RL. This applies when using pure planning with a dynamics model conditioned on the representation, but, also when utilizing the representation as policy and value function features in model-free RL. In experiments, our approach learns an accurate dynamics model to solve challenging high-dimensional locomotion tasks with online planners while being 4.1$\times$ faster to train compared to ensemble-based methods. With model-free RL without planning, especially on high-dimensional tasks, such as the Deepmind Control Suite Humanoid and Dog tasks, our approach outperforms model-free methods by a large margin and matches model-based methods’ sample efficiency while training 2.4$\times$ faster.
Yi Zhao 0014, Wenshuai Zhao, Rinu Boney, Juho Kannala, Joni Pajarinen
ICML2