VLDB 2026 Research / reviewers in the wild / expert
Chenyu Zhang 0002
dblp:136/1220-2
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0005-3612-4894ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 64% Efficient and distributed learning · 12% Learning theory · 12% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 54% Algorithmic game theory and mechanism design · 46% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › multi-agent reinforcement learning
mean field games |
1.6 | 2 | 2025 | Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function Approximation · ICLR 2025 Graphon Mean Field Games with a Representative Player: Analysis and Learning Algorithm · ICML 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
equilibrium learning |
0.9 | 1 | 2025 | Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function Approximation · ICLR 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function Approximation · ICLR 2025 |
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent |
0.9 | 1 | 2025 | Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function Approximation · ICLR 2025 |
Machine learning › Efficient and distributed learning › federated learning › federated sequential learning
federated reinforcement learning |
0.8 | 1 | 2024 | Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement Learning · ICLR 2024 |
Machine learning › Learning theory
finite-time analysis |
0.8 | 1 | 2024 | Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement Learning · ICLR 2024 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
graphon mean field games |
0.8 | 1 | 2024 | Graphon Mean Field Games with a Representative Player: Analysis and Learning Algorithm · ICML 2024 |
Machine learning › Reinforcement learning
on-policy reinforcement learning |
0.8 | 1 | 2024 | Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement Learning · ICLR 2024 |
Algorithmic game theory and mechanism design
equilibrium computation |
0.8 | 1 | 2024 | Graphon Mean Field Games with a Representative Player: Analysis and Learning Algorithm · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
linear function approximation · 2.5stochastic gradient descent · 1.7population-aware function approximation · 1.7sample complexity analysis · 1.5online learning algorithms · 0.8online learning algorithm · 0.8markovian sampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function ApproximationabstractMean field games (MFGs) model interactions in large-population multi-agent systems through population distributions. Traditional learning methods for MFGs are based on fixed-point iteration (FPI), where policy updates and induced population distributions are computed separately and sequentially. However, FPI-type methods may suffer from inefficiency and instability due to potential oscillations caused by this forward-backward procedure. In this work, we propose a novel perspective that treats the policy and population as a unified parameter controlling the game dynamics. By applying stochastic parameter approximation to this unified parameter, we develop SemiSGD, a simple stochastic gradient descent (SGD)-type method, where an agent updates its policy and population estimates simultaneously and fully asynchronously. Building on this perspective, we further apply linear function approximation (LFA) to the unified parameter, resulting in the first population-aware LFA (PA-LFA) for learning MFGs on continuous state-action spaces. A comprehensive finite-time convergence analysis is provided for SemiSGD with PA-LFA, including its convergence to the equilibrium for linear MFGs—a class of MFGs with a linear structure concerning the population—under the standard contractivity condition, and to a neighborhood of the equilibrium under a more practical condition. We also characterize the approximation error for non-linear MFGs. We validate our theoretical findings with six experiments on three MFGs. Chenyu Zhang 0002, Xu Chen 0033, Xuan Di |
ICLR | 1 |
| 2024 | A Single Online Agent Can Efficiently Learn Mean Field GamesabstractMean field games (MFGs) are a promising framework for modeling the behavior of large-population systems. However, solving MFGs can be challenging due to the coupling of forward population evolution and backward agent dynamics. Typically, obtaining mean field Nash equilibria (MFNE) involves an iterative approach where the forward and backward processes are solved alternately, known as fixed-point iteration (FPI). This method requires fully observed population propagation and agent dynamics over the entire spatial domain, which could be impractical in some real-world scenarios. To overcome this limitation, this paper introduces a novel online single-agent model-free learning scheme, which enables a single agent to learn MFNE using online samples, without prior knowledge of the state-action space, reward function, or transition dynamics. Specifically, the agent updates its policy through the value function (Q), while simultaneously evaluating the mean field state (M), using the same batch of observations. We develop two variants of this learning scheme: off-policy and on-policy QM iteration. We prove that they efficiently approximate FPI, and a sample complexity guarantee is provided. The efficacy of our methods is confirmed by numerical experiments. Chenyu Zhang 0002, Xu Chen 0033, Xuan Di |
ECAI | 1 |
| 2024 | Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement LearningabstractFederated reinforcement learning (FRL) has emerged as a promising paradigm for reducing the sample complexity of reinforcement learning tasks by exploiting information from different agents. However, when each agent interacts with a potentially different environment, little to nothing is known theoretically about the non-asymptotic performance of FRL algorithms. The lack of such results can be attributed to various technical challenges and their intricate interplay: Markovian sampling, linear function approximation, multiple local updates to save communication, heterogeneity in the reward functions and transition kernels of the agents' MDPs, and continuous state-action spaces. Moreover, in the on-policy setting, the behavior policies vary with time, further complicating the analysis. In response, we introduce FedSARSA, a novel federated on-policy reinforcement learning scheme, equipped with linear function approximation, to address these challenges and provide a comprehensive finite-time error analysis. Notably, we establish that FedSARSA converges to a policy that is near-optimal for all agents, with the extent of near-optimality proportional to the level of heterogeneity. Furthermore, we prove that FedSARSA leverages agent collaboration to enable linear speedups as the number of agents increases, which holds for both fixed and adaptive step-size configurations. Chenyu Zhang 0002, Han Wang 0016, Aritra Mitra, James Anderson 0001 |
ICLR | 1 |
| 2024 | Graphon Mean Field Games with a Representative Player: Analysis and Learning AlgorithmabstractWe propose a discrete time graphon game formulation on continuous state and action spaces using a representative player to study stochastic games with heterogeneous interaction among agents. This formulation admits both conceptual and mathematical advantages, compared to a widely adopted formulation using a continuum of players. We prove the existence and uniqueness of the graphon equilibrium with mild assumptions, and show that this equilibrium can be used to construct an approximate solution for the finite player game, which is challenging to analyze and solve due to curse of dimensionality. An online oracle-free learning algorithm is developed to solve the equilibrium numerically, and sample complexity analysis is provided for its convergence. Fuzhong Zhou, Chenyu Zhang 0002, Xu Chen 0033, Xuan Di |
ICML | 2 |