VLDB 2026 Research / reviewers in the wild / expert
Yuan Zhang 0027
dblp:48/2168-27
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 69% Multi-agent systems · 11% Motion planning and robot control · 9% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
2.8 | 5 | 2024 | Open Ad Hoc Teamwork with Cooperative Game Theory · ICML 2024 Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach · ICAPS 2024 SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
1.0 | 2 | 2022 | SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning · NeurIPS 2022 Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games · AAAI 2020 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
ad hoc teamwork |
0.8 | 1 | 2024 | Open Ad Hoc Teamwork with Cooperative Game Theory · ICML 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution |
0.8 | 1 | 2024 | Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach · ICAPS 2024 |
Robotics › Motion planning and robot control
robot learning |
0.8 | 1 | 2024 | Learning Continuous Control with Geometric Regularity from Robot Intrinsic Symmetry · ICRA 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
0.6 | 1 | 2022 | SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.5 | 1 | 2021 | Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework |
0.5 | 1 | 2021 | Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.5 | 1 | 2021 | Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021 |
Natural language and speech › Language models and text generation
text generation |
0.5 | 1 | 2021 | Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021 |
Algorithmic game theory and mechanism design
cooperative game theory |
0.4 | 1 | 2020 | Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games · AAAI 2020 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making |
0.2 | 1 | 2024 | Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach · ICAPS 2024 |
Methods — techniques the papers use, named apart from their topics
parameter sharing · 0.8offline reinforcement learning · 0.8local critic · 0.8graph neural network · 0.8geometric regularity · 0.8cooperative game theory · 0.8centralized training decentralized execution · 0.8PPO · 0.8shapley value · 0.6q-learning · 0.6shapley q-value · 0.4deep deterministic policy gradient · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | UDUC: An Uncertainty-Driven Approach for Learning-Based Robust ControlabstractLearning-based techniques have become popular in both model predictive control (MPC) and reinforcement learning (RL). Probabilistic ensemble (PE) models offer a promising approach for modelling system dynamics, showcasing the ability to capture uncertainty and scalability in high-dimensional control scenarios. However, PE models are susceptible to mode collapse, resulting in non-robust control when faced with environments slightly different from the training set. In this paper, we introduce the uncertainty-driven robust control (UDUC) loss as an alternative objective for training PE models, drawing inspiration from contrastive learning. We analyze the robustness of the UDUC loss through the lens of robust optimization and evaluate its performance on the challenging real-world reinforcement learning (RWRL) benchmark, which involves significant environmental mismatches between the training and testing environments. Yuan Zhang 0027, Jasper Hoffmann, Joschka Boedecker |
ECAI | 1 |
| 2024 | Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic ApproachabstractThe complex railway network is a challenging real-world multi-agent system usually involving thousands of agents. Current planning methods heavily depend on expert knowledge to formulate solutions for specific cases and are therefore hardly generalized to new scenarios, on which multi-agent reinforcement learning (MARL) draws significant attention. Despite some successful applications in multi-agent decision-making tasks, MARL is hard to scale to a large number of agents. This paper rethinks the curse of agents in the centralized-training-decentralized-execution (CTDE) paradigm and proposes a local-critic approach to address the issue. By combining the local critic with the PPO algorithm, we design a deep MARL algorithm denoted as local-critic PPO (LCPPO). In experiments, we evaluate the effectiveness of LCPPO on a complex railway network benchmark, Flatland, with various numbers of agents. Noticeably, LCPPO shows prominent generalizability and robustness under the changes of environments. Yuan Zhang 0027, Umashankar Deekshith, Joschka Boedecker |
ICAPS | 1 |
| 2024 | Open Ad Hoc Teamwork with Cooperative Game TheoryabstractAd hoc teamwork poses a challenging problem, requiring the design of an agent to collaborate with teammates without prior coordination or joint training. Open ad hoc teamwork (OAHT) further complicates this challenge by considering environments with a changing number of teammates, referred to as open teams. One promising solution in practice to this problem is leveraging the generalizability of graph neural networks to handle an unrestricted number of agents with various agent-types, named graph-based policy learning (GPL). However, its joint Q-value representation over a coordination graph lacks convincing explanations. In this paper, we establish a new theory to understand the representation of the joint Q-value for OAHT and its learning paradigm, through the lens of cooperative game theory. Building on our theory, we propose a novel algorithm named CIAO, based on GPL's framework, with additional provable implementation tricks that can facilitate learning. The demos of experimental results are available on https://sites.google.com/view/ciao2024, and the code of experiments is published on https://github.com/hsvgbkhgbv/CIAO. Yang Li 0116, Yuan Zhang 0027, Wei Pan 0004, Samuel Kaski |
ICML | 3 |
| 2024 | Learning Continuous Control with Geometric Regularity from Robot Intrinsic SymmetryabstractGeometric regularity, which leverages data symmetry, has been successfully incorporated into deep learning architectures such as CNNs, RNNs, GNNs, and Transformers. While this concept has been widely applied in robotics to address the curse of dimensionality when learning from high-dimensional data, the inherent reflectional and rotational symmetry of robot structures has not been adequately explored. Drawing inspiration from cooperative multi-agent reinforcement learning, we introduce novel network structures for single-agent control learning that explicitly capture these symmetries. Moreover, we investigate the relationship between the geometric prior and the concept of Parameter Sharing in multi-agent reinforcement learning. Last but not the least, we implement the proposed framework in online and offline learning methods to demonstrate its ease of use. Through experiments conducted on various challenging continuous control tasks on simulators and real robots, we highlight the significant potential of the proposed geometric regularity in enhancing robot learning capabilities. Shengchao Yan, Baohe Zhang, Yuan Zhang 0027, Joschka Boedecker, Wolfram Burgard |
ICRA | 3 |
| 2022 | SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-LearningabstractValue factorisation is a useful technique for multi-agent reinforcement learning (MARL) in global reward game, however, its underlying mechanism is not yet fully understood. This paper studies a theoretical framework for value factorisation with interpretability via Shapley value theory. We generalise Shapley value to Markov convex game called Markov Shapley value (MSV) and apply it as a value factorisation method in global reward game, which is obtained by the equivalence between the two games. Based on the properties of MSV, we derive Shapley-Bellman optimality equation (SBOE) to evaluate the optimal MSV, which corresponds to an optimal joint deterministic policy. Furthermore, we propose Shapley-Bellman operator (SBO) that is proved to solve SBOE. With a stochastic approximation and some transformations, a new MARL algorithm called Shapley Q-learning (SHAQ) is established, the implementation of which is guided by the theoretical results of SBO and MSV. We also discuss the relationship between SHAQ and relevant value factorisation methods. In the experiments, SHAQ exhibits not only superior performances on all tasks but also the interpretability that agrees with the theoretical analysis. The implementation of this paper is placed on https://github.com/hsvgbkhgbv/shapley-q-learning. Yuan Zhang 0027, Yunjie Gu, Tae-Kyun Kim 0001 |
NeurIPS | 2 |
| 2021 | Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System
Yuan Zhang 0027, Tae-Kyun Kim 0001, Yunjie Gu |
ICLR | 2 |
| 2020 | Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesabstractCooperative game is a critical research area in the multi-agent reinforcement learning (MARL). Global reward game is a subclass of cooperative games, where all agents aim to maximize the global reward. Credit assignment is an important problem studied in the global reward game. Most of previous works stood by the view of non-cooperative-game theoretical framework with the shared reward approach, i.e., each agent being assigned a shared global reward directly. This, however, may give each agent an inaccurate reward on its contribution to the group, which could cause inefficient learning. To deal with this problem, we i) introduce a cooperative-game theoretical framework called extended convex game (ECG) that is a superset of global reward game, and ii) propose a local reward approach called Shapley Q-value. Shapley Q-value is able to distribute the global reward, reflecting each agent's own contribution in contrast to the shared reward approach. Moreover, we derive an MARL algorithm called Shapley Q-value deep deterministic policy gradient (SQDDPG), using Shapley Q-value as the critic for each agent. We evaluate SQDDPG on Cooperative Navigation, Prey-and-Predator and Traffic Junction, compared with the state-of-the-art algorithms, e.g., MADDPG, COMA, Independent DDPG and Independent A2C. In the experiments, SQDDPG shows a significant improvement on the convergence rate. Finally, we plot Shapley Q-value and validate the property of fair credit assignment. Yuan Zhang 0027, Tae-Kyun Kim 0001, Yunjie Gu |
AAAI | 2 |