Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yuan Zhang 0027

dblp:48/2168-27 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 69% Multi-agent systems · 11% Motion planning and robot control · 9%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.852024
Open Ad Hoc Teamwork with Cooperative Game Theory · ICML 2024
Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach · ICAPS 2024
SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment
1.022022
SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning · NeurIPS 2022
Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games · AAAI 2020
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
ad hoc teamwork
0.812024
Open Ad Hoc Teamwork with Cooperative Game Theory · ICML 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.812024
Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach · ICAPS 2024
Robotics › Motion planning and robot control
robot learning
0.812024
Learning Continuous Control with Geometric Regularity from Robot Intrinsic Symmetry · ICRA 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
0.612022
SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning · NeurIPS 2022
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.512021
Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework
0.512021
Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.512021
Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021
Natural language and speech › Language models and text generation
text generation
0.512021
Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System · ICLR 2021
Algorithmic game theory and mechanism design
cooperative game theory
0.412020
Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games · AAAI 2020
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making
0.212024
Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach · ICAPS 2024

Methods — techniques the papers use, named apart from their topics

parameter sharing · 0.8offline reinforcement learning · 0.8local critic · 0.8graph neural network · 0.8geometric regularity · 0.8cooperative game theory · 0.8centralized training decentralized execution · 0.8PPO · 0.8shapley value · 0.6q-learning · 0.6shapley q-value · 0.4deep deterministic policy gradient · 0.4
YearPublicationVenuePosition
2024 UDUC: An Uncertainty-Driven Approach for Learning-Based Robust Control
abstract
Learning-based techniques have become popular in both model predictive control (MPC) and reinforcement learning (RL). Probabilistic ensemble (PE) models offer a promising approach for modelling system dynamics, showcasing the ability to capture uncertainty and scalability in high-dimensional control scenarios. However, PE models are susceptible to mode collapse, resulting in non-robust control when faced with environments slightly different from the training set. In this paper, we introduce the uncertainty-driven robust control (UDUC) loss as an alternative objective for training PE models, drawing inspiration from contrastive learning. We analyze the robustness of the UDUC loss through the lens of robust optimization and evaluate its performance on the challenging real-world reinforcement learning (RWRL) benchmark, which involves significant environmental mismatches between the training and testing environments.
Yuan Zhang 0027, Jasper Hoffmann, Joschka Boedecker
ECAI1
2024 Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic Approach
abstract
The complex railway network is a challenging real-world multi-agent system usually involving thousands of agents. Current planning methods heavily depend on expert knowledge to formulate solutions for specific cases and are therefore hardly generalized to new scenarios, on which multi-agent reinforcement learning (MARL) draws significant attention. Despite some successful applications in multi-agent decision-making tasks, MARL is hard to scale to a large number of agents. This paper rethinks the curse of agents in the centralized-training-decentralized-execution (CTDE) paradigm and proposes a local-critic approach to address the issue. By combining the local critic with the PPO algorithm, we design a deep MARL algorithm denoted as local-critic PPO (LCPPO). In experiments, we evaluate the effectiveness of LCPPO on a complex railway network benchmark, Flatland, with various numbers of agents. Noticeably, LCPPO shows prominent generalizability and robustness under the changes of environments.
Yuan Zhang 0027, Umashankar Deekshith, Joschka Boedecker
ICAPS1
2024 Open Ad Hoc Teamwork with Cooperative Game Theory
abstract
Ad hoc teamwork poses a challenging problem, requiring the design of an agent to collaborate with teammates without prior coordination or joint training. Open ad hoc teamwork (OAHT) further complicates this challenge by considering environments with a changing number of teammates, referred to as open teams. One promising solution in practice to this problem is leveraging the generalizability of graph neural networks to handle an unrestricted number of agents with various agent-types, named graph-based policy learning (GPL). However, its joint Q-value representation over a coordination graph lacks convincing explanations. In this paper, we establish a new theory to understand the representation of the joint Q-value for OAHT and its learning paradigm, through the lens of cooperative game theory. Building on our theory, we propose a novel algorithm named CIAO, based on GPL's framework, with additional provable implementation tricks that can facilitate learning. The demos of experimental results are available on https://sites.google.com/view/ciao2024, and the code of experiments is published on https://github.com/hsvgbkhgbv/CIAO.
Yang Li 0116, Yuan Zhang 0027, Wei Pan 0004, Samuel Kaski
ICML3
2024 Learning Continuous Control with Geometric Regularity from Robot Intrinsic Symmetry
abstract
Geometric regularity, which leverages data symmetry, has been successfully incorporated into deep learning architectures such as CNNs, RNNs, GNNs, and Transformers. While this concept has been widely applied in robotics to address the curse of dimensionality when learning from high-dimensional data, the inherent reflectional and rotational symmetry of robot structures has not been adequately explored. Drawing inspiration from cooperative multi-agent reinforcement learning, we introduce novel network structures for single-agent control learning that explicitly capture these symmetries. Moreover, we investigate the relationship between the geometric prior and the concept of Parameter Sharing in multi-agent reinforcement learning. Last but not the least, we implement the proposed framework in online and offline learning methods to demonstrate its ease of use. Through experiments conducted on various challenging continuous control tasks on simulators and real robots, we highlight the significant potential of the proposed geometric regularity in enhancing robot learning capabilities.
Shengchao Yan, Baohe Zhang, Yuan Zhang 0027, Joschka Boedecker, Wolfram Burgard
ICRA3
2022 SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning
abstract
Value factorisation is a useful technique for multi-agent reinforcement learning (MARL) in global reward game, however, its underlying mechanism is not yet fully understood. This paper studies a theoretical framework for value factorisation with interpretability via Shapley value theory. We generalise Shapley value to Markov convex game called Markov Shapley value (MSV) and apply it as a value factorisation method in global reward game, which is obtained by the equivalence between the two games. Based on the properties of MSV, we derive Shapley-Bellman optimality equation (SBOE) to evaluate the optimal MSV, which corresponds to an optimal joint deterministic policy. Furthermore, we propose Shapley-Bellman operator (SBO) that is proved to solve SBOE. With a stochastic approximation and some transformations, a new MARL algorithm called Shapley Q-learning (SHAQ) is established, the implementation of which is guided by the theoretical results of SBO and MSV. We also discuss the relationship between SHAQ and relevant value factorisation methods. In the experiments, SHAQ exhibits not only superior performances on all tasks but also the interpretability that agrees with the theoretical analysis. The implementation of this paper is placed on https://github.com/hsvgbkhgbv/shapley-q-learning.
Yuan Zhang 0027, Yunjie Gu, Tae-Kyun Kim 0001
NeurIPS2
2021 Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System
Yuan Zhang 0027, Tae-Kyun Kim 0001, Yunjie Gu
ICLR2
2020 Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games
abstract
Cooperative game is a critical research area in the multi-agent reinforcement learning (MARL). Global reward game is a subclass of cooperative games, where all agents aim to maximize the global reward. Credit assignment is an important problem studied in the global reward game. Most of previous works stood by the view of non-cooperative-game theoretical framework with the shared reward approach, i.e., each agent being assigned a shared global reward directly. This, however, may give each agent an inaccurate reward on its contribution to the group, which could cause inefficient learning. To deal with this problem, we i) introduce a cooperative-game theoretical framework called extended convex game (ECG) that is a superset of global reward game, and ii) propose a local reward approach called Shapley Q-value. Shapley Q-value is able to distribute the global reward, reflecting each agent's own contribution in contrast to the shared reward approach. Moreover, we derive an MARL algorithm called Shapley Q-value deep deterministic policy gradient (SQDDPG), using Shapley Q-value as the critic for each agent. We evaluate SQDDPG on Cooperative Navigation, Prey-and-Predator and Traffic Junction, compared with the state-of-the-art algorithms, e.g., MADDPG, COMA, Independent DDPG and Independent A2C. In the experiments, SQDDPG shows a significant improvement on the convergence rate. Finally, we plot Shapley Q-value and validate the property of fair credit assignment.
Yuan Zhang 0027, Tae-Kyun Kim 0001, Yunjie Gu
AAAI2