EDBT 2026 Demo / reviewers in the wild / expert
David Mguni
dblp:217/2369 · also David Henry Mguni
· DBLP profile ↗
17ranked-venue papers
6as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ensemble Value Functions for Efficient Exploration in Multi-Agent Reinforcement Learning
Lukas Schäfer 0001, Oliver Slumbers, Stephen McAleer, Yali Du 0001, Stefano V. Albrecht, David Mguni |
AAMAS | 6 |
| 2025 | Taming Multi-Agent Reinforcement Learning with Estimator Variance Reduction
Taher Jafferjee, Juliusz Krysztof Ziomek, Tianpei Yang, Zipeng Dai, Matthew E. Taylor, Kun Shao, Jun Wang 0012, David Mguni |
AAMAS | 9 |
| 2025 | A Bilevel Reinforcement Learning Framework with Language Prior Knowledge
Yan Song 0003, Filippos Christianos, David Mguni |
ECML/PKDD (6) | 7 |
| 2023 | Learning to Shape Rewards Using a Game of Two PartnersabstractReward shaping (RS) is a powerful method in reinforcement learning (RL) for overcoming the problem of sparse or uninformative rewards. However, RS typically relies on manually engineered shaping-reward functions whose construc- tion is time-consuming and error-prone. It also requires domain knowledge which runs contrary to the goal of autonomous learning. We introduce Reinforcement Learning Optimising Shaping Algorithm (ROSA), an automated reward shaping framework in which the shaping-reward function is constructed in a Markov game between two agents. A reward-shaping agent (Shaper) uses switching controls to determine which states to add shaping rewards for more efficient learning while the other agent (Controller) learns the optimal policy for the task using these shaped rewards. We prove that ROSA, which adopts existing RL algorithms, learns to construct a shaping-reward function that is beneficial to the task thus ensuring efficient convergence to high performance policies. We demonstrate ROSA’s properties in three didactic experiments and show its superior performance against state-of-the-art RS algorithms in challenging sparse reward environments. David Mguni, Taher Jafferjee, Nicolas Perez Nieves, Wenbin Song, Feifei Tong, Matthew E. Taylor, Tianpei Yang, Zipeng Dai, Jiangcheng Zhu, Kun Shao, Jun Wang 0012, Yaodong Yang 0001 |
AAAI | 1 |
| 2023 | Timing is Everything: Learning to Act Selectively with Costly Actions and Budgetary Constraints
David Mguni, Aivar Sootla, Juliusz Krysztof Ziomek, Oliver Slumbers, Zipeng Dai, Kun Shao, Jun Wang 0012 |
ICLR | 1 |
| 2023 | MANSA: Learning Fast and Slow in Multi-Agent SystemsabstractIn multi-agent reinforcement learning (MARL), independent learning (IL) often shows remarkable performance and easily scales with the number of agents. Yet, using IL can be inefficient and runs the risk of failing to successfully train, particularly in scenarios that require agents to coordinate their actions. Using centralised learning (CL) enables MARL agents to quickly learn how to coordinate their behaviour but employing CL everywhere is often prohibitively expensive in real-world applications. Besides, using CL in value-based methods often needs strong representational constraints (e.g. individual-global-max condition) that can lead to poor performance if violated. In this paper, we introduce a novel plug & play IL framework named Multi-Agent Network Selection Algorithm (MANSA) which selectively employs CL only at states that require coordination. At its core, MANSA has an additional agent that uses switching controls to quickly learn the best states to activate CL during training, using CL only where necessary and vastly reducing the computational burden of CL. Our theory proves MANSA preserves cooperative MARL convergence properties, boosts IL performance and can optimally make use of a fixed budget on the number CL calls. We show empirically in Level-based Foraging (LBF) and StarCraft Multi-agent Challenge (SMAC) that MANSA achieves fast, superior and more reliable performance while making 40% fewer CL calls in SMAC and using CL at only 1% CL calls in LBF. David Mguni, Haojun Chen, Taher Jafferjee, Longfei Yue, Xidong Feng, Stephen McAleer, Feifei Tong, Jun Wang 0012, Yaodong Yang 0001 |
ICML | 1 |
| 2023 | A Game-Theoretic Framework for Managing Risk in Multi-Agent SystemsabstractIn order for agents in multi-agent systems (MAS) to be safe, they need to take into account the risks posed by the actions of other agents. However, the dominant paradigm in game theory (GT) assumes that agents are not affected by risk from other agents and only strive to maximise their expected utility. For example, in hybrid human-AI driving systems, it is necessary to limit large deviations in reward resulting from car crashes. Although there are equilibrium concepts in game theory that take into account risk aversion, they either assume that agents are risk-neutral with respect to the uncertainty caused by the actions of other agents, or they are not guaranteed to exist. We introduce a new GT-based Risk-Averse Equilibrium (RAE) that always produces a solution that minimises the potential variance in reward accounting for the strategy of other agents. Theoretically and empirically, we show RAE shares many properties with a Nash Equilibrium (NE), establishing convergence properties and generalising to risk-dominant NE in certain cases. To tackle large-scale problems, we extend RAE to the PSRO multi-agent reinforcement learning (MARL) framework. We empirically demonstrate the minimum reward variance benefits of RAE in matrix games with high-risk outcomes. Results on MARL experiments show RAE generalises to risk-dominant NE in a trust dilemma game and that it reduces instances of crashing by 7x in an autonomous driving setting versus the best performing baseline. Oliver Slumbers, David Mguni, Stefano B. Blumberg, Stephen McAleer, Yaodong Yang 0001, Jun Wang 0012 |
ICML | 2 |
| 2023 | ChessGPT: Bridging Policy Learning and Language ModelingabstractWhen solving decision-making tasks, humans typically depend on information from two key sources: (1) Historical policy data, which provides interaction replay from the environment, and (2) Analytical insights in natural language form, exposing the invaluable thought process or strategic considerations. Despite this, the majority of preceding research focuses on only one source: they either use historical replay exclusively to directly learn policy or value functions, or engaged in language model training utilizing mere language corpus. In this paper, we argue that a powerful autonomous agent should cover both sources. Thus, we propose ChessGPT, a GPT model bridging policy learning and language modeling by integrating data from these two sources in Chess games. Specifically, we build a large-scale game and language dataset related to chess. Leveraging the dataset, we showcase two model examples ChessCLIP and ChessGPT, integrating policy learning and language modeling. Finally, we propose a full evaluation framework for evaluating language model's chess ability. Experimental results validate our model and dataset's effectiveness. We open source our code, model, and dataset at https://github.com/waterhorse1/ChessGPT. Xidong Feng, Yicheng Luo, Hongrui Tang, Mengyue Yang, Kun Shao, David Mguni, Yali Du 0001, Jun Wang 0012 |
NeurIPS | 7 |
| 2023 | Online Markov decision processes with non-oblivious strategic adversary
Le Cong Dinh, David Mguni, Long Tran-Thanh, Jun Wang 0012, Yaodong Yang 0001 |
Auton. Agents Multi Agent Syst. | 2 |
| 2022 | LIGS: Learnable Intrinsic-Reward Generation Selection for Multi-Agent Learning
David Mguni, Taher Jafferjee, Nicolas Perez Nieves, Oliver Slumbers, Feifei Tong, Yang Li 0116, Jiangcheng Zhu, Yaodong Yang 0001, Jun Wang 0012 |
ICLR | 1 |
| 2022 | Saute RL: Almost Surely Safe Reinforcement Learning Using State AugmentationabstractSatisfying safety constraints almost surely (or with probability one) can be critical for the deployment of Reinforcement Learning (RL) in real-life applications. For example, plane landing and take-off should ideally occur with probability one. We address the problem by introducing Safety Augmented (Saute) Markov Decision Processes (MDPs), where the safety constraints are eliminated by augmenting them into the state-space and reshaping the objective. We show that Saute MDP satisfies the Bellman equation and moves us closer to solving Safe RL with constraints satisfied almost surely. We argue that Saute MDP allows viewing the Safe RL problem from a different perspective enabling new features. For instance, our approach has a plug-and-play nature, i.e., any RL algorithm can be "Sauteed”. Additionally, state augmentation allows for policy generalization across safety constraints. We finally show that Saute RL algorithms can outperform their state-of-the-art counterparts when constraint satisfaction is of high importance. Aivar Sootla, Alexander I. Cowen-Rivers, Taher Jafferjee, David Mguni, Jun Wang 0012, Haitham Bou-Ammar |
ICML | 5 |
| 2022 | On the Convergence of Fictitious Play: A Decomposition ApproachabstractFictitious play (FP) is one of the most fundamental game-theoretical learning frameworks for computing Nash equilibrium in n-player games, which builds the foundation for modern multi-agent learning algorithms. Although FP has provable convergence guarantees on zero-sum games and potential games, many real-world problems are often a mixture of both and the convergence property of FP has not been fully studied yet. In this paper, we extend the convergence results of FP to the combinations of such games and beyond. Specifically, we derive new conditions for FP to converge by leveraging game decomposition techniques. We further develop a linear relationship unifying cooperation and competition in the sense that these two classes of games are mutually transferable. Finally, we analyse a non-convergent example of FP, the Shapley game, and develop sufficient conditions for FP to converge. Yurong Chen 0002, Xiaotie Deng, David Mguni, Jun Wang 0012, Yaodong Yang 0001 |
IJCAI | 4 |
| 2021 | Learning in Nonzero-Sum Stochastic Games with PotentialsabstractMulti-agent reinforcement learning (MARL) has become effective in tackling discrete cooperative game scenarios. However, MARL has yet to penetrate settings beyond those modelled by team and zero-sum games, confining it to a small subset of multi-agent systems. In this paper, we introduce a new generation of MARL learners that can handle \textit{nonzero-sum} payoff structures and continuous settings. In particular, we study the MARL problem in a class of games known as stochastic potential games (SPGs) with continuous state-action spaces. Unlike cooperative games, in which all agents share a common reward, SPGs are capable of modelling real-world scenarios where agents seek to fulfil their individual goals. We prove theoretically our learning method, $\ourmethod$, enables independent agents to learn Nash equilibrium strategies in \textit{polynomial time}. We demonstrate our framework tackles previously unsolvable tasks such as \textit{Coordination Navigation} and \textit{large selfish routing games} and that it outperforms the state of the art MARL baselines such as MADDPG and COMIX in such scenarios. David Mguni, Yutong Wu 0005, Yali Du 0001, Yaodong Yang 0001, Ziyi Wang 0004, Minne Li, Ying Wen 0001, Joel Jennings, Jun Wang 0012 |
ICML | 1 |
| 2021 | Modelling Behavioural Diversity for Learning in Open-Ended GamesabstractPromoting behavioural diversity is critical for solving games with non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). Yet, there is a lack of rigorous treatment for defining diversity and constructing diversity-aware learning dynamics. In this work, we offer a geometric interpretation of behavioural diversity in games and introduce a novel diversity metric based on \emph{determinantal point processes} (DPP). By incorporating the diversity metric into best-response dynamics, we develop \emph{diverse fictitious play} and \emph{diverse policy-space response oracle} for solving normal-form games and open-ended games. We prove the uniqueness of the diverse best response and the convergence of our algorithms on two-player games. Importantly, we show that maximising the DPP-based diversity metric guarantees to enlarge the \emph{gamescape} – convex polytopes spanned by agents’ mixtures of strategies. To validate our diversity-aware solvers, we test on tens of games that show strong non-transitivity. Results suggest that our methods achieve at least the same, and in most games, lower exploitability than PSRO solvers by finding effective and diverse strategies. Nicolas Perez Nieves, Yaodong Yang 0001, Oliver Slumbers, David Mguni, Ying Wen 0001, Jun Wang 0012 |
ICML | 4 |
| 2021 | Settling the Variance of Multi-Agent Policy GradientsabstractPolicy gradient (PG) methods are popular reinforcement learning (RL) methods where a baseline is often applied to reduce the variance of gradient estimates. In multi-agent RL (MARL), although the PG theorem can be naturally extended, the effectiveness of multi-agent PG (MAPG) methods degrades as the variance of gradient estimates increases rapidly with the number of agents. In this paper, we offer a rigorous analysis of MAPG methods by, firstly, quantifying the contributions of the number of agents and agents' explorations to the variance of MAPG estimators. Based on this analysis, we derive the optimal baseline (OB) that achieves the minimal variance. In comparison to the OB, we measure the excess variance of existing MARL algorithms such as vanilla MAPG and COMA. Considering using deep neural networks, we also propose a surrogate version of OB, which can be seamlessly plugged into any existing PG methods in MARL. On benchmarks of Multi-Agent MuJoCo and StarCraft challenges, our OB technique effectively stabilises training and improves the performance of multi-agent PPO and COMA algorithms by a significant margin. Code is released at \url{https://github.com/morning9393/Optimal-Baseline-for-Multi-agent-Policy-Gradients}. Jakub Grudzien Kuba, Muning Wen, Linghui Meng 0001, Shangding Gu, Haifeng Zhang 0002, David Mguni, Jun Wang 0012, Yaodong Yang 0001 |
NeurIPS | 6 |
| 2020 | Multi-Agent Determinantal Q-LearningabstractCentralized training with decentralized execution has become an important paradigm in multi-agent learning. Though practical, current methods rely on restrictive assumptions to decompose the centralized value function across agents for execution. In this paper, we eliminate this restriction by proposing multi-agent determinantal Q-learning. Our method is established on Q-DPP, a novel extension of determinantal point process (DPP) to multi-agent setting. Q-DPP promotes agents to acquire diverse behavioral models; this allows a natural factorization of the joint Q-functions with no need for \emph{a priori} structural constraints on the value function or special network architectures. We demonstrate that Q-DPP generalizes major solutions including VDN, QMIX, and QTRAN on decentralizable cooperative tasks. To efficiently draw samples from Q-DPP, we develop a linear-time sampler with theoretical approximation guarantee. Our sampler also benefits exploration by coordinating agents to cover orthogonal directions in the state space during training. We evaluate our algorithm on multiple cooperative benchmarks; its effectiveness has been demonstrated when compared with the state-of-the-art. Yaodong Yang 0001, Ying Wen 0001, Jun Wang 0012, Kun Shao, David Mguni, Weinan Zhang 0001 |
ICML | 6 |
| 2018 | Decentralised Learning in Systems With Many, Many Strategic AgentsabstractAlthough multi-agent reinforcement learning can tackle systems of strategically interacting entities, it currently fails in scalability and lacks rigorous convergence guarantees. Crucially, learning in multi-agent systems can become intractable due to the explosion in the size of the state-action space as the number of agents increases. In this paper, we propose a method for computing closed-loop optimal policies in multi-agent systems that scales independently of the number of agents. This allows us to show, for the first time, successful convergence to optimal behaviour in systems with an unbounded number of interacting adaptive learners. Studying the asymptotic regime of N-player stochastic games, we devise a learning protocol that is guaranteed to converge to equilibrium policies even when the number of agents is extremely large. Our method is model-free and completely decentralised so that each agent need only observe its local state information and its realised rewards. We validate these theoretical results by showing convergence to Nash-equilibrium policies in applications from economics and control theory with thousands of strategically interacting agents. David Mguni, Joel Jennings, Enrique Munoz de Cote |
AAAI | 1 |