Luke Marris

dblp:223/4422 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 45% Multi-agent systems · 21% Generative modeling · 10%
Theoretical computer science
5 papers
Algorithmic game theory and mechanism design · 83% Mathematical optimization · 17%

Topics — the 29 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
3.352025
Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025
NeuPL: Neural Population Learning · ICLR 2022
Algorithmic game theory and mechanism design
equilibrium computation
2.032024
Generative Adversarial Equilibrium Solvers · ICLR 2024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers · ICML 2021
Machine learning › Reinforcement learning
population-based learning
1.122022
Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games · ICML 2022
NeuPL: Neural Population Learning · ICLR 2022
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
1.022022
NeuPL: Neural Population Learning · ICLR 2022
A Generalized Training Approach for Multiagent Learning · ICLR 2020
Machine learning › Generative modeling
generative model
0.912025
Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Re-evaluating Open-ended Evaluation of Large Language Models · ICLR 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games
0.912025
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.912025
Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium
0.912025
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent modeling
0.912025
Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025
Algorithmic game theory and mechanism design
rating systems
0.912025
Re-evaluating Open-ended Evaluation of Large Language Models · ICLR 2025
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.812024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Machine learning › Generative modeling
generative adversarial network
0.812024
Generative Adversarial Equilibrium Solvers · ICLR 2024
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts › nash equilibrium
approximate nash equilibrium
0.812024
Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization · ICLR 2024
Algorithmic game theory and mechanism design › market equilibrium
competitive equilibrium
0.812024
Generative Adversarial Equilibrium Solvers · ICLR 2024
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
generalized nash equilibrium
0.812024
Generative Adversarial Equilibrium Solvers · ICLR 2024
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium
0.812024
Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization · ICLR 2024
Mathematical optimization › stochastic optimization
stochastic nonconvex optimization
0.812024
Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization · ICLR 2024
Mathematical optimization
stochastic optimization
0.812024
Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization · ICLR 2024
Algorithmic game theory and mechanism design › non-cooperative game
strategic game
0.812024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Knowledge, reasoning and agents › Multi-agent systems
equilibrium computation
0.612022
Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers · NeurIPS 2022
Machine learning › Deep learning architectures and training
equivariant neural network
0.612022
Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers · NeurIPS 2022
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
learning in games
0.612022
Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games · ICML 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
general-sum game
0.512021
Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers · ICML 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning › equilibrium learning
policy space response oracle
0.512021
Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers · ICML 2021
Algorithmic game theory and mechanism design › equilibrium computation
correlated equilibrium
0.512021
Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers · ICML 2021
Machine learning › Deep learning architectures and training › biologically plausible learning › feedback alignment
backpropagation alternative
0.312018
Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures · NeurIPS 2018
Knowledge, reasoning and agents › Multi-agent systems › game theory
game-theoretic reasoning
0.212024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Knowledge, reasoning and agents › Multi-agent systems
game theory
0.112020
A Generalized Training Approach for Multiagent Learning · ICLR 2020

Methods — techniques the papers use, named apart from their topics

equivariant neural network · 2.1game theory · 1.73-player game · 1.7transformer · 1.5generative adversarial learning · 1.5policy space response oracles · 0.9gradient descent · 0.9exploitability upper bound · 0.9best response · 0.9bargaining theory · 0.9stochastic gradient descent · 0.8neural network function approximation · 0.8monte carlo estimation · 0.8maximum gini correlated equilibrium · 0.5joint policy-space response oracles · 0.5
YearPublicationVenuePosition
2025 Re-evaluating Open-ended Evaluation of Large Language Models
abstract
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular solution. Despite their many advantages, we show that the current Elo-based rating systems can be susceptible to and even reinforce biases in data, intentional or accidental, due to their sensitivity to redundancies. To address this issue, we propose evaluation as a 3-player game, and introduce novel game-theoretic solution concepts to ensure robustness to redundancy. We show that our method leads to intuitive ratings and provide insights into the competitive landscape of LLM development.
Siqi Liu 0002, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, Marc Lanctot
ICLR3
2025 Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning
abstract
Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across time. We introduce the class of convex Markov games that allow general convex preferences over occupancy measures. Despite infinite time horizon and strictly higher generality than Markov games, pure strategy Nash equilibria exist. Furthermore, equilibria can be approximated empirically by performing gradient descent on an upper bound of exploitability. Our experiments reveal novel solutions to classic repeated normal-form games, find fair solutions in a repeated asymmetric coordination game, and prioritize safe long-term behavior in a robot warehouse environment. In the prisoner’s dilemma, our algorithm leverages transient imitation to find a policy profile that deviates from observed human play only slightly, yet achieves higher per-player utility while also being three orders of magnitude less exploitable.
Ian Gemp, Andreas Alexander Haupt, Luke Marris, Siqi Liu 0002, Georgios Piliouras
ICML3
2025 Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling
abstract
Opponent modeling methods typically involve two crucial steps: building a belief distribution over opponents' strategies, and exploiting this opponent model by playing a best response. However, existing approaches typically require domain-specific heurstics to come up with such a model, and algorithms for approximating best responses are hard to scale in large, imperfect information domains. In this work, we introduce a scalable and generic multiagent training regime for opponent modeling using deep game-theoretic reinforcement learning. We first propose Generative Best Respoonse (GenBR), a best response algorithm based on Monte-Carlo Tree Search (MCTS) with a learned deep generative model that samples world states during planning. This new method scales to large imperfect information domains and can be plug and play in a variety of multiagent algorithms. We use this new method under the framework of Policy Space Response Oracles (PSRO), to automate the generation of an offline opponent model via iterative game-theoretic reasoning and population-based training. We propose using solution concepts based on bargaining theory to build up an opponent mixture, which we find identifying profiles that are near the Pareto frontier. Then GenBR keeps updating an online opponent model and reacts against it during gameplay. We conduct behavioral studies where human participants negotiate with our agents in Deal-or-No-Deal, a class of bilateral bargaining games. Search with generative modeling finds stronger policies during both training time and test time, enables online Bayesian co-player prediction, and can produce agents that achieve comparable social welfare and Nash bargaining score negotiating with humans as humans trading among themselves.
Zun Li 0002, Marc Lanctot, Kevin R. McKee, Luke Marris, Ian Gemp, Daniel Hennes, Paul Muller, Kate Larson, Yoram Bachrach, Michael P. Wellman
IJCAI4
2024 NfgTransformer: Equivariant Representation Learning for Normal-form Games
abstract
Normal-form games (NFGs) are the fundamental model of *strategic interaction*. We study their representation using neural networks. We describe the inherent equivariance of NFGs --- any permutation of strategies describes an equivalent game --- as well as the challenges this poses for representation learning. We then propose the NfgTransformer architecture that leverages this equivariance, leading to state-of-the-art performance in a range of game-theoretic tasks including equilibrium-solving, deviation gain estimation and ranking, with a common approach to NFG representation. We show that the resulting model is interpretable and versatile, paving the way towards deep learning systems capable of game-theoretic reasoning when interacting with humans and with each other.
Siqi Liu 0002, Luke Marris, Georgios Piliouras, Ian Gemp, Nicolas Heess
ICLR2
2024 Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization
abstract
We propose the first loss function for approximate Nash equilibria of normal-form games that is amenable to unbiased Monte Carlo estimation. This construction allows us to deploy standard non-convex stochastic optimization techniques for approximating Nash equilibria, resulting in novel algorithms with provable guarantees. We complement our theoretical analysis with experiments demonstrating that stochastic gradient descent can outperform previous state-of-the-art approaches.
Ian Gemp, Luke Marris, Georgios Piliouras
ICLR2
2024 Generative Adversarial Equilibrium Solvers
abstract
We introduce the use of generative adversarial learning to compute equilibria in general game-theoretic settings, specifically the generalized Nash equilibrium (GNE) in pseudo-games, and its specific instantiation as the competitive equilibrium (CE) in Arrow-Debreu competitive economies. Pseudo-games are a generalization of games in which players' actions affect not only the payoffs of other players but also their feasible action spaces. Although the computation of GNE and CE is intractable in the worst-case, i.e., PPAD-hard, in practice, many applications only require solutions with high accuracy in expectation over a distribution of problem instances. We introduce Generative Adversarial Equilibrium Solvers (GAES): a family of generative adversarial neural networks that can learn GNE and CE from only a sample of problem instances. We provide computational and sample complexity bounds for Lipschitz-smooth function approximators in a large class of concave pseudo-games, and apply the framework to finding Nash equilibria in normal-form games, CE in Arrow-Debreu competitive economies, and GNE in an environmental economic model of the Kyoto mechanism.
Denizalp Goktas, David C. Parkes, Ian Gemp, Luke Marris, Georgios Piliouras, Romuald Elie, Guy Lever, Andrea Tacchetti
ICLR4
2022 NeuPL: Neural Population Learning
Siqi Liu 0002, Luke Marris, Daniel Hennes, Josh Merel, Nicolas Heess, Thore Graepel
ICLR2
2022 Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games
abstract
Learning to play optimally against any mixture over a diverse set of strategies is of important practical interests in competitive games. In this paper, we propose simplex-NeuPL that satisfies two desiderata simultaneously: i) learning a population of strategically diverse basis policies, represented by a single conditional network; ii) using the same network, learn best-responses to any mixture over the simplex of basis policies. We show that the resulting conditional policies incorporate prior information about their opponents effectively, enabling near optimal returns against arbitrary mixture policies in a game with tractable best-responses. We verify that such policies behave Bayes-optimally under uncertainty and offer insights in using this flexibility at test time. Finally, we offer evidence that learning best-responses to any mixture policies is an effective auxiliary task for strategic exploration, which, by itself, can lead to more performant populations.
Siqi Liu 0002, Marc Lanctot, Luke Marris, Nicolas Heess
ICML3
2022 Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers
abstract
Solution concepts such as Nash Equilibria, Correlated Equilibria, and Coarse Correlated Equilibria are useful components for many multiagent machine learning algorithms. Unfortunately, solving a normal-form game could take prohibitive or non-deterministic time to converge, and could fail. We introduce the Neural Equilibrium Solver which utilizes a special equivariant neural network architecture to approximately solve the space of all games of fixed shape, buying speed and determinism. We define a flexible equilibrium selection framework, that is capable of uniquely selecting an equilibrium that minimizes relative entropy, or maximizes welfare. The network is trained without needing to generate any supervised training data. We show remarkable zero-shot generalization to larger games. We argue that such a network is a powerful component for many possible multiagent algorithms.
Luke Marris, Ian Gemp, Thomas W. Anthony 0001, Andrea Tacchetti, Siqi Liu 0002, Karl Tuyls
NeurIPS1
2021 Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers
abstract
Two-player, constant-sum games are well studied in the literature, but there has been limited progress outside of this setting. We propose Joint Policy-Space Response Oracles (JPSRO), an algorithm for training agents in n-player, general-sum extensive form games, which provably converges to an equilibrium. We further suggest correlated equilibria (CE) as promising meta-solvers, and propose a novel solution concept Maximum Gini Correlated Equilibrium (MGCE), a principled and computationally efficient family of solutions for solving the correlated equilibrium selection problem. We conduct several experiments using CE meta-solvers for JPSRO and demonstrate convergence on n-player, general-sum games.
Luke Marris, Paul Muller, Marc Lanctot, Karl Tuyls, Thore Graepel
ICML1
2020 A Generalized Training Approach for Multiagent Learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland 0001, Karl Tuyls, Julien Pérolat, Siqi Liu 0002, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes 0001, Zhe Wang 0055, Guy Lever, Nicolas Heess, Thore Graepel, Rémi Munos
ICLR8
2018 Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures
abstract
The backpropagation of error algorithm (BP) is impossible to implement in a real brain. The recent success of deep networks in machine learning and AI, however, has inspired proposals for understanding how the brain might learn across multiple layers, and hence how it might approximate BP. As of yet, none of these proposals have been rigorously evaluated on tasks where BP-guided deep learning has proved critical, or in architectures more structured than simple fully-connected networks. Here we present results on scaling up biologically motivated models of deep learning on datasets which need deep networks with appropriate architectures to achieve good performance. We present results on the MNIST, CIFAR-10, and ImageNet datasets and explore variants of target-propagation (TP) and feedback alignment (FA) algorithms, and explore performance in both fully- and locally-connected architectures. We also introduce weight-transport-free variants of difference target propagation (DTP) modified to remove backpropagation from the penultimate layer. Many of these algorithms perform well for MNIST, but for CIFAR and ImageNet we find that TP and FA variants perform significantly worse than BP, especially for networks composed of locally connected units, opening questions about whether new architectures and algorithms are required to scale these approaches. Our results and implementation details help establish baselines for biologically motivated deep learning schemes going forward.
Sergey Bartunov, Adam Santoro, Blake A. Richards, Luke Marris, Geoffrey E. Hinton, Timothy P. Lillicrap
NeurIPS4