Siqi Liu 0002

dblp:60/9360-2 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-6381-4552ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 43% Multi-agent systems · 29% Language models and text generation · 9%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 24 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.932025
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025
NeuPL: Neural Population Learning · ICLR 2022
A Generalized Training Approach for Multiagent Learning · ICLR 2020
Machine learning › Reinforcement learning
population-based learning
1.122022
Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games · ICML 2022
NeuPL: Neural Population Learning · ICLR 2022
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
1.022022
NeuPL: Neural Population Learning · ICLR 2022
A Generalized Training Approach for Multiagent Learning · ICLR 2020
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Re-evaluating Open-ended Evaluation of Large Language Models · ICLR 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games
0.912025
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium
0.912025
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025
Algorithmic game theory and mechanism design
rating systems
0.912025
Re-evaluating Open-ended Evaluation of Large Language Models · ICLR 2025
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.812024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Algorithmic game theory and mechanism design
equilibrium computation
0.812024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Algorithmic game theory and mechanism design › non-cooperative game
strategic game
0.812024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Knowledge, reasoning and agents › Multi-agent systems
equilibrium computation
0.612022
Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers · NeurIPS 2022
Machine learning › Deep learning architectures and training
equivariant neural network
0.612022
Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers · NeurIPS 2022
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
learning in games
0.612022
Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games · ICML 2022
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Reinforcement learning › policy optimization
on-policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Reinforcement learning
policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning › competitive reinforcement learning
competitive multi-agent reinforcement learning
0.412019
Emergent Coordination Through Competition · ICLR (Poster) 2019
Knowledge, reasoning and agents › Multi-agent systems › swarm robotics
emergent coordination
0.412019
Emergent Coordination Through Competition · ICLR (Poster) 2019
Robotics › Motion planning and robot control › robot control
hierarchical control
0.412019
Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019
Robotics › Motion planning and robot control
humanoid robot control
0.412019
Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019
Computer vision › Vision and language
image captioning
0.312017
Improved Image Captioning via Policy Gradient optimization of SPIDEr · ICCV 2017
Natural language and speech › Language models and text generation
text generation evaluation
0.312017
Improved Image Captioning via Policy Gradient optimization of SPIDEr · ICCV 2017
Knowledge, reasoning and agents › Multi-agent systems › game theory
game-theoretic reasoning
0.212024
NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024
Knowledge, reasoning and agents › Multi-agent systems
game theory
0.112020
A Generalized Training Approach for Multiagent Learning · ICLR 2020

Methods — techniques the papers use, named apart from their topics

equivariant neural network · 2.1game theory · 1.73-player game · 1.7transformer · 1.5gradient descent · 0.9exploitability upper bound · 0.9policy gradient · 0.9variational objective · 0.6population-based training · 0.6conditional network · 0.6
YearPublicationVenuePosition
2025 Re-evaluating Open-ended Evaluation of Large Language Models
abstract
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular solution. Despite their many advantages, we show that the current Elo-based rating systems can be susceptible to and even reinforce biases in data, intentional or accidental, due to their sensitivity to redundancies. To address this issue, we propose evaluation as a 3-player game, and introduce novel game-theoretic solution concepts to ensure robustness to redundancy. We show that our method leads to intuitive ratings and provide insights into the competitive landscape of LLM development.
Siqi Liu 0002, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, Marc Lanctot
ICLR1
2025 Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning
abstract
Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across time. We introduce the class of convex Markov games that allow general convex preferences over occupancy measures. Despite infinite time horizon and strictly higher generality than Markov games, pure strategy Nash equilibria exist. Furthermore, equilibria can be approximated empirically by performing gradient descent on an upper bound of exploitability. Our experiments reveal novel solutions to classic repeated normal-form games, find fair solutions in a repeated asymmetric coordination game, and prioritize safe long-term behavior in a robot warehouse environment. In the prisoner’s dilemma, our algorithm leverages transient imitation to find a policy profile that deviates from observed human play only slightly, yet achieves higher per-player utility while also being three orders of magnitude less exploitable.
Ian Gemp, Andreas Alexander Haupt, Luke Marris, Siqi Liu 0002, Georgios Piliouras
ICML4
2024 NfgTransformer: Equivariant Representation Learning for Normal-form Games
abstract
Normal-form games (NFGs) are the fundamental model of *strategic interaction*. We study their representation using neural networks. We describe the inherent equivariance of NFGs --- any permutation of strategies describes an equivalent game --- as well as the challenges this poses for representation learning. We then propose the NfgTransformer architecture that leverages this equivariance, leading to state-of-the-art performance in a range of game-theoretic tasks including equilibrium-solving, deviation gain estimation and ranking, with a common approach to NFG representation. We show that the resulting model is interpretable and versatile, paving the way towards deep learning systems capable of game-theoretic reasoning when interacting with humans and with each other.
Siqi Liu 0002, Luke Marris, Georgios Piliouras, Ian Gemp, Nicolas Heess
ICLR1
2022 NeuPL: Neural Population Learning
Siqi Liu 0002, Luke Marris, Daniel Hennes, Josh Merel, Nicolas Heess, Thore Graepel
ICLR1
2022 Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games
abstract
Learning to play optimally against any mixture over a diverse set of strategies is of important practical interests in competitive games. In this paper, we propose simplex-NeuPL that satisfies two desiderata simultaneously: i) learning a population of strategically diverse basis policies, represented by a single conditional network; ii) using the same network, learn best-responses to any mixture over the simplex of basis policies. We show that the resulting conditional policies incorporate prior information about their opponents effectively, enabling near optimal returns against arbitrary mixture policies in a game with tractable best-responses. We verify that such policies behave Bayes-optimally under uncertainty and offer insights in using this flexibility at test time. Finally, we offer evidence that learning best-responses to any mixture policies is an effective auxiliary task for strategic exploration, which, by itself, can lead to more performant populations.
Siqi Liu 0002, Marc Lanctot, Luke Marris, Nicolas Heess
ICML1
2022 Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers
abstract
Solution concepts such as Nash Equilibria, Correlated Equilibria, and Coarse Correlated Equilibria are useful components for many multiagent machine learning algorithms. Unfortunately, solving a normal-form game could take prohibitive or non-deterministic time to converge, and could fail. We introduce the Neural Equilibrium Solver which utilizes a special equivariant neural network architecture to approximately solve the space of all games of fixed shape, buying speed and determinism. We define a flexible equilibrium selection framework, that is capable of uniquely selecting an equilibrium that minimizes relative entropy, or maximizes welfare. The network is trained without needing to generate any supervised training data. We show remarkable zero-shot generalization to larger games. We argue that such a network is a powerful component for many possible multiagent algorithms.
Luke Marris, Ian Gemp, Thomas W. Anthony 0001, Andrea Tacchetti, Siqi Liu 0002, Karl Tuyls
NeurIPS5
2020 A Generalized Training Approach for Multiagent Learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland 0001, Karl Tuyls, Julien Pérolat, Siqi Liu 0002, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes 0001, Zhe Wang 0055, Guy Lever, Nicolas Heess, Thore Graepel, Rémi Munos
ICLR6
2020 V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick
ICLR9
2019 Emergent Coordination Through Competition
Siqi Liu 0002, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, Thore Graepel
ICLR (Poster)1
2019 Hierarchical Visuomotor Control of Humanoids
Josh Merel, Arun Ahuja, Saran Tunyasuvunakool, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Greg Wayne
ICLR (Poster)5
2017 Improved Image Captioning via Policy Gradient optimization of SPIDEr
abstract
Current image captioning methods are usually trained via maximum likelihood estimation. However, the log-likelihood score of a caption does not correlate well with human assessments of quality. Standard syntactic evaluation metrics, such as BLEU, METEOR and ROUGE, are also not well correlated. The newer SPICE and CIDEr metrics are better correlated, but have traditionally been hard to optimize for. In this paper, we show how to use a policy gradient (PG) method to directly optimize a linear combination of SPICE and CIDEr (a combination we call SPIDEr): the SPICE score ensures our captions are semantically faithful to the image, while CIDEr score ensures our captions are syntactically fluent. The PG method we propose improves on the prior MIXER approach, by using Monte Carlo rollouts instead of mixing MLE training with PG. We show empirically that our algorithm leads to easier optimization and improved results compared to MIXER. Finally, we show that using our PG method we can optimize any of the metrics, including the proposed SPIDEr metric which results in image captions that are strongly preferred by human raters compared to captions generated by the same model but trained to optimize MLE or the COCO metrics.
Siqi Liu 0002, Zhenhai Zhu, Sergio Guadarrama, Kevin Murphy 0002
ICCV1