VLDB 2026 Research / reviewers in the wild / expert
Siqi Liu 0002
dblp:60/9360-2
· DBLP profile ↗
11ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-6381-4552ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Reinforcement learning · 43% Multi-agent systems · 29% Language models and text generation · 9% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 100% |
Topics — the 24 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.9 | 3 | 2025 | Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025 NeuPL: Neural Population Learning · ICLR 2022 A Generalized Training Approach for Multiagent Learning · ICLR 2020 |
Machine learning › Reinforcement learning
population-based learning |
1.1 | 2 | 2022 | Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games · ICML 2022 NeuPL: Neural Population Learning · ICLR 2022 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning |
1.0 | 2 | 2022 | NeuPL: Neural Population Learning · ICLR 2022 A Generalized Training Approach for Multiagent Learning · ICLR 2020 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Re-evaluating Open-ended Evaluation of Large Language Models · ICLR 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games |
0.9 | 1 | 2025 | Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium |
0.9 | 1 | 2025 | Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning · ICML 2025 |
Algorithmic game theory and mechanism design
rating systems |
0.9 | 1 | 2025 | Re-evaluating Open-ended Evaluation of Large Language Models · ICLR 2025 |
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning |
0.8 | 1 | 2024 | NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024 |
Algorithmic game theory and mechanism design
equilibrium computation |
0.8 | 1 | 2024 | NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024 |
Algorithmic game theory and mechanism design › non-cooperative game
strategic game |
0.8 | 1 | 2024 | NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024 |
Knowledge, reasoning and agents › Multi-agent systems
equilibrium computation |
0.6 | 1 | 2022 | Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
equivariant neural network |
0.6 | 1 | 2022 | Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers · NeurIPS 2022 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
learning in games |
0.6 | 1 | 2022 | Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games · ICML 2022 |
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
on-policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › competitive reinforcement learning
competitive multi-agent reinforcement learning |
0.4 | 1 | 2019 | Emergent Coordination Through Competition · ICLR (Poster) 2019 |
Knowledge, reasoning and agents › Multi-agent systems › swarm robotics
emergent coordination |
0.4 | 1 | 2019 | Emergent Coordination Through Competition · ICLR (Poster) 2019 |
Robotics › Motion planning and robot control › robot control
hierarchical control |
0.4 | 1 | 2019 | Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019 |
Robotics › Motion planning and robot control
humanoid robot control |
0.4 | 1 | 2019 | Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019 |
Computer vision › Vision and language
image captioning |
0.3 | 1 | 2017 | Improved Image Captioning via Policy Gradient optimization of SPIDEr · ICCV 2017 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.3 | 1 | 2017 | Improved Image Captioning via Policy Gradient optimization of SPIDEr · ICCV 2017 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
game-theoretic reasoning |
0.2 | 1 | 2024 | NfgTransformer: Equivariant Representation Learning for Normal-form Games · ICLR 2024 |
Knowledge, reasoning and agents › Multi-agent systems
game theory |
0.1 | 1 | 2020 | A Generalized Training Approach for Multiagent Learning · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
equivariant neural network · 2.1game theory · 1.73-player game · 1.7transformer · 1.5gradient descent · 0.9exploitability upper bound · 0.9policy gradient · 0.9variational objective · 0.6population-based training · 0.6conditional network · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Re-evaluating Open-ended Evaluation of Large Language ModelsabstractEvaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular solution. Despite their many advantages, we show that the current Elo-based rating systems can be susceptible to and even reinforce biases in data, intentional or accidental, due to their sensitivity to redundancies. To address this issue, we propose evaluation as a 3-player game, and introduce novel game-theoretic solution concepts to ensure robustness to redundancy. We show that our method leads to intuitive ratings and provide insights into the competitive landscape of LLM development. Siqi Liu 0002, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, Marc Lanctot |
ICLR | 1 |
| 2025 | Convex Markov Games: A New Frontier for Multi-Agent Reinforcement LearningabstractBehavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across time. We introduce the class of convex Markov games that allow general convex preferences over occupancy measures. Despite infinite time horizon and strictly higher generality than Markov games, pure strategy Nash equilibria exist. Furthermore, equilibria can be approximated empirically by performing gradient descent on an upper bound of exploitability. Our experiments reveal novel solutions to classic repeated normal-form games, find fair solutions in a repeated asymmetric coordination game, and prioritize safe long-term behavior in a robot warehouse environment. In the prisoner’s dilemma, our algorithm leverages transient imitation to find a policy profile that deviates from observed human play only slightly, yet achieves higher per-player utility while also being three orders of magnitude less exploitable. Ian Gemp, Andreas Alexander Haupt, Luke Marris, Siqi Liu 0002, Georgios Piliouras |
ICML | 4 |
| 2024 | NfgTransformer: Equivariant Representation Learning for Normal-form GamesabstractNormal-form games (NFGs) are the fundamental model of *strategic interaction*. We study their representation using neural networks. We describe the inherent equivariance of NFGs --- any permutation of strategies describes an equivalent game --- as well as the challenges this poses for representation learning. We then propose the NfgTransformer architecture that leverages this equivariance, leading to state-of-the-art performance in a range of game-theoretic tasks including equilibrium-solving, deviation gain estimation and ranking, with a common approach to NFG representation. We show that the resulting model is interpretable and versatile, paving the way towards deep learning systems capable of game-theoretic reasoning when interacting with humans and with each other. Siqi Liu 0002, Luke Marris, Georgios Piliouras, Ian Gemp, Nicolas Heess |
ICLR | 1 |
| 2022 | NeuPL: Neural Population Learning
Siqi Liu 0002, Luke Marris, Daniel Hennes, Josh Merel, Nicolas Heess, Thore Graepel |
ICLR | 1 |
| 2022 | Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum GamesabstractLearning to play optimally against any mixture over a diverse set of strategies is of important practical interests in competitive games. In this paper, we propose simplex-NeuPL that satisfies two desiderata simultaneously: i) learning a population of strategically diverse basis policies, represented by a single conditional network; ii) using the same network, learn best-responses to any mixture over the simplex of basis policies. We show that the resulting conditional policies incorporate prior information about their opponents effectively, enabling near optimal returns against arbitrary mixture policies in a game with tractable best-responses. We verify that such policies behave Bayes-optimally under uncertainty and offer insights in using this flexibility at test time. Finally, we offer evidence that learning best-responses to any mixture policies is an effective auxiliary task for strategic exploration, which, by itself, can lead to more performant populations. Siqi Liu 0002, Marc Lanctot, Luke Marris, Nicolas Heess |
ICML | 1 |
| 2022 | Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium SolversabstractSolution concepts such as Nash Equilibria, Correlated Equilibria, and Coarse Correlated Equilibria are useful components for many multiagent machine learning algorithms. Unfortunately, solving a normal-form game could take prohibitive or non-deterministic time to converge, and could fail. We introduce the Neural Equilibrium Solver which utilizes a special equivariant neural network architecture to approximately solve the space of all games of fixed shape, buying speed and determinism. We define a flexible equilibrium selection framework, that is capable of uniquely selecting an equilibrium that minimizes relative entropy, or maximizes welfare. The network is trained without needing to generate any supervised training data. We show remarkable zero-shot generalization to larger games. We argue that such a network is a powerful component for many possible multiagent algorithms. Luke Marris, Ian Gemp, Thomas W. Anthony 0001, Andrea Tacchetti, Siqi Liu 0002, Karl Tuyls |
NeurIPS | 5 |
| 2020 | A Generalized Training Approach for Multiagent Learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland 0001, Karl Tuyls, Julien Pérolat, Siqi Liu 0002, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes 0001, Zhe Wang 0055, Guy Lever, Nicolas Heess, Thore Graepel, Rémi Munos |
ICLR | 6 |
| 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick |
ICLR | 9 |
| 2019 | Emergent Coordination Through Competition
Siqi Liu 0002, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, Thore Graepel |
ICLR (Poster) | 1 |
| 2019 | Hierarchical Visuomotor Control of Humanoids
Josh Merel, Arun Ahuja, Saran Tunyasuvunakool, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Greg Wayne |
ICLR (Poster) | 5 |
| 2017 | Improved Image Captioning via Policy Gradient optimization of SPIDErabstractCurrent image captioning methods are usually trained via maximum likelihood estimation. However, the log-likelihood score of a caption does not correlate well with human assessments of quality. Standard syntactic evaluation metrics, such as BLEU, METEOR and ROUGE, are also not well correlated. The newer SPICE and CIDEr metrics are better correlated, but have traditionally been hard to optimize for. In this paper, we show how to use a policy gradient (PG) method to directly optimize a linear combination of SPICE and CIDEr (a combination we call SPIDEr): the SPICE score ensures our captions are semantically faithful to the image, while CIDEr score ensures our captions are syntactically fluent. The PG method we propose improves on the prior MIXER approach, by using Monte Carlo rollouts instead of mixing MLE training with PG. We show empirically that our algorithm leads to easier optimization and improved results compared to MIXER. Finally, we show that using our PG method we can optimize any of the metrics, including the proposed SPIDEr metric which results in image captions that are strongly preferred by human raters compared to captions generated by the same model but trained to optimize MLE or the COCO metrics. Siqi Liu 0002, Zhenhai Zhu, Sergio Guadarrama, Kevin Murphy 0002 |
ICCV | 1 |