VLDB 2026 Research / reviewers in the wild / expert
Stefanos Leonardos
dblp:192/1237
· DBLP profile ↗
16ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0002-1498-1490ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Asymptotic Extinction in Large Coordination GamesabstractWe study the exploration-exploitation trade-off for large multiplayer coordination games where players strategise via Q-Learning, a common learning framework in multi-agent reinforcement learning. Q-Learning is known to have two shortcomings, namely non-convergence and potential equilibrium selection problems, when there are multiple fixed points, called Quantal Response Equilibria (QRE). Furthermore, whilst QRE have full support for finite games, it is not clear how Q-Learning behaves as the game becomes large. In this paper, we characterise the critical exploration rate that guarantees convergence to a unique fixed point, addressing the two shortcomings above. Using a generating-functional method, we show that this rate increases with the number of players and the alignment of their payoffs. For many-player coordination games with perfectly aligned payoffs, this exploration rate is roughly twice that of p-player zero-sum games. As for large games, we provide a structural result for QRE, which suggests that as the game size increases, Q-Learning converges to a QRE near the boundary of the simplex of the action space, a phenomenon we term asymptotic extinction, where a constant fraction of the actions are played with zero probability at a rate o(1/N) for an N -action game. Desmond Chan, Bart de Keijzer, Tobias Galla, Stefanos Leonardos, Carmine Ventre |
AAAI | 4 |
| 2025 | From Competition to Centralization: The Oligopoly in Ethereum Block Building AuctionsabstractBlock production on the Ethereum blockchain has adopted an auction-based mechanism known as Proposer-Builder Separation (PBS), where validators outsource block creation to builders competing in MEV-Boost auctions for Maximal Extractable Value (MEV) rewards. We employ empirical game-theoretic analysis based on simulations to examine how advantages in latency and MEV access shape builder strategic bidding and auction outcomes. We find that a small set of dominant builders leverage these advantages, consolidating power, reducing auction efficiency, and heightening centralization. Our results underscore the need for fair MEV distribution and sustained efforts to promote decentralization in Ethereum’s block building market. Fei Wu 0030, Thomas Thiery, Stefanos Leonardos, Carmine Ventre |
ECAI | 3 |
| 2025 | Bitcoin's Edge: Embedded Sentiment in Blockchain Transactional Data
Charalampos Kleitsikas, Nikolaos Korfiatis, Stefanos Leonardos, Carmine Ventre |
ICBC | 3 |
| 2024 | Experience Report of the AWS+KCL Impact Accelerator for Public Sector EngagementabstractThis industry experience report chronicles the experience of developing an impact-focused group project module within a computer science master's programme at King's College London over two years. The module was set up in collaboration with Amazon Web Services to match student teams with public sector challenges requiring innovative technological solutions. An iterative process of modifications based on partner and student feedback aimed to enhance the learning experience and outcomes. Key benefits included providing authentic professional development for students, enabling innovation and entrepreneurship, building partnerships between academia and the public sector, and embedding responsible innovation into projects. However, challenges emerged around managing expectations, ensuring consistent partner engagement, providing support for spin-outs, and handling sensitive data issues. As more projects involved artificial intelligence applications in the second year, developing mechanisms to ethically provide access while protecting sensitive information was an increasingly crucial need. Moreover, understanding the value and impact of this model of software engineering project module requires additional research support. Overall, this collaborative module offers a promising model to deliver impact-driven solutions through coordinating academia, industry, and public sector partners. Further research can help optimise such partnerships for societal impact. Caitlin M. Bentley, Elena Simperl, Mike Bainbridge, Daisy Ogden, Stefanos Leonardos, Gunel Jahangirova, Joanna Walker, Christopher Hampson |
CSEE&T | 5 |
| 2024 | Strategic Bidding Wars in On-chain AuctionsabstractThe Ethereum block-building process has changed significantly since the emergence of Proposer-Builder Separation. Validators access blocks through a marketplace, where block builders bid for the right to construct the block and earn MEV (Maximal Extractable Value) rewards in an on-chain competition, known as the MEV-boost auction. While more than 90% of blocks are currently built via MEV-Boost, tradeoffs between builders’ strategic behaviors and auction design remain poorly understood. In this paper we address this gap. We introduce a game-theoretic model for MEV-Boost auctions and use simulations to study different builders’ bidding strategies observed in practice. We study various strategic interactions and auction setups and evaluate how the interplay between critical elements such as access to MEV opportunities and improved connectivity to relays impact bidding performance. Our results demonstrate the importance of latency on the effectiveness of builders’ strategies and the overall auction outcome from the proposer’s perspective. Fei Wu 0030, Thomas Thiery, Stefanos Leonardos, Carmine Ventre |
ICBC | 3 |
| 2024 | Beating Price of Anarchy and Gradient Descent without Regret in Potential GamesabstractArguably one of the thorniest problems in game theory is that of equilibrium selection. Specifically, in the presence of multiple equilibria do self-interested learning dynamics typically select the socially optimal ones? We study a rich class of continuous-time no-regret dynamics in potential games (PGs). Our class of dynamics, *Q-Replicator Dynamics* (QRD), include gradient descent (GD), log-barrier and replicator dynamics (RD) as special cases. We start by establishing *pointwise convergence* of all QRD to Nash equilibria in almost all PGs. In the case of GD, we show a tight average case performance within a factor of two of optimal, for a class of symmetric $2\times2$ potential games with unbounded Price of Anarchy (PoA). Despite this positive result, we show that GD is not always the optimal choice even in this restricted setting. Specifically, GD outperforms RD, if and only if *risk-* and *payoff-dominance* equilibria coincide. Finally, we experimentally show how these insights extend to all QRD dynamics and that unbounded gaps between average case performance and PoA analysis are common even in larger settings. Iosif Sakos, Stefanos Leonardos, Stelios Stavroulakis, Will Overman, Ioannis Panageas, Georgios Piliouras |
ICLR | 2 |
| 2024 | Introduction to the Special Issue on Mathematical Research for Blockchain EconomyabstractIntroduction to the Special Issue on Mathematical Research for Blockchain EconomyBlockchain Technology has been considered as the most revolutionizing invention since the Internet.Because of its immutable nature and the associated security and privacy benefits, it has widely attracted the attention of banks, governments, techno-corporations and venture investors.Blockchain applications range from finance to healthcare, from education and media to logistics, NFTs and many more.However, the theoretical limitations and technical barriers to the adoption of blockchain such as scalability, latency, privacy and security need to be further studied and addressed in high-quality research.This special issue of the ACM Distributed Ledger Technologies: Research and Practice (ACM DLT) journals contains selected and refereed papers on the topic of Mathematical Research in Blockchain Economies.Preliminary versions of some of the papers appeared in the 2022 edition of the International Conference on Mathematical Research for Blockchain Economy (MARBLE'22), which took place in Vilamoura, Portugal, from July 12 to 24, 2022.Following the paradigm of the conference, the current special issue provides a high-profile, cutting-edge platform for mathematicians, computer scientists and economists, from both industry and practice, to present the latest advances and innovations in key theories of blockchain.Having a broad international appeal, both the MARBLE conference and the current special issue focuses on the mathematics behind blockchain to bridge the gap between theory and practice.The three selected article in this special issue were selected from 10 submitted manuscripts, following the standard, rigorous ACM DLT review procedures.The articles cover topics in decentralized finance, smart contracts and game-theoretic modelling of blockchains.The content of the articles is as follows. Stefanos Leonardos, William J. Knottenbelt, Elise Alfieri, Panos M. Pardalos, Ilias S. Kotsireas |
Distributed Ledger Technol. Res. Pract. | 1 |
| 2023 | Optimality Despite Chaos in Fee Markets
Stefanos Leonardos, Daniël Reijsbergen, Barnabé Monnot, Georgios Piliouras |
FC | 1 |
| 2023 | AlberDICE: Addressing Out-Of-Distribution Joint Actions in Offline Multi-Agent RL via Alternating Stationary Distribution Correction EstimationabstractOne of the main challenges in offline Reinforcement Learning (RL) is the distribution shift that arises from the learned policy deviating from the data collection policy. This is often addressed by avoiding out-of-distribution (OOD) actions during policy improvement as their presence can lead to substantial performance degradation. This challenge is amplified in the offline Multi-Agent RL (MARL) setting since the joint action space grows exponentially with the number of agents.
To avoid this curse of dimensionality, existing MARL methods adopt either value decomposition methods or fully decentralized training of individual agents. However, even when combined with standard conservatism principles, these methods can still result in the selection of OOD joint actions in offline MARL. To this end, we introduce AlberDICE,
an offline MARL algorithm that alternatively performs centralized training of individual agents based on stationary distribution optimization. AlberDICE circumvents the exponential complexity of MARL by computing the best response of one agent at a time while effectively avoiding OOD joint action selection. Theoretically, we show that the alternating optimization procedure converges to Nash policies. In the experiments, we demonstrate that AlberDICE significantly outperforms baseline algorithms on a standard suite of MARL benchmarks. Daiki E. Matsunaga, Jongmin Lee 0004, Jaeseok Yoon, Stefanos Leonardos, Pieter Abbeel, Kee-Eung Kim |
NeurIPS | 4 |
| 2022 | Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games
Stefanos Leonardos, Will Overman, Ioannis Panageas, Georgios Piliouras |
ICLR | 1 |
| 2022 | Exploration-exploitation in multi-agent learning: Catastrophe theory meets game theory
Stefanos Leonardos, Georgios Piliouras |
Artif. Intell. | 1 |
| 2021 | Exploration-Exploitation in Multi-Agent Learning: Catastrophe Theory Meets Game TheoryabstractExploration-exploitation is a powerful and practical tool in multi-agent learning (MAL), however, its effects are far from understood. To make progress in this direction, we study a smooth analogue of Q-learning. We start by showing that our learning model has strong theoretical justification as an optimal model for studying exploration-exploitation. Specifically, we prove that smooth Q-learning has bounded regret in arbitrary games for a cost model that explicitly captures the balance between game and exploration costs and that it always converges to the set of quantal-response equilibria (QRE), the standard solution concept for games under bounded rationality, in weighted potential games with heterogeneous learning agents. In our main task, we then turn to measure the effect of exploration in collective system performance. We characterize the geometry of the QRE surface in low-dimensional MAL systems and link our findings with catastrophe (bifurcation) theory. In particular, as the exploration hyperparameter evolves over-time, the system undergoes phase transitions where the number and stability of equilibria can change radically given an infinitesimal change to the exploration parameter. Based on this, we provide a formal theoretical treatment of how tuning the exploration parameter can provably lead to equilibrium selection with both positive as well as negative (and potentially unbounded) effects to system performance. Stefanos Leonardos, Georgios Piliouras |
AAAI | 1 |
| 2021 | Dynamical analysis of the EIP-1559 Ethereum fee marketabstractParticipation in permissionless blockchains results in competition over system resources, which needs to be controlled with fees. Until recently, Ethereum's fee mechanism was implemented via a first-price auction that resulted in unpredictable fees as well as other inefficiencies. Launched on August 5, 2021, EIP-1559 is an improved proposal that introduces a number of innovative features such as a dynamically adaptive basefee that is burnt, instead of being paid to the miners. Despite intense interest in understanding its properties, several basic questions such as whether and under what conditions does this protocol self-stabilize have remained elusive thus far. Stefanos Leonardos, Barnabé Monnot, Daniël Reijsbergen, Stratis Skoulakis, Georgios Piliouras |
AFT | 1 |
| 2021 | Learning in Markets: Greed Leads to Chaos but Following the Price is RightabstractWe study learning dynamics in distributed production economies such as blockchain mining, peer-to-peer file sharing and crowdsourcing. These economies can be modelled as multi-product Cournot competitions or all-pay auctions (Tullock contests) when individual firms have market power, or as Fisher markets with quasi-linear utilities when every firm has negligible influence on market outcomes. In the former case, we provide a formal proof that Gradient Ascent (GA) can be Li-Yorke chaotic for a step size as small as Θ(1/n), where n is the number of firms. In stark contrast, for the Fisher market case, we derive a Proportional Response (PR) protocol that converges to market equilibrium. The positive results on the convergence of the PR dynamics are obtained in full generality, in the sense that they hold for Fisher markets with any quasi-linear utility functions. Conversely, the chaos results for the GA dynamics are established even in the simplest possible setting of two firms and one good, and they hold for a wide range of price functions with different demand elasticities. Our findings suggest that by considering multi-agent interactions from a market rather than a game-theoretic perspective, we can formally derive natural learning protocols which are stable and converge to effective outcomes rather than being chaotic. Yun Kuen Cheung, Stefanos Leonardos, Georgios Piliouras |
IJCAI | 2 |
| 2021 | Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded RationalityabstractThe interplay between exploration and exploitation in competitive multi-agent learning is still far from being well understood. Motivated by this, we study smooth Q-learning, a prototypical learning model that explicitly captures the balance between game rewards and exploration costs. We show that Q-learning always converges to the unique quantal-response equilibrium (QRE), the standard solution concept for games under bounded rationality, in weighted zero-sum polymatrix games with heterogeneous learning agents using positive exploration rates. Complementing recent results about convergence in weighted potential games [16,34], we show that fast convergence of Q-learning in competitive settings obtains regardless of the number of agents and without any need for parameter fine-tuning. As showcased by our experiments in network zero-sum games, these theoretical results provide the necessary guarantees for an algorithmic approach to the currently open problem of equilibrium selection in competitive multi-agent settings. Stefanos Leonardos, Georgios Piliouras, Kelly Spendlove |
NeurIPS | 1 |
| 2020 | Catastrophe by Design in Population Games: Destabilizing WastefulLocked-In Technologies
Stefanos Leonardos, Iosif Sakos, Costas Courcoubetis, Georgios Piliouras |
WINE | 1 |