EDBT 2026 Demo / reviewers in the wild / expert
Adrian Müller 0002
dblp:25/4454-2
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-2210-6845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 75% Optimization for machine learning · 25% | |
| Theoretical computer science
1 paper |
Approximation and online algorithms · 50% Algorithmic game theory and mechanism design · 50% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Approximation and online algorithms
online learning |
0.9 | 1 | 2025 | Best of Both Worlds: Regret Minimization versus Minimax Play · ICML 2025 |
Algorithmic game theory and mechanism design
zero-sum game |
0.9 | 1 | 2025 | Best of Both Worlds: Regret Minimization versus Minimax Play · ICML 2025 |
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.8 | 1 | 2024 | Truly No-Regret Learning in Constrained MDPs · ICML 2024 |
Machine learning › Reinforcement learning › regret minimization
no-regret learning |
0.8 | 1 | 2024 | Truly No-Regret Learning in Constrained MDPs · ICML 2024 |
Machine learning › Optimization for machine learning
primal-dual methods |
0.8 | 1 | 2024 | Truly No-Regret Learning in Constrained MDPs · ICML 2024 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.8 | 1 | 2024 | Truly No-Regret Learning in Constrained MDPs · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
regret analysis · 0.9bandit feedback · 0.9regularization · 0.8primal-dual method · 0.8last-iterate convergence · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Best of Both Worlds: Regret Minimization versus Minimax PlayabstractIn this paper, we investigate the existence of online learning algorithms with bandit feedback that simultaneously guarantee $O(1)$ regret compared to a given comparator strategy, and $\tilde{O}(\sqrt{T})$ regret compared to any fixed strategy, where $T$ is the number of rounds. We provide the first affirmative answer to this question whenever the comparator strategy supports every action. In the context of zero-sum games with min-max value zero, both in normal- and extensive form, we show that our results allow us to guarantee to risk at most $O(1)$ loss while being able to gain $\Omega(T)$ from exploitable opponents, thereby combining the benefits of both no-regret algorithms and minimax play. Adrian Müller 0002, Jon Schneider, Stratis Skoulakis, Luca Viano, Volkan Cevher |
ICML | 1 |
| 2024 | Truly No-Regret Learning in Constrained MDPsabstractConstrained Markov decision processes (CMDPs) are a common way to model safety constraints in reinforcement learning. State-of-the-art methods for efficiently solving CMDPs are based on primal-dual algorithms. For these algorithms, all currently known regret bounds allow for error cancellations — one can compensate for a constraint violation in one round with a strict constraint satisfaction in another. This makes the online learning process unsafe since it only guarantees safety for the final (mixture) policy but not during learning. As Efroni et al. (2020) pointed out, it is an open question whether primal-dual algorithms can provably achieve sublinear regret if we do not allow error cancellations. In this paper, we give the first affirmative answer. We first generalize a result on last-iterate convergence of regularized primal-dual schemes to CMDPs with multiple constraints. Building upon this insight, we propose a model-based primal-dual algorithm to learn in an unknown CMDP. We prove that our algorithm achieves sublinear regret without error cancellations. Adrian Müller 0002, Pragnya Alatur, Volkan Cevher, Giorgia Ramponi, Niao He |
ICML | 1 |