Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Adrian Müller 0002

dblp:25/4454-2 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-2210-6845ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 75% Optimization for machine learning · 25%
Theoretical computer science
1 paper
Approximation and online algorithms · 50% Algorithmic game theory and mechanism design · 50%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Approximation and online algorithms
online learning
0.912025
Best of Both Worlds: Regret Minimization versus Minimax Play · ICML 2025
Algorithmic game theory and mechanism design
zero-sum game
0.912025
Best of Both Worlds: Regret Minimization versus Minimax Play · ICML 2025
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.812024
Truly No-Regret Learning in Constrained MDPs · ICML 2024
Machine learning › Reinforcement learning › regret minimization
no-regret learning
0.812024
Truly No-Regret Learning in Constrained MDPs · ICML 2024
Machine learning › Optimization for machine learning
primal-dual methods
0.812024
Truly No-Regret Learning in Constrained MDPs · ICML 2024
Machine learning › Reinforcement learning
safe reinforcement learning
0.812024
Truly No-Regret Learning in Constrained MDPs · ICML 2024

Methods — techniques the papers use, named apart from their topics

regret analysis · 0.9bandit feedback · 0.9regularization · 0.8primal-dual method · 0.8last-iterate convergence · 0.8
YearPublicationVenuePosition
2025 Best of Both Worlds: Regret Minimization versus Minimax Play
abstract
In this paper, we investigate the existence of online learning algorithms with bandit feedback that simultaneously guarantee $O(1)$ regret compared to a given comparator strategy, and $\tilde{O}(\sqrt{T})$ regret compared to any fixed strategy, where $T$ is the number of rounds. We provide the first affirmative answer to this question whenever the comparator strategy supports every action. In the context of zero-sum games with min-max value zero, both in normal- and extensive form, we show that our results allow us to guarantee to risk at most $O(1)$ loss while being able to gain $\Omega(T)$ from exploitable opponents, thereby combining the benefits of both no-regret algorithms and minimax play.
Adrian Müller 0002, Jon Schneider, Stratis Skoulakis, Luca Viano, Volkan Cevher
ICML1
2024 Truly No-Regret Learning in Constrained MDPs
abstract
Constrained Markov decision processes (CMDPs) are a common way to model safety constraints in reinforcement learning. State-of-the-art methods for efficiently solving CMDPs are based on primal-dual algorithms. For these algorithms, all currently known regret bounds allow for error cancellations — one can compensate for a constraint violation in one round with a strict constraint satisfaction in another. This makes the online learning process unsafe since it only guarantees safety for the final (mixture) policy but not during learning. As Efroni et al. (2020) pointed out, it is an open question whether primal-dual algorithms can provably achieve sublinear regret if we do not allow error cancellations. In this paper, we give the first affirmative answer. We first generalize a result on last-iterate convergence of regularized primal-dual schemes to CMDPs with multiple constraints. Building upon this insight, we propose a model-based primal-dual algorithm to learn in an unknown CMDP. We prove that our algorithm achieves sublinear regret without error cancellations.
Adrian Müller 0002, Pragnya Alatur, Volkan Cevher, Giorgia Ramponi, Niao He
ICML1