Lucas Baudin

dblp:305/7860 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%
Artificial intelligence
2 papers
Reinforcement learning · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.122022
Smooth Fictitious Play in Stochastic Games with Perturbed Payoffs and Unknown Transitions · NeurIPS 2022
Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic Games · ICML 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games
0.612022
Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic Games · ICML 2022
Algorithmic game theory and mechanism design
equilibrium computation
0.612022
Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic Games · ICML 2022
Algorithmic game theory and mechanism design › learning in games
fictitious play
0.612022
Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic Games · ICML 2022
Algorithmic game theory and mechanism design › equilibrium computation
nash equilibrium computation
0.612022
Smooth Fictitious Play in Stochastic Games with Perturbed Payoffs and Unknown Transitions · NeurIPS 2022
Algorithmic game theory and mechanism design
stochastic games
0.612022
Smooth Fictitious Play in Stochastic Games with Perturbed Payoffs and Unknown Transitions · NeurIPS 2022
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium
0.212022
Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic Games · ICML 2022

Methods — techniques the papers use, named apart from their topics

stochastic approximation · 2.3smooth fictitious play · 1.1regularization · 1.1best response dynamics · 1.1
YearPublicationVenuePosition
2022 Fictitious Play and Best-Response Dynamics in Identical Interest and Zero-Sum Stochastic Games
abstract
This paper proposes an extension of a popular decentralized discrete-time learning procedure when repeating a static game called fictitious play (FP) (Brown, 1951; Robinson, 1951) to a dynamic model called discounted stochastic game (Shapley, 1953). Our family of discrete-time FP procedures is proven to converge to the set of stationary Nash equilibria in identical interest discounted stochastic games. This extends similar convergence results for static games (Monderer & Shapley, 1996a). We then analyze the continuous-time counterpart of our FP procedures, which include as a particular case the best-response dynamic introduced and studied by Leslie et al. (2020) in the context of zero-sum stochastic games. We prove the converge of this dynamics to stationary Nash equilibria in identical-interest and zero-sum discounted stochastic games. Thanks to stochastic approximations, we can infer from the continuous-time convergence some discrete time results such as the convergence to stationary equilibria in zero-sum and team stochastic games (Holler, 2020).
Lucas Baudin, Rida Laraki
ICML1
2022 Smooth Fictitious Play in Stochastic Games with Perturbed Payoffs and Unknown Transitions
abstract
Recent extensions to dynamic games of the well known fictitious play learning procedure in static games were proved to globally converge to stationary Nash equilibria in two important classes of dynamic games (zero-sum and identical-interest discounted stochastic games). However, those decentralized algorithms need the players to know exactly the model (the transition probabilities and their payoffs at every stage). To overcome these strong assumptions, our paper introduces regularizations of the recent algorithms which are moreover, model-free (players don't know the transitions and their payoffs are perturbed at every stage). Our novel procedures can be interpreted as extensions to stochastic games of the classical smooth fictitious play learning procedures in static games (where players best responses are regularized, thanks to a smooth perturbation of their payoff functions). We prove the convergence of our family of procedures to stationary regularized Nash equilibria in the same classes of dynamic games (zero-sum and identical interests discounted stochastic games). The proof uses the continuous smooth best-response dynamics counterparts, and stochastic approximation methods. In the case of a MDP (a one-player stochastic game), our procedures globally converge to the optimal stationary policy of the regularized problem. In that sense, they can be seen as an alternative to the well known Q-learning procedure.
Lucas Baudin, Rida Laraki
NeurIPS1