Jean-Baptiste Lespiau

dblp:230/4171 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 31% Reinforcement learning · 31% Question answering and dialogue systems · 11%
Theoretical computer science
3 papers
Algorithmic game theory and mechanism design · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
equilibrium computation
0.922021
From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization · ICML 2021
Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent · IJCAI 2019
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.822020
Fast computation of Nash Equilibria in Imperfect Information Games · ICML 2020
Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent · IJCAI 2019
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Information retrieval
document retrieval
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search
0.512021
Machine Translation Decoding beyond Beam Search · EMNLP (1) 2021
Natural language and speech › Language models and text generation
decoding
0.512021
Machine Translation Decoding beyond Beam Search · EMNLP (1) 2021
Machine learning › Learning theory › online learning › no-regret algorithms
follow-the-regularized-leader
0.512021
From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization · ICML 2021
Algorithmic game theory and mechanism design
imperfect information games
0.512021
From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization · ICML 2021
Knowledge, reasoning and agents › Multi-agent systems
imperfect information games
0.412020
Fast computation of Nash Equilibria in Imperfect Information Games · ICML 2020
Algorithmic game theory and mechanism design › equilibrium computation
nash equilibrium computation
0.412020
Fast computation of Nash Equilibria in Imperfect Information Games · ICML 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.412019
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search · ICLR (Poster) 2019
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412019
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search · ICLR (Poster) 2019
Machine learning › Reinforcement learning
policy search
0.412019
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search · ICLR (Poster) 2019
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts › nash equilibrium
approximate nash equilibrium
0.412019
Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent · IJCAI 2019

Methods — techniques the papers use, named apart from their topics

differentiable encoder · 1.1chunked cross-attention · 1.1regularization · 1.0poincaré recurrence · 1.0policy gradient · 0.9mirror ascent · 0.9best response · 0.9exploitability descent · 0.8reinforcement learning · 0.5beam search · 0.5policy optimization · 0.4function approximation · 0.4
YearPublicationVenuePosition
2022 Improving Language Models by Retrieving from Trillions of Tokens
abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre
ICML8
2021 Machine Translation Decoding beyond Beam Search
abstract
Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Lespiau Jean-Baptiste, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Jean-Baptiste Lespiau, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals
EMNLP (1)5
2021 From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization
abstract
In this paper we investigate the Follow the Regularized Leader dynamics in sequential imperfect information games (IIG). We generalize existing results of Poincar{é} recurrence from normal-form games to zero-sum two-player imperfect information games and other sequential game settings. We then investigate how adapting the reward (by adding a regularization term) of the game can give strong convergence guarantees in monotone games. We continue by showing how this reward adaptation technique can be leveraged to build algorithms that converge exactly to the Nash equilibrium. Finally, we show how these insights can be directly used to build state-of-the-art model-free algorithms for zero-sum two-player Imperfect Information Games (IIG).
Julien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland 0001, Pedro A. Ortega, Neil Burch, Thomas W. Anthony 0001, David Balduzzi, Bart De Vylder, Georgios Piliouras, Marc Lanctot, Karl Tuyls
ICML3
2020 Fast computation of Nash Equilibria in Imperfect Information Games
abstract
We introduce and analyze a class of algorithms, called Mirror Ascent against an Improved Opponent (MAIO), for computing Nash equilibria in two-player zero-sum games, both in normal form and in sequential form with imperfect information. These algorithms update the policy of each player with a mirror-ascent step to maximize the value of playing against an improved opponent. An improved opponent can be a best response, a greedy policy, a policy improved by policy gradient, or by any other reinforcement learning or search techniques. We establish a convergence result of the last iterate to the set of Nash equilibria and show that the speed of convergence depends on the amount of improvement offered by these improved policies. In addition, we show that under some condition, if we use a best response as improved policy, then an exponential convergence rate is achieved.
Rémi Munos, Julien Pérolat, Jean-Baptiste Lespiau, Mark Rowland 0001, Bart De Vylder, Marc Lanctot, Finbarr Timbers, Daniel Hennes, Shayegan Omidshafiei, Audrunas Gruslys, Mohammad Gheshlaghi Azar, Edward Lockhart, Karl Tuyls
ICML3
2019 Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
Lars Buesing, Theophane Weber, Yori Zwols, Nicolas Heess, Sébastien Racanière, Arthur Guez, Jean-Baptiste Lespiau
ICLR (Poster)7
2019 Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent
abstract
In this paper, we present exploitability descent, a new algorithm to compute approximate equilibria in two-player zero-sum extensive-form games with imperfect information, by direct policy optimization against worst-case opponents. We prove that when following this optimization, the exploitability of a player's strategy converges asymptotically to zero, and hence when both players employ this optimization, the joint policies converge to a Nash equilibrium. Unlike fictitious play (XFP) and counterfactual regret minimization (CFR), our convergence result pertains to the policies being optimized rather than the average policies. Our experiments demonstrate convergence rates comparable to XFP and CFR in four benchmark games in the tabular case. Using function approximation, we find that our algorithm outperforms the tabular version in two of the games, which, to the best of our knowledge, is the first such result in imperfect information games among this class of algorithms.
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, Karl Tuyls
IJCAI4