Yann Ollivier

dblp:63/343 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0007-3967-6808ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Reinforcement learning · 55% Efficient and distributed learning · 8% Generative modeling · 8%
Theoretical computer science
3 papers
Mathematical optimization · 85% Graph algorithms and graph theory · 15%

Topics — the 28 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity
0.812024
Simple Ingredients for Offline Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › imitation learning
few-shot imitation learning
0.812024
Fast Imitation via Behavior Foundation Models · ICLR 2024
Machine learning › Reinforcement learning
imitation learning
0.812024
Fast Imitation via Behavior Foundation Models · ICLR 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Simple Ingredients for Offline Reinforcement Learning · ICML 2024
Machine learning › Deep learning architectures and training
recurrent neural network
0.722018
Can recurrent neural networks warp time? · ICLR (Poster) 2018
Unbiased Online Recurrent Optimization · ICLR (Poster) 2018
Machine learning › Reinforcement learning › generalization in reinforcement learning
zero-shot reinforcement learning
0.712023
Does Zero-Shot Reinforcement Learning Exist? · ICLR 2023
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.712023
Does Zero-Shot Reinforcement Learning Exist? · ICLR 2023
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.512021
Learning One Representation to Optimize All Rewards · NeurIPS 2021
Machine learning › Reinforcement learning › unsupervised reinforcement learning
reward-free reinforcement learning
0.512021
Learning One Representation to Optimize All Rewards · NeurIPS 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412019
First-Order Adversarial Vulnerability of Neural Networks and Input Dimension · ICML 2019
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial vulnerability
0.412019
First-Order Adversarial Vulnerability of Neural Networks and Input Dimension · ICML 2019
Machine learning › Reinforcement learning › value-based reinforcement learning › q-learning
continuous-time q-learning
0.412019
Making Deep Q-learning methods robust to time discretization · ICML 2019
Machine learning › Reinforcement learning
deep reinforcement learning
0.412019
Making Deep Q-learning methods robust to time discretization · ICML 2019
Machine learning › Learning theory
loss function
0.412019
White-box vs Black-box: Bayes Optimal Strategies for Membership Inference · ICML 2019
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.412019
Making Deep Q-learning methods robust to time discretization · ICML 2019
Machine learning › Reinforcement learning
robust reinforcement learning
0.412019
Making Deep Q-learning methods robust to time discretization · ICML 2019
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction
0.412019
Separable value functions across time-scales · ICML 2019
Machine learning › Reinforcement learning
value function
0.412019
Separable value functions across time-scales · ICML 2019
Security and privacy of machine learning
membership inference
0.412019
White-box vs Black-box: Bayes Optimal Strategies for Membership Inference · ICML 2019
Machine learning › Generative modeling › generative adversarial network › GAN architecture
discriminator design
0.312018
Mixed batches and symmetric discriminators for GAN training · ICML 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
Mixed batches and symmetric discriminators for GAN training · ICML 2018
Machine learning › Learning theory › model selection
minimum description length
0.312018
The Description Length of Deep Learning models · NeurIPS 2018
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation
0.312018
Mixed batches and symmetric discriminators for GAN training · ICML 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
The Description Length of Deep Learning models · NeurIPS 2018
Machine learning › Optimization for machine learning
online optimization
0.312018
Unbiased Online Recurrent Optimization · ICLR (Poster) 2018
Machine learning › Time series and sequential data
time warping
0.312018
Can recurrent neural networks warp time? · ICLR (Poster) 2018
Machine learning › Optimization for machine learning › evolutionary computation
evolution strategies
0.312017
Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles · J. Mach. Learn. Res. 2017
Mathematical optimization
black-box optimization
0.312017
Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles · J. Mach. Learn. Res. 2017

Methods — techniques the papers use, named apart from their topics

reward-based imitation · 0.8feature matching · 0.8behavioral cloning · 0.8bayes optimal inference · 0.8IQL · 0.8AWAC · 0.8reinforcement learning · 0.7temporal difference learning · 0.5forward-backward representation · 0.5deep learning · 0.5regularization · 0.4gradient norm analysis · 0.4natural gradient · 0.3information geometry · 0.3evolution strategies · 0.3green measures · 0.1
YearPublicationVenuePosition
2024 Fast Imitation via Behavior Foundation Models
abstract
Imitation learning (IL) aims at producing agents that can imitate any behavior given a few expert demonstrations. Yet existing approaches require many demonstrations and/or running (online or offline) reinforcement learning (RL) algorithms for each new imitation task. Here we show that recent RL foundation models based on successor measures can imitate any expert behavior almost instantly with just a few demonstrations and no need for RL or fine-tuning, while accommodating several IL principles (behavioral cloning, feature matching, reward-based, and goal-based reductions). In our experiments, imitation via RL foundation models matches, and often surpasses, the performance of SOTA offline IL algorithms, and produces imitation policies from new demonstrations within seconds instead of hours.
Matteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric, Yann Ollivier
ICLR5
2024 Simple Ingredients for Offline Reinforcement Learning
abstract
Offline reinforcement learning algorithms have proven effective on datasets highly connected to the target downstream task. Yet, by leveraging a novel testbed (MOOD) in which trajectories come from heterogeneous sources, we show that existing methods struggle with diverse data: their performance considerably deteriorates as data collected for related but different tasks is simply added to the offline buffer. In light of this finding, we conduct a large empirical study where we formulate and test several hypotheses to explain this failure. Surprisingly, we find that targeted scale, more than algorithmic considerations, is the key factor influencing performance. We show that simple methods like AWAC and IQL with increased policy size overcome the paradoxical failure modes from the inclusion of additional data in MOOD, and notably outperform prior state-of-the-art algorithms on the canonical D4RL benchmark.
Edoardo Cetin, Andrea Tirinzoni, Matteo Pirotta, Alessandro Lazaric, Yann Ollivier, Ahmed Touati
ICML5
2023 Does Zero-Shot Reinforcement Learning Exist?
Ahmed Touati, Jérémy Rapin, Yann Ollivier
ICLR3
2021 Learning One Representation to Optimize All Rewards
abstract
We introduce the forward-backward (FB) representation of the dynamics of a reward-free Markov decision process. It provides explicit near-optimal policies for any reward specified a posteriori. During an unsupervised phase, we use reward-free interactions with the environment to learn two representations via off-the-shelf deep learning methods and temporal difference (TD) learning. In the test phase, a reward representation is estimated either from reward observations or an explicit reward description (e.g., a target state). The optimal policy for thatreward is directly obtained from these representations, with no planning. We assume access to an exploration scheme or replay buffer for the first phase.The corresponding unsupervised loss is well-principled: if training is perfect, the policies obtained are provably optimal for any reward function. With imperfect training, the sub-optimality is proportional to the unsupervised approximation error. The FB representation learns long-range relationships between states and actions, via a predictive occupancy map, without having to synthesize states as in model-based approaches.This is a step towards learning controllable agents in arbitrary black-box stochastic environments. This approach compares well to goal-oriented RL algorithms on discrete and continuous mazes, pixel-based MsPacman, and the FetchReach virtual robot arm. We also illustrate how the agent can immediately adapt to new tasks beyond goal-oriented RL.
Ahmed Touati, Yann Ollivier
NeurIPS2
2019 Separable value functions across time-scales
Joshua Romoff, Peter Henderson 0002, Ahmed Touati, Yann Ollivier, Joelle Pineau, Emma Brunskill
ICML4
2019 White-box vs Black-box: Bayes Optimal Strategies for Membership Inference
abstract
Membership inference determines, given a sample and trained parameters of a machine learning model, whether the sample was part of the training set. In this paper, we derive the optimal strategy for membership inference with a few assumptions on the distribution of the parameters. We show that optimal attacks only depend on the loss function, and thus black-box attacks are as good as white-box attacks. As the optimal strategy is not tractable, we provide approximations of it leading to several inference methods, and show that existing membership inference methods are coarser approximations of this optimal strategy. Our membership attacks outperform the state of the art in various settings, ranging from a simple logistic regression to more complex architectures and datasets, such as ResNet-101 and Imagenet.
Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, Yann Ollivier, Hervé Jégou
ICML4
2019 First-Order Adversarial Vulnerability of Neural Networks and Input Dimension
abstract
Over the past few years, neural networks were proven vulnerable to adversarial images: targeted but imperceptible image perturbations lead to drastically different predictions. We show that adversarial vulnerability increases with the gradients of the training objective when viewed as a function of the inputs. Surprisingly, vulnerability does not depend on network topology: for many standard network architectures, we prove that at initialization, the L1-norm of these gradients grows as the square root of the input dimension, leaving the networks increasingly vulnerable with growing image size. We empirically show that this dimension-dependence persists after either usual or robust training, but gets attenuated with higher regularization.
Carl-Johann Simon-Gabriel, Yann Ollivier, Léon Bottou, Bernhard Schölkopf, David Lopez-Paz
ICML2
2019 Making Deep Q-learning methods robust to time discretization
abstract
Despite remarkable successes, Deep Reinforce- ment Learning (DRL) is not robust to hyperparam- eterization, implementation details, or small envi- ronment changes (Henderson et al. 2017, Zhang et al. 2018). Overcoming such sensitivity is key to making DRL applicable to real world problems. In this paper, we identify sensitivity to time dis- cretization in near continuous-time environments as a critical factor; this covers, e.g., changing the number of frames per second, or the action frequency of the controller. Empirically, we find that Q-learning-based approaches such as Deep Q- learning (Mnih et al., 2015) and Deep Determinis- tic Policy Gradient (Lillicrap et al., 2015) collapse with small time steps. Formally, we prove that Q-learning does not exist in continuous time. We detail a principled way to build an off-policy RL algorithm that yields similar performances over a wide range of time discretizations, and confirm this robustness empirically.
Corentin Tallec, Léonard Blier, Yann Ollivier
ICML3
2019 Learning with Random Learning Rates
Léonard Blier, Pierre Wolinski, Yann Ollivier
ECML/PKDD (2)3
2018 Unbiased Online Recurrent Optimization
Corentin Tallec, Yann Ollivier
ICLR (Poster)2
2018 Can recurrent neural networks warp time?
Corentin Tallec, Yann Ollivier
ICLR (Poster)2
2018 Mixed batches and symmetric discriminators for GAN training
abstract
Generative adversarial networks (GANs) are pow- erful generative models based on providing feed- back to a generative network via a discriminator network. However, the discriminator usually as- sesses individual samples. This prevents the dis- criminator from accessing global distributional statistics of generated samples, and often leads to mode dropping: the generator models only part of the target distribution. We propose to feed the discriminator with mixed batches of true and fake samples, and train it to predict the ratio of true samples in the batch. The latter score does not depend on the order of samples in a batch. Rather than learning this invariance, we introduce a generic permutation-invariant discriminator ar- chitecture. This architecture is provably a uni- versal approximator of all symmetric functions. Experimentally, our approach reduces mode col- lapse in GANs on two synthetic datasets, and obtains good results on the CIFAR10 and CelebA datasets, both qualitatively and quantitatively.
Thomas Lucas 0002, Corentin Tallec, Yann Ollivier, Jakob Verbeek
ICML3
2018 The Description Length of Deep Learning models
abstract
Deep learning models often have more parameters than observations, and still perform well. This is sometimes described as a paradox. In this work, we show experimentally that despite their huge number of parameters, deep neural networks can compress the data losslessly even when taking the cost of encoding the parameters into account. Such a compression viewpoint originally motivated the use of variational methods in neural networks. However, we show that these variational methods provide surprisingly poor compression bounds, despite being explicitly built to minimize such bounds. This might explain the relatively poor practical performance of variational methods in deep learning. Better encoding methods, imported from the Minimum Description Length (MDL) toolbox, yield much better compression values on deep networks.
Léonard Blier, Yann Ollivier
NeurIPS2
2017 Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles
abstract
We present a canonical way to turn any smooth parametric family of probability distributions on an arbitrary search space $X$ into a continuous-time black-box optimization method on $X$, the information-geometric optimization (IGO) method. Invariance as a major design principle keeps the number of arbitrary choices to a minimum. The resulting IGO flow is the flow of an ordinary differential equation conducting the natural gradient ascent of an adaptive, time-dependent transformation of the objective function. It makes no particular assumptions on the objective function to be optimized. The IGO method produces explicit IGO algorithms through time discretization. It naturally recovers versions of known algorithms and offers a systematic way to derive new ones. In continuous search spaces, IGO algorithms take a form related to natural evolution strategies (NES). The cross-entropy method is recovered in a particular case with a large time step, and can be extended into a smoothed, parametrization-independent maximum likelihood update (IGO-ML). When applied to the family of Gaussian distributions on $\R^d$, the IGO framework recovers a version of the well-known CMA-ES algorithm and of xNES. For the family of Bernoulli distributions on $\{0,1\}^d$, we recover the seminal PBIL algorithm and cGA. For the distributions of restricted Boltzmann machines, we naturally obtain a novel algorithm for discrete optimization on $\{0,1\}^d$. All these algorithms are natural instances of, and unified under, the single information-geometric optimization framework. The IGO method achieves, thanks to its intrinsic formulation, maximal invariance properties: invariance under reparametrization of the search space $X$, under a change of parameters of the probability distribution, and under increasing transformation of the function to be optimized. The latter is achieved through an adaptive, quantile-based formulation of the objective. Theoretical considerations strongly suggest that IGO algorithms are essentially characterized by a minimal change of the distribution over time. Therefore they have minimal loss in diversity through the course of optimization, provided the initial diversity is high. First experiments using restricted Boltzmann machines confirm this insight. As a simple consequence, IGO seems to provide, from information theory, an elegant way to simultaneously explore several valleys of a fitness landscape in a single run.
Yann Ollivier, Ludovic Arnold, Anne Auger, Nikolaus Hansen
J. Mach. Learn. Res.1
2013 Objective improvement in information-geometric optimization
abstract
Information-Geometric Optimization (IGO) is a unified framework of stochastic algorithms for optimization problems. Given a family of probability distributions, IGO turns the original optimization problem into a new maximization problem on the parameter space of the probability distributions. IGO updates the parameter of the probability distribution along the natural gradient, taken with respect to the Fisher metric on the parameter manifold, aiming at maximizing an adaptive transform of the objective function. IGO recovers several known algorithms as particular instances: for the family of Bernoulli distributions IGO recovers PBIL, for the family of Gaussian distributions the pure rank-μ CMA-ES update is recovered, and for exponential families in expectation parametrization the cross-entropy/ML method is recovered.
Youhei Akimoto, Yann Ollivier
FOGA2
2012 A Curved Brunn-Minkowski Inequality on the Discrete Hypercube, Or: What Is the Ricci Curvature of the Discrete Hypercube?
abstract
We compare two approaches to Ricci curvature on nonsmooth spaces in the case of the discrete hypercube $\{0,1\}^N$. While the coarse Ricci curvature of the first author readily yields a positive value for curvature, the displacement convexity property of Lott, Sturm, and Villani could not be fully implemented. Yet along the way we get new results of a combinatorial and probabilistic nature, including a curved Brunn--Minkowski inequality on the discrete hypercube.
Yann Ollivier, Cédric Villani
SIAM J. Discret. Math.1
2007 Finding Related Pages Using Green Measures: An Illustration with Wikipedia
Yann Ollivier, Pierre Senellart
AAAI1