Sherjil Ozair

dblp:139/0736 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 38% Representation and self-supervised learning · 28% Probabilistic and Bayesian machine learning · 12%
Theoretical computer science
2 papers
Information theory · 35% Mathematical optimization · 35% Coding theory · 30%

Topics — the 27 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
1.632022
Planning in Stochastic Environments with a Learned Model · ICLR 2022
Procedural generalization by planning with self-supervised world models · ICLR 2022
Vector Quantized Models for Planning · ICML 2021
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
1.322024
Genie: Generative Interactive Environments · ICML 2024
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › sequential latent variable model
latent action models
0.812024
Genie: Generative Interactive Environments · ICML 2024
Machine learning › Representation and self-supervised learning
mutual information maximization
0.822019
Wasserstein Dependency Measure for Representation Learning · NeurIPS 2019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation
0.722019
On Variational Bounds of Mutual Information · ICML 2019
Mutual Information Neural Estimation · ICML 2018
Machine learning › Learning theory
generalization
0.612022
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
planning with learned models
0.612022
Planning in Stochastic Environments with a Learned Model · ICLR 2022
Machine learning › Reinforcement learning › generalization in reinforcement learning
procedural generalization
0.612022
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Generative modeling
generative adversarial network
0.522018
Mutual Information Neural Estimation · ICML 2018
Generative Adversarial Nets · NIPS 2014
Machine learning › Representation and self-supervised learning
vector quantization
0.512021
Vector Quantized Models for Planning · ICML 2021
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.522019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Wasserstein Dependency Measure for Representation Learning · NeurIPS 2019
Machine learning › Representation and self-supervised learning
contrastive learning
0.412019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
evidence lower bound
0.412019
On Variational Bounds of Mutual Information · ICML 2019
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.412019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Deep learning architectures and training
neural network estimator
0.312018
Mutual Information Neural Estimation · ICML 2018
Machine learning › Deep learning architectures and training
training objective
0.312018
Mutual Information Neural Estimation · ICML 2018
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Machine learning › Reinforcement learning › imitation learning
learning from observation
0.212024
Genie: Generative Interactive Environments · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming
0.212015
Efficient synthesis of probabilistic programs · PLDI 2015
Program synthesis and code generation
probabilistic program synthesis
0.212015
Efficient synthesis of probabilistic programs · PLDI 2015
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
probabilistic planning
0.212022
Planning in Stochastic Environments with a Learned Model · ICLR 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.112021
Vector Quantized Models for Planning · ICML 2021
Mathematical optimization › statistical estimation
bias-variance trade-off
0.112019
On Variational Bounds of Mutual Information · ICML 2019
Information theory › information measures
mutual information
0.112019
On Variational Bounds of Mutual Information · ICML 2019
Coding theory › source coding › rate-distortion theory
information bottleneck
0.112018
Mutual Information Neural Estimation · ICML 2018
High-performance computing
performance optimization at scale
0.112016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016

Methods — techniques the papers use, named apart from their topics

planning · 1.1unsupervised learning · 0.8spatiotemporal tokenizer · 0.8neural network parameterization · 0.8autoregressive dynamics models · 0.8self-supervised learning · 0.6learned dynamics model · 0.6monte carlo tree search · 0.5discrete autoencoder · 0.5variational bounds · 0.4variational bound · 0.4neural network estimation · 0.3gradient descent · 0.3batch dispatch · 0.2GPU-based inference · 0.2sketching · 0.2markov chain monte carlo · 0.2gaussian mixture approximation · 0.2
YearPublicationVenuePosition
2024 Genie: Generative Interactive Environments
abstract
We introduce Genie, the first *generative interactive environment* trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a *foundation world model*. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis *despite training without any ground-truth action labels* or other domain specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future.
Jake Bruce, Michael Dennis 0001, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes 0001, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie C. Y. Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott E. Reed, Jingwei Zhang 0001, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh 0001, Tim Rocktäschel
ICML18
2022 Procedural generalization by planning with self-supervised world models
Ankesh Anand, Jacob C. Walker, Yazhe Li, Eszter Vértes, Julian Schrittwieser, Sherjil Ozair, Theophane Weber, Jessica B. Hamrick
ICLR6
2022 Planning in Stochastic Environments with a Learned Model
Ioannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas Hubert, David Silver 0001
ICLR3
2021 Vector Quantized Models for Planning
abstract
Recent developments in the field of model-based RL have proven successful in a range of environments, especially ones where planning is essential. However, such successes have been limited to deterministic fully-observed environments. We present a new approach that handles stochastic and partially-observable environments. Our key insight is to use discrete autoencoders to capture the multiple possible effects of an action in a stochastic environment. We use a stochastic variant of Monte Carlo tree search to plan over both the agent’s actions and the discrete latent variables representing the environment’s response. Our approach significantly outperforms an offline version of MuZero on a stochastic interpretation of chess where the opponent is considered part of the environment. We also show that our approach scales to DeepMind Lab, a first-person 3D environment with large visual observations and partial observability.
Sherjil Ozair, Yazhe Li, Ali Razavi, Ioannis Antonoglou, Aäron van den Oord, Oriol Vinyals
ICML1
2020 SketchTransfer: A Challenging New Task for Exploring Detail-Invariance and the Abstractions Learned by Deep Networks
abstract
Deep networks have achieved excellent results in perceptual tasks, yet their ability to generalize to variations not seen during training has come under increasing scrutiny. In this work we focus on their ability to have invariance towards the presence or absence of details. For example, humans are able to watch cartoons, which are missing many visual details, without being explicitly trained to do so. As another example, 3D rendering software is a relatively recent development, yet people are able to understand such rendered scenes even though they are missing details (consider a film like Toy Story). The failure of ma- chine learning algorithms to do this indicates a significant gap in generalization between human abilities and the abilities of deep networks. We propose a dataset that will make it easier to study the detail-invariance problem concretely. We produce a concrete task for this: SketchTransfer, and we show that state-of-the-art domain transfer algorithms still struggle with this task. The state-of-the-art technique which achieves over 95% on MNIST → SVHN transfer only achieves 59% accuracy on the SketchTransfer task, which is much better than random (11% accuracy) but falls short of the 87% accuracy of a classifier trained directly on labeled sketches. This indicates that this task is approachable with today’s best methods but has substantial room for improvement.
Alex Lamb, Sherjil Ozair, Vikas Verma, David Ha
WACV2
2019 On Variational Bounds of Mutual Information
abstract
Estimating and optimizing Mutual Information (MI) is core to many problems in machine learning, but bounding MI in high dimensions is challenging. To establish tractable and scalable objectives, recent work has turned to variational bounds parameterized by neural networks. However, the relationships and tradeoffs between these bounds remains unclear. In this work, we unify these recent developments in a single framework. We find that the existing variational lower bounds degrade when the MI is large, exhibiting either high bias or high variance. To address this problem, we introduce a continuum of lower bounds that encompasses previous bounds and flexibly trades off bias and variance. On high-dimensional, controlled problems, we empirically characterize the bias and variance of the bounds and their gradients and demonstrate the effectiveness of these new bounds for estimation and representation learning.
Ben Poole, Sherjil Ozair, Aäron van den Oord, Alexander A. Alemi, George Tucker
ICML2
2019 Unsupervised State Representation Learning in Atari
abstract
State representation learning, or the ability to capture latent generative factors of an environment is crucial for building intelligent agents that can perform a wide variety of tasks. Learning such representations in an unsupervised manner without supervision from rewards is an open problem. We introduce a method that tries to learn better state representations by maximizing mutual information across spatially and temporally distinct features of a neural encoder of the observations. We also introduce a new benchmark based on Atari 2600 games where we evaluate representations based on how well they capture the ground truth state. We believe this new framework for evaluating representation learning models will be crucial for future representation learning research. Finally, we compare our technique with other state-of-the-art generative and contrastive representation learning methods.
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, R. Devon Hjelm
NeurIPS3
2019 Wasserstein Dependency Measure for Representation Learning
abstract
Mutual information maximization has emerged as a powerful learning objective for unsupervised representation learning obtaining state-of-the-art performance in applications such as object recognition, speech recognition, and reinforcement learning. However, such approaches are fundamentally limited since a tight lower bound on mutual information requires sample size exponential in the mutual information. This limits the applicability of these approaches for prediction tasks with high mutual information, such as in video understanding or reinforcement learning. In these settings, such techniques are prone to overfit, both in theory and in practice, and capture only a few of the relevant factors of variation. This leads to incomplete representations that are not optimal for downstream tasks. In this work, we empirically demonstrate that mutual information-based representation learning approaches do fail to learn complete representations on a number of designed and real-world tasks. To mitigate these problems we introduce the Wasserstein dependency measure, which learns more complete representations by using the Wasserstein distance instead of the KL divergence in the mutual information estimator. We show that a practical approximation to this theoretically motivated solution, constructed using Lipschitz constraint techniques from the GAN literature, achieves substantially improved results on tasks where incomplete representations are a major challenge.
Sherjil Ozair, Corey Lynch, Yoshua Bengio, Aäron van den Oord, Sergey Levine, Pierre Sermanet
NeurIPS1
2018 Mutual Information Neural Estimation
abstract
We argue that the estimation of mutual information between high dimensional continuous random variables can be achieved by gradient descent over neural networks. We present a Mutual Information Neural Estimator (MINE) that is linearly scalable in dimensionality as well as in sample size, trainable through back-prop, and strongly consistent. We present a handful of applications on which MINE can be used to minimize or maximize mutual information. We apply MINE to improve adversarially trained generative models. We also use MINE to implement the Information Bottleneck, applying it to supervised classification; our results demonstrate substantial improvement in flexibility and performance in these settings.
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, R. Devon Hjelm, Aaron C. Courville
ICML4
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML25
2015 Efficient synthesis of probabilistic programs
abstract
We show how to automatically synthesize probabilistic programs from real-world datasets. Such a synthesis is feasible due to a combination of two techniques: (1) We borrow the idea of ``sketching'' from synthesis of deterministic programs, and allow the programmer to write a skeleton program with ``holes''. Sketches enable the programmer to communicate domain-specific intuition about the structure of the desired program and prune the search space, and (2) we design an efficient Markov Chain Monte Carlo (MCMC) based synthesis algorithm to instantiate the holes in the sketch with program fragments. Our algorithm efficiently synthesizes a probabilistic program that is most consistent with the data. A core difficulty in synthesizing probabilistic programs is computing the likelihood L(P | D) of a candidate program P generating data D. We propose an approximate method to compute likelihoods using mixtures of Gaussian distributions, thereby avoiding expensive computation of integrals. The use of such approximations enables us to speed up evaluation of the likelihood of candidate programs by a factor of 1000, and makes Markov Chain Monte Carlo based search feasible. We have implemented our algorithm in a tool called PSKETCH, and our results are encouraging PSKETCH is able to automatically synthesize 16 non-trivial real-world probabilistic programs.
Aditya V. Nori, Sherjil Ozair, Sriram K. Rajamani, Deepak Vijaykeerthy
PLDI2
2014 Generative Adversarial Nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, Yoshua Bengio
NIPS6
2014 On the Equivalence between Deep NADE and Generative Stochastic Networks
Sherjil Ozair, Kyunghyun Cho, Yoshua Bengio
ECML/PKDD (3)2