Denis Yarats

dblp:200/8142 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
8since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Reinforcement learning · 44% Deep learning architectures and training · 18% Optimization for machine learning · 16%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
data augmentation
1.632022
Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning · ICLR 2022
Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels · ICLR 2021
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning
1.122022
Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning · ICLR 2022
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels · ICLR 2021
Machine learning › Reinforcement learning
continuous control
0.612022
Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning · ICLR 2022
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Unsupervised Reinforcement Learning with Contrastive Intrinsic Control · NeurIPS 2022
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.612022
Unsupervised Reinforcement Learning with Contrastive Intrinsic Control · NeurIPS 2022
Machine learning › Reinforcement learning
unsupervised reinforcement learning
0.612022
Unsupervised Reinforcement Learning with Contrastive Intrinsic Control · NeurIPS 2022
Machine learning › Reinforcement learning
actor-critic methods
0.512021
Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021
Machine learning › Optimization for machine learning
adaptive optimization
0.512021
On the Adequacy of Untuned Warmup for Adaptive Optimization · AAAI 2021
Machine learning › Deep learning architectures and training › data augmentation
automatic data augmentation
0.512021
Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.512021
Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning
0.512021
Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021
Machine learning › Optimization for machine learning › learning rate schedule
learning rate warmup
0.512021
On the Adequacy of Untuned Warmup for Adaptive Optimization · AAAI 2021
Machine learning › Reinforcement learning
model-free reinforcement learning
0.512021
Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021
Machine learning › Representation and self-supervised learning › prototype learning
prototype-based representation
0.512021
Reinforcement Learning with Prototypical Representations · ICML 2021
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
reinforcement learning from pixels
0.512021
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels · ICLR 2021
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.512021
Reinforcement Learning with Prototypical Representations · ICML 2021
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.512021
Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021
Machine learning › Optimization for machine learning › evolutionary computation
cross-entropy method
0.412020
The Differentiable Cross-Entropy Method · ICML 2020
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412019
Hierarchical Decision Making by Generating and Following Natural Language Instructions · NeurIPS 2019
Machine learning › Optimization for machine learning › gradient-based optimization
momentum methods
0.412019
Quasi-hyperbolic momentum and Adam for deep learning · ICLR (Poster) 2019
Machine learning › Optimization for machine learning
optimization
0.412019
Quasi-hyperbolic momentum and Adam for deep learning · ICLR (Poster) 2019
Machine learning › Optimization for machine learning
stochastic optimization
0.412019
Quasi-hyperbolic momentum and Adam for deep learning · ICLR (Poster) 2019
Natural language and speech › Language models and text generation › text generation › structured generation
hierarchical text generation
0.312018
Hierarchical Text Generation and Planning for Strategic Dialogue · ICML 2018
Machine learning › Reinforcement learning › multi-agent reinforcement learning › self-play
self-play reinforcement learning
0.312018
Hierarchical Text Generation and Planning for Strategic Dialogue · ICML 2018
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
negotiation dialogue
0.312017
Deal or No Deal? End-to-End Learning of Negotiation Dialogues · EMNLP 2017
Natural language and speech › Machine translation
neural machine translation
0.312017
Convolutional Sequence to Sequence Learning · ICML 2017
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning
0.312017
Convolutional Sequence to Sequence Learning · ICML 2017
Machine learning › Generative modeling
variational autoencoder
0.112021
Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
0.112019
Hierarchical Decision Making by Generating and Following Natural Language Instructions · NeurIPS 2019
Machine learning › Deep learning architectures and training
attention mechanism
0.112017
Convolutional Sequence to Sequence Learning · ICML 2017

Methods — techniques the papers use, named apart from their topics

data augmentation · 1.6adam · 0.9latent variable model · 0.7mutual information maximization · 0.6entropy maximization · 0.6contrastive learning · 0.6off-policy learning · 0.5linear warmup · 0.5auxiliary reconstruction loss · 0.5RAdam · 0.5
YearPublicationVenuePosition
2022 Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel Pinto
ICLR1
2022 Unsupervised Reinforcement Learning with Contrastive Intrinsic Control
abstract
We introduce Contrastive Intrinsic Control (CIC), an unsupervised reinforcement learning (RL) algorithm that maximizes the mutual information between state-transitions and latent skill vectors. CIC utilizes contrastive learning between state-transitions and skills vectors to learn behaviour embeddings and maximizes the entropy of these embeddings as an intrinsic reward to encourage behavioural diversity. We evaluate our algorithm on the Unsupervised RL Benchmark (URLB) in the asymptotic state-based setting, which consists of a long reward-free pre-training phase followed by a short adaptation phase to downstream tasks with extrinsic rewards. We find that CIC improves over prior exploration algorithms in terms of adaptation efficiency to downstream tasks on state-based URLB.
Michael Laskin, Hao Liu 0055, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, Pieter Abbeel
NeurIPS4
2021 On the Adequacy of Untuned Warmup for Adaptive Optimization
abstract
Adaptive optimization algorithms such as Adam (Kingma and Ba, 2014) are widely used in deep learning. The stability of such algorithms is often improved with a warmup schedule for the learning rate. Motivated by the difficulty of choosing and tuning warmup schedules, recent work proposes automatic variance rectification of Adam's adaptive learning rate, claiming that this rectified approach ("RAdam") surpasses the vanilla Adam algorithm and reduces the need for expensive tuning of Adam with warmup. In this work, we refute this analysis and provide an alternative explanation for the necessity of warmup based on the magnitude of the update term, which is of greater relevance to training stability. We then provide some "rule-of-thumb" warmup schedules, and we demonstrate that simple untuned warmup of Adam performs more-or-less identically to RAdam in typical practical settings. We conclude by suggesting that practitioners stick to linear warmup with Adam, with a sensible default being linear warmup over 2 / (1 - β₂) training iterations.
Jerry Ma, Denis Yarats
AAAI2
2021 Improving Sample Efficiency in Model-Free Reinforcement Learning from Images
abstract
Training an agent to solve control tasks directly from high-dimensional images with model-free reinforcement learning (RL) has proven difficult. A promising approach is to learn a latent representation together with the control policy. However, fitting a high-capacity encoder using a scarce reward signal is sample inefficient and leads to poor performance. Prior work has shown that auxiliary losses, such as image reconstruction, can aid efficient representation learning. However, incorporating reconstruction loss into an off-policy learning algorithm often leads to training instability. We explore the underlying reasons and identify variational autoencoders, used by previous investigations, as the cause of the divergence. Following these findings, we propose effective techniques to improve training stability. This results in a simple approach capable of matching state-of-the-art model-free and model-based algorithms on MuJoCo control tasks. Furthermore, our approach demonstrates robustness to observational noise, surpassing existing approaches in this setting. Code, results, and videos are anonymously available at https://sites.google.com/view/sac-ae/home.
Denis Yarats, Amy Zhang 0001, Ilya Kostrikov, Brandon Amos, Joelle Pineau, Rob Fergus
AAAI1
2021 Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Denis Yarats, Ilya Kostrikov, Rob Fergus
ICLR1
2021 Reinforcement Learning with Prototypical Representations
abstract
Learning effective representations in image-based environments is crucial for sample efficient Reinforcement Learning (RL). Unfortunately, in RL, representation learning is confounded with the exploratory experience of the agent – learning a useful representation requires diverse data, while effective exploration is only possible with coherent representations. Furthermore, we would like to learn representations that not only generalize across tasks but also accelerate downstream exploration for efficient task-specific training. To address these challenges we propose Proto-RL, a self-supervised framework that ties representation learning with exploration through prototypical representations. These prototypes simultaneously serve as a summarization of the exploratory experience of an agent as well as a basis for representing observations. We pre-train these task-agnostic representations and prototypes on environments without downstream task information. This enables state-of-the-art downstream policy learning on a set of difficult continuous control tasks.
Denis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel Pinto
ICML1
2021 Learning Navigation Skills for Legged Robots with Learned Robot Embeddings
abstract
Recent work has shown results on learning navigation policies for idealized cylinder agents in simulation and transferring them to real wheeled robots. Deploying such navigation policies on legged robots can be challenging due to their complex dynamics, and the large dynamical difference between cylinder agents and legged systems. In this work, we learn hierarchical navigation policies that account for the low-level dynamics of legged robots, such as maximum speed, slipping, contacts, and learn to successfully navigate cluttered indoor environments. To enable transfer of policies learned in simulation to new legged robots and hardware, we learn dynamics-aware navigation policies across multiple robots with robot-specific embeddings. The learned embedding is optimized on new robots, while the rest of the policy is kept fixed, allowing for quick adaptation. We train our policies across three legged robots in simulation - 2 quadrupeds (A1, AlienGo) and a hexapod (Daisy). At test time, we study the performance of our learned policy on two new legged robots in simulation (Laikago, 4-legged Daisy), and one real-world quadrupedal robot (A1). Our experiments show that our learned policy can sample-efficiently generalize to previously unseen robots, and enable sim-to-real transfer of navigation policies for legged robots.
Joanne Truong, Denis Yarats, Tianyu Li 0005, Franziska Meier, Sonia Chernova, Dhruv Batra, Akshara Rai
IROS2
2021 Automatic Data Augmentation for Generalization in Reinforcement Learning
abstract
Deep reinforcement learning (RL) agents often fail to generalize beyond their training environments. To alleviate this problem, recent work has proposed the use of data augmentation. However, different tasks tend to benefit from different types of augmentations and selecting the right one typically requires expert knowledge. In this paper, we introduce three approaches for automatically finding an effective augmentation for any RL task. These are combined with two novel regularization terms for the policy and value function, required to make the use of data augmentation theoretically sound for actor-critic algorithms. Our method achieves a new state-of-the-art on the Procgen benchmark and outperforms popular RL algorithms on DeepMind Control tasks with distractors. In addition, our agent learns policies and representations which are more robust to changes in the environment that are irrelevant for solving the task, such as the background.
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, Rob Fergus
NeurIPS3
2020 The Differentiable Cross-Entropy Method
abstract
We study the Cross-Entropy Method (CEM) for the non-convex optimization of a continuous and parameterized objective function and introduce a differentiable variant that enables us to differentiate the output of CEM with respect to the objective function’s parameters. In the machine learning setting this brings CEM inside of the end-to-end learning pipeline where this has otherwise been impossible. We show applications in a synthetic energy-based structured prediction task and in non-convex continuous control. In the control setting we show how to embed optimal action sequences into a lower-dimensional space. This enables us to use policy optimization to fine-tune modeling components by differentiating through the CEM-based controller.
Brandon Amos, Denis Yarats
ICML2
2019 Quasi-hyperbolic momentum and Adam for deep learning
Jerry Ma, Denis Yarats
ICLR (Poster)2
2019 Hierarchical Decision Making by Generating and Following Natural Language Instructions
abstract
We explore using latent natural language instructions as an expressive and compositional representation of complex actions for hierarchical decision making. Rather than directly selecting micro-actions, our agent first generates a latent plan in natural language, which is then executed by a separate model. We introduce a challenging real-time strategy game environment in which the actions of a large number of units must be coordinated across long time scales. We gather a dataset of 76 thousand pairs of instructions and executions from human play, and train instructor and executor models. Experiments show that models using natural language as a latent variable significantly outperform models that directly imitate human actions. The compositional structure of language proves crucial to its effectiveness for action representation. We also release our code, models and data.
Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, Mike Lewis
NeurIPS2
2018 Hierarchical Text Generation and Planning for Strategic Dialogue
abstract
End-to-end models for goal-orientated dialogue are challenging to train, because linguistic and strategic aspects are entangled in latent state vectors. We introduce an approach to learning representations of messages in dialogues by maximizing the likelihood of subsequent sentences and actions, which decouples the semantics of the dialogue utterance from its linguistic realization. We then use these latent sentence representations for hierarchical language generation, planning and reinforcement learning. Experiments show that our approach increases the end-task reward achieved by the model, improves the effectiveness of long-term planning using rollouts, and allows self-play reinforcement learning to improve decision making without diverging from human language. Our hierarchical latent-variable model outperforms previous work both linguistically and strategically.
Denis Yarats, Mike Lewis
ICML1
2017 Deal or No Deal? End-to-End Learning of Negotiation Dialogues
abstract
Much of human dialogue occurs in semicooperative settings, where agents with different goals attempt to agree on common decisions.Negotiations require complex communication and reasoning skills, but success is easy to measure, making this an interesting task for AI.We gather a large dataset of human-human negotiations on a multi-issue bargaining task, where agents who cannot observe each other's reward functions must reach an agreement (or a deal) via natural language dialogue.For the first time, we show it is possible to train end-to-end models for negotiation, which must learn both linguistic and reasoning skills with no annotated dialogue states.We also introduce dialogue rollouts, in which the model plans ahead by simulating possible complete continuations of the conversation, and find that this technique dramatically improves performance.Our code and dataset are publicly available.1
Mike Lewis, Denis Yarats, Yann N. Dauphin, Devi Parikh, Dhruv Batra
EMNLP2
2017 Convolutional Sequence to Sequence Learning
abstract
The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully parallelized during training to better exploit the GPU hardware and optimization is easier since the number of non-linearities is fixed and independent of the input length. Our use of gated linear units eases gradient propagation and we equip each decoder layer with a separate attention module. We outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT’14 English-German and WMT’14 English-French translation at an order of magnitude faster speed, both on GPU and CPU.
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin
ICML4