VLDB 2026 Research / reviewers in the wild / expert
Denis Yarats
dblp:200/8142
· DBLP profile ↗
14ranked-venue papers
5as first author
8since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Reinforcement learning · 44% Deep learning architectures and training · 18% Optimization for machine learning · 16% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
data augmentation |
1.6 | 3 | 2022 | Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning · ICLR 2022 Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021 Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels · ICLR 2021 |
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning |
1.1 | 2 | 2022 | Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning · ICLR 2022 Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels · ICLR 2021 |
Machine learning › Reinforcement learning
continuous control |
0.6 | 1 | 2022 | Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning · ICLR 2022 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | Unsupervised Reinforcement Learning with Contrastive Intrinsic Control · NeurIPS 2022 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.6 | 1 | 2022 | Unsupervised Reinforcement Learning with Contrastive Intrinsic Control · NeurIPS 2022 |
Machine learning › Reinforcement learning
unsupervised reinforcement learning |
0.6 | 1 | 2022 | Unsupervised Reinforcement Learning with Contrastive Intrinsic Control · NeurIPS 2022 |
Machine learning › Reinforcement learning
actor-critic methods |
0.5 | 1 | 2021 | Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021 |
Machine learning › Optimization for machine learning
adaptive optimization |
0.5 | 1 | 2021 | On the Adequacy of Untuned Warmup for Adaptive Optimization · AAAI 2021 |
Machine learning › Deep learning architectures and training › data augmentation
automatic data augmentation |
0.5 | 1 | 2021 | Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021 |
Machine learning › Reinforcement learning
generalization in reinforcement learning |
0.5 | 1 | 2021 | Automatic Data Augmentation for Generalization in Reinforcement Learning · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning |
0.5 | 1 | 2021 | Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021 |
Machine learning › Optimization for machine learning › learning rate schedule
learning rate warmup |
0.5 | 1 | 2021 | On the Adequacy of Untuned Warmup for Adaptive Optimization · AAAI 2021 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.5 | 1 | 2021 | Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021 |
Machine learning › Representation and self-supervised learning › prototype learning
prototype-based representation |
0.5 | 1 | 2021 | Reinforcement Learning with Prototypical Representations · ICML 2021 |
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
reinforcement learning from pixels |
0.5 | 1 | 2021 | Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels · ICLR 2021 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.5 | 1 | 2021 | Reinforcement Learning with Prototypical Representations · ICML 2021 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.5 | 1 | 2021 | Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021 |
Machine learning › Optimization for machine learning › evolutionary computation
cross-entropy method |
0.4 | 1 | 2020 | The Differentiable Cross-Entropy Method · ICML 2020 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.4 | 1 | 2019 | Hierarchical Decision Making by Generating and Following Natural Language Instructions · NeurIPS 2019 |
Machine learning › Optimization for machine learning › gradient-based optimization
momentum methods |
0.4 | 1 | 2019 | Quasi-hyperbolic momentum and Adam for deep learning · ICLR (Poster) 2019 |
Machine learning › Optimization for machine learning
optimization |
0.4 | 1 | 2019 | Quasi-hyperbolic momentum and Adam for deep learning · ICLR (Poster) 2019 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.4 | 1 | 2019 | Quasi-hyperbolic momentum and Adam for deep learning · ICLR (Poster) 2019 |
Natural language and speech › Language models and text generation › text generation › structured generation
hierarchical text generation |
0.3 | 1 | 2018 | Hierarchical Text Generation and Planning for Strategic Dialogue · ICML 2018 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › self-play
self-play reinforcement learning |
0.3 | 1 | 2018 | Hierarchical Text Generation and Planning for Strategic Dialogue · ICML 2018 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
negotiation dialogue |
0.3 | 1 | 2017 | Deal or No Deal? End-to-End Learning of Negotiation Dialogues · EMNLP 2017 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2017 | Convolutional Sequence to Sequence Learning · ICML 2017 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.3 | 1 | 2017 | Convolutional Sequence to Sequence Learning · ICML 2017 |
Machine learning › Generative modeling
variational autoencoder |
0.1 | 1 | 2021 | Improving Sample Efficiency in Model-Free Reinforcement Learning from Images · AAAI 2021 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination |
0.1 | 1 | 2019 | Hierarchical Decision Making by Generating and Following Natural Language Instructions · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2017 | Convolutional Sequence to Sequence Learning · ICML 2017 |
Methods — techniques the papers use, named apart from their topics
data augmentation · 1.6adam · 0.9latent variable model · 0.7mutual information maximization · 0.6entropy maximization · 0.6contrastive learning · 0.6off-policy learning · 0.5linear warmup · 0.5auxiliary reconstruction loss · 0.5RAdam · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel Pinto |
ICLR | 1 |
| 2022 | Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlabstractWe introduce Contrastive Intrinsic Control (CIC), an unsupervised reinforcement learning (RL) algorithm that maximizes the mutual information between state-transitions and latent skill vectors. CIC utilizes contrastive learning between state-transitions and skills vectors to learn behaviour embeddings and maximizes the entropy of these embeddings as an intrinsic reward to encourage behavioural diversity. We evaluate our algorithm on the Unsupervised RL Benchmark (URLB) in the asymptotic state-based setting, which consists of a long reward-free pre-training phase followed by a short adaptation phase to downstream tasks with extrinsic rewards. We find that CIC improves over prior exploration algorithms in terms of adaptation efficiency to downstream tasks on state-based URLB. Michael Laskin, Hao Liu 0055, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, Pieter Abbeel |
NeurIPS | 4 |
| 2021 | On the Adequacy of Untuned Warmup for Adaptive OptimizationabstractAdaptive optimization algorithms such as Adam (Kingma and Ba, 2014) are widely used in deep learning. The stability of such algorithms is often improved with a warmup schedule for the learning rate. Motivated by the difficulty of choosing and tuning warmup schedules, recent work proposes automatic variance rectification of Adam's adaptive learning rate, claiming that this rectified approach ("RAdam") surpasses the vanilla Adam algorithm and reduces the need for expensive tuning of Adam with warmup. In this work, we refute this analysis and provide an alternative explanation for the necessity of warmup based on the magnitude of the update term, which is of greater relevance to training stability. We then provide some "rule-of-thumb" warmup schedules, and we demonstrate that simple untuned warmup of Adam performs more-or-less identically to RAdam in typical practical settings. We conclude by suggesting that practitioners stick to linear warmup with Adam, with a sensible default being linear warmup over 2 / (1 - β₂) training iterations. Jerry Ma, Denis Yarats |
AAAI | 2 |
| 2021 | Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesabstractTraining an agent to solve control tasks directly from high-dimensional images with model-free reinforcement learning (RL) has proven difficult. A promising approach is to learn a latent representation together with the control policy. However, fitting a high-capacity encoder using a scarce reward signal is sample inefficient and leads to poor performance. Prior work has shown that auxiliary losses, such as image reconstruction, can aid efficient representation learning. However, incorporating reconstruction loss into an off-policy learning algorithm often leads to training instability. We explore the underlying reasons and identify variational autoencoders, used by previous investigations, as the cause of the divergence. Following these findings, we propose effective techniques to improve training stability. This results in a simple approach capable of matching state-of-the-art model-free and model-based algorithms on MuJoCo control tasks. Furthermore, our approach demonstrates robustness to observational noise, surpassing existing approaches in this setting. Code, results, and videos are anonymously available at https://sites.google.com/view/sac-ae/home. Denis Yarats, Amy Zhang 0001, Ilya Kostrikov, Brandon Amos, Joelle Pineau, Rob Fergus |
AAAI | 1 |
| 2021 | Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Denis Yarats, Ilya Kostrikov, Rob Fergus |
ICLR | 1 |
| 2021 | Reinforcement Learning with Prototypical RepresentationsabstractLearning effective representations in image-based environments is crucial for sample efficient Reinforcement Learning (RL). Unfortunately, in RL, representation learning is confounded with the exploratory experience of the agent – learning a useful representation requires diverse data, while effective exploration is only possible with coherent representations. Furthermore, we would like to learn representations that not only generalize across tasks but also accelerate downstream exploration for efficient task-specific training. To address these challenges we propose Proto-RL, a self-supervised framework that ties representation learning with exploration through prototypical representations. These prototypes simultaneously serve as a summarization of the exploratory experience of an agent as well as a basis for representing observations. We pre-train these task-agnostic representations and prototypes on environments without downstream task information. This enables state-of-the-art downstream policy learning on a set of difficult continuous control tasks. Denis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel Pinto |
ICML | 1 |
| 2021 | Learning Navigation Skills for Legged Robots with Learned Robot EmbeddingsabstractRecent work has shown results on learning navigation policies for idealized cylinder agents in simulation and transferring them to real wheeled robots. Deploying such navigation policies on legged robots can be challenging due to their complex dynamics, and the large dynamical difference between cylinder agents and legged systems. In this work, we learn hierarchical navigation policies that account for the low-level dynamics of legged robots, such as maximum speed, slipping, contacts, and learn to successfully navigate cluttered indoor environments. To enable transfer of policies learned in simulation to new legged robots and hardware, we learn dynamics-aware navigation policies across multiple robots with robot-specific embeddings. The learned embedding is optimized on new robots, while the rest of the policy is kept fixed, allowing for quick adaptation. We train our policies across three legged robots in simulation - 2 quadrupeds (A1, AlienGo) and a hexapod (Daisy). At test time, we study the performance of our learned policy on two new legged robots in simulation (Laikago, 4-legged Daisy), and one real-world quadrupedal robot (A1). Our experiments show that our learned policy can sample-efficiently generalize to previously unseen robots, and enable sim-to-real transfer of navigation policies for legged robots. Joanne Truong, Denis Yarats, Tianyu Li 0005, Franziska Meier, Sonia Chernova, Dhruv Batra, Akshara Rai |
IROS | 2 |
| 2021 | Automatic Data Augmentation for Generalization in Reinforcement LearningabstractDeep reinforcement learning (RL) agents often fail to generalize beyond their training environments. To alleviate this problem, recent work has proposed the use of data augmentation. However, different tasks tend to benefit from different types of augmentations and selecting the right one typically requires expert knowledge. In this paper, we introduce three approaches for automatically finding an effective augmentation for any RL task. These are combined with two novel regularization terms for the policy and value function, required to make the use of data augmentation theoretically sound for actor-critic algorithms. Our method achieves a new state-of-the-art on the Procgen benchmark and outperforms popular RL algorithms on DeepMind Control tasks with distractors. In addition, our agent learns policies and representations which are more robust to changes in the environment that are irrelevant for solving the task, such as the background. Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, Rob Fergus |
NeurIPS | 3 |
| 2020 | The Differentiable Cross-Entropy MethodabstractWe study the Cross-Entropy Method (CEM) for the non-convex optimization of a continuous and parameterized objective function and introduce a differentiable variant that enables us to differentiate the output of CEM with respect to the objective function’s parameters. In the machine learning setting this brings CEM inside of the end-to-end learning pipeline where this has otherwise been impossible. We show applications in a synthetic energy-based structured prediction task and in non-convex continuous control. In the control setting we show how to embed optimal action sequences into a lower-dimensional space. This enables us to use policy optimization to fine-tune modeling components by differentiating through the CEM-based controller. Brandon Amos, Denis Yarats |
ICML | 2 |
| 2019 | Quasi-hyperbolic momentum and Adam for deep learning
Jerry Ma, Denis Yarats |
ICLR (Poster) | 2 |
| 2019 | Hierarchical Decision Making by Generating and Following Natural Language InstructionsabstractWe explore using latent natural language instructions as an expressive and compositional representation of complex actions for hierarchical decision making. Rather than directly selecting micro-actions, our agent first generates a latent plan in natural language, which is then executed by a separate model. We introduce a challenging real-time strategy game environment in which the actions of a large number of units must be coordinated across long time scales. We gather a dataset of 76 thousand pairs of instructions and executions from human play, and train instructor and executor models. Experiments show that models using natural language as a latent variable significantly outperform models that directly imitate human actions. The compositional structure of language proves crucial to its effectiveness for action representation. We also release our code, models and data. Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, Mike Lewis |
NeurIPS | 2 |
| 2018 | Hierarchical Text Generation and Planning for Strategic DialogueabstractEnd-to-end models for goal-orientated dialogue are challenging to train, because linguistic and strategic aspects are entangled in latent state vectors. We introduce an approach to learning representations of messages in dialogues by maximizing the likelihood of subsequent sentences and actions, which decouples the semantics of the dialogue utterance from its linguistic realization. We then use these latent sentence representations for hierarchical language generation, planning and reinforcement learning. Experiments show that our approach increases the end-task reward achieved by the model, improves the effectiveness of long-term planning using rollouts, and allows self-play reinforcement learning to improve decision making without diverging from human language. Our hierarchical latent-variable model outperforms previous work both linguistically and strategically. Denis Yarats, Mike Lewis |
ICML | 1 |
| 2017 | Deal or No Deal? End-to-End Learning of Negotiation DialoguesabstractMuch of human dialogue occurs in semicooperative settings, where agents with different goals attempt to agree on common decisions.Negotiations require complex communication and reasoning skills, but success is easy to measure, making this an interesting task for AI.We gather a large dataset of human-human negotiations on a multi-issue bargaining task, where agents who cannot observe each other's reward functions must reach an agreement (or a deal) via natural language dialogue.For the first time, we show it is possible to train end-to-end models for negotiation, which must learn both linguistic and reasoning skills with no annotated dialogue states.We also introduce dialogue rollouts, in which the model plans ahead by simulating possible complete continuations of the conversation, and find that this technique dramatically improves performance.Our code and dataset are publicly available.1 Mike Lewis, Denis Yarats, Yann N. Dauphin, Devi Parikh, Dhruv Batra |
EMNLP | 2 |
| 2017 | Convolutional Sequence to Sequence LearningabstractThe prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully parallelized during training to better exploit the GPU hardware and optimization is easier since the number of non-linearities is fixed and independent of the input length. Our use of gated linear units eases gradient propagation and we equip each decoder layer with a separate attention module. We outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT’14 English-German and WMT’14 English-French translation at an order of magnitude faster speed, both on GPU and CPU. Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin |
ICML | 4 |