Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Junhyuk Oh

dblp:167/4825 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0003-4383-6396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
20 papers
Reinforcement learning · 72% Transfer learning and domain adaptation · 9% Deep learning architectures and training · 6%

Topics — the 30 heaviest of 56, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
1.242023
Contingency-Aware Exploration in Reinforcement Learning · ICLR (Poster) 2019
Self-Imitation Learning · ICML 2018
Control of Memory, Active Perception, and Action in Minecraft · ICML 2016
Machine learning › Reinforcement learning
hierarchical reinforcement learning
1.132021
Discovery of Options via Meta-Learned Subgoals · NeurIPS 2021
Hierarchical Reinforcement Learning for Zero-shot Generalization with Subtask Dependencies · NeurIPS 2018
Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning · ICML 2017
Machine learning › Reinforcement learning
deep reinforcement learning
0.922023
Deep Reinforcement Learning with Plasticity Injection · NeurIPS 2023
Control of Memory, Active Perception, and Action in Minecraft · ICML 2016
Natural language and speech › Language models and text generation
alignment
0.912025
Learning from negative feedback, or positive feedback or both · ICLR 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
learning from feedback
0.912025
Learning from negative feedback, or positive feedback or both · ICLR 2025
Machine learning › Reinforcement learning › meta-reinforcement learning
meta-gradient reinforcement learning
0.922020
A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020
Meta-Gradient Reinforcement Learning with an Objective Discovered Online · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
meta-learning
0.922020
Meta-Gradient Reinforcement Learning with an Objective Discovered Online · NeurIPS 2020
Discovering Reinforcement Learning Algorithms · NeurIPS 2020
Machine learning › Reinforcement learning
preference learning
0.912025
Learning from negative feedback, or positive feedback or both · ICLR 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Learning from negative feedback, or positive feedback or both · ICLR 2025
Machine learning › Reinforcement learning
actor-critic methods
0.822020
A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020
Self-Imitation Learning · ICML 2018
Machine learning › Reinforcement learning
value-based reinforcement learning
0.722020
Discovering Reinforcement Learning Algorithms · NeurIPS 2020
Value Prediction Network · NIPS 2017
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning
0.712023
In-context Reinforcement Learning with Algorithm Distillation · ICLR 2023
Machine learning › Deep learning architectures and training › training dynamics
plasticity loss
0.712023
Deep Reinforcement Learning with Plasticity Injection · NeurIPS 2023
Machine learning › Deep learning architectures and training › biologically plausible learning
synaptic plasticity
0.712023
Deep Reinforcement Learning with Plasticity Injection · NeurIPS 2023
Machine learning › Reinforcement learning
meta-reinforcement learning
0.612022
Introducing Symmetries to Black Box Meta Reinforcement Learning · AAAI 2022
Machine learning › Reinforcement learning
model-based reinforcement learning
0.522017
Value Prediction Network · NIPS 2017
Action-Conditional Video Prediction using Deep Networks in Atari Games · NIPS 2015
Machine learning › Reinforcement learning
constrained reinforcement learning
0.512021
Balancing Constraints and Rewards with Meta-Gradient D4PG · ICLR 2021
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning
0.512021
Balancing Constraints and Rewards with Meta-Gradient D4PG · ICLR 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.512021
Discovery of Options via Meta-Learned Subgoals · NeurIPS 2021
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.522020
On Learning Intrinsic Rewards for Policy Gradient Methods · NeurIPS 2018
Discovering Reinforcement Learning Algorithms · NeurIPS 2020
Machine learning › Optimization for machine learning
hyperparameter optimization
0.412020
A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
What Can Learned Intrinsic Rewards Capture? · ICML 2020
Machine learning › Reinforcement learning
learned objective functions
0.412020
Meta-Gradient Reinforcement Learning with an Objective Discovered Online · NeurIPS 2020
Machine learning › Reinforcement learning › meta-reinforcement learning
learned update rules
0.412020
Discovering Reinforcement Learning Algorithms · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-gradient
0.412020
What Can Learned Intrinsic Rewards Capture? · ICML 2020
Machine learning › Reinforcement learning
reward learning
0.412020
What Can Learned Intrinsic Rewards Capture? · ICML 2020
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks
0.412019
Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019
Machine learning › Reinforcement learning › value function estimation
general value functions
0.412019
Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019
Machine learning › Reinforcement learning
value function estimation
0.412019
Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019
Machine learning › Reinforcement learning › reward learning
intrinsic reward learning
0.312018
On Learning Intrinsic Rewards for Policy Gradient Methods · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

meta-gradient · 1.4actor-critic · 1.1preference optimization · 0.9meta-gradient descent · 0.9expectation-maximization · 0.9plasticity injection · 0.7symmetry · 0.6neural network · 0.6black-box meta RL · 0.6distributed distributional deterministic policy gradients · 0.5
YearPublicationVenuePosition
2025 Learning from negative feedback, or positive feedback or both
abstract
Existing preference optimization methods often assume scenarios where paired preference feedback (preferred/positive vs. dis-preferred/negative examples) is available. This requirement limits their applicability in scenarios where only unpaired feedback—for example, either positive or negative— is available. To address this, we introduce a novel approach that decouples learning from positive and negative feedback. This decoupling enables control over the influence of each feedback type and, importantly, allows learning even when only one feedback type is present. A key contribution is demonstrating stable learning from negative feedback alone, a capability not well-addressed by current methods. Our approach builds upon the probabilistic framework introduced in (Dayan and Hinton, 1997), which uses expectation-maximization (EM) to directly optimize the probability of positive outcomes (as opposed to classic expected reward maximization). We address a key limitation in current EM-based methods: they solely maximize the likelihood of positive examples, while neglecting negative ones. We show how to extend EM algorithms to explicitly incorporate negative examples, leading to a theoretically grounded algorithm that offers an intuitive and versatile way to learn from both positive and negative feedback. We evaluate our approach for training language models based on human feedback as well as training policies for sequential decision-making problems, where learned value functions are available.
Abbas Abdolmaleki, Bilal Piot, Bobak Shahriari, Jost Tobias Springenberg, Tim Hertweck, Michael Bloesch, Rishabh Joshi, Thomas Lampe, Junhyuk Oh, Nicolas Heess, Jonas Buchli, Martin A. Riedmiller
ICLR9
2025 DataRater: Meta-Learned Dataset Curation
abstract
The quality of foundation models depends heavily on their training data. Consequently, great efforts have been put into dataset curation. Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics. An approach that is ultimately more scalable (let alone more satisfying) is to \emph{learn} which data is actually valuable for training. This type of meta-learning could allow more sophisticated, fine-grained, and effective curation. Our proposed \emph{DataRater} is an instance of this idea. It estimates the value of training on any particular data point. This is done by meta-learning using `meta-gradients', with the objective of improving training efficiency on held out data. In extensive experiments across a range of model scales and datasets, we find that using our DataRater to filter data is highly effective, resulting in significantly improved compute efficiency.
Dan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf, Matteo Hessel, Jeremy Shar, Junhyuk Oh, András György 0001, Tom Schaul, Jeffrey Dean, Hado van Hasselt, David Silver 0001
NeurIPS7
2023 In-context Reinforcement Learning with Algorithm Distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen 0001, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh 0001, Volodymyr Mnih
ICLR3
2023 Deep Reinforcement Learning with Plasticity Injection
abstract
A growing body of evidence suggests that neural networks employed in deep reinforcement learning (RL) gradually lose their plasticity, the ability to learn from new data; however, the analysis and mitigation of this phenomenon is hampered by the complex relationship between plasticity, exploration, and performance in RL. This paper introduces plasticity injection, a minimalistic intervention that increases the network plasticity without changing the number of trainable parameters or biasing the predictions. The applications of this intervention are two-fold: first, as a diagnostic tool — if injection increases the performance, we may conclude that an agent's network was losing its plasticity. This tool allows us to identify a subset of Atari environments where the lack of plasticity causes performance plateaus, motivating future studies on understanding and combating plasticity loss. Second, plasticity injection can be used to improve the computational efficiency of RL training if the agent has to re-learn from scratch due to exhausted plasticity or by growing the agent's network dynamically without compromising performance. The results on Atari show that plasticity injection attains stronger performance compared to alternative methods while being computationally efficient.
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle, Razvan Pascanu, Will Dabney, André Barreto 0001
NeurIPS2
2022 Introducing Symmetries to Black Box Meta Reinforcement Learning
abstract
Meta reinforcement learning (RL) attempts to discover new RL algorithms automatically from environment interaction. In so-called black-box approaches, the policy and the learning algorithm are jointly represented by a single neural network. These methods are very flexible, but they tend to underperform compared to human-engineered RL algorithms in terms of generalisation to new, unseen environments. In this paper, we explore the role of symmetries in meta-generalisation. We show that a recent successful meta RL approach that meta-learns an objective for backpropagation-based learning exhibits certain symmetries (specifically the reuse of the learning rule, and invariance to input and output permutations) that are not present in typical black-box meta RL systems. We hypothesise that these symmetries can play an important role in meta-generalisation. Building off recent work in black-box supervised meta learning, we develop a black-box meta RL system that exhibits these same symmetries. We show through careful experimentation that incorporating these symmetries can lead to algorithms with a greater ability to generalise to unseen action & observation spaces, tasks, and environments.
Louis Kirsch, Sebastian Flennerhag, Hado van Hasselt, Abram L. Friesen, Junhyuk Oh, Yutian Chen 0001
AAAI5
2021 Balancing Constraints and Rewards with Meta-Gradient D4PG
Dan Andrei Calian, Daniel J. Mankowitz, Tom Zahavy, Zhongwen Xu, Junhyuk Oh, Nir Levine, Timothy A. Mann
ICLR5
2021 Discovery of Options via Meta-Learned Subgoals
abstract
Temporal abstractions in the form of options have been shown to help reinforcement learning (RL) agents learn faster. However, despite prior work on this topic, the problem of discovering options through interaction with an environment remains a challenge. In this paper, we introduce a novel meta-gradient approach for discovering useful options in multi-task RL environments. Our approach is based on a manager-worker decomposition of the RL agent, in which a manager maximises rewards from the environment by learning a task-dependent policy over both a set of task-independent discovered-options and primitive actions. The option-reward and termination functions that define a subgoal for each option are parameterised as neural networks and trained via meta-gradients to maximise their usefulness. Empirical analysis on gridworld and DeepMind Lab tasks show that: (1) our approach can discover meaningful and diverse temporally-extended options in multi-task RL domains, (2) the discovered options are frequently used by the agent while learning to solve the training tasks, and (3) that the discovered options help a randomly initialised manager learn faster in completely new tasks.
Vivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu, Junhyuk Oh, Iurii Kemaev, Hado van Hasselt, David Silver 0001, Satinder Singh 0001
NeurIPS5
2020 What Can Learned Intrinsic Rewards Capture?
abstract
The objective of a reinforcement learning agent is to behave so as to maximise the sum of a suitable scalar function of state: the reward. These rewards are typically given and immutable. In this paper, we instead consider the proposition that the reward function itself can be a good locus of learned knowledge. To investigate this, we propose a scalable meta-gradient framework for learning useful intrinsic reward functions across multiple lifetimes of experience. Through several proof-of-concept experiments, we show that it is feasible to learn and capture knowledge about long-term exploration and exploitation into a reward function. Furthermore, we show that unlike policy transfer methods that capture “how” the agent should behave, the learned reward functions can generalise to other kinds of agents and to changes in the dynamics of the environment by capturing “what” the agent should strive to do.
Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver 0001, Satinder Singh 0001
ICML2
2020 Discovering Reinforcement Learning Algorithms
abstract
Reinforcement learning (RL) algorithms update an agent’s parameters according to one of several possible rules, discovered manually through years of research. Automating the discovery of update rules from data could lead to more efficient algorithms, or algorithms that are better adapted to specific environments. Although there have been prior attempts at addressing this significant scientific challenge, it remains an open question whether it is feasible to discover alternatives to fundamental concepts of RL such as value functions and temporal-difference learning. This paper introduces a new meta-learning approach that discovers an entire update rule which includes both what to predict' (e.g. value functions) andhow to learn from it' (e.g. bootstrapping) by interacting with a set of environments. The output of this method is an RL algorithm that we call Learned Policy Gradient (LPG). Empirical results show that our method discovers its own alternative to the concept of value functions. Furthermore it discovers a bootstrapping mechanism to maintain and use its predictions. Surprisingly, when trained solely on toy environments, LPG generalises effectively to complex Atari games and achieves non-trivial performance. This shows the potential to discover general RL algorithms from data.
Junhyuk Oh, Matteo Hessel, Wojciech Czarnecki 0001, Zhongwen Xu, Hado van Hasselt, Satinder Singh 0001, David Silver 0001
NeurIPS1
2020 Meta-Gradient Reinforcement Learning with an Objective Discovered Online
abstract
Deep reinforcement learning includes a broad family of algorithms that parameterise an internal representation, such as a value function or policy, by a deep neural network. Each algorithm optimises its parameters with respect to an objective, such as Q-learning or policy gradient, that defines its semantics. In this work, we propose an algorithm based on meta-gradient descent that discovers its own objective, flexibly parameterised by a deep neural network, solely from interactive experience with its environment. Over time, this allows the agent to learn how to learn increasingly effectively. Furthermore, because the objective is discovered online, it can adapt to changes over time. We demonstrate that the algorithm discovers how to address several important issues in RL, such as bootstrapping, non-stationarity, and off-policy learning. On the Atari Learning Environment, the meta-gradient algorithm adapts over time to learn with greater efficiency, eventually outperforming the median score of a strong actor-critic baseline.
Zhongwen Xu, Hado van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh 0001, David Silver 0001
NeurIPS4
2020 A Self-Tuning Actor-Critic Algorithm
abstract
Reinforcement learning algorithms are highly sensitive to the choice of hyperparameters, typically requiring significant manual effort to identify hyperparameters that perform well on a new domain. In this paper, we take a step towards addressing this issue by using metagradients to automatically adapt hyperparameters online by meta-gradient descent (Xu et al., 2018). We apply our algorithm, Self-Tuning Actor-Critic (STAC), to self-tune all the differentiable hyperparameters of an actor-critic loss function, to discover auxiliary tasks, and to improve off-policy learning using a novel leaky V-trace operator. STAC is simple to use, sample efficient and does not require a significant increase in compute. Ablative studies show that the overall performance of STAC improved as we adapt more hyperparameters. When applied to the Arcade Learning Environment (Bellemare et al. 2012), STAC improved the median human normalized score in 200M steps from 243% to 364%. When applied to the DM Control suite (Tassa et al., 2018), STAC improved the mean score in 30M steps from 217 to 389 when learning with features, from 108 to 202 when learning from pixels, and from 195 to 295 in the Real-World Reinforcement Learning Challenge (Dulac-Arnold et al., 2020).
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver 0001, Satinder Singh 0001
NeurIPS5
2019 Contingency-Aware Exploration in Reinforcement Learning
Yijie Guo, Marcin Moczulski, Junhyuk Oh, Neal Wu, Mohammad Norouzi 0002, Honglak Lee
ICLR (Poster)4
2019 Discovery of Useful Questions as Auxiliary Tasks
abstract
Arguably, intelligent agents ought to be able to discover their own questions so that in learning answers for them they learn unanticipated useful knowledge and skills; this departs from the focus in much of machine learning on agents learning answers to externally defined questions. We present a novel method for a reinforcement learning (RL) agent to discover questions formulated as general value functions or GVFs, a fairly rich form of knowledge representation. Specifically, our method uses non-myopic meta-gradients to learn GVF-questions such that learning answers to them, as an auxiliary task, induces useful representations for the main task faced by the RL agent. We demonstrate that auxiliary tasks based on the discovered GVFs are sufficient, on their own, to build representations that support main task learning, and that they do so better than popular hand-designed auxiliary tasks from the literature. Furthermore, we show, in the context of Atari2600 videogames, how such auxiliary tasks, meta-learned alongside the main task, can improve the data efficiency of an actor-critic agent.
Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Janarthanan Rajendran, Richard L. Lewis, Junhyuk Oh, Hado van Hasselt, David Silver 0001, Satinder Singh 0001
NeurIPS6
2018 Self-Imitation Learning
abstract
This paper proposes Self-Imitation Learning (SIL), a simple off-policy actor-critic algorithm that learns to reproduce the agent’s past good decisions. This algorithm is designed to verify our hypothesis that exploiting past good experiences can indirectly drive deep exploration. Our empirical results show that SIL significantly improves advantage actor-critic (A2C) on several hard exploration Atari games and is competitive to the state-of-the-art count-based exploration methods. We also show that SIL improves proximal policy optimization (PPO) on MuJoCo tasks.
Junhyuk Oh, Yijie Guo, Satinder Singh 0001, Honglak Lee
ICML1
2018 Hierarchical Reinforcement Learning for Zero-shot Generalization with Subtask Dependencies
abstract
We introduce a new RL problem where the agent is required to generalize to a previously-unseen environment characterized by a subtask graph which describes a set of subtasks and their dependencies. Unlike existing hierarchical multitask RL approaches that explicitly describe what the agent should do at a high level, our problem only describes properties of subtasks and relationships among them, which requires the agent to perform complex reasoning to find the optimal subtask to execute. To solve this problem, we propose a neural subtask graph solver (NSGS) which encodes the subtask graph using a recursive neural network embedding. To overcome the difficulty of training, we propose a novel non-parametric gradient-based policy, graph reward propagation, to pre-train our NSGS agent and further finetune it through actor-critic method. The experimental results on two 2D visual domains show that our agent can perform complex reasoning to find a near-optimal way of executing the subtask graph and generalize well to the unseen subtask graphs. In addition, we compare our agent with a Monte-Carlo tree search (MCTS) method showing that our method is much more efficient than MCTS, and the performance of NSGS can be further improved by combining it with MCTS.
Sungryull Sohn, Junhyuk Oh, Honglak Lee
NeurIPS2
2018 On Learning Intrinsic Rewards for Policy Gradient Methods
abstract
In many sequential decision making tasks, it is challenging to design reward functions that help an RL agent efficiently learn behavior that is considered good by the agent designer. A number of different formulations of the reward-design problem, or close variants thereof, have been proposed in the literature. In this paper we build on the Optimal Rewards Framework of Singh et al. that defines the optimal intrinsic reward function as one that when used by an RL agent achieves behavior that optimizes the task-specifying or extrinsic reward function. Previous work in this framework has shown how good intrinsic reward functions can be learned for lookahead search based planning agents. Whether it is possible to learn intrinsic reward functions for learning agents remains an open problem. In this paper we derive a novel algorithm for learning intrinsic rewards for policy-gradient based learning agents. We compare the performance of an augmented agent that uses our algorithm to provide additive intrinsic rewards to an A2C-based policy learner (for Atari games) and a PPO-based policy learner (for Mujoco domains) with a baseline agent that uses the same policy learners but with only extrinsic rewards. Our results show improved performance on most but not all of the domains.
Junhyuk Oh, Satinder Singh 0001
NeurIPS2
2017 Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
abstract
As a step towards developing zero-shot task generalization capabilities in reinforcement learning (RL), we introduce a new RL problem where the agent should learn to execute sequences of instructions after learning useful skills that solve subtasks. In this problem, we consider two types of generalizations: to previously unseen instructions and to longer sequences of instructions. For generalization over unseen instructions, we propose a new objective which encourages learning correspondences between similar subtasks by making analogies. For generalization over sequential instructions, we present a hierarchical architecture where a meta controller learns to use the acquired skills for executing the instructions. To deal with delayed reward, we propose a new neural architecture in the meta controller that learns when to update the subtask, which makes learning more efficient. Experimental results on a stochastic 3D domain show that the proposed ideas are crucial for generalization to longer instructions as well as unseen instructions.
Junhyuk Oh, Satinder Singh 0001, Honglak Lee, Pushmeet Kohli
ICML1
2017 Value Prediction Network
abstract
This paper proposes a novel deep reinforcement learning (RL) architecture, called Value Prediction Network (VPN), which integrates model-free and model-based RL methods into a single neural network. In contrast to typical model-based RL methods, VPN learns a dynamics model whose abstract states are trained to make option-conditional predictions of future values (discounted sum of rewards) rather than of future observations. Our experimental results show that VPN has several advantages over both model-free and model-based baselines in a stochastic environment where careful planning is required but building an accurate observation-prediction model is difficult. Furthermore, VPN outperforms Deep Q-Network (DQN) on several Atari games even with short-lookahead planning, demonstrating its potential as a new way of learning a good state representation.
Junhyuk Oh, Satinder Singh 0001, Honglak Lee
NIPS1
2016 Learning Transferrable Knowledge for Semantic Segmentation with Deep Convolutional Neural Network
abstract
We propose a novel weakly-supervised semantic segmentation algorithm based on Deep Convolutional Neural Network (DCNN). Contrary to existing weakly-supervised approaches, our algorithm exploits auxiliary segmentation annotations available for different categories to guide segmentations on images with only image-level class labels. To make segmentation knowledge transferrable across categories, we design a decoupled encoder-decoder architecture with attention model. In this architecture, the model generates spatial highlights of each category presented in images using an attention model, and subsequently performs binary segmentation for each highlighted region using decoder. Combining attention model, the decoder trained with segmentation annotations in different categories boosts accuracy of weakly-supervised semantic segmentation. The proposed algorithm demonstrates substantially improved performance compared to the state-of-theart weakly-supervised techniques in PASCAL VOC 2012 dataset when our model is trained with the annotations in 60 exclusive categories in Microsoft COCO dataset.
Seunghoon Hong, Junhyuk Oh, Honglak Lee, Bohyung Han
CVPR2
2016 Control of Memory, Active Perception, and Action in Minecraft
abstract
In this paper, we introduce a new set of reinforcement learning (RL) tasks in Minecraft (a flexible 3D world). We then use these tasks to systematically compare and contrast existing deep reinforcement learning (DRL) architectures with our new memory-based DRL architectures. These tasks are designed to emphasize, in a controllable manner, issues that pose challenges for RL methods including partial observability (due to first-person visual observations), delayed rewards, high-dimensional visual observations, and the need to use active perception in a correct manner so as to perform well in the tasks. While these tasks are conceptually simple to describe, by virtue of having all of these challenges simultaneously they are difficult for current DRL architectures. Additionally, we evaluate the generalization performance of the architectures on environments not used during training. The experimental results show that our new architectures generalize to unseen environments better than existing DRL architectures.
Junhyuk Oh, Valliappa Chockalingam, Satinder Singh 0001, Honglak Lee
ICML1
2015 Action-Conditional Video Prediction using Deep Networks in Atari Games
abstract
Motivated by vision-based reinforcement learning (RL) problems, in particular Atari games from the recent benchmark Aracade Learning Environment (ALE), we consider spatio-temporal prediction problems where future (image-)frames are dependent on control variables or actions as well as previous frames. While not composed of natural scenes, frames in Atari games are high-dimensional in size, can involve tens of objects with one or more objects being controlled by the actions directly and many other objects being influenced indirectly, can involve entry and departure of objects, and can involve deep partial observability. We propose and evaluate two deep neural network architectures that consist of encoding, action-conditional transformation, and decoding layers based on convolutional neural networks and recurrent neural networks. Experimental results show that the proposed architectures are able to generate visually-realistic frames that are also useful for control over approximately 100-step action-conditional futures in some games. To the best of our knowledge, this paper is the first to make and evaluate long-term predictions on high-dimensional video conditioned by control inputs.
Junhyuk Oh, Honglak Lee, Richard L. Lewis, Satinder Singh 0001
NIPS1