Bradly C. Stadie

dblp:166/1368 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 56% Efficient and distributed learning · 13% Planning, search and constraint satisfaction · 10%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning
0.912025
Of Mice and Machines: A Comparison of Learning Between Real World Mice and RL Agents · ICML 2025
Machine learning › Transfer learning and domain adaptation
meta-learning
0.622018
Evolved Policy Gradients · NeurIPS 2018
One-Shot Imitation Learning · NIPS 2017
Machine learning › Reinforcement learning
imitation learning
0.622017
One-Shot Imitation Learning · NIPS 2017
Third Person Imitation Learning · ICLR (Poster) 2017
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
graph search
0.512021
World Model as a Graph: Learning Latent Landmarks for Planning · ICML 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning
0.512021
World Model as a Graph: Learning Latent Landmarks for Planning · ICML 2021
Machine learning › Reinforcement learning
model-based reinforcement learning
0.512021
World Model as a Graph: Learning Latent Landmarks for Planning · ICML 2021
Machine learning › Reinforcement learning › exploration › information-theoretic exploration
entropy-based exploration
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Reinforcement learning
exploration
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Reinforcement learning
long-horizon tasks
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Efficient and distributed learning
model compression
0.412020
One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation · ICLR 2020
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
multi-goal reinforcement learning
0.412020
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning · ICML 2020
Machine learning › Efficient and distributed learning › model compression › pruning › DNN pruning
one-shot pruning
0.412020
One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation · ICLR 2020
Machine learning › Efficient and distributed learning › model compression
pruning
0.412020
One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation · ICLR 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.412020
One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation · ICLR 2020
Machine learning › Reinforcement learning
meta-reinforcement learning
0.312018
The Importance of Sampling inMeta-Reinforcement Learning · NeurIPS 2018
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.312018
Evolved Policy Gradients · NeurIPS 2018
Machine learning › Learning theory
sampling distribution
0.312018
The Importance of Sampling inMeta-Reinforcement Learning · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.312017
One-Shot Imitation Learning · NIPS 2017
Machine learning › Reinforcement learning › imitation learning › few-shot imitation learning
one-shot imitation learning
0.312017
One-Shot Imitation Learning · NIPS 2017
Robotics › Motion planning and robot control
robot learning
0.312017
Third Person Imitation Learning · ICLR (Poster) 2017
Machine learning › Reinforcement learning
deep reinforcement learning
0.112021
World Model as a Graph: Learning Latent Landmarks for Planning · ICML 2021
Machine learning › Deep learning architectures and training › loss function design
loss function learning
0.112018
Evolved Policy Gradients · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

reward shaping · 0.9comparative study · 0.9q-function distillation · 0.5latent landmark learning · 0.5graph-structured world model · 0.5maximum entropy · 0.4jacobian spectrum evaluation · 0.4temporal convolution · 0.3evolutionary algorithm · 0.3MAML · 0.3
YearPublicationVenuePosition
2025 Of Mice and Machines: A Comparison of Learning Between Real World Mice and RL Agents
abstract
Recent advances in reinforcement learning (RL) have demonstrated impressive capabilities in complex decision-making tasks. This progress raises a natural question: how do these artificial systems compare to biological agents, which have been shaped by millions of years of evolution? To help answer this question, we undertake a comparative study of biological mice and RL agents in a predator-avoidance maze environment. Through this analysis, we identify a striking disparity: RL agents consistently demonstrate a lack of self-preservation instinct, readily risking ``death'' for marginal efficiency gains. These risk-taking strategies are in contrast to biological agents, which exhibit sophisticated risk-assessment and avoidance behaviors. Towards bridging this gap between the biological and artificial, we propose two novel mechanisms that encourage more naturalistic risk-avoidance behaviors in RL agents. Our approach leads to the emergence of naturalistic behaviors, including strategic environment assessment, cautious path planning, and predator avoidance patterns that closely mirror those observed in biological systems.
German Espinosa, Junda Huang, Daniel A. Dombeck, Malcolm A. MacIver, Bradly C. Stadie
ICML6
2021 World Model as a Graph: Learning Latent Landmarks for Planning
abstract
Planning, the ability to analyze the structure of a problem in the large and decompose it into interrelated subproblems, is a hallmark of human intelligence. While deep reinforcement learning (RL) has shown great promise for solving relatively straightforward control tasks, it remains an open problem how to best incorporate planning into existing deep RL paradigms to handle increasingly complex environments. One prominent framework, Model-Based RL, learns a world model and plans using step-by-step virtual rollouts. This type of world model quickly diverges from reality when the planning horizon increases, thus struggling at long-horizon planning. How can we learn world models that endow agents with the ability to do temporally extended reasoning? In this work, we propose to learn graph-structured world models composed of sparse, multi-step transitions. We devise a novel algorithm to learn latent landmarks that are scattered (in terms of reachability) across the goal space as the nodes on the graph. In this same graph, the edges are the reachability estimates distilled from Q-functions. On a variety of high-dimensional continuous control tasks ranging from robotic manipulation to navigation, we demonstrate that our method, named L3P, significantly outperforms prior work, and is oftentimes the only method capable of leveraging both the robustness of model-free RL and generalization of graph-search algorithms. We believe our work is an important step towards scalable planning in reinforcement learning.
Lunjun Zhang, Ge Yang 0003, Bradly C. Stadie
ICML3
2020 One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation
Matthew Shunshi Zhang, Bradly C. Stadie
ICLR2
2020 Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
abstract
What goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the agent should not pursue unobtainable goals. Instead, it should set its own intrinsic goals that maximize the entropy of the historical achieved goal distribution. We propose to optimize this objective by having the agent pursue past achieved goals in sparsely explored areas of the goal space, which focuses exploration on the frontier of the achievable goal set. We show that our strategy achieves an order of magnitude better sample efficiency than the prior state of the art on long-horizon multi-goal tasks including maze navigation and block stacking.
Silviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie, Jimmy Ba
ICML4
2020 Learning Intrinsic Rewards as a Bi-Level Optimization Problem
abstract
We reinterpret the problem of finding intrinsic rewards in reinforcement learning (RL) as a bilevel optimization problem. Using this interpretation, we can make use of recent advancements in the hyperparameter optimization literature, mainly from Self-Tuning Networks (STN), to learn intrinsic rewards. To facilitate our methods, we introduces a new general conditioning layer: Conditional Layer Normalization (CLN). We evaluate our method on several continuous control benchmarks in the Mujoco physics simulator. On all of these benchmarks, the intrinsic rewards learned on the fly lead to higher final rewards.
Bradly C. Stadie, Lunjun Zhang, Jimmy Ba
UAI1
2018 Evolved Policy Gradients
abstract
We propose a metalearning approach for learning gradient-based reinforcement learning (RL) algorithms. The idea is to evolve a differentiable loss function, such that an agent, which optimizes its policy to minimize this loss, will achieve high rewards. The loss is parametrized via temporal convolutions over the agent's experience. Because this loss is highly flexible in its ability to take into account the agent's history, it enables fast task learning. Empirical results show that our evolved policy gradient algorithm (EPG) achieves faster learning on several randomized environments compared to an off-the-shelf policy gradient method. We also demonstrate that EPG's learned loss can generalize to out-of-distribution test time tasks, and exhibits qualitatively different behavior from other popular metalearning algorithms.
Rein Houthooft, Phillip Isola, Bradly C. Stadie, Filip Wolski, Jonathan Ho, Pieter Abbeel
NeurIPS4
2018 The Importance of Sampling inMeta-Reinforcement Learning
abstract
We interpret meta-reinforcement learning as the problem of learning how to quickly find a good sampling distribution in a new environment. This interpretation leads to the development of two new meta-reinforcement learning algorithms: E-MAML and E-$\text{RL}^2$. Results are presented on a new environment we call `Krazy World': a difficult high-dimensional gridworld which is designed to highlight the importance of correctly differentiating through sampling distributions in meta-reinforcement learning. Further results are presented on a set of maze environments. We show E-MAML and E-$\text{RL}^2$ deliver better performance than baseline algorithms on both tasks.
Bradly C. Stadie, Ge Yang 0003, Rein Houthooft, Xi Chen 0022, Yan Duan, Yuhuai Wu, Pieter Abbeel, Ilya Sutskever
NeurIPS1
2017 Third Person Imitation Learning
Bradly C. Stadie, Pieter Abbeel, Ilya Sutskever
ICLR (Poster)1
2017 One-Shot Imitation Learning
abstract
Imitation learning has been commonly applied to solve different tasks in isolation. This usually requires either careful feature engineering, or a significant number of samples. This is far from what we desire: ideally, robots should be able to learn from very few demonstrations of any given task, and instantly generalize to new situations of the same task, without requiring task-specific engineering. In this paper, we propose a meta-learning framework for achieving such capability, which we call one-shot imitation learning. Specifically, we consider the setting where there is a very large (maybe infinite) set of tasks, and each task has many instantiations. For example, a task could be to stack all blocks on a table into a single tower, another task could be to place all blocks on a table into two-block towers, etc. In each case, different instances of the task would consist of different sets of blocks with different initial states. At training time, our algorithm is presented with pairs of demonstrations for a subset of all tasks. A neural net is trained that takes as input one demonstration and the current state (which initially is the initial state of the other demonstration of the pair), and outputs an action with the goal that the resulting sequence of states and actions matches as closely as possible with the second demonstration. At test time, a demonstration of a single instance of a new task is presented, and the neural net is expected to perform well on new instances of this new task. Our experiments show that the use of soft attention allows the model to generalize to conditions and tasks unseen in the training data. We anticipate that by training this model on a much greater variety of tasks and settings, we will obtain a general system that can turn any demonstrations into robust policies that can accomplish an overwhelming variety of tasks.
Yan Duan, Marcin Andrychowicz, Bradly C. Stadie, Jonathan Ho, Jonas Schneider 0002, Ilya Sutskever, Pieter Abbeel, Wojciech Zaremba
NIPS3