Philip J. Ball

dblp:244/1972 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 78% Learning paradigms · 7% Knowledge representation and reasoning · 3%

Topics — the 26 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
1.532022
Revisiting Design Choices in Offline Model Based Reinforcement Learning · ICLR 2022
Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment · ICML 2021
Ready Policy One: World Building Through Active Learning · ICML 2020
Machine learning › Reinforcement learning
sample efficiency
1.322023
Synthetic Experience Replay · NeurIPS 2023
Efficient Online Reinforcement Learning with Offline Data · ICML 2023
Machine learning › Reinforcement learning
off-policy reinforcement learning
1.222023
Efficient Online Reinforcement Learning with Offline Data · ICML 2023
Stabilizing Off-Policy Deep Reinforcement Learning from Pixels · ICML 2022
Machine learning › Reinforcement learning
exploration
1.232023
Learning General World Models in a Handful of Reward-Free Deployments · NeurIPS 2022
Ready Policy One: World Building Through Active Learning · ICML 2020
Efficient Online Reinforcement Learning with Offline Data · ICML 2023
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
1.122022
Learning General World Models in a Handful of Reward-Free Deployments · NeurIPS 2022
Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment · ICML 2021
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.712023
Synthetic Experience Replay · NeurIPS 2023
Machine learning › Reinforcement learning › online decision making
online reinforcement learning
0.712023
Efficient Online Reinforcement Learning with Offline Data · ICML 2023
Machine learning › Learning paradigms › data balancing
oversampling
0.712023
Synthetic Experience Replay · NeurIPS 2023
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.612022
Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning
0.612022
Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022
Machine learning › Learning theory
generalization
0.612022
Learning General World Models in a Handful of Reward-Free Deployments · NeurIPS 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
interference
0.612022
Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning
0.612022
Revisiting Design Choices in Offline Model Based Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning
multi-armed bandit
0.612022
Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
pixel-based reinforcement learning
0.612022
Stabilizing Off-Policy Deep Reinforcement Learning from Pixels · ICML 2022
Machine learning › Reinforcement learning
policy selection
0.612022
Same State, Different Task: Continual Reinforcement Learning without Interference · AAAI 2022
Machine learning › Reinforcement learning › exploration › exploration in markov decision processes
reward-free exploration
0.612022
Learning General World Models in a Handful of Reward-Free Deployments · NeurIPS 2022
Machine learning › Reinforcement learning › exploration › intrinsic motivation
self-supervised exploration
0.612022
Learning General World Models in a Handful of Reward-Free Deployments · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot task generalization
0.612022
Learning General World Models in a Handful of Reward-Free Deployments · NeurIPS 2022
Machine learning › Reinforcement learning › generalization in reinforcement learning
dynamics generalization
0.512021
Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment · ICML 2021
Machine learning › Reinforcement learning
offline reinforcement learning
0.512021
Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment · ICML 2021
Machine learning › Efficient and distributed learning
active learning
0.412020
Ready Policy One: World Building Through Active Learning · ICML 2020
Robotics › Motion planning and robot control
robot control
0.322021
Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment · ICML 2021
Ready Policy One: World Building Through Active Learning · ICML 2020
Machine learning › Generative modeling
diffusion model
0.212023
Synthetic Experience Replay · NeurIPS 2023
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
0.112021
Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment · ICML 2021
Machine learning › Reinforcement learning
continuous control
0.112020
Ready Policy One: World Building Through Active Learning · ICML 2020

Methods — techniques the papers use, named apart from their topics

off-policy learning · 0.7diffusion model · 0.7ablation study · 0.7temporal difference learning · 0.6replay buffer · 0.6multi-armed bandit · 0.6factorized policy · 0.6convolutional encoder · 0.6bayesian active learning · 0.6adaptive regularization · 0.6
YearPublicationVenuePosition
2023 Efficient Online Reinforcement Learning with Offline Data
abstract
Sample efficiency and exploration remain major challenges in online reinforcement learning (RL). A powerful approach that can be applied to address these issues is the inclusion of offline data, such as prior trajectories from a human expert or a sub-optimal exploration policy. Previous methods have relied on extensive modifications and additional complexity to ensure the effective use of this data. Instead, we ask: *can we simply apply existing off-policy methods to leverage offline data when learning online?* In this work, we demonstrate that the answer is yes; however, a set of minimal but important changes to existing off-policy RL algorithms are required to achieve reliable performance. We extensively ablate these design choices, demonstrating the key factors that most affect performance, and arrive at a set of recommendations that practitioners can readily apply, whether their data comprise a small number of expert demonstrations or large volumes of sub-optimal trajectories. We see that correct application of these simple recommendations can provide a $\mathbf{2.5\times}$ improvement over existing approaches across a diverse set of competitive benchmarks, with no additional computational overhead.
Philip J. Ball, Laura Smith 0001, Ilya Kostrikov, Sergey Levine
ICML1
2023 Synthetic Experience Replay
abstract
A key theme in the past decade has been that when large neural networks and large datasets combine they can produce remarkable results. In deep reinforcement learning (RL), this paradigm is commonly made possible through experience replay, whereby a dataset of past experiences is used to train a policy or value function. However, unlike in supervised or self-supervised learning, an RL agent has to collect its own data, which is often limited. Thus, it is challenging to reap the benefits of deep learning, and even small neural networks can overfit at the start of training. In this work, we leverage the tremendous recent progress in generative modeling and propose Synthetic Experience Replay (SynthER), a diffusion-based approach to flexibly upsample an agent's collected experience. We show that SynthER is an effective method for training RL agents across offline and online settings, in both proprioceptive and pixel-based environments. In offline settings, we observe drastic improvements when upsampling small offline datasets and see that additional synthetic data also allows us to effectively train larger networks. Furthermore, SynthER enables online agents to train with a much higher update-to-data ratio than before, leading to a significant increase in sample efficiency, without any algorithmic changes. We believe that synthetic training data could open the door to realizing the full potential of deep learning for replay-based RL algorithms from limited data. Finally, we open-source our code at https://github.com/conglu1997/SynthER.
Cong Lu, Philip J. Ball, Yee Whye Teh, Jack Parker-Holder
NeurIPS2
2022 Same State, Different Task: Continual Reinforcement Learning without Interference
abstract
Continual Learning (CL) considers the problem of training an agent sequentially on a set of tasks while seeking to retain performance on all previous tasks. A key challenge in CL is catastrophic forgetting, which arises when performance on a previously mastered task is reduced when learning a new task. While a variety of methods exist to combat forgetting, in some cases tasks are fundamentally incompatible with each other and thus cannot be learnt by a single policy. This can occur, in reinforcement learning (RL) when an agent may be rewarded for achieving different goals from the same observation. In this paper we formalize this "interference" as distinct from the problem of forgetting. We show that existing CL methods based on single neural network predictors with shared replay buffers fail in the presence of interference. Instead, we propose a simple method, OWL, to address this challenge. OWL learns a factorized policy, using shared feature extraction layers, but separate heads, each specializing on a new task. The separate heads in OWL are used to prevent interference. At test time, we formulate policy selection as a multi-armed bandit problem, and show it is possible to select the best policy for an unknown task using feedback from the environment. The use of bandit algorithms allows the OWL agent to constructively re-use different continually learnt policies at different times during an episode. We show in multiple RL environments that existing replay based CL methods fail, while OWL is able to achieve close to optimal performance when training sequentially.
Samuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren, Stephen J. Roberts
AAAI3
2022 Revisiting Design Choices in Offline Model Based Reinforcement Learning
Cong Lu, Philip J. Ball, Jack Parker-Holder, Michael A. Osborne, Stephen J. Roberts
ICLR2
2022 Stabilizing Off-Policy Deep Reinforcement Learning from Pixels
abstract
Off-policy reinforcement learning (RL) from pixel observations is notoriously unstable. As a result, many successful algorithms must combine different domain-specific practices and auxiliary losses to learn meaningful behaviors in complex environments. In this work, we provide novel analysis demonstrating that these instabilities arise from performing temporal-difference learning with a convolutional encoder and low-magnitude rewards. We show that this new visual deadly triad causes unstable training and premature convergence to degenerate solutions, a phenomenon we name catastrophic self-overfitting. Based on our analysis, we propose A-LIX, a method providing adaptive regularization to the encoder’s gradients that explicitly prevents the occurrence of catastrophic self-overfitting using a dual objective. By applying A-LIX, we significantly outperform the prior state-of-the-art on the DeepMind Control and Atari benchmarks without any data augmentation or auxiliary losses.
Edoardo Cetin, Philip J. Ball, Stephen J. Roberts, Oya Çeliktutan
ICML2
2022 Learning General World Models in a Handful of Reward-Free Deployments
abstract
Building generally capable agents is a grand challenge for deep reinforcement learning (RL). To approach this challenge practically, we outline two key desiderata: 1) to facilitate generalization, exploration should be task agnostic; 2) to facilitate scalability, exploration policies should collect large quantities of data without costly centralized retraining. Combining these two properties, we introduce the reward-free deployment efficiency setting, a new paradigm for RL research. We then present CASCADE, a novel approach for self-supervised exploration in this new setting. CASCADE seeks to learn a world model by collecting data with a population of agents, using an information theoretic objective inspired by Bayesian Active Learning. CASCADE achieves this by specifically maximizing the diversity of trajectories sampled by the population through a novel cascading objective. We provide theoretical intuition for CASCADE which we show in a tabular setting improves upon naïve approaches that do not account for population diversity. We then demonstrate that CASCADE collects diverse task-agnostic datasets and learns agents that generalize zero-shot to novel, unseen downstream tasks on Atari, MiniGrid, Crafter and the DM Control Suite. Code and videos are available at https://ycxuyingchen.github.io/cascade/
Yingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip J. Ball, Oleh Rybkin, Stephen J. Roberts, Tim Rocktäschel, Edward Grefenstette
NeurIPS4
2021 Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment
abstract
Reinforcement learning from large-scale offline datasets provides us with the ability to learn policies without potentially unsafe or impractical exploration. Significant progress has been made in the past few years in dealing with the challenge of correcting for differing behavior between the data collection and learned policies. However, little attention has been paid to potentially changing dynamics when transferring a policy to the online setting, where performance can be up to 90% reduced for existing methods. In this paper we address this problem with Augmented World Models (AugWM). We augment a learned dynamics model with simple transformations that seek to capture potential changes in physical properties of the robot, leading to more robust policies. We not only train our policy in this new setting, but also provide it with the sampled augmentation as a context, allowing it to adapt to changes in the environment. At test time we learn the context in a self-supervised fashion by approximating the augmentation which corresponds to the new environment. We rigorously evaluate our approach on over 100 different changed dynamics settings, and show that this simple approach can significantly improve the zero-shot generalization of a recent state-of-the-art baseline, often achieving successful policies where the baseline fails.
Philip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. Roberts
ICML1
2021 Towards tractable optimism in model-based reinforcement learning
abstract
The principle of optimism in the face of uncertainty is prevalent throughout sequential decision making problems such as multi-armed bandits and reinforcement learning (RL). To be successful, an optimistic RL algorithm must over-estimate the true value function (optimism) but not by so much that it is inaccurate (estimation error). In the tabular setting, many state-of-the-art methods produce the required optimism through approaches which are intractable when scaling to deep RL. We re-interpret these scalable optimistic model-based algorithms as solving a tractable noise augmented MDP. This formulation achieves a competitive regret bound: $\tilde{\mathcal{O}}( |\mathcal{S}|H\sqrt{|\mathcal{A}| T } )$ when augmenting using Gaussian noise, where $T$ is the total number of environment steps. We also explore how this trade-off changes in the deep RL setting, where we show empirically that estimation error is significantly more troublesome. However, we also show that if this error is reduced, optimistic model-based RL algorithms can match state-of-the-art performance in continuous control problems.
Aldo Pacchiano, Philip J. Ball, Jack Parker-Holder, Krzysztof Choromanski, Stephen J. Roberts
UAI2
2021 Active Inference: Demystified and Compared
abstract
Active inference is a first principle account of how autonomous agents operate in dynamic, nonstationary environments. This problem is also considered in reinforcement learning, but limited work exists on comparing the two approaches on the same discrete-state environments. In this letter, we provide (1) an accessible overview of the discrete-state formulation of active inference, highlighting natural behaviors in active inference that are generally engineered in reinforcement learning, and (2) an explicit discrete-state comparison between active inference and reinforcement learning on an OpenAI gym baseline. We begin by providing a condensed overview of the active inference literature, in particular viewing the various natural behaviors of active inference agents through the lens of reinforcement learning. We show that by operating in a pure belief-based setting, active inference agents can carry out epistemic exploration-and account for uncertainty about their environment-in a Bayes-optimal fashion. Furthermore, we show that the reliance on an explicit reward signal in reinforcement learning is removed in active inference, where reward can simply be treated as another observation we have a preference over; even in the total absence of rewards, agent behaviors are learned through preference learning. We make these properties explicit by showing two scenarios in which active inference agents can infer behaviors in reward-free environments compared to both Q-learning and Bayesian model-based reinforcement learning agents and by placing zero prior preferences over rewards and learning the prior preferences over the observations corresponding to reward. We conclude by noting that this formalism can be applied to more complex settings (e.g., robotic arm movement, Atari games) if appropriate generative models can be formulated. In short, we aim to demystify the behavior of active inference agents by presenting an accessible discrete state-space and time formulation and demonstrate these behaviors in a OpenAI gym environment, alongside reinforcement learning agents.
Noor Sajid, Philip J. Ball, Thomas Parr, Karl J. Friston
Neural Comput.2
2020 Ready Policy One: World Building Through Active Learning
abstract
Model-Based Reinforcement Learning (MBRL) offers a promising direction for sample efficient learning, often achieving state of the art results for continuous control tasks. However many existing MBRL methods rely on combining greedy policies with exploration heuristics, and even those which utilize principled exploration bonuses construct dual objectives in an ad hoc fashion. In this paper we introduce Ready Policy One (RP1), a framework that views MBRL as an active learning problem, where we aim to improve the world model in the fewest samples possible. RP1 achieves this by utilizing a hybrid objective function, which crucially adapts during optimization, allowing the algorithm to trade off reward v.s. exploration at different stages of learning. In addition, we introduce a principled mechanism to terminate sample collection once we have a rich enough trajectory batch to improve the model. We rigorously evaluate our method on a variety of continuous control tasks, and demonstrate statistically significant gains over existing approaches.
Philip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski, Stephen J. Roberts
ICML1
2019 The Sensitivity of Counterfactual Fairness to Unmeasured Confounding
Niki Kilbertus, Philip J. Ball, Matt J. Kusner, Adrian Weller, Ricardo Silva 0001
UAI2