Jacob Hilton

dblp:182/7972 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0002-2002-5929ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Estimating the Probabilities of Rare Outputs in Language Models
abstract
We consider the problem of *low probability estimation*: given a machine learning model and a formally-specified input distribution, how can we estimate the probability of a binary property of the model's output, even when that probability is too small to estimate by random sampling? This problem is motivated by the need to improve worst-case performance, which distribution shift can make much more likely. We study low probability estimation in the context of argmax sampling from small transformer language models. We compare two types of methods: importance sampling, which involves searching for inputs giving rise to the rare output, and activation extrapolation, which involves extrapolating a probability distribution fit to the model's logits. We find that importance sampling outperforms activation extrapolation, but both outperform naive sampling. Finally, we explain how minimizing the probability estimate of an undesirable behavior generalizes adversarial training, and argue that new methods for low probability estimation are needed to provide stronger guarantees about worst-case performance.
Gabriel Wu, Jacob Hilton
ICLR2
2025 Backdoor Defense, Learnability and Obfuscation
abstract
We introduce a formal notion of defendability against backdoors using a game between an attacker and a defender. In this game, the attacker modifies a function to behave differently on a particular input known as the "trigger", while behaving the same almost everywhere else. The defender then attempts to detect the trigger at evaluation time. If the defender succeeds with high enough probability, then the function class is said to be defendable. The key constraint on the attacker that makes defense possible is that the attacker's strategy must work for a randomly-chosen trigger. Our definition is simple and does not explicitly mention learning, yet we demonstrate that it is closely connected to learnability. In the computationally unbounded setting, we use a voting algorithm of Hanneke et al. (2022) to show that defendability is essentially determined by the VC dimension of the function class, in much the same way as PAC learnability. In the computationally bounded setting, we use a similar argument to show that efficient PAC learnability implies efficient defendability, but not conversely. On the other hand, we use indistinguishability obfuscation to show that the class of polynomial size circuits is not efficiently defendable. Finally, we present polynomial size decision trees as a natural example for which defense is strictly easier than learning. Thus, we identify efficient defendability as a notable intermediate concept in between efficient learnability and obfuscation.
Paul F. Christiano, Jacob Hilton, Victor Lecomte, Xianzhong Mark Xu
ITCS2
2023 Scaling Laws for Reward Model Overoptimization
abstract
In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy, optimizing its value too much can hinder ground truth performance, in accordance with Goodhart’s law. This effect has been frequently observed, but not carefully measured due to the expense of collecting human preference data. In this work, we use a synthetic setup in which a fixed “gold-standard” reward model plays the role of humans, providing labels used to train a proxy reward model. We study how the gold reward model score changes as we optimize against the proxy reward model using either reinforcement learning or best-of-$n$ sampling. We find that this relationship follows a different functional form depending on the method of optimization, and that in both cases its coefficients scale smoothly with the number of reward model parameters. We also study the effect on this relationship of the size of the reward model dataset, the number of reward model and policy parameters, and the coefficient of the KL penalty added to the reward in the reinforcement learning setup. We explore the implications of these empirical results for theoretical considerations in AI alignment.
Leo Gao, John Schulman, Jacob Hilton
ICML3
2022 TruthfulQA: Measuring How Models Mimic Human Falsehoods
abstract
We propose a benchmark to measure whether a language model is truthful in generating answers to questions.The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics.We crafted questions that some humans would answer falsely due to a false belief or misconception.To perform well, models must avoid generating false answers learned from imitating human texts.We tested GPT-3, GPT-Neo/J, GPT-2 and a T5-based model.The best model was truthful on 58% of questions, while human performance was 94%.Models generated many false answers that mimic popular misconceptions and have the potential to deceive humans.The largest models were generally the least truthful.This contrasts with other NLP tasks, where performance improves with model size.However, this result is expected if false answers are learned from the training distribution.We suggest that scaling up models alone is less promising for improving truthfulness than finetuning using training objectives other than imitation of text from the web."The enemy of truth is blind acceptance.
Stephanie Lin, Jacob Hilton, Owain Evans
ACL (1)2
2022 Batch size-invariance for policy optimization
abstract
We say an algorithm is batch size-invariant if changes to the batch size can largely be compensated for by changes to other hyperparameters. Stochastic gradient descent is well-known to have this property at small batch sizes, via the learning rate. However, some policy optimization algorithms (such as PPO) do not have this property, because of how they control the size of policy updates. In this work we show how to make these algorithms batch size-invariant. Our key insight is to decouple the proximal policy (used for controlling policy updates) from the behavior policy (used for off-policy corrections). Our experiments help explain why these algorithms work, and additionally show how they can make more efficient use of stale data.
Jacob Hilton, Karl Cobbe, John Schulman
NeurIPS1
2022 Training language models to follow instructions with human feedback
abstract
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through a language model API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.
Long Ouyang, Jeff Wu 0003, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, Ryan Lowe
NeurIPS12
2021 Phasic Policy Gradient
abstract
We introduce Phasic Policy Gradient (PPG), a reinforcement learning framework which modifies traditional on-policy actor-critic methods by separating policy and value function training into distinct phases. In prior methods, one must choose between using a shared network or separate networks to represent the policy and value function. Using separate networks avoids interference between objectives, while using a shared network allows useful features to be shared. PPG is able to achieve the best of both worlds by splitting optimization into two phases, one that advances training and one that distills features. PPG also enables the value function to be more aggressively optimized with a higher level of sample reuse. Compared to PPO, we find that PPG significantly improves sample efficiency on the challenging Procgen Benchmark.
Karl Cobbe, Jacob Hilton, Oleg Klimov, John Schulman
ICML2
2020 Leveraging Procedural Generation to Benchmark Reinforcement Learning
abstract
We introduce Procgen Benchmark, a suite of 16 procedurally generated game-like environments designed to benchmark both sample efficiency and generalization in reinforcement learning. We believe that the community will benefit from increased access to high quality training environments, and we provide detailed experimental protocols for using this benchmark. We empirically demonstrate that diverse environment distributions are essential to adequately train and evaluate RL agents, thereby motivating the extensive use of procedural content generation. We then use this benchmark to investigate the effects of scaling model size, finding that larger models significantly improve both sample efficiency and generalization.
Karl Cobbe, Christopher Hesse, Jacob Hilton, John Schulman
ICML3
2016 The Topological Pigeonhole Principle for Ordinals
abstract
Abstract Given a cardinal κ and a sequence ${\left( {{\alpha _i}} \right)_{i \in \kappa }}$ of ordinals, we determine the least ordinal β (when one exists) such that the topological partition relation $$\beta \to \left( {top\,{\alpha _i}} \right)_{i \in \kappa }^1$$ holds, including an independence result for one class of cases. Here the prefix “top” means that the homogeneous set must be of the correct homeomorphism class rather than the correct order type. The answer is linked to the nontopological pigeonhole principle of Milner and Rado.
Jacob Hilton
J. Symb. Log.1