Jack W. Rae

dblp:188/5991 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Deep learning architectures and training · 40% Efficient and distributed learning · 16% Language models and text generation · 12%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 84% Indexing and storage engines · 16%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
scaling laws
1.122022
An empirical analysis of compute-optimal large language model training · NeurIPS 2022
Unified Scaling Laws for Routed Language Models · ICML 2022
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.932018
Relational recurrent neural networks · NeurIPS 2018
Fast Parametric Learning with Activation Memorization · ICML 2018
Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes · NIPS 2016
Machine learning › Deep learning architectures and training
transformer
0.922020
Stabilizing Transformers for Reinforcement Learning · ICML 2020
Compressive Transformers for Long-Range Sequence Modelling · ICLR 2020
Machine learning › Transfer learning and domain adaptation
meta-learning
0.822020
Meta-Learning Deep Energy-Based Memory Models · ICLR 2020
Meta-Learning Neural Bloom Filters · ICML 2019
Machine learning › Deep learning architectures and training
attention mechanism
0.822020
Do Transformers Need Deep Long-Range Memory? · ACL 2020
Relational recurrent neural networks · NeurIPS 2018
Machine learning › Efficient and distributed learning
model compression
0.722020
Top-KAST: Top-K Always Sparse Training · NeurIPS 2020
Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes · NIPS 2016
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.612022
An empirical analysis of compute-optimal large language model training · NeurIPS 2022
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.612022
Unified Scaling Laws for Routed Language Models · ICML 2022
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Information retrieval
document retrieval
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Reinforcement learning
deep reinforcement learning
0.412020
Stabilizing Transformers for Reinforcement Learning · ICML 2020
Machine learning › Efficient and distributed learning › model compression › sparsity
dynamic sparsity
0.412020
Top-KAST: Top-K Always Sparse Training · NeurIPS 2020
Machine learning › Generative modeling
energy-based model
0.412020
Meta-Learning Deep Energy-Based Memory Models · ICLR 2020
Machine learning › Deep learning architectures and training › sequence modeling
long-range memory
0.412020
Do Transformers Need Deep Long-Range Memory? · ACL 2020
Machine learning › Deep learning architectures and training › sequence modeling
long sequence modeling
0.412020
Compressive Transformers for Long-Range Sequence Modelling · ICLR 2020
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Deep learning architectures and training › feature interaction
multiplicative interaction
0.412020
Multiplicative Interactions and Where to Find Them · ICLR 2020
Machine learning › Reinforcement learning › policy optimization
on-policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Reinforcement learning
policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Efficient and distributed learning › model compression
sparse training
0.412020
Top-KAST: Top-K Always Sparse Training · NeurIPS 2020
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model
0.412020
Do Transformers Need Deep Long-Range Memory? · ACL 2020
Natural language and speech › Language models and text generation
language modeling
0.422018
Fast Parametric Learning with Activation Memorization · ICML 2018
Relational recurrent neural networks · NeurIPS 2018
Natural language and speech › Language models and text generation › text generation › synthetic text generation
adversarial text generation
0.412019
Training Language GANs from Scratch · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412019
Meta-Learning Neural Bloom Filters · ICML 2019
Machine learning › Generative modeling
generative adversarial network
0.412019
Training Language GANs from Scratch · NeurIPS 2019
Natural language and speech › Language models and text generation › text generation › neural text generation
language GAN
0.412019
Training Language GANs from Scratch · NeurIPS 2019
Machine learning › Deep learning architectures and training › memory-augmented neural networks
neural data structure
0.412019
Meta-Learning Neural Bloom Filters · ICML 2019
Computer vision › Image recognition and object detection › object detection
training from scratch
0.412019
Training Language GANs from Scratch · NeurIPS 2019
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention
0.312018
Relational recurrent neural networks · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

differentiable encoder · 1.1chunked cross-attention · 1.1scaling law estimation · 0.6power-law scaling · 0.6effective parameter count · 0.6transformer · 0.4multiplicative interactions · 0.4intervention analysis · 0.4expectation-maximization · 0.4attention · 0.4meta-learning · 0.4memory-augmented neural network · 0.4
YearPublicationVenuePosition
2022 Improving Language Models by Retrieving from Trillions of Tokens
abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre
ICML26
2022 Unified Scaling Laws for Routed Language Models
abstract
The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters.
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan
ICML23
2022 An empirical analysis of compute-optimal large language model training
abstract
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more data. Chinchilla uniformly and significantly outperformsGopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, a 7% improvement over Gopher.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche 0002, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, Laurent Sifre
NeurIPS21
2020 Do Transformers Need Deep Long-Range Memory?
abstract
Deep attention models have advanced the modelling of sequential data across many domains.For language modelling in particular, the Transformer-XL -a Transformer augmented with a long-range memory of past activations -has been shown to be state-ofthe-art across a variety of well-studied benchmarks.The Transformer-XL incorporates a long-range memory at every layer of the network, which renders its state to be thousands of times larger than RNN predecessors.However it is unclear whether this is necessary.We perform a set of interventions to show that comparable performance can be obtained with 6X fewer long range memories and better performance can be obtained by limiting the range of attention in lower layers of the network.
Jack W. Rae, Ali Razavi
ACL1
2020 Meta-Learning Deep Energy-Based Memory Models
Sergey Bartunov, Jack W. Rae, Simon Osindero, Timothy P. Lillicrap
ICLR2
2020 Multiplicative Interactions and Where to Find Them
Siddhant M. Jayakumar, Wojciech Czarnecki 0001, Jacob Menick, Jonathan Schwarz, Jack W. Rae, Simon Osindero, Yee Whye Teh, Tim Harley, Razvan Pascanu
ICLR5
2020 Compressive Transformers for Long-Range Sequence Modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, Timothy P. Lillicrap
ICLR1
2020 V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick
ICLR6
2020 Stabilizing Transformers for Reinforcement Learning
abstract
Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown breakthrough success in natural language processing (NLP). Harnessing the transformer’s ability to process long time horizons of information could provide a similar performance boost in partially observable reinforcement learning (RL) domains, but the large-scale transformers used in NLP have yet to be successfully applied to the RL setting. In this work we demonstrate that the standard transformer architecture is difficult to optimize, which was previously observed in the supervised learning setting but becomes especially pronounced with RL objectives. We propose architectural modifications that substantially improve the stability and learning speed of the original Transformer and XL variant. The proposed architecture, the Gated Transformer-XL (GTrXL), surpasses LSTMs on challenging memory environments and achieves state-of-the-art results on the multi-task DMLab-30 benchmark suite, exceeding the performance of an external memory architecture. We show that the GTrXL has stability and performance that consistently matches or exceeds a competitive LSTM baseline, including on more reactive tasks where memory is less critical.
Emilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant M. Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matt M. Botvinick, Nicolas Heess, Raia Hadsell
ICML3
2020 Top-KAST: Top-K Always Sparse Training
abstract
Sparse neural networks are becoming increasingly important as the field seeks to improve the performance of existing models by scaling them up, while simultaneously trying to reduce power consumption and computational footprint. Unfortunately, most existing methods for inducing performant sparse models still entail the instantiation of dense parameters, or dense gradients in the backward-pass, during training. For very large models this requirement can be prohibitive. In this work we propose Top-KAST, a method that preserves constant sparsity throughout training (in both the forward and backward-passes). We demonstrate the efficacy of our approach by showing that it performs comparably to or better than previous works when training models on the established ImageNet benchmark, whilst fully maintaining sparsity. In addition to our ImageNet results, we also demonstrate our approach in the domain of language modeling where the current best performing architectures tend to have tens of billions of parameters and scaling up does not yet seem to have saturated performance. Sparse versions of these architectures can be run with significantly fewer resources, making them more widely accessible and applicable. Furthermore, in addition to being effective, our approach is straightforward and can easily be implemented in a wide range of existing machine learning frameworks with only a few additional lines of code. We therefore hope that our contribution will help enable the broader community to explore the potential held by massive models, without incurring massive computational cost.
Siddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero, Erich Elsen
NeurIPS3
2019 Meta-Learning Neural Bloom Filters
abstract
There has been a recent trend in training neural networks to replace data structures that have been crafted by hand, with an aim for faster execution, better accuracy, or greater compression. In this setting, a neural data structure is instantiated by training a network over many epochs of its inputs until convergence. In applications where inputs arrive at high throughput, or are ephemeral, training a network from scratch is not practical. This motivates the need for few-shot neural data structures. In this paper we explore the learning of approximate set membership over a set of data in one-shot via meta-learning. We propose a novel memory architecture, the Neural Bloom Filter, which is able to achieve significant compression gains over classical Bloom Filters and existing memory-augmented neural networks.
Jack W. Rae, Sergey Bartunov, Timothy P. Lillicrap
ICML1
2019 Training Language GANs from Scratch
abstract
Generative Adversarial Networks (GANs) enjoy great success at image generation, but have proven difficult to train in the domain of natural language. Challenges with gradient estimation, optimization instability, and mode collapse have lead practitioners to resort to maximum likelihood pre-training, followed by small amounts of adversarial fine-tuning. The benefits of GAN fine-tuning for language generation are unclear, as the resulting models produce comparable or worse samples than traditional language models. We show it is in fact possible to train a language GAN from scratch --- without maximum likelihood pre-training. We combine existing techniques such as large batch sizes, dense rewards and discriminator regularization to stabilize and improve language GANs. The resulting model, ScratchGAN, performs comparably to maximum likelihood training on EMNLP2017 News and WikiText-103 corpora according to quality and diversity metrics.
Cyprien de Masson d'Autume, Shakir Mohamed, Mihaela Rosca, Jack W. Rae
NeurIPS4
2018 Memory-based Parameter Adaptation
Pablo Sprechmann, Siddhant M. Jayakumar, Jack W. Rae, Alexander Pritzel, Adrià Puigdomènech Badia, Benigno Uria, Oriol Vinyals, Demis Hassabis, Razvan Pascanu, Charles Blundell
ICLR (Poster)3
2018 Fast Parametric Learning with Activation Memorization
abstract
Neural networks trained with backpropagation often struggle to identify classes that have been observed a small number of times. In applications where most class labels are rare, such as language modelling, this can become a performance bottleneck. One potential remedy is to augment the network with a fast-learning non-parametric model which stores recent activations and class labels into an external memory. We explore a simplified architecture where we treat a subset of the model parameters as fast memory stores. This can help retain information over longer time intervals than a traditional memory, and does not require additional space or compute. In the case of image classification, we display faster binding of novel classes on an Omniglot image curriculum task. We also show improved performance for word-based language models on news reports (GigaWord), books (Project Gutenberg) and Wikipedia articles (WikiText-103) - the latter achieving a state-of-the-art perplexity of 29.2.
Jack W. Rae, Chris Dyer, Peter Dayan, Timothy P. Lillicrap
ICML1
2018 Relational recurrent neural networks
abstract
Memory-based neural networks model temporal data by leveraging an ability to remember information for long periods. It is unclear, however, whether they also have an ability to perform complex relational reasoning with the information they remember. Here, we first confirm our intuitions that standard memory architectures may struggle at tasks that heavily involve an understanding of the ways in which entities are connected -- i.e., tasks involving relational reasoning. We then improve upon these deficits by using a new memory module -- a Relational Memory Core (RMC) -- which employs multi-head dot product attention to allow memories to interact. Finally, we test the RMC on a suite of tasks that may profit from more capable relational reasoning across sequential information, and show large gains in RL domains (BoxWorld & Mini PacMan), program evaluation, and language modeling, achieving state-of-the-art results on the WikiText-103, Project Gutenberg, and GigaWord datasets.
Adam Santoro, Ryan Faulkner 0001, David Raposo, Jack W. Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy P. Lillicrap
NeurIPS4
2018 Neural Arithmetic Logic Units
abstract
Neural networks can learn to represent and manipulate numerical information, but they seldom generalize well outside of the range of numerical values encountered during training. To encourage more systematic numerical extrapolation, we propose an architecture that represents numerical quantities as linear activations which are manipulated using primitive arithmetic operators, controlled by learned gates. We call this module a neural arithmetic logic unit (NALU), by analogy to the arithmetic logic unit in traditional processors. Experiments show that NALU-enhanced neural networks can learn to track time, perform arithmetic over images of numbers, translate numerical language into real-valued scalars, execute computer code, and count objects in images. In contrast to conventional architectures, we obtain substantially better generalization both inside and outside of the range of numerical values encountered during training, often extrapolating orders of magnitude beyond trained numerical ranges.
Andrew Trask, Felix Hill, Scott E. Reed, Jack W. Rae, Chris Dyer, Phil Blunsom
NeurIPS4
2016 Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes
abstract
Neural networks augmented with external memory have the ability to learn algorithmic solutions to complex tasks. These models appear promising for applications such as language modeling and machine translation. However, they scale poorly in both space and time as the amount of memory grows --- limiting their applicability to real-world domains. Here, we present an end-to-end differentiable memory access scheme, which we call Sparse Access Memory (SAM), that retains the representational power of the original approaches whilst training efficiently with very large memories. We show that SAM achieves asymptotic lower bounds in space and time complexity, and find that an implementation runs $1,\!000\times$ faster and with $3,\!000\times$ less physical memory than non-sparse models. SAM learns with comparable data efficiency to existing models on a range of synthetic tasks and one-shot Omniglot character recognition, and can scale to tasks requiring $100,\!000$s of time steps and memories. As well, we show how our approach can be adapted for models that maintain temporal associations between memories, as with the recently introduced Differentiable Neural Computer.
Jack W. Rae, Jonathan J. Hunt, Ivo Danihelka, Tim Harley, Andrew W. Senior, Greg Wayne, Alex Graves, Timothy P. Lillicrap
NIPS1