Jordan Hoffmann

dblp:243/3115 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Deep learning architectures and training · 26% Graph learning · 19% Efficient and distributed learning · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
scaling laws
1.122022
An empirical analysis of compute-optimal large language model training · NeurIPS 2022
Unified Scaling Laws for Routed Language Models · ICML 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.612022
A Systematic Investigation of Commonsense Knowledge in Large Language Models · EMNLP 2022
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.612022
An empirical analysis of compute-optimal large language model training · NeurIPS 2022
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.612022
Unified Scaling Laws for Routed Language Models · ICML 2022
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Information retrieval
document retrieval
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
independent mechanisms
0.512021
Recurrent Independent Mechanisms · ICLR 2021
Machine learning › Representation and self-supervised learning › representation learning
modular representation learning
0.512021
Recurrent Independent Mechanisms · ICLR 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.512021
Recurrent Independent Mechanisms · ICLR 2021
Machine learning › Graph learning
graph representation learning
0.412020
InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization · ICLR 2020
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412020
InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization · ICLR 2020
Machine learning › Graph learning
graph clustering
0.412019
vGraph: A Generative Model for Joint Community Detection and Node Representation Learning · NeurIPS 2019
Machine learning › Graph learning
network embedding
0.412019
vGraph: A Generative Model for Joint Community Detection and Node Representation Learning · NeurIPS 2019
Machine learning › Graph learning › graph representation learning
node representation learning
0.412019
vGraph: A Generative Model for Joint Community Detection and Node Representation Learning · NeurIPS 2019
Machine learning › Generative modeling › generative model
probabilistic generative model
0.412019
vGraph: A Generative Model for Joint Community Detection and Node Representation Learning · NeurIPS 2019
Machine learning › Learning paradigms
semi-supervised learning
0.112020
InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization · ICLR 2020

Methods — techniques the papers use, named apart from their topics

differentiable encoder · 1.1chunked cross-attention · 1.1scaling law estimation · 0.6probing · 0.6power-law scaling · 0.6effective parameter count · 0.6recurrent neural network · 0.5attention · 0.5mutual information maximization · 0.4contrastive learning · 0.4
YearPublicationVenuePosition
2022 A Systematic Investigation of Commonsense Knowledge in Large Language Models
abstract
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, Aida Nematzadeh. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xiang Li 0069, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh
EMNLP3
2022 Improving Language Models by Retrieving from Trillions of Tokens
abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre
ICML3
2022 Unified Scaling Laws for Routed Language Models
abstract
The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters.
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan
ICML6
2022 An empirical analysis of compute-optimal large language model training
abstract
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more data. Chinchilla uniformly and significantly outperformsGopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, a 7% improvement over Gopher.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche 0002, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, Laurent Sifre
NeurIPS1
2021 Recurrent Independent Mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, Bernhard Schölkopf
ICLR3
2020 InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization
Fan-Yun Sun, Jordan Hoffmann, Vikas Verma, Jian Tang 0005
ICLR2
2019 vGraph: A Generative Model for Joint Community Detection and Node Representation Learning
abstract
This paper focuses on two fundamental tasks of graph analysis: community detection and node representation learning, which capture the global and local structures of graphs respectively. In existing literature, these two tasks are usually independently studied while they are actually highly correlated. We propose a probabilistic generative model called vGraph to learn community membership and node representation collaboratively. Specifically, we assume that each node can be represented as a mixture of communities, and each community is defined as a multinomial distribution over nodes. Both the mixing coefficients and the community distribution are parameterized by the low-dimensional representations of the nodes and communities. We designed an effective variational inference algorithm for the optimization through backpropagation, which regularizes the community membership of neighboring nodes to be similar in the latent space. Experimental results on multiple real-world graphs show that vGraph is very effective in both community detection and node representation learning, outperforming many competitive baselines in both tasks. We show that the framework of vGraph is quite flexible and can be easily extended to detect hierarchical communities.
Fan-Yun Sun, Meng Qu, Jordan Hoffmann, Chin-Wei Huang, Jian Tang 0005
NeurIPS3