VLDB 2026 Research / reviewers in the wild / expert
Aidan Clark
dblp:204/1262
· DBLP profile ↗
6ranked-venue papers
1as first author
3since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 30% Reinforcement learning · 25% Efficient and distributed learning · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
scaling laws |
1.1 | 2 | 2022 | An empirical analysis of compute-optimal large language model training · NeurIPS 2022 Unified Scaling Laws for Routed Language Models · ICML 2022 |
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training |
0.6 | 1 | 2022 | An empirical analysis of compute-optimal large language model training · NeurIPS 2022 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.6 | 1 | 2022 | Unified Scaling Laws for Routed Language Models · ICML 2022 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Information retrieval
document retrieval |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2020 | Stabilizing Transformers for Reinforcement Learning · ICML 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
on-policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.4 | 1 | 2020 | High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020 |
Machine learning › Deep learning architectures and training
transformer |
0.4 | 1 | 2020 | Stabilizing Transformers for Reinforcement Learning · ICML 2020 |
Methods — techniques the papers use, named apart from their topics
differentiable encoder · 1.1chunked cross-attention · 1.1scaling law estimation · 0.6power-law scaling · 0.6effective parameter count · 0.6trust region optimization · 0.4gated transformer · 0.4expectation-maximization · 0.4adversarial network · 0.4LSTM · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Improving Language Models by Retrieving from Trillions of TokensabstractWe enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale. Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre |
ICML | 10 |
| 2022 | Unified Scaling Laws for Routed Language ModelsabstractThe performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters. Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan |
ICML | 1 |
| 2022 | An empirical analysis of compute-optimal large language model trainingabstractWe investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more data. Chinchilla uniformly and significantly outperformsGopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, a 7% improvement over Gopher. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche 0002, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, Laurent Sifre |
NeurIPS | 10 |
| 2020 | High Fidelity Speech Synthesis with Adversarial Networks
Mikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, Karen Simonyan |
ICLR | 4 |
| 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick |
ICLR | 4 |
| 2020 | Stabilizing Transformers for Reinforcement LearningabstractOwing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown breakthrough success in natural language processing (NLP). Harnessing the transformer’s ability to process long time horizons of information could provide a similar performance boost in partially observable reinforcement learning (RL) domains, but the large-scale transformers used in NLP have yet to be successfully applied to the RL setting. In this work we demonstrate that the standard transformer architecture is difficult to optimize, which was previously observed in the supervised learning setting but becomes especially pronounced with RL objectives. We propose architectural modifications that substantially improve the stability and learning speed of the original Transformer and XL variant. The proposed architecture, the Gated Transformer-XL (GTrXL), surpasses LSTMs on challenging memory environments and achieves state-of-the-art results on the multi-task DMLab-30 benchmark suite, exceeding the performance of an external memory architecture. We show that the GTrXL has stability and performance that consistently matches or exceeds a competitive LSTM baseline, including on more reactive tasks where memory is less critical. Emilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant M. Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matt M. Botvinick, Nicolas Heess, Raia Hadsell |
ICML | 9 |