Aidan N. Gomez

dblp:202/2262 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
4since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 40% Efficient and distributed learning · 36% Language models and text generation · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 100%

Topics — the 21 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
neural machine translation
0.622018
Depthwise Separable Convolutions for Neural Machine Translation · ICLR (Poster) 2018
Attention is All you Need · NIPS 2017
Natural language and speech › Language models and text generation › neural language model
autoregressive transformer
0.612022
Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval · ICML 2022
Machine learning › Deep learning architectures and training
backpropagation
0.612022
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning
data-efficient learning
0.612022
Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt · ICML 2022
Machine learning › Efficient and distributed learning
data selection
0.612022
Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt · ICML 2022
Machine learning › Efficient and distributed learning
distributed training
0.612022
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
Machine learning › Deep learning architectures and training › neural network training
local learning
0.612022
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.612022
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning › efficient training
training acceleration
0.612022
Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt · ICML 2022
Bioinformatics and computational biology › protein function prediction › protein variant effect prediction
protein fitness prediction
0.612022
Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval · ICML 2022
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.612022
Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval · ICML 2022
Machine learning › Generative modeling
generative adversarial network
0.312018
Unsupervised Cipher Cracking Using Discrete GANs · ICLR (Poster) 2018
Machine learning › Deep learning architectures and training
attention mechanism
0.312017
Attention is All you Need · NIPS 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
The Reversible Residual Network: Backpropagation Without Storing Activations · NIPS 2017
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network
0.312017
The Reversible Residual Network: Backpropagation Without Storing Activations · NIPS 2017
Machine learning › Efficient and distributed learning
memory-efficient training
0.312017
The Reversible Residual Network: Backpropagation Without Storing Activations · NIPS 2017
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.312017
The Reversible Residual Network: Backpropagation Without Storing Activations · NIPS 2017
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.312017
Attention is All you Need · NIPS 2017
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence transduction
0.312017
Attention is All you Need · NIPS 2017
Information retrieval
retrieval-augmented generation
0.212022
Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval · ICML 2022
Machine learning › Deep learning architectures and training
tabular data learning
0.112021
Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep Learning · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

retrieval · 1.7multiple sequence alignment · 1.7autoregressive transformer · 1.7backpropagation · 0.9local learning · 0.6interlocking backpropagation · 0.6importance sampling · 0.6holdout loss · 0.6curriculum learning · 0.6nonparametric estimation · 0.5discrete GANs · 0.3
YearPublicationVenuePosition
2022 Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt
abstract
Training on web-scale data can take months. But much computation and time is wasted on redundant and noisy points that are already learnt or not learnable. To accelerate training, we introduce Reducible Holdout Loss Selection (RHO-LOSS), a simple but principled technique which selects approximately those points for training that most reduce the model’s generalization loss. As a result, RHO-LOSS mitigates the weaknesses of existing data selection methods: techniques from the optimization literature typically select "hard" (e.g. high loss) points, but such points are often noisy (not learnable) or less task-relevant. Conversely, curriculum learning prioritizes "easy" points, but such points need not be trained on once learned. In contrast, RHO-LOSS selects points that are learnable, worth learning, and not yet learnt. RHO-LOSS trains in far fewer steps than prior art, improves accuracy, and speeds up training on a wide range of datasets, hyperparameters, and architectures (MLPs, CNNs, and BERT). On the large web-scraped image dataset Clothing-1M, RHO-LOSS trains in 18x fewer steps and reaches 2% higher final accuracy than uniform data shuffling.
Sören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma, Andreas Kirsch 0002, Winnie Xu, Benedikt Höltgen, Aidan N. Gomez, Adrien Morisot, Sebastian Farquhar, Yarin Gal
ICML8
2022 Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval
abstract
The ability to accurately model the fitness landscape of protein sequences is critical to a wide range of applications, from quantifying the effects of human variants on disease likelihood, to predicting immune-escape mutations in viruses and designing novel biotherapeutic proteins. Deep generative models of protein sequences trained on multiple sequence alignments have been the most successful approaches so far to address these tasks. The performance of these methods is however contingent on the availability of sufficiently deep and diverse alignments for reliable training. Their potential scope is thus limited by the fact many protein families are hard, if not impossible, to align. Large language models trained on massive quantities of non-aligned protein sequences from diverse families address these problems and show potential to eventually bridge the performance gap. We introduce Tranception, a novel transformer architecture leveraging autoregressive predictions and retrieval of homologous sequences at inference to achieve state-of-the-art fitness prediction performance. Given its markedly higher performance on multiple mutants, robustness to shallow alignments and ability to score indels, our approach offers significant gain of scope over existing approaches. To enable more rigorous model testing across a broader range of protein families, we develop ProteinGym – an extensive set of multiplexed assays of variant effects, substantially increasing both the number and diversity of assays compared to existing benchmarks.
Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado, Aidan N. Gomez, Debora S. Marks, Yarin Gal
ICML5
2022 Interlocking Backpropagation: Improving depthwise model-parallelism
abstract
The number of parameters in state of the art neural networks has drastically increased in recent years. This surge of interest in large scale neural networks has motivated the development of new distributed training strategies enabling such models. One such strategy is model-parallel distributed training. Unfortunately, model-parallelism can suffer from poor resource utilisation, which leads to wasted resources. In this work, we improve upon recent developments in an idealised model-parallel optimisation setting: local learning. Motivated by poor resource utilisation in the global setting and poor task performance in the local setting, we introduce a class of intermediary strategies between local and global learning referred to as interlocking backpropagation. These strategies preserve many of the compute-efficiency advantages of local optimisation, while recovering much of the task performance achieved by global optimisation. We assess our strategies on both image classification ResNets and Transformer language models, finding that our strategy consistently out-performs local learning in terms of task performance, and out-performs global learning in training efficiency.
Aidan N. Gomez, Oscar Key, Kuba Perlin, Stephen Gou, Nicholas Frosst, Jeffrey Dean, Yarin Gal
J. Mach. Learn. Res.1
2021 Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep Learning
abstract
We challenge a common assumption underlying most supervised deep learning: that a model makes a prediction depending only on its parameters and the features of a single input. To this end, we introduce a general-purpose deep learning architecture that takes as input the entire dataset instead of processing one datapoint at a time. Our approach uses self-attention to reason about relationships between datapoints explicitly, which can be seen as realizing non-parametric models using parametric attention mechanisms. However, unlike conventional non-parametric models, we let the model learn end-to-end from the data how to make use of other datapoints for prediction. Empirically, our models solve cross-datapoint lookup and complex reasoning tasks unsolvable by traditional deep learning models. We show highly competitive results on tabular data, early results on CIFAR-10, and give insight into how the model makes use of the interactions between points.
Jannik Kossen, Neil Band, Clare Lyle, Aidan N. Gomez, Tom Rainforth, Yarin Gal
NeurIPS4
2018 Unsupervised Cipher Cracking Using Discrete GANs
Aidan N. Gomez, Sicong Huang 0001, Ivan Zhang, Bryan M. Li, Lukasz Kaiser
ICLR (Poster)1
2018 Depthwise Separable Convolutions for Neural Machine Translation
Lukasz Kaiser, Aidan N. Gomez, François Chollet
ICLR (Poster)2
2017 The Reversible Residual Network: Backpropagation Without Storing Activations
abstract
Residual Networks (ResNets) have demonstrated significant improvement over traditional Convolutional Neural Networks (CNNs) on image classification, increasing in performance as networks grow both deeper and wider. However, memory consumption becomes a bottleneck as one needs to store all the intermediate activations for calculating gradients using backpropagation. In this work, we present the Reversible Residual Network (RevNet), a variant of ResNets where each layer's activations can be reconstructed exactly from the next layer's. Therefore, the activations for most layers need not be stored in memory during backprop. We demonstrate the effectiveness of RevNets on CIFAR and ImageNet, establishing nearly identical performance to equally-sized ResNets, with activation storage requirements independent of depth.
Aidan N. Gomez, Mengye Ren, Raquel Urtasun, Roger B. Grosse
NIPS1
2017 Attention is All you Need
abstract
The dominant sequence transduction models are based on complex recurrent orconvolutional neural networks in an encoder and decoder configuration. The best performing such models also connect the encoder and decoder through an attentionm echanisms. We propose a novel, simple network architecture based solely onan attention mechanism, dispensing with recurrence and convolutions entirely.Experiments on two machine translation tasks show these models to be superiorin quality while being more parallelizable and requiring significantly less timeto train. Our single model with 165 million parameters, achieves 27.5 BLEU onEnglish-to-German translation, improving over the existing best ensemble result by over 1 BLEU. On English-to-French translation, we outperform the previoussingle state-of-the-art with model by 0.7 BLEU, achieving a BLEU score of 41.1.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin
NIPS6