Andrew Brock

dblp:186/8021 · also Andy Brock · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2023
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Deep learning architectures and training · 39% Language models and text generation · 14% Generative modeling · 12%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
transformer
1.222023
Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation · ICLR 2023
Perceiver: General Perception with Iterative Attention · ICML 2021
Machine learning › Generative modeling
generative adversarial network
1.132020
Training Generative Adversarial Networks by Solving Ordinary Differential Equations · NeurIPS 2020
Large Scale GAN Training for High Fidelity Natural Image Synthesis · ICLR 2019
Neural Photo Editing with Introspective Adversarial Networks · ICLR (Poster) 2017
Machine learning › Deep learning architectures and training › neural network training
normalization-free training
1.022021
High-Performance Large-Scale Image Recognition Without Normalization · ICML 2021
Characterizing signal propagation to close the performance gap in unnormalized ResNets · ICLR 2021
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.822020
Evolving Normalization-Activation Layers · NeurIPS 2020
SMASH: One-Shot Model Architecture Search through HyperNetworks · ICLR (Poster) 2018
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.712023
Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation · ICLR 2023
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Natural language and speech › Language models and text generation
in-context learning
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Natural language and speech › Language models and text generation
multimodal language model
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Computer vision › Vision and language
vision-language model
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Information retrieval
document retrieval
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Computer vision › Video understanding and tracking
cross-modal perception
0.512021
Perceiver: General Perception with Iterative Attention · ICML 2021
Computer vision › Image recognition and object detection
image classification
0.512021
High-Performance Large-Scale Image Recognition Without Normalization · ICML 2021
Machine learning › Deep learning architectures and training › attention mechanism
iterative attention
0.512021
Perceiver: General Perception with Iterative Attention · ICML 2021
Machine learning › Deep learning architectures and training › training dynamics
signal propagation
0.512021
Characterizing signal propagation to close the performance gap in unnormalized ResNets · ICLR 2021
Machine learning › Optimization for machine learning › evolutionary computation
genetic algorithms
0.412020
Evolving Normalization-Activation Layers · NeurIPS 2020
Machine learning › Deep learning architectures and training › normalization
normalization layers
0.412020
Evolving Normalization-Activation Layers · NeurIPS 2020
Mathematical optimization › numerical analysis
ODE solvers
0.412020
Training Generative Adversarial Networks by Solving Ordinary Differential Equations · NeurIPS 2020
Mathematical optimization
ordinary differential equation
0.412020
Training Generative Adversarial Networks by Solving Ordinary Differential Equations · NeurIPS 2020
Machine learning › Generative modeling › image generation
natural image synthesis
0.412019
Large Scale GAN Training for High Fidelity Natural Image Synthesis · ICLR 2019
Machine learning › Efficient and distributed learning
parameter sharing
0.312018
SMASH: One-Shot Model Architecture Search through HyperNetworks · ICLR (Poster) 2018
Visual content generation and editing
image editing
0.312017
Neural Photo Editing with Introspective Adversarial Networks · ICLR (Poster) 2017
Computer vision › Vision and language
visual question answering
0.212022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

differentiable encoder · 1.1chunked cross-attention · 1.1resnet · 1.0self-attention modification · 0.7transformer architecture · 0.6pretrained vision encoder · 0.6pre-trained language model · 0.6interleaved multimodal pretraining · 0.6attention · 0.6dynamical systems analysis · 0.5runge-kutta · 0.4integration error control · 0.4latent space manipulation · 0.3introspective adversarial network · 0.3
YearPublicationVenuePosition
2023 Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation
Bobby He, James Martens, Aleksandar Botev, Andrew Brock, Samuel L. Smith, Yee Whye Teh
ICLR5
2022 Towards Learning Universal Audio Representations
abstract
The ability to learn universal audio representations that can solve diverse speech, music, and environment tasks can spur many applications that require general sound content understanding. In this work, we introduce a holistic audio representation evaluation suite (HARES) spanning 12 downstream tasks across audio domains and provide a thorough empirical study of recent sound representation learning systems on that benchmark. We discover that previous sound event classification or speech models do not generalize outside of their domains. We observe that more robust audio representations can be learned with the SimCLR objective; however, the model’s transferability depends heavily on the model architecture. We find the Slowfast architecture is good at learning rich representations required by different domains, but its performance is affected by the normalization scheme. Based on these findings, we propose a novel normalizer-free Slowfast NFNet and achieve state-of-the-art performance across all domains.
Luyu Wang, Pauline Luc, Yan Wu 0010, Adrià Recasens, Lucas Smaira, Andrew Brock, Andrew Jaegle, Jean-Baptiste Alayrac, Sander Dieleman, João Carreira 0001, Aäron van den Oord
ICASSP6
2022 Perceiver IO: A General Architecture for Structured Inputs & Outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier J. Hénaff, Matt M. Botvinick, Andrew Zisserman, Oriol Vinyals, João Carreira 0001
ICLR9
2022 Improving Language Models by Retrieving from Trillions of Tokens
abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre
ICML20
2022 Flamingo: a Visual Language Model for Few-Shot Learning
abstract
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan
NeurIPS20
2021 Characterizing signal propagation to close the performance gap in unnormalized ResNets
Andrew Brock, Soham De, Samuel L. Smith
ICLR1
2021 High-Performance Large-Scale Image Recognition Without Normalization
abstract
Batch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions between examples. Although recent work has succeeded in training deep ResNets without normalization layers, these models do not match the test accuracies of the best batch-normalized networks, and are often unstable for large learning rates or strong data augmentations. In this work, we develop an adaptive gradient clipping technique which overcomes these instabilities, and design a significantly improved class of Normalizer-Free ResNets. Our smaller models match the test accuracy of an EfficientNet-B7 on ImageNet while being up to 8.7x faster to train, and our largest models attain a new state-of-the-art top-1 accuracy of 86.5%. In addition, Normalizer-Free models attain significantly better performance than their batch-normalized counterparts when fine-tuning on ImageNet after large-scale pre-training on a dataset of 300 million labeled images, with our best models obtaining an accuracy of 89.2%.
Andrew Brock, Soham De, Samuel L. Smith, Karen Simonyan
ICML1
2021 Perceiver: General Perception with Iterative Attention
abstract
Biological systems understand the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. The perception models used in deep learning on the other hand are designed for individual modalities, often relying on domain-specific assumptions such as the local grid structures exploited by virtually all existing vision models. These priors introduce helpful inductive biases, but also lock models to individual modalities. In this paper we introduce the Perceiver {–} a model that builds upon Transformers and hence makes few architectural assumptions about the relationship between its inputs, but that also scales to hundreds of thousands of inputs, like ConvNets. The model leverages an asymmetric attention mechanism to iteratively distill inputs into a tight latent bottleneck, allowing it to scale to handle very large inputs. We show that this architecture is competitive with or outperforms strong, specialized models on classification tasks across various modalities: images, point clouds, audio, video and video+audio. The Perceiver obtains performance comparable to ResNet-50 and ViT on ImageNet without 2D convolutions by directly attending to 50,000 pixels. It is also competitive in all modalities in AudioSet.
Andrew Jaegle, Felix Gimeno, Andrew Brock, Oriol Vinyals, Andrew Zisserman, João Carreira 0001
ICML3
2020 Evolving Normalization-Activation Layers
abstract
Normalization layers and activation functions are fundamental components in deep networks and typically co-locate with each other. Here we propose to design them using an automated approach. Instead of designing them separately, we unify them into a single tensor-to-tensor computation graph, and evolve its structure starting from basic mathematical functions. Examples of such mathematical functions are addition, multiplication and statistical moments. The use of low-level mathematical functions, in contrast to the use of high-level modules in mainstream NAS, leads to a highly sparse and large search space which can be challenging for search methods. To address the challenge, we develop efficient rejection protocols to quickly filter out candidate layers that do not work well. We also use multi-objective evolution to optimize each layer's performance across many architectures to prevent overfitting. Our method leads to the discovery of EvoNorms, a set of new normalization-activation layers with novel, and sometimes surprising structures that go beyond existing design patterns. For example, some EvoNorms do not assume that normalization and activation functions must be applied sequentially, nor need to center the feature maps, nor require explicit activation functions. Our experiments show that EvoNorms work well on image classification models including ResNets, MobileNets and EfficientNets but also transfer well to Mask R-CNN with FPN/SpineNet for instance segmentation and to BigGAN for image synthesis, outperforming BatchNorm and GroupNorm based layers in many cases.
Hanxiao Liu, Andrew Brock, Karen Simonyan, Quoc V. Le
NeurIPS2
2020 Training Generative Adversarial Networks by Solving Ordinary Differential Equations
abstract
The instability of Generative Adversarial Network (GAN) training has frequently been attributed to gradient descent. Consequently, recent methods have aimed to tailor the models and training procedures to stabilise the discrete updates. In contrast, we study the continuous-time dynamics induced by GAN training. Both theory and toy experiments suggest that these dynamics are in fact surprisingly stable. From this perspective, we hypothesise that instabilities in training GANs arise from the integration error in discretising the continuous dynamics. We experimentally verify that well-known ODE solvers (such as Runge-Kutta) can stabilise training - when combined with a regulariser that controls the integration error. Our approach represents a radical departure from previous methods which typically use adaptive optimisation and stabilisation techniques that constrain the functional space (e.g. Spectral Normalisation). Evaluation on CIFAR-10 and ImageNet shows that our method outperforms several strong baselines, demonstrating its efficacy.
Chongli Qin, Yan Wu 0010, Jost Tobias Springenberg, Andrew Brock, Jeff Donahue, Timothy P. Lillicrap, Pushmeet Kohli
NeurIPS4
2019 Large Scale GAN Training for High Fidelity Natural Image Synthesis
Andrew Brock, Jeff Donahue, Karen Simonyan
ICLR1
2018 SMASH: One-Shot Model Architecture Search through HyperNetworks
Andrew Brock, Theodore Lim, James M. Ritchie, Nick Weston
ICLR (Poster)1
2017 Neural Photo Editing with Introspective Adversarial Networks
Andrew Brock, Theodore Lim, James M. Ritchie, Nick Weston
ICLR (Poster)1