VLDB 2026 Research / reviewers in the wild / expert
Andrew Brock
dblp:186/8021 · also Andy Brock
· DBLP profile ↗
13ranked-venue papers
5as first author
8since 2021 · last 2023
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Deep learning architectures and training · 39% Language models and text generation · 14% Generative modeling · 12% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
transformer |
1.2 | 2 | 2023 | Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation · ICLR 2023 Perceiver: General Perception with Iterative Attention · ICML 2021 |
Machine learning › Generative modeling
generative adversarial network |
1.1 | 3 | 2020 | Training Generative Adversarial Networks by Solving Ordinary Differential Equations · NeurIPS 2020 Large Scale GAN Training for High Fidelity Natural Image Synthesis · ICLR 2019 Neural Photo Editing with Introspective Adversarial Networks · ICLR (Poster) 2017 |
Machine learning › Deep learning architectures and training › neural network training
normalization-free training |
1.0 | 2 | 2021 | High-Performance Large-Scale Image Recognition Without Normalization · ICML 2021 Characterizing signal propagation to close the performance gap in unnormalized ResNets · ICLR 2021 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.8 | 2 | 2020 | Evolving Normalization-Activation Layers · NeurIPS 2020 SMASH: One-Shot Model Architecture Search through HyperNetworks · ICLR (Poster) 2018 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.7 | 1 | 2023 | Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation · ICLR 2023 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
in-context learning |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Natural language and speech › Language models and text generation
multimodal language model |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Computer vision › Vision and language
vision-language model |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Information retrieval
document retrieval |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Computer vision › Video understanding and tracking
cross-modal perception |
0.5 | 1 | 2021 | Perceiver: General Perception with Iterative Attention · ICML 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | High-Performance Large-Scale Image Recognition Without Normalization · ICML 2021 |
Machine learning › Deep learning architectures and training › attention mechanism
iterative attention |
0.5 | 1 | 2021 | Perceiver: General Perception with Iterative Attention · ICML 2021 |
Machine learning › Deep learning architectures and training › training dynamics
signal propagation |
0.5 | 1 | 2021 | Characterizing signal propagation to close the performance gap in unnormalized ResNets · ICLR 2021 |
Machine learning › Optimization for machine learning › evolutionary computation
genetic algorithms |
0.4 | 1 | 2020 | Evolving Normalization-Activation Layers · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › normalization
normalization layers |
0.4 | 1 | 2020 | Evolving Normalization-Activation Layers · NeurIPS 2020 |
Mathematical optimization › numerical analysis
ODE solvers |
0.4 | 1 | 2020 | Training Generative Adversarial Networks by Solving Ordinary Differential Equations · NeurIPS 2020 |
Mathematical optimization
ordinary differential equation |
0.4 | 1 | 2020 | Training Generative Adversarial Networks by Solving Ordinary Differential Equations · NeurIPS 2020 |
Machine learning › Generative modeling › image generation
natural image synthesis |
0.4 | 1 | 2019 | Large Scale GAN Training for High Fidelity Natural Image Synthesis · ICLR 2019 |
Machine learning › Efficient and distributed learning
parameter sharing |
0.3 | 1 | 2018 | SMASH: One-Shot Model Architecture Search through HyperNetworks · ICLR (Poster) 2018 |
Visual content generation and editing
image editing |
0.3 | 1 | 2017 | Neural Photo Editing with Introspective Adversarial Networks · ICLR (Poster) 2017 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
differentiable encoder · 1.1chunked cross-attention · 1.1resnet · 1.0self-attention modification · 0.7transformer architecture · 0.6pretrained vision encoder · 0.6pre-trained language model · 0.6interleaved multimodal pretraining · 0.6attention · 0.6dynamical systems analysis · 0.5runge-kutta · 0.4integration error control · 0.4latent space manipulation · 0.3introspective adversarial network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation
Bobby He, James Martens, Aleksandar Botev, Andrew Brock, Samuel L. Smith, Yee Whye Teh |
ICLR | 5 |
| 2022 | Towards Learning Universal Audio RepresentationsabstractThe ability to learn universal audio representations that can solve diverse speech, music, and environment tasks can spur many applications that require general sound content understanding. In this work, we introduce a holistic audio representation evaluation suite (HARES) spanning 12 downstream tasks across audio domains and provide a thorough empirical study of recent sound representation learning systems on that benchmark. We discover that previous sound event classification or speech models do not generalize outside of their domains. We observe that more robust audio representations can be learned with the SimCLR objective; however, the model’s transferability depends heavily on the model architecture. We find the Slowfast architecture is good at learning rich representations required by different domains, but its performance is affected by the normalization scheme. Based on these findings, we propose a novel normalizer-free Slowfast NFNet and achieve state-of-the-art performance across all domains. Luyu Wang, Pauline Luc, Yan Wu 0010, Adrià Recasens, Lucas Smaira, Andrew Brock, Andrew Jaegle, Jean-Baptiste Alayrac, Sander Dieleman, João Carreira 0001, Aäron van den Oord |
ICASSP | 6 |
| 2022 | Perceiver IO: A General Architecture for Structured Inputs & Outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier J. Hénaff, Matt M. Botvinick, Andrew Zisserman, Oriol Vinyals, João Carreira 0001 |
ICLR | 9 |
| 2022 | Improving Language Models by Retrieving from Trillions of TokensabstractWe enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale. Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre |
ICML | 20 |
| 2022 | Flamingo: a Visual Language Model for Few-Shot LearningabstractBuilding models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan |
NeurIPS | 20 |
| 2021 | Characterizing signal propagation to close the performance gap in unnormalized ResNets
Andrew Brock, Soham De, Samuel L. Smith |
ICLR | 1 |
| 2021 | High-Performance Large-Scale Image Recognition Without NormalizationabstractBatch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions between examples. Although recent work has succeeded in training deep ResNets without normalization layers, these models do not match the test accuracies of the best batch-normalized networks, and are often unstable for large learning rates or strong data augmentations. In this work, we develop an adaptive gradient clipping technique which overcomes these instabilities, and design a significantly improved class of Normalizer-Free ResNets. Our smaller models match the test accuracy of an EfficientNet-B7 on ImageNet while being up to 8.7x faster to train, and our largest models attain a new state-of-the-art top-1 accuracy of 86.5%. In addition, Normalizer-Free models attain significantly better performance than their batch-normalized counterparts when fine-tuning on ImageNet after large-scale pre-training on a dataset of 300 million labeled images, with our best models obtaining an accuracy of 89.2%. Andrew Brock, Soham De, Samuel L. Smith, Karen Simonyan |
ICML | 1 |
| 2021 | Perceiver: General Perception with Iterative AttentionabstractBiological systems understand the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. The perception models used in deep learning on the other hand are designed for individual modalities, often relying on domain-specific assumptions such as the local grid structures exploited by virtually all existing vision models. These priors introduce helpful inductive biases, but also lock models to individual modalities. In this paper we introduce the Perceiver {–} a model that builds upon Transformers and hence makes few architectural assumptions about the relationship between its inputs, but that also scales to hundreds of thousands of inputs, like ConvNets. The model leverages an asymmetric attention mechanism to iteratively distill inputs into a tight latent bottleneck, allowing it to scale to handle very large inputs. We show that this architecture is competitive with or outperforms strong, specialized models on classification tasks across various modalities: images, point clouds, audio, video and video+audio. The Perceiver obtains performance comparable to ResNet-50 and ViT on ImageNet without 2D convolutions by directly attending to 50,000 pixels. It is also competitive in all modalities in AudioSet. Andrew Jaegle, Felix Gimeno, Andrew Brock, Oriol Vinyals, Andrew Zisserman, João Carreira 0001 |
ICML | 3 |
| 2020 | Evolving Normalization-Activation LayersabstractNormalization layers and activation functions are fundamental components in deep networks and typically co-locate with each other. Here we propose to design them using an automated approach. Instead of designing them separately, we unify them into a single tensor-to-tensor computation graph, and evolve its structure starting from basic mathematical functions. Examples of such mathematical functions are addition, multiplication and statistical moments. The use of low-level mathematical functions, in contrast to the use of high-level modules in mainstream NAS, leads to a highly sparse and large search space which can be challenging for search methods. To address the challenge, we develop efficient rejection protocols to quickly filter out candidate layers that do not work well. We also use multi-objective evolution to optimize each layer's performance across many architectures to prevent overfitting. Our method leads to the discovery of EvoNorms, a set of new normalization-activation layers with novel, and sometimes surprising structures that go beyond existing design patterns. For example, some EvoNorms do not assume that normalization and activation functions must be applied sequentially, nor need to center the feature maps, nor require explicit activation functions. Our experiments show that EvoNorms work well on image classification models including ResNets, MobileNets and EfficientNets but also transfer well to Mask R-CNN with FPN/SpineNet for instance segmentation and to BigGAN for image synthesis, outperforming BatchNorm and GroupNorm based layers in many cases. Hanxiao Liu, Andrew Brock, Karen Simonyan, Quoc V. Le |
NeurIPS | 2 |
| 2020 | Training Generative Adversarial Networks by Solving Ordinary Differential EquationsabstractThe instability of Generative Adversarial Network (GAN) training has frequently been attributed to gradient descent. Consequently, recent methods have aimed to tailor the models and training procedures to stabilise the discrete updates. In contrast, we study the continuous-time dynamics induced by GAN training. Both theory and toy experiments suggest that these dynamics are in fact surprisingly stable. From this perspective, we hypothesise that instabilities in training GANs arise from the integration error in discretising the continuous dynamics. We experimentally verify that well-known ODE solvers (such as Runge-Kutta) can stabilise training - when combined with a regulariser that controls the integration error. Our approach represents a radical departure from previous methods which typically use adaptive optimisation and stabilisation techniques that constrain the functional space (e.g. Spectral Normalisation). Evaluation on CIFAR-10 and ImageNet shows that our method outperforms several strong baselines, demonstrating its efficacy. Chongli Qin, Yan Wu 0010, Jost Tobias Springenberg, Andrew Brock, Jeff Donahue, Timothy P. Lillicrap, Pushmeet Kohli |
NeurIPS | 4 |
| 2019 | Large Scale GAN Training for High Fidelity Natural Image Synthesis
Andrew Brock, Jeff Donahue, Karen Simonyan |
ICLR | 1 |
| 2018 | SMASH: One-Shot Model Architecture Search through HyperNetworks
Andrew Brock, Theodore Lim, James M. Ritchie, Nick Weston |
ICLR (Poster) | 1 |
| 2017 | Neural Photo Editing with Introspective Adversarial Networks
Andrew Brock, Theodore Lim, James M. Ritchie, Nick Weston |
ICLR (Poster) | 1 |