VLDB 2026 Research / reviewers in the wild / expert
Karen Simonyan
dblp:78/470 · also Karén Simonyan
· DBLP profile ↗
41ranked-venue papers
8as first author
8since 2021 · last 2022
0000-0003-0787-5507ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
33 papers |
Efficient and distributed learning · 19% Deep learning architectures and training · 16% Reinforcement learning · 10% | |
| Computer graphics and multimedia
3 papers |
Audio and music processing · 68% Visual content generation and editing · 32% |
Topics — the 30 heaviest of 79, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
generative adversarial network |
1.3 | 4 | 2021 | High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020 Large Scale Adversarial Representation Learning · NeurIPS 2019 Large Scale GAN Training for High Fidelity Natural Image Synthesis · ICLR 2019 |
Machine learning › Deep learning architectures and training
scaling laws |
1.1 | 2 | 2022 | An empirical analysis of compute-optimal large language model training · NeurIPS 2022 Unified Scaling Laws for Routed Language Models · ICML 2022 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
1.1 | 3 | 2020 | Evolving Normalization-Activation Layers · NeurIPS 2020 DARTS: Differentiable Architecture Search · ICLR (Poster) 2019 Hierarchical Representations for Efficient Architecture Search · ICLR (Poster) 2018 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 3 | 2018 | The challenge of realistic music generation: modelling raw audio at scale · NeurIPS 2018 Video Pixel Networks · ICML 2017 Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders · ICML 2017 |
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis |
0.8 | 2 | 2021 | End-to-end Adversarial Text-to-Speech · ICLR 2021 Efficient Neural Audio Synthesis · ICML 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 2 | 2020 | Fast Sparse ConvNets · CVPR 2020 Efficient Neural Audio Synthesis · ICML 2018 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.8 | 2 | 2020 | High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020 Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018 |
Computer vision › Image recognition and object detection
image classification |
0.7 | 3 | 2021 | High-Performance Large-Scale Image Recognition Without Normalization · ICML 2021 Deep Fisher Networks for Large-Scale Image Classification · NIPS 2013 A Compact and Discriminative Face Track Descriptor · CVPR 2014 |
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training |
0.6 | 1 | 2022 | An empirical analysis of compute-optimal large language model training · NeurIPS 2022 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
in-context learning |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.6 | 1 | 2022 | Unified Scaling Laws for Routed Language Models · ICML 2022 |
Natural language and speech › Language models and text generation
multimodal language model |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Computer vision › Vision and language
vision-language model |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Information retrieval
document retrieval |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.5 | 1 | 2021 | Machine Translation Decoding beyond Beam Search · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
decoding |
0.5 | 1 | 2021 | Machine Translation Decoding beyond Beam Search · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training › neural network training
normalization-free training |
0.5 | 1 | 2021 | High-Performance Large-Scale Image Recognition Without Normalization · ICML 2021 |
Machine learning › Deep learning architectures and training › recurrent neural network
real-time recurrent learning |
0.5 | 1 | 2021 | Practical Real Time Recurrent Learning with a Sparse Approximation · ICLR 2021 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.5 | 1 | 2021 | Practical Real Time Recurrent Learning with a Sparse Approximation · ICLR 2021 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.5 | 1 | 2021 | Practical Real Time Recurrent Learning with a Sparse Approximation · ICLR 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.4 | 2 | 2016 | Reading Text in the Wild with Convolutional Neural Networks · Int. J. Comput. Vis. 2016 Two-Stream Convolutional Networks for Action Recognition in Videos · NIPS 2014 |
Machine learning › Reinforcement learning
actor-critic methods |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Optimization for machine learning › evolutionary computation
genetic algorithms |
0.4 | 1 | 2020 | Evolving Normalization-Activation Layers · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › normalization
normalization layers |
0.4 | 1 | 2020 | Evolving Normalization-Activation Layers · NeurIPS 2020 |
Machine learning › Reinforcement learning › actor-critic methods
off-policy actor-critic |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Computer vision › 3D vision › point cloud processing
sparse convolution |
0.4 | 1 | 2020 | Fast Sparse ConvNets · CVPR 2020 |
Methods — techniques the papers use, named apart from their topics
autoregressive modeling · 1.6differentiable encoder · 1.1chunked cross-attention · 1.1adversarial training · 0.9reinforcement learning · 0.8v-trace · 0.8gradient-based optimization · 0.7convolutional neural network · 0.6power-law scaling · 0.6effective parameter count · 0.6sparse kernels · 0.4depthwise separable convolution · 0.4autoregressive discrete autoencoders · 0.3probabilistic video model · 0.3four-dimensional dependency chain · 0.3autoencoder · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Improving Language Models by Retrieving from Trillions of TokensabstractWe enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale. Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre |
ICML | 25 |
| 2022 | Unified Scaling Laws for Routed Language ModelsabstractThe performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters. Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan |
ICML | 26 |
| 2022 | Flamingo: a Visual Language Model for Few-Shot LearningabstractBuilding models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan |
NeurIPS | 27 |
| 2022 | An empirical analysis of compute-optimal large language model trainingabstractWe investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more data. Chinchilla uniformly and significantly outperformsGopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, a 7% improvement over Gopher. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche 0002, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, Laurent Sifre |
NeurIPS | 18 |
| 2021 | Machine Translation Decoding beyond Beam SearchabstractRémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Lespiau Jean-Baptiste, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Jean-Baptiste Lespiau, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals |
EMNLP (1) | 7 |
| 2021 | End-to-end Adversarial Text-to-Speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, Karen Simonyan |
ICLR | 5 |
| 2021 | Practical Real Time Recurrent Learning with a Sparse Approximation
Jacob Menick, Erich Elsen, Utku Evci, Simon Osindero, Karen Simonyan, Alex Graves |
ICLR | 5 |
| 2021 | High-Performance Large-Scale Image Recognition Without NormalizationabstractBatch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions between examples. Although recent work has succeeded in training deep ResNets without normalization layers, these models do not match the test accuracies of the best batch-normalized networks, and are often unstable for large learning rates or strong data augmentations. In this work, we develop an adaptive gradient clipping technique which overcomes these instabilities, and design a significantly improved class of Normalizer-Free ResNets. Our smaller models match the test accuracy of an EfficientNet-B7 on ImageNet while being up to 8.7x faster to train, and our largest models attain a new state-of-the-art top-1 accuracy of 86.5%. In addition, Normalizer-Free models attain significantly better performance than their batch-normalized counterparts when fine-tuning on ImageNet after large-scale pre-training on a dataset of 300 million labeled images, with our best models obtaining an accuracy of 89.2%. Andrew Brock, Soham De, Samuel L. Smith, Karen Simonyan |
ICML | 4 |
| 2020 | Fast Sparse ConvNetsabstractHistorically, the pursuit of efficient inference has been one of the driving forces behind the research into new deep learning architectures and building blocks. Some of the recent examples include: the squeeze-and-excitation module, depthwise separable convolutions in Xception, and the inverted bottleneck in MobileNet v2. Notably, in all of these cases, the resulting building blocks enabled not only higher efficiency, but also higher accuracy, and found wide adoption in the field. In this work, we further expand the arsenal of efficient building blocks for neural network architectures; but instead of combining standard primitives (such as convolution), we advocate for the replacement of these dense primitives with their sparse counterparts. While the idea of using sparsity to decrease the parameter count is not new, the conventional wisdom is that this reduction in theoretical FLOPs does not translate into real-world efficiency gains. We aim to correct this misconception by introducing a family of efficient sparse kernels for several hardware platforms, which we plan to open source for the benefit of the community. Equipped with our efficient implementation of sparse primitives, we show that sparse versions of MobileNet v1 and MobileNet v2 architectures substantially outperform strong dense baselines on the efficiency-accuracy curve. On Snapdragon 835 our sparse networks outperform their dense equivalents by 1.3 - 2.4× - equivalent to approximately one entire generation of improvement. We hope that our findings will facilitate wider adoption of sparsity as a tool for creating efficient and accurate deep learning architectures. Erich Elsen, Marat Dukhan, Trevor Gale, Karen Simonyan |
CVPR | 4 |
| 2020 | High Fidelity Speech Synthesis with Adversarial Networks
Mikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, Karen Simonyan |
ICLR | 8 |
| 2020 | Off-Policy Actor-Critic with Shared Experience ReplayabstractWe investigate the combination of actor-critic reinforcement learning algorithms with a uniform large-scale experience replay and propose solutions for two ensuing challenges: (a) efficient actor-critic learning with experience replay (b) the stability of off-policy learning where agents learn from other agents behaviour. To this end we analyze the bias-variance tradeoffs in V-trace, a form of importance sampling for actor-critic methods. Based on our analysis, we then argue for mixing experience sampled from replay with on-policy experience, and propose a new trust region scheme that scales effectively to data distributions where V-trace becomes unstable. We provide extensive empirical validation of the proposed solutions on DMLab-30 and further show the benefits of this setup in two training regimes for Atari: (1) a single agent is trained up until 200M environment frames per game (2) a population of agents is trained up until 200M environment frames each and may share experience. We demonstrate state-of-the-art data efficiency among model-free agents in both regimes. Simon Schmitt, Matteo Hessel, Karen Simonyan |
ICML | 3 |
| 2020 | Evolving Normalization-Activation LayersabstractNormalization layers and activation functions are fundamental components in deep networks and typically co-locate with each other. Here we propose to design them using an automated approach. Instead of designing them separately, we unify them into a single tensor-to-tensor computation graph, and evolve its structure starting from basic mathematical functions. Examples of such mathematical functions are addition, multiplication and statistical moments. The use of low-level mathematical functions, in contrast to the use of high-level modules in mainstream NAS, leads to a highly sparse and large search space which can be challenging for search methods. To address the challenge, we develop efficient rejection protocols to quickly filter out candidate layers that do not work well. We also use multi-objective evolution to optimize each layer's performance across many architectures to prevent overfitting. Our method leads to the discovery of EvoNorms, a set of new normalization-activation layers with novel, and sometimes surprising structures that go beyond existing design patterns. For example, some EvoNorms do not assume that normalization and activation functions must be applied sequentially, nor need to center the feature maps, nor require explicit activation functions. Our experiments show that EvoNorms work well on image classification models including ResNets, MobileNets and EfficientNets but also transfer well to Mask R-CNN with FPN/SpineNet for instance segmentation and to BigGAN for image synthesis, outperforming BatchNorm and GroupNorm based layers in many cases. Hanxiao Liu, Andrew Brock, Karen Simonyan, Quoc V. Le |
NeurIPS | 3 |
| 2020 | This time with feeling: learning expressive musical performanceabstractMusic generation has generally been focused on either creating scores or interpreting them. We discuss differences between these two problems and propose that, in fact, it may be valuable to work in the space of direct performance generation: jointly predicting the notes and also their expressive timing and dynamics. We consider the significance and qualities of the dataset needed for this. Having identified both a problem domain and characteristics of an appropriate dataset, we show an LSTM-based recurrent network model that subjectively performs quite well on this task. Critically, we provide generated examples. We also include feedback from professional composers and musicians about some of these examples. Sageev Oore, Ian Simon, Sander Dieleman, Douglas Eck, Karen Simonyan |
Neural Comput. Appl. | 5 |
| 2019 | Large Scale GAN Training for High Fidelity Natural Image Synthesis
Andrew Brock, Jeff Donahue, Karen Simonyan |
ICLR | 3 |
| 2019 | DARTS: Differentiable Architecture Search
Hanxiao Liu, Karen Simonyan, Yiming Yang 0002 |
ICLR (Poster) | 2 |
| 2019 | Large Scale Adversarial Representation LearningabstractAdversarially trained generative models (GANs) have recently achieved compelling image synthesis results. But despite early successes in using GANs for unsupervised representation learning, they have since been superseded by approaches based on self-supervision. In this work we show that progress in image generation quality translates to substantially improved representation learning performance. Our approach, BigBiGAN, builds upon the state-of-the-art BigGAN model, extending it to representation learning by adding an encoder and modifying the discriminator. We extensively evaluate the representation learning and generation capabilities of these BigBiGAN models, demonstrating that these generation-based models achieve the state of the art in unsupervised representation learning on ImageNet, as well as compelling results in unconditional image generation. Jeff Donahue, Karen Simonyan |
NeurIPS | 2 |
| 2018 | Hierarchical Representations for Efficient Architecture Search
Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, Koray Kavukcuoglu |
ICLR (Poster) | 2 |
| 2018 | IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner ArchitecturesabstractIn this work we aim to solve a large collection of tasks using a single reinforcement learning agent with a single set of parameters. A key challenge is to handle the increased amount of data and extended training time. We have developed a new distributed agent IMPALA (Importance Weighted Actor-Learner Architecture) that not only uses resources more efficiently in single-machine training but also scales to thousands of machines without sacrificing data efficiency or resource utilisation. We achieve stable learning at high throughput by combining decoupled acting and learning with a novel off-policy correction method called V-trace. We demonstrate the effectiveness of IMPALA for multi-task reinforcement learning on DMLab-30 (a set of 30 tasks from the DeepMind Lab environment (Beattie et al., 2016)) and Atari57 (all available Atari games in Arcade Learning Environment (Bellemare et al., 2013a)). Our results show that IMPALA is able to achieve better performance than previous agents with less data, and crucially exhibits positive transfer between tasks as a result of its multi-task approach. Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, Koray Kavukcuoglu |
ICML | 4 |
| 2018 | Learning to Search with MCTSnetsabstractPlanning problems are among the most important and well-studied problems in artificial intelligence. They are most typically solved by tree search algorithms that simulate ahead into the future, evaluate future states, and back-up those evaluations to the root of a search tree. Among these algorithms, Monte-Carlo tree search (MCTS) is one of the most general, powerful and widely used. A typical implementation of MCTS uses cleverly designed rules, optimised to the particular characteristics of the domain. These rules control where the simulation traverses, what to evaluate in the states that are reached, and how to back-up those evaluations. In this paper we instead learn where, what and how to search. Our architecture, which we call an MCTSnet, incorporates simulation-based search inside a neural network, by expanding, evaluating and backing-up a vector embedding. The parameters of the network are trained end-to-end using gradient-based optimisation. When applied to small searches in the well-known planning problem Sokoban, the learned search algorithm significantly outperformed MCTS baselines. Arthur Guez, Theophane Weber, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals, Daan Wierstra, Rémi Munos, David Silver 0001 |
ICML | 4 |
| 2018 | Efficient Neural Audio SynthesisabstractSequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating desired samples. Efficient sampling for this class of models at the cost of little to no loss in quality has however remained an elusive problem. With a focus on text-to-speech synthesis, we describe a set of general techniques for reducing sampling time while maintaining high output quality. We first describe a single-layer recurrent neural network, the WaveRNN, with a dual softmax layer that matches the quality of the state-of-the-art WaveNet model. The compact form of the network makes it possible to generate 24 kHz 16-bit audio 4 times faster than real time on a GPU. Secondly, we apply a weight pruning technique to reduce the number of weights in the WaveRNN. We find that, for a constant number of parameters, large sparse networks perform better than small dense networks and this relationship holds past sparsity levels of more than 96%. The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile phone CPU in real time. Finally, we describe a new dependency scheme for sampling that lets us trade a constant number of non-local, distant dependencies for the ability to generate samples in batches. The Batch WaveRNN produces 8 samples per step without loss of quality and offers orthogonal ways of further increasing sampling efficiency. Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, Koray Kavukcuoglu |
ICML | 3 |
| 2018 | Parallel WaveNet: Fast High-Fidelity Speech SynthesisabstractThe recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today’s massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, a 1000x speed up relative to the original WaveNet, and capable of serving multiple English and Japanese voices in a production setting. Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche 0002, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Daniel Belov, Demis Hassabis |
ICML | 4 |
| 2018 | The challenge of realistic music generation: modelling raw audio at scaleabstractRealistic music generation is a challenging task. When building generative models of music that are learnt from data, typically high-level representations such as scores or MIDI are used that abstract away the idiosyncrasies of a particular performance. But these nuances are very important for our perception of musicality and realism, so in this work we embark on modelling music in the raw audio domain. It has been shown that autoregressive models excel at generating raw audio waveforms of speech, but when applied to music, we find them biased towards capturing local signal structure at the expense of modelling long-range correlations. This is problematic because music exhibits structure at many different timescales. In this work, we explore autoregressive discrete autoencoders (ADAs) as a means to enable autoregressive models to capture long-range correlations in waveforms. We find that they allow us to unconditionally generate piano music directly in the raw audio domain, which shows stylistic consistency across tens of seconds. Sander Dieleman, Aäron van den Oord, Karen Simonyan |
NeurIPS | 3 |
| 2018 | Learning to Navigate in Cities Without a MapabstractNavigating through unstructured environments is a basic capability of intelligent creatures, and thus is of fundamental interest in the study and development of artificial intelligence. Long-range navigation is a complex cognitive task that relies on developing an internal representation of space, grounded by recognisable landmarks and robust visual processing, that can simultaneously support continuous self-localisation ("I am here") and a representation of the goal ("I am going there"). Building upon recent research that applies deep reinforcement learning to maze navigation problems, we present an end-to-end deep reinforcement learning approach that can be applied on a city scale. Recognising that successful navigation relies on integration of general policies with locale-specific knowledge, we propose a dual pathway architecture that allows locale-specific features to be encapsulated, while still enabling transfer to multiple cities. A key contribution of this paper is an interactive navigation environment that uses Google Street View for its photographic content and worldwide coverage. Our baselines demonstrate that deep reinforcement learning agents can learn to navigate in multiple cities and to traverse to target destinations that may be kilometres away. A video summarizing our research and showing the trained agent in diverse city environments as well as on the transfer task is available at: https://sites.google.com/view/learn-navigate-cities-nips18 Piotr Mirowski, Matthew Koichi Grimes, Mateusz Malinowski, Karl Moritz Hermann, Keith Anderson, Denis Teplyashin, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, Raia Hadsell |
NeurIPS | 7 |
| 2017 | Neural Audio Synthesis of Musical Notes with WaveNet AutoencodersabstractGenerative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in both these areas to enable similar progress in audio modeling. First, we detail a powerful new WaveNet-style autoencoder model that conditions an autoregressive decoder on temporal codes learned from the raw audio waveform. Second, we introduce NSynth, a large-scale and high-quality dataset of musical notes that is an order of magnitude larger than comparable public datasets. Using NSynth, we demonstrate improved qualitative and quantitative performance of the WaveNet autoencoder over a well-tuned spectral autoencoder baseline. Finally, we show that the model learns a manifold of embeddings that allows for morphing between instruments, meaningfully interpolating in timbre to create new types of sounds that are realistic and expressive. Jesse H. Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi 0002, Douglas Eck, Karen Simonyan |
ICML | 7 |
| 2017 | Video Pixel NetworksabstractWe propose a probabilistic video model, the Video Pixel Network (VPN), that estimates the discrete joint distribution of the raw pixel values in a video. The model and the neural architecture reflect the time, space and color structure of video tensors and encode it as a four-dimensional dependency chain. The VPN approaches the best possible performance on the Moving MNIST benchmark, a leap over the previous state of the art, and the generated videos show only minor deviations from the ground truth. The VPN also produces detailed samples on the action-conditional Robotic Pushing benchmark and generalizes to the motion of novel objects. Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, Koray Kavukcuoglu |
ICML | 3 |
| 2016 | Reading Text in the Wild with Convolutional Neural Networks
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
Int. J. Comput. Vis. | 2 |
| 2015 | Natural Neural NetworksabstractWe introduce Natural Neural Networks, a novel family of algorithms that speed up convergence by adapting their internal representation during training to improve conditioning of the Fisher matrix. In particular, we show a specific example that employs a simple and efficient reparametrization of the neural network weights by implicitly whitening the representation obtained at each layer, while preserving the feed-forward computation of the network. Such networks can be trained efficiently via the proposed Projected Natural Gradient Descent algorithm (PRONG), which amortizes the cost of these reparametrizations over many parameter updates and is closely related to the Mirror Descent online learning algorithm. We highlight the benefits of our method on both unsupervised and supervised learning tasks, and showcase its scalability by training on the large-scale ImageNet Challenge dataset. Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, Koray Kavukcuoglu |
NIPS | 2 |
| 2015 | Spatial Transformer NetworksabstractConvolutional Neural Networks define an exceptionallypowerful class of model, but are still limited by the lack of abilityto be spatially invariant to the input data in a computationally and parameterefficient manner. In this work we introduce a new learnable module, theSpatial Transformer, which explicitly allows the spatial manipulation ofdata within the network. This differentiable module can be insertedinto existing convolutional architectures, giving neural networks the ability toactively spatially transform feature maps, conditional on the feature map itself,without any extra training supervision or modification to the optimisation process. We show that the useof spatial transformers results in models which learn invariance to translation,scale, rotation and more generic warping, resulting in state-of-the-artperformance on several benchmarks, and for a numberof classes of transformations. Max Jaderberg, Karen Simonyan, Andrew Zisserman, Koray Kavukcuoglu |
NIPS | 2 |
| 2014 | Efficient On-the-fly Category Retrieval Using ConvNets and GPUs
Ken Chatfield, Karen Simonyan, Andrew Zisserman |
ACCV (1) | 2 |
| 2014 | Deep Convolutional Neural Networks for Efficient Pose Estimation in Gesture Videos
Tomas Pfister, Karen Simonyan, James Charles, Andrew Zisserman |
ACCV (1) | 2 |
| 2014 | Return of the Devil in the Details: Delving Deep into Convolutional Nets
Ken Chatfield, Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
BMVC | 2 |
| 2014 | A Compact and Discriminative Face Track DescriptorabstractOur goal is to learn a compact, discriminative vector representation of a face track, suitable for the face recognition tasks of verification and classification. To this end, we propose a novel face track descriptor, based on the Fisher Vector representation, and demonstrate that it has a number of favourable properties. First, the descriptor is suitable for tracks of both frontal and profile faces, and is insensitive to their pose. Second, the descriptor is compact due to discriminative dimensionality reduction, and it can be further compressed using binarization. Third, the descriptor can be computed quickly (using hard quantization) and its compact size and fast computation render it very suitable for large scale visual repositories. Finally, the descriptor demonstrates good generalization when trained on one dataset and tested on another, reflecting its tolerance to the dataset bias. In the experiments we show that the descriptor exceeds the state of the art on both face verification task (YouTube Faces without outside training data, and INRIA-Buffy benchmarks), and face classification task (using the Oxford-Buffy dataset). Omkar M. Parkhi, Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
CVPR | 2 |
| 2014 | Understanding Objects in Detail with Fine-Grained AttributesabstractWe study the problem of understanding objects in detail, intended as recognizing a wide array of fine-grained object attributes. To this end, we introduce a dataset of 7, 413 airplanes annotated in detail with parts and their attributes, leveraging images donated by airplane spotters and crowd-sourcing both the design and collection of the detailed annotations. We provide a number of insights that should help researchers interested in designing fine-grained datasets for other basic level categories. We show that the collected data can be used to study the relation between part detection and attribute prediction by diagnosing the performance of classifiers that pool information from different parts of an object. We note that the prediction of certain attributes can benefit substantially from accurate part detection. We also show that, differently from previous results in object detection, employing a large number of part templates can improve detection accuracy at the expenses of detection speed. We finally propose a coarse-to-fine approach to speed up detection through a hierarchical cascade algorithm. Andrea Vedaldi, Siddharth Mahendran, Stavros Tsogkas, Subhransu Maji, Ross B. Girshick, Juho Kannala, Esa Rahtu, Iasonas Kokkinos, Matthew B. Blaschko, David J. Weiss, Ben Taskar, Karen Simonyan, Naomi Saphra, Sammy Mohamed |
CVPR | 12 |
| 2014 | Two-Stream Convolutional Networks for Action Recognition in Videos
Karen Simonyan, Andrew Zisserman |
NIPS | 1 |
| 2014 | Learning Local Feature Descriptors Using Convex OptimisationabstractThe objective of this work is to learn descriptors suitable for the sparse feature detectors used in viewpoint invariant matching. We make a number of novel contributions towards this goal. First, it is shown that learning the pooling regions for the descriptor can be formulated as a convex optimisation problem selecting the regions using sparsity. Second, it is shown that descriptor dimensionality reduction can also be formulated as a convex optimisation problem, using Mahalanobis matrix nuclear norm regularisation. Both formulations are based on discriminative large margin learning constraints. As the third contribution, we evaluate the performance of the compressed descriptors, obtained from the learnt real-valued descriptors by binarisation. Finally, we propose an extension of our learning formulations to a weakly supervised case, which allows us to learn the descriptors from unannotated image collections. It is demonstrated that the new learning methods improve over the state of the art in descriptor learning on the annotated local patches data set of Brown et al. and unannotated photo collections of Philbin et al. Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Fisher Vector Faces in the WildabstractSeveral recent papers on automatic face verification have significantly raised the performance bar by developing novel, specialised representations that outperform standard features such as SIFT for this problem. This paper makes two contributions: first, and somewhat surprisingly, we show that Fisher vectors on densely sampled SIFT features, i.e. an off-the-shelf object recognition representation, are capable of achieving state-of-the-art face verification performance on the challenging “Labeled Faces in the Wild” benchmark; second, since Fisher vectors are very high dimensional, we show that a compact descriptor can be learnt from them using discriminative metric learning. This compact descriptor has a better recognition accuracy and is very well suited to large scale identification tasks. Karen Simonyan, Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman |
BMVC | 1 |
| 2013 | Deep Fisher Networks for Large-Scale Image ClassificationabstractAs massively parallel computations have become broadly available with modern GPUs, deep architectures trained on very large datasets have risen in popularity. Discriminatively trained convolutional neural networks, in particular, were recently shown to yield state-of-the-art performance in challenging image classification benchmarks such as ImageNet. However, elements of these architectures are similar to standard hand-crafted representations used in computer vision. In this paper, we explore the extent of this analogy, proposing a version of the state-of-the-art Fisher vector image encoding that can be stacked in multiple layers. This architecture significantly improves on standard Fisher vectors, and obtains competitive results with deep convolutional networks at a significantly smaller computational cost. Our hybrid architecture allows us to measure the performance improvement brought by a deeper image classification pipeline, while staying in the realms of conventional SIFT features and FV encodings. Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
NIPS | 1 |
| 2012 | Descriptor Learning Using Convex Optimisation
Karen Simonyan, Andrea Vedaldi, Andrew Zisserman |
ECCV (1) | 1 |
| 2011 | Immediate Structured Visual Search for Medical Images
Karen Simonyan, Andrew Zisserman, Antonio Criminisi |
MICCAI (3) | 1 |
| 2009 | Edge-Directed Interpolation in a Bayesian FrameworkabstractIn this paper we present a novel framework for Edge-Directed Interpolation (EDI) of still images. The problem is treated as finding maximum a posteriori estimates of each interpolated pixel type and intensity value. The pixel type may be one of the pre-defined edge directions or "non-edge". Instead of the separate steps of edge orientation detection and intensity interpolation, maximizing the joint probability density function of type and intensity provides a better fit to the local image structure. Such a technique allows an effective discrimination between edges and non-edges (uniform areas and texture), thus leading to the suppression of artifacts which are common to existing EDI methods. Objective and subjective comparisons with conventional EDI methods corroborate the advantages of the proposed one. Moreover, the locality and the low computational complexity of the method make it suitable for a hardware implementation. 1 Karen Simonyan, Dmitriy S. Vatolin |
BMVC | 1 |
| 2008 | Fast video super-resolution via classificationabstractIn this paper we propose a novel super-resolution algorithm based on motion compensation and edge-directed spatial interpolation succeeded by fusion via pixel classification. Two high-resolution images are constructed, the first by means of motion compensation and the second by means of edge-directed interpolation. The AdaBoost classifier is then used to fuse these images into an high-resolution frame. Experimental results show that the proposed method surpasses well-known resolution enhancement methods while maintaining moderate computational complexity. Karen Simonyan, Sergey Grishin, Dmitriy S. Vatolin, Dmitriy Popov |
ICIP | 1 |