Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Aäron van den Oord

dblp:118/3205 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
7since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
27 papers
Generative modeling · 26% Representation and self-supervised learning · 18% Speech recognition and synthesis · 10%
Computer graphics and multimedia
4 papers
Audio and music processing · 51% Visual content generation and editing · 30% Image and video coding · 20%

Topics — the 30 heaviest of 68, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
1.242018
The challenge of realistic music generation: modelling raw audio at scale · NeurIPS 2018
Parallel Multiscale Autoregressive Density Estimation · ICML 2017
Video Pixel Networks · ICML 2017
Machine learning › Representation and self-supervised learning
contrastive learning
1.022021
Divide and Contrast: Self-supervised Learning from Uncurated Data · ICCV 2021
Efficient Visual Pretraining with Contrastive Detection · ICCV 2021
Machine learning › Reinforcement learning
model-based reinforcement learning
0.922021
Vector Quantized Models for Planning · ICML 2021
Shaping Belief States with Generative Environment Models for RL · NeurIPS 2019
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.822021
Self-supervised Adversarial Robustness for the Low-label, High-data Regime · ICLR 2021
Adversarial Risk and the Dangers of Evaluating Against Weak Attacks · ICML 2018
Machine learning › Generative modeling
variational autoencoder
0.822019
Generating Diverse High-Fidelity Images with VQ-VAE-2 · NeurIPS 2019
Preventing Posterior Collapse with delta-VAEs · ICLR (Poster) 2019
Machine learning › Generative modeling › autoregressive model
autoregressive prior
0.722019
Generating Diverse High-Fidelity Images with VQ-VAE-2 · NeurIPS 2019
Neural Discrete Representation Learning · NIPS 2017
Machine learning › Generative modeling › variational autoencoder
vector-quantized variational autoencoder
0.722019
Generating Diverse High-Fidelity Images with VQ-VAE-2 · NeurIPS 2019
Neural Discrete Representation Learning · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.622018
Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018
Parallel Multiscale Autoregressive Density Estimation · ICML 2017
Machine learning › Generative modeling
diffusion model
0.612022
Step-unrolled Denoising Autoencoders for Text Generation · ICLR 2022
Natural language and speech › Language models and text generation
text generation
0.612022
Step-unrolled Denoising Autoencoders for Text Generation · ICLR 2022
Machine learning › Efficient and distributed learning › efficient training
efficient pre-training
0.512021
Efficient Visual Pretraining with Contrastive Detection · ICCV 2021
Machine learning › Representation and self-supervised learning › contrastive learning › negative sampling
hard negative mining
0.512021
Divide and Contrast: Self-supervised Learning from Uncurated Data · ICCV 2021
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
multimodal self-supervised learning
0.512021
Broaden Your Views for Self-Supervised Video Learning · ICCV 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Self-supervised Adversarial Robustness for the Low-label, High-data Regime · ICLR 2021
Machine learning › Representation and self-supervised learning
vector quantization
0.512021
Vector Quantized Models for Planning · ICML 2021
Computer vision › Video understanding and tracking
video representation learning
0.512021
Broaden Your Views for Self-Supervised Video Learning · ICCV 2021
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning › deep representation learning
autoencoder representation learning
0.412019
Unsupervised Speech Representation Learning Using WaveNet Autoencoders · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty › partially observable markov decision process
belief state learning
0.412019
Shaping Belief States with Generative Environment Models for RL · NeurIPS 2019
Machine learning › Reinforcement learning › model-based reinforcement learning
environment model
0.412019
Shaping Belief States with Generative Environment Models for RL · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
evidence lower bound
0.412019
On Variational Bounds of Mutual Information · ICML 2019
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation
0.412019
On Variational Bounds of Mutual Information · ICML 2019
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412019
Wasserstein Dependency Measure for Representation Learning · NeurIPS 2019
Machine learning › Generative modeling › variational autoencoder
posterior collapse
0.412019
Preventing Posterior Collapse with delta-VAEs · ICLR (Poster) 2019
Natural language and speech › Speech recognition and synthesis
speech representation learning
0.412019
Unsupervised Speech Representation Learning Using WaveNet Autoencoders · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech
0.412019
Sample Efficient Adaptive Text-to-Speech · ICLR (Poster) 2019
Machine learning › Learning paradigms › unsupervised learning
unsupervised speech representation
0.412019
Unsupervised Speech Representation Learning Using WaveNet Autoencoders · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Machine learning › Generative modeling › vector-quantized generative models
vector-quantized autoencoder
0.412019
Unsupervised Speech Representation Learning Using WaveNet Autoencoders · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Machine learning › Trustworthy machine learning › robustness evaluation
adversarial attack evaluation
0.312018
Adversarial Risk and the Dangers of Evaluating Against Weak Attacks · ICML 2018
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial risk
0.312018
Adversarial Risk and the Dangers of Evaluating Against Weak Attacks · ICML 2018
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.312018
Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018

Methods — techniques the papers use, named apart from their topics

autoregressive modeling · 1.9contrastive learning · 1.0vector quantization · 0.7variational autoencoder · 0.7recurrent neural network · 0.7step unrolling · 0.6denoising autoencoder · 0.6self-supervised pretraining · 0.5multi-view learning · 0.5contrastive detection · 0.5variational bounds · 0.4neural network parameterization · 0.4deep convolutional neural network · 0.3collaborative filtering · 0.3bag-of-words · 0.3autoregressive discrete autoencoders · 0.3probabilistic video model · 0.3four-dimensional dependency chain · 0.3
YearPublicationVenuePosition
2022 Towards Learning Universal Audio Representations
abstract
The ability to learn universal audio representations that can solve diverse speech, music, and environment tasks can spur many applications that require general sound content understanding. In this work, we introduce a holistic audio representation evaluation suite (HARES) spanning 12 downstream tasks across audio domains and provide a thorough empirical study of recent sound representation learning systems on that benchmark. We discover that previous sound event classification or speech models do not generalize outside of their domains. We observe that more robust audio representations can be learned with the SimCLR objective; however, the model’s transferability depends heavily on the model architecture. We find the Slowfast architecture is good at learning rich representations required by different domains, but its performance is affected by the normalization scheme. Based on these findings, we propose a novel normalizer-free Slowfast NFNet and achieve state-of-the-art performance across all domains.
Luyu Wang, Pauline Luc, Yan Wu 0010, Adrià Recasens, Lucas Smaira, Andrew Brock, Andrew Jaegle, Jean-Baptiste Alayrac, Sander Dieleman, João Carreira 0001, Aäron van den Oord
ICASSP11
2022 Step-unrolled Denoising Autoencoders for Text Generation
Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, Aäron van den Oord
ICLR5
2021 Efficient Visual Pretraining with Contrastive Detection
abstract
Self-supervised pretraining has been shown to yield powerful representations for transfer learning. These performance gains come at a large computational cost however, with state-of-the-art methods requiring an order of magnitude more computation than supervised pretraining. We tackle this computational bottleneck by introducing a new self-supervised objective, contrastive detection, which tasks representations with identifying object-level features across augmentations. This objective extracts a rich learning signal per image, leading to state-of-the-art transfer accuracy on a variety of downstream tasks, while requiring up to 10× less pretraining. In particular, our strongest ImageNet-pretrained model performs on par with SEER, one of the largest self-supervised systems to date, which uses 1000× more pretraining data. Finally, our objective seamlessly handles pretraining on more complex images such as those in COCO, closing the gap with supervised transfer learning from COCO to PASCAL.
Olivier J. Hénaff, Skanda Koppula, Jean-Baptiste Alayrac, Aäron van den Oord, Oriol Vinyals, João Carreira 0001
ICCV4
2021 Broaden Your Views for Self-Supervised Video Learning
abstract
Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views are similarly extracted by cropping and augmenting the resulting crop. However, these methods miss a crucial element in the video domain: time. We introduce BraVe, a self-supervised learning framework for video. In BraVe, one of the views has access to a narrow temporal window of the video while the other view has a broad access to the video content. Our models learn to generalise from the narrow view to the general content of the video. Furthermore, BraVe processes the views with different backbones, enabling the use of alternative augmentations or modalities into the broad view such as optical flow, randomly convolved RGB frames, audio or their combinations. We demonstrate that BraVe achieves state-of-the-art results in self-supervised representation learning on standard video and audio classification benchmarks including UCF101, HMDB51, Kinetics, ESC-50 and AudioSet.
Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altché, Michal Valko, Jean-Bastien Grill, Aäron van den Oord, Andrew Zisserman
ICCV12
2021 Divide and Contrast: Self-supervised Learning from Uncurated Data
abstract
Self-supervised learning holds promise in leveraging large amounts of unlabeled data, however much of its progress has thus far been limited to highly curated pre-training data such as ImageNet. We explore the effects of contrastive learning from larger, less-curated image datasets such as YFCC, and find there is indeed a large difference in the resulting representation quality. We hypothesize that this curation gap is due to a shift in the distribution of image classes—which is more diverse and heavy-tailed—resulting in less relevant negative samples to learn from. We test this hypothesis with a new approach, Divide and Contrast (DnC), which alternates between contrastive learning and clustering-based hard negative mining. When pretrained on less curated datasets, DnC greatly improves the performance of self-supervised learning on downstream tasks, while remaining competitive with the current state-of-the-art on curated datasets.
Yonglong Tian, Olivier J. Hénaff, Aäron van den Oord
ICCV3
2021 Self-supervised Adversarial Robustness for the Low-label, High-data Regime
Sven Gowal, Po-Sen Huang, Aäron van den Oord, Timothy A. Mann, Pushmeet Kohli
ICLR3
2021 Vector Quantized Models for Planning
abstract
Recent developments in the field of model-based RL have proven successful in a range of environments, especially ones where planning is essential. However, such successes have been limited to deterministic fully-observed environments. We present a new approach that handles stochastic and partially-observable environments. Our key insight is to use discrete autoencoders to capture the multiple possible effects of an action in a stochastic environment. We use a stochastic variant of Monte Carlo tree search to plan over both the agent’s actions and the discrete latent variables representing the environment’s response. Our approach significantly outperforms an offline version of MuZero on a stochastic interpretation of chess where the opponent is considered part of the environment. We also show that our approach scales to DeepMind Lab, a first-person 3D environment with large visual observations and partial observability.
Sherjil Ozair, Yazhe Li, Ali Razavi, Ioannis Antonoglou, Aäron van den Oord, Oriol Vinyals
ICML5
2020 Contrastive Predictive Coding of Audio with an Adversary
Luyu Wang, Kazuya Kawakami, Aäron van den Oord
INTERSPEECH3
2019 Low Bit-rate Speech Coding with VQ-VAE and a WaveNet Decoder
abstract
In order to efficiently transmit and store speech signals, speech codecs create a minimally redundant representation of the input signal which is then decoded at the receiver with the best possible perceptual quality. In this work we demonstrate that a neural network architecture based on VQ-VAE with a WaveNet decoder can be used to perform very low bit-rate speech coding with high reconstruction quality. A prosody-transparent and speaker-independent model trained on the LibriSpeech corpus coding audio at 1.6 kbps exhibits perceptual quality which is around halfway between the MELP codec at 2.4 kbps and AMR-WB codec at 23.05 kbps. In addition, when training on high-quality recorded speech with the test speaker included in the training set, a model coding speech at 1.6 kbps produces output of similar perceptual quality to that generated by AMR-WB at 23.05 kbps.
Cristina Garbacea, Aäron van den Oord, Yazhe Li, Felicia Lim, Alejandro Luebs, Oriol Vinyals, Thomas C. Walters
ICASSP2
2019 Sample Efficient Adaptive Text-to-Speech
Yutian Chen 0001, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, Nando de Freitas
ICLR (Poster)12
2019 Preventing Posterior Collapse with delta-VAEs
Ali Razavi, Aäron van den Oord, Ben Poole, Oriol Vinyals
ICLR (Poster)2
2019 On Variational Bounds of Mutual Information
abstract
Estimating and optimizing Mutual Information (MI) is core to many problems in machine learning, but bounding MI in high dimensions is challenging. To establish tractable and scalable objectives, recent work has turned to variational bounds parameterized by neural networks. However, the relationships and tradeoffs between these bounds remains unclear. In this work, we unify these recent developments in a single framework. We find that the existing variational lower bounds degrade when the MI is large, exhibiting either high bias or high variance. To address this problem, we introduce a continuum of lower bounds that encompasses previous bounds and flexibly trades off bias and variance. On high-dimensional, controlled problems, we empirically characterize the bias and variance of the bounds and their gradients and demonstrate the effectiveness of these new bounds for estimation and representation learning.
Ben Poole, Sherjil Ozair, Aäron van den Oord, Alexander A. Alemi, George Tucker
ICML3
2019 Shaping Belief States with Generative Environment Models for RL
abstract
When agents interact with a complex environment, they must form and maintain beliefs about the relevant aspects of that environment. We propose a way to efficiently train expressive generative models in complex environments. We show that a predictive algorithm with an expressive generative model can form stable belief-states in visually rich and dynamic 3D environments. More precisely, we show that the learned representation captures the layout of the environment as well as the position and orientation of the agent. Our experiments show that the model substantially improves data-efficiency on a number of reinforcement learning (RL) tasks compared with strong model-free baseline agents. We find that predicting multiple steps into the future (overshooting), in combination with an expressive generative model, is critical for stable representations to emerge. In practice, using expressive generative models in RL is computationally expensive and we propose a scheme to reduce this computational burden, allowing us to build agents that are competitive with model-free baselines.
Karol Gregor, Danilo Jimenez Rezende, Frederic Besse, Yan Wu 0010, Hamza Merzic, Aäron van den Oord
NeurIPS6
2019 Wasserstein Dependency Measure for Representation Learning
abstract
Mutual information maximization has emerged as a powerful learning objective for unsupervised representation learning obtaining state-of-the-art performance in applications such as object recognition, speech recognition, and reinforcement learning. However, such approaches are fundamentally limited since a tight lower bound on mutual information requires sample size exponential in the mutual information. This limits the applicability of these approaches for prediction tasks with high mutual information, such as in video understanding or reinforcement learning. In these settings, such techniques are prone to overfit, both in theory and in practice, and capture only a few of the relevant factors of variation. This leads to incomplete representations that are not optimal for downstream tasks. In this work, we empirically demonstrate that mutual information-based representation learning approaches do fail to learn complete representations on a number of designed and real-world tasks. To mitigate these problems we introduce the Wasserstein dependency measure, which learns more complete representations by using the Wasserstein distance instead of the KL divergence in the mutual information estimator. We show that a practical approximation to this theoretically motivated solution, constructed using Lipschitz constraint techniques from the GAN literature, achieves substantially improved results on tasks where incomplete representations are a major challenge.
Sherjil Ozair, Corey Lynch, Yoshua Bengio, Aäron van den Oord, Sergey Levine, Pierre Sermanet
NeurIPS4
2019 Generating Diverse High-Fidelity Images with VQ-VAE-2
abstract
We explore the use of Vector Quantized Variational AutoEncoder (VQ-VAE) models for large scale image generation. To this end, we scale and enhance the autoregressive priors used in VQ-VAE to generate synthetic samples of much higher coherence and fidelity than possible before. We use simple feed-forward encoder and decoder networks, making our model an attractive candidate for applications where the encoding and/or decoding speed is critical. Additionally, VQ-VAE requires sampling an autoregressive model only in the compressed latent space, which is an order of magnitude faster than sampling in the pixel space, especially for large images. We demonstrate that a multi-scale hierarchical organization of VQ-VAE, augmented with powerful priors over the latent codes, is able to generate samples with quality that rivals that of state of the art Generative Adversarial Networks on multifaceted datasets such as ImageNet, while not suffering from GAN's known shortcomings such as mode collapse and lack of diversity.
Ali Razavi, Aäron van den Oord, Oriol Vinyals
NeurIPS2
2019 Unsupervised Speech Representation Learning Using WaveNet Autoencoders
abstract
We consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms. The goal is to learn a representation able to capture high level semantic content from the signal, e.g. phoneme identities, while being invariant to confounding low level details in the signal such as the underlying pitch contour or background noise. Since the learned representation is tuned to contain only phonetic content, we resort to using a high capacity WaveNet decoder to infer information discarded by the encoder from previous samples. Moreover, the behavior of autoencoder models depends on the kind of constraint that is applied to the latent representation. We compare three variants: a simple dimensionality reduction bottleneck, a Gaussian Variational Autoencoder (VAE), and a discrete Vector Quantized VAE (VQ-VAE). We analyze the quality of learned representations in terms of speaker independence, the ability to predict phonetic content, and the ability to accurately reconstruct individual spectrogram frames. Moreover, for discrete encodings extracted using the VQ-VAE, we measure the ease of mapping them to phonemes. We introduce a regularization scheme that forces the representations to focus on the phonetic content of the utterance and report performance comparable with the top entries in the ZeroSpeech 2017 unsupervised acoustic unit discovery task.
Jan Chorowski, Ron J. Weiss, Samy Bengio, Aäron van den Oord
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Temporal Modeling Using Dilated Convolution and Gating for Voice-Activity-Detection
abstract
Voice activity detection (VAD) is the task of predicting which parts of an utterance contains speech versus background noise. It is an important first step to determine which samples to send to the decoder and when to close the microphone. The long short-term memory neural network (LSTM) is a popular architecture for sequential modeling of acoustic signals, and has been successfully used in several VAD applications. However, it has been observed that LSTMs suffer from state saturation problems when the utterance is long (i.e., for voice dictation tasks), and thus requires the LSTM state to be periodically reset. In this paper, we propose an alternative architecture that does not suffer from saturation problems by modeling temporal variations through a stateless dilated convolution neural network (CNN). The proposed architecture differs from conventional CNNs in three respects: it uses dilated causal convolution, gated activations and residual connections. Results on a Google Voice Typing task shows that the proposed architecture achieves 14% relative FA improvement at a FR of 1% over state-of-the-art LSTMs for VAD task. We also include detailed experiments investigating the factors that distinguish the proposed architecture from conventional convolution.
Shuo-Yiin Chang, Bo Li 0028, Gabor Simko, Tara N. Sainath, Anshuman Tripathi, Aäron van den Oord, Oriol Vinyals
ICASSP6
2018 Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions
Scott E. Reed, Yutian Chen 0001, Thomas Paine, Aäron van den Oord, S. M. Ali Eslami, Danilo Jimenez Rezende, Oriol Vinyals, Nando de Freitas
ICLR (Poster)4
2018 Efficient Neural Audio Synthesis
abstract
Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating desired samples. Efficient sampling for this class of models at the cost of little to no loss in quality has however remained an elusive problem. With a focus on text-to-speech synthesis, we describe a set of general techniques for reducing sampling time while maintaining high output quality. We first describe a single-layer recurrent neural network, the WaveRNN, with a dual softmax layer that matches the quality of the state-of-the-art WaveNet model. The compact form of the network makes it possible to generate 24 kHz 16-bit audio 4 times faster than real time on a GPU. Secondly, we apply a weight pruning technique to reduce the number of weights in the WaveRNN. We find that, for a constant number of parameters, large sparse networks perform better than small dense networks and this relationship holds past sparsity levels of more than 96%. The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile phone CPU in real time. Finally, we describe a new dependency scheme for sampling that lets us trade a constant number of non-local, distant dependencies for the ability to generate samples in batches. The Batch WaveRNN produces 8 samples per step without loss of quality and offers orthogonal ways of further increasing sampling efficiency.
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, Koray Kavukcuoglu
ICML8
2018 Parallel WaveNet: Fast High-Fidelity Speech Synthesis
abstract
The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today’s massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, a 1000x speed up relative to the original WaveNet, and capable of serving multiple English and Japanese voices in a production setting.
Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche 0002, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Daniel Belov, Demis Hassabis
ICML1
2018 Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
abstract
This paper investigates recently proposed approaches for defending against adversarial examples and evaluating adversarial robustness. We motivate adversarial risk as an objective for achieving models robust to worst-case inputs. We then frame commonly used attacks and evaluation metrics as defining a tractable surrogate objective to the true adversarial risk. This suggests that models may optimize this surrogate rather than the true adversarial risk. We formalize this notion as obscurity to an adversary, and develop tools and heuristics for identifying obscured models and designing transparent models. We demonstrate that this is a significant problem in practice by repurposing gradient-free optimization techniques into adversarial attacks, which we use to decrease the accuracy of several recently proposed defenses to near zero. Our hope is that our formulations and results will help researchers to develop more powerful defenses.
Jonathan Uesato, Brendan O'Donoghue, Pushmeet Kohli, Aäron van den Oord
ICML4
2018 The challenge of realistic music generation: modelling raw audio at scale
abstract
Realistic music generation is a challenging task. When building generative models of music that are learnt from data, typically high-level representations such as scores or MIDI are used that abstract away the idiosyncrasies of a particular performance. But these nuances are very important for our perception of musicality and realism, so in this work we embark on modelling music in the raw audio domain. It has been shown that autoregressive models excel at generating raw audio waveforms of speech, but when applied to music, we find them biased towards capturing local signal structure at the expense of modelling long-range correlations. This is problematic because music exhibits structure at many different timescales. In this work, we explore autoregressive discrete autoencoders (ADAs) as a means to enable autoregressive models to capture long-range correlations in waveforms. We find that they allow us to unconditionally generate piano music directly in the raw audio domain, which shows stylistic consistency across tens of seconds.
Sander Dieleman, Aäron van den Oord, Karen Simonyan
NeurIPS2
2018 Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video
Lionel Pigou, Aäron van den Oord, Sander Dieleman, Mieke Van Herreweghe, Joni Dambre
Int. J. Comput. Vis.2
2017 Video Pixel Networks
abstract
We propose a probabilistic video model, the Video Pixel Network (VPN), that estimates the discrete joint distribution of the raw pixel values in a video. The model and the neural architecture reflect the time, space and color structure of video tensors and encode it as a four-dimensional dependency chain. The VPN approaches the best possible performance on the Moving MNIST benchmark, a leap over the previous state of the art, and the generated videos show only minor deviations from the ground truth. The VPN also produces detailed samples on the action-conditional Robotic Pushing benchmark and generalizes to the motion of novel objects.
Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, Koray Kavukcuoglu
ICML2
2017 Count-Based Exploration with Neural Density Models
abstract
Bellemare et al. (2016) introduced the notion of a pseudo-count, derived from a density model, to generalize count-based exploration to non-tabular reinforcement learning. This pseudo-count was used to generate an exploration bonus for a DQN agent and combined with a mixed Monte Carlo update was sufficient to achieve state of the art on the Atari 2600 game Montezuma’s Revenge. We consider two questions left open by their work: First, how important is the quality of the density model for exploration? Second, what role does the Monte Carlo update play in exploration? We answer the first question by demonstrating the use of PixelCNN, an advanced neural density model for images, to supply a pseudo-count. In particular, we examine the intrinsic difficulties in adapting Bellemare et al.’s approach when assumptions about the model are violated. The result is a more practical and general algorithm requiring no special apparatus. We combine PixelCNN pseudo-counts with different agent architectures to dramatically improve the state of the art on several hard Atari games. One surprising finding is that the mixed Monte Carlo update is a powerful facilitator of exploration in the sparsest of settings, including Montezuma’s Revenge.
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, Rémi Munos
ICML3
2017 Parallel Multiscale Autoregressive Density Estimation
abstract
PixelCNN achieves state-of-the-art results in density estimation for natural images. Although training is fast, inference is costly, requiring one network evaluation per pixel; O(N) for N pixels. This can be sped up by caching activations, but still involves generating each pixel sequentially. In this work, we propose a parallelized PixelCNN that allows more efficient inference by modeling certain pixel groups as conditionally independent. Our new PixelCNN model achieves competitive density estimation and orders of magnitude speedup – O(log N) sampling instead of O(N) – enabling the practical generation of 512x512 images. We evaluate the model on class-conditional image generation, text-to-image synthesis, and action-conditional video generation, showing that our model achieves the best results among non-pixel-autoregressive density models that allow efficient sampling.
Scott E. Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gomez Colmenarejo, Ziyu Wang 0001, Yutian Chen 0001, Daniel Belov, Nando de Freitas
ICML2
2017 Neural Discrete Representation Learning
abstract
Learning useful representations without supervision remains a key challenge in machine learning. In this paper, we propose a simple yet powerful generative model that learns such discrete representations. Our model, the Vector Quantised-Variational AutoEncoder (VQ-VAE), differs from VAEs in two key ways: the encoder network outputs discrete, rather than continuous, codes; and the prior is learnt rather than static. In order to learn a discrete latent representation, we incorporate ideas from vector quantisation (VQ). Using the VQ method allows the model to circumvent issues of ``posterior collapse'' -— where the latents are ignored when they are paired with a powerful autoregressive decoder -— typically observed in the VAE framework. Pairing these representations with an autoregressive prior, the model can generate high quality images, videos, and speech as well as doing high quality speaker conversion and unsupervised learning of phonemes, providing further evidence of the utility of the learnt representations.
Aäron van den Oord, Oriol Vinyals, Koray Kavukcuoglu
NIPS1
2016 Pixel Recurrent Neural Networks
abstract
Modeling the distribution of natural images is a landmark problem in unsupervised learning. This task requires an image model that is at once expressive, tractable and scalable. We present a deep neural network that sequentially predicts the pixels in an image along the two spatial dimensions. Our method models the discrete probability of the raw pixel values and encodes the complete set of dependencies in the image. Architectural novelties include fast two-dimensional recurrent layers and an effective use of residual connections in deep recurrent networks. We achieve log-likelihood scores on natural images that are considerably better than the previous state of the art. Our main results also provide benchmarks on the diverse ImageNet dataset. Samples generated from the model appear crisp, varied and globally coherent.
Aäron van den Oord, Nal Kalchbrenner, Koray Kavukcuoglu
ICML1
2016 Conditional Image Generation with PixelCNN Decoders
abstract
This work explores conditional image generation with a new image density model based on the PixelCNN architecture. The model can be conditioned on any vector, including descriptive labels or tags, or latent embeddings created by other networks. When conditioned on class labels from the ImageNet database, the model is able to generate diverse, realistic scenes representing distinct animals, objects, landscapes and structures. When conditioned on an embedding produced by a convolutional network given a single image of an unseen face, it generates a variety of new portraits of the same person with different facial expressions, poses and lighting conditions. We also show that conditional PixelCNN can serve as a powerful decoder in an image autoencoder. Additionally, the gated convolutional layers in the proposed model improve the log-likelihood of PixelCNN to match the state-of-the-art performance of PixelRNN on ImageNet, with greatly reduced computational cost.
Aäron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Koray Kavukcuoglu, Oriol Vinyals, Alex Graves
NIPS1
2014 Factoring Variations in Natural Images with Deep Gaussian Mixture Models
Aäron van den Oord, Benjamin Schrauwen
NIPS1
2014 The student-t mixture as a natural image patch prior with application to image compression
Aäron van den Oord, Benjamin Schrauwen
J. Mach. Learn. Res.1
2013 Deep content-based music recommendation
abstract
Automatic music recommendation has become an increasingly relevant problem in recent years, since a lot of music is now sold and consumed digitally. Most recommender systems rely on collaborative filtering. However, this approach suffers from the cold start problem: it fails when no usage data is available, so it is not effective for recommending new and unpopular songs. In this paper, we propose to use a latent factor model for recommendation, and predict the latent factors from music audio when they cannot be obtained from usage data. We compare a traditional approach using a bag-of-words representation of the audio signals with deep convolutional neural networks, and evaluate the predictions quantitatively and qualitatively on the Million Song Dataset. We show that using predicted latent factors produces sensible recommendations, despite the fact that there is a large semantic gap between the characteristics of a song that affect user preference and the corresponding audio signal. We also show that recent advances in deep learning translate very well to the music recommendation setting, with deep convolutional neural networks significantly outperforming the traditional approach.
Aäron van den Oord, Sander Dieleman, Benjamin Schrauwen
NIPS1