VLDB 2026 Research / reviewers in the wild / expert
Nando de Freitas
dblp:42/631
· DBLP profile ↗
89ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0003-4770-7217ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 84 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
64 papers |
Reinforcement learning · 32% Probabilistic and Bayesian machine learning · 21% Optimization for machine learning · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 30 heaviest of 131, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
meta-learning |
1.8 | 5 | 2022 | Towards Learning Universal Hyperparameter Optimizers with Transformers · NeurIPS 2022 Modular Meta-Learning with Shrinkage · NeurIPS 2020 Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018 |
Machine learning › Reinforcement learning
offline reinforcement learning |
1.4 | 3 | 2021 | Active Offline Policy Selection · NeurIPS 2021 RL Unplugged: A Collection of Benchmarks for Offline Reinforcement Learning · NeurIPS 2020 Critic Regularized Regression · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable regression |
1.1 | 2 | 2022 | On Instrumental Variable Regression for Deep Offline Policy Evaluation · J. Mach. Learn. Res. 2022 Learning Deep Features in Instrumental Variable Regression · ICLR 2021 |
Machine learning › Reinforcement learning
off-policy evaluation |
1.1 | 2 | 2022 | On Instrumental Variable Regression for Deep Offline Policy Evaluation · J. Mach. Learn. Res. 2022 Active Offline Policy Selection · NeurIPS 2021 |
Machine learning › Optimization for machine learning
learned optimizer |
0.8 | 3 | 2017 | Learned Optimizers that Scale and Generalize · ICML 2017 Learning to Learn without Gradient Descent by Gradient Descent · ICML 2017 Learning to learn by gradient descent by gradient descent · NIPS 2016 |
Machine learning › Reinforcement learning
exploration |
0.8 | 2 | 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems · ICLR 2020 Playing hard exploration games by watching YouTube · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › sequential latent variable model
latent action models |
0.8 | 1 | 2024 | Genie: Generative Interactive Environments · ICML 2024 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.8 | 1 | 2024 | Genie: Generative Interactive Environments · ICML 2024 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.7 | 2 | 2022 | Towards Learning Universal Hyperparameter Optimizers with Transformers · NeurIPS 2022 Bayesian Optimization in High Dimensions via Random Embeddings · IJCAI 2013 |
Knowledge, reasoning and agents › Multi-agent systems
emergent communication |
0.7 | 2 | 2019 | Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning · ICML 2019 Compositional Obverter Communication Learning from Raw Visual Input · ICLR (Poster) 2018 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.7 | 2 | 2020 | Critic Regularized Regression · NeurIPS 2020 Dueling Network Architectures for Deep Reinforcement Learning · ICML 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.7 | 4 | 2016 | Herded Gibbs Sampling · J. Mach. Learn. Res. 2016 Distributed Parameter Estimation in Probabilistic Graphical Models · NIPS 2014 Linear and Parallel Learning of Markov Random Fields · ICML 2014 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.7 | 2 | 2022 | Learning Deep Features in Instrumental Variable Regression · ICLR 2021 On Instrumental Variable Regression for Deep Offline Policy Evaluation · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.7 | 2 | 2019 | Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning · ICML 2019 Learning to Perform Physics Experiments via Deep Reinforcement Learning · ICLR (Poster) 2017 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.6 | 2 | 2019 | Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning · ICML 2019 Learning to Communicate with Deep Multi-Agent Reinforcement Learning · NIPS 2016 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.6 | 2 | 2018 | Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018 Parallel Multiscale Autoregressive Density Estimation · ICML 2017 |
Machine learning › Reinforcement learning
imitation learning |
0.6 | 2 | 2018 | Playing hard exploration games by watching YouTube · NeurIPS 2018 Robust Imitation of Diverse Behaviors · NIPS 2017 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.6 | 2 | 2018 | Learning Awareness Models · ICLR (Poster) 2018 Learning to Perform Physics Experiments via Deep Reinforcement Learning · ICLR (Poster) 2017 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.5 | 4 | 2016 | Herded Gibbs Sampling · J. Mach. Learn. Res. 2016 Adaptive Hamiltonian and Riemann Manifold Monte Carlo · ICML (3) 2013 Bayesian Policy Learning with Trans-Dimensional MCMC · NIPS 2007 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.5 | 3 | 2017 | Taking the Human Out of the Loop: A Review of Bayesian Optimization · Proc. IEEE 2016 Bayesian Optimization in High Dimensions via Random Embeddings · IJCAI 2013 Learning to Learn without Gradient Descent by Gradient Descent · ICML 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.4 | 2 | 2016 | Herded Gibbs Sampling · J. Mach. Learn. Res. 2016 Linear and Parallel Learning of Markov Random Fields · ICML 2014 |
Machine learning › Reinforcement learning › regularization for reinforcement learning
critic regularization |
0.4 | 1 | 2020 | Critic Regularized Regression · NeurIPS 2020 |
Machine learning › Reinforcement learning › exploration › exploration in markov decision processes
hard exploration |
0.4 | 1 | 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems · ICLR 2020 |
Robotics › Robot manipulation
learning from demonstration |
0.4 | 1 | 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems · ICLR 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2020 | Critic Regularized Regression · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.4 | 2 | 2018 | Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018 Learning to Recognize Objects with Little Supervision · Int. J. Comput. Vis. 2008 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 2 | 2015 | Deep Fried Convnets · ICCV 2015 Predicting Parameters in Deep Learning · NIPS 2013 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.4 | 1 | 2019 | Hyperbolic Attention Networks · ICLR (Poster) 2019 |
Machine learning › Deep learning architectures and training › attention mechanism
hyperbolic attention |
0.4 | 1 | 2019 | Hyperbolic Attention Networks · ICLR (Poster) 2019 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning |
0.4 | 1 | 2019 | Hyperbolic Attention Networks · ICLR (Poster) 2019 |
Methods — techniques the papers use, named apart from their topics
gaussian process · 1.3bayesian optimization · 1.1reinforcement learning · 1.1unsupervised learning · 0.8spatiotemporal tokenizer · 0.8autoregressive dynamics models · 0.8meta-learning · 0.7uncertainty estimation · 0.6transformer · 0.6AGMM · 0.6neural programmer-interpreter · 0.4alphazero · 0.4sufficient statistics · 0.2maximum likelihood estimation · 0.2marginalization · 0.1discrete choice model · 0.1active learning · 0.1stochastic approximation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Genie: Generative Interactive EnvironmentsabstractWe introduce Genie, the first *generative interactive environment* trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a *foundation world model*. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis *despite training without any ground-truth action labels* or other domain specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future. Jake Bruce, Michael Dennis 0001, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes 0001, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie C. Y. Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott E. Reed, Jingwei Zhang 0001, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh 0001, Tim Rocktäschel |
ICML | 23 |
| 2023 | Machine Learning for Ancient Languages: A SurveyabstractAbstract Ancient languages preserve the cultures and histories of the past. However, their study is fraught with difficulties, and experts must tackle a range of challenging text-based tasks, from deciphering lost languages to restoring damaged inscriptions, to determining the authorship of works of literature. Technological aids have long supported the study of ancient texts, but in recent years advances in artificial intelligence and machine learning have enabled analyses on a scale and in a detail that are reshaping the field of humanities, similarly to how microscopes and telescopes have contributed to the realm of science. This article aims to provide a comprehensive survey of published research using machine learning for the study of ancient texts written in any language, script, and medium, spanning over three and a half millennia of civilizations around the ancient world. To analyze the relevant literature, we introduce a taxonomy of tasks inspired by the steps involved in the study of ancient documents: digitization, restoration, attribution, linguistic analysis, textual criticism, translation, and decipherment. This work offers three major contributions: first, mapping the interdisciplinary field carved out by the synergy between the humanities and machine learning; second, highlighting how active collaboration between specialists from both fields is key to producing impactful and compelling scholarship; third, highlighting promising directions for future work in this field. Thus, this work promotes and supports the continued collaborative impetus between the humanities and machine learning. Thea Sommerschield, Yannis M. Assael, John Pavlopoulos, Vanessa Stefanak, Andrew W. Senior, Chris Dyer, John Bodel, Jonathan Prag, Ion Androutsopoulos, Nando de Freitas |
Comput. Linguistics | 10 |
| 2022 | Towards Learning Universal Hyperparameter Optimizers with TransformersabstractMeta-learning hyperparameter optimization (HPO) algorithms from prior experiments is a promising approach to improve optimization efficiency over objective functions from a similar distribution. However, existing methods are restricted to learning from experiments sharing the same set of hyperparameters. In this paper, we introduce the OptFormer, the first text-based Transformer HPO framework that provides a universal end-to-end interface for jointly learning policy and function prediction when trained on vast tuning data from the wild, such as Google’s Vizier database, one of the world’s largest HPO datasets. Our extensive experiments demonstrate that the OptFormer can simultaneously imitate at least 7 different HPO algorithms, which can be further improved via its function uncertainty estimates. Compared to a Gaussian Process, the OptFormer also learns a robust prior distribution for hyperparameter response functions, and can thereby provide more accurate and better calibrated predictions. This work paves the path to future extensions for training a Transformer-based model as a general HPO optimizer. Yutian Chen 0001, Xingyou Song, Chansoo Lee, Qiuyi Zhang 0001, David Dohan, Kazuya Kawakami, Greg Kochanski, Arnaud Doucet, Marc'Aurelio Ranzato, Sagi Perel, Nando de Freitas |
NeurIPS | 12 |
| 2022 | On Instrumental Variable Regression for Deep Offline Policy EvaluationabstractWe show that the popular reinforcement learning (RL) strategy of estimating the state-action value (Q-function) by minimizing the mean squared Bellman error leads to a regression problem with confounding, the inputs and output noise being correlated. Hence, direct minimization of the Bellman error can result in significantly biased Q-function estimates. We explain why fixing the target Q-network in Deep Q-Networks and Fitted Q Evaluation provides a way of overcoming this confounding, thus shedding new light on this popular but not well understood trick in the deep RL literature. An alternative approach to address confounding is to leverage techniques developed in the causality literature, notably instrumental variables (IV). We bring together here the literature on IV and RL by investigating whether IV approaches can lead to improved Q-function estimates. This paper analyzes and compares a wide range of recent IV methods in the context of offline policy evaluation (OPE), where the goal is to estimate the value of a policy using logged data only. By applying different IV techniques to OPE, we are not only able to recover previously proposed OPE methods such as model-based techniques but also to obtain competitive new techniques. We find empirically that state-of-the-art OPE methods are closely matched in performance by some IV methods such as AGMM, which were not developed for OPE. We open-source all our code and datasets at https://github.com/liyuan9988/IVOPEwithACME. Yutian Chen 0001, Liyuan Xu, Caglar Gulcehre, Tom Le Paine, Arthur Gretton, Nando de Freitas, Arnaud Doucet |
J. Mach. Learn. Res. | 6 |
| 2021 | Learning Deep Features in Instrumental Variable Regression
Liyuan Xu, Yutian Chen 0001, Siddarth Srinivasan, Nando de Freitas, Arnaud Doucet, Arthur Gretton |
ICLR | 4 |
| 2021 | Active Offline Policy SelectionabstractThis paper addresses the problem of policy selection in domains with abundant logged data, but with a restricted interaction budget. Solving this problem would enable safe evaluation and deployment of offline reinforcement learning policies in industry, robotics, and recommendation domains among others. Several off-policy evaluation (OPE) techniques have been proposed to assess the value of policies using only logged data. However, there is still a big gap between the evaluation by OPE and the full online evaluation in the real environment. Yet, large amounts of online interactions are often not possible in practice. To overcome this problem, we introduce active offline policy selection --- a novel sequential decision approach that combines logged data with online interaction to identify the best policy. This approach uses OPE estimates to warm start the online evaluation. Then, in order to utilize the limited environment interactions wisely we decide which policy to evaluate next based on a Bayesian optimization method with a kernel function that represents policy similarity. We use multiple benchmarks with a large number of candidate policies to show that the proposed approach improves upon state-of-the-art OPE estimates and pure online policy evaluation. Ksenia Konyushkova, Yutian Chen 0001, Thomas Paine, Caglar Gulcehre, Cosmin Paduraru, Daniel J. Mankowitz, Misha Denil, Nando de Freitas |
NeurIPS | 8 |
| 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Caglar Gulcehre, Tom Le Paine, Bobak Shahriari, Misha Denil, Matt Hoffman 0001, Hubert Soyer, Richard Tanburn, Steven Kapturowski, Neil C. Rabinowitz, Duncan Williams, Gabriel Barth-Maron, Ziyu Wang 0001, Nando de Freitas |
ICLR | 13 |
| 2020 | Critic Regularized RegressionabstractOffline reinforcement learning (RL), also known as batch RL, offers the prospect of policy optimization from large pre-recorded datasets without online environment interaction. It addresses challenges with regard to the cost of data collection and safety, both of which are particularly pertinent to real-world applications of RL. Unfortunately, most off-policy algorithms perform poorly when learning from a fixed dataset. In this paper, we propose a novel offline RL algorithm to learn policies from data using a form of critic-regularized regression (CRR). We find that CRR performs surprisingly well and scales to tasks with high-dimensional state and action spaces -- outperforming several state-of-the-art offline RL algorithms by a significant margin on a wide range of benchmark tasks. Ziyu Wang 0001, Alexander Novikov 0001, Konrad Zolna, Josh Merel, Jost Tobias Springenberg, Scott E. Reed, Bobak Shahriari, Noah Y. Siegel, Caglar Gulcehre, Nicolas Heess, Nando de Freitas |
NeurIPS | 11 |
| 2020 | Modular Meta-Learning with ShrinkageabstractMany real-world problems, including multi-speaker text-to-speech synthesis, can greatly benefit from the ability to meta-learn large models with only a few task- specific components. Updating only these task-specific modules then allows the model to be adapted to low-data tasks for as many steps as necessary without risking overfitting. Unfortunately, existing meta-learning methods either do not scale to long adaptation or else rely on handcrafted task-specific architectures. Here, we propose a meta-learning approach that obviates the need for this often sub-optimal hand-selection. In particular, we develop general techniques based on Bayesian shrinkage to automatically discover and learn both task-specific and general reusable modules. Empirically, we demonstrate that our method discovers a small set of meaningful task-specific modules and outperforms existing meta- learning approaches in domains like few-shot text-to-speech that have little task data and long adaptation horizons. We also show that existing meta-learning methods including MAML, iMAML, and Reptile emerge as special cases of our method. Yutian Chen 0001, Abram L. Friesen, Feryal M. P. Behbahani, Arnaud Doucet, David Budden, Matt Hoffman 0001, Nando de Freitas |
NeurIPS | 7 |
| 2020 | RL Unplugged: A Collection of Benchmarks for Offline Reinforcement Learning
Caglar Gulcehre, Ziyu Wang 0001, Alexander Novikov 0001, Thomas Paine, Sergio Gomez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel J. Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Mohammad Norouzi 0002, Matt Hoffman 0001, Nicolas Heess, Nando de Freitas |
NeurIPS | 16 |
| 2019 | Sample Efficient Adaptive Text-to-Speech
Yutian Chen 0001, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, Nando de Freitas |
ICLR (Poster) | 14 |
| 2019 | Hyperbolic Attention Networks
Caglar Gulcehre, Misha Denil, Mateusz Malinowski, Ali Razavi, Razvan Pascanu, Karl Moritz Hermann, Peter W. Battaglia, Victor Bapst, David Raposo, Adam Santoro, Nando de Freitas |
ICLR (Poster) | 11 |
| 2019 | Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement LearningabstractWe propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents’ actions. Causal influence is assessed using counterfactual reasoning. At each timestep, an agent simulates alternate actions that it could have taken, and computes their effect on the behavior of other agents. Actions that lead to bigger changes in other agents’ behavior are considered influential and are rewarded. We show that this is equivalent to rewarding agents for having high mutual information between their actions. Empirical results demonstrate that influence leads to enhanced coordination and communication in challenging social dilemma environments, dramatically increasing the learning curves of the deep RL agents, and leading to more meaningful learned communication protocols. The influence rewards for all agents can be computed in a decentralized way by enabling agents to learn a model of other agents using deep neural networks. In contrast, key previous works on emergent communication in the MARL setting were unable to learn diverse policies in a decentralized manner and had to resort to centralized training. Consequently, the influence reward opens up a window of new opportunities for research in this area. Natasha Jaques, Angeliki Lazaridou, Edward Hughes 0001, Caglar Gulcehre, Pedro A. Ortega, DJ Strouse, Joel Z. Leibo, Nando de Freitas |
ICML | 8 |
| 2019 | Large-Scale Visual Speech RecognitionabstractThis work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking (3,886 hours of video). In tandem, we designed and trained an integrated lipreading system, consisting of a video processing pipeline that maps raw video to stable videos of lips and sequences of phonemes, a scalable deep neural network that maps the lip videos to sequences of phoneme distributions, and a production-level speech decoder that outputs sequences of words. The proposed system achieves a word error rate (WER) of 40.9% as measured on a held-out set. In comparison, professional lipreaders achieve either 86.4% or 92.9% WER on the same dataset when having access to additional types of contextual information. Our approach significantly improves on other lipreading approaches, including variants of LipNet and of Watch, Attend, and Spell (WAS), which are only capable of 89.8% and 76.8% WER respectively. Brendan Shillingford, Yannis M. Assael, Matt Hoffman 0001, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Misha Denil, Ben Coppin, Ben Laurie, Andrew W. Senior, Nando de Freitas |
INTERSPEECH | 16 |
| 2019 | Learning Compositional Neural Programs with Recursive Tree Search and PlanningabstractWe propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning. The experiments show that AlphaNPI can sort as well as previous strongly supervised NPI variants. The AlphaNPI agent is also trained on a Tower of Hanoi puzzle with two disks and is shown to generalize to puzzles with an arbitrary number of disks. The experiments also show that when deploying our neural network policies, it is advantageous to do planning with guided Monte Carlo tree search. Thomas Pierrot, Guillaume Ligner, Scott E. Reed, Olivier Sigaud, Nicolas Perrin-Gilbert, Alexandre Laterre, David Kas, Karim Beguir, Nando de Freitas |
NeurIPS | 9 |
| 2018 | Learning Awareness Models
Brandon Amos, Laurent Dinh, Serkan Cabi, Thomas Rothörl, Sergio Gomez Colmenarejo, Alistair Muldal, Tom Erez, Yuval Tassa, Nando de Freitas, Misha Denil |
ICLR (Poster) | 9 |
| 2018 | Compositional Obverter Communication Learning from Raw Visual Input
Edward Choi 0003, Angeliki Lazaridou, Nando de Freitas |
ICLR (Poster) | 3 |
| 2018 | Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions
Scott E. Reed, Yutian Chen 0001, Thomas Paine, Aäron van den Oord, S. M. Ali Eslami, Danilo Jimenez Rezende, Oriol Vinyals, Nando de Freitas |
ICLR (Poster) | 8 |
| 2018 | Playing hard exploration games by watching YouTubeabstractDeep reinforcement learning methods traditionally struggle with tasks where environment rewards are particularly sparse. One successful method of guiding exploration in these domains is to imitate trajectories provided by a human demonstrator. However, these demonstrations are typically collected under artificial conditions, i.e. with access to the agent’s exact environment setup and the demonstrator’s action and reward trajectories. Here we propose a method that overcomes these limitations in two stages. First, we learn to map unaligned videos from multiple sources to a common representation using self-supervised objectives constructed over both time and modality (i.e. vision and sound). Second, we embed a single YouTube video in this representation to learn a reward function that encourages an agent to imitate human gameplay. This method of one-shot imitation allows our agent to convincingly exceed human-level performance on the infamously hard exploration games Montezuma’s Revenge, Pitfall! and Private Eye for the first time, even if the agent is not presented with any environment rewards. Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang 0001, Nando de Freitas |
NeurIPS | 6 |
| 2017 | Sample Efficient Actor-Critic with Experience Replay
Ziyu Wang 0001, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Rémi Munos, Koray Kavukcuoglu, Nando de Freitas |
ICLR (Poster) | 7 |
| 2017 | Learning to Perform Physics Experiments via Deep Reinforcement Learning
Misha Denil, Pulkit Agrawal 0001, Tejas D. Kulkarni, Tom Erez, Peter W. Battaglia, Nando de Freitas |
ICLR (Poster) | 6 |
| 2017 | Learning to Learn without Gradient Descent by Gradient DescentabstractWe learn recurrent neural network optimizers trained on simple synthetic functions by gradient descent. We show that these learned optimizers exhibit a remarkable degree of transfer in that they can be used to efficiently optimize a broad range of derivative-free black-box functions, including Gaussian process bandits, simple control objectives, global optimization benchmarks and hyper-parameter tuning tasks. Up to the training horizon, the learned optimizers learn to trade-off exploration and exploitation, and compare favourably with heavily engineered Bayesian optimization packages for hyper-parameter tuning. Yutian Chen 0001, Matt Hoffman 0001, Sergio Gomez Colmenarejo, Misha Denil, Timothy P. Lillicrap, Matt M. Botvinick, Nando de Freitas |
ICML | 7 |
| 2017 | Parallel Multiscale Autoregressive Density EstimationabstractPixelCNN achieves state-of-the-art results in density estimation for natural images. Although training is fast, inference is costly, requiring one network evaluation per pixel; O(N) for N pixels. This can be sped up by caching activations, but still involves generating each pixel sequentially. In this work, we propose a parallelized PixelCNN that allows more efficient inference by modeling certain pixel groups as conditionally independent. Our new PixelCNN model achieves competitive density estimation and orders of magnitude speedup – O(log N) sampling instead of O(N) – enabling the practical generation of 512x512 images. We evaluate the model on class-conditional image generation, text-to-image synthesis, and action-conditional video generation, showing that our model achieves the best results among non-pixel-autoregressive density models that allow efficient sampling. Scott E. Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gomez Colmenarejo, Ziyu Wang 0001, Yutian Chen 0001, Daniel Belov, Nando de Freitas |
ICML | 8 |
| 2017 | Learned Optimizers that Scale and GeneralizeabstractLearning to learn has emerged as an important direction for achieving artificial intelligence. Two of the primary barriers to its adoption are an inability to scale to larger problems and a limited ability to generalize to new tasks. We introduce a learned gradient descent optimizer that generalizes well to new tasks, and which has significantly reduced memory and computation overhead. We achieve this by introducing a novel hierarchical RNN architecture, with minimal per-parameter overhead, augmented with additional architectural features that mirror the known structure of optimization tasks. We also develop a meta-training ensemble of small, diverse, optimization tasks capturing common properties of loss landscapes. The optimizer learns to outperform RMSProp/ADAM on problems in this corpus. More importantly, it performs comparably or better when applied to small convolutional neural networks, despite seeing no neural networks in its meta-training set. Finally, it generalizes to train Inception V3 and ResNet V2 architectures on the ImageNet dataset for thousands of steps, optimization problems that are of a vastly different scale than those it was trained on. Olga Wichrowska, Niru Maheswaranathan, Matt Hoffman 0001, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, Jascha Sohl-Dickstein |
ICML | 6 |
| 2017 | Robust Imitation of Diverse BehaviorsabstractDeep generative models have recently shown great promise in imitation learning for motor control. Given enough data, even supervised approaches can do one-shot imitation learning; however, they are vulnerable to cascading failures when the agent trajectory diverges from the demonstrations. Compared to purely supervised methods, Generative Adversarial Imitation Learning (GAIL) can learn more robust controllers from fewer demonstrations, but is inherently mode-seeking and more difficult to train. In this paper, we show how to combine the favourable aspects of these two approaches. The base of our model is a new type of variational autoencoder on demonstration trajectories that learns semantic policy embeddings. We show that these embeddings can be learned on a 9 DoF Jaco robot arm in reaching tasks, and then smoothly interpolated with a resulting smooth interpolation of reaching behavior. Leveraging these policy representations, we develop a new version of GAIL that (1) is much more robust than the purely-supervised controller, especially with few demonstrations, and (2) avoids mode collapse, capturing many diverse behaviors when GAIL on its own does not. We demonstrate our approach on learning diverse gaits from demonstration on a 2D biped and a 62 DoF 3D humanoid in the MuJoCo physics environment. Ziyu Wang 0001, Josh Merel, Scott E. Reed, Nando de Freitas, Greg Wayne, Nicolas Heess |
NIPS | 4 |
| 2017 | Cortical microcircuits as gated-recurrent neural networksabstractCortical circuits exhibit intricate recurrent architectures that are remarkably similar across different brain areas. Such stereotyped structure suggests the existence of common computational principles. However, such principles have remained largely elusive. Inspired by gated-memory networks, namely long short-term memory networks (LSTMs), we introduce a recurrent neural network in which information is gated through inhibitory cells that are subtractive (subLSTM). We propose a natural mapping of subLSTMs onto known canonical excitatory-inhibitory cortical microcircuits. Our empirical evaluation across sequential image classification and language modelling tasks shows that subLSTM units can achieve similar performance to LSTM units. These results suggest that cortical circuits can be optimised to solve complex contextual problems and proposes a novel view on their computational function. Overall our work provides a step towards unifying recurrent networks as used in machine learning with their biological counterparts. Rui Ponte Costa, Yannis M. Assael, Brendan Shillingford, Nando de Freitas, Tim P. Vogels |
NIPS | 4 |
| 2016 | Unbounded Bayesian Optimization via RegularizationabstractBayesian optimization has recently emerged as a powerful and flexible tool in machine learning for hyperparameter tuning and more generally for the efficient global optimization of expensive black box functions. The established practice requires a user-defined bounded domain, which is assumed to contain the global optimizer. However, when little is known about the probed objective function, it can be difficult to prescribe such a domain. In this work, we modify the standard Bayesian optimization framework in a principled way to allow for unconstrained exploration of the search space. We introduce a new alternative method and compare it to a volume doubling baseline on two common synthetic benchmarking test functions. Finally, we apply our proposed methods on the task of tuning the stochastic gradient descent optimizer for both a multi-layered perceptron and a convolutional neural network on the MNIST dataset. Bobak Shahriari, Alexandre Bouchard-Côté, Nando de Freitas |
AISTATS | 3 |
| 2016 | Dueling Network Architectures for Deep Reinforcement LearningabstractIn recent years there have been many successes of using deep representations in reinforcement learning. Still, many of these applications use conventional architectures, such as convolutional networks, LSTMs, or auto-encoders. In this paper, we present a new neural network architecture for model-free reinforcement learning. Our dueling network represents two separate estimators: one for the state value function and one for the state-dependent action advantage function. The main benefit of this factoring is to generalize learning across actions without imposing any change to the underlying reinforcement learning algorithm. Our results show that this architecture leads to better policy evaluation in the presence of many similar-valued actions. Moreover, the dueling architecture enables our RL agent to outperform the state-of-the-art on the Atari 2600 domain. Ziyu Wang 0001, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, Nando de Freitas |
ICML | 6 |
| 2016 | Learning to Learn and Compositionality with Deep Recurrent Neural Networks: Learning to Learn and CompositionalityabstractDeep neural network representations play an important role in computer vision, speech, computational linguistics, robotics, reinforcement learning and many other data-rich domains. In this talk I will show that learning-to-learn and compositionality are key ingredients for dealing with knowledge transfer so as to solve a wide range of tasks, for dealing with small-data regimes, and for continual learning. I will demonstrate this with several examples from my research team: learning to learn by gradient descent by gradient descent, neural programmers and interpreters, and learning communication. Nando de Freitas |
KDD | 1 |
| 2016 | Learning to learn by gradient descent by gradient descentabstractThe move from hand-designed features to learned features in machine learning has been wildly successful. In spite of this, optimization algorithms are still designed by hand. In this paper we show how the design of an optimization algorithm can be cast as a learning problem, allowing the algorithm to learn to exploit structure in the problems of interest in an automatic way. Our learned algorithms, implemented by LSTMs, outperform generic, hand-designed competitors on the tasks for which they are trained, and also generalize well to new tasks with similar structure. We demonstrate this on a number of tasks, including simple convex problems, training neural networks, and styling images with neural art. Marcin Andrychowicz, Misha Denil, Sergio Gomez Colmenarejo, Matt Hoffman 0001, David Pfau, Tom Schaul, Nando de Freitas |
NIPS | 7 |
| 2016 | Learning to Communicate with Deep Multi-Agent Reinforcement LearningabstractWe consider the problem of multiple agents sensing and acting in environments with the goal of maximising their shared utility. In these environments, agents must learn communication protocols in order to share information that is needed to solve the tasks. By embracing deep neural networks, we are able to demonstrate end-to-end learning of protocols in complex environments inspired by communication riddles and multi-agent computer vision problems with partial observability. We propose two approaches for learning in these domains: Reinforced Inter-Agent Learning (RIAL) and Differentiable Inter-Agent Learning (DIAL). The former uses deep Q-learning, while the latter exploits the fact that, during learning, agents can backpropagate error derivatives through (noisy) communication channels. Hence, this approach uses centralised learning but decentralised execution. Our experiments introduce new environments for studying the learning of communication protocols and present a set of engineering innovations that are essential for success in these domains. Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, Shimon Whiteson |
NIPS | 3 |
| 2016 | Bayesian Optimization in a Billion Dimensions via Random EmbeddingsabstractBayesian optimization techniques have been successfully applied to robotics, planning, sensor placement, recommendation, advertising, intelligent user interfaces and automatic algorithm configuration. Despite these successes, the approach is restricted to problems of moderate dimension, and several workshops on Bayesian optimization have identified its scaling to high-dimensions as one of the holy grails of the field. In this paper, we introduce a novel random embedding idea to attack this problem. The resulting Random EMbedding Bayesian Optimization (REMBO) algorithm is very simple, has important invariance properties, and applies to domains with both categorical and continuous variables. We present a thorough theoretical analysis of REMBO. Empirical results confirm that REMBO can effectively solve problems with billions of dimensions, provided the intrinsic dimensionality is low. They also show that REMBO achieves state-of-the-art performance in optimizing the 47 discrete parameters of a popular mixed integer linear programming solver. Ziyu Wang 0001, Frank Hutter, Masrour Zoghi, David Matheson, Nando de Freitas |
J. Artif. Intell. Res. | 5 |
| 2016 | Herded Gibbs SamplingabstractThe Gibbs sampler is one of the most popular algorithms for inference in statistical models. In this paper, we introduce a herding variant of this algorithm, called herded Gibbs, that is entirely deterministic. We prove that herded Gibbs has an $O(1/T)$ convergence rate for models with independent variables and for fully connected probabilistic graphical models. Herded Gibbs is shown to outperform Gibbs in the tasks of image denoising with MRFs and named entity recognition with CRFs. However, the convergence for herded Gibbs for sparsely connected probabilistic graphical models is still an open problem. Yutian Chen 0001, Luke Bornn, Nando de Freitas, Mareija Eskelin, Max Welling |
J. Mach. Learn. Res. | 3 |
| 2016 | Taking the Human Out of the Loop: A Review of Bayesian OptimizationabstractBig Data applications are typically associated with systems involving large numbers of users, massive complex software systems, and large-scale heterogeneous computing and storage architectures. The construction of such systems involves many distributed design choices. The end products (e.g., recommendation systems, medical analysis tools, real-time game engines, speech recognizers) thus involve many tunable configuration parameters. These parameters are often specified and hard-coded into the software by various developers or teams. If optimized jointly, these parameters can result in significant improvements. Bayesian optimization is a powerful tool for the joint optimization of design choices that is gaining great popularity in recent years. It promises greater automation so as to increase both product quality and human productivity. This review paper introduces Bayesian optimization, highlights some of its methodological aspects, and showcases a wide range of applications. Bobak Shahriari, Kevin Swersky, Ziyu Wang 0001, Ryan P. Adams, Nando de Freitas |
Proc. IEEE | 5 |
| 2015 | Deep Fried ConvnetsabstractThe fully-connected layers of deep convolutional neural networks typically contain over 90% of the network parameters. Reducing the number of parameters while preserving predictive performance is critically important for training big models in distributed systems and for deployment in embedded devices. In this paper, we introduce a novel Adaptive Fastfood transform to reparameterize the matrix-vector multiplication of fully connected layers. Reparameterizing a fully connected layer with d inputs and n outputs with the Adaptive Fastfood transform reduces the storage and computational costs costs from O(nd) to O(n) and O(n log d) respectively. Using the Adaptive Fastfood transform in convolutional networks results in what we call a deep fried convnet. These convnets are end-to-end trainable, and enable us to attain substantial reductions in the number of parameters without affecting prediction accuracy on the MNIST and ImageNet datasets. Marcin Moczulski, Misha Denil, Nando de Freitas, Alexander J. Smola, Ziyu Wang 0001 |
ICCV | 4 |
| 2015 | From Group to Individual Labels Using Deep FeaturesabstractIn many classification problems labels are relatively scarce. One context in which this occurs is where we have labels for groups of instances but not for the instances themselves, as in multi-instance learning. Past work on this problem has typically focused on learning classifiers to make predictions at the group level. In this paper we focus on the problem of learning classifiers to make predictions at the instance level. To achieve this we propose a new objective function that encourages smoothness of inferred instance-level labels based on instance-level similarity, while at the same time respecting group-level label constraints. We apply this approach to the problem of predicting labels for sentences given labels for reviews, using a convolutional neural network to infer sentence similarity. The approach is evaluated using three large review data sets from IMDB, Yelp, and Amazon, and we demonstrate the proposed approach is both accurate and scalable compared to various alternatives. Dimitrios Kotzias, Misha Denil, Nando de Freitas, Padhraic Smyth |
KDD | 3 |
| 2014 | On correlation and budget constraints in model-based bandit optimization with application to automatic machine learningabstractWe address the problem of finding the maximizer of a nonlinear function that can only be evaluated, subject to noise, at a finite number of query locations. Further, we will assume that there is a constraint on the total number of permitted function evaluations. We introduce a Bayesian approach for this problem and show that it empirically outperforms both the existing frequentist counterpart and other Bayesian optimization methods. The Bayesian approach places emphasis on detailed modelling, including the modelling of correlations among the arms. As a result, it can perform well in situations where the number of arms is much larger than the number of allowed function evaluation, whereas the frequentist counterpart is inapplicable. This feature enables us to develop and deploy practical applications, such as automatic machine learning toolboxes. The paper presents comprehensive comparisons of the proposed approach with many Bayesian and bandit optimization techniques, the first comparison of many of these methods in the literature. Matt Hoffman 0001, Bobak Shahriari, Nando de Freitas |
AISTATS | 3 |
| 2014 | Bayesian Multi-Scale Optimistic OptimizationabstractBayesian optimization is a powerful global optimization technique for expensive black-box functions. One of its shortcomings is that it requires auxiliary optimization of an acquisition function at each iteration. This auxiliary optimization can be costly and very hard to carry out in practice. Moreover, it creates serious theoretical concerns, as most of the convergence results assume that the exact optimum of the acquisition function can be found. In this paper, we introduce a new technique for efficient global optimization that combines Gaussian process confidence bounds and treed simultaneous optimistic optimization to eliminate the need for auxiliary optimization of acquisition functions. The experiments with global optimization benchmarks, as well as a novel application to automate information extraction, demonstrate that the resulting technique is more efficient than the two approaches from which it draws inspiration. Unlike most theoretical analyses of Bayesian optimization with Gaussian processes, our convergence rate proofs do not require exact optimization of an acquisition function. That is, our approach eliminates the unsatisfactory assumption that a difficult, potentially NP-hard, problem has to be solved in order to obtain vanishing regret rates. Ziyu Wang 0001, Babak Shakibi, Lin Jin, Nando de Freitas |
AISTATS | 4 |
| 2014 | Narrowing the Gap: Random Forests In Theory and In PracticeabstractDespite widespread interest and practical use, the theoretical properties of random forests are still not well understood. In this paper we contribute to this understanding in two ways. We present a new theoreti- cally tractable variant of random regression forests and prove that our algorithm is con- sistent. We also provide an empirical eval- uation, comparing our algorithm and other theoretically tractable random forest models to the random forest algorithm used in prac- tice. Our experiments provide insight into the relative importance of different simplifi- cations that theoreticians have made to ob- tain tractable models for analysis. Misha Denil, David Matheson, Nando de Freitas |
ICML | 3 |
| 2014 | Linear and Parallel Learning of Markov Random FieldsabstractWe introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike its competitors, our algorithm is fully parallel and for log-linear models it is also data efficient, requiring only the local sufficient statistics of the data to estimate parameters. Yariv Dror Mizrahi, Misha Denil, Nando de Freitas |
ICML | 3 |
| 2014 | Distributed Parameter Estimation in Probabilistic Graphical Models
Yariv Dror Mizrahi, Misha Denil, Nando de Freitas |
NIPS | 3 |
| 2014 | Bayesian Optimization with an Empirical Hardness Model for approximate Nearest Neighbour SearchabstractNearest Neighbour Search in high-dimensional spaces is a common problem in Computer Vision. Although no algorithm better than linear search is known, approximate algorithms are commonly used to tackle this problem. The drawback of using such algorithms is that their performance depends highly on parameter tuning. While this process can be automated using standard empirical optimization techniques, tuning is still time-consuming. In this paper, we propose to use Empirical Hardness Models to reduce the number of parameter configurations that Bayesian Optimization has to try, speeding up the optimization process. Evaluation on standard benchmarks of SIFT and GIST descriptors shows the viability of our approach. Julieta Martinez 0001, James J. Little, Nando de Freitas |
WACV | 3 |
| 2013 | Consistency of Online Random ForestsabstractAs a testament to their success, the theory of random forests has long been outpaced by their application in practice. In this paper, we take a step towards narrowing this gap by providing a consistency result for online random forests. Misha Denil, David Matheson, Nando de Freitas |
ICML (3) | 3 |
| 2013 | Adaptive Hamiltonian and Riemann Manifold Monte CarloabstractIn this paper we address the widely-experienced difficulty in tuning Hamiltonian-based Monte Carlo samplers. We develop an algorithm that allows for the adaptation of Hamiltonian and Riemann manifold Hamiltonian Monte Carlo samplers using Bayesian optimization that allows for infinite adaptation of the parameters of these samplers. We show that the resulting sampling algorithms are ergodic, and demonstrate on several models and data sets that the use of our adaptive algorithms makes it is easy to obtain more efficient samplers, in some precluding the need for more complex models. Hamiltonian-based Monte Carlo samplers are widely known to be an excellent choice of MCMC method, and we aim with this paper to remove a key obstacle towards the more widespread use of these samplers in practice. Ziyu Wang 0001, Shakir Mohamed, Nando de Freitas |
ICML (3) | 3 |
| 2013 | Bayesian Optimization in High Dimensions via Random Embeddings
Ziyu Wang 0001, Masrour Zoghi, Frank Hutter, David Matheson, Nando de Freitas |
IJCAI | 5 |
| 2013 | Predicting Parameters in Deep LearningabstractWe demonstrate that there is significant redundancy in the parameterization of several deep learning models. Given only a few weight values for each feature it is possible to accurately predict the remaining values. Moreover, we show that not only can the parameter values be predicted, but many of them need not be learned at all. We train several different architectures by learning only a small number of weights and predicting the rest. In the best case we are able to predict more than 95% of the weights of a network without any drop in accuracy. Misha Denil, Babak Shakibi, Laurent Dinh, Marc'Aurelio Ranzato, Nando de Freitas |
NIPS | 5 |
| 2012 | Prediction and Fault Detection of Environmental Signals with Uncharacterised FaultsabstractMany signals of interest are corrupted by faults of anunknown type. We propose an approach that uses Gaus-sian processes and a general “fault bucket” to capturea priori uncharacterised faults, along with an approxi-mate method for marginalising the potential faultinessof all observations. This gives rise to an efficient, flexible algorithm for the detection and automatic correction of faults. Our method is deployed in the domain of water monitoring and management, where it is able to solve several fault detection, correction, and prediction problems. The method works well despite the fact that the data is plagued with numerous difficulties, including missing observations, multiple discontinuities, nonlinearity and many unanticipated types of fault. Michael A. Osborne, Roman Garnett, Kevin Swersky, Nando de Freitas |
AAAI | 4 |
| 2012 | A Machine Learning Perspective on Predictive Coding with PAQ8abstractPAQ8 makes use of several simple machine learning models and algorithms. We show how understanding PAQ8 enables us to improve the algorithms. We also present a broad range of new applications of PAQ8 to machine learning tasks including language modeling and adaptive text prediction, adaptive game playing, classification, and lossy compression using features acquired via unsupervised learning. Byron Knoll, Nando de Freitas |
DCC | 2 |
| 2012 | Exponential Regret Bounds for Gaussian Process Bandits with Deterministic Observations
Nando de Freitas, Alexander J. Smola, Masrour Zoghi |
ICML | 1 |
| 2012 | Learning Where to Attend with Deep Architectures for Image TrackingabstractWe discuss an attentional model for simultaneous object tracking and recognition that is driven by gaze data. Motivated by theories of perception, the model consists of two interacting pathways, identity and control, intended to mirror the what and where pathways in neuroscience models. The identity pathway models object appearance and performs classification using deep (factored)-restricted Boltzmann machines. At each point in time, the observations consist of foveated images, with decaying resolution toward the periphery of the gaze. The control pathway models the location, orientation, scale, and speed of the attended object. The posterior distribution of these states is estimated with particle filtering. Deeper in the control pathway, we encounter an attentional mechanism that learns to select gazes so as to minimize tracking uncertainty. Unlike in our previous work, we introduce gaze selection strategies that operate in the presence of partial information and on a continuous action space. We show that a straightforward extension of the existing approach to the partial information setting results in poor performance, and we propose an alternative method based on modeling the reward surface as a gaussian process. This approach gives good performance in the presence of partial information and allows us to expand the action space from a small, discrete set of fixation points to a continuous domain. Misha Denil, Loris Bazzani, Hugo Larochelle, Nando de Freitas |
Neural Comput. | 4 |
| 2011 | Learning attentional policies for tracking and recognition in video with deep networks
Loris Bazzani, Nando de Freitas, Hugo Larochelle, Vittorio Murino, Jo-Anne Ting |
ICML | 2 |
| 2011 | On Autoencoders and Score Matching for Energy Based Models
Kevin Swersky, Marc'Aurelio Ranzato, David Buchman, Benjamin M. Marlin, Nando de Freitas |
ICML | 5 |
| 2011 | Portfolio Allocation for Bayesian Optimization
Matt Hoffman 0001, Eric Brochu, Nando de Freitas |
UAI | 3 |
| 2011 | Asymptotic Efficiency of Deterministic Estimators for Discrete Energy-Based Models: Ratio Matching and Pseudolikelihood
Benjamin M. Marlin, Nando de Freitas |
UAI | 2 |
| 2010 | Intracluster Moves for Constrained Discrete-Space MCMC
Firas Hamze, Nando de Freitas |
UAI | 2 |
| 2009 | New inference strategies for solving Markov Decision Processes using reversible jump MCMC
Matthias Hoffman, Hendrik Kück, Nando de Freitas, Arnaud Doucet |
UAI | 3 |
| 2008 | Target-directed attention: Sequential decision-making for gaze planningabstractIt is widely agreed that efficient visual search requires the integration of target-driven top-down information and image-driven bottom-up information. Yet the problem of gaze planning - that is, selecting the next best gaze location given the current observations - remains largely unsolved. We propose a probabilistic system that models the gaze sequence as a finite-horizon Bayesian sequential decision process. Direct policy search is used to reason about the next best gaze locations. The system integrates bottom-up saliency information, top-down target knowledge and additional context information through principled Bayesian priors. This results in proposal gaze locations that depend not only the featural visual saliency, but also on prior knowledge and the spatial likelihood of locating the target. The system has been implemented using state-of- the-art object detectors and evaluated on a real-world dataset by comparing it to gaze sequences proposed by a pure bottom-up saliency-based process and to an object detection approach that analyzes the full image. The target-directed attention system is shown to result in higher object detection precision than both competitors, to attend to more relevant targets than the bottom-up attention system, and to require significantly less computation time than the exhaustive approach. Julia Vogel, Nando de Freitas |
ICRA | 2 |
| 2008 | An interior-point stochastic approximation method and an L1-regularized delta ruleabstractThe stochastic approximation method is behind the solution to many important, actively-studied problems in machine learning. Despite its far-reaching application, there is almost no work on applying stochastic approximation to learning problems with constraints. The reason for this, we hypothesize, is that no robust, widely-applicable stochastic approximation method exists for handling such problems. We propose that interior-point methods are a natural solution. We establish the stability of a stochastic interior-point approximation method both analytically and empirically, and demonstrate its utility by deriving an on-line learning algorithm that also performs feature selection via L1 regularization. Peter Carbonetto, Mark Schmidt 0001, Nando de Freitas |
NIPS | 3 |
| 2008 | Learning to Recognize Objects with Little Supervision
Peter Carbonetto, Gyuri Dorkó, Cordelia Schmid, Hendrik Kück, Nando de Freitas |
Int. J. Comput. Vis. | 5 |
| 2007 | Analysis of Particle Methods for Simultaneous Robot Localization and Mapping and a New Algorithm: Marginal-SLAMabstractThis paper presents a new particle method, with stochastic parameter estimation, to solve the SLAM problem. The underlying algorithm is rooted on a solid probabilistic foundation and is guaranteed to converge asymptotically, unlike many existing popular approaches. Moreover, it is efficient in storage and computation. The new algorithm carries out filtering only in the marginal filtering space, thereby allowing for the recursive computation of low variance estimates of the map. The paper provides mathematical arguments and empirical evidence to substantiate the fact that the new method represents an improvement over the existing particle filtering approaches for SLAM, which work on the joint path state space. Ruben Martinez-Cantin, Nando de Freitas, José A. Castellanos 0001 |
ICRA | 2 |
| 2007 | Active Preference Learning with Discrete Choice DataabstractWe propose an active learning algorithm that learns a continuous valuation model from discrete preferences. The algorithm automatically decides what items are best presented to an individual in order to find the item that they value highly in as few trials as possible, and exploits quirks of human psychology to minimize time and cognitive burden. To do this, our algorithm maximizes the expected improvement at each query without accurately modelling the entire valuation surface, which would be needlessly expensive. The problem is particularly difficult because the space of choices is infinite. We demonstrate the effectiveness of the new algorithm compared to related active learning methods. We also embed the algorithm within a decision making tool for assisting digital artists in rendering materials. The tool finds the best parameters while minimizing the number of queries. Eric Brochu, Nando de Freitas, Abhijeet Ghosh |
NIPS | 2 |
| 2007 | Bayesian Policy Learning with Trans-Dimensional MCMCabstractA recently proposed formulation of the stochastic planning and control problem as one of parameter estimation for suitable artificial statistical models has led to the adoption of inference algorithms for this notoriously hard problem. At the algorithmic level, the focus has been on developing Expectation-Maximization (EM) algorithms. In this paper, we begin by making the crucial observation that the stochastic control problem can be reinterpreted as one of trans-dimensional inference. With this new interpretation, we are able to propose a novel reversible jump Markov chain Monte Carlo (MCMC) algorithm that is more efficient than its EM counterparts. Moreover, it enables us to implement full Bayesian policy search, without the need for gradients and with one single Markov chain. The new approach involves sampling directly from a distribution that is proportional to the reward and, consequently, performs better than classic simulations methods in situations where the reward is a rare event. Matt Hoffman 0001, Arnaud Doucet, Nando de Freitas, Ajay Jasra |
NIPS | 3 |
| 2007 | Large-Flip Importance Sampling
Firas Hamze, Nando de Freitas |
UAI | 2 |
| 2006 | Robust Visual Tracking for Multiple Targets
Yizheng Cai, Nando de Freitas, James J. Little |
ECCV (4) | 2 |
| 2006 | Fast particle smoothing: if I had a million particlesabstractWe propose efficient particle smoothing methods for generalized state-spaces models. Particle smoothing is an expensive O(N2) algorithm, where N is the number of particles. We overcome this problem by integrating dual tree recursions and fast multipole techniques with forward-backward smoothers, a new generalized two-filter smoother and a maximum a posteriori (MAP) smoother. Our experiments show that these improvements can substantially increase the practicality of particle smoothing. Mike Klaas, Mark Briers, Nando de Freitas, Arnaud Doucet, Simon Maskell, Dustin Lang |
ICML | 3 |
| 2006 | Conditional mean fieldabstractDespite all the attention paid to variational methods based on sum-product message passing (loopy belief propagation, tree-reweighted sum-product), these methods are still bound to inference on a small set of probabilistic models. Mean field approximations have been applied to a broader set of problems, but the solutions are often poor. We propose a new class of conditionally-specified variational approximations based on mean field theory. While not usable on their own, combined with sequential Monte Carlo they produce guaranteed improvements over conventional mean field. Moreover, experiments on a well-studied problem-- inferring the stable configurations of the Ising spin glass--show that the solutions can be significantly better than those obtained using sum-product-based methods. Peter Carbonetto, Nando de Freitas |
NIPS | 2 |
| 2005 | Fast Computational Methods for Visually Guided RobotsabstractThis paper proposes numerical algorithms for reducing the computational cost of semi-supervised and active learning procedures for visually guided mobile robots from O(M3to O(M), while reducing the storage requirements from M2to M . This reduction in cost is essential for real-time interaction with mobile robots. The considerable speed ups are achieved using Krylov subspace methods and the fast Gauss transform. Although these state-of-the-art numerical algorithms are known, their application to semi-supervised learning, active learning and mobile robotics is new and should be of interest and great value to the robotics community. We apply our fast algorithms to interactive object recognition on Sony’s ERS-7 Aibo. We provide comparisons that clearly demonstrate remarkable improvements in computational speed. Maryam Mahdaviani, Nando de Freitas, Robert Fraser, Firas Hamze |
ICRA | 2 |
| 2005 | Fast Krylov Methods for N-Body LearningabstractThis paper addresses the issue of numerical computation in machine learning domains based on similarity metrics, such as kernel methods, spectral techniques and Gaussian processes. It presents a general solution strategy based on Krylov subspace iteration and fast N-body learning methods. The experiments show significant gains in computation and storage on datasets arising in image segmentation, object detection and dimensionality reduction. The paper also presents theoretical bounds on the stability of these methods. Nando de Freitas, Yang Wang 0003, Maryam Mahdaviani, Dustin Lang |
NIPS | 1 |
| 2005 | Hot Coupling: A Particle Approach to Inference and Normalization on Pairwise Undirected GraphsabstractThis paper presents a new sampling algorithm for approximating func- tions of variables representable as undirected graphical models of arbi- trary connectivity with pairwise potentials, as well as for estimating the notoriously dif(cid:2)cult partition function of the graph. The algorithm (cid:2)ts into the framework of sequential Monte Carlo methods rather than the more widely used MCMC, and relies on constructing a sequence of in- termediate distributions which get closer to the desired one. While the idea of using (cid:147)tempered(cid:148) proposals is known, we construct a novel se- quence of target distributions where, rather than dropping a global tem- perature parameter, we sequentially couple individual pairs of variables that are, initially, sampled exactly from a spanning tree of the variables. We present experimental results on inference and estimation of the parti- tion function for sparse and densely-connected graphs. Firas Hamze, Nando de Freitas |
NIPS | 2 |
| 2005 | Nonparametric Bayesian Logic
Peter Carbonetto, Jacek Kisynski, Nando de Freitas, David Poole 0001 |
UAI | 3 |
| 2005 | Learning about Individuals from Group Statistics
Nando de Freitas, Hendrik Kück |
UAI | 1 |
| 2005 | Toward Practical N2 Monte Carlo: the Marginal Particle Filter
Mike Klaas, Nando de Freitas, Arnaud Doucet |
UAI | 2 |
| 2004 | A Statistical Model for General Contextual Object Recognition
Peter Carbonetto, Nando de Freitas, Kobus Barnard |
ECCV (1) | 2 |
| 2004 | A Constrained Semi-supervised Learning Approach to Data Association
Hendrik Kück, Peter Carbonetto, Nando de Freitas |
ECCV (3) | 3 |
| 2004 | A Boosted Particle Filter: Multitarget Detection and Tracking
Kenji Okuma, Ali Taleghani, Nando de Freitas, James J. Little, David G. Lowe |
ECCV (1) | 3 |
| 2004 | Beat Tracking the Graphical Model WayabstractWe present a graphical model for beat tracking in recorded music. Using a probabilistic graphical model allows us to incorporate local information and global smoothness constraints in a principled manner. We evaluate our model on a set of varied and difficult examples, and achieve impres- sive results. By using a fast dual-tree algorithm for graphical model in- ference, our system runs in less time than the duration of the music being processed. Dustin Lang, Nando de Freitas |
NIPS | 2 |
| 2004 | From Fields to Trees
Firas Hamze, Nando de Freitas |
UAI | 2 |
| 2004 | Diagnosis by a waiter and a Mars explorerabstractThis paper shows how state-of-the-art state estimation techniques can be used to provide efficient solutions to the difficult problem of real-time diagnosis in mobile robots. The power of the adopted estimation techniques resides in our ability to combine particle filters with classical algorithms, such as Kalman filters. We demonstrate these techniques in two scenarios: a mobile waiter robot and planetary rovers designed by NASA for Mars exploration. Nando de Freitas, Richard Dearden, Frank Hutter, Rubén Morales-Menéndez, Jim Mutch, David Poole 0001 |
Proc. IEEE | 1 |
| 2003 | Matching Words and Pictures
Kobus Barnard, Pinar Duygulu, David A. Forsyth, Nando de Freitas, David M. Blei, Michael I. Jordan |
J. Mach. Learn. Res. | 4 |
| 2003 | An Introduction to MCMC for Machine Learning
Christophe Andrieu, Nando de Freitas, Arnaud Doucet, Michael I. Jordan |
Mach. Learn. | 2 |
| 2002 | "Name That Song!" A Probabilistic Approach to Querying on Music and TextabstractWe present a novel, flexible statistical approach for modelling music and text jointly. The approach is based on multi-modal mixture models and maximum a posteriori estimation using EM. The learned models can be used to browse databases with documents containing music and text, to search for music using queries consisting of music and text (lyrics and other contextual information), to annotate text documents with music, and to automatically recommend or identify similar songs. Eric Brochu, Nando de Freitas |
NIPS | 2 |
| 2002 | Real-Time Monitoring of Complex Industrial Processes with Particle FiltersabstractThis paper discusses the application of particle filtering algorithms to fault diagnosis in complex industrial processes. We consider two ubiq- uitous processes: an industrial dryer and a level tank. For these appli- cations, we compared three particle filtering variants: standard parti- cle filtering, Rao-Blackwellised particle filtering and a version of Rao- Blackwellised particle filtering that does one-step look-ahead to select good sampling regions. We show that the overhead of the extra process- ing per particle of the more sophisticated methods is more than compen- sated by the decrease in error and variance. Rubén Morales-Menéndez, Nando de Freitas, David Poole 0001 |
NIPS | 2 |
| 2001 | Rao-Blackwellised Particle Filtering via Data AugmentationabstractEE Engineering University of Melbourne Parkville, Victoria 3052 Christophe Andrieu, Nando de Freitas, Arnaud Doucet |
NIPS | 2 |
| 2001 | Variational MCMC
Nando de Freitas, Pedro A. d. F. R. Højen-Sørensen, Stuart Russell 0001 |
UAI | 1 |
| 2001 | Robust Full Bayesian Learning for Radial Basis NetworksabstractWe propose a hierarchical full Bayesian model for radial basis networks. This model treats the model dimension (number of neurons), model parameters, regularization parameters, and noise parameters as unknown random variables. We develop a reversible-jump Markov chain Monte Carlo (MCMC) method to perform the Bayesian computation. We find that the results obtained using this method are not only better than the ones reported previously, but also appear to be robust with respect to the prior specification. In addition, we propose a novel and computationally efficient reversible-jump MCMC simulated annealing algorithm to optimize neural networks. This algorithm enables us to maximize the joint posterior distribution of the network parameters and the number of basis function. It performs a global search in the joint space of the parameters and number of parameters, thereby surmounting the problem of local minima to a large extent. We show that by calibrating the full hierarchical Bayesian prior, we can obtain the classical Akaike information criterion, Bayesian information criterion, and minimum description length model selection criteria within a penalized likelihood framework. Finally, we present a geometric convergence theorem for the algorithm with homogeneous transition kernel and a convergence theorem for the reversible-jump MCMC simulated annealing method. Christophe Andrieu, Nando de Freitas, Arnaud Doucet |
Neural Comput. | 2 |
| 2000 | Sequential Monte Carlo for model selection and estimation of neural networksabstractWe address the complex problem of sequential Bayesian learning and model selection for neural networks. This problem does not usually admit any type of closed-form analytical solution and, as a result, one has to resort to numerical methods. We propose here an original sequential simulation-based strategy to perform the necessary computations. It combines sequential importance sampling, a selection procedure, variance reduction techniques and reversible jump Markov chain Monte Carlo (MCMC) moves. We demonstrate the effectiveness of the method by applying it to radial basis function networks. The approach can be easily extended to other interesting on-line model selection problems. Christophe Andrieu, Nando de Freitas |
ICASSP | 2 |
| 2000 | The Unscented Particle FilterabstractIn this paper, we propose a new particle filter based on sequential importance sampling. The algorithm uses a bank of unscented fil(cid:173) ters to obtain the importance proposal distribution. This proposal has two very "nice" properties. Firstly, it makes efficient use of the latest available information and, secondly, it can have heavy tails. As a result, we find that the algorithm outperforms stan(cid:173) dard particle filtering and other nonlinear filtering methods very substantially. This experimental finding is in agreement with the theoretical convergence proof for the algorithm. The algorithm also includes resampling and (possibly) Markov chain Monte Carlo (MCMC) steps. Rudolph van der Merwe, Arnaud Doucet, Nando de Freitas, Eric A. Wan |
NIPS | 3 |
| 2000 | Reversible Jump MCMC Simulated Annealing for Neural Networks
Christophe Andrieu, Nando de Freitas, Arnaud Doucet |
UAI | 2 |
| 2000 | Rao-Blackwellised Particle Filtering for Dynamic Bayesian Networks
Arnaud Doucet, Nando de Freitas, Kevin Murphy 0002, Stuart Russell 0001 |
UAI | 2 |