Miguel Lázaro-Gredilla

dblp:77/4660 · DBLP profile ↗
← Back
36ranked-venue papers
13as first author
12since 2021 · last 2025
0000-0002-4528-5084ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 11 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Improving Transformer World Models for Data-Efficient RL
abstract
We present an approach to model-based RL that achieves a new state of the art performance on the challenging Craftax-classic benchmark, an open-world 2D survival game that requires agents to exhibit a wide range of general abilities---such as strong generalization, deep exploration, and long-term reasoning. With a series of careful design choices aimed at improving sample efficiency, our MBRL algorithm achieves a reward of 69.66% after only 1M environment steps, significantly outperforming DreamerV3, which achieves $53.2\%$, and, for the first time, exceeds human performance of 65.0%. Our method starts by constructing a SOTA model-free baseline, using a novel policy architecture that combines CNNs and RNNs. We then add three improvements to the standard MBRL setup: (a) "Dyna with warmup", which trains the policy on real and imaginary data, (b) "nearest neighbor tokenizer" on image patches, which improves the scheme to create the transformer world model (TWM) inputs, and (c) "block teacher forcing", which allows the TWM to reason jointly about the future tokens of the next timestep.
Antoine Dedieu, Joseph Ortiz, Xinghua Lou, Carter Wendelken, J. Swaroop Guntupalli, Wolfgang Lehrach, Miguel Lázaro-Gredilla, Kevin Murphy 0002
ICML7
2024 Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments
abstract
Despite their stellar performance on a wide range of tasks, including in-context tasks only revealed during inference, vanilla transformers and variants trained for next-token predictions (a) do not learn an explicit world model of their environment which can be flexibly queried and (b) cannot be used for planning or navigation. In this paper, we consider partially observed environments (POEs), where an agent receives perceptually aliased observations as it navigates, which makes path planning hard. We introduce a transformer with (multiple) discrete bottleneck(s), TDB, whose latent codes learn a compressed representation of the history of observations and actions. After training a TDB to predict the future observation(s) given the history, we extract interpretable cognitive maps of the environment from its active bottleneck(s) indices. These maps are then paired with an external solver to solve (constrained) path planning problems. First, we show that a TDB trained on POEs (a) retains the near-perfect predictive performance of a vanilla transformer or an LSTM while (b) solving shortest path problems exponentially faster. Second, a TDB extracts interpretable representations from text datasets, while reaching higher in-context accuracy than vanilla sequence models. Finally, in new POEs, a TDB (a) reaches near-perfect in-context accuracy, (b) learns accurate in-context cognitive maps (c) solves in-context path planning problems.
Antoine Dedieu, Wolfgang Lehrach, Dileep George, Miguel Lázaro-Gredilla
ICML5
2024 What type of inference is planning?
abstract
Multiple types of inference are available for probabilistic graphical models, e.g., marginal, maximum-a-posteriori, and even marginal maximum-a-posteriori. Which one do researchers mean when they talk about ``planning as inference''? There is no consistency in the literature, different types are used, and their ability to do planning is further entangled with specific approximations or additional constraints. In this work we use the variational framework to show that, just like all commonly used types of inference correspond to different weightings of the entropy terms in the variational problem, planning corresponds _exactly_ to a _different_ set of weights. This means that all the tricks of variational inference are readily applicable to planning. We develop an analogue of loopy belief propagation that allows us to perform approximate planning in factored-state Markov decisions processes without incurring intractability due to the exponentially large state space. The variational perspective shows that the previous types of inference for planning are only adequate in environments with low stochasticity, and allows us to characterize each type by its own merits, disentangling the type of inference from the additional approximations that its practical use requires. We validate these results empirically on synthetic MDPs and tasks posed in the International Planning Competition.
Miguel Lázaro-Gredilla, Li Yang Ku, Kevin Murphy 0002, Dileep George
NeurIPS1
2024 DMC-VB: A Benchmark for Representation Learning for Control with Visual Distractors
abstract
Learning from previously collected data via behavioral cloning or offline reinforcement learning (RL) is a powerful recipe for scaling generalist agents by avoiding the need for expensive online learning. Despite strong generalization in some respects, agents are often remarkably brittle to minor visual variations in control-irrelevant factors such as the background or camera viewpoint. In this paper, we present theDeepMind Control Visual Benchmark (DMC-VB), a dataset collected in the DeepMind Control Suite to evaluate the robustness of offline RL agents for solving continuous control tasks from visual input in the presence of visual distractors. In contrast to prior works, our dataset (a) combines locomotion and navigation tasks of varying difficulties, (b) includes static and dynamic visual variations, (c) considers data generated by policies with different skill levels, (d) systematically returns pairs of state and pixel observation, (e) is an order of magnitude larger, and (f) includes tasks with hidden goals. Accompanying our dataset, we propose three benchmarks to evaluate representation learning methods for pretraining, and carry out experiments on several recently proposed methods. First, we find that pretrained representations do not help policy learning on DMC-VB, and we highlight a large representation gap between policies learned on pixel observations and on states. Second, we demonstrate when expert data is limited, policy learning can benefit from representations pretrained on (a) suboptimal data, and (b) tasks with stochastic hidden goals. Our dataset and benchmark code to train and evaluate agents are available at https://github.com/google-deepmind/dmcvisionbenchmark.
Joseph Ortiz, Antoine Dedieu, Wolfgang Lehrach, J. Swaroop Guntupalli, Carter Wendelken, Ahmad Humayun, Sivaramakrishnan Swaminathan, Miguel Lázaro-Gredilla, Kevin Murphy 0002
NeurIPS9
2024 PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX
abstract
PGMax is an open-source Python/ JAX package for (a) easily specifying discrete Probabilistic Graphical Models (PGMs) as factor graphs; and (b) automatically running efficient and scalable differentiable Loopy Belief Propagation (LBP). PGMax supports general factor graphs with tractable factors, and leverages modern accelerators like GPUs for inference. Compared with alternative libraries, PGMax obtains higher-quality inference results with up to three orders-of-magnitude inference time speedups. PGMax interacts seamlessly with the growing JAX ecosystem, opening up new research possibilities. Our source code, examples and documentation are available at https://github.com/google-deepmind/PGMax
Antoine Dedieu, Nishanth Kumar, Wolfgang Lehrach, Shrinu Kushagra, Dileep George, Miguel Lázaro-Gredilla
J. Mach. Learn. Res.7
2023 3D Neural Embedding Likelihood: Probabilistic Inverse Graphics for Robust 6D Pose Estimation
abstract
The ability to perceive and understand 3D scenes is crucial for many applications in computer vision and robotics. Inverse graphics is an appealing approach to 3D scene understanding that aims to infer the 3D scene structure from 2D images. In this paper, we introduce probabilistic modeling to the inverse graphics framework to quantify uncertainty and achieve robustness in 6D pose estimation tasks. Specifically, we propose 3D Neural Embedding Likelihood (3DNEL) as a unified probabilistic model over RGBD images, and develop efficient inference procedures on 3D scene descriptions. 3DNEL effectively combines learned neural embeddings from RGB with depth information to improve robustness in sim-to-real 6D object pose estimation from RGB-D images. Performance on the YCB-Video dataset is on par with state-of-the-art yet is much more robust in challenging regimes. In contrast to discriminative approaches, 3DNEL’s probabilistic generative formulation jointly models multiple objects in a scene, quantifies uncertainty in a principled way, and handles object pose tracking under heavy occlusion. Finally, 3DNEL provides a principled framework for incorporating prior knowledge about the scene and objects, which allows natural extension to additional tasks like camera pose tracking from video.
Nishad Gothoskar, Lirui Wang, Josh Tenenbaum, Dan Gutfreund, Miguel Lázaro-Gredilla, Dileep George, Vikash Mansinghka 0001
ICCV6
2023 Learning Noisy OR Bayesian Networks with Max-Product Belief Propagation
abstract
Noisy-OR Bayesian Networks (BNs) are a family of probabilistic graphical models which express rich statistical dependencies in binary data. Variational inference (VI) has been the main method proposed to learn noisy-OR BNs with complex latent structures (Jaakkola & Jordan, 1999; Ji et al., 2020; Buhai et al., 2020). However, the proposed VI approaches either (a) use a recognition network with standard amortized inference that cannot induce "explaining-away"; or (b) assume a simple mean-field (MF) posterior which is vulnerable to bad local optima. Existing MF VI methods also update the MF parameters sequentially which makes them inherently slow. In this paper, we propose parallel max-product as an alternative algorithm for learning noisy-OR BNs with complex latent structures and we derive a fast stochastic training scheme that scales to large datasets. We evaluate both approaches on several benchmarks where VI is the state-of-the-art and show that our method (a) achieves better test performance than Ji et al. (2020) for learning noisy-OR BNs with hierarchical latent structures on large sparse real datasets; (b) recovers a higher number of ground truth parameters than Buhai et al. (2020) from cluttered synthetic scenes; and (c) solves the 2D blind deconvolution problem from Lazaro-Gredilla et al. (2021) and variants - including binary matrix factorization - while VI catastrophically fails and is up to two orders of magnitude slower.
Antoine Dedieu, Dileep George, Miguel Lázaro-Gredilla
ICML4
2023 Schema-learning and rebinding as mechanisms of in-context learning and emergence
abstract
In-context learning (ICL) is one of the most powerful and most unexpected capabilities to emerge in recent transformer-based large language models (LLMs). Yet the mechanisms that underlie it are poorly understood. In this paper, we demonstrate that comparable ICL capabilities can be acquired by an alternative sequence prediction learning method using clone-structured causal graphs (CSCGs). Moreover, a key property of CSCGs is that, unlike transformer-based LLMs, they are {\em interpretable}, which considerably simplifies the task of explaining how ICL works. Specifically, we show that it uses a combination of (a) learning template (schema) circuits for pattern completion, (b) retrieving relevant templates in a context-sensitive manner, and (c) rebinding of novel tokens to appropriate slots in the templates. We go on to marshall evidence for the hypothesis that similar mechanisms underlie ICL in LLMs. For example, we find that, with CSCGs as with LLMs, different capabilities emerge at different levels of overparameterization, suggesting that overparameterization helps in learning more complex template (schema) circuits. By showing how ICL can be achieved with small models and datasets, we open up a path to novel architectures, and take a vital step towards a more general understanding of the mechanics behind this important capability.
Sivaramakrishnan Swaminathan, Antoine Dedieu, Rajkumar Vasudeva Raju, Murray Shanahan, Miguel Lázaro-Gredilla, Dileep George
NeurIPS5
2022 DURableVS: Data-efficient Unsupervised Recalibrating Visual Servoing via online learning in a structured generative model
abstract
Visual servoing enables robotic systems to perform accurate closed-loop control, which is required in many applications. However, existing methods require either precise calibration of the robot kinematic model and cameras or use neural architectures that require large amounts of data to train. In this work, we present a method for unsupervised learning of visual servoing that does not require any prior calibration and is extremely data-efficient. Our key insight is that visual servoing does not depend on identifying the veridical kinematic and camera parameters, but instead only on an accurate generative model of image feature observations from the joint positions of the robot. We demonstrate that with our model architecture and learning algorithm, we can consistently learn accurate models from less than 50 training samples (which amounts to less than 1 min of unsupervised data collection), and that such data-efficient learning is not possible with standard neural architectures. Further, we show that by using the generative model in the loop and learning online, we can enable a robotic system to recover from calibration errors and to detect and quickly adapt to possibly unexpected changes in the robot-camera system (e.g. bumped camera, new objects).
Nishad Gothoskar, Miguel Lázaro-Gredilla, Yasemin Bekiroglu, Josh Tenenbaum, Vikash Mansinghka 0001, Dileep George
ICRA2
2021 Sample-Efficient L0-L2 Constrained Structure Learning of Sparse Ising Models
abstract
We consider the problem of learning the underlying graph of a sparse Ising model with p nodes from n i.i.d. samples. The most recent and best performing approaches combine an empirical loss (the logistic regression loss or the interaction screening loss) with a regularizer (an L1 penalty or an L1 constraint). This results in a convex problem that can be solved separately for each node of the graph. In this work, we leverage the cardinality constraint L0 norm, which is known to properly induce sparsity, and further combine it with an L2 norm to better model the non-zero coefficients. We show that our proposed estimators achieve an improved sample complexity, both (a) theoretically, by reaching new state-of-the-art upper bounds for recovery guarantees, and (b) empirically, by showing sharper phase transitions between poor and full recovery for graph topologies studied in the literature, when compared to their L1-based state-of-the-art methods.
Antoine Dedieu, Miguel Lázaro-Gredilla, Dileep George
AAAI2
2021 Query Training: Learning a Worse Model to Infer Better Marginals in Undirected Graphical Models with Hidden Variables
abstract
Probabilistic graphical models (PGMs) provide a compact representation of knowledge that can be queried in a flexible way: after learning the parameters of a graphical model once, new probabilistic queries can be answered at test time without retraining. However, when using undirected PGMS with hidden variables, two sources of error typically compound in all but the simplest models (a) learning error (both computing the partition function and integrating out the hidden variables is intractable); and (b) prediction error (exact inference is also intractable). Here we introduce query training (QT), a mechanism to learn a PGM that is optimized for the approximate inference algorithm that will be paired with it. The resulting PGM is a worse model of the data (as measured by the likelihood), but it is tuned to produce better marginals for a given inference algorithm. Unlike prior works, our approach preserves the querying flexibility of the original PGM: at test time, we can estimate the marginal of any variable given any partial evidence. We demonstrate experimentally that QT can be used to learn a challenging 8-connected grid Markov random field with hidden variables and that it consistently outperforms the state-of-the-art AdVIL when tested on three undirected models across multiple datasets.
Miguel Lázaro-Gredilla, Wolfgang Lehrach, Nishad Gothoskar, Antoine Dedieu, Dileep George
AAAI1
2021 Perturb-and-max-product: Sampling and learning in discrete energy-based models
abstract
Perturb-and-MAP offers an elegant approach to approximately sample from a energy-based model (EBM) by computing the maximum-a-posteriori (MAP) configuration of a perturbed version of the model. Sampling in turn enables learning. However, this line of research has been hindered by the general intractability of the MAP computation. Very few works venture outside tractable models, and when they do, they use linear programming approaches, which as we will show, have several limitations. In this work we present perturb-and-max-product (PMP), a parallel and scalable mechanism for sampling and learning in discrete EBMs. Models can be arbitrary as long as they are built using tractable factors. We show that (a) for Ising models, PMP is orders of magnitude faster than Gibbs and Gibbs-with-Gradients (GWG) at learning and generating samples of similar or better quality; (b) PMP is able to learn and sample from RBMs; (c) in a large, entangled graphical model in which Gibbs and GWG fail to mix, PMP succeeds.
Miguel Lázaro-Gredilla, Antoine Dedieu, Dileep George
NeurIPS1
2020 A Model of Fast Concept Inference with Object-Factorized Cognitive Programs
Daniel P. Sawyer, Miguel Lázaro-Gredilla, Dileep George
CogSci2
2018 Variational Rejection Sampling
abstract
Learning latent variable models with stochastic variational inference is challenging when the approximate posterior is far from the true posterior, due to high variance in the gradient estimates. We propose a novel rejection sampling step that discards samples from the variational posterior which are assigned low likelihoods by the model. Our approach provides an arbitrarily accurate approximation of the true posterior at the expense of extra computation. Using a new gradient estimator for the resulting unnormalized proposal distribution, we achieve average improvements of 3.71 nats and 0.31 nats over state-of-the-art single-sample and multi-sample alternatives respectively for estimating marginal log-likelihoods using sigmoid belief networks on the MNIST dataset. We show both theoretically and empirically how explicitly rejecting samples, while seemingly challenging to analyze due to the implicit nature of the resulting unnormalized proposal distribution, can have benefits over the competing state-of-the-art alternatives based on importance weighting. We demonstrate the effectiveness of the proposed approach via experiments on synthetic data and a benchmark density estimation task with sigmoid belief networks over the MNIST dataset.
Aditya Grover, Ramki Gummadi, Miguel Lázaro-Gredilla, Dale Schuurmans, Stefano Ermon
AISTATS3
2017 Schema Networks: Zero-shot Transfer with a Generative Causal Model of Intuitive Physics
abstract
The recent adaptation of deep neural network-based methods to reinforcement learning and planning domains has yielded remarkable progress on individual tasks. Nonetheless, progress on task-to-task transfer remains limited. In pursuit of efficient and robust generalization, we introduce the Schema Network, an object-oriented generative physics simulator capable of disentangling multiple causes of events and reasoning backward through causes to achieve goals. The richly structured architecture of the Schema Network can learn the dynamics of an environment directly from data. We compare Schema Networks with Asynchronous Advantage Actor-Critic and Progressive Networks on a suite of Breakout variations, reporting results on training efficiency and zero-shot generalization, consistently demonstrating faster, more robust learning and better transfer. We argue that generalizing from limited data and learning causal relationships are essential abilities on the path toward generally intelligent systems.
Ken Kansky, Tom Silver, David A. Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, D. Scott Phoenix, Dileep George
ICML5
2016 Laplace Approximation for Divisive Gaussian Processes for Nonstationary Regression
abstract
The standard Gaussian Process regression (GP) is usually formulated under stationary hypotheses: The noise power is considered constant throughout the input space and the covariance of the prior distribution is typically modeled as depending only on the difference between input samples. These assumptions can be too restrictive and unrealistic for many real-world problems. Although nonstationarity can be achieved using specific covariance functions, they require a prior knowledge of the kind of nonstationarity, not available for most applications. In this paper we propose to use the Laplace approximation to make inference in a divisive GP model to perform nonstationary regression, including heteroscedastic noise cases. The log-concavity of the likelihood ensures a unimodal posterior and makes that the Laplace approximation converges to a unique maximum. The characteristics of the likelihood also allow to obtain accurate posterior approximations when compared to the Expectation Propagation (EP) approximations and the asymptotically exact posterior provided by a Markov Chain Monte Carlo implementation with Elliptical Slice Sampling (ESS), but at a reduced computational load with respect to both, EP and ESS.
Luis Muñoz-González, Miguel Lázaro-Gredilla, Aníbal R. Figueiras-Vidal
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Biophysical parameter retrieval with warped Gaussian processes
abstract
This paper focuses on biophysical parameter retrieval based on Gaussian Processes (GPs). Very often an arbitrary transformation is applied to the observed variable (e.g. chlorophyll content) to better pose the problem. This standard practice essentially tries to linearize/uniformize the distribution by applying non-linear link functions like the logarithmic, the exponential or the logistic functions. In this paper, we propose to use a GP model that automatically learns the optimal transformation directly from the data. The so-called warped GP regression (WGPR) presented in [1] models output observations as a parametric nonlinear transformation of a GP. The parameters of such prior model are then learned via standard maximum likelihood. We show the good performance of the proposed model for the estimation of oceanic chlorophyll content, which outperforms the regular GPR and a more advanced heteroscedastic GPR model.
Jordi Muñoz-Marí, Jochem Verrelst, Miguel Lázaro-Gredilla, Gustau Camps-Valls
IGARSS3
2015 Local Expectation Gradients for Black Box Variational Inference
abstract
We introduce local expectation gradients which is a general purpose stochastic variational inference algorithm for constructing stochastic gradients by sampling from the variational distribution. This algorithm divides the problem of estimating the stochastic gradients over multiple variational parameters into smaller sub-tasks so that each sub-task explores intelligently the most relevant part of the variational distribution. This is achieved by performing an exact expectation over the single random variable that most correlates with the variational parameter of interest resulting in a Rao-Blackwellized estimate that has low variance. Our method works efficiently for both continuous and discrete random variables. Furthermore, the proposed algorithm has interesting similarities with Gibbs sampling but at the same time, unlike Gibbs sampling, can be trivially parallelized.
Michalis K. Titsias, Miguel Lázaro-Gredilla
NIPS2
2014 Doubly Stochastic Variational Bayes for non-Conjugate Inference
abstract
We propose a simple and effective variational inference algorithm based on stochastic optimisation that can be widely applied for Bayesian non-conjugate inference in continuous parameter spaces. This algorithm is based on stochastic approximation and allows for efficient use of gradient information from the model joint density. We demonstrate these properties using illustrative examples as well as in challenging and diverse Bayesian inference problems such as variable selection in logistic regression and fully Bayesian inference over kernel hyperparameters in Gaussian process regression.
Michalis K. Titsias, Miguel Lázaro-Gredilla
ICML2
2014 Retrieval of Biophysical Parameters With Heteroscedastic Gaussian Processes
abstract
An accurate estimation of biophysical variables is the key to monitor our Planet. Leaf chlorophyll content helps in interpreting the chlorophyll fluorescence signal from space, whereas oceanic chlorophyll concentration allows us to quantify the healthiness of the oceans. Recently, the family of Bayesian nonparametric methods has provided excellent results in these situations. A particularly useful method in this framework is the Gaussian process regression (GPR). However, standard GPR assumes that the variance of the noise process is independent of the signal, which does not hold in most of the problems. In this letter, we propose a nonstandard variational approximation that allows accurate inference in signal-dependent noise scenarios. We show that the so-called variational heteroscedastic GPR (VHGPR) is an excellent alternative to standard GPR in two relevant Earth observation examples, namely, Chl vegetation retrieval from hyperspectral images and oceanic Chl concentration estimation from in situ measured reflectances. The proposed VHGPR outperforms the tested empirical approaches, as well as statistical linear regression (both least squares and least absolute shrinkage and selection operator), neural nets, and kernel ridge regression, and the homoscedastic GPR, in terms of accuracy and bias, and proves more robust when a low number of examples is available.
Miguel Lázaro-Gredilla, Michalis K. Titsias, Jochem Verrelst, Gustau Camps-Valls
IEEE Geosci. Remote. Sens. Lett.1
2014 A Bayesian approach for adaptive multiantenna sensing in cognitive radio networks
Julio Manco-Vásquez, Miguel Lázaro-Gredilla, David Ramírez 0001, Javier Vía, Ignacio Santamaría
Signal Process.2
2014 A Gaussian Process Model for Data Association and a Semidefinite Programming Solution
abstract
In this paper, we propose a Bayesian model for the data association problem, in which trajectory smoothness is enforced through the use of Gaussian process priors. This model allows to score candidate associations using the evidence framework, thus casting the data association problem into an optimization problem. Under some additional mild assumptions, this optimization problem is shown to be equivalent to a constrained Max K-section problem. Furthermore, for K=2, a MaxCut formulation is obtained, to which an approximate solution can be efficiently found using an SDP relaxation. Solving this MaxCut problem is equivalent to finding the optimal association out of the combinatorially many possibilities. The obtained clustering depends only on two hyperparameters, which can also be selected by maximum evidence.
Miguel Lázaro-Gredilla, Steven Van Vaerenbergh
IEEE Trans. Neural Networks Learn. Syst.1
2014 Divisive Gaussian Processes for Nonstationary Regression
abstract
Standard Gaussian process regression (GPR) assumes constant noise power throughout the input space and stationarity when combined with the squared exponential covariance function. This can be unrealistic and too restrictive for many real-world problems. Nonstationarity can be achieved by specific covariance functions, though prior knowledge about this nonstationarity can be difficult to obtain. On the other hand, the homoscedastic assumption is needed to allow GPR inference to be tractable. In this paper, we present a divisive GPR model which performs nonstationary regression under heteroscedastic noise using the pointwise division of two nonparametric latent functions. As the inference on the model is not analytically tractable, we propose a variational posterior approximation using expectation propagation (EP) which allows for accurate inference at reduced cost. We have also made a Markov chain Monte Carlo implementation with elliptical slice sampling to assess the quality of the EP approximation. Experiments support the usefulness of the proposed approach.
Luis Muñoz-González, Miguel Lázaro-Gredilla, Aníbal R. Figueiras-Vidal
IEEE Trans. Neural Networks Learn. Syst.2
2013 Estimation of vegetation chlorophyll content with Variational Heteroscedastic Gaussian Processes
abstract
Accurate estimation of biophysical variables is the key to monitor our Planet. In particular, leaf chlorophyll content helps in interpreting the chlorophyll fluorescence signal from space, which is an accurate indicator of the actual state of the vegetation beyond greenness. Recently, the family of Bayesian nonparametric methods has provided excellent results in these situations. A particularly useful method in this framework is the Gaussian Processes regression (GP). However, standard GP assumes that the variance of the noise process is independent of the signal, which does not hold in most of the problems. In this paper, we propose a non-standard variational approximation that allows accurate inference in signal-dependent noise scenarios. We show that the so-called Variational Heteroscedastic Gaussian Process (VHGP) regression is an excellent alternative to standard GP for the retrieval of vegetation chlorophyll content from hyperspectral images. In general VHGP outperforms GP (and many other empirical and machine learning techniques) in accuracy and bias, and reveals more robust when a low number of examples is available.
Miguel Lázaro-Gredilla, Michalis K. Titsias, Jochem Verrelst, Gustau Camps-Valls
IGARSS1
2013 Variational Inference for Mahalanobis Distance Metrics in Gaussian Process Regression
abstract
We introduce a novel variational method that allows to approximately integrate out kernel hyperparameters, such as length-scales, in Gaussian process regression. This approach consists of a novel variant of the variational framework that has been recently developed for the Gaussian process latent variable model which additionally makes use of a standardised representation of the Gaussian process. We consider this technique for learning Mahalanobis distance metrics in a Gaussian process regression setting and provide experimental evaluations and comparisons with existing methods by considering datasets with high-dimensional inputs.
Michalis K. Titsias, Miguel Lázaro-Gredilla
NIPS2
2012 Bayesian Warped Gaussian Processes
abstract
Warped Gaussian processes (WGP) [1] model output observations in regression tasks as a parametric nonlinear transformation of a Gaussian process (GP). The use of this nonlinear transformation, which is included as part of the probabilistic model, was shown to enhance performance by providing a better prior model on several data sets. In order to learn its parameters, maximum likelihood was used. In this work we show that it is possible to use a non-parametric nonlinear transformation in WGP and variationally integrate it out. The resulting Bayesian WGP is then able to work in scenarios in which the maximum likelihood WGP failed: Low data regime, data with censored values, classification, etc. We demonstrate the superior performance of Bayesian warped GPs on several real data sets.
Miguel Lázaro-Gredilla
NIPS1
2012 Low-cost model selection for SVMs using local features
Miguel Lázaro-Gredilla, Vanessa Gómez-Verdejo, Emilio Parrado-Hernández
Eng. Appl. Artif. Intell.1
2012 Overlapping Mixtures of Gaussian Processes for the data association problem
Miguel Lázaro-Gredilla, Steven Van Vaerenbergh, Neil D. Lawrence
Pattern Recognit.1
2012 Kernel Recursive Least-Squares Tracker for Time-Varying Regression
abstract
In this paper, we introduce a kernel recursive least-squares (KRLS) algorithm that is able to track nonlinear, time-varying relationships in data. To this purpose, we first derive the standard KRLS equations from a Bayesian perspective (including a sensible approach to pruning) and then take advantage of this framework to incorporate forgetting in a consistent way, thus enabling the algorithm to perform tracking in nonstationary scenarios. The resulting method is the first kernel adaptive filtering algorithm that includes a forgetting factor in a principled and numerically stable manner. In addition to its tracking ability, it has a number of appealing properties. It is online, requires a fixed amount of memory and computation per time step, incorporates regularization in a natural manner and provides confidence intervals along with each prediction. We include experimental results that support the theory as well as illustrate the efficiency of the proposed algorithm.
Steven Van Vaerenbergh, Miguel Lázaro-Gredilla, Ignacio Santamaría
IEEE Trans. Neural Networks Learn. Syst.2
2011 Tracking performance of adaptively biased adaptive filters
abstract
Adaptive filters can improve their performance by exploiting the well known tradeoff between bias and variance of the estimated solution. In a previous work, a scheme for adaptively biasing the filter weights was introduced, multiplying the output of a filter of any kind by a shrinking factor a ∈ [0,1]. With an appropriate value a, such a scheme can reduce the steady-state error, especially for low signal-to-noise ra tio (SNR). Here, we extend such analysis for a tracking scenario in which the optimal solution follows a random walk-model. We briefly review a realizable scheme for learning a, based on recently proposed algorithms for adaptive filter combination. Our experiments validate the accurateness of the analysis, and illustrate the performance gains that can be expected from these biased configurations in stationary and tracking scenarios.
Jerónimo Arenas-García, Miguel Lázaro-Gredilla
ICASSP2
2011 Variational Heteroscedastic Gaussian Process Regression
Miguel Lázaro-Gredilla, Michalis K. Titsias
ICML1
2011 Spike and Slab Variational Inference for Multi-Task and Multiple Kernel Learning
abstract
We introduce a variational Bayesian inference algorithm which can be widely applied to sparse linear models. The algorithm is based on the spike and slab prior which, from a Bayesian perspective, is the golden standard for sparse inference. We apply the method to a general multi-task and multiple kernel learning model in which a common set of Gaussian process functions is linearly combined with task-specific sparse weights, thus inducing relation between tasks. This model unifies several sparse linear models, such as generalized linear models, sparse factor analysis and matrix factorization with missing values, so that the variational algorithm can be applied to all these cases. We demonstrate our approach in multi-output Gaussian process regression, multi-class classification, image processing applications and collaborative filtering.
Michalis K. Titsias, Miguel Lázaro-Gredilla
NIPS2
2011 Support Vector Machines With Constraints for Sparsity in the Primal Parameters
abstract
This paper introduces a new support vector machine (SVM) formulation to obtain sparse solutions in the primal SVM parameters, providing a new method for feature selection based on SVMs. This new approach includes additional constraints to the classical ones that drop the weights associated to those features that are likely to be irrelevant. A ν-SVM formulation has been used, where ν indicates the fraction of features to be considered. This paper presents two versions of the proposed sparse classifier, a 2-norm SVM and a 1-norm SVM, the latter having a reduced computational burden with respect to the first one. Additionally, an explanation is provided about how the presented approach can be readily extended to multiclass classification or to problems where groups of features, rather than isolated features, need to be selected. The algorithms have been tested in a variety of synthetic and real data sets and they have been compared against other state of the art SVM-based linear feature selection methods, such as 1-norm SVM and doubly regularized SVM. The results show the good feature selection ability of the approaches.
Vanessa Gómez-Verdejo, Manel Martínez-Ramón, Jerónimo Arenas-García, Miguel Lázaro-Gredilla, Harold Y. Molina-Bulla
IEEE Trans. Neural Networks4
2010 Sparse Spectrum Gaussian Process Regression
Miguel Lázaro-Gredilla, Joaquin Quiñonero Candela, Carl E. Rasmussen, Aníbal R. Figueiras-Vidal
J. Mach. Learn. Res.1
2010 Marginalized neural network mixtures for large-scale regression
abstract
For regression tasks, traditional neural networks (NNs) have been superseded by gaussian processes, which provide probabilistic predictions (input-dependent error bars), improved accuracy, and virtually no overfitting. Due to their high computational cost, in scenarios with massive data sets, one has to resort to sparse gaussian processes, which strive to achieve similar performance with much smaller computational effort. In this context, we introduce a mixture of NNs with marginalized output weights that can both provide probabilistic predictions and improve on the performance of sparse gaussian processes, at the same computational cost. The effectiveness of this approach is shown experimentally on some representative large data sets.
Miguel Lázaro-Gredilla, Aníbal R. Figueiras-Vidal
IEEE Trans. Neural Networks1
2009 Inter-domain Gaussian Processes for Sparse Inference using Inducing Features
abstract
We present a general inference framework for inter-domain Gaussian Processes (GPs), focusing on its usefulness to build sparse GP models. The state-of-the-art sparse GP model introduced by Snelson and Ghahramani in [1] relies on finding a small, representative pseudo data set of m elements (from the same domain as the n available data elements) which is able to explain existing data well, and then uses it to perform inference. This reduces inference and model selection computation time from O(n^3) to O(m^2n), where m << n. Inter-domain GPs can be used to find a (possibly more compact) representative set of features lying in a different domain, at the same computational cost. Being able to specify a different domain for the representative features allows to incorporate prior knowledge about relevant characteristics of data and detaches the functional form of the covariance and basis functions. We will show how previously existing models fit into this framework and will use it to develop two new sparse GP models. Tests on large, representative regression data sets suggest that significant improvement can be achieved, while retaining computational efficiency.
Miguel Lázaro-Gredilla, Aníbal R. Figueiras-Vidal
NIPS1