Daniel Hernández-Lobato

dblp:95/166 · DBLP profile ↗
← Back
62ranked-venue papers
21as first author
21since 2021 · last 2026
0000-0001-5845-437XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 19 first-author · 21 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2026 Improving the Linearized Laplace Approximation via Quadratic Approximations
abstract
Deep neural networks (DNNs) often produce overconfident out-of-distribution predictions, motivating Bayesian uncertainty quantification.The Linearized Laplace Approximation (LLA) achieves this by linearizing the DNN and applying Laplace inference to the resulting model.Importantly, the linear model is also used for prediction.We argue this linearization in the posterior may degrade fidelity to the true Laplace approximation.To alleviate this problem, without increasing significantly the computational cost, we propose the Quadratic Laplace Approximation (QLA).QLA approximates each second order factor in the approximate Laplace log-posterior using a rank-one factor obtained via efficient power iterations.QLA is expected to yield a posterior precision closer to that of the full Laplace without forming the full Hessian, which is typically intractable.For prediction, QLA also uses the linearized model.Empirically, QLA yields modest yet consistent uncertainty estimation improvements over LLA on five regression datasets.
Pedro Jiménez García-Ligero, Luis A. Ortega 0001, Pablo Morales-Alvarez, Daniel Hernández-Lobato
ESANN4
2026 Scalable Linearized Laplace Approximation via Surrogate Neural Kernel
abstract
We introduce a scalable method to approximate the kernel of the Linearized Laplace Approximation (LLA).For this, we use a surrogate deep neural network (DNN) that learns a compact feature representation whose inner product replicates the Neural Tangent Kernel (NTK).This avoids the need to compute large Jacobians.Training relies solely on efficient Jacobian-vector products, allowing to compute predictive uncertainty on large-scale pre-trained DNNs.Experimental results show similar or improved uncertainty estimation and calibration compared to existing LLA approximations.Notwithstanding, biasing the learned kernel significantly enhances out-of-distribution detection.This remarks the benefits of the proposed method for finding better kernels than the NTK in the context of LLA to compute prediction uncertainty given a pre-trained DNN.
Luis A. Ortega 0001, Simón Rodríguez Santana, Daniel Hernández-Lobato
ESANN3
2026 Robust-multi-task gradient boosting
abstract
The objective of this study is to develop a robust boosting framework capable of handling heterogeneous and outlier tasks in Multi-Task Learning (MTL). Conventional MTL methods assume strong relatedness among tasks, which often fails in real-world scenarios involving adversarial or unaligned tasks that degrade performance. To address this limitation, we propose Robust Multi-Task Gradient Boosting (R-MTGB), a novel ensemble framework that explicitly models task heterogeneity within the gradient boosting paradigm. The methodology structures learning into three sequential stages: (1) shared representation learning to extract common patterns across tasks, (2) outlier-aware partitioning using a learnable task-specific parameter to separate and reweight outlier and non-outlier tasks, and (3) task-specific fine-tuning to refine individual predictors. Extensive experiments on both synthetic and real-world datasets demonstrate that R-MTGB consistently improves predictive accuracy, effectively identifies outlier tasks, and enhances generalization compared to state-of-the-art methods. The achieved results confirm that R-MTGB not only ensures robust performance and interpretability through task-level outlier scores but also provides a scalable and principled framework for reliable multi-task learning in heterogeneous environments.
Seyedsaman Emami, Gonzalo Martínez-Muñoz, Daniel Hernández-Lobato
Expert Syst. Appl.3
2026 Robust multi-task boosting using clustering and local ensembling
abstract
Multi-Task Learning (MTL) aims to boost predictive performance by sharing information across related tasks, however, conventional methods often suffer from negative transfer when unrelated or noisy tasks are forced to share representations. We propose Robust Multi-Task Boosting using Clustering and Local Ensembling (RMB-CLE), a principled MTL framework that integrates error-based task clustering with local ensembling. Unlike prior work that assumes fixed clusters or hand-crafted similarity metrics, RMB-CLE derives inter-task similarity directly from cross-task errors , which admits a principled characterization of cross-task risk in terms of functional mismatch and uncertainty, providing a theoretically grounded mechanism to prevent negative transfer. Tasks are grouped adaptively via agglomerative clustering, and within each cluster, a local ensemble enables robust knowledge sharing while preserving task-specific patterns. Experiments show that RMB-CLE recovers ground-truth clusters in synthetic data and consistently outperforms multi-task, single-task, and pooling-based ensemble methods across diverse synthetic and real-world tabular benchmarks. These results demonstrate that RMB-CLE provides an effective boosting-based framework for robust multi-task learning on structured tabular datasets, where task heterogeneity and negative transfer are central challenges.
Seyedsaman Emami, Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz
Neurocomputing2
2026 Impact of evaluation noise in the context of multi-objective Bayesian optimization with the intrinsic coregionalization model
abstract
Multi-objective Bayesian optimization (MOBO) is a powerful framework for optimizing black-box functions with conflicting objectives and costly evaluations. A key component in MOBO is the probabilistic surrogate model used to guide the search for the optimum. A simple approach is to use independent Gaussian processes (GPs), one per objective. A more powerful method is to use multi-output GPs, such as those based on the intrinsic coregionalization model (ICM), to exploit correlations between objectives. However, with this model, often independent observation noise across objectives is assumed, which ignores part of the correlation structure. In this work, we isolate and analyze the role of observational noise in MOBO with ICM-based multi-output GPs. We consider models that can capture correlations in the observational noise, and study how the presence and structure of noise affect the joint predictive distribution, the considered acquisition function, and ultimately the optimization performance.Through theoretical analysis and extensive synthetic and real-world experiments, we show that multi-output GPs can significantly improve optimization efficiency when the evaluations are noisy. Importantly, however, when the evaluations are noiseless, we observe, both theoretically and empirically, only small gains by considering multi-output GPs, and cheaper independent GPs may be preferred. Furthermore, our synthetic experiments show that explicitly modeling noise correlations provides additional advantages when the objectives themselves are negatively correlated. In real-world problems, however, we have observed mixed results regarding modeling noise correlations. There are small gains in an ensemble benchmark and clearer gains, similar to those observed in the synthetic experiments, in a neural-network benchmark. These findings provide practical insights into when and how to use ICMs for MOBO.
Daniel Fernández-Sánchez, Carlos Sevilla-Salcedo, Vanessa Gómez-Verdejo, Daniel Hernández-Lobato
Neurocomputing4
2025 Joint entropy search for multi-objective Bayesian optimization with constraints and multiple fidelities
abstract
Bayesian optimization (BO) methods can be used to solve efficiently problems with several objectives and constraints. Each objective and constraint is considered a black-box function that is expensive to evaluate, lacking also a closed-form expression. BO methods use a model of each black-box to guide the search for the problem’s solution. Specifically, they make intelligent decisions about where each black-box function should be evaluated next with the goal of finding the solution using a few evaluations only. Sometimes, however, the black-boxes may be evaluated at different fidelity levels. A lower fidelity is simply a cheap proxy of the corresponding black-box. These lower fidelities correlate with the actual black-boxes to optimize and can, therefore, be used to reduce the overall cost of solving the optimization problem. Here, we propose Multi-fidelity Joint Entropy Search for Multi-objective Bayesian Optimization with Constraints (MF-JESMOC), a BO method for solving the aforementioned problems. MF-JESMOC chooses the next point, and fidelity level at which to evaluate the black-boxes, as the combination that is expected to reduce the most the joint entropy of the Pareto set and the Pareto front, normalized by the fidelity’s evaluation cost. We use Deep Gaussian processes to model each black-box and the dependencies between fidelities. These are powerful probabilistic models that can learn the dependency structure among fidelity levels of each black-box. Several experiments show that MF-JESMOC outperforms other state-of-the-art methods for multi-objective BO with constraints and different fidelity levels in both synthetic and real-world problems. • MF-JESMOC is a multi-fidelity method for constrained multi-objective BO. • It uses an MF-DGP to model linear and non-linear dependencies among fidelities. • It selects the next point and associated fidelity to reduce the solution’s entropy. • Synthetic and real-world experiments show that it outperforms previous SOTA methods.
Daniel Fernández-Sánchez, Daniel Hernández-Lobato
Neurocomputing2
2025 Alpha entropy search for new information-based Bayesian optimization
abstract
Bayesian optimization (BO) methods based on information theory have obtained state-of-the-art results in several tasks. These techniques rely on the Kullback–Leibler (KL) divergence to compute the acquisition function. We introduce a novel information-based class of acquisition functions for BO called Alpha Entropy Search (AES). AES is based on the alpha-divergence, which generalizes the KL-divergence. Iteratively, AES selects the next evaluation point as the one whose associated target value has the highest level of dependency with respect to the location and associated value of the global maximum of the optimization problem. Dependency is measured in terms of the alpha-divergence, as an alternative to the KL-divergence. Intuitively, this favors evaluating the objective function at the most informative points about the global maximum. The alpha-divergence has a free parameter α , which determines the behavior of the divergence, balancing local and global differences. Therefore, different values of α result in different acquisition functions. AES acquisition lacks a closed-form expression. However, we propose an efficient and accurate approximation using a truncated Gaussian distribution. In practice, the value of α can be chosen by the practitioner, but here we suggest using a combination of acquisition functions obtained by simultaneously considering a range of α values. We provide an implementation of AES in BOTorch and we evaluate its performance in synthetic, benchmark, and real-world experiments involving the tuning of the hyper-parameters of a deep neural network. These experiments show that AES performance is competitive with other information-based acquisition functions such as JES, MES, or PES.
Daniel Fernández-Sánchez, Eduardo C. Garrido-Merchán, Daniel Hernández-Lobato
Knowl. Based Syst.3
2024 Joint Entropy Search for Multi-objective Bayesian Optimization with Constraints and Multiple Fidelities
abstract
Bayesian optimization (BO) methods solve problems with several black-box objectives and constraints.Each black-box is expensive to evaluate and lacks a closed-form.They use a model of each black-box to guide the search for the problem's solution.Sometimes, however, the black-boxes may be evaluated at different fidelity levels.A lower fidelity is simply a cheap proxy of the corresponding black-box.Thus, lower fidelities that correlate with the actual black-box can be used to reduce the optimization cost.We propose Joint Entropy Search for Multi-Fidelity and Multi-objective Bayesian Optimization with Constraints (MF-JESMOC), a BO method for solving the aforementioned problems.It chooses the next point and fidelity level at which to evaluate the black-boxes as the one that is expected to reduce the most the joint entropy of the Pareto set and the Pareto front, normalized by the fidelity's cost.Deep Gaussian processes are used to model each black-box and dependencies between fidelities.In our experiments, MF-JESMOC outperforms other state-of-the-art methods for multi-objective BO with constraints and different fidelity levels.
Daniel Fernández-Sánchez, Daniel Hernández-Lobato
ESANN2
2024 Variational Linearized Laplace Approximation for Bayesian Deep Learning
abstract
The Linearized Laplace Approximation (LLA) has been recently used to perform uncertainty estimation on the predictions of pre-trained deep neural networks (DNNs). However, its widespread application is hindered by significant computational costs, particularly in scenarios with a large number of training points or DNN parameters. Consequently, additional approximations of LLA, such as Kronecker-factored or diagonal approximate GGN matrices, are utilized, potentially compromising the model’s performance. To address these challenges, we propose a new method for approximating LLA using a variational sparse Gaussian Process (GP). Our method is based on the dual RKHS formulation of GPs and retains as the predictive mean the output of the original DNN. Furthermore, it allows for efficient stochastic optimization, which results in sub-linear training time in the size of the training dataset. Specifically, its training cost is independent of the number of training points. We compare our proposed method against accelerated LLA (ELLA), which relies on the Nyström approximation, as well as other LLA variants employing the sample-then-optimize principle. Experimental results, both on regression and classification datasets, show that our method outperforms these already existing efficient variants of LLA, both in terms of the quality of the predictive distribution and in terms of total computational time.
Luis A. Ortega 0001, Simón Rodríguez Santana, Daniel Hernández-Lobato
ICML3
2023 Deep Variational Implicit Processes
Luis A. Ortega 0001, Simón Rodríguez Santana, Daniel Hernández-Lobato
ICLR3
2023 Efficient Transformed Gaussian Processes for Non-Stationary Dependent Multi-class Classification
abstract
This work introduces the Efficient Transformed Gaussian Process (ETGP), a new way of creating $C$ stochastic processes characterized by: 1) the $C$ processes are non-stationary, 2) the $C$ processes are dependent by construction without needing a mixing matrix, 3) training and making predictions is very efficient since the number of Gaussian Processes (GP) operations (e.g. inverting the inducing point's covariance matrix) do not depend on the number of processes. This makes the ETGP particularly suited for multi-class problems with a very large number of classes, which are the problems studied in this work. ETGP exploits the recently proposed Transformed Gaussian Process (TGP), a stochastic process specified by transforming a Gaussian Process using an invertible transformation. However, unlike TGP, ETGP is constructed by transforming a single sample from a GP using $C$ invertible transformations. We derive an efficient sparse variational inference algorithm for the proposed model and demonstrate its utility in 5 classification tasks which include low/medium/large datasets and a different number of classes, ranging from just a few to hundreds. Our results show that ETGP, in general, outperforms state-of-the-art methods for multi-class classification based on GPs, and has a lower computational cost (around one order of magnitude smaller).
Juan Maroñas Molano, Daniel Hernández-Lobato
ICML2
2023 Parallel predictive entropy search for multi-objective Bayesian optimization with constraints applied to the tuning of machine learning algorithms
Eduardo C. Garrido-Merchán, Daniel Fernández-Sánchez, Daniel Hernández-Lobato
Expert Syst. Appl.3
2023 Improved max-value entropy search for multi-objective bayesian optimization with constraints
abstract
We present MESMOC+, an improved version of Max-value Entropy search for Multi-Objective Bayesian optimization with Constraints (MESMOC). MESMOC+ can be used to solve constrained multi-objective problems when the objectives and the constraints are expensive to evaluate. It is based on minimizing the entropy of the solution of the optimization problem in function space (i.e., the Pareto front) to guide the search for the optimum. The cost of MESMOC+ is linear in the number of objectives and constraints. Furthermore, it is often significantly smaller than the cost of alternative methods based on minimizing the entropy of the Pareto set. The reason for this is that it is easier to approximate the required computations in MESMOC+. Moreover, MESMOC+’s acquisition function is expressed as the sum of one acquisition per each black-box (objective or constraint). Therefore, it can be used in a decoupled evaluation setting in which it is chosen not only the next input location to evaluate, but also which black-box to evaluate there. We compare MESMOC+ with related methods in synthetic, benchmark and real optimization problems. These experiments show that MESMOC+ has similar performance to that of state-of-the-art acquisitions based on entropy search, but it is faster to execute and simpler to implement. Moreover, our experiments also show that MESMOC+ is more robust with respect to the number of samples of the Pareto front.
Daniel Fernández-Sánchez, Eduardo C. Garrido-Merchán, Daniel Hernández-Lobato
Neurocomputing3
2023 Gaussian processes for missing value imputation
abstract
A missing value indicates that a particular attribute of an instance of a learning problem is not recorded. They are very common in many real-life datasets. In spite of this, however, most machine learning methods cannot handle missing values. Thus, they should be imputed before training. Gaussian Processes (GPs) are non-parametric models with accurate uncertainty estimates that combined with sparse approximations and stochastic variational inference scale to large data sets. Sparse GPs (SGPs) can be used to get a predictive distribution for missing values. We present a hierarchical composition of sparse GPs that is used to predict the missing values at each dimension using the observed values from the other dimensions. Importantly, we consider that the input attributes to each sparse GP used for prediction may also have missing values. The missing values in those input attributes are replaced by the predictions of the previous sparse GPs in the hierarchy. We call our approach missing GP (MGP). MGP can impute all observed missing values. It outputs a predictive distribution for each missing value that is then used in the imputation of other missing values. We evaluate MGP on one private clinical data set and on four UCI datasets with a different percentage of missing values. Furthermore, we compare the performance of MGP with other state-of-the-art methods for imputing missing values, including variants based on sparse GPs and deep GPs. Our results show that the performance of MGP is significantly better.
Bahram Jafrasteh, Daniel Hernández-Lobato, Simón Pedro Lubián-López, Isabel Benavente-Fernández
Knowl. Based Syst.2
2023 Inference over radiative transfer models using variational and expectation maximization methods
abstract
Earth observation from satellites offers the possibility to monitor our planet with unprecedented accuracy. Radiative transfer models (RTMs) encode the energy transfer through the atmosphere, and are used to model and understand the Earth system, as well as to estimate the parameters that describe the status of the Earth from satellite observations by inverse modeling. However, performing inference over such simulators is a challenging problem. RTMs are nonlinear, non-differentiable and computationally costly codes, which adds a high level of difficulty in inference. In this paper, we introduce two computational techniques to infer not only point estimates of biophysical parameters but also their joint distribution. One of them is based on a variational autoencoder approach and the second one is based on a Monte Carlo Expectation Maximization (MCEM) scheme. We compare and discuss benefits and drawbacks of each approach. We also provide numerical comparisons in synthetic simulations and the real PROSAIL model, a popular RTM that combines land vegetation leaf and canopy modeling. We analyze the performance of the two approaches for modeling and inferring the distribution of three key biophysical parameters for quantifying the terrestrial biosphere.
Daniel H. Svendsen, Daniel Hernández-Lobato, Luca Martino, Valero Laparra, Álvaro Moreno-Martínez, Gustau Camps-Valls
Mach. Learn.2
2022 Input Dependent Sparse Gaussian Processes
abstract
Gaussian Processes (GPs) are non-parametric models that provide accurate uncertainty estimates. Nevertheless, they have a cubic cost in the number of data instances $N$. To overcome this, sparse GP approximations are used, in which a set of $M \ll N$ inducing points is introduced. The location of the inducing points is learned by considering them parameters of an approximate posterior distribution $q$. Sparse GPs, combined with stochastic variational inference for inferring $q$ have a cost per iteration in $\mathcal{O}(M^3)$. Critically, the inducing points determine the flexibility of the model and they are often located in regions where the latent function changes. A limitation is, however, that in some tasks a large number of inducing points may be required to obtain good results. To alleviate this, we propose here to amortize the computation of the inducing points locations, as well as the parameters of $q$. For this, we use a neural network that receives a data instance as an input and outputs the corresponding inducing points locations and the parameters of $q$. We evaluate our method in several experiments, showing that it performs similar or better than other state-of-the-art sparse variational GPs. However, in our method the number of inducing points is reduced drastically since they depend on the input data. This makes our method scale to larger datasets and have faster training and prediction times.
Bahram Jafrasteh, Carlos Villacampa-Calvo, Daniel Hernández-Lobato
ICML3
2022 Function-space Inference with Sparse Implicit Processes
abstract
Implicit Processes (IPs) represent a flexible framework that can be used to describe a wide variety of models, from Bayesian neural networks, neural samplers and data generators to many others. IPs also allow for approximate inference in function-space. This change of formulation solves intrinsic degenerate problems of parameter-space approximate inference concerning the high number of parameters and their strong dependencies in large models. For this, previous works in the literature have attempted to employ IPs both to set up the prior and to approximate the resulting posterior. However, this has proven to be a challenging task. Existing methods that can tune the prior IP result in a Gaussian predictive distribution, which fails to capture important data patterns. By contrast, methods producing flexible predictive distributions by using another IP to approximate the posterior process cannot tune the prior IP to the observed data. We propose here the first method that can accomplish both goals. For this, we rely on an inducing-point representation of the prior IP, as often done in the context of sparse Gaussian processes. The result is a scalable method for approximate inference with IPs that can tune the prior IP parameters to the data, and that provides accurate non-Gaussian predictive distributions.
Simón Rodríguez Santana, Bryan Zaldivar, Daniel Hernández-Lobato
ICML3
2022 Alpha-divergence minimization for deep Gaussian processes
abstract
This paper proposes the minimization of α-divergences for approximate inference in the context of deep Gaussian processes (DGPs). The proposed method can be considered as a generalization of variational inference (VI) and expectation propagation (EP), two previously used methods for approximate inference in DGPs. Both VI and EP are based on the minimization of the Kullback-Leibler divergence. The proposed method is based on a scalable version of power expectation propagation, a method that introduces an extra parameter α that specifies the targeted α-divergence to be optimized. In particular, such a method can recover the VI solution when α→0 and the EP solution when α→1. An exhaustive experimental evaluation shows that the minimization of α-divergences via the proposed method is feasible in DGPs and that choosing intermediate values of the α parameter between 0 and 1 can give better results in some problems. This means that one can improve the results of VI and EP when training DGPs. Importantly, the proposed method allows for stochastic optimization techniques, making it able to address datasets with several millions of instances.
Carlos Villacampa-Calvo, Gonzalo Hernández-Muñoz, Daniel Hernández-Lobato
Int. J. Approx. Reason.3
2022 Adversarial α-divergence minimization for Bayesian approximate inference
Simón Rodríguez Santana, Daniel Hernández-Lobato
Neurocomputing2
2021 Activation-level uncertainty in deep neural networks
Pablo Morales-Alvarez, Daniel Hernández-Lobato, Rafael Molina 0001, José Miguel Hernández-Lobato
ICLR2
2021 Multi-class Gaussian Process Classification with Noisy Inputs
abstract
It is a common practice in the machine learning community to assume that the observed data are noise-free in the input attributes. Nevertheless, scenarios with input noise are common in real problems, as measurements are never perfectly accurate. If this input noise is not taken into account, a supervised machine learning method is expected to perform sub-optimally. In this paper, we focus on multi-class classification problems and use Gaussian processes (GPs) as the underlying classifier. Motivated by a data set coming from the astrophysics domain, we hypothesize that the observed data may contain noise in the inputs. Therefore, we devise several multi-class GP classifiers that can account for input noise. Such classifiers can be efficiently trained using variational inference to approximate the posterior distribution of the latent variables of the model. Moreover, in some situations, the amount of noise can be known before-hand. If this is the case, it can be readily introduced in the proposed methods. This prior information is expected to lead to better performance results. We have evaluated the proposed methods by carrying out several experiments, involving synthetic and real data. These include several data sets from the UCI repository, the MNIST data set and a data set coming from astrophysics. The results obtained show that, although the classification error is similar across methods, the predictive distribution of the proposed methods is better, in terms of the test log-likelihood, than the predictive distribution of a classifier based on GPs that ignores input noise.
Carlos Villacampa-Calvo, Bryan Zaldivar, Eduardo C. Garrido-Merchán, Daniel Hernández-Lobato
J. Mach. Learn. Res.4
2020 Deep Gaussian Processes Using Expectation Propagation and Monte Carlo Methods
Gonzalo Hernández-Muñoz, Carlos Villacampa-Calvo, Daniel Hernández-Lobato
ECML/PKDD (3)3
2020 Dealing with categorical and integer-valued variables in Bayesian Optimization with Gaussian processes
Eduardo C. Garrido-Merchán, Daniel Hernández-Lobato
Neurocomputing2
2020 Alpha divergence minimization in multi-class Gaussian process classification
Carlos Villacampa-Calvo, Daniel Hernández-Lobato
Neurocomputing2
2019 Predictive Entropy Search for Multi-objective Bayesian Optimization with Constraints
Eduardo C. Garrido-Merchán, Daniel Hernández-Lobato
Neurocomputing2
2018 Bayesian optimization of a hybrid system for robust ocean wave features prediction
Laura Cornejo-Bueno, Eduardo C. Garrido-Merchán, Daniel Hernández-Lobato, Sancho Salcedo-Sanz
Neurocomputing3
2017 Scalable Multi-Class Gaussian Process Classification using Expectation Propagation
abstract
This paper describes an expectation propagation (EP) method for multi-class classification with Gaussian processes that scales well to very large datasets. In such a method the estimate of the log-marginal-likelihood involves a sum across the data instances. This enables efficient training using stochastic gradients and mini-batches. When this type of training is used, the computational cost does not depend on the number of data instances N. Furthermore, extra assumptions in the approximate inference process make the memory cost independent of N. The consequence is that the proposed EP method can be used on datasets with millions of instances. We compare empirically this method with alternative approaches that approximate the required computations using variational inference. The results show that it performs similar or even better than these techniques, which sometimes give significantly worse predictive distributions in terms of the test log-likelihood. Besides this, the training process of the proposed approach also seems to converge in a smaller number of iterations.
Carlos Villacampa-Calvo, Daniel Hernández-Lobato
ICML2
2016 Scalable Gaussian Process Classification via Expectation Propagation
abstract
Variational methods have been recently considered for scaling the training process of Gaussian process classifiers to large datasets. As an alternative, we describe here how to train these classifiers efficiently using expectation propagation (EP). The proposed EP method allows to train Gaussian process classifiers on very large datasets, with millions of instances, that were out of the reach of previous implementations of EP. More precisely, it can be used for (i) training in a distributed fashion where the data instances are sent to different nodes in which the required computations are carried out, and for (ii) maximizing an estimate of the marginal likelihood using a stochastic approximation of the gradient. Several experiments involving large datasets show that the method described is competitive with the variational approach.
Daniel Hernández-Lobato, José Miguel Hernández-Lobato
AISTATS1
2016 Ambiguity Helps: Classification with Disagreements in Crowdsourced Annotations
abstract
Imagine we show an image to a person and ask her/him to decide whether the scene in the image is warm or not warm, and whether it is easy or not to spot a squirrel in the image. For exactly the same image, the answers to those questions are likely to differ from person to person. This is because the task is inherently ambiguous. Such an ambiguous, therefore challenging, task is pushing the boundary of computer vision in showing what can and can not be learned from visual data. Crowdsourcing has been invaluable for collecting annotations. This is particularly so for a task that goes beyond a clear-cut dichotomy as multiple human judgments per image are needed to reach a consensus. This paper makes conceptual and technical contributions. On the conceptual side, we define disagreements among annotators as privileged information about the data instance. On the technical side, we propose a framework to incorporate annotation disagreements into the classifiers. The proposed framework is simple, relatively fast, and outperforms classifiers that do not take into account the disagreements, especially if tested on high confidence annotations.
Viktoriia Sharmanska, Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Novi Quadrianto
CVPR2
2016 Deep Gaussian Processes for Regression using Approximate Expectation Propagation
abstract
Deep Gaussian processes (DGPs) are multi-layer hierarchical generalisations of Gaussian processes (GPs) and are formally equivalent to neural networks with multiple, infinitely wide hidden layers. DGPs are nonparametric probabilistic models and as such are arguably more flexible, have a greater capacity to generalise, and provide better calibrated uncertainty estimates than alternative deep models. This paper develops a new approximate Bayesian learning scheme that enables DGPs to be applied to a range of medium to large scale regression problems for the first time. The new method uses an approximate Expectation Propagation procedure and a novel and efficient extension of the probabilistic backpropagation algorithm for learning. We evaluate the new method for non-linear regression on eleven real-world datasets, showing that it always outperforms GP regression and is almost always better than state-of-the-art deterministic and sampling-based approximate inference methods for Bayesian neural networks. As a by-product, this work provides a comprehensive analysis of six approximate Bayesian methods for training neural networks.
Thang D. Bui, Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Yingzhen Li, Richard E. Turner
ICML2
2016 Predictive Entropy Search for Multi-objective Bayesian Optimization
abstract
We present \small PESMO, a Bayesian method for identifying the Pareto set of multi-objective optimization problems, when the functions are expensive to evaluate. \small PESMO chooses the evaluation points to maximally reduce the entropy of the posterior distribution over the Pareto set. The \small PESMO acquisition function is decomposed as a sum of objective-specific acquisition functions, which makes it possible to use the algorithm in \emphdecoupled scenarios in which the objectives can be evaluated separately and perhaps with different costs. This decoupling capability is useful to identify difficult objectives that require more evaluations. \small PESMO also offers gains in efficiency, as its cost scales linearly with the number of objectives, in comparison to the exponential cost of other methods. We compare \small PESMO with other methods on synthetic and real-world problems. The results show that \small PESMO produces better recommendations with a smaller number of evaluations, and that a decoupled evaluation can lead to improvements in performance, particularly when the number of objectives is large.
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Amar Shah 0001, Ryan P. Adams
ICML1
2016 Black-Box Alpha Divergence Minimization
abstract
Black-box alpha (BB-α) is a new approximate inference method based on the minimization of α-divergences. BB-αscales to large datasets because it can be implemented using stochastic gradient descent. BB-αcan be applied to complex probabilistic models with little effort since it only requires as input the likelihood function and its gradients. These gradients can be easily obtained using automatic differentiation. By changing the divergence parameter α, the method is able to interpolate between variational Bayes (VB) (α→0) and an algorithm similar to expectation propagation (EP) (α= 1). Experiments on probit regression and neural network regression and classification problems show that BB-αwith non-standard settings of α, such as α= 0.5, usually produces better predictions than with α→0 (VB) or α= 1 (EP).
José Miguel Hernández-Lobato, Yingzhen Li, Mark Rowland 0001, Thang D. Bui, Daniel Hernández-Lobato, Richard E. Turner
ICML5
2016 Non-linear Causal Inference using Gaussianity Measures
abstract
We provide theoretical and empirical evidence for a type of asymmetry between causes and effects that is present when these are related via linear models contaminated with additive non- Gaussian noise. Assuming that the causes and the effects have the same distribution, we show that the distribution of the residuals of a linear fit in the anti-causal direction is closer to a Gaussian than the distribution of the residuals in the causal direction. This Gaussianization effect is characterized by reduction of the magnitude of the high-order cumulants and by an increment of the differential entropy of the residuals. The problem of non-linear causal inference is addressed by performing an embedding in an expanded feature space, in which the relation between causes and effects can be assumed to be linear. The effectiveness of a method to discriminate between causes and effects based on this type of asymmetry is illustrated in a variety of experiments using different measures of Gaussianity. The proposed method is shown to be competitive with state-of-the-art techniques for causal inference.
Daniel Hernández-Lobato, Pablo Morales-Mombiela, David Lopez-Paz, Alberto Suárez 0001
J. Mach. Learn. Res.1
2015 A Probabilistic Model for Dirty Multi-task Feature Selection
abstract
Multi-task feature selection methods often make the hypothesis that learning tasks share relevant and irrelevant features. However, this hypothesis may be too restrictive in practice. For example, there may be a few tasks with specific relevant and irrelevant features (outlier tasks). Similarly, a few of the features may be relevant for only some of the tasks (outlier features). To account for this, we propose a model for multi-task feature selection based on a robust prior distribution that introduces a set of binary latent variables to identify outlier tasks and outlier features. Expectation propagation can be used for efficient approximate inference under the proposed prior. Several experiments show that a model based on the new robust prior provides better predictive performance than other benchmark methods.
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Zoubin Ghahramani
ICML1
2015 Special Issue on "Solving complex machine learning problems with ensemble methods"
Daniel Hernández-Lobato, Ioannis Katakis 0001, Gonzalo Martínez-Muñoz, Ioannis Partalas
Neurocomputing1
2015 Expectation propagation in linear regression models with spike-and-slab priors
José Miguel Hernández-Lobato, Daniel Hernández-Lobato, Alberto Suárez 0001
Mach. Learn.2
2014 Mind the Nuisance: Gaussian Process Classification using Privileged Noise
Daniel Hernández-Lobato, Viktoriia Sharmanska, Kristian Kersting, Christoph H. Lampert, Novi Quadrianto
NIPS1
2014 A Double Pruning Scheme for Boosting Ensembles
abstract
Ensemble learning consists of generating a collection of classifiers whose predictions are then combined to yield a single unified decision. Ensembles of complementary classifiers provide accurate and robust predictions, which are often better than the predictions of the individual classifiers in the ensemble. Nevertheless, ensembles also have some drawbacks: typically, all classifiers are queried to compute the final ensemble prediction. Therefore, all the classifiers need to be accessible to address potential queries. This entails larger storage requirements and slower predictions than a single classifier. Ensemble pruning techniques are useful to alleviate these drawbacks. Static pruning techniques reduce the ensemble size by selecting a sub-ensemble of classifiers from the original ensemble. In dynamic pruning, the querying process is halted when the partial ensemble prediction is sufficient to reach a stable final decision with a reasonable amount of confidence. In this paper, we present the results of a comprehensive analysis of static and dynamic pruning techniques applied to Adaboost ensembles. These ensemble pruning techniques are evaluated on a wide range of classification problems. From this analysis, one concludes that the combination of static and dynamic pruning techniques provides a notable reduction in the memory requirements and an improvement in the classification time without a significant loss of prediction accuracy.
Víctor Soto, Sergio García-Moratilla, Gonzalo Martínez-Muñoz, Daniel Hernández-Lobato, Alberto Suárez 0001
IEEE Trans. Cybern.4
2013 Statistical Tests for the Detection of the Arrow of Time in Vector Autoregressive Models
Pablo Morales-Mombiela, Daniel Hernández-Lobato, Alberto Suárez 0001
IJCAI2
2013 Learning Feature Selection Dependencies in Multi-task Learning
abstract
A probabilistic model based on the horseshoe prior is proposed for learning dependencies in the process of identifying relevant features for prediction. Exact inference is intractable in this model. However, expectation propagation offers an approximate alternative. Because the process of estimating feature selection dependencies may suffer from over-fitting in the model proposed, additional data from a multi-task learning scenario are considered for induction. The same model can be used in this setting with few modifications. Furthermore, the assumptions made are less restrictive than in other multi-task methods: The different tasks must share feature selection dependencies, but can have different relevant features and model coefficients. Experiments with real and synthetic data show that this model performs better than other multi-task alternatives from the literature. The experiments also show that the model is able to induce suitable feature selection dependencies for the problems considered, only from the training data.
Daniel Hernández-Lobato, José Miguel Hernández-Lobato
NIPS1
2013 Gaussian Process Conditional Copulas with Applications to Financial Time Series
abstract
The estimation of dependencies between multiple variables is a central problem in the analysis of financial time series. A common approach is to express these dependencies in terms of a copula function. Typically the copula function is assumed to be constant but this may be innacurate when there are covariates that could have a large influence on the dependence structure of the data. To account for this, a Bayesian framework for the estimation of conditional copulas is proposed. In this framework the parameters of a copula are non-linearly related to some arbitrary conditioning variables. We evaluate the ability of our method to predict time-varying dependencies on several equities and currencies and observe consistent performance gains compared to static copula models and other time-varying copula methods.
José Miguel Hernández-Lobato, James Robert Lloyd, Daniel Hernández-Lobato
NIPS3
2013 Generalized spike-and-slab priors for Bayesian group feature selection using expectation propagation
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Pierre Dupont
J. Mach. Learn. Res.1
2013 How large should ensembles of classifiers be?
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
Pattern Recognit.1
2012 On the Independence of the Individual Predictions in Parallel Randomized Ensembles
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
ESANN1
2011 Robust Multi-Class Gaussian Process Classification
abstract
Multi-class Gaussian Process Classifiers (MGPCs) are often affected by over-fitting problems when labeling errors occur far from the decision boundaries. To prevent this, we investigate a robust MGPC (RMGPC) which considers labeling errors independently of their distance to the decision boundaries. Expectation propagation is used for approximate inference. Experiments with several datasets in which noise is injected in the class labels illustrate the benefits of RMGPC. This method performs better than other Gaussian process alternatives based on considering latent Gaussian noise or heavy-tailed processes. When no noise is injected in the labels, RMGPC still performs equal or better than the other methods. Finally, we show how RMGPC can be used for successfully identifying data instances which are difficult to classify accurately in practice.
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Pierre Dupont
NIPS1
2011 Empirical analysis and evaluation of approximate techniques for pruning regression bagging ensembles
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
Neurocomputing1
2011 Network-based sparse Bayesian classification
José Miguel Hernández-Lobato, Daniel Hernández-Lobato, Alberto Suárez 0001
Pattern Recognit.2
2011 Inference on the prediction of ensembles of infinite size
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
Pattern Recognit.1
2010 Expectation Propagation for Bayesian Multi-task Feature Selection
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Thibault Helleputte, Pierre Dupont
ECML/PKDD (1)1
2010 Expectation Propagation for microarray data classification
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Alberto Suárez 0001
Pattern Recognit. Lett.1
2009 Statistical Instance-Based Ensemble Pruning for Multi-class Problems
Gonzalo Martínez-Muñoz, Daniel Hernández-Lobato, Alberto Suárez 0001
ICANN (1)2
2009 Statistical Instance-Based Pruning in Ensembles of Independent Classifiers
abstract
The global prediction of a homogeneous ensemble of classifiers generated in independent applications of a randomized learning algorithm on a fixed training set is analyzed within a Bayesian framework. Assuming that majority voting is used, it is possible to estimate with a given confidence level the prediction of the complete ensemble by querying only a subset of classifiers. For a particular instance that needs to be classified, the polling of ensemble classifiers can be halted when the probability that the predicted class will not change when taking into account the remaining votes is above the specified confidence level. Experiments on a collection of benchmark classification problems using representative parallel ensembles, such as bagging and random forests, confirm the validity of the analysis and demonstrate the effectiveness of the instance-based ensemble pruning method proposed.
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 An Analysis of Ensemble Pruning Techniques Based on Ordered Aggregation
abstract
Several pruning strategies that can be used to reduce the size and increase the accuracy of bagging ensembles are analyzed. These heuristics select subsets of complementary classifiers that, when combined, can perform better than the whole ensemble. The pruning methods investigated are based on modifying the order of aggregation of classifiers in the ensemble. In the original bagging algorithm, the order of aggregation is left unspecified. When this order is random, the generalization error typically decreases as the number of classifiers in the ensemble increases. If an appropriate ordering for the aggregation process is devised, the generalization error reaches a minimum at intermediate numbers of classifiers. This minimum lies below the asymptotic error of bagging. Pruned ensembles are obtained by retaining a fraction of the classifiers in the ordered ensemble. The performance of these pruned ensembles is evaluated in several benchmark classification tasks under different training conditions. The results of this empirical investigation show that ordered aggregation can be used for the efficient generation of pruned ensembles that are competitive, in terms of performance and robustness of classification, with computationally more costly methods that directly select optimal or near-optimal subensembles.
Gonzalo Martínez-Muñoz, Daniel Hernández-Lobato, Alberto Suárez 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Sparse Bayes Machines for Binary Classification
Daniel Hernández-Lobato
ICANN (1)1
2008 Class-switching neural network ensembles
Gonzalo Martínez-Muñoz, Aitor Sánchez-Martínez, Daniel Hernández-Lobato, Alberto Suárez 0001
Neurocomputing3
2008 Bayes Machines for binary classification
Daniel Hernández-Lobato, José Miguel Hernández-Lobato
Pattern Recognit. Lett.1
2007 GARCH Processes with Non-parametric Innovations for Market Risk Estimation
José Miguel Hernández-Lobato, Daniel Hernández-Lobato, Alberto Suárez 0001
ICANN (2)2
2007 Selection of Decision Stumps in Bagging Ensembles
Gonzalo Martínez-Muñoz, Daniel Hernández-Lobato, Alberto Suárez 0001
ICANN (1)2
2007 Out of Bootstrap Estimation of Generalization Error Curves in Bagging Ensembles
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
IDEAL1
2006 Building Ensembles of Neural Networks with Class-Switching
Gonzalo Martínez-Muñoz, Aitor Sánchez-Martínez, Daniel Hernández-Lobato, Alberto Suárez 0001
ICANN (1)3
2006 Pruning Adaptive Boosting Ensembles by Means of a Genetic Algorithm
Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Rubén Ruiz-Torrubiano, Angel Valle
IDEAL1
2006 Pruning in Ordered Regression Bagging Ensembles
abstract
An efficient procedure for pruning regression ensembles is introduced. Starting from a bagging ensemble, pruning proceeds by ordering the regressors in the original ensemble and then selecting a subset for aggregation. Ensembles of increasing size are built by including first the regressors that perform best when aggregated. This strategy gives an approximate solution to the problem of extracting from the original ensemble the minimum error subensemble, which we prove to be NP-hard. Experiments show that pruned ensembles with only 20% of the initial regressors achieve better generalization accuracies than the complete bagging ensembles. The performance of pruned ensembles is analyzed by means of the bias-variance decomposition of the error.
Daniel Hernández-Lobato, Gonzalo Martínez-Muñoz, Alberto Suárez 0001
IJCNN1