Michalis K. Titsias

dblp:19/5385 · DBLP profile ↗
← Back
51ranked-venue papers
20as first author
16since 2021 · last 2026
0000-0002-7603-0105ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 20 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Correction to: Personalized federated learning with exact stochastic gradient descent
Sotirios Nikoloutsopoulos, Iordanis Koutsopoulos, Michalis K. Titsias
Appl. Intell.3
2025 Learning-Order Autoregressive Models with Application to Molecular Graph Generation
abstract
Autoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction. While there appears to be a natural ordering for text (i.e., left-to-right), for many data types, such as graphs, the canonical ordering is less obvious. To address this problem, we introduce a variant of ARM that generates high-dimensional data using a probabilistic ordering that is sequentially inferred from data. This model incorporates a trainable probability distribution, referred to as an order-policy, that dynamically decides the autoregressive order in a state-dependent manner. To train the model, we introduce a variational lower bound on the exact log-likelihood, which we optimize with stochastic gradient estimation. We demonstrate experimentally that our method can learn meaningful autoregressive orderings in image and graph generation. On the challenging domain of molecular graph generation, we achieve state-of-the-art results on the QM9 and ZINC250k benchmarks, evaluated using the Fréchet ChemNet Distance (FCD), Synthetic Accessibility Score (SAS), Quantitative Estimate of Drug-likeness (QED).
Zhe Wang 0055, Jiaxin Shi, Nicolas Heess, Arthur Gretton, Michalis K. Titsias
ICML5
2025 New Bounds for Sparse Variational Gaussian Processes
abstract
Sparse variational Gaussian processes (GPs) construct tractable posterior approximations to GP models. At the core of these methods is the assumption that the true posterior distribution over training function values ${\bf f}$ and inducing variables ${\bf u}$ is approximated by a variational distribution that incorporates the conditional GP prior $p({\bf f} | {\bf u})$ in its factorization. While this assumption is considered as fundamental, we show that for model training we can relax it through the use of a more general variational distribution $q({\bf f} | {\bf u} )$ that depends on $N$ extra parameters, where $N$ is the number of training examples. In GP regression, we can analytically optimize the evidence lower bound over the extra parameters and express a tractable collapsed bound that is tighter than the previous bound. The new bound is also amenable to stochastic optimization and its implementation requires minor modifications to existing sparse GP code. Further, we also describe extensions to non-Gaussian likelihoods. On several datasets we demonstrate that our method can reduce bias when learning the hyperparameters and can lead to better predictive performance.
Michalis K. Titsias
ICML1
2025 Sparse Gaussian Processes: Structured Approximations and Power-EP Revisited
abstract
Inducing-point-based sparse variational Gaussian processes have become the standard workhorse for scaling up GP models. Recent advances show that these methods can be improved by introducing a diagonal scaling matrix to the conditional posterior density given the inducing points. This paper first considers an extension that employs a block-diagonal structure for the scaling matrix, provably tightening the variational lower bound. We then revisit the unifying framework of sparse GPs based on Power Expectation Propagation (PEP) and show that it can leverage and benefit from the new structured approximate posteriors. Through extensive regression experiments, we show that the proposed block-diagonal approximation consistently performs similarly to or better than existing diagonal approximations while maintaining comparable computational costs. Furthermore, the new PEP framework with structured posteriors provides competitive performance across various power hyperparameter settings, offering practitioners flexible alternatives to standard variational approaches.
Thang D. Bui, Michalis K. Titsias
NeurIPS2
2025 Personalized federated learning with exact stochastic gradient descent
Sotirios Nikoloutsopoulos, Iordanis Koutsopoulos, Michalis K. Titsias
Appl. Intell.3
2024 Kalman Filter for Online Classification of Non-Stationary Data
abstract
In Online Continual Learning (OCL) a learning system receives a stream of data and sequentially performs prediction and training steps. Key challenges in OCL include automatic adaptation to the specific non-stationary structure of the data and maintaining appropriate predictive uncertainty. To address these challenges we introduce a probabilistic Bayesian online learning approach that utilizes a (possibly pretrained) neural representation and a state space model over the linear predictor weights. Non-stationarity in the linear predictor weights is modelled using a “parameter drift” transition density, parametrized by a coefficient that quantifies forgetting. Inference in the model is implemented with efficient Kalman filter recursions which track the posterior distribution over the linear weights, while online SGD updates over the transition dynamics coefficient allow for adaptation to the non-stationarity observed in the data. While the framework is developed assuming a linear Gaussian model, we extend it to deal with classification problems and for fine-tuning the deep learning representation. In a set of experiments in multi-class classification using data sets such as CIFAR-100 and CLOC we demonstrate the model's predictive ability and its flexibility in capturing non-stationarity.
Michalis K. Titsias, Alexandre Galashov, Amal Rannen Triki, Razvan Pascanu, Yee Whye Teh, Jörg Bornschein
ICLR1
2024 Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
abstract
Neural networks are most often trained under the assumption that data come from a stationary distribution. However, settings in which this assumption is violated are of increasing importance; examples include supervised learning with distributional shifts, reinforcement learning, continual learning and non-stationary contextual bandits. Here, we introduce a novel learning approach that automatically models and adapts to non-stationarity by linking parameters through an Ornstein-Uhlenbeck process with an adaptive drift parameter. The adaptive drift draws the parameters towards the distribution used at initialisation, so the approach can be understood as a form of soft parameter reset. We show empirically that our approach performs well in non-stationary supervised, and off-policy reinforcement learning settings.
Alexandre Galashov, Michalis K. Titsias, András György 0001, Clare Lyle, Razvan Pascanu, Yee Whye Teh, Maneesh Sahani
NeurIPS2
2024 Simplified and Generalized Masked Diffusion for Discrete Data
abstract
Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and unclear relationships between different perspectives, leading to suboptimal parameterization, training objectives, and ad hoc adjustments to counteract these issues. In this work, we aim to provide a simple and general framework that unlocks the full potential of masked diffusion models. We show that the continuous-time variational objective of masked diffusion models is a simple weighted integral of cross-entropy losses. Our framework also enables training generalized masked diffusion models with state-dependent masking schedules. When evaluated by perplexity, our models trained on OpenWebText surpass prior diffusion language models at GPT-2 scale and demonstrate superior performance on 4 out of 5 zero-shot language modeling tasks. Furthermore, our models vastly outperform previous discrete diffusion models on pixel-level image modeling, achieving 2.75 (CIFAR-10) and 3.40 (ImageNet 64x64) bits per dimension that are better than autoregressive models of similar sizes.
Jiaxin Shi, Kehang Han, Arnaud Doucet, Michalis K. Titsias
NeurIPS5
2023 Optimal Preconditioning and Fisher Adaptive Langevin Sampling
abstract
We define an optimal preconditioning for the Langevin diffusion by analytically optimizing the expected squared jumped distance. This yields as the optimal preconditioning an inverse Fisher information covariance matrix, where the covariance matrix is computed as the outer product of log target gradients averaged under the target. We apply this result to the Metropolis adjusted Langevin algorithm (MALA) and derive a computationally efficient adaptive MCMC scheme that learns the preconditioning from the history of gradients produced as the algorithm runs. We show in several experiments that the proposed algorithm is very robust in high dimensions and significantly outperforms other methods, including a closely related adaptive MALA scheme that learns the preconditioning with standard adaptive MCMC as well as the position-dependent Riemannian manifold MALA sampler.
Michalis K. Titsias
NeurIPS1
2022 Double Control Variates for Gradient Estimation in Discrete Latent Variable Models
abstract
Stochastic gradient-based optimisation for discrete latent variable models is challenging due to the high variance of gradients. We introduce a variance reduction technique for score function estimators that makes use of double control variates. These control variates act on top of a main control variate, and try to further reduce the variance of the overall estimator. We develop a double control variate for the REINFORCE leave-one-out estimator using Taylor expansions. For training discrete latent variable models, such as variational autoencoders with binary latent variables, our approach adds no extra computational cost compared to standard training with the REINFORCE leave-one-out estimator. We apply our method to challenging high-dimensional toy examples and for training variational autoencoders with binary latent variables. We show that our estimator can have lower variance compared to other state-of-the-art estimators.
Michalis K. Titsias, Jiaxin Shi
AISTATS1
2022 Information-theoretic Online Memory Selection for Continual Learning
Shengyang Sun, Daniele Calandriello, Huiyi Hu, Michalis K. Titsias
ICLR5
2022 Gradient Estimation with Discrete Stein Operators
abstract
Gradient estimation---approximating the gradient of an expectation with respect to the parameters of a distribution---is central to the solution of many machine learning problems. However, when the distribution is discrete, most common gradient estimators suffer from excessive variance. To improve the quality of gradient estimation, we introduce a variance reduction technique based on Stein operators for discrete distributions. We then use this technique to build flexible control variates for the REINFORCE leave-one-out estimator. Our control variates can be adapted online to minimize variance and do not require extra evaluations of the target function. In benchmark generative modeling tasks such as training binary variational autoencoders, our gradient estimator achieves substantially lower variance than state-of-the-art estimators with the same number of function evaluations.
Jiaxin Shi, Jessica Hwang, Michalis K. Titsias, Lester Mackey
NeurIPS4
2021 Entropy-based adaptive Hamiltonian Monte Carlo
abstract
Hamiltonian Monte Carlo (HMC) is a popular Markov Chain Monte Carlo (MCMC) algorithm to sample from an unnormalized probability distribution. A leapfrog integrator is commonly used to implement HMC in practice, but its performance can be sensitive to the choice of mass matrix used therein. We develop a gradient-based algorithm that allows for the adaptation of the mass matrix by encouraging the leapfrog integrator to have high acceptance rates while also exploring all dimensions jointly. In contrast to previous work that adapt the hyperparameters of HMC using some form of expected squared jumping distance, the adaptation strategy suggested here aims to increase sampling efficiency by maximizing an approximation of the proposal entropy. We illustrate that using multiple gradients in the HMC proposal can be beneficial compared to a single gradient-step in Metropolis-adjusted Langevin proposals. Empirical evidence suggests that the adaptation method can outperform different versions of HMC schemes by adjusting the mass matrix to the geometry of the target distribution and by providing some control on the integration time.
Marcel Hirt, Michalis K. Titsias, Petros Dellaportas
NeurIPS2
2021 Unbiased gradient estimation for variational auto-encoders using coupled Markov chains
abstract
The variational auto-encoder (VAE) is a deep latent variable model that has two neural networks in an autoencoder-like architecture; one of them parameterizes the model’s likelihood. Fitting its parameters via maximum likelihood (ML) is challenging since the computation of the marginal likelihood involves an intractable integral over the latent space; thus the VAE is trained instead by maximizing a variational lower bound. Here, we develop a ML training scheme for VAEs by introducing unbiased estimators of the log-likelihood gradient. We obtain the estimators by augmenting the latent space with a set of importance samples, similarly to the importance weighted auto-encoder (IWAE), and then constructing a Markov chain Monte Carlo coupling procedure on this augmented space. We provide the conditions under which the estimators can be computed in finite time and with finite variance. We show experimentally that VAEs fitted with unbiased estimators exhibit better predictive performance.
Francisco J. R. Ruiz, Michalis K. Titsias, A. Taylan Cemgil, Arnaud Doucet
UAI2
2021 Information theoretic meta learning with Gaussian processes
abstract
We formulate meta learning using information theoretic concepts; namely, mutual information and the information bottleneck. The idea is to learn a stochastic representation or encoding of the task description, given by a training set, that is highly informative about predicting the validation set. By making use of variational approximations to the mutual information, we derive a general and tractable framework for meta learning. This framework unifies existing gradient-based algorithms and also allows us to derive new algorithms. In particular, we develop a memory-based algorithm that uses Gaussian processes to obtain non-parametric encoding representations. We demonstrate our method on a few-shot regression problem and on four few-shot classification problems, obtaining competitive accuracy when compared to existing baselines.
Michalis K. Titsias, Francisco J. R. Ruiz, Sotirios Nikoloutsopoulos, Alexandre Galashov
UAI1
2021 Large scale multi-label learning using Gaussian processes
abstract
Abstract We introduce a Gaussian process latent factor model for multi-label classification that can capture correlations among class labels by using a small set of latent Gaussian process functions. To address computational challenges, when the number of training instances is very large, we introduce several techniques based on variational sparse Gaussian process approximations and stochastic optimization. Specifically, we apply doubly stochastic variational inference that sub-samples data instances and classes which allows us to cope with Big Data. Furthermore, we show it is possible and beneficial to optimize over inducing points, using gradient-based methods, even in very high dimensional input spaces involving up to hundreds of thousands of dimensions. We demonstrate the usefulness of our approach on several real-world large-scale multi-label learning problems.
Aristeidis Panos, Petros Dellaportas, Michalis K. Titsias
Mach. Learn.3
2020 Sparse Orthogonal Variational Inference for Gaussian Processes
abstract
We introduce a new interpretation of sparse variational approximations for Gaussian processes using inducing points, which can lead to more scalable algorithms than previous methods. It is based on decomposing a Gaussian process as a sum of two independent processes: one spanned by a finite basis of inducing points and the other capturing the remaining variation. We show that this formulation recovers existing approximations and at the same time allows to obtain tighter lower bounds on the marginal likelihood and new stochastic variational inference algorithms. We demonstrate the efficiency of these algorithms in several Gaussian process models ranging from standard regression to multi-class classification using (deep) convolutional Gaussian processes and report state-of-the-art results on CIFAR-10 among purely GP-based models.
Jiaxin Shi, Michalis K. Titsias, Andriy Mnih
AISTATS2
2020 Functional Regularisation for Continual Learning with Gaussian Processes
Michalis K. Titsias, Jonathan Schwarz, Alexander G. de G. Matthews, Razvan Pascanu, Yee Whye Teh
ICLR1
2019 Augmented Ensemble MCMC sampling in Factorial Hidden Markov Models
abstract
Bayesian inference for Factorial Hidden Markov Models is challenging due to the exponentially sized latent variable space. Standard Monte Carlo samplers can have difficulties effectively exploring the posterior landscape and are often restricted to exploration around localised regions that depend on initialisation. We introduce a general purpose ensemble Markov Chain Monte Carlo (MCMC) technique to improve on existing poorly mixing samplers. This is achieved by combining parallel tempering and an auxiliary variable scheme to exchange information between the chains in an efficient way. The latter exploits a genetic algorithm within an augmented Gibbs sampler. We compare our technique with various existing samplers in a simulation study as well as in a cancer genomics application, demonstrating the improvements obtained by our augmented ensemble approach.
Kaspar Märtens, Michalis K. Titsias, Christopher Yau
AISTATS2
2019 Unbiased Implicit Variational Inference
abstract
We develop unbiased implicit variational inference (UIVI), a method that expands the applicability of variational inference by defining an expressive variational family. UIVI considers an implicit variational distribution obtained in a hierarchical manner using a simple reparameterizable distribution whose variational parameters are defined by arbitrarily flexible deep neural networks. Unlike previous works, UIVI directly optimizes the evidence lower bound (ELBO) rather than an approximation to the ELBO. We demonstrate UIVI on several models, including Bayesian multinomial logistic regression and variational autoencoders, and show that UIVI achieves both tighter ELBO and better predictive performance than existing approaches at a similar computational cost.
Michalis K. Titsias, Francisco J. R. Ruiz
AISTATS1
2019 A Contrastive Divergence for Combining Variational Inference and MCMC
abstract
We develop a method to combine Markov chain Monte Carlo (MCMC) and variational inference (VI), leveraging the advantages of both inference approaches. Specifically, we improve the variational distribution by running a few MCMC steps. To make inference tractable, we introduce the variational contrastive divergence (VCD), a new divergence that replaces the standard Kullback-Leibler (KL) divergence used in VI. The VCD captures a notion of discrepancy between the initial variational distribution and its improved version (obtained after running the MCMC steps), and it converges asymptotically to the symmetrized KL divergence between the variational distribution and the posterior of interest. The VCD objective can be optimized efficiently with respect to the variational parameters via stochastic optimization. We show experimentally that optimizing the VCD leads to better predictive performance on two latent variable models: logistic matrix factorization and variational autoencoders (VAEs).
Francisco J. R. Ruiz, Michalis K. Titsias
ICML2
2019 Gradient-based Adaptive Markov Chain Monte Carlo
abstract
We introduce a gradient-based learning method to automatically adapt Markov chain Monte Carlo (MCMC) proposal distributions to intractable targets. We define a maximum entropy regularised objective function, referred to as generalised speed measure, which can be robustly optimised over the parameters of the proposal distribution by applying stochastic gradient optimisation. An advantage of our method compared to traditional adaptive MCMC methods is that the adaptation occurs even when candidate state values are rejected. This is a highly desirable property of any adaptation strategy because the adaptation starts in early iterations even if the initial proposal distribution is far from optimum. We apply the framework for learning multivariate random walk Metropolis and Metropolis-adjusted Langevin proposals with full covariance matrices, and provide empirical evidence that our method can outperform other MCMC algorithms, including Hamiltonian Monte Carlo schemes.
Michalis K. Titsias, Petros Dellaportas
NeurIPS1
2018 Augment and Reduce: Stochastic Inference for Large Categorical Distributions
abstract
Categorical distributions are ubiquitous in machine learning, e.g., in classification, language models, and recommendation systems. However, when the number of possible outcomes is very large, using categorical distributions becomes computationally expensive, as the complexity scales linearly with the number of outcomes. To address this problem, we propose augment and reduce (A&R), a method to alleviate the computational complexity. A&R uses two ideas: latent variable augmentation and stochastic variational inference. It maximizes a lower bound on the marginal likelihood of the data. Unlike existing methods which are specific to softmax, A&R is more general and is amenable to other categorical models, such as multinomial probit. On several large-scale classification problems, we show that A&R provides a tighter bound on the marginal likelihood and has better predictive performance than existing approaches.
Francisco J. R. Ruiz, Michalis K. Titsias, Adji B. Dieng, David M. Blei
ICML2
2017 Bayesian Boolean Matrix Factorisation
abstract
Boolean matrix factorisation aims to decompose a binary data matrix into an approximate Boolean product of two low rank, binary matrices: one containing meaningful patterns, the other quantifying how the observations can be expressed as a combination of these patterns. We introduce the OrMachine, a probabilistic generative model for Boolean matrix factorisation and derive a Metropolised Gibbs sampler that facilitates efficient parallel posterior inference. On real world and simulated data, our method outperforms all currently existing approaches for Boolean matrix factorisation and completion. This is the first method to provide full posterior inference for Boolean Matrix factorisation which is relevant in applications, e.g. for controlling false positive rates in collaborative filtering and, crucially, improves the interpretability of the inferred patterns. The proposed algorithm scales to large datasets as we demonstrate by analysing single cell gene expression data in 1.3 million mouse brain cells across 11 thousand genes on commodity hardware.
Tammo Rukat, Christopher C. Holmes, Michalis K. Titsias, Christopher Yau
ICML3
2016 First learn then earn: optimizing mobile crowdsensing campaigns through data-driven user profiling
abstract
We study the optimal design of mobile crowdsensing campaigns in terms of the aggregate quality of contributions attracted for a set of tasks. The interaction of the campaign with users is realized through a mobile app interface that recommends tasks to users and offers them incentives. The main contribution is a novel perspective on the payment distribution problem faced by the crowdsensing campaign organizer in light of originally unknown individual user preferences. Contrary to common practice, we acknowledge that users exhibit high diversity in decision making because they assess differently attributes related to a task such as their proximity to the place of interest (PoI), the payment made for contributing data, or the task context/theme. We draw on logistic-regression techniques from machine learning to learn users' individual preferences from past data rather than hypothesizing about them. We then formulate non-linear (sigmoid) optimization problems to determine the tasks and incentives (payments) that should be optimally offered to each user. Our mechanism is validated against synthetic but also real data about the way users choose tasks, collected through an online questionnaire. It achieves very good approximations of the optimal solutions and substantially outperforms alternative preference-agnostic policies that do not exercise behavioral user profiling to target the provision of incentives.
Merkourios Karaliopoulos, Iordanis Koutsopoulos, Michalis K. Titsias
MobiHoc3
2016 The Generalized Reparameterization Gradient
abstract
The reparameterization gradient has become a widely used method to obtain Monte Carlo gradients to optimize the variational objective. However, this technique does not easily apply to commonly used distributions such as beta or gamma without further approximations, and most practical applications of the reparameterization gradient fit Gaussian distributions. In this paper, we introduce the generalized reparameterization gradient, a method that extends the reparameterization gradient to a wider class of variational distributions. Generalized reparameterizations use invertible transformations of the latent variables which lead to transformed distributions that weakly depend on the variational parameters. This results in new Monte Carlo gradients that combine reparameterization gradients and score function gradients. We demonstrate our approach on variational inference for two complex probabilistic models. The generalized reparameterization is effective: even a single sample from the variational distribution is enough to obtain a low-variance gradient.
Francisco J. R. Ruiz, Michalis K. Titsias, David M. Blei
NIPS2
2016 One-vs-Each Approximation to Softmax for Scalable Estimation of Probabilities
abstract
The softmax representation of probabilities for categorical variables plays a prominent role in modern machine learning with numerous applications in areas such as large scale classification, neural language modeling and recommendation systems. However, softmax estimation is very expensive for large scale inference because of the high cost associated with computing the normalizing constant. Here, we introduce an efficient approximation to softmax probabilities which takes the form of a rigorous lower bound on the exact probability. This bound is expressed as a product over pairwise probabilities and it leads to scalable estimation based on stochastic optimization. It allows us to perform doubly stochastic estimation by subsampling both training instances and class labels. We show that the new bound has interesting theoretical properties and we demonstrate its use in classification problems.
Michalis K. Titsias
NIPS1
2016 Overdispersed Black-Box Variational Inference
Francisco J. R. Ruiz, Michalis K. Titsias, David M. Blei
UAI2
2016 Variational Inference for Latent Variables and Uncertain Inputs in Gaussian Processes
abstract
The Gaussian process latent variable model (GP-LVM) provides a flexible approach for non-linear dimensionality reduction that has been widely applied. However, the current approach for training GP-LVMs is based on maximum likelihood, where the latent projection variables are maximised over rather than integrated out. In this paper we present a Bayesian method for training GP-LVMs by introducing a non-standard variational inference framework that allows to approximately integrate out the latent variables and subsequently train a GP-LVM by maximising an analytic lower bound on the exact marginal likelihood. We apply this method for learning a GP-LVM from i.i.d. observations and for learning non-linear dynamical systems where the observations are temporally correlated. We show that a benefit of the variational Bayesian procedure is its robustness to overfitting and its ability to automatically select the dimensionality of the non-linear latent space. The resulting framework is generic, flexible and easy to extend for other purposes, such as Gaussian process regression with uncertain or partially missing inputs. We demonstrate our method on synthetic data and standard machine learning benchmarks, as well as challenging real world datasets, including high resolution video data.
Andreas Damianou, Michalis K. Titsias, Neil D. Lawrence
J. Mach. Learn. Res.2
2015 Inference for determinantal point processes without spectral knowledge
abstract
Determinantal point processes (DPPs) are point process models thatnaturally encode diversity between the points of agiven realization, through a positive definite kernel $K$. DPPs possess desirable properties, such as exactsampling or analyticity of the moments, but learning the parameters ofkernel $K$ through likelihood-based inference is notstraightforward. First, the kernel that appears in thelikelihood is not $K$, but another kernel $L$ related to $K$ throughan often intractable spectral decomposition. This issue is typically bypassed in machine learning bydirectly parametrizing the kernel $L$, at the price of someinterpretability of the model parameters. We follow this approachhere. Second, the likelihood has an intractable normalizingconstant, which takes the form of large determinant in the case of aDPP over a finite set of objects, and the form of a Fredholm determinant in thecase of a DPP over a continuous domain. Our main contribution is to derive bounds on the likelihood ofa DPP, both for finite and continuous domains. Unlike previous work, our bounds arecheap to evaluate since they do not rely on approximating the spectrumof a large matrix or an operator. Through usual arguments, these bounds thus yield cheap variationalinference and moderately expensive exact Markov chain Monte Carlo inference methods for DPPs.
Rémi Bardenet, Michalis K. Titsias
NIPS2
2015 Local Expectation Gradients for Black Box Variational Inference
abstract
We introduce local expectation gradients which is a general purpose stochastic variational inference algorithm for constructing stochastic gradients by sampling from the variational distribution. This algorithm divides the problem of estimating the stochastic gradients over multiple variational parameters into smaller sub-tasks so that each sub-task explores intelligently the most relevant part of the variational distribution. This is achieved by performing an exact expectation over the single random variable that most correlates with the variational parameter of interest resulting in a Rao-Blackwellized estimate that has low variance. Our method works efficiently for both continuous and discrete random variables. Furthermore, the proposed algorithm has interesting similarities with Gibbs sampling but at the same time, unlike Gibbs sampling, can be trivially parallelized.
Michalis K. Titsias, Miguel Lázaro-Gredilla
NIPS1
2014 Doubly Stochastic Variational Bayes for non-Conjugate Inference
abstract
We propose a simple and effective variational inference algorithm based on stochastic optimisation that can be widely applied for Bayesian non-conjugate inference in continuous parameter spaces. This algorithm is based on stochastic approximation and allows for efficient use of gradient information from the model joint density. We demonstrate these properties using illustrative examples as well as in challenging and diverse Bayesian inference problems such as variable selection in logistic regression and fully Bayesian inference over kernel hyperparameters in Gaussian process regression.
Michalis K. Titsias, Miguel Lázaro-Gredilla
ICML1
2014 Hamming Ball Auxiliary Sampling for Factorial Hidden Markov Models
Michalis K. Titsias, Christopher Yau
NIPS1
2014 Retrieval of Biophysical Parameters With Heteroscedastic Gaussian Processes
abstract
An accurate estimation of biophysical variables is the key to monitor our Planet. Leaf chlorophyll content helps in interpreting the chlorophyll fluorescence signal from space, whereas oceanic chlorophyll concentration allows us to quantify the healthiness of the oceans. Recently, the family of Bayesian nonparametric methods has provided excellent results in these situations. A particularly useful method in this framework is the Gaussian process regression (GPR). However, standard GPR assumes that the variance of the noise process is independent of the signal, which does not hold in most of the problems. In this letter, we propose a nonstandard variational approximation that allows accurate inference in signal-dependent noise scenarios. We show that the so-called variational heteroscedastic GPR (VHGPR) is an excellent alternative to standard GPR in two relevant Earth observation examples, namely, Chl vegetation retrieval from hyperspectral images and oceanic Chl concentration estimation from in situ measured reflectances. The proposed VHGPR outperforms the tested empirical approaches, as well as statistical linear regression (both least squares and least absolute shrinkage and selection operator), neural nets, and kernel ridge regression, and the homoscedastic GPR, in terms of accuracy and bias, and proves more robust when a low number of examples is available.
Miguel Lázaro-Gredilla, Michalis K. Titsias, Jochem Verrelst, Gustau Camps-Valls
IEEE Geosci. Remote. Sens. Lett.2
2013 Estimation of vegetation chlorophyll content with Variational Heteroscedastic Gaussian Processes
abstract
Accurate estimation of biophysical variables is the key to monitor our Planet. In particular, leaf chlorophyll content helps in interpreting the chlorophyll fluorescence signal from space, which is an accurate indicator of the actual state of the vegetation beyond greenness. Recently, the family of Bayesian nonparametric methods has provided excellent results in these situations. A particularly useful method in this framework is the Gaussian Processes regression (GP). However, standard GP assumes that the variance of the noise process is independent of the signal, which does not hold in most of the problems. In this paper, we propose a non-standard variational approximation that allows accurate inference in signal-dependent noise scenarios. We show that the so-called Variational Heteroscedastic Gaussian Process (VHGP) regression is an excellent alternative to standard GP for the retrieval of vegetation chlorophyll content from hyperspectral images. In general VHGP outperforms GP (and many other empirical and machine learning techniques) in accuracy and bias, and reveals more robust when a low number of examples is available.
Miguel Lázaro-Gredilla, Michalis K. Titsias, Jochem Verrelst, Gustau Camps-Valls
IGARSS2
2013 Variational Inference for Mahalanobis Distance Metrics in Gaussian Process Regression
abstract
We introduce a novel variational method that allows to approximately integrate out kernel hyperparameters, such as length-scales, in Gaussian process regression. This approach consists of a novel variant of the variational framework that has been recently developed for the Gaussian process latent variable model which additionally makes use of a standardised representation of the Gaussian process. We consider this technique for learning Mahalanobis distance metrics in a Gaussian process regression setting and provide experimental evaluations and comparisons with existing methods by considering datasets with high-dimensional inputs.
Michalis K. Titsias, Miguel Lázaro-Gredilla
NIPS1
2012 Manifold Relevance Determination
Andreas Damianou, Carl Henrik Ek, Michalis K. Titsias, Neil D. Lawrence
ICML3
2011 Variational Heteroscedastic Gaussian Process Regression
Miguel Lázaro-Gredilla, Michalis K. Titsias
ICML2
2011 Variational Gaussian Process Dynamical Systems
abstract
High dimensional time series are endemic in applications of machine learning such as robotics (sensor data), computational biology (gene expression data), vision (video sequences) and graphics (motion capture data). Practical nonlinear probabilistic approaches to this data are required. In this paper we introduce the variational Gaussian process dynamical system. Our work builds on recent variational approximations for Gaussian process latent variable models to allow for nonlinear dimensionality reduction simultaneously with learning a dynamical prior in the latent space. The approach also allows for the appropriate dimensionality of the latent space to be automatically determined. We demonstrate the model on a human motion capture data set and a series of high resolution video sequences.
Andreas Damianou, Michalis K. Titsias, Neil D. Lawrence
NIPS2
2011 Spike and Slab Variational Inference for Multi-Task and Multiple Kernel Learning
abstract
We introduce a variational Bayesian inference algorithm which can be widely applied to sparse linear models. The algorithm is based on the spike and slab prior which, from a Bayesian perspective, is the golden standard for sparse inference. We apply the method to a general multi-task and multiple kernel learning model in which a common set of Gaussian process functions is linearly combined with task-specific sparse weights, thus inducing relation between tasks. This model unifies several sparse linear models, such as generalized linear models, sparse factor analysis and matrix factorization with missing values, so that the variational algorithm can be applied to all these cases. We demonstrate our approach in multi-output Gaussian process regression, multi-class classification, image processing applications and collaborative filtering.
Michalis K. Titsias, Miguel Lázaro-Gredilla
NIPS1
2009 Web Page Rank Prediction with PCA and EM Clustering
Polyxeni Zacharouli, Michalis K. Titsias, Michalis Vazirgiannis
WAW2
2008 Efficient Sampling for Gaussian Process Inference using Control Variables
abstract
Sampling functions in Gaussian process (GP) models is challenging because of the highly correlated posterior distribution. We describe an efficient Markov chain Monte Carlo algorithm for sampling from the posterior process of the GP model. This algorithm uses control variables which are auxiliary function values that provide a low dimensional representation of the function. At each iteration, the algorithm proposes new values for the control variables and generates the function from the conditional GP prior. The control variable input locations are found by continuously minimizing an objective function. We demonstrate the algorithm on regression and classification problems and we use it to estimate the parameters of a differential equation model of gene regulation.
Michalis K. Titsias, Neil D. Lawrence, Magnus Rattray
NIPS1
2007 The Infinite Gamma-Poisson Feature Model
abstract
We address the problem of factorial learning which associates a set of latent causes or features with the observed data. Factorial models usually assume that each feature has a single occurrence in a given data point. However, there are data such as images where latent features have multiple occurrences, e.g. a visual object class can have multiple instances shown in the same image. To deal with such cases, we present a probability model over non-negative integer valued matrices with possibly unbounded number of columns. This model can play the role of the prior in an nonparametric Bayesian learning scenario where both the latent features and the number of their occurrences are unknown. We use this prior together with a likelihood model for unsupervised learning from images using a Markov Chain Monte Carlo inference algorithm.
Michalis K. Titsias
NIPS1
2006 Bayesian Feature and Model Selection for Gaussian Mixture Models
abstract
We present a Bayesian method for mixture model training that simultaneously treats the feature selection and the model selection problem. The method is based on the integration of a mixture model formulation that takes into account the saliency of the features and a Bayesian approach to mixture learning that can be used to estimate the number of mixture components. The proposed learning algorithm follows the variational framework and can simultaneously optimize over the number of components, the saliency of the features, and the parameters of the mixture model. Experimental results using high-dimensional artificial and real data illustrate the effectiveness of the method.
Constantinos Constantinopoulos, Michalis K. Titsias, Aristidis Likas
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Fast Learning of Sprites using Invariant Features
abstract
A popular framework for the interpretation of image sequences is the layers or sprite model of e.g. Wang and Adelson (1994), Irani et al. (1994). Jojic and Frey (2001) provide a generative probabilistic model framework for this task, but their algorithm is slow as it needs to search over discretized transformations (e.g. translations, or affines) for each layer. In this paper we show that by using invariant features (e.g. Lowe’s SIFT features) and clustering their motions we can reduce or eliminate the search and thus learn the sprites much faster. We demonstrate our algorithm on two image sequences. 1
Moray Allan, Michalis K. Titsias, Christopher K. I. Williams
BMVC2
2004 Greedy Learning of Multiple Objects in Images Using Robust Statistics and Factorial Learning
abstract
We consider data that are images containing views of multiple objects. Our task is to learn about each of the objects present in the images. This task can be approached as a factorial learning problem, where each image must be explained by instantiating a model for each of the objects present with the correct instantiation parameters. A major problem with learning a factorial model is that as the number of objects increases, there is a combinatorial explosion of the number of configurations that need to be considered. We develop a method to extract object models sequentially from the data by making use of a robust statistical method, thus avoiding the combinatorial explosion, and present results showing successful extraction of objects from real images.
Christopher K. I. Williams, Michalis K. Titsias
Neural Comput.2
2003 Class Conditional Density Estimation Using Mixtures with Constrained Component Sharing
abstract
We propose a generative mixture model classifier that allows for the class conditional densities to be represented by mixtures having certain subsets of their components shared or common among classes. We argue that, when the total number of mixture components is kept fixed, the most efficient classification model is obtained by appropriately determining the sharing of components among class conditional densities. In order to discover such an efficient model, a training method is derived based on the EM algorithm that automatically adjusts component sharing. We provide experimental results with good classification performance.
Michalis K. Titsias, Aristidis Likas
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Learning About Multiple Objects in Images: Factorial Learning without Factorial Search
abstract
We consider data which are images containing views of multiple objects. Our task is to learn about each of the objects present in the images. This task can be approached as a factorial learning problem, where each image must be explained by instantiating a model for each of the objects present with the correct instantiation parameters. A major problem with learning a factorial model is that as the number of objects increases, there is a combinatorial explosion of the number of configurations that need to be considered. We develop a method to extract object models sequentially from the data by making use of a robust statistical method, thus avoid- ing the combinatorial explosion, and present results showing successful extraction of objects from real images.
Christopher K. I. Williams, Michalis K. Titsias
NIPS2
2002 Mixture of Experts Classification Using a Hierarchical Mixture Model
abstract
A three-level hierarchical mixture model for classification is presented that models the following data generation process: (1) the data are generated by a finite number of sources (clusters), and (2) the generation mechanism of each source assumes the existence of individual internal class-labeled sources (subclusters of the external cluster). The model estimates the posterior probability of class membership similar to a mixture of experts classifier. In order to learn the parameters of the model, we have developed a general training approach based on maximum likelihood that results in two efficient training algorithms. Compared to other classification mixture models, the proposed hierarchical model exhibits several advantages and provides improved classification performance as indicated by the experimental results.
Michalis K. Titsias, Aristidis Likas
Neural Comput.1
2001 Shared kernel models for class conditional density estimation
abstract
We present probabilistic models which are suitable for class conditional density estimation and can be regarded as shared kernel models where sharing means that each kernel may contribute to the estimation of the conditional densities of an classes. We first propose a model that constitutes an adaptation of the classical radial basis function (RBF) network (with full sharing of kernels among classes) where the outputs represent class conditional densities. In the opposite direction is the approach of separate mixtures model where the density of each class is estimated using a separate mixture density (no sharing of kernels among classes). We present a general model that allows for the expression of intermediate cases where the degree of kernel sharing can be specified through an extra model parameter. This general model encompasses both the above mentioned models as special cases. In all proposed models the training process is treated as a maximum likelihood problem and expectation-maximization algorithms have been derived for adjusting the model parameters.
Michalis K. Titsias, Aristidis Likas
IEEE Trans. Neural Networks1
2000 A Probabilistic RBF Network for Classification
abstract
We present a probabilistic neural network model which is suitable for classification problems. This model constitutes an adaptation of the classical RBF network where the outputs represent the class conditional distributions. Since the network outputs correspond to probability density functions, training process is treated as maximum likelihood problem and an expectation-maximization (EM) algorithm is proposed for adjusting the network parameters. Experimental results show that proposed architecture exhibits superior classification performance compared to the classical RBF network.
Michalis K. Titsias, Aristidis Likas
IJCNN (4)1