Thang D. Bui

dblp:155/1914 · also Thang Duc Bui · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
2since 2021 · last 2025
0000-0002-7878-9748ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Probabilistic and Bayesian machine learning · 83% Learning paradigms · 9% Time series and sequential data · 3%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
3.082025
Sparse Gaussian Processes: Structured Approximations and Power-EP Revisited · NeurIPS 2025
Variational Auto-Regressive Gaussian Processes for Continual Learning · ICML 2021
Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
1.432025
Sparse Gaussian Processes: Structured Approximations and Power-EP Revisited · NeurIPS 2025
A Unifying Framework for Gaussian Process Pseudo-Point Approximations using Power Expectation Propagation · J. Mach. Learn. Res. 2017
Deep Gaussian Processes for Regression using Approximate Expectation Propagation · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.132021
Variational Auto-Regressive Gaussian Processes for Continual Learning · ICML 2021
Variational Continual Learning · ICLR (Poster) 2018
Black-Box Alpha Divergence Minimization · ICML 2016
Machine learning › Learning paradigms
continual learning
0.822021
Variational Auto-Regressive Gaussian Processes for Continual Learning · ICML 2021
Variational Continual Learning · ICLR (Poster) 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.832017
A Unifying Framework for Gaussian Process Pseudo-Point Approximations using Power Expectation Propagation · J. Mach. Learn. Res. 2017
Black-Box Alpha Divergence Minimization · ICML 2016
Deep Gaussian Processes for Regression using Approximate Expectation Propagation · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.522020
Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights · NeurIPS 2020
Deep Gaussian Processes for Regression using Approximate Expectation Propagation · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
sparse approximation
0.522017
Streaming Sparse Gaussian Process Approximations · NIPS 2017
Tree-structured Gaussian Process Approximations · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
hierarchical gaussian process
0.412020
Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › prior modeling
weight prior
0.412020
Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational learning
variational continual learning
0.312018
Variational Continual Learning · ICLR (Poster) 2018
Machine learning › Time series and sequential data
streaming data
0.312017
Streaming Sparse Gaussian Process Approximations · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › hierarchical gaussian process
deep gaussian process
0.212016
Deep Gaussian Processes for Regression using Approximate Expectation Propagation · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning
divergence minimization
0.212016
Black-Box Alpha Divergence Minimization · ICML 2016
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
non-parametric kernel learning
0.212015
Learning Stationary Time Series using Gaussian Processes with Nonparametric Kernels · NIPS 2015
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
tree-structured approximation
0.212014
Tree-structured Gaussian Process Approximations · NIPS 2014
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.112021
Variational Auto-Regressive Gaussian Processes for Continual Learning · ICML 2021
Machine learning › Time series and sequential data
time series modeling
0.112015
Learning Stationary Time Series using Gaussian Processes with Nonparametric Kernels · NIPS 2015

Methods — techniques the papers use, named apart from their topics

variational inference · 1.3power expectation propagation · 1.2variational lower bound · 0.9block-diagonal approximation · 0.9stochastic gradient descent · 0.6sparse inducing point approximation · 0.5expectation propagation · 0.5unit embeddings · 0.4kernel design · 0.4label propagation · 0.3
YearPublicationVenuePosition
2025 Sparse Gaussian Processes: Structured Approximations and Power-EP Revisited
abstract
Inducing-point-based sparse variational Gaussian processes have become the standard workhorse for scaling up GP models. Recent advances show that these methods can be improved by introducing a diagonal scaling matrix to the conditional posterior density given the inducing points. This paper first considers an extension that employs a block-diagonal structure for the scaling matrix, provably tightening the variational lower bound. We then revisit the unifying framework of sparse GPs based on Power Expectation Propagation (PEP) and show that it can leverage and benefit from the new structured approximate posteriors. Through extensive regression experiments, we show that the proposed block-diagonal approximation consistently performs similarly to or better than existing diagonal approximations while maintaining comparable computational costs. Furthermore, the new PEP framework with structured posteriors provides competitive performance across various power hyperparameter settings, offering practitioners flexible alternatives to standard variational approaches.
Thang D. Bui, Michalis K. Titsias
NeurIPS1
2021 Variational Auto-Regressive Gaussian Processes for Continual Learning
abstract
Through sequential construction of posteriors on observing data online, Bayes’ theorem provides a natural framework for continual learning. We develop Variational Auto-Regressive Gaussian Processes (VAR-GPs), a principled posterior updating mechanism to solve sequential tasks in continual learning. By relying on sparse inducing point approximations for scalable posteriors, we propose a novel auto-regressive variational distribution which reveals two fruitful connections to existing results in Bayesian inference, expectation propagation and orthogonal inducing points. Mean predictive entropy estimates show VAR-GPs prevent catastrophic forgetting, which is empirically supported by strong performance on modern continual learning benchmarks against competitive baselines. A thorough ablation study demonstrates the efficacy of our modeling choices.
Sanyam Kapoor, Theofanis Karaletsos, Thang D. Bui
ICML3
2020 Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights
abstract
Probabilistic neural networks are typically modeled with independent weight priors, which do not capture weight correlations in the prior and do not provide a parsimonious interface to express properties in function space. A desirable class of priors would represent weights compactly, capture correlations between weights, facilitate calibrated reasoning about uncertainty, and allow inclusion of prior knowledge about the function space such as periodicity or dependence on contexts such as inputs. To this end, this paper introduces two innovations: (i) a Gaussian process-based hierarchical model for network weights based on unit embeddings that can flexibly encode correlated weight structures, and (ii) input-dependent versions of these weight priors that can provide convenient ways to regularize the function space through the use of kernels defined on contextual inputs. We show these models provide desirable test-time uncertainty estimates on out-of-distribution data, demonstrate cases of modeling inductive biases for neural networks with kernels which help both interpolation and extrapolation from training data, and demonstrate competitive predictive performance on an active learning benchmark.
Theofanis Karaletsos, Thang D. Bui
NeurIPS2
2018 Variational Continual Learning
Viet Cuong Nguyen, Yingzhen Li, Thang D. Bui, Richard E. Turner
ICLR (Poster)3
2018 Neural Graph Learning: Training Neural Networks Using Graphs
abstract
Label propagation is a powerful and flexible semi-supervised learning technique on graphs. Neural networks, on the other hand, have proven track records in many supervised learning tasks. In this work, we propose a training framework with a graph-regularised objective, namely Neural Graph Machines, that can combine the power of neural networks and label propagation. This work generalises previous literature on graph-augmented training of neural networks, enabling it to be applied to multiple neural architectures (Feed-forward NNs, CNNs and LSTM RNNs) and a wide range of graphs. The new objective allows the neural networks to harness both labeled and unlabeled data by: (a)~allowing the network to train using labeled data as in the supervised setting, (b)~biasing the network to learn similar hidden representations for neighboring nodes on a graph, in the same vein as label propagation. Such architectures with the proposed objective can be trained efficiently using stochastic gradient descent and scaled to large graphs, with a runtime that is linear in the number of edges. The proposed joint training approach convincingly outperforms many existing methods on a wide range of tasks (multi-label classification on social graphs, news categorization, document classification and semantic intent classification), with multiple forms of graph inputs (including graphs with and without node-level features) and using different types of neural networks.
Thang D. Bui, Sujith Ravi, Vivek Ramavajjala
WSDM1
2017 Streaming Sparse Gaussian Process Approximations
abstract
Sparse pseudo-point approximations for Gaussian process (GP) models provide a suite of methods that support deployment of GPs in the large data regime and enable analytic intractabilities to be sidestepped. However, the field lacks a principled method to handle streaming data in which both the posterior distribution over function values and the hyperparameter estimates are updated in an online fashion. The small number of existing approaches either use suboptimal hand-crafted heuristics for hyperparameter learning, or suffer from catastrophic forgetting or slow updating when new data arrive. This paper develops a new principled framework for deploying Gaussian process probabilistic models in the streaming setting, providing methods for learning hyperparameters and optimising pseudo-input locations. The proposed framework is assessed using synthetic and real-world datasets.
Thang D. Bui, Viet Cuong Nguyen, Richard E. Turner
NIPS1
2017 A Unifying Framework for Gaussian Process Pseudo-Point Approximations using Power Expectation Propagation
abstract
Gaussian processes (GPs) are flexible distributions over functions that enable high-level assumptions about unknown functions to be encoded in a parsimonious, flexible and general way. Although elegant, the application of GPs is limited by computational and analytical intractabilities that arise when data are sufficiently numerous or when employing non-Gaussian models. Consequently, a wealth of GP approximation schemes have been developed over the last 15 years to address these key limitations. Many of these schemes employ a small set of pseudo data points to summarise the actual data. In this paper we develop a new pseudo-point approximation framework using Power Expectation Propagation (Power EP) that unifies a large number of these pseudo-point approximations. Unlike much of the previous venerable work in this area, the new framework is built on standard methods for approximate inference (variational free- energy, EP and Power EP methods) rather than employing approximations to the probabilistic generative model itself. In this way all of the approximation is performed at `inference time' rather than at `modelling time', resolving awkward philosophical and empirical questions that trouble previous approaches. Crucially, we demonstrate that the new framework includes new pseudo-point approximation methods that outperform current approaches on regression and classification tasks.
Thang D. Bui, Josiah Yan, Richard E. Turner
J. Mach. Learn. Res.1
2016 Deep Gaussian Processes for Regression using Approximate Expectation Propagation
abstract
Deep Gaussian processes (DGPs) are multi-layer hierarchical generalisations of Gaussian processes (GPs) and are formally equivalent to neural networks with multiple, infinitely wide hidden layers. DGPs are nonparametric probabilistic models and as such are arguably more flexible, have a greater capacity to generalise, and provide better calibrated uncertainty estimates than alternative deep models. This paper develops a new approximate Bayesian learning scheme that enables DGPs to be applied to a range of medium to large scale regression problems for the first time. The new method uses an approximate Expectation Propagation procedure and a novel and efficient extension of the probabilistic backpropagation algorithm for learning. We evaluate the new method for non-linear regression on eleven real-world datasets, showing that it always outperforms GP regression and is almost always better than state-of-the-art deterministic and sampling-based approximate inference methods for Bayesian neural networks. As a by-product, this work provides a comprehensive analysis of six approximate Bayesian methods for training neural networks.
Thang D. Bui, Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Yingzhen Li, Richard E. Turner
ICML1
2016 Black-Box Alpha Divergence Minimization
abstract
Black-box alpha (BB-α) is a new approximate inference method based on the minimization of α-divergences. BB-αscales to large datasets because it can be implemented using stochastic gradient descent. BB-αcan be applied to complex probabilistic models with little effort since it only requires as input the likelihood function and its gradients. These gradients can be easily obtained using automatic differentiation. By changing the divergence parameter α, the method is able to interpolate between variational Bayes (VB) (α→0) and an algorithm similar to expectation propagation (EP) (α= 1). Experiments on probit regression and neural network regression and classification problems show that BB-αwith non-standard settings of α, such as α= 0.5, usually produces better predictions than with α→0 (VB) or α= 1 (EP).
José Miguel Hernández-Lobato, Yingzhen Li, Mark Rowland 0001, Thang D. Bui, Daniel Hernández-Lobato, Richard E. Turner
ICML4
2015 Learning Stationary Time Series using Gaussian Processes with Nonparametric Kernels
abstract
We introduce the Gaussian Process Convolution Model (GPCM), a two-stage nonparametric generative procedure to model stationary signals as the convolution between a continuous-time white-noise process and a continuous-time linear filter drawn from Gaussian process. The GPCM is a continuous-time nonparametric-window moving average process and, conditionally, is itself a Gaussian process with a nonparametric kernel defined in a probabilistic fashion. The generative model can be equivalently considered in the frequency domain, where the power spectral density of the signal is specified using a Gaussian process. One of the main contributions of the paper is to develop a novel variational free-energy approach based on inter-domain inducing variables that efficiently learns the continuous-time linear filter and infers the driving white-noise process. In turn, this scheme provides closed-form probabilistic estimates of the covariance kernel and the noise-free signal both in denoising and prediction scenarios. Additionally, the variational inference procedure provides closed-form expressions for the approximate posterior of the spectral density given the observed data, leading to new Bayesian nonparametric approaches to spectrum estimation. The proposed GPCM is validated using synthetic and real-world signals.
Felipe A. Tobar, Thang D. Bui, Richard E. Turner
NIPS2
2014 Tree-structured Gaussian Process Approximations
Thang D. Bui, Richard E. Turner
NIPS1