VLDB 2026 Research / reviewers in the wild / expert
Henry B. Moss
dblp:222/2926
· DBLP profile ↗
13ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0003-0427-7675ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Optimization for machine learning · 52% Probabilistic and Bayesian machine learning · 28% Reinforcement learning · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
3.3 | 5 | 2025 | Omnipresent Yet Overlooked: Heat Kernels in Combinatorial Bayesian Optimization · NeurIPS 2025 Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 GAUCHE: A Library for Gaussian Processes in Chemistry · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
2.0 | 3 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 GAUCHE: A Library for Gaussian Processes in Chemistry · NeurIPS 2023 BOSS: Bayesian Optimization over String Spaces · NeurIPS 2020 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
combinatorial bayesian optimization |
0.9 | 1 | 2025 | Omnipresent Yet Overlooked: Heat Kernels in Combinatorial Bayesian Optimization · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
kernel design |
0.9 | 1 | 2025 | Omnipresent Yet Overlooked: Heat Kernels in Combinatorial Bayesian Optimization · NeurIPS 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
latent space bayesian optimization |
0.9 | 1 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.9 | 1 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
acquisition function |
0.5 | 1 | 2021 | GIBBON: General-purpose Information-Based Bayesian Optimisation · J. Mach. Learn. Res. 2021 |
Machine learning › Reinforcement learning
bandit |
0.5 | 1 | 2021 | Scalable Thompson Sampling using Sparse Gaussian Process Models · NeurIPS 2021 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
information-theoretic acquisition function |
0.5 | 1 | 2021 | GIBBON: General-purpose Information-Based Bayesian Optimisation · J. Mach. Learn. Res. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
sparse gaussian process |
0.5 | 1 | 2021 | Scalable Thompson Sampling using Sparse Gaussian Process Models · NeurIPS 2021 |
Machine learning › Reinforcement learning
thompson sampling |
0.5 | 1 | 2021 | Scalable Thompson Sampling using Sparse Gaussian Process Models · NeurIPS 2021 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › structured kernel
string kernel |
0.4 | 1 | 2020 | BOSS: Bayesian Optimization over String Spaces · NeurIPS 2020 |
Mathematical optimization
combinatorial optimization |
0.3 | 1 | 2025 | Omnipresent Yet Overlooked: Heat Kernels in Combinatorial Bayesian Optimization · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
theoretical analysis · 1.7heat kernel derivation · 1.7closed-form expressions · 1.7decoupled surrogate modeling · 0.9bayesian update · 0.9gaussian process · 0.7bayesian optimization · 0.7sparse gaussian process · 0.5regret analysis · 0.5acquisition function maximization · 0.4bandit algorithm · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured SpacesabstractBayesian optimisation in the latent space of a VAE is a powerful framework for optimisation tasks over complex structured domains, such as the space of valid molecules. However, existing approaches tightly couple the surrogate and generative models, which can lead to suboptimal performance when the latent space is not tailored to specific tasks, which in turn has led to the proposal of increasingly sophisticated algorithms. In this work, we explore a new direction, instead proposing a decoupled approach that trains a generative model and a GP surrogate separately, then combines them via a simple yet principled Bayesian update rule. This separation allows each component to focus on its strengths— structure generation from the VAE and predictive modelling by the GP. We show that our decoupled approach improves our ability to identify high-potential candidates in molecular optimisation problems under constrained evaluation budgets. Henry B. Moss, Sebastian W. Ober, Tom Diethe |
ICML | 1 |
| 2025 | Omnipresent Yet Overlooked: Heat Kernels in Combinatorial Bayesian OptimizationabstractBayesian Optimization (BO) has the potential to solve various combinatorial tasks, ranging from materials science to neural architecture search. However, BO requires specialized kernels to effectively model combinatorial domains. Recent efforts have introduced several combinatorial kernels, but the relationships among them are not well understood. To bridge this gap, we develop a unifying framework based on heat kernels, which we derive in a systematic way and express as simple closed-form expressions. Using this framework, we prove that many successful combinatorial kernels are either related or equivalent to heat kernels, and validate this theoretical claim in our experiments. Moreover, our analysis confirms and extends the results presented in Bounce: certain algorithms' performance decreases substantially when the unknown optima of the function do not have a certain structure. In contrast, heat kernels are not sensitive to the location of the optima. Lastly, we show that a fast and simple pipeline, relying on heat kernels, is able to achieve state-of-the-art results, matching or even outperforming certain slow or complex algorithms. Colin Doumont, Victor Picheny, Viacheslav Borovitskiy, Henry B. Moss |
NeurIPS | 4 |
| 2023 | Inducing Point Allocation for Sparse Gaussian Processes in High-Throughput Bayesian OptimisationabstractSparse Gaussian processes are a key component of high-throughput Bayesian optimisation (BO) loops; however, we show that existing methods for allocating their inducing points severely hamper optimisation performance. By exploiting the quality-diversity decomposition of determinantal point processes, we propose the first inducing point allocation strategy designed specifically for use in BO. Unlike existing methods which seek only to reduce global uncertainty in the objective function, our approach provides the local high-fidelity modelling of promising regions required for precise optimisation. More generally, we demonstrate that our proposed framework provides a flexible way to allocate modelling capacity in sparse models and so is suitable for a broad range of downstream sequential decision making tasks. Henry B. Moss, Sebastian W. Ober, Victor Picheny |
AISTATS | 1 |
| 2023 | PF2ES: Parallel Feasible Pareto Frontier Entropy Search for Multi-Objective Bayesian OptimizationabstractWe present Parallel Feasible Pareto Frontier Entropy Search ($\{\mathrm{PF}\}^2$ES) — a novel information-theoretic acquisition function for multi-objective Bayesian optimization supporting unknown constraints and batch queries. Due to the complexity of characterizing the mutual information between candidate evaluations and (feasible) Pareto frontiers, existing approaches must either employ crude approximations that significantly hamper their performance or rely on expensive inference schemes that substantially increase the optimization’s computational overhead. By instead using a variational lower bound, $\{\mathrm{PF}\}^2$ES provides a low-cost and accurate estimate of the mutual information. We benchmark $\{\mathrm{PF}\}^2$ES against other information-theoretic acquisition functions, demonstrating its competitive performance for optimization across synthetic and real-world design problems. Jixiang Qing, Henry B. Moss, Tom Dhaene, Ivo Couckuyt |
AISTATS | 2 |
| 2023 | GAUCHE: A Library for Gaussian Processes in ChemistryabstractWe introduce GAUCHE, an open-source library for GAUssian processes in CHEmistry. Gaussian processes have long been a cornerstone of probabilistic machine learning, affording particular advantages for uncertainty quantification and Bayesian optimisation. Extending Gaussian processes to molecular representations, however, necessitates kernels defined over structured inputs such as graphs, strings and bit vectors. By providing such kernels in a modular, robust and easy-to-use framework, we seek to enable expert chemists and materials scientists to make use of state-of-the-art black-box optimization techniques. Motivated by scenarios frequently encountered in practice, we showcase applications for GAUCHE in molecular discovery, chemical reaction optimisation and protein design. The codebase is made available at https://github.com/leojklarner/gauche. Ryan-Rhys Griffiths, Leo Klarner, Henry B. Moss, Aditya Ravuri, Sang Truong, Yuanqi Du, Samuel Stanton, Gary Tom, Bojana Rankovic, Arian Rokkum Jamasb, Aryan Deshwal, Julius Schwartz, Austin Tripp, Gregory Kell, Simon Frieder, Anthony Bourached, Alex Chan, Jacob Moss, Chengzhi Guo, Johannes Peter Dürholt, Saudamini Chaurasia, Ji Won Park, Felix Strieth-Kalthoff, Alpha A. Lee, Bingqing Cheng, Alán Aspuru-Guzik, Philippe Schwaller, Jian Tang 0005 |
NeurIPS | 3 |
| 2022 | Bayesian quantile and expectile optimisationabstractBayesian optimisation (BO) is widely used to optimise stochastic black box functions. While most BO approaches focus on optimising conditional expectations, many applications require risk-averse strategies and alternative criteria accounting for the distribution tails need to be considered. In this paper, we propose new variational models for Bayesian quantile and expectile regression that are well-suited for heteroscedastic noise settings. Our models consist of two latent Gaussian processes accounting respectively for the conditional quantile (or expectile) and the scale parameter of an asymmetric likelihood functions. Furthermore, we propose two BO strategies based on max-value entropy search and Thompson sampling, that are tailored to such models and that can accommodate large batches of points. Contrary to existing BO approaches for risk-averse optimisation, our strategies can directly optimise for the quantile and expectile, without requiring replicating observations or assuming a parametric form for the noise. As illustrated in the experimental section, the proposed approach clearly outperforms the state of the art in the heteroscedastic, non-Gaussian case. Victor Picheny, Henry B. Moss, Léeonard Torossian, Nicolas Durrande |
UAI | 2 |
| 2021 | Scalable Thompson Sampling using Sparse Gaussian Process ModelsabstractThompson Sampling (TS) from Gaussian Process (GP) models is a powerful tool for the optimization of black-box functions. Although TS enjoys strong theoretical guarantees and convincing empirical performance, it incurs a large computational overhead that scales polynomially with the optimization budget. Recently, scalable TS methods based on sparse GP models have been proposed to increase the scope of TS, enabling its application to problems that are sufficiently multi-modal, noisy or combinatorial to require more than a few hundred evaluations to be solved. However, the approximation error introduced by sparse GPs invalidates all existing regret bounds. In this work, we perform a theoretical and empirical analysis of scalable TS. We provide theoretical guarantees and show that the drastic reduction in computational complexity of scalable TS can be enjoyed without loss in the regret performance over the standard TS. These conceptual claims are validated for practical implementations of scalable TS on synthetic benchmarks and as part of a real-world high-throughput molecular design task. Sattar Vakili, Henry B. Moss, Artem Artemev, Vincent Dutordoir, Victor Picheny |
NeurIPS | 2 |
| 2021 | GIBBON: General-purpose Information-Based Bayesian OptimisationabstractThis paper describes a general-purpose extension of max-value entropy search, a popular approach for Bayesian Optimisation (BO). A novel approximation is proposed for the information gain -- an information-theoretic quantity central to solving a range of BO problems, including noisy, multi-fidelity and batch optimisations across both continuous and highly-structured discrete spaces. Previously, these problems have been tackled separately within information-theoretic BO, each requiring a different sophisticated approximation scheme, except for batch BO, for which no computationally-lightweight information-theoretic approach has previously been proposed. GIBBON (General-purpose Information-Based Bayesian OptimisatioN) provides a single principled framework suitable for all the above, out-performing existing approaches whilst incurring substantially lower computational overheads. In addition, GIBBON does not require the problem's search space to be Euclidean and so is the first high-performance yet computationally light-weight acquisition function that supports batch BO over general highly structured input spaces like molecular search and gene design. Moreover, our principled derivation of GIBBON yields a natural interpretation of a popular batch BO heuristic based on determinantal point processes. Finally, we analyse GIBBON across a suite of synthetic benchmark tasks, a molecular search loop, and as part of a challenging batch multi-fidelity framework for problems with controllable experimental noise. Henry B. Moss, David S. Leslie, Javier González 0002, Paul Rayson |
J. Mach. Learn. Res. | 1 |
| 2020 | BOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian OptimizationabstractWe present BOFFIN TTS (Bayesian Optimization For FIne-tuning Neural Text To Speech), a novel approach for few-shot speaker adaptation. Here, the task is to fine-tune a pre-trained TTS model to mimic a new speaker using a small corpus of target utterances. We demonstrate that there does not exist a one-size-fits-all adaptation strategy, with convincing synthesis requiring a corpus-specific configuration of the hyper-parameters that control fine-tuning. By using Bayesian optimization to efficiently optimize these hyper-parameter values for a target speaker, we are able to perform adaptation with an average 30% improvement in speaker similarity over standard techniques. Results indicate, across multiple corpora, that BOFFIN TTS can learn to synthesize new speakers using less than ten minutes of audio, achieving the same naturalness as produced for the speakers used to train the base model. Henry B. Moss, Vatsal Aggarwal, Nishant Prateek, Javier González 0002, Roberto Barra-Chicote |
ICASSP | 1 |
| 2020 | BOSS: Bayesian Optimization over String SpacesabstractThis article develops a Bayesian optimization (BO) method which acts directly over raw strings, proposing the first uses of string kernels and genetic algorithms within BO loops. Recent applications of BO over strings have been hindered by the need to map inputs into a smooth and unconstrained latent space. Learning this projection is computationally and data-intensive. Our approach instead builds a powerful Gaussian process surrogate model based on string kernels, naturally supporting variable length inputs, and performs efficient acquisition function maximization for spaces with syntactic constraints. Experiments demonstrate considerably improved optimization over existing approaches across a broad range of constraints, including the popular setting where syntax is governed by a context-free grammar. Henry B. Moss, David S. Leslie, Daniel Beck, Javier González 0002, Paul Rayson |
NeurIPS | 1 |
| 2020 | MUMBO: MUlti-task Max-Value Bayesian Optimization
Henry B. Moss, David S. Leslie, Paul Rayson |
ECML/PKDD (3) | 1 |
| 2019 | FIESTA: Fast IdEntification of State-of-The-Art models using adaptive bandit algorithmsabstractWe present FIESTA, a model selection approach that significantly reduces the computational resources required to reliably identify state-of-the-art performance from large collections of candidate models.Despite being known to produce unreliable comparisons, it is still common practice to compare model evaluations based on single choices of random seeds.We show that reliable model selection also requires evaluations based on multiple train-test splits (contrary to common practice in many shared tasks).Using bandit theory from the statistics literature, we are able to adaptively determine appropriate numbers of data splits and random seeds used to evaluate each model, focusing computational resources on the evaluation of promising models whilst avoiding wasting evaluations on models with lower performance.Furthermore, our userfriendly Python implementation produces confidence guarantees of correctly selecting the optimal model.We evaluate our algorithms by selecting between 8 target-dependent sentiment analysis methods using dramatically fewer model evaluations than current model selection approaches. Henry B. Moss, Andrew Moore 0001, David S. Leslie, Paul Rayson |
ACL (1) | 1 |
| 2018 | Using J-K-fold Cross Validation To Reduce Variance When Tuning NLP ModelsabstractK-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so our performance estimates are in fact stochastic, with variability that can be substantial for natural language processing tasks. We demonstrate that these unstable estimates cannot be relied upon for effective parameter tuning. The resulting tuned parameters are highly sensitive to how our data is partitioned, meaning that we often select sub-optimal parameter choices and have serious reproducibility issues. Instead, we propose to use the less variable J-K-fold CV, in which J independent K-fold cross validations are used to assess performance. Our main contributions are extending J-K-fold CV from performance estimation to parameter tuning and investigating how to choose J and K. We argue that variability is more important than bias for effective tuning and so advocate lower choices of K than are typically seen in the NLP literature and instead use the saved computation to increase J. To demonstrate the generality of our recommendations we investigate a wide range of case-studies: sentiment classification (both general and target-specific), part-of-speech tagging and document classification. Henry B. Moss, David S. Leslie, Paul Rayson |
COLING | 1 |