Michael U. Gutmann

dblp:133/7772 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-5329-9910ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Probabilistic and Bayesian machine learning · 62% Generative modeling · 18% Deep learning architectures and training · 7%
Theoretical computer science
1 paper
Information theory · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 22 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference
1.742023
Is Learning Summary Statistics Necessary for Likelihood-free Inference? · ICML 2023
Neural Approximate Sufficient Statistics for Implicit Models · ICLR 2021
ELFI: Engine for Likelihood-Free Inference · J. Mach. Learn. Res. 2018
Machine learning › Generative modeling
variational autoencoder
0.922023
Variational Gibbs Inference for Statistical Model Estimation from Incomplete Data · J. Mach. Learn. Res. 2023
VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning · NIPS 2017
Machine learning › Deep learning architectures and training › deep generative model
implicit models
0.922021
Implicit Deep Adaptive Design: Policy-Based Experimental Design without Likelihoods · NeurIPS 2021
Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning
copula models
0.912025
Neural Mutual Information Estimation with Vector Copulas · NeurIPS 2025
Information theory › information measures › mutual information
mutual information estimation
0.912025
Neural Mutual Information Estimation with Vector Copulas · NeurIPS 2025
Machine learning › Generative modeling
generative adversarial network
0.722020
Generative Ratio Matching Networks · ICLR 2020
VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
approximate bayesian computation
0.722023
Neural Approximate Sufficient Statistics for Implicit Models · ICLR 2021
Is Learning Summary Statistics Necessary for Likelihood-free Inference? · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.712023
Variational Gibbs Inference for Statistical Model Estimation from Incomplete Data · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.712023
Variational Gibbs Inference for Statistical Model Estimation from Incomplete Data · J. Mach. Learn. Res. 2023
Computational science and engineering › information retrieval
rank aggregation
0.612022
Systematic comparison of ranking aggregation methods for gene lists in experimental results · Bioinform. 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
density ratio estimation
0.622020
Telescoping Density-Ratio Estimation · NeurIPS 2020
Generative Ratio Matching Networks · ICLR 2020
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
bayesian optimal experimental design
0.512021
Implicit Deep Adaptive Design: Policy-Based Experimental Design without Likelihoods · NeurIPS 2021
Machine learning › Learning theory › statistical estimation
sufficient statistics
0.512021
Neural Approximate Sufficient Statistics for Implicit Models · ICLR 2021
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.412020
Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation · ICML 2020
Machine learning › Generative modeling
energy-based model
0.412020
Telescoping Density-Ratio Estimation · NeurIPS 2020
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation
0.412020
Telescoping Density-Ratio Estimation · NeurIPS 2020
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412020
Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.412020
Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.312018
Conditional Noise-Contrastive Estimation of Unnormalised Models · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
noise contrastive estimation
0.312018
Conditional Noise-Contrastive Estimation of Unnormalised Models · ICML 2018
Machine learning › Generative modeling › generative adversarial network › GAN training
mode collapse
0.312017
VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning · NIPS 2017
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.212016
Bayesian Optimization for Likelihood-Free Inference of Simulator-Based Statistical Models · J. Mach. Learn. Res. 2016

Methods — techniques the papers use, named apart from their topics

neural network · 3.3copula theory · 1.7bayesian optimization · 0.8variational inference · 0.7sufficient statistics learning · 0.7maximum likelihood estimation · 0.7gibbs sampling · 0.7meta-analysis by information content · 0.6implicit models · 0.5design policy network · 0.5amortized inference · 0.5probabilistic modeling of discrepancy · 0.2gaussian process · 0.2
YearPublicationVenuePosition
2025 Neural Mutual Information Estimation with Vector Copulas
abstract
Estimating mutual information (MI) is a fundamental task in data science and machine learning. Existing estimators mainly rely on either highly flexible models (e.g., neural networks), which require large amounts of data, or overly simplified models (e.g., Gaussian copula), which fail to capture complex distributions. Drawing upon recent vector copula theory, we propose a principled interpolation between these two extremes to achieve a better trade-off between complexity and capacity. Experiments on state-of-the-art synthetic benchmarks and real-world data with diverse modalities demonstrate the advantages of the proposed method.
Yanzhi Chen, Zijing Ou, Adrian Weller, Michael U. Gutmann
NeurIPS4
2023 Is Learning Summary Statistics Necessary for Likelihood-free Inference?
abstract
Likelihood-free inference (LFI) is a set of techniques for inference in implicit statistical models. A longstanding question in LFI has been how to design or learn good summary statistics of data, but this might now seem unnecessary due to the advent of recent end-to-end (i.e. neural network-based) LFI methods. In this work, we rethink this question with a new method for learning summary statistics. We show that learning sufficient statistics may be easier than direct posterior inference, as the former problem can be reduced to a set of low-dimensional, easy-to-solve learning problems. This suggests us to explicitly decouple summary statistics learning from posterior inference in LFI. Experiments on diverse inference tasks with different data types validate our hypothesis.
Yanzhi Chen, Michael U. Gutmann, Adrian Weller
ICML2
2023 Variational Gibbs Inference for Statistical Model Estimation from Incomplete Data
abstract
Statistical models are central to machine learning with broad applicability across a range of downstream tasks. The models are controlled by free parameters that are typically estimated from data by maximum-likelihood estimation or approximations thereof. However, when faced with real-world data sets many of the models run into a critical issue: they are formulated in terms of fully-observed data, whereas in practice the data sets are plagued with missing data. The theory of statistical model estimation from incomplete data is conceptually similar to the estimation of latent-variable models, where powerful tools such as variational inference (VI) exist. However, in contrast to standard latent-variable models, parameter estimation with incomplete data often requires estimating exponentially-many conditional distributions of the missing variables, hence making standard VI methods intractable. We address this gap by introducing variational Gibbs inference (VGI), a new general-purpose method to estimate the parameters of statistical models from incomplete data. We validate VGI on a set of synthetic and real-world estimation tasks, estimating important machine learning models such as variational autoencoders and normalising flows from incomplete data. The proposed method, whilst general-purpose, achieves competitive or better performance than existing model-specific estimation methods.
Vaidotas Simkus, Benjamin Rhodes, Michael U. Gutmann
J. Mach. Learn. Res.3
2022 Systematic comparison of ranking aggregation methods for gene lists in experimental results
abstract
MOTIVATION: A common experimental output in biomedical science is a list of genes implicated in a given biological process or disease. The gene lists resulting from a group of studies answering the same, or similar, questions can be combined by ranking aggregation methods to find a consensus or a more reliable answer. Evaluating a ranking aggregation method on a specific type of data before using it is required to support the reliability since the property of a dataset can influence the performance of an algorithm. Such evaluation on gene lists is usually based on a simulated database because of the lack of a known truth for real data. However, simulated datasets tend to be too small compared to experimental data and neglect key features, including heterogeneity of quality, relevance and the inclusion of unranked lists. RESULTS: In this study, a group of existing methods and their variations that are suitable for meta-analysis of gene lists are compared using simulated and real data. Simulated data were used to explore the performance of the aggregation methods as a function of emulating the common scenarios of real genomic data, with various heterogeneity of quality, noise level and a mix of unranked and ranked data using 20 000 possible entities. In addition to the evaluation with simulated data, a comparison using real genomic data on the SARS-CoV-2 virus, cancer (non-small cell lung cancer) and bacteria (macrophage apoptosis) was performed. We summarize the results of our evaluation in a simple flowchart to select a ranking aggregation method, and in an automated implementation using the meta-analysis by information content algorithm to infer heterogeneity of data quality across input datasets. AVAILABILITY AND IMPLEMENTATION: The code for simulated data generation and running edited version of algorithms: https://github.com/baillielab/comparison_of_RA_methods. Code to perform an optimal selection of methods based on the results of this review, using the MAIC algorithm to infer the characteristics of an input dataset, can be downloaded here: https://github.com/baillielab/maic. An online service for running MAIC: https://baillielab.net/maic. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bo Wang 0118, Andy Law, Tim Regan, Nicholas Parkinson, Joby Cole, Clark D. Russell, David H. Dockrell, Michael U. Gutmann, J. Kenneth Baillie
Bioinform.8
2021 Bayesian Experimental Design for Intractable Models of Cognition
Simon Valentin, Steven Kleinegesse, Neil Bramley, Michael U. Gutmann, Christopher G. Lucas
CogSci4
2021 Neural Approximate Sufficient Statistics for Implicit Models
Yanzhi Chen, Dinghuai Zhang, Michael U. Gutmann, Aaron C. Courville, Zhanxing Zhu
ICLR3
2021 Implicit Deep Adaptive Design: Policy-Based Experimental Design without Likelihoods
abstract
We introduce implicit Deep Adaptive Design (iDAD), a new method for performing adaptive experiments in real-time with implicit models. iDAD amortizes the cost of Bayesian optimal experimental design (BOED) by learning a design policy network upfront, which can then be deployed quickly at the time of the experiment. The iDAD network can be trained on any model which simulates differentiable samples, unlike previous design policy work that requires a closed form likelihood and conditionally independent experiments. At deployment, iDAD allows design decisions to be made in milliseconds, in contrast to traditional BOED approaches that require heavy computation during the experiment itself. We illustrate the applicability of iDAD on a number of experiments, and show that it provides a fast and effective mechanism for performing adaptive design with implicit models.
Desi R. Ivanova, Adam Foster 0001, Steven Kleinegesse, Michael U. Gutmann, Tom Rainforth
NeurIPS4
2020 Robust Optimisation Monte Carlo
abstract
This paper is on Bayesian inference for parametric statistical models that are defined by a stochastic simulator which specifies how data is generated. Exact sampling is then possible but evaluating the likelihood function is typically prohibitively expensive. Approximate Bayesian Computation (ABC) is a framework to perform approximate inference in such situations. While basic ABC algorithms are widely applicable, they are notoriously slow and much research has focused on increasing their efficiency. Optimisation Monte Carlo (OMC) has recently been proposed as an efficient and embarrassingly parallel method that leverages optimisation to accelerate the inference. In this paper, we demonstrate an important previously unrecognised failure mode of OMC: It generates strongly overconfident approximations by collapsing regions of similar or near-constant likelihood into a single point. We propose an efficient, robust generalisation of OMC that corrects this. It makes fewer assumptions, retains the main benefits of OMC, and can be performed either as post-processing to OMC or as a stand-alone computation. We demonstrate the effectiveness of the proposed Robust OMC on toy examples and tasks in inverse-graphics where we perform Bayesian inference with a complex image renderer.
Borislav Ikonomov, Michael U. Gutmann
AISTATS2
2020 Generative Ratio Matching Networks
Akash Srivastava, Kai Xu 0016, Michael U. Gutmann, Charles Sutton
ICLR3
2020 Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation
abstract
Implicit stochastic models, where the data-generation distribution is intractable but sampling is possible, are ubiquitous in the natural sciences. The models typically have free parameters that need to be inferred from data collected in scientific experiments. A fundamental question is how to design the experiments so that the collected data are most useful. The field of Bayesian experimental design advocates that, ideally, we should choose designs that maximise the mutual information (MI) between the data and the parameters. For implicit models, however, this approach is severely hampered by the high computational cost of computing posteriors and maximising MI, in particular when we have more than a handful of design variables to optimise. In this paper, we propose a new approach to Bayesian experimental design for implicit models that leverages recent advances in neural MI estimation to deal with these issues. We show that training a neural network to maximise a lower bound on MI allows us to jointly determine the optimal design and the posterior. Simulation studies illustrate that this gracefully extends Bayesian experimental design for implicit models to higher design dimensions.
Steven Kleinegesse, Michael U. Gutmann
ICML2
2020 Stir to Pour: Efficient Calibration of Liquid Properties for Pouring Actions
abstract
Humans use simple probing actions to develop intuition about the physical behavior of common objects. Such intuition is particularly useful for adaptive estimation of favorable manipulation strategies of those objects in novel contexts. For example, observing the effect of tilt on a transparent bottle containing an unknown liquid provides clues on how the liquid might be poured. It is desirable to equip general-purpose robotic systems with this capability because it is inevitable that they will encounter novel objects and scenarios. In this paper, we teach a robot to use a simple, specified probing strategy - stirring with a stick- to reduce spillage when pouring unknown liquids. In the probing step, we continuously observe the effects of a real robot stirring a liquid, while simultaneously tuning the parameters to a model (simulator) until the two outputs are in agreement. We obtain optimal simulation parameters, characterizing the unknown liquid, via a Bayesian Optimizer that minimizes the discrepancy between real and simulated outcomes. Then, we optimize the pouring policy conditioning on the optimal simulation parameters determined via stirring. We show that using stirring as a probing strategy result in reduced spillage for three qualitatively different liquids when executed on a UR10 Robot, compared to probing via pouring. Finally, we provide quantitative insights into the reason for stirring being a suitable calibration task for pouring -a step towards automatic discovery of probing strategies.
Tatiana Lopez-Guevara, Rita Pucci, Nick K. Taylor, Michael U. Gutmann, Subramanian Ramamoorthy, Kartic Subr
IROS4
2020 Telescoping Density-Ratio Estimation
abstract
Density-ratio estimation via classification is a cornerstone of unsupervised learning. It has provided the foundation for state-of-the-art methods in representation learning and generative modelling, with the number of use-cases continuing to proliferate. However, it suffers from a critical limitation: it fails to accurately estimate ratios p/q for which the two densities differ significantly. Empirically, we find this occurs whenever the KL divergence between p and q exceeds tens of nats. To resolve this limitation, we introduce a new framework, telescoping density-ratio estimation (TRE), that enables the estimation of ratios between highly dissimilar densities in high-dimensional spaces. Our experiments demonstrate that TRE can yield substantial improvements over existing single-ratio methods for mutual information estimation, representation learning and energy-based modelling.
Benjamin Rhodes, Kai Xu 0016, Michael U. Gutmann
NeurIPS3
2019 Adaptive Gaussian Copula ABC
abstract
Approximate Bayesian computation (ABC) is a set of techniques for Bayesian inference when the likelihood is intractable but sampling from the model is possible. This work presents a simple yet effective ABC algorithm based on the combination of two classical ABC approaches — regression ABC and sequential ABC. The key idea is that rather than learning the posterior directly, we first target another auxiliary distribution that can be learned accurately by existing methods, through which we then subsequently learn the desired posterior with the help of a Gaussian copula. During this process, the complexity of the model changes adaptively according to the data at hand. Experiments on a synthetic dataset as well as three real-world inference tasks demonstrates that the proposed method is fast, accurate, and easy to use.
Yanzhi Chen, Michael U. Gutmann
AISTATS2
2019 Efficient Bayesian Experimental Design for Implicit Models
abstract
Bayesian experimental design involves the optimal allocation of resources in an experiment, with the aim of optimising cost and performance. For implicit models, where the likelihood is intractable but sampling from the model is possible, this task is particularly difficult and therefore largely unexplored. This is mainly due to technical difficulties associated with approximating posterior distributions and utility functions. We devise a novel experimental design framework for implicit models that improves upon previous work in two ways. First, we use the mutual information between parameters and data as the utility function, which has previously not been feasible. We achieve this by utilising Likelihood-Free Inference by Ratio Estimation (LFIRE) to approximate posterior distributions, instead of the traditional approximate Bayesian computation or synthetic likelihood methods. Secondly, we use Bayesian optimisation in order to solve the optimal design problem, as opposed to the typically used grid search or sampling-based methods. We find that this increases efficiency and allows us to consider higher design dimensions.
Steven Kleinegesse, Michael U. Gutmann
AISTATS2
2019 Variational Noise-Contrastive Estimation
abstract
Unnormalised latent variable models are a broad and flexible class of statistical models. However, learning their parameters from data is intractable, and few estimation techniques are currently available for such models. To increase the number of techniques in our arsenal, we propose variational noise-contrastive estimation (VNCE), building on NCE which is a method that only applies to unnormalised models. The core idea is to use a variational lower bound to the NCE objective function, which can be optimised in the same fashion as the evidence lower bound (ELBO) in standard variational inference (VI). We prove that VNCE can be used for both parameter estimation of unnormalised models and posterior inference of latent variables. The developed theory shows that VNCE has the same level of generality as standard VI, meaning that advances made there can be directly imported to the unnormalised setting. We validate VNCE on toy models and apply it to a realistic problem of estimating an undirected graphical model from incomplete data.
Benjamin Rhodes, Michael U. Gutmann
AISTATS2
2018 Conditional Noise-Contrastive Estimation of Unnormalised Models
abstract
Many parametric statistical models are not properly normalised and only specified up to an intractable partition function, which renders parameter estimation difficult. Examples of unnormalised models are Gibbs distributions, Markov random fields, and neural network models in unsupervised deep learning. In previous work, the estimation principle called noise-contrastive estimation (NCE) was introduced where unnormalised models are estimated by learning to distinguish between data and auxiliary noise. An open question is how to best choose the auxiliary noise distribution. We here propose a new method that addresses this issue. The proposed method shares with NCE the idea of formulating density estimation as a supervised learning problem but in contrast to NCE, the proposed method leverages the observed data when generating noise samples. The noise can thus be generated in a semi-automated manner. We first present the underlying theory of the new method, show that score matching emerges as a limiting case, validate the method on continuous and discrete valued synthetic data, and show that we can expect an improved performance compared to NCE when the data lie in a lower-dimensional manifold. Then we demonstrate its applicability in unsupervised deep learning by estimating a four-layer neural image model.
Ciwan Ceylan, Michael U. Gutmann
ICML2
2018 ELFI: Engine for Likelihood-Free Inference
abstract
Engine for Likelihood-Free Inference (ELFI) is a Python software library for performing likelihood-free inference (LFI). ELFI provides a convenient syntax for arranging components in LFI, such as priors, simulators, summaries or distances, to a network called ELFI graph. The components can be implemented in a wide variety of languages. The stand-alone ELFI graph can be used with any of the available inference methods without modifications. A central method implemented in ELFI is Bayesian Optimization for Likelihood-Free Inference (BOLFI), which has recently been shown to accelerate likelihood-free inference up to several orders of magnitude by surrogate-modelling the distance. ELFI also has an inbuilt support for output data storing for reuse and analysis, and supports parallelization of computation from multiple cores up to a cluster environment. ELFI is designed to be extensible and provides interfaces for widening its functionality. This makes the adding of new inference methods to ELFI straightforward and automatically compatible with the inbuilt features.
Jarno Lintusaari, Henri Vuollekoski, Antti Kangasrääsiö, Kusti Skytén, Marko Järvenpää, Pekka Marttinen, Michael U. Gutmann, Aki Vehtari, Jukka Corander, Samuel Kaski
J. Mach. Learn. Res.7
2017 VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning
abstract
Deep generative models provide powerful tools for distributions over complicated manifolds, such as those of natural images. But many of these methods, including generative adversarial networks (GANs), can be difficult to train, in part because they are prone to mode collapse, which means that they characterize only a few modes of the true distribution. To address this, we introduce VEEGAN, which features a reconstructor network, reversing the action of the generator by mapping from data to noise. Our training objective retains the original asymptotic consistency guarantee of GANs, and can be interpreted as a novel autoencoder loss over the noise. In sharp contrast to a traditional autoencoder over data points, VEEGAN does not require specifying a loss function over the data, but rather only over the representations, which are standard normal by assumption. On an extensive set of synthetic and real world image datasets, VEEGAN indeed resists mode collapsing to a far greater extent than other recent GAN variants, and produces more realistic samples.
Akash Srivastava, Lazar Valkov, Chris Russell 0001, Michael U. Gutmann, Charles Sutton
NIPS4
2016 Bayesian Optimization for Likelihood-Free Inference of Simulator-Based Statistical Models
abstract
Our paper deals with inferring simulator-based statistical models given some observed data. A simulator-based model is a parametrized mechanism which specifies how data are generated. It is thus also referred to as generative model. We assume that only a finite number of parameters are of interest and allow the generative process to be very general; it may be a noisy nonlinear dynamical system with an unrestricted number of hidden variables. This weak assumption is useful for devising realistic models but it renders statistical inference very difficult. The main challenge is the intractability of the likelihood function. Several likelihood-free inference methods have been proposed which share the basic idea of identifying the parameters by finding values for which the discrepancy between simulated and observed data is small. A major obstacle to using these methods is their computational cost. The cost is largely due to the need to repeatedly simulate data sets and the lack of knowledge about how the parameters affect the discrepancy. We propose a strategy which combines probabilistic modeling of the discrepancy with optimization to facilitate likelihood-free inference. The strategy is implemented using Bayesian optimization and is shown to accelerate the inference through a reduction in the number of required simulations by several orders of magnitude.
Michael U. Gutmann, Jukka Corander
J. Mach. Learn. Res.1
2014 Direct Learning of Sparse Changes in Markov Networks by Density Ratio Estimation
abstract
We propose a new method for detecting changes in Markov network structure between two sets of samples. Instead of naively fitting two Markov network models separately to the two data sets and figuring out their difference, we directly learn the network structure change by estimating the ratio of Markov network models. This density-ratio formulation naturally allows us to introduce sparsity in the network structure change, which highly contributes to enhancing interpretability. Furthermore, computation of the normalization term, a critical bottleneck of the naive approach, can be remarkably mitigated. We also give the dual formulation of the optimization problem, which further reduces the computation cost for large-scale Markov networks. Through experiments, we demonstrate the usefulness of our method.
Song Liu 0002, John A. Quinn, Michael U. Gutmann, Taiji Suzuki, Masashi Sugiyama
Neural Comput.3
2013 Direct Learning of Sparse Changes in Markov Networks by Density Ratio Estimation
Song Liu 0002, John A. Quinn, Michael U. Gutmann, Masashi Sugiyama
ECML/PKDD (2)3