Pierre Glaser

dblp:296/0262 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Probabilistic and Bayesian machine learning · 58% Learning theory · 16% Optimization for machine learning · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 65% Computational science and engineering · 35%
Theoretical computer science
3 papers
Mathematical optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
divergence minimization
1.422025
(De)-regularized Maximum Mean Discrepancy Gradient Flow · J. Mach. Learn. Res. 2025
KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support · NeurIPS 2021
Machine learning › Optimization for machine learning
gradient flow
0.912025
(De)-regularized Maximum Mean Discrepancy Gradient Flow · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.912025
Efficiently Vectorized MCMC on Modern Accelerators · ICML 2025
Computational science and engineering
latent variable model
0.912025
SIMPL: Scalable and hassle-free optimisation of neural representations from behaviour · ICLR 2025
Bioinformatics and computational biology › neuroscience › neuroinformatics
neural data analysis
0.912025
SIMPL: Scalable and hassle-free optimisation of neural representations from behaviour · ICLR 2025
Machine learning › Learning theory › statistical learning theory
asymptotic analysis
0.812024
Near-Optimality of Contrastive Divergence Algorithms · NeurIPS 2024
Machine learning › Generative modeling › energy-based model
contrastive divergence
0.812024
Near-Optimality of Contrastive Divergence Algorithms · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.812024
Near-Optimality of Contrastive Divergence Algorithms · NeurIPS 2024
Bioinformatics and computational biology
protein design
0.812024
Kernel-Based Evaluation of Conditional Biological Sequence Models · ICML 2024
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
gradient flow
0.512021
KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support · NeurIPS 2021
Mathematical optimization › optimal transport
wasserstein gradient flow
0.312025
(De)-regularized Maximum Mean Discrepancy Gradient Flow · J. Mach. Learn. Res. 2025
Mathematical optimization
stochastic optimization
0.212024
Near-Optimality of Contrastive Divergence Algorithms · NeurIPS 2024
Machine learning › Learning theory › probability metric
integral probability metric
0.112021
KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support · NeurIPS 2021
Machine learning › Learning theory › probability metric › integral probability metric
maximum mean discrepancy
0.112021
KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

finite state machine · 1.7de-regularization · 1.7automatic vectorization · 1.7adaptive schedule · 1.7JAX vmap · 1.7non-asymptotic analysis · 1.5maximum mean discrepancy · 1.5kernel methods · 1.5hypothesis testing · 1.5cramér-rao lower bound · 1.5latent variable model · 0.9EM algorithm · 0.9reproducing kernel hilbert space · 0.5particle method · 0.5fenchel duality · 0.5
YearPublicationVenuePosition
2025 SIMPL: Scalable and hassle-free optimisation of neural representations from behaviour
abstract
Neural activity in the brain is known to encode low-dimensional, time-evolving, behaviour-related variables. A long-standing goal of neural data analysis has been to identify these variables and their mapping to neural activity. A productive and canonical approach has been to simply visualise neural "tuning curves" as a function of behaviour. However, significant discrepancies between behaviour and the true latent variables -- such as an agent thinking of position Y whilst located at position X -- distort and blur the tuning curves, decreasing their interpretability. To address this, latent variable models propose to learn the latent variable from data; these are typically expensive, hard to tune, or scale poorly, complicating their adoption. Here we propose SIMPL (Scalable Iterative Maximization of Population-coded Latents), an EM-style algorithm which iteratively optimises latent variables and tuning curves. SIMPL is fast, scalable and exploits behaviour as an initial condition to further improve convergence and identifiability. It can accurately recover latent variables in spatial and non-spatial tasks. When applied to a large hippocampal dataset SIMPL converges on smaller, more numerous, and more uniformly sized place fields than those based on behaviour, suggesting the brain may encode space with greater resolution than previously thought.
Tom M. George, Pierre Glaser, Kimberly L. Stachenfeld, Caswell Barry, Claudia Clopath
ICLR2
2025 Efficiently Vectorized MCMC on Modern Accelerators
abstract
With the advent of automatic vectorization tools (e.g., JAX’s vmap), writing multi-chain MCMC algorithms is often now as simple as invoking those tools on single-chain code. Whilst convenient, for various MCMC algorithms this results in a synchronization problem—loosely speaking, at each iteration all chains running in parallel must wait until the last chain has finished drawing its sample. In this work, we show how to design single-chain MCMC algorithms in a way that avoids synchronization overheads when vectorizing with tools like vmap, by using the framework of finite state machines (FSMs). Using a simplified model, we derive an exact theoretical form of the obtainable speed-ups using our approach, and use it to make principled recommendations for optimal algorithm design. We implement several popular MCMC algorithms as FSMs, including Elliptical Slice Sampling, HMC-NUTS, and Delayed Rejection, demonstrating speed-ups of up to an order of magnitude in experiments.
Hugh Dance, Pierre Glaser, Peter Orbanz, Ryan P. Adams
ICML2
2025 (De)-regularized Maximum Mean Discrepancy Gradient Flow
abstract
We introduce a (de)-regularization of the Maximum Mean Discrepancy (DrMMD) and its Wasserstein gradient flow. Existing gradient flows that transport samples from source distribution to target distribution with only target samples, either lack tractable numerical implementation ($f$-divergence flows) or require strong assumptions and modifications, such as noise injection, to ensure convergence (Maximum Mean Discrepancy flows). In contrast, DrMMD flow can simultaneously (i) guarantee near-global convergence for a broad class of targets in both continuous and discrete time, and (ii) be implemented in closed form using only samples. The former is achieved by leveraging the connection between the DrMMD and the $\chi^2$-divergence, while the latter comes by treating DrMMD as MMD with a de-regularized kernel. Our numerical scheme employs an adaptive de-regularization schedule throughout the flow to optimally balance the trade-off between discretization errors and deviations from the $\chi^2$ regime. The potential application of the DrMMD flow is demonstrated across several numerical experiments, including a large-scale setting of training student/teacher networks.
Zonghao Chen, Aratrika Mustafi, Pierre Glaser, Anna Korba, Arthur Gretton, Bharath K. Sriperumbudur
J. Mach. Learn. Res.3
2024 Kernel-Based Evaluation of Conditional Biological Sequence Models
abstract
We propose a set of kernel-based tools to evaluate the designs and tune the hyperparameters of conditional sequence models, with a focus on problems in computational biology. The backbone of our tools is a new measure of discrepancy between the true conditional distribution and the model's estimate, called the Augmented Conditional Maximum Mean Discrepancy (ACMMD). Provided that the model can be sampled from, the ACMMD can be estimated unbiasedly from data to quantify absolute model fit, integrated within hypothesis tests, and used to evaluate model reliability. We demonstrate the utility of our approach by analyzing a popular protein design model, ProteinMPNN. We are able to reject the hypothesis that ProteinMPNN fits its data for various protein families, and tune the model's temperature hyperparameter to achieve a better fit.
Pierre Glaser, Steffanie Paul, Alissa M. Hummer, Charlotte M. Deane, Debora S. Marks, Alan Nawzad Amin
ICML1
2024 Near-Optimality of Contrastive Divergence Algorithms
abstract
We provide a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an $O(n^{-1 / 3})$ rate to the true parameter of the data distribution, we show that CD can achieve the parametric rate $O(n^{-1 / 2})$. Our analysis provides results for various data batching schemes, including fully online and minibatch. We additionally show that CD is near-optimal, in the sense that its asymptotic variance is close to the Cramér-Rao lower bound.
Pierre Glaser, Kevin Han Huang, Arthur Gretton
NeurIPS1
2023 Fast and scalable score-based kernel calibration tests
abstract
We introduce the Kernel Calibration Conditional Stein Discrepancy test (KCCSD test), a nonparametric, kernel-based test for assessing the calibration of probabilistic models with well-defined scores. In contrast to previous methods, our test avoids the need for possibly expensive expectation approximations while providing control over its type-I error. We achieve these improvements by using a new family of kernels for score-based probabilities that can be estimated without probability density samples, and by using a Conditional Goodness of Fit criterion for the KCCSD test’s U-statistic. We demonstrate the properties of our test on various synthetic settings.
Pierre Glaser, David Widmann, Fredrik Lindsten, Arthur Gretton
UAI1
2021 KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support
abstract
We study the gradient flow for a relaxed approximation to the Kullback-Leibler (KL) divergencebetween a moving source and a fixed target distribution.This approximation, termed theKALE (KL approximate lower-bound estimator), solves a regularized version ofthe Fenchel dual problem defining the KL over a restricted class of functions.When using a Reproducing Kernel Hilbert Space (RKHS) to define the functionclass, we show that the KALE continuously interpolates between the KL and theMaximum Mean Discrepancy (MMD). Like the MMD and other Integral ProbabilityMetrics, the KALE remains well defined for mutually singulardistributions. Nonetheless, the KALE inherits from the limiting KL a greater sensitivity to mismatch in the support of the distributions, compared with the MMD. These two properties make theKALE gradient flow particularly well suited when the target distribution is supported on a low-dimensional manifold. Under an assumption of sufficient smoothness of the trajectories, we show the global convergence of the KALE flow. We propose a particle implementation of the flow given initial samples from the source and the target distribution, which we use to empirically confirm the KALE's properties.
Pierre Glaser, Michael Arbel, Arthur Gretton
NeurIPS1