Marylou Gabrié

dblp:164/5772 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Probabilistic and Bayesian machine learning · 38% Generative modeling · 31% Learning theory · 13%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
sampling
1.622025
Learned Reference-based Diffusion Sampler for multi-modal distributions · ICLR 2025
Stochastic Localization via Iterative Posterior Sampling · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
1.222023
On Sampling with Approximate Transport Maps · ICML 2023
Local-Global MCMC kernels: the best of both worlds · NeurIPS 2022
Machine learning › Generative modeling
diffusion model
0.912025
Learned Reference-based Diffusion Sampler for multi-modal distributions · ICLR 2025
Machine learning › Generative modeling › diffusion model
diffusion sampling
0.912025
Learned Reference-based Diffusion Sampler for multi-modal distributions · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › sampling
multimodal sampling
0.912025
Learned Reference-based Diffusion Sampler for multi-modal distributions · ICLR 2025
Machine learning › Generative modeling
normalizing flow
0.822023
Local-Global MCMC kernels: the best of both worlds · NeurIPS 2022
On Sampling with Approximate Transport Maps · ICML 2023
Robotics › Robot navigation and mapping › localization
probabilistic localization
0.812024
Stochastic Localization via Iterative Posterior Sampling · ICML 2024
Machine learning › Learning theory
generalization
0.512021
On the interplay between data structure and loss function in classification problems · NeurIPS 2021
Machine learning › Learning theory › neural network theory
over-parameterized regime
0.512021
On the interplay between data structure and loss function in classification problems · NeurIPS 2021
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features
0.512021
On the interplay between data structure and loss function in classification problems · NeurIPS 2021
Machine learning › Learning theory
generalization bounds
0.312018
Entropy and mutual information in models of deep neural networks · NeurIPS 2018
Machine learning › Generative modeling › diffusion model
score-based generative model
0.212024
Stochastic Localization via Iterative Posterior Sampling · ICML 2024
Machine learning › Generative modeling › energy-based model
contrastive divergence
0.212015
Training Restricted Boltzmann Machine via the Thouless-Anderson-Palmer free energy · NIPS 2015
Machine learning › Generative modeling
energy-based model
0.212015
Training Restricted Boltzmann Machine via the Thouless-Anderson-Palmer free energy · NIPS 2015
Machine learning › Probabilistic and Bayesian machine learning › boltzmann machine
restricted boltzmann machine
0.212015
Training Restricted Boltzmann Machine via the Thouless-Anderson-Palmer free energy · NIPS 2015
Machine learning › Learning paradigms
unsupervised learning
0.212015
Training Restricted Boltzmann Machine via the Thouless-Anderson-Palmer free energy · NIPS 2015

Methods — techniques the papers use, named apart from their topics

markov chain monte carlo · 1.4normalizing flow · 1.2score-based diffusion · 0.9reference diffusion model · 0.9score-based learning · 0.8denoising · 0.8metropolis-hastings · 0.7geometric ergodicity · 0.6MCMC · 0.6statistical physics · 0.5
YearPublicationVenuePosition
2025 Learned Reference-based Diffusion Sampler for multi-modal distributions
abstract
Over the past few years, several approaches utilizing score-based diffusion have been proposed to sample from probability distributions, that is without having access to exact samples and relying solely on evaluations of unnormalized densities. The resulting samplers approximate the time-reversal of a noising diffusion process, bridging the target distribution to an easy-to-sample base distribution. In practice, the performance of these methods heavily depends on key hyperparameters that require ground truth samples to be accurately tuned. Our work aims to highlight and address this fundamental issue, focusing in particular on multi-modal distributions, which pose significant challenges for existing sampling methods. Building on existing approaches, we introduce *Learned Reference-based Diffusion Sampler* (LRDS), a methodology specifically designed to leverage prior knowledge on the location of the target modes in order to bypass the obstacle of hyperparameter tuning. LRDS proceeds in two steps by (i) learning a *reference* diffusion model on samples located in high-density space regions and tailored for multimodality, and (ii) using this reference model to foster the training of a diffusion-based sampler. We experimentally demonstrate that LRDS best exploits prior knowledge on the target distribution compared to competing algorithms on a variety of challenging distributions.
Maxence Noble, Louis Grenioux, Marylou Gabrié, Alain Durmus
ICLR3
2024 Stochastic Localization via Iterative Posterior Sampling
abstract
Building upon score-based learning, new interest in stochastic localization techniques has recently emerged. In these models, one seeks to noise a sample from the data distribution through a stochastic process, called observation process, and progressively learns a denoiser associated to this dynamics. Apart from specific applications, the use of stochastic localization for the problem of sampling from an unnormalized target density has not been explored extensively. This work contributes to fill this gap. We consider a general stochastic localization framework and introduce an explicit class of observation processes, associated with flexible denoising schedules. We provide a complete methodology, *Stochastic Localization via Iterative Posterior Sampling* (**SLIPS**), to obtain approximate samples of these dynamics, and as a by-product, samples from the target distribution. Our scheme is based on a Markov chain Monte Carlo estimation of the denoiser and comes with detailed practical guidelines. We illustrate the benefits and applicability of **SLIPS** on several benchmarks of multi-modal distributions, including Gaussian mixtures in increasing dimensions, Bayesian logistic regression and a high-dimensional field system from statistical-mechanics.
Louis Grenioux, Maxence Noble, Marylou Gabrié, Alain Durmus
ICML3
2023 On Sampling with Approximate Transport Maps
abstract
Transport maps can ease the sampling of distributions with non-trivial geometries by transforming them into distributions that are easier to handle. The potential of this approach has risen with the development of Normalizing Flows (NF) which are maps parameterized with deep neural networks trained to push a reference distribution towards a target. NF-enhanced samplers recently proposed blend (Markov chain) Monte Carlo methods with either (i) proposal draws from the flow or (ii) a flow-based reparametrization. In both cases, the quality of the learned transport conditions performance. The present work clarifies for the first time the relative strengths and weaknesses of these two approaches. Our study concludes that multimodal targets can be reliably handled with flow-based proposals up to moderately high dimensions. In contrast, methods relying on reparametrization struggle with multimodality but are more robust otherwise in high-dimensional settings and under poor training. To further illustrate the influence of target-proposal adequacy, we also derive a new quantitative bound for the mixing time of the Independent Metropolis-Hastings sampler.
Louis Grenioux, Alain Durmus, Eric Moulines, Marylou Gabrié
ICML4
2022 Adaptation of the Independent Metropolis-Hastings Sampler with Normalizing Flow Proposals
abstract
Markov Chain Monte Carlo (MCMC) methods are a powerful tool for computation with complex probability distributions. However the performance of such methods is critically dependent on properly tuned parameters, most of which are difficult if not impossible to know a priori for a given target distribution. Adaptive MCMC methods aim to address this by allowing the parameters to be updated during sampling based on previous samples from the chain at the expense of requiring a new theoretical analysis to ensure convergence. In this work we extend the convergence theory of adaptive MCMC methods to a new class of methods built on a powerful class of parametric density estimators known as normalizing flows. In particular, we consider an independent Metropolis-Hastings sampler where the proposal distribution is represented by a normalizing flow whose parameters are updated using stochastic gradient descent. We explore the practical performance of this procedure on both synthetic settings and in the analysis of a physical field system, and compare it against both adaptive and non-adaptive MCMC methods.
James A. Brofos, Marylou Gabrié, Marcus A. Brubaker, Roy R. Lederman
AISTATS2
2022 Local-Global MCMC kernels: the best of both worlds
abstract
Recent works leveraging learning to enhance sampling have shown promising results, in particular by designing effective non-local moves and global proposals. However, learning accuracy is inevitably limited in regions where little data is available such as in the tails of distributions as well as in high-dimensional problems. In the present paper we study an Explore-Exploit Markov chain Monte Carlo strategy ($\operatorname{Ex^2MCMC}$) that combines local and global samplers showing that it enjoys the advantages of both approaches. We prove $V$-uniform geometric ergodicity of $\operatorname{Ex^2MCMC}$ without requiring a uniform adaptation of the global sampler to the target distribution. We also compute explicit bounds on the mixing rate of the Explore-Exploit strategy under realistic conditions. Moreover, we propose an adaptive version of the strategy ($\operatorname{FlEx^2MCMC}$) where a normalizing flow is trained while sampling to serve as a proposal for global moves. We illustrate the efficiency of $\operatorname{Ex^2MCMC}$ and its adaptive version on classical sampling benchmarks as well as in sampling high-dimensional distributions defined by Generative Adversarial Networks seen as Energy Based Models.
Sergey Samsonov, Evgeny Lagutin, Marylou Gabrié, Alain Durmus, Alexey Naumov, Eric Moulines
NeurIPS3
2021 On the interplay between data structure and loss function in classification problems
abstract
One of the central features of modern machine learning models, including deep neural networks, is their generalization ability on structured data in the over-parametrized regime. In this work, we consider an analytically solvable setup to investigate how properties of data impact learning in classification problems, and compare the results obtained for quadratic loss and logistic loss. Using methods from statistical physics, we obtain a precise asymptotic expression for the train and test errors of random feature models trained on a simple model of structured data. The input covariance is built from independent blocks allowing us to tune the saliency of low-dimensional structures and their alignment with respect to the target function.Our results show in particular that in the over-parametrized regime, the impact of data structure on both train and test error curves is greater for logistic loss than for mean-squared loss: the easier the task, the wider the gap in performance between the two losses at the advantage of the logistic. Numerical experiments on MNIST and CIFAR10 confirm our insights.
Stéphane d'Ascoli, Marylou Gabrié, Levent Sagun, Giulio Biroli
NeurIPS2
2018 Entropy and mutual information in models of deep neural networks
abstract
We examine a class of stochastic deep learning models with a tractable method to compute information-theoretic quantities. Our contributions are three-fold: (i) We show how entropies and mutual informations can be derived from heuristic statistical physics methods, under the assumption that weight matrices are independent and orthogonally-invariant. (ii) We extend particular cases in which this result is known to be rigorously exact by providing a proof for two-layers networks with Gaussian random weights, using the recently introduced adaptive interpolation method. (iii) We propose an experiment framework with generative models of synthetic datasets, on which we train deep neural networks with a weight constraint designed so that the assumption in (i) is verified during learning. We study the behavior of entropies and mutual information throughout learning and conclude that, in the proposed setting, the relationship between compression and generalization remains elusive.
Marylou Gabrié, Andre Manoel, Clément Luneau, Jean Barbier, Nicolas Macris, Florent Krzakala, Lenka Zdeborová
NeurIPS1
2016 Inferring sparsity: Compressed sensing using generalized restricted Boltzmann machines
abstract
In this work, we consider compressed sensing reconstruction from M measurements of K-sparse structured signals which do not possess a writable correlation model. Assuming that a generative statistical model, such as a Boltzmann machine, can be trained in an unsupervised manner on example signals, we demonstrate how this signal model can be used within a Bayesian framework of signal reconstruction. By deriving a message-passing inference for general distribution restricted Boltzmann machines, we are able to integrate these inferred signal models into approximate message passing for compressed sensing reconstruction. Finally, we show for the MNIST dataset that this approach can be very effective, even for M <; K.
Eric W. Tramel, Andre Manoel, Francesco Caltagirone, Marylou Gabrié, Florent Krzakala
ITW4
2015 Training Restricted Boltzmann Machine via the Thouless-Anderson-Palmer free energy
abstract
Restricted Boltzmann machines are undirected neural networks which have been shown tobe effective in many applications, including serving as initializations fortraining deep multi-layer neural networks. One of the main reasons for their success is theexistence of efficient and practical stochastic algorithms, such as contrastive divergence,for unsupervised training. We propose an alternative deterministic iterative procedure based on an improved mean field method from statistical physics known as the Thouless-Anderson-Palmer approach. We demonstrate that our algorithm provides performance equal to, and sometimes superior to, persistent contrastive divergence, while also providing a clear and easy to evaluate objective function. We believe that this strategycan be easily generalized to other models as well as to more accurate higher-order approximations, paving the way for systematic improvements in training Boltzmann machineswith hidden units.
Marylou Gabrié, Eric W. Tramel, Florent Krzakala
NIPS1