Manon Michel

dblp:339/7114 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 48% Probabilistic and Bayesian machine learning · 30% Deep learning architectures and training · 22%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › statistical learning theory
asymptotic analysis
1.422024
Law of Large Numbers and Central Limit Theorem for Wide Two-layer Neural Networks: The Mini-Batch and Noisy Case · J. Mach. Learn. Res. 2024
Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference · COLT 2023
Machine learning › Deep learning architectures and training › training optimization
stochastic gradient descent dynamics
0.812024
Law of Large Numbers and Central Limit Theorem for Wide Two-layer Neural Networks: The Mini-Batch and Noisy Case · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.712023
Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference · COLT 2023
Machine learning › Learning theory › statistical learning theory › statistical physics of learning
mean-field analysis
0.712023
Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference · COLT 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.712023
Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference · COLT 2023
Machine learning › Deep learning architectures and training › feedforward neural network
two-layer neural network
0.212023
Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference · COLT 2023

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 0.8quadratic loss · 0.8mini-batch · 0.8reparametrization trick · 0.7monte carlo sampling · 0.7evidence lower bound · 0.7
YearPublicationVenuePosition
2024 Fixed-kinetic Neural Hamiltonian Flows for enhanced interpretability and reduced complexity
abstract
Normalizing Flows (NF) are Generative models which transform a simple prior distribution into the desired target. They however require the design of an invertible mapping whose Jacobian determinant has to be computable. Recently introduced, Neural Hamiltonian Flows (NHF) are Hamiltonian dynamics-based flows, which are continuous, volume-preserving and invertible and thus make for natural candidates for robust NF architectures. In particular, their similarity to classical Mechanics could lead to easier interpretability of the learned mapping. In this paper, we show that the current NHF architecture may still pose a challenge to interpretability. Inspired by Physics, we introduce a fixed-kinetic energy version of the model. This approach improves interpretability and robustness while requiring fewer parameters than the original model. We illustrate that on a 2D Gaussian mixture and on the MNIST and Fashion-MNIST datasets. Finally, we show how to adapt NHF to the context of Bayesian inference and illustrate the method on an example from cosmology.
Vincent Souveton, Arnaud Guillin, Jens Jasche, Guilhem Lavaux, Manon Michel
AISTATS5
2024 Law of Large Numbers and Central Limit Theorem for Wide Two-layer Neural Networks: The Mini-Batch and Noisy Case
abstract
In this work, we consider a wide two-layer neural network and study the behavior of its empirical weights under a dynamics set by a stochastic gradient descent along the quadratic loss with mini-batches and noise. Our goal is to prove a trajectorial law of large number as well as a central limit theorem for their evolution. When the noise is scaling as $1/N^\beta$ and $1/2<\beta\le\infty$, we rigorously derive and generalize the LLN obtained for example by Rotskoff and Van den Injden (Com. Pure. Appl. Math, 2022), Mei and Montanari and Nguyen (Pnas 2018) or Sirignano and Spiliopoulos (Siam. J. Appl. Math. 2020). When $3/4<\beta\le\infty$, we also generalize the CLT of Sirignano and Spiliopoulos (Stoch. Proc. Appl. 2020) and further exhibit the effect of mini-batching on the asymptotic variance which leads the fluctuations. The case $\beta=3/4$ is trickier and we give an example showing the divergence with time of the variance thus establishing the instability of the predictions of the neural network in this case. It is illustrated by simple numerical examples.
Arnaud Descours, Arnaud Guillin, Manon Michel, Boris Nectoux
J. Mach. Learn. Res.3
2023 Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference
abstract
We provide a rigorous analysis of training by variational inference (VI) of Bayesian neural networks in the two-layer and infinite-width case. We consider a regression problem with a regularized evidence lower bound (ELBO) which is decomposed into the expected log-likelihood of the data and the Kullback-Leibler (KL) divergence between the a priori distribution and the variational posterior. With an appropriate weighting of the KL, we prove a law of large numbers for three different training schemes: (i) the idealized case with exact estimation of a multiple Gaussian integral from the reparametrization trick, (ii) a minibatch scheme using Monte Carlo sampling, commonly known as Bayes by Backprop, and (iii) a new and computationally cheaper algorithm which we introduce as Minimal VI. An important result is that all methods converge to the same mean-field limit. Finally, we illustrate our results numerically and discuss the need for the derivation of a central limit theorem.
Arnaud Descours, Tom Huix, Arnaud Guillin, Manon Michel, Eric Moulines, Boris Nectoux
COLT4