EDBT 2026 Demo / reviewers in the wild / expert
Giulio Biroli
dblp:18/5547
· DBLP profile ↗
16ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-4115-2644ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Learning theory · 34% Deep learning architectures and training · 22% Optimization for machine learning · 9% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
training dynamics |
2.8 | 5 | 2025 | Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training · NeurIPS 2025 Cascade of phase transitions in the training of energy-based models · NeurIPS 2024 An analytic theory of shallow networks dynamics for hinge loss classification · NeurIPS 2020 |
Machine learning › Learning theory
generalization |
2.6 | 5 | 2025 | Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training · NeurIPS 2025 On the interplay between data structure and loss function in classification problems · NeurIPS 2021 Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features |
0.9 | 2 | 2021 | On the interplay between data structure and loss function in classification problems · NeurIPS 2021 Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training · NeurIPS 2025 |
Machine learning › Learning theory › over-parameterization
double descent |
0.9 | 2 | 2020 | Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 Double Trouble in Double Descent: Bias and Variance(s) in the Lazy Regime · ICML 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | A Differentiable Rank-Based Objective for Better Feature Learning · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
memorization |
0.9 | 1 | 2025 | Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training · NeurIPS 2025 |
Machine learning › Learning theory
overfitting |
0.9 | 2 | 2020 | Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 An analytic theory of shallow networks dynamics for hinge loss classification · NeurIPS 2020 |
Machine learning › Learning theory › model selection
variable selection |
0.9 | 1 | 2025 | A Differentiable Rank-Based Objective for Better Feature Learning · ICLR 2025 |
Machine learning › Optimization for machine learning
gradient flow |
0.8 | 2 | 2020 | Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval · NeurIPS 2020 Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor models · NeurIPS 2019 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.8 | 2 | 2020 | Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval · NeurIPS 2020 Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor models · NeurIPS 2019 |
Machine learning › Generative modeling
energy-based model |
0.8 | 1 | 2024 | Cascade of phase transitions in the training of energy-based models · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › boltzmann machine
restricted boltzmann machine |
0.8 | 1 | 2024 | Cascade of phase transitions in the training of energy-based models · NeurIPS 2024 |
Machine learning › Learning theory › statistical learning theory
statistical physics of learning |
0.6 | 2 | 2024 | Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor models · NeurIPS 2019 Cascade of phase transitions in the training of energy-based models · NeurIPS 2024 |
Machine learning › Learning theory
inductive bias |
0.6 | 1 | 2022 | Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual Tasks · ICML 2022 |
Machine learning › Efficient and distributed learning › model compression › pruning
iterative magnitude pruning |
0.6 | 1 | 2022 | Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual Tasks · ICML 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual Tasks · ICML 2022 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.6 | 1 | 2022 | Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual Tasks · ICML 2022 |
Mathematical optimization
nonconvex optimization |
0.5 | 2 | 2020 | Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval · NeurIPS 2020 Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.5 | 2 | 2021 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias · NeurIPS 2019 ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases · ICML 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases · ICML 2021 |
Machine learning › Learning theory › neural network theory
over-parameterized regime |
0.5 | 1 | 2021 | On the interplay between data structure and loss function in classification problems · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.5 | 1 | 2021 | ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases · ICML 2021 |
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff |
0.4 | 1 | 2020 | Double Trouble in Double Descent: Bias and Variance(s) in the Lazy Regime · ICML 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation |
0.4 | 1 | 2020 | An analytic theory of shallow networks dynamics for hinge loss classification · NeurIPS 2020 |
Machine learning › Optimization for machine learning
phase retrieval |
0.4 | 1 | 2020 | Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › feedforward neural network
shallow neural networks |
0.4 | 1 | 2020 | An analytic theory of shallow networks dynamics for hinge loss classification · NeurIPS 2020 |
Mathematical optimization › nonconvex optimization › optimization landscape
spurious local minima |
0.4 | 1 | 2020 | Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
loss landscape |
0.4 | 1 | 2019 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias · NeurIPS 2019 |
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry |
0.3 | 1 | 2018 | Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
statistical physics · 2.4mean-field theory · 1.9u-net · 0.9random features model · 0.9neural network regularizer · 0.9differentiable approximation · 0.9conditional independence · 0.9singular value decomposition · 0.8finite-size scaling · 0.8fully connected network · 0.6hessian analysis · 0.4BBP transition · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Differentiable Rank-Based Objective for Better Feature LearningabstractIn this paper, we leverage existing statistical methods to better understand feature learning from data. We tackle this by modifying the model-free variable selection method, Feature Ordering by Conditional Independence (FOCI), which is introduced in Azadkia & Chatterjee (2021). While FOCI is based on a non-parametric coefficient of conditional dependence, we introduce its parametric, differentiable approximation. With this approximate coefficient of correlation, we present a new algorithm called difFOCI, which is applicable to a wider range of machine learning problems thanks to its differentiable nature and learnable parameters. We present difFOCI in three contexts: (1) as a variable selection method with baseline comparisons to FOCI, (2) as a trainable model parametrized with a neural network, and (3) as a generic, widely applicable neural network regularizer, one that improves feature learning with better management of spurious correlations. We evaluate difFOCI on increasingly complex problems ranging from basic variable selection in toy examples to saliency map comparisons in convolutional networks. We then show how difFOCI can be incorporated in the context of fairness to facilitate classifications without relying on sensitive data. Krunoslav Lehman Pavasovic, Giulio Biroli, Levent Sagun |
ICLR | 2 |
| 2025 | Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingabstractDiffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time $\tau_\mathrm{gen}$ at which models begin to generate high-quality samples, and a later time $\tau_\mathrm{mem}$ beyond which memorization emerges. Crucially, we find that $\tau_\mathrm{mem}$ increases linearly with the training set size $n$, while $\tau_\mathrm{gen}$ remains constant. This creates a growing window of training times with $n$ where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when $n$ becomes larger than a model-dependent threshold that overfitting disappears at infinite training times.
These findings reveal a form of implicit dynamical regularization in the training dynamics, which allow to avoid memorization even in highly overparameterized settings. Our results are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, and by a theoretical analysis using a tractable random features model studied in the high-dimensional limit. Tony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc Mézard |
NeurIPS | 3 |
| 2024 | On the Impact of Overparameterization on the Training of a Shallow Neural Network in High DimensionsabstractWe study the training dynamics of a shallow neural network with quadratic activation functions and quadratic cost in a teacher-student setup. In line with previous works on the same neural architecture, the optimization is performed following the gradient flow on the population risk, where the average over data points is replaced by the expectation over their distribution, assumed to be Gaussian. We first derive convergence properties for the gradient flow and quantify the overparameterization that is necessary to achieve a strong signal recovery. Then, assuming that the teachers and the students at initialization form independent orthonormal families, we derive a high-dimensional limit for the flow and show that the minimal overparameterization is sufficient for strong recovery. We verify by numerical experiments that these results hold for more general initializations. Simon Martin 0006, Francis R. Bach, Giulio Biroli |
AISTATS | 3 |
| 2024 | Cascade of phase transitions in the training of energy-based modelsabstractIn this paper, we investigate the feature encoding process in a prototypical energy-based generative model, the Restricted Boltzmann Machine (RBM). We start with an analytical investigation using simplified architectures and data structures, and end with numerical analysis of real trainings on real datasets. Our study tracks the evolution of the model’s weight matrix through its singular value decomposition, revealing a series of thermodynamic phase transitions that shape the principal learning modes of the empirical probability distribution. We first describe this process analytically in several controlled setups that allow us to fully monitor the training dynamics until convergence. We then validate these findings by training the Bernoulli-Bernoulli RBM on real data sets. By studying the phase behavior over data sets of increasing dimension, we show that these phase transitions are genuine in the thermodynamic sense. Moreover, we propose a mean-field finite-size scaling hypothesis, confirming that the initial phase transition, reminiscent of the paramagnetic-to-ferromagnetic phase transition in mean-field ferromagnetism models, is governed by mean-field critical exponents. Dimitrios Bachtis, Giulio Biroli, Aurélien Decelle, Beatriz Seoane |
NeurIPS | 2 |
| 2022 | Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual TasksabstractPruning methods can considerably reduce the size of artificial neural networks without harming their performance and in some cases they can even uncover sub-networks that, when trained in isolation, match or surpass the test accuracy of their dense counterparts. Here, we characterize the inductive bias that pruning imprints in such "winning lottery tickets": focusing on visual tasks, we analyze the architecture resulting from iterative magnitude pruning of a simple fully connected network. We show that the surviving node connectivity is local in input space, and organized in patterns reminiscent of the ones found in convolutional networks. We investigate the role played by data and tasks in shaping the architecture of the pruned sub-network. We find that pruning performances, and the ability to sift out the noise and make local features emerge, improve by increasing the size of the training set, and the semantic value of the data. We also study different pruning procedures, and find that iterative magnitude pruning is particularly effective in distilling meaningful connectivity out of features present in the original task. Our results suggest the possibility to automatically discover new and efficient architectural inductive biases in other datasets and tasks. Franco Pellegrini, Giulio Biroli |
ICML | 2 |
| 2021 | ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesabstractConvolutional architectures have proven extremely successful for vision tasks. Their hard inductive biases enable sample-efficient learning, but come at the cost of a potentially lower performance ceiling. Vision Transformers (ViTs) rely on more flexible self-attention layers, and have recently outperformed CNNs for image classification. However, they require costly pre-training on large external datasets or distillation from pre-trained convolutional networks. In this paper, we ask the following question: is it possible to combine the strengths of these two architectures while avoiding their respective limitations? To this end, we introduce gated positional self-attention (GPSA), a form of positional self-attention which can be equipped with a “soft" convolutional inductive bias. We initialise the GPSA layers to mimic the locality of convolutional layers, then give each attention head the freedom to escape locality by adjusting a gating parameter regulating the attention paid to position versus content information. The resulting convolutional-like ViT architecture, ConViT, outperforms the DeiT on ImageNet, while offering a much improved sample efficiency. We further investigate the role of locality in learning by first quantifying how it is encouraged in vanilla self-attention layers, then analysing how it is escaped in GPSA layers. We conclude by presenting various ablations to better understand the success of the ConViT. Our code and models are released publicly at https://github.com/facebookresearch/convit. Stéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos, Giulio Biroli, Levent Sagun |
ICML | 5 |
| 2021 | On the interplay between data structure and loss function in classification problemsabstractOne of the central features of modern machine learning models, including deep neural networks, is their generalization ability on structured data in the over-parametrized regime. In this work, we consider an analytically solvable setup to investigate how properties of data impact learning in classification problems, and compare the results obtained for quadratic loss and logistic loss. Using methods from statistical physics, we obtain a precise asymptotic expression for the train and test errors of random feature models trained on a simple model of structured data. The input covariance is built from independent blocks allowing us to tune the saliency of low-dimensional structures and their alignment with respect to the target function.Our results show in particular that in the over-parametrized regime, the impact of data structure on both train and test error curves is greater for logistic loss than for mean-squared loss: the easier the task, the wider the gap in performance between the two losses at the advantage of the logistic. Numerical experiments on MNIST and CIFAR10 confirm our insights. Stéphane d'Ascoli, Marylou Gabrié, Levent Sagun, Giulio Biroli |
NeurIPS | 4 |
| 2020 | Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeabstractDeep neural networks can achieve remarkable generalization performances while interpolating the training data. Rather than the U-curve emblematic of the bias-variance trade-off, their test error often follows a "double descent"—a mark of the beneficial role of overparametrization. In this work, we develop a quantitative theory for this phenomenon in the so-called lazy learning regime of neural networks, by considering the problem of learning a high-dimensional function with random features regression. We obtain a precise asymptotic expression for the bias-variance decomposition of the test error, and show that the bias displays a phase transition at the interpolation threshold, beyond it which it remains constant. We disentangle the variances stemming from the sampling of the dataset, from the additive noise corrupting the labels, and from the initialization of the weights. We demonstrate that the latter two contributions are the crux of the double descent: they lead to the overfitting peak at the interpolation threshold and to the decay of the test error upon overparametrization. We quantify how they are suppressed by ensembling the outputs of $K$ independently initialized estimators. For $K\rightarrow \infty$, the test error is monotonously decreasing and remains constant beyond the interpolation threshold. We further compare the effects of overparametrizing, ensembling and regularizing. Finally, we present numerical experiments on classic deep learning setups to show that our results hold qualitatively in realistic lazy learning scenarios. Stéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent Krzakala |
ICML | 3 |
| 2020 | Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase RetrievalabstractDespite the widespread use of gradient-based algorithms for optimising high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retrieval from random measurements. When the ratio of the number of measurements over the input dimension is small the dynamics remains trapped in spurious minima with large basins of attraction. We find analytically that above a critical ratio those critical points become unstable developing a negative direction toward the signal. By numerical experiments we show that in this regime the gradient flow algorithm is not trapped; it drifts away from the spurious critical points along the unstable direction and succeeds in finding the global minimum. Using tools from statistical physics we characterise this phenomenon, which is related to a BBP-type transition in the Hessian of the spurious minima. Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Pierfrancesco Urbani, Lenka Zdeborová |
NeurIPS | 2 |
| 2020 | An analytic theory of shallow networks dynamics for hinge loss classificationabstractNeural networks have been shown to perform incredibly well in classification tasks over structured high-dimensional datasets. However, the learning dynamics of such networks is still poorly understood. In this paper we study in detail the training dynamics of a simple type of neural network: a single hidden layer trained to perform a classification task. We show that in a suitable mean-field limit this case maps to a single-node learning problem with a time-dependent dataset determined self-consistently from the average nodes population. We specialize our theory to the prototypical case of a linearly separable dataset and a linear hinge loss, for which the dynamics can be explicitly solved in the infinite dataset limit. This allow us to address in a simple setting several phenomena appearing in modern networks such as slowing down of training dynamics, crossover between feature and lazy learning, and overfitting. Finally, we asses the limitations of mean-field theory by studying the case of large but finite number of nodes and of training samples. Franco Pellegrini, Giulio Biroli |
NeurIPS | 2 |
| 2020 | Triple descent and the two kinds of overfitting: where & why do they appear?abstractA recent line of research has highlighted the existence of a ``double descent'' phenomenon in deep learning, whereby increasing the number of training examples N causes the generalization error of neural networks to peak when N is of the same order as the number of parameters P. In earlier works, a similar phenomenon was shown to exist in simpler models such as linear regression, where the peak instead occurs when N is equal to the input dimension D. Since both peaks coincide with the interpolation threshold, they are often conflated in the litterature. In this paper, we show that despite their apparent similarity, these two scenarios are inherently different. In fact, both peaks can co-exist when neural networks are applied to noisy regression tasks. The relative size of the peaks is then governed by the degree of nonlinearity of the activation function. Building on recent developments in the analysis of random feature models, we provide a theoretical ground for this sample-wise triple descent. As shown previously, the nonlinear peak at N=P is a true divergence caused by the extreme sensitivity of the output function to both the noise corrupting the labels and the initialization of the random features (or the weights in neural networks). This peak survives in the absence of noise, but can be suppressed by regularization. In contrast, the linear peak at N=D is solely due to overfitting the noise in the labels, and forms earlier during training. We show that this peak is implicitly regularized by the nonlinearity, which is why it only becomes salient at high noise and is weakly affected by explicit regularization. Throughout the paper, we compare the analytical results obtained in the random feature model with the outcomes of numerical experiments involving realistic neural networks. Stéphane d'Ascoli, Levent Sagun, Giulio Biroli |
NeurIPS | 3 |
| 2020 | Complex interactions can create persistent fluctuations in high-diversity ecosystemsabstractWhen can ecological interactions drive an entire ecosystem into a persistent non-equilibrium state, where many species populations fluctuate without going to extinction? We show that high-diversity spatially heterogeneous systems can exhibit chaotic dynamics which persist for extremely long times. We develop a theoretical framework, based on dynamical mean-field theory, to quantify the conditions under which these fluctuating states exist, and predict their properties. We uncover parallels with the persistence of externally-perturbed ecosystems, such as the role of perturbation strength, synchrony and correlation time. But uniquely to endogenous fluctuations, these properties arise from the species dynamics themselves, creating feedback loops between perturbation and response. A key result is that fluctuation amplitude and species diversity are tightly linked: in particular, fluctuations enable dramatically more species to coexist than at equilibrium in the very same system. Our findings highlight crucial differences between well-mixed and spatially-extended systems, with implications for experiments and their ability to reproduce natural dynamics. They shed light on the maintenance of biodiversity, and the strength and synchrony of fluctuations observed in natural systems. Felix Roy, Matthieu Barbier, Giulio Biroli, Guy Bunin |
PLoS Comput. Biol. | 3 |
| 2019 | Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor modelsabstractGradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in optimising high-dimensional non-convex functions and why they find good minima instead of being trapped in spurious ones.Here we present a quantitative theory explaining this behaviour in a spiked matrix-tensor model.Our framework is based on the Kac-Rice analysis of stationary points and a closed-form analysis of gradient-flow originating from statistical physics. We show that there is a well defined region of parameters where the gradient-flow algorithm finds a good global minimum despite the presence of exponentially many spurious local minima. We show that this is achieved by surfing on saddles that have strong negative direction towards the global minima, a phenomenon that is connected to a BBP-type threshold in the Hessian describing the critical points of the landscapes. Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Lenka Zdeborová |
NeurIPS | 2 |
| 2019 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural biasabstractDespite the phenomenal success of deep neural networks in a broad range of learning tasks, there is a lack of theory to understand the way they work. In particular, Convolutional Neural Networks (CNNs) are known to perform much better than Fully-Connected Networks (FCNs) on spatially structured data: the architectural structure of CNNs benefits from prior knowledge on the features of the data, for instance their translation invariance. The aim of this work is to understand this fact through the lens of dynamics in the loss landscape. We introduce a method that maps a CNN to its equivalent FCN (denoted as eFCN). Such an embedding enables the comparison of CNN and FCN training dynamics directly in the FCN space. We use this method to test a new training protocol, which consists in training a CNN, embedding it to FCN space at a certain ``relax time'', then resuming the training in FCN space. We observe that for all relax times, the deviation from the CNN subspace is small, and the final performance reached by the eFCN is higher than that reachable by a standard FCN of same architecture. More surprisingly, for some intermediate relax times, the eFCN outperforms the CNN it stemmed, by combining the prior information of the CNN and the expressivity of the FCN in a complementary way. The practical interest of our protocol is limited by the very large size of the highly sparse eFCN. However, it offers interesting insights into the persistence of architectural bias under stochastic gradient dynamics. It shows the existence of some rare basins in the FCN loss landscape associated with very good generalization. These can only be accessed thanks to the CNN prior, which helps navigate the landscape during the early stages of optimization. Stéphane d'Ascoli, Levent Sagun, Giulio Biroli, Joan Bruna |
NeurIPS | 3 |
| 2018 | Comparing Dynamics: Deep Neural Networks versus Glassy SystemsabstractWe analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are the complexity of the loss-landscape and of the dynamics within it, and to what extent DNNs share similarities with glassy systems. Our findings, obtained for different architectures and data-sets, suggest that during the training process the dynamics slows down because of an increasingly large number of flat directions. At large times, when the loss is approaching zero, the system diffuses at the bottom of the landscape. Despite some similarities with the dynamics of mean-field glassy systems, in particular, the absence of barrier crossing, we find distinctive dynamical behaviors in the two cases, thus showing that the statistical properties of the corresponding loss and energy landscapes are different. In contrast, when the network is under-parametrized we observe a typical glassy behavior, thus suggesting the existence of different phases depending on whether the network is under-parametrized or over-parametrized. Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gérard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, Giulio Biroli |
ICML | 9 |
| 2012 | Social Interaction, Noise and Antibiotic-Mediated Switches in the Intestinal MicrobiotaabstractThe intestinal microbiota plays important roles in digestion and resistance against entero-pathogens. As with other ecosystems, its species composition is resilient against small disturbances but strong perturbations such as antibiotics can affect the consortium dramatically. Antibiotic cessation does not necessarily restore pre-treatment conditions and disturbed microbiota are often susceptible to pathogen invasion. Here we propose a mathematical model to explain how antibiotic-mediated switches in the microbiota composition can result from simple social interactions between antibiotic-tolerant and antibiotic-sensitive bacterial groups. We build a two-species (e.g. two functional-groups) model and identify regions of domination by antibiotic-sensitive or antibiotic-tolerant bacteria, as well as a region of multistability where domination by either group is possible. Using a new framework that we derived from statistical physics, we calculate the duration of each microbiota composition state. This is shown to depend on the balance between random fluctuations in the bacterial densities and the strength of microbial interactions. The singular value decomposition of recent metagenomic data confirms our assumption of grouping microbes as antibiotic-tolerant or antibiotic-sensitive in response to a single antibiotic. Our methodology can be extended to multiple bacterial groups and thus it provides an ecological formalism to help interpret the present surge in microbiome data. Vanni Bucci, Serena Bradde, Giulio Biroli, João B. Xavier |
PLoS Comput. Biol. | 3 |