VLDB 2026 Research / reviewers in the wild / expert
Levent Sagun
dblp:155/9866
· DBLP profile ↗
12ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-5403-4124ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Trustworthy machine learning · 30% Learning theory · 22% Deep learning architectures and training · 15% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.5 | 3 | 2025 | An Effective Theory of Bias Amplification · ICLR 2025 A Differentiable Rank-Based Objective for Better Feature Learning · ICLR 2025 Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction · ICML 2024 |
Machine learning › Learning theory
generalization |
1.3 | 3 | 2021 | On the interplay between data structure and loss function in classification problems · NeurIPS 2021 Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias · NeurIPS 2019 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features |
0.9 | 2 | 2021 | On the interplay between data structure and loss function in classification problems · NeurIPS 2021 Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification |
0.9 | 1 | 2025 | An Effective Theory of Bias Amplification · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression |
0.9 | 1 | 2025 | An Effective Theory of Bias Amplification · ICLR 2025 |
Machine learning › Learning theory › model selection
variable selection |
0.9 | 1 | 2025 | A Differentiable Rank-Based Objective for Better Feature Learning · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness › fair graph learning
fairness in link prediction |
0.8 | 1 | 2024 | Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction · ICML 2024 |
Machine learning › Graph learning
graph neural network |
0.8 | 1 | 2024 | Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction · ICML 2024 |
Machine learning › Graph learning
link prediction |
0.8 | 1 | 2024 | Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction · ICML 2024 |
Machine learning › Trustworthy machine learning › fairness
within-group fairness |
0.8 | 1 | 2024 | Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction · ICML 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
0.7 | 2 | 2019 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias · NeurIPS 2019 Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.7 | 2 | 2019 | A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks · ICML 2019 Entropy-SGD: Biasing Gradient Descent Into Wide Valleys · ICLR (Poster) 2017 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.5 | 2 | 2021 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias · NeurIPS 2019 ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases · ICML 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases · ICML 2021 |
Machine learning › Learning theory › neural network theory
over-parameterized regime |
0.5 | 1 | 2021 | On the interplay between data structure and loss function in classification problems · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.5 | 1 | 2021 | ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases · ICML 2021 |
Machine learning › Learning theory › over-parameterization
double descent |
0.4 | 1 | 2020 | Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 |
Machine learning › Learning theory
overfitting |
0.4 | 1 | 2020 | Triple descent and the two kinds of overfitting: where & why do they appear? · NeurIPS 2020 |
Machine learning › Optimization for machine learning › stochastic optimization
heavy-tailed gradient noise |
0.4 | 1 | 2019 | A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks · ICML 2019 |
Machine learning › Deep learning architectures and training
loss landscape |
0.4 | 1 | 2019 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias · NeurIPS 2019 |
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry |
0.3 | 1 | 2018 | Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018 |
Machine learning › Reinforcement learning › regularization for reinforcement learning
entropy regularization |
0.3 | 1 | 2017 | Entropy-SGD: Biasing Gradient Descent Into Wide Valleys · ICLR (Poster) 2017 |
Machine learning › Probabilistic and Bayesian machine learning › continuous-time model
stochastic differential equations |
0.1 | 1 | 2019 | A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks · ICML 2019 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.1 | 1 | 2019 | A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks · ICML 2019 |
Mathematical optimization
nonconvex optimization |
0.1 | 1 | 2018 | Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
statistical physics · 1.2ridge regression · 0.9random projection · 0.9neural network regularizer · 0.9differentiable approximation · 0.9conditional independence · 0.9symmetric normalized graph filter analysis · 0.8graph convolutional network · 0.8knowledge distillation · 0.5gated positional self-attention · 0.5mean-field theory · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Differentiable Rank-Based Objective for Better Feature LearningabstractIn this paper, we leverage existing statistical methods to better understand feature learning from data. We tackle this by modifying the model-free variable selection method, Feature Ordering by Conditional Independence (FOCI), which is introduced in Azadkia & Chatterjee (2021). While FOCI is based on a non-parametric coefficient of conditional dependence, we introduce its parametric, differentiable approximation. With this approximate coefficient of correlation, we present a new algorithm called difFOCI, which is applicable to a wider range of machine learning problems thanks to its differentiable nature and learnable parameters. We present difFOCI in three contexts: (1) as a variable selection method with baseline comparisons to FOCI, (2) as a trainable model parametrized with a neural network, and (3) as a generic, widely applicable neural network regularizer, one that improves feature learning with better management of spurious correlations. We evaluate difFOCI on increasingly complex problems ranging from basic variable selection in toy examples to saliency map comparisons in convolutional networks. We then show how difFOCI can be incorporated in the context of fairness to facilitate classifications without relying on sensitive data. Krunoslav Lehman Pavasovic, Giulio Biroli, Levent Sagun |
ICLR | 3 |
| 2025 | An Effective Theory of Bias AmplificationabstractMachine learning models can capture and amplify biases present in data, leading to disparate test performance across social groups. To better understand, evaluate, and mitigate these biases, a deeper theoretical understanding of how model design choices and data distribution properties contribute to bias is needed. In this work, we contribute a precise analytical theory in the context of ridge regression, both with and without random projections, where the former models feedforward neural networks in a simplified regime. Our theory offers a unified and rigorous explanation of machine learning bias, providing insights into phenomena such as bias amplification and minority-group bias in various feature and parameter regimes. For example, we observe that there may be an optimal regularization penalty or training time to avoid bias amplification, and there can be differences in test error between groups that are not alleviated with increased parameterization. Importantly, our theoretical predictions align with empirical observations reported in the literature on machine learning bias. We extensively empirically validate our theory on synthetic and semi-synthetic datasets. Arjun Subramonian, Samuel J. Bell, Levent Sagun, Elvis Dohmatob |
ICLR | 3 |
| 2025 | On the Role of Speech Data in Reducing Toxicity Detection BiasabstractSamuel Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Samuel J. Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà |
NAACL (Long Papers) | 8 |
| 2024 | Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link PredictionabstractGraph neural network (GNN) link prediction is increasingly deployed in citation, collaboration, and online social networks to recommend academic literature, collaborators, and friends. While prior research has investigated the dyadic fairness of GNN link prediction, the within-group (e.g., queer women) fairness and "rich get richer" dynamics of link prediction remain underexplored. However, these aspects have significant consequences for degree and power imbalances in networks. In this paper, we shed light on how degree bias in networks affects Graph Convolutional Network (GCN) link prediction. In particular, we theoretically uncover that GCNs with a symmetric normalized graph filter have a within-group preferential attachment bias. We validate our theoretical analysis on real-world citation, collaboration, and online social networks. We further bridge GCN’s preferential attachment bias with unfairness in link prediction and propose a new within-group fairness metric. This metric quantifies disparities in link prediction scores within social groups, towards combating the amplification of degree and power disparities. Finally, we propose a simple training-time strategy to alleviate within-group unfairness, and we show that it is effective on citation, social, and credit networks. Arjun Subramonian, Levent Sagun, Yizhou Sun |
ICML | 2 |
| 2021 | ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesabstractConvolutional architectures have proven extremely successful for vision tasks. Their hard inductive biases enable sample-efficient learning, but come at the cost of a potentially lower performance ceiling. Vision Transformers (ViTs) rely on more flexible self-attention layers, and have recently outperformed CNNs for image classification. However, they require costly pre-training on large external datasets or distillation from pre-trained convolutional networks. In this paper, we ask the following question: is it possible to combine the strengths of these two architectures while avoiding their respective limitations? To this end, we introduce gated positional self-attention (GPSA), a form of positional self-attention which can be equipped with a “soft" convolutional inductive bias. We initialise the GPSA layers to mimic the locality of convolutional layers, then give each attention head the freedom to escape locality by adjusting a gating parameter regulating the attention paid to position versus content information. The resulting convolutional-like ViT architecture, ConViT, outperforms the DeiT on ImageNet, while offering a much improved sample efficiency. We further investigate the role of locality in learning by first quantifying how it is encouraged in vanilla self-attention layers, then analysing how it is escaped in GPSA layers. We conclude by presenting various ablations to better understand the success of the ConViT. Our code and models are released publicly at https://github.com/facebookresearch/convit. Stéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos, Giulio Biroli, Levent Sagun |
ICML | 6 |
| 2021 | On the interplay between data structure and loss function in classification problemsabstractOne of the central features of modern machine learning models, including deep neural networks, is their generalization ability on structured data in the over-parametrized regime. In this work, we consider an analytically solvable setup to investigate how properties of data impact learning in classification problems, and compare the results obtained for quadratic loss and logistic loss. Using methods from statistical physics, we obtain a precise asymptotic expression for the train and test errors of random feature models trained on a simple model of structured data. The input covariance is built from independent blocks allowing us to tune the saliency of low-dimensional structures and their alignment with respect to the target function.Our results show in particular that in the over-parametrized regime, the impact of data structure on both train and test error curves is greater for logistic loss than for mean-squared loss: the easier the task, the wider the gap in performance between the two losses at the advantage of the logistic. Numerical experiments on MNIST and CIFAR10 confirm our insights. Stéphane d'Ascoli, Marylou Gabrié, Levent Sagun, Giulio Biroli |
NeurIPS | 3 |
| 2020 | Triple descent and the two kinds of overfitting: where & why do they appear?abstractA recent line of research has highlighted the existence of a ``double descent'' phenomenon in deep learning, whereby increasing the number of training examples N causes the generalization error of neural networks to peak when N is of the same order as the number of parameters P. In earlier works, a similar phenomenon was shown to exist in simpler models such as linear regression, where the peak instead occurs when N is equal to the input dimension D. Since both peaks coincide with the interpolation threshold, they are often conflated in the litterature. In this paper, we show that despite their apparent similarity, these two scenarios are inherently different. In fact, both peaks can co-exist when neural networks are applied to noisy regression tasks. The relative size of the peaks is then governed by the degree of nonlinearity of the activation function. Building on recent developments in the analysis of random feature models, we provide a theoretical ground for this sample-wise triple descent. As shown previously, the nonlinear peak at N=P is a true divergence caused by the extreme sensitivity of the output function to both the noise corrupting the labels and the initialization of the random features (or the weights in neural networks). This peak survives in the absence of noise, but can be suppressed by regularization. In contrast, the linear peak at N=D is solely due to overfitting the noise in the labels, and forms earlier during training. We show that this peak is implicitly regularized by the nonlinearity, which is why it only becomes salient at high noise and is weakly affected by explicit regularization. Throughout the paper, we compare the analytical results obtained in the random feature model with the outcomes of numerical experiments involving realistic neural networks. Stéphane d'Ascoli, Levent Sagun, Giulio Biroli |
NeurIPS | 2 |
| 2019 | A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural NetworksabstractThe gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the classical central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochastic differential equation (SDE) driven by a Brownian motion. We argue that the Gaussianity assumption might fail to hold in deep learning settings and hence render the Brownian motion-based analyses inappropriate. Inspired by non-Gaussian natural phenomena, we consider the GN in a more general context and invoke the generalized CLT (GCLT), which suggests that the GN converges to a heavy-tailed $\alpha$-stable random variable. Accordingly, we propose to analyze SGD as an SDE driven by a Lévy motion. Such SDEs can incur ‘jumps’, which force the SDE transition from narrow minima to wider minima, as proven by existing metastability theory. To validate the $\alpha$-stable assumption, we conduct experiments on common deep learning scenarios and show that in all settings, the GN is highly non-Gaussian and admits heavy-tails. We investigate the tail behavior in varying network architectures and sizes, loss functions, and datasets. Our results open up a different perspective and shed more light on the belief that SGD prefers wide minima. Umut Simsekli, Levent Sagun, Mert Gürbüzbalaban |
ICML | 2 |
| 2019 | Finding the Needle in the Haystack with Convolutions: on the benefits of architectural biasabstractDespite the phenomenal success of deep neural networks in a broad range of learning tasks, there is a lack of theory to understand the way they work. In particular, Convolutional Neural Networks (CNNs) are known to perform much better than Fully-Connected Networks (FCNs) on spatially structured data: the architectural structure of CNNs benefits from prior knowledge on the features of the data, for instance their translation invariance. The aim of this work is to understand this fact through the lens of dynamics in the loss landscape. We introduce a method that maps a CNN to its equivalent FCN (denoted as eFCN). Such an embedding enables the comparison of CNN and FCN training dynamics directly in the FCN space. We use this method to test a new training protocol, which consists in training a CNN, embedding it to FCN space at a certain ``relax time'', then resuming the training in FCN space. We observe that for all relax times, the deviation from the CNN subspace is small, and the final performance reached by the eFCN is higher than that reachable by a standard FCN of same architecture. More surprisingly, for some intermediate relax times, the eFCN outperforms the CNN it stemmed, by combining the prior information of the CNN and the expressivity of the FCN in a complementary way. The practical interest of our protocol is limited by the very large size of the highly sparse eFCN. However, it offers interesting insights into the persistence of architectural bias under stochastic gradient dynamics. It shows the existence of some rare basins in the FCN loss landscape associated with very good generalization. These can only be accessed thanks to the CNN prior, which helps navigate the landscape during the early stages of optimization. Stéphane d'Ascoli, Levent Sagun, Giulio Biroli, Joan Bruna |
NeurIPS | 2 |
| 2018 | Comparing Dynamics: Deep Neural Networks versus Glassy SystemsabstractWe analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are the complexity of the loss-landscape and of the dynamics within it, and to what extent DNNs share similarities with glassy systems. Our findings, obtained for different architectures and data-sets, suggest that during the training process the dynamics slows down because of an increasingly large number of flat directions. At large times, when the loss is approaching zero, the system diffuses at the bottom of the landscape. Despite some similarities with the dynamics of mean-field glassy systems, in particular, the absence of barrier crossing, we find distinctive dynamical behaviors in the two cases, thus showing that the statistical properties of the corresponding loss and energy landscapes are different. In contrast, when the network is under-parametrized we observe a typical glassy behavior, thus suggesting the existence of different phases depending on whether the network is under-parametrized or over-parametrized. Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gérard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, Giulio Biroli |
ICML | 2 |
| 2017 | Early predictability of asylum court decisionsabstractIn the United States, foreign nationals who fear persecution in their home country can apply for asylum under the Refugee Act of 1980. Over the past decade, legal scholarship has uncovered significant disparities in asylum adjudication by judge, by region of the United States in which the application is filed, and by the applicant's nationality. These disparities raise concerns about whether applicants are receiving equal treatment under the law. Using machine learning to predict judges' decisions, we document another concern that may violate our notions of justice: we are able to predict the final outcome of a case with 80% accuracy at the time the case opens using only information on the identity of the judge handling the case and the applicant's nationality. Moreover, there is significant variation in the degree of predictability of judges at the time the case is assigned to a judge. We show that highly predictable judges tend to hold fewer hearing sessions before making their decision, which raises the possibility that early predictability is due to judges deciding based on snap or predetermined judgments rather than taking into account the specifics of each case. Early prediction of a case with 80% accuracy could assist asylum seekers in their applications. Matthew Dunn, Levent Sagun, Hale Sirin |
ICAIL | 2 |
| 2017 | Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, Riccardo Zecchina |
ICLR (Poster) | 8 |