Berfin Simsek

dblp:244/2455 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 45% Learning theory · 20% Kernel, tree and ensemble methods · 13%

Topics — the 24 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
loss landscape
2.232025
Flat Channels to Infinity in Neural Loss Landscapes · NeurIPS 2025
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding · ICLR 2025
Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances · ICML 2021
Machine learning › Optimization for machine learning
gradient flow
0.912025
Flat Channels to Infinity in Neural Loss Landscapes · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.922020
Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020
Implicit Regularization of Random Feature Models · ICML 2020
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression
0.922020
Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020
Implicit Regularization of Random Feature Models · ICML 2020
Machine learning › Graph learning
network embedding
0.912025
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding · ICLR 2025
Machine learning › Deep learning architectures and training › training dynamics
optimization dynamics
0.912025
Flat Channels to Infinity in Neural Loss Landscapes · NeurIPS 2025
Machine learning › Deep learning architectures and training › ReLU networks
shallow ReLU network
0.912025
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding · ICLR 2025
Machine learning › Representation and self-supervised learning
associative memory
0.812024
Learning Associative Memories with Gradient Descent · ICML 2024
Machine learning › Deep learning architectures and training › training dynamics
gradient descent dynamics
0.812024
Learning Associative Memories with Gradient Descent · ICML 2024
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability
0.812024
Expand-and-Cluster: Parameter Recovery of Neural Networks · ICML 2024
Machine learning › Learning theory › neural network theory
neural network analysis
0.812024
Expand-and-Cluster: Parameter Recovery of Neural Networks · ICML 2024
Machine learning › Learning theory › neural network theory
over-parameterized regime
0.812024
Learning Associative Memories with Gradient Descent · ICML 2024
Machine learning › Learning theory › statistical estimation
parameter recovery
0.812024
Expand-and-Cluster: Parameter Recovery of Neural Networks · ICML 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Learning Associative Memories with Gradient Descent · ICML 2024
Machine learning › Learning theory › approximation theory
neural network approximation
0.712023
Should Under-parameterized Student Networks Copy or Average Teacher Weights? · NeurIPS 2023
Machine learning › Deep learning architectures and training
teacher-student framework
0.712023
Should Under-parameterized Student Networks Copy or Average Teacher Weights? · NeurIPS 2023
Machine learning › Deep learning architectures and training
overparameterized neural network
0.512021
Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances · ICML 2021
Machine learning › Learning theory
generalization bounds
0.412020
Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020
Machine learning › Optimization for machine learning
implicit regularization
0.412020
Implicit Regularization of Random Feature Models · ICML 2020
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features
0.412020
Implicit Regularization of Random Feature Models · ICML 2020
Machine learning › Deep learning architectures and training
risk prediction
0.412020
Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020
Machine learning › Deep learning architectures and training
neural network expressivity
0.312025
Flat Channels to Infinity in Neural Loss Landscapes · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.212024
Learning Associative Memories with Gradient Descent · ICML 2024
Machine learning › Optimization for machine learning
non-convex optimization
0.112021
Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances · ICML 2021

Methods — techniques the papers use, named apart from their topics

gradient descent · 1.6random matrix theory · 0.9gradient flow analysis · 0.9directional stationary point analysis · 0.9adam · 0.9SGD · 0.9particle system analysis · 0.8overparameterized student networks · 0.8clustering · 0.8constrained optimization · 0.7
YearPublicationVenuePosition
2025 Learning Gaussian Multi-Index Models with Gradient Flow: Time Complexity and Directional Convergence
abstract
This work focuses on the gradient flow dynamics of a neural network model that uses correlation loss to approximate a multi-index function on high-dimensional standard Gaussian data. Specifically, the multi-index function we consider is a sum of neurons $f^*(x) = \sum_{j=1}^k \sigma^*(v_j^T x)$ where $v_1, ..., v_k$ are unit vectors, and $\sigma^*$ lacks the first and second Hermite polynomials in its Hermite expansion. It is known that, for the single-index case ($k=1$), overcoming the search phase requires polynomial time complexity. We first generalize this result to multi-index functions characterized by vectors in arbitrary directions. After the search phase, it is not clear whether the network neurons converge to the index vectors, or get stuck at a sub-optimal solution. When the index vectors are orthogonal, we give a complete characterization of the fixed points and prove that neurons converge to the nearest index vectors. Therefore, using $n \asymp k \log k$ neurons ensures finding the full set of index vectors with gradient flow with high probability over random initialization. When $v_i^T v_j = \beta \geq 0$ for all $i \neq j$, we prove the existence of a sharp threshold $\beta_c = c/(c+k)$ at which the fixed point that computes the average of the index vectors transitions from a saddle point to a minimum. Numerical simulations show that using a correlation loss and a mild overparameterization suffices to learn all of the index vectors when they are nearly orthogonal, however, the correlation loss fails when the dot product between the index vectors exceeds a certain threshold.
Berfin Simsek, Amire Bendjeddou, Daniel Hsu 0001
AISTATS1
2025 Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
abstract
In this paper, we study the loss landscape of one-hidden-layer neural networks with ReLU-like activation functions trained with the empirical squared loss using gradient descent (GD). We identify the stationary points of such networks, which significantly slow down loss decrease during training. To capture such points while accounting for the non-differentiability of the loss, the stationary points that we study are directional stationary points, rather than other notions like Clarke stationary points. We show that, if a stationary point does not contain "escape neurons", which are defined with first-order conditions, it must be a local minimum. Moreover, for the scalar-output case, the presence of an escape neuron guarantees that the stationary point is not a local minimum. Our results refine the description of the *saddle-to-saddle* training process starting from infinitesimally small (vanishing) initialization for shallow ReLU-like networks: By precluding the saddle escape types that previous works did not rule out, we advance one step closer to a complete picture of the entire dynamics. Moreover, we are also able to fully discuss how network embedding, which is to instantiate a narrower network with a wider network, reshapes the stationary points.
Frank Zhengqing Wu, Berfin Simsek, François Ged
ICLR2
2025 Flat Channels to Infinity in Neural Loss Landscapes
abstract
The loss landscapes of neural networks contain minima and saddle points that may be connected in flat regions or appear in isolation. We identify and characterize a special structure in the loss landscape: channels along which the loss decreases extremely slowly, while the output weights of at least two neurons, $a_i$ and $a_j$, diverge to $\pm$infinity, and their input weight vectors, $\mathbf{w_i}$ and $\mathbf{w_j}$, become equal to each other. At convergence, the two neurons implement a gated linear unit: $a_i\sigma(\mathbf{w_i} \cdot \mathbf{x}) + a_j\sigma(\mathbf{w_j} \cdot \mathbf{x}) \rightarrow c \sigma(\mathbf{w} \cdot \mathbf{x}) + (\mathbf{v} \cdot \mathbf{x}) \sigma'(\mathbf{w} \cdot \mathbf{x})$. Geometrically, these channels to infinity are asymptotically parallel to symmetry-induced lines of critical points. Gradient flow solvers, and related optimization methods like SGD or ADAM, reach the channels with high probability in diverse regression settings, but without careful inspection they look like flat local minima with finite parameter values. Our characterization provides a comprehensive picture of this quasi-flat region in terms of gradient dynamics, geometry, and functional interpretation. The emergence of gated linear units at the end of the channels highlights a surprising aspect of the computational capabilities of fully connected layers.
Flavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner, Johanni Brea
NeurIPS3
2024 Learning Associative Memories with Gradient Descent
abstract
This work focuses on the training dynamics of one associative memory module storing outer products of token embeddings. We reduce this problem to the study of a system of particles, which interact according to properties of the data distribution and correlations between embeddings. Through theory and experiments, we provide several insights. In overparameterized regimes, we obtain logarithmic growth of the ``classification margins.'' Yet, we show that imbalance in token frequencies and memory interferences due to correlated embeddings lead to oscillatory transitory regimes. The oscillations are more pronounced with large step sizes, which can create benign loss spikes, although these learning rates speed up the dynamics and accelerate the asymptotic convergence. We also find that underparameterized regimes lead to suboptimal memorization schemes. Finally, we assess the validity of our findings on small Transformer models.
Vivien Cabannes, Berfin Simsek, Alberto Bietti
ICML2
2024 Expand-and-Cluster: Parameter Recovery of Neural Networks
abstract
Can we identify the weights of a neural network by probing its input-output mapping? At first glance, this problem seems to have many solutions because of permutation, overparameterisation and activation function symmetries. Yet, we show that the incoming weight vector of each neuron is identifiable up to sign or scaling, depending on the activation function. Our novel method ’Expand-and-Cluster’ can identify layer sizes and weights of a target network for all commonly used activation functions. Expand-and-Cluster consists of two phases: (i) to relax the non-convex optimisation problem, we train multiple overparameterised student networks to best imitate the target function; (ii) to reverse engineer the target network’s weights, we employ an ad-hoc clustering procedure that reveals the learnt weight vectors shared between students – these correspond to the target weight vectors. We demonstrate successful weights and size recovery of trained shallow and deep networks with less than 10% overhead in the layer size and describe an ’ease-of-identifiability’ axis by analysing 150 synthetic problems of variable difficulty.
Flavio Martinelli, Berfin Simsek, Wulfram Gerstner, Johanni Brea
ICML2
2023 Should Under-parameterized Student Networks Copy or Average Teacher Weights?
abstract
Any continuous function $f^*$ can be approximated arbitrarily well by a neural network with sufficiently many neurons $k$. We consider the case when $f^*$ itself is a neural network with one hidden layer and $k$ neurons. Approximating $f^*$ with a neural network with $n< k$ neurons can thus be seen as fitting an under-parameterized "student" network with $n$ neurons to a "teacher" network with $k$ neurons. As the student has fewer neurons than the teacher, it is unclear, whether each of the $n$ student neurons should copy one of the teacher neurons or rather average a group of teacher neurons. For shallow neural networks with erf activation function and for the standard Gaussian input distribution, we prove that "copy-average" configurations are critical points if the teacher's incoming vectors are orthonormal and its outgoing weights are unitary. Moreover, the optimum among such configurations is reached when $n-1$ student neurons each copy one teacher neuron and the $n$-th student neuron averages the remaining $k-n+1$ teacher neurons. For the student network with $n=1$ neuron, we provide additionally a closed-form solution of the non-trivial critical point(s) for commonly used activation functions through solving an equivalent constrained optimization problem. Empirically, we find for the erf activation function that gradient flow converges either to the optimal copy-average critical point or to another point where each student neuron approximately copies a different teacher neuron. Finally, we find similar results for the ReLU activation function, suggesting that the optimal solution of underparameterized networks has a universal structure.
Berfin Simsek, Amire Bendjeddou, Wulfram Gerstner, Johanni Brea
NeurIPS1
2021 Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances
abstract
We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width $ r^*+ h =: m $ we explicitly describe the manifold of global minima: it consists of $ T(r^*, m) $ affine subspaces of dimension at least $ h $ that are connected to one another. For a network of width $m$, we identify the number $G(r,m)$ of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width $r Cite this Paper BibTeX @InProceedings{pmlr-v139-simsek21a, title = {Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances}, author = {Simsek, Berfin and Ged, Fran{\c{c}}ois and Jacot, Arthur and Spadaro, Francesco and Hongler, Clement and Gerstner, Wulfram and Brea, Johanni}, booktitle = {Proceedings of the 38th International Conference on Machine Learning}, pages = {9722--9732}, year = {2021}, editor = {Meila, Marina and Zhang, Tong}, volume = {139}, series = {Proceedings of Machine Learning Research}, month = {18--24 Jul}, publisher = {PMLR}, pdf = {http://proceedings.mlr.press/v139/simsek21a/simsek21a.pdf}, url = {https://proceedings.mlr.press/v139/simsek21a.html}, abstract = {We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width $ r^*+ h =: m $ we explicitly describe the manifold of global minima: it consists of $ T(r^*, m) $ affine subspaces of dimension at least $ h $ that are connected to one another. For a network of width $m$, we identify the number $G(r,m)$ of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width $r Copy to Clipboard Download Endnote %0 Conference Paper %T Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances %A Berfin Simsek %A François Ged %A Arthur Jacot %A Francesco Spadaro %A Clement Hongler %A Wulfram Gerstner %A Johanni Brea %B Proceedings of the 38th International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2021 %E Marina Meila %E Tong Zhang %F pmlr-v139-simsek21a %I PMLR %P 9722--9732 %U https://proceedings.mlr.press/v139/simsek21a.html %V 139 %X We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width $ r^*+ h =: m $ we explicitly describe the manifold of global minima: it consists of $ T(r^*, m) $ affine subspaces of dimension at least $ h $ that are connected to one another. For a network of width $m$, we identify the number $G(r,m)$ of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width $r Copy to Clipboard Download APA Simsek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W. & Brea, J.. (2021). Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances. Proceedings of the 38th International Conference on Machine Learning, in Proceedings of Machine Learning Research 139:9722-9732 Available from https://proceedings.mlr.press/v139/simsek21a.html. Copy to Clipboard Download Related Material Download PDF Supplementary PDF This site last compiled Sun, 05 Jul 2026 14:53:09 +0000 Github Account Copyright © The authors and PMLR 2026. MLResearchPress
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clément Hongler, Wulfram Gerstner, Johanni Brea
ICML1
2020 Implicit Regularization of Random Feature Models
abstract
Random Features (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with $P$ features, $N$ data points, and a ridge $\lambda$, we show that the average (i.e. expected) RF predictor is close to a KRR predictor with an \emph{effective ridge} $\tilde{\lambda}$. We show that $\tilde{\lambda} > \lambda$ and $\tilde{\lambda} \searrow \lambda$ monotonically as $P$ grows, thus revealing the \emph{implicit regularization effect} of finite RF sampling. We then compare the risk (i.e. test error) of the $\tilde{\lambda}$-KRR predictor with the average risk of the $\lambda$-RF predictor and obtain a precise and explicit bound on their difference. Finally, we empirically find an extremely good agreement between the test errors of the average $\lambda$-RF predictor and $\tilde{\lambda}$-KRR predictor.
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel
ICML2
2020 Kernel Alignment Risk Estimator: Risk Prediction from Training Data
abstract
We study the risk (i.e. generalization error) of Kernel Ridge Regression (KRR) for a kernel $K$ with ridge $\lambda>0$ and i.i.d. observations. For this, we introduce two objects: the Signal Capture Threshold (SCT) and the Kernel Alignment Risk Estimator (KARE). The SCT $\vartheta_{K,\lambda}$ is a function of the data distribution: it can be used to identify the components of the data that the KRR predictor captures, and to approximate the (expected) KRR risk. This then leads to a KRR risk approximation by the KARE $\rho_{K, \lambda}$, an explicit function of the training data, agnostic of the true data distribution. We phrase the regression problem in a functional setting. The key results then follow from a finite-size adaptation of the resolvent method for general Wishart random matrices. Under a natural universality assumption (that the KRR moments depend asymptotically on the first two moments of the observations) we capture the mean and variance of the KRR predictor. We numerically investigate our findings on the Higgs and MNIST datasets for various classical kernels: the KARE gives an excellent approximation of the risk. This supports our universality hypothesis. Using the KARE, one can compare choices of Kernels and hyperparameters directly from the training set. The KARE thus provides a promising data-dependent procedure to select Kernels that generalize well.
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel
NeurIPS2