EDBT 2026 Demo / reviewers in the wild / expert
Alexander B. Atanasov
dblp:207/8552
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-3338-0324ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Learning theory · 39% Deep learning architectures and training · 34% Optimization for machine learning · 17% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
neural network theory |
2.1 | 3 | 2025 | The Optimization Landscape of SGD Across the Feature Learning Strength · ICLR 2025 The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes · ICLR 2023 Neural Networks as Kernel Learners: The Silent Alignment Effect · ICLR 2022 |
Machine learning › Deep learning architectures and training › scaling laws
compute-optimal scaling |
1.6 | 2 | 2025 | How Feature Learning Can Improve Neural Scaling Laws · ICLR 2025 A Dynamical Model of Neural Scaling Laws · ICML 2024 |
Machine learning › Deep learning architectures and training
scaling laws |
1.6 | 2 | 2025 | How Feature Learning Can Improve Neural Scaling Laws · ICLR 2025 A Dynamical Model of Neural Scaling Laws · ICML 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
1.4 | 2 | 2024 | A Dynamical Model of Neural Scaling Laws · ICML 2024 Feature-Learning Networks Are Consistent Across Widths At Realistic Scales · NeurIPS 2023 |
Machine learning › Learning theory › model selection
cross-validation |
0.9 | 1 | 2025 | Risk and cross validation in ridge regression with correlated samples · ICML 2025 |
Machine learning › Optimization for machine learning › learning rate
learning rate scaling |
0.9 | 1 | 2025 | The Optimization Landscape of SGD Across the Feature Learning Strength · ICLR 2025 |
Machine learning › Optimization for machine learning
learning rate schedule |
0.9 | 1 | 2025 | The Optimization Landscape of SGD Across the Feature Learning Strength · ICLR 2025 |
Machine learning › Learning theory
model selection |
0.9 | 1 | 2025 | Risk and cross validation in ridge regression with correlated samples · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression |
0.9 | 1 | 2025 | Risk and cross validation in ridge regression with correlated samples · ICML 2025 |
Machine learning › Learning theory
statistical learning theory |
0.9 | 1 | 2025 | Risk and cross validation in ridge regression with correlated samples · ICML 2025 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.9 | 1 | 2025 | The Optimization Landscape of SGD Across the Feature Learning Strength · ICLR 2025 |
Machine learning › Learning theory › neural network theory
infinite-width limit |
0.7 | 1 | 2023 | Feature-Learning Networks Are Consistent Across Widths At Realistic Scales · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › training dynamics
lazy training regime |
0.7 | 1 | 2023 | The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes · ICLR 2023 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.7 | 1 | 2023 | Feature-Learning Networks Are Consistent Across Widths At Realistic Scales · NeurIPS 2023 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning |
0.6 | 1 | 2022 | Neural Networks as Kernel Learners: The Silent Alignment Effect · ICLR 2022 |
Machine learning › Kernel, tree and ensemble methods
model ensemble |
0.2 | 1 | 2023 | Feature-Learning Networks Are Consistent Across Widths At Realistic Scales · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
neural tangent kernel · 1.4stochastic gradient descent · 0.9random matrix theory · 0.9free probability · 0.9dynamical mean field theory · 0.9random feature model · 0.8gradient descent · 0.8spectral analysis · 0.7kernel theory · 0.7kernel methods · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Optimization Landscape of SGD Across the Feature Learning StrengthabstractWe consider neural networks (NNs) where the final layer is down-scaled by a fixed hyperparameter $\gamma$.
Recent work has identified $\gamma$ as controlling the strength of feature learning.
As $\gamma$ increases, network evolution changes from "lazy" kernel dynamics to "rich" feature-learning dynamics, with a host of associated benefits including improved performance on common tasks.
In this work, we conduct a thorough empirical investigation of the effect of scaling $\gamma$ across a variety of models and datasets in the online training setting.
We first examine the interaction of $\gamma$ with the learning rate $\eta$, identifying several scaling regimes in the $\gamma$-$\eta$ plane which we explain theoretically using a simple model.
We find that the optimal learning rate $\eta^*$ scales non-trivially with $\gamma$. In particular, $\eta^* \propto \gamma^2$ when $\gamma \ll 1$ and $\eta^* \propto \gamma^{2/L}$ when $\gamma \gg 1$ for a feed-forward network of depth $L$.
Using this optimal learning rate scaling, we proceed with an empirical study of the under-explored ``ultra-rich'' $\gamma \gg 1$ regime.
We find that networks in this regime display characteristic loss curves, starting with a long plateau followed by a drop-off, sometimes followed by one or more additional staircase steps.
We find networks of different large $\gamma$ values optimize along similar trajectories up to a reparameterization of time.
We further find that optimal online performance is often found at large $\gamma$ and could be missed if this hyperparameter is not tuned.
Our findings indicate that analytical study of the large-$\gamma$ limit may yield useful insights into the dynamics of representation learning in performant models. Alexander B. Atanasov, Alexandru Meterez, James B. Simon, Cengiz Pehlevan |
ICLR | 1 |
| 2025 | How Feature Learning Can Improve Neural Scaling LawsabstractWe develop a simple solvable model of neural scaling laws beyond the kernel limit. Theoretical analysis of this model predicts the performance scaling predictions with model size, training time and total amount of available data. From the scaling analysis we identify three relevant regimes: hard tasks, easy tasks, and super easy tasks. For easy and super-easy target functions, which are in the Hilbert space (RKHS) of the initial infinite-width neural tangent kernel (NTK), there is no change in the scaling exponents between feature learning models and models in the kernel regime. For hard tasks, which we define as tasks outside of the RKHS of the initial NTK, we show analytically and empirically that feature learning can improve the scaling with training time and compute, approximately doubling the exponent for very hard tasks. This leads to a new compute optimal scaling law for hard tasks in the feature learning regime. We support our finding that feature learning improves the scaling law for hard tasks with experiments of nonlinear MLPs fitting functions with power-law Fourier spectra on the circle and CNNs learning vision tasks. Blake Bordelon, Alexander B. Atanasov, Cengiz Pehlevan |
ICLR | 2 |
| 2025 | Risk and cross validation in ridge regression with correlated samplesabstractRecent years have seen substantial advances in our understanding of high-dimensional ridge regression, but existing theories assume that training examples are independent. By leveraging techniques from random matrix theory and free probability, we provide sharp asymptotics for the in- and out-of-sample risks of ridge regression when the data points have arbitrary correlations. We demonstrate that in this setting, the generalized cross validation estimator (GCV) fails to correctly predict the out-of-sample risk. However, in the case where the noise residuals have the same correlations as the data points, one can modify the GCV to yield an efficiently-computable unbiased estimator that concentrates in the high-dimensional limit, which we dub CorrGCV. We further extend our asymptotic analysis to the case where the test point has nontrivial correlations with the training set, a setting often encountered in time series forecasting. Assuming knowledge of the correlation structure of the time series, this again yields an extension of the GCV estimator, and sharply characterizes the degree to which such test points yield an overly optimistic prediction of long-time risk. We validate the predictions of our theory across a variety of high dimensional data. Alexander B. Atanasov, Jacob A. Zavatone-Veth, Cengiz Pehlevan |
ICML | 1 |
| 2024 | A Dynamical Model of Neural Scaling LawsabstractOn a variety of tasks, the performance of neural networks predictably improves with training time, dataset size and model size across many orders of magnitude. This phenomenon is known as a neural scaling law. Of fundamental importance is the compute-optimal scaling law, which reports the performance as a function of units of compute when choosing model sizes optimally. We analyze a random feature model trained with gradient descent as a solvable model of network training and generalization. This reproduces many observations about neural scaling laws. First, our model makes a prediction about why the scaling of performance with training time and with model size have different power law exponents. Consequently, the theory predicts an asymmetric compute-optimal scaling rule where the number of training steps are increased faster than model parameters, consistent with recent empirical observations. Second, it has been observed that early in training, networks converge to their infinite-width dynamics at a rate $1/\text{width}$ but at late time exhibit a rate $\text{width}^{-c}$, where $c$ depends on the structure of the architecture and task. We show that our model exhibits this behavior. Lastly, our theory shows how the gap between training and test loss can gradually build up over time due to repeated reuse of data. Blake Bordelon, Alexander B. Atanasov, Cengiz Pehlevan |
ICML | 2 |
| 2023 | The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes
Alexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz Pehlevan |
ICLR | 1 |
| 2023 | Feature-Learning Networks Are Consistent Across Widths At Realistic ScalesabstractWe study the effect of width on the dynamics of feature-learning neural networks across a variety of architectures and datasets. Early in training, wide neural networks trained on online data have not only identical loss curves but also agree in their point-wise test predictions throughout training. For simple tasks such as CIFAR-5m this holds throughout training for networks of realistic widths. We also show that structural properties of the models, including internal representations, preactivation distributions, edge of stability phenomena, and large learning rate effects are consistent across large widths. This motivates the hypothesis that phenomena seen in realistic models can be captured by infinite-width, feature-learning limits. For harder tasks (such as ImageNet and language modeling), and later training times, finite-width deviations grow systematically. Two distinct effects cause these deviations across widths. First, the network output has an initialization-dependent variance scaling inversely with width, which can be removed by ensembling networks. We observe, however, that ensembles of narrower networks perform worse than a single wide network. We call this the bias of narrower width. We conclude with a spectral perspective on the origin of this finite-width bias. Nikhil Vyas 0001, Alexander B. Atanasov, Blake Bordelon, Depen Morwani, Sabarish Sainathan, Cengiz Pehlevan |
NeurIPS | 2 |
| 2022 | Neural Networks as Kernel Learners: The Silent Alignment Effect
Alexander B. Atanasov, Blake Bordelon, Cengiz Pehlevan |
ICLR | 1 |