Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Raphaël Berthier

dblp:205/3030 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 44% Optimization for machine learning · 21% Learning theory · 20%
Theoretical computer science
2 papers
Mathematical optimization · 85% Distributed computing theory · 15%

Topics — the 24 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › attention mechanism › attention module
attention layer
0.912025
Attention layers provably solve single-location regression · ICLR 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Attention layers provably solve single-location regression · ICLR 2025
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.912025
Attention layers provably solve single-location regression · ICLR 2025
Machine learning › Deep learning architectures and training › transformer
transformer theory
0.912025
Attention layers provably solve single-location regression · ICLR 2025
Machine learning › Learning theory › generalization
generalization on the unseen
0.812024
On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions · ICML 2024
Machine learning › Learning theory › generalization
generalization theory
0.812024
On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions · ICML 2024
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.812024
On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions · ICML 2024
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features
0.812024
On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions · ICML 2024
Machine learning › Learning theory
statistical learning theory
0.722025
Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear Model · NeurIPS 2020
Attention layers provably solve single-location regression · ICLR 2025
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks
0.712023
Incremental Learning in Diagonal Linear Networks · J. Mach. Learn. Res. 2023
Machine learning › Optimization for machine learning
gradient flow
0.712023
Incremental Learning in Diagonal Linear Networks · J. Mach. Learn. Res. 2023
Machine learning › Optimization for machine learning › convergence analysis
gradient flow convergence
0.712023
Leveraging the two-timescale regime to demonstrate convergence of neural networks · NeurIPS 2023
Machine learning › Optimization for machine learning
implicit regularization
0.712023
Incremental Learning in Diagonal Linear Networks · J. Mach. Learn. Res. 2023
Machine learning › Optimization for machine learning
non-convex optimization
0.712023
Leveraging the two-timescale regime to demonstrate convergence of neural networks · NeurIPS 2023
Machine learning › Deep learning architectures and training › feedforward neural network
shallow neural networks
0.712023
Leveraging the two-timescale regime to demonstrate convergence of neural networks · NeurIPS 2023
Machine learning › Deep learning architectures and training
training dynamics
0.712023
Leveraging the two-timescale regime to demonstrate convergence of neural networks · NeurIPS 2023
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent
0.622021
Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms · NeurIPS 2021
Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear Model · NeurIPS 2020
Mathematical optimization
stochastic optimization
0.622021
Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms · NeurIPS 2021
Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear Model · NeurIPS 2020
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
accelerated gradient methods
0.512021
Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms · NeurIPS 2021
Mathematical optimization › continuous optimization › convex optimization
first-order methods
0.512021
Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms · NeurIPS 2021
Distributed computing theory › information dissemination
gossip protocols
0.512021
Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms · NeurIPS 2021
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization › accelerated gradient methods
nesterov acceleration
0.512021
Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms · NeurIPS 2021
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.412020
Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear Model · NeurIPS 2020
Machine learning › Learning theory › statistical learning theory › bayesian learning theory
bayes optimality
0.312025
Attention layers provably solve single-location regression · ICLR 2025

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 1.5gradient flow · 1.3gradient descent · 0.9asymptotic analysis · 0.9spectral dimension analysis · 0.9reproducing kernel hilbert space · 0.9transformer · 0.8random feature model · 0.8two-timescale analysis · 0.7ordinary differential equations · 0.5convex optimization · 0.5continuous-time analysis · 0.5
YearPublicationVenuePosition
2025 Attention layers provably solve single-location regression
abstract
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal linear representations. To address this gap, we introduce the single-location regression task, where only one token in a sequence determines the output, and its position is a latent random variable, retrievable via a linear projection of the input. To solve this task, we propose a dedicated predictor, which turns out to be a simplified version of a non-linear self-attention layer. We study its theoretical properties, by showing its asymptotic Bayes optimality and analyzing its training dynamics. In particular, despite the non-convex nature of the problem, the predictor effectively learns the underlying structure. This work highlights the capacity of attention mechanisms to handle sparse token information and internal linear structures.
Pierre Marion, Raphaël Berthier, Gérard Biau, Claire Boyer
ICLR2
2024 On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions
abstract
We investigate the out-of-domain generalization of random feature (RF) models and Transformers. We first prove that in the ‘generalization on the unseen (GOTU)’ setting, where training data is fully seen in some part of the domain but testing is made on another part, and for RF models in the small feature regime, the convergence takes place to interpolators of minimal degree as in the Boolean case (Abbe et al., 2023). We then consider the sparse target regime and explain how this regime relates to the small feature regime, but with a different regularization term that can alter the picture in the non-Boolean case. We show two different outcomes for the sparse regime with q-ary data tokens: (1) if the data is embedded with roots of unities, then a min-degree interpolator is learned like in the Boolean case for RF models, (2) if the data is not embedded as such, e.g., simply as integers, then RF models and Transformers may not learn minimal degree interpolators. This shows that the Boolean setting and its roots of unities generalization are special cases where the minimal degree interpolator offers a rare characterization of how learning takes place. For more general integer and real-valued settings, a more nuanced picture remains to be fully characterized.
Denys Pushkin, Raphaël Berthier, Emmanuel Abbe
ICML2
2023 Leveraging the two-timescale regime to demonstrate convergence of neural networks
abstract
We study the training dynamics of shallow neural networks, in a two-timescale regime in which the stepsizes for the inner layer are much smaller than those for the outer layer. In this regime, we prove convergence of the gradient flow to a global optimum of the non-convex optimization problem in a simple univariate setting. The number of neurons need not be asymptotically large for our result to hold, distinguishing our result from popular recent approaches such as the neural tangent kernel or mean-field regimes. Experimental illustration is provided, showing that the stochastic gradient descent behaves according to our description of the gradient flow and thus converges to a global optimum in the two-timescale regime, but can fail outside of this regime.
Pierre Marion, Raphaël Berthier
NeurIPS2
2023 Incremental Learning in Diagonal Linear Networks
abstract
Diagonal linear networks (DLNs) are a toy simplification of artificial neural networks; they consist in a quadratic reparametrization of linear regression inducing a sparse implicit regularization. In this paper, we describe the trajectory of the gradient flow of DLNs in the limit of small initialization. We show that incremental learning is effectively performed in the limit: coordinates are successively activated, while the iterate is the minimizer of the loss constrained to have support on the active coordinates only. This shows that the sparse implicit regularization of DLNs decreases with time. This work is restricted to the underparametrized regime with anti-correlated features for technical reasons.
Raphaël Berthier
J. Mach. Learn. Res.1
2021 Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms
abstract
We introduce the ``continuized'' Nesterov acceleration, a close variant of Nesterov acceleration whose variables are indexed by a continuous time parameter. The two variables continuously mix following a linear ordinary differential equation and take gradient steps at random times. This continuized variant benefits from the best of the continuous and the discrete frameworks: as a continuous process, one can use differential calculus to analyze convergence and obtain analytical expressions for the parameters; but a discretization of the continuized process can be computed exactly with convergence rates similar to those of Nesterov original acceleration. We show that the discretization has the same structure as Nesterov acceleration, but with random parameters. We provide continuized Nesterov acceleration under deterministic as well as stochastic gradients, with either additive or multiplicative noise. Finally, using our continuized framework and expressing the gossip averaging problem as the stochastic minimization of a certain energy function, we provide the first rigorous acceleration of asynchronous gossip algorithms.
Mathieu Even, Raphaël Berthier, Francis R. Bach, Nicolas Flammarion, Hadrien Hendrikx, Pierre Gaillard, Laurent Massoulié, Adrien B. Taylor
NeurIPS2
2020 Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear Model
abstract
In the context of statistical supervised learning, the noiseless linear model assumes that there exists a deterministic linear relation $Y = \langle \theta_*, \Phi(U) \rangle$ between the random output $Y$ and the random feature vector $\Phi(U)$, a potentially non-linear transformation of the inputs~$U$. We analyze the convergence of single-pass, fixed step-size stochastic gradient descent on the least-square risk under this model. The convergence of the iterates to the optimum $\theta_*$ and the decay of the generalization error follow polynomial convergence rates with exponents that both depend on the regularities of the optimum $\theta_*$ and of the feature vectors $\Phi(U)$. We interpret our result in the reproducing kernel Hilbert space framework. As a special case, we analyze an online algorithm for estimating a real function on the unit hypercube from the noiseless observation of its value at randomly sampled points; the convergence depends on the Sobolev smoothness of the function and of a chosen kernel. Finally, we apply our analysis beyond the supervised learning setting to obtain convergence rates for the averaging process (a.k.a. gossip algorithm) on a graph depending on its spectral dimension.
Raphaël Berthier, Francis R. Bach, Pierre Gaillard
NeurIPS1