Vittorio Erba

dblp:243/3504 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 64% Deep learning architectures and training · 32% Graph learning · 5%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Bayes optimal learning of attention-indexed models · NeurIPS 2025
Machine learning › Learning theory
empirical risk minimization
0.912025
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025
Machine learning › Learning theory
generalization
0.912025
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025
Machine learning › Learning theory
generalization error
0.912025
Bayes optimal learning of attention-indexed models · NeurIPS 2025
Machine learning › Learning theory › high-dimensional statistics
high-dimensional asymptotics
0.912025
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025
Machine learning › Deep learning architectures and training
overparameterized neural network
0.912025
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025
Machine learning › Graph learning › graph neural network › message passing
approximate message passing
0.312025
Bayes optimal learning of attention-indexed models · NeurIPS 2025
Mathematical optimization › continuous optimization
convex optimization
0.312025
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025
Mathematical optimization › continuous optimization › convex optimization › norm optimization
nuclear norm minimization
0.312025
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

spin-glass methods · 1.7matrix factorization · 1.7convex optimization · 1.7statistical mechanics · 0.9random matrix theory · 0.9approximate message passing · 0.9
YearPublicationVenuePosition
2025 Bayes optimal learning of attention-indexed models
abstract
We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures.
Fabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka Zdeborová
NeurIPS3
2025 The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
abstract
We study the high-dimensional asymptotics of empirical risk minimization (ERM) in over-parametrized two-layer neural networks with quadratic activations trained on synthetic data. We derive sharp asymptotics for both training and test errors by mapping the $\ell_2$-regularized learning problem to a convex matrix sensing task with nuclear norm penalization. This reveals that capacity control in such networks emerges from a low-rank structure in the learned feature maps. Our results characterize the global minima of the loss and yield precise generalization thresholds, showing how the width of the target function governs learnability. This analysis bridges and extends ideas from spin-glass methods, matrix factorization, and convex optimization and emphasizes the deep link between low-rank matrix sensing and learning in quadratic neural networks.
Vittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent Krzakala
NeurIPS1
2024 Asymptotic Characterisation of the Performance of Robust Linear Regression in the Presence of Outliers
abstract
We study robust linear regression in high-dimension, when both the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha=n/d$, and study a data model that includes outliers. We provide exact asymptotics for the performances of the empirical risk minimisation (ERM) using $\ell_2$-regularised $\ell_2$, $\ell_1$, and Huber losses, which are the standard approach to such problems. We focus on two metrics for the performance: the generalisation error to similar datasets with outliers, and the estimation error of the original, unpolluted function. Our results are compared with the information theoretic Bayes-optimal estimation bound. For the generalization error, we find that optimally-regularised ERM is asymptotically consistent in the large sample complexity limit if one perform a simple calibration, and compute the rates of convergence. For the estimation error however, we show that due to a norm calibration mismatch, the consistency of the estimator requires an oracle estimate of the optimal norm, or the presence of a cross-validation set not corrupted by the outliers. We examine in detail how performance depends on the loss function and on the degree of outlier corruption in the training set and identify a region of parameters where the optimal performance of the Huber loss is identical to that of the $\ell_2$ loss, offering insights into the use cases of different loss functions.
Matteo Vilucchio, Emanuele Troiani, Vittorio Erba, Florent Krzakala
AISTATS3