VLDB 2026 Research / reviewers in the wild / expert
Emanuele Troiani
dblp:267/5270
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-0968-7585ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning theory · 57% Deep learning architectures and training · 41% Graph learning · 2% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 54% Distributed computing theory · 23% Information theory · 23% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › high-dimensional statistics
high-dimensional asymptotics |
1.7 | 2 | 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025 Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds · ICML 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Bayes optimal learning of attention-indexed models · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
attention network |
0.9 | 1 | 2025 | Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds · ICML 2025 |
Machine learning › Learning theory
empirical risk minimization |
0.9 | 1 | 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025 |
Machine learning › Learning theory
generalization |
0.9 | 1 | 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025 |
Machine learning › Learning theory
generalization error |
0.9 | 1 | 2025 | Bayes optimal learning of attention-indexed models · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
overparameterized neural network |
0.9 | 1 | 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025 |
Machine learning › Learning theory
sample complexity |
0.9 | 1 | 2025 | Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds · ICML 2025 |
Machine learning › Learning theory › phase transition
sharp threshold |
0.9 | 1 | 2025 | Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds · ICML 2025 |
Machine learning › Learning theory
neural network theory |
0.8 | 1 | 2024 | Bayes-optimal learning of an extensive-width neural network from quadratically many samples · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
0.8 | 1 | 2024 | The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents · ICML 2024 |
Machine learning › Deep learning architectures and training › training dynamics
two-layer neural network training |
0.8 | 1 | 2024 | The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents · ICML 2024 |
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks |
0.8 | 1 | 2024 | Bayes-optimal learning of an extensive-width neural network from quadratically many samples · NeurIPS 2024 |
Machine learning › Graph learning › graph neural network › message passing
approximate message passing |
0.3 | 1 | 2025 | Bayes optimal learning of attention-indexed models · NeurIPS 2025 |
Mathematical optimization › continuous optimization
convex optimization |
0.3 | 1 | 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025 |
Mathematical optimization › continuous optimization › convex optimization › norm optimization
nuclear norm minimization |
0.3 | 1 | 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks · NeurIPS 2025 |
Information theory › signal processing › compressed sensing
approximate message passing |
0.2 | 1 | 2024 | Bayes-optimal learning of an extensive-width neural network from quadratically many samples · NeurIPS 2024 |
Distributed computing theory
message passing |
0.2 | 1 | 2024 | Bayes-optimal learning of an extensive-width neural network from quadratically many samples · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
approximate message passing · 3.3spin-glass methods · 1.7matrix factorization · 1.7convex optimization · 1.7statistical mechanics · 0.9random matrix theory · 0.9multi-index model · 0.9bayes-optimal learning · 0.9rotationally invariant matrix denoising · 0.8gradient flow analysis · 0.8gradient descent · 0.8dynamical mean field theory · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fundamental computational limits of weak learnability in high-dimensional multi-index modelsabstractMulti-index models - functions which only depend on the covariates through a non-linear transformation of their projection on a subspace - are a useful benchmark for investigating feature learning with neural networks. This paper examines the theoretical boundaries of efficient learnability in this hypothesis class, focusing particularly on the minimum sample complexity required for weakly recovering their low-dimensional structure with first-order iterative algorithms, in the high-dimensional regime where the number of samples is $n=\alpha d$ is proportional to the covariate dimension $d$. Our findings unfold in three parts: (i) first, we identify under which conditions a \textit{trivial subspace} can be learned with a single step of a first-order algorithm for any $\alpha>0$; (ii) second, in the case where the trivial subspace is empty, we provide necessary and sufficient conditions for the existence of an {\it easy subspace} consisting of directions that can be learned only above a certain sample complexity $\alpha>\alpha_c$. The critical threshold $\alpha_{c}$ marks the presence of a computational phase transition, in the sense that it is conjectured that no efficient iterative algorithm can succeed for $\alpha<\alpha_c$. In a limited but interesting set of really hard directions -akin to the parity problem- $\alpha_c$ is found to diverge. Finally, (iii) we demonstrate that interactions between different directions can result in an intricate hierarchical learning phenomenon, where some directions can be learned sequentially when coupled to easier ones. Our analytical approach is built on the optimality of approximate message-passing algorithms among first-order iterative methods, delineating the fundamental learnability limit across a broad spectrum of algorithms, including neural networks trained with gradient descent. Emanuele Troiani, Yatin Dandi, Leonardo Defilippis, Lenka Zdeborová, Bruno Loureiro, Florent Krzakala |
AISTATS | 1 |
| 2025 | Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholdsabstractIn this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index model to sequential covariates, for which we establish a number of general results. In the context of Bayes-optimal learning, in the limit of large dimension $D$ and proportionally large number of samples $N$, we derive a sharp asymptotic characterization of the optimal performance as well as the performance of the best-known polynomial-time algorithm for this setting –namely approximate message-passing–, and characterize sharp thresholds on the minimal sample complexity required for better-than-random prediction performance. Our analysis uncovers, in particular, how the different layers are learned sequentially. Finally, we discuss how this sequential learning can also be observed in a realistic setup. Emanuele Troiani, Hugo Cui, Yatin Dandi, Florent Krzakala, Lenka Zdeborová |
ICML | 1 |
| 2025 | Bayes optimal learning of attention-indexed modelsabstractWe introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures. Fabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka Zdeborová |
NeurIPS | 2 |
| 2025 | The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic NetworksabstractWe study the high-dimensional asymptotics of empirical risk minimization (ERM) in over-parametrized two-layer neural networks with quadratic activations trained on synthetic data. We derive sharp asymptotics for both training and test errors by mapping the $\ell_2$-regularized learning problem to a convex matrix sensing task with nuclear norm penalization. This reveals that capacity control in such networks emerges from a low-rank structure in the learned feature maps. Our results characterize the global minima of the loss and yield precise generalization thresholds, showing how the width of the target function governs learnability. This analysis bridges and extends ideas from spin-glass methods, matrix factorization, and convex optimization and emphasizes the deep link between low-rank matrix sensing and learning in quadratic neural networks. Vittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent Krzakala |
NeurIPS | 2 |
| 2024 | Asymptotic Characterisation of the Performance of Robust Linear Regression in the Presence of OutliersabstractWe study robust linear regression in high-dimension, when both the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha=n/d$, and study a data model that includes outliers. We provide exact asymptotics for the performances of the empirical risk minimisation (ERM) using $\ell_2$-regularised $\ell_2$, $\ell_1$, and Huber losses, which are the standard approach to such problems. We focus on two metrics for the performance: the generalisation error to similar datasets with outliers, and the estimation error of the original, unpolluted function. Our results are compared with the information theoretic Bayes-optimal estimation bound. For the generalization error, we find that optimally-regularised ERM is asymptotically consistent in the large sample complexity limit if one perform a simple calibration, and compute the rates of convergence. For the estimation error however, we show that due to a norm calibration mismatch, the consistency of the estimator requires an oracle estimate of the optimal norm, or the presence of a cross-validation set not corrupted by the outliers. We examine in detail how performance depends on the loss function and on the degree of outlier corruption in the training set and identify a region of parameters where the optimal performance of the Huber loss is identical to that of the $\ell_2$ loss, offering insights into the use cases of different loss functions. Matteo Vilucchio, Emanuele Troiani, Vittorio Erba, Florent Krzakala |
AISTATS | 2 |
| 2024 | The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap ExponentsabstractWe investigate the training dynamics of two-layer neural networks when learning multi-index target functions. We focus on multi-pass gradient descent (GD) that reuses the batches multiple times and show that it significantly changes the conclusion about which functions are learnable compared to single-pass gradient descent. In particular, multi-pass GD with finite stepsize is found to overcome the limitations of gradient flow and single-pass GD given by the information exponent (Ben Arous et al., 2021) and leap exponent (Abbe et al., 2023) of the target function. We show that upon re-using batches, the network achieves in just two time steps an overlap with the target subspace even for functions not satisfying the staircase property (Abbe et al., 2021). We characterize the (broad) class of functions efficiently learned in finite time. The proof of our results is based on the analysis of the Dynamical Mean-Field Theory (DMFT). We further provide a closed-form description of the dynamical process of the low-dimensional projections of the weights, and numerical experiments illustrating the theory. Yatin Dandi, Emanuele Troiani, Luca Arnaboldi 0002, Luca Pesce, Lenka Zdeborová, Florent Krzakala |
ICML | 2 |
| 2024 | Bayes-optimal learning of an extensive-width neural network from quadratically many samplesabstractWe consider the problem of learning a target function corresponding to a single
hidden layer neural network, with a quadratic activation function after the first layer,
and random weights. We consider the asymptotic limit where the input dimension
and the network width are proportionally large. Recent work [Cui et al., 2023]
established that linear regression provides Bayes-optimal test error to learn such
a function when the number of available samples is only linear in the dimension.
That work stressed the open challenge of theoretically analyzing the optimal test
error in the more interesting regime where the number of samples is quadratic in
the dimension. In this paper, we solve this challenge for quadratic activations and
derive a closed-form expression for the Bayes-optimal test error. We also provide an
algorithm, that we call GAMP-RIE, which combines approximate message passing
with rotationally invariant matrix denoising, and that asymptotically achieves the
optimal performance. Technically, our result is enabled by establishing a link
with recent works on optimal denoising of extensive-rank matrices and on the
ellipsoid fitting problem. We further show empirically that, in the absence of
noise, randomly-initialized gradient descent seems to sample the space of weights,
leading to zero training loss, and averaging over initialization leads to a test error
equal to the Bayes-optimal one. Antoine Maillard, Emanuele Troiani, Simon Martin 0008, Florent Krzakala, Lenka Zdeborová |
NeurIPS | 2 |