VLDB 2026 Research / reviewers in the wild / expert
Simone Bombari
dblp:317/4969
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0008-9209-349XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning theory · 37% Trustworthy machine learning · 23% Kernel, tree and ensemble methods · 23% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features |
3.0 | 4 | 2025 | Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization · ICML 2025 Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features · ICML 2024 How Spurious Features are Memorized: Precise Analysis for Random and NTK Features · ICML 2024 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
2.0 | 3 | 2024 | How Spurious Features are Memorized: Precise Analysis for Random and NTK Features · ICML 2024 Beyond the Universal Law of Robustness: Sharper Laws for Random Features and Neural Tangent Kernels · ICML 2023 Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterization · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › robustness
spurious correlation |
1.6 | 2 | 2025 | Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization · ICML 2025 How Spurious Features are Memorized: Precise Analysis for Random and NTK Features · ICML 2024 |
Machine learning › Learning theory
over-parameterization |
1.4 | 2 | 2025 | Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization · ICML 2025 Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterization · NeurIPS 2022 |
Machine learning › Learning theory
high-dimensional regression |
0.9 | 1 | 2025 | Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization · ICML 2025 |
Machine learning › Deep learning architectures and training › attention mechanism › attention module
attention layer |
0.8 | 1 | 2024 | Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features · ICML 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.7 | 1 | 2023 | Beyond the Universal Law of Robustness: Sharper Laws for Random Features and Neural Tangent Kernels · ICML 2023 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.6 | 1 | 2022 | Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterization · NeurIPS 2022 |
Machine learning › Learning theory › neural network theory › network capacity
memorization capacity |
0.6 | 1 | 2022 | Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterization · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
schur complement · 0.9ridge regularization · 0.9covariance analysis · 0.9stability analysis · 0.8softmax analysis · 0.8generalization bounds · 0.8feature alignment analysis · 0.8interaction matrix analysis · 0.7empirical risk minimization · 0.7gradient descent analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-ParameterizationabstractLearning models have been shown to rely on spurious correlations between non-predictive features and the associated labels in the training data, with negative implications on robustness, bias and fairness.
In this work, we provide a statistical characterization of this phenomenon for high-dimensional regression, when the data contains a predictive *core* feature $x$ and a *spurious* feature $y$. Specifically, we quantify the amount of spurious correlations $\mathcal C$ learned via linear regression, in terms of the data covariance and the strength $\lambda$ of the ridge regularization.
As a consequence, we first capture the simplicity of $y$ through the spectrum of its covariance, and its correlation with $x$ through the Schur complement of the full data covariance. Next, we prove a trade-off between $\mathcal C$ and the in-distribution test loss $\mathcal L$, by showing that the value of $\lambda$ that minimizes $\mathcal L$ lies in an interval where $\mathcal C$ is increasing. Finally, we investigate the effects of over-parameterization via the random features model, by showing its equivalence to regularized linear regression.
Our theoretical results are supported by numerical experiments on Gaussian, Color-MNIST, and CIFAR-10 datasets. Simone Bombari, Marco Mondelli |
ICML | 1 |
| 2024 | How Spurious Features are Memorized: Precise Analysis for Random and NTK FeaturesabstractDeep learning models are known to overfit and memorize spurious features in the training dataset. While numerous empirical studies have aimed at understanding this phenomenon, a rigorous theoretical framework to quantify it is still missing. In this paper, we consider spurious features that are uncorrelated with the learning task, and we provide a precise characterization of how they are memorized via two separate terms: (i) the stability of the model with respect to individual training samples, and (ii) the feature alignment between the spurious pattern and the full sample. While the first term is well established in learning theory and it is connected to the generalization error in classical work, the second one is, to the best of our knowledge, novel. Our key technical result gives a precise characterization of the feature alignment for the two prototypical settings of random features (RF) and neural tangent kernel (NTK) regression. We prove that the memorization of spurious features weakens as the generalization capability increases and, through the analysis of the feature alignment, we unveil the role of the model and of its activation function. Numerical experiments show the predictive power of our theory on standard datasets (MNIST, CIFAR-10). Simone Bombari, Marco Mondelli |
ICML | 1 |
| 2024 | Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random FeaturesabstractUnderstanding the reasons behind the exceptional success of transformers requires a better analysis of why attention layers are suitable for NLP tasks. In particular, such tasks require predictive models to capture contextual meaning which often depends on one or few words, even if the sentence is long. Our work studies this key property, dubbed _word sensitivity_ (WS), in the prototypical setting of random features. We show that attention layers enjoy high WS, namely, there exists a vector in the space of embeddings that largely perturbs the random attention features map. The argument critically exploits the role of the softmax in the attention layer, highlighting its benefit compared to other activations (e.g., ReLU). In contrast, the WS of standard random features is of order $1/\sqrt{n}$, $n$ being the number of words in the textual sample, and thus it decays with the length of the context. We then translate these results on the word sensitivity into generalization bounds: due to their low WS, random features provably cannot learn to distinguish between two sentences that differ only in a single word; in contrast, due to their high WS, random attention features have higher generalization capabilities. We validate our theoretical results with experimental evidence over the BERT-Base word embeddings of the imdb review dataset. Simone Bombari, Marco Mondelli |
ICML | 1 |
| 2023 | Beyond the Universal Law of Robustness: Sharper Laws for Random Features and Neural Tangent KernelsabstractMachine learning models are vulnerable to adversarial perturbations, and a thought-provoking paper by Bubeck and Sellke has analyzed this phenomenon through the lens of over-parameterization: interpolating smoothly the data requires significantly more parameters than simply memorizing it. However, this "universal" law provides only a necessary condition for robustness, and it is unable to discriminate between models. In this paper, we address these gaps by focusing on empirical risk minimization in two prototypical settings, namely, random features and the neural tangent kernel (NTK). We prove that, for random features, the model is not robust for any degree of over-parameterization, even when the necessary condition coming from the universal law of robustness is satisfied. In contrast, for even activations, the NTK model meets the universal lower bound, and it is robust as soon as the necessary condition on over-parameterization is fulfilled. This also addresses a conjecture in prior work by Bubeck, Li and Nagaraj. Our analysis decouples the effect of the kernel of the model from an "interaction matrix", which describes the interaction with the test data and captures the effect of the activation. Our theoretical results are corroborated by numerical evidence on both synthetic and standard datasets (MNIST, CIFAR-10). Simone Bombari, Shayan Kiyani, Marco Mondelli |
ICML | 1 |
| 2022 | Sharp asymptotics on the compression of two-layer neural networksabstractIn this paper, we study the compression of a target two-layer neural network with N nodes into a compressed network with M2loss between the outputs of the target and of the compressed network, under the assumption of Gaussian inputs. By using tools from high-dimensional probability, we show that this non-convex problem can be simplified when the target network is sufficiently over-parameterized, and provide the error rate of this approximation as a function of the input dimension and N. In this mean-field limit, the simplified objective, as well as the optimal weights of the compressed network, does not depend on the realization of the target network, but only on expected scaling factors. Furthermore, for networks with ReLU activation, we conjecture that the optimum of the simplified optimization problem is achieved by taking weights on the Equiangular Tight Frame (ETF), while the scaling of the weights and the orientation of the ETF depend on the parameters of the target network. Numerical evidence is provided to support this conjecture. Mohammad Hossein Amani, Simone Bombari, Marco Mondelli, Rattana Pukdee, Stefano Rini |
ITW | 2 |
| 2022 | Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterizationabstractThe Neural Tangent Kernel (NTK) has emerged as a powerful tool to provide memorization, optimization and generalization guarantees in deep neural networks. A line of work has studied the NTK spectrum for two-layer and deep networks with at least a layer with $\Omega(N)$ neurons, $N$ being the number of training samples. Furthermore, there is increasing evidence suggesting that deep networks with sub-linear layer widths are powerful memorizers and optimizers, as long as the number of parameters exceeds the number of samples. Thus, a natural open question is whether the NTK is well conditioned in such a challenging sub-linear setup. In this paper, we answer this question in the affirmative. Our key technical contribution is a lower bound on the smallest NTK eigenvalue for deep networks with the minimum possible over-parameterization: up to logarithmic factors, the number of parameters is $\Omega(N)$ and, hence, the number of neurons is as little as $\Omega(\sqrt{N})$. To showcase the applicability of our NTK bounds, we provide two results concerning memorization capacity and optimization guarantees for gradient descent training. Simone Bombari, Mohammad Hossein Amani, Marco Mondelli |
NeurIPS | 1 |