EDBT 2026 Demo / reviewers in the wild / expert
Alexander Tsigler
dblp:194/2673
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Learning theory · 50% Probabilistic and Bayesian machine learning · 50% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › overfitting
benign overfitting |
0.7 | 1 | 2023 | Benign overfitting in ridge regression · J. Mach. Learn. Res. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression |
0.7 | 1 | 2023 | Benign overfitting in ridge regression · J. Mach. Learn. Res. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression |
0.7 | 1 | 2023 | Benign overfitting in ridge regression · J. Mach. Learn. Res. 2023 |
Machine learning › Learning theory
statistical learning theory |
0.7 | 1 | 2023 | Benign overfitting in ridge regression · J. Mach. Learn. Res. 2023 |
Methods — techniques the papers use, named apart from their topics
effective rank analysis · 0.7bias-variance decomposition · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Benign overfitting in ridge regressionabstractIn many modern applications of deep learning the neural network has many more parameters than the data points used for its training. Motivated by those practices, a large body of recent theoretical research has been devoted to studying overparameterized models. One of the central phenomena in this regime is the ability of the model to interpolate noisy data, but still have test error lower than the amount of noise in that data. arXiv:1906.11300 characterized for which covariance structure of the data such a phenomenon can happen in linear regression if one considers the interpolating solution with minimum $\ell_2$-norm and the data has independent components: they gave a sharp bound on the variance term and showed that it can be small if and only if the data covariance has high effective rank in a subspace of small co-dimension. We strengthen and complete their results by eliminating the independence assumption and providing sharp bounds for the bias term. Thus, our results apply in a much more general setting than those of arXiv:1906.11300, e.g., kernel regression, and not only characterize how the noise is damped but also which part of the true signal is learned. Moreover, we extend the result to the setting of ridge regression, which allows us to explain another interesting phenomenon: we give general sufficient conditions under which the optimal regularization is negative. Alexander Tsigler, Peter L. Bartlett |
J. Mach. Learn. Res. | 1 |