Sara A. van de Geer

dblp:65/323 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
1since 2021 · last 2021
0000-0001-7181-2872ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-authorTheory of computation · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2021 De-Biased Sparse PCA: Inference for Eigenstructure of Large Covariance Matrices
abstract
Sparse principal component analysis has become one of the most widely used techniques for dimensionality reduction in high-dimensional datasets. While many methods are available for point estimation of eigenstructure in high-dimensional settings, in this paper we propose methodology for uncertainty quantification, such as construction of confidence intervals and tests for the principal eigenvector and the corresponding largest eigenvalue. We base our methodology on an M-estimator with Lasso penalty which achieves minimax optimal rates and is used to construct a de-biased sparse PCA estimator. The novel estimator has a Gaussian limiting distribution and can be used for hypothesis testing or support recovery of the first eigenvector. The empirical performance of the new estimator is demonstrated on synthetic data and we also show that the estimator compares favourably with the classical PCA in moderately high-dimensional regimes.
Jana Janková, Sara A. van de Geer
IEEE Trans. Inf. Theory2
2019 Sharp Oracle Inequalities for Stationary Points of Nonconvex Penalized M-Estimators
abstract
Many statistical estimation procedures lead to nonconvex optimization problems. Algorithms to solve these problems are often guaranteed to output a stationary point of the optimization problem. Oracle inequalities are an important theoretical instrument to assess the statistical performance of an estimator. Oracle results have focused on the theoretical properties of the uncomputable (global) minimum or maximum. In this paper, a general framework used for convex optimization problems to derive oracle inequalities for stationary points is extended. A main new ingredient of these oracle inequalities is that they are sharp: they show closeness to the best approximation within the model plus a remainder term. We apply this framework to different estimation problems.
Andreas Elsener, Sara A. van de Geer
IEEE Trans. Inf. Theory2
2019 A Framework for the Construction of Upper Bounds on the Number of Affine Linear Regions of ReLU Feed-Forward Neural Networks
abstract
We present a framework to derive upper bounds on the number of regions that feed-forward neural networks with ReLU activation functions are affine linear on. It is based on an inductive analysis that keeps track of the number of such regions per dimensionality of their images within the layers. More precisely, the information about the number regions per dimensionality is pushed through the layers starting with one region of the input dimension of the neural network and using a recursion based on an analysis of how many regions per output dimensionality a subsequent layer with a certain width can induce on an input region with a given dimensionality. The final bound on the number of regions depends on the number and widths of the layers of the neural network and on some additional parameters that were used for the recursion. It is stated in terms of the L1-norm of the last column of a product of matrices and provides a unifying treatment of several previously known bounds: Depending on the choice of the recursion parameters that determine these matrices, it is possible to obtain the bounds from Montúfar et al. [1] (2014), Montúfar [2] (2017), and Serra et al. [3] (2017) as special cases. For the latter, which is the strongest of these bounds, the formulation in terms of matrices provides new insight. In particular, by using explicit formulas for a Jordan-like decomposition of the involved matrices, we achieve new tighter results for the asymptotic setting, where the number of layers of the same fixed width tends to infinity.
Peter Hinz, Sara A. van de Geer
IEEE Trans. Inf. Theory2
2018 On Tight Bounds for the Lasso
abstract
We present upper and lower bounds for the prediction error of the Lasso. For the case of random Gaussian design, we show that under mild conditions the prediction error of the Lasso is up to smaller order terms dominated by the prediction error of its noiseless counterpart. We then provide exact expressions for the prediction error of the latter, in terms of compatibility constants. Here, we assume the active components of the underlying regression function satisfy some “betamin" condition. For the case of fixed design, we provide upper and lower bounds, again in terms of compatibility constants. As an example, we give an up to a logarithmic term tight bound for the least squares estimator with total variation penalty.
Sara A. van de Geer
J. Mach. Learn. Res.1
2017 Sharp Oracle Inequalities for Square Root Regularization
abstract
We study a set of regularization methods for high-dimensional linear regression models. These penalized estimators have the square root of the residual sum of squared errors as loss function, and any weakly decomposable norm as penalty function. This fit measure is chosen because of its property that the estimator does not depend on the unknown standard deviation of the noise. On the other hand, a generalized weakly decomposable norm penalty is very useful in being able to deal with different underlying sparsity structures. We can choose a different sparsity inducing norm depending on how we want to interpret the unknown parameter vector $\beta$. Structured sparsity norms, as defined in Micchelli et al. (2010), are special cases of weakly decomposable norms, therefore we also include the square root LASSO (Belloni et al., 2011), the group square root LASSO (Bunea et al., 2014) and a new method called the square root SLOPE (in a similar fashion to the SLOPE from Bogdan et al. 2015). For this collection of estimators our results provide sharp oracle inequalities with the Karush-Kuhn-Tucker conditions. We discuss some examples of estimators. Based on a simulation we illustrate some advantages of the square root SLOPE.
Benjamin Stucky, Sara A. van de Geer
J. Mach. Learn. Res.2
2009 Taking Advantage of Sparsity in Multi-Task Learning
Karim Lounici, Massimiliano Pontil, Alexandre B. Tsybakov, Sara A. van de Geer
COLT4
2005 Asymptotics in Empirical Risk Minimization
abstract
In this paper, we study a two-category classification problem. We indicate the categories by labels Y=1 and Y=-1. We observe a covariate, or feature, X ∈ X ⊂ ℜd. Consider a collection {ha} of classifiers indexed by a finite-dimensional parameter a, and the classifier ha* that minimizes the prediction error over this class. The parameter a* is estimated by the empirical risk minimizer ân over the class, where the empirical risk is calculated on a training sample of size n. We apply the Kim Pollard Theorem to show that under certain differentiability assumptions, ân converges to a* with rate n-1/3, and also present the asymptotic distribution of the renormalized estimator. For example, let V0 denote the set of x on which, given X=x, the label Y=1 is more likely (than the label Y=-1). If X is one-dimensional, the set V0 is the union of disjoint intervals. The problem is then to estimate the thresholds of the intervals. We obtain the asymptotic distribution of the empirical risk minimizer when the classifiers have K thresholds, where K is fixed. We furthermore consider an extension to higher-dimensional X, assuming basically that V0 has a smooth boundary in some given parametric class. We also discuss various rates of convergence when the differentiability conditions are possibly violated. Here, we again restrict ourselves to one-dimensional X. We show that the rate is n-1 in certain cases, and then also obtain the asymptotic distribution for the empirical prediction error.
Leila Mohammadi, Sara A. van de Geer
J. Mach. Learn. Res.2
2004 A global test for groups of genes: testing association with a clinical outcome
abstract
MOTIVATION: This paper presents a global test to be used for the analysis of microarray data. Using this test it can be determined whether the global expression pattern of a group of genes is significantly related to some clinical outcome of interest. Groups of genes may be any size from a single gene to all genes on the chip (e.g. known pathways, specific areas of the genome or clusters from a cluster analysis). RESULT: The test allows groups of genes of different size to be compared, because the test gives one p-value for the group, not a p-value for each gene. Researchers can use the test to investigate hypotheses based on theory or past research or to mine gene ontology databases for interesting pathways. Multiple testing problems do not occur unless many groups are tested. Special attention is given to visualizations of the test result, focussing on the associations between samples and showing the impact of individual genes on the test result. AVAILABILITY: An R-package globaltest is available from http://www.bioconductor.org
Jelle J. Goeman, Sara A. van de Geer, Floor de Kort, Hans C. van Houwelingen
Bioinform.2