VLDB 2026 Research / reviewers in the wild / expert
Ingo Steinwart
dblp:89/3492
· DBLP profile ↗
46ranked-venue papers
22as first author
11since 2021 · last 2025
0000-0002-4436-7109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 16 first-author · 10 since 2021Theory of computation · 6 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Empirical risk minimization in the interpolating regime with application to neural network learningabstractAbstract A common strategy to train deep neural networks (DNNs) is to use very large architectures and to train them until they (almost) achieve zero training error. Empirically observed good generalization performance on test data, even in the presence of lots of label noise, corroborate such a procedure. On the other hand, in statistical learning theory it is known that over-fitting models may lead to poor generalization properties, occurring in e.g. empirical risk minimization (ERM) over too large hypotheses classes. Inspired by this contradictory behavior, so-called interpolation methods have recently received much attention, leading to consistent and optimally learning methods for, e.g., some local averaging schemes with zero training error. We extend this analysis to ERM-like methods for least squares regression and show that for certain, large hypotheses classes called inflated histograms, some interpolating empirical risk minimizers enjoy very good statistical guarantees while others fail in the worst sense. Moreover, we show that the same phenomenon occurs for DNNs with zero training error and sufficiently large architectures. Nicole Mücke, Ingo Steinwart |
Mach. Learn. | 2 |
| 2024 | Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataabstractFor classification and regression on tabular data, the dominance of gradient-boosted decision trees (GBDTs) has recently been challenged by often much slower deep learning methods with extensive hyperparameter tuning. We address this discrepancy by introducing (a) RealMLP, an improved multilayer perceptron (MLP), and (b) strong meta-tuned default parameters for GBDTs and RealMLP. We tune RealMLP and the default parameters on a meta-train benchmark with 118 datasets and compare them to hyperparameter-optimized versions on a disjoint meta-test benchmark with 90 datasets, as well as the GBDT-friendly benchmark by Grinsztajn et al. (2022). Our benchmark results on medium-to-large tabular datasets (1K--500K samples) show that RealMLP offers a favorable time-accuracy tradeoff compared to other neural baselines and is competitive with GBDTs in terms of benchmark scores. Moreover, a combination of RealMLP and GBDTs with improved default parameters can achieve excellent results without hyperparameter tuning. Finally, we demonstrate that some of RealMLP's improvements can also considerably improve the performance of TabR with default parameters. David Holzmüller, Léo Grinsztajn, Ingo Steinwart |
NeurIPS | 3 |
| 2023 | Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimensionabstractThe success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has been established for some learning methods, current literature suggests that for regression with typical kernel methods and wide neural networks, benign overfitting requires a high-dimensional setting, where the dimension grows with the sample size. In this paper, we show that the smoothness of the estimators, and not the dimension, is the key: benign overfitting is possible if and only if the estimator's derivatives are large enough. We generalize existing inconsistency results to non-interpolating models and more kernels to show that benign overfitting with moderate derivatives is impossible in fixed dimension. Conversely, we show that benign overfitting is possible for regression with a sequence of spiky-smooth kernels with large derivatives. Using neural tangent kernels, we translate our results to wide neural networks. We prove that while infinite-width networks do not overfit benignly with the ReLU activation, this can be fixed by adding small high-frequency fluctuations to the activation function. Our experiments verify that such neural networks, while overfitting, can indeed generalize well even on low-dimensional data sets. Moritz Haas, David Holzmüller, Ulrike von Luxburg, Ingo Steinwart |
NeurIPS | 4 |
| 2023 | A Framework and Benchmark for Deep Batch Active Learning for RegressionabstractThe acquisition of labels for supervised learning can be expensive. To improve the sample efficiency of neural network regression, we study active learning methods that adaptively select batches of unlabeled data for labeling. We present a framework for constructing such methods out of (network-dependent) base kernels, kernel transformations, and selection methods. Our framework encompasses many existing Bayesian methods based on Gaussian process approximations of neural networks as well as non-Bayesian methods. Additionally, we propose to replace the commonly used last-layer features with sketched finite-width neural tangent kernels and to combine them with a novel clustering method. To evaluate different methods, we introduce an open-source benchmark consisting of 15 large tabular regression data sets. Our proposed method outperforms the state-of-the-art on our benchmark, scales to large data sets, and works out-of-the-box without adjusting the network architecture or training code. We provide open-source code that includes efficient implementations of all kernels, kernel transformations, and selection methods, and can be used for reproducing our results. David Holzmüller, Viktor Zaverkin, Johannes Kästner, Ingo Steinwart |
J. Mach. Learn. Res. | 4 |
| 2023 | Adaptive Clustering Using Kernel Density EstimatorsabstractWe derive and analyze a generic, recursive algorithm for estimating all splits in a finite cluster tree as well as the corresponding clusters. We further investigate statistical properties of this generic clustering algorithm when it receives level set estimates from a kernel density estimator. In particular, we derive finite sample guarantees, consistency, rates of convergence, and an adaptive data-driven strategy for choosing the kernel bandwidth. For these results we do not need continuity assumptions on the density such as Hölder continuity, but only require intuitive geometric assumptions of non-parametric nature. In addition, we compare our results to other guarantees found in the literature and also present some experiments comparing our algorithm to $k$-means and hierarchical clustering. Ingo Steinwart, Bharath K. Sriperumbudur, Philipp Thomann |
J. Mach. Learn. Res. | 1 |
| 2022 | SOSP: Efficiently Capturing Global Correlations by Second-Order Structured Pruning
Manuel Nonnenmacher, Thomas Pfeil, Ingo Steinwart, David Reeb |
ICLR | 3 |
| 2022 | Utilizing Expert Features for Contrastive Learning of Time-Series RepresentationsabstractWe present an approach that incorporates expert knowledge for time-series representation learning. Our method employs expert features to replace the commonly used data transformations in previous contrastive learning approaches. We do this since time-series data frequently stems from the industrial or medical field where expert features are often available from domain experts, while transformations are generally elusive for time-series data. We start by proposing two properties that useful time-series representations should fulfill and show that current representation learning approaches do not ensure these properties. We therefore devise ExpCLR, a novel contrastive learning approach built on an objective that utilizes expert features to encourage both properties for the learned representation. Finally, we demonstrate on three real-world time-series datasets that ExpCLR surpasses several state-of-the-art methods for both unsupervised and semi-supervised representation learning. Manuel Nonnenmacher, Lukas Oldenburg, Ingo Steinwart, David Reeb |
ICML | 3 |
| 2022 | Improved Classification Rates for Localized SVMsabstractLocalized support vector machines solve SVMs on many spatially defined small chunks and besides their computational benefit compared to global SVMs one of their main characteristics is the freedom of choosing arbitrary kernel and regularization parameter on each cell. We take advantage of this observation to derive global learning rates for localized SVMs with Gaussian kernels and hinge loss. It turns out that our rates outperform under suitable sets of assumptions known classification rates for localized SVMs, for global SVMs, and other learning algorithms based on e.g., plug-in rules or trees. The localized SVM rates are achieved under a set of margin conditions, which describe the behavior of the data-generating distribution, and no assumption on the existence of a density is made. Moreover, we show that our rates are obtained adaptively, that is without knowing the margin parameters in advance. The statistical analysis of the excess risk relies on a simple partitioning based technique, which splits the input space into a subset that is close to the decision boundary and into a subset that is sufficiently far away. A crucial condition to derive then improved global rates is a margin condition that relates the distance to the decision boundary to the amount of noise. Ingrid Blaschzyk, Ingo Steinwart |
J. Mach. Learn. Res. | 2 |
| 2022 | Training Two-Layer ReLU Networks with Gradient Descent is InconsistentabstractWe prove that two-layer (Leaky)ReLU networks initialized by e.g. the widely used method proposed by He et al. (2015) and trained using gradient descent on a least-squares loss are not universally consistent. Specifically, we describe a large class of one-dimensional data-generating distributions for which, with high probability, gradient descent only finds a bad local minimum of the optimization landscape, since it is unable to move the biases far away from their initialization at zero. It turns out that in these cases, the found network essentially performs linear regression even if the target function is non-linear. We further provide numerical evidence that this happens in practical situations, for some multi-dimensional distributions and that stochastic gradient descent exhibits similar behavior. We also provide empirical results on how the choice of initialization and optimizer can influence this behavior. David Holzmüller, Ingo Steinwart |
J. Mach. Learn. Res. | 2 |
| 2021 | Which Minimizer Does My Neural Network Converge To?
Manuel Nonnenmacher, David Reeb, Ingo Steinwart |
ECML/PKDD (3) | 3 |
| 2021 | A closer look at covering number bounds for Gaussian kernels
Ingo Steinwart, Simon Fischer 0004 |
J. Complex. | 1 |
| 2020 | Sobolev Norm Learning Rates for Regularized Least-Squares AlgorithmsabstractLearning rates for least-squares regression are typically expressed in terms of $L_2$-norms. In this paper we extend these rates to norms stronger than the $L_2$-norm without requiring the regression function to be contained in the hypothesis space. In the special case of Sobolev reproducing kernel Hilbert spaces used as hypotheses spaces, these stronger norms coincide with fractional Sobolev norms between the used Sobolev space and $L_2$. As a consequence, not only the target function but also some of its derivatives can be estimated without changing the algorithm. From a technical point of view, we combine the well-known integral operator techniques with an embedding property, which so far has only been used in combination with empirical process arguments. This combination results in new finite sample bounds with respect to the stronger norms. From these finite sample bounds our rates easily follow. Finally, we prove the asymptotic optimality of our results in many cases. Simon Fischer 0004, Ingo Steinwart |
J. Mach. Learn. Res. | 2 |
| 2019 | Learning rates for kernel-based expectile regression
Ingo Steinwart |
Mach. Learn. | 2 |
| 2018 | Kernel Density Estimation for Dynamical SystemsabstractWe study the density estimation problem with observations generated by certain dynamical systems that admit a unique underlying invariant Lebesgue density. Observations drawn from dynamical systems are not independent and moreover, usual mixing concepts may not be appropriate for measuring the dependence among these observations. By employing the $\mathcal{C}$-mixing concept to measure the dependence, we conduct statistical analysis on the consistency and convergence of the kernel density estimator. Our main results are as follows: First, we show that with properly chosen bandwidth, the kernel density estimator is universally consistent under $L_1$-norm; Second, we establish convergence rates for the estimator with respect to several classes of dynamical systems under $L_1$-norm. In the analysis, the density function $f$ is only assumed to be Hölder continuous or pointwise Hölder controllable which is a weak assumption in the literature of nonparametric density estimation and also more realistic in the dynamical system context. Last but not least, we prove that the same convergence rates of the estimator under $L_\infty$-norm and $L_1$-norm can be achieved when the density function is Hölder continuous, compactly supported, and bounded. The bandwidth selection problem of the kernel density estimator for dynamical system is also discussed in our study via numerical simulations. Hanyuan Hang, Ingo Steinwart, Yunlong Feng, Johan A. K. Suykens |
J. Mach. Learn. Res. | 2 |
| 2017 | Spatial Decompositions for Large Scale SVMsabstractAlthough support vector machines (SVMs) are theoretically well understood, their underlying optimization problem becomes very expensive if, for example, hundreds of thousands of samples and a non-linear kernel are considered. Several approaches have been proposed in the past to address this serious limitation. In this work we investigate a decomposition strategy that learns on small, spatially defined data chunks. Our contributions are two fold: On the theoretical side we establish an oracle inequality for the overall learning method using the hinge loss, and show that the resulting rates match those known for SVMs solving the complete optimization problem with Gaussian kernels. On the practical result we compare our approach to learning SVMs on small, randomly chosen chunks. Here it turns out that for comparable training times our approach is significantly faster during testing and also reduces the test error in most cases significantly. Furthermore, we show that our approach easily scales up to 10 million training samples: including hyper-parameter selection using cross validation, the entire training only takes a few hours on a single machine. Finally, we report an experiment on 32 million training samples. All experiments used liquidSVM (Steinwart and Thomann, 2017) Philipp Thomann, Ingrid Blaschzyk, Mona Meister, Ingo Steinwart |
AISTATS | 4 |
| 2016 | Optimal Learning Rates for Localized SVMsabstractOne of the limiting factors of using support vector machines (SVMs) in large scale applications are their super-linear computational requirements in terms of the number of training samples. To address this issue, several approaches that train SVMs on many small chunks separately have been proposed in the literature. With the exception of random chunks, which is also known as divide-and-conquer kernel ridge regression, however, these approaches have only been empirically investigated. In this work we investigate a spatially oriented method to generate the chunks. For the resulting localized SVM that uses Gaussian kernels and the least squares loss we derive an oracle inequality, which in turn is used to deduce learning rates that are essentially minimax optimal under some standard smoothness assumptions on the regression function. In addition, we derive local learning rates that are based on the local smoothness of the regression function. We further introduce a data-dependent parameter selection method for our local SVM approach and show that this method achieves the same almost optimal learning rates. Finally, we present a few larger scale experiments for our localized SVM showing that it achieves essentially the same test error as a global SVM for a fraction of the computational requirements. In addition, it turns out that the computational requirements for the local SVMs are similar to those of a vanilla random chunk approach, while the achieved test errors are significantly better. Mona Meister, Ingo Steinwart |
J. Mach. Learn. Res. | 2 |
| 2016 | Learning Theory Estimates with Observations from General Stationary Stochastic ProcessesabstractThis letter investigates the supervised learning problem with observations drawn from certain general stationary stochastic processes. Here by general, we mean that many stationary stochastic processes can be included. We show that when the stochastic processes satisfy a generalized Bernstein-type inequality, a unified treatment on analyzing the learning schemes with various mixing processes can be conducted and a sharp oracle inequality for generic regularized empirical risk minimization schemes can be established. The obtained oracle inequality is then applied to derive convergence rates for several learning schemes such as empirical risk minimization (ERM), least squares support vector machines (LS-SVMs) using given generic kernels, and SVMs using gaussian kernels for both least squares and quantile regression. It turns out that for independent and identically distributed (i.i.d.) processes, our learning rates for ERM recover the optimal rates. For non-i.i.d. processes, including geometrically [Formula: see text]-mixing Markov processes, geometrically [Formula: see text]-mixing processes with restricted decay, [Formula: see text]-mixing processes, and (time-reversed) geometrically [Formula: see text]-mixing processes, our learning rates for SVMs with gaussian kernels match, up to some arbitrarily small extra term in the exponent, the optimal rates. For the remaining cases, our rates are at least close to the optimal rates. As a by-product, the assumed generalized Bernstein-type inequality also provides an interpretation of the so-called effective number of observations for various mixing processes. Hanyuan Hang, Yunlong Feng, Ingo Steinwart, Johan A. K. Suykens |
Neural Comput. | 3 |
| 2015 | Towards an axiomatic approach to hierarchical clustering of measures
Philipp Thomann, Ingo Steinwart, Nico Schmid |
J. Mach. Learn. Res. | 2 |
| 2014 | Elicitation and Identification of PropertiesabstractProperties of distributions are real-valued functionals such as the mean, quantile or conditional value at risk. A property is elicitable if there exists a scoring function such that minimization of the associated risks recovers the property. We extend existing results to characterize the elicitability of properties in a general setting. We further relate elicitability to identifiability (a notion introduced by Osband) and provide a general formula describing all scoring functions for an elicitable property. Finally, we draw some connections to the theory of coherent risk measures. Ingo Steinwart, Chloé Pasin, Robert C. Williamson |
COLT | 1 |
| 2011 | Optimal learning rates for least squares SVMs using Gaussian kernelsabstractWe prove a new oracle inequality for support vector machines with Gaussian RBF kernels solving the regularized least squares regression problem. To this end, we apply the modulus of smoothness. With the help of the new oracle inequality we then derive learning rates that can also be achieved by a simple data-dependent parameter selection method. Finally, it turns out that our learning rates are asymptotically optimal for regression functions satisfying certain standard smoothness conditions. Mona Eberts, Ingo Steinwart |
NIPS | 2 |
| 2011 | Training SVMs Without Offset
Ingo Steinwart, Don R. Hush, Clint Scovel |
J. Mach. Learn. Res. | 1 |
| 2010 | Using support vector machines for anomalous change detectionabstractWe cast anomalous change detection as a binary classification problem, and use a support vector machine (SVM) to build a detector that does not depend on assumptions about the underlying data distribution. To speed up the computation, our SVM is implemented, in part, on a graphical processing unit. Results on real and simulated anomalous changes are used to compare performance to algorithms which effectively assume a Gaussian distribution. Ingo Steinwart, James Theiler, Daniel Llamocca |
IGARSS | 1 |
| 2010 | Universal Kernels on Non-Standard Input SpacesabstractDuring the last years support vector machines (SVMs) have been successfully applied even in situations where the input space $X$ is not necessarily a subset of $R^d$. Examples include SVMs using probability measures to analyse e.g. histograms or coloured images, SVMs for text classification and web mining, and SVMs for applications from computational biology using, e.g., kernels for trees and graphs. Moreover, SVMs are known to be consistent to the Bayes risk, if either the input space is a complete separable metric space and the reproducing kernel Hilbert space (RKHS) $H\subset L_p(P_X)$ is dense, or if the SVM is based on a universal kernel $k$. So far, however, there are no RKHSs of practical interest known that satisfy these assumptions on $\cH$ or $k$ if $X \not\subset R^d$. We close this gap by providing a general technique based on Taylor-type kernels to explicitly construct universal kernels on compact metric spaces which are not subset of $R^d$. We apply this technique for the following special cases: universal kernels on the set of probability measures, universal kernels based on Fourier transforms, and universal kernels for signal processing. Andreas Christmann, Ingo Steinwart |
NIPS | 2 |
| 2010 | Radial kernels and their reproducing kernel Hilbert spaces
Clint Scovel, Don R. Hush, Ingo Steinwart, James Theiler |
J. Complex. | 3 |
| 2009 | Optimal Rates for Regularized Least Squares Regression
Ingo Steinwart, Don R. Hush, Clint Scovel |
COLT | 1 |
| 2009 | Fast Learning from Non-i.i.d. ObservationsabstractWe prove an oracle inequality for generic regularized empirical risk minimization algorithms learning from $\a$-mixing processes. To illustrate this oracle inequality, we use it to derive learning rates for some learning methods including least squares SVMs. Since the proof of the oracle inequality uses recent localization ideas developed for independent and identically distributed (i.i.d.) processes, it turns out that these learning rates are close to the optimal rates known in the i.i.d. case. Ingo Steinwart, Andreas Christmann |
NIPS | 1 |
| 2009 | Oracle inequalities for support vector machines that are based on random entropy numbers
Ingo Steinwart |
J. Complex. | 1 |
| 2008 | Sparsity of SVMs that use the epsilon-insensitive lossabstractIn this paper lower and upper bounds for the number of support vectors are derived for support vector machines (SVMs) based on the epsilon-insensitive loss function. It turns out that these bounds are asymptotically tight under mild assumptions on the data generating distribution. Finally, we briefly discuss a trade-off in epsilon between sparsity and accuracy if the SVM is used to estimate the conditional median. Ingo Steinwart, Andreas Christmann |
NIPS | 1 |
| 2007 | Gaps in Support Vector Optimization
Nikolas List, Don R. Hush, Clint Scovel, Ingo Steinwart |
COLT | 4 |
| 2007 | How SVMs can estimate quantiles and the medianabstractWe investigate quantile regression based on the pinball loss and the ǫ-insensitive loss. For the pinball loss a condition on the data-generating distribution P is given that ensures that the conditional quantiles are approximated with respect to k · k1. This result is then used to derive an oracle inequality for an SVM based on the pinball loss. Moreover, we show that SVMs based on the ǫ-insensitive loss estimate the conditional median only under certain conditions on P . Andreas Christmann, Ingo Steinwart |
NIPS | 2 |
| 2007 | Stability of Unstable Learning Algorithms
Don R. Hush, Clint Scovel, Ingo Steinwart |
Mach. Learn. | 3 |
| 2006 | Function Classes That Approximate the Bayes Risk
Ingo Steinwart, Don R. Hush, Clint Scovel |
COLT | 1 |
| 2006 | An Oracle Inequality for Clipped Regularized Risk MinimizersabstractWe establish a general oracle inequality for clipped approximate minimizers of regularized empirical risks and apply this inequality to support vector machine (SVM) type algorithms. We then show that for SVMs using Gaussian RBF kernels for classification this oracle inequality leads to learning rates that are faster than the ones established in [9]. Finally, we use our oracle inequality to show that a simple parameter selection approach based on a validation set can yield the same fast learning rates without knowing the noise exponents which were required to be known a-priori in [9]. Ingo Steinwart, Don R. Hush, Clint Scovel |
NIPS | 1 |
| 2006 | QP Algorithms with Guaranteed Accuracy and Run Time for Support Vector MachinesabstractWe describe polynomial--time algorithms that produce approximate solutions with guaranteed accuracy for a class of QP problems that are used in the design of support vector machine classifiers. These algorithms employ a two--stage process where the first stage produces an approximate solution to a dual QP problem and the second stage maps this approximate dual solution to an approximate primal solution. For the second stage we describe an O(n log n) algorithm that maps an approximate dual solution with accuracy (2(2Km)1/2+8(λ)1/2)-2 λ εp2 to an approximate primal solution with accuracy εp where n is the number of data samples, Kn is the maximum kernel value over the data and λ > 0 is the SVM regularization parameter. For the first stage we present new results for decomposition algorithms and describe new decomposition algorithms with guaranteed accuracy and run time. In particular, for τ-rate certifying decomposition algorithms we establish the optimality of τ = 1/(n-1). In addition we extend the recent τ = 1/(n-1) algorithm of Simon (2004) to form two new composite algorithms that also achieve the τ = 1/(n-1) iteration bound of List and Simon (2005), but yield faster run times in practice. We also exploit the τ-rate certifying property of these algorithms to produce new stopping rules that are computationally efficient and that guarantee a specified accuracy for the approximate dual solution. Furthermore, for the dual QP problem corresponding to the standard classification problem we describe operational conditions for which the Simon and composite algorithms possess an upper bound of O(n) on the number of iterations. For this same problem we also describe general conditions for which a matching lower bound exists for any decomposition algorithm that uses working sets of size 2. For the Simon and composite algorithms we also establish an O(n2) bound on the overall run time for the first stage. Combining the first and second stages gives an overall run time of O(n2(ck + 1)) where ck is an upper bound on the computation to perform a kernel evaluation. Pseudocode is presented for a complete algorithm that inputs an accuracy εp and produces an approximate solution that satisfies this accuracy in low order polynomial time. Experiments are included to illustrate the new stopping rules and to compare the Simon and composite decomposition algorithms. Don R. Hush, Patrick Kelly, Clint Scovel, Ingo Steinwart |
J. Mach. Learn. Res. | 4 |
| 2006 | An Explicit Description of the Reproducing Kernel Hilbert Spaces of Gaussian RBF KernelsabstractAlthough Gaussian radial basis function (RBF) kernels are one of the most often used kernels in modern machine learning methods such as support vector machines (SVMs), little is known about the structure of their reproducing kernel Hilbert spaces (RKHSs). In this work, two distinct explicit descriptions of the RKHSs corresponding to Gaussian RBF kernels are given and some consequences are discussed. Furthermore, an orthonormal basis for these spaces is presented. Finally, it is discussed how the results can be used for analyzing the learning performance of SVMs. Ingo Steinwart, Don R. Hush, Clint Scovel |
IEEE Trans. Inf. Theory | 1 |
| 2005 | Fast Rates for Support Vector Machines
Ingo Steinwart, Clint Scovel |
COLT | 1 |
| 2005 | A Classification Framework for Anomaly DetectionabstractOne way to describe anomalies is by saying that anomalies are not concentrated. This leads to the problem of finding level sets for the data generating density. We interpret this learning problem as a binary classification problem and compare the corresponding classification risk with the standard performance measure for the density level problem. In particular it turns out that the empirical classification risk can serve as an empirical performance measure for the anomaly detection problem. This allows us to compare different anomaly detection algorithms empirically, i.e. with the help of a test set. Furthermore, by the above interpretation we can give a strong justification for the well-known heuristic of artificially sampling "labeled" samples, provided that the sampling plan is well chosen. In particular this enables us to propose a support vector machine (SVM) for anomaly detection for which we can easily establish universal consistency. Finally, we report some experiments which compare our SVM to other commonly used methods including the standard one-class SVM. Ingo Steinwart, Don R. Hush, Clint Scovel |
J. Mach. Learn. Res. | 1 |
| 2005 | Consistency of support vector machines and other regularized kernel classifiersabstractIt is shown that various classifiers that are based on minimization of a regularized risk are universally consistent, i.e., they can asymptotically learn in every classification task. The role of the loss functions used in these algorithms is considered in detail. As an application of our general framework, several types of support vector machines (SVMs) as well as regularization networks are treated. Our methods combine techniques from stochastics, approximation theory, and functional analysis Ingo Steinwart |
IEEE Trans. Inf. Theory | 1 |
| 2004 | Density Level Detection is ClassificationabstractWe show that anomaly detection can be interpreted as a binary classifi- cation problem. Using this interpretation we propose a support vector machine (SVM) for anomaly detection. We then present some theoret- ical results which include consistency and learning rates. Finally, we experimentally compare our SVM with the standard one-class SVM. Ingo Steinwart, Don R. Hush, Clint Scovel |
NIPS | 1 |
| 2004 | Fast Rates to Bayes for Kernel MachinesabstractWe establish learning rates to the Bayes risk for support vector machines (SVMs) with hinge loss. In particular, for SVMs with Gaussian RBF kernels we propose a geometric condition for distributions which can be used to determine approximation properties of these kernels. Finally, we compare our methods with a recent paper of G. Blanchard et al.. Ingo Steinwart, Clint Scovel |
NIPS | 1 |
| 2004 | On Robustness Properties of Convex Risk Minimization Methods for Pattern Recognition
Andreas Christmann, Ingo Steinwart |
J. Mach. Learn. Res. | 2 |
| 2003 | Sparseness of Support Vector Machines---Some Asymptotically Sharp BoundsabstractThe decision functions constructed by support vector machines (SVM’s) usually depend only on a subset of the training set—the so-called support vectors. We derive asymptotically sharp lower and upper bounds on the number of support vectors for several standard types of SVM’s. In par- ticular, we show for the Gaussian RBF kernel that the fraction of support vectors tends to twice the Bayes risk for the L1-SVM, to the probability of noise for the L2-SVM, and to 1 for the LS-SVM. Ingo Steinwart |
NIPS | 1 |
| 2003 | Sparseness of Support Vector Machines
Ingo Steinwart |
J. Mach. Learn. Res. | 1 |
| 2003 | On the Optimal Parameter Choice for v-Support Vector MachinesabstractWe determine the asymptotically optimal choice of the parameter /spl nu/ for classifiers of /spl nu/-support vector machine (/spl nu/-SVM) type which has been introduced by Scholkopf et al. (2000). It turns out that /spl nu/ should be a close upper estimate of twice the optimal Bayes risk provided that the classifier uses a so-called universal kernel such as the Gaussian RBF kernel. Moreover, several experiments show that this result can be used to implement some modified cross validation procedures which improve standard cross validation for /spl nu/-SVMs. Ingo Steinwart |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Support Vector Machines are Universally Consistent
Ingo Steinwart |
J. Complex. | 1 |
| 2001 | On the Influence of the Kernel on the Consistency of Support Vector Machines
Ingo Steinwart |
J. Mach. Learn. Res. | 1 |