Sandro Ridella

dblp:43/5919 · DBLP profile ↗
← Back
101ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0003-0612-8219ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 93 · 11 first-author · 10 since 2021Systems, architecture and hardware · 6 · 1 first-authorTheory of computation · 2
YearPublicationVenuePosition
2026 Reconciling grokking with statistical learning theory through the lens of norm- and stability-based generalization bounds
abstract
In recent years, Artificial Intelligence, particularly Machine Learning, has achieved remarkable success in solving complex problems. However, this progress has also revealed the emergence of unexpected, poorly understood, and elusive phenomena that characterize the behavior of machine intelligence and learning processes. These phenomena often challenge researchers to interpret them within the boundaries of existing Machine Learning theoretical frameworks, thereby motivating the development of new and more comprehensive theoretical foundations. One such phenomenon, known as grokking , refers to the sudden and substantial improvement in a model’s performance following a prolonged period of stagnant or even regressive learning. In this paper, we argue that it is possible to provide insights into grokking by leveraging the existing theoretical foundations of Machine Learning, in particular concepts from Statistical Learning Theory, such as norm-based and stability-based generalization bounds. We further show how these theories can help reconcile the phenomenon of grokking with established principles of learning and generalization. Furthermore, we demonstrate the practical applicability of these insights through concrete examples.
Luca Oneto, Sandro Ridella, Simone Minisi, Andrea Coraddu, Davide Anguita
Neurocomputing2
2025 Reconciling Grokking with Statistical Learning Theory
abstract
In recent years, Artificial Intelligence, particularly Machine Learning (ML), has demonstrated remarkable success in addressing complex problems.However, this progress has been accompanied by the emergence of unexpected, poorly understood, and elusive phenomena that characterize the behavior of machine intelligence and learning processes.Researchers are often challenged to interpret these phenomena within the existing theoretical frameworks of ML, fostering a search for more complex or technical explanations.One such phenomenon, known as "grokking", occurs when an ML model, after a long period of stagnant or even regressive learning, suddenly exhibits rapid and substantial improvement.In this paper, we argue that grokking can be explained with the theoretical foundations of ML by leveraging Statistical Learning Theory, i.e., Algorithmic Stability theory.We provide insights into how this theory can reconcile grokking with established principles of learning and generalization. * This work is partially supported by (i
Luca Oneto, Sandro Ridella, Andrea Coraddu, Davide Anguita
ESANN2
2025 Informed Machine Learning: Excess risk and generalization
abstract
Machine Learning (ML) has transformed both research and industry by offering powerful models capable of capturing complex phenomena. However, these models often require large, high-quality datasets and may struggle to generalize beyond the distributions on which they are trained. Informed Machine Learning (IML) tackles these challenges by incorporating domain knowledge at various stages of the ML pipeline, thereby reducing data requirements and enhancing generalization. Building on statistical learning theory, we present some theoretical comparison and insights about ML and IML excess risk and generalization performance. We then illustrate how these theoretical insights can be leveraged in practice through some practical examples. Our findings shed some light on the mechanisms and conditions under which IML can outperform traditional ML, offering valuable guidance for effective implementation in real-world settings. • ML-based predictive models have greatly reshaped research, industry and society. • Informed ML leverages prior knowledge to reduce data demands and boost extrapolation. • We compare ML and Informed ML in terms of excess risk and generalization. • Informed ML can surpass ML under conditions favoring domain-specific insights.
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing2
2024 Informed Machine Learning: Excess Risk and Generalization
abstract
Machine Learning (ML) based predictive models are impacting research, industry, and society at large thanks to their ability to model or surrogate real systems.Two of the main current limitations of ML are the need for large amounts of high quality data and low performance far away from the observed data.For this reason, in certain applications where prior knowledge is available, researchers have developed Informed ML (IML) to decrease ML high quality data voracity and increase ML extrapolation abilities.In this work we study the differences between ML and IML excess risk and generalization using also some examples to elucidate the theoretical discussions.Our findings shed some light on the mechanisms and the conditions under which IML outperforms ML.
Luca Oneto, Davide Anguita, Sandro Ridella
ESANN3
2024 Towards algorithms and models that we can trust: A theoretical perspective
abstract
In the last decade it became increasingly apparent the inability of technical metrics such as accuracy, sustainability, and non-regressiveness to well characterize the behavior of intelligent systems. In fact, they are nowadays requested to meet also ethical requirements such as explainability, fairness, robustness, and privacy increasing our trust in their use in the wild. Of course often technical and ethical metrics are in tension between each other but the final goal is to be able to develop a new generation of more responsible and trustworthy machine learning. In this paper, we focus our attention on machine learning algorithms and associated predictive models, questioning for the first time, from a theoretical perspective, if it is possible to simultaneously guarantee their performance in terms of both technical and ethical metrics towards machine learning algorithms that we can trust. In particular, we will investigate for the first time both theory and practice of deterministic and randomized algorithms and associated predictive models showing the advantages and disadvantages of the different approaches. For this purpose we will leverage the most recent advances coming from the statistical learning theory: Complexity-Based Methods, Distribution Stability, PAC-Bayes, and Differential Privacy. Results will show that it is possible to develop consistent algorithms which generate predictive models with guarantees on multiple trustworthiness metrics.
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing2
2023 Towards Randomized Algorithms and Models that We Can Trust: a Theoretical Perspective
abstract
In the last decade it became increasingly apparent the inability of technical metrics to well characterize the behavior of intelligent systems.In fact, they are nowadays requested to meet also ethical requirements such as explainability, fairness, robustness, and privacy increasing our trust in their use in the wild.The final goal is to be able to develop a new generation of more responsible and trustworthy machine learning.In this paper, we focus our attention on randomized machine learning algorithms and models questioning, from a theoretical perspective, if it is possible to simultaneously optimize multiple metrics that are in tension between each other towards randomized machine learning algorithms that we can trust.For this purpose we will leverage the most recent advances coming from the statistical learning theory: distribution stability and differential privacy.* This work is supported in
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2023 Do we really need a new theory to understand over-parameterization?
abstract
This century saw an unprecedented increase of public and private investments in Artificial Intelligence (AI) and especially in (Deep) Machine Learning (ML). This led to breakthroughs in their practical ability to solve complex real-world problems impacting research and society at large. Instead, our ability to understand the fundamental mechanism behind these breakthroughs has slowed down because of their increased complexity, while in the past breakthroughs often emerged from foundational research. This questioned researchers about the necessity for a new theoretical framework able to help researchers catch up on this lag. One of the still not well understood mechanisms is the so-called over-parametrization, namely the ability of certain models to increase their generalization performance (reduce test error) when the number of parameters is above the interpolating threshold (zero training error). In this paper we will show that this phenomenon can be better understood using both known theories (surveying them in the process) and empirical evidences for both shallow and deep learning algorithms.
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing2
2022 Do We Really Need a New Theory to Understand the Double-Descent?
abstract
This century saw an unprecedented increase of public and private investments in Artificial Intelligence (AI) and especially in Machine Learning (ML).This led to breakthroughs in their practical ability to solve complex real world problems impacting research and society at large.Instead, our ability to understand the fundamental mechanism behind these breakthroughs has slowed down because of their increased complexity.This questioned researchers about the necessity for a new theoretical framework able to help researchers catch up on this lag.One of the still not well understood mechanisms is the so called over-parametrization, namely the ability of certain models to increasing their generalization performance (reduce test error) when the number of parameters is above the interpolating threshold (zero training error), and the associated doubledescent curve.In this paper we will show that this phenomena can be better understood using both known theories, i.e., the algorithmic stability theory, and empirical evidence.
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2022 The benefits of adversarial defense in generalization
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing2
2021 The Benefits of Adversarial Defence in Generalisation
abstract
Recent researches have been shown that models induced by machine learning, in particular by deep learning, can be easily fooled by an adversary who carefully crafts imperceptible, at least from the human perspective, or physically plausible modifications of the input data.This discovery gave birth to a new field of research, the adversarial machine learning, where new methods of attacks and defence are developed continuously, mimicking what is happening from a long time in cybersecurity.In this paper we will show that the drawbacks of inducing models from data less prone to be misled actually provides some benefits when it comes to assess their generalisation abilities.
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2020 Improving the Union Bound: a Distribution Dependent Approach
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2019 Local Rademacher Complexity Machine
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing2
2018 Local Rademacher Complexity Machine
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2018 Randomized learning: Generalization performance of old and new theoretically grounded algorithms
Luca Oneto, Francesca Cipollini, Sandro Ridella, Davide Anguita
Neurocomputing3
2018 Learning With Kernels: A Local Rademacher Complexity-Based Analysis With Application to Graph Kernels
abstract
When dealing with kernel methods, one has to decide which kernel and which values for the hyperparameters to use. Resampling techniques can address this issue but these procedures are time-consuming. This problem is particularly challenging when dealing with structured data, in particular with graphs, since several kernels for graph data have been proposed in literature, but no clear relationship among them in terms of learning properties is defined. In these cases, exhaustive search seems to be the only reasonable approach. Recently, the global Rademacher complexity (RC) and local Rademacher complexity (LRC), two powerful measures of the complexity of a hypothesis space, have shown to be suited for studying kernels properties. In particular, the LRC is able to bound the generalization error of an hypothesis chosen in a space by disregarding those ones which will not be taken into account by any learning procedure because of their high error. In this paper, we show a new approach to efficiently bound the RC of the space induced by a kernel, since its exact computation is an NP-Hard problem. Then we show for the first time that RC can be used to estimate the accuracy and expressivity of different graph kernels under different parameter configurations. The authors' claims are supported by experimental results on several real-world graph data sets.
Luca Oneto, Nicolò Navarin, Michele Donini, Sandro Ridella, Alessandro Sperduti, Fabio Aiolli, Davide Anguita
IEEE Trans. Neural Networks Learn. Syst.4
2017 Generalization Performances of Randomized Classifiers and Algorithms built on Data Dependent Distributions
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2017 Differential privacy and generalization: Sharper bounds with applications
Luca Oneto, Sandro Ridella, Davide Anguita
Pattern Recognit. Lett.2
2016 Tuning the Distribution Dependent Prior in the PAC-Bayes Framework based on Empirical Data
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN2
2016 Tikhonov, Ivanov and Morozov regularization for support vector machine learning
Luca Oneto, Sandro Ridella, Davide Anguita
Mach. Learn.2
2016 A local Vapnik-Chervonenkis complexity
Luca Oneto, Davide Anguita, Sandro Ridella
Neural Networks3
2016 Global Rademacher Complexity Bounds: From Slow to Fast Convergence Rates
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
Neural Process. Lett.3
2016 PAC-bayesian analysis of distribution dependent priors: Tighter risk bounds and stability analysis
Luca Oneto, Davide Anguita, Sandro Ridella
Pattern Recognit. Lett.3
2016 Learning Hardware-Friendly Classifiers Through Algorithmic Stability
abstract
Most state-of-the-art machine-learning (ML) algorithms do not consider the computational constraints of implementing the learned model on embedded devices. These constraints are, for example, the limited depth of the arithmetic unit, the memory availability, or the battery capacity. We propose a new learning framework, the Algorithmic Risk Minimization (ARM), which relies on Algorithmic-Stability, and includes these constraints inside the learning process itself. ARM allows one to train advanced resource-sparing ML models and to efficiently deploy them on smart embedded systems. Finally, we show the advantages of our proposal on a smartphone-based Human Activity Recognition application by comparing it to a conventional ML approach.
Luca Oneto, Sandro Ridella, Davide Anguita
ACM Trans. Embed. Comput. Syst.2
2015 Shrinkage learning to improve SVM with hints
abstract
The Support Vector Machine (SVM) is one of the most effective and used algorithms, when targeting classification. Despite its large success, SVM is mainly afflicted by two issues: (i) some hyperparameters must be tuned in advance and are, in practice, identified through computationally intensive procedures; (ii) possible a-priori knowledge about the problem (e.g. doctor expertise in medical applications) cannot be straightforwardly exploited. In this paper, we introduce a new approach, able to cope with the two previous problems: several experiments, performed on real-world benchmarking datasets, show that our method outperforms, on average, other techniques proposed in the literature.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IJCNN3
2015 Support vector machines and strictly positive definite kernel: The regularization hyperparameter is more important than the kernel hyperparameters
abstract
When dealing with a Support Vector Machine (SVM) with a strictly positive definite kernel, a common misconception is that the main handle for controlling the nonlinearity of the classification surface is the set of kernel hyperparameters. We show here that this is not the case: in particular, we prove that, regardless of the value of the kernel hyperparameter, it is always possible to tune the nonlinearity of the classifier by acting only on the regularization hyperparameter C, even achieving perfect learning of any non-degenerate training set.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IJCNN3
2015 Fast convergence of extended Rademacher Complexity bounds
abstract
In this work we propose some new generalization bounds for binary classifiers, based on global Rademacher Complexity (RC), which exhibit fast convergence rates by combining state-of-the-art results by Talagrand on empirical processes and the exploitation of unlabeled patterns. In this framework, we are able to improve both the constants and the convergence rates of existing RC-based bounds. All the proposed bounds are based on empirical quantities, so that they can be easily computed in practice, and are provided both in implicit and explicit forms: the formers are the tightest ones, while the latter ones allow to get more insights about the impact of Talagrand's results and the exploitation of unlabeled patterns in the learning process. Finally, we verify the quality of the bounds, with respect to the theoretical limit, showing the room for further improvements in the common scenario of binary classification.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IJCNN3
2015 Learning Resource-Aware Classifiers for Mobile Devices: From Regularization to Energy Efficiency
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
Neurocomputing3
2015 Local Rademacher Complexity: Sharper risk bounds with and without unlabeled samples
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
Neural Networks3
2015 Fully Empirical and Data-Dependent Stability-Based Bounds
abstract
The purpose of this paper is to obtain a fully empirical stability-based bound on the generalization ability of a learning procedure, thus, circumventing some limitations of the structural risk minimization framework. We show that assuming a desirable property of a learning algorithm is sufficient to make data-dependency explicit for stability, which, instead, is usually bounded only in an algorithmic-dependent way. In addition, we prove that a well-known and widespread classifier, like the support vector machine (SVM), satisfies this condition. The obtained bound is then exploited for model selection purposes in SVM classification and tested on a series of real-world benchmarking datasets demonstrating, in practice, the effectiveness of our approach.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IEEE Trans. Cybern.3
2014 Learning with few bits on small-scale devices: From regularization to energy efficiency
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN4
2014 Smartphone battery saving by bit-based hypothesis spaces and local Rademacher Complexities
abstract
Smartphones emerge from the incorporation of new services and features into mobile phones, allowing to implement advanced functionalities for the final users. The implementation of Machine Learning (ML) algorithms on the smartphone itself, without resorting to remote computing systems, allow to achieve such goals without expensive data transmission. However, smartphones are resource-limited devices and, as such, suffer from many issues, which are typical of stand-alone devices, such as limited battery capacity and processing power. We show in this paper how to build a thrifty classifier by exploiting bit-based hypothesis spaces and local Rademacher Complexities. The resulting classifier is tested on a real-world Human Activity Recognition application, implemented on a Samsung Galaxy S II smartphone.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN4
2014 Unlabeled patterns to tighten Rademacher complexity error bounds for kernel classifiers
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
Pattern Recognit. Lett.4
2014 A Deep Connection Between the Vapnik-Chervonenkis Entropy and the Rademacher Complexity
abstract
In this paper, we derive a deep connection between the Vapnik-Chervonenkis (VC) entropy and the Rademacher complexity. For this purpose, we first refine some previously known relationships between the two notions of complexity and then derive new results, which allow computing an admissible range for the Rademacher complexity, given a value of the VC-entropy, and vice versa. The approach adopted in this paper is new and relies on the careful analysis of the combinatorial nature of the problem. The obtained results improve the state of the art on this research topic.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IEEE Trans. Neural Networks Learn. Syst.4
2013 A Learning Machine with a Bit-Based Hypothesis Space
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN4
2013 A Novel Procedure for Training L1-L2 Support Vector Machine Classifiers
Davide Anguita, Alessandro Ghio, Luca Oneto, Jorge Luis Reyes-Ortiz, Sandro Ridella
ICANN5
2013 Some results about the Vapnik-Chervonenkis entropy and the rademacher complexity
abstract
This paper deals with the problem of identifying a connection between the Vapnik-Chervonenkis (VC) Entropy, a notion of complexity introduced by Vapnik in his seminal work, and the Rademacher Complexity, a more powerful notion of complexity, which has been in the limelight of several works in the recent Machine Learning literature. In order to establish this connection, we refine some previously known relationships and derive a new result. Our proposal allows computing an admissible range for the Rademacher Complexity, given a value of the VC-Entropy, and vice versa, therefore opening new appealing research perspectives in the field of assessing the complexity of an hypothesis space.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN4
2013 A support vector machine classifier from a bit-constrained, sparse and localized hypothesis space
abstract
Choosing an appropriate hypothesis space in classification applications, according to the Structural Risk Minimization (SRM) principle, is of paramount importance to train effective models: in fact, properly selecting the the space complexity allows to optimize the learned functions performance. This selection is not straightforward, especially (though not solely) when few samples are available for deriving an effective model (e.g. in bioinformatics applications). In this paper, by exploiting a bit-based definition for Support Vector Machine (SVM) classifiers, selected from an hypothesis space described according to sparsity and locality principles, we show how the complexity of the corresponding space of functions can be effectively tuned through the number of bits used for the function representation. Real world datasets are exploited to show how the number of bits and the degree of sparsity/locality imposed to define the hypothesis space affect the complexity of the space of classifiers and, consequently, the performance of the model, picked up from this set.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN4
2013 An improved analysis of the Rademacher data-dependent bound using its self bounding property
Luca Oneto, Alessandro Ghio, Davide Anguita, Sandro Ridella
Neural Networks4
2012 The 'K' in K-fold Cross Validation
Davide Anguita, Luca Ghelardoni, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN5
2012 Structural Risk Minimization and Rademacher Complexity for Regression
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN4
2012 Nested Sequential Minimal Optimization for Support Vector Machines
Alessandro Ghio, Davide Anguita, Luca Oneto, Sandro Ridella, Carlotta Schatten
ICANN (2)4
2012 Rademacher Complexity and Structural Risk Minimization: An Application to Human Gene Expression Datasets
Luca Oneto, Davide Anguita, Alessandro Ghio, Sandro Ridella
ICANN (2)4
2012 Learning the mean: A neural network approach
Sergio Decherchi, Mauro Parodi, Sandro Ridella
Neurocomputing3
2012 In-sample Model Selection for Trimmed Hinge Loss Support Vector Machine
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
Neural Process. Lett.4
2012 In-Sample and Out-of-Sample Model Selection and Error Estimation for Support Vector Machines
abstract
In-sample approaches to model selection and error estimation of support vector machines (SVMs) are not as widespread as out-of-sample methods, where part of the data is removed from the training set for validation and testing purposes, mainly because their practical application is not straightforward and the latter provide, in many cases, satisfactory results. In this paper, we survey some recent and not-so-recent results of the data-dependent structural risk minimization framework and propose a proper reformulation of the SVM learning algorithm, so that the in-sample approach can be effectively applied. The experiments, performed both on simulated and real-world datasets, show that our in-sample approach can be favorably compared to out-of-sample methods, especially in cases where the latter ones provide questionable results. In particular, when the number of samples is small compared to their dimensionality, like in classification of microarray data, our proposal can outperform conventional out-of-sample approaches such as the cross validation, the leave-one-out, or the Bootstrap methods.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IEEE Trans. Neural Networks Learn. Syst.4
2011 Maximal Discrepancy vs. Rademacher Complexity for error estimation
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN4
2011 In-sample model selection for Support Vector Machines
abstract
In-sample model selection for Support Vector Machines is a promising approach that allows using the training set both for learning the classifier and tuning its hyperparameters. This is a welcome improvement respect to out-of-sample methods, like cross-validation, which require to remove some samples from the training set and use them only for model selection purposes. Unfortunately, in-sample methods require a precise control of the classifier function space, which can be achieved only through an unconventional SVM formulation, based on Ivanov regularization. We prove in this work that, even in this case, it is possible to exploit well-known Quadratic Programming solvers like, for example, Sequential Minimal Optimization, so improving the applicability of the in-sample approach.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN4
2011 Selecting the hypothesis space for improving the generalization ability of Support Vector Machines
abstract
The Structural Risk Minimization framework has been recently proposed as a practical method for model selection in Support Vector Machines (SVMs). The main idea is to effectively measure the complexity of the hypothesis space, as defined by the set of possible classifiers, and to use this quantity as a penalty term for guiding the model selection process. Unfortunately, the conventional SVM formulation defines a hypothesis space centered at the origin, which can cause undesired effects on the selection of the optimal classifier. We propose here a more flexible SVM formulation, which addresses this drawback, and describe a practical method for selecting more effective hypothesis spaces, leading to the improvement of the generalization ability of the final classifier.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN4
2011 The Impact of Unlabeled Patterns in Rademacher Complexity Theory for Kernel Classifiers
abstract
We derive here new generalization bounds, based on Rademacher Complexity theory, for model selection and error estimation of linear (kernel) classifiers, which exploit the availability of unlabeled samples. In particular, two results are obtained: the first one shows that, using the unlabeled samples, the confidence term of the conventional bound can be reduced by a factor of three; the second one shows that the unlabeled samples can be used to obtain much tighter bounds, by building localized versions of the hypothesis class containing the optimal classifier.
Luca Oneto, Davide Anguita, Alessandro Ghio, Sandro Ridella
NIPS4
2011 Maximal Discrepancy for Support Vector Machines
Davide Anguita, Alessandro Ghio, Sandro Ridella
Neurocomputing3
2010 Maximal Discrepancy for Support Vector Machines
Davide Anguita, Alessandro Ghio, Sandro Ridella
ESANN3
2010 Model selection for support vector machines: Advantages and disadvantages of the Machine Learning Theory
abstract
A common belief is that Machine Learning Theory (MLT) is not very useful, in pratice, for performing effective SVM model selection. This fact is supported by experience, because well-known hold-out methods like cross-validation, leave-one-out, and the bootstrap usually achieve better results than the ones derived from MLT. We show in this paper that, in a small sample setting, i.e. when the dimensionality of the data is larger than the number of samples, a careful application of the MLT can outperform other methods in selecting the optimal hyperparameters of a SVM.
Davide Anguita, Alessandro Ghio, Noemi Greco, Luca Oneto, Sandro Ridella
IJCNN5
2010 A neural model approach for regularization in the mean estimation case
abstract
Neural Networks are powerful tools for function approximation problems. A possible peculiar application of neural networks is that proposed here: estimating the univariate mean of a distribution from a finite sample. This problem characterizes a huge number of applicative and scientific problems. The Gaussian distribution case is analyzed, however the proposed analysis is of general validity and can be easily extended to other distributions. In particular the estimation problem is approached as a regularization problem and a solution to the selection of the regularization parameter is obtained via the employment of neural models. The paper, after introducing some theoretical results, presents two neural models, namely a MLP and a Circular Back Propagation Network, for the mean prediction. Experimental results show that neural networks can estimate the mean, in expectation, better than the usual sample mean formula.
Sergio Decherchi, Mauro Parodi, Sandro Ridella
IJCNN3
2010 Using unsupervised analysis to constrain generalization bounds for support vector classifiers
abstract
A crucial issue in designing learning machines is to select the correct model parameters. When the number of available samples is small, theoretical sample-based generalization bounds can prove effective, provided that they are tight and track the validation error correctly. The maximal discrepancy (MD) approach is a very promising technique for model selection for support vector machines (SVM), and estimates a classifier's generalization performance by multiple training cycles on random labeled data. This paper presents a general method to compute the generalization bounds for SVMs, which is based on referring the SVM parameters to an unsupervised solution, and shows that such an approach yields tight bounds and attains effective model selection. When one estimates the generalization error, one uses an unsupervised reference to constrain the complexity of the learning machine, thereby possibly decreasing sharply the number of admissible hypothesis. Although the methodology has a general value, the method described in the paper adopts vector quantization (VQ) as a representation paradigm, and introduces a biased regularization approach in bound computation and learning. Experimental results validate the proposed method on complex real-world data sets.
Sergio Decherchi, Sandro Ridella, Rodolfo Zunino, Paolo Gastaldo, Davide Anguita
IEEE Trans. Neural Networks2
2008 Using Variable Neighborhood Search to improve the Support Vector Machine performance in embedded automotive applications
abstract
In this work we show that a metaheuristic, the variable neighborhood search (VNS), can be effectively used in order to improve the performance of the hardware-friendly version of the support vector machine (SVM). Our target is the implementation of the feed-forward phase of SVM on resource-limited hardware devices, such as field programmable gate arrays (FPGAs) and digital signal processors (DSPs). The proposal has been tested on a machine-vision benchmark dataset for embedded automotive applications, showing considerable performance improvements respect to previously used techniques.
Enrique Alba 0001, Davide Anguita, Alessandro Ghio, Sandro Ridella
IJCNN4
2008 A support vector machine with integer parameters
Davide Anguita, Alessandro Ghio, Stefano Pischiutta, Sandro Ridella
Neurocomputing4
2007 A Hardware-friendly Support Vector Machine for Embedded Automotive Applications
abstract
We present here a hardware-friendly version of the support vector machine (SVM), which is useful to implement its feed-forward phase on limited-resources devices such as field programmable gate arrays (FPGAs) or microcontrollers, where a floating-point unit is seldom available. Our proposal is tested on a machine-vision benchmark dataset for automotive applications.
Davide Anguita, Alessandro Ghio, Stefano Pischiutta, Sandro Ridella
IJCNN4
2006 Testing the Augmented Binary Multiclass SVM on Microarray Data
abstract
In this paper we test a new multicategory SVM method, called augmented binary (AB), on microarray gene expression data. The AB SVM is one of the methods generating a multicategory classifier in one step, without dividing the multiclass problem into binary subproblems. This approach can be useful when the number of samples is very low, like in this kind of application. Furthermore, the use of a single SVM, instead of several binary ones, simplifies the search for optimal hyperparameters and allows a consistent output for all the classes.
Davide Anguita, Sandro Ridella, Dario Sterpi
IJCNN2
2006 Feed-Forward Support Vector Machine Without Multipliers
abstract
In this letter, we propose a coordinate rotation digital computer (CORDIC)-like algorithm for computing the feed-forward phase of a support vector machine (SVM) in fixed-point arithmetic, using only shift and add operations and avoiding resource-consuming multiplications. This result is obtained thanks to a hardware-friendly kernel, which greatly simplifies the SVM feed-forward phase computation and, at the same time, maintains good classification performance respect to the conventional Gaussian kernel.
Davide Anguita, Stefano Pischiutta, Sandro Ridella, Dario Sterpi
IEEE Trans. Neural Networks3
2005 K-fold generalization capability assessment for support vector classifiers
abstract
The problem of how to effectively implement k-fold cross-validation for support vector machines is considered. Indeed, despite the fact that this selection criterion is widely used due to its reasonable requirements in terms of computational resources and its good ability in identifying a well performing model, it is not clear how one should employ the committee of classifiers coming from the k folds for the task of on-line classification. Three methods are here described and tested, based respectively on: averaging, random choice and majority voting. Each of these methods is tested on a wide range of data-sets for different fold settings.
Davide Anguita, Sandro Ridella, Fabio Rivieccio
IJCNN2
2004 Unsupervised clustering and the capacity of support vector machines
abstract
In the framework of support vector machine (SVM) classifiers, an unsupervised analysis of empirical data supports an ordering criterion for the families of possible functions. The approach enhances the structural risk minimization paradigm by sharply reducing the number of admissible classifiers, thus tightening the associate generalization bound. The paper shows that kernel-based algorithms, allowing efficient optimization, can support both the unsupervised clustering process and the generalization-error estimation. The main result of this sample-based method may be a dramatic reduction in the predicted generalization error, as demonstrated by experiments on synthetic testbeds as well as real-world problems.
Davide Anguita, Sandro Ridella, Fabio Rivieccio, Rodolfo Zunino
IJCNN2
2004 A new method for multiclass support vector machines
abstract
In this paper we present a new method for solving multiclass problems with a support vector machine. Our method compares favorably with other proposals, appeared so far in the literature, both in terms of computational needs for the feedforward phase and of classification accuracy. The main result, however, is the mapping of the multiclass problem to a biclass one, which allows us to suggest a method for estimating the generalization error by using data-dependent error bounds.
Davide Anguita, Sandro Ridella, Dario Sterpi
IJCNN2
2004 Model selection in top quark tagging with a support vector classifier
abstract
The problem of tagging a top quark generation event in data coming from the collider detector at Fermilab is considered and tackled through the use of a support vector machine classifier. In order to select a fitting model, a twofold procedure has been adopted. The SVC hyperparameters have been selected through the bootstrap technique and then an additional tuning of the bias value and the error relevance has been performed by means both of a purity vs. efficiency curve and of the AUC value. The generalization capability of the model has been evaluated using the maximal discrepancy criterion.
Davide Anguita, Sandro Ridella, Silvia Amerio, Ignazio Lazzizzera
IJCNN2
2004 Vector quantization complexity and quantum computing
abstract
A dichotomy between 'analogue' modeling and 'digital' implementation is often encountered when designing vector quantizers. In the case of digital systems, the requirement of optimality can bring about NP-hard problems. The paper discusses the possibility of using advanced paradigms such as quantum computing for digital optimization processes in order to overcome the limitations of conventional machinery. The presented research provides analytical criteria determining the relative advantages of conventional over quantum-computing approaches.
Paolo Gastaldo, Sandro Ridella, Rodolfo Zunino
IJCNN2
2004 Using K-Winner Machines for domain analysis
Sandro Ridella, Rodolfo Zunino
Neurocomputing1
2003 SVM learning with fixed-point math
abstract
We present in this paper an algorithm for Support Vector Machine (SVM) learning, which can be implemented using fixed-point math. The advantages of the fixed-point representation, respect to the more common floating-point one, allows us to address digital VLSI implementations of SVM. In particular, simple algorithms and simple architectures can be exploited for targeting programmable devices like Field Programmable Gate Arrays (FPGAs), which are the basis of many embedded systems. This paper focuses on the SVM learning algorithm: for the complete version of this work, including an actual FPGA realization.
Davide Anguita, Andrea Boni, Sandro Ridella
IJCNN3
2003 Training support vector machines: a quantum-computing perspective
abstract
Recent advances in characterizing the generalization ability of support vector machines (SVMs) exploit refined concepts, such as Rademacher estimates of model complexity and nonlinear criteria for weighting empirical errors. Those methods improve the SVM representation ability and tighten generalization bounds. On the other hand, quadratic-programming algorithms are no longer applicable, hence the SVM-training process cannot benefit from the notable efficiency featured by those specialized techniques. The paper considers the possibility of using quantum computing to solve the resulting problem of effective optimization, especially in the case of digital SV implementations. The behavioral aspects of conventional and enhanced SVMs are compared, supported by experiments in both a synthetic and a real-world problem. Likewise, the related differences between quadratic-programming and quantum-based optimization techniques are analyzed.
Davide Anguita, Sandro Ridella, Fabio Rivieccio, Rodolfo Zunino
IJCNN2
2003 A model-selection approach to the VLSI design of vector quantizers
abstract
A formal methodology supports the design of digital services for hierarchical vector quantization (HVQ). A model-selection approach based on the minimum description length criterion is enhanced by circuit-related aspects allowing efficient design. The resulting parameters drive the subsequent digital VLSI realization, which yields a HVQ chip providing cost-effective, computationally efficient real-time performances. Real-world applications support the consistency of the VQ approach and the effectiveness of the HVQ device.
Massimiliano Bracco, Sandro Ridella, Rodolfo Zunino
IJCNN2
2003 Hyperparameter design criteria for support vector classifiers
Davide Anguita, Sandro Ridella, Fabio Rivieccio, Rodolfo Zunino
Neurocomputing2
2003 Quantum optimization for training support vector machines
Davide Anguita, Sandro Ridella, Fabio Rivieccio, Rodolfo Zunino
Neural Networks2
2003 A digital architecture for support vector machines: theory, algorithm, and FPGA implementation
abstract
In this paper, we propose a digital architecture for support vector machine (SVM) learning and discuss its implementation on a field programmable gate array (FPGA). We analyze briefly the quantization effects on the performance of the SVM in classification problems to show its robustness, in the feedforward phase, respect to fixed-point math implementations; then, we address the problem of SVM learning. The architecture described here makes use of a new algorithm for SVM learning which is less sensitive to quantization errors respect to the solution appeared so far in the literature. The algorithm is composed of two parts: the first one exploits a recurrent network for finding the parameters of the SVM; the second one uses a bisection process for computing the threshold. The architecture implementing the algorithm is described in detail and mapped on a real current-generation FPGA (Xilinx Virtex II). Its effectiveness is then tested on a channel equalization problem, where real-time performances are of paramount importance.
Davide Anguita, Andrea Boni, Sandro Ridella
IEEE Trans. Neural Networks3
2003 Digital implementation of hierarchical vector quantization
abstract
A formal methodology drives the design and realization of a digital very large-scale integration (VLSI) device supporting hierarchical vector quantization (HVQ) in computation-intensive coding applications. The hardware-oriented model-selection approach enhances the Minimum Description Length criterion with circuit-related aspects that allow consistent and efficient design. The resulting model parameters drive the subsequent realization in digital circuitry, which has first been implemented in field-programmable gate array (FPGA) technology to verify its correctness. The eventual VLSI realization results in an HVQ chip providing cost-effective, computationally efficient real-time performances. Real-world applications support the consistency of the vector quantization approach and the effectiveness of the HVQ device.
Massimiliano Bracco, Sandro Ridella, Rodolfo Zunino
IEEE Trans. Neural Networks2
2002 Automatic Hyperparameter Tuning for Support Vector Machines
Davide Anguita, Sandro Ridella, Fabio Rivieccio, Rodolfo Zunino
ICANN2
2001 K-winner machines for pattern classification
abstract
The paper describes the K-winner machine (KWM) model for classification. KWM training uses unsupervised vector quantization and subsequent calibration to label data-space partitions. A K-winner classifier seeks the largest set of best-matching prototypes agreeing on a test pattern, and provides a local-level measure of confidence. A theoretical analysis characterizes the growth function of a K-winner classifier, and the result leads to tight bounds to generalization performance. The method proves suitable for high-dimensional multiclass problems with large amounts of data. Experimental results on both a synthetic and a real domain (NIST handwritten numerals) confirm the approach effectiveness and the consistency of the theoretical framework.
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
IEEE Trans. Neural Networks1
2001 Empirical measure of multiclass generalization performance: the K-winner machine case
abstract
Combining the K-winner machine (KWM) model with empirical measurements of a classifier's Vapnik-Chervonenkis (VC)-dimension gives two major results. First, analytical derivations refine the theory that characterizes the generalization performances of binary classifiers. Second, a straightforward extension of the theoretical framework yields bounds to the generalization error for multiclass problems.
Sandro Ridella, Rodolfo Zunino
IEEE Trans. Neural Networks1
2000 The K-Winner Machine Model
abstract
A K-Winner Machine (KWM) selects among a family of classifiers the specific configuration that minimizes the expected generalization error. In training, KWM uses unsupervised vector quantization and subsequent calibration to label data-space partitions. At run time, KWM seeks the largest set of best-matching prototypes agreeing on a test sample, and provides a local-level measure of confidence. The VC-dim of a KWM classifier is worked out exactly; the resulting small values set tight bounds to generalization performance. The network can be applied to high-dimensional, multi-class problems with large data sets. Experimental results in both a synthetic and a real domain (NIST handwritten numerals) validate the consistency of the theoretical framework.
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
IJCNN (1)1
2000 Augmenting vector quantization with interval arithmetics for image-coding applications
abstract
Interval Arithmetic (IA) augments the basic Vector-Quantization (VQ) paradigm for image compression. The reformulated VQ scheme allows prototypes to assume ranges of admissible locations rather than be clamped to specific space positions. The image-reconstruction process exploits the resulting degrees of freedom to make up for the excessive discretization (such as blockiness) that often affects VQ-based coding. The paper describes the algorithms for both the training and the run-time use of IAVQ codebooks; the possibility of data-driven training endows the proposed methodology with the flexibility and adaptiveness of standard VQ methods, as confirmed by experimental results on real images.
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
ISCAS1
2000 Digital VLSI Algorithms and Architectures for Support Vector Machines
abstract
In this paper, we propose some very simple algorithms and architectures for a digital VLSI implementation of Support Vector Machines. We discuss the main aspects concerning the realization of the learning phase of SVMs, with special attention on the effects of fixed-point math for computing and storing the parameters of the network. Some experiments on two classification problems are described that show the efficiency of the proposed methods in reaching optimal solutions with reasonable hardware requirements.
Davide Anguita, Andrea Boni, Sandro Ridella
Int. J. Neural Syst.3
2000 Evaluating the Generalization Ability of Support Vector Machines through the Bootstrap
Davide Anguita, Andrea Boni, Sandro Ridella
Neural Process. Lett.3
1999 A VLSI friendly algorithm for support vector machines
abstract
We propose a VLSI friendly algorithm for the implementation of the learning phase of support vector machines (SVM). Differently from previous methods, that rely on sophisticated constrained nonlinear programming algorithms, our approach finds a simple updating rule that can be easily implemented in digital VLSI.
Davide Anguita, Andrea Boni, Sandro Ridella
IJCNN3
1999 Possibility and Necessity Pattern Classification using an Interval Arithmetic Perceptron
Gian Paolo Drago, Sandro Ridella
Neural Comput. Appl.2
1999 Worst case analysis of weight inaccuracy effects in multilayer perceptrons
abstract
We derive here a new method for the analysis of weight quantization effects in multilayer perceptrons based on the application of interval arithmetic. Differently from previous results, we find worst case bounds on the errors due to weight quantization, that are valid for every distribution of the input or weight values. Given a trained network, our method allows to easily compute the minimum number of bits needed to encode its weights.
Davide Anguita, Sandro Ridella, Stefano Rovetta
IEEE Trans. Neural Networks2
1999 Representation and generalization properties of class-entropy networks
abstract
Using conditional class entropy (CCE) as a cost function allows feedforward networks to fully exploit classification-relevant information. CCE-based networks arrange the data space into partitions, which are assigned unambiguous symbols and are labeled by class information. By this labeling mechanism the network can model the empirical data distribution at the local level. Region labeling evolves with the network-training process, which follows a plastic algorithm. The paper proves several theoretical properties about the performance of CCE-based networks, and considers both convergence during training and generalization ability at run-time. In addition, analytical criteria and practical procedures are proposed to enhance the generalization performance of the trained networks. Experiments on artificial and real-world domains confirm the accuracy of this class of networks and witness the validity of the described methods.
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
IEEE Trans. Neural Networks1
1999 Circular backpropagation networks embed vector quantization
abstract
This letter proves the equivalence between vector quantization (VQ) classifiers and circular backpropagation (CBP) networks. The calibrated prototypes for a VQ schema can be plugged in a CBP feedforward structure having the same number of hidden neurons and featuring the same mapping. The letter describes how to exploit such equivalence by using VQ prototypes to perform a meaningful initialization for BP optimization. The approach effectiveness was tested considering a real classification problem (NIST handwritten digits).
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
IEEE Trans. Neural Networks1
1998 Pruning with interval arithmetic perceptron
Gian Paolo Drago, Sandro Ridella
Neurocomputing2
1998 Plastic Algorithm for Adaptive Vector Quantisation
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
Neural Comput. Appl.1
1997 Circular backpropagation networks for classification
abstract
The class of mapping networks is a general family of tools to perform a wide variety of tasks. This paper presents a standardized, uniform representation for this class of networks, and introduces a simple modification of the multilayer perceptron with interesting practical properties, especially well suited to cope with pattern classification tasks. The proposed model unifies the two main representation paradigms found in the class of mapping networks for classification, namely, the surface-based and the prototype-based schemes, while retaining the advantage of being trainable by backpropagation. The enhancement in the representation properties and the generalization performance are assessed through results about the worst-case requirement in terms of hidden units and about the Vapnik-Chervonenkis dimension and cover capacity. The theoretical properties of the network also suggest that the proposed modification to the multilayer perceptron is in many senses optimal. A number of experimental verifications also confirm theoretical results about the model's increased performances, as compared with the multilayer perceptron and the Gaussian radial basis functions network.
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
IEEE Trans. Neural Networks1
1996 On the convergence of a growing topology neural algorithm
Gian Paolo Drago, Sandro Ridella
Neurocomputing2
1995 Learning the appropriate representation paradigm by circular processing units
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
ESANN1
1995 An Adaptive Momentum Back Propagation (AMBP)
Gian Paolo Drago, Mauro Morando, Sandro Ridella
Neural Comput. Appl.3
1995 Adaptive Internal Representation in Circular Back-Propagation Networks
Sandro Ridella, Stefano Rovetta, Rodolfo Zunino
Neural Comput. Appl.1
1994 Convergence Properties of Cascade Correlation in Function Approximation
Gian Paolo Drago, Sandro Ridella
Neural Comput. Appl.2
1994 Class-Entropy Minimisation Networks for Domain Analysis and Rule Extraction
Sandro Ridella, Gian Luca Speroni, Paolo Trebino, Rodolfo Zunino
Neural Comput. Appl.1
1993 Using chaos to generate keys for associative noise-like coding memories
Giancarlo Parodi, Sandro Ridella, Rodolfo Zunino
Neural Networks2
1992 A dedicated massively parallel architecture for the Boltzmann machine
Alessandro De Gloria, Paolo Faraboschi, Sandro Ridella
Parallel Comput.3
1992 Statistically controlled activation weight initialization (SCAWI)
abstract
An optimum weight initialization which strongly improves the performance of the back propagation (BP) algorithm is suggested. By statistical analysis, the scale factor, R (which is proportional to the maximum magnitude of the weights), is obtained as a function of the paralyzed neuron percentage (PNP). Also, by computer simulation, the performances on the convergence speed have been related to PNP. An optimum range for R is shown to exist in order to minimize the time needed to reach the minimum of the cost function. Normalization factors are properly defined, which leads to a distribution of the activations independent of the neurons, and to a single nondimensional quantity, R, the value of which can be quickly found by computer simulation.
Gian Paolo Drago, Sandro Ridella
IEEE Trans. Neural Networks2
1991 Efficient computation of the correlation dimension from a time series on a LIW computer
Angelo Corana, Aldo Casaleggio, Claudia Rolando, Sandro Ridella
Parallel Comput.4
1989 Corrigenda: "Minimizing Multimodal Functions of Continuous Variables with the 'Simulated Annealing' Algorithm"
abstract
No abstract available.
Angelo Corana, Claudio Martini, Sandro Ridella
ACM Trans. Math. Softw.3
1988 Solving linear equation systems on vector computers with maximum efficiency
A. Corona, Claudio Martini, Mauro Morando, Sandro Ridella, Claudia Rolando
Parallel Comput.4
1987 LU Factorization with Maximum Performances on FPS Architectures 38/64 BIT
Angelo Corana, Claudio Martini, Sandro Ridella, Claudia Rolando
ICS3
1987 Minimizing multimodal functions of continuous variables with the "simulated annealing" algorithm
abstract
A new global optimization algorithm for functions of continuous variables is presented, derived from the “Simulated Annealing” algorithm recently introduced in combinatorial optimization. The algorithm is essentially an iterative random search procedure with adaptive moves along the coordinate directions. It permits uphill moves under the control of a probabilistic criterion, thus tending to avoid the first local minima encountered. The algorithm has been tested against the Nelder and Mead simplex method and against a version of Adaptive Random Search. The test functions were Rosenbrock valleys and multiminima functions in 2,4, and 10 dimensions. The new method proved to be more reliable than the others, being always able to find the optimum, or at least a point very close to it. It is quite costly in term of function evaluations, but its cost can be predicted in advance, depending only slightly on the starting point.
Angelo Corana, Michele Marchesi, Claudio Martini, Sandro Ridella
ACM Trans. Math. Softw.4