Luca Oneto

dblp:91/10065 · DBLP profile ↗
← Back
141ranked-venue papers
66as first author
53since 2021 · last 2026
0000-0002-8445-395XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 130 · 62 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Theory of computation · 6 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models
abstract
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model’s latent space; e.g., computed as the difference between the centroids of harmful and harmless prompt representations. However, emerging evidence suggests that concepts in LLMs often appear to be encoded as a low-dimensional manifold embedded in the high-dimensional latent space. Motivated by these findings, we propose a novel method leveraging Self-Organizing Maps (SOMs) to extract multiple refusal directions. To this end, we first prove that SOMs generalize the prior work's difference-in-means technique. We then train SOMs on harmful prompt representations to identify multiple neurons. By subtracting the centroid of harmless representations from each neuron, we derive a set of multiple directions expressing the refusal concept. We validate our method on an extensive experimental setup, demonstrating that ablating multiple directions from models' internals outperforms not only the single-direction baseline but also specialized jailbreak algorithms, leading to an effective suppression of refusal. Finally, we conclude by analyzing the mechanistic implications of our approach.
Giorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, Battista Biggio
AAAI4
2026 Multi-label Complementary Labels Learning under Hard Logical Constraints
abstract
Two of the main challenges in multi-label classification are the need to collect labeled data, which can be costly or impractical, and the need to satisfy hard logical constraints between labels, which is often computationally expensive.In some applications, complementary labelsthat is, labels specifying a class to which a sample does not belong -are available and much less costly to obtain.Researchers have therefore developed methods to learn from such labels efficiently and effectively.Similar efforts have been made to address the problem of learning with hard logical constraints.Nevertheless, to the best of our knowledge, no prior work has investigated the problem of learning from complementary labels with hard logical constraints.In this work, we propose and compare methods to address this problem, showing that hard logical constraints, besides representing restrictions to be satisfied, can also serve as an additional source of weak supervision.The relationships between labels can help bridge the information gap between relevant and complementary labels.Experimental results on different datasets and scenarios support our claims.
Luca Oneto, Davide Anguita, Fabio Roli, Min-Ling Zhang, Fulvio Mastrogiovanni
ESANN1
2026 Neuro Symbolic AI and Complex Data
Luca Oneto, Nicolò Navarin, Luca Pasa, Davide Rigoni 0001, Davide Anguita
ESANN1
2026 Ensembling Post-Hoc Image Explanations: When It Works, When It Fails, and How to Tell the Difference
abstract
Post-hoc explanation methods of image recognition models often exhibit high variance or disagreement across explanations when the input data are perturbed, the underlying models are modified, or different explainability techniques are employed.To mitigate this issue, several approaches have been proposed, among which ensemble strategies that aggregate multiple explanations have attracted particular attention.Although some of these methods demonstrate good empirical performance, most existing works remain largely empirical, with limited theoretical justification or understanding of why ensemble strategies work and when they fail.In this paper, we analyze the factors that influence the success and failure of ensemble strategies that combines multiple explanations, using different datasets, convolutional neural network architectures, posthoc explanation techniques, and ensembling strategies to identify the most influential image patches.In particular, we compare various ensembling strategies based on distinct voting principles -namely, Borda Count, Kemeny-Young, Reciprocal Rank Fusion, and the Schulze method -and show that the performance of such ensemble methods depends on the degree of satisfaction of their underlying theoretical assumptions.
Luca Oneto, Jinhua Xu, Davide Anguita, Fabio Roli, Jing Yuan 0001
ESANN1
2026 Optimal In-Station Train Dispatching via Symbolic Pattern Planning
abstract
The Optimal In-Station Train Dispatching (InSTraDi) problem consists in commanding the movements of trains inside a railway station while both (i) respecting safety, time, and travel constraints and (ii) minimizing delays. In Symbolic Pattern Planning (SPP), a pattern, suggesting the sequence of happenings to reach the goal, is encoded in a logic formula whose models correspond to valid plans. If no valid plan is found, the pattern is extended until it covers a valid plan. However, plans of better quality could exist if we had continued extending the pattern. In this paper, we formalize the InSTraDi problem as a Temporal Planning Task with Intermediate Conditions and Effects, and we show an InSTraDi-dependent way to construct, in polynomial time, a pattern ensuring the optimal plan can be found by the SPP approach without never extending the pattern. Analysis on realistic railway data validate our approach.
Matteo Cardellini, Enrico Giunchiglia, Davide Anguita, Carmelo Lofiego, Luca Oneto, Pietro Ratto
KR5
2026 Informed machine learning for complex data
abstract
Machine Learning (ML) has become a central force in Artificial Intelligence, driving major breakthroughs in applications that handle increasingly complex data, from images and text sequences to graph structures. While new architectures such as Transformers and Graph Neural Networks continue to redefine performance benchmarks in various domains, these predominantly data-driven methods often neglect critical domain knowledge, practical constraints, and broader contextual factors. This oversight diminishes their trustworthiness and restricts their impact in real-world settings. In this paper, we discuss the need for a more informed approach to ML for complex data. Specifically, we advocate for solutions that explicitly integrate structural awareness to capture underlying relationships in the data, incorporate key technical requirements to ensure safety and compliance with industry standards, embed environmental considerations to promote sustainability and resource efficiency, adhere to established physical principles, and uphold ethical and societal values. By weaving these dimensions together, informed ML can bridge the gap between purely data-centric methods and the nuanced demands of practical applications. We show how this integrated framework not only strengthens model performance but also ensures that ML solutions remain trustworthy, efficient, and sensitive to human ecological, ethical, and regulatory imperatives. Our discussion underscores the transformative potential of Informed ML to drive innovation across diverse domains, setting a new benchmark for responsible and high-impact ML system design.
Luca Oneto, Nicolò Navarin, Alessio Micheli, Luca Pasa, Claudio Gallicchio, Davide Bacciu, Davide Anguita
Neurocomputing1
2026 Reconciling grokking with statistical learning theory through the lens of norm- and stability-based generalization bounds
abstract
In recent years, Artificial Intelligence, particularly Machine Learning, has achieved remarkable success in solving complex problems. However, this progress has also revealed the emergence of unexpected, poorly understood, and elusive phenomena that characterize the behavior of machine intelligence and learning processes. These phenomena often challenge researchers to interpret them within the boundaries of existing Machine Learning theoretical frameworks, thereby motivating the development of new and more comprehensive theoretical foundations. One such phenomenon, known as grokking , refers to the sudden and substantial improvement in a model’s performance following a prolonged period of stagnant or even regressive learning. In this paper, we argue that it is possible to provide insights into grokking by leveraging the existing theoretical foundations of Machine Learning, in particular concepts from Statistical Learning Theory, such as norm-based and stability-based generalization bounds. We further show how these theories can help reconcile the phenomenon of grokking with established principles of learning and generalization. Furthermore, we demonstrate the practical applicability of these insights through concrete examples.
Luca Oneto, Sandro Ridella, Simone Minisi, Andrea Coraddu, Davide Anguita
Neurocomputing1
2026 HORNET: Fast and minimal adversarial perturbations
abstract
Fixed-budget attacks aim to generate adversarial examples—carefully crafted inputs designed to induce misclassifications during inference—while adhering to a predefined perturbation budget. These attacks maximize misclassification confidence and benefit from the transferability property, enabling the generated adversarial examples to remain effective even against multiple unknown models. However, to preserve their transferability, such attacks often yield perceptible perturbations, compromising the visual integrity of the adversarial examples. In this paper, we introduce HORNET, an extension of gradient-based fixed-budget attacks designed to minimize the perturbation magnitude of adversarial examples while maintaining their transferability against the target model. HORNET utilizes a distinct source model to craft the adversarial examples and employs a limited number of queries to the unknown target model to further minimize perturbation magnitude. We evaluate HORNET empirically by integrating it with 41 existing attack implementations and testing it against 9 different models, resulting in a total of 1700 unique configurations. Our results demonstrate that HORNET outperforms the state of the art in generating minimally perturbed yet highly transferable adversarial examples across all tested models. Code available at: https://github.com/louiswup/HORNET .
Jiaping Wu, Antonio Emanuele Cinà, Francesco Villani, Zhaoqiang Xia, Luca Demetrio, Luca Oneto, Davide Anguita, Fabio Roli, Xiaoyi Feng
Inf. Sci.6
2026 Poison once, fool many: Practical poisoning attacks against text-to-image retrieval systems
abstract
Text-to-Image retrieval (IR) systems are widely used to match images to specific textual queries, often leveraging publicly available Vision-Language Pretrained models (VLPs) for their generalization capabilities. However, due to the diverse and open nature of the image data they rely on, these systems remain vulnerable to data poisoning attacks, where malicious images are injected into the database to manipulate retrieval results. Prior work has demonstrated the effectiveness of attacks when the exact user query is known at retrieval time. However, this assumption is often impractical, as users tend to express similar intents using varied, semantically equivalent queries (e.g., through synonyms), which reduces the effectiveness of existing attacks. In this paper, we address this gap by proposing an attack that remains effective even when users issue semantically varied queries. We introduce Collisio , a novel poisoning method that crafts a single poisoned image to be retrieved under any semantically equivalent form of a target query. To achieve this, Collisio leverages an Expectation over Queries (EoQ) strategy, generating a diverse set of synthetic and selectively transformed query variants, and then optimizes the poisoned image to align with them. We extensively evaluate Collisio on the Flickr30k and MSCOCO datasets across multiple VLPs, demonstrating the severity of Collisio under realistic query variations. Given the implications of this vulnerability, we examine countermeasures based on adversarially trained models and a data preprocessing defense, highlighting both their mitigation potential and the trade-offs involved.
Dario Lazzaro, Raffaele Mura, Antonio Emanuele Cinà, Giuseppe Laurita, Gianni Viardo Vercelli, Luca Oneto, Battista Biggio, Fabio Roli
Knowl. Based Syst.6
2025 Reconciling Grokking with Statistical Learning Theory
abstract
In recent years, Artificial Intelligence, particularly Machine Learning (ML), has demonstrated remarkable success in addressing complex problems.However, this progress has been accompanied by the emergence of unexpected, poorly understood, and elusive phenomena that characterize the behavior of machine intelligence and learning processes.Researchers are often challenged to interpret these phenomena within the existing theoretical frameworks of ML, fostering a search for more complex or technical explanations.One such phenomenon, known as "grokking", occurs when an ML model, after a long period of stagnant or even regressive learning, suddenly exhibits rapid and substantial improvement.In this paper, we argue that grokking can be explained with the theoretical foundations of ML by leveraging Statistical Learning Theory, i.e., Algorithmic Stability theory.We provide insights into how this theory can reconcile grokking with established principles of learning and generalization. * This work is partially supported by (i
Luca Oneto, Sandro Ridella, Andrea Coraddu, Davide Anguita
ESANN1
2025 TransferBench: Benchmarking Ensemble-based Black-box Transfer Attacks
abstract
Ensemble-based black-box transfer attacks optimize adversarial examples on a set of surrogate models, claiming to reach high success rates by querying the (unknown) target model only a few times. In this work, we show that prior evaluations are systematically biased, as such methods are tested only under overly optimistic scenarios, without considering (i) how the choice of surrogate models influences transferability, (ii) how they perform against robust target models, and (iii) whether querying the target to refine the attack is really required.To address these gaps, we introduce TransferBench, a framework for evaluating ensemble-based black-box transfer attacks under more realistic and challenging scenarios than prior work. Our framework considers 17 distinct settings on CIFAR-10 and ImageNet, including diverse surrogate-target combinations, robust targets, and comparisons to baseline methods that do not use any query-based refinement mechanism. Our findings reveal that existing methods fail to generalize to more challenging scenarios, and that query-based refinement offers little to no benefit, contradicting prior claims. These results highlight that building reliable and query-efficient black-box transfer attacks remains an open challenge. We release our benchmark and evaluation code at: https://github.com/pralab/transfer-bench.
Fabio Brau, Maura Pintor, Antonio Emanuele Cinà, Raffaele Mura, Luca Scionis, Luca Oneto, Fabio Roli, Battista Biggio
NeurIPS6
2025 Training-Free Constrained Generation With Stable Diffusion Models
abstract
Stable diffusion models represent the state-of-the-art in data synthesis across diverse domains and hold transformative potential for applications in science and engineering, e.g., by facilitating the discovery of novel solutions and simulating systems that are computationally intractable to model explicitly. While there is increasing effort to incorporate physics-based constraints into generative models, existing techniques are either limited in their applicability to latent diffusion frameworks or lack the capability to strictly enforce domain-specific constraints. To address this limitation this paper proposes a novel integration of stable diffusion models with constrained optimization frameworks, enabling the generation of outputs satisfying stringent physical and functional requirements. The effectiveness of this approach is demonstrated through material design experiments requiring adherence to precise morphometric properties, challenging inverse design tasks involving the generation of materials inducing specific stress-strain responses, and copyright-constrained content generation tasks. All code has been released at https://github.com/RAISELab-atUVA/Constrained-Stable-Diffusion.
Stefano Zampini, Jacob Christopher, Luca Oneto, Davide Anguita, Ferdinando Fioretto
NeurIPS3
2025 Informed Machine Learning: Excess risk and generalization
abstract
Machine Learning (ML) has transformed both research and industry by offering powerful models capable of capturing complex phenomena. However, these models often require large, high-quality datasets and may struggle to generalize beyond the distributions on which they are trained. Informed Machine Learning (IML) tackles these challenges by incorporating domain knowledge at various stages of the ML pipeline, thereby reducing data requirements and enhancing generalization. Building on statistical learning theory, we present some theoretical comparison and insights about ML and IML excess risk and generalization performance. We then illustrate how these theoretical insights can be leveraged in practice through some practical examples. Our findings shed some light on the mechanisms and conditions under which IML can outperform traditional ML, offering valuable guidance for effective implementation in real-world settings. • ML-based predictive models have greatly reshaped research, industry and society. • Informed ML leverages prior knowledge to reduce data demands and boost extrapolation. • We compare ML and Informed ML in terms of excess risk and generalization. • Informed ML can surpass ML under conditions favoring domain-specific insights.
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing1
2025 Robustness-Congruent Adversarial Training for Secure Machine Learning Model Updates
abstract
Machine-learning models demand periodic updates to improve their average accuracy, exploiting novel architectures and additional data. However, a newly updated model may commit mistakes the previous model did not make. Such misclassifications are referred to as negative flips, experienced by users as a regression of performance. In this work, we show that this problem also affects robustness to adversarial examples, hindering the development of secure model update practices. In particular, when updating a model to improve its adversarial robustness, previously ineffective adversarial attacks on some inputs may become successful, causing a regression in the perceived security of the system. We propose a novel technique, named robustness-congruent adversarial training, to address this issue. It amounts to fine-tuning a model with adversarial training, while constraining it to retain higher robustness on the samples for which no adversarial example was found before the update. We show that our algorithm and, more generally, learning with non-regression constraints, provides a theoretically-grounded framework to train consistent estimators. Our experiments on robust models for computer vision confirm that both accuracy and robustness, even if improved after model update, can be affected by negative flips, and our robustness-congruent adversarial training can mitigate the problem, outperforming competing baseline methods.
Daniele Angioni, Luca Demetrio, Maura Pintor, Luca Oneto, Davide Anguita, Battista Biggio, Fabio Roli
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Eight quick tips for biologically and medically informed machine learning
abstract
Machine learning has become a powerful tool for computational analysis in the biomedical sciences, with its effectiveness significantly enhanced by integrating domain-specific knowledge. This integration has give rise to informed machine learning, in contrast to studies that lack domain knowledge and treat all variables equally (uninformed machine learning). While the application of informed machine learning to bioinformatics and health informatics datasets has become more seamless, the likelihood of errors has also increased. To address this drawback, we present eight guidelines outlining best practices for employing informed machine learning methods in biomedical sciences. These quick tips offer recommendations on various aspects of informed machine learning analysis, aiming to assist researchers in generating more robust, explainable, and dependable results. Even if we originally crafted these eight simple suggestions for novices, we believe they are deemed relevant for expert computational researchers as well.
Luca Oneto, Davide Chicco
PLoS Comput. Biol.1
2025 Nine quick tips for trustworthy machine learning in the biomedical sciences
abstract
As machine learning (ML) becomes increasingly central to biomedical research, the need for trustworthy models is more pressing than ever. In this paper, we present nine concise and actionable tips to help researchers build ML systems that are technically sound but ethically responsible, and contextually appropriate for biomedical applications. These tips address the multifaceted nature of trustworthiness, emphasizing the importance of considering all potential consequences, recognizing the limitations of current methods, taking into account the needs of all involved stakeholders, and following open science practices. We discuss technical, ethical, and domain-specific challenges, offering guidance on how to define trustworthiness and how to mitigate sources of untrustworthiness. By embedding trustworthiness into every stage of the ML pipeline - from research design to deployment - these recommendations aim to support both novice and experienced practitioners in creating ML systems that can be relied upon in biomedical science.
Luca Oneto, Davide Chicco
PLoS Comput. Biol.1
2024 Informed Machine Learning: Excess Risk and Generalization
abstract
Machine Learning (ML) based predictive models are impacting research, industry, and society at large thanks to their ability to model or surrogate real systems.Two of the main current limitations of ML are the need for large amounts of high quality data and low performance far away from the observed data.For this reason, in certain applications where prior knowledge is available, researchers have developed Informed ML (IML) to decrease ML high quality data voracity and increase ML extrapolation abilities.In this work we study the differences between ML and IML excess risk and generalization using also some examples to elucidate the theoretical discussions.Our findings shed some light on the mechanisms and the conditions under which IML outperforms ML.
Luca Oneto, Davide Anguita, Sandro Ridella
ESANN1
2024 Informed Machine Learning for Complex Data
abstract
In the contemporary era of data-driven decision-making, the application of Machine Learning (ML) on complex data (e.g., images, text, sequences, trees, and graphs) has become increasingly pivotal (e.g., Large Language Models and Graph Neural Networks).In this context, there is a gap between purely data-driven models and domain-specific knowledge, requirements, and expertise.In particular, this domain specificity needs to be integrated into the ML models to improve learning generalization, sustainability, trustworthiness, reliability, security, and safety.This additional knowledge can assume different forms, e.g.: software developers require ML to comply with many technical requirements, companies require ML to comply with economic and environmental sustainability, domain experts require ML to be aligned with physical and logical laws, and society requires ML to be aligned with ethical principles.This special session gathers valuable contributions and early findings in the field of Informed ML for Complex Data.Our main objective is to showcase the potential and limitations of new ideas, improvements, or the blending of ML and other research areas in solving real-world problems.
Luca Oneto, Nicolò Navarin, Alessio Micheli, Luca Pasa, Claudio Gallicchio, Davide Bacciu, Davide Anguita
ESANN1
2024 Mitigating Unfair Regression in Machine Learning Model Updates
abstract
Machine learning systems often require updates for various reasons, such as the availability of new data or models and the need to optimize different technical or ethical metrics. Typically, these metrics reflect an average performance rather than sample-wise behavior. Indeed, improvements in metrics like accuracy can introduce negative flips, where the updated model makes errors that the previous model did not make. In certain applications, these negative flips can be perceived by developers or users as a regression in performance, contributing to the hidden technical debt of machine learning systems. Moreover, if the distribution of negative flips is biased with respect to some sensitive attribute (e.g., gender or race), it may be perceived as discrimination, termed unfair regression. In this paper we show, for the first time, the existence of the phenomenon of unfair regression and propose different ethical metrics to measure it. Additionally, we offer two mitigation strategies - one focused on modifying the learning algorithm and one focused on modifying the tuning phase - to address this issue. Our results on real-world datasets confirm the existence of the unfair regression phenomenon and demonstrate the effectiveness of the proposed mitigation strategies.
Irene Buselli, Anna Pallarès López, Eduard Martín Jiménez, Davide Anguita, Fabio Roli, Luca Oneto
ICMLA6
2024 Toward Measuring and Understanding the Overvalidation Phenomena
abstract
Over the past decade, advances in Machine Learning have significantly expanded both the number and complexity of algorithms. Consequently, solving specific tasks now requires making numerous choices, such as selecting the appropriate algorithm, architecture, and corresponding hyper-parameters. Although various methods have been proposed to expedite this search process, the final selection is typically made using a holdout set. While this practical approach is widely accepted, it can lead to the issue of overvalidation when the number of choices is large. Overvalidation refers to the bias in holdout performance, which can result in the incorrect selection of the optimal choice so to be not aligned with the technical needs. This issue can be mitigated by improving the quantity and quality of data in the validation set or through resampling, but it reappears as the number of choices increases. Thus, the challenge of better understanding and detecting overvalidation to avoid or further mitigate it remains unresolved. In this paper, we address this problem by measuring and understanding the overvalidation phenomenon using statistical learning theory and testing it with real-world examples.
Fabrizio Mori, Antonio Emanuele Cinà, Fabio Roli, Davide Anguita, Luca Oneto
ICMLA5
2024 Investigating over-parameterized randomized graph networks
abstract
In this paper, we investigate neural models based on graph random features for classification tasks. First, we aim to understand when over parameterization, namely generating more features than the ones necessary to interpolate, may be beneficial for the generalization abilities of the resulting models. We employ two measures: one from the algorithmic stability framework and another one based on information theory. We provide empirical evidence from several commonly adopted graph datasets showing that the considered measures, even without considering task labels, can be effective for this purpose. Additionally, we investigate whether these measures can aid in the process of hyperparameters selection. The results of our empirical analysis show that the considered measures have good correlations with the estimated generalization performance of the models with different hyperparameter configurations. Moreover, they can be used to identify good hyperparameters, achieving results comparable to the ones obtained with a classic grid search.
Giovanni Donghi, Luca Pasa, Luca Oneto, Claudio Gallicchio, Alessio Micheli, Davide Anguita, Alessandro Sperduti, Nicolò Navarin
Neurocomputing3
2024 Fair graph representation learning: Empowering NIFTY via Biased Edge Dropout and Fair Attribute Preprocessing
abstract
The increasing complexity and amount of data available in modern applications strongly demand Trustworthy Learning algorithms that can be fed directly with complex and large graphs data. In fact, on one hand, machine learning models must meet high technical standards (e.g., high accuracy with limited computational requirements), but, at the same time, they must be sure not to discriminate against subgroups of the population (e.g., based on gender or ethnicity). Graph Neural Networks (GNNs) are currently the most effective solution to meet the technical requirements, even if it has been demonstrated that they inherit and amplify the biases contained in the data as a reflection of societal inequities. In fact, when dealing with graph data, these biases can be hidden not only in the node attributes but also in the connections between entities. Several Fair GNNs have been proposed in the literature, with uNIfying Fairness and stabiliTY (NIFTY) (Agarwal et al., 2021) being one of the most effective. In this paper, we will empower NIFTY’s fairness with two new strategies. The first one is a Biased Edge Dropout, namely, we drop graph edges to balance homophilous and heterophilous sensitive connections, mitigating the bias induced by subgroup node cardinality. The second one is Attributes Preprocessing, which is the process of learning a fair transformation of the original node attributes. The effectiveness of our proposal will be tested on a series of datasets with increasingly challenging scenarios. These scenarios will deal with different levels of knowledge about the entire graph, i.e., how many portions of the graph are known and which sub-portion is labelled at the training and forward phases.
Danilo Franco, Vincenzo Stefano D'Amato, Luca Pasa, Nicolò Navarin, Luca Oneto
Neurocomputing5
2024 Advances in artificial neural networks, machine learning and computational intelligence
Nicolò Navarin, Dounia Mulders, Luca Oneto
Neurocomputing3
2024 Towards algorithms and models that we can trust: A theoretical perspective
abstract
In the last decade it became increasingly apparent the inability of technical metrics such as accuracy, sustainability, and non-regressiveness to well characterize the behavior of intelligent systems. In fact, they are nowadays requested to meet also ethical requirements such as explainability, fairness, robustness, and privacy increasing our trust in their use in the wild. Of course often technical and ethical metrics are in tension between each other but the final goal is to be able to develop a new generation of more responsible and trustworthy machine learning. In this paper, we focus our attention on machine learning algorithms and associated predictive models, questioning for the first time, from a theoretical perspective, if it is possible to simultaneously guarantee their performance in terms of both technical and ethical metrics towards machine learning algorithms that we can trust. In particular, we will investigate for the first time both theory and practice of deterministic and randomized algorithms and associated predictive models showing the advantages and disadvantages of the different approaches. For this purpose we will leverage the most recent advances coming from the statistical learning theory: Complexity-Based Methods, Distribution Stability, PAC-Bayes, and Differential Privacy. Results will show that it is possible to develop consistent algorithms which generate predictive models with guarantees on multiple trustworthiness metrics.
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing1
2023 Short-term Forecast and Long-term Simulation for Accurate Energy Consumption Prediction
abstract
Accurate energy consumption forecasting has become pivotal for many companies as a way to tailor the budget dedicated to energy purchase on their actual power demand, thus sustainably minimizing energy waste and expenses. For these companies, both short-term and long-term energy consumption forecasts are a matter of interest since they would like to both program last-minute buy and sell and also plan future investments for power optimization. For this purpose, in this paper, different Deep Neural Networks techniques will be tested to perform both a supervised short-term energy consumption forecasting and an unsupervised long-term simulation via generative learning since very long-term forecasting (i.e., more than 1 year) is usually too inaccurate. The first task will be performed by adopting both a Recurrent Neural Network and a Long Short-Term Memory Network, while the second one will be performed by adopting a Generative Adversarial Network. Result on public data from the Australian Energy Market Operator will support the proposal.
Daniele Giampaoli, Francesca Cipollini, Denise Maffione, Luca Oneto
DSAA4
2023 Improving Fairness via Intrinsic Plasticity in Echo State Networks
abstract
Artificial Intelligence, and in particular Machine Learning, has become ubiquitous in today's society, both revolutionizing and impacting society as a whole.However, it can also lead to algorithmic bias and unfair results, especially when sensitive information is involved.This paper addresses the problem of algorithmic fairness in Machine Learning for temporal data, focusing on ensuring that sensitive time-dependent information does not unfairly influence the outcome of a classifier.In particular, we focus on a class of training-efficient recurrent neural models called Echo State Networks, and show, for the first time, how to leverage local unsupervised adaptation of the internal dynamics in order to build fairer classifiers.Experimental results on real-world problems from physiological sensor data demonstrate the potential of the proposal.
Andrea Ceni, Davide Bacciu, Valerio De Caro, Claudio Gallicchio, Luca Oneto
ESANN5
2023 Mitigating Robustness Bias: Theoretical Results and Empirical Evidences
abstract
Recent research has shown that some learned classifiers can be more easily fooled by an adversary who carefully crafts imperceptible or physically plausible modifications of the input data regarding particular subgroups of the population (e.g., people with particular gender, ethnicity, or skin color).This form of unfairness has been just recently studied, noting the fact that classical fairness metrics, which only observe the model outputs, are not enough but robustness biases need to be measured and mitigated as well.For this reason, in this paper, we will first develop a new metric of fairness which generalizes the current ones and degenerates in the classical ones and then we will develop a theoretical mitigation framework with consistency results able to generate a new empirical mitigation strategy and explain why the current ones actually work. * This work is supported in part
Danilo Franco, Luca Oneto, Davide Anguita
ESANN2
2023 An Empirical Study of Over-Parameterized Neural Models based on Graph Random Features
abstract
In this paper, we investigate neural models based on graph random features.In particular, we aim to understand when over-parameterization, namely generating more features than the ones necessary to interpolate, may be beneficial for the generalization of the resulting models.Exploiting the algorithmic stability framework and based on empirical evidences from several commonly adopted graph datasets, we will shed some light on this issue.
Nicolò Navarin, Luca Pasa, Luca Oneto, Alessandro Sperduti
ESANN3
2023 Towards Randomized Algorithms and Models that We Can Trust: a Theoretical Perspective
abstract
In the last decade it became increasingly apparent the inability of technical metrics to well characterize the behavior of intelligent systems.In fact, they are nowadays requested to meet also ethical requirements such as explainability, fairness, robustness, and privacy increasing our trust in their use in the wild.The final goal is to be able to develop a new generation of more responsible and trustworthy machine learning.In this paper, we focus our attention on randomized machine learning algorithms and models questioning, from a theoretical perspective, if it is possible to simultaneously optimize multiple metrics that are in tension between each other towards randomized machine learning algorithms that we can trust.For this purpose we will leverage the most recent advances coming from the statistical learning theory: distribution stability and differential privacy.* This work is supported in
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2023 Physically plausible propeller noise prediction via recursive corrections leveraging prior knowledge and experimental data
abstract
For propeller-driven vessels, cavitation is the most dominant noise source producing both structure-borne and radiated noise impacting wildlife, passenger comfort, and underwater warfare. Physically plausible and accurate predictions of the underwater radiated noise at design stage, i.e., for previously untested geometries and operating conditions, are fundamental for designing silent and efficient propellers. State-of-the-art predictive models are based on physical, data-driven, and hybrid approaches. Physical models (PMs) meet the need for physically plausible predictions but are either too computationally demanding or not accurate enough at design stage. Data-driven models (DDMs) are computationally inexpensive ad accurate on average but sometimes produce physically implausible results. Hybrid models (HMs) combine PMs and DDMs trying to take advantage of their strengths while limiting their weaknesses but state-of-the-art hybridisation strategies do not actually blend them, failing to achieve the HMs full potential. In this work, for the first time, we propose a novel HM that recursively correct a state-of-the-art PM by means of a DDM which simultaneously exploits the prior physical knowledge in the definition of its feature set and the data coming from a vast experimental campaign at the Emerson Cavitation Tunnel on the Meridian standard propeller series behind different severities of the axial wake. Results in different extrapolating conditions, i.e., extrapolation with respect to propeller rotational speed, wakefield, and geometry, will support our proposal both in terms of accuracy and physical plausibility.
Miltiadis Kalikatzarakis, Andrea Coraddu, Mehmet Atlar, Stefano Gaggero, Giorgio Tani, Luca Oneto
Eng. Appl. Artif. Intell.6
2023 Do we really need a new theory to understand over-parameterization?
abstract
This century saw an unprecedented increase of public and private investments in Artificial Intelligence (AI) and especially in (Deep) Machine Learning (ML). This led to breakthroughs in their practical ability to solve complex real-world problems impacting research and society at large. Instead, our ability to understand the fundamental mechanism behind these breakthroughs has slowed down because of their increased complexity, while in the past breakthroughs often emerged from foundational research. This questioned researchers about the necessity for a new theoretical framework able to help researchers catch up on this lag. One of the still not well understood mechanisms is the so-called over-parametrization, namely the ability of certain models to increase their generalization performance (reduce test error) when the number of parameters is above the interpolating threshold (zero training error). In this paper we will show that this phenomenon can be better understood using both known theories (surveying them in the process) and empirical evidences for both shallow and deep learning algorithms.
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing1
2023 On the problem of recommendation for sensitive users and influential items: Simultaneously maintaining interest and diversity
abstract
Recommender systems, in real-world circumstances, tend to limit user exposure to certain topics and to overexpose them to others to maximize performance. However, repeated exposure to biased content could lead to the so-called echo chamber phenomenon: especially in social network environments, people encounter only information that reflects their previous beliefs and opinions, reinforcing them. This phenomenon could have worrying consequences for society, including the spread of aggressive, unhealthy, or risky behaviors. Some persons can be more affected than others by echo-chambers. We define as sensitive the users whose behavior could be influenced by the over- or under-exposure to certain items due to the echo-chamber effect, and as influential the items that could influence the behavior of such users. In this paper, we address the problem of recommending influential items to sensitive users. We formalize the problem and propose three techniques that can be used to diversify the distributions of influential items in order to positively affect sensitive users’ behavior. Recommendations that meet this diversity criterion could potentially avoid dangerous societal consequences and simultaneously promote healthier lifestyles. We tested the proposed techniques in a real-world dataset by considering two different case studies that involved potentially aggressive and potentially depressed users. All techniques have been proven to be effective and allow high performance to be maintained while diversifying recommendations.
Alvise De Biasio, Merylin Monaro, Luca Oneto, Lamberto Ballan, Nicolò Navarin
Knowl. Based Syst.3
2022 Biased Edge Dropout in NIFTY for Fair Graph Representation Learning
abstract
Graph Neural Networks (GNNs) are nowadays widely used in many real-world applications.Nonetheless, the data relationships can be a source of biases based on sensitive attributes (e.g., gender or ethnicity).Several methods have been proposed to learn fair graph node representations.In this work we extend NIFTY, an approach that exploits additional terms in the loss function based on perturbing the input data to enforce the fairness of the GNNs.In particular, we exploit a biased perturbation of the adjacency matrix of the graph able to reduce the edge homophily.We show the effectiveness of our approach in four real-world graph datasets.
Federico Caldart, Luca Pasa, Luca Oneto, Alessandro Sperduti, Nicolò Navarin
ESANN3
2022 Simple Non Regressive Informed Machine Learning Model for Predictive Maintenance of Railway Critical Assets
abstract
Signals, track circuits, switches, and relay rooms are simultaneously the most critical and most maintained railway assets.A fault of one of these assets may strongly reduce the railway network capacity or even disrupt the circulation.Effectively predicting what assets may need maintenance allows to anticipate the intervention thus avoiding a failure.Currently, this problem is tackled by infrastructure managers mostly relying on operators' experience and with limited support of decision supporting tools.In this paper, we propose a Simple Informed Machine Learning (ML) based model able to automatically predict what asset need to be maintained fully leveraging on the operator experience.However, ML models in modern industrial MLOps pipelines demand continuous data collection, model re-training, testing, and monitoring, creating a large technical debt.In fact, one of the main requirements of these pipelines is to not be regressive, i.e., not simply improve average performances but also not incorrectly predicting an output that was correctly classified by the reference model (negative flips).In this work we face this problem by empowering the proposed ML with Non Regressive properties.Results on real data coming from a portion of an Italian Railway Network managed by Rete Ferroviaria Italiana, the Italian Infrastructure Manager, will support our proposal. * This research
Luca Oneto, Simone Minisi, Andrea Garrone, Renzo Canepa, Carlo Dambra, Davide Anguita
ESANN1
2022 Do We Really Need a New Theory to Understand the Double-Descent?
abstract
This century saw an unprecedented increase of public and private investments in Artificial Intelligence (AI) and especially in Machine Learning (ML).This led to breakthroughs in their practical ability to solve complex real world problems impacting research and society at large.Instead, our ability to understand the fundamental mechanism behind these breakthroughs has slowed down because of their increased complexity.This questioned researchers about the necessity for a new theoretical framework able to help researchers catch up on this lag.One of the still not well understood mechanisms is the so called over-parametrization, namely the ability of certain models to increasing their generalization performance (reduce test error) when the number of parameters is above the interpolating threshold (zero training error), and the associated doubledescent curve.In this paper we will show that this phenomena can be better understood using both known theories, i.e., the algorithmic stability theory, and empirical evidence.
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2022 The Importance of Multiple Temporal Scales in Motion Recognition: when Shallow Model can Support Deep Multi Scale Models
abstract
The execution of a human movement involves different muscles that are activated and coordinated by the brain at different temporal scales in a complex cognitive process. For this reason, studying human motion requires to properly model multiple temporal scales that fully describe its complexity. Current approaches are not able to address this requirement properly or are based on oversimplified models with obvious limitations. Data-driven models represent research frontiers able to provide new insights. In this work we will investigate different data-driven approaches. The first one is based on shallow models that, while achieving reasonably good recognition performance, require to handcraft features according to the domain knowledge The second one is based on deep models that can be extended to manage multiple temporal scales but they are hard to exploit as too many architecture configurations exist. For this reason, we will propose a new deep multiple temporal scale data-driven model, based on Temporal Convolutional Network, capable of learning features from the data at different temporal scales, of outperforming state of the art deep and shallow models, and of exploiting shallow models to tune the architecture configuration. We designed, collected data and tested our proposal in a specially devised experiment, to prove the validity of our approach. In particular, we collected motion capture data about dyad actions where two people exchange a ball. As the weight of the ball and the throwing intentions change, we will show how it is possible to automatically detect either the weight of the ball or the intention behind the throw just based on motion data. Data regarding our experiment and code of the methods proposed in this work are also made freely available to the research community. Results support both the proposal and the need for the use of deep multi scale models as a tool to better understand human movement and its multiple time scale nature.
Vincenzo Stefano D'Amato, Luca Oneto, Antonio Camurri, Davide Anguita
IJCNN2
2022 The Importance of Multiple Temporal Scales in Motion Recognition: from Shallow to Deep Multi Scale Models
abstract
Studying human motion requires modelling its multiple temporal scale nature to fully describe its complexity since different muscles are activated and coordinated by the brain at different temporal scales in a complex cognitive process. Nevertheless, current approaches are not able to address this requirement properly, and are based on oversimplified models with obvious limitations. Data-driven methods represent a viable tool to address these limitations. Nevertheless, shallow data-driven models, while achieving reasonably good recognition performance, require to handcraft features based on domain-specific knowledge which, in this cases, is limited and does no allow to properly model motion- and subject-specific temporal scales. In this work, we propose a new deep multiple temporal scale data-driven model, based on Temporal Convolutional Networks, able to automatically learn features from the data at different temporal scales. Our proposal focuses first on over-performing state-of-the-art shallows and deep models in terms of recognition performance. Then, thanks to the use of feature ranking for shallow models and an attention map for deep models, we will give insights on what the different architectures actually learned from the data. We designed, collected data, and tested our proposal in custom experiment of motion recognition: detecting the person who draw a particular shape (i.e., an ellipse) on a graphics tablet, collecting data about his/her movement (e.g., pressure and speed) in different extrapolating scenarios (e.g., training with data collected from one hand and testing the model on the other one). Collected data regarding our experiment and code of the methods are also made freely available to the research community. Results, both in terms of accuracy and insight on the cognitive problem, support the proposal and support the use of the proposed technique as a support tool for better understanding the human movements and its multiple temporal scale nature.
Vincenzo Stefano D'Amato, Luca Oneto, Antonio Camurri, Davide Anguita, Zinat Zarandi, Luciano Fadiga, Alessandro D'Ausilio, Thierry Pozzo
IJCNN2
2022 Deep Learning for the Generation of Heuristics in Answer Set Programming: A Case Study of Graph Coloring
Carmine Dodaro, Davide Ilardi, Luca Oneto, Francesco Ricca
LPNMR3
2022 Assessing Emotions in Human-Robot Interaction Based on the Appraisal Theory
abstract
Emotions have always played a crucial role in human evolution, improving not only social contact but also their ability to adapt and react to a changing environment. In the field of social robotics, providing robots with the ability to recognize human emotions through the interpretation of non-verbal signals may represent the key to more effective and engaging interaction. However, the problem of emotion recognition has usually been addressed in limited and static scenarios, by classifying emotions using sensory data such as facial expressions, body postures, and voice. This work proposes a novel emotion recognition framework, based on the appraisal theory of emotion. According to the theory, the expected person’s appraisal of a given situation depending on their needs and goals (henceforth referred to as "appraisal information") is combined with sensory data. A pilot experiment was designed and conducted: participants were involved in spontaneous verbal interaction with the humanoid robot Pepper, programmed to elicit different emotions in various moments. Then, a Random Forest classifier was trained to classify positive and negative emotions using: (i) sensor data only; (ii) sensor data supplemented by appraisal information. Preliminary results confirm a performance improvement in emotion classification when appraisal information is considered.
Marco Demutti, Vincenzo Stefano D'Amato, Carmine Tommaso Recchiuto, Luca Oneto, Antonio Sgorbissa
RO-MAN4
2022 Deep fair models for complex data: Graphs labeling and explainable face recognition
Danilo Franco, Nicolò Navarin, Michele Donini, Davide Anguita, Luca Oneto
Neurocomputing5
2022 Advances in artificial neural networks, machine learning and computational intelligence
abstract
Learning machines for structured data (e.g., trees) are intrinsically based on their capacity to learn representations by aggregating information from the multi-way relationships emerging from the structure topology. While complex aggregation functions are desirable in this context to increase the expressiveness of the learned representations, the modelling of higher-order interactions among structure constituents is unfeasible, in practice, due to the exponential number of parameters required. Therefore, the common approach is to define models which rely only on first-order interactions among structure constituents.In this work, we leverage tensors theory to define a framework for learning in structured domains. Such a framework is built on the observation that more expressive models require a tensor parameterisation. This observation is the stepping stone for the application of tensor decompositions in the context of recursive models. From this point of view, the advantage of using tensor decompositions is twofold since it allows limiting the number of model parameters while injecting inductive biases that do not ignore higher-order interactions.We apply the proposed framework on probabilistic and neural models for structured data, defining different models which leverage tensor decompositions. The experimental validation clearly shows the advantage of these models compared to first-order and full-tensorial models.
Luca Oneto, Kerstin Bunte, Nicolò Navarin
Neurocomputing1
2022 Towards learning trustworthily, automatically, and with guarantees on graphs: An overview
Luca Oneto, Nicolò Navarin, Battista Biggio, Federico Errica, Alessio Micheli, Franco Scarselli, Monica Bianchini, Luca Demetrio, Pietro Bongini, Armando Tacchella, Alessandro Sperduti
Neurocomputing1
2022 Advances in artificial neural networks, machine learning and computational intelligence
Luca Oneto, Nicolò Navarin, Frank-Michael Schleif
Neurocomputing1
2022 The benefits of adversarial defense in generalization
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing1
2022 Eleven quick tips for data cleaning and feature engineering
abstract
Applying computational statistics or machine learning methods to data is a key component of many scientific studies, in any field, but alone might not be sufficient to generate robust and reliable outcomes and results. Before applying any discovery method, preprocessing steps are necessary to prepare the data to the computational analysis. In this framework, data cleaning and feature engineering are key pillars of any scientific study involving data analysis and that should be adequately designed and performed since the first phases of the project. We call "feature" a variable describing a particular trait of a person or an observation, recorded usually as a column in a dataset. Even if pivotal, these data cleaning and feature engineering steps sometimes are done poorly or inefficiently, especially by beginners and unexperienced researchers. For this reason, we propose here our quick tips for data cleaning and feature engineering on how to carry out these important preprocessing steps correctly avoiding common mistakes and pitfalls. Although we designed these guidelines with bioinformatics and health informatics scenarios in mind, we believe they can more in general be applied to any scientific area. We therefore target these guidelines to any researcher or practitioners wanting to perform data cleaning or feature engineering. We believe our simple recommendations can help researchers and scholars perform better computational analyses that can lead, in turn, to more solid outcomes and more reliable discoveries.
Davide Chicco, Luca Oneto, Erica Tavazzi
PLoS Comput. Biol.2
2022 Optimizing Fuel Consumption in Thrust Allocation for Marine Dynamic Positioning Systems
abstract
In offshore maritime operations, automated systems capable of maintaining the vessel’s position and heading using its own propellers and thrusters to compensate exogenous disturbances, like wind, waves, and currents, are referred to as marine dynamic positioning (DP) systems. DP systems play a central role in several marine operations, such as drilling, pipe-laying, coring, and ocean observation. These operations are the primary cause of fuel consumption, having a strong impact on the overall footprint of the vessel. For this reason, we will face the problem of optimal thrust allocation of an over-actuated vessel to maintain position and heading with minimal fuel consumption. State-of-the-art approaches simplify this problem by roughly approximating it and obtain a simple, mostly convex, optimization problem that can be solved in near-real time by the automation system. In this article, we improve current approaches with the following contributions. We will exploit a higher fidelity representation of the physical system, and we will manipulate the resulting optimization problem accordingly, to allow for near-real-time solutions on conventional computing platforms on-board. We evaluate the quality of the proposal with a case study on a drilling unit equipped with six thrusters. The results will show that it is possible to achieve up to 5% of fuel savings with respect to conventional approaches.Note to Practitioners—This article was motivated by the problem of minimizing fuel consumption in thrust allocation of DP systems. The current approaches simplify this issue by adopting simpler, yet related, optimization problems as surrogates, keeping the problem tractable for near-real-time control. We propose, instead, to solve the original problem with state-of-the-art modelization of the physical system and exploit reasonable and theoretical proprietaries to achieve optimal solutions in near-real-time. The results on a drilling unit will show additional fuel savings of up to 5% with respect to alternative state-of-the-art approaches.
Miltiadis Kalikatzarakis, Andrea Coraddu, Luca Oneto, Davide Anguita
IEEE Trans Autom. Sci. Eng.3
2021 In-Station Train Movements Prediction: from Shallow to Deep Multi Scale Models
abstract
Public railway transport systems play a crucial role in servicing the global society and are the transport backbone of a sustainable economy.While a significant effort has been devoted to predict inter-station trains movements to support stakeholders (i.e., infrastructure managers, train operators, and travellers) decisions, the problem of predicting instation movements, while being crucial to improve train dispatching (i.e., empowering human or automatic dispatchers), has been far more less investigated.In fact, stations are the most critical points in a railway network: even small improvements in the estimation of the duration of trains movements can remarkably enhance the dispatching efficiency in coping with the increase in capacity demand and with delays.In this work we will first leverage on state of the art shallow models, fed by domain experts with domain specific features, to improve the current predictive systems.Then, we will leverage on a customised deep multi scale model able to automatically learn the representation and improve the accuracy of the shallow models.Results on real-world data coming from the Italian railway network will support our proposal.* This work has been partially
Gianluca Boleto, Luca Oneto, Matteo Cardellini, Marco Maratea, Mauro Vallati, Renzo Canepa, Davide Anguita
ESANN2
2021 Complex Data: Learning Trustworthily, Automatically, and with Guarantees
abstract
Machine Learning (ML) achievements enabled automatic extraction of actionable information from data in a wide range of decisionmaking scenarios.This demands for improving both ML technical aspects (e.g., design and automation) and human-related metrics (e.g., fairness, robustness, privacy, and explainability), with performance guarantees at both levels.The aforementioned scenario posed three main challenges: (i) Learning from Complex Data (i.e., sequence, tree, and graph data), (ii) Learning Trustworthily, and (iii) Learning Automatically with Guarantees.The focus of this special session is on addressing one or more of these challenges with the final goal of Learning Trustworthily, Automatically, and with Guarantees from Complex Data.
Luca Oneto, Nicolò Navarin, Battista Biggio, Federico Errica, Alessio Micheli, Franco Scarselli, Monica Bianchini, Alessandro Sperduti
ESANN1
2021 The Benefits of Adversarial Defence in Generalisation
abstract
Recent researches have been shown that models induced by machine learning, in particular by deep learning, can be easily fooled by an adversary who carefully crafts imperceptible, at least from the human perspective, or physically plausible modifications of the input data.This discovery gave birth to a new field of research, the adversarial machine learning, where new methods of attacks and defence are developed continuously, mimicking what is happening from a long time in cybersecurity.In this paper we will show that the drawbacks of inducing models from data less prone to be misled actually provides some benefits when it comes to assess their generalisation abilities.
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2021 Learn and Visually Explain Deep Fair Models: an Application to Face Recognition
abstract
Trustworthiness, and in particular Algorithmic Fairness, is emerging as one of the most trending topics in Machine Learning (ML). In fact, ML is now ubiquitous in decision making scenarios, highlighting the necessity of discovering and correcting unfair treatments of (historically discriminated) subgroups in the population (e.g., based on gender, ethnicity, political and sexual orientation). This necessity is even more compelling and challenging when unexplainable black-box Deep Neural Networks (DNN) are exploited. An emblematic example of this necessity is provided by the detected unfair behavior of the ML-based face recognition systems exploited by law enforcement agencies in the United States. To tackle these issues, we first propose different (un)fairness mitigation regularizers in the training process of DNNs. We then study where these regularizers should be applied to make them as effective as possible. We finally measure, by means of different accuracy and fairness metrics and different visual explanation strategies, the ability of the resulting DNNs in learning the desired task while, simultaneously, behaving fairly. Results on the recent FairFace dataset prove the validity of our approach.
Danilo Franco, Luca Oneto, Nicolò Navarin, Davide Anguita
IJCNN2
2021 A Planning-based Approach for In-Station Train Dispatching
abstract
In-station train dispatching is the problem of optimising the effective utilisation of available railway infrastructures for mitigating incidents and delays. In this paper, we describe an approach for dealing with the in-station dispatching problem by means of automated planning techniques.
Matteo Cardellini, Marco Maratea, Mauro Vallati, Gianluca Boleto, Luca Oneto
SOCS5
2021 Marine dual fuel engines monitoring in the wild through weakly supervised data analytics
Andrea Coraddu, Luca Oneto, Davide Ilardi, Sokratis Stoumpos, Gerasimos Theotokatos
Eng. Appl. Artif. Intell.2
2021 An Enhanced Random Forests Approach to Predict Heart Failure From Small Imbalanced Gene Expression Data
abstract
Myocardial infarctions and heart failure are the cause of more than 17 million deaths annually worldwide. ST-segment elevation myocardial infarctions (STEMI) require timely treatment, because delays of minutes have serious clinical impacts. Machine learning can provide alternative ways to predict heart failure and identify genes involved in heart failure. For these scopes, we applied a Random Forests classifier enhanced with feature elimination to microarray gene expression of 111 patients diagnosed with STEMI, and measured the classification performance through standard metrics such as the Matthews correlation coefficient (MCC) and area under the receiver operating characteristic curve (ROC AUC). Afterwards, we used the same approach to rank all genes by importance, and to detect the genes more strongly associated with heart failure. We validated this ranking by literature review and gene set enrichment analysis. Our classifier employed to predict heart failure achieved MCC = +0.87 and ROC AUC = 0.918, and our analysis identified KLHL22, WDR11, OR4Q3, GPATCH3, and FAH as top five protein-coding genes related to heart failure. Our results confirm the effectiveness of machine learning feature elimination in predicting heart failure from gene expression, and the top genes found by our approach will be able to help biologists and cardiologists further our understanding of heart failure.
Davide Chicco, Luca Oneto
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 Learning Fair and Transferable Representations with Theoretical Guarantees
abstract
Developing learning methods which do not discriminate subgroups in the population is the central goal of algorithmic fairness. One way to reach this goal is by modifying the data representation in order to satisfy prescribed fairness constraints. This allows to reuse the same representation in other context (tasks) without discriminate subgroups. In this work we measure fairness according to demographic parity, requiring the probability of the possible model decisions to be independent of the sensitive information. We argue that the goal of imposing demographic parity can be substantially facilitated within a multi-task learning setting. We leverage task similarities by encouraging a shared fair representation across the tasks via low rank matrix factorization. We derive learning bounds establishing that the learned representation transfers well to novel tasks both in terms of prediction performance and fairness metrics. We present experiments on three real world datasets, showing that the proposed method outperforms state-of-the-art approaches by a significant margin.
Luca Oneto, Michele Donini, Massimiliano Pontil, Andreas Maurer
DSAA1
2020 Learning Deep Fair Graph Neural Networks
Luca Oneto, Nicolò Navarin, Michele Donini
ESANN1
2020 Improving the Union Bound: a Distribution Dependent Approach
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2020 Towards Online Discovery of Data-Aware Declarative Process Models from Event Streams
abstract
In recent years, several techniques have been made available to automatically discover declarative process models from event logs. These techniques are useful to provide a comprehensible picture of the process as opposed to full specifications of process behavior provided by procedural modeling languages. Since many modern systems produce "big data" from business process executions, in previous work, a framework for the discovery of LTL-based declarative process models from streaming event data has been proposed. This framework can be used to process events online, as they occur, as a way to deal with large and complex collections of datasets that are impossible to store and process altogether. However, the proposed framework does not take into account data attributes associated with events in the log, which can otherwise provide valuable insights into the rules that govern the process. This paper makes the first proposal to close this gap by presenting a technique for discovering declarative process models from event streams that incorporates both control-flow dependencies and data conditions. Specifically, we use Hoeffding trees to incrementally discover data-aware declarative process models, which are represented as conjunctions of first-order temporal logic expressions. The proposed technique has been validated on a synthetic event log, and on a real-life log of a cancer treatment process.
Nicolò Navarin, Matteo Cambiaso, Andrea Burattin, Fabrizio Maria Maggi, Luca Oneto, Alessandro Sperduti
IJCNN5
2020 Deep Learning for Cavitating Marine Propeller Noise Prediction at Design Stage
abstract
Reducing the noise impact of ships on the marine environment is one of the objectives of new propellers designs, since they represent the dominant source of underwater radiated noise, especially when cavitation occurs. Consequently, ship designers require new predictive tools able to verify the compliance with noise requirements and to compare the effectiveness of different design solutions. In this context, tools able to provide a reliable estimate of propeller noise spectra based just on the information available at design stage represent a fundamental tool to speed up the design process avoiding model scale tests. This work focus on developing such a tool, adopting methods coming from the world of Machine Learning and Deep Neural Networks, in order to create a model able to predict the cavitating marine propeller noise spectra. For this purpose authors will make use of a dataset collected by means of dedicated model scale measurements in a cavitation tunnel combined with the detailed flow characterization obtainable by calculations carried out with a Boundary Element Method. The performance of the proposed approaches are analyzed considering different definitions of the input and output variables used during the modelization.
Luca Oneto, Francesca Cipollini, Leonardo Miglianti, Giorgio Tani, Stefano Gaggero, Michele Viviani, Andrea Coraddu
IJCNN1
2020 General Fair Empirical Risk Minimization
abstract
We tackle the problem of algorithmic fairness, where the goal is to avoid the unfairly influence of sensitive information, in the general context of regression with possible continuous sensitive attributes. We extend the framework of fair empirical risk minimization of [1] to this general scenario, covering in this way the whole standard supervised learning setting. Our generalized fairness measure reduces to well known notions of fairness available in literature. We derive learning guarantees for our method, that imply in particular its statistical consistency, both in terms of the risk and the fairness measure. We then specialize our approach to kernel methods and propose a convex fair estimator in that setting. We test the estimator on a commonly used benchmark dataset (Communities and Crime) and on a new dataset collected at the University of Genoa1, containing the information of the academic career of five thousand students. The latter dataset provides a challenging real case scenario of unfair behaviour of standard regression methods that benefits from our methodology. The experimental results show that our estimator is effective at mitigating the trade-off between accuracy and fairness requirements.
Luca Oneto, Michele Donini, Massimiliano Pontil
IJCNN1
2020 Fair regression with Wasserstein barycenters
abstract
We study the problem of learning a real-valued function that satisfies the Demographic Parity constraint. It demands the distribution of the predicted output to be independent of the sensitive attribute. We consider the case that the sensitive attribute is available for prediction. We establish a connection between fair regression and optimal transport theory, based on which we derive a close form expression for the optimal fair predictor. Specifically, we show that the distribution of this optimum is the Wasserstein barycenter of the distributions induced by the standard regression function on the sensitive groups. This result offers an intuitive interpretation of the optimal fair prediction and suggests a simple post-processing algorithm to achieve fairness. We establish risk and distribution-free fairness guarantees for this procedure. Numerical experiments indicate that our method is very effective in learning fair models, with a relative increase in error rate that is inferior to the relative gain in fairness.
Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil
NeurIPS4
2020 Fair regression via plug-in estimator and recalibration with statistical guarantees
abstract
We study the problem of learning an optimal regression function subject to a fairness constraint. It requires that, conditionally on the sensitive feature, the distribution of the function output remains the same. This constraint naturally extends the notion of demographic parity, often used in classification, to the regression setting. We tackle this problem by leveraging on a proxy-discretized version, for which we derive an explicit expression of the optimal fair predictor. This result naturally suggests a two stage approach, in which we first estimate the (unconstrained) regression function from a set of labeled data and then we recalibrate it with another set of unlabeled data. The recalibration step can be efficiently performed via a smooth optimization. We derive rates of convergence of the proposed estimator to the optimal fair predictor both in terms of the risk and fairness constraint. Finally, we present numerical experiments illustrating that the proposed method is often superior or competitive with state-of-the-art methods.
Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil
NeurIPS4
2020 Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation Learning
abstract
Developing learning methods which do not discriminate subgroups in the population is a central goal of algorithmic fairness. One way to reach this goal is by modifying the data representation in order to meet certain fairness constraints. In this work we measure fairness according to demographic parity. This requires the probability of the possible model decisions to be independent of the sensitive information. We argue that the goal of imposing demographic parity can be substantially facilitated within a multitask learning setting. We present a method for learning a shared fair representation across multiple tasks, by means of different new constraints based on MMD and Sinkhorn Divergences. We derive learning bounds establishing that the learned representation transfers well to novel tasks. We present experiments on three real world datasets, showing that the proposed method outperforms state-of-the-art approaches by a significant margin.
Luca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto, Andreas Maurer, Massimiliano Pontil
NeurIPS1
2020 Advances in artificial neural networks, machine learning and computational intelligence
Luca Oneto, Kerstin Bunte, Alessandro Sperduti
Neurocomputing1
2020 Randomized learning and generalization of fair and private classifiers: From PAC-Bayes to stability and differential privacy
Luca Oneto, Michele Donini, Massimiliano Pontil, John Shawe-Taylor
Neurocomputing1
2020 Low-Resource Footprint, Data-Driven Malware Detection on Android
abstract
Resource-constrained systems are becoming more and more common as users migrate from PCs to mobile devices and as IoT systems enter the mainstream. At the same time, it is not acceptable to reduce the level of security hence it is necessary to accommodate the required security into the system-imposed resource constraints. This paper introduces BAdDroIds, a mobile application leveraging machine learning for detecting malware on resource constrained devices. BAdDroIds executes in background and transparently analyzes the applications as soon as they are installed, i.e., before infecting the device. BAdDroIds relies on static analysis techniques and features provided by the Android OS to build up sound and complete models of Android apps in terms of permissions and API invocations. It uses ad-hoc supervised classification techniques to allow resource-efficient malware detection. By exploiting the intrinsic nature of data, it has been possible to implement a state-of-the-art data-driven model which provides deep insights on the detection problem and can be efficiently executed on the device itself as it requires a very limited computational effort. Besides its limited resource footprint, BAdDroIds is extremely effective: An extensive experimental evaluation shows that it outperforms the currently available solutions in terms of accuracy, which is around 99 percent.
Simone Aonzo, Alessio Merlo, Mauro Migliardi, Luca Oneto, Francesco Palmieri 0002
IEEE Trans. Sustain. Comput.4
2019 Taking Advantage of Multitask Learning for Fair Classification
abstract
A central goal of algorithmic fairness is to reduce bias in automated decision making. An unavoidable tension exists between accuracy gains obtained by using sensitive information as part of a statistical model, and any commitment to protect these characteristics. Often, due to biases present in the data, using the sensitive information in the functional form of a classifier improves classification accuracy. In this paper we show how it is possible to get the best of both worlds: optimize model accuracy and fairness without explicitly using the sensitive feature in the functional form of the model, thereby treating different individuals equally. Our method is based on two key ideas. On the one hand, we propose to use Multitask Learning (MTL), enhanced with fairness constraints, to jointly learn group specific classifiers that leverage information between sensitive groups. On the other hand, since learning group specific models might not be permitted, we propose to first predict the sensitive features by any learning method and then to use the predicted sensitive feature to train MTL with fairness constraints. This enables us to tackle fairness with a three-pronged approach, that is, by increasing accuracy on each group, enforcing measures of fairness during training, and protecting sensitive information during testing. Experimental results on two real datasets support our proposal, showing substantial improvements in both accuracy and fairness.
Luca Oneto, Michele Donini, Amon Elders, Massimiliano Pontil
AIES1
2019 Societal Issues in Machine Learning: When Learning from Data is Not Enough
Davide Bacciu, Battista Biggio, Paulo J. G. Lisboa, José D. Martín, Luca Oneto, Alfredo Vellido
ESANN5
2019 Fairness and Accountability of Machine Learning Models in Railway Market: are Applicable Railway Laws Up to Regulate Them?
Charlotte Ducuing, Luca Oneto, Renzo Canepa
ESANN2
2019 PAC-Bayes and Fairness: Risk and Fairness Bounds on Distribution Dependent Fair Priors
Luca Oneto, Michele Donini, Massimiliano Pontil
ESANN1
2019 Hybrid Model for Cavitation Noise Spectra Prediction
abstract
In the latest years, models combining physical knowledge of a phenomenon and statistical inference are becoming of much interest in many real world applications. In this context, ship propeller underwater radiated noise is an interesting field of application for these so-called hybrid models, especially when the propeller cavitates. Nowadays, model scale tests are considered the state-of-the-art technique to predict the cavitation noise spectra. Unfortunately, they are negatively affected by scale effects which could alter the onset of some interesting cavitating phenomena respect to the full scale propeller; as a consequence, for some ship operational conditions it is not trivial to correctly reproduce the cavitation pattern in model scale tests. Moreover, model scale tests are quite expensive and time-consuming; it is not feasible to include them in the early stage of the design. Nevertheless, data collected during these tests can be adopted in order to tune a data-driven model while the physical equation describing the occurring phenomenon can be used to refine the prediction. In this work, the authors propose a hybrid model for the prediction of ships propeller underwater radiated noise, able to exploit both the physical knowledge of the problem and the real data obtained from cavitation tunnel experiments performed on different propellers in different working conditions. Results on real data will support the validity and the effectiveness of the proposal.
Francesca Cipollini, Fabiana Miglianti, Luca Oneto, Giorgio Tani, Michele Viviani
IJCNN3
2019 Ensemble Application of Transfer Learning and Sample Weighting for Stock Market Prediction
abstract
Forecasting stock market behavior is an interesting and challenging problem. Regression of prices and classification of daily returns have been widely studied with the main goal of supplying forecasts useful in real trading scenarios. Unfortunately, the outcomes are not directly related with the maximization of the financial gain. Firstly, the optimal strategy requires to invest on the most performing asset every period and trading accordingly is not trivial given the predictions. Secondly, price fluctuations of different magnitude are often treated as equals even if during market trading losses or gains of different intensities are derived. In this paper, the problem of stock market forecasting is formulated as regression of market returns. This approach is able to estimate the amount of price change and thus the most performing assets. Price fluctuations of different magnitude are treated differently through the application of different weights on samples and the scarcity of data is addressed using transfer learning. Results on a real simulation of trading show how, given a finite amount of capital, the predictions can be used to invest in high performing stocks and, hence, achieve higher profits with less trades.
Simone Merello, Andrea Picasso Ratto, Luca Oneto, Erik Cambria
IJCNN3
2019 Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification
abstract
We study the problem of fair binary classification using the notion of Equal Opportunity. It requires the true positive rate to distribute equally across the sensitive groups. Within this setting we show that the fair optimal classifier is obtained by recalibrating the Bayes classifier by a group-dependent threshold. We provide a constructive expression for the threshold. This result motivates us to devise a plug-in classification procedure based on both unlabeled and labeled datasets. While the latter is used to learn the output conditional probability, the former is used for calibration. The overall procedure can be computed in polynomial time and it is shown to be statistically consistent both in terms of the classification error and fairness measure. Finally, we present numerical experiments which indicate that our method is often superior or competitive with the state-of-the-art methods on benchmark datasets.
Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil
NeurIPS4
2019 Technical analysis and sentiment embeddings for market trend prediction
Andrea Picasso Ratto, Simone Merello, Luca Oneto, Erik Cambria
Expert Syst. Appl.4
2019 Simple continuous optimal regions of the space of data
Alessio Carrega, Francesca Cipollini, Luca Oneto
Neurocomputing3
2019 Advances in artificial neural networks, machine learning and computational intelligence: Selected papers from the 26th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN 2018)
Luca Oneto, Kerstin Bunte, Frank-Michael Schleif
Neurocomputing1
2019 Local Rademacher Complexity Machine
Luca Oneto, Sandro Ridella, Davide Anguita
Neurocomputing1
2018 Large-Scale Railway Networks Train Movements: A Dynamic, Interpretable, and Robust Hybrid Data Analytics System
abstract
We investigate the problem of analyzing the train movements in Large-Scale Railway Networks for the purpose of understanding and predicting their behaviour. We focus on different important aspects: the Running Time of a train between two stations, the Dwell Time of a train in a station, the Train Delay, and the Penalty Costs associated to a delay. Two main approaches exist in literature to study these aspects. One is based on the knowledge of the network and the experience of the operators. The other one is based on the analysis of the historical data about the network with advanced data analytics methods. In this paper, we will propose an hybrid approach in order to address the limitations of the current solutions. In fact, experience-based models are interpretable and robust but not really able to take into account all the factors which influence train movements resulting in low accuracy. From the other side, Data-Driven models are usually not easy to interpret, nor robust to infrequent events, and require a representative amount of data which is not always available if the phenomenon under examination changes too fast. Results on real world data coming from the Italian railway network will show that the proposed solution outperforms both state-of-the-art experience and Data-Driven based systems in terms of interpretability, robustness, ability to handle non recurrent events and changes in the behaviour of the network, and ability to consider complex and exogenous information.
Alessandro Lulli, Luca Oneto, Renzo Canepa, Simone Petralli, Davide Anguita
DSAA2
2018 Emerging trends in machine learning: beyond conventional methods and data
Luca Oneto, Nicolò Navarin, Michele Donini, Davide Anguita
ESANN1
2018 Local Rademacher Complexity Machine
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2018 Empirical Risk Minimization Under Fairness Constraints
abstract
We address the problem of algorithmic fairness: ensuring that sensitive information does not unfairly influence the outcome of a classifier. We present an approach based on empirical risk minimization, which incorporates a fairness constraint into the learning problem. It encourages the conditional risk of the learned classifier to be approximately constant with respect to the sensitive variable. We derive both risk and fairness bounds that support the statistical consistency of our methodology. We specify our approach to kernel methods and observe that the fairness requirement implies an orthogonality constraint which can be easily added to these methods. We further observe that for linear models the constraint translates into a simple data preprocessing step. Experiments indicate that the method is empirically effective and performs favorably against state-of-the-art approaches.
Michele Donini, Luca Oneto, Shai Ben-David, John Shawe-Taylor, Massimiliano Pontil
NeurIPS2
2018 Advances in artificial neural networks, machine learning and computational intelligence
Fabio Aiolli, Michael Biehl, Luca Oneto
Neurocomputing3
2018 Randomized learning: Generalization performance of old and new theoretically grounded algorithms
Luca Oneto, Francesca Cipollini, Sandro Ridella, Davide Anguita
Neurocomputing1
2018 Multilayer Graph Node Kernels: Stacking While Maintaining Convexity
Luca Oneto, Nicolò Navarin, Alessandro Sperduti, Davide Anguita
Neural Process. Lett.1
2018 Learning With Kernels: A Local Rademacher Complexity-Based Analysis With Application to Graph Kernels
abstract
When dealing with kernel methods, one has to decide which kernel and which values for the hyperparameters to use. Resampling techniques can address this issue but these procedures are time-consuming. This problem is particularly challenging when dealing with structured data, in particular with graphs, since several kernels for graph data have been proposed in literature, but no clear relationship among them in terms of learning properties is defined. In these cases, exhaustive search seems to be the only reasonable approach. Recently, the global Rademacher complexity (RC) and local Rademacher complexity (LRC), two powerful measures of the complexity of a hypothesis space, have shown to be suited for studying kernels properties. In particular, the LRC is able to bound the generalization error of an hypothesis chosen in a space by disregarding those ones which will not be taken into account by any learning procedure because of their high error. In this paper, we show a new approach to efficiently bound the RC of the space induced by a kernel, since its exact computation is an NP-Hard problem. Then we show for the first time that RC can be used to estimate the accuracy and expressivity of different graph kernels under different parameter configurations. The authors' claims are supported by experimental results on several real-world graph data sets.
Luca Oneto, Nicolò Navarin, Michele Donini, Sandro Ridella, Alessandro Sperduti, Fabio Aiolli, Davide Anguita
IEEE Trans. Neural Networks Learn. Syst.1
2017 Crack random forest for arbitrary large datasets
abstract
Random Forests (RF) of tree classifiers are a state-of-the-art method for classification purposes. RF show limited hyperparameter sensitivity, have high numerical robustness, possess native capacity of dealing with numerical and categorical features, and are quite effective in many real world problems with respect to other state-of-the-art techniques. In this work we show how to crack RF in order to be able to train them on arbitrary large datasets. In particular, we extend ReForeSt, an Apache Spark-based RF implementation. The new version of ReForeSt computation automatically adapts to two methodologies to distribute the data and the computation on the available machines and automatically chooses the one able to provide the result in less time. The new ReForeSt also supports Random Rotations, a quite recent randomization technique which can bust the accuracy of the original RF. We perform an extensive experimental evaluation between ReForeSt and MLlib by taking advantage of the Google Cloud Platform1. We test the performances and the scalability of ReForeSt and MLlib on several real world datasets. Results confirm that ReForeSt outperforms MLlib both in terms of memory and computational efficiency, and classification performances. ReForeSt is publicly available via GitHub2.
Alessandro Lulli, Luca Oneto, Davide Anguita
IEEE BigData2
2017 Generalization Performances of Randomized Classifiers and Algorithms built on Data Dependent Distributions
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2017 Dropout Prediction at University of Genoa: a Privacy Preserving Data Driven Approach
Luca Oneto, Anna Siri, Gianvittorio Luria, Davide Anguita
ESANN1
2017 ReForeSt: Random Forests in Apache Spark
Alessandro Lulli, Luca Oneto, Davide Anguita
ICANN (2)2
2017 Marine Safety and Data Analytics: Vessel Crash Stop Maneuvering Performance Prediction
Luca Oneto, Andrea Coraddu, Paolo Sanetti, Olena Karpenko, Francesca Cipollini, Toine Cleophas, Davide Anguita
ICANN (2)1
2017 Deep graph node kernels: A convex approach
abstract
Nowadays, developing effective techniques able to deal with data coming from structured domains is becoming crucial. In this context kernel methods are the state-of-the-art tool widely adopted in real-world applications that involve learning on structured data. Contrarily, when one has to deal with unstructured domains, deep learning methods represent a competitive, or even better, choice. In this paper we propose a new family of kernels for graphs which exploits a deep representation of the information. Our proposal exploits the advantages of the two worlds. From one side we exploit the potentiality of the state-of-the-art graph kernels. From the other side we develop a deep architecture through a series of stacked kernel pre-image estimators trained in an unsupervised fashion via convex optimization. The hidden layers of the proposed framework are trained in a forward manner and this allows us to avoid the greedy layerwise training of classical deep learning. Results on real world graph datasets confirm the quality of the proposal.
Luca Oneto, Nicolò Navarin, Alessandro Sperduti, Davide Anguita
IJCNN1
2017 Measuring the expressivity of graph kernels through Statistical Learning Theory
Luca Oneto, Nicolò Navarin, Michele Donini, Alessandro Sperduti, Fabio Aiolli, Davide Anguita
Neurocomputing1
2017 Differential privacy and generalization: Sharper bounds with applications
Luca Oneto, Sandro Ridella, Davide Anguita
Pattern Recognit. Lett.1
2017 Dynamic Delay Predictions for Large-Scale Railway Networks: Deep and Shallow Extreme Learning Machines Tuned via Thresholdout
abstract
Current train delay (TD) prediction systems do not take advantage of state-of-the-art tools and techniques for handling and extracting useful and actionable information from the large amount of endogenous (i.e., generated by the railway system itself) and exogenous (i.e., related to railway operation but generated by external phenomena) data available. Additionally, they are not designed in order to deal with the intrinsic time varying nature of the problem (e.g., regular changes in the nominal timetable, etc.). The purpose of this paper is to build a dynamic data-driven TD prediction system that exploits the most recent tools and techniques in the field of time varying big data analysis. In particular, we map the TD prediction problem into a time varying multivariate regression problem that allows exploiting both historical data about the train movements and exogenous data about the weather provided by the national weather services. The performance of these methods have been tuned through the state-of-the-art thresholdout technique, a very powerful procedure which relies on the differential privacy theory. Finally, the performance of two efficient implementations of shallow and deep extreme learning machines that fully exploit the recent in-memory large-scale data processing technologies have been compared with the current state-of-the-art TD prediction systems. Results on real-world data coming from the Italian railway network show that the proposal of this paper is able to remarkably improve the state-of-the-art systems.
Luca Oneto, Emanuele Fumeo, Giorgio Clerico, Renzo Canepa, Federico Papa, Carlo Dambra, Nadia Mazzino, Davide Anguita
IEEE Trans. Syst. Man Cybern. Syst.1
2016 Advanced Analytics for Train Delay Prediction Systems by Including Exogenous Weather Data
abstract
State-of-the-art train delay prediction systems neither exploit historical data about train movements, nor exogenous data about phenomena that can affect railway operations. They rely, instead, on static rules built by experts of the railway infrastructure based on classical univariate statistics. The purpose of this paper is to build a data-driven train delay prediction system that exploits the most recent analytics tools. The train delay prediction problem has been mapped into a multivariate regression problem and the performance of kernel methods, ensemble methods and feed-forward neural networks have been compared. Firstly, it is shown that it is possible to build a reliable and robust data-driven model based only on the historical data about the train movements. Additionally, the model can be further improved by including data coming from exogenous sources, in particular the weather information provided by national weather services. Results on real world data coming from the Italian railway network show that the proposal of this paper is able to remarkably improve the current state-of-the-art train delay prediction systems. Moreover, the performed simulations show that the inclusion of weather data into the model has a significant positive impact on its performance.
Luca Oneto, Emanuele Fumeo, Giorgio Clerico, Renzo Canepa, Federico Papa, Carlo Dambra, Nadia Mazzino, Davide Anguita
DSAA1
2016 Advances in Learning with Kernels: Theory and Practice in a World of growing Constraints
Luca Oneto, Nicolò Navarin, Michele Donini, Fabio Aiolli, Davide Anguita
ESANN1
2016 Measuring the Expressivity of Graph Kernels through the Rademacher Complexity
Luca Oneto, Nicolò Navarin, Michele Donini, Alessandro Sperduti, Fabio Aiolli, Davide Anguita
ESANN1
2016 Tuning the Distribution Dependent Prior in the PAC-Bayes Framework based on Empirical Data
Luca Oneto, Sandro Ridella, Davide Anguita
ESANN1
2016 Random Forests Model Selection
Ilenia Orlandi, Luca Oneto, Davide Anguita
ESANN2
2016 Transition-Aware Human Activity Recognition Using Smartphones
Jorge Luis Reyes-Ortiz, Luca Oneto, Albert Samà, Xavier Parra Llanas, Davide Anguita
Neurocomputing2
2016 Can machine learning explain human learning?
Mehrnoosh Vahdat, Luca Oneto, Davide Anguita, Mathias Funk, Matthias Rauterberg
Neurocomputing2
2016 Tikhonov, Ivanov and Morozov regularization for support vector machine learning
Luca Oneto, Sandro Ridella, Davide Anguita
Mach. Learn.1
2016 A local Vapnik-Chervonenkis complexity
Luca Oneto, Davide Anguita, Sandro Ridella
Neural Networks1
2016 Global Rademacher Complexity Bounds: From Slow to Fast Convergence Rates
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
Neural Process. Lett.1
2016 PAC-bayesian analysis of distribution dependent priors: Tighter risk bounds and stability analysis
Luca Oneto, Davide Anguita, Sandro Ridella
Pattern Recognit. Lett.1
2016 Learning Hardware-Friendly Classifiers Through Algorithmic Stability
abstract
Most state-of-the-art machine-learning (ML) algorithms do not consider the computational constraints of implementing the learned model on embedded devices. These constraints are, for example, the limited depth of the arithmetic unit, the memory availability, or the battery capacity. We propose a new learning framework, the Algorithmic Risk Minimization (ARM), which relies on Algorithmic-Stability, and includes these constraints inside the learning process itself. ARM allows one to train advanced resource-sparing ML models and to efficiently deploy them on smart embedded systems. Finally, we show the advantages of our proposal on a smartphone-based Human Activity Recognition application by comparing it to a conventional ML approach.
Luca Oneto, Sandro Ridella, Davide Anguita
ACM Trans. Embed. Comput. Syst.1
2015 Performance assessment and uncertainty quantification of predictive models for smart manufacturing systems
abstract
We review in this paper several methods from Statistical Learning Theory (SLT) for the performance assessment and uncertainty quantification of predictive models. Computational issues are addressed so to allow the scaling to large datasets and the application of SLT to Big Data analytics. The effectiveness of the application of SLT to manufacturing systems is exemplified by targeting the derivation of a predictive model for quality forecasting of products on an assembly line.
Luca Oneto, Ilenia Orlandi, Davide Anguita
IEEE BigData1
2015 A Learning Analytics Approach to Correlate the Academic Achievements of Students with Interaction Data from an Educational Simulator
abstract
This paper presents a Learning Analytics approach for understanding the learning behavior of students while interacting with Technology Enhanced Learning tools. In this work we show that it is possible to gain insight into the learning processes of students from their interaction data. We base our study on data collected through six laboratory sessions where first-year students of Computer Engineering at the University of Genoa were using a digital electronics simulator. We exploit Process Mining methods to investigate and compare the learning processes of students. For this purpose, we measure the understandability of their process models through a complexity metric. Then we compare the various clusters of students based on their academic achievements. The results show that the measured complexity has positive correlation with the final grades of students and negative correlation with the difficulty of the laboratory sessions. Consequently, complexity of process models can be used as an indicator of variations of student learning paths.
Mehrnoosh Vahdat, Luca Oneto, Davide Anguita, Mathias Funk, Matthias Rauterberg
EC-TEL2
2015 Model Selection for Big Data: Algorithmic Stability and Bag of Little Bootstraps on GPUs
Luca Oneto, Bernardo Pilarz, Alessandro Ghio, Davide Anguita
ESANN1
2015 Advances in learning analytics and educational data mining
Mehrnoosh Vahdat, Alessandro Ghio, Luca Oneto, Davide Anguita, Mathias Funk, Matthias Rauterberg
ESANN3
2015 Human Algorithmic Stability and Human Rademacher Complexity
Mehrnoosh Vahdat, Luca Oneto, Alessandro Ghio, Davide Anguita, Mathias Funk, Matthias Rauterberg
ESANN2
2015 Shrinkage learning to improve SVM with hints
abstract
The Support Vector Machine (SVM) is one of the most effective and used algorithms, when targeting classification. Despite its large success, SVM is mainly afflicted by two issues: (i) some hyperparameters must be tuned in advance and are, in practice, identified through computationally intensive procedures; (ii) possible a-priori knowledge about the problem (e.g. doctor expertise in medical applications) cannot be straightforwardly exploited. In this paper, we introduce a new approach, able to cope with the two previous problems: several experiments, performed on real-world benchmarking datasets, show that our method outperforms, on average, other techniques proposed in the literature.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IJCNN1
2015 Support vector machines and strictly positive definite kernel: The regularization hyperparameter is more important than the kernel hyperparameters
abstract
When dealing with a Support Vector Machine (SVM) with a strictly positive definite kernel, a common misconception is that the main handle for controlling the nonlinearity of the classification surface is the set of kernel hyperparameters. We show here that this is not the case: in particular, we prove that, regardless of the value of the kernel hyperparameter, it is always possible to tune the nonlinearity of the classifier by acting only on the regularization hyperparameter C, even achieving perfect learning of any non-degenerate training set.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IJCNN1
2015 Fast convergence of extended Rademacher Complexity bounds
abstract
In this work we propose some new generalization bounds for binary classifiers, based on global Rademacher Complexity (RC), which exhibit fast convergence rates by combining state-of-the-art results by Talagrand on empirical processes and the exploitation of unlabeled patterns. In this framework, we are able to improve both the constants and the convergence rates of existing RC-based bounds. All the proposed bounds are based on empirical quantities, so that they can be easily computed in practice, and are provided both in implicit and explicit forms: the formers are the tightest ones, while the latter ones allow to get more insights about the impact of Talagrand's results and the exploitation of unlabeled patterns in the learning process. Finally, we verify the quality of the bounds, with respect to the theoretical limit, showing the room for further improvements in the common scenario of binary classification.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IJCNN1
2015 Learning Resource-Aware Classifiers for Mobile Devices: From Regularization to Energy Efficiency
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
Neurocomputing1
2015 Local Rademacher Complexity: Sharper risk bounds with and without unlabeled samples
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
Neural Networks1
2015 Fully Empirical and Data-Dependent Stability-Based Bounds
abstract
The purpose of this paper is to obtain a fully empirical stability-based bound on the generalization ability of a learning procedure, thus, circumventing some limitations of the structural risk minimization framework. We show that assuming a desirable property of a learning algorithm is sufficient to make data-dependency explicit for stability, which, instead, is usually bounded only in an algorithmic-dependent way. In addition, we prove that a well-known and widespread classifier, like the support vector machine (SVM), satisfies this condition. The obtained bound is then exploited for model selection purposes in SVM classification and tested on a series of real-world benchmarking datasets demonstrating, in practice, the effectiveness of our approach.
Luca Oneto, Alessandro Ghio, Sandro Ridella, Davide Anguita
IEEE Trans. Cybern.1
2014 A Learning Analytics Methodology to Profile Students Behavior and Explore Interactions with a Digital Electronics Simulator
Mehrnoosh Vahdat, Luca Oneto, Alessandro Ghio, Giuliano Donzellini, Davide Anguita, Mathias Funk, Matthias Rauterberg
EC-TEL2
2014 Learning with few bits on small-scale devices: From regularization to energy efficiency
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN3
2014 Byte The Bullet: Learning on Real-World Computing Architectures
Alessandro Ghio, Luca Oneto
ESANN2
2014 Human Activity Recognition on Smartphones with Awareness of Basic Activities and Postural Transitions
Jorge Luis Reyes-Ortiz, Luca Oneto, Alessandro Ghio, Albert Samà, Davide Anguita, Xavier Parra Llanas
ICANN2
2014 Smartphone battery saving by bit-based hypothesis spaces and local Rademacher Complexities
abstract
Smartphones emerge from the incorporation of new services and features into mobile phones, allowing to implement advanced functionalities for the final users. The implementation of Machine Learning (ML) algorithms on the smartphone itself, without resorting to remote computing systems, allow to achieve such goals without expensive data transmission. However, smartphones are resource-limited devices and, as such, suffer from many issues, which are typical of stand-alone devices, such as limited battery capacity and processing power. We show in this paper how to build a thrifty classifier by exploiting bit-based hypothesis spaces and local Rademacher Complexities. The resulting classifier is tested on a real-world Human Activity Recognition application, implemented on a Samsung Galaxy S II smartphone.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN3
2014 Unlabeled patterns to tighten Rademacher complexity error bounds for kernel classifiers
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
Pattern Recognit. Lett.3
2014 A Deep Connection Between the Vapnik-Chervonenkis Entropy and the Rademacher Complexity
abstract
In this paper, we derive a deep connection between the Vapnik-Chervonenkis (VC) entropy and the Rademacher complexity. For this purpose, we first refine some previously known relationships between the two notions of complexity and then derive new results, which allow computing an admissible range for the Rademacher complexity, given a value of the VC-entropy, and vice versa. The approach adopted in this paper is new and relies on the careful analysis of the combinatorial nature of the problem. The obtained results improve the state of the art on this research topic.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IEEE Trans. Neural Networks Learn. Syst.3
2013 A Public Domain Dataset for Human Activity Recognition using Smartphones
Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra Llanas, Jorge Luis Reyes-Ortiz
ESANN3
2013 A Learning Machine with a Bit-Based Hypothesis Space
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN3
2013 Training Computationally Efficient Smartphone-Based Human Activity Recognition Models
Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra Llanas, Jorge Luis Reyes-Ortiz
ICANN3
2013 A Novel Procedure for Training L1-L2 Support Vector Machine Classifiers
Davide Anguita, Alessandro Ghio, Luca Oneto, Jorge Luis Reyes-Ortiz, Sandro Ridella
ICANN3
2013 Some results about the Vapnik-Chervonenkis entropy and the rademacher complexity
abstract
This paper deals with the problem of identifying a connection between the Vapnik-Chervonenkis (VC) Entropy, a notion of complexity introduced by Vapnik in his seminal work, and the Rademacher Complexity, a more powerful notion of complexity, which has been in the limelight of several works in the recent Machine Learning literature. In order to establish this connection, we refine some previously known relationships and derive a new result. Our proposal allows computing an admissible range for the Rademacher Complexity, given a value of the VC-Entropy, and vice versa, therefore opening new appealing research perspectives in the field of assessing the complexity of an hypothesis space.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN3
2013 A support vector machine classifier from a bit-constrained, sparse and localized hypothesis space
abstract
Choosing an appropriate hypothesis space in classification applications, according to the Structural Risk Minimization (SRM) principle, is of paramount importance to train effective models: in fact, properly selecting the the space complexity allows to optimize the learned functions performance. This selection is not straightforward, especially (though not solely) when few samples are available for deriving an effective model (e.g. in bioinformatics applications). In this paper, by exploiting a bit-based definition for Support Vector Machine (SVM) classifiers, selected from an hypothesis space described according to sparsity and locality principles, we show how the complexity of the corresponding space of functions can be effectively tuned through the number of bits used for the function representation. Real world datasets are exploited to show how the number of bits and the degree of sparsity/locality imposed to define the hypothesis space affect the complexity of the space of classifiers and, consequently, the performance of the model, picked up from this set.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN3
2013 An improved analysis of the Rademacher data-dependent bound using its self bounding property
Luca Oneto, Alessandro Ghio, Davide Anguita, Sandro Ridella
Neural Networks1
2012 The 'K' in K-fold Cross Validation
Davide Anguita, Luca Ghelardoni, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN4
2012 Structural Risk Minimization and Rademacher Complexity for Regression
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN3
2012 Nested Sequential Minimal Optimization for Support Vector Machines
Alessandro Ghio, Davide Anguita, Luca Oneto, Sandro Ridella, Carlotta Schatten
ICANN (2)3
2012 Rademacher Complexity and Structural Risk Minimization: An Application to Human Gene Expression Datasets
Luca Oneto, Davide Anguita, Alessandro Ghio, Sandro Ridella
ICANN (2)1
2012 In-sample Model Selection for Trimmed Hinge Loss Support Vector Machine
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
Neural Process. Lett.3
2012 In-Sample and Out-of-Sample Model Selection and Error Estimation for Support Vector Machines
abstract
In-sample approaches to model selection and error estimation of support vector machines (SVMs) are not as widespread as out-of-sample methods, where part of the data is removed from the training set for validation and testing purposes, mainly because their practical application is not straightforward and the latter provide, in many cases, satisfactory results. In this paper, we survey some recent and not-so-recent results of the data-dependent structural risk minimization framework and propose a proper reformulation of the SVM learning algorithm, so that the in-sample approach can be effectively applied. The experiments, performed both on simulated and real-world datasets, show that our in-sample approach can be favorably compared to out-of-sample methods, especially in cases where the latter ones provide questionable results. In particular, when the number of samples is small compared to their dimensionality, like in classification of microarray data, our proposal can outperform conventional out-of-sample approaches such as the cross validation, the leave-one-out, or the Bootstrap methods.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IEEE Trans. Neural Networks Learn. Syst.3
2011 Maximal Discrepancy vs. Rademacher Complexity for error estimation
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
ESANN3
2011 In-sample model selection for Support Vector Machines
abstract
In-sample model selection for Support Vector Machines is a promising approach that allows using the training set both for learning the classifier and tuning its hyperparameters. This is a welcome improvement respect to out-of-sample methods, like cross-validation, which require to remove some samples from the training set and use them only for model selection purposes. Unfortunately, in-sample methods require a precise control of the classifier function space, which can be achieved only through an unconventional SVM formulation, based on Ivanov regularization. We prove in this work that, even in this case, it is possible to exploit well-known Quadratic Programming solvers like, for example, Sequential Minimal Optimization, so improving the applicability of the in-sample approach.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN3
2011 Selecting the hypothesis space for improving the generalization ability of Support Vector Machines
abstract
The Structural Risk Minimization framework has been recently proposed as a practical method for model selection in Support Vector Machines (SVMs). The main idea is to effectively measure the complexity of the hypothesis space, as defined by the set of possible classifiers, and to use this quantity as a penalty term for guiding the model selection process. Unfortunately, the conventional SVM formulation defines a hypothesis space centered at the origin, which can cause undesired effects on the selection of the optimal classifier. We propose here a more flexible SVM formulation, which addresses this drawback, and describe a practical method for selecting more effective hypothesis spaces, leading to the improvement of the generalization ability of the final classifier.
Davide Anguita, Alessandro Ghio, Luca Oneto, Sandro Ridella
IJCNN3
2011 The Impact of Unlabeled Patterns in Rademacher Complexity Theory for Kernel Classifiers
abstract
We derive here new generalization bounds, based on Rademacher Complexity theory, for model selection and error estimation of linear (kernel) classifiers, which exploit the availability of unlabeled samples. In particular, two results are obtained: the first one shows that, using the unlabeled samples, the confidence term of the conventional bound can be reduced by a factor of three; the second one shows that the unlabeled samples can be used to obtain much tighter bounds, by building localized versions of the hypothesis class containing the optimal classifier.
Luca Oneto, Davide Anguita, Alessandro Ghio, Sandro Ridella
NIPS1
2010 Model selection for support vector machines: Advantages and disadvantages of the Machine Learning Theory
abstract
A common belief is that Machine Learning Theory (MLT) is not very useful, in pratice, for performing effective SVM model selection. This fact is supported by experience, because well-known hold-out methods like cross-validation, leave-one-out, and the bootstrap usually achieve better results than the ones derived from MLT. We show in this paper that, in a small sample setting, i.e. when the dimensionality of the data is larger than the number of samples, a careful application of the MLT can outperform other methods in selecting the optimal hyperparameters of a SVM.
Davide Anguita, Alessandro Ghio, Noemi Greco, Luca Oneto, Sandro Ridella
IJCNN4