Alessio Micheli

dblp:34/4759 · DBLP profile ↗
← Back
124ranked-venue papers
14as first author
47since 2021 · last 2026
0000-0001-5764-5238ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 115 · 14 first-author · 42 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 A method for the systematic generation of graph XAI benchmarks via Weisfeiler-Leman coloring
Michele Fontanesi, Alessio Micheli, Marco Podda, Domenico Tortorella
Data Min. Knowl. Discov.2
2026 Randomized Ising models for graph node representation
Maria Grazia Berni, Antonio Brau, Alessio Micheli, Domenico Tortorella
Neurocomputing3
2026 Informed machine learning for complex data
abstract
Machine Learning (ML) has become a central force in Artificial Intelligence, driving major breakthroughs in applications that handle increasingly complex data, from images and text sequences to graph structures. While new architectures such as Transformers and Graph Neural Networks continue to redefine performance benchmarks in various domains, these predominantly data-driven methods often neglect critical domain knowledge, practical constraints, and broader contextual factors. This oversight diminishes their trustworthiness and restricts their impact in real-world settings. In this paper, we discuss the need for a more informed approach to ML for complex data. Specifically, we advocate for solutions that explicitly integrate structural awareness to capture underlying relationships in the data, incorporate key technical requirements to ensure safety and compliance with industry standards, embed environmental considerations to promote sustainability and resource efficiency, adhere to established physical principles, and uphold ethical and societal values. By weaving these dimensions together, informed ML can bridge the gap between purely data-centric methods and the nuanced demands of practical applications. We show how this integrated framework not only strengthens model performance but also ensures that ML solutions remain trustworthy, efficient, and sensitive to human ecological, ethical, and regulatory imperatives. Our discussion underscores the transformative potential of Informed ML to drive innovation across diverse domains, setting a new benchmark for responsible and high-impact ML system design.
Luca Oneto, Nicolò Navarin, Alessio Micheli, Luca Pasa, Claudio Gallicchio, Davide Bacciu, Davide Anguita
Neurocomputing3
2025 Robustness in Protein-Protein Interaction Networks: A Link Prediction Approach
abstract
Protein-protein interaction networks (PPINs) are indispensable in exploring complex biological systems, facilitating advancements in fields like drug discovery, protein function annotation, and disease mechanism elucidation.So far, predicting the dynamical properties of biochemical pathways has relied on costly numerical simulations.In this paper, we propose exploiting the topological information in PPINs to restate the problem of predicting pathway robustness as a link prediction task.Our experiments show that the PPIN topology can supply information on inter-pathway relationships, significantly improving predictions of the graph-agnostic baseline relying only on protein sequence embeddings.
Alessandro Dipalma, Domenico Tortorella, Alessio Micheli
ESANN3
2025 Encoding Graph Topology with Randomized Ising Models
abstract
The increasing popularity of deep learning on graphs has motivated the need for the co-design of hardware and graph representation models.We propose Randomized Ising Model (RIM), a reservoir computing model for encoding topological information of graph nodes, that is amenable to physical implementation via neuromorphic hardware.Our experiments demonstrate that RIM's node embeddings are able to provide sufficient topological information to be suitable to address node classification tasks, exhibiting an accuracy in line with Graph Echo State Networks.
Domenico Tortorella, Antonio Brau, Alessio Micheli
ESANN3
2025 Graph Machine Learning for DNA Classification
abstract
Pangenomics is a rapidly evolving field in bioinformatics that enables the study of genetic diversity within populations by representing multiple genomes in a unified structure. Unlike traditional linear reference genomes, pangenomes provide a more comprehensive framework for analyzing genomic variations. Recent advancements in machine learning (ML) for DNA classification, based on large language models (LLMs) such as generative pre-trained transformers (GPT), have achieved state-of-the-art performance. However, these models often require extensive computational resources and do not leverage domain-specific bioinformatics techniques for pangenomic analysis. In this work, we introduce a novel Graph Machine Learning approach for DNA classification (GDNA) based on graph kernels. We represent DNA sequences using de Bruijn graphs, which allow a structured and information-rich representation of genomic data. We compute the similarity between graphs via the Weisfeiler-Lehman (WL) graph kernel and perform classification using a Support Vector Machine (SVM). Experimental evaluation on the Genomic Benchmarks dataset demonstrates that GDNA is competitive with state-of-the-art LLMs such as HyenaDNA, DNABERT, and GPT-based approaches on artificial and real-world DNA classification tasks, without the need for pre-training and with significantly higher efficiency. Overall, GDNA is an efficient and competitive DNA classifier that leverages pangenomes and graph-based representations, providing a valuable tool for genomic and pangenomic analysis in bioinformatics and precision medicine.
Luca Pedrelli, Veronica Guerrini, Nadia Pisanti, Alessio Micheli
IJCNN4
2025 Sensitivity analysis on protein-protein interaction networks through deep graph networks
abstract
BACKGROUND: Protein-protein interaction networks (PPINs) provide a comprehensive view of the intricate biochemical processes that take place in living organisms. In recent years, the size and information content of PPINs have grown thanks to techniques that allow for the functional association of proteins. However, PPINs are static objects that cannot fully describe the dynamics of the protein interactions; these dynamics are usually studied from external sources and can only be added to the PPIN as annotations. In contrast, the time-dependent characteristics of cellular processes are described in Biochemical Pathways (BP), which frame complex networks of chemical reactions as dynamical systems. Their analysis with numerical simulations allows for the study of different dynamical properties. Unfortunately, available BPs cover only a small portion of the interactome, and simulations are often hampered by the unavailability of kinetic parameters or by their computational cost. In this study, we explore the possibility of enriching PPINs with dynamical properties computed from BPs. We focus on the global dynamical property of sensitivity, which measures how a change in the concentration of an input molecular species influences the concentration of an output molecular species at the steady state of the dynamical system. RESULTS: We started with the analysis of BPs via ODE simulations, which enabled us to compute the sensitivity associated with multiple pairs of chemical species. The sensitivity information was then injected into a PPIN, using public ontologies (BioGRID, UniPROT) to map entities at the BP level with nodes at the PPIN level. The resulting annotated PPIN, termed the DyPPIN (Dynamics of PPIN) dataset, was used to train a DGN to predict the sensitivity relationships among PPIN proteins. Our experimental results show that this model can predict these relationships effectively under different use case scenarios. Furthermore, we show that the PPIN structure (i.e., the way the PPIN is "wired") is essential to infer the sensitivity, and that further annotating the PPIN nodes with protein sequence embeddings improves the predictive accuracy. CONCLUSION: To the best of our knowledge, the model proposed in this study is the first that allows performing sensitivity analysis directly on PPINs. Our findings suggest that, despite the high level of abstraction, the structure of the PPIN holds enough information to infer dynamic properties without needing an exact model of the underlying processes. In addition, the designed pipeline is flexible and can be easily integrated into drug design, repurposing, and personalized medicine processes.
Alessandro Dipalma, Michele Fontanesi, Alessio Micheli, Paolo Milazzo, Marco Podda
BMC Bioinform.3
2025 Bridging XAI and spectral analysis to investigate the inductive biases of deep graph networks
abstract
Understanding the inductive bias of Deep Graph Networks (DGNs) is crucial because it reveals how these models generalize from training data to unseen data. Discovering these learning assumptions, and their alignment to the task’s characteristics, allows informed architectural design choices and facilitates interpretation. With this goal, we analyze the different inductive biases of DGNs by relating the node-level explanations produced by explainable AI (XAI) methods to known network science measures of lower-order (local) and higher-order (increasingly global) connectivity. We then apply graph signal processing to refine this analysis at the granularity of the graph frequency spectrum spanned by the explanation signals. Our main finding is that different DGNs focus on different regions of the graph frequency spectrum, and in particular, high-frequency DGNs generalize by focusing on lower-order graph connectivity, while low-frequency DGNs generalize by recognizing higher-order graph structures. This characterization is first derived on synthetic benchmarks by showing that explanations align with network science measures sitting at the two extremes of the spectrum (Katz centrality in the high frequencies, and Fiedler eigenvector scores in the low frequencies). Moving to real-world chemical benchmarks, this result is generalized by showing that inductive biases do indeed lie on a continuum that corresponds to sub-regions of the frequency spectrum.
Michele Fontanesi, Alessio Micheli, Marco Podda, Domenico Tortorella
Mach. Learn.2
2025 Efficient quantification on large-scale networks
abstract
Network quantification (NQ) is the problem of estimating the proportions of nodes belonging to each class in subsets of unlabelled graph nodes. When prior probability shift is at play, this task cannot be effectively addressed by first classifying the nodes and then counting the class predictions. In addition, unlike non-relational quantification, NQ demands enhanced flexibility in order to capture a broad range of connectivity patterns, resilience to the challenge of heterophily, and scalability to large networks. In order to meet these stringent requirements, we introduce XNQ, a novel method that synergizes the flexibility and efficiency of the unsupervised node embeddings computed by randomized recursive Graph Neural Networks, with an Expectation-Maximization algorithm that provides a robust quantification-aware adjustment to the output probabilities of a calibrated node classifier. In an extensive evaluation, in which we also validate the design choices underpinning XNQ through comprehensive ablation experiments, we find that XNQ consistently and significantly improves on the best network quantification methods to date, thereby setting the new state of the art for this challenging task. XNQ also provides a training speed-up of up to 10x–100x over other methods based on graph learning.
Alessio Micheli, Alejandro Moreo, Marco Podda, Fabrizio Sebastiani 0001, William Simoni, Domenico Tortorella
Mach. Learn.1
2025 Enhancing antibody-antigen interaction prediction with atomic flexibility
abstract
Antibodies are indispensable components of the immune system, known for their specific binding to antigens. Beyond their natural immunological functions, they are fundamental in developing vaccines and therapeutic interventions for infectious diseases. The complex architecture of antibodies, particularly their variable regions responsible for antigen recognition, presents significant challenges for computational modeling. Recent advancements in deep learning have markedly improved protein structure prediction; however, accurately modeling antibody-antigen (Ab-Ag) interactions remains challenging due to the inherent flexibility of antibodies and the dynamic nature of binding processes. In this study, we examine the use of predicted Local Distance Difference Test (pLDDT) scores as indicators of residue and side-chain flexibility to model Ab-Ag interactions through a fingerprint-based approach. We demonstrate the significance of flexibility in different antibody-specific tasks, enhancing the predictive accuracy of Ab-Ag interaction models by 4%, resulting in an AUC-ROC of 92%. In addition, we showcase state-of-the-art performance in paratope prediction. These results emphasize the importance of accounting for conformational flexibility in modeling antibody-antigen interactions and show that pLDDT can serve as a coarse proxy for these dynamic features. By optimizing antibody flexibility using pLDDT, they can be engineered to improve affinity or breadth for a specific target. This approach is particularly beneficial for addressing highly variable pathogens like HIV and SARS-CoV-2, as greater flexibility enhances tolerance to sequence variations in target antigens.
Sara Joubbi, Alessio Micheli, Paolo Milazzo, Giorgio Ciano, Stéphane M. Gagné, Pietro Liò, Duccio Medini, Giuseppe Maccari
PLoS Comput. Biol.2
2025 An empirical evaluation of rewiring approaches in graph neural networks
Alessio Micheli, Domenico Tortorella
Pattern Recognit. Lett.1
2024 Analyzing Explanations of Deep Graph Networks Through Node Centrality and Connectivity
Michele Fontanesi, Alessio Micheli, Marco Podda, Domenico Tortorella
DS (1)2
2024 XAI and Bias of Deep Graph Networks
abstract
Generalization in machine learning involves introducing inductive biases that restrict the solution space of the learning problem, allowing for the inductive leap.In this paper, we show the existence of different inductive biases between convolutional and recursive Deep Graph Networks (DGN) by applying Explainable AI (XAI) methods as model inspection techniques.We show that different architectures can perfectly solve the given tasks by learning different labelling policies.Our results promote the usage of different architectures to address a task and raise warnings on the assessment of XAI techniques as their benchmarks may contain more ground truths than those provided.* Research partly funded by PNRR -M4C2 -Investimento 1.3, Partenariato Esteso PE00000013 -"FAIR -Future Artificial Intelligence Research" -Spoke 1 "Human-centered AI", funded by the European Commission under the NextGeneration EU programme.
Michele Fontanesi, Alessio Micheli, Marco Podda
ESANN2
2024 Informed Machine Learning for Complex Data
abstract
In the contemporary era of data-driven decision-making, the application of Machine Learning (ML) on complex data (e.g., images, text, sequences, trees, and graphs) has become increasingly pivotal (e.g., Large Language Models and Graph Neural Networks).In this context, there is a gap between purely data-driven models and domain-specific knowledge, requirements, and expertise.In particular, this domain specificity needs to be integrated into the ML models to improve learning generalization, sustainability, trustworthiness, reliability, security, and safety.This additional knowledge can assume different forms, e.g.: software developers require ML to comply with many technical requirements, companies require ML to comply with economic and environmental sustainability, domain experts require ML to be aligned with physical and logical laws, and society requires ML to be aligned with ethical principles.This special session gathers valuable contributions and early findings in the field of Informed ML for Complex Data.Our main objective is to showcase the potential and limitations of new ideas, improvements, or the blending of ML and other research areas in solving real-world problems.
Luca Oneto, Nicolò Navarin, Alessio Micheli, Luca Pasa, Claudio Gallicchio, Davide Bacciu, Davide Anguita
ESANN3
2024 Continual Learning with Graph Reservoirs: Preliminary experiments in graph classification
abstract
Continual learning aims to address the challenge of catastrophic forgetting in training models where data patterns are non-stationary.Previous research has shown that fully-trained graph learning models are particularly affected by this issue.One approach to lifting part of the burden is to leverage the representations provided by a training-free reservoir computing model.In this work, we evaluate for the first time different continual learning strategies in conjunction with Graph Echo State Networks, which have already demonstrated their efficacy and efficiency in graph classification tasks.Research partly supported by PNRR, PE00000013 -"FAIR -Future Artificial Intelligence Research" -Spoke 1, funded by European Commission under the NextGeneration EU programme.35
Domenico Tortorella, Alessio Micheli
ESANN2
2024 Hybrid CNN-MLP for Wastewater Quality Estimation
Marco Cardia, Stefano Chessa, Alessio Micheli, Antonella Giuliana Luminare, Francesca Gambineri
ICANN (9)3
2024 Onion Echo State Networks - A Preliminary Analysis of Dynamics
Domenico Tortorella, Alessio Micheli
ICANN (10)2
2024 Multitarget Wastewater Quality Assessment in a Smart Industry Context
abstract
This study addresses the need for a rapid and accurate process monitoring by developing an innovative approach for wastewater quality assessment to enhance the Industry 4.0’s vision. By integrating Ultraviolet-Visible (UV-Vis) spectroscopy with Machine Learning (ML), we focus on accurately determine key indicators such as Chemical Oxygen Demand, Total Suspended Solids (TSS), chlorides, and conductivity. Our findings demonstrate the efficacy of ML models in accurately predicting water quality from UV-Vis spectral data, underscoring their potential for real-time monitoring and analysis in industrial settings. The study also revealed the potential for both single and multitarget predictions. Additionally, the feature importance analysis provided valuable insights into the spectral regions most relevant for predicting each water quality indicator. This approach aligns with the goals of Industry 4.0, offering a smart, efficient solution for environmental monitoring and sustainable resource management.
Marco Cardia, Stefano Chessa, Alessio Micheli, Antonella Giuliana Luminare, Massimiliano Franceschi, Francesca Gambineri
IE3
2024 Continuously Deep Recurrent Neural Networks
Andrea Ceni, Peter Ford Dominey, Claudio Gallicchio, Alessio Micheli, Luca Pedrelli, Domenico Tortorella
ECML/PKDD (7)4
2024 Antibody design using deep learning: from sequence and structure design to affinity maturation
abstract
Deep learning has achieved impressive results in various fields such as computer vision and natural language processing, making it a powerful tool in biology. Its applications now encompass cellular image classification, genomic studies and drug discovery. While drug development traditionally focused deep learning applications on small molecules, recent innovations have incorporated it in the discovery and development of biological molecules, particularly antibodies. Researchers have devised novel techniques to streamline antibody development, combining in vitro and in silico methods. In particular, computational power expedites lead candidate generation, scaling and potential antibody development against complex antigens. This survey highlights significant advancements in protein design and optimization, specifically focusing on antibodies. This includes various aspects such as design, folding, antibody-antigen interactions docking and affinity maturation.
Sara Joubbi, Alessio Micheli, Paolo Milazzo, Giuseppe Maccari, Giorgio Ciano, Dario Cardamone, Duccio Medini
Briefings Bioinform.2
2024 Investigating over-parameterized randomized graph networks
abstract
In this paper, we investigate neural models based on graph random features for classification tasks. First, we aim to understand when over parameterization, namely generating more features than the ones necessary to interpolate, may be beneficial for the generalization abilities of the resulting models. We employ two measures: one from the algorithmic stability framework and another one based on information theory. We provide empirical evidence from several commonly adopted graph datasets showing that the considered measures, even without considering task labels, can be effective for this purpose. Additionally, we investigate whether these measures can aid in the process of hyperparameters selection. The results of our empirical analysis show that the considered measures have good correlations with the estimated generalization performance of the models with different hyperparameter configurations. Moreover, they can be used to identify good hyperparameters, achieving results comparable to the ones obtained with a classic grid search.
Giovanni Donghi, Luca Pasa, Luca Oneto, Claudio Gallicchio, Alessio Micheli, Davide Anguita, Alessandro Sperduti, Nicolò Navarin
Neurocomputing5
2024 Designs of graph echo state networks for node classification
abstract
Among the Graph Neural Network (GNN) models that address the task of node classification, Graph Echo State Networks (GESN) have proved particularly effective in addressing the challenge of heterophily, i.e. the presence of a significant fraction of inter-class edges in the learning task graph. The effectiveness of GESN is paired with its efficiency, owing to the reservoir computing paradigm. While previous literature has analyzed the design of reservoirs for sequence ESN and GESN for graph-level tasks, the problem of providing effective designs of reservoirs for node-level GESN is so far largely unexplored. In this paper we analyze the impact of different reservoir designs on node classification accuracy and on the quality of node embeddings computed by GESN, focusing both on dense and sparse reservoir layouts. As measures of embedding richness, we adopt both graph topology-dependent metrics previously employed in the analysis of embedding smoothing, and topology-independent metrics from the areas of information theory and numerical analysis. In particular, we propose the application of entropy measures for quantifying information in node embeddings.
Alessio Micheli, Domenico Tortorella
Neurocomputing1
2024 Guest Editorial: Deep Neural Networks for Graphs: Theory, Models, Algorithms, and Applications
abstract
Deep neural networks for graphs (DNNGs) represent an emerging field that studies how the deep learning method can be generalized to graph-structured data. Since graphs are a powerful and flexible tool to represent complex information in the form of patterns and their relationships, ranging from molecules to protein-to-protein interaction networks, to social or transportation networks, or up to knowledge graphs, potentially modeling systems at very different scales, these methods have been exploited for many application domains.
Ming Li 0065, Alessio Micheli, Yu Guang Wang 0001, Shirui Pan, Pietro Liò, Giorgio Gnecco, Marcello Sanguineti
IEEE Trans. Neural Networks Learn. Syst.2
2023 Graph Representation Learning
abstract
In a broad range of real-world machine learning applications, representing examples as graphs is crucial to avoid a loss of information.For this reason, in the last few years, the definition of machine learning methods, particularly neural networks, for graph-structured inputs has been gaining increasing attention.In particular, Deep Graph Networks (DGNs) are nowadays the most commonly adopted models to learn a representation that can be used to address different tasks related to nodes, edges, or even entire graphs.This tutorial paper reviews fundamental concepts and open challenges of graph representation learning and summarizes the contributions that have been accepted for publication to the ESANN 2023 special session on the topic.
Davide Bacciu, Federico Errica, Alessio Micheli, Nicolò Navarin, Luca Pasa, Marco Podda, Daniele Zambon
ESANN3
2023 Hidden Markov Models for Temporal Graph Representation Learning
abstract
We propose the Hidden Markov Model for temporal Graphs, a deep and fully probabilistic model for learning in the domain of dynamic time-varying graphs.We extend hidden Markov models for sequences to the graph domain by stacking probabilistic layers that perform efficient message passing and learn representations for the individual nodes.We evaluate the goodness of the learned representations on temporal node prediction tasks, and we observe promising results compared to neural approaches.
Federico Errica, Alessio Gravina, Davide Bacciu, Alessio Micheli
ESANN4
2023 Entropy Based Regularization Improves Performance in the Forward-Forward Algorithm
abstract
The forward-forward algorithm (FFA) is a recently proposed alternative to end-to-end backpropagation in deep neural networks.FFA builds networks greedily layer by layer, thus being of particular interest in applications where memory and computational constraints are important.In order to boost layers' ability to transfer useful information to subsequent layers, in this paper we propose a novel regularization term for the layerwise loss function that is based on Renyi's quadratic entropy.Preliminary experiments show accuracy is generally significantly improved across all network architectures.In particular, smaller architectures become more effective in addressing our classification tasks compared to the original FFA.
Matteo Pardi, Domenico Tortorella, Alessio Micheli
ESANN3
2023 Richness of Node Embeddings in Graph Echo State Networks
abstract
Graph Echo State Networks (GESN) have recently proved effective in node classification tasks, showing particularly able to address the issue of heterophily.While previous literature has analyzed the design of reservoirs for sequence ESN and GESN for graph-level tasks, the factors that contribute to rich node embeddings are so far unexplored.In this paper we analyze the impact of different reservoir designs on node classification accuracy and on the quality of node embeddings computed by GESN using tools from the areas of information theory and numerical analysis.In particular, we propose an entropy measure for quantifying information in node embeddings.
Domenico Tortorella, Alessio Micheli
ESANN2
2023 Analysis and Interpretation of ECG Time Series Through Convolutional Neural Networks in Brugada Syndrome Diagnosis
Alessio Micheli, Marco Natali, Luca Pedrelli, Lorenzo Simone, Maria-Aurora Morales, Marcello Piacenti, Federico Vozzi
ICANN (4)1
2023 Exploiting the structure of biochemical pathways to investigate dynamical properties with neural networks for graphs
abstract
MOTIVATION: Dynamical properties of biochemical pathways (BPs) help in understanding the functioning of living cells. Their in silico assessment requires simulating a dynamical system with a large number of parameters such as kinetic constants and species concentrations. Such simulations are based on numerical methods that can be time-expensive for large BPs. Moreover, parameters are often unknown and need to be estimated. RESULTS: We developed a framework for the prediction of dynamical properties of BPs directly from the structure of their graph representation. We represent BPs as Petri nets, which can be automatically generated, for instance, from standard SBML representations. The core of the framework is a neural network for graphs that extracts relevant information directly from the Petri net structure and exploits them to learn the association with the desired dynamical property. We show experimentally that the proposed approach reliably predicts a range of diverse dynamical properties (robustness, monotonicity, and sensitivity) while being faster than numerical methods at prediction time. In synergy with the neural network models, we propose a methodology based on Petri nets arc knock-out that allows the role of each molecule in the occurrence of a certain dynamical property to be better elucidated. The methodology also provides insights useful for interpreting the predictions made by the model. The results support the conjecture often considered in the context of systems biology that the BP structure plays a primary role in the assessment of its dynamical properties. AVAILABILITY AND IMPLEMENTATION: https://github.com/marcopodda/petri-bio (code), https://zenodo.org/record/7610382 (data).
Michele Fontanesi, Alessio Micheli, Paolo Milazzo, Marco Podda
Bioinform.2
2023 Addressing heterophily in node classification with graph echo state networks
abstract
Node classification tasks on graphs are addressed via fully-trained deep message-passing models that learn a hierarchy of node representations via multiple aggregations of a node’s neighbourhood. While effective on graphs that exhibit a high ratio of intra-class edges, this approach poses challenges in the opposite case, i.e. heterophily, where nodes belonging to the same class are usually further apart. In graphs with a high degree of heterophily, the smoothed representations based on close neighbours computed by convolutional models are no longer effective. So far, architectural variations in message-passing models to reduce excessive smoothing or rewiring the input graph to improve longer-range message passing have been proposed. In this paper, we address the challenges of heterophilic graphs with Graph Echo State Network (GESN) for node classification. GESN is a reservoir computing model for graphs, where node embeddings are recursively computed by an untrained message-passing function. Our experiments show that reservoir models are able to achieve better or comparable accuracy with respect to most fully trained deep models that implement ad hoc variations in the architectural bias or perform rewiring as a preprocessing step on the input graph, with an improvement in terms of efficiency/accuracy trade-off. Furthermore, our analysis shows that GESN is able to effectively encode the structural relationships of a graph node, by showing a correlation between iterations of the recursive embedding function and the distribution of shortest paths in a graph.
Alessio Micheli, Domenico Tortorella
Neurocomputing1
2023 Architectural richness in deep reservoir computing
Claudio Gallicchio, Alessio Micheli
Neural Comput. Appl.2
2022 Input Routed Echo State Networks
abstract
We introduce a novel Reservoir Computing (RC) approach for multi-dimensional temporal signals.Our proposal is based on routing the different dimensions of the driving input towards different dynamical sub-modules in a multi-reservoir architecture.At the same time, controllable interconnections among the sub-modules allow modeling the interplay between the different dynamics that might be required by the task.Experiments on synthetic and real-world time-series classification problems clearly show the advantages of the proposed approach in dealing with multi-dimensional signals in comparison to standard RC neural networks.
Luca Argentieri, Claudio Gallicchio, Alessio Micheli
ESANN3
2022 Beyond Homophily with Graph Echo State Networks
abstract
Graph Echo State Networks (GESN) have already demonstrated their efficacy and efficiency in graph classification tasks.However, semi-supervised node classification brought out the problem of oversmoothing in end-to-end trained deep models, which causes a bias towards high homophily graphs.We evaluate for the first time GESN on node classification tasks with different degrees of homophily, analyzing also the impact of the reservoir radius.Our experiments show that reservoir models are able to achieve better or comparable accuracy with respect to fully trained deep models that implement ad hoc variations in the architectural bias, with a gain in terms of efficiency.
Domenico Tortorella, Alessio Micheli
ESANN2
2022 Hierarchical Dynamics in Deep Echo State Networks
Domenico Tortorella, Claudio Gallicchio, Alessio Micheli
ICANN (3)3
2022 The Infinite Contextual Graph Markov Model
abstract
The Contextual Graph Markov Model (CGMM) is a deep, unsupervised, and probabilistic model for graphs that is trained incrementally on a layer-by-layer basis. As with most Deep Graph Networks, an inherent limitation is the need to perform an extensive model selection to choose the proper size of each layer’s latent representation. In this paper, we address this problem by introducing the Infinite Contextual Graph Markov Model (iCGMM), the first deep Bayesian nonparametric model for graph learning. During training, iCGMM can adapt the complexity of each layer to better fit the underlying data distribution. On 8 graph classification tasks, we show that iCGMM: i) successfully recovers or improves CGMM’s performances while reducing the hyper-parameters’ search space; ii) performs comparably to most end-to-end supervised methods. The results include studies on the importance of depth, hyper-parameters, and compression of the graph embeddings. We also introduce a novel approximated inference procedure that better deals with larger graph topologies.
Daniele Castellana, Federico Errica, Davide Bacciu, Alessio Micheli
ICML4
2022 Spectral Bounds for Graph Echo State Network Stability
abstract
Graph echo state networks (GESN) are a class of reservoir computing models for the efficient and effective processing of graphs. They compute graph embeddings by the convergence to a fixed point of a dynamical system, randomly initialized according to a generalization of the echo state property, called the graph embedding stability (GES) property. In this paper, we prove new and more accurate bounds for necessary and sufficient GES conditions. Experiments demonstrate how these bounds allow an easier parameter selection and better quality reservoirs.
Domenico Tortorella, Claudio Gallicchio, Alessio Micheli
IJCNN3
2022 Pyramidal Reservoir Graph Neural Network
Filippo Maria Bianchi, Claudio Gallicchio, Alessio Micheli
Neurocomputing3
2022 Discrete-time dynamic graph echo state networks
Alessio Micheli, Domenico Tortorella
Neurocomputing1
2022 Towards learning trustworthily, automatically, and with guarantees on graphs: An overview
Luca Oneto, Nicolò Navarin, Battista Biggio, Federico Errica, Alessio Micheli, Franco Scarselli, Monica Bianchini, Luca Demetrio, Pietro Bongini, Armando Tacchella, Alessandro Sperduti
Neurocomputing5
2022 Guest Editorial Special Issue on New Frontiers in Extremely Efficient Reservoir Computing
abstract
With the penetration of artificial intelligence (AI) technology into industrial applications, not only computational effectiveness but also computational efficiency in machine learning (ML) methods has been increasingly demanded. Reservoir computing (RC) is an ML framework leveraging a dynamicreservoirfor a nonlinear transformation of sequential inputs and areadoutfor mapping the reservoir state to a desired output. Since only the readout is trained with a simple learning algorithm, RC has attracted much attention as a promising approach to enhance compatibility between high computational performance and low learning cost. In addition, recent studies on physical reservoirs implemented with various physical substrates have boosted the potential of RC in the development of effective and efficient AI hardware. Therefore, it is time to further explore the new frontiers in extremely efficient RC.
Gouhei Tanaka, Claudio Gallicchio, Alessio Micheli, Juan-Pablo Ortega, Akira Hirose 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Robust Malware Classification via Deep Graph Networks on Call Graph Topologies
abstract
We propose a malware classification system that is shown to be robust to some common intra-procedural obfuscation techniques.Indeed, by training the Contextual Graph Markov Model on the call graph representation of a program, we classify it using only topological information, which is unaffected by such obfuscations.In particular, we show that the structure of the call graph is sufficient to achieve good accuracy on a multi-class classification benchmark.
Federico Errica, Giacomo Iadarola, Fabio Martinelli, Francesco Mercaldo, Alessio Micheli
ESANN5
2021 Complex Data: Learning Trustworthily, Automatically, and with Guarantees
abstract
Machine Learning (ML) achievements enabled automatic extraction of actionable information from data in a wide range of decisionmaking scenarios.This demands for improving both ML technical aspects (e.g., design and automation) and human-related metrics (e.g., fairness, robustness, privacy, and explainability), with performance guarantees at both levels.The aforementioned scenario posed three main challenges: (i) Learning from Complex Data (i.e., sequence, tree, and graph data), (ii) Learning Trustworthily, and (iii) Learning Automatically with Guarantees.The focus of this special session is on addressing one or more of these challenges with the final goal of Learning Trustworthily, Automatically, and with Guarantees from Complex Data.
Luca Oneto, Nicolò Navarin, Battista Biggio, Federico Errica, Alessio Micheli, Franco Scarselli, Monica Bianchini, Alessandro Sperduti
ESANN5
2021 Dynamic Graph Echo State Networks
abstract
Dynamic temporal graphs represent evolving relations between entities, e.g.interactions between social network users or infection spreading.We propose an extension of graph echo state networks for the efficient processing of dynamic temporal graphs, with a sufficient condition for their echo state property, and an experimental analysis of reservoir layout impact.Compared to temporal graph kernels that need to hold the entire history of vertex interactions, our model provides a vector encoding for the dynamic graph that is updated at each time-step without requiring training.Experiments show accuracy comparable to approximate temporal graph kernels on twelve dissemination process classification tasks.
Domenico Tortorella, Alessio Micheli
ESANN2
2021 Graph Mixture Density Networks
abstract
We introduce the Graph Mixture Density Networks, a new family of machine learning models that can fit multimodal output distributions conditioned on graphs of arbitrary topology. By combining ideas from mixture models and graph representation learning, we address a broader class of challenging conditional density estimation problems that rely on structured data. In this respect, we evaluate our method on a new benchmark application that leverages random graphs for stochastic epidemic simulations. We show a significant improvement in the likelihood of epidemic outcomes when taking into account both multimodality and structure. The empirical analysis is complemented by two real-world regression tasks showing the effectiveness of our approach in modeling the output prediction uncertainty. Graph Mixture Density Networks open appealing research opportunities in the study of structure-dependent phenomena that exhibit non-trivial conditional output distributions.
Federico Errica, Davide Bacciu, Alessio Micheli
ICML3
2021 Modeling Edge Features with Deep Bayesian Graph Networks
abstract
We propose an extension of the Contextual Graph Markov Model, a deep and probabilistic machine learning model for graphs, to model the distribution of edge features. Our approach is architectural, as we introduce an additional Bayesian network mapping edge features into discrete states to be used by the original model. In doing so, we are also able to build richer graph representations even in the absence of edge features, which is confirmed by the performance improvements on standard graph classification benchmarks. Moreover, we successfully test our proposal in a graph regression scenario where edge features are of fundamental importance, and we show that the learned edge representation provides substantial performance improvements against the original model on three link prediction tasks. By keeping the computational complexity linear in the number of edges, the proposed model is amenable to large-scale graph processing.
Daniele Atzeni, Davide Bacciu, Federico Errica, Alessio Micheli
IJCNN4
2021 Federated Reservoir Computing Neural Networks
abstract
A critical aspect in Federated Learning is the aggregation strategy for the combination of multiple models, trained on the edge, into a single model that incorporates all the knowledge in the federation. Common Federated Learning approaches for Recurrent Neural Networks (RNNs) do not provide guarantees on the predictive performance of the aggregated model. In this paper we show how the use of Echo State Networks (ESNs), which are efficient state-of-the-art RNN models for time-series processing, enables a form of federation that is optimal in the sense that it produces models mathematically equivalent to the corresponding centralized model. Furthermore, the proposed method is compliant with privacy constraints. The proposed method, which we denote as Incremental Federated Learning, is experimentally evaluated against an averaging strategy on two datasets for human state and activity recognition.
Davide Bacciu, Daniele Di Sarli, Pouria Faraji, Claudio Gallicchio, Alessio Micheli
IJCNN5
2021 Phase Transition Adaptation
abstract
Artificial Recurrent Neural Networks are a powerful information processing abstraction, and Reservoir Computing provides an efficient strategy to build robust implementations by projecting external inputs into high dimensional dynamical system trajectories. In this paper, we propose an extension of the original approach, a local unsupervised learning mechanism we call Phase Transition Adaptation, designed to drive the system dynamics towards the ‘edge of stability’. Here, the complex behavior exhibited by the system elicits an enhancement in its overall computational capacity. We show experimentally that our approach consistently achieves its purpose over several datasets.
Claudio Gallicchio, Alessio Micheli, Luca Silvestri
IJCNN2
2020 Fast and Deep Graph Neural Networks
abstract
We address the efficiency issue for the construction of a deep graph neural network (GNN). The approach exploits the idea of representing each input graph as a fixed point of a dynamical system (implemented through a recurrent neural network), and leverages a deep architectural organization of the recurrent units. Efficiency is gained by many aspects, including the use of small and very sparse networks, where the weights of the recurrent units are left untrained under the stability condition introduced in this work. This can be viewed as a way to study the intrinsic power of the architecture of a deep GNN, and also to provide insights for the set-up of more complex fully-trained models. Through experimental results, we show that even without training of the recurrent connections, the architecture of small deep GNN is surprisingly able to achieve or improve the state-of-the-art performance on a significant set of tasks in the field of graphs classification.
Claudio Gallicchio, Alessio Micheli
AAAI2
2020 A Deep Generative Model for Fragment-Based Molecule Generation
abstract
Molecule generation is a challenging open problem in cheminformatics. Currently, deep generative approaches addressing the challenge belong to two broad categories, differing in how molecules are represented. One approach encodes molecular graphs as strings of text, and learns their corresponding character-based language model. Another, more expressive, approach operates directly on the molecular graph. In this work, we address two limitations of the former: generation of invalid and duplicate molecules. To improve validity rates, we develop a language model for small molecular substructures called fragments, loosely inspired by the well-known paradigm of Fragment-Based Drug Design. In other words, we generate molecules fragment by fragment, instead of atom by atom. To improve uniqueness rates, we present a frequency-based masking strategy that helps generate molecules with infrequent fragments. We show experimentally that our model largely outperforms other language model-based competitors, reaching state-of-the-art performances typical of graph-based approaches. Moreover, generated molecules display molecular properties similar to those in the training sample, even in absence of explicit task-specific supervision.
Marco Podda, Davide Bacciu, Alessio Micheli
AISTATS3
2020 Efficient Embedded Machine Learning applications using Echo State Networks
abstract
The increasing role of Artificial Intelligence (AI) and Machine Learning (ML) in our lives brought a paradigm shift on how and where the computation is performed. Stringent latency requirements and congested bandwidth moved AI inference from Cloud space towards end-devices. This change required a major simplification of Deep Neural Networks (DNN), with memory-wise libraries or co-processors that perform fast inference with minimal power. Unfortunately, many applications such as natural language processing, time-series analysis and audio interpretation are built on a different type of Artifical Neural Networks (ANN), the so-called Recurrent Neural Networks (RNN), which, due to their intrinsic architecture, remains too complex and heavy to run efficiently on embedded devices. To solve this issue, the Reservoir Computing paradigm proposes sparse untrained non-linear networks, the Reservoir, that can embed temporal relations without some of the hindrances of Recurrent Neural Networks training, and with a lower memory usage. Echo State Networks (ESN) and Liquid State Machines are the most notable examples. In this scenario, we propose a performance comparison of a ESN, designed and trained using Bayesian Optimization techniques, against current RNN solutions. We aim to demonstrate that ESN have comparable performance in terms of accuracy, require minimal training time, and they are more optimized in terms of memory usage and computational efficiency. Preliminary results show that ESN are competitive with RNN on a simple benchmark, and both training and inference time are faster, with a maximum speed-up of 2.35x and 6.60x, respectively.
Luca Cerina, Marco D. Santambrogio, Giuseppe Franco, Claudio Gallicchio, Alessio Micheli
DATE5
2020 Pyramidal Graph Echo State Networks
Filippo Maria Bianchi, Claudio Gallicchio, Alessio Micheli
ESANN3
2020 Theoretically Expressive and Edge-aware Graph Learning
Federico Errica, Davide Bacciu, Alessio Micheli
ESANN3
2020 Simplifying Deep Reservoir Architectures
Claudio Gallicchio, Alessio Micheli, Antonio Sisbarra
ESANN2
2020 Biochemical Pathway Robustness Prediction with Graph Neural Networks
Marco Podda, Alessio Micheli, Davide Bacciu, Paolo Milazzo
ESANN2
2020 Time Series Clustering with Deep Reservoir Computing
Miguel A. Atencia Ruiz, Claudio Gallicchio, Gonzalo Joya Caparrós, Alessio Micheli
ICANN (2)4
2020 A Preliminary Investigation of Machine Learning Approaches for Mobility Monitoring from Smartphone Data
Claudio Gallicchio, Alessio Micheli, Massimiliano Petri, Antonio Pratelli
ICCSA (2)2
2020 A Fair Comparison of Graph Neural Networks for Graph Classification
Federico Errica, Marco Podda, Davide Bacciu, Alessio Micheli
ICLR4
2020 Ring Reservoir Neural Networks for Graphs
abstract
Machine Learning for graphs is nowadays a research topic of consolidated relevance. Common approaches in the field typically resort to complex deep neural network architectures and demanding training algorithms, highlighting the need for more efficient solutions. The class of Reservoir Computing (RC) models can play an important role in this context, enabling to develop fruitful graph embeddings through untrained recursive architectures. In this paper, we study progressive simplifications to the design strategy of RC neural networks for graphs. Our core proposal is based on shaping the organization of the hidden neurons to follow a ring topology. Experimental results on graph classification tasks indicate that ring-reservoirs architectures enable particularly effective network configurations, showing consistent advantages in terms of predictive performance.
Claudio Gallicchio, Alessio Micheli
IJCNN2
2020 Gated Echo State Networks: a preliminary study
abstract
Gating mechanisms are widely used in the context of Recurrent Neural Networks (RNNs) to improve the network's ability to deal with long-term dependencies within the data. The typical approach for training such networks involves the expensive algorithm of gradient descent and backpropagation. On the other hand, Reservoir Computing (RC) approaches like Echo State Networks (ESNs) are extremely efficient in terms of training time and resources thanks to their use of randomly initialized parameters that do not need to be trained. Unfortunately, basic ESNs are also unable to effectively deal with complex long-term dependencies. In this work, we start investigating the problem of equipping ESNs with gating mechanisms. Under rigorous experimental settings, we compare the behaviour of an ESN with randomized gate parameters (initialized with RC techniques) against several other models, among which a leaky ESN and a fully trained gated RNN. We observe that the use of randomized gates by itself can increase the predictive accuracy of a ESN, but this increase is not meaningful when compared with other techniques. Given these results, we propose a research direction for successfully designing ESN models with gating mechanisms.
Daniele Di Sarli, Claudio Gallicchio, Alessio Micheli
INISTA3
2020 Edge-based sequential graph generation with recurrent neural networks
Davide Bacciu, Alessio Micheli, Marco Podda
Neurocomputing2
2020 Probabilistic Learning on Graphs via Contextual Architectures
abstract
We propose a novel methodology for representation learning on graph-structured data, in which a stack of Bayesian Networks learns different distributions of a vertex's neighbourhood. Through an incremental construction policy and layer-wise training, we can build deeper architectures with respect to typical graph convolutional neural networks, with benefits in terms of context spreading between vertices. First, the model learns from graphs via maximum likelihood estimation without using target labels. Then, a supervised readout is applied to the learned graph embeddings to deal with graph classification and vertex classification tasks, showing competitive results against neural models for graphs. The computational complexity is linear in the number of edges, facilitating learning on large scale data sets. By studying how depth affects the performances of our model, we discover that a broader context generally improves performances. In turn, this leads to a critical analysis of some benchmarks used in literature.
Davide Bacciu, Federico Errica, Alessio Micheli
J. Mach. Learn. Res.3
2020 A gentle introduction to deep learning for graphs
Davide Bacciu, Federico Errica, Alessio Micheli, Marco Podda
Neural Networks3
2020 EchoBay: Design and Optimization of Echo State Networks under Memory and Time Constraints
abstract
The increase in computational power of embedded devices and the latency demands of novel applications brought a paradigm shift on how and where the computation is performed. Although AI inference is slowly moving from the cloud to end-devices with limited resources, time-centric recurrent networks like Long-Short Term Memory remain too complex to be transferred on embedded devices without extreme simplifications and limiting the performance of many notable applications. To solve this issue, the Reservoir Computing paradigm proposes sparse, untrained non-linear networks, the Reservoir, that can embed temporal relations without some of the hindrances of Recurrent Neural Networks training, and with a lower memory occupation. Echo State Networks (ESN) and Liquid State Machines are the most notable examples. In this scenario, we propose EchoBay , a comprehensive C++ library for ESN design and training. EchoBay is architecture-agnostic to guarantee maximum performance on different devices (whether embedded or not), and it offers the possibility to optimize and tailor an ESN on a particular case study, reducing at the minimum the effort required on the user side. This can be done thanks to the Bayesian Optimization (BO) process, which efficiently and automatically searches hyper-parameters that maximize a fitness function. Additionally, we designed different optimization techniques that take in consideration resource constraints of the device to minimize memory footprint and inference time. Our results in different scenarios show an average speed-up in training time of 119x compared to Grid and Random search of hyper-parameters, a decrease of 94% of trained models size and 95% in inference time, maintaining comparable performance for the given task. The EchoBay library is Open Source and publicly available at https://github.com/necst/Echobay.
Luca Cerina, Marco D. Santambrogio, Giuseppe Franco, Claudio Gallicchio, Alessio Micheli
ACM Trans. Archit. Code Optim.5
2019 Graph generation by sequential edge prediction
Davide Bacciu, Alessio Micheli, Marco Podda
ESANN2
2019 Comparison between DeepESNs and gated RNNs on multivariate time-series prediction
Claudio Gallicchio, Alessio Micheli, Luca Pedrelli
ESANN2
2019 Embeddings and Representation Learning for Structured Data
Benjamin Paaßen, Claudio Gallicchio, Alessio Micheli, Alessandro Sperduti
ESANN3
2019 An ambient intelligence approach for learning in smart robotic environments
abstract
Abstract Smart robotic environments combine traditional (ambient) sensing devices and mobile robots. This combination extends the type of applications that can be considered, reduces their complexity, and enhances the individual values of the devices involved by enabling new services that cannot be performed by a single device. To reduce the amount of preparation and preprogramming required for their deployment in real‐world applications, it is important to make these systems self‐adapting. The solution presented in this paper is based upon a type of compositional adaptation where (possibly multiple) plans of actions are created through planning and involve the activation of pre‐existing capabilities. All the devices in the smart environment participate in a pervasive learning infrastructure, which is exploited to recognize which plans of actions are most suited to the current situation. The system is evaluated in experiments run in a real domestic environment, showing its ability to proactively and smoothly adapt to subtle changes in the environment and in the habits and preferences of their user(s), in presence of appropriately defined performance measuring functions.
Davide Bacciu, Maurizio Di Rocco, Mauro Dragone, Claudio Gallicchio, Alessio Micheli, Alessandro Saffiotti
Comput. Intell.5
2019 Deep Reservoir Neural Networks for Trees
Claudio Gallicchio, Alessio Micheli
Inf. Sci.2
2019 Editorial: Booming of Neural Networks and Learning Systems
abstract
As you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community.
Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He
IEEE Trans. Neural Networks Learn. Syst.2
2018 Deep Echo State Networks for Diagnosis of Parkinson's Disease
Claudio Gallicchio, Alessio Micheli, Luca Pedrelli
ESANN2
2018 Randomized Recurrent Neural Networks
Claudio Gallicchio, Alessio Micheli, Peter Tiño
ESANN2
2018 Combining Memory and Non-linearity in Echo State Networks
Eleonora Di Gregorio, Claudio Gallicchio, Alessio Micheli
ICANN (2)3
2018 Contextual Graph Markov Model: A Deep and Generative Approach to Graph Processing
abstract
We introduce the Contextual Graph Markov Model, an approach combining ideas from generative models and neural networks for the processing of graph data. It founds on a constructive methodology to build a deep architecture comprising layers of probabilistic models that learn to encode the structured information in an incremental fashion. Context is diffused in an efficient and scalable way across the graph vertexes and edges. The resulting graph encoding is used in combination with discriminative models to address structure classification benchmarks.
Davide Bacciu, Federico Errica, Alessio Micheli
ICML3
2018 Tree Edit Distance Learning via Adaptive Symbol Embeddings
abstract
Metric learning has the aim to improve classification accuracy by learning a distance measure which brings data points from the same class closer together and pushes data points from different classes further apart. Recent research has demonstrated that metric learning approaches can also be applied to trees, such as molecular structures, abstract syntax trees of computer programs, or syntax trees of natural language, by learning the cost function of an edit distance, i.e. the costs of replacing, deleting, or inserting nodes in a tree. However, learning such costs directly may yield an edit distance which violates metric axioms, is challenging to interpret, and may not generalize well. In this contribution, we propose a novel metric learning approach for trees which we call embedding edit distance learning (BEDL) and which learns an edit distance indirectly by embedding the tree nodes as vectors, such that the Euclidean distance between those vectors supports class discrimination. We learn such embeddings by reducing the distance to prototypical trees from the same class and increasing the distance to prototypical trees from different classes. In our experiments, we show that BEDL improves upon the state-of-the-art in metric learning for trees on six benchmark data sets, ranging from computer science over biomedical data to a natural-language processing data set containing over 300,000 nodes.
Benjamin Paaßen, Claudio Gallicchio, Alessio Micheli, Barbara Hammer
ICML3
2018 Why Layering in Recurrent Neural Networks? A DeepESN Survey
abstract
The extension of Recurrent Neural Networks (RNNs) in the direction of deep learning is a topic that is gaining more and more attention in the neural networks community. The study of deep RNNs opened a number of intriguing research questions on the actual role played by layering in the architectural design of RNNs. Recently, the introduction of the Deep Echo State Network (DeepESN) model allowed to start addressing such open issues in literature, contributing to shed light on the intrinsic properties of state dynamics developed by hierarchical compositions of recurrent layers. This contribution intends to present a unified view over the major advancements in the study of DeepESNs, enabling to directly point out the natural advantages of a layered construction of recurrent networks for temporal data processing.
Claudio Gallicchio, Alessio Micheli
IJCNN2
2018 Deep Tree Echo State Networks
abstract
This work proposes a first study, through empirical assessment, of a deep recursive Neural Network (RecNN) architecture for tree structured data exploiting the efficient design of the Echo State Network (ESN) framework. Three benchmark tasks for trees allow us to assess the potentiality of the novel Deep Tree ESN (DeepTESN) model with respect to the shallow counterpart (Tree ESN) and literature results (including hidden tree Markov models and kernel based approaches) in different conditions and according to both efficiency and predictive performance.
Claudio Gallicchio, Alessio Micheli
IJCNN2
2018 Local Lyapunov exponents of deep echo state networks
Claudio Gallicchio, Alessio Micheli, Luca Silvestri
Neurocomputing2
2018 Design of deep echo state networks
Claudio Gallicchio, Alessio Micheli, Luca Pedrelli
Neural Networks2
2018 Generative Kernels for Tree-Structured Data
abstract
This paper presents a family of methods for the design of adaptive kernels for tree-structured data that exploits the summarization properties of hidden states of hidden Markov models for trees. We introduce a compact and discriminative feature space based on the concept of hidden states multisets and we discuss different approaches to estimate such hidden state encoding. We show how it can be used to build an efficient and general tree kernel based on Jaccard similarity. Furthermore, we derive an unsupervised convolutional generative kernel using a topology induced on the Markov states by a tree topographic mapping. This paper provides an extensive empirical assessment on a variety of structured data learning tasks, comparing the predictive accuracy and computational efficiency of state-of-the-art generative, adaptive, and syntactical tree kernels. The results show that the proposed generative approach has a good tradeoff between computational complexity and predictive performance, in particular when considering the soft matching introduced by the topographic mapping.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IEEE Trans. Neural Networks Learn. Syst.2
2017 Randomized Machine Learning Approaches: Recent Developments and Challenges
Claudio Gallicchio, José D. Martín-Guerrero, Alessio Micheli, Emilio Soria-Olivas
ESANN3
2017 Local Lyapunov Exponents of Deep RNN
Claudio Gallicchio, Alessio Micheli, Luca Silvestri
ESANN2
2017 A learning system for automatic Berg Balance Scale score estimation
Davide Bacciu, Stefano Chessa, Claudio Gallicchio, Alessio Micheli, Luca Pedrelli, Erina Ferro, Luigi Fortunati, Davide La Rosa, Filippo Palumbo, Federico Vozzi, Oberdan Parodi
Eng. Appl. Artif. Intell.4
2017 Deep reservoir computing: A critical experimental analysis
Claudio Gallicchio, Alessio Micheli, Luca Pedrelli
Neurocomputing2
2016 A reservoir activation kernel for trees
Davide Bacciu, Claudio Gallicchio, Alessio Micheli
ESANN3
2016 RSS-based Robot Localization in Critical Environments using Reservoir Computing
Mauro Dragone, Claudio Gallicchio, Roberto Guzmán, Alessio Micheli
ESANN4
2016 Deep Reservoir Computing: A Critical Analysis
Claudio Gallicchio, Alessio Micheli
ESANN2
2016 Detecting Socialization Events in Ageing People: The Experience of the DOREMI Project
abstract
The detection of socialization events is useful to build indicators about social isolation of people, which is an important indicator in e-health applications. On the other hand, it is rather difficult to achieve with non-invasive solutions. This paper reports about the currently work-in-progress on the technological solution for the detection of socialization events adopted in the DOREMI project.
Davide Bacciu, Stefano Chessa, Erina Ferro, Luigi Fortunati, Claudio Gallicchio, Davide La Rosa, Miguel Llorente, Alessio Micheli, Filippo Palumbo, Oberdan Parodi, Andrea Valenti, Federico Vozzi
Intelligent Environments8
2015 ESNigma: efficient feature selection for echo state networks
Davide Bacciu, Filippo Benedetti, Alessio Micheli
ESANN3
2015 A cognitive robotic ecology approach to self-configuring and evolving AAL systems
Mauro Dragone, Giuseppe Amato 0001, Davide Bacciu, Stefano Chessa, Sonya A. Coleman, Maurizio Di Rocco, Claudio Gallicchio, Claudio Gennaro, Héctor Lozano Peiteado, Liam P. Maguire, T. Martin McGinnity, Alessio Micheli, Gregory M. P. O'Hare, Arantxa Rentería, Alessandro Saffiotti, Claudio Vairo, Philip J. Vance
Eng. Appl. Artif. Intell.12
2015 Prediction of the Italian electricity price for smart grid applications
Emanuele Crisostomi, Claudio Gallicchio, Alessio Micheli, Marco Raugi, Mauro Tucci
Neurocomputing3
2014 Modeling Bi-directional Tree Contexts by Generative Transductions
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ICONIP (1)2
2014 Integrating bi-directional contexts in a generative kernel for trees
abstract
Context is essential to evaluate an atomic piece of information composing an articulated structured sample. A particular context captures different structural information with respect to an alternative context. The paper introduces a generative kernel that easily and effectively combines the structural information captured by generative tree models characterized by different contextual capabilities. The proposed approach exploits the idea of hidden states multisets to realize a tree encoding that takes into account both the summarized information on the path leading to a node (i.e. a top-down context) as well as the information on how substructures are composed to create a subtree rooted on a node (bottom-up context). An thorough experimental analysis is provided, showing that the bi-directional approach incorporating top-down and bottom-up contexts yields to superior performances with respect to the unidirectional contexts alone, achieving state of the art results on challenging tree classification benchmarks.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IJCNN2
2014 An experimental characterization of reservoir computing in ambient assisted living applications
Davide Bacciu, Paolo Barsocchi, Stefano Chessa, Claudio Gallicchio, Alessio Micheli
Neural Comput. Appl.5
2013 An input-output hidden Markov model for tree transductions
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
Neurocomputing2
2013 Tree Echo State Networks
Claudio Gallicchio, Alessio Micheli
Neurocomputing2
2013 Novel approaches in machine learning and computational intelligence
Alessio Micheli, Frank-Michael Schleif, Peter Tiño
Neurocomputing1
2013 Compositional Generative Mapping for Tree-Structured Data - Part II: Topographic Projection Model
abstract
We introduce GTM-SD (Generative Topographic Mapping for Structured Data), which is the first compositional generative model for topographic mapping of tree-structured data. GTM-SD exploits a scalable bottom-up hidden-tree Markov model that was introduced in Part I of this paper to achieve a recursive topographic mapping of hierarchical information. The proposed model allows efficient exploitation of contextual information from shared substructures by a recursive upward propagation on the tree structure which distributes substructure information across the topographic map. Compared to its noncompositional generative counterpart, GTM-SD is shown to allow the topographic mapping of the full sample tree, which includes a projection onto the lattice of all the distinct subtrees rooted in each of its nodes. Experimental results show that the continuous projection space generated by the smooth topographic mapping of GTM-SD yields a finer grained discrimination of the sample structures with respect to the state-of-the-art recursive neural network approach.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IEEE Trans. Neural Networks Learn. Syst.2
2012 Input-Output Hidden Markov Models for trees
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ESANN2
2012 Constructive Reservoir Computation with Output Feedbacks for Structured Domains
Claudio Gallicchio, Alessio Micheli, Giulio Visco
ESANN2
2012 A Generative Multiset Kernel for Structured Data
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ICANN (1)2
2012 Compositional Generative Mapping for Tree-Structured Data - Part I: Bottom-Up Probabilistic Modeling of Trees
abstract
We introduce a novel compositional (recursive) probabilistic model for trees that defines an approximated bottom-up generative process from the leaves to the root of a tree. The proposed model defines contextual state transitions from the joint configuration of the children to the parent nodes. We argue that the bottom-up context postulates different probabilistic assumptions with respect to a top-down approach, leading to different representational capabilities. We discuss classes of applications that are best suited to a bottom-up approach. In particular, the bottom-up context is shown to better correlate and model the co-occurrence of substructures among the child subtrees of internal nodes. A mixed memory approximation is introduced to factorize the joint children-to-parent state transition matrix as a mixture of pairwise transitions. The proposed approach is the first practical bottom-up generative model for tree-structured data that maintains the same computational class of its top-down counterpart. Comparative experimental analyses exploiting synthetic and real-world datasets show that the proposed model can deal with deep structures better than a top-down generative model. The model is also shown to better capture structural information from real-world data comprising trees with a large out-degree. The proposed bottom-up model can be used as a fundamental building block for the development of other new powerful models.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IEEE Trans. Neural Networks Learn. Syst.2
2011 Exploiting vertices states in GraphESN by weighted nearest neighbor
Claudio Gallicchio, Alessio Micheli
ESANN2
2011 Adaptive tree kernel by multinomial generative topographic mapping
abstract
Learning the kernel function from data is a challenging open issue in structured data processing. In the paper, we propose a novel adaptive kernel, defined over a generative learning model, that exploits a novel multinomial extension of the Generative Topographic Mapping for Structured Data (GTM-SD). We show how the proposed kernel effectively exploits the GTM-SD continuity and smoothness properties to provide dense kernels characterized by an high discriminative power even with small topographic maps. Experimental evaluations on challenging structured XML document repositories show the effectiveness of the proposed approach against state-of-the-art syntactic and adaptive convolutional kernels.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IJCNN2
2011 Architectural and Markovian factors of echo state networks
Claudio Gallicchio, Alessio Micheli
Neural Networks2
2010 A Markovian characterization of redundancy in echo state networks by PCA
Claudio Gallicchio, Alessio Micheli
ESANN2
2010 TreeESN: a Preliminary Experimental Analysis
Claudio Gallicchio, Alessio Micheli
ESANN2
2010 Bottom-Up Generative Modeling of Tree-Structured Data
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
ICONIP (1)2
2010 Compositional generative mapping of structured data
abstract
We introduce a compositional generative model for topographic mapping of tree-structured data. It exploits a scalable bottom-up hidden tree Markov model to achieve a recursive topographic mapping of hierarchical information. The model allows for an efficient exploitation of contextual information from shared substructures by recursive upward propagation on the tree structure and by allowing it to distribute across the map. Experimental results show that the model yields to a topographically ordered mapping of the substructures in the input data.
Davide Bacciu, Alessio Micheli, Alessandro Sperduti
IJCNN2
2010 Graph Echo State Networks
abstract
In this paper we introduce the Graph Echo State Network (GraphESN) model, a generalization of the Echo State Network (ESN) approach to graph domains. GraphESNs allow for an efficient approach to Recursive Neural Networks (RecNNs) modeling extended to deal with cyclic/acyclic, directed/undirected, labeled graphs. The recurrent reservoir of the network computes a fixed contractive encoding function over graphs and is left untrained after initialization, while a feed-forward readout implements an adaptive linear output function. Contractivity of the state transition function implies a Markovian characterization of state dynamics and stability of the state computation in presence of cycles. Due to the use of fixed (untrained) encoding, the model represents both an extremely efficient version and a baseline for the performance of recursive models with trained connections. The performance are shown on standard benchmark tasks from Chemical domains, allowing the comparison with both Neural Network and Kernel-based approaches for graphs.
Claudio Gallicchio, Alessio Micheli
IJCNN2
2009 Modeling adaptive kernels from probabilistic phylogenetic trees
Luca Nicotra, Alessio Micheli
Artif. Intell. Medicine2
2009 Neural Network for Graphs: A Contextual Constructive Approach
abstract
This paper presents a new approach for learning in structured domains (SDs) using a constructive neural network for graphs (NN4G). The new model allows the extension of the input domain for supervised neural networks to a general class of graphs including both acyclic/cyclic, directed/undirected labeled graphs. In particular, the model can realize adaptive contextual transductions, learning the mapping from graphs for both classification and regression tasks. In contrast to previous neural networks for structures that had a recursive dynamics, NN4G is based on a constructive feedforward architecture with state variables that uses neurons with no feedback connections. The neurons are applied to the input graphs by a general traversal process that relaxes the constraints of previous approaches derived by the causality assumption over hierarchical input data. Moreover, the incremental approach eliminates the need to introduce cyclic dependencies in the definition of the system state variables. In the traversal process, the NN4G units exploit (local) contextual information of the graphs vertices. In spite of the simplicity of the approach, we show that, through the compositionality of the contextual information developed by the learning, the model can deal with contextual information that is incrementally extended according to the graphs topology. The effectiveness and the generality of the new approach are investigated by analyzing its theoretical properties and providing experimental results.
Alessio Micheli
IEEE Trans. Neural Networks1
2007 Recursive Principal Component Analysis of Graphs
Alessio Micheli, Alessandro Sperduti
ICANN (2)1
2005 A preliminary empirical comparison of recursive neural networks and tree kernel methods on regression tasks for tree structured domains
Alessio Micheli, Filippo Portera, Alessandro Sperduti
Neurocomputing1
2005 Universal Approximation Capability of Cascade Correlation for Structures
abstract
Cascade correlation (CC) constitutes a training method for neural networks that determines the weights as well as the neural architecture during training. Various extensions of CC to structured data have been proposed: recurrent cascade correlation (RCC) for sequences, recursive cascade correlation (RecCC) for tree structures with limited fan-out, and contextual recursive cascade correlation (CRecCC) for rooted directed positional acyclic graphs (DPAGs) with limited fan-in and fan-out. We show that these models possess the universal approximation property in the following sense: given a probability measure P on the input set, every measurable function from sequences into a real vector space can be approximated by a sigmoidal RCC up to any desired degree of accuracy up to inputs of arbitrary small probability. Every measurable function from tree structures with limited fan-out into a real vector space can be approximated by a sigmoidal RecCC with multiplicative neurons up to any desired degree of accuracy up to inputs of arbitrary small probability. For sigmoidal CRecCC networks with multiplicative neurons, we show the universal approximation capability for functions on an important subset of all DPAGs with limited fan-in and fan-out for which a specific linear representation yields unique codes. We give one sufficient structural condition for the latter property, which can easily be tested: the enumeration of ingoing and outgoing edges should becom patible. This property can be fulfilled for every DPAG with fan-in and fan-out two via reenumeration of children and parents, and for larger fan-in and fan-out via an expansion of the fan-in and fan-out and reenumeration of children and parents. In addition, the result can be generalized to the case of input-output isomorphic transductions of structures. Thus, CRecCC networks consti-tute the first neural models for which the universal approximation ca-pability of functions involving fairly general acyclic graph structures is proved.
Barbara Hammer, Alessio Micheli, Alessandro Sperduti
Neural Comput.2
2004 A preliminary experimental comparison of recursive neural networks and a tree kernel method for QSAR/QSPR regression tasks
Alessio Micheli, Filippo Portera, Alessandro Sperduti
ESANN1
2004 Fisher kernel for tree structured data
abstract
We introduce a kernel for structured data, which is an extension of the Fisher kernel used for sequences. In our approach, we extract the Fisher score vectors from a Bayesian network, specifically a hidden tree Markov model, which can be constructed starting from the training data. Experiments on a QSPR (quantitative structure-property relationship) analysis, where instances are naturally represented as trees, allow a first test of the approach.
Luca Nicotra, Alessio Micheli, Antonina Starita
IJCNN2
2004 A general framework for unsupervised processing of structured data
Barbara Hammer, Alessio Micheli, Alessandro Sperduti, Marc Strickert
Neurocomputing2
2004 Recursive self-organizing network models
Barbara Hammer, Alessio Micheli, Alessandro Sperduti, Marc Strickert
Neural Networks2
2004 Contextual processing of structured data by recursive cascade correlation
abstract
This paper propose a first approach to deal with contextual information in structured domains by recursive neural networks. The proposed model, i.e., contextual recursive cascade correlation (CRCC), a generalization of the recursive cascade correlation (RCC) model, is able to partially remove the causality assumption by exploiting contextual information stored in frozen units. We formally characterize the properties of CRCC showing that it is able to compute contextual transductions and also some causal supersource transductions that RCC cannot compute. Experimental results on controlled sequences and on a real-world task involving chemical structures confirm the computational limitations of RCC, while assessing the efficiency and efficacy of CRCC in dealing both with pure causal and contextual prediction tasks. Moreover, results obtained for the real-world task show the superiority of the proposed approach versus RCC when exploring a task for which it is not known whether the structural causality assumption holds.
Alessio Micheli, Diego Sona, Alessandro Sperduti
IEEE Trans. Neural Networks1
2003 Formal Determination of Context in Contextual Recursive Cascade Correlation Networks
Alessio Micheli, Diego Sona, Alessandro Sperduti
ICANN1
2002 A general framework for unsupervised processing of structured data
Barbara Hammer, Alessio Micheli, Alessandro Sperduti
ESANN2
2000 Bi-Causal Recurrent Cascade Correlation
abstract
Recurrent neural networks fail to deal with prediction tasks which do not satisfy the causality assumption. We propose to exploit bi-causality to extend the recurrent cascade correlation model in order to deal with contextual prediction tasks. Preliminary results on artificial data show the ability of the model to preserve the prediction capability of recurrent cascade correlation on strict causal tasks, while extending this capability also to prediction tasks involving the future.
Alessio Micheli, Diego Sona, Alessandro Sperduti
IJCNN (3)1
2000 Building MLP Networks by Construction
abstract
We introduce two new models which are obtained through the modification of the well known methods MLP and cascade correlation. These two methods differ fundamentally as they employ learning techniques and produce network architectures that are not directly comparable. We extended the MLP architecture, and reduced the constructive method to obtain very comparable network architectures. The greatest benefit of these new models is that we can obtain an MLP-structured network through a constructive method based on the cascade correlation algorithm, and that we can train a cascade correlation structured network using the standard MLP learning technique. Additionally, we show that cascade correlation is a universal approximator, a fact that has not yet been discussed in literature.
Ah Chung Tsoi, Markus Hagenbuchner, Alessio Micheli
IJCNN (4)3
2000 Application of Cascade Correlation Networks for Structures to Chemistry
Anna Maria Bianucci, Alessio Micheli, Alessandro Sperduti, Antonina Starita
Appl. Intell.2