Cesare Alippi

dblp:84/6337 · DBLP profile ↗
← Back
133ranked-venue papers
61as first author
42since 2021 · last 2026
0000-0003-3819-0025ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 97 · 33 first-author · 39 since 2021Systems, architecture and hardware · 20 · 15 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Computer networks · 6 · 6 first-authorHuman-computer interaction and ubiquitous computing · 6 · 5 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Compensating Distribution Drifts in Continual Learning with Pre-trained Vision Transformers
abstract
Recent advances have shown that sequential fine-tuning (SeqFT) of pre-trained vision transformers (ViTs), followed by classifier refinement using approximate distributions of class features, can be an effective strategy for class-incremental learning (CIL). However, this approach is susceptible to distribution drift, caused by the sequential optimization of shared backbone parameters. This results in a mismatch between the distributions of the previously learned classes and that of the updated model, ultimately degrading the effectiveness of classifier performance over time. To address this issue, we introduce a latent space transition operator and propose Sequential Learning with Drift Compensation (SLDC). SLDC aims to align feature distributions across tasks to mitigate the impact of drift. First, we present a linear variant of SLDC, which learns a linear operator by solving a regularized least-squares problem that maps features before and after fine-tuning. Next, we extend this with a weakly nonlinear SLDC variant, which assumes that the ideal transition operator lies between purely linear and fully nonlinear transformations. This is implemented using learnable, weakly nonlinear mappings that balance flexibility and generalization. To further reduce representation drift, we apply knowledge distillation (KD) in both algorithmic variants. Extensive experiments on standard CIL benchmarks demonstrate that SLDC significantly improves the performance of SeqFT. Notably, by combining KD to address representation drift with SLDC to compensate distribution drift, SeqFT achieves performance comparable to joint training across all evaluated datasets.
Xuan Rao, Simian Xu, Bo Zhao 0015, Derong Liu 0001, Mingming Ha, Cesare Alippi
AAAI7
2026 Assessment of spatio-temporal predictors in the presence of missing and heterogeneous data
abstract
Deep learning methods achieve remarkable predictive performance in modeling complex, large-scale data. However, assessing the quality of derived models has become increasingly challenging, as more classical statistical assumptions may no longer apply. These difficulties are particularly pronounced for spatio-temporal data, which exhibit dependencies across both space and time and are often characterized by nonlinear dynamics, time variance, and missing observations, hence calling for new accuracy assessment methodologies. This paper introduces a residual correlation analysis framework for assessing the optimality of spatio-temporal relational-enabled neural predictive models, notably in settings with incomplete and heterogeneous data. By leveraging the principle that residual correlation indicates information not captured by the model, enabling the identification and localization of regions in space and time where predictive performance can be improved. A strength of the proposed approach is that it operates under minimal assumptions, allowing also for robust evaluation of deep learning models applied to multivariate time series, even in the presence of missing and heterogeneous data. In detail, the methodology constructs tailored spatio-temporal graphs to encode sparse spatial and temporal dependencies and employs asymptotically distribution-free summary statistics to detect time intervals and spatial regions where the model underperforms. The effectiveness of what proposed is demonstrated through experiments on both synthetic and real-world datasets using state-of-the-art predictive models. • Proposes a novel residual correlation analysis to assess the quality of deep spatio-temporal models and identify regions where the predictions can be improved. • Complements traditional accuracy-based evaluations by offering an independent, metric-agnostic assessment of model quality. • Operates under minimal assumptions and remains effective even with missing or heterogeneous data, making it broadly applicable to real-world scenarios and deep learning models.
Daniele Zambon, Cesare Alippi
Neurocomputing2
2026 Preference isolation forest for structure-based anomaly detection
Filippo Leveni, Luca Magri 0002, Cesare Alippi, Giacomo Boracchi
Pattern Recognit.3
2025 Relational Conformal Prediction for Correlated Time Series
abstract
We address the problem of uncertainty quantification in time series forecasting by exploiting observations at correlated sequences. Relational deep learning methods leveraging graph representations are among the most effective tools for obtaining point estimates from spatiotemporal data and correlated time series. However, the problem of exploiting relational structures to estimate the uncertainty of such predictions has been largely overlooked in the same context. To this end, we propose a novel distribution-free approach based on the conformal prediction framework and quantile regression. Despite the recent applications of conformal prediction to sequential data, existing methods operate independently on each target time series and do not account for relationships among them when constructing the prediction interval. We fill this void by introducing a novel conformal prediction method based on graph deep learning operators. Our approach, named Conformal Relational Prediction (CoRel), does not require the relational structure (graph) to be known a priori and can be applied on top of any pre-trained predictor. Additionally, CoRel includes an adaptive component to handle non-exchangeable data and changes in the input time series. Our approach provides accurate coverage and achieves state-of-the-art uncertainty quantification in relevant benchmarks.
Andrea Cini, Alexander Jenkins, Danilo P. Mandic, Cesare Alippi, Filippo Maria Bianchi
ICML4
2025 Learning Latent Graph Structures and their Uncertainty
abstract
Graph neural networks use relational information as an inductive bias to enhance prediction performance. Not rarely, task-relevant relations are unknown and graph structure learning approaches have been proposed to learn them from data. Given their latent nature, no graph observations are available to provide a direct training signal to the learnable relations. Therefore, graph topologies are typically learned on the prediction task alongside the other graph neural network parameters. In this paper, we demonstrate that minimizing point-prediction losses does not guarantee proper learning of the latent relational information and its associated uncertainty. Conversely, we prove that suitable loss functions on the stochastic model outputs simultaneously grant solving two tasks: (i) learning the unknown distribution of the latent graph and (ii) achieving optimal predictions of the target variable. Finally, we propose a sampling-based method that solves this joint learning task. Empirical results validate our theoretical claims and demonstrate the effectiveness of the proposed approach.
Alessandro Manenti, Daniele Zambon, Cesare Alippi
ICML3
2025 Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games
abstract
Equilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately solved. When the underlying graph structure varies, even the state-of-the-art RL methods require recomputation or at least fine-tuning, which can be time-consuming and impair real-time applicability. This paper proposes an Equilibrium Policy Generalization (EPG) framework to effectively learn a generalized policy with robust cross-graph zero-shot performance. In the context of PEGs, our framework is generally applicable to both pursuer and evader sides in both no-exit and multi-exit scenarios. These two generalizability properties, to our knowledge, are the first to appear in this domain. The core idea of the EPG framework is to train an RL policy across different graph structures against the equilibrium policy for each single graph. To construct an equilibrium oracle for single-graph policies, we present a dynamic programming (DP) algorithm that provably generates pure-strategy Nash equilibrium with near-optimal time complexity. To guarantee scalability with respect to pursuer number, we further extend DP and RL by designing a grouping mechanism and a sequence model for joint policy decomposition, respectively. Experimental results show that, using equilibrium guidance and a distance feature proposed for cross-graph PEG training, the EPG framework guarantees desirable zero-shot performance in various unseen real-world graphs. Besides, when trained under an equilibrium heuristic proposed for the graphs with exits, our generalized pursuer policy can even match the performance of the fine-tuned policies from the state-of-the-art PEG methods.
Runyu Lu, Peng Zhang 0127, Ruochuan Shi, Yuanheng Zhu, Dongbin Zhao, Yang Liu 0066, Dong Wang 0004, Cesare Alippi
NeurIPS8
2025 Over-squashing in Spatiotemporal Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have achieved remarkable success across various domains. However, recent theoretical advances have identified fundamental limitations in their information propagation capabilities, such as over-squashing, where distant nodes fail to effectively exchange information. While extensively studied in static contexts, this issue remains unexplored in Spatiotemporal GNNs (STGNNs), which process sequences associated with graph nodes. Nonetheless, the temporal dimension amplifies this challenge by increasing the information that must be propagated. In this work, we formalize the spatiotemporal over-squashing problem and demonstrate its distinct characteristics compared to the static case. Our analysis reveals that, counterintuitively, convolutional STGNNs favor information propagation from points temporally distant rather than close in time. Moreover, we prove that architectures that follow either time-and-space or time-then-space processing paradigms are equally affected by this phenomenon, providing theoretical justification for computationally efficient implementations. We validate our findings on synthetic and real-world datasets, providing deeper insights into their operational dynamics and principled guidance for more effective designs.
Ivan Marisca, Jacob Bamberger, Cesare Alippi, Michael M. Bronstein
NeurIPS3
2025 FX-DARTS: Designing Topology-Unconstrained Architectures With Differentiable Architecture Search and Entropy-BasedSuper-Network Shrinking
abstract
Strong priors are imposed on the search space of differentiable architecture search (DARTS), such that cells of the same type share the same topological structure and each intermediate node retains two operators from distinct nodes. While these priors reduce optimization difficulties and improve the applicability of searched architectures, they hinder the subsequent development of automated machine learning (auto-ML) and prevent the optimization algorithm from exploring more powerful neural networks through improved architectural flexibility. This article aims to reduce these prior constraints by eliminating restrictions on cell topology and modifying the discretization mechanism for super-networks. Specifically, the flexible DARTS (FX-DARTS) method, which leverages an entropy-based super-network shrinking (ESS) framework, is presented to address the challenges arising from the elimination of prior constraints. Notably, FX-DARTS enables the derivation of neural architectures without strict prior rules while maintaining the stability in the enlarged search space. Experimental results on image classification benchmarks demonstrate that FX-DARTS is capable of exploring a set of neural architectures with competitive trade-offs between performance and computational complexity within a single search procedure.
Xuan Rao, Bo Zhao 0015, Derong Liu 0001, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.4
2024 Graph-based Virtual Sensing from Sparse and Partial Multivariate Observations
abstract
Virtual sensing techniques allow for inferring signals at new unmonitored locations by exploiting spatio-temporal measurements coming from physical sensors at different locations. However, as the sensor coverage becomes sparse due to costs or other constraints, physical proximity cannot be used to support interpolation. In this paper, we overcome this challenge by leveraging dependencies between the target variable and a set of correlated variables (covariates) that can frequently be associated with each location of interest. From this viewpoint, covariates provide partial observability, and the problem consists of inferring values for unobserved channels by exploiting observations at other locations to learn how such variables can correlate. We introduce a novel graph-based methodology to exploit such relationships and design a graph deep learning architecture, named GgNet, implementing the framework. The proposed approach relies on propagating information over a nested graph structure that is used to learn dependencies between variables as well as locations. GgNet is extensively evaluated under different virtual sensing scenarios, demonstrating higher reconstruction accuracy compared to the state-of-the-art.
Giovanni de Felice, Andrea Cini, Daniele Zambon, Vladimir V. Gusev, Cesare Alippi
ICLR5
2024 Graph-based Time Series Clustering for End-to-End Hierarchical Forecasting
abstract
Relationships among time series can be exploited as inductive biases in learning effective forecasting models. In hierarchical time series, relationships among subsets of sequences induce hard constraints (hierarchical inductive biases) on the predicted values. In this paper, we propose a graph-based methodology to unify relational and hierarchical inductive biases in the context of deep learning for time series forecasting. In particular, we model both types of relationships as dependencies in a pyramidal graph structure, with each pyramidal layer corresponding to a level of the hierarchy. By exploiting modern - trainable - graph pooling operators we show that the hierarchical structure, if not available as a prior, can be learned directly from data, thus obtaining cluster assignments aligned with the forecasting objective. A differentiable reconciliation stage is incorporated into the processing architecture, allowing hierarchical constraints to act both as an architectural bias as well as a regularization element for predictions. Simulation results on representative datasets show that the proposed method compares favorably against the state of the art.
Andrea Cini, Danilo P. Mandic, Cesare Alippi
ICML3
2024 Graph-based Forecasting with Missing Data through Spatiotemporal Downsampling
abstract
Given a set of synchronous time series, each associated with a sensor-point in space and characterized by inter-series relationships, the problem of spatiotemporal forecasting consists of predicting future observations for each point. Spatiotemporal graph neural networks achieve striking results by representing the relationships across time series as a graph. Nonetheless, most existing methods rely on the often unrealistic assumption that inputs are always available and fail to capture hidden spatiotemporal dynamics when part of the data is missing. In this work, we tackle this problem through hierarchical spatiotemporal downsampling. The input time series are progressively coarsened over time and space, obtaining a pool of representations that capture heterogeneous temporal and spatial dynamics. Conditioned on observations and missing data patterns, such representations are combined by an interpretable attention mechanism to generate the forecasts. Our approach outperforms state-of-the-art methods on synthetic and real-world benchmarks under different missing data distributions, particularly in the presence of contiguous blocks of missing values.
Ivan Marisca, Cesare Alippi, Filippo Maria Bianchi
ICML2
2024 Temporal Graph ODEs for Irregularly-Sampled Time Series
Alessio Gravina, Daniele Zambon, Davide Bacciu, Cesare Alippi
IJCAI4
2024 Physics-Informed Graph Neural Cellular Automata: an Application to Compartmental Modelling
abstract
The recent outbreak of COVID-19 has spurred global collaborative research efforts to model and forecast the disease to improve preparation and control. Epidemiological models integrate experimental data and expert opinions to understand infection dynamics and control measures. Classical Machine Learning techniques often face challenges such as high data requirements, lack of interpretability, and difficulty integrating domain knowledge. A potential solution is to leverage Physically-Informed Machine Learning (PIML) models, which enhance models by incorporating known physical properties of viral spread. Additionally, epidemiological datasets are best represented as graphs, facilitating the modelling of interactions between individuals. In this paper, we propose a novel, interpretable graph-based PIML technique called SINDy-Graph to model infectious disease dynamics. Our approach is a Graph Cellular Automata architecture that combines the ability to identify dynamics for discovering the differential equations governing the physical phenomena under study using graphs modelling relationships between nodes (individuals). The experimental results demonstrate that integrating domain knowledge ensures better physical plausibility. In addition, our proposed model is easier to train and achieves a lower generalisation error compared to other baseline methods.
Nicolò Navarin, Paolo Frazzetto, Luca Pasa, Pietro Verzelli, Filippo Visentin, Alessandro Sperduti, Cesare Alippi
IJCNN7
2024 A Survey on Graph Neural Networks for Time Series: Forecasting, Classification, Imputation, and Anomaly Detection
abstract
Time series are the primary data type used to record dynamic system measurements and generated in great volume by both physical sensors and online processes (virtual sensors). Time series analytics is therefore crucial to unlocking the wealth of information implicit in available data. With the recent advancements in graph neural networks (GNNs), there has been a surge in GNN-based approaches for time series analysis. These approaches can explicitly model inter-temporal and inter-variable relationships, which traditional and other deep neural network-based methods struggle to do. In this survey, we provide a comprehensive review of graph neural networks for time series analysis (GNN4TS), encompassing four fundamental dimensions: forecasting, classification, anomaly detection, and imputation. Our aim is to guide designers and practitioners to understand, build applications, and advance research of GNN4TS. At first, we provide a comprehensive task-oriented taxonomy of GNN4TS. Then, we present and discuss representative research works and introduce mainstream applications of GNN4TS. A comprehensive discussion of potential future research directions completes the survey. This survey, for the first time, brings together a vast array of knowledge on GNN-based time series research, highlighting foundations, practical applications, and opportunities of graph neural networks for time series analysis.
Ming Jin 0005, Huan Yee Koh, Qingsong Wen, Daniele Zambon, Cesare Alippi, Geoffrey I. Webb, Irwin King, Shirui Pan
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Explainable Intelligent Fault Diagnosis for Nonlinear Dynamic Systems: From Unsupervised to Supervised Learning
abstract
The increased complexity and intelligence of automation systems require the development of intelligent fault diagnosis (IFD) methodologies. By relying on the concept of a suspected space, this study develops explainable data-driven IFD approaches for nonlinear dynamic systems. More specifically, we parameterize nonlinear systems through a generalized kernel representation for system modeling and the associated fault diagnosis. An important result obtained is a unified form of kernel representations, applicable to both unsupervised and supervised learning. More importantly, through a rigorous theoretical analysis, we discover the existence of a bridge (i.e., a bijective mapping) between some supervised and unsupervised learning-based entities. Notably, the designed IFD approaches achieve the same performance with the use of this bridge. In order to have a better understanding of the results obtained, both unsupervised and supervised neural networks are chosen as the learning tools to identify the generalized kernel representations and design the IFD schemes; an invertible neural network is then employed to build the bridge between them. This article is a perspective article, whose contribution lies in proposing and formalizing the fundamental concepts for explainable intelligent learning methods, contributing to system modeling and data-driven IFD designs for nonlinear dynamic systems.
Hongtian Chen, Zhigang Liu 0001, Cesare Alippi, Biao Huang 0001, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Understanding Pooling in Graph Neural Networks
abstract
Many recent works in the field of graph machine learning have introduced pooling operators to reduce the size of graphs. In this article, we present an operational framework to unify this vast and diverse literature by describing pooling operators as the combination of three functions: selection, reduction, and connection (SRC). We then introduce a taxonomy of pooling operators, based on some of their key characteristics and implementation differences under the SRC framework. Finally, we propose three criteria to evaluate the performance of pooling operators and use them to investigate the behavior of different operators on a variety of tasks.
Daniele Grattarola, Daniele Zambon, Filippo Maria Bianchi, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.4
2024 Guest Editorial: Special Issue on Explainable Representation Learning-Based Intelligent Inspection and Maintenance of Complex Systems
abstract
Over the past decade, representation learning has received particular attention in the intelligent inspection and maintenance of complex systems thanks to its overwhelming advantages in discovering and mining hidden knowledge representations. The room for in-depth investigations of representation learning-related topics remains open, especially explainable approaches for intelligent inspection and maintenance of complex systems. The primary objective of this special issue, entitled “Explainable Representation Learning-based Intelligent Inspection and Maintenance of Complex Systems,” of IEEE Transactions on Neural Networks and Learning Systems is to provide the related latest achievements made by researchers and practitioners on the one hand and to identify critical issues and challenges for future investigation on the other hand.
Zhigang Liu 0001, Cesare Alippi, Hongtian Chen, Derong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Scalable Spatiotemporal Graph Neural Networks
abstract
Neural forecasting of spatiotemporal time series drives both research and industrial innovation in several relevant application domains. Graph neural networks (GNNs) are often the core component of the forecasting architecture. However, in most spatiotemporal GNNs, the computational complexity scales up to a quadratic factor with the length of the sequence times the number of links in the graph, hence hindering the application of these models to large graphs and long temporal sequences. While methods to improve scalability have been proposed in the context of static graphs, few research efforts have been devoted to the spatiotemporal case. To fill this gap, we propose a scalable architecture that exploits an efficient encoding of both temporal and spatial dynamics. In particular, we use a randomized recurrent neural network to embed the history of the input time series into high-dimensional state representations encompassing multi-scale temporal dynamics. Such representations are then propagated along the spatial dimension using different powers of the graph adjacency matrix to generate node embeddings characterized by a rich pool of spatiotemporal features. The resulting node embeddings can be efficiently pre-computed in an unsupervised manner, before being fed to a feed-forward decoder that learns to map the multi-scale spatiotemporal representations to predictions. The training procedure can then be parallelized node-wise by sampling the node embeddings without breaking any dependency, thus enabling scalability to large networks. Empirical results on relevant datasets show that our approach achieves results competitive with the state of the art, while dramatically reducing the computational burden.
Andrea Cini, Ivan Marisca, Filippo Maria Bianchi, Cesare Alippi
AAAI4
2023 Anomaly Detection in Optical Spectra VIA Joint Optimization
abstract
Despite the remarkable progress of fiber optics in communication, little attention has been devoted to the automatic detection of anomalies in optical spectra, i.e., poorly transmitted channels. This task is typically addressed by ad-hoc heuristics that fall short in spectra presenting heavy distortions caused by optical amplifiers during transmission. We propose a method based on a joint optimization procedure for estimating the major trends that characterize the spectrum, enabling the detection of anomalies even in the presence of few channels and heavy distortions. Our experiments have shown that the proposed method can successfully localize anomalies achieving more than 98% accuracy, outperforming all competitors.
Antonino Maria Rizzo, Luca Magri 0001, Pietro Invernizzi, Enrico Sozio, Stefano Piciaccia, Alberto Tanzi, Stefano Binetti, Cesare Alippi, Giacomo Boracchi
ICASSP8
2023 Taming Local Effects in Graph-based Spatiotemporal Forecasting
abstract
Spatiotemporal graph neural networks have shown to be effective in time series forecasting applications, achieving better performance than standard univariate predictors in several settings. These architectures take advantage of a graph structure and relational inductive biases to learn a single (global) inductive model to predict any number of the input time series, each associated with a graph node. Despite the gain achieved in computational and data efficiency w.r.t. fitting a set of local models, relying on a single global model can be a limitation whenever some of the time series are generated by a different spatiotemporal stochastic process. The main objective of this paper is to understand the interplay between globality and locality in graph-based spatiotemporal forecasting, while contextually proposing a methodological framework to rationalize the practice of including trainable node embeddings in such architectures. We ascribe to trainable node embeddings the role of amortizing the learning of specialized components. Moreover, embeddings allow for 1) effectively combining the advantages of shared message-passing layers with node-specific parameters and 2) efficiently transferring the learned model to new node sets. Supported by strong empirical evidence, we provide insights and guidelines for specializing graph-based models to the dynamics of each time series and show how this aspect plays a crucial role in obtaining accurate predictions.
Andrea Cini, Ivan Marisca, Daniele Zambon, Cesare Alippi
NeurIPS4
2023 Sparse Graph Learning from Spatiotemporal Time Series
abstract
Outstanding achievements of graph neural networks for spatiotemporal time series analysis show that relational constraints introduce an effective inductive bias into neural forecasting architectures. Often, however, the relational information characterizing the underlying data-generating process is unavailable and the practitioner is left with the problem of inferring from data which relational graph to use in the subsequent processing stages. We propose novel, principled - yet practical - probabilistic score-based methods that learn the relational dependencies as distributions over graphs while maximizing end-to-end the performance at task. The proposed graph learning framework is based on consolidated variance reduction techniques for Monte Carlo score-based gradient estimation, is theoretically grounded, and, as we show, effective in practice. In this paper, we focus on the time series forecasting problem and show that, by tailoring the gradient estimators to the graph learning problem, we are able to achieve state-of-the-art performance while controlling the sparsity of the learned graph and the computational scalability. We empirically assess the effectiveness of the proposed method on synthetic and real-world benchmarks, showing that the proposed solution can be used as a stand-alone graph identification procedure as well as a graph learning component of an end-to-end forecasting architecture.
Andrea Cini, Daniele Zambon, Cesare Alippi
J. Mach. Learn. Res.3
2023 Graph Neural Networks for High-Level Synthesis Design Space Exploration
abstract
High-level Synthesis (HLS) Design-Space Exploration (DSE) aims at identifying Pareto-optimal synthesis configurations whose exhaustive search is unfeasible due to the design-space dimensionality and the prohibitive computational cost of the synthesis process. Within this framework, we address the design automation problem by proposing graph neural networks that jointly predict acceleration performance and hardware costs of a synthesized behavioral specification given optimization directives. Learned models can be used to rapidly approach the Pareto curve by guiding the DSE, taking into account performance and cost estimates. The proposed method outperforms traditional HLS-driven DSE approaches, by accounting for the arbitrary length of computer programs and the invariant properties of the input. We propose a novel hybrid control and dataflow graph representation that enables training the graph neural network on specifications of different hardware accelerators. Our approach achieves prediction accuracy comparable with that of state-of-the-art simulators without having access to analytical models of the HLS compiler. Finally, the learned representation can be exploited for DSE in unexplored configuration spaces by fine-tuning on a small number of samples from the new target domain. The outcome of the empirical evaluation of this transfer learning shows strong results against state-of-the-art baselines in relevant benchmarks.
Lorenzo Ferretti, Andrea Cini, Georgios Zacharopoulos 0001, Cesare Alippi, Laura Pozzi 0001
ACM Trans. Design Autom. Electr. Syst.4
2022 Filling the G_ap_s: Multivariate Time Series Imputation by Graph Neural Networks
Andrea Cini, Ivan Marisca, Cesare Alippi
ICLR3
2022 Spatio-Temporal Graph Neural Networks for Aggregate Load Forecasting
abstract
Accurate forecasting of electricity demand is a core component of the modern electricity infrastructure. Several approaches exist that tackle this problem by exploiting modern deep learning tools. However, most previous works focus on predicting the total load as a univariate time series forecasting task, ignoring all fine-grained information captured by the smart meters distributed across the power grid. We introduce a methodology to account for this information in the graph neural network framework. In particular, we consider spatio-temporal graphs where each node is associated with the aggregate load of a cluster of smart meters, and a global graph-level attribute indicates the total load on the grid. We propose two novel spatio-temporal graph neural network models to process this representation and take advantage of both the finer-grained information and the relationships existing between the different clusters of meters. We compare these models on a widely used, openly available, benchmark against a competitive baseline which only accounts for the total load profile. Within these settings, we show that the proposed methodology improves forecasting accuracy.
Simone Eandi, Andrea Cini, Slobodan Lukovic, Cesare Alippi
IJCNN4
2022 Understanding Catastrophic Forgetting of Gated Linear Networks in Continual Learning
abstract
In this paper, we consider the recently proposed family of continual learning models, called Gated Linear Networks (GLNs), and study two crucial aspects impacting on the amount of catastrophic forgetting affecting gated linear networks, namely, data standardization and gating mechanism. Data standardization is particularly challenging in the online/continual learning setting because data from future tasks is not available beforehand. The results obtained using an online standardization method show a considerably higher amount of forgetting compared to an offline -static- standardization. Interestingly, with the latter standardization, we observe that GLNs show almost no forgetting on the considered benchmark datasets. Secondly, for an effective GLNs, it is essential to tailor the hyperparameters of the gating mechanism to the data distribution. In this paper, we propose a gating strategy based on a set of prototypes and the resulting Voronoi tessellation. The experimental assessment shows that the proposed approach is more robust to different data standardizations compared to the original one, based on a halfspace gating mechanism, and shows improved predictive performance.
Matteo Munari, Luca Pasa, Daniele Zambon, Cesare Alippi, Nicolò Navarin
IJCNN4
2022 Graph iForest: Isolation of anomalous and outlier graphs
abstract
We present an anomaly and outlier detection method for graph data. The method relies on the consideration that anomalies and outliers are more easily isolated by certain incremental partitionings of the data space. Specifically, we build upon the isolation forest method and introduce a new incremental partitioning of the space of graphs that makes the isolation forest method applicable to generic attributed graphs, i.e., graphs where both nodes and edges can be associated with attributes. Within the considered general setup, the topology and the number of nodes can change from graph to graph, and a node correspondence between different graphs can be absent or unknown. Examples of applications of what proposed include the identification of frauds and fake news in communication networks, and breakage of systems monitored by sensor networks. The main novel contribution of the paper is a graph space partitioning which we prove to be expressive enough to identify anomalies and outlier graphs in a given dataset. An empirical analysis on synthetic and real-world graphs validates the effectiveness of the proposed method.
Daniele Zambon, Lorenzo Livi, Cesare Alippi
IJCNN3
2022 Learning to Reconstruct Missing Data from Spatiotemporal Graphs with Sparse Observations
abstract
Modeling multivariate time series as temporal signals over a (possibly dynamic) graph is an effective representational framework that allows for developing models for time series analysis. In fact, discrete sequences of graphs can be processed by autoregressive graph neural networks to recursively learn representations at each discrete point in time and space. Spatiotemporal graphs are often highly sparse, with time series characterized by multiple, concurrent, and long sequences of missing data, e.g., due to the unreliable underlying sensor network. In this context, autoregressive models can be brittle and exhibit unstable learning dynamics. The objective of this paper is, then, to tackle the problem of learning effective models to reconstruct, i.e., impute, missing data points by conditioning the reconstruction only on the available observations. In particular, we propose a novel class of attention-based architectures that, given a set of highly sparse discrete observations, learn a representation for points in time and space by exploiting a spatiotemporal propagation architecture aligned with the imputation task. Representations are trained end-to-end to reconstruct observations w.r.t. the corresponding sensor and its neighboring nodes. Compared to the state of the art, our model handles sparse data without propagating prediction errors or requiring a bidirectional model to encode forward and backward time dependencies. Empirical results on representative benchmarks show the effectiveness of the proposed method.
Ivan Marisca, Andrea Cini, Cesare Alippi
NeurIPS3
2022 AZ-whiteness test: a test for signal uncorrelation on spatio-temporal graphs
abstract
We present the first whiteness hypothesis test for graphs, i.e., a whiteness test for multivariate time series associated with the nodes of a dynamic graph; as such, the test represents an important model assessment tool for graph deep learning, e.g., in forecasting setups. The statistical test aims at detecting existing serial dependencies among close-in-time observations, as well as spatial dependencies among neighboring observations given the underlying graph. The proposed AZ-test can be intended as a spatio-temporal extension of traditional tests designed for system identification to graph signals. The AZ-test is versatile, allowing the underlying graph to be dynamic, changing in topology and set of nodes over time, and weighted, thus accounting for connections of different strength, as it is the case in many application scenarios like sensor and transportation networks. The asymptotic distribution of the designed test can be derived under the null hypothesis without assuming identically distributed data. We show the effectiveness of the test on both synthetic and real-world problems, and illustrate how it can be employed to assess the quality of spatio-temporal forecasting models by analyzing the prediction residuals appended to the graph stream.
Daniele Zambon, Cesare Alippi
NeurIPS2
2022 A transfer-learning approach for corrosion prediction in pipeline infrastructures
abstract
Abstract Pipeline infrastructures, carrying either gas or oil, are often affected by internal corrosion, which is a dangerous phenomenon that may cause threats to both the environment (due to potential leakages) and the human beings (due to accidents that may cause explosions in presence of gas leakages). For this reason, predictive mechanisms are needed to detect and address the corrosion phenomenon. Recently, we have seen a first attempt at leveraging Machine Learning (ML) techniques in this field thanks to their high ability in modeling highly complex phenomena. In order to rely on these techniques, we need a set of data, representing factors influencing the corrosion in a given pipeline, together with their related supervised information, measuring the corrosion level along the considered infrastructure profile. Unfortunately, it is not always possible to access supervised information for a given pipeline since measuring the corrosion is a costly and time-consuming operation. In this paper, we will address the problem of devising a ML-based predictive model for internal corrosion under the assumption that supervised information is unavailable for the pipeline of interest, while it is available for some other pipelines that can be leveraged through Transfer Learning (TL) to build the predictive model itself. We will cover all the methodological steps from data set creation to the usage of TL. The whole methodology will be experimentally validated on a set of real-world pipelines.
Giuseppe Canonaco, Manuel Roveri, Cesare Alippi, Fabrizio Podenzani, Antonio Bennardo, Marco Conti, Nicola Mancini
Appl. Intell.3
2022 Seizure localisation with attention-based graph neural networks
abstract
In this paper, we introduce a machine learning methodology for localising the seizure onset zone in subjects with epilepsy. We represent brain states as functional networks obtained from intracranial electroencephalography recordings, using correlation and the phase-locking value to quantify the coupling between different brain areas. Our method is based on graph neural networks (GNNs) and the attention mechanism, two of the most significant advances in artificial intelligence in recent years. Specifically, we train a GNN to distinguish between functional networks associated with interictal and ictal phases. The GNN is equipped with an attention-based layer that automatically learns to identify those regions of the brain (associated with individual electrodes) that are most important for a correct classification. The localisation of these regions does not require any prior information regarding the seizure onset zone. We show that the regions of interest identified by the GNN strongly correlate with the localisation of the seizure onset zone reported by electroencephalographers. We report results both for human patients and for simulators of brain activity. We also show that our GNN exhibits uncertainty for those patients for which the clinical localisation was unsuccessful, highlighting the robustness of the proposed approach.
Daniele Grattarola, Lorenzo Livi, Cesare Alippi, Richard Wennberg, Taufik A. Valiante
Expert Syst. Appl.3
2022 Known and unknown event detection in OTDR traces by deep learning networks
abstract
Abstract Optical fiber links are customarily monitored by Optical Time Domain Reflectometer (OTDR), an optoelectronic instrument that measures the scattered or reflected light along the fiber and returns a signal, namely the OTDR trace . OTDR traces are typically analyzed by experts in laboratories or by hand-crafted algorithms running in embedded systems to localize critical events occurring along the fiber. In this work, we address the problem of automatically detecting optical events in OTDR traces through a deep learning model that can be deployed in embedded systems. In particular, we take inspiration from Faster R-CNN and present the first 1D object-detection neural network for OTDR traces. Thanks to an ad-hoc preprocessing pipeline for OTDR traces, we can also identify unknown events , namely events that are not represented in training data but that might indicate rare and unforeseen situations that need to be reported. The resulting network brings several advantages with respect to existing solutions, as these typically classify fixed-size windows of OTDR traces, thus are less accurate in the localization. Moreover, existing solutions do not report events that cannot be safely associated to any label in the training set. Our experiments, performed on real OTDR traces, show very promising performance, and can be directly executed on embedded OTDR devices.
Antonino Maria Rizzo, Davide Rutigliano, Pietro Invernizzi, Enrico Sozio, Cesare Alippi, Stefano Binetti, Giacomo Boracchi
Neural Comput. Appl.6
2022 Graph Neural Networks With Convolutional ARMA Filters
abstract
Popular graph neural networks implement convolution operations on graphs based on polynomial spectral filters. In this paper, we propose a novel graph convolutional layer inspired by the auto-regressive moving average (ARMA) filter that, compared to polynomial ones, provides a more flexible frequency response, is more robust to noise, and better captures the global graph structure. We propose a graph neural network implementation of the ARMA filter with a recursive and distributed formulation, obtaining a convolutional layer that is efficient to train, localized in the node space, and can be transferred to new graphs at test time. We perform a spectral analysis to study the filtering effect of the proposed ARMA layer and report experiments on four downstream tasks: semi-supervised node classification, graph signal classification, graph classification, and graph regression. Results show that the proposed ARMA layer brings significant improvements over graph neural networks based on polynomial filters.
Filippo Maria Bianchi, Daniele Grattarola, Lorenzo Livi, Cesare Alippi
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Hierarchical Representation Learning in Graph Neural Networks With Node Decimation Pooling
abstract
In graph neural networks (GNNs), pooling operators compute local summaries of input graphs to capture their global properties, and they are fundamental for building deep GNNs that learn hierarchical representations. In this work, we propose the Node Decimation Pooling (NDP), a pooling operator for GNNs that generates coarser graphs while preserving the overall graph topology. During training, the GNN learns new node representations and fits them to a pyramid of coarsened graphs, which is computed offline in a preprocessing stage. NDP consists of three steps. First, a node decimation procedure selects the nodes belonging to one side of the partition identified by a spectral algorithm that approximates the MAXCUT solution. Afterward, the selected nodes are connected with Kron reduction to form the coarsened graph. Finally, since the resulting graph is very dense, we apply a sparsification procedure that prunes the adjacency matrix of the coarsened graph to reduce the computational cost in the GNN. Notably, we show that it is possible to remove many edges without significantly altering the graph structure. Experimental results show that NDP is more efficient compared to state-of-the-art graph pooling operators while reaching, at the same time, competitive performance on a significant variety of graph classification tasks.
Filippo Maria Bianchi, Daniele Grattarola, Lorenzo Livi, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.4
2022 Input-to-State Representation in Linear Reservoirs Dynamics
abstract
Reservoir computing is a popular approach to design recurrent neural networks, due to its training simplicity and approximation performance. The recurrent part of these networks is not trained (e.g., via gradient descent), making them appealing for analytical studies by a large community of researchers with backgrounds spanning from dynamical systems to neuroscience. However, even in the simple linear case, the working principle of these networks is not fully understood and their design is usually driven by heuristics. A novel analysis of the dynamics of such networks is proposed, which allows the investigator to express the state evolution using the controllability matrix. Such a matrix encodes salient characteristics of the network dynamics; in particular, its rank represents an input-independent measure of the memory capacity of the network. Using the proposed approach, it is possible to compare different reservoir architectures and explain why a cyclic topology achieves favorable results as verified by practitioners.
Pietro Verzelli, Cesare Alippi, Lorenzo Livi, Peter Tiño
IEEE Trans. Neural Networks Learn. Syst.2
2021 Event-Detection Deep Neural Network for OTDR Trace Analysis
Davide Rutigliano, Giacomo Boracchi, Pietro Invernizzi, Enrico Sozio, Cesare Alippi, Stefano Binetti
EANN5
2021 Deep learning for graphs
abstract
Deep learning for graphs encompasses all those neural models endowed with multiple layers of computation operating on data represented as graphs.The most common building blocks of these models are graph encoding layers, which compute a vector embedding for each node in a graph using message-passing operators.In this paper, we provide an overview of the key concepts in the field, point towards open questions, and frame the contributions of the ESANN 2021 special session into the broader context of deep learning for graphs.
Davide Bacciu, Filippo Maria Bianchi, Benjamin Paaßen, Cesare Alippi
ESANN4
2021 Graph Edit Networks
Benjamin Paaßen, Daniele Grattarola, Daniele Zambon, Cesare Alippi, Barbara Hammer
ICLR4
2021 Learning Graph Cellular Automata
abstract
Cellular automata (CA) are a class of computational models that exhibit rich dynamics emerging from the local interaction of cells arranged in a regular lattice. In this work we focus on a generalised version of typical CA, called graph cellular automata (GCA), in which the lattice structure is replaced by an arbitrary graph. In particular, we extend previous work that used convolutional neural networks to learn the transition rule of conventional CA and we use graph neural networks to learn a variety of transition rules for GCA. First, we present a general-purpose architecture for learning GCA, and we show that it can represent any arbitrary GCA with finite and discrete state space. Then, we test our approach on three different tasks: 1) learning the transition rule of a GCA on a Voronoi tessellation; 2) imitating the behaviour of a group of flocking agents; 3) learning a rule that converges to a desired target state.
Daniele Grattarola, Lorenzo Livi, Cesare Alippi
NeurIPS3
2021 Birdsong Detection at the Edge with Deep Learning
abstract
Understanding the distribution of bird species and populations and learning how birds behave and communicate are of great importance in wildlife biology, animal ecology, conservation of ecosystems, and assessing the effects of climate change and urbanization. The temporal and spatial limitations of human observation have motivated significant efforts to develop technology for bird song and vocalization detection and classification. While solutions based on signal processing and machine learning are extant, they are limited in various combinations of speed, computational complexity, and memory use, as well as in detection/classification capability in real-world conditions. This paper introduces ToucaNet, a deep neural network for birdsong detection based on transfer-learning, a deep learning mechanism allowing us to exploit knowledge acquired on various tasks: this enables us to speed up training and shows improved detection accuracy. ToucaNet provides birdsong detection accuracy in line with the best solutions in the literature but with much less computational complexity and memory demand. We also introduce BarbNet, an approximated version of ToucaNet tailored for Internet-of-Things (IoT) units. We show the proposed solution’s effectiveness and efficiency in terms of detection accuracy and the implementation feasibility in real-world IoT devices, with specific results for the STM32 Nucleo H7 board, which is based on an ARM Cortex-M7 processor. To our best knowledge, this is the first birdsong detection algorithm designed to take into account constraints on memory, computational speed, and power usage of embedded devices. Thus, this work points the way to cost-effective IoT technology for at-scale intelligent birdsong data collection and analysis in the field.
Simone Disabato, Giuseppe Canonaco, Paul G. Flikkema, Manuel Roveri, Cesare Alippi
SMARTCOMP5
2021 Gaussian Approximation for Bias Reduction in Q-Learning
abstract
Temporal-Difference off-policy algorithms are among the building blocks of reinforcement learning (RL). Within this family, Q-Learning is arguably the most famous one, which has been widely studied and extended. The update rule of Q-learning involves the use of the maximum operator to estimate the maximum expected value of the return. However, this estimate is positively biased, and may hinder the learning process, especially in stochastic environments and when function approximation is used. We introduce the Weighted Estimator as an effective solution to mitigate the negative effects of overestimation in Q-Learning. The Weighted Estimator estimates the maximum expected value as a weighted sum of the action values, with the weights being the probabilities that each action value is the maximum. In this work, we study the problem from the statistical perspective of estimating the maximum expected value of a set of random variables and provide bounds to the bias and the variance of the Weighted Estimator, showing its advantages over other estimators present in literature. Then, we derive algorithms to enable the use of the Weighted Estimator, in place of the Maximum Estimator, in online and batch RL, and we introduce a novel algorithm for deep RL. Finally, we empirically evaluate our algorithms in a large set of heterogeneous problems, encompassing discrete and continuous, low and high dimensional, deterministic and stochastic environments. Experimental results show the effectiveness of the Weighted Estimator in controlling the bias of the estimate, resulting in better performance than representative baselines and robust learning w.r.t. a large set of diverse environments.
Carlo D'Eramo, Andrea Cini, Alessandro Nuara, Matteo Pirotta, Cesare Alippi, Jan Peters 0001, Marcello Restelli
J. Mach. Learn. Res.5
2021 Distributed Deep Convolutional Neural Networks for the Internet-of-Things
abstract
Severe constraints on memory and computation characterizing the Internet-of-Things (IoT) units may prevent the execution of Deep Learning (DL)-based solutions, which typically demand large memory and high processing load. In order to support a real-time execution of the considered DL model at the IoT unit level, DL solutions must be designed having in mind constraints on memory and processing capability exposed by the chosen IoT technology. In this article, we introduce a design methodology aiming at allocating the execution of Convolutional Neural Networks (CNNs) on a distributed IoT application. Such a methodology is formalized as an optimization problem where the latency between the data-gathering phase and the subsequent decision-making one is minimized, within the given constraints on memory and processing load at the units level. The methodology supports multiple sources of data as well as multiple CNNs in execution on the same IoT system allowing the design of CNN-based applications demanding autonomy, low decision-latency, and high Quality-of-Service.
Simone Disabato, Manuel Roveri, Cesare Alippi
IEEE Trans. Computers3
2021 Sliding-Mode Surface-Based Approximate Optimal Control for Uncertain Nonlinear Systems With Asymptotically Stable Critic Structure
abstract
This article develops a novel sliding-mode surface (SMS)-based approximate optimal control scheme for a large class of nonlinear systems affected by unknown mismatched perturbations. The observer-based perturbation estimation procedure is employed to establish the online updated value function. The solution to the Hamilton-Jacobi-Bellman equation is approximated by an SMS-based critic neural network whose weights error dynamics is designed to be asymptotically stable by nested update laws. The sliding-mode control strategy is combined with the approximate optimal control design procedure to obtain a faster control action. The stability is proved based on the Lyapunov's direct method. The simulation results show the effectiveness of the developed control scheme.
Bo Zhao 0015, Derong Liu 0001, Cesare Alippi
IEEE Trans. Cybern.3
2020 Spectral Clustering with Graph Neural Networks for Graph Pooling
abstract
Spectral clustering (SC) is a popular clustering technique to find strongly connected communities on a graph. SC can be used in Graph Neural Networks (GNNs) to implement pooling operations that aggregate nodes belonging to the same cluster. However, the eigendecomposition of the Laplacian is expensive and, since clustering results are graph-specific, pooling methods based on SC must perform a new optimization for each new sample. In this paper, we propose a graph clustering approach that addresses these limitations of SC. We formulate a continuous relaxation of the normalized minCUT problem and train a GNN to compute cluster assignments that minimize this objective. Our GNN-based implementation is differentiable, does not require to compute the spectral decomposition, and learns a clustering function that can be quickly evaluated on out-of-sample graphs. From the proposed clustering method, we design a graph pooling operator that overcomes some important limitations of state-of-the-art graph pooling techniques and achieves the best performance in several supervised and unsupervised tasks.
Filippo Maria Bianchi, Daniele Grattarola, Cesare Alippi
ICML3
2020 Graph Random Neural Features for Distance-Preserving Graph Representations
abstract
We present Graph Random Neural Features (GRNF), a novel embedding method from graph-structured data to real vectors based on a family of graph neural networks. The embedding naturally deals with graph isomorphism and preserves the metric structure of the graph domain, in probability. In addition to being an explicit embedding method, it also allows us to efficiently and effectively approximate graph metric distances (as well as complete kernel functions); a criterion to select the embedding dimension trading off the approximation accuracy with the computational cost is also provided. GRNF can be used within traditional processing methods or as a training-free input layer of a graph neural network. The theoretical guarantees that accompany GRNF ensure that the considered graph distance is metric, hence allowing to distinguish any pair of non-isomorphic graphs.
Daniele Zambon, Cesare Alippi, Lorenzo Livi
ICML2
2020 PIF: Anomaly detection via preference embedding
abstract
We address the problem of detecting anomalies with respect to structured patterns. To this end, we conceive a novel anomaly detection method called PIF, that combines the advantages of adaptive isolation methods with the flexibility of preference embedding. Specifically, we propose to embed the data in a high dimensional space where an efficient tree-based method, PI-Forest, is employed to compute an anomaly score. Experiments on synthetic and real datasets demonstrate that PIF favorably compares with state-of-the-art anomaly detection techniques, and confirm that PI-Forest is better at measuring arbitrary distances and isolate points in the preference space.
Filippo Leveni, Luca Magri 0002, Giacomo Boracchi, Cesare Alippi
ICPR4
2020 Cluster-based Aggregate Load Forecasting with Deep Neural Networks
abstract
Highly accurate power demand forecasting represents one of key challenges of Smart Grid applications. In this setting, a large number of Smart Meters produces huge amounts of data that need to be processed to predict the load requested by the grid. Due to the high dimensionality of the problem, this often results in the adoption of simple aggregation strategies for the power that fail in capturing the relational information existing among the different types of user. A possible alternative, known as Cluster-based Aggregate Forecasting, consists in clustering the load profiles and, on top of that, building predictors of the aggregate at the cluster-level. In this work we explore the technique in the context of predictors based on deep recurrent neural networks and address the scalability issues presenting neural architectures adequate to process cluster-level aggregates. The proposed methods are finally evaluated both on a publicly available benchmark and a heterogenous dataset of Smart Meter data from an entire, medium-sized, Swiss town.
Andrea Cini, Slobodan Lukovic, Cesare Alippi
IJCNN3
2020 Advances in deep neural information processing
Dongbin Zhao, Shukai Duan 0001, Zheng Yan 0001, Cesare Alippi
Neurocomputing4
2020 Change Detection in Graph Streams by Learning Graph Embeddings on Constant-Curvature Manifolds
abstract
The space of graphs is often characterized by a nontrivial geometry, which complicates learning and inference in practical applications. A common approach is to use embedding techniques to represent graphs as points in a conventional Euclidean space, but non-Euclidean spaces have often been shown to be better suited for embedding graphs. Among these, constant-curvature Riemannian manifolds (CCMs) offer embedding spaces suitable for studying the statistical properties of a graph distribution, as they provide ways to easily compute metric geodesic distances. In this paper, we focus on the problem of detecting changes in stationarity in a stream of attributed graphs. To this end, we introduce a novel change detection framework based on neural networks and CCMs, which takes into account the non-Euclidean nature of graphs. Our contribution in this paper is twofold. First, via a novel approach based on adversarial learning, we compute graph embeddings by training an autoencoder to represent graphs on CCMs. Second, we introduce two novel change detection tests operating on CCMs. We perform experiments on synthetic data, as well as two real-world application scenarios: the detection of epileptic seizures using functional connectivity brain networks and the detection of hostility between two subjects, using human skeletal graphs. Results show that the proposed methods are able to detect even small changes in a graph-generating process, consistently outperforming approaches based on Euclidean embeddings.
Daniele Grattarola, Daniele Zambon, Lorenzo Livi, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.4
2019 Autoregressive Models for Sequences of Graphs
abstract
This paper proposes an autoregressive (AR) model for sequences of graphs, which generalises traditional AR models. A first novelty consists in formalising the AR model for a very general family of graphs, characterised by a variable topology, and attributes associated with nodes and edges. A graph neural network (GNN) is also proposed to learn the AR function associated with the graph-generating process (GGP), and subsequently predict the next graph in a sequence. The proposed method is compared with four baselines on synthetic GGPs, denoting a significantly better performance on all considered problems.
Daniele Zambon, Daniele Grattarola, Lorenzo Livi, Cesare Alippi
IJCNN4
2018 Security: the dark side of approximate computing?
abstract
Approximate computing promises significant advantages over more traditional computing architectures with respect to circuit area, performance, power efficiency, flexibility, and cost. Its use is suitable in applications where limited and controlled inaccuracies are tolerable or uncertainty is intrinsic in input or their data processing, e.g., as it happens in (deep-) machine learning, image and signal processing. This paper discusses a dimension of approximate computing that has been neglected so far, despite it represents nowadays a major asset, that of security. A number of hardware-related security threats are considered, and the implications of approximate circuits or systems designed to address these threats are discussed.
Francesco Regazzoni 0001, Cesare Alippi, Ilia Polian
ICCAD2
2018 Anomaly and Change Detection in Graph Streams through Constant-Curvature Manifold Embeddings
abstract
Mapping complex input data into suitable lower dimensional manifolds is a common procedure in machine learning. This step is beneficial mainly for two reasons: (1) it reduces the data dimensionality and (2) it provides a new data representation possibly characterised by convenient geometric properties. Euclidean spaces are by far the most widely used embedding spaces, thanks to their well-understood structure and large availability of consolidated inference methods. However, recent research demonstrated that many types of complex data (e.g., those represented as graphs) are actually better described by non-Euclidean geometries. Here, we investigate how embedding graphs on constant-curvature manifolds (hyper-spherical and hyperbolic manifolds) impacts on the ability to detect changes in sequences of attributed graphs. The proposed methodology consists in embedding graphs into a geometric space and perform change detection there by means of conventional methods for numerical streams. The curvature of the space is a parameter that we learn to reproduce the geometry of the original application-dependent graph space. Preliminary experimental results show the potential capability of representing graphs by means of curved manifold, in particular for change and anomaly detection problems.
Daniele Zambon, Lorenzo Livi, Cesare Alippi
IJCNN3
2018 Moving convolutional neural networks to embedded systems: the alexnet and VGG-16 case
abstract
Execution of deep learning solutions is mostly restricted to high performing computing platforms, e.g., those endowed with GPUs or FPGAs, due to the high demand on computation and memory such solutions require. Despite the fact that dedicated hardware is nowadays subject of research and effective solutions exist, we envision a future where deep learning solutions -here Convolutional Neural Networks (CNNs)- are mostly executed by low-cost off-the shelf embedded platforms already available in the market. This paper moves in this direction and aims at filling the gap between CNNs and embedded systems by introducing a methodology for the design and porting of CNNs to limited in resources embedded systems. In order to achieve this goal we employ approximate computing techniques to reduce the computational load and memory occupation of the deep learning architecture by compromising accuracy with memory and computation. The proposed methodology has been validated on two well-know CNNs, i.e., AlexNet and VGG-16, applied to an image-recognition application and ported to two relevant off-the-shelf embedded platforms.
Cesare Alippi, Simone Disabato, Manuel Roveri
IPSN1
2018 Investigating Echo-State Networks Dynamics by Means of Recurrence Analysis
abstract
In this paper, we elaborate over the well-known interpretability issue in echo-state networks (ESNs). The idea is to investigate the dynamics of reservoir neurons with time-series analysis techniques developed in complex systems research. Notably, we analyze time series of neuron activations with recurrence plots (RPs) and recurrence quantification analysis (RQA), which permit to visualize and characterize high-dimensional dynamical systems. We show that this approach is useful in a number of ways. First, the 2-D representation offered by RPs provides a visualization of the high-dimensional reservoir dynamics. Our results suggest that, if the network is stable, reservoir and input generate similar line patterns in the respective RPs. Conversely, as the ESN becomes unstable, the patterns in the RP of the reservoir change. As a second result, we show that an RQA measure, called , is highly correlated with the well-established maximal local Lyapunov exponent. This suggests that complexity measures based on RP diagonal lines distribution can quantify network stability. Finally, our analysis shows that all RQA measures fluctuate on the proximity of the so-called edge of stability, where an ESN typically achieves maximum computational capability. We leverage on this property to determine the edge of stability and show that our criterion is more accurate than two well-known counterparts, both based on the Jacobian matrix of the reservoir. Therefore, we claim that RPs and RQA-based analyses are valuable tools to design an ESN, given a specific problem.
Filippo Maria Bianchi, Lorenzo Livi, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.3
2018 A pdf-Free Change Detection Test Based on Density Difference Estimation
abstract
The ability to detect online changes in stationarity or time variance in a data stream is a hot research topic with striking implications. In this paper, we propose a novel probability density function-free change detection test, which is based on the least squares density-difference estimation method and operates online on multidimensional inputs. The test does not require any assumption about the underlying data distribution, and is able to operate immediately after having been configured by adopting a reservoir sampling mechanism. Thresholds requested to detect a change are automatically derived once a false positive rate is set by the application designer. Comprehensive experiments validate the effectiveness in detection of the proposed method both in terms of detection promptness and accuracy.
Li Bu, Cesare Alippi, Dongbin Zhao
IEEE Trans. Neural Networks Learn. Syst.2
2018 Determination of the Edge of Criticality in Echo State Networks Through Fisher Information Maximization
abstract
It is a widely accepted fact that the computational capability of recurrent neural networks (RNNs) is maximized on the so-called "edge of criticality." Once the network operates in this configuration, it performs efficiently on a specific application both in terms of: 1) low prediction error and 2) high short-term memory capacity. Since the behavior of recurrent networks is strongly influenced by the particular input signal driving the dynamics, a universal, application-independent method for determining the edge of criticality is still missing. In this paper, we aim at addressing this issue by proposing a theoretically motivated, unsupervised method based on Fisher information for determining the edge of criticality in RNNs. It is proved that Fisher information is maximized for (finite-size) systems operating in such critical regions. However, Fisher information is notoriously difficult to compute and requires the analytic form of the probability density function ruling the system behavior. This paper takes advantage of a recently developed nonparametric estimator of the Fisher information matrix and provides a method to determine the critical region of echo state networks (ESNs), a particular class of recurrent networks. The considered control parameters, which indirectly affect the ESN performance, are explored to identify those configurations lying on the edge of criticality and, as such, maximizing Fisher information and computational performance. Experimental results on benchmarks and real-world data demonstrate the effectiveness of the proposed method.
Lorenzo Livi, Filippo Maria Bianchi, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.3
2018 Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy
abstract
Detecting frauds in credit card transactions is perhaps one of the best testbeds for computational intelligence algorithms. In fact, this problem involves a number of relevant challenges, namely: concept drift (customers' habits evolve and fraudsters change their strategies over time), class imbalance (genuine transactions far outnumber frauds), and verification latency (only a small set of transactions are timely checked by investigators). However, the vast majority of learning algorithms that have been proposed for fraud detection rely on assumptions that hardly hold in a real-world fraud-detection system (FDS). This lack of realism concerns two main aspects: 1) the way and timing with which supervised information is provided and 2) the measures used to assess fraud-detection performance. This paper has three major contributions. First, we propose, with the help of our industrial partner, a formalization of the fraud-detection problem that realistically describes the operating conditions of FDSs that everyday analyze massive streams of credit card transactions. We also illustrate the most appropriate performance measures to be used for fraud-detection purposes. Second, we design and assess a novel learning strategy that effectively addresses class imbalance, concept drift, and verification latency. Third, in our experiments, we demonstrate the impact of class unbalance and concept drift in a real-world data stream containing more than 75 million transactions, authorized over a time window of three years.
Andrea Dal Pozzolo, Giacomo Boracchi, Olivier Caelen, Cesare Alippi, Gianluca Bontempi
IEEE Trans. Neural Networks Learn. Syst.4
2018 Concept Drift and Anomaly Detection in Graph Streams
abstract
Graph representations offer powerful and intuitive ways to describe data in a multitude of application domains. Here, we consider stochastic processes generating graphs and propose a methodology for detecting changes in stationarity of such processes. The methodology is general and considers a process generating attributed graphs with a variable number of vertices/edges, without the need to assume a one-to-one correspondence between vertices at different time steps. The methodology acts by embedding every graph of the stream into a vector domain, where a conventional multivariate change detection procedure can be easily applied. We ground the soundness of our proposal by proving several theoretical results. In addition, we provide a specific implementation of the methodology and evaluate its effectiveness on several detection problems involving attributed graphs representing biological molecules and drawings. Experimental results are contrasted with respect to suitable baseline methods, demonstrating the effectiveness of our approach.
Daniele Zambon, Cesare Alippi, Lorenzo Livi
IEEE Trans. Neural Networks Learn. Syst.2
2017 Detecting changes at the sensor level in cyber-physical systems: Methodology and technological implementation
abstract
Self-adaptive Cyber-Physical Systems (CPSs) enrich CPSs functionalities by introducing self-configuration, self-management, and self-healing skills. Such skills, which are crucial to support adaptation mechanisms, take advantage of the ability to detect changes in the acquired datastreams, e.g., induced by faults affecting sensors/actuators or time-variant environments. In turn, change detection permits CPSs to enable adaptive mechanisms such as reconfiguration of some functionalities to track or mitigate the effect of the change. This paper introduces a novel methodology together with a technological implementation specifically designed for detecting changes affecting the sensor acquisitions in units of CPSs. The methodology requires: 1) learning the signal model; 2) design a model-free change detection test; 3) design a change-point method to validate the detected change. A technological implementation of the proposed methodology encompassing linear predictive models, the ICI-based change detection test and the Mann-Whitney change-point method is introduced and tested on the ST STM32 Nucleo platform. The high detection accuracy altogether with the low computational load and memory occupation make the proposed methodology (and its technological implementation) well suited for self-adaptive CPSs.
Cesare Alippi, Viviana D'Alto, Mirko Falchetto, Danilo Pau, Manuel Roveri
IJCNN1
2017 Critical echo state network dynamics by means of Fisher information maximization
abstract
The computational capability of an Echo State Network (ESN), expressed in terms of low prediction error and high short-term memory capacity, is maximized on the so-called “edge of criticality”. In this paper we present a novel, unsupervised approach to identify this edge and, accordingly, we determine hyperparameters configuration that maximize network performance. The proposed method is application-independent and stems from recent theoretical results consolidating the link between Fisher information and critical phase transitions. We show how to identify optimal ESN hyperparameters by relying only on the Fisher information matrix (FIM) estimated from the activations of hidden neurons. In order to take into account the particular input signal driving the network dynamics, we adopt a recently proposed non-parametric FIM estimator. Experimental results on a set of standard benchmarks are provided and discussed, demonstrating the validity of the proposed method.
Filippo Maria Bianchi, Lorenzo Livi, Robert Jenssen, Cesare Alippi
IJCNN4
2017 A lightweight and energy-efficient Internet-of-birds tracking system
abstract
In this paper we introduce a novel engineering system for tracking animal movements. Study of animal movement has implications in several relevant and challenging research areas, i.e., from ornithology to global ecology and wildlife management. Inspired by the Internet-of-things vision, the proposed tracking system relies on GSM-based tracking devices that are directly connected to the GSM network to remotely transmit/exchange data with Application Servers/Cloud. The novel characteristic of the proposed GSM-based tracking system is the use of the GSM network for both localization and transmission. This allows to simplify the design of both hardware and software, hence reducing, size, weight and cost of the GSM-based tracking device as well as the fieldwork effort necessary to obtain tracking data. Moreover, this joint localization/transmission phase allows to reduce the energy consumption of the GSM-based tracking device, hence prolonging the life-time of the system and increasing the Quality-of-Service of the envisaged tracking application. We performed a detailed energy assessment of the GSM-based tracking device, while the whole GSM-based tracking system has been tested in a real deployment on four greater flamingos (Phoenicopterus roseus) for approximately 11 months in Northern Italy.
Cesare Alippi, Roberto Ambrosini, Violetta Longoni, Dario Cogliati, Manuel Roveri
PerCom1
2017 Solving Multiobjective Optimization Problems in Unknown Dynamic Environments: An Inverse Modeling Approach
abstract
Evolutionary multiobjective optimization in dynamic environments is a challenging task, as it requires the optimization algorithm converging to a time-variant Pareto optimal front. This paper proposes a dynamic multiobjective optimization algorithm which utilizes an inverse model set to guide the search toward promising decision regions. In order to reduce the number of fitness evalutions for change detection purpose, a two-stage change detection test is proposed which uses the inverse model set to check potential changes in the objective function landscape. Both static and dynamic multiobjective benchmark optimization problems have been considered to evaluate the performance of the proposed algorithm. Experimental results show that the improvement in optimization performance is achievable when the proposed inverse model set is adopted.
Sen Bong Gee, Kay Chen Tan, Cesare Alippi
IEEE Trans. Cybern.3
2017 Hierarchical Change-Detection Tests
abstract
We present hierarchical change-detection tests (HCDTs), as effective online algorithms for detecting changes in datastreams. HCDTs are characterized by a hierarchical architecture composed of a detection layer and a validation layer. The detection layer steadily analyzes the input datastream by means of an online, sequential CDT, which operates as a low-complexity trigger that promptly detects possible changes in the process generating the data. The validation layer is activated when the detection one reveals a change, and performs an offline, more sophisticated analysis on recently acquired data to reduce false alarms. Our experiments show that, when the process generating the datastream is unknown, as it is mostly the case in the real world, HCDTs achieve a far more advantageous tradeoff between false-positive rate and detection delay than their single-layered, more traditional counterpart. Moreover, the successful interplay between the two layers permits HCDTs to automatically reconfigure after having detected and validated a change. Thus, HCDTs are able to reveal further departures from the postchange state of the data-generating process.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IEEE Trans. Neural Networks Learn. Syst.1
2017 One-Class Classifiers Based on Entropic Spanning Graphs
abstract
One-class classifiers offer valuable tools to assess the presence of outliers in data. In this paper, we propose a design methodology for one-class classifiers based on entropic spanning graphs. Our approach also takes into account the possibility to process nonnumeric data by means of an embedding procedure. The spanning graph is learned on the embedded input data, and the outcoming partition of vertices defines the classifier. The final partition is derived by exploiting a criterion based on mutual information minimization. Here, we compute the mutual information by using a convenient formulation provided in terms of the -Jensen difference. Once training is completed, in order to associate a confidence level with the classifier decision, a graph-based fuzzy model is constructed. The fuzzification process is based only on topological information of the vertices of the entropic spanning graph. As such, the proposed one-class classifier is suitable also for data characterized by complex geometric structures. We provide experiments on well-known benchmarks containing both feature vectors and labeled graphs. In addition, we apply the method to the protein solubility recognition problem by considering several representations for the input samples. Experimental results demonstrate the effectiveness and versatility of the proposed method with respect to other state-of-the-art approaches.
Lorenzo Livi, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.2
2017 An Incremental Change Detection Test Based on Density Difference Estimation
abstract
We propose incremental least squares density difference (LSDD) change detection method, an incremental test to detect changes in stationarity based on the difference between the unknown prechange and the post-change probability density functions (pdfs). The method is computationally light and, hence, adequate to process continuous data streams, as those emerging from the Internet of Things and the big data framework. The incremental change detection test operates on two nonoverlapping data windows to estimate the LSDD between the two pdfs. We construct a theoretical framework that shows how the distribution of LSDD values follows a linear combination of χ2distributions and provides thresholds to control false positive rates. The proposed test can operate online, with needed estimates and thresholds computed incrementally as fresh samples come. Comprehensive experiments validate the effectiveness of the test both in detecting abrupt and drift types of changes.
Li Bu, Dongbin Zhao, Cesare Alippi
IEEE Trans. Syst. Man Cybern. Syst.3
2016 Change Detection in Multivariate Datastreams: Likelihood and Detectability Loss
Cesare Alippi, Giacomo Boracchi, Diego Carrera, Manuel Roveri
IJCAI1
2016 Online model-free sensor fault identification and dictionary learning in Cyber-Physical Systems
abstract
This paper presents a model-free method for the online identification of sensor faults and learning of their fault dictionary. The method, designed having in mind Cyber-Physical Systems (CPSs), takes advantage of functional relationships among the datastreams acquired by CPS sensing units. Existing model-free change detection mechanisms are proposed to detect faults and identify the fault type thanks to a fault dictionary which is built over time. The main features of the proposed algorithm are its ability to operate without requiring any a priori information about the system under inspection or the nature of the possibly occurring faults. As such, the method follows the model-free approach, characterized by the fact the fault dictionary is constructed online once faults are detected. Whenever available, humans can be considered in the loop to label a fault or a fault class in the dictionary as well as introduce fault instances generated thanks to a priori information. Experimental results on both synthetic and real datasets corroborate the effectiveness of the proposed fault diagnosis system.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IJCNN1
2016 Ensemble LSDD-based change detection tests
abstract
The least squares density difference change detection test (LSDD-CDT) has proven to be an effective method in detecting concept drift by inspecting features derived from the discrepancy between two probability density functions (pdfs). The first pdf is associated with the concept drift free case, the second to the possible post change one. Interestingly, the method permits to control the ratio of false positives. This paper introduces and investigates the performance of a family of LSDD methods constructed by exploring different ensemble options applied to the basic CDT procedure. Experiments show that most of proposed methods are characterized by improved performance in change detection once compared with the direct ensemble-free counterpart.
Li Bu, Cesare Alippi, Dongbin Zhao
IJCNN2
2016 One-class classification through mutual information minimization
abstract
In one-class classification problems, a model is synthesized by using only information coming from the nominal state of the data generating process. Many important applications can be cast in the one-class classification framework, such as anomaly and change in stationarity detection, and fault recognition. In this paper, we present a novel design methodology for one-class classifiers derived from graph-based entropy estimators. The entropic graph is used to generate a partition of the input nominal conditions, which corresponds to the classifier model. Here we propose a criterion based on mutual information minimization to learn such a partition. The α-Jensen difference is considered, which provides a convenient way for estimating the mutual information. The classifier incorporates also a fuzzy model, providing a confidence value for a generic test sample during operational modality, expressed as a membership degree of the sample to the nominal conditions class. The fuzzification mechanism is based only on topological properties of the entropic spanning graph vertices; as such, it allows to model clusters of arbitrary shapes. We show preliminary - yet very promising - results on both synthetic problems and real-world datasets for one-class classification.
Lorenzo Livi, Cesare Alippi
IJCNN2
2016 RTI Goes Wild: Radio Tomographic Imaging for Outdoor People Detection and Localization
abstract
In recent years, Radio frequency (RF) sensor networks have been used to localize people indoor without requiring them to wear invasive electronic devices. These wireless mesh networks, formed by low-power radio transceivers, continuously measure the received signal strength (RSS) of the links. Radio Tomographic Imaging (RTI) is a technique that generates, starting from these RSS measurements, 2D images of the change in the electromagnetic field inside the area covered by the radio transceivers to spot the presence and movements of animates (e.g., people, large animals) or large metallic objects (e.g., cars). Here, we present a RTI system for localizing and tracking people outdoors. Differently than in indoor environments where the RSS does not change significantly with time unless people are found in the monitored area, the outdoor RSS signal is time-variant, e.g., due to rainfall or wind-driven foliage. We present a novel outdoor RTI method that, despite the nonstationary noise introduced in the RSS data by the environment, achieves high localization accuracy and dramatically reduces the energy consumption of the sensing units. Experimental results demonstrate that the system accurately detects and tracks a person in real-time in a large forested area under varying environmental conditions, significantly reducing false positives, localization error and energy consumption compared to state-of-the-art RTI methods.
Cesare Alippi, Maurizio Bocca, Giacomo Boracchi, Neal Patwari, Manuel Roveri
IEEE Trans. Mob. Comput.1
2015 Credit card fraud detection and concept-drift adaptation with delayed supervised information
abstract
Most fraud-detection systems (FDSs) monitor streams of credit card transactions by means of classifiers returning alerts for the riskiest payments. Fraud detection is notably a challenging problem because of concept drift (i.e. customers' habits evolve) and class unbalance (i.e. genuine transactions far outnumber frauds). Also, FDSs differ from conventional classification because, in a first phase, only a small set of supervised samples is provided by human investigators who have time to assess only a reduced number of alerts. Labels of the vast majority of transactions are made available only several days later, when customers have possibly reported unauthorized transactions. The delay in obtaining accurate labels and the interaction between alerts and supervised information have to be carefully taken into consideration when learning in a concept-drifting environment. In this paper we address a realistic fraud-detection setting and we show that investigator's feedbacks and delayed labels have to be handled separately. We design two FDSs on the basis of an ensemble and a sliding-window approach and we show that the winning strategy consists in training two separate classifiers (on feedbacks and delayed labels, respectively), and then aggregating the outcomes. Experiments on large dataset of real-world transactions show that the alert precision, which is the primary concern of investigators, can be substantially improved by the proposed approach.
Andrea Dal Pozzolo, Giacomo Boracchi, Olivier Caelen, Cesare Alippi, Gianluca Bontempi
IJCNN4
2014 Change detection in streams of signals with sparse representations
abstract
We propose a novel approach to performing change-detection based on sparse representations and dictionary learning. We operate on observations that are finite support signals, which in stationary conditions lie within a union of low dimensional subspaces. We model changes as perturbations of these subspaces and provide an online and sequential monitoring solution to detect them. This approach allows extension of the change-detection framework to operate on streams of observations that are signals, rather than scalar or multi-variate measurements, and is shown to be effective for both synthetic data and on bursts acquired by rockfall monitoring systems.
Cesare Alippi, Giacomo Boracchi, Brendt Wohlberg
ICASSP1
2014 Dual Heuristic dynamic Programming for nonlinear discrete-time uncertain systems with state delay
Bin Wang 0034, Dongbin Zhao, Cesare Alippi, Derong Liu 0001
Neurocomputing3
2014 Full-range adaptive cruise control based on supervised adaptive dynamic programming
Dongbin Zhao, Zhaohui Hu, Zhongpu Xia, Cesare Alippi, Yuanheng Zhu, Ding Wang 0001
Neurocomputing4
2014 A Self-Building and Cluster-Based Cognitive Fault Diagnosis System for Sensor Networks
abstract
Cognitive fault diagnosis systems differentiate from more traditional solutions by providing online strategies to create and update the fault-free and the faulty classes directly from incoming data. This aspect is of paramount relevance within the big data framework, since measurements are there immediately processed to detect and identify the upsurge of potential faults. This paper introduces a novel cognitive fault diagnosis framework for processes described by nonlinear dynamic systems that inspects changes in the existing relationships among sensors. The proposed framework is based on an evolving clustering algorithm that operates in the parameter space of time invariant linear models approximating such relationships. During the operational life, parameter vectors associated with models thought not to belong to the nominal state are either labeled as outlier or fault. New classes of faults, here considered to propagate to the model parameters according to an abrupt profile, are created online as they appear. At the same time, existing classes can merge, depending on the information content carried by incoming data.
Cesare Alippi, Manuel Roveri, Francesco Trovò
IEEE Trans. Neural Networks Learn. Syst.1
2014 Guest Editorial Learning in Nonstationary and Evolving Environments
abstract
The papers in this special issue encompass a broad spectrum of scenarios and problems, including some of the new challenges and fundamental problems that are encountered in learning in nonstationary and evolving environments.
Robi Polikar, Cesare Alippi
IEEE Trans. Neural Networks Learn. Syst.2
2014 Detecting and Reacting to Changes in Sensing Units: The Active Classifier Case
abstract
The ability to detect concept drift, i.e., a structural change in the acquired datastream, and react accordingly is a major achievement for intelligent sensing units. This ability allows the unit, for actively tuning the application, to maintain high performance, changing online the operational strategy, detecting and isolating possible occurring faults to name a few tasks. In the paper, we consider a just-in-time strategy for adaptation; the sensing unit reacts exactly when needed, i.e., when concept drift is detected. Change detection tests (CDTs), designed to inspect structural changes in industrial and environmental data, are coupled here with adaptive k-nearest neighbor and support vector machine classifiers, and suitably retrained when the change is detected. Computational complexity and memory requirements of the CDT and the classifier, due to precious limited resources in embedded sensing, are taken into account in the application design. We show that a hierarchical CDT coupled with an adaptive resource-aware classifier is a suitable tool for processing and classifying sequential streams of data.
Cesare Alippi, Derong Liu 0001, Dongbin Zhao, Li Bu
IEEE Trans. Syst. Man Cybern. Syst.1
2013 A prior-free encode-decode change detection test to inspect datastreams for concept drift
abstract
Online change detection in datastreams has attracted many researchers and is becoming a very hot topic whose relevance will further increase with research on Big Data. Concept drift is induced by changes in stationarity of the process generating the data caused by faults, time variance of the environment and inaccuracy of the change detection mechanism. Here, we propose a recurrent auto-associative Encode-Decode machine trained to reconstruct input data. The generated residual is then inspected for structural changes with a Change Detection Test (CDT). Although any CDT can be used, in the paper we focus the attention on the Hierarchical Intersection of Confidence Intervals change detection test for its capability of controlling false positives with a two layered test and an online version of the Lepage Change Point Model. Once concept drift is detected, the designed Encode-Decode machine, globally acting as an Encode-Decode CDT, is retrained on new data to detect subsequent changes.
Cesare Alippi, Li Bu, Dongbin Zhao
IJCNN1
2013 Model ensemble for an effective on-line reconstruction of missing data in sensor networks
abstract
The literature has shown that model ensemble techniques are particularly effective to solve regression/classification applications by providing, given a suitable aggregation mechanism, a better generalization ability than the generic model of the ensemble. However, only few recent results consider the use of ensembles for a time-dependent framework, with focus on time-series forecasting. Here, we propose the use of ensemble of models to an on-line reconstruction of missing data coming from a sensor network. Reconstructing missing data is of paramount importance for any further data processing and must be carried out on-line not to introduce unnecessary latency when data lead to a decision or control action. The ensemble is designed by both exploiting temporal and spatial dependencies existing among the sensors composing the network. An effective aggregation mechanism is proposed for the considered models to improve the generalization ability of the ensemble. Results demonstrate the effectiveness of the proposed approach in reconstructing missing data.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IJCNN1
2013 Ensembles of change-point methods to estimate the change point in residual sequences
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
Soft Comput.1
2013 Special issue on intelligent control and information processing
Dongbin Zhao, Cesare Alippi, Derong Liu 0001, Huaguang Zhang
Soft Comput.2
2013 Just-In-Time Classifiers for Recurrent Concepts
abstract
Just-in-time (JIT) classifiers operate in evolving environments by classifying instances and reacting to concept drift. In stationary conditions, a JIT classifier improves its accuracy over time by exploiting additional supervised information coming from the field. In nonstationary conditions, however, the classifier reacts as soon as concept drift is detected; the current classification setup is discarded and a suitable one activated to keep the accuracy high. We present a novel generation of JIT classifiers able to deal with recurrent concept drift by means of a practical formalization of the concept representation and the definition of a set of operators working on such representations. The concept-drift detection activity, which is crucial in promptly reacting to changes exactly when needed, is advanced by considering change-detection tests monitoring both inputs and classes distributions.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IEEE Trans. Neural Networks Learn. Syst.1
2013 A Cognitive Fault Diagnosis System for Distributed Sensor Networks
abstract
This paper introduces a novel cognitive fault diagnosis system (FDS) for distributed sensor networks that takes advantage of spatial and temporal relationships among sensors. The proposed FDS relies on a suitable functional graph representation of the network and a two-layer hierarchical architecture designed to promptly detect and isolate faults. The lower processing layer exploits a novel change detection test (CDT) based on hidden Markov models (HMMs) configured to detect variations in the relationships between couples of sensors. HMMs work in the parameter space of linear time-invariant dynamic systems, approximating, over time, the relationship between two sensors; changes in the approximating model are detected by inspecting the HMM likelihood. Information provided by the CDT layer is then passed to the cognitive one, which, by exploiting the graph representation of the network, aggregates information to discriminate among faults, changes in the environment, and false positives induced by the model bias of the HMMs.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IEEE Trans. Neural Networks Learn. Syst.1
2013 A high-frequency sampling monitoring system for environmental and structural applications
abstract
High-frequency sampling is not only a prerogative of high-energy physics or machinery diagnostic monitoring: critical environmental and structural health monitoring applications also have such a challenging constraint. Moreover, such unique design constraints are often coupled with the requirement of high synchronism among the distributed acquisition units, minimal energy consumption, and large communication bandwidth. Such severe constraints have led scholars to suggest wired centralized monitoring solutions, which have only recently been complemented with wireless technologies. This article suggests a hybrid wireless-wired monitoring system combining the advantages of wireless and wired technologies within a distributed high-frequency-sampling framework. The suggested architecture satisfies the mentioned constraints, thanks to an ad-hoc design of the hardware, the availability of efficient energy management policies, and up-to-date harvesting mechanisms. At the same time, the architecture supports adaptation capabilities by relying on the remote reprogrammability of key application parameters. The proposed architecture has been successfully deployed in the Swiss-Italian Alps to monitor the collapse of rock faces in three geographical areas.
Cesare Alippi, Romolo Camplani, Cristian Galperti, Antonio Marullo, Manuel Roveri
ACM Trans. Sens. Networks1
2012 A "Learning from Models" Cognitive Fault Diagnosis System
Cesare Alippi, Manuel Roveri, Francesco Trovò
ICANN (2)1
2012 SVM-Based Just-in-Time Adaptive Classifiers
Cesare Alippi, Li Bu, Dongbin Zhao
ICONIP (2)1
2012 Just-in-time ensemble of classifiers
abstract
Handling dynamic environments and building up algorithms operating at low supervised-sample rates are two main challenges for classification systems designed to operate in real-life scenarios. Here, changes in the probability density function of classes characterizing the data-generating process (also called concept drift) should be detected as soon as possible to prevent the classifier from becoming obsolete. Moreover, when the rate of supervised samples during the operational life is low (as in those situations where the sample inspection is costly or destructive) both detecting the change and re-training the classifier become even more critical aspects. We present an adaptive classifier that exploits both supervised and unsupervised data to monitor the process stationarity. The classifier follows the just-in-time (JIT) approach and relies on two different change-detection tests (CDTs) to reveal changes in the environment and reconfigure the classifier accordingly. The proposed solution assesses the stationary in both the joint probability density function (CDT at the classification error) and the distribution of the inputs (CDT on unlabeled data). In addition, we integrate in the JIT adaptive classifier a procedure able to handle recurrent concepts within an ensemble of classifiers framework. Experiments show that monitoring unsupervised samples and handling recurrent concepts is essential for classifying in non-stationary environments when few supervised samples are available.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IJCNN1
2012 On-line reconstruction of missing data in sensor/actuator networks by exploiting temporal and spatial redundancy
abstract
Data streams from remote monitoring systems such as wireless sensor networks show immediately that the “you sample you get” statement is not always true. Not rarely, the data stream is interrupted by intermittent communication or sensors faults, resulting in missing data in the received sequence. This has a negative impact in many algorithms assuming continuous data stream; as such, the missing data must be suitably reconstructed, in order to guarantee continuous data availability. We suggest a general methodology for reconstructing missing data that exploits both temporal and spatial redundancy characterizing the phenomenon being monitored and the distributed system, a situation proper of many monitoring systems constituted by sensor and actuator networks. Temporal and spatial dependencies are learned through linear and non-linear non-parametric models, also encompassing neural -possibly recurrent- networks, which become the spatial transfer functions connecting the different views of the phenomenon under investigation. Missing data are finally reconstructed by exploiting the forecasting ability provided by such transfer functions. The experimental section shows the effectiveness of the proposed methodology.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IJCNN1
2012 An HMM-based change detection method for intelligent embedded sensors
abstract
In this work we address the problem of automatically detecting changes either induced by faults or concept drifts in data streams coming from multi-sensor units. The proposed methodology is based on the fact that the relationships among different sensor measurements follow a probabilistic pattern sequence when normal data, i.e. data which do not present a change, are observed. Differently, when a change in the process generating the data occurs the probabilistic pattern sequence is modified. The relationship between two generic data streams is modelled through a sequence of linear dynamic time-invariant models whose trained coefficients are used as features feeding a Hidden Markov Model (HMM) which, in turn, extracts the pattern structure. Change detection is achieved by thresholding the log-likelihood value associated with incoming new patterns, hence comparing the affinity between the structure of new acquisitions with that learned through the HMM. Experiments on both artificial and real data demonstrate the appreciable performance of the method both in terms of detection delay, false positive and false negative rates.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IJCNN1
2012 Netbrick: A high-performance, low-power hardware platform for wireless and hybrid sensor networks
abstract
The recent increase in number and complexity of wireless sensor networks (WSN)-based deployments made evident the limits of traditional hardware platforms designed to work in a controlled environment (laboratory) for a limited amount of time. The need for high-performance, low-power, flexible and scalable hardware platforms able to work in real-world (possibly harsh) environments led us to design and develop the NetBrick platform. The novelty and the advantages of the proposed platform w.r.t. other existing hardware platforms reside in: 1) flexibility at the board level (each module composing the board can be enabled/disabled by the software) 2) flexibility at the network level (the NetBrick natively allows for creating wireless, wired and hybrid networks); 3) the high performance guaranteed by the 32bit microprocessor at a very reasonable power consumption thanks to the ultra-low power Cortex M3 architecture; 4) the fine-grain energy management of the board modules (each module provides information about its power consumption), hence allowing the designer for defining advanced energy management policies. The NetBrick platform has been tested with success in a distributed monitoring system for landslide forecasting designed and developed by our group and deployed in the Alps (north Italy).
Cesare Alippi, Romolo Camplani, Manuel Roveri, Gabriele Viscardi
MASS1
2012 Data-driven optimal algorithms and their applications to pattern recognition
Huaguang Zhang, Cesare Alippi, Dongbin Zhao
Neurocomputing2
2012 A year of neural network research: Special Issue on the 2011 International Joint Conference on Neural Networks
Jean-Philippe Thivierge, Ali A. Minai, Hava T. Siegelmann, Cesare Alippi, Michael Georgiopoulos
Neural Networks4
2011 A Distributed Self-adaptive Nonparametric Change-Detection Test for Sensor/Actuator Networks
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
ICANN (2)1
2011 An effective just-in-time adaptive classifier for gradual concept drifts
abstract
Classification systems designed to work in nonstationary conditions rely on the ability to track the monitored process by detecting possible changes and adapting their knowledge-base accordingly. Adaptive classifiers present in the literature are effective in handling abrupt concept drifts (i.e., sudden variations), but, unfortunately, they are not able to adapt to gradual concept drifts (i.e., smooth variations) as these are, in the best case, detected as a sequence of abrupt concept drifts. To address this issue we introduce a novel adaptive classifier that is able to track and adapt its knowledge base to gradual concept drifts (modeled as polynomial trends in the expectations of the conditional probability density functions of input samples), while maintaining its effectiveness in dealing with abrupt ones. Experimental results show that the proposed classifier provides high classification accuracy both on synthetically generated datasets and measurements from real sensors.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IJCNN1
2011 A hierarchical, nonparametric, sequential change-detection test
abstract
Design of applications working in nonstationary environments requires the ability to detect and anticipate possible behavioral changes affecting the system under investigation. In this direction, the literature provides several tests aiming at assessing the stationarity of a data generating process; of particular interest are nonparametric sequential change-point detection tests that do not require any a-priori information regarding both process and change. Moreover, such tests can be made automatic through an on-line inspection of sequences of data, hence making them particularly interesting to address real applications. Following this approach, we suggest a novel two-level hierarchical change-detection test designed to detect possible occurrences of changes by observing incoming measurements. This hierarchical solution significantly reduces the number of false positives at the expenses of a negligible increase of false negatives and detection delays. Experiments show the effectiveness of the proposed approach both on synthetic dataset and measurements from real applications.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IJCNN1
2011 A just-in-time adaptive classification system based on the intersection of confidence intervals rule
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
Neural Networks1
2010 Adaptive Classifiers with ICI-Based Adaptive Knowledge Base Management
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
ICANN (2)1
2010 Change detection tests using the ICI rule
abstract
Designing tests able to effectively detect changes in the stationarity of a process generating data is a challenging problem, in particular when the process is unknown, and the only information available has to be extracted from a set of observations. This work proposes a novel approach for detecting changes in a process generating data whose distribution is unknown. Peculiarity of the approach is the use of the Intersection of Confidence Intervals (ICI) rule to monitor the process evolution. A change detection test derived from this approach is also presented. Experimental results show that the proposed test outperforms state-of-the art solutions, both in terms of efficiency and effectiveness, in particular when a reduced test configuration set is available.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IJCNN1
2010 Virtual k-fold cross validation: An effective method for accuracy assessment
abstract
LOO and k-fold cross validation are widely used validation methods assessing the accuracy of a model at the expenses of a high computational load (several models need to be trained and performance averaged). To mitigate such phenomenon a virtual LOO method has been suggested in which, by relying on the concept of leverages, provides the LOO estimate of the generalization error in a closed form without the need to re-training different models. In this paper, we extend and generalize such an approach by introducing the virtual k-fold cross validation method which provides a k-fold cross validation estimate without requiring training multiple models. Results, correct for linear models, are approximations for nonlinear ones. Simulation results show the effectiveness of the proposed virtual method which can be suitably extended to cover different figures of merit and performance assessment techniques.
Cesare Alippi, Manuel Roveri
IJCNN1
2010 An hybrid wireless-wired monitoring system for real-time rock collapse forecasting
abstract
Rock face collapses are one of the most dangerous and sudden natural risks in mountain environments. Traditional investigation techniques (e.g., strain gauge, inclinometer) cannot guarantee neither a non-invasive monitoring within the rocks nor a prompt alarm in case of possible collapse. Real-time monitoring systems, which might provide an effective evaluation of the dynamics of the phenomenon for a subsequent forecasting phase, generally rely on wired solutions that are unfeasible in environmental monitoring applications. In this paper, we present a real-time monitoring system for the rock collapse forecasting that exploits MEMS accelerometers and geophones (in addition to traditional sensors) for a non-invasive detection of micro-acoustic bursts associated with the formation and the evolution of the cracks within the rocks. The proposed monitoring system relies on an hybrid wireless-wired architecture that allows for detecting and localizing micro-acoustic emissions in real-time yet maintaining an high energy-efficiency by means of effective energy management policies and sophisticated energy harvesting mechanisms. The deployment area is the St. Martin mountain that dangerously insists on the town of Lecco (Italy).
Cesare Alippi, Romolo Camplani, Cristian Galperti, Antonio Marullo, Manuel Roveri
MASS1
2009 Just in time classifiers: Managing the slow drift case
abstract
A classifier expected to work in a non-stationary environment has to: (i) detect changes in the process generating the data; (ii) suitably react to the change by adapting to the new working condition. Just-in-time adaptive classifiers, a classification structure addressing stationary and nonstationary conditions, have been presented to the computational intelligence community. Such classifiers require a temporal detection of a (possible) process deviation followed by an adaptive management of the knowledge base characterizing the classifier to cope with the process change. This paper improves just-in-time adaptive classifiers by integrating temporal information about the state of the process under monitoring. An index for the process deviation is defined which, coupled with an adaptive weighted k-NN classifier, shows to be particularly effective in dealing with smooth process drifts and ageing phenomena.
Cesare Alippi, Giacomo Boracchi, Manuel Roveri
IJCNN1
2008 k-NN classifiers: Investigating the k=k(n) relationship
abstract
The paper proposes a theory-based method for estimating the optimal value of k in k-NN classifiers based on a n-sized training set. As expected, experiments show that the suggested k is such that k/n rarr 0 when both k and n tend to infinity, as required by the asymptotical consistency condition. Interestingly, it appears that the generalization error is robust w.r.t. to k when n becomes large (probably as a consequence of the k/n rarr 0 relationship); the immediate consequence is that there is no need to provide an accurate estimate for the optimal k and an approximated coarser value, e.g., provided with cross validation, 1-fold cross validation or leave one out is more than adequate.
Cesare Alippi, M. Fuhrman, Manuel Roveri
IJCNN1
2008 Just-in-Time Adaptive Classifiers - Part I: Detecting Nonstationary Changes
abstract
The stationarity requirement for the process generating the data is a common assumption in classifiers' design. When such hypothesis does not hold, e.g., in applications affected by aging effects, drifts, deviations, and faults, classifiers must react just in time, i.e., exactly when needed, to track the process evolution. The first step in designing effective just-in-time classifiers requires detection of the temporal instant associated with the process change, and the second one needs an update of the knowledge base used by the classification system to track the process evolution. This paper addresses the change detection aspect leaving the design of just-in-time adaptive classification systems to a companion paper. Two completely automatic tests for detecting nonstationarity phenomena are suggested, which neither require a priori information nor assumptions about the process generating the data. In particular, an effective computational intelligence-inspired test is provided to deal with multidimensional situations, a scenario where traditional change detection methods are generally not applicable or scarcely effective.
Cesare Alippi, Manuel Roveri
IEEE Trans. Neural Networks1
2008 Just-in-Time Adaptive Classifiers - Part II: Designing the Classifier
abstract
Aging effects, environmental changes, thermal drifts, and soft and hard faults affect physical systems by changing their nature and behavior over time. To cope with a process evolution adaptive solutions must be envisaged to track its dynamics; in this direction, adaptive classifiers are generally designed by assuming the stationary hypothesis for the process generating the data with very few results addressing nonstationary environments. This paper proposes a methodology based on k-nearest neighbor (NN) classifiers for designing adaptive classification systems able to react to changing conditions just-in-time (JIT), i.e., exactly when it is needed. k-NN classifiers have been selected for their computational-free training phase, the possibility to easily estimate the model complexity k and keep under control the computational complexity of the classifier through suitable data reduction mechanisms. A JIT classifier requires a temporal detection of a (possible) process deviation (aspect tackled in a companion paper) followed by an adaptive management of the knowledge base (KB) of the classifier to cope with the process change. The novelty of the proposed approach resides in the general framework supporting the real-time update of the KB of the classification system in response to novel information coming from the process both in stationary conditions (accuracy improvement) and in nonstationary ones (process tracking) and in providing a suitable estimate of k. It is shown that the classification system grants consistency once the change targets the process generating the data in a new stationary state, as it is the case in many real applications.
Cesare Alippi, Manuel Roveri
IEEE Trans. Neural Networks1
2007 Adaptive Classifiers in Stationary Conditions
abstract
Integrating new information in classification systems during their operational life requires adaptive mechanisms able to identify first the presence of valuable information and update then the knowledge base onto which the classifier is configured. In this paper we provide a design solution for adaptive classifiers operating in stationary environments; information provided (whenever available by a supervisor over time) is used to improve the performance of the classification system hence mimicking the asymptotical behavior suggested by the theory. The adaptive classifier relies on k -NNs, here chosen for their learning-free modality (hence easily supporting a real time adaptation mechanism); a novel method is proposed for matching the optimal k (measuring the complexity of the classifier) with the incremental knowledge acquired over time. A large experimental campaign shows the effectiveness of the proposed approach.
Cesare Alippi, Manuel Roveri
IJCNN1
2007 Just-in-time Adaptive Classifiers in Non-Stationary Conditions
abstract
In real world applications ageing effects, process drifts, soft and hard faults may affect the data generation mechanism and, as a consequence, data coming from it. Intelligent measurement systems developed for such processes (e.g., industrial quality assessment and control, environmental monitoring) require adaptive techniques which, by tracking the system evolution, allow the intelligent system for keeping acceptable performance. Here we focus on adaptive classifiers embedded in intelligent measurement systems designed to cope with non-stationary environments, yet well performing in stationary conditions. The novelty of the approach resides in the possibility to update in a just-in-time fashion, i.e., only when it is really needed, the knowledge base of the classifier. A large experimental campaign shows the effectiveness of the proposed design.
Cesare Alippi, Manuel Roveri
IJCNN1
2007 Adaptive Sampling for Energy Conservation in Wireless Sensor Networks for Snow Monitoring Applications
abstract
Energy conservation techniques for sensor networks typically rely on the assumption that data sensing and processing consume considerable less energy than communication. This assumption does not hold in some practical application scenarios, where ad hoc developed sensor units require power consumption comparable with, or even larger than, that of the radio. In this paper we focus on an embedded sensor for monitoring snow composition in mountain slopes for avalanche forecasting. To lower the sensor energy consumption we propose an adaptive sampling algorithm able to dynamically estimate the optimal sampling frequency of the signal to be monitored. In turn, this minimizes the activity of both the sensor and the radio (hence saving energy) while maintaining an acceptable accuracy on the acquired data. Simulation experiments show that the suggested solution can save up to 97% of the energy consumed for sensing when the sensor is always on, while maintaining the error at acceptable levels.
Cesare Alippi, Giuseppe Anastasi, Cristian Galperti, Francesca Mancini, Manuel Roveri
MASS1
2006 Particle Filters for Rss-Based Localization in Wireless Sensor Networks: An Experimental Study
abstract
This paper focuses on the development of a radio localization technique for a wireless sensor network infrastructure where a large number of simple power-aware nodes are spread in indoor environments. Fixed and moving nodes exchange radio messages but can only measure mutual power figures such as the received signal strength (RSS) indicator. Local maximum likelihood estimation from propagation models suffers from false alarm problems due to incorrect position information, complex indoor propagation effects and simple hardware radio architectures. Here, we propose a Bayesian approach to estimate and track the position of a moving node from power maps obtained through field measurements. To lower the computational power required by grid-based algorithms, we exploit particle filter techniques that implement an irregular sampling of the a-posteriori probability space. Finally, experimental results are presented and discussed
Carlo Morelli, Monica Nicoli, Vittorio Rampa, Umberto Spagnolini, Cesare Alippi
ICASSP (4)5
2006 A computational intelligence-based criterion to detect non-stationarity trends
abstract
The stationarity hypothesis is largely and implicitly assumed when designing classifiers (especially those for industrial applications) but it does not generally hold in practice. The paper goal is to provide an automatic, general purpose, easy to use and effective index for estimating deviations, drifts or ageing effects in the process generating the data (e.g., classifier inputs); in turns this will allow the designer for identifying when to intervene to update the knowledge space of adaptive classifiers. More specifically, we suggest a robust extension of the adaptive CUSUM test procedure which addresses a set of features (in contrast to the literature which considers a single feature) for detecting drifts. The application of the change detection test to real applications shows that its real additional value resides in the ability to detect continuous and small drifts, a critical situation for traditional tests.
Cesare Alippi, Manuel Roveri
IJCNN1
2006 A statistical approach to localize passive RFIDs
abstract
We suggest a novel method for localizing passive radio frequency identification (RFID) tags based on a Bayesian approach. The localization framework requires a suitable number of tag-readers deployed in the area under investigation and connected to antennas which, provided with rotation ability, perform the detection task by scanning each angular sector at different power levels. Since readers are programmed to detect tags at increasing transmitting power (and provide binary tag/no tag information), we take advantage of the reader interrogation results to develop a Bayesian model for tags localization. By testing the solution with off-the-shelf UHF tags were experimented the accuracy of the method that, in a 5mtimes4m environment, provides an average error in localization of about 0.6 meters. This without exploiting any a priori information of the location, orientation or the power delivered to the tag (since not available in actual tags) which would significantly improve localization accuracy
Cesare Alippi, D. Cogliati, Giovanni Vanini
ISCAS1
2006 An adaptive maximum power point tracker for maximising solar cell efficiency in wireless sensor nodes
abstract
The success of wireless sensor networks is somehow related to the energy supplies which, generally provided by batteries is a finite resource. Energy scavenge must be pursuit to grant long time operation and solar energy is surely the most effective one due to its ubiquitous distribution and relatively high power density. In this paper we propose a maximum power point tracker circuit (MPPT) for wireless nodes, i.e., a power converting circuit for optimally transferring solar energy to rechargeable batteries; efficiency is granted by an adaptive algorithm which keeps the MPPT in its optimal working point by tracking solar radiation so as to maximise the energy generation and transfer flow
Cesare Alippi, Cristian Galperti
ISCAS1
2006 An adaptive CUSUM-based test for signal change detection
abstract
Many applications, e.g., fault detection, quality of industrial process, monitoring and prediction of climatic phenomena assume the stationary hypothesis or require identification of the process change. Change detection tests satisfy the change detection necessity by identifying a drift, a different expected behavior, a deviation; their effectiveness is generally based on statistical confidence tests whose parameters are configured at design-time (generally through a trial-and-error approach). Here, we suggest an extension of the widely-used CUSUM change detection test which improves effectiveness and timeliness in detecting changes by adaptively configuring its test parameters
Cesare Alippi, Manuel Roveri
ISCAS1
2006 Application-based routing optimization in static/semi-staticWireless Sensor Networks
abstract
This paper presents a power aware routing algorithm for WSN exploiting application features and the radio frequency map of the deployed sensors. The suggested methodology can be applied to applications requiring data acquisition from semi-static medium-size networks. The improvement over traditional routing solutions derives from the fact that routing paths and optimal transmission powers are identified through an accurate simulation of the target application (not only at the routing layer) which uses real information such as the power response of the links. The methodology, tested on MICA2 wireless sensor networks, provided a reduction of power consumption of about 40% over the "multihoproute" routing algorithm considered in the TinyOS community
Cesare Alippi, Giovanni Vanini
PerCom1
2006 Exploiting application locality to design low-complexity, highly performing, and power-aware embedded classifiers
abstract
Temporal and spatial locality of the inputs, i.e., the property allowing a classifier to receive the same samples over time--or samples belonging to a neighborhood--with high probability, can be translated into the design of embedded classifiers. The outcome is a computational complexity and power aware design particularly suitable for implementation. A classifier based on the gated-parallel family has been found particularly suitable for exploiting locality properties: Subclassifiers are generally small, independent each other, and controlled by a master-enabling module granting that only a subclassifier is active at a time, the others being switched off. By exploiting locality properties we obtain classifiers with accuracy comparable with the ones designed without integrating locality but gaining a significant reduction in computational complexity and power consumption.
Cesare Alippi, Fabio Scotti
IEEE Trans. Neural Networks1
2006 Classification methods and inductive learning rules: what we may learn from theory
abstract
Inductive learning methods allow the system designer to infer a model of the relevant phenomena of an unknown process by extracting information from experimental data. A wide range of inductive learning methods is nowadays available, potentially ensuring different levels of accuracy on different problem domains. In this critical review of theoretic results gained in the last decade, we address the problem of designing an inductive classification system with optimal accuracy when domain knowledge is limited and the number of available experiments is-possibly-small. By analyzing the formal properties of consistent learning methods and of accuracy estimators, we wish to convey to the reader the message that the common practice of aggressively pursuing error minimization with different training algorithms and classification families is unjustified
Cesare Alippi, Pietro Braione
IEEE Trans. Syst. Man Cybern. Syst.1
2005 An adaptive system for automatic invoice-documents classification
abstract
The large amount of documents to be daily managed in modern offices requires development of automatic document classification tools aiming at (semi)automatically classifying the office documents into semantically similar classes. This paper presents an automatic invoice-documents classification system based on the analysis of the graphical information present in the document and able to perform both closed (the number of classes is fixed) and open world (the number of classes increases during operational life) classification. Invoice-documents of real companies prove that the classification system achieves a 99% correct classification in closed world and 79% in the open world case.
Cesare Alippi, F. Pessina, Manuel Roveri
ICIP (2)1
2004 A training-time analysis of robustness in feed-forward neural networks
abstract
The paper addresses the analysis of robustness over training time issue. Robustness is evaluated in the large, without assuming the small perturbation hypothesis, by means of randomised algorithms. We discovered that robustness is a strict property of the model -as it is accuracy- and, hence, it depends on the particular neural network family, application, training algorithm and training starting point. Complex neural networks are hence not necessarily more robust than less complex topologies. An early stopping algorithm is finally suggested which extends the one based on the test set inspection with robustness aspects.
Cesare Alippi, Daniele Sana, Fabio Scotti
IJCNN1
2003 An application-level synthesis methodology for multidimensional embedded processing systems
abstract
The implementation of multidimensional systems in embedded devices is a major design challenge due to the high algorithmic complexity of the applications. The authors suggest a novel application-level synthesis methodology for those parts of the embedded application which are characterized by being Lebesgue measurable (the computation involved in signal and image processing systems is Lebesgue measurable). The synthesis methodology, based on perturbation analysis, supports the design of analog, digital, or mixed implementations at the very high level of the system design cycle. The outputs of the methodology are quantitative indications regarding the maximum performance loss tolerable by the subsystems composing the application. Such information, augmented with a stochastic description of the tolerated perturbations, can be related to lower synthesis levels and guide the designer toward the final implementation of the embedded device. The perturbation analysis is based on randomized algorithms for an effective evaluation of the performance loss of the computational flow once affected by behavioral perturbations and a Tabu-search-inspired optimizing algorithm for distributing the tolerable performance loss at the system output along the computational subsystems composing the possibly multidimensional processing.
Cesare Alippi, Andrea Galbusera, Marco Stellini
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 A neural-network based control solution to air-fuel ratio control for automotive fuel-injection systems
abstract
Maximization of the catalyst efficiency in automotive fuel-injection engines requires the design of accurate control systems to keep the air-to-fuel ratio at the optimal stoichiometric value AF/sub S/. Unfortunately, this task is complex since the air-to-fuel ratio is very sensitive to small perturbations of the engine parameters. Some mechanisms ruling the engine and the combustion process are in fact unknown and/or show hard nonlinearities. These difficulties limit the effectiveness of traditional control approaches. In this paper, we suggest a neural based solution to the air-to-fuel ratio control in fuel injection systems. An indirect control approach has been considered which requires a preliminary modeling of the engine dynamics. The model for the engine and the final controller are based on recurrent neural networks with external feedbacks. Requirements for feasible control actions and the static precision of control have been integrated in the controller design to guide learning toward an effective control solution.
Cesare Alippi, Cosimo de Russis, Vincenzo Piuri
IEEE Trans. Syst. Man Cybern. Part C1
2002 Randomized Algorithms: A System-Level, Poly-Time Analysis of Robust Computation
abstract
Provides a methodology for analyzing the performance degradation of a computation once it has been affected by perturbations. The suggested methodology, by relaxing all assumptions made in the related literature, provides design guidelines for the subsequent implementation of complex computations in physical devices. Implementation issues, such as finite precision representation, fluctuations of the production parameters and aging effects, can be studied directly at the system level, independent of any technological aspect and quantization technique. Only the behavioral description of the computational flow, which is assumed to be Lebesgue-measurable, and the architecture to be investigated are needed. The suggested analysis is based on the theory of randomized algorithms, which transforms the computationally intractable problem of robustness investigation into a polynomial-time algorithm by resorting to probability.
Cesare Alippi
IEEE Trans. Computers1
2002 A probably approximately correct framework to estimate performancedegradation in embedded systems
abstract
Future design environments for embedded systems will require the development of sophisticated computer-aided design tools for compiling the high-level specifications of an application down to a final low-level language describing the embedded solution. This requires abstraction of technology-dependent aspects and requirements into behavioral entities. The paper takes a first step in this direction by introducing a high-level methodology for estimating the performance degradation of an application affected by perturbations; a special emphasis is given to accuracy performance. To grant generality it is uniquely assumed that the performance degradation function and the mathematical formulation describing the application are Lebesgue measurable. Perturbations affecting the application abstract details related to physical sources of uncertainties such as finite precision representation, faults, fluctuations of physical parameters, battery power variations, and aging effects whose impact on the computation can be treated within a high-level homogenous framework. A novel stochastic theory based on randomization is suggested to quantify the approximated nature of the perturbed environment. The outcomes are two algorithms which estimate in polynomial time the performance degradation of the application once affected by perturbations. Such information can then be exploited by HW/SW codesign methodologies to guide the subsequent partitioning between HW and SW, analog versus digital, fixed versus floating point, or used to validate architectural choices before any low-level design step takes place. The proposed method is finally applied to real designs involving neural and wavelet-based applications.
Cesare Alippi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2000 Determining the optimum extended instruction-set architecture for application specific reconfigurable VLIW CPUs (poster abstract)
abstract
No abstract available.
Cesare Alippi, William Fornaciari, Laura Pozzi 0001, Mariagiovanna Sami
FPGA1
2000 A Methodology for Example-Based Specification and Design
abstract
There is an ever-increasing use of embedded systems; fast prototyping, time to market and severe implementation constraints must be faced to provide an effective - low cost - solution for a given application. To this end, several algorithmic formalisms are available to describe and validate complex systems at a behavioural level in order to minimise development costs and facilitate the integration of design and implementation constraints. Unfortunately, a soft-computing paradigm cannot be directly manipulated by conventional development environments for embedded systems unless an algorithmic description is available. In general, such a description is the result of a training procedure which, by following the selection of the most suitable soft-computing paradigm, configures it. The paper addresses the issues related to the integration of soft-computing paradigms within conventional development environments for embedded systems. The analysis is carried out at a behavioural abstraction level.
Cesare Alippi, Stefano Ferrari, Vincenzo Piuri
IJCNN (3)1
1999 A DAG-Based Design Approach for Reconfigurable VLIW Processors
Cesare Alippi, William Fornaciari, Laura Pozzi 0001, Mariagiovanna Sami
DATE1
1998 Artificial neural networks
Vincenzo Piuri, Cesare Alippi
J. Syst. Archit.2
1998 Accuracy vs. Precision in Digital VLSI Architectures for Signal Processing
abstract
The paper provides a sensitivity analysis to measure the loss in accuracy induced by perturbations affecting acyclic computational flows composed of linear convolutions and nonlinear functions. We do not assume a large number of coefficients or input independence for the convolution module, nor strict requirements on the nonlinear function. The analysis is tailored to digital VLSI implementations where perturbations, associated with data quantization, affect the device inputs, coefficients, internal values, and outputs. The sensitivity analysis can be used to measure the loss in accuracy along the computational chain, to characterize the tolerated perturbations, and to dimension the whole architecture.
Cesare Alippi, Luciano Briozzo
IEEE Trans. Computers1
1998 Testability analysis and behavioral testing of the Hopfield neural paradigm
abstract
Testability analysis and test pattern generation for neural architectures can be performed at a very high abstraction level on the computational paradigm. In this paper, we consider the case of Hopfield's networks, as the simplest example of networks with feedback loops. A behavioral error model based on finite-state machines (FSM's) is introduced. Conditions for controllability, observability and global testability are derived to verify errors excitation and propagation to outputs. The proposed behavioral test pattern generator creates the minimum length test sequence for any digital implementation.
Cesare Alippi, Franco Fummi, Vincenzo Piuri, Mariagiovanna Sami, Donatella Sciuto
IEEE Trans. Very Large Scale Integr. Syst.1
1995 Off-Line Performance Maximisation in Feed-Forward Neural Networks by Applying Virtual Neurons and Covariance Transformations
abstract
Optimisation of a feed-forward neural paradigm for a given application involves problems such as maximisation of the generalisation ability (relevant to provide effectiveness) and structure minimisation (allowing for physical realisability by using dedicated VLSI devices). This paper proposes a contemporaneous solution of these conflicting goals. The globally-optimised structure is identified by using a covariance matrix transformation and layers of virtual neurons.
Cesare Alippi, Raffaele Petracca, Vincenzo Piuri
ISCAS1
1995 Real-time analysis of ships in radar images with neural networks
Cesare Alippi
Pattern Recognit.1
1994 Setting and Validating Precision Requirements in the Digital VLSI Implementation of a Neural Defect-Identifier for Machined Objects
abstract
In this paper we deal with the functional design of a dedicated digital VLSI neurochip for a real time demanding application: the identification of defects in machined parts of mechanical objects. Precision requirements analysis (in terms of the bits number to represent neural values) represents a critical point to be faced when choosing the final architecture since requirements on the chip's size may prevent parallelism exploitation. Application features and neural dynamics are thus carefully analysed to determine quasi-minimal bit requirements for the architectural components. Afterwards, once the hardware design has been defined, sensitivity analysis tools need to be applied to validate effective performances. It is shown that an accurate choice of precision requirements for hardware elements, despite poor signal to noise ratios, leads to a suitable architecture which solves the application and makes feasible the VLSI design.>
Cesare Alippi, Luciano Briozzo
ISCAS1
1994 Sensitivity to Errors in Artificial Neural Networks: a Behavioural Approach
abstract
A behavioral approach to the impact of errors due to faults in neural computation is analyzed. Starting from a geometrical description of errors affecting neural values, we derive the probability of error detection at the neuron's output and at the network's outputs.>
Cesare Alippi, Vincenzo Piuri, Mariagiovanna Sami
ISCAS1
1992 Galatea neural VLSI architectures: Communication and control considerations
Cesare Alippi, Marley M. B. R. Vellasco
Microprocess. Microprogramming1
1991 The determination of angular values and parameters in flat surfaces: from the mathematical approach to the CORDIC architecture
Cesare Alippi
Microprocessing and Microprogramming1
1991 Iiistological image understanding by error backpropagation
Apostolos Nikolaos Refenes, Cesare Alippi
Microprocessing and Microprogramming2