Fabio Fassetti

dblp:90/3977 · DBLP profile ↗
← Back
50ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-8416-906XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Reconstruction error-based anomaly detection with few outlying examples
abstract
Reconstruction error-based neural architectures constitute a classical deep learning approach to anomaly detection which has shown great performances. It consists in training an Autoencoder to reconstruct a set of examples deemed to represent the normality and then to point out as anomalies those data that show a sufficiently large reconstruction error. Unfortunately, these architectures often become able to well reconstruct also the anomalies in the data. This phenomenon is more evident when there are anomalies in the training set. In particular, when these anomalies are labeled, a setting called semi-supervised, the best way to train Autoencoders is to ignore anomalies and minimize the reconstruction error on normal data. When a sufficiently large and representative set of anomalous examples is available, the problem essentially shifts toward a classification task, where standard supervised strategies can be applied effectively. In this work, instead, we focus on the more challenging scenario in which only a limited number of anomalous examples is available, and these examples are not sufficiently representative of the wide variability that anomalies may exhibit. We propose , a novel reconstruction error-based architecture that explicitly leverages labeled anomalies to guide the model. Our method introduces a new loss formulation that forces anomalies to be reconstructed according to a transformation function, effectively pushing them outside the description of normal data. This strategy increases the separation between the reconstruction errors of normal and anomalous samples, thereby improving the detection of both seen and unseen anomalies. Extensive experiments demonstrate that consistently outperforms both standard Autoencoders and the most competitive deep learning techniques for semi-supervised anomaly detection, achieving state-of-the-art results. In particular, our method proves superior across a diverse set of benchmarks, including vectorial data, high-dimensional datasets, and image domains. Moreover, maintains its advantage even in challenging scenarios where the training data are polluted by anomalies that are incorrectly labeled as normal, further highlighting its robustness and practical applicability.
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina
Neurocomputing2
2025 Explaining Anomalous Data with Reinforcement Learning
Simone Amirato, Fabrizio Angiulli, Fabio Fassetti
DS3
2025 CF-OOD: Concentration-Free Density Estimation for Reliable Out-of-Distribution Detection
Fabrizio Angiulli, Fabio Fassetti, Maria Pia Zupi
SISAP2
2025 LLiMe: enhancing text classifier explanations with large language models
abstract
Abstract The widespread diffusion of text black-box classifiers necessitates explainable AI (XAI) techniques for this domain. A seminal XAI technique is Local Interpretable Model-agnostic Explanations (LIME). For text classification, LIME maps an input sentence and its neighbours into a bag of words, using a linear regressor as an interpretable model. However, this strategy has significant limitations. Neighbouring sentences are constructed solely by extracting subsets of the input sentence, which may fail to accurately capture the local decision boundary. Moreover, these subsets are not guaranteed to be representative of the classification classes, potentially leading to unbalanced or misleading interpretability. Additionally, such generated sentences might lack semantic coherence. Furthermore, the resulting explanation is often limited to confirming the relevance of a term or highlighting the impact of its removal, without providing deeper insights. This work tries to overcome these limitations by proposing LLiMean extension of LIME that exploits advances in Large Language Models (LLMs) to perform a classifier-driven generation of the neighbourhood. Our approach allows neighbours to employ a vocabulary larger than that of the input text. A generation procedure is introduced to more effectively capture the local decision boundary by ensuring generated samples span all classes involved in the classification. Additionally, an LLM-driven explanation and a counterfactual generation procedure are presented, returning the most relevant set of editing operations to influence the black-box predictor’s decision. Thus, the approach provides a richer, easier-to-interpret explanation and high-quality counterfactuals compared to standard LIME. Experiments on real datasets witness the technique’s effectiveness in providing suitable, relevant, and interpretable explanations.
Fabrizio Angiulli, Francesco De Luca, Fabio Fassetti, Simona Nisticò
Mach. Learn.3
2025 Improving local interpretable classifier explanations exploiting self-generated semantic features
abstract
Abstract Explaining predictions of classifiers is a fundamental problem in eXplainable Artificial Intelligence (XAI). LIME (for Local Interpretable Model-agnostic Explanations) is a popular XAI technique able to explain any classifier by providing an interpretable model which approximates the black box locally to the instance under consideration. In order to build interpretable local models, LIME requires the user to explicitly define a space of interpretable components, also called artefacts, associated with the input instance. To reconstruct local black-box behaviour, the instance neighbourhood is explored by generating instance neighbours as random subsets of the provided artefacts. In this work, we note that the above-depicted strategy has a limitation given by the fact that the local explanation is limited to be expressed only in terms of object artefacts. To overcome this limitation, in this work we propose $$\mathcal {S}{\text {-LIME}}$$ S -LIME , a variant of the basic LIME method exploiting unsupervised learning to replace object artefacts with self-generated semantic features in neighbourhood generation. This characteristic enables our approach to sample instance neighbours in a more semantic-driven fashion and greatly reduces the bias associated with explanations. We demonstrate the applicability and effectiveness of our proposal in the text classification domain. We also present a further extension for textual data in which word groups are used to obtain richer explanations. Comparison with the baseline highlights the superior quality of the explanations obtained by adopting our strategy.
Fabrizio Angiulli, Fabio Fassetti, Simona Nisticò
Neural Comput. Appl.2
2024 Large Language Models-Based Local Explanations of Text Classifiers
Fabrizio Angiulli, Francesco De Luca, Fabio Fassetti, Simona Nisticò
DS (1)3
2024 Indecision-Aware Deep Active Anomaly Detection
Simone Amirato, Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina
IDEAL (2)3
2024 Enhancing anomaly detectors with LatentOut
abstract
Abstract $${{\textbf{Latent}}\varvec{Out}}$$ Latent Out is a recently introduced algorithm for unsupervised anomaly detection which enhances latent space-based neural methods, namely (Variational) Autoencoders, GANomaly and ANOGan architectures. The main idea behind it is to exploit both the latent space and the baseline score of these architectures in order to provide a refined anomaly score performing density estimation in the augmented latent-space/baseline-score feature space. In this paper we investigate the performance of $${{\textbf{Latent}}\varvec{Out}}$$ Latent Out acting as a one-class classifier and we experiment the combination of $${{\textbf{Latent}}\varvec{Out}}$$ Latent Out with GAAL architectures, a novel type of Generative Adversarial Networks for unsupervised anomaly detection. Moreover, we show that the feature space induced by $${{\textbf{Latent}}\varvec{Out}}$$ Latent Out has the characteristic to enhance the separation between normal and anomalous data. Indeed, we prove that standard data mining outlier detection methods perform better when applied on this novel augmented latent space rather than on the original data space.
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina
J. Intell. Inf. Syst.2
2024 Explaining outliers and anomalous groups via subspace density contrastive loss
abstract
Abstract Explainable AI refers to techniques by which the reasons underlying decisions taken by intelligent artifacts are single out and provided to users. Outlier detection is the task of individuating anomalous objects within a given data population they belong to. In this paper we propose a new technique to explain why a given data object has been singled out as anomalous. The explanation our technique returns also includes counterfactuals, each of which denotes a possible way to “repair” the outlier to make it an inlier. Thus, given in input a reference data population and an object deemed to be anomalous, the aim is to provide possible explanations for the anomaly of the input object, where an explanation consists of a subset of the features, called choice, and an associated set of changes to be applied, called mask, in order to make the object “behave normally”. The paper presents a deep learning architecture exploiting a features choice module and mask generation module in order to learn both components of explanations. The learning procedure is guided by an ad-hoc loss function that simultaneously maximizes (minimizes, resp.) the isolation of the input outlier before applying the mask (resp., after the application of the mask returned by the mask generation module) within the subspace singled out by the features choice module, all that while also minimizing the number of features involved in the selected choice. We consider also the case in which a common explanation is required for a group of outliers provided together in input. We present experiments on both artificial and real data sets and a comparison with competitors validating the effectiveness of the proposed approach.
Fabrizio Angiulli, Fabio Fassetti, Simona Nisticò, Luigi Palopoli 0001
Mach. Learn.2
2023 Counterfactuals Explanations for Outliers via Subspaces Density Contrastive Loss
Fabrizio Angiulli, Fabio Fassetti, Simona Nisticò, Luigi Palopoli 0001
DS2
2023 Discriminative pattern discovery for the characterization of different network populations
abstract
MOTIVATION: An interesting problem is to study how gene co-expression varies in two different populations, associated with healthy and unhealthy individuals, respectively. To this aim, two important aspects should be taken into account: (i) in some cases, pairs/groups of genes show collaborative attitudes, emerging in the study of disorders and diseases; (ii) information coming from each single individual may be crucial to capture specific details, at the basis of complex cellular mechanisms; therefore, it is important avoiding to miss potentially powerful information, associated with the single samples. RESULTS: Here, a novel approach is proposed, such that two different input populations are considered, and represented by two datasets of edge-labeled graphs. Each graph is associated to an individual, and the edge label is the co-expression value between the two genes associated to the nodes. Discriminative patterns among graphs belonging to different sample sets are searched for, based on a statistical notion of 'relevance' able to take into account important local similarities, and also collaborative effects, involving the co-expression among multiple genes. Four different gene expression datasets have been analyzed by the proposed approach, each associated to a different disease. An extensive set of experiments show that the extracted patterns significantly characterize important differences between healthy and unhealthy samples, both in the cooperation and in the biological functionality of the involved genes/proteins. Moreover, the provided analysis confirms some results already presented in the literature on genes with a central role for the considered diseases, still allowing to identify novel and useful insights on this aspect. AVAILABILITY AND IMPLEMENTATION: The algorithm has been implemented using the Java programming language. The data underlying this article and the code are available at https://github.com/CriSe92/DiscriminativeSubgraphDiscovery.
Fabio Fassetti, Simona E. Rombo, Cristina Serrao
Bioinform.1
2023 Anomaly detection with correlation laws
Fabrizio Angiulli, Fabio Fassetti, Cristina Serrao
Data Knowl. Eng.2
2023 ${{\mathrm {Latent}}Out}$: an unsupervised deep anomaly detection approach exploiting latent space distribution
abstract
Abstract Anomaly detection methods exploiting autoencoders (AE) have shown good performances. Unfortunately, deep non-linear architectures are able to perform high dimensionality reduction while keeping reconstruction error low, thus worsening outlier detecting performances of AEs. To alleviate the above problem, recently some authors have proposed to exploit Variational autoencoders (VAE) and bidirectional Generative Adversarial Networks (GAN), which arise as a variant of standard AEs designed for generative purposes, both enforcing the organization of the latent space guaranteeing continuity. However, these architectures share with standard AEs the problem that they generalize so well that they can also well reconstruct anomalies. In this work we argue that the approach of selecting the worst reconstructed examples as anomalies is too simplistic if a continuous latent space autoencoder-based architecture is employed. We show that outliers tend to lie in the sparsest regions of the combined latent/error space and propose the $$\mathrm{VAE}Out$$ VAEOut and $${{\mathrm {Latent}}Out}$$ LatentOut unsupervised anomaly detection algorithms, identifying outliers by performing density estimation in this augmented feature space. The proposed approach shows sensible improvements in terms of detection performances over the standard approach based on the reconstruction error.
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina
Mach. Learn.2
2022 Outlier Explanation Through Masking Models
Fabrizio Angiulli, Fabio Fassetti, Simona Nisticò, Luigi Palopoli 0001
ADBIS2
2022 Cooperative Deep Unsupervised Anomaly Detection
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina, Rosaria Spada
DS2
2022 Detecting Anomalies with rmLatentOut: Novel Scores, Architectures, and Settings
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina
ISMIS2
2022 A Semi-automatic Data Generator for Query Answering
Fabrizio Angiulli, Alessandra Del Prete, Fabio Fassetti, Simona Nisticò
ISMIS3
2022 Graph-based construction of minimal models
Fabrizio Angiulli, Rachel Ben-Eliyahu-Zohary, Fabio Fassetti, Luigi Palopoli 0001
Artif. Intell.3
2022 A density estimation approach for detecting and explaining exceptional values in categorical data
abstract
Abstract In this work we deal with the problem of detecting and explaining anomalous values in categorical datasets. We take the perspective of perceiving an attribute value as anomalous if its frequency is exceptional within the overall distribution of frequencies. As a first main contribution, we provide the notion offrequency occurrence. This measure can be thought of as a form of Kernel Density Estimation applied to the domain of frequency values. As a second contribution, we define anoutliernessmeasure for categorical values that leverages the cumulated frequency distribution of the frequency occurrence distribution. This measure is able to identify two kinds of anomalies, calledlower outliersandupper outliers, corresponding to exceptionally low or high frequent values. Moreover, we provide interpretableexplanationsfor anomalous data values. We point out that providing interpretable explanations for the knowledge mined is a desirable feature of any knowledge discovery technique, though most of the traditional outlier detection methods do not provide explanations. Considering that when dealing with explanations the user could be overwhelmed by a huge amount of redundant information, as a third main contribution, we define a mechanism that allows us to single outoutstanding explanations. The proposed technique isknowledge-centric, since we focus on explanation-property pairs and anomalous objects are a by-product of the mined knowledge. This clearly differentiates the proposed approach from traditional outlier detection approaches which instead areobject-centric. The experiments highlight that the method is scalable and also able to identify anomalies of a different nature from those detected by traditional techniques.
Fabrizio Angiulli, Fabio Fassetti, Luigi Palopoli 0001, Cristina Serrao
Appl. Intell.2
2021 ODCA: An Outlier Detection Approach to Deal with Correlated Attributes
Fabrizio Angiulli, Fabio Fassetti, Cristina Serrao
DaWaK2
2021 A Stochastic Block Model Based Approach to Detect Outliers in Networks
Fabrizio Angiulli, Fabio Fassetti, Cristina Serrao
DEXA (1)2
2021 Local Interpretable Classifier Explanations with Self-generated Semantic Features
Fabrizio Angiulli, Fabio Fassetti, Simona Nisticò
DS2
2021 Meta-feature Extraction Strategies for Active Anomaly Detection
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina, Prospero Papaleo
IDEAL2
2021 Finding Local Explanations Through Masking Models
Fabrizio Angiulli, Fabio Fassetti, Simona Nisticò
IDEAL2
2021 Uncertain distance-based outlier detection with arbitrarily shaped data objects
abstract
Abstract Enabling information systems to face anomalies in the presence of uncertainty is a compelling and challenging task. In this work the problem of unsupervised outlier detection in large collections of data objects modeled by means of arbitrary multidimensional probability density functions is considered. We present a novel definition ofuncertain distance-based outlierunder the attribute level uncertainty model, according to which an uncertain object is an object that always exists but its actual value is modeled by a multivariate pdf. According to this definition an uncertain object is declared to be an outlier on the basis of the expected number of its neighbors in the dataset. To the best of our knowledge this is the first work that considers the unsupervised outlier detection problem on data objects modeled by means of arbitrarily shaped multidimensional distribution functions. We present the UDBOD algorithm which efficiently detects the outliers in an input uncertain dataset by taking advantages of three optimized phases, that are parameter estimation, candidate selection, and the candidate filtering. An experimental campaign is presented, including a sensitivity analysis, a study of the effectiveness of the technique, a comparison with related algorithms, also in presence of high dimensional data, and a discussion about the behavior of our technique in real case scenarios.
Fabrizio Angiulli, Fabio Fassetti
J. Intell. Inf. Syst.2
2020 Improving Deep Unsupervised Anomaly Detection by Exploiting VAE Latent Space Distribution
Fabrizio Angiulli, Fabio Fassetti, Luca Ferragina
DS2
2020 An Automatic Tool For Language Evaluation
abstract
The aim of evaluating children speech and language is to measure their communication skills. In particular, the speech language pathologist is interested in determining the child’s impairments in the areas of language, articulation, voice, fluency and swallowing. In literature some standardized tests have been proposed to assess and screen developmental language impairments but they require manual laborious transcription, annotation and calculation. This work is very time demanding and, also, may introduce several kinds of errors in the evaluation phase and non-uniform evaluations. In order to help therapists, a system performing automated evaluation is proposed. Providing as input the correct sentence and the sentence produced by patients, the technique evaluates the level of the verbal production and returns a score. The main phases of the method concern an ad-hoc transformation of the produced sentence in the reference sentence and in the evaluation of the cost of this transformation. Since the cost function is related to many weights, a learning phase is defined to automatically set such weights.
Fabio Fassetti, Ilaria Fassetti
LREC1
2019 A Density Estimation Approach for Detecting and Explaining Exceptional Values in Categorical Data
Fabrizio Angiulli, Fabio Fassetti, Luigi Palopoli 0001, Cristina Serrao
DS2
2019 FEDRO: a software tool for the automatic discovery of candidate ORFs in plants with c →u RNA editing
abstract
BACKGROUND: RNA editing is an important mechanism for gene expression in plants organelles. It alters the direct transfer of genetic information from DNA to proteins, due to the introduction of differences between RNAs and the corresponding coding DNA sequences. Software tools successful for the search of genes in other organisms not always are able to correctly perform this task in plants organellar genomes. Moreover, the available software tools predicting RNA editing events utilise algorithms that do not account for events which may generate a novel start codon. RESULTS: We present FEDRO, a Java software tool implementing a novel strategy to generate candidate Open Reading Frames (ORFs) resulting from Cytidine to Uridine (c→u) editing substitutions which occur in the mitochondrial genome (mtDNA) of a given input plant. The goal is to predict putative proteins of plants mitochondria that have not been yet annotated. In order to validate the generated ORFs, a screening is performed by checking for sequence similarity or presence in active transcripts of the same or similar organisms. We illustrate the functionalities of our framework on a model organism. CONCLUSIONS: The proposed tool may be used also on other organisms and genomes. FEDRO is publicly available at http://math.unipa.it/rombo/FEDRO .
Fabio Fassetti, Claudia Giallombardo, Ofelia Leone, Luigi Palopoli 0001, Simona E. Rombo, Adolfo Saiardi
BMC Bioinform.1
2017 Modular Construction of Minimal Models
Rachel Ben-Eliyahu-Zohary, Fabrizio Angiulli, Fabio Fassetti, Luigi Palopoli 0001
LPNMR3
2017 Outlying property detection with numerical attributes
Fabrizio Angiulli, Fabio Fassetti, Giuseppe Manco 0001, Luigi Palopoli 0001
Data Min. Knowl. Discov.2
2016 Anomaly Detection in Networks with Temporal Information
Fabrizio Angiulli, Fabio Fassetti, Estela Narvaez
DS2
2016 Toward Generalizing the Unification with Statistical Outliers: The Gradient Outlier Factor Measure
abstract
In this work, we introduce a novel definition of outlier, namely the Gradient Outlier Factor (or GOF), with the aim to provide a definition that unifies with the statistical one on some standard distributions but has a different behavior in the presence of mixture distributions. Intuitively, the GOF score measures the probability to stay in the neighborhood of a certain object. It is directly proportional to the density and inversely proportional to the variation of the density. We derive formal properties under which the GOF definition unifies the statistical outlier definition and show that the unification holds for some standard distributions, while the GOF is able to capture tails in the presence of different distributions even if their densities sensibly differ. Moreover, we provide a probabilistic interpretation of the GOF score, by means of the notion of density of the data density. Experimental results confirm that there are scenarios in which the novel definition can be profitably employed. To the best of our knowledge, except for distance-based outlier, no other data mining outlier definition has a so clearly established relationship with statistical outliers.
Fabrizio Angiulli, Fabio Fassetti
ACM Trans. Knowl. Discov. Data2
2014 On the tractability of minimal model computation for some CNF theories
Fabrizio Angiulli, Rachel Ben-Eliyahu-Zohary, Fabio Fassetti, Luigi Palopoli 0001
Artif. Intell.3
2014 Exploiting domain knowledge to detect outliers
Fabrizio Angiulli, Fabio Fassetti
Data Min. Knowl. Discov.2
2013 Principal Directions-Based Pivot Placement
Fabrizio Angiulli, Fabio Fassetti
SISAP2
2013 Nearest Neighbor-Based Classification of Uncertain Data
abstract
This work deals with the problem of classifying uncertain data. With this aim we introduce the Uncertain Nearest Neighbor (UNN) rule, which represents the generalization of the deterministic nearest neighbor rule to the case in which uncertain objects are available. The UNN rule relies on the concept of nearest neighbor class, rather than on that of nearest neighbor object. The nearest neighbor class of a test object is the class that maximizes the probability of providing its nearest neighbor. The evidence is that the former concept is much more powerful than the latter in the presence of uncertainty, in that it correctly models the right semantics of the nearest neighbor decision rule when applied to the uncertain scenario. An effective and efficient algorithm to perform uncertain nearest neighbor classification of a generic (un)certain test object is designed, based on properties that greatly reduce the temporal cost associated with nearest neighbor class probability computation. Experimental results are presented, showing that the UNN rule is effective and efficient in classifying uncertain data.
Fabrizio Angiulli, Fabio Fassetti
ACM Trans. Knowl. Discov. Data2
2013 Discovering Characterizations of the Behavior of Anomalous Subpopulations
abstract
We consider the problem of discovering attributes, or properties, accounting for the a priori stated abnormality of a group of anomalous individuals (the outliers) with respect to an overall given population (the inliers). To this aim, we introduce the notion of exceptional property and define the concept of exceptionality score, which measures the significance of a property. In particular, in order to single out exceptional properties, we resort to a form of minimum distance estimation for evaluating the badness of fit of the values assumed by the outliers compared to the probability distribution associated with the values assumed by the inliers. Suitable exceptionality scores are introduced for both numeric and categorical attributes. These scores are, both from the analytical and the empirical point of view, designed to be effective for small samples, as it is the case for outliers. We present an algorithm, called EXPREX, for efficiently discovering exceptional properties. The algorithm is able to reduce the needed computational effort by not exploring many irrelevant numerical intervals and by exploiting suitable pruning rules. The experimental results confirm that our technique is able to provide knowledge characterizing outliers in a natural manner.
Fabrizio Angiulli, Fabio Fassetti, Luigi Palopoli 0001
IEEE Trans. Knowl. Data Eng.2
2012 Indexing Uncertain Data in General Metric Spaces
abstract
In this study, we deal with the problem of efficiently answering range queries over uncertain objects in a general metric space. In this study, an uncertain object is an object that always exists but its actual value is uncertain and modeled by a multivariate probability density function. As a major contribution, this is the first work providing an effective technique for indexing uncertain objects coming from general metric spaces. We generalize the reverse triangle inequality to the probabilistic setting in order to exploit it as a discard condition. Then, we introduce a novel pivot-based indexing technique, called UP-index, and show how it can be employed to speed up range query computation. Importantly, the candidate selection phase of our technique is able to noticeably reduce the set of candidates with little time requirements. Finally, we provide a criterion to measure the quality of a set of pivots and study the problem of selecting a good set of pivots according to the introduced criterion. We report some intractability results and then design an approximate algorithm with statistical guarantees for selecting pivots. Experimental results validate the effectiveness of the proposed approach and reveal that the introduced technique may be even preferable to indexing techniques specifically designed for the euclidean space.
Fabrizio Angiulli, Fabio Fassetti
IEEE Trans. Knowl. Data Eng.2
2011 L-SME: A System for Mining Loosely Structured Motifs
Fabio Fassetti, Gianluigi Greco, Giorgio Terracina
ECML/PKDD (3)1
2010 Detection of Discriminating Rules
Fabrizio Angiulli, Fabio Fassetti, Luigi Palopoli 0001, Domenico Trimboli
ICAART (1)2
2010 Finding Distance-based Outliers in Subspaces through Both Positive and Negative Examples
Fabio Fassetti, Fabrizio Angiulli
ICAART (1)1
2010 Distance-based outlier queries in data streams: the novel task and algorithms
Fabrizio Angiulli, Fabio Fassetti
Data Min. Knowl. Discov.2
2010 On the complexity of identifying head-elementary-set-free programs
abstract
Abstract Head-elementary-set-free (HEF) programs were proposed in (Gebser et al. 2007) and shown to generalize over head-cycle-free programs while retaining their nice properties. It was left as an open problem in (Gebser et al. 2007) to establish the complexity of identifying HEF programs. This note solves the open problem by showing that the problem is complete for coNP.
Fabio Fassetti, Luigi Palopoli 0001
Theory Pract. Log. Program.1
2009 Outlier Detection Using Inductive Logic Programming
abstract
We present a novel definition of outlier in the context of inductive logic programming. Given a set of positive and negative examples, the definition aims at singling out the examples showing anomalous behavior. We note that the task here pursued is different from noise removal, and, in fact, the anomalous observations we discover are different in nature from noisy ones. We discuss pecularities of the novel approach, present an algorithm for detecting outliers, discuss some examples of knowledge mined, and compare it with alternative approaches.
Fabrizio Angiulli, Fabio Fassetti
ICDM2
2009 DOLPHIN: An efficient algorithm for mining distance-based outliers in very large datasets
abstract
In this work a novel distance-based outlier detection algorithm, named DOLPHIN, working on disk-resident datasets and whose I/O cost corresponds to the cost of sequentially reading the input dataset file twice, is presented. It is both theoretically and empirically shown that the main memory usage of DOLPHIN amounts to a small fraction of the dataset and that DOLPHIN has linear time performance with respect to the dataset size. DOLPHIN gains efficiency by naturally merging together in a unified schema three strategies, namely the selection policy of objects to be maintained in main memory, usage of pruning rules, and similarity search techniques. Importantly, similarity search is accomplished by the algorithm without the need of preliminarily indexing the whole dataset, as other methods do. The algorithm is simple to implement and it can be used with any type of data, belonging to either metric or nonmetric spaces. Moreover, a modification to the basic method allows DOLPHIN to deal with the scenario in which the available buffer of main memory is smaller than its standard requirements. DOLPHIN has been compared with state-of-the-art distance-based outlier detection algorithms, showing that it is much more efficient.
Fabrizio Angiulli, Fabio Fassetti
ACM Trans. Knowl. Discov. Data2
2009 Detecting outlying properties of exceptional objects
abstract
Assume you are given a data population characterized by a certain number of attributes. Assume, moreover, you are provided with the information that one of the individuals in this data population is abnormal, but no reason whatsoever is given to you as to why this particular individual is to be considered abnormal. In several cases, you will be indeed interested in discovering such reasons. This article is precisely concerned with this problem of discovering sets of attributes that account for the (a priori stated) abnormality of an individual within a given dataset. A criterion is presented to measure the abnormality of combinations of attribute values featured by the given abnormal individual with respect to the reference population. In this respect, each subset of attributes is intended to somehow represent a “property” of individuals. We distinguish between global and local properties. Global properties are subsets of attributes explaining the given abnormality with respect to the entire data population. With local ones, instead, two subsets of attributes are singled out, where the former one justifies the abnormality within the data subpopulation selected using the values taken by the exceptional individual on those attributes included in the latter one. The problem of individuating abnormal properties with associated explanations is formally stated and analyzed. Such a formal characterization is then exploited in order to devise efficient algorithms for detecting both global and local forms of most abnormal properties. The experimental evidence, which is accounted for in the article, shows that the algorithms are both able to mine meaningful information and to accomplish the computational task by examining a negligible fraction of the search space.
Fabrizio Angiulli, Fabio Fassetti, Luigi Palopoli 0001
ACM Trans. Database Syst.2
2008 Mining Loosely Structured Motifs from Biological Data
abstract
The discovery of information encoded in biological sequences is assuming a prominent role in identifying genetic diseases and in deciphering biological mechanisms. This information is usually encoded in patterns frequently occurring in the sequences, also called motifs. In fact, motif discovery has received much attention in the literature, and several algorithms have already been proposed, which are specifically tailored to deal with motifs exhibiting some kinds of "regular structure". Motivated by biological observations, this paper focuses on the mining of loosely structured motifs, i.e., of more general kinds of motif where several "exceptions" may be tolerated in pattern repetitions. To this end, an algorithm exploiting data structures conceived to efficiently handle pattern variabilities is presented and analyzed. Furthermore, a randomized variant with linear time and space complexity is introduced, and a theoretical guarantee on its performances is proven. Both algorithms have been implemented and tested on real data sets. Despite the ability of mining very complex kinds of pattern, performance results evidence a genome-wide applicability of the proposed techniques.
Fabio Fassetti, Gianluigi Greco, Giorgio Terracina
IEEE Trans. Knowl. Data Eng.1
2007 Very efficient mining of distance-based outliers
abstract
In this work a novel algorithm, named DOLPHIN, for detecting distance-based outliers is presented.
Fabrizio Angiulli, Fabio Fassetti
CIKM2
2007 Detecting distance-based outliers in streams of data
abstract
In this work a method for detecting distance-based outliers in data streams is presented. We deal with the sliding window model, where outlier queries are performed in order to detect anomalies in the current window. Two algorithms are presented. The first one exactly answers outlier queries, but has larger space requirements. The second algorithm is directly derived from the exact one, has limited memory requirements and returns an approximate answer based on accurate estimations with a statistical guarantee. Several experiments have been accomplished, confirming the effectiveness of the proposed approach and the high quality of approximate solutions.
Fabrizio Angiulli, Fabio Fassetti
CIKM2