Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Stefano Basta

dblp:42/717 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0002-9852-988XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data mining · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 67% High-performance computing · 20% Parallel and multicore computing · 13%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.532016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
Distance-Based Detection and Prediction of Outliers · IEEE Trans. Knowl. Data Eng. 2006
Data mining
anomaly detection
0.422016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
GPUs and heterogeneous computing › GPU computing
GPU algorithms
0.212016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Data mining › big data analytics › large-scale data mining
distributed data mining
0.212013
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
High-performance computing
parallel and distributed algorithms
0.112016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Data mining › anomaly detection
outlier detection
0.112006
Distance-Based Detection and Prediction of Outliers · IEEE Trans. Knowl. Data Eng. 2006
Parallel and multicore computing
parallel algorithms
0.012013
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013

Methods — techniques the papers use, named apart from their topics

bruteforce · 0.5solving set · 0.5solvingset · 0.4distributed computation · 0.3ROC analysis · 0.1
YearPublicationVenuePosition
2020 Reducing distance computations for distance-based outliers
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
Expert Syst. Appl.2
2016 GPU Strategies for Distance-Based Outlier Detection
abstract
The process of discovering interesting patterns in large, possibly huge, data sets is referred to as data mining, and can be performed in several flavours, known as “data mining functions.” Among these functions, outlier detection discovers observations which deviate substantially from the rest of the data, and has many important practical applications. Outlier detection in very large data sets is however computationally very demanding and currently requires high-performance computing facilities. We propose a family of parallel and distributed algorithms for graphic processing units (GPU) derived from two distance-based outlier detection algorithms: BruteForce and SolvingSet. The algorithms differ in the way they exploit the architecture and memory hierarchy of the GPU and guarantee significant improvements with respect to the CPU versions, both in terms of scalability and exploitation of parallelism. We provide a detailed discussion of their computational properties and measure performances with an extensive experimentation, comparing the several implementations and showing significant speedups.
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
IEEE Trans. Parallel Distributed Syst.2
2013 Distributed Strategies for Mining Outliers in Large Data Sets
abstract
We introduce a distributed method for detecting distance-based outliers in very large data sets. Our approach is based on the concept of outlier detection solving set [2], which is a small subset of the data set that can be also employed for predicting novel outliers. The method exploits parallel computation in order to obtain vast time savings. Indeed, beyond preserving the correctness of the result, the proposed schema exhibits excellent performances. From the theoretical point of view, for common settings, the temporal cost of our algorithm is expected to be at least three orders of magnitude faster than the classical nested-loop like approach to detect outliers. Experimental results show that the algorithm is efficient and that its running time scales quite well for an increasing number of nodes. We discuss also a variant of the basic strategy which reduces the amount of data to be transferred in order to improve both the communication cost and the overall runtime. Importantly, the solving set computed by our approach in a distributed environment has the same quality as that produced by the corresponding centralized method.
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
IEEE Trans. Knowl. Data Eng.2
2010 A Distributed Approach to Detect Outliers in Very Large Data Sets
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
Euro-Par (1)2
2006 Distance-Based Detection and Prediction of Outliers
abstract
A distance-based outlier detection method that finds the top outliers in an unlabeled data set and provides a subset of it, called outlier detection solving set, that can be used to predict the outlierness of new unseen objects, is proposed. The solving set includes a sufficient number of points that permits the detection of the top outliers by considering only a subset of all the pairwise distances from the data set. The properties of the solving set are investigated, and algorithms for computing it, with subquadratic time requirements, are proposed. Experiments on synthetic and real data sets to evaluate the effectiveness of the approach are presented. A scaling analysis of the solving set size is performed, and the false positive rate, that is, the fraction of new objects misclassified as outliers using the solving set instead of the overall data set, is shown to be negligible. Finally, to investigate the accuracy in separating outliers from inliers, ROC analysis of the method is accomplished. Results obtained show that using the solving set instead of the data set guarantees a comparable quality of the prediction, but at a lower computational cost.
Fabrizio Angiulli, Stefano Basta, Clara Pizzuti
IEEE Trans. Knowl. Data Eng.2
2004 Improving Prediction of Distance-Based Outliers
Fabrizio Angiulli, Stefano Basta, Clara Pizzuti
Discovery Science2