Stefano Lodi

dblp:61/3451 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0001-8861-6341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 95% Database theory · 2% Data integration and cleaning · 2%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 65% High-performance computing · 19% Parallel and multicore computing · 13%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.422016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.422016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
GPUs and heterogeneous computing › GPU computing
GPU algorithms
0.212016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Data mining › big data analytics › large-scale data mining
distributed data mining
0.212013
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
High-performance computing
parallel and distributed algorithms
0.112016
GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016
Parallel and multicore computing
parallel algorithms
0.012013
Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013
Data mining
clustering
0.012003
Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003
Data mining › clustering
density-based clustering
0.012003
Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003
Data mining › clustering › large-scale clustering
distributed clustering
0.012003
Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003
Data models and query languages
complex object databases
0.011998
Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998
Data integration and cleaning › data quality
inconsistency detection
0.011998
Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998
Database theory
integrity constraints
0.011998
Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998
Logic in computer science › proof systems
tableau method
0.011998
Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998
Distributed systems › distributed data processing
distributed data mining
0.012003
Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003

Methods — techniques the papers use, named apart from their topics

bruteforce · 0.5solvingset · 0.4solving set · 0.4distributed computation · 0.3sampling · 0.1density estimation · 0.1tableaux calculus · 0.0path relations · 0.0cyclic descriptions · 0.0
YearPublicationVenuePosition
2025 Quantum computing in the automotive industry: survey, challenges, and perspectives
Muhammad Waqas Arshad, Stefano Lodi
J. Supercomput.2
2021 Self-supervised Bernoulli Autoencoders for Semi-supervised Hashing
Ricardo Ñanculef, Francisco Alejandro Mena, Antonio Macaluso, Stefano Lodi, Claudio Sartori 0001
CIARP4
2020 Quantum splines for non-linear approximations
abstract
Quantum Computing offers a new paradigm for efficient computing and many AI applications could benefit from its potential boost in performance. However, the main limitation is the constraint to linear operations that hampers the representation of complex relationships in data. In this work, we propose an efficient implementation of quantum splines for non-linear approximation. In particular, we first discuss possible parametrisations, and select the most convenient for exploiting the HHL algorithm to obtain the estimates of spline coefficients. Then, we investigate QSpline performance as an evaluation routine for some of the most popular activation functions adopted in ML. Finally, a detailed comparison with classical alternatives to the HHL is also presented.
Antonio Macaluso, Luca Clissa, Stefano Lodi, Claudio Sartori 0001
CF3
2020 Reducing distance computations for distance-based outliers
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
Expert Syst. Appl.3
2017 Clustering Distributed Short Time Series with Dense Patterns
abstract
The clustering of genes with similar temporal profiles is an important task in gene expression data analysis. Current approaches to the clustering of sparse gene expression data with temporal information suffer from their at least quadratic complexity in the number of clusters, the number of genes, or both, and are not distributed. In this paper, we present the first distributed and density-based approach to short time series clustering, called DTSCluster, which is suitable for gene expression data. DTSCluster identifies dense patterns in the distributed datasets and uses them to generate the time series clusters. The comparative experimental results revealed that DTSCluster is scalable in the dataset size with linear complexity in time and space, and outperforms other representative approaches in terms of cluster validation with the silhouette index as well. The distributed scenario also opens up the opportunity for collaborative data mining between different gene expression data holders.
Josenildo Costa da Silva, Gustavo H. B. Oliveira, Stefano Lodi, Matthias Klusch
ICMLA3
2016 Privacy-Awareness of Distributed Data Clustering Algorithms Revisited
Josenildo Costa da Silva, Matthias Klusch, Stefano Lodi
IDA3
2016 Fast and scalable Lasso via stochastic Frank-Wolfe methods with a convergence guarantee
Emanuele Frandi, Ricardo Ñanculef, Stefano Lodi, Claudio Sartori 0001, Johan A. K. Suykens
Mach. Learn.3
2016 GPU Strategies for Distance-Based Outlier Detection
abstract
The process of discovering interesting patterns in large, possibly huge, data sets is referred to as data mining, and can be performed in several flavours, known as “data mining functions.” Among these functions, outlier detection discovers observations which deviate substantially from the rest of the data, and has many important practical applications. Outlier detection in very large data sets is however computationally very demanding and currently requires high-performance computing facilities. We propose a family of parallel and distributed algorithms for graphic processing units (GPU) derived from two distance-based outlier detection algorithms: BruteForce and SolvingSet. The algorithms differ in the way they exploit the architecture and memory hierarchy of the GPU and guarantee significant improvements with respect to the CPU versions, both in terms of scalability and exploitation of parallelism. We provide a detailed discussion of their computational properties and measure performances with an extensive experimentation, comparing the several implementations and showing significant speedups.
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
IEEE Trans. Parallel Distributed Syst.3
2013 Training Support Vector Machines using Frank-Wolfe Optimization Methods
abstract
Training a support vector machine (SVM) requires the solution of a quadratic programming problem (QP) whose computational complexity becomes prohibitively expensive for large scale datasets. Traditional optimization methods cannot be directly applied in these cases, mainly due to memory restrictions. By adopting a slightly different objective function and under mild conditions on the kernel used within the model, efficient algorithms to train SVMs have been devised under the name of core vector machines (CVMs). This framework exploits the equivalence of the resulting learning problem with the task of building a minimal enclosing ball (MEB) problem in a feature space, where data is implicitly embedded by a kernel function. In this paper, we improve on the CVM approach by proposing two novel methods to build SVMs based on the Frank–Wolfe algorithm, recently revisited as a fast method to approximate the solution of a MEB problem. In contrast to CVMs, our algorithms do not require to compute the solutions of a sequence of increasingly complex QPs and are defined by using only analytic optimization steps. Experiments on a large collection of datasets show that our methods scale better than CVMs in most cases, sometimes at the price of a slightly lower accuracy. As CVMs, the proposed methods can be easily extended to machine learning problems other than binary classification. However, effective classifiers are also obtained using kernels which do not satisfy the condition required by CVMs, and thus our methods can be used for a wider set of problems.
Emanuele Frandi, Ricardo Ñanculef, Maria Grazia Gasparo, Stefano Lodi, Claudio Sartori 0001
Int. J. Pattern Recognit. Artif. Intell.4
2013 Distributed Strategies for Mining Outliers in Large Data Sets
abstract
We introduce a distributed method for detecting distance-based outliers in very large data sets. Our approach is based on the concept of outlier detection solving set [2], which is a small subset of the data set that can be also employed for predicting novel outliers. The method exploits parallel computation in order to obtain vast time savings. Indeed, beyond preserving the correctness of the result, the proposed schema exhibits excellent performances. From the theoretical point of view, for common settings, the temporal cost of our algorithm is expected to be at least three orders of magnitude faster than the classical nested-loop like approach to detect outliers. Experimental results show that the algorithm is efficient and that its running time scales quite well for an increasing number of nodes. We discuss also a variant of the basic strategy which reduces the amount of data to be transferred in order to improve both the communication cost and the overall runtime. Importantly, the solving set computed by our approach in a distributed environment has the same quality as that produced by the corresponding centralized method.
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
IEEE Trans. Knowl. Data Eng.3
2010 A New Algorithm for Training SVMs Using Approximate Minimal Enclosing Balls
Emanuele Frandi, Maria Grazia Gasparo, Stefano Lodi, Ricardo Ñanculef, Claudio Sartori 0001
CIARP3
2010 A Distributed Approach to Detect Outliers in Very Large Data Sets
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001
Euro-Par (1)3
2010 Single-Pass Distributed Learning of Multi-class SVMs Using Core-Sets
abstract
We explore a technique to learn Support Vector Models (SVMs) when training data is partitioned among several data sources. The basic idea is to consider SVMs which can be reduced to Minimal Enclosing Ball (MEB) problems in an feature space. Computation of such SVMs can be efficiently achieved by finding a core-set for the image of the data in the feature space. Our main result is that the union of local core-sets provides a close approximation to a global core-set from which the SVM can be recovered. The method requires hence a single pass through each source of data in order to compute local core-sets and then to recover the SVM from its union. Extensive simulations in small and large datasets are presented in order to evaluate its classification accuracy, transmission efficiency and global complexity, comparing its results with a widely used single-pass heuristic to learn standard SVMs.
Stefano Lodi, Ricardo Ñanculef, Claudio Sartori 0001
SDM1
2008 Semantic peer, here are the neighbors you want!
abstract
Peer Data Management Systems (PDMSs) have been introduced as a solution to the problem of large-scale sharing of semantically rich data. A PDMS consists of semantic peers connected through semantic mappings. Querying a PDMS may lead to very poor results, because of the semantic degradation due to the approximations given by the traversal of the semantic mappings, thus leading to the problem of how to boost a network of mappings in a PDMS.In this paper we propose a strategy for the incremental maintenance of a flexible network organization that clusters together peers which are semantically related in Semantic Overlay Networks (SONs), while maintaining a high degree of node autonomy. Semantic features, a summarized representation of clusters, are stored in a light structure which effectively assists a newly entering peer when choosing its semantically closest overlay networks. Then, each peer is supported in the selection of its own neighbors within each overlay network according to two policies: Range-based selection and k-NN selection. For both policies, we introduce specific algorithms which exploit a distributed indexing mechanism for efficient network navigation. The proposed approach has been implemented in a prototype where its effectiveness and efficiency have been extensively tested.
Wilma Penzo, Stefano Lodi, Federica Mandreoli, Riccardo Martoglia, Simona Sassatelli
EDBT2
2007 W*-Grid: A Robust Decentralized Cross-layer Infrastructure for Routing and Multi-Dimensional Data Management in Wireless Ad-Hoc Sensor Networks
Gabriele Monti, Gianluca Moro, Stefano Lodi
Peer-to-Peer Computing3
2006 Stream Clustering Based on Kernel Density Estimation
Stefano Lodi, Gianluca Moro, Claudio Sartori 0001
ECAI1
2006 Inferences on Kernel Density Estimates by Solving Nonlinear Systems
abstract
Kernel density estimators are a popular family of nonparametric estimators with applications to exploratory statistics and data mining. Since kernel estimators must be constructed from the data, if the data are sensitive, only indirect representations of the estimate, such as graphs or tabulations, can be stored or transmitted. However, even such representations might contain enough information to allow for data reconstruction, yielding an inference problem for kernel estimates. The inference problem for kernel estimators can be described by a system of nonlinear equations that arises naturally from the kernel estimate of a multivariate dataset. The solution to the system is the set of data from which the kernel estimate was computed and, in practice a good approximation to the solution is not available. A serious threat to data privacy is posed by publicly available solvers for nonlinear systems. This paper investigates the numerical solution of the nonlinear systems arising from the kernel estimate of a multivariate dataset and shows that this task is challenging. In fact, the Jacobian matrix of the system is numerically singular and a large number of solvers for nonlinear equations fails as they have to solve linear systems whose coefficient matrix is given by the Jacobian. Further, up to date solvers for optimization problems that do not suffer from this drawback may fail to solve the nonlinear system. To show this fact, we tested a subspace trustregion method, a BFGS method and a gradient projection method on both a synthetic and a real dataset. These methods are able to find a solution to the optimization problem even starting far from it. However, the experimental results on both the synthetic and the real dataset show that, if the initial guess is not very close to the solution, all three meth- yielding an inference problem for kernel estimates. Consider for instance an investment bank database. Different customers have invested their savings in two funds in different amounts. If a bi-dimensional kernel estimate is graphically displayed as a three-dimensional graph, then it might be possible to derive a set of (x, y, z) triplets describing the input/output relationship of the estimate at given points on the plane. In the presentation of results, usually the parameters of the estimate are communicated, therefore the analytical form of the estimate is entirely known (except the data points). This knowledge could be exploited to recover the data points by attempting to solve a system of equations having the data points as variables. Then, the dataset could be compared to information leaked from other sources in order to assign the reconstructed points to individual customers.
Stefania Bellavia, Stefano Lodi, Benedetta Morini
SSDBM2
2006 Privacy-preserving agent-based distributed data clustering
Josenildo Costa da Silva, Matthias Klusch, Stefano Lodi, Gianluca Moro
Web Intell. Agent Syst.3
2004 Inference Attacks in Peer-to-Peer Homogeneous Distributed Data Mining
Josenildo Costa da Silva, Matthias Klusch, Stefano Lodi, Gianluca Moro
ECAI3
2003 Distributed Clustering Based on Sampling Local Density Estimates
Matthias Klusch, Stefano Lodi, Gianluca Moro
IJCAI2
2002 Detecting Outbreaks by Time Series Analysis
abstract
Exceptional events in a time series are observations which can be regarded as qualitatively significant anomalies. The detection of such events is an interesting problem in several domains, in particular for the generation of alarms in clinical microbiology. We propose an approach to the detection of exceptional events based on model selection. For each mathematical form of a model, we choose the parameters of the model by maximum likelihood techniques. Then we select, among the resulting instantiated models, the model which minimizes the mean square error. An exceptional event is detected with an assigned probability, if an observation lies outside the forecasting region defined by the selected model and a confidence interval.
Gianfranco Cellarosi, Stefano Lodi, Claudio Sartori 0001
CBMS2
1999 Efficient Shared Near Neighbours Clustering of Large Metric Data Sets
Stefano Lodi, Luisella Reami, Claudio Sartori 0001
PKDD1
1998 Consistency Checking in Complex Object Database Schemata with Integrity Constraints
abstract
Integrity constraints are rules that should guarantee the integrity of a database. Provided an adequate mechanism to express them is available, the following question arises: is there any way to populate a database which satisfies the constraints supplied by a database designer? That is, does the database schema, including constraints, admit at least a nonempty model? This work answers the above question in a complex object database environment, providing a theoretical framework, including the following ingredients: (1) two alternative formalisms, able to express a relevant set of state integrity constraints with a declarative style; (2) two specialized reasoners, based on the tableaux calculus, able to check the consistency of complex objects database schemata expressed with the two formalisms. The proposed formalisms share a common kernel, which supports complex objects and object identifiers, and which allow the expression of acyclic descriptions of: classes, nested relations and views, built up by means of the recursive use of record, quantified set, and object type constructors and by the intersection, union, and complement operators. Furthermore, the kernel formalism allows the declarative formulation of typing constraints and integrity rules. In order to improve the expressiveness and maintain the decidability of the reasoning activities, we extend the kernel formalism into two alternative directions. The first formalism, OLCP, introduces the capability of expressing path relations. Because cyclic schemas are extremely useful, we introduce a second formalism, OLCD, with the capability of expressing cyclic descriptions but disallowing the expression of path relations. In fact, we show that the reasoning activity in OLCDP (i.e., OLCP with cycles) is undecidable.
Domenico Beneventano, Sonia Bergamaschi, Stefano Lodi, Claudio Sartori 0001
IEEE Trans. Knowl. Data Eng.3
1994 The E/S Knowledge Representation System
Sonia Bergamaschi, Stefano Lodi, Claudio Sartori 0001
Data Knowl. Eng.2