VLDB 2026 Research / reviewers in the wild / expert
Stefano Lodi
dblp:61/3451
· DBLP profile ↗
24ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0001-8861-6341ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 95% Database theory · 2% Data integration and cleaning · 2% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 65% High-performance computing · 19% Parallel and multicore computing · 13% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
anomaly detection |
0.4 | 2 | 2016 | GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016 Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013 |
Data mining › anomaly detection › outlier detection
distance-based outlier detection |
0.4 | 2 | 2016 | GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016 Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013 |
GPUs and heterogeneous computing › GPU computing
GPU algorithms |
0.2 | 1 | 2016 | GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016 |
Data mining › big data analytics › large-scale data mining
distributed data mining |
0.2 | 1 | 2013 | Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013 |
High-performance computing
parallel and distributed algorithms |
0.1 | 1 | 2016 | GPU Strategies for Distance-Based Outlier Detection · IEEE Trans. Parallel Distributed Syst. 2016 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 2013 | Distributed Strategies for Mining Outliers in Large Data Sets · IEEE Trans. Knowl. Data Eng. 2013 |
Data mining
clustering |
0.0 | 1 | 2003 | Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003 |
Data mining › clustering
density-based clustering |
0.0 | 1 | 2003 | Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003 |
Data mining › clustering › large-scale clustering
distributed clustering |
0.0 | 1 | 2003 | Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003 |
Data models and query languages
complex object databases |
0.0 | 1 | 1998 | Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998 |
Data integration and cleaning › data quality
inconsistency detection |
0.0 | 1 | 1998 | Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998 |
Database theory
integrity constraints |
0.0 | 1 | 1998 | Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998 |
Logic in computer science › proof systems
tableau method |
0.0 | 1 | 1998 | Consistency Checking in Complex Object Database Schemata with Integrity Constraints · IEEE Trans. Knowl. Data Eng. 1998 |
Distributed systems › distributed data processing
distributed data mining |
0.0 | 1 | 2003 | Distributed Clustering Based on Sampling Local Density Estimates · IJCAI 2003 |
Methods — techniques the papers use, named apart from their topics
bruteforce · 0.5solvingset · 0.4solving set · 0.4distributed computation · 0.3sampling · 0.1density estimation · 0.1tableaux calculus · 0.0path relations · 0.0cyclic descriptions · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantum computing in the automotive industry: survey, challenges, and perspectives
Muhammad Waqas Arshad, Stefano Lodi |
J. Supercomput. | 2 |
| 2021 | Self-supervised Bernoulli Autoencoders for Semi-supervised Hashing
Ricardo Ñanculef, Francisco Alejandro Mena, Antonio Macaluso, Stefano Lodi, Claudio Sartori 0001 |
CIARP | 4 |
| 2020 | Quantum splines for non-linear approximationsabstractQuantum Computing offers a new paradigm for efficient computing and many AI applications could benefit from its potential boost in performance. However, the main limitation is the constraint to linear operations that hampers the representation of complex relationships in data. In this work, we propose an efficient implementation of quantum splines for non-linear approximation. In particular, we first discuss possible parametrisations, and select the most convenient for exploiting the HHL algorithm to obtain the estimates of spline coefficients. Then, we investigate QSpline performance as an evaluation routine for some of the most popular activation functions adopted in ML. Finally, a detailed comparison with classical alternatives to the HHL is also presented. Antonio Macaluso, Luca Clissa, Stefano Lodi, Claudio Sartori 0001 |
CF | 3 |
| 2020 | Reducing distance computations for distance-based outliers
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001 |
Expert Syst. Appl. | 3 |
| 2017 | Clustering Distributed Short Time Series with Dense PatternsabstractThe clustering of genes with similar temporal profiles is an important task in gene expression data analysis. Current approaches to the clustering of sparse gene expression data with temporal information suffer from their at least quadratic complexity in the number of clusters, the number of genes, or both, and are not distributed. In this paper, we present the first distributed and density-based approach to short time series clustering, called DTSCluster, which is suitable for gene expression data. DTSCluster identifies dense patterns in the distributed datasets and uses them to generate the time series clusters. The comparative experimental results revealed that DTSCluster is scalable in the dataset size with linear complexity in time and space, and outperforms other representative approaches in terms of cluster validation with the silhouette index as well. The distributed scenario also opens up the opportunity for collaborative data mining between different gene expression data holders. Josenildo Costa da Silva, Gustavo H. B. Oliveira, Stefano Lodi, Matthias Klusch |
ICMLA | 3 |
| 2016 | Privacy-Awareness of Distributed Data Clustering Algorithms Revisited
Josenildo Costa da Silva, Matthias Klusch, Stefano Lodi |
IDA | 3 |
| 2016 | Fast and scalable Lasso via stochastic Frank-Wolfe methods with a convergence guarantee
Emanuele Frandi, Ricardo Ñanculef, Stefano Lodi, Claudio Sartori 0001, Johan A. K. Suykens |
Mach. Learn. | 3 |
| 2016 | GPU Strategies for Distance-Based Outlier DetectionabstractThe process of discovering interesting patterns in large, possibly huge, data sets is referred to as data mining, and can be performed in several flavours, known as “data mining functions.” Among these functions, outlier detection discovers observations which deviate substantially from the rest of the data, and has many important practical applications. Outlier detection in very large data sets is however computationally very demanding and currently requires high-performance computing facilities. We propose a family of parallel and distributed algorithms for graphic processing units (GPU) derived from two distance-based outlier detection algorithms: BruteForce and SolvingSet. The algorithms differ in the way they exploit the architecture and memory hierarchy of the GPU and guarantee significant improvements with respect to the CPU versions, both in terms of scalability and exploitation of parallelism. We provide a detailed discussion of their computational properties and measure performances with an extensive experimentation, comparing the several implementations and showing significant speedups. Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Training Support Vector Machines using Frank-Wolfe Optimization MethodsabstractTraining a support vector machine (SVM) requires the solution of a quadratic programming problem (QP) whose computational complexity becomes prohibitively expensive for large scale datasets. Traditional optimization methods cannot be directly applied in these cases, mainly due to memory restrictions. By adopting a slightly different objective function and under mild conditions on the kernel used within the model, efficient algorithms to train SVMs have been devised under the name of core vector machines (CVMs). This framework exploits the equivalence of the resulting learning problem with the task of building a minimal enclosing ball (MEB) problem in a feature space, where data is implicitly embedded by a kernel function. In this paper, we improve on the CVM approach by proposing two novel methods to build SVMs based on the Frank–Wolfe algorithm, recently revisited as a fast method to approximate the solution of a MEB problem. In contrast to CVMs, our algorithms do not require to compute the solutions of a sequence of increasingly complex QPs and are defined by using only analytic optimization steps. Experiments on a large collection of datasets show that our methods scale better than CVMs in most cases, sometimes at the price of a slightly lower accuracy. As CVMs, the proposed methods can be easily extended to machine learning problems other than binary classification. However, effective classifiers are also obtained using kernels which do not satisfy the condition required by CVMs, and thus our methods can be used for a wider set of problems. Emanuele Frandi, Ricardo Ñanculef, Maria Grazia Gasparo, Stefano Lodi, Claudio Sartori 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2013 | Distributed Strategies for Mining Outliers in Large Data SetsabstractWe introduce a distributed method for detecting distance-based outliers in very large data sets. Our approach is based on the concept of outlier detection solving set [2], which is a small subset of the data set that can be also employed for predicting novel outliers. The method exploits parallel computation in order to obtain vast time savings. Indeed, beyond preserving the correctness of the result, the proposed schema exhibits excellent performances. From the theoretical point of view, for common settings, the temporal cost of our algorithm is expected to be at least three orders of magnitude faster than the classical nested-loop like approach to detect outliers. Experimental results show that the algorithm is efficient and that its running time scales quite well for an increasing number of nodes. We discuss also a variant of the basic strategy which reduces the amount of data to be transferred in order to improve both the communication cost and the overall runtime. Importantly, the solving set computed by our approach in a distributed environment has the same quality as that produced by the corresponding centralized method. Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | A New Algorithm for Training SVMs Using Approximate Minimal Enclosing Balls
Emanuele Frandi, Maria Grazia Gasparo, Stefano Lodi, Ricardo Ñanculef, Claudio Sartori 0001 |
CIARP | 3 |
| 2010 | A Distributed Approach to Detect Outliers in Very Large Data Sets
Fabrizio Angiulli, Stefano Basta, Stefano Lodi, Claudio Sartori 0001 |
Euro-Par (1) | 3 |
| 2010 | Single-Pass Distributed Learning of Multi-class SVMs Using Core-SetsabstractWe explore a technique to learn Support Vector Models (SVMs) when training data is partitioned among several data sources. The basic idea is to consider SVMs which can be reduced to Minimal Enclosing Ball (MEB) problems in an feature space. Computation of such SVMs can be efficiently achieved by finding a core-set for the image of the data in the feature space. Our main result is that the union of local core-sets provides a close approximation to a global core-set from which the SVM can be recovered. The method requires hence a single pass through each source of data in order to compute local core-sets and then to recover the SVM from its union. Extensive simulations in small and large datasets are presented in order to evaluate its classification accuracy, transmission efficiency and global complexity, comparing its results with a widely used single-pass heuristic to learn standard SVMs. Stefano Lodi, Ricardo Ñanculef, Claudio Sartori 0001 |
SDM | 1 |
| 2008 | Semantic peer, here are the neighbors you want!abstractPeer Data Management Systems (PDMSs) have been introduced as a solution to the problem of large-scale sharing of semantically rich data. A PDMS consists of semantic peers connected through semantic mappings. Querying a PDMS may lead to very poor results, because of the semantic degradation due to the approximations given by the traversal of the semantic mappings, thus leading to the problem of how to boost a network of mappings in a PDMS.In this paper we propose a strategy for the incremental maintenance of a flexible network organization that clusters together peers which are semantically related in Semantic Overlay Networks (SONs), while maintaining a high degree of node autonomy. Semantic features, a summarized representation of clusters, are stored in a light structure which effectively assists a newly entering peer when choosing its semantically closest overlay networks. Then, each peer is supported in the selection of its own neighbors within each overlay network according to two policies: Range-based selection and k-NN selection. For both policies, we introduce specific algorithms which exploit a distributed indexing mechanism for efficient network navigation. The proposed approach has been implemented in a prototype where its effectiveness and efficiency have been extensively tested. Wilma Penzo, Stefano Lodi, Federica Mandreoli, Riccardo Martoglia, Simona Sassatelli |
EDBT | 2 |
| 2007 | W*-Grid: A Robust Decentralized Cross-layer Infrastructure for Routing and Multi-Dimensional Data Management in Wireless Ad-Hoc Sensor Networks
Gabriele Monti, Gianluca Moro, Stefano Lodi |
Peer-to-Peer Computing | 3 |
| 2006 | Stream Clustering Based on Kernel Density Estimation
Stefano Lodi, Gianluca Moro, Claudio Sartori 0001 |
ECAI | 1 |
| 2006 | Inferences on Kernel Density Estimates by Solving Nonlinear SystemsabstractKernel density estimators are a popular family of nonparametric estimators with applications to exploratory statistics and data mining. Since kernel estimators must be constructed from the data, if the data are sensitive, only indirect representations of the estimate, such as graphs or tabulations, can be stored or transmitted. However, even such representations might contain enough information to allow for data reconstruction, yielding an inference problem for kernel estimates. The inference problem for kernel estimators can be described by a system of nonlinear equations that arises naturally from the kernel estimate of a multivariate dataset. The solution to the system is the set of data from which the kernel estimate was computed and, in practice a good approximation to the solution is not available. A serious threat to data privacy is posed by publicly available solvers for nonlinear systems. This paper investigates the numerical solution of the nonlinear systems arising from the kernel estimate of a multivariate dataset and shows that this task is challenging. In fact, the Jacobian matrix of the system is numerically singular and a large number of solvers for nonlinear equations fails as they have to solve linear systems whose coefficient matrix is given by the Jacobian. Further, up to date solvers for optimization problems that do not suffer from this drawback may fail to solve the nonlinear system. To show this fact, we tested a subspace trustregion method, a BFGS method and a gradient projection method on both a synthetic and a real dataset. These methods are able to find a solution to the optimization problem even starting far from it. However, the experimental results on both the synthetic and the real dataset show that, if the initial guess is not very close to the solution, all three meth- yielding an inference problem for kernel estimates. Consider for instance an investment bank database. Different customers have invested their savings in two funds in different amounts. If a bi-dimensional kernel estimate is graphically displayed as a three-dimensional graph, then it might be possible to derive a set of (x, y, z) triplets describing the input/output relationship of the estimate at given points on the plane. In the presentation of results, usually the parameters of the estimate are communicated, therefore the analytical form of the estimate is entirely known (except the data points). This knowledge could be exploited to recover the data points by attempting to solve a system of equations having the data points as variables. Then, the dataset could be compared to information leaked from other sources in order to assign the reconstructed points to individual customers. Stefania Bellavia, Stefano Lodi, Benedetta Morini |
SSDBM | 2 |
| 2006 | Privacy-preserving agent-based distributed data clustering
Josenildo Costa da Silva, Matthias Klusch, Stefano Lodi, Gianluca Moro |
Web Intell. Agent Syst. | 3 |
| 2004 | Inference Attacks in Peer-to-Peer Homogeneous Distributed Data Mining
Josenildo Costa da Silva, Matthias Klusch, Stefano Lodi, Gianluca Moro |
ECAI | 3 |
| 2003 | Distributed Clustering Based on Sampling Local Density Estimates
Matthias Klusch, Stefano Lodi, Gianluca Moro |
IJCAI | 2 |
| 2002 | Detecting Outbreaks by Time Series AnalysisabstractExceptional events in a time series are observations which can be regarded as qualitatively significant anomalies. The detection of such events is an interesting problem in several domains, in particular for the generation of alarms in clinical microbiology. We propose an approach to the detection of exceptional events based on model selection. For each mathematical form of a model, we choose the parameters of the model by maximum likelihood techniques. Then we select, among the resulting instantiated models, the model which minimizes the mean square error. An exceptional event is detected with an assigned probability, if an observation lies outside the forecasting region defined by the selected model and a confidence interval. Gianfranco Cellarosi, Stefano Lodi, Claudio Sartori 0001 |
CBMS | 2 |
| 1999 | Efficient Shared Near Neighbours Clustering of Large Metric Data Sets
Stefano Lodi, Luisella Reami, Claudio Sartori 0001 |
PKDD | 1 |
| 1998 | Consistency Checking in Complex Object Database Schemata with Integrity ConstraintsabstractIntegrity constraints are rules that should guarantee the integrity of a database. Provided an adequate mechanism to express them is available, the following question arises: is there any way to populate a database which satisfies the constraints supplied by a database designer? That is, does the database schema, including constraints, admit at least a nonempty model? This work answers the above question in a complex object database environment, providing a theoretical framework, including the following ingredients: (1) two alternative formalisms, able to express a relevant set of state integrity constraints with a declarative style; (2) two specialized reasoners, based on the tableaux calculus, able to check the consistency of complex objects database schemata expressed with the two formalisms. The proposed formalisms share a common kernel, which supports complex objects and object identifiers, and which allow the expression of acyclic descriptions of: classes, nested relations and views, built up by means of the recursive use of record, quantified set, and object type constructors and by the intersection, union, and complement operators. Furthermore, the kernel formalism allows the declarative formulation of typing constraints and integrity rules. In order to improve the expressiveness and maintain the decidability of the reasoning activities, we extend the kernel formalism into two alternative directions. The first formalism, OLCP, introduces the capability of expressing path relations. Because cyclic schemas are extremely useful, we introduce a second formalism, OLCD, with the capability of expressing cyclic descriptions but disallowing the expression of path relations. In fact, we show that the reasoning activity in OLCDP (i.e., OLCP with cycles) is undecidable. Domenico Beneventano, Sonia Bergamaschi, Stefano Lodi, Claudio Sartori 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1994 | The E/S Knowledge Representation System
Sonia Bergamaschi, Stefano Lodi, Claudio Sartori 0001 |
Data Knowl. Eng. | 2 |