VLDB 2026 Research / reviewers in the wild / expert
Diego Stucchi
dblp:263/7946
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-8285-5285ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MuRAL-CPD: Active Learning for Multiresolution Change Point DetectionabstractChange Point Detection (CPD) is a critical task in time series analysis, aiming to identify moments when the underlying data-generating process shifts. Traditional CPD methods often rely on unsupervised techniques, which lack adaptability to task-specific definitions of change and cannot benefit from user knowledge. To address these limitations, we propose MuRAL-CPD, a novel semi-supervised method that integrates active learning into a multiresolution CPD algorithm. MuRALCPD leverages a wavelet-based multiresolution decomposition to detect changes across multiple temporal scales and incorporates user feedback to iteratively optimize key hyperparameters. This interaction enables the model to align its notion of change with that of the user, improving both accuracy and interpretability. Our experimental results on several real-world datasets show the effectiveness of MuRAL-CPD against state-of-the-art methods, particularly in scenarios where minimal supervision is available. Stefano Bertolasi, Diego Carrera, Diego Stucchi, Pasqualina Fragneto, Luigi Amedeo Bianchi |
ICDM | 3 |
| 2025 | ClusterSSFDA: Clustered Semi-Supervised Federated Domain AdaptationabstractMost Federated Learning (FL) approaches assume a single global model is updated locally by clients and aggregated at the server. However, the single-model assumption is often too restrictive, especially in scenarios involving a large amount of different users, where adaptation to different domains and personalization are necessary to improve model performance. In this paper, we propose ClusterSSFDA, the first FL framework to leverage clustering for addressing domain shifts in Semi-Supervised Federated Learning (SSFL), where clients collect data without supervision. In particular, ClusterSSFDA clusters clients based on their models' agreement, and updates multiple models at the server, each tailored to a cluster of similar clients. ClusterSSFDA significantly improves adaptation in SSFL, striking a balance between a single global model, which may be suboptimal in the presence of domain shift among the clients, and fully personalized models, which would be trained on too small datasets. Our experiments on real-world scenarios with multiple levels of domain shift demonstrate that ClusterSSFDA outperforms existing methods, achieving superior performance in challenging SSFL settings. Michele Craighero, Taguhi Mesropyan, Diego Carrera, Beatrice Rossi, Diego Stucchi, Pasqualina Fragneto, Giacomo Boracchi |
ICDM | 5 |
| 2024 | SemiFDA: Domain Adaptation in Semi-Supervised Federated LearningabstractSemi-Supervised Federated Learning (SSFL) aims to improve a pretrained model using unlabeled data from clients. Traditional SSFL solutions relying on pseudo-labels or autoencoders often struggle in the presence of domain shift, i.e. a difference in data distributions between the server and the clients. In this paper we present SemiFDA, the first solution to effectively handle domain shift in SSFL. After training an initial classifier on the server's labeled data, we establish an unsupervised learning process at clients to train feature extractors based on encoders. This process adopts a custom unsupervised loss function that promotes the clients' encoders to align their feature distributions with those extracted by the encoder at server. The updated encoders are then aggregated at the server using Federated Aver-aging and sent back for the next iteration, while the classification head remains frozen to preserve the benefits of aligning features locally. Furthermore, we design an experimental framework to mimic various levels of domain shift and test SSFL methods in real-world scenarios, including HAR and Digit Classification. Our results also demonstrate the detrimental effects of domain shift in SSFL and show that SemiFDA outperforms other solutions under these challenging conditions. Michele Craighero, Giorgio Rossi, Beatrice Rossi, Diego Carrera, Diego Stucchi, Pasqualina Fragneto, Giacomo Boracchi |
ICDM | 5 |
| 2023 | Kernel QuantTreeabstractWe present Kernel QuantTree (KQT), a non-parametric change detection algorithm that monitors multivariate data through a histogram. KQT constructs a nonlinear partition of the input space that matches pre-defined target probabilities and specifically promotes compact bins adhering to the data distribution, resulting in a powerful detection algorithm. We prove two key theoretical advantages of KQT: i) statistics defined over the KQT histogram do not depend on the stationary data distribution $\phi_0$, so detection thresholds can be set a priori to control false positive rate, and ii) thanks to the kernel functions adopted, the KQT monitoring scheme is invariant to the roto-translation of the input data. Consequently, KQT does not require any preprocessing step like PCA. Our experiments show that KQT achieves superior detection power than non-parametric state-of-the-art change detection methods, and can reliably control the false positive rate. Diego Stucchi, Paolo Rizzo, Nicolò Folloni, Giacomo Boracchi |
ICML | 1 |
| 2023 | Multimodal Batch-Wise Change DetectionabstractWe address the problem of detecting distribution changes in a novel batch-wise and multimodal setup. This setup is characterized by a stationary condition where batches are drawn from potentially different modalities among a set of distributions in [Formula: see text] represented in the training set. Existing change detection (CD) algorithms assume that there is a unique-possibly multipeaked-distribution characterizing stationary conditions, and in batch-wise multimodal context exhibit either low detection power or poor control of false positives. We present MultiModal QuantTree (MMQT), a novel CD algorithm that uses a single histogram to model the batch-wise multimodal stationary conditions. During testing, MMQT automatically identifies which modality has generated the incoming batch and detects changes by means of a modality-specific statistic. We leverage the theoretical properties of QuantTree to: 1) automatically estimate the number of modalities in a training set and 2) derive a principled calibration procedure that guarantees false-positive control. Our experiments show that MMQT achieves high detection power and accurate control over false positives in synthetic and real-world multimodal CD problems. Moreover, we show the potential of MMQT in Stream Learning applications, where it proves effective at detecting concept drifts and the emergence of novel classes by solely monitoring the input distribution. Diego Stucchi, Luca Magri 0002, Diego Carrera, Giacomo Boracchi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Class Distribution Monitoring for Concept Drift DetectionabstractWe introduce Class Distribution Monitoring (CDM), an effective concept-drift detection scheme that monitors the class-conditional distributions of a datastream. In particular, our solution leverages multiple instances of an online and nonparametric change-detection algorithm based on QuantTree. CDM reports a concept drift after detecting a distribution change in any class, thus identifying which classes are affected by the concept drift. This can be precious information for diagnostics and adaptation. Our experiments on synthetic and real-world datastreams show that when the concept drift affects a few classes, CDM outperforms algorithms monitoring the overall data distribution, while achieving similar detection delays when the drift affects all the classes. Moreover, CDM outperforms comparable approaches that monitor the classification error, particularly when the change is not very apparent. Finally, we demonstrate that CDM inherits the properties of the underlying change detector, yielding an effective control over the expected time before a false alarm, or Average Run Length (ARL0). Diego Stucchi, Luca Frittoli, Giacomo Boracchi |
IJCNN | 1 |
| 2022 | Open-Set Recognition: an Inexpensive Strategy to Increase DNN ReliabilityabstractDeep Neural Networks (DNNs) are nowadays widely used in low-cost accelerators, characterized by limited computational resources. These models, and in particular DNNs for image classification, are becoming increasingly popular in safety-critical applications, where they are required to be highly reliable. Unfortunately, increasing DNNs reliability without computational overheads, which might not be affordable in low-power devices, is a non-trivial task. Our intuition is to detect network executions affected by faults as outliers with respect to the distribution of normal network’s output. To this purpose, we propose to exploit Open-Set Recognition (OSR) techniques to perform Fault Detection in an extremely low-cost manner. In particuar, we analyze the Maximum Logit Score (MLS), which is an established Open-Set Recognition technique, and compare it against other well-known OSR methods, namely OpenMax, energy-based outof-distribution detection and ODIN. Our experiments, performed on a ResNet-20 classifier trained on CIFAR-10 and SVHN datasets, demonstrate that MLS guarantees satisfactory detection performance while adding a negligible computational overhead. Most remarkably, MLS is extremely convenient to conFigure and deploy, as it does not require any modification or re-training of the existing network. A discussion of the advantages and limitations of the analysed solutions concludes the paper. Gabriele Gavarini, Diego Stucchi, Annachiara Ruospo, Giacomo Boracchi, Ernesto Sánchez 0001 |
IOLTS | 2 |
| 2022 | Fault Impact Estimation for Lightweight Fault Detection in Image FilteringabstractClassical redundancy-based fault detection techniques, such as Duplication with Comparison (DWC), rely on replicating the computation and comparing the replicas’ output at a bit-wise granularity. In many application environments these costs are prohibitive, especially when applications are characterized by an intrinsic level of tolerance. This article presents a novel fault-detection approach for the specific context of image filtering. Peculiarity of the proposed approach is that it estimates the impact of the fault on the processed output, in order to determine whether the image is usable or should be re-processed. To limit overheads, the proposed solution exploits Approximate Computing (AC), allowing the definition of disciplined AC strategies to trade-off between accuracy and costs. Core of our solution is the successful combination of Image Quality Assessment metrics and Machine Learning models to assess the visual impact of the fault in a lightweight manner. Extensive experimental campaigns demonstrate the effectiveness of the solution, achieving achieving a reduction in terms of execution time up to 44 percent with respect to the classical DWC, with a fault detection precision ranging from 94.58 to 96.70 percent, and recall ranging from 88.2 to 97.8 percent, depending on the adopted level of approximation. Cristiana Bolchini, Giacomo Boracchi, Luca Cassano, Antonio Miele, Diego Stucchi |
IEEE Trans. Computers | 5 |