EDBT 2026 Demo / reviewers in the wild / expert
Flavio Giobergia
dblp:235/0376
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-8806-7979ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FAME: Fictional Actors for Multilingual Erasure
Claudio Savelli, Moreno La Quatra, Alkis Koudounas, Flavio Giobergia |
LREC | 4 |
| 2026 | On the Evaluation of Machine Unlearning Methods: A Multi-domain Classification BenchmarkabstractAbstract Machine Unlearning (MU), the process of removing specific data influences from trained machine learning models, is critical for regulatory compliance (e.g., GDPR’s right to be forgotten) and for addressing copyright and privacy concerns in large-scale models. While a wide range of methods and metrics have been proposed, systematic evaluations remain fragmented, typically limited in scope by modality, metric coverage, or the number of methods considered. Moreover, the lack of standardized benchmarks leaves several gaps in evaluation protocols, including how to efficiently compare methods, identify optimal hyperparameters, and determine which experimental settings are appropriate for fair and meaningful benchmarking. To address these gaps, we present the most comprehensive MU benchmark to date, evaluating 12 unlearning methods across 8 classification datasets, 4 modalities, several hyperparameters and settings. Based on previous literature and our empirical results, we formalize evaluation protocol desiderata to guide future MU benchmarking. Following these guidelines, we report benchmark results highlighting the best methods within and across domains. To help with method comparison, we also introduce LUMA, a unified metric that aggregates core unlearning dimensions into a single score. Our code is reproducible and extensible to serve as a benchmark for MU research. Andrea D'Angelo, Claudio Savelli, Flavio Giobergia, Elena Baralis, Giovanni Stilo |
Mach. Learn. | 3 |
| 2026 | Enforcing domain constraints through Lagrangian primal-dual learningabstractAbstract While neural networks have demonstrated remarkable predictive capabilities in various scenarios, they typically struggle to learn to avoid regions of the output space that are considered off-limits due to known domain-specific constraints. This paper presents a novel primal-dual learning approach inspired by augmented Lagrangian methods to address a priori output constraints for neural network predictions. Our solution encodes domain constraints for the output space by using a static Implicit Neural Representation and penalising the violation of these constraints at the loss level. This choice allows full flexibility in incorporating the constraints. We conduct extensive evaluations on several 2D and 3D synthetic datasets with different constraint topologies and two real-world datasets to evaluate the effectiveness of the proposed method. Our approach consistently outperforms methods without constraints in all experiments, yielding higher accuracy and lower constraint violations. Finally, our method shows superior performance improvements in scenarios with limited data availability, opening up potential benefits for various applications. These include geo-localisation tasks, where accurate positioning is crucial, and physical problems with theoretically-imposed constraints. Simone Monaco, Flavio Giobergia, Alkis Koudounas, Daniele Apiletti |
Neural Comput. Appl. | 2 |
| 2025 | ERASURE: A Modular and Extensible Framework for Machine UnlearningabstractMachine Unlearning (MU) is an emerging research area that enables models to selectively forget specific data, a critical requirement for privacy compliance (e.g., GDPR, CCPA) and security. However, the lack of standardized benchmarks makes evaluating and developing unlearning methods difficult. To address this gap, we introduce ERASURE, a benchmarking and development framework designed to systematically assess MU techniques. ERASURE provides a modular, extensible, open-source environment with real-world datasets and standardized unlearning measures. The framework is designed with configuration-driven workflows and an inversion of control architecture, allowing integration of new datasets, models, and evaluation measures. ERASURE advances trustworthy AI research as a tool for researchers to develop and benchmark new MU methods. Andrea D'Angelo, Claudio Savelli, Gabriele Tagliente, Flavio Giobergia, Elena Baralis, Giovanni Stilo |
CIKM | 4 |
| 2025 | How to Make Reproducible Research in Machine Unlearning with ERASUREabstractMachine unlearning, the process of removing specific data influences from Machine Learning models, is critical for complying with regulations like the GDPR's right to be forgotten and addressing copyright disputes in large models. Despite its rising importance, the field still lacks standardized tools, hindering reproducibility and evaluation. Here, we present, in an extensive way, ERASURE, a unified framework enabling reproducibility by implementing common unlearning techniques, evaluation metrics, and dedicated datasets. ERASURE advances research, ensures solution comparability, and facilitates reproducibility, addressing future legal and ethical challenges in data management. Andrea D'Angelo, Claudio Savelli, Gabriele Tagliente, Flavio Giobergia, Elena Baralis, Giovanni Stilo |
IJCAI | 4 |
| 2025 | "Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
Alkis Koudounas, Claudio Savelli, Flavio Giobergia, Elena Baralis |
INTERSPEECH | 3 |
| 2025 | Detecting Interpretable Subgroup Drifts
Flavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena Baralis |
KDD (1) | 1 |
| 2025 | MAD: Multicriteria Anomaly Detection of Suspicious Financial Accounts from Billions of Cash TransactionsabstractThis paper presents a real-world deployment case study on using unsupervised anomaly detection for Anti-Money Laundering (AML).Using more than 2 billion anonymized bank transactions that Intesa Sanpaolo, a primary Italian financial institution, registered over 8 months, we developed, tuned and deployed a machine learning pipeline in production.Experts from Intesa Sanpaolo validated the performance of our approach against the institution's traditional rule-based system and checked new real-world cases the system allowed them to identify.Besides increasing both precision and recall by a factor of 6 in the detection of high-risk cases, our pipeline raises 200+ additional alerts during the 8-month period, manually identified by branch managers, but missed by the rulebased system.More importantly, a manual inspection of 100 new unseen cases revealed 28 significant previously unreported cases.The pipeline, now fully deployed in Intesa Sanpaolo's Transaction Monitoring system, highlights the advantages of machine learning over traditional approaches typically adopted in this traditionally very conservative sector. Giordano Paoletti, Flavio Giobergia, Danilo Giordano, Luca Cagliero, Silvia Ronchiadin, Dario Moncalvo, Marco Mellia, Elena Baralis |
KDD (2) | 2 |
| 2025 | Enhancing Software Maintainability Through LLM-Assisted Code Refactoring
Tommaso Fulcini, Riccardo Coppola, Flavio Giobergia, Amirali Changizi, Meelad Dashti, Kimia Dorrani, Domenico Amalfitano, Damiano Distante, Filippo Ricca |
PROFES | 3 |
| 2024 | A Contrastive Learning Approach to Mitigate Bias in Speech Models
Alkis Koudounas, Flavio Giobergia, Eliana Pastor, Elena Baralis |
INTERSPEECH | 2 |
| 2023 | Late Fusion-based Distributed Multimodal LearningabstractMultimodal artificial intelligence promises deeper insights by analyzing data from diverse sources such as text, images, audio and more. However, efficiently processing and fusing large multimodal datasets remains an open challenge. This paper presents a Spark-based approach to parallelize multimodal encoding tasks. A key aspect is the use of late fusion with frozen backbone encoders, allowing encodings to be processed independently across cluster nodes. The encoded vectors can then be used for a variety of supervised and unsupervised tasks, regardless of whether they are gradient-based or not. Experimental results on image, text and audio datasets show that Spark clusters can offer competitive performance compared to GPUs, especially for I/O-intensive modalities. While GPUs outperform the Spark cluster when sufficient CPU cores are available, Spark makes it possible to use already available commodity hardware. The presented architecture demonstrates how distributed computing platforms like Spark can be effectively “repurposed” for multimodal AI, enhancing scalability and making such systems more accessible. Flavio Giobergia, Elena Baralis |
IEEE Big Data | 1 |
| 2021 | Dissecting a data-driven prognostic pipeline: A powertrain use case
Danilo Giordano, Eliana Pastor, Flavio Giobergia, Tania Cerquitelli, Elena Baralis, Marco Mellia, Alessandra Neri, Davide Tricarico |
Expert Syst. Appl. | 3 |
| 2020 | DSLE: A Smart Platform for Designing Data Science CompetitionsabstractDuring the last years an increasing number of university-level and post-graduation courses on Data Science have been offered. Practices and assessments need specific learning environments where learners could play with data samples and run machine learning and data mining algorithms. To foster learner engagement many closed-and open-source platforms support the design of data science competitions. However, they show limitations on the ability to handle private data, customize the analytics and evaluation processes, and visualize learners' activities and outcomes. This paper presents Data Science Lab Environment (DSLE, in short), a new open-source platform to design and monitor data science competitions. DSLE offers a easily configurable interface to share training and test data, design group works or individual sessions, evaluate the competition runs according to customizable metrics, manage public and private leaderboards, monitor participants' activities and their progress over time. The paper describes also a real experience of usage of DSLE in the context of a 1st-year M.Sc. course, which has involved around 160 students. Giuseppe Attanasio, Flavio Giobergia, Andrea Pasini, Francesco Ventura, Elena Baralis, Luca Cagliero, Paolo Garza, Daniele Apiletti, Tania Cerquitelli, Silvia Chiusano |
COMPSAC | 2 |
| 2019 | Fast Self-Organizing Maps TrainingabstractSelf-organizing maps are an unsupervised machine learning technique that offers interpretable results by identifying topological properties in high-dimensional datasets and projecting them on a 2-dimensional grid. An important problem of self-organizing maps is the computational expensiveness of their training phase. In this paper, we propose a fast approach to train self-organizing maps. The approach consists of 2 steps. First, a small map identifies the most relevant areas from the entire high-dimensional input space. Then a larger map (initialized from the small one) is fine-tuned to further explore the local areas identified in the first step. The resulting map has performance (measured in terms of accuracy and quantization error) on par with self-organizing maps trained with the standard approach, but with a significantly reduced training time. Flavio Giobergia, Elena Baralis |
IEEE BigData | 1 |
| 2018 | Mining Sensor Data for Predictive Maintenance in the Automotive IndustryabstractPredictive maintenance is an ever-growing area of interest, spanning different fields and approaches. In the automotive industry faulty behaviors of the oxygen sensor are a key challenge to address. This paper presents OxyClog, a data-driven framework that, given a large number of time series collected from a vehicle's ECU (engine control unit), builds a model to predict if the oxygen sensor is currently unclogged, almost clogged (since the clogging of the sensor happens gradually), or clogged. OxyClog is characterized by a tailored preprocessing, which includes a custom and interpretable feature selection algorithm, along with a summarization strategy to transform a time-dependent problem into a time-independent one. Furthermore, a semi-supervised labeling methodology has been devised to use different data sources with different characteristics to define meaningful clogging labels. OxyClog integrates state-of-the-art classification algorithms - both interpretable and non-interpretable - to process real ECU data with good prediction performance. Flavio Giobergia, Elena Baralis, Maria Camuglia, Tania Cerquitelli, Marco Mellia, Alessandra Neri, Davide Tricarico, Alessia Tuninetti |
DSAA | 1 |