VLDB 2026 Research / reviewers in the wild / expert
Damian A. Tamburri
dblp:29/4648 · also Damian Andrew Tamburri, Damien A. Tamburri
· DBLP profile ↗
9ranked-venue papers in the field
1as first author
8since 2021 · last 2026
0000-0003-1230-8961ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 3Data Mining & Knowledge Discovery · 1Business Process & Enterprise Data · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How are MLOps Frameworks Used in Open Source Projects? An Empirical CharacterizationabstractMachine Learning (ML) Operations (MLOps) frameworks have been conceived to support developers and AI engineers in managing the lifecycle of their ML models. While such frameworks provide a wide range of features, developers may leverage only a subset of them, while missing some highly desired features. This paper investigates the practical use and desired feature enhancements of eight popular open-source MLOps frameworks. Specifically, we analyze their usage by dependent projects on GitHub, examining how they invoke the frameworks’ APIs and commands. Then, we qualitatively analyze feature requests and enhancements mined from the frameworks’ issue trackers, relating these desired improvements to the previously identified usage features. Results indicate that MLOps frameworks are rarely used out-of-the-box and are infrequently integrated into GitHub Workflows, but rather, developers use their APIs to implement custom functionality in their projects. Used features concern core ML phases and whole infrastructure governance, sometimes leveraging multiple frameworks with complementary features. The mapping with feature requests highlights that users mainly ask for enhancements to core features of the frameworks, but also better API exposure and CI/CD integration. Fiorella Zampetti, Federico Stocchetti, Federica Razzano, Damian A. Tamburri, Massimiliano Di Penta |
MSR | 4 |
| 2026 | LLMOps in Action: A Framework for Designing, Deploying, and Governing Advanced ChatbotsabstractThis paper presents a systematic literature review (SLR) focused on the implementation of chatbots using Large Language Models (LLMs), aimed at providing insights into the architectures, frameworks, best practices, and evaluation metrics that are shaping the field. By analyzing 39 primary studies, the review addresses six key research questions, exploring common architectures such as client-server and Retrieval-Augmented Generation (RAG), and identifying frequently utilized models, including the GPT family, BERT, and open-source models like LLaMA. The paper evaluates the performance of these models across various domains, emphasizing the impact of fine-tuning, prompt engineering, and embedding techniques on accuracy and domain-specific relevance. Additionally, it highlights the critical evaluation metrics used in LLM-based chatbot systems, including accuracy, user satisfaction, content quality, safety, and efficiency. Ethical considerations, including data governance, bias mitigation, and fairness audits, are also discussed to ensure responsible deployment of LLM chatbots. The review concludes with an exploration of the trade-offs between performance, cost-efficiency, and scalability, providing a comprehensive framework for future research and development of LLM-based chatbot applications. Pradheepan Raghavan, Damian A. Tamburri, Stefano Fossati, Willem-Jan van den Heuvel |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Copula-based Approaches for Anomaly Detection: a Case-study in Financial ForensicsabstractFraud detection is a critical challenge in various domains, necessitating accurate and reliable methods to distinguish between legitimate and fraudulent transactions. This work explores the application of copula-based models for anomaly detection in financial forensics. It focuses on their effectiveness in identifying fraudulent activities in a highly imbalanced dataset. Copula models are designed to capture the dependencies between continuous variables, providing a flexible framework for modeling the joint distribution of features. For instance, the variables in a dataset can follow different distributions and the copula is able to model how these variables jointly behave, particularly in extreme cases. In this work, we used a fraud dataset to calculate the copula-based probability of fraud and conditional Gaussian copula. Then, we derive the copula-based Generalized Linear Model (GLM) formula from the conditional copula which is essentially a GLM with probit link and transformed variables when the covariates are continuous. Finally, we compare the performance of these copula-based models with standard methods for predicting a binary variable like GLM with a probit link and logistic regression. Results indicate that copula-based probability and conditional copula formulas offer promising results, particularly in handling complex dependencies, but with a high computational time, while copula-based GLM, when combined with over-sampling, also outperforms traditional methods. Vanessa Tenconi, Damian A. Tamburri, Giovanni Quattrocchi, Corrado Pellegrino, Giuseppe Cascavilla, Willem-Jan van den Heuvel |
IEEE Big Data | 2 |
| 2024 | Few Images, Many Insights: Illicit Content Detection Using a Limited Number of ImagesabstractThe anonymity and untraceability benefits of the dark web increased its popularity exponentially. The cost of these technical benefits is that such anonymity has created a suitable womb for illicit activity. Hence—in collaboration with cybersecurity practitioners and law-enforcement agencies—the research community provided approaches for recognizing and classifying illicit activities. Most of these approaches exploit textual content from dark web markets, whereas few used images that originated from them. This article investigates alternative techniques for recognizing illegal activities from images. The significant contributions of our work are threefold: (a) We investigate label-agnostic learning techniques like One-Shot and Few-Shot learning that use Siamese Neural Networks. Our approach manages to handle small-scale datasets with promising accuracy. In particular, the Siamese Neural Network approach reaches 90.9% on 5-Shot experiments over a 10-class dataset. (b) This study’s satisfactory findings facilitate the creation of potent tools to assist authorities in identifying illicit content on the Web. Moreover, our proof-of-concept approach demonstrated the ability to recognize illegal images using a limited number of files, reducing the time constraint in collecting illegal images. (c) We provide a complete labeled dataset of 3,570 images from 55 different categories from dark web markets that can be used for future research activities. Giuseppe Cascavilla, Gemma Catolino, Mauro Conti, Dimos Mellios, Damian A. Tamburri |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | BigData Fusion for Trajectory Prediction of Multi-Sensor Surveillance Information SystemsabstractVideo surveillance information systems assist forensics to examine and analyze the evidence from crime scenes to develop objective findings in the investigation of crime. Often, the existing surveillance information systems exploit an array of security cameras and IoT devices monitoring the same crime scene from different points of view while the crime unfolds over a range of time. However, none can automatically and selectively merge big data streams connected to such systems to provide a holistic, end-to-end safety picture.This work proposes a trajectory prediction architecture framework within a multi-sensor surveillance system. We developed a novel position measurement technique using monocular depth perception networks with multi-camera setup using triangulation. We tested and compared our technique with a single camera sensor in our first experiment and as the multi-camera setup determines the position of our target more accurately, we used our measurement function in our second experiment. In our second experiment, we employed the Unscented Kalman Filter (UKF) for predicting the trajectory of the target, and proved that UKF has good potential for being used in surveillance systems. Lastly, we designed a general architecture framework for big data analysis in multi-sensor surveillance systems consisting the four layers: the Sensor Layer, the Single Sensor Computation Layer, the Data Fusion and Interpretation Layer, and the Human Interaction Layer. Giuseppe Cascavilla, Alfredo Cuzzocrea, Daniel De Pascale, Mandana Omidbakhsh, Damian A. Tamburri |
IEEE Big Data | 5 |
| 2023 | QuantumShare: Towards an Ontology for Bridging the Quantum Divide
Julian Martens, Indika Kumara, Geert Monsieur, Willem-Jan van den Heuvel, Damian A. Tamburri |
ER | 5 |
| 2023 | Real-world K-Anonymity applications: The KGen approach and its evaluation in fraudulent transactionsabstractK-Anonymity is a property for the measurement, management, and governance of the data anonymization. Many implementations of k-anonymity have been described in state of the art, but most of them are not practically usable over a large number of attributes in a “Big” dataset, i.e., a dataset drawing from Big Data. To address this significant shortcoming, we introduce and evaluate KGen, an approach to K-anonymity featuring meta-heuristics, specifically, Genetic Algorithms to compute a permutation of the dataset which is both K-anonymized and still usable for further processing, e.g., for private-by-design analytics. KGen promotes such a meta-heuristic approach since it can solve the problem by finding a pseudo-optimal solution in a reasonable time over a considerable load of input. KGen allows the data manager to guarantee a high anonymity level while preserving the usability and preventing loss of information entropy over the data. Differently from other approaches that provide optimal global solutions compatible with smaller datasets, KGen works properly also over Big datasets while still providing a good-enough K-anonymized but still processable dataset. Evaluation results show how our approach can still work efficiently on a real world dataset, provided by Dutch Tax Authority, with 47 attributes (i.e., the columns of the dataset to be anonymized) and over 1.5K+ observations (i.e., the rows of that dataset), as well as on a dataset with 97 attributes and over 3942 observations. Daniel De Pascale, Giuseppe Cascavilla, Damian A. Tamburri, Willem-Jan van den Heuvel |
Inf. Syst. | 3 |
| 2022 | Explaining IoT Attacks: An Effective and Efficient Semi-Supervised Learning FrameworkabstractCyber-attacks targeting Internet-of-Things (IoT) devices are prevalent due to the limited security resources of the target devices and their often limited connectivity. Explaining such attacks is therefore greatly important to construct countermeasures. Current methods of automated IoT attack analysis require either large amounts of labelled data for classification, or use clustering methods which can be inaccurate. However, when a desired grouping of the data, as well as some prior knowledge about some observations in the data is available, approximate semi-supervised learning methods may be used to create accurate cluster arrangements. We therefore investigated the use of semi-supervised clustering approaches for creating accurate clusters of IoT attack sessions based on their goals and characteristic commonalities. We first manually created a ground-truth grouping of recent IoT attacks based on their goal. We differentiated the goal of each session according to the purpose of the used commands and the taken approach, resulting in a total of five classes. We then automatically constructed a feature set suitable for clustering similar IoT attack sessions using a method proposed in recent literature, and passed it to two different semi-supervised clustering algorithms using either labelled data (SeededKMeans) or pairwise constraints (PCKMeans) as prior knowledge. We found that both semi-supervised approaches were able to create accurate cluster arrangements using only small amounts of prior knowledge. Moreover, they outperformed an entirely unsupervised KMeans algorithm in terms of accuracy. Giuseppe Cascavilla, Reinier Zwart, Damian A. Tamburri, Alfredo Cuzzocrea |
IEEE Big Data | 3 |
| 2020 | Design principles for the General Data Protection Regulation (GDPR): A formal concept analysis and its evaluationabstractData and software are nowadays one and the same: for this very reason, the European Union (EU) and other governments introduce frameworks for data protection — a key example being the General Data Protection Regulation (GDPR). However, GDPR compliance is not straightforward: its text is not written by software or information engineers but rather, by lawyers and policy-makers. As a design aid to information engineers aiming for GDPR compliance, as well as an aid to software users’ understanding of the regulation, this article offers a systematic synthesis and discussion of it, distilled by the mathematical analysis method known as Formal Concept Analysis (FCA). By its principles, GDPR is synthesised as a concept lattice, that is, a formal summary of the regulation, featuring 144372 records — its uses are manifold. For example, the lattice captures so-called attribute implications, the implicit logical relations across the regulation, and their intensity. These results can be used as drivers during systems and services (re-)design, development, operation, or information systems’ refactoring towards more GDPR consistency. Damian A. Tamburri |
Inf. Syst. | 1 |