EDBT 2026 Demo / reviewers in the wild / expert
Marco Arazzi
dblp:321/3418
· DBLP profile ↗
21ranked-venue papers
15as first author
21since 2021 · last 2026
0000-0002-3371-307XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SecureBreak: A Dataset towards Safe and Secure Models
Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera |
DATA (1) | 1 |
| 2026 | SD-RAG: A framework for secure selective disclosure in retrieval-augmented generation against single-turn prompt-leaking attacksabstractRetrieval-Augmented Generation (RAG) has attracted significant attention due to its ability to combine the generative capabilities of Large Language Models (LLMs) with knowledge obtained through efficient retrieval mechanisms over large-scale data collections. Currently, the majority of existing approaches overlook the risks associated with exposing sensitive or access-controlled information directly to the generation model. Only a few approaches propose techniques to instruct the generative model to refrain from disclosing sensitive information; however, recent studies have also demonstrated that such strategies remain vulnerable to prompt leaking attacks that can exfiltrate sensitive information via prompt injection. For these reasons, we propose a novel approach to Selective Disclosure in Retrieval-Augmented Generation, called SD-RAG, which decouples the enforcement of privacy constraints from the answer-generation process itself. SD-RAG relies on pre-redaction, applying sanitization and disclosure controls during the retrieval phase, prior to augmenting the question-answering LLM’s input with sensitive data. Moreover, we introduce a semantic mechanism to allow the ingestion of human-readable dynamic security and privacy constraints together with an optimized graph-based data model that supports fine-grained, policy-aware retrieval. In our experiments, we focus on the single-turn scenario, where an external attacker that relies on a malicious prompt template attempts to obtain sensitive information from the system by asking one question. Our experimental evaluation shows a promising improvement over the baseline in the single-turn prompt leaking scenario, achieving up to a 58% increase in the keyword-based privacy score metric that we introduce. Aiman Al Masoud, Marco Arazzi, Antonino Nocera |
Expert Syst. Appl. | 2 |
| 2026 | Let's focus: Focused backdoor attack against federated transfer learningabstractFederated Transfer Learning (FTL) is the most general variation of Federated Learning. According to this distributed paradigm, a feature learning pre-step is commonly carried out by only one party, typically the server, on publicly shared data. After that, the Federated Learning phase takes place to train a classifier collaboratively using the learned feature extractor. Each involved client contributes by locally training only the classification layers on a private training set. The peculiarity of an FTL scenario makes it hard to understand whether poisoning attacks can be developed to craft an effective backdoor. State-of-the-art attack strategies assume the possibility of shifting the model attention toward relevant features introduced by a forged trigger injected in the input data by some untrusted clients. Of course, this is not feasible in FTL, as the learned features are fixed once the server performs the pre-training step. Consequently, in this paper, we investigate this intriguing Federated Learning scenario to identify and exploit a vulnerability obtained by combining eXplainable AI (XAI) and dataset distillation. In particular, the proposed attack can be carried out by one of the clients during the Federated Learning phase of FTL by identifying the optimal local for the trigger through XAI and encapsulating compressed information of the backdoor class. Due to its behavior, we refer to our approach as a focused backdoor approach (FB-FTL for short) and test its performance by explicitly referencing an image classification scenario. With an average 80% attack success rate, obtained results show the effectiveness of our attack also against existing defenses for Federated Learning. Marco Arazzi, Stefanos Koffas, Antonino Nocera, Stjepan Picek |
Neurocomputing | 1 |
| 2025 | An IoE-based Framework Supporting Human-Centric IndustryabstractIndustry 5.0 envisions manufacturing systems that are human-centric, sustainable, and resilient. In this context, the Internet of Everything (IoE) enables integration of devices, people, and processes into a unified digital ecosystem. This paper presents a modular, semantically enriched framework that supports this transition by managing heterogeneous data sources—such as IoT sensors, wearable devices, and smart objects—through a layered architecture. The platform enables real-time data stream processing, semantic interoperability, and secure, context-aware access. Anomaly detection is enabled through a privacy-preserving mechanism based on behavioral fingerprinting and federated learning. The platform supports immersive human-machine interaction via gesture recognition, empowering workers to control and interact with industrial systems. Use cases demonstrate the system’s ability to support gesture-based control and intelligent monitoring, highlighting its potential to enhance adaptability, security, and worker empowerment in Industry 5.0 environments. Marco Arazzi, Alberto Belli, Claudio Cusano, Tullio Facchinetti, Marco Ferretti, Gabriele Galimberti, Monica Marconi Sciarroni, Paolo Napoletano, Antonino Nocera, Paola Pierleoni, Emanuele Storti, Domenico Ursino |
ETFA | 1 |
| 2025 | A Privacy-Preserving and Biometric-Aware Tasks Reallocation Strategy in Industry 5.0abstractIndustry 5.0 represents an emerging industrial paradigm that emphasizes seamless collaboration between human workers, collaborative robots (cobots), and smart objects. Its goal is to enable intelligent, adaptive manufacturing environments that not only boost operational efficiency and ensure regulatory compliance but also enhance workplace safety. In this context, we designed a complete framework based on a Reinforcement Learning (RL) strategy for intelligent and privacy-preserving task reallocation. Central to our vision is the prioritization of human well-being ensuring that both worker safety and privacy are protected, while the performance and reliability of machines and devices are optimized to support a truly human-centric manufacturing system. Our solution monitors workers’ physiological states and detects signs of fatigue, stress, or overload, ensuring that tasks can be dynamically reallocated to another worker or cobot to promote well-being without manual intervention. Moreover, by ensuring biometric data remains local to the worker’s device, the system respects data sovereignty and avoids unnecessary sharing of sensitive health information, guaranteeing compliance with regulations like GDPR. Our solution can adapt dynamically to the changing conditions and needs of human operators creating a privacy-preserving, safe, and efficient collaborative environment between people and machines. A comprehensive experimental analysis assesses the accuracy and performance of the proposed approach. Marco Arazzi, Mert Cihangiroglu, Serena Nicolazzo, Antonino Nocera |
ETFA | 1 |
| 2025 | Securing IoE Environments with Semantic Data Stream Analysis and Behavioral FingerprintingabstractIn the landscape of Industry 5.0, Internet of Everything (IoE) networks are emerging as crucial components for connecting diverse industrial sensors and devices, expanding beyond traditional IoT boundaries to integrate people, processes, and data. However, this increased connectivity raises significant security concerns, as the growing complexity of IoE environments introduces new attack vectors and privacy risks. Additionally, the integration of heterogeneous devices and data sources presents both technical and semantic interoperability challenges, requiring robust mechanisms for meaningful data interpretation and secure exchange. This paper, developed within the HOMEY project, presents an architecture for gathering and monitoring semantic data streams in IoE environments, addressing both interoperability and security challenges. Our approach leverages Knowledge Graphs to represent sensor metadata, locations, access rights, and operational contexts, enabling dynamic stream monitoring and data querying. An approach based on Federated Learning allows distributed behavioral fingerprinting of IoE devices, which is exploited on top of the platform to perform anomaly detection from real-time data streams. The approach enhances reliable, privacy-preserving anomaly detection, contributing to the security and resilience of next-generation industrial IoE ecosystems. Marco Arazzi, Monica Marconi Sciarroni, Serena Nicolazzo, Antonino Nocera, Emanuele Storti |
ETFA | 1 |
| 2025 | Augmented Knowledge Graph Querying leveraging LLMsabstractAdopting Knowledge Graphs (KGs) as a structured, semantic-oriented, data representation model has significantly improved data integration, reasoning, and querying capabilities across different domains. This is especially true in modern scenarios such as Industry 5.0, where the integration of data from humans, smart devices, and production processes is crucial, not only for industrial innovation, but also for supporting the digital transition of government administrations and organizations. However, the management, retrieval, and visualization of data from a KG using formal query languages can be difficult for non-expert users due to their technical complexity, thus limiting their usage inside industrial environments. For this reason, we introduce SparqLLM, a framework that utilizes a Retrieval-Augmented Generation (RAG) solution, to enhance the querying of Knowledge Graphs (KGs). SparqLLM executes the Extract, Transform, and Load (ETL) pipeline to construct KGs from raw data. It also features a natural language interface powered by Large Language Models (LLMs) to enable automatic SPARQL query generation. By integrating template-based methods as retrieved-context for the LLM, SparqLLM enhances query reliability and reduces semantic errors, ensuring more accurate and efficient KG interactions. Moreover, to improve usability, the system incorporates a dynamic visualization dashboard that adapts to the structure of the retrieved data, presenting the query results in an intuitive format. Rigorous experimental evaluations demonstrate that SparqLLM achieves high query accuracy, improved robustness, and user-friendly interaction with KGs, establishing it as a scalable solution to access semantic data. Marco Arazzi, Davide Ligari, Serena Nicolazzo, Antonino Nocera |
IJCNN | 1 |
| 2025 | Secure Federated Dataset DistillationabstractDataset Distillation (DD) is a powerful technique for reducing large datasets into compact, representative synthetic datasets, accelerating Machine Learning training. However, traditional DD methods operate in a centralized manner, which poses significant privacy threats and reduces its applicability. To mitigate these risks, we propose a Secure Federated Data Distillation (SFDD) framework to decentralize the distillation process while preserving privacy. Unlike existing Federated Distillation techniques that focus on training global models with distilled knowledge, our approach aims to produce a distilled dataset without exposing local contributions. We leverage the gradient-matching-based distillation method, adapting it for a distributed setting where clients contribute to the distillation process without sharing raw data. The central aggregator iteratively refines a synthetic dataset by integrating client-side updates while ensuring data confidentiality. To make our approach resilient to inference attacks perpetrated by the server that could exploit gradient updates to reconstruct private data, we create an optimized Local Differential Privacy approach, called LDPO-RLD (Label Differential Privacy Obfuscation via Randomized Linear Dispersion). Furthermore, we assess the framework’s resilience against malicious clients executing backdoor attacks (such as Doorping) and demonstrate robustness under the assumption of a sufficient number of participating clients. Our experimental results demonstrate the effectiveness of SFDD and that the proposed defense concretely mitigates the identified vulnerabilities, with minimal impact on the performance of the distilled dataset. By addressing the interplay between privacy and federation in dataset distillation, this work advances the field of privacy-preserving Machine Learning making our SFDD framework a viable solution for sensitive data-sharing applications. Marco Arazzi, Mert Cihangiroglu, Serena Nicolazzo, Antonino Nocera |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A defense mechanism against label inference attacks in Vertical Federated Learning
Marco Arazzi, Serena Nicolazzo, Antonino Nocera |
Neurocomputing | 1 |
| 2025 | DroidTTP: Mapping android applications with TTP for Cyber Threat IntelligenceabstractThe widespread use of Android devices for sensitive operations has made them prime targets for sophisticated cyber threats, including Advanced Persistent Threats (APT). Traditional malware detection methods focus primarily on malware classification, often failing to reveal the Tactics, Techniques, and Procedures (TTPs) used by attackers. To address this issue, we propose DroidTTP, a novel system for mapping Android malware to attack behaviors. We curated a dataset linking Android applications to Tactics and Techniques and developed an automated mapping approach using the Problem Transformation Approach and Large Language Models (LLMs). Our pipeline includes dataset construction, feature selection, data augmentation, model training, and explainability via SHAP. Furthermore, we explored the use of LLMs for TTP prediction using both Retrieval Augmented Generation and fine-tuning strategies. The Label Powerset XGBoost model achieved the best performance, with Jaccard Similarity scores of 0.9893 for Tactic classification and 0.9753 for Technique classification. The fine-tuned LLaMa model also performed competitively, achieving 0.9583 for Tactics and 0.9348 for Techniques. Although XGBoost slightly outperformed LLMs, the narrow performance gap highlights the potential of LLM-based approaches for Tactic and Technique prediction. Dincy R. Arikkat, P. Vinod 0001, Rafidha Rehiman K. A., Serena Nicolazzo, Marco Arazzi, Antonino Nocera, Mauro Conti |
J. Inf. Secur. Appl. | 5 |
| 2025 | Subject Data Auditing via Source Inference Attack in Cross-Silo Federated Learning
Marco Arazzi, Antonino Nocera, Mauro Conti |
J. Inf. Secur. Appl. | 2 |
| 2025 | RAG-IoE: IoT context-aware information retrieval with Large Language Models in Industry 5.0abstractHuman-centric design, intelligence, and seamless interconnectivity are key pillars of the Industry 5.0. A critical challenge in these scenarios is the efficient retrieval of relevant, context-aware information for workers within Internet of Everything (IoE) networks. Traditional information retrieval techniques struggle with the heterogeneous, dynamic data generated in industrial settings. To address this, we define a context-aware data model for IoE scenarios, on top of which we propose RAG-IoE, a novel Retrieval-Augmented Generation (RAG) solution to enable adaptive, scalable, and context-based information retrieval from both structured and unstructured data sources. Our approach organizes IoE data within a semantic framework, integrating hybrid retrieval methods. It combines structured search on a Knowledge Graph with unstructured data retrieval using embeddings stored in a vector database, followed by LLM-driven reasoning to refine results. This architecture enhances decision-making, reduces cognitive overload, and ensures precise guidance for industrial operators. We validate the efficiency and effectiveness of RAG-IoE using a novel dataset through both a user study and quantitative analysis, demonstrating its potential to optimize human-machine collaboration in Industry 5.0 environments. Marco Arazzi, Monica Marconi Sciarroni, Antonino Nocera, Emanuele Storti |
ACM Trans. Internet Things | 1 |
| 2024 | Privacy-preserving in Blockchain-based Federated Learning systems
K. M. Sameera, Serena Nicolazzo, Marco Arazzi, Antonino Nocera, Rafidha Rehiman K. A., P. Vinod 0001, Mauro Conti |
Comput. Commun. | 3 |
| 2024 | The SemIoE Ontology: A Semantic Model Solution for an IoE-Based IndustryabstractRecently, the Industry 5.0 is gaining attention as a novel paradigm, defining the next concrete steps toward more and more intelligent, green-aware, and user-centric digital systems. In an era in which smart devices typically adopted in the industry domain are more and more sophisticated and autonomous, the Internet of Things and its evolution, known as the Internet of Everything (IoE, for short), involving also people, robots, processes, and data in the network, represent the main driver to allow industries to put the experiences and needs of human beings at the center of their ecosystems. However, due to the extreme heterogeneity of the involved entities, their intrinsic need and capability to cooperate, and the aim to adapt to a dynamic user-centric context, special attention is required for the integration and processing of the data produced by such an IoE. This is the objective of the present paper, in which we propose a novel semantic model that formalizes the fundamental actors, elements and information of an IoE, along with their relationships. In our design, we focus on state-of-the-art design principles, in particular reuse, and abstraction, to build “SemIoE,” a lightweight ontology inheriting and extending concepts from well-known and consolidated reference ontologies. The defined semantic layer represents a core data model that can be extended to embrace any modern industrial scenario. It represents the base of an IoE knowledge graph, on the top of which, as an additional contribution, we analyze and define some essential services for an IoE-based industry. Marco Arazzi, Antonino Nocera, Emanuele Storti |
IEEE Internet Things J. | 1 |
| 2024 | A deep reinforcement learning approach for security-aware service acquisition in IoT
Marco Arazzi, Serena Nicolazzo, Antonino Nocera |
J. Inf. Secur. Appl. | 1 |
| 2024 | A novel IoT trust model leveraging fully distributed behavioral fingerprinting and secure delegationabstractThe pervasiveness and high number of Internet of Things (IoT) applications in people’s daily lives make this context a very critical attack surface for cyber threats. The high heterogeneity of involved entities, both in terms of hardware and software characteristics, does not allow the definition of uniform, global, and efficient security solutions. Therefore, researchers have started to investigate novel mechanisms, in which a super node (a gateway, a hub, or a router) analyzes the interactions of the target node with other peers in the network, to detect possible anomalies. The most recent of these strategies base such an analysis on the modeling of the fingerprint of a node behavior in an IoT; nevertheless, existing solutions do not cope with the fully distributed nature of the referring scenario. In this paper, we try to provide a contribution in this setting, by designing a novel and fully distributed trust model exploiting point-to-point devices’ behavioral fingerprints, a distributed consensus mechanism, and Blockchain technology. In our solution we tackle the non-trivial issue of equipping smart things with a secure mechanism to evaluate, also through their neighbors, the trustworthiness of an object in the network before interacting with it. Beyond the detailed description of our framework, we also illustrate the security model associated with it and the tests carried out to evaluate its correctness and performance. Marco Arazzi, Serena Nicolazzo, Antonino Nocera |
Pervasive Mob. Comput. | 1 |
| 2024 | PhotoStyle60: A Photographic Style Dataset for Photo Authorship Attribution and Photographic Style TransferabstractPhotography, like painting, allows artists to express themselves through their unique style. In digital photography, this is achieved not only with the choice of the subject and the composition but also by means of post-processing operations. The automatic identification of a photographer from the style of a photo is a challenging task, for many reasons, including the lack of suitable datasets including photos taken by a diverse panel of photographers with a clear photographic style. In this paper we present PhotoStyle60, a new dataset including 5708 photographs from 60 professional and semi-professional photographers. Additionally, we selected a reduced version of the dataset, called PhotoStyle10 containing images from 10 clearly distinguishable experts. We designed the dataset to address two tasks in particular: photo authorship attribution and photographic style transfer. In the former, we conducted an extensive analysis of the dataset through several classification experiments. In the latter, we explored the potential of our dataset to transfer a photographer's style to images from the Five-K dataset. Additionally, we propose also a simple but effective multi-image style transfer method that uses multiple samples of the target style. A user study demonstrated that such a method was able to reach accurate results, preserving the semantic content of the source photograph with very few artifacts. Marco Cotogni, Marco Arazzi, Claudio Cusano |
IEEE Trans. Multim. | 2 |
| 2023 | Turning Privacy-preserving Mechanisms against Federated LearningabstractRecently, researchers have successfully employed Graph Neural Networks (GNNs) to build enhanced recommender systems due to their capability to learn patterns from the interaction between involved entities. In addition, previous studies have investigated federated learning as the main solution to enable a native privacy-preserving mechanism for the construction of global GNN models without collecting sensitive data into a single computation unit. Still, privacy issues may arise as the analysis of local model updates produced by the federated clients can return information related to sensitive local data. For this reason, researchers proposed solutions that combine federated learning with Differential Privacy strategies and community-driven approaches, which involve combining data from neighbor clients to make the individual local updates less dependent on local sensitive data. Marco Arazzi, Mauro Conti, Antonino Nocera, Stjepan Picek |
CCS | 1 |
| 2023 | Predicting Tweet Engagement with Graph Neural NetworksabstractSocial Networks represent one of the most important online sources to share content across a world-scale audience. In this context, predicting whether a post will have any impact in terms of engagement is of crucial importance to drive the profitable exploitation of these media. In the literature, several studies address this issue by leveraging direct features of the posts, typically related to the textual content and the user publishing it. In this paper, we argue that the rise of engagement is also related to another key component, which is the semantic connection among posts published by users in social media. Hence, we propose TweetGage, a Graph Neural Network solution to predict the user engagement based on a novel graph-based model that represents the relationships among posts. To validate our proposal, we focus on the Twitter platform and perform a thorough experimental campaign providing evidence of its quality. Marco Arazzi, Marco Cotogni, Antonino Nocera, Luca Virgili |
ICMR | 1 |
| 2023 | The importance of the language for the evolution of online communities: An analysis based on Twitter and Reddit
Marco Arazzi, Serena Nicolazzo, Antonino Nocera, Manuel Zippo |
Expert Syst. Appl. | 1 |
| 2022 | An enhanced behavioral fingerprinting approach for the Internet of ThingsabstractWith the growing diffusion of the Internet of Things (IoT) technology across most of the aspects of people daily lives, security concerns have become critical to ensure the exploitation of advantages introduced by this technology. This is even more true in the context of Industry 4.0, for which the IoT is becoming an important driver for automation. The detection of anomalies in IoT systems to ensure the capability of such systems to tolerate attacks to single devices is a crucial aspect. Behavioral fingerprinting is a recent and promising security solution in this context, which still requires research efforts to embrace new challenges in such a complex environment. Existing solutions focus mostly on modeling the behavior of IoT devices by analyzing the information extracted from the header of exchanged networking packets. However, in many application contexts, also attacks on the content of the packets can lead to disruptive results. Our proposal focus on these approaches by addressing a fully distributed scenario in which computation is directly handled by IoT devices, also through delegation, and describes a novel behavioral fingerprinting approach based on features suitably engineered from packet payloads. The effective-ness of our proposed method is assessed by both simulated and experimental results. Alberico Aramini, Marco Arazzi, Tullio Facchinetti, Laurence S. Q. N. Ngankem, Antonino Nocera |
WFCS | 2 |