VLDB 2026 Research / reviewers in the wild / expert
Diletta Chiaro
dblp:336/4085
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-5145-4465ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LATInsights: LLM-Powered Conversational Agents for Geospatial Climate Analytics
Diletta Chiaro, Davide Piu, Antonio Elia Pascarella, Paolo De Piano, Giovanni Luca Giacco |
ICAART (1) | 1 |
| 2025 | FLAME: Federated Learning for Attack Mitigation and EvasionabstractIn today's interconnected cyber landscape, Distributed Denial of Service (DDoS) attacks represent a significant threat to the smooth functioning of online infrastructures. The nature of DDoS attacks, characterized by their distributed and dynamic nature, poses significant challenges for traditional centralized approaches to model training; however, the challenges of collaborative DDoS detection are compounded by stringent data privacy regulations, leaving mitigation efforts largely reliant on standalone and inflexible firewalls. Federated Learning (FL) represents a cutting-edge innovation in cybersecurity, presenting a revolutionary method for collectively training deep learning models without compromising sensitive data. Despite its promise, practical hurdles remain, particularly the reliance of most FL algorithms on centralized, server-side data for model evalu-ation-though some approaches avoid this centralized testing dependency. This limitation hinders the applicability of FL, especially in scenarios involving zero-day attacks on clients. Our paper examines a key hypothesis: whether the aggregated information from multiple clients can be effectively utilized to develop a global model that is inherently more resilient to zeroday attacks compared to models trained solely on individual client data. To investigate this, we introduce a methodology wherein FL models are trained on established DDoS attacks and subsequently evaluated against entirely novel, unencountered attacks, simulating zero-day scenarios at the client level. To ensure that each client contributes effectively to the training process, we utilize Jensen-Shannon Divergence (JSD) to evaluate and filter client updates based on their alignment with the global model. Building on this, we implement a kernel density estimation-based aggregation method to effectively mitigate feature distribution bias-a common issue in DDoS detection within FL environments. This approach forms a core component of our proposed framework, FLAME, which is built using the distributed framework Flower to realistically simulate FL in a decentralized setting. The code for our implementation can be found at: https://github.com/MODAL-UNINA/FLAME. Diletta Chiaro, Pian Qi, Edoardo Prezioso, Antonella Guzzo, Francesco Piccialli |
IPDPS | 1 |
| 2025 | CAPTURE - Computational Analysis and Predictive Techniques for Urban Resource EfficiencyabstractABSTRACT Municipal waste management (MWM) poses significant challenges in the context of rapid urbanisation and population growth. Accurate forecasting of waste production is crucial for designing sustainable waste management strategies. However, traditional forecasting methods often struggle to capture the complexities of waste generation dynamics. This paper proposes a novel methodology leveraging deep learning techniques to forecast municipal waste production. By harnessing the power of deep neural networks, our approach transcends the limitations of conventional models, providing more accurate and impactful predictions. We integrate heterogeneous data sources, including demographic and territorial information, into a comprehensive graph representation of municipalities. Graph Neural Networks are then employed to extract intricate spatial and temporal patterns from the graph structure. Empirical validation through a case study in the Apulia region demonstrates the effectiveness of our methodology in furnishing accurate forecasts for waste production. Our framework is adaptable and scalable, making it suitable for application across diverse geographical areas. This research contributes to advancing waste management practices by providing stakeholders with actionable insights for informed decision‐making. Marzia Canzaniello, Stefano Izzo, Diletta Chiaro, Antonella Longo, Francesco Piccialli |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | AGRIFOLD: AGRIculture Federated learning for Optimized Leaf disease DetectionabstractEfficient and accurate detection of plant leaf diseases is essential for protecting crop health and promoting sustainable and precision agriculture practices. However, the decentralized nature of agricultural data, combined with the inherent limitations of centralized Machine Learning (ML), presents significant challenges for developing scalable, privacy-preserving solutions. In this paper, we introduce AGRIFOLD, a Federated Learning (FL) framework designed to enable collaborative training of a lightweight Convolutional Neural Network (CNN) across diverse and distributed datasets while maintaining data privacy. By integrating an Efficient Channel Attention (ECA) mechanism into the VGG16 architecture, AGRIFOLD significantly improves classification accuracy and enhances interpretability through heatmaps that highlight regions affected by diseases. We evaluate the FL model using various aggregation methods, including FedAvg, FedProx, SCAFFOLD, FedBN, and FedDF, obtaining good accuracy levels for all tested aggregation strategies, with SCAFFOLD achieving the best overall performance. The model’s lightweight design, optimized through ablation and pruning techniques, facilitates deployment on resource-constrained edge devices. Additionally, to further support farmers’ decision-making, the framework incorporates a natural language processing-based recommender system that provides tailored treatment suggestions. Comprehensive experiments conducted on 12 heterogeneous datasets demonstrate high classification accuracy across 9 distinct leaf disease classes and healthy leaves, underscoring the practical potential of FL-based solutions for sustainable, real-world agricultural applications. The AGRIFOLD source code is available at https://github.com/MODAL-UNINA/AGRIFOLD . Francesco Piccialli, Ciro Della Bruna, Diletta Chiaro, Pian Qi, Martina Savoia |
Expert Syst. Appl. | 3 |
| 2025 | AgentAI: A comprehensive survey on autonomous agents in distributed AI for industry 4.0abstractAgentAI represents a transformative approach within distributed Artificial Intelligence (AI) in which autonomous agents work either individually or collaboratively in decentralized environments to address challenging problems. AgentAI enhances scalability, robustness, and flexibility by utilizing advanced communication, learning, and decision-making capabilities, making it integral to diverse applications in Industry 4.0. The ability of AI systems to interpret sensory data in open-world environments has seen significant advancements in recent years. This progress emphasizes the need to move beyond reductionist approaches and embrace more embodied and cohesive systems, which integrate foundational models into agent-driven actions. Existing surveys often focus on isolated domains or specific autonomy levels, lacking a cohesive analysis that spans the full spectrum of AgentAI development in Industry 4.0. This survey explicitly fills this gap by introducing a multi-domain taxonomy and by systematically analyzing both non-autonomous and fully autonomous AgentAI systems, offering a comprehensive synthesis not previously available in the literature. Additionally, the paper extends the discussion to Industry 5.0 and 6.0, exploring the evolution of AgentAI from automation to collaboration and, ultimately, to fully autonomous systems. This comprehensive analysis highlights the potential of AgentAI in driving industries toward a more efficient, sustainable, and adaptable future. Francesco Piccialli, Diletta Chiaro, Sundas Sarwar, Donato Cerciello, Pian Qi, Valeria Mele |
Expert Syst. Appl. | 2 |
| 2025 | Small models, big impact: A review on the power of lightweight Federated Learning
Pian Qi, Diletta Chiaro, Francesco Piccialli |
Future Gener. Comput. Syst. | 2 |
| 2025 | On the Road to AIoT: A Framework for Vehicle Road CooperationabstractThe paradigm of Augmented Intelligence of Things (AIoT) aims to empower Internet of Things (IoT) devices with intelligent capabilities to analyze data, make informed decisions, and execute actions autonomously. This study focuses on enhancing collaboration between vehicles and road infrastructure within the AIoT framework, particularly in the context of self-driving cars and smart city environments. A Proof of Concept (PoC) is presented, introducing a vehicle road cooperation framework tailored for online-vehicle-infrastructure cooperation (VIC) forecasting tasks. This framework enables real-time information exchange and trajectory prediction of target agents by leveraging IoT sensor technologies and incorporating two layers of cooperation: 1) ego-vehicles and 2) infrastructures. Experimental results demonstrate that the integration of information from both layers enhances prediction metrics compared to approaches focusing on individual layers. Comparative analysis with existing method, PP-VIC, underscores the superiority of the proposed framework in trajectory prediction. This research offers a promising avenue for enhancing communication and collaboration between infrastructure and autonomous vehicles, thereby contributing to the development of more efficient and safer transportation systems in smart cities. Daniela Annunziata, Diletta Chiaro, Pian Qi, Francesco Piccialli |
IEEE Internet Things J. | 2 |
| 2025 | FLAIR: Federated Learning for Augmented Industrial RetrievalabstractDeep learning (DL) has significantly advanced Industry 4.0 by leveraging data from the Industrial Internet of Things (IIoT) to enable smart manufacturing, predictive maintenance, and data-driven product marketing. However, multimodal industrial data presents challenges for traditional frameworks, including scalability, data privacy, and integration efficiency. This paper introduces an efficient product retrieval framework for e-commerce systems, addressing privacy and performance challenges through federated learning (FL). Specifically, we propose FLAIR (Federated Learning for Augmented Industrial Retrieval), a novel part retrieval system where distributed warehouses collaboratively train a multimodal foundation model, CLIP (Contrastive Language-Image Pre-Training), by fine-tuning only the Adapter module via FL, ensuring data privacy and efficiency. To address the limited availability of multimodal industrial data, our framework incorporates effective data augmentation strategies to enhance the diversity and quality of the training dataset. Comprehensive experiments on the Industrial Language-Image Dataset (ILID) highlight that FLAIR holds effective privacy safeguards and strong retrieval capabilities. Additionally, an advanced e-commerce recommendation system built on FLAIR showcases its practical effectiveness. FLAIR represents the first application of FL for industrial product retrieval, optimizing part searches, inventory management, and customer experience while maintaining data security. The complete code is available at https://github.com/MODAL-UNINA/FLAIR. Diletta Chiaro, Pian Qi, Valeria Mele, Francesco Piccialli |
IEEE Internet Things J. | 1 |
| 2025 | Generative AI-Empowered Digital Twin: A Comprehensive Survey With TaxonomyabstractGenerative artificial intelligence (GenAI) and digital twin (DT) technologies have individually demonstrated valuable capabilities across a range of fields. Their integration, however, offers a unique synergy with the potential to bring meaningful advancements in various sectors. In this survey, we explore the fusion of GenAI and DT, highlighting their combined ability to enhance insights, optimizations, and innovative solutions. We begin by clarifying the core principles of each technology and then outline their collaborative applications. Furthermore, we provide a detailed taxonomy of areas where GenAI has been leveraged within the realm of DT, and conversely, where DT technology has been enhanced through GenAI techniques. By systematically categorizing these applications, we aim to offer a clear perspective on the interplay between GenAI and DT across different sectors. Diletta Chiaro, Pian Qi, Antonio Pescapè, Francesco Piccialli |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Improving Energy Consumption Forecasting with Contextual Awareness: A Hybrid Deep Learning PerspectiveabstractAccurate energy consumption forecasting is becoming increasingly important due to rising global energy demands driven by economic development and population growth. Traditional forecasting models often overlook the impact of contextual factors, such as weather conditions and occupancy trends, which are essential for precise predictions. In this study, we propose a hybrid context-aware simulated scenario generation (CA-SSG) approach that integrates context space theory (CST) with deep learning techniques. This method leverages key contextual features to generate synthetic energy consumption data that more accurately mimics real-world patterns. Using the ASHRAE Great Energy Predictor III dataset, which includes diverse building types across various climates, we demonstrate the effectiveness of CA-SSG. The results show significant improvements in model performance, with reductions in Kullback-Leibler divergence (5%), increases in Pearson Correlation Coefficient (5%), and decreases in computation time compared to traditional approaches. These findings highlight the advantages of contextually enriched generative models for developing smarter energy management systems, enabling more accurate energy forecasting, and supporting strategic planning for energy consumption. Sundas Sarwar, Diletta Chiaro, Edoardo Prezioso, Sara Amitrano, Salvatore Cuomo, Francesco Piccialli |
IEEE Big Data | 2 |
| 2024 | FLOWS: Federated Learning Optimization With SinkhornabstractFederated learning (FL) enables the collaborative training of artificial intelligence models across multiple participating clients while preserving data privacy. Yet, the presence of statistical heterogeneity, characterized by non-independent and non-identically distributed (non-IID) data among clients, poses a significant hurdle in achieving optimal model convergence within the federated setting. In this study, we present FLOWS, a framework that seamlessly incorporates the Sinkhorn distance into each client’s local training process. This integration effectively tackles the well-known challenge by promoting a close alignment between local predictions and the global model’s predictions. Comprehensive experiments across diverse datasets were conducted to evaluate FLOWS’s performance against state-of-the-art FL algorithms. The results indicate that FLOWS enhances the performance of FL models without incurring in a substantial additional computational load. Diletta Chiaro, Fabio Giampaolo, Sara Amitrano, Francesco Piccialli |
ISCC | 1 |
| 2024 | On the Dynamics of Non-IID Data in Federated Learning and High-Performance ComputingabstractThis paper investigates the symbiosis of Federated Learning (FL) and High-Performance Computing (HPC) architectures, unraveling challenges introduced by the intricate interplay of heterogeneity and non-Independently and Identically Distributed (non-lID) data. By leveraging the Flower framework, our research delves into the nuanced implications of FL in diverse HPC environments. We provide a comprehensive exploration of the heterogeneity within contemporary HPC architectures, spanning node organizations, memory hierarchies, and special-ized accelerators, emphasizing adaptability to this complexity. Methodologically, we simulate a FL scenario within our research laboratory, leveraging Flower to orchestrate collaborative model training across heterogeneous nodes. The experiments involve variations in the Dirichlet beta parameter, offering insights into the effects of non-lID data. Results encompass communication efficiency, energy efficiency, and global model accuracy, providing a holistic understanding of the performances across diverse HPC infrastructures. This research contributes to the ongoing discourse on efficient and scalable algorithms, providing insights for collaborative learning in the era of diverse HPC architectures. Daniela Annunziata, Marzia Canzaniello, Diletta Chiaro, Stefano Izzo, Martina Savoia, Francesco Piccialli |
PDP | 3 |
| 2024 | KAFÈ: Kernel Aggregation for FEderated
Pian Qi, Diletta Chiaro, Fabio Giampaolo, Francesco Piccialli |
ECML/PKDD (4) | 2 |
| 2024 | Predictive maintenance for offshore oil wells by means of deep learning features extractionabstractAbstract Nowadays, the great diffusion of the Internet of Things and the improvements in Artificial Intelligence techniques have given a rise in the development and application of data‐driven approaches for Predictive Maintenance to reduce the costs linked to the maintenance of industrial machinery. Due to the wide real‐life applications and the strong interest by even more industries, this field is highly attractive for academics and practitioners. So, constructing efficient frameworks to address the Predictive Maintenance problem is an open debate. In this work, we propose a Deep Learning approach for the feature extraction in the offshore oil wells monitoring context, exploiting the public 3 W dataset, which is well‐known in the literature. The dataset is made up of about 2000 multivariate time series labelled according to the corresponding functioning of the well. So, there is a classification task with eight classes, each related to a particular machinery condition. Thanks to the peculiarities of the labels, the proposed framework is valid both for diagnostics and prognostics. In more detail, we compare two different approaches in feature extraction. The first is a statistical approach, widely used in the literature related to the considered dataset; the second is based on Convolutional 1D AutoEncoder. The extracted features are then used as input for several Machine Learning algorithms, namely the Random Forest, Nearest Neighbours, Gaussian Naive Bayes and Quadratic Discriminant Analysis. Different experiments on various time horizons prove the worthiness of the Convolutional AutoEncoder. Federico Gatta, Fabio Giampaolo, Diletta Chiaro, Francesco Piccialli |
Expert Syst. J. Knowl. Eng. | 3 |
| 2024 | Model aggregation techniques in federated learning: A comprehensive surveyabstractFederated learning (FL) is a distributed machine learning (ML) approach that enables models to be trained on client devices while ensuring the privacy of user data. Model aggregation, also known as model fusion, plays a vital role in FL. It involves combining locally generated models from client devices into a single global model while maintaining user data privacy. However, the accuracy and reliability of the resulting global model depend on the aggregation method chosen, making the selection of an appropriate method crucial. Initially, the simple averaging of model weights was the most commonly used method. However, due to its limitations in handling low-quality or malicious models, alternative techniques have been explored. As FL gains popularity in various domains, it is crucial to have a comprehensive understanding of the available model aggregation techniques and their respective strengths and limitations. However, there is currently a significant gap in the literature when it comes to systematic and comprehensive reviews of these techniques. To address this gap, this paper presents a systematic literature review encompassing 201 studies on model aggregation in FL. The focus is on summarizing the proposed techniques and the ones currently applied for model fusion. This survey serves as a valuable resource for researchers to enhance and develop new aggregation techniques, as well as for practitioners to select the most appropriate method for their FL applications. Pian Qi, Diletta Chiaro, Antonella Guzzo, Michele Ianni, Giancarlo Fortino, Francesco Piccialli |
Future Gener. Comput. Syst. | 2 |
| 2023 | Unveiling engagement in virtual classrooms: a multimodal analysisabstractOnline learning has yielded numerous advantages, notably enhanced accessibility and resource efficiency, which have played a vital role in sustaining educational continuity amidst unprecedented challenges, such as the COVID-19 pandemic. Despite the various benefits and opportunities provided by online learning, many challenges need to be addressed. For instance, the virtual learning environment may introduce potential barriers to effective communication and interaction. It has been established that genuine student engagement is pivotal for effective learning, surpassing the mere availability of high-quality educational materials. Utilizing deep learning (DL) architectures, we harness artificial intelligence (AI) to propose a multimodal approach for assessing and evaluating student engagement in online learning environments. Our results are promising, showcasing the potential impact of AI in enhancing online learning experiences for both students and educators. Additionally, we present an emotion classifier that outperforms the widely recognized DeepFace emotion recognition model on the test set, increasing accuracy from 54% to 72%. We aspire to stimulate further research in this direction, as the ongoing shift towards digital and online learning necessitates innovative solutions to ensure that educational outcomes remain robust and equitable for all learners. Diletta Chiaro, Daniela Annunziata, Stefano Izzo, Francesco Piccialli |
IEEE Big Data | 1 |
| 2023 | A blockchain-based secure Internet of medical things framework for stress detection
Pian Qi, Diletta Chiaro, Fabio Giampaolo, Francesco Piccialli |
Inf. Sci. | 2 |
| 2023 | Statistical arbitrage in the stock markets by the means of multiple time horizons clusteringabstractAbstract Nowadays, statistical arbitrage is one of the most attractive fields of study for researchers, and its applications are widely used also in the financial industry. In this work, we propose a new approach for statistical arbitrage based on clustering stocks according to their exposition on common risk factors. A linear multifactor model is exploited as theoretical background. The risk factors of such a model are extracted via Principal Component Analysis by looking at different time granularity. Furthermore, they are standardized to be handled by a feature selection technique, namely the Adaptive Lasso, whose aim is to find the factors that strongly drive each stock’s return. The assets are then clustered by using the information provided by the feature selection, and their exposition on each factor is deleted to obtain the statistical arbitrage. Finally, the Sequential Least SQuares Programming is used to determine the optimal weights to construct the portfolio. The proposed methodology is tested on the Italian, German, American, Japanese, Brazilian, and Indian Stock Markets. Its performances, evaluated through a Cross-Validation approach, are compared with three benchmarks to assess the robustness of our strategy. Federico Gatta, Carmela Iorio, Diletta Chiaro, Fabio Giampaolo, Salvatore Cuomo |
Neural Comput. Appl. | 3 |
| 2023 | Insight Extraction From E-Health Bookings by Means of Hypergraph and Machine LearningabstractNew technologies are transforming medicine, and this revolution starts with data. Usually, health services within public healthcare systems are accessed through a booking centre managed by local health authorities and controlled by the regional government. In this perspective, structuring e-health data through a Knowledge Graph (KG) approach can provide a feasible method to quickly and simply organize data and/or retrieve new information. Starting from raw health bookings data from the public healthcare system in Italy, a KG method is presented to support e-health services through the extraction of medical knowledge and novel insights. By exploiting graph embedding which arranges the various attributes of the entities into the same vector space, we are able to apply Machine Learning (ML) techniques to the embedded vectors. The findings suggest that KGs could be used to assess patients' medical booking patterns, either from unsupervised or supervised ML. In particular, the former can determine possible presence of hidden groups of entities that is not immediately available through the original legacy dataset structure. The latter, although the performance of the used algorithms is not very high, shows encouraging results in predicting a patient's likelihood to undergo a particular medical visit within a year. However, many technological advances remain to be made, especially in graph database technologies and graph embedding algorithms. Vincenzo Schiano Di Cola, Diletta Chiaro, Edoardo Prezioso, Stefano Izzo, Fabio Giampaolo |
IEEE J. Biomed. Health Informatics | 2 |