Chiara Rucco

dblp:323/1318 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-4067-0955ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Data Ingestion Efficiency in Cloud-Based Systems: A Design Pattern Approach
abstract
Abstract This paper aims to define design patterns specifically for data ingestion techniques within cloud-based architectures, addressing the challenges associated with high-volume data processing. The approach utilizes a flexible, metadata-driven framework that enhances adaptability and ease of use. This framework supports both incremental and full refresh methods, allowing for seamless changes to ingestion types, schema updates, table additions, and the incorporation of new data sources with minimal intervention from data engineers. The proposed design patterns were validated through experiments conducted on the Azure and Google Cloud platforms. The experiments demonstrate that the proposed design patterns significantly reduce data ingestion time, showcasing their effectiveness in managing high-volume data ingestion. This paper contributes to the field of data management by presenting a comprehensive definition of design patterns tailored for data ingestion in cloud-based architectures, effectively addressing key challenges in high-volume data processing.
Chiara Rucco, Antonella Longo, Motaz Saad
Data Sci. Eng.1
2026 Empirical evaluation of LLMs capabilities for data pipeline generation on Databricks platform
abstract
Building production-ready data pipelines on enterprise platforms requires expertise spanning distributed computing, data integration, and platform-specific tooling — a process that Large Language Models (LLMs) may help accelerate. This paper presents an empirical evaluation of seven leading LLMs — from the GPT-5, Claude 4, and Qwen3 families — on their ability to generate functional PySpark data pipelines for the Databricks platform. Each model was evaluated across three progressively complex ETL scenarios (data quality validation, temporal aggregation, and multi-source integration) using a fix-count metric measuring iterative LLM-driven error corrections. Results show that Claude Opus 4 and GPT-5 mini were the most reliable, requiring zero or one fix across all prompts, while Claude 3.5 Haiku required up to eight fixes on the most complex scenario. Code style compliance scores (Pylint/PEP8) proved orthogonal to functional correctness, with some models producing fully functional pipelines despite low style scores. These findings provide practitioners with empirical guidance on model selection for LLM-assisted pipeline development, while identifying current capability boundaries that require human oversight.
Tobia Martina, Motaz Saad, Chiara Rucco, Antonella Longo
Future Gener. Comput. Syst.3
2025 Multi-Agent Intelligence for e-Tourism: Lecce use Case and Enabling Digital Twins
Veronica Cretì, Chiara Rucco, Alessandro Stefano, Francesca Zampino, Motaz Saad, Antonella Longo, Alì Aghazadeh Ardebili
IEEE Big Data2
2024 Optimizing Data Ingestion for Big Data: A Cloud-Based Design Pattern Approach
abstract
The rapid growth of data has created significant challenges in managing and leveraging data effectively. Data engineering has emerged as a crucial discipline to address these challenges, providing frameworks for efficient data management. Data Engineering Patterns (DEP) and Data Engineering Design Patterns (DEDP) offer standardized practices and best practice solutions for data engineering tasks like ETL. While various DEPs and DEDPs exist, the issue of high-volume data ingestion remains insufficiently addressed. This paper focuses on defining design patterns specifically for data ingestion techniques within cloud-based architectures, covering both incremental and full refresh methods. The proposed approach utilizes a flexible, metadata-driven framework to enhance adaptability and ease of use, allowing for seamless changes to the ingestion type, schema updates, table additions, and incorporation of new data sources. Validated on the Azure cloud platform, the experiments demonstrate that the proposed design patterns significantly reduce data ingestion time, contributing to the field of data management by addressing key challenges in high-volume data processing.
Chiara Rucco, Antonella Longo, Motaz Saad
IEEE Big Data1