VLDB 2026 Research / reviewers in the wild / expert
Giovanni Quattrocchi
dblp:150/6029 · also Giovanni Ennio Quattrocchi
· DBLP profile ↗
36ranked-venue papers
6as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 29 · 6 first-author · 23 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DANTE: Data-Driven Test Case Selection and Prioritization for Long-Running Test Suites
Simone Reale, Elisabetta Di Nitto, Luciano Baresi, Massimiliano Di Penta, Giovanni Quattrocchi |
ICST | 5 |
| 2026 | Can LLMs Generate User Stories and Assess Their Quality?abstractRequirements elicitation is still one of the most challenging activities of the requirements engineering process due to the difficulty requirements analysts face in understanding and translating complex needs into concrete requirements that directly impact the quality of the software to be developed.Although automated tools allow for assessing the syntactic quality of requirements, evaluating semantic metrics (e.g., language clarity, internal consistency) remains a manual and time-consuming activity.This paper explores how LLMs can help automate requirements elicitation within agile frameworks, where requirements are defined as user stories.We used 10 state-of-theart LLMs to investigate their ability to generate user stories automatically by emulating customer interviews.We evaluated the quality of user stories generated by LLMs, comparing it with the quality of user stories generated by humans (domain experts and students).We also explored whether and how LLMs can be used to automatically evaluate the semantic quality of user stories.Our results indicate that LLMs can generate user stories similar to humans in terms of coverage and stylistic quality, but exhibit lower diversity and creativity.Although LLMgenerated user stories are generally comparable in quality to those created by humans, they tend to meet the acceptance quality criteria less frequently, regardless of the scale of the LLM model.Finally, LLMs can reliably assess the semantic quality of user stories when provided with clear evaluation criteria and have the potential to reduce human effort in large-scale assessments. Giovanni Quattrocchi, Liliana Pasquale, Paola Spoletini, Luciano Baresi |
IEEE Trans. Software Eng. | 1 |
| 2025 | Runtime defect prediction of industrial business processes: A focused look at real-life SAP systemsabstractBusiness process operations are the dominant logic underpinning most of the service-based applications currently in use. Situated in the field of SAP business processes — commonly referred to as iFlows — and their integration, this paper looks into the defectiveness of such flows with a Machine-Learning approach. We propose to cluster and classify at runtime the Integration Flows of business processes during their orchestration; we do so by using metrics extracted from the Integration of 400+ complex business interaction and service orchestration Flows along with their metadata. Through a combined ensemble-based, clustering, and supervised learning exercise, we conclude that an AI-based approach for runtime defect prediction of iFlows shows considerable promise in providing actionable insights for better orchestration intelligence, especially in sight of self-aware business processes of the future. Max Nijholt, Giovanni Quattrocchi, Damian A. Tamburri |
J. Syst. Softw. | 2 |
| 2025 | Engineering MLOps Pipelines With Data Quality: A Case Study on Tabular Datasets in KaggleabstractABSTRACT Ensuring high‐quality data is crucial for the successful deployment of machine learning models, thereby sustaining the operational pipelines around such models. However, a significant number of practitioners do not currently use data quality checks or measurements as gateways for their model construction and operationalization, indicating a need for greater awareness and adoption of these tools. In this study, we propose an automated approach for automating the process of architecting machine learning pipelines by means of (semi‐)automated data quality checks. We focus on tabular data as a representative of the most widely used structured data formats in said pipelines. Our work is based on a subset of metrics that are particularly relevant in MLOps pipelines, stemming from our engagement with expert practitioners in machine learning operations (MLOps). We selected Deepchecks, a well‐known tool for conducting data quality checks, from a cohort of similar tools to evaluate the quality of datasets collected from Kaggle, a widely used platform for machine learning competitions and data science projects. We also analyze the main features used by Kaggle to rank their datasets and used these features to validate the relevance of our approach. Our approach shows the potential for automated data quality checks to improve the efficiency and effectiveness of MLOps pipelines and their operation, by decreasing the risk of introducing errors and biases into machine learning models in production. Matteo Pancini, Matteo Camilli, Giovanni Quattrocchi, Damian A. Tamburri |
J. Softw. Evol. Process. | 3 |
| 2024 | Copula-based Approaches for Anomaly Detection: a Case-study in Financial ForensicsabstractFraud detection is a critical challenge in various domains, necessitating accurate and reliable methods to distinguish between legitimate and fraudulent transactions. This work explores the application of copula-based models for anomaly detection in financial forensics. It focuses on their effectiveness in identifying fraudulent activities in a highly imbalanced dataset. Copula models are designed to capture the dependencies between continuous variables, providing a flexible framework for modeling the joint distribution of features. For instance, the variables in a dataset can follow different distributions and the copula is able to model how these variables jointly behave, particularly in extreme cases. In this work, we used a fraud dataset to calculate the copula-based probability of fraud and conditional Gaussian copula. Then, we derive the copula-based Generalized Linear Model (GLM) formula from the conditional copula which is essentially a GLM with probit link and transformed variables when the covariates are continuous. Finally, we compare the performance of these copula-based models with standard methods for predicting a binary variable like GLM with a probit link and logistic regression. Results indicate that copula-based probability and conditional copula formulas offer promising results, particularly in handling complex dependencies, but with a high computational time, while copula-based GLM, when combined with over-sampling, also outperforms traditional methods. Vanessa Tenconi, Damian A. Tamburri, Giovanni Quattrocchi, Corrado Pellegrino, Giuseppe Cascavilla, Willem-Jan van den Heuvel |
IEEE Big Data | 3 |
| 2024 | Efficient and Dependency-Aware Placement of Serverless Functions on Edge Infrastructures
Luciano Baresi, Giovanni Quattrocchi, Inacio Gaspar Ticongolo |
ICSOC (1) | 2 |
| 2024 | A Conceptual Framework for Quality Assurance of LLM-based Socio-critical SystemsabstractRecent breakthroughs in Artificial Intelligence (AI) obfuscate the boundaries between digital, physical, and social spaces, a trend expected to continue in the foreseeable future. Traditionally, software engineering has prioritized technical aspects, focusing on functional correctness and reliability while often neglecting broader societal implications. With the rise of software agents enabled by Large Language Models (LLMs) and capable of emulating human intelligence and perception, there is a growing recognition of the need for addressing socio-critical issues. Unlike technical challenges, these issues cannot be resolved through traditional, deterministic approaches due to their subjective nature and dependence on evolving factors such as culture and demographics. This paper dives into this problem and advocates the need for revising existing engineering principles and methodologies. We propose a conceptual framework for quality assurance where AI is not only the driver of socio-critical systems but also a fundamental tool in their engineering process. Such framework encapsulates pre-production and runtime workflows where LLM-based agents, so-called artificial doppelgängers, continuously assess and refine socio-critical systems ensuring their alignment with established societal standards. Luciano Baresi, Matteo Camilli, Tommaso Dolci, Giovanni Quattrocchi |
ASE | 4 |
| 2024 | A qualitative and quantitative analysis of container enginesabstractContainerization is a virtualization technique that allows one to create and run executables consistently on any infrastructure. Compared to virtual machines, containers are lighter since they do not bundle a (guest) operating system but they share its kernel, and they only include the files, libraries, and dependencies that are required to properly execute a process. In the past few years, multiple container engines (i.e., tools for configuring, executing, and managing containers) have been developed ranging from some that are “general purpose”, and mostly employed for Cloud executions, to others that are built for specific contexts, namely Internet of Things and High-Performance Computing. Given the importance of this technology for many practitioners and researchers, this paper analyses six state-of-the-art container engines and compares them through a comprehensive study of their characteristics and performance. The results are organized around 0 findings that aim to help the readers understand the differences among the technologies and help them choose the best approach for their needs. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Luciano Baresi, Giovanni Quattrocchi, Nicholas Rasi |
J. Syst. Softw. | 2 |
| 2024 | NEPTUNE: A Comprehensive Framework for Managing Serverless Functions at the EdgeabstractApplications that are constrained by low-latency requirements can hardly be executed on cloud infrastructures, given the high network delay required to reach remote servers. Multi-access Edge Computing (MEC) is the reference architecture for executing applications on nodes that are located close to users (i.e., at the edge of the network). This way, the network overhead is reduced but new challenges emerge. The resources available on edge nodes are limited, workloads fluctuate since users can rapidly change location, and complex tasks are becoming widespread (e.g., machine learning inference). To address these issues, this article presents NEPTUNE , a serverless-based framework that automates the management of large-scale MEC infrastructures. In particular, NEPTUNE provides (i) the placement of serverless functions on MEC nodes according to users’ location, (ii) the resolution of resource contention scenarios by avoiding that single nodes be saturated, and (iii) the dynamic allocation of CPUs and GPUs to meet foreseen execution times. To assess NEPTUNE , we built a prototype based on K3S, an edge-dedicated version of Kubernetes, and executed a comprehensive set of experiments. Results show that NEPTUNE obtains a significant reduction in terms of response time, network overhead, and resource consumption compared with five state-of-the-art solutions. Luciano Baresi, Davide Yi Xian Hu, Giovanni Quattrocchi, Luca Terracciano |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2024 | Cromlech: Semi-Automated Monolith Decomposition Into MicroservicesabstractMicroservices architectures conceive an application as a composition of loosely-coupled sub-systems that are developed, deployed, maintained, updated, and scaled independently. Compared to monoliths, microservices speed up evolution and increase flexibility. For these reasons they are becoming the reference architecture for many practitioners. A key challenge to embrace a microservices architecture is how to decompose an application into microservices: a choice that deeply affects all subsequent development phases in ways that are difficult to foresee and evaluate. Without any tool to support their reasoning, developers may erroneously evaluate the various alternatives, leading to inaccurate decomposition choices that would result in increased development, operations, and maintenance costs. This paper tackles the problem with Cromlech, a semi-automatic tool to decompose a software system into microservices. Cromlech (i) takes in input a high-level model of the system in terms of functionalities and data entities accessed by those functionalities, (ii) formulates decomposition as an optimization problem, and (iii) outputs a proposed placement of functionalities and data onto microservices, using a visual representation that helps reasoning on the resulting architecture. Cromlech evaluates design concerns, communication overheads, data management requirements, opportunities and costs of data replication. Our evaluation on a real-world industrial application shows that Cromlech consistently delivers more efficient solutions than simple heuristics and state-of-the-art approaches, and provides useful insights to developers. Giovanni Quattrocchi, Davide Cocco, Simone Staffa, Alessandro Margara, Gianpaolo Cugola |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Autoscaling Solutions for Cloud Applications Under Dynamic WorkloadsabstractAutoscaling systems provide means to automatically change the resources allocated to a software system according to the incoming workload and its actual needs. Public cloud providers offer a variety of autoscaling solutions, ranging from those based on user-written rules to more sophisticated ones. Originally, these solutions were conceived to manage clusters of virtual machines, while more recently, they have also been employed in the operation of containers. This paper analyses the autoscaling solutions provided by three major cloud providers, namely Amazon Web Services, Google Cloud Platform, and Microsoft Azure, and compares them against two solutions we develop based on control theory (ScaleX) and queuing theory (QN-CTRL). We evaluate the different approaches using both an in-house simulation engine and cloud deployments by feeding them with various synthetic and real-world workloads. Our extensive evaluation collects both simulation results and real measurements by which we can assess that both scaleX and QN-CTRL outperform industrial techniques in most cases when considering the trade-offs between the service-level-agreement (SLA) violations and the optimal usage of resources. Giovanni Quattrocchi, Emilio Incerto, Riccardo Pinciroli, Catia Trubiani, Luciano Baresi |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Dependency-Aware Resource Allocation for Serverless Functions at the Edge
Luciano Baresi, Giovanni Quattrocchi, Inacio Gaspar Ticongolo |
ICSOC (1) | 2 |
| 2023 | A multi-faceted analysis of the performance variability of virtual machinesabstractAbstract Cloud computing and virtualization solutions allow one to rent the virtual machines (VMs) needed to run applications on a pay‐per‐use basis, but rented VMs do not offer any guarantee on their performance. Cloud platforms are known to be affected by performance variability , but a better understanding is still required. This article moves in that direction and presents an in‐depth, multi‐faceted study on the performance variability of VMs. Unlike previous studies, our assessment covers a wide range of factors: 16 VM types from 4 well‐known cloud providers, 10 benchmarks, and 28 different metrics. We present four new contributions. First, we introduce a new benchmark suite ( VMBS ) that let researchers and practitioners systematically collect a diverse set of performance data. Second, we present a new indicator, called V I , that allows for measuring variability in the performance of VMs. Third, we illustrate an analysis of the collected data across four different dimensions: resources , isolation , time , and cost . Fourth, we present multiple predictive models based on machine learning (ML) that aim to forecast future performance and detect time patterns. Our experiments provide important insights on the resource variability of VMs, highlighting differences and similarities between various cloud providers. To the best of our knowledge, this is the widest analysis ever conducted on the topic. Luciano Baresi, Tommaso Dolci, Giovanni Quattrocchi, Nicholas Rasi |
Softw. Pract. Exp. | 3 |
| 2023 | Making service continuity smarter with artificial intelligence: An approach and its evaluationabstractAbstract Service continuity entails establishing an observable and explainable continuum between customer experience and service operations. Such continuum is currently established manually, via service customer management operations (such as service incident management (IM)) often resulting in time‐consuming, human‐detrimental, and error‐prone activities. Conversely, artificial intelligence (AI) is rapidly emerging as an automated enabler towards handling the discontinuities in the aforementioned critical business tasks. Consequently, the emerging topic of AI‐driven incident management (AIIM) addresses practices and tools to resolve incidents by means of AI‐enabled organizational processes and methodologies. Our conjecture is that AIIM could reduce unplanned interruptions of service and let customers resume their work as quick as possible. While several techniques were presented in the literature to automatically identify the problems described in incident tickets by customers, this article focuses on the qualitative analysis and feature extraction off of the provided descriptions. When an incident ticket does not describe properly the problem, the analyst must ask the customer for additional details which could require several long‐lasting interactions. This article proposes ACQUA , an AIIM approach to automatically assess the quality of ticket descriptions with the goals of removing the need of additional communications and guiding the customers to properly describe the incident. A preliminary evaluation of ACQUA was performed on a dataset provided by a large bank in Europe, showing promising results and a boost of 13% in ticket resolution times and connected service continuity. Giovanni Quattrocchi, Damian A. Tamburri, Willem-Jan van den Heuvel |
Softw. Pract. Exp. | 1 |
| 2022 | Training and Serving Machine Learning Models at Scale
Luciano Baresi, Giovanni Quattrocchi |
ICSOC | 2 |
| 2022 | DeepThought: A Reputation and Voting-Based Blockchain Oracle
Marco Di Gennaro 0001, Lorenzo Italiano, Giovanni Meroni, Giovanni Quattrocchi |
ICSOC | 4 |
| 2022 | Blockchain-Oriented Services Computing in Action: Insights from a User Study
Giovanni Quattrocchi, Damian A. Tamburri, Willem-Jan van den Heuvel |
ICSOC | 1 |
| 2022 | A declarative modelling framework for the deployment and management of blockchain applicationsabstractThe deployment and management of Blockchain applications require non-trivial efforts given the unique characteristics of their infrastructure (i.e., immutability) and the complexity of the software systems being executed. The operation of Blockchain applications is still based on ad-hoc solutions that are error-prone, difficult to maintain and evolve, and do not manage their interactions with other infrastructures (e.g., a Cloud backend). Luciano Baresi, Giovanni Quattrocchi, Damian A. Tamburri, Luca Terracciano |
MoDELS | 2 |
| 2022 | NEPTUNE: Network- and GPU-aware Management of Serverless Functions at the EdgeabstractNowadays a wide range of applications is constrained by low-latency requirements that cloud infrastructures cannot meet. Multiaccess Edge Computing (MEC) has been proposed as the reference architecture for executing applications closer to users and reducing latency, but new challenges arise: edge nodes are resource-constrained, the workload can vary significantly since users are nomadic, and task complexity is increasing (e.g., machine learning inference). To overcome these problems, the paper presents NEPTUNE, a serverless-based framework for managing complex MEC solutions. NEPTUNE i) places functions on edge nodes according to user locations, ii) avoids the saturation of single nodes, iii) exploits GPUs when available, and iv) allocates resources (CPU cores) dynamically to meet foreseen execution times. A prototype, built on top of K3S, was used to evaluate NEPTUNE on a set of experiments that demonstrate a significant reduction in terms of response time, network overhead, and resource consumption compared to three well-known approaches. Luciano Baresi, Davide Yi Xian Hu, Giovanni Quattrocchi, Luca Terracciano |
SEAMS | 3 |
| 2022 | About the special issue on: "Distributed Complex Systems: Governance, Engineering, and Maintenance"abstractAbstract The volume at hand presents the Special Issue on “Distributed Complex Systems: Governance, Engineering, and Maintenance”. The Special Issue has been originally conceived within the context of the 2nd International Workshop on Governing Adaptive and Unplanned Systems of Systems (GAUSS 2020), one of the co‐located events of the 31st International Symposium on Software Reliability Engineering (ISSRE 2020). The authors of the best papers at GAUSS 2020 have been invited to submit an extended version of their previous work. In addition, the editors opened the submission also to all the other researches working on technical and managerial solution for governing Distributed Complex Systems. Ultimate objective of the editors with this Special Issue is to promote discussions focusing on anticipating, mitigating, or reacting to scenarios that were unplanned or under‐specified at design time. Pietro Braione, Daniela Briola, Guglielmo De Angelis, Francesco Gallo, Francesco Poggi, Giovanni Quattrocchi |
J. Softw. Evol. Process. | 6 |
| 2022 | Predictive maintenance of infrastructure code using "fluid" datasets: An exploratory study on Ansible defect pronenessabstractAbstract This work consolidates and compounds previous investigations in recognizing defects for infrastructure‐as‐code (IaC) scripts using general software development quality metrics with a focus on defect severity but adding to previous work an explorative look at creating datasets, which may boost the predictive power of provided models—we call this notion a fluid dataset. More specifically, we experiment with 50 different metrics harnessing a multiple dataset creation process whereby different versions of the same datasets are rigged with auto‐training facilities for model retraining and redeployment in a DataOps fashion. At this point, with a focus on the Ansible infrastructure code language—as a de facto standard for industrial‐strength infrastructure code—we build defect prediction models and manage to improve on the state of the art by finding an F1 score of 0.52 and a recall of 0.57 using a Naive–Bayes classifier. On the one hand, by improving state‐of‐the‐art defect prediction models using metrics generalizable for different IaC languages, we provide interesting leads for the future of infrastructure‐as‐code. On the other hand, we have barely scratched the surface on the novel approach of fluid‐datasets creation and automated retraining of Machine Learning (ML) defect prediction models, warranting for more research on the same direction in the future. Giovanni Quattrocchi, Damian A. Tamburri |
J. Softw. Evol. Process. | 1 |
| 2021 | KOSMOS: Vertical and Horizontal Resource Autoscaling for Kubernetes
Luciano Baresi, Davide Yi Xian Hu, Giovanni Quattrocchi, Luca Terracciano |
ICSOC | 3 |
| 2021 | Resource Management for TensorFlow Inference
Luciano Baresi, Giovanni Quattrocchi, Nicholas Rasi |
ICSOC | 2 |
| 2021 | Pangaea: Semi-automated Monolith Decomposition into Microservices
Simone Staffa, Giovanni Quattrocchi, Alessandro Margara, Gianpaolo Cugola |
ICSOC | 2 |
| 2021 | Intelligent re-deployment feedback loop for hybrid applicationsabstractWe propose enabling continuous performance optimisation of distributed hybrid applications in heterogeneous cloud, Edge, and HPC environments by employing an intelligent re-deployment feedback loop. Kalman Z. Meth, Indika Kumara, Giovanni Quattrocchi |
SYSTOR | 3 |
| 2021 | SODALITE@RT: Orchestrating Applications on Cloud-Edge InfrastructuresabstractAbstract IoT-based applications need to be dynamically orchestrated on cloud-edge infrastructures for reasons such as performance, regulations, or cost. In this context, a crucial problem is facilitating the work of DevOps teams in deploying, monitoring, and managing such applications by providing necessary tools and platforms. The SODALITE@RT open-source framework aims at addressing this scenario. In this paper, we present the main features of the SODALITE@RT: modeling of cloud-edge resources and applications using open standards and infrastructural code, and automated deployment, monitoring, and management of the applications in the target infrastructures based on such models. The capabilities of the SODALITE@RT are demonstrated through a relevant case study. Indika Kumara, Paul Mundt, Kamil Tokmakov, Dragan Radolovic, Alexander Maslennikov, Román Sosa, Jorge Fernández-Fabeiro, Giovanni Quattrocchi, Kalman Z. Meth, Elisabetta Di Nitto, Damian A. Tamburri, Willem-Jan van den Heuvel, Georgios Meditskos |
J. Grid Comput. | 8 |
| 2021 | Fine-Grained Dynamic Resource Allocation for Big-Data ApplicationsabstractMany big-data applications are batch applications that exploit dedicated frameworks to perform massively parallel computations across clusters of machines. The time needed to process the entirety of the inputs represents the application's response time, which can be subject to deadlines. Spark, probably the most famous incarnation of these frameworks today, allocates resources to applications statically at the beginning of the execution and deviations are not managed: to meet the applications' deadlines, resources must be allocated carefully. This paper proposes an extension to Spark, called dynaSpark, that is able to allocate and redistribute resources to applications dynamically to meet deadlines and cope with the execution of unanticipated applications. This work is based on two key enablers: containers, to isolate Spark's parallel executors and allow for the dynamic and fast allocation of resources, and control-theory to govern resource allocation at runtime and obtain required precision and speed. Our evaluation shows that dynaSpark can (i) allocate resources efficiently to execute single applications with respect to set deadlines and (ii) reduce deadline violations (w.r.t. Spark) when executing multiple concurrent applications. Luciano Baresi, Alberto Leva, Giovanni Quattrocchi |
IEEE Trans. Software Eng. | 3 |
| 2020 | COCOS: A Scalable Architecture for Containerized Heterogeneous SystemsabstractNowadays software systems are organized around several and heterogeneous components. For example, a modern application can be composed of different microservices, along with dedicated components for machine learning analytics and recurring batch processing jobs. While containers offer a means to deploy the system and tackle heterogeneity, these components have different execution models, can exploit different resource types (e.g., CPUs and GPUs) and result in completely different execution times (milliseconds vs hours). This complexity calls for a new, scalable architecture to allow the systems to operate efficiently. This paper presents COCOS, an architecture, based on containers and control-theory, that is able to manage large and heterogeneous software systems. The architecture is based on a three-level hierarchy of controllers that cooperatively enforce user-defined requirements on execution times and consumed resources. The paper also shows a prototype implementation of COCOS based on Kubernetes, a well-known container orchestrator. The evaluation shows the efficiency of COCOS when dealing with microservices, Spark jobs and machine learning applications. Luciano Baresi, Giovanni Quattrocchi |
ICSA | 2 |
| 2020 | Automated Quality Assessment of Incident Tickets for Smart Service Continuity
Luciano Baresi, Giovanni Quattrocchi, Damian A. Tamburri, Willem-Jan van den Heuvel |
ICSOC | 2 |
| 2020 | A Simulation-based Comparison between Industrial Autoscaling Solutions and COCOS for Cloud ApplicationsabstractDynamic resource allocation is the mechanism that allows one to change the resources associated with applications at runtime and match their actual needs. The autoscaling solutions offered by cloud infrastructures are probably the most widely-used incarnation of this concepts. Originally conceived to manage virtual machines according to user-defined rules, they are now much more sophisticated and can also allocate containers (lighter than virtual machines). This paper surveys the autoscaling solutions provided by the major cloud vendors and analyzes the services they provide. It also compares them against the solution we developed, called COCOS autoscaling. We simulated the different proposals and fed them with diverse workloads. Obtained results show that COCOS autoscaling outperforms its competitors in most of the cases: it optimizes resource allocation and keeps applications' response times under set thresholds. Luciano Baresi, Giovanni Quattrocchi |
ICWS | 2 |
| 2020 | Using formal verification to evaluate the execution time of Spark applicationsabstractAbstract Apache Spark is probably the most widely adopted framework for developing big-data batch applications and for executing them on a cluster of (virtual) machines. In general, the more resources (machines) one uses, the faster applications execute, but there is currently no adequate means to determine the proper size of a Spark cluster given time constraints, or to foresee execution times given the number of employed machines. One can only run these applications and use her/his experience to size the cluster and predict expected execution times. Wrong estimation of execution times can lead to costly overruns and overly long executions, thus calling for analytic sizing/prediction techniques that provide precise time guarantees. This paper addresses this problem by proposing a solution based on model-checking. The approach exploits a directed acyclic graph (DAG) to abstract the structure of the execution flows of Spark programs, annotates each node (Spark stage) with execution-related data, and formulates the identification of the global execution time as a reachability problem. To avoid the well-known state space explosion problem, the paper also proposes a technique to reduce the size of generated abstract models. This results in a significant decrease in used memory and/or verification time making our approach feasible for predicting the execution time of Spark applications given the resources available. The benefits of the proposed reduction technique are evaluated by using both timed automata and constraint LTL over clocks logic to formally encode and analyze generated models. The approach is also successfully validated on some realistic case studies. Since the optimization is not Spark-specific, we claim that it can be applied to a wide range of applications whose underlying model can be abstracted as a DAG. Luciano Baresi, Marcello M. Bersani, Francesco Marconi, Giovanni Quattrocchi, Matteo G. Rossi |
Formal Aspects Comput. | 4 |
| 2019 | PAPS: A Framework for Decentralized Self-management at the Edge
Luciano Baresi, Danilo Filgueira Mendonça, Giovanni Quattrocchi |
ICSOC | 3 |
| 2019 | Symbolic execution-driven extraction of the parallel execution plans of Spark applicationsabstractThe execution of Spark applications is based on the execution order and parallelism of the different jobs, given data and available resources. Spark reifies these dependencies in a graph that we refer to as the (parallel) execution plan of the application. All the approaches that have studied the estimation of the execution times and the dynamic provisioning of resources for this kind of applications have always assumed that the execution plan is unique, given the computing resources at hand. This assumption is at least simplistic for applications that include conditional branches or loops and limits the precision of the prediction techniques. Luciano Baresi, Giovanni Denaro, Giovanni Quattrocchi |
ESEC/SIGSOFT FSE | 3 |
| 2019 | A Unified Model for the Mobile-Edge-Cloud ContinuumabstractTechnologies such as mobile, edge, and cloud computing have the potential to form a computing continuum for new, disruptive applications. At runtime, applications can choose to execute parts of their logic on different infrastructures that constitute the continuum, with the goal of minimizing latency and battery consumption and maximizing availability. In this article, we propose A3-E, a unified model for managing the life cycle of continuum applications. In particular, A3-E exploits the Functions-as-a-Service model to bring computation to the continuum in the form of microservices. Furthermore, A3-E selects where to execute a certain function based on the specific context and user requirements. The article also presents a prototype framework that implements the concepts behind A3-E. Results show that A3-E is capable of dynamically deploying microservices and routing the application’s requests, reducing latency by up to 90% when using edge instead of cloud resources, and battery consumption by 74% when computation has been offloaded. Luciano Baresi, Danilo Filgueira Mendonça, Martin Garriga, Sam Guinea, Giovanni Quattrocchi |
ACM Trans. Internet Techn. | 5 |
| 2016 | A discrete-time feedback controller for containerized cloud applicationsabstractModern Web applications exploit Cloud infrastructures to scale their resources and cope with sudden changes in the workload. While the state of practice is to focus on dynamically adding and removing virtual machines, we advocate that there are strong benefits in containerizing the applications and in scaling the containers. Luciano Baresi, Sam Guinea, Alberto Leva, Giovanni Quattrocchi |
SIGSOFT FSE | 4 |
| 2015 | Wind fields from COSMO-SkyMed and RADARSAT-2 SAR in coastal areasabstractThis paper is aimed to outline the possible uses of SAR derived wind fields in coastal meteorology, taking advantage of two running projects: the first of meteorological flavor, aimed to measure and reproduce the surface wind field in a small coastal area; the second to investigate the possibilities offered by the combined use of simultaneous COSMO-SkyMed and RADARSAT-2 SAR images in the description of the wind field in the same area. The main issues faced in this study and partially solved are: the development of a suitable methodology to extract the wind field from SAR images at different bands, polarization, incidence angle; the understanding of the potentialities of the SAR derived winds to catch the patterns of the spatial variability of the wind over the coastal area; the assessment of the SAR derived wind from comparisons with the simultaneous experimental data. Of course, the results presented here are not exhaustive, but only aim to track the ongoing research on the topic of wind over local coastal areas. Stefano Zecchetto, Francesco De Biasio, Antonio Della Valle, Andrea Cucco, Giovanni Quattrocchi, Enrico Cadau |
IGARSS | 5 |