VLDB 2026 Research / reviewers in the wild / expert
Antonio Guerriero
dblp:230/2158
· DBLP profile ↗
19ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-8104-3832ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 2 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On-device training and pruning for energy saving and continuous learning in resource-constrained MCUs
Pietro Fusco, Gennaro Pio Rimoli, Antonio Guerriero, Francesco Palmieri 0002, Massimo Ficco |
Future Gener. Comput. Syst. | 3 |
| 2026 | Multivariate anomaly detection and root cause analysis of energy issues in microservice-based systemsabstractContext: Microservice-based systems have become the architecture style of choice for modern applications, offering scalability, flexibility, and resilience. However, their distributed nature leads to increased resource consumption and energy inefficiencies, posing challenges for maintaining sustainable operations. Accurate anomaly detection (AD) and root cause analysis (RCA) tools are critical for diagnosing energy consumption issues in these systems, yet existing solutions often lack focus on energy metrics. Goal: This study aims to evaluate the effectiveness of AD and RCA algorithms in identifying and diagnosing performance-related energy consumption anomalies in microservice-based systems. Method: Two representative systems, Sock Shop and Train Ticket, are deployed under controlled environments. Then, anomalies are deliberately introduced by stressing at the same time CPU, memory, and disk resources. The data collection is conducted using Prometheus for performance metrics and Scaphandre for energy metrics. Once normal and anomalous datasets are constructed for each system, the study evaluates five AD algorithms (Birch, iForest, KNN, LOF, and SVM) and four RCA algorithms (MicroRCA, CausalRCA, CIRCA, and RCD) based on their precision, recall, and scalability across varied scenarios and workloads. Results: The experiment reveals that overall, iForest is the most effective AD algorithms in detecting energy anomalies (0.59 F-Score in Sock Shop and 0.634 F-Score in Train Ticket). In particular, iForest performs better in precision when the user load is high (1000 concurrent users). For RCA, CIRCA performs well in identifying root causes in smaller systems, while RCD is more scalable for larger and more complex systems. Conclusions: The findings of this study provide insights for both researchers and practitioners. In the context of our experiment, AD algorithms tend to perform relatively well, whereas RCA algorithms tend to be imprecise in localizing energy issues. Berta Rodriguez Sanchez, Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Ivano Malavolta |
J. Syst. Softw. | 3 |
| 2025 | Adaptive Probabilistic Operational Testing for Large Language Models EvaluationabstractLarge Language Models (LLM) empower many modern software systems, and are required to be highly accurate and reliable. Evaluating LLM poses challenges due to the high costs of manual labeling and of validation of labeled data.This study investigates the suitability of probabilistic operational testing for effective and efficient evaluation of LLM, focusing on a case study with DistilBERT. To this aim, we adopt an existing framework (DeepSample) for Deep Neural Network (DNN) testing and adapt it to the LLM domain by introducing auxiliary variables tailored to LLM and classification tasks.Through a comprehensive evaluation, we demonstrate how sampling-based operational testing can yield reliable LLM accuracy estimates and effectively expose failures, or, under testing budget constraints, it can find a trade off between accuracy estimation and failure exposure. The experimental results, using DistilBERT on three sentiment analysis datasets, show that sampling-based methods can provide cost effective and reliable operational accuracy assessment for LLM. These findings offer practical insights for testers and help address critical gaps in current LLM evaluation practices. Ali Asgari, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
AST | 2 |
| 2025 | Learning-based Automated Generation of Critical Workload Configurations for Microservices Performance TestingabstractPerformance testing is an essential activity in the engineering of microservice applications to identify deviations from the specified ranges of relevant metrics and to analyse resources usage. It demands for high automation to fit within the short microservices development-operation cycles. Engineers are often interested in identifying critical workloads - ideally, in the “minimal“ load configurations causing tests to expose performance issues. Triggering performance issues is challenging, requiring proper workload characterization and test design. We present the microWave framework for learning-based automated generation of critical performance testing workloads for microservices. The framework can harness various learning strategies: we analyze a Deep Neural Network, a Large Language Model and a Causal Reasoning strategy. We evaluate them experimentally on four subjects, using a random approach and a manually-crafted ground truth as baselines. The results show that the strategies exhibit different behavior depending on the data they learn from. When inferring from past executions data including performance issues, the causal model performs better. The random predictor is preferable when no data is available; however, it is more costly as it requires more tests. The results allow to draw practical recommendations for testers on how to select the most suitable strategy depending on the needs. Cristian Mascia, Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ICWS | 3 |
| 2025 | Causal reasoning in Software Quality Assurance: A systematic reviewabstractContext: Software Quality Assurance (SQA) is a fundamental part of software engineering to ensure stakeholders that software products work as expected after release in operation. Machine Learning (ML) has proven to be able to boost SQA activities and contribute to the development of quality software systems. In this context, Causal Reasoning is gaining increasing interest as a methodology to go beyond a purely data-driven approach by exploiting the use of causality for more effective SQA strategies. Objective: Provide a broad and detailed overview of the use of causal reasoning for SQA activities, in order to support researchers to access this research field, identifying room for application, main challenges and research opportunities. Methods: A systematic review of the scientific literature on causal reasoning for SQA. The study has found, classified, and analyzed 86 articles, according to established guidelines for software engineering secondary studies. Results: Results highlight the primary areas within SQA where causal reasoning has been applied, the predominant methodologies used, and the level of maturity of the proposed solutions. Fault localization is the activity where causal reasoning is more exploited, especially in the web services/microservices domain, but other tasks like testing are rapidly gaining popularity. Both causal inference and causal discovery are exploited, with the Pearl’s graphical formulation of causality being preferred, likely due to its intuitiveness. Tools to favor their application are appearing at a fast pace — most of them after 2021. Conclusions: The findings show that causal reasoning is a valuable means for SQA tasks with respect to multiple quality attributes , especially during V&V, evolution and maintenance to ensure reliability, while it is not yet fully exploited for phases like requirements engineering and design. We give a picture of the current landscape, pointing out exciting possibilities for future research. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
Inf. Softw. Technol. | 2 |
| 2024 | Identifying Performance Issues in Microservice Architectures through Causal ReasoningabstractEvaluating the performance of Microservices Architectures (MSA) is essential to ensure their proper functioning and meet end-user satisfaction. For MSA performance analysts, one of the most challenging tasks is to determine the cause of any deviation of relevant metrics from the specified range. Luca Giamattei, Antonio Guerriero, Ivano Malavolta, Cristian Mascia, Roberto Pietrantuono, Stefano Russo 0001 |
AST | 2 |
| 2024 | DeepSample: DNN sampling-based testing for operational accuracy assessmentabstractDeep Neural Networks (DNN) are core components for classification and regression tasks of many software systems. Companies incur in high costs for testing DNN with datasets representative of the inputs expected in operation, as these need to be manually labelled. The challenge is to select a representative set of test inputs as small as possible to reduce the labelling cost, while sufficing to yield unbiased high-confidence estimates of the expected DNN accuracy. At the same time, testers are interested in exposing as many DNN mispredictions as possible to improve the DNN, ending up in the need for techniques pursuing a threefold aim: small dataset size, trustworthy estimates, mispredictions exposure. Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 1 |
| 2024 | Anomaly Detection and Root Cause Analysis of Microservices Energy ConsumptionabstractWith the expansion of cloud computing and data centers, the need has arisen to tackle their environmental impact. The increasing adoption of microservice architectures, while offering scalability and flexibility, poses new challenges in the effective management of systems’ energy consumption.This study analyzes experimentally the effectiveness, with respect to energy consumption, of algorithms for Anomaly Detection (AD) and Root Cause Analysis (RCA) for (containerized) microservices systems. The study analyzes five AD and three RCA algorithms. Metrics to assess the effectiveness of AD algorithms are Precision, Recall, and F-Score. For RCA algorithms, the chose metric is Precision at level k. Two subjects of different complexity are used: Sock Shop and UNI-Cloud. Experiments use a cross-over paired comparison design, involving multiple randomized runs for robust measures.The experiments show that AD algorithms exhibit a relatively moderate performance. The mean adjusted Precision for Sock Shop is 61.5%, while it is 75% for the best-performing algorithms (BIRCH, KNN, and SVM) on UNI-Cloud. The Recall and F-Score for UNI-Cloud, for the same algorithms, are 75%, while for Sock Shop KNN yields the best outcome at roughly 45%. MicroRCA and RCD emerge as the top-performing algorithms for RCA.We found that the effectiveness of AD algorithms is strongly influenced by anomaly thresholds, emphasizing the importance of careful tuning such algorithms. RCA algorithms reveal promising results, particularly RCD and MicroRCA, which showed robust performance. However, challenges remain, as seen with the ϵ-diagnosis algorithm, suggesting the need for further refinement.For DevOps engineers, the findings highlight the need to carefully select and tune AD and RCA algorithms for energy, and to take into account system topology and monitoring configurations. Maximilian Stefan Floroiu, Stefano Russo 0001, Luca Giamattei, Antonio Guerriero, Ivano Malavolta, Roberto Pietrantuono |
ICWS | 4 |
| 2024 | Automated functional and robustness testing of microservice architecturesabstractMicroservice Architectures (MSA) are nowadays largely adopted by companies in several domains to provide on-demand services. The reliability of microservices is fundamental to avoid failures compromising the business functionalities. MSA automated testing is possible thanks to well-defined service interfaces specified in open formats like OpenAPI/Swagger. To support automated MSA functional and non-functional testing, we define a framework that: (i) generates test cases with valid and invalid inputs, and executes and monitors tests; (ii) provides coverage and failure information not only on edge, but also on internal microservices; (iii) has the novel feature of identifying causal relations in observed chains of microservices failures. We abstract the testing process of MSA, present the MacroHive framework and its causal inference engine, compare it experimentally to state-of-the-art tools, and discuss its benefits in the MSA testing process. MacroHive exhibits performance comparable to advanced existing tools in terms of edge-level coverage. However, MacroHive has a better failure rate and provides the unique advantages of giving insights about internal coverage and failures, and of inferring causality in failure chains, evidencing microservices to be improved to increase the whole MSA reliability. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
J. Syst. Softw. | 2 |
| 2024 | Monitoring tools for DevOps and microservices: A systematic grey literature reviewabstractMicroservice-based systems are usually developed according to agile practices like DevOps, which enables rapid and frequent releases to promptly react and adapt to changes. Monitoring is a key enabler for these systems, as they allow to continuously get feedback from the field and support timely and tailored decisions for a quality-driven evolution. In the realm of monitoring tools available for microservices in the DevOps-driven development practice, each with different features, assumptions, and performance, selecting a suitable tool is an as much difficult as impactful task. This article presents the results of a systematic study of the grey literature we performed to identify, classify and analyze the available monitoring tools for DevOps and microservices. We selected and examined a list of 71 monitoring tools, drawing a map of their characteristics, limitations, assumptions, and open challenges, meant to be useful to both researchers and practitioners working in this area. Results are publicly available and replicable. Editor's note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001, Ivano Malavolta, Tanjina Islam, Madalina Dinga, Anne Koziolek, Snigdha Singh, Martin Armbruster, Jose-Maria Gutierrez-Martinez, Sergio Caro-Álvaro, Daniel Rodríguez-García, Sebastian Weber 0001, Jörg Henß, Estrella Fernández Vogelin, Fernando Simön Panojo |
J. Syst. Softw. | 2 |
| 2024 | Causality-driven Testing of Autonomous Driving SystemsabstractTesting Autonomous Driving Systems (ADS) is essential for safe development of self-driving cars. For thorough and realistic testing, ADS are usually embedded in a simulator and tested in interaction with the simulated environment. However, their high complexity and the multiple safety requirements lead to costly and ineffective testing. Recent techniques exploit many-objective strategies and ML to efficiently search the huge input space. Despite the indubitable advances, the need for smartening the search keep being pressing. This article presents CART ( CAusal-Reasoning-driven Testing ), a new technique that formulates testing as a causal reasoning task. Learning causation, unlike correlation, allows assessing the effect of actively changing an input on the output, net of possible confounding variables. CART first infers the causal relations between test inputs and outputs, then looks for promising tests by querying the learnt model. Only tests suggested by the model are run on the simulator. An extensive empirical evaluation, using Pylot as ADS and CARLA as simulator, compares CART with state-of-the-art algorithms used recently on ADS. CART shows a significant gain in exposing more safety violations and does so more efficiently. More broadly, the work opens to a wider exploitation of causal learning beside (or on top of) ML for testing-related tasks. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | An Empirical Evaluation of the Energy and Performance Overhead of Monitoring Tools on Docker-Based Systems
Madalina Dinga, Ivano Malavolta, Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono |
ICSOC (1) | 4 |
| 2023 | DevOpRET: Continuous reliability testing in DevOpsabstractAbstract To enter the production stage, in DevOps practices candidate software releases have to pass quality gates, where they are assessed to meet established target values for key indicators of interest. We believe software reliability should be an important such indicator, as it greatly contributes to the end‐user satisfaction. We propose DevOpRET , an approach for reliability testing as part of the acceptance testing stage in DevOps. DevOpRET relies on operational‐profile–based testing, a common reliability assessment technique. DevOpRET leverages usage and failure data monitored in operations to continuously refine its estimate. We evaluate accuracy and efficiency of DevOpRET through controlled experiments with a real‐world open source platform and with a microservice architectures benchmark. The results show that DevOpRET provides accurate and efficient estimates of the true reliability over subsequent DevOps cycles. Antonia Bertolino, Guglielmo De Angelis, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
J. Softw. Evol. Process. | 3 |
| 2022 | Microservices Integrated Performance and Reliability TestingabstractContinuous quality assurance for extra-functional properties of modern software systems is today a big challenge as their complexity is constantly increasing to satisfy market demands. This is the case of microservice systems. They provide high control on the scale of operation by means of fine-grained service decomposition, but this demands careful consideration of the relations between performance of individual microservices and service failures. Matteo Camilli, Antonio Guerriero, Andrea Janes, Barbara Russo, Stefano Russo 0001 |
AST | 2 |
| 2022 | Automated Grey-Box Testing of Microservice ArchitecturesabstractMicroservices Architectures (MSA) have found large adoption in companies delivering online services, often in conjunction with agile development practices. Microservices are distributed, independent and polyglot entities – all features favouring black-box testing. However, for real-scale MSA, a pure black-box strategy may not be able to exercise the system to properly cover the interactions involving internal microservices.We propose a grey-box strategy (MACROHIVE) for automated testing and monitoring of (internal) microservices interactions. It uses combinatorial testing to generate valid and invalid tests from microservices specification. Tests execution and monitoring are automated by a service mesh infrastructure. MACROHIVE runs the tests and traces the interactions among microservices, to report about internal coverage and failing behaviour.MACROHIVE is experimented on TrainTicket, an open-source MSA benchmark. It performs comparably to state-of-the-art techniques in terms of edge-level coverage, but exposes internal failures undetected by black-box testing, gives detailed internal coverage information, and requires fewer tests. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
QRS | 2 |
| 2021 | Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsabstractDeep Neural Networks (DNN) are typically tested for accuracy relying on a set of unlabelled real world data (operational dataset), from which a subset is selected, manually labelled and used as test suite. This subset is required to be small (due to manual labelling cost) yet to faithfully represent the operational context, with the resulting test suite containing roughly the same proportion of examples causing misprediction (i.e., failing test cases) as the operational dataset. However, while testing to estimate accuracy, it is desirable to also learn as much as possible from the failing tests in the operational dataset, since they inform about possible bugs of the DNN. A smart sampling strategy may allow to intentionally include in the test suite many examples causing misprediction, thus providing this way more valuable inputs for DNN improvement while preserving the ability to get trustworthy unbiased estimates. This paper presents a test selection technique (DeepEST) that actively looks for failing test cases in the operational dataset of a DNN, with the goal of assessing the DNN expected accuracy by a small and "informative" test suite (namely with a high number of mispredictions) for subsequent DNN improvement. Experiments with five subjects, combining four DNN models and three datasets, are described. The results show that DeepEST provides DNN accuracy estimates with precision close to (and often better than) those of existing sampling-based DNN testing techniques, while detecting from 5 to 30 times more mispredictions, with the same test suite size. Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 1 |
| 2020 | Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationabstractIn Continuous Integration (CI), regression testing is constrained by the time between commits. This demands for careful selection and/or prioritization of test cases within test suites too large to be run entirely. To this aim, some Machine Learning (ML) techniques have been proposed, as an alternative to deterministic approaches. Two broad strategies for ML-based prioritization are learning-to-rank and what we call ranking-to-learn (i.e., reinforcement learning). Various ML algorithms can be applied in each strategy. In this paper we introduce ten of such algorithms for adoption in CI practices, and perform a comprehensive study comparing them against each other using subjects from the Apache Commons project. We analyze the influence of several features of the code under test and of the test process. The results allow to draw criteria to support testers in selecting and tuning the technique that best fits their context. Antonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 2 |
| 2020 | Testing microservice architectures for operational reliabilityabstractSummary Microservice architectures (MSA) is an emerging software architectural paradigm for service‐oriented applications, well‐suited for dynamic contexts requiring loosely coupled independent services, frequent software releases and decentralized governance. A key problem in the engineering of MSA applications is the estimate of their reliability, which is difficult to perform prior to release due frequent releases/service upgrades, dynamic service interactions, and changes in the way customers use the applications. This paper presents an in vivo testing method, named EMART, to faithfully assess the reliability of an MSA application in operation. EMART is based on an adaptive sampling strategy, leveraging monitoring data about microservices usage and failure/success of user demands. We present results of evaluation of estimation accuracy, confidence and efficiency, through a set of controlled experiments with publicly available subjects. © 2019 John Wiley & Sons, Ltd. Roberto Pietrantuono, Stefano Russo 0001, Antonio Guerriero |
Softw. Test. Verification Reliab. | 3 |
| 2018 | Run-Time Reliability Estimation of Microservice ArchitecturesabstractMicroservices are gaining popularity as an architectural paradigm for service-oriented applications, especially suited for highly dynamic contexts requiring loosely-coupled independent services, frequent software releases, decentralized governance and data management. Because of the high flexibility and evolvability characterizing microservice architectures (MSAs), it is difficult to estimate their reliability at design time, as it changes continuously due to the services' upgrades and/or to the way applications are used by customers. This paper presents a testing method for on-demand reliability estimation of microservice applications in their operational phase. The method allows to faithfully assess, upon request, the reliability of a MSA-based application under a scarce testing budget, at any time when it is in operation, and exploit field data about microservice usage and failing/successful demands. A new in-vivo testing algorithm is developed based on an adaptive web sampling strategy, named Microservice Adaptive Reliability Testing (MART). The method is evaluated by simulation, as well as by experimentation on an example application based on the Netflix Open Source Software MSA stack, with encouraging results in terms of estimation accuracy and, especially, efficiency. Roberto Pietrantuono, Stefano Russo 0001, Antonio Guerriero |
ISSRE | 3 |