Luca Traini

dblp:225/0294 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-3676-0645ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 6 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Empirical Investigation on the Use of Large Language Models for Performance Bug Detection
Muhammad Imran 0026, Vittorio Cortellessa, Davide Di Ruscio, Riccardo Rubei, Luca Traini
SANER5
2026 A kernel-based approach for accurate steady-state detection in performance time series
abstract
This paper addresses the challenge of accurately detecting the transition from the warmup phase to the steady state in performance metric time series, which is a critical step for effective benchmarking. The goal is to introduce a method that avoids premature or delayed detection, which can lead to inaccurate or inefficient performance analysis. The proposed approach adapts techniques from the chemical reactors domain, detecting steady states online through the combination of kernel-based step detection and statistical methods. By using a window-based approach, it provides detailed information and improves the accuracy of identifying phase transitions, even in noisy or irregular time series. Results show that the new approach reduces total error by 14.5% compared to the best selected state-of-the-art method. It offers more reliable detection of the steady-state onset, delivering greater precision for benchmarking tasks. For users, the new approach enhances the accuracy and stability of performance benchmarking, efficiently handling diverse time series data. Its robustness and adaptability make it a valuable tool for real-world performance evaluation, ensuring consistent and reproducible results.
Martin Beseda, Vittorio Cortellessa, Daniele Di Pompeo, Luca Traini, Michele Tucci 0001
Future Gener. Comput. Syst.4
2026 Forecasting software runtime metrics: A comparative study of classical statistical, neural network, and foundation models
abstract
Modern software applications generate a wide range of runtime metrics, which are vital to many quality assurance activities. These data are often recorded and aggregated as time series to observe patterns and trends of various runtime aspects over time. In this context, Time Series Forecasting (TSF) offers unique opportunities for predicting software runtime behavior and identifying potential anomalies. Although TSF models have been successfully applied in fields such as economics and climatology, their capabilities for forecasting software runtime metrics remain relatively underexplored. In this paper, we conduct a comprehensive empirical evaluation of 8 TSF models on 110 real-world software runtime metrics recorded over the course of about one year. Our evaluation encompasses three classical statistical models, three neural network models, and two time series foundation models. Results show that the foundation models achieve state-of-the-art performance on TSF of software runtime metrics, outperforming other models with strong statistical significance. Our findings indicate that foundation models, despite being trained exclusively on time series data from other domains, can effectively generalize to software runtime metrics in a zero-shot setting. This makes them a convenient plug-and-play solution for practitioners and researchers aiming to integrate TSF into their software quality assurance processes. Yet, their performance is not uniformly superior across all the time series, underscoring the absence of a “ silver bullet ” solution.
Federico Di Menna, Luca Traini, Vittorio Cortellessa
J. Syst. Softw.2
2026 AMBER: An AI-enabled Java Microbenchmark Harness Extension to Dynamically Terminate Warm-up Iterations
abstract
Java Microbenchmark Harness ( JMH ) is the de facto standard framework for developing Java microbenchmarks—used to assess the performance of small code segments. A central challenge in microbenchmark design is determining the number of warm-up iterations required to reach steady-state execution: too few lead to inaccurate results, while too many introduce unnecessary overhead. This paper extends our previous contribution by providing a more detailed description of AMBER, an AI-enabled JMH extension that utilizes Time Series Classification to detect steady-state behavior at run-time and dynamically terminate warm-up iterations.
Antonio Trovato, Luca Traini, Federico Di Menna, Dario Di Nucci
Sci. Comput. Program.2
2025 AMBER: AI-Enabled Java Microbenchmark Harness
abstract
JMH is the standard framework for developing and running Java microbenchmarks-lightweight performance tests used to evaluate the execution time of small Java code segments. A key challenge in designing JMH microbenchmarks is determining the appropriate number of warm-up iterations- repeated executions needed to bring microbenchmarks to a performance steady state. Too few warm-up iterations can compromise result quality, as performance measurements may not accurately reflect steady-state behavior. Conversely, too many warm-up iterations can unnecessarily increase testing time. Here, we present AMBER, an AI-enabled extension of JMH, which leverages Time Series Classification algorithms to predict the beginning of the steady-state phase at run-time and dynamically halt warm-up iterations accordingly. Empirical results show the potential of Amber in enhancing the cost-effectiveness of Java microbenchmarks. A demo video of Amber is available at https://www.youtube.com/watch?v=7zOngDQ1z_k.
Antonio Trovato, Luca Traini, Federico Di Menna, Dario Di Nucci
ICST2
2025 Investigating Execution-Aware Language Models for Code Optimization
abstract
Code optimization is the process of enhancing code efficiency, while preserving its intended functionality. This process often requires a deep understanding of the code execution behavior at run-time to identify and address inefficiencies effectively. Recent studies have shown that language models can play a significant role in automating code optimization. However, these models may have insufficient knowledge of how code execute at run-time. To address this limitation, researchers have developed strategies that integrate code execution information into language models. These strategies have shown promise, enhancing the effectiveness of language models in various software engineering tasks. However, despite the close relationship between code execution behavior and efficiency, the specific impact of these strategies on code optimization remains largely unexplored. This study investigates how incorporating code execution information into language models affects their ability to optimize code. Specifically, we apply three different training strategies to incorporate four code execution aspects - line executions, line coverage, branch coverage, and variable states - into CodeT5+, a well-known language model for code. Our results indicate that executionaware models provide limited benefits compared to the standard CodeT5+ model in optimizing code.
Federico Di Menna, Luca Traini, Gabriele Bavota, Vittorio Cortellessa
ICPC2
2025 On the Compression of Language Models for Code: An Empirical Study on CodeBERT
abstract
Language models have proven successful across a wide range of software engineering tasks, but their significant computational costs often hinder their practical adoption. To address this challenge, researchers have begun applying various compression strategies to improve the efficiency of language models for code. These strategies aim to optimize inference latency and memory usage, though often at the cost of reduced model effectiveness. However, there is still a significant gap in understanding how these strategies influence the efficiency and effectiveness of language models for code. Here, we empirically investigate the impact of three well-known compression strategies - knowledge distillation, quantization, and pruning - across three different classes of software engineering tasks: vulnerability detection, code summarization, and code search. Our findings reveal that the impact of these strategies varies greatly depending on the task and the specific compression method employed. Practitioners and researchers can use these insights to make informed decisions when selecting the most appropriate compression strategy, balancing both efficiency and effectiveness based on their specific needs.
Giordano d'Aloisio, Luca Traini, Federica Sarro, Antinisca Di Marco
SANER2
2025 Is code coverage of performance tests related to source code features? An empirical study on open-source Java systems
abstract
Abstract Performance testing aims to ensure the operational efficiency of software systems. However, many factors influencing the efficacy and adoption of performance tests in practice are not yet fully understood. For instance, while code coverage is widely regarded as a key quality metric for evaluating the efficacy of functional testing suites, there is limited knowledge about the types and levels of coverage that performance tests specifically achieve. Another important factor, often perceived as a barrier to the broader adoption of performance tests yet remaining relatively unexplored, is their extended execution time. In this paper, we examine (i) the coverage of performance testing suites, (ii) the characteristics of source code associated with performance-tested components, and (iii) the time cost of executing performance tests. Our analysis on open-source Java systems reveals that performance tests achieve significantly lower code coverage than functional tests, as expected, and it highlights a significant trade-off between coverage and execution time. Our results also indicate a lack of generalizable characteristics in the source code covered by performance tests.
Muhammad Imran 0026, Vittorio Cortellessa, Davide Di Ruscio, Riccardo Rubei, Luca Traini
Empir. Softw. Eng.5
2024 An Empirical Study on Code Coverage of Performance Testing
abstract
Performance testing aims to ensure the operational efficiency of software systems. However, many factors influencing the efficacy and adoption of performance tests in practice are not yet fully understood. For instance, while code coverage is widely regarded as a key quality metric for evaluating the efficacy of functional testing suites, there is limited knowledge about the types and levels of coverage that performance tests specifically achieve. Another important factor, often perceived as a barrier to the broader adoption of performance tests yet remaining relatively unexplored, is their extended execution time. In this paper, we analyze the performance testing suites of 28 open-source systems to study (i) the magnitude of their code coverage, and (ii) their execution time. Our analysis shows that performance tests achieve significantly lower code coverage than functional tests, as expected, and it highlights a significant trade-off between coverage and execution time. Our results also suggest, in perspective, that automated test generation methods might not ensure affordable performance testing due to the associated time cost. This finding poses new challenges in the field of performance test generation.
Muhammad Imran 0026, Vittorio Cortellessa, Davide Di Ruscio, Riccardo Rubei, Luca Traini
EASE5
2024 AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
abstract
Performance testing aims at uncovering efficiency issues of software systems. In order to be both effective and practical, the design of a performance test must achieve a reasonable trade-off between result quality and testing time. This becomes particularly challenging in Java context, where the software undergoes a warm-up phase of execution, due to just-in-time compilation. During this phase, performance measurements are subject to severe fluctuations, which may adversely affect quality of performance test results. Both practitioners and researchers have proposed approaches to mitigate this issue. Practitioners typically rely on a fixed number of iterated executions that are used to warm-up the software before starting to collect performance measurements (state-of-practice). Researchers have developed techniques that can dynamically stop warm-up iterations at runtime (state-of-the-art). However, these approaches often provide suboptimal estimates of the warm-up phase, resulting in either insufficient or excessive warm-up iterations, which may degrade result quality or increase testing time. There is still a lack of consensus on how to properly address this problem. Here, we propose and study an AI-based framework to dynamically halt warm-up iterations at runtime. Specifically, our framework leverages recent advances in AI for Time Series Classification (TSC) to predict the end of the warm-up phase during test execution. We conduct experiments by training three different TSC models on half a million of measurement segments obtained from JMH microbenchmark executions. We find that our framework significantly improves the accuracy of the warm-up estimates provided by state-of-practice and state-of-the-art methods. This higher estimation accuracy results in a net improvement in either result quality or testing time for up to +35.3% of the microbenchmarks. Our study highlights that integrating AI to dynamically estimate the end of the warm-up phase can enhance the cost-effectiveness of Java performance testing.
Luca Traini, Federico Di Menna, Vittorio Cortellessa
ASE1
2024 RADig-X: a Tool for Regressions Analysis of User Digital Experience
abstract
The successful operation of a modern company re-lays on the dependability of its software infrastructure. However, ensuring a robust and dependable software infrastructure can be challenging, as software applications are subject to continuous updates that can introduce bugs and performance regressions. To mitigate this challenge, many companies use Application Performance Management (APM) tools to monitor their digital devices and identify potential issues that could affect business operability. However, the large volume and heterogeneity of the data collected by these tools can make it difficult to effectively analyze and exploit the rich source of information available. In this paper, we propose RADig-X, a tool designed to support the identification and analysis of digital experience issues. RADig-X leverages AI algorithms and a ranking heuristic to: (i) detect anomalies in runtime metrics collected by APM tools, (ii) assess the relevance of these anomalies based on their impact on the overall IT infrastructure, and (iii) rank problematic software updates that may be the root cause of relevant anomalies. We report on the adoption of RADig-X by a large company that monitors over 30,000 digital devices around the world. Our results demonstrate that RADig-X is able to improve the effectiveness of the identification process of digital experience issues, by enabling to identify and address potential anomalies that could impact business operations. RADig-X is currently used in production within the case company to support the diagnosis and problem resolution of digital experience issues.
Federico Di Menna, Vittorio Cortellessa, Maurizio Lucianelli, Luca Sardo, Luca Traini
SANER5
2024 Time Series Forecasting of Runtime Software Metrics: An Empirical Study
abstract
Software applications can produce a wide range of runtime software metrics (e.g., number of crashes, response times), which can be closely monitored to ensure operational efficiency and prevent significant software failures. These metrics are typically recorded as time series data. However, runtime software monitoring has become a high-effort task due to the growing complexity of today's software systems. In this context, time series forecasting (TSF) offers unique opportunities to enhance software monitoring and facilitate proactive issue resolution. While TSF methods have been widely studied in areas like economics and weather forecasting, our understanding of their effectiveness for software runtime metrics remains somewhat limited. In this paper, we investigate the effectiveness of four TSF methods on 25 real-world runtime software metrics recorded over a period of one and a half years. These methods comprise three recurrent neural network (RNN) models and one traditional time series analysis technique (i.e., SARIMA). The metrics are gathered from a large-scale IT infrastructure involving tens of thousands of digital devices. Our results indicate that, in general, RNN models are very effective in the runtime software metrics prediction, although in some scenarios and for certain specific metrics (e.g., waiting times) SARIMA proves to outperform RNN models. Additionally, our findings suggest that the advantages of using RNN models vanish when the prediction horizon becomes too wide, in our case when it exceeds one week.
Federico Di Menna, Luca Traini, Vittorio Cortellessa
ICPE2
2023 Towards effective assessment of steady state performance in Java software: are we there yet?
abstract
Abstract Microbenchmarking is a widely used form of performance testing in Java software. A microbenchmark repeatedly executes a small chunk of code while collecting measurements related to its performance. Due to Java Virtual Machine optimizations, microbenchmarks are usually subject to severe performance fluctuations in the first phase of their execution (also known as warmup). For this reason, software developers typically discard measurements of this phase and focus their analysis when benchmarks reach a steady state of performance. Developers estimate the end of the warmup phase based on their expertise, and configure their benchmarks accordingly. Unfortunately, this approach is based on two strong assumptions: (i) benchmarks always reach a steady state of performance and (ii) developers accurately estimate warmup. In this paper, we show that Java microbenchmarks do not always reach a steady state, and often developers fail to accurately estimate the end of the warmup phase. We found that a considerable portion of studied benchmarks do not hit the steady state, and warmup estimates provided by software developers are often inaccurate (with a large error). This has significant implications both in terms of results quality and time-effort. Furthermore, we found that dynamic reconfiguration significantly improves warmup estimation accuracy, but still it induces suboptimal warmup estimates and relevant side-effects. We envision this paper as a starting point for supporting the introduction of more sophisticated automated techniques that can ensure results quality in a timely fashion.
Luca Traini, Vittorio Cortellessa, Daniele Di Pompeo, Michele Tucci 0001
Empir. Softw. Eng.1
2023 DeLag: Using Multi-Objective Optimization to Enhance the Detection of Latency Degradation Patterns in Service-Based Systems
abstract
Performance debugging in production is a fundamental activity in modern service-based systems. The diagnosis of performance issues is often time-consuming, since it requires thorough inspection of large volumes of traces and performance indices. In this paper we present DeLag, a novel automated search-based approach for diagnosing performance issues in service-based systems. DeLag identifies subsets of requests that show, in the combination of their Remote Procedure Call execution times, symptoms of potentially relevant performance issues. We call such symptomsLatency Degradation Patterns. DeLag simultaneously searches for multiplelatency degradation patternswhile optimizing precision, recall and latency dissimilarity. Experimentation on 700 datasets of requests generated from two microservice-based systems shows that our approach provides better and more stable effectiveness than three state-of-the-art approaches and general purpose machine learning clustering algorithms. DeLag is more effective than all baseline techniques in at least one case study (with$p\leq 0.05$and non-negligible effect size). Moreover, DeLag outperforms in terms of efficiency the second and the third most effective baseline techniques on the largest datasets used in our evaluation (up to 22%).
Luca Traini, Vittorio Cortellessa
IEEE Trans. Software Eng.1
2022 Exploring Performance Assurance Practices and Challenges in Agile Software Development: An Ethnographic Study
Luca Traini
Empir. Softw. Eng.1
2022 How Software Refactoring Impacts Execution Time
abstract
Refactoring aims at improving the maintainability of source code without modifying its external behavior. Previous works proposed approaches to recommend refactoring solutions to software developers. The generation of the recommended solutions is guided by metrics acting as proxy for maintainability (e.g., number of code smells removed by the recommended solution). These approaches ignore the impact of the recommended refactorings on other non-functional requirements, such as performance, energy consumption, and so forth. Little is known about the impact of refactoring operations on non-functional requirements other than maintainability. We aim to fill this gap by presenting the largest study to date to investigate the impact of refactoring on software performance, in terms of execution time. We mined the change history of 20 systems that defined performance benchmarks in their repositories, with the goal of identifying commits in which developers implemented refactoring operations impacting code components that are exercised by the performance benchmarks. Through a quantitative and qualitative analysis, we show that refactoring operations can significantly impact the execution time. Indeed, none of the investigated refactoring types can be considered “safe” in ensuring no performance regression. Refactoring types aimed at decomposing complex code entities (e.g., Extract Class/Interface, Extract Method) have higher chances of triggering performance degradation, suggesting their careful consideration when refactoring performance-critical code.
Luca Traini, Daniele Di Pompeo, Michele Tucci 0001, Bin Lin 0008, Simone Scalabrino, Gabriele Bavota, Michele Lanza 0001, Rocco Oliveto, Vittorio Cortellessa
ACM Trans. Softw. Eng. Methodol.1
2020 Detecting Latency Degradation Patterns in Service-based Systems
abstract
Performance in heterogeneous service-based systems shows non-determistic trends. Even for the same request type, latency may vary from one request to another. These variations can occur due to several reasons on different levels of the software stack: operating system, network, software libraries, application code or others. Furthermore, a request may involve several Remote Procedure Calls (RPC), where each call can be subject to performance variation. Performance analysts inspect distributed traces and seek for recurrent patterns in trace attributes, such as RPCs execution time, in order to cluster traces in which variations may be induced by the same cause. Clustering "similar" traces is a prerequisite for effective performance debugging. Given the scale of the problem, such activity can be tedious and expensive. In this paper, we present an automated approach that detects relevant RPCs execution time patterns associated to request latency degradation, i.e. latency degradation patterns. The presented approach is based on a genetic search algorithm driven by an information retrieval relevance metric and an optimized fitness evaluation. Each latency degradation pattern identifies a cluster of requests subject to latency degradation with similar patterns in RPCs execution time. We show on a microservice-based application case study that the proposed approach can effectively detect clusters identified by artificially injected latency degradation patterns. Experimental results show that our approach outperforms in terms of F-score a state-of-art approach for latency profile analysis and widely popular machine learning clustering algorithms. We also show how our approach can be easily extended to trace attributes other than RPC execution time (e.g. HTTP headers, execution node, etc.).
Vittorio Cortellessa, Luca Traini
ICPE2
2018 A multi-objective framework for effective performance fault injection in distributed systems
abstract
Modern distributed systems should be built to anticipate performance degradation. Often requests in these systems involve ten to thousands Remote Procedure Calls, each of which can be a source of performance degradation. The PhD programme presented here intends to address this issue by providing automated instruments to effectively drive performance fault injection in distributed systems. The envisioned approach exploits multi-objective search-based techniques to automatically find small combinations of tiny performance degradations induced by specific RPCs,which have significant impacts on the user-perceived performance. Automating the search of these events will improve the ability to inject performance issues in production in order to force developers to anticipate and mitigate them.
Luca Traini
ASE1