Federico Di Menna

dblp:319/9491 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0003-5834-6389ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Forecasting software runtime metrics: A comparative study of classical statistical, neural network, and foundation models
abstract
Modern software applications generate a wide range of runtime metrics, which are vital to many quality assurance activities. These data are often recorded and aggregated as time series to observe patterns and trends of various runtime aspects over time. In this context, Time Series Forecasting (TSF) offers unique opportunities for predicting software runtime behavior and identifying potential anomalies. Although TSF models have been successfully applied in fields such as economics and climatology, their capabilities for forecasting software runtime metrics remain relatively underexplored. In this paper, we conduct a comprehensive empirical evaluation of 8 TSF models on 110 real-world software runtime metrics recorded over the course of about one year. Our evaluation encompasses three classical statistical models, three neural network models, and two time series foundation models. Results show that the foundation models achieve state-of-the-art performance on TSF of software runtime metrics, outperforming other models with strong statistical significance. Our findings indicate that foundation models, despite being trained exclusively on time series data from other domains, can effectively generalize to software runtime metrics in a zero-shot setting. This makes them a convenient plug-and-play solution for practitioners and researchers aiming to integrate TSF into their software quality assurance processes. Yet, their performance is not uniformly superior across all the time series, underscoring the absence of a “ silver bullet ” solution.
Federico Di Menna, Luca Traini, Vittorio Cortellessa
J. Syst. Softw.1
2026 AMBER: An AI-enabled Java Microbenchmark Harness Extension to Dynamically Terminate Warm-up Iterations
abstract
Java Microbenchmark Harness ( JMH ) is the de facto standard framework for developing Java microbenchmarks—used to assess the performance of small code segments. A central challenge in microbenchmark design is determining the number of warm-up iterations required to reach steady-state execution: too few lead to inaccurate results, while too many introduce unnecessary overhead. This paper extends our previous contribution by providing a more detailed description of AMBER, an AI-enabled JMH extension that utilizes Time Series Classification to detect steady-state behavior at run-time and dynamically terminate warm-up iterations.
Antonio Trovato, Luca Traini, Federico Di Menna, Dario Di Nucci
Sci. Comput. Program.3
2025 AMBER: AI-Enabled Java Microbenchmark Harness
abstract
JMH is the standard framework for developing and running Java microbenchmarks-lightweight performance tests used to evaluate the execution time of small Java code segments. A key challenge in designing JMH microbenchmarks is determining the appropriate number of warm-up iterations- repeated executions needed to bring microbenchmarks to a performance steady state. Too few warm-up iterations can compromise result quality, as performance measurements may not accurately reflect steady-state behavior. Conversely, too many warm-up iterations can unnecessarily increase testing time. Here, we present AMBER, an AI-enabled extension of JMH, which leverages Time Series Classification algorithms to predict the beginning of the steady-state phase at run-time and dynamically halt warm-up iterations accordingly. Empirical results show the potential of Amber in enhancing the cost-effectiveness of Java microbenchmarks. A demo video of Amber is available at https://www.youtube.com/watch?v=7zOngDQ1z_k.
Antonio Trovato, Luca Traini, Federico Di Menna, Dario Di Nucci
ICST3
2025 Investigating Execution-Aware Language Models for Code Optimization
abstract
Code optimization is the process of enhancing code efficiency, while preserving its intended functionality. This process often requires a deep understanding of the code execution behavior at run-time to identify and address inefficiencies effectively. Recent studies have shown that language models can play a significant role in automating code optimization. However, these models may have insufficient knowledge of how code execute at run-time. To address this limitation, researchers have developed strategies that integrate code execution information into language models. These strategies have shown promise, enhancing the effectiveness of language models in various software engineering tasks. However, despite the close relationship between code execution behavior and efficiency, the specific impact of these strategies on code optimization remains largely unexplored. This study investigates how incorporating code execution information into language models affects their ability to optimize code. Specifically, we apply three different training strategies to incorporate four code execution aspects - line executions, line coverage, branch coverage, and variable states - into CodeT5+, a well-known language model for code. Our results indicate that executionaware models provide limited benefits compared to the standard CodeT5+ model in optimizing code.
Federico Di Menna, Luca Traini, Gabriele Bavota, Vittorio Cortellessa
ICPC1
2024 AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
abstract
Performance testing aims at uncovering efficiency issues of software systems. In order to be both effective and practical, the design of a performance test must achieve a reasonable trade-off between result quality and testing time. This becomes particularly challenging in Java context, where the software undergoes a warm-up phase of execution, due to just-in-time compilation. During this phase, performance measurements are subject to severe fluctuations, which may adversely affect quality of performance test results. Both practitioners and researchers have proposed approaches to mitigate this issue. Practitioners typically rely on a fixed number of iterated executions that are used to warm-up the software before starting to collect performance measurements (state-of-practice). Researchers have developed techniques that can dynamically stop warm-up iterations at runtime (state-of-the-art). However, these approaches often provide suboptimal estimates of the warm-up phase, resulting in either insufficient or excessive warm-up iterations, which may degrade result quality or increase testing time. There is still a lack of consensus on how to properly address this problem. Here, we propose and study an AI-based framework to dynamically halt warm-up iterations at runtime. Specifically, our framework leverages recent advances in AI for Time Series Classification (TSC) to predict the end of the warm-up phase during test execution. We conduct experiments by training three different TSC models on half a million of measurement segments obtained from JMH microbenchmark executions. We find that our framework significantly improves the accuracy of the warm-up estimates provided by state-of-practice and state-of-the-art methods. This higher estimation accuracy results in a net improvement in either result quality or testing time for up to +35.3% of the microbenchmarks. Our study highlights that integrating AI to dynamically estimate the end of the warm-up phase can enhance the cost-effectiveness of Java performance testing.
Luca Traini, Federico Di Menna, Vittorio Cortellessa
ASE2
2024 RADig-X: a Tool for Regressions Analysis of User Digital Experience
abstract
The successful operation of a modern company re-lays on the dependability of its software infrastructure. However, ensuring a robust and dependable software infrastructure can be challenging, as software applications are subject to continuous updates that can introduce bugs and performance regressions. To mitigate this challenge, many companies use Application Performance Management (APM) tools to monitor their digital devices and identify potential issues that could affect business operability. However, the large volume and heterogeneity of the data collected by these tools can make it difficult to effectively analyze and exploit the rich source of information available. In this paper, we propose RADig-X, a tool designed to support the identification and analysis of digital experience issues. RADig-X leverages AI algorithms and a ranking heuristic to: (i) detect anomalies in runtime metrics collected by APM tools, (ii) assess the relevance of these anomalies based on their impact on the overall IT infrastructure, and (iii) rank problematic software updates that may be the root cause of relevant anomalies. We report on the adoption of RADig-X by a large company that monitors over 30,000 digital devices around the world. Our results demonstrate that RADig-X is able to improve the effectiveness of the identification process of digital experience issues, by enabling to identify and address potential anomalies that could impact business operations. RADig-X is currently used in production within the case company to support the diagnosis and problem resolution of digital experience issues.
Federico Di Menna, Vittorio Cortellessa, Maurizio Lucianelli, Luca Sardo, Luca Traini
SANER1
2024 Time Series Forecasting of Runtime Software Metrics: An Empirical Study
abstract
Software applications can produce a wide range of runtime software metrics (e.g., number of crashes, response times), which can be closely monitored to ensure operational efficiency and prevent significant software failures. These metrics are typically recorded as time series data. However, runtime software monitoring has become a high-effort task due to the growing complexity of today's software systems. In this context, time series forecasting (TSF) offers unique opportunities to enhance software monitoring and facilitate proactive issue resolution. While TSF methods have been widely studied in areas like economics and weather forecasting, our understanding of their effectiveness for software runtime metrics remains somewhat limited. In this paper, we investigate the effectiveness of four TSF methods on 25 real-world runtime software metrics recorded over a period of one and a half years. These methods comprise three recurrent neural network (RNN) models and one traditional time series analysis technique (i.e., SARIMA). The metrics are gathered from a large-scale IT infrastructure involving tens of thousands of digital devices. Our results indicate that, in general, RNN models are very effective in the runtime software metrics prediction, although in some scenarios and for certain specific metrics (e.g., waiting times) SARIMA proves to outperform RNN models. Additionally, our findings suggest that the advantages of using RNN models vanish when the prediction horizon becomes too wide, in our case when it exceeds one week.
Federico Di Menna, Luca Traini, Vittorio Cortellessa
ICPE1