Matteo Esposito 0001

dblp:326/5206 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-8451-3668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 9 first-author · 22 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Can AI Agents Generate Microservices? How Far are We?
abstract
Context. LLMs have advanced code generation, but their use for generating microservices with explicit dependencies and API contracts remains understudied.Goal. We examine whether AI agents can generate functional microservices and how different forms of contextual information influence their performance.Method. We assess 144 generated microservices across 3 agents, 4 projects, 2 prompting strategies, and 2 scenarios. Incremental generation operates within existing systems and is evaluated with unit tests. Clean state generation starts from requirements alone and is evaluated with integration tests. We analyze functional correctness, code quality, and efficiency.Results. Minimal prompts outperformed detailed ones in incremental generation, with 50-76% unit test pass rates. Clean state generation produced higher integration test pass rates (81-98%), indicating strong API contract adherence. Generated code showed lower complexity than human baselines. Generation times varied widely across agents, averaging 6-16 minutes per service.Conclusions. AI agents can produce microservices with maintainable code, yet inconsistent correctness and reliance on human oversight show that fully autonomous microservice generation is not yet achievable.
Bassam Adnan, Matteo Esposito 0001, Davide Taibi 0001, Karthik Vaidhyanathan
ICSA2
2026 Evaluating Large Language Models for Detecting Architectural Decision Violations
abstract
Architectural Decision Records (ADRs) play a central role in maintaining software architecture quality, yet many decision violations go unnoticed because projects lack both systematic documentation and automated detection mechanisms. Recent advances in Large Language Models (LLMs) open up new possibilities for automating architectural reasoning at scale. We investigated how effectively LLMs can identify decision violations in open-source systems by examining their agreement, accuracy, and inherent limitations. Our study analyzed 980 ADRs across 109 GitHub repositories using a multi-model pipeline in which one LLM primary screens potential decision violations, and three additional LLMs independently validate the reasoning. We assessed agreement, accuracy, precision, and recall, and complemented the quantitative findings with expert evaluation. The models achieved substantial agreement and strong accuracy for explicit, code-inferable decisions. Accuracy falls short for implicit or deployment-oriented decisions that depend on deployment configuration or organizational knowledge. Therefore, LLMs can meaningfully support validation of architectural decision compliance; however, they are not yet replacing human expertise for decisions not focused on code.
Ruoyu Su, Alexander Bakhtin, Noman Ahmad, Matteo Esposito 0001, Valentina Lenarduzzi, Davide Taibi 0001
ICSA4
2026 SQuaD: The Software Quality Dataset
abstract
Software quality research increasingly relies on large-scale datasets that measure both the product and process aspects of software systems. However, existing resources often focus on limited dimensions, such as code smells, technical debt, or refactoring activity, thereby restricting comprehensive analyses across isolated quality dimensions. To address this gap, we present the Software Quality Dataset (SQuaD), a multi-dimensional, time-aware collection of software quality metrics extracted from 450 mature open-source projects across diverse ecosystems, including Apache, Mozilla, FFmpeg, and the Linux kernel. By integrating nine state-of-the-art static analysis tools, i.e., SonarQube, CodeScene, PMD, Understand, CK, JaSoMe, RefactoringMiner, RefactoringMiner++, and PyRef, our dataset unifies over 700 unique metrics at method, class, file, and project levels. Covering a total of 63,586 analyzed project releases, SQuaD also provides version control and issue-tracking histories, software vulnerability data (CVE/CWE), and process metrics proven to enhance Just-In-Time (JIT) defect prediction. The SQuaD enables empirical research on maintainability, technical debt, software evolution, and quality assessment at unprecedented scale. We also outline emerging research directions, including automated dataset updates and cross-project quality modeling to support the continuous evolution of software analytics. The dataset is publicly available on ZENODO (DOI: 10.5281/zenodo.17566690).
Mikel Robredo, Matteo Esposito 0001, Davide Taibi 0001, Rafael Peñaloza, Valentina Lenarduzzi
MSR2
2026 Running Large Language Models at Scale for Mining Software Repositories: Lessons Learned from HPC-Based Batch Inference
abstract
The rapid diffusion of Large Language Models (LLMs) is fundamentally changing how Mining Software Repositories (MSR) research is conducted, particularly for studies that rely on unstructured textual artifacts such as commit messages, issue discussions, pull request reviews, and practitioner-generated content. While recent work has demonstrated the potential of LLMs to support classification, summarization, and qualitative analysis tasks, the majority of existing approaches rely on interactive or API-based executions [2, 3, 6]. Such execution models are poorly suited for large-scale empirical MSR studies, where thousands or hundreds of thousands of artifacts must be processed in a controlled, reproducible, and cost-aware manner.
Ruoyu Su, Matteo Esposito 0001, Davide Taibi 0001, Valentina Lenarduzzi
MSR2
2026 Generative AI as an infrastructure copilot: automating Infrastructure-As-Code across the DevSecOps lifecycle
abstract
Abstract Practitioners and researchers continuously focus on developing automation strategies to cope with the exponentially demanding need for the timely deployment of software projects in tight release schedules. Such automation techniques include Infrastructure-as-Code (IaC) and the DevOps and DevSecOps cycles. Recent studies investigated generative AI (GenAI) for generating infrastructure as code scripts. However, no studies have focused on using GenAI to generate IaC scripts based on DevSecOps stage artifacts. Different IaC tools serve varied purposes, requiring specific infrastructure setups for different project stages. We envision GenAI models leveraging artifacts from each DevSecOps stage to create and refine IaC scripts. We trust our approach to have an impact on practitioners to leverage it as an automatic copilot for infrastructure design and deployment, and for researchers to build on our vision and future empirical validation.
Matteo Esposito 0001, Mikel Robredo, Alexander Bakhtin, Davide Taibi 0001, Valentina Lenarduzzi
Autom. Softw. Eng.1
2026 Analyzing the ripple effects of refactoring
Mikel Robredo, Matteo Esposito 0001, Fabio Palomba, Rafael Peñaloza, Valentina Lenarduzzi
Empir. Softw. Eng.2
2026 Generative AI for software architecture. Applications, challenges, and future directions
Matteo Esposito 0001, Xiaozhou Li 0002, Sergio Moreschini, Noman Ahmad, Tomás Cerný, Karthik Vaidhyanathan, Valentina Lenarduzzi, Davide Taibi 0001
J. Syst. Softw.1
2026 Emerging trends in software architecture from the practitioner's perspective: A five-year review
Ruoyu Su, Noman Ahmad, Matteo Esposito 0001, Andrea Janes, Davide Taibi 0001, Valentina Lenarduzzi
J. Syst. Softw.3
2025 Centrality Change Proneness: An Early Indicator of Microservice Architectural Degradation
Alexander Bakhtin, Matteo Esposito 0001, Valentina Lenarduzzi, Davide Taibi 0001
ECSA2
2025 Network Centrality as a New Perspective on Microservice Architecture
abstract
Context: Over the past decade, the adoption of Microservice Architecture (MSA) has led to the identification of various patterns and anti-patterns, such as Nano/Mega/Hub services. Detecting these anti-patterns often involves modeling the system as a Service Dependency Graph (SDG) and applying graph-theoretic approaches. Aim: While previous research has explored software metrics (SMs) such as size, complexity, and quality for assessing MSAs, the potential of graph-specific metrics like network centrality remains largely unexplored. This study investigates whether centrality metrics (CMs) can provide new insights into MSA quality and facilitate the detection of architectural anti-patterns, complementing or extending traditional SMs. Method: We analyzed 24 open-source MSA projects, reconstructing their architectures to study 53 microservices. We measured SMs and CMs for each microservice and tested their correlation to determine the relationship between these metric types. Results and Conclusion: Among 902 computed metric correlations, we found weak to moderate correlation in 282 cases. These findings suggest that centrality metrics offer a novel perspective for understanding MSA properties. Specifically, ratio-based centrality metrics show promise for detecting specific anti-patterns, while subgraph centrality needs further investigation for its applicability in architectural assessments.
Alexander Bakhtin, Matteo Esposito 0001, Valentina Lenarduzzi, Davide Taibi 0001
ICSA2
2025 Solutions toCybersecurity Challenges in Secure Vehicle-to-Vehicle Communications: A Multivocal Literature Review
Siffat Ullah Khan, Mahmood Khan Niazi, Matteo Esposito 0001, Arif Ali Khan, Jamal Abdul Nasir
Inf. Softw. Technol.4
2025 Evaluating time-dependent methods and seasonal effects in code technical debt prediction
abstract
Background: Code Technical Debt (Code TD) prediction has gained significant attention in recent software engineering research. However, no standardized approach to Code TD prediction fully captures the factors influencing its evolution. Objective: Our study aims to assess the impact of time-dependent models and seasonal effects on Code TD prediction. It evaluates such models against widely used Machine Learning models also considering the influence of seasonality on prediction performance. Methods: We trained 11 prediction models with 31 Java open-source projects. To assess their performance, we predicted future observations of the SQALE index. To evaluate the practical usability of our TD forecasting model and their impact on practitioners, we surveyed 23 software engineering professionals. Results: Our study confirms the benefits of time-dependent techniques, with the ARIMAX model outperforming the others. Seasonal effects improved predictive performance, though the impact remained modest. ARIMAX/SARIMAX models demonstrated to provide well-balanced long-term forecasts. The survey highlighted strong industry interest in short- to medium-term TD forecasts. Conclusions: Our findings support using techniques that capture time dependence in historical software metric data, particularly for Code TD. Effectively addressing this evidence requires adopting methods that account for temporal patterns.
Mikel Robredo, Nyyti Saarimäki, Matteo Esposito 0001, Davide Taibi 0001, Rafael Peñaloza, Valentina Lenarduzzi
J. Syst. Softw.3
2025 On the correlation between architectural smells and static analysis warnings
abstract
Abstract Software quality assurance is essential during software development and maintenance. Static Analysis Tools (SAT) are widely used for assessing code quality. Architectural smells are becoming more daunting to address and evaluate among quality issues. We aim to understand the relationships between Static Analysis Warnings (“warnings”) and Architectural Smells (“smell”) to guide developers/maintainers in focusing their effort on warnings more prone to co-occurring with smell. We performed an empirical study on 103 Java projects totaling 72 million LOC belonging to projects from a vast set of domains, and 785 warnings were detected by three SAT, Checkstyle, Findbugs, PMD, SonarQube, and 4 architectural smells were detected by the ARCAN tool. We analyzed how warnings influence smell presence. Finally, we proposed a smell remediation effort prioritization based on warning severity and warning proneness to specific smells. Our study reveals a moderate correlation between warnings and smells. Different combinations of SATs and warnings significantly affect smell occurrence, with certain warnings more Likely to co-occur with specific smells. Conversely, 33.79% of warnings are “non-co-occurring” with any of the smells in our dataset. This provides an early indicator for potential architectural concerns before resource-intensive architectural analysis is performed. Practitioners can ignore about a third of warnings and focus on those most likely to be associated with smells. Prioritizing smell remediation based on warning severity or warning proneness to specific smells results in effective rankings like those based on smell severity. While not a substitute for specialized tools like ARCAN, warning-based prioritization provides a pragmatic bridge between low-level warnings and high-level architectural issues, particularly useful in contexts lacking full architectural visibility.
Matteo Esposito 0001, Mikel Robredo, Francesca Arcelli Fontana, Valentina Lenarduzzi
Softw. Qual. J.1
2024 An Extensive Comparison of Static Application Security Testing Tools
abstract
Context: Static Application Security Testing Tools (SASTTs) identify software vulnerabilities to support the security and reliability of software applications. Interestingly, several studies have suggested that alternative solutions may be more effective than SASTTs due to their tendency to generate false alarms, commonly referred to as low Precision. Aim: We aim to comprehensively evaluate SASTTs, setting a reliable benchmark for assessing and finding gaps in vulnerability identification mechanisms based on SASTTs or alternatives. Method: Our SASTTs evaluation is based on a controlled, though synthetic, Java codebase. It involves an assessment of 1.5 million test executions, and it features innovative methodological features such as effort-aware accuracy metrics and method-level analysis. Results: Our findings reveal that SASTTs detect a tiny range of vulnerabilities. In contrast to prevailing wisdom, SASTTs exhibit high Precision while falling short in Recall. Conclusions: Our findings suggest that enhancing Recall, alongside expanding the spectrum of detected vulnerability types, should be the primary focus for improving SASTTs or alternative approaches, such as machine learning-based vulnerability identification solutions.
Matteo Esposito 0001, Valentina Falaschi, Davide Falessi
EASE1
2024 Leveraging Large Language Models for Preliminary Security Risk Analysis: A Mission-Critical Case Study
abstract
Preliminary security risk analysis (PSRA) provides a quick approach to identify, evaluate, and propose remediation to potential risks in specific scenarios. The extensive expertise required for an effective PSRA and the substantial textual-related tasks hinders quick assessments in mission-critical contexts, where timely and prompt actions are essential. The speed and accuracy of human experts in PSRA significantly impact response time. A large language model can quickly summarise information in less time than a human. To our knowledge, no prior study has explored the capabilities of fine-tuned models (FTM) in PSRA. Our case study investigates the proficiency of FTM in assisting practitioners in PSRA. We manually curated 141 representative samples from over 50 mission-critical analyses archived by the industrial context team in the last five years. We compared the proficiency of the FTM versus seven human experts. Within the industrial context, our approach has proven successful in reducing errors in PSRA, hastening security risk detection, and minimizing false positives and negatives. This translates to cost savings for the company by averting unnecessary expenses associated with implementing unwarranted countermeasures. Therefore, experts can focus on more comprehensive risk analysis, leveraging LLMs for an effective preliminary assessment within a condensed timeframe.
Matteo Esposito 0001, Francesco Palagiano
EASE1
2024 Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
abstract
Context. Risk analysis assesses potential risks in specific scenarios. Risk analysis principles are context-less; the same methodology can be applied to a risk connected to health and information technology security. Risk analysis requires a vast knowledge of national and international regulations and standards and is time and effort-intensive. A large language model can quickly summarize information in less time than a human and can be fine-tuned to specific tasks.
Matteo Esposito 0001, Francesco Palagiano, Valentina Lenarduzzi, Davide Taibi 0001
ESEM1
2024 6GSoft: Software for Edge-to-Cloud Continuum
abstract
In the era of 6G, developing and managing software requires cutting-edge software engineering (SE) theories and practices tailored for such complexity across a vast number of connected edge devices. Our project aims to lead the development of sustainable methods and energy-efficient orchestration models specifically for edge environments, enhancing architectural support driven by AI for contemporary edge-to-cloud continuum computing. This initiative seeks to position Finland at the forefront of the 6G landscape, focusing on sophisticated edge orchestration and robust software architectures to optimize the performance and scalability of edge networks. Collaborating with leading Finnish universities and companies, the project emphasizes deep industry-academia collaboration and international expertise to address critical challenges in edge orchestration and software architecture, aiming to drive significant advancements in software productivity and market impact.
Muhammad Azeem Akbar, Matteo Esposito 0001, Sami Hyrynsalmi, Karthikeyan Dinesh Kumar, Valentina Lenarduzzi, Xiaozhou Li 0002, Ali Mehraj, Tommi Mikkonen, Sergio Moreschini, Niko Mäkitalo, Markku Oivo, Anna-Sofia Paavonen, Risha Parveen, Kari Smolander, Ruoyu Su, Kari Systä, Davide Taibi 0001, Zheying Zhang, Muhammad Zohaib
SEAA2
2024 VALIDATE: A deep dive into vulnerability prediction datasets
abstract
Vulnerabilities are an essential issue today, as they cause economic damage to the industry and endanger our daily life by threatening critical national security infrastructures. Vulnerability prediction supports software engineers in preventing the use of vulnerabilities by malicious attackers, thus improving the security and reliability of software. Datasets are vital to vulnerability prediction studies, as machine learning models require a dataset. Dataset creation is time-consuming, error-prone, and difficult to validate. This study aims to characterise the datasets of prediction studies in terms of availability and features. Moreover, to support researchers in finding and sharing datasets, we provide the first VulnerAbiLty predIction DatAseT rEpository (VALIDATE). We perform a systematic literature review of the datasets of vulnerability prediction studies. Our results show that out of 50 primary studies, only 22 studies (i.e., 38%) provide a reachable dataset. Of these 22 studies, only one study provides a dataset in a stable repository. Our repository of 31 datasets, 22 reachable plus nine datasets provided by authors via email, supports researchers in finding datasets of interest, hence avoiding reinventing the wheel; this translates into less effort, more reliability, and more reproducibility in dataset creation and use.
Matteo Esposito 0001, Davide Falessi
Inf. Softw. Technol.1
2023 Uncovering the Hidden Risks: The Importance of Predicting Bugginess in Untouched Methods
abstract
Bugs in untouched code can be a ticking time bomb, but are they worth predicting? Our study dives deep into the importance of predicting bugginess in untouched methods and its impact on bug prediction accuracy. Our analysis of six open-source projects reveals the hidden risks of dormant bugs and the benefits of predicting in isolation untouched methods. Our findings significantly increase prediction accuracy, decrease train dataset sizes, and prove that untouched methods are an overlooked yet vital aspect of bug prediction.
Matteo Esposito 0001, Davide Falessi
SCAM1
2023 Can We Trust the Default Vulnerabilities Severity?
abstract
As software systems become increasingly complex and interconnected, the risk of security debt has risen significantly, increasing cyber-attacks and data breaches. Vulnerability prioritization is a critical activity in software engineering as it helps identify and address security vulnerabilities in software systems promptly and effectively. With the increasing complexity of software systems and the growing number of potential threats, it is essential to have a systematic approach to vulnerability prioritization to ensure that the most critical vulnerabilities are addressed first. The present study aims to investigate the agreement between the default and the National Vulnerability Database (NVD) severity levels. We analyzed 1626 vulnerabilities encompassing 12 unique types of vulnerabilities associated with 125 Common Platform Enumeration identifiers belonging to 105 Apache projects. Our results show a scarce correlation between the default and NVD severity levels. Thus, the default severity of vulnerabilities is not trustworthy. Moreover, we discovered that, surprisingly, the same type of vulnerability has several NVD severity; therefore, no default prioritization can be accurate based only on the type of vulnerability. Future studies are needed to accurately estimate the priority of vulnerabilities by considering several aspects of vulnerabilities rather than only the type.
Matteo Esposito 0001, Sergio Moreschini, Valentina Lenarduzzi, David Hästbacka, Davide Falessi
SCAM1
2023 Enhancing the defectiveness prediction of methods and classes via JIT
abstract
Abstract Context Defect prediction can help at prioritizing testing tasks by, for instance, ranking a list of items (methods and classes) according to their likelihood to be defective. While many studies investigated how to predict the defectiveness of commits, methods, or classes separately, no study investigated how these predictions differ or benefit each other. Specifically, at the end of a release, before the code is shipped to production, testing can be aided by ranking methods or classes, and we do not know which of the two approaches is more accurate. Moreover, every commit touches one or more methods in one or more classes; hence, the likelihood of a method and a class being defective can be associated with the likelihood of the touching commits being defective. Thus, it is reasonable to assume that the accuracy of methods-defectiveness-predictions (MDP) and the class-defectiveness-predictions (CDP) are increased by leveraging commits-defectiveness-predictions (aka JIT). Objective The contribution of this paper is fourfold: (i) We compare methods and classes in terms of defectiveness and (ii) of accuracy in defectiveness prediction, (iii) we propose and evaluate a first and simple approach that leverages JIT to increase MDP accuracy and (iv) CDP accuracy. Method We analyse accuracy using two types of metrics (threshold-independent and effort-aware). We also use feature selection metrics, nine machine learning defect prediction classifiers, more than 2.000 defects related to 38 releases of nine open source projects from the Apache ecosystem. Our results are based on a ground truth with a total of 285,139 data points and 46 features among commits, methods and classes. Results Our results show that leveraging JIT by using a simple median approach increases the accuracy of MDP by an average of 17% AUC and 46% PofB10 while it increases the accuracy of CDP by an average of 31% AUC and 38% PofB20. Conclusions From a practitioner’s perspective, it is better to predict and rank defective methods than defective classes. From a researcher’s perspective, there is a high potential for leveraging statement-defectiveness-prediction (SDP) to aid MDP and CDP.
Davide Falessi, Simone Mesiano Laureani, Jonida Çarka, Matteo Esposito 0001, Daniel Alencar da Costa
Empir. Softw. Eng.4
2022 On effort-aware metrics for defect prediction
abstract
Abstract Context Advances in defect prediction models, aka classifiers, have been validated via accuracy metrics. Effort-aware metrics (EAMs) relate to benefits provided by a classifier in accurately ranking defective entities such as classes or methods. PofB is an EAM that relates to a user that follows a ranking of the probability that an entity is defective, provided by the classifier. Despite the importance of EAMs, there is no study investigating EAMs trends and validity. Aim The aim of this paper is twofold: 1) we reveal issues in EAMs usage, and 2) we propose and evaluate a normalization of PofBs (aka NPofBs), which is based on ranking defective entities by predicted defect density. Method We perform a systematic mapping study featuring 152 primary studies in major journals and an empirical study featuring 10 EAMs, 10 classifiers, two industrial, and 12 open-source projects. Results Our systematic mapping study reveals that most studies using EAMs use only a single EAM (e.g., PofB20) and that some studies mismatched EAMs names. The main result of our empirical study is that NPofBs are statistically and by orders of magnitude higher than PofBs. Conclusions In conclusion, the proposed normalization of PofBs: (i) increases the realism of results as it relates to a better use of classifiers, and (ii) promotes the practical adoption of prediction models in industry as it shows higher benefits. Finally, we provide a tool to compute EAMs to support researchers in avoiding past issues in using EAMs.
Jonida Çarka, Matteo Esposito 0001, Davide Falessi
Empir. Softw. Eng.2