Daniel Feitosa

dblp:20/8385 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0001-9371-232XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 41 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Mining Kubernetes Repositories: The Cloud was Not Built in a Day
abstract
We present MKR: Mining Kubernetes Repositories, a dataset capturing more than eleven years of development and community interaction in Kubernetes—an open-source platform for automating the deployment, scaling, and management of containerized applications. As the infrastructure backbone for running thousands of applications across diverse environments, Kubernetes has become one of the most widely adopted and influential projects in modern cloud-native computing. Spanning from June 2014 to July 2025, MKR integrates over two million artefacts from GitHub, including 130,832 commits (through July 2025), 83,368 pull requests, 46,768 issues, and 1,795,423 comments (through March 2025). With contributions from 28,890 unique GitHub commenters and 4,931 commit authors, MKR provides a longitudinal record of how Kubernetes has evolved, scaled, and been maintained over time. The dataset supports research on code evolution, long-term maintenance practices such as API deprecation, contributor retention, governance, and the role of automation in development. MKR allows analyses that connect technical change with decision-making, offering a resource for examining the social and technical dimensions of large-scale open source projects.
Giuseppe Destefanis, Silvia Bartolucci, Daniel Feitosa
MSR3
2026 Evolving Kubernetes: A Technical Debt Perspective
Jesse Maarleveld, Giuseppe Destefanis, Daniel Feitosa
MSR3
2026 Testing with AI Agents: An Empirical Study of Test Generation Frequency, Quality, and Coverage
abstract
Agent-based coding tools have transformed software development practices. Unlike prompt-based approaches that require developers to manually integrate generated code, these agent-based tools autonomously interact with repositories to create, modify, and execute code, including test generation. While many developers have adopted agent-based coding tools, little is known about how these tools generate tests in real-world development scenarios or how AI-generated tests compare to human-written ones.
Suzuka Yoshimoto, Shun Fujita, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, Hajimu Iida
MSR4
2026 Investigating CI/CD-based Technical Debt Management in Open-source Projects
abstract
Managing technical debt (TD) is critical to ensure the sustainability of long-term software projects. However, the time and cost involved in technical debt management (TDM) often discourage practitioners from performing this activity consistently. Continuous Integration and Continuous Delivery (CI/CD) pipelines offer an opportunity to support TDM by embedding automated practices directly into the development workflow. Despite this potential, it remains unclear how TDM tools could be integrated into CI/CD pipelines, and we still lack established best practices for this process. To address this problem, the objective of this study is to understand how TDM tools have been used in CI/CD pipelines and also identify potential configuration anti-patterns. To this end, we conducted a large-scale mining software repository (MSR) study on GitHub. In total, we collected around 600,000 Travis CI configuration files and 50,000 supporting scripts, and identified 3,684 pipelines that contain at least one TDM tool. We applied descriptive statistics to analyze the prevalence of tools and anti-patterns, and our findings show that most tools are executed and integrated using an external script; in addition, Absent Feedback is the most common configuration anti-pattern. We believe that researchers and practitioners can use the evidence of this study to further investigate how to improve both the tools that are integrated in CI/CD and the integration practices.
João Paulo Biazotto, Daniel Feitosa, Paris Avgeriou, Elisa Yumi Nakagawa
TechDebt@ICSE2
2026 How Do Practitioners Manage Traceability of Technical Debt in Continuous Software Engineering?
abstract
Continuous software engineering (CSE) has become essential for delivering flexible, market-driven software solutions by integrating development, operations, and business strategy under agile principles. CSE practices also lead to the accumulation of technical debt (TD), highlighting the importance of effective technical debt management (TDM). Traceability can play a key role in TDM by linking TD items to past decisions throughout the software life cycle. However, TD traceability remains underexplored in the literature. This study investigates how software practitioners manage the traceability of TD in CSE environments. We conducted eight semi-structured interviews to understand existing processes, tools used, and challenges. Findings reveal that TD traceability is generally ad hoc, lacks standardized practices, and is primarily supported by tools that focus on visualizing TD in backlogs without preserving decision rationale. These results point out to opportunities for future research to enhance TD traceability in CSE.
Lucas Carvalho, João Paulo Biazotto, Daniel Feitosa, Elisa Yumi Nakagawa
TechDebt@ICSE3
2026 TagDebt: a bot to support technical debt management
abstract
Abstract Context Technical debt (TD) is a widely studied metaphor that helps to explain how sub-optimal decisions, which usually have short-term benefits, can harm software maintainability over time. Although incurring TD is not intrinsically bad, tracking and managing TD are crucial to avoid its negative effects. Hence, researchers and practitioners have proposed and developed diverse approaches and tools for managing TD. However, we are still lacking specialized tools for technical debt management (TDM), specifically ones that can be easily integrated into existing development workflows. Objective We present and evaluate TagDebt, a bot that can be integrated within GitHub repositories and automatically assign labels to issues (i.e., SATD or non-SATD). TagDebt helps in the identification of TD (i.e., by looking for self-admitted technical debt (SATD)), leading to more efficient TDM. Methods We carried out a Design Science Research study to design and implement TagDebt. For its evaluation, we executed a Technology Acceptance Model (TAM) study through interviews with 16 practitioners, to check the bot’s usefulness, ease of use, and contextual factors that might impact the bot’s usage (such as team size and practitioners’ roles). Results Overall, practitioners found that TagDebt is useful, especially for organizing issues and reducing manual work. Furthermore, they pointed out that the bot is overall easy to use, and its documentation is clear. The analysis also revealed that contextual factors, such as team and codebase size, impact the decision to adopt TagDebt. Finally, several improvements were suggested, such as including features to check and update the source code. Conclusion TagDebt is a proof-of-concept for the development and usage of more specialized tools for TDM. It helps to make TD visible without disrupting existing workflows, which could lead to increased adoption of TDM tools and, consequently, help practitioners avoid the risks of unmanaged TD.
João Paulo Biazotto, Daniel Feitosa, Paris Avgeriou, Elisa Yumi Nakagawa
Empir. Softw. Eng.2
2026 Towards sustainable cloud deployments: A cost (anti)patterns catalog for terraform and CloudFormation
abstract
Infrastructure as Code (IaC) solutions such as HashiCorp’s Terraform and Amazon Web Services’ CloudFormation are invaluable tools in dealing with the increasing complexity of deploying software systems continuously on the cloud, and the need to manage larger and more complex infrastructures. However, not many existing works have examined the cost implications of IaC adoption for cloud-based software systems. In this work, we apply thematic analysis to 2289 file diffs, spanning 828 commits from 618 repositories, to identify recurring solutions and ineffective practices in cost management of Terraform and CloudFormation artifacts. We uncover a catalog of five patterns and seven antipatterns, and analyze their (co-)occurrences. Our results indicate that many teams address cost reactively by making incremental fixes, while others take proactive steps such as configuring budgets, integrating cost reports, and designing preventative templates. To aid practitioners in catching these issues early, we also present a linter as an extension of the Checkov tool that automates detection of selected (anti)patterns in Terraform and CloudFormation files, offering real-time feedback on potential oversights. We evaluated the utility of this tool by applying it to 182 active open-source repositories and soliciting practitioner feedback. Our results reveal that while cost-related misconfigurations are widespread, developer engagement is limited, suggesting that cost optimization is often prioritized lower than security or functionality and addressed reactively. Together, this catalog and detection tool provide actionable insights for optimizing cloud resource management, enhancing cost efficiency, and fostering informed decision-making in cloud deployments.
Koen Bolhuis, Allia Neamt, Daniel Feitosa, Vasilios Andrikopoulos
J. Syst. Softw.3
2025 PairSmell: A Novel Perspective Inspecting Software Modular Structure
abstract
Enhancing the modular structure of existing systems has attracted substantial research interest, focusing on two main methods: (1) software modularization and (2) identifying design issues (e.g., smells) as refactoring opportunities. However, remodularization solutions often require extensive modifications to the original modules, and the design issues identified are generally too coarse to guide refactoring strategies. Combining the above two methods, this paper introduces a novel concept, PairSmell, which exploits modularization to pinpoint design issues necessitating refactoring. We concentrate on a granular but fundamental aspect of modularity principles-modular relation (MR), i.e., whether a pair of entities are separated or collocated. The main assumption is that, if the actual MR of a pair violates its ‘apt MR’, i.e., an MR agreed on by multiple modularization tools (as raters), it can be deemed likely a flawed architectural decision that necessitates further examination. To quantify and evaluate PairSmell, we conduct an empirical study on 20 C/C++ and Java projects, using 4 established modularization tools to identify two forms of PairSmell: inapt separated pairs$InSep$and inapt collocated pairs$InCol$. Our study on 260,003 instances reveals that their architectural impacts are substantial: (1) on average, 14.60 % and 20.44 % of software entities are involved in$InSep$and$InCol$MRs respectively; (2)$InSep$pairs are associated with 190 % more co-changes than properly separated pairs, while$InCol$pairs are associated with 35% fewer co-changes than properly collocated pairs, both indicating a successful identification of modular structures detrimental to software quality; and (3) both forms of PairSmell persist across software evolution. This evidence strongly suggests that PairSmell can provide meaningful insights for inspecting modular structure, with the identified issues being both granular and fundamental, making the enhancement of modular design more efficient.
Chenxing Zhong, Daniel Feitosa, Paris Avgeriou, Yue Li 0047, He Zhang 0001
ICSE2
2025 Automating Technical Debt Management: Insights from Practitioner Discussions in Stack Exchange
abstract
Managing technical debt (TD) is essential for maintaining long-term software projects. Nonetheless, the time and cost involved in technical debt management (TDM) are often high, which may lead practitioners to omit TDM tasks. The adoption of tools, and particularly the usage of automated solutions, can potentially reduce the time, cost, and effort involved. However, the adoption of tools remains low, indicating the need for further research on TDM automation. To address this problem, this study aims at understanding which TDM activities practitioners are discussing with respect to automation in TDM, what tools they report for automating TDM, and the challenges they face that require automated solutions. To this end, we conducted a mining software repositories (MSR) study on three websites of Stack Exchange (Stack Overflow, Project Management, and Software Engineering) and collected 216 discussions, which were analyzed using both thematic synthesis and descriptive statistics. We found that identification and measurement are the most cited activities. Furthermore, 51 tools were reported as potential alternatives for TDM automation. Finally, a set of nine main challenges were identified and clustered into two main categories: challenges driving TDM automation and challenges related to tool usage. These findings highlight that tools for automating TDM are being discussed and used; however, several significant barriers persist, such as tool errors and poor explainability, hindering the adoption of these tools. Moreover, further research is needed to investigate the automation of other TDM activities such as TD prioritization.
João Paulo Biazotto, Daniel Feitosa, Paris Avgeriou, Elisa Yumi Nakagawa
TechDebt2
2025 Understanding practitioners' reasoning and requirements for efficient tool support in technical debt management
abstract
Abstract Context Maintaining software projects over the long term requires controlling the accumulation of technical debt (TD). However, the time and cost associated with technical debt management (TDM) are often high, hindering practitioners from performing TDM tasks. Using tools for TDM has the potential to reduce the effort involved. Despite this, the adoption of such tools remains low, indicating a need for more efficient tool support. Objective This study aims to understand practitioners’ perspectives on tool support for TDM, specifically regarding the selection and use of these tools. Additionally, we identified potential requirements that could be implemented into existing or new TD tools. Method We surveyed practitioners and received 103 answers, from which 89 valid answers were analyzed using thematic synthesis and descriptive statistics. Results Practitioners’ decision-making processes regarding adopting tools are primarily driven by ten main concerns identified from practitioners’ responses (e.g., the load of information provided by tools). Additionally, we elicited 46 requirements and classified them into two main categories (“Information to be provided” and “Tool Usage”). Conclusion Practitioners aim to maintain control over tool execution and outputs. Our study then highlights the necessity of human-centered approaches for TDM automation, i.e., not only tools are essential, but the interaction between tools and practitioners is critical for a more efficient TDM.
João Paulo Biazotto, Daniel Feitosa, Paris Avgeriou, Elisa Yumi Nakagawa
Empir. Softw. Eng.2
2025 A systematic mapping study on graph machine learning for static source code analysis
abstract
Context: In recent years, graph machine learning and particularly graph neural networks have seen successful and widespread applications in many fields, including static source code analysis. Such machine learning techniques enable learning on rich information networks capable of representing different relations and entities. However, there have been no comprehensive studies investigating the use of graph machine learning for static source code analysis. There is no complete systematic picture of what techniques may be considered tried and tested, and where opportunities for future improvements can still be found. The main goal of this study is to provide a broad overview of the state of the art of static source code analysis using graph machine learning. A systematic mapping was performed covering 4499 studies, presenting a final selection of 323 primary studies. Among the selected studies, seven major sub-domains were identified. The use and combinations of artefacts, different graph representations, different features, and different machine learning models used were collected and categorised. The use of graph learning, and in particular graph neural networks, has increased significantly since 2018. Although a wide variety of methods is used, across every dimension we investigated (artefacts, graphs, features, models), we found small sets of technologies which are used in the vast majority of studies. Future opportunities lie in exploring under-explored domains more thoroughly, exploring the use of additional artefacts alongside source code, and paying more attention to interpretability and explainability.
Jesse Maarleveld, Jiapan Guo, Daniel Feitosa
Inf. Softw. Technol.3
2024 A Catalog of Cost Patterns and Antipatterns for Infrastructure as Code
abstract
Cloud adoption is historically driven by cost considerations. As the complexity of the software systems deployed on the cloud continuously increases, and with it also the need to manage larger and more complex infrastructures, Infrastructure as Code (IaC) approaches become invaluable tools. However, not many existing works have looked into the cost implications of IaC use for cloud-based software. In this work we build on an existing dataset that has looked into cost-related commits on IaC artifacts in open-source repositories in order to identify recurring solutions and ineffective practices in cost management. We present a catalog of patterns and antipatterns organizing our findings, and discuss its implication for practitioners and researchers.
Koen Bolhuis, Daniel Feitosa, Vasilios Andrikopoulos
SEAA2
2024 Software Engineering Practices in Smart Contract Development: A Systematic Mapping Study
Antonios Giatzis, Elvira-Maria Arvanitou, Danai Papadopoulou, Theodoros Maikantis, Nikolaos Nikolaidis 0003, Daniel Feitosa, Christos K. Georgiadis, Apostolos Ampatzoglou, Alexander Chatzigeorgiou, Evdokimos I. Konstantinidis, Panagiotis D. Bamidis
PROFES6
2024 A metrics-based approach for selecting among various refactoring candidates
Nikolaos Nikolaidis 0003, Nikolaos Mittas, Apostolos Ampatzoglou, Daniel Feitosa, Alexander Chatzigeorgiou
Empir. Softw. Eng.4
2024 Technical debt management automation: State of the art and future perspectives
abstract
Technical debt (TD) refers to non-optimal decisions made in software projects that may lead to short-term benefits, but potentially harm the system’s maintenance in the long-term. Technical debt management (TDM) refers to a set of activities that are performed to handle TD, e.g., identification or measurement of TD. These activities typically entail tasks such as code and architectural analysis, which can be time-consuming if done manually. Thus, substantial research work has focused on automating TDM tasks (e.g., automatic identification of code smells). However, there is a lack of studies that summarize current approaches in TDM automation. This can hinder practitioners in selecting optimal automation strategies to efficiently manage TD. It can also prevent researchers from understanding the research landscape and addressing the research problems that matter the most. The main objective of this study is to provide an overview of the state of the art in TDM automation, analyzing the available tools, their use, and the challenges in automating TDM. We conducted a systematic mapping study (SMS), following the guidelines proposed by Kitchenham et al. From an initial set of 1086 primary studies, 178 were selected to answer three research questions covering different facets of TDM automation. We found 121 automation artifacts that can be used to automate TDM activities. The artifacts were classified in 4 different types (i.e., tools, plugins, scripts, and bots); the inputs/outputs and interfaces were also collected and reported. Finally, a conceptual model is proposed that synthesizes the results and allows to discuss the current state of TDM automation and related challenges. The research community has investigated to a large extent how to perform various TDM activities automatically, considering the number of studies and automation artifacts we identified. Nonetheless, more research is needed towards fully automated TDM, specially concerning the integration of the automation artifacts.
João Paulo Biazotto, Daniel Feitosa, Paris Avgeriou, Elisa Yumi Nakagawa
Inf. Softw. Technol.2
2024 Mining for cost awareness in the infrastructure as code artifacts of cloud-based applications: An exploratory study
abstract
Cloud computing’s rise as the primary platform for software development and delivery is largely driven by the potential cost savings. However, it is surprising that no empirical evidence has been collected to determine whether cost awareness permeates the development process and how it manifests in practice. This study aims to provide empirical evidence of cost awareness by mining open source repositories of cloud-based applications. The focus is on Infrastructure-as-Code artifacts that automate software (re)deployment on the cloud. A systematic examination of 152735 repositories yielded 2010 relevant hits. We then analyzed 538 relevant commits and 208 relevant issues using inductive and deductive coding and corroborated findings with discussions from Stack Overflow. The findings indicate that developers are not only concerned with the cost of their application deployments but also take actions to reduce these costs beyond selecting cheaper cloud services. We also identify research areas for future consideration. Although we focus on a particular Infrastructure-as-Code technology (Terraform), the findings can be applicable to cloud-based application development in general. The provided empirical grounding can serve developers seeking to reduce costs through service selection, resource allocation, deployment optimization, and other techniques.
Daniel Feitosa, Matei-Tudor Penca, Massimiliano Berardi, Rares-Dorian Boza, Vasilios Andrikopoulos
J. Syst. Softw.1
2024 Eclipse Open SmartCLIDE: An end-to-end framework for facilitating service reuse in cloud development
abstract
Service-Oriented Architectures (SOA) have become a standard for developing software applications, including but not limited to cloud-based ones and enterprise systems. When using SOA, software engineers organize the desired functionality into self-contained and independent services that are invoked through end-points (with API calls). The use of this emerging technology has changed drastically the way that software reuse is performed, in the sense that a “ service ” is a “ code chunk ” that is reusable (preferably in a black-box manner), but in many (especially “ in-house ”) cases, white-box reuse is also meaningful. To confront the reuse challenges opened-up by the rise of SOA, in the SmartCLIDE project 1 we have developed a framework (a methodology and a platform) to aid software engineers in systematic and more efficient (in terms of time, quality, defects, and process) reuse of services, when developing SOA-based cloud applications. In this work, we: (a) present the SmartCLIDE methodology and the Eclipse Open SmartCLIDE platform; and (b) evaluate the usefulness of the framework, in terms of relevance, usability, and obtained benefits. The results of the study have confirmed the relevance and rigor of the framework, unveiled some limitations, and pointed to interesting future work directions, but also provided some actionable implications for researchers and practitioners.
Nikolaos Nikolaidis 0003, Elvira-Maria Arvanitou, Christina Volioti, Theodoros Maikantis, Apostolos Ampatzoglou, Daniel Feitosa, Alexander Chatzigeorgiou, Phillipe Krief
J. Syst. Softw.6
2023 Uncovering Energy-Efficient Practices in Deep Learning Training: Preliminary Steps Towards Green AI
abstract
Modern AI practices all strive towards the same goal: better results. In the context of deep learning, the term "results" often refers to the achieved accuracy on a competitive problem set. In this paper, we adopt an idea from the emerging field of $\color{green}{\text{Green AI}}$ to consider energy consumption as a metric of equal importance to accuracy and to reduce any irrelevant tasks or energy usage. We examine the training stage of the deep learning pipeline from a sustainability perspective, through the study of hyperparameter tuning strategies and the model complexity, two factors vastly impacting the overall pipeline’s energy consumption. First, we investigate the effectiveness of grid search, random search and Bayesian optimisation during hyperparameter tuning, and we find that Bayesian optimisation significantly dominates the other strategies. Furthermore, we analyse the architecture of convolutional neural networks with the energy consumption of three prominent layer types: convolutional, linear and ReLU layers. The results show that convolutional layers are the most computationally expensive by a strong margin. Additionally, we observe diminishing returns in accuracy for more energy-hungry models. The overall energy consumption of training can be halved by reducing the network complexity. In conclusion, we highlight innovative and promising energy-efficient practices for training deep learning models. To expand the application of $\color{green}{\text{Green AI}}$, we advocate for a shift in the design of deep learning models, by considering the trade-off between energy efficiency and accuracy.
Tim Yarally, Luis Cruz 0002, Daniel Feitosa, June Sallou, Arie van Deursen
CAIN3
2023 Batching for Green AI - An Exploratory Study on Inference
abstract
The batch size is an essential parameter to tune during the development of new neural networks. Amongst other quality indicators, it has a large degree of influence on the model’s accuracy, generalisability, training times and parallelisability. This fact is generally known and commonly studied. However, during the application phase of a deep learning model, when the model is utilised by an end-user for inference, we find that there is a disregard for the potential benefits of introducing a batch size. In this study, we examine the effect of input batching on the energy consumption and response times of five fully-trained neural networks for computer vision that were considered state-of-the-art at the time of their publication. The results suggest that batching has a significant effect on both of these metrics. Furthermore, we present a timeline of the energy efficiency and accuracy of neural networks over the past decade. We find that in general, energy consumption rises at a much steeper pace than accuracy and question the necessity of this evolution. Additionally, we highlight one particular network, ShuffleNetV2 (2018), that achieved a competitive performance for its time while maintaining a much lower energy consumption. Nevertheless, we highlight that the results are model dependent.
Tim Yarally, Luis Cruz 0002, Daniel Feitosa, June Sallou, Arie van Deursen
SEAA3
2023 The lifecycle of Technical Debt that manifests in both source code and issue trackers
abstract
Context: Although Technical Debt (TD) has increasingly gained attention in recent years, most studies exploring TD are based on a single source (e.g., source code, code comments or issue trackers). Objective: Investigating information combined from different sources may yield insight that is more than the sum of its parts. In particular, we argue that exploring how TD items are managed in both issue trackers and software repositories (including source code and commit messages) can shed some light on what happens between the commits that incur TD and those that pay it back. Method: To this end, we randomly selected 3,000 issues from the trackers of five projects, manually analyzed 300 issues that contained TD information, and identified and investigated the lifecycle of 312 TD items. Results: The results indicate that most of the TD items marked as resolved in issue trackers are also paid back in source code, although many are not discussed after being identified in the issue tracker. Test Debt items are the least likely to be paid back in source code. We also learned that although TD items may be resolved a few days after being identified, it often takes a long time to be identified (around one year). In general, time is reduced if the same developer is involved in consecutive moments (i.e., introduction, identification, repayment decision-making and remediation), but whether the developer who paid back the item is involved in discussing the TD item does not seem to affect how quickly it is resolved. Conclusions: Investigating how developers manage TD across both source code repositories and issue trackers can lead to a more comprehensive oversight of this activity and support efforts to shorten the lifecycle of undesirable debt.
Jie Tan 0002, Daniel Feitosa, Paris Avgeriou
Inf. Softw. Technol.2
2023 On measuring coupling between microservices
abstract
In software quality management , the selection strategy for proper metrics varies depending on the application scenarios and measurement objectives. MicroService Architecture (MSA), despite being commonly employed nowadays, still cannot be reliably measured and compared if the microservices in a system are independent. Software managers and architects need to understand whether their microservices are “decoupled enough”, if not, which ones are over-coupled, and by how much. In this paper, we contribute a novel set of metrics – Microservice Coupling Index (MCI) – derived from the relative measurement theory. Instead of measuring coupling evidence with simple counts, we measure how dependent and coupled the microservices are relative to the possible couplings between them. We measured the MCI metrics for 15 open source projects that involve 113 distinct microservices. Empirical investigation confirmed that MCIs differ quite significantly from existing coupling measures and that they are more discriminative than existing ones for separating high and low degrees of microservice couplings and thus more useful in comparing design alternatives. A series of experimental studies were conducted, showing that the larger the MCIs, the less likely the bugs and changes can be localized and separated, and the less likely that the individual microservices in a system can be independently developed and evolved.
Chenxing Zhong, He Zhang 0001, Daniel Feitosa
J. Syst. Softw.5
2022 Service Classification through Machine Learning: Aiding in the Efficient Identification of Reusable Assets in Cloud Application Development
abstract
Developing software based on services is one of the most emerging programming paradigms in software development. Service-based software development relies on the composition of services (i.e., pieces of code already built and deployed in the cloud) through orchestrated API calls. Black-box reuse can play a prominent role when using this programming paradigm, in the sense that identifying and reusing already existing/deployed services can save substantial development effort. According to the literature, identifying reusable assets (i.e., components, classes, or services) is more successful and efficient when the discovery process is domain-specific. To facilitate domain-specific service discovery, we propose a service classification approach that can categorize services to an application domain, given only the service description. To validate the accuracy of our classification approach, we have trained a machine-learning model on thousands of open-source services and tested it on 67 services developed within two companies employing service-based software development. The study results suggest that the classification algorithm can perform adequately in a test set that does not overlap with the training set; thus, being (with some confidence) transferable to other industrial cases. Additionally, we expand the body of knowledge on software categorization by highlighting sets of domains that consist ‘grey-zones’ in service classification.
Zakieh Alizadehsani, Daniel Feitosa, Theodoros Maikantis, Apostolos Ampatzoglou, Alexander Chatzigeorgiou, David Berrocal-Macías, Alfonso González-Briones, Juan M. Corchado, Marcio Mateus, Johannes Groenewold
SEAA2
2022 Does it matter who pays back Technical Debt? An empirical study of self-fixed TD
abstract
Technical Debt (TD) can be paid back either by those that incurred it or by others. We call the former self-fixed TD, and it can be particularly effective, as developers are experts in their own code and are well-suited to fix the corresponding TD issues. The goal of our study is to investigate self-fixed technical debt, especially the extent in which TD is self-fixed, which types of TD are more likely to be self-fixed, whether the remediation time of self-fixed TD is shorter than non-self-fixed TD and how development behaviors are related to self-fixed TD. We report on an empirical study that analyzes the self-fixed issues of five types of TD (i.e., Code, Defect, Design, Documentation and Test), captured via static analysis, in more than 44,000 commits obtained from 20 Python and 16 Java projects of the Apache Software Foundation. The results show that about half of the fixed issues are self-fixed and that the likelihood of contained TD issues being self-fixed is negatively correlated with project size, the number of developers and total issues. Moreover, there is no significant difference of the survival time between self-fixed and non-self-fixed issues. Furthermore, developers are more keen to pay back their own TD when it is related to lower code level issues, e.g., Defect Debt and Code Debt. Finally, developers who are more dedicated to or knowledgeable about the project contribute to a higher chance of self-fixing TD. These results can benefit both researchers and practitioners by aiding the prioritization of TD remediation activities and refining strategies within development teams, and by informing the development of TD management tools.
Jie Tan 0002, Daniel Feitosa, Paris Avgeriou
Inf. Softw. Technol.2
2021 Do practitioners intentionally self-fix Technical Debt and why?
abstract
The impact of Technical Debt (TD) on software maintenance and evolution is of great concern, but recent evidence shows that a considerable amount of TD is fixed by the same developers who introduced it; this is termed self-fixed TD. This characteristic of TD management can potentially impact team dynamics and practices in managing TD. However, the initial evidence is based on low-level source code analysis; this casts some doubt whether practitioners repay their own debt intentionally and under what circumstances. To address this gap, we conducted an online survey on 17 well-known Java and Python open-source software communities to investigate practitioners' intent and rationale for self-fixing technical debt. We also investigate the relationship between human-related factors (e.g., experience) and self-fixing. The results, derived from the responses of 181 participants, show that a majority addresses their own debt consciously and often. Moreover, those with a higher level of involvement (e.g., more experience in the project and number of contributions) tend to be more concerned about self-fixing TD. We also learned that the sense of responsibility is a common self-fixing driver and that decisions to fix TD are not superficial but consider balancing costs and benefits, among other factors. The findings in this paper can lead to improving TD prevention and management strategies.
Jie Tan 0002, Daniel Feitosa, Paris Avgeriou
ICSME2
2021 Software reuse cuts both ways: An empirical analysis of its relationship with security vulnerabilities
Antonios Gkortzis, Daniel Feitosa, Diomidis Spinellis
J. Syst. Softw.2
2021 Evolution of technical debt remediation in Python: A case study on the Apache Software Ecosystem
abstract
Abstract In recent years, the evolution of software ecosystems and the detection of technical debt received significant attention by researchers from both industry and academia. While a few studies that analyze various aspects of technical debt evolution already exist, to the best of our knowledge, there is no large‐scale study that focuses on the remediation of technical debt over time in Python projects—that is, one of the most popular programming languages at the moment. In this paper, we analyze the evolution of technical debt in 44 Python open‐source software projects belonging to the Apache Software Foundation. We focus on the type and amount of technical debt that is paid back. The study required the mining of over 60K commits, detailed code analysis on 3.7K system versions, and the analysis of almost 43K fixed issues. The findings show that most of the repayment effort goes into testing, documentation, complexity, and duplication removal. Moreover, more than half of the Python technical debt is short term being repaid in less than 2 months. In particular, the observations that a minority of rules account for the majority of issues fixed and spent effort suggest that addressing those kinds of debt in the future is important for research and practice.
Jie Tan 0002, Daniel Feitosa, Paris Avgeriou, Mircea Lungu
J. Softw. Evol. Process.2
2020 Investigating the Relationship between Co-occurring Technical Debt in Python
abstract
Technical debt (TD) reflects issues that may negatively affect software maintenance and evolution. There is currently little evidence on how the different types of TD co-occur; for example, how code smells and design smells affect the same part of the system. This paper investigates how different types of TD co-occur, as well as the time period of the co-occurrence. To that end, we analyzed the co-occurring associations between five types of TD, captured in 42 SonarQube rules, in 3862 files of 20 Python projects from the Apache Software Foundation. We found that this phenomenon is dominant, affecting more than 90% of Python files. We also found that Documentation Debt and Test Debt appear in the majority of the files, although it seems to be mostly by coincidence. Finally, we noticed that co-occurrence of TD seems to happen very quickly: co-occurring issues tend to be introduced within the same week. But once it does happen, it is hard to get rid of. These results can benefit both researchers and practitioners by: aiding the prioritization of TD remediation; leading to novel tools for detecting co-occurring TD and warning potential issues; shedding further light on the explanation of how TD is introduced and can be mitigated.
Jie Tan 0002, Daniel Feitosa, Paris Avgeriou
SEAA2
2020 An empirical study on self-fixed technical debt
abstract
Technical Debt (TD) can be paid back either by those that incurred it or by others. We call the former self-fixed TD, and it is particularly effective, as developers are experts in their own code and are best-suited to fix the corresponding TD issues. To what extent is TD self-fixed, which types of TD are more likely to be self-fixed and is the remediation time of self-fixed TD shorter than non-self-fixed TD? This paper attempts to answer these questions. It reports on an empirical study that analyzes the self-fixed issues of five types of TD (i.e., Code, Defect, Design, Documentation and Test), captured via static analysis, in more than 17,000 commits from 20 Python projects of the Apache Software Foundation. The results show that more than two thirds of the issues are self-fixed and that the self-fixing rate is negatively correlated with the number of commits, developers and project size. Furthermore, the survival time of self-fixed issues is generally shorter than non-self-fixed issues. Moreover, the majority of Defect Debt tends to be self-fixed and has a shorter survival time, while Test Debt and Design Debt are likely to be fixed by other developers. These results can benefit both researchers and practitioners by aiding the prioritization of TD remediation activities within development teams, and by informing the development of TD management tools.
Jie Tan 0002, Daniel Feitosa, Paris Avgeriou
TechDebt@ICSE2
2020 CODE reuse in practice: Benefiting or harming technical debt
Daniel Feitosa, Apostolos Ampatzoglou, Antonios Gkortzis, Stamatia Bibi, Alexander Chatzigeorgiou
J. Syst. Softw.1
2020 Examining the reuse potentials of IoT application frameworks
abstract
The major challenge that a developer confronts when building IoT systems is the management of a plethora of technologies implemented with various constraints, from different manufacturers, that at the end need to cooperate. In this paper we argue that developers can benefit from IoT frameworks by reusing their components so as to build in less time and effort IoT systems that can easily integrate new technologies. In order to explore the reuse opportunities offered by IoT frameworks we have performed a case study and analyzed 503 components reused by 35 IoT projects. We examined (a) the types of functionality that are most facilitated for reuse (b) the reuse strategy that is most adopted (c) the quality of the reused components . The results of the case study suggest that the main functionality reused is the one related to the Device Management layer and that Black-box reuse is the main type. Moreover, the quality of the reused components is improved compared to the rest of the components built from scratch.
Paraskevi Smiari, Stamatia Bibi, Daniel Feitosa
J. Syst. Softw.3
2019 A Double-Edged Sword? Software Reuse and Potential Security Vulnerabilities
Antonios Gkortzis, Daniel Feitosa, Diomidis Spinellis
ICSR2
2019 Examining the Reusability of Smart Home Applications: A Case Study on Eclipse Smart Home
Paraskevi Smiari, Stamatia Bibi, Daniel Feitosa
ICSR3
2019 What can violations of good practices tell about the relationship between GoF patterns and run-time quality attributes?
Daniel Feitosa, Apostolos Ampatzoglou, Paris Avgeriou, Alexander Chatzigeorgiou, Elisa Yumi Nakagawa
Inf. Softw. Technol.1
2017 The Evolution of Design Pattern Grime: An Industrial Case Study
Daniel Feitosa, Paris Avgeriou, Apostolos Ampatzoglou, Elisa Yumi Nakagawa
PROFES1
2017 Investigating the effect of design patterns on energy consumption
abstract
Abstract Gang of Four (GoF) patterns are well‐known best practices for the design of object‐oriented systems. In this paper, we aim at empirically assessing their relationship to energy consumption, ie, a performance indicator that has recently attracted the attention of both researchers and practitioners. To achieve this goal, we investigate pattern‐participating methods (ie, those that play a role within the pattern) and compare their energy consumption to the consumption of functionally equivalent alternative (nonpattern) solutions. We obtained the alternative solution by refactoring the pattern instances using well‐known transformations (eg, replace polymorphism with conditional statements). The comparison is performed on 169 methods of 2 GoF patterns (namely, State/Strategy and Template Method), retrieved from 2 well‐known open source projects. The results suggest that for the majority of cases the alternative design excels in terms of energy consumption. However, in some cases (eg, when the method is large in size or invokes many methods) the pattern solution presents similar or lower energy consumption. The outcome of our study can be useful to both researchers and practitioners, because we: (1) provide evidence on a possible negative effect of GoF patterns, and (2) can provide guidance on which cases the use of the pattern is not hurting energy consumption.
Daniel Feitosa, Rutger Alders, Apostolos Ampatzoglou, Paris Avgeriou, Elisa Yumi Nakagawa
J. Softw. Evol. Process.1
2014 Consolidating a Process for the Design, Representation, and Evaluation of Reference Architectures
abstract
Reference architectures have emerged as a special type of software architecture that achieves well-recognized understanding of specific domains, promoting reuse of design expertise and facilitating the development, standardization, and evolution of software systems. Because of their advantages, several reference architectures have been proposed and have been also successfully used, including in the industry. However, the most of these architectures are still built using an ad-hoc approach, lacking of a systematization to their construction. If existing, these approaches could motivate and promote the building of new architectures and also support evolution of existing ones. In this scenario, the main contribution of this paper is to present the evolution of ProSA-RA, a process that systematizes the design, representation, and evaluation of reference architectures. ProSA-RA has been already applied in the establishment of reference architectures for different domains and this experience was used to evolve our process. In this paper, we illustrate an application of ProSA-RA in the robotics domain. Results achieved through the use of ProSA-RA have showed us that it is a viable, efficient process and, as a consequence, it could contribute to the reuse of knowledge in several applications domains, by promoting the establishment of new reference architectures.
Elisa Yumi Nakagawa, Milena Guessi, José Carlos Maldonado, Daniel Feitosa, Flávio Oquendo
WICSA4
2013 A Checklist for Evaluation of Reference Architectures of Embedded Systems (S)
José Filipe Marreiros Santos, Milena Guessi, Matthias Galster, Daniel Feitosa, Elisa Yumi Nakagawa
SEKE4
2011 Current State of Reference Architectures in the Context of Agile Methodologies
Vinícius Augusto Tagliatti Zani, Daniel Feitosa, Elisa Yumi Nakagawa
SEKE2
2010 An Approach Based on Visual Text Mining to Support Categorization and Classification in the Systematic Mapping
Kátia Romero Felizardo, Elisa Yumi Nakagawa, Daniel Feitosa, Rosane Minghim, José Carlos Maldonado
EASE3
2010 Reference Models and Reference Architectures Based on Service-Oriented Architecture: A Systematic Review
Lucas B. R. Oliveira, Kátia Romero Felizardo, Daniel Feitosa, Elisa Yumi Nakagawa
ECSA3
2010 Software Engineering in the Embedded Software and Mobile Robot Software Development: A Systematic Mapping
Daniel Feitosa, Kátia Romero Felizardo, Lucas B. R. Oliveira, Denis F. Wolf, Elisa Yumi Nakagawa
SEKE1