Zadia Codabux

dblp:150/3264 · also Zadia Codabux-Rossan · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-6715-3341ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 28 · 7 first-author · 23 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Investigating Technical Debt Types, Issues, and Solutions in Serverless Computing
abstract
Serverless computing is a cloud execution model where developers run code, and the server management is handled by the cloud provider. Serverless computing is increasingly gaining popularity as more systems adopt it to enhance scalability and reduce operational costs. While it has numerous benefits, it also embodies unique challenges inherent to serverless computing. One such challenge is Technical Debt (TD), which is exacerbated by the complexities of the serverless paradigm. While prior work has investigated the activities and bad practices that lead to TD in serverless computing, there remains a gap in understanding how TD manifests, the challenges it poses, and the solutions proposed to address TD issues in serverless systems. This study aims to investigate TD in the serverless context using Stack Overflow (SO) as a knowledge base. We collected 78,867 serverless questions on SO and labeled them as TD or non-TD using deep learning. Moreover, we conducted an in-depth analysis to identify types of TD in serverless settings, associated issues, and proposed solutions. We found that 37% of the serverless questions on SO are TD-related. We also identified six serverless-specific issues. Our research highlights the need for tools that can effectively detect TD in serverless applications.
Hasini Sumalee Perera, Zadia Codabux, Fabio Palomba
TechDebt@ICSE2
2026 Editorial of the special issue in the journal of systems and software on managing technical debt in software-intensive products and services
Zadia Codabux, Rodrigo O. Spínola, Carolyn B. Seaman, Matthias Galster
J. Syst. Softw.1
2025 Racing Against the Clock: Exploring the Impact of Scheduled Deadlines on Technical Debt
Joshua Aldrich Edbert, Zadia Codabux, Roberto Verdecchia
EASE2
2025 Investigating the Understandability of Review Comments on Code Change Requests
abstract
Code review is a widely adopted quality assurance practice in software engineering, where expert reviewers assess developers’ code changes before merging. While prior studies have explored review comment quality and usefulness, they often overlook the clarity and understandability of Code Change Request (CCR) comments. Unclear CCR comments can pose significant challenges for developers to address. Therefore, this study investigates the prevalence and impact of confusing or unclear CCR comments and proposes two approaches to enhance CCR communication during code review. Using a dataset of 182 open-source GitHub projects with over 55 K pull requests and 466 K CCR comments, we analyzed how often unclear comments occur and their effects on the review process. Our classifier, built from manually annotated developers’ replies in response to CCR comments, revealed that $24 \%$ of comments led to author confusion. Statistical analysis shows that unclear CCR comments significantly increase resolution time and discussion length, and that pull requests with clear CCR comments are more likely to be addressed and merged. A manual analysis of 400 confusing CCR comments identified six key characteristics, with lack of clarity and unclear rationale being the most common. Our first approach, the confusion classifier, flags authors’ confusion to enable reviewers to clarify ambiguities promptly (recall of 0.96), while the second classifier enables reviewers to evaluate the clarity and understandability of their CCR comments (recall of 0.93). This pioneering study further provides recommendations for enhancing CCR comments and offering a foundation for future research to streamline the review process.
Md Shamimur Rahman, Zadia Codabux, Chanchal Kumar Roy
MSR2
2025 Attributes of a great requirements engineer
Larissa Barbosa L. Pinheiro, Sávio Freire, Rita Suzana Pitangueira Maciel, Manoel G. Mendonça, Marcos Kalinowski, Zadia Codabux, Rodrigo O. Spínola
J. Syst. Softw.6
2024 Integrating Feedback From Application Reviews Into Software Development
abstract
In application (app) development, effectively harnessing user feedback is crucial for enhancing app quality and user feedback. However, the vast and unstructured nature of user reviews often complicates these efforts, posing challenges in accurately capturing and integrating this feedback into the development processes. We automate the classification of issues in app reviews and examine how these issues correlate with code quality metrics (code smells and bug reports) and development activities (additions, deletions, and time to merge in pull requests). We aim to provide evidence-based guidance for effectively prioritizing and addressing user feedback. Employing a Mining Software Repositories (MSR) approach, we gathered and analyzed reviews from seven open-source Android apps. We evaluated the efficacy of three machine learning models-Support Vector Machines (SVM), BERT, and a fine-tuned GPT-3.5-for classifying issues in app reviews. The GPT-3.5 model achieved the highest accuracy at 95.0%. We found statistically significant correlations between the classified issues, code quality metrics, and development activities. However, these relationships varied across applications, highlighting the complex relationship between user feedback and the development process. Our study highlights the effectiveness of automated tools in identifying and classifying feedback within app reviews. Our automated approach enhances developers' ability to manage feedback effectively and supports optimal resource allocation to improve app quality and user feedback.
Omar Abdelaziz, Zadia Codabux, Kevin Schneider
APSEC2
2024 Are Large Language Models a Threat to Programming Platforms? An Exploratory Study
abstract
Background: Competitive programming platforms such as LeetCode, Codeforces, and HackerRank provide challenges to evaluate programming skills. Technical recruiters frequently utilize these platforms as a criterion for screening resumes. With the recent advent of advanced Large Language Models (LLMs) like ChatGPT, Gemini, and Meta AI, there is a need to assess their problem-solving ability on the programming platforms. Aims: This study aims to assess LLMs’ capability to solve diverse programming challenges across programming platforms with varying difficulty levels, providing insights into their performance in real-time and offline scenarios, comparing them to human programmers, and identifying potential threats to established norms in programming platforms. Method: This study utilized 98 problems from LeetCode and 126 from Codeforces, covering 15 categories and varying difficulty levels. Then, we participated in nine online contests from Codeforces and LeetCode. Finally, two certification tests were attempted on HackerRank to gain insights into LLMs’ real-time performance. Prompts were used to guide LLMs in solving problems, and iterative feedback mechanisms were employed. We also tried to find any possible correlation among the LLMs in different scenarios. Results: LLMs generally achieved higher success rates on LeetCode (e.g., ChatGPT at 71.43%) but faced challenges on Codeforces. While excelling in HackerRank certifications, they struggled in virtual contests, especially on Codeforces. Despite diverse performance trends, ChatGPT consistently performed well across categories, yet all LLMs struggled with harder problems and lower acceptance rates. In LeetCode archive problems, LLMs generally outperformed users in time efficiency and memory usage but exhibited moderate performance in live contests, particularly in harder Codeforces contests compared to humans. Conclusions: While not necessarily a threat, the performance of LLMs on programming platforms is indeed a cause for concern. With the prospect of more efficient models emerging in the future, programming platforms need to address this issue promptly.
Md Mustakim Billah, Palash Ranjan Roy, Zadia Codabux, Banani Roy
ESEM3
2024 Decoding Android Permissions: A Study of Developer Challenges and Solutions on Stack Overflow
abstract
Background: The Android permission system is a set of controls to regulate access to sensitive data and platform resources (e.g., cameras). The fast-evolving nature of Android permissions and inadequate documentation result in numerous challenges for third-party developers. Aims: This study investigates the permission-related challenges developers face and the solutions provided to resolve them on the crowdsourcing platform Stack Overflow. Method: We conducted qualitative and quantitative analyses on 3,327 permission-related questions and 3,271 corresponding answers. Results: We found that most questions are related to non-evolving SDK permissions that remain constant across various Android versions, emphasizing the lack of documentation. We also classify developers’ challenges into several categories: Documentation-Related, Problems with Dependencies, Debugging, Conceptual Understanding, and Implementation Issues. Conclusions: Our study indicates the need for clear, consistent documentation to guide the use of permissions and reduce developer misunderstandings, which can lead to potential misuse of Android permissions.
Sahrima Jannat Oishwee, Zadia Codabux, Natalia Stakhanova
ESEM2
2024 Review-Pulse: A Dashboard for Managing User Feedback for Android Applications
abstract
Due to the large volume of data and its unstructured nature, managing user feedback via application (app) reviews is a significant challenge for Android developers. This study presents a dashboard to streamline this process using advanced machine learning and analysis techniques. The dashboard employs a fine-tuned Generative Pretrained Transformer (GPT-3.5) model to detect and categorize issues in user reviews automatically. Additional dashboard features include sentiment and toxicity analysis to provide insights into user emotions, potentially negative feedback, and code analysis to identify code smells across different app versions. We conducted a pilot study to evaluate the usability and effectiveness of the dashboard. The results indicate that the dashboard is user-friendly and effective in helping developers manage user feedback and monitor code quality. However, certain limitations were identified, such as dependency on the quality of training data and potential inaccuracies in sentiment and toxicity analysis. This dashboard aims to aid developers in effectively managing app reviews, prioritizing issues, and maintaining high app quality to improve user satisfaction. Tool URL: https://tdresearchgroup.github.io/Review-Pulseldashboard/ Demo Video: https://youtu.be/cT6su8dqh2g
Omar Abdelaziz, Zadia Codabux, Kevin Schneider
ICSME2
2024 Large Language Model vs. Stack Overflow in Addressing Android Permission Related Challenges
abstract
The Android permission system regulates access to sensitive mobile device resources such as camera and location. To access these resources, third-party developers need to request permissions. However, the Android permission system is complex and fast-evolving, presenting developers with numerous challenges surrounding compatibility issues, misuse of permissions, and vulnerabilities related to permissions. Our study aims to explore whether Large Language Models (LLMs) can serve as a reliable tool to assist developers in using Android permissions correctly and securely, thereby reducing the risks of misuse and security vulnerabilities in apps. In our study, we analyzed 1,008 Stack Overflow questions related to Android permissions and their accepted answers. In parallel, we generate answers to these questions using a popular LLM tool, ChatGPT. We focused on how well the ChatGPT's responses align with the accepted answers on Stack Overflow. Our findings show that above 50% of ChatGPT's answers align with Stack Overflow's accepted answers. ChatGPT offers better-aligned responses for challenges related to Documentation and Conceptual Understanding, while it provides less aligned answers for Debugging-related issues. In addition, we found that ChatGPT provides more consistent answers for 73.27% questions. Our study demonstrates the potential for using LLMs such as ChatGPT as a supporting tool to help developers navigate Android permission-related problems.
Sahrima Jannat Oishwee, Natalia Stakhanova, Zadia Codabux
MSR3
2024 A catalog of metrics at source code level for vulnerability prediction: A systematic mapping study
abstract
Abstract Industry practitioners assess software from a security perspective to reduce the risks of deploying vulnerable software. Besides following security best practice guidelines during the software development life cycle, predicting vulnerability before roll‐out is crucial. Software metrics are popular inputs for vulnerability prediction models. The objective of this study is to provide a comprehensive review of the source code‐level security metrics presented in the literature. Our systematic mapping study started with 1451 studies obtained by searching the four digital libraries from ACM, IEEE, ScienceDirect, and Springer. After applying our inclusion/exclusion criteria as well as the snowballing technique, we narrowed down 28 studies for an in‐depth study to answer four research questions pertaining to our goal. We extracted a total of 685 code‐level metrics. For each study, we identified the empirical methods, quality measures, types of vulnerabilities of the prediction models, and shortcomings of the work. We found that standard machine learning models, such as decision trees, regressions, and random forests, are most frequently used for vulnerability prediction. The most common quality measures are precision, recall, accuracy, and ‐measure. Based on our findings, we conclude that the list of software metrics for measuring code‐level security is not universal or generic yet. Nonetheless, the results of our study can be used as a starting point for future studies aiming at improving existing security prediction models and a catalog of metrics for vulnerability prediction for software practitioners.
Zadia Codabux, Kazi Zakia Sultana, Md. Naseef-Ur-Rahman Chowdhury
J. Softw. Evol. Process.1
2023 Exploring Technical Debt in Security Questions on Stack Overflow
abstract
Background: Software security is crucial to ensure that the users are protected from undesirable consequences such as malware attacks which can result in loss of data and, subsequently, financial loss. Technical Debt (TD) is a metaphor incurred by suboptimal decisions resulting in long-term consequences such as increased defects and vulnerabilities if not managed. Although previous studies have studied the relationship between security and TD, examining their intersection in developers' discussion on Stack Overflow (SO) is still unexplored. Aims: This study investigates the characteristics of security-related TD questions on SO. More specifically, we explore the prevalence of TD in security-related queries, identify the security tags most prone to TD, and investigate which user groups are more aware of TD. Method: We mined 117,233 security-related questions on SO and used a deep-learning approach to identify 45,078 security-related TD questions. Subsequently, we conducted quantitative and qualitative analyses of the collected security-related TD questions, including sentiment analysis. Results: Our analysis revealed that 38% of the security questions on SO are security-related TD questions. The most recurrent tags among the security-related TD questions emerged as “security” and “encryption.” The latter typically have a neutral sentiment, are lengthier, and are posed by users with higher reputation scores. Conclusions: Our findings reveal that developers implicitly discuss TD, suggesting developers have a potential knowledge gap regarding the TD metaphor in the security domain. Moreover, we identified the most common security topics mentioned in TD-related posts, providing valuable insights for developers and researchers to assist developers in prioritizing security concerns in order to minimize TD and enhance software security.
Joshua Aldrich Edbert, Sahrima Jannat Oishwee, Shubhashis Karmakar, Zadia Codabux, Roberto Verdecchia
ESEM4
2023 Integrating Visual Aids to Enhance the Code Reviewer Selection Process
abstract
Modern Code Review (MCR) is an integral part of a software development strategy that accelerates product quality by identifying defects, code smells, and other harmful practices. However, assigning appropriate reviewers to evaluate changed code during the review process remains challenging. While automated tools for reviewer assignments have limited impact in practice, the process often relies on manual investigation of project histories to retrieve knowledge of team members and their activities. Therefore, in this study, we present an approach to automatically assemble developers’ information and visualize it meaningfully, which helps to choose appropriate reviewers. First, we propose three metrics that measure developers’ collaboration, reviewers’ expertise, and reviewers’ workload and visualize them through networks. Second, we perform a case study of three popular open-source projects, where we compute and visualize each developers’ information according to the proposed metrics. Finally, we conducted two online surveys to assess the developers’ perceptions of the proposed visual benefits. The results show that the proposed method can assist in identifying relevant reviewers and be immensely helpful to new developers. Additionally, survey respondents expressed reliance on the efficacy of the visual aids in workload balancing and reducing review time.
Md Shamimur Rahman, Debajyoti Mondal, Zadia Codabux, Chanchal Kumar Roy
ICSME3
2023 Rubbing salt in the wound? A large-scale investigation into the effects of refactoring on security
abstract
Abstract Software refactoring is a behavior-preserving activity to improve the source code quality without changing its external behavior. Unfortunately, it is often a manual and error-prone task that may induce regressions in the source code. Researchers have provided initial compelling evidence of the relation between refactoring and defects, yet little is known about how much it may impact software security. This paper bridges this knowledge gap by presenting a large-scale empirical investigation into the effects of refactoring on the security profile of applications. We conduct a three-level mining software repository study to establish the impact of 14 refactoring types on (i) security-related metrics, (ii) security technical debt, and (iii) the introduction of known vulnerabilities. The study covers 39 projects and a total amount of 7,708 refactoring commits. The key results show that refactoring has a limited connection to security. However, Inline Method and Extract Interface statistically contribute to improving some security aspects connected to encapsulating security-critical code components. Extract Superclass and Pull Up Attribute refactoring are commonly found in commits violating specific security best practices for writing secure code. Finally, Extract Superclass and Extract & Move Method refactoring tend to occur more often in commits contributing to the introduction of vulnerabilities. We conclude by distilling lessons learned and recommendations for researchers and practitioners.
Emanuele Iannone, Zadia Codabux, Valentina Lenarduzzi, Andrea De Lucia, Fabio Palomba
Empir. Softw. Eng.2
2023 Towards a taxonomy of Roxygen documentation in R packages
abstract
Abstract Software documentation is often neglected, impacting maintenance and reuse and leading to technical issues. In particular, when working with scientific software, such issues in the documentation pose a risk to producing reliable scientific results as they may cause improper or incorrect use of the software. R is a popular programming language for scientific software with a prolific package-based ecosystem, where users contribute packages (i.e., libraries). R packages are intended to be reused, and their users rely extensively on the available documentation. Thus, understanding what information developers provide in their packages’ documentation (generally, through a system known as Roxygen, based on Javadoc) is essential to contribute to it. This study mined 379 GitHub repositories of R packages and analysed a sample to develop a taxonomy of natural language descriptions used in Roxygen documentation. This was done through hybrid card sorting, which included two experienced R developers. The resulting taxonomy covers parameters, returns, and descriptions, providing a baseline for further studies. Our taxonomy is the first of its kind for R. Based on previous studies in pure object-oriented languages, our taxonomy could be extensible to other dynamically-typed languages used in scientific programming.
Melina C. Vidoni, Zadia Codabux
Empir. Softw. Eng.2
2022 An Experience Report on Technical Debt in Pull Requests: Challenges and Lessons Learned
abstract
Background: GitHub is a collaborative platform for global software development, where Pull Requests (PRs) are essential to bridge code changes with version control. However, developers often trade software quality for faster implementation, incurring Technical Debt (TD). When developers undertake reviewers’ roles and evaluate PRs, they can often detect TD instances, leading to either PR rejection or discussions. Aims: We investigated whether Pull Request Comments (PRCs) indicate TD by assessing three large-scale repositories: Spark, Kafka, and React. Method: We combined manual classification with automated detection using machine learning and deep learning models. Results: We classified two datasets and found that 37.7 and 38.7% of PRCs indicate TD, respectively. Our best model achieved F1 = 0.85 when classifying TD during the validation phase. Conclusions: We faced several challenges during this process, which may hint that TD in PRCs is discussed differently from other software artifacts (e.g., code comments, commits, issues, or discussion forums). Thus, we present challenges and lessons learned to assist researchers in pursuing this area of research.
Shubhashis Karmakar, Zadia Codabux, Melina C. Vidoni
ESEM2
2022 On the Benefits of the Accelerate Metrics: An Industrial Survey at Vendasta
abstract
The popularity of the Accelerate metrics is increasing in the industry. The Accelerate metrics are four key metrics to evaluate the software delivery performance: lead time for changes, deployment frequency, mean time to recover, change fail rate. However, their benefits in monitoring the development process performance of microservice-based systems have not been evaluated. In this study, we analyze the case of Vendasta, a Canadian company that migrated to microservices two years ago and adopted the Accelerate metrics to monitor their development process. Our goal is to understand whether these metrics are beneficial in the microservices context from the practitioners' point of view. Therefore, we surveyed employees from different teams and obtained 62 responses. Our results show that the Accelerate metrics provide a good overview of the process issues and are particularly helpful for a high-level representation of the process performances. Furthermore, the Accelerate metrics also enabled the teams to improve their productivity, significantly reducing service outages.
Francesco Lomio, Zadia Codabux, Dale Birtch, Dale Hopkins, Davide Taibi 0001
SANER2
2022 Toward Understanding the Impact of Refactoring on Program Comprehension
abstract
Software refactoring is the activity associated with developers changing the internal structure of source code without modifying its external behavior. The literature argues that refactoring might have beneficial and harmful implications for software maintainability, primarily when performed without the support of automated tools. This paper continues the narrative on the effects of refactoring by exploring the dimension of program comprehension, namely the property that describes how easy it is for developers to understand source code. We start our investigation by assessing the basic unit of program comprehension, namely program readability. Next, we set up a large-scale empirical investigation – conducted on 156 open-source projects – to quantify the impact of refactoring on program readability. First, we mine refactoring data and, for each commit involving a refactoring, we compute (i) the amount and type(s) of refactoring actions performed and (ii) eight state-of-the-art program comprehension metrics. Afterwards, we build statistical models relating the various refactoring operations to each of the readability metrics considered to quantify the extent to which each refactoring impacts the metrics in either a positive or negative manner. The key results are that refactoring has a notable impact on most of the readability metrics considered.
Giulia Sellitto, Emanuele Iannone, Zadia Codabux, Valentina Lenarduzzi, Andrea De Lucia, Fabio Palomba, Filomena Ferrucci
SANER3
2022 Common Programming Mistakes Leading to Information Disclosure: A Preliminary Study
abstract
It is vital to engineer robust and secure software. Many security strategies and techniques have been proposed. However, technological growth increases security concerns and demands persistent software security analysis. The objective of our study is to analyze vulnerable code components of real-world software code repositories and mine developers' frequent programming mistakes, resulting in information disclosure in the software. Finding common programming mistakes during the implementation phase is a primary step towards building secure software. We investigate the published vulnerabilities in two open-source applications: Apache Tomcat and Android. We focus on the information disclosure vulnerability reported as security advisories and analyze the code to extract or mine the causes of the vulnerability. We found that improper or lack of bound checking is the most frequent programming mistake that can potentially cause information leakage. Our findings can help create awareness among developers of the common programming mistakes that lead to disclosing sensitive information to avoid it, or if such mistakes are already present in the code, they can be handled during the implementation phase. Moreover, our results can be incorporated in tools such as static analyzers to help detect information disclosure instances more accurately prior to software delivery.
Gowri Pandian Sundarapandi, Raiyan Hossain, Chandana Jasrai, Kazi Zakia Sultana, Zadia Codabux
SANER5
2022 Self-admitted technical debt in R: detection and causes
abstract
Abstract Self-Admitted Technical Debt (SATD) is primarily studied in Object-Oriented (OO) languages and traditionally commercial software. However, scientific software coded in dynamically-typed languages such as R differs in paradigm, and the source code comments’ semantics are different (i.e., more aligned with algorithms and statistics when compared to traditional software). Additionally, many Software Engineering topics are understudied in scientific software development, with SATD detection remaining a challenge for this domain. This gap adds complexity since prior works determined SATD in scientific software does not adjust to many of the keywords identified for OO SATD, possibly hindering its automated detection. Therefore, we investigated how classification models (traditional machine learning, deep neural networks, and deep neural Pre-Trained Language Models (PTMs)) automatically detect SATD in R packages. This study aims to study the capabilities of these models to classify different TD types in this domain and manually analyze the causes of each in a representative sample. Our results show that PTMs (i.e., RoBERTa) outperform other models and work well when the number of comments labelled as a particular SATD type has low occurrences. We also found that some SATD types are more challenging to detect. We manually identified sixteen causes, including eight new causes detected by our study. The most common cause was failure to remember, in agreement with previous studies. These findings will help the R package authors automatically identify SATD in their source code and improve their code quality. In the future, checklists for R developers can also be developed by scientific communities such as rOpenSci to guarantee a higher quality of packages before submission.
Rishab Sharma, Ramin Shahbazi, Fatemeh Hendijani Fard, Zadia Codabux, Melina C. Vidoni
Autom. Softw. Eng.4
2022 Infinite technical debt
Melina C. Vidoni, Zadia Codabux, Fatemeh Hendijani Fard
J. Syst. Softw.2
2021 A Preliminary Study on Common Programming Mistakes that Lead to Buffer Overflow Vulnerability
abstract
When vulnerabilities are exploited, the impact can be insignificant or detrimental, depending on the attack’s nature. Research found that buffer overflow is one of the most widespread and frequently reported vulnerabilities that result in system crashes. This study investigates the frequent errors in the source code of production software that lead to buffer overflow such that its causes can be determined. The findings of the study can help guide developers to avoid these programming errors. Therefore, our study’s primary objective is to analyze vulnerable code components of software repositories and extract the developers’ frequent programming mistakes that have resulted in a buffer overflow attack. Sixteen vulnerable code components and relevant resolutions were selected from three popular and well-known systems: Android, Eclipse, and Red Hat, to be analyzed. The results show that lack of input sanitization, improper checking of array bounds and parameters, and the lack of value and range checks on variables are the most common programming issues that lead to a buffer overflow in these systems. We also found improper use of "If" and "While" loop conditions frequently contributed to the errors in bounds and variable checks.
Giovanni George, Jeremiah Kotey, Megan Ripley, Kazi Zakia Sultana, Zadia Codabux
COMPSAC5
2021 Technical Debt in the Peer-Review Documentation of R Packages: a rOpenSci Case Study
abstract
Context: Technical Debt (TD) is a metaphor used to describe code that is "not quite right." Although TD studies have gained momentum, TD has yet to be studied as thoroughly in non-Object-Oriented (OO) or scientific software such as R. R is a multi-paradigm programming language, whose popularity in data science and statistical applications has amplified in recent years. Due to R's inherent ability to expand through user-contributed packages, several community-led organizations were created to organize and peer-review packages in a concerted effort to increase their quality. Nonetheless, it is well-known that most R users do not have a technical programming background, being from multiple disciplines. Objective: The goal of this study is to investigate TD in the documentation of the peer-review of R packages led by rOpenSci. Method: We collected over 5,000 comments from 157 packages that had been reviewed and approved to be published at rOpenSci. We manually analyzed a sample dataset of these comments posted by package authors, editors of rOpenSci, and reviewers during the review process to investigate the types of TD present in these reviews. Results: The findings of our study include (i) a taxonomy of TD derived from our analysis of the peer-reviews (ii) documentation debt as being the most prevalent type of debt (iii) different user roles are concerned with different types of TD. For instance, reviewers tend to report some types of TD more than other roles, and the types of TD they report are different from those reported by the authors of a package. Conclusion: TD analysis in scientific software or peer-review is almost non-existent. Our study is a pioneer but within the context of R packages. However, our findings can serve as a starting point for replication studies, given our public datasets, to perform similar analyses in other scientific software or to investigate the rationale behind our findings.
Zadia Codabux, Melina C. Vidoni, Fatemeh Hendijani Fard
MSR1
2020 Examining the Relationship of Code and Architectural Smells with Software Vulnerabilities
abstract
Context: Security is vital to software developed for commercial or personal use. Although more organizations are realizing the importance of applying secure coding practices, in many of them, security concerns are not known or addressed until a security failure occurs. The root cause of security failures is vulnerable code. While metrics have been used to predict software vulnerabilities, we explore the relationship between code and architectural smells with security weaknesses. As smells are surface indicators of a deeper problem in software, determining the relationship between smells and software vulnerabilities can play a significant role in vulnerability prediction models. Objective: This study explores the relationship between smells and software vulnerabilities to identify the smells. Method: We extracted the class, method, file, and package level smells for three systems: Apache Tomcat, Apache CXF, and Android. We then compared their occurrences in the vulnerable classes which were reported to contain vulnerable code and in the neutral classes (non-vulnerable classes where no vulnerability had yet been reported). Results: We found that a vulnerable class is more likely to have certain smells compared to a non-vulnerable class. God Class, Complex Class, Large Class, Data Class, Feature Envy, Brain Class have a statistically significant relationship with software vulnerabilities. We found no significant relationship between architectural smells and software vulnerabilities. Conclusion: We can conclude that for all the systems examined, there is a statistically significant correlation between software vulnerabilities and some smells.
Kazi Zakia Sultana, Zadia Codabux, Byron J. Williams
APSEC2
2020 Profiling Developers Through the Lens of Technical Debt
abstract
Context: Technical Debt needs to be managed to avoid disastrous consequences, and investigating developers' habits concerning technical debt management is invaluable information in software development. Objective: This study aims to characterize how developers manage technical debt based on the code smells they induce and the refactorings they apply. Method: We mined a publicly-available Technical Debt dataset for Git commit information, code smells, coding violations, and refactoring activities for each developer of a selected project. Results: By combining this information, we profile developers to recognize prolific coders, highlight activities that discriminate among developer roles (reviewer, lead, architect), and estimate coding maturity and technical debt tolerance.
Zadia Codabux, Christopher Dutchyn
ESEM1
2017 The Relationship between Traceable Code Patterns and Code Smells
abstract
Context: It is important to maintain software quality as a software system evolves.Managing code smells in source code contributes towards quality software.While metrics have been used to pinpoint code smells in source code, we present an empirical study on the correlation of code smells with class-level (micro pattern) and methodlevel (nano-pattern) traceable patterns of code.Objective: This study explores the relationship between code smells and class-level and method-level structural code constructs.Method: We extracted micro patterns at the class level and nano-patterns at the method level from three versions of Apache Tomcat and PersonalBlog and Roller from Standford SecuriBench and compared their distributions in code smell versus non-code smell classes and methods.Result: We found that DataM anager, Record and Outline micro patterns are more frequent in classes having code smell compared to non-code smell classes in the applications we analyzed.localReader, localW riter, Switcher, and ArrReader nano-patterns are more frequent in code smell methods compared to the non-code smell methods.Conclusion: We conclude that code smells are correlated with both micro and nano-patterns.
Zadia Codabux, Kazi Zakia Sultana, Byron J. Williams
SEKE1
2017 The Relationship Between Code Smells and Traceable Patterns - Are They Measuring the Same Thing?
abstract
It is important to maintain software quality as a software system evolves. Managing code smells in source code contributes towards quality software. While metrics have been used to pinpoint code smells in source code, we present an empirical study on the correlation of code smells with class-level (micro pattern) and method-level (nano-pattern) traceable code patterns. This study explores the relationship between code smells and class-level and method-level structural code constructs. We extracted micro patterns at the class level and nano-patterns at the method level from three versions of Apache Tomcat, three versions of Apache CXF and two J2EE web applications namely PersonalBlog and Roller from Stanford SecuriBench and then compared their distributions in code smell versus noncode smell classes and methods. We found that Immutable and Sink micro patterns are more frequent in classes having code smells compared to the noncode smell classes in the applications we analyzed. On the other hand, LocalReader and LocalWriter nano-patterns are more frequent in code smell methods compared to the noncode smell methods. We conclude that code smells are correlated with both micro and nano-patterns.
Zadia Codabux, Kazi Zakia Sultana, Byron J. Williams
Int. J. Softw. Eng. Knowl. Eng.1
2017 An empirical assessment of technical debt practices in industry
abstract
Abstract Context Technical debt refers to the consequences of taking shortcuts when developing software. These consequences can impede the software growth and have financial implications. The software engineering research community needs to explore technical debt further from a practitioner standpoint. Objective This study gathers insights from practitioners on key components of technical debt such as its definition, characterization, consequences, benefits, and how it is communicated. Method We conducted semi‐structured interviews with a convenience sample of 17 practitioners and a survey of 67 participants. Results Despite the lack of consensus, we identified the most commonly accepted definition, method to measure technical debt (as person hours), and method to reduce debt (by allocating time in iterations to address the debt). Defects were also identified as type of debt and the cost of technical debt is more than the cost to make changes to source code. Three distinct company profiles emerged. Conclusion Despite increasing research on technical debt, the field lacks consensus on its many facets. One interesting outcome of this study is how to assess the risks of technical debt by evaluating liabilities beyond costs of directly handling debt.
Zadia Codabux, Byron J. Williams, Gary L. Bradshaw, Murray Cantor
J. Softw. Evol. Process.1