Minhaz Fahim Zibran

dblp:91/814 · also Minhaz F. Zibran · DBLP profile ↗
← Back
35ranked-venue papers
3as first author
22since 2021 · last 2026
0009-0004-5353-5030ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 32 · 3 first-author · 21 since 2021Databases, data management, data science and information retrieval · 17 · 14 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 The Quiet Contributions: Insights into AI-Generated Silent Pull Requests
abstract
We present the first empirical study of AI-generated pull requests that are ‘silent,’ meaning no comments or discussions accompany them. This absence of any comments or discussions associated with such silent AI pull requests (SPRs) poses a unique challenge in understanding the rationale for their acceptance or rejection. Hence, we quantitatively study 4,762 SPRs of five AI agents made to popular Python repositories drawn from the AIDev public dataset. We examine SPRs impact on code complexity, other quality issues, and security vulnerabilities, especially to determine whether these insights can hint at the rationale for acceptance or rejection of SPRs.
S. M. Mahedy Hasan, Md. Fazle Rabbi, Minhaz Fahim Zibran
MSR3
2026 When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests
abstract
Autonomous coding agents increasingly contribute to software development by submitting pull requests on GitHub; yet, little is known about how these contributions integrate into human-driven review workflows. We present a large empirical study of agent-authored pull requests using the public AIDev dataset, examining integration outcomes, resolution speed, and review-time collaboration signals. Using logistic regression with repository-clustered standard errors, we find that reviewer engagement has the strongest correlation with successful integration, whereas larger change sizes and coordination-disrupting actions, such as force pushes, are associated with a lower likelihood of merging. In contrast, iteration intensity alone provides limited explanatory power once collaboration signals are considered. A qualitative analysis further shows that successful integration occurs when agents engage in actionable review loops that converge toward reviewer expectations. Overall, our results highlight that the effective integration of agent-authored pull requests depends not only on code quality but also on alignment with established review and coordination practices.
Costain Nachuma, Minhaz Fahim Zibran
MSR2
2026 A Task-Level Evaluation of AI Agents in Open-Source Projects
Shojibur Rahman, Md. Fazle Rabbi, Minhaz Fahim Zibran
MSR3
2026 The SBOM Gap: Adoption and Compliance in Open Source Software
Md. Fazle Rabbi, Asif Kamal Turzo, Arifa I. Champa, Minhaz Fahim Zibran
SANER4
2025 Insights into Vulnerability Trends in Maven Artifacts: Recurrence, Popularity, and User Behavior
abstract
Vulnerabilities in open-source software, particularly in ecosystems like Maven Central, propagate risks across projects. This paper examines vulnerability trends in Maven artifacts, focusing on recurrence patterns, user behavior after disclosures, and the link between popularity and exposure. Analyzing 24 vulnerable artifacts and $2,900+$ releases, we find recurring risks in previously vulnerable artifacts, significant intravs. extra-organizational differences in user behavior, and minimal correlation between popularity and vulnerability exposure. These results underscore the need for proactive security, effective disclosures, and better dependency management to strengthen ecosystem resilience.
Courtney Bodily, Eric Hill, Andreas Kramer, Leslie Kerby, Minhaz Fahim Zibran
MSR5
2025 Insights into Dependency Maintenance Trends in the Maven Ecosystem
abstract
As modern software development increasingly relies on reusable libraries and components, managing dependencies has become critical for ensuring software stability and security. However, challenges such as outdated dependencies, missed releases, and the complexity of interdependent libraries can significantly impact project maintenance. In this paper, we present a quantitative analysis of the Neo 4 j dataset using the Goblin framework to uncover patterns of freshness in projects with different numbers of dependencies. Our analysis reveals that releases with fewer dependencies have a higher number of missed releases. Additionally, our study shows that the dependencies in the latest releases have positive freshness scores, indicating better software management efficacy. These results can encourage better management practices and contribute to the overall health of software ecosystems.
Barisha Chowdhury, Md. Fazle Rabbi, S. M. Mahedy Hasan, Minhaz Fahim Zibran
MSR4
2025 Analyzing Dependency Clusters and Security Risks in the Maven Central Repository
abstract
We present a cluster analysis of the Maven Central Repository’s dependency structure to identify and assess vulnerability risks using the Goblin framework. Through analysis of over 15 million artifacts using the Leiden community detection algorithm, we identified approximately 67 thousand distinct clusters with a high modularity score. Our risk assessment framework combines CVE metrics, freshness scores, and inter-cluster connectivity patterns to evaluate cluster risk levels and potential vulnerability propagation paths. The analysis reveals that while individual clusters typically show low to moderate risk scores, the repository’s highly connected structure creates critical paths for vulnerability propagation through hub clusters, some containing over 1.5 million nodes. We provide recommendations for dependency risk monitoring, including tracking of bridge nodes and prioritizing high-connectivity clusters. Our systematic approach provides a framework to identify systemic dependency risks across the repository through targeted inspections at critical points in the dependency network.
George Lake, Minhaz Fahim Zibran
MSR2
2025 Decoding Dependency Risks: A Quantitative Study of Vulnerabilities in the Maven Ecosystem
abstract
This study investigates vulnerabilities within the Maven ecosystem by analyzing a comprehensive dataset of $14,459,139$ releases. Our analysis reveals the most critical weaknesses that pose significant threats to developers and their projects as they look to streamline their development tasks through code reuse. We show risky weaknesses, those unique to Maven, and emphasize those becoming increasingly dangerous over time. Furthermore, we reveal how vulnerabilities subtly propagate, impacting $31.39 \%$ of the 635,003 latest releases through direct dependencies and $62.89 \%$ through transitive dependencies. Our findings suggest that improper handling of input and mismanagement of resources pose the most risk. Additionally, Insufficient session-ID length in J2EE configuration and no throttling while allocating resources uniquely threaten the Maven ecosystem. We also find that weaknesses related to improper authentication and managing sensitive data without encryption have quickly gained prominence in recent years. These findings emphasize the need for proactive strategies to mitigate security risks in the Maven ecosystem.
Costain Nachuma, Md Mosharaf Hossan, Asif Kamal Turzo, Minhaz Fahim Zibran
MSR4
2025 Chasing the Clock: How Fast Are Vulnerabilities Fixed in the Maven Ecosystem?
abstract
This study investigates the software vulnerability resolution time in the Maven ecosystem, focusing on the influence of CVE severity, library popularity as measured by the number of dependents, and version release frequency. The results suggest that critical vulnerabilities are addressed slightly faster compared to lower-severity ones. Library popularity shows a positive impact on resolution times, while frequent version updates are associated with faster vulnerability fixes. These statistically significant findings are based on a thorough evaluation of over 14 million versions from 658,078 libraries using the dependency graph database of Goblin framework. These results emphasize the need for proactive maintenance strategies to improve vulnerability management in open-source ecosystems.
Md. Fazle Rabbi, Arifa I. Champa, Rajshakhar Paul, Minhaz Fahim Zibran
MSR4
2025 Understanding Software Vulnerabilities in the Maven Ecosystem: Patterns, Timelines, and Risks
abstract
Vulnerabilities in software libraries and reusable components cause major security challenges, particularly in dependency-heavy ecosystems such as Maven. This paper presents a large-scale analysis of vulnerabilities in the Maven ecosystem using the Goblin framework. Our analysis focuses on the aspects and implications of vulnerability types, documentation delays, and resolution timelines. We identify 77,393 vulnerable releases with 226 unique CWEs. On average, vulnerabilities take nearly half a decade to be documented and 4.4 years to be resolved, with some remaining unresolved for even over a decade. The delays in documenting and fixing vulnerabilities incur security risks for the library users emphasizing the need for more careful and efficient vulnerability management in the Maven ecosystem.
Md. Fazle Rabbi, Rajshakhar Paul, Arifa I. Champa, Minhaz Fahim Zibran
MSR4
2025 Faster Releases, Fewer Risks: A Study on Maven Artifact Vulnerabilities and Lifecycle Management
abstract
In modern software ecosystems, dependency management plays a critical role in ensuring secure and maintainable applications. However, understanding the relationship between release practices and their impact on vulnerabilities and update cycles remains a challenge. In this study, we analyze the release histories of $\mathbf{1 0, 0 0 0}$ Maven artifacts, covering over $\mathbf{2 0 3, 0 0 0}$ releases and 1.7 million dependencies. We evaluate how release speed affects software security and lifecycle. Our results show an inverse relationship between release speed and dependency outdatedness. Artifacts with more frequent releases maintain significantly shorter outdated times. We also find that faster release cycles are linked to fewer CVEs in dependency chains, indicating a strong negative correlation. These findings emphasize the importance of accelerated release strategies in reducing security risks and ensuring timely updates. Our research provides valuable insights for software developers, maintainers, and ecosystem managers.
Md Shafiullah Shafin, Md. Fazle Rabbi, S. M. Mahedy Hasan, Minhaz Fahim Zibran
MSR4
2025 A novel ensemble approach for crop disease detection by leveraging customized EfficientNets and interpretability
Nahrin Jannat, S. M. Mahedy Hasan, Minhaz Fahim Zibran
Pattern Recognit. Lett.3
2024 ChatGPT in Action: Analyzing Its Use in Software Development
abstract
The emergence of AI tools such as ChatGPT is being used to assist with software development, but little is known of how developers utilize these tools as well as the capabilities of these tools in software engineering tasks. Using the DevGPT dataset, we conduct quantitative analyses of the tasks developers seek assistance from ChatGPT and how effectively ChatGPT addresses them. We also examine the impact of initial prompt quality on conversation length. The findings reveal where ChatGPT is most and least suited to assist in the identified 12 software development tasks. The insights from this research would guide the software developers, researchers, and AI tool providers in optimizing these tools for more effective programming aid.
Arifa I. Champa, Md. Fazle Rabbi, Costain Nachuma, Minhaz Fahim Zibran
MSR4
2024 A Four-Dimension Gold Standard Dataset for Opinion Mining in Software Engineering
abstract
We present the first four-dimension gold standard dataset to advance opinion mining focused on the software engineering domain. Through a well-defined sampling and annotation strategy leveraging multiple coders, we construct a corpus of 2,000 Stack Overflow posts labeled with four dimensions/tuples, including sentiments, polar facts, aspects, and named entities. This multidimensional ground truth dataset opens up new research opportunities for opinion mining in domain-adapted NLP tools for software engineering by capturing existing relationships between extracted elements at a more granular level. It also facilitates investigating the effects of sentiments in the developers' social forums.
Md. Rakibul Islam 0002, Md. Fazle Rabbi, Youngeun Jo, Arifa I. Champa, Ethan Young, Camden Wilson, Gavin Scott, Minhaz Fahim Zibran
MSR8
2024 AI Writes, We Analyze: The ChatGPT Python Code Saga
abstract
In this study, we quantitatively analyze 1,756 AI-written Python code snippets in the DevGPT dataset and evaluate them for quality and security issues. We systematically distinguish the code snippets as either generated by ChatGPT from scratch (ChatGPT-generated) or modified user-provided code (ChatGPT-modified). The results reveal that ChatGPT-modified code more frequently displays quality issues compared to ChatGPT-generated code. The findings provide insights into the inherent limitations of AI-written code and emphasize the need for scrutiny before integrating such pieces of code into software systems.
Md. Fazle Rabbi, Arifa I. Champa, Minhaz Fahim Zibran, Md. Rakibul Islam 0002
MSR3
2023 Insights into Female Contributions in Open-Source Projects
abstract
This paper presents a large quantitative study of the contributions of females compared to males in open-source projects. Female participation is found substantially low and females are found more engaged in non-coding work compared to men. The findings are statistically significant and are derived from an in-depth analysis of over 10 thousand developers’ contributions to more than 81 million different projects in the World of Code (WoC) infrastructure. The insights from this study are useful in addressing gender disparity in the field.
Arifa I. Champa, Md. Fazle Rabbi, Minhaz Fahim Zibran, Md. Rakibul Islam 0002
MSR3
2023 Are We Aware? An Empirical Study on the Privacy and Security Awareness of Smartphone Sensors
abstract
Smartphones are equipped with a wide variety of sensors, which can pose significant security and privacy risks if not properly protected. To assess the privacy and security risks of smartphone sensors, we first systematically reviewed 55 research papers. Driven by the findings of the systematic review, we carried out a follow-up questionnaire-based survey on 23 human end-users. The results reflect that the participants have a varying level of familiarity with smartphone sensors, and there is a noticeable dearth of awareness about the potential threats and preventive measures associated with these sensors. The findings from this study will inform the development of effective solutions for addressing security and privacy in mobile devices and beyond.
Arifa I. Champa, Md. Fazle Rabbi, Farjana Z. Eishita, Minhaz Fahim Zibran
SERA4
2023 Phishy? Detecting Phishing Emails Using ML and NLP
abstract
Phishing emails, a type of cyberattack using fake emails, are difficult to recognize due to sophisticated techniques employed by attackers. In this paper, we use a natural language processing (NLP) and machine learning (ML) based approach for detecting phishing emails. We compare the efficacy of six different ML algorithms for the purpose. An empirical evaluation on two public datasets demonstrates that our approach detects phishing emails with high accuracy, precision, and recall. The findings from this work are useful in devising more efficient techniques for recognizing and preventing phishing attacks.
Md. Fazle Rabbi, Arifa I. Champa, Minhaz Fahim Zibran
SERA3
2022 Security Versus Performance Bugs: How Bugs are Handled in the Chromium Project
abstract
Bug fixing is a very important activity of software maintenance. Given the recent highlight on security and privacy, one may expect that the software vendors would give security bugs a higher priority in their bug fixing process. In this paper, we present an exploratory study of different categories (i.e., security, performance, and other) of bugs in the maintenance of the Chromium browser. In particular, we study the phenomena such as how much time is spent in bug triage, how fast different types of bugs are fixed, variations of developers' experiences who fix those bugs, and show often those fixed bugs are reopened. We find that the performance bugs are triaged and fixed faster. Security bugs, for fixing, are assigned to more experienced developers. All categories of bugs are almost equally reopened once closed.
Amrit Rajbhandari, Minhaz Fahim Zibran, Farjana Z. Eishita
SERA2
2022 Finding the Middle Ground: Measuring Passwords for Security and Memorability
abstract
Passwords are a ubiquitous element of our digital age, and the need for secure passwords is indispensable. However, traditionally secure passwords tends to be very difficult to remember, leading users to frustration in having to remake them or even abandoning secure ones for the sake of memorability. On the other hand, passwords that are considered memorable have a tendency to be less secure. Despite many studies on passwords, the process of the users' perceiving password memorability is still abstract. This security and memorability trade-off brings out the need to find a middle ground leaving the research question on the ground whether finding passwords that are both memorable and secure is a possibility or not. In this paper, we address this very question by conducting a survey and a user-study of memorability for certain categories of popular password styles. Next, we take the most memorable of these passwords and determine their security via a bits of entropy calculation commonly used to determine a password's strength against a traditional brute force attack. Finally, we find the best performers of both memorability and security to find the middle ground of password security and memorability. Our findings present a collection of both memorable and secure styles of passwords including an infrequent but effective tactic for password remembrance.
Joshua J. Rodriguez, Minhaz Fahim Zibran, Farjana Z. Eishita
SERA2
2021 Choosing the Weapon: A Comparative Study of Security Analyzers for Android Applications
abstract
This study compares security focused static code analyzers for Android applications. Android operated hand-held devices (e.g., smart phones, tablets) are used in the modern computing world for nearly every need. Banking, email, health care, and other sensitive dealings are completed through the Android applications. Hence, Android application security must be held to the same level of scrutiny as traditional application security. This study compares two open-source security analyzers, MobSF and MARA, against two benchmark datasets and 20 live Android applications. We highlight the strengths and weaknesses of each analyzer and reveal security vulnerabilities found in the Android applications.
Ryan B. Joseph, Minhaz Fahim Zibran, Farjana Z. Eishita
SERA2
2021 Plugins to Detect Vulnerable Plugins: An Empirical Assessment of the Security Scanner Plugins for WordPress
abstract
WordPress, possibly world's the most popular Content Management System (CMS), which supports around 455 million websites and claims 60.3% of all content management systems in use. The WordPress core is known to be relatively secure, but its plugin ecosystem is not. 92% of vulnerabilities found in WordPress powered websites are attributed to third-party plugins that those websites depend on.This paper presents an empirical study, where we examine the efficacy of 11 WordPress security scanner plugins in the detection of known vulnerabilities in another set of 51 insecure plugins. The results are mixed, with some security scanner plugins failing entirely and even the most effective plugins failing to identify significant vulnerabilities. The findings are derived based on both a quantitative analysis and a deeper qualitative analysis.
Daniel T. Murphy, Minhaz Fahim Zibran, Farjana Z. Eishita
SERA2
2019 Detecting Web Spam in Webgraphs with Predictive Model Analysis
abstract
Web spam is a serious threat for both end-users and search engines (w.r.t., query cost). Webgraphs can be exploited in detecting spam. In the past, several graph mining techniques were applied to measure metrics for pages and hyperlinks. In this paper, we justify the importance of webgraph to distinguish spam websites from non-spam ones based on several graph metrics computed for a labelled dataset (WEBSPAM-UK2007) and justify our model by testing on uk-2014 dataset, the most recently available dataset on the same (uk) domain. WEBSPAM-UK2007 dataset includes 0.1 million different hosts and four kinds of feature sets: Obvious, Link, Transformed Link and Content. We use five prominent machine learning (ML) techniques (i.e., Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Logistic Regression, Naïve Bayes and Random Forest) to build a ML-based classifier. To evaluate the performance of our classifier, we compute accuracy and F-1 score and perform 10-fold cross validation. We also compare graph based features with content based textual features and find that graph properties are similar or better than text properties. We achieve above 99% training accuracy for most of our machine learning models. We test our model with uk-2014 dataset with 4.7 million hosts for the graph-based feature sets and achieve accuracy in between 90-94% for most of the models. To the best of our knowledge, prior works on web spam detection with WEBSPAM-UK2007 dataset did not use different test dataset for their models. Our model classifier is capable of detecting web spam for any input webgraph based on its graph metrics features.
Naw Safrin Sattar, S. M. Arifuzzaman, Minhaz Fahim Zibran, Md Mohiuddin Sakib
IEEE BigData3
2018 A comparison of software engineering domain specific sentiment analysis tools
abstract
Sentiment Analysis (SA) in software engineering (SE) text has drawn immense interests recently. The poor performance of general-purpose SA tools, when operated on SE text, has led to recent emergence of domain-specific SA tools especially designed for SE text. However, these domain-specific tools were tested on single dataset and their performances were compared mainly against general-purpose tools. Thus, two things remain unclear: (i) how well these tools really work on other datasets, and (ii) which tool to choose in which context. To address these concerns, we operate three recent domain-specific SA tools on three separate datasets. Using standard accuracy measurement metrics, we compute and compare their accuracies in the detection of sentiments in SE text.
Md. Rakibul Islam 0002, Minhaz Fahim Zibran
SANER2
2018 SentiStrength-SE: Exploiting domain specificity for improved sentiment analysis in software engineering text
Md. Rakibul Islam 0002, Minhaz Fahim Zibran
J. Syst. Softw.2
2017 A Comparison of Dictionary Building Methods for Sentiment Analysis in Software Engineering Text
abstract
Sentiment Analysis (SA) in Software Engineering (SE) texts suffers from low accuracies primarily due to the lack of an effective dictionary. The use of a domain-specific dictionary can improve the accuracy of SA in a particular domain. Building a domain dictionary is not a trivial task. The performance of lexical SA also varies based on the method applied to develop the dictionary. This paper includes a quantitative comparison of four dictionaries representing distinct dictionary building methods to identify which methods have higher/lower potential to perform well in constructing a domain dictionary for SA in SE texts.
Md. Rakibul Islam 0002, Minhaz Fahim Zibran
ESEM2
2017 Security Vulnerabilities in Categories of Clones and Non-Cloned Code: An Empirical Study
abstract
Background: Software security has drawn immense importance in the recent years. While efforts are expected in minimizing security vulnerabilities in source code, the developers' practice of code cloning often causes multiplication of such vulnerabilities and program faults. Although previous studies examined the bug-proneness, stability, and changeability of clones against non-cloned code, the security aspects remained ignored. Aims: The objective of this work is to explore and understand the security vulnerabilities and their severity in different types of clones compared to non-clone code. Method: Using a state-of-the-art clone detector and two reputed security vulnerability detection tools, we detect clones and vulnerabilities in 8.7 million lines of code over 34 software systems. We perform a comparative study of the vulnerabilities identified in different types of clones and non-cloned code. The results are derived based on quantitative analyses with statistical significance. Results: Our study reveals that the security vulnerabilities found in code clones have higher severity of security risks compared to those in non-cloned code. However, the proportion (i.e., density) of vulnerabilities in clones and non-cloned code does not have any significant difference. Conclusion: The findings from this work add to our understanding of the characteristics and impacts of clones, which will be useful in clone-aware software development with improved software security.
Md. Rakibul Islam 0002, Minhaz Fahim Zibran, Aayush Nagpal
ESEM2
2017 Leveraging automated sentiment analysis in software engineering
abstract
Automated sentiment analysis in software engineering textual artifacts has long been suffering from inaccuracies in those few tools available for the purpose. We conduct an in-depth qualitative study to identify the difficulties responsible for such low accuracy. Majority of the exposed difficulties are then carefully addressed in developing SentiStrength-SE, a tool for improved sentiment analysis especially designed for application in the software engineering domain. Using a benchmark dataset consisting of 5,600 manually annotated JIRA issue comments, we carry out both quantitative and qualitative evaluations of our tool. SentiStrength-SE achieves 73.85% precision and 85% recall, which are significantly higher than a state-of-the-art sentiment analysis tool we compare with.
Md. Rakibul Islam 0002, Minhaz Fahim Zibran
MSR2
2017 Insights into continuous integration build failures
abstract
Continuous integration is prevalently used in modern software engineering to build software systems automatically. Broken builds hinder developers' work and delay project progress. We must identify the factors causing build failures. This paper presents a large empirical study to identify the factors such as, complexity of a task, build strategy and contribution models (i.e., push and pull request), and projects level attributes (i.e., sizes of projects and teams), which potentially have impacts on the build results. We have studied 3.6 million builds over 1,090 open-source projects. The derived results add to our understanding of the role of those factors on build results, which can be used in minimizing build failures.
Md. Rakibul Islam 0002, Minhaz Fahim Zibran
MSR2
2016 Towards understanding and exploiting developers' emotional variations in software engineering
abstract
Software development is highly dependent on human efforts and collaborations, which are immensely affected by emotions. This paper presents a quantitative empirical study of the emotional variations in different types of development activities (e.g., bug-fixing tasks) and development periods (i.e., days and times), in addition to in-depth investigation of emotions' impacts on software artifacts (i.e., commit messages) and exploration of scopes for exploiting emotional variations in software engineering activities. We study emotions in more than 490 thousand commit comments across 50 open-source projects. The findings add to our understanding of the role of emotions in software development, and expose scopes for exploitation of emotional awareness in improved task assignments and collaborations.
Md. Rakibul Islam 0002, Minhaz Fahim Zibran
SERA2
2011 Analyzing and Forecasting Near-Miss Clones in Evolving Software: An Empirical Study
abstract
Effort for development and maintenance of complex large software is believed to have dependency on the amount of duplicated code fragments (code clones) present in code-bases. For example, clones need to be carefully and consistently maintained and/or refactored for preventing accidental error propagation. Thus it is important to understand the proportion and evolution of clones in evolving software systems for cost estimation or the like. This paper presents a study on the evolution of near-miss clones at release level in medium to large open source software systems of different types (operating systems, database systems, editors, etc.) written in three different programming languages namely C, C#, and Java. Using a hybrid clone detector, NiCad, we detected both exact and near-miss clones at different levels of similarity. Applying statistical methods we investigated, from different dimensions, the evolution of both exact and near-miss clones, and also forecasted the amount of clones in future releases of the software systems. Our study offers significant insights into the existence and evolution of code clones and their relationships with programming language or paradigm and program size.
Minhaz Fahim Zibran, Ripon K. Saha, Muhammad Asaduzzaman, Chanchal Kumar Roy
ICECCS1
2011 Conflict-Aware Optimal Scheduling of Code Clone Refactoring: A Constraint Programming Approach
abstract
Duplicated code, also known as code clones, are one of the malicious `code smells' that often need to be removed through refactoring for enhancing maintainability. Among all the potential refactoring opportunities, the choice and order of a set of refactoring activities may have distinguishable effect on the design/code quality. Moreover, there may be dependencies and conflicts among those refactorings. The organization may also impose priorities on certain refactoring activities. Addressing all these conflicts, priorities, and dependencies, manual formulation of an optimal refactoring schedule is very expensive, if not impossible. Therefore, an automated refactoring scheduler is necessary, which will maximize benefit and minimize refactoring effort. In this paper, we present a refactoring effort model, and propose a constraint programming approach for conflict-aware optimal scheduling of code clone refactoring.
Minhaz Fahim Zibran, Chanchal Kumar Roy
ICPC1
2011 A Constraint Programming Approach to Conflict-Aware Optimal Scheduling of Prioritized Code Clone Refactoring
abstract
Duplicated code, also known as code clones, are one of the malicious ‘code smells' that often need to be removed through refactoring for enhancing maintainability. Among all the potential refactoring opportunities, the choice and order of a set of refactoring activities may have distinguishable effect on the design/code quality. Moreover, there may be dependencies and conflicts among those refactorings. The organization may also impose priorities on certain refactoring activities. Addressing all these conflicts, priorities, and dependencies, manual formulation of an optimal refactoring schedule is very expensive, if not impossible. Therefore, an automated refactoring scheduler is necessary, which will maximize benefit and minimize refactoring effort. In this paper, we present a refactoring effort model, and propose a constraint programming approach for conflict-aware optimal scheduling of code clone refactoring.
Minhaz Fahim Zibran, Chanchal Kumar Roy
SCAM1
2010 Evaluating Code Clone Genealogies at Release Level: An Empirical Study
abstract
Code clone genealogies show how clone groups evolve with the evolution of the associated software system, and thus could provide important insights on the maintenance implications of clones. In this paper, we provide an in-depth empirical study for evaluating clone genealogies in evolving open source systems at the release level. We develop a clone genealogy extractor, examine 17 open source C, Java, C++ and C# systems of diverse varieties and study different dimensions of how clone groups evolve with the evolution of the software systems. Our study shows that majority of the clone groups of the clone genealogies either propagate without any syntactic changes or change consistently in the subsequent releases, and that many of the genealogies remain alive during the evolution. These findings seem to be consistent with the findings of a previous study that clones may not be as detrimental in software maintenance as believed to be (at least by many of us), and that instead of aggressively refactoring clones, we should possibly focus on tracking and managing clones during the evolution of software systems.
Ripon K. Saha, Muhammad Asaduzzaman, Minhaz Fahim Zibran, Chanchal Kumar Roy, Kevin A. Schneider
SCAM3
2007 A multi-phase approach to the university course timetabling problem
Shahadat Hossain, Minhaz Fahim Zibran
CTW2