VLDB 2026 Research / reviewers in the wild / expert
Kevin A. Schneider
dblp:79/3635
· DBLP profile ↗
96ranked-venue papers
0as first author
22since 2021 · last 2026
0000-0003-1113-1754ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 89 · 20 since 2021Databases, data management, data science and information retrieval · 9Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRECT: A search-based developer recommendation approach for software crowdsourcing platforms
Nuri Almarimi, Ali Ouni 0001, Banani Roy, Moataz Chouchen, Chanchal Kumar Roy, Kevin A. Schneider |
Empir. Softw. Eng. | 6 |
| 2025 | Are Classical Clone Detectors Good Enough for the AI Era?abstractThe increasing adoption of AI-generated code has reshaped modern software development, introducing syntactic and semantic variations in cloned code. Unlike traditional human-written clones, AI-generated clones exhibit systematic syntactic patterns and semantic differences learned from largescale training data. This shift presents new challenges for classical code clone detection (CCD) tools, which have historically been validated primarily on human-authored codebases and optimized to detect syntactic (Type 1-3) and limited semantic clones. Given that AI-generated code can produce both syntactic and complex semantic clones, it is essential to evaluate the effectiveness of classical CCD tools within this new paradigm. In this paper, we systematically evaluate nine widely used CCD tools using GPTCloneBench, a benchmark containing GPT-3-generated clones. To contextualize and validate our results, we further test these detectors on established human-authored benchmarks, BigCloneBench and SemanticCloneBench, to measure differences in performance between traditional and AI-generated clones. Our analysis demonstrates that classical CCD tools, particularly those enhanced by effective normalization techniques, retain considerable effectiveness against AI-generated clones, while some exhibit notable performance variation compared to traditional benchmarks. This paper contributes by (1) evaluating classical CCD tools against AI-generated clones, providing critical insights into their current strengths and limitations; (2) highlighting the role of normalization techniques in improving detection accuracy; and (3) delivering detailed scalability and execution-time analyses to support practical CCD tool selection. The research underscores the continued relevance of classical CCD tools and suggests adopting a hybrid approach that combines both classical and AI-based methods to improve clone detection in the modern era. Ajmain Inqiad Alam, Palash Ranjan Roy, Farouq Al-Omari, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
ICSME | 6 |
| 2025 | Towards Enhancing IR-Based Bug Localization Leveraging Texts and Multimedia from Bug ReportsabstractSoftware bug reports often miss critical information, delaying their resolution. The last decade has seen a growing trend of combining textual and multimedia information from software artifacts (e.g., bug reports, and programming questions) to support various software engineering tasks (e.g., duplicate bug report detection, and bug reproduction). However, none of the studies performs a fine-grained analysis of the multimedia information attached to bug reports. Hence, it is not clear what their attached images or videos contain or whether they could help identify software bugs or errors automatically. In this paper, we conduct a preliminary study that investigates the presence or prevalence of key elements in 1,469 images attached to 1,000 bug reports and demonstrate their potential to support IR-based bug localization. We have several interesting findings. First, our analysis suggests that the attached images to visual bug reports contain a mix of UI components, programming components, and regular text. Second, our analysis using an LLM (e.g., GPT4o) suggests that it can extract the key elements from the attached images effectively, posing a suitable alternative to human annotators. Finally, our experiments suggest that the multimedia information extracted from the attached images can enhance the performance of a traditional technique for bug localization by improving 34.06 % of its search queries. Shamima Yeasmin, Chanchal Kumar Roy, Kevin A. Schneider, Mohammad Masudur Rahman 0001, Kartik Mittal, Ryder Hardy |
ICPC | 3 |
| 2025 | Quantum software engineering and potential of quantum computing in software engineering research: a review
Ashis Kumar Mandal, Md. Nadim, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
Autom. Softw. Eng. | 5 |
| 2025 | FSECAM: A contextual thematic approach for linking feature to multi-level software architectural components
Amit Kumar Mondal, Mainul Hossain, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
J. Syst. Softw. | 5 |
| 2024 | Take Loads Off Your Developers: Automated User Story Generation using Large Language ModelabstractSoftware Maintenance and Evolution (SME) is moving fast with the assistance of artificial intelligence (AI), especially Large Language Models (LLM). Researchers have already started automating various activities of the SME workflow. Un-derstanding the requirements for maintenance and development work i.e. Requirements Engineering (RE) is a crucial phase that kicks off the SME workflow through multiple discussions on a proposed scope of work documented in different forms. The RE phase ends with a list of user stories for each unit task and usually created and tracked on a project management tool such as GitHub, Jira, AzurDev, etc. In this research, we collaborated with Bell Mobility to develop a tool “Geneus” (Generate UserSory) using GPT-4-turbo to automatically create user stories from software requirements documents. Requirements documents are usually long and contain complex information. Since LLMs typically suffer from hallucination when the input is too complex, this paper proposes a new prompting strategy, “Refine and Thought” (RaT), to mitigate that issue and improve the performance of the LLM in prompts with large and noisy contexts. Along with manual evaluation using RUST (Readability, Understandability, Specificity, Technical-aspects) survey questionnaire, automatic evaluation with BERTScore, and AlignScore evaluation metrics are used to evaluate the results of the “Geneus” tool. Results show that our method with RaT performs consistently better in most of the cases of interactions compared to the single-shot baseline method. However, the BERTScore and AlignScore test results are not consistent. In the median case, Geneus performs significantly better in all three interactions (requirements specifi-cation, user story details, and test case specifications) according to AlignScorebut it shows slightly low performance in requirements specifications according to BERTScore. Distilling RE documents requires significant time & effort from the senior members of the team through multiple meetings with stakeholders. We believe automating this process will certainly reduce additional loads off the software engineers and increase the ultimate productivity allowing them to utilize their time on other prioritized tasks. Tajmilur Rahman, Yuecai Zhu, Lamyea Maha, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
ICSME | 6 |
| 2024 | Automated Derivation of UML Sequence Diagrams from User Stories: Unleashing the Power of Generative AI vs. a Rule-Based ApproachabstractUser stories are informal, non-technical descriptions of features from a user's perspective that guide collaboration and iterative development in Agile projects. However, ambiguities in user stories can lead to miscommunication among stakeholders. Design models, such as UML sequence diagrams, are essential for enhancing communication, clarifying system behavior, and improving the development process. This paper presents an automated approach for generating behavioral models specifically sequence diagrams from natural language requirements expressed as user stories. We also investigate the effectiveness of a Large Language Model (LLM) in using generative AI for this task. By applying our approach and ChatGPT to two benchmark datasets with the same set of user stories, we generated corresponding sequence diagrams for comparison. Expert evaluations in Software Engineering reveal that our approach effectively produces relevant, simplified diagrams for straightforward user stories, whereas the LLM tends to create more complex diagrams that sometimes go beyond the simplicity of the original user stories. Munima Jahan, Mohammad Mahdi Hassan, Reza Golpayegani, Golshid Ranjbaran, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
MODELS | 7 |
| 2024 | ReBack: recommending backports in social coding environments
Debasish Chakroborti, Kevin A. Schneider, Chanchal Kumar Roy |
Autom. Softw. Eng. | 2 |
| 2023 | GPTCloneBench: A comprehensive benchmark of semantic clones and cross-language clones using GPT-3 model and SemanticCloneBenchabstractWith the emergence of Machine Learning, there has been a surge in leveraging its capabilities for problem-solving across various domains. In the code clone realm, the identification of type-4 or semantic clones has emerged as a crucial yet challenging task. Researchers aim to utilize Machine Learning to tackle this challenge, often relying on the Big-CloneBench dataset. However, it’s worth noting that BigCloneBench, originally not designed for semantic clone detection, presents several limitations that hinder its suitability as a comprehensive training dataset for this specific purpose. Furthermore, CLCDSA dataset suffers from a lack of reusable examples aligning with real-world software systems, rendering it inadequate for cross-language clone detection approaches. In this work, we present a comprehensive semantic clone and cross-language clone benchmark, GPTCloneBench1by exploiting SemanticCloneBench and OpenAI’s GPT-3 model. In particular, using code fragments from SemanticCloneBench as sample inputs along with appropriate prompt engineering for GPT-3 model, we generate semantic and cross-language clones for these specific fragments and then conduct a combination of extensive manual analysis, tool-assisted filtering, functionality testing and automated validation in building the benchmark. From 79,928 clone pairs of GPT-3 output, we created a benchmark with 37,149 true semantic clone pairs, 19,288 false semantic pairs(Type-1/Type-2), and 20,770 cross-language clones across four languages (Java, C, C#, and Python). Our benchmark is 15-fold larger than SemanticCloneBench, has more functional code examples for software systems and programming language support than CLCDSA, and overcomes BigCloneBench’s qualities, quantification, and language variety limitations. GPTCloneBench can be found here1. Ajmain Inqiad Alam, Palash Ranjan Roy, Farouq Al-Omari, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
ICSME | 6 |
| 2023 | Release conventions of open-source software: An exploratory studyabstractAbstract Software engineering (SE) methodologies are widely used in both academia and industry to manage the software development life cycle. A number of studies of SE methodologies involve interviewing stakeholders to explore the real‐world practice. Although these interview‐based studies provide us with a user's perspective of an organization's practice, they do not describe the concrete summary of releases in open‐source social coding platforms. In particular, no existing studies investigated how releases are evolved in open‐source coding platforms, which assist release planners to a large extent. This study explores software development patterns followed in open‐source projects to see the overall management's reflection on software release decisions rather than concentrating on a particular methodology. Our experiments on 51 software origins (with 1777k revisions and 12k releases) from the Software Heritage Graph Dataset (SWHGD) and their GitHub project boards (with 23k cards) reveal reasonably active project management with phase simplicity can release software versions more frequently and can follow the small release conventions of Extreme Programming. Additionally, the study also reveals that a combination of development and management activities can be applied to predict the possible number of software releases in a month ( ). Debasish Chakroborti, Sristy Sumana Nath, Kevin A. Schneider, Chanchal Kumar Roy |
J. Softw. Evol. Process. | 3 |
| 2022 | SET-STAT-MAP: Extending Parallel Sets for Visualizing Mixed DataabstractMulti-attribute dataset visualizations are often designed based on attribute types, i.e., whether the attributes are categorical or numerical. Parallel Sets and Parallel Coordinates are two well-known techniques to visualize categorical and numerical data, respectively. A common strategy to visualize mixed data is to use multiple information linked view, e.g., Parallel Coordinates are often augmented with maps to explore spatial data with numeric attributes. In this paper, we design visualizations for mixed data, where the dataset may include numerical, categorical, and spatial attributes. The proposed solution SET-STAT-MAP is a harmonious combination of three interactive components: Parallel Sets (visualizes sets determined by the combination of categories or numeric ranges), statistics columns (visualizes numerical summaries of the sets), and a geospatial map view (visualizes the spatial information). We augment these components with colors and textures to enhance users' capability of analyzing distributions of pairs of attribute combinations. To improve scalability, we merge the sets to limit the number of possible combinations to be rendered on the display. We demonstrate the use of Set-stat-map using two different types of datasets: a meteorological dataset and an online vacation rental dataset (Airbnb). To examine the potential of the system, we collaborated with the meteorologists, which revealed both challenges and opportunities for Set-stat-map to be used for real-life visual analytics. Shisong Wang, Debajyoti Mondal, Sara Sadri, Chanchal Kumar Roy, James S. Famiglietti, Kevin A. Schneider |
PacificVis | 6 |
| 2022 | Backports: change types, challenges and strategiesabstractSource code repositories allow developers to manage multiple versions (or branches) of a software system. Pull-requests are used to modify a branch, and backporting is a regular activity used to port changes from a current development branch to other versions. In open-source software, backports are common and often need to be adapted by hand, which motivates us to explore backports and backporting challenges and strategies. In our exploration of 68,424 backports from 10 GitHub projects, we found that bug, test, document, and feature changes are commonly backported. We identified a number of backporting challenges, including that backports were inconsistently linked to their original pull-request (49%), that backports had incompatible code (13%), that backports failed to be accepted (10%), and that there were backporting delays (16 days to create, 5 days to merge). We identified some general strategies for addressing backporting issues. We also noted that backporting strategies depend on the project type and that further investigation is needed to determine their suitability. Furthermore, we created the first-ever backports dataset that can be used by other researchers and practitioners for investigating backports and backporting. Debasish Chakroborti, Kevin A. Schneider, Chanchal Kumar Roy |
ICPC | 2 |
| 2022 | Mining Software Information Sites to Recommend Cross-Language Analogical LibrariesabstractSoftware development is largely dependent on libraries to reuse existing functionalities instead of reinventing the wheel. Software developers often need to find analogical libraries (libraries similar to ones they are already familiar with) as an analogical library may offer improved or additional features. Developers also need to search for analogical libraries across programming languages when developing applications in different languages or for different platforms. However, manually searching for analogical libraries is a time-consuming and difficult task. This paper presents a technique, called XLibRec, that recommends analogical libraries across different programming languages. XLibRec collects Stack Overflow question titles containing library names, library usage information from Stack Overflow posts, and library descriptions from a third party website, Libraries.io. We generate word-vectors for each information and calculate a weight-based cosine similarity score from them to recommend analogical libraries. We performed an extensive evaluation using a large number of analogical libraries across four different programming languages. Results from our evaluation show that the proposed technique can recommend cross-language analogical libraries with great accuracy. The precision for the Top-3 recommendations ranges from 62-81% and has achieved 8-45% higher precision than the state-of-the-art technique. Kawser Wazed Nafi, Muhammad Asaduzzaman, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 5 |
| 2022 | The reproducibility of programming-related issues in Stack Overflow questions
Saikat Mondal, Mohammad Masudur Rahman 0001, Chanchal Kumar Roy, Kevin A. Schneider |
Empir. Softw. Eng. | 4 |
| 2022 | A survey of software architectural change detection and categorization techniques
Amit Kumar Mondal, Kevin A. Schneider, Banani Roy, Chanchal Kumar Roy |
J. Syst. Softw. | 2 |
| 2022 | Evaluating the performance of clone detection tools in detecting cloned co-change candidates
Md. Nadim, Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
J. Syst. Softw. | 4 |
| 2021 | Semantic Slicing of Architectural Change Commits: Towards Semantic Design ReviewabstractSoftware architectural changes involve more than one module or component and are complex to analyze compared to local code changes. Development teams aiming to review architectural aspects (design) of a change commit consider many essential scenarios such as access rules and restrictions on usage of program entities across modules. Moreover, design review is essential when proper architectural formulations are paramount for developing and deploying a system. Untangling architectural changes, recovering semantic design, and producing design notes are the crucial tasks of the design review process. To support these tasks, we construct a lightweight tool [4] that can detect and decompose semantic slices of a commit containing architectural instances. A semantic slice consists of a description of relational information of involved modules, their classes, methods and connected modules in a change instance, which is easy to understand to a reviewer. We extract various directory and naming structures (DANS) properties from the source code for developing our tool. Utilizing the DANS properties, our tool first detects architectural change instances based on our defined metric and then decomposes the slices (based on string processing). Our preliminary investigation with ten open-source projects (developed in Java and Kotlin) reveals that the DANS properties produce highly reliable precision and recall (93-100%) for detecting and generating architectural slices. Our proposed tool will serve as the preliminary approach for the semantic design recovery and design summary generation for the project releases. Amit Kumar Mondal, Chanchal Kumar Roy, Kevin A. Schneider, Banani Roy, Sristy Sumana Nath |
ESEM | 3 |
| 2021 | ContourDiff: Revealing Differential Trends in Spatiotemporal DataabstractChanges in spatiotemporal data may often go unnoticed due to their inherent noise and low variability (e.g., geological processes over years). Commonly used approaches such as side-by-side contour plots and spaghetti plots do not provide a clear idea about the temporal changes in such data. We propose ContourDiff, a vector-based visualization over contour plots to visualize the trends of change across spatial regions and temporal domain. Our approach first aggregates for each location, its value differences from the neighboring points over the temporal domain, and then creates a vector field representing the prominent changes. Finally, it overlays the vectors along the contour paths, revealing differential trends that the contour lines experienced over time. We evaluated our visualization using real-life datasets, consisting of millions of data points, where the visualizations were generated in less than a minute in a single-threaded execution. Our experimental results reveal that ContourDiff can reliably visualize the differential trends, and provide a new way to explore the change pattern in spatiotemporal data. Zonayed Ahmed, Michael Beyene, Debajyoti Mondal, Chanchal Kumar Roy, Christopher Dutchyn, Kevin A. Schneider |
IV | 6 |
| 2021 | FLeCCS: A Technique for Suggesting Fragment-Level Similar Co-change CandidatesabstractWhen a programmer changes a particular code fragment, the other similar code fragments in the code-base may also need to be changed together (i.e., co-changed) consistently to ensure that the software system remains consistent. Existing studies and tools apply clone detectors to identify these similar co-change candidates for a target code fragment. However, clone detectors suffer from a confounding configuration choice problem and it affects their accuracy in retrieving co-change candidates.In our research, we propose and empirically evaluate a lightweight co-change suggestion technique that can automatically suggest fragment level similar co-change candidates for a target code fragment using WA-DiSC (Weighted Average Dice-Sørensen Co-efficient) through a context-sensitive mining of the entire code-base. We apply our technique, FLeCCS (Fragment Level Co-change Candidate Suggester), on six subject systems written in three different programming languages (Java, C, and C#) and compare its performance with the existing state-of-the-art techniques. According to our experiment, our technique outperforms not only the existing code clone based techniques but also the association rule mining based techniques in detecting co-change candidates with a significantly higher accuracy (precision and recall). We also find that File Proximity Ranking performs significantly better than Similarity Extent Ranking when ranking the co-change candidates suggested by our proposed technique. Manishankar Mondal, Chanchal Kumar Roy, Banani Roy, Kevin A. Schneider |
ICPC | 4 |
| 2021 | ArchiNet: A Concept-token based Approach for Determining Architectural Change CategoriesabstractCauses of software architectural change are classified as perfective, preventive, corrective, and adaptive.Change classification is used to promote common approaches for addressing similar changes, produce appropriate design documentation for a release, construct a developer's profile, form a balanced team, support code review, etc.However, automated architectural change classification techniques are in their infancy, perhaps due to the lack of a benchmark dataset and the need for extensive human involvement.To address these shortcomings, we present a benchmark dataset and a text classifier for determining the architectural change rationale from commit descriptions.First, we explored source code properties for change classification independent of project activity descriptions and found poor outcomes.Next, through extensive analysis, we identified the challenges of classifying architectural change from text and proposed a new classifier that uses concept tokens derived from the concept analysis of change samples.We also studied the sensitivity of change classification of various types of tokens present in commit messages.The experimental outcomes employing 10fold and cross-project validation techniques with five popular open-source systems show that the F1 score of our proposed classifier is around 70%.The precision and recall are mostly consistent among all categories of change and more promising than competing methods for text classification. Amit Kumar Mondal, Banani Roy, Sristy Sumana Nath, Kevin A. Schneider |
SEKE | 4 |
| 2021 | A Testing Approach While Re-engineering Legacy Systems: An Industrial Case StudyabstractMany organizations use legacy systems as these systems contain their valuable business rules. However, these legacy systems answer the past requirements but are difficult to maintain and evolve due to old technology use. In this situation, stockholders decide to renovate the system with a minimum amount of cost and risk. Although the renovation process is a more affordable choice over redevelopment, it comes with its risks such as performance loss and failure to obtain quality goals. A proper test process can minimize risks incorporated with the renovation process. This work introduces a testing model tailored for the migration and re-engineering process and employs test automation, which results in early bug detection. Moreover, the automated tests ensure functional sameness between the old and the new system. This process enhances reliability, accuracy, and speed of testing. Hamid Khodabandehloo, Banani Roy, Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 5 |
| 2021 | ID-correspondence: a measure for detecting evolutionary coupling
Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
Empir. Softw. Eng. | 4 |
| 2020 | A Fine-Grained Analysis on the Inconsistent Changes in Code ClonesabstractExisting studies report that inconsistent changes in code clones can introduce bugs or inconsistencies in a software system's code-base. However, inconsistent changes can often be intentional and these might not lead to bugs. Thus, it would be beneficial if we could have an automatic mechanism for proactively determining which inconsistent changes are likely to introduce bugs in a code-base. With this focus, in our research we investigate the underlying factors affecting the possibility that an inconsistent change made to code clones will lead to bugs. We extract the evolutionary history of the clone fragments in open-source software systems and analyze whether clone-types, former clone evolutionary patterns, and clone proximity have impacts on the bug-proneness of the inconsistent changes to clones. We perform our investigation on six subject systems written in three different programming languages (Java, C, and C#) and find that inconsistent changes in Type 3 clones have the highest possibility of introducing bugs among all three clone-types (Type 1, 2, and 3). Moreover, similarity preserving inconsistent changes are significantly more likely to introduce bugs compared to diverging inconsistent changes. Proximity as well as granularity of code clones have significant impacts on their possibilities of experiencing bug-fixes after having inconsistent changes. Findings from our research can be important for minimizing bugs and inconsistencies in software systems during their evolution. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 3 |
| 2020 | Investigating Near-Miss Micro-Clones in Evolving SoftwareabstractCode clones are the same or nearly similar code fragments in a software system's code-base. While the existing studies have extensively studied regular code clones in software systems, micro-clones have been mostly ignored. Although an existing study investigated consistent changes in exact micro-clones, near-miss micro-clones have never been investigated. In our study, we investigate the importance of near-miss micro-clones in software evolution and maintenance by automatically detecting and analyzing the consistent updates that they experienced during the whole period of evolution of our subject systems. We compare the consistent co-change tendency of near-miss micro-clones with that of exact micro-clones and regular code clones. According to our investigation on thousands of revisions of six open-source subject systems written in two different programming languages, near-miss micro-clones have a significantly higher tendency of experiencing consistent updates compared to exact micro-clones and regular (both exact and near-miss) code clones. Consistent updates in near-miss micro-clones have a high tendency of being related with bug-fixes. Moreover, the percentage of commit operations where near-miss micro-clones experience consistent updates is considerably higher than that of regular clones and exact micro-clones. We finally observe that near-miss micro-clones staying in close proximity to each other have a high tendency of experiencing consistent updates. Our research implies that near-miss micro-clones should be considered equally important as of regular clones and exact micro-clones when making clone management decisions. Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
ICPC | 4 |
| 2020 | An Exploratory Study to Find Motives Behind Cross-platform Forks from Software Heritage DatasetabstractThe fork-based development mechanism provides the flexibility and the unified processes for software teams to collaborate easily in a distributed setting without too much coordination overhead. Currently, multiple social coding platforms support fork-based development, such as GitHub, GitLab, and Bitbucket. Although these different platforms virtually share the same features, they have different emphasis. As GitHub is the most popular platform and the corresponding data is publicly available, most of the current studies are focusing on GitHub hosted projects. However, we observed anecdote evidences that people are confused about choosing among these platforms, and some projects are migrating from one platform to another, and the reasons behind these activities remain unknown. With the advances of Software Heritage Graph Dataset (SWHGD), we have the opportunity to investigate the forking activities across platforms. In this paper, we conduct an exploratory study on 10 popular open-source projects to identify cross-platform forks and investigate the motivation behind. Preliminary result shows that cross-platform forks do exist. For the 10 subject systems used in this study, we found 81,357 forks in total among which 179 forks are on GitLab. Based on our qualitative analysis, we found that most of the cross-platform forks that we identified are mirrors of the repositories on another platform, but we still find cases that were created due to preference of using certain functionalities (e.g. Continuous Integration (CI)) supported by different platforms. This study lays the foundation of future research directions, such as understanding the differences between platforms and supporting cross-platform collaboration. Avijit Bhattacharjee, Sristy Sumana Nath, Shurui Zhou, Debasish Chakroborti, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
MSR | 7 |
| 2020 | Associating Code Clones with Association Rules for Change Impact AnalysisabstractWhen a programmer makes changes to a target program entity (files, classes, methods), it is important to identify which other entities might also get impacted. These entities constitute the impact set for the target entity. Association rules have been widely used for discovering the impact sets. However, such rules only depend on the previous co-change history of the program entities ignoring the fact that similar entities might often need to be updated together consistently even if they did not co-change before. Considering this fact, we investigate whether cloning relationships among program entities can be associated with association rules to help us better identify the impact sets. In our research, we particularly investigate whether the impact set detection capability of a clone detector can be utilized to enhance the capability of the state-of-the-art association rule mining technique, Tarmaq, in discovering impact sets. We use the well known clone detector called NiCad in our investigation and consider both regular and micro-clones. Our evolutionary analysis on thousands of commit operations of eight diverse subject systems reveals that consideration of code clones can enhance the impact set detection accuracy of Tarmaq with a significantly higher precision and recall. Micro-clones of 3LOC and 4LOC and regular code clones of 5LOC to 20LOC contribute the most towards enhancing the detection accuracy. Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 4 |
| 2020 | HistoRank: History-Based Ranking of Co-change CandidatesabstractEvolutionary coupling is a well investigated phenomenon during software evolution and maintenance. If two or more program entities co-change (i.e., change together) frequently during evolution, it is expected that the entities are coupled. This type of coupling is called evolutionary coupling or change coupling in the literature. Evolutionary coupling is realized using association rules and two measures: support and confidence. Association rules have been extensively used for predicting co-change candidates for a target program entity (i.e., an entity that a programmer attempts to change). However, association rules often predict a large number of co-change candidates with many false positives. Thus, it is important to rank the predicted co-change candidates so that the true positives get higher priorities. The predicted co-change candidates have always been ranked using the support and confidence measures of the association rules. In our research, we investigate five different ranking mechanisms on thousands of commits of ten diverse subject systems. On the basis of our findings, we propose a history-based ranking approach, HistoRank (History-based Ranking), that analyzes the previous ranking history to dynamically select the most appropriate one from those five ranking mechanisms for ranking co-change candidates of a target program entity. According to our experiment result, HistoRank outperforms each individual ranking mechanism with a significantly better MAP (mean average precision). We investigate different variants of HistoRank and realize that the variant that emphasizes the ranking in the most recent occurrence of co-change in the history performs the best. Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 4 |
| 2020 | CAPS: a supervised technique for classifying Stack Overflow posts concerning API issues
Md. Ahasanuzzaman, Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider |
Empir. Softw. Eng. | 4 |
| 2020 | CROKAGE: effective solution recommendation for programming tasks by leveraging crowd knowledge
Rodrigo Fernandes Gomes da Silva, Chanchal Kumar Roy, Mohammad Masudur Rahman 0001, Kevin A. Schneider, Klérisson Vinícius Ribeiro Paixão, Carlos Eduardo de Carvalho Dantas, Marcelo de Almeida Maia |
Empir. Softw. Eng. | 4 |
| 2020 | A survey on clone refactoring and tracking
Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
J. Syst. Softw. | 3 |
| 2020 | A machine learning based framework for code clone validation
Golam Mostaeen, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider, Jeffrey Svajlenko |
J. Syst. Softw. | 4 |
| 2020 | A universal cross language software similarity detector for open source software categorization
Kawser Wazed Nafi, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
J. Syst. Softw. | 4 |
| 2020 | VizSciFlow: A Visually Guided Scripting Framework for Supporting Complex Scientific Data AnalysisabstractScientific workflow management systems such as Galaxy, Taverna and Workspace, have been developed to automate scientific workflow management and are increasingly being used to accelerate the specification, execution, visualization, and monitoring of data-intensive tasks. For example, the popular bioinformatics platform Galaxy is installed on over 168 servers around the world and the social networking space myExperiment shares almost 4,000 Galaxy scientific workflows among its 10,665 members. Most of these systems offer graphical interfaces for composing workflows. However, while graphical languages are considered easier to use, graphical workflow models are more difficult to comprehend and maintain as they become larger and more complex. Text-based languages are considered harder to use but have the potential to provide a clean and concise expression of workflow even for large and complex workflows. A recent study showed that some scientists prefer script/text-based environments to perform complex scientific analysis with workflows. Unfortunately, such environments are unable to meet the needs of scientists who prefer graphical workflows. In order to address the needs of both types of scientists and at the same time to have script-based workflow models because of their underlying benefits, we propose a visually guided workflow modeling framework that combines interactive graphical user interface elements in an integrated development environment with the power of a domain-specific language to compose independently developed and loosely coupled services into workflows. Our domain-specific language provides scientists with a clean, concise, and abstract view of workflow to better support workflow modeling. As a proof of concept, we developed VizSciFlow, a generalized scientific workflow management system that can be customized for use in a variety of scientific domains. As a first use case, we configured and customized VizSciFlow for the bioinformatics domain. We conducted three user studies to assess its usability, expressiveness, efficiency, and flexibility. Results are promising, and in particular, our user studies show that VizSciFlow is more desirable for users to use than either Python or Galaxy for solving complex scientific problems. Muhammad M. Hossain, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2019 | Investigating Context Adaptation Bugs in Code ClonesabstractThe identical or nearly similar code fragments in a code-base are called code clones. There is a common belief that code cloning (copy/pasting code fragments) can introduce bugs in a software system if the copied code fragments are not properly adapted to their contexts (i.e., surrounding code). However, none of the existing studies have investigated whether such bugs are really present in code clones. We denote these bugs as Context Adaptation Bugs, or simply Context-Bugs, in our paper and investigate the extent to which they can be present in code clones. We define and automatically analyze two clone evolutionary patterns that indicate fixing of Context-Bugs. According to our analysis on thousands of revisions of six open-source subject systems written in Java, C, and C#, code cloning often introduces Context-Bugs in software systems. Around 50% of the clone related bug-fixes can occur for fixing Context-Bugs. Cloning (copy/pasting) a newly created code fragment (i.e., a code fragment that was not added in a former revision) is more likely to introduce Context-Bugs compared to cloning a preexisting fragment (i.e., a code fragment that was added in a former revision). Moreover, cloning across different files appears to have a significantly higher tendency of introducing Context-Bugs compared to cloning within the same file. Finally, Type 3 clones (gapped clones) have the highest tendency of containing Context-Bugs among the three major clone-types. Our findings can be important for early detection as well as removal of Context-Bugs in code clones. Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 4 |
| 2019 | Comparing bug replication in regular and micro code clonesabstractCopying and pasting source code during software development is known as code cloning. Clone fragments with a minimum size of 5 LOC were usually considered in previous studies. In recent studies, clone fragments which are less than 5 LOC are referred as micro-clones. It has been established by the literature that code clones are closely related with software bugs as well as bug replication. None of the previous studies have been conducted on bug-replication of micro-clones. In this paper we investigate and compare bug-replication in between regular and micro-clones. For the purpose of our investigation, we analyze the evolutionary history of our subject systems and identify occurrences of similarity preserving co-changes (SPCOs) in both regular and micro-clones where they experienced bug-fixes. From our experiment on thousands of revisions of six diverse subject systems written in three different programming languages, C, C# and Java we find that the percentage of clone fragments that take part in bug-replication is often higher in micro-clones than in regular code clones. The percentage of bugs that get replicated in micro-clones is almost the same as the percentage in regular clones. Finally, both regular and micro-clones have similar tendencies of replicating severe bugs according to our experiment. Thus, micro-clones in a code-base should not be ignored. We should rather consider these equally important as of the regular clones when making clone management decisions. Judith F. Islam, Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICPC | 4 |
| 2019 | Recommending comprehensive solutions for programming tasks by mining crowd knowledgeabstractDevelopers often search for relevant code examples on the web for their programming tasks. Unfortunately, they face two major problems. First, the search is impaired due to a lexical gap between their query (task description) and the information associated with the solution. Second, the retrieved solution may not be comprehensive, i.e., the code segment might miss a succinct explanation. These problems make the developers browse dozens of documents in order to synthesize an appropriate solution. To address these two problems, we propose CROKAGE (Crowd Knowledge Answer Generator), a tool that takes the description of a programming task (the query) and provides a comprehensive solution for the task. Our solutions contain not only relevant code examples but also their succinct explanations. Our proposed approach expands the task description with relevant API classes from Stack Overflow Q&A threads and then mitigates the lexical gap problems. Furthermore, we perform natural language processing on the top quality answers and then return such programming solutions containing code examples and code explanations unlike earlier studies. We evaluate our approach using 97 programming queries, of which 50% was used for training and 50% was used for testing, and show that it outperforms six baselines including the state-of-art by a statistically significant margin. Furthermore, our evaluation with 29 developers using 24 tasks (queries) confirms the superiority of CROKAGE over the state-of-art tool in terms of relevance of the suggested code examples, benefit of the code explanations and the overall solution quality (code + explanation). Rodrigo Fernandes Gomes da Silva, Chanchal Kumar Roy, Mohammad Masudur Rahman 0001, Kevin A. Schneider, Klérisson Vinícius Ribeiro Paixão, Marcelo de Almeida Maia |
ICPC | 4 |
| 2019 | CLCDSA: Cross Language Code Clone Detection using Syntactical Features and API DocumentationabstractSoftware clones are detrimental to software maintenance and evolution and as a result many clone detectors have been proposed. These tools target clone detection in software applications written in a single programming language. However, a software application may be written in different languages for different platforms to improve the application's platform compatibility and adoption by users of different platforms. Cross language clones (CLCs) introduce additional challenges when maintaining multi-platform applications and would likely go undetected using existing tools. In this paper, we propose CLCDSA, a cross language clone detector which can detect CLCs without extensive processing of the source code and without the need to generate an intermediate representation. The proposed CLCDSA model analyzes different syntactic features of source code across different programming languages to detect CLCs. To support large scale clone detection, the CLCDSA model uses an action filter based on cross language API call similarity to discard non-potential clones. The design methodology of CLCDSA is two-fold: (a) it detects CLCs on the fly by comparing the similarity of features, and (b) it uses a deep neural network based feature vector learning model to learn the features and detect CLCs. Early evaluation of the model observed an average precision, recall and F-measure score of 0.55, 0.86, and 0.64 respectively for the first phase and 0.61, 0.93, and 0.71 respectively for the second phase which indicates that CLCDSA outperforms all available models in detecting cross language clones. Kawser Wazed Nafi, Tonny Shekha Kar, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
ASE | 5 |
| 2019 | Software Engineering by Source Transformation - Experience with TXL (Most Influential Paper, SCAM 2001)
James R. Cordy, Thomas R. Dean, Andrew J. Malton, Kevin A. Schneider |
SCAM | 4 |
| 2019 | An Exploratory Study on Automatic Architectural Change Analysis Using Natural Language Processing TechniquesabstractContinuous architecture is vital for developing large, complex software systems and supporting continuous delivery, integration, and testing practices. Researchers and practitioners investigate models and rules for managing change to support architecture continuity. They employ manual techniques to analyze software change, categorizing the changes as perfective, corrective, adaptive, and preventive. However, a manual approach is impractical for analyzing systems involving thousands of artefacts as it is time-consuming, labor-intensive, and error-prone. In this paper, we investigate whether an automatic technique incorporating free-form natural language text (e.g., developers' communication and commit messages) is an effective solution for architectural change analysis. Our experiments with multiple projects showed encouraging results for detecting architectural messages using our proposed language model. Although architectural change categorization for the preventive class is moderate, the outcome for the random dataset is insignificant in general (around a 45% F1 score). We investigated the causes of the unpromising outcome. Overall, our study reveals that our automated architectural change analysis tool would be fruitful only if the developers provide considerable technical details in the commit messages or other text. Amit Kumar Mondal, Banani Roy, Kevin A. Schneider |
SCAM | 3 |
| 2019 | CloneCognition: machine learning based code clone validation toolabstractA code clone is a pair of similar code fragments, within or between software systems. To detect each possible clone pair from a software system while handling the complex code structures, the clone detection tools undergo a lot of generalization of the original source codes. The generalization often results in returning code fragments that are only coincidentally similar and not considered clones by users, and hence requires manual validation of the reported possible clones by users which is often both time-consuming and challenging. In this paper, we propose a machine learning based tool 'CloneCognition' (Open Source Codes: https://github.com/pseudoPixels/CloneCognition ; Video Demonstration: https://www.youtube.com/watch?v=KYQjmdr8rsw) to automate the laborious manual validation process. The tool runs on top of any code clone detection tools to facilitate the clone validation process. The tool shows promising clone classification performance with an accuracy of up to 87.4%. The tool also exhibits significant improvement in the results when compared with state-of-the-art techniques for code clone validation. Golam Mostaeen, Jeffrey Svajlenko, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
ESEC/SIGSOFT FSE | 5 |
| 2019 | An empirical study on bug propagation through code cloning
Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
J. Syst. Softw. | 4 |
| 2019 | Designing for Real-Time Groupware Systems to Support Complex Scientific Data AnalysisabstractScientific Workflow Management Systems (SWfMSs) have become popular for accelerating the specification, execution, visualization, and monitoring of data-intensive scientific experiments. Unfortunately, to the best of our knowledge no existing SWfMSs directly support collaboration. Data is increasing in complexity, dimensionality, and volume, and the efficient analysis of data often goes beyond the realm of an individual and requires collaboration with multiple researchers from varying domains. In this paper, we propose a groupware system architecture for data analysis that in addition to supporting collaboration, also incorporates features from SWfMSs to support modern data analysis processes. As a proof of concept for the proposed architecture we developed SciWorCS - a groupware system for scientific data analysis. We present two real-world use-cases: collaborative software repository analysis and bioinformatics data analysis. The results of the experiments evaluating the proposed system are promising. Our bioinformatics user study demonstrates that SciWorCS can leverage real-world data analysis tasks by supporting real-time collaboration among users. Golam Mostaeen, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2019 | Clone-World: A visual analytic system for large scale software clonesabstractWith the era of big data approaching, the number of software systems, their dependencies, as well as the complexity of the individual system is becoming larger and more intricate. Understanding these evolving software systems is thus a primary challenge for cost-effective software management and maintenance. In this paper we perform a case study with evolving code clones. The programmers often need to manually analyze the co-evolution of clone fragments to decide about refactoring, tracking, and bug removal. However, manual analysis is time consuming, and nearly infeasible for a large number of clones, e.g., with millions of similarity pairs, where clones are evolving over hundreds of software revisions. We propose an interactive visual analytics system, Clone-World , which leverages big data visualization approach to manage code clones in large software systems. Clone-World , gives an intuitive yet powerful solution to the clone analytic problems. Clone-World combines multiple information-linked zoomable views, where users can explore and analyze clones through interactive exploration in real time. User studies and experts’ reviews suggest that Clone-World may assist developers in many real-life software development and maintenance scenarios. We believe that Clone-World will ease the management and maintenance of clones, and inspire future innovation to adapt visual analytics to manage big software systems. Debajyoti Mondal, Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider, Shisong Wang |
Vis. Informatics | 4 |
| 2018 | Optimized Storing of Workflow Outputs through Mining Association RulesabstractWorkflows are frequently built and used to systematically process large datasets using workflow management systems (WMS). A workflow (i.e., a pipeline) is a finite set of processing modules organized as a series of steps that is applied to an input dataset to produce a desired output. In a workflow management system, users generally create workflows manually for their own investigations. However, workflows can sometimes be lengthy and the constituent processing modules might often be computationally expensive. In this situation, it would be beneficial if users could reuse intermediate stage results generated by previously executed workflows for executing their current workflow.In this paper, we propose a novel technique based on association rule mining for suggesting which intermediate stage results from a workflow that a user is going to execute should be stored for reusing in the future. We call our proposed technique, RISP (Recommending Intermediate States from Pipelines). According to our investigation on hundreds of workflows from two scientific workflow management systems, our proposed technique can efficiently suggest intermediate state results to store for future reuse. The results that are suggested to be stored have a high reuse frequency. Moreover, for creating around 51% of the entire pipelines, we can reuse results suggested by our technique. Finally, we can achieve a considerable gain (74% gain) in execution time by reusing intermediate results stored by the suggestions provided by our proposed technique. We believe that our technique (RISP) has the potential to have a significant positive impact on Big-Data systems, because it can considerably reduce execution time of the workflows through appropriate reuse of intermediate state results, and hence, can improve the performance of the systems. Debasish Chakroborti, Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
IEEE BigData | 5 |
| 2018 | [Research Paper] Detecting Evolutionary Coupling Using Transitive Association RulesabstractIf two or more program entities (such as files, classes, methods) co-change (i.e., change together) frequently during software evolution, then it is likely that these two entities are coupled (i.e., the entities are related). Such a coupling is termed as evolutionary coupling in the literature. The concept of traditional evolutionary coupling restricts us to assume coupling among only those entities that changed together in the past. The entities that did not co-change in the past might also have coupling. However, such couplings can not be retrieved using the current concept of detecting evolutionary coupling in the literature. In this paper, we investigate whether we can detect such couplings by applying transitive rules on the evolutionary couplings detected using the traditional mechanism. We call these couplings that we detect using our proposed mechanism as transitive evolutionary couplings. According to our research on thousands of revisions of four subject systems, transitive evolutionary couplings combined with the traditional ones provide us with 13.96% higher recall and 5.56% higher precision in detecting future co-change candidates when compared with a state-of-the-art technique. Md. Anaytul Islam, Md. Moksedul Islam, Manishankar Mondal, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
SCAM | 6 |
| 2018 | [Research Paper] On the Use of Machine Learning Techniques Towards the Design of Cloud Based Automatic Code Clone Validation ToolsabstractA code clone is a pair of code fragments, within or between software systems that are similar. Since code clones often negatively impact the maintainability of a software system, a great many numbers of code clone detection techniques and tools have been proposed and studied over the last decade. To detect all possible similar source code patterns in general, the clone detection tools work on syntax level (such as texts, tokens, AST and so on) while lacking user-specific preferences. This often means the reported clones must be manually validated prior to any analysis in order to filter out the true positive clones from task or user-specific considerations. This manual clone validation effort is very time-consuming and often error-prone, in particular for large-scale clone detection. In this paper, we propose a machine learning based approach for automating the validation process. In an experiment with clones detected by several clone detectors in several different software systems, we found our approach has an accuracy of up to 87.4% when compared against the manual validation by multiple expert judges. The proposed method shows promising results in several comparative studies with the existing related approaches for automatic code clone validation. We also present our experimental results in terms of different code clone detection tools, machine learning algorithms and open source software systems. Golam Mostaeen, Jeffrey Svajlenko, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
SCAM | 5 |
| 2018 | [Research Paper] CroLSim: Cross Language Software Similarity Detector Using API DocumentationabstractIn today's open source era, developers look forsimilar software applications in source code repositories for anumber of reasons, including, exploring alternative implementations, reusing source code, or looking for a better application. However, while there are a great many studies for finding similarapplications written in the same programming language, there isa marked lack of studies for finding similar software applicationswritten in different languages. In this paper, we fill the gapby proposing a novel modelCroLSimwhich is able to detectsimilar software applications across different programming lan-guages. In our approach, we use the API documentation tofind relationships among the API calls used by the differentprogramming languages. We adopt a deep learning based word-vector learning method to identify semantic relationships amongthe API documentation which we then use to detect cross-language similar software applications. For evaluating CroLSim, we formed a repository consisting of 8,956 Java, 7,658 C#, and 10,232 Python applications collected from GitHub. Weobserved thatCroLSimcan successfully detect similar softwareapplications across different programming languages with a meanaverage precision rate of 0.65, an average confidence rate of3.6 (out of 5) with 75% high rated successful queries, whichoutperforms all related existing approaches with a significantperformance improvement. Kawser Wazed Nafi, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider |
SCAM | 4 |
| 2018 | Classifying stack overflow posts on API issuesabstractThe design and maintenance of APIs are complex tasks due to the constantly changing requirements of its users. Despite the efforts of its designers, APIs may suffer from a number of issues (such as incomplete or erroneous documentation, poor performance, and backward incompatibility). To maintain a healthy client base, API designers must learn these issues to fix them. Question answering sites, such as Stack Overflow (SO), has become a popular place for discussing API issues. These posts about API issues are invaluable to API designers, not only because they can help to learn more about the problem but also because they can facilitate learning the requirements of API users. However, the unstructured nature of posts and the abundance of non-issue posts make the task of detecting SO posts concerning API issues difficult and challenging. In this paper, we first develop a supervised learning approach using a Conditional Random Field (CRF), a statistical modeling method, to identify API issue-related sentences. We use the above information together with different features of posts and experience of users to build a technique, called CAPS, that can classify SO posts concerning API issues. Evaluation of CAPS using carefully curated SO posts on three popular API types reveals that the technique outperforms all three baseline approaches we consider in this study. We also conduct studies to test the generalizability of CAPS results and to understand the effects of different sources of information on it. Muhammad Ahasanuzzaman, Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 4 |
| 2018 | Micro-clones in evolving softwareabstractDetection, tracking, and refactoring of code clones (i.e., identical or nearly similar code fragments in the code-base of a software system) have been extensively investigated by a great many studies. Code clones have often been considered bad smells. While clone refactoring is important for removing code clones from the code-base, clone tracking is important for consistently updating code clones that are not suitable for refactoring. In this research we investigate the importance of micro-clones (i.e., code clones of less than five lines of code) in consistent updating of the code-base. While the existing clone detectors and trackers have ignored micro clones, our investigation on thousands of commits from six subject systems imply that around 80% of all consistent updates during system evolution occur in micro clones. The percentage of consistent updates occurring in micro clones is significantly higher than that in regular clones according to our statistical significance tests. Also, the consistent updates occurring in micro-clones can be up to 23% of all updates during the whole period of evolution. According to our manual analysis, around 83% of the consistent updates in micro-clones are non-trivial. As micro-clones also require consistent updates like the regular clones, tracking or refactoring micro-clones can help us considerably minimize effort for consistently updating such clones. Thus, micro-clones should also be taken into proper consideration when making clone management decisions. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 3 |
| 2018 | Is cloned code really stable?
Manishankar Mondal, Md. Saidur Rahman 0002, Chanchal Kumar Roy, Kevin A. Schneider |
Empir. Softw. Eng. | 4 |
| 2018 | Bug-proneness and late propagation tendency of code clones: A Comparative study on different clone types
Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
J. Syst. Softw. | 3 |
| 2017 | Towards a Reference Architecture for Cloud-Based Plant Genotyping and Phenotyping Analysis FrameworksabstractThe domain of plant genotyping and phenotyping presents a number of challenges in the area of large data computation. Various tools and systems have been developed to automate the scientific workflows and support the computational needs of this domain. In this paper, we review a number of the popular systems (i.e., Galaxy, iPlant, GenAp and LemnaTec) in the domain of plant genotyping and phenotyping using the scenario-based architectural analysis method (SAAM). In particular, we focus on how different stakeholders are using these systems in a variety of scenarios and to what extent the systems support their needs. Our SAAM analysis shows that the existing systems have shortcomings. For example, they are limited in their support for high throughput processing of large amounts of heterogeneous types of data. Based on our findings we propose a reference architecture along with a preliminary evaluation in the subject domain. The reference architecture and its evaluation is aimed at helping developers/architects create suitable architectural designs and select appropriate technologies when developing plant phenotyping and genotyping systems. Banani Roy, Amit Kumar Mondal, Chanchal Kumar Roy, Kevin A. Schneider, Kawser Wazed Nafi |
ICSA | 4 |
| 2017 | Recommending Framework Extension ExamplesabstractThe use of software frameworks enables the delivery of common functionality but with significantly less effort than when developing from scratch. To meet application specific requirements, the behavior of a framework needs to be customized via extension points. A common way of customizing framework behavior is by passing a framework related object as an argument to an API call. Such an object can be created by subclassing an existing framework class or interface, or by directly customizing an existing framework object. However, to do this effectively requires developers to have extensive knowledge of the framework's extension points and their interactions. To aid the developers in this regard, we propose and evaluate a graph mining approach for extension point management. Specifically, we propose a taxonomy of extension patterns to categorize the various ways an extension point has been used in the code examples. Our approach mines a large amount of code examples to discover all extension points and patterns for each framework class. Given a framework class that is being used, our approach aids the developer by following a two-step recommendation process. First, it recommends all the extension points that are available in the class. Once the developer chooses an extension point, our approach then discovers all of its usage patterns and recommends the best code examples for each pattern. Using five frameworks, we evaluate the performance of our two-step recommendation, in terms of precision, recall, and F-measure. We also report several statistics related to framework extension points. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou |
ICSME | 3 |
| 2017 | Bug Propagation through Code Cloning: An Empirical StudyabstractCode clones are defined to be the identical or nearly similar code fragments in a code-base. According to a number of existing studies, code clones are directly related to bugs and inconsistencies in software systems. Code cloning (i.e., creating code clones) is suspected to propagate temporarily hidden bugs from one code fragment to another. However, there is no study on the intensity of bug-propagation through code cloning.In this paper we present our empirical study on bug-propagation through code cloning. We define two clone evolution patterns that reasonably indicate bug propagation through code cloning. We first identify code clones that experienced bug-fix changes by analyzing software evolution history, and then determine which of these code clones evolved following the bug propagation patterns. According to our study on thousands of commits of four open-source subject systems written in Java, up to 33% of the clone fragments that experience bug-fix changes can contain propagated bugs. Around 28.57% of the bug-fixes experienced by the code clones can occur for fixing propagated bugs. We also find that near-miss clones are primarily involved with bug-propagation rather than identical clones. The clone fragments involved with bug propagation are mostly method clones. Bug propagation is more likely to occur in the clone fragments that are created in the same commit operation rather than in different commits. Our findings are important for prioritizing code clones for refactoring and tracking from the perspective of bug propagation. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 3 |
| 2017 | Identifying code clones having high possibilities of containing bugsabstractCode cloning has emerged as a controversial term in software engineering research and practice because of its positive and negative impacts on software evolution and maintenance. Researchers suggest managing code clones through refactoring and tracking. Given the huge number of code clones in a software system's code-base, it is essential to identify the most important ones to manage. In our research, we investigate which clone fragments have high possibilities of containing bugs so that such clones can be prioritized for refactoring and tracking to help minimize future bug-fixing tasks. Existing studies on clone bug-proneness cannot pinpoint code clones that are likely to experience bug-fixes in the future. According to our analysis on thousands of revisions of four diverse subject systems written in Java, change frequency of code clones does not indicate their bug-proneness (i.e., does not indicate their tendencies of experiencing bug-fixes in future). Bug-proneness is mainly related with change recency of code clones. In other words, more recently changed code clones have a higher possibility of containing bugs. Moreover, for the code clones that were not changed previously we observed that clones that were created more recently have higher possibilities of experiencing bug-fixes. Thus, our research reveals the fact that bug-proneness of code clones mainly depends on how recently they were changed or created (for the ones that were not changed before). It invalidates the common intuition regarding the relatedness between high change frequency and bug-proneness. We believe that code clones should be prioritized for management considering their change recency or recency of creation (for the unchanged ones). Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICPC | 3 |
| 2017 | FEMIR: a tool for recommending framework extension examplesabstractSoftware frameworks enable developers to reuse existing well tested functionalities instead of taking the burden of implementing everything from scratch. However, to meet application specific requirements, the frameworks need to be customized via extension points. This is often done by passing a framework related object as an argument to an API call. To enable such customizations, the object can be created by extending a framework class, implementing an interface, or changing the properties of the object via API calls. However, it is both a common and non-trivial task to find all the details related to the customizations. In this paper, we present a tool, called FEMIR, that utilizes partial program analysis and graph mining technique to detect, group, and rank framework extension examples. The tool extends existing code completion infrastructure to inform developers about customization choices, enabling them to browse through extension points of a framework, and frequent usages of each point in terms of code examples. A video demo is made available at https://asaduzzamanparvez.wordpress.com/femir. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou |
ASE | 3 |
| 2017 | A Comparative Study of Software Bugs in Clone and Non-Clone CodeabstractCode cloning is a recurrent operation in everyday software development.Whether it is a good or bad practice is an ongoing debate among researchers and developers for the last few decades.In this paper, we conduct a comparative study on bugproneness in clone code and non-clone code by analyzing commit logs.According to our inspection on thousands of revisions of seven diverse subject systems, the percentage of changed files due to bug-fix commits is significantly higher in clone code compared with non-clone code.We perform a Mann-Whitney-Wilcoxon (MWW) test to show the statistical significance of our findings.Finally, the possibility of severe bugs occurring is higher in clone code than in non-clone code.Bug-fixing changes affecting clone code should be considered more carefully.According to our findings, clone code appears to be more bug-prone than non-clone code. Judith F. Islam, Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
SEKE | 4 |
| 2017 | Comparing Software Bugs in Clone and Non-clone Code: An Empirical StudyabstractCode cloning is a recurrent operation in everyday software development. Whether it is a good or bad practice is an ongoing debate among researchers and developers for the last few decades. In this paper, we conduct a comparative study on bug-proneness in clone code and non-clone code by analyzing commit logs. According to our inspection of thousands of revisions of seven diverse subject systems, the percentage of changed files due to bug-fix commits is significantly higher in clone code compared with non-clone code. We perform a Mann–Whitney–Wilcoxon (MWW) test to show the statistical significance of our findings. In addition, the possibility of occurrence of severe bugs is higher in clone code than in non-clone code. Bug-fixing changes affecting clone code should be considered more carefully. Finally, our manual investigation shows that clone code containing if-condition and if–else blocks has a high risk of having severing bugs. Changes to such types of clone fragments should be done carefully during software maintenance. According to our findings, clone code appears to be more bug-prone than non-clone code. Judith F. Islam, Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2016 | Mining duplicate questions in stack overflowabstractStack Overflow is a popular question answering site that is focused on programming problems. Despite efforts to prevent asking questions that have already been answered, the site contains duplicate questions. This may cause developers to unnecessarily wait for a question to be answered when it has already been asked and answered. The site currently depends on its moderators and users with high reputation to manually mark those questions as duplicates, which not only results in delayed responses but also requires additional efforts. In this paper, we first perform a manual investigation to understand why users submit duplicate questions in Stack Overflow. Based on our manual investigation we propose a classification technique that uses a number of carefully chosen features to identify duplicate questions. Evaluation using a large number of questions shows that our technique can detect duplicate questions with reasonable accuracy. We also compare our technique with DupPredictor, a state-of-the-art technique for detecting duplicate questions, and we found that our proposed technique has a better recall-rate than that technique. Muhammad Ahasanuzzaman, Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider |
MSR | 4 |
| 2016 | How developers use exception handling in Java?abstractException handling is a technique that addresses exceptional conditions in applications, allowing the normal flow of execution to continue in the event of an exception and/or to report on such events. Although exception handling techniques, features and bad coding practices have been discussed both in developer communities and in the literature, there is a marked lack of empirical evidence on how developers use exception handling in practice. In this paper we use the Boa language and infrastructure to analyze 274k open source Java projects in GitHub to discover how developers use exception handling. We not only consider various exception handling features but also explore bad coding practices and their relation to the experience of developers. Our results provide some interesting insights. For example, we found that bad exception handling coding practices are common in open source Java projects and regardless of experience all developers use bad exception handling coding practices. Muhammad Asaduzzaman, Muhammad Ahasanuzzaman, Chanchal Kumar Roy, Kevin A. Schneider |
MSR | 4 |
| 2016 | A Simple, Efficient, Context-sensitive Approach for Code CompletionabstractAbstract Code completion helps developers use application programming interfaces (APIs) and frees them from remembering every detail. In this paper, we first describe a novel technique called Context‐sensitive Code Completion (CSCC) for improving the performance of API method call completion. CSCC is context sensitive in that it uses new sources of information as the context of a target method call. CSCC indexes method calls in code examples by their context. To recommend completion proposals, CSCC ranks candidate methods by the similarities between their contexts and the context of the target call. Evaluation using a set of subject systems and five popular state‐of‐the‐art techniques suggests that CSCC performs better than existing type or example‐based code completion systems. We conduct experiments to find how different contextual elements of the target call benefit CSCC. Next, we investigate the adaptability of the technique to support another form of code completion, i.e., field completion. Evaluation with eight different subject systems suggests that CSCC can easily support field completion with high accuracy. Finally, we compare CSCC with four popular statistical language models that support code completion. Results of statistical tests from our study suggest that CSCC not only outperforms those techniques that are based on token level language models, but also in most cases performs better or equally well with GraLan, the state‐of‐the‐art graph‐based language model. Copyright © 2016 John Wiley & Sons, Ltd. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou |
J. Softw. Evol. Process. | 3 |
| 2016 | A comparative study on the intensity and harmfulness of late propagation in near-miss code clones
Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
Softw. Qual. J. | 3 |
| 2015 | Exploring API method parameter recommendationsabstractA number of techniques have been developed that support method call completion. However, there has been little research on the problem of method parameter completion. In this paper, we first present a study that helps us to understand how developers complete method parameters. Based on our observations, we developed a recommendation technique, called Parc, that collects parameter usage context using a source code localness property that suggests that developers tend to collocate related code fragments. Parc uses previous code examples together with contextual and static type analysis to recommend method parameters. Evaluating our technique against the only available state-of-the-art tool using a number of subject systems and different Java libraries shows that our approach has potential. We also explore the parameter recommendation support provided by the Eclipse Java Development Tools (JDT). Finally, we discuss limitations of our proposed technique and outline future research directions. Muhammad Asaduzzaman, Chanchal Kumar Roy, Samiul Monir, Kevin A. Schneider |
ICSME | 4 |
| 2015 | PARC: Recommending API methods parametersabstractAPIs have grown considerably in size. To free developers from remembering every detail of an API, code completion has become an integral part of modern IDEs. Most work on code completion targets completing API method calls and leaves the task of completing method parameters to the developers. However, parameter completion is also a non-trivial task. We present an Eclipse plugin, called PARC, that supports automatic completion of API method parameters. The tool is based on the localness property of source code, which states that developers tend to put related code fragments close together. PARC combines contextual and static type analysis to support a wide range of parameter expression types. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 3 |
| 2015 | A comparative study on the bug-proneness of different types of code clonesabstractCode clones are defined to be the exactly or nearly similar code fragments in a software system's code-base. The existing clone related studies reveal that code clones are likely to introduce bugs and inconsistencies in the code-base. However, although there are different types of clones, it is still unknown which types of clones have a higher likeliness of introducing bugs to the software systems and so, should be considered more important for managing with techniques such as refactoring or tracking. With this focus, we performed a study that compared the bug-proneness of the major clone-types: Type 1, Type 2, and Type 3. According to our experimental results on thousands of revisions of seven diverse subject systems, Type 3 clones exhibit the highest bug-proneness among the three clone-types. The bug-proneness of Type 1 clones is the lowest. Also, Type 3 clones have the highest likeliness of being co-changed consistently while experiencing bug-fixing changes. Moreover, the Type 3 clones that experience bug-fixes have a higher possibility of evolving following a Similarity Preserving Change Pattern (SPCP) compared to the bug-fix clones of the other two clone-types. From the experimental results it is clear that Type 3 clones should be given a higher priority than the other two clone-types when making clone management decisions. We believe that our study provides useful implications for ranking clones for refactoring and tracking. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 3 |
| 2015 | SPCP-Miner: A tool for mining code clones that are important for refactoring or trackingabstractCode cloning has both positive and negative impacts on software maintenance and evolution. Focusing on the issues related to code cloning, researchers suggest to manage code clones through refactoring and tracking. However, it is impractical to refactor or track all clones in a software system. Thus, it is essential to identify which clones are important for refactoring and also, which clones are important for tracking. In this paper, we present a tool called SPCP-Miner which is the pioneer one to automatically identify and rank the important refactoring as well as important tracking candidates from the whole set of clones in a software system. SPCP-Miner implements the existing techniques that we used to conduct a large scale empirical study on SPCP clones (i.e., the clones that evolved following a Similarity Preserving Change Pattern called SPCP). We believe that SPCP-Miner can help us in better management of code clones by suggesting important clones for refactoring or tracking. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
SANER | 3 |
| 2014 | CSCC: Simple, Efficient, Context Sensitive Code CompletionabstractCode Completion helps developers learn APIs and frees them from remembering every detail. In this paper, we describe a novel technique called CSCC (Context Sensitive Code Completion) for improving the performance of API method call completion. CSCC is context sensitive in that it uses new sources of information as the context of a target method call. CSCC indexes method calls in code examples by their contexts. To recommend completion proposals, CSCC ranks candidate methods by the similarities between their contexts and the context of the target call. Evaluation using a set of subject systems and five popular state of-the-art techniques suggests that CSCC performs better than existing type or example-based code completion systems. We also investigate how the different contextual elements of the target call benefit CSCC. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou |
ICSME | 3 |
| 2014 | Context-Sensitive Code Completion Tool for Better API UsabilityabstractDevelopers depend on APIs of frameworks and libraries to support the development process. Due to the large number of existing APIs, it is difficult to learn, remember, and use them during the development of a software. To mitigate the problem, modern integrated development environments provide code completion facilities that free developers from remembering every detail. In this paper, we introduce CSCC, a simple, efficient context-sensitive code completion tool that leverages previous code examples to support method completion. Compared to other existing code completion tools, CSCC uses new sources of contextual information together with lightweight source code analysis to better recommend API method calls. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou |
ICSME | 3 |
| 2014 | A Fine-Grained Analysis on the Evolutionary Coupling of Cloned CodeabstractCode clones are identical or similar code fragments in a code base. A group of code fragments that are similar to one another forms a clone class. Clone fragments from the same clone class often need to be changed together consistently and thus, they exhibit evolutionary coupling. Evolutionary coupling among clone fragments within a clone class has already been investigated and reported. However, a change to a clone fragment of a clone class may also trigger changes to non-cloned code as well as to clone fragments of other clone classes. Such coupling information is equally important for the proper management of clones during software maintenance. Unfortunately, there are no such studies reported in the literature. In this paper, we describe a large scale empirical study that we conduct to examine whether a clone fragment from a particular clone class exhibits evolutionary coupling with non-clone fragments and/or with clone fragments of other clone classes. Our experimental results on thousands of revisions of six diverse subject systems written in two programming languages indicate the presence of such couplings. We consider both exact and near-miss clones in our study. By analyzing the evolutionary couplings of a particular clone fragment from a particular clone class, we are able to predict its three types of co-change candidates with considerable accuracy in terms of precision and recall. These co-change candidates are: (1) non-clone fragments, (2) clone fragments from clone classes other than its own class, and (3) other clone fragments from its own clone class. Thus, we can improve existing clone tracking techniques so that they can also infer and suggest which non-clone fragments as well as which clone fragments from other clone classes might need to be co-changed correspondingly when modifying a clone fragment from a particular clone class. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 3 |
| 2014 | Interactive Visualization of Bug Reports Using Topic Evolution and Extractive SummariesabstractSoftware bug reports are important project artifacts that evolve throughout the life of a software project. Software bugs are issues that are reported by users when these issues hinder their work. Software projects evolve over time as bugs are addressed and new features are added. Managing bugs can be a significant challenge as a project manager generally needs to be aware of all the bug reports for the current version, and this can be even more challenging when the number of bug reports becomes large. It is preferable that a developer new to a project improves her knowledge with the project along with the bug reports during working on it, which is likely to help her avoid or handle the reported issues. In this paper, we propose a prototype that assists developers review a project's bug reports by interactively visualizing insightful information regarding the bug reports using topic analysis. In addition, in order to reduce developers' time and efforts when studying a bug report, the proposed prototype also provides an extractive summary visualization of each bug report. In this research, it is shown that our proposed prototype performs better in terms of precision, recall, and F-measure than a baseline approach that uses time-sensitive keyword extraction. Shamima Yeasmin, Chanchal Kumar Roy, Kevin A. Schneider |
ICSME | 3 |
| 2014 | Prediction and ranking of co-change candidates for clonesabstractCode clones are identical or similar code fragments scattered in a code-base. A group of code fragments that are similar to one another form a clone group. Clones in a particular group often need to be changed together (i.e., co-changed) consistently. However, all clones in a group might not require consistent changes, because some clone fragments might evolve independently. Thus, while changing a particular clone fragment, it is important for a programmer to know which other clone fragments in the same group should be consistently co-changed with that particular clone fragment. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
MSR | 3 |
| 2014 | Automatic Identification of Important Clones for Refactoring and TrackingabstractCode cloning is a controversial software engineering practice due to contradictory claims regarding its impacts on software evolution and maintenance. While a number of studies identify some positive aspects of code clones, there is strong empirical evidence of some negative impacts of clones too. Focusing on the issues related to clones researchers suggest to manage code clones through detection, refactoring, and tracking. However, all clones in a software system are not suitable for refactoring or tracking. Thus, it is important to identify which clones we should consider for refactoring and which clones should be considered for tracking. In this research work we apply the concept of evolutionary coupling to identify clones that are important for refactoring or tracking. By mining software evolution history, we determine and analyze constrained association rules of clone fragments that evolved following a particular change pattern called Similarity Preserving Change Pattern and are important from the perspective of refactoring and tracking. According to our investigation with rigorous manual analysis on thousands of revisions of six diverse subject systems covering two programming languages, overall 13.20% of all clones in a software system are important candidates for refactoring, and overall 10.27% of all clones are important candidates for tracking. Our implemented system can automatically identify these important candidates and thus, can help us in better maintenance of code clones in terms of refactoring and tracking. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
SCAM | 3 |
| 2014 | An insight into the dispersion of changes in cloned and non-cloned code: A genealogy based empirical study
Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
Sci. Comput. Program. | 3 |
| 2013 | LHDiff: A Language-Independent Hybrid Approach for Tracking Source Code LinesabstractTracking source code lines between two different versions of a file is a fundamental step for solving a number of important problems in software maintenance such as locating bug introducing changes, tracking code fragments or defects across versions, merging file versions, and software evolution analysis. Although a number of such approaches are available in the literature, their performance is sensitive to the kind and degree of source code changes. There is also a marked lack of study on the effect of change types on source location tracking techniques. In this paper, we propose a language-independent technique, LHDiff, for tracking source code lines across versions that leverages simhash technique together with heuristics to improve accuracy. We evaluate our approach against state-of-the- art techniques using benchmarks containing different degrees of changes where files are selected from real world applications. We further evaluate LHDiff with other techniques using a mutation based analysis to understand how different types of changes affect their performance. The results reveal that our technique is more effective than language-independent approaches and no worse than some language-dependent techniques. In our study LHDiff even shows better performance than a state-of-the-art language- dependent approach. In addition, we also discuss limitations of different line tracking techniques including ours and propose future research directions. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Massimiliano Di Penta |
ICSM | 3 |
| 2013 | LHDiff: Tracking Source Code Lines to Support Software Maintenance ActivitiesabstractTracking lines across versions of a file is a necessary step for solving a number of problems during software development and maintenance. Examples include, but are not limited to, locating bug-inducing changes, tracking code fragments or vulnerable instructions across versions, co-change analysis, merging file versions, reviewing source code changes, and software evolution analysis. In this tool demonstration, we present a language-independent line-level location tracker, named LHDiff, that can be used to track lines and analyze changes in various kinds of software artifacts, ranging from source code to arbitrary text files. The tool can effectively detect changed or moved lines across versions of a file, has the ability to detect line splits, and can easily be integrated with existing version control systems. It overcomes the limitations of existing language-independent techniques and is even comparable to tools that are language dependent. In addition to describing the tool, we also describe its effectiveness in analyzing source code artifacts. Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Massimiliano Di Penta |
ICSM | 3 |
| 2013 | gCad: A Near-Miss Clone Genealogy Extractor to Support Clone Evolution AnalysisabstractUnderstanding the evolution of code clones is important for both developers and researchers to understand the maintenance implications of clones and to design robust clone management systems. Generally, a study of clone evolution starts with extracting clone genealogies across multiple versions of a program and classifying them according to their change patterns. Although these tasks are straightforward for exact clones, extracting the history of near-miss clones and classifying their change patterns automatically is challenging due to the potential diverse variety of clone fragments even in the same clone class. In this tool demonstration paper we describe the design and implementation of a near-miss clone genealogy extractor, gCad, that can extract and classify both exact and near-miss clone genealogies. Developers and researchers can compute a wide range of popular metrics regarding clone evolution by simply post processing the gCad results. gCad scales well to large subject systems, works for different granularities of clones, and adapts easily to popular clone detection tools. Ripon K. Saha, Chanchal Kumar Roy, Kevin A. Schneider |
ICSM | 3 |
| 2013 | Insight into a method co-change pattern to identify highly coupled methods: An empirical studyabstractIn this paper, we describe an empirical study of a unique method co-change pattern that has the potential to pinpoint design deficiency in a software system. We automatically identify this pattern by inspecting the method co-change history using reasonable constraints on method association rules. We also investigate the effect of code clones on the method co-changes identified according to the pattern, because there is a common intuition that clone fragments from the same clone class often require corresponding changes to ensure they remain consistent with each other. According to our in-depth investigation on hundreds of revisions of seven open-source software systems considering three types of clones (Type 1, Type 2, Type 3), our identified pattern helps us detect methods that are logically coupled with multiple other methods and that exhibit a significantly higher modification frequency than other methods. We call the methods detected by the pattern MMCGs (Methods appearing in Multiple Commit Groups) considering the pattern semantic. MMCGs can be considered as the candidates for restructuring in order to minimize coupling as well as to reduce the change-proneness of a software system. According to our observation, code clones have a significant effect on method co-changes as well as on MMCGs. We believe that clone refactoring can help us minimize evolutionary coupling among methods. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICPC | 3 |
| 2013 | Improving the detection accuracy of evolutionary couplingabstractIf two or more program entities (e.g., files, classes, methods) co-change frequently during software evolution, these entities are said to have evolutionary coupling. The entities that frequently co-change (i.e., exhibit evolutionary coupling) are likely to have logical coupling (or dependencies) among them. Association rules and two related measurements, Support and Confidence, have been used to predict whether two or more co-changing entities are logically coupled. In this paper, we propose and investigate a new measurement, Significance, that has the potential to improve the detection accuracy of association rule mining techniques. Our preliminary investigation on four open-source subject systems implies that our proposed measurement is capable of extracting coupling relationships even from infrequently co-changed entity sets that might seem insignificant while considering only Support and Confidence. Our proposed measurement, Significance (in association with Support and Confidence), has the potential to predict logical coupling with higher precision and recall. Manishankar Mondal, Chanchal Kumar Roy, Kevin A. Schneider |
ICPC | 3 |
| 2013 | SimCad: An extensible and faster clone detection tool for large scale software systemsabstractCode cloning is an inevitable phenomenon in evolution of software systems. To reduce the harmful effects of clones in software evolution, they need to be identified correctly as well in a time efficient way. There might be various types of clones in a software system. Earlier research shows detection of near-miss clones in large datasets appears to be costly in terms of time and memory. Among the clone detection tools available in practice, not very many of them are found effective in that regard. In this paper we present a standalone clone detection tool SimCad. It is based on a highly scalable and faster clone detection algorithm designed to detect both exact and near-miss clones in large-scale software systems. One of the potential aspects of SimCad is that its clone detection function is made more portable by packaging it into a library called SimLib. Thus, SimLib now can be used as an off-the-shelf clone detection library that can be easily integrated into other applications that are designed to work based on detected clones. For example, a standalone tool or an Integrated Development Environment (IDE) plugin can use SimLib for realtime clone detection while providing its own services like clone visualization and/or clone management functionalities. We hope that both researchers and developers would enjoy and utilize the benefit of using these tools in different aspects of detection and management of clones in software. Md. Sharif Uddin, Chanchal Kumar Roy, Kevin A. Schneider |
ICPC | 3 |
| 2013 | Answering questions about unanswered questions of stack overflowabstractCommunity-based question answering services accumulate large volumes of knowledge through the voluntary services of people across the globe. Stack Overflow is an example of such a service that targets developers and software engineers. In general, questions in Stack Overflow are answered in a very short time. However, we found that the number of unanswered questions has increased significantly in the past two years. Understanding why questions remain unanswered can help information seekers improve the quality of their questions, increase their chances of getting answers, and better decide when to use Stack Overflow services. In this paper, we mine data on unanswered questions from Stack Overflow. We then conduct a qualitative study to categorize unanswered questions, which reveals characteristics that would be difficult to find otherwise. Finally, we conduct an experiment to determine whether we can predict how long a question will remain unanswered in Stack Overflow. Muhammad Asaduzzaman, Ahmed Shah Mashiyat, Chanchal Kumar Roy, Kevin A. Schneider |
MSR | 4 |
| 2013 | Understanding the evolution of type-3 clones: an exploratory studyabstractUnderstanding the evolution of clones is important both for understanding the maintenance implications of clones and building a robust clone management system. To this end, researchers have already conducted a number of studies to analyze the evolution of clones, mostly focusing on Type-1 and Type-2 clones. However, although there are a significant number of Type-3 clones in software systems, we know a little how they actually evolve. In this paper, we perform an exploratory study on the evolution of Type-1, Type-2, and Type-3 clones in six open source software systems written in two different programming languages and compare the result with a previous study to better understand the evolution of Type-3 clones. Our results show that although Type-3 clones are more likely to change inconsistently, the absolute number of consistently changed Type-3 clone classes is higher than that of Type-1 and Type-2. Type-3 clone classes also have a lifespan similar to that of Type-1 and Type-2 clones. In addition, a considerable number of Type-1 and Type-2 clones convert into Type-3 clones during evolution. Therefore, it is important to manage type-3 clones properly to limit their negative impact. However, various automated clone management techniques such as notifying developers about clone changes or linked editing should be chosen carefully due to the inconsistent nature of Type-3 clones. Ripon K. Saha, Chanchal Kumar Roy, Kevin A. Schneider, Dewayne E. Perry |
MSR | 3 |
| 2013 | A discriminative model approach for suggesting tags automatically for stack overflow questionsabstractAnnotating documents with keywords or `tags' is useful for categorizing documents and helping users find a document efficiently and quickly. Question and answer (Q&A) sites also use tags to categorize questions to help ensure that their users are aware of questions related to their areas of expertise or interest. However, someone asking a question may not necessarily know the best way to categorize or tag the question, and automatically tagging or categorizing a question is a challenging task. Since a Q&A site may host millions of questions with tags and other data, this information can be used as a training and test dataset for approaches that automatically suggest tags for new questions. In this paper, we mine data from millions of questions from the Q&A site Stack Overflow, and using a discriminative model approach, we automatically suggest question tags to help a questioner choose appropriate tags for eliciting a response. Avigit K. Saha, Ripon K. Saha, Kevin A. Schneider |
MSR | 3 |
| 2012 | Bug introducing changes: A case study with AndroidabstractChanges, a rather inevitable part of software development can cause maintenance implications if they introduce bugs into the system. By isolating and characterizing these bug introducing changes it is possible to uncover potential risky source code entities or issues that produce bugs. In this paper, we mine the bug introducing changes in the Android platform by mapping bug reports to the changes that introduced the bugs. We then use the change information to look for both potential problematic parts and dynamics in development that can cause maintenance implications. We believe that the results of our study can help better manage Android software development. Muhammad Asaduzzaman, Michael C. Bullock, Chanchal Kumar Roy, Kevin A. Schneider |
MSR | 4 |
| 2011 | An automatic framework for extracting and classifying near-miss clone genealogiesabstractExtracting code clone genealogies across multiple versions of a program and classifying them according to their change patterns underlies the study of code clone evolution. While there are a few studies in the area, the approaches do not handle near-miss clones well and the associated tools are often computationally expensive. To address these limitations, we present a framework for automatically extracting both exact and near-miss clone genealogies across multiple versions of a program and for identifying their change patterns using a few key similarity factors. We have developed a prototype clone genealogy extractor, applied it to three open source projects including the Linux Kernel, and evaluated its accuracy in terms of precision and recall. Our experience shows that the prototype is scalable, adaptable to different clone detection tools, and can automatically identify evolution patterns of both exact and near-miss clones by constructing their genealogies. Ripon K. Saha, Chanchal Kumar Roy, Kevin A. Schneider |
ICSM | 3 |
| 2011 | An Empirical Study of the Impacts of Clones in Software MaintenanceabstractThe impacts of clones on software maintenance is a long-lived debate on whether clones are beneficial or not. Some researchers argue that clones lead to additional changes during the maintenance phase and thus increase the overall maintenance effort. Moreover, they note that inconsistent changes to clones may introduce faults during evolution. On the other hand, other researchers argue that cloned code exhibits more stability than non-cloned code. Studies resulting in such contradictory outcomes may be a consequence of using different methodologies, using different clone detection tools, defining different impact assessment metrics, and evaluating different subject systems. In order to understand the conflicting results from the studies, we plan to conduct a comprehensive empirical study using a common framework incorporating nine existing methods that yielded mostly contradictory findings. Our research strategy involves implementing each of these methods using four clone detection tools and evaluating the methods on more than fifteen subject systems of different languages and of a diverse nature. We believe that our study will help eliminate tool and study biases to resolve conflicts regarding the impacts of clones on software maintenance. Manishankar Mondal, Md. Saidur Rahman 0002, Ripon K. Saha, Chanchal Kumar Roy, Jens Krinke, Kevin A. Schneider |
ICPC | 6 |
| 2010 | An Introduction to Experience RequirementsabstractWe consider the application of requirements engineering principles and techniques to the elicitation, capture, and representation of the output of the user experience design process. A stimulus-perception-response model is used to motivate experience requirements, defined as descriptions of user experiences that must be met (functional experiences) or are satisfaction goals (non-functional experiences). We identify potential benefits and look at experience requirements in video games. David Callele, Eric Neufeld, Kevin A. Schneider |
RE | 3 |
| 2010 | Evaluating Code Clone Genealogies at Release Level: An Empirical StudyabstractCode clone genealogies show how clone groups evolve with the evolution of the associated software system, and thus could provide important insights on the maintenance implications of clones. In this paper, we provide an in-depth empirical study for evaluating clone genealogies in evolving open source systems at the release level. We develop a clone genealogy extractor, examine 17 open source C, Java, C++ and C# systems of diverse varieties and study different dimensions of how clone groups evolve with the evolution of the software systems. Our study shows that majority of the clone groups of the clone genealogies either propagate without any syntactic changes or change consistently in the subsequent releases, and that many of the genealogies remain alive during the evolution. These findings seem to be consistent with the findings of a previous study that clones may not be as detrimental in software maintenance as believed to be (at least by many of us), and that instead of aggressively refactoring clones, we should possibly focus on tracking and managing clones during the evolution of software systems. Ripon K. Saha, Muhammad Asaduzzaman, Minhaz Fahim Zibran, Chanchal Kumar Roy, Kevin A. Schneider |
SCAM | 5 |
| 2009 | UI traces: Supporting the maintenance of interactive softwareabstractWe propose a method to support the maintenance of interactive software systems with user interface traces, that involves: (1) collecting execution traces of an interactive system, (2) segmenting execution traces into user interface traces according to user interface activity, and (3) mapping the user interface activity to the implementation activity. To support our approach, we developed a tool that uses aspect-oriented programming and load-time weaving to collect user interface traces from an interactive system. The tool allows us to browse the user interface traces and view user interface related data such as: user input, display updates, and thread activity. Using our tool, we demonstrate how developers can orient themselves and identify the slice of code relevant to performing common software maintenance tasks. Andrew Sutherland, Kevin A. Schneider |
ICSM | 2 |
| 2009 | Augmenting Emotional Requirements with Emotion Markers and Emotion PrototypesabstractA production-phase weakness in emotional requirements was identified and resolved during a follow-up study. The definition of emotional requirements was extended to include emotion prototypes and emotion markers. Improved practices for identifying media assets for emotional requirements were developed, enhancing their utility to the production process. David Callele, Eric Neufeld, Kevin A. Schneider |
RE | 3 |
| 2008 | Balancing Security Requirements and Emotional Requirements in Video GamesabstractA fundamental conflict exists between designers, players, and cheaters: who has control over how the game is played? Resolving this conflict, by balancing the associated emotional and security requirements is challenging. Emotional requirements can assist the development of security requirements and to prioritize their development. Failure to meet the playerpsilas emotional requirements can lead to market forces that override security requirements. We suggest that in-game justice systems would allow the players to act as a self-correcting mechanism for emotional requirement failures that lead to cheating or other threats to the integrity of the game experience. Further investigation into this form of just-in-time requirements negotiation is ongoing. David Callele, Eric Neufeld, Kevin A. Schneider |
RE | 3 |
| 2006 | Emotional Requirements in Video GamesabstractRequirements engineering for video games must address a wide range of functional and non-functional requirements. Video game designers are most concerned with capturing and representing the player experience: the means by which the player's consciousness is cognitively engaged while simultaneously inducing emotional responses. We show that emotional requirements can be expressed in two parts: as the emotional intent of the designer and the means by which the designer expects to induce the target emotional state. Spatial and temporal qualifiers on intent and means may also be required. We introduce emotional terrain maps, emotional intensity maps, and emotion time lines as visual mechanisms for capturing and expressing emotional requirements. Using a first-person shooter example, we show that these mechanisms can express the desired emotional requirements while providing support for spatial and temporal qualifiers David Callele, Eric Neufeld, Kevin A. Schneider |
RE | 3 |
| 2005 | Requirements Engineering and the Creative Process in the Video Game IndustryabstractThe software engineering process in video game development is not clearly understood, hindering the development of reliable practices and processes for this field. An investigation of factors leading to success or failure in video game development suggests that many failures can be traced to problems with the transition from preproduction to production. Three examples, drawn from real video games, illustrate specific problems: 1) how to transform documentation from its preproduction form to a form that can be used as a basis for production;, 2) how to identify implied information in preproduction documents; and 3) how to apply domain knowledge without hindering the creative process. We identify 3 levels of implication and show that there is a strong correlation between experience and the ability to identify issues at each level. The accumulated evidence clearly identifies the need to extend traditional requirements engineering techniques to support the creative process in video game development. David Callele, Eric Neufeld, Kevin A. Schneider |
RE | 3 |
| 2004 | Group awareness in distributed software developmentabstractOpen-source software development projects are almost always collaborative and distributed. Despite the difficulties imposed by distance, these projects have managed to produce large, complex, and successful systems. However, there is still little known about how open-source teams manage their collaboration. In this paper we look at one aspect of this issue: how distributed developers maintain group awareness. We interviewed developers, read project communication, and looked at project artifacts from three successful open source projects. We found that distributed developers do need to maintain awareness of one another, and that they maintain both a general awareness of the entire team and more detailed knowledge of people that they plan to work with. Although there are several sources of information, this awareness is maintained primarily through text-based communication (mailing lists and chat systems). These textual channels have several characteristics that help to support the maintenance of awareness, as long as developers are committed to reading the lists and to making their project communication public. Carl Gutwin, Reagan Penner, Kevin A. Schneider |
CSCW | 3 |
| 2003 | Agile Parsing in TXL
Thomas R. Dean, James R. Cordy, Andrew J. Malton, Kevin A. Schneider |
Autom. Softw. Eng. | 4 |
| 2002 | Source transformation in software engineering using the TXL transformation system
James R. Cordy, Thomas R. Dean, Andrew J. Malton, Kevin A. Schneider |
Inf. Softw. Technol. | 4 |
| 2001 | Using Design Recovery Techniques to Transform Legacy SystemsabstractThe year 2000 problem posed a difficult problem for many IT shops world wide. The most difficult part of the problem was not the actual changes to ensure compliance, but finding and classifying the data fields that represent dates. This is a problem well suited to design recovery. The paper presents an overview of LS/2000, a system that used design recovery to analyze source code for year 2000 risks and guide a source transformation that was able to automatically remediate over 99% of the year 2000 risks in over three billion lines of production IT source. Thomas R. Dean, James R. Cordy, Kevin A. Schneider, Andrew J. Malton |
ICSM | 3 |