Muhammad Asaduzzaman

dblp:88/9972 · DBLP profile ↗
← Back
32ranked-venue papers
12as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 12 first-author · 11 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SLICEFORMER: Static Program Slicing Using Language Models With Dataflow-Aware Pretraining and Constrained Decoding
abstract
Static program slicing is a fundamental software engineering technique for isolating code relevant to specific variables.While recent learning-based approaches using language models (LMs) show promise in automating slice prediction, they suffer from inaccurate dependency modeling and unconstrained generation, where LMs fail to capture precise data flow relations and produce slices containing hallucinated tokens and statements.To address these challenges, we propose SLICEFORMER, a novel approach that reformulates static program slicing as a sequence-to-sequence task using small language models such as CodeT5+.SLICEFORMER introduces two key innovations that directly target the identified limitations.First, to improve dependency modeling, we design dataflow-aware pretraining objectives that leverage data flow graphs (DFG) to teach models data dependencies through dataflowpreserving statement permutation and dataflowaware span corruption.Second, to eliminate hallucination, we develop a constrained decoding mechanism that enforces both lexical and syntactic constraints.We evaluate SLICE-FORMER on Java and Python program slicing benchmarks, demonstrating consistent improvements over state-of-the-art baselines with up to 22% gain in ExactMatch.
Shaowei Wang 0002, Tse-Hsun (Peter) Chen, Muhammad Asaduzzaman
ACL (1)4
2026 Typify: A Lightweight Usage-driven Static Analyzer for Precise Python Type Inference
Ali Aman, Muhammad Asaduzzaman, Shaowei Wang 0002
ICPC2
2026 How Do Agentic AI Systems Address Performance Optimizations? A BERTopic-Based Analysis of Pull Requests
Md. Nahidul Islam Opu, Shahidul Islam, Muhammad Asaduzzaman, Shaiful Alam Chowdhury
MSR3
2025 An Analysis of Task Offloading Approaches in Edge-Cloud Continuum
abstract
Edge computing has emerged as a vital paradigm for supporting latency-sensitive and computation-intensive applications by offloading tasks closer to end-user devices. In this work, we investigate the challenges of workload orchestration in dynamic edge environments characterized by heterogeneous devices, intermittent traffic, and fluctuating network conditions. We compare five offloading policies in different scenarios. Our evaluation encompasses a series of simulation experiments under varying configurations, including moderate device loads (200-1000 IoT devices) and extreme conditions with high device counts (2000+ IoT devices). The results highlight context-dependent policy performance: the FUZZY_BASED policy delivers service times below 2 seconds for up to 1000 devices, whereas the NETWORK_BASED policy outperforms others in high-bandwidth overload scenarios, achieving 2.5-second service times at 2000 devices compared to 3.1 seconds for FUZZY_BASED. The UTILIZATION_BASED strategy consistently underperforms, showing severe degradation (4.0s service time) in congested conditions. These findings reveal the complex trade-offs in edge orchestration, where network-awareness becomes crucial in bandwidth-abundant congestion scenarios, while fuzzy logic’s adaptability proves more effective in typical dynamic environments. Therefore, an intelligent AI-based decision-making policy should be employed to optimize performance across varying conditions.
Jenish Modi, A. B. M. Bodrul Alam, Muhammad Asaduzzaman
COMPSAC3
2025 Evidence is All We Need: Do Self-Admitted Technical Debts Impact Method-Level Maintenance?
abstract
Self-Admitted Technical Debt (SATD) refers to the phenomenon where developers explicitly acknowledge technical debt through comments in the source code. While considerable research has focused on detecting and addressing SATD, its true impact on software maintenance remains underexplored. The few studies that have examined this critical aspect have not provided concrete evidence linking SATD to negative effects on software maintenance. These studies, however, focused only on file- or class-level code granularity. This paper aims to empirically investigate the influence of SATD on various facets of software maintenance at the method level. We assess SATD’s effects on code quality, bug susceptibility, change frequency, and the time practitioners typically take to resolve SATD.By analyzing a dataset of 774,051 methods from 49 opensource projects, we discovered that methods containing SATD are not only larger and more complex but also exhibit lower readability and a higher tendency for bugs and changes. We also found that SATD often remains unresolved for extended periods, adversely affecting code quality and maintainability. Our results provide empirical evidence highlighting the necessity of early identification, resource allocation, and proactive management of SATD to mitigate its long-term impacts on software quality and maintenance costs.
Shaiful Alam Chowdhury, Hisham Kidwai, Muhammad Asaduzzaman
MSR3
2025 Understanding the Popularity of Packages in Maven Ecosystem
abstract
The widespread availability of open-source software packages in ecosystems like Maven has significantly improved developer productivity by promoting the reuse of pre-existing packages. However, the vast number of available packages often poses challenges in selecting suitable packages. This study investigates the role of popularity metrics in evaluating Maven packages by analyzing 103,315 packages, each at least two years old. Metrics were collected from the Maven Neo4j dataset and GitHub repositories to examine their relationships and importance in determining package popularity. Our analysis reveals strong interdependencies among community-driven GitHub metrics, such as stars, forks, pull requests, and contributors, which highlight their role in defining package popularity. Conversely, Maven-specific metrics, including dependencies and vulnerabilities, showed weak correlations with GitHub-based popularity indicators. Our analysis identified license status, commits count, presence of README files, and usages as the most significant predictors of package popularity, while vulnerabilities had limited statistical impact. These findings underscore the complementary nature of technical and community-driven metrics in assessing package popularity and provide actionable insights for developers and researchers to better evaluate and select open-source software packages.
Sadman Jashim Sakib, Muhammad Asaduzzaman, Curtis Bright, Cole Morgan
MSR2
2025 Dependency Dilemmas: A Comparative Study of Independent and Dependent Artifacts in Maven Central Ecosystem
abstract
Maven Central ecosystem forms the backbone of Java dependency management, hosting artifacts that vary significantly in their adoption, security, and ecosystem roles. Artifact reuse is fundamental in software development, and ecosystems like Maven facilitate this process. However, prior studies predominantly analyzed popular artifacts with numerous dependencies, leaving those without incoming dependencies (i.e., independent artifacts) unexplored. In this study, we analyzed 658,078 artifacts, of which 635,003 had at least one release. Among these, 93,101 artifacts (15.4%) were identified as independent (in-degree = 0), while the rest were classified as dependent. We looked at the impact of individual artifacts using PageRank and outdegree centrality and discovered that independent artifacts were very important to the ecosystem. Further analysis using 18 different metrics revealed several advantages and comparability of independent artifacts with dependent artifacts: comparable popularity ($\mathbf{2 5. 5 8}$ vs. 7.30), fewer vulnerabilities ($\mathbf{6 0}$ CVEs vs. 179 CVEs), and zero propagated vulnerabilities. These findings suggest that independent artifacts might be a beneficial choice for dependencies but have some maintenance issues. Therefore, developers should carefully incorporate independent artifacts into their projects, and artifact maintainers should prioritize this group of artifacts to mitigate the risk of transitive vulnerability propagation and improve software sustainability.
Mehedi Hasan Shanto, Muhammad Asaduzzaman, Manishankar Mondal, Shaiful Alam Chowdhury
MSR2
2025 ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets Using LLMs
abstract
Technical Q&A sites are valuable for software developers seeking knowledge, but the code snippets they provide are often uncompilable and incomplete due to unresolved types and missing libraries. This poses a challenge for users who wish to reuse or analyze these snippets. Existing methods either do not focus on creating compilable code or have low success rates. To address this, we propose ZS4C, a lightweight approach for zero-shot synthesis of compilable code from incomplete snippets using Large Language Models (LLMs). ZS4C operates in two stages: first, it uses an LLM, like GPT-3.5, to identify missing import statements in a snippet; second, it collaborates with a validator (e.g., compiler) to fix compilation errors caused by incorrect imports and syntax issues. We evaluated ZS4C on the StatType-SO benchmark and a new dataset, Python-SO, which includes 539 Python snippets from Stack Overflow across the 20 most popular Python libraries. ZS4C significantly outperforms existing methods, improving the compilation rate from 63% to 95.1% compared to the state-of-the-art SnR, marking a 50.1% improvement. On average, ZS4C can infer more accurate import statements (with an F1 score of 0.98) than SnR, with an improvement of 8.5% in the F1.
Azmain Kabir, Shaowei Wang 0002, Yuan Tian 0008, Tse-Hsun (Peter) Chen, Muhammad Asaduzzaman, Wenbin Zhang 0002
ACM Trans. Softw. Eng. Methodol.5
2024 On the Executability of R Markdown Files
abstract
R Markdown files are examples of literate programming documents that combine R code with results and explanations. Such dynamic documents are designed to execute easily and reproduce study results. However, little is known about the executability of R Markdown files which can cause frustration among its users who intend to reuse the document. This paper presents a large-scale study on the executability of R Markdown files collected from GitHub. Results from our study show that a significant number of R Markdown files (64.95%) are not executable, even after our best efforts. To better understand the challenges, we categorize the exceptions encountered while executing the documents into different categories. Finally, we develop a classifier to determine which Markdown files are likely to be executable. Such a classifier can be utilized by search engines in their ranking which helps developers to find literate programming documents as learning resources.
Md. Anaytul Islam, Muhammad Asaduzzaman, Shaowei Wang 0002
MSR2
2023 Selective EEG Signal Anonymization using Multi-Objective Autoencoders
abstract
The availability of low-cost brain-computer interfaces and related software has enabled application developers to access this technology seamlessly. However, unsupervised access to users' brain signals raises alarms for EEG data privacy and identity protection. This paper explores a new direction to selectively anonymize a person's brain signals resulting from a response to a stimulus. These time-locked potentials containing sensitive user information are masked to allow for the intended task/event classification (brain task activity) while minimizing the accuracy of subject classification (brain identity activity).We study the feasibility of an autoencoder architecture, enveloped with regularizers and a multi-objective loss function, to achieve an optimal utility-privacy trade-off for EEG data application. We observed a drop of 35% in the accuracy of the subject's classification, while suffering only a loss of 14% in the accuracy of the task classification. This algorithm can be applied to multi-channel and multi-subject scenarios, and our results demonstrate a proof-of-concept that we can generalize an anonymizing autoencoder architecture to be applicable to intricate stochastic data such as EEG.
Girijesh Singh, Palak Patel, Muhammad Asaduzzaman, Garima Bajwa
PST3
2023 Finding associations between natural and computer languages: A case-study of bilingual LDA applied to the bleeping computer forum posts
Kundi Yao, Gustavo Ansaldi Oliva, Ahmed E. Hassan, Muhammad Asaduzzaman, Andrew J. Malton, Andrew Walenstein
J. Syst. Softw.4
2022 Mining Software Information Sites to Recommend Cross-Language Analogical Libraries
abstract
Software development is largely dependent on libraries to reuse existing functionalities instead of reinventing the wheel. Software developers often need to find analogical libraries (libraries similar to ones they are already familiar with) as an analogical library may offer improved or additional features. Developers also need to search for analogical libraries across programming languages when developing applications in different languages or for different platforms. However, manually searching for analogical libraries is a time-consuming and difficult task. This paper presents a technique, called XLibRec, that recommends analogical libraries across different programming languages. XLibRec collects Stack Overflow question titles containing library names, library usage information from Stack Overflow posts, and library descriptions from a third party website, Libraries.io. We generate word-vectors for each information and calculate a weight-based cosine similarity score from them to recommend analogical libraries. We performed an extensive evaluation using a large number of analogical libraries across four different programming languages. Results from our evaluation show that the proposed technique can recommend cross-language analogical libraries with great accuracy. The precision for the Top-3 recommendations ranges from 62-81% and has achieved 8-45% higher precision than the state-of-the-art technique.
Kawser Wazed Nafi, Muhammad Asaduzzaman, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider
SANER2
2022 A Study of Bug Management Using the Stack Exchange Question and Answering Platform
abstract
Traditional bug management systems, like Bugzilla, are widely used in open source and commercial projects. Stack Exchange uses its online question and answer (Q&A) platform to collect and manage bugs, which brings several new unique features that are not offered in traditional bug management systems. Users can edit bug reports, use different communication channels, and vote on bug reports, answers, and their associated comments. Understanding how these features manage bug reports can provide insights to the designers of traditional bug management systems, like whether a feature should be introduced? and how would users leverage such a feature? We performed a large-scale analysis of 19,151 bug reports of the bug management system of Stack Exchange and studied the in-place editing, the answering and commenting, and the voting features. We find that: 1) The three features are used actively. 2) 57 percent of the edits improved the quality of bug reports. 3) Commenting provides a channel for discussing bug-related information, while answering offers a channel for explaining the causes of a bug and bug-fix information. 4) Downvotes are made due to the disagreement of the reported “bug” being a real bug and the low quality of bug reports. Based on our findings, we provide suggestions for traditional bug management systems.
Aaditya Bhatia, Shaowei Wang 0002, Muhammad Asaduzzaman, Ahmed E. Hassan
IEEE Trans. Software Eng.3
2020 Exploring Type Inference Techniques of Dynamically Typed Languages
abstract
Developers often prefer dynamically typed programming languages, such as JavaScript, because such languages do not require explicit type declarations. However, such a feature hinders software engineering tasks, such as code completion, type related bug fixes and so on. Deep learning-based techniques are proposed in the literature to infer the types of code elements in JavaScript snippets. These techniques are computationally expensive. While several type inference techniques have been developed to detect types in code snippets written in statically typed languages, it is not clear how effective those techniques are for inferring types in dynamically typed languages, such as JavaScript. In this paper, we investigate the type inference techniques of JavaScript to understand the above two issues further. While doing that we propose a new technique that considers the locally specific code tokens as the context to infer the types of code elements. The evaluation result shows that the proposed technique is 20-47% more accurate than the statically typed language-based techniques and 5–14 times faster than the deep learning techniques without sacrificing accuracy. Our analysis of sensitivity, overlapping of predicted types and the number of training examples justify the importance of our technique.
C. M. Khaled Saifullah, Muhammad Asaduzzaman, Chanchal Kumar Roy
SANER2
2020 CAPS: a supervised technique for classifying Stack Overflow posts concerning API issues
Md. Ahasanuzzaman, Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider
Empir. Softw. Eng.2
2019 Learning from Examples to Find Fully Qualified Names of API Elements in Code Snippets
abstract
Developers often reuse code snippets from online forums, such as Stack Overflow, to learn API usages of software frameworks or libraries. These code snippets often contain ambiguous undeclared external references. Such external references make it difficult to learn and use those APIs correctly. In particular, reusing code snippets containing such ambiguous undeclared external references requires significant manual efforts and expertise to resolve them. Manually resolving fully qualified names (FQN) of API elements is a non-trivial task. In this paper, we propose a novel context-sensitive technique, called COSTER, to resolve FQNs of API elements in such code snippets. The proposed technique collects locally specific source code elements as well as globally related tokens as the context of FQNs, calculates likelihood scores, and builds an occurrence likelihood dictionary (OLD). Given an API element as a query, COSTER captures the context of the query API element, matches that with the FQNs of API elements stored in the OLD, and rank those matched FQNs leveraging three different scores: likelihood, context similarity, and name similarity scores. Evaluation with more than 600K code examples collected from GitHub and two different Stack Overflow datasets shows that our proposed technique improves precision by 4-6% and recall by 3-22% compared to state-of-the-art techniques. The proposed technique significantly reduces the training time compared to the StatType, a state-of-the-art technique, without sacrificing accuracy. Extensive analyses on results demonstrate the robustness of the proposed technique.
C. M. Khaled Saifullah, Muhammad Asaduzzaman, Chanchal Kumar Roy
ASE2
2018 Classifying stack overflow posts on API issues
abstract
The design and maintenance of APIs are complex tasks due to the constantly changing requirements of its users. Despite the efforts of its designers, APIs may suffer from a number of issues (such as incomplete or erroneous documentation, poor performance, and backward incompatibility). To maintain a healthy client base, API designers must learn these issues to fix them. Question answering sites, such as Stack Overflow (SO), has become a popular place for discussing API issues. These posts about API issues are invaluable to API designers, not only because they can help to learn more about the problem but also because they can facilitate learning the requirements of API users. However, the unstructured nature of posts and the abundance of non-issue posts make the task of detecting SO posts concerning API issues difficult and challenging. In this paper, we first develop a supervised learning approach using a Conditional Random Field (CRF), a statistical modeling method, to identify API issue-related sentences. We use the above information together with different features of posts and experience of users to build a technique, called CAPS, that can classify SO posts concerning API issues. Evaluation of CAPS using carefully curated SO posts on three popular API types reveals that the technique outperforms all three baseline approaches we consider in this study. We also conduct studies to test the generalizability of CAPS results and to understand the effects of different sources of information on it.
Muhammad Ahasanuzzaman, Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider
SANER2
2017 Recommending Framework Extension Examples
abstract
The use of software frameworks enables the delivery of common functionality but with significantly less effort than when developing from scratch. To meet application specific requirements, the behavior of a framework needs to be customized via extension points. A common way of customizing framework behavior is by passing a framework related object as an argument to an API call. Such an object can be created by subclassing an existing framework class or interface, or by directly customizing an existing framework object. However, to do this effectively requires developers to have extensive knowledge of the framework's extension points and their interactions. To aid the developers in this regard, we propose and evaluate a graph mining approach for extension point management. Specifically, we propose a taxonomy of extension patterns to categorize the various ways an extension point has been used in the code examples. Our approach mines a large amount of code examples to discover all extension points and patterns for each framework class. Given a framework class that is being used, our approach aids the developer by following a two-step recommendation process. First, it recommends all the extension points that are available in the class. Once the developer chooses an extension point, our approach then discovers all of its usage patterns and recommends the best code examples for each pattern. Using five frameworks, we evaluate the performance of our two-step recommendation, in terms of precision, recall, and F-measure. We also report several statistics related to framework extension points.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou
ICSME1
2017 FEMIR: a tool for recommending framework extension examples
abstract
Software frameworks enable developers to reuse existing well tested functionalities instead of taking the burden of implementing everything from scratch. However, to meet application specific requirements, the frameworks need to be customized via extension points. This is often done by passing a framework related object as an argument to an API call. To enable such customizations, the object can be created by extending a framework class, implementing an interface, or changing the properties of the object via API calls. However, it is both a common and non-trivial task to find all the details related to the customizations. In this paper, we present a tool, called FEMIR, that utilizes partial program analysis and graph mining technique to detect, group, and rank framework extension examples. The tool extends existing code completion infrastructure to inform developers about customization choices, enabling them to browse through extension points of a framework, and frequent usages of each point in terms of code examples. A video demo is made available at https://asaduzzamanparvez.wordpress.com/femir.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou
ASE1
2016 Mining duplicate questions in stack overflow
abstract
Stack Overflow is a popular question answering site that is focused on programming problems. Despite efforts to prevent asking questions that have already been answered, the site contains duplicate questions. This may cause developers to unnecessarily wait for a question to be answered when it has already been asked and answered. The site currently depends on its moderators and users with high reputation to manually mark those questions as duplicates, which not only results in delayed responses but also requires additional efforts. In this paper, we first perform a manual investigation to understand why users submit duplicate questions in Stack Overflow. Based on our manual investigation we propose a classification technique that uses a number of carefully chosen features to identify duplicate questions. Evaluation using a large number of questions shows that our technique can detect duplicate questions with reasonable accuracy. We also compare our technique with DupPredictor, a state-of-the-art technique for detecting duplicate questions, and we found that our proposed technique has a better recall-rate than that technique.
Muhammad Ahasanuzzaman, Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider
MSR2
2016 How developers use exception handling in Java?
abstract
Exception handling is a technique that addresses exceptional conditions in applications, allowing the normal flow of execution to continue in the event of an exception and/or to report on such events. Although exception handling techniques, features and bad coding practices have been discussed both in developer communities and in the literature, there is a marked lack of empirical evidence on how developers use exception handling in practice. In this paper we use the Boa language and infrastructure to analyze 274k open source Java projects in GitHub to discover how developers use exception handling. We not only consider various exception handling features but also explore bad coding practices and their relation to the experience of developers. Our results provide some interesting insights. For example, we found that bad exception handling coding practices are common in open source Java projects and regardless of experience all developers use bad exception handling coding practices.
Muhammad Asaduzzaman, Muhammad Ahasanuzzaman, Chanchal Kumar Roy, Kevin A. Schneider
MSR1
2016 A Simple, Efficient, Context-sensitive Approach for Code Completion
abstract
Abstract Code completion helps developers use application programming interfaces (APIs) and frees them from remembering every detail. In this paper, we first describe a novel technique called Context‐sensitive Code Completion (CSCC) for improving the performance of API method call completion. CSCC is context sensitive in that it uses new sources of information as the context of a target method call. CSCC indexes method calls in code examples by their context. To recommend completion proposals, CSCC ranks candidate methods by the similarities between their contexts and the context of the target call. Evaluation using a set of subject systems and five popular state‐of‐the‐art techniques suggests that CSCC performs better than existing type or example‐based code completion systems. We conduct experiments to find how different contextual elements of the target call benefit CSCC. Next, we investigate the adaptability of the technique to support another form of code completion, i.e., field completion. Evaluation with eight different subject systems suggests that CSCC can easily support field completion with high accuracy. Finally, we compare CSCC with four popular statistical language models that support code completion. Results of statistical tests from our study suggest that CSCC not only outperforms those techniques that are based on token level language models, but also in most cases performs better or equally well with GraLan, the state‐of‐the‐art graph‐based language model. Copyright © 2016 John Wiley & Sons, Ltd.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou
J. Softw. Evol. Process.1
2015 Exploring API method parameter recommendations
abstract
A number of techniques have been developed that support method call completion. However, there has been little research on the problem of method parameter completion. In this paper, we first present a study that helps us to understand how developers complete method parameters. Based on our observations, we developed a recommendation technique, called Parc, that collects parameter usage context using a source code localness property that suggests that developers tend to collocate related code fragments. Parc uses previous code examples together with contextual and static type analysis to recommend method parameters. Evaluating our technique against the only available state-of-the-art tool using a number of subject systems and different Java libraries shows that our approach has potential. We also explore the parameter recommendation support provided by the Eclipse Java Development Tools (JDT). Finally, we discuss limitations of our proposed technique and outline future research directions.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Samiul Monir, Kevin A. Schneider
ICSME1
2015 PARC: Recommending API methods parameters
abstract
APIs have grown considerably in size. To free developers from remembering every detail of an API, code completion has become an integral part of modern IDEs. Most work on code completion targets completing API method calls and leaves the task of completing method parameters to the developers. However, parameter completion is also a non-trivial task. We present an Eclipse plugin, called PARC, that supports automatic completion of API method parameters. The tool is based on the localness property of source code, which states that developers tend to put related code fragments close together. PARC combines contextual and static type analysis to support a wide range of parameter expression types.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider
ICSME1
2014 CSCC: Simple, Efficient, Context Sensitive Code Completion
abstract
Code Completion helps developers learn APIs and frees them from remembering every detail. In this paper, we describe a novel technique called CSCC (Context Sensitive Code Completion) for improving the performance of API method call completion. CSCC is context sensitive in that it uses new sources of information as the context of a target method call. CSCC indexes method calls in code examples by their contexts. To recommend completion proposals, CSCC ranks candidate methods by the similarities between their contexts and the context of the target call. Evaluation using a set of subject systems and five popular state of-the-art techniques suggests that CSCC performs better than existing type or example-based code completion systems. We also investigate how the different contextual elements of the target call benefit CSCC.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou
ICSME1
2014 Context-Sensitive Code Completion Tool for Better API Usability
abstract
Developers depend on APIs of frameworks and libraries to support the development process. Due to the large number of existing APIs, it is difficult to learn, remember, and use them during the development of a software. To mitigate the problem, modern integrated development environments provide code completion facilities that free developers from remembering every detail. In this paper, we introduce CSCC, a simple, efficient context-sensitive code completion tool that leverages previous code examples to support method completion. Compared to other existing code completion tools, CSCC uses new sources of contextual information together with lightweight source code analysis to better recommend API method calls.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Daqing Hou
ICSME1
2013 LHDiff: A Language-Independent Hybrid Approach for Tracking Source Code Lines
abstract
Tracking source code lines between two different versions of a file is a fundamental step for solving a number of important problems in software maintenance such as locating bug introducing changes, tracking code fragments or defects across versions, merging file versions, and software evolution analysis. Although a number of such approaches are available in the literature, their performance is sensitive to the kind and degree of source code changes. There is also a marked lack of study on the effect of change types on source location tracking techniques. In this paper, we propose a language-independent technique, LHDiff, for tracking source code lines across versions that leverages simhash technique together with heuristics to improve accuracy. We evaluate our approach against state-of-the- art techniques using benchmarks containing different degrees of changes where files are selected from real world applications. We further evaluate LHDiff with other techniques using a mutation based analysis to understand how different types of changes affect their performance. The results reveal that our technique is more effective than language-independent approaches and no worse than some language-dependent techniques. In our study LHDiff even shows better performance than a state-of-the-art language- dependent approach. In addition, we also discuss limitations of different line tracking techniques including ours and propose future research directions.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Massimiliano Di Penta
ICSM1
2013 LHDiff: Tracking Source Code Lines to Support Software Maintenance Activities
abstract
Tracking lines across versions of a file is a necessary step for solving a number of problems during software development and maintenance. Examples include, but are not limited to, locating bug-inducing changes, tracking code fragments or vulnerable instructions across versions, co-change analysis, merging file versions, reviewing source code changes, and software evolution analysis. In this tool demonstration, we present a language-independent line-level location tracker, named LHDiff, that can be used to track lines and analyze changes in various kinds of software artifacts, ranging from source code to arbitrary text files. The tool can effectively detect changed or moved lines across versions of a file, has the ability to detect line splits, and can easily be integrated with existing version control systems. It overcomes the limitations of existing language-independent techniques and is even comparable to tools that are language dependent. In addition to describing the tool, we also describe its effectiveness in analyzing source code artifacts.
Muhammad Asaduzzaman, Chanchal Kumar Roy, Kevin A. Schneider, Massimiliano Di Penta
ICSM1
2013 Answering questions about unanswered questions of stack overflow
abstract
Community-based question answering services accumulate large volumes of knowledge through the voluntary services of people across the globe. Stack Overflow is an example of such a service that targets developers and software engineers. In general, questions in Stack Overflow are answered in a very short time. However, we found that the number of unanswered questions has increased significantly in the past two years. Understanding why questions remain unanswered can help information seekers improve the quality of their questions, increase their chances of getting answers, and better decide when to use Stack Overflow services. In this paper, we mine data on unanswered questions from Stack Overflow. We then conduct a qualitative study to categorize unanswered questions, which reveals characteristics that would be difficult to find otherwise. Finally, we conduct an experiment to determine whether we can predict how long a question will remain unanswered in Stack Overflow.
Muhammad Asaduzzaman, Ahmed Shah Mashiyat, Chanchal Kumar Roy, Kevin A. Schneider
MSR1
2012 Bug introducing changes: A case study with Android
abstract
Changes, a rather inevitable part of software development can cause maintenance implications if they introduce bugs into the system. By isolating and characterizing these bug introducing changes it is possible to uncover potential risky source code entities or issues that produce bugs. In this paper, we mine the bug introducing changes in the Android platform by mapping bug reports to the changes that introduced the bugs. We then use the change information to look for both potential problematic parts and dynamics in development that can cause maintenance implications. We believe that the results of our study can help better manage Android software development.
Muhammad Asaduzzaman, Michael C. Bullock, Chanchal Kumar Roy, Kevin A. Schneider
MSR1
2011 Analyzing and Forecasting Near-Miss Clones in Evolving Software: An Empirical Study
abstract
Effort for development and maintenance of complex large software is believed to have dependency on the amount of duplicated code fragments (code clones) present in code-bases. For example, clones need to be carefully and consistently maintained and/or refactored for preventing accidental error propagation. Thus it is important to understand the proportion and evolution of clones in evolving software systems for cost estimation or the like. This paper presents a study on the evolution of near-miss clones at release level in medium to large open source software systems of different types (operating systems, database systems, editors, etc.) written in three different programming languages namely C, C#, and Java. Using a hybrid clone detector, NiCad, we detected both exact and near-miss clones at different levels of similarity. Applying statistical methods we investigated, from different dimensions, the evolution of both exact and near-miss clones, and also forecasted the amount of clones in future releases of the software systems. Our study offers significant insights into the existence and evolution of code clones and their relationships with programming language or paradigm and program size.
Minhaz Fahim Zibran, Ripon K. Saha, Muhammad Asaduzzaman, Chanchal Kumar Roy
ICECCS3
2010 Evaluating Code Clone Genealogies at Release Level: An Empirical Study
abstract
Code clone genealogies show how clone groups evolve with the evolution of the associated software system, and thus could provide important insights on the maintenance implications of clones. In this paper, we provide an in-depth empirical study for evaluating clone genealogies in evolving open source systems at the release level. We develop a clone genealogy extractor, examine 17 open source C, Java, C++ and C# systems of diverse varieties and study different dimensions of how clone groups evolve with the evolution of the software systems. Our study shows that majority of the clone groups of the clone genealogies either propagate without any syntactic changes or change consistently in the subsequent releases, and that many of the genealogies remain alive during the evolution. These findings seem to be consistent with the findings of a previous study that clones may not be as detrimental in software maintenance as believed to be (at least by many of us), and that instead of aggressively refactoring clones, we should possibly focus on tracking and managing clones during the evolution of software systems.
Ripon K. Saha, Muhammad Asaduzzaman, Minhaz Fahim Zibran, Chanchal Kumar Roy, Kevin A. Schneider
SCAM2