EDBT 2026 Demo / reviewers in the wild / expert
Jafar M. Al-Kofahi
dblp:24/3120
· DBLP profile ↗
18ranked-venue papers
5as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 5 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
12 papers |
Software maintenance and evolution · 46% Empirical software engineering · 21% Program analysis · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering
mining software repositories |
0.6 | 3 | 2017 | Improving Automated Bug Triaging with Specialized Topic Model · IEEE Trans. Software Eng. 2017 Graph-based pattern-oriented, context-sensitive source code completion · ICSE 2012 Fuzzy set and cache-based approach for bug triaging · SIGSOFT FSE 2011 |
Software maintenance and evolution
bug triage |
0.5 | 3 | 2017 | Improving Automated Bug Triaging with Specialized Topic Model · IEEE Trans. Software Eng. 2017 Fuzzy set and cache-based approach for bug triaging · SIGSOFT FSE 2011 Fuzzy set-based automatic bug triaging · ICSE 2011 |
Software maintenance and evolution
code clone management |
0.3 | 3 | 2012 | Clone Management for Evolving Software · IEEE Trans. Software Eng. 2012 Clone-Aware Configuration Management · ASE 2009 Cleman: Comprehensive Clone Group Evolution Management · ASE 2008 |
Software maintenance and evolution › API usage
API usage pattern mining |
0.2 | 2 | 2012 | Graph-based pattern-oriented, context-sensitive source code completion · ICSE 2012 Graph-based mining of multiple object usage patterns · ESEC/SIGSOFT FSE 2009 |
Program analysis › symbolic execution
dynamic symbolic execution |
0.2 | 1 | 2015 | Poster: Static Detection of Configuration-Dependent Bugs in Configurable Software · ICSE (2) 2015 |
Program analysis
static analysis |
0.2 | 1 | 2015 | Poster: Static Detection of Configuration-Dependent Bugs in Configurable Software · ICSE (2) 2015 |
Program synthesis and code generation
code completion |
0.1 | 1 | 2012 | Graph-based pattern-oriented, context-sensitive source code completion · ICSE 2012 |
Software maintenance and evolution › bug triage
automatic bug triage |
0.1 | 1 | 2011 | Fuzzy set-based automatic bug triaging · ICSE 2011 |
Debugging and program repair › fault localization › bug localization
buggy file localization |
0.1 | 1 | 2011 | A topic-based approach for narrowing the search space of buggy files from a bug report · ASE 2011 |
Empirical software engineering › mining software repositories
bug report analysis |
0.1 | 1 | 2011 | Fuzzy set and cache-based approach for bug triaging · SIGSOFT FSE 2011 |
Debugging and program repair
fault localization |
0.1 | 1 | 2011 | A topic-based approach for narrowing the search space of buggy files from a bug report · ASE 2011 |
Debugging and program repair
automated program repair |
0.1 | 1 | 2010 | Recurring bug fixes in object-oriented programs · ICSE (1) 2010 |
Software maintenance and evolution › code clone detection
graph-based clone detection |
0.1 | 1 | 2009 | Complete and accurate clone detection in graph-based models · ICSE 2009 |
Software maintenance and evolution › code clone detection
model clone detection |
0.1 | 1 | 2009 | Complete and accurate clone detection in graph-based models · ICSE 2009 |
Requirements engineering and software design
model-driven engineering |
0.1 | 1 | 2009 | Complete and accurate clone detection in graph-based models · ICSE 2009 |
Data mining › text mining
text classification |
0.1 | 1 | 2017 | Improving Automated Bug Triaging with Specialized Topic Model · IEEE Trans. Software Eng. 2017 |
Software testing
test generation |
0.1 | 1 | 2015 | Poster: Static Detection of Configuration-Dependent Bugs in Configurable Software · ICSE (2) 2015 |
Data mining › text mining
topic model |
0.0 | 1 | 2011 | A topic-based approach for narrowing the search space of buggy files from a bug report · ASE 2011 |
Empirical software engineering › developer studies › developer expertise
developer expertise modeling |
0.0 | 1 | 2011 | Fuzzy set-based automatic bug triaging · ICSE 2011 |
Software maintenance and evolution
code clone detection |
0.0 | 1 | 2009 | Complete and accurate clone detection in graph-based models · ICSE 2009 |
Program analysis
dynamic analysis |
0.0 | 1 | 2009 | Graph-based mining of multiple object usage patterns · ESEC/SIGSOFT FSE 2009 |
Software maintenance and evolution
software configuration management |
0.0 | 1 | 2009 | Clone-Aware Configuration Management · ASE 2009 |
Methods — techniques the papers use, named apart from their topics
multi-feature topic model · 0.6latent dirichlet allocation · 0.6incremental learning · 0.6fuzzy set · 0.2dynamic symbolic execution · 0.2depth-first search · 0.2control flow graph · 0.2tree editing scripts · 0.1graph-based code completion · 0.1abstract syntax tree similarity · 0.1topic model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Four languages and lots of macros: analyzing autotools build systemsabstractBuild systems are crucial for software system development, however there is a lack of tool support to help with their high maintenance overhead. GNU Autotools are widely used in the open source community, but users face various challenges from its hard to comprehend nature and staging of multiple code generation steps, often leading to low quality and error-prone build code. In this paper, we present a platform, AutoHaven, to provide a foundation for developers to create analysis tools to help them understand, maintain, and migrate their GNU Autotools build systems. Internally it uses approximate parsing and symbolic analysis of the build logic. We illustrate the use of the platform with two tools: ACSense helps developers to better understand their build systems and ACSniff detects build smells to improve build code quality. Our evaluation shows that AutoHaven can support most GNU Autotools build systems and can detect build smells in the wild. Jafar M. Al-Kofahi, Suresh C. Kothari, Christian Kästner |
GPCE | 1 |
| 2017 | Improving Automated Bug Triaging with Specialized Topic ModelabstractBug triaging refers to the process of assigning a bug to the most appropriate developer to fix. It becomes more and more difficult and complicated as the size of software and the number of developers increase. In this paper, we propose a new framework for bug triaging, which maps the words in the bug reports (i.e., the term space) to their corresponding topics (i.e., the topic space). We propose a specialized topic modeling algorithm namedmulti-feature topic model (MTM)which extends Latent Dirichlet Allocation (LDA) for bug triaging.MTMconsiders product and component information of bug reports to map the term space to the topic space. Finally, we propose an incremental learning method namedTopicMinerwhich considers the topic distribution of a new bug report to assign an appropriate fixer based on the affinity of the fixer to the topics. We pairTopicMinerwith MTM (TopicMiner$^{MTM}$). We have evaluated our solution on 5 large bug report datasets including GCC, OpenOffice, Mozilla, Netbeans, and Eclipse containing a total of 227,278 bug reports. We show thatTopicMiner$^{MTM}$can achieve top-1 and top-5 prediction accuracies of 0.4831-0.6868, and 0.7686-0.9084, respectively. We also compareTopicMiner$^{MTM}$with Bugzie, LDA-KL, SVM-LDA, LDA-Activity, and Yang et al.'s approach. The results show thatTopicMiner$^{MTM}$on average improves top-1 and top-5 prediction accuracies of Bugzie by 128.48 and 53.22 percent, LDA-KL by 262.91 and 105.97 percent, SVM-LDA by 205.89 and 110.48 percent, LDA-Activity by 377.60 and 176.32 percent, and Yang et al.'s approach by 59.88 and 13.70 percent, respectively. Xin Xia 0001, David Lo 0001, Ying Ding 0005, Jafar M. Al-Kofahi, Tien N. Nguyen, Xinyu Wang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2015 | Poster: Static Detection of Configuration-Dependent Bugs in Configurable SoftwareabstractConfigurable software systems enable developers to configure at compile time a single variant of the system to tailor it towards specific environments and features. Although traditional static analysis tools can assist developers in software development and maintenance, they can only run on a concrete configuration of a configurable software system. Thus, it is necessary to derive many configurations so that the configuration-specific parts of the source code can be checked. To avoid this tedious and error-prone process, we propose an approach to automatically derive a set of configurations that cover as many combinations of configuration-specific blocks of code or source files as possible. We represent a C program with CPP directives (e.g., #ifdef) with a CPP control-flow graph (CPP-CFG) in which CPP expressions are condition nodes and #ifdef blocks are statement nodes. We then explore possible paths on CPP-CFG with dynamic symbolic execution and depth-first search algorithms, and correspondingly, producing possible combinations of concrete blocks of C code, on which an existing static analysis tool can run. Our preliminary evaluation on a benchmark of configuration-dependent bugs on Linux shows that our approach can detect more bugs than a state-of-the-art tool. Jafar M. Al-Kofahi, Lisong Guo, Hung Viet Nguyen, Hoan Anh Nguyen, Tien N. Nguyen |
ICSE (2) | 1 |
| 2014 | Fault Localization for Make-Based Build CrashesabstractIn large-scale software projects, build code has a high level of complexity, churn rate, and defect proneness. While it is desirable to have automated tools to help developers in localizing faults in build code, it is challenging to build such tools due to the dynamic nature of build code. Existing automatic fault localization methods focus on traditional code and none of them has such support for build code. This paper introduces MkFault, a novel automatic tool/method to localize faults in build code that cause run-time build failures. Given a test case that causes a run-time crash in the execution of a Make file, it returns a ranked list of statements in the Make file with their suspiciousness scores. MkFault records the evaluation traces from Make code to identify the corresponding concrete build rules and the execution traces of those rules. It then uses those traces and its novel Bayesian-like rating algorithm to give suspiciousness scores to the original statements in the Make file. Our empirical evaluation on real faults in several open-source projects has shown that MkFault can achieve high accuracy and help reduce a large percentage of the lines of code that developers need to examine. Jafar M. Al-Kofahi, Hung Viet Nguyen, Tien N. Nguyen |
ICSME | 1 |
| 2012 | Graph-based pattern-oriented, context-sensitive source code completionabstractCode completion helps improve developers' programming productivity. However, the current support for code completion is limited to context-free code templates or a single method call of the variable on focus. Using software libraries for development, developers often repeat API usages for certain tasks. Thus, a code completion tool could make use of API usage patterns. In this paper, we introduce GraPacc, a graph-based, pattern-oriented, context-sensitive code completion approach that is based on a database of such patterns. GraPacc represents and manages the API usage patterns of multiple variables, methods, and control structures via graph-based models. It extracts the context-sensitive features from the code under editing, e.g. the API elements on focus and their relations to other code elements. Those features are used to search and rank the patterns that are most fitted with the current code. When a pattern is selected, the current code will be completed via a novel graph-based code completion algorithm. Empirical evaluation on several real-world systems shows that GraPacc has a high level of accuracy in code completion. Anh Tuan Nguyen 0001, Tung Thanh Nguyen, Hoan Anh Nguyen, Ahmed Tamrawi, Hung Viet Nguyen, Jafar M. Al-Kofahi, Tien N. Nguyen |
ICSE | 6 |
| 2012 | Detecting semantic changes in Makefile build codeabstractBuild code in a Makefile represents the build rules with the dependencies among the files, and how they must be built together to produce a software system. As software evolves, its build code evolves as well to accommodate necessary changes in the build process. As part of software maintenance, it is crucial to understand how the build code is changed (e.g. changes in build rules or dependencies), and to verify and validate the correctness of the build process with different build configurations. Due to Make's dynamic nature, understanding and managing the changes to Makefiles is not trivial. In this paper, we introduce a set of semantic changes to build code in Makefiles. We also develop MkDiff, a tool to detect the changes to a Makefile at the semantic level. MkDiff uses symbolic dependency graphs (SDG) to find all possible concrete rules from a Makefile, and the dependencies among them. For two SDGs built from a Makefile at two versions, it first detects changed and unchanged nodes via its SDG matching algorithm. Then, from those results, it derives the semantic changes to the Makefile. Our empirical evaluation for MkDiff showed that it can accurately detect semantic changes in Makefiles. Jafar M. Al-Kofahi, Hung Viet Nguyen, Anh Tuan Nguyen 0001, Tung Thanh Nguyen, Tien N. Nguyen |
ICSM | 1 |
| 2012 | Clone Management for Evolving SoftwareabstractRecent research results suggest a need for code clone management. In this paper, we introduce JSync, a novel clone management tool. JSync provides two main functions to support developers in being aware of the clone relation among code fragments as software systems evolve and in making consistent changes as they create or modify cloned code. JSync represents source code and clones as (sub)trees in Abstract Syntax Trees, measures code similarity based on structural characteristic vectors, and describes code changes as tree editing scripts. The key techniques of JSync include the algorithms to compute tree editing scripts, to detect and update code clones and their groups, to analyze the changes of cloned code to validate their consistency, and to recommend relevant clone synchronization and merging. Our empirical study on several real-world systems shows that JSync is efficient and accurate in clone detection and updating, and provides the correct detection of the defects resulting from inconsistent changes to clones and the correct recommendations for change propagation across cloned code. Hoan Anh Nguyen, Tung Thanh Nguyen, Nam H. Pham, Jafar M. Al-Kofahi, Tien N. Nguyen |
IEEE Trans. Software Eng. | 4 |
| 2011 | Fuzzy set-based automatic bug triagingabstractAssigning a bug to the right developer is a key in reducing the cost, time, and efforts for developers in a bug fixing process. This assignment process is often referred to as bug triaging. In this paper, we propose Bugzie, a novel approach for automatic bug triaging based on fuzzy set-based modeling of bug-fixing expertise of developers. Bugzie considers a system to have multiple technical aspects, each is associated with technical terms. Then, it uses a fuzzy set to represent the developers who are capable/competent of fixing the bugs relevant to each term. The membership function of a developer in a fuzzy set is calculated via the terms extracted from the bug reports that (s)he has fixed, and the function is updated as new fixed reports are available. For a new bug report, its terms are extracted and corresponding fuzzy sets are union'ed. Potential fixers will be recommended based on their membership scores in the union'ed fuzzy set. Our preliminary results show that Bugzie achieves higher accuracy and efficiency than other state-of-the-art approaches. Ahmed Tamrawi, Tung Thanh Nguyen, Jafar M. Al-Kofahi, Tien N. Nguyen |
ICSE | 3 |
| 2011 | A topic-based approach for narrowing the search space of buggy files from a bug reportabstractLocating buggy code is a time-consuming task in software development. Given a new bug report, developers must search through a large number of files in a project to locate buggy code. We propose BugScout, an automated approach to help developers reduce such efforts by narrowing the search space of buggy files when they are assigned to address a bug report. BugScout assumes that the textual contents of a bug report and that of its corresponding source code share some technical aspects of the system which can be used for locating buggy source files given a new bug report. We develop a specialized topic model that represents those technical aspects as topics in the textual contents of bug reports and source files, and correlates bug reports and corresponding buggy files via their shared topics. Our evaluation shows that BugScout can recommend buggy files correctly up to 45% of the cases with a recommended ranked list of 10 files. Anh Tuan Nguyen 0001, Tung Thanh Nguyen, Jafar M. Al-Kofahi, Hung Viet Nguyen, Tien N. Nguyen |
ASE | 3 |
| 2011 | Fuzzy set and cache-based approach for bug triagingabstractBug triaging aims to assign a bug to the most appropriate fixer. That task is crucial in reducing time and efforts in a bug fixing process. In this paper, we propose Bugzie, a novel approach for automatic bug triaging based on fuzzy set and cache-based modeling of the bug-fixing expertise of developers. Bugzie considers a software system to have multiple technical aspects, each of which is associated with technical terms. For each technical term, it uses a fuzzy set to represent the developers who are capable/competent of fixing the bugs relevant to the corresponding aspect. The fixing correlation of a developer toward a technical term is represented by his/her membership score toward the corresponding fuzzy set. The score is calculated based on the bug reports that (s)he has fixed, and is updated as the newly fixed bug reports are available. For a new bug report, Bugzie combines the fuzzy sets corresponding to its terms and ranks the developers based on their membership scores toward that combined fuzzy set to find the most capable fixers. Our empirical results show that Bugzie achieves significantly higher accuracy and time efficiency than existing state-of-the-art approaches. Ahmed Tamrawi, Tung Thanh Nguyen, Jafar M. Al-Kofahi, Tien N. Nguyen |
SIGSOFT FSE | 3 |
| 2010 | Recurring bug fixes in object-oriented programsabstractPrevious research confirms the existence of recurring bug fixes in software systems. Analyzing such fixes manually, we found that a large percentage of them occurs in code peers, the classes/methods having the similar roles in the systems, such as providing similar functions and/or participating in similar object interactions. Based on graph-based representation of object usages, we have developed several techniques to identify code peers, recognize recurring bug fixes, and recommend changes for code units from the bug fixes of their peers. The empirical evaluation on several open-source projects shows that our prototype, FixWizard, is able to identify recurring bug fixes and provide fixing recommendations with acceptable accuracy. Tung Thanh Nguyen, Hoan Anh Nguyen, Nam H. Pham, Jafar M. Al-Kofahi, Tien N. Nguyen |
ICSE (1) | 4 |
| 2010 | Fuzzy set approach for automatic tagging in evolving softwareabstractSoftware tagging has been shown to be an efficient, lightweight social computing mechanism to improve different social and technical aspects of software development. Despite the importance of tags, there exists limited support for automatic tagging for software artifacts, especially during the evolutionary process of software development. We conducted an empirical study on IBM Jazz's repository and found that there are several missing tags in artifacts and more precise tags are desirable. This paper introduces a novel, accurate, automatic tagging recommendation tool that is able to take into account users' feedbacks on tags, and is very efficient in coping with software evolution. The core technique is an automatic tagging algorithm that is based on fuzzy set theory. Our empirical evaluation on the real-world IBM Jazz project shows the usefulness and accuracy of our approach and tool. Jafar M. Al-Kofahi, Ahmed Tamrawi, Tung Thanh Nguyen, Hoan Anh Nguyen, Tien N. Nguyen |
ICSM | 1 |
| 2009 | Accurate and Efficient Structural Characteristic Feature Extraction for Clone Detection
Hoan Anh Nguyen, Tung Thanh Nguyen, Nam H. Pham, Jafar M. Al-Kofahi, Tien N. Nguyen |
FASE | 4 |
| 2009 | Complete and accurate clone detection in graph-based modelsabstractModel-Driven Engineering (MDE) has become an important development framework for many large-scale software. Previous research has reported that as in traditional code-based development, cloning also occurs in MDE. However, there has been little work on clone detection in models with the limitations on detection precision and completeness. This paper presents ModelCD, a novel clone detection tool for Matlab/Simulink models, that is able to efficiently and accurately detect both exactly matched and approximate model clones. The core of ModelCD is two novel graph-based clone detection algorithms that are able to systematically and incrementally discover clones with a high degree of completeness, accuracy, and scalability. We have conducted an empirical evaluation with various experimental studies on many real-world systems to demonstrate the usefulness of our approach and to compare the performance of ModelCD with existing tools. Nam H. Pham, Hoan Anh Nguyen, Tung Thanh Nguyen, Jafar M. Al-Kofahi, Tien N. Nguyen |
ICSE | 4 |
| 2009 | Scalable and incremental clone detection for evolving softwareabstractCode clone management has been shown to have several benefits for software developers. When source code evolves, clone management requires a mechanism to efficiently and incrementally detect code clones in the new revision. This paper introduces an incremental clone detection tool, called ClemanX. Our tool represents code fragments as subtrees of abstract syntax trees (ASTs), measures their similarity levels based on their characteristic vectors of structural features, and solves the task of incrementally detecting similar code as an incremental distance based clustering problem. Our empirical evaluation on large-scale software projects shows the usefulness and good performance of ClemanX. Tung Thanh Nguyen, Hoan Anh Nguyen, Jafar M. Al-Kofahi, Nam H. Pham, Tien N. Nguyen |
ICSM | 3 |
| 2009 | Clone-Aware Configuration ManagementabstractRecent research results show several benefits of the management of code clones. In this paper, we introduce Clever, a novel clone-aware software configuration management (SCM) system. In addition to traditional SCM functionality, Clever provides clone management support, including clone detection and update, clone change management, clone consistency validating, clone synchronizing, and clone merging. Clever represents source code and clones as (sub)trees in Abstract Syntax Trees (ASTs), measures code similarity based on structural characteristic vectors, and describes code changes as tree editing scripts. The key techniques of Clever include the algorithms to compute tree editing scripts; to detect and update code clones and their groups; and to analyze the changes of cloned code to validate their consistency and recommend the relevant synchronization. Our empirical study on many real-world programs shows that Clever is highly efficient and accurate in clone detection and updating, and provides useful analysis of clone changes. Tung Thanh Nguyen, Hoan Anh Nguyen, Nam H. Pham, Jafar M. Al-Kofahi, Tien N. Nguyen |
ASE | 4 |
| 2009 | Graph-based mining of multiple object usage patternsabstractThe interplay of multiple objects in object-oriented programming often follows specific protocols, for example certain orders of method calls and/or control structure constraints among them that are parts of the intended object usages. Unfortunately, the information is not always documented. That creates long learning curve, and importantly, leads to subtle problems due to the misuse of objects. Tung Thanh Nguyen, Hoan Anh Nguyen, Nam H. Pham, Jafar M. Al-Kofahi, Tien N. Nguyen |
ESEC/SIGSOFT FSE | 4 |
| 2008 | Cleman: Comprehensive Clone Group Evolution ManagementabstractRecent research results have shown more benefits of the management of code clones, rather than detecting and removing them. However, existing management approaches for code clone group evolution are still ad hoc, unsatisfactory, and limited. In this paper, we introduce a novel method for comprehensive code clone group management in evolving software. The core of our method is Cleman, an algorithmic framework that allows for a systematic construction of efficient and accurate clone group management tools. Clone group management is rigorously formulated by a formal model, which provides the foundation for Cleman framework. We use Cleman framework to build a clone group management tool that is able to detect high-quality clone groups and efficiently manage them when the software evolves. We also conduct an empirical evaluation on real-world systems to show the flexibility of Cleman framework and the efficiency, completeness, and incremental updatability of our tool. Tung Thanh Nguyen, Hoan Anh Nguyen, Nam H. Pham, Jafar M. Al-Kofahi, Tien N. Nguyen |
ASE | 4 |