VLDB 2026 Research / reviewers in the wild / expert
Malcom Gethers
dblp:11/8025
· DBLP profile ↗
26ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 26 · 6 first-authorSystems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
10 papers |
Software maintenance and evolution · 71% Empirical software engineering · 23% Requirements engineering and software design · 6% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 71% Data mining · 29% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software maintenance and evolution
refactoring |
0.5 | 3 | 2014 | Methodbook: Recommending Move Method Refactorings via Relational Topic Models · IEEE Trans. Software Eng. 2014 Improving software modularization via automated analysis of latent topics and dependencies · ACM Trans. Softw. Eng. Methodol. 2014 Identifying method friendships to remove the feature envy bad smell · ICSE 2011 |
Software maintenance and evolution › code smell
feature envy |
0.3 | 2 | 2014 | Methodbook: Recommending Move Method Refactorings via Relational Topic Models · IEEE Trans. Software Eng. 2014 Identifying method friendships to remove the feature envy bad smell · ICSE 2011 |
Software maintenance and evolution › refactoring
move method refactoring |
0.3 | 2 | 2014 | Methodbook: Recommending Move Method Refactorings via Relational Topic Models · IEEE Trans. Software Eng. 2014 Identifying method friendships to remove the feature envy bad smell · ICSE 2011 |
Empirical software engineering
mining software repositories |
0.3 | 2 | 2013 | ExPort: Detecting and visualizing API usages in large source code repositories · ASE 2013 An adaptive approach to impact analysis from change requests to source code · ASE 2011 |
Software maintenance and evolution
change impact analysis |
0.3 | 2 | 2012 | Integrated impact analysis for managing software changes · ICSE 2012 An adaptive approach to impact analysis from change requests to source code · ASE 2011 |
Software maintenance and evolution › code smell
code smell detection |
0.2 | 1 | 2014 | Methodbook: Recommending Move Method Refactorings via Relational Topic Models · IEEE Trans. Software Eng. 2014 |
Information retrieval › retrieval models › latent semantic models
latent semantic indexing |
0.2 | 2 | 2012 | Concept location using formal concept analysis and information retrieval · ACM Trans. Softw. Eng. Methodol. 2012 Integrated impact analysis for managing software changes · ICSE 2012 |
Requirements engineering and software design
requirements traceability |
0.2 | 2 | 2012 | CodeTopics: which topic am I coding now? · ICSE 2011 Toward actionable, broadly accessible contests in Software Engineering · ICSE 2012 |
Empirical software engineering › mining software repositories › source-code mining
API usage mining |
0.2 | 1 | 2013 | ExPort: Detecting and visualizing API usages in large source code repositories · ASE 2013 |
Information retrieval › document retrieval › domain-specific retrieval
code search |
0.1 | 1 | 2012 | Concept location using formal concept analysis and information retrieval · ACM Trans. Softw. Eng. Methodol. 2012 |
Software maintenance and evolution
concept location |
0.1 | 1 | 2012 | Concept location using formal concept analysis and information retrieval · ACM Trans. Softw. Eng. Methodol. 2012 |
Software maintenance and evolution
traceability |
0.1 | 1 | 2012 | TraceLab: An experimental workbench for equipping researchers to innovate, synthesize, and comparatively evaluate traceability solutions · ICSE 2012 |
Software maintenance and evolution › traceability
traceability link recovery |
0.1 | 1 | 2012 | TraceLab: An experimental workbench for equipping researchers to innovate, synthesize, and comparatively evaluate traceability solutions · ICSE 2012 |
Empirical software engineering
developer studies |
0.1 | 1 | 2014 | Methodbook: Recommending Move Method Refactorings via Relational Topic Models · IEEE Trans. Software Eng. 2014 |
Software maintenance and evolution
software modularization |
0.1 | 1 | 2014 | Improving software modularization via automated analysis of latent topics and dependencies · ACM Trans. Softw. Eng. Methodol. 2014 |
Data mining › text mining
topic model |
0.0 | 1 | 2013 | ExPort: Detecting and visualizing API usages in large source code repositories · ASE 2013 |
Data mining › pattern mining › formal concept analysis
concept lattice |
0.0 | 1 | 2012 | Concept location using formal concept analysis and information retrieval · ACM Trans. Softw. Eng. Methodol. 2012 |
Data mining › pattern mining
formal concept analysis |
0.0 | 1 | 2012 | Concept location using formal concept analysis and information retrieval · ACM Trans. Softw. Eng. Methodol. 2012 |
Software maintenance and evolution
program comprehension |
0.0 | 1 | 2011 | CodeTopics: which topic am I coding now? · ICSE 2011 |
Methods — techniques the papers use, named apart from their topics
relational topic model · 0.8latent semantic indexing · 0.7dynamic analysis · 0.4data mining · 0.4probabilistic topic modeling · 0.4visualization · 0.3formal concept analysis · 0.3execution trace analysis · 0.3topic modeling · 0.1information retrieval · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | On the Equivalence of Information Retrieval Methods for Automated Traceability Link Recovery: A Ten-Year RetrospectiveabstractAt ICPC 2010 we presented an empirical study to statistically analyze the equivalence of several traceability recovery methods based on Information Retrieval (IR) techniques [1]. We experimented the Vector Space Model (VSM) [2], Latent Semantic Indexing (LSI) [3], the Jensen-Shannon (JS) method [4], and Latent Dirichlet Allocation (LDA) [5]. Unlike previous empirical studies we did not compare the different IR based traceability recovery methods only using the usual precision and recall metrics. We introduced some metrics to analyze the overlap of the set of candidate links recovered by each method. We also based our analysis on Principal Component Analysis (PCA) to analyze the orthogonality of the experimented methods. The results showed that while the accuracy of LDA was lower than previously used methods, LDA was able to capture some information missed by the other exploited IR methods. Instead, JS, VSM, and LSI were almost equivalent. This paved the way to possible integration of IR based traceability recovery methods [6]. Rocco Oliveto, Malcom Gethers, Denys Poshyvanyk, Andrea De Lucia |
ICPC | 2 |
| 2016 | An empirical study on how expert knowledge affects bug reportsabstractAbstract Bug reports are crucial software artifacts for both software maintenance researchers and practitioners. A typical use of bug reports by researchers is to evaluate automated software maintenance tools: a large repository of reports is used as input for a tool, and metrics are calculated from the tool's output. But this process is quite different from practitioners, who distinguish between reports written by experts, such as programmers, and reports written by non‐experts, such as users. Practitioners recognize that the content of a bug report depends on its author's expert knowledge. In this paper, we present an empirical study of the textual difference between bug reports written by experts and non‐experts. We find that a significant difference exists and that this difference has a significant impact on the results from a state‐of‐the‐art feature location tool. Through an additional study, we also found no evidence that these encountered differences were caused by the increased usage of terms from the source code in the expert bug reports. Our recommendation is that researchers evaluate maintenance tools using different sets of bug reports for experts and non‐experts. Copyright © 2016 John Wiley & Sons, Ltd. Paige Rodeghero, Collin McMillan, Malcom Gethers |
J. Softw. Evol. Process. | 5 |
| 2014 | An Empirical Study of the Effects of Expert Knowledge on Bug ReportsabstractBug reports are crucial software artifacts for both software maintenance researchers and practitioners. A typical use of bug reports by researchers is to evaluate automated software maintenance tools: a large repository of reports is used as input for a tool, and metrics are calculated from the tool's output. But this process is quite different from practitioners, who distinguish between reports written by experts such as programmers, and reports written by non-experts such as users. Practitioners recognize that the content of a bug report depends on its author's expert knowledge. In this paper, we present an empirical study of the textual difference between bug reports written by experts and non-experts. We find that a significance difference exists, and that this difference has a significant impact on the results from a state-of-the-art feature location tool. Our recommendation is that researchers evaluate maintenance tools using different sets of bug reports for experts and non-experts. Collin McMillan, Malcom Gethers |
ICSME | 4 |
| 2014 | Redacting sensitive information in software artifactsabstractIn the past decade, there have been many well-publicized cases of source code leaking from different well-known companies. These leaks pose a serious problem when the source code contains sensitive information encoded in its identifier names and comments. Unfortunately, redacting the sensitive information requires obfuscating the identifiers, which will quickly interfere with program comprehension. Program comprehension is key for programmers in understanding the source code, so sensitive information is often left unredacted. Mark Grechanik, Collin McMillan, Tathagata Dasgupta, Denys Poshyvanyk, Malcom Gethers |
ICPC | 5 |
| 2014 | Improving software modularization via automated analysis of latent topics and dependenciesabstractOftentimes, during software maintenance the original program modularization decays, thus reducing its quality. One of the main reasons for such architectural erosion is suboptimal placement of source-code classes in software packages. To alleviate this issue, we propose an automated approach to help developers improve the quality of software modularization. Our approach analyzes underlying latent topics in source code as well as structural dependencies to recommend (and explain) refactoring operations aiming at moving a class to a more suitable package. The topics are acquired via Relational Topic Models (RTM), a probabilistic topic modeling technique. The resulting tool, coined as R 3 (Rational Refactoring via RTM), has been evaluated in two empirical studies. The results of the first study conducted on nine software systems indicate that R 3 provides a coupling reduction from 10% to 30% among the software modules. The second study with 62 developers confirms that R 3 is able to provide meaningful recommendations (and explanations) for move class refactoring. Specifically, more than 70% of the recommendations were considered meaningful from a functional point of view. Gabriele Bavota, Malcom Gethers, Rocco Oliveto, Denys Poshyvanyk, Andrea De Lucia |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2014 | Methodbook: Recommending Move Method Refactorings via Relational Topic ModelsabstractDuring software maintenance and evolution the internal structure of the software system undergoes continuous changes. These modifications drift the source code away from its original design, thus deteriorating its quality, including cohesion and coupling of classes. Several refactoring methods have been proposed to overcome this problem. In this paper we propose a novel technique to identify Move Method refactoring opportunities and remove the Feature Envy bad smell from source code. Our approach, coined as Methodbook, is based on relational topic models (RTM), a probabilistic technique for representing and modeling topics, documents (in our case methods) and known relationships among these. Methodbook uses RTM to analyze both structural and textual information gleaned from software to better support move method refactoring. We evaluated Methodbook in two case studies. The first study has been executed on six software systems to analyze if the move method operations suggested by Methodbook help to improve the design quality of the systems as captured by quality metrics. The second study has been conducted with eighty developers that evaluated the refactoring recommendations produced by Methodbook. The achieved results indicate that Methodbook provides accurate and meaningful recommendations for move method refactoring operations. Gabriele Bavota, Rocco Oliveto, Malcom Gethers, Denys Poshyvanyk, Andrea De Lucia |
IEEE Trans. Software Eng. | 3 |
| 2013 | ExPort: Detecting and visualizing API usages in large source code repositoriesabstractThis paper presents a technique for automatically mining and visualizing API usage examples. In contrast to previous approaches, our technique is capable of finding examples of API usage that occur across several functions in a program. This distinction is important because of a gap between what current API learning tools provide and what programmers need: current tools extract relatively small examples from single files/functions, even though programmers use APIs to build large software. The small examples are helpful in the initial stages of API learning, but leave out details that are helpful in later stages. Our technique is intended to fill this gap. It works by representing software as a Relational Topic Model, where API calls and the functions that use them are modeled as a document network. Given a starting API, our approach can recommend complex API usage examples mined from a repository of over 14 million Java methods. Evan Moritz, Mario Linares-Vásquez, Denys Poshyvanyk, Mark Grechanik, Collin McMillan, Malcom Gethers |
ASE | 6 |
| 2013 | Integrating conceptual and logical couplings for change impact analysis in software
Huzefa H. Kagdi, Malcom Gethers, Denys Poshyvanyk |
Empir. Softw. Eng. | 2 |
| 2013 | Feature location in source code: a taxonomy and surveyabstractSUMMARY Feature location is the activity of identifying an initial location in the source code that implements functionality in a software system. Many feature location techniques have been introduced that automate some or all of this process, and a comprehensive overview of this large body of work would be beneficial to researchers and practitioners. This paper presents a systematic literature survey of feature location techniques. Eighty‐nine articles from 25 venues have been reviewed and classified within the taxonomy in order to organize and structure existing work in the field of feature location. The paper also discusses open issues and defines future directions in the field of feature location. Copyright © 2011 John Wiley & Sons, Ltd. Bogdan Dit, Meghan Revelle, Malcom Gethers, Denys Poshyvanyk |
J. Softw. Evol. Process. | 3 |
| 2012 | Toward actionable, broadly accessible contests in Software EngineeringabstractSoftware Engineering challenges and contests are becoming increasingly popular for focusing researchers' efforts on particular problems. Such contests tend to follow either an exploratory model, in which the contest holders provide data and ask the contestants to discover “interesting things” they can do with it, or task-oriented contests in which contestants must perform a specific task on a provided dataset. Only occasionally do contests provide more rigorous evaluation mechanisms that precisely specify the task to be performed and the metrics that will be used to evaluate the results. In this paper, we propose actionable and crowd-sourced contests: actionable because the contest describes a precise task, datasets, and evaluation metrics, and also provides a downloadable operating environment for the contest; and crowd-sourced because providing these features creates accessibility to Information Technology hobbyists and students who are attracted by the challenge. Our proposed approach is illustrated using research challenges from the software traceability area as well as an experimental workbench named TraceLab. Jane Cleland-Huang, Yonghee Shin, Ed Keenan, Adam Czauderna, Greg Leach, Evan Moritz, Malcom Gethers, Denys Poshyvanyk, Jane Huffman Hayes, Wenbin Li 0009 |
ICSE | 7 |
| 2012 | Integrated impact analysis for managing software changesabstractThe paper presents an adaptive approach to perform impact analysis from a given change request to source code. Given a textual change request (e.g., a bug report), a single snapshot (release) of source code, indexed using Latent Semantic Indexing, is used to estimate the impact set. Should additional contextual information be available, the approach configures the best-fit combination to produce an improved impact set. Contextual information includes the execution trace and an initial source code entity verified for change. Combinations of information retrieval, dynamic analysis, and data mining of past source code commits are considered. The research hypothesis is that these combinations help counter the precision or recall deficit of individual techniques and improve the overall accuracy. The tandem operation of the three techniques sets it apart from other related solutions. Automation along with the effective utilization of two key sources of developer knowledge, which are often overlooked in impact analysis at the change request level, is achieved. To validate our approach, we conducted an empirical evaluation on four open source software systems. A benchmark consisting of a number of maintenance issues, such as feature requests and bug fixes, and their associated source code changes was established by manual examination of these systems and their change history. Our results indicate that there are combinations formed from the augmented developer contextual information that show statistically significant improvement over standalone approaches. Malcom Gethers, Bogdan Dit, Huzefa H. Kagdi, Denys Poshyvanyk |
ICSE | 1 |
| 2012 | TraceLab: An experimental workbench for equipping researchers to innovate, synthesize, and comparatively evaluate traceability solutionsabstractTraceLab is designed to empower future traceability research, through facilitating innovation and creativity, increasing collaboration between researchers, decreasing the startup costs and effort of new traceability research projects, and fostering technology transfer. To this end, it provides an experimental environment in which researchers can design and execute experiments in TraceLab's visual modeling environment using a library of reusable and user-defined components. TraceLab fosters research competitions by allowing researchers or industrial sponsors to launch research contests intended to focus attention on compelling traceability challenges. Contests are centered around specific traceability tasks, performed on publicly available datasets, and are evaluated using standard metrics incorporated into reusable TraceLab components. TraceLab has been released in beta-test mode to researchers at seven universities, and will be publicly released via CoEST.org in the summer of 2012. Furthermore, by late 2012 TraceLab's source code will be released as open source software, licensed under GPL. TraceLab currently runs on Windows but is designed with cross platforming issues in mind to allow easy ports to Unix and Mac environments. Ed Keenan, Adam Czauderna, Greg Leach, Jane Cleland-Huang, Yonghee Shin, Evan Moritz, Malcom Gethers, Denys Poshyvanyk, Jonathan I. Maletic, Jane Huffman Hayes, Alex Dekhtyar, Daria Manukian, Shervin Hossein, Derek Hearn |
ICSE | 7 |
| 2012 | Triaging incoming change requests: Bug or commit history, or code authorship?abstractThere is a tremendous wealth of code authorship information available in source code. Motivated with the presence of this information, in a number of open source projects, an approach to recommend expert developers to assist with a software change request (e.g., a bug fixes or feature) is presented. It employs a combination of an information retrieval technique and processing of the source code authorship information. The relevant source code files to the textual description of a change request are first located. The authors listed in the header comments in these files are then analyzed to arrive at a ranked list of the most suitable developers. The approach fundamentally differs from its previously reported counterparts, as it does not require software repository mining. Neither does it require training from past bugs/issues, which is often done with sophisticated techniques such as machine learning, nor mining of source code repositories, i.e., commits. An empirical study to evaluate the effectiveness of the approach on three open source systems, ArgoUML, JEdit, and MuCommander, is reported. Our approach is compared with two representative approaches: (1) using machine learning on past bug reports, and (2) based on commit logs. The presented approach is found to provide recommendation accuracies that are equivalent or better than the two compared approaches. These findings are encouraging, as it opens up a promising and orthogonal possibility of recommending developers without the need of any historical change information. Mario Linares-Vásquez, Kamal Hossen, Hoang Dang, Huzefa H. Kagdi, Malcom Gethers, Denys Poshyvanyk |
ICSM | 5 |
| 2012 | Combining Conceptual and Domain-Based Couplings to Detect Database and Code DependenciesabstractKnowledge of software dependencies plays an important role in program comprehension and other maintenance activities. Traditionally, dependencies are derived by source code analysis, however, such an approach can be difficult to use in multi-tier hybrid software systems, or legacy applications where conventional code analysis tools simply do not work as is. In this paper, we propose a hybrid approach to detecting software dependencies by combining conceptual and domain-based coupling metrics. In recent years, a great deal of research focused on deriving various coupling metrics from these sources of information with the aim of assisting software maintainers. Conceptual metrics specify underlying relationships encoded by developers in identifiers and comments of source code classes whereas domain metrics exploit coupling manifested in domain-level information of software components and it is independent from software implementation. The proposed approach is independent from programming language, as such it can be used in multi-tier hybrid systems or legacy applications. We report the results of an empirical case study on a large-scale enterprise system where we demonstrate that the combined approach is able to detect database and source code dependencies with higher precision and recall as compared to its standalone constituents. Malcom Gethers, Amir Aryani, Denys Poshyvanyk |
SCAM | 1 |
| 2012 | Assigning change requests to software developersabstractSUMMARY The paper presents an approach to recommend a ranked list of expert developers to assist in the implementation of software change requests (e.g., bug reports and feature requests). An Information Retrieval (IR)‐based concept location technique is first used to locate source code entities, e.g., files and classes, relevant to a given textual description of a change request. The previous commits from version control repositories of these entities are then mined for expert developers. The role of the IR method in selectively reducing the mining space is different from previous approaches that textually index past change requests and/or commits. The approach is evaluated on change requests from three open‐source systems:ArgoUML,Eclipse, andKOffice, across a range of accuracy criteria. The results show that the overall accuracies of the correctly recommended developers are between 47 and 96% for bug reports, and between 43 and 60% for feature requests. Moreover, comparison results with two other recommendation alternatives show that the presented approach outperforms them with a substantial margin. Project leads or developers can use this approach in maintenance tasks immediately after the receipt of a change request in a free‐form text. Copyright © 2011 John Wiley & Sons, Ltd. Huzefa H. Kagdi, Malcom Gethers, Denys Poshyvanyk, Maen Hammad |
J. Softw. Maintenance Res. Pract. | 2 |
| 2012 | Concept location using formal concept analysis and information retrievalabstractThe article addresses the problem of concept location in source code by proposing an approach that combines Formal Concept Analysis and Information Retrieval. In the proposed approach, Latent Semantic Indexing, an advanced Information Retrieval approach, is used to map textual descriptions of software features or bug reports to relevant parts of the source code, presented as a ranked list of source code elements. Given the ranked list, the approach selects the most relevant attributes from the best ranked documents, clusters the results, and presents them as a concept lattice, generated using Formal Concept Analysis. The approach is evaluated through a large case study on concept location in the source code on six open-source systems, using several hundred features and bugs. The empirical study focuses on the analysis of various configurations of the generated concept lattices and the results indicate that our approach is effective in organizing different concepts and their relationships present in the subset of the search results. In consequence, the proposed concept location method has been shown to outperform a standalone Information Retrieval based concept location technique by reducing the number of irrelevant search results across all the systems and lattice configurations evaluated, potentially reducing the programmers' effort during software maintenance tasks involving concept location. Denys Poshyvanyk, Malcom Gethers, Andrian Marcus |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2011 | CodeTopics: which topic am I coding now?abstractRecent studies indicated that showing the similarity between the source code being developed and related high-level artifacts (HLAs), such as requirements, helps developers improve the quality of source code identifiers. In this paper, we present CodeTopics, an Eclipse plug-in that in addition to showing the similarity between source code and HLAs also highlights to what extent the code under development covers topics described in HLAs. Such views complement information derived by showing only the similarity between source code and HLAs helping (i) developers to identify functionality that are not implemented yet or (ii) newcomers to comprehend source code artifacts by showing them the topics that these artifacts relate to. Malcom Gethers, Trevor Savage, Massimiliano Di Penta, Rocco Oliveto, Denys Poshyvanyk, Andrea De Lucia |
ICSE | 1 |
| 2011 | Identifying method friendships to remove the feature envy bad smellabstractWe propose a novel approach to identify Move Method refactoring opportunities and remove the Feature Envy bad smell from source code. The proposed approach analyzes both structural and conceptual relationships between methods and uses Relational Topic Models to identify sets of methods that share several responsabilities, i.e., 'friend methods'. The analysis of method friendships of a given method can be used to pinpoint the target class (envied class) where the method should be moved in. The results of a preliminary empirical evaluation indicate that the proposed approach provides meaningful refactoring opportunities. Rocco Oliveto, Malcom Gethers, Gabriele Bavota, Denys Poshyvanyk, Andrea De Lucia |
ICSE | 2 |
| 2011 | On integrating orthogonal information retrieval methods to improve traceability recoveryabstractDifferent Information Retrieval (IR) methods have been proposed to recover traceability links among software artifacts. Until now there is no single method that sensibly outperforms the others, however, it has been empirically shown that some methods recover different, yet complementary traceability links. In this paper, we exploit this empirical finding and propose an integrated approach to combine orthogonal IR techniques, which have been statistically shown to produce dissimilar results. Our approach combines the following IR-based methods: Vector Space Model (VSM), probabilistic Jensen and Shannon (JS) model, and Relational Topic Modeling (RTM), which has not been used in the context of traceability link recovery before. The empirical case study conducted on six software systems indicates that the integrated method outperforms stand-alone IR methods as well as any other combination of non-orthogonal methods with a statistically significant margin. Malcom Gethers, Rocco Oliveto, Denys Poshyvanyk, Andrea De Lucia |
ICSM | 1 |
| 2011 | SE2 model to support software evolutionabstractThe paper proposes an integrated approach, namely SE2, to support three core software maintenance and evolution tasks: feature location, software change impact analysis, and expert developer recommendation. The approach is centered on the combinations of the conceptual and evolutionary relationships latent in structured and unstructured software artifacts. Information Retrieval (IR) and Mining Software Repositories (MSR) based techniques are used for analyzing and deriving these relationships. All the three tasks are supported under a single, common framework by providing systematic combinations of MSR and IR analyses on single and multiple versions of a software system. This combining ability of SE2sets it apart from previously reported relevant solutions in the literature. The outlined empirical assessment is aimed at identifying the exclusive and synergistic improvements offered by such combinations for each of the addressed tasks. Preliminary evaluation on a number of open source systems suggests that such combinations do offer improvements over individual approaches. Huzefa H. Kagdi, Malcom Gethers, Denys Poshyvanyk |
ICSM | 2 |
| 2011 | An adaptive approach to impact analysis from change requests to source codeabstractThe paper presents an adaptive approach to perform impact analysis from a given change request (e.g., a bug report) to source code. Given a textual change request, a single snapshot (release) of source code, indexed using Latent Semantic Indexing, is used to estimate the impact set. Additionally, the approach configures the best-fit combination of information retrieval, dynamic analysis, and data mining of past source code commits to produce an improved impact set. The tandem operation of the three techniques sets it apart from other related solutions. Malcom Gethers, Huzefa H. Kagdi, Bogdan Dit, Denys Poshyvanyk |
ASE | 1 |
| 2011 | Using structural and textual information to capture feature coupling in object-oriented software
Meghan Revelle, Malcom Gethers, Denys Poshyvanyk |
Empir. Softw. Eng. | 2 |
| 2010 | Exploiting statistical correlations for proactive prediction of program behaviorsabstractThis paper presents a finding and a technique on program behavior prediction. The finding is that surprisingly strong statistical correlations exist among the behaviors of different program components (e.g., loops) and among different types of program level behaviors (e.g., loop trip-counts versus data values). Furthermore, the correlations can be beneficially exploited: They help resolve the proactivity-adaptivity dilemma faced by existing program behavior predictions, making it possible to gain the strengths of both approaches--the large scope and earliness of offline-profiling--based predictions, and the cross-input adaptivity of runtime sampling-based predictions. Yunlian Jiang, Eddy Z. Zhang, Feng Mao, Malcom Gethers, Xipeng Shen, Yaoqing Gao |
CGO | 5 |
| 2010 | Using Relational Topic Models to capture coupling among classes in object-oriented software systemsabstractCoupling metrics capture the degree of interaction and relationships among source code elements in software systems. A vast majority of existing coupling metrics rely on structural information, which captures interactions such as usage relations between classes and methods or execute after associations. However, these metrics lack the ability to identify conceptual dependencies, which, for instance, specify underlying relationships encoded by developers in identifiers and comments of source code classes. We propose a new coupling metric for object-oriented software systems, namely Relational Topic based Coupling (RTC) of classes, which uses Relational Topic Models (RTM), generative probabilistic model, to capture latent topics in source code classes and relationships among them. A case study on thirteen open source software systems is performed to compare the new measure with existing structural and conceptual coupling metrics. The case study demonstrates that proposed metric not only captures new dimensions of coupling, which are not covered by the existing coupling metrics, but also can be used to effectively support impact analysis. Malcom Gethers, Denys Poshyvanyk |
ICSM | 1 |
| 2010 | TopicXP: Exploring topics in source code using Latent Dirichlet AllocationabstractAcquiring general understanding of large software systems and components from which they are built can be a time consuming task, but having such an understanding is an important prerequisite to adding features or fixing bugs. In this paper we propose the tool, namely TopicXP, to support developers during such software maintenance tasks by extracting and analyzing unstructured information in source code identifier names and comments using Latent Dirichlet Allocation. TopicXPenables developers to gain an overview of a software system under analysis by extracting and visualizing natural language topics, which generally correspond to concepts or features implemented in software classes. TopicXPis implemented as an open-source Eclipse plug-in, which proposes interactive visualization of topics along with structural dependencies between underlying classes implementing these topics. The paper also presents the results of a preliminary user study aimed at evaluating TopicXP. Trevor Savage, Bogdan Dit, Malcom Gethers, Denys Poshyvanyk |
ICSM | 3 |
| 2010 | On the Equivalence of Information Retrieval Methods for Automated Traceability Link RecoveryabstractWe present an empirical study to statistically analyze the equivalence of several traceability recovery methods based on Information Retrieval (IR) techniques. The analysis is based on Principal Component Analysis and on the analysis of the overlap of the set of candidate links provided by each method. The studied techniques are the Jensen-Shannon (JS) method, Vector Space Model (VSM), Latent Semantic Indexing (LSI), and Latent Dirichlet Allocation (LDA). The results show that while JS, VSM, and LSI are almost equivalent, LDA is able to capture a dimension unique to the set of techniques which we considered. Rocco Oliveto, Malcom Gethers, Denys Poshyvanyk, Andrea De Lucia |
ICPC | 2 |