Michael L. Collard

dblp:96/812 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-4271-1383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 39 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scalar: A Part-of-Speech Tagger for Identifiers
abstract
The paper presents the Source Code Analysis and Lexical Annotation Runtime (SCALAR), a tool specialized for mapping (annotating) source code identifier names to their corresponding part-of-speech tag sequence (grammar pattern). SCALAR's internal model is trained using scikit-learn's GradientBoostingClassifier in conjunction with a manually-curated oracle of identifier names and their grammar patterns. This specializes the tagger to recognize the unique structure of the natural language used by developers to create all types of identifiers (e.g., function names, variable names etc.). SCALAR's output is compared with a previous version of the tagger, as well as a modern off-the-shelf part-of-speech tagger to show how it improves upon other taggers' output for annotating identifiers. The code is available on Github11https://github.com/SCANL/scanl_tagger
Christian D. Newman, Brandon Scholten, Sophia Testa, Joshua Behler, Syreen Banabilah, Michael L. Collard, Michael John Decker, Mohamed Wiem Mkaouer, Marcos Zampieri, Eman Abdullah AlOmar, Reem S. Alsuhaibani, Anthony Peruma, Jonathan I. Maletic
ICPC6
2025 Impact of Gender on OSS File Contributions
abstract
We examine how gender impacts the use of specific programming languages, as analyzed across a stratified sample of 100k unique software developers from the World of Code (WoC) archive. A total of 50,000 male and 50,000 female developers are identified using the name-to-gender inference tool WikiGender-Sort. The top fifteen programming languages according to the 2024 StackOverflow Developer survey are considered. For each developer, we count the number of files that are edited in each programming language and compute the median across gender categories. Men and women tend to edit the same number of files among most programming languages, with the exception of developers using C#, C, Go, and Rust, which had more edits among men.
Leilani Torres, Heather M. Guarnera, Michael L. Collard, Amber Garcia
SIGCSE (2)3
2024 Stereocode: A Tool for Automatic Identification of Method and Class Stereotypes for Software Systems
abstract
We present Stereocode, a static analysis tool engineered to automatically identify, and re-document software systems written in C++, C#, and/or Java with method and class stereotypes. A stereotype is a simple abstraction that encapsulates the high-level behavior of a method or a class. The tool is built around the srcML infrastructure, an XML representation of source code. Stereocode annotates the srcML input with the computed stereotypes as XML attributes to the function and class tags. We showcase Stereocode's efficiency in conducting large-scale analysis of software systems, which involves using 1050 repositories from GitHub across C++, C#, and Java. The results provide valuable insights into the distribution of stereotypes. A demo video is available at: https://youtu.be/D90xwUIPbOI.
Ali F. Al-Ramadan, Joshua Behler, Michael John Decker, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSME5
2022 An approach to automatically assess method names
abstract
An approach is presented to automatically assess the quality of method names by providing a score and feedback. The approach implements ten method naming standards to evaluate the names. The naming standards are taken from work that validated the standards via a large survey of software professionals. Natural language processing techniques such as part-of-speech tagging, identifier splitting, and dictionary lookup are required to implement the standards. The approach is evaluated by first manually constructing a large golden set of method names. Each method name is rated by several developers and labeled as conforming to each standard or not. These ratings allow for comparing the results of the approach against expert assessment. Additionally, the approach is applied to several systems and the results are manually inspected for accuracy.
Reem S. Alsuhaibani, Christian D. Newman, Michael John Decker, Michael L. Collard, Jonathan I. Maletic
ICPC4
2021 On the Naming of Methods: A Survey of Professional Developers
abstract
This paper describes the results of a large (+1100 responses) survey of professional software developers concerning standards for naming source code methods. The various standards for source code method names are derived from and supported in the software engineering literature. The goal of the survey is to determine if there is a general consensus among developers that the standards are accepted and used in practice. Additionally, the paper examines factors such as years of experience and programming language knowledge in the context of survey responses. The survey results show that participants very much agree about the importance of various standards and how they apply to names and that years of experience and the programming language has almost no effect on their responses. The results imply that the given standards are both valid and to a large degree complete. The work provides a foundation for automated method name assessment during development and code reviews.
Reem S. Alsuhaibani, Christian D. Newman, Michael John Decker, Michael L. Collard, Jonathan I. Maletic
ICSE4
2021 Special Issue on Software Maintenance Tools at 35th International Conference on Software Maintenance and Evolution (ICSME 2019)
Shinpei Hayashi, Michael L. Collard
Sci. Comput. Program.2
2020 srcDiff: A syntactic differencing approach to improve the understandability of deltas
abstract
Abstract An efficient and scalable rule‐based syntactic differencing approach is presented. The tool srcDiff is built upon the srcML infrastructure. srcML adds abstract syntactic information into the code via an XML format. A syntactic difference of srcML documents is then taken. During this process, the differences are further refined using a set of rules that model typical editing patterns of source code by developers. Thus, the resulting deltas model edits that are programmer centric versus a purely syntactic tree edit view. Other syntactic differencing approaches focus on obtaining an optimal tree edit distance with the assumption that this will produce an accurate difference. While this may work well for small or simple changes, the differences quickly become unreadable for more complex changes. By contrast, the approach presented here purposely deviates from an optimal tree edit difference in order to create a delta that is both easier to understand and better models changes between the original and modified. To evaluate the approach, a comparison user study against a state‐of‐the‐art syntactic differencing approach and two line‐based differencing tools is conducted as an online within‐participant study with about 70 subjects on 14 sample changes. The results provide support that the rule‐based syntactic differencing produces more accurate and understandable deltas.
Michael John Decker, Michael L. Collard, L. Gwenn Volkert, Jonathan I. Maletic
J. Softw. Evol. Process.2
2019 srcPtr: a framework for implementing static pointer analysis approaches
abstract
A lightweight pointer-analysis framework, srcPtr, is presented to support the implementation and comparison of points-to analysis algorithms. It differentiates itself from existing tools by performing the analysis directly on the abstract syntax tree, as opposed to an intermediate representation (e.g., LLVM IR), by using srcML, an XML representation of source code. Working with srcML and the abstract syntax allows easy access to the actual source code as the programmer views it, thus better supporting comprehension. Currently the framework provides example implementations for both Andersen's and Steensgaard's pointer-analysis algorithms. It also allows for easy integration of other points-to algorithms for comparison of accuracy/speed. The approach is very scalable and can generate pointer dependencies for a 750 KLOC program in less than a minute.
Vlas Zyrianov, Christian D. Newman, Drew T. Guarnera, Michael L. Collard, Jonathan I. Maletic
ICPC4
2018 [Research Paper] Which Method-Stereotype Changes are Indicators of Code Smells?
abstract
A study of how method roles evolve during the lifetime of a software system is presented. Evolution is examined by analyzing when the stereotype of a method changes. Stereotypes provide a high-level categorization of a method's behavior and role, and also provide insight into how a method interacts with its environment and carries out tasks. The study covers 50 open-source systems and 6 closed-source systems. Results show that method behavior with respect to stereotype is highly stable and constant over time. Overall, out of all the history examined, only about 10% of changes to methods result in a change in their stereotype. Examples of methods that change stereotype are further examined. A select number of these types of changes are indicators of code smells.
Michael John Decker, Christian D. Newman, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic, Nicholas A. Kraft
SCAM4
2018 Introduction to the special issue on program comprehension
abstract
It is a pleasure to introduce the papers in this Special Issue based on the 24th International Conference on Program Comprehension (ICPC 2016). ICPC is the principal venue for works in the area of program comprehension. ICPC aims to provide a quality forum for researchers and practitioners from academia and industry to present and discuss state-of-the-art results and best practices in the field of program comprehension. The ICPC'16 call for papers attracted 67 submissions to the research track. Each submitted paper was reviewed by at least three members of the program committee (PC). Each PC member had four or five papers to review in 25 days. Then, all papers were discussed online among the Program Co-Chairs and the PC members to make a final decision. As the output of this process, 20 papers were accepted, leading to a ~30% acceptance rate. Context-based approach to prioritize code smells for prefactoring. By Natthawute Sae-Lim, Shinpei Hayashi, and Motoshi Saeki Code smells are widely recognized as important proxies to identify code components in need of refactoring. Code smell detectors can identify hundreds of problematic components in a system and, for this reason, it is important to prioritize these smell instances to properly focus the refactoring efforts. This work proposes an approach to prioritize code smells using the working context of software developers as mined from the issue tracker of the system under analysis. The evaluation of the technique involves a study performed with professional developers. A Comprehensive Model for Code Readability. By Simone Scalabrino, Mario Linares-Vásquez, Rocco Oliveto, and Denys Poshyvanyk Automatically assessing code readability can help in identifying refactoring opportunities as well as in estimating the effort required for implementation tasks. Current readability models are built on top of structural aspects of code (e.g., line length and indentation level), but miss to capture the quality of identifiers and comments. In this paper, the authors propose a new code readability model using both structural and textual features (e.g., the consistency between terms used in comments and identifiers). The evaluation involves more than 600 code snippets for which readability is manually assessed. We hope that the readers will enjoy these two great articles. We thank the reviewers for their rigor and dedication while reviewing the submissions for this special issue, as well as the authors for fulfilling all requests and providing such excellent work. In addition, we thank all ICPC'16 PC members for their careful and detailed reviews. We are also grateful for the continuous support by the Editorial board of the Journal of Software: Evolution and Process and in particular by the Editors-in-Chief Gerardo Canfora, Darren Dalcher and David Raffo.
Gabriele Bavota, Jonathan I. Maletic, Michael L. Collard
J. Softw. Evol. Process.3
2017 The Evaluation of an Approach for Automatic Generated Documentation
abstract
Two studies are conducted to evaluate an approach to automatically generate natural language documentation summaries for C++ methods. The documentation approach relies on a method's stereotype information. First, each method is automatically assigned a stereotype(s) based on static analysis and a set of heuristics. Then, the approach uses the stereotype information, static analysis, and predefined templates to generate a natural-language summary/documentation for each method. This documentation is automatically added to the code base as a comment for each method. The result of the first study reveals that the generated documentation is accurate, does not include unnecessary information, and does a reasonable job describing what the method does. Based on statistical analysis of the second study, the most important part of the documentation is the short description as it describes the intended behavior of a method.
Nahla J. Abid, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSME3
2017 srcQL: A syntax-aware query language for source code
abstract
A tool and domain specific language for querying source code is introduced and demonstrated. The tool, srcQL, allows for the querying of source code using the syntax of the language to identify patterns within source code documents. srcQL is built upon srcML, a widely used XML representation of source code, to identify the syntactic contexts being queried. srcML inserts XML tags into the source code to mark syntactic constructs. srcQL uses a combination of XPath on srcML, regular expressions, and syntactic patterns within a query. The syntactic patterns are snippets of source code that supports the use of logical variables which are unified during the query process. This allows for very complex patterns to be easily formulated and queried. The tool is implemented (in C++) and a number of queries are presented to demonstrate the approach. srcQL currently supports C++ and scales to large systems.
Brian Bartman, Christian D. Newman, Michael L. Collard, Jonathan I. Maletic
SANER3
2017 Lexical categories for source code identifiers
abstract
A set of lexical categories, analogous to part-of-speech categories for English prose, is defined for source-code identifiers. The lexical category for an identifier is determined from its declaration in the source code, syntactic meaning in the programming language, and static program analysis. Current techniques for assigning lexical categories to identifiers use natural-language part-of-speech taggers. However, these NLP approaches assign lexical tags based on how terms are used in English prose. The approach taken here differs in that it uses only source code to determine the lexical category. The approach assigns a lexical category to each identifier and stores this information along with each declaration. srcML is used as the infrastructure to implement the approach and so the lexical information is stored directly in the srcML markup as an additional XML element for each identifier. These lexical-category annotations can then be later used by tools that automatically generate such things as code summarization or documentation. The approach is applied to 50 open source projects and the soundness of the defined lexical categories evaluated. The evaluation shows that at every level of minimum support tested, categorization is consistent at least 79% of the time with an overall consistency (across all supports) of at least 88%. The categories reveal a correlation between how an identifier is named and how it is declared. This provides a syntax-oriented view (as opposed to English part-of-speech view) of developer intent of identifiers.
Christian D. Newman, Reem S. Alsuhaibani, Michael L. Collard, Jonathan I. Maletic
SANER3
2017 Simplifying the construction of source code transformations via automatic syntactic restructurings
abstract
Abstract A set of restructurings to systematically normalize selective syntax in C++ is presented. The objective is to convert variations in syntax of specific portions of code into a single form to simplify the construction of large, complex program transformation rules. Current approaches to constructing transformations require developers to account for a large number of syntactic cases, many of which are syntactically different but semantically equivalent. The work identifies classes of such syntactic variations and presents normalizing restructurings to simplify each variation to a single, consistent syntactic form. The normalizing restructurings for C++ are presented and applied to two open source systems for evaluation. The evaluation uses the system's test cases to validate that the normalizing restructurings do not affect the systems' tested behavior. In addition, a set of example transformations that benefit from the prior application of normalizing restructurings are presented along with a small survey to assess the effect of the readability of the resultant code.
Christian D. Newman, Brian Bartman, Michael L. Collard, Jonathan I. Maletic
J. Softw. Evol. Process.3
2016 srcML 1.0: Explore, Analyze, and Manipulate Source Code
abstract
Summary form only given. This technology briefing is intended for those interested in constructing custom software analysis and manipulation tools to support research or commercial applications. srcML (srcML.org) is an infrastructure consisting of an XML representation for C/C++/C#/Java source code along with efficient parsing technology to convert source code to-and-from the srcML format. The briefing describes srcML, the toolkit, and the application of XPath and XSLT to query and modify source code. Additionally, a short tutorial of how to use srcML and XML tools to construct custom analysis and manipulation tools will be conducted.
Michael L. Collard, Jonathan I. Maletic
ICSME1
2016 A Tool for Efficiently Reverse Engineering Accurate UML Class Diagrams
abstract
A tool that reverse engineers UML class diagrams from C++ source code is presented. The tool takes srcML as input and produces yUML as output. srcML is an XML representation of the abstract syntactic information of source code. The srcML parser (srcML.org) is highly scalable, efficient, and robust. yUML is a textual format for UML class diagrams that can be easily rendered into a graphical diagram via a web service (yUML.me) or a tool such as Graphvis. The approach utilizes efficient SAX (Simple API for XML) parsing to collect the information needed to construct the class diagram. Currently it supports the following UML features: differentiating between class, data type, or interface, identifying design level attributes, multiplicity and type, determining parameter direction, and identification of the relationships aggregation, composition, generalization, and realization. The tool produces yUML for all of Calligra (~1,144KLOC) in under 20 seconds (including translation into srcML). The tool is open source under a GPL license and available for download at srcML.org.
Michael John Decker, Kyle Swartz, Michael L. Collard, Jonathan I. Maletic
ICSME3
2016 Recovering Commit Branch of Origin from GitHub Repositories
abstract
An approach to automatically recover the name of the branch where a given commit is originally made within a GitHub repository is presented and evaluated. This is a difficult task because in Git, the commit object does not store the name of the branch when it is created. Here this is termed the commit's branch of origin. Developers typically use branches in Git to group sets of changes that are related by task or concern. The approach recovers the branch of origin only within the scope of a single repository. The recovery process first uses Git's default merge commit messages and then examines the relationships between neighboring commits. The evaluation includes a simulation, an empirical examination of 40 repositories of open-source systems, and a manual verification. The evaluations show that the average accuracy exceeds 97% of all commits and the average precision exceeds 80%.
Heather M. Guarnera, Drew T. Guarnera, Michael L. Collard, Jonathan I. Maletic
ICSME3
2016 srcType: A Tool for Efficient Static Type Resolution
abstract
An efficient, static type resolution tool is presented. The tool is implemented on top of srcML, an XML representation of source code and abstract syntax. The approach computes the type of every identifier (i.e., function names and variable names) within the provided body of code. The result is a dictionary that can be used to lookup the type of each name. Type information includes metadata such as constness, class membership, aliasing, line number, file, and namespace. The approach is highly scalable and can generate a dictionary for Linux (13 MLOC) in less than 7 minutes. The tool is open source under a GPL license and available for download at srcML.org.
Christian D. Newman, Jonathan I. Maletic, Michael L. Collard
ICSME3
2016 An empirical examination of the prevalence of inhibitors to the parallelizability of open source software systems
Saleh M. Alnaeli, Jonathan I. Maletic, Michael L. Collard
Empir. Softw. Eng.3
2015 Exploration, Analysis, and Manipulation of Source Code Using srcML
abstract
This technology briefing is intended for those interested in constructing custom software analysis and manipulation tools to support research or commercial applications. srcML (srcML.org) is an infrastructure consisting of an XML representation for C/C++/C#/Java source code along with efficient parsing technology to convert source code to-and-from the srcML format. The briefing describes srcML, the toolkit, and the application of XPath and XSLT to query and modify source code. Additionally, a hands-on tutorial of how to use srcML and XML tools to construct custom analysis and manipulation tools will be conducted.
Jonathan I. Maletic, Michael L. Collard
ICSE (2)2
2015 Using stereotypes in the automatic generation of natural language summaries for C++ methods
abstract
An approach to automatically generate natural language documentation summaries for C++ methods is presented. The approach uses prior work by the authors on stereotyping methods along with the source code analysis framework srcML. First, each method is automatically assigned a stereotype(s) based on static analysis and a set of heuristics. Then, the approach uses the stereotype information, static analysis, and predefined templates to generate a natural-language summary for each method. This summary is automatically added to the code base as a comment for each method. The predefined templates are designed to produce a generic summary for specific method stereotypes. Static analysis is used to extract internal details about the method (e.g., parameters, local variables, calls, etc.). This information is used to specialize the generated summaries.
Nahla J. Abid, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSME3
2014 A Slice-Based Estimation Approach for Maintenance Effort
abstract
Program slicing is used as a basis for an approach to estimate maintenance effort. A case study of the GNU Linux kernel with over 900 versions spanning 17 years of history is presented. For each version a system dictionary is built using a lightweight slicing approach and encodes the forward decomposition static slice profiles for all variables in all the files in the system. Changes to the system are then modeled at the behavioral level using the difference between the system dictionaries of two versions. The three different granularities of slice (i.e., line, function, and file) are analyzed. We use a direct extension of srcML to represent computed change information. The retrieved information reflects the fact that additional knowledge of the differences can be automatically derived to help maintainers understand code changes. We consider the hypotheses: (1) The structured format helps create traceability links between the changes and other software artifacts. (2) This model is predictive of maintenance effort. The results demonstrate that the approach accurately predicts effort in a scalable manner.
Hakam W. Alomari, Michael L. Collard, Jonathan I. Maletic
ICSME2
2014 srcSlice: very efficient and scalable forward static slicing
abstract
ABSTRACT A highly efficient lightweight forward static slicing approach is presented and evaluated. The approach does not compute the program/system dependence graph but instead dependence and control information is computed as needed while computing the slice on a variable. The result is a list of line numbers, dependent variables, aliases, and function calls that are part of the slice for all variables (both local and global) for the entire system. The method is implemented as a tool, calledsrcSlice, on top ofsrcML, an XML representation of source code. The approach is highly scalable and can generate the slices for all variables of the Linux kernel in approximately 20 min on a typical desktop. Benchmark results are compared with theCodeSurferslicing tool from GrammaTech Inc., and the approach compares well with regard to accuracy of slices. Copyright © 2014 John Wiley & Sons, Ltd.
Hakam W. Alomari, Michael L. Collard, Jonathan I. Maletic, Nouh Alhindawi, Omar Meqdadi
J. Softw. Evol. Process.2
2013 Improving Feature Location by Enhancing Source Code with Stereotypes
abstract
A novel approach to improve feature location by enhancing the corpus (i.e., source code) with static information is presented. An information retrieval method, namely Latent Semantic Indexing (LSI), is used for feature location. Adding stereotype information to each method/function enhances the corpus. Stereotypes are terms that describe the abstract role of a method, for example get, set, and predicate are well-known method stereotypes. Each method in the system is automatically stereotyped via a static-analysis approach. Experimental comparisons of using LSI for feature location with, and without, stereotype information are conducted on a set of open-source systems. The results show that the added information improves the recall and precision in the context of feature location. Moreover, the use of stereotype information decreases the total effort that a developer would need to expend to locate relevant methods of the feature.
Nouh Alhindawi, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM3
2013 srcML: An Infrastructure for the Exploration, Analysis, and Manipulation of Source Code: A Tool Demonstration
abstract
SrcML is an XML representation for C/C++/Java source code that forms a platform for the efficient exploration, analysis, and manipulation of large software projects. The lightweight format allows for round-trip transformation from source to srcML and back to source with no loss of information or formatting. The srcML toolkit consists of the src2srcml tool for robust translation to the srcML format and the srcml2src tool for querying via XPath, and transformation via XSLT. In this demonstration a guide of these features is provided along with the use of XPath for constructing source-code queries and XSLT for conducting simple transformations.
Michael L. Collard, Michael John Decker, Jonathan I. Maletic
ICSM1
2013 Towards Understanding Large-Scale Adaptive Changes from Version Histories
abstract
A case study of three open source systems undergoing large adaptive maintenance tasks is presented. The adaptive maintenance task involves migrating each system to a new version of a third party API. The changes to support the migration were spread out over multiple years for each system. The first two systems are both part of KDE, namely KOffice and Extragear/graphics. The adaptive maintenance task, for both systems, involves migrating to a new version of Qt. The third system is OpenSceneGraph that underwent a migration to a new version of OpenGL. The case study involves sifting through tens of thousands of commits to identify only those commits involved in the specific adaptive maintenance task. The object is to develop a data set that will be used for developing automated methods to identify/characterize adaptive maintenance commits.
Omar Meqdadi, Nouh Alhindawi, Michael L. Collard, Jonathan I. Maletic
ICSM3
2011 Using stereotypes to help characterize commits
abstract
Individual commits to a version control system are automatically characterized based on the stereotypes of added and deleted methods. The stereotype of each method is automatically reverse engineered using a previously defined taxonomy. Method stereotypes reflect intrinsic atomic behavior of a method and its role in the class. The stereotypes of the added and deleted methods form a descriptors are then used to categorize commits, into types, based on the impact of the changes to a class (or classes). The goal is to gain a better understanding of the design changes to a system over its history and provide a means for documenting the commit.
Natalia Dragan, Michael L. Collard, Maen Hammad, Jonathan I. Maletic
ICSM2
2011 Lightweight Transformation and Fact Extraction with the srcML Toolkit
abstract
The srcML toolkit for lightweight transformation and fact-extraction of source code is described. srcML is an XML format for C/C++/Java source code. The open source toolkit that includes the source-to-srcML and srcML-to-source translators for round-trip reverse engineering is freely available. The direct use of XPath and XSLT is supported, an archive format for large projects is included, and a rich set of input and output formats through a command-line interface is available. Applying transformations and formulating queries using srcML is very convenient. Application use-cases of transformations and fact-extraction are shown and demonstrated to be practical and scalable.
Michael L. Collard, Michael John Decker, Jonathan I. Maletic
SCAM1
2011 Automatically identifying changes that impact code-to-design traceability during evolution
Maen Hammad, Michael L. Collard, Jonathan I. Maletic
Softw. Qual. J.2
2010 A lightweight transformational approach to support large scale adaptive changes
abstract
An approach to automate adaptive maintenance changes on large-scale software systems is presented. This approach uses lightweight parsing and lightweight on-the-fly static analysis to support transformations that make corrections to source code in response to adaptive maintenance changes, such as platform changes. SrcML, an XML source code representation, is used and transformations can be performed using either XSLT or LINQ. A number of specific adaptive changes are presented, based on recent adaptive maintenance needs from products at ABB Inc. The transformations are described in detail and then demonstrated on a number of examples from the production systems. The results are compared with manual adaptive changes that were done by professional developers. The approach performed better than the manual changes, as it successfully transformed instances missed by the developers while not missing any instances itself. The work demonstrates that this lightweight approach is both efficient and accurate with an overall cost savings in development time and effort.
Michael L. Collard, Jonathan I. Maletic, Brian P. Robinson
ICSM1
2010 Automatic identification of class stereotypes
abstract
An approach is presented to automatically determine a class's stereotype. The stereotype is based on the frequency and distribution of method stereotypes in the class. Method stereotypes are automatically determined using a defined taxonomy given in previous work. The stereotypes, boundary, control and entity are used as a basis but refined based on an empirical investigation of 21 systems. A number of heuristics, derived from empirical evidence, are used to determine a class's stereotype. For example, the prominence of certain types of methods can indicate a class's main role. The approach is applied to five open source systems and evaluated. The results show that 95% of the classes are stereotyped by the approach. Additionally, developers (via manual inspection) agreed with the approach's results.
Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM2
2010 Measuring Class Importance in the Context of Design Evolution
abstract
A measure of how a class is impacted during design evolution is presented. The history of design changes that involve a given class is the basis for the measure. Classes that are often impacted by design changes are branded as important to the design of the system. Identifying these important classes helps reveal what parts of the system are regularly evolved (e.g., specific features or cross-cutting concerns). The design importance of a class is measured as the number of commits that impact both the design and the class. This is also measured for sets of classes that collaborate to realize a feature or concept in the system. Collaborating classes are identified using itemset mining on commits that impact the design. A small study is presented on two open source projects to illustrate the approach.
Maen Hammad, Michael L. Collard, Jonathan I. Maletic
ICPC2
2009 Using method stereotype distribution as a signature descriptor for software systems
abstract
Method stereotype distribution is used as a signature for software systems. The stereotype for each method is determined using a presented taxonomy. The counts of the different stereotypes form a signature of the system. Determining method stereotypes is done automatically and is based on language (C++) features, idioms, and the main role (purpose) of a method. The intent is to use the distribution of method stereotype is an indicator of system architecture.
Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM2
2009 Automatically identifying changes that impact code-to-design traceability
abstract
An approach is presented that automatically determines if a given source code change impacts the design (i.e., UML class diagram) of the system. This allows code-to-design traceability to be consistently maintained as the source code evolves. The approach uses lightweight analysis and syntactic differencing of the source code changes to determine if the change alters the class diagram in the context of abstract design. The intent is to support both the simultaneous updating of design documents with code changes and bringing old design documents up to date with current code given the change history. An efficient tool was developed to support the approach and is applied to an open source system (i.e., HippoDraw). The results are evaluated and compared against manual inspection by human experts. The tool performs better than (error prone) manual inspection.
Maen Hammad, Michael L. Collard, Jonathan I. Maletic
ICPC2
2007 Enforcing Constraints Between Documentary Comments and Source Code
abstract
An approach for enforcing constraints between program entities and their documentary comments is presented. The approach uses srcML to represent Java source code and introduces an XML format, namely srcDoc, for marking up Javadoc-style comments. The enforced constraints are specified with a combination of XML and XQuery. An Eclipse plugin is described that demonstrates the use of XML and related technologies to express and enforce constraints on documentary comments. Examples of constraints enforcing design rationale for methods in an API are shown.
C. Dylan Shearer, Michael L. Collard
ICPC2
2007 An approach to mining call-usage patternswith syntactic context
abstract
An approach to mine frequently appearing ordered sets of function-call usages, taking into account their proximal control constructs (e.g., if-statements), in the source code is presented. These ordered sets are termed as call-usage patterns. Additionally, variant usages, such as those with missing or out of order calls, are automatically identified along with their specific contextual location. The approach uses lightweight source code analysis and frequent sequential pattern mining. The hypothesis is that these call-usage patterns embody latent programming rules that developers commonly reuse, for example standard usages of API calls. The variants are an indicator of future changes such as the elimination of non-standard usages and/or bugs
Huzefa H. Kagdi, Michael L. Collard, Jonathan I. Maletic
ASE2
2007 A survey and taxonomy of approaches for mining software repositories in the context of software evolution
abstract
Abstract A comprehensive literature survey on approaches for mining software repositories (MSR) in the context of software evolution is presented. In particular, this survey deals with those investigations that examine multiple versions of software artifacts or other temporal information. A taxonomy is derived from the analysis of this literature and presents the work via four dimensions: the type of software repositories mined (what), the purpose (why), the adopted/invented methodology used (how), and the evaluation method (quality). The taxonomy is demonstrated to be expressive (i.e., capable of representing a wide spectrum of MSR investigations) and effective (i.e., facilitates similarities and comparisons of MSR investigations). Lastly, a number of open research issues in MSR that require further investigation are identified. Copyright © 2007 John Wiley & Sons, Ltd.
Huzefa H. Kagdi, Michael L. Collard, Jonathan I. Maletic
J. Softw. Maintenance Res. Pract.2
2006 Reverse Engineering Method Stereotypes
abstract
An approach to automatically identify the stereotypes of all the methods in an entire system is presented. A taxonomy for object-oriented class method stereotypes is given that unifies and extends the existing literature to address gaps and deficiencies. Based on this taxonomy, a set of definitions is given and method stereotypes are reverse engineered using lightweight static program analysis. Classification is done solely by programming language structures and idioms, in this case C++. The approach is used to automatically re-document each method by annotating the original source code with the stereotype information. A demonstration of the accuracy and scalability of the approach is given
Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM2
2004 Supporting Source Code Difference Analysis
abstract
The paper describes an approach to easily conduct analysis of source-code differences. The approach is termed meta-differencing to reflect the fact that additional knowledge of the differences can be automatically derived. Meta-differencing is supported by an underlying source-code representation developed by the authors. The representation, srcML, is an XML format that explicitly embeds abstract syntax within the source code while preserving the documentary structure as dictated by the developer. XML tools are leveraged together with standard differencing utilities (i.e., diff,) to generate a meta-difference. The meta-difference is also represented in an XML format called srcDiff. The meta-difference contains specific syntactic information regarding the source-code changes. In turn this can be queried and searched with XML tools for the purpose of extracting information about the specifics of the changes. A case study of using the meta-differencing approach on an open-source system is presented to demonstrate its usefulness and validity.
Jonathan I. Maletic, Michael L. Collard
ICSM2
2003 An Infrastructure to Support Meta-Differencing and Refactoring of Source Code
abstract
The proposed research aims to construct an underlying infrastructure to support (semi) automated construction of refactorings and system wide transformation via a fine grained syntax level differencing approach. We term this differencing approach meta-differencing as it has additional knowledge of the types of entities being differenced. The general approach is built on top of an XML representation of the source code, specifically srcXML by J. Maletic et al. (2002). This representation explicitly embeds high level syntactic information within the source code in such a way as to not interfere with program development and maintenance. Because both the source code and the difference are represented in XML, the transformational language, XSLT, can be used to model these changes. We propose to develop an environment (development/maintenance) that automatically generates XSLT programs based on changes to a program.
Michael L. Collard
ASE1
2002 Supporting document and data views of source code
abstract
The paper describes the use of an XML format to store and represent program source code. A new XML application, srcML (SouRCe Markup Language), is presented. srcML presumes a document view of source code where information about the syntactic structure is layered over the original source code document. The resultant multi-layered document has a base layer of all the original text (and formatting). The second layer is the syntactic information, derived from the grammar of the programming language, and is encoded in XML. This multi-layered view supports both the creation and viewing of the source code in its original form and the use of XML technologies (for tasks such as analysis and transformation of the source). Although directed at source code documents, (particularly C++) srcML is also applicable to other programming languages and to languages with a strict syntax. srcML represents a departure from the compiler centric manner in which source code is commonly stored, instead a document point of view is taken thus better supporting the manipulation and management of the large numbers of source documents typical in modern software systems.
Michael L. Collard, Jonathan I. Maletic, Andrian Marcus
ACM Symposium on Document Engineering1