VLDB 2026 Research / reviewers in the wild / expert
Michael John Decker
dblp:87/10404
· DBLP profile ↗
17ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-8505-1517ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalar: A Part-of-Speech Tagger for IdentifiersabstractThe paper presents the Source Code Analysis and Lexical Annotation Runtime (SCALAR), a tool specialized for mapping (annotating) source code identifier names to their corresponding part-of-speech tag sequence (grammar pattern). SCALAR's internal model is trained using scikit-learn's GradientBoostingClassifier in conjunction with a manually-curated oracle of identifier names and their grammar patterns. This specializes the tagger to recognize the unique structure of the natural language used by developers to create all types of identifiers (e.g., function names, variable names etc.). SCALAR's output is compared with a previous version of the tagger, as well as a modern off-the-shelf part-of-speech tagger to show how it improves upon other taggers' output for annotating identifiers. The code is available on Github11https://github.com/SCANL/scanl_tagger Christian D. Newman, Brandon Scholten, Sophia Testa, Joshua Behler, Syreen Banabilah, Michael L. Collard, Michael John Decker, Mohamed Wiem Mkaouer, Marcos Zampieri, Eman Abdullah AlOmar, Reem S. Alsuhaibani, Anthony Peruma, Jonathan I. Maletic |
ICPC | 7 |
| 2025 | On the structure and semantics of identifier names containing closed syntactic category wordsabstractAbstract Identifier names are crucial components of code, serving as primary clues for developers to understand program behavior. This paper investigates the linguistic structure of identifier names by extending the concept of grammar patterns, which represent the part-of-speech (PoS) sequences underlying identifier phrases. The specific focus is on closed syntactic categories (e.g., prepositions, conjunctions, determiners), which are rarely studied in software engineering despite their central role in general natural language. To study these categories, the Closed Category Identifier Dataset (CCID), a new manually annotated dataset of 1,275 identifiers drawn from 30 open-source systems, is constructed and presented. The relationship between closed-category grammar patterns and program behavior is then analyzed using grounded-theory-inspired coding, statistical, and pattern analysis. The results reveal recurring structures that developers use to express concepts such as control flow, data transformation, temporal reasoning, and other behavioral roles through naming. This work contributes an empirical foundation for understanding how linguistic resources encode behavior in identifier names and supports new directions for research in naming, program comprehension, and education. Christian D. Newman, Anthony Peruma, Eman Abdullah AlOmar, Mahie Crabbe, Syreen Banabilah, Reem S. Alsuhaibani, Michael John Decker, Farhad Akhbardeh, Marcos Zampieri, Mohamed Wiem Mkaouer, Jonathan I. Maletic |
Empir. Softw. Eng. | 7 |
| 2024 | Stereocode: A Tool for Automatic Identification of Method and Class Stereotypes for Software SystemsabstractWe present Stereocode, a static analysis tool engineered to automatically identify, and re-document software systems written in C++, C#, and/or Java with method and class stereotypes. A stereotype is a simple abstraction that encapsulates the high-level behavior of a method or a class. The tool is built around the srcML infrastructure, an XML representation of source code. Stereocode annotates the srcML input with the computed stereotypes as XML attributes to the function and class tags. We showcase Stereocode's efficiency in conducting large-scale analysis of software systems, which involves using 1050 repositories from GitHub across C++, C#, and Java. The results provide valuable insights into the distribution of stereotypes. A demo video is available at: https://youtu.be/D90xwUIPbOI. Ali F. Al-Ramadan, Joshua Behler, Michael John Decker, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic |
ICSME | 3 |
| 2022 | An approach to automatically assess method namesabstractAn approach is presented to automatically assess the quality of method names by providing a score and feedback. The approach implements ten method naming standards to evaluate the names. The naming standards are taken from work that validated the standards via a large survey of software professionals. Natural language processing techniques such as part-of-speech tagging, identifier splitting, and dictionary lookup are required to implement the standards. The approach is evaluated by first manually constructing a large golden set of method names. Each method name is rated by several developers and labeled as conforming to each standard or not. These ratings allow for comparing the results of the approach against expert assessment. Additionally, the approach is applied to several systems and the results are manually inspected for accuracy. Reem S. Alsuhaibani, Christian D. Newman, Michael John Decker, Michael L. Collard, Jonathan I. Maletic |
ICPC | 3 |
| 2022 | An Ensemble Approach for Annotating Source Code Identifiers With Part-of-Speech TagsabstractThis paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple part-of-speech taggers to annotate natural language text at a higher quality than the part-of-speech taggers are able to obtain independently. Our ensemble uses three state-of-the-art part-of-speech taggers: SWUM, POSSE, and Stanford. We study the quality of the ensemble’s annotations on five different types of identifier names: function, class, attribute, parameter, and declaration statement at the level of both individual words and full identifier names. We also study and discuss the weaknesses of our tagger to promote the future amelioration of these problems through further research. Our results show that the ensemble achieves 75 percent accuracy at the identifier level and 84-86 percent accuracy at the word level. This is an increase of +17% points at the identifier level from the closest independent part-of-speech tagger. Christian D. Newman, Michael John Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy J. Sheldon, Emily Hill 0001 |
IEEE Trans. Software Eng. | 2 |
| 2021 | On the Naming of Methods: A Survey of Professional DevelopersabstractThis paper describes the results of a large (+1100 responses) survey of professional software developers concerning standards for naming source code methods. The various standards for source code method names are derived from and supported in the software engineering literature. The goal of the survey is to determine if there is a general consensus among developers that the standards are accepted and used in practice. Additionally, the paper examines factors such as years of experience and programming language knowledge in the context of survey responses. The survey results show that participants very much agree about the importance of various standards and how they apply to names and that years of experience and the programming language has almost no effect on their responses. The results imply that the given standards are both valid and to a large degree complete. The work provides a foundation for automated method name assessment during development and code reviews. Reem S. Alsuhaibani, Christian D. Newman, Michael John Decker, Michael L. Collard, Jonathan I. Maletic |
ICSE | 3 |
| 2020 | On the generation, structure, and semantics of grammar patterns in source code identifiers
Christian D. Newman, Reem S. Alsuhaibani, Michael John Decker, Anthony Peruma, Dishant Kaushik, Mohamed Wiem Mkaouer, Emily Hill 0001 |
J. Syst. Softw. | 3 |
| 2020 | Contextualizing rename decisions using refactorings, commit messages, and data types
Anthony Peruma, Mohamed Wiem Mkaouer, Michael John Decker, Christian D. Newman |
J. Syst. Softw. | 3 |
| 2020 | srcDiff: A syntactic differencing approach to improve the understandability of deltasabstractAbstract An efficient and scalable rule‐based syntactic differencing approach is presented. The tool srcDiff is built upon the srcML infrastructure. srcML adds abstract syntactic information into the code via an XML format. A syntactic difference of srcML documents is then taken. During this process, the differences are further refined using a set of rules that model typical editing patterns of source code by developers. Thus, the resulting deltas model edits that are programmer centric versus a purely syntactic tree edit view. Other syntactic differencing approaches focus on obtaining an optimal tree edit distance with the assumption that this will produce an accurate difference. While this may work well for small or simple changes, the differences quickly become unreadable for more complex changes. By contrast, the approach presented here purposely deviates from an optimal tree edit difference in order to create a delta that is both easier to understand and better models changes between the original and modified. To evaluate the approach, a comparison user study against a state‐of‐the‐art syntactic differencing approach and two line‐based differencing tools is conducted as an online within‐participant study with about 70 subjects on 14 sample changes. The results provide support that the rule‐based syntactic differencing produces more accurate and understandable deltas. Michael John Decker, Michael L. Collard, L. Gwenn Volkert, Jonathan I. Maletic |
J. Softw. Evol. Process. | 1 |
| 2019 | An Empirical Study of Abbreviations and Expansions in Software ArtifactsabstractExpanding abbreviations is an important text normalization technique used for the purpose of either increasing developer comprehension or supporting the application of natural-language-based tools for source code identifiers. This paper closely studies abbreviations and where their expansions occur in different software artifacts. Without abbreviation expansion, developers will spend more time in comprehending the code they need to update, and tools analyzing software may obtain weak or non-generalizable results. There are numerous techniques for expanding abbreviations, most of which struggle to reach an average expansion accuracy of 59-62% on general source code identifiers. In this paper, we reveal some characteristics of abbreviations and their expansions through an empirical study of 861 abbreviation-expansion pairs extracted from 5 open-source systems in addition to analyzing previous literature. We use these characteristics to identify how current approaches may be complementary and how their results should be reported in the future to help maximize both our understanding of how they compare with other expansion techniques and their reproducibility. Christian D. Newman, Michael John Decker, Reem S. Alsuhaibani, Anthony Peruma, Dishant Kaushik, Emily Hill 0001 |
ICSME | 2 |
| 2019 | An Open Dataset of Abbreviations and ExpansionsabstractWe present a data set of abbreviations and expansions, derived from a set of five open source systems, for use by the research and development communities. Christian D. Newman, Michael John Decker, Reem S. Alsuhaibani, Anthony Peruma, Dishant Kaushik, Emily Hill 0001 |
ICSME | 2 |
| 2019 | Contextualizing Rename Decisions using Refactorings and Commit MessagesabstractIdentifier names are the atoms of comprehension; weak identifier names decrease productivity by increasing the chance that developers make mistakes and increasing the time taken to understand chunks of code. Therefore, it is vital to support developers in naming, and renaming, identifiers. In this paper, we study how terms in an identifier change during the application of rename refactorings and contextualize these changes using co-occurring refactorings and commit messages. The goal of this work is to understand how different development activities affect the type of changes applied to names during a rename. Results of this study can help researchers understand more about developers' naming habits and support developers in determining when to rename and what words to use. Anthony Peruma, Mohamed Wiem Mkaouer, Michael John Decker, Christian D. Newman |
SCAM | 3 |
| 2018 | Leveraging the agile development process for selecting invoking/excluding tests to support feature locationabstractA practical approach to feature location using agile unit tests is presented. The approach employs a modified software reconnaissance method for feature location, but in the context of an agile development methodology. Whereas a major drawback to software reconnaissance is the identification or development of invoking and excluding tests, the approach allows for the automatic identification of invoking and excluding tests by partially ordering existing agile unit tests via iteration information from the agile development process. The approach is validated in a comparison study with industry professionals, where the approach is shown to improve feature location speed, accuracy, and developer confidence over purely manual feature location. Gregory S. DeLozier, Michael John Decker, Christian D. Newman, Jonathan I. Maletic |
ICPC | 2 |
| 2018 | [Research Paper] Which Method-Stereotype Changes are Indicators of Code Smells?abstractA study of how method roles evolve during the lifetime of a software system is presented. Evolution is examined by analyzing when the stereotype of a method changes. Stereotypes provide a high-level categorization of a method's behavior and role, and also provide insight into how a method interacts with its environment and carries out tasks. The study covers 50 open-source systems and 6 closed-source systems. Results show that method behavior with respect to stereotype is highly stable and constant over time. Overall, out of all the history examined, only about 10% of changes to methods result in a change in their stereotype. Examples of methods that change stereotype are further examined. A select number of these types of changes are indicators of code smells. Michael John Decker, Christian D. Newman, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic, Nicholas A. Kraft |
SCAM | 1 |
| 2016 | A Tool for Efficiently Reverse Engineering Accurate UML Class DiagramsabstractA tool that reverse engineers UML class diagrams from C++ source code is presented. The tool takes srcML as input and produces yUML as output. srcML is an XML representation of the abstract syntactic information of source code. The srcML parser (srcML.org) is highly scalable, efficient, and robust. yUML is a textual format for UML class diagrams that can be easily rendered into a graphical diagram via a web service (yUML.me) or a tool such as Graphvis. The approach utilizes efficient SAX (Simple API for XML) parsing to collect the information needed to construct the class diagram. Currently it supports the following UML features: differentiating between class, data type, or interface, identifying design level attributes, multiplicity and type, determining parameter direction, and identification of the relationships aggregation, composition, generalization, and realization. The tool produces yUML for all of Calligra (~1,144KLOC) in under 20 seconds (including translation into srcML). The tool is open source under a GPL license and available for download at srcML.org. Michael John Decker, Kyle Swartz, Michael L. Collard, Jonathan I. Maletic |
ICSME | 1 |
| 2013 | srcML: An Infrastructure for the Exploration, Analysis, and Manipulation of Source Code: A Tool DemonstrationabstractSrcML is an XML representation for C/C++/Java source code that forms a platform for the efficient exploration, analysis, and manipulation of large software projects. The lightweight format allows for round-trip transformation from source to srcML and back to source with no loss of information or formatting. The srcML toolkit consists of the src2srcml tool for robust translation to the srcML format and the srcml2src tool for querying via XPath, and transformation via XSLT. In this demonstration a guide of these features is provided along with the use of XPath for constructing source-code queries and XSLT for conducting simple transformations. Michael L. Collard, Michael John Decker, Jonathan I. Maletic |
ICSM | 2 |
| 2011 | Lightweight Transformation and Fact Extraction with the srcML ToolkitabstractThe srcML toolkit for lightweight transformation and fact-extraction of source code is described. srcML is an XML format for C/C++/Java source code. The open source toolkit that includes the source-to-srcML and srcML-to-source translators for round-trip reverse engineering is freely available. The direct use of XPath and XSLT is supported, an archive format for large projects is included, and a rich set of input and output formats through a command-line interface is available. Applying transformations and formulating queries using srcML is very convenient. Application use-cases of transformations and fact-extraction are shown and demonstrated to be practical and scalable. Michael L. Collard, Michael John Decker, Jonathan I. Maletic |
SCAM | 2 |