Jonathan I. Maletic

dblp:m/JonathanIMaletic · DBLP profile ↗
← Back
101ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0001-5289-135XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 83 · 6 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Databases, data management, data science and information retrieval · 4
YearPublicationVenuePosition
2025 Extending Support for Analyzing Eye Tracking Studies on Python Source Code in iTrace
Joshua Behler, Zachary Kozak, Kang-Il Park, Bonita Sharif, Jonathan I. Maletic
ETRA5
2025 Scalar: A Part-of-Speech Tagger for Identifiers
abstract
The paper presents the Source Code Analysis and Lexical Annotation Runtime (SCALAR), a tool specialized for mapping (annotating) source code identifier names to their corresponding part-of-speech tag sequence (grammar pattern). SCALAR's internal model is trained using scikit-learn's GradientBoostingClassifier in conjunction with a manually-curated oracle of identifier names and their grammar patterns. This specializes the tagger to recognize the unique structure of the natural language used by developers to create all types of identifiers (e.g., function names, variable names etc.). SCALAR's output is compared with a previous version of the tagger, as well as a modern off-the-shelf part-of-speech tagger to show how it improves upon other taggers' output for annotating identifiers. The code is available on Github11https://github.com/SCANL/scanl_tagger
Christian D. Newman, Brandon Scholten, Sophia Testa, Joshua Behler, Syreen Banabilah, Michael L. Collard, Michael John Decker, Mohamed Wiem Mkaouer, Marcos Zampieri, Eman Abdullah AlOmar, Reem S. Alsuhaibani, Anthony Peruma, Jonathan I. Maletic
ICPC13
2025 On the structure and semantics of identifier names containing closed syntactic category words
abstract
Abstract Identifier names are crucial components of code, serving as primary clues for developers to understand program behavior. This paper investigates the linguistic structure of identifier names by extending the concept of grammar patterns, which represent the part-of-speech (PoS) sequences underlying identifier phrases. The specific focus is on closed syntactic categories (e.g., prepositions, conjunctions, determiners), which are rarely studied in software engineering despite their central role in general natural language. To study these categories, the Closed Category Identifier Dataset (CCID), a new manually annotated dataset of 1,275 identifiers drawn from 30 open-source systems, is constructed and presented. The relationship between closed-category grammar patterns and program behavior is then analyzed using grounded-theory-inspired coding, statistical, and pattern analysis. The results reveal recurring structures that developers use to express concepts such as control flow, data transformation, temporal reasoning, and other behavioral roles through naming. This work contributes an empirical foundation for understanding how linguistic resources encode behavior in identifier names and supports new directions for research in naming, program comprehension, and education.
Christian D. Newman, Anthony Peruma, Eman Abdullah AlOmar, Mahie Crabbe, Syreen Banabilah, Reem S. Alsuhaibani, Michael John Decker, Farhad Akhbardeh, Marcos Zampieri, Mohamed Wiem Mkaouer, Jonathan I. Maletic
Empir. Softw. Eng.11
2025 Automated Fixation Error Correction to Support Eye Tracking Studies on Source Code
abstract
A significant challenge in eye-tracking studies is detecting and fixing errors in data collection that happen for various reasons (drift, calibration issues, etc.). Many errors cannot be fully mitigated and require manual correction (which is intensively time-consuming) or automated correction. The work presented in this paper focuses on error correction, primarily on eye-tracking data on source code written in programming languages such as C++, Java, and C#. Many automated correction solutions are general-purpose, computationally inefficient, and use little information about the stimulus. To bridge this gap, we introduce srcGaze , a heuristic algorithm explicitly developed for correcting fixation gaze events in eye-tracking data from studies using source code as a stimulus. A golden dataset is manually constructed and verified to establish the heuristics. Results show a ≈40% improvement compared to no fixation correction. The approach has a multi-linear complexity and can correct over 44K fixations in approximately 6 seconds.
Drew T. Guarnera, Joshua Behler, Bonita Sharif, Jonathan I. Maletic
Proc. ACM Hum. Comput. Interact.4
2024 Stereocode: A Tool for Automatic Identification of Method and Class Stereotypes for Software Systems
abstract
We present Stereocode, a static analysis tool engineered to automatically identify, and re-document software systems written in C++, C#, and/or Java with method and class stereotypes. A stereotype is a simple abstraction that encapsulates the high-level behavior of a method or a class. The tool is built around the srcML infrastructure, an XML representation of source code. Stereocode annotates the srcML input with the computed stereotypes as XML attributes to the function and class tags. We showcase Stereocode's efficiency in conducting large-scale analysis of software systems, which involves using 1050 repositories from GitHub across C++, C#, and Java. The results provide valuable insights into the distribution of stereotypes. A demo video is available at: https://youtu.be/D90xwUIPbOI.
Ali F. Al-Ramadan, Joshua Behler, Michael John Decker, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSME6
2024 Extending iTrace-Visualize to Support Token-based Heatmaps and Region of Interest Scarf Plots for Source Code
abstract
The iTrace Infrastructure is a suite of community eye-tracking tools that enables researchers to conduct eye-tracking studies on software projects in real development environments. The infrastructure consists of tools providing support for data gathering, post processing, and visualization. iTrace-Visualize provides researchers with a way to visualize gathered and post-processed eye-movement data. iTrace involves the analysis of more than just eye-movement data, and includes information gathered from the development environment and the source code. This work describes additions to iTrace-Visualize that provide researchers with visualizations of the gathered source code data. Specifically, a tokenized heatmap of the source code is presented, which shows the source code tokens that are viewed the most. Additionally, a region of interest scarf plot that details the timeline of what parts of the code a participant views is added as a new feature. A usage example comparing student and industry developers is presented to demonstrate the use of these tools. Demo Video: https://youtu.be/0iZcCC8CK94
Joshua Behler, Giovanni Villalobos, Julia Pangonis, Bonita Sharif, Jonathan I. Maletic
VISSOFT5
2024 Exploring How Developers Layout UML Class Diagrams
abstract
The paper presents a video-based exploratory study that seeks to understand how developers modify UML class diagram layouts for better readability and comprehension of the system. Two diagram layouts showing a model subset from a Java open-source system are presented to six participants experienced in reading UML class diagrams. They are tasked to change the layout to make it easier for them to read and comprehend. The video is reviewed for major modifications to the layouts. Behaviors observed are presented. The eventual goal is to use this information to construct heuristics for automated layout algorithms based on semantics and architectural importance.
Bonita Sharif, Nathaniel Liess, Jonathan I. Maletic
VISSOFT3
2024 Examining the Effects of Layout and Working Memory on UML Class Diagram Defect Identification
abstract
A controlled experiment investigating the effect layout has on how students find defects in UML class diagrams with respect to requirements is presented. Two layout schemes from prior literature namely, multi-cluster and orthogonal layouts, are compared with respect to two open source systems, Doxygen and Qt. The experiment is conducted with 89 students from two universities in a classroom lab setting. Each participant is placed in one of two groups where each group are given 2 defect detection tasks (with five sub-parts) with each task using one of the two layouts in each subject system. The only difference between groups is that the layouts were flipped between the two tasks. Feedback is collected after each task. A mental rotation and object memory task is conducted at the end of the two tasks to correlate their spatial and working memory skills to the task performance. Results indicate that the multi-cluster layout performed better in terms of accuracy of finding defects, but not significantly. There is also not much difference in time to find them. Furthermore, it is found that the object memory skills are sometimes correlated with the performance of the defect detection tasks. These results can be used to help improve the teaching of UML class diagram defect detection skills by incorporating clustered layouts and object memory tasks. In addition, they can help identify people who are best suited for finding critical defects in design.
Bonita Sharif, Kang-Il Park, Michael P. DeJournett, Isaac Baysinger, Mohammed Aly, Jonathan I. Maletic
VISSOFT6
2024 Scoping Software Engineering for AI: The TSE Perspective
abstract
Advances in Artificial Intelligence (AI), and in particular in Machine Learning (ML), are introducing profound changes to scholarly submissions across publication venues, affecting in particular the contributions that are being submitted to Software Engineering (SE) conferences and journals. In this context, it is not always clear whether manuscripts submitted to SE venues under the umbrella term SE for AI are indeed relevant to SE, in the sense that they explicitly contain contributions to the SE body of knowledge. This leads to recurring discussions on whether certain AI-related submissions are appropriate to SE venues, or should instead be submitted to other journals and conferences, including AI or ML-specific ones. In this editorial, we discuss the kinds of AI-related contributions that are a better fit-and a less good fit-for publication in the IEEE Transactions on Software Engineering.
Sebastián Uchitel, Marsha Chechik, Massimiliano Di Penta, Bram Adams, Nazareno Aguirre, Gabriele Bavota, Domenico Bianculli, Kelly Blincoe, Ana Cavalcanti 0001, Yvonne Dittrich, Filomena Ferrucci, Rashina Hoda, LiGuo Huang, David Lo 0001, Michael R. Lyu, Lei Ma 0003, Jonathan I. Maletic, Leonardo Mariani, Collin McMillan, Tim Menzies, Martin Monperrus, Ana Moreno, Nachiappan Nagappan, Liliana Pasquale, Patrizio Pelliccione, Michael Pradel, Rahul Purandare, Sukyoung Ryu, Mehrdad Sabetzadeh, Alexander Serebrenik, Jun Sun 0001, Chakkrit Tantithamthavorn, Christoph Treude, Manuel Wimmer, Yingfei Xiong 0001, Tao Yue 0002, Andy Zaidman, Tao Zhang 0001, Hao Zhong 0001
IEEE Trans. Software Eng.17
2023 iTrace-Visualize: Visualizing Eye-Tracking Data for Software Engineering Studies
abstract
iTrace is community infrastructure that allows software engineering researchers to conduct eye-tracking studies on large realistic code bases. The iTrace infrastructure consists of a set of tools that assist with gathering, processing, and evaluating eye-tracking data on large software projects within an Integrated Development Environment (IDE). A typical eye-tracking study results in millions of raw gazes that are overwhelming to view and sort through. To help researchers view and comprehend this data, iTrace-Visualize is presented. This tool integrates information produced by the iTrace infrastructure into a dynamic video recording of the eye-tracking session. Eye fixations and the scan path between fixations are overlayed on the video. Additionally, the line being examined can be highlighted in the video. iTrace-Visualize lets a researcher replay eye fixations via a video overlay immediately after a study. This serves as quick validation of what was done during the study and can also provide quick insights into what the participants looked at. To illustrate iTrace-Visualize's capabilities, a small preliminary study is performed. Demo Video-https://youtu.be/c1hUFDmBM50
Joshua Behler, Gino Chiudioni, Alex Ely, Julia Pangonis, Bonita Sharif, Jonathan I. Maletic
VISSOFT6
2023 Studying Developer Eye Movements to Measure Cognitive Workload and Visual Effort for Expertise Assessment
abstract
Eye movement data provides valuable insights that help test hypotheses about a software developer's comprehension process. The pupillary response is successfully used to assess mental processing effort and attentional focus. Relatively little is known about the impact of expertise level in cognitive effort during programming tasks. This paper presents a quantitative analysis that compares the eye movements of 207 experts and novices collected while solving program comprehension tasks. The goal is to examine changes of developers' eye movement metrics in accordance with their expertise. The results indicate significant increase in pupil size with the novice group compared to the experts, explaining higher cognitive effort for novices. Novices also tend to have a significant number of fixations and higher gaze time compared to experts when they comprehend code. Moreover, a correlation study found that programming experience is still a powerful indicator when explaining expertise in this eye-tracking dataset among other expertise variables.
Salwa D. Aljehane, Bonita Sharif, Jonathan I. Maletic
Proc. ACM Hum. Comput. Interact.3
2022 An approach to automatically assess method names
abstract
An approach is presented to automatically assess the quality of method names by providing a score and feedback. The approach implements ten method naming standards to evaluate the names. The naming standards are taken from work that validated the standards via a large survey of software professionals. Natural language processing techniques such as part-of-speech tagging, identifier splitting, and dictionary lookup are required to implement the standards. The approach is evaluated by first manually constructing a large golden set of method names. Each method name is rated by several developers and labeled as conforming to each standard or not. These ratings allow for comparing the results of the approach against expert assessment. Additionally, the approach is applied to several systems and the results are manually inspected for accuracy.
Reem S. Alsuhaibani, Christian D. Newman, Michael John Decker, Michael L. Collard, Jonathan I. Maletic
ICPC5
2022 Deja Vu: semantics-aware recording and replay of high-speed eye tracking and interaction data to support cognitive studies of software engineering tasks - methodology and analyses
Vlas Zyrianov, Cole S. Peterson, Drew T. Guarnera, Joshua Behler, Praxis Weston, Bonita Sharif, Jonathan I. Maletic
Empir. Softw. Eng.7
2021 On the Naming of Methods: A Survey of Professional Developers
abstract
This paper describes the results of a large (+1100 responses) survey of professional software developers concerning standards for naming source code methods. The various standards for source code method names are derived from and supported in the software engineering literature. The goal of the survey is to determine if there is a general consensus among developers that the standards are accepted and used in practice. Additionally, the paper examines factors such as years of experience and programming language knowledge in the context of survey responses. The survey results show that participants very much agree about the importance of various standards and how they apply to names and that years of experience and the programming language has almost no effect on their responses. The results imply that the given standards are both valid and to a large degree complete. The work provides a foundation for automated method name assessment during development and code reviews.
Reem S. Alsuhaibani, Christian D. Newman, Michael John Decker, Michael L. Collard, Jonathan I. Maletic
ICSE5
2021 From Novice to Expert: Analysis of Token Level Effects in a Longitudinal Eye Tracking Study
abstract
Program comprehension is a vital skill in software development. This work investigates program comprehension by examining the eye movement of novice programmers as they gain programming experience over the duration of a Java course. Their eye movement behavior is compared to the eye movement of expert programmers. Eye movement studies of natural text show that word frequency and length influence eye movement duration and act as indicators of reading skill. The study uses an existing longitudinal eye tracking dataset with 20 novice and experienced readers of source code. The work investigates the acquisition of the effects of token frequency and token length in source code reading as an indication of program reading skill. The results show evidence of the frequency and length effects in reading source code and the acquisition of these effects by novices. These results are then leveraged in a machine learning model demonstrating how eye movement can be used to estimate programming proficiency and classify novices from experts with 72% accuracy.
Naser Al Madi, Cole S. Peterson, Bonita Sharif, Jonathan I. Maletic
ICPC4
2020 Integration of Program Slicing with Cognitive Complexity for Defect Prediction
abstract
Researchers have identified several quality metrics to predict defects that rely on different types of information. However, these approaches lack metrics to estimate the effort of program understandability of system artifacts. Code that is understandable is often considered more maintainable. This paper briefly describes a dissertation which presents novel metrics to compute the cognitive complexity based on program slicing. These metrics help identify code that is more likely to have defects due to being challenging to comprehension. A thorough empirical investigation into how cognitive complexity correlates with and predicts defects is performed. Finally, this paper discusses potential direction for future research based upon the work conducted, as well as a reflection on the PhD study, providing advice for current students.
Basma S. Alqadi, Jonathan I. Maletic
ICSME2
2020 Automated Recording and Semantics-Aware Replaying of High-Speed Eye Tracking and Interaction Data to Support Cognitive Studies of Software Engineering Tasks
abstract
The paper introduces a fundamental technological problem with collecting high-speed eye tracking data while studying software engineering tasks in an integrated development environment. The use of eye trackers is quickly becoming an important means to study software developers and how they comprehend source code and locate bugs. High quality eye trackers can record upwards of 120 to 300 gaze points per second. However, it is not possible to map each of these points to a line and column position in a source code file (in the presence of scrolling and file switching) in real time at data rates over 60 gaze points per second without data loss. Unfortunately, higher data rates are more desirable as they allow for finer granularity and more accurate study analyses. To alleviate this technological problem, a novel method for eye tracking data collection is presented. Instead of performing gaze analysis in real time, all telemetry (keystrokes, mouse movements, and eye tracker output) data during a study is recorded as it happens. Sessions are then replayed at a much slower speed allowing for ample time to map gaze point positions to the appropriate file, line, and column to perform additional analysis. A description of the method and corresponding tool, Déjà Vu, is presented. An evaluation of the method and tool is conducted using three different eye trackers running at four different speeds (60Hz, 120Hz, 150Hz, and 300 Hz). This timing evaluation is performed in Visual Studio and Eclipse IDEs. Results show that Déjà Vu can playback 100% of the data recordings, correctly mapping the gaze to corresponding elements, making it a well-founded and suitable post processing step for future eye tracking studies in software engineering.
Vlas Zyrianov, Drew T. Guarnera, Cole S. Peterson, Bonita Sharif, Jonathan I. Maletic
ICSME5
2020 Slice-Based Cognitive Complexity Metrics for Defect Prediction
abstract
Researchers have identified several quality metrics to predict defects, relying on different information however, these approaches lack metrics to estimate the effort of program understandability of system artifacts. In this paper, novel metrics to compute the cognitive complexity based on program slicing are introduced. These metrics help identify code that is more likely to have defects due to being challenging to comprehension. The metrics include such measures as the total number of slices in a file, the size, the average number of identifiers, and the average spatial distance of a slice. A scalable lightweight slicing tool is used to compute the necessary slicing data. A thorough empirical investigation into how cognitive complexity correlates with and predicts defects in the version histories of 10 datasets of 7 open source systems is performed. The results show that the increase of cognitive complexity significantly increases the number of defects in 94% of the cases. In a comparison study to metrics that have been shown to correlate with understandability and with defects, the addition of cognitive complexity metrics shows better prediction by up to 14% in F1, 16% in AUC, and 35% in R2.
Basma S. Alqadi, Jonathan I. Maletic
SANER2
2020 srcDiff: A syntactic differencing approach to improve the understandability of deltas
abstract
Abstract An efficient and scalable rule‐based syntactic differencing approach is presented. The tool srcDiff is built upon the srcML infrastructure. srcML adds abstract syntactic information into the code via an XML format. A syntactic difference of srcML documents is then taken. During this process, the differences are further refined using a set of rules that model typical editing patterns of source code by developers. Thus, the resulting deltas model edits that are programmer centric versus a purely syntactic tree edit view. Other syntactic differencing approaches focus on obtaining an optimal tree edit distance with the assumption that this will produce an accurate difference. While this may work well for small or simple changes, the differences quickly become unreadable for more complex changes. By contrast, the approach presented here purposely deviates from an optimal tree edit difference in order to create a delta that is both easier to understand and better models changes between the original and modified. To evaluate the approach, a comparison user study against a state‐of‐the‐art syntactic differencing approach and two line‐based differencing tools is conducted as an online within‐participant study with about 70 subjects on 14 sample changes. The results provide support that the rule‐based syntactic differencing produces more accurate and understandable deltas.
Michael John Decker, Michael L. Collard, L. Gwenn Volkert, Jonathan I. Maletic
J. Softw. Evol. Process.4
2019 Using developer eye movements to externalize the mental model used in code summarization tasks
abstract
Eye movements of developers are used to speculate the mental cognition model (i.e., bottom-up or top-down) applied during program comprehension tasks. The cognition models examine how programmers understand source code by describing the temporary information structures in the programmer's short term memory. The two types of models that we are interested in are top-down and bottom-up. The top-down model is normally applied as-needed (i.e., the domain of the system is familiar). The bottom-up model is typically applied when a developer is not familiar with the domain or the source code. An eye-tracking study of 18 developers reading and summarizing Java methods is used as our dataset for analyzing the mental cognition model. The developers provide a written summary for methods assigned to them. In total, 63 methods are used from five different systems. The results indicate that on average, experts and novices read the methods more closely (using the bottom-up mental model) than bouncing around (using top-down). However, on average novices spend longer gaze time performing bottom-up (66s.) compared to experts (43s.)
Nahla J. Abid, Jonathan I. Maletic, Bonita Sharif
ETRA2
2019 Factors influencing dwell time during source code reading: a large-scale replication experiment
abstract
The paper partially replicates and extends a previous study by Busjahn et al. [4] on the factors influencing dwell time during source code reading, where source code element type and frequency of gaze visits are studied as factors. Unlike the previous study, this study focuses on analyzing eye movement data in large open source Java projects. Five experts and thirteen novices participated in the study where the main task is to summarize methods. The results examine semantic line-level information that developers view during summarization. We find no correlation between the line length and the total duration of time spent looking on the line even though it exists between a token's length and the total fixation time on the token reported in prior work. The first fixations inside a method are more likely to be on a method's signature, a variable declaration, or an assignment compared to the other fixations inside a method. In addition, it is found that smaller methods tend to have shorter overall fixation duration for the entire method, but have significantly longer duration per line in the method. The analysis provides insights into how source code's unique characteristics can help in building more robust methods for analyzing eye movements in source code and overall in building theories to support program comprehension on realistic tasks.
Cole S. Peterson, Nahla J. Abid, Corey A. Bryant, Jonathan I. Maletic, Bonita Sharif
ETRA4
2019 Developer reading behavior while summarizing Java methods: size and context matters
abstract
An eye-tracking study of 18 developers reading and summarizing Java methods is presented. The developers provide a written summary for methods assigned to them. In total, 63 methods are used from five different systems. Previous studies on this topic use only short methods presented in isolation usually as images. In contrast, this work presents the study in the Eclipse IDE allowing access to all the source code in the system. The developer can navigate via scrolling and switching files while writing the summary. New eye-tracking infrastructure allows for this improvement in the study environment. Data collected includes eye gazes on source code, written summaries, and time to complete each summary. Unlike prior work that concluded developers focus on the signature the most, these results indicate that they tend to focus on the method body more than the signature. Moreover, both experts and novices tend to revisit control flow terms rather than reading them for a long period. They also spend a significant amount of gaze time and have higher gaze visits when they read call terms. Experts tend to revisit the body of the method significantly more frequently than its signature as the size of the method increases. Moreover, experts tend to write their summaries from source code lines that they read the most.
Nahla J. Abid, Bonita Sharif, Natalia Dragan, Hend Alrasheed, Jonathan I. Maletic
ICSE5
2019 srcPtr: a framework for implementing static pointer analysis approaches
abstract
A lightweight pointer-analysis framework, srcPtr, is presented to support the implementation and comparison of points-to analysis algorithms. It differentiates itself from existing tools by performing the analysis directly on the abstract syntax tree, as opposed to an intermediate representation (e.g., LLVM IR), by using srcML, an XML representation of source code. Working with srcML and the abstract syntax allows easy access to the actual source code as the programmer views it, thus better supporting comprehension. Currently the framework provides example implementations for both Andersen's and Steensgaard's pointer-analysis algorithms. It also allows for easy integration of other points-to algorithms for comparison of accuracy/speed. The approach is very scalable and can generate pointer dependencies for a 750 KLOC program in less than a minute.
Vlas Zyrianov, Christian D. Newman, Drew T. Guarnera, Michael L. Collard, Jonathan I. Maletic
ICPC5
2018 iTrace: eye tracking infrastructure for development environments
abstract
The paper presents iTrace, an eye tracking infrastructure, that enables eye tracking in development environments such as Visual Studio and Eclipse. Software developers work with software that is comprised of numerous source code files. This requires frequent switching between project artifacts during program understanding or debugging activities. Additionally, the amount of content contained within each artifact can be quite large and require scrolling or navigation of the content. Current approaches to eye tracking are meant for fixed stimuli and struggle to capture context during these activities. iTrace overcomes these limitations allowing developers to work in realistic settings during an eye tracking study. The iTrace architecture is presented along with several use cases of where it can be used by researchers. A short video demonstration is available at https://youtu.be/AmrLWgw4OEs
Drew T. Guarnera, Corey A. Bryant, Ashwin Mishra, Jonathan I. Maletic, Bonita Sharif
ETRA4
2018 Leveraging the agile development process for selecting invoking/excluding tests to support feature location
abstract
A practical approach to feature location using agile unit tests is presented. The approach employs a modified software reconnaissance method for feature location, but in the context of an agile development methodology. Whereas a major drawback to software reconnaissance is the identification or development of invoking and excluding tests, the approach allows for the automatic identification of invoking and excluding tests by partially ordering existing agile unit tests via iteration information from the agile development process. The approach is validated in a comparison study with industry professionals, where the approach is shown to improve feature location speed, accuracy, and developer confidence over purely manual feature location.
Gregory S. DeLozier, Michael John Decker, Christian D. Newman, Jonathan I. Maletic
ICPC4
2018 [Research Paper] Which Method-Stereotype Changes are Indicators of Code Smells?
abstract
A study of how method roles evolve during the lifetime of a software system is presented. Evolution is examined by analyzing when the stereotype of a method changes. Stereotypes provide a high-level categorization of a method's behavior and role, and also provide insight into how a method interacts with its environment and carries out tasks. The study covers 50 open-source systems and 6 closed-source systems. Results show that method behavior with respect to stereotype is highly stable and constant over time. Overall, out of all the history examined, only about 10% of changes to methods result in a change in their stereotype. Examples of methods that change stereotype are further examined. A select number of these types of changes are indicators of code smells.
Michael John Decker, Christian D. Newman, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic, Nicholas A. Kraft
SCAM5
2018 Introduction to the special issue on program comprehension
abstract
It is a pleasure to introduce the papers in this Special Issue based on the 24th International Conference on Program Comprehension (ICPC 2016). ICPC is the principal venue for works in the area of program comprehension. ICPC aims to provide a quality forum for researchers and practitioners from academia and industry to present and discuss state-of-the-art results and best practices in the field of program comprehension. The ICPC'16 call for papers attracted 67 submissions to the research track. Each submitted paper was reviewed by at least three members of the program committee (PC). Each PC member had four or five papers to review in 25 days. Then, all papers were discussed online among the Program Co-Chairs and the PC members to make a final decision. As the output of this process, 20 papers were accepted, leading to a ~30% acceptance rate. Context-based approach to prioritize code smells for prefactoring. By Natthawute Sae-Lim, Shinpei Hayashi, and Motoshi Saeki Code smells are widely recognized as important proxies to identify code components in need of refactoring. Code smell detectors can identify hundreds of problematic components in a system and, for this reason, it is important to prioritize these smell instances to properly focus the refactoring efforts. This work proposes an approach to prioritize code smells using the working context of software developers as mined from the issue tracker of the system under analysis. The evaluation of the technique involves a study performed with professional developers. A Comprehensive Model for Code Readability. By Simone Scalabrino, Mario Linares-Vásquez, Rocco Oliveto, and Denys Poshyvanyk Automatically assessing code readability can help in identifying refactoring opportunities as well as in estimating the effort required for implementation tasks. Current readability models are built on top of structural aspects of code (e.g., line length and indentation level), but miss to capture the quality of identifiers and comments. In this paper, the authors propose a new code readability model using both structural and textual features (e.g., the consistency between terms used in comments and identifiers). The evaluation involves more than 600 code snippets for which readability is manually assessed. We hope that the readers will enjoy these two great articles. We thank the reviewers for their rigor and dedication while reviewing the submissions for this special issue, as well as the authors for fulfilling all requests and providing such excellent work. In addition, we thank all ICPC'16 PC members for their careful and detailed reviews. We are also grateful for the continuous support by the Editorial board of the Journal of Software: Evolution and Process and in particular by the Editors-in-Chief Gerardo Canfora, Darren Dalcher and David Raffo.
Gabriele Bavota, Jonathan I. Maletic, Michael L. Collard
J. Softw. Evol. Process.2
2017 The Evaluation of an Approach for Automatic Generated Documentation
abstract
Two studies are conducted to evaluate an approach to automatically generate natural language documentation summaries for C++ methods. The documentation approach relies on a method's stereotype information. First, each method is automatically assigned a stereotype(s) based on static analysis and a set of heuristics. Then, the approach uses the stereotype information, static analysis, and predefined templates to generate a natural-language summary/documentation for each method. This documentation is automatically added to the code base as a comment for each method. The result of the first study reveals that the generated documentation is accurate, does not include unnecessary information, and does a reasonable job describing what the method does. Based on statistical analysis of the second study, the most important part of the documentation is the short description as it describes the intended behavior of a method.
Nahla J. Abid, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSME4
2017 An Empirical Study of Debugging Patterns Among Novices Programmers
abstract
Students taking introductory computer science courses often have difficulty with the debugging process. This work investigates a number of different logical errors that novice programmers encounter and the associated debugging behaviors. Data is collected and analyzed data in two different experiments from 142 subjects. The results show some errors are more difficult than others. Different types of bugs and novices' debugging behaviors are identified. Years of experience showed a significant role in the process of debugging in terms of correctness level and time required for debugging
Basma S. Alqadi, Jonathan I. Maletic
SIGCSE2
2017 srcQL: A syntax-aware query language for source code
abstract
A tool and domain specific language for querying source code is introduced and demonstrated. The tool, srcQL, allows for the querying of source code using the syntax of the language to identify patterns within source code documents. srcQL is built upon srcML, a widely used XML representation of source code, to identify the syntactic contexts being queried. srcML inserts XML tags into the source code to mark syntactic constructs. srcQL uses a combination of XPath on srcML, regular expressions, and syntactic patterns within a query. The syntactic patterns are snippets of source code that supports the use of logical variables which are unified during the query process. This allows for very complex patterns to be easily formulated and queried. The tool is implemented (in C++) and a number of queries are presented to demonstrate the approach. srcQL currently supports C++ and scales to large systems.
Brian Bartman, Christian D. Newman, Michael L. Collard, Jonathan I. Maletic
SANER4
2017 Lexical categories for source code identifiers
abstract
A set of lexical categories, analogous to part-of-speech categories for English prose, is defined for source-code identifiers. The lexical category for an identifier is determined from its declaration in the source code, syntactic meaning in the programming language, and static program analysis. Current techniques for assigning lexical categories to identifiers use natural-language part-of-speech taggers. However, these NLP approaches assign lexical tags based on how terms are used in English prose. The approach taken here differs in that it uses only source code to determine the lexical category. The approach assigns a lexical category to each identifier and stores this information along with each declaration. srcML is used as the infrastructure to implement the approach and so the lexical information is stored directly in the srcML markup as an additional XML element for each identifier. These lexical-category annotations can then be later used by tools that automatically generate such things as code summarization or documentation. The approach is applied to 50 open source projects and the soundness of the defined lexical categories evaluated. The evaluation shows that at every level of minimum support tested, categorization is consistent at least 79% of the time with an overall consistency (across all supports) of at least 88%. The categories reveal a correlation between how an identifier is named and how it is declared. This provides a syntax-oriented view (as opposed to English part-of-speech view) of developer intent of identifiers.
Christian D. Newman, Reem S. Alsuhaibani, Michael L. Collard, Jonathan I. Maletic
SANER4
2017 Simplifying the construction of source code transformations via automatic syntactic restructurings
abstract
Abstract A set of restructurings to systematically normalize selective syntax in C++ is presented. The objective is to convert variations in syntax of specific portions of code into a single form to simplify the construction of large, complex program transformation rules. Current approaches to constructing transformations require developers to account for a large number of syntactic cases, many of which are syntactically different but semantically equivalent. The work identifies classes of such syntactic variations and presents normalizing restructurings to simplify each variation to a single, consistent syntactic form. The normalizing restructurings for C++ are presented and applied to two open source systems for evaluation. The evaluation uses the system's test cases to validate that the normalizing restructurings do not affect the systems' tested behavior. In addition, a set of example transformations that benefit from the prior application of normalizing restructurings are presented along with a small survey to assess the effect of the readability of the resultant code.
Christian D. Newman, Brian Bartman, Michael L. Collard, Jonathan I. Maletic
J. Softw. Evol. Process.4
2016 srcML 1.0: Explore, Analyze, and Manipulate Source Code
abstract
Summary form only given. This technology briefing is intended for those interested in constructing custom software analysis and manipulation tools to support research or commercial applications. srcML (srcML.org) is an infrastructure consisting of an XML representation for C/C++/C#/Java source code along with efficient parsing technology to convert source code to-and-from the srcML format. The briefing describes srcML, the toolkit, and the application of XPath and XSLT to query and modify source code. Additionally, a short tutorial of how to use srcML and XML tools to construct custom analysis and manipulation tools will be conducted.
Michael L. Collard, Jonathan I. Maletic
ICSME2
2016 A Tool for Efficiently Reverse Engineering Accurate UML Class Diagrams
abstract
A tool that reverse engineers UML class diagrams from C++ source code is presented. The tool takes srcML as input and produces yUML as output. srcML is an XML representation of the abstract syntactic information of source code. The srcML parser (srcML.org) is highly scalable, efficient, and robust. yUML is a textual format for UML class diagrams that can be easily rendered into a graphical diagram via a web service (yUML.me) or a tool such as Graphvis. The approach utilizes efficient SAX (Simple API for XML) parsing to collect the information needed to construct the class diagram. Currently it supports the following UML features: differentiating between class, data type, or interface, identifying design level attributes, multiplicity and type, determining parameter direction, and identification of the relationships aggregation, composition, generalization, and realization. The tool produces yUML for all of Calligra (~1,144KLOC) in under 20 seconds (including translation into srcML). The tool is open source under a GPL license and available for download at srcML.org.
Michael John Decker, Kyle Swartz, Michael L. Collard, Jonathan I. Maletic
ICSME4
2016 Recovering Commit Branch of Origin from GitHub Repositories
abstract
An approach to automatically recover the name of the branch where a given commit is originally made within a GitHub repository is presented and evaluated. This is a difficult task because in Git, the commit object does not store the name of the branch when it is created. Here this is termed the commit's branch of origin. Developers typically use branches in Git to group sets of changes that are related by task or concern. The approach recovers the branch of origin only within the scope of a single repository. The recovery process first uses Git's default merge commit messages and then examines the relationships between neighboring commits. The evaluation includes a simulation, an empirical examination of 40 repositories of open-source systems, and a manual verification. The evaluations show that the average accuracy exceeds 97% of all commits and the average precision exceeds 80%.
Heather M. Guarnera, Drew T. Guarnera, Michael L. Collard, Jonathan I. Maletic
ICSME4
2016 srcType: A Tool for Efficient Static Type Resolution
abstract
An efficient, static type resolution tool is presented. The tool is implemented on top of srcML, an XML representation of source code and abstract syntax. The approach computes the type of every identifier (i.e., function names and variable names) within the provided body of code. The result is a dictionary that can be used to lookup the type of each name. Type information includes metadata such as constness, class membership, aliasing, line number, file, and namespace. The approach is highly scalable and can generate a dictionary for Linux (13 MLOC) in less than 7 minutes. The tool is open source under a GPL license and available for download at srcML.org.
Christian D. Newman, Jonathan I. Maletic, Michael L. Collard
ICSME2
2016 iTrace: Overcoming the Limitations of Short Code Examples in Eye Tracking Experiments
abstract
Summary form only given. Eye trackers are being used by software engineering researchers to study how developers work. In this technical briefing, we give an overview of eye tracking and how it can help researchers to conduct their own studies. Eye tracking studies are done on a single screen of text and there is no support for scrolling or switching between files. This scenario is impractical to study developers as they actually work on large software artifacts. To overcome this an Eclipse plugin, iTrace, is introduced that monitors developers eye movements even in the presence of scrolling and file switching within an IDE. In addition, it automatically maps the eye gaze to source code elements. Existing work using iTrace is presented followed by a scenario of how to setup and run an eye tracking study. Data filtering, data cleaning, and data analysis are also discussed.
Bonita Sharif, Jonathan I. Maletic
ICSME2
2016 Studying developer gaze to empower software engineering research and practice
abstract
A new research paradigm is proposed that leverages developer eye gaze to improve the state of the art in software engineering research and practice. The vision of this new paradigm for use on software engineering tasks such as code summarization, code recommendations, prediction, and continuous traceability is described. Based on this new paradigm, it is foreseen that new benchmarks will emerge based on developer gaze. The research borrows from cognitive psychology, artificial intelligence, information retrieval, and data mining. It is hypothesized that new algorithms will be discovered that work with eye gaze data to help improve current IDEs, thus improving developer productivity. Conducting empirical studies using an eye tracker will lead to inventing, evaluating, and applying innovative methods and tools that use eye gaze to support the developer. The implications and challenges of this paradigm for future software engineering research is discussed.
Bonita Sharif, Benjamin Clark, Jonathan I. Maletic
SIGSOFT FSE3
2016 An empirical examination of the prevalence of inhibitors to the parallelizability of open source software systems
Saleh M. Alnaeli, Jonathan I. Maletic, Michael L. Collard
Empir. Softw. Eng.2
2015 Exploration, Analysis, and Manipulation of Source Code Using srcML
abstract
This technology briefing is intended for those interested in constructing custom software analysis and manipulation tools to support research or commercial applications. srcML (srcML.org) is an infrastructure consisting of an XML representation for C/C++/C#/Java source code along with efficient parsing technology to convert source code to-and-from the srcML format. The briefing describes srcML, the toolkit, and the application of XPath and XSLT to query and modify source code. Additionally, a hands-on tutorial of how to use srcML and XML tools to construct custom analysis and manipulation tools will be conducted.
Jonathan I. Maletic, Michael L. Collard
ICSE (2)1
2015 Using stereotypes in the automatic generation of natural language summaries for C++ methods
abstract
An approach to automatically generate natural language documentation summaries for C++ methods is presented. The approach uses prior work by the authors on stereotyping methods along with the source code analysis framework srcML. First, each method is automatically assigned a stereotype(s) based on static analysis and a set of heuristics. Then, the approach uses the stereotype information, static analysis, and predefined templates to generate a natural-language summary for each method. This summary is automatically added to the code base as a comment for each method. The predefined templates are designed to produce a generic summary for specific method stereotypes. Static analysis is used to extract internal details about the method (e.g., parameters, local variables, calls, etc.). This information is used to specialize the generated summaries.
Nahla J. Abid, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSME4
2015 Guest editorial: special section on software maintenance and evolution
Massimiliano Di Penta, Jonathan I. Maletic
Empir. Softw. Eng.2
2014 A Slice-Based Estimation Approach for Maintenance Effort
abstract
Program slicing is used as a basis for an approach to estimate maintenance effort. A case study of the GNU Linux kernel with over 900 versions spanning 17 years of history is presented. For each version a system dictionary is built using a lightweight slicing approach and encodes the forward decomposition static slice profiles for all variables in all the files in the system. Changes to the system are then modeled at the behavioral level using the difference between the system dictionaries of two versions. The three different granularities of slice (i.e., line, function, and file) are analyzed. We use a direct extension of srcML to represent computed change information. The retrieved information reflects the fact that additional knowledge of the differences can be automatically derived to help maintainers understand code changes. We consider the hypotheses: (1) The structured format helps create traceability links between the changes and other software artifacts. (2) This model is predictive of maintenance effort. The results demonstrate that the approach accurately predicts effort in a scalable manner.
Hakam W. Alomari, Michael L. Collard, Jonathan I. Maletic
ICSME3
2014 srcSlice: very efficient and scalable forward static slicing
abstract
ABSTRACT A highly efficient lightweight forward static slicing approach is presented and evaluated. The approach does not compute the program/system dependence graph but instead dependence and control information is computed as needed while computing the slice on a variable. The result is a list of line numbers, dependent variables, aliases, and function calls that are part of the slice for all variables (both local and global) for the entire system. The method is implemented as a tool, calledsrcSlice, on top ofsrcML, an XML representation of source code. The approach is highly scalable and can generate the slices for all variables of the Linux kernel in approximately 20 min on a typical desktop. Benchmark results are compared with theCodeSurferslicing tool from GrammaTech Inc., and the approach compares well with regard to accuracy of slices. Copyright © 2014 John Wiley & Sons, Ltd.
Hakam W. Alomari, Michael L. Collard, Jonathan I. Maletic, Nouh Alhindawi, Omar Meqdadi
J. Softw. Evol. Process.3
2013 Improving Feature Location by Enhancing Source Code with Stereotypes
abstract
A novel approach to improve feature location by enhancing the corpus (i.e., source code) with static information is presented. An information retrieval method, namely Latent Semantic Indexing (LSI), is used for feature location. Adding stereotype information to each method/function enhances the corpus. Stereotypes are terms that describe the abstract role of a method, for example get, set, and predicate are well-known method stereotypes. Each method in the system is automatically stereotyped via a static-analysis approach. Experimental comparisons of using LSI for feature location with, and without, stereotype information are conducted on a set of open-source systems. The results show that the added information improves the recall and precision in the context of feature location. Moreover, the use of stereotype information decreases the total effort that a developer would need to expend to locate relevant methods of the feature.
Nouh Alhindawi, Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM4
2013 srcML: An Infrastructure for the Exploration, Analysis, and Manipulation of Source Code: A Tool Demonstration
abstract
SrcML is an XML representation for C/C++/Java source code that forms a platform for the efficient exploration, analysis, and manipulation of large software projects. The lightweight format allows for round-trip transformation from source to srcML and back to source with no loss of information or formatting. The srcML toolkit consists of the src2srcml tool for robust translation to the srcML format and the srcml2src tool for querying via XPath, and transformation via XSLT. In this demonstration a guide of these features is provided along with the use of XPath for constructing source-code queries and XSLT for conducting simple transformations.
Michael L. Collard, Michael John Decker, Jonathan I. Maletic
ICSM3
2013 Towards Understanding Large-Scale Adaptive Changes from Version Histories
abstract
A case study of three open source systems undergoing large adaptive maintenance tasks is presented. The adaptive maintenance task involves migrating each system to a new version of a third party API. The changes to support the migration were spread out over multiple years for each system. The first two systems are both part of KDE, namely KOffice and Extragear/graphics. The adaptive maintenance task, for both systems, involves migrating to a new version of Qt. The third system is OpenSceneGraph that underwent a migration to a new version of OpenGL. The case study involves sifting through tens of thousands of commits to identify only those commits involved in the specific adaptive maintenance task. The object is to develop a data set that will be used for developing automated methods to identify/characterize adaptive maintenance commits.
Omar Meqdadi, Nouh Alhindawi, Michael L. Collard, Jonathan I. Maletic
ICSM4
2013 A preliminary investigation of using age and distance measures in the detection of evolutionary couplings
abstract
An initial study of using two measures to improve the accuracy of evolutionary couplings uncovered from version history is presented. Two measures, namely the age of a pattern and the distance among items within a pattern, are defined and used with the traditional methods for computing evolutionary couplings. The goal is to reduce the number of false positives (i.e., inaccurate or irrelevant claims of coupling). Initial observations are presented that lend evidence that these measures may have the potential to improve the results of computing evolutionary couplings.
Abdulkareem Alali, Brian Bartman, Christian D. Newman, Jonathan I. Maletic
MSR4
2013 The impact of identifier style on effort and comprehension
Dave W. Binkley, Marcia Davis, Dawn J. Lawrie, Jonathan I. Maletic, Christopher Morrell, Bonita Sharif
Empir. Softw. Eng.4
2013 Emulating C++0x concepts
Andrew M. Sutton, Jonathan I. Maletic
Sci. Comput. Program.2
2012 An eye-tracking study on the role of scan time in finding source code defects
abstract
An eye-tracking study is presented that investigates how individuals find defects in source code. This work partially replicates a previous eye-tracking study by Uwano et al. [2006]. In the Uwano study, eye movements are used to characterize the performance of individuals in reviewing source code. Their analysis showed that subjects who did not spend enough time initially scanning the code tend to take more time finding defects. The study here follows a similar setup with added eye-tracking measures and analyses on effectiveness and efficiency of finding defects with respect to eye gaze. The subject pool is larger and is comprised of a varied skill level. Results indicate that scanning significantly correlates with defect detection time as well as visual effort on relevant defect lines. Results of the study are compared and contrasted to the Uwano study.
Bonita Sharif, Michael Falcone, Jonathan I. Maletic
ETRA3
2012 TraceLab: An experimental workbench for equipping researchers to innovate, synthesize, and comparatively evaluate traceability solutions
abstract
TraceLab is designed to empower future traceability research, through facilitating innovation and creativity, increasing collaboration between researchers, decreasing the startup costs and effort of new traceability research projects, and fostering technology transfer. To this end, it provides an experimental environment in which researchers can design and execute experiments in TraceLab's visual modeling environment using a library of reusable and user-defined components. TraceLab fosters research competitions by allowing researchers or industrial sponsors to launch research contests intended to focus attention on compelling traceability challenges. Contests are centered around specific traceability tasks, performed on publicly available datasets, and are evaluated using standard metrics incorporated into reusable TraceLab components. TraceLab has been released in beta-test mode to researchers at seven universities, and will be publicly released via CoEST.org in the summer of 2012. Furthermore, by late 2012 TraceLab's source code will be released as open source software, licensed under GPL. TraceLab currently runs on Windows but is designed with cross platforming issues in mind to allow easy ports to Unix and Mac environments.
Ed Keenan, Adam Czauderna, Greg Leach, Jane Cleland-Huang, Yonghee Shin, Evan Moritz, Malcom Gethers, Denys Poshyvanyk, Jonathan I. Maletic, Jane Huffman Hayes, Alex Dekhtyar, Daria Manukian, Shervin Hossein, Derek Hearn
ICSE9
2011 Using stereotypes to help characterize commits
abstract
Individual commits to a version control system are automatically characterized based on the stereotypes of added and deleted methods. The stereotype of each method is automatically reverse engineered using a previously defined taxonomy. Method stereotypes reflect intrinsic atomic behavior of a method and its role in the class. The stereotypes of the added and deleted methods form a descriptors are then used to categorize commits, into types, based on the impact of the changes to a class (or classes). The goal is to gain a better understanding of the design changes to a system over its history and provide a means for documenting the commit.
Natalia Dragan, Michael L. Collard, Maen Hammad, Jonathan I. Maletic
ICSM4
2011 Lightweight Transformation and Fact Extraction with the srcML Toolkit
abstract
The srcML toolkit for lightweight transformation and fact-extraction of source code is described. srcML is an XML format for C/C++/Java source code. The open source toolkit that includes the source-to-srcML and srcML-to-source translators for round-trip reverse engineering is freely available. The direct use of XPath and XSLT is supported, an archive format for large projects is included, and a rich set of input and output formats through a command-line interface is available. Applying transformations and formulating queries using srcML is very convenient. Application use-cases of transformations and fact-extraction are shown and demonstrated to be practical and scalable.
Michael L. Collard, Michael John Decker, Jonathan I. Maletic
SCAM3
2011 Automatically identifying changes that impact code-to-design traceability during evolution
Maen Hammad, Michael L. Collard, Jonathan I. Maletic
Softw. Qual. J.3
2010 The Effects of Layout on Detecting the Role of Design Patterns
abstract
A controlled experiment investigating the effect layout has on how students identify design pattern roles in UML class diagrams is presented. Two layout schemes, multi-cluster and orthogonal, are compared with respect to three open source systems and four design patterns. Seventeen students were asked a series of eight design pattern role detection (comprehension) questions for each layout, followed by eight preference rating questions. Results indicate a significant improvement in role detection accuracy with the multi-cluster layout for the strategy pattern and a significant improvement in detection time with the multi-cluster layout for all four patterns. Preference ratings significantly favored the multi-cluster layout for pattern role detection ease. These results can be used to help improve the teaching of design patterns.
Bonita Sharif, Jonathan I. Maletic
CSEE&T2
2010 A lightweight transformational approach to support large scale adaptive changes
abstract
An approach to automate adaptive maintenance changes on large-scale software systems is presented. This approach uses lightweight parsing and lightweight on-the-fly static analysis to support transformations that make corrections to source code in response to adaptive maintenance changes, such as platform changes. SrcML, an XML source code representation, is used and transformations can be performed using either XSLT or LINQ. A number of specific adaptive changes are presented, based on recent adaptive maintenance needs from products at ABB Inc. The transformations are described in detail and then demonstrated on a number of examples from the production systems. The results are compared with manual adaptive changes that were done by professional developers. The approach performed better than the manual changes, as it successfully transformed instances missed by the developers while not missing any instances itself. The work demonstrates that this lightweight approach is both efficient and accurate with an overall cost savings in development time and effort.
Michael L. Collard, Jonathan I. Maletic, Brian P. Robinson
ICSM2
2010 Automatic identification of class stereotypes
abstract
An approach is presented to automatically determine a class's stereotype. The stereotype is based on the frequency and distribution of method stereotypes in the class. Method stereotypes are automatically determined using a defined taxonomy given in previous work. The stereotypes, boundary, control and entity are used as a basis but refined based on an empirical investigation of 21 systems. A number of heuristics, derived from empirical evidence, are used to determine a class's stereotype. For example, the prominence of certain types of methods can indicate a class's main role. The approach is applied to five open source systems and evaluated. The results show that 95% of the classes are stereotyped by the approach. Additionally, developers (via manual inspection) agreed with the approach's results.
Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM3
2010 An eye tracking study on the effects of layout in understanding the role of design patterns
abstract
The effect of layout in the comprehension of design pattern roles in UML class diagrams is assessed. This work replicates and extends a previous study using questionnaires but uses an eye tracker to gather additional data. The purpose of the replication is to gather more insight into the eye gaze behavior not evident from questionnaire-based methods. Similarities and differences between the studies are presented. Four design patterns are examined in two layout schemes in the context of three open source systems. Fifteen participants answered a series of eight design pattern role detection questions. Results show a significant improvement in role detection accuracy and visual effort with a certain layout for the Strategy and Observer patterns and a significant improvement in role detection time for all four patterns. Eye gaze data indicates classes participating in a design pattern act like visual beacons when they are in close physical proximity and follow the canonical layout, even though they violate some general graph aesthetics.
Bonita Sharif, Jonathan I. Maletic
ICSM2
2010 Measuring Class Importance in the Context of Design Evolution
abstract
A measure of how a class is impacted during design evolution is presented. The history of design changes that involve a given class is the basis for the measure. Classes that are often impacted by design changes are branded as important to the design of the system. Identifying these important classes helps reveal what parts of the system are regularly evolved (e.g., specific features or cross-cutting concerns). The design importance of a class is measured as the number of commits that impact both the design and the class. This is also measured for sets of classes that collaborate to realize a feature or concept in the system. Collaborating classes are identified using itemset mining on commits that impact the design. A small study is presented on two open source projects to illustrate the approach.
Maen Hammad, Michael L. Collard, Jonathan I. Maletic
ICPC3
2010 An Eye Tracking Study on camelCase and under_score Identifier Styles
abstract
An empirical study to determine if identifier-naming conventions (i.e., camelCase and under_score) affect code comprehension is presented. An eye tracker is used to capture quantitative data from human subjects during an experiment. The intent of this study is to replicate a previous study published at ICPC 2009 (Binkley et al.) that used a timed response test method to acquire data. The use of eye-tracking equipment gives additional insight and overcomes some limitations of traditional data gathering techniques. Similarities and differences between the two studies are discussed. One main difference is that subjects were trained mainly in the underscore style and were all programmers. While results indicate no difference in accuracy between the two styles, subjects recognize identifiers in the underscore style more quickly.
Bonita Sharif, Jonathan I. Maletic
ICPC2
2010 Identification of Idiom Usage in C++ Generic Libraries
abstract
A tool supporting the automatic identification of programming idioms specific to the construction of C++ generic libraries is presented. The goal is to assist developers in understanding the complex syntactic elements of these libraries. Large C++ generic libraries are notorious for being extremely difficult to comprehend due to their use of advanced language features and idiomatic nature. To facilitate automated identification, the idioms are equated to micropatterns, which can be evaluated by a fact extractor. These micropattern instances act as beacons for the idioms being identified. The method is applied to study a number of widely used open source C++ generic libraries.
Andrew M. Sutton, Ryan Holeman, Jonathan I. Maletic
ICPC3
2009 Using method stereotype distribution as a signature descriptor for software systems
abstract
Method stereotype distribution is used as a signature for software systems. The stereotype for each method is determined using a presented taxonomy. The counts of the different stereotypes form a signature of the system. Determining method stereotypes is done automatically and is based on language (C++) features, idioms, and the main role (purpose) of a method. The intent is to use the distribution of method stereotype is an indicator of system architecture.
Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM3
2009 Abstracting the template instantiation relation in C++
abstract
A source code model that supports the static analysis of C++ templates and template metaprograms is presented. Analogous to techniques for object-oriented and procedural software (e.g., the abstraction of call graphs, inheritance hierarchies, etc.), this model provides a basis for maintenance concerns such as program comprehension, fact extraction, and impact analysis of generic code. The source code model is used to derive the template instantiation graph, and potential applications of this model discussed. An application to reverse engineer this model from source code is described.
Andrew M. Sutton, Ryan Holeman, Jonathan I. Maletic
ICSM3
2009 Working session: Using eye-tracking to understand program comprehension
abstract
The working session focuses on the use of eye-tracking technology to assess, understand, and evaluate tools and techniques for program comprehension. An introduction to the technology and tools of eye-tracking will be presented. A discussion of how these tools augment existing evaluation mechanism in the context of program comprehension will follow. Research directions and open problems will be a main topic.
Yann-Gaël Guéhéneuc, Huzefa H. Kagdi, Jonathan I. Maletic
ICPC3
2009 Automatically identifying changes that impact code-to-design traceability
abstract
An approach is presented that automatically determines if a given source code change impacts the design (i.e., UML class diagram) of the system. This allows code-to-design traceability to be consistently maintained as the source code evolves. The approach uses lightweight analysis and syntactic differencing of the source code changes to determine if the change alters the class diagram in the context of abstract design. The intent is to support both the simultaneous updating of design documents with code changes and bringing old design documents up to date with current code given the change history. An efficient tool was developed to support the approach and is applied to an open source system (i.e., HippoDraw). The results are evaluated and compared against manual inspection by human experts. The tool performs better than (error prone) manual inspection.
Maen Hammad, Michael L. Collard, Jonathan I. Maletic
ICPC3
2009 An empirical study on the comprehension of stereotyped UML class diagram layouts
abstract
An empirical study is presented that investigates how stereotype based layouts impact the comprehension of UML class diagrams. This work replicates a previous study using eye-tracking equipment but uses online questionnaires instead. Subjects were given two types of tasks: one addressing UML syntax and the other addressing software design. Three different layout strategies are compared. Along with general aesthetics, the layouts are primarily organized by class stereotypes of control, boundary, and entity. A confidence value for each question was collected from the subjects to help validate the categorization of subjects. Results of the study are compared and contrasted to the eye-tracking study done with the same tasks and layouts. Results show a significant improvement in performance in both types of tasks with the multi-cluster stereotyped layouts.
Bonita Sharif, Jonathan I. Maletic
ICPC2
2009 Introduction to the WCRE 2007 special issue
Massimiliano Di Penta, Jonathan I. Maletic
Softw. Qual. J.2
2008 Who can help me with this source code change?
abstract
An approach to recommend a ranked list of developers to assist in performing software changes to a particular file is presented. The ranking is based on change expertise, experience, and contributions of developers, as derived from the analysis of the previous commits involving the specific file in question. The commits are obtained from a software systempsilas version control repositories (e.g., Subversion). The basic premise is that a developer who has substantially contributed changes to specific files in the past is likely to best assist for their current or future change. Evaluation of the approach on a number of open source systems such as koffice, Apache httpd, and GNU gcc is also presented. The results show that the accuracy of the correctly recommended developers is between 43% and 82%. New developers to a long-lived software project, or project managers, can use this approach to assist them in undertaking maintenance tasks, e.g., bug fix or adding a new feature. The approach can be realized as a plug-in to development environments such as Eclipse.
Huzefa H. Kagdi, Maen Hammad, Jonathan I. Maletic
ICSM3
2008 Automatically identifying C++0x concepts in function templates
abstract
An automated approach to the identification of C++0x concepts in function templates is described. Concepts are part of a new language feature appearing in the next standard for C++ (i.e., C++0x). Concept identification is the enumeration of constraints on the sets of types over which templates can be instantiated. The approach analyzes template source code and computes a set of viable concept instances describing the implied data abstraction of the template parameters. The approach is evaluated on generic algorithms defined in the C++ Standard Template Library (STL). The evaluation demonstrates the effectiveness of the approach. The approach can be used to assist in reengineering existing generic libraries to C++0x. Additionally, it has the potential to assist in the validation of concept hierarchies and interface definition in generic libraries.
Andrew M. Sutton, Jonathan I. Maletic
ICSM2
2008 What's a Typical Commit? A Characterization of Open Source Software Repositories
abstract
The research examines the version histories of nine open source software systems to uncover trends and characteristics of how developers commit source code to version control systems (e.g., subversion). The goal is to characterize what a typical or normal commit looks like with respect to the number of files, number of lines, and number of hunks committed together. The results of these three characteristics are presented and the commits are categorized from extra small to extra large. The findings show that approximately 75% of commits are quite small for the systems examined along all three characteristics. Additionally, the commit messages are examined along with the characteristics. The most common words are extracted from the commit messages and correlated with the size categories of the commits. It is observed that sized categories can be indicative of the types of maintenance activities being performed.
Abdulkareem Alali, Huzefa H. Kagdi, Jonathan I. Maletic
ICPC3
2007 How We Manage Portability and Configuration with the C Preprocessor
abstract
An in-depth investigation of C preprocessor usage for portability and configuration management is presented. Three heavily-ported and widely used C++ libraries are examined. A core set of header files responsible for configuration management is identified in each system. Then macro usage is extracted and analyzed both manually and with the help of program analysis tools. The configuration structure of each library is discussed in details and commonalities between the systems, including conventions and patterns are discussed. A common configuration architecture for managing portability concerns is derived and presented.
Andrew M. Sutton, Jonathan I. Maletic
ICSM2
2007 Mining Software Repositories for Traceability Links
abstract
An approach to recover/discover traceability links between software artifacts via the examination of a software system's version history is presented. A heuristic-based approach that uses sequential-pattern mining is applied to the commits in software repositories for uncovering highly frequent co-changing sets of artifacts (e.g., source code and documentation). If different types of files are committed together with high frequency then there is a high probability that they have a traceability link between them. The approach is evaluated on a number of versions of the open source system KDE. As a validation step, the discovered links are used to predict similar changes in the newer versions of the same system. The results show highly precision predictions of certain types of traceability links.
Huzefa H. Kagdi, Jonathan I. Maletic, Bonita Sharif
ICPC2
2007 Assessing the Comprehension of UML Class Diagrams via Eye Tracking
abstract
Eye-tracking equipment is used to assess how well a subject comprehends UML class diagrams. The results of a study are presented in which eye movements are captured in a non-obtrusive manner as users performed various comprehension tasks on UML class diagrams. The goal of the study is to identify specific characteristics of UML class diagrams, such as layout, color, and stereotype usage that are most effective for supporting a given task. Results indicate subjects have a variation in the eye movements (i.e., how the subjects navigate the diagram) depending on their UML expertise and software-design ability to solve the given task. Layouts with additional semantic information about the design were found to be most effective and the use of class stereotypes seems to play a substantial role in comprehension of these diagrams.
Shehnaaz Yusuf, Huzefa H. Kagdi, Jonathan I. Maletic
ICPC3
2007 An approach to mining call-usage patternswith syntactic context
abstract
An approach to mine frequently appearing ordered sets of function-call usages, taking into account their proximal control constructs (e.g., if-statements), in the source code is presented. These ordered sets are termed as call-usage patterns. Additionally, variant usages, such as those with missing or out of order calls, are automatically identified along with their specific contextual location. The approach uses lightweight source code analysis and frequent sequential pattern mining. The hypothesis is that these call-usage patterns embody latent programming rules that developers commonly reuse, for example standard usages of API calls. The variants are an indicator of future changes such as the elimination of non-standard usages and/or bugs
Huzefa H. Kagdi, Michael L. Collard, Jonathan I. Maletic
ASE3
2007 Recovering UML class models from C++: A detailed explanation
Andrew M. Sutton, Jonathan I. Maletic
Inf. Softw. Technol.2
2007 A survey and taxonomy of approaches for mining software repositories in the context of software evolution
abstract
Abstract A comprehensive literature survey on approaches for mining software repositories (MSR) in the context of software evolution is presented. In particular, this survey deals with those investigations that examine multiple versions of software artifacts or other temporal information. A taxonomy is derived from the analysis of this literature and presents the work via four dimensions: the type of software repositories mined (what), the purpose (why), the adopted/invented methodology used (how), and the evaluation method (quality). The taxonomy is demonstrated to be expressive (i.e., capable of representing a wide spectrum of MSR investigations) and effective (i.e., facilitates similarities and comparisons of MSR investigations). Lastly, a number of open research issues in MSR that require further investigation are identified. Copyright © 2007 John Wiley & Sons, Ltd.
Huzefa H. Kagdi, Michael L. Collard, Jonathan I. Maletic
J. Softw. Maintenance Res. Pract.3
2007 Mining evolutionary dependencies from web-localization repositories
abstract
Abstract An approach to mining repositories of web‐based user documentation for patterns of evolutionary change in the context of internationalization and localization is presented. Localized web documents that are frequently co‐changed (i.e., an evolutionary dependency) during the natural language translation process are uncovered to support the future evolution of the system. A sequential‐pattern mining technique is used to uncover patterns from version histories. Characteristics of the uncovered patterns such as size, frequency, and occurrence within a single natural language or across multiple languages are discussed. Such patterns help provide an insight into the effort required in retranslation due to a change in the documentation. The approach is validated on the open source K Desktop Environment (KDE) system. KDE maintains documentation for over 50 different natural languages and presents a prime example of the problem. The technique accurately predicts which documents in KDE are retranslated or updated in future versions. Copyright © 2007 John Wiley & Sons, Ltd.
Huzefa H. Kagdi, Jonathan I. Maletic
J. Softw. Maintenance Res. Pract.2
2006 Reverse Engineering Method Stereotypes
abstract
An approach to automatically identify the stereotypes of all the methods in an entire system is presented. A taxonomy for object-oriented class method stereotypes is given that unifies and extends the existing literature to address gaps and deficiencies. Based on this taxonomy, a set of definitions is given and method stereotypes are reverse engineered using lightweight static program analysis. Classification is done solely by programming language structures and idioms, in this case C++. The approach is used to automatically re-document each method by annotating the original source code with the stereotype information. A demonstration of the accuracy and scalability of the approach is given
Natalia Dragan, Michael L. Collard, Jonathan I. Maletic
ICSM3
2006 Applying Dynamic Change Impact Analysis in Component-based Architecture Design
abstract
Change impact analysis plays an important role in maintenance and evolution of component-based software architecture. Viewing component replacement as a change to composition-based software architecture, this paper proposes a component interaction trace based approach to support dynamic change impact analysis at software architecture level. Given an architectural change, our approach determines the architecture elements causing the change and impacted by the change. Firstly, component-based software architecture and component interaction trace are defined. An algorithm for generating component interaction trace from static structure model of software architecture and UML sequence diagram is provided. Secondly, the taxonomy of changes on composition-based software architecture is presented, according to which a set of impact rules are suggested to determine the transfer of the changes in component and among components. Thirdly, by performing slicing on component interaction traces according to impact rules, the impact analysis results are obtained. Finally, the architecture design of SOCIAT, a tool supporting our approach, is developed and explained.
Tie Feng, Jonathan I. Maletic
SNPD2
2006 Guest editorial
James R. Cordy, Harald C. Gall, Jonathan I. Maletic
Softw. Qual. J.3
2005 Hybridizing evolutionary algorithms and clustering algorithms to find source-code clones
abstract
This paper presents a hybrid approach to detect source-code clones that combines evolutionary algorithms and clustering. A case-study is conducted on a small C++ code base. The preliminary investigation indicates that such an approach is effective in detecting groups of source-code clones.
Andrew M. Sutton, Huzefa H. Kagdi, Jonathan I. Maletic, L. Gwenn Volkert
GECCO3
2005 Context-Free Slicing of UML Class Models
abstract
The concept of model slicing is introduced as a means to support maintenance through the understanding, querying, and analysis of large UML models. The specific models being examined are class models as defined in the Unified Modeling Language (UML). Model slicing is analogous to classical program slicing. Since UML class models do not explicitly embody any behavioral aspect by themselves, models slices are computed in a context-free manner. The paper defines and formalizes the concept of context-free model slicing. A concrete application of model slicing in software maintenance is presented to support the usefulness and validity of the method.
Huzefa H. Kagdi, Jonathan I. Maletic, Andrew M. Sutton
ICSM2
2005 3rd international workshop on traceability in emerging forms of software engineering (TEFSE 2005)
abstract
Establishing and maintaining traceability links and consistency between software artifacts produced or modified in the software life-cycle are costly and tedious activities that are crucial but frequently neglected in practice. Traceability between the free text documentation associated with the development and maintenance cycle of a software system and its source code are crucial in a number of tasks such as program comprehension, software maintenance, and software verification & validation. Finally, maintaining traceability links between subsequent releases of a software system is important for evaluating relative source code deltas, highlighting effort/code variation inconsistencies, and assessing the change history. The main theme of the workshop is focused on understanding and defining the foundations for consistency and change management of software systems within the scope of artifact-to-artifact (model-to-model) traceability.The workshop will address the following issues:A formal definition of model to model traceabilityTraceability between artifacts and processesThe semantics of traceability linksRecovery of traceability linksVisualization of traceability linksInteroperable approaches to support traceabilityTraceability in emerging forms of software engineering including production lines, frameworks, components, etc..The goals of the workshop are to:Broaden awareness within the software engineering community of the potential for the application of traceabilityFacilitate the exchange of ideas and interaction between international researchersDefine open research problems faced in realizing usable approaches for traceabilityConstruct a foundation of materials for future research on traceability .For more information please visit the workshop web site is: http://re.cs.depaul.edu/tefse05/. The workshop proceedings are available through the ACM digital library.
Jonathan I. Maletic, Giuliano Antoniol, Jane Cleland-Huang, Jane Huffman Hayes
ASE1
2005 Recovery of Traceability Links between Software Documentation and Source Code
abstract
An approach for the semi-automated recovery of traceability links between software documentation and source code is presented. The methodology is based on the application of information retrieval techniques to extract and analyze the semantic information from the source code and associated documentation. A semi-automatic process is defined based on the proposed methodology. The paper advocates the use of latent semantic indexing (LSI) as the supporting information retrieval technique. Two case studies using existing software are presented comparing this approach with others. The case studies show positive results for the proposed approach, especially considering the flexibility of the methods used.
Andrian Marcus, Jonathan I. Maletic, Andrey Sergeyev
Int. J. Softw. Eng. Knowl. Eng.2
2004 Supporting Source Code Difference Analysis
abstract
The paper describes an approach to easily conduct analysis of source-code differences. The approach is termed meta-differencing to reflect the fact that additional knowledge of the differences can be automatically derived. Meta-differencing is supported by an underlying source-code representation developed by the authors. The representation, srcML, is an XML format that explicitly embeds abstract syntax within the source code while preserving the documentary structure as dictated by the developer. XML tools are leveraged together with standard differencing utilities (i.e., diff,) to generate a meta-difference. The meta-difference is also represented in an XML format called srcDiff. The meta-difference contains specific syntactic information regarding the source-code changes. In turn this can be queried and searched with XML tools for the purpose of extracting information about the specifics of the changes. A case study of using the meta-differencing approach on an open-source system is presented to demonstrate its usefulness and validity.
Jonathan I. Maletic, Michael L. Collard
ICSM1
2003 Source Viewer 3D (sv3D) - A Framework for Software Visualization
abstract
Source Viewer 3D is a software visualization framework that uses a 3D metaphor to represent software system and analysis data. The 3D representation is based on the SeeSoft pixel metaphor. It extends the original metaphor by rendering the visualization in a 3D space. New, object-based manipulation methods and simultaneous alternative mappings are available to the user.
Jonathan I. Maletic, Andrian Marcus, Louis Feng
ICSE1
2003 Recovering Documentation-to-Source-Code Traceability Links using Latent Semantic Indexing
abstract
An information retrieval technique, latent semantic indexing, is used to automatically identify traceability links from system documentation to program source code. The results of two experiments to identify links in existing software systems (i.e., the LEDA library, and Albergate) are presented. These results are compared with other similar type experimental results of traceability link identification using different types of information retrieval techniques. The method presented proves to give good results by comparison and additionally it is a low cost, highly flexible method to apply with regards to preprocessing and/or parsing of the source code and documentation.
Andrian Marcus, Jonathan I. Maletic
ICSE2
2002 Supporting document and data views of source code
abstract
The paper describes the use of an XML format to store and represent program source code. A new XML application, srcML (SouRCe Markup Language), is presented. srcML presumes a document view of source code where information about the syntactic structure is layered over the original source code document. The resultant multi-layered document has a base layer of all the original text (and formatting). The second layer is the syntactic information, derived from the grammar of the programming language, and is encoded in XML. This multi-layered view supports both the creation and viewing of the source code in its original form and the use of XML technologies (for tasks such as analysis and transformation of the source). Although directed at source code documents, (particularly C++) srcML is also applicable to other programming languages and to languages with a strict syntax. srcML represents a departure from the compiler centric manner in which source code is commonly stored, instead a document point of view is taken thus better supporting the manipulation and management of the large numbers of source documents typical in modern software systems.
Michael L. Collard, Jonathan I. Maletic, Andrian Marcus
ACM Symposium on Document Engineering2
2001 Ordinal Association Rules for Error Identification in Data Sets
abstract
A new extension of the Boolean association rules, ordinal association rules, that incorporates ordinal relationships among data items, is introduced. One use for ordinal rules is to identify possible errors in data. A method that finds these rules and identifies potential errors in data is proposed.
Andrian Marcus, Jonathan I. Maletic, King-Ip (David) Lin
CIKM2
2001 Incorporating PSP into a Traditional Software Engineering Course: An Experience Report
abstract
This paper presents an approach to incorporate PSP (Personal Software Process) into a traditional software engineering course that is typically contained within a computer science curriculum. Advantages and disadvantages of similar approaches are discussed. The approach has been implemented twice in an undergraduate course at the University of Memphis. This successful experience is described and gives support that the proposed approach is beneficial to both students and educators.
Jonathan I. Maletic, Anita Howald, Andrian Marcus
CSEE&T1
2001 Supporting Program Comprehension Using Semantic and Structural Information
abstract
Focuses on investigating the combined use of semantic and structural information of programs to support the comprehension tasks involved in the maintenance and reengineering of software systems. "Semantic information" refers to the domain-specific issues (both the problem and the development domains) of a software system. The other dimension, structural information, refers to issues such as the actual syntactic structure of the program, along with the control and data flow that it represents. An advanced information retrieval method, latent semantic indexing, is used to define a semantic similarity measure between software components. Components within a software system are then clustered together using this similarity measure. Simple structural information (i.e. the file organization) of the software system is then used to assess the semantic cohesion of the clusters and files with respect to each other. The measures are formally defined for general application. A set of experiments is presented which demonstrates how these measures can assist in the understanding of a nontrivial software system, namely a version of NCSA Mosaic.
Jonathan I. Maletic, Andrian Marcus
ICSE1
2001 Identification of High-Level Concept Clones in Source Code
abstract
Source code duplication occurs frequently within large software systems. Pieces of source code, functions, and data types are often duplicated in part or in whole, for a variety of reasons. Programmers may simply be reusing a piece of code via copy and paste or they may be "re-inventing the wheel". Previous research on the detection of clones is mainly focused on identifying pieces of code with similar (or nearly similar) structure. Our approach is to examine the source code text (comments and identifiers) and identify implementations of similar high-level concepts (e.g., abstract data types). The approach uses an information retrieval technique (i.e., latent semantic indexing) to statically analyze the software system and determine semantic similarities between source code documents (i.e., functions, files, or code segments). These similarity measures are used to drive the clone detection process. The intention of our approach is to enhance and augment existing clone detection methods that are based on structural analysis. This synergistic use of methods will improve the quality of clone detection. A set of experiments is presented that demonstrate the usage of semantic similarity measure to identify clones within a version of NCSA Mosaic.
Andrian Marcus, Jonathan I. Maletic
ASE2
2000 Using latent semantic analysis to identify similarities in source code to support program understanding
abstract
The paper describes the results of applying Latent Semantic Analysis (LSA), an advanced information retrieval method, to program source code and associated documentation. Latent semantic analysis is a corpus based statistical method for inducing and representing aspects of the meanings of words and passages (of natural language) reflective in their usage. This methodology is assessed for application to the domain of software components (i.e., source code and its accompanying documentation). Here LSA is used as the basis to cluster software components. This clustering is used to assist in the understanding of a nontrivial software system, namely a version of Mosaic. Applying latent semantic analysis to the domain of source code and internal documentation for the support of program understanding is a new application of this method and a departure from the normal application domain of natural language.
Jonathan I. Maletic, Andrian Marcus
ICTAI1
1999 Automatic Software Clustering via Latent Semantic Analysis
abstract
The paper describes the initial results of applying Latent Semantic Analysis (LSA) to program source code and associated documentation. Latent Semantic Analysis is a corpus based statistical method for inducing and representing aspects of the meanings of words and passages (of natural language) reflective in their usage. This methodology is assessed for application to the domain of software components (i.e., source code and its accompanying documentation). The intent of applying Latent Semantic Analysis to software components is to automatically induce a specific semantic meaning of a given component. Here LSA is used as the basis to cluster software components. Results of applying this method to the LEDA library and MINIX operating system are given. Applying Latent Semantic Analysis to the domain of source code and internal documentation for the support of software reuse is a new application of this method and a departure from the normal application domain of natural language.
Jonathan I. Maletic, Naveen Valluri
ASE1
1994 A Tool to Support Knowledge Based Software Maintenance: The Software Service Bay
abstract
A software maintenance methodology, The Software Service Bay, is introduced. This methodology is analogous to the automotive service bay which employs a number of experts for particular maintenance problems. Problems in maintenance are reformulated so they may be solved with current AI tools and technologies.>
Jonathan I. Maletic, Robert G. Reynolds
ICTAI1
1992 Operationalizing Software Reuse as a Problem in Inductive Learning
Robert G. Reynolds, Jonathan I. Maletic, Elena Zannoni
IEA/AIE2
1992 Extracting Procedural Knowledge from Software Systems Using Inductive Leaning in the PM system
abstract
The issue of software reuse has been found to be a much harder task than previously thought. Some of the problems are due to the lack of emphasis placed on non-functional requirements during the software development phase, such as maintainability and understandability. Other problems arise from the difficulty of defining precise criteria for considering a software module reusable. They are usually elusive, and vary dramatically from one domain to another. This paper presents PM, a software system the goal of which is the automation of the software reuse process. PM uses an incremental approach in performing analysis and storage of software modules, at different levels of granularity. Its fundamental characteristics are domain independence and flexibility, accomplished applying inductive learning techniques and analyzing reusable and nonreusable code examples.>
Robert G. Reynolds, Jonathan I. Maletic, Elena Zannoni
SEKE2
1991 The use of version space controlled genetic algorithms to solve the Boole problem
abstract
It is demonstrated that the VGA (version space guided genetic algorithm) is a particular instantiation of a more general class of systems, termed autonomous learning elements (ALEs). The basic components of an ALEs are discussed. The Boole problem posed by S. W. Wilson (1987) is introduced, and its expression in terms of the VGA framework is discussed. The details of the VGA system are given followed by a discussion of results. In particular, the performances of the VGA on two versions of the Boole problem are described and compared with those of classifier systems and decision trees.>
Robert G. Reynolds, Jonathan I. Maletic, Shan-Ping Chang
ICTAI2
1990 PM: A System to Support the Automatic Acquisition of Programming Knowledge
abstract
A system called partial metrics (PM) which utilizes chunking as a model for acquiring knowledge about program implementation is described. The chunking paradigm has three phases. The first phase partitions the object to be chunked into relatively independent parts called aggregates. The objects to be chunked in PM are code modules. Modules are separated into a collection of aggregates based on a model of stepwise refinement. A heuristic that generates a hierarchically structured collection of refinement steps describing how the program could have been developed as a set of independent refinement decisions (object-oriented stepwise implementation) is given. The second phase encodes (abstracts) each of the aggregates. Various techniques for symbolic learning can be applied to produce a frame-based encoding of information present in the code. This abstraction contains information about the aggregate's role in the refinement process as well as the code's functionality. The third phase inserts the chunked aggregate into a hierarchically structured library of cases based on the contents of its frame description. The storage of an aggregate enables its future use in problem-solving activities. An example of how this approach can be used to acquire knowledge from a sort module is described.>
Robert G. Reynolds, Jonathan I. Maletic, Stephen E. Porvin
IEEE Trans. Knowl. Data Eng.2
1989 PM: A Metrics Driven Plan Compiler
Robert G. Reynolds, Jonathan I. Maletic, Stephen E. Porvin
SEKE2