VLDB 2026 Research / reviewers in the wild / expert
Michael W. Godfrey
dblp:g/MichaelWGodfrey
· DBLP profile ↗
72ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0001-5500-025XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 70 · 9 first-author · 8 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Reviewer Experience in Code Review Comment GenerationabstractModern code review is a ubiquitous software quality assurance process aimed at identifying and resolving potential issues (e.g., functional, evolvability) within newly written code. Despite its effectiveness, the process demands large amounts of effort from the human reviewers involved. To help alleviate this workload, researchers have trained various deep learning-based language models to imitate human reviewers in providing natural language code reviews for submitted code. Formally, this automation task is known as code review comment generation. Prior work has demonstrated improvements in code review comment generation by leveraging machine learning techniques and neural models, such as transfer learning and the transformer architecture. However, the quality of the model-generated reviews remains sub-optimal due to the quality of the open-source code review data used in model training. This is in part due to the data obtained from open-source projects where code reviews are conducted in a public forum, and reviewers possess varying levels of software development experience, potentially affecting the quality of their feedback. To accommodate this variation, we propose a suite of experience-aware training methods that utilise the reviewers’ past authoring and reviewing experiences as signals for review quality. Specifically, we propose experience-aware loss functions (ELF), which use the reviewers’ authoring and reviewing ownership of a project as weights in the model’s loss function. Through this method, experienced reviewers’ code reviews yield larger influence over the model’s behaviour. Compared to the SOTA model, ELF was able to generate higher quality reviews in terms of accuracy (e.g., +29% applicable comments), informativeness (e.g., +56% suggestions), and issue types discussed (e.g., +129% functional issues identified). The key contribution of this work is the demonstration of how traditional software engineering concepts such as reviewer experience can be integrated into the design of AI-based automated code review models. Hong Yi Lin, Patanamon Thongtanunam, Christoph Treude, Michael W. Godfrey, Chunhua Liu, Wachiraphan Charoenwet 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Reevaluating the Defect Proneness of Atoms of Confusion in Java SystemsabstractBackground:Code confusion concerns source code characteristics that make code harder for authors and reviewers to comprehend. Atoms of Confusions (AoCss) are a set of low-level programming idioms for C-like languages that have been proposed as a potential source of code confusion; previous studies have empirically evaluated the extent to which they (i) are confusing to developers and (ii) introduce risk to software products. Guoshuai Shi, Farshad Kazemi, Michael W. Godfrey, Shane McIntosh |
ESEM | 3 |
| 2024 | Studying the impact of risk assessment analytics on risk awareness and code review performance
Xueyao Yu, Filipe Roseiro Côgo, Shane McIntosh, Michael W. Godfrey |
Empir. Softw. Eng. | 4 |
| 2024 | What is an app store? The software engineering perspective
Wenhan Zhu, Sebastian Proksch 0001, Daniel M. Germán, Michael W. Godfrey, Li Li 0029, Shane McIntosh |
Empir. Softw. Eng. | 4 |
| 2023 | Bash in the Wild: Language Usage, Code Smells, and BugsabstractThe Bourne-again shell (Bash) is a prevalent scripting language for orchestrating shell commands and managing resources in Unix-like environments. It is one of the mainstream shell dialects that is available on most GNU Linux systems. However, the unique syntax and semantics of Bash could easily lead to unintended behaviors if carelessly used. Prior studies primarily focused on improving the reliability of Bash scripts or facilitating writing Bash scripts; there is yet no empirical study on the characteristics of Bash programs written in reality, e.g., frequently used language features, common code smells, and bugs. In this article, we perform a large-scale empirical study of Bash usage, based on analyses over one million open source Bash scripts found in Github repositories. We identify and discuss which features and utilities of Bash are most often used. Using static analysis, we find that Bash scripts are often error-prone, and the error-proneness has a moderately positive correlation with the size of the scripts. We also find that the most common problem areas concern quoting, resource management, command options, permissions, and error handling. We envision that these findings can be beneficial for learning Bash and future research that aims to improve shell and command-line productivity and reliability. Yiwen Dong 0002, Zheyang Li, Yongqiang Tian 0001, Chengnian Sun, Michael W. Godfrey, Meiyappan Nagappan |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2022 | DocTer: documentation-guided fuzzing for testing deep learning API functionsabstractInput constraints are useful for many software development tasks. For example, input constraints of a function enable the generation of valid inputs, i.e., inputs that follow these constraints, to test the function deeper. API functions of deep learning (DL) libraries have DL-specific input constraints, which are described informally in the free-form API documentation. Existing constraint-extraction techniques are ineffective for extracting DL-specific input constraints. Danning Xie, Mijung Kim, Hung Viet Pham, Lin Tan 0001, Xiangyu Zhang 0001, Michael W. Godfrey |
ISSTA | 7 |
| 2022 | An empirical study of question discussions on Stack Overflow
Wenhan Zhu, Haoxiang Zhang 0001, Ahmed E. Hassan, Michael W. Godfrey |
Empir. Softw. Eng. | 4 |
| 2021 | Mea culpa: How developers fix their own simple bugs differently from other developersabstractIn this work, we study how the authorship of code affects bug-fixing commits using the SStuBs dataset, a collection of single-statement bug fix changes in popular Java Maven projects. More specifically, we study the differences in characteristics between simple bug fixes by the original author - that is, the developer who submitted the bug-inducing commit - and by different developers (i.e., non-authors). Our study shows that nearly half (i.e., 44.3%) of simple bugs are fixed by a different developer. We found that bug fixes by the original author and by different developers differed qualitatively and quantitatively. We observed that bug-fixing time by authors is much shorter than that of other developers. We also found that bug-fixing commits by authors tended to be larger in size and scope, and address multiple issues, whereas bug-fixing commits by other developers tended to be smaller and more focused on the bug itself. Future research can further study the different patterns in bug-fixing and create more tailored tools based on the developer's needs. Wenhan Zhu, Michael W. Godfrey |
MSR | 2 |
| 2020 | mel- model extractor language for extracting facts from modelsabstractThere is a large body of research on extracting models from code-related artifacts to enable model-based analyses of large software systems. However, engineers do not always have access to the entire code base of a system: some components may be procured from third-party suppliers based on a Model specification or their code may be generated automatically from Models. Robert Hackman, Joanne M. Atlee, Finn Hackett, Michael W. Godfrey |
MoDELS | 4 |
| 2019 | Detecting Feature-Interaction Symptoms in Automotive Software using Lightweight AnalysisabstractModern automotive software systems are large, complex, and feature rich; they can contain over 100 million lines of code, comprising hundreds of features distributed across multiple electronic control units (ECUs), all operating in parallel and communicating over a CAN bus. Because they are safety-critical systems, the problem of possible Feature Interactions (FIs) must be addressed seriously; however, traditional detection approaches using dynamic analyses are unlikely to scale to the size of these systems. We are investigating an approach that detects static source-code patterns that are symptomatic of FIs. The tools report Feature-Interaction warnings, which can be investigated further by engineers to determine if they represent true FIs and if those FIs are problematic. In this paper, we present our preliminary toolchain for FI detection. First, we extract a collection of static “facts” from the source code, such as function calls, variable assignments, and messages between features. Next, we perform relational algebra transformations on this factbase to infer additional “facts” that represent more complicated design information about the code, such as potential information flows and data dependencies; then, the full collection of “facts” is matched against a curated set of patterns for FI symptoms. We present a set of five patterns for FIs in automotive software as well a case study in which we applied our tools to the Autonomoose autonomous-driving software, developed at the University of Waterloo. Our approach identified 1,444 possible FIs in this codebase, of which 10% were classified as being probable interactions worthy of further investigation. Bryan J. Muscedere, Robert Hackman, Davood Anbarnam, Joanne M. Atlee, Ian J. Davis, Michael W. Godfrey |
SANER | 6 |
| 2017 | On-demand Developer DocumentationabstractWe advocate for a paradigm shift in supporting the information needs of developers, centered around the concept of automated on-demand developer documentation. Currently, developer information needs are fulfilled by asking experts or consulting documentation. Unfortunately, traditional documentation practices are inefficient because of, among others, the manual nature of its creation and the gap between the creators and consumers. We discuss the major challenges we face in realizing such a paradigm shift, highlight existing research that can be leveraged to this end, and promote opportunities for increased convergence in research on software documentation. Martin P. Robillard, Andrian Marcus, Christoph Treude, Gabriele Bavota, Oscar Chaparro, Neil A. Ernst, Marco Aurélio Gerosa, Michael W. Godfrey, Michele Lanza 0001, Mario Linares-Vásquez, Gail C. Murphy, Laura Moreno, David C. Shepherd, Edmund Wong |
ICSME | 8 |
| 2017 | Guest editor's introduction to the Special Issue on Source Code Analysis and Manipulation (SCAM 2015)abstractAbstract We are happy to introduce you to this special issue that presents selected papers from the 15th IEEE International Working Conference on Source Code Analysis and Manipulation (SCAM 2015). SCAM is a leading conference that brings together researchers and practitioners working on theory, techniques, and applications that concern analysis and/or manipulation of the source code of computer systems. While much attention in the wider software engineering community is properly directed towards other aspects of systems development and evolution, such as specification, design, and requirements engineering, it is the source code that contains the only precise description of the behavior of the system. The analysis and manipulation of source code thus remains a pressing concern. SCAM 2015 was held on September 27 to 28, 2015, in Bremen, Germany, together with 31st International Conference on Software Maintenance and Evolution (ICSME). Foutse Khomh, David Lo 0001, Michael W. Godfrey |
J. Softw. Evol. Process. | 3 |
| 2016 | Code review quality: how developers see itabstractIn a large, long-lived project, an effective code review process is key to ensuring the long-term quality of the code base. In this work, we study code review practices of a large, open source project, and we investigate how the developers themselves perceive code review quality. We present a qualitative study that summarizes the results from a survey of 88 Mozilla core developers. The results provide developer insights into how they define review quality, what factors contribute to how they evaluate submitted code, and what challenges they face when performing review tasks. We found that the review quality is primarily associated with the thoroughness of the feedback, the reviewer's familiarity with the code, and the perceived quality of the code itself. Also, we found that while different factors are perceived to contribute to the review quality, reviewers often find it difficult to keep their technical skills up-to-date, manage personal priorities, and mitigate context switching. Oleksii Kononenko, Olga Baysal, Michael W. Godfrey |
ICSE | 3 |
| 2016 | An empirical study on the practice of maintaining object-relational mapping code in Java systemsabstractDatabases have become one of the most important components in modern software systems. For example, web services, cloud computing systems, and online transaction processing systems all rely heavily on databases. To abstract the complexity of accessing a database, developers make use of Object-Relational Mapping (ORM) frameworks. ORM frameworks provide an abstraction layer between the application logic and the underlying database. Such abstraction layer automatically maps objects in Object-Oriented Languages to database records, which significantly reduces the amount of boilerplate code that needs to be written. Tse-Hsun (Peter) Chen, Weiyi Shang, Jinqiu Yang 0001, Ahmed E. Hassan, Michael W. Godfrey, Mohamed N. Nasser, Parminder Flora |
MSR | 5 |
| 2016 | Investigating technical and non-technical factors influencing modern code review
Olga Baysal, Oleksii Kononenko, Reid Holmes, Michael W. Godfrey |
Empir. Softw. Eng. | 4 |
| 2015 | Investigating code review quality: Do people and participation matter?abstractCode review is an essential element of any mature software development project; it aims at evaluating code contributions submitted by developers. In principle, code review should improve the quality of code changes (patches) before they are committed to the project's master repository. In practice, bugs are sometimes unwittingly introduced during this process. In this paper, we report on an empirical study investigating code review quality for Mozilla, a large open-source project. We explore the relationships between the reviewers' code inspections and a set of factors, both personal and social in nature, that might affect the quality of such inspections. We applied the SZZ algorithm to detect bug-inducing changes that were then linked to the code review information extracted from the issue tracking system. We found that 54% of the reviewed changes introduced bugs in the code. Our findings also showed that both personal metrics, such as reviewer workload and experience, and participation metrics, such as the number of involved developers, are associated with the quality of the code review process. Oleksii Kononenko, Olga Baysal, Latifa Guerrouj, Yaxin Cao, Michael W. Godfrey |
ICSME | 5 |
| 2015 | Going Green: An Exploratory Analysis of Energy-Related QuestionsabstractThe popularity of smartphones - small computers that run on battery power - has exploded in the last decade. Unsurprisingly, power consumption is an overarching concern for mobile app developers, who are anxious to learn about power-related problems that are encountered by others. In this paper, we present an empirical study exploring the characteristics of energy-related questions posed in Stack Overflow, issues faced by the developers, and the most significantly discussed APIs. We extracted a sample of 5009 Stack Overflow questions, and manually analyzed 1000 posts of Android-related energy questions. Our study shows that developers are most concerned about energy-related issues that concern improper implementations, sensor, and radio utilization. Haroon Malik, Michael W. Godfrey |
MSR | 3 |
| 2015 | Recommending Posts concerning API Issues in Developer Q&A SitesabstractAPI design is known to be a challenging craft, as API designers must balance their elegant ideals against "real-world" concerns, such as utility, performance, backwards compatibility, and unforeseen emergent uses. However, to date, there is no principled method to collect or analyze API usability information that incorporates input from typical developers. In practice, developers often turn to Q&A websites such as stackoverflow.com (SO) when seeking expert advice on API use, the popularity of such sites has thus led to a very large volume of unstructured information that can be searched with diligence for answers to specific questions. The collected wisdom within such sites could, in principle, be of great help to API designers to better support developer needs, if only it could be collected, analyzed, and distilled for practical use. In this paper, we present a methodology that combines several techniques, including social network analysis and topic mining, to recommend SO posts that are likely to concern API design-related issues. To establish a comparison baseline, we introduce two more recommendation approaches: a reputation-based recommender and a random recommender. We have found that when applied to Q&A discussion of two popular mobile platforms, Android and iOS, our methodology achieves up to 93% accuracy and is more stable with its recommendations when compared to the two baseline techniques. Wei Wang 0001, Haroon Malik, Michael W. Godfrey |
MSR | 3 |
| 2015 | Understanding software artifact provenance
Michael W. Godfrey |
Sci. Comput. Program. | 1 |
| 2014 | Compiling Clones: What Happens?abstractMost clone detection techniques have focused on the analysis of source code, however, sometimes stakeholders have access only to compiled code. To address this, some approaches have been developed for finding similarities at the binary level. However, the precise relationships between source-level and binary-level similarities remains unclear: While a compiler will preserve the semantics of the source code in the transformation to an executable, the resulting binary may differ significantly in structure, including the addition and deletion of entities in the source model. Also, compilation sometimes acts as a kind of normalization, transforming syntactically different but semantically similar structures into the same binary-level representation. In this paper, we describe a preliminary study into the effects of the javac Java compiler on the results of clone detection. We use CCFinderX -- which can perform clone detection on sequences of arbitrary tokens -- to find clones in both the source code and the corresponding byte code of four large Java systems. The study shows that source code and byte code clone detection can produce significantly different results, especially for large programs. We report on a few typical examples of differences, and analyze how they are introduced by the compiler. Finally, we discuss the greater significance of this work, and sketch plans for expanded study. Oleksii Kononenko, Michael W. Godfrey |
ICSME | 3 |
| 2014 | Recommending Clones for Refactoring Using Design, Context, and HistoryabstractDevelopers know that copy-pasting code (aka code cloning) is often a convenient shortcut to achieving a design goal, albeit one that carries risks to the code quality over time. However, deciding which, if any, clones should be eliminated within an existing system is a daunting task. Fixing a clone usually means performing an invasive refactoring, and not all clones may be worth the effort, cost, and risk that such a change entails. Furthermore, sometimes cloning fulfils a useful design role, and should not be refactored at al. And clone detection tools often return very large result sets, making it hard to choose which clones should be investigated and possibly removed. In this paper, we propose an automated approach to recommend clones for refactoring by training a decision tree-based classifier. We analyze more than 600 clone instances in three medium-to large-sized open source projects, and we collect features that are associated with the source code, the context, and the history of clone instances. Our approach achieves a precision of around 80% in recommending clone refactoring instances for each target system, and similarly good precision is achieved in cross-project evaluation. By recommending which clones are appropriate for refactoring, our approach allows for better resource allocation for refactoring itself after obtaining clone detection results, and can thus lead to improved clone management in practice. Wei Wang 0001, Michael W. Godfrey |
ICSME | 2 |
| 2014 | Mining modern repositories with elasticsearchabstractOrganizations are generating, processing, and retaining data at a rate that often exceeds their ability to analyze it effectively; at the same time, the insights derived from these large data sets are often key to the success of the organizations, allowing them to better understand how to solve hard problems and thus gain competitive advantage. Because this data is so fast-moving and voluminous, it is increasingly impractical to analyze using traditional offline, read-only relational databases. Oleksii Kononenko, Olga Baysal, Reid Holmes, Michael W. Godfrey |
MSR | 4 |
| 2014 | No issue left behind: reducing information overload in issue trackingabstractModern software development tools such as issue trackers are often complex and multi-purpose tools that provide access to an immense amount of raw information. Unfortunately, developers sometimes feel frustrated when they cannot easily obtain the particular information they need for a given task; furthermore, the constant influx of new data — the vast majority of which is irrelevant to their task at hand — may result in issues being "dropped on the floor". In this paper, we present a developer-centric approach to issue tracking that aims to reduce information overload and improve developers' situational awareness. Our approach is motivated by a grounded theory study of developer comments, which suggests that customized views of a project's repositories that are tailored to developer-specific tasks can help developers better track their progress and understand the surrounding technical context. From the qualitative study, we uncovered a model of the kinds of information elements that are essential for developers in completing their daily tasks, and from this model we built a tool organized around customized issue-tracking dashboards. Further quantitative and qualitative evaluation demonstrated that this dashboard-like approach to issue tracking can reduce the volume of irrelevant emails by over 99% and also improve support for specific issue-tracking tasks. Olga Baysal, Reid Holmes, Michael W. Godfrey |
SIGSOFT FSE | 3 |
| 2014 | Special issue on program comprehension
Michael W. Godfrey, Arie van Deursen |
Empir. Softw. Eng. | 1 |
| 2014 | On the evolution of Lehman's LawsabstractSUMMARY In this brief paper, we honor the contributions of the late Prof. Manny Lehman to the study of software evolution. We do so by means of a kind of evolutionary case study: First, we discuss his background in engineering and explore how this helped to shape his views on software systems and their development; next, we discuss the laws of software evolution that he postulated based on his industrial experiences; and finally, we examine how the nature of software systems and their development are undergoing radical change, and we consider what this means for future evolutionary studies of software. Copyright © 2013 John Wiley & Sons, Ltd. Michael W. Godfrey, Daniel M. Germán |
J. Softw. Evol. Process. | 1 |
| 2014 | An exploratory study of the evolution of communicated information about the execution of large software systemsabstractSUMMARY Substantial research in software engineering focuses on understanding the dynamic nature of software systems in order to improve software maintenance and program comprehension. This research typically makes use of automated instrumentation and profiling techniques after the fact, that is, without considering domain knowledge. In this paper, we examine another source of dynamic information that is generated from statements that have been inserted into the code base during development to draw the system administrators' attention to important run‐time phenomena. We call this source communicated information (CI). Examples of CI include execution logs and system events. The availability of CI has sparked the development of an ecosystem ofLog Processing Apps(LPAs) that surround the software system under analysis to monitor and document various run‐time constraints. The dependence of LPAs on the timeliness, accuracy and granularity of the CI means that it is important to understand the nature of CI and how it evolves over time, both qualitatively and quantitatively. Yet, to our knowledge, little empirical analysis has been performed on CI and its evolution. In a case study on two large open source and one industrial software systems, we explore the evolution of CI by mining the execution logs of these systems and the logging statements in the source code. Our study illustrates the need for better traceability between CI and the LPAs that analyze the CI. In particular, we find that the CI changes at a high rate across versions, which could lead to fragile LPAs. We found that up to 70% of these changes could have been avoided and the impact of 15% to 80% of the changes can be controlled through the use of robust analysis techniques by LPAs. We also found that LPAs that track implementation‐level CI (e.g. performance analysis) and the LPAs that monitor error messages (system health monitoring) are more fragile than LPAs that track domain‐level CI (e.g. workload modelling), because the latter CI tends to be long‐lived. Copyright © 2013 John Wiley & Sons, Ltd. Weiyi Shang, Zhen Ming (Jack) Jiang, Bram Adams, Ahmed E. Hassan, Michael W. Godfrey, Mohamed N. Nasser, Parminder Flora |
J. Softw. Evol. Process. | 5 |
| 2013 | Situational awareness: personalizing issue tracking systemsabstractIssue tracking systems play a central role in ongoing software development; they are used by developers to support collaborative bug fixing and the implementation of new features, but they are also used by other stakeholders including managers, QA, and end-users for tasks such as project management, communication and discussion, code reviews, and history tracking. Most such systems are designed around the central metaphor of the “issue” (bug, defect, ticket, feature, etc.), yet increasingly this model seems ill fitted to the practical needs of growing software projects; for example, our analysis of interviews with 20 Mozilla developers who use Bugzilla heavily revealed that developers face challenges maintaining a global understanding of the issues they are involved with, and that they desire improved support for situational awareness that is difficult to achieve with current issue management systems. In this paper we motivate the need for personalized issue tracking that is centered around the information needs of individual developers together with improved logistical support for the tasks they perform. We also describe an initial approach to implement such a system - extending Bugzilla - that enhances a developer's situational awareness of their working context by providing views that are tailored to specific tasks they frequently perform; we are actively improving this prototype with input from Mozilla developers. Olga Baysal, Reid Holmes, Michael W. Godfrey |
ICSE | 3 |
| 2013 | The MSR cookbook: mining a decade of researchabstractThe Mining Software Repositories (MSR) research community has grown significantly since the first MSR workshop was held in 2004. As the community continues to broaden its scope and deepens its expertise, it is worthwhile to reflect on the best practices that our community has developed over the past decade of research. We identify these best practices by surveying past MSR conferences and workshops. To that end, we review all 117 full papers published in the MSR proceedings between 2004 and 2012. We extract 268 comments from these papers, and categorize them using a grounded theory methodology. From this evaluation, four high-level themes were identified: data acquisition and preparation, synthesis, analysis, and sharing/replication. Within each theme we identify several common recommendations, and also examine how these recommendations have evolved over the past decade. In an effort to make this survey a living artifact, we also provide a public forum that contains the extracted recommendations in the hopes that the MSR community can engage in a continuing discussion on our evolving best practices. Hadi Hemmati, Sarah Nadi, Olga Baysal, Oleksii Kononenko, Wei Wang 0001, Reid Holmes, Michael W. Godfrey |
MSR | 7 |
| 2013 | Detecting API usage obstacles: a study of iOS and Android developer questionsabstractSoftware frameworks provide sets of generic functionalities that can be later customized for a specific task. When developers invoke API methods in a framework, they often encounter obstacles in finding the correct usage of the API, let alone to employ best practices. Previous research addresses this line of questions by mining API usage patterns to induce API usage templates, by conducting and compiling interviews of developers, and by inferring correlations among APIs. In this paper, we analyze API-related posts regarding iOS and Android development from a Q&A Web site, stackoverflow.com. Assuming that API-related posts are primarily about API usage obstacles, we find several iOS and Android API classes that appear to be particularly likely to challenge developers, even after we factor out API usage hotspots, inferred by modelling API usage of open source iOS and Android applications. For each API with usage obstacles, we further apply a topic mining tool to posts that are tagged with the API, and we discover several repetitive scenarios in which API usage obstacles occur. We consider our work as a stepping stone towards understanding API usage challenges based on forum-based input from a multitude of developers, input that is prohibitively expensive to collect through interviews. Our method helps to motivate future research in API usage, and can allow designers of platforms - such as iOS and Android - to better understand the problems developers have in using their platforms, and to make corresponding improvements. Wei Wang 0001, Michael W. Godfrey |
MSR | 2 |
| 2013 | Software Bertillonage - Determining the provenance of software development artifacts
Julius Davies, Daniel M. Germán, Michael W. Godfrey, Abram Hindle |
Empir. Softw. Eng. | 3 |
| 2013 | Automated topic naming - Supporting cross-project analysis of software maintenance activities
Abram Hindle, Neil A. Ernst, Michael W. Godfrey, John Mylopoulos |
Empir. Softw. Eng. | 3 |
| 2012 | Build system issues in multilanguage softwareabstractBuilding software from source is often viewed as a “solved problem” by software engineers, as there are many mature, well-known tools and techniques. However, anecdotal evidence suggests that these tools often do not effectively address the complexities of building multilanguage software. To investigate this apparent problem, we have performed a qualitative study on a set of five multilanguage open source software packages. Surprisingly, we found build system problems that prevented us from building many of these packages out-of-the-box. Our key finding is that there are commonalities among build problems that can be systematically addressed. In this paper, we describe the results of this exploratory study, identify a set of common build patterns and anti-patterns, and outline research directions for improving the build process. One such finding is that multilanguage packages avoid certain build problems by supporting compilation-free extension. As well, we find evidence that concerns from the application and implementation domains may “leak” into the build model, with both positive and negative effects on the resulting build systems. Andrew Neitsch, Kenny Wong, Michael W. Godfrey |
ICSM | 3 |
| 2012 | An industrial case study of Coman's automated task detection algorithm: What Worked, What Didn't, and WhyabstractProgrammers need explicit tool support for software maintenance tasks, and a prerequisite for this is an understanding of where the boundaries between distinct tasks lie. Asking developers to indicate manually when they switch tasks is disruptive to their normal work flow, so researchers have sought ways to infer task boundaries automatically based on the content of the interaction histories with the IDE. Coman previously reported a fully automated algorithm that achieved 80% accuracy in a lab validation study. In this paper, we evaluate the use of this algorithm within an industrial setting. We found two problems: first, a large number of the tasks identified are in fact only sessions or subparts of a larger task; second, the demonstrable effects of interruptions are not considered. We argue that the problem of task boundary detection consists of two sub-problems: first, detecting task sessions; and second, linking task sessions. Coman's algorithm only partially addresses the first, and ignores the second. Lijie Zou, Michael W. Godfrey |
ICSM | 2 |
| 2012 | Mining usage data and development artifactsabstractSoftware repository mining techniques generally focus on analyzing, unifying, and querying different kinds of development artifacts, such as source code, version control meta-data, defect tracking data, and electronic communication. In this work, we demonstrate how adding real-world usage data enables addressing broader questions of how software systems are actually used in practice, and by inference how development characteristics ultimately affect deployment, adoption, and usage. In particular, we explore how usage data that has been extracted from web server logs can be unified with product release history to study questions that concern both users' detailed dynamic behaviour as well as broad adoption trends across different deployment environments. To validate our approach, we performed a study of two open source web browsers: Firefox and Chrome. We found that while Chrome is being adopted at a consistent rate across platforms, Linux users have an order of magnitude higher rate of Firefox adoption. Also, Firefox adoption has been concentrated mainly in North America, while Chrome users appear to be more evenly distributed across the globe. Finally, we detected no evidence in age-specific differences in navigation behaviour among Chrome and Firefox users; however, we hypothesize that younger users are more likely to have more up-to-date versions than more mature users. Olga Baysal, Reid Holmes, Michael W. Godfrey |
MSR | 3 |
| 2012 | What Does Control Flow Really Look Like? Eyeballing the Cyclomatic Complexity MetricabstractAssessing the understandability of source code remains an elusive yet highly desirable goal for software developers and their managers. While many metrics have been suggested and investigated empirically, the McCabe cyclomatic complexity metric (CC) - which is based on control flow complexity - seems to hold enduring fascination within both industry and the research community despite its known limitations. In this work, we introduce the ideas of Control Flow Patterns (CFPs) and Compressed Control Flow Patterns (CCFPs), which eliminate some repetitive structure from control flow graphs in order to emphasize high-entropy graphs. We examine eight well-known open source Java systems by grouping the CFPs of the methods into equivalence classes, and exploring the results. We observed several surprising outcomes: first, the number of unique CFPs is relatively low, second, CC often does not accurately reflect the intricacies of Java control flow, and third, methods with high CC often have very low entropy, suggesting that they may be relatively easy to understand. These findings challenge the widely-held belief that there is a clear-cut causal relationship between CC and understandability, and suggest that CC and similar measures need to be reconsidered as metrics for code understandability. Jurgen J. Vinju, Michael W. Godfrey |
SCAM | 2 |
| 2012 | Introduction to the special issue on software repository mining in 2009
Michael W. Godfrey, E. James Whitehead Jr. |
Empir. Softw. Eng. | 1 |
| 2011 | Multifractal aspects of software developmentabstractSoftware development is difficult to model, particularly the noisy, non-stationary signals of changes per time unit, extracted from version control systems (VCSs). Currently researchers are utilizing timeseries analysis tools such as ARIMA to model these signals extracted from a project's VCS. Unfortunately current approaches are not very amenable to the underlying power-law distributions of this kind of signal. We propose modeling changes per time unit using multifractal analysis. This analysis can be used when a signal exhibits multi-scale self-similarity, as in the case of complex data drawn from power-law distributions. Specifically we utilize multifractal analysis to demonstrate that software development is multifractal, that is the signal is a fractal composed of multiple fractal dimensions along a range of Hurst exponents. Thus we show that software development has multi-scale self-similarity, that software development is multifractal. We also pose questions that we hope multifractal analysis can answer. Abram Hindle, Michael W. Godfrey, Richard C. Holt |
ICSE | 2 |
| 2011 | A Tale of Two BrowsersabstractWe explore the space of open source systems and their user communities by examining the development artifact histories of two popular web browsers — Firefox and Chrome — as well as usage data. By examining the data and addressing a number of research questions, two very different profiles emerge: Firefox, as the older and established system, with long product version cycles, longer bug fix cycles, and a user base that is slow to adopt newer versions; and Chrome, as the new and fast evolving system, with short version cycles, shorter bug fix cycles, and a user base that very quickly adopts new versions as they become available. 1. Olga Baysal, Ian J. Davis, Michael W. Godfrey |
MSR | 3 |
| 2011 | Software bertillonage: finding the provenance of an entityabstractDeployed software systems are typically composed of many pieces, not all of which may have been created by the main development team. Often, the provenance of included components -- such as external libraries or cloned source code -- is not clearly stated, and this uncertainty can introduce technical and ethical concerns that make it difficult for system owners and other stakeholders to manage their software assets. In this work, we motivate the need for the recovery of the provenance of software entities by a broad set of techniques that could include signature matching, source code fact extraction, software clone detection, call flow graph matching, string matching, historical analyses, and other techniques. We liken our provenance goals to that of Bertillonage, a simple and approximate forensic analysis technique based on bio-metrics that was developed in 19th century France before the advent of fingerprints. As an example, we have developed a fast, simple, and approximate technique called anchored signature matching for identifying library version information within a given Java application. This technique involves a type of structured signature matching performed against a database of candidates drawn from the Maven2 repository, a 150GB collection of open source Java libraries. An exploratory case study using a proprietary e-commerce Java application illustrates that the approach is both feasible and effective. Julius Davies, Daniel M. Germán, Michael W. Godfrey, Abram Hindle |
MSR | 3 |
| 2011 | Automated topic naming to support cross-project analysis of software maintenance activitiesabstractResearchers have employed a variety of techniques to extract underlying topics that relate to software development artifacts. Typically, these techniques use semi-unsupervised machine-learning algorithms to suggest candidate word-lists. However, word-lists are difficult to interpret in the absence of meaningful summary labels. Current topic modeling techniques assume manual labelling and do not use domainspecific knowledge to improve, contextualize, or describe results for the developers. We propose a solution: automated labelled topic extraction. Topics are extracted using Latent Dirichlet Allocation (LDA) from commit-log comments recovered from source control systems such as CVS and Bit-Keeper. These topics are given labels from a generalizable cross-project taxonomy, consisting of non-functional requirements. Our approach was evaluated with experiments and case studies on two large-scale RDBMS projects: MySQL and MaxDB. The case studies show that labelled topic extraction can produce appropriate, context-sensitive labels relevant to these projects, which provides fresh insight into their evolving software development activities. Abram Hindle, Neil A. Ernst, Michael W. Godfrey, John Mylopoulos |
MSR | 3 |
| 2011 | A Study of Cloning in the Linux SCSI DriversabstractTo date, most research on software code cloning has concentrated on detection and analysis techniques and their evaluation, and most empirical studies of cloning have investigated cloning within single system versions. In this paper, we present the results of a longitudinal study of cloning among the SCSI drivers for the Linux operating system that spans 16 years of evolution. We have chosen the SCSI driver subsystem as a test subject as it is known that cloning has been embraced by these developers as a design practice: when a new SCSI card comes out that is similar to an old one, but different enough to warrant its own implementation, a new driver may be cloned from an existing one. We discuss the results of our qualitative and quantitative analyses, including how the layered architecture of the SCSI subsystem seems to have affected the use of cloning as a design tool, the likelihood of consistent and inconsistent change over time, and the predictive power of using cloning between two independent driver implementations to model the similarity between two target devices. Wei Wang 0001, Michael W. Godfrey |
SCAM | 2 |
| 2010 | Software process recovery using Recovered Unified Process ViewsabstractThe development process for a given software system is a combination of an idealized, prescribed model and a messy set of ad hoc practices. To some degree, process compliance can be enforced by supporting tools that require various steps be followed in order; however, this approach is often perceived as heavyweight and inflexible by developers, who generally prefer that tools support their desired work habits rather than limit their choices. An alternative approach to monitoring process compliance is to instrument the various tools and repositories that developers use - such as version control systems, bug-trackers, and mailing-list archives - and to build models of the de facto development process through observation, analysis, and inference. In this paper, we present a technique for recovering a project's software development processes from a variety of existing artifacts. We first apply unsupervised and supervised techniques - including word-bags, topic analysis, summary statistics, and Bayesian classifiers - to annotate software artifacts by related topics, maintenance types, and non-functional requirements. We map the analysis results onto a time-line based view of the Unified Process development model, which we call Recovered Unified Process Views. We demonstrate our approach for extracting these process views on two case studies: FreeBSD and SQLite. Abram Hindle, Michael W. Godfrey, Richard C. Holt |
ICSM | 2 |
| 2009 | Understanding source package organization using the hybrid modelabstractWithin a large, object-oriented software system it is common to partition the classes into a set of packages, which implicitly serve as a set of coarsely-grained logical design units. However, as such a system evolves and design drift sets in, it becomes increasingly challenging for developers - especially those who are new to the project - to comprehend the underlying criteria behind the package-level design of the system. This problem is exacerbated by the fact that in most object-oriented programming languages the package (or namespace) construct has little semantics beyond that of a simple container, and so fails to capture the essential properties of the objects that its contained classes represent. In this paper, we propose an approach to uncovering package partitioning criteria by analyzing the collaboration patterns between packages. Our analysis approach is based on the Hybrid Model, a program model that describes the coarsely-grained structure and global behaviour of an object-oriented system. We present an exploratory case study to show how our approach can help maintainers to derive the design criteria related to coupling, cohesion, function reuse, and inheritance reuse. Xinyi Dong, Michael W. Godfrey |
ICSM | 2 |
| 2009 | What's hot and what's not: Windowed developer topic analysisabstractAs development on a software project progresses, developers shift their focus between different topics and tasks many times. Managers and newcomer developers often seek ways of understanding what tasks have recently been worked on and how much effort has gone into each; for example, a manager might wonder what unexpected tasks occupied their team's attention during a period when they were supposed to have been implementing new features. Tools such as Latent Dirichlet Allocation (LDA) and Latent Semantic Indexing (LSI) can be used to extract a set of independent topics from a corpus of commit-log comments. Previous work in the area has created a single set of topics by analyzing comments from the entire lifetime of the project. In this paper, we propose windowing the topic analysis to give a more nuanced view of the system's evolution. By using a defined time-window of, for example, one month, we can track which topics come and go over time, and which ones recur. We propose visualizations of this model that allows us to explore the evolving stream of topics of development occurring over time. We demonstrate that windowed topic analysis offers advantages over topic analysis applied to a project's lifetime because many topics are quite local. Abram Hindle, Michael W. Godfrey, Richard C. Holt |
ICSM | 2 |
| 2009 | A bug you like: A framework for automated assignment of bugsabstractAssigning bug reports to individual developers is typically a manual, time-consuming, and tedious task. In this paper, we present a framework for automated assignment of bug-fixing tasks. Our approach employs preference elicitation to learn developer predilections in fixing bugs within a given system. This approach infers knowledge about a developer's expertise by analyzing the history of bugs previously resolved by the developer. We apply a vector space model to recommend experts for resolving bugs. When a new bug report arrives, the system automatically assigns it to the appropriate developer considering his or her expertise, current workload, and preferences. We address the task allocation problem by proposing a set of heuristics that support accurate assignment of bug reports to the developers. Olga Baysal, Michael W. Godfrey, Robin Cohen |
ICPC | 2 |
| 2009 | Automatic classication of large changes into maintenance categoriesabstractLarge software systems undergo significant evolution during their lifespan, yet often individual changes are not well documented. In this work, we seek to automatically classify large changes into various categories of maintenance tasks - corrective, adaptive, perfective, feature addition, and non-functional improvement - using machine learning techniques. In a previous paper, we found that many commits could be classified easily and reliably based solely on the manual analysis of the commit metadata and commit messages (i.e., without reference to the source code). Our extension is the automation of classification by training machine learners on features extracted from the commit metadata, such as the word distribution of a commit message, commit author, and modules modified. We validated the results of the learners via 10-fold cross validation, which achieved accuracies consistently above 50%, indicating good to fair results. We found that the identity of the author of a commit provided much information about the maintenance class of a commit, almost as much as the words of the commit message. This implies that for most large commits, the Source Control System (SCS) commit messages plus the commit author identity is enough information to accurately and automatically categorize the nature of the maintenance task. Abram Hindle, Daniel M. Germán, Michael W. Godfrey, Richard C. Holt |
ICPC | 3 |
| 2009 | Editorial for Special Issue of JASE on Source Code Analysis and Manipulation
Michael W. Godfrey, Bogdan Korel |
Autom. Softw. Eng. | 1 |
| 2009 | Reading beside the lines: Using indentation to rank revisions by complexity
Abram Hindle, Michael W. Godfrey, Richard C. Holt |
Sci. Comput. Program. | 2 |
| 2008 | 2nd International Workshop on Advanced Software Development Tools and Techniques (WASDeTT): Tools for software maintenance, visualization, and reverse engineeringabstractThe objective of the 2nd international workshop on advanced software development tools and techniques (WASDeTT) is to provide interested researchers with a forum to share their tool building experiences and to explore how tools can be built more effectively and efficiently. This workshop specifically focuses on tools for software maintenance and comprehension and addresses issues such as tool-building in an industrial context, component-based tool building, and tool building in teams. Holger M. Kienle, Leon Moonen, Michael W. Godfrey, Hausi A. Müller |
ICSM | 3 |
| 2008 | Identifying Architectural Change Patterns in Object-Oriented SystemsabstractAs an object-oriented system evolves, its architecture tends to drift away from the original design. Knowledge of how the system has changed at coarse-grained levels is key to understanding the de facto architecture, as it helps to identify potential architectural decay and can provide guidance for further maintenance activities. However, current studies of object-oriented software changes are mostly targeted at the class or method level. In this paper, we propose a new approach to modeling object-oriented software changes at coarse-grained levels. We take snapshots of an object-oriented system, represent each version of the system as a hybrid model, and detect software changes at coarse-grained level by comparing two hybrid models. Based on this approach, we further identify a collection of change patterns, which help interpret how system changes at the architecture level. Finally, we present an exploratory case study to show how our approach can help maintainers capture and better comprehend architectural evolution of object-oriented software systems. Xinyi Dong, Michael W. Godfrey |
ICPC | 2 |
| 2008 | Reading Beside the Lines: Indentation as a Proxy for Complexity MetricabstractMaintainers face the daunting task of wading through a collection of both new and old revisions, trying to ferret out revisions which warrant personal inspection. One can rank revisions by size/lines of code (LOC), but often, due to the distribution of the size of changes, revisions will be of similar size. If we can't rank revisions by LOC perhaps we can rank by Halstead's and McCabe's complexity metrics? However, these metrics are problematic when applied to code fragments (revisions) written in multiple languages: special parsers are required which may not support the language or dialect used; analysis tools may not understand code fragments. We propose using the statistical moments of indentation as a lightweight, language independent, revision/diff friendly metric which actually proxies classical complexity metrics. We have extensively evaluated our approach against the entire CVS histories of the 278 of the most popular and most active SourceForge projects. We found that our results are linearly correlated and rank-correlated with traditional measures of complexity, suggesting that measuring indentation is a cheap and accurate proxy for code complexity of revisions. Thus ranking revisions by the standard deviation and summation of indentation will be very similar to ranking revisions by complexity. Abram Hindle, Michael W. Godfrey, Richard C. Holt |
ICPC | 2 |
| 2008 | Query Technologies and Applications for Program Comprehension (QTAPC 2008)abstractIndustrial software systems are large and complex, both in terms of the software entities and their relationships. Consequently, understanding how a software system works requires the ability to pose queries over the design-level entities of the system. Traditionally, this task has been supported by simple tools (e.g., grep) combined with the programmer's intuition and experience. Recently, however, specialized code query technologies have matured to the point where they can be used in industrial situations, providing more intelligent, timely, and efficient responses to developer queries. This working session aims to explore the state of the art in code query technologies, and discover new ways in which these technologies may be useful in program comprehension. The session brings together researchers and practitioners. We survey existing techniques and applications, trying to understand the strengths and weaknesses of the various approaches, and sketch out new frontiers that hold promise. Mathieu Verbaere, Michael W. Godfrey, Tudor Gîrba |
ICPC | 2 |
| 2008 | From Indentation Shapes to Code StructuresabstractIn a previous study, we showed that indentation was regular across multiple languages and the variance in the level of indentation of a block of revised code is correlated with metrics such as McCabe cyclomatic complexity. Building on that work the current paper investigates the relationship between the "shape'' of the indentation of the revised code block (the "revision'') and the corresponding syntactic structure of the code. We annotated revisions matching these three indentation shapes: "flat'' (all lines are equally indented), "slash'' (indentation becomes increasingly deep), or "bubble'' (indentation increases and then decreases). We then classified the code structure as one of: function definition, loop, expression, comment, etc. We studied thousands of revisions, coming from over 200 software projects, written in a variety of languages. Our study indicates that indentation shape correlates positively with code structure; that is, certain shapes typically correspond to certain code structures. For example, flat shapes commonly correspond to comments while bubble shapes commonly correspond to conditionals and function definitions. These results can form the basis of a tool framework that can analyze code in a language independent way to support browsing targeted to viewing particular code structures such as conditionals or comments. Abram Hindle, Michael W. Godfrey, Richard C. Holt |
SCAM | 2 |
| 2008 | "Cloning considered harmful" considered harmful: patterns of cloning in software
Cory Kapser, Michael W. Godfrey |
Empir. Softw. Eng. | 2 |
| 2007 | System-level Usage Dependency Analysis of Object-Oriented SystemsabstractUncovering, modelling, and understanding architectural level dependencies of software systems is a key task for software maintainers. However, current dependency analysis techniques for object-oriented software are targeted at the class or method level; this is because most dependencies—such as instantiates, references, and calls—must be interpreted in the context of one or more class hierarchies. In this paper, we propose an approach, called the High-level Object Dependency Graph (HODG), that captures all possible usage dependencies among coarse-grained entities. Based on the new model, we further propose a set of dependency analysis methods. Finally, we present an exploratory case study using HODGs—supported by an automated analysis tool—of the Apache Ant build system; we show how HODG analysis can help maintainers capture external properties of coarse-grained entities, and better understand the nature of their interdependencies. Xinyi Dong, Michael W. Godfrey |
ICSM | 2 |
| 2007 | Release Pattern Discovery: A Case Study of Database SystemsabstractStudying the release-time activities of a software project — that is, activities that occur around the time of a major or minor release — can provide insights into both the development processes used and the nature of the system itself. Although tools rarely record detailed logs of developer behavior, we can infer release-time activities from available data, such as logs from revision control systems, bug tracking systems, etc. In this paper, we discuss the results of a case study in mining patterns of release-time behavior from the revision control systems of four open source database systems. We partitioned the development artifacts into four classes — source code, tests, build files, and documentation — to be able to characterize the behavioral patterns more precisely. We found, for example, that there were consistent activity patterns around release time within each of the individual projects; we also found that these patterns did not persist across systems, leading us to hypothesize that the four projects follow different but consistent development patterns of activity around releases. Abram Hindle, Michael W. Godfrey, Richard C. Holt |
ICSM | 2 |
| 2007 | A Hybrid Program Model for Object-Oriented Reverse EngineeringabstractA commonly used strategy to address the scalability challenge in object-oriented reverse engineering is to synthesize coarse-grained representations, such as package diagrams. However, the traditional coarse-grained representations are poorly suited to object-oriented program comprehension as they can be difficult to map to the domain object models, contain little real detail, and provide few clues to the design decisions made during development. In this paper, we propose a hybrid model of objectoriented software that blends the use of classes and entities at different levels of granularity. Each coarse-grained entity represents a set of software objects, and contains the complete static description of the objects it represents. This hybrid model allows maintainers to understand objects as independent units, and focus on the their external properties and their interrelationships at different levels of granularity. We show the usefulness of the hybrid model to program comprehension by means of an exploratory case study. Michael W. Godfrey |
ICPC | 1 |
| 2007 | Detecting Interaction Coupling from Task Interaction HistoriesabstractA repository of task structures can reveal important latent knowledge about the development of a software system. Although approaches have been proposed to recover artifacts within a task structure, identifying relations that are relevant to a task remains a problem. In this work, we propose to detect "interaction coupling" from task interaction histories (i.e., records of when the artifacts were being used or modified in a task, as observed by the IDE), and use this information to mine patterns to aid in the comprehension of maintenance activities. In our case study, we found we were able to recover latent information about the development process; for example, our results suggest that restructuring is more costly than any other maintenance activity. Lijie Zou, Michael W. Godfrey, Ahmed E. Hassan |
ICPC | 2 |
| 2007 | Unified use case statecharts: case studies
Davor Svetinovic, Daniel M. Berry, Nancy A. Day, Michael W. Godfrey |
Requir. Eng. | 4 |
| 2006 | Supporting the analysis of clones in software systemsabstractAbstract Code duplication is a well‐documented problem in industrial software systems. There has been considerable research into techniques for detecting duplication in software, and there are several effective tools to perform this task. However, there have been few detailed qualitative studies into how cloning actually manifests itself within software systems. This is primarily due to the large result sets that many clone‐detection tools return; these result sets are very difficult to manage without complementary tool support that can scale to the size of the problem, and this kind of support does not currently exist. In this paper we present an in‐depth case study of cloning in a large software system that is in wide use, the Apache Web server; we provide insights into cloning as it exists in this system, and we demonstrate techniques to manage and make effective use of the large result sets of clone‐detection tools. In our case study, we found several interesting types of cloning occurrences, such as ‘cloning hotspots’, where a single subsystem comprising only 17% of the system code contained 38.8% of the clones. We also found several examples of cloning behavior that were beneficial to the development of the system, in particular cloning as a way to add experimental functionality. Copyright © 2006 John Wiley & Sons, Ltd. Cory Kapser, Michael W. Godfrey |
J. Softw. Maintenance Res. Pract. | 2 |
| 2005 | A Reference Architecture for Web BrowsersabstractA reference architecture for a domain captures the fundamental subsystems common to systems of that domain as well as the relationships between these subsystems. Having a reference architecture available can aid both during maintenance and at design time: it can improve understanding of a given system, it can aid in analyzing tradeoffs between different design options, and it can serve as a template for designing new systems and re-engineering existing ones. In this paper, we examine the history of the Web browser domain and identify several underlying phenomena that have contributed to its evolution. We develop a reference architecture for Web browsers based on two well known open source implementations, and we validate it against two additional implementations. Finally, we discuss our observations about this domain and its evolutionary history; in particular, we note that the significant reuse of open source components among different browsers and the emergence of extensive Web standards have caused the browsers to exhibit "convergent evolution". Alan Grosskurth, Michael W. Godfrey |
ICSM | 2 |
| 2005 | Improved Tool Support for the Investigation of Duplication in SoftwareabstractCode duplication is a well documented problem in software systems. There has been considerable research into techniques for detecting duplication in software, and there are several effective tools to perform this task. However, a common problem with such tools is that the result set returned can be too large to handle without complementary tool support. The goal of this paper is to describe the criteria for a complete tool that is designed to aid in the comprehension of cloning within a software system. Furthermore, we present a prototype of such a tool and demonstrate the value of its features through a case study on the Apache httpd Web server. For example, in our study we found that a single subsystem comprising only 17% of the system code contained 38.8% of the clones. Cory Kapser, Michael W. Godfrey |
ICSM | 2 |
| 2005 | Concept Identification in Object-Oriented Domain Analysis: Why Some Students Just Don't Get ItabstractAnyone who has taught object-oriented domain analysis or any other software process requiring concept identification has undoubtedly observed that some students just don't get it. Our evaluation of the work of over 740 University of Waterloo students on over 135 software requirements specifications during the last four years supports this same observation. The students' task was to specify a telephone exchange or a voice-over-IP telephone system and the related accounts management subsystem, based on models they developed using object-oriented analysis. A detailed comparative study of three much smaller specifications, all of an elevator system, suggests that object orientation is poorly suited to domain analysis, even of small-sized domains, and that the difficulties we have observed are independent both of the size of the system under specification and of the overall abilities of the students. Davor Svetinovic, Daniel M. Berry, Michael W. Godfrey |
RE | 3 |
| 2005 | Facilitating software evolution research with kenyonabstractSoftware evolution research inherently has several resource-intensive logistical constraints. Archived project artifacts, such as those found in source code repositories and bug tracking systems, are the principal source of input data. Analysis-specific facts, such as commit metadata or the location of design patterns within the code, must be extracted for each change or configuration of interest. The results of this resource-intensive "fact extraction" phase must be stored efficiently, for later use by more experimental types of research tasks, such as algorithm or model refinement. In order to perform any type of software evolution research, each of these logistical issues must be addressed and an implementation to manage it created. In this paper, we introduce Kenyon, a system designed to facilitate software evolution research by providing a common set of solutions to these common logistical problems. We have used Kenyon for processing source code data from 12 systems of varying sizes and domains, archived in 3 different types of software configuration management systems. We present our experiences using Kenyon with these systems, and also describe Kenyon's usage by students in a graduate seminar class. Jennifer Bevan, E. James Whitehead Jr., Sunghun Kim 0001, Michael W. Godfrey |
ESEC/SIGSOFT FSE | 4 |
| 2005 | Using Origin Analysis to Detect Merging and Splitting of Source Code EntitiesabstractMerging and splitting source code entities is a common activity during the lifespan of a software system; as developers rethink the essential structure of a system or plan for a new evolutionary direction, so must they be able to reorganize the design artifacts at various abstraction levels as seems appropriate. However, while the raw effects of such changes may be plainly evident in the new artifacts, the original context of the design changes is often lost. That is, it may be obvious which characters of which files have changed, but it may not be obvious where or why moving, renaming, merging, and/or splitting of design elements has occurred. In this paper, we discuss how we have extended origin analysis (Q. Tu et al., 2002), (M.W. Godfrey et al., 2002) to aid in the detection of merging and splitting of files and functions in procedural code; in particular, we show how reasoning about how call relationships have changed can aid a developer in locating where merges and splits have occurred, thereby helping to recover some information about the context of the design change. We also describe a case study of these techniques (as implemented in the Beagle tool) using the PostgreSQL database system as the subject. Michael W. Godfrey, Lijie Zou |
IEEE Trans. Software Eng. | 1 |
| 2004 | Analyzing the Evolution of Large-Scale SoftwareabstractAnalyzing the Evolution of Large-Scale SoftwareThis special issue reports on approaches that address the issue of analyzing and visualizing the evolution of large-scale software.These approaches were discussed during the IEEE Workshop on Evolution of Large-Scale Industrial Software Evolution (ELISA), which was co-located with the International Conference on Software Maintenance in Amsterdam in September 2003.After a thorough review process, two articles have been selected for this special issue.The techniques proposed and illustrated in these articles allow us to gain a deeper insight into the nature of the evolution of large-scale software. THE WORKSHOPThe IEEE ELISA workshop was co-located with the International Conference on Software Maintenance in Amsterdam in September 2003.It was organized by Tom Mens, Juan F. Ramil, Michael W. Godfrey and Brian Down.The main theme of the workshop was planned to be the evolution of large-scale industrial software.Wider themes, such as the evolution of open source software, were also well represented in the workshop submissions.The goal was to achieve a deeper and wider insight into the problems posed by the evolution of software and the possible technological and managerial solutions.This goal is particularly relevant in current-day software practice, since size issues compound the challenges to achieving disciplined software evolution.Industrial and open-source software generally consist of numerous software artifacts that need to be evolved in a harmonious fashion, change requests that arrive more quickly than can be reasonably implemented, and the involvement of several teams implementing the evolution.Typical examples of large-scale software include air traffic control, popular PC operating systems, and telephone switch software.The challenges posed by the continual evolution of these and similar software clearly challenge the current state-of-the-art.In total, the ELISA workshop received 19 workshop submissions, nine of which were short papers and 10 were full papers: 15 of the submissions were classified as research papers, the remaining four as experience papers.One submission was rejected because it was considered as being of insufficient quality. Tom Mens, Juan Fernández-Ramil, Michael W. Godfrey |
J. Softw. Maintenance Res. Pract. | 3 |
| 2001 | The Build-Time Software Architecture ViewabstractResearch and practice in the application of software architecture has reaffirmed the need to consider software systems from several distinct points of view. Previous work by P. Kruchten (1995) and C. Hofmeister et al. (2000) suggests that four or five points of view may be sufficient: the logical view (i.e., the domain object model), the (static) code view, the process/concurrency view, the deployment/execution view, plus scenarios and use-cases. We have found that some classes of software systems exhibit interesting and complex build-time properties that are not explicitly addressed by previous models. In this paper, we present the idea of build-time architectural views. We explain what they are, how to represent them, and how they fit into traditional models of software architecture. We present three case studies of software systems with interesting build-time architectural views, and show how modelling their build-time architectures can improve developer understanding of what the system is and how it is created. Finally, we introduce a new architectural style, the "code robot" that is often present in systems with interesting build-time views. Qiang Tu, Michael W. Godfrey |
ICSM | 2 |
| 2000 | Evolution in Open Source Software: A Case StudyabstractMost studies of software evolution have been performed on systems developed within a single company using traditional management techniques. With the widespread availability of several large software systems that have been developed using an "open source" development approach, we now have a chance to examine these systems in detail, and see if their evolutionary narratives are significantly different from commercially developed systems. The paper summarizes our preliminary investigations into the evolution of the best known open source system: the Linux operating system kernel. Because Linux is large (over two million lines of code in the most recent version) and because its development model is not as tightly planned and managed as most industrial software processes, we had expected to find that Linux was growing more slowly as it got bigger and more complex. Instead, we have found that Linux has been growing at a super-linear rate for several years. The authors explore the evolution of the Linux kernel both at the system level and within the major subsystems, and they discuss why they think Linux continues to exhibit such strong growth. Michael W. Godfrey, Qiang Tu |
ICSM | 1 |
| 2000 | Connecting architecture reconstruction frameworks
Ivan T. Bowman, Michael W. Godfrey, Richard C. Holt |
Inf. Softw. Technol. | 2 |
| 1999 | JDuck: building a software engineering tool in Java as a CS2 projectabstractThis paper describes our experiences in having students build a software engineering tool as a course project in a CS2 course. The tool, which we called JDuck (Java Documenter of Code, oK), was modelled on the javadoc tool that is part of Sun Microsystem's standard Java Development Kit (JDK). That is, a working version of JDuck would be able to read in Java source code and generate HTML files that summarize the basic structure of the provided classes. We discuss how we set up the project, what we think the students learned, what they told us they learned, and what we would do differently next time. Michael W. Godfrey, Dan Grossman |
SIGCSE | 1 |
| 1998 | Secure and Portable Database ExtensibilityabstractThe functionality of extensible database servers can be augmented by user-defined functions (UDFs). However, the server's security and stability are concerns whenever new code is incorporated. Recently, there has been interest in the use of Java for database extensibility. This raises several questions: Does Java solve the security problems? How does it affect efficiency? Michael W. Godfrey, Tobias Mayr 0001, Praveen Seshadri, Thorsten von Eicken |
SIGMOD Conference | 1 |
| 1998 | Teaching software engineering to a mixed audienceabstractThis paper describes some observations derived from teaching a course in software engineering to a mixed audience of undergraduates and professional Master's degree students at Cornell University. We describe our philosophical goals in teaching the course, some of the problems we encountered, some of the unexpected results, and what we intend to do differently next time. Michael W. Godfrey |
Inf. Softw. Technol. | 1 |