Michael Hilton 0001

dblp:116/7598 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0001-9195-6902ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 24 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures
abstract
Flaky tests produce inconsistent outcomes without code changes, creating major challenges for software developers. An industrial case study reported that developers spend 1.28% of their time repairing flaky tests at a monthly cost of $2,250. This paper reveals that flaky tests often exist in clusters, with co-occurring failures that share the same root causes, which we call systemic flakiness. This result suggests that developers can reduce test repair costs by addressing shared root causes, enabling them to fix multiple flaky tests at once rather than tackling them individually. This study represents an inflection point by challenging the deep-seated assumption that flaky test failures are isolated occurrences. We used an established dataset of 10,000 test suite runs from 24 Java projects on GitHub, spanning domains from data orchestration to job scheduling. Using a data set that contains 810 flaky tests, we performed a mixed-method empirical analysis of co-occurring flaky test failures, revealing that systemic flakiness is significant and widespread.
Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn
EASE3
2024 230,439 Test Failures Later: An Empirical Evaluation of Flaky Failure Classifiers
abstract
Flaky tests are tests that can non-deterministically pass or fail, even in the absence of code changes. Despite being a source of false alarms, flaky tests often remain in test suites once they are detected, as they also may be relied upon to detect true failures. Hence, a key open problem in flaky test research is: How to quickly determine if a test failed due to flakiness, or if it detected a bug? The state-of-the-practice is for developers to re-run failing tests: if a test fails and then passes, it is flaky by definition; if the test persistently fails, it is likely a true failure. However, this approach can be both ineffective and inefficient. An alternate approach that developers may already use for triaging test failures is failure de-duplication, which matches newly discovered test failures to previously witnessed flaky and true failures. However, because flaky test failure symptoms might resemble those of true failures, there is a risk of miss classifying a true test failure as a flaky failure to be ignored. Using a dataset of 498 flaky tests from 22 open-source Java projects, we collect a large dataset of 230,439 failure messages (both flaky and not), allowing us to empirically investigate the efficacy of failure de-duplication. We find that for some projects, this approach is extremely effective (with 100% specificity), while for other projects, the approach is entirely ineffective. By analyzing the characteristics of these flaky and non-flaky failures, we provide useful guidance on how developers should rely on this approach.
Abdulrahman Alshammari, Paul Ammann, Michael Hilton 0001, Jonathan Bell 0001
ICST3
2024 Trust in Generative AI among Students: An exploratory study
abstract
Generative Artificial Intelligence (GenAI) systems have experienced exponential growth in the last couple of years. These systems offer exciting capabilities for CS Education (CSEd), such as generating programs, that students can well utilize for their learning. Among the many dimensions that might affect the effective adoption of GenAI for CSEd, in this paper, we investigate students' trust. Trust in GenAI influences the extent to which students adopt GenAI, in turn affecting their learning. In this paper, we present results from a survey of 253 students at two large universities to understand how much they trust GenAI tools and their feedback on how GenAI impacts their performance in CS courses. Our results show that students have different levels of trust in GenAI. We also observe different levels of confidence and motivation, highlighting the need for further understanding of factors impacting trust.
Matin Amoozadeh, David Daniels, Daye Nam, Stella Chen, Michael Hilton 0001, Sruti Srinivasa Ragavan, Mohammad Amin Alipour
SIGCSE (1)6
2024 Improving Software Engineering Teamwork with Structured Feedback
abstract
Teamwork is a key learning outcome for our course, Foundations of Software Engineering. However, conflicts are inevitable in teams, and if students cannot resolve conflicts, this can lead to decreased satisfaction for everyone on the team. In this experience report, we present our approach to help students deal with conflicts that can occur in team projects. Working with faculty from the School of Business, following best practices of organizational behaviour research, we instructed students on how to provide high quality peer feedback, and designed activities where students provided feedback to each other following these principles. After this intervention, we compared the results of our teamwork survey with the results from the previous semester. We saw a meaningful drop in teamwork problems directly after the intervention, and the effect persisted for the rest of the semester.
Victor Weiqi Huang, Kori Krueger, Taya Cohen, Michael Hilton 0001
SIGCSE (1)4
2023 Towards Characterizing Trust in Generative Artificial Intelligence among Students
abstract
No abstract available.
Matin Amoozadeh, David Daniels, Stella Chen, Daye Nam, Michael Hilton 0001, Mohammad Amin Alipour, Sruti Srinivasa Ragavan
ICER (2)6
2023 Empirically evaluating flaky test detection techniques combining test case rerunning and machine learning models
abstract
Abstract A flaky test is a test case whose outcome changes without modification to the code of the test case or the program under test. These tests disrupt continuous integration, cause a loss of developer productivity, and limit the efficiency of testing. Many flaky test detection techniques are rerunning-based, meaning they require repeated test case executions at a considerable time cost, or are machine learning-based, and thus they are fast but offer only an approximate solution with variable detection performance. These two extremes leave developers with a stark choice. This paper introduces CANNIER, an approach for reducing the time cost of rerunning-based detection techniques by combining them with machine learning models. The empirical evaluation involving 89,668 test cases from 30 Python projects demonstrates that CANNIER can reduce the time cost of existing rerunning-based techniques by an order of magnitude while maintaining a detection performance that is significantly better than machine learning models alone. Furthermore, the comprehensive study extends existing work on machine learning-based detection and reveals a number of additional findings, including (1) the performance of machine learning models for detecting polluter test cases; (2) using the mean values of dynamic test case features from repeated measurements can slightly improve the detection performance of machine learning models; and (3) correlations between various test case features and the probability of the test case being flaky.
Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn
Empir. Softw. Eng.3
2022 What Do Developer-Repaired Flaky Tests Tell Us About the Effectiveness of Automated Flaky Test Detection?
abstract
Because they pass or fail without code changes, flaky tests cause serious problems such as spuriously failing builds and the eroding of developers' trust in tests. Many previous evaluations of automated flaky test detection techniques do not accurately assess their usefulness for the developers who identify the flaky tests to repair. This is because researchers evaluate detection techniques against baselines that are not derived from past developer behavior or against no baselines at all. To study the effectiveness of an automated test rerunning technique, a common baseline for other approaches to detection, this paper uses 75 commits --- authored by human software developers --- that repair test flakiness in 31 real-world Python projects. Surprisingly, automated rerunning detects the developer-repaired flaky tests in only 40% of the studied commits. This result suggests that automated rerunning does not often find those flaky tests that developers fix, implying that it makes an unsuitable baseline for assessing a detection technique's usefulness for developers.
Owain Parry, Michael Hilton 0001, Gregory M. Kapfhammer, Phil McMinn
AST2
2022 Evaluating Features for Machine Learning Detection of Order- and Non-Order-Dependent Flaky Tests
abstract
Flaky tests are test cases that can pass or fail without code changes. They often waste the time of software developers and obstruct the use of continuous integration. Previous work has presented several automated techniques for detecting flaky tests, though many involve repeated test executions and a lot of source code instrumentation and thus may be both intrusive and expensive. While this motivates researchers to evaluate machine learning models for detecting flaky tests, prior work on the features used to encode a test case is limited. Without further study of this topic, machine learning models cannot perform to their full potential in this domain. Previous studies also exclude a specific, yet prevalent and problematic, category of flaky tests: order-dependent (OD) flaky tests. This means that prior research only addresses part of the challenge of detecting flaky tests with machine learning. Closing this knowledge gap, this paper presents a new feature set for encoding tests, called Flake16. Using 54 distinct pipelines of data preprocessing, data balancing, and machine learning models for detecting both non-order-dependent (NOD) and OD flaky tests, this paper compares Flake16 to another well-established feature set. To assess the new feature set's effectiveness, this paper's experiments use the test suites of 26 Python projects, consisting of over 67,000 tests. Along with identifying the most impactful metrics for using machine learning to detect both types of flaky test, the empirical study shows how Flake16 is better than prior work, including (1) a 13% increase in overall F1 score when detecting NOD flaky tests and (2) a 17% increase in overall F1 score when detecting OD flaky tests.
Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn
ICST3
2022 A retrospective study of one decade of artifact evaluations
abstract
Most software engineering research involves the development of a prototype, a proof of concept, or a measurement apparatus. Together with the data collected in the research process, they are collectively referred to as research artifacts and are subject to artifact evaluation (AE) at scientific conferences. Since its initiation in the SE community at ESEC/FSE 2011, both the goals and the process of AE have evolved and today expectations towards AE are strongly linked with reproducible research results and reusable tools that other researchers can build their work on. However, to date little evidence has been provided that artifacts which have passed AE actually live up to these high expectations, i.e., to which degree AE processes contribute to AE's goals and whether the overhead they impose is justified.
Stefan Winter 0001, Christopher Steven Timperley, Ben Hermann, Jürgen Cito, Jonathan Bell 0001, Michael Hilton 0001, Dirk Beyer 0001
ESEC/SIGSOFT FSE6
2022 A Survey of Flaky Tests
abstract
Tests that fail inconsistently, without changes to the code under test, are described as flaky . Flaky tests do not give a clear indication of the presence of software bugs and thus limit the reliability of the test suites that contain them. A recent survey of software developers found that 59% claimed to deal with flaky tests on a monthly, weekly, or daily basis. As well as being detrimental to developers, flaky tests have also been shown to limit the applicability of useful techniques in software testing research. In general, one can think of flaky tests as being a threat to the validity of any methodology that assumes the outcome of a test only depends on the source code it covers. In this article, we systematically survey the body of literature relevant to flaky test research, amounting to 76 papers. We split our analysis into four parts: addressing the causes of flaky tests, their costs and consequences, detection strategies, and approaches for their mitigation and repair. Our findings and their implications have consequences for how the software-testing community deals with test flakiness, pertinent to practitioners and of interest to those wanting to familiarize themselves with the research area.
Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn
ACM Trans. Softw. Eng. Methodol.3
2021 FlakeFlagger: Predicting Flakiness Without Rerunning Tests
abstract
When developers make changes to their code, they typically run regression tests to detect if their recent changes (re) introduce any bugs. However, many tests are flaky, and their outcomes can change non-deterministically, failing without apparent cause. Flaky tests are a significant nuisance in the development process, since they make it more difficult for developers to trust the outcome of their tests, and hence, it is important to know which tests are flaky. The traditional approach to identify flaky tests is to rerun them multiple times: if a test is observed both passing and failing on the same code, it is definitely flaky. We conducted a very large empirical study looking for flaky tests by rerunning the test suites of 24 projects 10,000 times each, and found that even with this many reruns, some previously identified flaky tests were still not detected. We propose FlakeFlagger, a novel approach that collects a set of features describing the behavior of each test, and then predicts tests that are likely to be flaky based on similar behavioral features. We found that FlakeFlagger correctly labeled as flaky at least as many tests as a state-of-the-art flaky test classifier, but that FlakeFlagger reported far fewer false positives. This lower false positive rate translates directly to saved time for researchers and developers who use the classification result to guide more expensive flaky test detection processes. Evaluated on our dataset of 23 projects with flaky tests, FlakeFlagger outperformed the prior approach (by F1 score) on 16 projects and tied on 4 projects. Our results indicate that this approach can be effective for identifying likely flaky tests prior to running time-consuming flaky test detectors.
Abdulrahman Alshammari, Christopher Morris 0003, Michael Hilton 0001, Jonathan Bell 0001
ICSE3
2021 Combining Collaborative Reflection based on Worked-Out Examples with Problem-Solving Practice: Designing Collaborative Programming Projects for Learning at Scale
abstract
Computer science pedagogy has overwhelmingly favored problem-solving practice over methods of engagement like worked-out example study especially in advanced classes. This is due to the belief that while these alternative methods may improve student conceptual learning, they may leave them less able to perform on authentic problem-solving tasks from a lack of hands-on practice. In this paper, we perform a direct comparison of this trade-off in a synchronous collaborative programming project by adjusting the boundary between problem-solving and collaborative reflection based on a worked-out example while keeping the total time on task constant. We find that the more time students spent on worked example study, the more was the observed improvement in the pre- to post-test scores with no significant difference in performance on a subsequent problem-solving task. These results, therefore, challenge the dominant place of problem-solving practice in the advanced curricular context and inform the design of collaborative programming projects at scale.
Sreecharan Sankaranarayanan, Siddharth Reddy Kandimalla, Christopher Bogart, R. Charles Murray, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
L@S5
2021 Understanding and improving artifact sharing in software engineering research
abstract
In recent years, many software engineering researchers have begun to include artifacts alongside their research papers. Ideally, artifacts, including tools, benchmarks, and data, support the dissemination of ideas, provide evidence for research claims, and serve as a starting point for future research. However, in practice, artifacts suffer from a variety of issues that prevent the realization of their full potential. To help the software engineering community realize the full potential of artifacts, we seek to understand the challenges involved in the creation, sharing, and use of artifacts. To that end, we perform a mixed-methods study including a survey of artifacts in software engineering publications, and an online survey of 153 software engineering researchers. By analyzing the perspectives of artifact creators, users, and reviewers, we identify several high-level challenges that affect the quality of artifacts including mismatched expectations between these groups, and a lack of sufficient reward for both creators and reviewers. Using Diffusion of Innovations (DoI) as an analytical framework, we examine how these challenges relate to one another, and build an understanding of the factors that affect the sharing and success of artifacts. Finally, we make recommendations to improve the quality of artifacts based on our results and existing best practices.
Christopher Steven Timperley, Lauren Herckis, Claire Le Goues, Michael Hilton 0001
Empir. Softw. Eng.4
2020 Agent-in-the-Loop: Conversational Agent Support in Service of Reflection for Learning During Collaborative Programming
Sreecharan Sankaranarayanan, Siddharth Reddy Kandimalla, Sahil Hasan, Haokang An, Christopher Bogart, R. Charles Murray, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
AIED (2)7
2020 It Takes a Village to Build a Robot: An Empirical Study of The ROS Ecosystem
abstract
Over the past eleven years, the Robot Operating System (ROS), has grown from a small research project into the most popular framework for robotics development. Composed of packages released on the Rosdistro package manager, ROS aims to simplify development by providing reusable libraries, tools and conventions for building a robot. Still, developing a complete robot is a difficult task that involves bridging many technical disciplines. Experts who create computer vision packages, for instance, may need to rely on software designed by mechanical engineers to implement motor control. As building a robot requires domain expertise in software, mechanical, and electrical engineering, as well as artificial intelligence and robotics, ROS faces knowledge based barriers to collaboration.In this paper, we examine how the necessity of domain specific knowledge impacts the open source collaboration model. We create a comprehensive corpus of package metadata and dependencies over three years in the ROS ecosystem, analyze how collaboration is structured, and study the dependency network evolution. We find that the most widely used ROS packages belong to a small cluster of foundational working groups (FWGs), each organized around a different domain in robotics. We show that the FWGs are growing at a slower rate than the rest of the ecosystem, in terms of their membership and number of packages, yet the number of dependencies on FWGs is increasing at a faster rate. In addition, we mined all ROS packages on GitHub, and showed that 82% rely exclusively on functionality provided by FWGs. Finally, we investigate these highly influential groups and describe the unique model of collaboration they support in ROS.
Sophia Kolak, Afsoon Afzal, Claire Le Goues, Michael Hilton 0001, Christopher Steven Timperley
ICSME4
2020 A Study on Challenges of Testing Robotic Systems
abstract
Robotic systems are increasingly a part of everyday life. Characteristics of robotic systems such as interaction with the physical world, and integration of hardware and software components, differentiate robotic systems from conventional software systems. Although numerous studies have investigated the challenges of software testing in practice, no such study has focused on testing of robotic systems. In this paper, we conduct a qualitative study to better understand the testing practices used by the robotics community, and identify the challenges faced by practitioners when testing their systems. We identify a total of 12 testing practices and 9 testing challenges from our participants' responses. We group these challenges into 3 major themes: Real-world complexities, Community and standards, and Component integration. We believe that further research on addressing challenges described with these three major themes can result in higher adoption of robotics testing practices, more testing automation, and higher-quality robotic systems.
Afsoon Afzal, Claire Le Goues, Michael Hilton 0001, Christopher Steven Timperley
ICST3
2020 Empirical Study of Restarted and Flaky Builds on Travis CI
abstract
Continuous Integration (CI) is a development practice where developers frequently integrate code into a common codebase. After the code is integrated, the CI server runs a test suite and other tools to produce a set of reports (e.g., the output of linters and tests). If the result of a CI test run is unexpected, developers have the option to manually restart the build, re-running the same test suite on the same code; this can reveal build flakiness, if the restarted build outcome differs from the original build.
Thomas Durieux, Claire Le Goues, Michael Hilton 0001, Rui Abreu 0001
MSR3
2019 An Intelligent-Agent Facilitated Scaffold for Fostering Reflection in a Team-Based Project Course
Sreecharan Sankaranarayanan, Xu Wang 0016, Cameron Dashti, Marshall An, Clarence Ngoh, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
AIED (2)6
2019 Graph-based mining of in-the-wild, fine-grained, semantic code change patterns
abstract
Prior research exploited the repetitiveness of code changes to enable several tasks such as code completion, bug-fix recommendation, library adaption, etc. These and other novel applications require accurate detection of semantic changes, but the state-of-the-art methods are limited to algorithms that detect specific kinds of changes at the syntactic level. Existing algorithms relying on syntactic similarity have lower accuracy, and cannot effectively detect semantic change patterns. We introduce a novel graph-based mining approach, CPatMiner, to detect previously unknown repetitive changes in the wild, by mining fine-grained semantic code change patterns from a large number of repositories. To overcome unique challenges such as detecting meaningful change patterns and scaling to large repositories, we rely on fine-grained change graphs to capture program dependencies. We evaluate CPatMiner by mining change patterns in a diverse corpus of 5,000+ open-source projects from GitHub across a population of 170,000+ developers. We use three complementary methods. First, we sent the mined patterns to 108 open-source developers. We found that 70% of respondents recognized those patterns as their meaningful frequent changes. Moreover, 79% of respondents even named the patterns, and 44% wanted future IDEs to automate such repetitive changes. We found that the mined change patterns belong to various development activities: adaptive (9%), perfective (20%), corrective (35%) and preventive (36%, including refactorings). Second, we compared our tool with the state-of-the-art, AST-based technique, and reported that it detects 2.1x more meaningful patterns. Third, we use CPatMiner to search for patterns in a corpus of 88 GitHub projects with longer histories consisting of 164M SLOCs. It constructed 322K fine-grained change graphs containing 3M nodes, and detected 17K instances of change patterns from which we provide unique insights on the practice of change patterns among individuals and teams. We found that a large percentage (75%) of the change patterns from individual developers are commonly shared with others, and this holds true for teams. Moreover, we found that the patterns are not intermittent but spread widely over time. Thus, we call for a community-based change pattern database to provide important resources in novel applications.
Hoan Anh Nguyen, Tien N. Nguyen, Danny Dig, Hieu Tran, Michael Hilton 0001
ICSE6
2019 The Problem of Packaging Curricular Materials
abstract
Packaging materials is a generalized term to capture a broad array of tasks (creating, revising, sharing, finding, crediting, etc.) for materials such as assignments, teacher notes, and evaluation data. Substantial effort has gone into creating materials over the years, but the community still struggles to find ways to effectively manage these. This BoF provides an opportunity to identify needs, concerns, prior efforts, and future plans. A primary goal is the formation of a Working Group tasked to develop a standard for curricular material creation and sharing, joining with broader efforts of standardization (e.g., CSSPLICE) and existing initiatives for creating repositories, tools, and materials.
Austin Cory Bart, Michael Hilton 0001, Bob Edmison, Phillip T. Conrad
SIGCSE2
2019 Online Mob Programming: Effective Collaborative Project-Based Learning
abstract
This lightning talk presents an ongoing effort investigating the use of a collaborative programming paradigm originating in industry called Mob Programming, for effective collaborative learning in the classroom. In industry, Mob Programming involves the participation of a group of developers solving one problem at the same time and place. Developers take turns cycling through a structured process for collaboration having been assigned pre-defined roles responsible for brainstorming ideas, deciding on a path forward and implementing the consensus decision respectively. Pedagogically, there are several compelling reasons to motivate the adoption of Mob Programming in learning settings . First, the collaboration is well-structured meaning that the interaction between even a large group of students will not descend into chaos. Second, students are assigned to roles, which allows for the differentiation of responsibilities and development of skills in different aspects of solving the problem. Third, the rotation of assigned roles allows students to learn and exhibit multiple competencies as well as appreciate bringing different perspectives to bear on solving the problem. In order to investigate whether this promise is borne out in practice, the paradigm is currently being investigated in the context of a Cloud Computing course offered online to undergraduate and graduate students at a large American university and its satellite campuses. Since this effort is still underway, faculty who implement or are interested in implementing collaborative learning for this classrooms are invited to attend and provide feedback or consider joining the effort to investigate this paradigm for use in learning settings.
Michael Hilton 0001, Sreecharan Sankaranarayanan
SIGCSE1
2019 A conceptual replication of continuous integration pain points in the context of Travis CI
abstract
Continuous integration (CI) is an established software quality assurance practice, and the focus of much prior research with a diverse range of methods and populations. In this paper, we first conduct a literature review of 37 papers on CI pain points. We then conduct a conceptual replication study on results from these papers using a triangulation design consisting of a survey with 132 responses, 12 interviews, and two logistic regressions predicting Travis CI abandonment and switching on a dataset of 6,239 GitHub projects. We report and discuss which past results we were able to replicate, those for which we found conflicting evidence, those for which we did not find evidence, and the implications of these findings.
David Gray Widder, Michael Hilton 0001, Christian Kästner, Bogdan Vasilescu
ESEC/SIGSOFT FSE2
2018 DeFlaker: automatically detecting flaky tests
abstract
Developers often run tests to check that their latest changes to a code repository did not break any previously working functionality. Ideally, any new test failures would indicate regressions caused by the latest changes. However, some test failures may not be due to the latest changes but due to non-determinism in the tests, popularly called flaky tests. The typical way to detect flaky tests is to rerun failing tests repeatedly. Unfortunately, rerunning failing tests can be costly and can slow down the development cycle.
Jonathan Bell 0001, Owolabi Legunsen, Michael Hilton 0001, Lamyaa Eloussi, Tifany Yung, Darko Marinov
ICSE3
2018 A large-scale study of test coverage evolution
abstract
Statement coverage is commonly used as a measure of test suite quality. Coverage is often used as a part of a code review process: if a patch decreases overall coverage, or is itself not covered, then the patch is scrutinized more closely. Traditional studies of how coverage changes with code evolution have examined the overall coverage of the entire program, and more recent work directly examines the coverage of patches (changed statements). We present an evaluation much larger than prior studies and moreover consider a new, important kind of change --- coverage changes of unchanged statements. We present a large-scale evaluation of code coverage evolution over 7,816 builds of 47 projects written in popular languages including Java, Python, and Scala. We find that in large, mature projects, simply measuring the change to statement coverage does not capture the nuances of code evolution. Going beyond considering statement coverage as a simple ratio, we examine how the set of statements covered evolves between project revisions. We present and study new ways to assess the impact of a patch on a project's test suite quality that both separates coverage of the patch from coverage of the non-patch, and separates changes in coverage from changes in the set of statements covered.
Michael Hilton 0001, Jonathan Bell 0001, Darko Marinov
ASE1
2018 I'm leaving you, Travis: a continuous integration breakup story
abstract
Continuous Integration (CI) services, which can automatically build, test, and deploy software projects, are an invaluable asset in distributed teams, increasing productivity and helping to maintain code quality. Prior work has shown that CI pipelines can be sophisticated, and choosing and configuring a CI system involves tradeoffs. As CI technology matures, new CI tool offerings arise to meet the distinct wants and needs of software teams, as they negotiate a path through these tradeoffs, depending on their context. In this paper, we begin to uncover these nuances, and tell the story of open-source projects falling out of love with Travis, the earliest and most popular cloud-based CI system. Using logistic regression, we quantify the effects that open-source community factors and project technical factors have on the rate of Travis abandonment. We find that increased build complexity reduces the chances of abandonment, that larger projects abandon at higher rates, and that a project's dominant language has significant but varying effects. Finally, we find the surprising result that metrics of configuration attempts and knowledge dispersion in the project do not affect the rate of abandonment.
David Gray Widder, Michael Hilton 0001, Christian Kästner, Bogdan Vasilescu
MSR2
2017 Hazelnut: a bidirectionally typed structure editor calculus
abstract
Structure editors allow programmers to edit the tree structure of a program directly. This can have cognitive benefits, particularly for novice and end-user programmers. It also simplifies matters for tool designers, because they do not need to contend with malformed program text.
Cyrus Omar, Ian Voysey, Michael Hilton 0001, Jonathan Aldrich, Matthew A. Hammer
POPL3
2017 Trade-offs in continuous integration: assurance, security, and flexibility
abstract
Continuous integration (CI) systems automate the compilation, building, and testing of software. Despite CI being a widely used activity in software engineering, we do not know what motivates developers to use CI, and what barriers and unmet needs they face. Without such knowledge, developers make easily avoidable errors, tool builders invest in the wrong direction, and researchers miss opportunities for improving the practice of CI. We present a qualitative study of the barriers and needs developers face when using CI. We conduct semi-structured interviews with developers from different industries and development scales. We triangulate our findings by running two surveys. We find that developers face trade-offs between speed and certainty (Assurance), between better access and information security (Security), and between more configuration options and greater ease of use (Flexi- bility). We present implications of these trade-offs for developers, tool builders, and researchers.
Michael Hilton 0001, Nicholas Nelson 0002, Timothy Tunnell, Darko Marinov, Danny Dig
ESEC/SIGSOFT FSE1
2016 Usage, costs, and benefits of continuous integration in open-source projects
abstract
Continuous integration (CI) systems automate the compilation, building, and testing of software. Despite CI rising as a big success story in automated software engineering, it has received almost no attention from the research community. For example, how widely is CI used in practice, and what are some costs and benefits associated with CI? Without answering such questions, developers, tool builders, and researchers make decisions based on folklore instead of data. In this paper, we use three complementary methods to study the usage of CI in open-source projects. To understand which CI systems developers use, we analyzed 34,544 open-source projects from GitHub. To understand how developers use CI, we analyzed 1,529,291 builds from the most commonly used CI system. To understand why projects use or do not use CI, we surveyed 442 developers. With this data, we answered several key questions related to the usage, costs, and benefits of CI. Among our results, we show evidence that supports the claim that CI helps projects release more often, that CI is widely adopted by the most popular projects, as well as finding that the overall percentage of projects using CI continues to grow, making it important and timely to focus more research on CI.
Michael Hilton 0001, Timothy Tunnell, Darko Marinov, Danny Dig
ASE1
2016 Understanding and improving continuous integration
abstract
Continuous Integration (CI) has been widely adopted in the software development industry. However, the usage of CI in practice has been ignored for far too long by the research community. We propose to fill this blind spot by doing in- depth research into CI usage in practice. We will answer how questions by using using quantitative methods, such as investigating open source data that is publicly available. We will answer why questions using qualitative methods, such as semi-structured interviews and large scale surveys. In the course of our research, we plan on identifying barriers that developers face when using CI. We will develop techniques to overcome those barriers via automation. This work is advised by Professor Danny Dig.
Michael Hilton 0001
SIGSOFT FSE1
2016 API code recommendation using statistical learning from fine-grained changes
abstract
Learning and remembering how to use APIs is difficult. While code-completion tools can recommend API methods, browsing a long list of API method names and their documentation is tedious. Moreover, users can easily be overwhelmed with too much information. We present a novel API recommendation approach that taps into the predictive power of repetitive code changes to provide relevant API recommendations for developers. Our approach and tool, APIREC, is based on statistical learning from fine-grained code changes and from the context in which those changes were made. Our empirical evaluation shows that APIREC correctly recommends an API call in the first position 59% of the time, and it recommends the correct API call in the top five positions 77% of the time. This is a significant improvement over the state-of-the-art approaches by 30-160% for top-1 accuracy, and 10-30% for top-5 accuracy, respectively. Our result shows that APIREC performs well even with a one-time, minimal training dataset of 50 publicly available projects.
Anh Tuan Nguyen 0001, Michael Hilton 0001, Mihai Codoban, Hoan Anh Nguyen, Lily Mast, Eli Rademacher, Tien N. Nguyen, Danny Dig
SIGSOFT FSE2
2016 TDDViz: Using Software Changes to Understand Conformance to Test Driven Development
abstract
A bad software development process leads to wasted effort and inferior products. In order to improve a software process, it must be first understood. Our unique approach in this paper uses code and test changes to understand conformance to the Test Driven Development (TDD) process. We designed and implemented TDDViz , a tool that supports developers in better understanding how they conform to TDD. TDDViz supports this understanding by providing novel visualizations of developers’ TDD process. To enable TDDViz ’s visualizations, we developed a novel automatic inferencer that identifies the phases that make up the TDD process solely based on code and test changes. We evaluate TDDViz using two complementary methods: a controlled experiment with 35 participants to evaluate the visualization, and a case study with 2601 TDD Sessions to evaluate the inference algorithm. The controlled experiment shows that, in comparison to existing visualizations, participants performed significantly better when using TDDViz to answer questions about code evolution. In addition, the case study shows that the inferencing algorithm in TDDViz infers TDD phases with an accuracy (F-measure) of 87%.
Michael Hilton 0001, Nicholas Nelson 0002, Hugh McDonald, Sean McDonald, Ronald A. Metoyer, Danny Dig
XP1
2013 An evaluation of interactive test-driven labs with WebIDE in CS0
abstract
WebIDE is a framework that enables instructors to develop and deliver online lab content with interactive feedback. The ability to create lock-step labs enables the instructor to guide students through learning experiences, demonstrating mastery as they proceed. Feedback is provided through automated evaluators that vary from simple regular expression evaluation to syntactic parsers to applications that compile and run programs and unit tests. This paper describes WebIDE and its use in a CS0 course that taught introductory Java and Android programming using a test-driven learning approach. We report results from a controlled experiment that compared the use of dynamic WebIDE labs with more traditional static programming labs. Despite weaker performance on pre-study assessments, students who used WebIDE performed two to twelve percent better on all assessments than the students who used traditional labs. In addition, WebIDE students were consistently more positive about their experience in CS0.
David S. Janzen, John Clements, Michael Hilton 0001
ICSE3
2012 On teaching arrays with test-driven learning in WebIDE
abstract
Test-driven development (TDD) has been shown to reduce defects and to lead to better code, but can it help beginning students learn basic programming topics, specifically arrays? We performed a controlled experiment where we taught arrays to two CS0 classes, one using WebIDE, an intelligent tutoring system that enforced the use of Test-Driven Learning (TDL) methods, and one using more traditional static methods and a development environment that instructed, but did not enforce the use of TDD. Students who used the TDL approach with WebIDE performed significantly better in assessments and had significantly higher opinions of their experiences than students who used traditional methods and tools.
Michael Hilton 0001, David S. Janzen
ITiCSE1