Michael Sailer

dblp:140/4722 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Prioritizing Test Gaps by Risk in Industrial Practice: An Automated Approach and Multimethod Study
abstract
Context.Untested code changes, calledtest gaps, pose a significant risk for software projects. Since test gaps increase the probability of defects, managing test gaps and their individual risk is important, especially for rapidly changing software systems.Objective.This study aims at gaining an understanding of test gaps in industrial practice establishing criteria for precise prioritization of test gaps by their risk, informing practitioners that need to manage, review, and act on larger sets of test gaps.Method.We propose an automated approach for prioritizing test gaps based on key risk criteria. By means of an analysis of 31 historical test gap reviews from 8 industrial software systems of our industrial partners Munich Re and LV 1871, and by conducting semi-structured interviews with the 6 quality engineers that authored the historical test gap reviews, we validate the transferability of the identified risk criteria, such as code criticality and complexity metrics.Results.Our automated approach exhibits a ranking performance equivalent to expert assessments, in that test gaps labelled as risky in historical test gap reviews are prioritized correctly, on average, on the 30th percentile. In some scenarios, our automated ranking system even outpaces expert assessments, especially for test gaps in central code—for non-developers an opaque code property.Conclusion.This research underscores the industrial need of test gap risk estimation techniques to assist test management and quality assurance teams in identifying and addressing critical test gaps. Our multimethod study shows that even a lightweight prioritization approach helps practitioners to identify high-risk test gaps efficiently and to filter out low-risk test gaps.
Roman Haas, Michael Sailer, Mitchell Joblin, Elmar Jürgens, Sven Apel
IEEE Trans. Software Eng.2
2024 Cost of Flaky Tests in Continuous Integration: An Industrial Case Study
abstract
Researchers and practitioners alike increasingly often perceive flaky tests as a major challenge in software engineering. They spend a lot of effort trying to detect, repair, and mitigate the negative effects of flaky tests. However, it is yet unclear where and to what extent the costs of flaky tests manifest in industrial Continuous Integration (CI) development processes. In this study, we compile cost factors introduced by flaky tests in CI development from research and practice and derive a cost model that allows gaining insight into the costs incurred. We then instantiate this model in a case study of a large, commercial software project with ~30 developers and ~1M SLoC. We analyze five years of development history, including CI test logs, commits from the Version Control System (VCS), issue tickets, and tracked work time to quantify the cost factors implied by flaky tests. We find that the time spent dealing with flaky tests in the studied project represents at least 2.5% of the productive developer time. This effort is divided into investigating potentially flaky test failures, which accounts for 1.1% of the total time spent, repairing flaky tests adds another 1.3 %, and developing tools to monitor flaky tests adds 0.1 %. Contrary to most other studies, we find the cost for rerunning tests to be negligible and inexpensive. Automatically rerunning a test costs 0.02 cents, while not rerunning and thus letting the pipeline fail results in a manual investigation costing $5.67 in our context. The insights gained from our case study have led to the decision to shift effort from investigation and repair to automatically rerunning tests. Our cost model can help practitioners analyze the cost of flaky tests in their context and make informed decisions. Furthermore, our case study provides a first step to better understand the costs of flaky tests, which can lead researchers to industry-relevant problems.
Fabian Leinen, Daniel Elsner, Alexander Pretschner, Andreas Stahlbauer, Michael Sailer, Elmar Jürgens
ICST5
2019 Analysis of Automatic Annotation Suggestions for Hard Discourse-Level Tasks in Expert Domains
abstract
Claudia Schulz, Christian M. Meyer, Jan Kiesewetter, Michael Sailer, Elisabeth Bauer, Martin R. Fischer, Frank Fischer, Iryna Gurevych. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Claudia Schulz 0001, Christian M. Meyer, Jan Kiesewetter, Michael Sailer, Elisabeth Bauer, Martin R. Fischer, Frank Fischer 0001, Iryna Gurevych
ACL (1)4
2018 Automatic Recommendations for Data Coding: A Use Case from Medical and Teacher Education
abstract
Research in social sciences and humanities of ten involves analysing data to draw scientific conclusions. This however requires the manual coding of the data, which is highly time-consuming. A use case is the coding of students' essays in education to draw conclusions about students' reasoning and argumentation. The NeuralWeb API tackles this problem by providing automatic codings to other software components. These codings can for example be used in annotation platforms in terms of recommendations for expert coders from social sciences and humanities. After some initial manual annotations, the expert coders then merely need to verify the correctness of the automatic codings instead of manually annotating all data.
Claudia Schulz 0001, Michael Sailer, Jan Kiesewetter, Elisabeth Bauer, Frank Fischer 0001, Martin R. Fischer, Iryna Gurevych
eScience2