VLDB 2026 Research / reviewers in the wild / expert
Birgit Hofer
dblp:90/8164
· DBLP profile ↗
30ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0001-5144-059XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Choosing Abstraction Levels for Model-Based Software Debugging: A Theoretical and Empirical Analysis for Spreadsheet Programs (Abstract Reprint)abstractModel-based diagnosis is a generally applicable, principled approach to the systematic debugging of a wide range of system types such as circuits, knowledge bases, physical devices, or software. Based on a formal description of the system, it enables precise and deterministic reasoning about potential faults responsible for observed misbehavior. In software, such a formal system description can often even be extracted from the buggy program fully automatically. As logical reasoning is central to diagnosis, the performance of model-based debuggers is largely influenced by reasoning efficiency, which in turn depends on the complexity and expressivity of the system description. Since highly detailed models capturing exact semantics often exceed the capabilities of current reasoning tools, researchers have proposed more abstract representations. In this work, we thoroughly analyze system modeling techniques with a focus on fault localization in spreadsheets—one of the most widely used end-user programming paradigms. Specifically, we present three constraint model types characterizing spreadsheets at different abstraction levels, show how to extract them automatically from faulty spreadsheets, and provide theoretical and empirical investigations of the impact of abstraction on both diagnostic output and computational performance. Our main conclusions are that (i) for the model types, there is a trade-off between the conciseness of generated fault candidates and computation time, (ii) the exact model is often impractical, and (iii) a new model based on qualitative reasoning yields the same solutions as the exact one in up to more than half the cases while being orders of magnitude faster. Due to their ability to restrict the solution space in a sound way, the explored model-based techniques, rather than being used as standalone approaches, are expected to realize their full potential in combination with iterative sequential diagnosis or indeterministic but more performant statistical debugging methods. Patrick Rodler, Birgit Hofer, Dietmar Jannach, Iulia Nica, Franz Wotawa |
AAAI | 2 |
| 2025 | Choosing abstraction levels for model-based software debugging: A theoretical and empirical analysis for spreadsheet programs
Patrick Rodler, Birgit Hofer, Dietmar Jannach, Iulia Nica, Franz Wotawa |
Artif. Intell. | 2 |
| 2025 | Best practices for evaluating IRFL approachesabstractInformation retrieval fault localization (IRFL) is a popular research field and many IRFL approaches have been proposed recently. Unfortunately, the evaluation of some of these IRFL approaches is often too simplistic, which can cause an overestimation of performance of these approaches. In this paper, we discuss evaluation pitfalls and problems. Furthermore, we propose best practices to avoid them. In detail, we discuss evaluation strategies such as parameter tuning and temporal dependencies in the data, dataset issues, metrics, statistical significance testing, and the unavailability of supplemental material. To support our claim of the poor status quo of current evaluation practices in some research papers, we have performed a literature survey on 135 papers. We hope that this paper will help researchers to avoid the described pitfalls in their evaluation of IRFL approaches. • Discussion of common pitfalls in the evaluation of IRFL approaches. • Seven best practices that help to avoid these pitfalls. • Investigation of 135 IRFL papers w.r.t. these pitfalls and best practices. Thomas Hirsch, Birgit Hofer |
J. Syst. Softw. | 2 |
| 2024 | Detecting Soft Faults in Heat Pumps (Short Paper)
Birgit Hofer, Franz Wotawa |
DX | 1 |
| 2023 | The MAP Metric in Information Retrieval Fault LocalizationabstractThe MAP (Mean Average Precision) metric is one of the most popular performance metrics in the field of Information Retrieval Fault Localization (IRFL). However, there are problematic implementations of this MAP metric used in IRFL research. These implementations deviate from the text book definitions of MAP, rendering the metric sensitive to the truncation of retrieval results and inaccuracies and impurities of the used datasets. The application of such a deviating metric can lead to performance overestimation. This can pose a problem for comparability, transferability, and validity of IRFL performance results. In this paper, we discuss the definition and mathematical properties of MAP and common deviations and pitfalls in its implementation. We investigate and discuss the conditions enabling such overestimation: the truncation of retrieval results in combination with ground truths spanning multiple files and improper handling of undefined AP results. We demonstrate the overestimation effects using the Bench4BL benchmark and five well known IRFL techniques. Our results indicate that a flawed implementation of the MAP metric can lead to an overestimation of the IRFL performance, in extreme cases by up to 70 %. We argue for a strict adherence to the text book version of MAP with the extension of undefined AP values to be set to 0 for all IRFL experiments. We hope that this work will help to improve comparability and transferability in IRFL research. Thomas Hirsch, Birgit Hofer |
ASE | 2 |
| 2023 | Explaining software fault predictions to spreadsheet usersabstractA variety of automated software fault prediction techniques was proposed in recent years, in particular for the important class of spreadsheet programs. Software fault prediction techniques commonly create ranked lists of “suspicious” program statements for developers to inspect. Existing research, however, suggests that solely providing such ranked lists may not always be effective. In particular, it was found that developers often seek for explanations for the outcomes provided by a debugging tool and that such explanations may be key for developers to trust and rely on the tool. Research on how to explain the outcomes of fault prediction techniques, which are often based on complex machine learning models, is scarce, and little is known regarding how such explanations are perceived by developers. With this work, we aim to narrow this research gap and study the perception of different forms of explanations by spreadsheet users in the context of a machine learning based fault prediction tool. A between-subjects user study (N=120) revealed significant differences between the explored explanation styles. In particular, we found that well-designed natural language explanations can indeed help users better understand why certain spreadsheet cells were marked by the debugging tool and that such explanations can be effective to increase the users’ trust compared to a black box system. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Adil Mukhtar, Birgit Hofer, Dietmar Jannach, Franz Wotawa |
J. Syst. Softw. | 2 |
| 2022 | Boosting Spectrum-Based Fault Localization for Spreadsheets with Product Metrics in a Learning ApproachabstractFaults in spreadsheets are not uncommon and they can have significant negative consequences in practice. Various approaches for fault localization were proposed in recent years, among them techniques that transferred ideas from spectrum-based fault localization (SFL) to the spreadsheet domain. Applying SFL to spreadsheets proved to be effective, but has certain limitations. Specifically, the constrained computational structures of spreadsheets may lead to large sets of cells that have the same assumed fault probability according to SFL and thus have to be inspected manually. In this work, we propose to combine SFL with a fault prediction method based on spreadsheet metrics in a machine learning (ML) approach. In particular, we train supervised ML models using two orthogonal types of features: (i) variables that are used to compute similarity coefficients in SFL and (ii) spreadsheet metrics that have shown to be good predictors for faulty formulas in previous work. Experiments with a widely-used corpus of faulty spreadsheets indicate that the combined model helps to significantly improve fault localization performance in terms of wasted effort and accuracy. Adil Mukhtar, Birgit Hofer, Dietmar Jannach, Franz Wotawa, Konstantin Schekotihin |
ASE | 2 |
| 2022 | Pruning Boolean Expressions to Shorten Dynamic SlicesabstractThis paper presents a novel extension to dynamic slicing that we call pruned slicing. The proposed slicing approach produces smaller slices than traditional dynamic slicing. This is achieved by reasoning over Boolean expressions. We have implemented a prototype in Python and empirically evaluated its performance on three different benchmarks: TCAS, QuixBugs and the Refactory dataset. We show that pruned slicing reduces the size of dynamic slices on average by 10.96 percent for TCAS. For QuixBugs and the Refactory dataset, the slice size remains the same, but the number of Boolean expressions within the slice is reduced. Further, the empirical evaluation shows that pruned dynamic slicing comes with a low computational overhead compared to dynamic slicing. Pruned slicing can also be used in combination with relevant slicing. Thomas Hirsch, Birgit Hofer |
SCAM | 2 |
| 2022 | Detecting non-natural language artifacts for de-noising bug reportsabstractTextual documents produced in the software engineering process are a popular target for natural language processing (NLP) and information retrieval (IR) approaches. However, issue tickets often contain artifacts such as code snippets, log outputs and stack traces. These artifacts not only inflate the issue ticket sizes, but also can this noise constitute a real problem for some NLP approaches, and therefore has to be removed in the pre-processing of some approaches. In this paper, we present a machine learning based approach to classify textual content into natural language and non-natural language artifacts at line level. We show how data from GitHub issue trackers can be used for automated training set generation, and present a custom preprocessing approach for the task of artifact removal. The training sets are automatically created from Markdown annotated issue tickets and project documentation files. We use these generated training sets to train a Markdown agnostic model that is able to classify un-annotated content. We evaluate our approach on issue tickets from projects written in C++, Java, JavaScript, PHP, and Python. Our approach achieves ROC-AUC scores between 0.92 and 0.96 for language-specific models. A multi-language model trained on the issue tickets of all languages achieves ROC-AUC scores between 0.92 and 0.95. The provided models are intended to be used as noise reduction pre-processing steps for NLP and IR approaches working on issue tickets. Thomas Hirsch, Birgit Hofer |
Autom. Softw. Eng. | 2 |
| 2022 | A systematic literature review on benchmarks for evaluating debugging approachesabstractBug benchmarks are used in development and evaluation of debugging approaches, e.g. fault localization and automated repair. Quantitative performance comparison of different debugging approaches is only possible when they have been evaluated on the same dataset or benchmark. However, benchmarks are often specialized towards usage for certain debugging approaches in their contained data, metrics, and artifacts. Such benchmarks cannot be easily used on debugging approaches outside their scope as such approach may rely on specific data such as bug reports or code metrics that are not included in the dataset. Furthermore, benchmarks vary in their size w.r.t. the number of subject programs and the size of the individual subject programs. For these reasons, we have performed a systematic literature review where we have identified 73 benchmarks that can be used to evaluate debugging approaches. We compare the different benchmarks w.r.t. their size and the provided information such as bug reports, contained test cases, and other code metrics. This comparison is intended to help researchers to quickly identify all suitable benchmarks for evaluating their specific debugging approaches. Furthermore, we discuss reoccurring issues and challenges in selection, acquisition, and usage of such bug benchmarks, i.e., data availability, data quality, duplicated content, data formats, reproducibility, and extensibility. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Thomas Hirsch, Birgit Hofer |
J. Syst. Softw. | 2 |
| 2022 | Spreadsheet debugging: The perils of tool over-relianceabstractSpreadsheets are widely used in organizations for various purposes such as data aggregation, reporting and decision-making. Since spreadsheets, like other types of software, can contain faulty formulas, it is important to provide developers with appropriate methods to find and fix such faults. Recently, various heuristic and statistics-based fault identification methods were proposed, which point developers to potentially faulty parts of the spreadsheets. Due to their heuristic nature, these methods might, however, miss some faults. As a result, if spreadsheet developers rely too strongly on these methods, they might not pay sufficient attention to problems that are not pinpointed by the methods. In this research, we are the first to study this potential problem of over-reliance in spreadsheet debugging, which may lead to limited debugging effectiveness. We report the outcome of a controlled experiment where 59 participants were tasked to find faulty formulas in a given spreadsheet with and without support of a novel spreadsheet debugging tool. Our results indicate that tool over-reliance can indeed result as a phenomenon of using heuristic debugging techniques. However, the study also provides evidence that making users aware of potential tool limitations within the debugging environment may help to address this problem. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Adil Mukhtar, Birgit Hofer, Dietmar Jannach, Franz Wotawa |
J. Syst. Softw. | 2 |
| 2021 | Product metrics for spreadsheets - A systematic reviewabstractSoftware product metrics allow practitioners to improve their products and to optimize development processes based on quantifiable characteristics of source code. To facilitate similar benefits for spreadsheet programs, researchers proposed various product metrics for spreadsheets over the last decades. However, to our knowledge, no comprehensive overview of those efforts is currently available. In this paper, we close this gap by conducting a literature review of research works that either inherently or explicitly define product metrics for spreadsheets. We scanned five major digital libraries for scientific papers that define or use spreadsheet product metrics. Based on the identified 37 papers, we created a novel catalog of product metrics for spreadsheets. The catalog can be used by practitioners and researchers as a central reference for spreadsheet product metrics. In the paper, we (i) describe the proposed metrics in detail, (ii) report how often and for what purposes the metrics are used, (iii) identify significant discrepancies in the naming and definition of the metrics, and (iv) investigate how the appropriateness of the metrics was evaluated. Birgit Hofer, Dietmar Jannach, Patrick W. Koch, Konstantin Schekotihin, Franz Wotawa |
J. Syst. Softw. | 1 |
| 2021 | Metric-Based Fault Prediction for SpreadsheetsabstractElectronic spreadsheets are widely used in organizations for various data analytics and decision-making tasks. Even though faults within such spreadsheets are common and can have significant negative consequences, today's tools for creating and handling spreadsheets provide limited support for fault detection, localization, and repair. Being able to predict whether a certain part of a spreadsheet is faulty or not is often central for the implementation of such supporting functionality. In this work, we propose a novel approach to fault prediction in spreadsheet formulas, which combines an extensive catalog of spreadsheet metrics with modern machine learning algorithms. An analysis of the individual metrics from our catalog reveals that they are generally suited to discover a wide range of faults. Their predictive power is, however, limited when considered in isolation. Therefore, in our approach we apply supervised learning algorithms to obtain fault predictors that utilize all data provided by multiple spreadsheet metrics from our catalog. Experiments on different datasets containing faulty spreadsheets show that particularly Random Forests classifiers are often effective. As a result, the proposed method is in many cases able to make highly accurate predictions whether a given formula of a spreadsheet is faulty.11.Results of a preliminary study were published in[1]. Patrick W. Koch, Konstantin Schekotihin, Dietmar Jannach, Birgit Hofer, Franz Wotawa |
IEEE Trans. Software Eng. | 4 |
| 2019 | Fragment-based spreadsheet debuggingabstractFaults in spreadsheets can represent a major risk for businesses. To minimize such risks, various automated testing and debugging approaches for spreadsheets were proposed. In such approaches, often one main assumption is that the spreadsheet developer is able to indicate if the outcomes of certain calculations correspond to the intended values. This, however, might require that the user performs calculations manually, a process which can easily become tedious and error-prone for more complex spreadsheets. In this work, we propose an interactive spreadsheet algorithmic debugging method, which is based on partitioning the spreadsheet into fragments. Test cases can then be automatically or manually created for each of these smaller fragments, whose correctness or faultiness can be easier assessed by users than test cases that cover the entire spreadsheet. The annotated test cases are then fed into an algorithmic debugging technique, which returns a set of formulas that could have caused any observed failures, i.e., discrepancies between the expected and computed calculation outcomes. Simulation experiments demonstrate that the suggested decomposition approach can speed up the algorithmic debugging process and significantly reduce the number of fault candidates returned by the algorithm. An additional laboratory study shows that fragmenting a spreadsheet with our method furthermore reduces the time needed by users for creating test cases for a spreadsheet. Dietmar Jannach, Thomas Schmitz 0002, Birgit Hofer, Konstantin Schekotihin, Patrick W. Koch, Franz Wotawa |
Autom. Softw. Eng. | 3 |
| 2019 | On the refinement of spreadsheet smells by means of structure information
Patrick W. Koch, Birgit Hofer, Franz Wotawa |
J. Syst. Softw. | 2 |
| 2018 | Using LNT Formal Descriptions for Model-Based Diagnosis
Birgit Hofer, Radu Mateescu 0001, Wendelin Serwe, Franz Wotawa |
DX | 1 |
| 2017 | AI for Localizing Faults in Spreadsheets
Birgit Hofer, Iulia Nica, Franz Wotawa |
ICTSS | 1 |
| 2017 | Improving Spectrum-Based Fault Localization for Spreadsheet DebuggingabstractSpreadsheets often contain faults that are difficult to localize. Spectrum-based Fault Localization (SFL) assists users in the fault localization process by ranking cells by their suspiciousness to contain a fault. Since the ranking of the basic SFL approach is often imprecise, we propose three techniques to improve it, i.e., dynamic cones, grouping, and tie-breaking. We evaluate these techniques with three spreadsheet corpora comprising more than 1,000 faulty spreadsheets of different size and structure. While dynamic cones do not come with large improvements, grouping and tie-breaking do. Grouping has a positive impact on about 50% of the spreadsheets with almost no negative effects. The same holds for tie-breaking, where some of the strategies offer high positive impact while keeping the risk of negative influences low. Elisabeth Getzner, Birgit Hofer, Franz Wotawa |
QRS | 2 |
| 2017 | A decomposition-based approach to spreadsheet testing and debuggingabstractSpreadsheets serve as a basis for decision-making processes in many companies and bugs in spreadsheets can therefore represent a considerable risk to businesses. Systematic tests can help to locate such bugs, but providing test cases can be cumbersome and complex for large real-world spreadsheets. To make the specification of test cases easier, we propose to split spreadsheets into smaller logically connected parts (called fragments) which can be individually tested for correctness. We present an algorithmic approach to compute such fragments, which we validated with a laboratory study in the form of a spreadsheet debugging exercise involving 57 subjects. The results show that the fragmentation approach can help to significantly reduce the required efforts to test a spreadsheet. Thomas Schmitz 0002, Dietmar Jannach, Birgit Hofer, Patrick W. Koch, Konstantin Schekotihin, Franz Wotawa |
VL/HCC | 3 |
| 2017 | Combining Models for Improved Fault Localization in SpreadsheetsabstractSpreadsheets are the most prominent example of end-user programing, but they unfortunately are often erroneous, and thus, they compute wrong values. Localizing the true cause of such an observed misbehavior can be cumbersome and frustrating especially for large spreadsheets. Therefore, supporting techniques and tools for fault localization are highly required. Model-based software debugging (MBSD) is a well-known technique for fault localization in software written in imperative and object-oriented programing languages like C, C++, and Java. In this paper, we explain how to use MBSD for fault localization in spreadsheets and compare three types of models for MBSD, namely the value-based model (VBM), the dependency based model (DBM), and an improved version of the DBM. Whereas the VBM computes the lowest number of diagnoses, both DBMs convince by their low computational complexity. Hence, a combination of these two types of models is desired, and we present a solution that combines value-based and DBM in this paper. Moreover, we discuss a detailed evaluation of the models and the combined approach, which indicates that the combined approach computes the same number of diagnoses like the VBMs while requiring less computation time. Hence, the proposed approach is more appropriate to be used in tools for fault localization in spreadsheets. Birgit Hofer, Andrea Hofler, Franz Wotawa |
IEEE Trans. Reliab. | 1 |
| 2015 | Focused Diagnosis for Failing Software Tests
Birgit Hofer, Seema Jehan, Ingo Pill, Franz Wotawa |
IEA/AIE | 1 |
| 2015 | Testing for Distinguishing Repair Candidates in Spreadsheets - the Mussco Approach
Rui Abreu 0001, Simon Außerlechner, Birgit Hofer, Franz Wotawa |
ICTSS | 3 |
| 2015 | Fault Localization in the Light of Faulty User InputabstractSpreadsheets may be large, containing several thousand formulas, and thus they may be hard to comprehend and analyze. Unfortunately, they are also prone to errors. Identifying the cells which are responsible for an observed error is time-consuming, tedious, and frustrating. Spectrum-based Fault Localization (SFL) helps users to faster identify those cells that have to be modified in order to eliminate any observed misbehavior. SFL requires information about the correctness of certain cell values, and users might wrongly classify such cell values. A misclassification may influence the outcome of SFL substantially. In this paper, we investigate the influence of incorrect user information on the quality of SFL. In particular, we present a theoretical analysis of the impact of a misclassification on the Ochiai similarity coefficient and an empirical evaluation based on 33 spreadsheets with 218 faulty versions. Birgit Hofer, Franz Wotawa |
QRS | 1 |
| 2015 | On the empirical evaluation of similarity coefficients for spreadsheets fault localization
Birgit Hofer, Alexandre Perez, Rui Abreu 0001, Franz Wotawa |
Autom. Softw. Eng. | 1 |
| 2015 | Using constraints to diagnose faulty spreadsheets
Rui Abreu 0001, Birgit Hofer, Alexandre Perez, Franz Wotawa |
Softw. Qual. J. | 2 |
| 2014 | Generation of Relevant Spreadsheet Repair CandidatesabstractSpreadsheets are amongst the most successful examples of end user programming. Because of their, still increasing, importance for companies, spreadsheets have drastic economical and societal impact. Hence, locating and fixing spreadsheet faults is important and deserves attention from the research community. A state-of-the-art technique uses genetic programming for generating repair candidates, but a limitation that hinders real-world application is that it still computes too many repair candidates. In this paper, we discuss a novel technique based on constraint solving that uses distinguishing test cases to narrow down the number of repair candidates. Birgit Hofer, Rui Abreu 0001, Alexandre Perez, Franz Wotawa |
ECAI | 1 |
| 2014 | Comparing Models for Spreadsheet Fault LocalizationabstractLocating faults in spreadsheets can be difficult. Therefore, tools supporting the localization of faults are needed. This paper presents a novel dependency-based model that can be used in Model-based software debugging (MBSD). This model allows improvements of the diagnostic accuracy while keeping the computation times short. In an empirical evaluation, we show that dependency-based models of spreadsheets whose value-based models are often not solvable in an acceptable amount of time can be solved in less than one second. Furthermore, the amount of diagnoses is reduced by 15 % on average when using the novel instead of the original dependency-based model. Birgit Hofer, Franz Wotawa |
ECAI | 1 |
| 2014 | Why Does my Spreadsheet Compute Wrong Values?abstractSpreadsheets are by far the most used programs that are written by end-users. They often build the basis for decisions in companies and governmental organizations and therefore they have a high impact on our daily life. Ensuring correctness of spreadsheets is thus an important task. But what happens after detecting a faulty behavior? This question has not been sufficiently answered. Therefore, we focus on fault localization techniques for spreadsheets. In this paper, we introduce a novel dependency-based approach for model-based fault localization in spreadsheets. This approach improves diagnostic accuracy while keeping computation times short, thus making the automated fault localization more appropriate for practical applications. The presented approach allows for an acceptable fault localization time of less than a second, and reduces the number of computed root cause candidates by 15 % on average, when compared with another dependency-based approach. Birgit Hofer, Franz Wotawa |
ISSRE | 1 |
| 2014 | Avoiding, finding and fixing spreadsheet errors - A survey of automated approaches for spreadsheet QA
Dietmar Jannach, Thomas Schmitz 0002, Birgit Hofer, Franz Wotawa |
J. Syst. Softw. | 3 |
| 2013 | On the Empirical Evaluation of Fault Localization Techniques for Spreadsheets
Birgit Hofer, André Riboira, Franz Wotawa, Rui Abreu 0001, Elisabeth Getzner |
FASE | 1 |