Thomas Hirsch

dblp:40/5068 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 5 first-author · 6 since 2021Computer networks · 3Security and privacy · 1
YearPublicationVenuePosition
2025 Best practices for evaluating IRFL approaches
abstract
Information retrieval fault localization (IRFL) is a popular research field and many IRFL approaches have been proposed recently. Unfortunately, the evaluation of some of these IRFL approaches is often too simplistic, which can cause an overestimation of performance of these approaches. In this paper, we discuss evaluation pitfalls and problems. Furthermore, we propose best practices to avoid them. In detail, we discuss evaluation strategies such as parameter tuning and temporal dependencies in the data, dataset issues, metrics, statistical significance testing, and the unavailability of supplemental material. To support our claim of the poor status quo of current evaluation practices in some research papers, we have performed a literature survey on 135 papers. We hope that this paper will help researchers to avoid the described pitfalls in their evaluation of IRFL approaches. • Discussion of common pitfalls in the evaluation of IRFL approaches. • Seven best practices that help to avoid these pitfalls. • Investigation of 135 IRFL papers w.r.t. these pitfalls and best practices.
Thomas Hirsch, Birgit Hofer
J. Syst. Softw.1
2024 One-Class Classification and Cluster Ensembles for Anomaly Detection and Diagnosis in Multivariate Time Series Data
Adil Mukhtar, Thomas Hirsch, Gerald Schweiger
DX2
2023 The MAP Metric in Information Retrieval Fault Localization
abstract
The MAP (Mean Average Precision) metric is one of the most popular performance metrics in the field of Information Retrieval Fault Localization (IRFL). However, there are problematic implementations of this MAP metric used in IRFL research. These implementations deviate from the text book definitions of MAP, rendering the metric sensitive to the truncation of retrieval results and inaccuracies and impurities of the used datasets. The application of such a deviating metric can lead to performance overestimation. This can pose a problem for comparability, transferability, and validity of IRFL performance results. In this paper, we discuss the definition and mathematical properties of MAP and common deviations and pitfalls in its implementation. We investigate and discuss the conditions enabling such overestimation: the truncation of retrieval results in combination with ground truths spanning multiple files and improper handling of undefined AP results. We demonstrate the overestimation effects using the Bench4BL benchmark and five well known IRFL techniques. Our results indicate that a flawed implementation of the MAP metric can lead to an overestimation of the IRFL performance, in extreme cases by up to 70 %. We argue for a strict adherence to the text book version of MAP with the extension of undefined AP values to be set to 0 for all IRFL experiments. We hope that this work will help to improve comparability and transferability in IRFL research.
Thomas Hirsch, Birgit Hofer
ASE1
2022 Pruning Boolean Expressions to Shorten Dynamic Slices
abstract
This paper presents a novel extension to dynamic slicing that we call pruned slicing. The proposed slicing approach produces smaller slices than traditional dynamic slicing. This is achieved by reasoning over Boolean expressions. We have implemented a prototype in Python and empirically evaluated its performance on three different benchmarks: TCAS, QuixBugs and the Refactory dataset. We show that pruned slicing reduces the size of dynamic slices on average by 10.96 percent for TCAS. For QuixBugs and the Refactory dataset, the slice size remains the same, but the number of Boolean expressions within the slice is reduced. Further, the empirical evaluation shows that pruned dynamic slicing comes with a low computational overhead compared to dynamic slicing. Pruned slicing can also be used in combination with relevant slicing.
Thomas Hirsch, Birgit Hofer
SCAM1
2022 Detecting non-natural language artifacts for de-noising bug reports
abstract
Textual documents produced in the software engineering process are a popular target for natural language processing (NLP) and information retrieval (IR) approaches. However, issue tickets often contain artifacts such as code snippets, log outputs and stack traces. These artifacts not only inflate the issue ticket sizes, but also can this noise constitute a real problem for some NLP approaches, and therefore has to be removed in the pre-processing of some approaches. In this paper, we present a machine learning based approach to classify textual content into natural language and non-natural language artifacts at line level. We show how data from GitHub issue trackers can be used for automated training set generation, and present a custom preprocessing approach for the task of artifact removal. The training sets are automatically created from Markdown annotated issue tickets and project documentation files. We use these generated training sets to train a Markdown agnostic model that is able to classify un-annotated content. We evaluate our approach on issue tickets from projects written in C++, Java, JavaScript, PHP, and Python. Our approach achieves ROC-AUC scores between 0.92 and 0.96 for language-specific models. A multi-language model trained on the issue tickets of all languages achieves ROC-AUC scores between 0.92 and 0.95. The provided models are intended to be used as noise reduction pre-processing steps for NLP and IR approaches working on issue tickets.
Thomas Hirsch, Birgit Hofer
Autom. Softw. Eng.1
2022 A systematic literature review on benchmarks for evaluating debugging approaches
abstract
Bug benchmarks are used in development and evaluation of debugging approaches, e.g. fault localization and automated repair. Quantitative performance comparison of different debugging approaches is only possible when they have been evaluated on the same dataset or benchmark. However, benchmarks are often specialized towards usage for certain debugging approaches in their contained data, metrics, and artifacts. Such benchmarks cannot be easily used on debugging approaches outside their scope as such approach may rely on specific data such as bug reports or code metrics that are not included in the dataset. Furthermore, benchmarks vary in their size w.r.t. the number of subject programs and the size of the individual subject programs. For these reasons, we have performed a systematic literature review where we have identified 73 benchmarks that can be used to evaluate debugging approaches. We compare the different benchmarks w.r.t. their size and the provided information such as bug reports, contained test cases, and other code metrics. This comparison is intended to help researchers to quickly identify all suitable benchmarks for evaluating their specific debugging approaches. Furthermore, we discuss reoccurring issues and challenges in selection, acquisition, and usage of such bug benchmarks, i.e., data availability, data quality, duplicated content, data formats, reproducibility, and extensibility. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board.
Thomas Hirsch, Birgit Hofer
J. Syst. Softw.1
2019 Results from running an experiment as a service platform for mobile broadband networks in Europe
Vincenzo Mancuso, Miguel Peón-Quirós, Cise Midoglu, Mohamed Moulay, Vincenzo Comite, Andra Lutu, Özgü Alay, Stefan Alfredsson, Mohammad Rajiullah, Anna Brunström, Marco Mellia, Ali Safari Khatouni, Thomas Hirsch
Comput. Commun.13
2017 Experience: An Open Platform for Experimentation with Commercial Mobile Broadband Networks
abstract
Open experimentation with operational Mobile Broadband (MBB) networks in the wild is currently a fundamental requirement of the research community in its endeavor to address the need of innovative solutions for mobile communications. Even more, there is a strong need for objective data about stability and performance of MBB (e.g., 3G/4G) networks, and for tools that rigorously and scientifically assess their status. In this paper, we introduce the MONROE measurement platform: an open access and flexible hardware-based platform for measurements and custom experimentation on operational MBB networks. The MONROE platform enables accurate, realistic and meaningful assessment of the performance and reliability of 11 MBB networks in Europe. We report on our experience designing, implementing and testing the solution we propose for the platform. We detail the challenges we overcame while building and testing the MONROE testbed and argue our design and implementation choices accordingly. We describe and exemplify the capabilities of the platform and the wide variety of experiments that external users already perform using the system.
Özgü Alay, Andra Lutu, Miguel Peón-Quirós, Vincenzo Mancuso, Thomas Hirsch, Kristian Evensen, Audun Fosselie Hansen, Stefan Alfredsson, Jonas Karlsson 0001, Anna Brunström, Ali Safari Khatouni, Marco Mellia, Marco Ajmone Marsan
MobiCom5
2016 Measuring and assessing mobile broadband networks with MONROE
abstract
Mobile broadband (MBB) networks underpin numerous vital operations of the society and are arguably becoming the most important piece of the communications infrastructure. In this demo paper, our goal is to showcase the potential of a novel multi-homed MBB platform for measuring, monitoring and assessing the performance of MBB services in an objective manner. Our platform, MONROE, is composed of hundreds of nodes scattered over four European countries and a backend system that collects the measurement results. Through a user-friendly web client, the experimenters can schedule and deploy their experiments. The platform further embeds traffic analysis tools for real-time traffic flow analysis and a powerful visualization tool.
Özgü Alay, Andra Lutu, Rafael García, Miguel Peón-Quirós, Vincenzo Mancuso, Thomas Hirsch, Tobias Dely, Jonas Werme, Kristian Evensen, Audun Fosselie Hansen, Stefan Alfredsson, Jonas Karlsson 0001, Anna Brunström, Ali Safari Khatouni, Marco Mellia, Marco Ajmone Marsan, Roberto Monno, Håkon Lønsethagen
WoWMoM6
2008 Packet Sampling for Flow Accounting: Challenges and Limitations
Tanja Zseby, Thomas Hirsch, Benoit Claise
PAM2