VLDB 2026 Research / reviewers in the wild / expert
Alexander Berndt 0002
dblp:276/6208-2
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0009-5248-6405ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can We Classify Flaky Tests Using Only Test Code? an LLM-Based Empirical Study
Alexander Berndt 0002, Vekil Bekmyradov, Rainer Gemulla, Marcus Kessel, Thomas Bach 0001, Sebastian Baltes |
SANER | 1 |
| 2026 | Using Large Language Models to Support Automation of Failure Management in CI/CD Pipelines: A Case Study in SAP HANAabstractCI/CD pipeline failure management is time-consuming when performed manually. Automating this process is non-trivial because the information required for effective failure management is unstructured and cannot be automatically processed by traditional programs. With their ability to process unstructured data, large language models (LLMs) have shown promising results for automated failure management by previous work. Following these studies, we evaluated whether an LLM-based system could automate failure management in a CI/CD pipeline in the context of a large industrial software project, namely SAP HANA. We evaluated the ability of the LLM-based system to identify the error location and to propose exact solutions that contain no unnecessary actions. To support the LLM in generating exact solutions, we provided it with different types of domain knowledge, including pipeline information, failure management instructions, and data from historical failures. We conducted an ablation study to determine which type of domain knowledge contributed most to solution accuracy. The results show that data from historical failures contributed the most to the system's accuracy, enabling it to produce exact solutions in 92.1% of cases in our dataset. The system correctly identified the error location with 97.4% accuracy when provided with domain knowledge, compared to 84.2% accuracy without it. In conclusion, our findings indicate that LLMs, when provided with data from historical failures, represent a promising approach for automating CI/CD pipeline failure management. Duong Bui, Stefan Grintz, Alexander Berndt 0002, Thomas Bach 0001 |
SANER | 3 |
| 2024 | Do Test and Environmental Complexity Increase Flakiness? An Empirical Study of SAP HANAabstractBackground: Test flakiness is a major problem in the software industry. Flaky tests fail seemingly at random without changes to the code and thus impede continuous integration (CI). Some researchers argue that all tests can be considered flaky and that tests only differ in their frequency of flaky failures. This position implies that the definition of test flakiness includes failures caused by interruptions in the testing environment. Alexander Berndt 0002, Thomas Bach 0001, Sebastian Baltes |
ESEM | 1 |
| 2023 | The Vocabulary of Flaky Tests in the Context of SAP HANAabstractBackground. Automated test execution is an important activity to gather information about the quality of a software project. So-called flaky tests, however, negatively affect this process. Such tests fail seemingly at random without changes to the code and thus do not provide a clear signal. Previous work proposed to identify flaky tests based on the source code identifiers in the test code. So far, these approaches have not been evaluated in a large-scale industrial setting. Aims. We evaluate approaches to identify flaky tests and their root causes based on source code identifiers in the test code in a large-scale industrial project. Method. First, we replicate previous work by Pinto et al. in the context of SAP HANA. Second, we assess different feature extraction techniques, namely TF-IDF and TF-IDFC-RF. Third, we evaluate CodeBERT and XGBoost as classification models. For a sound comparison, we utilize both the data set from previous work and two data sets from SAP HANA. Results. Our replication shows similar results on the original data set and on one of the SAP HANA data sets. While the original approach yielded an F1-Score of 0.94 on the original data set and 0.92 on the SAP HANA data set, our extensions achieve F1-Scores of 0.96 and 0.99, respectively. The reliance on external data sources is a common root cause for test flakiness in the context of SAP HANA. Conclusions. The vocabulary of a large industrial project seems to be slightly different with respect to the exact terms, but the categories for the terms, such as remote dependencies, are similar to previous empirical findings. However, even with rather large F1-Scores, both finding source code identifiers for flakiness and a black box prediction have limited use in practice as the results are not actionable for developers. Alexander Berndt 0002, Zoltán Nochta, Thomas Bach 0001 |
ESEM | 1 |