VLDB 2026 Research / reviewers in the wild / expert
Katharina Großer
dblp:315/4056 · also Katharina Naujokat
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-4532-0270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Too Many Issues: Automatically Prioritizing Analyzer Findings by Tracing Security ImportanceabstractCode-based analyzers often find too many potentially security-related issues to address them all. Therefore, issues likely to lead to vulnerabilities should be fixed first. Such prioritization requires project-specific knowledge, such as quality requirements, security-related decisions, and design, which is not accessible to code analyzers. We present TraceSEC, an automated technique for prioritizing issues according to their security-related importance to the project. Its core concept is to incorporate available design artifacts and trace links between them, thus considering the project context that the code lacks. We reduce the problem of issue prioritization to a maximum flow problem and quantify the importance of each issue by the flow from user-defined quality aspects to the issue, i.e., quantifying its impact on project-specific security preferences. Our evaluation shows that TraceSEC effectively provides automated prioritization and can be tailored to project-specific quality goals. Its prioritization correlates stronger with manual expert prioritization than SonarQube rule severities, which are commonly used in practice. In particular, TraceSEC has a higher similarity for identifying high-priority issues. TraceSEC scales reasonably well for codebases up to four million lines of code, and the initial setup overhead is likely to be recouped after the first automated prioritization. Sven Peldszus, Katharina Großer, Marco Konersmann, Wasja Brunotte, Maike Ahrens, Kurt Schneider, Jan Jürjens |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | From missile warhead to smart fridge: Interviews with industry experts on tracing safety- and security-relevant artifactsabstractEnsuring traceability of safety- and security-related artifacts is vital in software development to comply with standards and mitigate risks. Despite its importance, the practical implementation of defining and tracing safety- and security-relevant artifacts remains ambiguous. Based on eight semi-structured interviews with industry experts, this work explores the definitions, methods, processes, and challenges of tracing safety- and security-related artifacts. The interviews revealed that definitions of safety- and security-relevant artifacts are highly context-dependent, shaped by regulatory standards, internal processes, technical characteristics, and practitioner judgment. Rather than signaling a deficiency, this variability reflects the inherently multifaceted nature of safety and security work, where artifact classification emerges from practical reasoning rather than strict or universal criteria. Tools play a key role in supporting traceability, and cross-team alignment remains a concern in practice. Our findings provide actionable insights for organizations seeking to strengthen traceability. The recommendations encourage the development of internal classification criteria, support effective collaboration with external partners, support guidance, onboarding, and training, and help align practices with across teams, fostering more reliable and transparent management of safety- and security-relevant artifacts. Marc Herrmann, Alexander Specht, Abdurrahman Sekerci, Martin Obaidi, Marco Ehl, Duaa Adel Ali Elsofi, Katharina Großer, Jil Klünder, Jan Jürjens, Kurt Schneider |
J. Syst. Softw. | 7 |
| 2024 | Benchmarking requirement template systems: comparing appropriateness, usability, and expressivenessabstractAbstract Various semi-formal syntax templates for natural language requirements foster to reduce ambiguity while preserving human readability. Existing studies on their effectiveness focus on individual notations only and do not allow to systematically investigate quality benefits. We strive for a comparative benchmark and evaluation of template systems to assist practitioners in selecting appropriate ones and enable researchers to work on pinpoint improvements and domain-specific adaptions. We conduct comparative experiments with five popular template systems—EARS, Adv-EARS, Boilerplates, MASTeR , and SPIDER. First, we compare a control group of free-text requirements and treatment groups of their variants following the different templates. Second, we compare MASTeR and EARS in user experiments for reading and writing. Third, we analyse all five meta-models’ formality and ontological expressiveness based on the Bunge-Wand-Weber reference ontology. The comparison of the requirement phrasings across seven relevant quality characteristics and a dataset of 1764 requirements indicates that, except SPIDER, all template systems have positive effects on all characteristics. In a user experiment with 43 participants, mostly students, we learned that templates are a method that requires substantial prior training and that profound domain knowledge and experience is necessary to understand and write requirements in general. The evaluation of templates systems’ meta-models suggests different levels of formality, modularity, and expressiveness. MASTeR and Boilerplates provide high numbers of variants to express requirements and achieve the best results with respect to completeness. Templates can generally improve various quality factors compared to free text. Although MASTeR leads the field, there is no conclusive favourite choice, as most effect sizes are relatively similar. Katharina Großer, Amir Shayan Ahmadian, Marina Rukavitsyna, Qusai Ramadan, Jan Jürjens |
Requir. Eng. | 1 |
| 2023 | A Comparative Evaluation of Requirement Template SystemsabstractContext: Multiple semi-formal syntax templates for natural language requirements foster to reduce ambiguity while preserving readability. Yet, existing studies on their effectiveness do not allow to systematically investigate quality benefits and compare different notations. Objectives: We strive for a comparative benchmark and evaluation of template systems to support practitioners in selecting template systems and enable researchers to work on pinpoint improvements and domain-specific adaptions. Methods: We conduct a comparative experiment with a control group of free-text requirements and treatment groups of their variants following different templates. We compare effects on metrics systematically derived from quality guidelines. Results: We present a benchmark consisting of a systematically derived metric suite over seven relevant quality categories and a dataset of 1764 requirements, comprising 249 free-text forms from five projects and variants in five template systems. We evaluate effects in comparison to free text. Except for one template system, all have solely positive effects in all categories. Conclusions: The proposed benchmark enables the identification of the relative strengths and weaknesses of different template systems. Results show that templates can generally improve quality compared to free text. Although MASTER leads the field, there is no conclusive favourite choice, as overall effect sizes are relatively similar. Katharina Großer, Marina Rukavitsyna, Jan Jürjens |
RE | 1 |
| 2023 | ILLOD Replication Package: An Open-Source Framework for Abbreviation-Expansion Pair Detection and Term Consolidation in RequirementsabstractILLOD is a tool for detecting abbreviation-expansion pairs (AEPs) in requirement sets. It utilizes syntactic features such as Initial Letters, term Lengths, Order, and Distribution of characters to determine if a term is a potential long form to a given abbreviation. The artifact bundles all source code and data resources to replicate evaluation results presented for ILLOD in two research papers published at the REFSQ2022 Conference and in the Information and Software Technology (IST) journal. In addition, ILLOD can be used to detect AEPs, perform abbreviation detection, and the input data-set can be used for further research in requirements engineering or other related fields. The repository is organized into different directories containing data, Python sources, and notebooks for experiments and evaluations. Detailed instructions are provided to load and use the tool on a local system, and the results generated by ILLOD are stored in output files. The tool demonstrates its effectiveness in detecting AEPs and consolidating glossary terms, and the evaluation results provide insights into the performance of different classifiers. The artifact repository is a valuable resource for researchers and practitioners in the field of requirements engineering and related areas. Hussein Hasso, Katharina Großer, Iliass Aymaz, Hanna Geppert, Jan Jürjens |
RE | 2 |
| 2023 | Enhanced abbreviation-expansion pair detection for glossary term extraction
Hussein Hasso, Katharina Großer, Iliass Aymaz, Hanna Geppert, Jan Jürjens |
Inf. Softw. Technol. | 2 |
| 2022 | Evaluation Methods and Replicability of Software Architecture Research ObjectsabstractContext: Software architecture (SA) as research area experienced an increase in empirical research, as identified by Galster and Weyns in 2016 [1]. Empirical research builds a sound foundation for the validity and comparability of the research. A current overview on the evaluation and replicability of SA research objects could help to discuss our empirical standards as a community. However, no such current overview exists.Objective: We aim at assessing the current state of practice of evaluating SA research objects and replication artifact provision in full technical conference papers from 2017 to 2021.Method: We first create a categorization of papers regarding their evaluation and provision of replication artifacts. In a systematic literature review (SLR) with 153 papers we then investigate how SA research objects are evaluated and how artifacts are made available.Results: We found that technical experiments (28%) and case studies (29%) are the most frequently used evaluation methods over all research objects. Functional suitability (46% of evaluated properties) and performance (29%) are the most evaluated properties. 17 papers (11%) provide replication packages and 97 papers (63%) explicitly state threats to validity. 17% of papers reference guidelines for evaluations and 14% of papers reference guidelines for threats to validity.Conclusions: Our results indicate that the generalizability and repeatability of evaluations could be improved to enhance the maturity of the field; although, there are valid reasons for contributions to not publish their data. We derive from our findings a set of four proposals for improving the state of practice in evaluating software architecture research objects. Researchers can use our results to find recommendations on relevant properties to evaluate and evaluation methods to use and to identify reusable evaluation artifacts to compare their novel ideas with other research. Reviewers can use our results to compare the evaluation and replicability of submissions with the state of the practice. Marco Konersmann, Angelika Kaplan, Thomas Kühn 0001, Robert Heinrich, Anne Koziolek, Ralf Reussner, Jan Jürjens, Mahmood al-Doori, Nicolas Boltz, Marco Ehl, Dominik Fuchß, Katharina Großer, Sebastian Hahner, Jan Keim, Matthias Lohr, Timur Saglam, Sophie Corallo, Jan-Philipp Töberg |
ICSA | 12 |
| 2022 | Abbreviation-Expansion Pair Detection for Glossary Term Extraction
Hussein Hasso, Katharina Großer, Iliass Aymaz, Hanna Geppert, Jan Jürjens |
REFSQ | 2 |
| 2022 | Requirements document relationsabstractAbstract Relations between requirements are part of nearly every requirements engineering approach. Yet, relations of views, such as requirements documents, are scarcely considered. This is remarkable as requirements documents and their structure are a key factor in requirements reuse, which is still challenging. Explicit formalized relations between documents can help to ensure consistency, improve completeness, and facilitate review activities in general. For example, this is relevant in space engineering, where many challenges related to complex document dependencies occur: 1. Several contractors contribute to a project. 2. Requirements from standards have to be applied in several projects. 3. Requirements from previous phases have to be reused. We exploit the concept of “layered traceability”, explicitly considering documents as views on sets of individual requirements and specific traceability relations on and between all of these representation layers. Different types of relations and their dependencies are investigated with a special focus on requirement reuse through standards and formalized in an Object-Role Modelling (ORM) conceptual model. Automated analyses of requirement graphs based on this model are able to reveal document inconsistencies. We show examples of such queries in Neo4J/Cypher for the EagleEye case study. This work aims to be a step toward a better support to handle highly complex requirement document dependencies in large projects with a special focus on requirements reuse and to enable automated quality checks on dependent documents to facilitate requirements reviews. Katharina Großer, Volker Riediger, Jan Jürjens |
Softw. Syst. Model. | 1 |