VLDB 2026 Research / reviewers in the wild / expert
Jan Keim
dblp:172/4597
· DBLP profile ↗
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-8899-7081ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 4 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Architecture in the Cradle: Early Warning of Architectural Decay with ArchGuard
Dominik Fuchß, Sophie Corallo, Maximilian Hummel, Jan Keim, Tobias Hey 0001 |
ICSA | 5 |
| 2025 | LLMs for Software Architecture Knowledge: A Comparative Analysis Among Seven LLMs
Mohamed Soliman 0001, Elia Ashraf, Kamel M. K. Abdelsalam, Jan Keim, Ashwin Prasad Shivarpatna Venkatesh |
ECSA | 4 |
| 2025 | Enabling Architecture Traceability by LLM-based Architecture Component Name Extraction
Dominik Fuchß, Tobias Hey 0001, Jan Keim, Anne Koziolek |
ICSA | 4 |
| 2025 | Do Large Language Models Contain Software Architectural Knowledge? : An Exploratory Case Study with GPTabstractArchitectural knowledge (AK) of existing systems is essential for software engineers to make design decisions. Recently, Large Language Models (LLMs) trained on large-scale datasets, including software repositories, have shown promise in embedding knowledge and answering questions. However, LLMs have not been evaluated for their abilities to answer questions about AK, leaving doubts about their accuracy. This paper assesses GPT, a leading LLM, by evaluating its responses’ accuracy, quality, and trustworthiness on the AK of the large-scale open-source system HDFS. We conducted an exploratory case study with 14 software engineers who posed questions to GPT and compared its responses to a predefined ground truth. The engineers rated GPT’s answers with moderate quality and trustworthiness. Our findings on GPT’s accuracy indicates moderate recall but lower precision, especially in identifying quality attribute solutions and design rationales. These results suggest that while GPT and similar models can provide initial insights into AK, expert validation remains necessary for reliability. This study underscores LLMs’ potential and limitations to discover software AK. Mohamed Soliman 0001, Jan Keim |
ICSA | 2 |
| 2025 | LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationabstractThere are a multitude of software artifacts which need to be handled during the development and maintenance of a software system. These artifacts interrelate in multiple, complex ways. Therefore, many software engineering tasks are enabled - and even empowered - by a clear understanding of artifact interrelationships and also by the continued advancement of techniques for automated artifact linking. However, current approaches in automatic Traceability Link Recovery (TLR) target mostly the links between specific sets of artifacts, such as those between requirements and code. Fortu-nately, recent advancements in Large Language Models (LLMs) can enable TLR approaches to achieve broad applicability. Still, it is a nontrivial problem how to provide the LLMs with the specific information needed to perform TLR. In this paper, we present LiSSA, a framework that har-nesses LLM performance and enhances them through Retrieval-Augmented Generation (RAG). We empirically evaluate LiSSA on three different TLR tasks, requirements to code, documentation to code, and architecture documentation to architecture models, and we compare our approach to state-of-the-art approaches. Our results show that the RAG-based approach can signifi-cantly outperform the state-of-the-art on the code-related tasks. However, further research is required to improve the performance of RAG-based approaches to be applicable in practice. Dominik Fuchß, Tobias Hey 0001, Jan Keim, Niklas Ewald, Tobias Thirolf, Anne Koziolek |
ICSE | 3 |
| 2025 | Requirements Traceability Link Recovery via Retrieval-Augmented Generation
Tobias Hey 0001, Dominik Fuchß, Jan Keim, Anne Koziolek |
REFSQ | 3 |
| 2024 | Recovering Trace Links Between Software Documentation And CodeabstractIntroduction Software development involves creating various artifacts at different levels of abstraction and establishing relationships between them is essential. Traceability link recovery (TLR) automates this process, enhancing software quality by aiding tasks like maintenance and evolution. However, automating TLR is challenging due to semantic gaps resulting from different levels of abstraction. While automated TLR approaches exist for requirements and code, architecture documentation lacks tailored solutions, hindering the preservation of architecture knowledge and design decisions. Methods This paper presents our approach TransArC for TLR between architecture documentation and code, using component-based architecture models as intermediate artifacts to bridge the semantic gap. We create transitive trace links by combining the existing approach ArDoCo for linking architecture documentation to models with our novel approach ArCoTL for linking architecture models to code. Jan Keim, Sophie Corallo, Dominik Fuchß, Tobias Hey 0001, Tobias Telge, Anne Koziolek |
ICSE | 1 |
| 2024 | Requirements Classification for Traceability Link RecoveryabstractBeing aware of and understanding the relations between the requirements of a software system to its other artifacts is crucial for their successful development, maintenance and evolution. There are approaches to automatically recover this traceability information, but they fail to identify the actual relevant parts of the requirements. Recent large language model-based requirements classification approaches have shown to be able to identify aspects and concerns of requirements with promising accuracy. Therefore, we investigate the potential of those classification approaches for identifying irrelevant requirement parts for traceability link recovery between requirements and code. We train the large language model-based requirements classification approach NoRBERT on a new dataset of requirements and their entailed aspects and concerns. We use the results of the classification to filter irrelevant parts of the requirements before recovering trace links with the fine-grained word embedding-based FTLR approach. Two empirical studies show promising results regarding the quality of classification and the impact on traceability link recov-ery. NoRBERT can identify functional and user-related aspects in the requirements with an F I-score of 84 %. With the classification and requirements filtering, the performance of FTLR could be improved significantly and FTLR performs better than state-of-the-art unsupervised traceability link recovery approaches. Tobias Hey 0001, Jan Keim, Sophie Corallo |
RE | 2 |
| 2023 | Automated Reverse Engineering of the Technology-Induced Software System Structure
Yves Richard Kirschner, Jan Keim, Nico Peter, Anne Koziolek |
ECSA | 2 |
| 2023 | Detecting Inconsistencies in Software Architecture Documentation Using Traceability Link RecoveryabstractDocumenting software architecture is important for a system’s success. Software architecture documentation (SAD) makes information about the system available and eases comprehensibility. There are different forms of SADs like natural language texts and formal models with different benefits and different purposes. However, there can be inconsistent information in different SADs for the same system. Inconsistent documentation then can cause flaws in development and maintenance. To tackle this, we present an approach for inconsistency detection in natural language SAD and formal architecture models. We make use of traceability link recovery (TLR) and extend an existing approach. We utilize the results from TLR to detect unmentioned (i.e., model elements without natural language documentation) and missing model elements (i.e., described but not modeled elements). In our evaluation, we measure how the adaptations on TLR affected its performance. Moreover, we evaluate the inconsistency detection. We use a benchmark with multiple open source projects and compare the results with existing and baseline approaches. For TLR, we achieve an excellent F1-score of 0.81, significantly outperforming the other approaches by at least 0.24. Our approach also achieves excellent results (accuracy: 0.93) for detecting unmentioned model elements and good results for detecting missing model elements (accuracy: 0.75). These results also significantly outperform competing baselines. Although we see room for improvements, the results show that detecting inconsistencies using TLR is promising. Jan Keim, Sophie Corallo, Dominik Fuchß, Anne Koziolek |
ICSA | 1 |
| 2022 | Introducing an Evaluation Method for TaxonomiesabstractBackground: Taxonomies are crucial for the development of a research field, as they play a major role in structuring a complex body of knowledge and help to classify processes, approaches, and solutions. While there is an increasing interest in taxonomies in the software engineering (SE) research field, we observe that SE taxonomies are rarely evaluated. Aim: To raise awareness and provide operational guidance on how to evaluate a taxonomy, this paper presents a three step evaluation method evaluating its structure, applicability, and purpose. Method: To show the feasibility and applicability of our approach, we provide a running example and additionally illustrate our approach to a practical case study in SE research. Results and Conclusion: Our method with operational guidance enables SE researchers to systematically evaluate and improve the quality of their taxonomies and support reviewers to systematically assess a taxonomy’s quality. Angelika Kaplan, Thomas Kühn 0001, Sebastian Hahner, Niko Benkler, Jan Keim, Dominik Fuchß, Sophie Corallo, Robert Heinrich |
EASE | 5 |
| 2022 | Evaluation Methods and Replicability of Software Architecture Research ObjectsabstractContext: Software architecture (SA) as research area experienced an increase in empirical research, as identified by Galster and Weyns in 2016 [1]. Empirical research builds a sound foundation for the validity and comparability of the research. A current overview on the evaluation and replicability of SA research objects could help to discuss our empirical standards as a community. However, no such current overview exists.Objective: We aim at assessing the current state of practice of evaluating SA research objects and replication artifact provision in full technical conference papers from 2017 to 2021.Method: We first create a categorization of papers regarding their evaluation and provision of replication artifacts. In a systematic literature review (SLR) with 153 papers we then investigate how SA research objects are evaluated and how artifacts are made available.Results: We found that technical experiments (28%) and case studies (29%) are the most frequently used evaluation methods over all research objects. Functional suitability (46% of evaluated properties) and performance (29%) are the most evaluated properties. 17 papers (11%) provide replication packages and 97 papers (63%) explicitly state threats to validity. 17% of papers reference guidelines for evaluations and 14% of papers reference guidelines for threats to validity.Conclusions: Our results indicate that the generalizability and repeatability of evaluations could be improved to enhance the maturity of the field; although, there are valid reasons for contributions to not publish their data. We derive from our findings a set of four proposals for improving the state of practice in evaluating software architecture research objects. Researchers can use our results to find recommendations on relevant properties to evaluate and evaluation methods to use and to identify reusable evaluation artifacts to compare their novel ideas with other research. Reviewers can use our results to compare the evaluation and replicability of submissions with the state of the practice. Marco Konersmann, Angelika Kaplan, Thomas Kühn 0001, Robert Heinrich, Anne Koziolek, Ralf Reussner, Jan Jürjens, Mahmood al-Doori, Nicolas Boltz, Marco Ehl, Dominik Fuchß, Katharina Großer, Sebastian Hahner, Jan Keim, Matthias Lohr, Timur Saglam, Sophie Corallo, Jan-Philipp Töberg |
ICSA | 14 |
| 2021 | Towards an Automated Classification Approach for Software Engineering ResearchabstractThe rapid growth of software engineering research publications forces an amount of scholarly knowledge that needs to be managed, organized and communicated in digital libraries and scientific search engines. Thus, there is a need for classified papers to accomplish these tasks, but the classification process is cumbersome. Moreover, in case of new schemas, one would need to reclassify previously published research. We propose to automate the classification and present different possible techniques for doing so: Using natural language models, a rule-based approach, or an approach based on topic-labeling. In this proposal paper, we initially implemented a prototype for text classification of software engineering research papers. Angelika Kaplan, Jan Keim |
EASE | 2 |
| 2021 | Trace Link Recovery for Software Architecture Documentation
Jan Keim, Sophie Corallo, Dominik Fuchß, Claudius Kocher, Janek Speit, Anne Koziolek |
ECSA | 1 |
| 2020 | Does BERT Understand Code? - An Exploratory Study on the Detection of Architectural Tactics in Code
Jan Keim, Angelika Kaplan, Anne Koziolek, Mehdi Mirakhorli |
ECSA | 1 |
| 2020 | NoRBERT: Transfer Learning for Requirements ClassificationabstractClassifying requirements is crucial for automatically handling natural language requirements. The performance of existing automatic classification approaches diminishes when applied to unseen projects because requirements usually vary in wording and style. The main problem is poor generalization. We propose NoRBERT that fine-tunes BERT, a language model that has proven useful for transfer learning. We apply our approach to different tasks in the domain of requirements classification. We achieve similar or better results F1-scores of up to 94%) on both seen and unseen projects for classifying functional and non-functional requirements on the PROMISE NFR dataset. NoRBERT outperforms recent approaches at classifying non-functional requirements subclasses. The most frequent classes are classified with an average F1-score of 87%. In an unseen project setup on a relabeled PROMISE NFR dataset, our approach achieves an improvement of 15 percentage points in average F1score compared to recent approaches. Additionally, we propose to classify functional requirements according to the included concerns, i.e., function, data, and behavior. We labeled the functional requirements in the PROMISE NFR dataset and applied our approach. NoRBERT achieves an F1-score of up to 92%. Overall, NoRBERT improves requirements classification and can be applied to unseen projects with convincing results. Tobias Hey 0001, Jan Keim, Anne Koziolek, Walter F. Tichy |
RE | 2 |