VLDB 2026 Research / reviewers in the wild / expert
Johannes Härtel
dblp:150/3820
· DBLP profile ↗
14ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-7461-2320ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The sampling threat when mining generalizable inter-library usage patterns
Yunior Pacheco Correa, Coen De Roover, Johannes Härtel |
Sci. Comput. Program. | 3 |
| 2025 | An AI Security Testbed for the 5G CoreabstractThe 5G core network is the backbone of modern mobile communication, providing high-speed, low-latency, and diverse services for users and industries. Artificial Intelligence (AI) plays an important role in this network by optimizing per-formance, supporting dynamic resource scaling, and improving security through anomaly detection and threat mitigation. Testing AI in 5G environments is difficult because of the complexity of the network and the many possible attack vectors. In this paper, we present a modular and reproducible testbed for evaluating AI-based security mechanisms in the 5G core. The testbed emulates key 5G components and traffic patterns, enabling systematic experiments under realistic conditions. It also provides reliable measurements of Key Performance Indicators (KPIs) to evaluate the effectiveness, robustness, and operational impact of AI solutions, including their ability to detect and mitigate threats. Our work provides a structured framework for testing AI solutions and supports the development of secure, resilient, and AI -enhanced 5G networks. Clément Legrand-Duchesne, Johannes Härtel, Fabio Massacci, Mengyuan Zhang 0001, Agathe Blaise |
CloudCom | 2 |
| 2025 | Improved Labeling of Security Defects in Code Review by Active Learning with LLMsabstractMining high-quality datasets of security defects is important for cybersecurity. In this paper, we focus on mining a dataset of reviews that discuss potential security defects in code or other artifacts. Mining such datasets often involves labeling, and this is challenging because security defects are rare. Johannes Härtel |
EASE | 1 |
| 2024 | Threats to Instrument Validity Within "in Silico" Research: Software Engineering to the RescueabstractAbstract “In Silico” research drives the world around us, as illustrated by the way our society handles climate change, controls the COVID-19 pandemic and governs economic growth. Unfortunately, the code embedded in the underlying data processing is mostly written by scientists lacking formal training in software engineering. The resulting code is vulnerable, suffering from what is known as threats to instrument validity. This position paper aims to understand and remedy threats to instrument validity in current “in silico” research. To achieve this goal, we specify a research agenda listing how recent software engineering achievements may improve “in silico” research (SE4Silico) and, conversely, how software engineering may strengthen its applicability (Silico4SE). Serge Demeyer, Coen De Roover, Mutlu Beyazit, Johannes Härtel |
ISoLA (4) | 4 |
| 2023 | Symbolic Execution to Detect Semantic Merge ConflictsabstractCollaborative software development depends on managing multiple versions of a program which requires mechanisms to merge program versions to eventually deploy a single executable. Merging program versions can be challenging as conflicts can arise. The most challenging form is a semantic conflict, which introduces unintended behaviour in the resulting executable while merging. In this paper, we develop an approach that detects such semantic merge conflicts by symbolic execution. We define the program semantics as path conditions, produced by a symbolic executor, and check whether the conditions satisfy established rules that reflect a merge conflict. Our usage of symbolic execution to check these rules is novel. We evaluate the correctness of our approach through mutation testing, and evaluate it empirically by applying the approach to real-world merges sampled from GitHub. We also discuss what challenges arise in the empirical evaluation, including the problems i) that semantic merge conflicts are rare in the wild, ii) and, even in retrospection, hard to find using standard search mechanisms. Our evaluation shows that in specific cases, our approach using symbolic execution is a promising extension to existing mechanisms to merge conflict detection. Ward Muylaert, Johannes Härtel, Coen De Roover |
SCAM | 2 |
| 2023 | Operationalizing validity of empirical software engineering studies
Johannes Härtel, Ralf Lämmel |
Empir. Softw. Eng. | 1 |
| 2022 | Operationalizing Threats to MSR Studies by Simulation-Based TestingabstractQuantitative studies on the border between Mining Software Repository (MSR) and Empirical Software Engineering (ESE) apply data analysis methods, like regression modeling, statistic tests or correlation analysis, to commits or pulls to better understand the software development process. Such studies assure the validity of the reported results by following a sound methodology. However, with increasing complexity, parts of the methodology can still go wrong. This may result in MSR/ESE studies with undetected threats to validity. In this paper, we propose to systematically protect against threats by operationalizing their treatment using simulations. A simulation substitutes observed and unobserved data, related to an MSR/ESE scenario, with synthetic data, carefully defined according to plausible assumptions on the scenario. Within a simulation, unobserved data becomes transparent, which is the key difference to a real study, necessary to detect threats to an analysis methodology. Running an analysis methodology on synthetic data may detect basic technical bugs and misinterpretations, but it also improves the trust in the methodology. The contribution of a simulation is to operationalize testing the impact of important assumptions. Assumptions still need to be rated for plausibility. We evaluate simulation-based testing by operationalizing undetected threats in the context of four published MSR/ESE studies. We recommend that future research uses such more systematic treatment of threats, as a contribution against the reproducibility crisis. Johannes Härtel, Ralf Lämmel |
MSR | 1 |
| 2020 | Incremental Map-Reduce on Repository HistoryabstractWork on Mining Software Repositories typically involves processing abstractions of resources on individual revisions. A corresponding processing of abstractions of resource changes often depends on working with all revisions of the repository history to guarantee a high resolution of the measured changes. Abstractions of resources and abstractions of resource changes are often very related up to the point that they can be used interchangeably in the processing. In practice, approaches working with abstractions processed over high revision counts face a scalability challenge. In this work, we contribute to the challenge by incrementalizing the processing of repository resources and the corresponding abstractions. Our work is inspired by incrementalization theory including insights on Abelian groups, group homomorphisms and indexing. We provide a map-reduce interface that enables calls to foreign functionality and convenient operations for processing abstractions, such as mapping, filtering, group-wise aggregation and joining. Apache Spark is used for distribution. We compare the scalability of our approach with available MSR approaches, i.e., with LISA that reduces redundancy and with DJ-Rex that migrates an analysis to a distributed map-reduce framework. Johannes Härtel, Ralf Lämmel |
SANER | 1 |
| 2020 | Understanding MDE projects: megamodels to the rescue for architecture recovery
Juri Di Rocco, Davide Di Ruscio, Johannes Härtel, Ludovico Iovino, Ralf Lämmel, Alfonso Pierantonio |
Softw. Syst. Model. | 3 |
| 2019 | Empirical study on the usage of graph query languages in open source Java projectsabstractGraph data models are interesting in various domains, in part because of the intuitiveness and flexibility they offer compared to relational models. Specialized query languages, such as Cypher for property graphs or SPARQL for RDF, facilitate their use. In this paper, we present an empirical study on the usage of graph-based query languages in open-source Java projects on GitHub. We investigate the usage of SPARQL, Cypher, Gremlin and GraphQL in terms of popularity and their development over time. We select repositories based on dependencies related to these technologies and employ various popularity and source-code based filters and ranking features for a targeted selection of projects. For the concrete languages SPARQL and Cypher, we analyze the activity of repositories over time. For SPARQL, we investigate common application domains, query use and existence of ontological data modeling in applications that query for concrete instance data. Our results show, that the usage of graph query languages in open-source projects increased over the last years, with SPARQL and Cypher being by far the most popular. SPARQL projects are more active in terms of query related artifact changes and unique developers involved, but Cypher is catching up. Relatively few applications use SPARQL to query for concrete instance data: A majority of those applications employ multiple different ontologies, including project and domain specific ones. Common application domains are management systems and data visualization tools. Philipp Seifer, Johannes Härtel, Martin Leinberger, Ralf Lämmel, Steffen Staab |
SLE | 2 |
| 2018 | EMF Patterns of Usage on GitHub
Johannes Härtel, Marcel Heinz, Ralf Lämmel |
ECMFA | 1 |
| 2018 | Classification of APIs by hierarchical clusteringabstractAPIs can be classified according to the programming domains (e.g., GUIs, databases, collections, or security) that they address. Such classification is vital in searching repositories (e.g., the Maven Central Repository for Java) and for understanding the technology stack used in software projects. We apply hierarchical clustering to a curated suite of Java APIs to compare the computed API clusters with preexisting API classifications. Clustering entails various parameters (e.g., the choice of IDF versus LSI versus LDA). We describe the corresponding variability in terms of a feature model. We exercise all possible configurations to determine the maximum correlation with respect to two baselines: i) a smaller suite of APIs manually classified in previous research; ii) a larger suite of APIs from the Maven Central Repository, thereby taking advantage of crowd-sourced classification while relying on a threshold-based approach for identifying important APIs and versions thereof, subject to an API dependency analysis on GitHub. We discuss the configurations found in this way and we examine the influence of particular features on the correlation between computed clusters and baselines. To this end, we also leverage interactive exploration of the parameter space and the resulting dendrograms. In this manner, we can also identify issues with the use of classifiers (e.g., missing classifiers) in the baselines and limitations of the clustering approach. Johannes Härtel, Hakan Aksu, Ralf Lämmel |
ICPC | 1 |
| 2017 | A chrestomathy of DSL implementationsabstractSelecting and properly using approaches for DSL implementation can be challenging, given their variety and complexity. To support developers, we present the software chrestomathy MetaLib, a well-organized and well-documented collection of DSL implementations useful for learning. We focus on basic metaprogramming techniques for implementing DSL syntax and semantics. The DSL implementations are organized and enhanced by feature modeling, semantic annotation, and model-based documentation. The chrestomathy enables side-by-side exploration of different implementation approaches for DSLs. Source code, feature model, feature configurations, semantic annotations, and documentation are publicly available online, explorable through a web application, and maintained by a collaborative process. Simon Schauss, Ralf Lämmel, Johannes Härtel, Marcel Heinz, Kevin Klein 0001, Lukas Härtel, Thorsten Berger |
SLE | 3 |
| 2014 | Test-Data Generation for Xtext - Tool Paper
Johannes Härtel, Lukas Härtel, Ralf Lämmel |
SLE | 1 |