VLDB 2026 Research / reviewers in the wild / expert
Marcus Kessel
dblp:133/2140
· DBLP profile ↗
9ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0003-3088-2166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can We Classify Flaky Tests Using Only Test Code? an LLM-Based Empirical Study
Alexander Berndt 0002, Vekil Bekmyradov, Rainer Gemulla, Marcus Kessel, Thomas Bach 0001, Sebastian Baltes |
SANER | 4 |
| 2026 | Towards Observation Lakehouses: Living, Interactive Archives of Software BehaviorabstractAccompanying dataset for "Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior" (accepted @ SANER Tool Demo Track 2026). Preprint is available on arxiv: https://arxiv.org/abs/2512.02795 The observation lakehouse project and its code is hosted on GitHub: https://github.com/SoftwareObservatorium/observation-lakehouse Marcus Kessel |
SANER | 1 |
| 2025 | Morescient GAI for Software EngineeringabstractThe ability of Generative AI (GAI) technology to automatically check, synthesize, and modify software engineering artifacts promises to revolutionize all aspects of software engineering. Using GAI for software engineering tasks is consequently one of the most rapidly expanding fields of software engineering research, with over a hundred LLM-based code models having been published since 2021. However, the overwhelming majority of existing code models share a major weakness—they are exclusively trained on the syntactic facet of software, significantly lowering their trustworthiness in tasks dependent on software semantics. To address this problem, a new class of “Morescient” GAI is needed that is “aware” of (i.e., trained on) both the semantic and static facets of software. This, in turn, will require a new generation of software observation platforms capable of generating large quantities of execution observations in a structured and readily analyzable way. In this article, we present a vision and roadmap for how such “Morescient” GAI models can be engineered, evolved, and disseminated according to the principles of open science. Marcus Kessel, Colin Atkinson 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Promoting open science in test-driven software experimentsabstractA core principle of open science is the clear, concise and accessible publication of empirical data, including “raw” observational data as well as processed results. However, in empirical software engineering there are no established standards (de jure or de facto) for representing and “opening” observations collected in test-driven software experiments — that is, experiments involving the execution of software subjects in controlled scenarios. Execution data is therefore usually represented in ad hoc ways, often making it abstruse and difficult to access without significant manual effort. In this paper we present new data structures designed to address this problem by clearly defining, correlating and representing the stimuli and responses used to execute software subjects in test-driven experiments. To demonstrate their utility, we show how they can be used to promote the repetition, replication and reproduction of experimental evaluations of AI-based code completion tools. We also show how the proposed data structures facilitate the incremental expansion of execution data sets, and thus promote their repurposing for new experiments addressing new research questions. Marcus Kessel, Colin Atkinson 0001 |
J. Syst. Softw. | 1 |
| 2024 | Code search engines for the next generationabstractGiven the abundance of software in open source repositories, code search engines are increasingly turning to “big data” technologies such as natural language processing and machine learning, to deliver more useful search results. However, like the syntax-based approaches traditionally used to analyze and compare code in the first generation of code search engines, big data technologies are essentially static analysis processes. When dynamic properties of software, such as run-time behavior (i.e., semantics) and performance, are among the search criteria, the exclusive use of static algorithms has a significant negative impact on the precision and recall of the search results as well as other key usability factors such as ranking quality. Therefore, to address these weaknesses and provide a more reliable and usable service, the next generation of code search engines needs to complement static code analysis techniques with equally large-scale, dynamic analysis techniques based on its execution and observation. In this paper we describe a new software platform specifically developed to achieve this by simplifying and largely automating the dynamic analysis (i.e., observation) of code at a large scale. We show how this platform can combine dynamically observed properties of code modules with static properties to improve the quality and usability of code search results. Marcus Kessel, Colin Atkinson 0001 |
J. Syst. Softw. | 1 |
| 2022 | Diversity-driven unit test generation
Marcus Kessel, Colin Atkinson 0001 |
J. Syst. Softw. | 1 |
| 2019 | Automatically Curated Data Setsabstracto validate hypotheses and tools that depend on the semantics of software, it is necessary to assemble, prepare and maintain (i.e. curate) large, high-quality corpora of executable software systems exhibiting certain desired behavior and/or properties. Today this is a highly tedious and laborious activity requiring significant human time and effort. In this paper we therefore present a prototype platform that supports the notion of “live data sets” where almost all aspects of the data set curation process are automated. Instead of curating data sets by hand, or writing dedicated tools to select and check software samples on a case-by-case basis, a live data set allows users to simply describe their requirements as abstract scripts written in a declarative domain specific language. After explaining the approach and the key ideas behind its implementation, in this paper we present two examples of executable corpora generated automatically from a live data set populated from Maven Central. The first illustrates a “semantics agnostic” use case where the actual behavior of the software is unimportant, while the second illustrates a “semantics specific” use case where software implementing a specific functional abstraction is selected. Marcus Kessel, Colin Atkinson 0001 |
SCAM | 1 |
| 2019 | On the Efficacy of Dynamic Behavior Comparison for Judging Functional EquivalenceabstractSince it was first proposed in 1992 under the name of "behavior sampling", the idea of judging whether software systems are functionally equivalent by observing their responses to common stimuli (i.e. tests) has been used for a range of tasks such as software retrieval, functional redundancy measurement and semantic clone detection. However, its efficacy has only been studied in one small experiment, with limited generalizability, described in the original paper proposing the approach. The results of that experiment suggest that a relatively small number of randomly generated tests (i.e. 4) is sufficient to recognize non-functional-equivalent software 85% of the time. This number has therefore been adopted as "sufficient" in numerous applications of the approach. In this paper we present a much larger study which suggests at least 39 randomly generated tests are actually needed to achieve this level of effectiveness, but that a far fewer number of tests generated using coverage-based heuristics are sufficient. Since these results are much more generalizable, they have implications for future applications of behavioral sampling for dynamic behavior comparison. Marcus Kessel, Colin Atkinson 0001 |
SCAM | 1 |
| 2013 | On the Synergy between Search-Based and Search-Driven Software Engineering
Colin Atkinson 0001, Marcus Kessel, Marcus Schumacher |
SSBSE | 2 |