EDBT 2026 Demo / reviewers in the wild / expert
Marcus Gerhold
dblp:161/9832
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-2655-9617ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward Automated UML Diagram Assessment: Comparing LLM-Generated Scores with Teaching AssistantsabstractThis paper investigates the feasibility of using Large Language Models (LLMs) to automate the grading of Unified Modeling Language (UML) class diagrams in a software design course. Our method involves carefully designing case studies with constraints that guide students’ design choices, converting visual diagrams to textual descriptions, and leveraging LLMs’ natural language processing capabilities to evaluate submissions. We evaluated our approach using 92 student submissions, comparing grades assigned by three teaching assistants with those generated by three LLMs (Llama, GPT o1-mini, and Claude). Our results show that GPT o1-mini and Claude Sonnet achieved strong alignment with human graders, reaching correlation coefficients above 0.76 and Mean Absolute Errors below 4 points on a 40-point scale. The findings suggest that LLM-based grading can provide consistent, scalable assessment of UML diagrams while matching the grading quality of human assessors. This approach offers a promising solution for managing growing student numbers while ensuring fair and timely feedback. Nacir Bouali, Marcus Gerhold, Tosif Ul Rehman |
CSEDU (1) | 2 |
| 2025 | Affective Mirroring in Video Game NPCs: A Pilot Study Evaluating Player EngagementabstractAdvancements in Artificial Intelligence (AI) technologies and Affective Computing have created new possibilities in video game design, particularly in enhancing the interactivity and emotional depth of non-player characters (NPCs). This study explores the potential of AI-driven affective mirroring to enhance player attachment to non-player characters (NPCs) in video games by dynamically recognizing and responding to player emotions. Using a visual novel dating game prototype, the research employs facial expression analysis to recognize and match player emotions with NPC reactions, aiming to create emotionally responsive interactions. Ten participants, aged 21 to 24, including both gamers and non-gamers, were involved in the testing. Results show that while affective mirroring can strengthen player attachment, subtle and context-aware responses are key to maintaining immersion and believability. Ethical concerns are discussed in the context of privacy issues related to the collection and use of emotional data, with recommendations for transparent data practices and safeguards against emotional manipulation. The study suggests that, with further development, affective mirroring could improve player engagement in games and expand to other fields that involve virtual agents such as healthcare, education, or customer service. Hana Sinkovic, Marcello A. Gómez Maureira, Marcus Gerhold |
FDG | 3 |
| 2025 | Time for Quiescence: Modelling Quiescent Behaviour in Testing via Time-Outs in Timed Automata
Laura Brandán Briones, Marcus Gerhold, Petra van den Bos, Mariëlle Stoelinga |
ICTSS | 2 |
| 2025 | Conformance in the railway industry: Single-Input-Change testing a EULYNX controllerabstractAbstract We propose a novel framework for model-based testing against specifications from EULYNX, a SysML-based standard from the railway industry for the controllers of systems such as points, signals, sensors, and crossings. The main challenge here is the sheer complexity: with state spaces exceeding $10^{10}$ 10 10 states, it is hard to derive test suites that achieve a meaningful type of coverage. We tackle this problem by moving away from the traditional interleaving semantics for SysML. Instead, we propose a synchronous semantics in terms of Finite State Machines (FSMs), leveraging the fact that EULYNX is implemented on Programmable Logic Controllers (PLCs). Then, we deploy Single-Input-Change Deterministic Finite State Machines (SIC-DFSMs), which ensures fully deterministic tests, thus minimizing scalability issues. Our focus lies on the EULYNX specification for point controllers. The generated test suite achieves maximal transition coverage, but test execution time remains substantial. We introduce an additional test suite that achieves maximal transition label coverage. Remarkably, this smaller suite successfully identifies the same four faults as the larger suite. Djurre van der Wal, Marcus Gerhold, Mariëlle Stoelinga, Arend Rensink |
Int. J. Softw. Tools Technol. Transf. | 2 |
| 2024 | Teaching Assistants as Assessors: An Experience Based NarrativeabstractThis study explores the role of teaching assistants (TAs) as assessors in a university’s computer science program. It examines the challenges and implications of TAs in grading, with a focus on their expertise and grading consistency. The paper analyzes grading experiences in various exam settings and investigates the impact on assessment quality. We adopt an empirical methodology and answer the research question by analyzing the data from two exams. The chosen exams have similar learning objectives but they differ in how TAs graded them, thus providing an opportunity to reflect on different grading styles. It concludes with recommendations for enhancing TA grading effectiveness, emphasizing the need for detailed rubrics, training, and monitoring to ensure fair and reliable assessment in higher education. Nacir Bouali, Marcus Gerhold |
CSEDU (1) | 3 |
| 2024 | The Limits of the Identifiable: Challenges in Python Version Identification with Deep LearningabstractThe evolution of Python requires accurate version identification to facilitate compatibility and ongoing support. We extend previous work on deep learning models for Python version identification, where LSTM and CodeBERT achieved a 92% accuracy on short code snippets. We further expand these results to larger realistic files, utilising code segmentation techniques for varying input granularities. These techniques ranged from per-line analysis to larger code segments. Our findings show that while LSTM with CodeBERT embeddings maintained high accuracy on short snippets, performance significantly drops on longer segments, particularly in balancing information retention and misclassification risks. Notably, import-statement analysis, despite being the most intuitive indicator of version requirements, reached only a 30% accuracy. This exposes the limitations of our approach when encountering rare or user-defined modules. The findings expose the limitations of deep learning for language version identification, and suggest that alternative approaches may be necessary for high accuracy on larger datasets. Marcus Gerhold, Lola Solovyeva, Vadim Zaytsev |
SANER | 1 |
| 2024 | Deriving modernity signatures of codebases with static analysisabstractThis paper addresses the problem of determining the modernity of software systems by analysing the use of new language features and their adoption over time. We propose the concept of modernity signatures to estimate the age of a codebase, naturally adjusted for maintenance practices, such that the modernity of a regularly updated system would be above that of a more recently created one which neglects current features and best practices. This can provide insights into coding practices, codebase health and the evolution of software languages. We present case studies on PHP and Python code, demonstrating the effectiveness of modernity signatures in determining the age of a codebase without executing the code or performing extensive human inspection. The paper describes the technical implementation details of generating the modernity signature for both of these languages, including the use of existing tools like the PHP parser and Vermin. The findings suggest that modernity signatures can aid developers in many ways from choosing whether to use a system or how to approach its maintenance, to assessing usefulness of a language feature, thus providing a valuable tool for source code analysis and manipulation. Chris Admiraal, Wouter Van den Brink, Marcus Gerhold, Vadim Zaytsev, Cristian Zubcu |
J. Syst. Softw. | 3 |
| 2023 | Computer Aided Content Generation - A Gloomhaven Case StudyabstractWe present how an evolutionary algorithm can be used to generate scenarios for the board game Gloomhaven. The scenarios are evaluated according to size, difficulty, thematic coherence, complexity and layout. We encode the game’s default scenarios into textual descriptions and use them as initial population for the algorithm. Our dungeon generation works within the confines given by the physical board game, i.e., special attention is given to availability of game pieces and map tiles. The generated dungeons can be constructed without overlapping tiles. Marcus Gerhold, Kristian Tijben |
FDG | 1 |
| 2023 | Conformance in the Railway Industry: Single-Input-Change Testing a EULYNX Controller
Djurre van der Wal, Marcus Gerhold, Mariëlle Stoelinga |
FMICS | 2 |
| 2022 | Deriving Modernity Signatures for PHP Systems with Static AnalysisabstractThe PHP language has undergone many changes in its syntax and grammar, with respect to both features the language has to offer as well as the distribution of language features used by programmers in their projects. We present a novel method of using grammar usage statistics to calculate a modernity signature for a PHP system, so that we can determine its age. The system will aid developers in choosing whether or not to execute or use a PHP system, without having to perform an extensive inspection. Wouter Van den Brink, Marcus Gerhold, Vadim Zaytsev |
SCAM | 2 |
| 2018 | A Hierarchy of Scheduler Classes for Stochastic AutomataabstractStochastic automata are a formal compositional model for concurrent stochastic timed systems, with general distributions and nondeterministic choices. Measures of interest are defined over schedulers that resolve the nondeterminism. In this paper we investigate the power of various theoretically and practically motivated classes of schedulers, considering the classic complete-information view and a restriction to non-prophetic schedulers. We prove a hierarchy of scheduler classes w.r.t. unbounded probabilistic reachability. We find that, unlike Markovian formalisms, stochastic automata distinguish most classes even in this basic setting. Verification and strategy synthesis methods thus face a tradeoff between powerful and efficient classes. Using lightweight scheduler sampling, we explore this tradeoff and demonstrate the concept of a useful approximative verification technique for stochastic automata. Pedro R. D'Argenio, Marcus Gerhold, Arnd Hartmanns, Sean Sedwards |
FoSSaCS | 2 |
| 2018 | Model-based testing of probabilistic systemsabstractAbstract This work presents an executable model-based testing framework for probabilistic systems with non-determinism. We provide algorithms to automatically generate, execute and evaluate test cases from a probabilistic requirements specification. The framework connects input/output conformance-theory with hypothesis testing: our algorithms handle functional correctness, while statistical methods assess, if the frequencies observed during the test process correspond to the probabilities specified in the requirements. At the core of our work lies the conformance relation for probabilistic input/output conformance, enabling us to pin down exactly when an implementation should pass a test case. We establish the correctness of our framework alongside this relation as soundness and completeness; Soundness states that a correct implementation indeed passes a test suite, while completeness states that the framework is powerful enough to discover each deviation from a specification up to arbitrary precision for a sufficiently large sample size. The underlying models are probabilistic automata that allow invisible internal progress. We incorporate divergent systems into our framework by phrasing four rules that each well-formed system needs to adhere to. This enables us to treat divergence as the absence of output, or quiescence, which is a well-studied formalism in model-based testing. Lastly, we illustrate the application of our framework on three case studies. Marcus Gerhold, Mariëlle Stoelinga |
Formal Aspects Comput. | 1 |
| 2016 | Model-Based Testing of Probabilistic Systems
Marcus Gerhold, Mariëlle Stoelinga |
FASE | 1 |