VLDB 2026 Research / reviewers in the wild / expert
Stefan Klikovits
dblp:167/2898
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2027
0000-0003-4212-7029ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 4 first-author · 12 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Does road diversity really matter in testing automated driving systems?abstractAbstract Context The use of automated driving systems (ADSs) in the real world requires rigorous testing to ensure safety. To increase trust, ADSs should be tested on a large set of diverse road scenarios. Literature suggests that if a vehicle is driven along a set of geometrically diverse roads—measured using various diversity measures (DMs)—it will react in a wide range of behaviours, thereby increasing the chances of observing failures, or strengthening the confidence in its safety, if no failures are observed. However, this assumption has never been tested before, nor have road DMs been assessed for their properties. Objective Our goal was to perform an exploratory study on 53 currently used and new, potentially promising road DMs. Specifically, our research questions looked into the road DMs themselves, to analyse their properties (e.g. monotonicity , computation efficiency ), and to test correlation between DMs. Furthermore, we investigated the use of road DMs to determine whether the assumption that diverse test suites of roads expose diverse driving behaviour holds. Method Our empirical analysis relies on a state-of-the-art, open-source ADS testing infrastructure and uses a data set containing over 97,000 individual road geometries and matching simulation data that were collected using two driving agents. By considering test suites of various sizes and measuring their roads’ geometric diversity, we studied road DM properties, the correlation between road DMs, and the correlation between road DMs and the observed behaviour. Results Our findings reveal a strong correlation between road diversity and behavioural diversity, confirming that geometrically diverse test suites systematically exercise diverse driving behaviours. We identified and aggregations as most effective, with achieving the strongest correlation of 0.95 while requiring minimal computation time. The analysed measures maintain robust correlation with behavioural diversity across test suites containing roads of varying lengths, eliminating the need for length normalisation. Conclusions These results empirically validate the fundamental assumption underlying diversity-driven ADS testing: road geometry diversity serves as a reliable proxy for behavioural diversity. For practitioners, we recommend or as optimal choices, whilst -based measures should be avoided entirely. The near-identical correlation patterns observed across architecturally different driving agents indicate that our findings generalise beyond specific ADS implementations, providing a solid foundation for diversity-driven test generation and selection. Stefan Klikovits, Vincenzo Riccio, Ezequiel Castellano, Ahmet Cetinkaya, Alessio Gambi, Paolo Arcaini |
Empir. Softw. Eng. | 1 |
| 2026 | AugmentAble: Teaming Web Augmentation and LLMs for Instant Accessibility Improvements
Sarah Mayrhofer, Stefan Klikovits, Manuel Wimmer |
ICWE | 2 |
| 2025 | ICST Tool Competition 2025 - Self-Driving Car Testing TrackabstractThis is the first edition of the tool competition on testing self-driving cars (SDCs) at the International Conference on Software Testing, Verification and Validation (ICST). The aim is to provide a platform for software testers to submit their tools addressing the test selection problem for simulation-based testing of SDCs, which is considered an emerging and vital domain. The competition provides an advanced software platform and representative case studies to ease participants' entry into SDC regression testing, enabling them to develop their initial test generation tools for SDCS. In this first edition, the competition includes five tools from different authors. All tools were evaluated using (regression) metrics for test selection as well as compared with a baseline approache. This paper provides an overview of the competition, detailing its context, framework, participating tools, evaluation methodology, and key findings. Christian Birchler, Stefan Klikovits, Mattia Fazzini, Sebastiano Panichella |
ICST | 2 |
| 2025 | Inclusive Model-Driven Engineering for Accessible SoftwareabstractWhile model-driven engineering (MDE) claims to be a good development approach for cross-cutting concerns, this raises the question of why not every application created with MDE is accessible. Moreover, why are the MDE development processes and the tools themselves not more inclusive? In this vision paper, we sketch an ideal picture of how inclusivity in MDE - considered throughout tools, methods, artifacts, and processes - would intrinsically lead to more accessible software systems. We review the state-of-the-art, discuss current challenges, and present a possible future of inclusive MDE. Dominik Bork, Stefan Klikovits, Judith Michael, Lukas Netz, Bernhard Rumpe |
MODELS | 2 |
| 2025 | Towards LLM-enhanced Conflict Detection and Resolution in Model VersioningabstractIn the past two decades, a range of model versioning workflows have been proposed. Standard workflows are based on three-way model merging, which allows reasoning on potentially conflicting changes in concurrently developed model versions. However, the considered conflicts that can be detected are mostly targeting the syntactic level of models, such as update/update or delete/usage conflicts. In contrast, unintended semantic inconsistencies often remain unnoticed as detection mechanisms lack the semantic awareness of the modeling language or modeled domain. The resolution of such conflicts remains a manual task. In this paper, we explore how Large Language Models (LLMs) can augment model versioning workflows by supporting conflict detection and resolution. In particular, we present an LLM-enhanced solution for detecting conflicts in the three-way model merging setting. Drawing on a collection of conflict types from prior literature, we demonstrate how an LLM assistant can 1) pinpoint conflicting changes and 2) provide resolution options with clear rationales and explanations of their implications. Our results indicate that the LLMs' access to a broad range of domains and modeling languages can help find and resolve complex versioning conflicts. Our implementation combines the industrial tool LemonTree for analyzing models and model changes, with a GPT-4o (LLM) assistant primed with relevant context to detect and resolve conflicts. We conclude by discussing directions for future research to improve model versioning workflows using LLMs. Martin Eisenberg, Stefan Klikovits, Manuel Wimmer, Konrad Wieland |
MODELS | 2 |
| 2025 | GeQuPI: Quantum Program Improvement with Multi-Objective Genetic ProgrammingabstractProcessing quantum information poses novel challenges regarding the debugging of faulty quantum programs. Notably, the lack of accessible information on intermediate states during quantum processing, renders traditional debugging techniques infeasible. Moreover, even correct quantum programs might not be processable, as current quantum computers are limited in computation capacity. Thus, quantum program developers have to consider trade-offs between accuracy (i.e., probabilistically correct functionality) and computational cost of the proposed solutions. Manually finding sufficiently accurate and efficient solutions is a challenging task, even for quantum computing experts. To tackle these challenges, we propose a quantum program improvement framework for an automated generation of accurate and efficient solutions, coined Genetic Quantum Program Improver ( GeQuPI ). In particular, we focus on the tasks of debugging and optimization of quantum programs. Our framework uses techniques from quantum information theory and applies multi-objective genetic programming, which can be further hybridized with quantum-aware optimizers. To demonstrate the benefits of GeQuPI , it is applied to 47 quantum programs reused from literature and openly published libraries. The results show that our approach is capable of correcting faulty programs and optimize inefficient ones for the majority of the studied cases, showing average optimizations of 35% with respect to computational cost. • A framework for genetic improvement of quantum programs is proposed. • The framework provides debugging and optimization facilities. • It applies methods from quantum information theory and multi-objective optimization. • Its possible configurations include, a.o., varying levels of hardware closeness. • The evaluation on 47 quantum programs shows promising results. Felix Gemeinhardt, Stefan Klikovits, Manuel Wimmer |
J. Syst. Softw. | 2 |
| 2025 | Continuous Evolution of Digital Twins using the DarTwin NotationabstractAbstract Despite best efforts, various challenges remain in the creation and maintenance processes of digital twins (DTs). One of those primary challenges is the constant, continuous and omnipresent evolution of systems, their user’s needs and their environment, demanding the adaptation of the developed DT systems. DTs are developed for a specific purpose, which generally entails the monitoring, analysis, simulation or optimisation of a specific aspect of an actual system, referred to as the actual twin (AT). As such, when the twin system changes, that is either the AT itself changes, or the scope/purpose of a DT is modified, the DTs usually evolve in close synchronicity with the AT. As DTs are software systems, the best practices or methodologies for software evolution can be leveraged. This paper tackles the challenge of maintaining a (set of) DT(s) throughout the evolution of the user’s requirements and priorities and tries to understand how this evolution takes place. In doing so, we provide two contributions: (i) we develop , a visual notation form that enables reasoning on a twin system, its purposes, properties and implementation, and (ii) we introduce a set of architectural transformations that describe the evolution of DT systems. The development of these transformations is driven and illustrated by the evolution and transformations of a family home’s DT, whose purpose is expanded, changed and re-prioritised throughout its ongoing lifecycle. Additionally, we evaluate the transformations on a laboratory-scale gantry crane’s DT. Joost Mertens, Stefan Klikovits, Francis Bordeleau, Joachim Denil, Øystein Haugen |
Softw. Syst. Model. | 2 |
| 2023 | Frenetic-lib: An extensible framework for search-based generation of road structures for ADS testing
Stefan Klikovits, Ezequiel Castellano, Ahmet Cetinkaya, Paolo Arcaini |
Sci. Comput. Program. | 1 |
| 2023 | Parameter Coverage for Testing of Autonomous Driving Systems under UncertaintyabstractAutonomous Driving Systems (ADSs) are promising, but must show they are secure and trustworthy before adoption. Simulation-based testing is a widely adopted approach, where the ADS is run in a simulated environment over specific scenarios. Coverage criteria specify what needs to be covered to consider the ADS sufficiently tested. However, existing criteria do not guarantee to exercise the different decisions that the ADS can make, which is essential to assess its correctness. ADSs usually compute their decisions using parameterised rule-based systems and cost functions, such as cost components or decision thresholds. In this article, we argue that the parameters characterise the decision process, as their values affect the ADS’s final decisions. Therefore, we propose parameter coverage, a criterion requiring to cover the ADS’s parameters. A scenario covers a parameter if changing its value leads to different simulation results, meaning it is relevant for the driving decisions made in the scenario. Since ADS simulators are slightly uncertain, we employ statistical methods to assess multiple simulation runs for execution difference and coverage. Experiments using the Autonomoose ADS show that the criterion discriminates between different scenarios and that the cost of computing coverage can be managed with suitable heuristics. Thomas Laurent 0003, Stefan Klikovits, Paolo Arcaini, Fuyuki Ishikawa, Anthony Ventresque |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2022 | Dynamic Shielding for Reinforcement Learning in Black-Box Environments
Masaki Waga, Ezequiel Castellano, Sasinee Pruekprasert, Stefan Klikovits, Toru Takisaka, Ichiro Hasuo |
ATVA | 4 |
| 2021 | Handling Noise in Search-Based Scenario Generation for Autonomous Driving SystemsabstractThis paper presents the first evaluation of k-nearest neighbours-Averaging (kNN-Avg) on a real-world case study. kNN-Avg is a novel technique that tackles the challenges of noisy multi-objective optimisation (MOO). Existing studies suggest the use of repetition to overcome noise. In contrast, kNN-Avg approximates these repetitions and exploits previous executions, thereby avoiding the cost of re-running. We use kNN-Avg for the scenario generation of a real-world autonomous driving system (ADS) and show that it is better than the noisy baseline. Furthermore, we compare it to the repetition-method and outline indicators as to which approach to choose in which situations. Stefan Klikovits, Paolo Arcaini |
PRDC | 1 |
| 2021 | Pragmatic reuse for DSML developmentabstractAbstract By bridging the semantic gap, domain-specific language (DSLs) serve an important role in the conquest to allow domain experts to model their systems themselves. In this publication we present a case study of the development of the Continuous REactive SysTems language (CREST), a DSL for hybrid systems modeling. The language focuses on the representation of continuous resource flows such as water, electricity, light or heat. Our methodology follows a very pragmatic approach, combining the syntactic and semantic principles of well-known modeling means such as hybrid automata, data-flow languages and architecture description languages into a coherent language. The borrowed aspects have been carefully combined and formalised in a well-defined operational semantics. The DSL provides two concrete syntaxes: CREST diagrams, a graphical language that is easily understandable and serves as a model basis, and , an internal DSL implementation that supports rapid prototyping—both are geared towards usability and clarity. We present the DSL’s semantics, which thoroughly connect the various language concerns into an executable formalism that enables sound simulation and formal verification in , and discuss the lessons learned throughout the project. Stefan Klikovits, Didier Buchs |
Softw. Syst. Model. | 1 |
| 2018 | A Model Checker Collection for the Model Checking Contest Using Docker and Machine Learning
Didier Buchs, Stefan Klikovits, Alban Linard, Romain Mencattini, Dimi Racordon |
Petri Nets | 2 |