VLDB 2026 Research / reviewers in the wild / expert
Jan Arne Sparka
dblp:347/2826
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0002-5886-4595ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Validity constraints for data analysis workflowsabstractPorting a scientific data analysis workflow (DAW) to a cluster infrastructure, a new software stack, or even only a new dataset with some notably different properties is often challenging. Despite the structured definition of the steps (tasks) and their interdependencies during a complex data analysis in the DAW specification, relevant assumptions may remain unspecified and implicit. Such hidden assumptions often lead to crashing tasks without a reasonable error message, poor performance in general, non-terminating executions, or silent wrong results of the DAW, to name only a few possible consequences. Searching for the causes of such errors and drawbacks in a distributed compute cluster managed by a complex infrastructure stack, where DAWs for large datasets typically are executed, can be tedious and time-consuming. We propose validity constraints (VCs) as a new concept for DAW languages to alleviate this situation. A VC is a constraint specifying logical conditions that must be fulfilled at certain times for DAW executions to be valid. When defined together with a DAW, VCs help to improve the portability, adaptability, and reusability of DAWs by making implicit assumptions explicit. Once specified, VCs can be controlled automatically by the DAW infrastructure, and violations can lead to meaningful error messages and graceful behaviour (e.g., termination or invocation of repair mechanisms). We provide a broad list of possible VCs, classify them along multiple dimensions, and compare them to similar concepts one can find in related fields. We also provide a proof-of-concept implementation for the workflow system Nextflow. Florian Schintke, Khalid Belhajjame, Ninon De Mecquenem, David Frantz, Vanessa Emanuela Guarino, Marcus Hilbrich, Fabian Lehmann, Paolo Missier, Rebecca Sattler, Jan Arne Sparka, Daniel T. Speckhard, Hermann Stolte, Duc Anh Vu 0001, Ulf Leser |
Future Gener. Comput. Syst. | 10 |
| 2024 | Grammar-based fuzzing of data integration parsers in computational materials scienceabstractAbstract Context Computational materials science (CMS) focuses on in silico experiments to compute the properties of known and novel materials, where many software packages are used in the community. The NOMAD Laboratory (Draxl C, Scheffler) offers to store the input and output files in its FAIR data repository. Since the file formats of these software packages are non‐standardized, parsers are used to provide the results in a normalized format. Objective The main goal of this article is to report experience and findings of using grammar‐based fuzzing on these parsers. Method We have constructed an input grammar for four common software packages in the CMS domain and performed an experimental evaluation on the capabilities of grammar‐based fuzzing to detect failures in the Novel Materials Discovery (NOMAD) parsers. Results With our approach, we were able to identify three unique critical bugs concerning service availability, as well as several additional syntactic, semantic, logical, and downstream bugs in the investigated NOMAD parsers. We reported all issues to the developer team prior to publication. Conclusion Based on the experience gained, we can recommend grammar‐based fuzzing also for other research software packages to improve the trust level in the correctness of the produced results. Sebastian Müller 0007, Jan Arne Sparka, Martin Kuban, Claudia Ambrosch-Draxl, Lars Grunske |
Softw. Pract. Exp. | 2 |
| 2023 | Design by Contract Revisited in the Context of Scientific Data Analysis WorkflowsabstractSoftware systems enabling large-scale data analysis workflows (DAWs) are a key technology for modern science as they allow extracting new insights from experimental results. DAWs are pipelines composed of interdependent tasks that are executed in a distributed fashion on large compute clusters. Typically, the individual task implementations are developed by research groups all over the world and usually not tested outside a narrow scope of possible inputs, parameters, and infrastructures. As a result, the operations' correctness depends on many implicit assumptions, such as the completeness and suitability of input data, infrastructure properties such as available cores, etc. This makes quality assurance of DAWs a critical issue. We propose to address this problem by introducing a contract-driven approach to DAW design and implementation. Following the well-known principle of Design by Contract, DAW developers specify contracts in the form of requirements and promises for each task of a DAW. These contracts serve as guards to ensure that tasks run in a proper environment and produce correct results. The detection of contract violations allows to halt the execution of a DAW and to identify the culprit that caused the violation. Thus, the integration of contracts into DAW design provides opportunities for efficiency improvements by reducing computation and debugging time in case of errors. Duc Anh Vu 0001, Jan Arne Sparka, Ninon De Mecquenem, Timo Kehrer, Ulf Leser, Lars Grunske |
e-Science | 2 |
| 2023 | Contract-Driven Design of Scientific Data Analysis WorkflowsabstractSoftware systems enabling large-scale data analysis workflows (DAWs) are a key technology for many scientific disciplines, as they allow extracting new insights from experimental results. DAWs are (non-)linear pipelines composed of multiple interdependent tasks that are executed in a distributed fashion on large compute clusters. In science, the individual task implementations are developed by research groups all over the world and usually not tested outside a narrow scope of possible inputs, parameters, and infrastructures. As a result, the operations' correctness depends on many implicit assumptions. Among others this includes the completeness and suitability of input data, infrastructure properties such as available cores or main memory, etc. This combination of complexity, distribution and untested components makes quality assurance of DAWs a critical issue. In this paper, we propose to address this problem by introducing a contract-driven approach to DAW design and implementation. Following this method, DAW developers specify contracts in the form of requirements and promises for each task of a DAW. These contracts serve as guards to ensure that tasks run in a proper environment and produce correct results. We provide the first formal definition of contracts for DAWs and show how they are connected to DAW scheduling and execution. As a proof of concept, we extended Nextflow, a popular scientific workflow system, with contracts and defined a light-weight DSL for their specification. We exemplify the power of a contract-driven approach to DAW development by enhancing several real-world DAWs from Bioinformatics to capture typical problems during their execution and show how the specific notifications issued by broken contracts help debugging the DAWs. Duc Anh Vu 0001, Jan Arne Sparka, Ninon De Mecquenem, Timo Kehrer, Ulf Leser, Lars Grunske |
e-Science | 2 |