Daniel Seybold

dblp:177/4838 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0002-7973-5485ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 A Methodology and Framework to Determine the Isolation Capabilities of Virtualisation Technologies
abstract
The capability to isolate system resources is an essential characteristic of virtualisation technologies and is therefore important for research and industry alike. It allows the co-location of experiments and workloads, the partitioning of system resources and enables multi-tenant business models such as cloud computing. Poor isolation among tenants bears the risk of noisy-neighbour and contention effects which negatively impacts all of those use-cases. These effects describe the negative impact of one tenant onto another by utilising shared resources. Both industry and research provide many different concepts and technologies to realise isolation. Yet, the isolation capabilities of all these different approaches are not well understood; nor is there an established way to measure the quality of their isolation capabilities. Such an understanding, however, is of uttermost importance in practice to elaborately decide on a suited implementation. Hence, in this work, we present a novel methodology to measure the isolation capabilities of virtualisation technologies for system resources, that fulfils all requirements to benchmarking including reliability. It relies on an immutable approach, based on Experiment-as-Code. The complete process holistically includes everything from bare metal resource provisioning to the actual experiment enactment.
Simon Volpert, Benjamin Erb, Georg Eisenhart, Daniel Seybold, Stefan Wesner, Jörg Domaschka
ICPE4
2022 Same, Same, but Dissimilar: Exploring Measurements for Workload Time-series Similarity
abstract
Benchmarking is a core element in the toolbox of most systems researchers and is used for analyzing, comparing, and validating complex systems. In the quest for reliable benchmark results, a consensus has formed that a significant experiment must be based on multiple runs. To interpret these runs, mean and standard deviation are often used. In case of experiments where each run produces a time series, applying and comparing the mean is not easily applicable and not necessarily statistically sound. Such an approach ignores the possibility of significant differences between runs with a similar average. In order to verify this hypothesis, we conducted a survey of 1,112 publications of selected performance engineering and systems conferences canvassing open data sets from performance experiments. The identified 3 data sets purely rely on average and standard deviation. Therefore, we propose a novel analysis approach based on similarity analysis to enhance the reliability of performance evaluations. Our approach evaluates 12 (dis-)similarity measures with respect to their applicability in analysing performance measurements and identifies four suitable similarity measures. We validate our approach by demonstrating the increase in reliability for the data sets found in the survey.
Mark Leznik, Johannes Grohmann, Nina Kliche, André Bauer 0001, Daniel Seybold, Simon Eismann, Samuel Kounev, Jörg Domaschka
ICPE5
2020 Baloo: Measuring and Modeling the Performance Configurations of Distributed DBMS
abstract
Correctly configuring a distributed database management system (DBMS) deployed in a cloud environment for maximizing performance poses many challenges to operators. Even if the entire configuration spectrum could be measured directly, which is often infeasible due to the multitude of parameters, single measurements are subject to random variations and need to be repeated multiple times. In this work, we propose Baloo, a framework for systematically measuring and modeling different performance-relevant configurations of distributed DBMS in cloud environments. Baloo dynamically estimates the required number of measurement configurations, as well as the number of required measurement repetitions per configuration based on a desired target accuracy. We evaluate Baloo based on a data set consisting of 900 DBMS configuration measurements conducted in our private cloud setup. Our evaluation shows that the highly configurable framework is able to achieve a prediction error of up to 12 %, while saving over 80 % of the measurement effort. We also publish all code and the acquired data set to foster future research.
Johannes Grohmann, Daniel Seybold, Simon Eismann, Mark Leznik, Samuel Kounev, Jörg Domaschka
MASCOTS2
2019 Kaa: Evaluating Elasticity of Cloud-Hosted DBMS
abstract
Auto-scaling is able to change the scale of an application at runtime. Understanding the application characteristics, scaling impact as well as the workload, an auto-scaler aligns the acquired resources to match the current workload. For distributed Database Management Systems (DBMS) forming the backend of many large-scale cloud applications, it is currently an open question to what extent they support scaling at run-time. In particular, elasticity properties of existing distributed DBMS are widely unknown and difficult to evaluate and compare. This paper presents a comprehensive methodology for the evaluation of the elasticity of distributed DBMS. On the basis of this methodology, we introduce a framework that automates the full evaluation process. We validate the framework by defining significant elasticity scenarios for a case study that comprises two DBMS for write-heavy and read-heavy workloads of different intensities. The results show that scalable distributed DBMS are not necessarily elastic and that adding more instances to a cluster at run-time may even decrease the experienced performance.
Daniel Seybold, Simon Volpert, Stefan Wesner, André Bauer 0001, Nikolas Herbst, Jörg Domaschka
CloudCom1
2019 Mowgli: Finding Your Way in the DBMS Jungle
abstract
Big Data and IoT applications require highly-scalable database management system (DBMS), preferably operated in the cloud to ensure scalability also on the resource level. As the number of existing distributed DBMS is extensive, the selection and operation of a distributed DBMS in the cloud is a challenging task. While DBMS benchmarking is a supportive approach, existing frameworks do not cope with the runtime constraints of distributed DBMS and the volatility of cloud environments. Hence, DBMS evaluation frameworks need to consider DBMS runtime and cloud resource constraints to enable portable and reproducible results. In this paper we present Mowgli, a novel evaluation framework that enables the evaluation of non-functional DBMS features in correlation with DBMS runtime and cloud resource constraints. Mowgli fully automates the execution of cloud and DBMS agnostic evaluation scenarios, including DBMS cluster adaptations. The evaluation of Mowgli is based on two IoT-driven scenarios, comprising the DBMSs Apache Cassandra and Couchbase, nine DBMS runtime configurations, two cloud providers with two different storage backends. Mowgli automates the execution of the resulting 102 evaluation scenarios, verifying its support for portable and reproducible DBMS evaluations. The results provide extensive insights into the DBMS scalability and the impact of different cloud resources. The significance of the results is validated by the correlation with existing DBMS evaluation results.
Daniel Seybold, Moritz Keppler, Daniel Gründler, Jörg Domaschka
ICPE1
2018 A Provider-Agnostic Approach to Multi-cloud Orchestration Using a Constraint Language
abstract
Cloud computing and its computing as an utility paradigm provides on-demand resources allowing the seamless adaptation of applications to fluctuating demands. While the Cloud's ongoing commercialisation has lead to a vast provider landscape, vendor lock-in is still a major hindrance. Recent outages demonstrate that relying exclusively on one provider is not sufficient. While existing cloud orchestration tools promise to solve the problems by supporting deployments across multiple cloud providers, they typically rely on provider dependent models forcing prior knowledge of offers and obstructing flexibility in case of errors. We propose a cloud provider-agnostic application and resource description using a constraint language. It allows users to express resource requirements of an application without prior knowledge of existing offers. Additionally, we propose a discovery service automatically collecting available offers. We combine this with a matchmaking algorithm representing the discovery model and the user-given constraints in a constraint satisfaction problem (CSP) that is then solved. Finally, we manipulate this discovery model during runtime to react on errors. Our evaluation shows that using a constraint-based language is a feasible approach to the provider selection problem, and that it helps to overcome vendor lock-in.
Daniel Baur, Daniel Seybold, Frank Griesinger, Hynek Masata, Jörg Domaschka
CCGrid2
2017 A Cloud-driven View on Business Process as a Service
Jörg Domaschka, Frank Griesinger, Daniel Seybold, Stefan Wesner
CLOSER3
2016 Is elasticity of scalable databases a Myth?
abstract
The age of cloud computing has introduced all the mechanisms needed to elastically scale distributed, cloud-enabled applications. At roughly the same time, NoSQL databases have been proclaimed as the scalable alternative to relational databases. Since then, NoSQL databases are a core component of many large-scale distributed applications. This paper evaluates the scalability and elasticity features of the three widely used NoSQL database systems Couchbase, Cassandra and MongoDB under various workloads and settings using throughput and latency as metrics. The numbers show that the three database systems have dramatically different baselines with respect to both metrics and also behave unexpected when scaling out. For instance, while Couchbase's throughput increases by 17% when scaled out from 1 to 4 nodes, MongoDB's throughput decreases by more than 50%. These surprising results show that not all tested NoSQL databases do scale as expected and even worse, in some cases scaling harms performances.
Daniel Seybold, Benjamin Erb, Jörg Domaschka
IEEE BigData1
2016 Experiences of models@run-time with EMF and CDO
Daniel Seybold, Jörg Domaschka, Alessandro Rossini, Christopher B. Hauser, Frank Griesinger, Athanasios Tsitsipas
SLE1