David Bermbach

dblp:83/10278 · DBLP profile ↗
← Back
60ranked-venue papers
11as first author
45since 2021 · last 2026
0000-0002-7524-3256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 3 first-author · 8 since 2021Systems, architecture and hardware · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Spatial Analysis on Value-Based Quadtrees of Rasterized Vector Data
Diana Baumann, Nils Japke, Tim C. Rese, David Bermbach
MDM4
2026 GeoBenchr: An Application-Centric Benchmarking Suite for Spatiotemporal Database Platforms
Tim C. Rese, Nils Japke, Diana Baumann, Natalie Carl, David Bermbach
MDM5
2026 Exploring Influence Factors on LLM Suitability for No-Code Development of End User Applications
abstract
ABSTRACT Context/Problem Statement No‐Code Development Platforms (NCDPs) empower non‐technical end users to build applications tailored to their specific demands without writing code. While NCDPs lower technical barriers, users still require some technical knowledge, for example, to structure process steps or define event‐action rules. Large Language Models (LLMs) offer a promising solution to further reduce technical requirements by supporting natural language interaction and dynamic code generation. By integrating LLMs, NCDPs can be more accessible to non‐technical users, enabling application development truly without requiring any technical expertise. Despite growing interest in LLM‐powered NCDPs, a systematic investigation into the factors influencing LLM suitability and performance remains absent. Understanding these factors is critical to effectively leveraging LLMs capabilities and maximizing their impact. Objective In this paper, we aim to investigate key factors influencing the effectiveness of LLMs in supporting end‐user application development within NCDPs. Methods We conducted comprehensive experiments evaluating four key factors, i.e., model selection, prompt language, training data background, and an error‐informed few‐shot setup, on the quality of generated applications. Specifically, we selected a range of LLMs based on architecture, scale, design focus, and training data, and evaluated them across four real‐world smart home automation scenarios implemented on a representative open‐source LLM‐powered NCDP. Results Model selection emerged as the most critical factor influencing performance. General‐purpose LLMs with strong natural language understanding generally outperformed others. Prompt language effects varied by model and task complexity: original prompts worked best for advanced multilingual LLMs, whereas translation steps improved performance for lighter or less capable models. LLMs showcased outperforming performance when their linguistic background aligned with the prompt language. In addition, incorporating an error‐informed few‐shot approach enhanced LLM performance, particularly for coding‐oriented and medium‐performing models, though its benefits were secondary to model choice and required additional engineering effort. Conclusion Our findings provide practical insights into how LLMs can be effectively integrated into NCDPs, informing both platform design and the selection of suitable LLMs for end‐user application development.
Minghe Wang, Alexandra Kapp, Trever Schirmer, Tobias Pfandzelter, David Bermbach
Softw. Pract. Exp.5
2025 Towards Serverless Processing of Spatiotemporal Big Data Queries
abstract
Spatiotemporal data are being produced in continuously growing volumes by a variety of data sources and a variety of application fields rely on rapid analysis of such data. Existing systems such as PostGIS or MobilityDB usually build on relational database systems, thus, inheriting their scale-out characteristics. As a consequence, big spatiotemporal data scenarios still have limited support even though many query types can easily be parallelized. In this paper, we propose our vision of a native serverless data processing approach for spatiotemporal data: We break down queries into small subqueries which then leverage the near-instant scaling of Function-as-a-Service platforms to execute them in parallel. With this, we partially solve the scalability needs of big spatiotemporal data processing.
Diana Baumann, Tim C. Rese, David Bermbach
IC2E3
2025 Multi-Event Triggers for Serverless Computing
Valentin Carl, Trever Schirmer, Joshua Adamek, Niklas Kowallik, Tobias Pfandzelter, Sergio Lucia, David Bermbach
IC2E7
2025 Towards an Optimized Benchmarking Platform for CI/CD Pipelines
abstract
Performance regressions in large-scale software systems can lead to substantial resource inefficiencies, making their early detection critical. Frequent benchmarking is essential for identifying these regressions and maintaining service-level agreements (SLAs). Performance benchmarks, however, are resource-intensive and time-consuming, which is a major challenge for integration into Continuous Integration / Continuous Deployment (CI/CD) pipelines. Although numerous benchmark optimization techniques have been proposed to accelerate benchmark execution, there is currently no practical system that integrates these optimizations seamlessly into real-world CI/CD pipelines. In this vision paper, we argue that the field of benchmark optimization remains under-explored in key areas that hinder its broader adoption. We identify three central challenges to enabling frequent and efficient benchmarking: (a) the composability of benchmark optimization strategies, (b) automated evaluation of benchmarking results, and (c) the usability and complexity of applying these strategies as part of CI/CD systems in practice. We also introduce a conceptual cloud-based benchmarking framework handling these challenges transparently. By presenting these open problems, we aim to stimulate research toward making performance regression detection in CI/CD systems more practical and effective.
Nils Japke, Helmut Lukasczyk, David Bermbach
IC2E4
2025 Evaluating the Impact of Spatial Features of Mobility Data and Index Choice on Database Performance
abstract
The growing number of moving Internet-of-Things (IoT) devices has led to a surge in moving object data, powering applications such as traffic routing, hotspot detection, or weather forecasting. When managing such data, spatial database systems offer various index options and data formats, e.g., point-based or trajectory-based. Likewise, dataset characteristics such as geographic overlap and skew can vary severely. All three can strongly affect database performance. While this has been studied in existing papers, none of them explore the effects and trade-offs resulting from a combination of all three aspects.In this paper, we evaluate the performance impact of index choice, data format, and dataset characteristics on a popular spatial database system, PostGIS. We focus on two aspects of dataset characteristics, the degree of overlap and the degree of skew, and propose novel metrics and approximation methods to determine these features. We design a benchmark that compares a variety of spatial indexing strategies and data formats, while also considering the impact of dataset characteristics on database performance. We include a variety of real-world and synthetic datasets, and write and read operations to cover a broad range of scenarios that might occur during application runtime.Our results offer practical guidance for developers looking to optimize spatial storage and querying, while also providing insights into dataset characteristics and their impact on database performance.
Tim C. Rese, Alexan, dra Kapp, David Bermbach
IC2E4
2025 Towards an Application-Centric Benchmark Suite for Spatiotemporal Database Systems
abstract
Spatiotemporal data play a key role for mobility-based applications and are their produced volume is growing continuously, among others, due to the increased availability of IoT devices. When working with spatiotemporal data, developers rely on spatiotemporal database systems such as PostGIS or MobilityDB. For better understanding their quality of service behavior and then choosing the best system, benchmarking is the go-to approach. Unfortunately, existing work in this field studies only small isolated aspects and a comprehensive application-centric benchmark suite is still missing.In this paper, we argue that an application-centric benchmark suite for spatiotemporal database systems is urgently needed. We identify requirements for such a benchmark suite, discuss domain-specific challenges, and sketch-out the architecture of a modular benchmarking suite.
Tim C. Rese, David Bermbach
IC2E2
2025 Towards a Testbed for Scalable FaaS Platforms
abstract
Most cloud platforms have a Function-as-a-Service (FaaS) offering that enables users to easily write highly scalable applications. To better understand how the platform’s architecture impacts its performance, we present a research-focused testbed that can be adapted to quickly evaluate the impact of different architectures and technologies on the characteristics of scalability-focused FaaS platforms.
Trever Schirmer, David Bermbach
IC2E2
2025 Minos: Exploiting Cloud Performance Variation with Function-as-a-Service Instance Selection
abstract
Serverless Function-as-a-Service (FaaS) is a popular cloud paradigm to quickly and cheaply implement complex applications. Because the function instances cloud providers start to execute user code run on shared infrastructure, their performance can vary. From a user perspective, slower instances not only take longer to complete, but also increase cost due to the pay-per-use model of FaaS services where execution duration is billed with microsecond accuracy. In this paper, we present MINOS, a system to take advantage of this performance variation by intentionally terminating instances that are slow. Fast instances are not terminated, so that they can be re-used for subsequent invocations. One use case for this are data processing and machine learning workflows, which often download files as a first step, during which MINOS can run a short benchmark. Only if the benchmark passes, the main part of the function is actually executed. Otherwise, the request is re-queued and the instance crashes itself, so that the platform has to assign the request to another (potentially faster) instance. In our experiments, this leads to a speedup of up to 13% in the resource intensive part of a data processing workflow, resulting in up to 4% faster overall performance (and consequently 4% cheaper prices). Longer and complex workflows lead to increased savings, as the pool of fast instances is re-used more often. For platforms exhibiting this behavior, users get better performance and save money by wasting more of the platforms resources.
Trever Schirmer, Valentin Carl, Nils Höller, Tobias Pfandzelter, David Bermbach
IC2E5
2025 Umbilical Choir: Automated Live Testing for Edge-to-Cloud FaaS Applications
abstract
Application users react negatively to performance regressions or availability issues across software releases. To address this, modern cloud-based applications with their multiple daily releases rely on live testing techniques such as A/B testing or canary releases. In edge-to-cloud applications, however, which have similar problems, developers currently still have to hard-code custom live testing tooling as there is no general framework for edge-to-cloud live testing. With Umbilical Choir, we partially close this gap for serverless edge-to-cloud applications. Umbilical Choir is compatible with all Function-as-a-Service platforms and (extensively) supports various live testing techniques, including canary releases with various geo-aware strategies, A/B testing, and gradual roll-outs. We evaluate Umbilical Choir through a complex release scenario showcasing various live testing techniques in a mixed edge-cloud deployments and discuss different geo-aware strategies.
Mohammadreza Malekabbasi, Tobias Pfandzelter, David Bermbach
ICFEC3
2025 µOpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
abstract
Performance regressions have a tremendous impact on the quality of software. One way to catch regressions before they reach production is executing performance tests before deployment, e.g., using microbenchmarks, which measure performance at subroutine level. In projects with many microbenchmarks, this may take several hours due to repeated execution to get accurate results, disqualifying them from frequent use in CI/CD pipelines. We propose µOpTime, a static approach to reduce the execution time of microbenchmark suites by configuring the number of repetitions for each microbenchmark. Based on the results of a full, previous microbenchmark suite run, µOpTime determines the minimal number of (measurement) repetitions with statistical stability metrics that still lead to accurate results. We evaluate µOpTime with an experimental study on 14 open-source projects written in two programming languages and five stability metrics. Our results show that (i) µOpTime reduces the total suite execution time (measurement phase) by up to 95.83% (Go) and 94.17% (Java), (ii) the choice of stability metric depends on the project and programming language, (iii) microbenchmark warmup phases have to be considered for Java projects (potentially leading to higher reductions), and (iv) µOpTime can be used to reliably detect performance regressions in CI/CD pipelines.
Nils Japke, Martin Grambow, Christoph Laaber, David Bermbach
ACM Trans. Softw. Eng. Methodol.4
2024 Komet: A Serverless Platform for Low-Earth Orbit Edge Services
abstract
Low-Earth orbit satellite networks can provide global broadband Internet access using constellations of thousands of satellites. Integrating edge computing resources in such networks can enable global low-latency access to compute services, supporting end users in rural areas, remote industrial applications, or the IoT. To achieve this, resources must be carefully allocated to various services from multiple tenants. Moreover, applications must navigate the dynamic nature of satellite networks, where orbital mechanics necessitate frequent client hand-offs. Therefore, managing applications on the low-Earth orbit edge will require the right platform abstractions.
Tobias Pfandzelter, David Bermbach
SoCC2
2024 GeoFF: Federated Serverless Workflows with Data Pre-Fetching
abstract
Function-as-a-Service (FaaS) is a popular cloud computing model in which applications are implemented as workflows of multiple independent functions. While cloud providers usually offer composition services for such workflows, they do not support cross-platform workflows forcing developers to hardcode the composition logic. Furthermore, FaaS workflows tend to be slow due to cascading cold starts, inter-function latency, and data download latency on the critical path. In this paper, we propose GEOFF, a serverless choreography middleware that executes FaaS workflows across different public and private FaaS platforms, including ad-hoc workflow recomposition. Furthermore, GEOFF supports function pre-warming and data pre-fetching. This minimizes end-to-end workflow latency by taking cold starts and data download latency off the critical path. In experiments with our proof-of-concept prototype and a realistic application, we were able to reduce end-to-end latency by more than 50%.
Valentin Carl, Trever Schirmer, Tobias Pfandzelter, David Bermbach
IC2E4
2024 GeoFaaS: An Edge-to-Cloud FaaS Platform
abstract
The massive growth of mobile and IoT devices demands geographically distributed computing systems for optimal performance, privacy, and scalability. However, existing edge-to-cloud serverless platforms lack location awareness, resulting in inefficient network usage and increased latency. In this paper, we propose GeoFaaS, a novel edge-to-cloud Function-as-a-Service (FaaS) platform that leverages real-time client location information for transparent request execution on the nearest available FaaS node. If needed, GeoFaaS transparently offloads requests to the cloud when edge resources are overloaded, thus, ensuring consistent execution without user intervention. GeoFaaS has a modular and decentralized architecture: building on the single-node FaaS system tinyFaaS, GeoFaaS works as a stand-alone edge-to-cloud FaaS platform but can also integrate and act as a routing layer for existing FaaS services, e.g., in the cloud. To evaluate our approach, we implemented an open-source proof-of-concept prototype and studied performance and fault-tolerance behavior in experiments.
Mohammadreza Malekabbasi, Tobias Pfandzelter, Trever Schirmer, David Bermbach
IC2E4
2024 Are Unikernels Ready for Serverless on the Edge?
abstract
Function-as-a-Service (FaaS) is a promising edge computing execution model but requires secure sandboxing mechanisms to isolate workloads from multiple tenants on constrained infrastructure. Although Docker containers are lightweight and popular in open-source FaaS platforms, they are generally considered insufficient for executing untrusted code and providing sandbox isolation. Commercial cloud FaaS platforms thus rely on Linux microVMs or hardened container runtimes, which are secure but come with a higher resource footprint.Unikernels combine application code and limited operating system primitives into a single purpose appliance, reducing the footprint of an application and its sandbox while providing full Linux compatibility. In this paper, we study the suitability of unikernels as an edge FaaS execution environment using the Nanos and OSv unikernel tool chains. We compare performance along several metrics such as cold start overhead and idle footprint against sandboxes such as Firecracker Linux microVMs, Docker containers, and secure gVisor containers. We find that unikernels exhibit desirable cold start performance, yet lag behind Linux microVMs in stability. Nevertheless, we show that unikernels are a promising candidate for further research on Linux-compatible FaaS isolation.
Felix Moebius, Tobias Pfandzelter, David Bermbach
IC2E3
2024 Towards Anti-Collision Coordination for UAVs with Serverless Edge Computing
Tobias Pfandzelter, David Bermbach, Robert Vilter, Ingo Friese, Sergiy Melnyk, Qiuheng Zhou, Hans D. Schotten
IC2E2
2024 Increasing Efficiency and Result Reliability of Continuous Benchmarking for FaaS Applications
abstract
In a continuous deployment setting, Function-as-aService (FaaS) applications frequently receive updated releases, each of which can cause a performance regression. While continuous benchmarking, i.e., comparing benchmark results of the updated and the previous version, can detect such regressions, performance variability of FaaS platforms necessitates thousands of function calls, thus, making continuous benchmarking time intensive and expensive. In this paper, we propose DuetFaaS, an approach which adapts duet benchmarking to FaaS applications. With DuetFaaS, we deploy two versions of FaaS function in a single cloud function instance and execute them in parallel to reduce the impact of platform variability. We evaluate our approach against state-of-the-art approaches, running on AWS Lambda. Overall, DuetFaaS requires fewer invocations to accurately detect performance regressions than other state-of-the-art approaches. In $98.41 \%$ of evaluated cases, our approach provides equal or smaller confidence interval size. DuetFaaS achieves an interval size reduction in $59.06 \%$ of all evaluated sample sizes when compared to the competitive approaches.
Tim C. Rese, Nils Japke, Tobias Pfandzelter, David Bermbach
IC2E5
2024 ElastiBench: Scalable Continuous Benchmarking on Cloud FaaS Platforms
abstract
Running microbenchmark suites often and early in the development process enables developers to identify performance issues in their application. Microbenchmark suites of complex applications can comprise hundreds of individual benchmarks and take multiple hours to evaluate meaningfully, making running those benchmarks as part of CI/CD pipelines infeasible. In this paper, we reduce the total execution time of microbenchmark suites by leveraging the massive scalability and elasticity of FaaS (Function-as-a-Service) platforms. While using FaaS enables users to quickly scale up to thousands of parallel function instances to speed up microbenchmarking, the performance variation and low control over the underlying computing resources complicate reliable benchmarking. We present ElastiBench, an architecture for executing microbenchmark suites on cloud FaaS platforms, and evaluate it on code changes from an open-source time series database. Our evaluation shows that our prototype can produce reliable results ($\sim 95 \%$ of performance changes accurately detected) in a quarter of the time ($\leq 15 \mathrm{~min}$ vs. $\sim 4 \mathrm{~h}$) and at lower cost ($\$0.49$ vs. $\$ 1.18$) compared to cloud-based virtual machines.
Trever Schirmer, Tobias Pfandzelter, David Bermbach
IC2E3
2024 FUSIONIZE++: Improving Serverless Application Performance Using Dynamic Task Inlining and Infrastructure Optimization
abstract
The Function-as-a-Service (FaaS) execution model increases developer productivity by removing operational concerns such as managing hardware or software runtimes. Developers, however, still need to partition their applications into FaaS functions, which is error-prone and complex: Encapsulating only the smallest logical unit of an application as a FaaS function maximizes flexibility and reusability. Yet, it also leads to invocation overheads, additional cold starts, and may increase cost due to double billing during synchronous invocations. Conversely, deploying an entire application as a single FaaS function avoids these overheads but decreases flexibility. In this paper we presentFusionize, a framework that automates optimizing for this trade-off by automatically fusing application code into an optimized multi-function composition. Developers only need to write fine-grained application code following the serverless model, whileFusionizeautomatically fuses different parts of the application into FaaS functions, manages their interactions, and configures the underlying infrastructure. At runtime, it monitors application performance and adapts it to minimize request-response latency and costs. Real-world use cases show thatFusionizecan improve the deployment artifacts of the application, reducing both median request-response latency and cost of an example IoT application by more than 35%.
Trever Schirmer, Joel Scheuner, Tobias Pfandzelter, David Bermbach
IEEE Trans. Cloud Comput.4
2023 Identifying Nearest Fog Nodes With Network Coordinate Systems
abstract
Identifying the closest fog node is crucial for mobile clients to benefit from fog computing. In this paper, we analyze the performance of the Meridian and Vivaldi network coordinate systems for this task. To that end, we simulate a dense fog environment with mobile clients. We find that while network coordinate systems really find fog nodes in close network proximity, a purely latency-oriented identification approach ignores the larger problem of balancing load across fog nodes.
Simon Huber, Tobias Pfandzelter, David Bermbach
IC2E3
2023 A Crowdsensing Approach for Deriving Surface Quality of Cycling Infrastructure
abstract
Cities worldwide are trying to increase the modal share of bicycle traffic to address traffic and carbon emission problems. Aside from safety, a key factor for this is the cycling comfort, including the surface quality of cycle paths. In this paper, we propose a novel edge-based crowdsensing method for analyzing the surface quality of bicycle paths using smartphone sensor data: Cyclists record their rides which after preprocessed on their phones before being uploaded to a private cloud backend. There, additional analysis modules aggregate data from all available rides to derive surface quality information which can then used for surface quality-aware routing and planning of infrastructure maintenance.
Ahmet-Serdar Karakaya, Leonard Thomas, Denis Koljada, David Bermbach
IC2E4
2023 Towards a Benchmark for Fog Data Processing
abstract
Fog data processing systems provide key abstractions to manage data and event processing in the geo-distributed and heterogeneous fog environment. The lack of standardized benchmarks for such systems, however, hinders their development and deployment, as different approaches cannot be compared quantitatively. Existing cloud data benchmarks are inadequate for fog computing, as their focus on workload specification ignores the tight integration of application and infrastructure inherent in fog computing.In this paper, we outline an approach to a fog-native data processing benchmark that combines workload specifications with infrastructure specifications. This holistic approach allows researchers and engineers to quantify how a software approach performs for a given workload on given infrastructure. Further, by basing our benchmark in a realistic IoT sensor network scenario, we can combine paradigms such as low-latency event processing, machine learning inference, and offline data analytics, and analyze the performance impact of their interplay in a fog data processing system.
Tobias Pfandzelter, David Bermbach
IC2E2
2023 Evaluating LEO Edge Software in the Cloud with Celestial
abstract
Celestial is a toolkit for building LEO edge testbeds on cloud VMs. This allows testing and benchmarking LEO edge software systems such as applications and platforms without access to real LEO satellite infrastructure. Celestial scales to emulate thousands of satellites servers on few cloud VMs using microVM technology. Its low footprint makes it cost-efficient for research and educational purposes without sacrificing flexibility and isolation. We show considerations for designing LEO edge software and demonstrate evaluating it on a Celestial testbed.
Tobias Pfandzelter, David Bermbach
IC2E2
2023 Eventually Consistent Configuration Management in Fog Systems With CRDTs
abstract
Fog systems rely on centralized and strongly consistent services for configuration management originally designed for cloud systems. Considering geo-distribution, such systems can exhibit high communication latency or become unavailable in case of network partition. In this paper, we examine the drawbacks of strong consistency for fog configuration management and propose an alternative based on CRDTs. We prototypically implement our approach for the FReD fog data management platform with promising early results.
Nick Stender, Tobias Pfandzelter, David Bermbach
IC2E3
2023 Kernel-as-a-Service: A Serverless Programming Model for Heterogeneous Hardware Accelerators
abstract
With the slowing of Moore's law and decline of Dennard scaling, computing systems increasingly rely on specialized hardware accelerators in addition to general-purpose compute units. Increased hardware heterogeneity necessitates disaggregating applications into workflows of fine-grained tasks that run on a diverse set of CPUs and accelerators. Current accelerator delivery models cannot support such applications efficiently, as (1) the overhead of managing accelerators erases performance benefits for fine-grained tasks; (2) exclusive accelerator use per task leads to underutilization; and (3) specialization increases complexity for developers.
Tobias Pfandzelter, Aditya Dhakal, Eitan Frachtenberg, Sai Rahul Chalamalasetti, Darel Emmot, Ninad Hogade, Rolando P. Hong Enriquez, Gourav Rattihalli, David Bermbach, Dejan S. Milojicic
Middleware9
2023 Achieving realistic cyclist behavior in SUMO using the SimRa dataset
Ahmet-Serdar Karakaya, Ioan-Alexandru Stef, Konstantin Köhler, Julian Heinovski, Falko Dressler, David Bermbach
Comput. Commun.6
2023 CycleSense: Detecting near miss incidents in bicycle traffic from mobile motion sensors
Ahmet-Serdar Karakaya, Thomas Ritter, Felix Bießmann, David Bermbach
Pervasive Mob. Comput.4
2023 Managing data replication and distribution in the fog with FReD
abstract
Summary The heterogeneous, geographically distributed infrastructure of fog computing poses challenges in data replication, data distribution, and data mobility for fog applications. Fog computing is still missing the necessary abstractions to manage application data, and fog application developers need to re‐implement data management for every new piece of software. Proposed solutions are limited to certain application domains, such as the IoT, are not flexible in regard to network topology, or do not provide the means for applications to control the movement of their data. In this paper, we present FReD, a data replication middleware for the fog. FReD serves as a building block for configurable fog data distribution and enables low‐latency, high‐bandwidth, and privacy‐sensitive applications. FReD is a common data access interface across heterogeneous infrastructure and network topologies, provides transparent and controllable data distribution, and can be integrated with applications from different domains. To evaluate our approach, we present a prototype implementation of FReD and show the benefits of developing with FReD using three case studies of fog computing applications.
Tobias Pfandzelter, Nils Japke, Trever Schirmer, Jonathan Hasenburg, David Bermbach
Softw. Pract. Exp.5
2023 Special Issue on benchmarking, experimentation tools, and reproducible practices for data-intensive systems from edge to cloud
abstract
As data analytics and machine learning increasingly permeate our cities, factories, and homes, the computing infrastructure for data-intensive systems becomes more challenging. That is, the vision of pervasive, intelligent, and cyber-physical IoT systems will not be realized with centralized cloud resources alone. Such resources are simply too far away from sensor-equipped devices and users, resulting in high latency, bandwidth bottlenecks, and unnecessary energy consumption. In addition, there are often privacy and security requirements that mandate distributed architectures. As a result, new distributed computing paradigms are emerging that promise to bring computing and storage closer to data sources and users. The emerging distributed computing environments of edge and fog computing provide additional resources within mobile networks, ISP infrastructure, and even LEO satellites. These diverse and dynamic computing environments pose significant challenges to the performance, dependability, and efficiency of data-intensive systems running on such infrastructure. At the same time, it is far less clear how to properly benchmark, evaluate, and test the behavior of systems that span IoT devices, edge nodes, and cloud resources. For example, IoT sensor data stream processing systems can be leveraged to continuously optimize the operation of urban infrastructures (such as public transportation systems, water networks, or medical infrastructures). The behavior of such systems must be thoroughly assessed before they can be deployed to edge and fog infrastructure. In addition, these systems must be evaluated reproducibly under the expected computing environment conditions, including variations of those conditions, given the inherently unsteady nature of IoT environments. In addition, there is growing concern about the energy consumption and greenhouse gas emissions of ICT (and especially distributed ML-based applications), which further warrants close examination of the behavior of new data-intensive applications. Despite significant research and development efforts to improve benchmarking, experimentation tools, and reproducible practices for data-intensive systems spanning from the edge to the cloud, more research is urgently needed. We therefore invited high-quality research papers on this topic for this special issue of Software: Practice and Experience, and we were able to select two out of four submissions for this special issue with the help of our reviewers. The first accepted paper is titled “faas-sim: A Trace-Driven Simulation Framework for Serverless Edge Computing Platforms”.1 It is co-authored by Philipp Raith, Thomas Rausch, Alireza Furutanpey, and Schahram Dustdar. The paper presents the design and implementation of a new simulation framework, “faas-sim,” for modeling and evaluating serverless software architectures spanning the edge-cloud continuum based on a scenario description, a given network topology, and workload traces. The new simulator is demonstrated by using it for performance estimation, resource planning, co-simulation, and scientific evaluation. The authors also evaluate faas-sim's network simulation and resource utilization. Furthermore, they highlight traces that come with faas-sim and provide an overview of published research that has used faas-sim. The second accepted paper is titled “Software-in-the-Loop Simulation for Developing and Testing Carbon-Aware Applications”.2 It is co-authored by Philipp Wiesner, Marvin Steinke, Henrik Nickel, Yazan Kitana, and Odej Kao. As an alternative to relying on purely simulated or purely real testbeds, the paper proposes the use of software-in-the-loop simulation and hybrid testbeds for testing carbon-aware software applications in the context of energy simulations. The paper describes the design and implementation of a prototype, “Vessim,” as well as two experiments demonstrating the capabilities and features of the novel tool. In this way, the paper shows how a message broker can reliably and realistically connect currently running applications under test to real-time simulations, while the energy demand is continuously measured or modeled. We are grateful to the Editor-in-Chief of the Journal, Dr. Rajkumar Buyya, for inviting us to organize this special issue. We are also grateful for the valuable support from the administrative office of the journal. In addition, we are very grateful for the thorough and thoughtful reviews provided by our reviewers. Finally, we appreciate the hard work and trust of the authors who submitted papers to our special issue.
Lauritz Thamsen, David Bermbach, Demetris Trihinas
Softw. Pract. Exp.2
2023 Using Microbenchmark Suites to Detect Application Performance Changes
abstract
Software performance changes are costly and often hard to detect pre-release. Similar to software testing frameworks, either application benchmarks or microbenchmarks can be integrated into quality assurance pipelines to detect performance changes before releasing a new application version. Unfortunately, extensive benchmarking studies usually take several hours which is problematic when examining dozens of daily code changes in detail; hence, trade-offs have to be made. Optimized microbenchmark suites, which only include a small subset of the full suite, are a potential solution for this problem, given that they still reliably detect the majority of the application performance changes such as an increased request latency. It is, however, unclear whether microbenchmarks and application benchmarks detect the same performance problems and one can be a proxy for the other. In this paper, we explore whether microbenchmark suites can detect the same application performance changes as an application benchmark. For this, we run extensive benchmark experiments with both the complete and the optimized microbenchmark suites of two time-series database systems, i.e.,InfluxDBandVictoriaMetrics, and compare their results to the results of corresponding application benchmarks. We do this for 70 and 110 commits, respectively. Our results show that it is not trivial to detect application performance changes using an optimized microbenchmark suite. The detection (i) is only possible if the optimized microbenchmark suite covers all application-relevant code sections, (ii) is prone to false alarms, and (iii) cannot precisely quantify the impact on application performance. For certain software projects, an optimized microbenchmark suite can, thus, provide fast performance feedback to developers (e.g., as part of a local build process), help estimating the impact of code changes on application performance, and support a detailed analysis while a daily application benchmark detects major performance problems. Thus, although a regular application benchmark cannot be substituted for both studied systems, our results motivate further studies to validate and optimize microbenchmark suites.
Martin Grambow, Denis Kovalev, Christoph Laaber, Philipp Leitner 0001, David Bermbach
IEEE Trans. Cloud Comput.5
2023 MockFog 2.0: Automated Execution of Fog Application Experiments in the Cloud
abstract
Fog computing is an emerging computing paradigm that uses processing and storage capabilities located at the edge, in the cloud, and possibly in between. Testing and benchmarking fog applications, however, is hard since runtime infrastructure will typically be in use or may not exist, yet. While approaches for the emulation of infrastructure testbeds do exist, their focus is typically the emulation of edge devices. Other approaches also emulate infrastructure within the core network or the cloud, but they miss support for automated experiment orchestration. In this article, we propose to evaluate fog applications on an emulated infrastructure testbed created in the cloud which can be manipulated based on a pre-defined orchestration schedule. Developers can freely design the infrastructure, configure performance characteristics, manage application components, and orchestrate their experiments. We also present our proof-of-concept implementation MockFog 2.0. We use MockFog 2.0 to evaluate a fog-based smart factory application and showcase how its features can be used to study the impact of infrastructure changes and workload variations. With these experiments, we also show that MockFog can achieve good experiment reproducibility, even in a public cloud environment.
Jonathan Hasenburg, Martin Grambow, David Bermbach
IEEE Trans. Cloud Comput.3
2022 Towards Distributed Coordination for Fog Platforms
abstract
Distributed fog and edge applications communicate over unreliable networks and are subject to high communication delays. This makes using existing distributed coordination technologies from cloud applications infeasible, as they are built on the assumption of a highly reliable, low-latency datacenter network to achieve strict consistency with low overheads. To help implement configuration and state management for fog platforms and applications, we propose a novel decentralized approach that lets systems specify coordination strategies and membership for different sets of coordination data.
Tobias Pfandzelter, Trever Schirmer, David Bermbach
CCGRID3
2022 Network Emulation in Large-Scale Virtual Edge Testbeds: A Note of Caution and the Way Forward
abstract
The growing research and industry interest in the Internet of Things and the edge computing paradigm has increased the need for cost-efficient virtual testbeds for large-scale distributed applications. Researchers, students, and practitioners need to test and evaluate the interplay of hundreds or thousands of real software components and services connected with a realistic edge network without access to phvsical infrastructure. While advances in virtualization technologies have enabled parts of this, network emulation as a crucial part in the development of edge testbeds is lagging behind: As we show in this paper, NetEm, the current state-of-the-art network emulation tooling included in the Linux kernel, imposes prohibitive scalability limits. We quantify these limits, investigate possible causes, and present a way forward for network emulation in large-scale virtual edge testbeds based on eBPFs.
Sören Becker 0001, Tobias Pfandzelter, Nils Japke, David Bermbach, Odej Kao
IC2E4
2022 Streaming vs. Functions: A Cost Perspective on Cloud Event Processing
abstract
In cloud event processing, data generated at the edge is processed in real-time by cloud resources. Both distributed stream processing (DSP) and Function-as-a-Service (FaaS) have been proposed to implement such event processing applications. FaaS emphasizes fast development and easy operation, while DSP emphasizes efficient handling of large data volumes. Despite their architectural differences, both can be used to model and implement loosely-coupled job graphs. In this paper, we consider the selection of FaaS and DSP from a cost perspective. We implement stateless and stateful workflows from the Theodolite benchmarking suite using cloud FaaS and DSP. In an extensive evaluation, we show how application type, cloud service provider, and runtime environment can influence the cost of application deployments and derive decision guidelines for cloud engineers.
Tobias Pfandzelter, Sören Henning, Trever Schirmer, Wilhelm Hasselbring, David Bermbach
IC2E5
2022 Fusionize: Improving Serverless Application Performance through Feedback-Driven Function Fusion
abstract
Serverless computing increases developer productivity by removing operational concerns such as managing hardware or software runtimes. Developers, however, still need to partition their application into functions, which can be error-prone and adds complexity: Using a small function size where only the smallest logical unit of an application is inside a function maximizes flexibility and reusability. Yet, having small functions leads to invocation overheads, additional cold starts, and may increase cost due to double billing during synchronous invocations. In this paper we present Fusionize, a framework that removes these concerns from developers by automatically fusing the application code into a multi-function orchestration with varying function size. Developers only need to write the application code following a lightweight programming model and do not need to worry how the application is turned into functions. Our framework automatically fuses different parts of the application into functions and manages their interactions. Leveraging monitoring data, the framework optimizes the distribution of application parts to functions to optimize deployment goals such as end-to-end latency and cost. Using two example applications, we show that Fusionizecan automatically and iteratively improve the deployment artifacts of the application.
Trever Schirmer, Joel Scheuner, Tobias Pfandzelter, David Bermbach
IC2E4
2022 QoS-Aware Resource Placement for LEO Satellite Edge Computing
abstract
With the advent of large LEO satellite communication networks to provide global broadband Internet access, interest in providing edge computing resources within LEO networks has emerged. The LEO Edge promises low-latency, high-bandwidth access to compute and storage resources for a global base of clients and IoT devices regardless of their geographical location.Current proposals assume compute resources or service replicas at every LEO satellite, which requires high upfront investments and can lead to over-provisioning. To implement and use the LEO Edge efficiently, methods for server and service placement are required that help select an optimal subset of satellites as server or service replica locations. In this paper, we show how the existing research on resource placement on a 2D torus can be applied to this problem by leveraging the unique topology of LEO satellite networks. Further, we extend the existing discrete resource placement methods to allow placement with QoS constraints. In simulation of proposed LEO satellite communication networks, we show how QoS depends on orbital parameters and that our proposed method can take these effects into account where the existing approach cannot.
Tobias Pfandzelter, David Bermbach
ICFEC2
2022 Celestial: Virtual Software System Testbeds for the LEO Edge
abstract
As private space companies such as SpaceX and Telesat are building large LEO satellite constellations to provide global broadband Internet access, researchers have proposed to embed compute services within satellite constellations to provide computing services on the *LEO edge*. While the LEO edge is merely theoretical at the moment, providers are expected to rapidly develop their satellite technologies to keep the upper hand in the new space race.
Tobias Pfandzelter, David Bermbach
Middleware2
2022 Special issue on co-design of data and computation management in Fog Computing
Monica Vitali, Pierluigi Plebani, David Bermbach, Erik Elmroth
Future Gener. Comput. Syst.3
2022 AuctionWhisk: Using an auction-inspired approach for function placement in serverless fog platforms
abstract
Abstract The Function‐as‐a‐Service (FaaS) paradigm has a lot of potential as a computing model for fog environments comprising both cloud and edge nodes, as compute requests can be scheduled across the entire fog continuum in a fine‐grained manner. When the request rate exceeds capacity limits at the resource‐constrained edge, some functions need to be offloaded toward the cloud. In this article, we present an auction‐inspired approach in which application developers bid on resources while fog nodes decide locally which functions to execute and which to offload in order to maximize revenue. Unlike many current approaches to function placement in the fog, our approach can work in an online and decentralized manner. We also present our proof‐of‐concept prototype AuctionWhisk that illustrates how such an approach can be implemented in a real FaaS platform. Through a number of simulation runs and system experiments, we show that revenue for overloaded nodes can be maximized without dropping function requests.
David Bermbach, Jonathan Bader, Jonathan Hasenburg, Tobias Pfandzelter, Lauritz Thamsen
Softw. Pract. Exp.1
2021 Edge (of the Earth) Replication: Optimizing Content Delivery in Large LEO Satellite Communication Networks
abstract
Large low earth orbit (LEO) satellite networks such as SpaceX's Starlink constellation promise to deliver low-latency, high-bandwidth Internet access with global coverage. As an alternative to terrestrial fiber as a global Internet backbone, they could potentially serve billions of Internet-connected devices. Currently, operators of CDNs exploit the hierarchical topology of the Internet to place points-of-presence near users, yet this approach is no longer possible when the topology changes to a single, wide-area, converged access and backhaul network.In this paper, we explore the opportunities of points-of-presence for CDNs within the satellite network itself, as it could provide better access latency for users while reducing operational costs for the satellite Internet service providers. We propose four strategies for selecting points-of-presence in satellite constellations that we evaluate through extensive simulation. In one case, we find that replicating web content within satellites can reduce bandwidth usage in the constellation by 93% over an approach without replication in the network, while storing only 0.01% of all content in individual satellites.
Tobias Pfandzelter, David Bermbach
CCGRID2
2021 On the Future of Cloud Engineering
abstract
Ever since the commercial offerings of the Cloud started appearing in 2006, the landscape of cloud computing has been undergoing remarkable changes with the emergence of many different types of service offerings, developer productivity enhancement tools, and new application classes as well as the manifestation of cloud functionality closer to the user at the edge. The notion of utility computing, however, has remained constant throughout its evolution, which means that cloud users always seek to save costs of leasing cloud resources while maximizing their use. On the other hand, cloud providers try to maximize their profits while assuring service-level objectives of the cloud-hosted applications and keeping operational costs low. All these outcomes require systematic and sound cloud engineering principles. The aim of this paper is to highlight the importance of cloud engineering, survey the landscape of best practices in cloud engineering and its evolution, discuss many of the existing cloud engineering advances, and identify both the inherent technical challenges and research opportunities for the future of cloud computing in general and cloud engineering in particular.
David Bermbach, Abhishek Chandra, Chandra Krintz, Aniruddha S. Gokhale, Aleksander Slominski, Lauritz Thamsen, Everton Cavalcante, Tian Guo 0001, Ivona Brandic, Richard Wolski
IC2E1
2021 BeFaaS: An Application-Centric Benchmarking Framework for FaaS Platforms
abstract
Following the increasing interest and adoption of FaaS systems, benchmarking frameworks for determining nonfunctional properties have also emerged. While existing (microbenchmark) frameworks only evaluate single aspects of FaaS platforms, a more holistic, application-driven approach is still missing. In this paper, we design and present BeFaaS, an extensible application-centric benchmarking framework for FaaS environments that focuses on the evaluation of FaaS platforms through realistic and typical examples of FaaS applications. BeFaaS includes a built-in e-commerce benchmark, is extensible for new workload profiles and new platforms, supports federated benchmark runs in which the benchmark application is distributed over multiple providers, and supports a fine-grained result analysis. Our evaluation compares three major FaaS providers in single cloud provider setups and shows that BeFaaS is capable of running each benchmark automatically with minimal configuration effort and providing detailed insights for each interaction.
Martin Grambow, Tobias Pfandzelter, Luk Burchard, Carsten Schubert, Max Xiaohang Zhao, David Bermbach
IC2E6
2021 Towards Predictive Replica Placement for Distributed Data Stores in Fog Environments
abstract
Mobile clients that consume and produce data are abundant in fog environments. Low latency access to this data can only be achieved by storing it in close physical proximity to the clients. Current data store systems fall short as they do not replicate data based on client movement. We propose an approach to predictive replica placement that autonomously and proactively replicates data close to likely client locations.
Tobias Pfandzelter, David Bermbach
IC2E2
2021 From zero to fog: Efficient engineering of fog-based Internet of Things applications
abstract
Abstract In Internet of Things (IoT) data processing, cloud computing alone does not suffice due to latency constraints, bandwidth limitations, and privacy concerns. By introducing intermediary nodes closer to the edge of the network that offer compute services in proximity to IoT devices, fog computing can reduce network strain and high access latency to application services. While this is the only viable approach to enable efficient IoT applications, the issue of component placement among cloud and intermediary nodes in the fog adds a new dimension to system design. State‐of‐the‐art solutions to this issue rely on simulation or solving a formalized assignment problem through heuristics only, which both have their drawbacks. In this article, we present a five‐step process for designing practical fog‐based IoT applications that combines best practices, simulation, and testbed analysis to converge towards an efficient system architecture. We then apply this process in a smart factory case study. By deploying filtered options to a physical testbed, we show that each step of our process converges towards more efficient application designs.
Tobias Pfandzelter, Jonathan Hasenburg, David Bermbach
Softw. Pract. Exp.3
2020 GeoBroker: Leveraging geo-contexts for IoT data distribution
Jonathan Hasenburg, David Bermbach
Comput. Commun.2
2020 SimRa: Using crowdsourcing to identify near miss hotspots in bicycle traffic
Ahmet-Serdar Karakaya, Jonathan Hasenburg, David Bermbach
Pervasive Mob. Comput.3
2019 Continuous Benchmarking: Using System Benchmarking in Build Pipelines
abstract
Continuous integration and deployment are established paradigms in modern software engineering. Both intend to ensure the quality of software products and to automate the testing and release process. Today's state of the art, however, focuses on functional tests or small microbenchmarks such as single method performance while the overall quality of service (QoS) is ignored. In this paper, we propose to add a dedicated benchmarking step into the testing and release process which can be used to ensure that QoS goals are met and that new system releases are at least as "good" as the previous ones. For this purpose, we present a research prototype which automatically deploys the system release, runs one or more benchmarks, collects and analyzes results, and decides whether the release fulfills predefined QoS goals. We evaluate our approach by replaying two years of Apache Cassandra's commit history.
Martin Grambow, Fabian Lehmann, David Bermbach
IC2E3
2017 Audio-visual Cues for Cloud Service Monitoring
David Bermbach, Jacob Eberhardt
CLOSER1
2017 BenchFoundry: A Benchmarking Framework for Cloud Storage Services
David Bermbach, Jörn Kuhlenkamp, Akon Dey, Arunmoezhi Ramachandran, Alan D. Fekete, Stefan Tai
ICSOC1
2016 Pick your choice in HBase: Security or performance
abstract
When analyzing sensitive data in a cloud-deployed Hadoop stack, data-in-transit security needs to be enabled, especially in the underlying storage tier. This, however, will affect the performance of the system and may partially offset the cost benefits of the cloud. In this paper, we discuss two strategies for securing HBase deployments in the cloud. For both, we present benchmarking results which show performance impacts that significantly exceed the suggested 10% from the official documentation. These results demonstrate (i) that security configurations should follow a rational decision process based on benchmarking results and (ii) that the security architecture of HBase/HDFS should be redesigned with an emphasis on performance.
Frank Pallas, Johannes Günther 0003, David Bermbach
IEEE BigData3
2016 Towards Audio-Visual Cues for Cloud Infrastructure Monitoring
abstract
When monitoring their systems' states, DevOps engineers and operations teams alike, today, have to choose whether they want to dedicate their full attention to a visual dashboard showing monitoring results or whether they want to rely on threshold-or algorithm-based alarms which always come with false positive and false negative signals. In this work, we propose an alternative approach which translates a stream of cloud monitoring data into a continuous, normalized stream of score changes. Based on the score level, we propose to gradually change environment factors, e.g., music output or ambient lighting. We do this with the goal of enabling developers to subconsciously become aware of changes in monitoring data while dedicating their full attention to their primary task.
David Bermbach, Jacob Eberhardt
IC2E1
2016 Benchmarking Web API Quality
David Bermbach, Erik Wittern
ICWE1
2015 An Introduction to Cloud Benchmarking
abstract
Over the last few years, more and more Cloud Computing offerings have emerged ranging from compute, data storage, and middleware services over platform environments up to ready-to-use applications. Choosing the best offering for a particular use case, is a complex task which involves comparison and trade-off analysis of functional and non-functional service properties; for non-functional quality of service (QoS) properties, this is typically done via benchmarking. Today, a plethora of benchmarking solutions exist for different layers in the cloud stack (IaaS, PaaS, SaaS) which typically address a single QoS dimension - a holistic cloud benchmark even for a single layer in the cloud stack is still missing. In this tutorial, we will give an overview of existing cloud benchmarking solutions and point-out ways in which these different benchmarks could be used in concert to actually compare clouds as a whole (i.e., for instance Amazon cloud vs. Google cloud) instead of analyzing isolated QoS dimensions of single cloud services.
David Bermbach
IC2E1
2015 AISLE: Assessment of Provisioned Service Levels in Public IaaS-Based Database Systems
Jörn Kuhlenkamp, Kevin Rudolph, David Bermbach
ICSOC3
2014 Benchmarking Eventual Consistency: Lessons Learned from Long-Term Experimental Studies
abstract
Cloud storage services and NoSQL systems typically guarantee only Eventual Consistency. Knowing the degree of inconsistency increases transparency and comparability, it also eases application development. As every change to the system implementation, configuration, and deployment may affect the consistency guarantees of a storage system, long-term experiments are necessary to analyze how consistency behavior evolves over time. Building on our original publication on consistency benchmarking, we describe extensions to our benchmarking approach and report the surprising development of consistency behavior in Amazon S3 over the last two years. Based on our findings, we argue that consistency behavior should be monitored continuously and that deployment decisions should be reconsidered periodically. For this purpose, we propose a new method called Indirect Consistency Monitoring which allows to track all application-relevant changes in consistency behavior in a much more cost-efficient way compared to continuously running consistency benchmarks.
David Bermbach, Stefan Tai
IC2E1
2014 Benchmarking the Performance Impact of Transport Layer Security in Cloud Database Systems
abstract
Cloud storage services and NoSQL systems are optimized for performance and availability. Hence, enterprise-grade features like security mechanisms are typically neglected even though there is a need for them with increased cloud adoption by enterprises. Only Transport Layer Security (TLS) is frequently supported. Furthermore, the standard Transport Layer Security (TLS) protocol offers many configuration options which are usually chosen purely based on chance. We argue that in cloud database systems, configuration options should be chosen based on the degree of vulnerability to attacks and security threats as well as on the performance overhead of the respective algorithms. Our contributions are a benchmarking approach for transparent analysis of the performance impact of various TLS configuration options and a custom TLS socket implementation which offers more fine-grained control over the configuration options chosen. We also use our benchmarking approach to study the performance impact of TLS in Amazon DynamoDB and Apache Cassandra.
Steffen Müller 0002, David Bermbach, Stefan Tai, Frank Pallas
IC2E2
2013 A Middleware Guaranteeing Client-Centric Consistency on Top of Eventually Consistent Datastores
abstract
Applications often have consistency requirements beyond those guaranteed by the underlying eventually consistent storage system. In this work, we present an approach that guarantees monotonic read consistency and read your writes consistency by running a special middleware component on the same server as the application. We evaluate our approach using both simulation and real world experiments on Cloud storage systems.
David Bermbach, Jörn Kuhlenkamp, Bugra Derre, Markus Klems, Stefan Tai
IC2E1
2013 Cloud Federation: Effects of Federated Compute Resources on Quality of Service and Cost
abstract
Cloud Federation is one concept to confront challenges that still persist in Cloud Computing, such as vendor lock-in or compliance requirements. The lack of a standardized meaning for the term Cloud Federation has led to multiple conflicting definitions and an unclear prospect of its possible benefits. Taking a client-side perspective on federated compute services, we analyse how choosing a certain federation strategy affects Quality of Service and cost of the resulting service or application. Based on a use case, we experimentally prove our analysis to be correct and describe the different trade-offs that exist within each of the strategies.
David Bermbach, Tobias Kurze, Stefan Tai
IC2E1
2011 MetaStorage: A Federated Cloud Storage System to Manage Consistency-Latency Tradeoffs
abstract
Cost and scalability benefits of Cloud storage services are apparent. However, selecting a single storage service provider limits availability and scalability to the selected provider and may further cause a vendor lock-in effect. In this paper, we present MetaStorage, a federated Cloud storage system that can integrate diverse Cloud storage providers. MetaStorage is a highly available and scalable distributed hash table that replicates data on top of diverse storage services. MetaStorage reuses mechanisms from Amazon's Dynamo for cross-provider replication and hence introduces a novel approach to manage consistency-latency tradeoffs by extending the traditional quorum (N,R,W) configurations to an (N_P,R,W) scheme that includes different providers as an additional dimension. With MetaStorage, new means to control consistency-latency tradeoffs are introduced.
David Bermbach, Markus Klems, Stefan Tai, Michael Menzel 0002
IEEE CLOUD1