EDBT 2026 Demo / reviewers in the wild / expert
André Bauer 0001
dblp:17/3178
· DBLP profile ↗
40ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0002-5582-8812ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 4 first-author · 15 since 2021Systems, architecture and hardware · 10 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Kubernetes Performance for GenAI Inference: From Automatic Speech Recognition to LLM Summarization
Sai Sindhur Malleni, Raul Sevilla Canavate, Aleksei Vasilevskii, José Castillo Lema, André Bauer 0001 |
ICPE | 5 |
| 2025 | TACO: A Lightweight Tree-Based Approximate Compression Method for Time Series
André Bauer 0001 |
DATA | 1 |
| 2025 | EvoAgents: Generative AI for Rapid Architecture Search in Cell Nuclei DetectionabstractScience relies on image data across several disciplines like microbiology, but developing effective deep learning models for segmentation and detection tasks present challenges, including limited time, expertise, or financial resources. We introduce an evolutionary agent that autonomously generates, evaluates, and evolves convolutional neural networks (CNNs) based on natural language descriptions. Powered by large language models (LLMs) and guided by a genetic algorithm, our agent allows scientists to obtain trainable, reproducible model code for their specific tasks without manual development. A case study on cell nuclei detection demonstrates the pipeline’s feasibility and therefore establishes a baseline for architectures driven by LLMs. Designed for generalizability to diverse scientific imaging problems, this agent acts as a collaborative interface between deep learning and science and therefore contributing to scientific discovery and reduced economic and technical barriers. Samantha Alonso, André Bauer 0001 |
eScience | 2 |
| 2025 | Efficient Privacy-Preserving Recommendation on Sparse Data using Fully Homomorphic EncryptionabstractIn today’s data-driven world, recommendation systems personalize user experiences across industries but rely on sensitive data, raising privacy concerns. Fully homomorphic encryption (FHE) can secure these systems, but a significant challenge in applying FHE to recommendation systems is efficiently handling the inherently large and sparse user-item rating matrices. FHE operations are computationally intensive, and naively processing various sparse matrices in recommendation systems would be prohibitively expensive. Additionally, the communication overhead between parties remains a critical concern in encrypted domains. We propose a novel approach combining Compressed Sparse Row (CSR) representation with FHE-based matrix factorization that efficiently handles matrix sparsity in the encrypted domain while minimizing communication costs. Our experimental results demonstrate high recommendation accuracy with encrypted data while achieving the lowest communication costs, effectively preserving user privacy. Moontaha Nishat Chowdhury, André Bauer 0001, Minxuan Zhou |
eScience | 2 |
| 2025 | Finding the Sweet Spot: Speed vs. Fidelity in Encrypted LLMs via Proxy SimulationabstractFully homomorphic encryption (FHE) enables arithmetic operations to be performed directly on ciphertext, but this encrypted arithmetic is orders of magnitude slower than plaintext. When securing a large language model (LLM) using the CKKS encryption scheme, the slowdown is compounded by the need to replace each non-linear layer with a polynomial, which also injects approximation noise that can erode accuracy. Higher-degree polynomials reduce this error yet increase multiplicative depth and latency. Therefore, designers need to pinpoint where additional depth no longer pays off. To this end, we introduce a lightweight noise-injection proxy that measures CKKS residuals once, stores them, and replays the noise inside the plaintext model while sweeping polynomial degrees. Our proxy produces an accuracy-versus-latency curve and reveals the "sweet spot": The polynomial degree beyond which extra depth yields only marginal accuracy gains for a steep time cost. This curve guides developers in choosing an appropriate polynomial degree, eliminating the need for time-consuming encryption benchmarks. Rajan Savani, Anwar Benhnini, André Bauer 0001 |
eScience | 3 |
| 2025 | An Empirical Study on Transient Phases of Microservice ApplicationsabstractMicroservice applications in practice often face runtime adaptations, like autoscaling or updates. These adaptations cause these applications to leave their steady-state performance, leading to so-called transient phases. There is a lack of understanding and quantitative data about the transient phases of microservice applications. Several use cases, including microbenchmarking and autoscaling, can profit from empirical data on transient phases. To this end, we present the most comprehensive empirical study of transient phases to date. We investigated 92 microservices from 10 reference applications, analyzing the influence of programming languages, workloads, and container properties on the duration of transient phases. Ivo Rohwer, Martin Sträßer, Yannik Lubas, Samuel Kounev, André Bauer 0001 |
MASCOTS | 5 |
| 2025 | Addressing Reproducibility Challenges in HPC with Continuous IntegrationabstractThe high-performance computing (HPC) community has adopted incentive structures to motivate reproducible research, with major conferences awarding badges to papers that meet reproducibility requirements. Yet, many papers do not meet such requirements. The uniqueness of HPC infrastructure and software, coupled with strict access requirements, may limit opportunities for reproducibility. In the absence of resource access, we believe that regular documented testing, through continuous integration (CI), coupled with complete provenance information, can be used as a substitute. Here, we argue that better HPC-compliant CI solutions will improve reproducibility of applications. We present a survey of reproducibility initiatives and describe the barriers to reproducibility in HPC. To address existing limitations, we present a GitHub Action, CORRECT, that enables secure execution of tests on remote HPC resources. We evaluate CORRECT’s usability across three different types of HPC applications, demonstrating the effectiveness of using CORRECT for automating and documenting reproducibility evaluations. Valérie Hayot-Sasson, Nathaniel Hudson 0001, André Bauer 0001, Maxime Gonthier, Ian T. Foster, Kyle Chard |
SC | 3 |
| 2025 | Core Hours and Carbon Credits: Incentivizing Sustainability in HPCabstractEfforts to reduce the environmental impact of HPC often focus on resource providers, but choices made by users, e.g., concerning where to run, can be equally consequential. Here we present evidence that new accounting methods that charge users for energy used can incentivize significantly more efficient behavior. We first survey 300 HPC users and find that fewer than 30% are aware of their energy consumption, and that energy efficiency is a low priority concern. We then propose two new multi-resource accounting methods that charge for computations based on their energy consumption or carbon footprint, respectively. Finally, we conduct both simulation studies and a user study to evaluate the impact of these two methods on user behavior. We find that while only providing users feedback on their energy use had no impact on their behavior, associating energy with cost incentivized users to select more efficient resources, and use 40% less energy. Alok Kamatar, Maxime Gonthier, Valérie Hayot-Sasson, André Bauer 0001, Marcin Copik, Raul Castro Fernandez, Torsten Hoefler, Kyle Chard, Ian T. Foster |
SC | 4 |
| 2025 | Quantifying Data Leakage in Failure Prediction TasksabstractWith the ever increasing importance of cloud computing and a strong focus on reliable data centers, a high amount of research has been done on failure prediction for hard disk drives. The collection of monitoring data, such as SMART statistics (Self-Monitoring, Analysis, and Reporting Technology) from operational HDDs, enables operators to obtain predictions about the expected remaining useful life. Numerous methods for HDD failure prediction have been published in recent years, and their evaluation has shown decent results. However, a naive splitting into training and test sets can lead to data leakage and, thus, over-optimistic results that cannot be achieved on the data of scientific interest. In this paper, we propose a novel data leakage measure for quantifying the amount of data leakage in training and test datasets. Further, we define four splitting techniques and evaluate our measure in terms of the performance optimism of classification models with respect to these different splitting strategies. Our results consistently show that splitting techniques prone to data leakage induce an overestimation of predictive performance. Overall, we were able to show the usefulness of the defined data leakage measure, as well as its connection with different splitting techniques and the performance optimism of prediction models. Daniel Grillmeyer, Marius Hadry, Veronika Lesch, Vanessa Borst, Robert Leppich, André Bauer 0001, Samuel Kounev |
ICPE | 6 |
| 2025 | Generating Executable Microservice Applications for Performance BenchmarkingabstractMicroservice applications are the building blocks of modern cloud applications. As such, their performance aspects have been receiving increasing attention in the software engineering community. However, many microservice performance studies use only a small set of popular microservice test applications for experiments, questioning the applicability of their approaches in practice. Researchers currently lack the opportunity to collect large and diverse datasets containing performance metrics of microservices. This is because popular test applications only represent specific technology stacks and often come with custom benchmark tooling (e.g., load generation and monitoring). In this paper, we present Creo, a framework for generating microservice applications that (1) are fully executable, (2) have configurable properties and resource usage profiles, and (3) have built-in support for standardized monitoring, load generation, and deployment. Our approach enables researchers to run experiments with diverse microservice applications with minimal effort. We demonstrate the value of our approach in the context of two use cases. First, we show that using generated applications when training machine learning models for predicting performance degradation can improve the prediction accuracy. Second, we evaluate a recent approach for performance anomaly classification on a set of generated applications highlighting strengths and weaknesses not discussed in the original work. Yannik Lubas, Martin Sträßer, André Bauer 0001, Samuel Kounev |
ICPE | 3 |
| 2025 | Bridging Clusters: A Comparative Look at Multi-Cluster Networking Performance in KubernetesabstractMicroservices and containers have transformed the way applications are developed, tested, deployed, scaled, and managed. Several container orchestration platforms, like Kubernetes, have emerged, streamlining container management at scale and providing enterprise-grade support for application modernization. Driven by application, compliance, and end-user requirements, companies opt to deploy multiple Kubernetes clusters across public and private clouds. However, deploying applications in multi-cluster environments presents distinct challenges, especially managing the communication between the microservices spread across clusters. Traditionally, custom configurations, like VPNs or firewall rules, were required to connect such complex setups of clusters spanning the public cloud and on-premise infrastructure. This industry paper presents a comprehensive analysis of network performance characteristics for three popular open-source multi-cluster networking solutions (namely, Skupper, Submariner, and Istio), addressing the challenges of microservices connectivity across clusters. We evaluate key factors such as latency, throughput, and resource utilization using established tools and benchmarks, offering valuable insights for organizations aiming to optimize the network performance of their multi-cluster deployments. Our experiments revealed that each solution involves unique trade-offs in performance and resource efficiency: Submariner offers low latency and consistency, Istio excels in throughput with moderate resource consumption, and Skupper stands out for its ease of configuration while maintaining balanced performance. Sai Sindhur Malleni, Raul Sevilla Canavate, José Castillo Lema, André Bauer 0001 |
ICPE | 4 |
| 2025 | Telling fortunes? Evaluation of traffic forecasting models using traffic and context featuresabstractAbstract The need for efficient and reliable logistics solutions has increased significantly in the last decade. Traffic forecasts are a promising source of information that can be used to improve the planning of delivery schedules. However, most existing traffic forecasting approaches only support a forecasting horizon of up to an hour, which is insufficient for per-day-based schedule planning. In this paper, we focus on short-term traffic forecasting for up to four hours. We first propose a data collection process integrating traffic speed, incidents, weather, and holiday information. We have used this process to collect real-world traffic data for 115 days. We then define and evaluate twelve models for vehicle traffic forecasting, including well-known time series forecasting approaches and state-of-the-art deep learning models. Our results show that the best model in our comparison improved the accuracy by approximately 30% compared to a naive forecaster that repeats the last known value. The evaluation also shows that LSTM-based approaches are competitive to state-of-the-art models. Overall, the proposed deep-learning-based models perform best while requiring a smaller input timeframe than statistical models. Marius Hadry, André Bauer 0001, Robert Leppich, Veronika Lesch, Samuel Kounev |
Appl. Intell. | 2 |
| 2025 | Quo Vadis CKKS: Comparison of the realization of basic mathematical functions for the homomorphic cryptosystem CKKS using De Bello and polynomial approximationsabstractAs data storage and processing increasingly shift to the cloud, the risk of data breaches also rises. One way to address this is using Homomorphic Encryption (HE), which allows for data processing while the data remains encrypted, unlike traditional methods. However, current HE libraries support only addition and multiplication, requiring users to implement other mathematical functions themselves. To this end, we developed and analyzed basic mathematical functions in a previous work. Since polynomial approximations are more common in HE, this paper expands on that by examining and comparing polynomial approximations of these functions with the previously implemented methods. Our findings indicate that while polynomial approximations offer the benefit of low multiplication depth, the previously implemented methods generally outperform them in most scenarios despite their higher computational cost. Thomas Prantl, Lukas Horn, Simon Engel, André Bauer 0001, Samuel Kounev |
J. Inf. Secur. Appl. | 4 |
| 2025 | Object Proxy Patterns for Accelerating Distributed ApplicationsabstractWorkflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. We evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage. J. Gregory Pauloski, Valérie Hayot-Sasson, Logan T. Ward, Alex Brace, André Bauer 0001, Kyle Chard, Ian T. Foster |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | Security Analysis of a Decentralized, Revocable and Verifiable Attribute-Based Encryption SchemeabstractIn recent years, digital services have experienced significant growth, exemplified by platforms like Netflix achieving unprecedented revenue levels. Some of these services employ subscription models, with certain content requiring additional payments or offering third-party products. To ensure the widespread availability of diverse digital services anytime and anywhere, providers must have control over content accessibility. To address the multifaceted challenges in this domain, one promising solution is the adoption of attribute-based encryption (ABE). Over the years, various approaches have been proposed in the literature, offering a wide range of features. In a prior study [18], we assessed the security of one of these proposed approaches and identified one that did not meet its promised security standards. In this research we focuses on conducting a security analysis for another ABE scheme to pinpoint its shortcomings and emphasize the critical importance of evaluating the safety and effectiveness of newly proposed schemes. Specifically, we uncover an attack vector within this ABE scheme, which enables malicious users to decrypt content without the required permissions or attributes. Furthermore, we propose a solution to rectify this identified vulnerability. Thomas Prantl, Marco Lauer, Lukas Horn, Simon Engel, David Dingel, André Bauer 0001, Christian Krupitzer, Samuel Kounev |
ARES | 6 |
| 2024 | An Empirical Investigation of Container Building Strategies and Warm Times to Reduce Cold Starts in Scientific Computing Serverless FunctionsabstractServerless computing has revolutionized application development and deployment by abstracting infrastructure management, allowing developers to focus on writing code. To do so, serverless platforms dynamically create execution environments, often using containers. The cost to create and deploy these environments is known as "cold start" latency, and this cost can be particularly detrimental to scientific computing workloads characterized by sporadic and dynamic demands. We investigate methods to mitigate cold start issues in scientific computing applications by pre-installing Python packages in container images. Using data from Globus Compute and Binder, we empirically analyze cold start behavior and evaluate four strategies for building containers, including fully pre-built environments and dynamic, on-demand installations. Our results show that pre-installing all packages reduces initial cold start time but requires significant storage. Conversely, dynamic installation offers lower storage requirements but incurs repetitive delays. Additionally, we implemented a simulator and assessed the impact of different warm times, finding that moderate warm times significantly reduce cold starts without the excessive overhead of maintaining always-hot states. André Bauer 0001, Maxime Gonthier, Haochen Pan, Ryan Chard, Daniel Grzenda, Martin Sträßer, J. Gregory Pauloski, Alok Kamatar, Matt Baughman, Nathaniel Hudson 0001, Ian T. Foster, Kyle Chard |
e-Science | 1 |
| 2024 | The globus compute dataset: An open function-as-a-service dataset from the edge to the cloud
André Bauer 0001, Haochen Pan, Ryan Chard, Yadu N. Babuji, Josh Bryan, Devesh Tiwari, Ian T. Foster, Kyle Chard |
Future Gener. Comput. Syst. | 1 |
| 2023 | Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and VisionabstractDeep learning methods are transforming research, enabling new techniques, and ultimately leading to new discoveries. As the demand for more capable AI models continues to grow, we are now entering an era of Trillion Parameter Models (TPM), or models with more than a trillion parameters---such as Huawei's PanGu-Σ. We describe a vision for the ecosystem of TPM users and providers that caters to the specific needs of the scientific community. We then outline the significant technical challenges and open problems in system design for serving TPMs to enable scientific research and discovery. Specifically, we describe the requirements of a comprehensive software stack and interfaces to support the diverse and flexible requirements of researchers. Nathaniel Hudson 0001, J. Gregory Pauloski, Matt Baughman, Alok Kamatar, Mansi Sakarvadia, Logan T. Ward, Ryan Chard, André Bauer 0001, Maksim Levental, Will Engler, Owen Price Skelly, Ben Blaiszik, Rick L. Stevens, Kyle Chard, Ian T. Foster |
BDCAT | 8 |
| 2023 | An Empirical Study of Container Image Configurations and Their Impact on Start TimesabstractA core selling point of application containers is their fast start times compared to other virtualization approaches like virtual machines. Predictable and fast container start times are crucial for improving and guaranteeing the performance of containerized cloud, serverless, and edge applications. While previous work has investigated container starts, there remains a lack of understanding of how start times may vary across container configurations. We address this shortcoming by presenting and analyzing a dataset of approximately 200,000 open-source Docker Hub images featuring different image configurations (e.g., image size and exposed ports). Leveraging this dataset, we investigate the start times of containers in two environments and identify the most influential features. Our experiments show that container start times can vary between hundreds of milliseconds and tens of seconds in the same environment. Moreover, we conclude that no single dominant configuration feature determines a container's start time, and hardware and software parameters must be considered together for an accurate assessment. Martin Sträßer, André Bauer 0001, Robert Leppich, Nikolas Herbst, Kyle Chard, Ian T. Foster, Samuel Kounev |
CCGrid | 2 |
| 2023 | Performance Impact Analysis of Homomorphic Encryption: A Case Study Using Linear Regression as an Example
Thomas Prantl, Simon Engel, Lukas Horn, Dennis Kaiser, Lukas Iffländer, André Bauer 0001, Christian Krupitzer, Samuel Kounev |
ISPEC | 6 |
| 2023 | Autoscaler Evaluation and Configuration: A Practitioner's GuidelineabstractAutoscalers are indispensable parts of modern cloud deployments and determine the service quality and cost of a cloud application in dynamic workloads. The configuration of an autoscaler strongly influences its performance and is also one of the biggest challenges and showstoppers for the practical applicability of many research autoscalers. Many proposed cloud experiment methodologies can only be partially applied in practice, and many autoscaling papers use custom evaluation methods and metrics. This paper presents a practical guideline for obtaining meaningful and interpretable results on autoscaler performance with reasonable overhead. We provide step-by-step instructions for defining realistic usage behaviors and traffic patterns. We divide the analysis of autoscaler performance into a qualitative antipattern-based analysis and a quantitative analysis. To demonstrate the applicability of our guideline, we conduct several experiments with a microservice of our industry partner in a realistic test environment. Martin Sträßer, Simon Eismann, Jóakim von Kistowski, André Bauer 0001, Samuel Kounev |
ICPE | 4 |
| 2023 | A Systematic Approach for Benchmarking of Container Orchestration FrameworksabstractContainer orchestration frameworks play a critical role in modern cloud computing paradigms such as cloud-native or serverless computing. They significantly impact the quality and cost of service deployment as they manage many performance-critical tasks such as container provisioning, scheduling, scaling, and networking. Consequently, a comprehensive performance assessment of container orchestration frameworks is essential. However, until now, there is no benchmarking approach that covers the many different tasks implemented in such platforms and supports evaluating different technology stacks. In this paper, we present a systematic approach that enables benchmarking of container orchestrators. Based on a definition of container orchestration, we define the core requirements and benchmarking scope for such platforms. Each requirement is then linked to metrics and measurement methods, and a benchmark architecture is proposed. With COFFEE, we introduce a benchmarking tool supporting the definition of complex test campaigns for container orchestration frameworks. We demonstrate the potential of our approach with case studies of the frameworks Kubernetes and Nomad in a self-hosted environment and on the Google Cloud Platform. The presented case studies focus on container startup times, crash recovery, rolling updates, and more. Martin Sträßer, Jonas Mathiasch, André Bauer 0001, Samuel Kounev |
ICPE | 3 |
| 2023 | A literature review of IoT and CPS - What they are, and what they are not
Veronika Lesch, Marwin Züfle, André Bauer 0001, Lukas Iffländer, Christian Krupitzer, Samuel Kounev |
J. Syst. Softw. | 3 |
| 2022 | An Experience Report on the Suitability of a Distributed Group Encryption Scheme for an IoT Use CaseabstractThe critical component in any IoT application is the communication between devices. This must not only function smoothly, but also be secured. An important step in securing IoT communication is its encryption. However, in order for IoT devices to encrypt their communication with each other, they must first agree on appropriate cryptographic keys. In practice, the generation and distribution of such keys is usually managed by a central authority. However, this centralized approach has the disadvantages that (i) a central authority must be trusted, (ii) the central authority represents a single point of failure, and (iii) the central authority may be far away and thus communication with it takes a long time. To overcome these drawbacks, distributed group key agreement approaches have also been proposed. Since these distributed approaches were not originally developed for IoT devices, their performance on such devices is unknown. Therefore, in this work, we investigate the performance of a distributed group encryption scheme on IoT devices. To this end, we have built a measurement environment for distributed group encryption schemes and compare centralized and distributed group encryption schemes for IoT. Our measurements show that under perfect network conditions, the distributed approach performs worse than the centralized approaches in terms of time and memory requirements. However, our measurements also show that the distributed approach allows a group of 5 members to agree on a key in less than a minute. Thus, the distributed method can be used for small IoT groups if agreeing on a key is not time-sensitive. Thomas Prantl, Simon Engel, André Bauer 0001, Ala Eddine Ben Yahya, Stefan Herrnleben, Lukas Iffländer, Alexandra Dmitrienko, Samuel Kounev |
VTC Spring | 3 |
| 2022 | Same, Same, but Dissimilar: Exploring Measurements for Workload Time-series SimilarityabstractBenchmarking is a core element in the toolbox of most systems researchers and is used for analyzing, comparing, and validating complex systems. In the quest for reliable benchmark results, a consensus has formed that a significant experiment must be based on multiple runs. To interpret these runs, mean and standard deviation are often used. In case of experiments where each run produces a time series, applying and comparing the mean is not easily applicable and not necessarily statistically sound. Such an approach ignores the possibility of significant differences between runs with a similar average. In order to verify this hypothesis, we conducted a survey of 1,112 publications of selected performance engineering and systems conferences canvassing open data sets from performance experiments. The identified 3 data sets purely rely on average and standard deviation. Therefore, we propose a novel analysis approach based on similarity analysis to enhance the reliability of performance evaluations. Our approach evaluates 12 (dis-)similarity measures with respect to their applicability in analysing performance measurements and identifies four suitable similarity measures. We validate our approach by demonstrating the increase in reliability for the data sets found in the survey. Mark Leznik, Johannes Grohmann, Nina Kliche, André Bauer 0001, Daniel Seybold, Simon Eismann, Samuel Kounev, Jörg Domaschka |
ICPE | 4 |
| 2022 | Why Is It Not Solved Yet?: Challenges for Production-Ready AutoscalingabstractAutoscaling is a task of major importance in the cloud computing domain as it directly affects both operating costs and customer experience. Although there has been active research in this area for over ten years now, there is still a significant gap between the proposed methods in the literature and the deployed autoscalers in practice. Hence, many research autoscalers do not find their way into production deployments. This paper describes six core challenges that arise in production systems that are still not solved by most research autoscalers. We illustrate these problems through experiments in a realistic cloud environment with a real-world multi-service business application and show that commonly used autoscalers have various shortcomings. In addition, we analyze the behavior of overloaded services and show that these can be problematic for existing autoscalers. Generally, we analyze that these challenges are only insufficiently addressed in the literature and conclude that future scaling approaches should focus on the needs of production systems. Martin Sträßer, Johannes Grohmann, Jóakim von Kistowski, Simon Eismann, André Bauer 0001, Samuel Kounev |
ICPE | 5 |
| 2021 | Benchmarking of Pre- and Post-Quantum Group Encryption Schemes with Focus on IoTabstractIn the next few years, both the number of IoT devices and the performance of quantum computers will increase. Both technologies pose a challenge to our current crypto-strategies. Therefore, post-quantum n-to-n communication encryption is a crucial field of research. Here, the development of new schemes and the analysis, and comparison of existing schemes is necessary. However, current work only investigates the performance of post-quantum schemes only for 1-to-1 communication. Therefore, in this paper, we analyze existing post-quantum schemes concerning n-to-n communication and compare them with pre-quantum schemes. Our results show that the pre-quantum schemes perform better regarding computation times than the post-quantum schemes, but the differences are sometimes only marginal. However, these marginal differences in computation times lead to the lower energy efficiency of the post-quantum schemes. In terms of features, there is no difference between both scheme classes. We show that the post-quantum schemes require unicast, whereas some pre-quantum schemes also support broadcast. Deciding whether to use pre- or post-quantum schemes for n-to-n encryption in IoT use cases depends on (i) whether energy efficiency is essential – e.g., in case of limited power supply – and (ii) whether unicast or broadcast is available. Thomas Prantl, Dominik Prantl, André Bauer 0001, Lukas Iffländer, Alexandra Dmitrienko, Samuel Kounev, Christian Krupitzer |
IPCCC | 3 |
| 2021 | Libra: A Benchmark for Time Series Forecasting MethodsabstractIn many areas of decision making, forecasting is an essential pillar. Consequently, there are many different forecasting methods. According to the "No-Free-Lunch Theorem", there is no single forecasting method that performs best for all time series. In other words, each method has its advantages and disadvantages depending on the specific use case. Therefore, the choice of the forecasting method remains a mandatory expert task. However, expert knowledge cannot be fully automated. To establish a level playing field for evaluating the performance of time series forecasting methods in a broad setting, we propose Libra, a forecasting benchmark that automatically evaluates and ranks forecasting methods based on their performance in a diverse set of evaluation scenarios. The benchmark comprises four different use cases, each covering 100 heterogeneous time series taken from different domains. The data set was assembled from publicly available time series and was designed to exhibit much higher diversity than existing forecasting competitions. Based on this benchmark, we perform a comprehensive evaluation to compare different existing time series forecasting methods. André Bauer 0001, Marwin Züfle, Simon Eismann, Johannes Grohmann, Nikolas Herbst, Samuel Kounev |
ICPE | 1 |
| 2021 | SARDE: A Framework for Continuous and Self-Adaptive Resource Demand EstimationabstractResource demands are crucial parameters for modeling and predicting the performance of software systems. Currently, resource demand estimators are usually executed once for system analysis. However, the monitored system, as well as the resource demand itself, are subject to constant change in runtime environments. These changes additionally impact the applicability, the required parametrization as well as the resulting accuracy of individual estimation approaches. Over time, this leads to invalid or outdated estimates, which in turn negatively influence the decision-making of adaptive systems. In this article, we present SARDE , a framework for self-adaptive resource demand estimation in continuous environments. SARDE dynamically and continuously tunes, selects, and executes an ensemble of resource demand estimation approaches to adapt to changes in the environment. This creates an autonomous and unsupervised ensemble estimation technique, providing reliable resource demand estimations in dynamic environments. We evaluate SARDE using two realistic datasets. One set of different micro-benchmarks reflecting different possible system states and one dataset consisting of a continuously running application in a changing environment. Our results show that by continuously applying online optimization, selection and estimation, SARDE is able to efficiently adapt to the online trace and reduce the model error using the resulting ensemble technique. Johannes Grohmann, Simon Eismann, André Bauer 0001, Simon Spinner, Johannes Blum 0001, Nikolas Herbst, Samuel Kounev |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2021 | Methodological Principles for Reproducible Performance Evaluation in Cloud ComputingabstractThe rapid adoption and the diversification of cloud computing technology exacerbate the importance of a sound experimental methodology for this domain. This work investigates how to measure and report performance in the cloud, and how well the cloud research community is already doing it. We propose a set of eight important methodological principles that combine best-practices from nearby fields with concepts applicable only to clouds, and with new ideas about the time-accuracy trade-off. We show how these principles are applicable using a practical use-case experiment. To this end, we analyze the ability of the newly released SPEC Cloud IaaS benchmark to follow the principles, and showcase real-world experimental studies in common cloud environments that meet the principles. Last, we report on a systematic literature review including top conferences and journals in the field, from 2012 to 2017, analyzing if the practice of reporting cloud performance measurements follows the proposed eight principles. Worryingly, this systematic survey and the subsequent two-round human reviews, reveal that few of the published studies follow the eight experimental principles. We conclude that, although these important principles are simple and basic, the cloud community is yet to adopt them broadly to deliver sound measurement of cloud environments. Alessandro Vittorio Papadopoulos, Laurens Versluis, André Bauer 0001, Nikolas Herbst, Jóakim von Kistowski, Ahmed Ali-Eldin, Cristina L. Abad, José Nelson Amaral, Petr Tuma 0001, Alexandru Iosup |
IEEE Trans. Software Eng. | 3 |
| 2020 | Telescope: An Automatic Feature Extraction and Transformation Approach for Time Series Forecasting on a Level-Playing FieldabstractOne central problem of machine learning is the inherent limitation to predict only what has been learned -stationarity. Any time series property that eludes stationarity poses a challenge for the proper model building. Furthermore, existing forecasting methods lack reliable forecast accuracy and time-to-result if not applied in their sweet spot. In this paper, we propose a fully automated machine learning-based forecasting approach. Our Telescope approach extracts and transforms features from an input time series and uses them to generate an optimized forecast model. In a broad competition including the latest hybrid forecasters, established statistical, and machine learning-based methods, our Telescope approach shows the best forecast accuracy coupled with a lower and reliable time-to-result. André Bauer 0001, Marwin Züfle, Nikolas Herbst, Samuel Kounev, Valentin Curtef |
ICDE | 1 |
| 2020 | An Automated Forecasting Framework based on Method Recommendation for Seasonal Time SeriesabstractDue to the fast-paced and changing demands of their users, computing systems require autonomic resource management. To enable proactive and accurate decision-making for changes causing a particular overhead, reliable forecasts are needed. In fact, choosing the best performing forecasting method for a given time series scenario is a crucial task. Taking the "No-Free-Lunch Theorem" into account, there exists no forecasting method that performs best on all types of time series. To this end, we propose an automated approach that (i) extracts characteristics from a given time series, (ii) selects the best-suited machine learning method based on recommendation, and finally, (iii) performs the forecast. Our approach offers the benefit of not relying on a single method with its possibly inaccurate forecasts. In an extensive evaluation, our approach achieves the best forecasting accuracy. André Bauer 0001, Marwin Züfle, Johannes Grohmann, Norbert Schmitt, Nikolas Herbst, Samuel Kounev |
ICPE | 1 |
| 2020 | Time Series Forecasting for Self-Aware SystemsabstractModern distributed systems and Internet-of-Things applications are governed by fast living and changing requirements. Moreover, they have to struggle with huge amounts of data that they create or have to process. To improve the self-awareness of such systems and enable proactive and autonomous decisions, reliable time series forecasting methods are required. However, selecting a suitable forecasting method for a given scenario is a challenging task. According to the “No-Free-Lunch Theorem,” there is no general forecasting method that always performs best. Thus, manual feature engineering remains to be a mandatory expert task to avoid trial and error. Furthermore, determining the expected time-to-result of existing forecasting methods is a challenge. In this article, we extensively assess the state-of-the-art in time series forecasting. We compare existing methods and discuss the issues that have to be addressed to enable their use in a self-aware computing context. To address these issues, we present a step-by-step approach to fully automate the feature engineering and forecasting process. Then, following the principles from benchmarking, we establish a level-playing field for evaluating the accuracy and time-to-result of automated forecasting methods for a broad set of application scenarios. We provide results of a benchmarking competition to guide in selecting and appropriately using existing forecasting methods for a given self-aware computing context. Finally, we present a case study in the area of self-aware data-center resource management to exemplify the benefits of fully automated learning and reasoning processes on time series data. André Bauer 0001, Marwin Züfle, Nikolas Herbst, Albin Zehe, Andreas Hotho, Samuel Kounev |
Proc. IEEE | 1 |
| 2019 | Kaa: Evaluating Elasticity of Cloud-Hosted DBMSabstractAuto-scaling is able to change the scale of an application at runtime. Understanding the application characteristics, scaling impact as well as the workload, an auto-scaler aligns the acquired resources to match the current workload. For distributed Database Management Systems (DBMS) forming the backend of many large-scale cloud applications, it is currently an open question to what extent they support scaling at run-time. In particular, elasticity properties of existing distributed DBMS are widely unknown and difficult to evaluate and compare. This paper presents a comprehensive methodology for the evaluation of the elasticity of distributed DBMS. On the basis of this methodology, we introduce a framework that automates the full evaluation process. We validate the framework by defining significant elasticity scenarios for a case study that comprises two DBMS for write-heavy and read-heavy workloads of different intensities. The results show that scalable distributed DBMS are not necessarily elastic and that adding more instances to a cluster at run-time may even decrease the experienced performance. Daniel Seybold, Simon Volpert, Stefan Wesner, André Bauer 0001, Nikolas Herbst, Jörg Domaschka |
CloudCom | 4 |
| 2019 | Chamulteon: Coordinated Auto-Scaling of Micro-ServicesabstractNowadays, in order to keep track of the fast changing requirements of Internet applications, auto-scaling is used as an essential mechanism for adapting the number of provisioned resources to the resource demand. The straightforward approach is to deploy a set of common and opensource single-service auto-scalers for each service independently. However, this deployment leads to problems such as bottleneck-shifting and increased oscillations. Existing auto-scalers that scale applications consisting of multiple services are kept closed-source. To face these challenges, we first survey existing auto-scalers and highlight current challenges. Then, we introduce Chamulteon, a redesign of our previously introduced mechanism, which can scale applications consisting of multiple services in a coordinated manner. We evaluate Chamulteon against four different well-cited auto-scalers in four sets of measurement-based experiments where we use diverse environments (VM vs. Docker), real-world traces, and vary the scale of the demanded resources. Overall, Chamulteon achieves the best auto-scaling performance based on established user-oriented and endorsed elasticity metrics. André Bauer 0001, Veronika Lesch, Laurens Versluis, Alexey Ilyushkin, Nikolas Herbst, Samuel Kounev |
ICDCS | 1 |
| 2019 | Modeling of Aggregated IoT Traffic and Its Application to an IoT CloudabstractAs the Internet of Things (IoT) continues to gain traction in telecommunication networks, a very large number of devices are expected to be connected and used in the near future. In order to appropriately plan and dimension the network, as well as the back-end cloud systems and the resulting signaling load, traffic models are employed. These models are designed to accurately capture and predict the properties of IoT traffic in a concise manner. To achieve this, Poisson process approximations, based on the Palm–Khintchine theorem, have often been used in the past. Due to the scale (and the difference in scales in various IoT networks) of the modeled systems, the fidelity of this approximation is crucial, as, in practice, it is very challenging to accurately measure or simulate large-scale IoT deployments. The main goal of this paper is to understand the level of accuracy of the Poisson approximation model. To this end, we first survey both common IoT network properties and network scales as well as traffic types. Second, we explain and discuss the Palm–Khintiche theorem, how it is applied to the problem, and which inaccuracies can occur when using it. Based on this, we derive guidelines as to when a Poisson process can be assumed for aggregated periodic IoT traffic. Finally, we evaluate our approach in the context of an IoT cloud scaler use case. Florian Metzger, Tobias Hoßfeld, André Bauer 0001, Samuel Kounev, Poul E. Heegaard |
Proc. IEEE | 3 |
| 2019 | Chameleon: A Hybrid, Proactive Auto-Scaling Mechanism on a Level-Playing FieldabstractAuto-scalers for clouds promise stable service quality at low costs when facing changing workload intensity. The major public cloud providers provide trigger-based auto-scalers based on thresholds. However, trigger-based auto-scaling has reaction times in the order of minutes. Novel auto-scalers from literature try to overcome the limitations of reactive mechanisms by employing proactive prediction methods. However, the adoption of proactive auto-scalers in production is still very low due to the high risk of relying on a single proactive method. This paper tackles the challenge of reducing this risk by proposing a new hybrid auto-scaling mechanism, called Chameleon, combining multiple different proactive methods coupled with a reactive fallback mechanism. Chameleon employs on-demand, automated time series-based forecasting methods to predict the arriving load intensity in combination with run-time service demand estimation to calculate the required resource consumption per work unit without the need for application instrumentation. We benchmark Chameleon against five different state-of-the-art proactive and reactive auto-scalers one in three different private and public cloud environments. We generate five different representative workloads each taken from different real-world system traces. Overall, Chameleon achieves the best scaling behavior based on user and elasticity performance metrics, analyzing the results from 400 hours aggregated experiment time. André Bauer 0001, Nikolas Herbst, Simon Spinner, Ahmed Ali-Eldin, Samuel Kounev |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | TeaStore: A Micro-Service Reference Application for Benchmarking, Modeling and Resource Management ResearchabstractModern distributed applications offer complex performance behavior and many degrees of freedom regarding deployment and configuration. Researchers employ various methods of analysis, modeling, and management that leverage these degrees of freedom to predict or improve non-functional properties of the software under consideration. In order to demonstrate and evaluate their applicability in the real world, methods resulting from such research areas require test and reference applications that offer a range of different behaviors, as well as the necessary degrees of freedom. Existing production software is often inaccessible for researchers or closed off to instrumentation. Existing testing and benchmarking frameworks, on the other hand, are either designed for specific testing scenarios, or they do not offer the necessary degrees of freedom. Further, most test applications are difficult to deploy and run, or are outdated. In this paper, we introduce the TeaStore, a state-of-the-art micro-service-based test and reference application. TeaStore offers services with different performance characteristics and many degrees of freedom regarding deployment and configuration to be used as a benchmarking framework for researchers. The TeaStore allows evaluating performance modeling and resource management techniques; it also offers instrumented variants to enable extensive run-time analysis. We demonstrate TeaStore's use in three contexts: performance modeling, cloud resource management, and energy efficiency analysis. Our experiments show that TeaStore can be used for evaluating novel approaches in these contexts and also motivates further research in the areas of performance modeling and resource management. Jóakim von Kistowski, Simon Eismann, Norbert Schmitt, André Bauer 0001, Johannes Grohmann, Samuel Kounev |
MASCOTS | 4 |
| 2018 | FOX: Cost-Awareness for Autonomic Resource Management in Public CloudsabstractNowadays, to keep track with the fast changing requirements of internet applications, auto-scaling is an essential mechanism for adapting the number of provisioned resources to the resource demand. In the context of public clouds, there exist different natures of cost-models for charging resources. However, the accounted resource units and charged resource units may differ significantly due to the applied cost model. This can lead to a significant increase of charged costs when using an auto-scaler as it tries to match the demand of the application as close as possible. In the literature, several auto-scalers exist that support cost-aware scaling decisions but they introduce inherent drawbacks. In this work, this lack of existing cost-aware mechanisms is addressed by introducing a mediator between an application and the auto-scaler. This cost-aware mechanism is called FOX. It leverages knowledge of the charging model of the public cloud and reviews the scaling decisions found by the auto-scaler to reduce the charged costs to a minimum. More precisely, FOX delays or omits releases of resources to avoid additional charging costs if the resource is required in the future. Hereby, FOX is not restricted to use one specific auto-scaler but offers interfaces to use any auto-scaler. For an evalation under controlled conditions, FOX scales a multi-tier application deployed in a private cloud that is stressed with two real world workloads: BibSonomy and IBM CICS. As FOX provides an interface for auto-scalers, we evaluate the cost-aware mechanism with three state of the art auto-scalers: React, Adapt, and Reg. The experiments show that FOX is able to reduce the charged costs by 34% at maximum for the Amazon EC2 charging model. According to the cost model, FOX provisions more resources than required. This results in a decreased SLO violation rate from 28% to 2% at maximum. The accounted instance time increases at max. by 30%. Veronika Lesch, André Bauer 0001, Nikolas Herbst, Samuel Kounev |
ICPE | 2 |
| 2017 | Design and Evaluation of a Proactive, Application-Aware Auto-Scaler: Tutorial PaperabstractSimple, threshold-based auto-scaling mechanisms as mainly used in practice bring no features to overcome resource provisioning delays and non-linear scalability of a software service. In this tutorial paper, we guide the reader step-by-step through the design and evaluation of a proactive and application-aware auto-scaling mechanism. André Bauer 0001, Nikolas Herbst, Samuel Kounev |
ICPE | 1 |