VLDB 2026 Research / reviewers in the wild / expert
Nikolas Herbst
dblp:130/4084 · also Nikolas Roman Herbst
· DBLP profile ↗
30ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0003-3462-6426ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 1 first-author · 8 since 2021Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The FISHNET Case Study on Implementing and Scaling a Complex Earth Observation WorkflowabstractEarth observation (EO) scientists use sophisticated algorithms, large datasets, and high-performance data analytics (HPDA) clusters to develop and execute complex EO workflows. To improve portability and reusability within the domain, the Open Geospatial Consortium (OGC) published a set of best practices for developing EO workflows. Even though EO data products are often shared openly, representative EO workflow implementations are hardly available.In this paper, we contribute a case study on implementing and scaling a complex EO workflow, adhering to the OGC best practices and demonstrating its portability by deploying it both locally and on the HPDA platform "terrabyte" at the Leibniz Supercomputing Centre. Contentwise, the contributed workflow analyzes settlement patterns by first delineating coherent settlements derived from leveraged satellite data that maps builtup areas and subsequently calculating centrality measures to characterize the spatial arrangement and hierarchy of settlements in a given region.We demonstrate the workflow’s scalability and variance in resource demands by analyzing its time-to-result, total CPU time, and resource efficiency under different inputs and configurations. Interestingly, implicit parameters hidden in the input data semantics, like the number and area of settlements in the region of interest, significantly impact the time required to complete the processing. Concurrent processing is restricted to connected components of the settlement graph in the analysis stage, leading to an unbalanced workload distribution when analyzing large urban areas, showcasing the scalability challenges EO scientists face. Lorenz Gruber, Nikolas Herbst, Peter Friedl, Thomas Esch, Samuel Kounev |
eScience | 2 |
| 2025 | PARAGRAPH: Phase-Aware Resource Demand Profiling for HPDA/HPC JobsabstractThe processing of large amounts of data in central high performance data analytics (HPDA) systems is playing an increasingly important role in science and business. However, many HPDA systems exhibit a low utilization of their available resources during normal operation. An important reason for this underutilization is that too many resources are reserved for individual jobs. This is often a consequence of the common practice of reserving a uniform amount of resources such as CPU or memory for the entire execution time of a job. Given that many data intensive (DI) jobs consist of different phases with different resource demands, resources are normally reserved according to the demand of the most resource-intensive phase. This results in more resources being reserved over a long period of time than are actually needed. Ivo Rohwer, Nikolas Herbst, Maximilian Schwinger, Peter Friedl, Michael Stephan, Samuel Kounev |
ICPE | 2 |
| 2024 | ExDe: Design space exploration of scheduler architectures and mechanisms for serverless data-processingabstractServerless computing is increasingly used for data-processing applications in both science and business domains. At the core of serverless data-processing systems is the scheduler, which ensures dynamic decisions about task and data placement. Due to the variety of user, cluster, and workload properties, the design space for high-performance and cost-effective scheduling architectures and mechanisms is vast. The large design space is difficult to explore and characterize. To help the system designer disentangle this complexity, we present ExDe, a framework to systematically explore the design space of scheduling architectures and mechanisms. The framework includes a conceptual model and a simulator to assist in design space exploration. We use the framework, and real-world workloads, to characterize the performance of three scheduling architectures and two mechanisms. Our framework is open-source software available on Zenodo. Sacheendra Talluri, Nikolas Herbst, Cristina L. Abad, Tiziano De Matteis, Alexandru Iosup |
Future Gener. Comput. Syst. | 2 |
| 2023 | An Empirical Study of Container Image Configurations and Their Impact on Start TimesabstractA core selling point of application containers is their fast start times compared to other virtualization approaches like virtual machines. Predictable and fast container start times are crucial for improving and guaranteeing the performance of containerized cloud, serverless, and edge applications. While previous work has investigated container starts, there remains a lack of understanding of how start times may vary across container configurations. We address this shortcoming by presenting and analyzing a dataset of approximately 200,000 open-source Docker Hub images featuring different image configurations (e.g., image size and exposed ports). Leveraging this dataset, we investigate the start times of containers in two environments and identify the most influential features. Our experiments show that container start times can vary between hundreds of milliseconds and tens of seconds in the same environment. Moreover, we conclude that no single dominant configuration feature determines a container's start time, and hardware and software parameters must be considered together for an accurate assessment. Martin Sträßer, André Bauer 0001, Robert Leppich, Nikolas Herbst, Kyle Chard, Ian T. Foster, Samuel Kounev |
CCGrid | 4 |
| 2023 | A Trace-driven Performance Evaluation of Hash-based Task Placement Algorithms for Cache-enabled Serverless ComputingabstractData-driven interactive computation is widely used for business analytics, search-based decision-making, and log mining. These applications' short duration and bursty nature makes them a natural fit for serverless computing. Data processing serverless applications are composed of many small tasks. Application tasks that use remote storage encounter bottlenecks in the form of high latency, performance variability, and throttling. Caching has been used to mitigate this bottleneck for intermediate data. However, the use of caching for input data, albeit widely used in industry, has yet to be studied. We present the first performance study of scaling, a key feature of serverless computing, on serverless clusters with input data caches. We compare 8 task placement algorithms and quantify their impact on task slowdown and resource usage before and after scaling. We quantify the consequences of using work stealing. We quantify the performance impact of scaling in the buffer period immediately after scaling. We find up to a 420% increase in task slowdown after scaling without work stealing and a 22% slowdown with work stealing. We also find that cache misses after scaling can lead to an additional 21% resource usage. Sacheendra Talluri, Nikolas Herbst, Cristina L. Abad, Animesh Trivedi, Alexandru Iosup |
CF | 2 |
| 2022 | The State of Serverless Applications: Collection, Characterization, and Community ConsensusabstractOver the last five years, all major cloud platform providers have increased their serverless offerings. Many early adopters report significant benefits for serverless-based over traditional applications, and many companies are considering moving to serverless themselves. However, currently there exist only few, scattered, and sometimes even conflicting reports on when serverless applications are well suited and what the best practices for their implementation are. We address this problem in the present study about the state of serverless applications. We collect descriptions of 89 serverless applications from open-source projects, academic literature, industrial literature, and domain-specific feedback. We analyze 16 characteristics that describe why and when successful adopters are using serverless applications, and how they are building them. We further compare the results of our characterization study to 10 existing, mostly industrial, studies and datasets; this allows us to identify points of consensus across multiple studies, investigate points of disagreement, and overall confirm the validity of our results. The results of this study can help managers to decide if they should adopt serverless technology, engineers to learn about current practices of building serverless applications, and researchers and platform providers to better understand the current landscape of serverless applications. Simon Eismann, Joel Scheuner, Erwin Van Eyk, Maximilian Schwinger, Johannes Grohmann, Nikolas Herbst, Cristina L. Abad, Alexandru Iosup |
IEEE Trans. Software Eng. | 6 |
| 2021 | Sizeless: predicting the optimal size of serverless functionsabstractServerless functions are an emerging cloud computing paradigm that is being rapidly adopted by both industry and academia. In this cloud computing model, the provider opaquely handles resource management tasks such as resource provisioning, deployment, and auto-scaling. The only resource management task that developers are still in charge of is selecting how much resources are allocated to each worker instance. However, selecting the optimal size of serverless functions is quite challenging, so developers often neglect it despite its significant cost and performance benefits. Existing approaches aiming to automate serverless functions resource sizing require dedicated performance tests, which are time-consuming to implement and maintain. Simon Eismann, Long Bui, Johannes Grohmann, Cristina L. Abad, Nikolas Herbst, Samuel Kounev |
Middleware | 5 |
| 2021 | The Fourth Workshop on Hot Topics in Cloud Computing Performance (HotCloudPerf'21): Benchmarking in the CloudabstractThe HotCloudPerf workshop is a meeting venue for academics and practitioners, from experts to trainees, in the field of cloud computing performance. The workshop aims to engage this community, and to lead to the development of new methodological aspects for gaining deeper understanding not only of cloud performance, but also of cloud operation and behavior, through diverse quantitative evaluation tools, including benchmarks, metrics, and workload generators. The workshop focuses on novel cloud properties such as elasticity, performance isolation, dependability, and other non-functional system properties, in addition to classical performance-related metrics such as response time, throughput, scalability, and efficiency. The theme for the 2021 edition is "Benchmarking in the Cloud". HotCloudPerf 2021, co-located with the 12th ACM/SPEC International Conference on Performance Engineering (ICPE 2021), is held on April 19-20th, 2021. Cristina L. Abad, Nikolas Herbst, Alexandru Uta, Alexandru Iosup |
ICPE | 2 |
| 2021 | Libra: A Benchmark for Time Series Forecasting MethodsabstractIn many areas of decision making, forecasting is an essential pillar. Consequently, there are many different forecasting methods. According to the "No-Free-Lunch Theorem", there is no single forecasting method that performs best for all time series. In other words, each method has its advantages and disadvantages depending on the specific use case. Therefore, the choice of the forecasting method remains a mandatory expert task. However, expert knowledge cannot be fully automated. To establish a level playing field for evaluating the performance of time series forecasting methods in a broad setting, we propose Libra, a forecasting benchmark that automatically evaluates and ranks forecasting methods based on their performance in a diverse set of evaluation scenarios. The benchmark comprises four different use cases, each covering 100 heterogeneous time series taken from different domains. The data set was assembled from publicly available time series and was designed to exhibit much higher diversity than existing forecasting competitions. Based on this benchmark, we perform a comprehensive evaluation to compare different existing time series forecasting methods. André Bauer 0001, Marwin Züfle, Simon Eismann, Johannes Grohmann, Nikolas Herbst, Samuel Kounev |
ICPE | 5 |
| 2021 | SuanMing: Explainable Prediction of Performance Degradations in Microservice ApplicationsabstractApplication performance management (APM) tools are useful to observe the performance properties of an application during production. However, APM is normally purely reactive, that is, it can only report about current or past performance degradation. Although some approaches capable of predictive application monitoring have been proposed, they can only report a predicted degradation but cannot explain its root-cause, making it hard to prevent the expected degradation. Johannes Grohmann, Martin Sträßer, Avi Chalbani, Simon Eismann, Yair Arian, Nikolas Herbst, Noam Peretz, Samuel Kounev |
ICPE | 6 |
| 2021 | SARDE: A Framework for Continuous and Self-Adaptive Resource Demand EstimationabstractResource demands are crucial parameters for modeling and predicting the performance of software systems. Currently, resource demand estimators are usually executed once for system analysis. However, the monitored system, as well as the resource demand itself, are subject to constant change in runtime environments. These changes additionally impact the applicability, the required parametrization as well as the resulting accuracy of individual estimation approaches. Over time, this leads to invalid or outdated estimates, which in turn negatively influence the decision-making of adaptive systems. In this article, we present SARDE , a framework for self-adaptive resource demand estimation in continuous environments. SARDE dynamically and continuously tunes, selects, and executes an ensemble of resource demand estimation approaches to adapt to changes in the environment. This creates an autonomous and unsupervised ensemble estimation technique, providing reliable resource demand estimations in dynamic environments. We evaluate SARDE using two realistic datasets. One set of different micro-benchmarks reflecting different possible system states and one dataset consisting of a continuously running application in a changing environment. Our results show that by continuously applying online optimization, selection and estimation, SARDE is able to efficiently adapt to the online trace and reduce the model error using the resulting ensemble technique. Johannes Grohmann, Simon Eismann, André Bauer 0001, Simon Spinner, Johannes Blum 0001, Nikolas Herbst, Samuel Kounev |
ACM Trans. Auton. Adapt. Syst. | 6 |
| 2021 | Methodological Principles for Reproducible Performance Evaluation in Cloud ComputingabstractThe rapid adoption and the diversification of cloud computing technology exacerbate the importance of a sound experimental methodology for this domain. This work investigates how to measure and report performance in the cloud, and how well the cloud research community is already doing it. We propose a set of eight important methodological principles that combine best-practices from nearby fields with concepts applicable only to clouds, and with new ideas about the time-accuracy trade-off. We show how these principles are applicable using a practical use-case experiment. To this end, we analyze the ability of the newly released SPEC Cloud IaaS benchmark to follow the principles, and showcase real-world experimental studies in common cloud environments that meet the principles. Last, we report on a systematic literature review including top conferences and journals in the field, from 2012 to 2017, analyzing if the practice of reporting cloud performance measurements follows the proposed eight principles. Worryingly, this systematic survey and the subsequent two-round human reviews, reveal that few of the published studies follow the eight experimental principles. We conclude that, although these important principles are simple and basic, the cloud community is yet to adopt them broadly to deliver sound measurement of cloud environments. Alessandro Vittorio Papadopoulos, Laurens Versluis, André Bauer 0001, Nikolas Herbst, Jóakim von Kistowski, Ahmed Ali-Eldin, Cristina L. Abad, José Nelson Amaral, Petr Tuma 0001, Alexandru Iosup |
IEEE Trans. Software Eng. | 4 |
| 2020 | Telescope: An Automatic Feature Extraction and Transformation Approach for Time Series Forecasting on a Level-Playing FieldabstractOne central problem of machine learning is the inherent limitation to predict only what has been learned -stationarity. Any time series property that eludes stationarity poses a challenge for the proper model building. Furthermore, existing forecasting methods lack reliable forecast accuracy and time-to-result if not applied in their sweet spot. In this paper, we propose a fully automated machine learning-based forecasting approach. Our Telescope approach extracts and transforms features from an input time series and uses them to generate an optimized forecast model. In a broad competition including the latest hybrid forecasters, established statistical, and machine learning-based methods, our Telescope approach shows the best forecast accuracy coupled with a lower and reliable time-to-result. André Bauer 0001, Marwin Züfle, Nikolas Herbst, Samuel Kounev, Valentin Curtef |
ICDE | 3 |
| 2020 | An Automated Forecasting Framework based on Method Recommendation for Seasonal Time SeriesabstractDue to the fast-paced and changing demands of their users, computing systems require autonomic resource management. To enable proactive and accurate decision-making for changes causing a particular overhead, reliable forecasts are needed. In fact, choosing the best performing forecasting method for a given time series scenario is a crucial task. Taking the "No-Free-Lunch Theorem" into account, there exists no forecasting method that performs best on all types of time series. To this end, we propose an automated approach that (i) extracts characteristics from a given time series, (ii) selects the best-suited machine learning method based on recommendation, and finally, (iii) performs the forecast. Our approach offers the benefit of not relying on a single method with its possibly inaccurate forecasts. In an extensive evaluation, our approach achieves the best forecasting accuracy. André Bauer 0001, Marwin Züfle, Johannes Grohmann, Norbert Schmitt, Nikolas Herbst, Samuel Kounev |
ICPE | 5 |
| 2020 | Predicting the Costs of Serverless WorkflowsabstractFunction-as-a-Service (FaaS) platforms enable users to run arbitrary functions without being concerned about operational issues, while only paying for the consumed resources. Individual functions are often composed into workflows for complex tasks. However, the pay-per-use model and nontransparent reporting by cloud providers make it challenging to estimate the expected cost of a workflow, which prevents informed business decisions. Existing cost-estimation approaches assume a static response time for the serverless functions, without taking input parameters into account. In this paper, we propose a methodology for the cost prediction of serverless workflows consisting of input-parameter sensitive function models and a monte-carlo simulation of an abstract workflow model. Our approach enables workflow designers to predict, compare, and optimize the expected costs and performance of a planned workflow, which currently requires time-intensive experimentation. In our evaluation, we show that our approach can predict the response time and output parameters of a function based on its input parameters with an accuracy of 96.1%. In a case study with two audio-processing workflows, our approach predicts the costs of the two workflows with an accuracy of 96.2%. Simon Eismann, Johannes Grohmann, Erwin Van Eyk, Nikolas Herbst, Samuel Kounev |
ICPE | 4 |
| 2020 | 3rd Workshop on Hot Topics in Cloud Computing Performance (HotCloudPerf'20): Performance VariabilityabstractNo abstract available. Alexandru Uta, Dmitry Duplyakin, Cristina L. Abad, Nikolas Herbst, Alexandru Iosup |
ICPE | 4 |
| 2020 | Time Series Forecasting for Self-Aware SystemsabstractModern distributed systems and Internet-of-Things applications are governed by fast living and changing requirements. Moreover, they have to struggle with huge amounts of data that they create or have to process. To improve the self-awareness of such systems and enable proactive and autonomous decisions, reliable time series forecasting methods are required. However, selecting a suitable forecasting method for a given scenario is a challenging task. According to the “No-Free-Lunch Theorem,” there is no general forecasting method that always performs best. Thus, manual feature engineering remains to be a mandatory expert task to avoid trial and error. Furthermore, determining the expected time-to-result of existing forecasting methods is a challenge. In this article, we extensively assess the state-of-the-art in time series forecasting. We compare existing methods and discuss the issues that have to be addressed to enable their use in a self-aware computing context. To address these issues, we present a step-by-step approach to fully automate the feature engineering and forecasting process. Then, following the principles from benchmarking, we establish a level-playing field for evaluating the accuracy and time-to-result of automated forecasting methods for a broad set of application scenarios. We provide results of a benchmarking competition to guide in selecting and appropriately using existing forecasting methods for a given self-aware computing context. Finally, we present a case study in the area of self-aware data-center resource management to exemplify the benefits of fully automated learning and reasoning processes on time series data. André Bauer 0001, Marwin Züfle, Nikolas Herbst, Albin Zehe, Andreas Hotho, Samuel Kounev |
Proc. IEEE | 3 |
| 2019 | Kaa: Evaluating Elasticity of Cloud-Hosted DBMSabstractAuto-scaling is able to change the scale of an application at runtime. Understanding the application characteristics, scaling impact as well as the workload, an auto-scaler aligns the acquired resources to match the current workload. For distributed Database Management Systems (DBMS) forming the backend of many large-scale cloud applications, it is currently an open question to what extent they support scaling at run-time. In particular, elasticity properties of existing distributed DBMS are widely unknown and difficult to evaluate and compare. This paper presents a comprehensive methodology for the evaluation of the elasticity of distributed DBMS. On the basis of this methodology, we introduce a framework that automates the full evaluation process. We validate the framework by defining significant elasticity scenarios for a case study that comprises two DBMS for write-heavy and read-heavy workloads of different intensities. The results show that scalable distributed DBMS are not necessarily elastic and that adding more instances to a cluster at run-time may even decrease the experienced performance. Daniel Seybold, Simon Volpert, Stefan Wesner, André Bauer 0001, Nikolas Herbst, Jörg Domaschka |
CloudCom | 5 |
| 2019 | Chamulteon: Coordinated Auto-Scaling of Micro-ServicesabstractNowadays, in order to keep track of the fast changing requirements of Internet applications, auto-scaling is used as an essential mechanism for adapting the number of provisioned resources to the resource demand. The straightforward approach is to deploy a set of common and opensource single-service auto-scalers for each service independently. However, this deployment leads to problems such as bottleneck-shifting and increased oscillations. Existing auto-scalers that scale applications consisting of multiple services are kept closed-source. To face these challenges, we first survey existing auto-scalers and highlight current challenges. Then, we introduce Chamulteon, a redesign of our previously introduced mechanism, which can scale applications consisting of multiple services in a coordinated manner. We evaluate Chamulteon against four different well-cited auto-scalers in four sets of measurement-based experiments where we use diverse environments (VM vs. Docker), real-world traces, and vary the scale of the demanded resources. Overall, Chamulteon achieves the best auto-scaling performance based on established user-oriented and endorsed elasticity metrics. André Bauer 0001, Veronika Lesch, Laurens Versluis, Alexey Ilyushkin, Nikolas Herbst, Samuel Kounev |
ICDCS | 5 |
| 2019 | Chameleon: A Hybrid, Proactive Auto-Scaling Mechanism on a Level-Playing FieldabstractAuto-scalers for clouds promise stable service quality at low costs when facing changing workload intensity. The major public cloud providers provide trigger-based auto-scalers based on thresholds. However, trigger-based auto-scaling has reaction times in the order of minutes. Novel auto-scalers from literature try to overcome the limitations of reactive mechanisms by employing proactive prediction methods. However, the adoption of proactive auto-scalers in production is still very low due to the high risk of relying on a single proactive method. This paper tackles the challenge of reducing this risk by proposing a new hybrid auto-scaling mechanism, called Chameleon, combining multiple different proactive methods coupled with a reactive fallback mechanism. Chameleon employs on-demand, automated time series-based forecasting methods to predict the arriving load intensity in combination with run-time service demand estimation to calculate the required resource consumption per work unit without the need for application instrumentation. We benchmark Chameleon against five different state-of-the-art proactive and reactive auto-scalers one in three different private and public cloud environments. We generate five different representative workloads each taken from different real-world system traces. Overall, Chameleon achieves the best scaling behavior based on user and elasticity performance metrics, analyzing the results from 400 hours aggregated experiment time. André Bauer 0001, Nikolas Herbst, Simon Spinner, Ahmed Ali-Eldin, Samuel Kounev |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Using Machine Learning for Recommending Service Demand Estimation Approaches - Position Paper
Johannes Grohmann, Nikolas Herbst, Simon Spinner, Samuel Kounev |
CLOSER | 2 |
| 2018 | FOX: Cost-Awareness for Autonomic Resource Management in Public CloudsabstractNowadays, to keep track with the fast changing requirements of internet applications, auto-scaling is an essential mechanism for adapting the number of provisioned resources to the resource demand. In the context of public clouds, there exist different natures of cost-models for charging resources. However, the accounted resource units and charged resource units may differ significantly due to the applied cost model. This can lead to a significant increase of charged costs when using an auto-scaler as it tries to match the demand of the application as close as possible. In the literature, several auto-scalers exist that support cost-aware scaling decisions but they introduce inherent drawbacks. In this work, this lack of existing cost-aware mechanisms is addressed by introducing a mediator between an application and the auto-scaler. This cost-aware mechanism is called FOX. It leverages knowledge of the charging model of the public cloud and reviews the scaling decisions found by the auto-scaler to reduce the charged costs to a minimum. More precisely, FOX delays or omits releases of resources to avoid additional charging costs if the resource is required in the future. Hereby, FOX is not restricted to use one specific auto-scaler but offers interfaces to use any auto-scaler. For an evalation under controlled conditions, FOX scales a multi-tier application deployed in a private cloud that is stressed with two real world workloads: BibSonomy and IBM CICS. As FOX provides an interface for auto-scalers, we evaluate the cost-aware mechanism with three state of the art auto-scalers: React, Adapt, and Reg. The experiments show that FOX is able to reduce the charged costs by 34% at maximum for the Amazon EC2 charging model. According to the cost model, FOX provisions more resources than required. This results in a decreased SLO violation rate from 28% to 2% at maximum. The accounted instance time increases at max. by 30%. Veronika Lesch, André Bauer 0001, Nikolas Herbst, Samuel Kounev |
ICPE | 3 |
| 2017 | Design and Evaluation of a Proactive, Application-Aware Auto-Scaler: Tutorial PaperabstractSimple, threshold-based auto-scaling mechanisms as mainly used in practice bring no features to overcome resource provisioning delays and non-linear scalability of a software service. In this tutorial paper, we guide the reader step-by-step through the design and evaluation of a proactive and application-aware auto-scaling mechanism. André Bauer 0001, Nikolas Herbst, Samuel Kounev |
ICPE | 2 |
| 2017 | An Experimental Performance Evaluation of Autoscaling Policies for Complex WorkflowsabstractSimplifying the task of resource management and scheduling for customers, while still delivering complex Quality-of-Service (QoS), is key to cloud computing. Many autoscaling policies have been proposed in the past decade to decide on behalf of cloud customers when and how to provision resources to a cloud application utilizing cloud elasticity features. However, in prior work, when a new policy is proposed, it is seldom compared to the state-of-the-art, and is often compared only to static provisioning using a predefined QoS target. This reduces the ability of cloud customers and of cloud operators to choose and deploy an autoscaling policy. In our work, we conduct an experimental performance evaluation of autoscaling policies, using as application model workflows, a commonly used formalism for automating resource management for applications with well-defined yet complex structure. We present a detailed comparative study of general state-of-the-art autoscaling policies, along with two new workflow-specific policies. To understand the performance differences between the 7 policies, we conduct various forms of pairwise and group comparisons. We report both individual and aggregated metrics. Our results highlight the trade-offs between the suggested policies, and thus enable a better understanding of the current state-of-the-art. Alexey Ilyushkin, Ahmed Ali-Eldin, Nikolas Herbst, Alessandro Vittorio Papadopoulos, Bogdan Ghit, Dick H. J. Epema, Alexandru Iosup |
ICPE | 3 |
| 2017 | Modeling and Extracting Load Intensity ProfilesabstractToday’s system developers and operators face the challenge of creating software systems that make efficient use of dynamically allocated resources under highly variable and dynamic load profiles, while at the same time delivering reliable performance. Autonomic controllers, for example, an advanced autoscaling mechanism in a cloud computing context, can benefit from an abstracted load model as knowledge to reconfigure on time and precisely. Existing workload characterization approaches have limited support to capture variations in the interarrival times of incoming work units over time (i.e., a variable load profile). For example, industrial and scientific benchmarks support constant or stepwise increasing load, or interarrival times defined by statistical distributions or recorded traces. These options show shortcomings either in representative character of load variation patterns or in abstraction and flexibility of their format. In this article, we present the Descartes Load Intensity Model (DLIM) approach addressing these issues. DLIM provides a modeling formalism for describing load intensity variations over time. A DLIM instance is a compact formal description of a load intensity trace. DLIM-based tools provide features for benchmarking, performance, and recorded load intensity trace analysis. As manually obtaining and maintaining DLIM instances becomes time consuming, we contribute three automated extraction methods and devised metrics for comparison and method selection. We discuss how these features are used to enhance system management approaches for adaptations during runtime, and how they are integrated into simulation contexts and enable benchmarking of elastic or adaptive behavior. We show that automatically extracted DLIM instances exhibit an average modeling error of 15.2% over 10 different real-world traces that cover between 2 weeks and 7 months. These results underline DLIM model expressiveness. In terms of accuracy and processing speed, our proposed extraction methods for the descriptive models are comparable to existing time series decomposition methods. Additionally, we illustrate DLIM applicability by outlining approaches of workload modeling in systems engineering that employ or rely on our proposed load intensity modeling formalism. Jóakim von Kistowski, Nikolas Herbst, Samuel Kounev, Henning Groenda, Christian Stier, Sebastian Lehrig |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2016 | Which Cloud Auto-Scaler Should I Use for my Application?: Benchmarking Auto-Scaling AlgorithmsabstractRapid elasticity is one of the essential characteristics of cloud computing identified by NIST [17]. Elasticity allows resources to be provisioned and released to scale rapidly out ward and in ward according to demand. Tens -- if not hundreds -- of algorithms have been proposed in the literature to automatically achieve elastic provisioning [15, 23, 14, 21, 13, 20, 6, 12, 16, 10]. These algorithms are typically referred to as elasticity algorithms, dynamic provisioning techniques or autoscalers. While trying to solve the same problem, sometimes with differing assumption, many of these algorithms are either compared to static provisioning or to a predefined QoS target, e.g., predefined response time target, with very little -- or no -- comparison to previously published work. This reduces the ability of an application owner or a cloud operator to choose and deploy a suitable algorithm from the literature. Many of these algorithms have been tested with one single -- real or synthetic -- workload in a specific use-case [13, 14, 10]. While all published algorithms are shown to work in the specific use-case they were designed for with the, typically short, workloads tested with, it is seldom the case that the real scenarios will be any thing close to the test cases for which the algorithms are shown to work. Bursts occur in workloads occasionally. Workload dynamics change over time and the load-mix of an application significantly affects how provisioning should be done [21]. Ahmed Ali-Eldin, Alexey Ilyushkin, Bogdan Ghit, Nikolas Herbst, Alessandro Vittorio Papadopoulos, Alexandru Iosup |
ICPE | 4 |
| 2015 | Proactive Memory Scaling of Virtualized ApplicationsabstractEnterprise applications in virtualized environments are often subject to time-varying workloads with multiple seasonal patterns and trends. In order to ensure quality of service for such applications while avoiding over-provisioning, resources need to be dynamically adapted to accommodate the current workload demands. Many memory-intensive applications are not suitable for the traditional horizontal scaling approach often used for runtime performance management, as it relies on complex and expensive state replication. On the other hand, vertical scaling of memory often requires a restart of the application. In this paper, we propose a proactive approach to memory scaling for virtualized applications. It uses statistical forecasting to predict the future workload and reconfigure the memory size of the virtual machine of an application automatically. To this end, we propose an extended forecasting technique that leverages meta-knowledge, such as calendar information, to improve the forecast accuracy. In addition, we develop an application controller to adjust settings associated with application memory management during memory reconfiguration. Our evaluation using real-world traces shows that the forecast accuracy quantified with the MASE error metric can be improved by 11 - 59%. Furthermore, we demonstrate that the proactive approach can reduce the impact of reconfiguration on application availability by over 80% and significantly improve performance relative to a reactive controller. Simon Spinner, Nikolas Herbst, Samuel Kounev, Xiaoyun Zhu, Mustafa Uysal, Rean Griffith |
CLOUD | 2 |
| 2014 | LIMBO: a tool for modeling variable load intensitiesabstractModern software systems are expected to deliver reliable performance under highly variable load intensities while at the same time making efficient use of dynamically allocated resources. Conventional benchmarking frameworks provide limited support for emulating such highly variable and dynamic load profiles and workload scenarios. Industrial benchmarks typically use workloads with constant or stepwise increasing load intensity, or they simply replay recorded workload traces. In this paper, we present LIMBO - an Eclipse-based tool for modeling variable load intensity profiles based on the Descartes Load Intensity Model as an underlying modeling formalism. Jóakim von Kistowski, Nikolas Herbst, Samuel Kounev |
ICPE | 2 |
| 2014 | Self-adaptive workload classification and forecasting for proactive resource provisioningabstractSUMMARY As modern enterprise software systems become increasingly dynamic, workload forecasting techniques are gaining an importance as a foundation for online capacity planning and resource management. Time series analysis offers a broad spectrum of methods to calculate workload forecasts based on history monitoring data. Related work in the field of workload forecasting mostly concentrates on evaluating specific methods and their individual optimisation potential or on predicting QoS metrics directly. As a basis, we present a survey on established forecasting methods of the time series analysis concerning their benefits and drawbacks and group them according to their computational overheads. In this paper, we propose a novel self‐adaptive approach that selects suitable forecasting methods for a given context based on a decision tree and direct feedback cycles together with a corresponding implementation. The user needs to provide only his general forecasting objectives. In several experiments and case studies based on real‐world workload traces, we show that our implementation of the approach provides continuous and reliable forecast results at run‐time. The results of this extensive evaluation show that the relative error of the individual forecast points is significantly reduced compared with statically applied forecasting methods, for example, in an exemplary scenario on average by 37%. In a case study, between 55 and 75% of the violations of a given service level objective can be prevented by applying proactive resource provisioning based on the forecast results of our implementation. Copyright © 2014 John Wiley & Sons, Ltd. Nikolas Herbst, Nikolaus Huber, Samuel Kounev, Erich Amrehn |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Self-adaptive workload classification and forecasting for proactive resource provisioningabstractAs modern enterprise software systems become increasingly dynamic, workload forecasting techniques are gaining in importance as a foundation for online capacity planning and resource management. Time series analysis offers a broad spectrum of methods to calculate workload forecasts based on history monitoring data. Related work in the field of workload forecasting mostly concentrates on evaluating specific methods and their individual optimisation potential or on predicting Quality-of-Service (QoS) metrics directly. As a basis, we present a survey on established forecasting methods of the time series analysis concerning their benefits and drawbacks and group them according to their computational overheads. In this paper, we propose a novel self-adaptive approach that selects suitable forecasting methods for a given context based on a decision tree and direct feedback cycles together with a corresponding implementation. The user needs to provide only his general forecasting objectives. In several experiments and case studies based on real-world workload traces, we show that our implementation of the approach provides continuous and reliable forecast results at run-time. The results of this extensive evaluation show that the relative error of the individual forecast points is significantly reduced compared to statically applied forecasting methods, e.g. in an exemplary scenario on average by 37%. In a case study, between 55% and 75% of the violations of a given service level agreement can be prevented by applying proactive resource provisioning based on the forecast results of our implementation. Nikolas Herbst, Nikolaus Huber, Samuel Kounev, Erich Amrehn |
ICPE | 1 |