EDBT 2026 Demo / reviewers in the wild / expert
Philipp Leitner 0001
dblp:03/5268
· DBLP profile ↗
79ranked-venue papers
12as first author
21since 2021 · last 2026
0000-0003-2777-528XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 50 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Computer networks · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Let's trace it: Fine-grained serverless benchmarking for synchronous and asynchronous applicationsabstractMaking serverless computing widely applicable requires detailed understanding of performance. Although benchmarking approaches exist, their insights are coarse-grained and typically insufficient for (root cause) analysis of realistic serverless applications, which often consist of asynchronously coordinated functions and services. Addressing this gap, we design and implement ServiTrace, an approach for fine-grained distributed trace analysis and an application-level benchmarking suite for diverse serverless-application architectures. ServiTrace (i) analyzes distributed serverless traces using a novel algorithm and heuristics for extracting a detailed latency breakdown , (ii) leverages a suite of serverless applications representative of production usage, including synchronous and asynchronous serverless applications with external service integrations, and (iii) automates comprehensive, end-to-end experiments to capture application-level performance. Using our ServiTrace reference implementation, we conduct a large-scale empirical performance study in the market-leading AWS environment, collecting over 7.5 million execution traces. We make four main observations enabled by our latency breakdown analysis of median latency, cold starts, and tail latency for different application types and invocation patterns. For example, the median end-to-end latency of serverless applications is often dominated not by function computation but by external service calls, orchestration, and trigger-based coordination; all of which could be hidden without ServiTrace-like benchmarking. We release empirical data under FAIR principles and ServiTrace as a tested, extensible, open-source tool at https://github.com/ServiTrace/ReplicationPackage . Joel Scheuner, Simon Eismann, Sacheendra Talluri, Erwin Van Eyk, Cristina L. Abad, Philipp Leitner 0001, Alexandru Iosup |
Future Gener. Comput. Syst. | 6 |
| 2025 | A Brief Journey Through History: From Distributed Objects over SOA to Microservices
Philipp Leitner 0001 |
CLOSER | 1 |
| 2025 | BLAFS: A Bloat-Aware Container File SystemabstractContainers have become the standard for deploying applications in many cloud systems due to its convenience. However, this convenience leads to significant container bloat, i.e., unused files that inflate container image sizes, increase provisioning times, waste resources and introduce security vulnerabilities. Bloat is particularly problematic in serverless and edge computing scenarios, where resources are constrained, and performance is critical, and for microservice applications where rapid scaling is key to meet performance targets. However, existing container debloating tools are often limited in both effectiveness and robustness. In this paper, we propose BLAFS, a bloat-aware container filesystem that removes bloat while guaranteeing the correct operation of the debloated containers. BLAFS addresses bloat at the filesystem level by introducing new layers in the filesystem to enable debloating. During runtime, accessed files are moved to the debloating layers, and then similar to garbage collection mechanisms, BLAFS removes files that are not accessed during runtime. An optional reloading layer fetches files from a remote cloud cache on-demand if the files are mistakenly removed. We discuss how BLAFS can be used in different deployment scenarios and for different use-cases including container security-hardened and a dynamic deployment mode where the target is improved provisioning performance. We evaluate BLAFS performance using the top 20 downloaded containers from DockerHub, four ML containers, and SEBS, a Serverless Benchmark containing 10 serverless functions and compare its performance against two state-of-the-art debloating tools. Our evaluation shows that BLAFS reduces container sizes by up to 95% and cold-starts by up to 68%. In the security-hardened mode, BLAFS removes up to 89% of CVEs while the two state-of-the-art debloating tools fail on most of the workloads. We identify their limitations, and show how BLAFS provides a more principled approach to debloating. Additionally, when combined with lazy-loading snapshotters, BLAFS improves provisioning efficiency, reducing conversion times by up to 93% and provisioning times by up to 19%. Huaifeng Zhang, Mohannad Alhanahnah, Philipp Leitner 0001, Ahmed Ali-Eldin |
SoCC | 3 |
| 2025 | A Brief Journey Through History: From Distributed Objects Over SOA to Microservices
Philipp Leitner 0001 |
ENASE | 1 |
| 2025 | The Impact of Prompt Programming on Function-Level Code GenerationabstractLarge Language Models (LLMs) are increasingly used by software engineers for code generation. However, limitations of LLMs such as irrelevant or incorrect code have highlighted the need for prompt programming (or prompt engineering) where engineers apply specific prompt techniques (e.g., chain-of-thought or input-output examples) to improve the generated code. While some prompt techniques have been studied, the impact of different techniques—and their interactions— on code generation is still not fully understood. In this study, we introduce CodePromptEval, a dataset of 7072 prompts designed to evaluate five prompt techniques (few-shot, persona, chain-of-thought, function signature, list of packages) and their effect on the correctness, similarity, and quality of complete functions generated by three LLMs (GPT-4o, Llama3, and Mistral). Our findings show that while certain prompt techniques significantly influence the generated code, combining multiple techniques does not necessarily improve the outcome. Additionally, we observed a trade-off between correctness and quality when using prompt techniques. Our dataset and replication package enable future research on improving LLM-generated code and evaluating new prompt techniques. Ranim Khojah, Francisco Gomes de Oliveira Neto, Mazen Mohamad, Philipp Leitner 0001 |
IEEE Trans. Software Eng. | 4 |
| 2024 | The Impact of Compiler Warnings on Code Quality in C++ ProjectsabstractModern compilers often offer a variety of warning flags, which developers can enable to get feedback on code that, while syntactically correct, may be problematic. In the case of C++, one example of such "correct but problematic" code is code that leads to undefined behavior (UB). The usage of compiler warnings has long been suspected as a way to decrease bugs and increase code quality. However, empirical evidence that supports this hypothesis is rare. In this study, we present evidence from a study of 127 open source C++ projects. We categorize their usage of compiler warnings into five groups based on which warning flags are being used, and analyse the relationship between compiler warnings and five quality metrics (bugs, critical issues, vulnerabilities, code smells, and technical debt) using Bayesian analysis. We conclude that, in general, compiler warnings indeed correlate with, and potentially cause, higher code quality, with the clearest impact being on the number of critical issues in a project. Using stricter warning flags expectantly correlates with higher code quality in our study objects. However, there are substantial differences between projects, which we attribute to the project's individual development culture. That is, while warnings matter, other factors such as quality culture, are likely to be even more important to source code quality. Albin Johansson, Carl Holmberg, Francisco Gomes de Oliveira Neto, Philipp Leitner 0001 |
ICPC | 4 |
| 2024 | A unified active learning framework for annotating graph data for regression tasksabstractIn many domains, effectively applying machine learning models requires a large number of annotations and labelled data, which might not be available in advance. Acquiring annotations often requires significant time, effort, and computational resources, making it challenging. Active learning strategies are pivotal in addressing these challenges, particularly for diverse data types such as graphs. Although active learning has been extensively explored for node-level classification, its application to graph-level learning, especially for regression tasks, is not well-explored. We develop a unified active learning framework specializing in graph annotating and graph-level learning for regression tasks on both standard and expanded graphs, which are more detailed representations. We begin with graph collection and construction. Then, we construct various graph embeddings (unsupervised and supervised) into a latent space. Given such an embedding, the framework becomes task agnostic and active learning can be performed using any regression method and query strategy suited for regression. Within this framework, we investigate the impact of using different levels of information for active and passive learning, e.g., partially available labels and unlabelled test data. Despite our framework being domain agnostic, we validate it on a real-world application of software performance prediction, where the execution time of the source code is predicted. Thus, the graph is constructed as an intermediate source code representation. We support our methodology with a real-world dataset to underscore the applicability of our approach. Our real-world experiments reveal that satisfactory performance can be achieved by querying labels for only a small subset of all the data. A key finding is that Graph2Vec (an unsupervised embedding approach for graph data) performs the best, but only when all train and test features are used. However, Graph Neural Networks (GNNs) are the most flexible embedding techniques when used for different levels of information with and without label access. In addition, we find that the benefit of active learning increases for larger datasets (more graphs) and when the graphs are more complex, which is arguably when active learning is the most important. Peter Samoaa, Linus Aronsson, Antonio Longa, Philipp Leitner 0001, Morteza Haghir Chehreghani |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | An empirical investigation on the competences and roles of practitioners in Microservices-based ArchitecturesabstractMicroservices-based Architectures (MSAs) are gaining popularity since, among others, they enable rapid and independent delivery of software at scale, facilitating the delivery of business value. Additionally, there are attempts towards understanding practitioners’ roles and technical knowledge. MSAs call for affinity in several technologies as well as business domains. This diversity makes it challenging to scope and describe the roles of practitioners. In addition, practitioners often do not receive training and contents of MSA training remain largely undefined, even though there are challenges in finding or developing relevant technical expertise. In this research, we determine the different technical roles that are required in MSAs, along with their detailed competences. We use public online forums (e.g., StackOverflow), where developers share technical knowledge. We analyze 13,517 public profiles of software engineers, deriving their technical competences. Our taxonomy of technical competences in MSAs, contains 11 competences clusters, organized in 3 collections of competences — Web Technologies, DevOps, and Data Technologies. In addition, we derive the roles of microservice practitioners and the characteristics of their roles. Our findings organize the technical competences of MSAs practitioners and determine the training topics and combination of topics that can prepare engineers for MSAs. Hamdy Michael Ayas, Regina Hebig, Philipp Leitner 0001 |
J. Syst. Softw. | 3 |
| 2023 | Batch Mode Deep Active Learning for Regression on Graph DataabstractAcquiring labelled data for machine learning tasks, for example, for software performance prediction, remains a resource-intensive task. This study extends our previous work by introducing a batch-mode deep active learning approach tailored for regression in graph-structured data. Our framework leverages the source code conversion into Flow Augmented-AST graphs (FA-AST), subsequently utilizing both supervised and unsupervised graph embeddings. In contrast to single-instance querying, the batch-mode paradigm adaptively selects clusters of unlabeled data for labelling. We deploy an array of base kernels, kernel transformations, and selection methods, informed by both Bayesian and non-Bayesian strategies, to enhance the sample efficiency of neural network regression. Our experimental evaluation, conducted on multiple real-world software performance datasets, demonstrates the efficacy of the batch mode deep active learning approach in achieving robust performance with a reduced labelling budget. The methodology scales effectively to larger datasets and requires minimal alterations to existing neural network architectures. Peter Samoaa, Linus Aronsson, Philipp Leitner 0001, Morteza Haghir Chehreghani |
IEEE Big Data | 3 |
| 2023 | An empirical study of the systemic and technical migration towards microservicesabstractContext: As many organizations modernize their software architecture and transition to the cloud, migrations towards microservices become more popular. Even though such migrations help to achieve organizational agility and effectiveness in software development, they are also highly complex, long-running, and multi-faceted. Objective: In this study we aim to comprehensively map the journey towards microservices and describe in detail what such a migration entails. In particular, we aim to discuss not only the technical migration, but also the long-term journey of change, on a systemic level. Method: Our research method is an inductive, qualitative study on two data sources. Two main methodological steps take place - interviews and analysis of discussions from StackOverflow. The analysis of both, the 19 interviews and 215 StackOverflow discussions, is based on techniques found in grounded theory. Results: Our results depict the migration journey, as it materializes within the migrating organization, from structural changes to specific technical changes that take place in the work of engineers. We provide an overview of how microservices migrations take place as well as a deconstruction of high level modes of change to specific solution outcomes. Our theory contains 2 modes of change taking place in migration iterations, 14 activities and 53 solution outcomes of engineers. One of our findings is on the architectural change that is iterative and needs both a long and short term perspective, including both business and technical understanding. In addition, we found that a big proportion of the technical migration has to do with setting up supporting artifacts and changing the paradigm that software is developed. Hamdy Michael Ayas, Philipp Leitner 0001, Regina Hebig |
Empir. Softw. Eng. | 2 |
| 2023 | Using Microbenchmark Suites to Detect Application Performance ChangesabstractSoftware performance changes are costly and often hard to detect pre-release. Similar to software testing frameworks, either application benchmarks or microbenchmarks can be integrated into quality assurance pipelines to detect performance changes before releasing a new application version. Unfortunately, extensive benchmarking studies usually take several hours which is problematic when examining dozens of daily code changes in detail; hence, trade-offs have to be made. Optimized microbenchmark suites, which only include a small subset of the full suite, are a potential solution for this problem, given that they still reliably detect the majority of the application performance changes such as an increased request latency. It is, however, unclear whether microbenchmarks and application benchmarks detect the same performance problems and one can be a proxy for the other. In this paper, we explore whether microbenchmark suites can detect the same application performance changes as an application benchmark. For this, we run extensive benchmark experiments with both the complete and the optimized microbenchmark suites of two time-series database systems, i.e.,InfluxDBandVictoriaMetrics, and compare their results to the results of corresponding application benchmarks. We do this for 70 and 110 commits, respectively. Our results show that it is not trivial to detect application performance changes using an optimized microbenchmark suite. The detection (i) is only possible if the optimized microbenchmark suite covers all application-relevant code sections, (ii) is prone to false alarms, and (iii) cannot precisely quantify the impact on application performance. For certain software projects, an optimized microbenchmark suite can, thus, provide fast performance feedback to developers (e.g., as part of a local build process), help estimating the impact of code changes on application performance, and support a detailed analysis while a daily application benchmark detects major performance problems. Thus, although a regular application benchmark cannot be substituted for both studied systems, our results motivate further studies to validate and optimize microbenchmark suites. Martin Grambow, Denis Kovalev, Christoph Laaber, Philipp Leitner 0001, David Bermbach |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Automated Generation and Evaluation of JMH Microbenchmark Suites From Unit TestsabstractPerformance is a crucial non-functional requirement of many software systems. Despite the widespread use of performance testing, developers still struggle to construct and evaluate the quality of performance tests. To address these two major challenges, we implement a framework, dubbedju2jmh, to automatically generate performance microbenchmarks from JUnit tests and use mutation testing to study the quality of generated microbenchmarks. Specifically, we compare ourju2jmhgenerated benchmarks to manually written JMH benchmarks and to automatically generated JMH benchmarks using the AutoJMH framework, as well as directly measuring system performance with JUnit tests. For this purpose, we have conducted a study on three subjects (Rxjava,Eclipse-collections, andZipkin) with$\sim$454Ksource lines of code(SLOC), 2,417 JMH benchmarks (including manually written and generated AutoJMH benchmarks) and 35,084 JUnit tests. Our results show that theju2jmhgenerated JMH benchmarks consistently outperform using the execution time and throughput of JUnit tests as a proxy of performance and JMH benchmarks automatically generated using the AutoJMH framework while being comparable to JMH benchmarks manually written by developers in terms of tests’ stability and ability to detect performance bugs. Nevertheless,ju2jmhbenchmarks are able to cover more of the software applications than manually written JMH benchmarks during the microbenchmark execution. Furthermore,ju2jmhbenchmarks are generated automatically, while manually written JMH benchmarks require many hours of hard work and attention; therefore our study can reduce developers’ effort to construct microbenchmarks. In addition, we identify three factors (too low test workload, unstable tests and limited mutant coverage) that affect a benchmark's ability to detect performance bugs. To the best of our knowledge, this is the first study aimed at assisting developers in fully automated microbenchmark creation and assessing microbenchmark quality for performance testing. Mostafa Jangali, Yiming Tang 0002, Niclas Alexandersson, Philipp Leitner 0001, Jinqiu Yang 0001, Weiyi Shang |
IEEE Trans. Software Eng. | 4 |
| 2022 | An Empirical Analysis of Microservices Systems Using Consumer-Driven Contract TestingabstractTesting has a prominent role in revealing faults in software based on microservices. One of the most important discussion points in MSAs is the granularity of services, often in different levels of abstraction. Similarly, the granularity of tests in MSAs is reflected in different test types. However, it is challenging to conceptualize how the overall testing architecture comes together when combining testing in different levels of abstraction for microservices. There is no empirical evidence on the overall testing architecture in such microservices implementations. Furthermore, there is a need to empirically understand how the current state of practice resonates with existing best practices on testing. In this study, we mine Github to find different candidate projects for an in-depth, qualitative assessment of their test artifacts. We analyze 16 repositories that use microservices and include various test artifacts. We focus on four projects that use consumer-driven-contract testing. Our results demonstrate how these projects cover different levels of testing. This study (i) drafts a testing architecture including activities and artifacts, and (ii) demonstrates how these align with best practices and guidelines. Our proposed architecture helps the categorization of system and test artifacts in empirical studies of microservices. Finally, we showcase a view of the boundaries between different levels of testing in systems using microservices. Hamdy Michael Ayas, Hartmut Fischer, Philipp Leitner 0001, Francisco Gomes de Oliveira Neto |
SEAA | 3 |
| 2022 | TriggerBench: A Performance Benchmark for Serverless Function TriggersabstractServerless computing offers a scalable event-based paradigm for deploying managed cloud-native applications. Function triggers are essential building blocks in serverless, as they initiate any function execution. However, function triggering is insufficiently studied and inherently hard to measure given the distributed, ephemeral, and asynchronous nature of event-based function coordination. To address this gap, we present TriggerBench, a cross-provider benchmark for evaluating serverless function triggers based on distributed tracing. We evaluate the trigger latency (i.e., time to transition between two functions) of eight types of triggers in Microsoft Azure and three in AWS. Our results show that all triggers suffer from long tail latency, storage triggers introduce variable multi-second delays, and HTTP triggers are most suitable for interactive applications. Our insights can guide developers in choosing optimal event or messaging triggers for latency-sensitive applications. Researchers can extend TriggerBench to study the latency, scalability, and reliability of further trigger types and cloud providers. Joel Scheuner, Marcus Bertilsson, Oskar Grönqvist, Henrik Tao, Henrik Lagergren, Jan-Philipp Steghöfer, Philipp Leitner 0001 |
IC2E | 7 |
| 2022 | TEP-GNN: Accurate Execution Time Prediction of Functional Tests Using Graph Neural Networks
Hazem Peter Samoaa, Antonio Longa, Mazen Mohamad, Morteza Haghir Chehreghani, Philipp Leitner 0001 |
PROFES | 5 |
| 2022 | A systematic mapping study of source code representation for deep learning in software engineeringabstractAbstract The usage of deep learning (DL) approaches for software engineering has attracted much attention, particularly in source code modelling and analysis. However, in order to use DL, source code needs to be formatted to fit the expected input form of DL models. This problem is known as source code representation. Source code can be represented via different approaches, most importantly, the tree‐based, token‐based, and graph‐based approaches. We use a systematic mapping study to investigate i detail the representation approaches adopted in 103 studies that use DL in the context of software engineering. Thus, studies are collected from 2014 to 2021 from 14 different journals and 27 conferences. We show that each way of representing source code can provide a different, yet orthogonal view of the same source code. Thus, different software engineering tasks might require different (combinations of) code representation approaches, depending on the nature and complexity of the task. Particularly, we show that it is crucial to define whether the DL approach requires lexical, syntactical, or semantic code information. Our analysis shows that a wide range of different representations and combinations of representations (hybrid representations) are used to solve a wide range of common software engineering problems. However, we also observe that current research does not generally attempt to transfer existing representations or models to other studies even though there are other contexts in which these representations and models may also be useful. We believe that there is potential for more reuse and the application of transfer learning when applying DL to software engineering tasks. Hazem Peter Samoaa, Firas Bayram, Pasquale Salza, Philipp Leitner 0001 |
IET Softw. | 4 |
| 2021 | Facing the Giant: a Grounded Theory Study of Decision-Making in Microservices MigrationsabstractBackground: Microservices migrations are challenging and expensive projects with many decisions that need to be made in a multitude of dimensions. Existing research tends to focus on technical issues and decisions (e.g., how to split services). Equally important organizational or business issues and their relations with technical aspects often remain out of scope or on a high level of abstraction. Hamdy Michael Ayas, Philipp Leitner 0001, Regina Hebig |
ESEM | 2 |
| 2021 | The Migration Journey Towards Microservices
Hamdy Michael Ayas, Philipp Leitner 0001, Regina Hebig |
PROFES | 2 |
| 2021 | An Exploratory Study of the Impact of Parameterization on JMH Measurement Results in Open-Source ProjectsabstractThe Java Microbenchmarking Harness (JMH) is a widely used tool for testing performance-critical code on a low level. One of the key features of JMH is the support for user-defined parameters, which allows executing the same benchmark with different workloads. However, a benchmark configured with n parameters with m different values each requires JMH to execute the benchmark mn times (once for each combination of configured parameter values). Consequently, even fairly modest parameterization leads to a combinatorial explosion of benchmarks that have to be executed, hence dramatically increasing execution time. However, so far no research has investigated how this type of parameterization is used in practice, and how important different parameters are to benchmarking results. In this paper, we statistically study how strongly different user parameters impact benchmark measurements for 126 JMH benchmarks from five well-known open source projects. We show that 40% of the studied metric parameters have no correlation with the resulting measurement, i.e., testing with different values in these parameters does not lead to any insights. If there is a correlation, it is often strongly predictable following a power law, linear, or step function curve. Our results provide a first understanding of practical usage of user-defined JMH parameters, and how they correlate with the measurements produced by benchmarks. We further show that a machine learning model based on Random Forest ensembles can be used to predict the measured performance of an untested metric parameter value with an accuracy of 93% or higher for all but one benchmark class, demonstrating that given sufficient training data JMH performance test results for different parameterizations are highly predictable. Hazem Peter Samoaa, Philipp Leitner 0001 |
ICPE | 2 |
| 2021 | Applying test case prioritization to software microbenchmarksabstractAbstract Regression testing comprises techniques which are applied during software evolution to uncover faults effectively and efficiently. While regression testing is widely studied for functional tests, performance regression testing, e.g., with software microbenchmarks, is hardly investigated. Applying test case prioritization (TCP), a regression testing technique, to software microbenchmarks may help capturing large performance regressions sooner upon new versions. This may especially be beneficial for microbenchmark suites, because they take considerably longer to execute than unit test suites. However, it is unclear whether traditional unit testing TCP techniques work equally well for software microbenchmarks. In this paper, we empirically study coverage-based TCP techniques, employing total and additional greedy strategies, applied to software microbenchmarks along multiple parameterization dimensions, leading to 54 unique technique instantiations. We find that TCP techniques have a mean APFD-P (average percentage of fault-detection on performance) effectiveness between 0.54 and 0.71 and are able to capture the three largest performance changes after executing 29% to 66% of the whole microbenchmark suite. Our efficiency analysis reveals that the runtime overhead of TCP varies considerably depending on the exact parameterization. The most effective technique has an overhead of 11% of the total microbenchmark suite execution time, making TCP a viable option for performance regression testing. The results demonstrate that the total strategy is superior to the additional strategy. Finally, dynamic-coverage techniques should be favored over static-coverage techniques due to their acceptable analysis overhead; however, in settings where the time for prioritzation is limited, static-coverage techniques provide an attractive alternative. Christoph Laaber, Harald C. Gall, Philipp Leitner 0001 |
Empir. Softw. Eng. | 3 |
| 2021 | What's Wrong with My Benchmark Results? Studying Bad Practices in JMH BenchmarksabstractMicrobenchmarking frameworks, such as Java's Microbenchmark Harness (JMH), allow developers to write fine-grained performance test suites at the method or statement level. However, due to the complexities of the Java Virtual Machine, developers often struggle with writing expressive JMH benchmarks which accurately represent the performance of such methods or statements. In this paper, we empirically study bad practices of JMH benchmarks. We present a tool that leverages static analysis to identify 5 bad JMH practices. Our empirical study of 123 open source Java-based systems shows that each of these 5 bad practices are prevalent in open source software. Further, we conduct several experiments to quantify the impact of each bad practice in multiple case studies, and find that bad practices often significantly impact the benchmark results. To validate our experimental results, we constructed seven patches that fix the identified bad practices for six of the studied open source projects, of which six were merged into the main branch of the project. In this paper, we show that developers struggle with accurate Java microbenchmarking, and provide several recommendations to developers of microbenchmarking frameworks on how to improve future versions of their framework. Diego Costa 0001, Cor-Paul Bezemer, Philipp Leitner 0001, Artur Andrzejak 0001 |
IEEE Trans. Software Eng. | 3 |
| 2020 | Topology-Aware Continuous Experimentation in Microservice-Based Applications
Gerald Schermann, Fábio Oliveira, Erik Wittern, Philipp Leitner 0001 |
ICSOC | 4 |
| 2020 | An empirical study of bots in software development: characteristics and challenges from a practitioner's perspectiveabstractSoftware engineering bots – automated tools that handle tedious tasks – are increasingly used by industrial and open source projects to improve developer productivity. Current research in this area is held back by a lack of consensus of what software engineering bots (DevBots) actually are, what characteristics distinguish them from other tools, and what benefits and challenges are associated with DevBot usage. In this paper we report on a mixed-method empirical study of DevBot usage in industrial practice. We report on findings from interviewing 21 and surveying a total of 111 developers. We identify three different personas among DevBot users (focusing on autonomy, chat interfaces, and “smartness”), each with different definitions of what a DevBot is, why developers use them, and what they struggle with.We conclude that future DevBot research should situate their work within our framework, to clearly identify what type of bot the work targets, and what advantages practitioners can expect. Further, we find that there currently is a lack of general purpose “smart” bots that go beyond simple automation tools or chat interfaces. This is problematic, as we have seen that such bots, if available, can have a transformative effect on the projects that use them. Linda Erlenhov, Francisco Gomes de Oliveira Neto, Philipp Leitner 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2020 | Dynamically reconfiguring software microbenchmarks: reducing execution time without sacrificing result qualityabstractExecuting software microbenchmarks, a form of small-scale performance tests predominantly used for libraries and frameworks, is a costly endeavor. Full benchmark suites take up to multiple hours or days to execute, rendering frequent checks, e.g., as part of continuous integration (CI), infeasible. However, altering benchmark configurations to reduce execution time without considering the impact on result quality can lead to benchmark results that are not representative of the software’s true performance. Christoph Laaber, Stefan Würsten, Harald C. Gall, Philipp Leitner 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2020 | Function-as-a-Service performance evaluation: A multivocal literature reviewabstractFunction-as-a-Service (FaaS) is one form of the serverless cloud computing paradigm and is defined through FaaS platforms (e.g., AWS Lambda) executing event-triggered code snippets (i.e., functions). Many studies that empirically evaluate the performance of such FaaS platforms have started to appear but we are currently lacking a comprehensive understanding of the overall domain. To address this gap, we conducted a multivocal literature review (MLR) covering 112 studies from academic (51) and grey (61) literature. We find that existing work mainly studies the AWS Lambda platform and focuses on micro-benchmarks using simple functions to measure CPU speed and FaaS platform overhead (i.e., container cold starts). Further, we discover a mismatch between academic and industrial sources on tested platform configurations, find that function triggers remain insufficiently studied, and identify HTTP API gateways and cloud storages as the most used external service integrations. Following existing guidelines on experimentation in cloud systems, we discover many flaws threatening the reproducibility of experiments presented in the surveyed studies. We conclude with a discussion of gaps in literature and highlight methodological suggestions that may serve to improve future FaaS performance evaluation studies. Joel Scheuner, Philipp Leitner 0001 |
J. Syst. Softw. | 2 |
| 2019 | Interactive production performance feedback in the IDEabstractBecause of differences between development and production environments, many software performance problems are detected only after software enters production. We present PerformanceHat, a new system that uses profiling information from production executions to develop a global performance model suitable for integration into interactive development environments. PerformanceHat's ability to incrementally update this global model as the software is changed in the development environment enables it to deliver near real-time predictions of performance consequences reflecting the impact on the production environment. We implement PerformanceHat as an Eclipse plugin and evaluate it in a controlled experiment with 20 professional software developers implementing several software maintenance tasks using our approach and a representative baseline (Kibana). Our results indicate that developers using PerformanceHat were significantly faster in (1) detecting the performance problem, and (2) finding the root-cause of the problem. These results provide encouraging evidence that our approach helps developers detect, prevent, and debug production performance problems during development before the problem manifests in production. Jürgen Cito, Philipp Leitner 0001, Martin C. Rinard, Harald C. Gall |
ICSE | 2 |
| 2019 | Cachematic - Automatic Invalidation in Application-Level Caching SystemsabstractCaching is a common method for improving the performance of modern web applications. Due to the varying architecture of web applications, and the lack of a standardized approach to cache management, ad-hoc solutions are common. These solutions tend to be hard to maintain as a code base grows, and are a common source of bugs. We present Cachematic, a general purpose application-level caching system with an au- tomatic cache management strategy. Cachematic provides a simple programming model, allowing developers to explic- itly denote a function as cacheable. The result of a cacheable function will transparently be cached without the developer having to worry about cache management. We present algo- rithms that automatically handle cache management, han- dling the cache dependency tree, and cache invalidation. Our experiments showed that the deployment of Cachematic decreased response time for read requests, compared to a manual cache management strategy for a representative case study conducted in collaboration with Bison, an US-based business intelligence company. We also found that, com- pared to the manual strategy, the cache hit rate was in- creased with a factor of around 1.64x. However, we observe a significant increase in response time for write requests. We conclude that automatic cache management as implemented in Cachematic is attractive for read-domminant use cases, but the substantial write overhead in our current proof-of- concept implementation represents a challenge. Viktor Holmqvist, Jonathan Nilsfors, Philipp Leitner 0001 |
ICPE | 3 |
| 2019 | Software microbenchmarking in the cloud. How bad is it really?
Christoph Laaber, Joel Scheuner, Philipp Leitner 0001 |
Empir. Softw. Eng. | 3 |
| 2019 | A mixed-method empirical study of Function-as-a-Service software development in industrial practice
Philipp Leitner 0001, Erik Wittern, Josef Spillner, Waldemar Hummer |
J. Syst. Softw. | 1 |
| 2018 | Estimating Cloud Application Performance Based on Micro-Benchmark ProfilingabstractThe continuing growth of the cloud computing market has led to an unprecedented diversity of cloud services. To support service selection, micro-benchmarks are commonly used to identify the best performing cloud service. However, it remains unclear how relevant these synthetic micro-benchmarks are for gaining insights into the performance of real-world applications. Therefore, this paper develops a cloud benchmarking methodology that uses micro-benchmarks to profile applications and subsequently predicts how an application performs on a wide range of cloud services. A study with a real cloud provider (Amazon EC2) has been conducted to quantitatively evaluate the estimation model with 38 metrics from 23 micro-benchmarks and 2 applications from different domains. The results reveal remarkably low variability in cloud service performance and show that selected micro-benchmarks can estimate the duration of a scientific computing application with a relative error of less than 10% and the response time of a Web serving application with a relative error between 10% and 20%. In conclusion, this paper emphasizes the importance of cloud benchmarking by substantiating the suitability of micro-benchmarks for estimating application performance in comparison to common baselines but also highlights that only selected micro-benchmarks are relevant to estimate the performance of a particular application. Joel Scheuner, Philipp Leitner 0001 |
IEEE CLOUD | 2 |
| 2018 | Search-Based Scheduling of Experiments in Continuous DeploymentabstractContinuous experimentation involves practices for testing new functionality on a small fraction of the user base in production environments. Running multiple experiments in parallel requires handling user assignments (i.e., which users are part of which experiments) carefully as experiments might overlap and influence each other. Furthermore, experiments are prone to change, get canceled, or are adjusted and restarted, and new ones are added regularly. We formulate this as an optimization problem, fostering the parallel execution of experiments and making sure that enough data is collected for every experiment avoiding overlapping experiments. We propose a genetic algorithm that is capable of (re-)scheduling experiments and compare with other search-based approaches (random sampling, local search, and simulated annealing). Our evaluation shows that our genetic implementation outperforms the other approaches by up to 19% regarding the fitness of the solutions identified and up to a factor three in execution time in our evaluation scenarios. Gerald Schermann, Philipp Leitner 0001 |
ICSME | 2 |
| 2018 | An evaluation of open-source software microbenchmark suites for continuous performance assessmentabstractContinuous integration (CI) emphasizes quick feedback to developers. This is at odds with current practice of performance testing, which predominantely focuses on long-running tests against entire systems in production-like environments. Alternatively, software microbenchmarking attempts to establish a performance baseline for small code fragments in short time. This paper investigates the quality of microbenchmark suites with a focus on suitability to deliver quick performance feedback and CI integration. We study ten open-source libraries written in Java and Go with benchmark suite sizes ranging from 16 to 983 tests, and runtimes between 11 minutes and 8.75 hours. We show that our study subjects include benchmarks with result variability of 50% or higher, indicating that not all benchmarks are useful for reliable discovery of slowdowns. We further artificially inject actual slowdowns into public API methods of the study subjects and test whether test suites are able to discover them. We introduce a performance-test quality metric called the API benchmarking score (ABS). ABS represents a benchmark suite's ability to find slowdowns among a set of defined core API methods. Resulting benchmarking scores (i.e., fraction of discovered slowdowns) vary between 10% and 100% for the study subjects. This paper's methodology and results can be used to (1) assess the quality of existing microbenchmark suites, (2) select a set of tests to be run as part of CI, and (3) suggest or generate benchmarks for currently untested parts of an API. Christoph Laaber, Philipp Leitner 0001 |
MSR | 2 |
| 2018 | We're doing it live: A multi-method empirical study on continuous experimentation
Gerald Schermann, Jürgen Cito, Philipp Leitner 0001, Uwe Zdun, Harald C. Gall |
Inf. Softw. Technol. | 3 |
| 2017 | An Approach and Case Study of Cloud Instance Type Selection for Multi-Tier Web ApplicationsabstractA challenging problem for users of Infrastructure-as-a-Service (IaaS) clouds is selecting cloud providers, regions, and instance types cost-optimally for a given desired service level. Issues such as hardware heterogeneity, contention, and virtual machine (VM) placement can result in considerably differing performance across supposedly equivalent cloud resources. Existing research on cloud benchmarking helps, but often the focus is on providing low-level microbenchmarks (e.g., CPU or network speed), which are hard to map to concrete business metrics of enterprise cloud applications, such as request throughput of a multi-tier Web application. In this paper, we propose Okta, a general approach for fairly and comprehensively benchmarking the performance and cost of a multi-tier Web application hosted in an IaaS cloud. We exemplify our approach for a case study based on the two-tier AcmeAir application, which we evaluate for 11 real-life deployment configurations on Amazon EC2 and Google Compute Engine. Our results show that for this application, choosing compute-optimized instance types in the Web layer and small bursting instances for the database tier leads to the overall most cost-effective deployments. This result held true for both cloud providers. The least cost-effective configuration in our study provides only about 67% of throughput per US dollar spent. Our case study can serve as a blueprint for future industrial or academic application benchmarking projects. Christian Davatz, Christian Inzinger, Joel Scheuner, Philipp Leitner 0001 |
CCGrid | 4 |
| 2017 | A Tale of CI Build Failures: An Open Source and a Financial Organization PerspectiveabstractContinuous Integration (CI) and Continuous Delivery (CD) are widespread in both industrial and open-source software (OSS) projects. Recent research characterized build failures in CI and identified factors potentially correlated to them. However, most observations and findings of previous work are exclusively based on OSS projects or data from a single industrial organization. This paper provides a first attempt to compare the CI processes and occurrences of build failures in 349 Java OSS projects and 418 projects from a financial organization, ING Nederland. Through the analysis of 34,182 failing builds (26% of the total number of observed builds), we derived a taxonomy of failures that affect the observed CI processes. Using cluster analysis, we observed that in some cases OSS and ING projects share similar build failure patterns (e.g., few compilation failures as compared to frequent testing failures), while in other cases completely different patterns emerge. In short, we explain how OSS and ING CI processes exhibit commonalities, yet are substantially different in their design and in the failures they report. Carmine Vassallo, Gerald Schermann, Fiorella Zampetti, Daniele Romano, Philipp Leitner 0001, Andy Zaidman, Massimiliano Di Penta, Sebastiano Panichella |
ICSME | 5 |
| 2017 | Extraction of Microservices from Monolithic Software ArchitecturesabstractDriven by developments such as mobile computing, cloud computing infrastructure, DevOps and elastic computing, the microservice architectural style has emerged as a new alternative to the monolithic style for designing large software systems. Monolithic legacy applications in industry undergo a migration to microservice-oriented architectures. A key challenge in this context is the extraction of microservices from existing monolithic code bases. While informal migration patterns and techniques exist, there is a lack of formal models and automated support tools in that area. This paper tackles that challenge by presenting a formal microservice extraction model to allow algorithmic recommendation of microservice candidates in a refactoring and migration scenario. The formal model is implemented in a web-based prototype. A performance evaluation demonstrates that the presented approach provides adequate performance. The recommendation quality is evaluated quantitatively by custom microservice-specific metrics. The results show that the produced microservice candidates lower the average development team size down to half of the original size or lower. Furthermore, the size of recommended microservice conforms with microservice sizing reported by empirical surveys and the domain-specific redundancy among different microservices is kept at a low rate. Genc Mazlami, Jürgen Cito, Philipp Leitner 0001 |
ICWS | 3 |
| 2017 | An empirical analysis of the docker container ecosystem on GitHubabstractDocker allows packaging an application with its dependencies into a standardized, self-contained unit (a so-called container), which can be used for software development and to run the application on any system. Dockerfiles are declarative definitions of an environment that aim to enable reproducible builds of the container. They can often be found in source code repositories and enable the hosted software to come to life in its execution environment. We conduct an exploratory empirical study with the goal of characterizing the Docker ecosystem, prevalent quality issues, and the evolution of Dockerfiles. We base our study on a data set of over 70000 Dockerfiles, and contrast this general population with samplings that contain the Top-100 and Top-1000 most popular Docker-using projects. We find that most quality issues (28.6%) arise from missing version pinning (i.e., specifying a concrete version for dependencies). Further, we were not able to build 34% of Dockerfiles from a representative sample of 560 projects. Integrating quality checks, e.g., to issue version pinning warnings, into the container build process could result into more reproducible builds. The most popular projects change more often than the rest of the Docker population, with 5.81 revisions per year and 5 lines of code changed on average. Most changes deal with dependencies, that are currently stored in a rather unstructured manner. We propose to introduce an abstraction that, for instance, could deal with the intricacies of different package managers and could improve migration to more light-weight images. Jürgen Cito, Gerald Schermann, Erik Wittern, Philipp Leitner 0001, Sali Zumberi, Harald C. Gall |
MSR | 4 |
| 2017 | An empirical analysis of build failures in the continuous integration workflows of Java-based open-source softwareabstractContinuous Integration (CI) has become a common practice in both industrial and open-source software development. While CI has evidently improved aspects of the software development process, errors during CI builds pose a threat to development efficiency. As an increasing amount of time goes into fixing such errors, failing builds can significantly impair the development process and become very costly. We perform an indepth analysis of build failures in CI environments. Our approach links repository commits to data of corresponding CI builds. Using data from 14 open-source Java projects, we first identify 14 common error categories. Besides test failures, which are by far the most common error category (up to >80% per project), we also identify noisy build data, e.g., induced by transient Git interaction errors, or general infrastructure flakiness. Second, we analyze which factors impact the build results, taking into account general process and specific CI metrics. Our results indicate that process metrics have a significant impact on the build outcome in 8 of the 14 projects on average, but the strongest influencing factor across all projects is overall stability in the recent build history. For 10 projects, more than 50% (up to 80%) of all failed builds follow a previous build failure. Moreover, the fail ratio of the last k=10 builds has a significant impact on build results for all projects in our dataset. Thomas Rausch, Waldemar Hummer, Philipp Leitner 0001, Stefan Schulte 0002 |
MSR | 3 |
| 2017 | (h|g)opper: Performance History Mining and AnalysisabstractPerformance changes of software systems, and especially performance regressions, have a tremendous impact on users of that system. Historical data can help developers to reason about how performance has changed over the course of a software's lifetime. In this demo paper we present two tools: hopper to mine historical performance metrics based on benchmarks and unit tests, and gopper to analyse the data with respect to performance changes. Christoph Laaber, Philipp Leitner 0001 |
ICPE | 2 |
| 2017 | An Exploratory Study of the State of Practice of Performance Testing in Java-Based Open Source ProjectsabstractThe usage of open source (OS) software is wide-spread across many industries. While the functional quality of OS projects is considered to be similar to closed-source software, much is unknown about the quality in terms of performance. One challenge for OS developers is that, unlike for functional testing, there is a lack of accepted best practices for performance testing. To reveal the state of practice of performance testing in OS projects, we conduct an exploratory study on 111 Java-based OS projects from GitHub. We study the performance tests of these projects from five perspectives: (1) developers, (2) size, (3) test organization, (4) types of performance tests and (5) used tooling. We show that writing performance tests is not a popular task in OS projects: performance tests form only a small portion of the test suite, are rarely updated, and are usually maintained by a small group of core project developers. Further, even though many projects are aware that they need performance tests, developers appear to struggle implementing them. We argue that future performance testing frameworks should provider better support for low-friction testing, for instance via non-parameterized methods or performance test generation, as well as focus on a tight integration with standard continuous integration tooling. Philipp Leitner 0001, Cor-Paul Bezemer |
ICPE | 1 |
| 2017 | Optimized IoT service placement in the fogabstractThe Internet of Things (IoT) leads to an ever-growing presence of ubiquitous networked computing devices in public, business, and private spaces. These devices do not simply act as sensors, but feature computational, storage, and networking resources. Being located at the edge of the network, these resources can be exploited to execute IoT applications in a distributed manner. This concept is known as fog computing. While the theoretical foundations of fog computing are already established, there is a lack of resource provisioning approaches to enable the exploitation of fog-based computational resources. To resolve this shortcoming, we present a conceptual fog computing framework. Then, we model the service placement problem for IoT applications over fog resources as an optimization problem, which explicitly considers the heterogeneity of applications and resources in terms of Quality of Service attributes. Finally, we propose a genetic algorithm as a problem resolution heuristic and show, through experiments, that the service execution can achieve a reduction of network communication delays when the genetic algorithm is used, and a better utilization of fog resources when the exact optimization method is applied. Olena Skarlat, Matteo Nardelli 0001, Stefan Schulte 0002, Michael Borkowski, Philipp Leitner 0001 |
Serv. Oriented Comput. Appl. | 5 |
| 2016 | Towards quality gates in continuous delivery and deploymentabstractQuality gates, steps required to ensure the reliability of code changes, are supposed to increase the confidence stakeholders have in a release. In today's fast paced environments, we have less time to perform the necessary precautions to minimize the risk of a faulty release. This leads to an inherent trade-off between risk of lower release quality and time to market. We provide a model for this trade-off of release “confidence” and “velocity” that led to the formulation of 4 categories (cautious, balanced, problematic, madness), in which companies can be classified in. We showcase real examples of these categories as case studies based on previous empirical studies. We close by presenting possible transitions between categories that guide future research. Gerald Schermann, Jürgen Cito, Philipp Leitner 0001, Harald C. Gall |
ICPC | 3 |
| 2016 | Bifrost: Supporting Continuous Deployment with Automated Enactment of Multi-Phase Live Testing Strategies
Gerald Schermann, Dominik Schöni, Philipp Leitner 0001, Harald C. Gall |
Middleware | 3 |
| 2016 | Patterns in the Chaos - A Study of Performance Variation and Predictability in Public IaaS CloudsabstractBenchmarking the performance of public cloud providers is a common research topic. Previous work has already extensively evaluated the performance of different cloud platforms for different use cases, and under different constraints and experiment setups. In this article, we present a principled, large-scale literature review to collect and codify existing research regarding the predictability of performance in public Infrastructure-as-a-Service (IaaS) clouds. We formulate 15 hypotheses relating to the nature of performance variations in IaaS systems, to the factors of influence of performance variations, and how to compare different instance types. In a second step, we conduct extensive real-life experimentation on four cloud providers to empirically validate those hypotheses. We show that there are substantial differences between providers. Hardware heterogeneity is today less prevalent than reported in earlier research, while multitenancy has a dramatic impact on performance and predictability, but only for some cloud providers. We were unable to discover a clear impact of the time of the day or the day of the week on cloud performance. Philipp Leitner 0001, Jürgen Cito |
ACM Trans. Internet Techn. | 1 |
| 2015 | Discovering loners and phantoms in commit and issue dataabstractThe interlinking of commit and issue data has become a de-facto standard in software development. Modern issue tracking systems, such as JIRA, automatically interlink commits and issues by the extraction of identifiers (e.g., Issue key) from commit messages. However, the conventions for the use of interlinking methodologies vary between software projects. For example, some projects enforce the use of identifiers for every commit while others have less restrictive conventions. In this work, we introduce a model called PaLiMod to enable the analysis of interlinking characteristics in commit and issue data. We surveyed 15 Apache projects to investigate differences and commonalities between linked and non-linked commits and issues. Based on the gathered information, we created a set of heuristics to interlink the residual of non-linked commits and issues. We present the characteristics of Loners and Phantoms in commit and issue data. The results of our evaluation indicate that the proposed PaLiMod model and heuristics enable an automatic interlinking and can indeed reduce the residual of non-linked commits and issues in software projects. Gerald Schermann, Martin Brandtner, Sebastiano Panichella, Philipp Leitner 0001, Harald C. Gall |
ICPC | 4 |
| 2015 | Intent, tests, and release dependencies: Pragmatic recipes for source code integrationabstractContinuous integration of source code changes, for example, via pull-request driven contribution channels, has become standard in many software projects. However, the decision to integrate source code changes into a release is complex and has to be taken by a software manager. In this work, we identify a set of three pragmatic recipes plus variations to support the decision making of integrating code contributions into a release. These recipes cover the isolation of source code changes, contribution of test code, and the linking of commits to issues. We analyze the development history of 21 open-source software projects, to evaluate whether, and to what extent, those recipes are followed in open-source projects. The results of our analysis showed that open-source projects largely follow recipes on a compliance level of > 75%. Hence, we conclude that the identified recipes plus variations can be seen as wide-spread relevant best-practices for source code integration. Martin Brandtner, Philipp Leitner 0001, Harald C. Gall |
SCAM | 2 |
| 2015 | SPEEDL - A Declarative Event-Based Language to Define the Scaling Behavior of Cloud ApplicationsabstractContemporary cloud providers offer out-of-the-box auto-scaling solutions. However, defining a non-trivial scaling behavior that goes beyond the feature set provided by existing solutions is still challenging. In this paper we present SPEEDL, a declarative and extensible domain-specific language that simplifies the creation of elastic scaling behavior on top of IaaS clouds. SPEEDL simplifies the creation of event-driven policies for resource management (How many resources, and what resource types, are needed?), as well as task mapping (Which tasks should be handled by which resources?). Based on a dataset of real-life scaling policies, we demonstrate that SPEEDL can cover most scaling behaviors real-life developers want to express, and that the resulting SPEEDL policies are at the same time substantially more compact, easier to read, and less error-prone than the same behavior expressed via a general-purpose programming language. Rostyslav Zabolotnyi, Philipp Leitner 0001, Stefan Schulte 0002, Schahram Dustdar |
SERVICES | 2 |
| 2015 | The making of cloud applications: an empirical study on software development for the cloudabstractCloud computing is gaining more and more traction as a deployment and provisioning model for software. While a large body of research already covers how to optimally operate a cloud system, we still lack insights into how professional software engineers actually use clouds, and how the cloud impacts development practices. This paper reports on the first systematic study on how software developers build applications for the cloud. We conducted a mixed-method study, consisting of qualitative interviews of 25 professional developers and a quantitative survey with 294 responses. Our results show that adopting the cloud has a profound impact throughout the software development process, as well as on how developers utilize tools and data in their daily work. Among other things, we found that (1) developers need better means to anticipate runtime problems and rigorously define metrics for improved fault localization and (2) the cloud offers an abundance of operational data, however, developers still often rely on their experience and intuition rather than utilizing metrics. From our findings, we extracted a set of guidelines for cloud development and identified challenges for researchers and tool vendors. Jürgen Cito, Philipp Leitner 0001, Thomas Fritz 0001, Harald C. Gall |
ESEC/SIGSOFT FSE | 2 |
| 2015 | SQA-Profiles: Rule-based activity profiles for Continuous Integration environmentsabstractContinuous Integration (CI) environments cope with the repeated integration of source code changes and provide rapid feedback about the status of a software project. However, as the integration cycles become shorter, the amount of data increases, and the effort to find information in CI environments becomes substantial. In modern CI environments, the selection of measurements (e.g., build status, quality metrics) listed in a dashboard does only change with the intervention of a stakeholder (e.g., a project manager). In this paper, we want to address the shortcoming of static views with so-called Software Quality Assessment (SQA) profiles. SQA-Profiles are defined as rule-sets and enable a dynamic composition of CI dashboards based on stakeholder activities in tools of a CI environment (e.g., version control system). We present a set of SQA-Profiles for project management committee (PMC) members: Bandleader, Integrator, Gatekeeper, and Onlooker. For this, we mined the commit and issue management activities of PMC members from 20 Apache projects. We implemented a framework to evaluate the performance of our rule-based SQA-Profiles in comparison to a machine learning approach. The results showed that project-independent SQA-Profiles can be used to automatically extract the profiles of PMC members with a precision of 0.92 and a recall of 0.78. Martin Brandtner, Sebastian C. Müller, Philipp Leitner 0001, Harald C. Gall |
SANER | 3 |
| 2015 | Identifying Web Performance Degradations through Synthetic and Real-User Monitoring
Jürgen Cito, Devan Gotowka, Philipp Leitner 0001, Ryan Pelette, Dritan Suljoti, Schahram Dustdar |
J. Web Eng. | 3 |
| 2015 | JCloudScale: Closing the Gap Between IaaS and PaaSabstractBuilding Infrastructure-as-a-Service (IaaS) applications today is a complex, repetitive, and error-prone endeavor, as IaaS does not provide abstractions on top of virtual machines. This article presents JC loud S cale , a Java-based middleware for moving elastic applications to IaaS clouds, with minimal adjustments to the application code. We discuss the architecture and technical features, as well as evaluate our system with regard to user acceptance and performance overhead. Our user study reveals that JC loud S cale allows many participants to build IaaS applications more efficiently, compared to industrial Platform-as-a-Service (PaaS) solutions. Additionally, unlike PaaS, JC loud S cale does not lead to a control loss and vendor lock-in. Rostyslav Zabolotnyi, Philipp Leitner 0001, Waldemar Hummer, Schahram Dustdar |
ACM Trans. Internet Techn. | 2 |
| 2015 | Comparing and Combining Predictive Business Process Monitoring TechniquesabstractPredictive business process monitoring aims at forecasting potential problems during process execution before they occur so that these problems can be handled proactively. Several predictive monitoring techniques have been proposed in the past. However, so far those prediction techniques have been assessed only independently from each other, making it hard to reliably compare their applicability and accuracy. We empirically analyze and compare three main classes of predictive monitoring techniques, which are based on machine learning, constraint satisfaction, and Quality-of-Service (QoS) aggregation. Based on empirical evidence from an industrial case study in the area of transport and logistics, we assess those techniques with respect to five accuracy indicators. We further determine the dependency of accuracy on the point in time during process execution when a prediction is made in order to determine lead-times for accurate predictions. Our evidence suggests that, given a lead-time of half of the process duration, all predictive monitoring techniques consistently provide an accuracy of at least 70%. Yet, it also becomes evident that the techniques differ in terms of how accurately they may predict violations and nonviolations. To improve the prediction process, we thus exploit the characteristics of the individual techniques and propose their combination. Based on our case study data, evidence indicates that certain combinations of techniques may outperform individual techniques with respect to specific accuracy indicators. Combining constraint satisfaction with QoS aggregation, for instance, improves precision by 14%; combining machine learning with constraint satisfaction shows an improvement in recall by 23%. Andreas Metzger, Philipp Leitner 0001, Dragan Ivanovic, Eric Schmieders, Rod Franklin, Manuel Carro, Schahram Dustdar, Klaus Pohl |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2014 | Cloud Work Bench - Infrastructure-as-Code Based Cloud BenchmarkingabstractIn order to optimally deploy their applications, users of Infrastructure-as-a-Service clouds are required to evaluate the costs and performance of different combinations of cloud configurations to find out which combination provides the best service level for their specific application. Unfortunately, benchmarking cloud services is cumbersome and error-prone. In this paper, we propose an architecture and concrete implementation of a cloud benchmarking Web service, which fosters the definition of reusable and representative benchmarks. In distinction to existing work, our system is based on the notion of Infrastructure-as-Code, which is a state of the art concept to define IT infrastructure in a reproducible, well-defined, and testable way. Joel Scheuner, Philipp Leitner 0001, Jürgen Cito, Harald C. Gall |
CloudCom | 2 |
| 2014 | WPress: An Application-Driven Performance Benchmark for Cloud-Based Virtual MachinesabstractApproaching a comprehensive performance benchmark for on-line transaction processing (OLTP) applications in a cloud environment is a challenging task. Fundamental features of clouds, such as the pay-as-you-go pricing model and unknown underlying configuration of the system, are contrary to the basic assumptions of available benchmarks such as TPC-W or RUBiS. In this paper, we introduce a systematic performance benchmark approach for OLTP applications on public clouds that use virtual machines(VMs). We propose WPress benchmark, which is based on the widespread blogging software, WordPress, as a representative OLTP application and implement an open source workload generator. Furthermore, we utilize a CPU micro-benchmark to investigate CPU performance of cloud-based VMs in greater detail. Average response time and total VM cost are the performance metrics measured by WPress. We evaluate small and large instance types of three real-life cloud providers, Amazon EC2, Microsoft Azure and Rackspace cloud. Results imply that Rackspace cloud has better average response times and total VM cost on small instances. However, Microsoft Azure is preferable for large instance type. Amir Hossein Borhani, Philipp Leitner 0001, Bu-Sung Lee, Xiaorong Li, Terence Hung |
EDOC | 2 |
| 2014 | Identifying Root Causes of Web Performance Degradation Using Changepoint Analysis
Jürgen Cito, Dritan Suljoti, Philipp Leitner 0001, Schahram Dustdar |
ICWE | 3 |
| 2014 | CloudWave: Where adaptive cloud management meets DevOpsabstractThe transition to cloud computing offers a large number of benefits, such as lower capital costs and a highly agile environment. Yet, the development of software engineering practices has not kept pace with this change. Moreover, the design and runtime behavior of cloud based services and the underlying cloud infrastructure are largely decoupled from one another.This paper describes the innovative concepts being developed by CloudWave to utilize the principles of DevOps to create an execution analytics cloud infrastructure where, through the use of programmable monitoring and online data abstraction, much more relevant information for the optimization of the ecosystem is obtained. Required optimizations are subsequently negotiated between the applications and the cloud infrastructure to obtain coordinated adaption of the ecosystem. Additionally, the project is developing the technology for a Feedback Driven Development Standard Development Kit which will utilize the data gathered through execution analytics to supply developers with a powerful mechanism to shorten application development cycles. Dario Bruneo, Thomas Fritz 0001, Sharon Barner, Philipp Leitner 0001, Francesco Longo 0001, Clarissa Cassales Marquezan, Andreas Metzger, Klaus Pohl, Antonio Puliafito, Danny Raz, Andreas Roth 0001, Eliot E. Salant, Itai Segall, Massimo Villari, Yaron Wolfsthal, Chris Woods |
ISCC | 4 |
| 2014 | Generic event-based monitoring and adaptation methodology for heterogeneous distributed systemsabstractSUMMARY The Cloud computing paradigm provides the basis for a class of platforms and applications that face novel challenges related to multi‐tenancy, adaptivity, and elasticity. To account for service delivery guarantees in the face of ever increasing levels of heterogeneity, scale, and dynamism, service provisioning in the Cloud has raised the demand for systematic and flexible approaches to monitoring and adaptation of applications. In this paper, we tackle this issue and present a framework for efficient runtime management of Cloud environments and distributed heterogeneous systems in general. A novel domain‐specific language termed MONINA is introduced that allows to define integrated monitoring and adaptation functionality for controlling such systems. We propose a mechanism for optimal deployment of the defined control operators onto available computing resources. Deployment is based on solving a quadratic programming problem, which aims at achieving minimized reaction times, low overhead, and scalable monitoring and adaptation. The monitoring infrastructure is based on a distributed messaging middleware, providing high level of decoupling and allowing new monitoring nodes to join the system dynamically. We provide a detailed formalization of the problem domain, discuss architectural details, highlight the implementation of the developed prototype, and put our work into perspective with existing work in the field. Copyright © 2014 John Wiley & Sons, Ltd. Christian Inzinger, Waldemar Hummer, Benjamin Satzger, Philipp Leitner 0001, Schahram Dustdar |
Softw. Pract. Exp. | 4 |
| 2014 | A note on software tools and techniques for monitoring and prediction of cloud servicesabstractCloud computing is the latest computing paradigm that transparently delivers Information and Communication Technology resources as services, freeing the users of Cloud applications from dealing with low-level implementation and system administration details. Cloud provides the promise of on-demand access to affordable large-scale computing (e.g., multi-core CPUs, GPUs, and clusters of GPUs), storage (such as disks), and software (e.g., databases, application servers, and data processing frameworks) resources without substantial up-front investment. Cloud resources are hosted in large datacenters, often referred to as virtualized data farms, operated by companies such as Amazon, Apple, GoGrid, and Microsoft. While the growing ubiquity of Cloud computing is having a significant impact in many applications domains, there are still significant problems that exist with regard to efficient provisioning and delivery of applications using its Information and Communication Technology resources. These barriers are due to resource uncertainties 1 that have degradable effect on the run-time Quality of Service (e.g., access latency and number of requests being successfully served per second) of software applications deployed in the Cloud. There are many reasons for such uncertainties including (i) unpredictable application workload types (enterprise, scientific, and streaming big data analytics), (ii) fluctuations in resource capacity demands (i.e., bandwidth, memory, storage, and CPU), (iii) abrupt failures (e.g., failure of a network link), (iv) stochastic access patterns (e.g., number of end-users and their geo-location), (v) heterogeneity in device types (e.g., mobile phone, laptop, and smart TV), (v) heterogeneous resource types and their providers, and (vi) heterogeneity in data types (3D images, videos, audios, text, etc.) and network types (e.g., wired and wireless). These Cloud resource uncertainties need to be managed optimally to maintain contractual requirements defined in Service-Level Agreements (SLAs) that underlie most Cloud computing contracts. Basically, SLAs are legal documents (paper and/or electronic) that encode the nature and scope of QoS parameters (e.g., ensure availability 99.99% and ensure web application server latency to be less than 100 ms). To tackle uncertainties, recent research and industry efforts 2 have focused on developing monitoring techniques and frameworks that can assist cloud providers and application owners in (i) keeping their resources and applications operating at peak efficiency, (ii) detecting variations in resource and application performance, (iii) accounting the SLA violations of certain QoS parameters, and (iv) tracking the leave and join operations of cloud resources due to failures and other dynamic configuration changes. The rest of this editorial note is organized as follows: Section 2 gives a brief overview of the research and development work carried out for monitoring application QoS over cloud resources; Section 3 summarizes the research contributions that were accepted for this special issue; Section 3 concludes the paper with some future remarks. In last 20 years, a large body of research has focused on developing tools and techniques for monitoring the QoS status of resources and applications over distributed systems (e.g., grids, clusters, and clouds). Some QoS monitoring techniques have been investigated and implemented in computational grids, such as Network Weather Service (NWS) 3, which monitors the network and computing resource QoS and periodically forecast the QoS in a future arrival of a application workload. The current version of NWS gathers the operating system level metrics such as available CPU percentage, available non-paged memory, and TCP/IP Performance. Other monitoring tools 4, 5 that were popular in grid and cluster computing era included R-GMA, Hawkeye, Ganglia, MDS-I, and MDS-II. Aforementioned monitoring techniques and tools were designed for managing static system configuration, where numbers of hardware and software resource types were assumed to remain constant over lifecycle of an application. In other words, these tools did not consider the issue of auto scaling and de-scaling primitives supported by virtualized cloud resources. These tools were only concerned about monitoring the QoS parameters for the hardware resources (CPU, storage, and network), while being completely agnostic to application-specific QoS parameters and SLA requirements. The performance of these tools was optimized for monitoring the QoS of only one type of application (e.g., high performance computing application). On the other hand, in cloud computing datacenters, multiple application instances can be multiplexed and co-allocated on single physical resource. Clearly, the monitoring tools developed in grid and cluster computing era (while being innovative and useful) is not suitable to tackle the challenges on cloud computing environments and hosted application types. Current cloud resource and application QoS monitoring frameworks (e.g., Amazon CloudWatch 6, Azure Fabric Controller) typically monitor the entire virtual machine (VM, a software implementation of a physical CPU resource) as a black box and lacks ability to inter-operate across cloud datacenters managed by different providers (e.g., Amazon, Microsoft, GoGrid, and CA). This means that QoS of software resources (e.g., web server, and database server) contained in the application stack is not properly monitored and managed. While frameworks such as Monitis 7 and Nimsoft 8 overcome the aforementioned limitations of CloudWatch and Fabric Controller, they lack ability to monitor and enforce application-specific QoS requirements. Further, all of the aforementioned frameworks lack ability to predict and detect faults before they occur. Some of the recent research works 9 have also focused on applying large-scale data and pattern mining to the QoS monitoring history and event log data. Authors in 10 evaluated the prediction capability of Support Vector Machine, Neural Network, and Linear Regression techniques for learning the QoS behavior of cloud hosted applications. To predict the CPU usage of VMs, authors in 11 applied Markov Chain model. Authors in 12 applied prediction techniques such as Moving Average, Auto Regression, Neural Networks, Support Vector Machines, and Gene Expression Programming for predictive VM QoS monitoring and provisioning. Most of these techniques focused on monitoring and predicting QoS of VMs rather than individual application components. Further, these approaches did not reason about the interplay of QoS parameters and SLA requirements across multiple layers (software as a service, platform as a service, and infrastructure as a service) of cloud application stack. In this special issue, we present seven articles that tackle several aspects of the aforementioned resource uncertainties for monitoring QoS of applications hosted on Cloud resources. In particular, Ryckbosch and Diwan propose a Temporal Pattern Analyzer system in their paper 13 Analyzing Performance Traces Using Temporal Formulas that uses formulas in linear-temporal logic extended with variables to analyze traces to investigate long-tail performance problems at Google and reduce the manual labor involved in analyzing traces. The technique is applied on user request logs, which contain events at each stage of processing of a user request to Gmail. The authors show that the system can scale to large traces, a prerequisite considering that Gmail produces a million or more events a second. Two of the case studies presented in the paper have directly contributed to improving the performance of Google. Cao et al. also use execution trace information, in this case, CPU load traces and propose 14 a novel method for CPU load prediction for cloud environment based on a dynamic ensemble model to obtain better performances. The ensemble model proposed consists of two layers, a predictor optimization layer that can continuously incorporate new predictor instances and remove those ones with a poor performance and an ensemble layer that is responsible for producing the final prediction based on the results of multiple predictor instances. The four papers are all concerned with monitoring Cloud applications, ranging from a model and language to define design-time adaption techniques in the paper 15 by Inzinger et al. on a Generic Event-Based Monitoring and Adaptation Methodology for Heterogeneous Distributed Systems, to better visualization techniques in the monitoring process in the paper 16 A Novel Monitoring Mechanism by Event Trigger for Hadoop System Performance Analysis by Chang et al., to adapting to failed application service in a distributed environment by introducing fault avoidance service that can be called instead of the failed service by Gülcü et al. in their paper 17 Fault Masking as a Service, to a feature-based high availability mechanism that monitors data streams for a quantile feature in the paper 18 by Ding et al. on a Feature-based High Availability Mechanism for Quantile Tasks in Real-time Data Stream Processing. In particular, Inzinger et al. present 15 a novel domain-specific language termed MONINA that allows specification of system components and their monitoring and adaptation-relevant behavior for controlling Cloud systems. The authors propose a mechanism for optimal deployment of the defined control operators onto available computing resources by monitoring the cloud environment with complex-event processing queries and adapt to problems by condition action rules performed on top of a distributed knowledge base. Chang et al. propose 16 a system called Event Trigger that provides an automatic recording mechanism on the Hadoop Cloud Computing system, to check the system performance at every static time interval, and compares the variation. The performance parameters are collected during the system monitoring process and are applied onto an easy-understandable visual graph for users to adjust the hardware deployment in order to refine the Hadoop system. Gülcü et al. propose 17 an approach to prevent the occurrence of errors that result from the unavailability of partner services in the first place. They introduce a fault avoidance service to which composite services can register at will. After registration, this fault avoidance service periodically checks the partner links, detects unavailable partner services, and updates the composite service with available alternatives. Thus, in case of a partner service error, the composite service will have been updated before attempting an ill-destined request. Ding et al. focus 18 on the monitoring of data streams on the quantile tasks, a typical summary-oriented operation for aggregation, and propose a feature-based high availability mechanism to reduce related overhead and latency. With the help of a monitor module, the quantile feature is maintained incrementally through histogram synopsis over a time-based sliding window. Consequently, failed tasks can be recovered precisely with a high probability in an efficient way. Finally, the special issue is rounded off by a paper 19 on Design and Implementation of Task Scheduling Strategies for Massive Remote Sensing Data Processing Across Multiple Data Centers by Zhang et al. that proposes scheduling strategies for data processing workflows. In particular, they propose scheduling strategies in massive remote sensing data processing to reduce the total task execution time. The authors divided the data processing workflows into two categories, namely, Bag of Tasks applications that consist of a large number of independent tasks and Direct Acyclic Graph applications that contain a large number of interdependent tasks. They propose two strategies to deal with issues in either of the two categories, a Partitioning Group based on Hypergraph algorithm that partitions data into several groups to minimize the amount of sharing data transferring and an Optimized Task Tree strategy to find the key workflow path, which would be endowed with a high priority in the execution. This special issue presents, through these seven papers, several techniques that can dynamically predict and capture the relationship between an application performance targets, current hardware resource allocation, and changes in workload patterns, in order to adjust resource configuration at design-time and run-time. More work in this area is rapidly emerging, further improving the availability of massively distributed Cloud applications. This will further improve the economies of scale of Cloud applications making Cloud Computing an even more compelling paradigm in comparison to traditional in-house hosted applications. Application QoS monitoring will continue to remain an important research area for cloud-based systems. More tangible efforts are needed for developing monitoring tools and techniques that can specify, reason, and monitor QoS related to a variety of application (enterprise, scientific, and streaming big data analytics) and cloud datacenter types (private and public). Further, research should also aim to correlate events with data from many different sources (e.g., holiday schedules, job schedules, and trends from social media about application usage sentiment) in order to predict how external events can impact an application QoS. In this special issue, we have selected research papers that aim to address some of these challenges. We hope that the readers will find the articles of this special issue informative and useful. Rajiv Ranjan 0001, Rajkumar Buyya, Philipp Leitner 0001, Armin Haller, Stefan Tai |
Softw. Pract. Exp. | 3 |
| 2013 | Decisions, Models, and Monitoring - A Lifecycle Model for the Evolution of Service-Based SystemsabstractThe process of engineering and provisioning service-based systems (SBS) follows a complex and dynamic lifecycle with different phases and levels of abstraction. We tackle the problem of making this lifecycle explicit, providing development time and runtime support for evolutionary changes in such systems. SBSs are modeled as integrated ecosystems consisting of four conceptual layers (or phases): design, implementation, deployment, and runtime. Our work is driven by the notion that identifying the right changes (monitoring) and effecting of these changes (adaptation) usually takes place individually on each layer. While considering changes on a single layer (e.g., runtime adaptation) is often sufficient, some cases require systematic escalation to adjacent layers. We present a generic lifecycle model that provides an abstracted view of the problem domain and can be mapped to concrete artifacts on each individual layer. We introduce a real-life scenario taken from the telecommunications domain, which serves as the basis for discussion of the challenges and our solution. Based on the scenario and our experience from a research project on Virtual Service Platforms, we evaluate three concrete use cases which illustrate the diversity of evolutionary changes supported by the approach. Christian Inzinger, Waldemar Hummer, Ioanna Lytra, Philipp Leitner 0001, Uwe Zdun, Schahram Dustdar |
EDOC | 4 |
| 2013 | Model-based Adaptation of Cloud Computing ApplicationsabstractIn this paper we propose a provider-managed, model-based adaptation approach for cloud computing applications, allowing customers to easily specify application behavior goals or adaptation rules. Delegating control over corrective actions to the cloud provider will pose advantages for both, customers and providers. Customers are relieved of effort and expertise requirements necessary to build sophisticated adaptation solutions, while providers can incorporate and analyze data from a multitude of customers to improve adaptation decisions. The envisioned approach will enable increased application performance, as well as cost savings for customers, whereas providers can manage their infrastructure more efficiently. Christian Inzinger, Benjamin Satzger, Philipp Leitner 0001, Waldemar Hummer, Schahram Dustdar |
MODELSWARD | 3 |
| 2013 | Data-driven and automated prediction of service level agreement violations in service compositions
Philipp Leitner 0001, Johannes Ferner, Waldemar Hummer, Schahram Dustdar |
Distributed Parallel Databases | 1 |
| 2013 | Testing of data-centric and event-based dynamic service compositionsabstractSUMMARY This paper addresses integration testing of data‐centric and event‐based dynamic service compositions. The compositions under test define abstract services that are replaced by concrete candidate services at runtime. Testing all possible instantiations of a composition leads to combinatorial explosion and is often infeasible. We consider data dependencies between services as potential points of failure and introduce the k‐node data flow test coverage metric, which helps to significantly reduce the number of test combinations. We formulate a combinatorial optimization problem for generating minimal sets of test cases. On the basis of this formalization, we present a mapping to the model of FoCuS, a coverage analysis tool. FoCuS efficiently computes near‐optimal solutions, which are used to automatically generate test instances. The proposed approach is applicable to various composition paradigms. We illustrate the end‐to‐end practicability based on an integrated scenario, which uses two diverse composition techniques: on the one hand, the Web Services Business Process Execution Language and on the other hand, WS‐Aggregation, a platform for event‐based service composition. Copyright © 2013 John Wiley & Sons, Ltd. Waldemar Hummer, Orna Raz, Onn Shehory, Philipp Leitner 0001, Schahram Dustdar |
Softw. Test. Verification Reliab. | 4 |
| 2013 | Cost-Based Optimization of Service CompositionsabstractFor providers of composite services, preventing cases of SLA violations is crucial. Previous work has established runtime adaptation of compositions as a promising tool to achieve SLA conformance. However, to get a realistic and complete view of the decision process of service providers, the costs of adaptation need to be taken into account. In this paper, we formalize the problem of finding the optimal set of adaptations, which minimizes the total costs arising from SLA violations and the adaptations to prevent them. We present possible algorithms to solve this complex optimization problem, and detail an end-to-end system based on our earlier work on the PREvent (prediction and prevention based on event monitoring) framework, which clearly indicates the usefulness of our model. We discuss experimental results that show how the application of our approach leads to reduced costs for the service provider, and explain the circumstances in which different algorithms lead to more or less satisfactory results. Philipp Leitner 0001, Waldemar Hummer, Schahram Dustdar |
IEEE Trans. Serv. Comput. | 1 |
| 2012 | Cost-Efficient and Application SLA-Aware Client Side Request Scheduling in an Infrastructure-as-a-Service CloudabstractProviders of applications deployed in an Infrastructure-as-a-Service cloud permanently face the decision of whether it is more cost-efficient to scale up(i.e., rent more resources from the cloud) or to delay incoming requests, even though doing so may lead to dissatisfied customers and broken service level agreements. This decision is further complicated by the fact that not all customers have the same agreements, and not all requests require the same amount of resources devoted to them. In this paper, we present an approach for optimally scheduling incoming requests to virtual computing resources in the cloud, so that the sum of payments for resources and loss incurred by service level agreement violations is minimized. We discuss our approach based on an illustrative use case. Furthermore, we present a numerical evaluation based on real-life request data, which shows that our agreement-aware algorithm improves upon earlier work, which does not take service level agreements into account. Philipp Leitner 0001, Waldemar Hummer, Benjamin Satzger, Christian Inzinger, Schahram Dustdar |
IEEE CLOUD | 1 |
| 2012 | Towards Identifying Root Causes of Faults in Service-Based ApplicationsabstractIn this paper we study fault localization techniques for identification of incompatible configurations and implementations in service-based applications. We propose an approach using pooled decision trees for localization of faulty service parameter and binding configurations, explicitly addressing temporary and changing fault conditions. Christian Inzinger, Waldemar Hummer, Benjamin Satzger, Philipp Leitner 0001, Schahram Dustdar |
SRDS | 4 |
| 2011 | Esc: Towards an Elastic Stream Computing Platform for the CloudabstractToday, most tools for processing big data are batch-oriented. However, many scenarios require continuous, online processing of data streams and events. We present ESC, a new stream computing engine. It is designed for computations with real-time demands, such as online data mining. It offers a simple programming model in which programs are specified by directed acyclic graphs (DAGs). The DAG defines the data flow of a program, vertices represent operations applied to the data. The data which are streaming through the graph are expressed as key/value pairs. ESC allows programmers to focus on the problem at hand and deals with distribution and fault tolerance. Furthermore, it is able to adapt to changing computational demands. In the cloud, ESC can dynamically attach and release machines to adjust the computational capacities to the current needs. This is crucial for stream computing since the amount of data fed into the system is not under the platform's control. We substantiate the concepts we propose in this paper with an evaluation based on a high-frequency trading scenario. Benjamin Satzger, Waldemar Hummer, Philipp Leitner 0001, Schahram Dustdar |
IEEE CLOUD | 3 |
| 2011 | Test Coverage of Data-Centric Dynamic Compositions in Service-Based SystemsabstractThis paper addresses the problem of integration testing of data-centric dynamic compositions in service-based systems. These compositions define abstract services, which are replaced by invocations to concrete candidate services at runtime. Testing all possible runtime instances of a composition is often unfeasible. We regard data dependencies between services as potential points of failure, and introduce the k-node data flow test coverage metric. Limiting the level of desired coverage helps to significantly reduce the search space of service combinations. We formulate the problem of generating a minimum set of test cases as a combinatorial optimization problem. Based on the formalization we present a mapping of the problem to the data model of FoCuS, a coverage analysis tool developed at IBM. FoCuS can efficiently compute near-optimal solutions, which we then use to automatically generate and execute test instances of the composition. We evaluate our prototype implementation using an illustrative scenario to show the end-to-end practicability of the approach. Waldemar Hummer, Orna Raz, Onn Shehory, Philipp Leitner 0001, Schahram Dustdar |
ICST | 4 |
| 2011 | Stepwise and Asynchronous Runtime Optimization of Web Service Compositions
Philipp Leitner 0001, Waldemar Hummer, Benjamin Satzger, Schahram Dustdar |
WISE | 1 |
| 2011 | SEPL - a domain-specific language and execution environment for protocols of stateful Web services
Waldemar Hummer, Philipp Leitner 0001, Schahram Dustdar |
Distributed Parallel Databases | 2 |
| 2010 | Preventing SLA Violations in Service Compositions Using Aspect-Based Fragment Substitution
Philipp Leitner 0001, Branimir Wetzstein, Dimka Karastoyanova, Waldemar Hummer, Schahram Dustdar, Frank Leymann |
ICSOC | 1 |
| 2010 | Monitoring, Prediction and Prevention of SLA Violations in Composite ServicesabstractWe propose the PREvent framework, which is a system that integrates event-based monitoring, prediction of SLA violations using machine learning techniques, and automated runtime prevention of those violations by triggering adaptation actions in service compositions. PREvent improves on related work in that it can be used to prevent violations ex ante, before they have negatively impacted the provider's SLAs. We explain PREvent in detail and show the impact on SLA violations based on a case study. Philipp Leitner 0001, Anton Michlmayr, Florian Rosenberg, Schahram Dustdar |
ICWS | 1 |
| 2010 | End-to-End Support for QoS-Aware Service Selection, Binding, and Mediation in VRESCoabstractService-Oriented Computing has recently received a lot of attention from both academia and industry. However, current service-oriented solutions are often not as dynamic and adaptable as intended because the publish-find-bind-execute cycle of the Service-Oriented Architecture triangle is not entirely realized. In this paper, we highlight some issues of current web service technologies, with a special emphasis on service metadata, Quality of Service, service querying, dynamic binding, and service mediation. Then, we present the Vienna Runtime Environment for Service-Oriented Computing (VRESCo) that addresses these issues. We give a detailed description of the different aspects by focusing on service querying and service mediation. Finally, we present a performance evaluation of the different components, together with an end-to-end evaluation to show the applicability and usefulness of our system. Anton Michlmayr, Florian Rosenberg, Philipp Leitner 0001, Schahram Dustdar |
IEEE Trans. Serv. Comput. | 3 |
| 2009 | Selecting Web services based on past user experiencesabstractSince the Internet of Services (IoS) is becoming reality, there is an inherent need for novel service selection mechanisms, which work in spite of large numbers of alternative services and take the user-centric nature of services in the IoS into account. One way to do this is to incorporate feedback from previous service users. However, practical issues such as trust aspects, interaction contexts or synonymous feedbacks have to be taken into account. In this paper we discuss a service selection mechanism which makes use of structured and unstructured feedback to capture the Quality of Experience that services have provided in the past. We have implemented our approach within the SOA runtime VRESCO, where we use freeform tags for unstructured and numerical ratings for structured user feedback. We discuss the general process of feedback-based service selection, and explain how the problems described above can be tackled. We conclude the paper with an illustrative case study and discussion of the presented ideas. Philipp Leitner 0001, Anton Michlmayr, Florian Rosenberg, Schahram Dustdar |
APSCC | 1 |
| 2009 | An End-to-End Approach for QoS-Aware Service CompositionabstractA simple and effective composition of software services into higher-level composite services is still a very challenging task. Especially in enterprise environments, quality of service (QoS) concerns play a major role when building software systems following the service-oriented architecture (SOA) paradigm. Inthis paper we present a composition approach based on a domain-specific language(DSL) for specifying functional requirements of services and the expected QoS inform of constraint hierarchies by leveraging hard and soft constraints. Acomposition runtime will resolve the user's constraints to find an optimize dcomposition semi-automatically. To this end we leverage data flow analysis to generate a structured composition model and use two different techniques for the optimization, a constraint programming and an integer programming approach. Florian Rosenberg, Predrag Celikovic, Anton Michlmayr, Philipp Leitner 0001, Schahram Dustdar |
EDOC | 4 |
| 2009 | Monitoring and Analyzing Influential Factors of Business Process PerformanceabstractBusiness activity monitoring enables continuous observation of key performance indicators (KPIs). However, if things go wrong, a deeper analysis of process performance becomes necessary. Business analysts want to learn about the factors that influence the performance of business processes and most often contribute to the violation of KPI target values, and how they relate to each other. We provide a framework for performance monitoring and analysis of WS-BPEL processes, which consolidates process events and Quality of Service measurements. The framework uses machine learning techniques in order to construct tree structures, which represent the dependencies of a KPI on process and QoS metrics. These dependency trees allow business analysts to analyze how the process KPIs depend on lower-level process metrics and QoS characterisitics of the IT infrastructure. Deeper knowledge about the structure of dependencies can be gained by drill-down analysis of single factors of influence. Branimir Wetzstein, Philipp Leitner 0001, Florian Rosenberg, Ivona Brandic, Schahram Dustdar, Frank Leymann |
EDOC | 2 |
| 2009 | Towards Composition as a Service - A Quality of Service Driven ApproachabstractSoftware as a Service (SaaS) and the possibility to compose Web services provisioned over the Internet are important assets for a service-oriented architecture (SOA). However, the complexity and time for developing and provisioning a composite service is very high and it is generally an error-prone task. In this paper we address these issues by describing a semi-automated "Composition as a Service'' (CaaS) approach combined with a domain-specific language called VCL (Vienna composition language). The proposed approach facilitates rapid development and provisioning of composite services by specifying what to compose in a constraint-hierarchy based way using VCL. Invoking the composition service triggers the composition process and upon success the newly composed service is immediately deployed and available. This solution requires no client-side composition infrastructure because it is transparently encapsulated in the CaaS infrastructure. Florian Rosenberg, Philipp Leitner 0001, Anton Michlmayr, Predrag Celikovic, Schahram Dustdar |
ICDE | 2 |
| 2009 | Service Provenance in QoS-Aware Web Service RuntimesabstractIn general, provenance of electronic data represents an important issue in information systems. So far, service-oriented computing research has mainly focused on provenance of data. However, service provenance also plays a central role since service providers and consumers want to be aware of the service's origin and history. In this paper, we present an approach for service provenance that builds on service metadata and various service runtime events. In addition, access control mechanisms are implemented to restrict access to this information. Besides being able to query and subscribe to provenance information, provenance graphs can be used to illustrate the history of services. We give some usage examples of service provenance and show how our approach was integrated into the VRESCo Web service runtime environment. Anton Michlmayr, Florian Rosenberg, Philipp Leitner 0001, Schahram Dustdar |
ICWS | 3 |
| 2007 | Fault Management based on peer-to-peer paradigms; A case study report from the CELTIC project MadeiraabstractWe present an approach to fault management based on an architecture for distributed and collaborative network management as developed in the CELTIC project Madeira. It uses peer-to-peer communication facilities and a logical overlay network facilitating decentralized and iterative alarm processing and correlation. We argue that such an approach might help to overcome key challenges that are posed by NGN scenarios to traditional centralized network management systems. Its feasibility is demonstrated by means of a case study from the area of wireless mesh networks, where an application prototype has been developed. Markus Leitner, Philipp Leitner 0001, Martin Zach, Sandra Collins, Claire Fahy |
Integrated Network Management | 2 |
| 2006 | A Distributed Policy Based Solution in a Fault Management ScenarioabstractThe Madeira project, part of the Celtic Initiative1, investigates the use of a fully distributed, policy-based network management framework that exploits the peer-to-peer paradigm with the aim of providing a successful solution to Next Generation Networks (NGN) challenges. This paper is focused on the distributed policy-based approach adopted in the project. Thanks to this approach, the management system is flexible and adaptable to different management applications of networks with time-varying topologies. Currently available results coming from the execution of specific scenarios in the area of Fault Management reveal that the policy-based system architecture works properly in the highly distributed peer-to-peer environment where it is deployed. Ricardo Marin, Julio Vivero, Joan Serrat 0001, Philipp Leitner 0001, Martin Zach, Claire Fahy |
GLOBECOM | 5 |