Giuliano Casale

dblp:16/2365 · DBLP profile ↗
← Back
114ranked-venue papers
34as first author
44since 2021 · last 2026
0000-0003-4548-7951ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 48 · 18 first-author · 15 since 2021Software engineering, systems software and programming languages · 28 · 13 first-author · 7 since 2021Computer networks · 22 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Security and privacy · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fine-grained Tracing for Performance Anomaly Diagnosis of Serverless Functions
abstract
Serverless function compositions subject to unpredictable faults are challenging to evaluate for root cause analysis. Even though distributed tracing provides observations at multiple levels of granularity for troubleshooting, excessive code instrumentation increases the tracing overheads in terms of both computation and storage. Therefore, developers face the challenge of where and how to instrument serverless functions to maximize the likelihood of locating faults based on tracing data while minimizing tracing overhead and costs. In this article, we propose a methodology to instrument an application with code-level tracing to infer the location of faults, taking into account constraints in terms of the maximum cost of the instrumentation and testing. We encode the tracing probe placement based on the control flow graph of the application and devise heuristics-based tracing data collection strategies to relate possible probe placements with their ability to locate a fault. Then we train novelty detection models to identify the internal anomalies and present an enhanced global search algorithm that automatically computes a probe placement with optimal fault localization ability versus cost. Experimental results show high performance in locating single and multiple faults with over 90% recall score for up to 15% latency anomalies, with minimal instrumentation overhead.
Runan Wang, Guangba Yu, Giuliano Casale, Pengfei Chen 0002, Antonio Filieri
ACM Trans. Auton. Adapt. Syst.3
2025 CEED: Collaborative Early Exit Neural Network Inference at the Edge
Yichong Chen, Zifeng Niu, Manuel Roveri, Giuliano Casale
INFOCOM4
2025 DeepBAT: Performance and Cost Optimization of Serverless Inference Using Transformers
abstract
Serverless computing is an autoscaling pay-as-yougo paradigm that can efficiently support machine learning inference especially under bursty workload conditions. Within the serverless paradigm, batching ML inference requests before serving them is widely adopted. Thanks to its parallelism properties, batching can highly improve inference performance while reducing the monetary cost of serverless. Identifying the correct serverless parameterization to simultaneously meet conflicting targets (i.e., keep monetary cost at a minimum while meeting pre-defined service level objectives, SLO) may be cast as a resource allocation problem. In this paper, we illustrate that a deep surrogate model can quickly discover optimized serverless configurations by learning the relationship among the workload patterns and achieve performance measures. We develop DeepBAT, an SLO-aware framework that leverages the Transformer encoder and multi-head attention mechanism to optimize the performance of serverless inference subject to bursty and previously unobserved workloads. We illustrate the effectiveness of DeepBAT on a set of case studies and show that for the problem of inference serving on AWS Lambda, DeepBAT can speed up the solution time of state-of-the-art analytic solutions by over 55 times while generalizing remarkably well on unseen workloads.
Riccardo Pinciroli, Giuliano Casale, Evgenia Smirni
IPDPS3
2025 Enhanced Training of Response Time Anomaly Detectors Using Diffusion Models
abstract
Machine learning (ML) approaches have grown in popularity in recent years as a way to identify anomalies in microservice-based applications. However, training ML models may suffer from time constraints and limited real-world failure data availability. In this paper, we propose a method to mitigate this issue using diffusion models for training data augmentation. Response time data collected using distributed traces is used to train diffusion models leveraging a customized UNet proposed in the paper. The resulting diffusion models can then generate new data to improve the training of response time anomaly detectors. Experiments using the DeathStarBench microservices architecture demonstrate that the proposed approach increases the accuracy after training of anomaly detection models by 20%. We further show that the response time data generated by our diffusion models cannot be distinguished by classic discriminators, which confirms that the generated data are of high quality.
Wenxiang Luo, Giuliano Casale
MASCOTS2
2025 Competitive Analysis of Vehicle-Sharing Systems With Cournot Queueing Games
abstract
Product-form closed queueing networks are a useful formalism for analyzing the availability of vehicles within a vehicle sharing system with stochastic user behavior. However, existing models assume that the provider has full control over the system. In reality, competition between providers within a single geographic area is quite common. In this paper, we introduce a non-cooperative game that extends a closed queueing model into a competitive environment, in which players decide the number of jobs that they wish to submit. This can be used to model vehicle sharing systems with multiple providers, in which they receive fares from trips but are responsible for the costs of their own fleet. The core technical results of this paper include conditions that guarantee the existence of a pure Nash equilibrium, and an efficient equilibrium-finding algorithm. We then present a case study, using our model, of a vehicle sharing system in Oslo with multiple providers. In this case study, we find that adding an additional competitor can increase the number of trips by up to 18.9%, and a highly competitive market can increase this by up to 30%.
Matthew Sheldon, Dario Paccagnan, Giuliano Casale
IEEE Trans. Intell. Transp. Syst.3
2025 TBI: Transient Hierarchical Modeling of Large-Scale Vehicle Sharing Systems
abstract
Due to rush-hour effects, transient analysis may be necessary to model and optimize vehicle sharing systems. However, existing methods for transient modeling face computational challenges at scale. To solve this, we present trajectory-based iteration (TBI), a method for the efficient decomposed modeling of vehicle-sharing systems, and we extend it to a scalable method for non-convex, transient optimization problems. The proposed TBI method works by iteratively solving spatial submodels, then periodically updating the vehicle flows between submodels. We then present experimental results where our method results in a 37% improvement in runtime for solving fluid models for mobility systems. Then, we present a case study in which our optimization method results in an estimated 22% improvement in the objective function for a vehicle-sharing system with offline rebalancing.
Matthew Sheldon, Daphné Tuncer, Giuliano Casale
IEEE Trans. Intell. Transp. Syst.3
2025 Towards Cost-Optimal Policies for DAGs to Utilize IaaS Clouds With Online Learning
abstract
Premier cloud service providers (CSPs) offer on-demand and spot instances (i.e., virtual machines) with time-varying features in availability and price. While interacting with a CSP, what concerns users is the process of cost-effectively utilizing these instances, possibly in addition to self-owned instances. A job in data-intensive applications is represented by a directed acyclic graph (DAG) whose nodes represent tasks and the DAG can further be transformed into a chain of tasks. The key to achieving cost efficiency is determining the allocation of a specific deadline to each task, as well as the allocation of different types of instances to the task. In this paper, we make some mild assumptions on the usage of various instances and give an analytical understanding of the expected behaviors while utilizing instances. Based on this, we propose informed heuristic policies to determine the allocation of deadlines and instances. The policies are parametric to support the usage of online learning to infer the optimal values against the dynamics of cloud markets. Finally, intuitive greedy policies are used as baselines to validate the effectiveness of the proposed analytical solutions. The cost improvement is up to 23.58% when spot and on-demand instances are considered and up to 50.08% when self-owned instances are also considered.
Han Yu 0001, Giuliano Casale, Guanyu Gao
IEEE Trans. Serv. Comput.3
2024 ChainNet: A Customized Graph Neural Network Model for Loss-Aware Edge AI Service Deployment
abstract
Edge AI seeks for the deployment of deep neural network (DNN) based services across distributed edge devices, embedding intelligence close to data sources. Due to capacity constraints at the edge, a difficult challenge lies in planning a dependable deployment that minimizes the data loss rate so as to meet application Quality-of-Service (QoS) goals. In this paper, we present ChainNet, a customized graph neural network (GNN) model serving as a surrogate to assess the reliability of alternative deployments and guide the loss-aware search for an optimal edge AI deployment plan. Extensive results show that ChainNet delivers a substantial improvement in loss prediction accuracy by over 50% compared to established GNN models, such as graph attention networks (GATs). Moreover, we show that ChainNet provides significantly more dependable deployment decisions under a fixed time budget compared to simulation-based search across a spectrum of systems from small to large-scale.
Zifeng Niu, Manuel Roveri, Giuliano Casale
DSN3
2024 Approximating Closed Queueing Networks in Semi-Markov Random Environments
abstract
Modeling queueing networks in random environments is a challenging problem that finds application in routing, reliability analysis, and evaluation of systems processing bursty workloads. In this paper, we introduce a method to evaluate closed queueing networks in random environments that can deal with non-Markovian transitions among environment stages and accurately approximate time-averaged performance metrics. Our technique supports both state-independent and state-dependent random environments, the latter allowing us to analyze challenging queueing network models with bursty service processes. Using simulation on representative case studies, we show that our technique provides low approximation errors. We further demonstrate the applicability of our result to a mobility case study demonstrating the applicability of the proposed approach.
Yaqi Zhou, Matthew Sheldon, Giuliano Casale
MASCOTS3
2024 Optimizing Edge AI: Performance Engineering in Resource-Constrained Environments
abstract
Recent years have witnessed the growth of Edge AI, a transformative paradigm that integrates neural networks with edge computing, bringing computational intelligence closer to end users. However, this innovation is not without its challenges, especially in environments with limited computing, network, and memory constraints, where resource-hungry AI models often need to be partitioned for distributed execution. This issue becomes even more acute in scenarios where post-deployment updates are infeasible or costly, posing a need to accurately reason about the interplay between resource constraints and Quality-of-Service (QoS) in Edge AI systems, so as to optimally design and operate them.
Giuliano Casale
ICPE1
2024 RADF: Architecture decomposition for function as a service
abstract
Abstract As the most successful realization of serverless, function as a service (FaaS) brings in a novel cloud computing paradigm that can save operating costs, reduce management effort, enable seamless scalability, and augment development productivity. Migration of an existing application to the serverless architecture is, however, an intricate task as a great number of decisions need to be made along the way. We propose in this paper RADF, a semi‐automatic approach that decomposes a monolith into serverless functions by analyzing the business logic inherent in the interface of the application. The proposed approach adopts a two‐stage refactoring strategy, where a coarse‐grained decomposition is performed at first, followed by a fine‐grained one. As such, the decomposition process is simplified into smaller steps and adaptable to generate a solution at either microservice or function level. We have implemented RADF in a holistic DevOps methodology and evaluated its capability for microservice identification and feasibility for code refactoring. In the evaluation experiments, RADF achieves lower coupling and relatively balanced cohesion, compared to previous decomposition approaches.
Lulai Zhu, Damian A. Tamburri, Giuliano Casale
Softw. Pract. Exp.3
2024 Scheduling Inputs in Early Exit Neural Networks
abstract
Early exit neural networks (EENs) reduce the processing times of deep convolutional neural networks by means of internal classifiers (ICs) that allow jobs, being the input of the EEN, to exit early from the processing pipeline. However, the current designs used in pervasive systems ignore variability in data arrival rates, exposing EEN-based services to potential loss of the incoming jobs, due to finite input buffer capacity. Motivated by this issue, we introduce and study theearly exitscheduling problem, which aims at dynamically configuring IC thresholds at runtime to achieve effective trade-offs between job classification accuracy, processing time, and job loss ratio. We argue that deciding the EEN exit layer for a job at the start of its processing makes the problem mathematically tractable, allowing us to develop policies to control buffer backlog, classification accuracy, and processing time across the EEN layers. The main contribution of the paper is the introduction of single-exit IC threshold configurations as a mechanism to allow the scheduling policy to reliably predict the best EEN exit layer of each input job. Three scheduling policies that leverage this idea are proposed to dynamically schedule job arrivals to an EEN-based service. The proposed solution, here tailored to EENs based on convolutional neural networks (CNNs), is fairly general and can be applied to different use cases. The two application scenarios considered in this paper focus on image classification and intrusion detection. Experiments on some popular CNNs for the two aforementioned application scenarios indicate that the proposed policies can achieve significant savings in processing times and improve job loss ratio compared to both ordinary EENs and CNNs while still providing high mean classification accuracy.
Giuliano Casale, Manuel Roveri
IEEE Trans. Computers1
2024 PreGAN+: Semi-Supervised Fault Prediction and Preemptive Migration in Dynamic Mobile Edge Environments
abstract
Typical mobile edge computing infrastructures have to contend with unreliable computing devices at their end-points. The limited resource capacities of mobile edge devices gives rise to frequent contentions, node overloads or failures. This is exacerbated by the strict deadlines of modern applications. To avoid failures, fault-tolerant approaches utilize preemptive migration to transfer active tasks across nodes and prevent nodes running at capacity. However, prior work struggles to dynamically adapt in settings with highly volatile workloads or even accurately detect and diagnose anomalies for optimal remediation. To meet the strict service level objectives of contemporary workloads, there is a need for dynamic fault-tolerant methods that can quickly adapt to changes in edge environments while having parsimonious remediation in the form of preemptive migration to avoid stressing the system network. This work proposes PreGAN, featuring a Generative Adversarial Network (GAN) based approach to predict contentions, pinpoint specific resource types with high chance of overload, and generate migration decisions to proactively avoid system downtime. PreGAN leverages coupled-simulations to train the GAN model at run-time and a few-shot fault classifier to update decisions of an underpinning scheduler. We also extend it to PreGAN+ that also periodically tunes the decision model using semi-supervised training and a Transformer based neural network for low tuning time, albeit with higher memory overheads. Experiments on a Raspberry-Pi based edge environment demonstrate that both models outperform state-of-the-art baselines in fault detection and diagnosis scores by up to 12.5% and 31.2% respectively. This also translates in improvements in Quality of Service against baseline approaches.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Mob. Comput.2
2024 Neural Density Estimation of Response Times in Layered Software Systems
abstract
Layered queueing networks (LQNs) are a class of performance models for software systems in which multiple distributed resources may be possessed simultaneously by a job. Estimating response times in a layered system is an essential but challenging analysis dimension in Quality of Service (QoS) assessment. Current analytic methods are capable of providing accurate estimates of mean response times. However, accurately approximating response time distributions used in service-level objective analysis is a demanding task. This paper proposes a novel hybrid framework that leverages phase-type (PH) distributions and neural networks to provide accurate density estimates of response times in layered queueing networks. The core step of this framework is to recursively obtain response time distributions in the submodels that are used to analyze the network by means of decomposition. We describe these response time distributions as a mixture of density functions for which we learn the parameters through a Mixture Density Network (MDN). The approach recursively propagates MDN predictions across software layers using PH distributions and performs repeated moment-matching based refitting to efficiently estimate end-to-end response time densities. Extensive numerical experiment results show that our scheme significantly improves density estimations compared to the state-of-the-art.
Zifeng Niu, Giuliano Casale
IEEE Trans. Software Eng.2
2023 Coupling QoS Co-Simulation with Online Adaptive Arrival Forecasting
abstract
Coupled simulation, also known as co-simulation, has been proposed to provide more information to a task scheduler by simulating at runtime the Quality of Service (QoS) arising from a scheduling action. To do so, co-simulation algorithms run the simulation assuming a static set of arrival time series, restricting the diversity of the traffic scenarios. To ensure the co-simulator can provide valuable and representative results, we present an online adaptive arrival forecasting framework that contains a change-point detection module and a probabilistic transformer model to couple co-simulators with arrival series forecasting. The framework can also update the prediction model to adapt to dynamic environments. Our experiments show that our online adaptive forecasting framework has lower forecasting errors than established prediction models, such as autoregressive processes, and lower on real-world traces the co-simulator prediction error by up to 27 % on average response time and 39% on average service-level agreement (SLA) violation.
Yichong Chen, Manuel Roveri, Shreshth Tuli, Giuliano Casale
CNSM4
2023 DeepFT: Fault-Tolerant Edge Computing using a Self-Supervised Deep Surrogate Model
abstract
The emergence of latency-critical AI applications has been supported by the evolution of the edge computing paradigm. However, edge solutions are typically resource-constrained, posing reliability challenges due to heightened contention for compute capacities and faulty application behavior in the presence of overload conditions. Although a large amount of generated log data can be mined for fault prediction, labeling this data for training is a manual process and thus a limiting factor for automation. Due to this, many companies resort to unsupervised fault-tolerance models. Yet, failure models of this kind can incur a loss of accuracy when they need to adapt to non-stationary workloads and diverse host characteristics. Thus, we propose a novel modeling approach, DeepFT, to proactively avoid system overloads and their adverse effects by optimizing the task scheduling decisions. DeepFT uses a deep-surrogate model to accurately predict and diagnose faults in the system and co-simulation based self-supervised learning to dynamically adapt the model in volatile settings. Experimentation on an edge cluster shows that DeepFT can outperform state-of-the-art methods in fault-detection and QoS metrics. Specifically, DeepFT gives the highest F1 scores for fault-detection, reducing service deadline violations by up to 37% while also improving response time by up to 9%.
Shreshth Tuli, Giuliano Casale, Ludmila Cherkasova, Nicholas R. Jennings
INFOCOM2
2023 SampleHST: Efficient On-the-Fly Selection of Distributed Traces
abstract
Since only a small number of traces generated from distributed tracing helps in troubleshooting, its storage requirement can be significantly reduced by biasing the selection towards anomalous traces. To aid in this scenario, we propose SampleHST, a novel approach to sample on-the-fly from a stream of traces in an unsupervised manner. SampleHST adjusts the storage quota of normal and anomalous traces depending on the size of its budget. Initially, it utilizes a forest of Half Space Trees (HSTs) for trace scoring. This is based on the distribution of the mass scores across the trees, which characterizes the probability of observing different traces. The mass distribution from HSTs is subsequently used to cluster the traces online leveraging a variant of the mean-shift algorithm. This trace-cluster association eventually drives the sampling decision. We have compared the performance of SampleHST with a recently suggested method using data from a cloud data center and demonstrated that SampleHST improves sampling performance up to by 9.5$\times$.
Alim Ul Gias, Matthew Sheldon, José A. Perusquía, Owen O'Brien, Giuliano Casale
NOMS6
2023 AI augmented Edge and Fog computing: Trends and challenges
abstract
In recent years, the landscape of computing paradigms has witnessed a gradual yet remarkable shift from monolithic computing to distributed and decentralized paradigms such as Internet of Things (IoT), Edge, Fog, Cloud, and Serverless. The frontiers of these computing technologies have been boosted by shift from manually encoded algorithms to Artificial Intelligence (AI)-driven autonomous systems for optimum and reliable management of distributed computing resources. Prior work focuses on improving existing systems using AI across a wide range of domains, such as efficient resource provisioning, application deployment, task placement, and service management. This survey reviews the evolution of data-driven AI-augmented technologies and their impact on computing systems. We demystify new techniques and draw key insights in Edge, Fog and Cloud resource management-related uses of AI methods and also look at how AI can innovate traditional applications for enhanced Quality of Service (QoS) in the presence of a continuum of resources. We present the latest trends and impact areas such as optimizing AI models that are deployed on or for computing systems. We layout a roadmap for future research directions in areas such as resource management for QoS optimization and service reliability. Finally, we discuss blue-sky ideas and envision this work as an anchor point for future research on AI-driven computing systems.
Shreshth Tuli, Fatemeh Mirhakimi, Samodha Pallewatta, Syed Zawad, Giuliano Casale, Bahman Javadi, Feng Yan 0001, Rajkumar Buyya, Nicholas R. Jennings
J. Netw. Comput. Appl.5
2023 SciNet: Codesign of Resource Management in Cloud Computing Environments
abstract
The rise of distributed cloud computing technologies has been pivotal for the large-scale adoption of Artificial Intelligence (AI) based applications for high fidelity and scalable service delivery. Systematic resource management is central in maintaining optimal Quality of Service (QoS) in cloud platforms and is divided into three fundamental types: resource provisioning, AI model deployment and workload placement. To exploit the synergy among these decision types, it becomes imperative to concurrently design (co-design) the provisioning, deployment and placement decisions for optimal QoS. As users and cloud service providers shift to non-stationary AI-based workloads, frequent decision making imposes severe time constraints on the resource management models. Existing AI-based solutions often optimize decision types independently and tend to ignore the dependencies across various system performance aspects such as energy consumption and CPU utilization, making them perform poorly in large-scale cloud systems. To address this, we propose a novel method, called SciNet, that leverages a co-simulated digital-twin of the infrastructure to capture inter-metric dependencies and accurately estimate QoS scores. To avoid expensive simulation overheads at test time, SciNet trains a neural network based imitation learner that aims to mimic an oracle, which takes optimal decisions based on co-simulated QoS estimates. Offline model training and online decision making based on the imitation learner, enables SciNet to take optimal decisions while being time-efficient. Experiments with real-life AI-based benchmark applications on a public cloud testbed show that SciNet gives up to 48% lower execution cost, 79% higher inference accuracy, 71% lower energy consumption and 56% lower response times compared to the current state-of-the-art methods.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Computers2
2023 Redundancy Planning for Cost Efficient Resilience to Cyber Attacks
abstract
We investigate the extent to which redundancy (including with diversity) can help mitigate the impact of cyber attacks that aim to reduce system performance. Using analytical techniques, we estimate impacts, in terms of monetary costs, of penalties from breaching Service Level Agreements (SLAs), and find optimal resource allocations to minimize the overall costs arising from attacks. Our approach combines attack impact analysis, based on performance modeling using queueing networks, with an attack model based on attack graphs. We evaluate our approach using a case study of a website, and show how resource redundancy and diversity can improve the resilience of a system by reducing the likelihood of a fully disruptive attack. We find that the cost-effectiveness of redundancy depends on the SLA terms, the probability of attack detection, the time to recover, and the cost of maintenance. In our case study, redundancy with diversity achieved a saving of up to around 50 percent in expected attack costs relative to no redundancy. The overall benefit over time depends on how the saving during attacks compares to the added maintenance costs due to redundancy.
Jukka Soikkeli, Giuliano Casale, Luis Muñoz-González, Emil C. Lupu
IEEE Trans. Dependable Secur. Comput.2
2023 SplitPlace: AI Augmented Splitting and Placement of Large-Scale Neural Networks in Mobile Edge Environments
abstract
In recent years, deep learning models have become ubiquitous in industry and academia alike. Deep neural networks can solve some of the most complex pattern-recognition problems today, but come with the price of massive compute and memory requirements. This makes the problem of deploying such large-scale neural networks challenging in resource-constrained mobile edge computing platforms, specifically in mission-critical domains like surveillance and healthcare. To solve this, a promising solution is to split resource-hungry neural networks into lightweight disjoint smaller components for pipelined distributed processing. At present, there are two main approaches to do this: semantic and layer-wise splitting. The former partitions a neural network into parallel disjoint models that produce a part of the result, whereas the latter partitions into sequential models that produce intermediate results. However, there is no intelligent algorithm that decides which splitting strategy to use and places such modular splits to edge nodes for optimal performance. To combat this, this work proposes a novel AI-driven online policy, SplitPlace, that uses Multi-Armed-Bandits to intelligently decide between layer and semantic splitting strategies based on the input task's service deadline demands. SplitPlace places such neural network split fragments on mobile edge devices using decision-aware reinforcement learning for efficient and scalable computing. Moreover, SplitPlace fine-tunes its placement engine to adapt to volatile environments. Our experiments on physical mobile-edge environments with real-world workloads show that SplitPlace can significantly improve the state-of-the-art in terms of average response time, deadline violation rate, inference accuracy, and total reward by up to 46, 69, 3 and 12 percent respectively.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Mob. Comput.2
2023 DRAGON: Decentralized Fault Tolerance in Edge Federations
abstract
Edge Federation is a new computing paradigm that seamlessly interconnects the resources of multiple edge service providers. A key challenge in such systems is the deployment of latency-critical and AI based resource-intensive applications in constrained devices. To address this challenge, we propose a novel memory-efficient deep learning based model, namely generative optimization networks (GON). Unlike GANs, GONs use a single network to both discriminate input and generate samples, significantly reducing their memory footprint. Leveraging the low memory footprint of GONs, we propose a decentralized fault-tolerance method called DRAGON that runs simulations (as per a digital modeling twin) to quickly predict and optimize the performance of the edge federation. Extensive experiments with real-world edge computing benchmarks on multiple Raspberry-Pi based federated edge configurations show that DRAGON can outperform the baseline methods in fault-detection and Quality of Service (QoS) metrics. Specifically, the proposed method gives higher F1 scores for fault-detection than the best deep learning (DL) method, while consuming lower memory than the heuristic methods. This allows for improvement in energy consumption, response time and service level agreement violations by up to 74, 63 and 82 percent, respectively.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Netw. Serv. Manag.2
2023 CILP: Co-Simulation-Based Imitation Learner for Dynamic Resource Provisioning in Cloud Computing Environments
abstract
Intelligent Virtual Machine (VM) provisioning is central to cost and resource efficient computation in cloud computing environments. As bootstrapping VMs is time-consuming, a key challenge for latency-critical tasks is to predict future workload demands to provision VMs proactively. However, existing AI-based solutions tend to not holistically consider all crucial aspects such as provisioning overheads, heterogeneous VM costs and Quality of Service (QoS) of the cloud system. To address this, we propose a novel method, called CILP, that formulates the VM provisioning problem as two sub-problems of prediction and optimization, where the provisioning plan is optimized based on predicted workload demands. CILP leverages a neural network as a surrogate model to predict future workload demands with a co-simulated digital-twin of the infrastructure to compute QoS scores. We extend the neural network to also act as an imitation learner that dynamically decides the optimal VM provisioning plan. A transformer based neural model reduces training and inference overheads while our novel two-phase decision making loop facilitates in making informed provisioning decisions. Crucially, we address limitations of prior work by including resource utilization, deployment costs and provisioning overheads to inform the provisioning decisions in our imitation learning framework. Experiments with three public benchmarks demonstrate that CILP gives up to 22% higher resource utilization, 14% higher QoS scores and 44% lower execution costs compared to the current online and offline optimization based state-of-the-art methods.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Netw. Serv. Manag.2
2023 Guest Editorial: Special Section on Machine Learning and Artificial Intelligence for Managing Networks, Systems, and Services - Part II
abstract
Machine learning and artificial intelligence can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, machine learning and artificial intelligence have found ways to improve operations and management of information technology and communications.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Giuliano Casale, Khalil El-Khatib, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova, Farhana Zulkernine
IEEE Trans. Netw. Serv. Manag.4
2023 START: Straggler Prediction and Mitigation for Cloud Computing Environments Using Encoder LSTM Networks
abstract
A common performance problem in large-scale cloud systems is dealing with straggler tasks that are slow running instances which increase the overall response time. Such tasks impact the system's QoS and the SLA. There is a need for automatic straggler detection and mitigation mechanisms that execute jobs without violating the SLA. Prior work typically builds reactive models that focus first on detection and then mitigation of straggler tasks, which leads to delays. Other works use prediction based proactive mechanisms, but ignore volatile task characteristics. We propose a Straggler Prediction and Mitigation Technique (START) that is able to predict which tasks might be stragglers and dynamically adapt scheduling to achieve lower response times. START analyzes all tasks and hosts based on compute and network resource consumption using an Encoder LSTM network to predict and mitigate expected straggler tasks. This reduces the SLA violation rate and execution time without compromising QoS. Specifically, we use the CloudSim toolkit to simulate START and compare it with IGRU-SD, SGC, Dolly, GRASS, NearestFit and Wrangler in terms of QoS parameters. Experiments show that START reduces execution time, resource contention, energy and SLA violations by 13%, 11%, 16%, 19%, compared to the state-of-the-art.
Shreshth Tuli, Sukhpal Singh, Peter Garraghan, Rajkumar Buyya, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Serv. Comput.5
2022 MetaNet: Automated Dynamic Selection of Scheduling Policies in Cloud Environments
abstract
Task scheduling is a well-studied problem in the context of optimizing the Quality of Service (QoS) of cloud computing environments. In order to sustain the rapid growth of computational demands, one of the most important QoS metrics for cloud schedulers is the execution cost. In this regard, several data-driven deep neural networks (DNNs) based schedulers have been proposed in recent years to allow scalable and efficient resource management in dynamic workload settings. However, optimal scheduling frequently relies on sophisticated DNNs with high computational needs implying higher execution costs. Further, even in non-stationary environments, sophisticated schedulers might not always be required and we could briefly rely on low-cost schedulers in the interest of cost-efficiency. Therefore, this work aims to solve the non-trivial meta problem of online dynamic selection of a scheduling policy using a surrogate model called MetaNet. Unlike traditional solutions with a fixed scheduling policy, MetaNet on-the-fly chooses a scheduler from a large set of DNN based methods to optimize task scheduling and execution costs in tandem. Compared to state-of-the-art DNN schedulers, this allows for improvement in execution costs, energy consumption, response time and service level agreement violations by up to 11, 43, 8 and 13 percent, respectively.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
CLOUD2
2022 CAROL: Confidence-Aware Resilience Model for Edge Federations
abstract
In recent years, the deployment of large-scale Inter-net of Things (IoT) applications has given rise to edge federations that seamlessly interconnect and leverage resources from multiple edge service providers. The requirement of supporting both latency-sensitive and compute-intensive IoT tasks necessitates service resilience, especially for the broker nodes in typical broker-worker deployment designs. Existing fault-tolerance or resilience schemes often lack robustness and generalization capability in non-stationary workload settings. This is typically due to the expensive periodic fine-tuning of models required to adapt them in dynamic scenarios. To address this, we present a confidence aware resilience model, CAROL, that utilizes a memory-efficient generative neural network to predict the Quality of Service (QoS) for a future state and a confidence score for each prediction. Thus, whenever a broker fails, we quickly recover the system by executing a local-search over the broker-worker topology space and optimize future QoS. The confidence score enables us to keep track of the prediction performance and run parsimonious neural network fine-tuning to avoid excessive overheads, further improving the QoS of the system. Experiments on a Raspberry-Pi based edge testbed with IoT benchmark applications show that CAROL outperforms state-of-the-art resilience schemes by reducing the energy consumption, deadline violation rates and resilience overheads by up to 16, 17 and 36 percent, respectively.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
DSN2
2022 Enhancing Performance Modeling of Serverless Functions via Static Analysis
Runan Wang, Giuliano Casale, Antonio Filieri
ICSOC2
2022 PreGAN: Preemptive Migration Prediction Network for Proactive Fault-Tolerant Edge Computing
abstract
Building a fault-tolerant edge system that can quickly react to node overloads or failures is challenging due to the unreliability of edge devices and the strict service deadlines of modern applications. Moreover, unnecessary task migrations can stress the system network, giving rise to the need for a smart and parsimonious failure recovery scheme. Prior approaches often fail to adapt to highly volatile workloads or accurately detect and diagnose faults for optimal remediation. There is thus a need for a robust and proactive fault-tolerance mechanism to meet service level objectives. In this work, we propose PreGAN, a composite AI model using a Generative Adversarial Network (GAN) to predict preemptive migration decisions for proactive fault-tolerance in containerized edge deployments. PreGAN uses co-simulations in tandem with a GAN to learn a few-shot anomaly classifier and proactively predict migration decisions for reliable computing. Extensive experiments on a Raspberry-Pi based edge environment show that PreGAN can outperform state-of-the-art baseline methods in fault-detection, diagnosis and classification, thus achieving high quality of service. PreGAN accomplishes this by 5.1% more accurate fault detection, higher diagnosis scores and 23.8% lower overheads compared to the best method among the considered baselines.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
INFOCOM2
2022 JCSP: Joint Caching and Service Placement for Edge Computing Systems
abstract
With constrained resources, what, where, and how to cache at the edge is one of the key challenges for edge computing systems. The cached items include not only the application data contents but also the local caching of edge services that handle incoming requests. However, current systems separate the contents and services without considering the latency interplay of caching and queueing. Therefore, in this paper, we propose a novel class of stochastic models that enable the optimization of content caching and service placement decisions jointly. We first explain how to apply layered queueing networks (LQNs) models for edge service placement and show that combining this with genetic algorithms provides higher accuracy in resource allocation than an established baseline. Next, we extend LQNs with caching components to establish a joint modeling method for content caching and service placement (JCSP) and present analytical methods to analyze the resulting model. Finally, we simulate real-world Azure traces to evaluate the JCSP method and find that JCSP achieves up to 35% improvement in response time and 500MB reduction in memory usage than baseline heuristics for edge caching resource allocation.
Giuliano Casale
IWQoS2
2022 HUNTER: AI based holistic resource management for sustainable cloud computing
Shreshth Tuli, Sukhpal Singh, Minxian Xu, Peter Garraghan, Rami Bahsoon, Schahram Dustdar, Rizos Sakellariou, Omer F. Rana, Rajkumar Buyya, Giuliano Casale, Nicholas R. Jennings
J. Syst. Softw.10
2022 TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series Data
abstract
Efficient anomaly detection and diagnosis in multivariate time-series data is of great importance for modern industrial applications. However, building a system that is able to quickly and accurately pinpoint anomalous observations is a challenging problem. This is due to the lack of anomaly labels, high data volatility and the demands of ultra-low inference times in modern applications. Despite the recent developments of deep learning approaches for anomaly detection, only a few of them can address all of these challenges. In this paper, we propose TranAD, a deep transformer network based anomaly detection and diagnosis model which uses attention-based sequence encoders to swiftly perform inference with the knowledge of the broader temporal trends in the data. TranAD uses focus score-based self-conditioning to enable robust multi-modal feature extraction and adversarial training to gain stability. Additionally, model-agnostic meta learning (MAML) allows us to train the model using limited data. Extensive empirical studies on six publicly available datasets demonstrate that TranAD can outperform state-of-the-art baseline methods in detection and diagnosis performance with data and time-efficient training. Specifically, TranAD increases F1 scores by up to 17%, reducing training times by up to 99% compared to the baselines.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
Proc. VLDB Endow.2
2022 Guest Editorial: Special Issue on Machine Learning and Artificial Intelligence for Managing Networks, Systems, and Services - Part I
abstract
Machine learning and artificial intelligence can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, machine learning and artificial intelligence have found ways to improve operations and management of information technology and communications.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Giuliano Casale, Khalil El-Khatib, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova, Farhana Zulkernine
IEEE Trans. Netw. Serv. Manag.4
2022 MCDS: AI Augmented Workflow Scheduling in Mobile Edge Cloud Computing Systems
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Parallel Distributed Syst.2
2022 GOSH: Task Scheduling Using Deep Surrogate Models in Fog Computing Environments
abstract
Recently, intelligent scheduling approaches using surrogate models have been proposed to efficiently allocate volatile tasks in heterogeneous fog environments. Advances like deterministic surrogate models, deep neural networks (DNN) and gradient-based optimization allow low energy consumption and response times to be reached. However, deterministic surrogate models, which estimate objective values for optimization, do not consider the uncertainties in the distribution of the Quality of Service (QoS) objective function that can lead to high Service Level Agreement (SLA) violation rates. Moreover, the brittle nature of DNN training and the limited exploration with low agility in gradient-based optimization prevent such models from reaching minimal energy or response times. To overcome these difficulties, we present a novel scheduler that we call GOSH for Gradient Based Optimization using Second Order derivatives and Heteroscedastic Deep Surrogate Models. GOSH uses a second-order gradient based optimization approach to obtain better QoS and reduce the number of iterations to converge to a scheduling decision, subsequently lowering the scheduling time. Instead of a vanilla DNN, GOSH uses a Natural Parameter Network (NPN) to approximate objective scores. Further, a Lower Confidence Bound (LCB) optimization approach allows GOSH to find an optimal trade-off between greedy minimization of the mean latency and uncertainty reduction by employing error-based exploration. Thus, GOSH and its co-simulation based extension GOSH*, can adapt quickly and reach better objective scores than baseline methods. We show that GOSH* reaches better objective scores than GOSH, but it is suitable only for high resource availability settings, whereas GOSH is apt for limited resource settings. Real system experiments for both GOSH and GOSH* show significant improvements against the state-of-the-art in terms of energy consumption, response time and SLA violations by up to 18, 27 and 82 percent, respectively.
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Parallel Distributed Syst.2
2022 COSCO: Container Orchestration Using Co-Simulation and Gradient Based Optimization for Fog Computing Environments
abstract
Intelligent task placement and management of tasks in large-scale fog platforms is challenging due to the highly volatile nature of modern workload applications and sensitive user requirements of low energy consumption and response time. Container orchestration platforms have emerged to alleviate this problem with prior art either using heuristics to quickly reach scheduling decisions or AI driven methods like reinforcement learning and evolutionary approaches to adapt to dynamic scenarios. The former often fail to quickly adapt in highly dynamic environments, whereas the latter have run-times that are slow enough to negatively impact response time. Therefore, there is a need for scheduling policies that are both reactive to work efficiently in volatile environments and have low scheduling overheads. To achieve this, we propose a Gradient Based Optimization Strategy using Back-propagation of gradients with respect to Input (GOBI). Further, we leverage the accuracy of predictive digital-twin models and simulation capabilities by developing a Coupled Simulation and Container Orchestration Framework (COSCO). Using this, we create a hybrid simulation driven decision approach, GOBI*, to optimize Quality of Service (QoS) parameters. Co-simulation and the back-propagation approaches allow these methods to adapt quickly in volatile environments. Experiments conducted using real-world data on fog applications using the GOBI and GOBI* methods, show a significant improvement in terms of energy consumption, response time, Service Level Objective and scheduling time by up to 15, 40, 4, and 82 percent respectively when compared to the state-of-the-art algorithms.
Shreshth Tuli, Shivananda R. Poojara, Satish Narayana Srirama, Giuliano Casale, Nicholas R. Jennings
IEEE Trans. Parallel Distributed Syst.4
2021 RDOF: Deployment Optimization for Function as a Service
abstract
Function as a service (FaaS) simplifies the runtime resource management of cloud applications and enables fine-grained scaling and billing at the function level, thus becoming the most widespread serverless paradigm today. Cost-effective use of FaaS entails appropriately deploying individual functions. We propose in this paper RDOF1, a model-driven approach to deployment optimization for FaaS. RDOF predicts the performance of a FaaS-based application by instantiating a layered queueing network and finds the optimal configuration of each function such that the total operating cost is minimized under the specified performance requirements. We have validated RDOF on Amazon Web Services (AWS) and implemented it in an online tool that operates on TOSCA metamodels.
Lulai Zhu, Giorgos Giotis, Vasilios Tountopoulos, Giuliano Casale
CLOUD4
2021 MEAD: Model-Based Vertical Auto-Scaling for Data Stream Processing
abstract
The unpredictable variability of Data Stream Processing (DSP) application workloads calls for advanced mechanisms and policies for elastically scaling the processing capacity of DSP operators. Whilst many different approaches have been used to devise policies, most of the solutions have focused on data arrival rate and operator resource utilization as key metrics for auto-scaling. We here show that, under burstiness in the data flows, overly simple characterizations of the input stream can yet lead to very inaccurate performance estimations that affect such policies, resulting in sub-optimal resource allocation.We then present MEAD, a vertical auto-scaling solution that relies on online state-based representation of burstiness to drive resource allocation. We use in particular Markovian Arrival Processes (MAPs), which are composable with analytical queueing models, allowing us to efficiently predict performance at run-time under burstiness. We integrate MEAD in Apache Flink, and evaluate its benefits over simpler yet popular auto-scaling solutions, using both synthetic and real-world workloads. Differently from existing approaches, MEAD satisfies response time requirements under burstiness, while saving up to 50% CPU resources with respect to a static allocation.
Gabriele Russo Russo, Valeria Cardellini, Giuliano Casale, Francesco Lo Presti
CCGRID3
2021 Deep Learning Models for Automated Identification of Scheduling Policies
abstract
Queueing network models are commonly used as performance models of distributed software applications and service-based systems. Although several methods exist for learning their parameters, such as demand estimation methods, little research has been carried out in the literature on automatically identifying scheduling policies from empirical datasets. Scheduling policies and their parameters have an impact on the model’s stationary distribution in general, thus their correct determination is important for model accuracy. They are particularly relevant for correctly estimating percentiles and higher-order moments of performance indexes such as response times. We propose a deep learning technique based on transformer models - a common technique in natural language processing, to address the lack of methods for this parameter identification problem. From a sample path of the joint network state, or an aggregate thereof, our approach can classify the scheduling policy of the stations in a queueing network. We show that the transformer model delivers good-classification precision and recall, improving significantly over support vector machines or simpler recurrent neural networks.
Yichong Chen, Giuliano Casale
MASCOTS2
2021 A Mixture Density Network Approach to Predicting Response Times in Layered Systems
abstract
Layering is a common feature in modern service-based systems. The characterization of response times in a layered system is an important but challenging analysis dimension in Quality of Service (QoS) assessment. In this paper, we develop a novel approach to estimate the mean and variance of response time in systems that may be abstracted as layered queueing networks. The core step of the method is to obtain the response time distributions in the submodels that are used to analyze the layered queueing networks by means of decomposition. We model the conditional response time distribution as a mixture of Gamma density functions for which we learn the parameters by means of a Mixture Density Network (MDN). The scheme recursively propagates the MDN predictions through the layers using phase-type distributions and performs convolutions to gain the approximation of the system delay. The experimental results show an accurate match between simulations and MDN predictions and also verify the effectiveness of the approach.
Zifeng Niu, Giuliano Casale
MASCOTS2
2021 Facilitating load-dependent queueing analysis through factorization
Giuliano Casale, Peter G. Harrison, Wai Hong Ong
Perform. Evaluation1
2021 Guest Editorial: Special Section on Embracing Artificial Intelligence for Network and Service Management
abstract
Artificial Intelligence (AI) has the potential to leverage the immense amount of operational data of clouds, services, and social and communication networks. As a concrete example, AI techniques have been adopted by telcom operators to develop virtual assistants based on advances in natural language processing (NLP) for interaction with customers and machine learning (ML) to enhance the customer experience by improving customer flow. Machine learning has also been applied to finding fraud patterns which enables operators to focus on dealing with the activity as opposed to the previous focus on detecting fraud.
Hanan Lutfiyya, Robert Birke, Giuliano Casale, Amogh Dhamdhere, Jinho Hwang, Takeru Inoue, Neeraj Kumar 0001, Deepak Puthal, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.3
2021 Guest Editorial: Special Issue on Data Analytics and Machine Learning for Network and Service Management - Part II
abstract
Network and Service analytics can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, analytics and machine learning have found ways to improve reliability, configuration, performance, fault and security management. In particular, we see a growing trend towards using machine learning, artificial intelligence and data analytics to improve operations and management of information technology services, systems and networks.
Nur Zincir-Heywood, Giuliano Casale, David Carrera 0001, Lydia Y. Chen, Amogh Dhamdhere, Takeru Inoue, Hanan Lutfiyya, Taghrid Samak
IEEE Trans. Netw. Serv. Manag.2
2021 Performance Analysis Methods for List-Based Caches With Non-Uniform Access
abstract
List-based caches can offer lower miss rates than single-list caches, but their analysis is challenging due to state space explosion. In this setting, we propose novel methods to analyze performance for a general class of list-based caches with tree structure, non-uniform access to items and lists, and random or first-in first-out replacement policies. Even though the underlying Markov process is shown to admit a product-form solution, this is difficult to exploit for large caches. Thus, we develop novel approximations for cache performance metrics, in particular by means of a singular perturbation method and a refined mean field approximation. We compare the accuracy of these approaches to simulations, finding that our new methods rapidly converge to the equilibrium distribution as the number of items and the cache capacity grow in a fixed ratio. We find that they are much more accurate than fixed point methods similar to prior work, with mean average errors typically below 1.5% even for very small caches. Our models are also generalized to account for synchronous requests, fetch latency, and item sizes, extending the applicability of approximations for list-based caches.
Giuliano Casale, Nicolas Gast
IEEE/ACM Trans. Netw.1
2020 COCOA: Cold Start Aware Capacity Planning for Function-as-a-Service Platforms
abstract
Function-as-a-Service (FaaS) has become increasingly popular in the software industry due to the implied cost-savings in event-driven workloads and its synergy with DevOps. To size an on-premise FaaS platform, it is important to estimate the required CPU and memory capacity to serve the expected loads. Given the service-level agreements, it is however challenging to take the cold start issue into account during the sizing process. We have investigated the similarity of this problem with the hit rate improvement problem in Time to Live (TTL) caches and concluded that solutions for TTL cache, although potentially applicable, lead to over-provisioning in FaaS. Thus, we propose a novel approach, COCOA, to solve this issue. COCOA uses a queueing-based approach to assess the effect of cold starts on FaaS response times. It also considers different memory consumption values depending on whether the function is idle or in execution. Using an event-driven FaaS simulator, FaasSim, that we have developed, we show that COCOA can reduce overprovisioning by over 70% under some of the workloads we have considered, while satisfying the service-level agreements.
Alim Ul Gias, Giuliano Casale
MASCOTS2
2020 Guest editor's forewords: Special issue on Valuetools 2017
Andrea Marin, Giuliano Casale, Dorina C. Petriu, Sabina Rossi
Perform. Evaluation2
2020 Fluid approximation of closed queueing networks with discriminatory processor sharing
Lulai Zhu, Giuliano Casale, Iker Perez
Perform. Evaluation2
2020 Guest Editorial: Special Section on Data Analytics and Machine Learning for Network and Service Management-Part I
Nur Zincir-Heywood, Giuliano Casale, David Carrera 0001, Lydia Y. Chen, Amogh Dhamdhere, Takeru Inoue, Hanan Lutfiyya, Taghrid Samak
IEEE Trans. Netw. Serv. Manag.2
2019 ATOM: Model-Driven Autoscaling for Microservices
abstract
Microservices based architectures are increasingly widespread in the cloud software industry. Still, there is a shortage of auto-scaling methods designed to leverage the unique features of these architectures, such as the ability to independently scale a subset of microservices, as well as the ease of monitoring their state and reciprocal calls. We propose to address this shortage with ATOM, a model-driven autoscaling controller for microservices. ATOM instantiates and solves at run-time a layered queueing network model of the application. Computational optimization is used to dynamically control the number of replicas for each microservice and its associated container CPU share, overall achieving a fine-grained control of the application capacity at run-time. Experimental results indicate that for heavy workloads ATOM offers around 30%-37% higher throughput than baseline model-agnostic controllers based on simple static rules. We also find that model-driven reasoning reduces the number of actions needed to scale the system as it reduces the number of bottleneck shifts that we observe with model-agnostic controllers.
Alim Ul Gias, Giuliano Casale, C. Murray Woodside
ICDCS2
2019 Holistic resource management for sustainable and reliable cloud computing: An innovative solution to global challenge
Sukhpal Singh, Peter Garraghan, Vlado Stankovski, Giuliano Casale, Ruppa K. Thulasiram, Soumya K. Ghosh 0001, Kotagiri Ramamohanarao, Rajkumar Buyya
J. Syst. Softw.4
2019 Guest Editorial: Special Issue on Novel Techniques in Big Data Analytics for Management
abstract
Cloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, configuration, performance, fault and security management. In particular, we see a growing trend towards using statistical analysis, Artificial Intelligence (AI) and machine learning to improve operations and management of IT systems and networks.
David Carrera 0001, Giuliano Casale, Takeru Inoue, Hanan Lutfiyya, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.2
2018 Analyzing Replacement Policies in List-Based Caches with Non-Uniform Access Costs
abstract
List-based caches can offer lower miss rates than single-list caches, but their analysis is challenging due to state-space explosion. We analyze in this setting randomized replacement policies for caches with non-uniform access costs. In our model, costs can depend on the stream a request originated from, the target item, and the list that contains it. We first show that, similarly to the uniform-cost case, the random replacement (RR) and first-in first-out (FIFO) policies can be exactly analyzed using a product-form expression for the equilibrium state probabilities of the cache. We then tackle the state space explosion by means of the singular perturbation method, deriving limiting expressions for the equilibrium performance measures as the number of items and the cache capacity grow in a fixed ratio. Simulations indicate that our asymptotic formulas rapidly converge to the cache equilibrium distribution.
Giuliano Casale
INFOCOM1
2018 PAX: Partition-aware autoscaling for the Cassandra NoSQL database
abstract
Apache Cassandra has emerged as one of the most widely adopted NoSQL databases. However, there is still a limited understanding on how to optimally operate Cassandra in the cloud using autoscaling methods, by which resources can be scaled up or down to reduce operational costs and meet service-level objectives (SLOs). To address this limitation, we present PAX, a partition-aware elastic resource management system for Apache Cassandra. PAX uses low-overhead query sampling and knowledge of the data-partitioning across the nodes to automatically adapt capacity in Cassandra clusters. Differently from existing autoscaling methods for Cassandra, which incur large acquisition times for new nodes, PAX exploits Cassandra's hinted handoff mechanism and a shared hints storage to minimize the time needed to acquire a node into the cluster. We propose a reactive and a proactive implementation of PAX and compare their performance against different workloads with varying intensities and item popularity distributions, finding that the proactive version significantly reduces SLO violations.
Salvatore Dipietro, Rajkumar Buyya, Giuliano Casale
NOMS3
2018 Guest Editorial: Special Section on Advances in Big Data Analytics for Management
abstract
Cloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, automated configuration, performance, and optimized network management in general. In this area, we have witnessed a growing trend towards using statistical analysis and machine learning techniques to improve operations and management of IT systems and networks.
Giuliano Casale, Yixin Diao, Marco Mellia, Rajiv Ranjan 0001, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.1
2017 How to Supercharge the Amazon T2: Observations and Suggestions
abstract
Cloud service providers adopt a credit system to allow users to obtain periods of performance bursts without additional cost. For example, the Amazon EC2 T2 instance offers low baseline performance and the capability to achieve short periods of high performance using CPU credits. Once a T2 instance is created and assigned some initial credits, while its CPU utilization is above the baseline threshold, there is a transient period where performance is boosted and the assigned CPU credits are used. After all credits are used, the maximum achievable performance drops to baseline. Credits accrue periodically, when the instance utilization is below the baseline threshold. This paper proposes a methodology to increase the performance benefits of T2 by seamlessly extending the duration of the transient period while maintaining high performance. This extension of the high performance transient period is combined with proactive migration to further take advantage of the initially assigned credits. We conduct experiments to demonstrate the benefits of this methodology for both single-tier and multi-tier applications.
Feng Yan 0001, Lihua Ren, Daniel J. Dubois, Giuliano Casale, Jiawei Wen, Evgenia Smirni
CLOUD4
2017 Energy-efficient resource allocation and provisioning for in-memory database clusters
abstract
Systems for processing large scale analytical workloads are increasingly moving from on-premise setups to on-demand configurations deployed on scalable cloud infrastructures. To reduce the cost of such infrastructures, existing research focuses on developing novel methods for workload and server consolidation. In this paper, we combine analytical modeling and non-linear optimization to help cloud providers increase the energy-efficiency of in-memory database clusters in cloud environments. We model this scenario as a multi-dimensional bin-packing problem and propose a new approach based on a hybrid genetic algorithm that efficiently handles resource allocation and server assignment for a given set of in-memory databases. Our trace-driven evaluation is based on measurements from an SAP HANA in-memory system and indicates improvements between 6% and 32% over the popular best-fit decreasing heuristic.
Karsten Molka, Giuliano Casale
IM2
2017 Generalized Synchronizations and Capacity Constraints for Java Modelling Tools
abstract
Java Modelling Tools (JMT) is a suite of performance evaluation tools based on queueing network models. Recently {\em JSIMgraph}, the JMT discrete-event simulation tool, has been extended to express features of current computing systems such as Big data applications. The goal of this demonstration is to showcase novel support in JMT for fork-join synchronization, dynamic scaling of parallelism levels, memory and group capacity constraints.
Giuliano Casale, Mattia Cazzoli, Vitor S. Lopes, Giuseppe Serazzi, Lulai Zhu
ICPE1
2017 Line: Evaluating Software Applications in Unreliable Environments
abstract
Cloud computing has paved the way to the flexible deployment of software applications. This flexibility offers service providers a number of options to tailor their deployments to the observed and foreseen customer workloads, without incurring in large capital costs. However, cloud deployments pose novel challenges regarding application reliability and performance. Examples include managing the reliability of deployments that make use of spot instances, or coping with the performance variability caused by multiple tenants in a virtualized environment. In this paper, we introduce Line, a tool for performance and reliability analysis of software applications. Line solves layered queueing network (LQN) models, a popular class of stochastic models in software performance engineering, by setting up and solving an associated system of ordinary differential equations. A key differentiator of Line compared to existing solvers for LQNs is that Line incorporates a model of the environment the application operates in. This enables the modeling of reliability and performance issues such as resource failures, server breakdowns and repairs, slow start-up times, resource interference due to multitenancy, among others. This paper describes the Line tool, its support for performance and reliability modeling, and illustrates its potential by comparing Line predictions against data obtained from a cloud deployment. We also illustrate the applicability of Line with a case study on reliability-aware resource provisioning.
Juan F. Pérez, Giuliano Casale
IEEE Trans. Reliab.2
2016 Model-Driven Application Refactoring to Minimize Deployment Costs in Preemptible Cloud Resources
abstract
Performance assessment of cloud-based applications requires new methodologies to deal with the complexity of software systems and the variability of cloud resources. In this paper, we address the problem of reducing the total costs for running cloud-based applications while fulfilling service-level objectives (SLOs). To this end, we define an approach to refactor a cloud application in such a way that, when it is deployed, it requires less computational capacity and therefore less resources. We experimented our approach on top of a modified optimal provisioning heuristic designed for preemptible cloud resources and the results show that it reduces deployment costs, up to 60% when compared to the same approach, but without model-driven application refactoring.
Daniel J. Dubois, Catia Trubiani, Giuliano Casale
CLOUD3
2016 An Uncertainty-Aware Approach to Optimal Configuration of Stream Processing Systems
abstract
Finding optimal configurations for Stream Processing Systems (SPS) is a challenging problem due to the large number of parameters that can influence their performance and the lack of analytical models to anticipate the effect of a change. To tackle this issue, we consider tuning methods where an experimenter is given a limited budget of experiments and needs to carefully allocate this budget to find optimal configurations. We propose in this setting Bayesian Optimization for Configuration Optimization (BO4CO), an auto-tuning algorithm that leverages Gaussian Processes (GPs) to iteratively capture posterior distributions of the configuration spaces and sequentially drive the experimentation. Validation based on Apache Storm demonstrates that our approach locates optimal configurations within a limited experimental budget, with an improvement of SPS performance typically of at least an order of magnitude compared to existing configuration algorithms.
Pooyan Jamshidi, Giuliano Casale
MASCOTS2
2016 Efficient Memory Occupancy Models for In-memory Databases
abstract
Predicting memory occupancy during the execution of large-scale analytical workloads becomes critical for in-memory databases. In particular, probabilistic performance measures for such systems are of interest, but difficult to model with analytical methods due to the highly variable threading levels in corresponding workloads. Since literature with queueing theoretic background largely ignores the memory modeling part, we propose a new probabilistic model to capture the memory occupancy distribution in such systems. We further combine this model with our analytical formulation TP-AMVA for greater efficiency compared to simulation and evaluate against experiments using SAP HANA.
Karsten Molka, Giuliano Casale
MASCOTS2
2016 Quantifying the Impact of Replication on the Quality-of-Service in Cloud Databases
abstract
Cloud databases achieve high availability by automatically replicating data on multiple nodes. However, the overhead caused by the replication process can lead to an increase in the mean and variance of transaction response times, causing unforeseen impacts on the offered quality-of-service (QoS). In this paper, we propose a measurement-driven methodology to predict the impact of replication on Database-as-a-Service (DBaaS) environments. Our methodology uses operational data to parameterize a closed queueing network model of the database cluster together with a Markov model that abstracts the dynamic replication process. Experiments on Amazon RDS show that our methodology predicts response time mean and percentiles with errors of just 1% and 15% respectively, and under operational conditions that are significantly different from the ones used for model parameterization. We show that our modeling approach surpasses standard modeling methods and illustrate the applicability of our methodology for automated DBaaS provisioning.
Rasha Osman, Juan F. Pérez, Giuliano Casale
QRS3
2016 Automated Parameterization of Performance Models from Measurements
abstract
Estimating parameters of performance models from empirical measurements is a critical task, which often has a major influence on the predictive accuracy of a model. This tutorial presents the problem of parameter estimation in queueing systems and queueing networks. The focus is on reliable estimation of the arrival rates of the requests and of the service demands they place at the servers. The tutorial covers common estimation techniques such as regression methods, maximum-likelihood estimation, and moment-matching, discussing their sensitivity with respect to data and model characteristics. The tutorial also demonstrates the automated estimation of model parameters using new open source tools.
Giuliano Casale, Simon Spinner, Weikun Wang
ICPE1
2016 Maximum Likelihood Estimation of Closed Queueing Network Demands from Queue Length Data
abstract
Resource demand estimation is essential for the application of analyical models, such as queueing networks, to real-world systems. In this paper, we investigate maximum likelihood (ML) estimators for service demands in closed queueing networks with load-independent and load-dependent service times. Stemming from a characterization of necessary conditions for ML estimation, we propose new estimators that infer demands from queue-length measurements, which are inexpensive metrics to collect in real systems. One advantage of focusing on queue-length data compared to response times or utilizations is that confidence intervals can be rigorously derived from the equilibrium distribution of the queueing network model. Our estimators and their confidence intervals are validated against simulation and real system measurements for a multi-tier application.
Weikun Wang, Giuliano Casale, Ajay Kattepur, Manoj Nambiar 0001
ICPE2
2016 Guest Editors' Introduction: Special Issue on Big Data Analytics for Management
abstract
Cloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, configuration, performance, and security management. In particular, we see a growing trend towards using statistical analysis and machine learning to improve operations and management of IT systems and networks.
Giuliano Casale, Yixin Diao, Hanan Lutfiyya, Philippe Owezarski, Danny Raz
IEEE Trans. Netw. Serv. Manag.1
2015 Less Can Be More: Micro-managing VMs in Amazon EC2
abstract
Micro instances (t1. micro) are the class of Amazon EC2 virtual machines (VMs) offering the lowest operational costs for applications with short bursts in their CPU requirements. As processing proceeds, EC2 throttles CPU capacity of micro instances in a complex, unpredictable, manner. This paper aims at making micro instances more predictable and efficient to use. First, we present a characterization of EC2 micro instances that evaluates the complex interactions between cost, performance, idleness and CPU throttling. Next, we define adaptive algorithms to manage CPU consumption by learning the workload characteristics at runtime and by injecting idleness to diminish host-level throttling. We show that a gradient-hill strategy leads to favorable results. For CPU bound workloads, we observe that a significant portion of jobs (up to 65%) can have end-to-end times that are even four times shorter than those of the more expensive m1. small class. Our algorithms drastically reduce the long tails of job execution times on the micro instances, resulting to favorable comparisons against even small instances.
Jiawei Wen, Giuliano Casale, Evgenia Smirni
CLOUD3
2015 Experiments or simulation? A characterization of evaluation methods for in-memory databases
abstract
The recent growth of interest for in-memory databases poses the question on whether established prediction methods such as response surfaces and simulation are effective to describe the performance of these systems. In particular, the limited dependence of in-memory technologies on the disk makes methods such as simulation more appealing than in the past, since disks are difficult to simulate. To answer this question, we study an in-memory commercial solution, SAP HANA, deployed on a high-end server with 120 physical cores. First, we apply experimental design methods to generate response surfaces that describe database performance as a function of workload and hardware parameters. Next, we develop a class-switching queueing network model to predict in-memory database performance under similar scenarios. By comparing the applicability of the two approaches to modeling multi-tenancy, we find that both queueing and response surface models yield mean prediction errors in the range 5%-22% with respect to mean memory occupancy and response times, but the accuracy for the latter deteriorates in response surfaces as the number of experiments are reduced, whereas simulation is effective in all cases. This suggests that simulation can be very effective in performance prediction for in-memory database management.
Karsten Molka, Giuliano Casale
CNSM2
2015 DICE: Quality-Driven Development of Data-Intensive Cloud Applications
abstract
Model-driven engineering (MDE) often features quality assurance (QA) techniques to help developers creating software that meets reliability, efficiency, and safety requirements. In this paper, we consider the question of how quality-aware MDE should support data-intensive software systems. This is a difficult challenge, since existing models and QA techniques largely ignore properties of data such as volumes, velocities, or data location. Furthermore, QA requires the ability to characterize the behavior of technologies such as Hadoop/MapReduce, NoSQL, and stream-based processing, which are poorly understood from a modeling standpoint. To foster a community response to these challenges, we present the research agenda of DICE, a quality-aware MDE methodology for data-intensive cloud applications. DICE aims at developing a quality engineering tool chain offering simulation, verification, and architectural optimization for Big Data applications. We overview some key challenges involved in developing these tools and the underpinning models.
Giuliano Casale, Danilo Ardagna, Matej Artac, Franck Barbier, Elisabetta Di Nitto, Alexis Henry, Gabriel Iuhasz, Christophe Joubert, José Merseguer, Victor Ion Munteanu, Juan F. Pérez, Dana Petcu, Matteo G. Rossi, Craig Sheridan, Ilias Spais, Daniel Vladuic
MiSE@ICSE1
2015 QD-AMVA: Evaluating systems with queue-dependent service requirements
abstract
Workload measurements in enterprise systems often lead to observe a dependence between the number of requests running at a resource and their mean service requirements. However, multiclass performance models that feature these dependences are challenging to analyze, a fact that discourages practitioners from characterizing workload dependences. We here focus on closed multiclass queueing networks and introduce QD-AMVA, the first approximate mean-value analysis (AMVA) algorithm that can efficiently and robustly analyze queue-dependent service times in a multiclass setting. A key feature of QD-AMVA is that it operates on mean values, avoiding the computation of state probabilities. This property is an innovative result for state-dependent models, which increases the computational efficiency and numerical robustness of their evaluation. Extensive validation on random examples, a cloud load-balancing case study and comparison with a fluid method and an existing AMVA approximation prove that QD-AMVA is efficient, robust and easy to apply, thus enhancing the tractability of queue-dependent models.
Giuliano Casale, Juan F. Pérez, Weikun Wang
Perform. Evaluation1
2015 Evaluating approaches to resource demand estimation
Simon Spinner, Giuliano Casale, Fabian Brosig, Samuel Kounev
Perform. Evaluation2
2015 Estimating Computational Requirements in Multi-Threaded Applications
abstract
Performance models provide effective support for managing quality-of-service (QoS) and costs of enterprise applications. However, expensive high-resolution monitoring would be needed to obtain key model parameters, such as the CPU consumption of individual requests, which are thus more commonly estimated from other measures. However, current estimators are often inaccurate in accounting for scheduling in multi-threaded application servers. To cope with this problem, we propose novel linear regression and maximum likelihood estimators. Our algorithms take as inputs response time and resource queue measurements and return estimates of CPU consumption for individual request types. Results on simulated and real application datasets indicate that our algorithms provide accurate estimates and can scale effectively with the threading levels.
Juan F. Pérez, Giuliano Casale, Sergio Pacheco-Sanchez
IEEE Trans. Software Eng.2
2014 Memory-aware sizing for in-memory databases
abstract
In-memory database systems are among the technological drivers of big data processing. In this paper we apply analytical modeling to enable efficient sizing of in-memory databases. We present novel response time approximations under online analytical processing workloads to model thread-level fork-join and per-class memory occupation.We combine these approximations with a non-linear optimization program to minimize memory swapping in in-memory database clusters. We compare our approach with state-of-the-art response time approximations and trace-driven simulation using real data from an SAP HANA in-memory system and show that our optimization model is significantly more accurate than existing approaches at similar computational costs.
Karsten Molka, Giuliano Casale, Thomas Molka, Laura Moore
NOMS2
2014 LibReDE: a library for resource demand estimation
abstract
When creating a performance model, it is necessary to quantify the amount of resources consumed by an application serving individual requests. In distributed enterprise systems, these resource demands usually cannot be observed directly, their estimation is a major challenge. Different statistical approaches to resource demand estimation based on monitoring data have been proposed, e.g., using linear regression or Kalman filtering techniques. In this paper, we present LibReDE, a library of ready-to-use implementations of approaches to resource demand estimation that can be used for online and offline analysis. It is the first publicly available tool for this task and aims at supporting performance engineers during performance model construction. The library enables the quick comparison of the estimation accuracy of different approaches in a given context and thus helps to select an optimal one.
Simon Spinner, Giuliano Casale, Xiaoyun Zhu, Samuel Kounev
ICPE2
2014 Special Issue on "Quantitative Evaluation of SysTems" (QEST 2012)
Giuliano Casale, Ludmila Cherkasova, Holger Hermanns
Perform. Evaluation1
2014 Blending randomness in closed queueing network models
Giuliano Casale, Mirco Tribastone, Peter G. Harrison
Perform. Evaluation1
2013 A Feasibility Study of Host-Level Contention Detection by Guest Virtual Machines
abstract
We investigate the feasibility of detecting host-level CPU contention from inside a guest virtual machine (VM). Our methodology involves running benchmarks with deterministic and randomized execution times inside a guest VM in a private cloud testbed. Simultaneously, using the recently proposed COCOMA tool, we expose the guest VM to host-level CPU stealing events of increasing intensity. This leads us to observe that the use of hyper-threading in the host can hinder detection of CPU contention, which otherwise can be done accurately using the CPU steal metric. For systems where hyper-threading is enabled, we investigate the performance of some basic detection algorithms. We find that thresholding often outperforms more sophisticated statistical tests.
Giuliano Casale, Carmelo Ragusa, Panos Parpas
CloudCom (2)1
2013 Fitting second-order acyclic Marked Markovian Arrival Processes
abstract
Markovian Arrival Processes (MAPs) are a tractable class of point-processes useful to model correlated time series, such as those commonly found in network traces and system logs used in performance analysis and reliability evaluation. Marked MAPs (MMAPs) generalize MAPs by further allowing the modeling of multi-class traces, possibly with cross-correlation between multi-class arrivals. In this paper, we present analytical formulas to fit second-order acyclic MMAPs with an arbitrary number of classes. We initially define closed-form formulas to fit second-order MMAPs with two classes, where the underlying MAP is in canonical form. Our approach leverages forward and backward moments, which have recently been defined, but never exploited jointly for fitting. Then, we show how to sequentially apply these formulas to fit an arbitrary number of classes. Representative examples and trace-driven simulation using storage traces show the effectiveness of our approach for fitting empirical datasets.
Andrea Sansottera, Giuliano Casale, Paolo Cremonesi
DSN2
2013 RPO: Runtime web server optimization under simultaneous multithreading
Samira Musabbir, Diwakar Krishnamurthy, Giuliano Casale
IM3
2013 An Offline Demand Estimation Method for Multi-threaded Applications
abstract
Parameterizing performance models for multi-threaded enterprise applications requires finding the service rates offered by worker threads to the incoming requests. Statistical inference on monitoring data is here helpful to reduce the overheads of application profiling and to infer missing information. While linear regression of utilization data is often used to estimate service rates, it suffers erratic performance and also ignores a large part of application monitoring data, e.g., response times. Yet inference from other metrics, such as response times or queue-length samples, is complicated by the dependence on scheduling policies. To address these issues, we propose novel scheduling-aware estimation approaches for multi-threaded applications based on linear regression and maximum likelihood estimators. The proposed methods estimate demands from samples of the number of requests in execution in the worker threads at the admission instant of a new request. Validation results are presented on simulated and real application datasets for systems with multi-class requests, class switching, and admission control.
Juan F. Pérez, Sergio Pacheco-Sanchez, Giuliano Casale
MASCOTS3
2013 Bayesian Service Demand Estimation Using Gibbs Sampling
abstract
Performance modelling of web applications involves the task of estimating service demands of requests at physical resources, such as CPUs. In this paper, we propose a service demand estimation algorithm based on a Markov Chain Monte Carlo (MCMC) technique, Gibbs sampling. Our methodology is widely applicable as it requires only queue length samples at each resource, which are simple to measure. Additionally, since we use a Bayesian approach, our method can use prior information on the distribution of parameters, a feature not always available with existing demand estimation approaches. The main challenge of Gibbs sampling is to efficiently evaluate the conditional expression required to sample from the posterior distribution of the demands. This expression is shown to be the equilibrium solution of a multiclass closed queueing network. We define a novel approximation to efficiently obtain the normalising constant to make the cost of its evaluation acceptable for MCMC applications. Experimental evaluation based on simulation data with different model sizes demonstrates the effectiveness of Gibbs sampling for service demand estimation.
Weikun Wang, Giuliano Casale
MASCOTS2
2013 Heavy-traffic revenue maximization in parallel multiclass queues
Jonatha Anselmi, Giuliano Casale
Perform. Evaluation2
2013 Performance models of storage contention in cloud environments
abstract
We propose simple models to predict the performance degradation of disk requests due to storage device contention in consolidated virtualized environments. Model parameters can be deduced from measurements obtained inside Virtual Machines (VMs) from a system where a single VM accesses a remote storage server. The parameterized model can then be used to predict the effect of storage contention when multiple VMs are consolidated on the same server. We first propose a trace-driven approach that evaluates a queueing network with fair share scheduling using simulation. The model parameters consider Virtual Machine Monitor level disk access optimizations and rely on a calibration technique. We further present a measurement-based approach that allows a distinct characterization of read/write performance attributes. In particular, we define simple linear prediction models for I/O request mean response times, throughputs and read/write mixes, as well as a simulation model for predicting response time distributions. We found our models to be effective in predicting such quantities across a range of synthetic and emulated application workloads.
Stephan Kraft, Giuliano Casale, Diwakar Krishnamurthy, Des Greer, Peter Kilpatrick
Softw. Syst. Model.2
2012 WIQ: Work-Intensive Query Scheduling for In-Memory Database Systems
abstract
We propose a novel admission control policy for database queries. Our methodology uses system measurements of CPU utilization and query backlogs to determine interference between queries in execution on the same database server. Query interference may arise due to the concurrent access of hardware and software resources and can affect performance in positive and negative ways. Specifically our admission control considers the mix of jobs in service and prioritizes the query classes consuming CPU resources more efficiently. The policy ignores I/O subsystems and is therefore highly appropriate for in-memory databases. We validate our approach in trace-driven simulation and show performance increases of query slowdowns and throughputs compared to first-come first-served and shortest expected processing time first scheduling. Simulation experiments are parameterized from system traces of a SAP HANA in-memory database installation with TPC-H type workloads.
Stephan Kraft, Giuliano Casale, Alin Jula, Peter Kilpatrick, Des Greer
IEEE CLOUD2
2012 MODAClouds: a model-driven approach for the design and execution of applications on multiple clouds
abstract
Cloud computing is emerging as a major trend in the ICT industry. While most of the attention of the research community is focused on considering the perspective of the Cloud providers, offering mechanisms to support scaling of resources and interoperability and federation between Clouds, the perspective of developers and operators willing to choose the Cloud without being strictly bound to a specific solution is mostly neglected. We argue that Model-Driven Development can be helpful in this context as it would allow developers to design software systems in a cloud-agnostic way and to be supported by model transformation techniques into the process of instantiating the system into specific, possibly, multiple Clouds. The MODAClouds (MOdel-Driven Approach for the design and execution of applications on multiple Clouds) approach we present here is based on these principles and aims at supporting system developers and operators in exploiting multiple Clouds for the same system and in migrating (part of) their systems from Cloud to Cloud as needed. MODAClouds offers a quality-driven design, development and operation method and features a Decision Support System to enable risk analysis for the selection of Cloud providers and for the evaluation of the Cloud adoption impact on internal business processes. Furthermore, MODAClouds offers a run-time environment for observing the system under execution and for enabling a feedback loop with the design environment. This allows system developers to react to performance fluctuations and to re-deploy applications on different Clouds on the long term.
Danilo Ardagna, Elisabetta Di Nitto, Giuliano Casale, Dana Petcu, Parastoo Mohagheghi, Sébastien Mosser 0001, Peter Matthews, Anke Gericke, Cyril Ballagny, Francesco D'Andria, Cosmin-Septimiu Nechifor, Craig Sheridan
MiSE3
2012 A class of tractable models for run-time performance evaluation
abstract
Run-time resource allocation requires the availability of system performance models that are both accurate and inexpensive to solve. We here propose a new methodology for run-time performance evaluation based on a class of closed queueing networks. Compared to exponential product-form models, the proposed queueing networks also support the inclusion of resources having first-come first-served scheduling under non-exponential service times. Motivated by the lack of an exact solution for these networks, we propose a fixed-point algorithm that approximates performance indexes in linear time and linear space with respect to the number of requests considered in the model. Numerical evaluation shows that, compared to simulation, the proposed models solved by fixed-point iteration have errors of about 1%-6%, while, on the same test cases, exponential product-form models suffer errors even in excess of 100%. Execution times on commodity hardware are of the order of a few seconds or less, making the proposed methodology practical for run-time decision-making.
Giuliano Casale, Peter G. Harrison
ICPE1
2012 ASIdE: Using Autocorrelation-Based Size Estimation for Scheduling Bursty Workloads
abstract
Temporal dependence in workloads creates peak congestion that can make service unavailable and reduce system performance. To improve system performability under conditions of temporal dependence, a server should quickly process bursts of requests that may need large service demands. In this paper, we propose and evaluateASIdE, an Autocorrelation-based SIze Estimation, that selectively delays requests which contribute to the workload temporal dependence. ASIdE implicitly approximates the shortest job first (SJF) scheduling policy but without any prior knowledge of job service times. Extensive experiments show that (1) ASIdE achieves good service time estimates from the temporal dependence structure of the workload to implicitly approximate the behavior of SJF; and (2) ASIdE successfully counteracts peak congestion in the workload and improves system performability under a wide variety of settings. Specifically, we show that system capacity under ASIdE is largely increased compared to the first-come first-served (FCFS) scheduling policy and is highly-competitive with SJF.
Ningfang Mi, Giuliano Casale, Evgenia Smirni
IEEE Trans. Netw. Serv. Manag.2
2012 BURN: Enabling Workload Burstiness in Customized Service Benchmarks
abstract
We introduce BURN, a methodology to create customized benchmarks for testing multitier applications under time-varying resource usage conditions. Starting from a set of preexisting test workloads, BURN finds a policy that interleaves their execution to stress the multitier application and generate controlled burstiness in resource consumption. This is useful to study, in a controlled way, the robustness of software services to sudden changes in the workload characteristics and in the usage levels of the resources. The problem is tackled by a model-based technique which first generates Markov models to describe resource consumption patterns of each test workload. Then, a policy is generated using an optimization program which sets as constraints a target request mix and user-specified levels of burstiness at the different resources in the system. Burstiness is quantified using a novel metric called overdemand, which describes in a natural way the tendency of a workload to keep a resource congested for long periods of time and across multiple requests. A case study based on a three-tier application testbed shows that our method is able to control and predict burstiness for session service demands at a fine-grained scale. Furthermore, experiments demonstrate that for any given request mix our approach can expose latency and throughput degradations not found with nonbursty workloads having the same request mix.
Giuliano Casale, Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia
IEEE Trans. Software Eng.1
2012 Dealing with Burstiness in Multi-Tier Applications: Models and Their Parameterization
abstract
Workloads and resource usage patterns in enterprise applications often show burstiness resulting in large degradation of the perceived user performance. In this paper, we propose a methodology for detecting burstiness symptoms in multi-tier applications but, rather than identifying the root cause of burstiness, we incorporate this information into models for performance prediction. The modeling methodology is based on the index of dispersion of the service process at a server, which is inferred by observing the number of completions within the concatenated busy times of that server. The index of dispersion is used to derive a Markov-modulated process that captures burstiness and variability of the service process at each resource well and that allows us to define queueing network models for performance prediction. Experimental results and performance model predictions are in excellent agreement and argue for the effectiveness of the proposed methodology under both bursty and nonbursty workloads. Furthermore, we show that the methodology extends to modeling flash crowds that create burstiness in the stream of requests incoming to the application.
Giuliano Casale, Ningfang Mi, Ludmila Cherkasova, Evgenia Smirni
IEEE Trans. Software Eng.1
2011 Markovian Workload Characterization for QoS Prediction in the Cloud
abstract
Resource allocation in the cloud is usually driven by performance predictions, such as estimates of the future incoming load to the servers or of the quality-of-service(QoS) offered by applications to end users. In this context, characterizing web workload fluctuations in an accurate way is fundamental to understand how to provision cloud resources under time-varying traffic intensities. In this paper, we investigate the Markovian Arrival Processes (MAP) and the related MAP/MAP/1 queueing model as a tool for performance prediction of servers deployed in the cloud. MAPs are a special class of Markov models used as a compact description of the time-varying characteristics of workloads. In addition, MAPs can fit heavy-tail distributions, that are common in HTTP traffic, and can be easily integrated within analytical queueing models to efficiently predict system performance without simulating. By comparison with traced riven simulation, we observe that existing techniques for MAP parameterization from HTTP log files often lead to inaccurate performance predictions. We then define a maximum likelihood method for fitting MAP parameters based on data commonly available in Apache log files, and a new technique to cope with batch arrivals, which are notoriously difficult to model accurately. Numerical experiments demonstrate the accuracy of our approach for performance prediction of web systems.
Sergio Pacheco-Sanchez, Giuliano Casale, Bryan W. Scotney, Sally I. McClean, Gerard P. Parr, Stephen Dawson
IEEE CLOUD2
2011 Approximate analysis of blocking queueing networks with temporal dependence
abstract
In this paper we extend the class of MAP queueing networks to include blocking models, which are useful to describe the performance of service instances which have a limited concurrency level. We consider two different blocking mechanisms: Repetitive Service-Random Destination (RS-RD) and Blocking After Service (BAS). We propose a methodology to evaluate MAP queueing networks with blocking based on the recently proposed Quadratic Reduction (QR), a state space transformation that decreases the number of states in the Markov chain underlying the queueing network model. From this reduced state space, we obtain boundable approximations on average performance indexes such as throughput, response time, utilizations. The two approximations that dramatically enhance the QR bounds are based on maximum entropy and on a novel minimum mutual information principle, respectively. Stress cases of increasing complexity illustrate the excellent accuracy of the proposed approximations on several models of practical interest.
Vittoria de Nitto Persone, Giuliano Casale, Evgenia Smirni
DSN2
2011 Introduction
Shirley Moore, Derrick Kondo, Brian J. N. Wylie, Giuliano Casale
Euro-Par (1)4
2011 Building accurate workload models using Markovian arrival processes
abstract
No abstract available.
Giuliano Casale
SIGMETRICS1
2011 Quantitative system evaluation with Java modeling tools
abstract
Java Modelling Tools (JMT) is a suite of open source applications for performance evaluation and workload characterization of computer and communication systems based on queueing networks. JMT includes tools for workload characterization (JWAT), solution of queueing networks with analytical algorithms (JMVA), simulation of general-purpose queueing models (JSIM), bottleneck identification (JABA), and teaching support for Markov chain models underlying queueing systems (JMCH). This tutorial summarizes the main features of the tools that compose the suite. Furthermore, using a composite case study, we provide intuition on the versatility of JMT in dealing with the different aspects of quality-of-service (QoS) evaluation, what-if analysis, and software performance tuning.
Giuliano Casale, Giuseppe Serazzi
ICPE1
2011 IO performance prediction in consolidated virtualized environments
abstract
We propose a trace-driven approach to predict the performance degradation of disk request response times due to storage device contention in consolidated virtualized environments. Our performance model evaluates a queueing network with fair share scheduling using trace-driven simulation. The model parameters can be deduced from measurements obtained inside Virtual Machines (VMs) from a system where a single VM accesses a remote storage server. The parameterized model can then be used to predict the effect of storage contention when multiple VMs are consolidated on the same virtualized server. The model parameter estimation relies on a search technique that tries to estimate the splitting and merging of blocks at the the Virtual Machine Monitor (VMM) level in the case of multiple competing VMs. Simulation experiments based on traces of the Postmark and FFSB disk benchmarks show that our model is able to accurately predict the impact of workload consolidation on VM disk IO response times.
Stephan Kraft, Giuliano Casale, Diwakar Krishnamurthy, Des Greer, Peter Kilpatrick
ICPE2
2011 A generalized method of moments for closed queueing networks
Giuliano Casale
Perform. Evaluation1
2011 Exact analysis of performance models by the Method of Moments
Giuliano Casale
Perform. Evaluation1
2010 CWS: a model-driven scheduling policy for correlated workloads
abstract
We define CWS, a non-preemptive scheduling policy for workloads with correlated job sizes. CWS tackles the scheduling problem by inferring the expected sizes of upcoming jobs based on the structure of correlations and on the outcome of past scheduling decisions. Size prediction is achieved using a class of Hidden Markov Models (HMM) with continuous observation densities that describe job sizes. We show how the forward-backward algorithm of HMMs applies effectively in scheduling applications and how it can be used to derive closed-form expressions for size prediction. This is particularly simple to implement in the case of observation densities that are phase-type (PH-type) distributed, where existing fitting methods for Markovian point processes may also simplify the parameterization of the HMM workload model.Based on the job size predictions, CWS emulates size-based policies which favor short jobs, with accuracy depending mainly on the HMM used to parametrize the scheduling algorithm. Extensive simulation and analysis illustrate that CWS is competitive with policies that assume exact information about the workload.
Giuliano Casale, Ningfang Mi, Evgenia Smirni
SIGMETRICS1
2010 Approximating passage time distributions in queueing models by Bayesian expansion
Giuliano Casale
Perform. Evaluation1
2010 Trace data characterization and fitting for Markov modeling
Giuliano Casale, Eddy Z. Zhang, Evgenia Smirni
Perform. Evaluation1
2010 KPC-Toolbox: Best recipes for automatic trace fitting using Markovian Arrival Processes
Giuliano Casale, Eddy Z. Zhang, Evgenia Smirni
Perform. Evaluation1
2010 Model-Driven System Capacity Planning under Workload Burstiness
abstract
In this paper, we define and study a new class of capacity planning models called MAP queueing networks. MAP queueing networks provide the first analytical methodology to describe and predict accurately the performance of complex systems operating under bursty workloads, such as multitier architectures or storage arrays. Burstiness is a feature that significantly degrades system performance and that cannot be captured explicitly by existing capacity planning models. MAP queueing networks address this limitation by describing computer systems as closed networks of servers whose service times are Markovian Arrival Processes (MAPs), a class of Markov-modulated point processes that can model general distributions and burstiness. In this paper, we show that MAP queueing networks provide reliable performance predictions even if the service processes are bursty. We propose a methodology to solve MAP queueing networks by two state space transformations, which we call Linear Reduction (LR) and Quadratic Reduction (QR). These transformations dramatically decrease the number of states in the underlying Markov chain of the queueing network model. From these reduced state spaces, we obtain two classes of bounds on arbitrary performance indexes, e.g., throughput, response time, and utilizations. Numerical experiments show that LR and QR bounds achieve good accuracy. We also illustrate the high effectiveness of the LR and QR bounds in the performance analysis of a real multitier architecture subject to TPC-W workloads that are characterized as bursty. These results promote MAP queueing networks as a new class of robust capacity planning models.
Giuliano Casale, Ningfang Mi, Evgenia Smirni
IEEE Trans. Computers1
2009 MAP-AMVA: Approximate mean value analysis of bursty systems
abstract
MAP queueing networks are recently proposed models for performance assessment of enterprise systems, such as multi-tier applications, where workloads are significantly affected by burstiness. Although MAP networks do not admit a simple product-form solution, performance metrics can be estimated accurately by linear programming bounds, yet these are expensive to compute under large populations. In this paper, we introduce an approximate mean value analysis (AMVA) approach to MAP network solution that significantly reduces the computational cost of model evaluation. We define a number of balance equations that relate mean performance indices such as utilizations and response times. We show that the quality of a MAP-AMVA solution is competitive with much more complex bounds which evaluate the state space of the underlying Markov chain. Numerical results on stress cases indicate that the MVA approach is much more scalable than existing evaluation methods for MAP networks.
Giuliano Casale, Evgenia Smirni
DSN1
2009 Autocorrelation-driven load control in distributed systems
abstract
In this paper, we propose a new approach for the development of load control policies in autonomic multitier systems. We control system load in a completely new way compared to existing policies: we leverage on the autocorrelation of service times and show that autocorrelation can be used to forecast future service requirements of requests and adaptively control system load. To the best of our knowledge, this is the first direct application of autocorrelation of service times to autonomic load control. We propose ALoC and D ALoC, two autocorrelation-driven policies that drop a percentage of the load in order to meet pre-defined quality-of-service levels in a distributed system. Both policies are easy to implement and rely on minimal assumptions. In particular, D ALoC is a fully no-knowledge measurement-based policy that self-adjusts its load control parameters based only on policy targets and on statistical information of requests served in the past. We illustrate the effectiveness of these new policies in a distributed multi-server setting via detailed trace driven simulations. We show that if these policies are employed in the server with a temporal dependent service process, then end-to-end response time, across all servers, reduces up to 80% by only dropping at most 13% of the incoming requests. Using real traces, we also show that, in the constrained case of being able to drop only from a portion of the incoming workload, our policy still improves request response time by up to 30%.
Ningfang Mi, Giuliano Casale, Qi Zhang 0012, Alma Riska, Evgenia Smirni
MASCOTS2
2009 Automatic Stress Testing of Multi-tier Systems by Dynamic Bottleneck Switch Generation
Giuliano Casale, Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia
Middleware1
2009 CoMoM: Efficient Class-Oriented Evaluation of Multiclass Performance Models
abstract
We introduce the class-oriented method of moments (CoMoM), a new exact algorithm to compute performance indexes in closed multiclass queuing networks. Closed models are important for performance evaluation of multitier applications, but when the number of service classes is large, they become too expensive to solve with exact methods such as mean value analysis (MVA). CoMoM addresses this limitation by a new recursion that scales efficiently with the number of classes. Compared to the MVA algorithm, which recursively computes mean queue lengths, CoMoM also carries on in the recursion information on higher-order moments of queue lengths. We show that this additional information greatly reduces the number of operations needed to solve the model and makes CoMoM the best-available algorithm for networks with several classes. We conclude the paper by generalizing CoMoM to the efficient computation of marginal queue-length probabilities, which finds application in the evaluation of state-dependent attributes such as quality-of-service metrics.
Giuliano Casale
IEEE Trans. Software Eng.1
2008 Scheduling for performance and availability in systems with temporal dependent workloads
abstract
Temporal locality in workloads creates conditions in which a server, in order to remain available, should quickly process bursts of requests with large service requirements. In this paper, we show how to counteract the resulting peak congestions and maintain high availability by delaying selected requests that contribute to the temporal locality. We propose and evaluate SWAP, a measurement-based scheduling policy that approximates the shortest job first (SJF) scheduling without requiring any knowledge of job service times. We show that good service time estimates can be obtained from the temporal dependence structure of the workload and allow to closely approximate the behavior of SJF. Experimental results indicate that SWAP significantly improves system performability. In particular, we show that system capacity under SWAP is largely increased compared to first-come first-served (FCFS) scheduling and is highly-competitive with SJF, but without requiring a priori information of job service times.
Ningfang Mi, Giuliano Casale, Evgenia Smirni
DSN2
2008 Versatile models of systems using map queueing networks
abstract
Analyzing the performance impact of temporal dependent workloads on hardware and software systems is a challenging task that yet must be addressed to enhance performance of real applications. For instance, existing matrix-analytic queueing models can capture temporal dependence only in systems that can be described by one or two queues, but the capacity planning of real multi-tier architectures requires larger models with arbitrary topology. To address the lack of a proper modeling technique for systems subject to temporal dependent workloads, we introduce a class of closed queueing networks where service times can have non-exponential distribution and accurately approximate temporal dependent features such as short or long range dependence. We describe these service processes using Markovian arrival processes (MAPs), which include the popular Markov-modulated Poisson processes (MMPPs) as special cases. Using a linear programming approach, we obtain for MAP closed networks tight upper and lower bounds for arbitrary performance indexes (e.g., throughput, response time, utilization). Numerical experiments indicate that our bounds achieve a mean accuracy error of 2% and promote our modeling approach for the accurate performance analysis of real multi-tier architectures.
Giuliano Casale, Ningfang Mi, Evgenia Smirni
IPDPS1
2008 Burstiness in Multi-tier Applications: Symptoms, Causes, and New Models
Ningfang Mi, Giuliano Casale, Ludmila Cherkasova, Evgenia Smirni
Middleware2
2008 Robust Workload Estimation in Queueing Network Performance Models
abstract
Traditional approaches for capacity planning are based on queueing network models. However, modeling with queueing networks requires the knowledge of the service demands of each class of workloads at each device described in the model. In real systems, such service demands can be very difficult to measure. In this paper, we present an optimization-based technique to address the problem. The technique is formulated as a robust linear parameter estimation that can be used with both closed and open queueing network models. We consider the case where aggregate measurements (throughput and utilization) are available. Such measurements are typically much easier to obtain than the service demands. We present experimental results which prove the effectiveness of the constrained and robust linear estimation.
Giuliano Casale, Paolo Cremonesi, Roberto Turrin
PDP1
2008 Bound analysis of closed queueing networks with workload burstiness
abstract
Burstiness and temporal dependence in service processes are often found in multi-tier architectures and storage devices and must be captured accurately in capacity planning models as these features are responsible of significant performance degradations. However, existing models and approximations for networks of first-come first-served (FCFS) queues with general independent (GI) service are unable to predict performance of systems with temporal dependence in workloads.
Giuliano Casale, Ningfang Mi, Evgenia Smirni
SIGMETRICS1
2008 Geometric Bounds: A Noniterative Analysis Technique for Closed Queueing Networks
abstract
We propose the Geometric Bounds (GBs), a new family of fast and accurate noniterative bounds on closed queueing network performance metrics that can be used in the online optimization of distributed applications. Compared to state-of-the-art techniques such as the Balanced Job Bounds (BJBs), GB achieves higher accuracy at similar computational costs, limiting the worst- case bounding error typically within 5-13 percent when, for the BJB, it is usually in the range of 15-35 percent. Optimization problems that are solved with GBs return solutions that are much closer to the global optimum than with existing bounds. We also show that the GB technique generalizes as an accurate approximation to closed fork-join networks commonly used in disk, parallel, and database models, thus extending the applicability of the method beyond the optimization of basic product-form networks.
Giuliano Casale, Richard R. Muntz, Giuseppe Serazzi
IEEE Trans. Computers1
2007 New Results on the Performance Effects of Autocorrelated Flows in Systems
abstract
Temporal dependence within the workload of any computing or networking system has been widely recognized as a significant factor affecting performance. More specifically, burstiness, as a form of temporal dependency, is catastrophic for performance. We use the autocorrelation function in a workload flow to formalize burstiness and also to characterize temporal dependence within a flow. We present results from two application areas: load balancing in a homogeneous cluster environment and capacity planning in a multi-tiered e-commerce system. For the load balancing problem, we show that if autocorrelation exists in the arrival stream to the cluster, classic load balancing policies become ineffective and solutions that focus on "unbalancing" the load offer superior performance. For the case of multi-tiered systems, we show that if there is autocorrelation in the flows, we observe the surprising result that in spite of the fact that the bottleneck resource in the system is far from saturation and that the measured throughput and utilizations of other resources are also modest, user response times are very high. For multi-tired systems, this underutilization of resources falsely indicates that the system can sustain higher capacities. We present analysis of the above phenomena that aims at the development better scheduling policies under auto correlated flows.
Evgenia Smirni, Qi Zhang 0012, Ningfang Mi, Alma Riska, Giuliano Casale
IPDPS5
2007 Approximate Solution of Multiclass Queuing Networks with Region Constraints
abstract
Among existing modeling techniques, queueing networks with "finite capacity regions" have largely proven to be effective in characterizing push-back effects and simultaneous resource possession in which a request holds more resources simultaneously. Queueing network models with finite capacity regions impose upper bounds on the number of jobs that can simultaneously reside in a set of service centers. For this reason they can be used to model application constraints. However, since they do not satisfy product-form assumptions, they are difficult to treat. In this paper we propose a novel approximate method for closed multiclass queueing networks containing finite capacity regions and shared constraints. Our approach is based on Norton's theorem for queueing networks where a region is replaced by a single flow equivalent service center (FESC). We propose a population-mix driven definition of FESCs service rates which provides increased accuracy with respect to existing methods. We solve the resulting non-product-form network with a new approximate variant of the convolution algorithm proposed in the paper. A comparison with simulation shows that the algorithm typically has a 4% approximation error.
Jonatha Anselmi, Giuliano Casale, Paolo Cremonesi
MASCOTS2
2006 A New Class of Non-Iterative Bounds for Closed Queueing Networks
abstract
A new emerging class of problems related to the online configuration and optimization of computer systems and networks requires the solution in a very short amount of time of a large number of analytical performance models, often based on queueing networks. In this paper we propose the Geometric Bounds (GB), a new family of fast noniterative bounds on performance metrics of closed productform queueing networks. In spite of their simplicity, the proposed bounds are more accurate than the popular Balanced Job Bounds (BJB), even in the difficult case of networks with multiple bottlenecks or large delays.
Giuliano Casale, Richard R. Muntz, Giuseppe Serazzi
MASCOTS1