VLDB 2026 Research / reviewers in the wild / expert
Peter Garraghan
dblp:69/10837 · also Peter Michael Garraghan
· DBLP profile ↗
33ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-7103-2515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CECF: A DNN-Based Energy-Efficient Cloud-Edge Collaboration Framework for Intelligent Workload Scheduling in 6G-Enabled Transportation SystemsabstractThe rapid growth of Internet of Vehicle (IoV) devices and Artificial Intelligence (AI) applications has accelerated the adoption of Cloud and Edge Computing. The advent of sixth-generation mobile communication technology (6G) further facilitates the deployment of Cloud-Edge collaborative computing in large-scale Intelligent Transportation Systems (ITS). Effective ITS must efficiently handle both latency-sensitive tasks (e.g., obstacle detection, traffic signal recognition) and computationally intensive tasks (e.g., path optimization, traffic flow prediction). However, existing Cloud-Edge collaborative frameworks struggle to accurately classify diverse workloads and provide efficient low-latency processing, leading to energy inefficiencies and task failures. To address these challenges, this paper introduces a Deep Learning-based Cloud-Edge Collaboration Framework (CECF) designed to optimize energy conservation in Cloud and Edge environments. CECF employs a DNN-based classifier to categorize workloads for processing in the Cloud or Edge. The classified tasks are managed by a dedicated Cloud scheduler (DSGA) and an Edge scheduler (EA-DFPSO), respectively. To enhance scheduling efficiency for highly variable Cloud tasks, DSGA incorporates a novel self-adaptive mutation algorithm and a random point fixed distance crossover method. Extensive evaluations using real-world workload traces demonstrate that CECF achieves up to a 8.5% improvement in system reliability and reduces energy consumption by 35.88% compared to baseline approaches. Yao Lu 0021, Lu Liu 0001, John Panneerselvam, Jiayan Gu, Peter Garraghan, Geyong Min |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Energy-adaptive Network Switching via Intra-device ScalingabstractWe propose horizontal intra-device scaling for network switches. Our approach allows for a network device to dynamically scale energy use in response to changing network utilization at a finer-grain in comparison to existing monolithic approaches, and enables a reduction in cost and environmental impact via reduced network energy use outside peak operating periods. We demonstrate the feasibility of intra-device switch scaling by designing a network switch architecture comprising multiple, less powerful, network devices leveraging a Multiple Spanning Tree Protocol (MSTP) to operate in parallel in place of a singular powerful device. Our preliminary results demonstrate that our approach can reduce total network energy use by 66.3% in comparison to established approaches with minimal performance penalty, and outlines future work for further improvement for this new form of network switch architecture for reducing energy use within core network infrastructure. Matthew Alexander Hodkin, Louise Krug, Peter Garraghan |
ICC | 3 |
| 2023 | DOPpler: Parallel Measurement Infrastructure for Auto-Tuning Deep Learning Tensor ProgramsabstractThe heterogeneity of Deep Learning models, libraries, and hardware poses an important challenge for improving model inference performance. Auto-tuners address this challenge via automatic tensor program optimization towards a target-device. However, auto-tuners incur a substantial time cost to complete given their design necessitates performing tensor program candidate measurements serially within an isolated target-device to minimize latency measurement inaccuracy. In this article we propose DOPpler, a parallel auto-tuning measurement infrastructure. DOPpler allows for considerable auto-tuning speedup over conventional approaches whilst maintaining high-quality tensor program optimization. DOPpler accelerates the auto-tuning process by proposing a parallel execution engine to efficiently execute candidate tensor programs in parallel across the CPU-host and GPU target-device, and overcomes measurement inaccuracy by introducing a high-precision on-device measurement technique when measuring tensor program kernel latency. DOPpler is designed to automatically calculate the optimal degree of parallelism to provision fast and accurate auto-tuning for different tensor programs, auto-tuners and target-devices. Experiment results show that DOPpler reduces total auto-tuning time by 50.5% on average whilst achieving optimization gains equivalent to conventional auto-tuning infrastructure. Damian Borowiec, Gingfung Yeung, Adrian Friday, Richard Harper 0001, Peter Garraghan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | START: Straggler Prediction and Mitigation for Cloud Computing Environments Using Encoder LSTM NetworksabstractA common performance problem in large-scale cloud systems is dealing with straggler tasks that are slow running instances which increase the overall response time. Such tasks impact the system's QoS and the SLA. There is a need for automatic straggler detection and mitigation mechanisms that execute jobs without violating the SLA. Prior work typically builds reactive models that focus first on detection and then mitigation of straggler tasks, which leads to delays. Other works use prediction based proactive mechanisms, but ignore volatile task characteristics. We propose a Straggler Prediction and Mitigation Technique (START) that is able to predict which tasks might be stragglers and dynamically adapt scheduling to achieve lower response times. START analyzes all tasks and hosts based on compute and network resource consumption using an Encoder LSTM network to predict and mitigate expected straggler tasks. This reduces the SLA violation rate and execution time without compromising QoS. Specifically, we use the CloudSim toolkit to simulate START and compare it with IGRU-SD, SGC, Dolly, GRASS, NearestFit and Wrangler in terms of QoS parameters. Experiments show that START reduces execution time, resource contention, energy and SLA violations by 13%, 11%, 16%, 19%, compared to the state-of-the-art. Shreshth Tuli, Sukhpal Singh, Peter Garraghan, Rajkumar Buyya, Giuliano Casale, Nicholas R. Jennings |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Trimmer: Cost-Efficient Deep Learning Auto-tuning for Cloud DatacentersabstractCloud datacenters capable of provisioning high performance Machine Learning-as-a-Service (MLaaS) at reduced resource cost is achieved via auto-tuning: automated tensor program optimization of Deep Learning models to minimize inference latency within a hardware device. However given the extensive heterogeneity of Deep Learning models, libraries, and hardware devices, performing auto-tuning within Cloud datacenters incurs a significant time, compute resource, and energy cost of which state-of-the-art auto-tuning is not designed to mitigate. In this paper we propose Trimmer, a high performance and cost-efficient Deep Learning auto-tuning framework for Cloud datacenters. Trimmer maximizes DL model performance and tensor program cost-efficiency by preempting tensor program implementations exhibiting poor optimization improvement; and applying an ML-based filtering method to replace expensive low performing tensor programs to provide greater likelihood of selecting low latency tensor programs. Through an empirical study exploring the cost of DL model optimization techniques, our analysis indicates that 26–43% of total energy is expended on measuring tensor program implementations that do not positively contribute towards auto-tuning. Experiment results show that Trimmer achieves high auto-tuning cost-efficiency across different DL models, and reduces auto-tuning energy use by 21.8–40.9% for Cloud clusters whilst achieving DL model latency equivalent to state-of-the-art techniques. Damian Borowiec, Gingfung Yeung, Adrian Friday, Richard Harper 0001, Peter Garraghan |
CLOUD | 5 |
| 2022 | HUNTER: AI based holistic resource management for sustainable cloud computing
Shreshth Tuli, Sukhpal Singh, Minxian Xu, Peter Garraghan, Rami Bahsoon, Schahram Dustdar, Rizos Sakellariou, Omer F. Rana, Rajkumar Buyya, Giuliano Casale, Nicholas R. Jennings |
J. Syst. Softw. | 4 |
| 2022 | Cross-VM Network Channel Attacks and Countermeasures Within Cloud Computing EnvironmentsabstractCloud providers attempt to maintain the highest levels of isolation between Virtual Machines (VMs) and inter-user processes to keep co-located VMs and processes separate. This logical isolation creates an internal virtual network to separate VMs co-residing within a shared physical network. However, as co-residing VMs share their underlying VMM (Virtual Machine Monitor), virtual network, and hardware are susceptible to cross VM attacks. It is possible for a malicious VM to potentially access or control other VMs through network connections, shared memory, other shared resources, or by gaining the privilege level of its non-root machine. This research presents a two novel zero-day cross-VM network channel attacks. In the first attack, a malicious VM can redirect the network traffic of target VMs to a specific destination by impersonating the Virtual Network Interface Controller (VNIC). The malicious VM can extract the decrypted information from target VMs by using open source decryption tools such as Aircrack. The second contribution of this research is a privilege escalation attack in a cross VM cloud environment with Xen hypervisor. An adversary having limited privileges rights may execute Return-Oriented Programming (ROP), establish a connection with the root domain by exploiting the network channel, and acquiring the tool stack (root domain) which it is not authorized to access directly. Countermeasures against this attacks are also presented Atif Saeed, Peter Garraghan, Syed Asad Hussain |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Horus: Interference-Aware and Prediction-Based Scheduling in Deep Learning SystemsabstractTo accelerate the training of Deep Learning (DL) models, clusters of machines equipped with hardware accelerators such as GPUs are leveraged to reduce execution time. State-of-the-art resource managers are needed to increase GPU utilization and maximize throughput. While co-locating DL jobs on the same GPU has been shown to be effective, this can incur interference causing slowdown. In this article we propose Horus: an interference-aware and prediction-based resource manager for DL systems. Horus proactively predicts GPU utilization of heterogeneous DL jobs extrapolated from the DL model's computation graph features, removing the need for online profiling and isolated reserved GPUs. Through micro-benchmarks and job co-location combinations across heterogeneous GPU hardware, we identify GPU utilization as a general proxy metric to determine good placement decisions, in contrast to current approaches which reserve isolated GPUs to perform online profiling and directly measure GPU utilization for each unique submitted job. Our approach promotes high resource utilization and makespan reduction; via real-world experimentation and large-scale trace driven simulation, we demonstrate that Horus outperforms other DL resource managers by up to 61.5 percent for GPU resource utilization, 23.7-30.7 percent for makespan reduction and 68.3 percent in job wait time reduction. Gingfung Yeung, Damian Borowiec, Renyu Yang, Adrian Friday, Richard Harper 0001, Peter Garraghan |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | An Empirical Study of Inter-cluster Resource Orchestration within Federated Cloud ClustersabstractFederated clusters are composed of multiple independent clusters of machines interconnected by a resource management system, and possess several advantages over centralized cloud datacenter clusters including seamless provisioning of applications across large geographic regions, greater fault tolerance, and increased cluster resource utilization. However, while existing resource management systems for federated clusters are capable of improving application intra-cluster performance, they do not capture inter-cluster performance in their decision making. This is important given federated clusters must execute a wide variety of applications possessing heterogeneous system architectures, which are a impacted by unique inter-cluster performance conditions such as network latency and localized cluster resource contention. In this work we present an empirical study demonstrating how inter-cluster performance conditions negatively impact federated cluster orchestration systems. We conduct a series of micro-benchmarks under various cluster operational scenarios showing the critical importance in capturing inter-cluster performance for resource orchestration in federated clusters. From this benchmark, we determine precise limitations in existing federated orchestration, and highlight key insights to design future orchestration systems. Findings of notable interest entail different application types exhibiting innate performance affinities across various federated cluster operational conditions, and experience substantial performance degradation from even minor increases to latency (8.7x) and resource contention (12.0x) in comparison to centralized cluster architectures. Dominic Lindsay, Gingfung Yeung, Yehia El-khatib, Peter Garraghan |
JCC | 4 |
| 2020 | Horus: An Interference-Aware Resource Manager for Deep Learning Systems
Gingfung Yeung, Damian Borowiec, Renyu Yang, Adrian Friday, Richard Harper 0001, Peter Garraghan |
ICA3PP (2) | 6 |
| 2020 | Integrating clustering and regression for workload estimation in the cloudabstractAbstract Workload prediction has been widely researched in the literature. However, existing techniques are per‐job based and useful for service‐like tasks whose workloads exhibit seasonality and trend. But cloud jobs have many different workload patterns and some do not exhibit recurring workload patterns. We consider job‐pool‐based workload estimation, which analyzes the characteristics of existing tasks' workloads to estimate the currently running tasks' workload. First cluster existing tasks based on their workloads. For a new task J, collect the initial workload of J and determine which cluster J may belong to, then use the cluster's characteristics to estimate J′s workload. Based on the Google dataset, the algorithm is experimentally evaluated and its effectiveness is confirmed. However, the workload patterns of some tasks do have seasonality and trend, and conventional per‐job‐based regression methods may yield better workload prediction results. Also, in some cases, some new tasks may not follow the workload patterns of existing tasks in the pool. Thus, develop an integrated scheme which combines clustering and regression and utilize the best of them for workload prediction. Experimental study shows that the combined approach can further improve the accuracy of workload prediction. Yongjia Yu, Vasu Jindal, I-Ling Yen, Farokh B. Bastani, Jie Xu 0007, Peter Garraghan |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | ThermoSim: Deep learning based framework for modeling and simulation of thermal-aware resource management for cloud computing environments
Sukhpal Singh, Shreshth Tuli, Adel Nadjaran Toosi, Félix Cuadrado, Peter Garraghan, Rami Bahsoon, Hanan Lutfiyya, Rizos Sakellariou, Omer F. Rana, Schahram Dustdar, Rajkumar Buyya |
J. Syst. Softw. | 5 |
| 2020 | Tails in the cloud: a survey and taxonomy of straggler management within large-scale cloud data centres
Sukhpal Singh, Xue Ouyang 0003, Peter Garraghan |
J. Supercomput. | 3 |
| 2020 | Performance-Aware Speculative Resource Oversubscription for Large-Scale ClustersabstractIt is a long-standing challenge to achieve a high degree of resource utilization in cluster scheduling. Resource oversubscription has become a common practice in improving resource utilization and cost reduction. However, current centralized approaches to oversubscription suffer from the issue with resource mismatch and fail to take into account other performance requirements, e.g., tail latency. In this article we present ROSE, a new resource management platform capable of conducting performance-aware resource oversubscription. ROSE allows latency-sensitive long-running applications (LRAs) to co-exist with computation-intensive batch jobs. Instead of waiting for resource allocation to be confirmed by the centralized scheduler, job managers in ROSE can independently request to launch speculative tasks within specific machines according to their suitability for oversubscription. Node agents of those machines can however, avoid any excessive resource oversubscription by means of a mechanism for admission control using multi-resource threshold control and performance-aware resource throttle. Experiments show that in case of mixed co-location of batch jobs and latency-sensitive LRAs, the CPU utilization and the disk utilization can reach 56.34 and 43.49 percent, respectively, but the 95th percentile of read latency in YCSB workloads only increases by 5.4 percent against the case of executing the LRAs alone. Renyu Yang, Chunming Hu, Peter Garraghan, Tianyu Wo, Zhenyu Wen, Hao Peng 0001, Jie Xu 0007 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | ROUTER: Fog enabled cloud based intelligent resource management approach for smart home IoT devices
Sukhpal Singh, Peter Garraghan, Rajkumar Buyya |
J. Syst. Softw. | 2 |
| 2019 | Holistic resource management for sustainable and reliable cloud computing: An innovative solution to global challenge
Sukhpal Singh, Peter Garraghan, Vlado Stankovski, Giuliano Casale, Ruppa K. Thulasiram, Soumya K. Ghosh 0001, Kotagiri Ramamohanarao, Rajkumar Buyya |
J. Syst. Softw. | 2 |
| 2019 | Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud DatacentersabstractIncreased complexity and scale of virtualized distributed systems has resulted in the manifestation of emergent phenomena substantially affecting overall system performance. This phenomena is known as “Long Tail”, whereby a small proportion of task stragglers significantly impede job completion time. While work focuses on straggler detection and mitigation, there is limited work that empirically studies straggler root-cause and quantifies its impact upon system operation. Such analysis is critical to ascertain in-depth knowledge of straggler occurrence for focusing developmental and research efforts towards solving the Long Tail challenge. This paper provides an empirical analysis of straggler root-cause within virtualized Cloud datacenters; we analyze two large-scale production systems to quantify the frequency and impact stragglers impose, and propose a method for conducting root-cause analysis. Results demonstrate approximately 5 percent of task stragglers impact 50 percent of total jobs for batch processes, and 53 percent of stragglers occur due to high server resource utilization. We leverage these findings to propose a method for extreme straggler detection through a combination of offline execution patterns modeling and online analytic agents to monitor tasks at runtime. Experiments show the approach is capable of detecting stragglers less than 11 percent into their execution lifecycle with 95 percent accuracy for short duration jobs. Peter Garraghan, Xue Ouyang 0003, Renyu Yang, David McKee 0001, Jie Xu 0007 |
IEEE Trans. Serv. Comput. | 1 |
| 2018 | A Cross-Virtual Machine Network Channel Attack via Mirroring and TAP ImpersonationabstractData privacy and security is a leading concern for providers and customers of cloud computing, where Virtual Machines (VMs) can co-reside within the same underlying physical machine. Side channel attacks within multi-tenant virtualized cloud environments are an established problem, where attackers are able to monitor and exfiltrate data from co-resident VMs. Virtualization services have attempted to mitigate such attacks by preventing VM-to-VM interference on shared hardware by providing logical resource isolation between co-located VMs via an internal virtual network. However, such approaches are also insecure, with attackers capable of performing network channel attacks which bypass mitigation strategies using vectors such as ARP Spoofing, TCP/IP steganography, and DNS poisoning. In this paper we identify a new vulnerability within the internal cloud virtual network, showing that through a combination of TAP impersonation and mirroring, a malicious VM can successfully redirect and monitor network traffic of VMs co-located within the same physical machine. We demonstrate the feasibility of this attack in a prominent cloud platform - OpenStack - under various security requirements and system conditions, and propose countermeasures for mitigation. Atif Saeed, Peter Garraghan, Barnaby Craggs, Dirk van der Linden, Awais Rashid, Syed Asad Hussain |
IEEE CLOUD | 2 |
| 2018 | ROSE: Cluster Resource Scheduling via Speculative Over-SubscriptionabstractA long-standing challenge in cluster scheduling is to achieve a high degree of utilization of heterogeneous resources in a cluster. In practice there exists a substantial disparity between perceived and actual resource utilization. A scheduler might regard a cluster as fully utilized if a large resource request queue is present, but the actual resource utilization of the cluster can be in fact very low. This disparity results in the formation of idle resources, leading to inefficient resource usage and incurring high operational costs and an inability to provision services. In this paper we present a new cluster scheduling system, ROSE, that is based on a multi-layered scheduling architecture with an ability to over-subscribe idle resources to accommodate unfulfilled resource requests. ROSE books idle resources in a speculative manner: instead of waiting for resource allocation to be confirmed by the centralized scheduler, it requests intelligently to launch tasks within machines according to their suitability to oversubscribe resources. A threshold control with timely task rescheduling ensures fully-utilized cluster resources without generating potential task stragglers. Experimental results show that ROSE can almost double the average CPU utilization, from 36.37% to 65.10%, compared with a centralized scheduling scheme, and reduce the workload makespan by 30.11%, with an 8.23% disk utilization improvement over other scheduling strategies. Chunming Hu, Renyu Yang, Peter Garraghan, Tianyu Wo, Jie Xu 0007, Jianyong Zhu |
ICDCS | 4 |
| 2018 | Holistic energy and failure aware workload scheduling in Cloud datacenters
Xiang Li 0017, Xiaohong Jiang 0002, Peter Garraghan, Zhaohui Wu 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | Adaptive Speculation for Efficient Internetware Application Execution in CloudsabstractModern Cloud computing systems are massive in scale, featuring environments that can execute highly dynamic Internetware applications with huge numbers of interacting tasks. This has led to a substantial challenge—the straggler problem, whereby a small subset of slow tasks significantly impede parallel job completion. This problem results in longer service responses, degraded system performance, and late timing failures that can easily threaten Quality of Service (QoS) compliance. Speculative execution (or speculation) is the prominent method deployed in Clouds to tolerate stragglers by creating task replicas at runtime. The method detects stragglers by specifying a predefined threshold to calculate the difference between individual tasks and the average task progression within a job. However, such a static threshold debilitates speculation effectiveness as it fails to capture the intrinsic diversity of timing constraints in Internetware applications, as well as dynamic environmental factors, such as resource utilization. By considering such characteristics, different levels of strictness for replica creation can be imposed to adaptively achieve specified levels of QoS for different applications. In this article, we present an algorithm to improve the execution efficiency of Internetware applications by dynamically calculating the straggler threshold, considering key parameters including job QoS timing constraints, task execution progress, and optimal system resource utilization. We implement this dynamic straggler threshold into the YARN architecture to evaluate it’s effectiveness against existing state-of-the-art solutions. Results demonstrate that the proposed approach is capable of reducing parallel job response time by up to 20% compared to the static threshold, as well as a higher speculation success rate, achieving up to 66.67% against 16.67% in comparison to the static method. Xue Ouyang 0003, Peter Garraghan, Bernhard Primas, David McKee 0001, Paul Townend, Jie Xu 0007 |
ACM Trans. Internet Techn. | 2 |
| 2018 | Holistic Virtual Machine Scheduling in Cloud Datacenters towards Minimizing Total EnergyabstractEnergy consumed by Cloud datacenters has dramatically increased, driven by rapid uptake of applications and services globally provisioned through virtualization. By applying energy-aware virtual machine scheduling, Cloud providers are able to achieve enhanced energy efficiency and reduced operation cost. Energy consumption of datacenters consists of computing energy and cooling energy. However, due to the complexity of energy and thermal modeling of realistic Cloud datacenter operation, traditional approaches are unable to provide a comprehensive in-depth solution for virtual machine scheduling which encompasses both computing and cooling energy. This paper addresses this challenge by presenting an elaborate thermal model that analyzes the temperature distribution of airflow and server CPU. We propose GRANITE - a holistic virtual machine scheduling algorithm capable of minimizing total datacenter energy consumption. The algorithm is evaluated against other existing workload scheduling algorithms MaxUtil, TASA, IQR and Random using real Cloud workload characteristics extracted from Google datacenter tracelog. Results demonstrate that GRANITE consumes 4.3-43.6 percent less total energy in comparison to the state-of-the-art, and reduces the probability of critical temperature violation by 99.2 with 0.17 percent SLA violation rate as the performance penalty. Xiang Li 0017, Peter Garraghan, Xiaohong Jiang 0002, Zhaohui Wu 0001, Jie Xu 0007 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | A Framework and Task Allocation Analysis for Infrastructure Independent Energy-Efficient Scheduling in Cloud Data CentersabstractCloud computing represents a paradigm shift in provisioning on-demand computational resources underpinned by data center infrastructure, which now constitutes 1.5% of worldwide energy consumption. Such consumption is not merely limited to operating IT devices, but encompasses cooling systems representing 40% total data center energy usage. Given the substantive complexity and heterogeneity of data center operation spanning both computing and cooling components, obtaining analytical models for optimizing data center energy-efficiency is an inherently difficult challenge. Specifically, difficulties arise pertaining to the non-intuitive relationship between computing and cooling energy in the data center, computationally complex energy modeling, as well as cooling models restricted to a specific class of data center facility geometry - all of which arise from the interdisciplinary nature of this research domain.In this paper we propose a framework for energy-efficient scheduling to alleviate these challenges. It is applicable to any type of data center infrastructure and does not require complex modeling of energy.Instead, the concept of a target workload distribution is proposed. If the workload is assigned to nodes according to the target workload distribution, then the energy consumption is minimized. The exact target workload distribution is unknown, but an approximated distribution is delivered by the framework. The scheduling objective is to assign workload to nodes such that the workload distribution becomes as similar as possible to the target distribution in order to reduce energy consumption.Several mathematically sound algorithms have been designed to address this novel type of scheduling problem. Simulation results demonstrate that our algorithms reduce the relative deviation by at least 16.9% and the relative variance by at least 22.67% in comparison to (asymmetric) load balancing algorithms. Bernhard Primas, Peter Garraghan, David McKee 0001, Jon Summers, Jie Xu 0007 |
CloudCom | 2 |
| 2017 | Reliable Computing Service in Massive-Scale Systems through Rapid Low-Cost FailoverabstractLarge-scale distributed systems deployed as Cloud datacenters are capable of provisioning service to consumers with diverse business requirements. Providers face pressure to provision uninterrupted reliable services while reducing operational costs due to significant software and hardware failures. A widely adopted means to achieve such a goal is using redundant system components to implement user-transparent failover, yet its effectiveness must be balanced carefully without incurring heavy overhead when deployed—an important practical consideration for complex large-scale systems. Failover techniques developed for Cloud systems often suffer serious limitations, including mandatory restart leading to poor cost-effectiveness, as well as solely focusing on crash failures, omitting other important types, such as timing failures and simultaneous failures. This paper addresses these limitations by presenting a new approach to user-transparent failover for massive-scale systems. The approach uses soft-state inference to achieve rapid failure recovery and avoid unnecessary restart, with minimal system resource overhead. It also copes with different failures, including correlated and simultaneous events. The proposed approach was implemented, deployed and evaluated within Fuxi system, the underlying resource management system used within Alibaba Cloud. Results demonstrate that our approach tolerates complex failure scenarios while incurring at worst 228.5 microsecond instance overhead with 1.71 percent additional CPU usage. Renyu Yang, Peter Garraghan, Yihui Feng, Jin Ouyang, Jie Xu 0007, Zhuo Zhang 0015 |
IEEE Trans. Serv. Comput. | 3 |
| 2016 | Straggler Detection in Parallel Computing Systems through Dynamic Threshold CalculationabstractCloud computing systems face the substantial challenge of the Long Tail problem: a small subset of straggling tasks significantly impede parallel jobs completion. This behavior results in longer service response times and degraded system utilization. Speculative execution, which create task replicas at runtime, is a typical method deployed in large-scale distributed systems to tolerate stragglers. This approach defines stragglers by specifying a static threshold value, which calculates the temporal difference between an individual task and the average task progression for a job. However, specifying static threshold debilitates speculation effectiveness as it fails to consider the intrinsic diversity of job timing constraints within modern day Cloud computing systems. Capturing such heterogeneity enables the ability to impose different levels of strictness for replica creation while achieving specified levels of QoS for different application types. Furthermore, a static threshold also fails to consider system environmental constraints in terms of replication overheads and optimal system resource usage. In this paper we present an algorithm for dynamically calculating a threshold value to identify task stragglers, considering key parameters including job QoS timing constraints, task execution characteristics, and optimal system resource utilization. We study and demonstrate the effectiveness of our algorithm through simulating a number of different operational scenarios based on real production cluster data against state-of-the-art solutions. Results demonstrate that our approach is capable of creating 58.62% less replicas under high resource utilization while reducing response time up to 17.86% for idle periods compared to a static threshold. Xue Ouyang 0003, Peter Garraghan, David McKee 0001, Paul Townend, Jie Xu 0007 |
AINA | 2 |
| 2016 | Virtual Machine Level Temperature Profiling and Prediction in Cloud DatacentersabstractTemperature prediction can enhance datacenter thermal management towards minimizing cooling power draw. Traditional approaches achieve this through analyzing task-temperature profiles or resistor-capacitor circuit models to predict CPU temperature. However, they are unable to capture task resource heterogeneity within multi-tenant environments and make predictions under dynamic scenarios such as virtual machine migration, which is one of the main characteristics of Cloud computing. This paper proposes virtual machine level temperature prediction in Cloud datacenters. Experiments show that the mean squared error of stable CPU temperature prediction is within 1.10, and dynamic CPU temperature prediction can achieve 1.60 in most scenarios. Zhaohui Wu 0001, Xiang Li 0017, Peter Garraghan, Xiaohong Jiang 0002, Kejiang Ye, Albert Y. Zomaya |
ICDCS | 3 |
| 2016 | Tolerating Transient Late-Timing Faults in Cloud-Based Real-Time Stream ProcessingabstractReal-time stream processing is a frequently deployed application within Cloud datacenters that is required to provision high levels of performance and reliability. Numerous fault-tolerant approaches have been proposed to effectively achieve this objective in the presence of crash failures. However, such systems struggle with transient late-timing faults - a fault classification challenging to effectively tolerate - that manifests increasingly within large-scale distributed systems. Such faults represent a significant threat towards minimizing soft real-time execution of streaming applications in the presence of failures. This work proposes a fault-tolerant approach for QoS-aware data prediction to tolerate transient late-timing faults. The approach is capable of determining the most effective data prediction algorithm for imposed QoS constraints on a failed stream processor at run-time. We integrated our approach into Apache Storm with experiment results showing its ability to minimize stream processor end-to-end execution time by 61% compared to other fault-tolerant approaches. The approach incurs 12% additional CPU utilization while reducing network usage by 44%. Peter Garraghan, Stuart Perks, Xue Ouyang 0003, David McKee 0001, Ismael Solís Moreno |
ISORC | 1 |
| 2016 | SEED: A Scalable Approach for Cyber-Physical System SimulationabstractSimulation is critical when studying real operational behavior of increasingly complex Cyber-Physical Systems, forecasting future behavior, and experimenting with hypothetical scenarios. A critical aspect of simulation is the ability to evaluate large-scale systems within a reasonable time frame while modeling complex interactions between millions of components. However, modern simulations face limitations in provisioning this functionality for CPSs in terms of balancing simulation complexity with performance, resulting in substantial operational costs required for completing simulation execution. Moreover, users are required to have expertise in modeling and configuring simulations to infrastructure which is time consuming. In this paper we present Simulation EnvironmEnt Distributor (SEED), a novel approach for simulating large-scale CPSs across a loosely-coupled distributed system requiring minimal user configuration. This is achieved through automated simulation partitioning and instantiation while enforcing tight event messaging across the system. SEED operates efficiently within both small and large-scale OTS hardware, agnostic of cluster heterogeneity and OS running, and is capable of simulating the full system and network stack of a CPS. Our approach is validated through experiments conducted in a cluster to simulate CPS operation. Results demonstrate that SEED is capable of simulating CPSs containing 2,000,000 tasks across 2,000 nodes with only 6.89× slow down relative to real time, and executes effectively across distributed infrastructure. Peter Garraghan, David McKee 0001, Xue Ouyang 0003, David Webster, Jie Xu 0007 |
IEEE Trans. Serv. Comput. | 1 |
| 2015 | Workload Estimation for Improving Resource Management Decisions in the CloudabstractIn cloud computing, good resource management can benefit both cloud users as well as cloud providers. Workload prediction is a crucial step towards achieving good resource management. While it is possible to estimate the workloads of long-running tasks based on the periodicity in their historical workloads, it is difficult to do so for tasks which do not have such recurring workload patterns. In this paper, we present an innovative clustering based resource estimation approach which groups tasks that have similar characteristics into the same cluster. The historical workload data for tasks in a cluster are used to estimate the resources needed by new tasks based on the cluster(s) to which they belong. In particular, for a new task T, we measure T's initial workload and predict to which cluster(s) it may belong. Then, the workload information of the cluster(s) is used to estimate the workload of T. The approach is experimentally evaluated using Google dataset, including resource usage data of over half a million tasks. We develop a workload model based on the dataset which is then used to estimate the workload patterns of several randomly selected tasks from the trace log. The results confirm the effectiveness of this cluster-based method for estimating the resources required by each task. Jemishkumar Patel, Vasu Jindal, I-Ling Yen, Farokh B. Bastani, Jie Xu 0007, Peter Garraghan |
ISADS | 6 |
| 2015 | Timely Long Tail Identification through Agent Based Monitoring and AnalyticsabstractThe increasing complexity and scale of distributed systems has resulted in the manifestation of emergent behavior which substantially affects overall system performance. A significant emergent property is that of the "Long Tail", whereby a small proportion of task stragglers significantly impact job execution completion times. To mitigate such behavior, straggling tasks occurring within the system need to be accurately identified in a timely manner. However, current approaches focus on mitigation rather than identification, which typically identify stragglers too late in the execution lifecycle. This paper presents a method and tool to identify Long Tail behavior within distributed systems in a timely manner, through a combination of online and offline analytics. This is achieved through historical analysis to profile and model task execution patterns, which then inform online analytic agents that monitor task execution at runtime. Furthermore, we provide an empirical analysis of two large-scale production Cloud data enters that demonstrate the challenge of data skew within modern distributed systems, this analysis shows that approximately 5% of task stragglers caused by data skew impact 50% of the total jobs for batch processes. Our results demonstrate that our approach is capable of identifying task stragglers less than 11% into their execution lifecycle with 98% accuracy, signifying significant improvement over current state-of-the-art practice and enables far more effective mitigation strategies in large-scale distributed systems worldwide. Peter Garraghan, Xue Ouyang 0003, Paul Townend, Jie Xu 0007 |
ISORC | 1 |
| 2014 | Fault-Tolerant Dynamic Deduplication for Utility ComputingabstractUtility computing is an increasingly important paradigm, whereby computing resources are provided on-demand as utilities. An important component of utility computing is storage, data volumes are growing rapidly, and mechanisms to mitigate this growth need to be developed. Data deduplication is a promising technique for drastically reducing the amount of data stored in such system systems, however, current approachs are static in nature, using an amount of redundancy fixed at design time. This is inappropriate for truly dynamic modern systems. We propose a real-time adaptive deduplication system for Cloud and Utility computing that monitors in real-time for changing system, user, and environmental behaviour in order to fulfill a balance between changing storage efficiency, performance, and fault tolerance requirements. We evaluate our system through simulation, with experimental results showing that our system is both efficient and sclable. We also perform experimentation to evaluate the fault tolerance of the system by measuring Mean Time to Repair (MTTR), and using these values to calculate availability of the system. The results show that higher replication levels result in higher system availability, however, the number of files in the system also effects recovery time. We show that the tradeoff between replication levels and recovery time when the system overloads needs further investigation. Waraporn Leesakul, Paul Townend, Peter Garraghan, Jie Xu 0007 |
ISORC | 3 |
| 2014 | Analysis, Modeling and Simulation of Workload Patterns in a Large-Scale Utility CloudabstractUnderstanding the characteristics and patterns of workloads within a Cloud computing environment is critical in order to improve resource management and operational conditions while Quality of Service (QoS) guarantees are maintained. Simulation models based on realistic parameters are also urgently needed for investigating the impact of these workload characteristics on new system designs and operation policies. Unfortunately there is a lack of analyses to support the development of workload models that capture the inherent diversity of users and tasks, largely due to the limited availability of Cloud tracelogs as well as the complexity in analyzing such systems. In this paper we present a comprehensive analysis of the workload characteristics derived from a production Cloud data center that features over 900 users submitting approximately 25 million tasks over a time period of a month. Our analysis focuses on exposing and quantifying the diversity of behavioral patterns for users and tasks, as well as identifying model parameters and their values for the simulation of the workload created by such components. Our derived model is implemented by extending the capabilities of the CloudSim framework and is further validated through empirical comparison and statistical hypothesis tests. We illustrate several examples of this work's practical applicability in the domain of resource management and energy-efficiency. Ismael Solís Moreno, Peter Garraghan, Paul Townend, Jie Xu 0007 |
IEEE Trans. Cloud Comput. | 2 |
| 2013 | An Analysis of the Server Characteristics and Resource Utilization in Google CloudabstractUnderstanding the resource utilization and server characteristics of large-scale systems is crucial if service providers are to optimize their operations whilst maintaining Quality of Service. For large-scale data enters, identifying the characteristics of resource demand and the current availability of such resources, allows system managers to design and deploy mechanisms to improve data enter utilization and meet Service Level Agreements with their customers, as well as facilitating business expansion. In this paper, we present a large-scale analysis of server resource utilization and a characterization of a production Cloud data enter using the most recent data enter trace logs made available by Google. We present their statistical properties, and a comprehensive coarse-grain analysis of the data, including submission rates, server classification, and server resource utilization. Additionally, we perform a fine-grained analysis to quantify the resource utilization of servers wasted due to the early termination of tasks. Our results show that data enter resource utilization remains relatively stable at between 40 - 60%, that the degree of correlation between server utilization and Cloud workload environment varies by server architecture, and that the amount of resource utilization wasted varies between 4.53 - 14.22% for different server architectures. This provides invaluable real-world empirical data for Cloud researchers in many subject areas. Peter Garraghan, Paul Townend, Jie Xu 0007 |
IC2E | 1 |