VLDB 2026 Research / reviewers in the wild / expert
J. Lakshmi
dblp:36/573
· DBLP profile ↗
22ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0001-5484-9622ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Serverless Platform for Smart Deployment of Service Function ChainsabstractCloud Data Centres deal with dynamic changes all the time. Networks in particular, need to adapt their configurations to changing workloads. Given these expectations, Network Function Virtualization (NFV) using Software Defined Networks (SDNs) has realized the aspect of programmability in networks. NFVs allow network services to be programmed as software entities that can be deployed on commodity clusters in the Cloud. Being software, they inherently carry the ability to be customized to specific tenants’ requirements and thus support multi-tenant variations with ease. However, the ability to exploit scaling in alignment with changing demands with minimal loss of service, and improving resource usage efficiency still remains a challenge. Several recent works in literature have proposed platforms to realize Virtual Network functions (VNFs) on the Cloud using service offerings such as Infrastructure as a Service (IaaS) and serverless computing. These approaches are limited by deployment difficulties (configuration and sizing), adaptability to performance requirements (elastic scaling), and changing workload dynamics (scaling and customization). In the current work, we propose a Hybrid Serverless Platform (HSP) to address these identified lacunae. The HSP is implemented using a combination of persistent IaaS, and FaaS components. The IaaS components handle the steady state load, whereas the FaaS components activate during the dynamic change associated with scaling to minimize service loss. The HSP controller takes provisioning decisions based on Quality of Service (QoS) rules and flow statistics using an auto recommender, alleviating users of sizing decisions for function deployment. HSP controller design exploits data locality in SFC realization, reducing data-transfer times between VNFs. It also enables the usage of application characteristics to offer higher control over SFC deployment. A proof-of-concept realization of HSP is presented in the paper and is evaluated for a representative Service Function Chain (SFC) for a dynamic workload, which shows minimal loss in flowlet service, up to 35% resource savings as compared to a pure IaaS deployment and up to 55% lower end-to-end times as compared to a baseline FaaS implementation. Sheshadri K. R, J. Lakshmi |
IEEE Trans. Cloud Comput. | 2 |
| 2024 | Ataru: A Lightweight VMM + Runtime for Low Latency Serverless FunctionsabstractServerless computing is a paradigm that allows application developers to focus on defining functions triggered by events, while the service provider handles resource allocation, isolation, scalability, and orchestration. Function-as-a-Service (FaaS) is a popular implementation of serverless coputing, characterized by short-lived stateless functions. However, current FaaS platforms suffer from high overheads of boot-strapping the process for function execution, which degrade the performance of applications that require low latency and high throughput. Moreover, current FaaS platforms do not support sharing memory among function instances of the same workflow, which limits the efficiency and functionality of applications that rely on data dependencies. In this paper, we propose Ataru, a native function execution virtualization construct that exploits full hardware potential with the provision for shared memory and minimal process bootstrapping overhead, without sacrificing the security offered by virtualization technologies. Ataru consists of two components: Ataru-KVM, a lightweight Virtual Machine Monitor (VMM) based on KVM that supports fast bootup and dynamic suspension of virtual CPUs (vCPUs), and Ataru Runtime, a runtime system that manages the execution of an application defined by Directed Acyclic Graph (DAG) of functions and memory sharing within the virtual machine (VM). We evaluate Ataru against a process-based solution over Firecracker that offers similar VM isolation as Ataru, and show that Ataru outperforms Firecracker significantly in terms of bootup time, function execution time, and vCPU utilization, especially when the functions have execution times in the order of tens of microseconds. Prakhar Gupta, J. Lakshmi |
CloudCom | 2 |
| 2024 | RightFusion: Enabling QoS driven Function Fusion in Edge-Cloud FaaSabstractFunction-as-a-Service (FaaS), a form of serverless computing, allows applications to be composed as a workflow of stateless, event-driven, short-lived functions that are deployed automatically by the service provider. FaaS has its origins in the Cloud, but recently, it has seen considerable adoption in the Edge-Computing too. FaaS exhibits significant overheads between invocations of consecutive functions and function fusion is being used to resolve these. The fused function groups are assigned the highest resource of all the functions in the fusion group and are limited by the fixed resource sizes associated with the functions; this translates to high function execution costs. In this trade-off, the function size and execution time, which is associated with input size and user-performance expectation, are not considered for fusion decisions on Heterogeneous resources. In this paper, we present ‘RightFusion”, a novel fusion technique that performs function-fusion on an Edge-Cloud resource pool based on application input and QoS requirements. The platform right-sizes resource instances and performs cost-efficient function fusion while meeting the QoS requirements. It also exploits heterogeneous resource cost difference to choose resources for fused functions deployment. Evaluations based on a diverse range of applications show that RightFusion reduces resource usage by above $35 \%$ at Edge and Cloud while meeting the QoS requirements compared to the PeakFusion baseline. Sheshadri K. R, J. Lakshmi |
IC2E | 2 |
| 2023 | Hybrid Serverless Platform for Service Function ChainsabstractCloud Data Centres deal with dynamic changes all the time. Networks in particular, need to adapt their configurations to changing workloads. Given these expectations, Network Function Virtualization (NFV) using Software Defined Networks (SDNs) have realized the aspect of programmability in networks, bringing in the necessary fidelity to managing network resources. NFVs allows network services to be programmed as software entities that can be deployed on commodity clusters in the Cloud. Being software, they inherently carry the ability to be customized to specific tenants' requirements and thus support multi-tenant variations with ease. However, the ability to exploit scaling in alignment with changing demands with minimal loss of service and improving resource usage efficiency still remains a challenge. Several recent works in literature have proposed platforms to realize Virtual Network functions (VNFs) on the Cloud using different services such as Infrastructure as a Service (IaaS) and serverless computing. These approaches suffer from deployment difficulties (configuration and sizing) or adaptability to performance requirements or changing dynamics. In the current work, we propose a Hybrid Serverless Platform (HSP) to address these identified lacunae. The HSP is implemented using a combination of persistent IaaS and FaaS components. The IaaS components handle the steady state load, whereas the FaaS components activate during the dynamic change associated with scaling to minimize service loss. The HSP's controller takes provisioning decisions on Quality of Service (QoS) attributes derived from flow statistics, thereby alleviating sizing decisions for deployment. A proof-of-concept realization of HSP is presented in the paper and is evaluated for an example Service Function Chain (SFC) scenario for a dynamic workload, which shows minimal loss in flowlet service, up to 35% resource savings as compared to a pure IaaS deployment and up to 55% lower end-to-end SFC execution times as compared to a baseline FaaS implementation. Sheshadri K. R, J. Lakshmi |
CLOUD | 2 |
| 2023 | End-to-end Resiliency Analysis Framework for Cloud Storage ServicesabstractCloud storage services are becoming increasingly complex with the stored data volume and operational scale. The complexity is the result of the ever-growing components with various functionalities being involved for service rendition, exposing the service to numerous failures or disruptions. In such scenarios, resiliency becomes one of the most crucial parameters to evaluate how well-equipped the system is to withstand the effect of disruptions and maintain a reliable service. The existing evaluations primarily focus on the resiliency of stored user data, which is insufficient to project the storage service level resiliency. This work proposes an end-to-end resiliency framework for cloud storage services that enables the assessment of the overall resiliency. The framework is based on three properties – service expectations, crucial components for service rendition, and their resiliency to meet the expectations while facing various disruptions. The framework is used to model the resiliency of two distinct and well-known storage services, OpenStack Swift and CephFS, as Stochastic Petri Nets. The models enable the effective quantification of system resiliency through the achieved service reliability. Archita Ghosh, J. Lakshmi |
PRDC | 2 |
| 2022 | Towards More Effective and Explainable Fault Management Using Cross-Layer Service TopologyabstractAs microservice architecture becomes prominent, existing fault management techniques to deal with service disruption become limiting mainly due to the amount of data needed to be analyzed. This paper emphasizes the need to consider the cross-layer topology of the cloud service to intelligently identify and correlate the observability data and assist in implementing efficient and more accurate fault management techniques that can provide better explainability. Towards this goal, the paper presents a tool that discovers the cross-layer topology for a cloud microservice application and discusses the benefits of using cross-layer service topology to implement effective fault management. Dhanya R. Mathews, Mudit Verma, J. Lakshmi, Pooja Aggarwal |
CLOUD | 3 |
| 2022 | QoS aware FaaS for Heterogeneous Edge-Cloud continuumabstractFunction as a Service (FaaS) is one of the widely used serverless computing service offerings to build and deploy applications on the Cloud. The platform is popular for its "pay-as-you-go" billing model, microservice-based design, event-driven executions, and autonomous scaling. Although it has its firm roots in Cloud computing service offerings, it is considerably explored in the Edge computing layer. The efficient resource management of FaaS is attractive to Edge computing because of the limited nature of resources. Existing literature on Edge-Cloud FaaS platforms orchestrates compute workloads based on factors such as data locality, resource availability, network costs, and bandwidth. However, the state-of-the-art platforms lack a comprehensive way to address the challenges of managing heterogeneous resources in the FaaS platform. The resource specification in a heterogeneous setting, lack of Quality of Service (QoS) driven resource provisioning, and function deployment exacerbate the problem of resource selection, and function deployment in FaaS platforms with a heterogeneous resource pool. To address these gaps, the current work presents a novel heterogeneous FaaS platform that deduces function resource specification using Machine Learning (ML) methods, performs smart function placement on Edge/Cloud based on a user-specified QoS requirement, and exploit data locality by caching appropriate data for function executions. Experimental results based on real-world workloads on a video surveillance application show that the proposed platform brings efficient resource utilization and cost savings at the Cloud by reducing the resource usage by up to 30%, while improving the performance of function executions by up to 25% at Edge and Cloud. Sheshadri K. R, J. Lakshmi |
CLOUD | 2 |
| 2022 | Understanding the Resiliency of Cloud Storage ServicesabstractA cloud storage system requires multiple functional and management layers to render a global-scale storage solution. Providing a reliable service through this complex architecture becomes a challenging task. Moreover, the mere complexity of the system makes it difficult to identify the critical components for maintaining a reliable service. While data redundancy has been the predominant factor in improving reliability, it becomes crucial to understand if it is sufficient for the reliability of a cloud storage service. This work proposes a resiliency evaluation method that identifies the components necessary for storage service rendition and the ability of the system to absorb the effect of their failures and minimize the impact on service. Using this method, the resiliency of two distinct cloud storage services, OpenStack Swift and CephFS, are evaluated. The evaluation has revealed that data access requests may get delayed (up to 8x of mean response time) and even fail due to the lack of resiliency for access path components, even when there is enough user data redundancy. The lack of effort is evident while maintaining the consistency of critical internal data components resulting in reduced reliability. Even for stored data, some common failures result in loss of recoverability. The work concludes with some general observations and possible solutions that will help to improve storage service reliability. Archita Ghosh, J. Lakshmi |
PRDC | 2 |
| 2021 | Insights into Multi-Layered Fault Propagation and Analysis in a Cloud StackabstractEmerging application modernisation efforts are pushing new application services to be built and existing monoliths to be refactored as loosely coupled distributed components (e.g, microservices) for independent scaling and management in cloud. With dynamic operating conditions, component failures, complex component interconnections across the cloud stack, etc., it becomes a challenge to develop effective fault management techniques at the granularity of a multi-layered cloud application service. This paper emphasises on considering faults, errors and failure across the components in different layers of a cloud stack for effective fault management. Dhanya R. Mathews, Mudit Verma, Pooja Aggarwal, J. Lakshmi |
CLOUD | 4 |
| 2021 | QoS aware FaaS platformabstractFunction as a Service (FaaS), a form of serverless computing, is one of the recent cloud computing service offerings, which abstracts and automates the management and provisioning of resources, and deployment of applications. It provides powerful abstractions to compose applications as stateless functions and triggers their executions through events. The platform offers autonomous scaling for applications and pay-as-you-go sub-second billing model. However, contemporary FaaS platforms provide limited features in stating resource requirements. They often lack specifications to express application specificities and resource requirements associated with Quality of Service (QoS). Such specifications can effectively guide the resource provisioning and function deployment at the resource provider, leading to efficient resource utilization and cost savings. This research exploration motivates the need for a QoS specification framework for FaaS and proposes ideas for realizing an initial QoS aware FaaS platform. Experimental results based on real-world workload trace show the cost savings and efficient resource utilization that QoS can bring in FaaS platforms. Sheshadri K. R, J. Lakshmi |
CCGRID | 2 |
| 2020 | A Hierarchical Control Plane Framework for Integrated SDN-SFC Management in Multi-tenant Cloud Data CentersabstractService Functions have become a prominent part of cloud data center networks. Providing qualitative services to the hosted applications, the provision and management of service function chain components require an autonomous and reactive management framework that can respond to the dynamics of network state and application requirements. While software defined network controllers enable programmable network configuration, current implementations do not extend its capability to manage service function chains. Hence, a common network control entity that manages not only forwarding devices, but also the service function chain components is required. With a global view of the network which includes the service functions, the integrated controller can make network configuration decisions that are globally optimized. However, this integration gives rise to the problems of controller scalability and multi-tenancy management. To this end, this paper proposes a hierarchical controller framework for software defined network controller integrated with service function chain management that addresses both these problems in cloud data centers. The proposed controller is implemented and evaluated over a simulated data center network, and a comparative evaluation is made with existing controller architectures. The proposed controller is observed to reduce packet loss by 19.81% and 9.37% in comparison to centralized and distributed controller frameworks with integrated SFC management, respectively. Also, while the centralized controller gives the least flow setup times for smaller number of tenants, beyond 70 tenants the proposed hierarchical controller gives the least flow setup times. B. S. Lakshmi, J. Lakshmi |
CLOUD | 2 |
| 2019 | QuADD: QUantifying Accelerator Disaggregated Datacenter EfficiencyabstractIn the current era of data explosion accelerators such as GPUs facilitate data-driven applications with requisite compute boost. Availability of GPUs in Public Cloud offerings has expedited their mass adoption. Consequently, varied customer demands and exclusive allocation of GPUs to VMs can leave stranded GPUs across the datacenter. Hardware disaggregation can alleviate this issue to enable a powerefficient datacenter. However, it is important to first quantify the gains associated with this new deployment paradigm. In the absence of real deployments, simulations can be helpful to evaluate the benefits of disaggregation at scale. In this paper, we evaluate the gains associated primarily with disaggregated GPU deployments. For this, we use QUADD-SIM, a simulator we built to model, quantify, and contrast different facets of these emerging GPU deployments. Using QUADD-SIM we model different VM and resource provisioning aspects of disaggregated GPU deployments. We simulate realistic AI workload requests for a period of 3 months with characteristics derived from recent public datacenter traces. Our results attest that disaggregated GPU deployment strategies outperform traditional GPU deployments in terms of failed VM requests and GPU Watt-hours consumption. Overall, 5.14% and 7.90% additional failed VM requests were serviced by disaggregated GPU deployments consuming 10.92% and 3.30% lesser GPU Watt-hours compared to traditional deployment. Anubhav Guleria, J. Lakshmi, Chakri Padala |
CLOUD | 2 |
| 2019 | Integrating Service Function Chain Management into Software Defined Network ControllerabstractMiddleboxes play a crucial role in networks in the context of performance and security aspects. Large scale networks such as data center networks employ a wide range of middleboxes to service application traffic. Handling of dynamic traffic by means of static middlebox configurations is a complex task. To this end, a combination of Network Function Virtualization and Software Defined Networks offer promising approaches to enable provisioning and chaining of network functions to create service function chains dynamically, in response to the changing application workload and policies. These paradigms together enable elastic and scalable provisioning of service function chains which are programmable at a high, abstract level. This opens new opportunities for co-ordinated working of cloud manager and Software Defined Network controller to jointly manage service function instances as flexibly as the application Virtual Machines, based on the temporal network characteristics of the data center network. On the other hand, the practical implementation of such a controller requires overcoming various hurdles. In this paper we enumerate and motivate the need for such a setup in large scale data center networks followed by the challenges in realizing it. As a proof of concept we realize such a setup using simulation environment and present some early results on the effect on Software Defined Network controller design. B. S. Lakshmi, J. Lakshmi |
SERVICES | 2 |
| 2018 | Migrating VM Workloads to Containers: Issues and ChallengesabstractVirtualization technologies such as KVM and XEN have served the purpose of workload consolidation while providing required isolation. With the advent of Linux Docker platform, micro-service architectures have gained popularity. Adaptation of containers have changed the way, the applications are architected, developed, deployed and managed. Compared to Virtualization, containers provide lightweight alternative to co-host multiple applications on single server at the cost of isolation. This paper highlights the issues and challenges associated with migrating VM based workloads to Container platforms. A systematic analysis, illustrated through a representative application benchmark chosen from a real-life e-commerce private cloud setup is used to present these aspects. Specifically, the application workflow that is currently hosted on VMs is chosen and re-casted on to containers platform and a critical study is conducted based on resource sharing, concurrency, isolation and dependability parameters. Surya Kant Garg, J. Lakshmi, Jain Johny |
IEEE CLOUD | 2 |
| 2018 | Towards Improving Data Center Utilisation by Reducing FragmentationabstractMany enterprise organizations, like large e-commerce platforms, are consolidating their ICT infrastructure on private clouds to harness cloud properties of scalability and elasticity. Administratively many of these cloud deployments use pre-defined VM sizes to ease effort of use and placement. With this sizing, workloads’ requirements poorly match with the offered sizes leading to poor VM and host utilisations. This paper proposes alternate VM sizing approaches to overcome this and compares internal, external resource fragmentation in each case. To evaluate this, utilisation traces from the private cloud of Flipkart, having thousands of servers and VMs are used. Evaluations show that applications as well as host utilisation benefits accrue by customising VM sizes to application requirements. We also show how VM sizes and host configurations influence consolidation ratios and utilisation. Shravan S. K, J. Lakshmi, Neeraj Bisht |
IEEE CLOUD | 2 |
| 2018 | Cost-Benefit Analysis of Public Clouds for Offloading In-House HPC JobsabstractThe Supercomputing Education and Research Centre (SERC) at the Indian Institute of Science (IISc), located in the South East Asian Region, has been providing state-of-the-art services and support for High Performance Computing (HPC) for the academic users of the institute since 1990. This centre, facilitates OpenMP, MPI, CUDA and OpenCL based HPC jobs for academic research. To augment or support the existing demand for compute intensive, tightly coupled HPC workloads in the institute, it was desirable to explore cloud based HPC services. This study is a comparative cost-benefit analysis of the in-house compute facility at SERC with that of the commercial cloud providers like Amazon, Google Cloud, Microsoft Azure and Sabalcore. The Total Cost of Ownership (TCO) of SahasraT (the in-house supercomputer) is computed and used to identify the cost of running a small job for 24 hours on the in-house facility. This is then compared with the estimated cost of running the same job on comparative cloud instances. Several research papers that provide a detailed evaluation of HPC on the cloud with performance benchmarks and the comparison to in-house HPC facilities, were studied. These references have been used to compare performance of compute instance in the public cloud to the performance of HPC at SahasraT. The results obtained in this study, advocate that current cloud platforms are expensive in both cost and performance as compared to the in-house facility at SERC Akhila Prabhakaran, J. Lakshmi |
IEEE CLOUD | 2 |
| 2015 | Location Obfuscation for Location Data PrivacyabstractAdvances in wireless internet, sensor technologies, mobile technologies, and global positioning technologies have renewed interest in location based services (LBSs) among mobile users. LBSs on smartphones allow consumers to locate nearby products and services, in exchange of their location information. Precision of location data helps for accurate query processing of LBSs but it may lead to severe security violations and several privacy threats, as intruders can easily determine user's common paths or actual locations. Encryption is the most explored approach for ensuring security. It can give protection against third party attacks but it cannot provide protection against privacy threats on the server which can still obtain user location and use it for malicious purposes. Location obfuscation is a technique to protect user privacy by altering the location of the users while preserving capability of server to compute few mathematical functions which are useful for the user over the obfuscated location information. This work mainly concentrates on LBSs which wants to know the distance travelled by user for providing their services and compares encryption and obfuscation techniques. This study proposes various methods of location obfuscation for GPS location data which are used to obfuscate user's path and location from service provider. Our work shows that user privacy can be maintained without affecting LBSs results, and without incurring significant overheads. Vaibhav Ankush Kachore, J. Lakshmi, S. K. Nandy 0001 |
SERVICES | 2 |
| 2015 | PriDyn: Enabling Differentiated I/O Services in Cloud Using Dynamic PrioritiesabstractVirtualization is one of the key enabling technologies for Cloud computing. Although it facilitates improved utilization of resources, virtualization can lead to performance degradation due to the sharing of physical resources like CPU, memory, network interfaces, disk controllers, etc. Multi-tenancy can cause highly unpredictable performance for concurrent I/O applications running inside virtual machines that share local disk storage in Cloud. Disk I/O requests in a typical Cloud setup may have varied requirements in terms of latency and throughput as they arise from a range of heterogeneous applications having diverse performance goals. This necessitates providing differential performance services to different I/O applications. In this paper, we present PriDyn, a novel scheduling framework which is designed to consider I/O performance metrics of applications such as acceptable latency and convert them to an appropriate priority value for disk access based on the current system state. This framework aims to provide differentiatedI/O service to various applications and ensures predictable performance for critical applications in multi-tenant Cloud environment. We demonstrate through experimental validations on real world I/O traces that this framework achieves appreciable enhancements in I/O performance, indicating that this approach is a promising step towards enabling QoS guarantees on Cloud storage. Nitisha Jain, J. Lakshmi |
IEEE Trans. Serv. Comput. | 2 |
| 2014 | PriDyn: Framework for Performance Specific QoS in Cloud StorageabstractVirtualization is one of the key enabling technologies for cloud computing. Although it facilitates improved utilization of resources, virtualization can lead to performance degradation due to the sharing of physical resources like CPU, memory, network interfaces, disk controllers, etc. Multi-tenancy can cause highly unpredictable performance for concurrent I/O applications running inside virtual machines that share local disk storage in cloud. Disk I/O requests in a typical cloud setup may have varied requirements in terms of latency and throughput as they arise from a range of heterogeneous applications having diverse performance goals. This necessitates providing differential performance services to different I/O applications. In this paper, we present PriDyn, a novel scheduling framework which is designed to consider I/O performance metrics of applications such as acceptable latency and convert them to an appropriate priority value for disk access based on the current system state. This framework aims to provide differentiated I/O service to various applications and ensures predictable performance for critical applications in multi-tenant cloud environment. We demonstrate that this framework achieves appreciable enhancements in I/O performance indicating that this approach is a promising step towards enabling QoS guarantees on cloud storage. Nitisha Jain, J. Lakshmi |
IEEE CLOUD | 2 |
| 2013 | Elastic Resources Framework in IaaS, Preserving Performance SLAsabstractElasticity in cloud systems provides the flexibility to acquire and relinquish computing resources on demand. However, in current virtualized systems resource allocation is mostly static. Resources are allocated during VM instantiation and any change in workload leading to significant increase or decrease in resources is handled by VM migration. Hence, cloud users tend to characterize their workloads at a coarse grained level which potentially leads to under-utilized VM resources or under performing application. A more flexible and adaptive resource allocation mechanism would benefit variable workloads, such as those characterized by web servers. In this paper, we present an elastic resources framework for IaaS cloud layer that addresses this need. The framework provisions for application workload forecasting engine, that predicts at run-time the expected demand, which is input to the resource manager to modulate resource allocation based on the predicted demand. Based on the prediction errors, resources can be over-allocated or under-allocated as compared to the actual demand made by the application. Over-allocation leads to unused resources and under allocation could cause under performance. To strike a good trade-off between over-allocation and under-performance we derive an excess cost model. In this model excess resources allocated are captured as over-allocation cost and under-allocation is captured as a penalty cost for violating application service level agreement (SLA). Confidence interval for predicted workload is used to minimize this excess cost with minimal effect on SLA violations. An example case-study for an academic institute web server workload is presented. Using the confidence interval to minimize excess cost, we achieve significant reduction in resource allocation requirement while restricting application SLA violations to below 2-3%. Mohit Dhingra, J. Lakshmi, S. K. Nandy 0001, Chiranjib Bhattacharyya, K. Gopinath |
IEEE CLOUD | 2 |
| 2013 | Virtual Machine Placement Optimization Supporting Performance SLAsabstractCloud computing model separates usage from ownership in terms of control on resource provisioning. Resources in the cloud are projected as a service and are realized using various service models like IaaS, PaaS and SaaS. In IaaS model, end users get to use a VM whose capacity they can specify but not the placement on a specific host or with which other VMs it can be co-hosted. Typically, the placement decisions happen based on the goals like minimizing the number of physical hosts to support a given set of VMs by satisfying each VMs capacity requirement. However, the role of the VMM usage to support I/O specific workloads inside a VM can make this capacity requirement incomplete. I/O workloads inside VMs require substantial VMM CPU cycles to support their performance. As a result, placement algorithms need to include the VMM's usage on a per VM basis. Secondly, cloud centers encounter situations wherein change in existing VM's capacity or launching of new VMs need to be considered during different placement intervals. Usually, this change is handled by migrating existing VMs to meet the goal of optimal placement. We argue that VM migration is not a trivial task and does include loss of performance during migration. We quantify this migration overhead based on the VM's workload type and include the same in placement problem. One of the goals of the placement algorithm is to reduce the VM's migration prospects, thereby reducing chances of performance loss during migration. This paper evaluates the existing ILP and First Fit Decreasing (FFD) algorithms to consider these constraints to arrive at placement decisions. We observe that ILP algorithm yields optimal results but needs long computing time even with parallel version. However, FFD heuristics are much faster and scalable algorithms that generate a sub-optimal solution, as compared to ILP, but in time-scales that are useful in real-time decision making. We also observe that including VM migration overheads in the placement algorithm results in a marginal increase in the number of physical hosts but a significant, of about 84 percent reduction in VM migration. Ankit Anand, J. Lakshmi, S. K. Nandy 0001 |
CloudCom (1) | 2 |
| 2006 | Framework for Enabling Highly Available Distributed Applications for Utility Computing
J. Lakshmi, S. K. Nandy 0001, Ranjani Narayan, Keshavan Varadarajan |
ISPA | 1 |