Purushottam Kulkarni

dblp:48/2303 · also Puru Kulkarni · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
8since 2021 · last 2026
0009-0008-0272-9299ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 1 first-author · 5 since 2021Computer networks · 9 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 FLYT: Transparent and Elastic GPU Provisioning for Multi-Tenant Cloud Services
abstract
Modern cloud services such as AI inference, video analytics, and scientific computing exhibit highly variable and bursty GPU demand patterns that static provisioning and coarse-grained sharing mechanism struggle to accommodate efficiently. Existing GPU multiplexing approaches, including NVIDIA MPS and MIG, provide limited flexibility in multi-tenant environments, often leading to resource fragmentation, under-utilization, or unpredictable latency. We present Flyt, a transparent, latency-0aware GPU orchestration framework for virtualized cloud services. Flyt enables fine-grain runtime scaling of Streaming Multiprocessors (SMs) and breaks the traditional VM–GPUs binding by allowing applications inside a VM to execute on different GPUs over time. This design supports elastic scaling and live inter–node GPU migration without application or guest OS modifications, by virtualizing GPU memory through address translation and enforcing elastic SM execution caps.
Santhosh M. Kumar, Sameer Ahmad, Armaan Chowfin, Purushottam Kulkarni, Anand Eswaran, Praveen Jayachandran
ICPE4
2025 FLASH: Fast Linked AF_XDP Sockets for High Performance Network Function Chains
abstract
Recent advances in the Linux kernel, such as eXpress Data Path (XDP) and AF_XDP sockets, enable high-speed packet processing for software NFs while preserving access to kernel features. However, the default AF_XDP implementation in the Linux kernel does not permit easy and performant NF chaining, e.g., zero-copy transfer of packets across NFs co-located on the same host. While prior work has proposed solutions for optimized NF chaining in the context of kernel bypass frameworks like DPDK that operate entirely in userspace, such solutions do not extend easily to AF_XDP, because the AF_XDP datapath is fragmented across the kernel driver and userspace. This paper introduces FLASH, a low-overhead inkernel chaining mechanism for AF_XDP sockets. FLASH enables zero-copy packet transfers for FLASH-native NFs, and single-copy packet transfers for legacy AF_XDP NFs. Further, via integration with K8s, FLASH supports the deployment of unprivileged containerized NFs on cloud platforms. Our work contributes several novel modifications to the AF_XDP datapath in the kernel to implement optimized NF chaining, and provides userspace libraries/APIs to easily build NFs that leverage FLASH. Our evaluations show that FLASH matches the performance of userspace DPDK-based NF chaining frameworks, while outperforming the best available AF_XDP-based alternatives by up to 2.5× in throughput, and achieving the lowest latency among all NF chaining frameworks.
Debojeet Das, Kevin Prafull Baua, Aditya Kansara, Arghyadip Chakraborty, Dheeraj Kurukunda, Mythili Vutukuru, Purushottam Kulkarni
SoCC7
2023 DAGit: A Platform For Enabling Serverless Applications
abstract
Serverless computing is rapidly gaining popularity for provisioning composable, auto-scalable and cost-effective applications. An important mechanism for deploying serverless applications is specification of function workflows (via DAGs). The end-to-end life cycle of this process being DAG specification, DAG orchestration, execution of DAG components and persistent storage of application outputs. To the best of our knowledge, an open-source platform that offers functionality along all these components does not exist. Towards this, our primary contribution is DAGit, an open-source solution for serverless applications-as-a-service. The main features of DAGit are interfaces and specifications to register serverless functions, applications (via DAGs) and triggers to instantiate the serverless applications. DAGit provides a rich set of DAG primitives to enable a varied set of applications and also implements a scalable orchestrator for application execution. As part of this work, we present the architecture and design details of DAGit, and demonstrate its feature set via showcasing the specification and execution of a varied set of serverless applications. Further, we also present a performance and resource costs characterization of executing applications on the DAGit platform.
Anubhav Jana, Purushottam Kulkarni, Umesh Bellur
HiPC2
2022 Portkey: hypervisor-assisted container migration in nested cloud environments
abstract
Derivative cloud service providers use nesting to provision virtual computational entities (VCE) within VCEs, e.g., containers runtimes within virtual machines. As part of resource management and ensuring application performance, migration of nested containers is an important and useful mechanism. Checkpoint Restore In Userspace (CRIU) is the dominant method for migration, used by Docker and other container technologies. While CRIU works well for container migration from host to host, it suffers from significant increase in resource requirements in nested setups. The overheads are primarily due to the high network virtualization overhead in nested environments. While techniques such as SR-IOV can mitigate the overheads, they require additional hardware features and tight coupling of network endpoints. Based on our insights of network virtualization being the main bottleneck, we present Portkey - a software-based solution for efficient nested container migration that significantly reduces CPU utilization at both the source and destination hosts. Our solution relies on interposing a layer that directly coordinates network IO from within a virtual machine with the hypervisor. A new set of hypercalls provide this interfacing along with a control loop that minimizes the hypercall path usage. Extensive evaluation of our solution shows that Portkey reduces CPU usage by up to 75% and 82% at the source and destination hosts, respectively.
Debadatta Mishra, Purushottam Kulkarni, Umesh Bellur
VEE3
2021 Speedo: Fast dispatch and orchestration of serverless workflows
abstract
Structuring cloud applications as collections of interacting fine-grained microservices makes them scalable and affords the flexibility of hot upgrading parts of the application. The current avatar of serverless computing (FaaS) with its dynamic resource allocation and auto-scaling capabilities make it the deployment model of choice for such applications. FaaS platforms operate with user space dispatchers that receive requests over the network and make a dispatch decision to one of multiple workers (usually a container) distributed in the data center. With the granularity of microservices approaching execution times of a few milliseconds combined with loads approaching tens of thousands of requests a second, having a low dispatch latency of less than one millisecond becomes essential to keep up with line rates. When these microservices are part of a workflow making up an application, the orchestrator that coordinates the sequence in which microservices execute also needs to operate with microsecond latency. Our observations reveal that the most significant component of the dispatch/orchestration latency is the time it takes for the request to traverse into and out of the user space from the network. Motivated by the presence of a multitude of low power cores on today's SmartNICs, one approach to keeping up with these high line rates and the stringent latency expectations is to run both the dispatcher and the orchestrator close to the network on a SmartNIC. Doing so will save valuable cycles spent in transferring requests to and back from the user space. The operating characteristics of short-lived ephemeral state and low CPU burst requirements of FaaS dispatcher/orchestrator make them ideal candidates for offloading from the server to the NIC cores. This also brings other benefit of freeing up the server CPU. In this paper, we present Speedo--- a design for offloading of FaaS dispatch and orchestration services to the SmartNIC from the user space. We implemented Speedo on ASIC based Netronome Agilio SmartNICs and our comprehensive evaluation shows that Speedo brings down the dispatch latency from ~150ms to ~140μs at a load of 10K requests per second.
Nilanjan Daw, Umesh Bellur, Purushottam Kulkarni
SoCC3
2021 FaaSter: Accelerated Functions-as-a-Service with Heterogeneous GPUs
abstract
In this work, we present FaaSter, an Accelerated Functions as a Service (AFaaS) offering that unifies the function-as-a-service model with GPU acceleration resources. FaaSter provides an acceleration function library as a service, which in turn is provisioned on heterogeneous GPUs. To provide seamless access to these accelerated functions and ensure that each function has the best possible response time, we utilize GPU kernel slicing to split and execute an accelerated function instance across multiple heterogeneous GPUs. The central challenge is to be able to quickly decide the number of slices to split each function into and then map the slices to the right GPUs. To this end, we present a scheduling heuristic that is able to significantly reduce the average turn-around time of functions when compared to a non-sliceable full GPU scheduling approach. Our evaluation results show that the FaaSter scheduler achieves 62 % mean and up to 80 % improvement in average turn-around time and always performs equal to or better than non-sliceable GPU scheduling.
Anshuj Garg, Purushottam Kulkarni, Umesh Bellur, Sriram Yenamandra
HiPC2
2021 Optimizing Goodput of Real-time Serverless Functions using Dynamic Slicing with vGPUs
abstract
As the popularity and relevance of the Function-as-a-Service (FaaS) model keeps growing, we believe newer avatars of the service will support computationally intensive SIMT functions that will execute on GPUs. With hardware-assisted virtualization of GPUs now possible, cloud offerings including GPUs usually bind a virtual GPU (vGPU) to a VM. While there is a choice of scheduling algorithms to multiplex vGPUs on to the physical GPU, the work-conserving best-effort scheduler helps to maintain a high level of utilization of the GPU. With this, we observe that the total share of the GPU per VM is non-deterministic and depends on how different VMs load the GPU via their vGPUs. As a result, any function-to-vGPU scheduler that does not explicitly account for this nondeterministic vGPU capacity will suffer from lower than optimal goodput - particularly when these functions are deadline bound as is the case with FaaS offerings today. In this work, we exploit a software based task slicing technique to dynamically determine task sizes for scheduling on vGPUs to maximize successful completion of functions within their deadlines. Our solution extends the conventional earliest-deadline first (EDF) scheduling algorithm by balancing scheduling opportunities (via kernel slicing) and maximizing the chances of functions finishing before their deadline. The work is motivated by the fact that static decisions that consider entire tasks as scheduling units or use a fixed, statically decided slice size as scheduling units cannot adapt to the non-deterministic vGPU capacity. A comparison of our solution with well-known deadline aware scheduling approach (earliest deadline first), yielded an improvement of up to 2.9x in goodput.
Anshuj Garg, Umesh Bellur, Purushottam Kulkarni, Uday Kurkure, Hari Sivaraman, Lan Vu
IC2E4
2021 SymFlex: Elastic, Persistent and Symbiotic SSD Caching in Virtualization Environments
abstract
Hypervisor managed SSD caching is an often used technique for improving IO performance in virtualization based hosting solutions. Such caches are either explicitly managed by the hypervisor which approximate the access semantics of the applications for improving cache utilization, or operate as statically partitioned devices (which are utilized as caches) by virtual machines. We reason that both these broad directions do not exploit the potential of SSD based IO caches to the fullest, in terms of generalized management policies and performance. We propose SymFlex, a novel method to perform symbiotic management of IO caches by enabling elastic SSD devices. Each virtual machine is configured with an elastic virtual SSD whose contents can be managed according to guest OS and application semantics and requirements. Furthermore, the SSD sizing is managed by the hypervisor with a ballooning-like mechanism to dynamically adjust SSD provisioning to VMs based on performance and usage fairness policies. Our primary contribution of this work is to design and engineer the mechanism for elastic SSD disks to be virtualized, and demonstrate usage models and effectiveness of the symbiotic management of SSD caches across virtual machines. Through our empirical evaluation, we show that the overhead of implementing a virtio-based elastic SSD device is minimal (within 5% of virtio based device virtualization techniques). Further, we demonstrate using dm-cache and Fatcache, the applicability and benefits of SymFlex for enhancing IO throughput and enforcing VM-level SSD allocation policies.
Muhammed Unais P, Purushottam Kulkarni
ICPE2
2020 Xanadu: Mitigating cascading cold starts in serverless function chain deployments
abstract
Organization of tasks as workflows are an essential feature to expand the applicability of the serverless computing framework. Existing serverless platforms are either agnostic to function chains (workflows as a composition of functions) or rely on naive provisioning and management mechanisms of the serverless framework---an example is that they provision resources after the trigger to each function in a workflow arrives thereby forcing a setup latency for each function in the workflow. In this work, we focus on mitigating the cascading cold start problem--- the latency overheads in triggering a sequence of serverless functions according to a workflow specification. We first establish the nature and extent of the cascading effects in cold start situations across multiple commercial server platforms and cloud providers. Towards mitigating these cascading overheads, we design and develop several optimizations, that are built into our tool Xanadu. Xanadu offers multiple instantiation options based on the desired runtime isolation requirements and supports function chaining with or without explicit workflow specifications. Xanadu's optimizations to address the cascading cold start problem are built on speculative and just-in-time provisioning of resources. Our evaluation of the Xanadu system reveals almost complete elimination of cascading cold starts at minimal cost overheads, outperforming the available state of the art platforms. For even relatively short workflows, Xanadu reduces platform overheads by almost 18x compared to Knative and 10x compared to Apache Openwhisk.
Nilanjan Daw, Umesh Bellur, Purushottam Kulkarni
Middleware3
2019 Empirical Analysis of Hardware-Assisted GPU Virtualization
abstract
The increasing use of Graphics Processing Unit (GPUs) for accelerating compute intensive tasks and graphics-related computations has led to their inclusion in High Performance Clusters and Cloud setups. Several cloud vendors provide virtual machine instances with GPU capabilities. With the advent of virtualization aware GPU hardware (NVIDIA vGPUs), allocating and sharing physical GPU among virtual machines has become easier and cost-efficient. The sharing mechanism and extent of sharing are determined by the vGPU scheduling algorithm and a configurable vGPU profile. As part of this work, we present a thorough empirical study of the hardware-assisted virtualized GPU setups. In particular, we quantify the virtualization overheads, study the interference effects of concurrently executing homogeneous and heterogeneous workloads, and impact of vGPU scheduling algorithms. We also demonstrate that the best vGPU configuration parameters are sensitive to the mix of workload characteristics that share the GPU. Our study also compares the performance of vGPU based virtualization with PCI passthrough based direct GPU assignment to virtual machines. Based on our evaluation using heterogeneous workloads setups and varying vGPU configurations, we observe that the virtualization overheads are up to 7% (in terms of reduced memory availability) and up to 20% increase in execution times. Further, we demonstrate using our setups, that co-placing of heterogeneous workloads can improve the efficiency of GPU multiplexing and decrease execution times to as low as 20% as compared to homogeneous workload placements.
Anshuj Garg, Purushottam Kulkarni, Uday Kurkure, Hari Sivaraman, Lan Vu
HiPC2
2019 Dynamic Memory Management for GPU-Based Training of Deep Neural Networks
abstract
Deep learning has been widely adopted for different applications of artificial intelligence-speech recognition, natural language processing, computer vision etc. The growing size of Deep Neural Networks (DNNs) has compelled the researchers to design memory efficient and performance optimal algorithms. Apart from algorithmic improvements, specialized hardware like Graphics Processing Units (GPUs) are being widely employed to accelerate the training and inference phases of deep networks. However, the limited GPU memory capacity limits the upper bound on the size of networks that can be offloaded to and trained using GPUs. vDNN addresses the GPU memory bottleneck issue and provides a solution which enables training of deep networks that are larger than GPU memory. In our work, we characterize and identify multiple bottlenecks with vDNN like delayed computation start, high pinned memory requirements and GPU memory fragmentation. We present vDNN++ which extends vDNN and resolves the identified issues. Our results show that the performance of vDNN++ is comparable or better (up to 60% relative improvement) than vDNN. We propose different heuristics and order for memory allocation, and empirically evaluate the extent of memory fragmentation with them. We are also able to reduce the pinned memory requirement by up to 60%.
Shriram S. B, Anshuj Garg, Purushottam Kulkarni
IPDPS3
2019 Synergy: A Hypervisor Managed Holistic Caching System
abstract
Efficient system-wide memory management is an important challenge for over-commitment based hosting in virtualized systems. Due to the limitation of memory domains considered for sharing, current deduplication solutions simply cannot achieve system-wide deduplication. Popular memory management techniques like sharing and ballooning enable important memory usage optimizations individually. However, they do not complement each other and, in fact, may degrade individual benefits when combined. We propose $\mathsf{Synergy}$Synergy, a hypervisor managed caching system to improve memory efficiency in over-commitment scenarios. $\mathsf{Synergy}$Synergy builds on an exclusive caching framework to achieve, for the first time, system-wide memory deduplication. $\mathsf{Synergy}$Synergy also enables the co-existence of the mutually agnostic ballooning and sharing techniques within hypervisor managed systems. Finally, $\mathsf{Synergy}$Synergy implements a novel file-level eviction policy that prevents hypervisor caching benefits from being squandered away due to partial cache hits. $\mathsf{Synergy}$Synergy's cache is flexible with configuration knobs for cache sizing and data storage options, and a utility-based cache partitioning scheme. Our evaluation shows that $\mathsf{Synergy}$Synergy consistently uses 10 to 75 percent lesser memory by exploiting system-wide deduplication as compared to inclusive caching techniques and achieves application speedup of 2x to 23x. We also demonstrate the capabilities of $\mathsf{Synergy}$Synergy to increase VM packing density and support for dynamic reconfiguration of cache partitioning policies.
Debadatta Mishra, Purushottam Kulkarni, Raju Rangaswami
IEEE Trans. Cloud Comput.2
2018 Deterministic Container Resource Management in Derivative Clouds
abstract
IaaS providers offer virtual machines of fix granularity which has prompted the evolution of derivative clouds. With a derivative setup where containers are provisioned within virtual machines, the guest OS manages virtual resources inside a VM whereas the hypervisor manages the physical resources distributed among VMs. This results in two control centers over the set of resources used by the containers. The hypervisor takes control actions such as memory ballooning or the withdrawal of a virtual CPU to manage over-provisioning without being aware of the effect these actions will have on individual containers inside the VM. The derivative cloud provider executing containers in the VM needs a mechanism to react to such changes—based on resource management policies setup for this purpose. In this work we first show via experimental results that hypervisor actions used to manage over-commitment such as ballooning and vCPU stealing have unpredictable and non-deterministic effects on nested containers. Based on this analysis, we design a policy driven controller that smoothes over the effect of these hypervisor actions on these nested containers. We expose several useful policies for each resource type (CPU and Memory) that can help derivative cloud providers better manage container instances.
Prashanth, Umesh Bellur, Purushottam Kulkarni
IC2E4
2018 pcube: Primitives for Network Data Plane Programming
abstract
P4 is a domain specific language to configure packet processing pipelines in programmable dataplane switches, and is a powerful idea towards realizing the goal of flexible software-defined networks. This paper presents pcube, a framework that provides a set of primitives to simplify the development of P4-based dataplane applications. pcube provides primitives for loops, summations, and other common operations on indexed state variables, which can be embedded within P4 code and unrolled by the pcube preprocessor. pcube also provides primitives to synchronize state variables across switches in distributed dataplane applications, which are automatically translated into P4 code to send and receive synchronization messages across multiple switches by pcube. We build example dataplane applications such as a distributed load balancer in our framework, and show that using pcube reduces the programming effort (in term of lines of code) significantly-by a factor of up to 5.4x.
Rinku Shah, Aniket Shirke, Akash Trehan, Mythili Vutukuru, Purushottam Kulkarni
ICNP5
2018 Cuttlefish: Hierarchical SDN Controllers with Adaptive Offload
abstract
Offloading computation to local controllers (closer to switches) has been a popular approach to designing scalable SDN controllers. We observe that, in addition to the offload of local switch-specific state, a subset of global state can also be offloaded to, and accessed at local controllers with suitable synchronization. We present the design and implementation of Cuttlefish, an SDN controller framework that adaptively offloads a portion of the application state (and computation) to local controllers. Cuttlefish uses developer-specified input to identify control messages that can be correctly processed at local controllers, and makes offloading decisions based on the cost of synchronizing the offloaded state across controllers. SDN applications use the Cuttlefish API to access the offloaded state, and Cuttlefish transparently manages the state synchronization, and redirection of control messages to the appropriate (central or local) controller. We have implemented Cuttlefish using the Floodlight SDN controller. Our evaluation shows that Cuttlefish applications achieve ~2X higher control plane throughput and ~50% lower control plane latency as compared to the traditional SDN design.
Rinku Shah, Mythili Vutukuru, Purushottam Kulkarni
ICNP3
2017 Devolve-Redeem: Hierarchical SDN Controllers with Adaptive Offloading
abstract
Towards improving SDN control plane scalability, past work has proposed SDN controller frameworks that offload computation which depends on local state to controllers residing on the switches. Our work identifies another type of computation that can be offloaded to local controllers: that which depends on state that is generated globally but can be used within local controllers with loose synchronization. Because using such state locally incurs a synchronization cost, such offload makes sense only when the benefits of the offload out-weigh the synchronization cost. We present the design and implementation of Devolve-Redeem, an SDN controller framework that can offload computation to local controllers depending on the mix of various control messages in the incoming traffic. The offload decision in our framework is made by computing a cost metric that captures the relative costs of processing every control message at the central and local controllers, taking into account synchronization costs. The SDN application developer using our framework writes a single application that runs at both the central and local controllers, using our state management API to access offloadable state. Our framework migrates between various offload modes using the computed cost metric, by manipulating the rules in the SDN switches that forward control messages to the controllers. Our framework also transparently handles state synchronization between central and local controllers in a manner that is consistent with the offload mode. We have implemented the SDN-based LTE EPC application in our framework, and experiments with our prototype demonstrate the effectiveness of our adaptive offload framework.
Rinku Shah, Mythili Vutukuru, Purushottam Kulkarni
APNet3
2017 Mitigating Nesting-Agnostic Hypervisor Policies in Derivative Clouds
abstract
The fixed granularity of virtual machines offered by IaaS providers has prompted the evolution of derivative clouds where resources are repackaged into smaller containers and leased out typically in PaaS mode. In such a setup, containers are provisioned within virtual machines. Such a nested setup results in two control centers for the resources used by those containers—the guest OS and the Hypervisor. The latter’s control actions are agnostic of the application executing within a VM. This lack of visibility may result in hypervisor control that has a non-uniform effect on the VM’s nested containers which is undesirable. In this work, we propose policy based control of the effect of the hypervisor’s control actions amongst the containers nested in the affected VM.
Prashanth, Purushottam Kulkarni, Umesh Bellur
ICDCS3
2017 DoubleDecker: a cooperative disk caching framework for derivative clouds
abstract
Derivative clouds, light weight application containers provisioned in virtual machines, are becoming viable and cost-effective options for infrastructure and software-based services. Ubiquitous dynamic memory management techniques in virtualized systems are centralized at the hypervisor and are ineffective in nested derivative cloud setups. In this paper, we highlight the challenges in management of memory resources in derivative cloud systems. Hypervisor caching, an enabler of centralized disk cache management, provides flexible memory or non-volatile memory management at the hypervisor to improve the resource usage efficiency and performance of applications. Existing hypervisor caching solutions have limited effectiveness in nested setups due to their nesting agnostic design, centralized management model and lack of holistic view of memory management. We propose DoubleDecker, a decentralized disk caching framework, realized through guest OS and hypervisor cooperation, with support for efficient memory management in derivative clouds. The DoubleDecker hypervisor caching framework, an integral part of our proposed solution, provides interfaces for differentiated cache partitioning and management in nested setups and is equipped to handle both memory and SSD based caching stores. We demonstrate the flexibility of DoubleDecker to handle dynamic and changing memory provisioning requirements and its capability to simultaneously provision memory across multiple levels. Such multi-level configurations cannot be explored by centralized designs and are a key feature of DoubleDecker. Our experimentation with DoubleDecker demonstrates that application performance can be consistently improved due to the flexible policy framework for disk caching. With our setup, we report an average performance improvement of 4x and a maximum of 11x.
Debadatta Mishra, Prashanth, Purushottam Kulkarni
Middleware3
2017 Catalyst: GPU-assisted rapid memory deduplication in virtualization environments
abstract
Content based page sharing techniques improve memory efficiency in virtualized systems by identifying and merging identical pages. Kernel Same-page Merging (KSM), a Linux kernel utility for page sharing, sequentially scans memory pages of virtual machines to deduplicate pages. Sequential scanning of pages has several undesirable side effects---wasted CPU cycles when no sharing opportunities exist, and rate of discovery of sharing being dependent on the scanning rate and corresponding CPU availability. In this work, we exploit presence of GPUs on modern systems to enable rapid memory sharing through targeted scanning of pages. Our solution, Catalyst, works in two phases, the first where pages of virtual machines are processed by the GPU to identify likely pages for sharing and a second phase that performs page-level similarity checks on a targeted set of shareable pages. Opportunistic usage of the GPU to produce sharing hints enables rapid and low-overhead duplicate detection, and sharing of memory pages in virtualization environments. We evaluate Catalyst against various benchmarks and workloads to demonstrate that Catalyst can achieve higher memory sharing in lesser time compared to different scan rate configurations of KSM, at lower or comparable compute costs.
Anshuj Garg, Debadatta Mishra, Purushottam Kulkarni
VEE3
2016 On Selecting the Right Optimizations for Virtual Machine Migration
abstract
To reduce the migration time of a virtual machine and network traffic generated during migration, existing works have proposed a number of optimizations to pre-copy live migration. These optimizations are delta compression, page skip, deduplication, and data compression. The cost-benefit analysis of these optimizations may preclude the use of certain optimizations in specific scenarios. However, no study has compared the performance & cost of these optimizations, and identified the impact of application behaviour on performance gain. Hence, it is not clear for a given migration scenario and an application, what is the best optimization that one must employ?
Senthil Nathan, Umesh Bellur, Purushottam Kulkarni
VEE3
2015 Towards a comprehensive performance model of virtual machine live migration
abstract
Although many models exist to predict the time taken to migrate a virtual machine from one physical machine to another, our empirical validation of these models has shown the 90th percentile error to be 46% (43 secs) and 159% (112 secs) for KVM and Xen live migration, respectively. Our analysis reveals that these models are fundamentally flawed as they all fail to take into account the following three critical parameters: (i) the writable working set size, (ii) the number of pages eligible for the skip technique, (iii) the relation of the number of skipped pages with the page dirty rate and the page transfer rate, and incorrectly model the key parameter---the number of new pages dirtied per unit time. In this paper, we propose a novel model that takes all these parameters into account. We present a thorough validation with 53 workloads and show that the 90th percentile error in the estimated migration times is only 12% (8 secs) and 19% (14 secs) for KVM and Xen live migration, respectively.
Senthil Nathan, Umesh Bellur, Purushottam Kulkarni
SoCC3
2014 Vagabond: Dynamic Network Endpoint Reconfiguration in Virtualized Environments
abstract
One of the biggest challenges of virtualization today is to efficiently share and manage network devices among different virtual machines (VMs). Software-based network virtualization solutions like device emulation and split driver device models have advantages of resource sharing and fine grained hypervisor resource control. However, software based approaches have performance and scalability impediments due to the software interventions for every I/O activity. Recent hardware advancements in network devices allow in-device partitioning and assignment of network functions to different guest operating systems. The nature of the assignment is static which gives rise to inflexibility in efficient network resource management. Additionally, fine grained hypervisor control on the network device is compromised because of the direct hardware assignment to the guest virtual machine.
Kallol Dey, Debadatta Mishra, Purushottam Kulkarni
SoCC3
2014 DRIVE: Using implicit caching hints to achieve disk I/O reduction in virtualized environments
abstract
Co-hosting of virtualized applications results in similar content across multiple blocks on disk, which are fetched into memory (the host's page cache). Content similarity can be harnessed both to avoid duplicate disk I/O requests that fetch the same content repeatedly, as well as to prevent multiple occurrences of duplicate content in cache. Typically, caches store the most recently or frequently accessed blocks to reduce the number of disk read accesses. These caches are referenced by block number, and can not recognize content similarity across multiple blocks. Existing work in memory deduplication merges cache pages after multiple identical blocks have already been fetched from disk into cache, while existing work in I/O deduplication reserves a portion of the host-cache to be maintained as a content-aware cache. We propose a disk I/O reduction system for the virtualization environment that addresses the dual problems of duplicate I/O and duplicate content in the host-cache, without being invasive. We build a disk read-access optimization called DRIVE, that identifies content similarity across multiple blocks, and performs hint-based read I/O redirection to improve cache effectiveness, thus reducing the number of disk reads further. A metadata store is maintained based on the virtual machine's disk accesses and implicit caching hints are collected for future read I/O redirection. The read I/O redirection is performed from within the virtual block device in the virtualized system, to manipulate the entire host-cache as a content-deduplicated cache implicitly. Our trace-based evaluation using a custom simulator, reveals that DRIVE always performs equal to or better than the Vanilla system, achieving up to 20% better cache-hit ratios and reducing the number of disk reads by up to 80%. The results also indicate that our system is able to achieve up to 97% content deduplication in the host-cache.
Sujesha Sudevalayam, Purushottam Kulkarni
HiPC2
2014 Comparative Analysis of Page Cache Provisioning in Virtualized Environments
abstract
Efficient management of system memory plays a critical role in provisioning virtual machines, as it impacts levels of over-commitment and associated application performance. Typically, file accesses from a virtual machine traverse through different levels of page caches, which consume memory. Different configurations of page cache provisioning are possible, each providing different levels of memory utilization and performance levels. In this work, we study different page cache provisioning options with the KVM (Kernel Virtual Machine) virtual machine monitor solution. Our goal is to systematically understand possible provisioning use cases to compare their cost-benefit tradeoffs. Towards this we implement and evaluate tmem, an exclusive caching model (based on the transcendent memory model) for file blocks. Together with the tmem-caching model and existing page cache provisioning options, we present an empirical analysis of all cases. Our evaluation focuses on identifying actual caching needs, overheads and benefits for different combinations and identifies the relative benefits of each. We find that there is up-to 10x increase in disk read throughput with tmem-based caching and the CPU overheads for this technique are proportional to the gain in throughput.
Debadatta Mishra, Purushottam Kulkarni
MASCOTS2
2013 Share-o-meter: An empirical analysis of KSM based memory sharing in virtualized systems
abstract
Content based memory sharing in virtualized environments has proven to be a useful technique for over-commitment based placement of virtual machines. Kernel-based Virtual Machine (KVM) on Linux uses Kernel SamePage Merging (KSM) to identify and exploit sharing opportunities. In this paper, we present an analysis of page sharing across virtual machines by comparing page sharing achieved by KSM to total sharing opportunities presented by virtual machines. We study the impact of different KSM configurations, system resources, and workload characteristics on page sharing achieved by KSM. We also study the cost of sharing in terms of CPU utilization overhead from Copy-On-Write page breaks that occur on KSM shared pages. Our analysis is aimed at exploring the KSM configuration space towards obtaining desired sharing levels with minimal overheads for a given amount of system resources and workload characteristics. Our empirical analysis shows that for workloads exhibiting different memory usage patterns, different KSM configuration parameters are required to achieve maximum savings. We quantify the levels of savings and associated costs for several (individual and combinations) of workloads, exhibiting different sharing opportunities and memory usage characteristics. Further, we demonstrate the need for adaptive configuration of KSM's aggressiveness based on changes in total memory available for sharing and change in memory usage characteristics.
Shashank Rachamalla, Debadatta Mishra, Purushottam Kulkarni
HiPC3
2013 Resource availability based performance benchmarking of virtual machine migrations
abstract
Virtual machine migration enables load balancing, hot spot mitigation and server consolidation in virtualized environments. Live VM migration can be of two types - adaptive, in which the rate of page transfer adapts to virtual machine behaviour (mainly page dirty rate), and non-adaptive, in which the VM pages are transferred at a maximum possible network rate. In either method, migration requires a significant amount of CPU and network resources, which can seriously impact the performance of both the VM being migrated as well as other VMs. This calls for building a good understanding of the performance of migration itself and the resource needs of migration. Such an understanding can help select the appropriate VMs for migration while at the same time allocating the appropriate amount of resources for migration. While several empirical studies exist, a comprehensive evaluation of migration techniques with resource availability constraints is missing. As a result, it is not clear as to which migration technique to employ under a given set of conditions. In this work, we conduct a comprehensive empirical study to understand the sensitivity of migration performance to resource availability and other system parameters (like page dirty rate and VM size). The empirical study (with the Xen Hypervisor) reveals several shortcomings of the migration process. We propose several fixes and develop the Improved Live Migration technique (ILM) to overcome these shortcomings. Over a set of workloads used to evaluate ILM, the network traffic for migration was reduced by 14-93% and the migration time was reduced by 34-87% compared to the vanilla live migration technique. We also quantified the impact of migration on the performance of applications running on the migrating VM and other co-located VMs.
Senthil Nathan, Purushottam Kulkarni, Umesh Bellur
ICPE2
2013 Affinity-aware modeling of CPU usage with communicating virtual machines
Sujesha Sudevalayam, Purushottam Kulkarni
J. Syst. Softw.2
2012 Risk Aware Provisioning and Resource Aggregation Based Consolidation of Virtual Machines
abstract
Server consolidation has emerged as an important technique to save on energy costs in virtualized datacenters. The issue of instantiation of a given set of Virtual Machines (VMs) on a set of Physical Machines (PMs) can be thought of as consisting of a provisioning step where we determine the amount of resources to be allocated to a VM and a placement step which decides which VMs can be placed together on a physical machines thereby allocating VMs to PMs. In this paper, we introduce a provisioning scheme which takes into account acceptable intensity of violation of provisioned resources. In addition we identify a serious shortcoming of existing placement schemes that correct in our correlation aware placement scheme. We consider correlation among aggregated resource demands of VMs while finding the VM-PM mapping. Experimental results reveal that our approach leads to a significant amount of reduction in the number of servers (up to 32% in our settings) required to host 1000 VMs and thus enables us to turn off unnecessary servers. It achieves this by packing VMs more tightly by correlating resource requirements across the entire set of VMs to be placed. We present a comprehensive set of experimental results comparing our scheme with the existing provisioning and placement schemes.
Kishaloy Halder, Umesh Bellur, Purushottam Kulkarni
IEEE CLOUD3
2012 Singleton: system-wide page deduplication in virtual environments
abstract
We investigate memory-management in hypervisors and propose Singleton, a KVM-based system-wide page deduplication solution to increase memory usage efficiency. We address the problem of double-caching that occurs in KVM---the same disk blocks are cached at both the host(hypervisor) and the guest(VM) page caches. Singleton's main components are identical-page sharing across guest virtual machines and an implementation of an exclusive-cache for the host and guest page cache hierarchy. We use and improve KSM--Kernel SamePage Merging to identify and share pages across guest virtual machines. We utilize guest memory-snapshots to scrub the host page cache and maintain a single copy of a page across the host and the guests. Singleton operates on a completely black-box assumption---we do not modify the guest or assume anything about its behaviour. We show that conventional operating system cache management techniques are sub-optimal for virtual environments, and how Singleton supplements and improves the existing Linux kernel memory-management mechanisms. Singleton is able to improve the utilization of the host cache by reducing its size(by upto an order of magnitude), and increasing the cache-hit ratio(by factor of 2x). This translates into better VM performance(40% faster I/O). Singleton's unified page deduplication and host cache scrubbing is able to reclaim large amounts of memory and facilitates higher levels of memory overcommitment. The optimizations to page deduplication we have implemented keep the overhead down to less than 20% CPU utilization.
Prateek Sharma 0001, Purushottam Kulkarni
HPDC2
2012 A survey of sensory data boundary estimation, covering and tracking techniques using collaborating sensors
Sumana Srinivasan, Subhasri Duttagupta, Purushottam Kulkarni, Krithi Ramamritham
Pervasive Mob. Comput.3
2011 VirtPerf: A Performance Profiling Tool for Virtualized Environments
abstract
Several applications in the "physical'' world are being consolidated in "virtual'' environments using different virtualization technologies. An important criteria for this exercise is to understand potential resource requirements and performance levels achieved in virtual environments. Empirical evidence of these can be gotten by benchmarking the application's performance in a controlled manner in virtual environments. These measurements can be used for a variety of purposes from virtual machine capacity planning to building sophisticated performance models to predict performance for loads that cannot be practically tested. In this paper, we present VirtPerf, an integrated workload generator and measurement tool to capture resource utilization levels and performance metrics of applications executing under controlled circumstances in virtualized environments. The tool aims to provide comprehensive measurement-based analysis for applications in different virtualization settings. Additionally, a configurable workload generator can be used to stress and profile applications under different load conditions. We present the detailed design of VirtPerf and a comprehensive empirical study to demonstrate its correctness and capabilities.
Prajakta Patil, Purushottam Kulkarni, Umesh Bellur
IEEE CLOUD2
2011 Affinity-Aware Modeling of CPU Usage for Provisioning Virtualized Applications
abstract
While virtualization-based systems become a reality, an important issue is that of virtual machine migration-enabled consolidation and dynamic resource provisioning. Mutually communicating virtual machines, as part of migration and consolidation strategies, may get colocated on the same physical machine or placed on different machines. In this work, we argue the need for network affinity-awareness not only in placement but also in resource provisioning for virtual machines. First, we empirically quantify the resource savings due to colocation of communicating virtual machines. We also discuss the increase in resource usage due to dispersion of previously colocated virtual machines. Next, we build models based on different resource-usage micro-benchmarks to predict the resource usages when transitioning from non-colocated placements to colocated placements and vice-versa. These resource usage prediction models are usable along-with consolidation and migration procedures to determine requirements of VMs in colocated and non colocated scenarios. Via extensive experimentation, we evaluate the applicability of our models for synthetic and benchmark application workloads. We find that the models have high prediction accuracy - 90th percentile prediction error within 3% absolute CPU usage for both synthetic and application workloads.
Sujesha Sudevalayam, Purushottam Kulkarni
IEEE CLOUD2
2011 Tracking Dynamic Boundaries Using Sensor Network
abstract
We examine the problem of tracking dynamic boundaries occurring in natural phenomena using a network of range sensors. Two main challenges of the boundary tracking problem are accurate boundary estimation from noisy observations and continuous tracking of the boundary. We propose Dynamic Boundary Tracking (DBTR), an algorithm that combines the spatial estimation and temporal estimation techniques. The regression-based spatial estimation technique determines discrete points on the boundary and estimates a confidence band around the entire boundary. In addition, a Kalman Filter-based temporal estimation technique tracks changes in the boundary and aperiodically updates the spatial estimate to meet accuracy requirements. DBTR provides a low energy solution compared to similar periodic update techniques to track boundaries without requiring prior knowledge about the dynamics. Experimental results demonstrate the effectiveness of our algorithm; estimated confidence bands indicate a loss of coverage of less than 2 to 5 percent for a variety of boundaries with different spatial characteristics.
Subhasri Duttagupta, Krithi Ramamritham, Purushottam Kulkarni
IEEE Trans. Parallel Distributed Syst.3
2010 LokVaani: demonstrating interactive voice in Lo3
abstract
In this work, we consider the goal of enabling effective voice communication in a TDMA, multi-hop mesh network, using low cost and low power platforms. We consider two primary usage scenarios: (1) enabling a local voice communication within a village-like setting, in developing regions (2) supporting an on-site local communication among a team of users e.g. during emergency response systems.
Vijay Gabale, Bhaskaran Raman, Kameswari Chebrolu, Purushottam Kulkarni
SIGCOMM4
2010 Vehicular wifi access and rate adaptation
abstract
Vehicular WiFi access is distinct in two respects, (i) continuous mobility of clients and (ii) possibility of predictable link quality. As part of this study, we aim to comprehensively evaluate existing rate adaptation algorithms in real environments. Further, if required, we aim to develop a simple, low-overhead rate adaptation algorithm suited for vehicular WiFi access.
Ajinkya Uday Joshi, Purushottam Kulkarni
SIGCOMM2
2010 Road traffic estimation using in-situ acoustic sensing
abstract
In this paper, we explore the efficacy of curb-side acoustic sensing to estimate road traffic conditions. We formulated a set of hypotheses which attempted to correlate traffic conditions with the ambient traffic noise. We present the evaluation of our hypotheses under various traffic conditions. Our threshold-based-classification yields 70-90% accuracy in distinguishing congested from free-flowing traffic.
C. Viven Rajendra, Purushottam Kulkarni
SIGCOMM2
2008 Tracking Dynamic Boundary Fronts Using Range Sensors
Subhasri Duttagupta, Krithi Ramamritham, Purushottam Kulkarni, Kannan M. Moudgalya
EWSN3
2008 ACE in the Hole: Adaptive Contour Estimation Using Collaborating Mobile Sensors
abstract
This paper focuses on the use of mobile sensors to estimate contours in a field. In particular, we focus on strategies to estimate the contour with minimum latency and maximum precision. We propose a novel algorithm, ACE (adaptive contour estimation), that (a) estimates and exploits information regarding the gradients in the field to move towards the contour and (b) uses a spread component to surround the contour in order to optimize latency. While it is possible for sensors to spread as they approach the contour, it is crucial to judiciously determine when and how much to spread. Spreading too early or too much may result in increasing the latency or affecting the precision. ACE dynamically makes this decision using local sensor measurements, history of measurements as well as collaboration between sensors while adapting to different types of deployment, distance from the contour and shapes of the contour. We demonstrate that ACE, in the absence of energy constraints precisely determines the contour with a lower latency than when only gradients are used for movement or when the sensors spread out right from the start of estimation. Additionally, we show that ACE significantly improves precision of contour estimation in the presence of energy constraints. We also demonstrate a proof of concept implementation on a mobile robot testbed.
Sumana Srinivasan, Krithi Ramamritham, Purushottam Kulkarni
IPSN3
2007 Approximate Initialization of Camera Sensor Networks
Purushottam Kulkarni, Prashant J. Shenoy, Deepak Ganesan
EWSN1
2006 Snapshot: A Self-Calibration Protocol for Camera Sensor Networks
abstract
A camera sensor network is a wireless network of cameras designed for ad-hoc deployment. The camera sensors in such a network need to be properly calibrated by determining their location, orientation, and range. This paper presents Snapshot, an automated calibration protocol that is explicitly designed and optimized for camera sensor networks. Snapshot uses the inherent imaging abilities of the cameras themselves for calibration and can determine the location and orientation of a camera sensor using only four reference points. Our techniques draw upon principles from computer vision, optics, and geometry and are designed to work with low-fidelity, low-power camera sensors that are typical in sensor networks. An experimental evaluation of our prototype implementation shows that Snapshot yields an error of 1-2.5 degrees when determining the camera orientation and 5-10cm when determining the camera location. We show that this is a tolerable error in practice since a Snapshot-calibrated sensor network can track moving objects to within 11cm of their actual locations. Finally, our measurements indicate that Snapshot can calibrate a camera sensor within 20 seconds, enabling it to calibrate a sensor network containing tens of cameras within minutes.
Purushottam Kulkarni, Prashant J. Shenoy, Deepak Ganesan
BROADNETS2
2005 SensEye: a multi-tier camera sensor network
abstract
This paper argues that a camera sensor network containing heterogeneous elements provides numerous benefits over traditional homogeneous sensor networks. We present the design and implementation of senseye---a multi-tier network of heterogeneous wireless nodes and cameras. To demonstrate its benefits, we implement a surveillance application using senseye comprising three tasks: object detection, recognition and tracking. We propose novel mechanisms for low-power low-latency detection, low-latency wakeups, efficient recognition and tracking. Our techniques show that a multi-tier sensor network can reconcile the traditionally conflicting systems goals of latency and energy-efficiency. An experimental evaluation of our prototype shows that, when compared to a single-tier prototype, our multi-tier senseye can achieve an order of magnitude reduction in energy usage while providing comparable surveillance accuracy.
Purushottam Kulkarni, Deepak Ganesan, Prashant J. Shenoy, Qifeng Lu
ACM Multimedia1
2005 The case for multi-tier camera sensor networks
abstract
In this position paper, we examine recent technology trends that have resulted in a broad spectrum of camera sensors, wireless radio technologies, and embedded sensor platforms with varying capabilities. We argue that future sensor applications will be hierarchical with multiple tiers, where each tier employs sensors with different characteristics. We argue that multi-tier networks are not only scalable, they offer a number of advantages over simpler, single-tier unimodal networks: lower cost, better coverage, higher functionality, and better reliability. However, the design of such mixed networks raises a number of new challenges that are not adequately addressed by current research. We discuss several of these challenges and illustrate how they can be addressed in the context of SensEye, a multi-tier video surveillance application that we are designing in our research group.
Purushottam Kulkarni, Deepak Ganesan, Prashant J. Shenoy
NOSSDAV1
2004 Redundancy Elimination Within Large Collections of Files
Purushottam Kulkarni, Fred Douglis, Jason D. LaVoie, John M. Tracey
USENIX ATC, General Track1
2003 Handling Client Mobility and Intermittent Connectivity in Mobile Web Accesses
Purushottam Kulkarni, Prashant J. Shenoy, Krithi Ramamritham
Mobile Data Management1
2003 Scalable techniques for memory-efficient CDN simulations
abstract
Since CDN simulations are known to be highly memory-intensive, in this paper, we argue the need for reducing the memory requirements of such simulations. We propose a novel memory-efficient data structure that stores cache state for a small subset of popular objects accurately and uses approximations for storing the state for the remaining objects. Since popular objects receive a large fraction of the requests while less frequently accessed objects consume much of the memory space, this approach yields large memory savings and reduces errors. We use bloom filters to store approximate state and show that careful choice of parameters can substantially reduce the probability of errors due to approximations. We implement our techniques into a user library for constructing proxy caches in CDN simulators. Our experimental results show up to an order of magnitude reduction in memory requirements of CDN simulations, while incurring a 5-10% error.
Purushottam Kulkarni, Prashant J. Shenoy, Weibo Gong
WWW1
2003 Scalable Consistency Maintenance in Content Distribution Networks Using Cooperative Leases
abstract
We argue that cache consistency mechanisms designed for stand-alone proxies do not scale to the large number of proxies in a content distribution network and are not flexible enough to allow consistency guarantees to be tailored to object needs. To meet the twin challenges of scalability and flexibility, we introduce the notion of cooperative consistency along with a mechanism, called cooperative leases, to achieve it. By supporting /spl Delta/-consistency semantics and by using a single lease for multiple proxies, cooperative leases allow the notion of leases to be applied in a flexible, scalable manner to CDNs. Further, the approach employs application-level multicast to propagate server notifications to proxies in a scalable manner. We implement our approach in the Apache Web server and the Squid proxy cache and demonstrate its efficacy using a detailed experimental evaluation. Our results show a factor of 2.5 reduction in server message overhead and a 20 percent reduction in server state space overhead when compared to original leases albeit at an increased interproxy communication overhead.
Anoop George Ninan, Purushottam Kulkarni, Prashant J. Shenoy, Krithi Ramamritham, Renu Tewari
IEEE Trans. Knowl. Data Eng.2
2002 Cooperative leases: scalable consistency maintenance in content distribution networks
abstract
In this paper, we argue that cache consistency mechanisms designed for stand-alone proxies do not scale to the large number of proxies in a content distribution network and are not flexible enough to allow consistency guarantees to be tailored to object needs. To meet the twin challenges of scalability and flexibility, we introduce the notion of cooperative consistency along with a mechanism, called cooperative leases, to achieve it. By supporting Δ-consistency semantics and by using a single lease for multiple proxies, cooperative leases allows the notion of leases to be applied in a flexible, scalable manner to CDNs. Further, the approach employs application-level multicast to propagate server notifications to proxies in a scalable manner. We implement our approach in the Apache web server and the Squid proxy cache and demonstrate its efficacy using a detailed experimental evaluation. Our results show a factor of 2.5 reduction in server message overhead and a 20% reduction in server state space overhead when compared to original leases albeit at an increased inter-proxy communication overhead.
Anoop George Ninan, Purushottam Kulkarni, Prashant J. Shenoy, Krithi Ramamritham, Renu Tewari
WWW2
2002 Implications of proxy caching for provisioning networks and servers
abstract
In this paper, we examine the potential benefits of Web proxy caches in improving the effective capacity of servers and networks. Since networks and servers are typically provisioned based on a high percentile of the load, we focus on the effects of proxy caching on the tail of the load distribution. We find that, unlike their substantial impact on the average load, proxies have a diminished impact on the tail of the load distribution. The exact reduction in the tail and the corresponding capacity savings depend on the nature of the workload and the percentile of the load distribution chosen for provisioning networks and servers-the higher the percentile, the smaller the savings. For workloads considered in this study, compared with over a 50% reduction in the average load, the savings in network and server capacity was only 20%-35% for the 99th percentile of the load distribution. We also find that while proxies can be somewhat useful in smoothing out some of the burstiness in Web workloads; the resulting workload continues, however, to exhibit substantial burstiness and a heavy-tailed nature. We identify one-time requests for large objects to be the limiting factor that diminishes the impact of proxies on the tail of load distribution. We conclude that, while proxies are immensely useful to users due to the reduction in the average response time, they are less effective in improving the capacities of networks and servers.
M. S. Raunak 0001, Prashant J. Shenoy, Pawan Goyal 0001, Krithi Ramamritham, Purushottam Kulkarni
IEEE J. Sel. Areas Commun.5