EDBT 2026 Demo / reviewers in the wild / expert
Umesh Deshpande
dblp:83/1394
· DBLP profile ↗
20ranked-venue papers
13as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 11 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Symbiosis: Multi-Adapter Inference and Fine-TuningabstractParameter-efficient fine-tuning (PEFT) allows model builders to capture the task-specific parameters into adapters, which are a fraction of the size of the original base model. Popularity of PEFT technique for fine-tuning has led to the creation of a large number of adapters for popular Large Language Models (LLMs). However, existing frameworks fall short in supporting inference or fine-tuning with multiple adapters in the following ways. 1) For fine-tuning, each job needs to deploy its dedicated base model instance, which results in excessive GPU memory consumption and poor GPU utilization. 2) While popular inference platforms can serve multiple PEFT adapters, they do not allow independent resource management or mixing of different PEFT methods. 3) They cannot make effective use of heterogeneous accelerators. 4) They do not provide privacy to users who may not wish to expose their fine-tuned parameters to service providers. In Symbiosis, we address the above problems by enabling the as-a-service deployment of the base model. The base model layers can be shared across multiple inference or fine-tuning processes. Our split-execution technique decouples the execution of client-specific adapters and layers from the frozen base model layers offering them flexibility to manage their resources, to select their fine-tuning method, to achieve their performance goals. Our approach is transparent to models and works out-of-the-box for most models in the transformers library. We demonstrate the use of Symbiosis to simultaneously fine-tune 20 Gemma2-27B LoRA adapters on 8 GPUs. Saransh Gupta, Umesh Deshpande, Travis Janssen, Swaminathan Sundararaman |
SoCC | 2 |
| 2024 | MoEsaic: Shared Mixture of ExpertsabstractMixture of Expert (MoE) models consist of several experts, each specializing in a specific task. During inference, a subset of the experts is invoked based on their relevance to the request. MoE's modular architecture lets users compose their model from popular off-the-shelf experts. This leads to multiple MoE deployments with identical experts. The duplication of experts across model instances results in excessive GPU memory consumption and increased model serving cost. Moreover, since all experts are not invoked for each request, individual experts rarely receive enough requests to exploit the GPUs' computational capabilities, resulting in low GPU utilization. To address these problems, we propose Shared Mixture of Experts in MoEsaic. MoEsaic automatically identifies and deduplicates identical experts across model instances, thus reducing their memory footprint. Moreover, it batches the requests directed toward the identical experts belonging to different clients, which also improves the processing efficiency. We show that for Mixtral-8x7B model, when compared to deploying dedicated MoE instances, MoEsaic can serve 7X more model instances with little impact on inference performance. Umesh Deshpande, Travis Janssen, Mudhakar Srivatsa, Swaminathan Sundararaman |
SoCC | 1 |
| 2024 | Atlas: Hybrid Cloud Migration Advisor for Interactive MicroservicesabstractHybrid cloud provides an attractive solution to microservices for better resource elasticity. A subset of application components can be offloaded from the on-premises cluster to the cloud, where they can readily access additional resources. However, the selection of this subset is challenging because of the large number of possible combinations. A poor choice degrades the application performance, disrupts the critical services, and increases the cost to the extent of making the use of hybrid cloud unviable. This paper presents Atlas, a hybrid cloud migration advisor. Atlas uses a data-driven approach to learn how each user-facing API utilizes different components and their network footprints to drive the migration decision. It learns to accelerate the discovery of high-quality migration plans from millions and offers recommendations with customizable trade-offs among three quality indicators: end-to-end latency of user-facing APIs representing application performance, service availability, and cloud hosting costs. Atlas continuously monitors the application even after the migration for proactive recommendations. Our evaluation shows that Atlas can achieve 21% better API performance (latency) and 11% cheaper cost with less service disruption than widely used solutions. Ka-Ho Chow 0001, Umesh Deshpande, Veera Deenadhayalan, Sangeetha Seshadri, Ling Liu 0001 |
EuroSys | 2 |
| 2023 | Two stream multi-layer convolutional network for keyframe-based video summarization
Khushboo Khurana, Umesh Deshpande |
Multim. Tools Appl. | 2 |
| 2022 | Storage Capacity Prediction using Population AnalyticsabstractCapacity forecasting is a common feature in popular storage systems. The storage system vendors use prediction methods based on the historical usage for forecasting future utilization. However, we observed that such methods do not perform well under situations when the storage usage may not follow the historical trend. To address the problem, we propose a prediction method based on the data observed over 18,000 storage system. An analysis of a large scale dataset allows us to identify the popular trends across the storage systems at various stages of their lifecycle and further improve the prediction. Our evaluation shows that using the population data reduces the prediction error by 45% over linear regression. Umesh Deshpande, Thanh Pham |
IEEE Big Data | 1 |
| 2022 | DeepRest: deep resource estimation for interactive microservicesabstractInteractive microservices expose API endpoints to be invoked by users. For such applications, precisely estimating the resources required to serve specific API traffic is challenging. This is because an API request can interact with different components and consume different resources for each component. The notion of API traffic is vital to application owners since the API endpoints often reflect business logic, e.g., a customer transaction. The existing systems that simply rely on historical resource utilization are not API-aware and thus cannot estimate the resource requirement accurately. This paper presents DeepRest, a deep learning-driven resource estimation system. DeepRest formulates resource estimation as a function of API traffic and learns the causality between user interactions and resource utilization directly in a production environment. Our evaluation shows that DeepRest can estimate resource requirements with over 90% accuracy, even if the API traffic to be estimated has never been observed (e.g., 3× more users than ever or unseen traffic shape). We further apply resource estimation for application sanity checks. DeepRest identifies system anomalies by verifying whether the resource utilization is justifiable by how the application is being used. It can successfully identify two major cyber threats: ransomware and cryptojacking attacks. Ka-Ho Chow 0001, Umesh Deshpande, Sangeetha Seshadri, Ling Liu 0001 |
EuroSys | 2 |
| 2021 | Self-service data protection for stateful containersabstractData protection in containerized environments poses several challenges arising from the need for self-service and the resulting churn in the environment. We present a self-managing backup system for containerized frameworks designed to work for users with little knowledge of the underlying infrastructure. Our system presents users with an interface which allows them to interact with data protection service in the same way used for managing their applications. Additionally, we present a backup scheduler to honor user expressed data protection guarantees while reacting to resource fluctuations on the underlying shared infrastructure. We demonstrate the effectiveness of our system with thousands of request having different data protection guarantees in an environment with various bandwidth and IO patterns. Umesh Deshpande, Nick Linck, Sangeetha Seshadri |
HotStorage | 1 |
| 2021 | SRA: Smart Recovery Advisor for Cyber AttacksabstractContinuous Data Protection (CDP) is becoming instrumental in recovering applications from crypto-ransomware attacks. It enables fine-grained recovery through journaling, allowing the applications (its volumes) to recover to any previous state. While zero data loss can be achieved during recovery with CDP, the timestamp of the desired restore point, i.e., the one just prior to the attack, needs to be provided to reconstruct the volume. Such information is often unavailable in practice, and system administrators can only adopt a trial-and-error strategy to narrow down the time range of desired restore points by making multiple time-consuming recovery attempts. The recovery systems offer little guidance in pointing to the restore points containing a valid application state and reducing data loss. To address this problem, we equip the CDP-based recovery with machine intelligence. This demonstration showcases Smart Recovery Advisor (SRA), which offers interpretable, data-driven, and feedback-aware restore point recommendations that reduce the number of recovery attempts while minimizing data loss. Ka-Ho Chow 0001, Umesh Deshpande, Sangeetha Seshadri, Ling Liu 0001 |
SIGMOD Conference | 2 |
| 2021 | Self managed data protection for containersabstractContainer frameworks have been gaining popularity in recent years, with container native storage being one of the fastest growing segment. According to IDC report [1], 90% of applications on cloud platforms and over 95% of new microservices are being deployed in containers. The growth of container native storage is largely driven by stateful applications [2, 3], the mainstay of enterprise IT environments. As organizations are increasingly adopting containerized deployments, they must also deal with data protection to maintain business continuity. Umesh Deshpande, Nick Linck, Sangeetha Seshadri |
SYSTOR | 1 |
| 2019 | Caravel: Burst Tolerant Scheduling for Containerized Stateful ApplicationsabstractIn a containerized environment, the applications are generally categorized as either stateful or stateless, each consisting of multiple containers. Their co-scheduling in a cluster presents unique challenges for the container orchestration frameworks. The two types of applications differ from each other in how they handle the temporary load spikes. The stateless applications can scale out by instantly spawning new identical instances, whereas the stateful applications require deliberate planning to scale, as each application instance is unique. Instead, the stateful applications can more conveniently acquire the additional resources needed during a spike by scaling up on the same node. However, when an application's container uses more than requested resources, it risks being evicted from the node. The evictions are particularly detrimental for stateful applications because of their longer start up time and the resulting degradation. Moreover, the existing container orchestration frameworks schedule or evict containers without any knowledge of its impact on their owning applications. For instance, an eviction of the application's multiple containers in a short period of time could compromise its availability and severely degrade its performance. To address these challenges, we present Caravel, a scheduling approach that provides better experience to stateful applications in dealing with load spikes. It allows them to overstep the resource request during a burst and use the resources on the same node while minimizing their evictions. Moreover, the scheduler provides a fair opportunity to all the stateful applications to use the spare resources in the cluster. The evaluation shows that our approach reduces the eviction of stateful applications by up to 90% over the traditional approach. Umesh Deshpande |
ICDCS | 1 |
| 2018 | Scatter-Gather Live Migration of Virtual MachinesabstractWe introduce a new metric for live migration of virtual machines (VM) called eviction time defined as the time to evict the state of one or more VMs from the source host. Eviction time determines how quickly the source can be taken offline or its resources repurposed for other VMs. In traditional live migration, such as pre-copy and post-copy, eviction time equals the total migration time because the source is tied up until the destination receives the entire VM. We present Scatter-Gather live migration which decouples the source and destination during migration to reduce eviction time when the destination is slow. The source scatters the memory of VMs to multiple nodes, including the destination and one or more intermediaries. Concurrently, the destination gathers the VMs' memory from the intermediaries and the source. Thus eviction from the source is no longer bottlenecked by the reception speed of the destination. We support simultaneous live eviction of multiple VMs and exploit deduplication to reduce network overhead. Our Scatter-Gather implementation in the KVM/QEMU platform reduces the eviction time by up to a factor of 6 against traditional pre-copy and post-copy while maintaining comparable total migration time when the destination is slower than the source. Umesh Deshpande, Danny Chan, Kartik Gopalan, Nilton Bila |
IEEE Trans. Cloud Comput. | 1 |
| 2017 | Traffic-sensitive Live Migration of Virtual Machines
Umesh Deshpande, Kate Keahey |
Future Gener. Comput. Syst. | 1 |
| 2016 | Agile Live Migration of Virtual MachinesabstractA key attraction of virtual machines (VMs) is live migration - the ability to move their execution state across physical machines even as the VMs continue to run. Unfortunately, the traditional pre-copy and post-copy techniques are not agile in the face of resource pressures at the source host, since it takes a long time to transfer the memory state of a VM. Consequently, the performance suffers for all VMs - those being migrated as well as those being left behind. Prior works have attempted to optimize indirect measures of migration effectiveness such as downtime, total migration time, and network overhead. However, none have treated the performance of VMs impacted by migration as the primary metric of migration effectiveness. We propose an Agile live migration technique that quickly recovers the performance of all VMs under resource pressure by eliminating resource pressure faster than traditional live migration. The working set of a VM is typically much smaller than its full memory footprint. Our approach works by transparently tracking the working set of each VM and offloading the non-working set (cold pages) in advance to portable per-VM swap devices. We present a new hybrid pre/post-copy technique that reduces the performance impact on the VM's workload by transferring only the working set of the VM while enabling destination to remotely access cold pages from the per-VM swap device. We describe the challenges in the design and implementation of Agile live migration in the KVM/QEMU platform without modifying the guest OS in the VM. When live migrating under memory pressure, we demonstrate a reduction in the performance impact on VMs by a up to factor of 2, reduction in migration time by up to factor of 4 besides reduction in memory pressure on both the source and destination hosts. Umesh Deshpande, Danny Chan, Ten-Young Guh, James Edouard, Kartik Gopalan, Nilton Bila |
IPDPS | 1 |
| 2016 | Enabling Efficient Hypervisor-as-a-Service Clouds with Eemeral VirtualizationabstractWhen considering a hypervisor, cloud providers must balance conflicting requirements for simple, secure code bases with more complex, feature-filled offerings. This paper introduces Dichotomy, a new two-layer cloud architecture in which the roles of the hypervisor are split. The cloud provider runs a lean hyperplexor that has the sole task of multiplexing hardware and running more substantial hypervisors (called featurevisors) that implement features. Cloud users choose featurevisors from a selection of lightly-modified hypervisors potentially offered by third-parties in an "as-a-service" model for each VM. Rather than running the featurevisor directly on the hyperplexor using nested virtualization, Dichotomy uses a new virtualization technique called eemeral virtualization which efficiently (and repeatedly) transfers control of a VM between the hyperplexor and featurevisor using memory mapping techniques. Nesting overhead is only incurred when the VM is accessed by the featurevisor. We have implemented Dichotomy in KVM/QEMU and demonstrate average switching times of 80 ms, two to three orders of magnitude faster than live VM migration. We show that, for the featurevisor applications we evaluated, VMs hosted in Dichotomy deliver up to 12% better performance than those hosted on nested hypervisors, and continue to show benefit even when the featurevisor applications run as often as every 2.5~seconds. Dan Williams 0001, Yaohui Hu, Umesh Deshpande, Piush K. Sinha, Nilton Bila, Kartik Gopalan, Hani Jamjoom |
VEE | 3 |
| 2015 | Performance Analysis of Encryption in Securing the Live Migration of Virtual MachinesabstractVirtual machine (VM) migration is a technique for transferring the execution state of a VM from one physical host to another. While VM migration is critical for load balancing, consolidation, and server maintenance in virtualized data centers, it can also increase security risks. During VM migration, an attacker with sufficient privileges can compromise a VM by modifying its memory contents during transit to subvert its applications or the guest operating system. One could maintain dedicated, and presumably more secure, control networks to carry the migration traffic, but at significant hardware and administrative complexity. Alternatively, one could encrypt the migration traffic, which eliminates the need for dedicated control networks, but might introduce performance overheads. To date, there has been no systematic study of how encryption affects VM migration, especially in high-bandwidth low-delay networks that are common within data centers. In this paper, we present a study of the impact of AES and 3DES encryption algorithms on two widely used live VM migration approaches - pre-copy and post-copy. Our key findings are as follows. The encryption algorithm used can have a significant impact on the total migration time. The impact of encryption on downtime varies with the type of the migration technique. The overhead of encryption also depends upon the relative speeds of source and target machines. Finally, an application's performance within a VM during encrypted migration varies with the type of the application and the migration mechanism. Yaohui Hu, Sanket Panhale, Tianlin Li, Emine Kaynar, Danny Chan, Umesh Deshpande, Ping Yang 0002, Kartik Gopalan |
CLOUD | 6 |
| 2015 | Traffic-Sensitive Live Migration of Virtual MachinesabstractIn this paper we address the problem of network contention between the migration traffic and the VM application traffic for the live migration of co-located Virtual Machines (VMs). When VMs are migrated with pre-copy, they run at the source host during the migration. Therefore the VM applications with predominantly outbound traffic contend with the outgoing migration traffic at the source host. Similarly, during post-copy migration, the VMs run at the destination host. Therefore the VM applications with predominantly inbound traffic contend with the incoming migration traffic at the destination host. Such a contention increases the total migration time of the VMs and degrades the performance of VM application. Here, we propose traffic-sensitive live VM migration technique to reduce the contention of migration traffic with the VM application traffic. It uses a combination of pre-copy and post-copy techniques for the migration of the co-located VMs, instead of relying upon any single pre-determined technique for the migration of all the VMs. We base the selection of migration techniques on VMs' network traffic profiles so that the direction of migration traffic complements the direction of the most VM application traffic. We have implemented a prototype of traffic-sensitive migration on the KVM/QEMU platform. In the evaluation, we compare traffic-sensitive migration against the approaches that use only pre-copy or only post-copy for VM migration. We show that our approach minimizes the network contention for migration, thus reducing the total migration time and the application degradation. Umesh Deshpande, Kate Keahey |
CCGRID | 1 |
| 2014 | Fast Server Deprovisioning through Scatter-Gather Live Migration of Virtual MachinesabstractTraditional metrics for live migration of virtual machines (VM) include total migration time, downtime, network overhead, and application degradation. In this paper, we introduce a new metric, "eviction time", defined as the time to evict the entire state of a VM from the source host. Eviction time determines how quickly the source host can be taken offline, or the freed resources re-purposed for other VMs. In traditional approaches for live VM migration, such as pre-copy and post-copy, eviction time is equal to the total migration time, because the source and destination hosts are coupled for the duration of the migration. Eviction time increases if the destination host is slow to receive the incoming VM, such as due to insufficient memory or network bandwidth, thus tying up the source host. We present a new approach, called "Scatter-Gather" live migration, which reduces the eviction time when the destination host is resource constrained. The key idea is to decouple the source and the destination hosts. The source scatters the VM's memory state quickly to multiple intermediaries (hosts or middleboxes) in the cluster. Concurrently, the destination gathers the VM's memory from the intermediaries using a variant of post-copy VM migration. We have implemented a prototype of Scatter-Gather in the KVM/QEMU platform. In our evaluations, Scatter-Gather reduces the VM eviction time by up to a factor of 6 while maintaining comparable total migration time against traditional pre-copy and post-copy for a resource constrained destination. Umesh Deshpande, Danny Chan, Nilton Bila, Kartik Gopalan |
IEEE CLOUD | 1 |
| 2013 | Gang Migration of Virtual Machines Using Cluster-wide DeduplicationabstractGang migration refers to the simultaneous live migration of multiple Virtual Machines (VMs) from one set of physical machines to another in response to events such as load spikes and imminent failures. Gang migration generates a large volume of network traffic and can overload the core network links and switches in a data center. In this paper, we present an approach to reduce the network overhead of gang migration using global deduplication (GMGD). GMGD identifies and eliminates the retransmission of duplicate memory pages among VMs running on multiple physical machines in the cluster. We describe the design, implementation and evaluation of a GMGD prototype using QEMU/KVM VMs. Evaluations on a 30-node Gigabit Ethernet cluster having 10GigE core links shows that GMGD can reduce the network traffic on core links by up to 65% and the total migration time of VMs by up to 42% when compared to the default migration technique in QEMU/KVM. Furthermore, GMGD has a smaller adverse performance impact on network-bound applications. Umesh Deshpande, Brandon Schlinker, Eitan Adler, Kartik Gopalan |
CCGRID | 1 |
| 2011 | Live gang migration of virtual machinesabstractThis paper addresses the problem of simultaneously migrating a group of co-located and live virtual machines (VMs), i.e, VMs executing on the same physical machine. We refer to such a mass simultaneous migration of active VMs as live gang migration. Cluster administrators may often need to perform live gang migration for load balancing, system maintenance, or power savings. Application performance requirements may dictate that the total migration time, network traffic overhead, and service downtime, be kept minimal when migrating multiple VMs. State-of-the-art live migration techniques optimize the migration of a single VM. In this paper, we optimize the simultaneous live migration of multiple co-located VMs. We present the design, implementation, and evaluation of a de-duplication based approach to perform concurrent live migration of co-located VMs. Our approach transmits memory content that is identical across VMs only once during migration to significantly reduce both the total migration time and network traffic. Using the QEMU/KVM platform, we detail a proof-of-concept prototype implementation of two types of de-duplication strategies (at page level and sub-page level) and a differential compression approach to exploit content similarity across VMs. Evaluations over Gigabit Ethernet with various types of VM workloads demonstrate that our prototype for live gang migration can achieve significant reductions in both network traffic and total migration time. Categories andSubjectDescriptors Umesh Deshpande, Xiaoshuang Wang, Kartik Gopalan |
HPDC | 1 |
| 2010 | MemX: Virtualization of Cluster-Wide MemoryabstractWe present MemX -- a distributed system that virtualizes cluster-wide memory to support data-intensive and large memory workloads in virtual machines (VMs). MemX provides a number of benefits in virtualized settings: (1) VM workloads that access large datasets can perform low-latency I/O over virtualized cluster-wide memory; (2) VMs can transparently execute very large memory applications that require more memory than physical DRAM present in the host machine; (3) MemX reduces the effective memory usage of the cluster by de-duplicating pages that have identical content; (4) existing applications do not require any modifications to benefit from MemX such as the use of special APIs, libraries, recompilation, or relinking; and (5) MemX supports live migration of large-footprint VMs by eliminating the need to migrate part of their memory footprint resident on other nodes. Detailed evaluations of our MemX prototype show that large dataset applications and multiple concurrent VMs achieve significant performance improvements using MemX compared against virtualized local and iSCSI disks. Umesh Deshpande, Beilan Wang, Shafee Haque, Michael R. Hines, Kartik Gopalan |
ICPP | 1 |