VLDB 2026 Research / reviewers in the wild / expert
Prateek Sharma 0001
dblp:88/8903-1
· DBLP profile ↗
23ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MQGPU: A Multi-Queue Scheduling Framework For GPU Accelerated Serverless FunctionsabstractHardware accelerators like GPUs are now ubiquitous in data centers, but are not fully supported by common cloud abstractions such as Functions as a Service (FaaS). Many popular and emerging FaaS applications such as machine learning and scientific computing can benefit from GPU acceleration. However, FaaS frameworks (such as OpenWhisk) are not capable of providing this acceleration because of the impedance mismatch between GPUs and the FaaS programming model, which requires virtualization and sandboxing of each function. The challenges are amplified due to the highly dynamic and heterogeneous FaaS workloads. Alexander Fuerst, Siddharth Anil, Prateek Sharma 0001 |
ICPE | 3 |
| 2025 | OmniLearn: A Framework for Distributed Deep Learning Over Heterogeneous ClustersabstractDeep learning systems are optimized for clusters with homogeneous resources. However, heterogeneity is prevalent in computing infrastructure across edge, cloud and HPC. When training neural networks using stochastic gradient descent techniques on heterogeneous resources, performance degrades due to stragglers and stale updates. In this work, we develop an adaptive batch-scaling framework calledOmniLearnto mitigate the effects of heterogeneity in distributed training. Our approach is inspired by proportional controllers to balance computation across heterogeneous servers, and works under varying resource availability. By dynamically adjusting worker mini-batches at runtime,OmniLearnreduces training time by 14-85%. We also investigate asynchronous training, where our techniques improve accuracy by up to 6.9%. Sahil Tyagi, Prateek Sharma 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Accountable Carbon Footprints and Energy Profiling For Serverless FunctionsabstractCloud computing is a significant and growing cause of carbon emissions. Understanding the energy consumption and carbon footprints of cloud applications is a fundamental prerequisite to raising awareness, designing sustainability metrics, and creating targeted system optimizations. In this paper, we address the challenges of providing accurate and full-system (not just CPU) carbon footprints for serverless (FaaS) functions. To the best of our knowledge, this is the first work which develops an energy and carbon metrology framework for FaaS. Prateek Sharma 0001, Alexander Fuerst |
SoCC | 1 |
| 2023 | Scavenger: A Cloud Service For Optimizing Cost and Performance of ML TrainingabstractCloud computing platforms can provide the compu-tational resources required for training large machine learning models such as deep neural networks. While the pay-as-you- go nature of cloud virtual machines (VMs) makes it easy to spin-up large clusters for training models, it can also lead to ballooning costs. The 100s of virtual machine sizes provided by cloud platforms also makes it extremely challenging to select the “right” cloud cluster configuration for training. Furthermore, the training time and cost of distributed model training is highly sensitive to the cluster configurations, and presents a large and complex tradeoff-space. In this paper, we develop principled and practical techniques for optimizing the training time and cost of distributed ML model training on the cloud. Our key insight is that both the parallel and statistical efficiency must be considered when selecting the optimum job configuration parameters such as the number of workers and the batch size. By combining conventional parallel scaling concepts and new insights into SGD noise, we develop models for estimating the time and cost on different cluster configurations. Using the repetitive nature of training and our performance models, our Scavenger cloud service can search for optimum cloud configurations in a black-box, online manner. Our approach reduces training times by 2 x and costs by more than 50 %. Our performance models are accurate to within 2 %, and our search imposes only a 10% overhead compared to an ideal oracle- based approach. Sahil Tyagi, Prateek Sharma 0001 |
CCGrid | 2 |
| 2023 | Ilúvatar: A Fast Control Plane for Serverless ComputingabstractProviding efficient Functions as a Service (FaaS) is challenging due to the serverless programming model and highly heterogeneous and dynamic workloads. Great strides have been made in optimizing FaaS performance through scheduling, caching, virtualization, and other resource management techniques. The combination of these advances and growing FaaS workloads have pushed the performance bottleneck into the control plane itself. Current FaaS control planes like OpenWhisk introduce 100s of milliseconds of latency overhead, and are becoming unsuitable for high performance FaaS research and deployments. Alexander Fuerst, Abdul Rehman 0005, Prateek Sharma 0001 |
HPDC | 3 |
| 2022 | Memory-harvesting VMs in cloud platformsabstractloud platforms monetize their spare capacity by renting “Spot” virtual machines (VMs) that can be evicted in favor of higher-priority VMs. Recent work has shown that resource-harvesting VMs are more effective at exploiting spare capacity than Spot VMs, while also reducing the number of evictions. However, the prior work focused on harvesting CPU cores while keeping memory size fixed. This wastes a substantial monetization opportunity and may even limit the ability of harvesting VMs to leverage spare cores. Thus, in this paper, we explore memory harvesting and its challenges in real cloud platforms, namely its impact on VM creation time, NUMA spanning, and page fragmentation. We start by characterizing the amount and dynamics of the spare memory in Azure. We then design and implement memory-harvesting VMs (MHVMs), introducing new techniques for memory buffering, batching, and pre-reclamation. To demonstrate the use of MHVMs, we also extend a popular cluster scheduling framework (Hadoop) and a FaaS platform to adapt to them. Our main results show that (1) there is plenty of scope for memory harvesting in real platforms; (2) MHVMs are effective at mitigating the negative impacts of harvesting; and (3) our extensions of Hadoop and FaaS successfully hide the MHVMs’ varying memory size from the users’ data-processing jobs and functions. We conclude that memory harvesting has great potential for practical deployment and users can save up to 93% of their costs when running workloads on MHVMs. Alexander Fuerst, Stanko Novakovic, Íñigo Goiri, Gohar Irfan Chaudhry, Prateek Sharma 0001, Kapil Arya, Kevin Broas, Eugene Bak, Mehmet Iyigun, Ricardo Bianchini |
ASPLOS | 5 |
| 2022 | Locality-aware Load-Balancing For Serverless ClustersabstractWhile serverless computing provides more convenient abstractions for developing and deploying applications, the Function-as-a-Service (FaaS) programming model presents new resource management challenges for the FaaS provider. In this paper, we investigate load-balancing policies for serverless clusters. Locality, i.e., running repeated invocations of a function on the same server, is a key determinant of performance because it increases warm-starts and reduces cold-start overheads. We find that the locality vs. load tradeoff is crucial and presents a large design space. Alexander Fuerst, Prateek Sharma 0001 |
HPDC | 2 |
| 2022 | SciSpot: Scientific Computing On Temporally Constrained Cloud Preemptible VMsabstractScientific computing applications are being increasingly deployed on cloud computing platforms. Transient servers such as EC2 spot instances and Google Preemptible VMs, can be used to lower the costs of running applications on the cloud by up to$10\times$. However, the frequent preemptions and resource heterogeneity of these transient servers introduces many challenges in their effective and efficient use. In this paper, we develop techniques for modeling and mitigating preemptions of transient servers, and present SciSpot, a software framework that enables low-cost scientific computing on the cloud. SciSpot deploys applications on Google Cloud Preemptible Virtual Machines that exhibit temporally constrained preemptions: VMs are always preempted in a 24 hour interval. Our empirical analysis shows that the preemption rate is generally bathtub shaped, which raises multiple fundamental challenges in performance modeling and policy design. We develop a new reliability model for temporally constrained preemptions, and use statistical mechanics to show why the bathtub shape is generally exhibited. SciSpot’s design is guided by our observation that many emerging scientific computing applications that integrate machine learning with simulations, can be deployed as “bags” of jobs, which represent multiple instantiations of the same computation with different physical model parameters. For a bag of jobs, SciSpot finds the optimal transient server on-the-fly, by taking into account the price, performance, and preemption rates of different servers. SciSpot reduces costs by$5\times$compared to conventional cloud deployments, and reduces makespans by up to$10\times$compared to conventional high performance computing clusters. J. C. S. Kadupitiya, Vikram Jadhao, Prateek Sharma 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Molecular Dynamics Simulations on Cloud Computing and Machine Learning PlatformsabstractScientific computing applications have benefited greatly from high performance computing infrastructure such as supercomputers. However, we are seeing a paradigm shift in the computational structure, design, and requirements of these applications. Increasingly, data-driven and machine learning approaches are being used to support, speed-up, and enhance scientific computing applications, especially molecular dynamics simulations. Concurrently, cloud computing platforms are increasingly appealing for scientific computing, providing “infi-nite” computing powers, easier programming and deployment models, and access to computing accelerators such as TPUs (Tensor Processing Units). This confluence of machine learning (ML) and cloud computing represents exciting opportunities for cloud and systems researchers. ML-assisted molecular dynamics simulations are a new class of workload, and exhibit unique computational patterns. These simulations present new challenges for low-cost and high-performance execution. We argue that transient cloud resources, such as low-cost preemptible cloud VMs, can be a viable platform for this new workload. Finally, we present some low-hanging fruits and long-term challenges in cloud resource management, and the integration of molecular dynamics simulations into ML platforms (such as TensorFlow). Prateek Sharma 0001, Vikram Jadhao |
CLOUD | 1 |
| 2021 | FaasCache: keeping serverless computing alive with greedy-dual cachingabstractFunctions as a Service (also called serverless computing) promises to revolutionize how applications use cloud resources. However, functions suffer from cold-start problems due to the overhead of initializing their code and data dependencies before they can start executing. Keeping functions alive and warm after they have finished execution can alleviate the cold-start overhead. Keep-alive policies must keep functions alive based on their resource and usage characteristics, which is challenging due to the diversity in FaaS workloads. Alexander Fuerst, Prateek Sharma 0001 |
ASPLOS | 2 |
| 2020 | Cloud-scale VM-deflation for Running Interactive Applications On Transient ServersabstractTransient computing has become popular in public cloud environments for running delay-insensitive batch and data processing applications at low cost. Since transient cloud servers can be revoked at any time by the cloud provider, they are considered unsuitable for running interactive application such as web services. In this paper, we present VM deflation as an alternative mechanism to server preemption for reclaiming resources from transient cloud servers under resource pressure. Using real traces from top-tier cloud providers, we show the feasibility of using VM deflation as a resource reclamation mechanism for interactive applications in public clouds. We show how current hypervisor mechanisms can be used to implement VM deflation and present cluster deflation policies for resource management of transient and on-demand cloud VMs. Experimental evaluation of our deflation system on a Linux cluster shows that microservice-based applications can be deflated by up to 50% with negligible performance overhead. Our cluster-level deflation policies allow overcommitment levels as high as 50%, with less than a 1% decrease in application throughput, and can enable cloud platforms to increase revenue by 30%. Alexander Fuerst, Ahmed Ali-Eldin, Prashant J. Shenoy, Prateek Sharma 0001 |
HPDC | 4 |
| 2020 | Modeling The Temporally Constrained Preemptions of Transient Cloud VMsabstractTransient cloud servers such as Amazon Spot instances, Google Preemptible VMs, and Azure Low-priority batch VMs, can reduce cloud computing costs by as much as 10x, but can be unilaterally preempted by the cloud provider. Understanding preemption characteristics (such as frequency) is a key first step in minimizing the effect of preemptions on application performance, availability, and cost. However, little is understood about temporally constrained preemptions---wherein preemptions must occur in a given time window. We study temporally constrained preemptions by conducting a large scale empirical study of Google's Preemptible VMs (that have a maximum lifetime of 24 hours), develop a new preemption probability model, new model-driven resource management policies, and implement them in a batch computing service for scientific computing workloads. Our statistical and experimental analysis indicates that temporally constrained preemptions are not uniformly distributed but are time-dependent and have a bathtub shape. We find that existing memoryless models and policies are not suitable for temporally constrained preemptions. We develop a new probability model for bathtub preemptions and analyze it through the lens of reliability theory. To highlight the effectiveness of our model, we develop optimized policies for job scheduling and checkpointing. Compared to existing techniques, our model-based policies can reduce the probability of job failure by more than 2x. We also implement our policies as part of a batch computing service for scientific computing applications, which reduces cost by 5x compared to conventional cloud deployments and keeps performance overheads under 3%. J. C. S. Kadupitige, Vikram Jadhao, Prateek Sharma 0001 |
HPDC | 3 |
| 2019 | Resource Deflation: A New Approach For Transient Resource ReclamationabstractData centers and clouds are increasingly offering low-cost computational resources in the form of transient virtual machines. Whenever demand for computational resources exceeds their availability, transient resources can reclaimed by preempting the transient VMs. Conventionally, these transient VMs are used by low-priority applications that can tolerate the disruption caused by preemptions. Prateek Sharma 0001, Ahmed Ali-Eldin, Prashant J. Shenoy |
EuroSys | 1 |
| 2019 | SpotWeb: Running Latency-sensitive Distributed Web Services on Transient Cloud ServersabstractMany cloud providers offer servers with transient availability at a reduced cost. These servers can be unilaterally revoked by the provider, usually after a warning period to the user. Until recently, it has been thought that these servers are not suitable to run latency-sensitive workloads due to their transient availability. In this paper, we introduce SpotWeb, a framework for running latency-sensitive web workloads on transient computing platforms while maintaining the Quality-of-Service (QoS) of the running applications. SpotWeb is based on three novel concepts; using multi-period optimization---a novel approach developed in finance---for server selection; transiency-aware load-balancing; and using intelligent capacity over-provisioning. We implement SpotWeb and evaluate its performance in both simulations and testbed experiments. Our results show that SpotWeb reduces costs by up to 50% compared to state-of-the-art solutions while being scalable to hundreds of cloud server configurations. Ahmed Ali-Eldin, Jonathan Westin, Prateek Sharma 0001, Prashant J. Shenoy |
HPDC | 4 |
| 2019 | The Price Is (Not) Right: Reflections on Pricing for Transient Cloud ServersabstractAmazon introduced spot instances in December 2009, enabling "customers to bid on unused Amazon EC2 capacity and run those instances for as long as their bid exceeds the current Spot Price.'' Amazon's real-time computational spot market was novel in multiple respects. For example, it was the first (and to date only) large-scale public implementation of market-based resource allocation based on dynamic pricing after decades of research, and it provided users with useful information, control knobs, and options for optimizing the cost of running cloud applications. Spot instances also introduced the concept of transient cloud servers derived from variable idle capacity that cloud platforms could revoke at any time. Transient servers have since become central to efficient resource management of modern clusters and clouds. As a result, Amazon's spot market was the motivation for substantial research over the past decade. Yet, in November 2017, Amazon effectively ended its realtime spot market by announcing that users no longer needed to place bids and that spot prices will "...adjust more gradually, based on longer-term trends in supply and demand.'' The changes made spot instances more similar to the fixed-price transient servers offered by other cloud platforms. Unfortunately, while these changes made spot instances less complex, they eliminated many benefits to sophisticated users in optimizing their applications. This paper provides a retrospective on Amazon's real-time spot market, including its advantages and disadvantages for allocating transient servers compared to current fixed-price approaches. We also discuss some fundamental problems with Amazon's spot market, which we identified in prior work (from 2016), that predicted its eventual end. We then discuss potential options for allocating transient servers that combine the advantages of Amazon's real-time spot market, while also addressing the problems that likely led to its elimination. David Irwin 0001, Prashant J. Shenoy, Lurdh Pradeep Reddy Ambati, Prateek Sharma 0001, Supreeth Shastri, Ahmed Ali-Eldin |
ICCCN | 4 |
| 2019 | Performance Evaluation of Multi-Path TCP for Data Center and Cloud WorkloadsabstractToday's cloud data centers host a wide range of applications including data analytics, batch processing, and interactive processing. These applications require high throughput, low latency, and high reliability from the network. Satisfying these requirements in the face of dynamically varying network conditions remains a challenging problem. Multi-Path TCP (MPTCP) is a recently proposed IETF extension to TCP that divides a conventional TCP flow into multiple subflows so as to utilize multiple paths over the network. Despite the theoretical and practical benefits of MPTCP, its effectiveness for cloud applications and environments remains unclear as there has been little work to quantify the benefits of MPTCP for real cloud applications. We present a broad empirical study of the effectiveness and feasibility of MPTCP for data center and cloud applications, under different network conditions. Our results show that while MPTCP provides useful bandwidth aggregation, congestion avoidance, and improved resiliency for some cloud applications, these benefits do not apply uniformly across applications, especially in cloud settings. Lucas Chaufournier, Ahmed Ali-Eldin, Prateek Sharma 0001, Prashant J. Shenoy, Don Towsley |
ICPE | 3 |
| 2018 | Managing Risk in a Derivative IaaS CloudabstractInfrastructure-as-a-Service (IaaS) cloud platforms rent computing resources with different cost and availability tradeoffs. For example, users may acquire virtual machines (VMs) in the spot market-that are cheap, but can be unilaterally terminated by the cloud operator. Because of this revocation risk, spot servers have been conventionally used for delay and risk tolerant batch jobs. In this paper, we develop risk mitigation policies which allow even interactive applications to run on spot servers. Our System, SpotCheck is a derivative cloud platform, and provides the illusion of an IaaS platform that offers always-available VMs on demand for a cost near that of spot servers, and supports unmodified applications. SpotCheck's design combines virtualization-based mechanisms for fault-tolerance, and bidding and server selection policies for managing the risk and cost. We implement SpotCheck on EC2 and show that it i) provides nested VMs with 99.9989 percent availability, ii) achieves upto 2-5x cost savings compared to using on-demand VMs, and iii) eliminates any risk of losing VM state. Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | The Financialization of Cloud Computing: Opportunities and ChallengesabstractUnder competitive pressure to maximize their infrastructure's utilization and revenue, modern cloud platforms are quickly evolving into server markets that offer increasingly sophisticated contracts beyond simple on-demand servers, such as spot, preemptible, burstable, and reserved servers. In parallel, continuing advances in system and network virtualization are making server-time a more fungible commodity. These trends have motivated calls for open cloud commodity markets akin to other commodity markets, e.g., for oil, gold, corn, etc. However, such open cloud markets have not yet materialized due to key differences between cloud resources and other commodities. In particular, the relationship between applications and their underlying server resources is fundamentally different and more complex than other commodities. Unfortunately, software developers generally do not have the necessary background to effectively manage this complexity as part of their applications. Financial cloud computing is an emerging area that focuses on adapting and extending concepts from economics and finance to explicitly manage applications' tradeoffs between cost, risk, availability, and performance in cloud markets. A key goal of financial cloud computing is to develop systems-level abstractions and mechanisms that manage the market's complexity. This paper introduces this emerging area and its potential benefits, surveys related work, discusses challenges to realizing a cloud commodity market, and then outlines future research directions. David Irwin 0001, Prateek Sharma 0001, Supreeth Shastri, Prashant J. Shenoy |
ICCCN | 2 |
| 2016 | Flint: batch-interactive data-intensive processing on transient serversabstractCloud providers now offer transient servers, which they may revoke at anytime, for significantly lower prices than on-demand servers, which they cannot revoke. The low price of transient servers is particularly attractive for executing an emerging class of workload, which we call Batch-Interactive Data-Intensive (BIDI), that is becoming increasingly important for data analytics. BIDI workloads require large sets of servers to cache massive datasets in memory to enable low latency operation. In this paper, we illustrate the challenges of executing BIDI workloads on transient servers, where revocations (akin to failures) are the common case. To address these challenges, we design Flint, which is based on Spark and includes automated checkpointing and server selection policies that i) support batch and interactive applications and ii) dynamically adapt to application characteristics. We evaluate a prototype of Flint using EC2 spot instances, and show that it yields cost savings of up to 90% compared to using on-demand servers, while increasing running time by < 2%. Prateek Sharma 0001, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 1 |
| 2016 | Containers and Virtual Machines at Scale: A Comparative Study
Prateek Sharma 0001, Lucas Chaufournier, Prashant J. Shenoy, Y. C. Tay |
Middleware | 1 |
| 2015 | SpotOn: a batch computing service for the spot marketabstractCloud spot markets enable users to bid for compute resources, such that the cloud platform may revoke them if the market price rises too high. Due to their increased risk, revocable resources in the spot market are often significantly cheaper (by as much as 10×) than the equivalent non-revocable on-demand resources. One way to mitigate spot market risk is to use various fault-tolerance mechanisms, such as checkpointing or replication, to limit the work lost on revocation. However, the additional performance overhead and cost for a particular fault-tolerance mechanism is a complex function of both an application's resource usage and the magnitude and volatility of spot market prices. Supreeth Subramanya, Tian Guo 0001, Prateek Sharma 0001, David Irwin 0001, Prashant J. Shenoy |
SoCC | 3 |
| 2015 | SpotCheck: designing a derivative IaaS cloud on the spot marketabstractInfrastructure-as-a-Service (IaaS) cloud platforms rent resources, in the form of virtual machines (VMs), under a variety of contract terms that offer different levels of risk and cost. For example, users may acquire VMs in the spot market that are often cheap but entail significant risk, since their price varies over time based on market supply and demand and they may terminate at any time if the price rises too high. Currently, users must manage all the risks associated with using spot servers. As a result, conventional wisdom holds that spot servers are only appropriate for delay-tolerant batch applications. In this paper, we propose a derivative cloud platform, called SpotCheck, that transparently manages the risks associated with using spot servers for users. Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 1 |
| 2012 | Singleton: system-wide page deduplication in virtual environmentsabstractWe investigate memory-management in hypervisors and propose Singleton, a KVM-based system-wide page deduplication solution to increase memory usage efficiency. We address the problem of double-caching that occurs in KVM---the same disk blocks are cached at both the host(hypervisor) and the guest(VM) page caches. Singleton's main components are identical-page sharing across guest virtual machines and an implementation of an exclusive-cache for the host and guest page cache hierarchy. We use and improve KSM--Kernel SamePage Merging to identify and share pages across guest virtual machines. We utilize guest memory-snapshots to scrub the host page cache and maintain a single copy of a page across the host and the guests. Singleton operates on a completely black-box assumption---we do not modify the guest or assume anything about its behaviour. We show that conventional operating system cache management techniques are sub-optimal for virtual environments, and how Singleton supplements and improves the existing Linux kernel memory-management mechanisms. Singleton is able to improve the utilization of the host cache by reducing its size(by upto an order of magnitude), and increasing the cache-hit ratio(by factor of 2x). This translates into better VM performance(40% faster I/O). Singleton's unified page deduplication and host cache scrubbing is able to reclaim large amounts of memory and facilitates higher levels of memory overcommitment. The optimizations to page deduplication we have implemented keep the overhead down to less than 20% CPU utilization. Prateek Sharma 0001, Purushottam Kulkarni |
HPDC | 1 |