VLDB 2026 Research / reviewers in the wild / expert
Vivek M. Bhasi
dblp:304/9913
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-6326-0665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FLEXI: Phase-Aware Function Resizing for Heterogeneous Serverless GPU Workloads
Shruti Mohanty, Vivek M. Bhasi, Jashwant Raj Gunasekaran, Prashanth Thinakaran, Mahmut T. Kandemir, Chita R. Das |
IEEE Big Data | 2 |
| 2025 | Dally: A Network-Placement Sensitive Cluster Scheduler for Deep Learning
Aakash Sharma, Vivek M. Bhasi, Sonali Singh, Mahmut T. Kandemir, George Kesidis, Chita R. Das |
IEEE Big Data | 2 |
| 2025 | GSCoder: Enabling Fast and Efficient Encoding for Game Streaming ApplicationsabstractThe recent proliferation of cloud gaming (also referred to as game streaming), enabling high-fidelity gaming quality on edge devices without high-end hardware, promises a transformative and democratized gaming experience across diverse populations. Yet, streaming high definition (4K UHD) game frames to thin-client devices, especially mobile, requires substantially higher bandwidth than traditional video streaming and often results in frame drops, degrading user experience. We identified that this high bandwidth demand arises from inefficient compression of game frames with real-time compression requirement (60 frames per second (FPS)) because the rapid, irregular motion in game frames violates the predictability assumptions baked into standard video motion estimation algorithms.To address this issue, we propose and evaluate GSCoder, a realtime efficient game frame compression framework that utilizes motion cues from the game’s rendering pipeline to directly acquire and optimize encoding-compliant motion vectors precisely, bypassing the costly motion estimation step inherent in standard encoders. Our evaluation, conducted in five open-source games with varying motion complexities, demonstrates that GSCoder achieves on average 49% and 19% higher compression efficiency than state-of-the-art (SOTA) game and video encoders, respectively, while maintaining real-time performance and delivering high-quality streams. Moreover, GSCoder delivers at least $3.5 \times$ encoding speedup compared to SOTA video encoders. Sandeepa Bhuyan, Ziyu Ying 0001, Vivek M. Bhasi, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 3 |
| 2024 | FAAStloop: Optimizing Loop-Based Applications for Serverless ComputingabstractServerless Computing has garnered significant interest for executing High-Performance Computing (HPC) applications in recent years, attracting attention for its elastic scalability, reduced entry barriers, and pay-per-use pricing model. Specifically, highly parallel HPC apps can be divided and offloaded to multiple Serverless Functions (SFs) that execute their respective tasks concurrently and, finally, their results are stored/aggregated. While state-of-the-art userside serverless frameworks have attempted to fine-tune task division amongst the SFs to optimize for performance and/or cost, they have either used static task division parameters or have only focused on minimizing the number of SFs through task packing. However, these methods treat the HPC code as a black-box and usually require significant manual intervention to find the optimal task division. Since a significant portion of the HPC applications have a loop structure, in this work, we try to answer the following two questions: (i) Can modifying the loop structure in the HPC code, originally optimized for monolithic (non-serverless) frameworks, enhance performance and reduce costs in a serverless architecture?, and (ii) Can we develop a framework that allows for an efficient transition of monolithic code to serverless, with minimum user input? Shruti Mohanty, Vivek M. Bhasi, Myungjun Son, Mahmut T. Kandemir, Chita R. Das |
SoCC | 2 |
| 2024 | Paldia: Enabling SLO-Compliant and Cost-Effective Serverless Computing on Heterogeneous HardwareabstractAmong the variety of applications (apps) being deployed on serverless platforms, apps such as Machine Learning (ML) inference serving can achieve better performance from leveraging accelerators like GPUs. Yet, major serverless providers, despite having GPU-equipped servers, do not offer GPU support for their serverless functions. Given that serverless functions are deployed on various generations of CPUs already, extending this to various (typically more expensive) GPU generations can offer providers a greater range of hardware to serve incoming requests according to the functions and request traffic. Here, providers are faced with the challenge of selecting hardware to reach a well-proportioned trade-off point between cost and performance. While recent works have attempted to address this, they often fail to do so as they overlook optimization opportunities arising from intelligently leveraging existing GPU sharing mechanisms. To address this point, we devise a heterogeneous serverless framework, PALDIA, which uses a prudent Hardware selection policy to acquire capable, cost-effective hardware and perform intelligent request scheduling on it to yield high performance and cost savings. Specifically, our scheduling algorithm employs hybrid spatio-temporal GPU sharing that intelligently trades off job queueing delays and interference to allow the chosen cost-effective hardware to also be highly performant. We extensively evaluate PALDIA using 16 ML inference workloads with real-world traces on a 6 node heterogeneous cluster. Our results show that PALDIA significantly outperforms state-of-the-art works in terms of Service Level Objective (SLO) compliance (up to 13.3% more) and tail latency (up to ∼50% less), with cost savings up to 86%. Vivek M. Bhasi, Aakash Sharma, Shruti Mohanty, Mahmut T. Kandemir, Chita R. Das |
IPDPS | 1 |
| 2024 | Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsabstractPersonalized recommendation is a ubiquitous appli-cation on the internet, with many industries and hyperscalers extensively leveraging Deep Learning Recommendation Models (DLRMs) for their personalization needs (like ad serving or movie suggestions). With growing model and dataset sizes pushing computation and memory requirements, GPUs are being increasingly preferred for executing DLRM inference. However, serving newer DLRMs, while meeting acceptable latencies, continues to remain challenging, making traditional deployments increasingly more GPU-hungry, resulting in higher inference serving costs. In this paper, we show that the embedding stage continues to be the primary bottleneck in the GPU inference pipeline, leading up to a 3.2 x embedding-only performance slowdown. To thoroughly grasp the problem, we conduct a detailed microarchitecture characterization and highlight the presence of low occupancy in the standard embedding kernels. By leveraging direct compiler optimizations, we achieve optimal occupancy, pushing the performance by up to 53 %. Yet, long memory latency stalls continue to exist. To tackle this challenge, we propose spe-cialized plug-and-play-based software prefetching and L2 pinning techniques, which help in hiding and decreasing the latencies. Further, we propose combining them, as they complement each other. Experimental evaluations using AI00 GPUs with large models and datasets show that our proposed techniques improve performance by up to 103% for the embedding stage, and up to 77 % for the overall D LRM inference pipeline. Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
MICRO | 2 |
| 2024 | Towards SLO-Compliant and Cost-Effective Serverless Computing on Emerging GPU ArchitecturesabstractServerless platforms are supporting an increasing variety of applications (apps). Among these, apps such as Machine Learning (ML) inference serving can benefit significantly from leveraging accelerators like GPUs. Yet, major serverless providers, despite having GPU-equipped servers, do not offer GPU support for their serverless functions. While recent works have attempted to bridge this gap, they are agnostic to the capabilities of new-generation GPUs, thereby, overlooking several performance optimization opportunities. Vivek M. Bhasi, Aakash Sharma, Jashwant Raj Gunasekaran, Ashutosh Pattnaik, Mahmut T. Kandemir, Chita R. Das |
Middleware | 1 |
| 2023 | Stash: A Comprehensive Stall-Centric Characterization of Public Cloud VMs for Distributed Deep LearningabstractDeep neural networks (DNNs) are increasingly popular owing to their ability to solve complex problems such as image recognition, autonomous driving, and natural language processing. Their growing complexity coupled with the use of larger volumes of training data (to achieve acceptable accuracy) has warranted the use of GPUs and other accelerators. Such accelerators are typically expensive, with users having to pay a high upfront cost to acquire them. For infrequent use, users can, instead, leverage the public cloud to mitigate the high acquisition cost. However, with the wide diversity of hardware instances (particularly GPU instances) available in public cloud, it becomes challenging for a user to make an appropriate choice from a cost/performance standpoint. In this work, we try to address this problem by (i) introducing a comprehensive distributed deep learning (DDL) profiler Stash, which determines the various execution stalls that DDL suffers from, and (ii) using Stash to extensively characterize various public cloud GPU instances by running popular DNN models on them. Specifically, it estimates two types of communication stalls, namely, interconnect and network stalls, that play a dominant role in DDL execution time. Stash is implemented on top of prior work, DS-analyzer, that computes only the CPU and disk stalls. Using our detailed stall characterization, we list the advantages and shortcomings of public cloud GPU instances for users to help them make an informed decision(s). Our characterization results indicate that the more expensive GPU instances may not be the most performant for all DNN models and that AWS can sometimes sub-optimally allocate hardware interconnect resources. Specifically, the intra-machine interconnect can introduce communication overheads of up to 90% of DNN training time and the network-connected instances can suffer from up to 5× slowdown compared to training on a single instance. Furthermore, (iii) we also model the impact of DNN macroscopic features such as the number of layers and the number of gradients on communication stalls, and finally, (iv) we briefly discuss a cost comparison with existing work. Aakash Sharma, Vivek M. Bhasi, Sonali Singh, Jashwant Raj Gunasekaran, Subrata Mitra, Mahmut T. Kandemir, George Kesidis, Chita R. Das |
ICDCS | 2 |
| 2022 | Cypress: input size-sensitive container provisioning and request scheduling for serverless platformsabstractThe growing popularity of the serverless platform has seen an increase in the number and variety of applications (apps) being deployed on it. The majority of these apps process user-provided input to produce the desired results. Existing work in the area of input-sensitive profiling has empirically shown that many such apps have input size-dependent execution times which can be determined through modelling techniques. Nevertheless, existing serverless resource management frameworks are agnostic to the input size-sensitive nature of these apps. We demonstrate in this paper that this can potentially lead to container over-provisioning and/or end-to-end Service Level Objective (SLO) violations. To address this, we propose Cypress, an input size-sensitive resource management framework, that minimizes the containers provisioned for apps, while ensuring a high degree of SLO compliance. We perform an extensive evaluation of Cypress on top of a Kubernetes-managed cluster using 5 apps from the AWS Serverless Application Repository and/or Open-FaaS Function Store with real-world traces and varied input size distributions. Our experimental results show that Cypress spawns up to 66% fewer containers, thereby, improving container utilization and saving cluster-wide energy by up to 2.95X and 23%, respectively, versus state-of-the-art frameworks, while remaining highly SLO-compliant (up to 99.99%). Vivek M. Bhasi, Jashwant Raj Gunasekaran, Aakash Sharma, Mahmut T. Kandemir, Chita R. Das |
SoCC | 1 |
| 2021 | Kraken: Adaptive Container Provisioning for Deploying Dynamic DAGs in Serverless PlatformsabstractThe growing popularity of microservices has led to the proliferation of online cloud service-based applications, which are typically modelled as Directed Acyclic Graphs (DAGs) comprising of tens to hundreds of microservices. The vast majority of these applications are user-facing, and hence, have stringent SLO requirements. Serverless functions, having short resource provisioning times and instant scalability, are suitable candidates for developing such latency-critical applications. However, existing serverless providers are unaware of the workflow characteristics of application DAGs, leading to container over-provisioning in many cases. This is further exacerbated in the case of dynamic DAGs, where the function chain for an application is not known a priori. Motivated by these observations, we propose Kraken, a workflow-aware resource management framework that minimizes the number of containers provisioned for an application DAG while ensuring SLO-compliance. We design and implement Kraken on OpenFaaS and evaluate it on a multi-node Kubernetes-managed cluster. Our extensive experimental evaluation using DeathStarbench workload suite and real-world traces demonstrates that Kraken spawns up to 76% fewer containers, thereby improving container utilization and saving cluster-wide energy by up to 4x and 48%, respectively, when compared to state-of-the art schedulers employed in serverless platforms. Vivek M. Bhasi, Jashwant Raj Gunasekaran, Prashanth Thinakaran, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das |
SoCC | 1 |