EDBT 2026 Demo / reviewers in the wild / expert
Anshuj Garg
dblp:144/4850
· DBLP profile ↗
8ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0002-1579-9631ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Hummingbird: Leveraging Heterogeneous System Architecture for deploying dynamic NFV chainsabstractNetwork Function Virtualization has gained traction as a network function deployment alternative due to its flexibility and cost benefits. The telecommunication (telecom) operators and infrastructure providers are looking for high throughput, low latency NFV deployment model to avail the benefits of NFV. Moreover, NFV is one of the core technology for the next-generation communication network such as 5G. Furthermore, telecom operators employ groups of network functions(NFs) that process packets in linear order so that the output of one NF becomes an input for another, thus forming the network function chain (NFC). However, these NFCs should be flexible, as all telecom packets do not necessarily need to be processed by the same set of NFs. It has been earlier shown that GPU increases the throughput of NFV chains. To the best of our knowledge, none of the GPU-based frameworks supports dynamic NFV chains. Furthermore, discrete GPUs are expensive and consume a fair amount of energy. This paper presents the design and evaluation of Hummingbird, a framework to support high throughput, dynamically routed NFV chain on Heterogeneous System Architecture (HSA). Though HSAs are affordable and power-efficient, they lack high throughput GPU-CPU synchronization. Furthermore, current technology does not provide a zero-copy mechanism for network IO between GPU and NIC for HSAs. Hummingbird addressed those challenges. As per our knowledge, this is the first such framework that provides high throughput dynamic NFV chaining, with NFs chained across GPU and CPU and designed in conformance to OpenCL 2.0 standard. Hummingbird achieves 6x throughput per-core and 3.5x throughput per unit of energy consumption compared to state-of-the-art NFV deployment framework G-net, which uses powerful and costly discrete GPU. Avinash Kumar Chaurasia, Bhaskaran Raman, Omkar Prabhu, Shashank P, Anshuj Garg |
CCGRID | 6 |
| 2022 | Simmer: Rate proportional scheduling to reduce packet drops in vGPU based NF chainsabstractNetwork Function Virtualization (NFV) paradigm offers flexibility, cost benefits, and ease of deployment by decoupling network function from hardware middleboxes. The service function chains (SFC) deployed using the NFV platform require efficient sharing of resources among various network functions in the chain. Graphics Processing Units (GPUs) have been used to improve various network functions’ performance. However, sharing a single GPU among multiple virtualized network functions (virtual machines) in a service function chain has been challenging due to their proprietary hardware and software stack. Earlier GPU architectures had a limitation: a single physical GPU can only be allocated to one virtual machine (VM) and cannot be shared among multiple VMs. The newer GPUs are virtualization-aware (hardware-assisted virtualization) and allow multiple virtual machines to share a single physical GPU. Although virtualization-aware, these GPUs still lack support for custom scheduling policies and do not expose the preemption control to users. When network functions (hosted within virtual machines) with different processing requirements share the same GPU, virtualization-aware GPUs’ default round-robin scheduling mechanism proves to be inefficient, resulting in packet drops and lower throughput. This paper presents Simmer, an efficient mechanism for scheduling a network function service chain on virtualization-aware GPUs. Our scheduling solution considers the processing requirement of NFs in a GPU-based SFC, thus improving overall throughput by up to 29% and reducing the packet drop to zero compared to vanilla setup. Avinash Kumar Chaurasia, Anshuj Garg, Bhaskaran Raman, Uday Kurkure, Hari Sivaraman, Lan Vu, Sairam Veeraswamy |
ICPP | 2 |
| 2021 | FaaSter: Accelerated Functions-as-a-Service with Heterogeneous GPUsabstractIn this work, we present FaaSter, an Accelerated Functions as a Service (AFaaS) offering that unifies the function-as-a-service model with GPU acceleration resources. FaaSter provides an acceleration function library as a service, which in turn is provisioned on heterogeneous GPUs. To provide seamless access to these accelerated functions and ensure that each function has the best possible response time, we utilize GPU kernel slicing to split and execute an accelerated function instance across multiple heterogeneous GPUs. The central challenge is to be able to quickly decide the number of slices to split each function into and then map the slices to the right GPUs. To this end, we present a scheduling heuristic that is able to significantly reduce the average turn-around time of functions when compared to a non-sliceable full GPU scheduling approach. Our evaluation results show that the FaaSter scheduler achieves 62 % mean and up to 80 % improvement in average turn-around time and always performs equal to or better than non-sliceable GPU scheduling. Anshuj Garg, Purushottam Kulkarni, Umesh Bellur, Sriram Yenamandra |
HiPC | 1 |
| 2021 | Optimizing Goodput of Real-time Serverless Functions using Dynamic Slicing with vGPUsabstractAs the popularity and relevance of the Function-as-a-Service (FaaS) model keeps growing, we believe newer avatars of the service will support computationally intensive SIMT functions that will execute on GPUs. With hardware-assisted virtualization of GPUs now possible, cloud offerings including GPUs usually bind a virtual GPU (vGPU) to a VM. While there is a choice of scheduling algorithms to multiplex vGPUs on to the physical GPU, the work-conserving best-effort scheduler helps to maintain a high level of utilization of the GPU. With this, we observe that the total share of the GPU per VM is non-deterministic and depends on how different VMs load the GPU via their vGPUs. As a result, any function-to-vGPU scheduler that does not explicitly account for this nondeterministic vGPU capacity will suffer from lower than optimal goodput - particularly when these functions are deadline bound as is the case with FaaS offerings today. In this work, we exploit a software based task slicing technique to dynamically determine task sizes for scheduling on vGPUs to maximize successful completion of functions within their deadlines. Our solution extends the conventional earliest-deadline first (EDF) scheduling algorithm by balancing scheduling opportunities (via kernel slicing) and maximizing the chances of functions finishing before their deadline. The work is motivated by the fact that static decisions that consider entire tasks as scheduling units or use a fixed, statically decided slice size as scheduling units cannot adapt to the non-deterministic vGPU capacity. A comparison of our solution with well-known deadline aware scheduling approach (earliest deadline first), yielded an improvement of up to 2.9x in goodput. Anshuj Garg, Umesh Bellur, Purushottam Kulkarni, Uday Kurkure, Hari Sivaraman, Lan Vu |
IC2E | 2 |
| 2019 | Empirical Analysis of Hardware-Assisted GPU VirtualizationabstractThe increasing use of Graphics Processing Unit (GPUs) for accelerating compute intensive tasks and graphics-related computations has led to their inclusion in High Performance Clusters and Cloud setups. Several cloud vendors provide virtual machine instances with GPU capabilities. With the advent of virtualization aware GPU hardware (NVIDIA vGPUs), allocating and sharing physical GPU among virtual machines has become easier and cost-efficient. The sharing mechanism and extent of sharing are determined by the vGPU scheduling algorithm and a configurable vGPU profile. As part of this work, we present a thorough empirical study of the hardware-assisted virtualized GPU setups. In particular, we quantify the virtualization overheads, study the interference effects of concurrently executing homogeneous and heterogeneous workloads, and impact of vGPU scheduling algorithms. We also demonstrate that the best vGPU configuration parameters are sensitive to the mix of workload characteristics that share the GPU. Our study also compares the performance of vGPU based virtualization with PCI passthrough based direct GPU assignment to virtual machines. Based on our evaluation using heterogeneous workloads setups and varying vGPU configurations, we observe that the virtualization overheads are up to 7% (in terms of reduced memory availability) and up to 20% increase in execution times. Further, we demonstrate using our setups, that co-placing of heterogeneous workloads can improve the efficiency of GPU multiplexing and decrease execution times to as low as 20% as compared to homogeneous workload placements. Anshuj Garg, Purushottam Kulkarni, Uday Kurkure, Hari Sivaraman, Lan Vu |
HiPC | 1 |
| 2019 | Dynamic Memory Management for GPU-Based Training of Deep Neural NetworksabstractDeep learning has been widely adopted for different applications of artificial intelligence-speech recognition, natural language processing, computer vision etc. The growing size of Deep Neural Networks (DNNs) has compelled the researchers to design memory efficient and performance optimal algorithms. Apart from algorithmic improvements, specialized hardware like Graphics Processing Units (GPUs) are being widely employed to accelerate the training and inference phases of deep networks. However, the limited GPU memory capacity limits the upper bound on the size of networks that can be offloaded to and trained using GPUs. vDNN addresses the GPU memory bottleneck issue and provides a solution which enables training of deep networks that are larger than GPU memory. In our work, we characterize and identify multiple bottlenecks with vDNN like delayed computation start, high pinned memory requirements and GPU memory fragmentation. We present vDNN++ which extends vDNN and resolves the identified issues. Our results show that the performance of vDNN++ is comparable or better (up to 60% relative improvement) than vDNN. We propose different heuristics and order for memory allocation, and empirically evaluate the extent of memory fragmentation with them. We are also able to reduce the pinned memory requirement by up to 60%. Shriram S. B, Anshuj Garg, Purushottam Kulkarni |
IPDPS | 2 |
| 2017 | Catalyst: GPU-assisted rapid memory deduplication in virtualization environmentsabstractContent based page sharing techniques improve memory efficiency in virtualized systems by identifying and merging identical pages. Kernel Same-page Merging (KSM), a Linux kernel utility for page sharing, sequentially scans memory pages of virtual machines to deduplicate pages. Sequential scanning of pages has several undesirable side effects---wasted CPU cycles when no sharing opportunities exist, and rate of discovery of sharing being dependent on the scanning rate and corresponding CPU availability. In this work, we exploit presence of GPUs on modern systems to enable rapid memory sharing through targeted scanning of pages. Our solution, Catalyst, works in two phases, the first where pages of virtual machines are processed by the GPU to identify likely pages for sharing and a second phase that performs page-level similarity checks on a targeted set of shareable pages. Opportunistic usage of the GPU to produce sharing hints enables rapid and low-overhead duplicate detection, and sharing of memory pages in virtualization environments. We evaluate Catalyst against various benchmarks and workloads to demonstrate that Catalyst can achieve higher memory sharing in lesser time compared to different scan rate configurations of KSM, at lower or comparable compute costs. Anshuj Garg, Debadatta Mishra, Purushottam Kulkarni |
VEE | 1 |
| 2013 | GAGM: Genome assembly on GPU using mate pairsabstractGenome fragment assembly has long been a time and computation intensive problem in the field of bioinformatics. Many parallel assemblers have been proposed to accelerate the process but there hasn't been any effective approach proposed for GPUs. Also with the increasing power of GPUs, applications from various research fields are being parallelized to take advantage of the massive number of “cores” available in GPUs. In this paper we present the design and development of a GPU based assembler (GAGM) for sequence assembly using Nvidia's GPUs with the CUDA programming model. Our assembler utilizes the mate pair reads produced by the current NGS technologies to build paired de Bruijn graph. Every paired read is broken into paired k-mers and l-mers. Every paired k-mer represents a vertex and paired l-mers are mapped as edges. Contigs are formed by grouping the regions of graph which can be unambiguously connected. We present parallel algorithms for k - mer extraction, paired de Bruijn graph construction and grouping of edges. We have benchmarked GAGM on four bacterial genomes. Our results show that the design on GPU is effective in terms of time as well as the quality of assembly produced. Ashutosh Jain, Anshuj Garg, Kolin Paul |
HiPC | 2 |