VLDB 2026 Research / reviewers in the wild / expert
Kaustubh Shivdikar
dblp:254/2627
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-4449-7974ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMabstractProcessing-in-memory (PIM), where compute is moved closer to memory or data, has been explored to accelerate emerging workloads. Different PIM-based systems have been announced, each offering a unique microarchitectural organization of their compute units, ranging from fixed functional units to programmable general-purpose compute cores near memory. However, one fundamental limitation of PIM is that each compute unit can only access its local memory; access to “remote” memory must occur through the host CPU - potentially limiting application performance scalability. In this work, we first characterize the scalability of real PIM architectures using the UPMEM PIM system. We analyze how the overhead of communicating through the host (instead of providing direct communication between the PIM compute units) can become a bottleneck for collective communications that are commonly used in many workloads. To overcome this inter-PIM bank communication, we propose PIMnet - a PIM interconnection network for PIM banks that provides direct connectivity between compute units and removes the overhead of communicating through the host. PIMnet exploits bandwidth parallelism where communication across the different PIM bank/chips can occur in parallel to maximize communication performance. PIMnet also matches the DRAM packaging hierarchy with a multi-tier network architecture. Unlike traditional interconnection networks, PIMnet is a PIMcontrolled network where communication is managed by the PIM logic, optimizing collective communications and minimizing the hardware overhead of PIMnet. Our evaluation of PIMnet shows that it provides up to $85 \times$ speedup on collective communications and achieves a $11.8 \times$ improvement on real applications compared to the baseline PIM. Hyojun Son, Gilbert Jonatan, Haeyoon Cho 0002, Kaustubh Shivdikar, José L. Abellán, Ajay Joshi, David R. Kaeli, John Kim 0001 |
HPCA | 5 |
| 2025 | FIDESlib: A Fully-Fledged Open-Source FHE Library for Efficient CKKS on GPUsabstractWord-wise Fully Homomorphic Encryption (FHE) schemes, such as CKKS, are gaining significant traction due to their ability to provide post-quantum-resistant, privacypreserving approximate computing-an especially desirable feature in the Machine-Learning-as-a-Service (MLaaS) paradigm. In this work, we introduce FIDESlib, the first open-source server-side CKKS GPU library that is fully interoperable with well-established client-side OpenFHE operations. Unlike other existing open-source GPU libraries, FIDESlib provides the first implementation featuring heavily optimized GPU kernels for all CKKS primitives, including bootstrapping. Our library also integrates robust benchmarking and testing, ensuring it remains adaptable to further optimization. Comparing our scheme against Phantom (the previously top open-source CKK library, we show that FIDESlib offers superior performance and scalability. For bootstrapping, FIDESlib achieves no less than$70 \times$speedup over the AVX-optimized OpenFHE implementation. FIDESlib is available on Github11https://github.com/CAPS-UMU/FIDESlib. Carlos Agulló-Domingo, Óscar Vera-López, Seyda Nur Güzelhan, Lohit Daksha, Aymane El Jerari, Kaustubh Shivdikar, Rashmi S. Agrawal 0001, David R. Kaeli, Ajay Joshi, José L. Abellán |
ISPASS | 6 |
| 2024 | MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks TrainingabstractIn the acceleration of deep neural network training, the graphics processing unit (GPU) has become the mainstream platform. GPUs face substantial challenges on Graph Neural Networks (GNNs), such as workload imbalance and memory access irregularities, leading to underutilized hardware. Existing solutions such as PyG, DGL with cuSPARSE, and GNNAdvisor frameworks partially address these challenges. However, the memory traffic involved with Sparse-Dense Matrix Matrix Multiplication (SpMM) is still significant. Hongwu Peng, Kaustubh Shivdikar, Amit Hasan 0001, Shaoyi Huang, Omer Khan, David R. Kaeli, Caiwen Ding |
ASPLOS (2) | 3 |
| 2024 | NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial AcceleratorabstractGraph Neural Networks (GNNs) are emerging as a formidable tool for processing non-euclidean data across various domains, ranging from social network analysis to bioinformatics. Despite their effectiveness, their adoption has not been pervasive because of scalability challenges associated with large-scale graph datasets, particularly when leveraging message passing. They exhibit irregular sparsity patterns, resulting in unbalanced compute resource utilization. Prior accelerators investigating Gustavson’s technique adopted look-ahead buffers for prefetching data, aiming to prevent compute stalls. However, these solutions lead to inefficient use of the on-chip memory, leading to redundant data residing in cache.To tackle these challenges, we introduce NeuraChip, a novel GNN spatial accelerator based on Gustavson’s algorithm. NeuraChip decouples the multiplication and addition computations in sparse matrix multiplication. This separation allows for independent exploitation of their unique data dependencies, facilitating efficient resource allocation. We introduce a rolling eviction strategy to mitigate data idling in on-chip memory as well as address the prevalent issue of memory bloat in sparse graph computations. Furthermore, the compute resource load balancing is achieved through a dynamic reseeding hash-based mapping, ensuring uniform utilization of computing resources agnostic of sparsity patterns. Finally, we present NeuraSim, an open-source, cycle-accurate, multi-threaded, modular simulator for comprehensive performance analysis.Overall, NeuraChip presents a significant improvement, yielding an average speedup of $22.1 \times$ over Intel’s MKL, $17.1 \times$ over NVIDIA’s cuSPARSE, $16.7 \times$ over AMD’s hipSPARSE, and $1.5 \times$ over prior state-of-the-art SpGEMM accelerator and $1.3 \times$ over GNN accelerator. The source code for our open-sourced simulator and performance visualizer is publicly accessible on GitHub1. CCS CONCEPTS • Computer systems organization → Multicore architectures; Interconnection architectures; • Computing methodologies → Neural networks; • Theory of computation → Graph algorithms analysis; • Hardware → Hardware accelerators.1https://github.com/NeuraChip/neurachip Kaustubh Shivdikar, Nicolas Bohm Agostini, Malith Jayaweera, Gilbert Jonatan, José L. Abellán, Ajay Joshi, John Kim 0001, David R. Kaeli |
ISCA | 1 |
| 2023 | GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables the processing of encrypted data without decrypting it. FHE has garnered significant attention over the past decade as it supports secure outsourcing of data processing to remote cloud services. Despite its promise of strong data privacy and security guarantees, FHE introduces a slowdown of up to five orders of magnitude as compared to the same computation using plaintext data. This overhead is presently a major barrier to the commercial adoption of FHE. Kaustubh Shivdikar, Yuhui Bao, Rashmi S. Agrawal 0001, Michael Tian Shen, Gilbert Jonatan, Evelio Mora, Alexander Ingare, Neal Livesay, José L. Abellán, John Kim 0001, Ajay Joshi, David R. Kaeli |
MICRO | 1 |
| 2021 | GNNMark: A Benchmark Suite to Characterize Graph Neural Network Training on GPUsabstractGraph Neural Networks (GNNs) have emerged as a promising class of Machine Learning algorithms to train on non-euclidean data. GNNs are widely used in recommender systems, drug discovery, text understanding, and traffic forecasting. Due to the energy efficiency and high-performance capabilities of GPUs, GPUs are a natural choice for accelerating the training of GNNs. Thus, we want to better understand the architectural and system-level implications of training GNNs on GPUs. Presently, there is no benchmark suite available designed to study GNN training workloads. In this work, we address this need by presenting GNNMark, a feature-rich benchmark suite that covers the diversity present in GNN training workloads, datasets, and GNN frameworks. Our benchmark suite consists of GNN workloads that utilize a variety of different graph-based data structures, including homogeneous graphs, dynamic graphs, and heterogeneous graphs commonly used in a number of application domains that we mentioned above. We use this benchmark suite to explore and characterize GNN training behavior on GPUs. We study a variety of aspects of GNN execution, including both compute and memory behavior, highlighting major bottlenecks observed during GNN training. At the system level, we study various aspects, including the scalability of training GNNs across a multi-GPU system, as well as the sparsity of data, encountered during training. The insights derived from our work can be leveraged by both hardware and software developers to improve both the hardware and software performance of GNN training on GPUs. Trinayan Baruah, Kaustubh Shivdikar, Shi Dong 0002, Yifan Sun 0002, Saiful A. Mojumder, Kihoon Jung, José L. Abellán, Yash Ukidave, Ajay Joshi, John Kim 0001, David R. Kaeli |
ISPASS | 2 |
| 2019 | Student cluster competition 2018, team northeastern university: Reproducing performance of a multi-physics simulations of the Tsunamigenic 2004 Sumatra Megathrust earthquake on the AMD EPYC 7551 architecture
Christopher C. Bunn, Harrison Barclay, A. Lazarev, F. Yusuf, J. Fitch, J. Booth, Kaustubh Shivdikar, David R. Kaeli |
Parallel Comput. | 7 |