Aditya K. Kamath

dblp:245/7176 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-6565-3764ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Reducing the GPU Memory Bottleneck with Lossless Compression for ML
abstract
Machine learning (ML) training and inference often process data sets far exceeding GPU memory capacity, forcing them to rely on PCIe for on-demand tensor transfers, causing critical transfer bottlenecks. Lossy compression has been proposed to relieve bottlenecks but introduces workload-dependent accuracy loss, making it complex or even prohibitive to use in existing ML deployments.
Aditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon Peter 0001
EuroSys1
2025 POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 0001, Ramachandran Ramjee, Ashish Panwar
ASPLOS (2)1
2024 Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V Manycore
abstract
Existing tiled manycore architectures propose to convert abundant silicon resources into general-purpose parallel processors with unmatched computational density and programmability. However, as we approach 100 K cores in one chip, conventional manycore architectures struggle to navigate three key axes: scalability, programmability, and density. Many manycores sacrifice programmability for density; or scalability for programmability. In this paper, we explore HammerBlade, which simultaneously achieves scalability, programmability and density. HammerBlade is a fully open-source RISC-V manycore architecture, which has been silicon-validated with a 2048-core ASIC implementation using a 14/16nm process. We evaluate the system using a suite of parallel benchmarks that captures a broad spectrum of computation and communication patterns.
Dai Cheol Jung, Max Ruttenberg, Paul Gao 0001, Scott Davidson 0004, Daniel Ruelas-Petrisko, Kangli Li, Aditya K. Kamath, Shaolin Xie, Peitian Pan, Zhongyuan Zhao 0004, Zichao Yue, Bandhav Veluri, Sripathi Muralitharan, Adrian Sampson, Andrew Lumsdaine, Zhiru Zhang, Christopher Batten, Mark Oskin, Dustin Richmond, Michael B. Taylor
ISCA7
2024 (MC)2: Lazy MemCopy at the Memory Controller
abstract
(MC)2is a lazy memory copy mechanism which can be used within memcpy-like functions to significantly reduce the CPU overhead for copies that are sparsely accessed. It can also hide copy latencies by enhancing the CPU’s ability to execute them asynchronously. (MC)2,s lazy memcpy avoids copying data at the time of invocation. Instead, (MC)2tracks prospective copies. If copied data is later accessed by a CPU or the cache, (MC)2uses the tracking information to lazily execute a copy, when necessary. Placing (MC)2at the memory controller puts it at the perfect vantage point to eliminate the largest source of memcpy overhead–CPU stalls due to cache misses in the critical path–while imposing minimal overhead itself. (MC)2consists of three main components: memory controller extensions that implement a lazy memcpy operation, a new instruction exposing the lazy memcpy, and a flexible software wrapper with semantics identical to memcpy. We implement and evaluate (MC)2in the gem5 simulator using a variety of microbenchmarks and workloads, including Google’s Protobuf, where (MC)2provides a $43 \%$ speedup and Linux huge page copy-on-write faults, where (MC)2provides $250 \times$ lower latency.
Aditya K. Kamath, Simon Peter 0001
ISCA1
2023 Scoped Buffered Persistency Model for GPUs
abstract
While the implications of persistent memory (PM) on CPU hardware and software are well-explored, the same is not true for GPUs (Graphics Processing Units). A recent work, GPM, demonstrated how GPU programs can benefit from the fine-grain persistence of PM. However, in the absence of a persistency model, one cannot reason about the correctness of PM-aware GPU programs. Persistency models define the order in which writes to PM are persisted. We explore persistency models for GPUs.
Shweta Pandey 0001, Aditya K. Kamath, Arkaprava Basu
ASPLOS (2)2
2022 GPM: leveraging persistent memory from a GPU
abstract
The GPU is a key computing platform for many application domains. While the new non-volatile memory technology has brought the promise of byte-addressable persistence (a.k.a., persistent memory, or PM) to CPU applications, the same, unfortunately, is beyond the reach of GPU programs.
Shweta Pandey 0001, Aditya K. Kamath, Arkaprava Basu
ASPLOS2
2021 iGUARD: In-GPU Advanced Race Detection
abstract
Newer use cases of GPU (Graphics Processing Unit) computing, e.g., graph analytics, look less like traditional bulk-synchronous GPU programs. To cater to the needs of emerging applications with semantically richer and finer grain sharing patterns, GPU vendors have been introducing advanced programming features, e.g., scoped synchronization and independent thread scheduling. While these features can speed up many applications and enable newer use cases, they can also introduce subtle synchronization errors if used incorrectly.
Aditya K. Kamath, Arkaprava Basu
SOSP1
2020 ScoRD: A Scoped Race Detector for GPUs
abstract
GPUs have emerged as a key computing platform for an ever-growing range of applications. Unlike traditional bulk-synchronous GPU programs, many emerging GPU-accelerated applications, such as graph processing, have irregular interaction among the concurrent threads. Consequently, they need complex synchronization. To enable both high performance and adequate synchronization, GPU vendors have introduced scoped synchronization operations that allow a programmer to synchronize within a subset of concurrent threads (a.k. a., scope) that she deems adequate. Scoped-synchronization avoids the performance overhead of synchronization across thousands of GPU threads while ensuring correctness when used appropriately. This flexibility, however, could be a new source of incorrect synchronization where a race can occur due to insufficient scope of the synchronization operation, and not due to missing synchronization as in a typical race. We introduce ScoRD, a race detector that enables hardware support for efficiently detecting global memory races in a GPU program, including those that arise due to insufficient scopes of synchronization operations. We show that ScoRD can detect a variety of races with a modest performance overhead (on average, 35%). In the process of this study, we also created a benchmark suite consisting of seven applications and three categories of microbenchmarks that use scoped synchronization operations.
Aditya K. Kamath, Alvin A. George, Arkaprava Basu
ISCA1
2018 Sobriety Testing Based on Thermal Infrared Images Using Convolutional Neural Networks
abstract
This paper proposes a method to test the sobriety of an individual using infrared images of the persons eyes, face, hand, and facial profile. The database we used consisted of images of forty different individuals. The process is broken down into two main stages. In the first stage, the data set was divided according to body part and each one was run through its own Convolutional Neural Network (CNN). We then tested the resulting network against a validation data set. The results obtained gave us an indication of which body parts were better suited for identifying signs of drunken state and sobriety. In the second stage, we took the weights of CNN giving best validation accuracy from the first stage. We then grouped the body parts according to the person they belong to. The body parts were fed together into a CNN using the weights obtained in the first stage. The result for each body part was passed to a simple back-propagation neural network (BPNN) to get final results. We tried to identify the most optimal configuration of neural networks for each stage of the process. The results we obtained showed that facial profile images tend to give very good indications of sobriety. The results also showed that combining the results of multiple body parts using a simple BPNN gives a higher accuracy than that of individual ones.
Aditya K. Kamath, A. Tarun Karthik, Leslie Monis, Manjunath Mulimani, Shashidhar G. Koolagudi
TENCON1