EDBT 2026 Demo / reviewers in the wild / expert
Ajay Nayak
dblp:294/4252
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-5313-1328ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 46% Memory systems · 34% Cloud and datacenter computing · 20% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 50% Concurrent programming · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › inference serving
large language model serving |
0.9 | 1 | 2025 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025 |
Memory systems › memory management › memory allocation
dynamic memory allocation |
0.9 | 1 | 2025 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025 |
GPUs and heterogeneous computing
GPU memory management |
0.9 | 1 | 2025 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025 |
Memory systems › cache management
KV cache management |
0.9 | 1 | 2025 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025 |
GPUs and heterogeneous computing
GPU programming |
0.8 | 1 | 2024 | Over-Synchronization in GPU Programs · MICRO 2024 |
GPUs and heterogeneous computing › GPU computing
GPU synchronization |
0.8 | 1 | 2024 | Over-Synchronization in GPU Programs · MICRO 2024 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.5 | 1 | 2021 | Faastlane: Accelerating Function-as-a-Service Workflows · USENIX ATC 2021 |
Cloud and datacenter computing
serverless computing |
0.5 | 1 | 2021 | Faastlane: Accelerating Function-as-a-Service Workflows · USENIX ATC 2021 |
Program analysis
static analysis |
0.2 | 1 | 2024 | Over-Synchronization in GPU Programs · MICRO 2024 |
Concurrent programming › concurrency analysis
synchronization analysis |
0.2 | 1 | 2024 | Over-Synchronization in GPU Programs · MICRO 2024 |
Methods — techniques the papers use, named apart from their topics
virtual memory management · 2.6static analysis · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionabstractPagedAttention is a popular approach for dynamic memory allocation in LLM serving systems. It enables on-demand allocation of GPU memory to mitigate KV cache fragmentation - a phenomenon that crippled the batch size (and consequently throughput) in prior systems. However, in trying to allocate physical memory at runtime, PagedAttention ends up changing the virtual memory layout of the KV cache from contiguous to non-contiguous. Such a design leads to non-trivial programming and performance overheads. Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar |
ASPLOS (1) | 2 |
| 2024 | Over-Synchronization in GPU ProgramsabstractThe performance of GPU (Graphics Processing Unit)-accelerated functions affects a large spectrum of modern software. Efficiently synchronizing across thousands of concurrent threads is critical to the performance of GPU programs. GPU vendors have introduced advanced programming constructs, e.g., scopes, for efficiently synchronizing within a chosen subset of threads. However, programmers must explicitly employ them, where applicable, to benefit from such features. We demonstrate how GPU programs can leave performance on the table by failing to fully harness advanced synchronization features in modern GPUs - leading to over-synchronization. We discover three different variants of over-synchronization observed in real-world applications. We then build a tool, ScopeAdvice, to find cases of over-synchronization in CUDA programs. Avoiding reported over-synchronization improves the performance of several GPU applications by up to 55%. Ajay Nayak, Arkaprava Basu |
MICRO | 1 |
| 2021 | (Mis)managed: A Novel TLB-based Covert Channel on GPUsabstractGPUs are now commonly available in most modern computing platforms. They are increasingly being adopted in cloud platforms and data centers due to their immense computing capability. In response to this growth in usage, manufacturers continuously try to improve GPU hardware by adding new features. However, this increase in usage and the addition of utility-improving features can create new, unexpected attack channels. In this paper, we show that two such features-unified virtual memory (UVM) and multi-process service (MPS)-primarily introduced to improve the programmability and efficiency of GPU kernels have an unexpected consequence-that of creating a novel covert-timing channel via the GPU's translation lookaside buffer (TLB) hierarchy. To enable this covert channel, we first perform experiments to understand the characteristics of TLBs present on a GPU. The use of UVM allows fine-grained management of translations, and helps us discover several idiosyncrasies of the TLB hierarchy, such as three-levels of TLB, coalesced entries. We use this newly-acquired understanding to demonstrate a novel covert channel via the shared TLB. We then leverage MPS to increase the bandwidth of this channel by 40×. Finally, we demonstrate the channel's utility by leaking data from a GPU-accelerated database application. Ajay Nayak, Pratheek B, Vinod Ganapathy, Arkaprava Basu |
AsiaCCS | 1 |
| 2021 | Faastlane: Accelerating Function-as-a-Service Workflows
Swaroop Kotni, Ajay Nayak, Vinod Ganapathy, Arkaprava Basu |
USENIX ATC | 2 |