EDBT 2026 Demo / reviewers in the wild / expert
Anurag Khandelwal
dblp:162/1936
· DBLP profile ↗
25ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-2199-6391ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 12 since 2021Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Computer networks · 6 · 2 first-author · 3 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural AssumptionsabstractHardware event counters offer the potential to reveal not only performance bottlenecks but also detailed microarchitectural behavior. In practice, this promise is undermined by their vague specifications, opaque designs, and multiplexing noise, making event counter data hard to interpret. We introduce CounterPoint, a framework that tests user-specified microarchitectural models — expressed as ?path Decision Diagrams — for consistency with performance counter data. When mismatches occur, CounterPoint pinpoints plausible microarchitectural features that could explain them, using multi-dimensional counter confidence regions to mitigate multiplexing noise. We apply CounterPoint to the Haswell Memory Management Unit as a case study, shedding light on multiple undocumented and underdocumented microarchitectural behaviors. These include a load–store queue-side TLB prefetcher, merging page table walkers, abortable page table walks, and more. Overall, CounterPoint helps experts reconcile noisy hardware performance counter measurements with their mental model of the microarchitecture - uncovering subtle, previously hidden hardware features along the way. Nick Lindsay, Caroline Trippel, Anurag Khandelwal, Abhishek Bhattacharjee |
ASPLOS (2) | 3 |
| 2026 | BulletTime: Time Dilation for High-Fidelity Tracing
Sibren Isaacman, Abhishek Bhattacharjee, Anurag Khandelwal |
ISCA | 4 |
| 2026 | TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited AgentsabstractLarge Language Models (LLMs) are increasingly integrated into Physical-I/O limited agents, such as robots and voice assistants, which execute outputs sequentially. However, existing LLM serving systems typically employ a throughput-oriented batching mechanism, ignoring the large gap between LLM generation speed and the constrained physical I/O rates of agents, thus wasting execution slack and worsening resource contention. Besides, they treat all tokens equally and cannot anticipate the execution implications of different content, preventing scheduling aligned with agent-side behavior. To address it, we propose a new system named TimelyLLM that coordinates LLM generation with the physical behavior of agents. TimelyLLM introduces a novel segmented generation and scheduling mechanism, strategically leveraging the time gap between agent plan generation and execution to reduce contention and improve response latency under multi-agent workloads. We implement TimelyLLM on top of a widely-used LLM serving framework. We also build a dataset collection system to construct serving workloads from real-world robots, including drones, robot arms, and quadruped robots. Our evaluation demonstrates that TimelyLLM improves the time utility up to 1.52×, and reduces the overall waiting time by 84%. Neiwen Ling, Anurag Khandelwal, Lin Zhong 0001 |
MobiSys | 3 |
| 2025 | pulse: Accelerating Distributed Pointer-Traversals on Disaggregated MemoryabstractCaches at CPU nodes in disaggregated memory architectures amortize the high data access latency over the network. However, such caches are fundamentally unable to improve performance for workloads requiring pointer traversals across linked data structures. We argue for accelerating these pointer traversals closer to disaggregated memory in a manner that preserves expressiveness for supporting various linked structures, ensures energy efficiency and performance, and supports distributed execution. We design pulse, a distributed pointer-traversal framework for rack-scale disaggregated memory to meet all the above requirements. Our evaluation of pulse shows that it enables low-latency, high-throughput, and energy-efficient execution for a wide range of pointer traversal workloads on disaggregated memory that fare poorly with caching alone. Yupeng Tang, SeungSeob Lee, Abhishek Bhattacharjee, Anurag Khandelwal |
ASPLOS (1) | 4 |
| 2025 | CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory OrderingabstractIncreasingly, multi-processing unit (PU) systems (e.g., CPU-GPU, multi-CPU, multi-GPU, etc.) are embracing cache-coherent shared memory to facilitate inter-PU communication.The coherence protocols in these systems support write-through accesses that place the data directly at the LLC to enable efficient producer-consumer communications pervasive in AI/ML workloads.Moreover, release consistency has emerged as the standard memory model in such systems due to its programming simplicity and ability to support high performance.In today's multi-PU systems, the source processor that issues the writes also orders them to enforce release consistency, even for write-through accesses.Unfortunately, such source ordering of write-through operations results in unnecessary communications between the source processor and the LLC directory, incurring significant performance, interconnect traffic, and energy overheads for multi-PU applications.To eliminate such communication, we present cord 1 , a novel cache coherence protocol that orders write-through accesses directly at the cache directory.cord employs several novel mechanisms to minimize the metadata required for ordering traffic while efficiently scaling to multiple directories.Evaluations atop the gem5 simulator show that compared to source ordering, cord improves application performance by 24% and reduces traffic by 13% on average while incurring < 1% storage, area, and power overheads.Compared to hand-optimized message-passing implementations, cord observes a mere 3% performance overhead and 6% more traffic on average with a significantly simpler programming model. Yanpeng Yu, Nicolai Oswald, Anurag Khandelwal |
ISCA | 3 |
| 2025 | Weave: Efficient and Expressive Oblivious Analytics at Scale
Mahdi Soleimani, Grace Jia, Anurag Khandelwal |
OSDI | 3 |
| 2025 | Spirit: Fair Allocation of Interdependent Resources in Remote Memory SystemsabstractWe address the problem of fair resource allocation in multiuser remote memory systems. Allocating local memory (used as cache) and network bandwidth to remote memory in such systems is challenging due to the complex interdependence between the two resources and application performance. A larger cache may reduce the need for fetching data over the network, while a larger bandwidth may permit more concurrent network requests, avoiding the need for large caches. As a result, applications can achieve the same data access throughput for a wide range of cache and bandwidth allocations. Such interdependence is unique to each application and hard to capture offline. SeungSeob Lee, Jachym Putta, Ziming Mao, Anurag Khandelwal |
SOSP | 4 |
| 2025 | Scalable Far Memory: Balancing Faults and EvictionsabstractPage-based far memory systems transparently expand an application's memory capacity beyond a single machine without modifying application code. However, existing systems are tailored to scenarios with low application thread counts, and fail to scale on today's multi-core machines. This makes them unsuitable for data-intensive applications that both rely on far memory support and scale with increasing thread count. Our analysis reveals that this poor scalability stems from inefficient holistic coordination between page fault-in and eviction operations. As thread count increases, current systems encounter scalability bottlenecks in TLB shootdowns, page accounting, and memory allocation. Yueyang Pan, Yash Lala, Musa Unal, Yujie Ren, SeungSeob Lee, Abhishek Bhattacharjee, Anurag Khandelwal, Sanidhya Kashyap |
SOSP | 7 |
| 2025 | Found in Translation: A Generative Language Modeling Approach to Memory Access Pattern Attacks
Grace Jia, Alex Wong 0001, Anurag Khandelwal |
USENIX Security Symposium | 3 |
| 2024 | Trinity: A Fast Compressed Multi-attribute Data StoreabstractWith the proliferation of attribute-rich machine-generated data, emerging real-time monitoring, diagnosis, and visualization tools ingest and analyze such data across multiple attributes simultaneously. Due to the sheer volume of the data, applications need storage-efficient and performant data representations to analyze them efficiently. Ziming Mao, Kiran Srinivasan, Anurag Khandelwal |
EuroSys | 3 |
| 2024 | Length Leakage in Oblivious Data Access Mechanisms
Grace Jia, Rachit Agarwal 0001, Anurag Khandelwal |
USENIX Security Symposium | 3 |
| 2023 | Prefetching Using Principles of Hippocampal-Neocortical InteractionabstractMemory prefetching improves performance across many systems layers. However, achieving high prefetch accuracy with low overhead is challenging, as memory hierarchies and application memory access patterns become more complicated. Furthermore, a prefetcher's ability to adapt to new access patterns as they emerge is becoming more crucial than ever. Recent work has demonstrated the use of deep learning techniques to improve prefetching accuracy, albeit with impractical compute and storage overheads. This paper suggests taking inspiration from the learning mechanisms and memory architecture of the human brain---specifically, the hippocampus and neocortex---to build resource-efficient, accurate, and adaptable prefetchers. Ketaki Joshi, Andrew Sheinberg, Guilherme Cox, Anurag Khandelwal, Raghavendra Pradyumna Pothukuchi, Abhishek Bhattacharjee |
HotOS | 5 |
| 2023 | SCALO: An Accelerator-Rich Distributed System for Scalable Brain-Computer InterfacingabstractSCALO is the first distributed brain-computer interface (BCI) consisting of multiple wireless-networked implants placed on different brain regions. SCALO unlocks new treatment options for debilitating neurological disorders and new research into brain-wide network behavior. Achieving the fast and low-power communication necessary for real-time processing has historically restricted BCIs to single brain sites. SCALO also adheres to tight power constraints, but enables fast distributed processing. Central to SCALO's efficiency is its realization as a full stack distributed system of brain implants with accelerator-rich compute. SCALO balances modular system layering with aggressive cross-layer hardware-software co-design to integrate compute, networking, and storage. The result is a lesson in designing energy-efficient networked distributed systems with hardware accelerators from the ground up. Karthik Sriram, Raghavendra Pradyumna Pothukuchi, Michal Gerasimiuk, Muhammed Ugur, Oliver Ye, Rajit Manohar, Anurag Khandelwal, Abhishek Bhattacharjee |
ISCA | 7 |
| 2023 | SHEPHERD: Serving DNNs in the Wild
Hong Zhang 0025, Yupeng Tang, Anurag Khandelwal, Ion Stoica |
NSDI | 3 |
| 2023 | Karma: Resource Allocation for Dynamic Demands
Midhul Vuppalapati, Giannis Fikioris, Rachit Agarwal 0001, Asaf Cidon, Anurag Khandelwal, Éva Tardos |
OSDI | 5 |
| 2022 | Jiffy: elastic far-memory for stateful serverless analyticsabstractStateful serverless analytics can be enabled using a remote memory system for inter-task communication, and for storing and exchanging intermediate data. However, existing systems allocate memory resources at job granularity---jobs specify their memory demands at the time of the submission; and, the system allocates memory equal to the job's demand for the entirety of its lifetime. This leads to resource underutilization and/or performance degradation when intermediate data sizes vary during job execution. Anurag Khandelwal, Yupeng Tang, Rachit Agarwal 0001, Aditya Akella, Ion Stoica |
EuroSys | 1 |
| 2022 | SHORTSTACK: Distributed, Fault-tolerant, Oblivious Data Access
Midhul Vuppalapati, Kushal Babel, Anurag Khandelwal, Rachit Agarwal 0001 |
OSDI | 3 |
| 2021 | Caerus: NIMBLE Task Scheduling for Serverless Analytics
Hong Zhang 0025, Yupeng Tang, Anurag Khandelwal, Jingrong Chen 0002, Ion Stoica |
NSDI | 3 |
| 2021 | MIND: In-Network Memory Management for Disaggregated Data CentersabstractMemory disaggregation promises transparent elasticity, high resource utilization and hardware heterogeneity in data centers by physically separating memory and compute into network-attached resource "blades". However, existing designs achieve performance at the cost of resource elasticity, restricting memory sharing to a single compute blade to avoid costly memory coherence traffic over the network. SeungSeob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong 0001, Abhishek Bhattacharjee |
SOSP | 4 |
| 2020 | Le Taureau: Deconstructing the Serverless Landscape & A Look ForwardabstractAkin to the natural evolution of programming in assembly language to high-level languages, serverless computing represents the next frontier in the evolution of cloud computing: bare metal -> virtual machines -> containers -> serverless. The genesis of serverless computing can be traced back to the fundamental need of enabling a programmer to singularly focus on writing application code in a high-level language and isolating all facets of system management (for example, but not limited to, instance selection, scaling, deployment, logging, monitoring, fault tolerance and so on). This is particularly critical in light of today's, increasingly tightening, time-to-market constraints. Currently, serverless computing is supported by leading public cloud vendors, such as AWS Lambda, Google Cloud Functions, Azure Cloud Functions and others. While this is an important step in the right direction, there are many challenges going forward. For instance, but not limited to, how to enable support for dynamic optimization, how to extend support for stateful computation, how to efficiently bin-pack applications, how to support hardware heterogeneity (this will be key especially in light of the emergence of hardware accelerators for deep learning workloads). Inspired by Picasso's Le Taureau, in the tutorial proposed herein, we shall deconstruct evolution of serverless --- the overarching intent being to facilitate better understanding of the serverless landscape. This, we hope, would help push the innovation frontier on both fronts, the paradigm itself and the applications built atop of it. Anurag Khandelwal, Arun Kejariwal, Karthikeyan Ramasamy |
SIGMOD Conference | 1 |
| 2020 | Pancake: Frequency Smoothing for Encrypted Data Stores
Paul Grubbs, Anurag Khandelwal, Marie-Sarah Lacharité, Lloyd Brown, Lucy Li, Rachit Agarwal 0001, Thomas Ristenpart |
USENIX Security Symposium | 2 |
| 2019 | Confluo: Distributed Monitoring and Diagnosis Stack for High-speed Networks
Anurag Khandelwal, Rachit Agarwal 0001, Ion Stoica |
NSDI | 1 |
| 2017 | ZipG: A Memory-efficient Graph Store for Interactive QueriesabstractWe present ZipG, a distributed memory-efficient graph store for serving interactive graph queries. ZipG achieves memory efficiency by storing the input graph data using a compressed representation. What differentiates ZipG from other graph stores is its ability to execute a wide range of graph queries directly on this compressed representation. ZipG can thus execute a larger fraction of queries in main memory, achieving query interactivity. ZipG exposes a minimal API that is functionally rich enough to implement published functionalities from several industrial graph stores. We demonstrate this by implementing and evaluating graph queries from Facebook TAO, LinkBench, Graph Search and several other workloads on top of ZipG. On a single server with 244GB memory, ZipG executes tens of thousands of queries from these workloads for raw graph data over half a TB; this leads to an order of magnitude (sometimes as much as 23×) higher throughput than Neo4j and Titan. We get similar gains in distributed settings compared to Titan. Anurag Khandelwal, Zongheng Yang, Evan Ye, Rachit Agarwal 0001, Ion Stoica |
SIGMOD Conference | 1 |
| 2016 | BlowFish: Dynamic Storage-Performance Tradeoff in Data Stores
Anurag Khandelwal, Rachit Agarwal 0001, Ion Stoica |
NSDI | 1 |
| 2015 | Succinct: Enabling Queries on Compressed Data
Rachit Agarwal 0001, Anurag Khandelwal, Ion Stoica |
NSDI | 2 |