VLDB 2026 Research / reviewers in the wild / expert
Aravinda Prasad
dblp:177/8610
· DBLP profile ↗
8ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0002-8558-2814ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TierScape: Harnessing Multiple Compressed Tiers to Tame Server Memory TCO
Aravinda Prasad, Sreenivas Subramoney |
EuroSys | 2 |
| 2025 | TierTrain: Proactive Memory Tiering for CPU-Based DNN TrainingabstractDeep neural networks (DNNs) are one of the popular models for learning relationships between complex data. Training a DNN model is a compute- and memory-intensive operation. The size of modern DNN models spans into the terabyte region, requiring multiple accelerators to train -- driving up the training cost. Such humongous memory requirements shift the focus toward memory rather than computation. Sathvik Swaminathan, Aravinda Prasad, Sreenivas Subramoney |
ISMM | 3 |
| 2024 | Telescope: Telemetry for Gargantuan Memory Footprint Applications
Alan Nair, Aravinda Prasad, Andy Rudoff, Sreenivas Subramoney |
USENIX ATC | 3 |
| 2021 | Radiant: efficient page table management for tiered memory systemsabstractModern enterprise servers are increasingly embracing tiered memory systems with a combination of low latency DRAMs and large capacity but high latency non-volatile main memories (NVMMs) such as Intel’s Optane DC PMM. Prior works have focused on the efficient placement and migration of data on a tiered memory system, but have not studied the optimal placement of page tables. Aravinda Prasad, Smruti R. Sarangi, Sreenivas Subramoney |
ISMM | 2 |
| 2018 | Making Huge Pages Actually UsefulabstractThe virtual-to-physical address translation overhead, a major performance bottleneck for modern workloads, can be effectively alleviated with huge pages. However, since huge pages must be mapped contiguously, OSs have not been able to use them well because of the memory fragmentation problem despite hardware support for huge pages being available for nearly two decades. This paper presents a comprehensive study of the interaction of fragmentation with huge pages in the Linux kernel. We observe that when huge pages are used, problems such as high CPU utilization and latency spikes occur because of unnecessary work (e.g., useless page migration) performed by memory management related subsystems due to the poor handling of unmovable (i.e., kernel) pages. This behavior is even more harmful in virtualized systems where unnecessary work may be performed in both guest and host OSs. We present Illuminator, an efficient memory manager that provides various subsystems, such as the page allocator, the ability to track all unmovable pages. It allows subsystems to make informed decisions and eliminate unnecessary work which in turn leads to cost-effective huge page allocations. Illuminator reduces the cost of compaction (up to 99%), improves application performance (up to 2.3x) and reduces the maximum latency of MySQL database server (by 30x). Importantly, this work shows the effectiveness of a simple solution for long-standing huge page related problems. Ashish Panwar, Aravinda Prasad, K. Gopinath |
ASPLOS | 2 |
| 2018 | A frugal approach to reduce RCU grace period overheadabstractGrace period computation is a core part of the Read-Copy-Update (RCU) synchronization technique that determines the safe time to reclaim the deferred objects' memory. We first show that the eager grace period computation employed in the Linux kernel is appropriate only for enterprise workloads such as web and database servers where a large amount of reclaimable memory awaits the completion of a grace period. However, such memory is negligible in High-Performance Computing (HPC) and mostly idling environments due to limited OS kernel activity. Hence an eager approach is not only futile but also detrimental as the CPU cycles consumed to compute a grace period leads to jitter in HPC and frequent CPU wake-ups in idle environments. Aravinda Prasad, K. Gopinath |
EuroSys | 1 |
| 2017 | The RCU-Reader Preemption Problem in VMs
Aravinda Prasad, K. Gopinath, Paul E. McKenney |
USENIX ATC | 1 |
| 2016 | Prudent Memory Reclamation in Procrastination-Based SynchronizationabstractProcrastination is the fundamental technique used in synchronization mechanisms such as Read-Copy-Update (RCU) where writers, in order to synchronize with readers, defer the freeing of an object until there are no readers referring to the object. The synchronization mechanism determines when the deferred object is safe to reclaim and when it is actually reclaimed. Hence, such memory reclamations are completely oblivious of the memory allocator state. This induces poor memory allocator performance, for instance, when the reclamations are ill-timed. Furthermore, deferred objects provide hints about the future that inform memory regions that are about to be freed. Although useful, hints are not exploited as deferred objects are not visible to memory allocators. We introduce Prudence, a dynamic memory allocator, that is tightly integrated with the synchronization mechanism to ensure visibility of deferred objects to the memory allocator. Such an integration enables Prudence to (i) identify the safe time to reclaim deferred objects' memory, (ii) have an inclusive view of the allocated, free and about-to-be-freed objects, and (iii) exploit optimizations based on the hints about the future during important state transitions. Our evaluation in the Linux kernel shows that Prudence integrated with RCU performs 3.9X to 28X better in micro-benchmarks compared to SLUB, a recent memory allocator in the Linux kernel. It also improves the overall performance perceptibly (4%-18%) for a mix of widely used synthetic and application benchmarks. Further, it performs better (up to 98%) in terms of object hits in caches, object cache churns, slab churns, peak memory usage and total fragmentation, when compared with the SLUB allocator. Aravinda Prasad, K. Gopinath |
ASPLOS | 1 |