Ramya Prabhu

dblp:356/4824 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0002-7266-5522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 67% GPUs and heterogeneous computing · 33%
Artificial intelligence
2 papers
Language models and text generation · 56% Efficient and distributed learning · 44%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model inference
0.912025
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference · ASPLOS (2) 2025
Machine learning › Efficient and distributed learning › inference serving
large language model serving
0.912025
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025
Memory systems › memory management › memory allocation
dynamic memory allocation
0.912025
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025
GPUs and heterogeneous computing
GPU memory management
0.912025
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025
Memory systems › cache management
KV cache management
0.912025
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention · ASPLOS (1) 2025
Natural language and speech › Language models and text generation
large language model
0.312025
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference · ASPLOS (2) 2025

Methods — techniques the papers use, named apart from their topics

virtual memory management · 2.6attention · 0.9
YearPublicationVenuePosition
2025 POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 0001, Ramachandran Ramjee, Ashish Panwar
ASPLOS (2)2
2025 vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
abstract
PagedAttention is a popular approach for dynamic memory allocation in LLM serving systems. It enables on-demand allocation of GPU memory to mitigate KV cache fragmentation - a phenomenon that crippled the batch size (and consequently throughput) in prior systems. However, in trying to allocate physical memory at runtime, PagedAttention ends up changing the virtual memory layout of the KV cache from contiguous to non-contiguous. Such a design leads to non-trivial programming and performance overheads.
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar
ASPLOS (1)1