Prasoon Sinha

dblp:163/8446 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-5538-8829ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
abstract
KV cache accelerates LLM inference by avoiding redundant computation, at the expense of memory. To support larger KV caches, prior work extends GPU memory with CPU memory via CPU-offloading. This involves swapping KV cache between GPU and CPU memory. However, because the cache updates dynamically, such swapping incurs high CPU memory traffic. We make a key observation that model parameters remain constant during runtime, unlike the dynamically updated KV cache. Building on this, we introduce Oneiros, which avoids KV cache swapping by remapping, and thereby repurposing, the memory allocated to model parameters for KV cache. This parameter remapping is especially beneficial in multi-tenant environments, where the memory used for the parameters of the inactive models can be more aggressively reclaimed. Exploiting the high CPU-GPU bandwidth offered by the modern hardware, such as the NVIDIA Grace Hopper Superchip, we show that Oneiros significantly outperforms state-of-the-art solutions, achieving a reduction of 44.8%-82.5% in tail time-between-token latency, 20.7%-99.3% in tail time-to-first-token latency, and 6.6%-86.7% higher throughput compared to vLLM. Source code of Oneiros is available at https://github.com/UT-SysML/Oneiros/.
Ruihao Li 0002, Shagnik Pal, Vineeth Narayan Pullu, Prasoon Sinha, Jeeho Ryoo, Lizy Kurian John, Neeraja J. Yadwadkar
SoCC4
2025 ALAP: Intent-Based Serverless Computing via Delayed Decision-Making
abstract
The lack of an interface that allows users to express their intent, to trade off latency versus cost, continues to be a hurdle in the adoption of serverless computing for latency-critical and cost-constrained applications. Existing systems shy away from providing a user-intent knob mainly because they make resource allocation decisions as soon as a function is registered, when inputs to the function are mostly unavailable. We find that function performance and resource utilization can vary greatly (up to 6× in latency and 4.78× in utilization) across inputs. Thus, to enable intent-based serverless computing, our key insight is to delay making resource allocation decisions until after function inputs are available so that resources can be allocated independently for each input and resource type. We introduce ALAP, a resource management framework for serverless systems that makes decisions As Late As Possible to right-size each invocation and meet user-intent efficiently. ALAP uses an online learning agent to predict the required amount of resources to meet an invocation's constraints and schedules these right-sized containers in a cold-start-aware manner. For a range of functions and inputs, ALAP adapts to user intents: it reduces latency and cost by up to 17.5× (3× on avg.) and 13× (3.7× on avg.), respectively, and latency and cost SLO violations by 1.3-2.3× and 1.6-2.2×, respectively, while nearly eliminating CPU and memory waste compared to four state-of-the-art systems.
Prasoon Sinha, Kostis Kaffes, Neeraja J. Yadwadkar
SoCC1
2025 MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM Serving
abstract
Large Language Models (LLMs) are becoming ubiquitous across industries, where applications demand they fulfill diverse user intents. However, developers currently face the challenge of manually exploring numerous deployment configurations—combinations of parallelism and compression techniques that impact resource usage, latency, cost, and accuracy—to meet these intents. Previous works automate configuration selection and deployment, but, they (a) rely on expensive profiling, and (b) suboptimally utilize the fragmented resource availability in multi-tenant GPU clusters, inflating operational costs for the provider. Moreover, none of these solutions tailors deployment configuration decisions to diverse user intents.
Dimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee, Neeraja J. Yadwadkar
SC2
2022 Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich Systems
abstract
Scientists are increasingly exploring and utilizing the massive parallelism of general-purpose accelerators such as GPUs for scientific breakthroughs. As a result, datacenters, hyperscalers, national computing centers, and supercomputers have procured hardware to support this evolving application paradigm. These systems contain hundreds to tens of thousands of accelerators, enabling peta- and exa-scale levels of compute for scientific workloads. Recent work demonstrated that power management (PM) can impact application performance in CPU-based HPC systems, even when machines have the same architecture and SKU (stock keeping unit). This variation occurs due to manufacturing variability and the chip's PM. However, while modern HPC systems widely employ accelerators such as GPUs, it is unclear how much this variability affects applications. Accordingly, we seek to characterize the extent of variation due to GPU PM in modern HPC and supercomputing systems. We study a variety of applications that stress different GPU components on five large-scale computing centers with modern GPUs: Oak Ridge's Summit, Sandia's Vortex, TACC's Frontera and Longhorn, and Livermore's Corona. These clusters use a variety of cooling methods and GPU vendors. In total, we collect over 18,800 hours of data across more than 90% of the GPUs in these clusters. Regardless of the application, cluster, GPU vendor, and cooling method, our results show significant variation: 8% (max 22%) average performance variation even though the GPU architecture and vendor SKU are identical within each cluster, with outliers up to 1.5× slower than the median GPU. These results highlight the difficulty in efficiently using existing GPU clusters for modern HPC and scientific workloads, and the need to embrace variability in future accelerator-based systems.
Prasoon Sinha, Akhil Guliani, Rutwik Jain, Brandon Tran, Matthew D. Sinclair, Shivaram Venkataraman
SC1