Jayashree Mohan

dblp:168/3415 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0005-5260-3203ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 QoServe: Breaking the Silos of LLM Inference Serving
Kanishk Goel, Jayashree Mohan, Nipun Kwatra, Ravi Shreyas Anupindi, Ramachandran Ramjee
ASPLOS (2)2
2025 POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 0001, Ramachandran Ramjee, Ashish Panwar
ASPLOS (2)3
2025 vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
abstract
PagedAttention is a popular approach for dynamic memory allocation in LLM serving systems. It enables on-demand allocation of GPU memory to mitigate KV cache fragmentation - a phenomenon that crippled the batch size (and consequently throughput) in prior systems. However, in trying to allocate physical memory at runtime, PagedAttention ends up changing the virtual memory layout of the KV cache from contiguous to non-contiguous. Such a design leads to non-trivial programming and performance overheads.
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar
ASPLOS (1)3
2025 ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
abstract
Large multimodal models (LMMs) demonstrate impressive capabilities in understanding images, videos, and audio beyond text. However, efficiently serving LMMs in production environments poses significant challenges due to their complex model architectures and heterogeneous characteristics across their multi-stage inference pipelines and modalities.
Haoran Qiu, Anish Biswas, Jayashree Mohan, Alind Khare, Esha Choukse, Íñigo Goiri, Zeyu Zhang 0005, Haiying Shen, Chetan Bansal, Ramachandran Ramjee, Rodrigo Fonseca
SoCC4
2025 Project Silica: Towards Sustainable Cloud Archival Storage in Glass
abstract
Sustainable and cost-effective long-term storage remains an unsolved problem. The most widely used storage technologies today are magnetic (hard disk drives and tape). They use media that degrades over time and has a limited lifetime, which leads to inefficient, wasteful, and costly solutions for long-lived data. This article presents Silica: the first cloud storage system for archival data underpinned by quartz glass, an extremely resilient media that allows data to be left in situ indefinitely. The hardware and software of Silica have been co-designed and co-optimized from the media up to the service level with sustainability as a primary objective. The design follows a cloud-first, data-driven methodology underpinned by principles derived from analyzing the archival workload of a large public cloud service. Silica can support a wide range of archival storage workloads and ushers in a new era of sustainable, cost-effective storage.
Patrick Anderson 0001, Erika Blancada Aranas, Youssef Assaf, Raphael Behrendt, Richard Black, Marco Caballero, Pashmina Cameron, Burcu Canakci, Andromachi Chatzieleftheriou, Rebekah Storan Clarke, James Clegg, Daniel Cletheroe, Bridgette Cooper, Thales De Carvalho, Tim Deegan, Austin Donnelly, Rokas Drevinskas, Alexander L. Gaunt, Christos Gkantsidis, Ariel Gomez Diaz, István Haller, Freddie Hong, Teodora Ilieva, Shashidhar Joshi, Russell Joyce, Mint Kunkel, David Lara Alabazares, Sergey Legtchenko, Fanglin Linda Liu, Bruno Magalhães, Alana Marzoev, Marvin McNett, Jayashree Mohan, Michael Myrah, Sebastian Nowozin, Aaron Ogus, Hiske Overweg, Antony I. T. Rowstron, Maneesh Sah, Masaaki Sakakura, Peter Scholtz, Nina Schreiner, Omer Sella, Ioan A. Stefanovici, David Sweeney, Benn C. Thomsen, Govert Verkes, Phil Wainman, Jonathan Westcott, Luke Weston, Charles Whittaker, Pablo Wilke Berenguer, Hugh Williams, Stefan Winzeck
ACM Trans. Storage33
2024 Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, Ramachandran Ramjee
OSDI4
2023 Project Silica: Towards Sustainable Cloud Archival Storage in Glass
abstract
Sustainable and cost-effective long-term storage remains an unsolved problem. The most widely used storage technologies today are magnetic (hard disk drives and tape). They use media that degrades over time and has a limited lifetime, which leads to inefficient, wasteful, and costly solutions for long-lived data. This paper presents Silica: the first cloud storage system for archival data underpinned by quartz glass, an extremely resilient media that allows data to be left in situ indefinitely. The hardware and software of Silica have been co-designed and co-optimized from the media up to the service level with sustainability as a primary objective. The design follows a cloud-first, data-driven methodology underpinned by principles derived from analyzing the archival workload of a large public cloud service. Silica can support a wide range of archival storage workloads and ushers in a new era of sustainable, cost-effective storage.
Patrick Anderson 0001, Erika Blancada Aranas, Youssef Assaf, Raphael Behrendt, Richard Black, Marco Caballero, Pashmina Cameron, Burcu Canakci, Thales De Carvalho, Andromachi Chatzieleftheriou, Rebekah Storan Clarke, James Clegg, Daniel Cletheroe, Bridgette Cooper, Tim Deegan, Austin Donnelly, Rokas Drevinskas, Alexander L. Gaunt, Christos Gkantsidis, Ariel Gomez Diaz, István Haller, Freddie Hong, Teodora Ilieva, Shashidhar Joshi, Russell Joyce, Mint Kunkel, David Lara Alabazares, Sergey Legtchenko, Fanglin Linda Liu, Bruno Magalhães, Alana Marzoev, Marvin McNett, Jayashree Mohan, Michael Myrah, Sebastian Nowozin, Aaron Ogus, Hiske Overweg, Antony I. T. Rowstron, Maneesh Sah, Masaaki Sakakura, Peter Scholtz, Nina Schreiner, Omer Sella, Ioan A. Stefanovici, David Sweeney, Benn C. Thomsen, Govert Verkes, Phil Wainman, Jonathan Westcott, Luke Weston, Charles Whittaker, Pablo Wilke Berenguer, Hugh Williams, Stefan Winzeck
SOSP33
2022 Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters
Jayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay Chidambaram
OSDI1
2021 CheckFreq: Frequent, Fine-Grained DNN Checkpointing
Jayashree Mohan, Amar Phanishayee, Vijay Chidambaram
FAST1
2021 Memory Optimization for Deep Networks
Aashaka Shah, Chao-Yuan Wu, Jayashree Mohan, Vijay Chidambaram, Philipp Krähenbühl
ICLR3
2021 Analyzing and Mitigating Data Stalls in DNN Training
abstract
Training Deep Neural Networks (DNNs) is resource-intensive and time-consuming. While prior research has explored many different ways of reducing DNN training time, the impact of input data pipeline , i.e., fetching raw data items from storage and performing data pre-processing in memory, has been relatively unexplored. This paper makes the following contributions: (1) We present the first comprehensive analysis of how the input data pipeline affects the training time of widely-used computer vision and audio Deep Neural Networks (DNNs), that typically involve complex data pre-processing. We analyze nine different models across three tasks and four datasets while varying factors such as the amount of memory, number of CPU threads, storage device, GPU generation etc on servers that are a part of a large production cluster at Microsoft. We find that in many cases, DNN training time is dominated by data stall time : time spent waiting for data to be fetched and pre-processed. (2) We build a tool, DS-Analyzer to precisely measure data stalls using a differential technique, and perform predictive what-if analysis on data stalls. (3) Finally, based on the insights from our analysis, we design and implement three simple but effective techniques in a data-loading library, CoorDL, to mitigate data stalls. Our experiments on a range of DNN tasks, models, datasets, and hardware configs show that when PyTorch uses CoorDL instead of the state-of-the-art DALI data loading library, DNN training time is reduced significantly (by as much as 5X on a single server).
Jayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay Chidambaram
Proc. VLDB Endow.1
2020 INSTalytics: Cluster Filesystem Co-design for Big-data Analytics
abstract
We present the design, implementation, and evaluation of INSTalytics , a co-designed stack of a cluster file system and the compute layer, for efficient big-data analytics in large-scale data centers. INSTalytics amplifies the well-known benefits of data partitioning in analytics systems; instead of traditional partitioning on one dimension, INSTalytics enables data to be simultaneously partitioned on four different dimensions at the same storage cost, enabling a larger fraction of queries to benefit from partition filtering and joins without network shuffle. To achieve this, INSTalytics uses compute-awareness to customize the three-way replication that the cluster file system employs for availability. A new heterogeneous replication layout enables INSTalytics to preserve the same recovery cost and availability as traditional replication. INSTalytics also uses compute-awareness to expose a new sliced-read API that improves performance of joins by enabling multiple compute nodes to read slices of a data block efficiently via co-ordinated request scheduling and selective caching at the storage nodes. We have built a prototype implementation of INSTalytics in a production analytics stack, and we show that recovery performance and availability is similar to physical replication, while providing significant improvements in query performance, suggesting a new approach to designing cloud-scale big-data analytics systems.
Muthian Sivathanu, Midhul Vuppalapati, Bhargav S. Gulavani, Kaushik Rajan, Jyoti Leeka, Jayashree Mohan, Piyus Kedia
ACM Trans. Storage6
2019 INSTalytics: Cluster Filesystem Co-design for Big-data Analytics
Muthian Sivathanu, Midhul Vuppalapati, Bhargav S. Gulavani, Kaushik Rajan, Jyoti Leeka, Jayashree Mohan, Piyus Kedia
FAST6
2019 Recipe: converting concurrent DRAM indexes to persistent-memory indexes
abstract
We present Recipe, a principled approach for converting concurrent DRAM indexes into crash-consistent indexes for persistent memory (PM). The main insight behind Recipe is that isolation provided by a certain class of concurrent in-memory indexes can be translated with small changes to crash-consistency when the same index is used in PM. We present a set of conditions that enable the identification of this class of DRAM indexes, and the actions to be taken to convert each index to be persistent. Based on these conditions and conversion actions, we modify five different DRAM indexes based on B+ trees, tries, radix trees, and hash tables to their crash-consistent PM counterparts. The effort involved in this conversion is minimal, requiring 30--200 lines of code. We evaluated the converted PM indexes on Intel DC Persistent Memory, and found that they outperform state-of-the-art, hand-crafted PM indexes in multi-threaded workloads by up-to 5.2x. For example, we built P-CLHT, our PM implementation of the CLHT hash table by modifying only 30 LOC. When running YCSB workloads, P-CLHT performs up to 2.4x better than Cacheline-Conscious Extendible Hashing (CCEH), the state-of-the-art PM hash table.
Se Kwon Lee, Jayashree Mohan, Sanidhya Kashyap, Taesoo Kim, Vijay Chidambaram
SOSP2
2019 CrashMonkey and ACE: Systematically Testing File-System Crash Consistency
abstract
We present C rash M onkey and A ce , a set of tools to systematically find crash-consistency bugs in Linux file systems. C rash M onkey is a record-and-replay framework which tests a given workload on the target file system by simulating power-loss crashes while the workload is being executed, and checking if the file system recovers to a correct state after each crash. A ce automatically generates all the workloads to be run on the target file system. We build C rash M onkey and A ce based on a new approach to test file-system crash consistency: bounded black-box crash testing ( B 3 ). B 3 tests the file system in a black-box manner using workloads of file-system operations. Since the space of possible workloads is infinite, B 3 bounds this space based on parameters such as the number of file-system operations or which operations to include, and exhaustively generates workloads within this bounded space. B 3 builds upon insights derived from our study of crash-consistency bugs reported in Linux file systems in the last 5 years. We observed that most reported bugs can be reproduced using small workloads of three or fewer file-system operations on a newly created file system, and that all reported bugs result from crashes after fsync()-related system calls. C rash M onkey and A ce are able to find 24 out of the 26 crash-consistency bugs reported in the last 5 years. Our tools also revealed 10 new crash-consistency bugs in widely used, mature Linux file systems, 7 of which existed in the kernel since 2014. Additionally, our tools found a crash-consistency bug in a verified file system, FSCQ. The new bugs result in severe consequences like broken rename atomicity, loss of persisted files and directories, and data loss.
Jayashree Mohan, Ashlie Martinez, Soujanya Ponnapalli, Pandian Raju, Vijay Chidambaram
ACM Trans. Storage1
2018 Finding Crash-Consistency Bugs with Bounded Black-Box Crash Testing
Jayashree Mohan, Ashlie Martinez, Soujanya Ponnapalli, Pandian Raju, Vijay Chidambaram
OSDI1
2017 Storage on Your SmartPhone Uses More Energy Than You Think
Jayashree Mohan, Dhathri Purohith, Matthew Halpern, Vijay Chidambaram, Vijay Janapa Reddi
HotStorage1
2016 Optimizing Downloads over Random Duration Links in Mobile Networks
abstract
Short range vehicle to vehicle and device to device communications are of growing interest due to their utility for vehicular safety and infotainment applications as well as for improving the capacity of cellular networks. These mobile systems are characterized by ephemeral, stochastic links. We consider a fundamental problem in this domain -- how to maximize the amount of useful content downloaded by a client from a server over an encounter that lasts a random amount of time. We assume that the distribution of link duration is known or estimated \emph{a priori} based on historical as well as real-time measurements. We present MERLIN (Maximum Expected download over Random LINks), a single-phase file request protocol that is provably optimal. We evaluate MERLIN comprehensively via simulations based on both ideal link duration distributions as well as empirical distributions obtained from real vehicular mobility traces (from Taxis in Shanghai and Buses in Chicago). We also present two Contiki OS-based implementations of MERLIN (with local and remote calculations) evaluated on the Tmote Sky wireless embedded platform.
Amber Bhargava, Spencer Congero, Timothy Ferrell, Leo Linsky, Jayashree Mohan, Bhaskar Krishnamachari
ICCCN6