Dushyanth Narayanan

dblp:45/2517 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
3since 2021 · last 2025
0009-0000-1194-2958ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 2 since 2021Computer networks · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-author
YearPublicationVenuePosition
2025 Good things come in small packages: Should we build AI clusters with Lite-GPUs?
abstract
To match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages. However, there is growing uncertainty about the scalability of individual GPUs and thus AI clusters, as state-of-the-art GPUs are already displaying packaging, yield, and cooling limitations. We propose to rethink the design and scaling of AI clusters through efficiently-connected large clusters of Lite-GPUs, GPUs with single, small dies and a fraction of the capabilities of larger GPUs. We think recent advances in co-packaged optics can enable distributing AI workloads onto many Lite-GPUs through high bandwidth and efficient communication. In this paper, we present the key benefits of Lite-GPUs on manufacturing cost, blast radius, yield, and power efficiency; and discuss systems opportunities and challenges around resource, workload, memory, and network management.
Burcu Canakci, Xingbo Wu, Nathanael Cheriere, Paolo Costa, Sergey Legtchenko, Dushyanth Narayanan, Antony I. T. Rowstron
HotOS7
2025 Storage Class Memory is Dead, All Hail Managed-Retention Memory: Rethinking Memory for the AI Era
abstract
AI clusters today are one of the major uses of High Bandwidth Memory (HBM). However, HBM is suboptimal for AI workloads for several reasons. Analysis shows HBM is overprovisioned on write performance, but underprovisioned on density and read bandwidth, and also has significant energy per bit overheads. It is also expensive, with lower yield than DRAM due to manufacturing complexity. We propose a new memory class: Managed-Retention Memory (MRM), which is more optimized to store key data structures for AI inference workloads. We believe that MRM may finally provide a path to viability for technologies that were originally proposed to support Storage Class Memory (SCM). These technologies traditionally offered long-term persistence (10+ years) but provided poor IO performance and/or endurance. MRM makes different trade-offs, and by understanding the workload IO patterns, MRM foregoes long-term data retention and write performance for better potential performance on the metrics important for these workloads.
Sergey Legtchenko, Ioan A. Stefanovici, Richard Black, Antony I. T. Rowstron, Paolo Costa, Burcu Canakci, Dushyanth Narayanan, Xingbo Wu
HotOS8
2025 Holographic Storage for the Cloud: advances and challenges
abstract
Holographic Storage is an old idea that has always promised high density and fast random access, but has never been commercially competitive with Hard Disk Drives (HDDs) and Solid State Devices (SSDs). In Project HSD at Microsoft Research we asked the question: “Does holographic storage finally make sense for cloud storage?” This article describes our journey toward answering this question. We achieved 1.8× higher density than the previous state-of-the-art, using commodity components available today and leveraging machine learning to compensate for the noise and distortions introduced by commodity components. This uncovered two new challenges which are the focus of this article: achieving high end-to-end energy efficiency without sacrificing capacity, and spatial multiplexing without mechanical movement. Improving end-to-end energy efficiency requires joint optimization across low-level media parameters and higher-level system parameters that govern background maintenance operations such as read refresh and garbage collection. We developed new physics models of the media; analytic and simulation models of the media access and background media maintenance; and workload-driven optimization to find optimal parameter combinations. These techniques resulted in a 14× improvement over the previous approach for typical workloads without sacrificing capacity. We also designed the first scalable and mechanical movement free spatial multiplexing system for holographic storage. Despite these advances, we conclude that currently, holographic storage is still far from the combination of density, capacity scaling, and energy efficiency needed to compete with the incumbent technologies. We need fundamental advances in the physical media that improve energy efficiency by another 1–2 orders of magnitude without reducing data density. Further advances in optics are also required to achieve spatial multiplexing that is simultaneously scalable, low-loss, and high-density.
Nathanael Cheriere, Jiaqi Chu, Grace Brennan, Pashmina Cameron, Pedro Da Costa, Jannes Gladrow, Guilherme Ilunga, Douglas J. Kelly, Joowon Lim, Giorgio Maltese, Tony Mason, Greg O'Shea, Soujanya Ponnapalli, Michael Rudow, Alan Sanders, Theano Stavrinos, Xingbo Wu, Mengyang Yang, Dushyanth Narayanan, Benn C. Thomsen, Antony I. T. Rowstron
ACM Trans. Storage20
2020 Could cloud storage be disrupted in the next decade?
Andromachi Chatzieleftheriou, Ioan A. Stefanovici, Dushyanth Narayanan, Benn C. Thomsen, Antony I. T. Rowstron
HotStorage3
2019 Fast General Distributed Transactions with Opacity
abstract
Transactions can simplify distributed applications by hiding data distribution, concurrency, and failures from the application developer. Ideally the developer would see the abstraction of a single large machine that runs transactions sequentially and never fails. This requires the transactional subsystem to provide opacity (strict serializability for both committed and aborted transactions), as well as transparent fault tolerance with high availability. As even the best abstractions are unlikely to be used if they perform poorly, the system must also provide high performance. Existing distributed transactional designs either weaken this abstraction or are not designed for the best performance within a data center. This paper extends the design of FaRM --- which provides strict serializability only for committed transactions --- to provide opacity while maintaining FaRM's high throughput, low latency, and high availability within a modern data center. It uses timestamp ordering based on real time with clocks synchronized to within tens of microseconds across a cluster, and a failover protocol to ensure correctness across clock master failures. FaRM with opacity can commit 5.4 million neworder transactions per second when running the TPC-C transaction mix on 90 machines with 3-way replication.
Alex Shamis, Matthew Renzelmann, Stanko Novakovic, Georgios Chatzopoulos, Aleksandar Dragojevic, Dushyanth Narayanan, Miguel Castro 0001
SIGMOD Conference6
2016 Filo: Consolidated Consensus as a Cloud Service
Parisa Jalili Marandi, Christos Gkantsidis, Flavio Paiva Junqueira, Dushyanth Narayanan
USENIX ATC4
2015 No compromises: distributed transactions with consistency, availability, and performance
abstract
Transactions with strong consistency and high availability simplify building and reasoning about distributed systems. However, previous implementations performed poorly. This forced system designers to avoid transactions completely, to weaken consistency guarantees, or to provide single-machine transactions that require programmers to partition their data. In this paper, we show that there is no need to compromise in modern data centers. We show that a main memory distributed computing platform called FaRM can provide distributed transactions with strict serializability, high performance, durability, and high availability. FaRM achieves a peak throughput of 140 million TATP transactions per second on 90 machines with a 4.9 TB database, and it recovers from a failure in less than 50 ms. Key to achieving these results was the design of new transaction, replication, and recovery protocols from first principles to leverage commodity networks with RDMA and a new, inexpensive approach to providing non-volatile DRAM.
Aleksandar Dragojevic, Dushyanth Narayanan, Ed Nightingale, Matthew Renzelmann, Alex Shamis, Anirudh Badam, Miguel Castro 0001
SOSP2
2014 FaRM: Fast Remote Memory
Aleksandar Dragojevic, Dushyanth Narayanan, Miguel Castro 0001, Orion Hodson
NSDI2
2013 Scale-up vs scale-out for Hadoop: time to rethink?
abstract
In the last decade we have seen a huge deployment of cheap clusters to run data analytics workloads. The conventional wisdom in industry and academia is that scaling out using a cluster of commodity machines is better for these workloads than scaling up by adding more resources to a single server. Popular analytics infrastructures such as Hadoop are aimed at such a cluster scale-out environment.
Raja Appuswamy, Christos Gkantsidis, Dushyanth Narayanan, Orion Hodson, Antony I. T. Rowstron
SoCC3
2013 Rhea: Automatic Filtering for Unstructured Cloud Storage
Christos Gkantsidis, Dimitrios Vytiniotis, Orion Hodson, Dushyanth Narayanan, Florin Dinu, Antony I. T. Rowstron
NSDI4
2012 Whole-system persistence
abstract
Today's databases and key-value stores commonly keep all their data in main memory. A single server can have over 100 GB of memory, and a cluster of such servers can have 10s to 100s of TB. However, a storage back end is still required for recovery from failures. Recovery can last for minutes for a single server or hours for a whole cluster, causing heavy load on the back end. Non-volatile main memory (NVRAM) technologies can help by allowing near-instantaneous recovery of in-memory state. However, today's software does not support this well. Block-based approaches such as persistent buffer caches suffer from data duplication and block transfer overheads. Recently, user-level persistent heaps have been shown to have much better performance than these. However they require substantial application modification and still have significant runtime overheads. This paper proposes whole-system persistence (WSP) as an alternative. WSP is aimed at systems where all memory is non-volatile. It transparently recovers an application's entire state, making a failure appear as a suspend/resume event. Runtime overheads are eliminated by using "flush on fail": transient state in processor registers and caches is flushed to NVRAM only on failure, using the residual energy from the system power supply. Our evaluation shows that this approach has 1.6--13 times better runtime performance than a persistent heap, and that flush-on-fail can complete safely within 2--35\% of the residual energy window provided by standard power supplies.
Dushyanth Narayanan, Orion Hodson
ASPLOS1
2011 Sierra: practical power-proportionality for data center storage
abstract
Online services hosted in data centers show significant diurnal variation in load levels. Thus, there is significant potential for saving power by powering down excess servers during the troughs. However, while techniques like VM migration can consolidate computational load, storage state has always been the elephant in the room preventing this powering down. Migrating storage is not a practical way to consolidate I/O load.
Eno Thereska, Austin Donnelly, Dushyanth Narayanan
EuroSys3
2010 Hermes: clustering users in large-scale e-mail services
abstract
Hermes is an optimization engine for large-scale enterprise e-mail services. Such services could be hosted by a virtualized e-mail service provider, or by dedicated enterprise data centers. In both cases we observe that the pattern of e-mails between employees of an enterprise forms an implicit social graph. Hermes tracks this implicit social graph, periodically identifies clusters of strongly connected users within the graph, and co-locates such users on the same server. Co-locating the users reduces storage requirements: senders and recipients who reside on the same server can share a single copy of an e-mail. Co-location also reduces inter-server bandwidth usage.
Thomas Karagiannis, Christos Gkantsidis, Dushyanth Narayanan, Antony I. T. Rowstron
SoCC3
2009 Migrating server storage to SSDs: analysis of tradeoffs
abstract
Recently, flash-based solid-state drives (SSDs) have become standard options for laptop and desktop storage, but their impact on enterprise server storage has not been studied. Provisioning server storage is challenging. It requires optimizing for the performance, capacity, power and reliability needs of the expected workload, all while minimizing financial costs. In this paper we analyze a number of workload traces from servers in both large and small data centers, to decide whether and how SSDs should be used to support each. We analyze both complete replacement of disks by SSDs, as well as use of SSDs as an intermediate tier between disks and DRAM. We describe an automated tool that, given device models and a block-level trace of a workload, determines the least-cost storage configuration that will support the workload's performance, capacity, and fault-tolerance requirements. We found that replacing disks by SSDs is not a costeffective option for any of our workloads, due to the low capacity per dollar of SSDs. Depending on the workload, the capacity per dollar of SSDs needs to increase by a factor of 3-3000 for an SSD-based solution to break even with a diskbased solution. Thus, without a large increase in SSD capacity per dollar, only the smallest volumes, such as system boot volumes, can be cost-effectively migrated to SSDs. The benefit of using SSDs as an intermediate caching tier is also limited: fewer than 10% of our workloads can reduce provisioning costs by using an SSD tier at today's capacity per dollar, and fewer than 20% can do so at any SSD capacity per dollar. Although SSDs are much more energy-efficient than enterprise disks, the energy savings are outweighed by the hardware costs, and comparable energy savings are achievable with low-power SATA disks.
Dushyanth Narayanan, Eno Thereska, Austin Donnelly, Sameh Elnikety, Antony I. T. Rowstron
EuroSys1
2008 Write Off-Loading: Practical Power Management for Enterprise Storage
Dushyanth Narayanan, Austin Donnelly, Antony I. T. Rowstron
FAST1
2008 Everest: Scaling Down Peak Loads Through I/O Off-Loading
Dushyanth Narayanan, Austin Donnelly, Eno Thereska, Sameh Elnikety, Antony I. T. Rowstron
OSDI1
2008 Write off-loading: Practical power management for enterprise storage
abstract
In enterprise data centers power usage is a problem impacting server density and the total cost of ownership. Storage uses a significant fraction of the power budget and there are no widely deployed power-saving solutions for enterprise storage systems. The traditional view is that enterprise workloads make spinning disks down ineffective because idle periods are too short. We analyzed block-level traces from 36 volumes in an enterprise data center for one week and concluded that significant idle periods exist, and that they can be further increased by modifying the read/write patterns using write off-loading . Write off-loading allows write requests on spun-down disks to be temporarily redirected to persistent storage elsewhere in the data center. The key challenge is doing this transparently and efficiently at the block level, without sacrificing consistency or failure resilience. We describe our write off-loading design and implementation that achieves these goals. We evaluate it by replaying portions of our traces on a rack-based testbed. Results show that just spinning disks down when idle saves 28--36% of energy, and write off-loading further increases the savings to 45--60%.
Dushyanth Narayanan, Austin Donnelly, Antony I. T. Rowstron
ACM Trans. Storage1
2008 Delay aware querying with Seaweed
Dushyanth Narayanan, Austin Donnelly, Richard Mortier, Antony I. T. Rowstron
VLDB J.1
2006 Delay Aware Querying with Seaweed
Dushyanth Narayanan, Austin Donnelly, Richard Mortier, Antony I. T. Rowstron
VLDB1
2005 Continuous resource monitoring for self-predicting DBMS
abstract
Administration tasks increasingly dominate the total cost of ownership of database management systems. A key task, and a very difficult one for an administrator, is to justify upgrades of CPU, memory and storage resources with quantitative predictions of the expected improvement in workload performance. Current database systems are not designed with such prediction in mind and hence offer only limited help to the administrator. This paper proposes changes to database system design that enable a Resource Advisor to answer "what-if" questions about resource upgrades. A prototype Resource Advisor built to work with a commercial DBMS shows the efficacy of our approach in predicting the effect of upgrading a key resource-buffer pool size-on OLTP workloads in a highly concurrent system.
Dushyanth Narayanan, Eno Thereska, Anastasia Ailamaki
MASCOTS1
2003 Magpie: Online Modelling and Performance-aware Systems
Paul Barham 0001, Rebecca Isaacs, Richard Mortier, Dushyanth Narayanan
HotOS4
2003 Predictive Resource Management for Wearable Computing
abstract
Achieving crisp interactive response in resource-intensive applications such as augmented reality, language translation, and speech recognition is a major challenge on resource-poor wearable hardware. In this paper we describe a solution based on multi-fidelity computation supported by predictive resource management. We show that such an approach can substantially reduce both the mean and the variance of response time. On a benchmark representative of augmented reality, we demonstrate a 60% reduction in mean latency and a 30% reduction in the coefficient of variation. We also show that a history-based approach to demand prediction is the key to this performance improvement: by applying simple machine learning techniques to logs of measured resource demand, we are able to accurately model resource demand as a function of fidelity.
Dushyanth Narayanan, Mahadev Satyanarayanan
MobiSys1
2001 Self-Tuned Remote Execution for Pervasive Computing
abstract
Pervasive computing creates environments saturated with computing and communication capability, yet gracefully integrated with human users. Remote execution has a natural role to play, in such environments, since it lets applications simultaneously leverage the mobility of small devices and the greater resources of large devices. In this paper, we describe Spectra, a remote execution system designed for pervasive environments. Spectra monitors resources such as battery, energy and file cache state which are especially important for mobile clients. It also dynamically balances energy use and quality goals with traditional performance concerns to decide where to locate functionality. Finally, Spectra is self-tuning-it does not require applications to explicitly specify intended resource usage. Instead, it monitors application behavior, learns functions predicting their resource usage, and uses the information to anticipate future behavior.
Jason Flinn, Dushyanth Narayanan, Mahadev Satyanarayanan
HotOS2
2001 Hinting for Goodness' Sake
abstract
Modern operating systems and adaptive applications offer an overwhelming number of parameters affecting application latency, throughput, image resolution, audio quality, and so on. We are designing a system to automatically tune resource allocation and application parameters at runtime, with the aim of maximizing user happiness or goodness. Consider a 3D graphics application that operates at variable resolution, trading output fidelity for processor time. Simultaneously, a data mining application adapts to network and processor load by migrating computation between the client and storage node. We must allocate resources between these applications and select their adaptive parameters to meet the user's overall goals. Since the user lacks the time and expertise to translate his preferences into parameter values, we would like the system to do this. Existing systems lack the right abstractions for applications to expose information for automated parameter tuning. Goodness hints are the solution to this problem. Applications use these hints to tell the operating system how resource allocations will affect their goodness (utility). Goodness hints are used by the operating system to make resource allocation decisions and by applications to tune their adaptive parameters. Our contribution is a decomposition of goodness hints into manageable and independent pieces and a methodology to automatically generate them.
David Petrou, Dushyanth Narayanan, Gregory R. Ganger, Garth A. Gibson, Elizabeth A. M. Shriver
HotOS2
2001 Multi-Fidelity Algorithms for Interactive Mobile Applications
Mahadev Satyanarayanan, Dushyanth Narayanan
Wirel. Networks2
1997 Agile Application-Aware Adaptation for Mobility
abstract
In this paper we show that application-aware adaptation, a collaborative partnership between the operating system and applications, offers the most general and effective approach to mobile information access.We describe the design of Odyssey, a prototype implementing this approach, and show how it supports concurrent execution of diverse mobile applications.We identify agility as a key attribute of adaptive systems, and describe how to quantify and measure it.We present the results of our evaluation of Odyssey, indicating performance improvements up to a factor of 5 on a benchmark of three applications concurrently using remote services over a network with highly variable bandwidth.This research was supported by the
Brian D. Noble, Mahadev Satyanarayanan, Dushyanth Narayanan, J. Eric Tilton, Jason Flinn, Kevin R. Walker
SOSP3