VLDB 2026 Research / reviewers in the wild / expert
Ryan Stutsman
dblp:02/1650
· DBLP profile ↗
46ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0001-5446-9603ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 15 · 1 first-author · 8 since 2021Computer networks · 5 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Remote Memory Ordering for Non-Coherent SystemsabstractSoftware using non-coherent interconnects like PCI Express requires fine-grained memory ordering, but current hardware mandates the use of costly source-side serialization. We show that this architectural mismatch severely limits the performance of two critical applications: (1) the transmission of network packets from a CPU to a NIC (requiring write-to-write ordering) and (2) key-value store lookups by an RDMA-enabled NIC (requiring read-to-read ordering). We address this by proposing a new destination-based ordering model and the hardware-software co-design comprising PCIe extensions and ISA extensions that allow software to express ordering intent efficiently. Novel microarchitecture at the Root Complex enforces these expressed semantics, eliminating source-side stalls. Our approach significantly improves the throughput of these application kernels and enables new, simpler protocols that outperform the state-of-the-art. Wei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil 0002, Ryan Stutsman, Vijay Nagarajan |
ASPLOS (2) | 4 |
| 2026 | Consistency and Coherence of the NVIDIA Grace-Hopper SuperchipabstractModern heterogeneous processors like the NVIDIA Grace-Hopper Superchip tightly integrate CPU and GPU cores across a cache-coherent interconnect, with an implicit assumption that independently compiled CPU and GPU code can safely interact via shared memory. Yet the memory consistency and coherence of such systems remain empirically unvalidated. This paper presents the first systematic study of consistency and coherence on the Grace-Hopper. We empirically validate that the system enforces the Compound Memory Consistency Model (CMCM)---a theoretical prerequisite for correct independent compilation---using a novel heterogeneous litmus testing methodology spanning 1,960 test variants. We further introduce Value Propagation tests to reverse-engineer the underlying coherence mechanisms, revealing that internal GPU coherence relies on write-throughs and self-invalidations rather than classical writer-initiated invalidations, while global CPU-GPU coherence is maintained via directory-based invalidations consistent with an AMBA CHI-like protocol. These results establish the CMCM as a concrete architectural target for heterogeneous systems and provide the first empirical characterization of GPU and CPU-GPU coherence mechanisms in a commercial heterogeneous processor. Soham Bagchi, Sanya Srivastava, Reese Levine, Tyler Sorensen 0001, Ryan Stutsman, Vijay Nagarajan |
ISMM | 5 |
| 2026 | DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM ServingabstractThe rapid adoption of large language models (LLMs) has increased the need for efficient multi-tenant inference systems that maximize GPU utilization. However, existing frameworks struggle to scale due to the high memory demands of model weights and key-value (KV) caches. We present DynamoServe, a multi-tenant LLM serving framework that addresses these challenges through three key innovations: (1) leveraging stranded GPU memory to offload model weights and KV caches, (2) mitigating resource fragmentation in multi-workload environments, and (3) improving memory locality through coordinated data placement and demand-driven weight migration across GPUs. Together, these techniques enable high-throughput, low-latency inference. Experiments on state-of-the-art models show that DynamoServe significantly improves memory efficiency without sacrificing latency. Diman Zad Tootaghaj, Khaled Diab 0001, Bob Lantz, Hanjiang Wu, K. K. Ramakrishnan, Md Ashfaqur Rahaman, Ryan Stutsman, Puneet Sharma 0001, Tushar Krishna |
SIGCOMM | 7 |
| 2025 | Water Footprint of Datacenter Applications: Methodological Implications of Manufacturing, Operational, and Decommissioning PhasesabstractRising computational demands have made cloud datacenters' water footprint a critical concern. We demonstrate how different water footprint accounting methodologies - incorporating operational, manufacturing, and decommissioning water consumption, impact measurements and highlight the need for methodology standardization for water-aware operations. Our analysis reveals opportunities for water-aware scheduling in datacenters by considering regional water variations and lifecycle impacts. Amit Samanta 0001, Yankai Jiang 0002, Ryan Stutsman, Rohan Basu Roy |
SoCC | 3 |
| 2025 | GridGreen: Integrating Serverless Computing in HPC Systems for Performance and SustainabilityabstractWe present GridGreen, a scheduling framework that improves the sustainability and performance of scientific workflow execution by integrating serverless computing with traditional on-premise high performance computing (HPC) clusters. GridGreen allocates workflow components across HPC and serverless environments leveraging spatio-temporal variation of carbon intensity and component execution characteristics. It incorporates component-level optimization, speculative pre-warming, I/O-aware data management, and fallback adaptation to jointly minimize carbon footprint and service time under user-defined cost constraints. Our evaluations on large-scale bioinformatics workflows across leadership-class HPC facilities and cloud-based serverless regions demonstrate that GridGreen achieves robust, cost-effective execution while improving carbon efficiency. Amit Samanta 0001, Ryan Stutsman, Rohan Basu Roy |
SoCC | 2 |
| 2025 | Stop Taking the Scenic Route: the Shortest Distance Between the CPU and the NIC is MMIOabstractWhat is the fastest way to transfer data from the CPU to a network interface card (NIC)? Conventional wisdom suggests that the answer is direct memory access (DMA), and recent works have advocated for leveraging cache-coherent I/O interconnects. However, in this paper we examine the arguments against memory-mapped I/O (MMIO) and show that high write throughput can be achieved with MMIO by relaxing ordering constraints. We also propose efficient hardware for recovering ordering at the NIC to ensure correctness while taking advantage of the performance benefits of unordered MMIO writes. Wei Siew Liew, Md Ashfaqur Rahaman, James Patrick Mcmahon, Ryan Stutsman, Vijay Nagarajan |
HotOS | 4 |
| 2025 | Auto-reconfiguration for Latency Minimization in CPU-based DNN ServingabstractIn this paper, we investigate how to push the performance limits of serving Deep Neural Network (DNN) models on CPU-based servers. Specifically, we observe that while intra-operator parallelism across multiple threads is an effective way to reduce inference latency, it provides diminishing returns. Our primary insight is that instead of running a single instance of a model with all available threads on a server, running multiple instances each with smaller batch sizes and fewer threads for intra-op parallelism can provide lower inference latency. However, the right configuration is hard to determine manually since it is workload- (DNN model and batch size used by the serving system) and deployment-dependent (number of CPU cores on server). We present Packrat, a new serving system for online inference that given a model and batch size (𝐵) algorithmically picks the optimal number of instances (𝑖), the number of threads each should be allocated (𝑡), and the batch sizes each should operate on (𝑏) that minimizes latency. Packrat is built as an extension to TorchServe and supports online reconfigurations to avoid serving downtime. Averaged across a range of batch sizes, Packrat improves inference latency by 1.43× to 1.83× on a range of commonly used DNNs. Ankit Bhardwaj 0002, Amar Phanishayee, Deepak Narayanan, Ryan Stutsman |
ICML | 4 |
| 2024 | Fair, Efficient Multi-Resource Scheduling for Stateless Serverless Functions with AnubisabstractAlthough serverless platforms have been extremely popular recently, certain kinds of applications are not well-supported by these platforms. For example, they work well for workloads that transform or aggregate bulk data but not for applications where predictable response times are crucial. These workloads have fine-grained tasks and stringent service level agreements (SLAs). Current serverless platforms’ overheads and resource contention lead to variable performance, making using them for real-time online applications impossible. We demonstrate Anubis, a new platform built on top of OpenWhisk, that helps meet SLAs for response-time-sensitive serverless workloads, even when multiple concurrent workloads compete for various resource types. Anubis’s approach centers on multi-resource fair queuing. Even as stateless functions compete for CPU, storage, and network resources, Anubis ensures each gets its fair share according to its dominant resource needs. We show that when various serverless workloads are intermixed, Anubis’s approach can reduce SLA violations by 15–34%, improving max-min fairness by 2× compared to competing scheduling policies. Amit Samanta 0001, Ryan Stutsman |
CCGrid | 2 |
| 2024 | Poster: Topocloud: Getting Datacenter Network Experiments Into the Right ShapeabstractThe physical topology of a datacenter network is fundamental as it determines latency, bisection bandwidth, the location and properties of congestion, and the network's ability to tolerate failures. Yet, topology is one of the most difficult factors to control during experimentation: when using a production datacenter or a testbed, the only topology available is the one already deployed. If running on an experimenter's own equipment, rewiring may be possible but is cumbersome and of limited scale. This presents a challenge for controlled experimentation, because such experiments should ideally be run on a range of realistic topologies. To tackle this, we present our work on Topocloud, a system that enables experimenters to construct and modify network topologies in a shared public testbed. Topocloud uses hardware P4 switches to fulfill a dual role: each port can act as either a typical Ethernet MAC-learning switch or as a “virtual wire” that connects ports transparently. We demonstrate that virtual wires in Topocloud can be used to construct low latency network topologies incorporating L2 switches. Pavani Kuppili, Aleksander Maricq, Brent E. Stephens, Ryan Stutsman, Robert Ricci |
ICNP | 4 |
| 2023 | Sharding the State Machine: Automated Modular Reasoning for Complex Concurrent Systems
Travis Hance, Yi Zhou 0025, Andrea Lattuada 0001, Reto Achermann, Alexander Conway 0001, Ryan Stutsman, Gerd Zellweger, Chris Hawblitzel, Jon Howell, Bryan Parno |
OSDI | 6 |
| 2023 | Efficient linearizability checking for actor-based systemsabstractAbstract Recent demand for distributed software had led to a surge in popularity in actor‐based frameworks. However, even with the stylized message passing model of actors, writing correct distributed software is still difficult. We present our work on linearizability checking in DS2, an integrated framework for specifying, synthesizing, and testing distributed actor systems. The key insight of our approach is that often subcomponents of distributed actor systems represent common algorithms or data structures (e.g., a distributed hash table or tree) that can be validated against a simple sequential model of the system. This makes it easy for developers to validate their concurrent actor systems without complex specifications. DS2 automatically explores the concurrent schedules that system could arrive at, and it compares observed output of the system to ensure it is equivalent to what the sequential implementation could have produced. We describe DS2's linearizability checking and test it on several concurrent replication algorithms from the literature. We explore in detail how different algorithms for enumerating the model schedule space fare in finding bugs in actor systems, and we present our own refinements on algorithms for exploring actor system schedules that we show are effective in finding bugs. Mohammed Al-Mahfoudh, Ryan Stutsman, Ganesh Gopalakrishnan |
Softw. Pract. Exp. | 2 |
| 2022 | Cache-coherent accelerators for persistent memory crash consistencyabstractBuilding persistent memory (PM) data structures is difficult because crashes interrupt operations, leaving data structures in an inconsistent state. Solving this requires augmenting code that modifies PM state to ensure that interrupted operations can be completed or undone. Today, this is done using careful, hand-crafted code, a compiler pass, or page faults. We propose a new, easy way to transform volatile data structure code to work with PM that uses a cache-coherent accelerator to do this augmentation, and we show that it may outperform existing approaches for building PM structures. Ankit Bhardwaj 0002, Todd Thornley, Vinita Pawar, Reto Achermann, Gerd Zellweger, Ryan Stutsman |
HotStorage | 6 |
| 2022 | XRP: In-Kernel Storage Functions with eBPF
Yuhong Zhong, Yu Jian Wu, Ioannis Zarkadas, Jeffrey Tao, Evan Mesterhazy, Michael Makris, Amy Tai, Ryan Stutsman, Asaf Cidon |
OSDI | 10 |
| 2021 | BPF for storage: an exokernel-inspired approachabstractThe overhead of the kernel storage path accounts for half of the access latency for new NVMe storage devices. We explore using BPF to reduce this overhead, by injecting user-defined functions deep in the kernel's I/O processing stack. When issuing a series of dependent I/O requests, this approach can increase IOPS by over 2.5X and cut latency by half, by bypassing kernel layers and avoiding user-kernel boundary crossings. However, we must avoid losing important properties when bypassing the file system and block layer such as the safety guarantees of the file system and translation between physical blocks addresses and file offsets. We sketch potential solutions to these problems, inspired by exokernel file systems from the late 90s, whose time, we believe, has finally come! Yuhong Zhong, Hongyi Wang 0007, Yu Jian Wu, Asaf Cidon, Ryan Stutsman, Amy Tai |
HotOS | 5 |
| 2021 | NrOS: Effective Replication and Sharing in an Operating System
Ankit Bhardwaj 0002, Chinmay Kulkarni 0002, Reto Achermann, Irina Calciu, Sanidhya Kashyap, Ryan Stutsman, Amy Tai, Gerd Zellweger |
OSDI | 6 |
| 2021 | Achieving High Throughput and Elasticity in a Larger-than-Memory StoreabstractMillions of sensors, mobile applications and machines now generate billions of events. Specialized many-core key-value stores (KVSs) can ingest and index these events at high rates (over 100 Mops/s on one machine) if events are generated on the same machine; however, to be practical and cost-effective they must ingest events over the network and scale across cloud resources elastically. We present Shadowfax, a new distributed KVS based on FASTER, that transparently spans DRAM, SSDs, and cloud blob storage while serving 130 Mops/s/VM over commodity Azure VMs using conventional Linux TCP. Beyond high single-VM performance, Shadowfax uses a unique approach to distributed reconfiguration that avoids any server-side key ownership checks or cross-core coordination both during normal operation and migration. Hence, Shadowfax can shift load in 17 s to improve system throughput by 10 Mops/s with little disruption. Compared to the state-of-the-art, it has 8x better throughput (than Seastar+memcached) and avoids costly I/O to move cold data during migration. On 12 machines, Shadowfax retains its high throughput to perform 930 Mops/s, which, to the best of our knowledge, is the highest reported throughput for a distributed KVS used for large-scale data ingestion and indexing. Chinmay Kulkarni 0002, Badrish Chandramouli, Ryan Stutsman |
Proc. VLDB Endow. | 3 |
| 2020 | Auto-Scaling Cloud-Based Memory-Intensive ApplicationsabstractToday, Cloud providers offer simplistic scaling policies that rely on thresholds that force tenants to have a priori knowledge of their workloads. We develop a new method for scaling memory-intensive workloads that needs no thresholds. This makes it worry-free for tenants, and it adapts even as workloads evolve. This is especially hard for memory-bound applications where even a small decrease in the amount of memory available can have a dramatic, almost unbounded impact on performance. Hence, sizing a machine's physical memory correctly is critical to application performance and operating cost. To determine a natural threshold for memory-intensive applications, our approach automatically analyzes an application's miss ratio curve (MRC) and models it as a hyperbola. Intuitively, a memory scaling policy should operate at the point where the curve flattens: that is, at its intersection with its latus rectum (LR). Our system uses a new approach to constructing and analyzing MRCs at run time that captures memory references from a slice of any scalable application as it executes on standard virtual machines from any major Cloud provider. We demonstrate with multiple applications running on Amazon Web Services (AWS) and Microsoft Azure. Our implementation and evaluation show that, though the LR doesn't require tenants to set thresholds, it is effective in scaling memory-intensive workloads to save on operating costs while avoiding queuing, thrashing, or collapse. It increases throughput by 1.5× and reduces queuing delay by 2× in our evaluation. Joe H. Novak, Sneha Kumar Kasera, Ryan Stutsman |
CLOUD | 3 |
| 2020 | Compact Leakage-Free Support for Integrity and ReliabilityabstractThe memory system is vulnerable to a number of security breaches, e.g., an attacker can interfere with program execution by disrupting values stored in memory. Modern Intel® Software Guard Extension (SGX) systems already support integrity trees to detect such malicious behavior. However, in spite of recent innovations, the bandwidth overhead of integrity+replay protection is non-trivial; state-of-the-art solutions like Synergy introduce average slowdowns of 2.3× for memory-intensive benchmarks. Prior work also implements a tree that is shared by multiple applications, thus introducing a potential side channel. In this work, we build on the Synergy and SGX baselines, and introduce three new techniques. First, we isolate each application by implementing a separate integrity tree and metadata cache for each application; this improves metadata cache efficiency and improves performance by 39%, while eliminating the potential side channel. Second, we reduce the footprint of the metadata. Synergy uses a combination of integrity and error correction metadata to provide low-overhead support for both. We share error correction metadata across multiple blocks, thus lowering its footprint (by 16×) while preventing error correction only in rare corner cases. However, we discover that shared error correction metadata, even with caching, does not improve performance. Third, we observe that thanks to its lower footprint, the error correction metadata can be embedded into the integrity tree. This reduces the metadata blocks that must be accessed to support both integrity verification and chipkill reliability. The proposed Isolated Tree with Embedded Shared Parity (ITESP) yields an overall performance improvement of 64%, relative to baseline Synergy. Meysam Taassori, Rajeev Balasubramonian, Siddhartha Chhabra, Alaa R. Alameldeen, Manjula Peddireddy, Rajat Agarwal, Ryan Stutsman |
ISCA | 7 |
| 2020 | Adaptive Placement for In-memory Storage Functions
Ankit Bhardwaj 0002, Chinmay Kulkarni 0002, Ryan Stutsman |
USENIX ATC | 3 |
| 2019 | DPI: The Data Processing Interface for Modern Networks
Gustavo Alonso, Carsten Binnig, Ippokratis Pandis, Kenneth Salem, Jan Skrzypczak, Ryan Stutsman, Lasse Thostrup, Tianzheng Wang 0001, Zeke Wang, Tobias Ziegler 0001 |
CIDR | 6 |
| 2019 | Narrowing the Gap Between Serverless and its State with Storage FunctionsabstractServerless computing has gained attention due to its fine-grained provisioning, large-scale multi-tenancy, and on-demand scaling. However, it also forces applications to externalize state in remote storage, adding substantial overheads. To fix this "data shipping problem" we built Shredder, a low-latency multi-tenant cloud store that allows small units of computation to be performed directly within storage nodes. Storage tenants provide Shredder with JavaScript functions (or WebAssembly programs), which can interact directly with data without moving them over the network. Tian Zhang 0005, Dong Xie 0001, Feifei Li 0001, Ryan Stutsman |
SoCC | 4 |
| 2019 | GenCache: Leveraging In-Cache Operators for Efficient Sequence AlignmentabstractPrecision Medicine will rely on frequent genomic analysis, especially for patients undergoing cancer treatments or suffering from rare diseases. Sequence alignment is invoked in multiple stages of the genomic analysis pipeline. Recent projects have introduced accelerators, GenAx and Darwin, for 2nd and 3rd generation sequencers respectively. In this work, we improve upon the GenAx design by increasing its parallelism and reducing its memory bandwidth demands. This is achieved with a combination of hardware and software innovations. We first integrate in-cache operators from prior work into the GenAx memory hierarchy; we then augment the in-cache peripheral circuit to support additional new operators. We then re-structure the sequence alignment algorithm to (i) leverage the many in-cache operators, (ii) exploit the common case in genomic datasets, (iii) use Bloom Filters to reduce futile accesses, and (iv) maximize data reuse within a re-organized memory hierarchy. While the baseline GenAx accelerator processes a batch of reads in 194 seconds while nearly saturating the 153.6 GB/s memory bandwidth, the proposed GenCache architecture processes the same batch of reads in 37 seconds at an improved energy efficiency of 8.6×, while demanding 20 GB/s average memory bandwidth. Our hardware and software techniques thus interact synergistically to target both memory and compute bottlenecks, while not affecting the outputs of the application. We show that the basic principles in GenCache can also be exploited by 3rd generation sequence aligners. Anirban Nag, C. N. Ramachandra, Rajeev Balasubramonian, Ryan Stutsman, Edouard Giacomin, Hari Kambalasubramanyam, Pierre-Emmanuel Gaillardon |
MICRO | 4 |
| 2019 | Flashield: a Hybrid Key-value Cache that Controls Flash Write Amplification
Assaf Eisenman, Asaf Cidon, Evgenya Pergament, Or Haimovich, Ryan Stutsman, Mohammad Alizadeh, Sachin Katti |
NSDI | 5 |
| 2019 | SolarDB: Toward a Shared-Everything Database on Distributed Log-Structured StorageabstractEfficient transaction processing over large databases is a key requirement for many mission-critical applications. Although modern databases have achieved good performance through horizontal partitioning, their performance deteriorates when cross-partition distributed transactions have to be executed. This article presents SolarDB, a distributed relational database system that has been successfully tested at a large commercial bank. The key features of SolarDB include (1) a shared-everything architecture based on a two-layer log-structured merge-tree; (2) a new concurrency control algorithm that works with the log-structured storage, which ensures efficient and non-blocking transaction processing even when the storage layer is compacting data among nodes in the background; and (3) find-grained data access to effectively minimize and balance network communication within the cluster. According to our empirical evaluations on TPC-C, Smallbank, and a real-world workload, SolarDB outperforms the existing shared-nothing systems by up to 50x when there are close to or more than 5% distributed transactions. Tao Zhu 0004, Zhuoyue Zhao 0001, Feifei Li 0001, Weining Qian, Aoying Zhou, Dong Xie 0001, Ryan Stutsman, HaiNing Li, Huiqi Hu |
ACM Trans. Storage | 7 |
| 2018 | MobileStream: a scalable, programmable and evolvable mobile core control plane platformabstractControl planes in future mobile core networks face two new challenges. First, they must scale to process the growing control traffic generated by an ever increasing number of mobile devices. Second, they must be flexible and evolvable to support the range of emerging service abstractions and to realize customized network slices to meet the broad range of requirements of these networks. To address these challenges, we propose MobileStream, a scalable, programmable, and evolvable mobile core control plane platform. MobileStream provides a set of refactored basic building blocks, functionally decomposed from existing monolithic control plane components. It leverages realtime streaming frameworks to assemble, execute, and scale these blocks as streaming control plane applications. Moreover, it allows users to add their own functions to customize and optimize streaming control plane applications. We present several streaming control plane applications to showcase the flexibility and generality of MobileStream. We describe our extensive functional testing, with a variety of mobile devices and base stations, to validate the MobileStream prototype, and present the results of large-scale experiments demonstrating its scalability. Junguk Cho, Ryan Stutsman, Jacobus E. van der Merwe |
CoNEXT | 2 |
| 2018 | Hybrid network clusters using common gameplay for massively multiplayer online gamesabstractWith advancements in network technology and cost-efficient hardware, developers have begun placing servers throughout the world. These servers reside at the edge of cloud network infrastructures and vastly improved network quality. However, many locations in the world are still distant when accessing these edge servers. Further, massively multiplayer online games can strain edge server resources with additional hardware not available at the desired location. Jared N. Plumb, Sneha Kumar Kasera, Ryan Stutsman |
FDG | 3 |
| 2018 | Exploiting Google's Edge Network for Massively Multiplayer Online GamesabstractMassively multiplayer online game servers are challenging to build and maintain; they require always-on availability, low-latency, and high predictability. In many ways, these systems could benefit from pushing application logic to the connected clients. However, historically peer-to-peer networks have not worked when applied to this domain. Pushing game logic to the clients complicates development since client end hosts cannot trust their peers and they have unreliably persistent connections, increasing the possibility of cheating or network failure. Google's Edge Network changes everything we have concluded about peer-to-peer networks over the past decade. Having access to thousands of servers all over the world solves problems of availability, data access, link saturation, security, and control. This new system allows for the inclusion of trusted peers in otherwise untrusted node clusters. Developers can explore peer-to-peer or decentralized algorithms while allowing complete control over their data, code, and players, in their developed virtual worlds. They can offload application game logic to a low-latency scalable system, which provides a cost-effective solution residing closer to the connected clients. In this paper, we investigate the gains game developers may obtain by exploring emerging edge cloud technology. We demonstrate how massively multiplayer online games benefit from this new system though simulation. We compare area-of-interest latency, which is the latency between client interactions within the virtual world for known World of Warcraft and Google data centers locations. We show the benefits to area-of-interest latency gained from using Google's Edge Network, while minimally affecting client to server latency. Finally, we present a novel approach to maximize latency reductions for area-of- interest latency by moving players to optimal peering edge servers, thus reducing distance between clients. Jared N. Plumb, Ryan Stutsman |
ICFEC | 2 |
| 2018 | ECHO: A Reliable Distributed Cellular Core Network for Hyper-scale Public CloudsabstractEconomies of scale associated with hyper-scale public cloud platforms offer flexibility and cost-effectiveness, resulting in various services and businesses moving to the cloud. One area with little progress in this direction is cellular core networks. A cellular core network manages the state of cellular clients; it is essentially a large distributed state machine with very different virtualization challenges compared to typical cloud services. In this paper we present a novel cellular core network architecture, called ECHO, particularly suited to public cloud deployments, where the availability guarantees might be an order of magnitude worse compared to existing (redundant) hardware platforms. We present the design and implementation of our approach and evaluate its functionality on a public cloud platform. Analysis shows ECHO promises higher availability than existing telco solutions. Binh Nguyen 0003, Tian Zhang 0005, Bozidar Radunovic, Ryan Stutsman, Thomas Karagiannis, Jakub Kocur, Jacobus E. van der Merwe |
MobiCom | 4 |
| 2018 | Splinter: Bare-Metal Extensions for Multi-Tenant Low-Latency Storage
Chinmay Kulkarni 0002, Sara Moore, Mazhar Naqvi, Tian Zhang 0005, Robert Ricci, Ryan Stutsman |
OSDI | 6 |
| 2018 | Taming Performance Variability
Aleksander Maricq, Dmitry Duplyakin, Ivo Jimenez, Carlos Maltzahn, Ryan Stutsman, Robert Ricci, Ana Klimovic |
OSDI | 5 |
| 2018 | Tailwind: Fast and Atomic RDMA-based Replication
Yacine Taleb, Ryan Stutsman, Gabriel Antoniu, Toni Cortes |
USENIX ATC | 2 |
| 2018 | Solar: Towards a Shared-Everything Database on Distributed Log-Structured Storage
Tao Zhu 0004, Zhuoyue Zhao 0001, Feifei Li 0001, Weining Qian, Aoying Zhou, Dong Xie 0001, Ryan Stutsman, HaiNing Li, Huiqi Hu |
USENIX ATC | 7 |
| 2018 | anexVis: visual analytics framework for analysis of RNA expressionabstractSummary: Although RNA expression data are accumulating at a remarkable speed, gaining insights from them still requires laborious analyses, which hinder many biological and biomedical researchers. This report introduces a visual analytics framework that applies several well-known visualization techniques to leverage understanding of an RNA expression dataset. Our analyses on glycosaminoglycan-related genes have demonstrated the broad application of this tool, anexVis (analysis of RNA expression), to advance the understanding of tissue-specific glycosaminoglycan regulation and functions, and potentially other biological pathways. Availability and implementation: The application is accessible at https://anexvis.chpc.utah.edu/, source codes deposited on GitHub. Supplementary information: Supplementary data are available at Bioinformatics online. Diem-Trang T. Tran, Tian Zhang 0005, Ryan Stutsman, Matthew Might, Umesh R. Desai, Balagurunathan Kuberan |
Bioinform. | 3 |
| 2017 | Rocksteady: Fast Migration for Low-latency In-memory StorageabstractScalable in-memory key-value stores provide low-latency access times of a few microseconds and perform millions of operations per second per server. With all data in memory, these systems should provide a high level of reconfigurability. Ideally, they should scale up, scale down, and rebalance load more rapidly and flexibly than disk-based systems. Rapid reconfiguration is especially important in these systems since a) DRAM is expensive and b) they are the last defense against highly dynamic workloads that suffer from hot spots, skew, and unpredictable load. However, so far, work on in-memory key-value stores has generally focused on performance and availability, leaving reconfiguration as a secondary concern. Chinmay Kulkarni 0002, Aniraj Kesavan, Tian Zhang 0005, Robert Ricci, Ryan Stutsman |
SOSP | 5 |
| 2017 | Memshare: a Dynamic Multi-tenant Key-value Cache
Asaf Cidon, Daniel Rushton, Stephen M. Rumble, Ryan Stutsman |
USENIX ATC | 4 |
| 2015 | High Performance Transactions in Deuteronomy
Justin J. Levandoski, David B. Lomet, Sudipta Sengupta, Ryan Stutsman, Rui Wang 0002 |
CIDR | 4 |
| 2015 | Experience with Rules-Based Programming for Distributed, Concurrent, Fault-Tolerant Code
Ryan Stutsman, Collin Lee, John K. Ousterhout |
USENIX ATC | 1 |
| 2015 | Multi-Version Range Concurrency Control in DeuteronomyabstractThe Deuteronomy transactional key value store executes millions of serializable transactions/second by exploiting multi-version timestamp order concurrency control. However, it has not supported range operations, only individual record operations (e.g., create, read, update, delete). In this paper, we enhance our multi-version timestamp order technique to handle range concurrency and prevent phantoms. Importantly, we maintain high performance while respecting the clean separation of duties required by Deuteronomy, where a transaction component performs purely logical concurrency control (including range support), while a data component performs data storage and management duties. Like the rest of the Deuteronomy stack, our range technique manages concurrency information in a latch-free manner. With our range enhancement, Deuteronomy can reach scan speeds of nearly 250 million records/s (more than 27 GB/s) on modern hardware, while providing serializable isolation complete with phantom prevention. Justin J. Levandoski, David B. Lomet, Sudipta Sengupta, Ryan Stutsman, Rui Wang 0002 |
Proc. VLDB Endow. | 4 |
| 2015 | To Lock, Swap, or Elide: On the Interplay of Hardware Transactional Memory and Lock-Free IndexingabstractThe release of hardware transactional memory (HTM) in commodity CPUs has major implications on the design and implementation of main-memory databases, especially on the architecture of high-performance lock-free indexing methods at the core of several of these systems. This paper studies the interplay of HTM and lock-free indexing methods. First, we evaluate whether HTM will obviate the need for crafty lock-free index designs by integrating it in a traditional B-tree architecture. HTM performs well for simple data sets with small fixed-length keys and payloads, but its benefits disappear for more complex scenarios (e.g., larger variable-length keys and payloads), making it unattractive as a general solution for achieving high performance. Second, we explore fundamental differences between HTM-based and lock-free B-tree designs. While lock-freedom entails design complexity and extra mechanism, it has performance advantages in several scenarios, especially high-contention cases where readers proceed uncontested (whereas HTM aborts readers). Finally, we explore the use of HTM as a method to simplify lock-free design. We find that using HTM to implement a multi-word compare-and-swap greatly reduces lock-free programming complexity at the cost of only a 10-15% performance degradation. Our study uses two state-of-the-art index implementations: a memory-optimized B-tree extended with HTM to provide multi-threaded concurrency and the Bw-tree lock-free B-tree used in several Microsoft production environments. Darko Makreshanski, Justin J. Levandoski, Ryan Stutsman |
Proc. VLDB Endow. | 3 |
| 2015 | The RAMCloud Storage SystemabstractRAMCloud is a storage system that provides low-latency access to large-scale datasets. To achieve low latency, RAMCloud stores all data in DRAM at all times. To support large capacities (1PB or more), it aggregates the memories of thousands of servers into a single coherent key-value store. RAMCloud ensures the durability of DRAM-based data by keeping backup copies on secondary storage. It uses a uniform log-structured mechanism to manage both DRAM and secondary storage, which results in high performance and efficient memory usage. RAMCloud uses a polling-based approach to communication, bypassing the kernel to communicate directly with NICs; with this approach, client applications can read small objects from any RAMCloud storage server in less than 5μs, durable writes of small objects take about 13.5μs. RAMCloud does not keep multiple copies of data online; instead, it provides high availability by recovering from crashes very quickly (1 to 2 seconds). RAMCloud’s crash recovery mechanism harnesses the resources of the entire cluster working concurrently so that recovery performance scales with cluster size. John K. Ousterhout, Arjun Gopalan, Ankita Kejriwal, Collin Lee, Behnam Montazeri, Diego Ongaro, Seo Jin Park, Henry Qin, Mendel Rosenblum, Stephen M. Rumble, Ryan Stutsman |
ACM Trans. Comput. Syst. | 12 |
| 2013 | Toward Common Patterns for Distributed, Concurrent, Fault-Tolerant Code
Ryan Stutsman, John K. Ousterhout |
HotOS | 1 |
| 2013 | Copysets: Reducing the Frequency of Data Loss in Cloud Storage
Asaf Cidon, Stephen M. Rumble, Ryan Stutsman, Sachin Katti, John K. Ousterhout, Mendel Rosenblum |
USENIX ATC | 3 |
| 2011 | Energy management in mobile devices with the cinder operating systemabstractWe argue that controlling energy allocation is an increasingly useful and important feature for operating systems, especially on mobile devices. We present two new low-level abstractions in the Cinder operating system, reserves and taps, which store and distribute energy for application use. We identify three key properties of control -- isolation, delegation, and subdivision -- and show how using these abstractions can achieve them. We also show how the architecture of the HiStar information-flow control kernel lends itself well to energy control. We prototype and evaluate Cinder on a popular smartphone, the Android G1. Stephen M. Rumble, Ryan Stutsman, Philip Alexander Levis, David Mazières, Nickolai Zeldovich |
EuroSys | 3 |
| 2011 | It's Time for Low Latency
Stephen M. Rumble, Diego Ongaro, Ryan Stutsman, Mendel Rosenblum, John K. Ousterhout |
HotOS | 3 |
| 2011 | Fast crash recovery in RAMCloudabstractRAMCloud is a DRAM-based storage system that provides inexpensive durability and availability by recovering quickly after crashes, rather than storing replicas in DRAM. RAMCloud scatters backup data across hundreds or thousands of disks, and it harnesses hundreds of servers in parallel to reconstruct lost data. The system uses a log-structured approach for all its data, in DRAM as well as on disk: this provides high performance both during normal operation and during recovery. RAMCloud employs randomized techniques to manage the system in a scalable and decentralized fashion. In a 60-node cluster, RAMCloud recovers 35 GB of data from a failed server in 1.6 seconds. Our measurements suggest that the approach will scale to recover larger memory sizes (64 GB or more) in less time with larger clusters. Diego Ongaro, Stephen M. Rumble, Ryan Stutsman, John K. Ousterhout, Mendel Rosenblum |
SOSP | 3 |
| 2009 | Translation-based steganographyabstractThis paper investigates systems that steganographically embed information in the “noise” created by automatic translation of natural language documents. The main thrust of the work focuses on two problems – generation of plausible steganographic text Christian Grothoff, Krista Bennett, Ryan Stutsman, Ludmila Alkhutova, Mikhail J. Atallah |
J. Comput. Secur. | 3 |