Amy Ousterhout

dblp:149/9258 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-6590-8392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Extended User Interrupts (xUI): Fast and Flexible Notification without Polling
abstract
Extended user interrupts (xUI) is a set of processor extensions that builds on Intel's UIPI model of user interrupts, for enhanced performance and flexibility. This paper deconstructs Intel's current UIPI design through analysis and measurement, and uses this to develop an accurate model of its timing. It then introduces four novel enhancements to user interrupts: tracked interrupts, hardware safepoints, a kernel bypass timer, and interrupt forwarding. xUI is modeled in gem5 simulation and evaluated on three use cases -- preemption in a high-performance user-level runtime, IO notification in a layer3 router using DPDK, and IO notification in a synthetic workload with a streaming accelerator modeled after Intel's Data Streaming Accelerator. This work shows that xUI offers the performance of shared memory polling with the efficiency of asynchronous notification.
Berk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel, Dean M. Tullsen, Amy Ousterhout, Mohammadkazem Taram
ASPLOS (2)6
2025 The Benefits and Limitations of User Interrupts for Preemptive Userspace Scheduling
Linsong Guo, Danial Zuberi, Tal Garfinkel, Amy Ousterhout
NSDI4
2025 Eden: Developer-Friendly Application-Integrated Far Memory
Anil Yelam, Stewart Grant, Saarth Deshpande, Nadav Amit, Radhika Niranjan Mysore, Amy Ousterhout, Marcos K. Aguilera, Alex C. Snoeren
NSDI6
2024 Efficient Microsecond-scale Blind Scheduling with Tiny Quanta
abstract
A longstanding performance challenge in datacenter-based applications is how to efficiently handle incoming client requests that spawn many very short (μs scale) jobs that must be handled with high throughput and low tail latency. When no assumptions are made about the duration of individual jobs, or even about the distribution of their durations, this requires blind scheduling with frequent and efficient preemption, which is not scalably supported for μs-level tasks. We present Tiny Quanta (TQ), a system that enables efficient blind scheduling of μs-level workloads. TQ performs fine-grained preemptive scheduling and does so with high performance via a novel combination of two mechanisms: forced multitasking and two-level scheduling. Evaluations with a wide variety of μs-level workloads show that TQ achieves low tail latency while sustaining 1.2x to 6.8x the throughput of prior blind scheduling systems.
Zhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro, Amy Ousterhout, Sylvia Ratnasamy, Scott Shenker
ASPLOS (2)5
2023 Zed: Leveraging Data Types to Process Eclectic Data
Amy Ousterhout, Steven McCanne, Henri Dubois-Ferrière, Silvery D. Fu, Sylvia Ratnasamy, Noah Treuhaft
CIDR1
2023 Out of Hand for Hardware? Within Reach for Software!
abstract
Events that take 10s to 100s of ns like cache misses increasingly cause CPU stalls. However, hiding the latency of these events is challenging: hardware mechanisms suffer from the lack of flexibility, whereas prior software mechanisms fall short due to large overhead and limited event visibility. In this paper, we argue that with a combination of two emerging techniques - light-weight coroutines and sample-based profiling, hiding these events in software is within reach.
Zhihong Luo, Silvery D. Fu, Emmanuel Amaro, Amy Ousterhout, Sylvia Ratnasamy, Scott Shenker
HotOS4
2023 Cornflakes: Zero-Copy Serialization for Microsecond-Scale Networking
abstract
Data serialization is critical for many datacenter applications, but the memory copies required to move application data into packets are costly. Recent zero-copy APIs expose NIC scatter-gather capabilities, raising the possibility of offloading this data movement to the NIC. However, as the memory coordination required for scatter-gather adds bookkeeping overhead, scatter-gather is not always useful. We describe Cornflakes, a hybrid serialization library stack that uses scatter-gather for serialization when it improves performance and falls back to memory copies otherwise. We have implemented Cornflakes within a UDP and TCP networking stack, across Mellanox and Intel NICs. On a Twitter cache trace, Cornflakes achieves 15.4% higher throughput than prior software approaches on a custom key-value store and 8.8% higher throughput than Redis serialization within Redis.
Deepti Raghavan, Shreya Ravi, Gina Yuan, Pratiksha Thaker, Sanjari Srivastava, Micah Murray, Pedro Henrique de Mello Morado Penna, Amy Ousterhout, Philip Alexander Levis, Matei Zaharia, Irene Zhang
SOSP8
2022 Efficient Scheduling Policies for Microsecond-Scale Tasks
Sarah McClure, Amy Ousterhout, Scott Shenker, Sylvia Ratnasamy
NSDI2
2020 Can far memory improve job throughput?
abstract
As memory requirements grow, and advances in memory technology slow, the availability of sufficient main memory is increasingly the bottleneck in large compute clusters. One solution to this is memory disaggregation, where jobs can remotely access memory on other servers, or far memory. This paper first presents faster swapping mechanisms and a far memory-aware cluster scheduler that make it possible to support far memory at rack scale. Then, it examines the conditions under which this use of far memory can increase job throughput. We find that while far memory is not a panacea, for memory-intensive workloads it can provide performance improvements on the order of 10% or more even without changing the total amount of memory available.
Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout, Marcos K. Aguilera, Aurojit Panda, Sylvia Ratnasamy, Scott Shenker
EuroSys4
2020 Remote Memory Calls
abstract
In this paper we propose an extension to RDMA, called Remote Memory Calls (RMCs), that allows applications to install a customized set of 1-sided RDMA operations. We then explain how RMCs can be implemented on the forthcoming generation of SmartNICs and discuss the resulting tradeoffs between RMCs, 1-sided and 2-sided RDMA operations.
Emmanuel Amaro, Zhihong Luo, Amy Ousterhout, Arvind Krishnamurthy, Aurojit Panda, Sylvia Ratnasamy, Scott Shenker
HotNets3
2020 Caladan: Mitigating Interference at Microsecond Timescales
Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam Belay
OSDI3
2019 Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads
Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, Hari Balakrishnan
NSDI1
2017 Flexplane: An Experimentation Platform for Resource Management in Datacenters
Amy Ousterhout, Jonathan Perry 0001, Hari Balakrishnan, Petr Lapukhov
NSDI1
2014 Fastpass: a centralized "zero-queue" datacenter network
abstract
An ideal datacenter network should provide several properties, including low median and tail latency, high utilization (throughput), fair allocation of network resources between users or applications, deadline-aware scheduling, and congestion (loss) avoidance. Current datacenter networks inherit the principles that went into the design of the Internet, where packet transmission and path selection decisions are distributed among the endpoints and routers. Instead, we propose that each sender should delegate control---to a centralized arbiter---of when each packet should be transmitted and what path it should follow.
Jonathan Perry 0001, Amy Ousterhout, Hari Balakrishnan, Devavrat Shah, Hans Fugal
SIGCOMM2