Rebecca Isaacs

dblp:89/4883 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0002-6737-1503ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1
YearPublicationVenuePosition
2026 Analyzing Metastable Failures
Rebecca Isaacs
CIDR1
2025 Analyzing Metastable Failures
abstract
A metastable failure is a self-sustaining congestive collapse in which a system degrades in response to a transient stressor (e.g., a load surge) but fails to recover after the stressor is removed. These rare but potentially catastrophic events are notoriously hard to diagnose and mitigate, sometimes causing prolonged outages affecting millions of users.
Rebecca Isaacs, Peter Alvaro, Rupak Majumdar, Kiran Kumar, Muniswamy Reddy, Mahmoud Salamati, Sadegh Esmaeil Zadeh Soudjani
HotOS1
2023 LatenSeer: Causal Modeling of End-to-End Latency Distributions by Harnessing Distributed Tracing
abstract
End-to-end latency estimation in web applications is crucial for system operators to foresee the effects of potential changes, helping ensure system stability, optimize cost, and improve user experience. However, estimating latency in microservices-based architectures is challenging due to the complex interactions between hundreds or thousands of loosely coupled microservices. Current approaches either track only latency-critical paths or require laborious bespoke instrumentation, which is unrealistic for end-to-end latency estimation in complex systems.
Yazhuo Zhang, Rebecca Isaacs, Yao Yue, Juncheng Yang, Lei Zhang 0223, Ymir Vigfusson
SoCC2
2022 Metastable Failures in the Wild
Lexiang Huang, Matt Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, Aleksey Charapko
OSDI5
2017 Thinking about Availability in Large Service Infrastructures
abstract
Our company has learned to design and operate planetaryscale services with reasonably high availability.Historically, these have been Software as a Service (SaaS) systems (search, YouTube, GMail, etc.), implemented as scale-out distributed systems that could tolerate all sorts of failures in lower layers, through the use of traditional techniques such as replication, distributed consensus algorithms (e.g.Paxos), and transactions,
Jeffrey C. Mogul, Rebecca Isaacs, Brent Welch
HotOS2
2013 Differential Dataflow
Frank McSherry, Derek Gordon Murray, Rebecca Isaacs, Michael Isard
CIDR3
2013 Naiad: a timely dataflow system
abstract
Naiad is a distributed system for executing data parallel, cyclic dataflow programs. It offers the high throughput of batch processors, the low latency of stream processors, and the ability to perform iterative and incremental computations. Although existing systems offer some of these features, applications that require all three have relied on multiple platforms, at the expense of efficiency, maintainability, and simplicity. Naiad resolves the complexities of combining these features in one framework.
Derek Gordon Murray, Frank McSherry, Rebecca Isaacs, Michael Isard, Paul Barham 0001, Martín Abadi
SOSP3
2011 More Intervention Now!
Moisés Goldszmidt, Rebecca Isaacs
HotOS2
2011 AC: composable asynchronous IO for native languages
abstract
This paper introduces AC, a set of language constructs for composable asynchronous IO in native languages such as C/C++. Unlike traditional synchronous IO interfaces, AC lets a thread issue multiple IO requests so that they can be serviced concurrently, and so that long-latency operations can be overlapped with computation. Unlike traditional asynchronous IO interfaces, AC retains a sequential style of programming without requiring code to use multiple threads, and without requiring code to be "stack-ripped" into chains of callbacks. AC provides an "async" statement to identify opportunities for IO operations to be issued concurrently, a "do..finish" block that waits until any enclosed "async" work is complete, and a "cancel" statement that requests cancellation of unfinished IO within an enclosing "do..finish". We give an operational semantics for a core language. We describe and evaluate implementations that are integrated with message passing on the Barrelfish research OS, and integrated with asynchronous file and network IO on Microsoft Windows. We show that AC offers comparable performance to existing C/C++ interfaces for asynchronous IO, while providing a simpler programming model.
Tim Harris 0001, Martín Abadi, Rebecca Isaacs, Ross McIlroy
OOPSLA3
2009 Your computer is already a distributed system. Why isn't your OS?
Andrew Baumann, Simon Peter 0001, Adrian Schüpbach, Akhilesh Singhania, Timothy Roscoe, Paul Barham 0001, Rebecca Isaacs
HotOS7
2009 The multikernel: a new OS architecture for scalable multicore systems
abstract
Commodity computer systems contain more and more processor cores and exhibit increasingly diverse architectural tradeoffs, including memory hierarchies, interconnects, instruction sets and variants, and IO configurations. Previous high-performance computing systems have scaled in specific cases, but the dynamic nature of modern client and server workloads, coupled with the impossibility of statically optimizing an OS for all workloads and hardware variants pose serious challenges for operating system structures.
Andrew Baumann, Paul Barham 0001, Pierre-Évariste Dagand, Tim Harris 0001, Rebecca Isaacs, Simon Peter 0001, Timothy Roscoe, Adrian Schüpbach, Akhilesh Singhania
SOSP5
2008 30 seconds is not enough!: a study of operating system timer usage
abstract
The basic system timer facilities used by applications and OS kernels for scheduling timeouts and periodic activities have remained largely unchanged for decades, while hardware architectures and application loads have changed radically. This raises concerns with CPU overhead power management and application responsiveness.
Simon Peter 0001, Andrew Baumann, Timothy Roscoe, Paul Barham 0001, Rebecca Isaacs
EuroSys5
2008 CT-NOR: Representing and Reasoning About Events in Continuous Time
Aleksandr Simma, Moisés Goldszmidt, John MacCormick, Paul Barham 0001, Richard Black, Rebecca Isaacs, Richard Mortier
UAI6
2006 Discovering Dependencies for Network Management
Paramvir Bahl, Paul Barham 0001, Richard Black, Ranveer Chandra, Moisés Goldszmidt, Rebecca Isaacs, Srikanth Kandula, John MacCormick, David A. Maltz, Richard Mortier, Michal Wawrzoniak, Ming Zhang 0005
HotNets6
2006 Reclaiming Network-wide Visibility Using Ubiquitous Endsystem Monitors
Evan Cooke, Richard Mortier, Austin Donnelly, Paul Barham 0001, Rebecca Isaacs
USENIX ATC, General Track5
2004 Using Magpie for Request Extraction and Workload Modelling
Paul Barham 0001, Austin Donnelly, Rebecca Isaacs, Richard Mortier
OSDI3
2003 Magpie: Online Modelling and Performance-aware Systems
Paul Barham 0001, Rebecca Isaacs, Richard Mortier, Dushyanth Narayanan
HotOS2