Irene Zhang

dblp:68/1122 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 14 · 5 first-author · 6 since 2021Systems, architecture and hardware · 10 · 3 first-author · 1 since 2021Computer networks · 4 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Capybara: Dynamic Load Balancing with Microsecond-Scale TCP Migration
abstract
Layer-4 load balancers are a popular solution to high tail latencies but perform poorly under unpredictable skewed workloads because they statically assign connections to servers. We present Capybara, a new load balancer architecture that enables dynamic rebalancing of established connections. Capybara divides load balancing responsibility into a fast L4 load balancer, a host-switch co-designed connection migration protocol, and a transport interface for application-level connection state migration. Capybara leverages two trends - programmable switches and kernel-bypass - to efficiently implement connection migration without disruption, while maintaining transparency to clients. Under realistic workloads, Capybara achieves up to 149× lower tail latency and more than 2× higher throughput for scale-out services compared to state-of-the-art load balancing approaches.
Inho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi, Joshua Fried, Omar S. Navarro Leija, Dan R. K. Ports, Irene Zhang, Jialin Li 0001
SIGCOMM8
2025 Apiary: An OS for the Modern FPGA
abstract
Many datacenter operators have deployed FPGAs as hardware accelerators because their reconfigurability allows them to be repurposed as the application mix changes. Directly attaching the FPGA to the network further reduces latency, improves cost-performance, and reduces energy use relative to mediating network communications with CPUs. However, building accelerated applications or services for direct-attached FPGAs is challenging, especially with the complex I/O and multi-accelerator capacity of modern FPGAs. To address this, we propose Apiary, a microkernel operating system for direct-attached FPGA accelerators. The key idea in Apiary is to raise the level of abstraction for accelerated application code, with security, virtualization, threaded execution, and interprocess communication provided by the hardware OS layer.
Katie Lim, Matthew Giordano, Irene Zhang, Baris Kasikci, Thomas E. Anderson
HotOS3
2024 Beehive: A Flexible Network Stack for Direct-Attached Accelerators
abstract
Direct-attached accelerators, where application accelerators are directly connected to the datacenter network via a hardware network stack, offer substantial benefits in terms of reduced latency, CPU overhead, and energy use. However, a key challenge is that modern datacenter network stacks are complex, with interleaved protocol layers, network management functions, and virtualization support. To operators, network feature agility, diagnostics, and manageability are often considered just as important as raw performance. By contrast, existing hardware network stacks only support basic protocols and are often difficult to extend since they use fixed processing pipelines. We propose Beehive, a new, open-source FPGA network stack for direct-attached accelerators designed to enable flexible and adaptive construction of complex network functionality in hardware. Application and network protocol elements are modularized as tiles over a network-on-chip substrate. Elements can be added or scaled up/down to match workload characteristics with minimal effort or changes to other elements. Flexible diagnostics and control are integral, with tooling to ensure deadlock safety. Our implementation interoperates with standard Linux TCP and UDP clients, with a 4x improvement in end-to-end RPC tail latency for Linux UDP clients versus a CPU-attached accelerator. Beehive is available at https://github:com/beehive-fpga/beehive
Katie Lim, Matthew Giordano, Theano Stavrinos, Irene Zhang, Jacob Nelson 0001, Baris Kasikci, Thomas E. Anderson
MICRO4
2023 Cornflakes: Zero-Copy Serialization for Microsecond-Scale Networking
abstract
Data serialization is critical for many datacenter applications, but the memory copies required to move application data into packets are costly. Recent zero-copy APIs expose NIC scatter-gather capabilities, raising the possibility of offloading this data movement to the NIC. However, as the memory coordination required for scatter-gather adds bookkeeping overhead, scatter-gather is not always useful. We describe Cornflakes, a hybrid serialization library stack that uses scatter-gather for serialization when it improves performance and falls back to memory copies otherwise. We have implemented Cornflakes within a UDP and TCP networking stack, across Mellanox and Intel NICs. On a Twitter cache trace, Cornflakes achieves 15.4% higher throughput than prior software approaches on a custom key-value store and 8.8% higher throughput than Redis serialization within Redis.
Deepti Raghavan, Shreya Ravi, Gina Yuan, Pratiksha Thaker, Sanjari Srivastava, Micah Murray, Pedro Henrique de Mello Morado Penna, Amy Ousterhout, Philip Alexander Levis, Matei Zaharia, Irene Zhang
SOSP11
2021 Breakfast of champions: towards zero-copy serialization with NIC scatter-gather
abstract
Microsecond I/O will make data serialization a major bottleneck for datacenter applications. Serialization is fundamentally about data movement: serialization libraries coalesce and flatten in-memory data structures into a single transmittable buffer. CPU-based serialization approaches will hit a performance limit due to data movement overheads and be unable to keep up with modern networks.
Deepti Raghavan, Philip Alexander Levis, Matei Zaharia, Irene Zhang
HotOS4
2021 PRISM: Rethinking the RDMA Interface for Distributed Systems
abstract
Remote Direct Memory Access (RDMA) has been used to accelerate a variety of distributed systems, by providing low-latency, CPU-bypassing access to a remote host's memory. However, most of the distributed protocols used in these systems cannot easily be expressed in terms of the simple memory READs and WRITEs provided by RDMA. As a result, designers face a choice between introducing additional protocol complexity (e.g., additional round trips) or forgoing the benefits of RDMA entirely.
Matthew Burke 0001, Sowmya Dharanipragada, Shannon Joyner, Adriana Szekeres, Jacob Nelson 0001, Irene Zhang, Dan R. K. Ports
SOSP6
2021 When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with Perséphone
abstract
This paper introduces Perséphone, a kernel-bypass OS scheduler designed to minimize tail latency for applications executing at microsecond-scale and exhibiting wide service time distributions. Perséphone integrates a new scheduling policy, Dynamic Application-aware Reserved Cores (DARC), that reserves cores for requests with short processing times. Unlike existing kernel-bypass schedulers, DARC is not work conserving. DARC profiles application requests and leaves a small number of cores idle when no short requests are in the queue, so when short requests do arrive, they are not blocked by longer-running ones. Counter-intuitively, leaving cores idle lets DARC maintain lower tail latencies at higher utilization, reducing the overall number of cores needed to serve the same workloads and consequently better utilizing the datacenter resources.
Henri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias, Boon Thau Loo, Linh T. X. Phan, Irene Zhang
SOSP7
2021 The Demikernel Datapath OS Architecture for Microsecond-scale Datacenter Systems
abstract
Datacenter systems and I/O devices now run at single-digit microsecond latencies, requiring ns-scale operating systems. Traditional kernel-based operating systems impose an unaffordable overhead, so recent kernel-bypass OSes [73] and libraries [23] eliminate the OS kernel from the I/O datapath. However, none of these systems offer a general-purpose datapath OS replacement that meet the needs of μs-scale systems.' [email protected] paper proposes Demikernel, a flexible datapath OS and architecture designed for heterogenous kernel-bypass devices and μs-scale datacenter systems. We build two prototype Demikernel OSes and show that minimal effort is needed to port existing μs-scale systems. Once ported, Demikernel lets applications run across heterogenous kernel-bypass devices with ns-scale overheads and no code changes.
Irene Zhang, Amanda Raybuck, Pratyush Patel, Kirk Olynyk, Jacob Nelson 0001, Omar S. Navarro Leija, Ashlie Martinez, Anna Kornfeld Simpson, Sujay Jayakar, Pedro Henrique de Mello Morado Penna, Max Demoulin, Piali Choudhury, Anirudh Badam
SOSP1
2020 Talek: Private Group Messaging with Hidden Access Patterns
abstract
Talek is a private group messaging system that sends messages through potentially untrustworthy servers, while hiding both data content and the communication patterns among its users. Talek explores a new point in the design space of private messaging; it guarantees access sequence indistinguishability, which is among the strongest guarantees in the space, while assuming an anytrust threat model, which is only slightly weaker than the strongest threat model currently found in related work. Our results suggest that this is a pragmatic point in the design space, since it supports strong privacy and good performance: we demonstrate a 3-server Talek cluster that achieves throughput of 9,433 messages/second for 32,000 active users with 1.7-second end-to-end latency. To achieve its security goals without coordination between clients, Talek relies on information-theoretic private information retrieval. To achieve good performance and minimize server-side storage, Talek introduces new techniques and optimizations that may be of independent interest, e.g., a novel use of blocked cuckoo hashing and support for private notifications. The latter provide a private, efficient mechanism for users to learn, without polling, which logs have new messages.
Raymond Cheng 0001, William Scott 0002, Elisaweta Masserova, Irene Zhang, Vipul Goyal, Thomas E. Anderson, Arvind Krishnamurthy, Bryan Parno
ACSAC4
2020 LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCs
abstract
Today's cloud storage stack is extremely resource hungry, burning 10-20% of datacenter x86 cores, a major "storage tax" that cloud providers must pay. Yet, the complex cloud storage stack is not completely offload-ready to today's IO accelerators. We present LeapIO, a new cloud storage stack that leverages ARM-based co-processors to offload complex storage services. LeapIO addresses many deployment challenges, such as hardware fungibility, software portability, virtualizability, composability, and efficiency. It uses a set of OS/software techniques and new hardware properties that provide a uni- form address space across the x86 and ARM cores and ex- pose virtual NVMe storage to unmodified guest VMs, at a performance that is competitive with bare-metal servers.
Huaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte, Sriram Govindan, Dan R. K. Ports, Irene Zhang, Ricardo Bianchini, Haryadi S. Gunawi, Anirudh Badam
ASPLOS7
2020 Meerkat: multicore-scalable replicated transactions following the zero-coordination principle
abstract
Traditionally, the high cost of network communication between servers has hidden the impact of cross-core coordination in replicated systems. However, new technologies, like kernel-bypass networking and faster network links, have exposed hidden bottlenecks in distributed systems.
Adriana Szekeres, Michael J. Whittaker, Jialin Li 0001, Naveen Kr. Sharma, Arvind Krishnamurthy, Dan R. K. Ports, Irene Zhang
EuroSys7
2020 Persistent State Machines for Recoverable In-memory Storage Systems with NVRam
Scott Shenker, Irene Zhang
OSDI3
2020 End the Senseless Killing: Improving Memory Management for Mobile Operating Systems
Niel Lebeck, Arvind Krishnamurthy, Henry M. Levy, Irene Zhang
USENIX ATC4
2020 A.M.B.R.O.S.I.A: Providing Performant Virtual Resiliency for Distributed Applications
abstract
When writing today's distributed programs, which frequently span both devices and cloud services, programmers are faced with complex decisions and coding tasks around coping with failure, especially when these distributed components are stateful. If their application can be cast as pure data processing, they benefit from the past 40--50 years of work from the database community, which has shown how declarative database systems can completely isolate the developer from the possibility of failure in a performant manner. Unfortunately, while there have been some attempts at bringing similar functionality into the more general distributed programming space, a compelling general-purpose system must handle non-determinism, be performant, support a variety of machine types with varying resiliency goals, and be language agnostic, allowing distributed components written in different languages to communicate. This paper introduces Ambrosia, the first system to satisfy all these requirements. We coin the term "virtual resiliency", analogous to virtual memory, for the platform feature which allows failure oblivious code to run in a failure resilient manner. We also introduce novel programming language constructs for resiliently handling non-determinism. Of further interest is the effective reapplication of much database performance optimization technology to make Ambrosia more performant than many of today's non-resilient cloud solutions.
Jonathan Goldstein, Ahmed S. Abdelhamid, Michael Barnett 0001, Sebastian Burckhardt, Badrish Chandramouli, Darren Gehring, Niel Lebeck, Christopher Meiklejohn, Umar Farooq Minhas, Ryan Newton, Rahee Peshawaria, Tal Zaccai, Irene Zhang
Proc. VLDB Endow.13
2019 TMC: Pay-as-you-Go Distributed Communication
abstract
We revisit the gap between what distributed systems need from the transport layer and what protocols in wide deployment provide. Such a gap complicates the implementation of distributed systems and impacts their performance. We introduce Tunable Multicast Communication (TMC), an abstraction that allows developers to easily specialize communication channels in distributed systems. TMC is presented as a deployable and extensible user-space library that exposes high-level tunable guarantees. TMC has the potential of improving the performance of distributed applications with minimal-to-zero development and deployment effort.
Henri Maxime Demoulin, Nikos Vasilakis, John Sonchack, Isaac Pedisich, Vincent Liu 0001, Boon Thau Loo, Linh T. X. Phan, Jonathan M. Smith, Irene Zhang
APNet9
2019 I'm Not Dead Yet!: The Role of the Operating System in a Kernel-Bypass Era
abstract
Researchers have long predicted the demise of the operating system [21, 26, 41]. As datacenter servers increasingly incorporate I/O devices that let applications bypass the OS kernel (e.g., RDMA [12] and DPDK [15] network devices or SPDK storage devices), this prediction may finally come true. While kernel-bypass devices do eliminate the OS kernel from the I/O path, they do not handle the kernel's most important job: offering higher-level abstractions. This paper argues for a new high-level, device-agnostic I/O abstraction for kernel-bypass devices. We propose the Demikernel, a new library OS architecture for kernel-bypass devices. It defines a high-level, kernel-bypass I/O abstraction and provides user-space library OSes to implement that abstraction across a range of kernel-bypass devices. The Demikernel makes applications easier to build, portable across devices, and unmodified as devices continue to evolve.
Irene Zhang, Jing Liu 0074, Amanda Austin, Michael Lowell Roberts, Anirudh Badam
HotOS1
2017 Building Consistent Transactions with Inconsistent Replication
abstract
Application programmers increasingly prefer distributed storage systems with strong consistency and distributed transactions (e.g., Google’s Spanner) for their strong guarantees and ease of use. Unfortunately, existing transactional storage systems are expensive to use—in part, because they require costly replication protocols, like Paxos, for fault tolerance. In this article, we present a new approach that makes transactional storage systems more affordable: We eliminate consistency from the replication protocol, while still providing distributed transactions with strong consistency to applications. We present the Transactional Application Protocol for Inconsistent Replication (TAPIR), the first transaction protocol to use a novel replication protocol, called inconsistent replication , that provides fault tolerance without consistency. By enforcing strong consistency only in the transaction protocol, TAPIR can commit transactions in a single round-trip and order distributed transactions without centralized coordination. We demonstrate the use of TAPIR in a transactional key-value store, TAPIR-KV . Compared to conventional systems, TAPIR-KV provides better latency and better throughput.
Irene Zhang, Naveen Kr. Sharma, Adriana Szekeres, Arvind Krishnamurthy, Dan R. K. Ports
ACM Trans. Comput. Syst.1
2016 Disciplined Inconsistency with Consistency Types
abstract
Distributed applications and web services, such as online stores or social networks, are expected to be scalable, available, responsive, and fault-tolerant. To meet these steep requirements in the face of high round-trip latencies, network partitions, server failures, and load spikes, applications use eventually consistent datastores that allow them to weaken the consistency of some data. However, making this transition is highly error-prone because relaxed consistency models are notoriously difficult to understand and test.
Brandon Holt, James Bornholt, Irene Zhang, Dan R. K. Ports, Mark Oskin, Luis Ceze
SoCC3
2016 Diamond: Automating Data Management and Storage for Wide-Area, Reactive Applications
Irene Zhang, Niel Lebeck, Pedro Fonseca 0001, Brandon Holt, Raymond Cheng 0001, Ariadna Norberg, Arvind Krishnamurthy, Henry M. Levy
OSDI1
2016 Arrakis: The Operating System Is the Control Plane
abstract
Recent device hardware trends enable a new approach to the design of network server operating systems. In a traditional operating system, the kernel mediates access to device hardware by server applications to enforce process isolation as well as network and disk security. We have designed and implemented a new operating system, Arrakis, that splits the traditional role of the kernel in two. Applications have direct access to virtualized I/O devices, allowing most I/O operations to skip the kernel entirely, while the kernel is re-engineered to provide network and disk protection without kernel mediation of every operation. We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2 to 5 × in latency and 9 × throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation.
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe
ACM Trans. Comput. Syst.3
2015 Building consistent transactions with inconsistent replication
abstract
Application programmers increasingly prefer distributed storage systems with strong consistency and distributed transactions (e.g., Google's Spanner) for their strong guarantees and ease of use. Unfortunately, existing transactional storage systems are expensive to use -- in part because they require costly replication protocols, like Paxos, for fault tolerance. In this paper, we present a new approach that makes transactional storage systems more affordable: we eliminate consistency from the replication protocol while still providing distributed transactions with strong consistency to applications.
Irene Zhang, Naveen Kr. Sharma, Adriana Szekeres, Arvind Krishnamurthy, Dan R. K. Ports
SOSP1
2014 Towards High-Performance Application-Level Storage Management
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Thomas E. Anderson, Arvind Krishnamurthy, Mark Zbikowski, Doug Woos
HotStorage3
2014 Arrakis: The Operating System is the Control Plane
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe
OSDI3
2014 Customizable and Extensible Deployment for Mobile/Cloud Applications
Irene Zhang, Adriana Szekeres, Dana Van Aken, Isaac Ackerman, Steve D. Gribble, Arvind Krishnamurthy, Henry M. Levy
OSDI1
2013 Optimizing VM Checkpointing for Restore Performance in VMware ESXi
Irene Zhang, Tyler Denniston, Yury Baskakov, Alex Garthwaite
USENIX ATC1
2011 Fast restore of checkpointed memory using working set estimation
abstract
In order to make save and restore features practical, saved virtual machines (VMs) must be able to quickly restore to normal operation. Unfortunately, fetching a saved memory image from persistent storage can be slow, especially as VMs grow in memory size. One possible solution for reducing this time is to lazily restore memory after the VM starts. However, accesses to unrestored memory after the VM starts can degrade performance, sometimes rendering the VM unusable for even longer. Existing performance metrics do not account for performance degradation after the VM starts, making it difficult to compare lazily restoring memory against other approaches. In this paper, we propose both a better metric for evaluating the performance of different restore techniques and a better scheme for restoring saved VMs.
Irene Zhang, Alex Garthwaite, Yury Baskakov, Kenneth C. Barr
VEE1
2010 Transactional Consistency and Automatic Management in an Application Data Cache
Dan R. K. Ports, Austin T. Clements, Irene Zhang, Samuel Madden 0001, Barbara Liskov
OSDI3
2009 Flexible, Wide-Area Storage for Distributed Systems with WheelFS
Jeremy Stribling, Yair Sovran, Irene Zhang, Xavid Pretzer, Jinyang Li 0001, M. Frans Kaashoek, Robert Morris 0005
NSDI3
1995 Delay performance evaluation of high speed protocols for multimedia communications
abstract
Multimedia applications integrate a variety of media namely, audio, video, images, graphics, text, and data, each of which has different quality of service (QoS) requirements. To study the support of multimedia traffic on high speed protocols, we examine two significant metrics: delay fairness and worst-case delay performance. Four members of reservation-based high speed protocols are studied DQDB (distributed queue dual bus), CRMA (cyclic reservation multiple access), DQMA (distributed queue multiple access), and FDQ (fair distributed queue). The first part of this work presents delay fairness study of the protocols under various input traffic. Both access delay and message delay experienced by individual network nodes are measured to illustrate their fairness performance. The second part conducts worst-case delay experiments. Too worst-case scenarios of a test message are studied. The delay experienced by this message under varied inter-node distance is measured. The simulation results show that DQDB has the worst fairness performance. CRMA and DQMA present mixed results. FDQ proves itself to be a very fair protocol. The results also suggest that to there is no singe metric to justify a new protocol. Comparing our results with previous results on throughput evaluation of multimedia traffic support, we found that the two results are strongly correlated the fairer a protocol is, the better it is for supporting heterogeneous traffic under heavy network load.
Melody Moh, Yu-Jen Chien, Irene Zhang, Teng-Sheng Moh
ICCCN3