Qizhe Cai

dblp:217/1009 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Understanding Host Network Stack Latency
Tianyu Zuo, Jae-Hyun Hwang, Ao Tang, Rachit Agarwal 0001, Qizhe Cai
SIGCOMM5
2024 Harmony: A Congestion-free Datacenter Architecture
Saksham Agarwal, Qizhe Cai, Rachit Agarwal 0001, David B. Shmoys, Amin Vahdat
NSDI2
2024 High-throughput and Flexible Host Networking for Accelerated Computing
Athinagoras Skiadopoulos, Mark Zhao, Qizhe Cai, Saksham Agarwal, Jacob Adelmann, David Ahern, Carlo Contavalli, Michael D. Goldflam, Vitaly Mayatskikh, Raghu Raja, Daniel Walton, Rachit Agarwal 0001, Shrijeet Mukherjee, Christoforos E. Kozyrakis
OSDI4
2024 Fast & Safe IO Memory Protection
abstract
IO Memory protection mechanisms prevent malicious and/or buggy IO devices from executing errant transfers into memory. Modern servers achieve this using an IOMMU---IO devices operate on virtual addresses, and IOMMU translates virtual addresses to physical addresses (potentially speeding up translations using a cache called IOTLB) before executing memory transfers. Despite their importance, design of memory protection mechanisms that can provide strong safety properties while achieving high performance has remained elusive. Indeed, recent studies from production datacenters demonstrate that inefficiencies within state-of-the-art memory protection mechanisms result in significant throughput degradation, orders-of-magnitude tail latency inflation, and violation of isolation guarantees.
Benny Rubin, Saksham Agarwal, Qizhe Cai, Rachit Agarwal 0001
SOSP3
2022 dcPIM: near-optimal proactive datacenter transport
abstract
Datacenter Parallel Iterative Matching (dcPIM) is a proactive data-center transport design that simultaneously achieves near-optimal tail latency for short flows and near-optimal network utilization, without requiring any specialized network hardware.
Qizhe Cai, Mina Tahmasbi Arashloo, Rachit Agarwal 0001
SIGCOMM1
2022 Towards μs tail latency and terabit ethernet: disaggregating the host network stack
abstract
Dedicated, tightly integrated, and static packet processing pipelines in today's most widely deployed network stacks preclude them from fully exploiting capabilities of modern hardware.
Qizhe Cai, Midhul Vuppalapati, Jae-Hyun Hwang, Christoforos E. Kozyrakis, Rachit Agarwal 0001
SIGCOMM1
2021 Understanding host network stack overheads
abstract
Traditional end-host network stacks are struggling to keep up with rapidly increasing datacenter access link bandwidths due to their unsustainable CPU overheads. Motivated by this, our community is exploring a multitude of solutions for future network stacks: from Linux kernel optimizations to partial hardware offload to clean-slate userspace stacks to specialized host network hardware. The design space explored by these solutions would benefit from a detailed understanding of CPU inefficiencies in existing network stacks.
Qizhe Cai, Shubham Chaudhary 0004, Midhul Vuppalapati, Jae-Hyun Hwang, Rachit Agarwal 0001
SIGCOMM1
2020 TCP ≈ RDMA: CPU-efficient Remote Storage Access with i10
Jae-Hyun Hwang, Qizhe Cai, Ao Tang, Rachit Agarwal 0001
NSDI2