Gefei Zuo

dblp:220/6764 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-0861-5010ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Power Sloshing in Compound Servers for Large-Scale AI Inference Workloads
Albert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia, Sultan Mahmud Sajal, Gefei Zuo, Krishna T. Malladi, Devon Akers, Kalyan Subramanian, Shobhit O. Kanaujia, Alexandros Daglis
ISCA6
2025 NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yilong Zhao 0002, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye 0001, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci
OSDI5
2023 Vidi: Record Replay for Reconfigurable Hardware
abstract
Developers are turning to heterogeneous computing devices, such as Field Programmable Gate Arrays (FPGAs), to accelerate data center and cloud computing workloads. FPGAs enable rapid prototyping and should facilitate an agile software-like development workflow to fix correctness bugs, performance issues, and security vulnerabilities. Unfortunately, hardware development still does not have a vast ecosystem of tools needed to support the agile hardware development vision. The capability to record and replay FPGA executions would constitute a key building block that will inspire the development of many tools, similar to what record/replay did for software. However, building a practical record/replay tool for FPGA is challenging; existing approaches either record too much or too little information and cannot support real-world executions.
Gefei Zuo, Jiacheng Ma 0001, Andrew Quinn 0001, Baris Kasikci
ASPLOS (3)1
2022 Debugging in the brave new world of reconfigurable hardware
abstract
Software and hardware development cycles have traditionally been quite distinct. Software allows post-deployment patches, which leads to a rapid development cycle. In contrast, hardware bugs that are found after fabrication are extremely costly to fix (and sometimes even unfixable), so the traditional hardware development cycle involves massive investment in extensive simulation and formal verification. Reconfigurable hardware, such as a Field Programmable Gate Array (FPGA), promises to propel hardware development towards an agile software-like development approach, since it enables a hardware developer to patch bugs that are detected during on-chip testing or in production. Unfortunately, FPGA programmers lack bug localization tools amenable to this rapid development cycle, since past tools mainly find bugs via simulation and verification. To develop hardware bug localization tools for a rapid development cycle, a thorough understanding of the symptoms, root causes, and fixes of hardware bugs is needed.
Jiacheng Ma 0001, Gefei Zuo, Kevin Loughlin, Andrew Quinn 0001, Baris Kasikci
ASPLOS2
2021 Rethinking File Mapping for Persistent Memory
Ian Neal, Gefei Zuo, Eric Shiple, Tanvir Ahmed Khan 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci
FAST2
2021 Execution reconstruction: harnessing failure reoccurrences for failure reproduction
abstract
Reproducing production failures is crucial for software reliability. Alas, existing bug reproduction approaches are not suitable for production systems because they are not simultaneously efficient, effective, and accurate. In this work, we survey prior techniques and show that existing approaches over-prioritize a subset of these properties, and sacrifice the remaining ones. As a result, existing tools do not enable the plethora of proposed failure reproduction use-cases (e.g., debugging, security forensics, fuzzing) for production failures.
Gefei Zuo, Jiacheng Ma 0001, Andrew Quinn 0001, Pramod Bhatotia, Pedro Fonseca 0001, Baris Kasikci
PLDI1
2021 1Pipe: scalable total order communication in data center networks
abstract
This paper proposes 1Pipe, a novel communication abstraction that enables different receivers to process messages from senders in a consistent total order. More precisely, 1Pipe provides both unicast and scattering (i.e., a group of messages to different destinations) in a causally and totally ordered manner. 1Pipe provides a best effort service that delivers each message at most once, as well as a reliable service that guarantees delivery and provides restricted atomic delivery for each scattering. 1Pipe can simplify and accelerate many distributed applications, e.g., transactional key-value stores, log replication, and distributed data structures.
Bojie Li, Gefei Zuo, Wei Bai 0001
SIGCOMM2
2020 A Hypervisor for Shared-Memory FPGA Platforms
abstract
Cloud providers widely deploy FPGAs as application-specific accelerators for customer use. These providers seek to multiplex their FPGAs among customers via virtualization, thereby reducing running costs. Unfortunately, most virtualization support is confined to FPGAs that expose a restrictive, host-centric programming model in which accelerators cannot issue direct memory accesses (DMAs). The host-centric model incurs high runtime overhead for workloads that exhibit pointer chasing. Thus, FPGAs are beginning to support a shared-memory programming model in which accelerators can issue DMAs. However, virtualization support for shared-memory FPGAs is limited. This paper presents Optimus, the first hypervisor that supports scalable shared-memory FPGA virtualization. Optimus offers both spatial multiplexing and temporal multiplexing to provide efficient and flexible sharing of each accelerator on an FPGA. To share the FPGA-CPU interconnect at a high clock frequency, Optimus implements a multiplexer tree. To isolate each guest's address space, Optimus introduces the technique of page table slicing as a hardware-software co-design. To support preemptive temporal multiplexing, Optimus provides an accelerator preemption interface. We show that Optimus supports eight physical accelerators on a single FPGA and improves the aggregate throughput of twelve real-world benchmarks by 1.98x-7x.
Jiacheng Ma 0001, Gefei Zuo, Kevin Loughlin, Xiaohe Cheng, Yanqiang Liu, Abel Mulugeta Eneyew, Zhengwei Qi, Baris Kasikci
ASPLOS2