Xiangfeng Zhu

dblp:233/6265 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0009-7557-2090ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Programmable and Adaptive Scheduling for Distributed Systems
abstract
Existing frameworks for managing distributed systems hard-code scheduling policies and their implementations (e.g., centralized vs. decentralized), limiting customization and hurting performance across diverse applications and workloads. We argue for an adaptive scheduling approach, where developers express policies in a high-level, framework-agnostic DSL, and a compiler generates optimized implementations based on policy semantics, workload characteristics, and execution environments. We demonstrate that our compiler-guided approach can significantly improve both scheduling quality and performance.
Xiangfeng Zhu, Ratul Mahajan, Stephanie Wang
HotNets2
2025 Rethinking RPC Communication for Microservices-based Applications
abstract
Fast and efficient RPCs are key to the performance of applications based on microservices. But RPC communication suffers from significant overhead today because it relies on the standard, layered protocol stack and loose coupling between the end host and in-network proxies that process RPCs. We propose delayering the RPC communication stack and tightly coupling the end host and in-network processing using high-level abstractions. This approach leads to more efficient and performant RPC communication because it eliminates many sources of overhead.
Xiangfeng Zhu, Arvind Krishnamurthy, Sam Kumar, Ratul Mahajan, Danyang Zhuo
HotOS1
2025 High-level Programming for Application Networks
Xiangfeng Zhu, Banruo Liu, Yongtong Wu, Nikola Bojanic, Jingrong Chen 0002, Gilbert Louis Bernstein, Arvind Krishnamurthy, Sam Kumar, Ratul Mahajan, Danyang Zhuo
NSDI1
2025 Netherite: efficient execution of serverless workflows
Sebastian Burckhardt, Badrish Chandramouli, Chris Gillum, David Justo, Konstantinos Kallas, Connor McMahon, Christopher Meiklejohn, Xiangfeng Zhu
VLDB J.8
2023 Dissecting Overheads of Service Mesh Sidecars
abstract
Service meshes play a central role in the modern application ecosystem by providing an easy and flexible way to connect microservices of a distributed application. However, because of how they interpose on application traffic, they can substantially increase application latency and its resource consumption. We develop a tool called MeshInsight to help developers quantify the overhead of service meshes in deployment scenarios of interest and make informed trade-offs about their functionality vs. overhead. Using MeshInsight, we confirm that service meshes can have high overhead---up to 269% higher latency and up to 163% more virtual CPU cores for our benchmark applications---but the severity is intimately tied to how they are configured and the application workload. IPC (inter-process communication) and socket writes dominate when the service mesh operates as a TCP proxy, but protocol parsing dominates when it operates as an HTTP proxy. MeshInsight also enables us to study the end-to-end impact of optimizations to service meshes. We show that not all seemingly-promising optimizations lead to a notable overhead reduction in realistic settings.
Xiangfeng Zhu, Guozhen She, Yu Zhang 0209, Yongsu Zhang, Xuan Kelvin Zou, Xiongchun Duan, Peng He 0003, Arvind Krishnamurthy, Matthew Lentz, Danyang Zhuo, Ratul Mahajan
SoCC1
2023 Application Defined Networks
abstract
With the rise of microservices, the execution environment of many cloud applications has become a set of virtual machines or containers connected by a flexible and feature-rich virtual network. We argue that the implementation of such virtual networks should be completely application-specific and not layered on top of general-purpose network abstractions from the Internet age. Such layering tends to more than double the latency and CPU usage of applications. We propose application-defined networks in which developers specify network functionality in a high-level language and a controller generates a custom distributed implementation that runs across available hardware and software resources. Experiments with a preliminary prototype suggest that, compared to the state of the art, ADN reduces latency by up to 20x and increases the throughput by up to 6x.
Xiangfeng Zhu, Weixin Deng, Banruo Liu, Jingrong Chen 0002, Thomas E. Anderson, Arvind Krishnamurthy, Ratul Mahajan, Danyang Zhuo
HotNets1
2022 FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
abstract
We present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging from image classification and object detection to language modeling and speech recognition. Each dataset comes with a unified evaluation protocol using real-world data splits and evaluation metrics. To reproduce realistic FL behavior, FedScale contains a scalable and extensible runtime. It provides high-level APIs to implement FL algorithms, deploy them at scale across diverse hardware and software backends, and evaluate them at scale, all with minimal developer efforts. We combine the two to perform systematic benchmarking experiments and highlight potential opportunities for heterogeneity-aware co-optimizations in FL. FedScale is open-source and actively maintained by contributors from different institutions at http://fedscale.ai. We welcome feedback and contributions from the community.
Fan Lai 0001, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury
ICML5
2022 Netherite: Efficient Execution of Serverless Workflows
abstract
Serverless is a popular choice for cloud service architects because it can provide scalability and load-based billing with minimal developer effort. Functions-as-a-service (FaaS) are originally stateless, but emerging frameworks add stateful abstractions. For instance, the widely used Durable Functions (DF) allow developers to write advanced serverless applications, including reliable workflows and actors, in a programming language of choice. DF implicitly and continuosly persists the state and progress of applications, which greatly simplifies development, but can create an IOps bottleneck. To improve efficiency, we introduce Netherite, a novel architecture for executing serverless workflows on an elastic cluster. Netherite groups the numerous application objects into a smaller number of partitions, and pipelines the state persistence of each partition. This improves latency and throughput, as it enables workflow steps to group commit, even if causally dependent. Moreover, Netherite leverages FASTER's hybrid log approach to support larger-than-memory application state, and to enable efficient partition movement between compute hosts. Our evaluation shows that (a) Netherite achieves lower latency and higher throughput than the original DF engine, by more than an order of magnitude in some cases, and (b) that Netherite has lower latency than some commonly used alternatives, like AWS Step Functions or cloud storage triggers.
Sebastian Burckhardt, Badrish Chandramouli, Chris Gillum, David Justo, Konstantinos Kallas, Connor McMahon, Christopher Meiklejohn, Xiangfeng Zhu
Proc. VLDB Endow.8
2021 Oort: Efficient Federated Learning via Guided Participant Selection
Fan Lai 0001, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury
OSDI2
2020 Sol: Fast Distributed Computation Over Slow Networks
Fan Lai 0001, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury
NSDI3
2019 Fixed It For You: Protocol Repair Using Lineage Graphs
Lennart Oldenburg, Xiangfeng Zhu, Kamala Ramasubramanian, Peter Alvaro
CIDR2