Kefei Liu 0004

dblp:286/4112 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0008-5874-6610ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 5 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Skyline: A Cloud Centric Internet Monitoring Engine
Shixian Guo, Yangyang Bai, Kefei Liu 0004, Zhenyang Zhong, Sisi Wen, Yongbin Dong, Anjian Chen, Jiale Feng, Lingpei Meng, Siwan Chen, Juntao Zhong, Chaoran Hu, Yibo Huang 0005, Yiming Qiu 0001
NSDI5
2026 Horizon: A Hyper-Edge Observability Engine for Live Streaming Networks
abstract
Live streaming services power mainstream real-time interactions on top of dedicated live streaming networks (LiveNets). Yet making LiveNets reliable at scale is challenging: failures arise on the userfacing delivery path and within streaming protocol and application logic, so operators need both continuous runtime monitoring to detect and localize incidents quickly and proactive preflight testing to exercise changes under representative environments and sustained playback behavior. Meeting these goals hinges on the right vantage point: the observability workflow must traverse the same network paths and delivery stacks as users while remaining controllable and non-intrusive. We present Horizon, which leverages near-user, provider-managed hyper-edge devices and orchestrates them into a shared fleet that supports both always-on monitoring and customizable, scenario-driven validation. Horizon has been deployed in production for over three years; in 2025, it identified 2,000+ major network incidents using 100,000+ hyper-edge agents.
Daqian Ding, Shixian Guo, Zhendong Xie, Aifang Xu, Changqian Wang, Kefei Liu 0004, Jialin Li 0001, Yunming Xiao, Heming Cui, Yiming Qiu 0001
SIGCOMM8
2025 ByteTracker: An Agentless and Real-time Path-aware Network Probing System
abstract
As the number of data center servers grows into the millions and due to the demand for more accurate, rapid and powerful network fault detection and location, the existing Pingmesh-centric monitoring and diagnostic system is not efficient enough. In this paper, we propose ByteTracker, the first agentless probing and diagnostic system for large-scale data center networks. It does not need to deploy probe processes or make any configurations on end hosts, and all probes are launched by a small number of centralized Probers. ByteTracker achieves accurate, real-time probe path tracking with packet mirroring on switches. By reducing end-host probe noise, precisely identifying network timeout probes, accurately tracking probe paths, and marking the failed switch with multiple network timeout probes, ByteTracker can locate network failures with nearly 100% accuracy. We have deployed ByteTracker in all of our data centers for over half a year. During deployment, ByteTracker can detect almost all network anomalies and locate them within 5 seconds with 100% accuracy.
Shixian Guo, Kefei Liu 0004, Yulin Lai, Yangyang Bai, Jianghang Ning, Yongbin Dong, Sisi Wen, Jiale Feng, Chengcai Yao, Zhuo Jiang, Jiao Zhang 0002, Tao Huang 0005
SIGCOMM2
2025 ByteTuning: Watermark Tuning for RoCEv2
abstract
RDMA over Converged Ethernet v2 (RoCEv2) is one of the most popular high-speed datacenter networking solutions. Watermark is the general term for various trigger and release thresholds of RoCEv2 flow control protocols, and its reasonable configuration is an important factor affecting RoCEv2 performance. In this paper, we propose ByteTuning, a centralized watermark tuning system for RoCEv2. First, three real cases of network performance degradation caused by non-optimal or improper watermark configuration are reported, and the network performance results of different watermark configurations in three typical scenarios are traversed, indicating the necessity of watermark tuning. Then, based on the RDMA Fluid model, the influence of watermark on the RoCEv2 performance is modeled and evaluated. Next, the design of the ByteTuning is introduced, which includes three mechanisms. They are (1) using simulated annealing algorithm to make the real-time watermark converge to the near-optimal configuration, (2) using network telemetry to optimize the feedback overhead, (3) compressing the search space to improve the tuning efficiency. Finally, We validate the performance of ByteTuning in multiple real datacenter networking environments, and the results show that ByteTuning outperforms existing solutions.
Lizhuang Tan, Zhuo Jiang, Kefei Liu 0004, Pengfei Huo, Huiling Shi, Wei Zhang 0049, Wei Su 0006
IEEE Trans. Cloud Comput.3
2025 RHCC: Revisiting Intra-Host Congestion Control in RDMA Networks
abstract
RDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with CPU-to-memory traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to revisit the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines CPU-to-memory traffic congestion avoidance with sub-RTT granularity and proactive RNIC traffic adjustment. RHCC ensures fast congestion avoidance and can work with different inter-host congestion control methods. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2$\times$and 1.4$\times$, respectively.
Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Yongchen Pan, Tao Huang 0005
IEEE Trans. Netw.4
2024 Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE Clusters
abstract
RoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary.
Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Zicheng Wang 0004, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
APNet1
2024 Rethinking Intra-host Congestion Control in RDMA Networks
abstract
RDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with intra-host traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to rethink the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines sub-RTT granularity intra-host traffic congestion avoidance and proactive RNIC traffic adjustment. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2 × and 1.4 ×, respectively.
Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Tao Huang 0005
APNet4
2024 R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic System
abstract
RoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related.
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Haohan Xu, Dongyang Song, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
SIGCOMM1
2024 Diagnosing End-Host Network Bottlenecks in RDMA Servers
abstract
In RDMA (Remote Direct Memory Access) networks, end-host networks, including intra-host networks and RNICs (RDMA NIC), were considered robust and have received little attention. However, as the RNIC line rate rapidly increases to multi-hundred gigabits, the intra-host network becomes a potential performance bottleneck for network applications. Intra-host network bottlenecks can result in degraded intra-host bandwidth and increased intra-host latency. In addition, RNIC network problems can result in connection failures and packet drops. Host network problems can severely degrade network performance. However, when host network problems occur, they can hardly be noticed due to the lack of a monitoring system. Furthermore, existing diagnostic mechanisms cannot efficiently diagnose host network problems. In this paper, we analyze the symptom of host network problems based on our long-term troubleshooting experience and propose Hostping, the first monitoring and diagnostic system dedicated to host networks. The core idea of Hostping is to conduct 1) loopback tests between RNICs and endpoints within the host to measure intra-host latency and bandwidth, and 2) mutual probing between RNICs on a host to measure RNIC connectivity. We have deployed Hostping on thousands of servers in our distributed machine learning system. Not only can Hostping detect and diagnose host network problems we already knew in minutes, but it also reveals eight problems we did not notice before.
Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Xiaolong Zhong, Lizhuang Tan, Tian Pan 0001, Tao Huang 0005
IEEE/ACM Trans. Netw.1
2023 Hostping: Diagnosing Intra-host Network Bottlenecks in RDMA Servers
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Xiaolong Zhong, Lizhuang Tan, Tian Pan 0001, Tao Huang 0005
NSDI1
2020 PLB: Adaptive Partial Congestion-aware Load Balancing for Datacenter Networks
abstract
In order to accommodate ever-increasing new tenants and applications, datacenter networks (DCNs) require an efficient load balancing scheme to fully utilize their bisection bandwidth. Equal-cost MultiPath routing (ECMP) is a widely used load-balancing mechanism in the DCN. However, ECMP blindly hashes traffic to parallel paths and results in imbalance and collisions. Motivated by ECMP's shortcomings, some recent schemes provide more visibility into networks via active probing. They could be broadly classified as probing all the paths or a fixed number of paths (e.g., 3 paths) each probe interval. However, they all suffer from some limitations. Probing all paths introduces high probing overhead while probing a fixed number of paths is suboptimal when the network topology and traffic load change. To our best knowledge, none of the existing schemes adapt the number of paths being probed to the network conditions. Enlightened by the defects of previous work, we introduce PLB, an adaptive partial congestion-aware load-balancing mechanism. At its heart, PLB randomly probes partial paths each probe interval and the number of them changes according to the network topology and the traffic load. Besides, PLB splits flow into flowlets and makes careful routing/rerouting decisions for them. Through analysis, we formulate the correlations between the number of paths being probed and the network conditions. Furthermore, simulations with realistic workloads validate our conclusions and show that PLB reduces overall flow completion times compared to the state-of-the-art load balancing schemes both in symmetric and asymmetric topologies.
Kefei Liu 0004, Jiao Zhang 0002, Dehui Wei, Tao Huang 0005
GLOBECOM1