EDBT 2026 Demo / reviewers in the wild / expert
Kaihui Gao
dblp:249/7745
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0001-9013-5993ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 7 first-author · 16 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001 |
APNet | 3 |
| 2026 | CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter Networks
Tianfeng Liu, Kaihui Gao, Li Chen 0008, Dan Li 0001, Jin Guang, Vincent Liu 0001, Yiwei Zhang 0016, Ni Jin |
NSDI | 2 |
| 2026 | Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding
Kaihui Gao, Li Chen 0008, Dan Li 0001, Yiwei Zhang 0016, Fei Gui, Yitao Xing, Wenjia Wei, Bingyang Liu |
NSDI | 2 |
| 2026 | PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload ReallocationabstractModern collective communication libraries (CCLs) execute a collective communication task (CCT) by decomposing it into multiple sub-tasks, each mapped to a specific Virtual Topology (VT), which is an ordered graph of GPUs (e.g., a ring or a tree), to maximize parallelism and link utilization. As AI training scales to larger clusters, network anomalies (congestion and failures) are unavoidable, and a single straggling VT can delay the entire CCT. Existing solutions either rely on low-level transport-layer solutions which lacks a cross-sub-task perspective, or static CCL scheduling, failing to adapt to the dynamic and heterogeneous networks. Kaihui Gao, Li Chen 0008, Fei Gui, Dan Li 0001, Jiamin Cao |
SIGCOMM | 2 |
| 2026 | Networked Agent Memory and Causality Representation: Experiences towards Interpretable Cloud-Scale Root-Causing
Yanyu Ren, Xianshang Lin, Chenxu Wang 0007, Li Chen 0008, Shuai Wang 0028, Kaihui Gao, Dan Li 0001, Chen Tian 0001, Yunguang Li, Ennan Zhai |
SIGCOMM | 6 |
| 2025 | Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel Simulation
Fei Gui, Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Hongbing Yang, Dian Xiong |
NSDI | 2 |
| 2025 | From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model TrainingabstractThe development of large language models (LLMs) poses new challenges in data center network topology design. To assist in exploring topology design, we propose ATOP, an Automated Topology Optimization Pipeline, which models network topology as a set of hyperparameters, enabling the discovery of potential topologies. With various optimization algorithms and customizable optimization objectives, ATOP achieves automated topology optimization on a scale of tens of thousands of GPUs. We apply ATOP on network topologies for 256, 1024, 4096, and 16384 GPUs, optimizing performance under LLMs training traffic patterns, collective communication performance, fault tolerance, and network cost. We also evaluate ATOP in different scenarios: building, optimizing, and expanding a data center. From ATOP's results, we discover a new topology — ZCube, which reaches the highest cost-effectiveness across various GPU scales. Simulation results show that ZCube, compared to the previous state-of-the-art topologies, including Rail-optimized Fat-tree (ROFT), Rail-only, and HPN, improves end-to-end LLM training speed by 3% to 7% and reduces network hardware costs by 26% to 46%. We also construct ZCube on a real-world testbed. Results show that ZCube reduces hardware costs by 25% compared to Rail-Optimized Topology while maintaining the same all-reduce and all-to-all performance. Dan Li 0001, Li Chen 0008, Dian Xiong, Kaihui Gao, Yiwei Zhang 0016, Menglei Zhang, Bochun Zhang, Zhuo Jiang, Jianxi Ye, Haibin Lin |
SIGCOMM | 5 |
| 2024 | RedTE: Mitigating Subsecond Traffic Bursts with Real-time and Distributed Traffic EngineeringabstractInternet traffic bursts usually happen within a second, thus conventional burst mitigation methods ignore the potential of Traffic Engineering (TE). However, our experiments indicate that a TE system, with a sub-second control loop latency, can effectively alleviate burst-induced congestion. TE-based methods can leverage network-wide tunnel-level information to make globally informed decisions (e.g., balancing traffic bursts among multiple paths). Our insight in reducing control loop latency is to let each router make local TE decisions, but this introduces the key challenge of minimizing performance loss compared to centralized TE systems. Fei Gui, Dan Li 0001, Li Chen 0008, Kaihui Gao, Congcong Min, Yi Wang 0004 |
SIGCOMM | 5 |
| 2023 | Demo: NetVision: Efficient Visualization Front-End for Packet-level Discrete-Event Network SimulationabstractVisualization of network simulation is an essential tool for network practitioners. However, the front-end of existing network simulators often fails to deliver satisfactory performance when dealing with modern network scales and interface speed. In this paper, we propose NetVision, an efficient visualization front-end of network simulation based on the Unity engine, which is commonly used for video game and virtual reality development. NetVision offers flow-level visualization of network behavior and performances. Then, through parallel optimization, NetVision supports real-time visualization for large-scale high-speed networks. Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016 |
SIGCOMM | 1 |
| 2023 | DONS: Fast and Affordable Discrete Event Network Simulation with Automatic ParallelizationabstractDiscrete Event Simulation (DES) is an essential tool for network practitioners. Unfortunately, existing DES simulators cannot achieve satisfactory performance at the scale of modern networks. Recent work has attempted to address these challenges by reducing the traffic processed via novel approximation techniques; however, we argue in this paper that much of the slowdown of existing DES simulators is due to their underlying software architecture. Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016 |
SIGCOMM | 1 |
| 2023 | Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the CloudabstractRequest latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline. Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Bandwidth-efficient Microburst Measurement in Large-scale Datacenter NetworksabstractMicroburst measurement is essential for diagnosing and mitigating performance problems in datacenter networks. The key is to efficiently identify the flows that contribute the most to queue buildup. However, because the existing microburst measurement systems capture packet-level information, they incur significant bandwidth overhead. We present BurstScope, a bandwidth-efficient microburst measurement system that can profile the microburst characteristics and the contributing flows. BurstScope detects the microburst-involved packets in egress pipeline, then aggregates the measurement granularity from packet level to flow level by an invertible sketch. Finally, by carefully partitioning the measurement and statistic tasks between the data and control plane, we generate only one telemetry packet for each microburst. We have implemented BurstScope on Barefoot Tofino switches. Testbed-based evaluations show that BurstScope keeps low bandwidth overhead (< 0.02%) and high identification accuracy (> 97%). Compared with the state-of-the-art system, BurstScope can reduce 60 × bandwidth overhead. Kaihui Gao, Dan Li 0001, Shuai Wang 0028 |
APNet | 1 |
| 2022 | Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005 |
NSDI | 1 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 2 |
| 2021 | Fast and Robust Online Traffic Classification Supporting Unseen ApplicationsabstractOnline traffic classification is a fundamental toolkit in network management, such as QoS and network security. The speed and generalization ability of online classification are two requirements that need to be satisfied simultaneously. However, existing methods may suffer from generalization degradation on the traffic with unseen applications which are constantly emerging in the network, due to the feature distribution drift (FDD) caused by their non-robust feature engineering approaches. Based on Deep Metric Learning which can restrict the distances between samples explicitly and clustering algorithm which can learn multiple clusters within each category, this paper presents Robot, a fast and robust online traffic classification system. At its core, Robot leverages two building blocks to classify high-speed traffic flows: 1) For fast classification, Fast model classifies traffic and detects FDD samples simultaneously based on only one packet. 2) For robust classification, once FDD samples are detected, a flow collector will be triggered to collect flows and then Robust model, a multi-center model, will further identify them based on the hybrid of packet-level and flow-level features. Our comprehensive experiments demonstrate that Robot can achieve comparative classification speed and better generalization ability on the mixed traffic datasets with seen and unseen applications (with the FDD detection accuracy of up to 85.8% and with nearly 10% improvement in classification accuracy), compared with the state-of-the-art methods. Dan Li 0001, Kaihui Gao |
GLOBECOM | 3 |
| 2020 | Incorporating Intra-flow Dependencies and Inter-flow Correlations for Traffic Matrix PredictionabstractTraffic matrix (TM) prediction is essential for effective traffic engineering and network management. Based on our analysis of real traffic traces from Wide Area Network, the traffic flows in TM are both time-varying (i.e. with intra-flow dependencies) and correlated with each other (i.e. with inter-flow correlations). However, most existing works in TM prediction ignore inter-flow correlations. In this paper, we propose a novel Attention-based Convolutional Recurrent Neural Network (ACRNN) model to capture both intra-flow dependencies and inter-flow correlations. ACRNN mainly contains two components: 1) Correlational Modeling employs attention-based convolutional structures to capture the correlation of any two flows in TMs; 2) Temporal Modeling uses attention-based recurrent structures to model the long-term temporal dependencies of each flow, and then predicts TMs according inter-flow correlations and intra-flow dependencies. Experiments on two real-world datasets show that, when predicting the next TM, ACRNN model reduces the Mean Squared Error by up to 44.8% and reduces the Mean Absolute Error by up to 30.6%, compared to state-of-the-art method; and the gap is even larger when predicting the next multiple TMs. Besides, simulation results demonstrate that ACRNN's accurate prediction can help traffic engineering to mitigate traffic congestion. Kaihui Gao, Dan Li 0001, Li Chen 0008, Jinkun Geng, Fei Gui |
IWQoS | 1 |
| 2019 | HiPower: A High-Performance RDMA Acceleration Solution for Distributed Transaction Processing
Runhua Zhang 0002, Jinkun Geng, Shuai Wang 0028, Kaihui Gao, Guowei Shen |
NPC | 5 |