EDBT 2026 Demo / reviewers in the wild / expert
Dennis Cai
dblp:130/8438
· DBLP profile ↗
35ranked-venue papers
1as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 28 · 1 first-author · 25 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan 0005, Zhiyu Yin, Sheng Cheng 0002, Chaojie Yang, Kun Qian 0004, Tianyin Xu, Yang Zhang 0102, Yong Li 0008, Dennis Cai, Ennan Zhai |
NSDI | 12 |
| 2026 | Balancing and Beyond: Communication-Centric Optimizations in Expert ParallelismabstractThe Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to preserve kernel efficiency. However, production EP deployments often suffer from two bottlenecks: (1) expert workload imbalance, which creates computation and communication stragglers, and (2) communication inefficiency, where inter-GPU transfers dominate latency even after balancing. We present EPIC, an experience-driven EP inference system that addresses these issues progressively for real deployments. EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap. EPIC has been deployed at scale across O(10K) GPUs in our online inference service for both open-source models (e.g., Qwen3-Coder and DeepSeek-R1) and internal models, reducing communication time and per-token latency by up to 40% and 21%, respectively. Jiamin Cao, Qingxu Li, Yaozhong Liu, Shangfeng Shi, Kunling He, Ennan Zhai, Jianbo Dong, Binzhang Fu, Dennis Cai |
SIGCOMM | 14 |
| 2026 | From Nimitz to NetPila: The Evolution of Production-Scale Container Network
Sheng Cheng 0002, Jiamin Cao, Shuhong Zhu, Ennan Zhai, Dennis Cai |
SIGCOMM | 9 |
| 2026 | Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001 |
SIGCOMM | 21 |
| 2026 | Retriever: A Distributed Intrusion Detection System for NOS-Enabled NetworksabstractNetwork Operating Systems (NOS) are being widely deployed on edge devices by cloud service providers to perform fast configurations and offer high availability for new network protocols. However, NOS-enabled networks open the door to intruders that can stealthily corrupt less-guarded programmable switches to launch attacks on the entire network. Traditional centralized intrusion detection systems may neglect anomalous events on NOS-equipped switches and fail to detect such attacks. In this paper, we make the first attempt towards intrusion detection for NOS-enabled networks by designingRetriever.Retrieverfeatures a lightweight local anomaly detection module on programmable switches and a central anomaly assessment module on the central server. The local anomaly detection module selectively traces both system and network events on switches, based on which a provenance graph of events is established. Upcoming events unmatched by the provenance graph are aggregated to construct a suspicious subgraph to report to the central server. The central anomaly assessment module extracts semantic representations from reported suspicious subgraphs and computes their anomaly scores. Large-scale experiments show thatRetrievercan achieve high intrusion detection accuracy (nearly 100%) with low overheads. Runmin Ou, Yijie Bai, Yanjiao Chen, Bingchuan Tian, Zhiming Ji, Ennan Zhai, Dennis Cai, Wenyuan Xu 0001 |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2025 | Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication OptimizationabstractThe emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Specifically, GPUs involved in the same job require periodic synchronization to exchange necessary data, such as gradients, parameters, or activations. As a result, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. Moreover, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C 4. The key insights of C 4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, $\mathbf{C} 4$ can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C 4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The $\mathbf{C 4}$ has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to $\mathbf{4 5 \%}$. This enhancement is attributed to a $\mathbf{3 0 \%}$ reduction in error-induced overhead and a 15% reduction in communication costs. Jianbo Dong, Yikai Zhu, Hairong Jiao, Ennan Zhai, Wencong Xiao, Man Yuan, Siran Yang, Jiamang Wang, Rui Men, Dennis Cai, Binzhang Fu |
HPCA | 23 |
| 2025 | Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian 0021, Zhilong Zheng, Liang Chen 0001, Yichi Xu, Yikai Zhu, Xue Li 0024, Zhihui Ren, Yang Liu 0245, Yu Guan 0005, Chaojie Yang, Yang Zhang 0102, Man Yuan, Yong Li 0008, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin 0016, Dennis Cai |
NSDI | 32 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 17 |
| 2025 | New Evolution of Hoyan: Enhancing Scalability, Usability, and Accuracy for Alibaba's Global WAN VerificationabstractThe network verification system Hoyan has been deployed for Alibaba Cloud's wide-area network (WAN) for years and achieved considerable success in preventing misconfiguration-caused network incidents. However, recent years have seen the emergence of new challenges in scalability, usability, and accuracy for Hoyan. This paper presents the new evolution of Hoyan to address these challenges. First, to support the large increase in the number of routers and prefixes on our WAN, Hoyan's simulation has evolved from a centralized fashion to a distributed framework, which improves the efficiency by 5 times and can scale to O(104) routers, millions of prefixes, and billions of flows. Second, to improve Hoyan's usability in checking route change intents, we developed a specification language RCL, which supports the easy specification and automatic verification of route change intents. Third, to ensure high accuracy we enhanced Hoyan's accuracy diagnosis framework, which helped us identify and fix dozens of implementation and modeling issues. Hoyan is used on a daily basis for our WAN. It supports O(100) verification requests each week, prevents O(10) incidents each year, and helps reduce the percentage of misconfiguration-caused network incidents from 56% to 5%. Yifei Yuan 0001, Fangdan Ye, Jingkai Zhang, Mengqi Liu 0001, Yuyang Sang, Ruizhen Yang, Duncheng She, Zhiqing Ye, Tianchen Guo, Xinji Tang, Zhongyu Guan, Lingpeng Su, Ci Wang, Ruiyang Feng, Zhonghui Xie, Xianlong Zeng, Dennis Cai, Ennan Zhai |
SIGCOMM | 24 |
| 2025 | SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingabstractThe performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts. Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 16 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 39 |
| 2025 | Towards LLM-Based Failure Localization in Production-Scale NetworksabstractRoot causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it extremely challenging even for experienced operators. Large language models (LLMs) have shown great potential in text understanding and reasoning. In this paper, we present BiAn, an LLM-based framework designed to assist operators in efficient incident investigation. BiAn processes monitoring data and generates error device rankings with detailed explanations. To date, BiAn has been deployed in our network infrastructure for 10 months and it has successfully assisted operators in identifying error devices more quickly, reducing time to root causing by 20.5% (55.2% for high-risk incidents). Extensive performance evaluations based on 17 months of real cases further demonstrate that BiAn achieves accurate and fast failure localization. It improves accuracy by 9.2% compared to the baseline approach. Chenxu Wang 0007, Xumiao Zhang, Runwei Lu, Xianshang Lin, Xuan Zeng 0002, Zhe An, Gongwei Wu, Chen Tian 0001, Guihai Chen, Guyue Liu, Yuhong Liao, Dennis Cai, Ennan Zhai |
SIGCOMM | 15 |
| 2025 | SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud InfrastructuresabstractFor providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network failures. In practice, there is a gap between the flooding raw alerts data collected by network monitoring tools and the readable information needed for failure diagnosis. Existing solutions using limited network monitoring data sources and heuristic diagnostic rules, lack comprehensive coverage and the capability to address severe failures, especially which network operators have never handled a similar one before. This paper presents SkyNet, a network analysis system to extract scope and severity information from alert floods. SkyNet ensures comprehensive coverage by integrating multiple monitoring data sources through a uniform input format, enhancing extensibility for new network monitoring tools. During alert floods, SkyNet groups alerts, assesses their severity, and filters out insignificant ones to aid network operators in mitigating network failures. To date, SkyNet has been running stably on our network for one and a half years without any false negatives and has successfully reduced the time-to-mitigation for over 80% of network failures since its deployment in production. Huanwu Hu, Yunguang Li, Xiangyu Tang, Bingchuan Tian, Gongwei Wu, Xumiao Zhang, Ennan Zhai, Yuhong Liao, Dennis Cai |
SIGCOMM | 14 |
| 2025 | Roaming Free in the VR World with MP2
Xumiao Zhang, Yuning Chen, Xuan Zeng 0002, Zhilong Zheng, Xianshang Lin, Yanmei Liu, Songwu Lu, Z. Morley Mao, Wan Du, Dennis Cai, Ennan Zhai |
USENIX ATC | 12 |
| 2024 | LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge Clouds
Tian Pan 0001, Xionglie Wei, Yisong Qiao, Tiesheng Cheng, Wenqiang Su, Yuke Hong, Zhengzhong Wang, Chongjing Dai, Peiqiao Wang, Xuetao Jia, Jianyuan Lu, Enge Song, Biao Lyu, Ennan Zhai, Jiao Zhang 0002, Tao Huang 0005, Dennis Cai, Shunmin Zhu |
NSDI | 24 |
| 2024 | Sirius: Composing Network Function Chains into P4-Capable Edge Gateways
Jiamin Cao, Mengqi Liu 0001, Dennis Cai, Ennan Zhai |
NSDI | 6 |
| 2024 | Reasoning about Network Traffic Load Property at Production Scale
Fangdan Ye, Yifei Yuan 0001, Ruizhen Yang, Bingchuan Tian, Tianchen Guo, Zhongyu Guan, Xianlong Zeng, Chenren Xu, Dennis Cai, Ennan Zhai |
NSDI | 13 |
| 2024 | Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingabstractDeep learning training (DLT), e.g., large language model (LLM) training, has become one of the most important services in multitenant cloud computing. By deeply studying in-production DLT jobs, we observed that communication contention among different DLT jobs seriously influences the overall GPU computation utilization, resulting in the low efficiency of the training cluster. In this paper, we present Crux, a communication scheduler that aims to maximize GPU computation utilization by mitigating the communication contention among DLT jobs. Maximizing GPU computation utilization for DLT, nevertheless, is NP-Complete; thus, we formulate and prove a novel theorem to approach this goal by GPU intensity-aware communication scheduling. Then, we propose an approach that prioritizes the DLT flows with high GPU computation intensity, reducing potential communication contention. Our 96-GPU testbed experiments show that Crux improves 8.3% to 14.8% GPU computation utilization. The large-scale production trace-based simulation further shows that Crux increases GPU computation utilization by up to 23% compared with alternatives including Sincronia, TACCL, and CASSINI. Jiamin Cao, Yu Guan 0005, Kun Qian 0021, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 8 |
| 2024 | A General and Efficient Approach to Verifying Traffic Load Properties under Arbitrary k FailuresabstractThis paper presents YU, the first verification system for checking traffic load properties under arbitrary failure scenarios that can scale to production Wide Area Networks (WANs). Building a practical YU requires us to address two challenges in terms of generality and efficiency. The state-of-the-art efforts either assume shortest-path-based forwarding (e.g., QARC) or only target single-failure reasoning (e.g., Jingubang). As a result, the former inherently cannot generalize to widely used protocols (e.g., SR and iBGP) that are beyond shortest-path forwarding, while the latter cannot efficiently handle arbitrary failure scenarios. For the generality challenge, we propose an approach inspired by symbolic execution, called symbolic traffic execution, to model the forwarding behavior of a range of practically deployed protocols (e.g., eBGP, iBGP, iGP, and SR) under failure scenarios. For the efficiency challenge, we propose diverse equivalence classification techniques (i.e., k-failure-equivalence and link-local-equivalence reduction) to reduce the symbolic traffic execution overhead caused by both the large size of the production WAN and the huge number of traffic flows traversing it. YU has been used in the daily verification of our WAN for several months and has successfully identified potential failure scenarios that would lead to traffic load violations. Yifei Yuan 0001, Fangdan Ye, Mengqi Liu 0001, Ruizhen Yang, Tianchen Guo, Xianlong Zeng, Chenren Xu, Dennis Cai, Ennan Zhai |
SIGCOMM | 11 |
| 2024 | Alibaba HPN: A Data Center Network for Large Language Model TrainingabstractThis paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production. Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai |
SIGCOMM | 18 |
| 2024 | Diagnosing Application-network Anomalies for Millions of IPs in Production Clouds
Zhe Wang 0015, Huanwu Hu, Linghe Kong, Xinlei Kang, Qiao Xiang, Peihao Yang, Jiejian Wu, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Xianlong Zeng, Dennis Cai, Guihai Chen |
USENIX ATC | 15 |
| 2023 | Norma: Towards Practical Network Load Testing
Bingchuan Tian, Chen Tian 0001, Yu Zhou 0008, Mengjing Ma, Zhewen Yang, Guihai Chen, Dennis Cai, Ennan Zhai |
NSDI | 11 |
| 2023 | Flor: An Open High Performance RDMA Framework Over Heterogeneous RNICs
Qiang Li 0045, Yixiao Gao, Xiaoliang Wang 0001, Haonan Qiu, Yanfang Le, Derui Liu, Qiao Xiang, Bo Li 0061, Jianbo Dong, Lingbo Tang, Hongqiang Harry Liu, Shaozong Liu, Rui Miao 0001, Yaohui Wu, Zhiwu Wu, Zheng Cao 0003, Zhongjie Wu, Chen Tian 0001, Guihai Chen, Dennis Cai, Jiaji Zhu, Jiesheng Wu, Jiwu Shu |
OSDI | 25 |
| 2023 | CellFusion: Multipath Vehicle-to-Cloud Video Streaming with Network Coding in the WildabstractThis paper presents CellFusion, a system designed for high-quality, real-time video streaming from vehicles to the cloud. It leverages an innovative blend of multipath QUIC transport and network coding. Surpassing the limitations of individual cellular carriers, CellFusion uses a unique last-mile overlay that integrates multiple cellular networks into a single, unified cloud connection. This integration is made possible through the use of in-vehicle Customer Premises Equipment (CPEs) and edge-cloud proxy servers. Yunzhe Ni, Zhilong Zheng, Xianshang Lin, Fengyu Gao, Xuan Zeng 0002, Yirui Liu 0001, Senlang Du, Guang Yang 0006, Yuanchao Su, Dennis Cai, Hongqiang Harry Liu, Chenren Xu, Ennan Zhai |
SIGCOMM | 13 |
| 2023 | XRON: A Hybrid Elastic Cloud Overlay Network for Video Conferencing at Planetary ScaleabstractQuality and cost are two key considerations for video conferencing services. Service providers face a dilemma when selecting network tiers to build their infrastructure---relying on Internet links has poor quality, while using premium links brings excessive cost. Bingyang Wu, Kun Qian 0021, Bo Li 0061, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 8 |
| 2023 | Automated Verification of an In-Production DNS Authoritative EngineabstractThis paper presents DNS-V, a verification framework for our in-production DNS authoritative engine, which is the core of our DNS service. The key idea for automated verification in general is based on the layered verification principle. However, we face the challenge that our in-production DNS authoritative engine lacks modularity, more specifically, as can be seen with unclean interfaces and poor data structure encapsulation. This makes the layered verification hard to apply. To address this challenge, we propose a summarization approach that performs full-path symbolic execution to accumulate all path conditions and computation effects, and then represents a module's behavior in an abstract form as a set of input-effect pairs. In addition, for portability to future iterated versions of our DNS authoritative engine, we identify common dependency library modules that remain stable across different versions, and carefully design their abstractions to make them amenable to automated reasoning. Our framework has been successful in identifying and preventing tens of critical bugs in different versions of our DNS authoritative engine from reaching production, with a porting effort of less than one person-week. Naiqian Zheng, Mengqi Liu 0001, Yuxing Xiang, Linjian Song, Nan Wang 0041, Zhuo Liang, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SOSP | 10 |
| 2023 | DxPU: Large-scale Disaggregated GPU Pools in the DatacenterabstractThe rapid adoption of AI and convenience offered by cloud services have resulted in the growing demands for GPUs in the cloud. Generally, GPUs are physically attached to host servers as PCIe devices. However, the fixed assembly combination of host servers and GPUs is extremely inefficient in resource utilization, upgrade, and maintenance. Due to these issues, the GPU disaggregation technique has been proposed to decouple GPUs from host servers. It aggregates GPUs into a pool and allocates GPU node(s) according to user demands. However, existing GPU disaggregation systems have flaws in software-hardware compatibility, disaggregation scope, and capacity. In this article, we present a new implementation of datacenter-scale GPU disaggregation, named DxPU. DxPU efficiently solves the above problems and can flexibly allocate as many GPU node(s) as users demand. To understand the performance overhead incurred by DxPU, we build up a performance model for AI specific workloads. With the guidance of modeling results, we develop a prototype system, which has been deployed into the datacenter of a leading cloud provider for a test run. We also conduct detailed experiments to evaluate the performance overhead caused by our system. The results show that the overhead of DxPU is less than 10%, compared with native GPU servers, in most of user scenarios. Weinan Li, Yajin Zhou, Linquan Jiang, Qiang Liu 0036, Dennis Cai |
ACM Trans. Archit. Code Optim. | 11 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 14 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 14 |
| 2022 | GSO-simulcast: global stream orchestration in simulcast video conferencing systemsabstractWe present GSO-Simulcast, a new architecture designed for large-scale multi-party video-conferencing systems. GSO-Simulcast is currently deployed at full-scale in Alibaba's Dingtalk video conferencing that serves more than 500 million users. It marks a fundamental shift from today's Simulcast, where a media server locally decides how to switch and forward video streams based on a fragmented network view. Instead, GSO-Simulcast globally orchestrates the publishing, subscribing, as well as the resolution and bitrate of video streams for each participant using a centralized controller that is aware of all network constraints in a meeting. The controller automatically modifies stream configurations to meet the participants' real-time network changes and updates. In doing so, GSO-Simulcast achieves multiple goals: (1) reducing video and network mismatch, (2) less path congestion, and (3) automated stream policy management. With the deployment of GSO-Simulcast, we observed more than a 35% reduction in the average video stall, 50% reduction in the average voice stall, and 6% improvement in the average video framerate. We describe the principle, design, deployment, and lessons learned. Xianshang Lin, Junshao Zhang, Yao Cui, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 8 |
| 2022 | From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudabstractThis paper presents the two generations of storage network stacks that reduced the average I/O latency of Alibaba Cloud's EBS service by 72% in the last five years: Luna, a user-space TCP stack that corresponds the latency of network to the speed of SSD; and Solar, a storage-oriented UDP stack that enables both storage and network hardware accelerations. Rui Miao 0001, Lingjun Zhu, Kun Qian 0021, Shujun Zhuang, Bo Li 0061, Shuguang Cheng, Binzhang Fu, Jiaji Zhu, Jiesheng Wu, Dennis Cai, Hongqiang Harry Liu |
SIGCOMM | 16 |
| 2021 | When Cloud Storage Meets RDMA
Yixiao Gao, Qiang Li 0045, Lingbo Tang, Yongqing Xi, Wenwen Peng, Bo Li 0061, Yaohui Wu, Shaozong Liu, Xingkui Liu, Zhongjie Wu, Junping Wu, Zheng Cao 0003, Chen Tian 0001, Jiaji Zhu, Haiyong Wang, Dennis Cai, Jiesheng Wu |
NSDI | 23 |
| 2020 | Flow Event Telemetry on Programmable Data PlaneabstractNetwork performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic. Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001 |
SIGCOMM | 12 |
| 2020 | Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsabstractProgrammable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs. Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu |
SIGCOMM | 8 |
| 2014 | Evolve carrier ethernet architecture with SDN and segment routingabstractEthernet technology has been evolving to become the main Wide Area Network transport technology. To address the rising CAPEX and OPEX issues, service providers are deploying Unified MPLS technology to consolidate various networks into one an integrated carrier ethernet network. Although Unified MPLS has simplified MPLS deployment in many aspects, it is still deemed complex and it is not as agile as what clouding computing demands. This paper proposes a further evolution of carrier ethernet architecture by coupling the emerging segment routing and Software-Defined Network (SDN) technologies. This new architecture will significantly simplify the network infrastructure while providing rich converged services with embedded high availability and agility. Dennis Cai, Anna Wielosz, Songbin Wei |
WoWMoM | 1 |