EDBT 2026 Demo / reviewers in the wild / expert
Jiamin Cao
dblp:224/2165
· DBLP profile ↗
25ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 23 · 7 first-author · 16 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters
Chenyang Hei, Jiamin Cao, Chengxi Gao, Xiuzhu Sha, Tongrui Liu, Dengke Zhang, Ennan Zhai, Xingwei Wang 0001 |
NSDI | 3 |
| 2026 | Balancing and Beyond: Communication-Centric Optimizations in Expert ParallelismabstractThe Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to preserve kernel efficiency. However, production EP deployments often suffer from two bottlenecks: (1) expert workload imbalance, which creates computation and communication stragglers, and (2) communication inefficiency, where inter-GPU transfers dominate latency even after balancing. We present EPIC, an experience-driven EP inference system that addresses these issues progressively for real deployments. EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap. EPIC has been deployed at scale across O(10K) GPUs in our online inference service for both open-source models (e.g., Qwen3-Coder and DeepSeek-R1) and internal models, reducing communication time and per-token latency by up to 40% and 21%, respectively. Jiamin Cao, Qingxu Li, Yaozhong Liu, Shangfeng Shi, Kunling He, Ennan Zhai, Jianbo Dong, Binzhang Fu, Dennis Cai |
SIGCOMM | 1 |
| 2026 | PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload ReallocationabstractModern collective communication libraries (CCLs) execute a collective communication task (CCT) by decomposing it into multiple sub-tasks, each mapped to a specific Virtual Topology (VT), which is an ordered graph of GPUs (e.g., a ring or a tree), to maximize parallelism and link utilization. As AI training scales to larger clusters, network anomalies (congestion and failures) are unavoidable, and a single straggling VT can delay the entire CCT. Existing solutions either rely on low-level transport-layer solutions which lacks a cross-sub-task perspective, or static CCL scheduling, failing to adapt to the dynamic and heterogeneous networks. Kaihui Gao, Li Chen 0008, Fei Gui, Dan Li 0001, Jiamin Cao |
SIGCOMM | 8 |
| 2026 | Theseus: Runtime-Adaptive GPU Collective Communication with Hot-Swappable SchedulesabstractCurrent GPU Collective Communication Libraries (CCLs) employ predefined schedules optimized for stable environments. Their supported schedules and selection logic are fixed at communicator initialization, which fails to account for evolving runtime conditions, such as workload characteristics and hardware health status. Consequently, long-running GPU jobs experience suboptimal performance after hours or days of execution, which translates into longer job completion times and wasted GPU cluster resources. To address this problem, we present Theseus, a novel CCL backend that provides schedule-level runtime adaptivity. It admits user-defined schedules and selection policies. As runtime conditions change, Theseus selects suitable schedules using cluster-wide runtime attributes beyond CCL-internal metrics. Moreover, it hot-swaps from the previous schedule consistently across GPUs with low overhead. Theseus acts as a drop-in replacement to facilitate integration. We evaluate Theseus extensively on various GPU workloads with intuitive policies. Compared with NCCL, Theseus achieves up to 1.61X speedup of communication time in stable environments and 2.46X in dynamic environments. It improves end-to-end job completion time by up to 1.84X while incurring comparable or lower overhead. Rui Ding 0014, Xiandong Lu, Xunpeng Liu, Xuran Hao, Houyuan Zhu, Anyi Xu, Sinuo Cao, Haifeng Sun 0004, Qun Huang 0001, Jiamin Cao |
SIGCOMM | 12 |
| 2026 | From Nimitz to NetPila: The Evolution of Production-Scale Container Network
Sheng Cheng 0002, Jiamin Cao, Shuhong Zhu, Ennan Zhai, Dennis Cai |
SIGCOMM | 4 |
| 2025 | SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingabstractThe performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts. Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 1 |
| 2025 | ResCCL: Resource-Efficient Scheduling for Collective CommunicationabstractAs distributed deep learning training (DLT) systems scale, collective communication has become a significant performance bottleneck. While current approaches optimize bandwidth utilization and task completion time, existing communication libraries (CCLs) backends fail to efficiently manage GPU resources during algorithm execution, limiting the performance of advanced algorithms. This paper proposes ResCCL, a novel CCL backend designed for Resource-Efficient Scheduling to address key limitations in current systems. ResCCL enhances execution efficiency by optimizing scheduling at the primitive level (e.g., send and recvReduceCopy), enabling flexible thread block (TB) allocation, and generating lightweight communication kernels to minimize runtime overhead. Our approach tackles the global scheduling problem, reduces idle TB resources, and enhances communication bandwidth. Evaluation results demonstrate that ResCCL achieves up to 2.5× improvement in bandwidth performance compared to both NCCL and MSCCL. It reduces SM resource overhead by 77.8% and increases TB utilization by 41.6% while running the same algorithms. In end-to-end DLT, ResCCL boosts Megatron's throughput by up to 39%. Tongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiamin Cao, Ennan Zhai, Xingwei Wang 0001 |
SIGCOMM | 5 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 20 |
| 2024 | Sirius: Composing Network Function Chains into P4-Capable Edge Gateways
Jiamin Cao, Mengqi Liu 0001, Dennis Cai, Ennan Zhai |
NSDI | 2 |
| 2024 | Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingabstractDeep learning training (DLT), e.g., large language model (LLM) training, has become one of the most important services in multitenant cloud computing. By deeply studying in-production DLT jobs, we observed that communication contention among different DLT jobs seriously influences the overall GPU computation utilization, resulting in the low efficiency of the training cluster. In this paper, we present Crux, a communication scheduler that aims to maximize GPU computation utilization by mitigating the communication contention among DLT jobs. Maximizing GPU computation utilization for DLT, nevertheless, is NP-Complete; thus, we formulate and prove a novel theorem to approach this goal by GPU intensity-aware communication scheduling. Then, we propose an approach that prioritizes the DLT flows with high GPU computation intensity, reducing potential communication contention. Our 96-GPU testbed experiments show that Crux improves 8.3% to 14.8% GPU computation utilization. The large-scale production trace-based simulation further shows that Crux increases GPU computation utilization by up to 23% compared with alternatives including Sincronia, TACCL, and CASSINI. Jiamin Cao, Yu Guan 0005, Kun Qian 0021, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 1 |
| 2024 | Alibaba HPN: A Data Center Network for Large Language Model TrainingabstractThis paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production. Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai |
SIGCOMM | 3 |
| 2024 | P4runpro: Enabling Runtime Programmability for RMT Programmable SwitchesabstractProgrammable switches have revolutionized network operations by enabling the flexible customization of packet processing logic using language like P4. However, changing the programs running on the switch requires disturbing traffic and suspending other unrelated programs. In this paper, we present P4runpro, enabling runtime data plane updates with dynamic resource allocation. The P4runpro data plane abstracts hardware resources and defines dynamically reconfigurable atomic operations that form packet processing logic. P4runpro provides runtime programming interfaces called P4runpro primitives for the operator to write high-level programs. We have designed the P4runpro compiler to automatically and consistently link the P4runpro programs to the running data plane. We implement our prototype on a Tofino switch. We implement 15 example runtime programs using P4runpro to demonstrate its generality and expressiveness. Our evaluation results show that compared to the state-of-the-art, P4runpro can respond within hundreds of milliseconds, achieve an average of 60% to 80% dynamic resource utilization, concurrently run ≈0.6K to ≈2.8K programs, and introduce lower overhead. Our case studies illustrate the benefit of runtime programming and prove the same functionality between P4runpro and conventional P4 programs. Yifan Yang 0009, Lin He 0004, Xiaoyi Shi, Jiamin Cao, Ying Liu 0024 |
SIGCOMM | 5 |
| 2023 | HyperClassifier: Accurate, Extensible and Scalable Traffic Classification with Programmable SwitchesabstractTraffic classification provides substantial benefits for service differentiation, security policy enforcement, and traffic engineering. However, accurately classifying large volumes of network traffic using existing solutions is pretty challenging, as they are typically implemented on commodity servers with slow CPUs for packet processing. To address this, we leverage the opportunity provided by emerging programmable switches and propose HyperClassifier as a solution to achieve accurate, extensible, and scalable traffic classification. HyperClassifier designs an efficient classifying table with an effective flow expiration mechanism that enables lightweight packet inspection on resource-limited switches. We implement an open-source prototype of HyperClassifier on a hardware Tofino switch and conduct extensive evaluations. The results of our evaluation demonstrate that, compared to existing solutions, HyperClassifier can provide orders of magnitude higher classification throughput with comparable classification accuracy. Yichi Xu, Jiamin Cao, Menghao Zhang 0001, Ying Liu 0024, Mingwei Xu 0001 |
ICC | 3 |
| 2022 | Firebolt: Finding Bugs in Programmable Data Plane Generators
Jiamin Cao, Yu Zhou 0008, Chen Sun 0005, Lin He 0004, Zhaowei Xi, Ying Liu 0024 |
USENIX ATC | 1 |
| 2022 | TurboNet: Faithfully Emulating Networks With Programmable SwitchesabstractFaithfully emulating networks is critical for verifying the correctness and effectiveness of new networking-related designs. Existing network experiment platforms either cannot faithfully emulate the functionality and performance of production networks or cannot scale well due to cost constraints. In this paper, we proposeTurboNet, a new network emulator that utilizes one or more programmable switches to achieve faithful emulation of the network data plane and control plane. For data plane emulation, we propose a series of key designs, such as port mapper, queue mapper, and delayed queue, to emulate network topologies and performance metrics with high flexibility and accuracy. For control plane emulation, we support static routing configurations, distributed routing agents, and the centralized routing controllers. Meanwhile, we provide APIs for operators to simplify network emulation tasks. We implementTurboNeton Tofino switches. Evaluation results show that: (1) On the data plane,TurboNetcan flexibly emulate various topologies, such as an 8-ary fat-tree with only one programmable switch and a 10-ary fat-tree with four programmable switches; (2) On the control plane,TurboNetsupports about 200 BGP agents on a single programmable switch with a CPU usage of 25%; (3)TurboNetcan accurately emulate different network performance metrics such as 10−8link loss, and microsecond to millisecond link delay. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Newton: Intent-Driven Network Traffic MonitoringabstractNetwork monitoring systems are designed to fulfill operators’ intents and serve as essential tools to modern networks. As a result of rapidly increasing network bandwidth and scale nowadays, network monitors should satisfy on-demand network monitoring for continuously growing traffic volumes. However, existing monitoring systems either cannot satisfy flexible intents on demand or produce significant overheads. In this paper, we presentNewton, an intent-driven traffic monitor that is able to specify operators’ intents with traffic monitoring queries and conduct dynamic and scalable network-wide queries deployment.Newtonenables operators to customize and modify queries dynamically without interrupting the network workflow. Besides,Newtonproposes systematic optimizations at device level and network-wide level to reduce resource consumption while deploying queries.Newtoncan combine the resources across switches to deploy complex queries with high resilience to dynamic network status. Evaluations prove thatNewtonis of high flexibility, scalability, and resource efficiency, which demonstratesNewtonis promising to be deployed in large-scale programmable networks. Zhaowei Xi, Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | CoFilter: High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal ServersabstractAs one of the most critical cloud services, Bare-Metal Servers (BMS) introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for BMS. However, the off-the-shelf stateful packet filters either are costly for cloud DCNs or introduce significant performance bottlenecks. In this article, we presentCoFilter, which leverages low-cost programmable switches to accelerate the stateful packet filter for BMS.CoFilteruses (1)stateful process partitionto enable complex stateful packet filtering logic on programmability-limited switching ASICs, (2)state compressionto track tens of millions of connections with constrained hardware memory, and (3)per-tenant packet rate limit and tenant-aware flow migrationto achieve efficient performance isolation among different tenants. Overall,CoFilterimplements a high-performance stateful packet filter via the co-design of programmable switching ASIC and CPU. We evaluateCoFilterunder various data center traffic traces with real-world flow distributions. The evaluation results show thatCoFilterremarkably outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay within 1us, and freeing a significant quantity of CPU cores, with rather small memory usage, i.e., accommodating over$10^7$connections with only 16MB SRAM. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Chen Sun 0005, Yangyang Wang 0001, Mingwei Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | pSAV: A Practical and Decentralized Inter-AS Source Address Validation Service FrameworkabstractSource IP address spoofing has been a major vulnerability of the Internet for many years. Although much work has been done to study the problem extensively, spoofing continues to occur frequently and has led to many serious network attacks. Inter-AS source address validation (SAV) is considered an important defense method for AS to filter spoofed packets. However, existing work has been unable to drive inter-AS SAV deployment into practice due to the lack of deployment incentives and trust foundation.In this paper, we propose a practical and decentralized inter-AS SAV service framework, pSAV, to promote inter-AS SAV deployment. pSAV increases deployment incentives by treating SAV as a payable service and dividing the participant ASes into service subscribers, providers, and auditors. On the control plane, pSAV leverages blockchain as a trust foundation to provide service subscriptions and audits with automatic incentive allocation. On the data plane, pSAV leverages P4-programmable switches to provide flexible and high-performance SAV services. We prototype the pSAV control plane based on Hyperledger Fabric and implement various SAV techniques on Barefoot Tofino switches. The evaluation results show that (1) on the control plane, pSAV blockchain can provide high-performance service transactions (hundreds of transactions per second with second latency), and (2) on the data plane, pSAV can provide various high-throughput (hundreds of Gbps) SAV services using only one programmable switch. Jiamin Cao, Ying Liu 0024, Mingxing Liu, Lin He 0004, Yihao Jia |
IWQoS | 1 |
| 2020 | Newton: intent-driven network traffic monitoringabstractMonitoring network traffic based on operators' intents is essential to today's networks. As the bandwidth and size of networks increase steeply, monitoring systems shall fulfill the requirements of on-demand network monitoring for ever-growing traffic volumes. However, existing monitoring systems either cannot satisfy operators' intents on demand or introduce substantial monitoring overheads. In this paper, we present Newton, an intent-driven traffic monitor that enables specifying operators' intents with traffic monitoring queries and supports dynamic and scalable network-wide queries. Specifically, Newton 1) empowers operators to dynamically create, remove, and update on-data-plane queries without interrupting normal packet forwarding, 2) conducts systematic optimizations to achieve precise network traffic monitoring, and 3) executes network-wide queries with high resilience to dynamic network status. Evaluation results show that Newton improves the flexibility, scalability, and resource efficiency of traffic monitoring, demonstrating its great potential to be deployed in large-scale programmable networks. Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001 |
CoNEXT | 5 |
| 2020 | TurboNet: Faithfully Emulating Networks with Programmable SwitchesabstractFaithfully emulating networks is critical for verifying the correctness and effectiveness of new networking-related designs. Existing network experiment platforms either cannot faithfully emulate functionality and performance of production networks or cannot scale well because of cost limitations. In this paper, we propose TurboNet, a new network emulator that leverages one programmable switch to enable faithful emulation of both network data plane and control plane. For data plane emulation, we present a series of key designs such as port mapper, queue mapper, and delayed queue to emulate network topologies and performance metrics with high flexibility and accuracy. For control plane emulation, we support static routing configurations, distributed routing agents, and the centralized routing controller. Meanwhile, we provide API for operators to simplify network emulation tasks. We implement TurboNet on a Tofino switch. The evaluation results show that: (1) TurboNet can flexibly emulate various topologies such as the 8-ary fat-tree on the data plane and support about 200 BGP agents with 25% CPU usage on the control plane; (2) TurboNet can accurately emulate different network performance metrics, including 400Gbps linerate background traffic injection, as small as 10-8link loss, and microsecond-level to millisecond-level link delay. Jiamin Cao, Yu Zhou 0008, Ying Liu 0024, Mingwei Xu 0001, Yongkai Zhou |
ICNP | 1 |
| 2020 | Martini: Bridging the Gap between Network Measurement and Control Using Switching ASICsabstractAdvanced network management systems, including network measurement and traffic control, rely on a remote controller to make control decisions. However, this approach incurs a long control loop of a few seconds to minutes. Even if we switch to switch-local controller, the latency is still tens of milliseconds and is unacceptable for many latency-sensitive tasks. In this paper, we propose Martini, a general framework that supports measurement-based timely control. The key idea is to perform measurement, control decision, and control entirely in the switch data plane. This could shorten the control loop of management tasks that require timely control based on only locally measured statistics in the switch. First, Martini introduces a set of primitives to describe management tasks. Next, Martini provides an innovative network-wide task placement mechanism to exploit resources of all switches to accommodate massive management tasks. Finally, Martini provides a code library and a compiler to support measurement and control on a state-of-the-art switching ASIC. Evaluation results show that Martini can effectively support a wide range of fine-timescale management tasks such as microburst detection and fast load balancing by reducing the control loop from seconds to nanoseconds. Shuhe Wang, Chen Sun 0005, Zili Meng, Minhu Wang, Jiamin Cao, Mingwei Xu 0001, Jun Bi, Qun Huang 0001, Masoud Moshref, Tong Yang 0003, Hongxin Hu, Gong Zhang 0001 |
ICNP | 5 |
| 2020 | HyperSight: Towards Scalable, High-Coverage, and Dynamic Network Monitoring QueriesabstractPerforming fine-grained and real-time network monitoring is the core logic of various data center operation applications, such as traffic engineering, network troubleshooting, and anomaly detecting. However, the state-of-the-art network monitoring solutions either fall short of completely detecting all network incidents (i.e., congestion), yielding limited monitoring coverage, or introduce large overheads, yielding limited scalability. In this paper, we present HyperSight, a network traffic monitor with both high coverage and low overheads. The key idea of HyperSight is to monitor networks at the behavior level via tracking packet behavior changes. HyperSight proposes three designs for behavior-level monitoring. First, to facilitate expressing various network monitoring tasks, HyperSight presents a declarative query language based on the streaming processing model. Second, HyperSight proposes Bloom Filter Queue (BFQ), a memory-efficient algorithm to empower in-network capability for monitoring packet behavior changes. BFQ can be implemented on commodity programmable switches. Third, to support dynamic deployment and execution of packet behavior change monitoring tasks without interrupting on-service switches, HyperSight proposes virtual BFQ to support dynamic query compilation. We build a prototype of HyperSight and deploy it on commodity programmable switches. Evaluation results show that HyperSight supports a wide range of network event queries and can monitor over 99% packet behavior changes while keeping remarkably low overheads. Yu Zhou 0008, Jun Bi, Tong Yang 0003, Kai Gao 0001, Jiamin Cao, Yangyang Wang 0001, Cheng Zhang 0012 |
IEEE J. Sel. Areas Commun. | 5 |
| 2019 | CoFilter: A High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal ServersabstractAs one of the most critical cloud services, Bare-metal Servers introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for bare-metal servers. However, the off-the-shelf hardware-based and software-based stateful packet filters either are prohibitively costly for cloud DCNs or introduce significant performance bottlenecks. In this paper, we present CoFilter, which employs cheap programmable switches to accelerate the stateful packet filter for bare-metal servers. CoFilter consists of two key designs. First, to support complex stateful packet filtering logic in programmability-limited switching ASICs, CoFilter partitions the stateful packet filtering logic between programmable ASICs and switch CPU. Most packets are directly processed in switching ASICs to achieve high performance, while only a small number of packets go to switch CPU for connection tracking. Second, to track massive connections with constrained hardware memory, CoFilter employs hash to compress connection states and provides an efficient settlement for hash collisions. We build a prototype of CoFilter and evaluate it on the Tofino switch under various data center traffic traces with real-world flow distribution. The evaluation shows that CoFilter largely outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay at 1us, and freeing a significant quantity of CPU cores. Furthermore, CoFilter presents great scalability and accommodates over ten million connections with only 16MB SRAM. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Chen Sun 0005, Yangyang Wang 0001, Jun Bi |
ICCCN | 1 |
| 2019 | P4Tester: efficient runtime rule fault detection for programmable data planesabstractP4 and programmable data planes bring significant flexibility to network operation but are inevitably prone to various faults. Some faults, like P4 program bugs, can be verified statically, while some faults, like runtime rule faults, only happen to running network devices, and they are hardly possible to handle before deployment. Existing network testing systems can troubleshoot runtime rule faults via injecting probes, but are insufficient for programmable data planes due to large overheads or limited fault coverage. In this paper, we propose P4Tester, a new network testing system for troubleshooting runtime rule faults on programmable data planes. First, P4Tester proposes a new intermediate representation based on Binary Decision Diagram, which enables efficient probe generation for various P4-defined data plane functions. Second, P4Tester offers a new probe model that uses source routing to forward probes. This probe model largely reduces rule fault detection overheads, i.e. requiring only one server to generate probes for large networks and minimizing the number of probes. Moreover, this probe model can test all table rules in a network, achieving full fault coverage. Evaluation based on real-world data sets indicates that P4Tester can efficiently check all rules in programmable data planes, generate 59% fewer probes than ATPG and Pronto, be faster than ATPG by two orders of magnitude, and troubleshoot multiple rule faults within one second on BMv2 and Tofino. Yu Zhou 0008, Jun Bi, Yunsenxiao Lin, Yangyang Wang 0001, Zhaowei Xi, Jiamin Cao, Chen Sun 0005 |
IWQoS | 7 |
| 2018 | KeySight: Troubleshooting Programmable Switches via Scalable High-Coverage Behavior TrackingabstractThe rise of programmable switches and P4 brings much flexibility to networks, but this flexibility comes with increased risks of bugs. Diagnosing these bugs is essential for network operation but is non-trivial. A potential approach is to track packet behaviors through postcards, but existing tools either generate substantial postcards (limited scalability) or only track a small proportion of packet behaviors (low coverage). In this paper, we present KeySight, a platform that troubleshoots programmable switches with high scalability and high coverage. The key idea is based on the Packet Equivalence Class (PEC) abstraction that aggregates packets with identical behaviors and generates one postcard per behavior. The PEC abstraction minimizes the number of postcards while tracking all packet behaviors. We design novel algorithms to analyze PECs of P4 programs and to implement the PEC abstraction on programmable switches. We deploy KeySight on Tofino and SmartNIC, and evaluate it against 80 P4 programs and real packet traces of over 5TB. Results show that in the premise of overseeing over 99.9% packet behaviors, KeySight reduces the number of postcards by one to two orders of magnitude when comparing with NetSight. Yu Zhou 0008, Jun Bi, Tong Yang 0003, Kai Gao 0001, Cheng Zhang 0012, Jiamin Cao, Yangyang Wang 0001 |
ICNP | 6 |