Yu Zhou 0008

dblp:36/2728-8 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-3469-0906ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 34 · 8 first-author · 14 since 2021Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Dynamic Compute and Network Orchestration for Disaggregated RL
abstract
Disaggregating the generation and training stages in RL is widely adopted to scale LLM post-training. There are two critical challenges here. First, the generation stage often becomes a bottleneck due to dynamic workload shifts and severe execution imbalances. Second, the decoupled stages result in diverse and dynamic network traffic patterns that strain the conventional static fabric.
Xin Tan 0004, Yicheng Feng, Yu Zhou 0008, Yibo Zhu 0001, Hong Xu 0001
SIGCOMM3
2025 InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
abstract
Scaling Large Language Model (LLM) training relies on multidimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like Tensor Parallelism. However, existing HBD architectures face fundamental limitations in scalability, cost, and fault resiliency: switch-centric HBDs (e.g., NVL-72) incur prohibitive scaling costs, while GPU-centric HBDs (e.g., TPUv3/Dojo) suffer from severe fault propagation. Switch-GPU hybrid HBDs (e.g., TPUv4) take a middle-ground approach, but the fault explosion radius remains large.
Chenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng, Yu Zhou 0008, Wenqing Lv, Yelong Xu, Yuanwei Lu, Yanbo Yu, Yichen Shen 0001, Yibo Zhu 0001, Daxin Jiang
SIGCOMM5
2024 Exploring Dynamic Rule Caching Under Dependency Constraints for Programmable Switches: Theory, Algorithm, and Implementation
abstract
Ternary Content Addressable Memory (TCAM) enables fast lookup and is widely used by routers and switches to support policy-based forwarding. Due to high cost and small capacity, only a small subset of important rules can be cached in TCAM, so determining it is critical to increasing the hit ratio. This is more challenging than traditional caching problems because of complicated rule dependency relationships. Existing works are based on heuristics and they don’t work well under all practical scenarios. Worse still, the lack of fundamental understanding of the design space, complexity, and optimality makes all explorations in mystery. In this paper, we use a modeling-based method to formulate the problem, prove its complexity, and propose DROPS, a dynamic rule caching framework with a much higher hit ratio. In particular, we deduce the rule selection problem into a multi-dimensional rule space transformation problem. Thus, we are no longer limited by using the intrinsic rules; rather, we can transform original rules into “new rules” equivalently without rule dependency. We design non-trivial rule placement and update algorithms and implement them in programmable switches. In the experimental evaluation, we show that our method outperforms all existing methods.
Xinhao Deng 0001, Mingwei Xu 0001, Qi Li 0002, Weijie Wu, Yuan Yang 0001, Menghao Zhang 0001, Yu Zhou 0008
IEEE Trans. Netw. Serv. Manag.7
2023 Norma: Towards Practical Network Load Testing
Bingchuan Tian, Chen Tian 0001, Yu Zhou 0008, Mengjing Ma, Zhewen Yang, Guihai Chen, Dennis Cai, Ennan Zhai
NSDI5
2023 Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the Cloud
abstract
Request latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline.
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016
IEEE/ACM Trans. Netw.5
2023 Dependable Virtualized Fabric on Programmable Data Plane
abstract
In modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives.
Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010
IEEE/ACM Trans. Netw.7
2022 Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005
NSDI5
2022 Predictable vFabric on informative data plane
abstract
In multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works.
Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM7
2022 Firebolt: Finding Bugs in Programmable Data Plane Generators
Jiamin Cao, Yu Zhou 0008, Chen Sun 0005, Lin He 0004, Zhaowei Xi, Ying Liu 0024
USENIX ATC2
2022 TurboNet: Faithfully Emulating Networks With Programmable Switches
abstract
Faithfully emulating networks is critical for verifying the correctness and effectiveness of new networking-related designs. Existing network experiment platforms either cannot faithfully emulate the functionality and performance of production networks or cannot scale well due to cost constraints. In this paper, we proposeTurboNet, a new network emulator that utilizes one or more programmable switches to achieve faithful emulation of the network data plane and control plane. For data plane emulation, we propose a series of key designs, such as port mapper, queue mapper, and delayed queue, to emulate network topologies and performance metrics with high flexibility and accuracy. For control plane emulation, we support static routing configurations, distributed routing agents, and the centralized routing controllers. Meanwhile, we provide APIs for operators to simplify network emulation tasks. We implementTurboNeton Tofino switches. Evaluation results show that: (1) On the data plane,TurboNetcan flexibly emulate various topologies, such as an 8-ary fat-tree with only one programmable switch and a 10-ary fat-tree with four programmable switches; (2) On the control plane,TurboNetsupports about 200 BGP agents on a single programmable switch with a CPU usage of 25%; (3)TurboNetcan accurately emulate different network performance metrics such as 10−8link loss, and microsecond to millisecond link delay.
Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Mingwei Xu 0001
IEEE/ACM Trans. Netw.3
2022 Newton: Intent-Driven Network Traffic Monitoring
abstract
Network monitoring systems are designed to fulfill operators’ intents and serve as essential tools to modern networks. As a result of rapidly increasing network bandwidth and scale nowadays, network monitors should satisfy on-demand network monitoring for continuously growing traffic volumes. However, existing monitoring systems either cannot satisfy flexible intents on demand or produce significant overheads. In this paper, we presentNewton, an intent-driven traffic monitor that is able to specify operators’ intents with traffic monitoring queries and conduct dynamic and scalable network-wide queries deployment.Newtonenables operators to customize and modify queries dynamically without interrupting the network workflow. Besides,Newtonproposes systematic optimizations at device level and network-wide level to reduce resource consumption while deploying queries.Newtoncan combine the resources across switches to deploy complex queries with high resilience to dynamic network status. Evaluations prove thatNewtonis of high flexibility, scalability, and resource efficiency, which demonstratesNewtonis promising to be deployed in large-scale programmable networks.
Zhaowei Xi, Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001
IEEE/ACM Trans. Netw.2
2022 CoFilter: High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal Servers
abstract
As one of the most critical cloud services, Bare-Metal Servers (BMS) introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for BMS. However, the off-the-shelf stateful packet filters either are costly for cloud DCNs or introduce significant performance bottlenecks. In this article, we presentCoFilter, which leverages low-cost programmable switches to accelerate the stateful packet filter for BMS.CoFilteruses (1)stateful process partitionto enable complex stateful packet filtering logic on programmability-limited switching ASICs, (2)state compressionto track tens of millions of connections with constrained hardware memory, and (3)per-tenant packet rate limit and tenant-aware flow migrationto achieve efficient performance isolation among different tenants. Overall,CoFilterimplements a high-performance stateful packet filter via the co-design of programmable switching ASIC and CPU. We evaluateCoFilterunder various data center traffic traces with real-world flow distributions. The evaluation results show thatCoFilterremarkably outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay within 1us, and freeing a significant quantity of CPU cores, with rather small memory usage, i.e., accommodating over$10^7$connections with only 16MB SRAM.
Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Chen Sun 0005, Yangyang Wang 0001, Mingwei Xu 0001
IEEE Trans. Parallel Distributed Syst.3
2022 NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation
abstract
In distributed storage systems, Erasure Coding (EC) is a crucial technology to enable high data availability. By downloading parity data from survived machines, EC can reconstruct lost data with much lower storage overheads than data replication. However, this reduction in storage cost comes at the expense of extra performance problems:low reconstruction rate,high degraded read latency, andhigh host CPU utilization. Our analysis shows that these performance problems are deeply rooted in thehost-basedEC processing. To resolve these problems, we present NetEC, an in-network accelerating framework that fully offloads EC to the new generation programmable switching ASICs. We propose Explicit Buffer Size Notification (EBSN) to constrain decoding buffer usage, and design an on-switch one-to-many TCP proxy to integrate EBSN with TCP. We also design two parallel Galois Field (GF) offloading methods—table lookup and bitmatrix methods—to maximize parsable bytes. We implement NetEC on programmable switches and integrate it with HDFS. Extensive evaluations show that NetEC improves the reconstruction rate by 2.7x-6.8x, reduces the degraded read latency significantly, and removes the host CPU overhead completely. We also emulate multi-rack scenarios and show that NetEC is able to support$\sim$∼GB/s reconstruction rate and tens of concurrent tasks.
Yi Qiao, Menghao Zhang 0001, Yu Zhou 0008, Han Zhang 0009, Mingwei Xu 0001, Jun Bi, Jilong Wang 0001
IEEE Trans. Parallel Distributed Syst.3
2021 DOVE: Diagnosis-driven SLO Violation Detection
abstract
Service-level objectives (SLOs), as network performance requirements for delay and packet loss typically, should be guaranteed for increasing high-performance applications, e.g., telesurgery and cloud gaming. However, SLO violations are common and destructive in today’s network operation. Detection and diagnosis, meaning monitoring performance to discover anomalies and analyzing causality of SLO violations respectively, are crucial for fast recovery. Unfortunately, existing diagnosis approaches require exhaustive causal information to function. Meanwhile, existing detection tools incur large overhead or are only able to provide limited information for diagnosis. This paper presents DOVE, a diagnosis-driven SLO detection system with high accuracy and low overhead. The key idea is to identify and report the information needed by diagnosis along with SLO violation alerts from the data plane selectively and efficiently. Network segmentation is introduced to balance scalability and accuracy. Novel algorithms to measure packet loss and percentile delay are implemented completely on the data plane without the involvement of the control plane for fine-grained SLO detection. We implement and deploy DOVE on Tofino and P4 software switch (BMv2) and show the effectiveness of DOVE with a use case. The reported SLO violation alerts and diagnosis-needing information are compared with ground truth and show high accuracy (>97%). Our evaluation also shows that DOVE introduces up to two orders of magnitude less traffic overhead than NetSight. In addition, memory utilization and required processing ability are low to be deployable in real network topologies.
Yiran Lei, Yu Zhou 0008, Yunsenxiao Lin, Mingwei Xu 0001, Yangyang Wang 0001
ICNP2
2021 SketchINT: Empowering INT with TowerSketch for Per-flow Per-switch Measurement
abstract
1Network measurement is indispensable to network operations. Two most promising measurement solutions are In-band Network Telemetry (INT) solutions and sketching solutions. INT solutions provide fine-grained per-switch per-packet information at the cost of high network overhead. Sketching solutions have low network overhead but fail to achieve both simplicity and accuracy for per-flow measurement. To keep their advantages, and at the same time, overcome their shortcomings, we first design SketchINT to combine INT and sketches, aiming to obtain all per-flow per-switch information with low network overhead. Second, for deployment flexibility and measurement accuracy, we design a new sketch for SketchINT, namely TowerSketch, which achieves both simplicity and accuracy. The key idea of TowerSketch is to use different-sized counters for different arrays under the property that the number of bits used for different arrays stays the same. TowerSketch can automatically record larger flows in larger counters and smaller flows in smaller counters. We have fully implemented our SketchINT prototype on a testbed consisting of 10 switches. We also implement our TowerSketch on P4, single-core CPU, multi-core CPU, and FPGA platforms to verify its deployment flexibility. Extensive experimental results verify that 1) TowerSketch achieves better accuracy than prior art on various tasks, outperforming the state-of-the-art ElasticSketch up to 13.9 times in terms of error; 2) Compared to INT, SketchINT reduces the number of packets in the collection process by 3 4 orders of magnitude with an error smaller than 5%.
Kaicheng Yang 0001, Yuanpeng Li 0002, Zirui Liu 0002, Tong Yang 0003, Yu Zhou 0008, Jintao He, Jing'an Xue, Zhengyi Jia, Yongqiang Yang
ICNP5
2021 ECRaft: A Raft Based Consensus Protocol for Highly Available and Reliable Erasure-Coded Storage Systems
abstract
Erasure-coded redundancy is a fault-tolerant method with low-cost storage overhead. It only stores data fragments and parity fragments rather than full data across the cluster. The write process of erasure-coded data can be asynchronous or synchronous. For synchronous write process, data are encoded when written to servers. The common method doing the process needs to confirm that each coded-fragment of the data is stored in a different server to maintain the best fault tolerance. This method underperforms in terms of availability, and also fails to achieve good performance because any failure of servers will shortly disturb the write process. Some consensus protocols such as RS- Paxos and CRaft, which are based on Paxos and Raft, can solve above problems by providing fault-tolerant ability for systems. However, RS-Paxos cannot achieve the same liveness as Paxos. CRaft still adopts full data redundancy to keep the same liveness as Raft when there are not enough healthy servers. Therefore, to solve the availability problem during synchronous erasure-coded data write process, we present a novel protocol ECRaft based on Raft. It always uses erasure-coded redundancy when the ratio of erasure-coded data fragments to parity fragments is bigger than 1. It also can reach the same liveness as Raft. With state machine purge, storage redundancy can be reduced to the extent that typical erasure-coded storage systems can achieve. We build a key-value store based on ECRaft to evaluate it. In our experiments, compared with CRaft using complete-entry replication, ECRaft can save 63 % of storage, increase write throughput by 28.2 %, and reduce write latency by 19 %.
Mingwei Xu 0001, Yu Zhou 0008, Yuanyuan Qiao 0002, Yu Wang 0096, Jie Yang 0023
ICPADS2
2021 Aquila: a practically usable verification system for production-scale programmable data planes
abstract
This paper presents Aquila, the first practically usable verification system for Alibaba's production-scale programmable data planes. Aquila addresses four challenges in building a practically usable verification: (1) specification complexity; (2) verification scalability; (3) bug localization; and (4) verifier self validation. Specifically, first, Aquila proposes a high-level language that facilitates easy expression of specifications, reducing lines of specification codes by tenfold compared to the state-of-the-art. Second, Aquila constructs a sequential encoding algorithm to circumvent the exponential growth of states associated with the upscaling of data plane programs to production level. Third, Aquila adopts an automatic and accurate bug localization approach that can narrow down suspects based on reported violations and pinpoint the culprit by simulating a fix for each suspect. Fourth and finally, Aquila can perform self validation based on refinement proof, which involves the construction of an alternative representation and subsequent equivalence checking. To this date, Aquila has been used in the verification of our production-scale programmable edge networks for over half a year, and it has successfully prevented many potential failures resulting from data plane bugs.
Bingchuan Tian, Mengqi Liu 0001, Ennan Zhai, Yu Zhou 0008, Mengjing Ma, Xionglie Wei, Hongqiang Harry Liu, Ming Zhang 0005, Chen Tian 0001, Minlan Yu
SIGCOMM6
2021 HyperTester: High-Performance Network Testing Driven by Programmable Switches
abstract
Modern network devices and systems are raising higher requirements on network testers that are regularly used to evaluate performance and assess correctness. These requirements include high scale, high accuracy, flexibility and low cost, which existing testers cannot fulfill at the same time. In this paper, we propose HyperTester, a network tester leveraging new-generation programmable switches and achieving all of the above goals simultaneously. Programmable switches are born with features like high throughput and linerate, deterministic processing pipelines and nanosecond-level hardware timestamps, the P4 programming model as well as comparable pricing with commodity servers, but they come with limited programmability and memory resources. HyperTester uses template-based packet generation to overcome the limitations of the switch ASIC in programmability and designs a stateless connection mechanism as well as counter-based state compression algorithms to overcome the memory resource constraints in the data plane. We have implemented HyperTester on Tofino, and the evaluations on the hardware testbed show that HyperTester supports high-scale packet generation (more than 1.6Tbps) and achieves highly accurate rate control and timestamping. We demonstrate that programmable switches can be potential and attractive targets for realizing network testers.
Yu Zhou 0008, Zhaowei Xi, Yangyang Wang 0001, Mingwei Xu 0001
IEEE/ACM Trans. Netw.2
2020 Newton: intent-driven network traffic monitoring
abstract
Monitoring network traffic based on operators' intents is essential to today's networks. As the bandwidth and size of networks increase steeply, monitoring systems shall fulfill the requirements of on-demand network monitoring for ever-growing traffic volumes. However, existing monitoring systems either cannot satisfy operators' intents on demand or introduce substantial monitoring overheads. In this paper, we present Newton, an intent-driven traffic monitor that enables specifying operators' intents with traffic monitoring queries and supports dynamic and scalable network-wide queries. Specifically, Newton 1) empowers operators to dynamically create, remove, and update on-data-plane queries without interrupting normal packet forwarding, 2) conducts systematic optimizations to achieve precise network traffic monitoring, and 3) executes network-wide queries with high resilience to dynamic network status. Evaluation results show that Newton improves the flexibility, scalability, and resource efficiency of traffic monitoring, demonstrating its great potential to be deployed in large-scale programmable networks.
Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001
CoNEXT1
2020 NetView: Towards On-Demand Network-Wide Telemetry in the Data Center
abstract
Network telemetry is to collect information (e.g., hop latency, throughput) from network devices. Network-wide telemetry is critical for operators to understand the quality of network performance and to diagnose on-going failures. The state-of-the-art telemetry approaches are far from ideal as they are unable to fully satisfy diverse requirements of operators, specifically for on-demand, full coverage, and scalable telemetry. In this paper, we provide a new framework of network telemetry for data center networks, called NetView. NetView can support various telemetry applications and frequencies on demand, monitoring each device via proactively sending dedicated probes. Technically, NetView leverages source routing to forward probes, achieving full coverage. Besides, a series of probe generation algorithms largely reduce probe number, providing high scalability. The evaluation shows that NetView reduces the bandwidth occupancy by more than two orders of magnitude compared with Pingmesh and INT-path, and conducts network-wide telemetry for large-scale data center network using only one vantage server, without bringing about resources bottleneck.
Yunsenxiao Lin, Yu Zhou 0008, Zhengzheng Liu, Yangyang Wang 0001, Mingwei Xu 0001, Jun Bi, Ying Liu 0024
ICC2
2020 TurboNet: Faithfully Emulating Networks with Programmable Switches
abstract
Faithfully emulating networks is critical for verifying the correctness and effectiveness of new networking-related designs. Existing network experiment platforms either cannot faithfully emulate functionality and performance of production networks or cannot scale well because of cost limitations. In this paper, we propose TurboNet, a new network emulator that leverages one programmable switch to enable faithful emulation of both network data plane and control plane. For data plane emulation, we present a series of key designs such as port mapper, queue mapper, and delayed queue to emulate network topologies and performance metrics with high flexibility and accuracy. For control plane emulation, we support static routing configurations, distributed routing agents, and the centralized routing controller. Meanwhile, we provide API for operators to simplify network emulation tasks. We implement TurboNet on a Tofino switch. The evaluation results show that: (1) TurboNet can flexibly emulate various topologies such as the 8-ary fat-tree on the data plane and support about 200 BGP agents with 25% CPU usage on the control plane; (2) TurboNet can accurately emulate different network performance metrics, including 400Gbps linerate background traffic injection, as small as 10-8link loss, and microsecond-level to millisecond-level link delay.
Jiamin Cao, Yu Zhou 0008, Ying Liu 0024, Mingwei Xu 0001, Yongkai Zhou
ICNP2
2020 FlexMesh: Flexibly Chaining Network Functions on Programmable Data Planes at Runtime
Yu Zhou 0008, Jun Bi, Cheng Zhang 0012, Mingwei Xu 0001, Jinaping Wu
Networking1
2020 Flow Event Telemetry on Programmable Data Plane
abstract
Network performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic.
Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001
SIGCOMM1
2020 Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICs
abstract
Programmable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs.
Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu
SIGCOMM5
2020 NetView: Towards on-demand network-wide telemetry in the data center
Yunsenxiao Lin, Yu Zhou 0008, Zhengzheng Liu, Yangyang Wang 0001, Mingwei Xu 0001, Jun Bi, Ying Liu 0024
Comput. Networks2
2020 VMS: Load Balancing Based on the Virtual Switch Layer in Datacenter Networks
abstract
There have been many load balancing solutions for datacenter networks. Almost all of them require modifications to the network fabric or/and virtual machines. Recently, the virtual switch layer becomes an ideal location for datacenter operators to deal with the load balancing problem. In this paper, we propose Virtual Multi-channel Scatter (VMS), a packet-level load balancing design in the virtual switch layer. VMS scatters packets in one TCP flow to several different forwarding paths (channels). VMS has several noteworthy properties. First, VMS is low cost and transparent to tenants. It can be deployed when the datacenter operators do not attempt to change the network fabric or cannot control the transport protocol inside VMs. Second, by employing window-based channel selection, VMS is adaptive to network congestion and topology asymmetry. Third, VMS works well with Generic Segmentation Offload/Generic Receive Offload (GRO/GSO) mechanism in the Linux kernel, unlike other packet-level load balancing schemes. Finally, VMS can also be offloaded to SmartNIC to reduce CPU overhead further. Our evaluations show that VMS achieves comparable performance to the ideal packet-level scheme in normal cases and well handles topology asymmetries, while only modifies the virtual switch layer. In the symmetric topology, VMS achieves up to 47% and 22% better flow completion time (FCT) than Equal Cost MultiPath (ECMP) and the best-of-breed flowlet-level CONGA. When there is topology asymmetry, VMS outperforms the ideal packet-level scheme and CONGA by up to $3.0\times $ and $1.4\times $ respectively. Further, the overhead of VMS is tolerable.
Jun Bi, Zhaogeng Li, Yu Zhou 0008, Yangyang Wang 0001
IEEE J. Sel. Areas Commun.4
2020 HyperSight: Towards Scalable, High-Coverage, and Dynamic Network Monitoring Queries
abstract
Performing fine-grained and real-time network monitoring is the core logic of various data center operation applications, such as traffic engineering, network troubleshooting, and anomaly detecting. However, the state-of-the-art network monitoring solutions either fall short of completely detecting all network incidents (i.e., congestion), yielding limited monitoring coverage, or introduce large overheads, yielding limited scalability. In this paper, we present HyperSight, a network traffic monitor with both high coverage and low overheads. The key idea of HyperSight is to monitor networks at the behavior level via tracking packet behavior changes. HyperSight proposes three designs for behavior-level monitoring. First, to facilitate expressing various network monitoring tasks, HyperSight presents a declarative query language based on the streaming processing model. Second, HyperSight proposes Bloom Filter Queue (BFQ), a memory-efficient algorithm to empower in-network capability for monitoring packet behavior changes. BFQ can be implemented on commodity programmable switches. Third, to support dynamic deployment and execution of packet behavior change monitoring tasks without interrupting on-service switches, HyperSight proposes virtual BFQ to support dynamic query compilation. We build a prototype of HyperSight and deploy it on commodity programmable switches. Evaluation results show that HyperSight supports a wide range of network event queries and can monitor over 99% packet behavior changes while keeping remarkably low overheads.
Yu Zhou 0008, Jun Bi, Tong Yang 0003, Kai Gao 0001, Jiamin Cao, Yangyang Wang 0001, Cheng Zhang 0012
IEEE J. Sel. Areas Commun.1
2019 HyperTester: high-performance network testing driven by programmable switches
abstract
Modern network research and operations are inseparable from network testers to evaluate performance limits of proofs-of-concept, troubleshoot failures, etc. Existing network testers suffer from either constrained flexibility or a low performance-cost ratio. In this paper, we propose a new network tester, HyperTester. The core of HyperTester is to leverage new-generation programmable switches for generating and capturing test traffic with high performance, low cost, and remarkable flexibility. We design a series of efficient mechanisms, including template-based packet generation, false-positive-free counter-based queries, and stateless connections to realize various network testing tasks upon switches with limited programmability and resources. Meanwhile, to facilitate developing testing tasks upon HyperTester, we provide a high-level network testing API. We have implemented HyperTester on the Tofino switch and built dozens of network testing tasks. The evaluations on the hardware testbed show that HyperTester supports line-rate packet generation (400Gbps in the testbed) with highly-accurate rate control, while HyperTester can save $40150 per Tps and 9225W per Tbps when compared with the software network testers.
Yu Zhou 0008, Zhaowei Xi, Yangyang Wang 0001, Jinqiu Wang, Mingwei Xu 0001
CoNEXT1
2019 CoFilter: A High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal Servers
abstract
As one of the most critical cloud services, Bare-metal Servers introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for bare-metal servers. However, the off-the-shelf hardware-based and software-based stateful packet filters either are prohibitively costly for cloud DCNs or introduce significant performance bottlenecks. In this paper, we present CoFilter, which employs cheap programmable switches to accelerate the stateful packet filter for bare-metal servers. CoFilter consists of two key designs. First, to support complex stateful packet filtering logic in programmability-limited switching ASICs, CoFilter partitions the stateful packet filtering logic between programmable ASICs and switch CPU. Most packets are directly processed in switching ASICs to achieve high performance, while only a small number of packets go to switch CPU for connection tracking. Second, to track massive connections with constrained hardware memory, CoFilter employs hash to compress connection states and provides an efficient settlement for hash collisions. We build a prototype of CoFilter and evaluate it on the Tofino switch under various data center traffic traces with real-world flow distribution. The evaluation shows that CoFilter largely outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay at 1us, and freeing a significant quantity of CPU cores. Furthermore, CoFilter presents great scalability and accommodates over ten million connections with only 16MB SRAM.
Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Chen Sun 0005, Yangyang Wang 0001, Jun Bi
ICCCN3
2019 P4Tester: efficient runtime rule fault detection for programmable data planes
abstract
P4 and programmable data planes bring significant flexibility to network operation but are inevitably prone to various faults. Some faults, like P4 program bugs, can be verified statically, while some faults, like runtime rule faults, only happen to running network devices, and they are hardly possible to handle before deployment. Existing network testing systems can troubleshoot runtime rule faults via injecting probes, but are insufficient for programmable data planes due to large overheads or limited fault coverage. In this paper, we propose P4Tester, a new network testing system for troubleshooting runtime rule faults on programmable data planes. First, P4Tester proposes a new intermediate representation based on Binary Decision Diagram, which enables efficient probe generation for various P4-defined data plane functions. Second, P4Tester offers a new probe model that uses source routing to forward probes. This probe model largely reduces rule fault detection overheads, i.e. requiring only one server to generate probes for large networks and minimizing the number of probes. Moreover, this probe model can test all table rules in a network, achieving full fault coverage. Evaluation based on real-world data sets indicates that P4Tester can efficiently check all rules in programmable data planes, generate 59% fewer probes than ATPG and Pronto, be faster than ATPG by two orders of magnitude, and troubleshoot multiple rule faults within one second on BMv2 and Tofino.
Yu Zhou 0008, Jun Bi, Yunsenxiao Lin, Yangyang Wang 0001, Zhaowei Xi, Jiamin Cao, Chen Sun 0005
IWQoS1
2019 HyperVDP: High-Performance Virtualization of the Programmable Data Plane
abstract
With the advent of P4-specific programmable data plane (PDP), network functions (NFs) can be offloaded into the PDP to achieve high performance guaranteed by hardware. Meanwhile, CPU powers consumed by NFs can be released to user applications. However, as more and more NFs can be offloaded, several problems rooted inside the PDP severely hinder it from facilitating this offloading trend. 1) The existing PDP provides the exclusive data plane abstraction where different NFs cannot operate the same data plane. 2) The PDP is hardly able to deploy NFs in a “hitless” manner. In this paper, we propose HyperVDP as a high-performance data plane hypervisor to provision non-exclusive abstraction and uninterrupted reconfigurability on the P4-specific PDP. To achieve virtualization, we design several innovative techniques to equally express functions of all programmable elements in the P4-specific PDP. We implement the prototype of HyperVDP on different target platforms, and evaluate different target-based prototypes by comparing with their counterparts. Results show that BMv2-target HyperVDP averagely prevails over its counterpart 2.5× in performance and 4× in resource efficiency. DPDK-target HyperVDP performs comparably to its counterparts while offering virtualization features which neither of its counterparts could provide.
Cheng Zhang 0012, Jun Bi, Yu Zhou 0008
IEEE J. Sel. Areas Commun.3
2019 P4DB: On-the-Fly Debugging for Programmable Data Planes
abstract
While extending network programmability to a more considerable extent, P4 raises the difficulty of detecting and locating bugs, e.g., P4 program bugs and missed table rules, in runtime. These runtime bugs, without prompt disposal, can ruin the functionality and performance of networks. Unfortunately, the absence of efficient debugging tools makes runtime bug troubleshooting intricate for operators. This paper is devoted to on-the-fly debugging of runtime bugs for programmable data planes. We propose P4DB, a general debugging platform that empowers operators to debug P4 programs in three levels of visibility with rich primitives. By P4DB, operators can use the watch primitive to quickly narrow the debugging scope from the network level or the device level to the table level, then use the break and next primitives to decompose match-action tables and finely locate bugs. We implement a prototype of P4DB and evaluate the prototype on two widely-used P4 targets. On the software target, P4DB merely introduces a small throughput penalty (1.3% to 13.8%) and a little delay increase (0.6% to 11.9%). Notably, P4DB almost introduces no performance overhead on Tofino, the hardware P4 target.
Yu Zhou 0008, Jun Bi, Cheng Zhang 0012, Bingyang Liu, Zhaogeng Li, Yangyang Wang 0001, Mingli Yu
IEEE/ACM Trans. Netw.1
2018 NetVision: Towards Network Telemetry as a Service
abstract
In-band Network Telemetry (INT) can provide fine-grained and accurate device-level telemetry metrics. Nonetheless, INT can track only a small ratio of devices and links and embedding telemetry data into normal packets brings high overhead and high operation complexity. Hence, we present NetVision, a powerful proactive network telemetry platform with high coverage and high scalability.
Zhengzheng Liu, Jun Bi, Yu Zhou 0008, Yangyang Wang 0001, Yunsenxiao Lin
ICNP3
2018 KeySight: Troubleshooting Programmable Switches via Scalable High-Coverage Behavior Tracking
abstract
The rise of programmable switches and P4 brings much flexibility to networks, but this flexibility comes with increased risks of bugs. Diagnosing these bugs is essential for network operation but is non-trivial. A potential approach is to track packet behaviors through postcards, but existing tools either generate substantial postcards (limited scalability) or only track a small proportion of packet behaviors (low coverage). In this paper, we present KeySight, a platform that troubleshoots programmable switches with high scalability and high coverage. The key idea is based on the Packet Equivalence Class (PEC) abstraction that aggregates packets with identical behaviors and generates one postcard per behavior. The PEC abstraction minimizes the number of postcards while tracking all packet behaviors. We design novel algorithms to analyze PECs of P4 programs and to implement the PEC abstraction on programmable switches. We deploy KeySight on Tofino and SmartNIC, and evaluate it against 80 P4 programs and real packet traces of over 5TB. Results show that in the premise of overseeing over 99.9% packet behaviors, KeySight reduces the number of postcards by one to two orders of magnitude when comparing with NetSight.
Yu Zhou 0008, Jun Bi, Tong Yang 0003, Kai Gao 0001, Cheng Zhang 0012, Jiamin Cao, Yangyang Wang 0001
ICNP1
2018 B-Cache: A Behavior-Level Caching Framework for the Programmable Data Plane
abstract
By enabling operators to program behaviors of the packet processing pipeline, P4, a domain-specific language, unleashes new opportunities for offloading network functions onto the programmable data plane (PDP) and enhancing network performance. However, recent research shows that as P4 programs and the corresponding packet processing pipeline grow in size and complexity, the performance of the PDP will decrease significantly, which compromises the programmability and flexibility brought by P4. To overcome this performance degradation, we propose B-Cache, a general behavior-level caching framework for both stateful and stateless behaviors on the PDP. The basic idea of B-Cache is to compile packet processing behaviors that were once distributed across multiple tables into one synthetic cache table, thus guarantee the performance on various P4 targets. Our experiment results indicate that B-Cache comparably yields significant performance benefits including a 49% delay decrease and a 200% throughput increase on the software target, and a 60% throughput increase on the hardware target.
Cheng Zhang 0012, Jun Bi, Yu Zhou 0008, Keyao Zhang, Zijun Ma
ISCC3
2017 HyperV: A High Performance Hypervisor for Virtualization of the Programmable Data Plane
abstract
P4 is a domain specific language designed to define the behavior of a programmable data plane. It facilitates offloading hardware-suitable Network Functions (NFs) to a data plane. Consequently, NFs can maximally benefit from high performance of hardware devices, meanwhile more CPU power can be reserved for user applications. However, since the programmable data plane provides an NF with an exclusive network context, different NFs cannot operate on the same data plane simultaneously. Besides, it is hardly possible to dynamically reconfigure programmable network devices without interrupting the operation of a data plane. Therefore, we propose HyperV, a high performance hypervisor for virtualization of a P4 specific data plane, to provide both non-exclusive and uninterrupted features.We implemented HyperV based on a P4-BMv2 target and a DPDK target respectively. Then we evaluated BMv2-target HyperV by comparing with Hyper4, a recently proposed hypervisor, and evaluated DPDK- target HyperV by comparing with PISCES and Open vSwitch. Results show that BMv2- target HyperV averagely prevails over Hyper4 2.5x in performance while reducing resource usage by 4x. DPDK-target HyperV performs comparably to Open vSwitch and PISCES, with the worst case of a throughput penalty in less than 7\%, while providing a powerful capability of virtualization which neither of them provides.
Cheng Zhang 0012, Jun Bi, Yu Zhou 0008, Abdul Basit Dogar
ICCCN3
2017 P4DB: On-the-fly debugging of the programmable data plane
abstract
While extending network programmability to a larger degree, P4 also raises the risks of incurring runtime bugs after the deployment of P4 programs. These runtime bugs, if not handled promptly and properly, can ruin the functionality and performance of networks. Unfortunately, the absence of runtime debuggers makes troubleshooting of P4 program bugs challenging and intricate for operators. This paper is devoted to the on-the-fly debugging of runtime bugs in P4-enabled networks. We propose P4DB, a general debugging platform that empowers operators to debug P4 programs in three levels of visibility by provisioning operator-friendly primitives. By P4DB, operators can use the watch primitive to quickly narrow the debugging scope from network level or device level to table level, then use the break and next primitives to decompose the match-action table into three steps and troubleshoot the runtime bugs step by step. We implemented a prototype of P4DB and evaluated the performance in terms of the data plane, control plane and control channel. On P4-specific programmable data plane, P4DB merely introduces a small throughput penalty (1.3%~13.8%) and imposes a little-increased delay (0.6%~11.9%).
Cheng Zhang 0012, Jun Bi, Yu Zhou 0008, Bingyang Liu, Zhaogeng Li, Abdul Basit Dogar, Yangyang Wang 0001
ICNP3
2016 Source Address Validation in Software Defined Networks
abstract
In this paper, we present the preliminary design and implementation of SDN-SAVI, an SDN application that enables SAVI functionalities in SDN networks. In this proposal, all the functionalities are implemented on the controller without modifying SDN switches. To enforce SAVI on packets in the data plane, the controller installs binding tables in switches using existing SDN techniques, such as OpenFlow. With SDN-SAVI, a network administrator can now enforce SAVI in her network by merely integrating a module on the controller, rather than purchasing SAVI-capable switches and replacing legacy ones.
Bingyang Liu, Jun Bi, Yu Zhou 0008
SIGCOMM3