Yangyang Wang 0001

dblp:20/8967-1 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-2713-0748ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 28 · 4 first-author · 8 since 2021Security and privacy · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 JitterSketch: Finding Jittery Flows in Network Streams
abstract
In the modern internet, with the proliferation of real-time applications such as online gaming and video conferencing, the timely detection of network jitter has become a critical task in network measurement. Network jitter is defined as the abrupt fluctuations in packet inter-arrival times within network flows, which severely degrade the Quality of Service for these applications. Traditional jitter detection methods primarily focus on macro-level end-to-end or hop-by-hop latency variations, neglecting the fine-grained jitter that occurs within specific flows. In this paper, we present JitterSketch, the first sketch-based algorithm specifically designed for detecting jittery flows. JitterSketch employs a novel three-stage structure to efficiently filter out infrequent and stable flows, thereby identifying and reporting the jittery flows that have the most significant impact on network quality. Extensive experiments demonstrate that JitterSketch achieves an improvement of up to 50 percentage points in both recall and precision rates compared to baseline solutions, while maintaining high processing throughput. Furthermore, we deployed JitterSketch in a QoS simulation system, where it yielded significant improvements in QoS.
Zhongxian Liang, Qilong Shi, Xiyan Liang, Wenjun Li 0004, Tong Yang 0003, Yangyang Wang 0001, Mingwei Xu 0001, Weizhe Zhang
WWW7
2026 A large-scale measurement study of region-based web access restrictions: The case of China
Yuying Du, Jiahao Cao 0001, Junrui Xu, Yangyang Wang 0001, Renjie Xie, Changliyun Liu, Mingwei Xu 0001
Comput. Secur.4
2025 Right the Ship: Assessing the Legitimacy of Invalid Routes in RPKI
abstract
Resource Public Key Infrastructure (RPKI) aims to prevent prefix hijacking by providing secure mappings between IP prefixes and their authorized origin Autonomous Systems (ASes). In recent years, there has been notable growth in the deployment of RPKI and Route Origin Validation (ROV). Nonetheless, over 40% of the routes in the global routing table still lack the protection of RPKI. One of the critical reasons some networks are reluctant to deploy RPKI is the concern that some ROV-invalid routes may be legitimate, and filtering these routes will harm network service quality, especially affecting network connectivity.
Yangyang Wang 0001, Jia Zhang 0010, Mingwei Xu 0001
CCS2
2025 Poster: Uncovering Hidden ASes in ROV Deployment via Temporal Fingerprinting
abstract
BGP, the Internet’s inter-domain routing protocol, is vulnerable to prefix hijacks due to tamperable prefix–origin bindings. RPKI addresses this by cryptographically binding prefixes to authorized ASes, enabling Route Origin Validation (ROV). Given RPKI’s critical role in Internet security, identifying the proportion of ASes that deploy ROV is a key research question. However, measuring ROV deployment is difficult due to limited visibility into private AS configurations. Existing methods suffer from low accuracy, restricted coverage, and the inability to detect hidden ASes whose behavior is masked by upstream ROV deployment. To address these limitations, we propose RIFT, a novel inference method based on temporal fingerprinting. RIFT leverages the insight that periodic ROA retrieval creates distinctive temporal patterns in the routing behavior of ROV-deployed ASes. Experiments show that RIFT achieves 94% accuracy and an F1 score of 0.88 in identifying ROV deployment.
Shucan Yang, Jiahao Cao 0001, Mingwei Xu 0001, Renjie Xie, Yangyang Wang 0001
ICNP9
2025 GeoWatch: Measuring Geoblocking Practices Towards China at Scale
abstract
This paper presents GeoWatch to conduct the first large-scale measurement study of geoblocking practices towards China. Although prior studies have examined geoblocking in specific contexts as well as China's censorship, GeoWatch focuses on identifying which websites proactively block users from China. It employs advanced domain mining techniques and globally distributed vantage points to identify geoblocking websites, analyzing a total of 97.78 million domains worldwide and identifying 4.54 million geoblocking domains.
Yuying Du, Jiahao Cao 0001, Junrui Xu, Yangyang Wang 0001, Changliyun Liu
IWQoS4
2025 HeavyFinder: Efficient and Fine-Grained Heavy Hitters Detection with Instantaneous Flow Rate
abstract
The detection of Heavy Hitters (HH) of flows in a network plays a crucial role in a variety of critical applications, including network topology optimization, congestion control, and network security (e.g., DDoS mitigation). In high-performance data centers, tasks such as optimizing model training and detecting anomalies require microsecond-level granularity for burst and congestion detection, demanding higher accuracy and fine granularity in HH detection. However, existing HH detection schemes primarily focus on traffic accumulation over longer periods and ignore instantaneous flow rates, making them incapable of detecting short-duration heavy hitters (at the granularity of milliseconds to microseconds) that can significantly affect network performance. This paper proposes a rate-sensitive definition of HHs and introduces HeavyFinder, a framework for enabling efficient HH detection of flows at microsecond granularity and accurately describing their traffic changes. It detects HHs based on inherent characteristics of both traffic volume and instantaneous flow rates, and improves the existing elephant flow filtering mechanism. Furthermore, this framework provides an efficient information aggregation method for reporting HH information, which can reduce overhead further. We deployed and tested HeavyFinder on x86 CPUs and evaluated its performance by using real network trace data. Results show that HeavyFinder achieves sub-$\mathbf{1 0}$-microsecond accuracy in detection for the start and end times of HHs, with$3-10 \times$lower reporting overhead compared to existing frameworks.
Yangyang Wang 0001, Jiahao Cao 0001, Qilong Shi, Mingwei Xu 0001, Lihua Miao
IWQoS2
2025 Cooled-KLL: Enhancing Quantile Estimation by Filtering Hot Item
abstract
Quantile estimation is critical for diverse applications, including database management and network traffic monitoring. Probabilistic quantile sketches are widely employed in practice, with the KLL sketch (introduced in 2016) being particularly notable for its theoretically space-optimal properties. However, KLL overlooks the inherent repetition of elements often present in real-world data streams. Such streams are frequently highly skewed, characterized by ''hot items''-items that appear with high frequency. The KLL sketch processes these hot items without accounting for their prevalence, resulting in suboptimal space utilization due to redundant insertions and storage. To overcome this limitation, we propose Cooled-KLL, an enhanced KLL sketch. Cooled-KLL introduces a novel ''Hot Filter'' structure that efficiently identifies and stores hot items as compact key-value pairs. This mechanism ensures that only ''cold'' (less frequent) items are subsequently processed by the core KLL sketch. Our approach significantly reduces memory consumption without compromising processing speed. Extensive experiments demonstrate that Cooled-KLL consistently outperforms five other state-of-the-art algorithms, achieving up to 2.5 orders of magnitude higher accuracy compared to the standard KLL sketch.
Qilong Shi, Wei Zhou 0077, Yizhuo Zheng, Xinye Xu, Yuanyuan Zhang 0006, Long Yao, Yangyang Wang 0001, Mingwei Xu 0001
KDD (2)8
2025 HeavyLocker: Lock Heavy Hitters in Distributed Data Streams
abstract
In recent years, sketching has emerged as a pivotal technique for identifying heavy hitters (items with high frequency) in large-scale data streams. Despite this progress, the majority of existing sketch algorithms are tailored primarily for detecting local heavy hitters within a single data stream, with only a few capable of extending their application to global heavy hitters across distributed data streams. A common challenge encountered by these algorithms is balancing performance with accuracy. To address this challenge, we introduce HeavyLocker, a novel sketch algorithm that takes advantage of a distinct feature of real data streams: the separability of heavy hitters. By leveraging this attribute, HeavyLocker precisely locks and protects potential heavy hitters during the data stream processing, ensuring accuracy in local heavy hitter detection without compromising on speed. This unique capability also facilitates its application to global detection tasks. Through theoretical analysis, we validate the efficacy of HeavyLocker's locking mechanism. Our extensive experiments show that HeavyLocker outperforms five benchmarked algorithms in accuracy and maintains fast speed for both local and global heavy hitter detection, significantly reducing errors by up to an order of magnitude compared to the renowned Double-Anonymous Sketch.
Qilong Shi, Hanyue Zheng, Tong Yang 0003, Yangyang Wang 0001, Mingwei Xu 0001
KDD (1)5
2025 VPGFuzz: Vulnerable Path-Guided Greybox Fuzzing
abstract
Fuzzing is a prevalent technology for identifying software vulnerabilities. Existing fuzzing techniques predominantly focus on maximizing code coverage to unearth potential security issues. However, the mere expansion of explored code does not necessarily correlate with an increased discovery of vulnerabilities. Additionally, existing fuzzers often neglect comprehensive execution path information in code exploration. Consequently, potential vulnerabilities may be delayed or overlooked in the fuzzing process. To address this, we propose VPGFUZZ, a vulnerable path-guided fuzzer that can not only explore new code but also exploit known vulnerability path knowledge for vulnerability discovery. It employs a vulnerable path recognition model to identify test cases with potentially vulnerable paths. This model is trained with various execution paths derived from real-world vulnerability PoCs (Proof of Concepts). Based on this model, VPGFUZZ applies an explore-exploit seed selection strategy to effectively choose test cases for testing. Unlike traditional seed selection methods that maintain a single queue for exploring new code, this strategy includes a separate queue for retaining test cases identified as potentially vulnerable, allowing for more thorough testing. Experimental results demonstrate that VPGFUZZ discovers 24 zero-day vulnerabilities, with 18 receiving vulnerability identifiers from third-party organizations such as CVE. Our evaluation also shows VPGFUZZ’s superior efficiency by uncovering the first vulnerability approximately 1.2 to 70 times faster than popular fuzzers in most programs.
Zhechao Lin, Jiahao Cao 0001, Xinda Wang 0001, Renjie Xie, Yuxi Zhu, Xiao Li 0044, Qi Li 0002, Yangyang Wang 0001, Mingwei Xu 0001
IEEE Trans. Inf. Forensics Secur.8
2024 Poster: Few-Shot Inter-Domain Routing Threat Detection with Large-Scale Multi-Modal Pre-Training
abstract
Border Gateway Protocol (BGP) plays a pivotal role as the de facto inter-domain routing protocol on the Internet. However, BGP threats continually emerge and undermine the Internet reliability. Existing BGP threat detection methods based on machine learning require substantial labeled data and expert involvement, making them costly and labor-intensive. Moreover, they fail to learn rich information from massive unlabeled BGP data consistently generated on the Internet. In this paper, we propose FIRE that enables few-shot inter-domain routing threat detection with large-scale multi-modal pre-training. FIRE conducts domain-specific pre-training tasks to acquire rich BGP implicit knowledge from massive unlabeled BGP data for few-shot learning. Our experiments show that FIRE can be fine-tuned to precisely identify BGP threats with only a few labeled samples, e.g., a 93.2% precision in route leak detection with merely 8 events for fine-tuning.
Jiahao Cao 0001, Renjie Xie, Yangyang Wang 0001, Mingwei Xu 0001
CCS5
2024 Bubble Sketch: A High-performance and Memory-efficient Sketch for Finding Top-k Items in Data Streams
abstract
Sketch algorithms are crucial for identifying top-k items in large-scale data streams. Existing methods often compromise between performance and accuracy, unable to efficiently handle increasing data volumes with limited memory. We present Bubble Sketch, a compact algorithm that excels in both performance and accuracy. Bubble Sketch achieves this by (1) Recording only full keys of hot items, significantly reducing memory usage, and (2) Using threshold relocation to resolve conflicts, enhancing detection accuracy. Unlike traditional methods, Bubble Sketch eliminates the need for a Min-Heap, ensuring fast processing speeds. Experiments show Bubble Sketch outperforms the other seven algorithms compared, with the highest throughput and precision, and surpasses HeavyKeeper in accuracy by up to two orders of magnitude.
Qilong Shi, Yuxi Liu 0017, Hanyue Zheng, Yao Xin, Wenjun Li 0004, Tong Yang 0003, Yangyang Wang 0001, Yang Xu 0010, Weizhe Zhang, Mingwei Xu 0001
CIKM8
2023 SD-INT: Towards Lightweight Network-Wide Passive INT in the Self-Driving Way
abstract
The In-band Network Telemetry (INT) provides unprecedented network visibility by encapsulating fine-grained device-internal status into packet headers. Existing efforts proposed to reduce the telemetry cost can either monitor a small part of flows and links, or require complex calculation supported by programmable switches. In this paper, we propose SD-INT (Self-Driving INT), a lightweight network-wide passive INT system, which can be readily deployed based on commodity switches. The key idea of SD-INT is to reduce two dimensions of redundancy in the INT data which we have observed. Furthermore, the optimized INT is conducted in a self-driving system, adjusting the INT strategy according to the feedback of INT result history. In addition to the mechanism, we design the latency-hiding technology and a suite of efficient algorithms for flow selection and sampling ratio adaptation. These designs enable us to keep the monitoring overhead under a reasonable level, while still collecting network-wide fine-grained telemetry data. Further, we prototype and evaluate SD-INT with both large-scale simulations and testbed deployment. Experiments for both TCP(Transmission Control Protocol) and RDMA(Remote Direct Memory Access) flows show that SD-INT reduces orders of magnitude data volume while achieves similar link coverage compared with INT. Besides, compared with the relevant technologies based on programmable switches, SD-INT is competitive in reducing the data volume.
Yunsenxiao Lin, Yangyang Wang 0001, Mingwei Xu 0001, Yongfeng Ni, Shaowen Zheng, Kehan Yao
ICNP4
2022 BGP-Multipath Routing in the Internet
abstract
BGP-Multipath (BGP-M) is a multipath routing technique for load balancing. Distinct from other techniques deployed at a router inside an Autonomous System (AS), BGP-M is deployed at a border router that has installed multiple inter-domain border links to a neighbor AS. It uses the equal-cost multi-path (ECMP) function of a border router to share traffic to a destination prefix on different border links. Despite recent research interests in multipath routing, there is little study on BGP-M. Here we provide the first measurement and a comprehensive analysis of BGP-M routing in the Internet. We extracted information on BGP-M from query data collected from Looking Glass (LG) servers. We revealed that BGP-M has already been extensively deployed and used in the Internet. A particular example is Hurricane Electric (AS6939), a Tier-1 network operator, which has implemented >1,000 cases of BGP-M at 69 of its border routers to prefixes in 611 of its neighbor ASes, including many hyper-giant ASes and large content providers, on both IPv4 and IPv6 Internet. We examined the distribution and operation of BGP-M. We also ran traceroute using RIPE Atlas to infer the routing paths, the schemes of traffic allocation, and the delay on border links. This study provided the state-of-the-art knowledge on BGP-M with novel insights into the unique features and the distinct advantages of BGP-M as an effective and readily available technique for load balancing.
Jie Li 0117, Vasileios Giotsas, Yangyang Wang 0001, Shi Zhou
IEEE Trans. Netw. Serv. Manag.3
2022 Newton: Intent-Driven Network Traffic Monitoring
abstract
Network monitoring systems are designed to fulfill operators’ intents and serve as essential tools to modern networks. As a result of rapidly increasing network bandwidth and scale nowadays, network monitors should satisfy on-demand network monitoring for continuously growing traffic volumes. However, existing monitoring systems either cannot satisfy flexible intents on demand or produce significant overheads. In this paper, we presentNewton, an intent-driven traffic monitor that is able to specify operators’ intents with traffic monitoring queries and conduct dynamic and scalable network-wide queries deployment.Newtonenables operators to customize and modify queries dynamically without interrupting the network workflow. Besides,Newtonproposes systematic optimizations at device level and network-wide level to reduce resource consumption while deploying queries.Newtoncan combine the resources across switches to deploy complex queries with high resilience to dynamic network status. Evaluations prove thatNewtonis of high flexibility, scalability, and resource efficiency, which demonstratesNewtonis promising to be deployed in large-scale programmable networks.
Zhaowei Xi, Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001
IEEE/ACM Trans. Netw.7
2022 CoFilter: High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal Servers
abstract
As one of the most critical cloud services, Bare-Metal Servers (BMS) introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for BMS. However, the off-the-shelf stateful packet filters either are costly for cloud DCNs or introduce significant performance bottlenecks. In this article, we presentCoFilter, which leverages low-cost programmable switches to accelerate the stateful packet filter for BMS.CoFilteruses (1)stateful process partitionto enable complex stateful packet filtering logic on programmability-limited switching ASICs, (2)state compressionto track tens of millions of connections with constrained hardware memory, and (3)per-tenant packet rate limit and tenant-aware flow migrationto achieve efficient performance isolation among different tenants. Overall,CoFilterimplements a high-performance stateful packet filter via the co-design of programmable switching ASIC and CPU. We evaluateCoFilterunder various data center traffic traces with real-world flow distributions. The evaluation results show thatCoFilterremarkably outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay within 1us, and freeing a significant quantity of CPU cores, with rather small memory usage, i.e., accommodating over$10^7$connections with only 16MB SRAM.
Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Chen Sun 0005, Yangyang Wang 0001, Mingwei Xu 0001
IEEE Trans. Parallel Distributed Syst.6
2021 DOVE: Diagnosis-driven SLO Violation Detection
abstract
Service-level objectives (SLOs), as network performance requirements for delay and packet loss typically, should be guaranteed for increasing high-performance applications, e.g., telesurgery and cloud gaming. However, SLO violations are common and destructive in today’s network operation. Detection and diagnosis, meaning monitoring performance to discover anomalies and analyzing causality of SLO violations respectively, are crucial for fast recovery. Unfortunately, existing diagnosis approaches require exhaustive causal information to function. Meanwhile, existing detection tools incur large overhead or are only able to provide limited information for diagnosis. This paper presents DOVE, a diagnosis-driven SLO detection system with high accuracy and low overhead. The key idea is to identify and report the information needed by diagnosis along with SLO violation alerts from the data plane selectively and efficiently. Network segmentation is introduced to balance scalability and accuracy. Novel algorithms to measure packet loss and percentile delay are implemented completely on the data plane without the involvement of the control plane for fine-grained SLO detection. We implement and deploy DOVE on Tofino and P4 software switch (BMv2) and show the effectiveness of DOVE with a use case. The reported SLO violation alerts and diagnosis-needing information are compared with ground truth and show high accuracy (>97%). Our evaluation also shows that DOVE introduces up to two orders of magnitude less traffic overhead than NetSight. In addition, memory utilization and required processing ability are low to be deployable in real network topologies.
Yiran Lei, Yu Zhou 0008, Yunsenxiao Lin, Mingwei Xu 0001, Yangyang Wang 0001
ICNP5
2021 HyperTester: High-Performance Network Testing Driven by Programmable Switches
abstract
Modern network devices and systems are raising higher requirements on network testers that are regularly used to evaluate performance and assess correctness. These requirements include high scale, high accuracy, flexibility and low cost, which existing testers cannot fulfill at the same time. In this paper, we propose HyperTester, a network tester leveraging new-generation programmable switches and achieving all of the above goals simultaneously. Programmable switches are born with features like high throughput and linerate, deterministic processing pipelines and nanosecond-level hardware timestamps, the P4 programming model as well as comparable pricing with commodity servers, but they come with limited programmability and memory resources. HyperTester uses template-based packet generation to overcome the limitations of the switch ASIC in programmability and designs a stateless connection mechanism as well as counter-based state compression algorithms to overcome the memory resource constraints in the data plane. We have implemented HyperTester on Tofino, and the evaluations on the hardware testbed show that HyperTester supports high-scale packet generation (more than 1.6Tbps) and achieves highly accurate rate control and timestamping. We demonstrate that programmable switches can be potential and attractive targets for realizing network testers.
Yu Zhou 0008, Zhaowei Xi, Yangyang Wang 0001, Mingwei Xu 0001
IEEE/ACM Trans. Netw.4
2020 Newton: intent-driven network traffic monitoring
abstract
Monitoring network traffic based on operators' intents is essential to today's networks. As the bandwidth and size of networks increase steeply, monitoring systems shall fulfill the requirements of on-demand network monitoring for ever-growing traffic volumes. However, existing monitoring systems either cannot satisfy operators' intents on demand or introduce substantial monitoring overheads. In this paper, we present Newton, an intent-driven traffic monitor that enables specifying operators' intents with traffic monitoring queries and supports dynamic and scalable network-wide queries. Specifically, Newton 1) empowers operators to dynamically create, remove, and update on-data-plane queries without interrupting normal packet forwarding, 2) conducts systematic optimizations to achieve precise network traffic monitoring, and 3) executes network-wide queries with high resilience to dynamic network status. Evaluation results show that Newton improves the flexibility, scalability, and resource efficiency of traffic monitoring, demonstrating its great potential to be deployed in large-scale programmable networks.
Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001
CoNEXT6
2020 NetView: Towards On-Demand Network-Wide Telemetry in the Data Center
abstract
Network telemetry is to collect information (e.g., hop latency, throughput) from network devices. Network-wide telemetry is critical for operators to understand the quality of network performance and to diagnose on-going failures. The state-of-the-art telemetry approaches are far from ideal as they are unable to fully satisfy diverse requirements of operators, specifically for on-demand, full coverage, and scalable telemetry. In this paper, we provide a new framework of network telemetry for data center networks, called NetView. NetView can support various telemetry applications and frequencies on demand, monitoring each device via proactively sending dedicated probes. Technically, NetView leverages source routing to forward probes, achieving full coverage. Besides, a series of probe generation algorithms largely reduce probe number, providing high scalability. The evaluation shows that NetView reduces the bandwidth occupancy by more than two orders of magnitude compared with Pingmesh and INT-path, and conducts network-wide telemetry for large-scale data center network using only one vantage server, without bringing about resources bottleneck.
Yunsenxiao Lin, Yu Zhou 0008, Zhengzheng Liu, Yangyang Wang 0001, Mingwei Xu 0001, Jun Bi, Ying Liu 0024
ICC5
2020 NetView: Towards on-demand network-wide telemetry in the data center
Yunsenxiao Lin, Yu Zhou 0008, Zhengzheng Liu, Yangyang Wang 0001, Mingwei Xu 0001, Jun Bi, Ying Liu 0024
Comput. Networks5
2020 VMS: Load Balancing Based on the Virtual Switch Layer in Datacenter Networks
abstract
There have been many load balancing solutions for datacenter networks. Almost all of them require modifications to the network fabric or/and virtual machines. Recently, the virtual switch layer becomes an ideal location for datacenter operators to deal with the load balancing problem. In this paper, we propose Virtual Multi-channel Scatter (VMS), a packet-level load balancing design in the virtual switch layer. VMS scatters packets in one TCP flow to several different forwarding paths (channels). VMS has several noteworthy properties. First, VMS is low cost and transparent to tenants. It can be deployed when the datacenter operators do not attempt to change the network fabric or cannot control the transport protocol inside VMs. Second, by employing window-based channel selection, VMS is adaptive to network congestion and topology asymmetry. Third, VMS works well with Generic Segmentation Offload/Generic Receive Offload (GRO/GSO) mechanism in the Linux kernel, unlike other packet-level load balancing schemes. Finally, VMS can also be offloaded to SmartNIC to reduce CPU overhead further. Our evaluations show that VMS achieves comparable performance to the ideal packet-level scheme in normal cases and well handles topology asymmetries, while only modifies the virtual switch layer. In the symmetric topology, VMS achieves up to 47% and 22% better flow completion time (FCT) than Equal Cost MultiPath (ECMP) and the best-of-breed flowlet-level CONGA. When there is topology asymmetry, VMS outperforms the ideal packet-level scheme and CONGA by up to $3.0\times $ and $1.4\times $ respectively. Further, the overhead of VMS is tolerable.
Jun Bi, Zhaogeng Li, Yu Zhou 0008, Yangyang Wang 0001
IEEE J. Sel. Areas Commun.5
2020 HyperSight: Towards Scalable, High-Coverage, and Dynamic Network Monitoring Queries
abstract
Performing fine-grained and real-time network monitoring is the core logic of various data center operation applications, such as traffic engineering, network troubleshooting, and anomaly detecting. However, the state-of-the-art network monitoring solutions either fall short of completely detecting all network incidents (i.e., congestion), yielding limited monitoring coverage, or introduce large overheads, yielding limited scalability. In this paper, we present HyperSight, a network traffic monitor with both high coverage and low overheads. The key idea of HyperSight is to monitor networks at the behavior level via tracking packet behavior changes. HyperSight proposes three designs for behavior-level monitoring. First, to facilitate expressing various network monitoring tasks, HyperSight presents a declarative query language based on the streaming processing model. Second, HyperSight proposes Bloom Filter Queue (BFQ), a memory-efficient algorithm to empower in-network capability for monitoring packet behavior changes. BFQ can be implemented on commodity programmable switches. Third, to support dynamic deployment and execution of packet behavior change monitoring tasks without interrupting on-service switches, HyperSight proposes virtual BFQ to support dynamic query compilation. We build a prototype of HyperSight and deploy it on commodity programmable switches. Evaluation results show that HyperSight supports a wide range of network event queries and can monitor over 99% packet behavior changes while keeping remarkably low overheads.
Yu Zhou 0008, Jun Bi, Tong Yang 0003, Kai Gao 0001, Jiamin Cao, Yangyang Wang 0001, Cheng Zhang 0012
IEEE J. Sel. Areas Commun.7
2019 HyperTester: high-performance network testing driven by programmable switches
abstract
Modern network research and operations are inseparable from network testers to evaluate performance limits of proofs-of-concept, troubleshoot failures, etc. Existing network testers suffer from either constrained flexibility or a low performance-cost ratio. In this paper, we propose a new network tester, HyperTester. The core of HyperTester is to leverage new-generation programmable switches for generating and capturing test traffic with high performance, low cost, and remarkable flexibility. We design a series of efficient mechanisms, including template-based packet generation, false-positive-free counter-based queries, and stateless connections to realize various network testing tasks upon switches with limited programmability and resources. Meanwhile, to facilitate developing testing tasks upon HyperTester, we provide a high-level network testing API. We have implemented HyperTester on the Tofino switch and built dozens of network testing tasks. The evaluations on the hardware testbed show that HyperTester supports line-rate packet generation (400Gbps in the testbed) with highly-accurate rate control, while HyperTester can save $40150 per Tps and 9225W per Tbps when compared with the software network testers.
Yu Zhou 0008, Zhaowei Xi, Yangyang Wang 0001, Jinqiu Wang, Mingwei Xu 0001
CoNEXT4
2019 CoFilter: A High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal Servers
abstract
As one of the most critical cloud services, Bare-metal Servers introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for bare-metal servers. However, the off-the-shelf hardware-based and software-based stateful packet filters either are prohibitively costly for cloud DCNs or introduce significant performance bottlenecks. In this paper, we present CoFilter, which employs cheap programmable switches to accelerate the stateful packet filter for bare-metal servers. CoFilter consists of two key designs. First, to support complex stateful packet filtering logic in programmability-limited switching ASICs, CoFilter partitions the stateful packet filtering logic between programmable ASICs and switch CPU. Most packets are directly processed in switching ASICs to achieve high performance, while only a small number of packets go to switch CPU for connection tracking. Second, to track massive connections with constrained hardware memory, CoFilter employs hash to compress connection states and provides an efficient settlement for hash collisions. We build a prototype of CoFilter and evaluate it on the Tofino switch under various data center traffic traces with real-world flow distribution. The evaluation shows that CoFilter largely outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay at 1us, and freeing a significant quantity of CPU cores. Furthermore, CoFilter presents great scalability and accommodates over ten million connections with only 16MB SRAM.
Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Chen Sun 0005, Yangyang Wang 0001, Jun Bi
ICCCN5
2019 P4Tester: efficient runtime rule fault detection for programmable data planes
abstract
P4 and programmable data planes bring significant flexibility to network operation but are inevitably prone to various faults. Some faults, like P4 program bugs, can be verified statically, while some faults, like runtime rule faults, only happen to running network devices, and they are hardly possible to handle before deployment. Existing network testing systems can troubleshoot runtime rule faults via injecting probes, but are insufficient for programmable data planes due to large overheads or limited fault coverage. In this paper, we propose P4Tester, a new network testing system for troubleshooting runtime rule faults on programmable data planes. First, P4Tester proposes a new intermediate representation based on Binary Decision Diagram, which enables efficient probe generation for various P4-defined data plane functions. Second, P4Tester offers a new probe model that uses source routing to forward probes. This probe model largely reduces rule fault detection overheads, i.e. requiring only one server to generate probes for large networks and minimizing the number of probes. Moreover, this probe model can test all table rules in a network, achieving full fault coverage. Evaluation based on real-world data sets indicates that P4Tester can efficiently check all rules in programmable data planes, generate 59% fewer probes than ATPG and Pronto, be faster than ATPG by two orders of magnitude, and troubleshoot multiple rule faults within one second on BMv2 and Tofino.
Yu Zhou 0008, Jun Bi, Yunsenxiao Lin, Yangyang Wang 0001, Zhaowei Xi, Jiamin Cao, Chen Sun 0005
IWQoS4
2019 P4DB: On-the-Fly Debugging for Programmable Data Planes
abstract
While extending network programmability to a more considerable extent, P4 raises the difficulty of detecting and locating bugs, e.g., P4 program bugs and missed table rules, in runtime. These runtime bugs, without prompt disposal, can ruin the functionality and performance of networks. Unfortunately, the absence of efficient debugging tools makes runtime bug troubleshooting intricate for operators. This paper is devoted to on-the-fly debugging of runtime bugs for programmable data planes. We propose P4DB, a general debugging platform that empowers operators to debug P4 programs in three levels of visibility with rich primitives. By P4DB, operators can use the watch primitive to quickly narrow the debugging scope from the network level or the device level to the table level, then use the break and next primitives to decompose match-action tables and finely locate bugs. We implement a prototype of P4DB and evaluate the prototype on two widely-used P4 targets. On the software target, P4DB merely introduces a small throughput penalty (1.3% to 13.8%) and a little delay increase (0.6% to 11.9%). Notably, P4DB almost introduces no performance overhead on Tofino, the hardware P4 target.
Yu Zhou 0008, Jun Bi, Cheng Zhang 0012, Bingyang Liu, Zhaogeng Li, Yangyang Wang 0001, Mingli Yu
IEEE/ACM Trans. Netw.6
2018 NetVision: Towards Network Telemetry as a Service
abstract
In-band Network Telemetry (INT) can provide fine-grained and accurate device-level telemetry metrics. Nonetheless, INT can track only a small ratio of devices and links and embedding telemetry data into normal packets brings high overhead and high operation complexity. Hence, we present NetVision, a powerful proactive network telemetry platform with high coverage and high scalability.
Zhengzheng Liu, Jun Bi, Yu Zhou 0008, Yangyang Wang 0001, Yunsenxiao Lin
ICNP4
2018 KeySight: Troubleshooting Programmable Switches via Scalable High-Coverage Behavior Tracking
abstract
The rise of programmable switches and P4 brings much flexibility to networks, but this flexibility comes with increased risks of bugs. Diagnosing these bugs is essential for network operation but is non-trivial. A potential approach is to track packet behaviors through postcards, but existing tools either generate substantial postcards (limited scalability) or only track a small proportion of packet behaviors (low coverage). In this paper, we present KeySight, a platform that troubleshoots programmable switches with high scalability and high coverage. The key idea is based on the Packet Equivalence Class (PEC) abstraction that aggregates packets with identical behaviors and generates one postcard per behavior. The PEC abstraction minimizes the number of postcards while tracking all packet behaviors. We design novel algorithms to analyze PECs of P4 programs and to implement the PEC abstraction on programmable switches. We deploy KeySight on Tofino and SmartNIC, and evaluate it against 80 P4 programs and real packet traces of over 5TB. Results show that in the premise of overseeing over 99.9% packet behaviors, KeySight reduces the number of postcards by one to two orders of magnitude when comparing with NetSight.
Yu Zhou 0008, Jun Bi, Tong Yang 0003, Kai Gao 0001, Cheng Zhang 0012, Jiamin Cao, Yangyang Wang 0001
ICNP7
2017 P4DB: On-the-fly debugging of the programmable data plane
abstract
While extending network programmability to a larger degree, P4 also raises the risks of incurring runtime bugs after the deployment of P4 programs. These runtime bugs, if not handled promptly and properly, can ruin the functionality and performance of networks. Unfortunately, the absence of runtime debuggers makes troubleshooting of P4 program bugs challenging and intricate for operators. This paper is devoted to the on-the-fly debugging of runtime bugs in P4-enabled networks. We propose P4DB, a general debugging platform that empowers operators to debug P4 programs in three levels of visibility by provisioning operator-friendly primitives. By P4DB, operators can use the watch primitive to quickly narrow the debugging scope from network level or device level to table level, then use the break and next primitives to decompose the match-action table into three steps and troubleshoot the runtime bugs step by step. We implemented a prototype of P4DB and evaluated the performance in terms of the data plane, control plane and control channel. On P4-specific programmable data plane, P4DB merely introduces a small throughput penalty (1.3%~13.8%) and imposes a little-increased delay (0.6%~11.9%).
Cheng Zhang 0012, Jun Bi, Yu Zhou 0008, Bingyang Liu, Zhaogeng Li, Abdul Basit Dogar, Yangyang Wang 0001
ICNP8
2017 A tool for tracing network data plane via SDN/OpenFlow
Yangyang Wang 0001, Jun Bi, Keyao Zhang
Sci. China Inf. Sci.1
2017 A SDN-Based Framework for Fine-Grained Inter-domain Routing Diversity
Yangyang Wang 0001, Jun Bi, Keyao Zhang
Mob. Networks Appl.1
2015 MLV: A Multi-dimension Routing Information Exchange Mechanism for Inter-domain SDN
abstract
Software Defined Networking (SDN) separates the tightly coupled network control and data forwarding functions. During the past years, it has been applied in all kinds of intra-domain networks, such as enterprise networks, data centers and content provider networks. However, it is a big challenge to extend SDN to inter-domain networks. In this paper, we extend the advantage of SDN to an inter-domain network federation to improve the Internet routing flexibility. To achieve this goal, we propose a Multi-dimension Link Vector network view exchange mechanism (MLV) to exchange the fine-grained inter-domain routing information and enable programmable inter-domain routing. MLV can support flexible inter-domain routing control by exchanging multiple fields of the IP header. Based on MLV, innovations in inter-domain routing can be deployed as applications over SDN controllers. In order to validate the MLV design, we analyzed its performance with BGP-derived Internet AS topology. Finally, we implemented a prototype of MLV and tested it on an internationally collaborative inter-domain SDN testbed.
Ze Chen 0006, Jun Bi, Yonghong Fu, Yangyang Wang 0001, Anmin Xu
ICNP4
2015 WEBridge: west-east bridge for distributed heterogeneous SDN NOSes peering
abstract
Abstract Large networks are often partitioned by the network operators into several smaller networks when deploying software‐defined networks (SDNs). Additionally, a dedicated network operating system (NOS) is deployed for each of these SDNs. Each NOS can learn the local network view that enables control of how data packets are forwarded within its network. Controlling the flow of data packets in an entire network requires each NOS to have a global network view to determine the next NOS hop. Hence, NOSes are required to share or exchange reachability and topological information. How such information is efficiently exchanged has not been well addressed so far, especially in the case of multi‐vendor NOSes. This paper proposes a west–east bridge mechanism for distributed heterogeneous NOSes to cooperate in enterprise/data center/intra‐autonomous system networks. We propose to simplify physical networks into virtual networks and only exchange the simplified virtual network information to construct the global network view. To achieve a resilient peer‐to‐peer control plane of distributed heterogeneous NOSes, we propose a “maximum connection degree”‐based connection algorithm. Considering the security issue, we adopt controller identity authentication. We implement the west–east bridge and analyze the performance obtained: about 100% of enterprises and data centers, and about 99.5% of autonomous systems can adopt to this solution. The deployment in three SDNs (CERNET, Internet2, and CSTNET) proves the feasibility. Copyright © 2014 John Wiley & Sons, Ltd.
Pingping Lin, Jun Bi, Yangyang Wang 0001
Secur. Commun. Networks3
2013 Refining IP-to-AS Mappings for AS-Level Traceroute
abstract
It is of great significance for network operators and researchers to obtain accurate AS-level traceroute paths, for which mapping IP addresses to correct AS numbers is critical. Thus, there have been a lot of efforts to improve the original IP-to-AS mapping table, which was extracted from BGP routing tables. One of these efforts is called pair matching, which refines the original mapping table by maximizing the number of matched pairs of traceroute and BGP AS paths. However, the existing pair-matching-based methods refine the original IP-to-AS mapping table only with the prefix granularity, i.e., IP addresses in the same /24 prefix are mapped to the same AS or the same set of ASes, which does not fit reality. In this paper, we attempt to refine the IP-to-AS mapping table with the IP address granularity, i.e., allowing IP addresses in the same prefix to be mapped to different ASes. The results show that our fine-grained method can produce a more accurate IP-to-AS mapping table. In addition, this paper also provides a better understanding for the pair-matching-based methods.
Baobao Zhang, Jun Bi, Yangyang Wang 0001, Yu Zhang 0036
ICCCN3
2012 Towards an Aggregation-Aware Internet Routing
abstract
Internet is composed of a large amount of autonomous systems (ASes). Border Gateway Protocol (BGP) is the de facto standard used to connect these ASes and exchange reachability information between them. The global BGP routing table size in default free zone (DFZ) grows fast due to many factors including IP address allocation, multihoming, and traffic engineering, etc. Increasing prefix fragments consume more memory space and computational capacity in network forwarding devices. It has been known that the Internet has a potential routing scalability issue along with large address space (e.g., IPv6) deployment in the future. Route aggregation is a practical approach to reduce route entries. In this paper, we propose an innovation based on BGP, named Aggregation-aware Inter-Domain Routing (AIDR). It takes advantage of the redundant paths to the same destination in the Internet, and takes route aggregation into account in route selection to get more aggregation for forwarding table (FIB). We give a detailed analysis and evaluation on the effect of AIDR using the public BGP traces from RouteViews and RIPE. It shows that AIDR can produce aggregated FIBs of size roughly 20%~36% of the original routing table size with allowing 2.0 AS path stretch, and 25%~40% without AS path tretch.
Yangyang Wang 0001, Jun Bi
ICCCN1
2011 IPv6 evolution, stability and deployment
abstract
Our subject focuses on IPv6 network, which develops for more than 10 years. How IPv6 evolve in those years? Is IPv6 network mature enough to undertake the load produced by users? Can we find some principles to guide IPv6 deployment, which make the whole network more robust and efficiency? This paper tries to answer these questions with in-depth statistics. Good news is that network is growing at a speed of O(d2) (d is time) after 2006, moreover, network itself and its routing system become more and more stable. And we explore special properties of this preliminary network, We find that distribution of AS degree follows "Power-Law Distribution", but AS-level topology cannot be described as "Small-World Model" properly. We also propose a method to define the importance of AS and give a simple principle of IPv6 deployment. We even build "6Stats Project"[1] to provide data which help deploy IPv6.
Xiaoke Jiang, Jun Bi, Yangyang Wang 0001, Zhijie He, Wei Zhang 0040, Hongcheng Tian
ICNP3
2011 AIDR: Aggregation of BGP routing table with AS path stretch
abstract
As Internet growth, more and more prefix fragments are announced into the global routing system due to operational reasons of inconsecutive address allocation, multihoming, and traffic engineering. The BGP routing table size in Default Free Zone (DFZ) fast growth will consume more memory space and computational capacity. It has been known that Internet will face with routing scalability issue, especially in the large address space (e.g., IPv6) deployment. In this paper, we propose an innovation to BGP, named Aggregation-aware Inter-Domain Routing (AIDR). It will take the prefix aggregation into account to make tradeoff in the best route selection. We evaluate the effect of AIDR on global routing system using the BGP traces from RouteViews and RIPE. It shows that, averagely, AIDR-based aggregation can reduce to roughly 15%~35% of original routing table size under the 2.0 AS path stretch constraint, and to 25%~40% with no AS path stretch.
Yangyang Wang 0001, Jun Bi
ICNP1
2011 A Framework to Quantify the Pitfalls of Using Traceroute in AS-Level Topology Measurement
abstract
Although traceroute has the potential to discover AS links that are invisible to existing BGP monitors, it is well known that the common approach for mapping router IP addresses to AS numbers based on BGP routing tables is highly error-prone. We develop a systematic framework to quantify the potential errors of traceroute measurement in AS-level topology inference. In comparing traceroute-derived AS paths with BGP AS paths, we take a novel approach to identifying mismatched path segments and then inferring the causes of these mismatches through a set of tests. Our results show that about 60% of mismatches are due to routers using IP addresses belonging to peering neighbors. This result helps settle a debate in previous works regarding the major cause of errors in traceroute measurement. With the approximate ground truth of the ASes with BGP monitors inside, we identify the inaccuracy of publicly available traceroute-derived topology datasets and find that between 8% and 42% of AS adjacencies on the monitored ASes are false. With a new method to characterize AS links, we show that the derived (false) links between Tier-1/large ISPs and their customers' customers appear more frequently than real links do.
Yu Zhang 0036, Ricardo V. Oliveira, Yangyang Wang 0001, Shen Su, Baobao Zhang, Jun Bi, Hongli Zhang 0001, Lixia Zhang 0001
IEEE J. Sel. Areas Commun.3
2010 Empirical Analysis of Core-Edge Separation by Decomposing Internet Topology Graph
abstract
Border Gateway Protocol (BGP) is the de facto standard protocol for the inter-domain routing. Due to multi-homing and traffic engineering, the BGP routing table size of default free zone (DFZ) is growing rapidly. Inter-domain routing is facing the scaling challenge. Many solutions have been proposed. Among them, the core-edge separation scheme gets more attentions than others due to its practical advantages. It separates specific prefixes of edge networks from entering into transit core, and reduces the DFZ BGP routing table size. However, there has less evaluation on how much scalability can be improved from core-edge separation. In this paper, we take the further step to quantifying the impact of the core-edge separation on Internet inter-domain routing. We find that separation at stub-transit can reduce 43% routing table size and prevent more than half of BGP updates. We decompose the topology graph by k-core and customer-provider based decomposition methods, and analyze the impact of deploying separation at different level of topological hierarchy. We believe that complicated separation deployment strategies (not the simple stub-transit split) are feasible to approaching an optimal effect.
Yangyang Wang 0001, Jun Bi
GLOBECOM1