Jilong Wang 0001

dblp:18/6008-1 · DBLP profile ↗
← Back
100ranked-venue papers
2as first author
77since 2021 · last 2026
0000-0002-4493-5145ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 65 · 1 first-author · 51 since 2021Security and privacy · 10 · 8 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Systems, architecture and hardware · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 ECOTE: Priority-Aware Optical Restoration for WAN Traffic Engineering
abstract
Fiber cuts are among the most common and disruptive failures in cloud networks. They can prevent cloud providers from maintaining committed service availability, causing Service Level Agreement (SLA) violations that directly translate into monetary penalties. Existing traffic engineering (TE) approaches enhance failure resilience, and recent systems further incorporate optical restoration to recover lost bandwidth after failures. However, they still treat services largely uniformly and optimize primarily for network throughput rather than the economic impact of heterogeneous SLA penalties. In this paper, we present ECOTE, the first priority-aware TE system with optical restoration that explicitly minimizes revenue loss. Specifically, ECOTE introduces a new optical restoration formulation with a dedicated capacity restoration solver to compute an optimal restoration plan that maximizes restorable capacity under physical constraints. ECOTE also designs a priority-aware TE algorithm that allocates the restored bandwidth capacity according to SLA penalties, thereby reducing monetary cost. We evaluate ECOTE using a production-level WAN testbed and through large-scale simulations. The testbed evaluation demonstrates ECOTE achieves zero loss for high-priority services and more than 10× revenue loss reduction compared to state-of-the-art. Our large-scale simulation results show that ECOTE can support at least 2.5× and 2.0× more demand for different high priority services compared to the state-of-the-art solutions. Meanwhile, ECOTE reduces the revenue loss by at least an order of magnitude less than existing solutions.
Kunling He, Ran Shu 0001, Jilong Wang 0001, Congcong Miao
EuroSys5
2026 AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory Constraints
Yumeng Cui, Hui Wang 0011, Najila Liu, Ling Deng, Chuxuan Zeng, Jilong Wang 0001
INFOCOM7
2026 Comprehensive Network Configuration Verification via Effective Environment Reduction
Han Zhang 0009, Renrui Tian, Xia Yin 0001, Xingang Shi, Gang Ren 0003, Jilong Wang 0001, Jiangyuan Yao
INFOCOM8
2026 Crack in the Armor: Underlying Infrastructure Threats to RPKI Publication Point Reachability
Yunhao Liu 0001, Hui Wang 0011, Yuedong Xu 0001, Zongpeng Li, Jilong Wang 0001
NDSS6
2026 DDoS Detection at the Scale of One Hundred Tbps
Yunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang, Heng Yu 0005, Jiahao Cao 0001, Yong Jiang 0001, Jilong Wang 0001, Mingwei Xu 0001, Congcong Miao
NSDI10
2026 Dorado: Scaling SmartNIC Session Tables on Commodity DDRs
Heng Yu 0005, Jiajun Liang, Baozeng Zhang, Guozhi Lin, Xinyi Zhang 0004, Jian Zhao 0006, Ziyue Zhai, Chao Pei, Jilong Wang 0001, Gaogang Xie, Ang Chen 0001, Congcong Miao
SIGCOMM12
2026 Towards High-Performance Intrusion Detection with Robustness Guarantees on Programmable Switches at ISP Scale
abstract
In order to provide security connections to the enterprise campus sites, internet service providers are offering comprehensive intrusion detection services at the network layer. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for high-speed network protection, especially for encrypted traffic analysis. In this paper, we design and implement SiteGuard, an inline network intrusion detection system with programmable switches specifically developed to protect enterprise campus sites connecting to ISP. SiteGuard proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. SiteGuard also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify malicious traffic. In addition, SiteGuard introduces an online update mechanism that aims to dynamically adjust the detection model in response to environmental changes. SiteGuard has been in production for more than three years. Our production and testbed evaluations demonstrate SiteGuard can detect malicious traffic with approximately 90% accuracy in minutes.
Han Zhang 0009, Linqiang Qian, Guyue Liu, Kaiyang Zhao 0004, Yantu Tong, Zeji Xiao, Dongbiao He, Ke Ruan, Jilong Wang 0001, Xia Yin 0001
SIGCOMM12
2026 One Char to Rule Them All: Systematically Exploring and Exploiting DNS Silent Vulnerabilities in Domain Name Resolution
Fasheng Miao, Xiang Li 0108, Changqing An, Jilong Wang 0001
SP5
2026 A comprehensive survey on encrypted network traffic classification
Shangbin Han, Han Zhang 0009, Mengmeng Lu, Sifang Guo, Boyuan Tian, Jilong Wang 0001
Comput. Networks6
2026 Clover: Workload Verification for Real-Time Detection of Contention-Induced Slowdowns in Serverless Platforms
abstract
Serverless computing, or Function-as-a-Service, continues to gain popularity due to its pay-as-you-go billing model, flexibility, and cost efficiency. However, these same features introduce significant security risks, such as the Denial-of-Wallet (DoW) attack. In this paper, we conduct real-world DoW attacks on commercial serverless platforms to evaluate their severity. To detect such attacks, we design, implement, and evaluate Clover, an accurate and user-friendly DoW detection system with negligible performance overhead. Clover addresses information ambiguity in serverless environments by deploying a request-oriented metric collection agent. At its core, Clover proposes a workload verification approach to bridge performance metrics and execution duration. Specifically, Clover uses a multivariate linear model to learn the benign relationship between metrics and execution duration, effectively characterizing normal workload behavior. It then continuously monitors runtime workloads by calculating their Mahalanobis distance from this learned benign model. Deviations identified through this distance indicate potential DoW attacks. Implemented as a practical system, Clover introduces performance overhead of less than 3.2%, maintains an average model execution time of only 0.84 microseconds, and achieves an accuracy of 92.7% under the most challenging scenario.
Junxian Shen, Han Zhang 0009, Weiwei Lin 0001, Yantao Geng, Jilong Wang 0001, Mingwei Xu 0001
IEEE Trans. Netw.6
2026 Link Prediction-Based Measurement Strategy for Efficient Topology Completeness Improvements
abstract
The Autonomous System (AS) level topology observed from current measurement infrastructures is far from complete. Although Looking Glass (LG) vantage points (VPs) that support BGP route queries can provide valuable topology information, the query rate limitations of LG VPs imply that blindly using the VPs to conduct more measurements to improve the topology completeness is inefficient, if not infeasible. In this paper, we try to improve the efficiency by designing a link prediction based measurement strategy, whose basic idea is to first predict where unseen AS links are likely to be located and then use the prediction results to guide the measurements toward a more complete AS-level topology. We formulate the prediction of unseen AS links as a matrix completion problem and develop a side-information assisted learning-based matrix completion method. The method exploits a neural network and utilizes carefully chosen AS attributes based on our understanding on Internet peering practices, thereby learning more expressive latent vectors and achieving outstanding prediction performance in our scenario. We then develop a measurement strategy which takes the link prediction results as guidance to achieve efficient topology completeness improvements. The strategy leverages several heuristics to estimate the utilities of different measurements and takes a greedy algorithm to select the most valuable measurements. Experiments show that our link prediction method can achieve a high AUC (Area Under the Receiver Operating Characteristic Curve) of 0.834 and the link-guided measurement strategy can discover 1.82 times more unseen links than those discovered from non-guided measurement strategies with an equal number of measurements.
Shuying Zhuang, Hui Wang 0011, Jilong Wang 0001, Changqing An, Yuedong Xu 0001, Tianhao Wu 0010
IEEE Trans. Netw.3
2025 Poster: ERIS: Evaluating ROV via ICMPv6 Rate Limiting Side Channels
abstract
The Resource Public Key Infrastructure (RPKI) plays a crucial role in securing BGP against prefix hijacking by enabling Route Origin Validation (ROV). However, the limited adoption of ROV in the real world undermines the effectiveness of RPKI. Hence, measuring ROV deployment in practice is essential for assessing the impact of RPKI. Existing measurement efforts either suffer from limited coverage and accuracy due to reliance on control-plane data, or require controlled IP prefixes or large-scale deployment of vantage points. Furthermore, most studies focused on IPv4, leaving ROV status in IPv6 largely underexplored.
Renrui Tian, Han Zhang 0009, Xia Yin 0001, Xingang Shi, Jilong Wang 0001
CCS8
2025 MADE: a Masked Autoencoder Based Desensitization Framework for Encrypted Traffic
abstract
In recent years, encryption algorithms have become widely adopted as a trusted means of securing data transmission, effectively safeguarding against unauthorized access. However, providing raw encrypted traffic directly to researchers can still pose privacy risks, making it essential to implement desensitization measures before sharing such data. The main challenge is to desensitize the encrypted traffic in a privacyprotection way while preserving as much of the original traffic characteristics as possible to ensure the downstream tasks run properly. In this paper, we propose MADE, an encrypted traffic desensitization framework based on Masked AutoEncoder (MAE). This framework transforms the traffic desensitization task into a classifier-guided image reconstruction problem, using supervised signals to achieve state-aware traffic reconstruction, which ensures the utility as well as the safety of the desensitized traffic for downstream tasks. Moreover, MADE is highly scalable, permitting users to customize which data attributes to preserve in the traffic being processed, simply by modifying the classifier. Experimental results reveal that our proposed desensitization framework not only outperforms baseline methods but also exhibits high applicability across various scenarios.
Shilin Xie, Nan Jiang Jiang, Suxiang Wu, Jilong Wang 0001
ICC5
2025 Flexnetic: Cost-Effective and Smooth Evolution of Optical Backbone
abstract
The increasing traffic on WANs due to growing number of applications imposes significant strain on the infrastructure of cloud service providers. The primary expense in augmenting network capacity entails high costs associated with procuring expensive transponders that facilitate inter-regional optical signal transmission. The advent of spacing variable transponders facilitates cost-effective network upgrades and increased transmission efficiency, however, implementing such upgrades in one step may incur high costs and service disruptions. Achieving a cost-effective and smooth network upgrade presents considerable challenges in identifying critical IP links and minimizing costs within existing architecture. We introduce Flexnetic, a planning tool which utilizes a hybrid approach of both modern and legacy transponders, along with establishment of optical bypass, to accommodate the escalating traffic demands while minimizing the costs during network upgrades. Flexnetic incorporates two novel algorithms: a NLP model to maximize IP link capacity utilization, and an MIP model for efficient IP layer implementation at optical layer, emphasizing reuse of existing transponder and minimizes new device requirement. Our simulation of upgrade plans on common WAN topologies revealed superior performance, with up to 91.9% cost savings and 2.33× capacity increase over existing state-of-the-art solutions, highlighting Flexnetic’s potential for cost-efficient and capacity-optimized network upgrades.
Congcong Miao, Kunling He, Jilong Wang 0001
ICCCN5
2025 InternetSim: A Fast and Memory-Efficient Internet-Scale Inter-Domain Routing Simulator
abstract
Existing inter-domain routing simulators often suffer from low simulation accuracy, slow processing speeds, and high memory consumption, hindering their ability to perform large-scale Internet simulations. In this paper, we present InternetSim, a multi-threaded simulator for Internet-scale inter-domain routing. InternetSim enables incremental computation to deal with policy changes by some ASes, thus significantly accelerating multi-iteration context-continuous routing simulations. The memory efficiency is also improved by designing compact data structures for routing tables and out-of-memory errors are prevented by offloading the data to disk when necessary. Our simulator can complete an iteration of Internet-scale simulation within 13 hours (21× speedup compared to C-BGP) and can complete an ''incremental computation'' iteration within 15 seconds. The memory requirement is only 45% of C-BGP (without offloading) and the simulator can work well for Internet-scale simulations on a server with only 64GB if offloading is always enabled.
Jiahong Lai, Hui Wang 0011, Yunhao Liu 0001, Jilong Wang 0001
IMC4
2025 Poster: RMap: Uncovering Risky DNS Resolution Chains and Misconfigurations
abstract
In recent years, large-scale network outages caused by DNS misconfigurations have become increasingly common. The intricate inter-domain dependencies, along with emerging mechanisms (Such as DNSSEC, EDNS, and 0x20), have made DNS resolution increasingly complex and fault localization more challenging. We present RMap, a tool that rapidly probes all potential resolution chains of a domain, reveals its resolution dependency topology, and detects security risks. We experimentally demonstrate the effectiveness of RMap and its broad applicability. Our findings reveal that domain configurations in real-world environments remain concerning, with potential issues observed even in several well-known top-level domains. RMap is avaliable in https://github.com/ahlien/rmap.
Fasheng Miao, Shuying Zhuang, Xiang Li 0108, Changqing An, Deliang Chang, Baojun Liu 0002, Jia Zhang 0004, Jilong Wang 0001
IMC8
2025 Probabilistic Analysis of Overload-Free Property for Critical Traffic
abstract
Network structures are sophisticated and hence vulnerable to errors. Link failures and traffic load fluctuations lead to complexity in network states. Different failure scenarios can result in varying network traffic distribution patterns. Meanwhile, the load on links within the same failure scenario dynamically changes with fluctuations in traffic. Network administrators are particularly concerned about whether links along the paths traversed by critical traffic are overload-free guaranteed when link failures occur. Yet, no attention was ever paid to overload-free property analysis for critical traffic. We propose Offaela, an efficient and accurate probabilistic analysis framework that verifies an overload-free property for critical traffic. We prudently formulate the problem and prove its computational hardness, then storm this fortification by proffering a failure scenario merging algorithm and adopting a randomized approximation method. Evaluations on real networks show that Offaela outperforms the state-of-the-art solution by 4.83 x and can provide availability analysis assistance such as identifying vulnerable failure scenarios.
Zhiyun Tang, Ke Ruan, Yingjun Ye, Jilong Wang 0001, Xia Yin 0001, Xingang Shi, Han Zhang 0009
IWQoS6
2025 Argus Lens: Innovating Internet Measurement Infrastructure via Relay Services
Changqing An, Jilong Wang 0001
Networking3
2025 Unlocking ECMP Programmability for Precise Traffic Control
Yunming Xiao, Weizhen Dang, Xiang Li 0223, Zekun He, Jilong Wang 0001, Aleksandar Kuzmanovic, Ang Chen 0001, Congcong Miao
NSDI8
2025 Fornax: A Hardware-Centric Session Management in Large Public Cloud Network
abstract
SmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%.
Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao
SIGCOMM15
2025 Achieving High-Speed and Robust Encrypted Traffic Anomaly Detection with Programmable Switches
abstract
Attacks against data centers are becoming more common as a result of the fast expansion of applications. In order to keep pace with the growing amount of data centers connected to their networks, internet service providers must offer comprehensive security services. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for the high-speed encrypted network traffic. In this paper, we design and implement Mazu, an inline network intrusion detection system with programmable switches specifically developed to protect data centers connecting to the internet service provider. Mazu proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. Mazu also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify the malicious traffic. In addition, Mazu introduces an online update mechanism aimed at dynamically adjusting the detection model in response to environmental changes. Mazu has been in production for two years, during which time it has identified over 10 critical attack events and protect more than 10 million servers for two ISPs. Our production and testbed evaluations demonstrate that Mazu can detect malicious traffic entering the data center sites with approximately 90% accuracy within minutes.
Han Zhang 0009, Guyue Liu, Xingang Shi, Dongbiao He, Jilong Wang 0001, Ke Ruan, Xia Yin 0001
SIGCOMM6
2025 Low-Overhead Distributed Application Observation with DeepTrace: Achieving Accurate Tracing in Production Systems
abstract
As microservices grow in scale and complexity, their operation and debugging become increasingly challenging. Even a single user request can involve interactions across hundreds of components. In such intricate systems, distributed tracing, which tracks the end-to-end execution flow of requests, has become a critical monitoring tool. Among these, non-intrusive tracing frameworks that do not require code modification are particularly valued for their convenience. However, existing non-intrusive solutions either have limited applicability or lack sufficient accuracy under high concurrency. To address these challenges, we propose DeepTrace, a transaction-based, non-intrusive distributed tracing framework designed for microservices. DeepTrace leverages API endpoints and transaction fields embedded within request content to categorize requests into distinct transactions, thereby reducing the likelihood of incorrectly merging traces from different transactions. Compared to state-of-the-art frameworks, DeepTrace maintains an accuracy rate of over 95% even under high concurrency. It has also been adopted by dozens of companies in their production systems for tasks such as failure diagnosis and resource optimization.
Yantao Geng, Han Zhang 0009, Jilong Wang 0001, Xia Yin 0001
SIGCOMM5
2025 PreTE: Traffic Engineering with Predictive Failures
abstract
Fiber links in wide-area networks (WANs) are exposed to complicated environments and hence are vulnerable to failures like fiber cuts. The conventional approach of using static probabilistic failures falls short in fiber-cut scenarios because these fiber cuts are rare but disruptive, making it difficult for network operators to balance network utilization and availability in WAN traffic engineering. Our large-scale measurements of per-second optical-layer data reveal that the fiber's failure probability increases by several orders of magnitude when experiencing a rare and ephemeral degradation state. Therefore, we present a novel traffic engineering (TE) system called PreTE to factor in the dynamic fiber cut probabilities directly into TE systems. At the core of the PreTE system, fiber degradation facilitates failure predictions and traffic tunnels to be proactively updated, followed by traffic allocation optimizations among updated tunnels. We evaluate PreTE using a production-level WAN testbed and large-scale simulations. The testbed evaluation quantifies PreTE's runtime to demonstrate the feasibility to implement in large-scale WANs. Our large-scale simulation results show that PreTE can support up to 2× more demand at the same level of availability as compared to existing TE schemes.
Congcong Miao, Zhizhen Zhong, Arpit Gupta, Ying Zhang 0022, Zekun He, Xianneng Zou, Jilong Wang 0001
SIGCOMM9
2025 Ares: Comprehensive Path Hijacking Detection via Routing Tree
Yinxiang Tao, Chengwan Zhang, Changqing An, Shuying Zhuang, Jilong Wang 0001, Congcong Miao
USENIX Security Symposium5
2025 ACME++: A Secure Authorization Mechanism for ACME Clients in the Web PKI Ecosystem
abstract
The Automatic Certificate Management Environment (ACME) protocol automates the issuance and renewal of secure socket layer certificates, simplifying the management of large-scale certificate deployments. To reduce the load on Certificate Authority (CA) servers, ACME employs a caching mechanism that stores domain validation (DV) results for 30 days. However, this mechanism allows attackers to reuse previously authorized results, potentially bypassing the DV process. In this paper, we introduce the ACME Authz Cache Attack, whereby an attacker can obtain fraudulent certificates without domain control. We demonstrate that even the prominent CA, Let's Encrypt, is vulnerable to this attack. To mitigate this, we propose ACME++, an enhanced protocol that binds the client's IP address and a unique identifier to the ACME account, ensuring secure authorization for each new client and effectively preventing the ACME Authz Cache Attack. Our implementation of ACME++ shows that it introduces little overhead to the CA server.
Han Zhang 0009, Yunze Wei, Xingang Shi, Jilong Wang 0001, Xia Yin 0001
WWW6
2025 Cost-Efficient FEC Scheme for Time-Sensitive Multi-Hop Transmissions in Overlay Networks
abstract
In pursuit of low latency, real-time communication (RTC) service providers usually use multi-hop overlay links worldwide to bypass congested links, especially for medium- and long-distance transmissions. In such multi-hop long-distance transmission scenarios, utilizing retransmission to recover lost packets can result in increased end-to-end latency. Therefore, Forward Error Correction (FEC) is viewed as a promising way to solve the loss problem. However, for multi-hop overlay transmission, existing FEC schemes either introduce a non-negligible processing delay at each hop or reduce the processing delay at the cost of a high coefficient overhead. In this work, we propose a multi-hop FEC scheme, i.e., FEC-OEM, which considers both processing delay and coefficient overhead. FEC-OEM is designed based on two observations we obtained from measurements. First, coefficient overhead can only be reduced through an implicit transmission way. Therefore, we design a modulation-based recoding module that enables implicit coefficient transmission and hop-by-hop recoding at the same time. Second, using on-the-fly computation is a promising way to reduce processing delay. Accordingly, we design an elimination method to make the modulation-based recoding can be carried out on-the-fly. Real-world experiments demonstrate that FEC-OEM can reduce the processing delay by up to 88% without increasing the coefficient overhead compared to state-of-the-art schemes. We also use FEC-OEM to transmit packets for applications with different loss tolerances, and the results show that FEC-OEM can improve the QoE more effectively than state-of-the-art coding schemes.
Chao Xu 0015, Hui Wang 0011, Jilong Wang 0001, Jun Zhang 0004
IEEE Trans. Mob. Comput.5
2025 Predictive Configuration on DHCP in WLANs
abstract
DHCP is widely deployed in WLANs to automatically assign IP addresses to WiFi devices when users connect to the WLANs. However, frequent user mobility brings big challenges to the DHCP performance. Recently proposed IP configuration (e.g., IP lease time, size of IP address pool) decisions on DHCP are based on traditional models to study user mobility patterns which lead to poor DHCP performance since the online time of individuals varies due to their personal pReferences and the number of crowds differs spatially and temporally. In this paper, we propose PredHCP, a predictive configuration framework on DHCP to improve the DHCP performance. Specifically, PredHCP utilizes an attention-based recurrent neural network (ARNN) to learn sequential patterns of individual mobility and accurately predicts user online time to ensure the effective IP lease time configuration. Meanwhile, PredHCP introduces a spatio-temporal graph neural network (STGNN) to learn both spatial and temporal dependencies of crowd migration and accurately predict crowd size in each area to ensure effective IP pool configuration. We conduct comprehensive experiments on real network traces for a month to evaluate the performance of PredHCP. Experimental results show that PredHCP can accurately predict user mobility patterns by achieving lower prediction errors. By accurately modeling mobility patterns, PredHCP makes effective IP configuration to ensure high DHCP performance. Large-scale simulation results show that PredHCP can save up to 69% IP addresses and the IP efficiency is 41% which outperforms existing methods by 6%.
Pei Zhang 0003, Hanyan Yin, Botong Wu, Xiaohong Huang 0003, Yan Ma 0003, Jilong Wang 0001, Congcong Miao
IEEE Trans. Netw.8
2024 LogRAG: Semi-Supervised Log-based Anomaly Detection with Retrieval-Augmented Generation
abstract
Log-based anomaly detection is critical in monitoring the operation of microservice systems and in the realtime reporting of system failures. Utilizing deep learning-based log anomaly detection methods facilitates effective detection of anomalies within logs. However, existing methods are greatly dependent on log parsers, and parsing errors can considerably affect downstream anomaly detection tasks. Additionally, methods that predict the next log event in a sequence are susceptible to the instability of sequences and the emergence of unseen logs as systems evolve, resulting in a higher false positive rate. In this paper, we propose a semi-supervised log anomaly detection framework based on retrieval-augmented generation (RAG). This framework conducts phased detection using both Log Tokens and Log Templates to mitigate the impact of log parsing errors. It also utilizes a single-class classifier to model the normal behavior of the system, thereby circumventing the effects of unstable sequences. Finally, it employs large language model (LLM) empowered by RAG to reevaluate detected anomalous logs.
Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, Jilong Wang 0001
ICWS7
2024 Collecting Self-reported Semantics of BGP Communities and Investigating Their Consistency with Real-world Usage
abstract
People can extract various kinds of information about the Internet from BGP routes tagged with BGP community values with known semantics. In this paper, we conduct a study on the following three issues related to BGP community semantics. First, we design a method to automatically collect self-reported semantics from the Internet and assemble the collected semantics described in natural language into a structured dictionary. The comparison with prior dictionaries shows many community values are exclusively covered by ours and many of them had been used when prior dictionaries were constructed, which confirms the effectiveness of our method. Second, based on this large-size dictionary, we are able to re-evaluate two recent algorithms designed for categorizing community values with unknown semantics, which is a task that, while easier than inferring the detailed semantics, is also very valuable. Our evaluation uncovers some issues within the algorithms that can contribute to their performance improvement. Third, we investigate the fundamental issue in extracting information using community semantics: whether ISPs' behavior is consistent with the published semantics. Our preliminary best-effort investigation reveals the potential risks of using the semantics of some categories of community values.
Yunhao Liu 0001, Tianhao Wu 0010, Hui Wang 0011, Jilong Wang 0001, Shuying Zhuang
IMC4
2024 Rumors Stop with the Wise: Unveiling Inbound SAV Deployment through Spoofed ICMP Messages
abstract
In the era of increasing network-based threats, particularly IP spoofing, Source Address Validation (SAV) is paramount for network security. The effective deployment of Inbound Source Address Validation (ISAV) is crucial yet often inadequate, posing significant risks to Internet infrastructure. This study presents ICMP_Sonar, a measurement system that deploys "rumors" -carefully crafted spoofed ICMP packets-to probe the network's defenses, revealing the "wise" networks with their robust ISAV implementations. ICMP_Sonar introduces two novel approaches that exploit the characteristics of ICMP unreachable messages and ICMP fragment needed messages, and exhibits the advantages of high coverage, fine granularity, low error rates, and the ability to measure in both IPv4 and IPv6. We also evaluate the applicability and security risks of ICMP error messages. Through large-scale measurements, ICMP_Sonar successfully covers 86M IPv4 hosts (0.8M IPv6 hosts), 3.5M IPv4 /24 subnets (24K IPv6 /40 subnets), and 59K IPv4 ASes (8.3K IPv6 ASes), surpassing the state-of-the-art dual-stack method's coverage by 16.2 (51.6), 2.9 (2.34), and 1.7 (1.7) times, respectively. The broad coverage across multiple granularities enables us to capture a more comprehensive and fine-grained view of ISAV deployment. Measurements show that while the percentage of ASes with no ISAV deployment is lower than previously identified, the percentage of ASes with partial ISAV deployment is much higher, indicating significant gaps in overall security. The analysis also reveals that ISAV deployment practices vary across different networks and between IPv4 and IPv6.
Shuaicong Yu, Shuying Zhuang, Changqing An, Jilong Wang 0001
IMC5
2024 NetFEC: In-network FEC Encoding Acceleration for Latency-sensitive Multimedia Applications
abstract
In face of packet loss, latency-sensitive multimedia applications cannot afford re-transmission because loss detection and re-transmission could lead to extra latency or otherwise compromised media quality. Alternatively, forward error correction (FEC) ensures reliability by adding redundancy and it is able to achieve lower latency at the cost of bandwidth and computational overheads. We propose to re-locate FEC encoding to hardware that better suits the computational pattern of FEC encoding than CPUs. In this paper, we present NetFEC, an in-network acceleration system that offloads the entire FEC encoding process on the emergent programmable switching ASICs, eliminating all CPU involvement. We design the ghost packet mechanism so that NetFEC can be compatible with important media transport functionalities, including congestion control, pacing and statistics. We integrate NetFEC with WebRTC and conduct extensive experiments with real hardwares. Our evaluations demonstrate that NetFEC is able to eliminate server CPU burden and adds negligible overheads.
Yi Qiao, Han Zhang 0009, Jilong Wang 0001
INFOCOM3
2024 ROV-GD: Improving the Measurement of ROV Deployment Using Graph Difference
abstract
BGP has been threatened by prefix hijacking attacks due to the lack of authentication mechanisms. In recent years, many ASes have begun to participate in RPKI deployment to improve Internet security jointly. Some researchers have studied how to measure the actual deployment of ROV globally, but such works suffer from the shortcomings of small measurement coverage and insufficient accuracy. Therefore, we propose RO-VGD, a ROV measurement method based on graph difference. We collect routing path data from the control plane and data plane. Then, we rely on prefix reachability and propagation edges for ROV inference. The results show that about half of the tested ASes have ROV filtering behaviors. In addition, our method can be extended to the ROV measurement of IXPs. Through validation and analysis, we prove that our method leads to convincing results and, at the same time, has a broader coverage and better applicability than other existing methods.
Han Zhang 0009, Changqing An, Jilong Wang 0001
ISCC5
2024 Leveraging RAG-Enhanced Large Language Model for Semi-Supervised Log Anomaly Detection
abstract
Log-based anomaly detection is critical in monitoring the operations of information systems and in the real-time reporting of system failures. Utilizing deep learning-based log anomaly detection methods facilitates effective detection of anomalies within logs. However, existing methods are greatly dependent on log parsers, and parsing errors can considerably affect downstream anomaly detection tasks. Additionally, methods that predict the next log event in a sequence are susceptible to the instability of sequences and the emergence of unseen logs as systems evolve, resulting in a higher false positive rate. In this paper, we put forward LogRAG, a semi-supervised log anomaly detection framework based on retrieval-augmented generation (RAG). This framework conducts phased detection using both Log Tokens and Log Templates to mitigate the impact of log parsing errors. It also utilizes a single-class classifier to model the normal behavior of the system, thereby circumventing the effects of unstable sequences. Finally, it employs large language model (LLM) empowered by RAG to reevaluate detected anomalous logs, thereby improving accuracy. LogRAG demonstrates a 15% improvement in F1 Score on the BGL dataset and a 60% improvement on the Spirit dataset when compared to the previous best semi-supervised learning algorithm.
Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, Jilong Wang 0001
ISSRE7
2024 An Efficient FEC Scheme with SLA Consideration for Low Latency Transmissions
abstract
Forward Error Correction (FEC) is the preferred method for recovering lost packets in time-sensitive applications. The key for FEC to recover lost packets successfully is whether the number of redundant packets is sufficient. Due to imperfect loss prediction algorithms, existing FEC schemes, which set the number of redundant packets according to the prediction results, make transport service providers likely to fall into the awkward situation of either failing to recover lost packets or wasting a large amount of bandwidth. In this work, we propose P-FEC, an FEC scheme that can take SLA into consideration and empirically achieve the targeted decoding success rate while minimizing bandwidth waste. P-FEC combines intra- and inter-generation coding to balance decoding success rate and bandwidth waste, in which intra-generation provides quick but conservative recovery and inter-generation coding provides delayed but more efficient loss recovery services. We future profile the loss prediction errors to derive cumulative distribution functions of the errors for diverse network conditions, and then determine the parameters of intra- and inter-generation coding according to these distribution functions. The real-world transmission experiments empirically demonstrate that P-FEC can achieve the targeted decoding success rate while its bandwidth waste is only 1%-16% of the FEC with the code rate that can theoretically guarantee the target success rate. Furthermore, P-FEC can work well with computation light-weighted prediction algorithms although these algorithms have low accuracy, which makes it extremely useful for the transmission environment with limited computing resources.
Chao Xu 0015, Hui Wang 0011, Zongpeng Li, Jilong Wang 0001
NOMS6
2024 Investigate and Improve the Certificate Revocation in Web PKI
abstract
The validity and efficiency of certificate revocation in today's web Public Key Infrastructure (PKI) are consistently overlooked. In this paper, we analyse current certificate revocation schemes from the perspective of Certificate Authorities (CAs) and browsers. We find that the average size of CRL files collected from popular CAs can be as large as 2MB and the average response time of OCSP is around 430ms, which means that the time overhead brought by current certificate revocation schemes is not negligible. Moreover, browsers often fail to perform the revocation check correctly and allow websites to use revoked certificates, which can help attackers to launch man-in-the-middle attacks using fraudulent certificates.We also summarise existing problems and propose a novel certificate revocation scheme utilizing DNS resource records for efficient and privacy-preserving distribution. We have implemented a prototype to evaluate the performance of our scheme, and test results show that our scheme is time-efficient and able to withstand realistic workloads. We hope that our work can stimulate further discussion of the problems in Web PKI.
Changqing An, Zhiyan Zheng, Jilong Wang 0001
NOMS5
2024 Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster
abstract
Big data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions.
Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao
SIGCOMM13
2024 A First Look at FEC Code Rate Determination from a Computational Cost Perspective
abstract
Forward Error Correction (FEC) is the preferred method for recovering lost packets in time-sensitive applications. However, setting the code rate to adapt to changing network environment is still a significant but challenging problem. Recently, researchers explored deep learning (DL) methods to solve this problem, but DL-based methods come with increased computation overhead and decision-making latency compared to traditional methods, which may make these methods infeasible or less effective. Thus, real-time network (RTN) providers usually have trouble with these two questions when considering whether to adopt a DL-based method: (1). How much additional computing resources are required by each DL-based method? (2). If resource competition occurs, can these DL-based methods be able to make timely code rate decisions? In this paper, we evaluate existing proposed DL-based methods for code rate determination in terms of computational cost to answer these two questions. Particularly, we measure the computational resources and the time required by these DL-based methods for making one code rate decision. We find that among existing DL-based methods, some methods require several times or even tens of times more computational resources to achieve timely decision-making. Furthermore, we identify that three methods are prone to resource competition with existing FEC schemes, which indicates that RTN providers need to configure their computing resources carefully when selecting these methods. We further analyze the reasons behind this phenomenon and provide some design suggestions for RTN nroviders.
Chao Xu 0015, Hui Wang 0011, Shilin Xie, Jilong Wang 0001
WCNC4
2024 Proactively Verifying Quantitative Network Policy Across Unsafe and Unreliable Environments
abstract
Network managers configure networks to enforce various high-level policies, and to respond to the wide range of network events (e.g., attacks, intrusions, malicious route announcements from neighbors) that may occur. It is incredibly difficult to specify these high-level policies in terms of distributed low-level configuration. These high-level policies hold only if the distributed configurations are well equipped to react to unsafe and unreliable environments (e.g., malicious route announcements, unsafe components and devices). Therefore, it is important to proactively verify whether network policies hold across continually changing environments in terms of current network configurations. State-of-the-art policy verification techniques are limited because they can check only the Boolean policies (e.g., forwarding reachability, waypoint or blackhole-freeness). However, many policy violations express themselves in quantitative ways (e.g., a link becomes overloaded). In this paper, we propose quantitative network verification (QNV) analyzing the quantitative policies of networks across unsafe and unreliable environments. QNV translates network configurations into a symbolic simulation model that captures the stable states to which the network forwarding will converge as a result of interactions between routing protocols. It then generates a logical formula matrix that describes network forwarding in the event of failures and verifies quantitative policies based on the formula matrix. We implement QNV and evaluate it on realistic and synthetic configurations. Our evaluation shows that QNV can precisely verify quantitative policies in only a few minutes, even in large networks.
Han Zhang 0009, Jilong Wang 0001, Xingang Shi, Xia Yin 0001, Jiankun Hu, Congcong Miao
IEEE Trans. Inf. Forensics Secur.3
2024 Graph Structure Reshaping Against Adversarial Attacks on Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have achieved impressive performance in many tasks on graph data. Recent studies show that they are vulnerable to adversarial attacks. Deliberate and unnoticeable perturbations on topology structure could render them near-useless in applications. How to design effective methods to improve the robustness of GNNs is a crucial problem. To solve this problem, some works attempt to design more robust GNN models, while others attempt to remove perturbations from the poisoned graph. Different from the previous works, this paper proposes a general framework termed asGraphReshapeto enhance the robustness of GNNs via directly correcting the shifted classification boundary of GNN models in the presence of adversarial attacks.GraphReshapeconsists of two modules:locating tractive nodesthat could correct GNNs andreshaping local structureto improve their representations in the latent space. Extensive experiments on four real-world datasets show thatGraphReshapeachieves significant performance gain compared with state-of-the-art baselines against different adversarial attacks.
Haibo Wang 0004, Chuan Zhou 0001, Jia Wu 0001, Shirui Pan, Zhao Li 0007, Jilong Wang 0001, Philip S. Yu
IEEE Trans. Knowl. Data Eng.7
2024 Live Migration of Video Analytics Applications in Edge Computing
abstract
In order to schedule resources efficiently or maintain applications' continuity for mobile customers, edge platforms often need to adaptively migrate the applications on them. However, our measurement shows that existing migration solutions cannot solve the issue of migrating video analytics applications in edge computing because the memory states of video analytics applications have different characteristics from other applications. We conduct a breakdown analysis of the memory states of video analytics applications, and propose to treat three types of states separately with three different techniques,i.e., warm-up, sync, and replay, to minimize the negative influence of migrations on application performance. Based on this idea, we implement a prototype system in which two new components,i.e.,state storeandsidecar, are designed to achieve near-transparent live migration with minimal application code modifications. Evaluation experiments demonstrate that the time of application interruption caused by migrating a video analytics application with our solution is less than 405ms, and our solution does not consume much resources.
Chenghao Rong, Hui Wang 0011, Jilong Wang 0001, Yipeng Zhou, Jun Zhang 0004
IEEE Trans. Mob. Comput.3
2023 TinyG: Accurate IP Geolocation Using a Tiny Number of Probers
abstract
IP geolocation is essential for various applications. However, the reliability of IP geolocation databases has been proven to be inadequate. In recent years, the growing number of public probers has offered the potential for more accurate geolocation results through active measurement. The conventional practice is to probe the target IP address using all available probers and feed the measurement results to the active geolocation method. However, this practice is cost-inefficient and may trigger the anti-flood mechanism. Moreover, public probers typically impose user-level limits on the frequency and quantity of measurements. Therefore, it is important to reduce the average number of probers (ANP) selected for successfully probing each target. Researchers have discovered that geolocation accuracy primarily depends on the minimum delay between probers and the target. Inspired by that, we propose TinyG, a prober selection algorithm designed to reduce the ANP needed to find probers within a sufficiently small delay from the target. TinyG divides the probing process into multiple rounds and leverages previous measurement results to guide the selection of probers for subsequent rounds. Experimental results show that when the associated minimum delay is within 2 ms, various active geolocation methods can provide credible geolocation results. TinyG outperforms other algorithms in reducing the ANP needed to obtain credible results. Compared to using more than 1,300 probers, TinyG can achieve an ANP of 6.7 with only a 6% coverage loss of credible results.
Hui Wang 0011, Jilong Wang 0001, Peiran Wang
CNSM3
2023 Top AS Router Geolocation in Databases: Performance and Techniques
abstract
Autonomous systems (ASes) at the top of the global transit hierarchy play core roles in the Internet. Thus, accurately geolocating their routers is crucial for drawing credible conclusions on topics such as network resilience and traffic censorship. IP geolocation databases (DBs) are frequently used for this purpose. However, the accuracy of DBs on top AS routers has not been fully evaluated due to the lack of a comprehensive ground truth dataset. In this study, we address this gap by constructing a geolocation ground truth dataset that contains more than 12,000 router interfaces in the top 10 ASes, utilizing delay measurements from carefully selected Looking Glass vantage points. We evaluate the coverage and accuracy of 6 DBs, including 4 widely used ones and 2 new ones. Our evaluation shows that most DBs exhibit poor accuracy when geolocating top AS routers. We conduct an in-depth analysis to uncover the primary techniques behind DBs. Our investigation reveals the reasons behind the poor performance of certain DBs. Moreover, we discover that the best-performing DB heavily relies on hostnames, and all DBs perform poorly in geolocating routers without hostnames. Our dataset, which is the largest ground truth dataset of top AS routers to the best of our knowledge, will be publicly accessible to the Internet research community.
Hui Wang 0011, Jilong Wang 0001, Peiran Wang
GLOBECOM3
2023 Be Careful of Your Neighbors: Injected Sub-Prefix Hijacking Invisible to Public Monitors
abstract
Prefix hijackings have always been a significant security issue in BGP and have continued to occur in recent years. Detecting prefix hijackings is a vital part of defending against them. Most detection approaches mainly rely on the feed from the monitors of public route collector infrastructures. We propose an injected sub-prefix hijacking that utilizes the BGP communities attribute and AS path poisoning to control the propagation of invalid sub-prefix routes. This attack only pollutes neighboring ASes, thus guaranteeing the invisibility to monitors. Then the attacker can stealthily hijack traffic passing through the polluted ASes. Through extensive simulations, we show that this attack has an enormous impact and propose the crucial indicator affecting the attacker's capability. Finally, we demonstrate that existing defenses are difficult to handle this attack and then propose several defense strategies against it.
Han Zhang 0009, Changqing An, Jilong Wang 0001
ICC6
2023 PUVAR:Minimize Idle Resource SLO Violations by Uncertainty-Aware Scheduling in Cloud Platforms
abstract
Nowadays, idle resource makes up a non-negligible fraction of datacenter capacity in mainstream cloud platforms. Cloud platforms offer idle resource with low service level objectives (SLOs) at low prices to attract cost-sensitive users. Despite their fault tolerance, these users still want some SLO guarantees for idle resource. Cloud platforms have started to provide statistical SLO for idle resource, but the violations of statistical SLOs have not been modeled and minimized. In this paper, we propose PUVAR, a scheduling policy based on prediction+optimization, to explicitly model and minimize the statistical SLO violation for idle resource in cloud platforms. The design of PUVAR possesses two major innovations: (1) explicitly quantify prediction uncertainty of idle resource capacity and iteratively reduce its impact on scheduling decisions; (2) treat the SLO of idle resources as a soft constraint and minimize its violation by two-stage scheduling of regular requests and idle requests. We provide a theoretical convergence rate for the parameter optimization of PUVAR. Promising results of ablation analysis by comparison with popular baseline algorithms on the trace of a real cloud platform indicate that PUVAR can significantly reduce the SLO violation of idle resource with little additional cost on the current scheduling of regular requests.
Han Zhang 0009, Jilong Wang 0001
ICWS3
2023 Knocking Cells: Latency-Based Identification of IPv6 Cellular Addresses on the Internet
abstract
IPv6 mobile networks are becoming increasingly important. Many jobs rely on understanding IPv6 mobile networks at the IP level. Previous works on cellular identification suffer from coarse identification granularity, proprietary data, or not working for IPv6. The high latency in mobile networks makes identifying cellular addresses possible using Round-Trip Time (RTT) difference. However, due to the impact of packet loss on the measurement of RTT difference, identifying cellular addresses with less overhead is challenging. In this paper, by triggering the non-zero RTT difference of the cellular /48 subnets with probes, we propose an accurate latency-based method to identify cellular /48 subnets from fixed subnets. Experiments demonstrate that the method can identify cellular subnets with a precision of 93.52% to 99.95% and a recall of 99.96% on a worldwide dataset. The overhead of measuring RTT difference reduces to at least 1/10th compared to the existing methods while robust to packet loss.
Han Zhang 0009, Anlun Hong, Jilong Wang 0001
ISCC6
2023 Metis: Detecting Fake AS-PATHs Based on Link Prediction
abstract
BGP route hijacking is a critical threat to the Internet. Existing works on path hijacking detection firstly monitor the routes of the whole network and then directly trigger a suspicious alarm if the link has not been seen before. However, these naive approaches will cause false positive identification and introduce unnecessary verification overhead. In this work, we propose Metis, a matching-and-prediction system to filter out normal unseen links. We first use a matching method with three rules to find out suspicious links if there is an unseen AS. Otherwise, we propose using a neural network to make a prediction based on the AS information at each end of the link and further quantify the suspicion level. Our large-scale simulation results show that Metis can achieve precision and recall of over 80% for detecting fake AS-PATHs. Moreover, our deployment experiences show that compared to state-of-the-art system, Metis can save 80% overhead.
Chengwan Zhang, Congcong Miao, Changqing An, Anlun Hong, Ning Wang 0001, Jilong Wang 0001
ISCC7
2023 Delay Based Congestion Control for Cross-Datacenter Networks
abstract
Numerous distributed applications are deployed in the cross-datacenter networks (Cross-DC) where geographically distributed data centers (DC) are connected by wide area network (WAN). These online applications will generate both intra-datacenter and inter-datacenter traffic, each with distinct requirements and characteristics. We find that existing combined congestion control schemes ignore the interaction of the two types of traffic and the hybrid congestion control schemes fail to accurately estimate cross data center network congestion extent. In this paper, we propose IDCC a delay based congestion control scheme that uses delay to handle congestion inside the DC and in the WAN, respectively. We respectively utilize In-band network telemetry (INT) and round trip time (RTT) to measure the queuing delay inside DC and in the WAN and guarantee the stability of the algorithm by Proportional Integral Derivative (PID). Simultaneously, we demonstrate the empirical results of optimizing flow completion time (FCT) of intra-DC short flow by weight function in cross-DC. We implemented IDCC in simulation platform ns-3 and have performed extensive large scale simulation evaluations. Results show that IDCC decreases the FCT of intra-DC traffic by 3.6× to 12× and improves the throughput of inter-DC traffic by 9% to 16% compared to Gemini, Annulus.
Yantao Geng, Han Zhang 0009, Xingang Shi, Jilong Wang 0001, Xia Yin 0001, Dongbiao He
IWQoS4
2023 FTM-RCA: A Fast Two-Stage Multi-dimensional Root-Cause Analysis of Network Anomalies
abstract
Multi-dimensional Root Cause Analysis (RCA) is often applied to identify abnormal traffic patterns, i.e., localizing the abnormal combination of traffic header fields. Several techniques have been proposed recently, but they were mainly designed for smaller-scaled datasets and were not feasible in the real network due to the high computational overhead. To overcome the aforementioned limitations, we propose FTM-RCA, which accelerates RCA by breaking the analysis procedure into two stages: coarse-grained rules filtering and fine-grained localization. In the first stage, an optimized frequent itemset mining (FIM) technique called CUSC is proposed, which can detect high-volume combinations faster based on the mutual exclusion of dimension values. Experiments on CUSC show that it can speed up by 44.87% and reduce memory consumption by 21.89% compared to the best previous FIM algorithms. In the second stage, a dimension-based search method is proposed to identify the root cause combinations, which consists of two key components: 1) drill-down strategy, which utilizes Contributive Power to measure the correlation between the combination and anomaly. 2) pruning strategy, which adopts the Shannon entropy to avoid generating trivial results. As a result, the overall diagnostic time of FTM-RCA is at least 25 times faster than the previous best research while improving accuracy by an average of 21.6%. Also, our practical application in real network also illustrates the applicability of FTM-RCA.
Yeqing Meng, Qianli Zhang, Xiangyu Tang, Wanhao Zhang, Jilong Wang 0001
IWQoS5
2023 Evaluating and Improving Regional Network Robustness from an AS TOPO Perspective
abstract
Currently, regional networks are subject to various security attacks and threats, which can cause the network to fail. This paper borrows the quantitative ranking idea from the fields of statistics and proposes a ranking method for evaluating regional resilience. Large-scale simulated failure events based on probabilistic sampling is performed, and a significance tester that measures the impact of events from the overall level and variance aspect is also implemented. To improve a region’s robustness, this paper proposes a greedy algorithm to optimize the resilience of regions by adding key links among AS. This paper selects the AS topology of 50 countries/regions for research and ranking, evaluating the topology robustness from connectivity, user, and domain influence perspectives, clustering the results and get typical region types, and adding optimal links to improve the network resilience. Experimental results illustrate that the resilience of regional networks can be greatly improved by establishing a few new connections, which demonstrates the effectiveness of the optimization method.
Changqing An, Zhiyan Zheng, Zidong Pei, Jilong Wang 0001, Chalermpol Charnsripinyo
NOMS6
2023 TENSOR: Lightweight BGP Non-Stop Routing
abstract
As the solitary inter-domain protocol, BGP plays an important role in today's Internet. Its failures threaten network stability and will usually result in large-scale packet losses. Thus, the non-stop routing (NSR) capability that protects inter-domain connectivity from being disrupted by various failures, is critical to any Autonomous System (AS) operator. Replicating the BGP and underlying TCP connection status is key to realizing NSR. But existing NSR solutions, which heavily rely on OS kernel modifications, have become impractical due to providers' adoption of virtualized network gateways for better scalability and manageability.
Congcong Miao, Yunming Xiao, Marco Canini, Ruiqiang Dai, Shengli Zheng, Jilong Wang 0001, Jiwu Bu, Aleksandar Kuzmanovic, Yachen Wang
SIGCOMM6
2023 FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical Backbones
abstract
The rising demand for WAN capacity driven by the rapid growth of inter-data center traffic poses new challenges for costly optical networks. Today cloud providers rely on fixed optical backbones, where all hardware devices operate on a rigid spectrum grid, leading to the waste of expensive optical resources and subpar performance in handling failures. In this paper, we introduce FlexWAN, a novel flexible WAN infrastructure designed to provision cost-effective WAN capacity while ensuring resilience to optical failures. FlexWAN achieves this by incorporating spacing-variable hardware at the optical layer, enabling the generated wavelength to optimize the utilization of limited spectrum resources for the WAN capacity. The configuration of spacing-variable hardware in a multi-vendor optical backbone presents challenges related to spectrum management. To address this, FlexWAN leverages a centralized controller to achieve coordinated control of network-wide optical devices in a vendor-agnostic manner. Moreover, the flexibility at the optical layer introduces new algorithmic problems. FlexWAN formulates the problem of provisioning WAN capacity with the goal of minimizing hardware costs. We evaluate the system performance in production and share insights from years of production experience. Compared to existing optical backbones, FlexWAN can save at least 57% of transponders and reduce 36% of spectrum usage while continuing to meet up to 8× the present-day demands using existing hardware and fiber deployments. FlexWAN further incorporates failure resilience that revives 15% more bandwidth capacity in the overloaded optical backbone.
Congcong Miao, Zhizhen Zhong, Ying Zhang 0022, Kunling He, Fangchao Li, Minggang Chen, Xiang Li 0223, Zekun He, Xianneng Zou, Jilong Wang 0001
SIGCOMM11
2023 Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero Code
abstract
Microservices are becoming more complicated, posing new challenges for traditional performance monitoring solutions. On the one hand, the rapid evolution of microservices places a significant burden on the utilization and maintenance of existing distributed tracing frameworks. On the other hand, complex infrastructure increases the probability of network performance problems and creates more blind spots on the network side. In this paper, we present DeepFlow, a network-centric distributed tracing framework for troubleshooting microservices. DeepFlow provides out-of-the-box tracing via a network-centric tracing plane and implicit context propagation. In addition, it eliminates blind spots in network infrastructure, captures network metrics in a low-cost way, and enhances correlation between different components and layers. We demonstrate analytically and empirically that DeepFlow is capable of locating microservice performance anomalies with negligible overhead. DeepFlow has already identified over 71 critical performance anomalies for more than 26 companies and has been utilized by hundreds of individual developers. Our production evaluations demonstrate that DeepFlow is able to save users hours of instrumentation efforts and reduce troubleshooting time from several hours to just a few minutes.
Junxian Shen, Han Zhang 0009, Xingang Shi, Yunxi Shen, Yongxiang Wu, Xia Yin 0001, Jilong Wang 0001, Mingwei Xu 0001, Jiping Yin, Jianchang Song, Zhuofeng Li, Runjie Nie
SIGCOMM10
2023 Real-Time Malicious Traffic Detection With Online Isolation Forest Over SD-WAN
abstract
Software Defined Network (SDN) has been widely used in modern network architecture. The SD-WAN is considered as a technology that has a potential to revolutionize the WAN service usage by utilizing the SDN philosophy. Attacking SDN router and controller can affect the network and block the entire services. In this paper, we propose a machine learning based anomalous traffic detection framework named OADSD over SD-WAN that can achieve task independent and has the ability of adapting to the environment. The OADSD adopts Distributed Dynamic Feature Extraction (DDFE) to extract representative features directly from the raw traffic, and proposes the On-demand Evolving Isolation Forest (OEIF) to make the system adapt to an environment. We provide a theoretical analysis of the performance of the OADSD. We also conduct comprehensive experiments to evaluate the performance of the OADSD with real world public datasets as well as a small real testbed. Our experiments under real world public datasets show that, the OADSD can accurately detect various kinds of attacks with a high performance. Compared with the state-of-the-art systems, the OADSD can achieve up to 60% accuracy improvement.
Pei Zhang 0003, Fangzhou He, Han Zhang 0009, Jiankun Hu, Xiaohong Huang 0003, Jilong Wang 0001, Xia Yin 0001, Huahong Zhu
IEEE Trans. Inf. Forensics Secur.6
2023 Offloading Elastic Transfers to Opportunistic Vehicular Networks Based on Imperfect Trajectory Prediction
abstract
Due to the high cost of cellular networks, vehicle users would like to offload elastic traffic through vehicular networks as much as possible. This demand prompts researchers to consider how to make the vehicular network system achieve better performance for requests coming online, such as maximizing throughput. The traffic in vehicular networks is transferred through opportunistic contacts between vehicles and infrastructures. When making scheduling decisions, the scheduler must be aware of vehicles’ future trajectories. Vehicles’ future trajectories are usually predicted by trajectory prediction algorithms when users are unwilling to report their future trips. Unfortunately, no trajectory prediction algorithm can be completely accurate, and these inaccurate prediction results will degrade the throughput achieved by scheduling algorithms. In this paper, we focus on reducing the negative impact of inaccurate predictions. Specifically, we measure two data-driven trajectory prediction algorithms that have been widely used for trajectory predictions and understand the characteristics of the accuracy of predicted contacts. Based on the enlightenment from the measurement, we design a system, i.e., i-Offload, to offload elastic traffic under imperfect trajectory predictions. The experimental results show that our system has good throughput and high scheduling efficiency even under imperfect trajectory predictions. Compared with existing scheduling algorithms, our method improves the throughput by about one time.
Chao Xu 0015, Hui Wang 0011, Jilong Wang 0001, Yipeng Zhou, Yuedong Xu 0001, Di Wu 0001, Changqing An
IEEE/ACM Trans. Netw.3
2023 Achieving High Availability in Inter-DC WAN Traffic Engineering
abstract
Inter-DataCenter Wide Area Network (Inter-DC WAN) that connects geographically distributed data centers is becoming one of the most critical network infrastructures. Due to limited bandwidth and inevitable link failures, it is highly challenging to guarantee network availability for services, especially those with stringent bandwidth demands, over inter-DC WAN. We present$\mathsf {TEDAT}$, a novel Traffic Engineering (TE) framework for Diverse Availability Targets (DAT), where a Service Level Agreement (SLA) is defined to ensure that each bandwidth demand must be satisfied with a stipulated probability, when subjected to the network capacity and possible failures of the inter-DC WAN.$\mathsf {TEDAT}$has two core components, i.e., traffic scheduling and failure recovery, which are crystalized through different mathematical models and theoretically analyzed. They are also extensively compared against state-of-the-art TE schemes, using a testbed as well as real trace driven simulations across different topologies, traffic matrices and failure scenarios. Our evaluations show that, compared with the optimal admission strategy,$\mathsf {TEDAT}$can speed up the online admission control by$30\times $at the expense of less than 4% false rejections. On the other hand, compared with the latest TE schemes like FFC and TEAVAR,$\mathsf {TEDAT}$can meet the bandwidth availability SLAs for 23%~60% more demands under normal loads, and when network failure causes SLA violations, it can retain 10%~20% more profit under a pricing and refunding model.
Han Zhang 0009, Xia Yin 0001, Xingang Shi, Jilong Wang 0001, Yingya Guo, Tian Lan 0001, Ke Ruan, Haijun Geng
IEEE/ACM Trans. Netw.4
2023 A General Approach to Generate Test Packets With Network Configurations
abstract
The correctness and reliability of modern networks are often the greatest concerns. A myriad network events like software update, device crash and resource exhaustion, inevitably lead to liveness errors on data plane. This paper focuses on fault detection of the network data plane using test packets. Existing test packet generation techniques are limited in two aspects: i) it is difficult to collect the input data plane snapshot through SNMP or terminals ii) it may rise false negatives due to inconsistent snapshot. In this paper, we propose a new framework, SWIFT, that automatically generates test packets with network configurations. SWIFT minimizes the number of test packets by allowing a packet to go through multiple links or interfaces. For network updates, SWIFT updates test packets in an incremental way to revalidate the network. We evaluate its performance using hundreds of benchmark network configurations. The results show that it takes few seconds to generate test packets to exercise all links and interfaces, and updates the test packets in few seconds for configuration changes. We also deployed a SWIFT prototype in a university network, and successfully detected many network outages.
Han Zhang 0009, Jilong Wang 0001, Xia Yin 0001, Xingang Shi
IEEE Trans. Parallel Distributed Syst.3
2023 Serpens: A High Performance FaaS Platform for Network Functions
abstract
More and more enterprises deploy applications on Function-as-a-Service (FaaS) platforms to improve resource efficiency and save monetary costs. Network Functions (NFs) suffer from staggered peaks of traffic patterns and could benefit from fine-grained resource multiplexing in FaaS platform. However, naively exploring existing FaaS platforms to support NFs can introduce significant performance overheads in three aspects, including slow instance startup, remote state access for NFs, and costly packet delivery between NFs. To address these problems, we propose${\sf Serpens}$, a high performance FaaS platform for NFs. First,${\sf Serpens}$proposes a reusable NF runtime design to slash instance startup overhead. Second,${\sf Serpens}$designs a novel state management mechanism to support local state access. Third,${\sf Serpens}$introduces an advanced service chaining approach to avoid extra packet delivery. Besides,${\sf Serpens}$designs an NF scaling mechanism to minimize performance fluctuation. We have implemented a prototype of${\sf Serpens}$and conducted comprehensive experiments. Compared with the NFs and Service Function Chains (SFCs) that run on existing FaaS platforms,${\sf Serpens}$can improve the throughput by more than 10× and reduce the latency by more than 90%.
Heng Yu 0005, Han Zhang 0009, Junxian Shen, Yantao Geng, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001
IEEE Trans. Parallel Distributed Syst.5
2022 Gringotts: Fast and Accurate Internal Denial-of-Wallet Detection for Serverless Computing
abstract
Serverless computing, or Function-as-a-Service, is gaining continuous popularity due to its pay-as-you-go billing model, flexibility, and low costs. These characteristics, however, bring additional security risks, such as the Denial-of-Wallet (DoW) attack, to serverless tenants. In this paper, we perform a real-world DoW attack on commodity serverless platforms to evaluate its severity. To identify such attacks, we design, implement, and evaluate Gringotts, an accurate, easy-to-use DoW detection system with a negligible performance overhead. Gringotts addresses the information ambiguity inherent in serverless functions by introducing a well-designed performance metrics collection agent. Then, Gringotts uses the Mahalanobis distance to discover anomalies in the distribution of the metrics. We implement Gringotts as a real system and conduct extensive experiments using a testbed to evaluate the performance of Gringotts. Our results indicate that Gringotts has a performance overhead of less than 1.1%, with an average detection delay of 1.86 seconds and an average accuracy of over 95.75%.
Junxian Shen, Han Zhang 0009, Yantao Geng, Jilong Wang 0001, Mingwei Xu 0001
CCS5
2022 Distributed Routing Controller for Large-scale Live Video Streams in Real-Time Networks
abstract
Overlay routing control for live video delivery has become an important yet challenging task for Real-Time Networks (RTNs), but existing approaches designed for traditional Content Delivery Networks (CDNs) fall short of meeting the challenge. In this paper, we develop a distributed overlay routing controller for RTNs to deliver massive high-quality live video streams in low latencies. We first formulate a joint optimization that offers rich control flexibility and can yield low-latency and cost-effective routing solutions. To obtain the appealing potential value of the optimal solution in the context of large-scale live videos, we develop a distributed control framework that can find near-optimal solutions at fine-grained timescales. Evaluations on real-world live video traces show that our distributed controller derives high-quality (in terms of both performance and cost) overlay routing solutions while reducing the decision latency by 38%-89% compared to the state-of-the-art centralized controller.
Hui Wang 0011, Chao Xu 0015, Jilong Wang 0001
GLOBECOM5
2022 PipeCompress: Accelerating Pipelined Communication for Distributed Deep Learning
abstract
Distributed learning is widely used to accelerate the training of deep learning models, but it is known that communication efficiency limits the scalability of distributed learning systems. Current gradient compression techniques provide promising methods to reduce communication time, but the extra time incurred by compression is not negligible. After compression techniques are applied, the communication time is significantly reduced because the data size needed to communicate becomes much smaller, but compressing gradients is time-consuming and it becomes a new bottleneck. In this paper, we design and implement PipeCompress, a system to decouple compression and backpropagation operations into two processes and pipeline the two processes to hide compression time. We also propose a specialized inter-process communication mechanism based on the characteristics of DNN distributed training to improve the efficiency of passing messages between the two processes, which makes sure that the decoupling does not bring much extra inter-process communication time cost. As far as we know, this is the first work that notices the overhead of compression and pipelines backpropagation and compression operations to hide compression time in distributed learning. Experiments show that PipeCompress can significantly hide compression time, reduce iteration time, and accelerate the training process on various DNN models.
Juncai Liu, Hui Wang 0011, Chenghao Rong, Jilong Wang 0001
ICC4
2022 Phishing Detection Based on Multi-Feature Neural Network
abstract
Phishing detection methods are used to protect Internet users from leaking private information to phishing websites. However, the passive phishing detection method, which is widely used and based on blacklists, has limitations on timeliness and defense against zero-day phishing attacks and active phishing detection models with single-feature can be easily targeted by attackers. It is necessary to design and establish an active phishing detection model with high timeliness and strong adaptability. We propose a phishing detection method based on multi-feature extraction and deep learning technology. The model is constructed of a multilayer perceptron (MLP) for self-defined feature, a convolutional neural network (CNN) for image feature, a recurrent neural network (RNN) for text feature to extract feature vectors, and a classification network to fuse features and make the judgement. Our model’s accuracy achieves 0.9775 and recall reaches up to 0.9901. Experiment results of our model prove superior performance to those of other classification algorithms, demonstrating our model’s ability to deal with complex and changeable phishing detection tasks at this stage.
Shuaicong Yu, Changqing An, Tianshu Li, Jilong Wang 0001
IPCCC6
2022 Scorpius: Proactive Code Preparation to Accelerate Function Startup
abstract
Massive enterprises deploy their applications on public clouds to relieve infrastructure management burden. However, applications are faced with highly fluctuating workloads, while clouds provision exclusive resources at coarse time granularity, resulting in severely low resource efficiency. Function-as-a-Service (FaaS) platform enables fine-grained resource multiplexing, which has the potential to improve efficiency. However, FaaS platforms could consume several seconds to start functions and the long startup latency can severely hurt the performance of applications. In this paper, we measure the FaaS platforms and find that most startup latency is occupied by code preparation. To reduce the code preparation latency with little resource overhead, we propose Scorpius, a FaaS platform that proactively prepares code based on the historical data of functions. It combines two optimization categories: (1) To reduce the code size, Scorpius proposes to proactively prepare partial libraries over servers and run functions on the server with most library sharing. (2) To advance the start time, Scorpius proposes to predict the function overload with a simple model and proactively scale code to more servers. We have implemented a prototype of Scorpius and conducted extensive experiments. Evaluation results demonstrate that compared with state-of-the-art methods, Scorpius can reduce the code preparation latency by 87.6% with only 9.3% storage overhead.
Heng Yu 0005, Junxian Shen, Han Zhang 0009, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001
IWQoS4
2022 Predicting Unseen Links Using Learning-based Matrix Completion
abstract
Researchers have noticed the AS-level Internet topology that can be observed from the current measurement infrastructure is far from complete, which means researchers have to deploy more measurement vantage points (VPs) and conduct measurements for more source/destination pairs to fully understand the whole Internet. Unfortunately, it is known that blindly deploying more points and conducting more measurements to achieve the goal is inefficient, if not infeasible. In this paper, we try to improve the efficiency by predicting where unseen AS links might be located from the observed AS paths to guide the measurements towards a more complete AS-level topology. We formulate the prediction of unseen links as a matrix completion problem. However, the traditional matrix completion methods have limited learning capacities and cannot deal with the complex constraints on the underlying topology. We develop a learning-based matrix completion method specifically for the unseen AS link prediction problem. The method exploits a neural network and utilizes side-information which is carefully chosen from AS attributes based on our understanding on Internet peering practices, therefore our method is able to learn more expressive latent vectors and achieves outstanding prediction performance in our scenario. Experiments performed on a real-world dataset show the prediction results can achieve a high AUC (Area Under the Receiver Operating Characteristic Curve) of 0.834.
Shuying Zhuang, Hui Wang 0011, Jilong Wang 0001, Changqing An, Yuedong Xu 0001, Tianhao Wu 0010
NOMS3
2022 Detecting Ephemeral Optical Events with OpTel
Congcong Miao, Minggang Chen, Arpit Gupta, Zili Meng, Lianjin Ye, Jingyu Xiao, Zekun He, Xulong Luo, Jilong Wang 0001, Heng Yu 0005
NSDI10
2022 RouteInfer: Inferring Interdomain Paths by Capturing ISP Routing Behavior Diversity and Generality
Tianhao Wu 0010, Hui Wang 0011, Jilong Wang 0001, Shuying Zhuang
PAM3
2022 Exploring the Layered Structure of Containers for Design of Video Analytics Application Migration
abstract
The existing solutions to the migration of container-based applications are not suitable for live video analytics applications because these solutions can result in excessive migration time. Intuitively, it is possible to exploit the layered structure of containers to improve the migration performance, but we still need a good understanding of the characteristics of the containers related to video analytics applications to justify the intuitive idea and design a high-performance solution. In this paper, we pull 3735 representative images from Docker Hub. We analyze these images and get three main findings: (1) the images of video analytics applications have more layers and larger sizes than that of general applications; (2) we can cache images in the destination servers and reuse the same layers cached in the destination server to reduce the migration time when an image is migrated from its source server to its destination server; (3) the size of the remaining data to be transferred during migrations is still large and a high-performance migration scheme is still necessary. Based on the above findings, we propose a pipelined migration scheme to optimize migration performance. Evaluations show that pipelined migration performs significantly better than other migration schemes.
Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001
WCNC5
2022 Predicting Human Mobility via Graph Convolutional Dual-attentive Networks
abstract
Human mobility prediction is of great importance for various applications such as smart transportation and personalized recommender systems. Although many traditional pattern-based methods and deep models ($e.g.,$ recurrent neural networks) based methods have been developed for this task, they essentially do not well cope with the sparsity and inaccuracy of trajectory data and the complicated high-order nature of the sequential dependency, which are typical challenges in mobility prediction. To solve the problems, this paper proposes a novel framework named G raph C onvolutional D ual-a ttentive N etworks (GCDAN), which consists of two modules: spatio-temporal embedding and trajectory encoder-decoder. The first module employs a bidirectional diffusion graph convolution to preserve the spatial dependency in the location embedding. The second module employs a dual-attentive mechanism based on a Sequence to Sequence architecture to effectively extract the long-range sequential dependency within a trajectory and the correlation between different trajectories for predictions. Extensive experiments on three real-world datasets show that GCDAN achieves significant performance gain compared with state-of-the-art baselines.
Weizhen Dang, Haibo Wang 0004, Shirui Pan, Pei Zhang 0003, Chuan Zhou 0001, Jilong Wang 0001
WSDM7
2022 USST: A two-phase privacy-preserving framework for personalized recommendation with semi-distributed training
Yipeng Zhou, Jun Liu 0001, Hui Wang 0011, Jilong Wang 0001, Guanfeng Liu 0001, Di Wu 0001, Chao Li 0067, Shui Yu 0001
Inf. Sci.4
2022 Scheduling Massive Camera Streams to Optimize Large-Scale Live Video Analytics
abstract
In smart cities, more and more government departments will make use of live analytics of videos from surveillance cameras in their tasks, such as vehicle traffic monitoring and criminal detection. Obviously, it is costly for each individual department to deploy its own infrastructure,i.e., cameras and analytics system. In this paper, we consider a scenario in which a city deploys an infrastructure and departments submit requests to access and analyze videos for their own purposes. The live analytics of massive streams is computation-intensive and the tasks might be latency-critical, which makes scheduling massive streams to optimize all tasks an essential and challenging work. We exploit an end-edge-cloud architecture and propose an adaptive system to schedule the massive camera streams and tasks, which considers all factors affecting the computation and networking resource consumption,e.g., sharing of model computation, video quality, model partition, and task placement. Particularly, the resource consumption ofFaster R-CNN + ResNet101under each partition scheme is profiled for the first time and we notice the partition must be used together with lossless compression techniques to be beneficial. Furthermore, sometimes tasks might be required to migrate because the scheduling decision made by the system changes to adapt to the changing resource supply and demand. In order to avoid the performance degradation during migration, we propose a non-destructive migration scheme and implement it in the system. Simulations demonstrate our system achieves a total utility close to the maximum and our analytics system performs better than state-of-the-art solutions.
Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001, Sharon X. Huang
IEEE/ACM Trans. Netw.4
2022 NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation
abstract
In distributed storage systems, Erasure Coding (EC) is a crucial technology to enable high data availability. By downloading parity data from survived machines, EC can reconstruct lost data with much lower storage overheads than data replication. However, this reduction in storage cost comes at the expense of extra performance problems:low reconstruction rate,high degraded read latency, andhigh host CPU utilization. Our analysis shows that these performance problems are deeply rooted in thehost-basedEC processing. To resolve these problems, we present NetEC, an in-network accelerating framework that fully offloads EC to the new generation programmable switching ASICs. We propose Explicit Buffer Size Notification (EBSN) to constrain decoding buffer usage, and design an on-switch one-to-many TCP proxy to integrate EBSN with TCP. We also design two parallel Galois Field (GF) offloading methods—table lookup and bitmatrix methods—to maximize parsable bytes. We implement NetEC on programmable switches and integrate it with HDFS. Extensive evaluations show that NetEC improves the reconstruction rate by 2.7x-6.8x, reduces the degraded read latency significantly, and removes the host CPU overhead completely. We also emulate multi-rack scenarios and show that NetEC is able to support$\sim$∼GB/s reconstruction rate and tens of concurrent tasks.
Yi Qiao, Menghao Zhang 0001, Yu Zhou 0008, Han Zhang 0009, Mingwei Xu 0001, Jun Bi, Jilong Wang 0001
IEEE Trans. Parallel Distributed Syst.8
2021 Boosting bandwidth availability over inter-DC WAN
abstract
Inter-DataCenter Wide Area Network (Inter-DC WAN) that connects geographically distributed data centers is becoming one of the most critical network infrastructures. Due to limited bandwidth and inevitable link failures, it is highly challenging to guarantee network availability for services, especially those with stringent bandwidth demands, over inter-DC WAN. We present BATE, a novel Traffic Engineering (TE) framework for bandwidth availability (BA) provision, which aims to ensure that each bandwidth demand must be satisfied with a stipulated probability, when subjected to the network capacity and possible failures of the inter-DC WAN. The three core components of BATE, i.e., admission control, traffic scheduling and failure recovery, are formulated through different mathematical models and theoretically analyzed. They are also extensively compared against state-of-the-art TE schemes, using a testbed as well as real trace driven simulations across different topologies, traffic matrices and failure scenarios. Our evaluations show that, compared with the optimal admission strategy, BATE can speed up the online admission control by 30x at the expense of less than 4% false rejections. On the other hand, compared with the latest TE schemes like FFC and TEAVAR, BATE can meet the bandwidth availability targets for 23%~60% more demands under normal loads, and when network failure causes BA targets violations.
Han Zhang 0009, Xingang Shi, Xia Yin 0001, Jilong Wang 0001, Yingya Guo, Tian Lan 0001
CoNEXT4
2021 Discovering obscure looking glass sites on the web to facilitate internet measurement research
abstract
Despite researchers have noticed that Looking Glass (LG) vantage points (VPs) are valuable for Internet measurement researches, they can only exploit VPs from well-known LG sites published on several LG portal pages. There should be a lot of LG sites that are not published in these portal pages, namely obscure LG sites, which are not easy to be found and exploited by researchers. In this paper, we design an efficient focused crawler to discover as many LG sites as possible which can avoid unnecessary resource consumption on analyzing irrelevant pages. Our designed focused crawler takes a similarity-guided search that exploits the well-developed search engines and comprehensively mines the common features shared by known LG sites to discover more LG pages. Moreover, the focused crawler takes a two-step PU learning classifier based on carefully selected LG features to efficiently discard irrelevant URLs, thus avoiding a lot of unnecessary resource consumption. As far as we know, we are the first to develop a method to discover obscure LG sites on the web. Experimental results show the effectiveness of our focused crawler. To facilitate practical applications, we further develop an automation tool, which can successfully retrieve 910 obscure automatable LG VPs from relevant pages obtained through our focused crawler. The 910 LG VPs significantly increase the geographic and network coverage of available VPs and we show their potential values in improving the completeness of AS-level Internet topology by a simple case study. Our method and the final VP list are beneficial to the measurement community.
Shuying Zhuang, Hui Wang 0011, Jilong Wang 0001, Zujiang Pan, Tianhao Wu 0010
CoNEXT3
2021 Hopping on Spectrum: Measuring and Boosting a Large-scale Dual-band Wireless Network
abstract
In recent years, more and more wireless networks support both 2.4GHz and 5GHz bands. However, in large-scale dual-band wireless networks, lack of understanding on the behavior and performance makes the network diagnosis and optimization extremely challenging. In this paper, we conduct a comprehensive measurement to characterize the behavior and performance in a large-scale dual-band wireless network (TD WLAN). We make several meaningful observations. (1) Although the 5GHz band outperforms the 2.4GHz band, 60% of devices tend to be associated with the 2.4GHz band. The device association behavior has a large impact on the performance. (2) Rogue and non-WiFi devices are prevalent, wherein hidden terminal interference increases the average loss rate by 8%, carrier sense interference increases the average WiFi latency by 45%, and RF interference further aggravates both packet loss and channel contention. (3) The dynamic channel assignment strategy is not always effective. On this basis, we propose a novel and easy-to-implement strategy to improve the wireless performance by intelligent band navigation and heuristic channel optimization. The actual deployment in TD WLAN shows the packet loss reduces by 40% on average and the WiFi latency for more than 60% of devices is below 5ms.
Haibo Wang 0004, Weizhen Dang, Jing'an Xue, Jiahao Cao 0001, Jilong Wang 0001
ICNP7
2021 A Distributed Hybrid Load Management Model for Anycast CDNs
abstract
Anycast content delivery networks rely on the underlying routing to schedule clients to their nearby service nodes, which however is not natively aware of server load or path latency. Requests burst from some regions may cause overload and hurt user experience. This scenario demands quickly adjusting clients to other nearby servers with available capacity. However, state-of-the-art solutions do not work well. On one hand, native routing-based scheduling is not flexible and precise enough, which may cause cascading damage and interrupt ongoing sessions. On the other hand, centralized algorithm is vulnerable and not responsive due to high complexity. We propose a practical distributed hybrid load management model to solve load burst problem. First, the hybrid mechanism leverages flexible DNS-based redirection, which can schedule at per-request granularity without interrupting ongoing sessions. Second, the distributed model is responsive by reducing computation overhead and theoretically guarantees to converge to the optimal solution. Based on the model, we further propose an cooperative and two heuristic distributed algorithms. At last, using a measurement dataset, we demonstrate their effectiveness and scalability, and illustrate how to adapt them to different scenarios.
Jing'an Xue, Haibo Wang 0004, Jilong Wang 0001, Tong Li 0014
MSN3
2021 Predicting Crowd Flows via Pyramid Dilated Deeper Spatial-temporal Network
abstract
Predicting crowd flows is crucial for urban planning, traffic management and public safety. However, predicting crowd flows is not trivial because of three challenges: 1) highly heterogeneous mobility data collected by various services; 2) complex spatio-temporal correlations of crowd flows, including multi-scale spatial correlations along with non-linear temporal correlations. 3) diversity in long-term temporal patterns. To tackle these challenges, we proposed an end-to-end architecture, called pyramid dilated spatial-temporal network (PDSTN), to effectively learn spatial-temporal representations of crowd flows with a novel attention mechanism. Specifically, PDSTN employs the ConvLSTM structure to identify complex features that capture spatial-temporal correlations simultaneously, and then stacks multiple ConvLSTM units for deeper feature extraction. For further improving the spatial learning ability, a pyramid dilated residual network is introduced by adopting several dilated residual ConvLSTM networks to extract multi-scale spatial information. In addition, a novel attention mechanism, which considers both long-term periodicity and the shift in periodicity, is designed to study diverse temporal patterns. Extensive experiments were conducted on three highly heterogeneous real-world mobility datasets to illustrate the effectiveness of PDSTN beyond the state-of-the-art methods. Moreover, PDSTN provides intuitive interpretation into the prediction.
Congcong Miao, Jiajun Fu, Jilong Wang 0001, Heng Yu 0005, Botao Yao, Anqi Zhong, Zekun He
WSDM3
2021 FedPA: An adaptively partial model aggregation strategy in Federated Learning
Juncai Liu, Hui Wang 0011, Chenghao Rong, Yuedong Xu 0001, Jilong Wang 0001
Comput. Networks6
2021 Octans: Optimal Placement of Service Function Chains in Many-Core Systems
abstract
Network Function Virtualization (NFV) offers service delivery flexibility and reduces overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance degradation due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2 percent lower throughput comparing to an optimal placement solution. In this article, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) Model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e., performance drop index, and propose a formal and accurate prediction model based on system level performance metrics. Finally, we propose two online algorithms to quickly find near-optimal placement solutions for one-time and incremental deployment. Extensive evaluation on a prototype implementation shows that Octans significantly improves the aggregate throughput comparing to two state-of-the-art placement solutions by 27.1 ~ 45.2 percent for one-time deployment and by 20.9 ~ 38.1 percent for incremental deployment, with very low prediction errors. Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap.
Heng Yu 0005, Zhilong Zheng, Junxian Shen, Congcong Miao, Chen Sun 0005, Hongxin Hu, Jun Bi, Jilong Wang 0001
IEEE Trans. Parallel Distributed Syst.9
2020 Serpens: A High-Performance Serverless Platform for NFV
abstract
Many enterprises run Network Function Virtualization (NFV) services on public clouds to relieve management burdens and reduce costs. However, NFV operators still face the burden of choosing the right types of virtual machines (VMs) for various network functions (NFs), as well as the cost of renting VMs at a granularity of months or years while many VMs remain idle during valley hours. A recent computing model named serverless computing automatically executes user-defined functions on requests arrival, and charges users based on the number of processed requests. For NFV operators, serverless computing has the potential of completely relieving NF management burden and significantly reducing costs. Nevertheless, naively exploring existing serverless platforms for NFV introduces significant performance overheads in three aspects, including high remote state access latency, long NF launching time, and high packet delivery latency between NFs. To address these problems, we propose Serpens, a high-performance serverless platform for NFV. Firstly, Serpens designs a novel state management mechanism to support local state access. Secondly, Serpens proposes an efficient NF execution model to provide fast NF launching and avoid extra packet delivery. We have implemented a prototype of Serpens. Evaluation results demonstrate that Serpens could significantly improve performance for NFs and service function chains (SFCs) comparing to existing serverless platforms.
Junxian Shen, Heng Yu 0005, Zhilong Zheng, Chen Sun 0005, Mingwei Xu 0001, Jilong Wang 0001
IWQoS6
2020 Graph Stochastic Neural Networks for Semi-supervised Learning
abstract
Graph Neural Networks (GNNs) have achieved remarkable performance in the task of the semi-supervised node classification. However, most existing models learn a deterministic classification function, which lack sufficient flexibility to explore better choices in the presence of kinds of imperfect observed data such as the scarce labeled nodes and noisy graph structure. To improve the rigidness and inflexibility of deterministic classification functions, this paper proposes a novel framework named Graph Stochastic Neural Networks (GSNN), which aims to model the uncertainty of the classification function by simultaneously learning a family of functions, i.e., a stochastic function. Specifically, we introduce a learnable graph neural network coupled with a high-dimensional latent variable to model the distribution of the classification function, and further adopt the amortised variational inference to approximate the intractable joint posterior for missing labels and the latent variable. By maximizing the lower-bound of the likelihood for observed node labels, the instantiated models can be trained in an end-to-end manner effectively. Extensive experiments on three real-world datasets show that GSNN achieves substantial performance gain in different scenarios compared with stat-of-the-art baselines.
Haibo Wang 0004, Chuan Zhou 0001, Jia Wu 0001, Shirui Pan, Jilong Wang 0001
NeurIPS6
2020 Predicting Human Mobility via Attentive Convolutional Network
abstract
Predicting human mobility is an important trajectory mining task for various applications, ranging from smart city planning to personalized recommendation system. While most of previous works adopt GPS tracking data to model human mobility, the recent fast-growing geo-tagged social media (GTSM) data brings new opportunities to this task. However, predicting human mobility on GTSM data is not trivial because of three challenges: 1) extreme data sparsity; 2) high order sequential patterns of human mobility and 3) evolving preference of users for tagging.
Congcong Miao, Ziyan Luo, Fengzhu Zeng, Jilong Wang 0001
WSDM4
2020 Understanding the latency to visit websites in China: An infrastructure perspective
Shuying Zhuang, Hui Wang 0011, Pei Zhang 0003, Jilong Wang 0001
Comput. Networks4
2020 Squeezing the Gap: An Empirical Study on DHCP Performance in a Large-Scale Wireless Network
abstract
Dynamic Host Configuration Protocol (DHCP) is widely used to dynamically assign IP addresses to users. However, due to little knowledge on the behavior and performance of DHCP, it is challenging to configure lease time and divide IP addresses for address pools properly in large-scale wireless networks. In this paper, we conduct the largest known measurement on the behavior and performance of DHCP in the wireless network of T University (TWLAN). We find the performance of DHCP is far from satisfactory: (1) The non-authenticated devices lead to a waste of 25% of addresses at the rush hour. (2) Address pool utilization varies greatly under the current address division strategy. (3) A device does not generate traffic for 67% of the lease time on average. Meanwhile, we observe devices of different locations and operating systems show diverse online patterns. A unified lease time setting could result in an inefficient usage of addresses. To address the problems, taking account of authentication information and online patterns, we propose a new leasing strategy. The results show it outperforms three state-of-the-art baselines and reduces the number of assigned addresses by 24% and the average total lease time by 17% without significantly increasing the DHCP server load. Besides, we further propose an adaptive address division strategy to balance the address utilization of pools, which can be deployed in parallel with the new leasing strategy and reduce the risk of address exhaustion.
Haibo Wang 0004, Hui Wang 0011, Jilong Wang 0001, Weizhen Dang, Jing'an Xue, Jinzhe Shan
IEEE/ACM Trans. Netw.3
2019 Buffet: Enabling Multi-Tenant Network Functions
abstract
Many enterprises outsource traffic processing to third- party Network Function (NF) service providers to relieve management burden and reduce cost. NF providers have to process packets from multiple tenants simultaneously. However, most existing software based NFs are designed for one single tenant without internal state isolation mechanisms. These NFs cannot be securely shared across multiple tenants. Existing solutions that support multitenancy are either inefficient or ad-hoc for specific NFs. In this paper, we propose Buffet, a general and efficient framework that enables multitenancy for a wide range of NFs. First, Buffet introduces a general programming abstraction for various NFs to relieve NF developers from considering isolation details. Second, Buffet proposes dynamic tenant-level affinity to achieve high performance and resource efficiency. Finally, Buffet exploits SmartNIC offloading to eliminate host CPU overhead. We have implemented a prototype of Buffet. Evaluation results demonstrate that Buffet can effectively enable multitenancy for a wide range of NFs with high performance and resource efficiency.
Heng Yu 0005, Junxian Shen, Chen Sun 0005, Zhilong Zheng, Jilong Wang 0001
GLOBECOM5
2019 WRT: Constructing Users' Web Request Trees from HTTP Header Logs
abstract
As web traffic has already dominated Internet, massive web logs are being generated ceaselessly. It is essential and meaningful for operators to mine valuable information and knowledge from the log data. However, the state-less feature of HTTP and increasing dynamics and complexities of web services bring a challenge to web mining in web logs. To solve the problem, in this paper, we introduce Web Request Tree (WRT) to reconstruct users' web request behaviors from web logs with HTTP header information, which can be applied in many fields. Compared to previous related work, our method pays special attentions to modern technologies to handle cases caused by these technologies such as PJAX. We evaluate the feasibility of our method with a measurement study on referrer policies of Alexa top websites and results show that our method can achieve a high accuracy for most websites. We also conduct experiments on a real-world dataset collected from official websites of a top university in China. We find that a lot of abnormal requests exist in log data by analyzing WRT and WRT is rich in information that is valuable for operators to analyze user behaviors, detect anomalies, optimize performance and so on.
Shengchao Liu, Jilong Wang 0001, Hui Wang 0011, Haibo Wang 0004
ICC2
2019 BDAC: A Behavior-aware Dynamic Adaptive Configuration on DHCP in Wireless LANs
abstract
DHCP is widely used to dynamically allocate IP addresses to the devices on local area networks, but the explosive increases of WiFi devices and their frequent mobility pose great challenges on DHCP performance in wireless LANs. In this paper, by analyzing large scale real network traces, we observe that the dynamic WiFi user behavior (e.g., online time pattern and spatio-temporal mobility pattern) leads to the poor DHCP performance. The IP pools in some VLANs have been exhausted in rush hours although the total IP utilization in WLAN is only 24%. Therefore, we have to configure IP lease times and IP pools dynamically and make sure that they are adaptive to the WiFi user behavior. In order to achieve this goal, we characterize and model the user behavior across online time pattern and spatiotemporal mobility pattern. Then we propose BDAC, a behaviour-aware dynamic adaptive configuration, which is combined of two strategies: adaptive IP lease time configuration and dynamic IP pool configuration. The former is to set adaptive lease times across user roles and area types based on online time pattern to reclaim IP addresses in time and reduce the peak IP usage, while the latter dynamically migrates the IP addresses across VLANs based on spatio-temporal mobility correlation to save the IP addresses. Using the real network traces of a different week, we conduct experiments to evaluate the performance of BDAC. Results show that BDAC can save up to 60% of IP addresses and the actual IP utilization rises from 24% to 59%. Furthermore, BDAC maintains high IP utilization when the number of VLANs in a WLAN increases.
Congcong Miao, Jilong Wang 0001, Tianying Ji, Hui Wang 0011, Chao Xu 0015, Fengyuan Ren
ICNP2
2019 Evaluating performance and inefficient routing of an anycast CDN
abstract
Anycast has been increasingly deployed for content delivery networks to map clients to their nearby replicas, which relies on the underlying routing. However, the simplicity of operation comes at cost of less precise client-mapping control. Although many works have measured anycast DNS, anycast CDNs, with different service goals and engineering, are still not fully understood. In this paper, we design novel methods and combine large-scale traceroute and HTTP measurement to evaluate the overall client-proximity and inefficient routing of the largest anycast CDN, Cloudflare. We find that 90% paths traverse only 2-4 ASes, which highlights its direct networks providers. By further identifying and characterizing direct providers at finer granularity of facilities, we quantitatively shows that Cloudflare unevenly uses few large transit providers to delivery the majority of contents. Inspired by the observations, we propose an anycast routing pathology and diagnosis methodology. Investigation reveals that few huge providers have outsized impact in that they are not only related to many inter-domain inflations, but also have path inflation inside their own networks, thus deserving priority focus when troubleshooting.
Jing'an Xue, Weizhen Dang, Haibo Wang 0004, Jilong Wang 0001, Hui Wang 0011
IWQoS4
2019 A survey on resource scheduling for data transfers in inter-datacenter WANs
Hui Wang 0011, Jilong Wang 0001, Changqing An, Qianli Zhang
Comput. Networks2
2019 An Adaptive Online Scheme for Scheduling and Resource Enforcement in Storm
abstract
As more and more applications need to analyze unbounded data streams in a real-time manner, data stream processing platforms, such as Storm, have drawn the attention of many researchers, especially the scheduling problem. However, there are still many challenges unnoticed or unsolved. In this paper, we propose and implement an adaptive online scheme to solve three important challenges of scheduling. First, how to make a scaling decision in a real-time manner to handle the fluctuant load without congestion? Second, how to minimize the number of affected workers during rescheduling while satisfying the resource demand of each instance? We also point out that the stateful instances should not be placed on the same worker with stateless instances. Third, currently, the application performance cannot be guaranteed because of resource contention even if the computation platform implements an optimal scheduling algorithm. In this paper, we realize resource isolation using Cgroup, and then the performance interference caused by resource contention is mitigated. We implement our scheduling scheme and plug it into Storm, and our experiments demonstrate in some respects our scheme achieves better performance than the state-of-the-art solutions.
Shengchao Liu, Jianping Weng, Hui Wang 0011, Changqing An, Yipeng Zhou, Jilong Wang 0001
IEEE/ACM Trans. Netw.6
2018 Deep Structure Learning for Fraud Detection
abstract
Fraud detection is of great importance because fraudulent behaviors may mislead consumers or bring huge losses to enterprises. Due to the lockstep feature of fraudulent behaviors, fraud detection problem can be viewed as finding suspicious dense blocks in the attributed bipartite graph. In reality, existing attribute-based methods are not adversarially robust, because fraudsters can take some camouflage actions to cover their behavior attributes as normal. More importantly, existing structural information based methods only consider shallow topology structure, making their effectiveness sensitive to the density of suspicious blocks. In this paper, we propose a novel deep structure learning model named DeepFD to differentiate normal users and suspicious users. DeepFD can preserve the non-linear graph structure and user behavior information simultaneously. Experimental results on different types of datasets demonstrate that DeepFD outperforms the state-of-the-art baselines.
Haibo Wang 0004, Chuan Zhou 0001, Jia Wu 0001, Weizhen Dang, Xingquan Zhu 0001, Jilong Wang 0001
ICDM6
2018 Squeezing the Gap: An Empirical Study on DHCP Performance in a Large-scale Wireless Network
abstract
Dynamic Host Configuration Protocol (DHCP) is widely used to dynamically assign IP addresses. However, due to little knowledge on the behavior and performance of DHCP, it is challenging to configure a proper lease time in complicated wireless network. In this paper, we conduct the largest known measurement on the behavior and performance of DHCP based on the wireless network of T University (TWLAN). TWLAN has more than 59,000 users, 10,000 wireless access points and 130,000 IP addresses. We find the performance of DHCP is far from satisfactory: (1) The non-authenticated devices lead to a waste of 25% of IP addresses at the rush hour. (2) A device does not generate traffic for 67 % of the lease time on average. Meanwhile, we find devices of different locations and operating systems show diverse online patterns. A unified lease time setting could result in an inefficient utilization of addresses. To address the problems, taking account of authentication information and device online patterns, we propose a new leasing strategy. The results show it reduces the number of assigned addresses by 24 % and reduces the time during which an IP address is occupied by 17 % without sianificantly Increasing the DHCP server load.
Haibo Wang 0004, Jilong Wang 0001, Weizhen Dang, Jing'an Xue
INFOCOM2
2018 A Multi-dimension Measurement Study of a Large Scale Campus WiFi Network
abstract
The growing trend of wireless devices and WiFi networks poses significant management challenges to network administrators. Characterizing WiFi user behavior and understanding WiFi network usage pattern are helpful to identify the management challenges so that network administrators could manage WiFi networks more efficiently. In this work, we collect comprehensive datasets, i.e., DHCP dataset, AAA dataset, SNMP dataset of ACs in a large campus WiFi network. We provide a detailed measurement study from multiple dimensions, i.e., server plane, temporal plane, spatial plane and traffic plane. We observe that the WiFi network under study is far from optimal. First, the phenomenon of IP waste is severe due to the isolation between DHCP server and AAA server. Second, current deployment of network infrastructure resources is based on network administrators' experience and it results in that the WiFi performance varies a lot across different areas. Furthermore, we also study the user behavior with different types of devices and in different kinds of buildings. Our observations indicate that the WiFi network could be improved and managed more efficiently from multiple dimensions. We believe that this measurement study is helpful for network administrators and researchers to understand more about large scale WiFi networks.
Congcong Miao, Jilong Wang 0001, Hui Wang 0011, Jun Zhang 0004, Shengchao Liu
LCN2
2018 RNN-SM: Fast Steganalysis of VoIP Streams Using Recurrent Neural Network
abstract
Quantization index modulation (QIM) steganography makes it possible to hide secret information in voice-over IP (VoIP) streams, which could be utilized by unauthorized entities to set up covert channels for malicious purposes. Detecting short QIM steganography samples, as is required by real circumstances, remains an unsolved challenge. In this paper, we propose an effective online steganalysis method to detect QIM steganography. We find four strong codeword correlation patterns in VoIP streams, which will be distorted after embedding with hidden data. To extract those correlation features, we propose the codeword correlation model, which is based on recurrent neural network (RNN). Furthermore, we propose the feature classification model to classify those correlation features into cover speech and stego speech categories. The whole RNN-based steganalysis model (RNN-SM) is trained in a supervised learning framework. Experiments show that on full embedding rate samples, RNN-SM is of high detection accuracy, which remains over 90% even when the sample is as short as 0.1 s, and is significantly higher than other state-of-the-art methods. For the challenging task of conducting steganalysis towards low embedding rate samples, RNN-SM also achieves a high accuracy. The average testing time for each sample is below 0.15% of sample length. These clues show that RNN-SM meets the short sample detection demand and is a state-of-the-art algorithm for online VoIP steganalysis.
Zinan Lin 0001, Yongfeng Huang 0001, Jilong Wang 0001
IEEE Trans. Inf. Forensics Secur.3
2017 Impact of Development and Governance Factors on IPv4 Address Ownership
abstract
The worldwide distribution of IPv4 addresses is highly non-uniform. The factors that contribute to IPv4 address resources owned by a country remain unknown. In this paper, we systematically study the relationship between the number of allocated IPv4 addresses and 30 factors extracted from political, economic, ecological, social, and cultural dimensions. We observe that GDP has the largest influence on the number of IPv4 addresses while the influence of political factors is very weak. To further quantify the complex relationship among them, we fit data into a Decision Tree model. Based on the pruned model, we can extract precise rules that contribute to IPv4 address resources owned by a country. To the best of our knowledge, this is the first work to evaluate IPv4 address ownership by quantitative factors. Our work is helpful to better understand the network development of a country.
Haibo Wang 0004, Jilong Wang 0001, Jing'an Xue
LCN2
2016 Understanding the Impact of AP Density on WiFi Performance Through Real-World Deployment
abstract
802.11 (WiFi) networks have become increasingly important for our daily lives. However, previous work has shown that enterprise WiFi performance is often unsatisfactory and that over-utilization and interference from rogue APs are the two primary reasons. To address the above problem, this paper proposes to improve the capacity of WiFi infrastructures by increasing the enterprise AP deployment density, as well as disabling the wired Internet access in buildings to eliminate rogue APs and their interference. We deployed several WiFi networks with different AP density and vendors on Tsinghua campus. Based on the measurement results from our real-world deployments, we made three main observations: 1) in general, higher AP density improves WiFi performance; 2) over-dense deployment with unnecessarily high transmission power can worsen WiFi performance. 3) choice of AP vendors also has an impact on WiFi performance.
Kaixin Sui, Yousef Azzabi, Xiaoping Zhang 0004, Youjian Zhao, Jilong Wang 0001, Zimu Li, Dan Pei
LANMAN6
2014 UDP traffic classification using most distinguished port
abstract
Comparing to TCP traffic, the composition of UDP traffic is still unclear. Although it is observed that a large fraction of UDP traffic appears to be P2P applications, application level classification of UDP traffic is still very hard since most of these applications are private protocols based. In this paper, a novel method is proposed to classify UDP traffic. Based on the assumption that traffic from two communicating half-tuples identified by theis from the same application, all half-tuples can be grouped into several connected subgraphs. The port numbers which are adopted by most links or half-tuples in each subgroup can thus be used to characterize the application types of the whole subgroup. Experiment results show that this approach is feasible and can classify UDP traffic only using flow level information. The port numbers adopted by most links or half-tuples are surprisingly stable among different time periods, for example, for Youku application remain the same for more than 90% of periods in all the 1429 periods.
Qianli Zhang, Jilong Wang 0001, Xing Li 0001
APNOMS3
2011 Rumor Riding: Anonymizing Unstructured Peer-to-Peer Systems
abstract
Although anonymizing Peer-to-Peer (P2P) systems often incurs extra traffic costs, many systems try to mask the identities of their users for privacy considerations. Existing anonymity approaches are mainly path-based: peers have to pre-construct an anonymous path before transmission. The overhead of maintaining and updating such paths is significantly high. We propose Rumor Riding (RR), a lightweight and non-path-based mutual anonymity protocol for decentralized P2P systems. Employing a random walk mechanism, RR takes advantage of lower overhead by mainly using the symmetric cryptographic algorithm. We conduct comprehensive trace-driven simulations to evaluate the effectiveness and efficiency of this design, and compare it with previous approaches. We also introduce some early experiences on RR implementations.
Yunhao Liu 0001, Jinsong Han, Jilong Wang 0001
IEEE Trans. Parallel Distributed Syst.3
2008 DRAGON-Lab - Next generation internet technology experiment platform
Jilong Wang 0001, ZhongHui Li, Guohan Lu, Caiping Jiang, Xing Li 0001, Qianli Zhang
Sci. China Ser. F Inf. Sci.1
2007 Internet Management Network
Jilong Wang 0001, Jiahai Yang 0001
APNOMS1
2007 On the Design of Fast Prefix-Preserving IP Address Anonymization Scheme
Qianli Zhang, Jilong Wang 0001, Xing Li 0001
ICICS2
2005 Traffic Measurement and Analysis of TUNET
abstract
Traffic measurement and analysis, as one of the important methods of understanding and characterizing network, can provide significant support for network management. After a brief introduction of a novel NP (network processor)-based architecture of traffic measurement, the paper presents the detailed analysis results of the traffic collected from the gigabit link connecting Tsinghua University campus network (in short, TUNET) to its upstream ISP, China Education and Research NETwork (in short, CERNET). Then, the paper comprehensively analyzes the traffic from multi-dimension viewpoints, including temporal distribution, packet length distribution, port-based distribution, protocol-based distribution, and TopN statistics. Such analysis not only provides support for the study of user behavior, but also enriches traffic measurement technology
Jun Zhang 0004, Jiahai Yang 0001, Changqing An, Jilong Wang 0001
CW4