VLDB 2026 Research / reviewers in the wild / expert
Han Zhang 0009
dblp:26/4189-9
· DBLP profile ↗
69ranked-venue papers
12as first author
52since 2021 · last 2026
0000-0003-4429-9959ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 46 · 11 first-author · 30 since 2021Security and privacy · 12 · 12 since 2021Systems, architecture and hardware · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Comprehensive Network Configuration Verification via Effective Environment Reduction
Han Zhang 0009, Renrui Tian, Xia Yin 0001, Xingang Shi, Gang Ren 0003, Jilong Wang 0001, Jiangyuan Yao |
INFOCOM | 3 |
| 2026 | Towards High-Performance Intrusion Detection with Robustness Guarantees on Programmable Switches at ISP ScaleabstractIn order to provide security connections to the enterprise campus sites, internet service providers are offering comprehensive intrusion detection services at the network layer. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for high-speed network protection, especially for encrypted traffic analysis. In this paper, we design and implement SiteGuard, an inline network intrusion detection system with programmable switches specifically developed to protect enterprise campus sites connecting to ISP. SiteGuard proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. SiteGuard also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify malicious traffic. In addition, SiteGuard introduces an online update mechanism that aims to dynamically adjust the detection model in response to environmental changes. SiteGuard has been in production for more than three years. Our production and testbed evaluations demonstrate SiteGuard can detect malicious traffic with approximately 90% accuracy in minutes. Han Zhang 0009, Linqiang Qian, Guyue Liu, Kaiyang Zhao 0004, Yantu Tong, Zeji Xiao, Dongbiao He, Ke Ruan, Jilong Wang 0001, Xia Yin 0001 |
SIGCOMM | 1 |
| 2026 | GlassMiner: Mining Looking Glass Services via Structure-Semantics Fusion for Web Observability
Yunze Wei, Xingang Shi, Han Zhang 0009, Xia Yin 0001 |
WWW | 3 |
| 2026 | A comprehensive survey on encrypted network traffic classification
Shangbin Han, Han Zhang 0009, Mengmeng Lu, Sifang Guo, Boyuan Tian, Jilong Wang 0001 |
Comput. Networks | 2 |
| 2026 | Glint: Localization of Gray Violations in Untrusted and Unreliable SRv6 NetworksabstractIn the Segment Routing over IPv6 (SRv6) network, a wide range of network events (e.g., attacks, intrusions, violations, malicious route announcements) may occur. Network management requires real-time monitoring of untrusted and unreliable environments (e.g., unsafe components and devices). Early localization of abnormal links causing violations in the SRv6 network helps minimize the compensation required for service unavailability. However, the overhead of the state-of-the-art methods does not scale efficiently to large-scale SRv6 networks and exhibit poor robustness to addressing various disturbances from unreliable networks. To cope with these challenges, we propose Glint, an in-band network telemetry framework to localize abnormal links in SRv6 networks. The key idea of Glint is sampling part of the information while the overall information is known. Glint provides probabilistic in-band collection to gather segment-level telemetry data, reducing overhead and improving efficiency. Glint also proposes distributed verification-based detection to enhance the trustworthiness of security assessments, further improving robustness against disturbances. In addition, we design selective telemetry that reduces telemetry reports while preserving security-relevant visibility. Our evaluations demonstrate that, compared to the state-of-the-art frameworks, Glint significantly reduces header bandwidth overhead by 75.6% and memory overhead by 48.7% while reducing false positives. We also implement Glint on the Intel Tofino switch, achieving over a 50% reduction in hardware resource consumption compared to existing methods. Kaiyang Zhao 0004, Han Zhang 0009, Xingang Shi, Xia Yin 0001, Jiankun Hu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Interference-Aware UAV Path Planning on Grid SINR Maps With Event-Triggered UpdatesabstractUrban unmanned aerial vehicle (UAV) navigation operates under tight bandwidth, compute, and latency budgets. Interference and building blockage cause rapid link fluctuations, making interference-aware path planning on grid signal-to-interference-plus-noise ratio (SINR) maps essential for reliable communication. In this paper, we formulate the problem as a joint update–navigation decision: map uncertainty triggers on-demand partial refreshes that are co-optimized with motion under a bandwidth budget. We introduce UT-Grid, an uncertainty-triggered grid-update framework, that refreshes conditionally triggered upon necessity to reduce overall map-update traffic, and MoE-D3QN, a Dueling Double DQN with sparse Mixture-of-Experts (Top-1 routing) that activates a single expert per step, cutting per-decision active parameters and FLOPs while matching or surpassing comparable dense D3QN planners. In an urban simulation with multi-source interference, the framework outperforms static-map, periodic-refresh, and dense D3QN baselines, increasing reaching probability and path efficiency while markedly reducing communication overhead. Lantu Guo, Mengchen Yao, Han Zhang 0009, Weiqing Mu, Yun Lin 0005 |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | Clover: Workload Verification for Real-Time Detection of Contention-Induced Slowdowns in Serverless PlatformsabstractServerless computing, or Function-as-a-Service, continues to gain popularity due to its pay-as-you-go billing model, flexibility, and cost efficiency. However, these same features introduce significant security risks, such as the Denial-of-Wallet (DoW) attack. In this paper, we conduct real-world DoW attacks on commercial serverless platforms to evaluate their severity. To detect such attacks, we design, implement, and evaluate Clover, an accurate and user-friendly DoW detection system with negligible performance overhead. Clover addresses information ambiguity in serverless environments by deploying a request-oriented metric collection agent. At its core, Clover proposes a workload verification approach to bridge performance metrics and execution duration. Specifically, Clover uses a multivariate linear model to learn the benign relationship between metrics and execution duration, effectively characterizing normal workload behavior. It then continuously monitors runtime workloads by calculating their Mahalanobis distance from this learned benign model. Deviations identified through this distance indicate potential DoW attacks. Implemented as a practical system, Clover introduces performance overhead of less than 3.2%, maintains an average model execution time of only 0.84 microseconds, and achieves an accuracy of 92.7% under the most challenging scenario. Junxian Shen, Han Zhang 0009, Weiwei Lin 0001, Yantao Geng, Jilong Wang 0001, Mingwei Xu 0001 |
IEEE Trans. Netw. | 2 |
| 2026 | Efficient Slice-Parallel Distributed Probing in SRv6 NetworksabstractSegment Routing over IPv6 (SRv6) is widely deployed, where operators construct numerous parallel SR-based network slices. An accurate diagnosis of latency bottlenecks in SRv6 tunnels is essential to maintain service-level objectives. However, building a distributed, at-scale probing system is non-trivial: SRv6 priority policies, SR-aware multipath forwarding, and slice isolation collectively invalidate assumptions made by existing diagnostic methods. In this paper, we present SRmesh, a distributed system for diagnosing latency bottleneck links in SRv6 networks. First, we adopt distributed SR-based probing agents to control routing paths and ensure that probe packets emulate per-slice production traffic. Second, we employ a latency-based multipath inference that runs in parallel across slices to resolve SR-induced routing ambiguities. Third, we introduce a sliceparallel progressive diagnosis that incrementally reuses probe results to reduce redundant measurements, optimizing diagnostic overhead for large-scale SRv6 overlay networks. We implement a prototype of SRmesh and conduct extensive evaluations on 247 real network topologies. The results indicate that SRmesh achieves high diagnostic accuracy with a 93.4% reduction in probe overhead, demonstrating its practicality and scalability in large-scale SRv6 environments. Kaiyang Zhao 0004, Han Zhang 0009, Xingang Shi, Xia Yin 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | Poster: ERIS: Evaluating ROV via ICMPv6 Rate Limiting Side ChannelsabstractThe Resource Public Key Infrastructure (RPKI) plays a crucial role in securing BGP against prefix hijacking by enabling Route Origin Validation (ROV). However, the limited adoption of ROV in the real world undermines the effectiveness of RPKI. Hence, measuring ROV deployment in practice is essential for assessing the impact of RPKI. Existing measurement efforts either suffer from limited coverage and accuracy due to reliance on control-plane data, or require controlled IP prefixes or large-scale deployment of vantage points. Furthermore, most studies focused on IPv4, leaving ROV status in IPv6 largely underexplored. Renrui Tian, Han Zhang 0009, Xia Yin 0001, Xingang Shi, Jilong Wang 0001 |
CCS | 3 |
| 2025 | Log-Based Anomaly Detection with Multi-level Progressive Temporal-Semantic FusionabstractLog anomaly detection plays a pivotal role in ensuring system stability and security, particularly in largescale environments characterized by the generation of log data at exceptionally high volumes and velocities. Conventional approaches often struggle to effectively filter log information and fully leverage temporal dynamics, resulting in challenges such as information loss, semantic drift, and heightened computational overhead. To overcome these limitations, we present QYXLAD, an innovative log anomaly detection framework. QYXLAD enhances log representation accuracy by seamlessly integrating temporal and semantic information. It introduces a MPMM(Multi-level Progressive Masking Mechanism)-based Feature Fusion designed to capture temporal dependencies and semantic features across diverse pattern combinations, thereby significantly improving the sensitivity and precision of anomaly detection. Furthermore, QYXLAD utilizes a Mamba-based classifier for anomaly identification. Comprehensive theoretical analysis and empirical evaluations demonstrate that QYXLAD achieves state-of-the-art performance on multiple public log datasets, surpassing existing methods in key metrics such as precision, recall, and F1-score. These results underscore the framework’s efficacy and superiority in addressing log anomaly detection challenges. Zhiyu Wen, Pei Zhang 0003, Yanxu Fu, Xiaohong Huang 0003, Yan Ma 0003, Han Zhang 0009 |
ICCCN | 7 |
| 2025 | SRmesh: Deterministic and Efficient Diagnosis of Latency Bottleneck Links in SRv6 NetworksabstractSegment Routing over IPv6 (SRv6) has attracted more attention from network operators. Diagnosing performance bottlenecks for SRv6 tunnels is critical to maintaining network quality. However, SRv6 introduces priority policies, special forms of multipath routing, and SR-based network slicing, all of which make existing methods difficult to apply. In this paper, we present SRmesh, a framework for diagnosing latency bottleneck links specifically tailored for SRv6 tunnel performance analysis. First, we adopt SR-based probing to deterministically control routing paths and ensure that probe packets emulate real production traffic. Second, we employ a latency-based multipath inference to resolve routing ambiguities caused by SR. Third, we introduce a topology-independent progressive diagnosis that incrementally reuses probe results to reduce redundant measurements, optimizing diagnostic overhead for large-scale SRv6 overlay networks. We implement a prototype of SRmesh and conduct extensive evaluations on real network topologies. The results indicate that SRmesh achieves high diagnostic accuracy with up to a 91.9% reduction in probe overhead, demonstrating its practicality and scalability in large-scale SRv6 environments. Kaiyang Zhao 0004, Han Zhang 0009, Xingang Shi, Xia Yin 0001 |
ICNP | 2 |
| 2025 | On Non-Commutative RoutingabstractThe complexity of routing requirements leads to increasingly intricate routing metrics. Existing routing algebra theories have demonstrated that convergent and optimal routing algorithms can be designed only when path metrics satisfy certain properties such as monotonicity and isotonicity. Furthermore, some non-isotonic metrics can be converted into isotonic forms on partial orders through reduction. However, practical scenarios often involve non-commutative algebraic properties, which are overlooked by existing theories. For these problems, there lacks a unified framework to study their solvability, a systematic method for their reduction, and an efficient algorithm to compute optimal routes. In this work, we extend routing algebra to accommodate non-commutative routing problems, propose general reduction methods for them, and explore their solvability. In addition, we design a link state algorithm that converge fast on a reduced partial order. All these discussions are supported by concrete examples, theoretical proofs, and simulations on various network topologies. Zhaozhen Wang, Xingang Shi, Haijun Geng, Zitong Jin, Han Zhang 0009, Xia Yin 0001 |
INFOCOM | 5 |
| 2025 | Probabilistic Analysis of Overload-Free Property for Critical TrafficabstractNetwork structures are sophisticated and hence vulnerable to errors. Link failures and traffic load fluctuations lead to complexity in network states. Different failure scenarios can result in varying network traffic distribution patterns. Meanwhile, the load on links within the same failure scenario dynamically changes with fluctuations in traffic. Network administrators are particularly concerned about whether links along the paths traversed by critical traffic are overload-free guaranteed when link failures occur. Yet, no attention was ever paid to overload-free property analysis for critical traffic. We propose Offaela, an efficient and accurate probabilistic analysis framework that verifies an overload-free property for critical traffic. We prudently formulate the problem and prove its computational hardness, then storm this fortification by proffering a failure scenario merging algorithm and adopting a randomized approximation method. Evaluations on real networks show that Offaela outperforms the state-of-the-art solution by 4.83 x and can provide availability analysis assistance such as identifying vulnerable failure scenarios. Zhiyun Tang, Ke Ruan, Yingjun Ye, Jilong Wang 0001, Xia Yin 0001, Xingang Shi, Han Zhang 0009 |
IWQoS | 9 |
| 2025 | Achieving High-Speed and Robust Encrypted Traffic Anomaly Detection with Programmable SwitchesabstractAttacks against data centers are becoming more common as a result of the fast expansion of applications. In order to keep pace with the growing amount of data centers connected to their networks, internet service providers must offer comprehensive security services. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for the high-speed encrypted network traffic. In this paper, we design and implement Mazu, an inline network intrusion detection system with programmable switches specifically developed to protect data centers connecting to the internet service provider. Mazu proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. Mazu also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify the malicious traffic. In addition, Mazu introduces an online update mechanism aimed at dynamically adjusting the detection model in response to environmental changes. Mazu has been in production for two years, during which time it has identified over 10 critical attack events and protect more than 10 million servers for two ISPs. Our production and testbed evaluations demonstrate that Mazu can detect malicious traffic entering the data center sites with approximately 90% accuracy within minutes. Han Zhang 0009, Guyue Liu, Xingang Shi, Dongbiao He, Jilong Wang 0001, Ke Ruan, Xia Yin 0001 |
SIGCOMM | 1 |
| 2025 | Low-Overhead Distributed Application Observation with DeepTrace: Achieving Accurate Tracing in Production SystemsabstractAs microservices grow in scale and complexity, their operation and debugging become increasingly challenging. Even a single user request can involve interactions across hundreds of components. In such intricate systems, distributed tracing, which tracks the end-to-end execution flow of requests, has become a critical monitoring tool. Among these, non-intrusive tracing frameworks that do not require code modification are particularly valued for their convenience. However, existing non-intrusive solutions either have limited applicability or lack sufficient accuracy under high concurrency. To address these challenges, we propose DeepTrace, a transaction-based, non-intrusive distributed tracing framework designed for microservices. DeepTrace leverages API endpoints and transaction fields embedded within request content to categorize requests into distinct transactions, thereby reducing the likelihood of incorrectly merging traces from different transactions. Compared to state-of-the-art frameworks, DeepTrace maintains an accuracy rate of over 95% even under high concurrency. It has also been adopted by dozens of companies in their production systems for tasks such as failure diagnosis and resource optimization. Yantao Geng, Han Zhang 0009, Jilong Wang 0001, Xia Yin 0001 |
SIGCOMM | 2 |
| 2025 | ACME++: A Secure Authorization Mechanism for ACME Clients in the Web PKI EcosystemabstractThe Automatic Certificate Management Environment (ACME) protocol automates the issuance and renewal of secure socket layer certificates, simplifying the management of large-scale certificate deployments. To reduce the load on Certificate Authority (CA) servers, ACME employs a caching mechanism that stores domain validation (DV) results for 30 days. However, this mechanism allows attackers to reuse previously authorized results, potentially bypassing the DV process. In this paper, we introduce the ACME Authz Cache Attack, whereby an attacker can obtain fraudulent certificates without domain control. We demonstrate that even the prominent CA, Let's Encrypt, is vulnerable to this attack. To mitigate this, we propose ACME++, an enhanced protocol that binds the client's IP address and a unique identifier to the ACME account, ensuring secure authorization for each new client and effectively preventing the ACME Authz Cache Attack. Our implementation of ACME++ shows that it introduces little overhead to the CA server. Han Zhang 0009, Yunze Wei, Xingang Shi, Jilong Wang 0001, Xia Yin 0001 |
WWW | 2 |
| 2025 | Adaptive traffic engineering with segment routing through deep reinforcement learning
Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009 |
Comput. Networks | 6 |
| 2025 | End-to-End Attack Scene Reconstruction in a Host With Rules and Anomaly-Based Detection ModelsabstractCritical devices on the Internet are frequently targeted by skilled and advanced network attackers. These attackers often orchestrate complex and persistent intrusion campaigns, which involve multiple stages of attacks. In the context of host-based threat detection, the reconstruction of the entire attack scenario is crucial for tracing threats and fixing system vulnerabilities. Prior anomaly-based studies lack the capability to interpret the attack scenario, while rule-based approaches struggle with detecting novel attack patterns. We introduce eaGle, an end-to-end framework that takes original host-based data as input and reconstructs the potential attack scenario as output. It leverages an anomaly-based algorithm and a fine-grained misuse detection module to assign anomalous scores to host data, constructs the potential attack scenario using a novel anomalous subtree detection algorithm, and generates the interpretable attack scenario graph through a coarse-grained rule matching method. We assess the performance of eaGle using three attack scenarios from the DARPA TC dataset and three deployment scenarios. The results demonstrate that eaGle can effectively uncover the hidden attack scenario within the host data and outperforms three state-of-the-art attack scenario reconstruction systems. Xia Yin 0001, Han Zhang 0009, Xingang Shi, Jiahai Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | MFFGCN: Multimodal Feature Fusion Graph Convolution Network for Radio Map Estimation With Uneven Spatial SamplingabstractRadio map estimation (RME) is a crucial method for analyzing spectrum space utilization and network coverage, serving as an essential tool for the mobile communication. However, physical constraints, security, privacy, and other issues often render some areas inaccessible, resulting in extremely sparse and unevenly distributed measurement data. To address these challenges, we propose a multimodal feature fusion graph convolution network (MFFGCN). The model incorporates a dual-encoder architecture with an adaptive multi-feature fusion module to exploit environmental information and learn the shadowing effects of radio-signal propagation. We then convert the coarse estimation into regional feature patches and construct a graph over these patches. A graph neural network aggregates contextual information among them, thereby alleviating the impact of uneven spatial sampling. Extensive experiments on open datasets demonstrate that our method achieves state-of-the-art performance, effectively reducing the effects of uneven sampling. Han Zhang 0009, Yu Han 0003, Lingxin Meng, Guan Gui 0001, Wei Xiang 0001, Yun Lin 0005 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | HELA: Inferring AS Relationships With a Hybrid of Empirical and Learning AlgorithmsabstractKnowledge of the business relationships between Autonomous Systems (ASes) is the basis for studying many aspects of the Internet. Despite the significant progress achieved by the latest inference algorithms, their inference results still suffer from errors on many special or critical links, thus hindering many relationship-related applications. We take an in-depth analysis on the challenges inherent in inferring AS relationships, including complex routing policies, limited and biased vantage point (VP) coverage, as well as a lack of accurate validation data. To address these challenges, we introduce HELA, a framework for inferring AS relationships with a hybrid of empirical and machine learning algorithms. HELA incorporates an array of grouping, voting, and machine learning algorithms and allows flexible substitution of each. We systematically evaluate various combinations of them to determine the most effective one for HELA. Furthermore, we describe the collection of varied validation datasets, including BGP community and RPSL records from Internet Routing Registries (IRRs), as well as OneStep community. Our up-to-date dataset corrects errors in previous published validation sets, contains 95% more labelled links, and exhibits a closer alignment to actual link distribution. Using routing data and validation datasets composed for each month during$2021\sim 2023$, we access HELA’s superiority in both inference accuracy and stability compared to the state-of-the-art inference algorithms, i.e., AS-Rank, ProbLink, and HELA’s predecessor TopoScope. In particular, HELA achieves up to$2.9\times $reduction on error rates across overall datasets, up to$2.6\times $reduction with a$4.5\times $decrease on standard deviation on incomplete and biased datasets, and up to$1.7\times $reduction on various sources of validation datasets. Xingang Shi, Zitong Jin, Bin Xiong, Xinyao Huang, Xiaotian Xi, Han Zhang 0009, Xia Yin 0001 |
IEEE Trans. Netw. | 7 |
| 2025 | Centralized Network Utility Maximization With Accelerated Gradient MethodabstractNetwork utility maximization (NUM) is a fundamental problem for network traffic management and resource allocation. Due to the inherent decentralization and complexity of networks, much of the existing research has focused on developing decentralized algorithms for NUM. However, with the rise of Software-Defined Networking (SDN), especially in cloud networks and inter-datacenter networks managed by large enterprises, there has been growing interest in centralized NUM algorithms. To cope with the large and increasing number of flows in such SDN networks, existing studies on centralized NUM focus on the scalability of the algorithm with respect to the number of flows, but the efficiency is ignored. In this paper, we propose a centralized, efficient and scalable algorithm for the NUM problem. By designing smooth utility and penalty functions, we formulate the NUM problem with a smooth objective function, which enables the use of Nesterov’s accelerated gradient method (AGM). We prove that the proposed method achieves an$O(d/t^{2})$convergence rate, demonstrating superior convergence speed with respect to the number of iterationst, and our method is scalable with respect to the number of flowsdin the network. Our smooth objective NUM formulation and AGM are effective not only in simple network scenarios with non-prioritized flows routed on one simple paths, but also in more complex and practical scenarios involving prioritized flows routed across multiple complex paths. Experiment results confirm that our method obtains accurate solutions with fewer iterations, and achieves close-to-optimal network utility. Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009 |
IEEE Trans. Netw. | 6 |
| 2024 | NetFEC: In-network FEC Encoding Acceleration for Latency-sensitive Multimedia ApplicationsabstractIn face of packet loss, latency-sensitive multimedia applications cannot afford re-transmission because loss detection and re-transmission could lead to extra latency or otherwise compromised media quality. Alternatively, forward error correction (FEC) ensures reliability by adding redundancy and it is able to achieve lower latency at the cost of bandwidth and computational overheads. We propose to re-locate FEC encoding to hardware that better suits the computational pattern of FEC encoding than CPUs. In this paper, we present NetFEC, an in-network acceleration system that offloads the entire FEC encoding process on the emergent programmable switching ASICs, eliminating all CPU involvement. We design the ghost packet mechanism so that NetFEC can be compatible with important media transport functionalities, including congestion control, pacing and statistics. We integrate NetFEC with WebRTC and conduct extensive experiments with real hardwares. Our evaluations demonstrate that NetFEC is able to eliminate server CPU burden and adds negligible overheads. Yi Qiao, Han Zhang 0009, Jilong Wang 0001 |
INFOCOM | 2 |
| 2024 | ROV-GD: Improving the Measurement of ROV Deployment Using Graph DifferenceabstractBGP has been threatened by prefix hijacking attacks due to the lack of authentication mechanisms. In recent years, many ASes have begun to participate in RPKI deployment to improve Internet security jointly. Some researchers have studied how to measure the actual deployment of ROV globally, but such works suffer from the shortcomings of small measurement coverage and insufficient accuracy. Therefore, we propose RO-VGD, a ROV measurement method based on graph difference. We collect routing path data from the control plane and data plane. Then, we rely on prefix reachability and propagation edges for ROV inference. The results show that about half of the tested ASes have ROV filtering behaviors. In addition, our method can be extended to the ROV measurement of IXPs. Through validation and analysis, we prove that our method leads to convincing results and, at the same time, has a broader coverage and better applicability than other existing methods. Han Zhang 0009, Changqing An, Jilong Wang 0001 |
ISCC | 3 |
| 2024 | Cost-efficient flow migration for SFC dynamical scheduling in geo-distributed clouds
Weihan Chen, Han Zhang 0009, Xia Yin 0001, Xingang Shi |
Comput. Networks | 3 |
| 2024 | Proactively Verifying Quantitative Network Policy Across Unsafe and Unreliable EnvironmentsabstractNetwork managers configure networks to enforce various high-level policies, and to respond to the wide range of network events (e.g., attacks, intrusions, malicious route announcements from neighbors) that may occur. It is incredibly difficult to specify these high-level policies in terms of distributed low-level configuration. These high-level policies hold only if the distributed configurations are well equipped to react to unsafe and unreliable environments (e.g., malicious route announcements, unsafe components and devices). Therefore, it is important to proactively verify whether network policies hold across continually changing environments in terms of current network configurations. State-of-the-art policy verification techniques are limited because they can check only the Boolean policies (e.g., forwarding reachability, waypoint or blackhole-freeness). However, many policy violations express themselves in quantitative ways (e.g., a link becomes overloaded). In this paper, we propose quantitative network verification (QNV) analyzing the quantitative policies of networks across unsafe and unreliable environments. QNV translates network configurations into a symbolic simulation model that captures the stable states to which the network forwarding will converge as a result of interactions between routing protocols. It then generates a logical formula matrix that describes network forwarding in the event of failures and verifies quantitative policies based on the formula matrix. We implement QNV and evaluate it on realistic and synthetic configurations. Our evaluation shows that QNV can precisely verify quantitative policies in only a few minutes, even in large networks. Han Zhang 0009, Jilong Wang 0001, Xingang Shi, Xia Yin 0001, Jiankun Hu, Congcong Miao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Be Careful of Your Neighbors: Injected Sub-Prefix Hijacking Invisible to Public MonitorsabstractPrefix hijackings have always been a significant security issue in BGP and have continued to occur in recent years. Detecting prefix hijackings is a vital part of defending against them. Most detection approaches mainly rely on the feed from the monitors of public route collector infrastructures. We propose an injected sub-prefix hijacking that utilizes the BGP communities attribute and AS path poisoning to control the propagation of invalid sub-prefix routes. This attack only pollutes neighboring ASes, thus guaranteeing the invisibility to monitors. Then the attacker can stealthily hijack traffic passing through the polluted ASes. Through extensive simulations, we show that this attack has an enormous impact and propose the crucial indicator affecting the attacker's capability. Finally, we demonstrate that existing defenses are difficult to handle this attack and then propose several defense strategies against it. Han Zhang 0009, Changqing An, Jilong Wang 0001 |
ICC | 3 |
| 2023 | GapReplay: A High-Accuracy Packet ReplayerabstractNetwork traffic has become increasingly complicated and diverse with the continuing development of the Internet. It is challenging for the traffic generators to construct a large variety of traffic while maintaining the similarity to the real traffic. Given that real traffic from the production network is handy to retrieve, replaying traces from the production network can maximize the authenticity of traffic, contributing to experiments to evaluate the effectiveness of algorithms. Recently, many pieces of research in the networking field, especially algorithms based on machine learning (ML), are extremely sensitive to statistical characteristics of the traffic, such as ML-based traffic classification, traffic detection, traffic prediction, etc. However, there are only a few available statistics for the growing encrypted traffic, such as packet size and inter-arrival time (IAT). To replay traffic accurately, in this paper, we propose the GapReplay, a packet replayer that can remain identical with the original nanosecond-precision pcap trace in packet contents and achieve high accuracy in timestamps. GapReplay reaches line rate while transmitting packets by appending and extending packets, which will be dropped and truncated by a programmable device, to guarantee satisfying accuracy. We implement our mechanism on top of DPDK and P4. The evaluation results demonstrate that GapReplay can achieve a nanosecond-level accuracy, much better than the state-of-the-art such as MoonGen and tcpreplay, where the best of them is at least 1000 times less accurate than our mechanism. Shuanghong Yu, Han Zhang 0009, Kunling He, Xingkun Yao |
ICC | 3 |
| 2023 | PUVAR:Minimize Idle Resource SLO Violations by Uncertainty-Aware Scheduling in Cloud PlatformsabstractNowadays, idle resource makes up a non-negligible fraction of datacenter capacity in mainstream cloud platforms. Cloud platforms offer idle resource with low service level objectives (SLOs) at low prices to attract cost-sensitive users. Despite their fault tolerance, these users still want some SLO guarantees for idle resource. Cloud platforms have started to provide statistical SLO for idle resource, but the violations of statistical SLOs have not been modeled and minimized. In this paper, we propose PUVAR, a scheduling policy based on prediction+optimization, to explicitly model and minimize the statistical SLO violation for idle resource in cloud platforms. The design of PUVAR possesses two major innovations: (1) explicitly quantify prediction uncertainty of idle resource capacity and iteratively reduce its impact on scheduling decisions; (2) treat the SLO of idle resources as a soft constraint and minimize its violation by two-stage scheduling of regular requests and idle requests. We provide a theoretical convergence rate for the parameter optimization of PUVAR. Promising results of ablation analysis by comparison with popular baseline algorithms on the trace of a real cloud platform indicate that PUVAR can significantly reduce the SLO violation of idle resource with little additional cost on the current scheduling of regular requests. Han Zhang 0009, Jilong Wang 0001 |
ICWS | 2 |
| 2023 | Knocking Cells: Latency-Based Identification of IPv6 Cellular Addresses on the InternetabstractIPv6 mobile networks are becoming increasingly important. Many jobs rely on understanding IPv6 mobile networks at the IP level. Previous works on cellular identification suffer from coarse identification granularity, proprietary data, or not working for IPv6. The high latency in mobile networks makes identifying cellular addresses possible using Round-Trip Time (RTT) difference. However, due to the impact of packet loss on the measurement of RTT difference, identifying cellular addresses with less overhead is challenging. In this paper, by triggering the non-zero RTT difference of the cellular /48 subnets with probes, we propose an accurate latency-based method to identify cellular /48 subnets from fixed subnets. Experiments demonstrate that the method can identify cellular subnets with a precision of 93.52% to 99.95% and a recall of 99.96% on a worldwide dataset. The overhead of measuring RTT difference reduces to at least 1/10th compared to the existing methods while robust to packet loss. Han Zhang 0009, Anlun Hong, Jilong Wang 0001 |
ISCC | 3 |
| 2023 | Delay Based Congestion Control for Cross-Datacenter NetworksabstractNumerous distributed applications are deployed in the cross-datacenter networks (Cross-DC) where geographically distributed data centers (DC) are connected by wide area network (WAN). These online applications will generate both intra-datacenter and inter-datacenter traffic, each with distinct requirements and characteristics. We find that existing combined congestion control schemes ignore the interaction of the two types of traffic and the hybrid congestion control schemes fail to accurately estimate cross data center network congestion extent. In this paper, we propose IDCC a delay based congestion control scheme that uses delay to handle congestion inside the DC and in the WAN, respectively. We respectively utilize In-band network telemetry (INT) and round trip time (RTT) to measure the queuing delay inside DC and in the WAN and guarantee the stability of the algorithm by Proportional Integral Derivative (PID). Simultaneously, we demonstrate the empirical results of optimizing flow completion time (FCT) of intra-DC short flow by weight function in cross-DC. We implemented IDCC in simulation platform ns-3 and have performed extensive large scale simulation evaluations. Results show that IDCC decreases the FCT of intra-DC traffic by 3.6× to 12× and improves the throughput of inter-DC traffic by 9% to 16% compared to Gemini, Annulus. Yantao Geng, Han Zhang 0009, Xingang Shi, Jilong Wang 0001, Xia Yin 0001, Dongbiao He |
IWQoS | 2 |
| 2023 | Anomaly Detection in the Open World: Normality Shift Detection, Explanation, and Adaptation
Rui Yu 0003, Han Zhang 0009, Minghui Jin, Jiahai Yang 0001, Xingang Shi, Xia Yin 0001 |
NDSS | 7 |
| 2023 | Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero CodeabstractMicroservices are becoming more complicated, posing new challenges for traditional performance monitoring solutions. On the one hand, the rapid evolution of microservices places a significant burden on the utilization and maintenance of existing distributed tracing frameworks. On the other hand, complex infrastructure increases the probability of network performance problems and creates more blind spots on the network side. In this paper, we present DeepFlow, a network-centric distributed tracing framework for troubleshooting microservices. DeepFlow provides out-of-the-box tracing via a network-centric tracing plane and implicit context propagation. In addition, it eliminates blind spots in network infrastructure, captures network metrics in a low-cost way, and enhances correlation between different components and layers. We demonstrate analytically and empirically that DeepFlow is capable of locating microservice performance anomalies with negligible overhead. DeepFlow has already identified over 71 critical performance anomalies for more than 26 companies and has been utilized by hundreds of individual developers. Our production evaluations demonstrate that DeepFlow is able to save users hours of instrumentation efforts and reduce troubleshooting time from several hours to just a few minutes. Junxian Shen, Han Zhang 0009, Xingang Shi, Yunxi Shen, Yongxiang Wu, Xia Yin 0001, Jilong Wang 0001, Mingwei Xu 0001, Jiping Yin, Jianchang Song, Zhuofeng Li, Runjie Nie |
SIGCOMM | 2 |
| 2023 | From the Dialectical Perspective: Modeling and Exploiting of Hybrid Worm PropagationabstractThe hierarchical network is the more effective platform, which provides multiple channels for various worm propagation. Thus, emerging worms can infect vulnerable hosts by scanning strategy and social media. However, the spread of scan-based worm is restrained due to uneven distribution of vulnerable hosts and NAT (Network Address Translation) technique. Meanwhile, topological dependency dictates to topology-based worm only infecting those hosts in social networks. To avoid their respective disadvantages, modern hybrid worm, which combines the above two propagation mechanisms, can implement efficient IP-address scanning by enhanced combination-scanning strategy, and spread more aggressively in social networks using enhanced reinfection mechanism. This paper presents a Hierarchical-Stochastic Propagation model to understand hybrid worm propagation. Inspired by hybrid worm, we design a new vaccine based on the Hierarchical-Measure Immunization strategy. For physical networking layer, we can estimate vulnerable-host distribution to find vulnerable hosts effectively through Maximum Likelihood estimation. For social networking layer, we use a novel propagation centrality measure to discover vital social nodes accurately. The experimental results show that our model can characterize the propagation mechanism of hybrid worms more comprehensively, and greatly outperforms state of the art models in terms of estimation accuracy. Meanwhile, our strategy is more effective to restrain the hybrid worm from spreading in networks. Tianbo Wang 0001, Huacheng Li, Chunhe Xia, Han Zhang 0009, Pei Zhang 0003 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Real-Time Malicious Traffic Detection With Online Isolation Forest Over SD-WANabstractSoftware Defined Network (SDN) has been widely used in modern network architecture. The SD-WAN is considered as a technology that has a potential to revolutionize the WAN service usage by utilizing the SDN philosophy. Attacking SDN router and controller can affect the network and block the entire services. In this paper, we propose a machine learning based anomalous traffic detection framework named OADSD over SD-WAN that can achieve task independent and has the ability of adapting to the environment. The OADSD adopts Distributed Dynamic Feature Extraction (DDFE) to extract representative features directly from the raw traffic, and proposes the On-demand Evolving Isolation Forest (OEIF) to make the system adapt to an environment. We provide a theoretical analysis of the performance of the OADSD. We also conduct comprehensive experiments to evaluate the performance of the OADSD with real world public datasets as well as a small real testbed. Our experiments under real world public datasets show that, the OADSD can accurately detect various kinds of attacks with a high performance. Compared with the state-of-the-art systems, the OADSD can achieve up to 60% accuracy improvement. Pei Zhang 0003, Fangzhou He, Han Zhang 0009, Jiankun Hu, Xiaohong Huang 0003, Jilong Wang 0001, Xia Yin 0001, Huahong Zhu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Achieving High Availability in Inter-DC WAN Traffic EngineeringabstractInter-DataCenter Wide Area Network (Inter-DC WAN) that connects geographically distributed data centers is becoming one of the most critical network infrastructures. Due to limited bandwidth and inevitable link failures, it is highly challenging to guarantee network availability for services, especially those with stringent bandwidth demands, over inter-DC WAN. We present$\mathsf {TEDAT}$, a novel Traffic Engineering (TE) framework for Diverse Availability Targets (DAT), where a Service Level Agreement (SLA) is defined to ensure that each bandwidth demand must be satisfied with a stipulated probability, when subjected to the network capacity and possible failures of the inter-DC WAN.$\mathsf {TEDAT}$has two core components, i.e., traffic scheduling and failure recovery, which are crystalized through different mathematical models and theoretically analyzed. They are also extensively compared against state-of-the-art TE schemes, using a testbed as well as real trace driven simulations across different topologies, traffic matrices and failure scenarios. Our evaluations show that, compared with the optimal admission strategy,$\mathsf {TEDAT}$can speed up the online admission control by$30\times $at the expense of less than 4% false rejections. On the other hand, compared with the latest TE schemes like FFC and TEAVAR,$\mathsf {TEDAT}$can meet the bandwidth availability SLAs for 23%~60% more demands under normal loads, and when network failure causes SLA violations, it can retain 10%~20% more profit under a pricing and refunding model. Han Zhang 0009, Xia Yin 0001, Xingang Shi, Jilong Wang 0001, Yingya Guo, Tian Lan 0001, Ke Ruan, Haijun Geng |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | A General Approach to Generate Test Packets With Network ConfigurationsabstractThe correctness and reliability of modern networks are often the greatest concerns. A myriad network events like software update, device crash and resource exhaustion, inevitably lead to liveness errors on data plane. This paper focuses on fault detection of the network data plane using test packets. Existing test packet generation techniques are limited in two aspects: i) it is difficult to collect the input data plane snapshot through SNMP or terminals ii) it may rise false negatives due to inconsistent snapshot. In this paper, we propose a new framework, SWIFT, that automatically generates test packets with network configurations. SWIFT minimizes the number of test packets by allowing a packet to go through multiple links or interfaces. For network updates, SWIFT updates test packets in an incremental way to revalidate the network. We evaluate its performance using hundreds of benchmark network configurations. The results show that it takes few seconds to generate test packets to exercise all links and interfaces, and updates the test packets in few seconds for configuration changes. We also deployed a SWIFT prototype in a university network, and successfully detected many network outages. Han Zhang 0009, Jilong Wang 0001, Xia Yin 0001, Xingang Shi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Serpens: A High Performance FaaS Platform for Network FunctionsabstractMore and more enterprises deploy applications on Function-as-a-Service (FaaS) platforms to improve resource efficiency and save monetary costs. Network Functions (NFs) suffer from staggered peaks of traffic patterns and could benefit from fine-grained resource multiplexing in FaaS platform. However, naively exploring existing FaaS platforms to support NFs can introduce significant performance overheads in three aspects, including slow instance startup, remote state access for NFs, and costly packet delivery between NFs. To address these problems, we propose${\sf Serpens}$, a high performance FaaS platform for NFs. First,${\sf Serpens}$proposes a reusable NF runtime design to slash instance startup overhead. Second,${\sf Serpens}$designs a novel state management mechanism to support local state access. Third,${\sf Serpens}$introduces an advanced service chaining approach to avoid extra packet delivery. Besides,${\sf Serpens}$designs an NF scaling mechanism to minimize performance fluctuation. We have implemented a prototype of${\sf Serpens}$and conducted comprehensive experiments. Compared with the NFs and Service Function Chains (SFCs) that run on existing FaaS platforms,${\sf Serpens}$can improve the throughput by more than 10× and reduce the latency by more than 90%. Heng Yu 0005, Han Zhang 0009, Junxian Shen, Yantao Geng, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Gringotts: Fast and Accurate Internal Denial-of-Wallet Detection for Serverless ComputingabstractServerless computing, or Function-as-a-Service, is gaining continuous popularity due to its pay-as-you-go billing model, flexibility, and low costs. These characteristics, however, bring additional security risks, such as the Denial-of-Wallet (DoW) attack, to serverless tenants. In this paper, we perform a real-world DoW attack on commodity serverless platforms to evaluate its severity. To identify such attacks, we design, implement, and evaluate Gringotts, an accurate, easy-to-use DoW detection system with a negligible performance overhead. Gringotts addresses the information ambiguity inherent in serverless functions by introducing a well-designed performance metrics collection agent. Then, Gringotts uses the Mahalanobis distance to discover anomalies in the distribution of the metrics. We implement Gringotts as a real system and conduct extensive experiments using a testbed to evaluate the performance of Gringotts. Our results indicate that Gringotts has a performance overhead of less than 1.1%, with an average detection delay of 1.86 seconds and an average accuracy of over 95.75%. Junxian Shen, Han Zhang 0009, Yantao Geng, Jilong Wang 0001, Mingwei Xu 0001 |
CCS | 2 |
| 2022 | FedRFID: Federated Learning for Radio Frequency Fingerprint Identification of WiFi SignalsabstractWith the rapid development of the cognitive radio networks, the number of terminal devices has exploded. Massive devices generate a large amount of privacy-sensitive data, typically WiFi signals. This paper proposes a method for Radio frequency (RF) fingerprinting identification of WiFi signals based on federated learning, which trains a cooperative model to complete RF fingerprinting identification without transmitting privacy-sensitive data. The experimental findings on a real-world dataset validate that the strategy described in this study increases the RF fingerprinting identification accuracy in a variety of size circumstances, and ensures that data privacy will not be compromised. Jibo Shi, Han Zhang 0009, Sen Wang 0006, Shiwen Mao, Yun Lin 0005 |
GLOBECOM | 2 |
| 2022 | Centralized Network Utility Maximization with Accelerated Gradient MethodabstractNetwork utility maximization (NUM) is a well-studied problem for network traffic management and resource allocation. Because of the inherent decentralization and complexity of networks, most researches develop decentralized NUM algorithms. In recent years, the Software Defined Networking (SDN) architecture has been widely used, especially in cloud networks and inter-datacenter networks managed by large enterprises, promoting the design of centralized NUM algorithms. To cope with the large and increasing number of flows in such SDN networks, existing researches about centralized NUM focus on the scalability of the algorithm with respect to the number of flows, however the efficiency is ignored. In this paper, we focus on the SDN scenario, and derive a centralized, efficient and scalable algorithm for the NUM problem. By the designing of a smooth utility function and a smooth penalty function, we formulate the NUM problem with a smooth objective function, which enables the use of Nesterov's accelerated gradient method. We prove that the proposed method has$O(d/t^{2})$convergence rate, which is the fastest with respect to the number of iterations$t$, and our method is scalable with respect to the number of flows$d$in the network. Experiments show that our method obtains accurate solutions with less iterations, and achieves close-to-optimal network utility. Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009 |
ICNP | 6 |
| 2022 | Scorpius: Proactive Code Preparation to Accelerate Function StartupabstractMassive enterprises deploy their applications on public clouds to relieve infrastructure management burden. However, applications are faced with highly fluctuating workloads, while clouds provision exclusive resources at coarse time granularity, resulting in severely low resource efficiency. Function-as-a-Service (FaaS) platform enables fine-grained resource multiplexing, which has the potential to improve efficiency. However, FaaS platforms could consume several seconds to start functions and the long startup latency can severely hurt the performance of applications. In this paper, we measure the FaaS platforms and find that most startup latency is occupied by code preparation. To reduce the code preparation latency with little resource overhead, we propose Scorpius, a FaaS platform that proactively prepares code based on the historical data of functions. It combines two optimization categories: (1) To reduce the code size, Scorpius proposes to proactively prepare partial libraries over servers and run functions on the server with most library sharing. (2) To advance the start time, Scorpius proposes to predict the function overload with a simple model and proactively scale code to more servers. We have implemented a prototype of Scorpius and conducted extensive experiments. Evaluation results demonstrate that compared with state-of-the-art methods, Scorpius can reduce the code preparation latency by 87.6% with only 9.3% storage overhead. Heng Yu 0005, Junxian Shen, Han Zhang 0009, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001 |
IWQoS | 3 |
| 2022 | Dynamic and Diverse Transformations for Defending Against Adversarial ExamplesabstractIt is demonstrated that deep neural networks can be easily fooled by adversarial examples. To improve the robustness of neural networks against adversarial attacks, substantial research on adversarial defenses is being carried out, of which input transformation is a typical category of defenses. However, because the transformation also has an impact on the accuracy of clean examples, the existing transformation-based defenses usually adopt minor transformations such as shift and scaling, which limits the defense effect of the transformation to some extent. To this end, we propose a method by using dynamic and diverse transformations for defending against adversarial attacks. Firstly, we constructed a transformation pool that contains both minor and major transformations (e.g., flip, rotate). Secondly, we retrained the model with the data transformed by major transformations to ensure that the performance of model itself is not affected. Finally, we dynamically select transformations to preprocess the input of the model to defend against adversarial examples. We conducted extensive experiments on MNIST and CIFAR-10 datasets and compared our method with the state-of-the-art adversarial training and transformation-based defenses. The experimental results show that our proposed method outperforms the existing methods, improving the robustness of the model against adversarial examples greatly while maintaining high accuracy on clean examples. Our code is available at https://github.com/byerose/DynamicDiverseTransformations. Ming Zhang 0021, Xiaohui Kuang, Xuhong Zhang 0002, Han Zhang 0009 |
TrustCom | 6 |
| 2022 | Multi channel spectrum prediction algorithm based on GCN and LSTMabstractWith the increasingly serious shortage of spectrum resources, spectrum dynamic access based on spectrum prediction technology is widely recognized. Due to the high burstiness and complex intrinsic correlation of spectrum monitoring data, high-precision multi-channel spectrum prediction is challenging. This paper constructs spectrum monitoring data as a kind of graph structure data based on the correlation of spectrum itself, and designs a graph network model combining Graph convolution network(GCN) and Long-short term memory network(LSTM) for multi-channel spectrum prediction. This paper creatively introduces the method of graph network. And GCN is used instead of CNN to extract the correlation of channels, so as to improve the accuracy of multi-channel prediction. Experiments are conducted based on a real-world spectrum measurement dataset. The results show that the model proposed in this paper has better predictive performance compared with other methods. Han Zhang 0009, Qiao Tian 0002, Yu Han 0003 |
VTC Fall | 1 |
| 2022 | THREATRACE: Detecting and Tracing Host-Based Threats in Node Level Through Provenance Graph LearningabstractHost-based threats such as Program Attack, Malware Implantation, and Advanced Persistent Threats (APT), are commonly adopted by modern attackers. Recent studies propose leveraging the rich contextual information in data provenance to detect threats in a host. Data provenance is a directed acyclic graph constructed from system audit data. Nodes in a provenance graph represent system entities (e.g.,processesandfiles) and edges represent system calls in the direction of information flow. However, previous studies, which extract features of the provenance graph, are not sensitive to the small quantity of threat-related entities and thus result in low performance when hunting stealthy threats. We present THREATRACE, an anomaly-based detector that detects host-based threats at system entity level without prior knowledge of attack patterns. We tailor GraphSAGE, an inductive graph neural network, to learn every benign entity’s role in a provenance graph. THREATRACE is a real-time system, which is scalable of monitoring a long-term running host and capable of detecting host-based intrusion in their early phase. We evaluate THREATRACE on five public datasets. The results show that THREATRACE outperforms seven state-of-the-art host intrusion detection systems. Xia Yin 0001, Han Zhang 0009, Xingang Shi, Jiahai Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2022 | NetEC: Accelerating Erasure Coding Reconstruction With In-Network AggregationabstractIn distributed storage systems, Erasure Coding (EC) is a crucial technology to enable high data availability. By downloading parity data from survived machines, EC can reconstruct lost data with much lower storage overheads than data replication. However, this reduction in storage cost comes at the expense of extra performance problems:low reconstruction rate,high degraded read latency, andhigh host CPU utilization. Our analysis shows that these performance problems are deeply rooted in thehost-basedEC processing. To resolve these problems, we present NetEC, an in-network accelerating framework that fully offloads EC to the new generation programmable switching ASICs. We propose Explicit Buffer Size Notification (EBSN) to constrain decoding buffer usage, and design an on-switch one-to-many TCP proxy to integrate EBSN with TCP. We also design two parallel Galois Field (GF) offloading methods—table lookup and bitmatrix methods—to maximize parsable bytes. We implement NetEC on programmable switches and integrate it with HDFS. Extensive evaluations show that NetEC improves the reconstruction rate by 2.7x-6.8x, reduces the degraded read latency significantly, and removes the host CPU overhead completely. We also emulate multi-rack scenarios and show that NetEC is able to support$\sim$∼GB/s reconstruction rate and tens of concurrent tasks. Yi Qiao, Menghao Zhang 0001, Yu Zhou 0008, Han Zhang 0009, Mingwei Xu 0001, Jun Bi, Jilong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Efficient and Accurate Flow Record Collection With HashFlowabstractTraditional tools like NetFlow face great challenges as both the speed and the complexity of the network traffic increase. To keep the pace up, we propose HashFlow for more efficient and accurate collection of flow records. HashFlow keeps large flows in its main flow table and uses an ancillary table to summarize the other flows when the main table is full. With ourflow collision resolutionandflow record promotionschemes, a flow in the ancillary table is promoted back to the main flow table with a guaranteed probability when it becomes large enough. These operations can be performed highly efficiently, so HashFlow can keep up with ultra-high traffic speed. We implement HashFlow in a Tofino switch, and using traces from different operational networks, we compare its performance against some state-of-the-art flow measurement algorithms. Our experiments show that, for various types of traffic analysis applications, HashFlow consistently demonstrates clearly better performance than its competitors. For example, the performance of HashFlow in flow size estimation, flow size distribution estimation and heavy hitter detection is up to 21, 60 and 35 percent better than those of the best competitors respectively, and these merits of HashFlow come with almost no degradation of throughput. Zongyi Zhao, Xingang Shi, Qing Li 0006, Han Zhang 0009, Xia Yin 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security ApplicationsabstractUnsupervised Deep Learning (DL) techniques have been widely used in various security-related anomaly detection applications, owing to the great promise of being able to detect unforeseen threats and superior performance provided by Deep Neural Networks (DNN). However, the lack of interpretability creates key barriers to the adoption of DL models in practice. Unfortunately, existing interpretation approaches are proposed for supervised learning models and/or non-security domains, which are unadaptable for unsupervised DL models and fail to satisfy special requirements in security domains. Ying Zhong 0008, Han Zhang 0009, Jiahai Yang 0001, Xingang Shi, Xia Yin 0001 |
CCS | 6 |
| 2021 | Cost-Efficient Dynamic Service Function Chain Embedding in Edge CloudsabstractEdge Computing (EC) provides delay protection for some delay-sensitive network services by deploying cloud infrastructure with limited resources at the edge of the network. In addition, Network Function Virtualization (NFV) implements network functions by replacing traditional dedicated hardware devices with Virtual Network Function (VNF) that can run on general servers. In NFV environment, Service Function Chaining (SFC) is regarded as a promising way to reduce the cost of configuring network services. NFV therefore allows to deploy network functions in a more flexible and cost-efficient manner, and schedule network resources according to the dynamical variation of network traffic in EC. For service providers, seeking an optimal SFC embedding scheme can improve service performance and reduce embedding cost. In this paper, we study the problem of how to dynamically embed SFC in geo-distributed edge clouds network to serve user requests with different delay requirements, and formulate this problem as a Mixed Integer Linear Programming (MILP) which aims to minimize the total embedding cost. Furthermore, a novel SFC Cost-Efficient emBedding (SFC-CEB) algorithm has been proposed to efficiently embed required SFC and optimize the embedding cost. Based on the results of trace-driven simulations, the proposed algorithm can reduce SFC embedding cost by up to 37% compared with state-of-the-art schemes (e.g., RDIP). Weihan Chen, Han Zhang 0009, Xia Yin 0001, Xingang Shi |
CNSM | 3 |
| 2021 | Traffic Engineering with Segment Routing Considering Probabilistic FailuresabstractSegment Routing (SR) is a source routing paradigm that routes a packet through an ordered list of instructions called segments. It is widely used in Traffic Engineering (TE) because of its simplicity and scalability. Although there are lots of research about TE with SR (SR-TE), fewer consider network failures. The reactive approaches may suffer from latency and update issues, and the proactive approaches don't perform very well because the objectives aren't carefully designed. Besides, although different types of failures are considered, the failure probabilities are ignored. In this paper, we take failure probabilities in to consideration, and propose a proactive 2-SR model 2SRPF to handle SR-TE problem with network failures, aiming at minimizing maximum link utilization (MLU). Considering that severe failures are more noteworthy, we use probability as a severity threshold, and minimize the expectation of the larger MLUs whose corresponding failure states have probabilities sum to a specific threshold value. We solve it with probabilistic risk management. Experiments show that 2SRPF performs well with one threshold setting for different topologies consistently, and gets close to optimal results when network fails. Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009, Yingya Guo, Haijun Geng |
CNSM | 6 |
| 2021 | Boosting bandwidth availability over inter-DC WANabstractInter-DataCenter Wide Area Network (Inter-DC WAN) that connects geographically distributed data centers is becoming one of the most critical network infrastructures. Due to limited bandwidth and inevitable link failures, it is highly challenging to guarantee network availability for services, especially those with stringent bandwidth demands, over inter-DC WAN. We present BATE, a novel Traffic Engineering (TE) framework for bandwidth availability (BA) provision, which aims to ensure that each bandwidth demand must be satisfied with a stipulated probability, when subjected to the network capacity and possible failures of the inter-DC WAN. The three core components of BATE, i.e., admission control, traffic scheduling and failure recovery, are formulated through different mathematical models and theoretically analyzed. They are also extensively compared against state-of-the-art TE schemes, using a testbed as well as real trace driven simulations across different topologies, traffic matrices and failure scenarios. Our evaluations show that, compared with the optimal admission strategy, BATE can speed up the online admission control by 30x at the expense of less than 4% false rejections. On the other hand, compared with the latest TE schemes like FFC and TEAVAR, BATE can meet the bandwidth availability targets for 23%~60% more demands under normal loads, and when network failure causes BA targets violations. Han Zhang 0009, Xingang Shi, Xia Yin 0001, Jilong Wang 0001, Yingya Guo, Tian Lan 0001 |
CoNEXT | 1 |
| 2021 | Traffic Engineering in Hybrid Software Defined Network via Reinforcement Learning
Yingya Guo, Han Zhang 0009, Wenzhong Guo, Xia Yin 0001 |
J. Netw. Comput. Appl. | 3 |
| 2021 | Log-Based Anomaly Detection With Robust Feature Extraction and Online LearningabstractCloud technology has brought great convenience to enterprises as well as customers. System logs record notable events and are becoming valuable resources to track and investigate system status. Detecting anomaly from logs as fast as possible can improve the quality of service significantly. Although many machine learning algorithms (e.g., SVM, Logistic Regression) have high detection accuracy, we find that they assume data are clean and might have high training time. Facing these challenges, in this paper, we propose Robust Online Evolving Anomaly Detection (ROEAD) framework which adopts Robust Feature Extractor (RFE) to remove the effects of noise and Online Evolving Anomaly Detection (OEAD) to dynamic update parameters. We propose Online Evolving SVM (OES) algorithm as the example of online anomaly detection methods. We analyze the performance of OES in theory and prove the performance difference between OES and the best hypothesis tends to zero as time goes infinity. We compare the performance of ROEAD against state-of-the-art anomaly detection algorithms using public log datasets. The results demonstrate that ROEAD is able to remove the effects of noise and OES can improve the detection accuracy by more than 40%. Shangbin Han, Qianhong Wu, Han Zhang 0009, Jiankun Hu, Xingang Shi, Linfeng Liu 0001, Xia Yin 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Assisting reachability verification of network configurations updates with NUV
Xia Yin 0001, Xingang Shi, Fangdan Ye, Jiangyuan Yao, Han Zhang 0009 |
Comput. Networks | 8 |
| 2020 | Efficient computation of loop-free alternates
Haijun Geng, Han Zhang 0009, Xingang Shi, Xia Yin 0001 |
J. Netw. Comput. Appl. | 2 |
| 2019 | MSAID: Automated detection of interference in multiple SDN applications
Jiangyuan Yao, Xia Yin 0001, Xingang Shi, Han Zhang 0009 |
Comput. Networks | 7 |
| 2019 | Joint optimization of tasks placement and routing to minimize Coflow Completion Time
Yingya Guo, Han Zhang 0009, Xia Yin 0001, Xingang Shi |
J. Netw. Comput. Appl. | 3 |
| 2019 | DA&FD-Deadline-Aware and Flow Duration-Based Rate Control for Mixed Flows in DCNsabstractData center has become an important facility for hosting various applications. For data center networks, deadline missing rate and average flow completion time are two main metrics for the performance of applications. In this paper, we find deadline-aware methods can only reduce the percentage of flows missing deadline, while flowsize-aware and information-cumulative methods can only optimize the average flow completion time. However, traffic in data center is the mixture of various flows and focusing on the single goal is not enough. We advocate to incorporate deadline and flow duration time into flow rate control. Then we design DA&FD (Deadline-Aware and Flow Duration) based rate control mechanism and analyze its performance in theory. At last, we evaluate DA&FD under different topologies, real world traffic and load scenarios, both by simulation and in real testbed. Our results show that DA&FD performs close to D2TCP and about 15%, 25%, 30%, 35% better than Ameon, L2DCT, Karuna, DCTCP on deadline missing rate. For average FCT, the performance of DA&FD is similar to L2DCT and compared with Ameon, D2TCP, Karuna, DCTCP, DA&FD can reduce average FCT by 10%, 15%, 20%, 25%. Han Zhang 0009, Haijun Geng, Xia Yin 0001, Xingang Shi, Qianhong Wu, Jianwei Liu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Efficient Scheduling of Weighted Coflows in Data CentersabstractTraditional network resource management mechanisms are mainly flow or packet based. Recently, coflow has been proposed as a new abstraction to capture the communication patterns in a rich set of data parallel applications in data centers. Coflows effectively model the application-level semantics of network resource usage, so high-level optimization goals, such as reducing the transfer latency of applications, can be better achieved by taking coflows as the basic elements in network resource allocation or scheduling. Although efficient coflow scheduling methods have been studied, in this paper, we advocate to schedule weighted coflows as a further step in this direction, where weights are used to express the importances or priorities of different coflows or their corresponding applications. We propose the Weighted Coflow Completion Time (WCCT) minimization problem and a (2-2/n+1)-approximate optimal offline algorithm, where n is the concurrent number of coflows. We then design an information-agnostic online algorithm named IAOA to dynamically schedule coflows according to their weights and the instantaneous network condition. We also design and implement a coflow scheduling system named FlyTransfer, which can use the online algorithm as its scheduling method. We test the performance of FlyTransfer by trace-driven simulations as well as real deployment in openstack. Our evaluation results show that, compared to the latest information-agnostic coflow scheduling algorithms, FlyTransfer can reduce more than 40 percent of the WCCT, and more than 30 percent of the completion time for coflows with above-the-average level of importance. It even outperforms the most efficient clairvoyant coflow scheduling method by reducing around 30 percent WCCT, and 25- 30 percent of the completion time for coflows with above-the-average importance, respectively. Han Zhang 0009, Xingang Shi, Xia Yin 0001, Haijun Geng, Qianhong Wu, Jianwei Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Yosemite: Efficient scheduling of weighted coflows in data centersabstractRecently, coflow has been proposed as a new abstraction to capture the communication patterns in a rich set of data parallel applications in data centers. Coflows effectively model the application-level semantics of network resource usage, so high-level optimization goals, such as reducing the transfer latency of applications, can be better achieved by taking coflows as the basic elements in network resource allocation or scheduling. Although efficient coflow scheduling methods have been studied, in this paper, we propose to schedule weighted coflows as a further step in this direction, where weights are used to express the emergences or priorities of different coflows or their corresponding applications. We design an information-agnostic online algorithm to dynamically schedule coflows according to their weights and the instantaneous network condition. Then We implement the algorithm in a scheduling system named Yosemite. Our evaluation results show that, compared to the latest information-agnostic coflow scheduling algorithms, Yosemite can reduce more than 40% of the WCCT (Weighted Coflow Completion Time), and more than 30% of the completion time for coflows with above-the-average level of emergence. It even outperforms the most efficient clairvoyant coflow scheduling method by reducing around 30% WCCT, and 25%~30% of the completion time for coflows with above-the-average emergence, respectively. Han Zhang 0009, Xingang Shi, Xia Yin 0001 |
ICNP | 1 |
| 2017 | Joint source selection and transfer optimization for erasure coding storage systemabstractWith the deployment of big data applications, more and more data are stored in the online storage. Erasure coding storage system has been widely used by companies such as Google and Facebook, since it provides space-optimal data redundancy to protect against data loss. In erasure coding storage system, (n, k) MDS erasure code is used to divide file into n chunks. When a user want to access the file, any subset of k out of n chunks will be needed to reconstruct the file. In this case, how to select k out of n chunks and how to let the chunks transfer quickly become important problems. In this paper, we joint the two problems together to optimize. Our optimization goal is to minimize average file access time (FAT). To achieve this, we propose smallest load first heuristic to do source selection and design an online algorithm to reduce chunks transfer latency. Base on this, we design and implement D-Target, a centralized scheduler that tries to minimize average FAT in distributed erasure coding storage system. We then test D-Target's performance by trace-driven simulation. Results show that, for the trace of AT&T, D-Target performs 2.5×, 1.7×, 1.8×, 3.6× better than TCP, Aalo, Barrat and pFabric respectively. Han Zhang 0009, Xingang Shi, Yingya Guo, Haijun Geng, Xia Yin 0001 |
IPCCC | 1 |
| 2017 | More load, more differentiation - Let more flows finish before deadline in data center networks
Han Zhang 0009, Xingang Shi, Yingya Guo, Xia Yin 0001 |
Comput. Networks | 1 |
| 2016 | Reduce completion time and guarantee throughput by transport with slight congestionabstractIn typical data center networks, an overwhelming majority of the flows are smaller than 200 KB in size, while most transmitted bytes are from a small fraction of large flows. The small flows are usually from the applications interacting with end users, thus they require small completion times. Meanwhile, the data center owners hope to keep the high throughput of the network to make full use of their investments on the network devices. To reduce the completion times of small flows while maintaining the high throughput of the network, we propose a novel transport algorithm, SCT (Transport with Slight Congestion), in this paper. SCT gives small flows higher priority by increasing their congestion windows at a higher rate. Moreover, SCT keeps the network to be in high utilization, thus the throughput of network is guaranteed. Extensive simulations show that SCT can reduce the average completion time of small flows by up to 48% at the expense of degrading the throughput of network by 5% only, compared with DCTCP. Zongyi Zhao, Qing Li 0006, Mingwei Xu 0001, Xingang Shi, Han Zhang 0009 |
ICC | 5 |
| 2016 | FDRC - Flow duration time based rate control in data center networksabstractData Center is now becoming an important facility for many applications (e.g, web search and retail). As TCP can't meet applications' demands for latency and throughput, many tcp-based protocols (e.g, DCTCP, D2TCP, L2DCT) have been proposed. Among them, protocols such as D2TCP incorporate explicit deadline into congestion window adjustment procedure to guarantee flows' latency and protocols such as L2DCT consider flow size when computing congestion window adjustment factor to guarantee the throughput of short flows. These two methods work well at some scenery but they have some deficiencies on two aspects. Firstly, we find that they can only reduce the percentage of flows missing deadline or reduce flow completion time, but can not meet both the goals simultaneously. Secondly, most of these methods need the user to know flow information (e.g, deadline, flow size), which may be hard to know exact value beforehand. In this paper, we advocate to use flow duration time into congestion window adjustment procedure. Based on this, we propose FDRC-Flow Duration Time based Rate Control algorithm. We find that without knowing flow information beforehand, FDRC can achieve the goal of reducing the percentage of flows missing deadline and cutting average flow completion time simultaneously. We theoretically analyze FDRC's behavior and implement FDRC into ns-2 as well as linux kernel. Our experiments show that FDRC performs better than D2TCP and L2DCT at nearly all the scenarios. On average, it performs 30% better than the state-of-art deadline-aware congestion control protocol D2TCP and 10% better than the state-of-art flowsize-aware protocol L2DCT. Han Zhang 0009, Xingang Shi, Xia Yin 0001, Yingya Guo |
IWQoS | 1 |
| 2015 | An efficient link protection scheme for link-state routing networksabstractTo enhance the network reliability without incurring significant extra overhead, we propose a novel link protection scheme, Hybrid Link Protection (HLP), to achieve failure resilient routing. Compared to previous schemes, HLP ensures high network availability in a more efficient way, and also provides other features such as load balancing. HLP is implemented in two stages. Stage one provides Multiple Next-hop Protection (MNP), where only one single Shortest Path Tree (SPT) needs to be constructed on each node to find multiple next hops for any destination. Stage two provides Backup Path Protection (BPP), where only a minimum number of links need to be protected, using special paths and packet headers, to meet the network availability requirement. We evaluate these algorithms in a wide spread of relevant topologies, both real and synthetic, and the results reveal that HLP can achieve high network availability without introducing conspicuous overhead. Haijun Geng, Xingang Shi, Xia Yin 0001, Han Zhang 0009 |
ICC | 5 |
| 2015 | More load, more differentiation - A design principle for deadline-aware congestion controlabstractData center network has become an important facility for hosting various online services and applications, and thus its performance and underlying technologies are attracting more and more interests. In order to achieve better network performance, recent studies have proposed to tailor data center network traffic management in different aspects, devising various routing and transport schemes. In particular, for applications that must serve users in a timely manner, strict deadlines for their internal traffic flows should be met, and are explicitly taken into consideration in some latest flow rate control or scheduling algorithms in data center networks. In this paper, we advocate that when designing such deadline-aware rate control schemes, a simple principle should be followed: flows with different deadlines should be differentiated in their bandwidth allocation/occupation, and the more traffic load, the more differentiation should be made. We derive sufficient and necessary conditions for a flow rate control scheme to follow this principle, and present a simple congestion control algorithm called Load Proportional Differentiation (LPD) as its application. We have evaluated LPD under different topologies and load scenarios, both by simulation and in real testbed. Our results show that LPD nearly always outperforms D2TCP, a latest deadline-aware rate control scheme, and often reduces the number of flows missing their deadlines by more than 50%. We also give some other applications of this principle, for example, in reducing flow completion time. Han Zhang 0009, Xingang Shi, Xia Yin 0001, Fengyuan Ren |
INFOCOM | 1 |
| 2015 | Incremental deployment for traffic engineering in hybrid SDN networkabstractTraffic engineering is a method to balance the flows and optimize the routing in the network. Software defined networking is a new network architecture and we can gain great benefit by migrating the traditional IP network to the SDN-enabled network from the perspective of traffic engineering. However, due to the economical, organizational and technical challenges, migrating to the network with a full deployment of SDN routers is impractical in the short term. It is a desirable choice to deploy SDN incrementally. In this paper, we seek to search for an optimal migration sequence of the legacy routers to SDN-enabled routers so that we can decide where and how many routers to migrate firstly. Our main contribution is that we propose a heuristic algorithm, i.e., genetic algorithm, to seek a migration sequence of the routers that obtains the most of the benefit from the perspective of traffic engineering. We evaluate the algorithm by conducting simulation experiments, making comparison to the greedy migration algorithm and static migration algorithms that we propose. The experiments exhibit that the genetic algorithm, outperforms the other migration algorithms in searching for a migration sequence. When properly deployed, about a migration of 40% of routers reaps most of the benefit. Yingya Guo, Xia Yin 0001, Xingang Shi, Han Zhang 0009 |
IPCCC | 6 |
| 2015 | Algebra and algorithms for efficient and correct multipath QoS routing in link state networksabstractThe diversity of QoS (Quality-of-Service) requirements of Internet applications motivates various QoS routing algorithms that take different QoS metrics into consideration. Routing algebra has been proposed as a framework to study the fundamental properties of QoS routing algorithms, such as their optimality and loop-freeness. However, for multipath QoS routing, little has been done in these aspects. Existing multipath QoS routing algorithms often take a rather conservative approach to guarantee loop-freeness, at the cost of efficiency. On the other hand, simply adapting existing efficient multipath routing algorithms to support various QoS metrics cannot guarantee correctness. In face of that, we propose a routing metric algebra for multipath QoS routing in link state networks, where a key property of the routing metrics called isotonicity, which plays an important role. To let routers efficiently and correctly find multiple next-hops for each destination, we also develop two distributed multipath QoS routing algorithms. The algorithms are run locally and independently, without exchanging messages other than the basic link states. They are specifically tailored for algebras with strict or non strict isotonicity, and their correctness are formally proved. Haijun Geng, Xingang Shi, Xia Yin 0001, Han Zhang 0009 |
IWQoS | 5 |
| 2014 | Let more nodes have a second choiceabstractCurrent intra-domain routing protocols computes only shortest paths for any pair of nodes which cannot provide good fast reroute when network failures occur. Multipath routing can be fundamentally more efficient than the currently used single path routing protocols. It can significantly reduce congestion in network by shifting traffic to unused network resources. This improves network utilization and provides load balancing. To enhance failure resiliency we propose a new scheme More Nodes Have At Least Two Choices (MNTC) where the goal is how to maximize the number of nodes that have at least two next-hops towards their destinations. We evaluate the algorithm in a wide space of relevant topologies and the results show that it can achieve good reliability while keeping low stretch. Haijun Geng, Xingang Shi, Xia Yin 0001, Han Zhang 0009, Jiangyuan Yao |
IPCCC | 5 |
| 2014 | A hybrid link protection scheme for link-state routing networksabstractThe Internet is playing an increasingly crucial role in both personal and business activities. Handling link failures is an important task in designing routing protocols. To enhance the network availability without incurring significant extra overhead, we propose a novel link protection algorithm, Hybrid Link Protection Scheme (HLP) to achieve failure resilient routing. Haijun Geng, Xingang Shi, Xia Yin 0001, Han Zhang 0009, Jiangyuan Yao |
IPCCC | 5 |