Jiahai Yang 0001

dblp:62/2814-1 · also Jiahai John Yang · DBLP profile ↗
← Back
165ranked-venue papers
2as first author
89since 2021 · last 2026
0000-0001-6109-6737ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 98 · 46 since 2021Security and privacy · 31 · 1 first-author · 25 since 2021Systems, architecture and hardware · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Understanding the IPv6 Address Usage Strategies of Top Internet Services
Lin He 0004, Zedong Jia, Daguo Cheng, Jinlong E, Yuhan Du, Guanglei Song, Ying Liu 0024, Xingang Shi, Shenglin Zhang, Jiahai Yang 0001, Mingwei Xu 0001
ICC10
2026 FASST-LLM: Implementing Service Scanning Tools Automatically with LLM
Lichao Qin, Yirui Luo, Chenglong Li 0006, Jiahai Yang 0001
INFOCOM5
2026 Breaking the Seed Barrier: Discovering Active IPv6 Addresses in Seedless Scenarios
Wenjian Zhang, Guanglei Song, Binkai Ma, Lin He 0004, Songyun Wu, Jiahai Yang 0001
INFOCOM8
2026 ISP or Customer? Inferring the Ownership of Public IPs of Non-Cooperative Satellite Internet via Internet Measurements
Enhuan Dong, Jiahai Yang 0001, Wenjian Zhang, Guanglei Song, Kexin Qiang, Hui Zhang 0141, Xiaowen Quan
IWQoS3
2026 F2D: Detection of resolver DNS hijacking based on filtration funnel strategy
abstract
Abstract In recent years, DNS hijacking represents a significant security threat to the infrastructure of the Domain Name System (DNS). A prevalent form of DNS hijacking involves exploiting open resolvers to manipulate DNS records. Such attacks undermine the availability and confidentiality of network services, posing serious risks to legitimate users. Current DNS hijacking detection methods tend to focus on specific domains, leveraging the unique characteristics of domain-specific hijacking to identify attacks. Consequently, these methods are often limited in applicability and may lack accuracy when dealing with diverse hijacking scenarios. Additionally, many existing approaches face challenges related to efficiency, making them less effective for long-term monitoring of hijacking activities. To address these challenges, this paper introduces an efficient detection method F2D tailored for general DNS hijacking. First, F2D uses an accurate and efficient filtration funnel strategy for targeted resolver hijacking detection. Second, two optimized detection algorithms are proposed for comprehensive filtration. Third, the method includes an efficient mechanism for identifying CDN domains, enabling the filtration of a large number of content replication servers and enhancing overall detection efficiency. During the validation phase, we monitor around 36k domains and around 600 resolvers over a one-month period. The effectiveness of our method is validated using manually labeled sample data. Experimental results demonstrate that our method can improve the F1 performance by 10% with the same false alert level, and time efficiency by 39% compared to the state-of-the-arts. Furthermore, we conduct an in-depth analysis of the captured hijacking incidents and deduce the motivation of the hijacking.
Cong Dong, Haoran Jiao, Jiahai Yang 0001, Chenglong Li 0006, Xia Yin 0001
Cybersecur.3
2026 Autonomous penetration testing using reinforcement learning: A review and perspectives
abstract
Penetration testing (pentesting) assesses cybersecurity through controlled, authorized attacks, but traditional manual methods demand considerable human and time resources. Reinforcement learning (RL), with its agent-environment interaction paradigm, offers a promising approach for autonomous pentesting. Despite remarkable advancements in this field, there is a lack of comprehensive reviews and perspectives on RL-based autonomous pentesting. To address this gap, this paper presents a systematic review of RL-based autonomous pentesting research. We outline the key challenges faced when applying RL in autonomous pentesting and categorize the existing literature into two main areas: attack path planning and autonomous pentesting frameworks, based on the research objectives and hypotheses. Additionally, we offer an in-depth analysis of the latest advancements and limitations in this field, while proposing a perspective on future research directions in the field of RL-based autonomous pentesting. We hope that our work will provide valuable insights for researchers, contributing to the advancement of autonomous pentesting and its practical application in the complex and diverse scenarios of the real world.
Jingju Liu, Yue Zhang 0049, Shicheng Zhou, Jiahai Yang 0001, Yuliang Lu, Xiaofeng Zhong
Expert Syst. Appl.4
2026 Alert2Vec: Eliminating Alert Fatigue by Embedding Security Alerts Through Subgraph Learning
Songyun Wu, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001
IEEE Trans. Dependable Secur. Comput.6
2026 HINHJ: Hierarchical Attention-Based Heterogeneous Graph Neural Network for DNS Hijacking Detection
abstract
The Domain Name System (DNS) is a critical internet infrastructure that translates human-readable domain names into machine-routable IP addresses. However, DNS is inherently vulnerable to manipulation, with hijacking attacks growing in both frequency and sophistication. Existing detection methods primarily rely on traffic analysis at specific network points. However, they suffer from limited coverage and low accuracy in complex environments, such as when CDN is employed. While recent approaches employ graph-based techniques, they still suffer from detection inaccuracy issues due to their failure to account for the complex interdependencies among multiple types of nodes. To address these limitations, we propose a novel heterogeneous graph-based detection framework. Based on the collected DNS records from distributed scanners, our method extracts activity and security features and constructs a heterogeneous graph to capture resolution patterns and cross-entity relationships. We further design a time-decay graph neural network TNHAN that enhances traditional Heterogeneous Graph Attention Networks (HAN) by dynamically weighting recent records. This network improves adaptability to legitimate DNS changes. For evaluation, we conduct experiments on real-world resolvers and domain datasets. Experiment results demonstrate the effectiveness of our method. Our method can achieve an F1-score of 0.96, outperforming the best baseline by 0.057 on average, and up to 0.113 under low label proportion. Moreover, we conduct several case studies on detected incidents, including cases related to geopolitical conflicts, censorship-related hijacking, and manipulation by malicious resolvers. These cases demonstrate the method’s effectiveness in identifying diverse hijacking behaviors in practice.
Haoran Jiao, Cong Dong, Chenglong Li 0006, Jiahai Yang 0001, Leyao Nie, Changzhi Zhao, Xia Yin 0001
IEEE Trans. Inf. Forensics Secur.4
2026 OwnerHunter: Multilingual Website Owner Identification Powered by Large Language Model
abstract
As cyberspace continues to expand, identifying the organization or individual behind a website has become increasingly vital in security incident response, phishing website detection, and other cybersecurity subfields. An existing solution for it involves analyzing webpage content and extracting owner names using named entity recognition techniques. However, since these techniques operate on a sentence-by-sentence basis, they struggle to identify the true owner when multiple individual or organizational names appear on a webpage. Moreover, they often perform poorly on non-English websites. To address these limitations, we propose OwnerHunter, a novel multilingual framework powered by large language models, which formulates website owner identification as a multilingual document-level information extraction task and utilizes global information from webpages to identify the owner. In OwnerHunter, we first craft prompts that fully leverage the capabilities of large language models to effectively recognize potential owners on webpages in different languages with minimal examples. To enhance the comprehensiveness and accuracy of recognition, we further design a multimodal augmentation strategy, an example pool strategy, and a self-verification strategy. Then, we devise a semantic and string similarity aggregation-based entity disambiguation technique to eliminate ambiguities among multiple potential owners recognized by large language models and a position-based hybrid ranking technique to exactly select the true owner. To evaluate OwnerHunter, we refine the publicly available English dataset ONER and construct the Chinese dataset WOI-cn with 16,036 real websites. Experimental results show that OwnerHunter achieves F1 scores of 0.9505 on ONER and 0.9621 on WOI-cn, setting new state-of-the-art performance on both datasets.
Cheng Tu, Enhuan Dong, Zexiang Zhang, Min Zhang 0054, Yang Li 0215, Jiahai Yang 0001
IEEE Trans. Inf. Forensics Secur.9
2026 AddrProbe: An Internet-Wide Active IPv6 Address Probing System With Limited Seeds
abstract
With the large-scale deployment of IPv6, it is becoming more and more important to probe active IPv6 addresses on the global Internet. However, the vast address space and the random distribution of active addresses make the probing process full of challenges, especially for the probing of IPv6 prefixes without seed addresses. Furthermore, the widespread existence of IPv6 aliased prefixes also causes significant trouble for probing. In this paper, we presentAddrProbe, an active IPv6 address probing system, which dynamically probes all global routing prefixes based on learned fine-grained address patterns from limited seed addresses and quickly detects aliased prefixes during probing. The evaluation results show thatAddrProbeachieves a hit rate of 23%-45% with all routing prefixes announced by the BGP system, which is 6.6-13× that of current state-of-the-art approaches (no more than 4%). Moreover, we find 1.2×1033aliased addresses characterized by the detected aliased prefixes, covering 6,412 routing prefixes, which is a 107× and 5.9× improvement over existing methods, respectively. Finally, an IPv6 Hitlist is constructed based on the long-term probing results, which contains 562M addresses covering 190K routing prefixes and 29K ASes. These widely distributed addresses are meaningful for analyzing IPv6 address assignments and some other IPv6 measurement activities.
Daguo Cheng, Lin He 0004, Qilei Yin, Guangxing Han, Boran Jin, Ying Liu 0024, Guanglei Song, Jinlong E, Tiankai Yang 0001, Jiahai Yang 0001
IEEE Trans. Netw.11
2025 ScannerGrouper: A Generalizable and Effective Scanning Organization Identification System Toward the Open World
abstract
In recent years, many scanning organizations deploy large numbers of scanners to actively probe the Internet. Identifying the organizations behind these scanners is of significant value. The problem of analyzing the sources of scanners has been investigated in various studies. However, as far as we know, the problem of effectively and generally identifying scanner organizations in real-world scenarios remains unsolved.
Enhuan Dong, Jiyuan Han, Hui Zhang 0141, Lianyi Sun, Supei Zhang, Guanglei Song, Xiaowen Quan, Jiahai Yang 0001
CCS13
2025 ZVDetector: State-Guided Vulnerability Detection System for Zigbee Devices
abstract
Nowadays, Zigbee devices are widely used in smart home, smart agriculture and other industries. However, there are many vulnerabilities in Zigbee devices that could compromise their normal functionality. Existing research either analyzes firmware or fuzzes devices through Zigbee networks to discover potential vulnerabilities. However, they overlook the impact of device state and protocol state on firmware or explore only a limited state space. Thus, they fail to identify many vulnerabilities caused by hidden states within each of the two states, especially vulnerabilities triggered by the combination of these two states. In this paper, we design a state-guided fuzzing system, named ZVDetector, aimed at uncovering firmware vulnerabilities caused by hidden and combined states. Specifically, we design two state-aware modules that explore richer unknown protocol state transitions based on message relationships and gain a more complete understanding of the intrinsic device state attributes. We develop a fuzzing algorithm that incorporates message semantics awareness and correlation state analysis. By integrating the perceived state information, it can explore the combined state space more efficiently. We validate the performance of ZVDetector on 10 Zigbee devices and find 25 vulnerabilities (19 zero-day). Our experiments also demonstrate the ability to explore more device state attributes and discover more message relationships related to unknown protocol states.
Chenglong Li 0006, Jiahai Yang 0001
CCS3
2025 TopoMiner: Efficient IPv6 Topology Discovery
abstract
Topology discovery can be used to obtain the connectivity and operational status of network devices by discovered interfaces. By performing topology discovery on networks, administrators can better understand and manage networks and improve network reliability and security. However, the large IPv6 address space, sparse address distribution, and unknown address assignment policies make it infeasible to simply and roughly probe the IPv6 network topology. To this end, we design an efficient IPv6 interface-level topology discovery method called TopoMiner. It performs multiple rounds of probing the IPv6 topology with the given set of IPv6 prefixes and locates the high-density interface regions based on the results of each round of topology discovery. TopoMiner then performs prefix expansion and address generation for the high-density interface address regions. At the same time, TopoMiner also uses some of the prefixes that have been probed to regenerate the target addresses. In this way, TopoMiner can dynamically adjust the breadth and depth of topology discovery and thus complete multiple rounds of topology discovery within the packetsending budget. In real-word probing, TopoMiner achieves a$3 \sim 7 \times$enhancement in probing efficiency compared to state-of-the-art topology discovery methods.
Hongwei Li 0021, Lin He 0004, Guanglei Song, Daguo Cheng, Jiahai Yang 0001, Ying Liu 0024
ICC7
2025 Lightning in the Dark: Uncovering Global IPv6 Router Interfaces and Their Security Implications
abstract
The IPv6 routing infrastructure is an important part of the modern Internet, and the collection of its interface addresses is greatly significant in network security, performance optimization, and measurement analysis. However, existing methods suffer from two major problems: the lack of flexibility in budget allocation across probing rounds and the absence of a dynamic hop limit adjustment mechanism based on feedback. These problems lead to the low hit rate and inefficiency of existing methods for discovering router interfaces, which seriously hinders the comprehensive knowledge of IPv6 routing infrastructure.To this end, we propose Helixir, a feedback-based, high hit-rate, and efficient IPv6 router interface discovery system. Helixir’s core design includes a dynamic budget allocation mechanism across probing rounds, an inter-prefix budget allocation strategy that adequately trades off exploration and exploitation, and a hop limit selection method based on Thompson sampling. Real-world experiments show that with a 100M budget, Helixir achieves a hit rate 3.64× that of state-of-the-art methods on the BGP prefixes dataset, and Helixir successfully discovers over 31 million IPv6 router interface addresses in total within half an hour. In addition, a systematic security analysis of the discovered router interfaces shows that many devices open sensitive ports and expose hundreds of potential CVE vulnerabilities, highlighting the security risks in the IPv6 network.
Ying Liu 0024, Lin He 0004, Xiaoyi Shi, Yifan Yang 0009, Chentian Wei, Daguo Cheng, Jiahai Yang 0001
ICNP9
2025 Poster: TopoHunter: Enabling Efficient and High-Coverage Active IPv6 Topology Discovery
abstract
We introduce TopoHunter, an efficient IPv6 Internet topology discovery system. The central concept of TopoHunter is to allocate more probing resources to target prefix spaces that yield greater topological benefits, as well as to their surrounding areas. To achieve this, we design a feedback-based target generation module comprised of a Target Prefix Probing Value Forest that maintains the estimated probing values of hierarchical target prefix spaces. Our system has successfully discovered the most extensive and complete IPv6 topology map to date, comprising over 144 million router interfaces and 251 million edges, covering 72.83% of autonomous systems and 43.36% of routing prefixes announced by the BGP system.
Lin He 0004, Hongwei Li 0021, Guanglei Song, Wentong Wang, Daguo Cheng, Enhuan Dong, Chenglong Li 0006, Hui Zhang 0141, Jinlong E, Ying Liu 0024, Jiahai Yang 0001
IMC13
2025 6RV: Incremental Learning-Based Continuous Identification of IPv6 Router Vendors
abstract
The growth of IPv6 networks has led to an expanding number of network routers, but there is not enough research on vendors of these devices. Existing router vendor identification algorithms are based on IPv4, and these static analysis algorithms cannot adapt to dynamically changing IPv6 networks. In this paper, we develop 6RV, a framework for IPv6 router vendors continuous identification based on incremental learning. First, we count the addresses of newly discovered router interfaces every month and obtain router device fingerprints through active probing, and then analyze these data fingerprints through an incremental learning approach to identify the router vendors of new nodes. We validate our identification framework on the ITDK dataset over a period of 8 months, obtaining more than 500K router vendor labels with 94% correctness. Finally, we also analyze the IPv6 router vendor dataset from different perspectives and draw some interesting conclusions.
Shenao Li, Jiahai Yang 0001, Enhuan Dong, Chenglong Li 0006, Lin He 0004, Hui Zhang 0052, Guanglei Song
NOMS3
2025 LMGeo6: A Comprehensive IPv6 Landmark Mining Methodology to Facilitate IPv6 Geolocation
abstract
IP geolocation is essential for various location-aware Internet applications. With the rapid development of the IPv6 protocol, IPv6 geolocation is becoming increasingly important. High-accuracy IPv6 geolocation relies heavily on high-quality landmarks. However, existing landmark mining methods mainly focus on IPv4 landmarks, and cannot mine IPv6 landmarks or are extremely inefficient, making IPv6 geolocation accuracy hard to improve. In this paper, we present a novel IPv6 geolocation framework, LMGeo6, proposing a comprehensive IPv6 landmark mining methodology for high-accuracy IPv6 geolocation. LMGeo6 migrates IPv4 landmark mining methods and designs specific IPv6 landmark mining methods, improving the number and coverage of IPv6 landmarks significantly. Based on these designs, we implement the LMGeo6 system and conduct real-world experiments to validate its efficacy. The experiment results show that LMGeo6 can mine 62,697 IPv6 landmarks, improving 15× over the total of other methods. LMGeo6 landmarks cover 165 countries, 5,805 cities, and 4,030 ASes, improving over 300% and reaching the same scale as IPv4 landmarks.
Jinlei Lin, Chenglong Li 0006, Hui Zhang 0052, Wentong Wang, Jiahai Yang 0001
NOMS6
2025 Post-Standardization Analysis of DoQ: Deployment, Certificates Ecosystem and Implementation
abstract
To address the security issues caused by traditional plaintext DNS transmission, encrypted protocols were introduced to protect DNS traffic. DNS over QUIC (DoQ) is the most recent DNS encryption protocol standardized in 2022. While earlier protocols like DoT and DoH have been extensively studied, research on DoQ remains limited, focusing primarily on basic deployment and performance. There is a lack of comprehensive research on the DoQ ecosystem after its standardization, and the compliance and security of its deployment remain unclear. This paper presents the first in-depth measurement of DoQ deployment across IPv4, IPv6, and authoritative servers. Our findings offer an early view of the DoQ ecosystem, covering its deployment, certificate ecology, and practical implementations. Overall, the progress of DoQ standardization is satisfactory. Since standardization, DoQ adoption has tripled, and its certificate ecosystem shows a promising trend, with fewer than 10% of certificates being invalid. However, potential security concerns persist. First, the centralization issue in DoQ is more pronounced compared to DoH and DoT. Second, about 30% of DoQ authority servers support recursive parsing, facing the risk of cache poisoning or DDoS attacks. In addition, 2% of DoQ deployments fail to meet RFC requirements, potentially enabling amplification attacks. Therefore, we highlight the need for stricter compliance with standards in future DoQ implementations to enhance security and reliability.
Chenglong Li 0006, Wenchong Dong, Cong Dong, Jiahai Yang 0001, Hui Zhang 0052
NOMS5
2025 ALM: A Two-Stage Traffic Anomaly Detection and Analysis System via the Large Language Model
abstract
In recent years, deep learning-based traffic anomaly detection has proven very promising. Although current methods achieve high accuracy in detecting anomalies, they struggle to accurately classify attack types of anomalies due to the imbalanced distribution of attack samples. To address the issue, we propose a highly intelligent system, ALM, which can simultaneously provide accurate traffic anomaly detection and attack types classification with the aid of the Large Language Model (LLM)'s few-shot learning ability. To tackle the challenges of training cost and inference efficiency associated with large models, ALM adopts a two-stage solution, i.e., AnomalyDetector and Anomaly Analyzer, that combines the fine-tuned LLM with small models. In the first stage, AnomalyDetector ensembles a set of lightweight models to handle high-concurrency real-time network traffic anomaly detection. In the second stage, Anomaly Analyzer leverages the LLM's powerful fitting and few-shot learning abilities for traffic anomaly analysis through three processes: LLM task adaption, traffic to sequence, and LLM fine-tuning. This allows Anomaly Analyzer to accurately identify the attack types and potential false positives. Experimental results indicate that ALM achieves over 90% Micro-F1 on four public datasets, with a maximum of 99.94 %, surpassing the baseline. Additionally, it requires minimal training costs while significantly improving inference efficiency compared to the pure LLM mode.
Songyun Wu, Enhuan Dong, Haina Hu, Jiahai Yang 0001
NOMS7
2025 Hawkeye: Diagnosing RDMA Network Performance Anomalies with PFC Provenance
abstract
RDMA is becoming increasingly prevalent from private data centers to public multi-tenant clouds, due to its remarkable performance improvement. However, its lossless traffic control, i.e., PFC, introduces new complexities in network performance anomalies (NPAs) due to its cascading congestion spreading property, which usually incurs complaints from customers/applications about certain flows' performance degradation. Existing studies fall short in fine-grained visibility of PFC impact and traceability of PFC causality, and are thus ineffective in diagnosing the root causes for RDMA NPAs. In this paper, we propose Hawkeye, an accurate and efficient RDMA NPA diagnosis system based on PFC provenance. Hawkeye comprises 1) a fine-grained PFC-aware telemetry mechanism to record the PFC impact on flows; 2) an in-network PFC causality analysis and tracing mechanism to quickly and efficiently collect causal telemetry for diagnosis; and 3) a provenance-based diagnosis algorithm to comprehensively present the anomaly breakdown, identifying the anomaly type and root causes accurately. Through extensive evaluations on both NS-3 simulations and a Tofino testbed, Hawkeye can quickly and accurately diagnose multiple RDMA NPAs with over 90% precision and 1–4 orders of magnitude lower overhead than baselines.
Menghao Zhang 0001, Xiao Li 0044, Qiyang Peng, Mingwei Xu 0001, Xiaohe Hu, Jiahai Yang 0001, Xingang Shi
SIGCOMM9
2025 IPdb: A High-Precision IP Level Industry Categorization of Web Services
abstract
IP addresses with web services are crucial in the Internet ecosystem. Classifying these addresses by industry and organization offers valuable insights into the entities utilizing them, enabling more efficient network management and enhanced security. Previous work in website classification and Internet management struggles to offer an IP-level perspective of the industries of web services due to their limited industry categories or potential industry inconsistencies between IP address owners and AS owners. To this end, we present IPdb, an IP-level industry categorization dataset. To construct the dataset, we developed LLMIC, a Large Language Model-based Industry Categorization framework with a precision of nearly 96%. IPdb serves as a labeled database for future endeavors in developing IP-level industry classifiers, encompassing over 200 million IP addresses. Furthermore, our study indicates that 30% ~ 50% of organizations within critical infrastructure industries deploy web servers across multiple ASes. Our study also validates the problem of mismatched granularity in industry categorization at the AS level with 87.83% ASes in IPv4 and 72.96% ASes in IPv6 containing IP addresses from different industries.
Guanglei Song, Jiahai Yang 0001, Songyun Wu, Jinlei Lin, Lin He 0004, Chenglong Li 0006
WWW4
2025 Adaptive traffic engineering with segment routing through deep reinforcement learning
Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009
Comput. Networks5
2025 E-DoH: elegantly detecting the depths of open DoH service on the internet
abstract
Abstract In recent years, DoE methods have been regarded as a novel trend within the realm of the DNS ecosystem. Measuring these DoE services in the wild can promote improvements in DoE methods and facilitate their widespread adoption. A primary requirement for measuring DoE methods is the discovery of these services. The discovery is relatively straightforward for DoT and DoQ, but complex for DoH since it shares port 443 with web services as suggested in RFC 8484. Although previous works primarily analyze the surface of the DoH service, they (1) result in long detection time and large traffic volume by adopting an enumeration strategy to discover the DoH service; (2) lack an in-depth analysis of the status of upper-layer DNS services. In this paper, we propose the E-DoH method for elegant, efficient, and in-depth DoH service measurement. First, we propose a measurement mechanism to enable a single DoH connection to accomplish multiple tasks including service discovery, correctness validation, and dependency construction with minimal backend configuration. Second, we propose a dynamic protocol negotiation strategy to enhance probing efficiency while significantly reducing the required traffic volume. Based on the above optimization methods, we conducted an exploration of the IPv4 space and performed an in-depth analysis of DoH based on the collected information. Through experiments, our approach demonstrates a remarkable 80% improvement in time efficiency and only requires 4–20% traffic volume to complete the detection task. In wild detection, our approach discovered 46k DoH services, which nearly doubles the number discovered by the state-of-the-art. This indicates the growing trend of DoH services. Based on the collected information, we present several intriguing conclusions about the current DoH service ecosystem.
Cong Dong, Jiahai Yang 0001, Haoran Jiao, Chenglong Li 0006, Xia Yin 0001
Cybersecur.2
2025 SCRIPT: A Scalable Continual Reinforcement Learning Framework for Autonomous Penetration Testing
Shicheng Zhou, Jingju Liu, Yuliang Lu, Jiahai Yang 0001, Yue Zhang 0049, Bo Lin 0011, Xiaofeng Zhong, Shulong Hu
Expert Syst. Appl.4
2025 Mind the Gap: towards generalizable autonomous penetration testing via domain randomization and meta-reinforcement learning
abstract
With the increasing number of vulnerabilities exposed on the Internet, autonomous penetration testing (pentesting) has emerged as a promising research area. Reinforcement learning (RL) is a natural fit for studying this topic. However, two key challenges limit the applicability of RL-based autonomous pentesting in real-world scenarios: the training environment dilemma—training agents in simulated environments is sample-efficient while ensuring that their realism remains challenging; poor generalization ability—agents’ policies often perform poorly when transferred to unseen scenarios, with even slight changes potentially causing a significant generalization gap. To address both challenges, we propose a generalizable autonomous pentesting framework termed GAP, which aims to achieve efficient policy training in realistic environments and train generalizable agents capable of drawing inferences about other cases from one instance. GAP introduces a real-to-sim-to-real pipeline that enables end-to-end policy learning in unknown real environments while constructing realistic simulations and improves agents’ generalization ability by leveraging domain randomization and meta-RL learning. We are among the first to apply domain randomization in autonomous pentesting and propose a large language model-powered domain randomization method for synthetic environment generation. We further apply meta-RL to improve agents’ generalization ability in unseen environments by leveraging synthetic environments. Combining the two methods effectively bridges the generalization gap and improves agents’ policy adaptation performance. Simulations are conducted on various vulnerable virtual machines, with results showing that GAP can enable policy learning in various realistic environments, achieve zero-shot policy transfer in similar environments, and achieve rapid policy adaptation in dissimilar environments.
Shicheng Zhou, Jingju Liu, Yuliang Lu, Jiahai Yang 0001, Yue Zhang 0049, Jie Chen 0079
Frontiers Inf. Technol. Electron. Eng.4
2025 APRIL: Towards Scalable and Transferable Autonomous Penetration Testing in Large Action Space via Action Embedding
abstract
Penetration testing (pentesting) assesses cybersecurity through simulated attacks, while the conventional manual-based method is costly, time-consuming, and personnel-constrained. Reinforcement learning (RL) provides an agent-environment interaction learning paradigm, making it a promising way for autonomous pentesting. However, agents’ scalability in large action spaces and policy transferability across scenarios limit the applicability of RL-based autonomous pentesting. To address these challenges, we present a novel autonomous pentesting framework based on reinforcement learning (namely APRIL) to train agents that are scalable and transferable in large action spaces. In APRIL, we construct realistic, bounded, host-level state space via embedding techniques to avoid the complexities of dealing with unbounded network-level information. We employ semantic correlations between pentesting actions as prior knowledge to represent discrete action space into a continuous and semantically meaningful embedding space. Agents are then trained to reason over actions within the action embedding space, where two key methods are applied: an upper-confidence bound-based action refinement method to encourage efficient exploration, and a distance-aware loss to improve learning efficiency and generalization performance. We conduct experiments in simulated scenarios constructed based on virtualized vulnerable environments. The results demonstrate APRIL's scalability in large action spaces and its ability to facilitate policy transfer across diverse scenarios.
Shicheng Zhou, Jingju Liu, Yuliang Lu, Jiahai Yang 0001, Dongdong Hou, Yue Zhang 0049, Shulong Hu
IEEE Trans. Dependable Secur. Comput.4
2025 End-to-End Attack Scene Reconstruction in a Host With Rules and Anomaly-Based Detection Models
abstract
Critical devices on the Internet are frequently targeted by skilled and advanced network attackers. These attackers often orchestrate complex and persistent intrusion campaigns, which involve multiple stages of attacks. In the context of host-based threat detection, the reconstruction of the entire attack scenario is crucial for tracing threats and fixing system vulnerabilities. Prior anomaly-based studies lack the capability to interpret the attack scenario, while rule-based approaches struggle with detecting novel attack patterns. We introduce eaGle, an end-to-end framework that takes original host-based data as input and reconstructs the potential attack scenario as output. It leverages an anomaly-based algorithm and a fine-grained misuse detection module to assign anomalous scores to host data, constructs the potential attack scenario using a novel anomalous subtree detection algorithm, and generates the interpretable attack scenario graph through a coarse-grained rule matching method. We assess the performance of eaGle using three attack scenarios from the DARPA TC dataset and three deployment scenarios. The results demonstrate that eaGle can effectively uncover the hidden attack scenario within the host data and outperforms three state-of-the-art attack scenario reconstruction systems.
Xia Yin 0001, Han Zhang 0009, Xingang Shi, Jiahai Yang 0001
IEEE Trans. Inf. Forensics Secur.9
2025 Centralized Network Utility Maximization With Accelerated Gradient Method
abstract
Network utility maximization (NUM) is a fundamental problem for network traffic management and resource allocation. Due to the inherent decentralization and complexity of networks, much of the existing research has focused on developing decentralized algorithms for NUM. However, with the rise of Software-Defined Networking (SDN), especially in cloud networks and inter-datacenter networks managed by large enterprises, there has been growing interest in centralized NUM algorithms. To cope with the large and increasing number of flows in such SDN networks, existing studies on centralized NUM focus on the scalability of the algorithm with respect to the number of flows, but the efficiency is ignored. In this paper, we propose a centralized, efficient and scalable algorithm for the NUM problem. By designing smooth utility and penalty functions, we formulate the NUM problem with a smooth objective function, which enables the use of Nesterov’s accelerated gradient method (AGM). We prove that the proposed method achieves an$O(d/t^{2})$convergence rate, demonstrating superior convergence speed with respect to the number of iterationst, and our method is scalable with respect to the number of flowsdin the network. Our smooth objective NUM formulation and AGM are effective not only in simple network scenarios with non-prioritized flows routed on one simple paths, but also in more complex and practical scenarios involving prioritized flows routed across multiple complex paths. Experiment results confirm that our method obtains accurate solutions with fewer iterations, and achieves close-to-optimal network utility.
Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009
IEEE Trans. Netw.5
2024 Rules Refine the Riddle: Global Explanation for Deep Learning-Based Anomaly Detection in Security Applications
abstract
Deep learning (DL) based anomaly detection has shown great promise in the field of security due to its remarkable performance in various tasks. However, the issue of poor interpretability in DL models has significantly impeded their deployment in practical security applications. Despite the progress made in existing studies on DL explanations, the majority of them focus on providing local explanations for individual samples, neglecting the global understanding of the model knowledge. Furthermore, most explanations for supervised models fail to apply to anomaly detection due to their different learning mechanisms.
Minghui Jin, Jiahai Yang 0001, Xingang Shi, Xia Yin 0001, Yang Liu 0003
CCS8
2024 ChatScam: Unveiling the Rising Impact of ChatGPT on Domain Name Abuse
abstract
Since 2022, ChatGPT has been a big breakthrough in technology, creating lots of discussions online. It has had big effects in different areas, but in cybersecurity, it is both good and bad. There has been a lot of misuse, especially with squatting domains. Our research aims to understand this misuse and the potential threats it poses. We develop a novel method that looks at historical Passive DNS (PDNS) data. Based on the two-stage identification, our method can efficiently and accurately collect ChatGPT-related squatting domains. In the end, we found over 1.3 million ChatGPT-related squatting domains, part of which were shared with the security community. Our findings show that these squatting domains are increasing quickly. This is the case whether the keywords related to ChatG PT are registered with the domain registrar or set up on sub domains. Even though the number of domains is increasing, only 5.3 % set up meaningful content on their websites. After digging into their web contents, we found that these web sites show various signs of misuse, such as promotion on illegal underground websites and emerging fraudulent activities related to dialogue features. The security community is not fully aware of these threats yet. We are the first to conduct a large-scale quantitative analysis of ChatGPT-related abusive behavior. We believe that our work unveils the abuse ecosystem surrounding ChatGPT-related squatting domains. We hope to underscore the urgent need for increased attention and protective measures against ChatGPT-related domain abuse.
Mingxuan Liu 0006, Zhenglong Jin, Jiahai Yang 0001, Baoiun Liu, Hai-Xin Duan, Ying Liu 0024, Ximeng Liu, Shujun Tang
DSN3
2024 6Vision: Image-Encoding-Based IPv6 Target Generation in Few-Seed Scenarios
abstract
Efficient global Internet scanning is crucial for network measurement and security analysis. While existing target generation algorithms verify remarkable performance in largescale detection, their efficiency notably diminishes in few-seed scenarios. This decline is primarily attributed to the intricate configuration rules and sampling bias of seed addresses. Moreover, instances where BGP prefixes have few seed addresses are widespread, constituting$63.65 \%$of occurrences. We introduce 6 Vision to tackle this challenge by introducing a novel approach to encoding IPv6 addresses into images, facilitating comprehensive analysis of intricate configuration rules. Through feature stitching, 6 Vision not only improves the learnable features but also amalgamates addresses associated with configuration patterns for enhanced learning. Moreover, it integrates an environmental feedback mechanism to refine model parameters based on identified active addresses, thereby alleviating the sampling bias inherent in seed addresses. As a result, 6Vision achieves high-accuracy detection even in few-seed scenarios. The HitRate of 6 Vision is improved by$181 \% \sim 2,490 \%$compared to existing algorithms, while the CoverNum is$1.18 \sim 11.20$times that of them. Additionally, 6Vision can function as a preliminary detection module for existing algorithms, yielding a conversion gain (CG) ranging from$242 \% \sim 2,081 \%$. Ultimately, we achieve a conversion rate (CR) of$28.97 \%$for few-seed scenarios. We enrich the IPv6 hitlist, not only enhancing current target generation algorithms for large-scale address detection in few-seed scenarios but also effectively supporting IPv6 network measurement and security analysis.
Wenjian Zhang, Guanglei Song, Lin He 0004, Jinlei Lin, Songyun Wu, Chenglong Li 0006, Jiahai Yang 0001
ICNP8
2024 CloudPlanner: Minimizing Upgrade Risk of Virtual Network Devices for Large-Scale Cloud Networks
abstract
Cloud networks continuously upgrade softwarized virtual network devices (VNDs) to meet evolving tenant demands. However, such upgrades may result in unexpected failures. An intuitive idea to prevent upgrade failures is to resolve all compatibility issues before deployment, but it is impractical to replicate all deployed VND cases and test them with lots of replayed real traffic for the VND developers. As a result, the operations team takes upgrade risk to test upgrades by gradually deploying them. Although careful upgrade schedule planning is the most common method to minimize upgrade risk, to the best of our knowledge, no VND upgrade schedule planning scheme has been adequately studied for large-scale cloud networks. To fill this gap, we propose CloudPlanner, the first VND upgrade schedule planning scheme aiming to minimize the VND upgrade risk for large-scale cloud networks. CloudPlanner prioritizes upgrading VNDs that are more likely to trigger failures based on expert knowledge and historical failure-trigger VND properties and limits the number of tenants associated with simultaneously upgraded VNDs. We also propose a heuristic solver which can quickly and greedily plan schedules. Using real-world data from production environments, we demonstrate the benefits of CloudPlanner through extensive experiments.
Enhuan Dong, Jiahai Yang 0001, Shize Zhang, Zejie Wang, Xiaoqing Sun, Enge Song, Jianyuan Lu, Biao Lyu, Shunmin Zhu
INFOCOM3
2024 CP-IoT: A Cross-Platform Monitoring System for Smart Home
Chenglong Li 0006, Jiahai Yang 0001, Linna Fan, Chenxin Duan
NDSS3
2024 LoRDMA: A New Low-Rate DoS Attack in RDMA Networks
Menghao Zhang 0001, Yuying Du, Ziteng Chen, Mingwei Xu 0001, Renjie Xie, Jiahai Yang 0001
NDSS8
2024 TrafficSiam: More Realistic Few-shot Website Fingerprinting Attack with Contrastive Learning
abstract
Website fingerprinting (WF) attacks pose a serious threat to users’ online privacy, even when using privacy-enhancing tools like Tor. Previous attacks mostly rely on supervised learning and only a few studies have explored the few-shot setting which is more realistic. In this paper, we propose TrafficSiam, a novel WF attack based on few-shot learning with self-supervised learning, which enables our model to be pretrained with unlabeled tor traffic and transferred to new tasks using a small number of labeled samples which is more realistic and yield stronger generalization ability. We conduct a series of experiments, using only a small amount of labeled samples, and find that our model achieves 92.32% accuracy in the closed-world setting, compared to the highest accuracy 88.73%, using previous methods. Furthermore, our model also outperforms previous attacks in the open-world setting.
Shangdong Wang, Chenglong Li 0006, Jiahai Yang 0001, Hui Zhang 0052
NOMS5
2024 AudiTrim: A Real-time, General, Efficient, and Low-overhead Data Compaction System for Intrusion Detection
abstract
Recently enterprises and governments face escalating APT attacks, leading to significant economic losses. APT attacks often persist for extended periods, necessitating the storage of extensive audit logs for effective detection. To reduce data storage overhead, enterprises commonly adopt compression strategies. However, efficient compression strategies may introduce additional query overhead. Existing approaches propose data reduction algorithms, but these methods can compromise data integrity, rendering current attack investigation and anomaly-based intrusion detection ineffective.
Zheyu Jiang, Jiahai Yang 0001
RAID6
2024 Network anomaly detection via similarity-aware ensemble learning with ADSim
abstract
The last decade has seen the increasing application of machine learning to various tasks, including network anomaly detection . But anomaly detection methods based on a single machine learning algorithm usually fail to achieve good results, since network traffic have complex and changeable patterns. Therefore, many solutions based on ensemble learning have been proposed to address this problem. However, most previous studies have the main drawback that they overlook the similarity between the weak classifiers , which may degrade the detection performance. What is more, most existing works use offline and supervised algorithms, which means a large number of computing resources and reliable labels are necessary during the training period. In this paper, we propose ADSim , an online, unsupervised, and similarity-aware network anomaly detection algorithm based on ensemble learning. For a similarity-aware scheme, the target of ADSim can be intuitively described as recognizing the similar weak classifiers during the training phase and treat them as a whole. To achieve this, ADSim first incrementally maintains a distance matrix to record the similarity between the classifiers in the training phase and uses Hierarchy Clustering to group the similar classifiers. In the detecting phase, each cluster will be assigned a weight depending on the consistency of the detection results of the classifiers within it. Moreover, the working procedure of ADSim is online and unsupervised, which significantly improves its practicality. We test ADSim on two datasets, MAWILab and CIC-IDS-2017. The results show that ADSim outperforms the state-of-the-art ensemble learning methods and has ideal runtime performance.
Liyuan Chang, Ying Zhong 0008, Chenxin Duan, Xia Yin 0001, Jiahai Yang 0001, Xingang Shi
Comput. Networks9
2024 IoTa: Fine-Grained Traffic Monitoring for IoT Devices via Fully Packet-Level Models
abstract
With Internet-of-Things (IoT) devices gaining popularity, dedicated monitoring systems which accurately detect intrusion traffic for them are in high demand. Existing methods mainly use statistical spatial-temporal traffic features and machine learning models. Their practicality has been limited due to the lack of detection ability for stealthy and tricky attacks, diagnostic utility and long-term performance. To address these problems and motivated by the simplicity of mini IoT devices, we propose to construct fully packet-level models to profile traffic patterns for IoT devices by constructing automaton for short flow and long flow, where the length and direction of each packet are the representative features. We apply these fine-grained models to design and develop a traffic monitoring system, namelyIoTa, to detect intrusion traffic for IoT devices.IoTamatches the ongoing traffic with patterns extracted from normal traffic traces. With visible and interactive traffic profiles,IoTacan generate interpretable alerts and is available for long-term use under reasonable human efforts. Evaluations on dozens of common IoT devices show thatIoTacan achieve excellent detection accuracy (nearly perfect recalls and always over 0.999 precisions) for various intrusion traffic covering the complete kill chains. Incorrect detection results can be compensated for by error recovery mechanisms and the understandable alert context can be used by the operator to enhance the system. The diagnostic utility and little alert weariness are recognized by the experienced operators.
Chenxin Duan, Sainan Li, Guanglei Song, Chenglong Li 0006, Jiahai Yang 0001
IEEE Trans. Dependable Secur. Comput.7
2024 RFG-HELAD: A Robust Fine-Grained Network Traffic Anomaly Detection Model Based on Heterogeneous Ensemble Learning
abstract
Fine-grained attack detection is an important network security task. A large number of machine learning/deep learning( ML/DL) based algorithms have been proposed. However, attacks not present in the training set pose a challenge to the model (openset problem). Further, ML/DL based models face the problem of adversarial attacks. Despite the large amount of work attempting to address these problems, there are still some challenges as follows. First, the open-set problem in fine-grained attack detection is difficult to solve because there is no effective representation of the distribution of unknown attacks. Second, in the open set environment, how the fine-grained attack detection model resists the adversarial attack is a more difficult problem. For example, the presence of unknown attacks poses a challenge for adversarial defense. For these reasons, we propose the RFG-HELAD model, which consists of aKclassification model based on deep neural network (DNN) with contrastive learning (CL), and aK+ 1 classification model combining a generative adversarial networks (GAN) with two discriminators and deepk-nearest neighbors (Deep kNN). Among them, Deep kNN uses latent features from GAN and contrastive learning as input, which is essentially a distance-based out-of-distribution detection algorithm used to determine unknown attacks. The large category of unknown attacks has been added to theKclassification, so it is aK+ 1 classification. To further improve the robustness of the RFG-HELAD model, we perform Fourier transform as well as feature fusion on the features, and also conduct adversarial training on theKclassification model. Generative adversarial training of our GAN model can implicitly defend against adversarial attack. Experiments show that our model is superior to other state-of-the-art (SOTA) models in the presence of unknown attacks as well as under adversarial attacks. Especially, our model improves the accuracy by at least 18.7% over the corresponding SOTA model with adversarial defense. Further, we discuss the grounded deployment of the model and demonstrate its feasibility.
Ying Zhong 0008, Xingang Shi, Jiahai Yang 0001, Keqin Li 0001
IEEE Trans. Inf. Forensics Secur.4
2024 ProbeGeo: A Comprehensive Landmark Mining Framework Based on Web Content
abstract
IP geolocation is essential for various location-aware Internet applications. High-quality IP geolocation landmarks play a decisive role in IP geolocation accuracy. However, the previous research works focusing on mining landmarks from the Internet are hampered by limited quantity, poor coverage, and insufficient landmark quality. In this paper, we present a new framework called ProbeGeo to mine high-quality landmarks automatically. We divide landmarks into common landmarks and probe landmarks, providing systematic mining methods based on online retrieval and web content. ProbeGeo expands traditional common landmarks by taking advantage of the exposure of multiple IoT (Internet of Things) devices on the Internet, mining them based on search engines and webpage contents. Common landmarks, consisting of multi-type devices, significantly improve landmark quantity and coverage. Furthermore, ProbeGeo establishes a methodology for acquiring new probe landmarks from Internet VPs (Vantage Points) webpages, extracting geographical locations from heterogeneous webpages and utilizing active probe functions. Probe landmarks enhance landmark quality and functions, bringing new geolocation frameworks and breaking through the geolocation accuracy bottleneck. We develop the ProbeGeo as a continuously running system and conduct real-world experiments to validate its efficacy. Our results show that ProbeGeo can detect 89,849 high-quality landmarks, including 6,874 probe landmarks and 82,975 common landmarks. ProbeGeo landmarks are about 10x more than existing work, distributed in 181 countries and 7,094 cities. ProbeGeo landmarks cover more than 8 types of devices, and more than 60% of them remain stable over one month. Moreover, the landmark accuracy of more than 58% of ProbeGeo landmarks is above street level, which has not been achieved in previous works. ProbeGeo can provide geolocation services with higher landmark accuracy and broader coverage by correlating a large scale of landmarks.
Jinlei Lin, Chenglong Li 0006, Guanglei Song, Linna Fan, Jiahai Yang 0001
IEEE/ACM Trans. Netw.7
2024 PMap: Reinforcement Learning-Based Internet-Wide Port Scanning
abstract
Internet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the most open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap performs better on hit rate, coverage, and intrusiveness. Our experiments over real networks show that PMap can find 90% open ports by only scanning 125 ports (90%@125) to each address, which is 99.3% less than the state-of-the-art port scanning methods. It reduces the number of scanned ports to decrease the intrusive nature of port scanning. In addition, PMap is highly parallel and lightweight. It scans 500 networks in parallel, achieving a port recommendation rate of up to 18 million per second, consuming only 7GB of memory. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research.
Guanglei Song, Lin He 0004, Jinlei Lin, Linna Fan, Jiahai Yang 0001
IEEE/ACM Trans. Netw.8
2024 AddrMiner: A Fast, Efficient, and Comprehensive Global Active IPv6 Address Detection System
abstract
Fast Internet-wide scanning is essential for network situational awareness and asset evaluation. However, the vast IPv6 address space makes brute-force scanning infeasible. Despite advancements in state-of-the-art methods, they do not work in seedless regions and suffer low detection efficiency and speed in regions with known active IPv6 addresses (i.e., seed addresses). Moreover, the collected active address list (i.e., IPv6 hitlist) with low coverage cannot truly represent the active IPv6 address landscape of the Internet. This paper introduces AddrMiner, a fast, efficient, and comprehensive global active IPv6 address detection system. We design a systematic active IPv6 address detection strategy that divides the IPv6 space into two detection scenarios based on the presence or absence of seed addresses to discover active IPv6 addresses from scratch and from few to many. In the seedless regions, we present AddrMiner-N, leveraging a multi-level association policy to probe active addresses. It fills the gap of address detection in seedless regions and successfully discovers active addresses in 39,899 BGP prefixes without seed addresses, with a$1.03\times $higher hit rate,$30\sim 911\times $higher speed, and$2.7\times $broader coverage, compared to existing solutions. In the regions with seed addresses, our method AddrMiner-S dynamically generates target addresses using reinforcement learning. Compared to state-of-the-art methods, AddrMiner-S achieves an impressive 56.3% hit rate and a discovery speed of 839.0/s, which is$1.9\sim 2153\times $and$1.5\sim 755\times $of existing works, respectively. Finally, we deploy AddrMiner and discover 2.1B active IPv6 addresses, including 1.7B de-aliased active addresses and 0.4B aliased addresses, through continuous probing for three years.
Guanglei Song, Lin He 0004, Feiyu Zhu 0002, Jinlei Lin, Wenjian Zhang, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001
IEEE/ACM Trans. Netw.9
2024 Proactive Telemetry in Large-Scale Multi-Tenant Cloud Overlay Networks
abstract
At present, public clouds have served millions of tenants. To provide reliable services, cloud vendors need to perceive health status of the cloud network by building a telemetry system to detect possible network failures. While telemetry systems for physical networks have been extensively studied, research on telemetry systems for virtual networks is still insufficient. Different from physical networks, we conclude that building a virtual network telemetry system faces new challenges of feasibility, efficiency, and effectiveness. Specifically, we need to 1) protect privacy of tenants and adapt to heterogeneous middleboxes at the data plane; 2) handle frequent virtual network topology updates and compress large-scale measurement paths for millions of tenants at the control plane; 3) analyze telemetry results to locate network failures at the analysis plane. To address these challenges, we present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. At the data plane, Zoonet uses host agent and arp-ping to protect tenants’ privacy and defines an elegant generalization of ping and traceroute, which can work on heterogeneous middleboxes. At the control plane, Zoonet conducts update batch processing and substantial probing path pruning to lessen the overhead. At the analysis plane, Zoonet reduces noises and aggregates alerts based on temporal and spatial correlation and conducts the hop-by-hop telemetry mode to locate failures. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Shize Zhang, Xiaoqing Sun, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Xihui Yang, Jiahai Yang 0001
IEEE/ACM Trans. Netw.15
2024 CouldPin-Fast: Effient and Effective Root Cause Localization for Shared Bandwidth Package Traffic Anomalies in Public Cloud Networks
abstract
As cloud services become increasingly widespread, many public cloud tenants opt for Shared Bandwidth Package (sBwp) services for inbound/outbound communication. The sBwp service allows tenants to purchase shared bandwidth for multiple virtual machines (VMs) instead of buying it individually, which is a convenient and cost-effective traffic management mode. However, the sBwp service presents new challenges for operators to identify the root cause of abnormal sBwp traffic, especially in large-scale, globally distributed public clouds with millions of users. Developing a localization system in public cloud faces several challenges, including dynamic scalability, hyper-scale data efficiently obtaining, and complex application scenarios. To address these challenges, we propose a two-stage localization method calledCloudPin-Fast. First,CloudPin-Fastemploys a cold-start mode to meet dynamic requirements. Second,CloudPin-Fastimplements a pre-filter to reduce the transmission and processing of hyper-scale data. Finally,CloudPin-Fastuses an anomaly localization algorithm based on multi-dimensional statistics fusion in the second stage to cover complex scenarios. The evaluation results on four production datasets have shown superior efficiency and effectiveness. We also share lessons learned from deployingCloudPin-Fastfor over a year in a world-renowned public cloud vendor.
Shize Zhang, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Enhuan Dong, Jiahai Yang 0001
IEEE Trans. Serv. Comput.9
2023 Anomaly Detection in Heterogeneous Time Series Data for Server-Monitoring Tasks
abstract
When conducting anomaly detection on server monitoring data, it is important to consider the heterogeneity of the data, which is characterized by the diverse and irregular nature of events. The event values can vary widely, encompassing both continuous and discrete values, and there may be a multitude of randomly occurring events. However, many commonly used anomaly detection methods tend to overlook or discard this heterogeneous data, resulting in a significant loss of valuable information. As such, we propose a novel method, called Heterogeneous Time Series Anomaly Detection (HTSAD), to overcome this difficulty. The approach introduces event gates in the Long Short-Term Memory (LSTM) model while using unsupervised learning to overcome the challenges mentioned above. The results of our experiments on real-world datasets show that HTSAD could achieve an f-score of 0.958, which demonstrates the effectiveness of our approach in detecting anomalies in heterogeneous time series data.
Rui Yu 0003, Jiahai Yang 0001, Minghui Jin, Chenglong Li 0006, Enhuan Dong, Shutao Xia
ISCC4
2023 GraphIoT: Accurate IoT Identification based on Heterogeneous Graph
abstract
IoT devices deployed on campus and enterprise networks facilitate people's lives and work. However, these devices also bring serious network asset management and security management problems. IoT device identification is the premise to solve these problems. Although current IoT identification methods can identify devices with relatively high accuracy in ideal environments, it is difficult to accurately identify devices in real-world complex environments (e.g., campus networks, enterprise networks). Therefore, we propose to use exact features. To solve the problem of different dimensions of exact features, we creatively model the IoT identification problem as a heterogeneous graph representation learning problem and design a new representation learning algorithm. We are the first to propose an approach to accurately identify IoT devices in real-world complex environments and solve this problem through heterogeneous graphs. The evaluation shows that GraphIoT's macro F1 is on average 13.58% and 12.77% higher than the other methods on two public datasets.
Linna Fan, Lin He 0004, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001, Jinlei Lin, Guanglei Song
IWQoS5
2023 Which Doors Are Open: Reinforcement Learning-based Internet-wide Port Scanning
abstract
Internet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the majority of open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap achieves better performance on hit rate, coverage, and intrusiveness. Our experiments over real-world networks show that PMap can find 90% open ports by only scanning 125 ports (90% @125) to each active address with 136× less than the state-of-the-art port probing methods. PMap reduces the number of scanned ports to decrease the intrusive nature of port scanning. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research.
Guanglei Song, Lin He 0004, Tianyun Zhao, Yirui Luo, Yichao Wu, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001
IWQoS9
2023 Anomaly Detection in the Open World: Normality Shift Detection, Explanation, and Adaptation
Rui Yu 0003, Han Zhang 0009, Minghui Jin, Jiahai Yang 0001, Xingang Shi, Xia Yin 0001
NDSS10
2023 Your Router is My Prober: Measuring IPv6 Networks via ICMP Rate Limiting Side Channels
Long Pan, Jiahai Yang 0001, Lin He 0004, Leyao Nie, Guanglei Song, Yaozhong Liu
NDSS2
2023 BARS: Local Robustness Certification for Deep Learning based Traffic Analysis Systems
Jiahai Yang 0001, Xingang Shi, Xia Yin 0001
NDSS5
2023 Multi-stage Location for Root-Cause Metrics in Online Service Systems
abstract
The failure of the online service system will seriously affect the user experience and bring huge economic losses. Therefore, the operators usually monitor service-level metrics and machine-level metrics to help quickly find failures, locate root-cause metrics, and reduce MTTR(mean time to repair). Many methods have emerged in recent years to automatically locate root-cause metrics. However, the existing methods cannot meet the requirements of efficiency, accuracy, and ease of deployment at the same time, and are difficult to use in practice. To overcome their limitations, we propose MetricMiner- a multi-stage location method for root-cause metrics in online service systems. Our approach is based on a key observation from numerous real-world cases: root-cause metrics tend to be unique in both the time dimension and the machine dimension. Therefore, we divide the root-cause metrics localization into three stages: first, quickly filter out normal metrics with limited historical data; second, obtain sufficient historical data to eliminate abnormal metrics; finally, according to the clustering of abnormal metrics between machines to sort and locate root-cause metrics. Experimental results on two real-world datasets with 194 cases show that our method can significantly outperform the state-of-the-art methods. Moreover, MetricMiner has been deployed to multiple banking services for more than six months, and we also shared some lessons learned from real deployment.
Wenchi Zhang, Shize Zhang, Kaixin Sui, Enhuan Dong, Jiahai Yang 0001
NOMS7
2023 SmartSBD: Smart shared bottleneck detection for efficient multipath congestion control over heterogeneous networks
Enhuan Dong, Yuan Yang 0001, Mingwei Xu 0001, Xiaoming Fu 0001, Jiahai Yang 0001
Comput. Networks6
2023 AutoIoT: Automatically Updated IoT Device Identification With Semi-Supervised Learning
abstract
IoT devices bring great convenience to a person's life and industrial production. However, their rapid proliferation also troubles device management and network security. Network administrators usually need to know how many IoT devices are in the network and whether they behave normally. IoT device identification is the first step to achieving these goals. Previous IoT device identification methods reach high accuracy in a closed environment. But they are not applicable in the continuously changing environment. When new types of devices are plugged in, they cannot update themselves automatically. Besides, they usually rely on supervised learning and need lots of labeled data, which is costly. To solve these problems, we propose a novel IoT device identification model namedAutoIoT, updating itself automatically when new types of devices are plugged in. Besides, it only needs a few labeled data and identifies IoT devices with high accuracy. The evaluation on two public datasets shows thatAutoIoTcan identify new device types only using 1.5$\sim$2.5 hours’ traffic and still have high accuracy after updating. Moreover, it has a better performance than other works when there are only a few labeled data, especially in an environment with scanning traffic.
Linna Fan, Lin He 0004, Yichao Wu, Shize Zhang, Jia Li 0033, Jiahai Yang 0001, Chaocan Xiang, Xiaoqian Ma
IEEE Trans. Mob. Comput.7
2022 PerfTrace: A New Multi-metric Network Performance Monitoring Tool
abstract
We present PerfTrace, an end-to-end tool for efficient, real-time, and multi-metric network performance monitoring. PerfTrace provides a high integration of different existing measurement functions, supporting the measurement of essential metrics such as latency, jitter, packet loss, and available bandwidth. More importantly, innovative schemes and algorithms are proposed to address the weaknesses of existing tools.After conducting comprehensive evaluations, we find that (i) PerfTrace measures one-way and two-way latency, jitter, and packet loss ∼9.4× faster and ∼3.6× more data-efficiently; (ii) PerfTrace measures available bandwidth in our testbed with minimal mean relative error (5.22%), outperforming all the tools compared (ranging from 8.17% to 37.24%). Meanwhile, PerfTrace consumes a more constant percentage of bandwidth resources than other tools when monitoring available bandwidth. PerfTrace’s data overhead is always only about 1/600 of the total bandwidth for a measurement frequency once per minute.
Yaozhong Liu, Long Pan, Chenglong Li 0006, Lin He 0004, Yirui Luo, Guanglei Song, Jiahai Yang 0001
CNSM7
2022 Towards a Behavioral and Privacy Analysis of ECS for IPv6 DNS Resolvers
abstract
The Domain Name System (DNS) is critical to Internet communications. EDNS Client Subnet (ECS), a DNS extension, allows recursive resolvers to include client subnet information in DNS queries to improve CDN end-user mapping, extending the visibility of client information to a broader range. Major content delivery network (CDN) vendors, content providers (CP), and public DNS service providers (PDNS) are accelerating their IPv6 infrastructure development. With the increasing deployment of IPv6-enabled services and DNS being the most foundational system of the Internet, it becomes important to analyze the behavioral and privacy status of IPv6 resolvers. However, there is a lack of research on ECS for IPv6 DNS resolvers.In this paper, we study the ECS deployment and compliance status of IPv6 resolvers. Our measurement shows that 11.12% IPv6 open resolvers implement ECS. We discuss abnormal noncompliant scenarios that exist in both IPv6 and IPv4 that raise privacy and performance issues. Additionally, we measured if the sacrifice of clients’ privacy can enhance IPv6 CDN performance. We find that in some cases ECS helps end-user mapping but with an unnecessary privacy loss. And even worse, the exposure of client address information can sometimes backfire, which deserves attention from both Internet users and PDNSes.
Leyao Nie, Lin He 0004, Guanglei Song, Chenglong Li 0006, Jiahai Yang 0001
CNSM7
2022 Zoonet: a proactive telemetry system for large-scale cloud networks
abstract
We present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. The requirements are to (1) cover hyper-scale virtual networks with millions of tenants and millions of VMs for top tenants; (2) handle frequent virtual topology changes due to tenants' configuration through flexible APIs; (3) adapt to heterogeneous middleboxes along the probing paths; (4) achieve VM-to-VM telemetry without breaking tenant privacy; (5) differentiate virtual and physical network problems. We argue existing physical network telemetry solutions fail to satisfy our needs due to either incomplete telemetry coverage or outrageous telemetry overhead. Zoonet sets an ambitious goal to provide VM-to-VM hop-by-hop telemetry for each tenant, which is achieved based on self-developed, customizable middleboxes via hundreds of person-months under close team collaboration. At the data plane, Zoonet defines an elegant generalization of ping and traceroute, but made to work on multi-tenant clouds with heterogeneous middleboxes. At the control plane, Zoonet conducts substantial probing path pruning and update batch processing to lessen the overhead. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Jiahai Yang 0001
CoNEXT12
2022 Both Efficient and Accurate: A Large-scale One-way Delay Measurement Scheme
abstract
One-way delay (OWD) is one of the essential network performance metrics. In large-scale resilient overlay networks (RONs), OWD measurements can be used for shortest path selection and troubleshooting. However, OWD measurements remain difficult because of the need for precise time synchronization. Especially in large-scale networks, clock synchronization of all nodes has always been a considerable challenge. Therefore, in many cases, people use half of the round-trip time (RTT/2) as a rough substitute for the OWD. This paper presents an efficient and easy-to-deploy scheme for large-scale OWD measurements with the algorithm ClockConverger at its core. The scheme consists of three steps: Firstly, we perform low-precision time synchronization for all the measured nodes relying on network time protocol daemons (ntpd); Then, we use the open-source tool OWPing to perform OWD measurements; Finally, we correct the errors of the measured OWDs with our proposed ClockConverger. The theory and experiments show that our scheme's accuracy is significantly better than RTT/2. Meanwhile, the complexity of ClockConverger is$O(n^{2})$, which is much lower than the exponential complexity of the existing Maximum-Entropy algorithm.
Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006
GLOBECOM3
2022 What Causes Delay Asymmetry: A Large-scale One-way Delay Measurement and Empirical Study
abstract
In global communications, severe one-way delay (OWD) asymmetry often occurs. Due to the difficulties of OWD measurement (need to control both ends and synchronize their clocks), now RTT/2 is commonly used to estimate OWD. However, OWD asymmetry can lead to large errors in the halving RTT method, which in turn affects the end-to-end quality of service (QoS) guarantees. In this paper, we investigate OWD asymmetry through large-scale OWD measurements on a global scale. The measurements show that more than 11% of network paths have OWDs with a relative difference of more than 10% compared to RTT/2. By analyzing the measurement results in depth, we try to explain why the delay asymmetry occurs. We find that 67% is caused by hop inflation or a significant increase in propagation distance, and 33% is caused by variable queuing delays. We also find AS-level paths between node pairs with significant delay asymmetry are much more likely (~ 10 ×) to violate the well-known valley-free rule.
Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006
GLOBECOM2
2022 Centralized Network Utility Maximization with Accelerated Gradient Method
abstract
Network utility maximization (NUM) is a well-studied problem for network traffic management and resource allocation. Because of the inherent decentralization and complexity of networks, most researches develop decentralized NUM algorithms. In recent years, the Software Defined Networking (SDN) architecture has been widely used, especially in cloud networks and inter-datacenter networks managed by large enterprises, promoting the design of centralized NUM algorithms. To cope with the large and increasing number of flows in such SDN networks, existing researches about centralized NUM focus on the scalability of the algorithm with respect to the number of flows, however the efficiency is ignored. In this paper, we focus on the SDN scenario, and derive a centralized, efficient and scalable algorithm for the NUM problem. By the designing of a smooth utility function and a smooth penalty function, we formulate the NUM problem with a smooth objective function, which enables the use of Nesterov's accelerated gradient method. We prove that the proposed method has$O(d/t^{2})$convergence rate, which is the fastest with respect to the number of iterations$t$, and our method is scalable with respect to the number of flows$d$in the network. Experiments show that our method obtains accurate solutions with less iterations, and achieves close-to-optimal network utility.
Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009
ICNP5
2022 HetGLM: Lateral Movement Detection by Discovering Anomalous Links with Heterogeneous Graph Neural Network
abstract
As a critical stage in the Advanced Persistent Threat (APT) lifecycle, lateral movement (LM) has become a major concern in cybersecurity due to its stealthy nature. Recent authentication graph-based LM detection systems have achieved promising results. However, these methods have some unpractical requirements on data collection and model deployment, which severely affects their performance in real-world scenarios. In this paper, we propose HetGLM, a more accurate and practical LM detection system. Specifically, to fully explore the scenario, HetGLM constructs a heterogeneous graph with various network entities like users, devices, processes, etc. On this basis, we design MADR, a Graph neural network (GNN)-based anomaly link detection algorithm, to spot lateral movements. With the metapath-based sampling strategy, attention mechanism, the dual-decoder structure, and a mutual information regularization term, MADR can detect anomaly links on heterogeneous graphs, requiring neither labeled or purely benign training datasets nor manually preset thresholds. We implement a prototype of HetGLM and evaluate its performance via comprehensive experiments over public datasets. Comparison results show that HetGLM outperforms the state-of-the-art approaches in accuracy and practicality.
Xiaoqing Sun, Jiahai Yang 0001
IPCCC2
2022 WebIoT: Classifying Internet of Things Devices at Internet Scale through Web Characteristics
abstract
The number of Internet of Things (IoT) devices connected to the Internet has been growing rapidly. Such a large number of IoT devices bring significant challenges to device man-agement and cyberspace security. The discovery and classification of IoT devices are the prerequisites for monitoring and protecting them. However, existing Internet-scale IoT device classification methods mainly rely on textual analysis of the device response data, whose performance can be affected by the complexity or the multilingualism of the response texts. In this paper, we propose WebIoT, which mainly utilizes the image characteristics of the IoT devices' web interfaces to classify them for the first time. We leverage the observation that many IoT devices have web interfaces for device configuration and device status display, whose visual presentations contain abundant characteristics for device classification. Experiment results show that our method achieves 95.4% precision and 91.5% recall, which significantly outperforms other text analysis-based methods.
Yichao Wu, Chenglong Li 0006, Jiahai Yang 0001, Ang Xia, Yong Jiang 0001, Liuli Wu
ISCC3
2022 ROV-MI: Large-Scale, Accurate and Efficient Measurement of ROV Deployment
Chenxin Duan, Xia Yin 0001, Jiahai Yang 0001, Xingang Shi
NDSS6
2022 Monitoring Smart Home Traffic under Differential Privacy
abstract
Recent years have witnessed the proliferation of smart home ecosystems. Well-characterized traffic generated by smart home devices has promoted the development of security enhancing techniques for smart homes but exposes users to the privacy disclosure risk at the same time. Malicious eavesdroppers can infer working states of smart home devices and user activities based on spatial-temporal traffic characteristics. Existing countermeasures towards this kind of side channel attack ignore the utility of smart home traffic profiles and signatures for network management and attempt to completely eliminate them. In this paper, we give a comprehensive study on the trade offs between the usability of smart home traffic for security monitoring and its privacy threat. We propose to monitor the smart homes under differential privacy. Based on our solution, decoy traffic can be generated in a controlled manner so as to confound the attackers without disturbing the running monitoring systems. We prototyped our proposal and demonstrate its effectiveness empirically. An interview study is also conducted to learn the user acceptance of the proposed privacy preserving mechanism.
Chenxin Duan, Guanglei Song, Jiahai Yang 0001
NOMS5
2022 AddrMiner: A Comprehensive Global Active IPv6 Address Discovery System
Guanglei Song, Jiahai Yang 0001, Lin He 0004, Chenxin Duan, Yaozhong Liu, Zhongxiang Sun
USENIX ATC2
2022 EvoIoT: An evolutionary IoT and non-IoT classification model in open environments
Linna Fan, Lin He 0004, Enhuan Dong, Jiahai Yang 0001, Chenglong Li 0006, Jinlei Lin
Comput. Networks4
2022 Joint prediction on security event and time interval through deep learning
Songyun Wu, Bo Wang 0066, Shuhan Fan, Jiahai Yang 0001, Jia Li 0033
Comput. Secur.5
2022 THREATRACE: Detecting and Tracing Host-Based Threats in Node Level Through Provenance Graph Learning
abstract
Host-based threats such as Program Attack, Malware Implantation, and Advanced Persistent Threats (APT), are commonly adopted by modern attackers. Recent studies propose leveraging the rich contextual information in data provenance to detect threats in a host. Data provenance is a directed acyclic graph constructed from system audit data. Nodes in a provenance graph represent system entities (e.g.,processesandfiles) and edges represent system calls in the direction of information flow. However, previous studies, which extract features of the provenance graph, are not sensitive to the small quantity of threat-related entities and thus result in low performance when hunting stealthy threats. We present THREATRACE, an anomaly-based detector that detects host-based threats at system entity level without prior knowledge of attack patterns. We tailor GraphSAGE, an inductive graph neural network, to learn every benign entity’s role in a provenance graph. THREATRACE is a real-time system, which is scalable of monitoring a long-term running host and capable of detecting host-based intrusion in their early phase. We evaluate THREATRACE on five public datasets. The results show that THREATRACE outperforms seven state-of-the-art host intrusion detection systems.
Xia Yin 0001, Han Zhang 0009, Xingang Shi, Jiahai Yang 0001
IEEE Trans. Inf. Forensics Secur.9
2022 IntStream: Towards Flexible, Expressive, and Scalable Network Telemetry
abstract
Due to the complexity of the network structure and the high growth of the transmission speed, the measurement and management of the network are facing serious challenges. The traditional bottom-up network telemetry methods are no longer applicable to complex network scenarios. To bridge this gap, we propose IntStream, a flexible, expressive and scalable network telemetry framework to allow network operators to measure and analyze network through passive stream processing and active probing. However, there are three key challenges to building an intent-based telemetry system: (1) The diversity of network data sources. (2) The complexity of the measurement tasks. (3) The low overhead requirements of the telemetry system. IntStream introduces a lightweight component to extract and parse data from various data sources and divides the data stream processing into local and global stages. IntStream provides a set of rich expressive primitives to support operators to write telemetry tasks based on intent. By performing part of the telemetry task on the local stage, the transmission overhead of intermediate data can be effectively reduced. The evaluation results conducted on a large campus network show that IntStream can support a wide range of telemetry tasks while reducing the intermediate data transmission overhead by 99.31% on average.
Xin Cheng 0022, Shize Zhang, Jiahai Yang 0001
IEEE Trans. Netw. Serv. Manag.5
2022 ByteIoT: A Practical IoT Device Identification System Based on Packet Length Distribution
abstract
A tremendous amount of Internet-of-Things (IoT) devices have been deployed in recent years, bringing new challenges for network management and cyber security. It is important for network managers to know what types of IoT devices are connecting to the network. Despite much research efforts, previous works place more emphasis on accuracy but ignore some other performance indicators also in high demand, like efficiency, robustness, adaptability to special scenarios and extensibility for new devices. In this paper, we propose a practical IoT device identification system, namely ByteIoT, based on a simple but well-organized traffic feature, i.e., the frequency distribution of bidirectional packet lengths. ByteIoT applies k-nearest neighbors algorithm as the classifier to gain extensibility and adaptability. We evaluate ByteIoT on several datasets and the results show that ByteIoT can outperform other state-of-the-art methods in the aspects of accuracy, efficiency, extensibility and adaptability.
Chenxin Duan, Guanglei Song, Jiahai Yang 0001
IEEE Trans. Netw. Serv. Manag.4
2022 Weighted NSFIB Aggregation With Generalized Next Hop of Strict Partial Order
abstract
The size of the global routing table has been growing at an alarming rate. With the exhaustion of IPv4 addresses and the gradual deployment of IPv6 networks, the growth rate will continue to accelerate in the future. Although modern high performance routers provide enough line-card memory, Internet Service Providers (ISPs) cannot afford to upgrade their routers as fast as the growth of global routing tables. In this paper, we propose an algorithm to calculate the generalized next hops with strict partial order (GSPO next hops) of a network prefix and use them for the aggregation of the Nexthop-Selectable Forwarding Information Base (NSFIB). Since the existing NSFIB aggregation algorithm may introduce path stretch, we also propose a weighted NSFIB aggregation algorithm to effectively control path stretch under a given upper limit of the FIB size. Experiment results show that our algorithm can shrink the FIB size by at most 97% under IPv4 networks, and at most 95% under IPv6 networks. Under a given upper limit of the FIB size, our algorithm can reduce the path stretch by at least 22%.
Qing Li 0006, Yichao Wu, Jingpu Duan, Jiahai Yang 0001, Yong Jiang 0001
IEEE Trans. Netw. Serv. Manag.4
2022 DET: Enabling Efficient Probing of IPv6 Active Addresses
abstract
Fast IPv4 scanning significantly improves network measurement and security research. Nevertheless, it is infeasible to perform brute-force scanning of the IPv6 address space. Alternatively, one can find active IPv6 addresses through scanning the candidate addresses generated by state-of-the-art algorithms. However, the probing efficiency of such algorithms is often very low. In this paper, our objective is to improve the probing efficiency of IPv6 addresses. We first perform a longitudinal active measurement study and build a high-quality dataset, hitlist, including more than 1.95B IPv6 addresses distributed in 58.2K BGP prefixes and collected over 17 months period. Different from the previous works, we probe the announced BGP prefixes using a pattern-based algorithm. This results in a dataset without uneven address distribution and low active rates. Further, we propose an efficient address generation algorithm, DET, which builds a density space tree to learn high-density address regions of the seed addresses with linear time complexity and improves the active addresses’ probing efficiency. We then compare our algorithm DET against state-of-the-art algorithms on the public hitlist and our hitlist by scanning 50M addresses. Our analysis shows that DET increases the de-aliased active address ratio and active address (including aliased addresses) ratio by 10%, and 14%, respectively. Furthermore, we develop a fingerprint-based method to detect aliased prefixes. The proposed method for the first time directly verifies whether the prefix is aliased or not. Our method finds that 10.64% of the public aliased prefixes are false positive.
Guanglei Song, Jiahai Yang 0001, Lin He 0004, Jinlei Lin, Long Pan, Chenxin Duan, Xiaowen Quan
IEEE/ACM Trans. Netw.2
2021 MineHunter: A Practical Cryptomining Traffic Detection Algorithm Based on Time Series Tracking
abstract
With the development of cryptocurrencies’ market, the problem of cryptojacking, which is an unauthorized control of someone else’s computer to mine cryptocurrency, has been more and more serious. Existing cryptojacking detection methods require to install anti-virus software on the host or load plug-in in the browser, which are difficult to deploy on enterprise or campus networks with a large number of hosts and servers. To bridge the gap, we propose MineHunter, a practical cryptomining traffic detection algorithm based on time series tracking. Instead of being deployed at the hosts, MineHunter detects the cryptomining traffic at the entrance of enterprise or campus networks. Minehunter has taken into account the challenges faced by the actual deployment environment, including extremely unbalanced datasets, controllable alarms, traffic confusion, and efficiency. The accurate network-level detection is achieved by analyzing the network traffic characteristics of cryptomining and investigating the association between the network flow sequence of cryptomining and the block creation sequence of cryptocurrency. We evaluate our algorithm at the entrance of a large office building in a campus network for a month. The total volumes exceed 28 TeraBytes. Our experimental results show that MineHunter can achieve precision of 97.0% and recall of 99.7%.
Shize Zhang, Jiahai Yang 0001, Xin Cheng 0022, Xiaoqian Ma, Hui Zhang 0052, Bo Wang 0066, Zimu Li
ACSAC3
2021 DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications
abstract
Unsupervised Deep Learning (DL) techniques have been widely used in various security-related anomaly detection applications, owing to the great promise of being able to detect unforeseen threats and superior performance provided by Deep Neural Networks (DNN). However, the lack of interpretability creates key barriers to the adoption of DL models in practice. Unfortunately, existing interpretation approaches are proposed for supervised learning models and/or non-security domains, which are unadaptable for unsupervised DL models and fail to satisfy special requirements in security domains.
Ying Zhong 0008, Han Zhang 0009, Jiahai Yang 0001, Xingang Shi, Xia Yin 0001
CCS7
2021 IntStream: An Intent-driven Streaming Network Telemetry Framework
abstract
Due to the complexity of the network structure and the high growth of the transmission speed, the measurement and management of the network are facing serious challenges. The traditional bottom-up network telemetry methods are no longer applicable to complex network scenarios. To bridge this gap, we propose IntStream, an intent-driven streaming network telemetry framework to allow network operators to measure and analyze network traffic. However, there are three key challenges to building an intent-based telemetry system: (1) The diversity of network data sources. (2) The complexity of the measurement tasks. (3) The low overhead requirements of the telemetry system. IntStream introduces a lightweight component to extract and parse data from various types of data sources to form a data stream and divides the data stream conversation process into local and global stages. IntStream provides a set of rich expressive primitives to support users to write telemetry tasks based on intent. By performing part of the telemetry task on the local stage, the transmission overhead of intermediate data can be effectively reduced. The evaluation results conducted on a large campus network show that IntStream can support a wide range of telemetry tasks while reducing the intermediate data transmission overhead by 99.64% on average.
Xin Cheng 0022, Shize Zhang, Jiahai Yang 0001
CNSM5
2021 Traffic Engineering with Segment Routing Considering Probabilistic Failures
abstract
Segment Routing (SR) is a source routing paradigm that routes a packet through an ordered list of instructions called segments. It is widely used in Traffic Engineering (TE) because of its simplicity and scalability. Although there are lots of research about TE with SR (SR-TE), fewer consider network failures. The reactive approaches may suffer from latency and update issues, and the proactive approaches don't perform very well because the objectives aren't carefully designed. Besides, although different types of failures are considered, the failure probabilities are ignored. In this paper, we take failure probabilities in to consideration, and propose a proactive 2-SR model 2SRPF to handle SR-TE problem with network failures, aiming at minimizing maximum link utilization (MLU). Considering that severe failures are more noteworthy, we use probability as a severity threshold, and minimize the expectation of the larger MLUs whose corresponding failure states have probabilities sum to a specific threshold value. We solve it with probabilistic risk management. Experiments show that 2SRPF performs well with one threshold setting for different topologies consistently, and gets close to optimal results when network fails.
Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Han Zhang 0009, Yingya Guo, Haijun Geng
CNSM5
2021 Slider: Towards Precise, Robust and Updatable Sketch-based DDoS Flooding Attack Detection
abstract
Distributed Denial of Service (DDoS) flooding attacks have been a severe threat to the Internet for decades. These attacks usually are launched by exhausting bandwidth, network resources or server resources. Since most of these attacks are launched abruptly and severely, it is crucial to develop an efficient DDoS flooding attack detection system. In this paper, we present Slider, an online sketch-based DDoS flooding attack detection system. Slider utilizes a new type of sketch structure, namely Rotation Sketch, to effectively detect DDoS flooding attacks and efficiently identify the malicious hosts. Meanwhile, Slider also learns the characteristics of the current network during the time specified by the network operator to periodically update the parameters of its detection model. We have developed a prototype of Slider and the evaluation results on real-world traffic and public DDoS/DoS attack datasets demonstrate that Slider can effectively detect various DDoS flooding attacks with high precision and robustness.
Xin Cheng 0022, Shize Zhang, Jia Li 0033, Jiahai Yang 0001
GLOBECOM5
2021 Deception Maze: A Stackelberg Game-Theoretic Defense Mechanism for Intranet Threats
abstract
The intranets in modern organizations are facing severe data breaches and critical resource misuses. By reusing user credentials from compromised systems, Advanced Persistent Threat (APT) attackers can move laterally within the internal network. A promising new approach called deception technology makes the network administrator (i.e., defender) able to deploy decoys to deceive the attacker in the intranet and trap him into a honeypot. Then the defender ought to reasonably allocate decoys to potentially insecure hosts. Unfortunately, existing APT-related defense resource allocation models are infeasible because of the neglect of many realistic factors.In this paper, we make the decoy deployment strategy feasible by proposing a game-theoretic model called the APT Deception Game to describe interactions between the defender and the attacker. More specifically, we decompose the decoy deployment problem into two subproblems and make the problem solvable. Considering the best response of the attacker who is aware of the defender’s deployment strategy, we provide an elitist reservation genetic algorithm to solve this game. Simulation results demonstrate the effectiveness of our deployment strategy compared with other heuristic strategies.
Jieling Liu, Jiahai Yang 0001, Bo Wang 0066, Lin He 0004, Guanglei Song
ICC3
2021 Unsupervised IoT Fingerprinting Method via Variational Auto-encoder and K-means
abstract
With the rapid growth of the number of IoT devices on the Internet, security problems of IoT devices are becoming more and more serious, which bring more challenges to network administrators. The first task to solve these problems for network administrators is being aware of IoT devices in the network. Previous IoT device identification methods typically use supervised machine learning methods, which require a large amount of labeled sample data. However, it is difficult to obtain a large number of labeled samples effectively. In order to address this problem, we propose an unsupervised IoT device fingerprinting method at the network level, which can effectively cluster IoT devices without labeled samples. We deeply analyze the temporal and spatial dimension characteristics of network traffic, which can adequately reflect the differences between different IoT devices. By using these features, we develop a clustering framework based on variational autoencoder and K-means algorithms. We conduct evaluation experiments on a public dataset including 24 different IoT devices. The experimental results show that our clustering algorithm can achieve accuracy of 86.7% outperforming a k-NN based state-of-art supervised approach.
Shize Zhang, Jiahai Yang 0001, Dongbin Bai, Fuliang Li, Zimu Li
ICC3
2021 ADSIM: Network Anomaly Detection via Similarity-aware Heterogeneous Ensemble Learning
Ying Zhong 0008, Chenxin Duan, Xia Yin 0001, Jiahai Yang 0001, Xingang Shi
IM7
2021 PINBALL: Universal and Robust Signature Extraction for Smart Home Devices
Chenxin Duan, Shize Zhang, Jiahai Yang 0001, Yang Yang 0004, Jia Li 0033
IM3
2021 STRAD: Network Intrusion Detection Algorithm Based on Zero-Positive Learning in Real Complex Network Environment
abstract
With the increasing network security risks, network intrusion detection technology has become more important. At present, machine learning is applied in most advanced traffic anomaly detection algorithms, but these algorithms have three main shortcomings. First, algorithms using deep neural network are highly complex and not suitable for real-time online processing. Second, algorithms based on supervised learning require training on huge labeled data sets, which are limited and insufficient. Third, most algorithms have such poor generalization ability and portability that they are less suitable for real-world environments. Therefore, we propose a novel network anomaly detection model, STRAD. We use Word2vec and Damped Incremental Statistics algorithm for spatiotemporal features extraction, latent space compression (LSC) for feature vectors compression and an unsupervised one-class classifier for anomaly detection. Our evaluations show that STRAD has a better performance than other state of the art algorithms.
Ying Zhong 0008, Rui Li 0019, Citong Que, Jiahai Yang 0001, Xia Yin 0001, Xingang Shi, Keqin Li 0001
ISCC7
2021 CloudPin: A Root Cause Localization Framework of Shared Bandwidth Package Traffic Anomalies in Public Cloud Networks
abstract
Due to the sharing nature of public cloud, most of the cloud services use a sharing bandwidth package (sBwp) model to conduct inbound/outbound communication. The sBwp model allows users to purchase a sharing bandwidth for plenty of virtual machines instead of purchasing bandwidth for each virtual machine separately. The advantage of sBwp is that it can provide users with convenient configuration and lower economic cost. However, the sBwp model brings new challenges for operators to localize the root cause of traffic anomalies of a sharing bandwidth, especially for a globally distributed large-scale public cloud with millions of users. In this paper, we first formalize the sBwp problem on the cloud and propose CloudPin, a root cause localization framework for this problem. Our framework solves all the challenges by employing a multi-dimensional algorithm with three sub-models of prediction deviation, anomaly ampli-tude, and shape similarity, and an overall ranking algorithm. Evaluations on real-world data, from one of the world-renowned public cloud vendors, show that our algorithm precision reaches 97.8% for the top 1 of the ranking list, outperforming multiple baseline algorithms.
Shize Zhang, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Jiahai Yang 0001, Lin He 0004
ISSRE7
2021 Distributed and Adaptive Traffic Engineering with Deep Reinforcement Learning
abstract
Lots of studies focus on distributed traffic engineering (TE) where routers make routing decisions independently. Existing approaches usually tackle distributed TE problems through traditional optimization methods. However, due to the intrinsic complexity of the distributed TE problems, routing decisions cannot be obtained efficiently, which leads to significant performance degradation, especially for highly dynamic traffic. Emerging machine learning technologies like deep reinforcement learning (DRL) provide a new choice to address TE problems in an experience-driven method. In this paper, we propose DATE, a distributed and adaptive TE framework with DRL. DATE distributes well-trained agents to the routers in the located network. Each agent makes local routing decisions independently based on link utilization ratios flooded by each router periodically. To coordinate the distributed agents to achieve the global optimization in different traffic conditions, we construct candidate paths, develop the agents carefully, and realize a virtual environment to train the agents with a DRL algorithm. We do extensive simulations and experiments using real-world network topologies with both real and synthetic traffic traces. The results show that DATE outperforms some existing approaches and yields near-optimal performance with superior robustness.
Nan Geng, Mingwei Xu 0001, Yuan Yang 0001, Chenyi Liu, Jiahai Yang 0001, Qi Li 0002, Shize Zhang
IWQoS5
2021 Towards securing Duplicate Address Detection using P4
Lin He 0004, Peng Kuang, Ying Liu 0024, Gang Ren 0003, Jiahai Yang 0001
Comput. Networks5
2021 A novel workload scheduling framework for intrusion detection system in NFV scenario
Jia Li 0033, Jiahai Yang 0001, Jinlei Lin
Comput. Secur.3
2021 Evaluating and Improving Adversarial Robustness of Machine Learning-Based Network Intrusion Detectors
abstract
Machine learning (ML), especially deep learning (DL) techniques have been increasingly used in anomaly-based network intrusion detection systems (NIDS). However, ML/DL has shown to be extremely vulnerable to adversarial attacks, especially in such security-sensitive systems. Many adversarial attacks have been proposed to evaluate the robustness of ML-based NIDSs. Unfortunately, existing attacks mostly focused on feature-space and/or white-box attacks, which make impractical assumptions in real-world scenarios, leaving the study on practical gray/black-box attacks largely unexplored. To bridge this gap, we conduct the first systematic study of the gray/black-box traffic-space adversarial attacks to evaluate the robustness of ML-based NIDSs. Our work outperforms previous ones in the following aspects: (i) practical -the proposed attack can automatically mutate original traffic with extremely limited knowledge and affordable overhead while preserving its functionality; (ii) generic -the proposed attack is effective for evaluating the robustness of various NIDSs using diverse ML/DL models and non-payload-based features; (iii) explainable -we propose an explanation method for the fragile robustness of ML-based NIDSs. Based on this, we also propose a defense scheme against adversarial attacks to improve system robustness. We extensively evaluate the robustness of various NIDSs using diverse feature sets and ML/DL models. Experimental results show our attack is effective (e.g., >97% evasion rate in half cases for Kitsune, a state-of-the-art NIDS) with affordable execution cost and the proposed defense method can effectively mitigate such attacks (evasion rate is reduced by >50% in most cases).
Ying Zhong 0008, Jiahai Yang 0001, Shuqiang Lu, Xingang Shi, Xia Yin 0001
IEEE J. Sel. Areas Commun.5
2021 Improving the Performance of Online Bitrate Adaptation with Multi-Step Prediction Over Cellular Networks
abstract
Video streaming over mobile is flourishing, and most commercial players use adaptive bitrate (ABR) streaming to deliver video in varying network conditions. Using network capacity and buffer occupancy as system states, ABR algorithms adjust bitrate based on the instantaneous system states, which is able to adapt to network changes in real-time and ensure high quality of experience (QoE). However, they are incapable of providing good QoE over mobile. Due to the high dynamic characteristics of cellular network, the system states change rapidly over time. The instantaneous state-based adaptation can induce significant video quality fluctuation which greatly degrades QoE. In this paper, we propose an online ABR algorithm called MSPC to provide good QoE in cellular network. To balance the conflict between rapid adaptation and smooth bitrate, MSPC utilizes the multi-step prediction of future system states to select bitrates instead of the instantaneous current states. At the same time, it controls the buffer occupancy to eliminate the impact of prediction error on performance. We implement MSPC on a reference video player with performance evaluated based on realistic cellular traces. Experimental results show that MSPC reduces the bitrate change of existing online algorithms by 62.4 percent on average while maintaining high bitrates and achieving zero rebuffering over 97.83 percent of all tested sessions.
Bo Wang 0066, Fengyuan Ren, Jiahai Yang 0001, Chao Zhou 0003
IEEE Trans. Mob. Comput.3
2021 PAVI: Bootstrapping Accountability and Privacy to IPv6 Internet
abstract
Accountability and privacy are considered valuable but conflicting properties in the Internet, which at present does not provide native support for either. Past efforts to balance accountability and privacy in the Internet have unsatisfactory deployability due to the introduction of new communication identifiers, and because of large-scale modifications to fully deployed infrastructures and protocols. The IPv6 is being deployed around the world and this trend will accelerate. In this paper, we propose a private and accountable proposal based on IPv6 called PAVI that seeks to bootstrap accountability and privacy to the IPv6 Internet without introducing new communication identifiers and large-scale modifications to the deployed base. A dedicated quantitative analysis shows that the proposed PAVI achieves satisfactory levels of accountability and privacy. The results of the evaluation of a PAVI prototype show that it incurs little performance overhead, and is widely deployable.
Lin He 0004, Gang Ren 0003, Ying Liu 0024, Jiahai Yang 0001
IEEE/ACM Trans. Netw.4
2021 vSFC: Generic and Agile Verification of Service Function Chains in the Cloud
abstract
With the advent of network function virtualization (NFV), outsourcing network functions (NFs) to the cloud is becoming increasingly popular for enterprises since it brings significant benefits for NF deployment and maintenance, such as improved scalability and reduced overhead. However, NF outsourcing limits the control of customer enterprises over NF deployment and management, consequently raising serious security concerns. Enterprises cannot ensure whether their outsourced NFs and associated service function chains (SFCs) are correctly enforced according to their specifications. In this paper, we propose vSFC, an SFC verification scheme that allows an enterprise to accurately verify the correctness of SFC enforcement in real time. Specifically, it can detect a wide range of SFC violations including forwarding path incompliance, packet dropping, and flow dropping attacks. Meanwhile, it is generic and agile, which can be applied to arbitrary cloud architectures without requiring any modification to NFs. To demonstrate the feasibility and performance of vSFC, we implement a vSFC prototype on top of Linux kernel-based virtual machines (KVM) and conduct extensive experiments with real traffic. The experimental results show that vSFC can accurately detect SFC violations with negligible overhead.
Xiaoli Zhang 0003, Qi Li 0002, Jiahai Yang 0001
IEEE/ACM Trans. Netw.5
2020 An IoT Device Identification Method based on Semi-supervised Learning
abstract
With the rapid proliferation of IoT devices, device management and network security are becoming significant challenges. Knowing how many IoT devices are in the network and whether they are behaving normally is significant. IoT device identification is the first step to achieve these goals. Previous IoT identification works mainly use supervised learning and need lots of labeled data. Considering collecting labeled data is time-consuming and cannot be scaled, in this paper, we propose an IoT identification model based on semi-supervised learning. The model can differentiate IoT and non-IoT and classify specific IoT devices based on time interval features, traffic volume features, protocol features and TLS related features. The evaluation in a public dataset shows that our model only needs 5% labeled data and gets accuracy over 99%.
Linna Fan, Shize Zhang, Yichao Wu, Chenxin Duan, Jia Li 0033, Jiahai Yang 0001
CNSM7
2020 Domain-Embeddings Based DGA Detection with Incremental Training Method
abstract
DGA-based botnet, which uses Domain Generation Algorithms (DGAs) to evade supervision, has become a part of the most destructive threats to network security. Over the past decades, a wealth of defense mechanisms focusing on domain features have emerged to address the problem. Nonetheless, DGA detection remains a daunting and challenging task due to the big data nature of Internet traffic and the potential fact that the linguistic features extracted only from the domain names are insufficient and the enemies could easily forge them to disturb detection. In this paper, we propose a novel DGA detection system which employs an incremental word-embeddings method to capture the interactions between end hosts and domains, characterize time-series patterns of DNS queries for each IP address and therefore explore temporal similarities between domains. We carefully modify the Word2Vec algorithm and leverage it to automatically learn dynamic and discriminative feature representations for over 1.9 million domains, and develop an simple classifier for distinguishing malicious domains from the benign. Given the ability to identify temporal patterns of domains and update models incrementally, the proposed scheme makes the progress towards adapting to the changing and evolving strategies of DGA domains. Our system is evaluated and compared with the state-of-art system FANCI and two deep-learning methods CNN and LSTM, with data from a large university’s network named TUNET. The results suggest that our system outperforms the strong competitors by a large margin on multiple metrics and meanwhile achieves a remarkable speed-up on model updating.
Xiaoqing Sun, Jiahai Yang 0001
ISCC3
2020 Towards the Construction of Global IPv6 Hitlist and Efficient Probing of IPv6 Address Space
abstract
Fast IPv4 scanning has made sufficient progress in network measurement and security research. However, it is infeasible to perform brute-force scanning of the IPv6 address space. We can find active IPv6 addresses through scanning candidate addresses generated by the state-of-the-art algorithms, whose probing efficiency of active IPv6 addresses, however, is still very low. In this paper, we aim to improve the probing efficiency of IPv6 addresses in two ways. Firstly, we perform a longitudinal active measurement study over four months, building a high-quality dataset called hitlist with more than 1.3 billion IPv6 addresses distributed in 45.2k BGP prefixes. Different from previous work, we probe the announced BGP prefixes using a pattern-based algorithm, which makes our dataset overcome the problems of uneven address distribution and low active rate. Secondly, we propose an efficient address generation algorithm DET, which builds a density space tree to learn high-density address regions of the seed addresses in linear time and improves the probing efficiency of active addresses. On the public hitlist and our hitlist, we compare our algorithm DET against state-of-the-art algorithms and find that DET increases the de-aliased active address ratio by 10%, and active address (including aliased addresses) ratio by 14%, by scanning 50 million addresses.
Guanglei Song, Lin He 0004, Jiahai Yang 0001, Jieling Liu
IWQoS4
2020 HGDom: Heterogeneous Graph Convolutional Networks for Malicious Domain Detection
abstract
As a fundamental component of the Internet, Domain Name System (DNS) is widely abused by attackers in various cybercrimes, making malicious domain detection an essential task in network defenses. However, some well-crafted attacks with tricky techniques can not only bypass blacklists but also make some machine learning-based detection systems infeasible. In this paper, we design HGDom, an accurate and robust malicious domain detection system based on a heterogeneous graph convolutional network method. First, we jointly analyze domain features as well as the complex relations among domains, clients, and IP addresses. To capture richer information, we introduce a Heterogeneous Information Network (HIN) to model the DNS scene. Then, we propose a novel representation method named MAGCN. With a meta-path-based attention mechanism, it can handle node features and the graph structure in HIN at the same time. To our best knowledge, this is the first work to apply GCN in cyber security analysis. Comprehensive experiments over DNS data from TUNET and CERNET2 are conducted to validate the effectiveness and superiority of our proposed methods. The comparison results show that HGDom outperforms state-of-the-art approaches with promising performance. Besides, the system is decided to be deployed in production to assist with network security management for CERNET2.
Xiaoqing Sun, Jiahai Yang 0001
NOMS2
2020 Far from classification algorithm: dive into the preprocessing stage in DGA detection
abstract
Domain-Flux technique has been widely used by attackers to maintain a botnet for many years and the core of it is the adoption of domain generation algorithm (DGA). To combat attackers, there are lots of works in DGA domain detection area recently. But they usually collect quite limited data and conduct experiments in a closed dataset, meaning that the DGA data and the benign data they collected can not well represent the real distribution between them. Moreover, they handle the domains roughly and use the origin data to train the classifier directly, which is also not adequate to classify these two types of domains with lots of false positives and false negatives happening during the real-world deployment. In this paper, we conduct the first large-scale DGA domain analysis in traffic level and argue that the preprocessing stage is also vital for the final classifier, which is usually ignored by the existing works. We collect the largest amount of DGA domain data than prior works and collect DNS log offered by a big company, whose DNS data covers most important industries in China. Based on this data, we analyze the distribution of DGA domains in traffic and give quantifiable results showing that NXDomain (domain not exist) is more suitable for DGA detection. Moreover, we give detailed preprocessing steps to handle the original domains. Our experiment shows that with the preprocessing stage mentioned above, classifier performs better in DGA detection task. Our research indicates that improving the classification algorithm is far from enough in DGA detection and the preprocessing stage is also the key component in bringing the DGA detection methods from lab to product.
Mingkai Tong, Runzi Zhang, Jianxin Xue, Wenmao Liu, Jiahai Yang 0001
TrustCom6
2020 HELAD: A novel network anomaly detection model based on heterogeneous ensemble learning
Ying Zhong 0008, Xia Yin 0001, Xingang Shi, Jiahai Yang 0001, Keqin Li 0001
Comput. Networks9
2020 Deepdom: Malicious domain detection with scalable and heterogeneous graph convolutional networks
Xiaoqing Sun, Jiahai Yang 0001
Comput. Secur.3
2020 Fine-Grained Cloud Resource Provisioning for Virtual Network Function
abstract
The deployment of Virtualized Network Functions is expected to be dynamic and swift when using Network Function Virtualization technology. The dynamic nature of workload from users requires the resource allocation of underlying infrastructure to be flexible to cope with the changes. Existing works investigated elastic NFV solutions by dynamically creating and dismantling Virtual Machine (VM) replicas, while maintaining balanced workload among VMs. However, those solutions are coarse-grained which may cause unnecessary resource over-provisioning as different network functions consume different amount of resources. In this paper, we present ElasticNFV, a dynamic and fine-grained cloud resource provisioning solution for VNF. ElasticNFV takes real-time resource demand of multiple service chains and allocates resources through an elastic provision mechanism. When a scaling conflict occurs, ElasticNFV provides a two-phase minimal migration algorithm to optimize the migration time and embedding cost of VNF instances. We implement ElasticNFV on top of the KVM platform to provide elastic VM for each VNF instance and Open vSwitch to form elastic intra-cloud network with virtual links between VNF instances. Our evaluation results show that ElasticNFV can improve VNF performance significantly, and achieve high resource utilization and fast migration time with low cost.
Hui Yu 0004, Jiahai Yang 0001, Carol J. Fung
IEEE Trans. Netw. Serv. Manag.2
2020 Traffic Engineering in Partially Deployed Segment Routing Over IPv6 Network With Deep Reinforcement Learning
abstract
Segment Routing (SR) is a source routing paradigm which is widely used in Traffic Engineering (TE). By using SR, a node steers a packet through an ordered list of instructions called segments. By some extensions of interior gateway protocol, SR can be applied to IP/MPLS or IPv6 network without signal protocol. SR over IPv6 (SRv6) is attracting wide attention because of its interoperation ability with IPv6. However, upgrading the existing IPv6 network directly to a full SRv6 one can be difficult, because large-scale equipment replacement or software upgrade may cause economic and technical problems. TE in partially deployed SR network is becoming a hot research topic. In this paper, we propose the TE algorithm Weight Adjustment-SRTE (WA-SRTE) in partially deployed SRv6 network, in which SRv6 capable nodes are dispersedly deployed. Our objective is to minimize the network's maximum link utilization. WA-SRTE converts the TE problem into a Deep Reinforcement Learning problem and optimizes the OSPF weight, SRv6 node deployment and traffic paths simultaneously. Besides, traffic variation is also considered and we use a representative Traffic Matrix (TM) to epitomize the traffic characteristics over a period of time. Experiments demonstrate that with 20% to 40% of the SRv6 nodes deployed, we can achieve TE performance as good as in a full SR network for the experiment topologies. The results with WA remarkably outperform the results without it. Our algorithm also gets near-optimal results with changing traffic.
Xia Yin 0001, Xingang Shi, Yingya Guo, Haijun Geng, Jiahai Yang 0001
IEEE/ACM Trans. Netw.7
2019 Measurement and Analysis of Adult Websites in IPv6 Networks
abstract
The Internet is in the transition from IPv4 to IPv6. At present, researches on IPv6 networks mainly focus on architectural issues, such as routing, addressing, and security; there are few studies on the operational issues of IPv6 networks. Our preliminary observation shows that there are a large amount adult websites and traffic in IPv6 networks. Adult websites can damage health of teenagers and bring operational issues in IPv6 networks. This paper conducts a comprehensive measurement and analysis of the adult websites and traffic in IPv6 networks to help solve these operational issues. The data used in this paper is the raw packet traffic from CNGI-CERNET2 which is a pure IPv6 academic network in China. The duration of the data is from July 2017 to January 2018 and the total amount is 40+ terabytes. We detected about 3000 adult websites in the global IPv6 network. This paper analyzes these adult websites and traffic from the perspectives of websites, users and ISPs respectively. We find that adult websites are still in the developing stage in IPv6 networks and only 30% adult websites with full resources can be accessed in IPv6-only networks. But due to the IPv6-first policy in RFC 4038, adult traffic will continue to migrate to IPv6 networks from IPv4 networks. On the other hand, we find that CDN vendors promote the development of adult websites in IPv6 networks and many adult website owners use muti-domain policies to escape ISPs restricting. Our findings may help ISPs effectively understand adult websites and enhance the restriction of adult content in IPv6 networks.
Shize Zhang, Hui Zhang 0052, Jiahai Yang 0001, Guanglei Song
APNOMS3
2019 ALEAP: Attention-based LSTM with Event Embedding for Attack Projection
abstract
Cyberattacks have developed rapidly in diversity and complexity in recent years. Despite the existence of various defense systems, it cannot provide early warnings and prevent catastrophic consequences in advance. Therefore, the need for prediction becomes more and more urgent, especially for those multiple step attacks in which several steps are required for achieving the attack successfully. In this paper, we focus on attack projection that is aimed to predict the next step of the attack based on historical information and gained knowledge of similar events happened in the past. Previous models on attack projection based on probability graph model or simple RNN models, which may limit their capability of noise tolerance and sequence association analysis. To remedy this, we propose a method called ALEAP which incorporates event embedding and attention mechanism into LSTM models to better predict the future events. We test ALEAP on a dataset of millions of security events collected from the multi-source security devices, and show that our approach is effective in event prediction. ALEAP also provides a useful method for security specialists and all computer environment-related parties to better predict attack projection and defend known attacks.
Shuhan Fan, Songyun Wu, Zimu Li, Jiahai Yang 0001
IPCCC5
2019 D3N: DGA Detection with Deep-Learning Through NXDomain
Mingkai Tong, Xiaoqing Sun, Jiahai Yang 0001, Hui Zhang 0052, Shuang Zhu
KSEM (1)3
2019 HinDom: A Robust Malicious Domain Detection System based on Heterogeneous Information Network with Transductive Classification
Xiaoqing Sun, Mingkai Tong, Jiahai Yang 0001
RAID3
2019 Towards predictable performance via two-layer bandwidth allocation in cloud datacenter
Hui Yu 0004, Jiahai Yang 0001, Hui Wang 0011, Hui Zhang 0052
J. Parallel Distributed Comput.2
2018 Elastic Network Service Chain with Fine-Grained Vertical Scaling
abstract
By moving network functions from dedicated hardware to software, Network Function Virtualization (NFV) is expected to bring the advantages of cloud computing to network management. Frequent workload changes require the underlying infrastructure to be dynamic and agile to cope with the changes. Some existing studies have investigated elastic virtual machine (VM) positioning solutions by dynamically creating and destroying VM replicas, while maintaining balanced workload among VMs. However, those solutions are coarse-grained which may cause unnecessary resource over-provisioning and low resource utilization. In this paper, we propose ElasticNFV, a dynamic solution that achieves fine-grained cloud resource provisioning for Virtual Network Functions (VNFs). ElasticNFV analyzes realtime resource demand of multiple service chains and allocates resource through an elastic provision mechanism. When a scaling conflict occurs, ElasticNFV provides a Two-Phase Minimal Migration (TPMM) algorithm to optimize migration time and embedding cost of VNFs based on prediction. We implemented ElasticNFV on top of a KVM virtualization platform and Open vSwitch. Through simulation and testbed evaluation, we show that ElasticNFV can achieve high resource utilization and short migration time with low cost.
Hui Yu 0004, Jiahai Yang 0001, Carol J. Fung
GLOBECOM2
2018 ENSC: Multi-Resource Hybrid Scaling for Elastic Network Service Chain in Clouds
abstract
Software-based network service chains in Network Function Virtualization (NFV) need to be dynamically allocated and scaled on hardware resources. This is because the resource demand of virtual network functions (VNFs) typically varies as a results of network flow volume. NFV elastic solutions by coarse-grained horizontal scaling or fine-grained vertical scaling have been investigated in recent years. However, none of the existing solutions can achieve both efficiency and scalability. To address this challenge, we propose elastic network service chain (ENSC), which utilizes a fine-grained hybrid scaling method to achieve both NFV efficiency and scalability. We systematically compare horizontal scaling with vertical scaling from six aspects and determine the priority within hybrid scaling. We formulate the resource allocation problem in the cloud datacenter as an integer linear programming (ILP) model and develop a heuristic algorithm called Rubik. Our evaluation results show that ENSC achieves higher acceptance ratios and resource utilization than horizontal scaling and vertical scaling methods.
Hui Yu 0004, Jiahai Yang 0001, Carol J. Fung, Raouf Boutaba
ICPADS2
2018 Root Cause Analysis of Anomalies of Multitier Services in Public Clouds
Jianping Weng, Hui Wang 0011, Jiahai Yang 0001, Yang Yang 0004
IEEE/ACM Trans. Netw.3
2017 Robust regression for anomaly detection
abstract
In our previous work, we have applied ordinary linear regression equation to network anomaly detection. However, the performance of ordinary linear regression equation is susceptible to outliers. Unfortunately, it is almost impossible to obtain a “clean” traffic data set for ordinary regression model due to the burstiness of network traffic and the pervasive network attacks. In this paper, we make use of robust regression techniques to mitigate the impact of outliers in the training data set. The experiment results show that the robust regression based method is more reliable than the ordinary regression based method in the face of outliers.
Ziyu Wang 0007, Jiahai Yang 0001, Shize Zhang
ICC2
2017 Root cause analysis of anomalies of multitier services in public clouds
abstract
Anomalies of multitier services running in cloud platform can be caused by components of the same tenant or performance interference from other tenants. If the performance of a multitier service degrades, we need to find out the root causes precisely to recover the service as soon as possible. In this paper, we argue that cloud providers are in a better position than tenants to solve this problem, and the solution should be non-intrusive to tenants' services or applications. Based on these two considerations, we propose a solution for cloud providers to help tenants to localize root causes of any anomaly. We design a non-intrusive method to capture the dependency relationships of components, which improves the feasibility of root cause localization system. Our solution can find out root causes no matter they are in the same tenant as the anomaly or from other tenants. Our proposed two-step localization algorithm exploits measurement data of both application layer and underlay infrastructure and a random walk procedure to improve its accuracy. Our realworld experiments of a three-tier web application running in a small-scale cloud platform show a 38.9% improvement in mean average precision compared to current methods.
Jianping Weng, Hui Wang 0011, Jiahai Yang 0001, Yang Yang 0004
IWQoS3
2017 Generic and agile service function chain verification on cloud
abstract
Network Function Virtualization (NFV) is an emerging technology to enable network functions (NFs) outsourcing on cloud so as to reduce the costs of deploying and maintaining NFs. However, NF outsourcing poses a serious gap between the expected service function chains (SFCs) and the real enforcement because SFC deployment and management on cloud is invisible to NF customers (i.e., enterprises). In this paper, we propose verifiable SFC, i.e., vSFC, the first scheme that allows an enterprise to accurately verify the correct enforcement of SFC in realtime. In particular, different from the-state-ofthe-art network function verification schemes, vSFC is generic and agile, which can be deployed on various clouds, while not requiring modifications to any NFs on cloud. vSFC detects a wide range of SFC violations including forwarding path incompliance, flow dropping, and packet injection attacks. To demonstrate the feasibility and performance of vSFC, we implement a vSFC prototype built on top of KVM and conduct experiments with real traces. Our experiment results show that vSFC detects various SFC violations with a negligible overhead.
Xiaoli Zhang 0003, Qi Li 0002, Jiahai Yang 0001
IWQoS4
2017 Optimal construction of virtual networks for Cloud-based MapReduce workflows
Jiahai Yang 0001, Kevin Yin, Hui Yu 0004
Comput. Networks2
2016 Load-aware hybrid scheduling in large compute clusters
abstract
With the increasing of workloads in large scale heterogeneous compute clusters, distributed scheduling has won support from the academia and industry because of its inherent scalability and flexibility. However, existing schedulers cannot guarantee that all the jobs are acceptable and the average latency is extremely large. In particular, when the schedulers apply gang scheduling and non-preemptive policy, the serious job starvation problem will be triggered especially in heavily loaded clusters. In this paper, we introduce a novel hierarchical hybrid design of schedulers to address this problem, called En-Omega. In En-Omega, we enhance the fully distributed schedulers with a central scheduler, which can provide global fairness to the jobs from different schedulers and simultaneously reduce the average latency of all the jobs sharply. To reduce the overhead, in our En-Omega design, we activate the central scheduler only when the cluster is heavily loaded. Furthermore, the cache used for central queuing and the scoring policy used in central scheduling are all load-aware. We evaluate En-Omega based on Google trace and experimental results show that, compared to the baseline design, our method can reduce the average latency of starving jobs up to 90% with reasonable overhead.
Di Fu, Jiahai Yang 0001, Hui Zhang 0052
ISCC2
2016 Tetris: Optimizing cloud resource usage unbalance with elastic VM
abstract
Recently, the cloud systems face an increasing number of big data applications. It becomes an important issue for the cloud providers to allocate resources so as to accommodate as many of these big data applications as possible. In current cloud service, e.g., Amazon EMR, a job runs on a fixed cluster. This means that a fixed amount of resources (e.g. CPU, memory) is allocated to the life cycle of this job. We observe that the resources are inefficiently used in such services because of resources usage unbalance. Therefore, we propose a runtime elastic VM approach where the cloud system can increase or decrease the number of CPUs at different time periods for the jobs. There is little change to such services as Amazon EMR, yet the cloud system can accommodate many more jobs. In this paper, we first present a measurement study to show the feasibility and the quantitative impact of adjusting VM configurations dynamically. We then model the task and job completion time of big data applications, which are used for elastic VM adjustment decisions. We validate our models through experiments. We present Tetris, an elastic VM strategy based on cloud system that can better optimize resource utilization to support big data applications. We further implement a Tetris prototype and comprehensively evaluate Tetris on a real private cloud platform using Facebook trace and Wikipedia dataset. We observe that with Tetris, the cloud system can accommodate 31.3% more jobs.
Yi Yuan 0005, Dan Wang 0002, Jiahai Yang 0001
IWQoS4
2016 Towards online anomaly detection by combining multiple detection methods and Storm
abstract
In this paper, we illustrate the significance and advantage of combining the results of multiple detection methods. We implement these methods as bolts in a Apache Storm cluster which is a famous real-time computation framework. We simulate two kinds of anomalies — one involving large number of small network flows and the other involving small number of large network flows. The experiments show that combining multiple methods outperforms any single detection method from the point of view of statistics. Besides, we observe that all the results are outputted in real time without delay, which means that our detection platform is indeed an effective online system.
Ziyu Wang 0007, Jiahai Yang 0001, Hui Zhang 0052, Shize Zhang, Hui Wang 0011
NOMS2
2016 SpongeNet: Towards bandwidth guarantees of cloud datacenter with two-phase VM placement
abstract
In today's production-grade cloud datacenter, cloud service providers do not offer any bandwidth guarantees between VMs, which results in unpredictable performance of tenants' applications. To address this issue, we present SpongeNet, a solution that provides bandwidth guarantees for tenants with a novel network abstraction model and a two-phase VM placement algorithm. Prior solutions have significant limitations: 1) the existing coarse-grained network abstraction models cannot fully express tenants' network requirements and waste a lot of bandwidth resources in demand level; 2) the prior VM placement algorithms, take neither the two scheduling phases nor the tenants' requirements into consideration. As an extension of the existing studies, the proposed network abstraction model in this paper, called Fine-grained Virtual Cluster or FGVC, provides a more precise and flexible way for tenants to specify network requirements and realizes bandwidth saving. SpongeNet also proposes a novel two-phase VM placement algorithm that provides the optimal combinations of ordering policies and dispatching policies in consideration of different goals. Extensive simulations based on real application traces and 3-level tree topology show that SpongeNet provides 48% bandwidth saving than the state-of-art solutions (e.g., the Oktopus system), while significantly improving the throughput rates by 18% and response times by 92%.
Hui Yu 0004, Jiahai Yang 0001, Hui Wang 0011, Zi Liang
NOMS2
2016 Characteristics analysis at prefix granularity: A case study in an IPv6 network
Fuliang Li, Jiahai Yang 0001, Xingwei Wang 0001, Tian Pan 0001, Changqing An
J. Netw. Comput. Appl.2
2016 Joint scheduling of MapReduce jobs with servers: Performance bounds and experiments
abstract
MapReduce-like frameworks have achieved tremendous success for large-scale data processing in data centers. A key feature distinguishing MapReduce from previous parallel models is that it interleaves parallel and sequential computation. Past schemes, and especially their theoretical bounds, on general parallel models are therefore, unlikely to be applied to MapReduce directly. There are many recent studies on MapReduce job and task scheduling. These studies assume that the servers are assigned in advance. In current data centers, multiple MapReduce jobs of different importance levels run together. In this paper, we investigate a schedule problem for MapReduce taking server assignment into consideration as well. We formulate a MapReduce server-job organizer problem (MSJO) and show that it is NP-complete. We develop a 3-approximation algorithm and a fast heuristic design. Moreover, we further propose a novel fine-grained practical algorithm for general MapReduce-like task scheduling problem. Finally, we evaluate our algorithms through both simulations and experiments on Amazon EC2 with an implementation with Hadoop. The results confirm the superiority of our algorithms.
Yi Yuan 0005, Dan Wang 0002, Jiangchuan Liu, Jiahai Yang 0001
J. Parallel Distributed Comput.5
2016 Towards Zero-Time Wakeup of Line Cards in Power-Aware Routers
abstract
As the network infrastructure has been consuming more and more power, various schemes have been proposed to improve the power efficiency of network devices. Many schemes put links to sleep when idle and wake them up when needed. A presumption in these schemes, though, is that router's line cards can be waken up very quickly. However, through systematic measurement of a major vendor's high-end routers, we find that it takes minutes to get a line card ready under the current design. To address this issue, we propose a new line card design that 1) keeps the host processor in a line card standby, which only consumes a small fraction of power but will save considerable wakeup time, and 2) downloads a slim slot of popular prefixes with higher priority, so that the line card will be ready for forwarding most of the traffic much earlier. We design algorithms as well as architecture that ensure fast and correct longest prefix match during prioritized routing prefix download. Experiments on an FPGA-based prototype show that the customized hardware can be ready to forward packets in 127.27 ms, which is 0.3% of the time the original design takes. This can better support numerous power-saving schemes based on the sleep/wakeup mechanism.
Tian Pan 0001, Ting Zhang 0010, Junxiao Shi, Yang Li 0062, Linxiao Jin, Fuliang Li, Jiahai Yang 0001, Beichuan Zhang 0001, Xueren Yang, Mingui Zhang, Huichen Dai, Bin Liu 0001
IEEE/ACM Trans. Netw.7
2015 MOE-A framework integrating network performance monitoring, optimization and evaluation
abstract
Due to network dynamics, performance tuning is often indispensable in network management. In this paper, we propose MOE, a framework integrating network performance monitoring, optimization and evaluation. This is a trial towards the top-down and systematic management of network performance. We validate MOE based on a typical scenario in the real network environment. Results show that MOE can collect many kinds of network information, based on which it could conduct performance tuning automatically. In addition, MOE has the ability of evaluating the effect during and after performance tuning. Evaluation results are further analyzed and could fed back to provide positive advices to minimize the influence caused by network adjustments and maximize the performance profits.
Fuliang Li, Jiahai Yang 0001, Xingwei Wang 0001
APNOMS2
2015 A Lightweight DDoS Flooding Attack Detection Algorithm Based on Synchronous Long Flows
abstract
DDoS flooding attack is one of the top threats to the Internet. However, due to the fast development of the Internet, current detection algorithms are already inadequate to meet the growth of network traffic. In this paper, we propose a lightweight algorithm. We first observe the real Internet traffic, and find that flows of DDoS flooding attack traffic are persistent and synchronous while most flows of normal traffic are short-lived and non- synchronous. According to this difference, we propose our detection algorithm. We label the alarms firstly and then confirm the attack. Our algorithm is lightweight and sensitive to the ongoing attack. However, randomly spoofing the IP address of the attack source to different IP addresses can hide the synchronization of attack flows. Thus, we add a spoofing IP detection algorithm called hop-count filter (HCF) to our algorithm to strengthen the robustness. At last, we evaluate our detection algorithm based on the real Internet traffic from CAIDA. Results show that our detection algorithm has a high accuracy (93.3%), no false positive in attack confirmation and just 1.1% false positive rate in labeling alarms. In addition, we analyze the challenges we may face when dealing with distributed LDoS attack.
Jiahai Yang 0001, Ziyu Wang 0007, Fuliang Li, Yang Yang 0004
GLOBECOM2
2015 MuLTI: Multiple location tags inference for users in social networks
abstract
Social networks, with tremendous popularity all over the world, have become the most important platform for many services in the past years. Location, as part of users' basic information, is always the key to many recommendation services in social networks. Most of the previous research works focus on inferring on the users' home locations. However, it is not enough as many people in social networks have multiple location tags, including home location, work location and on. In this paper, we propose a multiple location tags inference algorithm, i.e. MuLTI to build complete location profiles for users in social networks. We formulate the correlations between the users' location tags and their friendships, tweets, and then infer the users' locations in each of their friendships and tweets. It reflects the activity level of users to be in different locations. Apart from the activity level, we also consider the time span of users to be in different locations, so as to infer the users' long-term location tags better, as we find that users may also be active in their temporal locations. Experiments show that MuLTI improves the precision by about 15%, and the recall by about 25% compared with the state-of-the-art algorithms.
Zejia Chen, Jiahai Yang 0001, Hui Wang 0011
ISCC2
2015 Nexthop-Selectable FIB aggregation: An instant approach for internet routing scalability
Qing Li 0006, Mingwei Xu 0001, Dan Wang 0002, Jun Li 0001, Yong Jiang 0001, Jiahai Yang 0001
Comput. Commun.6
2015 Solving multicast problem in cloud networks using overlay routing
Hui Wang 0011, Jeffrey Cai, Jerry Lu, Kevin Yin, Jiahai Yang 0001
Comput. Commun.5
2014 Evolution of network configurations: High-level analysis of an operational IP backbone network
abstract
In this paper, we gather the weekly reports of an operational IP backbone network from January 2006 to January 2013, according to which, we can restore the truth and uncover the evolution of network configurations of the studied network. Our high-level analyses illustrate that rate limiting and launching routes for new customers are most frequently configured. We can identify and construct configuration templates by correlating each task to a certain set of commands in configuration files, based on which, automated configuration provisioning for an operational backbone network is feasible. In addition, we can configure redundant links for those with higher rate of failures according to our detailed analyses of link failures, which will enhance the stability and reliability of data transmission.
Fuliang Li, Jiahai Yang 0001, Huijing Zhang, Suogang Li, Xingwei Wang 0001
APNOMS2
2014 An on-line anomaly detection method based on LMS algorithm
abstract
Anomaly detection has been a hot topic in recent years due to its capability of detecting zero attacks. In this paper, we propose a new on-line anomaly detection method based on LMS algorithm. The basic idea of the LMS-based detector is to predict IGTE using IGFE, given the high linear correlation between them. Using the artificial synthetic data, it is shown that the LMS-based detector possesses strong detection capability, and its false positive rate is within acceptable scope.
Ziyu Wang 0007, Jiahai Yang 0001, Fuliang Li
APNOMS2
2014 Two-stage detection algorithm for RoQ attack based on localized periodicity analysis of traffic anomaly
abstract
Reduction of Quality (RoQ) attack is a stealthy denial of service attack. It can decrease or inhibit normal TCP flows in network. Victims are hard to perceive it as the final network throughput is decreasing instead of increasing during the attack. Therefore, the attack is strongly hidden and it is difficult to be detected by existing detection systems. Based on the principle of Time-Frequency analysis, we propose a two-stage detection algorithm which combines anomaly detection with misuse detection. In the first stage, we try to detect the potential anomaly by analyzing network traffic through Wavelet multiresolution analysis method. According to different time-domain characteristics, we locate the abrupt change points. In the second stage, we further analyze the local traffic around the abrupt change point. We extract the potential attack characteristics by autocorrelation analysis. By the two-stage detection, we can ultimately confirm whether the network is affected by the attack. Results of simulations and real network experiments demonstrate that our algorithm can detect RoQ attacks, with high accuracy and high efficiency.
Jiahai Yang 0001, Fengjuan Cheng, Ziyu Wang 0007
ICCCN2
2014 Towards zero-time wakeup of line cards in power-aware routers
abstract
As the network infrastructure has been consuming more and more power, various schemes have been proposed to improve power efficiency of network devices. Many schemes put links to sleep when idle and wake them up when needed. A presumption in these schemes, though, is that router's line cards can be waken up quickly. However, through systematic measurement of a major vender's high-end router, we find that it takes minutes to get a line card ready under the current implementation. To address this issue, we propose a new line card design that (1) keeps the host processor in a line card always up, which only consumes a small fraction of power, and (2) downloads a slim slot of popular prefixes with higher priority, so that the line card will be ready for forwarding most of the traffic much earlier. We design algorithms that ensure fast and correct longest prefix match lookup during prioritized routing prefix download. Experiments on real hardware show that the wakeup time can be reduced to 127.27ms, which is 0.3% of the original line card wakeup time, well supporting many power-saving schemes.
Tian Pan 0001, Ting Zhang 0010, Junxiao Shi, Yang Li 0062, Linxiao Jin, Fuliang Li, Jiahai Yang 0001, Beichuan Zhang 0001, Bin Liu 0001
INFOCOM7
2014 A cascading framework for uncovering spammers in social networks
abstract
With tremendous popularity, OSNs have become the most important platform for marketing and advertising during the past years. Meanwhile, spamming has already become a very serious problem in OSNs, drawing the attention of both academic and industry communities. In this paper, we investigate the problem of spammer detection from the perspective of user behaviors, including relation creation, user activeness, user interaction and tweet content. We quantitatively explore their correlations with spammer detection and find that tweet content is the most important factor for spammer detection, followed by relation creation. Based on these behavior factors, we propose a novel cascading framework CWB-SPAM for spammer detection in OSNs. Experiments on dataset crawled from Sina Microblog show that the proposed algorithm outperforms over all classical algorithms we investigated in terms of F-score1• Experiments also demonstrate that as a probabilistic classification model, the proposed CWB-SPAM has a good ranking quality. It enables the OSN operators to make tradeoff between precision and recall easily so that the proposed algorithm can be used in different scenarios. Besides, we also note that the proposed framework can be used in other probabilistic binary classification models and thus applied in more scenarios.
Zejia Chen, Jiahai Yang 0001, Hui Wang 0011
Networking2
2014 A value based framework for provider selection of regional ISPs
abstract
To access to the Internet, regional Internet Service Providers (ISPs) have to buy transit service from global ISPs. Provider selection strategies are related closely to ISPs' economic interests. With the growing number of potential transit provider and the flattening topology of the Internet, it's getting harder for ISPs to select upstream provider empirically as before. In this paper, we propose a concept of bargaining power as an important decision-making criterion of ISPs during their provider selection process, and design a value based framework to help ISPs' provider selection based on it. The bargaining power of each involved ISP is computed by applying the Shapley Value based transit value distribution mechanism to each involved traffic flow, taking into consideration the cooperative possibility among ISPs and the market roles these ISPs play, i.e., potential providers or potential competitors. It reflects not only the cost and link level transit performance, geographical constraints, but also includes the influence of interconnection impacts, demand/supply relationships by analyzing the traffic content and commercial relationships among ISPs. We then design a quantitative provider selection framework and instantiate our framework using the operation data of a real-world network, CERNET, a national ISP in China. In addition, we evaluate our provider selection results for CERNET and the experimental results show the effectiveness and practicability of our solution in this paper.
Hui Wang 0011, Jiahai Yang 0001
NOMS3
2014 Towards Optimal Collaboration of Policies in the Two-Phase Scheduling of Cloud Tasks
Jiahai Yang 0001, Di Fu, Hui Zhang 0052
NPC2
2014 A New Anomaly Detection Method Based on IGTE and IGFE
Ziyu Wang 0007, Jiahai Yang 0001, Fuliang Li
SecureComm (2)2
2014 An On-Line Anomaly Detection Method Based on a New Stationary Metric - Entropy-Ratio
abstract
Anomaly detection has been a hot topic in recent years due to its capability of detecting zero day attacks. In this paper, we propose a new metric called Entropy-Ratio. We validate that the Entropy-Ratio is stationary. Making use of this observation, we combine the Least Mean Square algorithm and the Forward Linear Predictor to propose a new on-line detector called LMS-FLP detector. Using the two synthetic data sets - CEGI-6IX synthetic data and CERNET2 synthetic data, we validate that the LMS-FLP detector is very effective in detecting both anomalies involving many small IP flows and anomalies involving a few large IP flows.
Ziyu Wang 0007, Jiahai Yang 0001, Fuliang Li
TrustCom2
2014 A study of traffic from the perspective of a large pure IPv6 ISP
Fuliang Li, Changqing An, Jiahai Yang 0001, Hui Zhang 0052
Comput. Commun.3
2014 Configuration analysis and recommendation: Case studies in IPv6 networks
Fuliang Li, Jiahai Yang 0001, Zhiyan Zheng, Huijing Zhang, Xingwei Wang 0001
Comput. Commun.2
2013 CSS-VM: A centralized and semi-automatic system for VLAN management
Fuliang Li, Jiahai Yang 0001, Changqing An
IM2
2013 IPv6 network topology discovery method based on novel graph mapping algorithms
abstract
As a crucial function of network management, network topology discovery provides a basis for lots of network analysis, such as network monitoring and performance management, etc. With the undergoing deployment of IPv6, the importance of precise topology discovery method in IPv6 networks becomes more and more evident. However, IPv6 network topology discovery faces new challenges due to different characteristics between IPv4 and IPv6, and the lack of well support of IPv6 related MIBs from device manufacturers in current state. At present, there are no well-accepted topology discovery methods for pure IPv6 networks with high accuracy, high coverage and less reliance on network configuration and device support. In this paper, we propose an IPv6 network topology discovery solution combining the advantages of two discovery methods, based on ICMP and routing protocol respectively. We model the mapping process of topology results from the two methods above into a graph mapping problem, which is the key point of the entire solution, and design novel mapping algorithms. We focus on the mapping coverage and accuracy and validate the mapping algorithms by large scale simulation. We also implement and test the proposed algorithms on the real network CERNET2. The experiments and simulation results verify the practicability and excellent performance of our solutions, with 100% discovery accuracy and over 99% discovery coverage while spending less time and producing lower overhead.
Jiahai Yang 0001, Changqing An, Fuliang Li
ISCC2
2013 MBST: Detecting Packet-Level Traffic Anomalies by Feature Stability
abstract
In this paper, we present a statistical analysis of six traffic features based on entropy and distinct feature number at the packet level, and we find that, although these traffic features are unstable and show seasonal patterns like traffic volume in a long-time period, they are stable and consistent with Gaussian distribution in a short-time period. However, this equilibrium property will be violated by some anomalies. Based on this observation, we propose a Multi-dimensional Box plot method for Short-time scale Traffic (MBST) to classify abnormal and normal traffic. We compare our new method with the MCST method proposed in our prior work and the well-known wavelet-based and A Short-Timescale Uncorrelated-Traffic Equilibrium (ASTUTE) techniques. The detection result on synthetic anomaly traffic shows that MBST can better detect the low-rate attacks than wavelet-based and MCST methods, and detection result on real traffic demonstrates that MBST can detect more anomalies with lower false alarm rate than the two methods. Especially compared with ASTUTE, MBST performs much better for detecting anomalies involving a few large flows despite a little poor for detecting anomalies involving large number of small flows.
Bin Zhang 0016, Jiahai Yang 0001, Ziyu Wang 0007
Comput. J.2
2013 An efficient critical protection scheme for intra-domain routing using link characteristics
Mingwei Xu 0001, Meijia Hou, Dan Wang 0002, Jiahai Yang 0001
Comput. Networks4
2012 AMIR: Another Multipath Interdomain Routing
abstract
Multipath routing is an important and promising technique to increase the Internet's reliability and to give users greater control over the service they receive. Currently the interdomain routing protocol limits each router to using a single route for a destination network, which does not satisfy the diverse requirements of end users. In this paper, in order to support the effective and efficient multipath routing, we propose a multipath interdomain routing system(AMIR), which not only provides more novel paths but also realizes a new AS-level routing scheme. In the control plane, the topology information is collected from neighboring ASes around the primary path, and based on this topology, multipath of the special node pairs is calculated by our multipath discovery algorithm. In the data plane, we use the interdomain source routing to forward the packets. Experiments with Internet topology and routing data demonstrate that AMIR is practical and feasible, and offers tremendous flexibility and diversity for path selection with reasonable overhead.
Donghong Qin, Jiahai Yang 0001, Zhuolin Liu, Hui Wang 0011, Bin Zhang 0016
AINA2
2012 Flattening and preferential attachment in the internet evolution
abstract
Understanding of the Internet evolution is important for many research topics, such as network planning, optimal routing design, etc. In this paper, we try to analyze CAIDA AS-level topology dataset from 2004 to 2010 to validate two conjectures on the Internet evolution, i.e., the Internet flattening trend and the preferential attachment rule. Our analysis shows that the evolvement of the Internet core is different from the edge of Internet. We classify the Internet into several layers using different layering methods, i.e., Rich Club coefficient based method, k-core decomposition method and SARK hierarchy model, and then study the changes of the features of these layers. Under all of these laying methods, we find that the boundaries between neighboring layers in the Internet core are more and more blurred; ASes in the core distribute more evenly and different layers are closer to each other in size, while the Internet edge still has a distinct hierarchical characteristic. It is more evident in Asia and Europe than North America. The other difference between Internet core and Internet edge is that link births/deaths in the Internet core follow the “Preferential Attachment/de-attachment” rule, while link births/deaths in the Internet edge follow a super linear preferential attachment/de-attachement rule. On the other hand, in both Internet core and Internet edge, link births caused by AS births present stronger preference than link rewiring.
Hui Wang 0011, Jiahai Yang 0001
APNOMS3
2012 Separating identifier from locator with extended DNS
abstract
Although most researchers have agreed that the locator/identifier separation is beneficial for the Internet, there is no consensus on how to define the “identifier”. In this paper, we propose a scheme in which identifiers are distributed by authorities to endpoints, and the authorities are responsible for maintaining real-time locators of endpoints with identifiers they distributed. This scheme can be helpful to accounting, security and other network management tasks. We also present in details how to implement this scheme with extended DNS and a new infrastructure, i.e., ID Mapping System (IDMS).
Hui Wang 0011, Mingwei Xu 0001, Jiahai Yang 0001
ICC4
2012 Unravel the characteristics and development of current IPv6 network
abstract
In this paper, many aspects related to characteristics and development of IPv6 network are investigated. Additionally, in order to gain a deep view of IPv6 network, we correlate our system with a user authentication system, so we explore some meaningful user behaviors. According to the analysis, we obtain a comprehensive knowledge of current operating situation of IPv6 network which, we believe, can provide an experimental basis for IPv6 network operators and researchers.
Fuliang Li, Changqing An, Jiahai Yang 0001, Zejia Chen
LCN3
2012 What's going on in Chinese IPv6 world
abstract
With IPv4 addresses quickly dwindling, the Internet is forcing an evolution of itself. During the long term transition from IPv4 to IPv6, what's going on in IPv6 world becomes unknown for network operators and researchers. In this paper, we propose a heuristic algorithm to identify p2p traffic accurately and implement traffic classification based on Netflow v9 exports to illustrate what applications Chinese IPv6 users are really running. Additionally, we present a detailed study of p2p traffic over IPv6 and advice ISPs to localize p2p traffic at the AS level for future IPv6 traffic management and network resources planning, leaving modeling traffic behavior and deeper classification of IPv6 traffic as our future work.
Jiahai Yang 0001, Hui Zhang 0052, Donghong Qin, Bin Zhang 0016
NOMS2
2012 PCA-subspace method - Is it good enough for network-wide anomaly detection
abstract
PCA-subspace method has been proposed for network-wide anomaly detection. Normal subspace contamination is still a great challenge for PCA although some methods are proposed to reduce the contamination. In this paper, we apply PCA-subspace method to six-month Origin-Destination (OD) flow data from the Abilene. The result shows that normal subspace contamination is mainly caused by anomalies from a few strongest OD flows, and seems unavoidable for subspace method. Further comparison of anomalies detected by subspace method and manually tagged anomalies from each OD flows, we find that anomalies detected by subspace method are mainly caused by anomalies from medium and a few large OD flows, and most anomalies of minor OD flows are buried in abnormal subspace and hard to be detected by PCA-subspace method. We analyze the reason for those anomalies undetected by subspace method and suggest to use normal subspace to detect anomalies caused by a few strongest OD flows, and to further divide abnormal subspace to detect more anomalies from minor OD flows. The goal of this paper is to address limitations neglected by prior works and further improve the subspace method on one hand, also call for novel detection methods for network-wide traffic on another hand.
Bin Zhang 0016, Jiahai Yang 0001, Donghong Qin
NOMS2
2012 Diagnosing Traffic Anomalies Using a Two-Phase Model
Bin Zhang 0016, Jiahai Yang 0001, Ying-Wu Zhu
J. Comput. Sci. Technol.2
2011 FlowInfra: A fault-resilient scalable infrastructure for network-wide flow level measurement
abstract
The fine-grained flow level measurement is getting increasing demand in recent years. Though it fails to be a generic solution for its biased sampling, NetFlow is promising for its compatibility with major routers and its convenience to perform direct flow level measurement of both IPv4 and IPv6 traffic. Traditional flow level measurement systems based on NetFlow are mostly centralized and each of them independently performs traffic analysis of its local flow records without any coordination in a large-scale network, suffering from unbalancing workload and bad scalability. In this paper we present the design, implementation and evaluation of FlowInfra which is a fault-resilient scalable infrastructure for network-wide flow measurement of pure IPv6 flow records from NetFlow v9 exports. Through the assessment of its performance and flexible features, we show that FlowInfra achieved enhanced ability and robustness to perform network-wide flow level measurement and satisfied the goal for IPv6 network operation and management with better scalability.
Jiahai Yang 0001, Hui Zhang 0052, Bin Zhang 0016, Donghong Qin
APNOMS2
2011 Adaptive tuning of operation parameters for automatically learned filter table
abstract
Automatically learned filter table is used in many network security mechanisms to validate packets. Building filter item for each IP address in access networks can prevent IP spoofing at fine granularity but may consume large amount of filter table which is limited due to the expensive storage which is usually TCAM for high speed access. It is an urgent problem to use filter table effectively and keep network available. We analyze the change of filter table size and find that setting proper lifetime for filter item can significantly improve the utilization of filter table and avoid denial of service. In this paper, we take SAVI (source address validation improvement) switch as an example, and propose a dynamic adjustment method. It has two phases. Firstly it calculates out an optimal lifetime value for each switch based on one week user online logs, and then adjusts it dynamically to capture the bursts of filter table size. We deploy our prototype in a real campus network which has about 1000 SAVI switches providing network accessing service for nearly 20000 users. Based on the analysis of one month user online logs, we verify that our algorithm can reduce 92% of the duplicate confirming processes and guarantee the availability of network.
Changqing An, Jiahai Yang 0001
APNOMS3
2011 Investigating the efficiency of fine granularity source address validation in IPv6 networks
abstract
IPv6 protocol has been widely deployed in the world. As the IANA pool of IPv4 addresses has run out, IPv6 will become increasingly important. Although the IPv6 protocol stack presents considerable advantages compared with the IPv4 protocol stack, IP source address spoofing is still exploited in IPv6 to initiate malicious attacks. Some techniques are proposed and deployed to implement source address validation at fine granularity. In this paper, we investigate the efficiency of fine granularity IP source address validation, e.g. whether filtering technology is deployed to prevent hosts from using forged IP address. We develop a detection tool with controlled spoofing ability which can infer whether the function of filtering spoofing address packets is enabled. We run this tool in 12 famous universities in China and collect the testing data. We gather a total of 41373 probes from 324 clients, and each probe includes sending at least 5 packets with the same spoofing source address to the control server. Results reveal that, 77.02% of the spoofing probes are completely filtered, 0.29% of the spoofing probes are partly filtered and the rest spoofing probes are not filtered at all. Overall, this illustrates that techniques of source address validation have been widely deployed in campus networks. Our statistical results provide practical basis for the deployment and further development of source address validation protocols in IPv6 networks.
Fuliang Li, Changqing An, Jiahai Yang 0001
APNOMS3
2011 MCST: Anomaly detection using feature stability for packet-level traffic
abstract
In this paper, we present a statistical analysis of six traffic features based on entropy and distinct feature number at the packet level, and we find that, although these traffic features are unstable and show seasonal patterns like traffic volume for a long period, they are stable and consistent with Gaussian distribution in a short time period. However, this equilibrium property will be violated by some anomalies. Based on this observation, we propose a Multi-dimensional Clustering method for Short-time scale Traffic(MCST) to classify abnormal and normal traffic. We compare our new method to the well known wavelet technique. The detection result on synthetic anomaly traffic shows MCST can better detect the low-rate attacks than wavelet-based method, and detection result on real traffic demonstrates that MCST can detect more anomalies with low false alarm rate.
Bin Zhang 0016, Jiahai Yang 0001, Donghong Qin
APNOMS2
2011 An efficient parallel TCAM scheme for the forwarding engine of the next-generation router
abstract
Ternary Content-Addressable Memory (TCAM) is a popular hardware device for fast IP address lookup. High link transmission speed of Internet backbone demands more powerful IP address lookup engine. Restricted by the memory access speed, the lookup engine for next-generation routers demands exploiting parallelism among multiple TCAM chips. However, most existing schemes improve lookup performance and reduce power consumption but ignore the update efficiency. In this paper, we propose a crossed address range division and shared caching scheme. We improve the update efficiency significantly by buddy update method while keep low power dissipation by decreasing the number of the triggered TCAMs access in each lookup operation. The lookup throughput is ultra high through adaptive load balance. Our simulation results show that the proposed scheme can achieve an average lookup speedup factor greater than 11 with 12 TCAM chips, on the cost of 10% more memory space and an additional cache chip.
Bin Zhang 0016, Jiahai Yang 0001, Qi Li 0002, Donghong Qin
Integrated Network Management2
2011 On the scalability of router forwarding tables: Nexthop-Selectable FIB aggregation
abstract
In recent years, the core-net routing table, e.g., Forwarding Information Base (FIB), is growing at an alarming speed and this has become a major concern for Internet Service Providers. One effective solution for this routing scalability problem, which requires only upgrades on individual routers, is FIB aggregation. Intrinsically, IP prefixes with numerical prefix matching and the same next hop can be aggregated. Very commonly, all previous studies assume that each IP prefix has one corresponding next hop, i.e., towards one optimal path. In this paper, we argue that a packet can be delivered to its destination through a path other than the one optimal path. Based on this observation, we for the first time propose Nexthop-Selectable FIB Aggregation that is fundamentally different from all previous aggregation schemes. IP prefixes are aggregated if they have numerical prefix matching and share one common next hop. Consequently, IP prefixes that cannot be aggregated, due to lack of the same next hop, are aggregated; and we achieve a substantially higher aggregation ratio. In this paper, we provide a systematic study on this Nexthop-Selectable FIB Aggregation problem. We present several practical choices to build the sets of selectable next hops for the IP prefixes. To maximize the aggregation, we formulate the problem as an optimization problem. We show that the problem can be solved by dynamic programming. While the straightforward application of dynamic programming has exponential complexity, we propose a novel algorithm that is O(N). We then develop an optimal online algorithm with constant running time. We evaluate our algorithms through a comprehensive set of simulations with BRITE with RIBs collected from RouteViews. Our evaluation shows that we can reduce more than an order of the FIB size.
Qing Li 0006, Dan Wang 0002, Mingwei Xu 0001, Jiahai Yang 0001
INFOCOM4
2011 MIB design and application for source address validation improvement protocol
abstract
In this paper, we present SAVI-MIB, a management information base (MIB) designed to support configuration and monitoring of SAVI protocol which can provide fine granularity source address validation. Objects are defined to meet the detailed management requirement of local networks and accommodate different scenarios. SAVI-MIB is implemented in switches and deployed in some campus networks. Objects of SAVI-MIB are retrieved and used to help find configuration errors in SAVI deployment and profile behavior of end hosts. SAVI-MIB can also be used in parameter optimization, auto configuration, anomaly detection, etc.
Changqing An, Hui Wang 0011, Jiahai Yang 0001
ISCC3
2011 A study of traffic, user behavior and pricing policies in a large campus network
Hui Wang 0011, Changqing An, Jiahai Yang 0001
Comput. Commun.3
2011 A study on key strategies in P2P file sharing systems and ISPs' P2P traffic management
Hui Wang 0011, Chungang Wang, Jiahai Yang 0001, Changqing An
Peer-to-Peer Netw. Appl.3
2010 A measure of growth of user community in OSNs
abstract
Online Social Networks (OSNs) become more and more popular in recent years, with almost several hundred million users involved. A lot of research efforts have been done on features like analysis of the structure of user network, statistics of static properties and dynamic characteristics in OSNs, while little is known about the way how the communities of users grow bigger. This paper proposes a new methodology in OSN research to study the long-run evolution of user community and give guidance for managing network traffic and web caching policy, and a few discussions are also presented.
Jiahai Yang 0001, Hui Wang 0011, Hui Zhang 0052
IWQoS2
2010 Efficient parallel searching with TCAMs
abstract
Chip-level parallel TCAMs were deployed to circumvent the limitation of a single TCAM. In this paper, we propose a crossed address range division and integral caching scheme. We improve the update efficiency significantly by buddy update method while keep low power dissipation by decreasing the number of the triggered TCAMs access in each lookup operation. The lookup throughput is ultra high through adaptive load balance. Our simulation results show that the proposed scheme can achieve an average lookup speedup factor greater than 11 with 12 TCAM chips, given 10% more memory space.
Bin Zhang 0016, Qi Li 0002, Jiahai Yang 0001
IWQoS4
2010 Efficient searching with parallel TCAM chips
abstract
Ternary Content-Addressable Memory (TCAM) is a popular hardware device for fast IP address lookup. High link transmission speed of Internet backbone demands more powerful IP address lookup engine. Restricted by the memory access speed, the lookup engine for next-generation routers demands exploiting parallelism among multiple TCAM chips. In this paper, we propose a fast lookup and power-saving scheme which can almost make full use of TCAM chips' capability. With N parallel TCAM chips, the scheme can achieve a worst-case speedup factor of (N-1)*90% in cache-update state, and a worst-case speedup factor of N*90% in normal work state.
Bin Zhang 0016, Jiahai Yang 0001, Qi Li 0002
LCN2
2009 Understanding Web Hosting Utility of Chinese ISPs
Guanqun Zhang, Hui Wang 0011, Jiahai Yang 0001
APNOMS3
2009 Selective Protection: A Cost-Efficient Backup Scheme for Link State Routing
abstract
In recent years, there are substantial demands to reduce packet loss in the Internet. Among the schemes proposed, finding backup paths in advance is considered to be an effective method to reduce the reaction time. Very commonly, a backup path is chosen to be a most disjoint path from the primary path, or in the network level, backup paths are computed for all links (e.g., IPRFF). The validity of this straightforward choice is based on 1) all the links may fail with equal probability; and 2) facing the high protection requirement today, having links not protected or sharing links between the primary and backup paths just simply look weird. Nevertheless, indications from many research studies have confirmed that the vulnerability of the links in the Internet is far from equality. In addition, we have seen that full protection schemes may introduce high costs. In this paper, we argue that such approaches may not be cost effective. We first analyze the failure characteristics based on real world traces from CERNET2, the China education and Research NETwork 2. We observe that the failure probabilities of the links is heavy-tail, i.e., a small set of links caused most of the failures. We thus propose a selective protection scheme. We carefully analyze the implementation details and the overhead for general backup path schemes of the Internet today. We formulate an optimization problem where the routing performance (in terms of network level availability) should be guaranteed and the backup cost should be minimized. This cost is special as it involves computation overhead. Consequently, we propose a novel Critical-Protection Algorithm which is fast itself. We evaluate our scheme systematically, using real world topologies and randomly generated topologies. We show significant gain even when the network availability requirement is 99.99\% as compared to that of the full protection scheme.
Meijia Hou, Dan Wang 0002, Mingwei Xu 0001, Jiahai Yang 0001
ICDCS4
2009 Towards Next Generation Internet Management: CNGI-CERNET2 Experiences
Jiahai Yang 0001, Hui Zhang 0052, Jinxiang Zhang, Changqing An
J. Comput. Sci. Technol.1
2008 Understanding IPv6 Usage: Communities and Behaviors
Shaojun Huang, Changqing An, Hui Wang 0011, Jiahai Yang 0001
APNOMS4
2008 Traffic Matrix Estimation Using Square Root Filtering/Smoothing Algorithm
Jiahai Yang 0001, Yang Yang 0004, Guanqun Zhang
APNOMS2
2007 Internet Management Network
Jilong Wang 0001, Jiahai Yang 0001
APNOMS3
2005 Traffic Measurement and Analysis of TUNET
abstract
Traffic measurement and analysis, as one of the important methods of understanding and characterizing network, can provide significant support for network management. After a brief introduction of a novel NP (network processor)-based architecture of traffic measurement, the paper presents the detailed analysis results of the traffic collected from the gigabit link connecting Tsinghua University campus network (in short, TUNET) to its upstream ISP, China Education and Research NETwork (in short, CERNET). Then, the paper comprehensively analyzes the traffic from multi-dimension viewpoints, including temporal distribution, packet length distribution, port-based distribution, protocol-based distribution, and TopN statistics. Such analysis not only provides support for the study of user behavior, but also enriches traffic measurement technology
Jun Zhang 0004, Jiahai Yang 0001, Changqing An, Jilong Wang 0001
CW2
2003 A nonstationary traffic train model for fine scale inference from coarse scale counts
abstract
The self-similarity of network traffic has been convincingly established based on detailed packet traces. This fundamental result promises the possibility of solving on-line and off-line traffic engineering problems using easily collectible coarse time-scale data, such as simple network management protocol measurements. This paper proposes a statistical model that supports predicting fine time-scale behavior of network traffic from coarse time-scale aggregate measurements. The model generalizes the commonly used fractional Gaussian noise process in two important ways: (1) it accommodates the recurring daily load patterns commonly observed on backbone links and (2) features of long range dependence and self-similarity are modeled only at fine time scales and are progressively damped as the time period increases. Using the data we collected on the Chinese Education and Research Network, we demonstrate that the proposed model fits 5-min data and generates 10-s aggregates that are similar to actual 10-s data.
Chuanhai Liu, Scott A. Vander Wiel, Jiahai Yang 0001
IEEE J. Sel. Areas Commun.3
2000 CARDS: A Distributed System for Detecting Coordinated Attacks
Jiahai Yang 0001, Peng Ning, Xiaoyang Sean Wang, Sushil Jajodia
SEC1