Xiaoyan Hu 0007

dblp:66/1557-7 · DBLP profile ↗
← Back
71ranked-venue papers
24as first author
61since 2021 · last 2026
0000-0002-4172-1977ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 58 · 21 first-author · 49 since 2021Security and privacy · 11 · 2 first-author · 11 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LARSS: A Hardware-Software Co-designed Framework for Load-Aware Receive Side Scaling
Guang Cheng 0001, Hua Wu 0004, Deyu Zhao, Yuyu Zhao, Xiaoyan Hu 0007
IWQoS6
2026 FSG-NID: Early network intrusion detection via flow segment graph analysis
Bayi Xu, Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
Comput. Networks2
2026 Toward Evolvable IoT Device-Type Identification Using Few-Shot Incremental Learning
abstract
The proliferation of Internet of Things (IoT) devices has profoundly transformed various industries. However, with the continual emergence of new IoT device types, their extensive heterogeneity and weak security have posed significant challenges for network management and security. Despite existing research excelling at classifying IoT traffic, they have yet to consider incremental updates and timely identification, which are critical for early device management and security in IoT networks. As a solution, we present EAPN, a novel and evolvable IoT device identification model that supports adapting to new IoT devices with limited traffic. EAPN extracts traffic features from only a few dozen packets and resorts to a metric-learning-based triplet network to capture accurate, discriminative behavioral representations among IoT devices. Then, inspired by Few-Shot Incremental Learning (FSIL), EAPN further considers advancing incremental model updates based on the relationship between traffic features of new and existing IoT devices. Extensive evaluations demonstrate that EAPN not only surpasses state-of-the-art adaptive IoT device-type identification methods but also achieves over 90% accuracy under almost all incremental conditions, maintaining satisfactory performance even after multiple updates.
Bowen Ouyang, Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
IEEE Internet Things J.2
2026 STMWF: Multi-Tab Website Fingerprinting via Spatial-Temporal Sequence Analysis
abstract
Website fingerprinting (WF) attacks are employed to identify websites that utilize Tor encryption. Although State-Of-The-Art (SOTA) WF attacks demonstrate strong performance in single-tab scenarios, they face challenges in multi-tab scenarios. Many multi-tab WF attacks rely solely on direction sequence or process directional and temporal sequence separately. They ignore the coupling between directional and temporal features, which reflects distinct resource-loading processes for different websites. To address the limitations of existing approaches, this paper proposes a new multi-tab WF attack, STMWF. It leverages spatial-temporal sequence analysis and jointly models Inter Arrival Time (IAT) with the direction sequence. STMWF utilizes an SE-attention-based feature extractor to derive features from various website resources within the spatial-temporal sequence. It then employs correlation self-attention to integrate these resource features into their respective websites, ultimately constructing distinct fingerprints for each site. Additionally, the method incorporates correlation denoising to suppress noise in the website fingerprints, thereby enhancing the discriminability of the extracted features. We collected single-tab traces to synthesize a dataset with controlled overlap ratios. We also captured real-world multi-tab traffic with varying tab-opening intervals, evaluating performance under authentic conditions. The experimental results indicate that STMWF significantly outperforms the SOTA multi-tab attacks in both dynamic and static settings. Specifically, it achieves an average F1-score improvement of approximately 14.87% under static conditions and 34.81% under dynamic conditions compared to the SOTA multi-tab WF attack, ARES. Furthermore, STMWF exhibits greater robustness against WF defenses than SOTA attacks and consistently surpasses them across varying overlapping scenarios.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
IEEE Trans. Inf. Forensics Secur.2
2026 Data-Efficient Cross-Domain Few-Shot Website Fingerprinting With Unsupervised Domain Adaptation
abstract
Website fingerprinting (WF) attacks identify Torencrypted websites but struggle with cross-domain scenarios due to traffic distribution shifts. The existing few-shot WF attacks address the cross-domain problem with excessive auxiliary data, significantly reducing deployment efficiency. This work proposes UDA-WF, a data-efficient few-shot WF with Unsupervised Domain Adaptation (UDA). UDA-WF first pre-trains the feature extractor with limited auxiliary data in the source website domain. Then, it extracts the invariant feature space by computing the intersection of the source and target feature spaces through the unsupervised domain adaptation with the softmatch mechanism. Finally, UDA-WF fine-tunes the feature extractor and a single-layer perceptron to extract the discriminative unique feature space of the target website domain. We evaluate UDA-WF on our WF dataset collected over multiple months. UDA-WF significantly overcomes the cross-domain problem while reducing auxiliary data requirements by 95% and pre-training bootstrap time by 99% compared to the State-Of-The-Art (SOTA) methods. UDA-WF achieves an accuracy of 97.37% under the 20-shot setting in the closed-world scenario and outperforms SOTA methods. To further demonstrate the model’s adaptability to diverse real-world requirements, we validate it on the DF and Wang datasets, achieving accuracies exceeding 92% and 94%, respectively. Moreover, the results show that our UDA-WF is more resilient to concept drift and robust to WF defense.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
IEEE Trans. Inf. Forensics Secur.2
2026 A Generalized Video Platform Identification Method Over Obfuscated Encrypted Protocols in Real-World Networks
Hua Wu 0004, Anting Lu, Guang Cheng 0001, Xiaoyan Hu 0007
IEEE Trans. Netw. Serv. Manag.6
2026 VeCroToken: An Efficient, Verifiable, and Privacy-Preserving Cross-Chain Model for Consortium Blockchains Based on zk-SNARKs
abstract
Consortium blockchains enable secure economic applications through privacy-preserving architectures and efficient processing. Growing cross-chain demands require value-exchange mechanisms, yet expose privacy risks during external interactions. Encrypting cross-chain information is necessary, requiring third-party verification of relations within the encrypted content. Existing privacy-preserving cross-chain research for consortium chains struggles to balance transaction efficiency, transaction validity verification, and complex trust assumptions for relays. We present VeCroToken, an efficient, verifiable, and privacy-preserving cross-chain model for consortium blockchains. VeCroToken introduces a dual-balance mechanism and designs four types of cross-chain zero-knowledge transactions based on zk-SNARKs. These transactions encrypt two types of balances and transaction amounts, effectively protecting participant privacy. The encrypted cross-chain data and zero-knowledge proof credentials are stored on participants’ consortium blockchains and the relay chain. Relay nodes and third parties can validate transaction proofs with public parameters, preserving privacy while ensuring compliance and validity. We give a security analysis in the UC framework that proves verifiability and balance safety. We also provide a privacy analysis establishing the amount, balance, and fund-correlation privacy. We implement a prototype on Hyperledger Fabric. Our experimental results show that VeCroToken has a lower overall zero-knowledge proof overhead than the state-of-the-art models and performs well in transaction performance.
Xiaoyan Hu 0007, Weicheng Zhou, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
IEEE Trans. Netw. Serv. Manag.2
2025 Encrypted Yet Leaking: Analyzing Side-Channel Vulnerabilities in Location Privacy of LBS
abstract
With the widespread use of smartphones and the rapid growth in mobile users worldwide, Location-Based Services (LBS) have become indispensable in daily life. The utilization of such services inevitably results in the generation of a considerable quantity of geolocation data. Most of these data have been encrypted to safeguard user privacy. However, the risk of side-channel information leakage still exists. In order to reveal the vulnerabilities in geolocation privacy, we propose a novel attack method, ETLA, leveraging encrypted LBS traffic analysis techniques. Specifically, we design a feature extraction algorithm, TPFC, to accurately restore the application-layer LBS transmission patterns and construct highly recognizable location combined features. Finally, we conduct extensive experiments based on real traffic datasets, and the results show that the average attack accuracy of ETLA reaches 97.68%. Additionally, we validate the application agnosticism, temporal stability and anti-interference capability of the model, further emphasizing the threat posed by the attack in real-world application scenarios.
Xuqiong Bian, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
ICC6
2025 Accurate and Early Detection of Iot Malware Via Dns Traffic Analysis with Deep Learning
abstract
Malware increasingly targets current Internet of Things (IoT) devices, causing significant economic losses. Accurate and early detection of malware is essential for defense. Existing IoT malware detection methods primarily analyze interactive traffic between compromised devices and C&C servers. Such detection needs to be performed while IoT devices are undergoing attacks, which still exposes IoT devices to danger. By analyzing real-world DNS traffic generated by IoT devices, we uncover that the DNS behavior patterns of benign IoT devices and malwareinfected devices differ. Therefore, this work proposes a method to accurately and early detect IoT malware via DNS traffic analysis with deep learning before attacks are launched, referred to as IoTMD-2D. IoTMD-2D first extracts a comprehensive set of DNS traffic features of IoT devices that effectively characterize DNS traffic behavioral patterns. Then, it integrates an attention-based LSTM to capture hidden relationships within domain names and a 1D-CNN to explore hidden patterns in DNS behavior-level features for generating feature representations that discriminate DNS traffic of benign IoT devices and malware-infected devices. Finally, IoTMD-2D accurately detects IoT malware based on the generated feature representation. Our experimental study on public IoT datasets demonstrates that our IoTMD-2D achieves an accuracy of 97.63 % in detecting IoT malware at an early stage via DNS traffic analysis.
Chenxing Zhang, Xiaoyan Hu 0007, Xuanlin Pan, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
ICC2
2025 BTG-RF: Recognizing Douyin payment behaviors based on behavioral traffic graph analysis
Xiaoyan Hu 0007, Xinghai Chen, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
Comput. Networks2
2025 IEA-DMS: An Interpretable feature-driven, Efficient and Accurate Detection Method for Slow HTTP DoS in high-speed networks
abstract
Slow HTTP DoS (SHD) is a novel DoS attack that exploits HTTP/HTTPS. SHD often operates at the application layer with encryption and has long packet intervals due to its slow transmission rate, making it more concealed and difficult to detect. Therefore, traditional detection methods for high-speed DDoS are ineffective against SHD. Meanwhile, Existing SHD detection approaches need many generic features or complex models, thus becoming less interpretable and more resource-intensive to meet real-time demands in high-speed networks. Moreover, most methods rely on bidirectional traffic, neglecting the prevalent issue of asymmetric routing in high-speed networks. To overcome these shortcomings, this paper proposes IEA-DMS, an Interpretable feature-driven, Efficient and Accurate Detection Method for Slow HTTP DoS in high-speed networks. We first analyze SHD mechanisms and construct a representative feature set based on its traffic characteristics to perform effectively under sampling and asymmetric routing. Then, to fast and accurately record the features, we employ Slow HTTP DoS Sketch and provide a detailed error analysis and suggest appropriate parameters. Experiments using public datasets show that the proposed features are efficient and interpretable. Even with numerous unidirectional flows and a 1/64 sampling rate , IEA-DMS detects SHD accurately within 2 min with low memory usage. Besides, IEA-DMS’s processing performance reaches 13.1 Mpps and can continuously process more than 100 days of traffic without clearing memory.
Hua Wu 0004, Suyue Wang, Guang Cheng 0001, Xiaoyan Hu 0007
Comput. Secur.6
2025 WEDoHTool: Word embedding based early identification of DoH tunnel tool traffic in dynamic network environments
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
Comput. Secur.2
2025 The secret behind instant messaging: video identification attack against complex protocols
abstract
Abstract People conveniently share and watch videos through Instant Messaging(IM) software, which is likely to reveal their preferences. Identifying IM video content can enable attackers to snoop on user privacy. Existing methods identify videos based on the features embodied in the DASH stream. However, IM software does not transmit video using DASH. IM software uses various transmission protocols or even private protocols for video transmission, which poses a challenge for video content identification. In this paper, we propose a video content identification framework for IM software, which obtains video content by extracting unique and stable features of videos as transmission fingerprints and matching them in a video fingerprint database. We evaluate the method on two popular IM software. The experimental results show that our method has an accuracy of 98.05% and 99.49% for ciphertext videos transmitted over GQUIC protocol and HTTPS protocol, respectively, and even reaches 100% identification accuracy for plaintext videos. Furthermore, the experimental results outperform the existing methods.
Ruiqi Huang, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
Cybersecur.5
2025 A Detection Scheme for Multiplexed Asymmetric Workload DDoS Attacks in High-Speed Networks
abstract
The asymmetric workload attack is an application layer attack that aims to exhaust the Central Processing Unit (CPU) resources of a server. Some attackers exploit new features of the Hypertext Transfer Protocol version 2 (HTTP/2) to launch Multiplexed Asymmetric Workload DDoS (MAWD) attacks using a small number of bots, which can cause denial of service on HTTP/2 servers. Data centers in high-speed networks host a large number of web applications. However, most of the detection methods for asymmetric workload attacks rely on request semantic analysis, which cannot be applied to encrypted MAWD attack traffic in high-speed networks. Besides, traditional rate-based DDoS detection methods are ineffective in detecting MAWD because the MAWD attacks use legitimate HTTP requests, and HTTP/2 traffic is bursty in nature. This paper proposes a practical scheme to detect MAWD attacks in high-speed networks. We construct an effective feature set based on the characteristics of MAWD attacks in high-speed networks and design MAWD-HashTable (MAWD-HT) to extract features quickly. Experimental results on real traffic traces with speeds reaching Gbps demonstrate that our scheme can detect MAWD attacks within 3 seconds, with a recall rate of more than 99%, a FPR of less than 0.1%, and an acceptable resource consumption.
Fuhao Yang, Hua Wu 0004, Xiaoyan Hu 0007, Jing Ren 0002
IEEE Trans. Netw. Serv. Manag.5
2024 Unveiling the Unseen: Video Recognition Attacks on Social Software
Hangyu Zhao, Hua Wu 0004, Xuqiong Bian, Guang Cheng 0001, Xiaoyan Hu 0007, Zhiyi Tian
ACISP (2)6
2024 Enhancing Unknown Encrypted Traffic Clustering with Self-Supervised Learning
abstract
Many malicious attacks are launched through encrypted traffic from unknown proprietary network protocols. Timely identification of such malicious unknown encrypted traffic is essential for the defense. However, it is challenging to acquire labels for unknown protocols in the context of encrypted traffic. Due to the lack of prior knowledge, unsupervised learning is adopted to cluster unknown encrypted traffic. The existing unsupervised encrypted traffic clustering methods do not customize feature extraction and representation of unknown encrypted traffic, resulting in imperfect clustering results. This work innovatively proposes BiFR-SSL to accurately cluster unknown encrypted traffic without prior knowledge. BiFRSSL extracts features of each unknown encrypted bidirectional network flow based on the lengths, arrival time, and directions of packets within the flow to construct a Bidirectional Flowpic Representation (BiFR). Subsequently, it exploits Self-Supervised Learning (SSL) to pre-train a feature extractor that produces feature representations from BiFRs, guaranteeing the closeness of encrypted traffic flows from the same protocol in the representation space. Finally, it clusters unknown encrypted traffic based on their feature representations generated by the pre-trained feature extractor. Our experimental studies demonstrate that BiFR-SSL can effectively cluster encrypted traffic of unknown protocols and outperforms state-of-the-art methods.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
GLOBECOM3
2024 Efficient Short Video Identification Attack for Scenarios with Hybrid Transmission Modes and Preloading Mechanism
abstract
To protect user privacy, video traffic is usually encrypted during transmission. Some research has been con-ducted to implement video identification attacks by analyzing the features of video traffic. However, short video platforms use hybrid transmission modes and a preloading mechanism to improve user experience. These new characteristics make video identification attacks targeted at long videos not appliable to short videos. In this paper, we proposed a method for extracting and correcting the hybrid fingerprints of application layer video resources, and designed an SVP-DOHMM matching method for video identification. Our method can identify each video for scenarios with hybrid transmission modes and preloading mechanism. We implemented our method on the dataset con-taining more than 100,000 video fingerprints of a short video platform. The experimental results show that the accuracy can reach 98.16% by sniffing the traffic for only 10 seconds in the closed world and can identify videos with 97.64 % accuracy in the open world.
Jingwen Quan, Jiajia Du, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
MSN5
2024 Identifying Video Resolution from Encrypted QUIC Streams in Segment-combined Transmission Scenarios
abstract
With the rapid rise of video services, Internet Service Providers (ISPs) need to better monitor the Quality of Experience (QoE). Video resolution is a crucial factor affecting QoE. However, with the widespread use of QUIC based on UDP for video transmission, existing resolution identification methods based on the TCP header information cannot extract information from the UDP header. Moreover, in recent years, video platforms have begun to send video segments using random combinations to avoid side-channel attacks, leading to the failure of existing machine learning-based methods. To address this problem, we propose a method to identify the video resolution from QUIC traffic. The method takes the length of the video segment sequence as fingerprints, uses the features of QUIC to accurately extract and correct the length of the video segment sequence from the encrypted video stream, and then uses a combinatorial matching method to identify the corresponding video segment, thus accurately identifying the resolution of the video segment. Experimental results using YouTube videos show that the accuracy of this method for video resolution identification is more than 97%, and the average identification time is 0.13 seconds. Using this method, ISPs can accurately identify the resolution of videos transmitted via QUIC in real-time, which provides a basis for monitoring users' QoE.
Yuanjie Zhao, Hua Wu 0004, Liujinhan Chen, Guang Cheng 0001, Xiaoyan Hu 0007
NOSSDAV6
2024 AHDom: Algorithmically generated domain detection using attribute heterogeneous graph neural network
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
Comput. Networks1
2024 SD-MDN-TM: A traceback and mitigation integrated mechanism against DDoS attacks with IP spoofing
Suyue Wang, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007, Jing Ren 0002
Comput. Networks4
2024 RT-CBCH: Real-Time VPN Traffic Service Identification Based on Sampled Data in High-Speed Networks
abstract
Virtual Private Network (VPN) technology can bypass censorship and access geographically locked services. Some harmful information may be hidden in VPN traffic and circumvent the surveillance systems, bringing a significant challenge to network security. Considering the increasing richness of service types in VPN traffic, identifying traffic service facilitates further targeting harmful VPN traffic. Therefore, VPN traffic service identification is critical in network management. The existing identification methods use complete traffic for analysis. However, massive data analysis in high-speed networks consumes enormous resources, limiting the real-time processing of traffic identification. This paper proposes a real-time VPN traffic service identification method named RT-CBCH. We construct features that are still available after sampling and design a fast traffic processing structure based on Counting Bloom Filter and Chained Hash Table (CBCH). Experimental results validate the real-time capability, stability and accuracy of our method. At the sampling ratio of 1/256, it takes only 23.63 seconds to process the mixed traffic of 900-second traffic generated on a 10 Gbps link and our collected V2Ray traffic, which is increasingly common in VPN traffic. Under different sampling ratios, the identification results remain respectable, with an overall accuracy of about 90% for application service and over 99% for V2Ray proxy service. Furthermore, comparisons with similar work illustrate the high accuracy and low resource consumption of RT-CBCH. Experimental results show that our method can stably implement real-time VPN traffic service identification from sampled data in high-speed networks.
Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
IEEE Trans. Netw. Serv. Manag.4
2023 Accurate Identification of Encrypted Videos in Asymmetric Routing Scenarios
Hua Wu 0004, Jingwen Quan, Guang Cheng 0001, Xiaoyan Hu 0007
APNOMS5
2023 Towards Early and Accurate IoT Device-Type Identification with Global Attention Mechanism
abstract
With the rapid development of Internet of Things (loT) technology, there is explosive growth in the number of loT devices. Meanwhile, the low security and network heterogeneity of loT networks have brought new challenges to implementing network management and security strategies in smart homes and small offices. Early and accurate loT device-type identification is the first step towards the security management of loT networks. The existing machine learning-based and deep learning-based models for loT traffic classification have achieved decent results. However, most of these methods rely on a long-term window to collect loT device traffic for identification, resulting in limited real-time performance. This work proposes 10T-GFCN, an early and accurate loT device-type identification model with global attention mechanism. 10T-GFCN first constructs a multi-feature sequence for each device from a small packet window. Then 10T-GFCN resorts to the global attention mechanism to efficiently mine temporal information and feature relationships and obtain an updated embedding of each multi-feature sequence. Finally, a fully convolutional neural network is trained based on the updated embeddings of traffic features to identify loT device types. Our experimental study suggests that 10T-GFCN can efficiently capture distinguishable representations for packet-level features of loT traffic and outperforms state-of-the-art loT identification methods. It achieves an average accuracy of 98.88 % with a window size of 75 packets (the traffic of about three minutes) on the UNSW dataset.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
GLOBECOM1
2023 Website Fingerprinting with Packet Sampling: A More Realistic Approach in Real-World Networks
abstract
Website fingerprinting (WF) attack enables an eavesdropper to spy on users' browsing activity for malicious purposes, which poses a critical threat to Internet users' privacy. Prior research mainly focuses on attack performance in local area networks. Meanwhile, with the upgrade of network infrastruc-ture, a notable transition to high-speed connectivity in real-world network nodes has emerged. While in high-speed networks, there has been a mismatch between the overall traffic in transmission and the upper limit of the attacker's processing capabilities. Prior attacks based on full-traffic collection will face a sharp increase in resource overhead and a decrease in attack efficiency. In response to the emerging challenges, we first apply sampling techniques to WF attack to reduce the amount of data that needs to be processed. In addition, we devise an effective attack model with good performance on sampled traffic. Evaluations indicate that our attack model achieves 94.9 % accuracy in 1/8 packet sampling scenario and 98% accuracy in non-sampling scenario, outperforming the state-of-the-art in both cases. The compatibility in packet sampling environments helps extend the WF attack from the laboratory setting targeting at a few users to real-world high-speed networks capable of massive surveillance. The code of this paper is publicly available at https://github.com/code-flyerISAPWF.
Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
GLOBECOM4
2023 NFlowGAN: High-Utility Privacy-Preserving Network Flow Synthesis Based on GAN
abstract
The sensitivity of network traffic data has led to the scarcity of public traffic datasets, hindering the development of data-driven research in this field. Researchers proposed publishing synthetic network traffic instead of the original dataset. However, existing traffic synthesis methods are inadequate in data utility and seldom consider privacy protection. For this reason, we propose NFlowGAN for high-utility privacy-preserving network flow synthesis. We introduce spectral normalization in the network structure to improve training stability, thus improving the data utility. In addition, we add a Gaussian noise layer to the discriminator of NFlowGAN to provide higher privacy guarantees for the synthesized flow. The experimental evaluation results on the Darknet2020 dataset demonstrate that our proposed NFlowGAN achieves a significant improvement in data utility with privacy preservation compared to the two baselines. The synthesized high-utility dataset can be widely shared for research and educational purposes.
Zhaoxu Ge, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
ICC4
2023 Detecting Cryptomining Traffic Over an Encrypted Proxy Based on K-S Test
abstract
In recent years, the good revenue generated by cryptocurrency mining has attracted a lot of people to participate in it. It has also caught the attention of hackers, and cryptojacking attacks are becoming more common. Detecting cryptomining behavior can effectively reduce the lost caused by cryptojacking attacks. Existing host-based cryptomining detection methods can protect only end devices and violate users' privacy. Besides, network-based solutions can not better handle anti-reconnaissance means of encrypted proxy. To bridge this gap, we propose a cryptomining traffic detection model based on K-S Test(CMD-KST). Our traffic analysis study confirms that the feature distributions of cryptomining traffic over an encrypted proxy are still stable and unique. CMD-KST compares the feature distributions of a network flow segment with that of cryptomining traffic over the encrypted proxy to complete the detection task. CMD-KST is easily deployable and can detect cryptomining traffic at the entrance of the managed network. Our experimental results demonstrate that CMD-KST achieves a recall of 98.84% without generating false positives and takes only 6 minutes of analyzing mining traffic to complete the detection. CMD-KST is faster than other network-based cryptomining traffic detection methods and achieves a higher precision. Furthermore, the adversarial evaluation shows that it is challenging for the attackers to counteract our detection.
Xiaoyan Hu 0007, Boquan Lin, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
ICC1
2023 A Novel Darknet Traffic Classification Method Based on Knowledge Graph with Dynamic Embedding Learning
abstract
Darknet is described as an individual encrypted part of the Internet that can only be accessed with specific anonymity tools. Achieving accurate classification of darknet traffic is crucial for identifying anonymous network applications and combating cybercrimes. Machine learning-based and deep learning-based classifiers have achieved decent results in darknet traffic classification. However, these methods can not learn global and distinctive darknet flow embedding representations, resulting in limited classification performance. To tackle these issues, we propose Dark-DKGC, a novel darknet traffic classification method based on Knowledge Graph (KG) with Dynamic Knowledge Graph (DKG) embedding learning. Dark-DKGC first constructs Darknet Traffic Dynamic Knowledge Graph (Dark-DKG). Then Dark-DKGC utilizes the DKG embedding method to effectively learn the embedding representations of all flows. Finally, machine learning-based classifiers are trained based on the embedding representations of flows to identify darknet traffic. Our experimental studies suggest that Dark-DKGC can effectively capture distinguishable embedding representations for darknet flows. In multiclass classification scenario, its average accuracy is about 7%-13% higher than state-of-the-art methods and 1% higher than the static KG embedding-based classifier. Besides, compared to the static KG embedding method, Dark-DKGC takes advantage of its online embedding learning to improve test efficiency significantly. Moreover, the visualization of Dark-DKG allows a certain degree of interpretability for the classification results.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
ICC1
2023 Real-Time Phishing Detection Based on URL Multi-Perspective Features: Aiming at the Real Web Environment
abstract
Phishing deceives users' trust through subtle URL and HTML disguises, stealing sensitive data or spreading malicious viruses. Phishing detection from URLs has been the focus of research in recent years, which balances the performance and time compared to list-based and content-based approaches. The approaches using neural networks to extract semantic information from URLs to detect phishing websites can avoid feature engineering. However, the features' plausibility cannot be verified. Heuristic features designed artificially can reflect URL differences more reasonably, but the current features lack diversity and have poor generalization in the real web environment. In this paper, we propose a phishing detection model combining heuristic features and machine learning, which extracts features from URL components and linguistics perspectives, leading to lightweight and feature diversity. Three datasets with significant differences are used to verify the model's generalizability. Eventually, the average accuracy of the three datasets reaches 98.68%, and the average precision, recall, and F1-score are all above 98%, which shows good generalizability and outperforms the baselines.
Shiyue Liu, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
ICC4
2023 An Accurate and Real-Time Detection Method for Concealed Slow HTTP DoS in Backbone Network
Hua Wu 0004, Suyue Wang, Guang Cheng 0001, Xiaoyan Hu 0007
SEC5
2023 Real-Time Platform Identification of VPN Video Streaming Based on Side-Channel Attack
Anting Lu, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
SEC5
2023 A Stable Fine-Grained Webpage Fingerprinting: Aiming at the Unstable Realistic Network
Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
SecureComm (2)5
2023 Fine-grained Ethereum behavior identification via encrypted traffic analysis with serialized backward inference
Xiaoyan Hu 0007, Zhuozhuo Shu, Zhongqi Tong, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
Comput. Networks1
2023 PD-CPS: A practical scheme for detecting covert port scans in high-speed networks
Hua Wu 0004, Ziling Shao 0002, Fuhao Yang, Guang Cheng 0001, Xiaoyan Hu 0007, Jing Ren 0002, Wei Wang 0171
Comput. Networks5
2023 Batch classifier with adaptive update for backbone traffic classification
Hua Wu 0004, Weina Li, Xiying Chen, Guang Cheng 0001, Xiaoyan Hu 0007, Youqiong Zhuang
Comput. Commun.5
2023 A Deep Subdomain Adaptation Network With Attention Mechanism for Malware Variant Traffic Identification at an IoT Edge Gateway
abstract
The prevailing of malware variants in ubiquitous Internet of Things (IoT) devices causes enormous losses. Accurate and timely identification of malware variant traffic at an IoT edge gateway can effectively reduce the loss. TransNet, the state-of-the-art technology for malware variant traffic detection, considers only global domain adaptation and ignores the alignment of distributions between different subdomains, which fails to capture the fine-grained information of classification targets. Besides, TransNet converges very slowly, which may use up precious resources in IoT devices. This article proposes a deep subdomain adaptation network with attention mechanism (DSAN-AT) to accurately and efficiently identify malware variant traffic at an IoT edge gateway. DSAN-AT utilizes local maximum mean discrepancy (LMMD) to align the traffic feature distributions of subdomains in the source and target domains. It also exploits channel and spatial attention mechanisms to accelerate learning traffic features between different subdomains to save precious computing resources at the IoT edge gateway. Our experimental study demonstrates that DSAN-AT achieves an average accuracy of 97.15% (96.37% for TransNet) and converges fast without using a large target domain training data set. DSAN-AT has strong practicality for identifying malware variant traffic at an edge IoT gateway.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
IEEE Internet Things J.1
2023 Towards verifiable and privacy-preserving account model on a consortium blockchain based on zk-SNARKs
Xiaoyan Hu 0007, Weicheng Zhou, Guang Cheng 0001, Shen Yan 0005, Hua Wu 0004
Peer Peer Netw. Appl.1
2023 ReplaceDGA: BiLSTM-Based Adversarial DGA With High Anti-Detection Ability
abstract
Botnets extensively leverage Domain Generation Algorithms (DGAs) to establish reliable communication channels between bots and Command and Control (C&C) servers. Numerous character-level DGA classifiers have been extensively studied to detect and classify domain names generated by DGAs. Meanwhile, a series of adversarial domain generation algorithms have been proposed to evade DGA classifiers. Although the existing domain name generation algorithms have progressed against DGA classifier, their anti-detection abilities are still weak. This paper proposes a Bidirectional Long Short-Term Memory (BiLSTM) network-based adversarial DGA with high anti-detection ability, referred to as ReplaceDGA. ReplaceDGA requires no knowledge of the targeted DGA classifiers. It first builds a prediction model for benign domain names using the BiLSTM network to model the semantic relationship hidden within benign domain names and then replaces two characters of each input benign domain name based on the prediction model to maximize the similarity between the benign and generated domain names. Our experimental results validate that ReplaceDGA successfully evades various character-level DGA classifiers even after they are retrained by domain names generated by ReplaceDGA and outperforms the state-of-the-art adversarial DGAs in anti-detection ability, repetition rate, and collision rate. Our study of ReplaceDGA promotes the urgent need for developing more comprehensive and robust DGA classifiers that consider other factors besides character-level information of domain names.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004, Yali Yuan
IEEE Trans. Inf. Forensics Secur.1
2023 Toward Early and Accurate Network Intrusion Detection Using Graph Embedding
abstract
Early and accurate detection of network intrusions is crucial to ensure network security and stability. Existing network intrusion detection methods mainly use conventional machine learning or deep learning technology to classify intrusions based on the statistical features of network flows. The feature extraction relies on expert experience and cannot be performed until the end of network flows, which delays intrusion detection. The existing graph-based intrusion detection methods require global network traffic to construct communication graphs, which is complex and time-consuming. Besides, the existing deep learning-based and graph-based intrusion detection methods resort to massive training samples. This paper proposes Graph2vec+RF, an early and accurate network intrusion detection method based on graph embedding technology. We construct a flow graph from the initial several interactive packets for each bidirectional network flow instead, adopt graph embedding technology, graph2vec, to learn the vector representation of the flow graph and classify the graph vectors with Random Forest (RF). Graph2vec+RF automatically extracts flow graph features using subgraph structures and relies on only a small number of the initial interactive packets per bidirectional network flow without requiring massive training samples to achieve early and accurate network intrusion detection. Our experimental results on the CICIDS2017 and CICIDS2018 datasets show that our proposed Graph2vec+RF outperforms the state-of-the-art methods in terms of accuracy, recall, precision, and F1-score.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
IEEE Trans. Inf. Forensics Secur.1
2023 AF-FDS: An Accurate, Fast, and Fine-Grained Detection Scheme for DDoS Attacks in High-Speed Networks With Asymmetric Routing
abstract
Distributed Denial of Service (DDoS) attacks have posed severe threats to the Internet. Although researchers have proposed many DDoS detection schemes, there are still some challenging issues. Traditional per-flow-based DDoS methods are impractical for massive amounts of high-speed network traffic due to the huge resource consumption. In addition, existing methods are not designed to take into account the widespread asymmetric routing in high-speed networks, resulting in false positives when these methods are deployed on the Internet. Furthermore, existing methods can not achieve a good trade-off between detection accuracy and granularity when detecting hybrid DDoS attacks. This paper proposes an Accurate, Fast, and Fine-grained Detection Scheme (AF-FDS) for DDoS attacks in high-speed networks with asymmetric routing. We select features based on the characteristics of DDoS attacks and design a data structure Double Composite Structure Sketch (DCSS). DCSS can achieve fast recording and extraction of the selected features from the sampled traffic. Experimental results using real-world traces in a 10Gbps network with asymmetric routing show that AF-FDS can detect nine types of DDoS attacks at a fine-grained level within 15 seconds with over 98.0% precision and recall, even at a sampling rate of 1/1024. Furthermore, the comparison with several state-of-the-art methods illustrates that AF-FDS can detect DDoS attacks with a lower false positive rate (FPR) and shorter alarm time in asymmetric routing scenarios.
Ziling Shao 0002, Tingzheng Chen, Guang Cheng 0001, Xiaoyan Hu 0007, Weina Li, Hua Wu 0004
IEEE Trans. Netw. Serv. Manag.4
2023 LossDetection: Real-Time Packet Loss Monitoring System for Sampled Traffic Data
abstract
Packet loss is common in networks, which leads to network quality of service degradation. Packet loss is an essential and concerning symptom when the quality of service is degraded. Therefore, real-time passive packet loss detection is conducive to estimating network services. Existing passive packet loss detection methods mainly study the packet loss for TCP using header information from full traffic. However, it cannot infer packet loss status for UDP due to its limited header information and is too costly to perform full acquisition in real networks. To address these problems, we propose a framework called LossDetection based on packet sampling and Feature-Sketch to detect packet loss in real time for both TCP and UDP. The result shows that our methodology can detect packet loss with an accuracy of 98%-100% at a sampling rate of 1/16. Furthermore, our extensive evaluation demonstrates that LossDetection is easy to implement in a software router and achieves low memory and detection latency while providing real-time information about packet loss.
Hua Wu 0004, Shanshan Ni, Guang Cheng 0001, Xiaoyan Hu 0007
IEEE Trans. Netw. Serv. Manag.5
2023 Resolution Identification of Encrypted Video Streaming Based on HTTP/2 Features
abstract
With the inevitable dominance of video traffic on the Internet, Internet service providers (ISP) are striving to deliver video streaming with high quality. Video resolution, as a direct reflection of video quality, is a key factor of the video quality of experience (QoE). Since the displayed information of video cannot be observed by ISPs, ISPs can only measure the video resolution from traffic. However, with HTTP/2 being gradually adopted in video services, the multiplexing feature of HTTP/2 allows audio and video chunks to be mixed during transmission, making existing monitoring approaches unusable. In this article, we propose a method called H2CI to monitor resolution for adaptive encrypted video traffic under HTTP/2. We consider the size of the mixed data for identification. Specifically, H2CI consists of a length restoration method to extract restored fingerprints and a fingerprint-matching method for fine-grained resolution identification. The experimental results show that H2CI can achieve more than 98% accuracy for fine-grained resolution identification. Our method can be effectively applied to infer the adaptation behavior of encrypted video streaming and monitor the QoE of video services under HTTP/2.
Hua Wu 0004, Xin Li 0194, Guang Cheng 0001, Xiaoyan Hu 0007
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Service classification of high-speed network traffic based on Two-Stage Clustering
abstract
Service classification of high-speed network traffic is critical for Internet Service Providers (ISPs) to ensure network Quality of Service (QoS). As high-speed network transmission accelerates, ISPs can only obtain unlabeled and sampled traffic from high-speed networks, making supervised learning methods difficult to apply. Some existing methods use unsupervised learning to classify services to reduce the need for labeled data. However, when these methods are applied, fluctuations in the feature vector lead to a certain percentage of the same class of services being grouped into different clusters. We proposes a practical method for classifying traffic services in high-speed networks. Specifically, we propose a method called Two-Stage Clustering (TSC), which automatically implements merging clusters of the same service. Validation experiments on publicly available datasets show that our classifier achieves an accuracy of 90.07% and a recall of 91.81% even with a sampling rate of 1:64, which is higher than the classification methods that also use unsupervised learning.
Hua Wu 0004, Yuping Sui, Guang Cheng 0001, Xiaoyan Hu 0007, Qinghua Shang
APNOMS4
2022 HDS: A Hierarchical Scheme for Accurate and Efficient DDoS Flooding Attack Detection
abstract
As the scale of Distributed Denial of Service (DDoS) flooding attacks has increased significantly, many detection methods have applied sketch data structures to compress the IP traffic for storage saving. However, due to the large IP address space, these methods need to flush the sketch frequently to reduce the hash collisions. Besides, few of them can be applied to detect attacks in the high-speed network where sampling is usually adopted. This paper proposes a hierarchical system named HDS for efficient and continuous DDoS flooding attack detection in high-speed networks. Rather than directly processing the IP traffic, HDS uses sketches to track sampled traffic at different levels of aggregation: interface level, area level, and host level. Then traffic classifiers are trained for each level for attack detection. The main advantage of our approach is that each detection level only tracks a small set of traffic, which can identify the attack victim fastly and hardly causes hash collisions. Experimental results on the real-world 10Gbps network traffic datasets show that HDS can effectively detect various DDoS flooding attacks with high accuracy and identify the victim within an average of 10s when the sampling rate exceeds 1/2048.
Youqiong Zhuang, Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
APNOMS5
2022 An Adversarial Learning-based Tor Malware Traffic Detection Model
abstract
Attackers often use Tor to launch cyberattacks and conduct illegal transactions, threatening cyberspace's security and people's daily lives. Existing methods for malware traffic detection on Tor can be classified as rule-based and network-based, both of which apply machine learning extensively. Tor malware traffic detection systems are often deployed in open network environments. Their machine learning systems are the first to be attacked by adversarial samples. To ensure that Tor is not abused, this paper proposes an Adversarial Learning-based Tor Malware Traffic Detection model, AL-TMTD. We generate realistic attack samples that can evade detection and use these samples to produce an augmented training set for producing hardened detectors. In such a way, we obtain a more resilient Tor malware traffic detection model that achieves adversarial robustness. We validate our proposal through an extensive experimental campaign that considers multiple machine learning algorithms and shadow models. We simulate the adversary to construct functionally approximate shadow models through black-box model extraction and generate adversarial samples to validate the adversarial robustness of our proposed AL-TMTD model. Our experimental results demonstrate that the average accuracy of AL-TMTD after the adversarial retraining is as high as 0.995 in detecting adversarial samples, which is 0.314 without the adversarial retraining, a significant improvement.
Xiaoyan Hu 0007, Yishu Gao, Guang Cheng 0001, Hua Wu 0004, Ruidong Li 0001
GLOBECOM1
2022 A Dynamic Access Control Model Based on Attributes and Intro VAE
abstract
Affected by the COVID-19 pandemic, teleworking is becoming more popular, with the exposed attack surface of the internal network expanding. Once outsiders personate accounts or insiders conduct illegal operations, the data security in teleworking with traditional border protection will be broken. Therefore, it is necessary to implement fine-grained and dynamic access control to protect data from malicious access. Attribute-based access control (ABAC) is ideal, where authorization is performed through attributes and rules. On this basis, risk assessment, context awareness, and machine learning are supplemented for dynamic access control. However, these methods have their limitations due to the requirement of sufficient prior knowledge and massive label-classified data. Moreover, it is challenging to obtain the samples of attack behaviors, and the attack behaviors may change frequently to evade detection. In contrast, the normal behaviors are relatively stable except for the update of network services. We propose a dynamic access control model, ABAC-IntroVAE, to address the above issues. ABAC-IntroVAE judges users' requests through rule matching and behavior analysis based on the attributes of the requests. It first filters out requests against the rules by rule matching. Then, the introspective variational autoencoder (IntroVAE) is used for behavior analysis to realize dynamic access decisions. Requests classified as normal can be authorized for access. ABAC-IntroVAE only needs samples of normal requests for training, avoiding the difficult task of collecting massive and frequently changing samples of attack requests. Meanwhile, the IntroVAE model is updated through continual learning to adapt to new-style normal behaviors due to the update of network services. Our experiment study suggests that our proposed ABAC-IntroVAE can effectively perform dynamic access control. It achieves an accuracy of 97.2% in abnormal detection and maintains an accuracy of over 97% through continual learning, despite the addition of new-style user behavior patterns.
Xiaoyan Hu 0007, Yuelin Hu, Guang Cheng 0001, Hua Wu 0004, Yifei Qin
GLOBECOM1
2022 Detecting Slow Port Scans of Long Duration in High-Speed Networks
abstract
Port scanning is an extensively used technique by attackers to probe for vulnerabilities in network systems. Since fast port scans can be effectively detected by many existing methods, some advanced attackers perform slow port scans in order not to be suspected. A highly stealthy slow scan can last for dozens of days, which brings significant challenges to current intrusion detection approaches. Besides, the existing port scan detection methods are all based on full traffic. They are not suitable for high-speed networks because of huge computational and storage resource consumption. According to the protocol characteristics and the connection patterns of port scans, we construct a traffic feature set that can not only distinguish the specific scan types, but also remain effective for the sampled traffic. Furthermore, we customize a data structure Scan Detection Sketch (SDS) for feature extraction. Experimental results using public datasets show that our method can detect slow port scans in a 10Gbps high-speed network with high accuracy and acceptable memory consumption. And the proposed method works well even for slow port scans lasting more than 60 days.
Hua Wu 0004, Ziling Shao 0002, Guang Cheng 0001, Xiaoyan Hu 0007, Jing Ren 0002, Wei Wang 0171
GLOBECOM4
2022 Real-time Identification of VPN Traffic based on Counting Bloom Filter and Chained Hash Table from Sampled Data in High-speed Networks
abstract
Virtual Private Network (VPN) can bypass censorship and access services that are geographically locked. Therefore, VPN traffic identification has become an urgent problem in traffic classification. The existing VPN traffic identification methods use complete traffic for analysis. However, massive data analysis in high-speed networks consumes many resources, limiting the real-time processing of traffic identification. The management of high-speed networks is mainly based on sampled traffic. As VPN traffic accounts for a relatively low proportion, it is particularly challenging to identify VPN traffic from sampled data. This paper proposes a real-time identification method for VPN traffic from sampled data in high-speed networks. In our method, we construct features that are still available after sampling and design a fast traffic processing structure based on Counting Bloom Filter and Chained Hash Table (CBCH). To validate the usability of our method, we use 900 seconds of traffic generated on a 10 Gbps link as background traffic, mixed with V2Ray traffic, which is increasingly common in VPN traffic. With the VPN traffic proportion of 0.03%, it takes only 50.79 seconds to complete the processing at the sampling ratio of 1/256. This time is significantly less than the traffic generation time. For the effective flows extracted from sampled backbone traffic, the identification results are maintained at a high level with 97% precision, 93% recall, and 95% F1 score. In addition, our method can achieve fine-grained VPN traffic identification of different V2Ray tools.
Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
ICC4
2022 IM-Shield: A Novel Defense System against DDoS Attacks under IP Spoofing in High-speed Networks
abstract
DDoS attacks are usually accompanied by IP spoofing, but the availability of existing DDoS defense systems for high-speed networks decreases when facing DDoS attacks with IP spoofing. Although IP traceback technologies are proposed to focus on IP spoofing in DDoS attacks, there are problems in practical application such as the need to change existing protocols and extensive infrastructure support. To defend against DDoS attacks under IP spoofing in high-speed networks, we propose a novel DDoS defense system, IM-Shield. IM-Shield uses the address pair consisting of the upper router interface MAC address and the destination IP address for DDoS attack detection. IM-Shield implements fine-grained defense against DDoS attacks under IP spoofing by filtering the address pairs of attack traffic without requiring protocol and infrastructure extensions to be applied on the Internet. Detection experiments using the public dataset show that in a 10Gbps high-speed network, the detection precision of IM-Shield for DDoS attacks under IP spoofing is higher than 99.9%; and defense experiments simulating real-time processing in a 10Gbps high-speed network show that IM-Shield can effectively defend against DDoS attacks under IP spoofing.
Hua Wu 0004, Xuange Zhang, Tingzheng Chen, Guang Cheng 0001, Xiaoyan Hu 0007
ICC5
2022 Towards Accurate DGA Detection based on Siamese Network with Insufficient Training Samples
abstract
Domain Generation Algorithms (DGAs) are widely applied in diversified malicious attack patterns such as botnets. Attacks utilize DGAs to dynamically create pseudorandom domains to evade security detection and successfully connect bots with Command and Controls (C&C) servers. The detection of Algorithmically Generated Domains (AGDs) plays an essential role in network attack detection. Most of the existing DGA detectors are machine learning or deep learning-based methods. However, these DGA detectors perform relatively poorly with insufficient training samples, such as small-scale DGA families and emerging DGA variants. Besides, machine learning-based detectors require sophisticated and time-consuming artificial feature extraction, and attackers can circumvent the extracted features. This paper focuses on accurately detecting DGAs based on siamese network with insufficient training samples. Our proposed DGA detection method is referred to as DGAD-SN. DGAD-SN first introduces contrastive learning and adopts the siamese network framework to construct the feature extractor, which excavates the implicit relationship information between characters in the domain name strings using limited training samples. Then machine learning-based DGA classifiers are trained based on the extracted neural feature vectors of domain names to identify AGDs. Our experimental studies suggest that DGAD-SN can efficiently extract distinguishable neural feature vectors for domain names and outperforms state-of-the-art DGA detectors in identifying small-scale DGA families or emerging DGA variants. Its average accuracy is 10%−15% higher than conventional machine learning-based detection methods and about 1%−2% higher than deep learning-based detection methods using limited training samples.
Xiaoyan Hu 0007, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
ICC1
2022 Real-time Application Identification of RTC Media Streams via Encrypted Traffic Analysis
abstract
The globalization of the economy and the increase in network bandwidth have contributed significantly to the development and popularity of real-time communication (RTC) social applications. RTC media streams, such as video meetings and calls, require more network resources and real-time performance than other services. In order to meet the requirements of RTC application providers to offer a higher level of service to their subscribers, Internet Service Providers (ISPs) need to identify the application to which the RTC media stream belongs. There are already some studies on traffic identification. However, the extant work is not yet able to distinguish the corresponding applications from the same type of media streams in real time. In addition, most of the work is not validated with actual data containing massive background traffic. Hence, we propose a real-time application identification method for meeting and calling RTC media streams in social networks. By analyzing the encrypted traffic, the method extracts features from the unit-time traffic aggregation without using payload and related information fields. The generated feature sequences are fed to our lightweight model. Our proposed method does not depend on initial packets or whole flows, and only an arbitrary 3-second traffic block is needed to achieve over 99% accuracy. Moreover, experiments using high-speed network traffic reflect that our approach can identify corresponding applications from RTC media streams in real time. Besides, comparisons with similar work show that this method requires only 1/160th of the memory and 1/10th of the processing time.
Hua Wu 0004, Cheng-Fei Zhu, Guang Cheng 0001, Xiaoyan Hu 0007
ICCCN4
2022 PSCM: Towards Practical Encrypted Unknown Protocol Classification
abstract
Network traffic classification is the basis for network management, Quality of Service and intrusion detection. As the number of Internet applications increases, the variety of unknown protocols grows, posing a significant challenge to network traffic classification. Traditional rule-based traffic classification methods are currently limited by the rise of dynamic ports and encryption protocols. Statistical methods using statistical features have good recognition of protocols with public formats. However, there is no public protocol format for unknown protocols, making it challenging to extract useful features. This paper proposes a practical Probability Statistics and Cluster Merging (PSCM) method to automatically extract encrypted unknown protocol features and map the clustering results to the actual protocols. Experimental results on real-world network traffic show that the method achieves an accuracy of 99.28% and performs well in the sampling scenarios.
Hua Wu 0004, Chaoqun Cui, Guang Cheng 0001, Xiaoyan Hu 0007
ISCC4
2022 Identify IoT Devices from Backbone Networks Using Lightweight Neural Networks
abstract
Due to the heterogeneity, fragmentation, and lack of visibility, Internet of Things has become the new target for attacks. Therefore, it is necessary for Internet Service Providers to identify IoT devices to prevent attacks and protect the entire network in time. In this paper, we propose an IoT device identification approach based on lightweight deep learning models using a single feature. Specifically, we analyze the traffic pattern specific to IoT devices and use one feature to characterize this pattern, reducing the time consumption. Moreover, we select multiple time scales to extract this feature for different IoT devices, achieving an accurate characterization and improving the accuracy. Furthermore, we use unidirectional flows as analysis objects, suitable for backbone networks. The evaluation results on real-world datasets show that our approach achieves an accuracy of over 99%, with one-seventeenth of the time consumption of the state-of-the-art approach, realizing the lightweight and real-time requirements.
Hua Wu 0004, Xingmeng Fan, Guang Cheng 0001, Xiaoyan Hu 0007
LCN4
2022 Service-Based Identification of Highly Coupled Mobile Applications
abstract
Identifying mobile applications from network traffic is important for Internet service providers (ISPs) to manage their networks at a fine-grained level. However, the rise of public services has led to a gradual increase in service coupling among applications, making it more difficult to identify applications. Existing methods produce classification ambiguities when identifying highly coupled mobile applications, resulting in low application identification accuracy. In this paper, we propose a service-based method to quickly identify highly service coupling applications after the applications are launched. It can accurately identify highly coupled mobile applications based on the features of the services accessed by the applications. Experiments on a real network traffic dataset of highly coupled mobile applications verify that our method can identify applications within 25s after the mobile applications are launched, and the identification accuracy is over 99%.
Hua Wu 0004, Guang Cheng 0001, Xiaoyan Hu 0007
LCN4
2022 Verifying Privacy-Preserving Financing Orders on a Consortium Blockchain Based on zk-SNARKs
abstract
Due to its efficiency, low overhead, and high scalability, consortium blockchain has been deeply applied in various fields of society. Order financing is one of the scenarios of applying consortium blockchain. Since data on the consortium blockchain is available to the blockchain members, information of a financing order written directly to the blockchain will leak the commercial privacy of the purchaser and supplier. Therefore, the financing order data should be encrypted when published as a transaction on the consortium blockchain. However, the investor needs to verify the financing order data on a consortium blockchain before loaning money to the supplier. It is tricky to efficiently satisfy the verifiability of encrypted financing order data on the consortium blockchain. This work proposes VmppOrder, a verifiable model for privacy-preserving financing orders on a consortium blockchain based on zero-knowledge Succinct Non-interactive ARguments of Knowledge (zk-SNARKs). By the supplier publishing zero-knowledge proofs generated from the financing order, the investor can verify the encrypted financing order published on the consortium blockchain without decrypting it. We elaborate on the specific construction of VmppOrder and analyze the security of the constructed circuit with zero-knowledge proof. We implement a prototype of the model on Hyperledger Fabric based on Libsnark and conduct comprehensive experiments to evaluate its performance. Our experimental results validate the efficiency of the proposed model. Its order proof generation takes about 6.31 seconds, the order verification takes only 2.58 milliseconds, and the transaction processing speed is about 660 transactions per second on a moderately equipped machine.
Xiaoyan Hu 0007, Guang Cheng 0001, Honggang Chen, Zhichao Liang
WCNC1
2022 Efficient sharing of privacy-preserving sensing data on consortium blockchain via group key agreement
Xiaoyan Hu 0007, Xiaoyi Song, Guang Cheng 0001, Hua Wu 0004
Comput. Commun.1
2022 Identifying Ethereum traffic based on an active node library and DEVp2p features
Xiaoyan Hu 0007, Zhongqi Tong, Guang Cheng 0001, Ruidong Li 0001, Hua Wu 0004
Future Gener. Comput. Syst.1
2021 Detecting Cryptojacking Traffic Based on Network Behavior Features
abstract
Bitcoin and other digital cryptocurrencies have de-veloped rapidly in recent years. To reduce hardware and power costs, many criminals use the botnet to infect other hosts to mine cryptocurrency for themselves, which has led to the proliferation of mining botnets and is referred to as cryptojacking. At present, the mechanisms specific to cryptojacking detection include host-based, Deep Packet Inspection (DPI) based, and dynamic network characteristics based. Host-based detection requires detection installation and running at each host, and the other two are heavyweight. Besides, DPI-based detection is a breach of privacy and loses efficacy if encountering encrypted traffic. This paper de-signs a lightweight cryptojacking traffic detection method based on network behavior features for an ISP, without referring to the payload of network traffic. We set up an environment to collect cryptojacking traffic and conduct a cryptojacking traffic study to obtain its discriminative network traffic features extracted from only the first four packets in a flow. Our experimental study suggests that the machine learning classifier, random forest, based on the extracted discriminative network traffic features can accurately and efficiently detect cryptojacking traffic.
Xiaoyan Hu 0007, Zhuozhuo Shu, Xiaoyi Song, Guang Cheng 0001
GLOBECOM1
2021 Accurate and Fast Detection of DDoS Attacks in High-Speed Network with Asymmetric Routing
abstract
The existing DDoS attack detection methods based on a single monitoring point only consider symmetric routing scenarios, which may not be practical. Such schemes will produce high false positives when facing the asymmetric routing scenarios. Besides, few of them are applicable in high-speed networks. The paper designs a DDoS detection scheme customized for high-speed networks and takes asymmetric routing scenarios into account. Systematic sampling is applied to high-speed incoming traffic, and a proposed Double Composite Structure Sketch (DCSS) is utilized for fast recording and extraction of features based on the characteristics of DDoS attacks in both symmetric and asymmetric routing scenarios. Then classifiers are trained for online DDoS detection. Our experimental results using the public dataset show that in a 10Gbps network with asymmetric routing, our approach can accurately detect UDP Flood and SYN Flood attacks within 20 seconds when the sampling rate is set to 1/2048.
Hua Wu 0004, Tingzheng Chen, Ziling Shao 0002, Guang Cheng 0001, Xiaoyan Hu 0007
GLOBECOM5
2021 BCAC: Batch Classifier based on Agglomerative Clustering for traffic classification in a backbone network
abstract
Backbone network is the core part of the Internet. Due to the high transmission speed of traffic in the backbone network, Quality of Service (QoS) monitoring of services in the backbone network becomes a highly important and challenging issue. Traffic classification is the basis of QoS monitoring. The existing traffic classification is based on full traffic, which is impractical in high-speed backbone network traffic. This paper presents a method to classify the sampled traffic and gives an example of its application in QoS monitoring. Specifically, we design the Multiple Counter Sketch (MC Sketch) to quickly extract features from the sampled data stream in a backbone, propose the Batch Classifier based on Agglomerative Clustering (BCAC) for unsupervised clustering of traffic, and combine with the supervised machine learning method to train the labeled data in the clustering results to get the classification model. The experimental results of sampled traffic collected on a 10Gbps link show that even when the sampling ratio is 1:1024, the accuracy of our classification model reaches 96.3%. When different block sizes are set, the average clustering time of BCAC is only about one-third of the traditional agglomerative classifier. Moreover, we give an example of applying our traffic classification method to monitor the QoS, and the results show that our method can efficiently and accurately monitor the QoS dynamics of backbone network traffic.
Hua Wu 0004, Xiying Chen, Guang Cheng 0001, Xiaoyan Hu 0007, Youqiong Zhuang
IWQoS4
2021 Towards Efficient Co-audit of Privacy-Preserving Data on Consortium Blockchain via Group Key Agreement
abstract
Blockchain is well known for its storage consistency, decentralization and tamper-proof, but the privacy disclosure and difficulty in auditing discourage the innovative application of blockchain technology. As compared to public blockchain and private blockchain, consortium blockchain is widely used across different industries and use cases due to its privacy-preserving ability, auditability and high transaction rate. However, the present co-audit of privacy-preserving data on consortium blockchain is inefficient. Private data is usually encrypted by a session key before being published on a consortium blockchain for privacy preservation. The session key is shared with transaction parties and auditors for their access. For decentralizing auditorial power, multiple auditors on the consortium blockchain jointly undertake the responsibility of auditing. The distribution of the session key to an auditor requires individually encrypting the session key with the public key of the auditor. The transaction initiator needs to be online when each auditor asks for the session key, and one encryption of the session key for each auditor consumes resources. This work proposes GAChain and applies group key agreement technology to efficiently co-audit privacy-preserving data on consortium blockchain. Multiple auditors on the consortium blockchain form a group and utilize the blockchain to generate a shared group encryption key and their respective group decryption keys. The session key is encrypted only once by the group encryption key and stored on the consortium blockchain together with the encrypted private data. Auditors then obtain the encrypted session key from the chain and decrypt it with their respective group decryption key for co-auditing. The group key generation is involved only when the group forms or group membership changes, which happens very infrequently on the consortium blockchain. We implement the prototype of GAChain based on Hyperledger Fabric framework. Our experimental studies demonstrate that GAChain improves the co-audit efficiency of transactions containing private data on Fabric, and its incurred overhead is moderate.
Xiaoyan Hu 0007, Xiaoyi Song, Guang Cheng 0001, Honggang Chen, Zhichao Liang
MSN1
2021 SFIM: Identify user behavior based on stable features
Hua Wu 0004, Qiuyan Wu, Guang Cheng 0001, Shuyi Guo, Xiaoyan Hu 0007, Shen Yan 0005
Peer-to-Peer Netw. Appl.5
2020 Towards Network Coding and Request Pipelining Enabled NDN for Big Data Transmission
abstract
Two intrinsic features of Named Data Networking(NDN), in-network caching and multipath communication, offer the potential for fast and reliable big data transmissions via multisource content delivery. Network coding has recently been utilized to achieve efficient multisource content delivery in NDN. On the other hand, request pipelining is essential for efficient multisource content delivery in network coding enabled NDN. However, it is a challenge to simultaneously support network coding and request pipelining in NDN in a cost-effective way. To address this problem, we propose NCP-NDN, a network coding and request pipelining enabled named data networking architecture. NCP-NDN supports reasonably efficient Interest aggregation when enabling request pipelining without undermining the privacy of content retrieval by extending each Interest with a session and generation oriented requester identifier. Besides, NCP-NDN provides consumers with linearly independent blocks while request pipelining is enabled by combining rank-based matching, one forwarding of each block for per requester and face, and recoding the matching cached blocks before replying an Interest. Our experimental studies illustrate that NCP-NDN improves the performance of content delivery and reduces the overhead of big data transmissions as compared to the existing schemes.
Xiaoyan Hu 0007, Xiaoyi Song, Shaoqi Zheng, Ruidong Li 0001, Guang Cheng 0001
GLOBECOM1
2020 A Demand and Responsiveness-based Caching Strategy for Network Coding Enabled NDN
abstract
In-network caching and multipath forwarding are prominent features of Named Data Networking (NDN). Network coding enabled NDN (NC-NDN) coordinates the in-network caching and multipath forwarding to improve content delivery performance. However, the existing NC-NDN lacks the considerations on the incorporation with caching schemes. It basically adopts the caching scheme of Caching Everything Everywhere (CEE) leading to cache redundancy and unnecessary and frequent cache replacement. On the other hand, the existing caching schemes for native NDN do not make use of the characteristic of network coded Data packets. To address these problems, this paper proposes a demand and responsiveness-based caching strategy specific to NC-NDN to enable cost-effective caching for network coded Data packets. In our proposed caching strategy, the caching decision at a caching node takes into account three factors, its present demand on the network coded Data packets of the requested generation of content, its distance to the original content provider, and its potential responsiveness to the future requests for the same generation based on the number of network coded Data packets locally cached and that would return for the requested generation. It commits to cache network coded Data packets of a generation of content at more valuable nodes along the transmission path. Our experimental studies show that the proposed caching strategy offers high-performance content delivery and reduces the caching overhead as compared to the existing strategies.
Xiaoyan Hu 0007, Shaoqi Zheng, Ruidong Li 0001, Guang Cheng 0001
GLOBECOM1
2020 An on-demand off-path cache exploration based multipath forwarding strategy
Xiaoyan Hu 0007, Shaoqi Zheng, Guoqiang Zhang 0004, Lixia Zhao, Guang Cheng 0001, Ruidong Li 0001
Comput. Networks1
2019 Exploration and Exploitation of Off-path Cached Content in Network Coding Enabled Named Data Networking
abstract
Named Data Networking (NDN) intrinsically supports in-network caching and multipath forwarding. The two salient features offer the potential to simultaneously transmit content segments that comprise the requested content from original content publishers and in-network caches. However, due to the complexity of maintaining the reachability information of off-path cached content at the fine-grained packet level of granularity, the multipath forwarding and off-path cached copies are significantly underutilized in NDN so far. Network coding enabled NDN, referred to as NC-NDN, was proposed to effectively utilize multiple on-path routes to transmit content, but off-path cached copies are still unexploited. This work enhances NC-NDN with an On-demand Off-path Cache Exploration based Multipath Forwarding strategy, dubbed as O2CEMF, to take full advantage of the multipath forwarding to efficiently utilize off-path cached content. In O2CEMF, each network node reactively explores the reachability information of nearby off-path cached content when consumers begin to request a generation of content, and maintains the reachability at the coarse-grained generation level of granularity instead. Then the consumers simultaneously retrieve content from the original content publisher(s) and the explored capable off-path caches. Our experimental studies validate that this strategy improves the content delivery performance efficiently as compared to that in the present NC-NDN.
Xiaoyan Hu 0007, Shaoqi Zheng, Lixia Zhao, Guang Cheng 0001
ICNP1
2019 VATE: A trade-off between memory and preserving time for high accurate cardinality estimation under sliding time window
Jie Xu 0014, Wei Ding 0001, Xiaoyan Hu 0007, Qiushi Gong
Comput. Commun.3
2018 Most Memory Efficient Distributed Super Points Detection on Core Networks
Jie Xu 0014, Wei Ding 0001, Xiaoyan Hu 0007
ICA3PP (1)3
2016 Social relationship discovery of IP addresses in the managed IP networks by observing traffic at network boundary
Ahmad Jakalan, Xiaoyan Hu 0007, Abdeldime M. S. Abdelgader
Comput. Networks4
2015 Enhancing in-network caching by coupling cache placement, replacement and location
abstract
As a distinctive feature of Information Centric Networking (ICN), in-network caching plays a fundamental role on system performance. The line-speed requirement of in-network caching invalidates the employ of complex collaborative caching schemes. The cache management of in-network caching includes three components - cache placement, cache replacement and cache location. Coupling the three pieces would make more efficient use of in-network caches, but existing in-network caching schemes consider only one or two of the three pieces. This work enhances in-network caching with a low complexity cache placement scheme that takes into account content popularity, hop reduction gains, cache space contention and replacement penalty and couples with cache replacement and location, here dubbed PRL (coupling cache Placement, Replacement and Location). PRL keeps data chunks that are more popular and farther away to fetch at an en-route router with less cache space contention. And PRL locates cached copies so as to serve a higher proportion of requests from in-network caches. Our preliminary simulation results suggest that PRL increases cache hit ratio and reduces the average hop count traversed by users' requests and the caching operations at routers as compared to existing representative innetwork caching schemes.
Xiaoyan Hu 0007, Guang Cheng 0001, Chengyu Fan
ICC1
2013 Not So Cooperative Caching in Named Data Networking
abstract
This work designs and implements a Not So Cooperative Caching system for Information Centric Networking (ICN). We consider a network comprised of selfish nodes; each is with caching capability and an objective of reducing its own access cost by fetching data from its local cache or from neighboring caches. These nodes would cooperate in caching and sharing content if and only if they each benefit. The challenges are to determine what objects to cache at each node and to implement the system in the context of Named Data Networking (NDN), a large effort that exemplifies ICN. Our results include both a solution for the Not So Cooperative Caching problem and an NDN design and implementation. We evaluate our approach by deploying the system we developed on PlanetLab and show that it improves the content hit ratio by up to 13%.
Xiaoyan Hu 0007, Christos Papadopoulos, Daniel Massey
GLOBECOM1
2013 PhD forum: Not so cooperative caching
abstract
This work proposes a scheme to promote autonomous and selfish NDN (Named Data Networking) peering domains to cooperate in caching, here dubbed Not So Cooperative Caching (NSCC). We consider a network comprised of selfish nodes; each is with a caching capability and an objective of reducing its own access cost by fetching data from local cache or from neighboring caches. The challenge is to determine what objects to cache at each node so as to induce low individual node access costs, and the realistic access “price” model which allows various access “prices” of different node pairs further complicates the decision making. NSCC attempts to identify mistreatment-free object placement to incur implicit cooperation even among these selfishly behaving domains, and to further identify Nash equilibrium object placement from mistreatment-free object placements so that no domain can unilaterally change its placement and benefit while the others keep theirs unchanged, and to improve the cooperation performance with respect to fairness So far, using a game-theoretic approach NSCC seeks a global object placement in which the individual node access costs are reduced as compared to that when they operate in isolation and achieves Nash equilibrium. Our preliminary experiments with IBR verified its effectiveness. And we discuss the specific issues of NSCC's implementation in NDN.
Xiaoyan Hu 0007
ICNP1