Haiyang Jiang 0001

dblp:82/5266-1 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 3 first-author · 11 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Not All Pretrained Representation has The Sweet Danger of Sugar: Robust and Trustworthy Representation Learning for Encrypted Malicious Traffic Identification
abstract
Identifying encrypted malicious traffic is a key challenge in network security. Although pre-training techniques based on self-supervised learning are a trend towards reducing dependence on labeled data, existing mainstream methods that directly apply the architecture of natural language processing and masked language modeling tasks have been shown to be "high sugar" by recent research. They may rely on specific "shortcuts" rather than robust traffic behavioral representations to achieve inflated performance. How to extract more robust and generalizable features from covert and sparse encrypted malicious traffic is a major challenge. Therefore, we propose SUGARLESS, a robust representation learning for encrypted malicious traffic identification. To depart from the BERT-based paradigm, we first propose a spatial-temporal contrastive learning as the pre-training task that aligns temporal and spatial modal features without using NLP-style objectives, encouraging the model to learn cross-modal traffic correlations between modalities. We also propose a traffic-specific prompt-tuning mechanism to bridge the gap between pre-training and downstream tasks. Meanwhile, we develop a spatial-temporal feature fusion module to maintain this alignment during fine-tuning. Experiments on two public malware traffic datasets show that SUGARLESS achieves the best precision, recall, and F1-score with competitive accuracy, improves the average F1-score by 9.26% over YaTC, and is more robust under packet loss and reordering.
Mingyu Qiao, Zulong Diao, Guangxing Zhang, Wanhua Li, Haiyang Jiang 0001, Zhenyu Li 0001, Gaogang Xie
APNet6
2026 Not All Flows are Worth N Packets: Robust Encrypted Traffic Classification via Dynamic Patch-level Feature Learning
abstract
Network traffic classification is significant for modern network security. The widespread use of encryption protocols, such as TLS, has resulted in fewer identifiable features of traffic, rendering encrypted traffic classification a challenging task. However, existing flow-level classification methods typically rely on the features of the first N packets. This fixed-length paradigm, truncating long flows and padding short flows with invalid data, is difficult to adapt to the dynamic distribution of traffic length, which limits robustness in interference environments such as packet loss. Considering this limitation, we propose DART, a robust encrypted traffic classification method via dynamic patch-level feature learning. We first design a robust feature representation method to improve robustness against interference, which divides a flow into multiple sub-flows named patches to extract patch-level features, replacing the interference-prone packet-by-packet sequence and significantly reducing the length of the feature sequence. We also propose a sample-adaptive dynamic inference mechanism in response to traffic heterogeneity. The input scale is adaptively selected based on traffic complexity, and simple samples are allowed to exit early in shallow layers, achieving a balance between accuracy and efficiency. The experimental results on two public malware traffic datasets demonstrate that DART exhibits excellent anti-interference robustness and improves classification accuracy by 6.59%-30.51%, while maintaining high classification efficiency compared to existing mainstream methods.
Mingyu Qiao, Zulong Diao, Guangxing Zhang, Haiyang Jiang 0001, Zhenyu Li 0001, Gaogang Xie
APNet8
2026 Task-Aware Network Traffic Label Reuse With Schema-Gated Evidence
abstract
Network traffic labels are built from heterogeneous evidence: capture metadata, port rules, service names, model predictions, and expert decisions. These sources are not interchangeable across tasks: HTTP can support an application label but cannot prove attack behavior. We propose TaskLabel, a schema-gated framework that makes traffic-label reuse traceable and reviewable. TaskLabel stores each label as a schema-bound artifact with traffic unit, target schema, confidence, provenance, evidence requirement, and reuse boundary. Its compatibility gate admits weak-source votes only when source semantics match the target dataset schema; otherwise, the cue remains provenance and is routed for review. In a retrospective pilot on 1652 held-out RT-IoT2022 flows, non-model cues reach 64.6% of flows, but direct reuse disagrees with held-out target labels on 32.5% of reached flows. Schema-agnostic weak fusion falls to 58.5% macro-F1, whereas TaskLabel blocks incompatible votes and preserves the feature-only target predictor at 93.3% macro-F1 while routing a 16.4% queue that captures an oracle-estimated 86.8% of target-label errors.
Ming You, Guangxing Zhang, Mingyu Qiao, Haiyang Jiang 0001, Jinsheng Zhao
APNet5
2026 Graph-based fast-flux domain detection using graph neural networks
Haiyang Jiang 0001, Hongtao Guan
Comput. Networks3
2023 Towards Diagnosing Accurately the Performance Bottleneck of Software-Based Network Function Implementation
Ru Jia, Haiyang Jiang 0001, Serge Fdida, Gaogang Xie
PAM3
2023 Cognition: Accurate and Consistent Linear Log Parsing Using Template Correction
Zulong Diao, Haiyang Jiang 0001, Gaogang Xie
J. Comput. Sci. Technol.3
2022 LogDAC: A Universal Efficient Parser-based Log Compression Approach
abstract
A large volume of logs provides a reliable data source for online services. For instance, 5G, Big 5G or Beyong 5G cloud core network and its User Plane Functions require analysis built on massive log storage. The storage cost of log data remains a problem for the industry. Compressing the logs before archives remains the most popular way to reduce it. However, structured logs, which usually present in tabular format, and unstructured logs, which build on variable templates, have to use different compression strategies. Using log parsers can eliminate this difference by converting unstructured logs into structured ones. Based on this idea, we introduce our universal efficient parser-based log compression approach LogDAC. It also uses a divide-and-conquer preprocess strategy to improve the compatibility with successive dictionary-based general compressors. The evaluation against both structured and unstructured public datasets show a solid improvement up to 45.4% and 489.53% of compression ratio respectively, compared with existing parser-based compressors.
Zulong Diao, Haiyang Jiang 0001, Gaogang Xie
ICC3
2022 A deep dive into DNS behavior and query failures
Donghui Yang, Zhenyu Li 0001, Haiyang Jiang 0001, Gareth Tyson, Gaogang Xie
Comput. Networks3
2021 Compact-index: an efficient index algorithm for network traffic
abstract
In many network security systems, network packets will be archived with no loss for the purpose of forensic, troubleshooting and so on. In order to achieve fast retrieval for these stored packets, index is essential. However, with the rapid increase of network link bandwidth, indexing network traffic traces is facing the challenges of index construction speed, index storage overhead and retrieval efficiency. In this paper, we propose an efficient index scheme for packet index, named Compact-Index, that not only effectively reduces the storage cost of index, but also greatly improves index insertion rate and retrieval efficiency. Experimental results show that our scheme can achieve 7.14Mpps index insertion rate for IPv4 traffic and 6.11Mpps index insertion rate for IPv6 traffic, and supports millisecond scale response to most queries. In addition, the average space cost of index is only about 5% of the raw network traffic. All the experimental results above indicate that Compact-Index significantly outperforms the existing state-of-art index schemes.
Guangxing Zhang, Haiyang Jiang 0001, Gaogang Xie
CoNEXT3
2021 An Accuracy Network Anomaly Detection Method Based on Ensemble Model
abstract
Identifying network anomaly detection is important since they may carry critical information in circumstances such as a burst of intrusions, privacy theft, system damage and fraudulent activities. In recent years, there are many detection methods for network anomalies are proposed, however, a single model always faces the problems of over or under-fitting, high bias and variance. An improved method is to comprehensively use the results of multiple models and then reform the final predictions. This paper introduces an ensemble model, which is a powerful technique to increase accuracy on network anomaly detection. By combining three base models Xgboost, LightGBM and Catboost into one anomaly detector, we successfully detect different DDOS-smurf and Probing activities. This ensemble model is verified on ZYELL-NCTU net traffic, which is a large-scale dataset for read-world network anomaly detection. All code are open source in Github and can be directly run on Colab Jupyter Notebook.
Haiyang Jiang 0001, Gaogang Xie
ICASSP4
2021 CyCo: A Temporal Cycle Consistency Based Labeling Method for Time Series Data
Haiyang Jiang 0001, Zulong Diao, Yanbiao Li 0001, Gaogang Xie
IJCNN2
2021 C2QoS: CPU-Cycle based Network QoS Strategy in vSwitch of Public Cloud
Haiyang Jiang 0001, Yulei Wu, Yilong Lv, Xing Li 0007, Gaogang Xie
IM2
2021 NeVe: A Log-based Fast Incremental Network Feature Embedding Approach
abstract
Similarity (distance) measurement among network features (e.g. IP address, MAC address, port number, and protocol, etc.) based on network logs is a critical step for data mining in intrusion detection, anomaly prediction, and log analysis. A practical approach is necessarily accurate, fast, and incremental due to the dynamic network environment. However, existing solutions fail to satisfy these demands simultaneously. Therefore, we propose a novel unsupervised network feature embedding approach: Network Vector (NeVe). It learns the similarity from context information by introducing a natural language processing algorithm GloVe. Since the network data is more timeliness with an almost infinite corpus size, we adjust the algorithm to adapt the input data format and design a fast scalable online update mechanism. Our evaluation demonstrates that NeVe can achieve the highest accuracy with minimal time consumption (13 ~ 15 times faster) compared with the state-of-the-art approach.
Zulong Diao, Haiyang Jiang 0001, Gaogang Xie
ISCC3
2021 DSQNet: Domain SeQuence based Deep Neural Network for AGDs Detection
abstract
Modern botnets widely rely on Algorithmically Generated Domains (AGDs) to contact with Command-and-Control (C&C) servers. Existing AGD detection solutions check the domains one by one based on the structural differences between AGD and benign ones, e.g., some AGD families show much more random character composition than legitimate ones. These methods can hardly deal with the newly emerged camouflage technology based AGD types, as each individual AGD seems benign in domain structure features of itself. In this work, the structural correlations among AGDs are analyzed and we find the inter-AGD correlation can be adopted for the AGD detection. We then propose DSQNet, a Domain SeQuence based Deep Neural Network AGD detection model, that simultaneously checks the domains in batch to take the inter-AGD correlation into consideration during the detection. Experiments on the public and real-world dataset show the superiority of the proposed approach.
Haiyang Jiang 0001, Hongtao Guan
ISCC2
2021 S2H: Hypervisor as a setter within Virtualized Network I/O for VM isolation on cloud platform
Haiyang Jiang 0001, Guangxing Zhang, Xin Wang 0001, Yilong Lv, Xing Li 0007, Serge Fdida, Gaogang Xie
Comput. Networks2
2021 C2QoS: Network QoS guarantee in vSwitch through CPU-cycle management
Haiyang Jiang 0001, Yulei Wu, Chunjing Han, Yilong Lv, Xing Li 0007, Serge Fdida, Gaogang Xie
J. Syst. Archit.2
2019 A Massively Multi-Tenant Virtualized Network Intrusion Prevention Service on NFV Platform
abstract
Multi-Tenancy (MT) is critical for Network Function Virtualization (NFV) platform as it reduces the cost of having network services by sharing expensive server resource among customers. This is especially critical for memory and CPU intensive services like Network Intrusion Prevention System (NIPS). In this work, we explore the issue of deploying a large-scale virtualized NIPS service on a commercial NFV platform. We observe that the scalability of NIPS service is not good when based on independent Virtual Machines (VMs). We propose a Multi-Tenant Aho-Corasick state machine data structure (MT-AC) and adapt it into NIPS to solve the issue. One MT-AC based NIPS service simultaneously checks traffic belonging to different tenants against a merged ruleset. The MT-AC data structure is very efficient as it eliminates the redundancies among tenants' signatures during the rulesets merging. Experimental results with real-world ruleset show that, in comparison with an independent VM-based solution, the MT-AC based NIPS service can support 2 to 4 times more tenants. Moreover, the throughput and latency performance of MT-AC based NIPS engine only degrades by 1%, when the tenant count increases from 8 to 128. The results validate that, the proposed MT-AC based NIPS service on NFV platform can support a large amount of tenants with a very low cost.
Haiyang Jiang 0001, Hongtao Guan, Gaogang Xie, Kavé Salamatian
ICCCN1
2013 Scalable high-performance parallel design for Network Intrusion Detection Systems on many-core processors
abstract
Network Intrusion Detection Systems (NIDSes) face significant challenges coming from the relentless network link speed growth and increasing complexity of threats. Both hardware accelerated and parallel software-based NIDS solutions, based on commodity multi-core and GPU processors, have been proposed to overcome these challenges. This work explores new parallel opportunities afforded by many-core processors for high performance, scalable and inexpensive NIDS. We exploit the huge many-core computational power by adopting a hybrid parallel architecture combining data and pipeline parallelism. We also design a hybrid load balancing scheme, using both ruleset and flow space partitioning. Furthermore, the proposed design leverages particular features of the processor to break the bottlenecks. We have integrated the open source NIDS Suricata into our proposed design and evaluated its performance with synthetic traffic. The prototype exhibits almost linear speedup and can handle up to 7.2 Gbps traffic with 100-bytes packets.
Haiyang Jiang 0001, Guangxing Zhang, Gaogang Xie, Kavé Salamatian, Laurent Mathy
ANCS1
2013 Efficient fingerprint extraction for high performance Intrusion Detection System
abstract
Deep Packet Inspection (DPI) module in Intrusion Detection Systems (IDSes) consists of two components: Pre-filter and Rule Verification (RV). Pre-filter adopts Multi-Pattern Matching (MPM) engine to filter out the vast majority of benign packets and then leave a few suspicious packets with false positives into RV component. These false positives are due to the scanning process in the pre-filter: it detects the traffic in a single pass against a set of fingerprints, which are extracted from the given ruleset by selecting only a small portion of the patterns in each signature. RV component precisely checks the suspicious packets and eliminates these false positives. The performance of DPI module is related to the extracted fingerprint set. An efficient fingerprint set should improve the pre-filter throughput, and at the same time decrease the count of checking activities in RV component. We show in this paper that these two requirements cannot be simultaneously satisfied in the existing fingerprint extraction strategies. Pre-filter performance greatly benefits from smaller fingerprint set because of the more compact MPM engine. But RV component suffers from the higher rate of false positives caused by the smaller fingerprint set. We optimally trade off these two requirements with a new extraction method in this work. Through analysing a small amount of training traffic in the initial phase, our strategy gives each fingerprint candidate an empirical weight for the subsequent extraction. Experimental results obtained by integrating our proposed method into the Snort IDS show that our strategy improves the IDS average throughput by at least 69% over the latest real ruleset and real traffic.
Haiyang Jiang 0001, Gaogang Xie, Kavé Salamatian
ISCC1
2011 Exploring and Enhancing the Performance of Parallel IDS on Multi-core Processors
abstract
With the advancement of multi-core processor, it is highly desired to use parallel design to improve the IDS throughput in nowadays. However, existing parallel schemes often fail to achieve linear speedup in IDS. The throughput even deteriorates severely for some network traffics. Hence, exploring and addressing the problems have become one of the most urgent issues in parallel IDS. In this work, an IDS model adopting software pipeline is developed as the test bed of parallel performance evaluation. The contribution of this paper is two-fold. First, we explore the performance of parallel IDS through substantive experiments and quantitative analyses. We find Heavy Rule Fingerprint (HRF) in the pre-filter, which has not been mentioned in previous papers as we know, causes severe performance deterioration mentioned in existing studies. The experiments illustrate that the time consumption on HRF in parallel system is at least 6-7 times longer than that in normal system. Second, we propose a new fingerprint extraction strategy to deal with HRF. Experimental results show that the throughput deterioration is resolved completely and the throughput is enhanced by 45% by integrating our proposed method into the test bed and with making use of DARPA evaluation dataset.
Haiyang Jiang 0001, Gaogang Xie
TrustCom1