EDBT 2026 Demo / reviewers in the wild / expert
Ruijie Zhao 0001
dblp:92/10854-1
· DBLP profile ↗
28ranked-venue papers
14as first author
28since 2021 · last 2026
0000-0001-6168-8687ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 11 · 2 first-author · 11 since 2021Computer networks · 10 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Flow Semantics for Encrypted Traffic Analysis: A Contrastive Pre-Training ApproachabstractEncrypted traffic analysis is crucial for cyberspace security. Self-supervised learning shows great promise to enhance traffic analysis with the pre-trained traffic encoder, which is constructed using large-scale, readily available unlabeled traffic data. However, existing approaches struggle to handle the increasingly prevalent encrypted traffic, as their generative reconstruction tasks cannot process encrypted content. To this end, we propose TACO, a robust and flexible encrypted traffic analysis system based on flow semantics learning. Specifically, we first design several feasible traffic data augmentation strategies to prepare flow semantics knowledge from the unlabeled traffic. Then, our traffic encoder with a traffic partition module learns the semantics knowledge based on the contrastive pre-training paradigm. It serves as a traffic foundation encoder that can comprehend flow semantics and extract effective semantic representations. Finally, we fine-tune the traffic encoder to leverage flow semantics for various downstream encrypted traffic analysis tasks. The experimental results illustrate that TACO outperforms the optimal baseline by 7.5% in average F1 score on four traffic classification datasets and achieves an improvement of at least 11.62% in average F1 score on the three transfer tasks, while indicating superior efficiency. We will release the source code as well as the experiment data upon publication to foster future research. Ruijie Zhao 0001, Mingwei Zhan, Qi Li 0002, Zhuotao Liu, Xianwen Deng, Guang Cheng 0001, Zhi Xue, Ke Xu 0002 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Robust Training of Efficient Traffic Classifier with Noisy Labels
Zuoyu Qiu, Mingwei Zhan, Xianwen Deng, Zhi Xue, Ruijie Zhao 0001 |
Inscrypt (2) | 6 |
| 2025 | Detecting Malicious Encrypted Traffic with Multimodal RepresentationsabstractThe rapid advancement of encryption technology enhances network security while enabling hidden attackers to avoid detection. Traditional methods for malicious encrypted traffic detection, which predominantly rely on a single modality such as statistical features or content representations, often fall short of adapting to dynamic network environments. Methods based on graph representations grapple with challenges such as insufficient modeling of the encryption properties and substantial computational resource requirements. Multimodal-based methods seldom consider the graph-based dynamic representation and often overlook the differences in feature spaces. Moreover, these methods are not evaluated for universality across platforms. To solve challenges above, we propose M2D, a multimodal-based framework for malicious encrypted traffic detection suitable for all versions of TLS protocols. M2D extracts (a) heterogeneous graph representation from spatial and temporal features to capture both dynamic patterns and complex interactions between different entities; (b) ciphertext visual representation to enhance content encapsulation; and (c) plaintext representation to explore semantics, then fuses them through the multi-head attention mechanism to emphasize more effective components. Furthermore, we set up an encrypted network traffic dataset generated by sandbox, with session keys embedded for decryption. Experimental results on both public and proposed datasets demonstrate the superior performance of M2D in binary and multi-class classification tasks. Additionally, ablation studies confirm the effectiveness of each component. Ruijie Zhao 0001, Libo Chen 0001, Lingyun Ying, Zhengguang Han, Zhi Xue |
ICC | 2 |
| 2025 | Multi-modal Datagram Representation with Spatial-Temporal State Space Models and Inter-flow Contrastive Learning for Encrypted Traffic Classification
Xianwen Deng, Ruijie Zhao 0001, Mingwei Zhan, Shaoqian Wu, Zhi Xue |
ICICS (3) | 2 |
| 2025 | FlowRefiner: A Robust Traffic Classification Framework against Label NoiseabstractNetwork traffic classification is essential for network management and security. In recent years, deep learning (DL) algorithms have emerged as essential tools for classifying complex traffic. However, they rely heavily on high-quality labeled training data. In practice, traffic data is often noisy due to human error or inaccurate automated labeling, which could render classification unreliable and lead to severe consequences. Although some studies have alleviated the label noise issue in specific scenarios, they are difficult to generalize to general traffic classification tasks due to the inherent semantic complexity of traffic data. In this paper, we propose FlowRefiner, a robust and general traffic classification framework against label noise. FlowRefiner consists of three core components: a traffic semantics-driven noise detector, a confidence-guided label correction mechanism, and a cross-granularity robust classifier. First, the noise detector utilizes traffic semantics extracted from a pre-trained encoder to identify mislabeled flows. Next, the confidence-guided label correction module fine-tunes a label predictor to correct noisy labels and construct refined flows. Finally, the cross-granularity robust classifier learns generalized patterns of both flow-level and packet-level, improving classification robustness against noisy labels. We evaluate our method on four traffic datasets with various classification scenarios across varying noise ratios. Experimental results demonstrate that FlowRefiner mitigates the impact of label noise and consistently outperforms state-of-the-art baselines by a large margin. The code is available at https://github.com/NSSL-SJTU/FlowRefiner. Mingwei Zhan, Ruijie Zhao 0001, Xianwen Deng, Zhi Xue, Qi Li 0002, Zhuotao Liu, Guang Cheng 0001, Ke Xu 0002 |
NeurIPS | 2 |
| 2025 | Countmamba: A Generalized Website Fingerprinting Attack via Coarse-Grained Representation and Fine-Grained PredictionabstractTor is the leading low-latency anonymous communication network, widely used to protect users' privacy through mechanisms such as random relay selection. However, despite these defenses, Tor traffic remains susceptible to website finger-printing (WF) attacks, where attackers analyze side-channel information (e.g., packet size, direction, inter-packet timing) to infer visited websites. Although WF attacks have shown high success rates in controlled settings, they rely on complete, unperturbed traffic, making them vulnerable to real-world de-fense mechanisms. Traditional WF approaches, which typically employ Machine Learning (ML) or Deep Learning (DL) to classify packet sequences as a single-label prediction, struggle to generalize in practical scenarios, especially under defenses that alter packet patterns or in environments requiring multi-label, early-stage analysis. In this work, we introduce Countmamba, a robust and adaptable WF attack framework designed to address the challenges posed by real-world defenses, early-stage traffic analysis, and multi-tab browsing. Countmamba employs a Windowed Traffic Counting Matrix (WTCM) to create re-silient, coarse-grained traffic representations by aggregating packet events within fixed time intervals, allowing it to with-stand moderate perturbations from defenses. Additionally, a state-space-oriented (SSO) classifier incrementally generates fine-grained predictions from partial traffic data, maintaining high attack accuracy while enabling early-stage and multi-tab attack capabilities. Unlike prior WF methods, Countmamba iteratively updates predictions as new data arrives, eliminating the need for complete traffic capture and enabling reliable inference even in complex, multi-tab environments. Extensive experiments demonstrate that Countmamba outperforms state-of-the-art WF attacks across robust, early-stage, and multi-tab scenarios, highlighting its applicability for realistic, adaptive WF analysis in Tor networks. The source code as well as the experiment data is available at https://github.com/SJTU-dxw/CountMamba-WF. Xianwen Deng, Ruijie Zhao 0001, Mingwei Zhan, Zhi Xue |
SP | 2 |
| 2024 | CaptchaSAM: Segment Anything in Text-based CaptchasabstractWhile text-based captchas, designed to distinguish between human users and bots, have encountered numerous attack methods, they remain a prevalent security mechanism employed by various websites. Some deep learning-based approaches can recognize captcha character sequences end-to-end; however, the labor-intensive and time-consuming labeling process severely restricts their feasibility. In this study, we introduce CaptchaSAM, to segment anything in text-based captchas. Our insight lies in the fact that identifying individual characters is a simpler task compared to recognizing character sequences, leading to a substantial reduction in labeling dependency. To accomplish this, we utilize the Segment Anything Model (SAM) for character-level semi-automatic annotation. Subsequently, we leverage the annotated data to train a semantic segmentation model. Our experiments with real-world captcha systems demonstrate that CaptchaSAM significantly outperforms state-of-the-art methods with just a few labeled captchas. We anticipate that our research will encourage security experts to reconsider the design and deployment of text-based captchas. The source code is accessible at https://github.com/SJTU-dxw/CaptchaSAM. Weiqi Bai, Ruijie Zhao 0001, Xianwen Deng |
TrustCom | 4 |
| 2024 | Vulnerability-oriented Testing for RESTful APIs
Wenlong Du, Libo Chen 0001, Ruijie Zhao 0001, Junmin Zhu, Zhengguang Han, Zhi Xue |
USENIX Security Symposium | 5 |
| 2024 | Code is not Natural Language: Unlock the Power of Semantics-Oriented Graph Representation for Binary Code Similarity Detection
Haojie He, Xingwei Lin, Ziang Weng, Ruijie Zhao 0001, Shuitao Gan, Libo Chen 0001, Yuede Ji, Jiashui Wang, Zhi Xue |
USENIX Security Symposium | 4 |
| 2024 | A Novel Self-Supervised Framework Based on Masked Autoencoder for Traffic ClassificationabstractTraffic classification is a critical task in network security and management. Recent research has demonstrated the effectiveness of the deep learning-based traffic classification method. However, the following limitations remain: (1) the traffic representation is simply generated from raw packet bytes, resulting in the absence of important information; (2) the model structure of directly applying deep learning algorithms does not take traffic characteristics into account; and (3) scenario-specific classifier training usually requires a labor-intensive and time-consuming process to label data. In this paper, we introduce a masked autoencoder (MAE) based traffic transformer with multi-level flow representation to tackle these problems. To model raw traffic data, we design a formatted traffic representation matrix with hierarchical flow information. After that, we develop an efficient Traffic Transformer, in which packet-level and flow-level attention mechanisms implement more efficient feature extraction with lower complexity. At last, we utilize MAE paradigm to pre-train our classifier with a large amount of unlabeled data, and perform fine-tuning with a few labeled data for a series of traffic classification tasks. Experiment findings reveal that our method outperforms state-of-the-art methods on five real-world traffic datasets by a large margin. The code is available at https://github.com/NSSL-SJTU/YaTC. Ruijie Zhao 0001, Mingwei Zhan, Xianwen Deng, Fangqi Li 0001, Guan Gui 0001, Zhi Xue |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Yet Another Traffic Classifier: A Masked Autoencoder Based Traffic Transformer with Multi-Level Flow RepresentationabstractTraffic classification is a critical task in network security and management. Recent research has demonstrated the effectiveness of the deep learning-based traffic classification method. However, the following limitations remain: (1) the traffic representation is simply generated from raw packet bytes, resulting in the absence of important information; (2) the model structure of directly applying deep learning algorithms does not take traffic characteristics into account; and (3) scenario-specific classifier training usually requires a labor-intensive and time-consuming process to label data. In this paper, we introduce a masked autoencoder (MAE) based traffic transformer with multi-level flow representation to tackle these problems. To model raw traffic data, we design a formatted traffic representation matrix with hierarchical flow information. After that, we develop an efficient Traffic Transformer, in which packet-level and flow-level attention mechanisms implement more efficient feature extraction with lower complexity. At last, we utilize the MAE paradigm to pre-train our classifier with a large amount of unlabeled data, and perform fine-tuning with a few labeled data for a series of traffic classification tasks. Experiment findings reveal that our method outperforms state-of-the-art methods on five real-world traffic datasets by a large margin. The code is available at https://github.com/NSSL-SJTU/YaTC. Ruijie Zhao 0001, Mingwei Zhan, Xianwen Deng, Guan Gui 0001, Zhi Xue |
AAAI | 1 |
| 2023 | DHBE: Data-free Holistic Backdoor Erasing in Deep Neural Networks via Restricted Adversarial DistillationabstractBackdoor attacks have emerged as an urgent threat to Deep Neural Networks (DNNs), where victim DNNs are furtively implanted with malicious neurons that could be triggered by the adversary. To defend against backdoor attacks, many works establish a staged pipeline to remove backdoors from victim DNNs: inspecting, locating, and erasing. However, in a scenario where a few clean data can be accessible, such pipeline is fragile and cannot erase backdoors completely without sacrificing model accuracy. To address this issue, in this paper, we propose a novel data-free holistic backdoor erasing (DHBE) framework. Instead of the staged pipeline, the DHBE treats the backdoor erasing task as a unified adversarial procedure, which seeks equilibrium between two different competing processes: distillation and backdoor regularization. In distillation, the backdoored DNN is distilled into a proxy model, transferring its knowledge about clean data, yet backdoors are simultaneously transferred. In backdoor regularization, the proxy model is holistically regularized to prevent from infecting any possible backdoor transferred from distillation. These two processes jointly proceed with data-free adversarial optimization until a clean, high-accuracy proxy model is obtained. With the novel adversarial design, our framework demonstrates its superiority in three aspects: 1) minimal detriment to model accuracy, 2) high tolerance for hyperparameters, and 3) no demand for clean data. Extensive experiments on various backdoor attacks and datasets are performed to verify the effectiveness of the proposed framework. Code is available at https://github.com/yanzhicong/DHBE Zhicong Yan, Shenghong Li 0001, Ruijie Zhao 0001, Yuan Tian 0017 |
AsiaCCS | 3 |
| 2023 | SAWD: Structural-Aware Webshell Detection System with Control Flow GraphabstractWith the increasing prevalence of web servers, protecting them from cyber attacks has become a crucial task for online service providers.Webshells, which are backdoors to websites, are commonly used by hackers to gain unauthorized access to web servers.However, traditional methods for detecting webshells often fail to produce satisfactory results due to the use of obfuscation or encryption to conceal their characteristics.In recent years, webshell detection methods based on deep learning (DL) have received significant attention, but they struggle to preserve the syntax and semantic information contained in the source code.In this paper, we propose a structuralaware webshell detection system to address these problems, denoted as SAWD.Specifically, we first generate the control flow graph (CFG) with syntax and semantic information from the PHP source code.Then, we leverage CFG to build our graph representation, which consists of the adjacency matrix and keywords-based basic block features.Finally, based on our graph representation, we adopt convolutional neural networks (GCN) combined with graph pooling to detect webshells more efficiently.Experimental results demonstrate that our method outperforms state-of-the-art webshell detection systems on the collected dataset. Junmin Zhu, Yizhao Yao, Xianwen Deng, Yaoguang Yong, Libo Chen 0001, Zhi Xue, Ruijie Zhao 0001 |
SEKE | 8 |
| 2023 | GeeSolver: A Generic, Efficient, and Effortless Solver with Self-Supervised Learning for Breaking Text CaptchasabstractAlthough text-based captcha, which is used to differentiate between human users and bots, has faced many attack methods, it remains a widely used security mechanism and is employed by some websites. Some deep learning-based text captcha solvers have shown excellent results, but the labor-intensive and time-consuming labeling process severely limits their viability. Previous works attempted to create easy-to-use solvers using a limited collection of labeled data. However, they are hampered by inefficient preprocessing procedures and inability to recognize the captchas with complicated security features.In this paper, we propose GeeSolver, a generic, efficient, and effortless solver for breaking text-based captchas based on self-supervised learning. Our insight is that numerous difficult-to-attack captcha schemes that "damage" the standard font of characters are similar to image masks. And we could leverage masked autoencoders (MAE) to improve the captcha solver to learn the latent representation from the "unmasked" part of the captcha images. Specifically, our model consists of a ViT encoder as latent representation extractor and a well-designed decoder for captcha recognition. We apply MAE paradigm to train our encoder, which enables the encoder to extract latent representation from local information (i.e., without masking part) that can infer the corresponding character. Further, we freeze the parameters of the encoder and leverage a few labeled captchas and many unlabeled captchas to train our captcha decoder with semi-supervised learning.Our experiments with real-world captcha schemes demonstrate that GeeSolver outperforms the state-of-the-art methods by a large margin using a few labeled captchas. We also show that GeeSolver is highly efficient as it can solve a captcha within 25 ms using a desktop CPU and 9 ms using a desktop GPU. Besides, thanks to latent representation extraction, we successfully break the hard-to-attack captcha schemes, proving the generality of our solver. We hope that our work will help security experts to revisit the design and availability of text-based captchas. The code is available at https://github.com/NSSL-SJTU/GeeSolver. Ruijie Zhao 0001, Xianwen Deng, Zhicong Yan, Zhengguang Han, Libo Chen 0001, Zhi Xue |
SP | 1 |
| 2023 | Semisupervised Federated-Learning-Based Intrusion Detection Method for Internet of ThingsabstractFederated learning (FL) has become an increasingly popular solution for intrusion detection to avoid data privacy leakage in Internet of Things (IoT) edge devices. Existing FL-based intrusion detection methods, however, suffer from three limitations: 1) model parameters transmitted in each round may be used to recover private data, which leads to security risks; 2) not independent and identically distributed (non-IID) private data seriously adversely affect the training of FL (especially distillation-based FL); and 3) high communication overhead caused by the large model size greatly hinders the actual deployment of the solution. To address these problems, this article develops an intrusion detection method based on a semisupervised FL scheme via knowledge distillation. First, our proposed method leverages unlabeled data via distillation method to enhance the classifier performance. Second, we build a model based on convolutional neural networks (CNNs) for extracting deep features of the traffic packets, and take this model as both the classifier network and discriminator network. Third, the discriminator is designed to improve the quality of each client’s predicted labels, and to avoid the failure of distillation training caused by a large number of incorrect predictions under private non-IID data. Moreover, the combination of the hard-label strategy and voting mechanism further reduces communication overhead. The experiments on the real-world traffic data set with three non-IID scenarios show that our proposed method can achieve better detection performance as well as lower communication overhead than state-of-the-art methods. Ruijie Zhao 0001, Zhi Xue, Tomoaki Ohtsuki, Bamidele Adebisi, Guan Gui 0001 |
IEEE Internet Things J. | 1 |
| 2023 | A Novel Traffic Classifier With Attention Mechanism for Industrial Internet of ThingsabstractWith the development of the Industrial Internet of Things (IIoT), the complex traffic generated by large-scale IIoT devices presents challenges for traffic analysis. Most of existing deep learning-based traffic analysis methods use a single flow for classification, resulting in being misled by the irrelevant flow. Thus, it is necessary to use flow sequences for traffic analysis. However, existing models fail to effectively distinguish unimportant flows in flow sequence, which affects the classification performance. To address the aforementioned challenges, we propose a novel traffic classifier called flow transformer to perform traffic analysis with flow sequences, which leverages multihead attention mechanism to strengthen the information interaction between related flows. Besides, the RF-based feature selection method is designed to select the optimal feature combination, avoiding insignificant features from reducing the performance of the classifier. Experimental results on three real-world traffic datasets demonstrate that our method outperforms state-of-the-art methods with a large margin. Ruijie Zhao 0001, Yiteng Huang, Xianwen Deng, Yong Shi 0009, Jiabin Li, Zijing Huang, Zhi Xue |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | A Lightweight Semi-Supervised Learning Method Based on Consistency Regularization for Intrusion DetectionabstractWith the development of the Industrial Internet of Things (IIoT), more frequent attacks occur to intrude IIoT devices. A reasonably designed intrusion detection method can effectively guarantee the security of IIoT. Over the past decade, different methods of intrusion detection based on deep learning (DL) have been proposed, which helps intrusion detection keep evolving and become more robust. However, these previous researches usually require the participation of a large number of experts, and gradually become invalid with the continuous development of intrusion methods. The limited compute capability of IIoT devices also greatly hinder the deployment of overly complex DL models. To address these challenges, this paper proposes a lightweight semi-supervised learning (LSSL) method based on consistency regularization for intrusion detection. Our proposed method enhances the detection performance by using unlabeled traffic data for consistency training. Besides, we adopt separable convolutions for efficient feature extraction. Experimental results on two widely-used benchmark datasets show that the detection performance of our model is significantly improved by the consistency training, and it can effectively detect various attacks in complex networks. Ruijie Zhao 0001, Tiantian Tang, Guan Gui 0001, Zhi Xue |
ICC | 1 |
| 2022 | A Semi-Supervised Federated Learning Scheme via Knowledge Distillation for Intrusion DetectionabstractFederated learning (FL) has become an increasingly popular solution for intrusion detection to avoid data privacy leakage in Internet of Things (IoT) edge devices. However, most of the current FL-based intrusion detection methods still suffer from three limitations: (1) model parameters transmitted in each round may be used to recover private data which leads to security risks, (2) not independent and identically distributed (non-IID) private data seriously adversely affects the training of FL (especially distillation-based FL), and (3) high communication overhead caused by the large model size greatly hinders the actual deployment of the solution. To address these problems, this paper develops an intrusion detection method based on semi-supervised FL scheme via knowledge distillation. First, our proposed method leverages unlabeled data via distillation method to enhance the classifier performance. Second, we build a CNN-based model for extracting deep features of the traffic packets, and take this model as both the classifier network and discriminator network. Third, discriminator is designed to improve the quality of each client’s predicted labels, to avoid the failure of distillation training caused by a large number of incorrect predictions under private non-IID data. Moreover, the combination of hard-label strategy and voting mechanism further reduces communication overhead. Experimental results on the real-world traffic dataset show that our proposed method can achieve better classification performance as well as lower communication overhead than state-of-the-art methods. Ruijie Zhao 0001, Linbo Yang, Zhi Xue, Guan Gui 0001, Tomoaki Ohtsuki |
ICC | 1 |
| 2022 | 3E-Solver: An Effortless, Easy-to-Update, and End-to-End Solver with Semi-Supervised Learning for Breaking Text-Based CaptchasabstractText-based captchas are the most widely used security mechanism currently. Due to the limitations and specificity of the segmentation algorithm, the early segmentation-based attack method has been unable to deal with the current captchas with newly introduced security features (e.g., occluding lines and overlapping). Recently, some works have designed captcha solvers based on deep learning methods with powerful feature extraction capabilities, which have greater generality and higher accuracy. However, these works still suffer from two main intrinsic limitations: (1) many labor costs are required to label the training data, and (2) the solver cannot be updated with unlabeled data to recognize captchas more accurately. In this paper, we present a novel solver using improved FixMatch for semi-supervised captcha recognition to tackle these problems. Specifically, we first build an end-to-end baseline model to effectively break text-based captchas by leveraging encoder-decoder architecture and attention mechanism. Then we construct our solver with a few labeled samples and many unlabeled samples by improved FixMatch, which introduces teacher forcing, adaptive batch normalization, and consistency loss to achieve more effective training. Experiment results show that our solver outperforms state-of-the-arts by a large margin on current captcha schemes. We hope that our work can help security experts to revisit the design and usability of text-based captchas. The source code of this work is available at https://github.com/SJTU-dxw/3E-Solver-CAPTCHA. Xianwen Deng, Ruijie Zhao 0001, Libo Chen 0001, Zhi Xue |
IJCAI | 2 |
| 2022 | Flow Sequence-Based Anonymity Network Traffic Identification with Residual Graph Convolutional NetworksabstractIdentifying anonymity services from network traffic is a crucial task for network management and security. Currently, some works based on deep learning have achieved excellent performance for traffic analysis, especially those based on flow sequence (FS), which utilizes information and features of the traffic flow. However, these models still face a serious challenge because of lacking a mechanism to take into account relationships between flows, resulting in mistakenly recognizing irrelevant flows in FS as clues for identifying traffic. In this paper, we propose a novel FS-based anonymity network traffic identification framework to tackle this problem, which leverages Residual Graph Convolutional Network (ResGCN) to exploit relationships between flows for FS feature extraction. Moreover, we design a practical scheme to preprocess the raw data of real-world traffic, which further improves identification performance and efficiency. Experimental results on two real-world traffic datasets demonstrate that our method outperforms state-of-the-art methods by a large margin. Ruijie Zhao 0001, Xianwen Deng, Libo Chen 0001, Zhi Xue |
IWQoS | 1 |
| 2022 | MT-FlowFormer: A Semi-Supervised Flow Transformer for Encrypted Traffic ClassificationabstractWith the increasing demand for the protection of personal network meta-data, encrypted networks have grown in popularity, so do the challenge of monitoring and analyzing encrypted network traffic. Currently, some deep learning-based methods have been proposed to leverage statistical features for encrypted traffic classification, which are barely affected by encryption techniques. However, these works still suffer from two main intrinsic limitations: (1) the feature extraction process lacks a mechanism to take into account correlations between flows in the flow sequence; and (2) a large volume of manually-labeled data is required for training an effective deep classifier. In this paper, we propose a novel semi-supervised framework to address these problems. To be specific, an efficient classifier with attention mechanism is proposed to extract features from flow sequences with low computational cost. Then, a Mean Teacher-style semi-supervised framework is adopted to exploit the unlabeled traffic data, where a spatiotemporal data augmentation method is designed as the key component to explore the spatial and temporal relationship within the unlabeled traffic data. Experimental results on two real-world traffic datasets demonstrate that our method outperforms state-of-the-art methods with a large margin. Ruijie Zhao 0001, Xianwen Deng, Zhicong Yan, Zhi Xue |
KDD | 1 |
| 2022 | A Novel Intrusion Detection Method Based on Lightweight Neural Network for Internet of ThingsabstractThe purpose of a network intrusion detection (NID) is to detect intrusions in the network, which plays a critical role in ensuring the security of the Internet of Things (IoT). Recently, deep learning (DL) has achieved a great success in the field of intrusion detection. However, the limited computing capabilities and storage of IoT devices hinder the actual deployment of DL-based high-complexity models. In this article, we propose a novel NID method for IoT based on the lightweight deep neural network (LNN). In the data preprocessing stage, to avoid high-dimensional raw traffic features leading to high model complexity, we use the principal component analysis (PCA) algorithm to achieve feature dimensionality reduction. Besides, our classifier uses the expansion and compression structure, the inverse residual structure, and the channel shuffle operation to achieve effective feature extraction with low computational cost. For the multiclassification task, we adopt the NID loss that acts as a better loss function to replace the standard cross-entropy loss for dealing with the problem of uneven distribution of samples. The results of experiments on two real-world NID data sets demonstrate that our method has excellent classification performance with low model complexity and small model size, and it is suitable for classifying the IoT traffic of normal and attack scenarios. Ruijie Zhao 0001, Guan Gui 0001, Zhi Xue, Tomoaki Ohtsuki, Bamidele Adebisi, Haris Gacanin |
IEEE Internet Things J. | 1 |
| 2022 | SEAF: A Scalable, Efficient, and Application-independent Framework for container security detectionabstractContainer technology has become a popular development that can conveniently accelerate building, running, and sharing applications. However, a container image packaging a collection of software usually lurks various defects threatening consumer safety, such as embedded malware, software vulnerability, privacy leakage, etc. Moreover, developers and users share container images through a centralized, public, and massive repository (e.g., Docker Hub), which can magnify the impact of these security defects in a fast-spreading way. Unfortunately, existing detection methods cannot effectively or efficiently discover such hidden flaws among the numerous images. This paper proposes a novel method to effectively detect and measure container security flaws embedded in images. Based on the crucial insight that container images are constructed hierarchically, each image depends on layers of forwarding image and adds updated content in layers of itself. Our work mines a Global Relationship Tree (GRT) based on dependency among the images that contain common layers. Meanwhile, by traversing the GRT and leveraging content differential analysis, we can locate the changing content in an image corresponding to defects. Therefore, when checking flaws among numerous images, we make a layer-sensitive detection by reusing common layers’ detection results in iterative processes to boost detection and accurately measure the influence scope of defects. Finally, we summarize and develop a set of detection primitives for scaling our approach to handle various flaws that may lead to multiple risks in potential. Depending upon this method, we implemented SEAF, a Scalable, Efficient, and Application-independent Framework, and evaluated it on popular images of diverse applications in Docker Hub. The experiment result shows that SEAF can discover different security flaws fast. Compared to the state-of-the-art tool, Clair, SEAF is more efficient and can find significantly more types of defects. Libo Chen 0001, Yihang Xia, Zhenbang Ma, Ruijie Zhao 0001, Wenqi Sun, Zhi Xue |
J. Inf. Secur. Appl. | 4 |
| 2022 | Online Intrusion Detection for Internet of Things Systems With Full Bayesian Possibilistic Clustering and Ensembled Fuzzy ClassifiersabstractThe pervasive deployment of the Internet of Things (IoT) has significantly facilitated manufacturing and living. The diversity and continual updates of IoT systems make their security a crucial challenge, among which the detection of malicious network traffic turns out to be the most common yet destructive threat. Despite the efforts on feature engineering and classification backend designing, established intrusion detection systems sometimes lack robustness and are inflexible against the shift of the traffic distribution. To deal with these disadvantages, we design a fuzzy system for the online defense of IoT. Our framework incorporates a full Bayesian possibilistic clustering module for feature processing and an ensemble module motivated by reinforcement learning and adaptive boosting that dynamically fits the streaming data. The proposed clustering module overcomes the issue of determining the number of clusters and can dynamically identify new patterns. The classifier backend combines a collection of fuzzy decision trees that provide readable decision boundaries. The ensembled classifiers can accommodate the drift of data distribution to optimize the long-time performance. Our proposal is tested on settings including one dataset collected from real IoT systems and is compared to numerous competitors. Experimental results verified the advantage of our system regarding accuracy and stability. Fangqi Li 0001, Ruijie Zhao 0001, Shi-Lin Wang, Libo Chen 0001, Alan Wee-Chung Liew, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | An Efficient and Lightweight Approach for Intrusion Detection based on Knowledge DistillationabstractNetwork intrusion detection (NID) is an important cyber security scheme to identify attacks in network traffic. Recent years, a large amount of studies try to improve the accuracy of the NID by kinds of deep learning approaches. However, these models always require a lot of calculation and space, which constitutes a major hurdle to practical implementation of DL-based models. Thus, lightweight model is imperative, but there are very few applications of DL-based lightweight algorithms in NID models. In this paper, we propose a lightweight knowledge distillation (LKD) model for NID using the idea of knowledge distillation and separable convolution. To the best of our knowledge, it is the first system to use the knowledge distillation approach for NID. The experiment results show that the accuracy of the proposed approach reaches 91.46% and 94.30% on the KDD-CUP99 and UNSW-NB15 datasets respectively. The performance of our model is superior to some approaches based on deep neural network or some machine learning methods. Moreover, both the computational cost and model size of our model are reduced by about 99% compared to the original model. Ruijie Zhao 0001, Yong Shi 0009, Zhi Xue |
ICC | 1 |
| 2021 | Flow Transformer: A Novel Anonymity Network Traffic Classifier with Attention MechanismabstractSupervising anonymity network is a critical issue in the field of network security, and traditional traffic analysis methods cannot cope with complex anonymity traffic. In recent years, the traffic analysis method based on deep learning has achieved good performance. However, most of the existing studies do not consider the temporal-spatial correlation of the traffic, and only use a single flow for classification. A few works take continuous flows as flow sequence for traffic classification, but they do not distinguish the different importance of each flow. To tackle this issue, we propose a novel flow-based traffic classifier called FLOW TRANSFORMER to classify anonymity network traffic. FLOW TRANSFORMER uses multi-head attention mechanism to set higher weights for important flows, and extracts flow sequence features according to the importance weights. Besides, the RF-based feature selection method is designed to select the optimal feature combination, which can effectively avoid the insignificant features from reducing the performance and efficiency of the classifier. Experimental results on two real-world traffic datasets demonstrate that the proposed method outperforms state-of-the-art methods with a large margin. Ruijie Zhao 0001, Yiteng Huang, Xianwen Deng, Zhi Xue, Jiabin Li, Zijing Huang |
MSN | 1 |
| 2021 | A Semi-supervised Deep Learning-Based Solver for Breaking Text-Based CAPTCHAsabstractText-based CAPTCHAs are still the most widely used CAPTCHA mode. Many researchers have proposed attack methods to break them. In previous attacks, segmentation-based methods require at least three steps: preprocessing, segmentation, and recognition, which means that different modes of CAPTCHA require various preprocessing and segmentation algorithms. In recent years, a series of deep learning (DL) models have been designed for cracking text-based CAPTCHAs. However, these methods require annotating numerous images, which are time-consuming and labor-intensive. In this paper, we propose a semi-supervised DL-based solver for breaking text-based CAPTCHAs, which can use a small number of labeled CAPTCHAs to achieve a high-performance attack model. The CNN module and the attention-based Seq2Seq module are two key components for effective feature extraction and character recognition. The experimental results show that our solver successfully attacked 9 types of most popular text-based CAPTCHAs, and the attack success rate is better than the four latest attack models. In addition, our model does not perform any data preprocessing and has a fast attack speed, making it more suitable for real-time attacks. The code and dataset are available on the github. Xianwen Deng, Ruijie Zhao 0001, Zhi Xue, Libo Chen 0001 |
TrustCom | 2 |
| 2021 | A Novel Approach based on Lightweight Deep Neural Network for Network Intrusion DetectionabstractWith the ubiquitous network applications and the continuous development of network attack technology, all social circles have paid close attention to the cyberspace security. Intrusion detection systems (IDS) plays a very important role in ensuring computer and communication systems security. Recently, deep learning has achieved a great success in the field of intrusion detection. However, the high computational complexity poses a major hurdle for the practical deployment of DL-based models. In this paper, we propose a novel approach based on a lightweight deep neural network (LNN) for IDS. We design a lightweight unit that can fully extract data features while reducing the computational burden by expanding and compressing feature maps. In addition, we use inverse residual structure and channel shuffle operation to achieve more effective training. Experiment results show that our proposed model for intrusion detection not only reduces the computational cost by 61.99% and the model size by 58.84%, but also achieves satisfactory accuracy and detection rate. Ruijie Zhao 0001, Zhaojie Li, Zhi Xue, Tomoaki Ohtsuki, Guan Gui 0001 |
WCNC | 1 |