Zelin Cui

dblp:240/2586 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-7229-3231ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 No train, no pain: a training-free few-shot traffic classifier based on LLMs
abstract
Abstract Encrypted web traffic and evolving Internet technologies pose an increasing challenge to network traffic analysis. However, existing traffic classification methods, though effective, require large labeled datasets and complex training. This makes sustaining them prohibitively expensive and difficult in real-world scenarios. To narrow this gap, we propose a novel training-free few-shot network traffic classification framework based on large language models (LLMs). By integrating meta-learning with LLMs, it reduces reliance on labeled data, eliminates task-specific training, and improves performance. Specifically, we first apply an efficient feature extraction method to extract features from traffic flows. We then design meta-tasks that combine task descriptions with textualized features to produce natural language meta-task formulations. Building on these meta-tasks, the LLM performs reasoning to carry out traffic classification. Finally, to mitigate hallucination in the LLM outputs, we exploit the temporal characteristics of network traffic and aggregate predictions over samples within a defined time window. Extensive experiments on three widely-used encrypted traffic datasets demonstrate that our proposed framework outperforms the state-of-the-art methods, achieving an average absolute improvement in F1 score of 9.75, 9.82, and 12.06 percentage points on the three datasets, respectively.
Xingmao Guan, Xueying Han, Jinlai Huang, Tao Wang 0029, Zelin Cui, Zhigang Lu 0002, Baoxu Liu
Cybersecur.7
2025 Cross Page Recognition Methods for Encrypted Web Application Fingerprinting
abstract
The widespread implementation of the HTTPS protocol has greatly bolstered user privacy and data security. However, the widespread use of the HTTPS protocol has also provided criminals with a cloak to disseminate harmful content through websites, thereby undermining the integrity of the online environment. Web page fingerprinting has emerged as a highly popular method for web application identification. Yet, due to the frequent updates of web applications, existing methods struggle with accuracy issues. To tackle these challenges, this paper introduces a novel approach called CrossWP, which leverages cross-web page fingerprinting to enhance the security of the network environment. CrossWP aims at classifying web applications, which novelly constructs the cross web pages behavior sequences on handshake, request and response sequence. CrossWP uses the transformer model based on multi-sequence fusion to thoroughly learn and integrate these unique spatio-temporal sequence characteristics. This model can also capture the internal similarity of behavior sequence, and achieving high accuracy. The effectiveness of CrossWP is validated through closed-world and open-world evaluations, which involve identifying and classifying news websites, social websites, online video websites and e-commerce websites. The results indicate that CrossWP outperforms existing algorithms in terms of both robustness and accuracy.
Zelin Cui, Pu Dong, Dongxu Han, Bo Jiang 0013, Zhigang Lu 0002, Huamin Feng
CSCWD1
2025 Dynamic Behavior-Based Detection Techniques for Encrypted Variant Webshells
abstract
Webshell, as a common type of malicious script, is frequently utilized by cyber attackers who execute unauthorized commands on the victim's server to carry out attacks. Strengthening research on Webshell detection techniques is crucial for building a robust cybersecurity defense. Despite significant progress in the field of Webshell detection, these techniques still face numerous challenges. Firstly, the continuous evolution of attacker techniques has enhanced the adversarial capabilities of Webshell, including rapid updates to version variants, as well as the use of advanced obfuscation techniques. Secondly, the use of HTTPS has grown dramatically from 40% in 2014 to 98% in 2023, rendering techniques based on plaintext rules ineffective for detecting Webshell. These technological updates make it difficult for traditional detection methods to effectively identify and defend against Webshell variants based on the HTTPS protocol. To address this problems, we propose a novel dynamic behavior-based detection techniques called DBBDdetect, aiming at detecting encrypted variant Webshell in order to protect critical infrastructure. DBBDdetect delves into the interaction process between the Webshell and the server, extracting three types of feature information. It utilizes CNNs to obtain vector features of the traffic payload and concat statistical features and similar sequential byte behavior features. Then, it uses DBSCAN clustering model to analyze behavioral similarities to detect variant Webshell attacks. This method captures the intrinsic similarities in behavior, and experiments have shown that it achieves a high level of accuracy.
Zelin Cui, Pu Dong, Mengchuan Shang, Bo Jiang 0013, Zhigang Lu 0002, Huamin Feng
CSCWD1
2025 CotexFinger: Enhancing IoT Device Identification with Context-Packet Fingerprinting and Lightweight Mamba
abstract
With the growth of the Internet of Things (IoT) market, IoT devices have become major targets for attacks. For network administrators, effectively addressing these attacks requires accurately identifying IoT devices. Due to the rich communication protocols and complex network environments of IoT, current fingerprinting methods lack robustness. This paper proposes a fingerprinting method based on Context-Packet, which effectively reduces the impact of traffic noise and fingerprint redundancy by focusing only on session packets that carry payload data. Additionally, a lightweight Efficient VMamba model is trained to extract fingerprints from the packets. We design a method that utilizes the Context-Packet length sequences to classify traffic samples from the same device category, while retaining rare traffic samples during the random sampling process. Experiments show that our method outperforms existing approaches in classification performance, improving the F1 score to 97.08% ( 0.60% ↑ ), 98.07% ( 2.7% ↑ ), 95.55% ( 1.27% ↓ ), and 92.22% ( 3.71% ↑ ) in the AA, II, AI, and IA experiments, respectively. Moreover, it offers faster inference speed and lower resource overhead.
Zelin Cui, Bo Jiang 0013, Zhigang Lu 0002, Huamin Feng
IJCNN3
2025 TCP-Awareness Augmented Robust TLS Traffic Classification: A Hybrid Deep Learning Approach
Yewa Li, Yitan Huang, Bo Jiang 0013, Zhigang Lu 0002, Zelin Cui
TrustCom5
2025 Towards effective black-box attacks on DoH tunnel detection systems
Linghao Li, Wei Qiao 0005, Zelin Cui, Susu Cui, Bo Jiang 0013, Zhigang Lu 0002
Comput. Networks5
2024 HBGraph: a Host Behavior Graph Model for C&C Traffic Detection
abstract
The command and control (C&C) mechanism is the key to realizing many network attack activities. Most available C&C traffic detection methods rely on machine learning. However, the methods often detect specific attacks and struggle to adapt to the complex needs of real-world network traffic environments due to high training costs and limited transferability. To address the problems, we propose a C&C communication traffic detection model called HBGraph, a host behavior graph model based on directed packet payload length sequences. The model can preprocess the original traffic datasets, extract directed payload length sequences, and then integrate them into weighted directed graphs that represent different host communication behaviors. We also propose a method of merging graphs with the same label. In HBGraph, we measure the similarity between the unknown instance and the signature with both node and edge similarity scores. Finally, the model can predict whether the unknown test instance belongs to the C&C communication traffic. After experimental evaluation, we prove that our model has good performance, strong generalizability, and detection ability for C&C communication behaviors.
Yitan Huang, Zelin Cui, Bo Jiang 0013, Zhigang Lu 0002
CSCWD3
2024 Multi-language Webshell Detection based on Abstract Syntax Tree and TreeLSTM
abstract
Webshell is a command execution environment existing in web containers, which is used by attackers to remotely control servers and illegally access website resources. Accurately detecting Webshells is of great significance for maintaining web security. Current research faces several challenges. On the one hand, in order to evade detection, Webshells use a large amount of obfuscation, and existing research methods often use source code or opcode, which cannot fully utilize the semantic and syntactic information of Webshell code. On the other hand, Webshells can be constructed using any web application programming language, while most existing methods only detect one or a few types of Webshells. This paper proposes a novel approach called WS-Tree, which effectively utilizes the semantics and syntax of Webshells by using abstract syntax tree as input features. The TreeLSTM model is used as an encoder to handle node relationships in the syntax tree, thereby achieving the detection of obfuscated and multi-language Webshells. We also propose a new dataset of Webshells containing obfuscated and non-obfuscated to prevent dataset leakage. Extensive experiments demonstrate that our proposed model performs better than the state-of-theart baselines under different webshell programming languages and improves model generalizability.
Mengchuan Shang, Xueying Han, Changzhi Zhao, Zelin Cui, Dan Du, Bo Jiang 0013
CSCWD4
2022 IV-IDM: Reliable Intrusion Detection Method based on Involution and Voting
abstract
Intrusion detection is critical in the area of cyberspace security. Deep learning methods, especially CNN, have been widely used in intrusion detection in recent years. Network traffic is usually converted into images for processing. However, images converted from network traffic do not have multi-channel features like real-world pictures and have explicit long-distance dependencies between pixels. These characteristics will cause weak performance and poor explanation, making images converted from network traffic unsuitable to be processed by CNN. Besides, most works only consider the first few packets (named head packets) in a flow, which contains the information about connection establishment and interaction between two parts. However, the last few packets (named tail packets) are omitted, resulting in the loss of information about disconnection. To handle the above problems, we propose a reliable intrusion detection model called IV-IDM. Instead of convolution, IV-IDM uses a new structure, involution. Involution has the properties of spatial-specific and channel-agnostic and is more suitable for intrusion detection tasks than convolution. We also propose I-Res, which is constructed based on involution and is used as the base classifier of IV-IDM. We use head and tail packets of a flow as the inputs to two I-Res respectively to learn richer information and employ a voting algorithm to integrate the results of these two parts to promote the robustness of the model. Finally, IV-IDM is evaluated by the ISCX-IDS-2012 and the CIC-IDS-2017 datasets. The experimental results demonstrate that IV-IDM outperforms the state-of-the-art models and is qualified for intrusion detection.
Xueying Han, Pu Dong, Bo Jiang 0013, Zhigang Lu 0002, Zelin Cui
ICC6
2022 PUMD: a PU learning-based malicious domain detection framework
abstract
Abstract Domain name system (DNS), as one of the most critical internet infrastructure, has been abused by various cyber attacks. Current malicious domain detection capabilities are limited by insufficient credible label information, severe class imbalance, and incompact distribution of domain samples in different malicious activities. This paper proposes a malicious domain detection framework named PUMD, which innovatively introduces Positive and Unlabeled (PU) learning solution to solve the problem of insufficient label information, adopts customized sample weight to improve the impact of class imbalance, and effectively constructs evidence features based on resource overlapping to reduce the intra-class distance of malicious samples. Besides, a feature selection strategy based on permutation importance and binning is proposed to screen the most informative detection features. Finally, we conduct experiments on the open source real DNS traffic dataset provided by QI-ANXIN Technology Group to evaluate the PUMD framework’s ability to capture potential command and control (C&C) domains for malicious activities. The experimental results prove that PUMD can achieve the best detection performance under different label frequencies and class imbalance ratios.
Zhaoshan Fan, Qing Wang 0041, Haoran Jiao, Zelin Cui
Cybersecur.5
2021 MBTree: Detecting Encryption RATs Communication Using Malicious Behavior Tree
abstract
Network trace signature matching is one reliable approach to detect active Remote Control Trojan, (RAT). Compared to statistical-based detection of malicious network traces in the face of known RATs, the signature-based method can achieve more stable performance and thus more reliability. However, with the development of encrypted technologies and disguise tricks, current methods suffer inaccurate signature descriptions and inflexible matching mechanisms. In this paper, we propose to tackle above problems by presenting MBTree, an approach to detect encryption RATs Command and Control (C&C) communication based on host-level network trace behavior. MBTree first models the RAT network behaviors as the malicious set by automatically building the multiple level tree, MLTree from distinctive network traces of each sample. Then, MBTree employs a detection algorithm to detect malicious network traces that are similar to any MLTrees in the malicious set. To illustrate the effectiveness of our proposed method, we adopt theoretical analysis of MBTree from the probability perspective. In addition, we have implemented MBTree to evaluate it on five datasets which are reorganized in a sophisticated manner for comprehensive assessment. The experimental results demonstrate the accurate and robust of MBTree, especially in the face of new emerging benign applications.
Cong Dong, Zhigang Lu 0002, Zelin Cui, Baoxu Liu, Kai Chen 0012
IEEE Trans. Inf. Forensics Secur.3