Qilei Yin

dblp:230/1810 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 6 since 2021Security and privacy · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AddrProbe: An Internet-Wide Active IPv6 Address Probing System With Limited Seeds
abstract
With the large-scale deployment of IPv6, it is becoming more and more important to probe active IPv6 addresses on the global Internet. However, the vast address space and the random distribution of active addresses make the probing process full of challenges, especially for the probing of IPv6 prefixes without seed addresses. Furthermore, the widespread existence of IPv6 aliased prefixes also causes significant trouble for probing. In this paper, we presentAddrProbe, an active IPv6 address probing system, which dynamically probes all global routing prefixes based on learned fine-grained address patterns from limited seed addresses and quickly detects aliased prefixes during probing. The evaluation results show thatAddrProbeachieves a hit rate of 23%-45% with all routing prefixes announced by the BGP system, which is 6.6-13× that of current state-of-the-art approaches (no more than 4%). Moreover, we find 1.2×1033aliased addresses characterized by the detected aliased prefixes, covering 6,412 routing prefixes, which is a 107× and 5.9× improvement over existing methods, respectively. Finally, an IPv6 Hitlist is constructed based on the long-term probing results, which contains 562M addresses covering 190K routing prefixes and 29K ASes. These widely distributed addresses are meaningful for analyzing IPv6 address assignments and some other IPv6 measurement activities.
Daguo Cheng, Lin He 0004, Qilei Yin, Guangxing Han, Boran Jin, Ying Liu 0024, Guanglei Song, Jinlong E, Tiankai Yang 0001, Jiahai Yang 0001
IEEE Trans. Netw.3
2026 Toward Robust Multi-Tab Website Fingerprinting
abstract
Website fingerprinting enables an eavesdropper to determine which websites a user is visiting over an encrypted connection. State-of-the-art website fingerprinting (WF) attacks have demonstrated effectiveness even against Tor-protected network traffic. However, existing WF attacks have critical limitations on accurately identifying websites in multi-tab browsing sessions, where the holistic pattern of individual websites is no longer preserved, and the number of tabs opened by a client is unknown a priori. In this paper, we propose ARES, a novel WF framework natively designed for multi-tab WF attacks. ARES formulates the multi-tab attack as a multi-label classification problem and solves it using the novel Transformer-based models. Specifically, ARES extracts local patterns based on multi-level traffic aggregation features and utilizes the improved self-attention mechanism to analyze the correlations between these local patterns, effectively identifying websites. We implement a prototype of ARES and extensively evaluate its effectiveness using our large-scale datasets collected over multiple months. The experimental results illustrate that ARES achieves optimal performance in several realistic scenarios. Further, ARES remains robust even against various WF defenses.
Xinhao Deng 0001, Qilei Yin, Zhuotao Liu, Qi Li 0002, Mingwei Xu 0001, Ke Xu 0002
IEEE Trans. Netw.3
2026 Do Not Fall Into the Trap: Efficiently Discovering IPv6 Fully Responsive Prefixes in the Wild
Lin He 0004, Chentian Wei, Daguo Cheng, Qilei Yin, Boran Jin, Zhaoan Wang, Xiaoteng Pan, Sixu Zhou, Ying Liu 0024, Shenglin Zhang, Fuchao Tan, Wenmao Liu
IEEE Trans. Netw.4
2026 Toward Robust Detection of Malicious Encrypted Traffic Using Only Low-Quality Training Data
abstract
Machine learning (ML) is promising in accurately detecting malicious flows in encrypted network traffic; however, it is challenging to collect a training dataset that contains a sufficient amount of encrypted malicious data with correct labels. When ML models are trained with low-quality training data, they suffer degraded performance. In this paper, we aim to address a real-world low-quality training dataset problem, namely, detecting encrypted malicious traffic generated by continuously evolving malware. We develop RAPIER+ that fully utilizes different distributions of normal and malicious traffic data in the feature space, where normal data is tightly distributed in a certain area, and the malicious data is scattered over the entire feature space to augment training data for model training. RAPIER+ includes two pre-processing modules to convert traffic into feature vectors and correct label noises. We evaluate our system on two public datasets and one combined dataset. With 1000 samples and 45% noise from each dataset, our system achieves the F1 scores of 0.78, 0.84, and 0.87, respectively, achieving average improvements of 358.5%, 314.0%, and 221.1% over the existing methods, respectively. Furthermore, we evaluate RAPIER+ with a real-world dataset obtained from a security enterprise. RAPIER+ effectively achieves encrypted malicious traffic detection with the best F1 score of 0.81 and improves the F1 score of existing methods by an average of 288.7%.
Yuqi Qing, Qilei Yin, Xinhao Deng 0001, Zhuotao Liu, Kun Sun 0001, Ke Xu 0002, Jia Zhang 0004, Qi Li 0002
IEEE Trans. Netw.2
2025 Training Robust Classifiers for Classifying Encrypted Traffic under Dynamic Network Conditions
abstract
Most existing DL-based encrypted traffic classification methods suffer performance degradation in real-world deployments due to dynamic network conditions, e.g., network environment changes and traffic obfuscation. Dynamic network conditions cause encrypted traffic to exhibit distinct feature patterns during training and testing phases. To address this issue, we propose MetaTraffic, a novel and general DL training framework built upon meta-learning that enhances the performance of supervised DL models designed for encrypted traffic classification against dynamic network conditions. Our key observation is that the traffic of the same network behaviors share the same semantic features even under different network conditions, which can be considered as stable feature representations. Therefore, MetaTraffic helps DL models learn stable feature representations by minimizing the discrepancies in how the models represent traffic features under different network conditions, thereby achieving robust classification under dynamic network conditions. We implement MetaTraffic based on meta-learning with three innovative facilitate modules to enhance its performance. We evaluate MetaTraffic using three public datasets and three new large-scale encrypted traffic datasets that cover multiple types of network conditions. Experimental results show that, under dynamic multiple types of network conditions, our framework improves the accuracy of DL models by 8.94% and the F1-Macro score by 12.55%, while existing robust training methods decrease the accuracy by 28.85% and the F1-Macro score by 33.52%.
Yuqi Qing, Qilei Yin, Xinhao Deng 0001, Xiaoli Zhang 0003, Zhuotao Liu, Kun Sun 0001, Ke Xu 0002, Qi Li 0002
CCS2
2025 Detecting and Adapting to Stealthy Label-Inversion Drifts via Conditional Distribution Inference
abstract
Deep learning (DL) based malicious traffic detectors have been widely developed to detect diverse network attacks, yet they are suffering from significant performance degradation due to concept drift. Existing anti-concept drift arts focus on combating the drifting traffic whose features significantly diverge from training traffic. However, they neglect a stealthy yet common situation where the testing traffic has similar features to the training traffic but with opposite ground truth labels. As a result, the DL-based detectors would always make incorrect predictions for the stealthy drifting traffic, insufficient to perform long-term real-world intrusion detection. In this paper, we propose Chameleon, a novel active learning framework that combats stealthy drifting traffic by inferring the conditional distribution of the testing traffic with small manual labeling overhead. Specifically, Chameleon measures the fine-grained correlations between the high-dimensional and heterogeneous testing traffic and selects a small number of highly representative testing traffic samples for manual labeling, to accurately infer other testing samples’ labels. With the inferred labels, Chameleon checks the conditional distribution shift from the training to testing traffic to detect concept drift and incrementally trains the DL-based detectors to make them effectively adapt to the shifted distribution. Extensive experiments with six supervised and unsupervised DL-based detectors on three public and four synthetic datasets show that, under stealthy drifting traffic, Chameleon improves the AUT of the DL-based detectors by a range of $18.53 \%$ to $23.89 \%$, while the improvement of SOTA baselines is only between $0.06 \%$ and $1.86 \%$.
Xiaoli Zhang 0003, Qilei Yin, Jianrong Zhang, Ke Xu 0002, Qi Li 0002, Xu-Cheng Yin
RAID3
2024 Luori: Active Probing and Evaluation of Internet-Wide IPv6 Fully Responsive Prefixes
abstract
With the large-scale deployment and application of IPv6, IPv6 network measurements will become increasingly important. However, a special type of IPv6 prefix called Fully Responsive Prefix (FRP) is having a significant impact on IPv6 measurement campaigns, which is defined as all addresses under a prefix responding to scans. Obviously, there cannot be a real responder behind each of these addresses. To reveal the current status and impact of Internet-wide IPv6 FRPs, we propose for the first time an active probing method for Internet-wide IPv6 FRPs, Luori, which transforms the active probing process under IPv6 huge prefix space (potential range of prefix presence) into a dynamic search process in a tree based on reinforcement learning, achieving efficient probing of arbitrary routing prefixes. The evaluation results show that Luori found 31.7K largest FRPs in a single Internet-wide probing with 11 M budget, covering$1.5 \times 10^{30}$address space, which is$10^{6} \times$that of existing methods. More importantly, after six months of Internet-wide probing, we have found 516 K largest FRPs, which covers$1.3 \times 10^{33}$address space and 795 ASes, making it the largest publicly known FRP list. Based on this list, we screen out$20 \%$of the addresses covered by FRPs from a well-known IPv6 active address dataset. Furthermore, we further analyze and find that the distribution of these FRPs is extensive and their implementation methods are diverse, which can provide beneficial references for the practical application of FRPs. We also make this list publicly available and maintain it long-term for use and study by relevant researchers.
Daguo Cheng, Lin He 0004, Chentian Wei, Qilei Yin, Boran Jin, Zhaoan Wang, Xiaoteng Pan, Sixu Zhou, Ying Liu 0024, Shenglin Zhang, Fuchao Tan, Wenmao Liu
ICNP4
2024 SHTree: A Structural Encrypted Traffic Fingerprint Generation Method for Multiple Classification Tasks
abstract
In recent years, encrypted traffic classification has been found widespread applications in the field of cybersecurity. Its main challenge lies in accurately represent traffic when features are obscured due to encryption. To address this, researchers utilize fingerprint construction methods based on statistical information or employ Deep Learning (DL) for traffic representation. However, in previous methods of feature selection, flat key-value pair features, or raw packet bytes are often used, ignoring the structured information embedded in packets and flows. Therefore, We propose a novel structured encrypted traffic fingerprint generation method called SHTree. It constructs traffic fingerprints using a set of tree-based structures to represent traffic, encapsulating structural features from the traffic, enhancing the representation of traffic. This enables it to adapt to various classification tasks through general feature selection. The experiments demonstrate that our method achieves comparable accuracy to state-of-the-art Large Language Models (LLMs), with an F1 score higher by 0.5% on specific tasks. Meanwhile, it outperforms by three orders of magnitude in classification speed. In unsupervised abnormal detection tasks, the True Positive Rate (TPR) exceeds 99%, while maintaining a False Positive Rate (FPR) of 0.5%.
Minghao Ma, Zhixin Shi, Qilei Yin, Yangyang Zong
ISCC3
2024 Low-Quality Training Data Only? A Robust Framework for Detecting Encrypted Malicious Network Traffic
Yuqi Qing, Qilei Yin, Xinhao Deng 0001, Zhuotao Liu, Kun Sun 0001, Ke Xu 0002, Jia Zhang 0004, Qi Li 0002
NDSS2
2024 Learning with Semantics: Towards a Semantics-Aware Routing Anomaly Detection System
Qilei Yin, Qi Li 0002, Zhuotao Liu, Ke Xu 0002, Mingwei Xu 0001
USENIX Security Symposium2
2024 Multi-objective niching quantum genetic algorithm-based optimization method for pneumatic hammer structure
Jine Cao, Pinlu Cao, Chengda Wen, Hongyu Cao, Qilei Yin
Expert Syst. Appl.6
2023 Robust Multi-tab Website Fingerprinting Attacks in the Wild
abstract
Website fingerprinting enables an eavesdropper to determine which websites a user is visiting over an encrypted connection. State-of-the-art website fingerprinting (WF) attacks have demonstrated effectiveness even against Tor-protected network traffic. However, existing WF attacks have critical limitations on accurately identifying websites in multi-tab browsing sessions, where the holistic pattern of individual websites is no longer preserved, and the number of tabs opened by a client is unknown a priori. In this paper, we propose ARES, a novel WF framework natively designed for multi-tab WF attacks. ARES formulates the multi-tab attack as a multi-label classification problem and solves it using a multi-classifier framework. Each classifier, designed based on a novel transformer model, identifies a specific website using its local patterns extracted from multiple traffic segments. We implement a prototype of ARES and extensively evaluate its effectiveness using our large-scale dataset collected over multiple months (by far the largest multi-tab WF dataset studied in academic papers.) The experimental results illustrate that ARES effectively achieves the multi-tab WF attack with the best F1-score of 0.907. Further, ARES remains robust even against various WF defenses.
Xinhao Deng 0001, Qilei Yin, Zhuotao Liu, Qi Li 0002, Mingwei Xu 0001, Ke Xu 0002
SP2
2022 Improving the Semantic Consistency of Textual Adversarial Attacks via Prompt
abstract
Adversarial examples can expose the vulnerabilities of neural networks. State-of-the-art textual adversarial attacks have demonstrated their effectiveness in triggering errors in the output of natural language processing models. However, these attacks are limited to ensuring the semantic consistency between the adversarial example and the original input, increasing the possibility of the attacks being detected by human judges. In this paper, we propose a novel textual adversarial attack, Prompt-Attack, which aims to generate the adversarial examples having consistent semantics with the original input. Specifically, Prompt-Attack enhances the input's semantics with the prompts that represent the semantics of the different target segments extracted from the original input, to predict the substitutions having consistent semantics with the target segments. Then, it crafts the adversarial examples by replacing the important segments with their substitutions that can most affect the victim model's output. Besides, Prompt-Attack proposes a span-level segment identification strategy to extract more target segments from the input and a novel masking strategy to ensure the grammatical correctness of the generated adversarial examples. Extensive experiments on public datasets illustrate that Prompt-Attack significantly improves the semantic consistency score of the baseline attacks by an average of 48%. Further, Prompt-Attack achieves the best attack success rate of 0.906, showing an average improvement of 40% to the baselines. Moreover, the experimental results demonstrate that Prompt-Attack can achieve good performance in attacking different language models and Prompt-Attack is not sensitive to different settings.
Qilei Yin, Zhixin Shi, Yuru Ma
IJCNN2
2022 Encrypted Malware Traffic Detection via Graph-based Network Analysis
abstract
Malicious activities on the Internet continue to grow in volume and damage, posing a serious risk to society. Malware with remote control capabilities is considered one of the most threatening malicious activities, as it can enable arbitrary types of cyber-attacks. As a countermeasure, many malware detection methods are proposed to identify malicious behaviours based on traffic characteristics. However, the emerging encryption and evasion techniques pose substantial barriers to the full exploitation of network information. This significantly impairs the effectiveness of existing malware detection methods relying on a singular type of characteristics. In this paper, we propose ST-Graph to resolve this issue. In addition to traditional stream attributes, ST-Graph explores spatial and temporal characteristics of network behaviours based on a graph representation learning algorithm and integrates all available information to boost the detection decision. To illustrate the effectiveness of ST-Graph, we evaluate it on two datasets. Experimental results demonstrate that ST-Graph outperforms state-of-the-art malware detection systems and also shows good performance in efficiency, generalizability, and robustness. Specifically, it achieves over 99% precision and recall, and its False Positive Rate is even two orders of magnitude lower than (nearly 0.02 times) that of baseline models. Meanwhile, the deployment of ST-Graph in two real network scenarios for around one year shows an outstanding efficiency with only 160 seconds time cost for 5-minute traffic in 1.7 Gbps bandwidth.
Zhuoqun Fu, Mingxuan Liu 0006, Jia Zhang 0004, Yuan Zou, Qilei Yin, Qi Li 0002, Hai-Xin Duan
RAID6
2022 Unsupervised Contextual Anomaly Detection for Database Systems
abstract
Abnormal data access operations in database systems always hap-pen, which are typically incurred by misoperations or attacks, though these systems are enforced with strict access control policies. However, prior arts only focus on detecting abnormal data accesses by utilizing known attack patterns or identifying behaviors significantly deviated from normal behaviors. They cannot capture stealthy abnormal data access operations that are similar to normal ones. In this paper, we propose a novel unsupervised anomaly detection system UCAD, which aims to detect abnormal data access operations, by comparing operation's semantics with their contextual intent. However, it is non-trivial to obtain accurate semantics of operations for intent analysis because (i) the same operation may exhibit diverse semantics under different operation contexts and (ii) different operation sequences could have identical semantics due to heterogeneous user access patterns. To address this issue, we develop a new transformer model called Trans-DAS for UCAD. Trans-DAS learns the semantics of individual operations by utilizing the attention mechanism that analyzes the relevance between any pair of operations in sequence, and captures the contextual intent of operations inferred from the contexts. Specifically, Trans-DAS utilizes a particular embedding layer to embed the semantics of individual operations without the operation order information and a masking mechanism that allows Trans-DAS to learn the semantics according to the bidirectional contexts. Also, we define a new training objective for Trans-DAS to enlarge the difference among the embedded semantics. Furthermore, in order to effectively utilize Trans-DAS for detection, we develop two modules in UCAD, i.e., a data preprocessing module that allows Trans-DAS to accurately learn the normal semantic information by removing noisy data, and an anomaly detection module that learns the semantic information for intent comparison. We evaluate the performance of UCAD on real-world data traces under different settings (e.g., varied parameters and hybrid datasets). The results demonstrate that UCAD achieves the average F1-score of 0.94 in two scenarios, which significantly outperform baselines, and shows robustness to hybrid data and good transferability to different tasks.
Sainan Li, Qilei Yin, Guoliang Li 0001, Qi Li 0002, Zhuotao Liu
SIGMOD Conference2
2022 An Automated Multi-Tab Website Fingerprinting Attack
abstract
In Website Fingerprinting (WF) attack, a local passive eavesdropper utilizes network flow information to identify which web pages a user is browsing. Previous researchers have demonstrated the feasibility and effectiveness of WF attacks under a strong Single Page Assumption: the network flow extracted by the adversary belongs to a single web page. In reality, the assumption may not hold because users tend to open multiple tabs simultaneously (or within a short period of time) so that their network traffic is mixed. In this article, we propose an automated multi-tab Website Fingerprinting attack that is able to accurately classify websites regardless of the number of simultaneously opened pages. Our design is powered by two innovative designs. First, we develop a split point classification method to dynamically identify the split point between the first page and its subsequent pages. As a result, the network traffic before the split point is solely generated for the first page. Then, we propose a new chunk-based WF classifier to infer the websites based on the initial chunk of clean traffic. For both classifiers, we apply automated feature selection to select a concise yet representative feature set. We implement a prototype of our design and perform extensive evaluations using SSH and Tor-based datasets to demonstrate the effectiveness of both our system components individually and the integrated system as a whole.
Qilei Yin, Zhuotao Liu, Qi Li 0002, Tao Wang 0012, Qian Wang 0002, Chao Shen 0001, Yixiao Xu
IEEE Trans. Dependable Secur. Comput.1
2019 A Behavior-Based Method for Distinguishing the Type of C&C Channel
Qilei Yin, Zhixin Shi, Guokun Xu, Xiaoyu Kang
ICA3PP (1)2
2019 A New C&C Channel Detection Framework Using Heuristic Rule and Transfer Learning
abstract
A great many of botnet detection methods focus on recognizing the significant C&C channels. Most of them require a C&C training set to build a behavior detection model. However, when lacking such training set for new or unknown botnets, these methods may become inefficient or even invalid.To overcome it, we propose a new general framework for C&C channel detection. It neither needs us to know the families of bots or prepare a training set nor requires deploying malicious activity monitors. Also, it is capable of mining useful knowledge from the historical dataset to boost its detection performance. In our framework, we put forward a clustering method and several heuristic rules to aggregate and label partial C&C traffic, a sample selection function to mine useful historical knowledge and a transfer learning based model to find other C&C channels. We evaluated our framework on two datasets and achieved the best C&C F-measure of about 0.886 and 0.960 respectively. Moreover, the comparison result further indicates its performance advantage and better behavior learning ability.
Qilei Yin, Zhixin Shi, Meimei Li
IPCCC2
2018 Comprehensive Behavior Profiling Model for Malware Classification
abstract
In view of the great threat posed by malware and the rapid growing trend about malware variants, it is necessary to determine the category of new samples accurately for further analysis and taking appropriate countermeasures. The network behavior based classification methods have become more popular now. However, the behavior profiling models they used usually only depict partial network behavior of samples or require specific traffic selection in advance, which may lead to adverse effects on categorizing advanced malware with complex activities. In this paper, to overcome the shortages of traditional models, we raise a comprehensive behavior model for profiling the behavior of malware network activities. And we also propose a corresponding malware classification method which can extract and compare the major behavior of samples. The experimental and comparison results not only demonstrate our method can categorize samples accurately in both criteria, but also prove the advantage of our profiling model to two other approaches in accuracy performance, especially under scenario based criteria.
Qilei Yin, Zhixin Shi, Meimei Li
ISCC2