EDBT 2026 Demo / reviewers in the wild / expert
Qingya Yang
dblp:276/7391
· DBLP profile ↗
11ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TA: A Chinese Adversarial Samples Generation Approach Based on Multi-Strategy PerturbationsabstractWhile adversarial attacks on English texts have been extensively investigated, research on Chinese language scenarios remains comparatively under-developed. To address this gap, we present TextAttack (TA), a novel Chinese adversarial sample generation approach that leverages the linguistic and pictographic characteristics of Chinese through multi-strategy perturbations. we propose a three-phase methodology to deal with the challenge systematically: The Key Token Extractor (KTE) employs confidence variation to perform segmented identification of critical sentences and pivotal tokens; The Weighted Multi-strategy Perturbation (WMP) module implements targeted weighted perturbations through a variety of strategies such as glyph-based substitution and synonym replacement; The Phonetic-Visual-Semantic (PVS) constrained iterative search mechanism optimizes perturbation selection while maintaining textual naturalness through multimodal constraint integration. Experiments show that TA performs well on six classification models across two datasets, reducing the classification accuracy of hotel review dataset and news topic dataset by over 44.1% and 72.2%, respectively, with a small cost and high transfer ability. Yangyang Ding, Gaopeng Gou, Qingya Yang |
CSCWD | 3 |
| 2025 | SwCC: A Swapped-Contrastive Clustering Learning for Few-shot Website Fingerprinting AttacksabstractWebsite fingerprinting (WF) attacks exploit distinctive traffic patterns to identify the specific web page a user visits over anonymized connections. While traditional WF attacks have achieved impressive results, they are typically evaluated in abundant labeled data settings and assume that website traffic features remain static. This assumption is often unrealistic in real-world scenarios. Recent methods either rely on deep learning, which still requires large amounts of labeled data, or employ self-supervised pre-training to ease this demand. However, they leave clustering information crucial for few-shot WF attacks underexplored, leaving ample room for performance gains. In this paper, we propose a novel self-supervised pre-training model for few-shot WF attacks, called Swapped Contrastive Clustering (SwCC). SwCC proposes a comprehensive and principled data augmentation scheme, combining Tor-tailored transformations with statistical procedures to foster robust and discriminative feature learning for WF attacks. SwCC further introduces an innovative dual-level contrastive learning framework that jointly leverages instance-level and prototype-based objectives, which can further perform latent-space clustering on extracted features in WF attacks. Our pre-trained model can be fine-tuned in few-shot learning scenarios and achieves state-of-the-art(SOTA) performance on few-shot WF attack tasks. Under a 5-shot learning setting in a closed-world scenario, our SwCC achieves up to 83.5% accuracy when the evaluation traces are collected from an environment unseen by the WF adversary, outperforming the SOTA methods. Gaopeng Gou, Wenqi Dong, Gang Xiong 0001, Zhen Li 0011, Qingya Yang |
TrustCom | 8 |
| 2024 | OSN Bots Traffic Transformer : MAE-Based Multimodal Social Bots Behavior Pattern MiningabstractIn recent years, online social networks (OSN) have rapidly gained popularity worldwide, becoming important platforms for information dissemination. Cyber manipulators use OSN bots to disseminate harmful information and manipulate public opinion, which can engage in cyber violence and conduct financial crimes. Therefore, it is crucial to propose an effective detection solution for OSN bots as a matter of urgency. Different OSN bots exhibit distinct behavioral patterns compared to regular users due to varying behavioral preferences. Analyzing network behavior patterns can reveal the fundamental rules and anomalies of OSN bots, providing support for effective detection in order to gather evidence of any illegal activities. Traditional social bot detection methods based on user profiles or social relationships pose risks of infringing on user privacy. Therefore, we propose a new detection framework for OSN bots——OBTT model, which demonstrates significant advantages in identifying bot traffic to OSN and discovering behavior patterns of different types of bots. OBTT adopts a multimodal approach, integrating graph embeddings from raw traffic with sequential features, while incorporating temporal information to explore the regularities in bot action sequences. Using large-scale unlabeled data, we pretrain a Masked Autoencoder (MAE) and fine-tune it with a small amount of labeled data to enhance the model capacity to detect various bot behavior patterns. Experiments conducted on our OSNBotTraffic5 dataset show that OBTT achieved an accuracy of 0.95, demonstrating excellent performance. Notably, this is the first time that different OSN bot behavior patterns have been identified in quasi-real time from the perspective of network traffic. Haonan Zhai, Ruiqi Liang, Zhen Li 0011, Bingxu Wang, Qingya Yang |
TrustCom | 7 |
| 2023 | Identifying DoH Tunnel Traffic Using Core Feathers and Machine Learning MethodabstractDNS protocol is a plaintext domain name resolution protocol, which has the risk of privacy disclosure. DNS over HTTPS (DOH) protocol is designed to encrypt DNS traffic, which solves the privacy problem. However, many network attackers use the DOH tunnel for malicious transmission. From the passive traffic, there is no obvious difference between normal DOH traffic and DOH tunnel traffic, which brings great challenges to identify them. At present, researches mainly focus on the plaintext DNS covert tunnel, but less on the encrypted DOH tunnel. In this paper, we propose DOH covert tunnel detection method based on core features and machine learning method using two steps. Firstly, we detect DOH traffic according to the threshold of features. On this basis, we use core features and machine learning methods to detect tunnel traffic in all DOH traffic. Finally, we use self collected and public datasets to verify our method. The results show that the method achieves up to 99 % precision and recall that is superior to state of the art method. Bingxu Wang, Gang Xiong 0001, Gaopeng Gou, Jiaying Song, Zhen Li 0011, Qingya Yang |
CSCWD | 6 |
| 2023 | A Recurrent Self-learning Labeler for Building Network Traffic Ground TruthabstractWith the increasing number of traffic category, machine learning-based methods have gradually become the mainstream way of traffic classification to support network security and management. Machine learning-based methods require a large amount of high quality labelled data to learn network behavior patterns to achieve better recognition results. In the field of network traffic labeling, manual labeling can achieve more accurate labeling results, but the labeling efficiency is low and the labor cost is high. Deep packet inspection (DPI) technology can greatly improve labeling efficiency and reduce labor costs, but DPI labeling suffers from the problem of inaccurate and incomplete labeling. In this paper, we propose a recurrent self-learning framework (RSL-Labeler) for traffic labeling, which can solve the problem of inaccurate and incomplete DPI labeling. This framework consists of three components: high-quality data generation, class behavior pattern learning, and confidence filtering. Based on high-quality data labeled by multiple DPIs, we build three classifiers to learn the behavior patterns of DPI labeling intelligently from three perspectives. Then, we propose the idea of confidence filtering, which combines the pseudo-labeled data and the confidence values of three learning models to filter the credible samples by combined voting. These samples are added to the self-learning model for recurrent training. Experiments show that our method is able to label application traffic with accuracy of 99%, which is at least 8% better than single DPI. Qingya Yang, Chang Liu 0049, Peipei Fu, Bingxu Wang, Gaopeng Gou, Gang Xiong 0001 |
CSCWD | 1 |
| 2023 | Multi-Feature Fusion Based Approach for Classifying Encrypted Mobile Application TrafficabstractWith rapid development of mobile Internet, a great number of mobile applications has emerged, presenting a great explosion in mobile Internet traffic. Therefore, accurate classification of application traffic is necessary to more effectively manage mobile Internet traffic. However, the encryption of mobile application traffic gradually eliminates traditional classification approaches based on specific signatures, greatly increasing the difficulty of the classification of mobile application traffic. Therefore, we propose a novel multi-feature fusion (MFF)- based approach to enhance the accuracy of mobile application traffic classification. We also extract packet length sequence, byte sequence, statistical feature, etc. Then, we perform weighted fusions of features based on Relief-F algorithm to achieve the best set of features. Finally, we use machine learning techniques for application classification. Compared to several other feature extraction methods, MFF achieves an excellent performance with an accuracy of 97.6% for 16 mobile applications and a F1-score of over 99% for VPN-nonVPN. Qingya Yang, Peipei Fu, Junzheng Shi, Bingxu Wang, Zhen Li 0011, Gang Xiong 0001 |
CSCWD | 1 |
| 2022 | Anomaly Detection in Encrypted Identity Resolution Traffic based on Machine LearningabstractIdentity resolution is an emerging network resource widely applied in Industrial Internet of Things. Although encryption improves the privacy of identity resolution, it also challenges DPI-based anomaly detection. Therefore, it is imperative to recognize and supplement the encrypted information of IDS. In this paper, we design a machine learning-based framework to automatically extract critical information of identity resolution system from network traffic. According to the characteristics of traffic, we use the hybrid feature of statistics and sequences to describe encrypted traffic. Besides, a supervised classification algorithm is applied to explore the effective classification of two communication processes, which are service attribution information for node addressing and operation behavior for data management. We tested this method based on the encrypted traffic collected from a realistic identity resolution system. The results indicate that our approach exhibits good performance, outperforms related works, and can be applied in resource-constrained industrial scenario. This is the first work analysing the identity resolution system from the perspective of traffic analysis. Zhishen Zhu, Qingya Yang, Chonghua Wang, Zhen Li 0011 |
QRS | 3 |
| 2022 | BSBA: Burst Series Based Approach for Identifying Fake Free-trafficabstractIn recent years, mobile traffic has gradually become a major part of network traffic. To attract customers, mobile network operators provide free-traffic, which is a preferential policy that is free of charge for specific application traffic. Since the emergence of free-traffic, fake free-traffic also appeared soon. Fake free-traffic is a malicious behavior, which helps attackers illegally use network resources and evade network resource charging. The appearance of fake free-traffic maliciously harms the interests of operators and disrupts the rules of network resource charging. Because of the uniqueness of free-traffic, it encapsulates a layer of the HTTP protocol in addition to the actual application communication protocol, existing studies on encrypted traffic analysis are not applicable to identify fake free-traffic. In this paper, we propose Burst Series Based Approach (BSBA), a novel method for identifying fake free-traffic. The key idea behind BSBA is to construct effective features by capturing the differences of burst series among fake free-traffic, free-traffic and non-free traffic, and combine the constructed features with machine learning algorithms to identify fake free-traffic. We collect a real-world traffic dataset and conduct evaluations to verify the effectiveness of the BSBA. Experiment results demonstrate that the BSBA achieves excellent performances (96.82% Accuracy, 96.46% Precision, 96.57% Recall and 96.51% F1-score) and is superior to the state-of-the-art methods. Chang Liu 0049, Zhen Li 0011, Qingya Yang, Anlin Xu, Gaopeng Gou |
WoWMoM | 4 |
| 2021 | Towards Multi-source Extension: A Multi-classification Method Based on Sampled NetFlow RecordsabstractWith the rapid development of the Internet, network traffic is growing explosively. It brings great challenges to the traditional traffic identification technology using full traffic analysis, which requires more resources to achieve the collection and analysis of full traffic. And, handling the raw traffic may lead to the compromise of user privacy. NetFlow has good compatibility with the existing routing or switching devices, can aggregate network traffic information, support traffic sampling, reduce the invasion of user privacy, and can effectively deal with the challenges. However, as NetFlow is usually output after traffic sampling to ensure the performance of network devices and only contains session-level statistical information, existing NetFlow research mostly focuses on the binary classification problems (e.g., specific anomaly traffic detection), and less exploration has been conducted on traffic multi-classification problems. And NetFlow is even less involved in the currently popular field of encrypted traffic classification. In this paper, we focus on how to perform encrypted traffic multi-classification research based on sampled NetFlow records and propose a multi-classification method based on the multi-source extension of sampled NetFlow records. To improve the distinguishability and applicability of the sampled NetFlow records, we extend and enrich the records with full consideration of the head or payload information in traffic data, including TTL values, Cipher Suites, etc. For different application scenarios, the methods based on head information extension and payload information extension are proposed, respectively. Through comprehensive experiments, the results show that the proposed method is more applicable and effective than the method based on a single N etFlow record in dealing with multi-classification problems in different encryption application scenarios. Peipei Fu, Qingya Yang, Yangyang Guan, Bingxu Wang, Gaopeng Gou, Zhen Li 0011, Gang Xiong 0001 |
TrustCom | 2 |
| 2021 | Attack versus Attack: Toward Adversarial Example Defend Website Fingerprinting AttackabstractWebsite Fingerprinting (WF) attack is a side channel attack against encrypted tunnels which infers network activities of encrypted tunnels users. WF attack has been successfully applied to the Tor network, which poses a huge threat to the privacy of Tor visitors. A lot of countermeasures are therefore proposed to defend against such attacks. However, the newest attack successfully undermined the existing defense leveraging deep learning technique. In this paper, we propose an defense named Attack to Attack (A2A) that leverages adversarial example to attack the attacker's classifier. A2A treats website fingerprinting model as a black box. In order to find effective adversarial examples for the attacker's model, A2A manipulates traffic iteratively according to the output of a substitute model which is an elaborate model intentionally learning a similar classification boundary with the attacker's model. We evaluate the effectiveness of A2A on a public tor traffic dataset and the newest WF attack. The experimental results show that the proposed method provides effective defense with a bandwidth overhead of 2.2%, which significantly outperforms the manually designed defense (typically has a bandwidth overhead of 31%). Chengshang Hou, Junzheng Shi, Mingxin Cui, Qingya Yang |
TrustCom | 4 |
| 2020 | NSA-Net: A NetFlow Sequence Attention Network for Virtual Private Network Traffic Detection
Peipei Fu, Chang Liu 0049, Qingya Yang, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
WISE (1) | 3 |