VLDB 2026 Research / reviewers in the wild / expert
Junzheng Shi
dblp:201/6203
· DBLP profile ↗
34ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0003-4653-1686ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 9 since 2021Security and privacy · 9 · 8 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ATOPOS: Dynamic Path Exploration with Adaptive Probe Construction for Extensive and Efficient Network Topology Discovery
Yaochen Ren, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Tianyu Cui, Junzheng Shi |
INFOCOM | 7 |
| 2026 | Odysseus: A Context-Level Pre-training Framework for Out-of-Distribution Encrypted Traffic Classification
Wenqi Dong, Longtao He, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Jianshuo Liu, Gang Xiong 0001 |
IWQoS | 6 |
| 2026 | CDWF: Few-Shot Learning for Cross-Domain Multi-Tab Website Fingerprinting
Xinlei Ju, Zhen Li 0011, Lihua Yin, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001 |
IWQoS | 8 |
| 2026 | MDDB-AETB: Malicious domain detection boosting based on alignment with encrypted traffic behavior in restricted scenarios
Mengrui Cao, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001, Zhen Li 0011 |
Comput. Secur. | 3 |
| 2026 | BAPTISM: A Robust Framework for Encrypted Malicious Traffic Identification With Low-Quality Training DataabstractMachine learning (ML) is highly effective for accurate encrypted malicious traffic identification by using highquality training data. In fact, obtaining such data is costly and challenging. As a result, many ML-based models are inevitably trained on low-quality data and perform poorly. To enhance performance, some methods utilize various sample selection techniques to choose confident samples for model training. However, they often rely on a single metric for this selection, which restricts their adaptability across diverse datasets and noise conditions. In this paper, we propose a robust framework BAPTISM for identifying encrypted malicious traffic with low-quality training data. Particularly, BAPTISM selects a suitable base model for each task, and trains it with early stopping to generate traffic representation before overfitting occurs. Then, we devise an adaptive metric selection strategy to select confident samples. By employing two metrics (JSD and CSD) to assess the characteristic of traffic representation from distinct perspective, we find the more proper metric for each class and apply it for confident sample selection. According to the confident samples and selected metric for each class, we develop a label correction tactic which adapts to class nature to improve the quality of training data. Finally, we employ parallel training strategy to train the base model with the corrected data, further mitigating the impact of low-quality data. We conduct experiments across three real-world malicious traffic datasets with various noise settings. The results demonstrate that BAPTISM is compatible with different base models and outperforms across noise ratios ranging from 20% to 90%. Meanwhile, BAPTISM consistently selects the confident samples with the highest purity and volume under each setting. Chang Liu 0049, Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Li Guo 0001, Binxing Fang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | PromptFuzz: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMsabstractLarge Language Models (LLMs) have gained widespread use in various applications due to their powerful capability to generate human-like text. However, prompt injection attacks, which involve overwriting a model’s original instructions with malicious prompts to manipulate the generated text, have raised significant concerns about the security and reliability of LLMs. In this paper, we propose PromptFuzz, a novel testing framework that leverages fuzzing techniques to systematically assess the robustness of LLMs against prompt injection attacks. Inspired by software fuzzing, PromptFuzz selects promising seed prompts and generates a diverse set of prompt injections to evaluate the target LLM’s resilience. PromptFuzz operates in two stages: theprepare phase, which involves selecting promising initial seeds and collecting few-shot examples, and thefocus phase, which uses the collected examples to generate diverse, high-quality prompt injections. By deploying the generated attack prompts from PromptFuzz in a real-world competition, we achieved the 7th ranking out of over 4000 participants (top 0.14%) within 2 hours, demonstrating PromptFuzz’s effectiveness compared to experienced human attackers. Additionally, we also deploy the generated attack prompts on 50 popular LLM-integrated online applications, including those from Coze and OpenAI, and found that 92% of them can be exploited by PromptFuzz. We also run PromptFuzz on 15 online LLM-based resume judging applications and found that 13 of these applications’ responses can be hijacked by PromptFuzz. Yangguang Shao, Jiahao Yu 0001, Hanwen Miao, Gaopeng Gou, Zhen Li 0011, Junzheng Shi |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Splash: Adversarial Defense with Short Perturbation Blocks Against Adversarial Training Aided Website FingerprintingabstractAdversarial perturbation generation allows network users to mislead website fingerprinting (WF) classifiers without compromising real-time transmission or data integrity, causing misclassification. However, adversarial perturbations are vulnerable to adversarial training (AT), which enables attackers to improve their classifiers using perturbed adversarial samples, rendering user defenses ineffective. Due to reliance on real and non-redundant data, existing defenses against AT fail to scale to large-scale user scenarios. This paper proposes an improved adversarial perturbation generation method named Splash, which mitigates performance degradation caused by defense configuration collisions in traditional AT-aided attack defenses by applying two real-time traffic obfuscation steps using both global adversarial perturbations and Short Perturbation Blocks placed at random positions. Evaluation shows that Splash performs better traffic obfuscation than three other representative defenses, causing attacker classifiers to misclassify over 97% of traffic. In addition, it offers enhanced functionality by causing 45-60% of traffic to be misclassified into arbitrary target classes. Splash outperforms SOTA defenses such as AWA and ALERT against AT-aided attacks, reducing success rates to below 30%. Furthermore, it demonstrates significantly stronger resilience when attackers adopt the same defense configurations as users. Runsheng Ma, Chengshang Hou, Gaopeng Gou, Junzheng Shi, Zhen Li 0011, Gang Xiong 0001 |
ACSAC | 4 |
| 2025 | PBC-MWF: Robust Multi-Tab Website Fingerprinting via Interval-Aggregated Packet-Burst CountsabstractWebsite fingerprinting (WF) attacks undermine the privacy promised by anonymizing networks such as Tor by inferring the websites a user visits from encrypted-traffic side-channels. Recent criticisms of the single-tab assumption have shifted attention to the more realistic multi-tab setting, where concurrent page loads create severe noise. Existing multi-tab studies rely on direction sequences that ignore temporal structure and therefore provide only limited discriminative power. Our experiments show that packet-level timestamps do carry extra signal, yet their raw form is fragile under overlapping tabs and timing-obfuscation defences. We propose the PacketBurst Counts (PBC) feature—a$4 \times L$matrix that, for each time interval, stores the counts of upstream packets, downstream packets, upstream bursts and downstream bursts. PBC preserves coarse temporal structure while discarding noisy finegrained timings, striking a balance between expressiveness and robustness. Building on PBC, we design PBC-MWF, an end-to-end framework that (i) uses a residual CNN to learn local embeddings and (ii) applies an adaptive sparse transformer to capture global correlations while suppressing tab-overlap noise. Unlike prior work, which evaluates on datasets with a fixed number of tabs and reports top-k accuracy, we additionally merge datasets with varying tab-count and perform threshold-based inference; the threshold is tuned on the validation set before testing. To the best of our knowledge, PBCMWF is the first WF framework to simultaneously address the multi-tab setting's challenges of fine-grained webpage identification and resilience against WF defences. Evaluations on three public multi-tab datasets demonstrate PBC-MWF's enhanced robustness: compared against nine baselines, it surpasses the best prior method-improving F1 by up to 12.4 % on site-level, 8.4 % on page-level, and over 10 % under defenses. Yuhao Wei, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou, Junzheng Shi, Yingchao Qin |
IPCCC | 5 |
| 2025 | T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video RetrievalabstractText-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing work has primarily focused on extending CLIP knowledge for video-text tasks. However, videos typically contain richer information than images. In current video-text datasets, textual descriptions can only reflect a portion of the video content, leading to partial misalignment in video-text matching. Therefore, directly aligning text representations with video representations can result in incorrect supervision, ignoring the inequivalence of information. In this work, we propose T2VParser to extract multiview semantic representations from text and video, achieving adaptive semantic alignment rather than aligning the entire representation. To extract corresponding representations from different modalities, we introduce Adaptive Decomposition Tokens, which consist of a set of learnable tokens shared across modalities. The goal of T2VParser is to emphasize precise alignment between text and video while retaining the knowledge of pretrained models. Experimental results demonstrate that T2VParser achieves accurate partial alignment through effective cross-modal content decomposition. The code is available at https://github.com/Lilidamowang/T2VParser. Yili Li, Gang Xiong 0001, Gaopeng Gou, Xiangyan Qu, Jiamin Zhuang, Zhen Li 0011, Junzheng Shi |
ACM Multimedia | 7 |
| 2025 | SSRCorr: A Self-Supervised Robust Flow Representation Learning Framework for Flow Correlation Attacks on TorabstractTor is one of the most widely adopted anonymity networks, yet its anonymity can be undermined by adversaries through flow correlation attacks. Current mainstream technologies focus on exploiting the sequence characteristics of packet lengths and timestamps to execute attacks. However, the padding mechanism of the Tor network and time delays caused by multi-hop relays obscure these single-modal features. Additionally, the diversity of network services and the randomness of user behavior result in sparse packet distributions, which impact model training and inference. In this paper, we propose SSRCorr, a novel self-supervised learning framework for flow correlation attacks, incorporating the Flow Feature Aggregation (FFA) module and Global-Local Fusion (GLoF) Encoder to address these challenges. Firstly, we construct a Byte-based Traffic Aggregation Matrix (BTAM) by integrating time and length sequences and applying two data augmentation methods tailored for Tor flow correlation, thereby reducing the impact of Tor network noise on attack effectiveness. Secondly, we employ GLoF to extract features from the output by FFA and fuse the global context information of the traffic, thus mitigating the impact of low-information traffic on model performance. Experiments show that SSRCorr achieves a TPR of 96%, surpassing other methods, and maintains robust performance under temporal drift and obfuscation, supporting future research on countering anonymity system defenses. Mengyan Liu, Yaochen Ren, Yanbo Wu, Yangyang Guan, Zhen Li 0011, Gaopeng Gou, Junzheng Shi |
TrustCom | 9 |
| 2025 | ProxyCorr: robust traffic correlation attacks via mixed spatio-temporal analysis in encrypted proxy networks
Mengyan Liu, Gaopeng Gou, Gang Xiong 0001, Junzheng Shi, Hanwen Miao |
Comput. Networks | 4 |
| 2025 | Enhanced detection of obfuscated HTTPS tunnel traffic using heterogeneous information network
Mengyan Liu, Gaopeng Gou, Gang Xiong 0001, Junzheng Shi, Hanwen Miao, Yang Li 0002 |
Comput. Networks | 4 |
| 2024 | WebPromptM2: A Website Classification Method Leveraging Prompt-Based Learning with Multimodal FeaturesabstractWebsite classification proves crucial for tasks like malicious website detection and information management. Current methods typically focus on effective feature extraction and algorithm selection to create balanced website datasets, often leading to decreased performance due to data imbalance. In this study, we propose an intelligent website classification method(WebPromptM2) based on prompt-based learning with multimodal features. We design a prompt template which incorporates the textual and visual elements of the website, thereby facilitating a multimodal representation of the website, then leverage domain-specific expertise to establish mapping relationships between website categories and a label word set. Finally, we fine-tune the masked pre-trained language model (PLM) and map the prediction results to the categories. We find that our method increases recognition accuracy of tail classes and achieves superior performance on long-tail and short-tail datasets. Mengyan Liu, Gaopeng Gou, Gang Xiong 0001, Junzheng Shi, Chang Liu 0049 |
CSCWD | 4 |
| 2023 | Multi-Feature Fusion Based Approach for Classifying Encrypted Mobile Application TrafficabstractWith rapid development of mobile Internet, a great number of mobile applications has emerged, presenting a great explosion in mobile Internet traffic. Therefore, accurate classification of application traffic is necessary to more effectively manage mobile Internet traffic. However, the encryption of mobile application traffic gradually eliminates traditional classification approaches based on specific signatures, greatly increasing the difficulty of the classification of mobile application traffic. Therefore, we propose a novel multi-feature fusion (MFF)- based approach to enhance the accuracy of mobile application traffic classification. We also extract packet length sequence, byte sequence, statistical feature, etc. Then, we perform weighted fusions of features based on Relief-F algorithm to achieve the best set of features. Finally, we use machine learning techniques for application classification. Compared to several other feature extraction methods, MFF achieves an excellent performance with an accuracy of 97.6% for 16 mobile applications and a F1-score of over 99% for VPN-nonVPN. Qingya Yang, Peipei Fu, Junzheng Shi, Bingxu Wang, Zhen Li 0011, Gang Xiong 0001 |
CSCWD | 3 |
| 2023 | Bitcoin Mixing Service Detection Based on Spatio-Temporal Information Representation of Transaction GraphabstractCoin mixing is a technique used to enhance Bitcoin’s anonymity and can be used to obfuscate the relationship among transaction input addresses. Due to this property, much of the criminal activity on Bitcoin uses coin-mixing techniques to launder money, making these illicit funds difficult to trace. Therefore, it is important to implement the detection of Bitcoin mixing services. Several methods for identifying bitcoin mixing services have been proposed, but balancing their efficiency and generality at the same time is a challenging task. In this paper, We propose STMD (Spatio-Temporal Mixing Detector), which combines local features and global features of Bitcoin transactions to identify coin-mixing transactions. On one hand, we extract and process the statistical features of neighboring nodes of the transaction as local features. On the other hand, we construct a global position encoding (GPE) containing spatio-temporal information of the transaction as global features. Additionally, we employ the attention mechanism to handle these two types of features, effectively combining them. Finally, we utilize linear layers to achieve the detection of coin-mixing transactions. The experimental results show that STMD performs better than existing methods on the same dataset; it also has a higher recall on the test set of other types of coin-mixing transactions, which reflects the generality of the model. In particular, We apply local features and global features for experiments separately and verify the necessity of the two features. The results of the model trained using only global features also outperform the existing methods, which shows that the global position encoding (GPE) we constructed is effective for mixed currency transaction identification. Hanzhi Yang, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001, Zhen Li 0011 |
IPCCC | 4 |
| 2023 | Few-shot encrypted traffic classification via multi-task representation enhanced meta-learning
Gang Xiong 0001, Junzheng Shi, Gaopeng Gou, Zhen Li 0011, Chang Liu 0049 |
Comput. Networks | 4 |
| 2022 | GALG: Linking Addresses in Tracking Ecosystem Using Graph Autoencoder with Link Generation
Tianyu Cui, Gang Xiong 0001, Chang Liu 0049, Junzheng Shi, Peipei Fu, Gaopeng Gou |
ECML/PKDD (6) | 4 |
| 2022 | MCFM: Discover Sensitive Behavior from Encrypted Traffic in Industrial Control SystemabstractTo tackle with advanced persistent threats against industrial control system, Siemens has developed S7CommPlus- TLS, a new version of the encrypted protocol challenging traditional DPI-based anomaly detection methods. However, the communication mode of industrial control system leads to the overlapping of periodic traffic and sensitive behavior traffic, and thus makes mainstream encrypted traffic classification methods exhibit a poor performance in S7CommPlus-TLS protocol. Therefore, we design a multiple clustering framework called MCFM, which can automatically extract sensitive behavior of S7CommPlus-TLS from network traffic. The first-clustering is used as a pre-processing model to separate and remove periodic traffic from overlapping flows according to the communication mode of industrial control system. Besides, we employ the second- clustering as a generator to extract the fingerprint of sensitive behaviors. Our comprehensive experiments on the simulation dataset covering six sensitive behaviors indicate that MCFM achieves an excellent performance, and outperforms present cutting-edge methods. To the best of our knowledge, this is the first work analyzing industrial control system from the perspective of encrypted traffic analysis. Zhishen Zhu, Junzheng Shi, Chonghua Wang, Gang Xiong 0001, Zhiqiang Hao, Gaopeng Gou |
TrustCom | 2 |
| 2022 | ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationabstractEncrypted traffic classification requires discriminative and robust traffic representation captured from content-invisible and imbalanced traffic data for accurate classification, which is challenging but indispensable to achieve network security and network management. The major limitation of existing solutions is that they highly rely on the deep features, which are overly dependent on data size and hard to generalize on unseen data. How to leverage the open-domain unlabeled traffic data to learn representation with strong generalization ability remains a key challenge. In this paper, we propose a new traffic representation model called Encrypted Traffic Bidirectional Encoder Representations from Transformer (ET-BERT), which pre-trains deep contextualized datagram-level representation from large-scale unlabeled data. The pre-trained model can be fine-tuned on a small number of task-specific labeled data and achieves state-of-the-art performance across five encrypted traffic classification tasks, remarkably pushing the F1 of ISCX-VPN-Service to 98.9% (5.2%↑), Cross-Platform (Android) to 92.5% (5.4%↑), CSTNET-TLS 1.3 to 97.4% (10.0%↑). Notably, we provide explanation of the empirically powerful pre-training model by analyzing the randomness of ciphers. It gives us insights in understanding the boundary of classification ability over encrypted traffic. The code is available at: https://github.com/linwhitehat/ET-BERT. Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Jing Yu 0007 |
WWW | 5 |
| 2021 | Let Imbalance Have Nowhere to Hide: Class-Sensitive Feature Extraction for Imbalanced Traffic ClassificationabstractWith the full encryption of network traffic, traffic classification schemes based on machine learning emerge in endlessly. Class imbalance, as a widely-studied challenge in machine learning, has not attracted enough attention in traffic classification researches. The uneven distribution hidden in the real-world traffic will cause performance degradation of the existing schemes. In existing methods, data pre-sampling is easy to introduce noise or lose massive information; the cost matrix of cost-sensitive methods is difficult to design; feature selection methods will filter out lots of “redundant” features and cause unsatisfactory results. In this paper, we propose an effective end-to-end framework for imbalanced traffic classification which avoids the above weaknesses, called DeepFE. We adopt deep neural networks for feature extraction, and model features from the perspective of channels. It can learn class-sensitive feature representation, which is quite helpful to distinguish the minority traffic classes. Moreover, DeepFE can be applied to various tasks because of its unlimited input format, i.e., both the raw bytes and the packet length sequence can be used. We conducted experiments on the public dataset ISCXVPN2016 and a realworld traffic dataset covering 27 applications. The results show that DeepFE achieves excellent results, significantly alleviating the performance degradation caused by imbalance, and surpasses several state-of-the-art methods. Gaopeng Gou, Gang Xiong 0001, Junzheng Shi |
IJCNN | 5 |
| 2021 | TA-GAN: GAN based Traffic Augmentation for Imbalanced Network Traffic ClassificationabstractAs the mainstream in network traffic classification (NTC), machine learning (ML) based methods suffer performance degradation due to the imbalance distribution of Internet traffic. Data augmentation methods including the traditional oversampling techniques and the Generative Adversarial Network (GAN) based generation methods are most commonly used to counter the imbalance problem in NTC. However, the former is prone to overfitting and introducing noise. The latter overcomes the above weaknesses, but the quality of the generated traffic samples is difficult to judge. Besides, these methods all divide the imbalanced traffic classification problem into two subproblems, which cannot guarantee the global optimality. In this paper, we propose a GAN based Traffic Augmentation (TA-GAN) for imbalanced traffic classification. TA-GAN is an end-to-end framework that integrates the generation of the minority traffic samples with the training of the target classifier. We design the feedback mechanism to better guide the direction of the sample generation and simultaneously indicate the quality of the synthesized samples. Moreover, the existing deep learning-based NTC methods can be easily adapted to imbalance scenarios with TA-GAN. Comprehensive experiments on the public ISCXVPN2016 dataset demonstrate that TA-GAN effectively mitigates the influence of traffic imbalance (a maximum 14.64% improvement to the minority class'$F_{1}$score) and outperforms the state-of-the-art methods. Gang Xiong 0001, Zhen Li 0011, Junzheng Shi, Mingxin Cui, Gaopeng Gou |
IJCNN | 4 |
| 2021 | Combating Imbalance in Network Traffic Classification Using GAN Based OversamplingabstractWith the proliferation of encrypted traffic, machine learning (ML) based network traffic classification (NTC) has become the mainstream method. However, most studies ignored two issues. On the one hand, Internet traffic presents a natural uneven distribution. On the other hand, machine learning algorithms generally aim to achieve the highest overall accuracy without considering class imbalance. This leads to severe performance degradation of existing ML-based NTC schemes when facing imbalanced scenarios. In this paper, we design a novel Generative Adversarial Network (GAN) architecture to generate traffic samples, in which the addition of the classifier and the pretraining module makes the generation process more stable and effective. We propose an end-to-end framework for imbalanced traffic classification, named ITCGAN, which can generate traffic samples for minority classes to adaptively rebalance the original traffic and simultaneously train the optimal classifier. We evaluate its effectiveness on the public ISCXVPN2016 dataset based on the global metrics and individual metrics. The results show that our method performs well in imbalanced NTC tasks, fully alleviating the performance degradation (a 10.27-percentage-point improvement to the precision of the most minority class). Meanwhile, it surpasses five state-of-the-art oversampling methods. Gang Xiong 0001, Zhen Li 0011, Junzheng Shi, Mingxin Cui, Gaopeng Gou |
Networking | 4 |
| 2021 | Universal Website Fingerprinting Defense Based on Adversarial ExamplesabstractWebsite fingerprinting (WF) attacks pose a threat to privacy of web activity, especially on anonymity networks such as Tor. Recent studies show that the deep neural network (DNN) significantly improves the impact of website fingerprinting attacks. Especially, DNN-based attack undermines the existing defense methods which are mainly rely on the manually designed rule. In this paper, we present a novel defense that generates universal perturbation that can transform original examples to adversarial examples which is effectively defending against a specific WF model. The proposed defense is evaluated on state-of-the-art DNN attack over a public Tor traffic dataset. The experimental results show our adversarial example generation method performs better than the baseline methods. The proposed defense defeats all existing WF attacks based on deep neural networks with a low overhead. Comparing with state-of-the-art defenses such as Walkie-Talkie and WTF-PAD with a lower bound of 31% and 64% overheads, the proposed defense achieves identical defense performance with at least 50% bandwidth overhead saving. Chengshang Hou, Junzheng Shi, Mingxin Cui, Mengyan Liu, Jing Yu 0007 |
TrustCom | 2 |
| 2021 | Attack versus Attack: Toward Adversarial Example Defend Website Fingerprinting AttackabstractWebsite Fingerprinting (WF) attack is a side channel attack against encrypted tunnels which infers network activities of encrypted tunnels users. WF attack has been successfully applied to the Tor network, which poses a huge threat to the privacy of Tor visitors. A lot of countermeasures are therefore proposed to defend against such attacks. However, the newest attack successfully undermined the existing defense leveraging deep learning technique. In this paper, we propose an defense named Attack to Attack (A2A) that leverages adversarial example to attack the attacker's classifier. A2A treats website fingerprinting model as a black box. In order to find effective adversarial examples for the attacker's model, A2A manipulates traffic iteratively according to the output of a substitute model which is an elaborate model intentionally learning a similar classification boundary with the attacker's model. We evaluate the effectiveness of A2A on a public tor traffic dataset and the newest WF attack. The experimental results show that the proposed method provides effective defense with a bandwidth overhead of 2.2%, which significantly outperforms the manually designed defense (typically has a bandwidth overhead of 31%). Chengshang Hou, Junzheng Shi, Mingxin Cui, Qingya Yang |
TrustCom | 2 |
| 2020 | PST: a More Practical Adversarial Learning-based Defense Against Website FingerprintingabstractTo prevent serious privacy leakage from website fingerprinting (WF) attacks, many traditional or adversarial WF defenses have been released. However, traditional WF defenses such as Walkie-Talkie (W-T) still generate patterns that might be captured by the deep learning (DL) based WF attacks, which are not effective. Adversarial perturbation based WF defenses better confuse WF attacks, but their requirements for the entire original traffic trace and perturbating any points including historical packets or cells of the network traffic are not practical. To deal with the effectiveness and practicality issues of existing defenses, we proposed a novel WF defense in this paper, called PST. Given a few past bursts of a trace as input, PST Predicts subsequent fuzzy bursts with a neural network, then Searches small but effective adversarial perturbation directions based on observed and predicted bursts, and finally Transfers the perturbation directions to the remaining bursts. Our experimental results over a public closed-world dataset demonstrate that PST can successfully break the network traffic pattern and achieve a high evasion rate of 87.6%, beating W-T by more than 31.59% at the same bandwidth overhead, with only observing 10 transferred bursts. Moreover, our defense adapts to WF attacks dynamically, which could be retrained or updated. Yong Wang 0032, Gaopeng Gou, Wei Cai 0007, Gang Xiong 0001, Junzheng Shi |
GLOBECOM | 6 |
| 2020 | Not Afraid of the Unseen: a Siamese Network based Scheme for Unknown Traffic DiscoveryabstractAs an essential task for network management and security, network traffic classification has attracted increasing attention in recent years. Traditional traffic classification methods achieve certain success in identifying specific application traffic but fail with un-predefined unknown classes. Existing unknown traffic discovery methods commonly pick out some unlabeled testing data as part of training data to train the classification models, which is not in line with the real-world open environments. In this paper, we propose a novel scheme named SEEN to achieve unknown traffic detection in network traffic classification. There are three crucial phases in the SEEN: unknown discovery, unknown clustering, and system update. In the first step, using a metric-based approach with siamese network, SEEN identifies unknown traffic as well as accurately classifies the traffic generated by pre-defined application classes. After discovery, unknown traffic is automatically clustered into more fine-grained categories in the unknown clustering step. In the system update step, inspired by low-shot learning, SEEN allows new classes to be added or unnecessary known classes to be deleted quickly without retraining from the sketch, which can complement the system’s knowledge. Experimental results exhibit that SEEN can achieve outstanding performances both on known and unknown traffic identification on two open real-world datasets, and the proposed scheme can address the problem of unknown traffic effectively. Zhen Li 0011, Junzheng Shi, Gaopeng Gou, Chang Liu 0049, Gang Xiong 0001 |
ISCC | 3 |
| 2020 | WF-GAN: Fighting Back Against Website Fingerprinting Attack Using Adversarial LearningabstractWebsite Fingerprinting (WF) attack is an side-channel attack which aims at encrypted web traffic. WF attackers recognize encrypted website traffic through constructing fingerprinting for each website using the flow-based features extracted from encrypted traffic. WF defense typically aims at modifying the features of the encrypted websites. However, those countermeasures either cause high overhead or fail to counter the subsequent WF attacks. Especially, the newest WF attacks, which are based on deep neural network, is able to classify the defended traffic by directly learning from the labeled defended traffic. In this paper, we propose an novel defense through making use of the trick that machine learning models are vulnerable to adversarial exmaples. We design WF-GAN, a GAN with an additional WF classifier component, to generate adversarial examples for WF classifiers through adversarial learning. As the website set is divided into source and target website, WF-GAN are trained to map websites features from source set to adversarial examples and make adversarial examples more similar to the website features in the target set. The experimental result shows that WF-GAN achieves 90% success rate with at most 15% overhead for untargeted defense, which outperforms previous defense. In addition, adversarial examples based defense support targeted defense, which is not support by traditional defense. The result shows that WF-GAN achieves over 90% targeted defense success rate when the target websites set is twice as many as the source website set. Chengshang Hou, Gaopeng Gou, Junzheng Shi, Peipei Fu, Gang Xiong 0001 |
ISCC | 3 |
| 2020 | 6VecLM: Language Modeling in Vector Space for IPv6 Target Generation
Tianyu Cui, Gang Xiong 0001, Gaopeng Gou, Junzheng Shi |
ECML/PKDD (4) | 4 |
| 2019 | A Comprehensive Study of Accelerating IPv6 DeploymentabstractSince the lack of IPv6 network development, China is currently accelerating IPv6 deployment. In this scenario, traffic and network structure show a huge shift. However, due to the long-term prosperity, we are ignorant of the problems behind such outbreak of traffic and performance improvement events in accelerating deployment. IPv6 development in some regions will still face similar challenges in the future. To contribute to solving this problem, in this paper, we produce a new measurement framework and implement a 5-month passive measurement on the IPv6 network during the accelerating deployment in China. We combine 6 global-scale datasets to form the normal status of IPv6 network, which is against to the accelerating status formed by the passive traffic. Moreover, we compare with the traffic during World IPv6 Day 2011 and Launch 2012 to discuss the common nature of accelerating deployment. Finally, the results indicate that the IPv6 accelerating deployment is often accompanied by an unbalanced network status. It exposes unresolved security issues including the challenge of user privacy and inappropriate access methods. According to the investigation, we point the future IPv6 development after accelerating deployment. Tianyu Cui, Chang Liu 0049, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001 |
IPCCC | 4 |
| 2019 | Identify OS from encrypted traffic with TCP/IP stack fingerprintingabstractMore and more security vulnerabilities are closely related to operating system (OS) information, but how to accurately identify OS versions on a real-world dynamic network in encrypted traffic is still a challenge. In this paper, we propose a comprehensive passive OS identification method based on encrypted traffic. It takes advantage of several features in TLS headers and TCP/IP headers. Moreover, we also consider flow statistic features for each session. We collect a large dataset of more than 2 million samples to evaluate the performance of our approach. According to the experimental results, the performance of the proposed method is preferable to the traditional method. Xinlei Fan, Gaopeng Gou, Cuicui Kang, Junzheng Shi, Gang Xiong 0001 |
IPCCC | 4 |
| 2019 | I Know What You Are Doing With Remote DesktopabstractRemote desktop enables users to remotely access their computers via the Internet, which is widely used as a basic tool in areas such as remote work, remote assistance and remote administration. However, existing remote desktop is designed to work in the mode of updating user's real-time command and remote screen's state interactively for a better user experience, such working mode may cause serious side-channel information leakage problem in spite of encryption of the traffic, as revealed in this paper. We carry out an experimental research to assess the side-channel information leakage of six most popular remote desktop softwares in Windows 10 & 7 platforms: Anydesk, ConnectWise, MicroRDS, RealVNC, Teamviewer, and Zoho Assist. With the help of machine learning techniques including logistic regression, support vector machine, gradient boosting decision tree, random forest as well as statistic features of flow burst, we observe that an adversary can excellently uncover (top at 99.26% TPR, 0.57% FPR, 97.17% F1-score) 5 rough kinds of daily activities covering editing documents, reading documents, surfing webs, watching videos and installing softwares and even worse precisely classify 4 fine activities predefined as editing documents with Microsoft Office Word and the other three edit tools with high true positive rate and low false positive rate. Our results prove the fact for remote desktop traffic encryption mechanism is nothing sufficient to prevent side-channel information leakage and both users and providers of remote desktop should pay more attention to such serious privacy leakage problem. Gaopeng Gou, Junzheng Shi, Gang Xiong 0001 |
IPCCC | 3 |
| 2018 | Classifying User Activities in the Encrypted WeChat TrafficabstractThe security and privacy of encrypted mobile applications have attracted the attention of researchers. However, most of the existing researches focus on analysis of SSL/TLS traffic, while few studies focus on proprietary encrypted traffic, which is also important and challenging. In this paper, we make a deep study of WeChat, which is one of the most popular social applications in the world with over one billion active users. The application uses a proprietary encryption protocol called as MMTLS for most of its communications. It is designed based on Transport Layer Security (TLS) 1.3 drafts for both performance and security. We explore the fine-grained classification of typical user activities inside the MMTLS encrypted channels and compare the MMTLS with the HTTPS (e.g. flow duration and packet size), which are jointly used in WeChat. It is found that MMTLS is suitable for scenarios of low latency and lightweight messaging. With the WeChat traffic collected from different platforms (Android, iOS) and devices (Huawei, Samsung, iPhone, iPad, etc.) by different users, we classify seven typical activities, encrypted by MMTLS protocol such as payment, advertisement click, browsing moments and so on. The experimental results show that both of the average precision and recall can reach over 92%. Our work is the first to perform classification on this proprietary encrypted protocol and understanding the difference between MMTLS and TLS. It is believed that the work will benefit the security and privacy of WeChat and other proprietary encryption applications. Chengshang Hou, Junzheng Shi, Cuicui Kang, Zigang Cao, Xiong Gang |
IPCCC | 2 |
| 2017 | POSTER: A Comprehensive Study of Forged Certificates in the WildabstractWith the widespread use of SSL, many issues have been exposed as well. Forged certificates used for MITM attacks or proxies can make SSL encryption useless easily, leading to privacy disclosure and property loss of careless victims. In this paper, we implement a large scale of passive measurement of SSL/TLS and analyze the forged certificates in the wild comprehensively. We measured SSL/TLS connection for 16 months on two large research networks, which provided a total of 100 Gbps bandwidth. We gathered nearly 135 million leaf certificates and studied the forged ones. Our findings reveal main reasons of signing forged certificates, and show the preference of them. Finally, we find out several suspicious servers that might be used for MITM. Mingxin Cui, Zigang Cao, Gang Xiong 0001, Junzheng Shi |
CCS | 4 |
| 2017 | Auto-identification of background traffic based on autonomous periodic interactionabstractBackground traffic of web applications refers to the traffic not generated directly due to user activities (e.g. user behavior profiling) that is usually useful to the application providers, but not the users. A recent study indicated that background traffic, contributing 51.8% bandwidth, has exceeded user-generated traffic. Accurate identification of background traffic can help network managers to optimize network resource allocation and avoid network congestion. However, identification of background traffic is not easy and the solution must be robust enough for all applications. In this paper, we propose the first method that can self-learn background traffic rules from unlabeled data and automatically identify online background traffic. The accuracy of the extracted rules is 90.51%. When applying our method in a real enterprise network, the false positive rate (FPR) is only 3% showing that our method is accurate and effective. Our method is derived from a critical observation that the background traffic exhibits a periodic behavior (referred as autonomous periodic interaction (AuPI)). Technically, we propose two indexes, Time Regularity Factor (TRF) and Time Interval Factor (TIF), to capture this AuPI pattern from unlabeled communication traffic. As a side contribution, we created a public benchmark dataset of 45 hot applications with 97,000+ background traffic flows that can be used by researchers to further investigate background traffic. Chang Liu 0049, Lingwu Zeng, Junzheng Shi, Gang Xiong 0001, Siu-Ming Yiu |
IPCCC | 3 |