EDBT 2026 Demo / reviewers in the wild / expert
Zhen Li 0011
dblp:74/2397-11
· DBLP profile ↗
97ranked-venue papers
1as first author
74since 2021 · last 2026
0000-0002-3892-4909ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 36 · 25 since 2021Security and privacy · 31 · 25 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnhanCorr: Stable and Enhanced Flow Correlation under Bursty Traffic
Xinlei Ju, Zhen Li 0011, Yuguo Wang, Gaopeng Gou, Gang Xiong 0001 |
ICC | 2 |
| 2026 | ATOPOS: Dynamic Path Exploration with Adaptive Probe Construction for Extensive and Efficient Network Topology Discovery
Yaochen Ren, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Tianyu Cui, Junzheng Shi |
INFOCOM | 5 |
| 2026 | Odysseus: A Context-Level Pre-training Framework for Out-of-Distribution Encrypted Traffic Classification
Wenqi Dong, Longtao He, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Jianshuo Liu, Gang Xiong 0001 |
IWQoS | 5 |
| 2026 | CDWF: Few-Shot Learning for Cross-Domain Multi-Tab Website Fingerprinting
Xinlei Ju, Zhen Li 0011, Lihua Yin, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001 |
IWQoS | 2 |
| 2026 | Dive into streaming: efficient identification of encrypted dynamic DASH video traffic
Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Binxing Fang |
Sci. China Inf. Sci. | 4 |
| 2026 | EN-Fusion: Malware detection through end-net fusion representation
Ziqian Chen, Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Haikuo Li |
Comput. Networks | 5 |
| 2026 | MDDB-AETB: Malicious domain detection boosting based on alignment with encrypted traffic behavior in restricted scenarios
Mengrui Cao, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001, Zhen Li 0011 |
Comput. Secur. | 5 |
| 2026 | TranCTC: Transformer-based IPv6 covert timing channel detection with heterogeneous features fusionabstractNetwork Covert Timing Channels (NCTCs) pose a serious threat to network security, through which attackers transmit hidden information by manipulating inter-packet delays (IPDs). Existing methods have shown strong performance in IPv4 networks by relying solely on timing features. However, more flexible routing mechanisms may introduce additional timing jitter into benign traffic in IPv6; meanwhile, the heterogeneous processing of optional IPv6 extension headers by routers make per-hop processing delays across routers increasingly unpredictable. These factors collectively increase the complexity of IPv6 IPD patterns, making timing-only detection methods insufficient for accurately modeling legitimate baseline behavior. To address such challenges, we propose TranCTC , a Transformer-based anomaly detection method. By modeling temporal-structural alignment, TranCTC effectively overcomes the limitations of timing-only baseline construction and learns a multidimensional representation of legitimate IPv6 traffic. On the public CAIDA dataset, TranCTC significantly outperforms existing approaches across multiple NCTC types in the IPv6/TCP setting, and ablation studies show that incorporating structural features and contextual modeling substantially improve detection accuracy. TranCTC fills the gap in detecting covert timing channels in IPv6 networks. Yuguo Wang, Gaopeng Gou, Xinlei Ju, Zhen Li 0011, Gang Xiong 0001 |
Comput. Secur. | 6 |
| 2026 | BAPTISM: A Robust Framework for Encrypted Malicious Traffic Identification With Low-Quality Training DataabstractMachine learning (ML) is highly effective for accurate encrypted malicious traffic identification by using highquality training data. In fact, obtaining such data is costly and challenging. As a result, many ML-based models are inevitably trained on low-quality data and perform poorly. To enhance performance, some methods utilize various sample selection techniques to choose confident samples for model training. However, they often rely on a single metric for this selection, which restricts their adaptability across diverse datasets and noise conditions. In this paper, we propose a robust framework BAPTISM for identifying encrypted malicious traffic with low-quality training data. Particularly, BAPTISM selects a suitable base model for each task, and trains it with early stopping to generate traffic representation before overfitting occurs. Then, we devise an adaptive metric selection strategy to select confident samples. By employing two metrics (JSD and CSD) to assess the characteristic of traffic representation from distinct perspective, we find the more proper metric for each class and apply it for confident sample selection. According to the confident samples and selected metric for each class, we develop a label correction tactic which adapts to class nature to improve the quality of training data. Finally, we employ parallel training strategy to train the base model with the corrected data, further mitigating the impact of low-quality data. We conduct experiments across three real-world malicious traffic datasets with various noise settings. The results demonstrate that BAPTISM is compatible with different base models and outperforms across noise ratios ranging from 20% to 90%. Meanwhile, BAPTISM consistently selects the confident samples with the highest purity and volume under each setting. Chang Liu 0049, Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Li Guo 0001, Binxing Fang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | PromptFuzz: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMsabstractLarge Language Models (LLMs) have gained widespread use in various applications due to their powerful capability to generate human-like text. However, prompt injection attacks, which involve overwriting a model’s original instructions with malicious prompts to manipulate the generated text, have raised significant concerns about the security and reliability of LLMs. In this paper, we propose PromptFuzz, a novel testing framework that leverages fuzzing techniques to systematically assess the robustness of LLMs against prompt injection attacks. Inspired by software fuzzing, PromptFuzz selects promising seed prompts and generates a diverse set of prompt injections to evaluate the target LLM’s resilience. PromptFuzz operates in two stages: theprepare phase, which involves selecting promising initial seeds and collecting few-shot examples, and thefocus phase, which uses the collected examples to generate diverse, high-quality prompt injections. By deploying the generated attack prompts from PromptFuzz in a real-world competition, we achieved the 7th ranking out of over 4000 participants (top 0.14%) within 2 hours, demonstrating PromptFuzz’s effectiveness compared to experienced human attackers. Additionally, we also deploy the generated attack prompts on 50 popular LLM-integrated online applications, including those from Coze and OpenAI, and found that 92% of them can be exploited by PromptFuzz. We also run PromptFuzz on 15 online LLM-based resume judging applications and found that 13 of these applications’ responses can be hijacked by PromptFuzz. Yangguang Shao, Jiahao Yu 0001, Hanwen Miao, Gaopeng Gou, Zhen Li 0011, Junzheng Shi |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Splash: Adversarial Defense with Short Perturbation Blocks Against Adversarial Training Aided Website FingerprintingabstractAdversarial perturbation generation allows network users to mislead website fingerprinting (WF) classifiers without compromising real-time transmission or data integrity, causing misclassification. However, adversarial perturbations are vulnerable to adversarial training (AT), which enables attackers to improve their classifiers using perturbed adversarial samples, rendering user defenses ineffective. Due to reliance on real and non-redundant data, existing defenses against AT fail to scale to large-scale user scenarios. This paper proposes an improved adversarial perturbation generation method named Splash, which mitigates performance degradation caused by defense configuration collisions in traditional AT-aided attack defenses by applying two real-time traffic obfuscation steps using both global adversarial perturbations and Short Perturbation Blocks placed at random positions. Evaluation shows that Splash performs better traffic obfuscation than three other representative defenses, causing attacker classifiers to misclassify over 97% of traffic. In addition, it offers enhanced functionality by causing 45-60% of traffic to be misclassified into arbitrary target classes. Splash outperforms SOTA defenses such as AWA and ALERT against AT-aided attacks, reducing success rates to below 30%. Furthermore, it demonstrates significantly stronger resilience when attackers adopt the same defense configurations as users. Runsheng Ma, Chengshang Hou, Gaopeng Gou, Junzheng Shi, Zhen Li 0011, Gang Xiong 0001 |
ACSAC | 5 |
| 2025 | HDFG: Ethereum Smart Contract Honeypot Detection Based on Pre-Training TechniquesabstractIn recent years, a new fraud method, namely smart contract honeypots, has emerged on the famous blockchain platform Ethereum. The difference from smart contract vulnerabilities is that the contract honeypot essentially has no vulnerabilities, luring victims to call in a seemingly vulnerable form. However, the victims ultimately cannot obtain the desired benefits and will lose certain funds. Deep learning algorithms are preferred among current contract honeypot detection methods because they can learn more general characteristics and do not rely on expert experience. Most previous works use natural language models to learn the opcodes of contract honeypots but overlook the relevant structural features of the source code. We propose a novel method called the Smart Contract Honey-pot Data Flow Graph, which utilizes a data flow graph to extract the calling relationships of critical source code within contract honeypots and employs a pre-trained model for representation learning. First, contract honeypots generally have a code that transfers money to the calling address, which is critical information for constructing a source code data flow graph. Then, the pre-trained model is used to learn the source code representation and perform downstream classification tasks. The F1-score of our model significantly outperforms the state-of-the-art approaches in the contract honeypot classification task and is close to the highest performance in the detection task. In addition, this model is an end-to-end model that can detect unknown-type contract honeypots. Jiaying Song, Zhen Li 0011, Yingchao Qin, Bingxu Wang, Gang Xiong 0001, Hanwen Miao |
CSCWD | 2 |
| 2025 | ANASETC: Automatic Neural Architecture Search for Encrypted Traffic ClassificationabstractThe widespread adoption of encrypted network protocols has made traffic encryption ubiquitous, creating substantial challenges for network management and security. This paper introduces a novel encrypted traffic classification system, ANASETC, which combines traffic burst features with Neural Architecture Search (NAS) to automatically design efficient neural network architectures. ANASETC autonomously generates high-performance classification models, significantly reducing manual intervention while maintaining high classification accuracy. To enhance search efficiency, we introduce a new search space called ETNasnet, which optimizes the training process through parameter sharing among sub-models. We evaluate ANASETC’s performance on three public datasets and a real-world satellite network traffic dataset. The results show that ANASETC achieves an optimal balance between classification accuracy and search efficiency, demonstrating strong robustness and adaptability across various task scenarios, outperforming state-of-the-art methods. Ziqian Chen, Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Guangyan Huang |
ICASSP | 6 |
| 2025 | MalSE: Malware Detection Based on Multi-Dimensional API Call Sensitivity EstimationabstractMalware poses a significant threat to the security of cyberspace. For malware detection, utilizing machine learning or deep learning techniques to analyze API sequences has been proven to be effective. However, the existing methods fail in mitigating the interference caused by redundant information when processing excessively lengthy or behavior-masking sequences. To address this issue, we propose MalSE, a novel malware detection framework based on API call sensitivity estimation. MalSE aims to highlight key information in the sequence and minimize the interference of redundant information, thereby increasing detection performance. Firstly, MalSE uses a novel statistical-based method to annotate parameter sensitivity labels, which provide a foundation for subsequent module training. Secondly, MalSE employs a Bert-based estimator to transform the parameters into the semantic space and then predict the parameters’ sensitivities. Thirdly, MalSE assesses the sensitivity of each API call by aggregating the sensitivities of parameters, thus providing powerful features for detection tasks. Finally, MalSE employs an attention-based sequence model, which can concentrate on crucial information within the sequence to enhance the detection performance. We evaluate MalSE on 2 binary classification tasks and 1 multi-classification task. MalSE outperforms other methods across all tasks, demonstrating superior and robust detection capabilities under different scenarios. Ziqian Chen, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou, Haikuo Li |
IJCNN | 3 |
| 2025 | Exploring the Potential and Boundaries of KAN in Encrypted Traffic ClassificationabstractTraditional deep learning-based encrypted traffic classification models generally suffer from weak interpretability and struggle to effectively model the nonlinear relationships between features, resulting in poor generalization when faced with out-of-distribution data. Kolmogorov-Arnold Networks (KAN), known for their high interpretability and nonlinear modeling capabilities, have not been fully explored in the field of encrypted traffic. We proposed the KAN-ResNet model, drawing on the working principle of KAN, to enable end-to-end interpretable encrypted traffic classification by replacing different parts of ResNet with KAN layers and KAN convolutional layers. Experiments on the ISCX-Botnet-2014 and ISCX-Tor-2016 datasets demonstrated that our method outperforms traditional classification models (such as ViT) in terms of accuracy and F1 score. By visualizing the features extracted by the KAN network and the B-spline functions within the network, we interpreted the feature extraction of the KAN network from an interpretability perspective, showcasing its great potential in encrypted traffic classification. However, when the Residual Block organized by the KAN network is too deep, the increased training complexity and overfitting issues have become current limitations of its capabilities. Shituo Ma, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
IJCNN | 3 |
| 2025 | IPv6 Prefix Target Generation through Pattern and Distribution Learning using Vision-Transformer and Guided-Diffusion
Yaochen Ren, Gaopeng Gou, Chengshang Hou, Tianyu Cui, Zhen Li 0011, Gang Xiong 0001, Chang Liu 0049 |
INFOCOM | 5 |
| 2025 | PBC-MWF: Robust Multi-Tab Website Fingerprinting via Interval-Aggregated Packet-Burst CountsabstractWebsite fingerprinting (WF) attacks undermine the privacy promised by anonymizing networks such as Tor by inferring the websites a user visits from encrypted-traffic side-channels. Recent criticisms of the single-tab assumption have shifted attention to the more realistic multi-tab setting, where concurrent page loads create severe noise. Existing multi-tab studies rely on direction sequences that ignore temporal structure and therefore provide only limited discriminative power. Our experiments show that packet-level timestamps do carry extra signal, yet their raw form is fragile under overlapping tabs and timing-obfuscation defences. We propose the PacketBurst Counts (PBC) feature—a$4 \times L$matrix that, for each time interval, stores the counts of upstream packets, downstream packets, upstream bursts and downstream bursts. PBC preserves coarse temporal structure while discarding noisy finegrained timings, striking a balance between expressiveness and robustness. Building on PBC, we design PBC-MWF, an end-to-end framework that (i) uses a residual CNN to learn local embeddings and (ii) applies an adaptive sparse transformer to capture global correlations while suppressing tab-overlap noise. Unlike prior work, which evaluates on datasets with a fixed number of tabs and reports top-k accuracy, we additionally merge datasets with varying tab-count and perform threshold-based inference; the threshold is tuned on the validation set before testing. To the best of our knowledge, PBCMWF is the first WF framework to simultaneously address the multi-tab setting's challenges of fine-grained webpage identification and resilience against WF defences. Evaluations on three public multi-tab datasets demonstrate PBC-MWF's enhanced robustness: compared against nine baselines, it surpasses the best prior method-improving F1 by up to 12.4 % on site-level, 8.4 % on page-level, and over 10 % under defenses. Yuhao Wei, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou, Junzheng Shi, Yingchao Qin |
IPCCC | 2 |
| 2025 | 6RIS: IPv6 Address Correlation Attacks on TLS Encrypted Traffic Using Joint Representation of Interaction and Sequential BehaviorabstractIPv6 address correlation attacks determine whether two temporary addresses belong to the same user, compromising user privacy. Particularly, existing works have shown that methods based on TLS traffic analysis can be used to perform correlation attacks. However, they suffer from inaccurate differentiation of complex user behaviors and low correlation efficiency, leading to limitations in practical applications. In this paper, we propose a 6RIS model to improve IPv6 address correlation attacks on TLS-encrypted traffic. 6RIS learns the joint representation of interaction and sequential behavior from traffic, which is used to construct a KD-Tree for efficient correlation. Statistical aggregation and semantic preference modules are designed to extract generalized features from complex interaction behavior. To model sequential behavior, we utilize a sequence learning module to capture service dependencies, enhancing behavior representation. Experiments on a real-world IPv6 dataset show that 6RIS ($\mathbf{9 1. 8 6 \%}$TPR,$\mathbf{0. 8 3 \%}$FPR) outperforms state-of-theart methods. The correlation efficiency of 6RIS improves by at least 57 % compared to existing methods. Additionally, we further confirm through 6RIS that persistent session IDs in TLS session resumption can directly expose IPv6 temporary addresses to correlation attacks. Yang Li 0002, Chang Liu 0049, Gaopeng Gou, Tianyu Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
IWQoS | 6 |
| 2025 | Beneath the Heavens: A Thorough Measurement Study of the Starlink Terrestrial NetworkabstractThe emerging low earth orbit (LEO) satellite Internet has gained worldwide popularity. Starlink, a prominent LEO satellite Internet service provider, has attracted the most users due to its low latency, wide coverage, and strong usability. Current research on Starlink primarily focuses on the space segment and its impact on network performance. However, as an essential component, the architecture and unique features of the Starlink Terrestrial Network (SLTN) are not well-explored, which significantly affects the performance, security, and development of the whole network. In this paper, we fill this gap by conducting a thorough measurement study to profile the SLTN from various aspects and reveal its potential effects, with specific attention to the network assets, topology, and routing strategies. We developed a novel framework including active and passive measurement methods for collecting multiple network assets, tracing different route paths, and scanning active service of the SLTN. The open source intelligence was utilized for the first time to collect extensive real-user network status. Leveraging these techniques, we observed a rapid expansion of the Starlink service with the latest network assets, collected in January 2025, including over 23.5K/28 IP prefixes residing in 143 countries. We mapped the consistent topology of the global Starlink Internet access service and identified the specific IP addresses associated with the four types of network routing nodes. Multiple internal routing strategies were uncovered, which facilitate direct user-to-user interactions. Particularly, we revealed the switches of terrestrial infrastructure that users connected and the changes in routing strategies, which should be considered in future quality of service (QoS) evaluations. Our measurement also provides methods for improving users' perceptions and serves as a basis for studies like security risk evaluation and service discovery. Yanbo Wu, Mingxin Cui, Gaopeng Gou, Yuhao Wei, Gang Xiong 0001, Zhen Li 0011, Xinlei Ju |
IWQoS | 6 |
| 2025 | T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video RetrievalabstractText-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing work has primarily focused on extending CLIP knowledge for video-text tasks. However, videos typically contain richer information than images. In current video-text datasets, textual descriptions can only reflect a portion of the video content, leading to partial misalignment in video-text matching. Therefore, directly aligning text representations with video representations can result in incorrect supervision, ignoring the inequivalence of information. In this work, we propose T2VParser to extract multiview semantic representations from text and video, achieving adaptive semantic alignment rather than aligning the entire representation. To extract corresponding representations from different modalities, we introduce Adaptive Decomposition Tokens, which consist of a set of learnable tokens shared across modalities. The goal of T2VParser is to emphasize precise alignment between text and video while retaining the knowledge of pretrained models. Experimental results demonstrate that T2VParser achieves accurate partial alignment through effective cross-modal content decomposition. The code is available at https://github.com/Lilidamowang/T2VParser. Yili Li, Gang Xiong 0001, Gaopeng Gou, Xiangyan Qu, Jiamin Zhuang, Zhen Li 0011, Junzheng Shi |
ACM Multimedia | 6 |
| 2025 | Smart Contract Vulnerability Detection via Fusion of Sequence and Graph FeaturesabstractSmart contracts control critical financial assets on blockchains, with potential weaknesses risking substantial losses. Thus, smart contract vulnerability detection is essential for maintaining blockchain ecosystem stability. Traditional methods depend extensively on expert-driven patterns, resulting in poor scalability. Although deep learning-based approaches have made significant progress, they still suffer from issues such as inflexible representations, insufficient feature modalities, and limited model capabilities. In this paper, we propose FSGDec, a novel smart contract vulnerability detection framework that fuses sequential information and structural features at the bytecode level. Firstly, an efficient node embedding method is developed for contract control flow graphs, flexibly processing node sequences and incorporating node-specific semantic information associated with weaknesses. Then, by modeling node features as time series signals, an adaptive graph wave network is introduced to automatically capture vulnerability-related structural features. Finally, a classifier is deployed to perform bug detection utilizing the extracted graph-level features that integrate semantic information. Evaluated on two real-world smart contract datasets, the experimental results demonstrate that FSGDec achieves superior performance compared to state-of-the-art baselines. Haikuo Li, Gang Xiong 0001, Juwei Yue, Ziqian Chen, Gaopeng Gou, Zhen Li 0011 |
SMC | 7 |
| 2025 | FakeApp: A High-Precision Method for Domain Fronting Detection in Real Networks with Neuro-Symbolic IntegrationabstractDomain fronting is a covert communication technique, which evades detection by connecting with legitimate domains to imitate normal network traffic. But the imitation is flawed, so current detection methods usually treat domain fronting as abnormal traffic. However, these methods show very low precision due to normal-abnormal traffic imbalance in practice. In the paper, we find that domain fronting’s imitation is limited to popular applications (apps), such as Chrome and Firefox browser. Thus, we mitigate traffic imbalance by defining domain fronting detection as an app discrimination problem, rather than previous anomaly detection task.According to the revised definition, we propose FakeApp, a high-precision method for domain fronting detection in real networks. Using frequent item analysis, FakeApp first extracts the imitated app information from domain fronting tools as symbolic features. Then, through deep neural networks, it discriminates whether traffic belongs to genuine or spoofed apps. Finally, FakeApp integrates symbolic features and neural networks together to identify domain fronting in real networks. Evaluations over 2 million flows show that the precision of FakeApp is over 95%, far surpassing state-of-the-art methods on four domain fronting tools. These results also indicate that we have effectively mitigated the traffic imbalance issue. Yibo Xie, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
SMC | 4 |
| 2025 | SwCC: A Swapped-Contrastive Clustering Learning for Few-shot Website Fingerprinting AttacksabstractWebsite fingerprinting (WF) attacks exploit distinctive traffic patterns to identify the specific web page a user visits over anonymized connections. While traditional WF attacks have achieved impressive results, they are typically evaluated in abundant labeled data settings and assume that website traffic features remain static. This assumption is often unrealistic in real-world scenarios. Recent methods either rely on deep learning, which still requires large amounts of labeled data, or employ self-supervised pre-training to ease this demand. However, they leave clustering information crucial for few-shot WF attacks underexplored, leaving ample room for performance gains. In this paper, we propose a novel self-supervised pre-training model for few-shot WF attacks, called Swapped Contrastive Clustering (SwCC). SwCC proposes a comprehensive and principled data augmentation scheme, combining Tor-tailored transformations with statistical procedures to foster robust and discriminative feature learning for WF attacks. SwCC further introduces an innovative dual-level contrastive learning framework that jointly leverages instance-level and prototype-based objectives, which can further perform latent-space clustering on extracted features in WF attacks. Our pre-trained model can be fine-tuned in few-shot learning scenarios and achieves state-of-the-art(SOTA) performance on few-shot WF attack tasks. Under a 5-shot learning setting in a closed-world scenario, our SwCC achieves up to 83.5% accuracy when the evaluation traces are collected from an environment unseen by the WF adversary, outperforming the SOTA methods. Gaopeng Gou, Wenqi Dong, Gang Xiong 0001, Zhen Li 0011, Qingya Yang |
TrustCom | 6 |
| 2025 | SSRCorr: A Self-Supervised Robust Flow Representation Learning Framework for Flow Correlation Attacks on TorabstractTor is one of the most widely adopted anonymity networks, yet its anonymity can be undermined by adversaries through flow correlation attacks. Current mainstream technologies focus on exploiting the sequence characteristics of packet lengths and timestamps to execute attacks. However, the padding mechanism of the Tor network and time delays caused by multi-hop relays obscure these single-modal features. Additionally, the diversity of network services and the randomness of user behavior result in sparse packet distributions, which impact model training and inference. In this paper, we propose SSRCorr, a novel self-supervised learning framework for flow correlation attacks, incorporating the Flow Feature Aggregation (FFA) module and Global-Local Fusion (GLoF) Encoder to address these challenges. Firstly, we construct a Byte-based Traffic Aggregation Matrix (BTAM) by integrating time and length sequences and applying two data augmentation methods tailored for Tor flow correlation, thereby reducing the impact of Tor network noise on attack effectiveness. Secondly, we employ GLoF to extract features from the output by FFA and fuse the global context information of the traffic, thus mitigating the impact of low-information traffic on model performance. Experiments show that SSRCorr achieves a TPR of 96%, surpassing other methods, and maintains robust performance under temporal drift and obfuscation, supporting future research on countering anonymity system defenses. Mengyan Liu, Yaochen Ren, Yanbo Wu, Yangyang Guan, Zhen Li 0011, Gaopeng Gou, Junzheng Shi |
TrustCom | 7 |
| 2025 | DecETT: Accurate App Fingerprinting Under Encrypted Tunnels via Dual Decouple-based Semantic EnhancementabstractDue to the growing demand for privacy protection, encrypted tunnels have become increasingly popular among mobile app users, which brings new challenges to app fingerprinting (AF)-based network management. Existing methods primarily transfer traditional AF methods to encrypted tunnels directly, ignoring the core obfuscation and re-encapsulation mechanism of encrypted tunnels, thus resulting in unsatisfactory performance. In this paper, we propose DecETT, a dual decouple-based semantic enhancement method for accurate AF under encrypted tunnels. Specifically, DecETT improves AF under encrypted tunnels from two perspectives: app-specific feature enhancement and irrelevant tunnel feature decoupling. Considering the obfuscated app-specific information in encrypted tunnel traffic, DecETT introduces TLS traffic with stronger app-specific information as a semantic anchor to guide and enhance the fingerprint generation for tunnel traffic. Furthermore, to address the app-irrelevant tunnel feature introduced by the re-encapsulation mechanism, DecETT is designed with a dual decouple-based fingerprint enhancement module, which decouples the tunnel feature and app semantic feature from tunnel traffic separately, thereby minimizing the impact of tunnel features on accurate app fingerprint extraction. Evaluation under five prevalent encrypted tunnels indicates that DecETT outperforms state-of-the-art methods in accurate AF under encrypted tunnels, and further demonstrates its superiority under tunnels with more complicated obfuscation. Project page: https://github.com/DecETT/DecETT Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
WWW | 7 |
| 2025 | HoleMal: A lightweight IoT malware detection framework based on efficient host-level traffic processing
Ziqian Chen, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou, Haikuo Li, Junchao Xiao |
Comput. Secur. | 3 |
| 2025 | Respond to Change With Constancy: Instruction-Tuning With LLM for Non-I.I.D. Network Traffic ClassificationabstractEncrypted traffic classification is highly challenging in network security due to the need for extracting robust features from content-agnostic traffic data. Existing approaches face critical issues: (i) Distribution drift, caused by reliance on the closedworld assumption, limits adaptability to real-world, shifting patterns; (ii) Dependence on labeled data restricts applicability where such data is scarce or unavailable. Large language models (LLMs) have demonstrated remarkable potential in offering generalizable solutions across a wide range of tasks, achieving notable success in various specialized fields. However, their effectiveness in traffic analysis remains constrained by challenges in adapting to the unique requirements of the traffic domain. In this paper, we introduce a novel traffic representation model named Encrypted Traffic Out-of-Distribution Instruction Tuning with LLM (ETooL), which integrates LLMs with knowledge of traffic structures through a self-supervised instruction tuning paradigm. This framework establishes connections between textual information and traffic interactions. ETooL demonstrates more robust classification performance and superior generalization in both supervised and zero-shot traffic classification tasks. Notably, it achieves significant improvements in F1 scores: APP53 (I.I.D.) to 93.19%(6.62%↑) and 92.11%(4.19%↑), APP53 (O.O.D.) to 74.88%(18.17%↑) and 72.13%(15.15%↑), and ISCX-Botnet (O.O.D.) to 95.03%(9.16%↑) and 81.95%(12.08%↑). Additionally, we construct NETD, a traffic dataset designed to support dynamic distributional shifts, and use it to validate ETooL’s effectiveness under varying distributional conditions. Furthermore, we evaluate the efficiency gains achieved through ETooL’s instruction tuning approach. Gang Xiong 0001, Gaopeng Gou, Wenqi Dong, Jing Yu 0007, Zhen Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | TMGAN: A GAN-Based Traffic Morphing Defense Against Website FingerprintingabstractWith the rapid growth of encrypted traffic, methods that use side-channel information to monitor online user behavior have emerged, known as Website Fingerprinting (WF) attacks. These attacks pose a significant threat to the privacy of users’ online activities. To address the threat posed by various WF attacks to network behavior privacy, current methods lack defenses based on the source/target misclassification of traffic. Our goal is to design a WF defense method that transforms the side-channel information of the given original category traffic into that of another category to counter existing DNN-based WF attacks. We utilize adversarial examples and employ a Generative Adversarial Network (GAN) incorporating a WF model to generate perturbations. These perturbations are overlaid onto the given traffic, morphing it into traffic of another category, thereby enhancing the privacy of network behavior. We refer to the proposed method as the Traffic Morphing Generative Adversarial Network (TMGAN). Experimental results demonstrate that this method effectively counters most WF attacks, enhancing the privacy of users’ online behavior. In white-box scenarios, it achieves average perturbation rates of 90-99% and morphing rates of 31-75%. In black-box scenarios, it achieves average perturbation rates of 95-98% and morphing rates of 5-28%. Shukan Huang, Junchao Xiao, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
HPCC | 5 |
| 2024 | RecoSelector: Cost-Sensitive Feature Selection for Network Intrusion Detection in Resource-Constrained Internet of ThingsabstractDetecting malware in Internet of Things (IoT) networks is crucial for ensuring IoT security. Machine learning based Network Intrusion Detection System (NIDS) has been proven to be effective, but it faces the challenge of achieving high computational efficiency. Previous feature selection methods improve the efficiency of NIDS by removing redundant features. However, these methods fail to consider the significant disparity in computational cost among different traffic features, so they are not fully applicable for resource-constrained environment. To address this issue, we propose a novel framework RecoSelector to select traffic features in a cost-sensitive manner for NIDS. RecoSelector aims to effectively select feature subsets with strong detection capability and low computational cost. Firstly, we quantify and analyze the feature construction cost among IoT short flows, IoT long flows and cross platform flows from 6 scenarios. Based on the flows, we generate computational-loss of 69 flow features through flow transmission frequency. Secondly, we propose Particle Initialization based on Orthogonal Sparse Vector (PIOSV) to optimize search direction and increase the possibility of finding a global optimal solution. Finally, we design an elaborate fitness function RecoFitness, aiming to carry out multi-objective optimization, for iterative selection. We obtain flow feature cost in a real resource-constrained environment. Experiments demonstrate that RecoSelector exhibits superior performance with spending only 3% to 60% of the computational time cost while achieving comparable F1 scores compared to existing methods. Ziqian Chen, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
IPCCC | 3 |
| 2024 | Smart Contract Vulnerability Detection Based on AST-Augmented Heterogeneous GraphsabstractSmart contracts have been increasingly deployed and applied on various blockchain platforms. Nevertheless, vulnerabilities may cause significant financial losses due to the involvement of substantial funds in smart contracts. Traditional analysis tools heavily rely on manually predefined rules. Recent studies have demonstrated the promising potential of deep learning techniques in smart contract vulnerability detection. However, existing approaches often disregard cross-function and cross-contract vulnerability scenarios, focusing primarily on characterization or detection tasks at the function level. In this study, we propose CL-HGAN, a novel framework for smart contract vulnerability detection at the contract level. Firstly, we construct a contract-level heterogeneous graph to embody the relationships between contracts and functions. Specifically, we build the backbone of the heterogeneous graph based on the abstract syntax tree (AST) and multiple types of edges and then incorporate two additional categories of edges to augment its structural information. Subsequently, we design a two-phase feature learning method to automatically generate graph-level representations based on a heterogeneous graph attention network and meta-paths specific to the constructed graph. Finally, we employ a classifier to perform vulnerability detection tasks. In particular, the proposed CL-HGAN comprehensively captures vulnerability features and accurately identifies vulnerabilities at the contract level. Furthermore, we evaluate the CL-HGAN framework on an Ethereum smart contract dataset containing thirty types of vulnerabilities. The experimental results show that the average metrics of our approach outperform the state-of-the-art baselines. Haikuo Li, Gang Xiong 0001, Chengshang Hou, Gaopeng Gou, Ziqian Chen, Zhen Li 0011 |
IPCCC | 6 |
| 2024 | HoneyRank: A Low-Cost Discovering Method of 0-Day Ethereum Smart Contract HoneypotsabstractOne attack method that actively deploys smart contract honeypots has recently become popular. A contract honeypot is a smart contract that pretends to have vulnerabilities, enticing victims who call the contract to lose funds. However, previous works detected contract honeypots by individual characteristics, such as codes and ledger details. They overlooked the connection between the two parties in the transaction. Therefore, we propose the HoneyRank algorithm, which uses known honeypots as initial seeds to construct a contract honeypot transaction relationship network (HoneyNet) and source code text similarity detection to discover 0-day honeypots that previous work missed in the same detected block height range. This low-cost method detects only a few highly suspicious smart contracts and does not require machine learning training. Specifically, we trace transaction history data to collect the accounts and relationships of honeypot seeds, attackers, and victims and construct a HoneyNet. Based on transaction behavior inference, we label and calculate the source code similarity between high-risk smart contracts and ground truth honeypots. Finally, we select the high-similarity smart contracts to confirm honeypots manually. Besides, we analyze the criminal associations in a HoneyNet. As far as we know, we are the first to construct a HoneyNet and use it to find new honeypots. These honeypots visually reveal the potential connections between the attackers (creators of the honeypot) and the victims. We discovered 54 0-day honeypots never found by previous methods and mined 11 attacker communities composed of attackers and puppet accounts for the first time. Jiaying Song, Zhen Li 0011, Gaopeng Gou, Bingxu Wang, Gang Xiong 0001, Yingchao Qin |
MSN | 2 |
| 2024 | OSN Bots Traffic Transformer : MAE-Based Multimodal Social Bots Behavior Pattern MiningabstractIn recent years, online social networks (OSN) have rapidly gained popularity worldwide, becoming important platforms for information dissemination. Cyber manipulators use OSN bots to disseminate harmful information and manipulate public opinion, which can engage in cyber violence and conduct financial crimes. Therefore, it is crucial to propose an effective detection solution for OSN bots as a matter of urgency. Different OSN bots exhibit distinct behavioral patterns compared to regular users due to varying behavioral preferences. Analyzing network behavior patterns can reveal the fundamental rules and anomalies of OSN bots, providing support for effective detection in order to gather evidence of any illegal activities. Traditional social bot detection methods based on user profiles or social relationships pose risks of infringing on user privacy. Therefore, we propose a new detection framework for OSN bots——OBTT model, which demonstrates significant advantages in identifying bot traffic to OSN and discovering behavior patterns of different types of bots. OBTT adopts a multimodal approach, integrating graph embeddings from raw traffic with sequential features, while incorporating temporal information to explore the regularities in bot action sequences. Using large-scale unlabeled data, we pretrain a Masked Autoencoder (MAE) and fine-tune it with a small amount of labeled data to enhance the model capacity to detect various bot behavior patterns. Experiments conducted on our OSNBotTraffic5 dataset show that OBTT achieved an accuracy of 0.95, demonstrating excellent performance. Notably, this is the first time that different OSN bot behavior patterns have been identified in quasi-real time from the perspective of network traffic. Haonan Zhai, Ruiqi Liang, Zhen Li 0011, Bingxu Wang, Qingya Yang |
TrustCom | 4 |
| 2024 | Identifying malicious traffic under concept drift based on intraclass consistency enhanced variational autoencoder
Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Binxing Fang |
Sci. China Inf. Sci. | 5 |
| 2024 | Let gambling hide nowhere: Detecting illegal mobile gambling apps via heterogeneous graph-based encrypted traffic analysis
Gaopeng Gou, Chang Liu 0049, Zhen Li 0011, Gang Xiong 0001 |
Comput. Networks | 6 |
| 2024 | A blind flow fingerprinting and correlation method against disturbed anonymous traffic based on pattern reconstruction
Chang Liu 0049, Gaopeng Gou, Zhen Li 0011, Gang Xiong 0001, Yangyang Ding, Chengshang Hou |
Comput. Networks | 4 |
| 2024 | DomEye: Detecting network covert channel of domain fronting with throughput fluctuation
Yibo Xie, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
Comput. Secur. | 4 |
| 2024 | Traffic spills the beans: A robust video identification attack against YouTube
Gang Xiong 0001, Zhen Li 0011, Gaopeng Gou, Binxing Fang |
Comput. Secur. | 3 |
| 2023 | TGC: Transaction Graph Contrast Network for Ethereum Phishing Scam DetectionabstractPhishing scams have become the most serious type of crime involved in Ethereum. However, existing methods ignore the natural camouflage and sparse distribution of phishing scams in Ethereum leading to unsatisfactory performance, and they are also limited by the data scale which cannot be applied to real-world dynamic scenarios. In this paper, we propose a Transaction Graph Contrast network (TGC) to enhance phishing scam detection performance on Ethereum. TGC inputs subgraphs instead of the entire graph for training, which eases the model’s requirements for machine configuration and data connectivity. Motivated by phishing nodes are surrounded by normal nodes, we design the comparison between node-level to help phishing nodes learn the unique properties of themselves different from their neighbors. Observing the small number and sparse distribution of phishing nodes, we narrow the distance between phishing nodes by comparing node context-level structures, so as to learn universal transaction patterns. We further combine the obtained features with common statistics to identify phishing addresses. Evaluated on real-world Ethereum phishing scams datasets, our TGC outperforms the state-of-the-art methods in detecting phishing addresses and has obvious advantages in large-scale and dynamic scenarios. Gaopeng Gou, Chang Liu 0049, Gang Xiong 0001, Zhen Li 0011, Junchao Xiao, Xinyu Xing 0001 |
ACSAC | 5 |
| 2023 | PTC: Prompt-based Continual Encrypted Traffic ClassificationabstractEncrypted traffic classification (ETC) is necessary for network security, which is the process of identifying encrypted network traffic into a specific class, thus there are numerous applications in the security of network. The rapid development of network web services (applications) makes it attractive to tackle classification of encrypted traffic in a continual learning environment. However, the traffic ambiguity and privacy leakage, restrict existing incremental approaches from achieving satisfactory results in the traffic. We introduce a prompt-based continual encrypted traffic classification method (PTC) in this research to progressively learn tasks under multiple process transitions. Prompts are tiny, learnable parameters that are stored in ram according to our suggested structure. The objective is to find the best way to use prompts to help models make predictions, keep both task-peculiar and task-constant knowledge in model, and prevent catastrophic forgetting. We carry out extensive tests using both real-world and open datasets. PTC method can strengthen the existing offline traffic classification works, make them adapt to online scenarios, and outperforms the SOTA online traffic classification method in three datasets. (by 4.54 %, 7.28 % and 11.68 % on three datasets, respectively) Wei Cai 0007, Chengshang Hou, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
CSCWD | 6 |
| 2023 | Identifying DoH Tunnel Traffic Using Core Feathers and Machine Learning MethodabstractDNS protocol is a plaintext domain name resolution protocol, which has the risk of privacy disclosure. DNS over HTTPS (DOH) protocol is designed to encrypt DNS traffic, which solves the privacy problem. However, many network attackers use the DOH tunnel for malicious transmission. From the passive traffic, there is no obvious difference between normal DOH traffic and DOH tunnel traffic, which brings great challenges to identify them. At present, researches mainly focus on the plaintext DNS covert tunnel, but less on the encrypted DOH tunnel. In this paper, we propose DOH covert tunnel detection method based on core features and machine learning method using two steps. Firstly, we detect DOH traffic according to the threshold of features. On this basis, we use core features and machine learning methods to detect tunnel traffic in all DOH traffic. Finally, we use self collected and public datasets to verify our method. The results show that the method achieves up to 99 % precision and recall that is superior to state of the art method. Bingxu Wang, Gang Xiong 0001, Gaopeng Gou, Jiaying Song, Zhen Li 0011, Qingya Yang |
CSCWD | 5 |
| 2023 | Covertness Analysis of Snowflake Proxy RequestabstractSnowflake is a special proxy system against IP-based network blocking. As its IP addresses refresh frequently, faster than IP blacklist’s update, users can exploit it to access blocked websites. To block snowflake, existing methods focus on detecting snowflake proxies. But they are susceptible to various factors, for example, proxy’s location and version. In the paper, we propose a new manner to block snowflake. We observe that to adapt fast IP changes, users need to request latest proxies from proxy database before using snowflake. Thus, adversaries can block snowflake by detecting proxy request instead of proxy itself. To verify our method, we analyse covertness of snowflake proxy requests, that has been protected by imitating normal web requests. After comparing with typical web requests, we find the imitation is vulnerable in packet size, direction, time and network speed, such as, the latency time is higher than normal obviously. Using the four vulnerabilities, we train machine learning algorithm to detect snowflake proxy requests in reality. Experimental results demonstrate that proxy request can be detected accurately across different versions at the beginning of connection. In conclusion, our work paves a new way to block snowflake. Yibo Xie, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui |
CSCWD | 4 |
| 2023 | Multi-Feature Fusion Based Approach for Classifying Encrypted Mobile Application TrafficabstractWith rapid development of mobile Internet, a great number of mobile applications has emerged, presenting a great explosion in mobile Internet traffic. Therefore, accurate classification of application traffic is necessary to more effectively manage mobile Internet traffic. However, the encryption of mobile application traffic gradually eliminates traditional classification approaches based on specific signatures, greatly increasing the difficulty of the classification of mobile application traffic. Therefore, we propose a novel multi-feature fusion (MFF)- based approach to enhance the accuracy of mobile application traffic classification. We also extract packet length sequence, byte sequence, statistical feature, etc. Then, we perform weighted fusions of features based on Relief-F algorithm to achieve the best set of features. Finally, we use machine learning techniques for application classification. Compared to several other feature extraction methods, MFF achieves an excellent performance with an accuracy of 97.6% for 16 mobile applications and a F1-score of over 99% for VPN-nonVPN. Qingya Yang, Peipei Fu, Junzheng Shi, Bingxu Wang, Zhen Li 0011, Gang Xiong 0001 |
CSCWD | 5 |
| 2023 | Analysing Covertness of Tor Bridge RequestabstractTor bridges are hidden entrances of Tor network. Users can exploit bridges to hide their visits of Tor. To restrict hidden Tor visits, many attacks focus on bridge information discovery or bridge traffic detection. But these attacks are less effective because bridges' information cannot be discovered thoroughly and its traffic are often obfuscated. In the paper, we present a novel attack to stop hidden Tor visits. We observe that users need to request information of bridges from a database before visiting Tor network. Thus, attackers can stop Tor visits by detecting the process of bridge requests rather than bridge itself. To verify our attack's feasibility, we analyse covertness of the most widely-used bridge request tool, which imitates normal network request when communicating with bridge database. After comparing with five types of typical web request, we find that this tool fails to imitate in packet time, size and direction, for example, the variation of simulated packet sizes are more dynamic than normal. Based on the three imitation vulnerabilities, we train machine learning algorithms to detect bridge request. Extensive experiments demonstrate that bridge request can be identified with high accuracy and very low false-positive rates in real-world. In conclusion, our work paves a new way to block evasive Tor visits. Yibo Xie, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui |
ICC | 4 |
| 2023 | FedMP: Robust and Communication-Efficient Federated Multi-Prototype Intrusion Detection Framework in IoTabstractDue to its excellent performance in privacy protection, federated learning (FL) technology is gradually introduced into the IoT environment to build a distributed intrusion detection framework. However, the previous frameworks have two limitations: 1) high communication overhead caused by the frequent exchange of model parameters is not friendly for resource-constrained IoT devices; 2) single global model hardly handles not independent and identically distributed (Non-IID) intrusion data on different IoT clients. In this paper, we propose a Federated Multi-Prototype intrusion detection framework (FedMP) to address the above limitations. Specifically, FedMP includes a k-means clustering module that extracts low-dimensional local prototypes for IoT clients and a novel aggregation algorithm that fairly aggregates all local prototypes on the central server to generate global prototypes. By exchanging prototypes instead of model parameters between IoT clients and the central server, FedMP aims to train a unique personalized model for each IoT client to adapt to its local intrusion detection tasks. Experimental results on real-world intrusion detection datasets show that FedMP achieves the best detection performance in Non-IID scenarios while significantly reducing communication overhead compared to other state-of-the-art methods. Minsheng Le, Zhen Li 0011, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001 |
ICPADS | 2 |
| 2023 | FA-Net: More Accurate Encrypted Network Traffic Classification Based on Burst with Self-AttentionabstractEncrypted network traffic classification (ENTC) is crucial in fields including network cyberspace security, network administration and service quality. Combining the machine learning algorithms with manual-designed burst features has been studied extensively in the ENTC community. However, these features depend on professional experience heavily, which needs lots of human effort. These hand-crafted features are task-oriented and incomplete in various complex tasks. What's more, they are also affected by the potential network jitters. In this paper, we propose a novel encrypted traffic classification method FA-Net to mine burst features. We adopt two hierarchical multi-head self-attention encoders to enumerate all potential intra-burst features and inter-burst dependencies completely, and select the optimal associations automatically. For more robust against network jitter, we design an additional burst positional encoding to loose the model's sensitivity about out-of-order packets within bursts. We evaluate the FA-Net on multiple datasets, including website and mobile application classification tasks. The results show the FA-Net model outperforms other state-of-the-art methods in all the datasets, even gains more than 5% absolute improvement in accuracy. Additionally, the quantitative measurements about burst feature similarity show that the burst features learned by FA-Net exhibits more intraclass similarity and more inter-class separation. Mingxin Cui, Chengshang Hou, Wei Cai 0007, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
IJCNN | 5 |
| 2023 | The Potential Utility of Image Descriptions: User Identity Linkage across Social Networks Based on MultiModal Self-Attention FusionabstractThe task of user identity linkage across social networks aims to predict whether users from different social networks refer to the same person. This task plays a crucial role in cross-social network information dissemination and intelligent recommendations. However, existing user identity linkage tasks suffer from several challenges: 1) excessive reliance on social network topology, neglecting users’ visual modality information; 2) inadequate handling of noise in user feature data; and 3) ineffective fusion of users’ multimodal information. To address these issues, we investigated a method that utilizes heterogeneous multimodal posts, including user-generated text, images, and check-in messages, to achieve user identity linkage across social networks. We innovatively leveraged a pre-trained model for image-to-text conversion to further explore users’ image data and proposed an adversarial learning model based on the multimodal self-attention mechanism (AMSA). The AMSA model consists of four components: user feature extraction, user feature processing, user feature fusion, and adversarial learning. Specifically, AMSA initially employed advanced pre-trained models to extract features from multiple modalities of users, including images and text. Subsequently, it utilized multiple mechanisms, such as multi-head self-attention, to process data from each modality separately and then fused them into user representation vectors. Finally, AMSA employed adversarial learning to enhance the model’s learning capacity and mitigate semantic disparities in user information across different platforms. We conducted model performance evaluations on publicly available datasets, and experimental results demonstrated the superiority of the proposed AMSA model. Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui |
IPCCC | 4 |
| 2023 | Identifying Exposed ICS Remote Management Device using Multimodal Feature in the WildabstractIndustrial Control System (ICS) devices with Internet-accessible IP addresses are critical to the smooth functioning of industries, power grids, and other critical infrastructures. Previous methods used to identify ICS devices exposed to the Internet often ignored these remotely managed devices. Specifically, these systems, which do not openly provide ICS-specific port services, remain undetected during Internet-wide scans for such services. The existing method for scanning and discovering this part of remote management devices has a single feature extraction, and discovering such remote management devices is inefficient. In this paper, we propose a novel strategy dedicated to identifying exposed remote managed devices on the Internet by using multidimensional approaches, such as traffic periodicity analysis, device customized field identification, key content extraction via image-to-text conversion, and remote management device access HTTP traffic feature analysis. We have effectively identified 26 different types of remote management devices in Japan, comprising a total of 983 exposed devices, in a shorter timeframe. When juxtaposed with previous methods, our strategy has identified more devices faster. Therefore, our method holds considerable potential for identifying and reducing the attack surface of critical infrastructures on the Internet. Furthermore, it also has substantial significance for protecting global network security. Liuxing Su, Gaopeng Gou, Zhen Li 0011, Gang Xiong 0001, Chengshang Hou |
IPCCC | 4 |
| 2023 | Bitcoin Mixing Service Detection Based on Spatio-Temporal Information Representation of Transaction GraphabstractCoin mixing is a technique used to enhance Bitcoin’s anonymity and can be used to obfuscate the relationship among transaction input addresses. Due to this property, much of the criminal activity on Bitcoin uses coin-mixing techniques to launder money, making these illicit funds difficult to trace. Therefore, it is important to implement the detection of Bitcoin mixing services. Several methods for identifying bitcoin mixing services have been proposed, but balancing their efficiency and generality at the same time is a challenging task. In this paper, We propose STMD (Spatio-Temporal Mixing Detector), which combines local features and global features of Bitcoin transactions to identify coin-mixing transactions. On one hand, we extract and process the statistical features of neighboring nodes of the transaction as local features. On the other hand, we construct a global position encoding (GPE) containing spatio-temporal information of the transaction as global features. Additionally, we employ the attention mechanism to handle these two types of features, effectively combining them. Finally, we utilize linear layers to achieve the detection of coin-mixing transactions. The experimental results show that STMD performs better than existing methods on the same dataset; it also has a higher recall on the test set of other types of coin-mixing transactions, which reflects the generality of the model. In particular, We apply local features and global features for experiments separately and verify the necessity of the two features. The results of the model trained using only global features also outperform the existing methods, which shows that the global position encoding (GPE) we constructed is effective for mixed currency transaction identification. Hanzhi Yang, Gaopeng Gou, Junzheng Shi, Gang Xiong 0001, Zhen Li 0011 |
IPCCC | 6 |
| 2023 | MENDER: Multi-level Feature Fusion Discovery Framework for Exposed ICS Remote Management Devices in the WildabstractWith the development of the Internet, many industrial control system (ICS) remote management devices for key infrastructure, such as solar power plants, sewage treatment, and buildings, are easily exposed to the Internet through network connections. Existing studies on ICS detection can not detect these remote management devices, which are not open to specific industrial control protocol services. Effectively identifying exposed real-world ICS remote management devices while minimizing the attack surface remains an enormous challenge. To address this challenge, we propose a Multi-level fEature fusioN DiscovEry fRamework (MENDER) for discovering neglected ICS remote management devices. First, we conduct a comprehensive and multi-level data collection in the detection process, including the traffic generated by website access, web resource files and HTML. We build an efficient and comprehensive data detection and acquisition module. Second, we design a novel multi-level feature extraction and fusion model to mine key features from raw data. We perform hierarchical clustering based on HTML features and combine the extracted multi-layered key features to filter potential ICS remote management devices. Third, we use the Random Forest model to classify and predict ICS devices based on the extracted multi-level features, aiming to learn inherent features profoundly for enhanced detection of these remote management devices. In a month, we detect 1, 069 devices in Japan, some of devices are insecure, i.e. allowing access to the status or even the control industrial devices without proper authentication. Compared with existing method, MENDER’s time spent on device discovery has been reduced by 94.1%, the number of device discovery is increased by 20.1%, and 26 different types of devices are found. Our MENDER’s device discover ability is superior to it both in time and quantity. Liuxing Su, Gaopeng Gou, Zhen Li 0011, Gang Xiong 0001, Chengshang Hou |
TrustCom | 4 |
| 2023 | Zero-relabelling mobile-app identification over drifted encrypted network traffic
Mingxin Cui, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
Comput. Networks | 6 |
| 2023 | Few-shot encrypted traffic classification via multi-task representation enhanced meta-learning
Gang Xiong 0001, Junzheng Shi, Gaopeng Gou, Zhen Li 0011, Chang Liu 0049 |
Comput. Networks | 6 |
| 2023 | FlowTracker: Improved flow correlation attacks with denoising and contrastive learning
Chang Liu 0049, Gang Xiong 0001, Zhen Li 0011, Gaopeng Gou |
Comput. Secur. | 4 |
| 2022 | Shoot Before You Escape: Dynamic Behavior Monitor of Bitcoin Users via Bi-Temporal Network Analytics
Jianing Ding, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
ACISP | 4 |
| 2022 | Anomaly Detection in Encrypted Identity Resolution Traffic based on Machine LearningabstractIdentity resolution is an emerging network resource widely applied in Industrial Internet of Things. Although encryption improves the privacy of identity resolution, it also challenges DPI-based anomaly detection. Therefore, it is imperative to recognize and supplement the encrypted information of IDS. In this paper, we design a machine learning-based framework to automatically extract critical information of identity resolution system from network traffic. According to the characteristics of traffic, we use the hybrid feature of statistics and sequences to describe encrypted traffic. Besides, a supervised classification algorithm is applied to explore the effective classification of two communication processes, which are service attribution information for node addressing and operation behavior for data management. We tested this method based on the encrypted traffic collected from a realistic identity resolution system. The results indicate that our approach exhibits good performance, outperforms related works, and can be applied in resource-constrained industrial scenario. This is the first work analysing the identity resolution system from the perspective of traffic analysis. Zhishen Zhu, Qingya Yang, Chonghua Wang, Zhen Li 0011 |
QRS | 5 |
| 2022 | BSBA: Burst Series Based Approach for Identifying Fake Free-trafficabstractIn recent years, mobile traffic has gradually become a major part of network traffic. To attract customers, mobile network operators provide free-traffic, which is a preferential policy that is free of charge for specific application traffic. Since the emergence of free-traffic, fake free-traffic also appeared soon. Fake free-traffic is a malicious behavior, which helps attackers illegally use network resources and evade network resource charging. The appearance of fake free-traffic maliciously harms the interests of operators and disrupts the rules of network resource charging. Because of the uniqueness of free-traffic, it encapsulates a layer of the HTTP protocol in addition to the actual application communication protocol, existing studies on encrypted traffic analysis are not applicable to identify fake free-traffic. In this paper, we propose Burst Series Based Approach (BSBA), a novel method for identifying fake free-traffic. The key idea behind BSBA is to construct effective features by capturing the differences of burst series among fake free-traffic, free-traffic and non-free traffic, and combine the constructed features with machine learning algorithms to identify fake free-traffic. We collect a real-world traffic dataset and conduct evaluations to verify the effectiveness of the BSBA. Experiment results demonstrate that the BSBA achieves excellent performances (96.82% Accuracy, 96.46% Precision, 96.57% Recall and 96.51% F1-score) and is superior to the state-of-the-art methods. Chang Liu 0049, Zhen Li 0011, Qingya Yang, Anlin Xu, Gaopeng Gou |
WoWMoM | 3 |
| 2022 | ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationabstractEncrypted traffic classification requires discriminative and robust traffic representation captured from content-invisible and imbalanced traffic data for accurate classification, which is challenging but indispensable to achieve network security and network management. The major limitation of existing solutions is that they highly rely on the deep features, which are overly dependent on data size and hard to generalize on unseen data. How to leverage the open-domain unlabeled traffic data to learn representation with strong generalization ability remains a key challenge. In this paper, we propose a new traffic representation model called Encrypted Traffic Bidirectional Encoder Representations from Transformer (ET-BERT), which pre-trains deep contextualized datagram-level representation from large-scale unlabeled data. The pre-trained model can be fine-tuned on a small number of task-specific labeled data and achieves state-of-the-art performance across five encrypted traffic classification tasks, remarkably pushing the F1 of ISCX-VPN-Service to 98.9% (5.2%↑), Cross-Platform (Android) to 92.5% (5.4%↑), CSTNET-TLS 1.3 to 97.4% (10.0%↑). Notably, we provide explanation of the empirically powerful pre-training model by analyzing the randomness of ciphers. It gives us insights in understanding the boundary of classification ability over encrypted traffic. The code is available at: https://github.com/linwhitehat/ET-BERT. Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Jing Yu 0007 |
WWW | 4 |
| 2022 | Accurate mobile-app fingerprinting using flow-level relationship with graph neural networks
Zhen Li 0011, Peipei Fu, Wei Cai 0007, Mingxin Cui, Gang Xiong 0001, Gaopeng Gou |
Comput. Networks | 2 |
| 2022 | Privacy protection of China's top websites: A Multi-layer privacy measurement via network behaviours and privacy policies
Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
Comput. Secur. | 3 |
| 2021 | BAPM: Block Attention Profiling Model for Multi-tab Website Fingerprinting Attacks on TorabstractWebsite fingerprinting attacks on Tor pose an security issue in anonymity privacy, in which attackers can identify websites visited by victims through passively capturing and analyzing encrypted packet traces. Although related works have been studied over a long period, most of them focus on single-tab packet traces which only contain one page tab’s data. However, users often open multiple page tabs successively when browsing the web, and multi-tab packet traces generated will corrupt common single-tab attacks. Existing multi-tab attacks still depend on an elaborate feature engineering, besides, they fail to exploit the overlapping area which contains the mixed data of two adjacent page tabs, thus suffering from the information lost or confusion. In this paper, we propose a Block Attention Profiling Model named BAPM as a new multi-tab attacking model. Specifically, BAPM fully utilizes the whole multi-tab packet trace including the overlapping area to avoid information lost. It generates a tab-aware representation from direction sequences and performs the block division to separate mixed page tabs as clearly as possible, thus relieving the information confusion. Then the attention-based profiling is used to group blocks belonging to the same page tab and finally multiple websites are simultaneously identified under a global view. We compare BAPM with state of the art multi-tab attacks, and BAPM outperforms comparison methods even with larger overlapping area. The effectiveness of model design is also validated through ablation, sensitivity and generalization analysis. Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Mingxin Cui, Chang Liu 0049 |
ACSAC | 4 |
| 2021 | Multi-scene Classification of Blockchain Encrypted Traffic
Yu Wang 0134, Chencheng Wang, Gang Xiong 0001, Zhen Li 0011 |
BlockSys | 4 |
| 2021 | EthSniffer: A Global Passive Perspective on Ethereum
Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
BlockSys | 3 |
| 2021 | A Hashgraph-Based Knowledge Sharing Approach for Mobile Robot Swarm
Xiao Shu, Bo Ding 0001, Xiang Fu 0002, Zhen Li 0011 |
CollaborateCom (2) | 6 |
| 2021 | TA-GAN: GAN based Traffic Augmentation for Imbalanced Network Traffic ClassificationabstractAs the mainstream in network traffic classification (NTC), machine learning (ML) based methods suffer performance degradation due to the imbalance distribution of Internet traffic. Data augmentation methods including the traditional oversampling techniques and the Generative Adversarial Network (GAN) based generation methods are most commonly used to counter the imbalance problem in NTC. However, the former is prone to overfitting and introducing noise. The latter overcomes the above weaknesses, but the quality of the generated traffic samples is difficult to judge. Besides, these methods all divide the imbalanced traffic classification problem into two subproblems, which cannot guarantee the global optimality. In this paper, we propose a GAN based Traffic Augmentation (TA-GAN) for imbalanced traffic classification. TA-GAN is an end-to-end framework that integrates the generation of the minority traffic samples with the training of the target classifier. We design the feedback mechanism to better guide the direction of the sample generation and simultaneously indicate the quality of the synthesized samples. Moreover, the existing deep learning-based NTC methods can be easily adapted to imbalance scenarios with TA-GAN. Comprehensive experiments on the public ISCXVPN2016 dataset demonstrate that TA-GAN effectively mitigates the influence of traffic imbalance (a maximum 14.64% improvement to the minority class'$F_{1}$score) and outperforms the state-of-the-art methods. Gang Xiong 0001, Zhen Li 0011, Junzheng Shi, Mingxin Cui, Gaopeng Gou |
IJCNN | 3 |
| 2021 | UMVD-FSL: Unseen Malware Variants Detection Using Few-Shot LearningabstractAs the tool for launching cyber attacks, the ever-increasing malware variants pose a significant threat to the interconnected network community. The detection methods based on conventional machine learning techniques require lots of samples for training. However, in real-world scenarios, such as in the early stage of novel attacks appearance, only a small number of malicious samples can be obtained. Applying data-intensive traditional methods in the above scenarios will cause serious overfitting problems. Therefore, there is a need for few-shot detection. In his paper, we propose UMVD-FSL, a framework based on few-shot learning to detect unseen malware variants with a small set of data. We start with network traffic data generated by malware variants and benign applications and then convert them to grayscale images. The prototype-based few-shot learning model takes the grayscale images as the input and utilizes meta-training to generalize the meta-learner for adapting new tasks. When a new sample appears, the model performs classification by computing distances to prototype representation of each class. We evaluate different methods through a series of comparative experiments. Our method has the best performance on all subtasks. The experimental results indicate that our method is universal and robust in detecting malware variants from the same network environment and different network environments. The above points prove that our method can accomplish the task of few-shot unseen malware variants detection. Candong Rong, Gaopeng Gou, Chengshang Hou, Zhen Li 0011, Gang Xiong 0001, Li Guo 0001 |
IJCNN | 4 |
| 2021 | 6GAN: IPv6 Multi-Pattern Target Generation via Generative Adversarial Nets with Reinforcement LearningabstractGlobal IPv6 scanning has always been a challenge for researchers because of the limited network speed and computational power. Target generation algorithms are recently proposed to overcome the problem for Internet assessments by predicting a candidate set to scan. However, IPv6 custom address configuration emerges diverse addressing patterns discouraging algorithmic inference. Widespread IPv6 alias could also mislead the algorithm to discover aliased regions rather than valid host targets. In this paper, we introduce 6GAN, a novel architecture built with Generative Adversarial Net (GAN) and reinforcement learning for multi-pattern target generation. 6GAN forces multiple generators to train with a multi-class discriminator and an alias detector to generate non-aliased active targets with different addressing pattern types. The rewards from the discriminator and the alias detector help supervise the address sequence decision-making process. After adversarial training, 6GAN's generators could keep a strong imitating ability for each pattern and 6GAN's discriminator obtains outstanding pattern discrimination ability with a 0.966 accuracy. Experiments indicate that our work outperformed the state-of-the-art target generation algorithms by reaching a higher-quality candidate set. Tianyu Cui, Gaopeng Gou, Gang Xiong 0001, Chang Liu 0049, Peipei Fu, Zhen Li 0011 |
INFOCOM | 6 |
| 2021 | RecGraph: Graph Recovery Attack using Variational Graph AutoencodersabstractGraph-structured data contains a lot of sensitive information about individuals. In order to protect users’ privacy, many anonymization mechanisms for graph-structured data are proposed. However, one common drawback of these mechanisms is that they only consider to hide the local characteristics, such as the degree of nodes or their neighbors. They lack the consideration for the nodes’ attribute features and the features of potential global graph structure, which leads to the failure of these mechanisms to provide sufficient security.To address this shortcoming, we propose RecGraph, a framework for graph recovery attack based on variational graph autoencoders. We use RecGraph to perform graph recovery attack on three real social network datasets, and compare it with five existing baselines, to prove the effectiveness of our method. We also evaluate the privacy wastage after performing the graph recovery attack using RecGraph to demonstrate the serious security risks faced by the existing graph anonymization mechanisms. Chang Liu 0049, Gaopeng Gou, Zhen Li 0011, Gang Xiong 0001, Yangyang Guan |
IPCCC | 4 |
| 2021 | Combating Imbalance in Network Traffic Classification Using GAN Based OversamplingabstractWith the proliferation of encrypted traffic, machine learning (ML) based network traffic classification (NTC) has become the mainstream method. However, most studies ignored two issues. On the one hand, Internet traffic presents a natural uneven distribution. On the other hand, machine learning algorithms generally aim to achieve the highest overall accuracy without considering class imbalance. This leads to severe performance degradation of existing ML-based NTC schemes when facing imbalanced scenarios. In this paper, we design a novel Generative Adversarial Network (GAN) architecture to generate traffic samples, in which the addition of the classifier and the pretraining module makes the generation process more stable and effective. We propose an end-to-end framework for imbalanced traffic classification, named ITCGAN, which can generate traffic samples for minority classes to adaptively rebalance the original traffic and simultaneously train the optimal classifier. We evaluate its effectiveness on the public ISCXVPN2016 dataset based on the global metrics and individual metrics. The results show that our method performs well in imbalanced NTC tasks, fully alleviating the performance degradation (a 10.27-percentage-point improvement to the precision of the most minority class). Meanwhile, it surpasses five state-of-the-art oversampling methods. Gang Xiong 0001, Zhen Li 0011, Junzheng Shi, Mingxin Cui, Gaopeng Gou |
Networking | 3 |
| 2021 | CQNet: A Clustering-Based Quadruplet Network for Decentralized Application Classification via Encrypted Traffic
Yu Wang 0134, Gang Xiong 0001, Chang Liu 0049, Zhen Li 0011, Mingxin Cui, Gaopeng Gou |
ECML/PKDD (4) | 4 |
| 2021 | Towards Multi-source Extension: A Multi-classification Method Based on Sampled NetFlow RecordsabstractWith the rapid development of the Internet, network traffic is growing explosively. It brings great challenges to the traditional traffic identification technology using full traffic analysis, which requires more resources to achieve the collection and analysis of full traffic. And, handling the raw traffic may lead to the compromise of user privacy. NetFlow has good compatibility with the existing routing or switching devices, can aggregate network traffic information, support traffic sampling, reduce the invasion of user privacy, and can effectively deal with the challenges. However, as NetFlow is usually output after traffic sampling to ensure the performance of network devices and only contains session-level statistical information, existing NetFlow research mostly focuses on the binary classification problems (e.g., specific anomaly traffic detection), and less exploration has been conducted on traffic multi-classification problems. And NetFlow is even less involved in the currently popular field of encrypted traffic classification. In this paper, we focus on how to perform encrypted traffic multi-classification research based on sampled NetFlow records and propose a multi-classification method based on the multi-source extension of sampled NetFlow records. To improve the distinguishability and applicability of the sampled NetFlow records, we extend and enrich the records with full consideration of the head or payload information in traffic data, including TTL values, Cipher Suites, etc. For different application scenarios, the methods based on head information extension and payload information extension are proposed, respectively. Through comprehensive experiments, the results show that the proposed method is more applicable and effective than the method based on a single N etFlow record in dealing with multi-classification problems in different encryption application scenarios. Peipei Fu, Qingya Yang, Yangyang Guan, Bingxu Wang, Gaopeng Gou, Zhen Li 0011, Gang Xiong 0001 |
TrustCom | 6 |
| 2021 | LFETT2021: A Large-scale Fine-grained Encrypted Tunnel Traffic DatasetabstractWith the widespread use of tunnel technology, the volume of encrypted tunnel traffic rises sharply, which brings a new challenge to traditional encrypted traffic identification. A number of real-world application scenarios, including Quality of Service and intrusion detection, have put forward new re-quirements for identifying numerous tunnels, applications, and fine-grained behavior. However, previous studies and datasets on encrypted tunnel traffic identification fail to meet these requirements due to their low dataset coverage and coarse label granularity. These weaknesses further affect the extracted features based on these datasets, making them unable to adequately characterize encrypted tunnel traffic. In this paper, we refine the previous tunnel traffic identification granularity from prevalent application identification to behavior identification, and propose LFETT2021, a large-scale fine-grained encrypted tunnel traffic dataset. Our dataset expands the coverage to two operating system platforms, five tunnels, 23 applications, and 76 behaviors. Furthermore, we propose a set of Time-Packet-Related features to better characterize encrypted tunnel traffic. Our comprehensive experiments on LFETT2021 and Time-Packet-Related features show the best average precision of 85% and recall of 88% in 3 different granularity identification scenarios. Gaopeng Gou, Chengshang Hou, Gang Xiong 0001, Zhen Li 0011 |
TrustCom | 5 |
| 2021 | Old Habits Die Hard: A Sober Look at TLS Client Certificates in the Real WorldabstractCertificates play a key role in TLS, which is by far the most widely used security protocol for protecting network traffic. Studies have shown that inappropriate usage of certificates may incur security and privacy risks, most of which are focused on the server-side certificates. However, with the rapid development of the Internet of Things that interconnects countless nodes over the world, as well as the Zero Trust philosophy that stresses authentication of every entity, the adoption of client certificates could be a lot more vital. According to our observation, many practical problems and security risks still exist in the deployment and use of client certificates. In this paper, we present a passive measurement of over 24 million client certificates, collected by a framework deployed on the CSTNET, one of the major academic backbone networks in China. By performing a comprehensive analysis of the large scale real-world data, we give a big picture of the client certificates usage in current network, and disclose implementation flaws of these certificates which may possibly harm transport layer security and user privacy. As many as 342,699 defective client certificates are unearthed, which is an important reminder that never should we neglect the correct use of certificates on the client side. Wei Wang 0314, Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011 |
TrustCom | 7 |
| 2021 | SiamHAN: IPv6 Address Correlation Attacks on TLS Encrypted Traffic via Siamese Heterogeneous Graph Attention Network
Tianyu Cui, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui, Chang Liu 0049 |
USENIX Security Symposium | 4 |
| 2021 | Survey of security supervision on blockchain from the perspective of technology
Yu Wang 0134, Gaopeng Gou, Chang Liu 0049, Mingxin Cui, Zhen Li 0011, Gang Xiong 0001 |
J. Inf. Secur. Appl. | 5 |
| 2021 | Classifying encrypted traffic using adaptive fingerprints with multi-level attributes
Chang Liu 0049, Gang Xiong 0001, Gaopeng Gou, Siu-Ming Yiu, Zhen Li 0011, Zhihong Tian 0001 |
World Wide Web | 5 |
| 2020 | Joint Analysis of Port and Protocol via Endpoint Measurement: An Empirical StudyabstractAs network services continuously evolving, accurately classifying traffic is important for network operators to optimize QoS and customize policy. Network service uses non-standard ports and protocol obfuscation causing damage to the accurate port-based and payload-based traffic classification. However, Deep Packet Inspection (DPI) technique, which combines the payload-based method and port-based method, is still adopted by practitioners from the academic and industrial community. In this paper, we investigate the DPI classification result on a large network to estimate the impact of two factors. We qualify the popularity of non-standard port among different protocols. By endpoint filtering, we discover a large proportion of non-standard ports are opened temporally. We show there still is strong association between P2P protocols and camouflaged protocol. In particular, using both host and label association between endpoints, we find camouflaged protocols exhibit an abnormal port span that is different with the original protocol and are similar to the port span of P2P protocols. Chengshang Hou, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
APNOMS | 4 |
| 2020 | FLAGB: Focal Loss based Adaptive Gradient Boosting for Imbalanced Traffic ClassificationabstractMachine learning (ML) is widely applied to network traffic classification (NTC), which is an essential component for network management and security. While the imbalance distribution exhibiting in real-world network traffic degrades the classifier's performance and leads to prediction bias towards majority classes, which is always ignored by exiting ML-based NTC studies. Some researches have proposed solutions such as resampling for imbalanced traffic classification. However, most methods don't take traffic characteristics into account and consume much time, resulting in unsatisfactory results. In this paper, we analyze the imbalanced traffic data and propose the focal loss based adaptive gradient boosting framework (FLAGB) for imbalanced traffic classification. FLAGB can automatically adapt to NTC tasks with different imbalance levels and overcome imbalance without the prior knowledge of data distribution. Our comprehensive experiments on two network traffic datasets covering binary and multiple classes prove that FLAGB outperforms the state-of-the-art methods. Its low time consumption during training also makes it an excellent choice for highly imbalanced traffic classification. Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
IJCNN | 3 |
| 2020 | JumpEstimate: a Novel Black-box Countermeasure to Website Fingerprint Attack Based on Decision-boundary ConfusionabstractRecent research shows that website fingerprinting (WF) is a growing threat to privacy-sensitive web users, especially when using machine learning techniques such as deep learning or machine learning (DL / ML) to attack website fingerprint, reducing the effectiveness of the previous defense strategies. The reason is that the features targeted by the previous defense countermeasures are manually extracted, the range of it can’t be large enough to cover the range of features automatically extracted by DL / ML-based attacks. This paper proposes a black box defense countermeasure based on decision boundary confusion. Instead of manually extracting features, it uses the classification results of the classifier to determine the decision boundary of the classifier then automatically find the adversarial traffic that may cause the classifier to be confused. At the same time, to solve the retraining problem caused by adversarial traffic, we also utilize Monte Carlo estimation to modify adversarial traffic, to confuse decision boundary, improve the retraining resistance of adversarial traffic. Therefore, it is difficult for the classifier to form a stable and effective decision boundary after training the adversarial traffic. Results shows that our method gets a average defense success rate of 78.2% when facing the baseline WF Attacks, outperforming existing SOTA method Walkie-Talkie’s 63.6% average defense success rate. At the same time, our method improves the ability of the adversarial traffic to resist retrain, increased the retrain defense success rate from 12% to 78.2% under 31% overhead. Wei Cai 0007, Gaopeng Gou, Peipei Fu, Zhen Li 0011, Gang Xiong 0001 |
ISCC | 5 |
| 2020 | Not Afraid of the Unseen: a Siamese Network based Scheme for Unknown Traffic DiscoveryabstractAs an essential task for network management and security, network traffic classification has attracted increasing attention in recent years. Traditional traffic classification methods achieve certain success in identifying specific application traffic but fail with un-predefined unknown classes. Existing unknown traffic discovery methods commonly pick out some unlabeled testing data as part of training data to train the classification models, which is not in line with the real-world open environments. In this paper, we propose a novel scheme named SEEN to achieve unknown traffic detection in network traffic classification. There are three crucial phases in the SEEN: unknown discovery, unknown clustering, and system update. In the first step, using a metric-based approach with siamese network, SEEN identifies unknown traffic as well as accurately classifies the traffic generated by pre-defined application classes. After discovery, unknown traffic is automatically clustered into more fine-grained categories in the unknown clustering step. In the system update step, inspired by low-shot learning, SEEN allows new classes to be added or unnecessary known classes to be deleted quickly without retraining from the sketch, which can complement the system’s knowledge. Experimental results exhibit that SEEN can achieve outstanding performances both on known and unknown traffic identification on two open real-world datasets, and the proposed scheme can address the problem of unknown traffic effectively. Zhen Li 0011, Junzheng Shi, Gaopeng Gou, Chang Liu 0049, Gang Xiong 0001 |
ISCC | 2 |
| 2020 | ResTor: A Pre-Processing Model for Removing the Noise Pattern in Flow CorrelationabstractFlow correlation is a common approach to break the anonymity of anonymous communication. However, unpredictable network noise caused by multiple factors in open Internet raises the bar for existing correlation methods. Traditional methods comparing statistical distance of data flows and deep learning methods such as convolutional neural network behave worse because network noise changes traffic shape. In this paper, we design a pre-processing model called ResTor to perform the noise reduction before actually correlating entering and exiting flows. ResTor treats the byte accumulation sequences smoothed at fixed intervals as fitting targets, and takes advantage of the stacked auto-encoder architecture to remove noise in two phases. Experiment results show that the exiting Tor flows processed by ResTor are closer to their corresponding entering flows, thus the correlation task can be finished effectively even using traditional correlation ways: cosine distance and other statistical metrics assisted by ResTor achieve less computational overhead and higher correlation accuracy on Tor compared to the state-of-the-art method of DeepCorr, especially when traffic is obfuscated. Gang Xiong 0001, Zhen Li 0011, Gaopeng Gou |
ISCC | 3 |
| 2020 | MalFinder: An Ensemble Learning-based Framework For Malicious Traffic DetectionabstractMalicious events pose a significant threat to the current increasingly interconnected Internet community. Detection based on features of network traffic and machine learning algorithms is a common approach to identify malicious events. The performance of approaches is associated with the used features and algorithms. In this paper, we propose MalFinder, an ensemble learning-based framework for malicious traffic detection. Considering the trend of network traffic encryption and the complexity of decrypting traffic, we utilize statistical features and sequence features to describe network traffic. We extend the dimensions of these two types of features to enhance their capability for representing traffic data. Feature importance analysis and contrast experiments illustrate the effectiveness of our new features. Among our selected classifiers suitable for malicious traffic detection, boosting-based classifiers XGBoost and LightGBM can reduce bias, and bagging-based classifier Random Forest can reduce variance. Stacking, which is the integration method of the classification results used in our framework, can improve the generalization ability of the method. MalFinder can achieve 96.58% F-measure and 95.44% accuracy in the malicious traffic detection task on a real-world dataset, whose results are better than those of comparison methods. In terms of unseen malicious traffic discovery, MalFinder still provides good performance with 93.46% F-measure and 91.04% accuracy, which even surpasses the results in the task of known malicious traffic detection of other comparative methods. With consideration of the scarcity of public data sets used for malicious traffic detection, we have exposed our self-built dataset for more extensive researches. Candong Rong, Gaopeng Gou, Mingxin Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
ISCC | 5 |
| 2020 | TransNet: Unseen Malware Variants Detection Using Deep Transfer Learning
Candong Rong, Gaopeng Gou, Mingxin Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
SecureComm (2) | 5 |
| 2020 | Identifying DApps and User Behaviors on Ethereum via Encrypted Traffic
Yu Wang 0134, Gaopeng Gou, Gang Xiong 0001, Chencheng Wang, Zhen Li 0011 |
SecureComm (2) | 6 |
| 2020 | NSA-Net: A NetFlow Sequence Attention Network for Virtual Private Network Traffic Detection
Peipei Fu, Chang Liu 0049, Qingya Yang, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
WISE (1) | 7 |
| 2020 | A Survey of Key Technologies for Constructing Network Covert ChannelabstractIn order to protect user privacy or guarantee free access to the Internet, the network covert channel has become a hot research topic. It refers to an information channel in which the messages are covertly transmitted under the network environment. In recent years, many new construction schemes of network covert channels are proposed. But at the same time, network covert channel has also received the attention of censors, leading to many attacks. The network covert channel refers to an information channel in which the messages are covertly transmitted under the network environment. Many users exploit the network covert channel to protect privacy or guarantee free access to the Internet. Previous construction schemes of the network covert channel are based on information steganography, which can be divided into CTCs and CSCs. In recent years, there are some covert channels constructed by changing the transmission network architecture. On the other side, some research work promises that the characteristics of emerging network may better fit the construction of the network covert channel. In addition, the covert channel can also be constructed by changing the transmission network architecture. The proxy and anonymity communication technology implement this construction scheme. In this paper, we divide the key technologies for constructing network covert channels into two aspects: communication content level (based on information steganography) and transmission network level (based on proxy and anonymity communication technology). We give an comprehensively summary about covert channels at each level. We also introduce work for the three new types of network covert channels (covert channels based on streaming media, covert channels based on blockchain, and covert channels based on IPv6). In addition, we present the attacks against the network covert channel, including elimination, limitation, and detection. Finally, the challenge and future research trend in this field are discussed. Gang Xiong 0001, Zhen Li 0011, Gaopeng Gou |
Secur. Commun. Networks | 3 |
| 2019 | Deep Forest with LRRS Feature for Fine-grained Website Fingerprinting with Encrypted SSL/TLSabstractWith the development of encryption protocol, such as Secure Sockets Layer (SSL) and Transport Layer Security (TLS), the traditional fingerprinting approaches based on packet content and special field are difficult to fingerprint the websites. Therefore, recent research imported machine learning algorithms to deal with this problem, and various features are extracted for the machine learning algorithms. However, previous approaches of fingerprinting encrypted websites are based on HTTP/1.1, which are not applicable to the widely used HTTP/2. In addition, most of the work only fingerprints the home page of each website, but in fact, users also visit other web pages of the website. To solve the feature compatibility problem, we propose to use the local request and response sequence (LRRS) as features. LRRS can represent the patterns of the encrypted Internet traffic not only based on HTTP/1.1 but also based on HTTP/2 using local packet sequences. In order to fingerprint different web pages in the same website, we import Deep Forest to extract fine-grained features. It utilizes a convolution structure to make full use of LRRS sequential features and multi-layer structure to enhance the ability of feature representation. The experimental results show the proposed algorithm has achieved the best overall performance on four datasets. Especially on the bidirectional encrypted traffic dataset with HTTP/2, the proposed approach achieved 55% higher of f1 score than the state-of-the-art method KFP with Random Forest. Cuicui Kang, Gang Xiong 0001, Zhen Li 0011 |
CIKM | 4 |
| 2019 | DLchain: A Covert Channel over Blockchain Based on Dynamic Labels
Gaopeng Gou, Chang Liu 0049, Gang Xiong 0001, Zhen Li 0011 |
ICICS | 6 |
| 2019 | Vision Information and Laser Module Based UAV Target TrackingabstractThis paper investigates the target tracking mission of an Unmanned Aerial Vehicle (UAV) equipped with a camera and a laser module. Firstly, utilizing Deep Neural Network (DNN) and Kernelized Correlation Filters (KCF), target recognition and location in the pixel coordinate system is achieved based on vision. Furthermore, by combining the laser ranging information and the distance estimation algorithm based on image, the distance between the UAV and the target is well estimated. To ensure the target tracking, a PID controller based on the distance error is applied to the UAV. The effectiveness of the system is verified on an actual UAV target tracking scenario. Chang Liu 0049, Yansui Song, Yuyan Guo, Bin Xu 0003, Yu Zhang 0018, Zhen Li 0011 |
IECON | 7 |
| 2019 | FS-Net: A Flow Sequence Network For Encrypted Traffic ClassificationabstractWith more attention paid to user privacy and communication security, the volume of encrypted traffic rises sharply, which brings a huge challenge to traditional rule-based traffic classification methods. Combining machine learning algorithms and manual-design features has become the mainstream methods to solve this problem. However, these features depend on professional experience heavily, which needs lots of human effort. And these methods divide the encrypted traffic classification problem into piece-wise sub-problems, which could not guarantee the optimal solution. In this paper, we apply the recurrent neural network to the encrypted traffic classification problem and propose the Flow Sequence Network (FS-Net). The FS-Net is an end-to-end classification model that learns representative features from the raw flows, and then classifies them in a unified framework. Moreover, we adopt a multi-layer encoder-decoder structure which can mine the potential sequential characteristics of flows deeply, and import the reconstruction mechanism which can enhance the effectiveness of features. Our comprehensive experiments on the real-world dataset covering 18 applications indicate that FS-Net achieves an excellent performance (99.14% TPR, 0.05% FPR and 0.9906 FTF) and outperforms the state-of-the-art methods. Chang Liu 0049, Longtao He, Gang Xiong 0001, Zigang Cao, Zhen Li 0011 |
INFOCOM | 5 |
| 2019 | Malicious Domain Detection via Domain Relationship and Graph ModelsabstractMalicious domain is a vital component of various cyber attacks. Recent techniques detect malicious domains by building classifiers based on domain character features which may be easily evaded by attackers. In this paper, we propose a malicious domain detection approach based on domain relationship features, PDNS features, and domain character features. The key insight is that malicious domains deploy on IP that is loosely regulated and the domains on such IP have similar network characteristics including domain relationships, resolution characteristics, and network behaviors. We find that the relationship of malicious domains is different from that of benign domains. Take this into account, we build meaningful associations among domains and extract the domains relationship features by a modified graph embedding algorithm from Passive DNS data. Besides, we mine more features from PDNS which have not been mentioned in previous work. These PDNS features can enhance the effectiveness of the classifier. Finally, we combine domain character features, PDNS features and relationship features as the feature set. We evaluate the performance of our model on a real-world dataset from DNS servers. We achieve excellent performance by applying several classifiers based on domain character features, PDNS features and relationship features with an accuracy of 94.0%, a recall of 94.3% and a precision of 93.8% in the challenging scenario where domains deploy on the same IP and malicious domains share similar character features with benign domains. We also compare our method with two state-of-the-art detection approaches and find that our approach outperforms those SOTA approaches. Based on the comparison results, we point out that our way to construct a domain relationship graph can effectively mine the domain association features and the features combined with PDNS features and domain character features can effectively identify malicious domains which are similar to benign domains. Gaopeng Gou, Cuicui Kang, Chang Liu 0049, Zhen Li 0011, Gang Xiong 0001 |
IPCCC | 5 |
| 2018 | SSL/TLS Security Exploration Through X.509 Certificate's Life Cycle MeasurementabstractWith the popular use of SSL/TLS, more and more web applications, such as online banking, e-mail, and ecommerce, turn to secured channels for communication, which rely on X.509 certificate for authentication. Generally, every certificate has a theoretical validity period when it is issued. However, the used period in practice is often different from the theoretical validity, namely, before or after the validity, for a long or short time. If a certificate is expired, it is easily to be exploited by cyber-attackers, leading to web users' personal information at risk. To explore the security flaws of the SSL/TLS certificate, we conduct a large-scale measurement study of X.509 certificate life cycle from the view of leaf certificates. Based on a passive data set collected over one year, we investigate the certificate validity period in a fine-grained manner, and uncover that the actual usage of the certificates are not satisfactory. Meanwhile, we discover several security-related issues that may leave the web communication at risk. The recommendations are summarized to ensure the long-term security for certificate use in practice. We believe that the work will be beneficial to web security and improve the certificate utilization in the future. Peipei Fu, Zhen Li 0011, Gang Xiong 0001, Zigang Cao, Cuicui Kang |
ISCC | 2 |
| 2018 | LaFFT: Length-Aware FFT Based Fingerprinting for Encrypted Network Traffic ClassificationabstractEncrypted traffic classficiation has become an emergent and challenging task for network monitoring and management. Traditional classification methods for encrypted traffic rely on complex statistical characteristic construction and in-depth packet resolution, which produce huge loads. In this paper, we develop Length-aware FFT (LaFFT) fingerprinting to identify different encrypted application traffic with packet length sequences. We apply FFT to packet length sequences to generate the frequency domain vectors as LaFFT features. We verify the distinguishability of LaFFT fingerprinting by data analysis. Furthermore, the linear inseparability and the front superiority of LaFFT fingerprinting are demonstrated by comprehensive experiments. In the real-world dataset, the LaFFT fingerprinting with random forest classifier can achieve 96.8% TPR, 0.32% FPR and 0.959 FFT, which significantly outperform the state-of-the-art methods. Chang Liu 0049, Zigang Cao, Zhen Li 0011, Gang Xiong 0001 |
ISCC | 3 |
| 2017 | Learning Deep Semantic Embeddings for Cross-Modal RetrievalabstractDeep learning methods have been actively researched for cross-modal retrieval, with the softmax cross-entropy loss commonly applied for supervised learning. However, the softmax cross-entropy loss is known to result in large intra-class variances, which is not not very suited for cross-modal matching. In this paper, a deep architecture called Deep Semantic Embedding (DSE) is proposed, which is trained in an end-to-end manner for image-text cross-modal retrieval. With images and texts mapped to a feature embedding space, class labels are used to guide the embedding learning, so that the embedding space has a semantic meaning common for both images and texts. This way, the difference between different modalities is eliminated. Under this framework, the center loss is introduced beyond the commonly used softmax cross-entropy loss to achieve both inter-class separation and intra-class compactness. Besides, a distance based softmax cross-entropy loss is proposed to jointly consider the softmax cross-entropy and center losses in fully gradient based learning. Experiments have been done on three popular image-text cross-modal retrieval databases, showing that the proposed algorithms have achieved the best overall performances. Cuicui Kang, Shengcai Liao, Zhen Li 0011, Zigang Cao, Gang Xiong 0001 |
ACML | 3 |
| 2017 | POSTER: An Empirical Measurement Study on Multi-tenant Deployment Issues of CDNsabstractContent delivery network (CDN) has been playing an important role in accelerating users' visit speed, bring good experience for popular web sites around the world. It has become a common security enhance service for CDN providers to offer HTTPS support to tenants. When several tenants are deployed to share a same IP address due to resource efficiency and cost, CDN providers should make comprehensive settings to ensure that all tenants' sites work correctly on users' requests. Otherwise, issues can take place such as denial of service (DOS) and privacy leakage, causing very bad user experience to users as well as potential economic loss for tenants, especially under the situation of hybrid deployment of HTTP and HTTPS. We examine the deployments of typical multi-tenant CDN providers by active measurement and find that CDN providers, namely Akaimai and ChinaCenter, have configuration problems which can result in DOS by certificate name mismatch error. Several advices are given to help to mitigate the issue. We believe that our study is meaningful for improving the security and the robustness of CDN. Zixi Cai, Zigang Cao, Gang Xiong 0001, Zhen Li 0011 |
CCS | 4 |
| 2017 | Identifying malware with HTTP content type inconsistency via header-payload comparisonabstractMalware is one of the most severe security threats on the Internet. A key challenge for attackers is to install their malware programs on as many victim machines as possible. HTTP protocol, being the most popular protocol and occupying a significant portion of network traffic, is an obvious target for attackers to exploit for malware distribution. Advanced attackers would even hide the malicious executable program behind a benign file such as text, image. The existence of malware becomes harder to detect and the distribution channels become more evasive (i.e., not clear to identify). However, the exploited and hidden behavior often leads to an inconsistency between the actual content type and the declared content type. In this paper, we conduct a detailed study on a seven-month traffic of content type inconsistency executable program downloaded from an ISP of CSTNET (China Science and Technology Network). We found that 99.78% (891/893) of PE (portable executable) files declared to be images are malicious and 100% of PE files declared to be text with typical file extensions, “.pdf”, “.doc”, “.css” are malware. So, content type inconsistency can be used to detect evasive network attacks as well as effectively discover unknown malware from the traffic. Haiqing Pan, Zigang Cao, Zhen Li 0011, Gang Xiong 0001, Yangyang Guan, Siu-Ming Yiu |
IPCCC | 4 |
| 2017 | Metrie learning with statistical features for network traffic classificationabstractWith the development of Internet techniques, such as the Secure Sockets Layer and Transport Layer Security encryption protocol, the traditional internet traffic classification approaches based on port, IP and packet content is difficult to identify the traffic flows. Therefore, many researches imported Machine Learning algorithm to deal with the problem, and the statistical features are extracted for the machine learning algorithms. However, the features are often constructed of various features in different spaces, such as the port ID, packets number, one-hot encodings and statistical properties. The traditional machine learning algorithms usually use Euclidean metric for the distance computing, which is unable to make the best use of the artificial features with various Internet traffic flow attributes. Considering this, the paper proposed to utilize Metric Learning algorithms to learn the adaptive distance metric for the multiple features. As a result, the proposed algorithm can take better advantage of the artificial features and make full use of the characteristics. Finally, the evaluation is conducted on the encrypted web sites traffic database with the comparison of several state-of-the-art algorithms, and experimental results show that the proposed algorithm has achieved the best performance with 8% higher of accuracy than Decision Tree which is the second best algorithm. Cuicui Kang, Peipei Fu, Zigang Cao, Zhen Li 0011, Gang Xiong 0001 |
IPCCC | 5 |
| 2017 | A network attack forensic platform against HTTP evasive behavior
Zhen Li 0011, Haiqing Pan, Zigang Cao, Gang Xiong 0001 |
J. Supercomput. | 1 |
| 2014 | POSTER: Mining Elephant Applications in Unknown Traffic by Service ClusteringabstractNetwork traffic classification is of great importance for fine-grained network management and network security. However, with the rapid development of new network applications in recent years, traffic that cannot be identified by classifiers accounts for an increasing ratio, which brings a great challenge for network operators. Most of the unknown traffic is usually generated by only a few or some certain kinds of applications. We call this kind of traffic as the elephant traffic. It is generally recognized that traffic sharing the same server IP and server port is generated by the same application. In this paper, we say that they are belonging to the same service. Therefore, we propose a novel method, in which service-based statistical features are used for cluster analysis, to classify these elephant traffic. Preliminary results on a real network traffic dataset show that our method is able to automatically identify similar unknown applications. We believe that classifying unknown traffic in service perspective is a promising direction. Gang Xiong 0001, Li Guo 0001, Zhen Li 0011, Yong Wang 0032 |
CCS | 5 |