Tianning Zang

dblp:07/6573 · DBLP profile ↗
← Back
49ranked-venue papers
2as first author
33since 2021 · last 2026
0000-0003-3583-6249ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 7 since 2021Security and privacy · 13 · 8 since 2021Human-computer interaction and ubiquitous computing · 12 · 10 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CARVE: Vulnerability Entity Alignment via Semantic Focusing and Staged Selection
Haohui Peng, Senhao Wen, Tianning Zang, Yueyue Hu
ICIC (11)4
2025 Cultivating Gaming Sense for Yourself: Making VLMs Gaming Experts
abstract
Developing agents capable of fluid gameplay in first/third-person games without API access remains a critical challenge in Artificial General Intelligence (AGI). Recent efforts leverage Vision Language Models (VLMs) as direct controllers, frequently pausing the game to analyze screens and plan action through language reasoning. However, this inefficient paradigm fundamentally restricts agents to basic and non-fluent interactions: relying on isolated VLM reasoning for each action makes it impossible to handle tasks requiring high reactivity (e.g., FPS shooting) or dynamic adaptability (e.g., ACT combat). To handle this, we propose a paradigm shift in gameplay agent design: instead of direct control, VLM serves as a developer, creating specialized execution modules tailored for tasks like shooting and combat. These modules handle real-time game interactions, elevating VLM to a high-level developer. Building upon this paradigm, we introduce GameSense, a gameplay agent framework where VLM develops task-specific game sense modules by observing task execution and leveraging vision tools and neural network training pipelines. These modules encapsulate action-feedback logic, ranging from direct action rules to neural network-based decisions. Experiments demonstrate that our framework is the first to achieve fluent gameplay in diverse genres, including ACT, FPS, and Flappy Bird, setting a new benchmark for game-playing agents.
Wenxuan Lu, Jiangyang He, Zhanqiu Zhang, Steven Y. Guo, Tianning Zang
ACL (1)5
2025 Global Eye: Breaking the "Fixed Thinking Pattern" during the Instruction Expansion Process
abstract
An extensive high-quality instruction dataset is crucial for the instruction tuning process of Large Language Models (LLMs).Recent instruction expansion methods have demonstrated their capability to improve the quality and quantity of existing datasets, by prompting high-performance LLM to generate multiple new instructions from the original ones.However, existing methods focus on constructing multi-perspective prompts (e.g., increasing complexity or difficulty) to expand instructions, overlooking the "Fixed Thinking Pattern" issue of LLMs.This issue arises when repeatedly using the same set of prompts, causing LLMs to rely on a limited set of certain expressions to expand all instructions, potentially compromising the diversity of the final expanded dataset.This paper theoretically analyzes the causes of the "Fixed Thinking Pattern", and corroborates this phenomenon through multi-faceted empirical research.Furthermore, we propose a novel method based on dynamic prompt updating: Global Eye.Specifically, after a fixed number of instruction expansions, we analyze the statistical characteristics of newly generated instructions and then update the prompts.Experimental results show that our method enables LLaMA3-8B and LLaMA2-13B to surpass the performance of open-source LLMs and GPT3.5 across various metrics.
Wenxuan Lu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Songhao Jiang, Tianning Zang
ACL (1)6
2025 Pegasus: Accelerating Provenance Graph-based Intrusion Detection Methods
abstract
Provenance-based Endpoint Detection and Response (P-EDR) systems are considered as the key to future Advanced Persistent Threat (APT) defense. Building provenance graphs that consider causal relationships between software behaviors can better provide contextual information of cyber attacks, which is capable of effectively reconstructing complex cyber attack scenarios represented by APT. Although promising to assist in attack investigation, existing methods for attack detection using provenance graphs adopt a centralized detection architecture, sending all system audit logs to servers for processing, resulting in unbearable costs in terms of data transmission, data storage, and computation. To address the above fundamental challenges, we propose Pegasus, a distributed detection system that can reduce memory consumption during training through a distributed system. Our system is evaluated on a large public dataset, and experimental results show that our system reduces memory consumption by 47%–65% compared with existing provenance-based EDR. And the above processing has little impact on attack detection performance, and our EDR system can still achieve sufficiently good detection results.
Pengcheng Bi, Tianning Zang, Xiao-chun Yun
CSCWD4
2025 Towards Open-World DoH Tunnel Detection: A Dual-View Contrastive Learning Framework with Adaptive Feature Boundaries
abstract
The emergence of DNS-over-HTTPS (DoH) tunnels poses significant challenges to network security, particularly when encountering unknown traffic patterns not seen during training. Existing approaches struggle to effectively identify novel DoH tunnel variants while maintaining accurate classi-fication of known traffic patterns. In this paper, we propose DualConBound, a novel framework that combines dual-view contrastive learning with dynamic feature boundary estimation to address open-set network traffic classification. Our approach leverages complementary traffic representations and adaptive decision boundaries to better distinguish between known and unknown traffic patterns. Through extensive experiments on real-world network traffic datasets, we demonstrate that our framework achieves 91 % accuracy on known traffic patterns and 71 % accuracy on unknown variants, outperforming other methods by 6% in open-set scenarios. Our work provides a robust solution for identifying emerging DoH tunnel threats in real-world network environments and establishes a new paradigm for open-set network traffic classification.
Beibei Feng, Zhefeng Nan, Jiang Xie 0004, Tianning Zang, Jingrun Ma
CSCWD5
2025 Enhanced Encrypted Traffic Classification Through Packet-Flow Dual Views
abstract
In the field of cybersecurity, encrypted traffic classification is of paramount importance, serving as a crucial technology for ensuring data privacy and defending against cyber attacks. Existing ETC (Encrypted Traffic Classification) methods commonly face three challenges in real-world Mobile Internet environments, particularly pertinent to the multimedia context: 1) singular entity traffic information, 2) fixed-size feature filters in learning models, and 3) reliance on traditional encryption protocols. To address these, we introduce the innovative Packet-Flow Dual Views (PFDV) framework. PFDV aims to push the performance boundaries of ETC tasks in multimedia traffic scenarios by examining encrypted payloads separately from the packet and flow views. More specifically, PFDV employs a dual attention framework, comprising both payload attention and gram attention, to accurately identify nuanced differences in network traffic. Such an approach underscores granular packet-level details while offering a comprehensive flow view with packet positional data, striving for an optimal balance in multimedia traffic analysis. PFDV demonstrated superior performance across diverse tasks, achieving accuracy rates and F1 scores significantly higher than competing methods, with results of up to 98.34% accuracy and 95.69 % F1 score in multimedia traffic classification. Furthermore, ablation studies demonstrate each component in PFDV contributes significantly to the overall enhancement of ETC performance in the model.
Beibei Feng, Tianning Zang, Jingrun Ma
WCNC6
2024 MTS-IoT: A Robust Encrypted IoT Traffic Classification via Multi-dimensional Time Series
abstract
The rapid proliferation of IoT devices presents a two-fold challenge: a pronounced deficiency in security measures and a diverse array of Quality of Service (QoS) requirements. For network providers to address these challenges effectively, it is paramount first accurately to identify IoT devices. Current methodologies fall short in their robustness, owing to encrypted traffic and the intricate network environments. This paper applies multi-feature time series to the problem of encrypted IoT device traffic classification, proposing a multi-dimensional time series-based IoT traffic classification method, MTS-IoT. MTS-IoT constructs multi-dimensional time series samples from raw traffic using sliding windows of fixed packet numbers, preserving abundant information. It then utilizes a "Global-Local-Spacial" framework to deeply extract sequence features and introduces sparse self-attention to reduce the training overhead caused by multi-dimensional temporal features. Comprehensive experiments were conducted on a renowned dataset, in which we retained the traffic of non-IoT devices to simulate the real world. Experiments indicate that MTS-IoT outperforms existing methods in classification performance, pushing the F1 to 98.63%(4.46%↑). Moreover, it can achieve accurate detection throughout the entire traffic cycle, resist network congestion, and resist traffic shaping, underscoring its robustness in diverse scenarios.
Tianye Gao, Kehong Liu, ShengBao Li, Ruihai Ge, Tianning Zang
CSCWD6
2024 IMTCDF: A Multi-Module-Based Internet Malicious Traffic Classification and Detection Framework
abstract
Recently, research has shown that neural networks can be utilized for identifying malicious traffic. However, there are shortcomings in existing methods, such as detection rate bottleneck and fewer applicable scenarios. Furthermore, the time-consuming data preprocessing methods negatively impact efficiency of the models and the ability to learn feature information. Besides, the complex fingerprint extraction in traffic is required to solve urgently. Therefore, this paper proposes IMTCDF, a multi-module-based Internet Malicious Traffic Classification and Detection Framework, designed for fast and accurate classification and detection of malicious traffic. The preprocessing module adopts a segmentation method based on global threshold to simplify data processing. In the malicious traffic detection module, this paper designs a Depthwise Separable Convolution with Global Composite Attention Model (DSC-GCA model), benefiting from better capture and learning capability of feature information. We used the publicly available USTF-TFC2016 dataset and the TCD-2022 dataset obtained from autonomous collection in a real Internet environment for our experiments. Multiple groups of experiments show that IMTCDF has outstanding detection capabilities, and the performance remains stable in different scenarios. We selected some models and methods proposed in recent years, which are also used for malicious traffic detection tasks, as a control group for comparative experiments, and the results show that IMTCDF has lower time cost and significant progress in evaluation metrics such as accuracy, precision, recall, and F1-Score.
Zhenyu Cheng 0001, Tianning Zang
CSCWD4
2024 A Machine Learning-based Method for Clustering the Traffic of Linux NATed Network Entities with TCP/IP Feature
abstract
It is crucial to distinguish the network entities (NEs) behind NAT devices for tasks such as information gathering, asset management, and legal interception. The current mainstream approach has the problem of erroneously clustering traffic from a sparsely communicating NE as traffic from multiple NE. This leads to current algorithms not being able to be applied effectively in real environment. To address this issue, we propose a Linux NEs traffic clustering method based on TCP/IP features in this paper. We integrate multiple hidden level features such as operating system and timestamp and freely cluster the data set based on the density of spatial distribution. We use the datasets from the real campus network and public dataset to evaluate the performance of our method, and the results show that our method overcomes the problem erroneously clustering traffic from a sparsely communicating NE as traffic from multiple NEs. For the current mainstream Linux types, the purity of each version exceeds 0.91, RI exceeds 0.98, ARI exceeds 0.88, and FMI exceeds 0.89. Our method can accurately calculate the number of hosts behind NAT and cluster traffic accordingly.
Kehong Liu, Tianye Gao, Tianxing Ma, Tianning Zang
CSCWD5
2024 Revisiting Open DNS Resolver Vulnerabilities to Reflection-Based DDoS Threats
abstract
DNS, as a vital component of the Internet, is frequently exploited for malicious activities. Millions of open DNS resolvers are exposed with public access, posing significant risks. Amplification vulnerability in UDP-based DNS protocol has been abused by miscreants to launch reflection amplification Distributed Denial of Service (DDoS) attacks. In reflection amplification attacks, forged DNS request packets are continuously sent to open resolvers, triggering amplified attack traffic against the targeted victim. To defend against such attacks, resolvers can take measures to reject anomalous requests and limit the size of responses. Measures such as source address verification and response rate limiting prove effective in mitigating the risk of resolver exploitation. However, implementing these measures requires software and hardware updates or configuration changes, potentially incurring additional costs. Currently, it remains unclear how many open resolvers are adequately protected and how many still pose the potential for exploitation. In this paper, we conducted a thorough measurement on open resolvers about the actual potential of abuse. Our measurement results indicated that 14.9% of open resolvers are susceptible to exploitation for reflection-based DDoS attacks and thousands of resolvers are still exposed to reflection amplification attacks with no mitigation measure.
Kehong Liu, Junnan Yin, Letian Du, Tianning Zang
CSCWD5
2024 Efficiently Adapting Traffic Pre-trained Models for Encrypted Traffic Classification
abstract
Classification of encrypted traffic is essential to the security of collaborative systems. Traffic pre-trained models (TPTMs) have shown promising results on this task. Existing methods of TPTMs for encrypted traffic classification typically fine-tune the parameters of TPTM to adapt to different downstream tasks. However, this approach requires updating all parameters of the large-scale TPTM during the training stage and storing an entire new model for every task, which consumes significant computational and storage resources. This paper proposes the Efficient Adapter of Traffic (EAT), which improves the parameter efficiency of TPTMs. Specifically, it fixes the original parameters of the TPTM and inserts a learnable module between each layer of the TPTM. In the process of adapting to all downstream tasks, only the learnable module needs to be trained and stored, which makes our method more efficient in terms of computation and storage. To validate the effectiveness of our approach, we conducted experiments on various encrypted traffic datasets. The results demonstrate that our method achieves similar performance to the fine-tuning method by only adjusting 3.4% of the TPTM parameter quantity, which significantly reduces computational and storage resources.
Wenxuan Lu, Zhuohang Lv, Lanqi Yang, Tianning Zang
CSCWD5
2024 Sky-Eye: Detect Multi-stage Cyber Attacks at the Bigger Picture
Pengcheng Bi, Zhuohang Lv, Xiao-chun Yun, Tianning Zang
ICDF2C (1)5
2024 AppFineGraph: Hierarchical Mobile Encrypted Traffic Classification at Multi-Granularities via a Branch Graph Neural Network
abstract
Most of existing methods for classifying mobile encrypted traffic are primarily designed for coarse-grained scenarios, focusing on classification at the app name level. This implies that these methods can effectively categorize traffic into app names, but possibly performing poorly when it comes to finer-grained classification of in-app activities such as posting comments or location navigation. Classifying in-app activities poses a greater challenge compared to app name-level classification, as different activities within the same app often adopt similar API libraries, resulting in closely resembling traffic patterns. A more formidable challenge is achieving generic classification that supports both the granularity of app names and in-app activities. It is complex for a single classifier to effectively integrate feature information at multiple levels. Addressing these issues, we propose a novel encrypted traffic classification approach – AppFineGraph. AppFineGraph employs a hierarchical classification framework based on Branch Neural Network (B-NN), enabling the capture of both global and local information at different levels. This approach facilitates multi-granular traffic classification for both app names and in-app activities. Additionally, AppFineGraph introduces a graph neural network for feature representation, effectively mining flow features and correlations from encrypted traffic. In comparison to state-of-the-art approaches, experiments demonstrate that AppFineGraph achieves superior performance across all granularities.
ShengBao Li, Zhuohang Lv, Tianning Zang, Lanqi Yang, Qian Qiang
IJCNN3
2024 Vicinal Data Augmentation for Classification Model via Feature Weaken
Songhao Jiang, Yan Chu 0001, Tianxing Ma, Xiaochen Miao, Zhengkui Wang, Tianning Zang
KSEM (1)6
2024 Multilingual Temporal Answer Grounding in Video Corpus with Enhanced Visual-Textual Integration
Tianxing Ma, Yueyue Hu, Shuang Jiang, Zhenhao Yin, Tianning Zang
NLPCC (5)5
2024 From Scarcity to Clarity: Few-Shot Learning for DoH Tunnel Detection Through Prototypical Network
abstract
The widespread adoption of DNS over HTTPS (DoH) has introduced significant challenges in network security, particularly the emergence of DoH tunnels. Existing methods, reliant on large labeled datasets, struggle to adapt to novel DoH tunnel variants. We present ProtoDoH, a novel framework for DoH tunnel detection that uniquely integrates prototypical networks with meta-learning. Our method leverages few-shot learning principles to detect DoH tunnels with minimal training samples, addressing the challenge of sample scarcity in new attack scenarios. By employing a metric-based meta-learning framework, ProtoDoH enables rapid adaptation to novel DoH tunnel variants, significantly reducing the detection time for new DoH tunnels. Experimental results demonstrate that our approach achieves over 99% accuracy in detecting DoH tunnels with as few as 5 samples, notably outperforming traditional machine learning and deep learning methods. Furthermore, our model exhibits strong generalization capabilities across different network environments and DoH tunnel tools.
Beibei Feng, Tianning Zang, Jingrun Ma
TrustCom5
2024 Path Generation Method of Anti-Tracking Network based on Dynamic Asymmetric Hierarchical Architecture
abstract
The continuous advancement of digital processes has significantly increased the risk to users’ data privacy. To protect private data, it is crucial to provide anonymous, anti-tracking secure communication services. However, Existing anti-tracking data transmission technologies has problems such as centralized node, unstable links, low transmission efficiency, and vulnerable static transmission, which fail to meet users’ privacy and security requirements. This paper proposes a dynamic asymmetric hierarchical architecture path generation method(DAHP) for anti-tracking network. The method adopts a hierarchical architecture design, comprising a Secure Access Layer responsible for link construction and a Secret Transmission Layer handling the actual transmission. A node hybrid selection strategy, combined with a dynamic asymmetric path generation strategy, minimizes the risk of node exposure and enables dynamic variation in path generation. The efficacy of the DAHP method is validated through extensive testing on representative network and existing path generation methods. Experimental results demonstrate that the proposed method exhibits superior scalability and anti-tracking performance.
Zhefeng Nan, Changbo Tian, Tianning Zang, Dongwei Zhu
TrustCom5
2024 FineNet: Few-Shot Mobile Encrypted Traffic Classification via a Deep Triplet Learning Network Based on Transformer
abstract
The existing encrypted mobile traffic classification methods often require a large number of training sets to ensure the effect of the model. When it is difficult to collect enough samples, these methods may not yield satisfactory results. Therefore, it is crucial to study an effective few-shot learning method to achieve accurate traffic classification with very few samples. In this paper, we propose FineNet, a few-shot fine-grained classification technology based on a triplet deep learning network with Transformer. The triplet deep learning network is able to discover subtle differences between traffic flows and effectively transfer existing classification knowledge to few-shot scenarios. Moreover, we introduce Transformer as the base model for the triplet network, fully leveraging Transformer's powerful representation capabilities for sequential data to enhance the express ability of FineNet. We conduct multiple comparative experiments, and the result proves that the accuracy rate is at least 2.4% higher than state-of-the-art approaches in few-shot environment.
ShengBao Li, Qian Qiang, Tianning Zang, Lanqi Yang, Tianye Gao
WCNC3
2024 Let model keep evolving: Incremental learning for encrypted traffic classification
Xiang Li 0135, Jiang Xie 0004, Qige Song, Yafei Sang, Yongzheng Zhang 0002, Tianning Zang
Comput. Secur.7
2023 A Semi-supervised Learning Method for Malware Traffic Classification with Raw Bitmaps
Jingrun Ma, Tianning Zang, Beibei Feng
CollaborateCom (2)3
2023 TCCN: A Network Traffic Classification and Detection Model Based on Capsule Network
abstract
Privacy information theft traffic is usually detected using traffic classification methods, and deep learning-based detection methods are effective for this task. However, these methods have complex preprocessing processes as well as tend to ignore the deep features of network traffic, while the generalization ability is not outstanding. In this paper, a network traffic classification and detection model TCCN (Traffic Classification Capsule Network) based on capsule network is proposed. Meanwhile, PacketCGAN-based data balancing method is introduced to assist TCCN in traffic classification. A new feature graph vectorization method is used to improve the efficiency of TCCN. In addition, TCCN uses dynamic routing mechanism that can retain more valid traffic characteristics. In parallel, this paper proposes a new loss function to improve the generalization performance of TCCN. The results show that TCCN can show good performance in different experimental scenarios. After effective training, TCCN can also show high detection accuracy in new datasets, and the generalization ability of the model reaches a relatively excellent level.
Yafei Sang, Zhenyu Cheng 0001, Tianning Zang
ICC4
2023 Explainable Text Classification via Attentive and Targeted Mixing Data Augmentation
abstract
Mixing data augmentation methods have been widely used in text classification recently. However, existing methods do not control the quality of augmented data and have low model explainability. To tackle these issues, this paper proposes an explainable text classification solution based on attentive and targeted mixing data augmentation, ATMIX. Instead of selecting data for augmentation without control, ATMIX focuses on the misclassified training samples as the target for augmentation to better improve the model's capability. Meanwhile, to generate meaningful augmented samples, it adopts a self-attention mechanism to understand the importance of the subsentences in a text, and cut and mix the subsentences between the misclassified and correctly classified samples wisely. Furthermore, it employs a novel dynamic augmented data selection framework based on the loss function gradient to dynamically optimize the augmented samples for model training. In the end, we develop a new model explainability evaluation method based on subsentence attention and conduct extensive evaluations over multiple real-world text datasets. The results indicate that ATMIX is more effective with higher explainability than the typical classification models, hidden-level, and input-level mixup models.
Songhao Jiang, Yan Chu 0001, Zhengkui Wang, Tianxing Ma, Wenxuan Lu, Tianning Zang
IJCAI7
2023 MTCD-Model: A Two-Layer Model for Malicious Traffic Classification and Detection Based on Hierarchical Feature Learning
abstract
The rapid growth of cyber world and higher awareness of security in recent years have contributed to a significant demand in classification and detection of malicious traffic. Neural network is considered one of the effective methods. However, existing methods need to be improved. For example, the model performance is influenced by over dependance on manual design and extraction of feature. In addition, truncating or zero-complementing the traffic data results in loss of key traffic information or irrelevant input, which in turn affects the model performance in classifying and detecting malicious traffic. Motivated by these considerations and demands, this paper proposes a Malicious Traffic Classification and Detection Model (MTCD-Model), a two-layer model based on hierarchical feature learning. This model exploits both the CNN and Bi-SRU to learn the features of raw traffic data in indefinite length by hierarchical learning method, and achieve the classification and detection of malicious traffic with capsule network. The experimental results, based on the primary dataset TCD-2022, show that the F1-Score of MTCD-Model can reach 98.62, while the performance remains stable in different experimental scenarios. In addition, MTCD-Model generates different degrees of improvement in various evaluation metrics compared with the control model.
Zhenyu Cheng 0001, Tianning Zang
IJCNN3
2023 Facing Unknown: Open-World Encrypted Traffic Classification Based on Contrastive Pre-Training
abstract
Traditional Encrypted Traffic Classification (ETC) methods face a significant challenge in classifying large volumes of encrypted traffic in the open-world assumption, i.e., simultaneously classifying the known applications and detecting unknown applications. We propose a novel Open-World Contrastive Pre-training (OWCP) framework for this. OWCP performs contrastive pre-training to obtain a robust feature representation. Based on this, we determine the spherical mapping space to find the marginal flows for each known class, which are used to train GANs to synthesize new flows similar to the known parts but do not belong to any class. These synthetic flows are assigned to Softmax's unknown node to modify the classifier, effectively enhancing sensitivity towards known flows and significantly suppressing unknown ones. Extensive experiments on three datasets show that OWCP significantly outperforms existing ETC and generic open-world classification methods. Furthermore, we conduct comprehensive ablation studies and sensitivity analyses to validate each integral component of OWCP.
Xiang Li 0135, Beibei Feng, Tianning Zang, Jingrun Ma
ISCC3
2023 Topology construction method of anti-tracking network based on cross-domain decentralized gravity model
abstract
With the increasing threats of network tracking and information leakage, privacy protection has become a widely concern in the field of network security. As an important means of protecting the privacy of network users, anti-tracking networks have gradually become one of the important research directions. However, the existing topology structures of anti-tracking networks still have problems such as intra-domain aggregation and key nodes, which are vulnerable to attack, tracking and destruction, and can not meet the privacy requirements. Therefore, this paper proposes a topology construction method based on cross-domain decentralized gravity model (CDTC). Firstly, the model comprehensively considers the local neighbor information, the location information of nodes, and the path information between nodes to calculate the node attraction. Secondly, each node determines its link status with other nodes through the model by ranking its local nodes, achieving optimization of the network topology. Finally, experiments on typical network structures and open network datasets, the experimental results show that the proposed model has better decentralization, cross-domain, and anti-tracking performance.
Zhefeng Nan, Qian Qiang, Tianning Zang, Changbo Tian, Shuhe Liu
TrustCom3
2022 Malicious Blockchain Domain Detection Based on Heterogeneous Information Network
abstract
With the popularity of Blockchain Domain Name System (BDNS), more and more cybercriminals integrate Blockchain Domain Names (BDNs) into their infrastructure. Due to the anonymity and anti-censorship of BDNs, it is difficult to detect malicious activities based on BDNs, posing a serious threat to network security. In this paper, we propose a novel method to detect malicious BDNs. First, we extract 16 statistical features of domain names. Second, we construct a Heterogeneous Information Network (HIN) of BDNS, which can use malicious traditional domain names as supplementary data. Then we associate domain names by meta-paths in the HIN and build an association graph of domain names. To better characterize domain names, we use the graph convolutional network algorithm to fuse the statistical features of domain names in the association graph. Finally, we detect malicious BDNs by the neural network algorithm. Compared with the existing methods, the experimental results show that our method can accurately detect malicious BDNs with the F1 score of 0.9901 and discover more unknown malicious BDNs from the dataset.
Songhao Jiang, Tianning Zang
GLOBECOM4
2022 Cost-Effective Malware Classification Based on Deep Active Learning
Qian Qiang, Tianning Zang, Mian Cheng, Quanbo Pan, Zisen Qi
SecureComm4
2022 Hidden Path: Understanding the Intermediary in Malicious Redirections
abstract
URL redirection has become an important tool for adversaries to cover up their malicious campaigns. In this paper, we conduct the first large-scale measurement study on how adversaries leverage URL redirection to circumvent security checks and distribute malicious content in practice. To this end, we design an iteratively running framework to mine the domains used for malicious redirections constantly. First, we use a bipartite graph-based method to dig out the domains potentially involved in malicious redirections from real-world DNS traffic. Then, we dynamically crawl these suspicious domains and recover the corresponding redirection chains from the crawler’s performance log. Based on the collected redirection chains, we analyze the working mechanism of various malicious redirections, involving the abused modes and methods, and highlight the pervasiveness of node sharing. Notably, we find a new redirection abuse, redirection fluxing, which is abused to enhance the concealment of malicious sites by introducing randomness into the redirection. Our case studies reveal the adversary’s preference for abusing JavaScript methods to conduct redirection, even by introducing time-delay and fabricating user clicks to simulate normal users.
Yuwei Zeng, Xunxun Chen, Tianning Zang
IEEE Trans. Inf. Forensics Secur.4
2021 Mobile Encrypted Traffic Classification Based on Message Type Inference
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008
CollaborateCom (1)2
2021 Inspector: A Semantics-Driven Approach to Automatic Protocol Reverse Engineering
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Yipeng Wang 0001
CollaborateCom (1)2
2021 Incremental Learning for Mobile Encrypted Traffic Classification
abstract
With the rising popularity of mobile networks and applications, network traffic classification has gradually become essential to mobile network management and cyberspace security. Existing state-of-the-art methods have achieved high accuracy in the closed-world mobile encrypted traffic classification, where the classifier only needs to process the classes seen in the training. When we update the dataset with new mobile applications, these methods must retrain a new classifier from scratch to learn the knowledge of all applications because directly fine-tuning the existing classifier would lead to the catastrophic forgetting problem. Thus, it is challenging to incrementally add new applications to the classification system while preserving the learned knowledge of the existing classifier. To tackle this issue, we propose an incremental learning framework based on the one vs rest (OvR) strategy and neural network classifiers. Moreover, we adopt a sample selection algorithm to balance the conflict between the growing training effort caused by new applications and the high classification accuracy. The experimental results demonstrate that our proposed framework achieves incremental learning with high classification accuracy like the closed-world method, and the selection algorithm significantly reduces training efforts to meet the dataset scale control and classification accuracy requirement in the lifetime incremental learning.
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Linshu Ouyang
ICC2
2021 Winding Path: Characterizing the Malicious Redirection in Squatting Domain Names
Yuwei Zeng, Xunxun Chen, Tianning Zang, Haiwei Tsang
PAM3
2021 Finding disposable domain names: A linguistics-based stacking approach
Yuwei Zeng, Xiao-chun Yun, Xunxun Chen, Boquan Li 0002, Haiwei Tsang, Yipeng Wang 0001, Tianning Zang, Yongzheng Zhang 0002
Comput. Networks7
2020 Leopard: Understanding the Threat of Blockchain Domain Name Based Malware
Zhangrong Huang, Tianning Zang
PAM3
2020 MalPortrait: Sketch Malicious Domain Portraits Based on Passive DNS Data
abstract
Malicious domain detection is of great significance for cybersecurity. Most prior works detect malicious domains based on individual features, which are only related to the attributes of domains themselves and can be easily changed to avoid detection. To solve the problem, we propose a novel system called MalPortrait, which combines individual features and association information of domains to detect malicious domains. In MalPortrait, we show the association information among domains by a domain association graph where vertices represent domains and edges connect domains resolved to the same IP. Based on the graph, we combine individual features (e.g., string-based, network-based) of each domain and its association information to generate new features. Compared with individual features, the new features are harder to be tampered with and can help determine whether a domain is malicious from a more comprehensive perspective. We evaluate MalPortrait on the passive DNS traffic collected from real-world large ISP networks. Our experimental results show that MalPortrait can accurately identify malicious domain names with a precision of 96.8% and a recall of 95.5%. Compared with prior works, MalPortrait performs better and hardly relies on additional knowledge (e.g., IP reputation, Domain whois).
Zhizhou Liang, Tianning Zang, Yuwei Zeng
WCNC2
2020 Khaos: An Adversarial Neural Network DGA With High Anti-Detection Ability
abstract
A botnet is a network of remote-controlled devices that are infected with malware controlled by botmasters in order to launch cyber attacks. To evade detection, the botmaster frequently changes the domain name of his Command and Control (C&C) server. Notice that most of these types of domain names are generated by domain generation algorithms (DGAs). In this paper, we propose Khaos, a novel DGA with high anti-detection ability based on neural language models and the Wasserstein Generative Adversarial Network (WGAN). The key insight of our research is that real domain names are composed of readable syllables and acronyms, and thus we can arrange syllables and acronyms using neural language models to mimic real domain names. In Khaos, we first find the most common n-grams in real domain names, then tokenize these domain names into n-grams, and finally synthesize new domain names after learning arrangements of n-grams from real domain names. We carry out experiments using a variety of state-of-the-art DGA detection approaches: the statistics-based, the distribution-based, the LSTM-based and the graph-based detection approach. Our experimental results show that the average distance for detecting Khaos under the distribution-based detection approach is 0.64, the AUCs of Khaos under the statistics-based and the LSTM-based detection approach are 0.76 and 0.57, respectively, and the precision of Khaos under the graph-based detection approach is 0.68. Our work proves that the existing detection approaches have big troubles in detecting Khaos, and Khaos has better anti-detection ability than state-of-the-art DGAs. In addition, we find that training the existing detection approach on a dataset including the domain names generated by Khaos can improve its detection ability.
Xiao-chun Yun, Yipeng Wang 0001, Tianning Zang, Yuan Zhou 0008, Yongzheng Zhang 0002
IEEE Trans. Inf. Forensics Secur.4
2019 Nemesis: Detecting Algorithmically Generated Domains with an LSTM Language Model
Dunsheng Yuan, Tianning Zang
CollaborateCom3
2019 Pontus: A Linguistics-Based DGA Detection System
abstract
Many botmasters use domain generation algorithms (DGA) to generate a host of malicious algorithmically- generated domains (mAGDs) and then choose several mAGDs for actual command and control (C2) communication. The botmasters use different seeds (e.g., timestamp) to generate different mAGDs, which makes the communication mechanism resilient to blacklisting. Thus the botmasters can hide C2 channels very well. If we can quickly detect these mAGDs from DNS traffic, we will effectively block the communication. In this paper, we propose a novel system, called Pontus, to detect mAGDs from DNS traffic. Pontus extract features exclusively from the individual domain names. We compare Pontus with the state-of-the-art system and find that Pontus improves the precision by at least 4.7%.
Dingkui Yan, Huilin Zhang, Yipeng Wang 0001, Tianning Zang, Yuwei Zeng
GLOBECOM4
2019 A Comprehensive Measurement Study of Domain-Squatting Abuse
abstract
Domain-squatting abuse refers to the premeditated attempt by an attacker to register perceptively confusing domain names thereby tricking visitors into querying them. There are totally five squatting types have been investigated so far, namely typo-squatting, bit-squatting, homograph-squatting, sound-squatting, and combo-squatting. Existing researches only focus on one specific squatting type and never explore the relationship among them. In this paper, we perform the first comprehensive measurement study of domain-squatting abuse. We select 786 the most queried domains, and hunt for squatting abuses against them in ISP-level DNS traffic. We find that although typo-squatting accounts for most of squatting domains, combo-squatting are able to attract more traffic. Our further case studies show that parking ads is still the most important way for attackers to make profits. The only exception is combo-squatting, in which squatters tend to leverage the reputation of squatted domains to develop their own business. It is worth noting that some squatting domains are even used to deliver malware. Moreover, the Alexa ranks of certain squatting domains have already surpassed the original domains. These results clearly call for the need to better protect the intellectual property of domain names.
Yuwei Zeng, Tianning Zang, Yongzheng Zhang 0002, Xunxun Chen, Yipeng Wang 0001
ICC2
2019 Rethinking Encrypted Traffic Classification: A Multi-Attribute Associated Fingerprint Approach
abstract
With the unprecedented prevalence of mobile network applications, cryptographic protocols, such as the Secure Socket Layer/Transport Layer Security (SSL/TLS), are widely used in mobile network applications for communication security. The proven methods for encrypted video stream classification or encrypted protocol detection are unsuitable for the SSL/TLS traffic. Consequently, application-level traffic classification based networking and security services are facing severe challenges in effectiveness. Existing encrypted traffic classification methods exhibit unsatisfying accuracy for applications with similar state characteristics. In this paper, we propose a multiple-attribute-based encrypted traffic classification system named Multi-Attribute Associated Fingerprints (MAAF). We develop MAAF based on the two key insights that the DNS traces generated during the application runtime contain classification guidance information and that the handshake certificates in the encrypted flows can provide classification clues. Apart from the exploitation of key insights, MAAF employs the context of the encrypted traffic to overcome the attribute-lacking problem during the classification. Our experimental results demonstrate that MAAF achieves 98.69% accuracy on the real-world traceset that consists of 16 applications, supports the early prediction, and is robust to the scale of the training traceset. Besides, MAAF is superior to the state-of-the-art methods in terms of both accuracy and robustness.
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Yipeng Wang 0001
ICNP2
2019 A Linguistics-based Stacking Approach to Disposable Domains Detection
abstract
More Internet services tend to collect the one-time information from clients via DNS queries. Notably, the uncertainty of such transient information makes these domain names be queried only once in their lifetime. This type of domain is called disposable domain. Although they are not malicious, the efficiency of DNS infrastructures will still be affected by their ever-increasing number. In this paper, we propose Vogers, a linguistics-based stacking model, to detect the disposable domains. Our evaluation demonstrates that Vogers decreases the false positive rate by more than 19%, compared with the prior art, while maintaining the true positive rate above 98.9%.
Yuwei Zeng, Yongzheng Zhang 0002, Tianning Zang, Xunxun Chen, Yipeng Wang 0001
ICNP3
2019 Themis: A Novel Detection Approach for Detecting Mixed Algorithmically Generated Domains
abstract
As DGA (Domain Generation Algorithm) detection technologies and systems become more and more complex, more types of AGD (Algorithmically Generated Domain) appear: Dictionary-based AGD, Hash-based AGD, etc. This paper applies deep learning to the field of network security, proposes a lightweight AGD detection approach, Themis, which can classify domain names into legitimate domain names or AGDs through domain name strings. Themis combines WordNet and GRU to capture the different characteristics of legitimate domain name and AGD for classification. Compared with the prior art, Themis has two differences: 1) Themis is the first approach to detect mixed AGD (Arithmetic-based and Dictionary-based); 2) Themis performs well in detecting unknowns AGD.
Chaoyi Zheng, Qian Qiang, Tianning Zang, Wen-Han Chao
MSN3
2018 A Stacking Approach to Objectionable-Related Domain Names Identification by Passive DNS Traffic (Short Paper)
Yongzheng Zhang 0002, Tianning Zang, Zhizhou Liang, Yipeng Wang 0001
CollaborateCom3
2018 The Rapid Extraction of Suspicious Traffic from Passive DNS
Tianning Zang, Yuqing Lan
ICISSP2
2018 Information Propagation Prediction Based on Key Users Authentication in Microblogging
abstract
In microblogging, key users are a significant factor for information propagation. Key users can affect information propagation size while retweeting the information. In this paper, to predict information propagation, we propose a novel linear model based on key users authentication. This model mines key users to dynamically improve the linear model while predicting information propagation. So our model can not only predict information propagation but also mine key users. Experimental results show that our model can achieve remarkable efficiency on predicting information propagation problem in real microblogging networks. At the same time, our model can find the key users who affect information propagation.
Miao Yu 0006, Yongzheng Zhang 0002, Tianning Zang, Yipeng Wang 0001
Secur. Commun. Networks3
2017 Rethinking robust and accurate application protocol identification
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Tianning Zang
Comput. Networks5
2014 Detecting Malicious Behaviors in Repackaged Android Apps with Loosely-Coupled Payloads Filtering Scheme
Yongzheng Zhang 0002, Tianning Zang
SecureComm (1)3
2010 Cooperative Work Systems for the Security of Digital Computing Infrastructure
abstract
On open digital computing infrastructure, various large-scale and complicated malicious behaviors are increasingly threatening the security of digital computing infrastructure. In this paper, a Cooperative Work Model (CRM) is presented by extending the conceptions of the Universal Turing Machine to deal with the threats. Then the Cooperative Work System Framework (CWSF) is derived from the model. Based on the framework, two practical Cooperative Work Systems (CWSs) are developed to track and analyze the Botnet and DDoS on digital computing infrastructure respectively. The systems collectively use and coordinate various monitoring systems distributed in the back-bone network of the infrastructure. The experimental results of analyzing typical security events show that the framework and systems are efficient and effective to collaboratively use diverse related network systems for monitoring and analyzing the large-scale network events. Currently, the systems are running steadily in the monitoring environment of a large-scale back-bone network.
Tianning Zang, Xiao-chun Yun, Tianyi Zang, Yongzheng Zhang 0002, Chaoguang Men
ICPADS1
2008 A Survey of Alert Fusion Techniques for Security Incident
abstract
Security incident have been imposing tremendous threats on todaypsilas network information system. To protect this information system from the increasing threat of intrusion, various kinds of detection systems and sensors for security incident have been developed. The main disadvantages of current systems and sensors are a high false detection rate and the lack of post-incident decision support capability. To minimize these drawbacks, various alert fusion technologies have been proposed in the recent years. This paper presents a general summary of these technologies. Basic models and key technologies of alert fusion are analyzed and discussed. Moreover, important aggregation and correlation algorithms are discussed. Finally, we make concluding remarks by predicting the development tendencies of alert correlation technologies.
Tianning Zang, Xiao-chun Yun, Yongzheng Zhang 0002
WAIM1