Yongzheng Zhang 0002

dblp:z/YZZhang2 · DBLP profile ↗
← Back
94ranked-venue papers
1as first author
29since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 27 · 1 first-author · 11 since 2021Computer networks · 25 · 7 since 2021Human-computer interaction and ubiquitous computing · 15 · 7 since 2021Systems, architecture and hardware · 10Artificial intelligence and machine learning · 7 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 Let model keep evolving: Incremental learning for encrypted traffic classification
Xiang Li 0135, Jiang Xie 0004, Qige Song, Yafei Sang, Yongzheng Zhang 0002, Tianning Zang
Comput. Secur.5
2024 SepBIN: Binary Feature Separation for Better Semantic Comparison and Authorship Verification
abstract
Binary semantic comparison and authorship verification are critical in many security applications. They respectively focus on the functional semantic features and developers’ programming style features of binary code, which are usually mixed without clear demarcation. Recently, researchers have proposed learning-based approaches for intelligent binary analysis. They generally addressed single tasks with hand-crafted feature sets or neural binary encoders, which suffer performance bottlenecks due to the noise in mixed features. This paper proposesSepBIN, a novel neural network framework that exploits the intrinsic correlation of binary semantic comparison and authorship verification tasks and automatically separates semantic and stylistic binary features. We first construct a strong backbone binary encoder, then utilize preliminary decomposition subnets and the flexible gating-based feature fusion mechanism to distill pure semantic-related and style-related binary representations, and further improve their quality by a feature reconstruction module. The overallSepBINmodel is optimized by a multi-objective joint optimization strategy. We conduct extensive experiments on Google Code Jam (GCJ) datasets in different languages and scales. Results show thatSepBINsimultaneously benefits binary semantic comparison and authorship verification tasks through the effective binary semantic-style feature separation mechanism, and provides multi-perspectives interpretability for the performance gains. For state-of-the-art approaches with different binary encoders,SepBINcan adaptively improve them with the designed separation modules. Furthermore, we adopt a pretraining-finetuning strategy to effectively transferSepBIN’s separation capability in real-world applications, including APT malware homology detection and binary semantic comparison against code obfuscations.
Qige Song, Yafei Sang, Yongzheng Zhang 0002
IEEE Trans. Inf. Forensics Secur.3
2023 Listen to Minority: Encrypted Traffic Classification for Class Imbalance with Contrastive Pre-Training
abstract
Mobile Internet has profoundly reshaped modern lifestyles in various aspects. Encrypted Traffic Classification (ETC) naturally plays a crucial role in managing mobile Internet, especially with the explosive growth of mobile apps using encrypted communication. Despite some existing learning-based ETC methods showing promising results, three-fold limitations still remain in real-world network environments, i) label bias caused by traffic class imbalance, ii) traffic homogeneity caused by component sharing, and iii) training with reliance on sufficient labeled traffic. None of the existing ETC methods can address all these limitations. In this paper, we propose a novel Pre-trAining Semi-Supervised ETC framework, dubbed PASS. Our key insight is to resample the original train dataset and perform contrastive pre-training without using individual app labels directly to avoid label bias issues caused by class imbalance, while obtaining a robust feature representation to differentiate overlapping homogeneous traffic by pulling positive traffic pairs closer and pushing negative pairs away. Meanwhile, PASS designs a semi-supervised optimization strategy based on pseudo-label iteration and dynamic loss weighting algorithms in order to effectively utilize massive unlabeled traffic data and alleviate manual train dataset annotation workload. PASS outperforms state-of-the-art ETC methods and generic sampling approaches on four public datasets with significant class imbalance and traffic homogeneity, remarkably pushing the F1 of Cross-Platform215 with 1.31%$\uparrow$, ISCX-17 with 9.12%$\uparrow$. Furthermore, we validate the generality of the contrastive pre-training and pseudo-label iteration components of PASS, which can adaptively benefit ETC methods with diverse feature extractors.
Xiang Li 0135, Juncheng Guo, Qige Song, Jiang Xie 0004, Yafei Sang, Yongzheng Zhang 0002
SECON7
2023 Toward IoT device fingerprinting from proprietary protocol traffic via key-blocks aware approach
Yafei Sang, Jisong Yang, Yongzheng Zhang 0002
Comput. Secur.3
2023 Encrypted TLS Traffic Classification on Cloud Platforms
abstract
Nowadays, encryption technology has been widely used to protect user privacy. With the explosive growth of mobile Internet, encrypted TLS traffic rises sharply and occupies a great share of current Internet traffic. In reality, the classification of encrypted TLS traffic on cloud platforms brings a new challenge to traditional encrypted traffic classification methods, because some information such as certificates in the TLS flows is no longer effective. In this paper, we apply deep learning technology to the problem of encrypted TLS traffic classification on cloud platforms, and propose NeuTic, which takes the packet sequence of each TLS flow as the input, and effectively classifies raw TLS flows generated by many “cloud” applications. Our approach is able to automatically capture the long-range dependencies between elements in the packet sequences for robust and accurate encrypted TLS traffic classification. In NeuTic, we first convert each TLS flow into three attribute sequences. Then, we train a multi-application traffic classification model using our newly designed deep learning model. Finally, we use the well-trained classification model to classify new incoming TLS flows. We conduct comprehensive experiments on real-world application traces covering multiple “cloud” applications from three different companies. In addition, we compare our experimental results of NeuTic with two deep learning-based methods for encrypted traffic classification. NeuTic outperforms the state-of-the-art approaches in classification accuracy.
Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002
IEEE/ACM Trans. Netw.3
2022 An Adaptive Ensembled Neural Network-Based Approach to IoT Device Identification
Jingrun Ma, Yafei Sang, Yongzheng Zhang 0002, Beibei Feng, Yuwei Zeng
CollaborateCom (2)3
2022 Evading Encrypted Traffic Classifiers by Transferable Adversarial Traffic
Hanwu Sun, Chengwei Peng, Yafei Sang, Yongzheng Zhang 0002, Yujia Zhu
CollaborateCom (2)5
2022 A Longitudinal Measurement and Analysis of Pink, a Hybrid P2P IoT Botnet
Binglai Wang, Yafei Sang, Yongzheng Zhang 0002, Ruihai Ge
CollaborateCom (2)3
2022 A Dual-Branch Self-attention Method for Mobile Malware Detection via Network Traffic
abstract
The desperate increase of mobile malware has constituted a severe threat to user privacy, economic life, and cyberspace security. Existing anti-malware solutions have no-ticeable weaknesses due to the adoption of content analysis-based approaches. The main limitation of these approaches is that they rely on careful expert engineering and professional handcrafted input features. Some researchers have tried to solve this limitation by using deep learning models to automatically learn feature representations from raw traffic. In this paper, we explore a deep learning detection framework based on self-attention to discriminate between malicious and benign network traffic. As a major advantage with respect to the state-of-the-art methods, we point out that the attention mechanism can better learn the underlying features of malicious traffic in terms of flows and bytes. We design a dual-branch deep learning method that consists of a flow importance-discrimination branch and a byte importance-discrimination branch. The flow importance-discrimination branch calculates the attentions between flows to obtain the feature contributions of different flows, and the byte importance-discrimination branch builds the feature contributions of diverse bytes by considering the connections among all bytes of network payload. Both flow features and byte features are combined to enhance the representation ability of network traffic behaviors generated by mobile applications (apps). We evaluate proposed method using a publicly available dataset including 55,992 malicious traffic traces and 47,779 benign traffic traces. The experimental results demonstrate that our method is able to identify malicious apps with high accuracy, outperforming the baseline methods and the popular deep-learning models.
Ruihai Ge, Yongzheng Zhang 0002, Guoqiao Zhou
IJCNN2
2022 Autonomous Anti - interference Identification of $\text{IoT}$ Device Traffic based on Convolutional Neural Network
abstract
Network traffic classification plays a vital role in many fields such as intrusion detection, network management, and network security. As the proportion of IoT device traffic increases, many approaches to identifying IoT device types through traffic have emerged. Specifically, Deep Learning (DL) has been proven to be a more efficient approach for encrypted traffic identification than other traditional methods. However, most existing classification models are created in static datasets from the closed world, so they can only classify within a limited domain. In this case, interfering traffic in the open world is easily misidentified by classifiers as IoT device traffic. An autonomous framework is proposed to tackle this issue, effectively identifying the device type according to the grayscale graph generated by packet payload and automatically updating to adapt to the unknown environment in the open world. The core of the proposed framework consists of a packet graph-vector transformer, a CNN-based classifier, and an autonomous optimizer. The optimizer can filter interfering data and optimize the model by updating the training dataset. We comprehensively evaluated the proposed framework on two datasets, one taken from the UNSW IoT traces and the other collected by our experiments, containing traffic generated from two devices and three open-world scenarios. The results demonstrate that the proposed framework can update the training dataset by unsupervised filtering interference packets, enabling the model to automatically suit complex environments for accurate and robust IoT device type identification in the open world.
Shuhe Liu, Yongzheng Zhang 0002, Yipeng Wang 0001
IJCNN3
2022 Multi-relational Instruction Association Graph for Cross-Architecture Binary Similarity Comparison
Qige Song, Yongzheng Zhang 0002
SecureComm2
2022 A longitudinal Measurement and Analysis Study of Mozi, an Evolving P2P IoT Botnet
abstract
Due to the remarkable diversity and ubiquity of IoT devices, many IoT botnets like Mirai and Hajime are inclined to infect vulnerable embedded devices to achieve the purpose of scale expansion and future attacks. Nowadays, a novel, evolving P2P botnet in the wild, named Mozi, targets numerous devices but substantially differs in many aspects involving self-proliferation, network control, and distribution situations. To uncover the curtain, throughout detailed measurement–a novel active probing approach (namely Bawa) for tracking Mozi’s dynamic topology infrastructure and passive collection of Mozi binaries and its download URLs, we provide a holistic view of Mozi, including its scale and geographic distribution. The dataset1collected in a passive mode is available to assist in tracking and mitigating the Mozi botnet’s expansion.
Binglai Wang, Yafei Sang, Yongzheng Zhang 0002
TrustCom3
2022 BinMLM: Binary Authorship Verification with Flow-aware Mixture-of-Shared Language Model
abstract
Binary authorship analysis is a significant problem in many software engineering applications. In this paper, we formulate a binary authorship verification task to accurately reflect the real-world working process of software forensic experts. It aims to determine whether an anonymous binary is developed by a specific programmer with a small set of support samples, and the actual developer may not belong to the known candidate set but from the wild. We propose an effective binary authorship verification framework, BinMLM. BinMLM trains the RNN language model on consecutive opcode traces extracted from the control-flow-graph (CFG) to characterize the candidate developers' programming styles. We build a mixture-of-shared architecture with multiple shared encoders and author-specific gate layers, which can learn the developers' combination preferences of universal programming patterns and alleviate the problem of low training resources. Through an optimization pipeline of external pre-training, joint training, and fine-tuning, our framework can eliminate additional noise and accurately distill developers' unique styles. Extensive experiments show that BinMLM achieves promising results on Google Code Jam (GCJ) and Codeforces datasets with different numbers of programmers and supporting samples. It significantly outperforms the baselines built on the state-of-the-art feature set (4.73% to 19.46% improvement) and remains robust in multi-author collaboration scenarios. Furthermore, Bin-MLM can perform organization-level verification on a real-world APT malware dataset, which can provide valuable auxiliary information for exploring the group behind the APT attack.
Qige Song, Yongzheng Zhang 0002, Linshu Ouyang
SANER2
2022 Analysis and Detection against Network Attacks in the Overlapping Phenomenon of Behavior Attribute
Jiang Xie 0004, Yongzheng Zhang 0002, Peishuai Sun
Comput. Secur.3
2022 Detecting unknown HTTP-based malicious communication behavior via generated adversarial flows and hierarchical traffic features
Xiao-chun Yun, Jiang Xie 0004, Yongzheng Zhang 0002, Peishuai Sun
Comput. Secur.4
2022 Inter-BIN: Interaction-Based Cross-Architecture IoT Binary Similarity Comparison
abstract
The big wave of Internet of Things (IoT) malware reflects the fragility of the current IoT ecosystem. Research has found that IoT malware can spread quickly on devices of different processer architectures, which leads our attention to cross-architecture binary similarity comparison technology. The goal of binary similarity comparison is to determine whether the semantics of two binary snippets is similar. Existing learning-based approaches usually learn the representations of binary code snippets individually and perform similarity matching based on the distance metric, without considering interbinary semantic interactions. Moreover, they often rely on the large-scale external code corpus for instruction embeddings pretraining, which is heavyweight and easy to suffer the out-of-vocabulary (OOV) problem. In this article, we propose an interaction-based cross-architecture IoT binary similarity comparison system,Inter-BIN. Our key insight is to introduce interaction between instruction sequences by co-attention mechanism, which can flexibly perform soft alignment of semantically related instructions from different architectures. And we design a lightweight multifeature fusion-based instruction embedding method, which can avoid the heavy workload and the OOV problem of previous approaches. Extensive experiments show thatInter-BINcan significantly outperform state-of-the-art approaches on cross-architecture binary similarity comparison tasks of different input granularities. Furthermore, we present an IoT malware function matching data set from real network environments, CrossMal, containing 1878437 cross-architecture reuse function pairs. Experimental results on CrossMal prove thatInter-BINis practical and scalable on real-world binary similarity comparison collections.
Qige Song, Yongzheng Zhang 0002, Binglai Wang
IEEE Internet Things J.2
2022 A Multi-Scale Feature Attention Approach to Network Traffic Classification and Its Model Explanation
abstract
Network traffic classification, the task of associating network traffic with their generating application protocols or applications, is valuable for the control, allocation, and management of resources in today’s TCP/IP networks. In this paper, we propose Ulfar, a multi-scale feature attention approach to network traffic classification, which uses convolutional neural networks (CNN) as the building block of the deep packet analysis model. In Ulfar, we take only one packet per flow for network traffic classification. Ulfar is based on the key insight that format-related bytes appear at fixed offsets or in a specific pattern in the IP packet, and these format-related bytes are important for accurate network traffic classification. Our neural network model can automatically recover the format-related bytes by building high-level, multi-scale${n}$-gram features from raw byte sequences. In addition, at the representation learning side, we try to understand what patterns and signatures our neural network model learns from network traffic. We evaluate Ulfar using two publicly available datasets, and our experimental results show that Ulfar can conduct accurate network traffic classification. Also, we compare the results of Ulfar with four state-of-the-art approaches, and find that Ulfar has the ability to classify network traffic more accurately.
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Xin Liu 0002
IEEE Trans. Netw. Serv. Manag.3
2021 Mobile Encrypted Traffic Classification Based on Message Type Inference
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008
CollaborateCom (1)3
2021 Inspector: A Semantics-Driven Approach to Automatic Protocol Reverse Engineering
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Yipeng Wang 0001
CollaborateCom (1)3
2021 Topology Self-optimization for Anti-tracking Network via Nodes Distributed Computing
Changbo Tian, Yongzheng Zhang 0002
CollaborateCom (1)2
2021 Exploiting Heterogeneous Information for IoT Device Identification Using Graph Convolutional Network
Jisong Yang, Yafei Sang, Yongzheng Zhang 0002, Chengwei Peng
CollaborateCom (1)3
2021 Incremental Learning for Mobile Encrypted Traffic Classification
abstract
With the rising popularity of mobile networks and applications, network traffic classification has gradually become essential to mobile network management and cyberspace security. Existing state-of-the-art methods have achieved high accuracy in the closed-world mobile encrypted traffic classification, where the classifier only needs to process the classes seen in the training. When we update the dataset with new mobile applications, these methods must retrain a new classifier from scratch to learn the knowledge of all applications because directly fine-tuning the existing classifier would lead to the catastrophic forgetting problem. Thus, it is challenging to incrementally add new applications to the classification system while preserving the learned knowledge of the existing classifier. To tackle this issue, we propose an incremental learning framework based on the one vs rest (OvR) strategy and neural network classifiers. Moreover, we adopt a sample selection algorithm to balance the conflict between the growing training effort caused by new applications and the high classification accuracy. The experimental results demonstrate that our proposed framework achieves incremental learning with high classification accuracy like the closed-world method, and the selection algorithm significantly reduces training efforts to meet the dataset scale control and classification accuracy requirement in the lifetime incremental learning.
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Linshu Ouyang
ICC3
2021 LIFH: Learning Interactive Features from HTTP Payload using Image Reconstruction
abstract
The complexity and intelligence of the attacks towards the application layer have raised to an unprecedented level. HyperText Transfer Protocol (HTTP), as the widely used application layer protocol, is part of the main vectors for various malicious attacks. The previous detection based on Deep Packet Inspection (DPI) relies heavily on packets, which leads to insufficient detection and a high false alarm rate. In this paper, we propose LIFH, a deep neural network model equipped with interactive information for detecting application-layer attacks. Firstly, the image reconstruction method is designed to reconstruct the HTTP traffic session into an image. Then, the latent features, instead of explicit features which are typically used in machine learning models, are extracted by HTTP-CNN in order to respond against forgery attacks. Finally, the high-level features are further fed to multi-classifiers to identify the traffic involved in malicious activities. We make exclusive experiments and evaluate the performance of LIFH on the standard dataset CICIDS_2017 and IIE_data collected from critical web servers. The results demonstrate that the proposed model can significantly improve the performance of malicious traffic detection with an accuracy of 99.07% and a false positive rate of 0.40% which is superior to the state of the arts.
Jinbu Geng, Yongzheng Zhang 0002, Zhenyu Cheng 0001
ICC3
2021 Phishing Web Page Detection with Semi-Supervised Deep Anomaly Detection
Linshu Ouyang, Yongzheng Zhang 0002
SecureComm (2)2
2021 Phishing Web Page Detection with HTML-Level Graph Neural Network
abstract
Phishing web page is one of the most serious threats to the users of the Internet. Traditional phishing web page detection methods rely on manually designed features. Recently, deep learning-based methods using HTML as input have achieved significant detection performance improvement. They usually treat HTML codes as sequences of characters and utilize Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) for classification. However, CNN and RNN typically can only extract local features in the HTML code sequences while failing to model the long-range semantics that is crucial for phishing detection. In this paper, we propose a novel Graph Neural Network (GNN) based phishing web page detection method that can effectively utilize the inherent structural information of HTML to capture the long-range semantics. We first naturally represent an HTML as a graph according to its Document Object Model (DOM) and utilize RNN to extract the local features of node attributes. Then we adopt GNN to model the long-range relations between nodes based on these local features and the graph structure. Our proposed model combines the advantage of RNN and GNN to better understand the intention of HTML codes. Extensive experiments on a real-world dataset demonstrate that the accuracy of our method outperforms other state-of-the-art methods by a large margin.
Linshu Ouyang, Yongzheng Zhang 0002
TrustCom2
2021 DroidRadar: Android Malware Detection Based on Global Sensitive Graph Embedding
abstract
Android application markets face severe threats of malware attacks. Existing learning-based malware detection approaches rely on easily obfuscated features or unscalable sophisticated graph analysis techniques. In this paper, we propose DroidRadar, an accurate Android malware detection system based on lightweight graph embedding. The key insight of our method is constructing an entire Android application collection as a global graph schema and using sensitive APIs as bridge nodes to propagate inter-application information. We conduct statistical correlation analysis from different perspectives to model the application's usage pattern of sensitive APIs, then apply graph convolution network (GCN) to perform node embedding and malware detection. We evaluate DroidRadar on large scale datasets spanning nine years. Results show that DroidRadar has an average detection accuracy of 98.57% and a false-positive rate of 1.4 % on different time periods, which outperforms the state-of-the-art approaches, and it has strong robustness when detecting obfuscated malware variants.
Qige Song, Yongzheng Zhang 0002, Junliang Yao
TrustCom2
2021 A Feature-Flux Traffic Camouflage Method based on Twin Gaussian Process
abstract
Recent work has shown that the properties of network traffic may reveal some patterns (such as, payload size, packet interval, etc.) that can expose users' identities and their private information. The existing defense approaches, such as traffic morphing, protocol tunneling, still suffer from revealing the special traffic pattern. To address this problem, we propose a feature-flux traffic camouflage method (FFTC). FFTC forecasts the pattern of normal traffic via twin Gaussian process(TGP), and dynamically change the on-going traffic feature based on the learned traffic pattern to conceal the camouflaged traffic in the normal traffic. TGP-based traffic forecasting makes FFTC more sensitive to the feature dynamics of normal traffic. Then, the camouflaged traffic can always synchronize with the normal traffic pattern in real time. Furthermore, FFTC can learn multiple traffic patterns from different kinds of normal traffic, and dynamically change the camouflaged traffic pattern to achieve the feature-flux ability. From the experimental results, FFTC improves the indistinguishability of the camouflaged traffic and the normal traffic, and the feature-flux of the camouflaged traffic mitigates the traffic analysis attack effectively.
Changbo Tian, Yongzheng Zhang 0002
TrustCom2
2021 Finding disposable domain names: A linguistics-based stacking approach
Yuwei Zeng, Xiao-chun Yun, Xunxun Chen, Boquan Li 0002, Haiwei Tsang, Yipeng Wang 0001, Tianning Zang, Yongzheng Zhang 0002
Comput. Networks8
2021 A Novel Method to Prevent Misconfigurations of Industrial Automation and Control Systems
abstract
Configuration errors are among the dominant causes of system faults for the industrial automation and control systems (IACS). It is difficult to detect and correct such errors of IACS as there are various kinds of systems and devices with miscellaneous configuration specifications. In this article, we first propose a streaming algorithm to keep all the configuration changes in the limited memory space. When making a new configuration change, another novel streaming algorithm is proposed to search and return all the similar historical changes, which can be used to validate this new one. So far, we are the first to model the configuration changes of IACS as a data stream and apply the streaming similarity search in correcting configuration errors while overcoming the inherent unbounded-memory bottleneck. The theoretical correctness and complexity analyses are presented. Experiments with real and synthetic datasets confirm the theoretical analyses and demonstrate the effectiveness of the proposed method in preventing misconfigurations of IACS.
Yu Zhang 0095, Yani Ge, Peiran Yu, Jianzhong Zhang 0003, Yongzheng Zhang 0002, Thar Baker
IEEE Trans. Ind. Informatics5
2020 Joint Character-Level Word Embedding and Adversarial Stability Training to Defend Adversarial Text
abstract
Text classification is a basic task in natural language processing, but the small character perturbations in words can greatly decrease the effectiveness of text classification models, which is called character-level adversarial example attack. There are two main challenges in character-level adversarial examples defense, which are out-of-vocabulary words in word embedding model and the distribution difference between training and inference. Both of these two challenges make the character-level adversarial examples difficult to defend. In this paper, we propose a framework which jointly uses the character embedding and the adversarial stability training to overcome these two challenges. Our experimental results on five text classification data sets show that the models based on our framework can effectively defend character-level adversarial examples, and our models can defend 93.19% gradient-based adversarial examples and 94.83% natural adversarial examples, which outperforms the state-of-the-art defense models.
Yongzheng Zhang 0002, Yipeng Wang 0001, Zheng Lin 0001
AAAI2
2020 CgNet: Predicting Urban Congregations from Spatio-Temporal Data Using Deep Neural Networks
abstract
Predicting urban congregations can help in monitoring a variety of unusual group events, which is of great importance to public safety and traffic management in smart cities. However, it is very challenging because of complicated spatio-temporal correlations. In this article, we propose a deep neural network-based model, entitled CgNet, for urban congregations prediction. Firstly, we design three types of flows to present dependencies between regions among different timestamps to model the mobility of individuals. Secondly, CgNet utilizes four components, including spatial feature extraction, temporal feature extraction, external factors fusion and congregation feature fusion to collaboratively predict congregations. The combination of these components is capable of not only capturing the spatial and temporal correlations simultaneously, but also learning the essential relationships between three flows and congregations in each region at different stages. Finally, we evaluated the effectiveness of CgNet with extensive experimental study on real taxi trajectory data. The results demonstrate the advantages of our model beyond several baselines.
Tianran Chen, Yongzheng Zhang 0002, Yupeng Tuo
GLOBECOM2
2020 IncreAIBMF: Incremental Learning for Encrypted Mobile Application Identification
Yafei Sang, Mao Tian, Yongzheng Zhang 0002
ICA3PP (3)3
2020 IoTCMal: Towards A Hybrid IoT Honeypot for Capturing and Analyzing Malware
abstract
Nowadays, the emerging Internet-of-Things (IoT) emphasize the need for the security of network-connected devices. Additionally, there are two types of services in IoT devices that are easily exploited by attackers, weak authentication services (e.g., SSH/Telnet) and exploited services using command injection. Based on this observation, we propose IoTCMal, a hybrid IoT honeypot framework for capturing more comprehensive malicious samples aiming at IoT devices. The key novelty of IoTC-MAL is three-fold: (i) it provides a high-interactive component with common vulnerable service in real IoT device by utilizing traffic forwarding technique; (ii) it also contains a low-interactive component with Telnet/SSH service by running in virtual environment. (iii) Distinct from traditional low-interactive IoT honeypots[1], which only analyze family categories of malicious samples, IoTCMal primarily focuses on homology analysis of malicious samples. We deployed IoTCMal on 36 VPS1instances distributed in 13 cities of 6 countries. By analyzing the malware binaries captured from IoTCMal, we discover 8 malware families controlled by at least 11 groups of attackers, which mainly launched DDoS attacks and digital currency mining. Among them, about 60% of the captured malicious samples ran in ARM or MIPs architectures, which are widely used in IoT devices.
Binglai Wang, Yu Dou, Yafei Sang, Yongzheng Zhang 0002
ICC4
2020 TDAE: Autoencoder-based Automatic Feature Learning Method for the Detection of DNS tunnel
abstract
The DNS protocol is one of the most important network infrastructure protocols. The encrypted information based on this protocol will not be intercepted by the firewall, so the attacker uses this vulnerability to pass private data through the establishment of DNS tunnels and avoids the security inspection. In order to detect the DNS tunnel conveniently and effectively, we present a novel method that uses Autoencoder to learn latent representation of different datasets. Because the feature is not extracted manually, we show how Autoencoder(AE) can automatically learn the concept of semantic similarity among features of normal traffic. We propose a novel method named TDAE which can detect DNS tunnel traffics using Autoencoder algorithms. To verify the validity of our method, we select a labeled dataset and a public and unlabeled dataset as our training set. The experimental results show that the recall rate can exceed 0.9834 on the labeled dataset and 0.9313 on the SINGH-data [1].
Kemeng Wu, Yongzheng Zhang 0002
ICC2
2020 Gated POS-Level Language Model for Authorship Verification
abstract
Authorship verification is an important problem that has many applications. The state-of-the-art deep authorship verification methods typically leverage character-level language models to encode author-specific writing styles. However, they often fail to capture syntactic level patterns, leading to sub-optimal accuracy in cross-topic scenarios. Also, due to imperfect cross-author parameter sharing, it's difficult for them to distinguish author-specific writing style from common patterns, leading to data-inefficient learning. This paper introduces a novel POS-level (Part of Speech) gated RNN based language model to effectively learn the author-specific syntactic styles. The author-agnostic syntactic information obtained from the POS tagger pre-trained on large external datasets greatly reduces the number of effective parameters of our model, enabling the model to learn accurate author-specific syntactic styles with limited training data. We also utilize a gated architecture to learn the common syntactic writing styles with a small set of shared parameters and let the author-specific parameters focus on each author's special syntactic styles. Extensive experimental results show that our method achieves significantly better accuracy than state-of-the-art competing methods, especially in cross-topic scenarios (over 5\% in terms of AUC-ROC).
Linshu Ouyang, Yongzheng Zhang 0002, Yipeng Wang 0001
IJCAI2
2020 Unified Graph Embedding-Based Anomalous Edge Detection
abstract
Detecting anomalous edges in graph-structured data plays an important role in many fields such as finance, social network, and network security. Recently, graph embedding based anomaly detection methods show promising results. These methods typically encode graph structure information into vector representation and apply general anomaly detection methods. However, since the parameters in these two parts are learned separately with different objectives, the learned representation may contain some information irrelevant to the task. It would be ideal if we can combine representation learning and anomaly detection into one objective function to force the model to focus on learning task relevant patterns. In this paper, we propose a novel end-to-end neural network architecture that can accurately estimate the probability distribution of edges in the graph based on its local structure. An edge has a high chance to be considered an anomaly if the probability of its existence is low. Extensive experiments on several public datasets at different scales show that the accuracy and scalability of our method outperform other methods by a large margin.
Linshu Ouyang, Yongzheng Zhang 0002, Yipeng Wang 0001
IJCNN2
2020 A Feature Ensemble-based Approach to Malicious Domain Name Identification from Valid DNS Responses
abstract
Identifying malicious domain names in Internet activities has become an effective method to protect Internet users. Previous works have achieved great identification results, but they highly rely on historical Domain Name System (DNS) responses and external intelligence sources. Thus, they may fail to identify unknown domain name without any prior knowledge. In this paper, we propose Glacier, a feature ensemble-based approach to identifying malicious domain names from valid DNS responses. Glacier addresses the aforementioned problem by utilizing two types of features in domain name strings: the linguistical features and the statistical features. (1) Linguistical features are vector representations generated from the character sequences of domain names by a bidirectional long short-term memory (BiLSTM) neural network. It is worthy to notice that we modify the last BiLSTM layer to enhance the expressiveness of the linguistical features. (2) Statistical features are six manually designed statistics that represent the structural information of a domain name. Structural information can hardly be learnt by a BiLSTM neural network directly. Thus, combining statistical features with linguistical features can improve the effectiveness of malicious domain name identification. We evaluate the identification ability of Glacier on a real-world domain name data set. The best metrics of Glacier are an average accuracy of 90.86% and an average F1-score of 84.37%. Our experimental results show that Glacier can accurately identify resolvable malicious domain names without any DNS traffic data or prior knowledge about unknown domain names.
Yongzheng Zhang 0002, Yipeng Wang 0001
IJCNN2
2020 Efficient Malware Originated Traffic Classification by Using Generative Adversarial Networks
abstract
With the booming of malware-based cyber-security incidents and the sophistication of attacks, previous detections based on malware sample analysis appear powerless due to time-consuming and labor-intensive analysis process. The existing detection methods based on traffic analysis rely heavily on the available traffic patterns, which hinder detecting the zero-day attacks caused by malware variants. In this paper, we propose an approach based on deep learning referred to as TrafficGAN, which analyzes (HTTP) traffic sessions to distinguish between malware-related and normal traffic. We first try to explore traffic patterns of malware variants by adding noise and category condition to the Generative Adversarial Networks (GAN), thus generating various similar but slightly different traffic. And then, we use discriminative model to seek the deviation between abnormal traffic and normal traffic by extracting the essential difference. Notablely, we increase the diversity of data by generating samples adversarially, which enhances the robustness of the system to detect zero-day attacks and highlights the lack of sensitive data in the security community. We conduct extensive experiments on the public dataset and our data collected for specific targets. The results demonstrate that our method achieves superior performance to other methods and protects specific targets from the susceptibility of malware.
Yongzheng Zhang 0002, Xiao-chun Yun, Zhenyu Cheng 0001
ISCC3
2020 Exploit Internal Structural Information for IoT Malware Detection Based on Hierarchical Transformer Model
abstract
The number of IoT devices continues to increase, but the security of IoT devices cannot be guaranteed. Many IoT devices are infected with malware, forming huge botnets, which could launch DDoS attacks and cause heavy losses. In recent years, the IoT malware family has a tendency to be centralized on ARM-based IoT devices. The most widely spread families are the Mirai family and Gafgyt family. In this paper, we automatically extract the instruction sequences of these two families' samples and use the instruction sequences as language to describe these samples. We transfer instruction sequences to word vector space by Word2Vec. Then exploiting internal hierarchical structure of functions in malware to construct a hierarchical language model based on transformer-encoder to classify the samples. And the results obtained after visualizing the weights of the model can reflect the correlation of the functions in the sample, which can help the sample analyst find the key function. We use IoT software samples including Mirai samples, Gafgyt samples and benign samples to train our model. In the experiments, our model achieves 99.12% recall rate of malware and 94.67% family classification accuracy rate, which is better than other methods.
Kejia Xu, Yongzheng Zhang 0002
TrustCom4
2020 FTPB: A Three-stage DNS Tunnel Detection Method Based on Character Feature Extraction
abstract
The domain name system(DNS) protocol is one of the most versatile protocols in the world. If a hacker can control the DNS protocol to pass messages and control the host of victim, then no firewall can effectively intercept the DNS protocol. In recent years, the DNS tunnel is one of the most dangerous threats in the field of steganography, which supports a wide range of criminal activities. In order to detect and distinguish different types of DNS tunnels, we present a three-stage DNS tunnel detection method based on character feature extraction. This novel method named FTPB which uses feature extraction to filter out the domain names of the DNS tunnels by feature extraction classifier and then converts them into a high-dimensional vector by term frequency-inverse document frequency(TF-IDF), and reduces the dimension to 2 by principal components analysis(PCA). Ultimately, the data is classified by a binary vector classifier. Our method only needs to extract the character information of the domain names, and can effectively discover different types of DNS tunnels. Compared with traditional detection methods that can only detect DNS tunnels based on content, our research method can also have the ability to detect DNS tunnels based on codebooks, and the outcome of the evaluation substantially prove the efficacy of our method with accuracy and precision are more than 99%.
Kemeng Wu, Yongzheng Zhang 0002
TrustCom2
2020 HSTF-Model: An HTTP-based Trojan detection model via the Hierarchical Spatio-temporal Features of Traffics
Jiang Xie 0004, Xiao-chun Yun, Yongzheng Zhang 0002
Comput. Secur.4
2020 Khaos: An Adversarial Neural Network DGA With High Anti-Detection Ability
abstract
A botnet is a network of remote-controlled devices that are infected with malware controlled by botmasters in order to launch cyber attacks. To evade detection, the botmaster frequently changes the domain name of his Command and Control (C&C) server. Notice that most of these types of domain names are generated by domain generation algorithms (DGAs). In this paper, we propose Khaos, a novel DGA with high anti-detection ability based on neural language models and the Wasserstein Generative Adversarial Network (WGAN). The key insight of our research is that real domain names are composed of readable syllables and acronyms, and thus we can arrange syllables and acronyms using neural language models to mimic real domain names. In Khaos, we first find the most common n-grams in real domain names, then tokenize these domain names into n-grams, and finally synthesize new domain names after learning arrangements of n-grams from real domain names. We carry out experiments using a variety of state-of-the-art DGA detection approaches: the statistics-based, the distribution-based, the LSTM-based and the graph-based detection approach. Our experimental results show that the average distance for detecting Khaos under the distribution-based detection approach is 0.64, the AUCs of Khaos under the statistics-based and the LSTM-based detection approach are 0.76 and 0.57, respectively, and the precision of Khaos under the graph-based detection approach is 0.68. Our work proves that the existing detection approaches have big troubles in detecting Khaos, and Khaos has better anti-detection ability than state-of-the-art DGAs. In addition, we find that training the existing detection approach on a dataset including the domain names generated by Khaos can improve its detection ability.
Xiao-chun Yun, Yipeng Wang 0001, Tianning Zang, Yuan Zhou 0008, Yongzheng Zhang 0002
IEEE Trans. Inf. Forensics Secur.6
2019 GeoCET: Accurate IP Geolocation via Constraint-Based Elliptical Trajectories
Xiuguo Bao, Yongzheng Zhang 0002, Huanhuan Yang
CollaborateCom3
2019 A Smart Topology Construction Method for Anti-tracking Network Based on the Neural Network
Changbo Tian, Yongzheng Zhang 0002, Yupeng Tuo, Ruihai Ge
CollaborateCom2
2019 Achieving Dynamic Communication Path for Anti-Tracking Network
abstract
The increasingly rampant network monitoring and tracing bring the huge challenge on the protection of netizens' privacy. The anonymous networks mitigate the threat of network monitoring and tracing to a certain degree, but the static communication path has become the weakness. To address the problem, we propose a tracking-resistant communication mechanism with dynamic paths(TresMep). Different with the stepping stone chain like Tor, TresMep provides a chain of node groups which include at least one honest node. The message is transferred between groups. Each group uses asynchronous DC-net to hide the exit node which deliver the message to the honest node of the next group, and each group randomly chooses the exit node in each round of transmission through lagrange interpolation. Then, the transmission path would be changed dynamically and randomly to provide stronger tracking-resistance. The experimental results show that TresMep has a stronger performance of trackingresistance than the stepping stones based anti-tracking network with static communication path. The communication efficiency of TresMep is also satisfactory. But when message load is big, the communication efficiency of TresMep becomes worse. TresMep takes a tradeoff between tracking- resistance and communication efficiency.
Changbo Tian, Yongzheng Zhang 0002, Yupeng Tuo, Ruihai Ge
GLOBECOM2
2019 A Comprehensive Measurement Study of Domain-Squatting Abuse
abstract
Domain-squatting abuse refers to the premeditated attempt by an attacker to register perceptively confusing domain names thereby tricking visitors into querying them. There are totally five squatting types have been investigated so far, namely typo-squatting, bit-squatting, homograph-squatting, sound-squatting, and combo-squatting. Existing researches only focus on one specific squatting type and never explore the relationship among them. In this paper, we perform the first comprehensive measurement study of domain-squatting abuse. We select 786 the most queried domains, and hunt for squatting abuses against them in ISP-level DNS traffic. We find that although typo-squatting accounts for most of squatting domains, combo-squatting are able to attract more traffic. Our further case studies show that parking ads is still the most important way for attackers to make profits. The only exception is combo-squatting, in which squatters tend to leverage the reputation of squatted domains to develop their own business. It is worth noting that some squatting domains are even used to deliver malware. Moreover, the Alexa ranks of certain squatting domains have already surpassed the original domains. These results clearly call for the need to better protect the intellectual property of domain names.
Yuwei Zeng, Tianning Zang, Yongzheng Zhang 0002, Xunxun Chen, Yipeng Wang 0001
ICC3
2019 Rethinking Encrypted Traffic Classification: A Multi-Attribute Associated Fingerprint Approach
abstract
With the unprecedented prevalence of mobile network applications, cryptographic protocols, such as the Secure Socket Layer/Transport Layer Security (SSL/TLS), are widely used in mobile network applications for communication security. The proven methods for encrypted video stream classification or encrypted protocol detection are unsuitable for the SSL/TLS traffic. Consequently, application-level traffic classification based networking and security services are facing severe challenges in effectiveness. Existing encrypted traffic classification methods exhibit unsatisfying accuracy for applications with similar state characteristics. In this paper, we propose a multiple-attribute-based encrypted traffic classification system named Multi-Attribute Associated Fingerprints (MAAF). We develop MAAF based on the two key insights that the DNS traces generated during the application runtime contain classification guidance information and that the handshake certificates in the encrypted flows can provide classification clues. Apart from the exploitation of key insights, MAAF employs the context of the encrypted traffic to overcome the attribute-lacking problem during the classification. Our experimental results demonstrate that MAAF achieves 98.69% accuracy on the real-world traceset that consists of 16 applications, supports the early prediction, and is robust to the scale of the training traceset. Besides, MAAF is superior to the state-of-the-art methods in terms of both accuracy and robustness.
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Yipeng Wang 0001
ICNP3
2019 A Linguistics-based Stacking Approach to Disposable Domains Detection
abstract
More Internet services tend to collect the one-time information from clients via DNS queries. Notably, the uncertainty of such transient information makes these domain names be queried only once in their lifetime. This type of domain is called disposable domain. Although they are not malicious, the efficiency of DNS infrastructures will still be affected by their ever-increasing number. In this paper, we propose Vogers, a linguistics-based stacking model, to detect the disposable domains. Our evaluation demonstrates that Vogers decreases the false positive rate by more than 19%, compared with the prior art, while maintaining the true positive rate above 98.9%.
Yuwei Zeng, Yongzheng Zhang 0002, Tianning Zang, Xunxun Chen, Yipeng Wang 0001
ICNP2
2019 CLR: A Classification of DNS Tunnel Based on Logistic Regression
abstract
In order to detect DNS tunnel conveniently and effectively, and to classify different DNS tunneling tools, in this paper we conceive a method that does not require manual extraction of features, but converts packet bytes into high-latitude vectors by ASCII and classifies mainstream DNS tunneling tools through a wide range of machine learning algorithms. This method is named CLR which is The experimental results show that our method not only has excellent performance of classification of DNS tunnel, but also is superior to the proposed approach in terms of accuracy, recall rate, training time and testing time, and these values are 0.9999, 0.9999, 27.3 sec, 1 sec respectively.
Kemeng Wu, Yongzheng Zhang 0002
IPCCC2
2019 A Method Based on Hierarchical Spatiotemporal Features for Trojan Traffic Detection
abstract
Trojans are one of the most threatening network attacks currently. HTTP-based Trojan, in particular, accounts for a considerable proportion of them. Moreover, as the network environment becomes more complex, HTTP-based Trojan is more concealed than others. At present, many intrusion detection systems (IDSs) are increasingly difficult to effectively detect such Trojan traffic due to the inherent shortcomings of the methods used and the backwardness of training data. Classical anomaly detection and traditional machine learning-based (TML-based) anomaly detection are highly dependent on expert knowledge to extract features artificially, which is difficult to implement in HTTP-based Trojan traffic detection. Deep learning-based (DL-based) anomaly detection has been locally applied to IDSs, but it cannot be transplanted to HTTP-based Trojan traffic detection directly. To solve this problem, in this paper, we propose a neural network detection model (HSTF-Model) based on hierarchical spatiotemporal features of traffic. Meanwhile, we combine deep learning algorithms with expert knowledge through feature encoders and statistical characteristics to improve the self-learning ability of the model. Experiments indicate that F1of HSTF-Model can reach 99.4% in real traffic. In addition, we present a dataset BTHT consisting of HTTP-based benign and Trojan traffic to facilitate related research in the field.
Jiang Xie 0004, Yongzheng Zhang 0002, Xiao-chun Yun
IPCCC3
2019 A Loss-Tolerant Mechanism of Message Segmentation and Reconstruction in Multi-path Communication of Anti-tracking Network
Changbo Tian, Yongzheng Zhang 0002, Yupeng Tuo, Ruihai Ge
SecureComm (1)2
2019 A Method of HTTP Malicious Traffic Detection on Mobile Networks
abstract
Aiming at solving the problem of HTTP malicious traffic detection on mobile networks, we propose a method of HMTD(HTTP Malicious Traffic Detection) based on the spatiotemporal sequence characteristics of traffic data. The traditional malicious traffic detection methods are relatively simple and mainly biased towards misuse detection or abnormal detection and probably suffer from a high false positive rate or false negative rate, so they are difficult to adapt to the current rapid development of the Internet. HMTD uses neural networks for malicious traffic identification, and extracts features from malicious and normal HTTP traffic, which can produce excellent detection results. HMTD utilizes CNN to extract the packet spatial characteristics in the traffic, and utilizes LSTM to extract the temporal characteristics between the packets in the traffic. The experimental results demonstrate that the proposed method can achieve an accuracy of more than 99.4% in the actual network environment and has excellent performance in terms of Precision and Recall.
Xiao-chun Yun, Mao Tian, Jiang Xie 0004, Yongzheng Zhang 0002, Yu Zhou 0028
WCNC6
2018 Important Member Discovery of Attribution Trace Based on Relevant Circle (Short Paper)
Jian Xu 0010, Xiao-chun Yun, Yongzheng Zhang 0002, Zhenyu Cheng 0001
CollaborateCom3
2018 MUI-defender: CNN-Driven, Network Flow-Based Information Theft Detection for Mobile Users
Zhenyu Cheng 0001, Xunxun Chen, Yongzheng Zhang 0002, Jian Xu 0010
CollaborateCom3
2018 GeoBLR: Dynamic IP Geolocation Method Based on Bayesian Linear Regression
Xiuguo Bao, Yongzheng Zhang 0002
CollaborateCom3
2018 MalShoot: Shooting Malicious Domains Through Graph Embedding on Passive DNS Data
Chengwei Peng, Xiao-chun Yun, Yongzheng Zhang 0002
CollaborateCom3
2018 A Stacking Approach to Objectionable-Related Domain Names Identification by Passive DNS Traffic (Short Paper)
Yongzheng Zhang 0002, Tianning Zang, Zhizhou Liang, Yipeng Wang 0001
CollaborateCom2
2018 MalHunter: Performing a Timely Detection on Malicious Domains via a Single DNS Query
Chengwei Peng, Xiao-chun Yun, Yongzheng Zhang 0002
ICICS3
2018 Community Discovery of Attribution Trace Based on Deep Learning Approach
Jian Xu 0010, Xiao-chun Yun, Yongzheng Zhang 0002, Zhenyu Cheng 0001
ICICS3
2018 Online Discovery of Congregate Groups on Sparse Spatio-temporal Data
abstract
The pervasiveness of location-acquisition technologies leads to large amounts of spatio-temporal data, which brings us opportunities and challenges to discover interesting group patterns from these individual's trajectories. In this work, firstly, we propose a novel group pattern called congregate group, which captures various congregations by exploiting trajectory streams. Then, we design a discovery framework which contains three main stages including trajectory preprocessing, crowds generation and congregate groups discovery to detect congregations. Meanwhile, an interpolation method is proposed to handle missing points on sparse data. Besides, a set of optimization techniques is applied to reduce computational costs. Finally, our extensive experiments based on real cellular network dataset and real taxicab trajectory dataset demonstrate the effectiveness, efficiency and scalability of our proposed approach.
Tianran Chen, Yongzheng Zhang 0002, Yupeng Tuo
PIMRC2
2018 Information Propagation Prediction Based on Key Users Authentication in Microblogging
abstract
In microblogging, key users are a significant factor for information propagation. Key users can affect information propagation size while retweeting the information. In this paper, to predict information propagation, we propose a novel linear model based on key users authentication. This model mines key users to dynamically improve the linear model while predicting information propagation. So our model can not only predict information propagation but also mine key users. Experimental results show that our model can achieve remarkable efficiency on predicting information propagation problem in real microblogging networks. At the same time, our model can find the key users who affect information propagation.
Miao Yu 0006, Yongzheng Zhang 0002, Tianning Zang, Yipeng Wang 0001
Secur. Commun. Networks2
2017 ProNet: Toward Payload-Driven Protocol Fingerprinting via Convolutions and Embeddings
Yafei Sang, Yongzheng Zhang 0002, Chengwei Peng
CollaborateCom2
2017 Fingerprinting Protocol at Bit-Level Granularity: A Graph-Based Approach Using Cell Embedding
abstract
Traffic identification is defined as the act of ascertaining which application or service or protocol is contributing to the network traffic by using a fingerprint, which is a distinguishable unique pattern representing a particular applications traffic. The continual appearance of new applications and their frequent updates emphasize the need for automatic protocol fingerprints generation. In this paper, we propose BitGrapher, a novel graph-based approach that accurately infers protocol fingerprints at bit-level granularity for accurate traffic identification. The proposal is designed to accommodate to various protocol traces including text-based and binary-based protocols, and even possible unknown proprietary communication protocols. Our proposed approach introduces a new concept of cell that allows BitGrapher to encode protocol payloads as a graphical model, and then converts fingerprinting protocol problem as a series of graph operations (e.g., graph construction, pruning, partition). The key insight of the graphical model use cell embeddings that captures the distinguishable positions with their values and distinguishable correlation among them. We implement and evaluate BitGrapher on real-world traces, including DNS, QQLive, SopCast, SMB, HTTP, and SMTP, and our experimental results show that BitGrapher can accurately identify the protocol trace with an average precision of about 97.28% and an average recall of about 99.12%. We also compare the results of BitGrapher to two state-of-the-art approaches ProWord and ProDigger, which shows that BitGrapher provides significant improvements in precision and recall for protocol identification task.
Yafei Sang, Yongzheng Zhang 0002
ICPADS2
2017 Detecting Information Theft Based on Mobile Network Flows for Android Users
abstract
With the widespread use of smartphones, more and more malicious attacks happen with information leakage from apps installed on users' devices. The adversary always uses a malware as the client to take remote control of smartphones, and leverages the vulnerability of operation systems to send back the collected information without users' permissions. All the information has to be transferred by network traffic. In this paper, we consider that different apps maybe generate different network flows by different operations, and the "shapes" of the benign flows and malicious ones will be diverse. Thus we propose a detection model based on the analysis of relationships between behavior patterns and network flows, which achieves our goal by using the Random Forest machine learning algorithm to classify the network flows into benign or malicious. To further improve the controllability of the experiment, we design an app called Moledroid to simulate malwares by uploading the user's privacy without authorization, in addition, we can change the behavior pattern of the app to complete our evaluation. Finally, we run this app and several benign apps to generate traffic to detect the malicious network flows, and it shows that our detection model can achieve precision and accuracy higher than 95%, which demonstrates that our model is suitable for detecting the network flows of information theft.
Zhenyu Cheng 0001, Xunxun Chen, Yongzheng Zhang 0002, Yafei Sang
NAS3
2017 Towards Robust and Accurate Similar Trajectory Discovery: Weak-Parametric Approaches
abstract
Trajectory analysis is crucial and has been more and more widely used in various fields, such as location-based services (LBS), urban traffic control, user classification and route planner, etc. In this paper, we propose GSIM and ASIM, two novel approaches that are weak-parametric and can effectively measure and discover similar trajectories. The proposed methods are based on the key insight that the similarity can be reflected by observing the growth rate of specific indicators. (1) GSIM defines a 3-layer grid structure and statistics the total overlapping points for all grids between trajectories in each layer, it finally calculates the growth rate of the total counts as the grid radius grows from layer 1 to layer 3. (2) ASIM assumes that any two trajectories are similar and calculates the area of the minimum boundary rectangle that contains all the points. Then it cuts the rectangle from four directions one point by one to get the maximum boundary rectangle that contains the other two percentage of total points. Finally it utilizes the average change rate of the areas as the similarity. Further, we design parameter-learning modules to learn the setting of corresponding parameters automatically. Extensive experiments on real-world dataset show that, compared with typical approaches like LCSS, EDIT, DTW, etc., the proposed methods can significantly improve the effectiveness and achieve better efficiency in most test cases. Meanwhile, they are not sensitive to parameter settings.
Yupeng Tuo, Xiao-chun Yun, Yongzheng Zhang 0002
NAS3
2017 NSIM: A robust method to discover similar trajectories on cellular network location data
abstract
Trajectory analysis is crucial and has been more and more widely used in various fields, such as location-based services, urban traffic control, route plan, etc. The existing methods have certain limitations when applied to cellular network location data. In this paper, we propose NSIM, a novel approach that can effectively discover similar trajectories. In NSIM, we first design an algorithm that can discover all the common moving patterns among trajectories, and then we adopt a vectorization method to abstract each trajectory as a summary vector that composed of specific common moving patterns. Finally we measure the similarity of trajectories by computing the distance between the summary vectors. Extensive experiments on real-world dataset show that, compared with three other approaches, NSIM achieves good effectiveness in most test cases and achieves better efficiency when applied to small or medium length trajectories.
Yupeng Tuo, Xiao-chun Yun, Yongzheng Zhang 0002
PIMRC3
2017 MSTM: A novel map matching approach for low-sampling-rate trajectories
abstract
Map matching is an important technique that matches user trajectories to the real road networks on a digital map. It is crucial and has been more and more widely used in various fields, such as route plan, traffic forecast, location-based services and so on. However, most existing algorithms are less effective when applied to low-sampling-rate trajectories. In this paper, we propose MSTM, a novel approach that can effectively match the low-sampling-rate trajectory to road networks. In MSTM, we first partition the trajectories into trajectory segments according to the stay points. Then we construct a map-searching tree by conditional extend and prune operations, which contains all the candidate paths. Finally, by considering the spatial and temporal information of trajectories, we evaluate each branch path in the map-searching tree and choose the one with the highest score as the result. Extensive experiments on real-world datasets show that, compared with two classic approaches, MSTM outperforms ST-Matching and IVMM in terms of matching accuracy as well as efficiency.
Yupeng Tuo, Xiao-chun Yun, Yongzheng Zhang 0002
PIMRC3
2017 Rethinking robust and accurate application protocol identification
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Tianning Zang
Comput. Networks3
2017 A nonparametric approach to the automated protocol fingerprint inference
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Guangjun Wu
J. Netw. Comput. Appl.3
2016 DMNS: A Framework to Dynamically Monitor Simulated Network
abstract
With rapid development of network simulation technology, monitoring system has become an essential tool for the researching and testing of network space activities. However, the current monitoring technologies of simulated networks cannot satisfy the requirements in terms of flexibility and efficiency. This paper proposes a framework called DMNS to dynamically monitor simulated network. With DMNS, users are able to customize the monitored objects and monitoring actions to meet the requirements of flexibility. In addition, the administrators could dynamically change the monitoring rules for saving resources based on callback mechanism. DMNS also considers the requirements of large-scale distributed simulation. Specifically, it leverages message oriented middleware to achieve efficient monitoring message information transmission to guarantee the robustness of the monitoring system. We implement a prototype of DMNS in a network range system and demonstrate the effectiveness by a case study.
Zhiyu Hao, Yongzheng Zhang 0002, Yaqiong Peng, Zhenxi Sun
ICPADS3
2016 Quantitative threat situation assessment based on alert verification
abstract
Abstract Traditional network threat situational assessment is based on raw alerts, not combined with contextual information, which influences the accuracy of assessment. In this paper, we propose a method to quantitatively assess network threat situation based on not only alerts but also contextual information. It firstly verifies alerts by matching alerts with contextual information to determine the successful probability of attacks, then analyzes the impact caused by attacks according to the severity and the corresponding asset value of them, and finally quantitatively assesses network threat situation based on the successful probability and the impact of attacks. Case studies show that the method can assess network threat situations more reasonably. Copyright © 2016 John Wiley & Sons, Ltd.
Rongrong Xi, Xiao-chun Yun, Zhiyu Hao, Yongzheng Zhang 0002
Secur. Commun. Networks4
2016 A Semantics-Aware Approach to the Automated Network Protocol Identification
abstract
Traffic classification, a mapping of traffic to network applications, is important for a variety of networking and security issues, such as network measurement, network monitoring, as well as the detection of malware activities. In this paper, we propose Securitas, a network trace-based protocol identification system, which exploits the semantic information in protocol message formats. Securitas requires no prior knowledge of protocol specifications. Deeming a protocol as a language between two processes, our approach is based upon the new insight that the n-grams of protocol traces, just like those of natural languages, exhibit highly skewed frequency-rank distribution that can be leveraged in the context of protocol identification. In Securitas, we first extract the statistical protocol message formats by clustering n-grams with the same semantics, and then use the corresponding statistical formats to classify raw network traces. Our tool involves the following key features: 1) applicable to both connection oriented protocols and connection less protocols; 2) suitable for both text and binary protocols; 3) no need to assemble IP packets into TCP or UDP flows; and 4) effective for both long-live flows and short-live flows. We implement Securitas and conduct extensive evaluations on real-world network traces containing both textual and binary protocols. Our experimental results on BitTorrent, CIFS/SMB, DNS, FTP, PPLIVE, SIP, and SMTP traces show that Securitas has the ability to accurately identify the network traces of the target application protocol with an average recall of about 97.4% and an average precision of about 98.4%. Our experimental results prove Securitas is a robust system, and meanwhile displaying a competitive performance in practice.
Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002, Yu Zhou 0015
IEEE/ACM Trans. Netw.3
2015 Traffic Replay in Virtual Network Based on IP-Mapping
Zhiyu Hao, Yongzheng Zhang 0002, Zhenquan Ding, Haiqiang Fei
ICA3PP (4)3
2015 Rethinking Robust and Accurate Application Protocol Identification: A Nonparametric Approach
abstract
Protocol traffic analysis is important for a variety of networking and security infrastructures, such as intrusion detection and prevention systems, network management systems, and protocol specification parsers. In this paper, we propose ProHacker, a nonparametric approach that extracts robust and accurate protocol keywords from network traces and effectively identifies the protocol trace from mixed Internet traffic. ProHacker is based on the key insight that the n-grams of protocol traces have highly predictable statistical nature that can be effectively captured by statistical language models and leveraged for robust and accurate protocol identification. In ProHacker, we first extract protocol keywords using a nonparametric Bayesian statistical model, and then use the corresponding protocol keywords to classify protocol traces by a semi-supervised learning algorithm. We implement and evaluate ProHacker on real-world traces, including SMTP, FTP, PPLive, SopCast, and PPStream, and our experimental results show that ProHacker can accurately identify the protocol trace with an average precision of about 99.42% and an average recall of about 98.64%. We also compare the results of ProHacker to two state-of-the-art approaches ProWord and Securitas using backbone traffic. We show that ProHacker provides significant improvements on precision and recall for online protocol identification.
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002
ICNP3
2015 A Markov Random Field Approach to Automated Protocol Signature Inference
Yongzheng Zhang 0002, Yipeng Wang 0001, Jianliang Sun, Xiaoyu Zhang 0002
SecureComm1
2015 Unsupervised adaptive sign language recognition based on hypothesis comparison guided cross validation and linguistic prior filtering
Yu Zhou 0015, Xiaokang Yang 0001, Yongzheng Zhang 0002, Yipeng Wang 0001, Xiujuan Chai, Weiyao Lin
Neurocomputing3
2015 SMS Worm Propagation Over Contact Social Networks: Modeling and Validation
abstract
Nowadays, short message service (SMS) worms have been discovered to propagate themselves via victims' contact lists by sending malicious text messages. Correspondingly, defenders need to analyze and model the dynamics of these worms to lessen their potential threat. However, the existing worm propagation models, which almost generate the similar curves of an exponential smooth rise, cannot well explain the infection dynamics of real-world SMS worms, which exhibits an uneven wave-like uplift. Motivated by this observation, we formalize the general infection process of SMS worms in contact social networks, and propose a novel analytical model based on stochastic processes. In contrast to previous models, our model not only considers the different and asymmetrical relationships between mobile users by modeling the node reputation and the edge trust degree, but also describes the user behavior of checking messages by introducing two susceptible states. Moreover, the strong assumptions in previous works are eliminated by determining related components based on extensive statistical investigations. Afterward, both real-world SMS worms and artificial ones are utilized in validation and comparison experiments, and the results show that our model is more suitable for describing the propagation of these sophisticated worms, compared with the state-of-the-art models. In addition, we study on the impacts of key factors and give some interesting discoveries.
Xiao-chun Yun, Yongzheng Zhang 0002
IEEE Trans. Inf. Forensics Secur.3
2014 DR-SNBot: A Social Network-Based Botnet with Strong Destroy-Resistance
abstract
Social network-based botnets have become an important research direction of botnets. To avoid the single-point failure of existing centralized botnets, we propose a Social Network-based Botnet with strong Destroy-Resistance (DR-SNBot). By enhancing the security of the Command and Control (C&C) channel and introducing a divide-and-conquer and automatic reconstruction mechanism, we greatly improve the destroy-resistance of DR-SNBot. Moreover, we design the pseudo code for nickname generation algorithm, botmaster and bot respectively. Then, we construct the DR-SNBot via sin a blog and make simulated experiments to evaluate it. Furthermore, we make comparisons of controllability between botnets Mrrbot and DR-SNBot. The experimental results indicate that DRSNBot is more resilient. It is not only available in real-world environment, but also resistant enough to varying degrees of C&C-server removals in simulated environment.
Yongzheng Zhang 0002
NAS2
2014 A Segmentation Pattern Based Approach to Automated Protocol Identification
abstract
In-depth understanding of network traffic is important for a variety of applications, such as network management and network security. In this paper, we propose a novel protocol identification system PSKS, which relies on the statistical signatures of network packet payloads. The proposed approach is based on the key insight that message segmentation patterns can be leveraged for accurate application identification. Specifically, the segmentation possibility for every position of protocol messages exhibits highly skewed frequency distribution due to the reason that different protocols have different message formats (i.e., Distinct message segmentation patterns). Motivated by this observation, we want to extract statistical application fingerprints by exploiting the message segmentation patterns. In PSKS, we first extract the message segmentation patterns by scoring the segmentation possibility scale for each position of messages, and then extract statistical signatures by Kolmogorov-Smirnov test and feed the signatures to tri-training, a collaborative learning algorithm. The tri-training can improve the generalization ability of our final classifier. We implemented and evaluated PSKS, and the experimental results show that PSKS achieves an average precision and recall of approximately 98%.
Yafei Sang, Yongzheng Zhang 0002, Yipeng Wang 0001, Yu Zhou 0015
PDCAT2
2014 Detecting Malicious Behaviors in Repackaged Android Apps with Loosely-Coupled Payloads Filtering Scheme
Yongzheng Zhang 0002, Tianning Zang
SecureComm (1)2
2014 Visual Similarity Based Anti-phishing with the Combination of Local and Global Features
abstract
Phishing uses a fake Web page to steal personal sensitive information such as credit card numbers and passwords. Generally, the fake Web page is visually similar to the legitimate target Web page. The phishers can obtain financial benefits through these information. Anti-phishing is very important for a variety of applications such as phishing attacks, online transaction security, and user privacy protection. In this paper, we propose a novel and effective visual similarity based phishing detection approach that compares the snapshot image pair of the suspected Web page and the protected Web page. The proposed approach is based on the key insight that both the local and the global features of the Web page image can be used to represent the visual characteristics of the Web page together. This approach is purely on the image level, and thus can effectively deal with the non-text phishing tricks including images or Flashes objects in the HTML contents. For the local feature, the existence of the target logo is detected. For the global feature, the similarity of the visible part of the Web page is considered. We implemented and evaluated the proposed approach on a large scale dataset consisting of 2,129 real world phishing Web pages and 1,367 irrelevant legitimate Web pages. The experimental results show that the proposed approach can achieve over 90.00% true positive rate and 97.00% true negative rate. Our approach has been applied in the anti-phishing project of a major Internet Service Provider and gives a periodical reports to the potential users.
Yu Zhou 0015, Yongzheng Zhang 0002, Yipeng Wang 0001, Weiyao Lin
TrustCom2
2013 Counting sort for the live migration of virtual machines
abstract
The live migration of virtual machines is an important technique in the area of virtualization, and it has been used for load balancing, fault tolerance, and system maintenance in modern data centers, clusters, and cloud computing. The pre-copy algorithm is the most used method for the live migration of virtual machines. However, the existing problem of repeatedly transferring dirty memory pages leads the increase of the transferring data amount, delays of the total migration time as well as the downtime. By analyzing the iteration process of the pre-copy algorithm, we find that the transferring order of memory pages during every middle round has a huge impact on the generation and transferring of dirty memory pages. Further we put forward the concept of the live migration of virtual machines based on the counting sort. During every middle round of the iteration process, we do not transfer the memory pages according to their original order, instead we transfer the memory pages according to their times of being dirty. Experiment results show that with different workloads the counting sort method could simultaneously decrease the transferring data amount, the total migration time, and the downtime to improve the performance of the live migration.
Qingxin Zou, Zhiyu Hao, Xiao-chun Yun, Yongzheng Zhang 0002
CLUSTER5
2013 CADM: A Centralized Administration and Dynamic Monitoring Framework for Network Intrusion Detection Based on Virtualization
abstract
Virtualization technology, which has the characteristic of producing dynamic change, enables the virtual network structure to no longer depend strictly on the underlying hardware environment. With virtualization platform administrators tasked with preventing attacks in order to provide uninterrupted service, existing intrusion detection technologies are continuously challenged. Consequently, this paper proposes a Centralized Administration and Dynamic Monitoring framework (CADM) based on virtualization for network intrusion detection. CADM is able to centrally administrate, and monitor network behavior in the virtual computing environment by automatically deploying and updating intrusion detection processes and rules. In the aspect of monitoring capability, CADM allows the monitoring locations in intrusion detection to be automatically adjusted in real time, thus adapting to the dynamic changes (such as migration) of virtual machines (VMs). Moreover, the monitoring processes involved in intrusion detection could also be automatically updated by dynamically updating security strategies. In the aspect of monitoring granularity, CADM is able to monitor network interfaces of each virtual machine (VM) for fine-grained network intrusion detection and network traffic acquisition. Our experimental results demonstrate that more convenient and efficient monitoring and administrating capabilities are available with CADM for virtualization platform administrators.
Zhenquan Ding, Zhiyu Hao, Yongzheng Zhang 0002
PDCAT3
2012 A semantics aware approach to automated reverse engineering unknown protocols
abstract
Extracting the protocol message format specifications of unknown applications from network traces is important for a variety of applications such as application protocol parsing, vulnerability discovery, and system integration. In this paper, we propose ProDecoder, a network trace based protocol message format inference system that exploits the semantics of protocol messages without the executable code of application protocols. ProDecoder is based on the key insight that the n-grams of protocol traces exhibit highly skewed frequency distribution that can be leveraged for accurate protocol message format inference. In ProDecoder, we first discover the latent relationship among n-grams by first grouping protocol messages with the same semantics and then inferring message formats by keyword based clustering and cluster sequence alignment. We implemented and evaluated ProDecoder to infer message format specifications of SMB (a binary protocol) and SMTP (a textual protocol). Our experimental results show that ProDecoder accurately parses and infers SMB protocol with 100% precision and recall. For SMTP, ProDecoder achieves approximately 95% precision and recall.
Yipeng Wang 0001, Xiao-chun Yun, Zubair Shafiq, Alex X. Liu, Danfeng Yao, Yongzheng Zhang 0002, Li Guo 0001
ICNP8
2012 A General Framework of Trojan Communication Detection Based on Network Traces
abstract
Because of the widespread Trojan, Internet users become more and more vulnerable to the threat of information leakage. Traditional techniques of Trojan detection were classified into two main categories: host-based and network-based. Unfortunately, existing techniques are insufficient and limited, because of the following reasons: (1)only uncover the known Trojan while inefficiently detecting novel samples, (2) should be adjusted in a timely fashion even a trivial change is applied, and (3)become computationally more expensive. In our work, we focus on a network behavior based method to address the limitations of previous network-based approaches. We analyze the profile of network behavior at two levels: (i)flow-level, (ii)IP-level. Our approach present two main advantages: (1)capture more detailed information to describe the network behavior profile, (2)consume lower computational overhead. We proposed a system, Manto, which detects Trojan communication with high accuracy using clustering technique. We implement Manto on real-world traces. The evaluation results exhibit that Manto is suitable for detecting Trojan communication amongst the vast amount of network traffic, with over 91% accuracy and less than 3.2% false positive ratio. We confidently regard our approach as a complementary way to the existing network-based techniques for we could address their main shortcomings.
Shicong Li, Xiao-chun Yun, Yongzheng Zhang 0002, Yipeng Wang 0001
NAS3
2012 Modeling Social Engineering Botnet Dynamics across Multiple Social Networks
Xiao-chun Yun, Zhiyu Hao, Yongzheng Zhang 0002, Xiang Cui, Yipeng Wang 0001
SEC4
2011 Network Threat Assessment Based on Alert Verification
abstract
In face of overwhelming alerts produced by firewalls or intrusion detection devices, it is difficult to assess network threats that we face. In this paper, we propose a threat assessment approach to estimate the impact of attacks on network. The approach employs the Common Vulnerability Scoring System to quantitatively assess network threats and further correlates alerts with contextual information to improve the accuracy of assessment. In the case studies, we demonstrate how the approach is applied in real networks. The experimental results show that the approach can make an accurate assessment of network threats.
Rongrong Xi, Xiao-chun Yun, Shuyuan Jin, Yongzheng Zhang 0002
PDCAT4
2011 CNSSA: A Comprehensive Network Security Situation Awareness System
abstract
With tremendous attacks in the Internet, there is a high demand for network analysts to know about the situations of network security effectively. Traditional network security tools lack the capability of analyzing and assessing network security situations comprehensively. In this paper, we introduce a novel network situation awareness tool CNSSA (Comprehensive Network Security Situation Awareness) to perceive network security situations comprehensively. Based on the fusion of network information, CNSSA makes a quantitative assessment on the situations of network security. It visualizes the situations of network security in its multiple and various views, so that network analysts can know about the situations of network security easily and comprehensively. The case studies demonstrate how CNSSA can be deployed into a real network and how CNSSA can effectively comprehend the situation changes of network security in real time.
Rongrong Xi, Shuyuan Jin, Xiao-chun Yun, Yongzheng Zhang 0002
TrustCom4
2011 Parallelizing weighted frequency counting in high-speed network monitoring
Yu Zhang 0036, Binxing Fang, Yongzheng Zhang 0002
Comput. Commun.3
2010 Cooperative Work Systems for the Security of Digital Computing Infrastructure
abstract
On open digital computing infrastructure, various large-scale and complicated malicious behaviors are increasingly threatening the security of digital computing infrastructure. In this paper, a Cooperative Work Model (CRM) is presented by extending the conceptions of the Universal Turing Machine to deal with the threats. Then the Cooperative Work System Framework (CWSF) is derived from the model. Based on the framework, two practical Cooperative Work Systems (CWSs) are developed to track and analyze the Botnet and DDoS on digital computing infrastructure respectively. The systems collectively use and coordinate various monitoring systems distributed in the back-bone network of the infrastructure. The experimental results of analyzing typical security events show that the framework and systems are efficient and effective to collaboratively use diverse related network systems for monitoring and analyzing the large-scale network events. Currently, the systems are running steadily in the monitoring environment of a large-scale back-bone network.
Tianning Zang, Xiao-chun Yun, Tianyi Zang, Yongzheng Zhang 0002, Chaoguang Men
ICPADS4
2010 Identifying heavy hitters in high-speed network monitoring
Yu Zhang 0036, Binxing Fang, Yongzheng Zhang 0002
Sci. China Inf. Sci.3
2008 UBSF: A novel online URL-Based Spam Filter
abstract
Spam fighting is a classic puzzle in network security. In the past decades, many filtering spam solutions have been proposed. However, currently the conventional techniques still suffer from high false positives and false negatives, especially the former which usually cannot be accepted by the end users. This paper proposes a novel online URL-based spam filter (UBSF) on the basis of analyses over the conventional especially the URL-based anti-spam techniques. UBSF identifies spam by comparing the similarity of the extracted URLs (universal resource locator) from them with the URLs from our user-oriented standard URL Library (SUL). Experimental evaluations both from the contrast experiment and the prototype based on UBSF demonstrate it can significantly raise the filtering accuracy, effectively reduce false positives and can be applied to online process and heavy traffic environment by reducing the computational cost than the state-of-the-art techniques.
Yang Li 0002, Binxing Fang, Li Guo 0001, Zhihong Tian 0001, Yongzheng Zhang 0002, Zhi-Gang Wu
ISCC5
2008 A Survey of Alert Fusion Techniques for Security Incident
abstract
Security incident have been imposing tremendous threats on todaypsilas network information system. To protect this information system from the increasing threat of intrusion, various kinds of detection systems and sensors for security incident have been developed. The main disadvantages of current systems and sensors are a high false detection rate and the lack of post-incident decision support capability. To minimize these drawbacks, various alert fusion technologies have been proposed in the recent years. This paper presents a general summary of these technologies. Basic models and key technologies of alert fusion are analyzed and discussed. Moreover, important aggregation and correlation algorithms are discussed. Finally, we make concluding remarks by predicting the development tendencies of alert correlation technologies.
Tianning Zang, Xiao-chun Yun, Yongzheng Zhang 0002
WAIM3
2005 Computer Vulnerability Evaluation Using Fault Tree Analysis
Mingzeng Hu, Xiao-chun Yun, Yongzheng Zhang 0002
ISPEC4