VLDB 2026 Research / reviewers in the wild / expert
Yipeng Wang 0001
dblp:36/4627-1
· DBLP profile ↗
59ranked-venue papers
11as first author
26since 2021 · last 2026
0000-0002-8188-4666ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 6 since 2021Security and privacy · 12 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advances and challenges of multi-task learning method in recommender systems: A survey
Ruiping Yin, Zhen Yang 0004, Yipeng Wang 0001 |
Neurocomputing | 4 |
| 2026 | EO-EPTC: End-to-End Original Traffic-Based Encrypted Proxy Traffic Classification FrameworkabstractMachine learning-based methods for encrypted traffic classification can be effectively applied to analyze encrypted proxy traffic generated by proxy protocols, which are intermediary protocols used to route network traffic through a remote server. Nonetheless, different encrypted proxy protocols generate distinct traffic patterns, even when they handle the same network behavior. To address these distribution differences, a straight-forward approach is to collect datasets specific to each proxy protocol. However, typical proxy protocols repackage original traffic by encrypting it without payload padding or compression. This leads to a definite characteristic correlation between original and encrypted proxy traffic. We propose an End-to-end Original traffic-based Encrypted Proxy Traffic Classification framework (EO-EPTC) to bridge the distribution gap between original traffic and proxied traffic, enabling the classification of encrypted proxy traffic using a original traffic dataset. EO-EPTC conducts sequence feature alignment to reduce distribution bias and employs a Seq2Seq model to capture the underlying semantics of the proxy protocol, creating a sequence feature transformation model. We apply EO-EPTC to existing encrypted traffic classification models, training them on original traffic to classify proxied traffic. This achieves up to 99.70% accuracy on encrypted proxy traffic, comparable to models trained directly on proxied traffic. Huajie Jia, Zhenzhou Tang, Yipeng Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | MUFFIN: A Meta-Knowledge Decoupling-Based Approach to Few-Shot IoT Traffic ClassificationabstractTraditional machine learning-based approaches to Internet of Things (IoT) traffic classification are hindered by the need for a large number of labeled flows in real-world network environments. To cope with this problem, some few-shot traffic classification approaches have been proposed to achieve accurate classification by comparing the similarity of flows to the given labeled flows. These approaches train neural networks to concurrently learn feature extraction and feature comparison meta-knowledge from sufficient labeled flows. However, they conflate these two types of meta-knowledge and ignore their differences. We propose MUFFIN, a meta-knowledge decoupling-based approach to few-shot IoT traffic classification. MUFFIN is based on our key insight that feature extraction and feature comparison meta-knowledge are two different types of meta-knowledge with distinct learning goals. Therefore, MUFFIN decouples their learning in the meta-knowledge phase and introduces a masked autoencoder as its feature extractor to enhance the learning of feature extraction meta-knowledge. Extensive experiments conducted on three publicly available datasets demonstrate the classification ability and robustness of MUFFIN. Furthermore, a comparison with five state-of-the-art few-shot traffic classification approaches reveals that MUFFIN outperforms them in classification accuracy and requires only 20 labeled flows. Yipeng Wang 0001, Yingxu Lai, Shui Yu 0001, Jinqiao Shi |
IEEE Trans. Netw. | 2 |
| 2025 | Critical Node-aware Augmentation for Hypergraph Contrastive LearningabstractHypergraph contrastive learning enables effective representation learning for hypergraphs without requiring labels. However, existing methods typically rely on randomly deleting or replacing nodes during hypergraph augmentation, which may lead to the absence of critical nodes and further disrupt the higher-order structural relationships within augmented hypergraphs. To address this issue, we propose a Critical Node-aware hypergraph contrastive learning method, which is the first attempt to leverage hyperedge prediction to retain critical nodes and accordingly maintain the reliable higher-order structural relationships within augmented hypergraphs. Specifically, we first employ contrastive learning to align the augmented hypergraphs, and then generate hyperedge embeddings to characterize node representations and their structural correlations. During the hyperedge embedding encoding process, we introduce a hyperedge prediction discriminator to score these embeddings, which quantifies the nodes' contributions to identify the critical nodes and maintain the higher-order structural relationships within augmented hypergraphs. Compared with previous studies, our proposed method can effectively alleviate the erroneous deletion or replacement of critical nodes and steadily maintain the inherent structural relationships between original hypergraph and augmented hypergraphs, naturally guiding better hypergraph representations for downstream tasks. Extensive experiments on various tasks demonstrate that our method is significantly superior to state-of-the-art methods. Yuena Lin, Yipeng Wang 0001, Wenmao Liu, Mingliang Yu, Zhen Yang 0004, Gengyu Lyu |
IJCAI | 3 |
| 2025 | Federated Multi-View Multi-Label ClassificationabstractMulti-view multi-label classification is a crucial machine learning paradigm aimed at building robust multi-label predictors by integrating heterogeneous features from various sources while addressing multiple correlated labels. However, in real-world applications, concerns over data confidentiality and security often prevent data exchange or fusion across different sources, leading to the challenging issue of data islands. To tackle this problem, we propose a general federated multi-view multi-label classification method, FMVML, which integrates a novel multi-view multi-label classification technique into a federated learning framework. This approach enables cross-view feature fusion and multi-label semantic classification while preserving the data privacy of each independent source. Within this federated framework, we first extract view-specific information from each individual client to capture unique characteristics and then consolidate consensus information from different views on the global server to represent shared features. Unlike previous methods, our approach enhances cross-view fusion and semantic expression by jointly capturing both feature and semantic aspects of specificity and commonality. The final label predictions are generated by combining the view-specific predictions from individual clients and the consensus predictions from the global server. Extensive experiments across various applications demonstrate that FMVML fully leverages multi-view data in a privacy-preserving manner and consistently outperforms state-of-the-art methods. Hongdao Meng, Yongjian Deng, Qiyu Zhong, Yipeng Wang 0001, Zhen Yang 0004, Gengyu Lyu |
IEEE Trans. Big Data | 4 |
| 2025 | Proxied Traffic Fingerprinting for Hidden Service De-Anonymization With Burst ReshapingabstractTraffic fingerprinting attack is a promising approach for Tor hidden services (HS) de-anonymization. However, it is inherently difficult to acquire traffic of target HSs (HST) for fingerprinting model training, because the physical location of the services is hidden due to the design of Tor protocol. In order to solve this problem, some alternatives such as mirrored HST (MHST) and client-side HST (CHST) have been proposed for training fingerprinting model. These alternatives are easy to acquire and aim to closely match the characteristics of the target HST. However, they cannot perfectly replace the target HST for the aspects of consistency of both response and protocol. In this paper, we propose a proxied fingerprinting approach called PF. A Proxy HS is deployed to acquire proxied HS traffic (PHST) as an alternative to conduct traffic fingerprinting attack, which satisfies both response and protocol consistency and is easy to acquire. In order to mitigate the impact introduced by Proxy HS, PF also introduces Burst Reshaping which includes burst reconstruction and pseudo-label learning to enhance the similarities between PHST and target HST. Experiments show that, PHST is a superior alternative to target HST, fingerprinting model trained using PF achieved an accuracy of 92.2%, surpassing the models trained with MHST and CHST by 72% and 34%, respectively. Additionally, PF is an add-on approach capable of improving the HS deanonymization effectiveness of any fingerprinting model architecture. The source code and dataset are available at https://github.com/Lzreal/BurstReshapedPHST. Yipeng Wang 0001, Haoting Liu, Jiapeng Zhao, Jinqiao Shi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Reliable Open-Set Network Traffic ClassificationabstractThe widespread use of modern network communications necessitates effective resource control and management in TCP/IP networks. However, most existing network traffic classification methods are limited to labeled known classes and struggle to handle open-set scenarios, where known classes coexist with significant volumes of unknown classes of traffic. To solve this problem more accurately and reliably, we propose RoNeTC. This method achieves high-precision classification by enhancing feature extraction and quantifying the reliability of classification decisions through uncertainty estimation. For feature extraction, we divide each packet of a flow into three views for parallel training, integrating both local and global feature representations across multiple packets to enhance accuracy. We devise a second-order classification probability to quantify the reliability of the classifier’s results and to visualize the reliability of open-set flow classification in terms of uncertainty. Additionally, we dynamically fuse classification decisions from multiple views, evaluating decision uncertainty to classify known and unknown flows and ensure robust, reliable results. We compare RoNeTC with four state-of-the-art (SOTA) methods in six open-set scenarios. RoNeTC outperforms the other methods by an average of 25.94% in F1 across all open-set scenarios, indicating its superior performance in open-set network traffic classification. Xueman Wang, Yipeng Wang 0001, Yingxu Lai, Zhiyu Hao, Alex X. Liu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Joint Communication-Motion Planning for UAV Swarm against Jamming with Multi-Agent Deep Reinforcement LearningabstractIn this paper, we investigate the joint communication-motion planning problem for unmanned aerial vehicle (UAV) swarm in the presence of jammers. Specifically, we consider a cluster-based UAV swarm architecture, where multiple cluster member (CM) UAVs transmit messages to a cluster head (CH) UAV through air-to-air links affected by malicious jammers. Our objective is to maximize the sum uplink rate of the UAV swarm by optimizing the trajectories and the transmit power of all UAVs. To achieve this goal, we formulate a joint multi-UAV trajectories and transmit power optimization problem under speed, transmit power, trajectories and received signal-to-interference-plus-noise ratio (SINR) constraints. In order to solve the problem, we establish a Markov decision process (MDP). For the multi-agent environment and the high-dimensional continuous action space, we adopt a multi-agent twin delayed deep deterministic (MATD3) policy gradient-based algorithm. Simulation results show that the proposed scheme can effectively improve the sum uplink rate of the UAV swarm compared to the baseline schemes. Zhenxin Guo, Yiming Liu 0002, Yipeng Wang 0001, Baoling Liu |
PIMRC | 3 |
| 2024 | An adaptive classification and updating method for unknown network traffic in open environments
Siqi Le, Yingxu Lai, Yipeng Wang 0001, Huijie He |
Comput. Networks | 3 |
| 2024 | MPAF: Encrypted Traffic Classification With Multi-Phase Attribute FingerprintabstractThe widespread use of cryptographic protocols such as Transport Layer Security (TLS) has necessitated the development of effective methods for encrypted traffic classification. The existing methods relying on a single feature source face challenges in achieving high accuracy and efficiency simultaneously. Additionally, there is a decrease in accuracy in complex scenarios, posing significant challenges for networks and security services based on application-level traffic classification. In this paper, we propose Multi-Phase Attribute Fingerprint (MPAF), an encrypted traffic classification system that overcomes these limitations. MPAF leverages three phases to separately leverage attributes that emerge at different time periods of encrypted traffic communication. Additionally, we transform discrete attributes into computable vectors through embedding and design a classifier for the multi-phase mechanism based on a leaf node masking tree. The experimental results show that MPAF achieves a classification accuracy ranging from 96.33% to 99.42% and an average waiting time (AWT) ranging from 0.18s to 0.45s. MPAF outperforms other approaches in scenarios with high robustness requirements, including small-scale training datasets, cross-dataset classification, and unknown application recognition. Yipeng Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Joint Trajectory Optimization and Task Offloading for UAV-Assisted Mobile Edge ComputingabstractUnmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) has been touted as a promising solution for providing computing services in disaster relief and other settings due to its flexibility and ease of deployment. Nevertheless, providing computing services for a large number of mobile devices is challenged by UAVs’ limited computation and energy resources. To this end, we propose a scheme for joint trajectory optimization and task offloading that aims to minimize the total delay of all computing tasks. Our proposed scheme involves formulating the scheduling of mobile devices and computing tasks, adjustment of UAV flight angle and speed, and transmission power control as a non-convex mixed integer programming problem. In order to address the issue, we establish a Markov decision process (MDP) for UAV-assisted MEC systems. Given the high-dimensional continuous action space, we adopt a reinforcement learning algorithm based on Deep Deterministic Policy Gradient (DDPG). The results of the simulation indicate that our suggested scheme outperforms the baseline schemes in processing delay, and the DDPG-based algorithm exhibits rapid convergence. Yipeng Wang 0001, Yiming Liu 0002, Baoling Liu |
PIMRC | 1 |
| 2023 | Relational reasoning-based approach for network protocol reverse engineeringabstractExtracting the protocol format specifications from packets plays a critical role in many applications, such as application protocol parsing, vulnerability scanning, as well as malware behaviour analysis. In this study, we propose RelaNet, a novel relational reasoning-based method for network protocol reverse engineering. It is based on the key insight that n-grams of packets have context relations. Such relations are especially informative between keywords and can be used to infer protocol formats. RelaNet contains three modules to mimic the analysis strategy used by experts: coarse structure generation, relation learning and fine structure generation. In coarse structure generation, RelaNet first constructs a coarse-grained structure based on the occurrence frequency of n-grams. Relation learning is then performed to discover the context relations between the n-grams in the structure. By using such relations, we can eventually discover the strong context relations that exist between keywords, which can be used to accurately generate a fine-grained structure, namely the protocol format. We implement RelaNet and evaluate it on two publicly available datasets, the experimental results demonstrate the effectiveness and efficiency of RelaNet for protocol format inference. Furthermore, we compare RelaNet with state-of-the-art methods, and the results show our approach outperforms those methods. Yingxu Lai, Yipeng Wang 0001 |
Comput. Networks | 3 |
| 2023 | A data skew-based unknown traffic classification approach for TLS applications
Huijie He, Yingxu Lai, Yipeng Wang 0001, Siqi Le |
Future Gener. Comput. Syst. | 3 |
| 2023 | MSGAN: multi-stage generative adversarial network-based data recovery in cyber-attacks
Bitao Tian, Yingxu Lai, Samuel S. M. Sun, Yipeng Wang 0001, Jing Liu 0028 |
Neural Comput. Appl. | 4 |
| 2023 | Intrusion Detection System Based on In-Depth Understandings of Industrial Control LogicabstractIn industrial control systems (ICSs), intrusion detection is a vital task. Conventional intrusion detection systems (IDSs) rely on manually designed rules. These rules heavily depend on professional experience, thereby making it challenging to represent the increasingly complicated industrial control logic. Although deep learning-based approaches provide better accuracy than other methods, they can only provide alerts. However, they cannot provide administrators with detailed information. In this study, we propose the logic understanding IDS (LU-IDS), which is a rule-based IDS with in-depth understandings of industrial control logic. Our proposed LU-IDS uses a specially designed deep learning-based model to capture features automatically and carry out attack classification. More importantly, it analyzes the knowledge learned from the classification of attacks to understand the abnormal industrial control logic and generate rules. The experimental results indicate that our proposed LU-IDS demonstrates excellent performance on intrusion detection. The rules generated by our proposed LU-IDS can be used to successfully detect all types of attacks on two public datasets. Motong Sun, Yingxu Lai, Yipeng Wang 0001, Jing Liu 0028, Beifeng Mao, Haoran Gu |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | A Two-Phase Approach to Fast and Accurate Classification of Encrypted TrafficabstractEncryption technology has been widely used in today’s network communications. The early classification of encrypted flows is of great value to the control, allocation and management of resources in TCP/IP networks. In this paper, we propose TaTic, an early classification method for encrypted traffic, which aims to reduce the time spent observing the encrypted flows to be classified, and at the same time ensure the flow classification accuracy. TaTic is based on our key observation that the majority of encrypted flows can be classified accurately using only the first few packets, and we call such flows “easy flows”, whereas the rest of encrypted flows requires more packets for fine-grained analysis to achieve accurate traffic classification, and we call such flows “hard flows”. Given an encrypted flow, in the first phase, we use only the first few packets to quickly determine whether it is an easy flow or a hard flow; if it is an easy flow, we directly classify it in this phase; otherwise, we use more packets to perform traffic classification in the second phase. Therefore, we can greatly reduce the time spent in observing the flows without sacrificing the classification accuracy. Our experimental results show that TaTic can greatly reduce the unnecessary time spent in observing the flow to be classified, and at the same time ensure high classification accuracy. We compare our experimental results of TaTic with four existing methods. TaTic is superior to the existing methods in terms of both classification accuracy and average waiting time. Yipeng Wang 0001, Huijie He, Yingxu Lai, Alex X. Liu |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Encrypted TLS Traffic Classification on Cloud PlatformsabstractNowadays, encryption technology has been widely used to protect user privacy. With the explosive growth of mobile Internet, encrypted TLS traffic rises sharply and occupies a great share of current Internet traffic. In reality, the classification of encrypted TLS traffic on cloud platforms brings a new challenge to traditional encrypted traffic classification methods, because some information such as certificates in the TLS flows is no longer effective. In this paper, we apply deep learning technology to the problem of encrypted TLS traffic classification on cloud platforms, and propose NeuTic, which takes the packet sequence of each TLS flow as the input, and effectively classifies raw TLS flows generated by many “cloud” applications. Our approach is able to automatically capture the long-range dependencies between elements in the packet sequences for robust and accurate encrypted TLS traffic classification. In NeuTic, we first convert each TLS flow into three attribute sequences. Then, we train a multi-application traffic classification model using our newly designed deep learning model. Finally, we use the well-trained classification model to classify new incoming TLS flows. We conduct comprehensive experiments on real-world application traces covering multiple “cloud” applications from three different companies. In addition, we compare our experimental results of NeuTic with two deep learning-based methods for encrypted traffic classification. NeuTic outperforms the state-of-the-art approaches in classification accuracy. Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | Zen-tor: A Zero Knowledge Known-Unknown Traffic Classification MethodabstractHome smart devices are widely used today, but due to their inherent drawbacks, one's home might be in danger if the home network has too many unknown traffic where malicious traffic hides. We present a conceptually new method called Zero Knowledge Known-unknown Traffic Classification(Zen-tor) which utilizes Generative Adversarial Networks(GAN) and con-volutional network. Zen-tor can classify unknown network traffic from known one under the situation of knowing zero knowledge of unknown traffic, thus it can be trained to protect private home network with only known traffic. We evaluate Zen-tor on a publicly available dataset, and the results show that Zen-tor has excellent unknown classification accuracy and outperforms the state-of-the-art unknown traffic classification methods. Yizhe Gu, Yingxu Lai, Yipeng Wang 0001 |
GLOBECOM | 3 |
| 2022 | FITIC: A Few-shot Learning Based IoT Traffic Classification MethodabstractWith the rapid development and wide application of Internet of Things (IoT) technology, Internet Service Providers need to accurately classify IoT traffic to provide hierarchical network management and network protection for highly het-erogeneous IoT devices. Currently, popular traditional machine learning and deep learning-based approaches to IoT traffic classification require large amounts of labeled traffic to build classification models. However, in practice simple IoT traffic with simple operating modes can be identified with only a small amount of labeled traffic and some classes of IoT devices only generate a limited amount of traffic, therefore, the aforementioned methods is not applicable in such scenarios. In this paper, we propose FITIC, a novel IoT traffic classification method based on few-shot learning. FITIC proposes a feature construction method for IoT traffic characteristics and can classify IoT traffic with only a limited number of labeled traffic samples. We evaluate FITIC on two publicly available datasets, and the experimental results show that FITIC has excellent classification accuracy and outperforms the state-of-the-art traffic classification methods. Wenxu Jia, Yipeng Wang 0001, Yingxu Lai, Huijie He, Ruiping Yin |
ICCCN | 2 |
| 2022 | Autonomous Anti - interference Identification of $\text{IoT}$ Device Traffic based on Convolutional Neural NetworkabstractNetwork traffic classification plays a vital role in many fields such as intrusion detection, network management, and network security. As the proportion of IoT device traffic increases, many approaches to identifying IoT device types through traffic have emerged. Specifically, Deep Learning (DL) has been proven to be a more efficient approach for encrypted traffic identification than other traditional methods. However, most existing classification models are created in static datasets from the closed world, so they can only classify within a limited domain. In this case, interfering traffic in the open world is easily misidentified by classifiers as IoT device traffic. An autonomous framework is proposed to tackle this issue, effectively identifying the device type according to the grayscale graph generated by packet payload and automatically updating to adapt to the unknown environment in the open world. The core of the proposed framework consists of a packet graph-vector transformer, a CNN-based classifier, and an autonomous optimizer. The optimizer can filter interfering data and optimize the model by updating the training dataset. We comprehensively evaluated the proposed framework on two datasets, one taken from the UNSW IoT traces and the other collected by our experiments, containing traffic generated from two devices and three open-world scenarios. The results demonstrate that the proposed framework can update the training dataset by unsupervised filtering interference packets, enabling the model to automatically suit complex environments for accurate and robust IoT device type identification in the open world. Shuhe Liu, Yongzheng Zhang 0002, Yipeng Wang 0001 |
IJCNN | 4 |
| 2022 | Identifying malicious nodes in wireless sensor networks based on correlation detectionabstractThe wireless sensor network (WSN) is a multi-hop wireless network that comprises multiple sensor nodes arranged in a self-organized manner. It is usually deployed in unattended areas where sensor nodes can easily be infiltrated by attackers who can affect the detection results by injecting false data. This paper proposes a malicious-node identification method based on correlation theory that prevents fault data injection attacks. First, anomalies among similar types of sensor data are detected based on time correlation. Second, malicious nodes are identified based on spatial correlation. Third, the identified malicious nodes are verified based on event correlation. The experimental results and their comparison with those of existing methods show that the proposed scheme has better recall with lower false-positive and false-negative rates than those of the traditional fuzzy reputation model and weighted-trust-based methods. Yingxu Lai, Liyao Tong, Jing Liu 0028, Yipeng Wang 0001, Hua Qin |
Comput. Secur. | 4 |
| 2022 | DEIDS: a novel intrusion detection system for industrial control systemsabstractAbstract Owing to the development of industrial production, the hidden danger in industrial control systems (ICSs) has considerably increased, causing challenges in traditional safety defense methods. The combination of machine-learning or deep-learning algorithms and intrusion detection systems (IDSs) has become the mainstream method for solving this problem. However, these methods depend on a massive amount of high-quality attack traffic data, which cannot be obtained easily owing to the independence and unique characteristics of ICSs. In this study, we apply the reconstructed convolutional neural network and a data expansion algorithm named CenterBorderline_SMOTE (CB_SMOTE) to an IDS and propose data expansion intrusion detection system (DEIDS). The DEIDS is an end-to-end detection model that learns representative attack features from raw traffic and classifies them in a unified framework. Moreover, we adopt the classification activation map structure, which can deeply mine the potential characteristics of traffic and enhance the effectiveness of attack features. While enhancing the data quality, we introduce the designed CB_SMOTE algorithm into DEIDS to expand the data and solve the problem of insufficient attack data in the system. Our comprehensive experiments on different open datasets indicate that DEIDS achieves an excellent performance (97 $$\%$$ % detection accuracy) and outperforms the state-of-the-art methods. The experimental results also show that our method has high efficiency and high accuracy in processing ICSs datasets. Haoran Gu, Yingxu Lai, Yipeng Wang 0001, Jing Liu 0028, Motong Sun, Beifeng Mao |
Neural Comput. Appl. | 3 |
| 2022 | Correction to: DEIDS: a novel intrusion detection system for industrial control systems
Haoran Gu, Yingxu Lai, Yipeng Wang 0001, Jing Liu 0028, Motong Sun, Beifeng Mao |
Neural Comput. Appl. | 3 |
| 2022 | A Multi-Scale Feature Attention Approach to Network Traffic Classification and Its Model ExplanationabstractNetwork traffic classification, the task of associating network traffic with their generating application protocols or applications, is valuable for the control, allocation, and management of resources in today’s TCP/IP networks. In this paper, we propose Ulfar, a multi-scale feature attention approach to network traffic classification, which uses convolutional neural networks (CNN) as the building block of the deep packet analysis model. In Ulfar, we take only one packet per flow for network traffic classification. Ulfar is based on the key insight that format-related bytes appear at fixed offsets or in a specific pattern in the IP packet, and these format-related bytes are important for accurate network traffic classification. Our neural network model can automatically recover the format-related bytes by building high-level, multi-scale${n}$-gram features from raw byte sequences. In addition, at the representation learning side, we try to understand what patterns and signatures our neural network model learns from network traffic. We evaluate Ulfar using two publicly available datasets, and our experimental results show that Ulfar can conduct accurate network traffic classification. Also, we compare the results of Ulfar with four state-of-the-art approaches, and find that Ulfar has the ability to classify network traffic more accurately. Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Xin Liu 0002 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Inspector: A Semantics-Driven Approach to Automatic Protocol Reverse Engineering
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Yipeng Wang 0001 |
CollaborateCom (1) | 6 |
| 2021 | Finding disposable domain names: A linguistics-based stacking approach
Yuwei Zeng, Xiao-chun Yun, Xunxun Chen, Boquan Li 0002, Haiwei Tsang, Yipeng Wang 0001, Tianning Zang, Yongzheng Zhang 0002 |
Comput. Networks | 6 |
| 2020 | Joint Character-Level Word Embedding and Adversarial Stability Training to Defend Adversarial TextabstractText classification is a basic task in natural language processing, but the small character perturbations in words can greatly decrease the effectiveness of text classification models, which is called character-level adversarial example attack. There are two main challenges in character-level adversarial examples defense, which are out-of-vocabulary words in word embedding model and the distribution difference between training and inference. Both of these two challenges make the character-level adversarial examples difficult to defend. In this paper, we propose a framework which jointly uses the character embedding and the adversarial stability training to overcome these two challenges. Our experimental results on five text classification data sets show that the models based on our framework can effectively defend character-level adversarial examples, and our models can defend 93.19% gradient-based adversarial examples and 94.83% natural adversarial examples, which outperforms the state-of-the-art defense models. Yongzheng Zhang 0002, Yipeng Wang 0001, Zheng Lin 0001 |
AAAI | 3 |
| 2020 | Gated POS-Level Language Model for Authorship VerificationabstractAuthorship verification is an important problem that has many applications. The state-of-the-art deep authorship verification methods typically leverage character-level language models to encode author-specific writing styles. However, they often fail to capture syntactic level patterns, leading to sub-optimal accuracy in cross-topic scenarios. Also, due to imperfect cross-author parameter sharing, it's difficult for them to distinguish author-specific writing style from common patterns, leading to data-inefficient learning. This paper introduces a novel POS-level (Part of Speech) gated RNN based language model to effectively learn the author-specific syntactic styles. The author-agnostic syntactic information obtained from the POS tagger pre-trained on large external datasets greatly reduces the number of effective parameters of our model, enabling the model to learn accurate author-specific syntactic styles with limited training data. We also utilize a gated architecture to learn the common syntactic writing styles with a small set of shared parameters and let the author-specific parameters focus on each author's special syntactic styles. Extensive experimental results show that our method achieves significantly better accuracy than state-of-the-art competing methods, especially in cross-topic scenarios (over 5\% in terms of AUC-ROC). Linshu Ouyang, Yongzheng Zhang 0002, Yipeng Wang 0001 |
IJCAI | 5 |
| 2020 | Unified Graph Embedding-Based Anomalous Edge DetectionabstractDetecting anomalous edges in graph-structured data plays an important role in many fields such as finance, social network, and network security. Recently, graph embedding based anomaly detection methods show promising results. These methods typically encode graph structure information into vector representation and apply general anomaly detection methods. However, since the parameters in these two parts are learned separately with different objectives, the learned representation may contain some information irrelevant to the task. It would be ideal if we can combine representation learning and anomaly detection into one objective function to force the model to focus on learning task relevant patterns. In this paper, we propose a novel end-to-end neural network architecture that can accurately estimate the probability distribution of edges in the graph based on its local structure. An edge has a high chance to be considered an anomaly if the probability of its existence is low. Extensive experiments on several public datasets at different scales show that the accuracy and scalability of our method outperform other methods by a large margin. Linshu Ouyang, Yongzheng Zhang 0002, Yipeng Wang 0001 |
IJCNN | 3 |
| 2020 | A Feature Ensemble-based Approach to Malicious Domain Name Identification from Valid DNS ResponsesabstractIdentifying malicious domain names in Internet activities has become an effective method to protect Internet users. Previous works have achieved great identification results, but they highly rely on historical Domain Name System (DNS) responses and external intelligence sources. Thus, they may fail to identify unknown domain name without any prior knowledge. In this paper, we propose Glacier, a feature ensemble-based approach to identifying malicious domain names from valid DNS responses. Glacier addresses the aforementioned problem by utilizing two types of features in domain name strings: the linguistical features and the statistical features. (1) Linguistical features are vector representations generated from the character sequences of domain names by a bidirectional long short-term memory (BiLSTM) neural network. It is worthy to notice that we modify the last BiLSTM layer to enhance the expressiveness of the linguistical features. (2) Statistical features are six manually designed statistics that represent the structural information of a domain name. Structural information can hardly be learnt by a BiLSTM neural network directly. Thus, combining statistical features with linguistical features can improve the effectiveness of malicious domain name identification. We evaluate the identification ability of Glacier on a real-world domain name data set. The best metrics of Glacier are an average accuracy of 90.86% and an average F1-score of 84.37%. Our experimental results show that Glacier can accurately identify resolvable malicious domain names without any DNS traffic data or prior knowledge about unknown domain names. Yongzheng Zhang 0002, Yipeng Wang 0001 |
IJCNN | 3 |
| 2020 | Khaos: An Adversarial Neural Network DGA With High Anti-Detection AbilityabstractA botnet is a network of remote-controlled devices that are infected with malware controlled by botmasters in order to launch cyber attacks. To evade detection, the botmaster frequently changes the domain name of his Command and Control (C&C) server. Notice that most of these types of domain names are generated by domain generation algorithms (DGAs). In this paper, we propose Khaos, a novel DGA with high anti-detection ability based on neural language models and the Wasserstein Generative Adversarial Network (WGAN). The key insight of our research is that real domain names are composed of readable syllables and acronyms, and thus we can arrange syllables and acronyms using neural language models to mimic real domain names. In Khaos, we first find the most common n-grams in real domain names, then tokenize these domain names into n-grams, and finally synthesize new domain names after learning arrangements of n-grams from real domain names. We carry out experiments using a variety of state-of-the-art DGA detection approaches: the statistics-based, the distribution-based, the LSTM-based and the graph-based detection approach. Our experimental results show that the average distance for detecting Khaos under the distribution-based detection approach is 0.64, the AUCs of Khaos under the statistics-based and the LSTM-based detection approach are 0.76 and 0.57, respectively, and the precision of Khaos under the graph-based detection approach is 0.68. Our work proves that the existing detection approaches have big troubles in detecting Khaos, and Khaos has better anti-detection ability than state-of-the-art DGAs. In addition, we find that training the existing detection approach on a dataset including the domain names generated by Khaos can improve its detection ability. Xiao-chun Yun, Yipeng Wang 0001, Tianning Zang, Yuan Zhou 0008, Yongzheng Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Translating with Bilingual Topic Knowledge for Neural Machine TranslationabstractThe dominant neural machine translation (NMT) models that based on the encoder-decoder architecture have recently achieved the state-of-the-art performance. Traditionally, the NMT models only depend on the representations learned during training for mapping a source sentence into the target domain. However, the learned representations often suffer from implicit and inadequately informed properties. In this paper, we propose a novel bilingual topic enhanced NMT (BLTNMT) model to improve translation performance by incorporating bilingual topic knowledge into NMT. Specifically, the bilingual topic knowledge is included into the hidden states of both encoder and decoder, as well as the attention mechanism. With this new setting, the proposed BLT-NMT has access to the background knowledge implied in bilingual topics which is beyond the sequential context, and enables the attention mechanism to attend to topic-level attentions for generating accurate target words during translation. Experimental results show that the proposed model consistently outperforms the traditional RNNsearch and the previous topic-informed NMT on Chinese-English and EnglishGerman translation tasks. We also introduce the bilingual topic knowledge into the newly emerged Transformer base model on English-German translation and achieve a notable improvement. Xiangpeng Wei, Yue Hu 0002, Luxi Xing, Yipeng Wang 0001 |
AAAI | 4 |
| 2019 | Pontus: A Linguistics-Based DGA Detection SystemabstractMany botmasters use domain generation algorithms (DGA) to generate a host of malicious algorithmically- generated domains (mAGDs) and then choose several mAGDs for actual command and control (C2) communication. The botmasters use different seeds (e.g., timestamp) to generate different mAGDs, which makes the communication mechanism resilient to blacklisting. Thus the botmasters can hide C2 channels very well. If we can quickly detect these mAGDs from DNS traffic, we will effectively block the communication. In this paper, we propose a novel system, called Pontus, to detect mAGDs from DNS traffic. Pontus extract features exclusively from the individual domain names. We compare Pontus with the state-of-the-art system and find that Pontus improves the precision by at least 4.7%. Dingkui Yan, Huilin Zhang, Yipeng Wang 0001, Tianning Zang, Yuwei Zeng |
GLOBECOM | 3 |
| 2019 | A Comprehensive Measurement Study of Domain-Squatting AbuseabstractDomain-squatting abuse refers to the premeditated attempt by an attacker to register perceptively confusing domain names thereby tricking visitors into querying them. There are totally five squatting types have been investigated so far, namely typo-squatting, bit-squatting, homograph-squatting, sound-squatting, and combo-squatting. Existing researches only focus on one specific squatting type and never explore the relationship among them. In this paper, we perform the first comprehensive measurement study of domain-squatting abuse. We select 786 the most queried domains, and hunt for squatting abuses against them in ISP-level DNS traffic. We find that although typo-squatting accounts for most of squatting domains, combo-squatting are able to attract more traffic. Our further case studies show that parking ads is still the most important way for attackers to make profits. The only exception is combo-squatting, in which squatters tend to leverage the reputation of squatted domains to develop their own business. It is worth noting that some squatting domains are even used to deliver malware. Moreover, the Alexa ranks of certain squatting domains have already surpassed the original domains. These results clearly call for the need to better protect the intellectual property of domain names. Yuwei Zeng, Tianning Zang, Yongzheng Zhang 0002, Xunxun Chen, Yipeng Wang 0001 |
ICC | 5 |
| 2019 | Rethinking Encrypted Traffic Classification: A Multi-Attribute Associated Fingerprint ApproachabstractWith the unprecedented prevalence of mobile network applications, cryptographic protocols, such as the Secure Socket Layer/Transport Layer Security (SSL/TLS), are widely used in mobile network applications for communication security. The proven methods for encrypted video stream classification or encrypted protocol detection are unsuitable for the SSL/TLS traffic. Consequently, application-level traffic classification based networking and security services are facing severe challenges in effectiveness. Existing encrypted traffic classification methods exhibit unsatisfying accuracy for applications with similar state characteristics. In this paper, we propose a multiple-attribute-based encrypted traffic classification system named Multi-Attribute Associated Fingerprints (MAAF). We develop MAAF based on the two key insights that the DNS traces generated during the application runtime contain classification guidance information and that the handshake certificates in the encrypted flows can provide classification clues. Apart from the exploitation of key insights, MAAF employs the context of the encrypted traffic to overcome the attribute-lacking problem during the classification. Our experimental results demonstrate that MAAF achieves 98.69% accuracy on the real-world traceset that consists of 16 applications, supports the early prediction, and is robust to the scale of the training traceset. Besides, MAAF is superior to the state-of-the-art methods in terms of both accuracy and robustness. Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Yipeng Wang 0001 |
ICNP | 5 |
| 2019 | A Linguistics-based Stacking Approach to Disposable Domains DetectionabstractMore Internet services tend to collect the one-time information from clients via DNS queries. Notably, the uncertainty of such transient information makes these domain names be queried only once in their lifetime. This type of domain is called disposable domain. Although they are not malicious, the efficiency of DNS infrastructures will still be affected by their ever-increasing number. In this paper, we propose Vogers, a linguistics-based stacking model, to detect the disposable domains. Our evaluation demonstrates that Vogers decreases the false positive rate by more than 19%, compared with the prior art, while maintaining the true positive rate above 98.9%. Yuwei Zeng, Yongzheng Zhang 0002, Tianning Zang, Xunxun Chen, Yipeng Wang 0001 |
ICNP | 5 |
| 2018 | A Stacking Approach to Objectionable-Related Domain Names Identification by Passive DNS Traffic (Short Paper)
Yongzheng Zhang 0002, Tianning Zang, Zhizhou Liang, Yipeng Wang 0001 |
CollaborateCom | 5 |
| 2018 | A cascade forest approach to application classification of mobile tracesabstractWith the rapid development of mobile networks, mobile traffic classification, a mapping of mobile traffic to mobile applications, becomes more and more important for variant networking and security issues, such as network management, monitoring and the detection of malware activities. In this paper, we propose CFMTC (Cascade Forest for Mobile Traces Classification), a mobile network trace-based traffic classification system, which exploits flow statistical features extracted from mobile traces. Compared to other classification approaches, our system is based upon the key insight that deep learning techniques and the statistical features of bidirectional flows of mobile traces can be combined together for accurate mobile application classification. In CFMTC, we first filter UDP and TCP flows from mobile traces according to the flow attributes (Source IP, Destination IP, Source port, Destination port, Protocol), and then train Cascade Forest to classify raw mobile traces. We use a feature selection method to find the optimal feature set and determine the influence of different features. Our approach involves the following key features: 1) suitable for mobile traces classification; 2) adapted Cascade Forest algorithm for mobile traffic classification; 3) applicable to both connection-oriented protocols and connection-less protocols; 4) effective for both encrypted and non-encrypted flows. We implement CFMTC and conduct extensive evaluations on mobile network traces containing text, audio and video flows generated by Kuwo Music, WeChat, PPTV Live traces. Our experimental results show that CFMTC has the ability to accurately classify the mobile traces of the target mobile applications with an average accuracy of about 88.71%. Our experimental results prove that CFMTC is a robust system, and meanwhile displays competitive performance in practice. Shuzhuang Zhang, Yipeng Wang 0001 |
WCNC | 5 |
| 2018 | Information Propagation Prediction Based on Key Users Authentication in MicrobloggingabstractIn microblogging, key users are a significant factor for information propagation. Key users can affect information propagation size while retweeting the information. In this paper, to predict information propagation, we propose a novel linear model based on key users authentication. This model mines key users to dynamically improve the linear model while predicting information propagation. So our model can not only predict information propagation but also mine key users. Experimental results show that our model can achieve remarkable efficiency on predicting information propagation problem in real microblogging networks. At the same time, our model can find the key users who affect information propagation. Miao Yu 0006, Yongzheng Zhang 0002, Tianning Zang, Yipeng Wang 0001 |
Secur. Commun. Networks | 4 |
| 2017 | Supporting Real-Time Analytic Queries in Big and Fast Data Environments
Guangjun Wu, Xiao-chun Yun, Chao Li 0062, Yipeng Wang 0001, Xiaoyu Zhang 0002, Siyu Jia, Guangyan Zhang |
DASFAA (2) | 5 |
| 2017 | Rethinking robust and accurate application protocol identification
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Tianning Zang |
Comput. Networks | 1 |
| 2017 | A nonparametric approach to the automated protocol fingerprint inference
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Guangjun Wu |
J. Netw. Comput. Appl. | 1 |
| 2016 | Multiple-combinational-channel: A network architecture for workload balance and deadlock free
Yipeng Wang 0001, Huandong Wang, Wenxiang Wang, Hua Jing, Guangfei Zhang |
Future Gener. Comput. Syst. | 2 |
| 2016 | A Semantics-Aware Approach to the Automated Network Protocol IdentificationabstractTraffic classification, a mapping of traffic to network applications, is important for a variety of networking and security issues, such as network measurement, network monitoring, as well as the detection of malware activities. In this paper, we propose Securitas, a network trace-based protocol identification system, which exploits the semantic information in protocol message formats. Securitas requires no prior knowledge of protocol specifications. Deeming a protocol as a language between two processes, our approach is based upon the new insight that the n-grams of protocol traces, just like those of natural languages, exhibit highly skewed frequency-rank distribution that can be leveraged in the context of protocol identification. In Securitas, we first extract the statistical protocol message formats by clustering n-grams with the same semantics, and then use the corresponding statistical formats to classify raw network traces. Our tool involves the following key features: 1) applicable to both connection oriented protocols and connection less protocols; 2) suitable for both text and binary protocols; 3) no need to assemble IP packets into TCP or UDP flows; and 4) effective for both long-live flows and short-live flows. We implement Securitas and conduct extensive evaluations on real-world network traces containing both textual and binary protocols. Our experimental results on BitTorrent, CIFS/SMB, DNS, FTP, PPLIVE, SIP, and SMTP traces show that Securitas has the ability to accurately identify the network traces of the target application protocol with an average recall of about 97.4% and an average precision of about 98.4%. Our experimental results prove Securitas is a robust system, and meanwhile displaying a competitive performance in practice. Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002, Yu Zhou 0015 |
IEEE/ACM Trans. Netw. | 2 |
| 2015 | Rethinking Robust and Accurate Application Protocol Identification: A Nonparametric ApproachabstractProtocol traffic analysis is important for a variety of networking and security infrastructures, such as intrusion detection and prevention systems, network management systems, and protocol specification parsers. In this paper, we propose ProHacker, a nonparametric approach that extracts robust and accurate protocol keywords from network traces and effectively identifies the protocol trace from mixed Internet traffic. ProHacker is based on the key insight that the n-grams of protocol traces have highly predictable statistical nature that can be effectively captured by statistical language models and leveraged for robust and accurate protocol identification. In ProHacker, we first extract protocol keywords using a nonparametric Bayesian statistical model, and then use the corresponding protocol keywords to classify protocol traces by a semi-supervised learning algorithm. We implement and evaluate ProHacker on real-world traces, including SMTP, FTP, PPLive, SopCast, and PPStream, and our experimental results show that ProHacker can accurately identify the protocol trace with an average precision of about 99.42% and an average recall of about 98.64%. We also compare the results of ProHacker to two state-of-the-art approaches ProWord and Securitas using backbone traffic. We show that ProHacker provides significant improvements on precision and recall for online protocol identification. Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002 |
ICNP | 1 |
| 2015 | A Markov Random Field Approach to Automated Protocol Signature Inference
Yongzheng Zhang 0002, Yipeng Wang 0001, Jianliang Sun, Xiaoyu Zhang 0002 |
SecureComm | 3 |
| 2015 | Update vs. upgrade: Modeling with indeterminate multi-class active learning
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Xiao-chun Yun, Guangjun Wu, Yipeng Wang 0001 |
Neurocomputing | 6 |
| 2015 | Unsupervised adaptive sign language recognition based on hypothesis comparison guided cross validation and linguistic prior filtering
Yu Zhou 0015, Xiaokang Yang 0001, Yongzheng Zhang 0002, Yipeng Wang 0001, Xiujuan Chai, Weiyao Lin |
Neurocomputing | 5 |
| 2014 | A Segmentation Pattern Based Approach to Automated Protocol IdentificationabstractIn-depth understanding of network traffic is important for a variety of applications, such as network management and network security. In this paper, we propose a novel protocol identification system PSKS, which relies on the statistical signatures of network packet payloads. The proposed approach is based on the key insight that message segmentation patterns can be leveraged for accurate application identification. Specifically, the segmentation possibility for every position of protocol messages exhibits highly skewed frequency distribution due to the reason that different protocols have different message formats (i.e., Distinct message segmentation patterns). Motivated by this observation, we want to extract statistical application fingerprints by exploiting the message segmentation patterns. In PSKS, we first extract the message segmentation patterns by scoring the segmentation possibility scale for each position of messages, and then extract statistical signatures by Kolmogorov-Smirnov test and feed the signatures to tri-training, a collaborative learning algorithm. The tri-training can improve the generalization ability of our final classifier. We implemented and evaluated PSKS, and the experimental results show that PSKS achieves an average precision and recall of approximately 98%. Yafei Sang, Yongzheng Zhang 0002, Yipeng Wang 0001, Yu Zhou 0015 |
PDCAT | 3 |
| 2014 | Visual Similarity Based Anti-phishing with the Combination of Local and Global FeaturesabstractPhishing uses a fake Web page to steal personal sensitive information such as credit card numbers and passwords. Generally, the fake Web page is visually similar to the legitimate target Web page. The phishers can obtain financial benefits through these information. Anti-phishing is very important for a variety of applications such as phishing attacks, online transaction security, and user privacy protection. In this paper, we propose a novel and effective visual similarity based phishing detection approach that compares the snapshot image pair of the suspected Web page and the protected Web page. The proposed approach is based on the key insight that both the local and the global features of the Web page image can be used to represent the visual characteristics of the Web page together. This approach is purely on the image level, and thus can effectively deal with the non-text phishing tricks including images or Flashes objects in the HTML contents. For the local feature, the existence of the target logo is detected. For the global feature, the similarity of the visible part of the Web page is considered. We implemented and evaluated the proposed approach on a large scale dataset consisting of 2,129 real world phishing Web pages and 1,367 irrelevant legitimate Web pages. The experimental results show that the proposed approach can achieve over 90.00% true positive rate and 97.00% true negative rate. Our approach has been applied in the anti-phishing project of a major Internet Service Provider and gives a periodical reports to the potential users. Yu Zhou 0015, Yongzheng Zhang 0002, Yipeng Wang 0001, Weiyao Lin |
TrustCom | 4 |
| 2012 | A semantics aware approach to automated reverse engineering unknown protocolsabstractExtracting the protocol message format specifications of unknown applications from network traces is important for a variety of applications such as application protocol parsing, vulnerability discovery, and system integration. In this paper, we propose ProDecoder, a network trace based protocol message format inference system that exploits the semantics of protocol messages without the executable code of application protocols. ProDecoder is based on the key insight that the n-grams of protocol traces exhibit highly skewed frequency distribution that can be leveraged for accurate protocol message format inference. In ProDecoder, we first discover the latent relationship among n-grams by first grouping protocol messages with the same semantics and then inferring message formats by keyword based clustering and cluster sequence alignment. We implemented and evaluated ProDecoder to infer message format specifications of SMB (a binary protocol) and SMTP (a textual protocol). Our experimental results show that ProDecoder accurately parses and infers SMB protocol with 100% precision and recall. For SMTP, ProDecoder achieves approximately 95% precision and recall. Yipeng Wang 0001, Xiao-chun Yun, Zubair Shafiq, Alex X. Liu, Danfeng Yao, Yongzheng Zhang 0002, Li Guo 0001 |
ICNP | 1 |
| 2012 | A General Framework of Trojan Communication Detection Based on Network TracesabstractBecause of the widespread Trojan, Internet users become more and more vulnerable to the threat of information leakage. Traditional techniques of Trojan detection were classified into two main categories: host-based and network-based. Unfortunately, existing techniques are insufficient and limited, because of the following reasons: (1)only uncover the known Trojan while inefficiently detecting novel samples, (2) should be adjusted in a timely fashion even a trivial change is applied, and (3)become computationally more expensive. In our work, we focus on a network behavior based method to address the limitations of previous network-based approaches. We analyze the profile of network behavior at two levels: (i)flow-level, (ii)IP-level. Our approach present two main advantages: (1)capture more detailed information to describe the network behavior profile, (2)consume lower computational overhead. We proposed a system, Manto, which detects Trojan communication with high accuracy using clustering technique. We implement Manto on real-world traces. The evaluation results exhibit that Manto is suitable for detecting Trojan communication amongst the vast amount of network traffic, with over 91% accuracy and less than 3.2% false positive ratio. We confidently regard our approach as a complementary way to the existing network-based techniques for we could address their main shortcomings. Shicong Li, Xiao-chun Yun, Yongzheng Zhang 0002, Yipeng Wang 0001 |
NAS | 5 |
| 2012 | Modeling Social Engineering Botnet Dynamics across Multiple Social Networks
Xiao-chun Yun, Zhiyu Hao, Yongzheng Zhang 0002, Xiang Cui, Yipeng Wang 0001 |
SEC | 6 |
| 2011 | Inferring Protocol State Machine from Network Traces: A Probabilistic Approach
Yipeng Wang 0001, Danfeng Yao, Buyun Qu, Li Guo 0001 |
ACNS | 1 |
| 2011 | Performance evaluation of Xunlei peer-to-peer network: A measurement studyabstractXunlei is a P2P file sharing application that is popular in China. The performance of previous P2P applications is limited by the selfishness of peers and the widely use of firewall. Xunlei applies implicit uploading strategy and firewall bypassing technologies to conquer above drawbacks. To evaluate the effects of above solutions, we perform a series of measurements on Xunlei and BitTorrent networks. Our study provides three new findings by comparing the situation in the two networks. (1)Compared with the free control strategy in BitTorrent, the implicit uploading strategy improves not only the number of active peers but also the seed ratio. Benefit from the strategy, Xunlei has much better downloading performance in our measurement. (2)The strategy extends the seed service time that is related to swarm lifespan. According to our result, Xunlei swarm lives longer than BitTorrent one. (3)The connectivity influences the performance greatly. The result of our measurement shows that only 5% peers can be connected directly. Our theoretical analysis and measurement results show that better connectivity also leads to higher downloading speed and longer lifespan. Yipeng Wang 0001, Li Guo 0001, Binxing Fang |
CCNC | 3 |
| 2011 | Using Entropy to Classify Traffic More DeeplyabstractThe network community always pays its attention to find better methods for traffic classification, which is crucial for Internet Service Providers (ISPs) to provide better QoS for users. Prior works on traffic classification mainly focus their attentions on dividing Internet traffic into different categories based on application layer protocols (such as HTTP, Bit Torrent etc.). Making traffic classification from another point of view, we divide Internet traffic into different content types. Our technology is an attempt to solve the classification problem of network traffic, which contains unknown and proprietary protocols (i.e., no publicly available protocol specification). In this paper, we design a classifier which can distinguish Internet traffic into different content types using machine learning techniques. Features of our classifier are entropy of consecutive bytes and frequencies of characters. Our method is capable of classifying real-world traces into different content types (including Text, Picture, Audio, Video, Compressed, Base 64-encoded image, Base 64-encoded text and Encrypted). The chief features of our classifier are small computing space (about 1K Bytes) and high classification accuracy (about 81%). Yipeng Wang 0001, Li Guo 0001 |
NAS | 1 |
| 2011 | A Propagation Model for Social Engineering Botnets in Social NetworksabstractWith the rapid development of social networking services and the diversification of social engineering attacks, new high-infection botnet (called SE-botnet by us), which exploits social engineering attacks to spread bots in social networks, has become an underlying threat. Predicting the threat of SE-botnet can help defenders mitigate it effectively. In this paper, we focus on SE-botnet's infection and defense, presenting a propagation model for it. We take full account of social networks' characteristics and human dynamics, and abstract the general process of social engineering attacks used by SE-botnet. Our preliminary simulation results demonstrate that the SE-botnet can capture tens of thousands of bots in one day with a great infection capacity. our propagation model can accurately predict this process with less than 5% deviation. Xiao-chun Yun, Zhiyu Hao, Xiang Cui, Yipeng Wang 0001 |
PDCAT | 5 |
| 2011 | Biprominer: Automatic Mining of Binary Protocol FeaturesabstractApplication-level protocol specifications are helpful for network security management, including intrusion detection and intrusion prevention which rely on monitoring technologies such as deep packet inspection. Moreover, detailed knowledge of protocol specifications is also an effective way of detecting malicious code. However, current methods for obtaining unknown and proprietary protocol message formats (i.e., no publicly available protocol specification), especially binary protocols, highly rely on manual operations, such as reverse engineering which is time-consuming and laborious. In this paper, we propose Biprominer, a tool that can automatically extract binary protocol message formats of an application from its real-world network trace. In addition, we present a transition probability model for a better description of the protocol. The chief feature of Biprominer is that it does not need to have any priori knowledge of protocol formats, because Biprominer is based on the statistical nature of the protocol format. We evaluate the efficacy of Biprominer over three binary protocols, with an average precision more than 99% and a recall better than 96.7%. Yipeng Wang 0001, Xingjian Li 0002, Jiao Meng, Li Guo 0001 |
PDCAT | 1 |
| 2010 | Inferring Protocol State Machine from Real-World Trace
Yipeng Wang 0001, Li Guo 0001 |
RAID | 1 |