VLDB 2026 Research / reviewers in the wild / expert
Xiao-chun Yun
dblp:21/3928 · also Xiaochun Yun
· DBLP profile ↗
83ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 22 · 4 first-author · 7 since 2021Computer networks · 15 · 2 first-author · 5 since 2021Systems, architecture and hardware · 13 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 13 · 1 since 2021Artificial intelligence and machine learning · 6Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CTISum: A new benchmark dataset for Cyber Threat Intelligence summarization
Wei Peng 0008, Junmei Ding, Wei Wang 0428, Lei Cui 0003, Zhiyu Hao, Xiao-chun Yun |
Comput. Secur. | 7 |
| 2025 | Pegasus: Accelerating Provenance Graph-based Intrusion Detection MethodsabstractProvenance-based Endpoint Detection and Response (P-EDR) systems are considered as the key to future Advanced Persistent Threat (APT) defense. Building provenance graphs that consider causal relationships between software behaviors can better provide contextual information of cyber attacks, which is capable of effectively reconstructing complex cyber attack scenarios represented by APT. Although promising to assist in attack investigation, existing methods for attack detection using provenance graphs adopt a centralized detection architecture, sending all system audit logs to servers for processing, resulting in unbearable costs in terms of data transmission, data storage, and computation. To address the above fundamental challenges, we propose Pegasus, a distributed detection system that can reduce memory consumption during training through a distributed system. Our system is evaluated on a large public dataset, and experimental results show that our system reduces memory consumption by 47%–65% compared with existing provenance-based EDR. And the above processing has little impact on attack detection performance, and our EDR system can still achieve sufficiently good detection results. Pengcheng Bi, Tianning Zang, Xiao-chun Yun |
CSCWD | 5 |
| 2025 | BTRFormer: Hierarchical Learning of Encrypted Traffic Using a Masked Autoencoder with Block-Based Traffic RepresentationabstractEncrypted traffic classification (ETC) is essential for ensuring network security and efficient management. Despite advances in deep learning, ETC remains challenging as existing models struggle to learn robust, discriminative representations from content-encrypted, highly imbalanced traffic.To address these challenges, we propose BTRFormer, a novel ETC approach that capitalizes on the inherent properties of encryption algorithms to enhance classification accuracy. At the core of BTRFormer lies a block-based, multi-layer traffic representation that adopts a 4×4 block as the fundamental unit, inspired by the encryption algorithm’s use of 16-byte blocks for encryption operations. This representation preserves the intrinsic structure of encrypted payloads, facilitating the model’s ability to learn deep semantic features. Subsequently, a transformer-based model is employed to learn from the multi-layer representation, capturing intra-block, inter-block, and inter-packet dependencies through block-wise attention mechanisms. Finally, BTRFormer leverages a pre-training phase on large-scale unlabeled data, followed by fine-tuning with a minimal amount of labeled samples to improve generalization and adaptability. Experimental results show that BTRFormer significantly outperforms SOTA methods on six real-world datasets, highlighting its effectiveness in encrypted traffic classification and secure network management. Junnan Yin, Lei Cui 0003, Zhiyu Hao, Peng Liu 0044, Xiao-chun Yun |
ICNP | 6 |
| 2025 | AdvTG: An Adversarial Traffic Generation Framework to Deceive DL-Based Malicious Traffic Detection ModelsabstractDeep learning-based (DL-based) malicious traffic detection models are effective but vulnerable to adversarial attacks. Existing adversarial attacks have shown promising results when targeting traffic detection models based on statistics and sequence features. However, these attacks are less effective against models that rely on payload analysis. The main reason is the difficulty in generating semantic, compliant, and functional payloads, which limits their practical application. Peishuai Sun, Xiao-chun Yun, Chengxiang Si, Jiang Xie 0004 |
WWW | 2 |
| 2025 | Sample analysis and multi-label classification for malicious sample datasets
Jiang Xie 0004, Xiao-chun Yun, Chengxiang Si |
Comput. Networks | 3 |
| 2025 | Bottom Aggregating, Top Separating: An Aggregator and Separator Network for Encrypted Traffic UnderstandingabstractEncrypted traffic classification refers to the task of identifying the application, service or malware associated with network traffic that is encrypted. Previous methods mainly have two weaknesses. Firstly, from the perspective of word-level (namely, byte-level) semantics, current methods use pre-training language models like BERT, learned general natural language knowledge, to directly process byte-based traffic data. However, understanding traffic data is different from understanding words in natural language, using BERT directly on traffic data could disrupt internal word sense information so as to affect the performance of classification. Secondly, from the perspective of packet-level semantics, current methods mostly implicitly classify traffic using abstractive semantic features learned at the top layer, without further explicitly separating the features into different space of categories, leading to poor feature discriminability. In this paper, we propose a simple but effective Aggregator and Separator Network (ASNet) for encrypted traffic understanding, which consists of two core modules. Specifically, a parameter-free word sense aggregator enables BERT to rapidly adapt to understanding traffic data and keeping the complete word sense without introducing additional model parameters. And a category-constrained semantics separator with task-aware prompts (as the stimulus) is introduced to explicitly conduct feature learning independently in semantic spaces of different categories. Experiments on five datasets across seven tasks demonstrate that our proposed model achieves the current state-of-the-art results without pre-training in both the public benchmark and real-world collected traffic dataset. Statistical analyses and visualization experiments also validate the interpretability of the core modules. Furthermore, what is important is that ASNet does not need pre-training, which dramatically reduces the cost of computing power and time. The model code and dataset will be released inhttps://github.com/pengwei-iie/ASNET. Wei Peng 0008, Lei Cui 0003, Wei Wang 0428, Xiaoyu Cui, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Traffic2Chain: Revealing Covert Multi-Step Attacks Through Unsupervised Traffic Behaviour CorrelationabstractWith the continuous development of network technology, covert multi-step attacks have become one of the significant attack methods. It is a multi-step attack with the intention of destroying the system- or data-privacy, such as network stealing. Current methods usually generate single-step alerts first and then perform correlation analysis. However, it is difficult for these methods to perform fine-grained annotation and alert amount control for single-step alerts, as well as to completely correlate the alerts of different phases into a chain due to alert fatigue. In this paper, we propose Traffic2Chain, an innovative unsupervised traffic behaviour correlation method to detect covert multi-step attacks from the network side. Traffic2Chain (1) generates alerts at different phases in real-time and annotates to sub-techniques based on MITRE ATT&CK knowledge database; (2) performs alert clustering based on SIMCSE and automatically generates event descriptions based on the Large Language Model (LLM) technique, and (3) extracts the attack chain through multi-dimensional information correlation to reveal the complete attack process. Experimental results demonstrate that the F1 score of Traffic2Chain reaches 98.36%, which has a significant advantage over other methods. In the real-world network, the detection speed can reach 40 Gbps. Most importantly, we discovered an unknown attack pattern based on Traffic2Chain - attackers delivered a variant of the Silver Fox Trojan by impersonating VPN services, eventually building a botnet with stealing capabilities and a node size of more than one million. Jiang Xie 0004, Xiao-chun Yun, Peishuai Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Digital Scapegoat: An Incentive Deception Model for Resisting Unknown APT Stealing Attacks on Critical Data ResourceabstractIt is a challenging problem to resist unknown advanced persistent threats (APTs) on stealing data resources in an information system of critical infrastructures, because APT attackers have very specific objectives and compromise the system stealthily and slowly. We observe that it is a necessary condition for APT attackers to achieve their campaigns via controlling unknown Trojans to access and exfiltrate critical files. We present a theoretical model called Digital Scapegoat (abbreviated as DS-IDep) that constructs an Incentive Deception defense schema to hijack the attacker’s access to critical files and redirect it to avatar files without awareness. We propose a FlipIDep Game model (GF) and a Markov Game model (GM) to characterize completely the payoffs, equilibria, and best strategies from the perspective of the attacker and the defender respectively. We also design an exponential risk propagation model to evaluate the ability of DS-IDep to eliminate stealing impact when the risk is propagated between states. Theoretically, we can achieve the objective of stealing impact elimination (LK0.7) and the probability of an attack operation bypassing the defense surface is less than 0.1 (r* × μ <0.1) under Stackelberg strategies. We develop a kernel-level incentive deception defense surface according to the theoretical parameters of the DS-IDep. The experimental results show that DS-IDep can resist APT stealing attacks from unknown Trojans. We also evaluate the DS-IDep in five well-known software applications. It demonstrates that DS-IDep can address unknown attacks from compromised software with less than 10% performance overhead. Xiao-chun Yun, Guangjun Wu, Qige Song, Zixian Tang, Zhenyu Cheng 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Sky-Eye: Detect Multi-stage Cyber Attacks at the Bigger Picture
Pengcheng Bi, Zhuohang Lv, Xiao-chun Yun, Tianning Zang |
ICDF2C (1) | 4 |
| 2024 | APIBeh: Learning Behavior Inclination of APIs for Malware ClassificationabstractMalware classification involves categorizing mal-ware samples based on their characteristics. While deep learning techniques applied to malware execution traces, mainly API calls, have shown potential in this field, they still perform poorly. This is primarily because they treat all APIs equally and train classifiers directly on native APIs, which inadequately capture the under-lying family-related semantics. In this paper, we first investigate the behaviors of multiple malware families and observe that different families exhibit divergent behaviors, with each family consistently favoring certain behaviors over time. Motivated by this, we propose APIBeh, a new embedding method designed to enhance malware classification. APIBeh first utilizes Benignity Degree Algorithm to identify and exclude insignificant, likely benign APIs from sequences. Then, it introduces the concept of Behavior Inclination, which quantifies the association between an API and malicious behaviors, facilitating high-level behavior encoding for each API. This Behavior Inclination embedding is then concatenated with raw embedding to represent an API, and fed into a DL model for classifier training. Experimental results show that APIBeh outperforms existing embedding methods in classification performance, e.g., 3.18% boost in weighted f1-score over a recent study using word2vec. In addition, it offers robustness to concept drift and adversarial attacks. Lei Cui 0003, Yiran Zhu, Junnan Yin, Zhiyu Hao, Wei Wang 0428, Peng Liu 0044, Xiao-chun Yun |
ISSRE | 8 |
| 2024 | API2Vec++: Boosting API Sequence Representation for Malware Detection and ClassificationabstractAnalyzing malware based on API call sequences is an effective approach, as these sequences reflect the dynamic execution behavior of malware. Recent advancements in deep learning have facilitated the application of these techniques to mine valuable information from API call sequences. However, these methods typically operate on raw sequences and may not effectively capture crucial information, especially in the case of multi-process malware, due to theAPI call interleaving problem. Furthermore, they often fail to capture contextual behaviors within or across processes, which is particularly important for identifying and classifying malicious activities. Motivated by this, we present API2Vec++, a graph-based API embedding method for malware detection and classification. First, we construct a graph model to represent the raw sequence. Specifically, we design the Temporal Process Graph (TPG) to model inter-process behaviors and the Temporal API Property Graph (TAPG) to model intra-process behaviors. Compared to our previous graph model, the TAPG model exposes operations with associated behaviors within the process through node properties and thus enhances detection and classification abilities. Using these graphs, we develop a heuristic random walk algorithm to generate numerous paths that can capture fine-grained malicious familial behavior. By pre-training these paths using the BERT model, we generate embeddings of paths and APIs, which can then be used for malware detection and classification. Experiments on a real-world malware dataset demonstrate that API2Vec++ outperforms state-of-the-art embedding methods and detection/classification methods in both accuracy and robustness, particularly for multi-process malware. Lei Cui 0003, Junnan Yin, Jiancong Cui, Yuede Ji, Peng Liu 0044, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Software Eng. | 7 |
| 2023 | Encrypted TLS Traffic Classification on Cloud PlatformsabstractNowadays, encryption technology has been widely used to protect user privacy. With the explosive growth of mobile Internet, encrypted TLS traffic rises sharply and occupies a great share of current Internet traffic. In reality, the classification of encrypted TLS traffic on cloud platforms brings a new challenge to traditional encrypted traffic classification methods, because some information such as certificates in the TLS flows is no longer effective. In this paper, we apply deep learning technology to the problem of encrypted TLS traffic classification on cloud platforms, and propose NeuTic, which takes the packet sequence of each TLS flow as the input, and effectively classifies raw TLS flows generated by many “cloud” applications. Our approach is able to automatically capture the long-range dependencies between elements in the packet sequences for robust and accurate encrypted TLS traffic classification. In NeuTic, we first convert each TLS flow into three attribute sequences. Then, we train a multi-application traffic classification model using our newly designed deep learning model. Finally, we use the well-trained classification model to classify new incoming TLS flows. We conduct comprehensive experiments on real-world application traces covering multiple “cloud” applications from three different companies. In addition, we compare our experimental results of NeuTic with two deep learning-based methods for encrypted traffic classification. NeuTic outperforms the state-of-the-art approaches in classification accuracy. Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002 |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Detecting unknown HTTP-based malicious communication behavior via generated adversarial flows and hierarchical traffic features
Xiao-chun Yun, Jiang Xie 0004, Yongzheng Zhang 0002, Peishuai Sun |
Comput. Secur. | 1 |
| 2022 | A Sketching Approach for Obtaining Real-Time Statistics Over Data Streams in CloudabstractMany applications of complex event processing (CEP) in Cloud can tolerate analytical errors to some extent, and it provides us an opportunity to optimize real-time analytics using methods of approximate query processing over big data streams. In this article, we present a novel rules-based sampling technique, which supports to construct sketch over one-pass and high-speed asynchronous data streams and provides accurate answers for different types of analytical queries. Moreover, we propose two methods of distributed sketching implementation, i.e., D-AQP$_b$and D-AQP$_i$, to make our approach to be compatible with batch processing and interactive processing architectures respectively, and be appropriate for stream processing systems in Cloud. Experimental results with real-world and synthetic datasets indicate that our approach can obtain more accurate estimates and improve two times of system throughput when compared with state-of-the-art Hadoop-based approximate engine BlinkDB. When compared with current batch processing systems Spark and stream processing system Spark-Streaming, our methods of D-AQP$_b$and D-AQP$_i$can achieve 2 and 4 orders of magnitude improvement on query response time respectively. Guangjun Wu, Xiao-chun Yun, Yong Wang 0032, Binbin Li 0001, Yong Liu 0018 |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | A Multi-Scale Feature Attention Approach to Network Traffic Classification and Its Model ExplanationabstractNetwork traffic classification, the task of associating network traffic with their generating application protocols or applications, is valuable for the control, allocation, and management of resources in today’s TCP/IP networks. In this paper, we propose Ulfar, a multi-scale feature attention approach to network traffic classification, which uses convolutional neural networks (CNN) as the building block of the deep packet analysis model. In Ulfar, we take only one packet per flow for network traffic classification. Ulfar is based on the key insight that format-related bytes appear at fixed offsets or in a specific pattern in the IP packet, and these format-related bytes are important for accurate network traffic classification. Our neural network model can automatically recover the format-related bytes by building high-level, multi-scale${n}$-gram features from raw byte sequences. In addition, at the representation learning side, we try to understand what patterns and signatures our neural network model learns from network traffic. We evaluate Ulfar using two publicly available datasets, and our experimental results show that Ulfar can conduct accurate network traffic classification. Also, we compare the results of Ulfar with four state-of-the-art approaches, and find that Ulfar has the ability to classify network traffic more accurately. Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Xin Liu 0002 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | iConSnap: An Incremental Continuous Snapshots System for Virtual MachinesabstractThe reliability of data and services hosted on a virtual machine (VM) is a top concern in cloud environments. The Continuous Snapshots can reduce the data loss in case of failures and thus is prevailing for protecting long-running systems. However, existing methods suffer from long VM downtime, long snapshot interval and significant performance loss. In this article, we present iConSnap, a system designed to take fine-grained continuous snapshots of virtual machines without compromising VM performance. First, iConSnap adopts the copy-on-write (COW) mechanism to save the memory pages on-demand, and thus decreases the VM downtime to about 200 milliseconds. Second, we extend the idea of COW and propose a lazily incremental approach to save the delta data between two successive snapshots only once, thereby reducing the snapshot duration and snapshot data a lot. Third, we propose a scheduling mechanism to mitigate the VM performance penalty issue. Last, we introduce a method combined of compression and time-aware multi-granularity reclamation strategy to reduce the storage costs without losing performance and availability. We implement iConSnap on QEMU/KVM and evaluate it through a set of experiments. The experimental results show that iConSnap outperforms existing approaches in terms of VM downtime, snapshot duration, storage costs and VM performance. Zhiyu Hao, Wei Wang 0428, Lei Cui 0003, Xiao-chun Yun, Zhenquan Ding |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Finding disposable domain names: A linguistics-based stacking approach
Yuwei Zeng, Xiao-chun Yun, Xunxun Chen, Boquan Li 0002, Haiwei Tsang, Yipeng Wang 0001, Tianning Zang, Yongzheng Zhang 0002 |
Comput. Networks | 2 |
| 2021 | VulDetector: Detecting Vulnerabilities Using Weighted Feature Graph ComparisonabstractCode similarity is one promising approach to detect vulnerabilities hidden in software programs. However, due to the complexity and diversity of source code, current methods suffer low accuracy, high false negative and poor performance, especially in analyzing a large program. In this paper, we propose to tackle these problems by presenting VulDetector, a static-analysis tool to detect C/C++ vulnerabilities based on graph comparison at the granularity of function. At the key of VulDetector is a weighted feature graph (WFG) model which characterizes function with a small yet semantically rich graph. It first pinpoints vulnerability-sensitive keywords to slice the control flow graph of a function, thereby reducing the graph size without compromising security-related semantics. Then, each sliced subgraph is characterized using WFG, which provides both syntactic and semantic features in varying degrees of security. As for graph comparison, we take full usage of vulnerability graph and patch graph to improve accuracy. In addition, we propose two optimization methods based on analysis of vulnerabilities. We have implemented VulDetector to automatically detect vulnerabilities in software programs with known vulnerabilities. The experimental results prove the effectiveness and efficiency of VulDetector. Lei Cui 0003, Zhiyu Hao, Haiqiang Fei, Xiao-chun Yun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Efficient Malware Originated Traffic Classification by Using Generative Adversarial NetworksabstractWith the booming of malware-based cyber-security incidents and the sophistication of attacks, previous detections based on malware sample analysis appear powerless due to time-consuming and labor-intensive analysis process. The existing detection methods based on traffic analysis rely heavily on the available traffic patterns, which hinder detecting the zero-day attacks caused by malware variants. In this paper, we propose an approach based on deep learning referred to as TrafficGAN, which analyzes (HTTP) traffic sessions to distinguish between malware-related and normal traffic. We first try to explore traffic patterns of malware variants by adding noise and category condition to the Generative Adversarial Networks (GAN), thus generating various similar but slightly different traffic. And then, we use discriminative model to seek the deviation between abnormal traffic and normal traffic by extracting the essential difference. Notablely, we increase the diversity of data by generating samples adversarially, which enhances the robustness of the system to detect zero-day attacks and highlights the lack of sensitive data in the security community. We conduct extensive experiments on the public dataset and our data collected for specific targets. The results demonstrate that our method achieves superior performance to other methods and protects specific targets from the susceptibility of malware. Yongzheng Zhang 0002, Xiao-chun Yun, Zhenyu Cheng 0001 |
ISCC | 4 |
| 2020 | HSTF-Model: An HTTP-based Trojan detection model via the Hierarchical Spatio-temporal Features of Traffics
Jiang Xie 0004, Xiao-chun Yun, Yongzheng Zhang 0002 |
Comput. Secur. | 3 |
| 2020 | Khaos: An Adversarial Neural Network DGA With High Anti-Detection AbilityabstractA botnet is a network of remote-controlled devices that are infected with malware controlled by botmasters in order to launch cyber attacks. To evade detection, the botmaster frequently changes the domain name of his Command and Control (C&C) server. Notice that most of these types of domain names are generated by domain generation algorithms (DGAs). In this paper, we propose Khaos, a novel DGA with high anti-detection ability based on neural language models and the Wasserstein Generative Adversarial Network (WGAN). The key insight of our research is that real domain names are composed of readable syllables and acronyms, and thus we can arrange syllables and acronyms using neural language models to mimic real domain names. In Khaos, we first find the most common n-grams in real domain names, then tokenize these domain names into n-grams, and finally synthesize new domain names after learning arrangements of n-grams from real domain names. We carry out experiments using a variety of state-of-the-art DGA detection approaches: the statistics-based, the distribution-based, the LSTM-based and the graph-based detection approach. Our experimental results show that the average distance for detecting Khaos under the distribution-based detection approach is 0.64, the AUCs of Khaos under the statistics-based and the LSTM-based detection approach are 0.76 and 0.57, respectively, and the precision of Khaos under the graph-based detection approach is 0.68. Our work proves that the existing detection approaches have big troubles in detecting Khaos, and Khaos has better anti-detection ability than state-of-the-art DGAs. In addition, we find that training the existing detection approach on a dataset including the domain names generated by Khaos can improve its detection ability. Xiao-chun Yun, Yipeng Wang 0001, Tianning Zang, Yuan Zhou 0008, Yongzheng Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Lock-Free Parallelization for Variance-Reduced Stochastic Gradient Descent on Streaming DataabstractStochastic Gradient Descent (SGD) is an iterative algorithm for fitting a model to the training dataset in machine learning problems. With low computation cost, SGD is especially suited for learning from large datasets. However, the variance of SGD tends to be high because it uses only a single data point to determine the update direction at each iteration of gradient descent, rather than all available training data points. Recent research has proposed variance-reduced variants of SGD by incorporating a correction term to approximate full-data gradients. However, it is difficult to parallelize such variants with high performance and accuracy, especially on streaming data. As parallelization is a crucial requirement for large-scale applications, this article focuses on the parallel setting in a multicore machine and presents LFS-STRSAGA, a lock-free approach to parallelizing variance-reduced SGD on streaming data. LFS-STRSAGA embraces a lock-free data structure to process the arrival of streaming data in parallel, and asynchronously maintains the essential information to approximate full-data gradients with low cost. Both our theoretical and empirical results show that LFS-STRSAGA matches the accuracy of the state-of-the-art variance-reduced SGD on streaming data under sparsity assumption (common in machine learning problems), and that LFS-STRSAGA reduces the model update time by over 98 percent. Yaqiong Peng, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Accelerating Real-Time Tracking Applications over Big Data Stream with Constrained Space
Guangjun Wu, Xiao-chun Yun, Ge Fu, Chao Li 0062, Yong Liu 0018, Binbin Li 0001, Yong Wang 0032 |
DASFAA (1) | 2 |
| 2019 | A Method Based on Hierarchical Spatiotemporal Features for Trojan Traffic DetectionabstractTrojans are one of the most threatening network attacks currently. HTTP-based Trojan, in particular, accounts for a considerable proportion of them. Moreover, as the network environment becomes more complex, HTTP-based Trojan is more concealed than others. At present, many intrusion detection systems (IDSs) are increasingly difficult to effectively detect such Trojan traffic due to the inherent shortcomings of the methods used and the backwardness of training data. Classical anomaly detection and traditional machine learning-based (TML-based) anomaly detection are highly dependent on expert knowledge to extract features artificially, which is difficult to implement in HTTP-based Trojan traffic detection. Deep learning-based (DL-based) anomaly detection has been locally applied to IDSs, but it cannot be transplanted to HTTP-based Trojan traffic detection directly. To solve this problem, in this paper, we propose a neural network detection model (HSTF-Model) based on hierarchical spatiotemporal features of traffic. Meanwhile, we combine deep learning algorithms with expert knowledge through feature encoders and statistical characteristics to improve the self-learning ability of the model. Experiments indicate that F1of HSTF-Model can reach 99.4% in real traffic. In addition, we present a dataset BTHT consisting of HTTP-based benign and Trojan traffic to facilitate related research in the field. Jiang Xie 0004, Yongzheng Zhang 0002, Xiao-chun Yun |
IPCCC | 4 |
| 2019 | A Method of HTTP Malicious Traffic Detection on Mobile NetworksabstractAiming at solving the problem of HTTP malicious traffic detection on mobile networks, we propose a method of HMTD(HTTP Malicious Traffic Detection) based on the spatiotemporal sequence characteristics of traffic data. The traditional malicious traffic detection methods are relatively simple and mainly biased towards misuse detection or abnormal detection and probably suffer from a high false positive rate or false negative rate, so they are difficult to adapt to the current rapid development of the Internet. HMTD uses neural networks for malicious traffic identification, and extracts features from malicious and normal HTTP traffic, which can produce excellent detection results. HMTD utilizes CNN to extract the packet spatial characteristics in the traffic, and utilizes LSTM to extract the temporal characteristics between the packets in the traffic. The experimental results demonstrate that the proposed method can achieve an accuracy of more than 99.4% in the actual network environment and has excellent performance in terms of Precision and Recall. Xiao-chun Yun, Mao Tian, Jiang Xie 0004, Yongzheng Zhang 0002, Yu Zhou 0028 |
WCNC | 2 |
| 2019 | Framework for risk assessment in cyber situational awarenessabstractA large number of data is generated to help network analysts to evaluate the network security situation in traditional detection and prevention measures, but it is not used fully and effectively, there is not a holistic view of the network situation on it for now. To address this issue, a framework is proposed to evaluate the security situation of the network from three dimensions: threat, vulnerability and stability, and merge the results at decision level to measure the security situation of the overall network. In the case studies, the authors demonstrate how the framework is deployed in the network and how to use it to reflect the security situation of the network in real time. Results of the case study show that the framework can evaluate the security situation of the network accurately and reasonably. Rongrong Xi, Xiao-chun Yun, Zhiyu Hao |
IET Inf. Secur. | 2 |
| 2019 | Fast Wait-Free Construction for Pool-Like Objects with Weakened Internal Order: Stacks as an ExampleabstractThis paper focuses on a large class of concurrent data structures that we call pool-like objects (e.g., stack, double-ended queue, and queue). Performance and progress guarantee are two important characteristics for concurrent data structures. In the aspect of performance, weakening the internal order in a pool-like object is an effective technique to reduce the synchronization cost among threads accessing the object, but no objects with weakened internal order provide a progress guarantee as strong as wait-freedom. Meanwhile, wait-free algorithms tend to be inefficient, which is mainly attributed to the helping mechanisms. Based on the philosophy of existing helping mechanisms, a wait-free pool-like object with weakened internal order would suffer from unnecessary process of getting the latest object state and synchronization. This paper takes a state-of-the-art implementation of stacks with weakened internal order as an example, and transforms it into a highly-efficient wait-free stack named WF-TS-Stack. The transformation method includes a helping mechanism with state reuse and a relaxed removal scheme. In addition, we use a simple and effective scheme to further improve the performance of WF-TS-Stack in Non-Uniform Memory Access (NUMA) architectures. Our evaluation with representative benchmarks shows that WF-TS-Stack outperforms its original building blocks by up to 1.45× at maximum concurrency. We also discuss how to yield an efficient double-ended queue (deque) variant of WF-TS-Stack, because deque is a more generalized pool-like object. Yaqiong Peng, Xiao-chun Yun, Zhiyu Hao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Important Member Discovery of Attribution Trace Based on Relevant Circle (Short Paper)
Jian Xu 0010, Xiao-chun Yun, Yongzheng Zhang 0002, Zhenyu Cheng 0001 |
CollaborateCom | 2 |
| 2018 | MalShoot: Shooting Malicious Domains Through Graph Embedding on Passive DNS Data
Chengwei Peng, Xiao-chun Yun, Yongzheng Zhang 0002 |
CollaborateCom | 2 |
| 2018 | MalHunter: Performing a Timely Detection on Malicious Domains via a Single DNS Query
Chengwei Peng, Xiao-chun Yun, Yongzheng Zhang 0002 |
ICICS | 2 |
| 2018 | Community Discovery of Attribution Trace Based on Deep Learning Approach
Jian Xu 0010, Xiao-chun Yun, Yongzheng Zhang 0002, Zhenyu Cheng 0001 |
ICICS | 2 |
| 2018 | SnapFiner: A Page-Aware Snapshot System for Virtual MachinesabstractVirtual machine (VM) snapshot, enabling a VM to be resumed from a previously recorded state, is an essential part of cloud infrastructures. Unfortunately, the snapshot data are likely to be lost due to the high rate of disk failures, so that the associated VM fails to recover properly. To enhance data availability without compromising application performance upon rollback recovery, it is desired to place multiple replicas of snapshot across disperse disks. However, due to the large size of replica, it induces non-trivial storage cost when managing massive snapshots in clouds. In this paper, we investigate this problem and find out that the semantic gap existed between snapshot creation and snapshot storing is one key factor inducing high storage cost. To this end, we propose SnapFiner, a page-aware snapshot system for creating and storing massive snapshot files efficiently. First, SnapFiner acquires a fine-grained page categorization with an in-depth page exploration from three orthogonal views, thereby discovering more pages that can be excluded from the snapshot. Second, SnapFiner varies the number of replicas for different page categories based on a page-aware replication policy, achieving low storage cost without compromising availability and performance. Third, SnapFiner handles the loss of pages either intentionally dropped upon snapshot creation or unexpectedly damaged due to disk failures, enabling proper system execution after rollback recovery. We have implemented SnapFiner on QEMU/KVM to justify its practicality for Linux guests. The experimental results demonstrate that SnapFiner reduces the storage cost by 33 and 69.5 percent respectively compared to our previous work PARS and the naive approach on QEMU/KVM and HDFS. Lei Cui 0003, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | NOR: Towards Non-intrusive, Real-Time and OS-agnostic Introspection for Virtual Machines in Cloud Environment
Chonghua Wang, Zhiyu Hao, Xiao-chun Yun |
Inscrypt | 3 |
| 2017 | Supporting Real-Time Analytic Queries in Big and Fast Data Environments
Guangjun Wu, Xiao-chun Yun, Chao Li 0062, Yipeng Wang 0001, Xiaoyu Zhang 0002, Siyu Jia, Guangyan Zhang |
DASFAA (2) | 2 |
| 2017 | Efficient Data Blocking and Skipping Framework Applying Heuristic RulesabstractData blocking has been an effective technique of data skipping to reduce data access and shorten query response time in query engines. By generating fine-grained, balanced blocks and corresponding metadata, a query may skip a block if the metadata indicates that the block does not contain relevant data. Obviously, the deciding factor of a promising blocking strategy depends on how to produce effective data layout in reasonable time that is expected to skip most data. In this paper, we propose several algorithms that drastically reduce the time complexity of existent blocking strategies based on workload analysis, at the cost of relatively small loss of estimated tuples could be skipped. Via theoretical analysis, we prove that the time complexity of our algorithms is apparently lower than that of ward algorithm. Afterwards, we demonstrate the whole blocking and skipping workflow, install it into Spark SQL and obtain experimental evaluation results. Experimental results show that our technique gains significant improvement in aspect of blocking efficiency compared to ward algorithm, while keeping almost the same level of skipping ability. Yong Wang 0032, Xiao-chun Yun, Yongshang Wu |
ICPADS | 2 |
| 2017 | Towards Robust and Accurate Similar Trajectory Discovery: Weak-Parametric ApproachesabstractTrajectory analysis is crucial and has been more and more widely used in various fields, such as location-based services (LBS), urban traffic control, user classification and route planner, etc. In this paper, we propose GSIM and ASIM, two novel approaches that are weak-parametric and can effectively measure and discover similar trajectories. The proposed methods are based on the key insight that the similarity can be reflected by observing the growth rate of specific indicators. (1) GSIM defines a 3-layer grid structure and statistics the total overlapping points for all grids between trajectories in each layer, it finally calculates the growth rate of the total counts as the grid radius grows from layer 1 to layer 3. (2) ASIM assumes that any two trajectories are similar and calculates the area of the minimum boundary rectangle that contains all the points. Then it cuts the rectangle from four directions one point by one to get the maximum boundary rectangle that contains the other two percentage of total points. Finally it utilizes the average change rate of the areas as the similarity. Further, we design parameter-learning modules to learn the setting of corresponding parameters automatically. Extensive experiments on real-world dataset show that, compared with typical approaches like LCSS, EDIT, DTW, etc., the proposed methods can significantly improve the effectiveness and achieve better efficiency in most test cases. Meanwhile, they are not sensitive to parameter settings. Yupeng Tuo, Xiao-chun Yun, Yongzheng Zhang 0002 |
NAS | 2 |
| 2017 | NSIM: A robust method to discover similar trajectories on cellular network location dataabstractTrajectory analysis is crucial and has been more and more widely used in various fields, such as location-based services, urban traffic control, route plan, etc. The existing methods have certain limitations when applied to cellular network location data. In this paper, we propose NSIM, a novel approach that can effectively discover similar trajectories. In NSIM, we first design an algorithm that can discover all the common moving patterns among trajectories, and then we adopt a vectorization method to abstract each trajectory as a summary vector that composed of specific common moving patterns. Finally we measure the similarity of trajectories by computing the distance between the summary vectors. Extensive experiments on real-world dataset show that, compared with three other approaches, NSIM achieves good effectiveness in most test cases and achieves better efficiency when applied to small or medium length trajectories. Yupeng Tuo, Xiao-chun Yun, Yongzheng Zhang 0002 |
PIMRC | 2 |
| 2017 | MSTM: A novel map matching approach for low-sampling-rate trajectoriesabstractMap matching is an important technique that matches user trajectories to the real road networks on a digital map. It is crucial and has been more and more widely used in various fields, such as route plan, traffic forecast, location-based services and so on. However, most existing algorithms are less effective when applied to low-sampling-rate trajectories. In this paper, we propose MSTM, a novel approach that can effectively match the low-sampling-rate trajectory to road networks. In MSTM, we first partition the trajectories into trajectory segments according to the stay points. Then we construct a map-searching tree by conditional extend and prune operations, which contains all the candidate paths. Finally, by considering the spatial and temporal information of trajectories, we evaluate each branch path in the map-searching tree and choose the one with the highest score as the result. Extensive experiments on real-world datasets show that, compared with two classic approaches, MSTM outperforms ST-Matching and IVMM in terms of matching accuracy as well as efficiency. Yupeng Tuo, Xiao-chun Yun, Yongzheng Zhang 0002 |
PIMRC | 2 |
| 2017 | A Hypervisor Level Provenance System to Reconstruct Attack Story Caused by Kernel Malware
Chonghua Wang, Shiqing Ma, Xiangyu Zhang 0001, Junghwan Rhee, Xiao-chun Yun, Zhiyu Hao |
SecureComm | 5 |
| 2017 | Rethinking robust and accurate application protocol identification
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Tianning Zang |
Comput. Networks | 2 |
| 2017 | A nonparametric approach to the automated protocol fingerprint inference
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Guangjun Wu |
J. Netw. Comput. Appl. | 2 |
| 2017 | Piccolo: A Fast and Efficient Rollback System for Virtual Machine ClustersabstractRollback is an effective technique to resume the system execution from a recorded intermediate state upon failures, without having to restart the entire system. However, in virtualized environments, rollback of a virtual machine cluster (VMC) produces high network traffic and long service disruption, particularly for a large cluster used for scientific computing, thereby imposing significant overhead both on network and applications. This paper proposes Piccolo, a fast and efficient rollback system, to restore a VMC from snapshot files over data center network. First, we exploit the similarity among VMC snapshots and leverage multicast to deliver the identical pages across VMs placed on disperse hosts, thereby bypassing unnecessary transmission of a large number of pages. Second, we analyze the impact on network traffic of varying VM placements in data center network, formulate the traffic aware placement as an optimization problem, and design a two-tier approximation algorithm that efficiently solves the problem. In addition to presenting Piccolo, we detail its implementation, and evaluate it by a set of experiments. The results show that Piccolo could achieve a significant reduction in terms of total sent data, network traffic and rollback latency compared to the existing generic techniques. Lei Cui 0003, Zhiyu Hao, Yaqiong Peng, Xiao-chun Yun |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Retweeting behavior prediction using probabilistic matrix factorizationabstractRetweeting is an important mechanism for information diffusion, popular event prediction, and so on. Due to the increasing requirements, in recent years, the task has attracted extensive attentions. In this paper, we propose a novel framework using probabilistic matrix factorization technique to predict retweeting behavior. Our study consists of three components. First, we convert retweeting behavior problem to a matrix factorization problem. Second, following the intuition that a user's social network will affect his retweeting behavior, we extensively study how to model social information to improve the prediction accuracy. Finally, message semantic embedding information is employed in designing a semantic regularization term to constrain the matrix factorization objective function. We also propose a set of metrics to construct the embeddings among messages based on messages' structural and textual features. The empirical results and analysis demonstrate that our methods perform better than the state-of-the-art approaches. Kai Zhang 0079, Xiao-chun Yun, Jiguang Liang, Xiaoyu Zhang 0002, Chao Li 0062 |
ISCC | 2 |
| 2016 | Weighted hierarchical geographic information description model for social relation estimation
Kai Zhang 0079, Xiao-chun Yun, Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Chao Li 0062 |
Neurocomputing | 2 |
| 2016 | Quantitative threat situation assessment based on alert verificationabstractAbstract Traditional network threat situational assessment is based on raw alerts, not combined with contextual information, which influences the accuracy of assessment. In this paper, we propose a method to quantitatively assess network threat situation based on not only alerts but also contextual information. It firstly verifies alerts by matching alerts with contextual information to determine the successful probability of attacks, then analyzes the impact caused by attacks according to the severity and the corresponding asset value of them, and finally quantitatively assesses network threat situation based on the successful probability and the impact of attacks. Case studies show that the method can assess network threat situations more reasonably. Copyright © 2016 John Wiley & Sons, Ltd. Rongrong Xi, Xiao-chun Yun, Zhiyu Hao, Yongzheng Zhang 0002 |
Secur. Commun. Networks | 2 |
| 2016 | A Semantics-Aware Approach to the Automated Network Protocol IdentificationabstractTraffic classification, a mapping of traffic to network applications, is important for a variety of networking and security issues, such as network measurement, network monitoring, as well as the detection of malware activities. In this paper, we propose Securitas, a network trace-based protocol identification system, which exploits the semantic information in protocol message formats. Securitas requires no prior knowledge of protocol specifications. Deeming a protocol as a language between two processes, our approach is based upon the new insight that the n-grams of protocol traces, just like those of natural languages, exhibit highly skewed frequency-rank distribution that can be leveraged in the context of protocol identification. In Securitas, we first extract the statistical protocol message formats by clustering n-grams with the same semantics, and then use the corresponding statistical formats to classify raw network traces. Our tool involves the following key features: 1) applicable to both connection oriented protocols and connection less protocols; 2) suitable for both text and binary protocols; 3) no need to assemble IP packets into TCP or UDP flows; and 4) effective for both long-live flows and short-live flows. We implement Securitas and conduct extensive evaluations on real-world network traces containing both textual and binary protocols. Our experimental results on BitTorrent, CIFS/SMB, DNS, FTP, PPLIVE, SIP, and SMTP traces show that Securitas has the ability to accurately identify the network traces of the target application protocol with an average recall of about 97.4% and an average precision of about 98.4%. Our experimental results prove Securitas is a robust system, and meanwhile displaying a competitive performance in practice. Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002, Yu Zhou 0015 |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | Exploring Efficient and Robust Virtual Machine Introspection Techniques
Chonghua Wang, Xiao-chun Yun, Zhiyu Hao, Lei Cui 0003, Yandong Han, Qingxin Zou |
ICA3PP (3) | 2 |
| 2015 | Rethinking Robust and Accurate Application Protocol Identification: A Nonparametric ApproachabstractProtocol traffic analysis is important for a variety of networking and security infrastructures, such as intrusion detection and prevention systems, network management systems, and protocol specification parsers. In this paper, we propose ProHacker, a nonparametric approach that extracts robust and accurate protocol keywords from network traces and effectively identifies the protocol trace from mixed Internet traffic. ProHacker is based on the key insight that the n-grams of protocol traces have highly predictable statistical nature that can be effectively captured by statistical language models and leveraged for robust and accurate protocol identification. In ProHacker, we first extract protocol keywords using a nonparametric Bayesian statistical model, and then use the corresponding protocol keywords to classify protocol traces by a semi-supervised learning algorithm. We implement and evaluate ProHacker on real-world traces, including SMTP, FTP, PPLive, SopCast, and PPStream, and our experimental results show that ProHacker can accurately identify the protocol trace with an average precision of about 99.42% and an average recall of about 98.64%. We also compare the results of ProHacker to two state-of-the-art approaches ProWord and Securitas using backbone traffic. We show that ProHacker provides significant improvements on precision and recall for online protocol identification. Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002 |
ICNP | 2 |
| 2015 | Update vs. upgrade: Modeling with indeterminate multi-class active learning
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Xiao-chun Yun, Guangjun Wu, Yipeng Wang 0001 |
Neurocomputing | 4 |
| 2015 | FastRAQ: A Fast Approach to Range-Aggregate Queries in Big Data EnvironmentsabstractRange-aggregate queries are to apply a certain aggregate function on all tuples within given query ranges. Existing approaches to range-aggregate queries are insufficient to quickly provide accurate results in big data environments. In this paper, we propose FastRAQ-a fast approach to range-aggregate queries in big data environments. FastRAQ first divides big data into independent partitions with a balanced partitioning algorithm, and then generates a local estimation sketch for each partition. When a range-aggregate query request arrives, FastRAQ obtains the result directly by summarizing local estimates from all partitions. FastRAQ has O(1) time complexity for data updates and O(N/P×B) time complexity for range-aggregate queries, where N is the number of distinct tuples for all dimensions, P is the partition number, and B is the bucket number in the histogram. We implement the FastRAQ approach on the Linux platform, and evaluate its performance with about 10 billions data records. Experimental results demonstrate that FastRAQ provides range-aggregate query results within a time period two orders of magnitude lower than that of Hive, while the relative error is less than 3 percent within the given confidence interval. Xiao-chun Yun, Guangjun Wu, Guangyan Zhang, Keqin Li 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2015 | SMS Worm Propagation Over Contact Social Networks: Modeling and ValidationabstractNowadays, short message service (SMS) worms have been discovered to propagate themselves via victims' contact lists by sending malicious text messages. Correspondingly, defenders need to analyze and model the dynamics of these worms to lessen their potential threat. However, the existing worm propagation models, which almost generate the similar curves of an exponential smooth rise, cannot well explain the infection dynamics of real-world SMS worms, which exhibits an uneven wave-like uplift. Motivated by this observation, we formalize the general infection process of SMS worms in contact social networks, and propose a novel analytical model based on stochastic processes. In contrast to previous models, our model not only considers the different and asymmetrical relationships between mobile users by modeling the node reputation and the edge trust degree, but also describes the user behavior of checking messages by introducing two susceptible states. Moreover, the strong assumptions in previous works are eliminated by determining related components based on extensive statistical investigations. Afterward, both real-world SMS worms and artificial ones are utilized in validation and comparison experiments, and the results show that our model is more suitable for describing the propagation of these sophisticated worms, compared with the state-of-the-art models. In addition, we study on the impacts of key factors and give some interesting discoveries. Xiao-chun Yun, Yongzheng Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2015 | Bidirectional Active Learning: A Two-Way Exploration Into Unlabeled and Labeled Data SetabstractIn practical machine learning applications, human instruction is indispensable for model construction. To utilize the precious labeling effort effectively, active learning queries the user with selective sampling in an interactive way. Traditional active learning techniques merely focus on the unlabeled data set under a unidirectional exploration framework and suffer from model deterioration in the presence of noise. To address this problem, this paper proposes a novel bidirectional active learning algorithm that explores into both unlabeled and labeled data sets simultaneously in a two-way process. For the acquisition of new knowledge, forward learning queries the most informative instances from unlabeled data set. For the introspection of learned knowledge, backward learning detects the most suspiciously unreliable instances within the labeled data set. Under the two-way exploration framework, the generalization ability of the learning model can be greatly improved, which is demonstrated by the encouraging experimental results. Xiaoyu Zhang 0002, Xiao-chun Yun |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | MMD: An Approach to Improve Reading Performance in Deduplication SystemsabstractThe approach of data deduplication has been widely used in backup systems and primary storage such as virtual machine platform. However, the reading speed in those systems suffers due to chunk fragmentation in deduplication. So it has become an important problem to improve reading performance in deduplication systems. In this paper, firstly we propose a new storage method using multiple disks to boost reading performance, which is called MMD. MMD takes advantage of the multiple parallelized disks, each of which is used as independent logical device. Then we present a deduplication model based on MMD, which focuses on optimization of data layout on disks to improve reading speed. Two I/O scheduling algorithms in that model are discussed, which aim at assigning the containers in deduplication systems to appropriate disks. Experiments show that MMD can achieve an obvious reading performance improvement than RAID in deduplication systems. Chao Li 0062, Xiao-chun Yun, Guangjun Wu |
NAS | 3 |
| 2013 | Counting sort for the live migration of virtual machinesabstractThe live migration of virtual machines is an important technique in the area of virtualization, and it has been used for load balancing, fault tolerance, and system maintenance in modern data centers, clusters, and cloud computing. The pre-copy algorithm is the most used method for the live migration of virtual machines. However, the existing problem of repeatedly transferring dirty memory pages leads the increase of the transferring data amount, delays of the total migration time as well as the downtime. By analyzing the iteration process of the pre-copy algorithm, we find that the transferring order of memory pages during every middle round has a huge impact on the generation and transferring of dirty memory pages. Further we put forward the concept of the live migration of virtual machines based on the counting sort. During every middle round of the iteration process, we do not transfer the memory pages according to their original order, instead we transfer the memory pages according to their times of being dirty. Experiment results show that with different workloads the counting sort method could simultaneously decrease the transferring data amount, the total migration time, and the downtime to improve the performance of the live migration. Qingxin Zou, Zhiyu Hao, Xiao-chun Yun, Yongzheng Zhang 0002 |
CLUSTER | 4 |
| 2012 | A semantics aware approach to automated reverse engineering unknown protocolsabstractExtracting the protocol message format specifications of unknown applications from network traces is important for a variety of applications such as application protocol parsing, vulnerability discovery, and system integration. In this paper, we propose ProDecoder, a network trace based protocol message format inference system that exploits the semantics of protocol messages without the executable code of application protocols. ProDecoder is based on the key insight that the n-grams of protocol traces exhibit highly skewed frequency distribution that can be leveraged for accurate protocol message format inference. In ProDecoder, we first discover the latent relationship among n-grams by first grouping protocol messages with the same semantics and then inferring message formats by keyword based clustering and cluster sequence alignment. We implemented and evaluated ProDecoder to infer message format specifications of SMB (a binary protocol) and SMTP (a textual protocol). Our experimental results show that ProDecoder accurately parses and infers SMB protocol with 100% precision and recall. For SMTP, ProDecoder achieves approximately 95% precision and recall. Yipeng Wang 0001, Xiao-chun Yun, Zubair Shafiq, Alex X. Liu, Danfeng Yao, Yongzheng Zhang 0002, Li Guo 0001 |
ICNP | 2 |
| 2012 | A General Framework of Trojan Communication Detection Based on Network TracesabstractBecause of the widespread Trojan, Internet users become more and more vulnerable to the threat of information leakage. Traditional techniques of Trojan detection were classified into two main categories: host-based and network-based. Unfortunately, existing techniques are insufficient and limited, because of the following reasons: (1)only uncover the known Trojan while inefficiently detecting novel samples, (2) should be adjusted in a timely fashion even a trivial change is applied, and (3)become computationally more expensive. In our work, we focus on a network behavior based method to address the limitations of previous network-based approaches. We analyze the profile of network behavior at two levels: (i)flow-level, (ii)IP-level. Our approach present two main advantages: (1)capture more detailed information to describe the network behavior profile, (2)consume lower computational overhead. We proposed a system, Manto, which detects Trojan communication with high accuracy using clustering technique. We implement Manto on real-world traces. The evaluation results exhibit that Manto is suitable for detecting Trojan communication amongst the vast amount of network traffic, with over 91% accuracy and less than 3.2% false positive ratio. We confidently regard our approach as a complementary way to the existing network-based techniques for we could address their main shortcomings. Shicong Li, Xiao-chun Yun, Yongzheng Zhang 0002, Yipeng Wang 0001 |
NAS | 2 |
| 2012 | Online Traffic Classification Based on Co-training MethodabstractOnline traffic classification has been widely used in quality of service measurements, network management and security monitoring. Currently, more and more research works tend to apply machine learning techniques to online traffic classification, and most of them are based on supervised learning and unsupervised learning techniques. Although supervised learning method has exhibited good classification performance, it needs lots of labeled training samples which are difficult to collect. The co-training method is a semi-supervised learning method, which can use little labeled samples and plenty of unlabeled samples to enhance the performance of supervised learning method. In this paper, we investigate the co-training algorithm for online traffic classification. The co-training algorithm needs two separate features which are sufficient to train a good classifier. We choose packet size and inter-packet time of the first packets of a traffic flow as two features. However, the inter-packet time is dependent to network conditions and will be impacted by network jitter. This paper constructs a robust inter-packet time feature named "Netipt" which is resilient to network jitter, and we integrate Netipt feature to co-training algorithm. We test our co-training algorithm based on two real-world traffic datasets. The results show that the co-training algorithm can enhance the accuracy of traffic classification drastically even when there are very few training samples. Jinghua Yan, Xiao-chun Yun, Zhi-Gang Wu, Hao Luo 0010, Shuzhuang Zhang, Shuyuan Jin |
PDCAT | 2 |
| 2012 | Modeling Social Engineering Botnet Dynamics across Multiple Social Networks
Xiao-chun Yun, Zhiyu Hao, Yongzheng Zhang 0002, Xiang Cui, Yipeng Wang 0001 |
SEC | 2 |
| 2011 | A Propagation Model for Social Engineering Botnets in Social NetworksabstractWith the rapid development of social networking services and the diversification of social engineering attacks, new high-infection botnet (called SE-botnet by us), which exploits social engineering attacks to spread bots in social networks, has become an underlying threat. Predicting the threat of SE-botnet can help defenders mitigate it effectively. In this paper, we focus on SE-botnet's infection and defense, presenting a propagation model for it. We take full account of social networks' characteristics and human dynamics, and abstract the general process of social engineering attacks used by SE-botnet. Our preliminary simulation results demonstrate that the SE-botnet can capture tens of thousands of bots in one day with a great infection capacity. our propagation model can accurately predict this process with less than 5% deviation. Xiao-chun Yun, Zhiyu Hao, Xiang Cui, Yipeng Wang 0001 |
PDCAT | 2 |
| 2011 | Network Threat Assessment Based on Alert VerificationabstractIn face of overwhelming alerts produced by firewalls or intrusion detection devices, it is difficult to assess network threats that we face. In this paper, we propose a threat assessment approach to estimate the impact of attacks on network. The approach employs the Common Vulnerability Scoring System to quantitatively assess network threats and further correlates alerts with contextual information to improve the accuracy of assessment. In the case studies, we demonstrate how the approach is applied in real networks. The experimental results show that the approach can make an accurate assessment of network threats. Rongrong Xi, Xiao-chun Yun, Shuyuan Jin, Yongzheng Zhang 0002 |
PDCAT | 2 |
| 2011 | CNSSA: A Comprehensive Network Security Situation Awareness SystemabstractWith tremendous attacks in the Internet, there is a high demand for network analysts to know about the situations of network security effectively. Traditional network security tools lack the capability of analyzing and assessing network security situations comprehensively. In this paper, we introduce a novel network situation awareness tool CNSSA (Comprehensive Network Security Situation Awareness) to perceive network security situations comprehensively. Based on the fusion of network information, CNSSA makes a quantitative assessment on the situations of network security. It visualizes the situations of network security in its multiple and various views, so that network analysts can know about the situations of network security easily and comprehensively. The case studies demonstrate how CNSSA can be deployed into a real network and how CNSSA can effectively comprehend the situation changes of network security in real time. Rongrong Xi, Shuyuan Jin, Xiao-chun Yun, Yongzheng Zhang 0002 |
TrustCom | 3 |
| 2011 | Graph-based multi-space semantic correlation propagation for video retrieval
Bailan Feng, Juan Cao 0001, Xiuguo Bao, Yongdong Zhang 0001, Shouxun Lin, Xiao-chun Yun |
Vis. Comput. | 7 |
| 2010 | A Pseudo-Random Number Generator Based on LZSSabstractA pseudo-random sequence generator (PRNG), L12RC4, inspired by the LZSS compression algorithm and RC4 stream cipher, was presented and implemented. The result of the NIST and Diehard test suite indicate that the L12RC4 is a good PRNG, and so it seems to be sound and may be suitable for use in some cryptographic applications. We also found that the probability distribution of the index value frequency is associated with the compression pass and INDEX_BIT_COUNT value. As for one pass mode, the greater INDEX_BIT_COUNT value, the more uniformly distributed, and the double pass mode has better uniformity than the one pass mode. Wei-ling Chang, Binxing Fang, Xiao-chun Yun, Xiangzhan Yu |
DCC | 3 |
| 2010 | Cooperative Work Systems for the Security of Digital Computing InfrastructureabstractOn open digital computing infrastructure, various large-scale and complicated malicious behaviors are increasingly threatening the security of digital computing infrastructure. In this paper, a Cooperative Work Model (CRM) is presented by extending the conceptions of the Universal Turing Machine to deal with the threats. Then the Cooperative Work System Framework (CWSF) is derived from the model. Based on the framework, two practical Cooperative Work Systems (CWSs) are developed to track and analyze the Botnet and DDoS on digital computing infrastructure respectively. The systems collectively use and coordinate various monitoring systems distributed in the back-bone network of the infrastructure. The experimental results of analyzing typical security events show that the framework and systems are efficient and effective to collaboratively use diverse related network systems for monitoring and analyzing the large-scale network events. Currently, the systems are running steadily in the monitoring environment of a large-scale back-bone network. Tianning Zang, Xiao-chun Yun, Tianyi Zang, Yongzheng Zhang 0002, Chaoguang Men |
ICPADS | 2 |
| 2010 | Adapting information bottleneck method for automatic construction of domain-oriented sentiment lexiconabstractDomain-oriented sentiment lexicons are widely used for fine-grained sentiment analysis on reviews; therefore, the automatic construction of domain-oriented sentiment lexicon is a fundamental and important task for sentiment analysis research. Most of existing construction approaches take only the kind of relationships between words into account, which makes them have a lot of room for improvement. This paper proposes an adapted information bottleneck method for the construction of domain-oriented sentiment lexicon. This approach can naturally make full use of the mutual reinforcement between documents and words by fusing three kinds of relationships either from words to documents or from words to words; either homogeneous or heterogeneous; either within-domain or cross-domain. The experimental results demonstrate that proposed method could dramatically improve the accuracy of the baseline approach on the construction of out-of-domain sentiment lexicon. Weifu Du, Songbo Tan, Xueqi Cheng 0001, Xiao-chun Yun |
WSDM | 4 |
| 2009 | The Block LZSS Compression AlgorithmabstractIn this paper, we studied the block LZSS algorithm and investigated the relationship between the compression ratio of block LZSS and the value of index or length. We found that as the block size increases, the compression ratio becomes better. We also found that the bit of length has little effect on the compression performance, and the bit of index has a significant effect on the compression ratio. We showed that the more the bit of index is set, the bigger optimal block size is obtained. Wei-ling Chang, Xiao-chun Yun, Binxing Fang |
DCC | 2 |
| 2009 | A review of classification methods for network vulnerabilityabstractClassification of network vulnerability is critical to detection and risk analysis of network vulnerability. A broad range of classification methods have been proposed in literature. This paper reviews a total of 25 selected approaches and identifies the differences and relations among them. It also points out some open issues for research in this field. Shuyuan Jin, Yong Wang 0032, Xiang Cui, Xiao-chun Yun |
SMC | 4 |
| 2008 | A Quasi Word-Based Compression Method of English Text Using Byte-Oriented Coding SchemeabstractIn this paper we present a universal compression algorithm for English text, ERecode. The proposed scheme highlights the importance of pre-processing work for English text, and employs one or two bytes code values to recode the 511 most common used English words, sequences of symbols and ASCII codes based on their occurrence frequency. Acting as a pre-processing tool for English text by the popular compression utilities, ERecode can improve their compression ratio from 0.89% to 19.65%. The proposed method also is applicable to text files for other languages. Wei-ling Chang, Xiao-chun Yun, Binxing Fang |
WAIM | 2 |
| 2008 | Optimizing Traffic Classification Using Hybrid Feature SelectionabstractThe identification of network applications is of fundamental important to numerous network activities. Unfortunately, traditional port-based classification and packet payload-based analysis exhibit a number of shortfalls. A promising alternative is to use Machine Learning (ML) techniques and identify network applications based on per-flow features. Since a lot of flow features can be used for flow classification, the flow classifier may deal with huge amount of data, which contains irrelevant and redundant features causing slower training and testing process, higher resource consumption as well as poor classification accuracy. Therefore, feature selection plays a vital role in performance optimizing. In this paper, we propose a hybrid feature selection method for flow classification using Chi-Squared and C4.5 algorithm (ChiSquared-C4.5). The experiments demonstrate our approach can greatly improve computational performance without negative impact on classification accuracy. Xiao-chun Yun |
WAIM | 2 |
| 2008 | Design and Implementation of Multi-Version Disk Backup Data Merging AlgorithmabstractMulti-version data management in disk backup and recovery is to manage the temporal attribute of backuped data. It can support to retrieve timestamp (time slice) disk data according to different query type. Exiting multi-version data management algorithms have two shortcomings. First, they are inefficient in multi-time point data query and updating which are adopted by data backup and recovery usually. Second, they use centralized data indexes which are not suitable for backup data management. To overcome these limitations, Backup Data Merging (BDM) algorithm is proposed in this paper, which uses distributed storage structure according to disk data format. By range operation, BDM algorithm can generate timestamp (time slice) data index dynamically. By comparing with traditional algorithms, BDM algorithm achieves high performance in storage utilization and query efficiency. Guangjun Wu, Xiao-chun Yun |
WAIM | 2 |
| 2008 | A Survey of Alert Fusion Techniques for Security IncidentabstractSecurity incident have been imposing tremendous threats on todaypsilas network information system. To protect this information system from the increasing threat of intrusion, various kinds of detection systems and sensors for security incident have been developed. The main disadvantages of current systems and sensors are a high false detection rate and the lack of post-incident decision support capability. To minimize these drawbacks, various alert fusion technologies have been proposed in the recent years. This paper presents a general summary of these technologies. Basic models and key technologies of alert fusion are analyzed and discussed. Moreover, important aggregation and correlation algorithms are discussed. Finally, we make concluding remarks by predicting the development tendencies of alert correlation technologies. Tianning Zang, Xiao-chun Yun, Yongzheng Zhang 0002 |
WAIM | 2 |
| 2007 | Optimizing IP Flow Classification Using Feature SelectionabstractThe identification of network applications is essential to numerous network activities. Unfortunately, traditional port-based classification and packet payload-based analysis exhibit a number of shortfalls. An alternative is to use Machine Learning (ML) techniques and identify network applications based on per-flow features. Since a lot of flow features can be used for flow classification and there are many irrelevant and redundant features among them, feature selection plays a vital role in performance optimizing. In this paper, we propose a wrapper-based feature selection method for IP flow classification using modified random-mutation hill-climbing (RMHC) and C4.5 algorithm (MRMHC-C4.5). The experiments show our approach can greatly improve computational performance without negative impact on classification accuracy. Xiao-chun Yun |
PDCAT | 3 |
| 2007 | Handover Cost Optimization in Traffic Management for Multi-homed Mobile Networks
Jianping Wang 0001, Mei Yang 0001, Xiao-chun Yun, Yingtao Jiang |
UIC | 4 |
| 2007 | How to construct secure proxy cryptosystem
Yuan Zhou 0008, Binxing Fang, Zhenfu Cao, Xiao-chun Yun, Xiaoming Cheng |
Inf. Sci. | 4 |
| 2006 | A User Habit Based Approach to Detect and Quarantine WormsabstractIn the long term usage of the network, users will form certain types of habit according to their specific characteristics, individual hobbies and given restrictions. On the burst-out of worms, the overwhelming flow caused by worm's scanning will temporarily alter the behavior representation of users. Therefore, it is reasonable to conclude that the statistics and classification of user habits can contribute to the detection of worms. We observe that number of destinations accessed in a long time range by a user is approximately limited, it means possible to record the access habit of users. Based on the analysis of both users and worms, we construct the patterns of user-habit and propose a new approach for the early warning of worms. And a better quarantine strategy is proposed to insure the normal access of user. Binxing Fang, Xiao-chun Yun |
ICC | 3 |
| 2006 | Adaptive Method for Monitoring Network and Early Detection of Internet Worms
Chen Bo, Binxing Fang, Xiao-chun Yun |
ISI | 3 |
| 2006 | A Counting-Based Method for Massive Spam Mail Classification
Hao Luo 0010, Binxing Fang, Xiao-chun Yun |
ISPEC | 3 |
| 2006 | Model and Estimation of Worm Propagation Under Network Partition
Binxing Fang, Xiao-chun Yun |
ISPEC | 3 |
| 2005 | Using Boosting Learning Method for Intrusion Detection
Wu Yang 0001, Xiao-chun Yun, Yongtian Yang |
ADMA | 2 |
| 2005 | Computer Vulnerability Evaluation Using Fault Tree Analysis
Mingzeng Hu, Xiao-chun Yun, Yongzheng Zhang 0002 |
ISPEC | 3 |
| 2005 | A New Approach to Automatically Detect WormsabstractWorms have seriously harmed computer and network systems due to their rapid spread rate. Therefore, it is necessary to research automatic worm detection systems in large networks. In this paper, data stream based anomaly detection is used to screen out anomalous network data flow, subsequently, the signature is extracted. After analyzed, the signature is updated to the misuse detection pattern. Based on an automatic worm defense, a system could discover an epidemic situation effectively and detect an unknown worm. Binxing Fang, Xiao-chun Yun |
PDCAT | 3 |
| 2005 | Worm Detection in Large Scale Network by TrafficabstractNowadays, worms have been one of the leading threats to information security and service availability. Current operational practices have not been able to manage the threat effectively. So it is very important to make early warning of the burst of worm in large scale network. In this paper we analyze the real network traffic in large scale network. Based on long time statistic, we construct a network traffic model which concern two parameters: the traffic volume and curve of traffic function. And then we propose a method to computer the function curve of normal traffic function in ideal condition. We deployed them in our campus network (more than 20000 computers, 400M/s bandwidth to internet).It is shown that the worms are detected automatically and efficiently. Binxing Fang, Xiao-chun Yun, Hai-Yong Chen |
PDCAT | 3 |
| 2004 | Defending Against Flash Crowds and Malicious Traffic Attacks with An Auction-Based MethodabstractFlash crowd events (FCEs) and malicious traffic including DDoS and worm attacks present a real threat to the stability of Web services. In this paper, we design a practical defense system that can provide some needed relief from the two types of events and protect the availability of Web services. A novel method of dynamic bandwidth arbitration using Generalized Vickrey auction based on microeconomics is proposed. By adopting this approach, not only the availability of Web services is improved but also the total utility of users can be maximized. Initial simulations have shown that this mechanism is promising direction to control both FCEs and malicious traffic. The presentation in this paper is a first step towards a more rigorous evaluation. Zhihong Tian 0001, Binxing Fang, Xiao-chun Yun |
Web Intelligence | 3 |