Yu Wang 0017

dblp:02/5889-17 · DBLP profile ↗
← Back
62ranked-venue papers
12as first author
15since 2021 · last 2025
0000-0002-9807-2293ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 23 · 1 first-author · 4 since 2021Systems, architecture and hardware · 16 · 4 first-author · 2 since 2021Computer networks · 11 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Consistency Regularization Semisupervised Learning for PolSAR Image Classification
abstract
Polarimetric Synthetic Aperture Radar (PolSAR) images have emerged as an important data source for land cover classification research due to their all‐weather, all‐day monitoring capabilities. Deep learning‐based classification methods have recently gained significant attention in PolSAR image classification since they have demonstrated excellent performance in the computer vision field. However, the main issue with deep learning‐based methods is that they require large amounts of training data. Additionally, the scarcity of labeled data is a significant challenge in the PolSAR image field. Therefore, in this article, we proposed an advanced semisupervised deep self‐training algorithm for PolSAR image classification, which utilized both labeled and unlabeled data in a semisupervised way. Then, a training optimization method and a high‐confidence sample selection strategy are proposed by integrating consistency regularization. In addition, to achieve stronger feature extraction capabilities, we designed a deep learning‐based classifier that combines residual blocks with an efficient multiscale attention module. We have conducted experiments on three popular real PolSAR datasets: 1989 Flevoland, 1991 Flevoland, and Oberpfaffenhofen. The classification results on these datasets demonstrated that the proposed method outperforms several other comparison algorithms, with overall accuracy up to 99.3%, 99.15%, and 94.12%, respectively. These results demonstrated the effectiveness of the proposed method for PolSAR image classification.
Yu Wang 0017, Shan Jiang 0023
Int. J. Intell. Syst.1
2024 Towards Alarm Reduction in Intrusion Detection: A Recurrent Neural Network Approach
abstract
Intrusion detection systems (IDSs) are one of the most commonly deployed systems to detect cyber attacks. One major problem of IDSs is the large amount of false alarms (or false positives) generated in a practical network. In recent decades, many researchers utilized machine learning techniques to help reduce false alarms, i.e., deep learning is quite popular in recent years. However, most research studies mainly consider one dataset in the evaluation, and it is unclear whether the approach can be scalable for other network settings. Motivated by such challenges, in this work, we introduce a hybrid deep learning-based false alarm filtration method, by combining a modern recurrent neural network (RNN) and a concise alert correlation method to reduce false positives for an IDS. To evaluate the performance, we consider three publicly available datasets such as DARPA, UNSW-NB and CSE-CIC-IDS. Experimental results demonstrate that our approach can reduce up to 99.35% false alarms in the best case.
Alexander Gödeke, Weizhi Meng 0001, Yu Wang 0017
HPCC3
2024 An eID-Based Privacy-Enhanced Public Transportation Ticket System
Kanagaratnam Anojjan, Weizhi Meng 0001, Brooke Kidmose, Yu Wang 0017
ISPEC4
2024 ARDST: An Adversarial-Resilient Deep Symbolic Tree for Adversarial Learning
abstract
The advancement of intelligent systems, particularly in domains such as natural language processing and autonomous driving, has been primarily driven by deep neural networks (DNNs). However, these systems exhibit vulnerability to adversarial attacks that can be both subtle and imperceptible to humans, resulting in arbitrary and erroneous decisions. This susceptibility arises from the hierarchical layer‐by‐layer learning structure of DNNs, where small distortions can be exponentially amplified. While several defense methods have been proposed, they often necessitate prior knowledge of adversarial attacks to design specific defense strategies. This requirement is often unfeasible in real‐world attack scenarios. In this paper, we introduce a novel learning model, termed “immune” learning, known as adversarial‐resilient deep symbolic tree (ARDST), from a neurosymbolic perspective. The ARDST model is semiparametric and takes the form of a tree, with logic operators serving as nodes and learned parameters as weights of edges. This model provides a transparent reasoning path for decision‐making, offering fine granularity, and has the capacity to withstand various types of adversarial attacks, all while maintaining a significantly smaller parameter space compared to DNNs. Our extensive experiments, conducted on three benchmark datasets, reveal that ARDST exhibits a representation learning capability similar to DNNs in perceptual tasks and demonstrates resilience against state‐of‐the‐art adversarial attacks.
Shengda Zhuo, Di Wu 0056, Xin Hu 0008, Yu Wang 0017
Int. J. Intell. Syst.4
2023 Online Semi-supervised Learning with Mix-Typed Streaming Features
abstract
Online learning with feature spaces that are not fixed but can vary over time renders a seemingly flexible learning paradigm thus has drawn much attention. Unfortunately, two restrictions prohibit a ubiquitous application of this learning paradigm in practice. First, whereas prior studies mainly assume a homogenous feature type, data streams generated from real applications can be heterogeneous in which Boolean, ordinal, and continuous co-exist. Existing methods that prescribe parametric distributions such as Gaussians would not suffice to model the correlation among such mixtyped features. Second, while full supervision seems to be a default setup, providing labels to all arriving data instances over a long time span is tangibly onerous, laborious, and economically unsustainable. Alas, a semi-supervised online learner that can deal with mix-typed, varying feature spaces is still missing. To fill the gap, this paper explores a novel problem, named Online Semi-supervised Learning with Mixtyped streaming Features (OSLMF), which strives to relax the restrictions on the feature type and supervision information. Our key idea to solve the new problem is to leverage copula model to align the data instances with different feature spaces so as to make their distance measurable. A geometric structure underlying data instances is then established in an online fashion based on their distances, through which the limited labeling information is propagated, from the scarce labeled instances to their close neighbors. Experimental results are documented to evidence the viability and effectiveness of our proposed approach. Code is released in https://github.com/wudi1989/OSLMF.
Di Wu 0056, Shengda Zhuo, Yu Wang 0017, Zhong Chen 0003, Yi He 0007
AAAI3
2023 Application-Layer DDoS Attack Detection Using Explicit Duration Recurrent Network-Based Application-Layer Protocol Communication Models
abstract
Existing application‐layer distributed denial of service (AL‐DDoS) attack detection methods are mainly targeted at specific attacks and cannot effectively detect other types of AL‐DDoS attacks. This study presents an application‐layer protocol communication model for AL‐DDoS attack detection, based on the explicit duration recurrent network (EDRN). The proposed method includes model training and AL‐DDoS attack detection. In the AL‐DDoS attack detection phase, the output of each observation sequence is updated in real time. The observation sequences are based on application‐layer protocol keywords and time intervals between adjacent protocol keywords. Protocol keywords are extracted based on their identification using regular expressions. Experiments are conducted using datasets collected from a real campus network and the CICDDoS2019 dataset. The results of the experiments show that EDRN is superior to several popular recurrent neural networks in accuracy, F 1, recall, and loss values. The proposed model achieves an accuracy of 0.996, F 1 of 0.992, recall of 0.993, and loss of 0.041 in detecting HTTP DDoS attacks on the CICDDoS2019 dataset. The results further show that our model can effectively detect multiple types of AL‐DDoS attacks. In a comparison test, the proposed method outperforms several state‐of‐the‐art approaches.
Bailin Xie, Yu Wang 0017, Guogui Wen
Int. J. Intell. Syst.2
2023 Analysis of hybrid attack and defense based on block withholding strategy
Binjie Liao, Yu Wang 0017, Weizhi Meng 0001, Jun Zhang 0010
J. Inf. Secur. Appl.3
2023 ASSBert: Active and semi-supervised bert for smart contract vulnerability detection
Xiaobing Sun 0001, Liangqiong Tu, Jiale Zhang 0001, Jie Cai 0006, Bin Li 0006, Yu Wang 0017
J. Inf. Secur. Appl.6
2022 Enhancing blockchain-based filtration mechanism via IPFS for collaborative intrusion detection in IoT networks
Wenjuan Li 0001, Yu Wang 0017, Jin Li 0002
J. Syst. Archit.2
2022 DCUS: Evaluating Double-Click-Based Unlocking Scheme on Smartphones
Wenjuan Li 0001, Yu Wang 0017, Jiao Tan
Mob. Networks Appl.2
2021 Online Learning in Variable Feature Spaces with Mixed Data
abstract
This paper explores a new online learning problem where the data streams are generated from an over-time varying feature space, in which the random variables are of mixed data types including Boolean, ordinal, and continuous. The crux of this setting lies in how to establish the relationship among features, such that the learner can enjoy 1) reconstructed information of the missed-out old features and 2) a jump-start of learning new features with educated weight initialization. Unfortunately, existing methods mainly assume a linear mapping relationship among features or that the multivariate joint distribution could be modeled as Gaussians, limiting their applicability to the mixed data streams. To fill the gap, we in this paper propose to model the complex joint distribution underlying mixed data with Gaussian copula, where the observed features with arbitrary marginals are mapped onto a latent normal space. The feature correlation is approximated in the latent space through an online EM process. Two base learners trained on the observed and latent features are ensembled to expedite convergence, thereby minimizing prediction risk in an online setting. Theoretical and empirical studies substantiate the effectiveness of our proposed approach. Code is released in https://github.com/xiexvying/OVFM.
Yi He 0007, Jiaxian Dong, Bo-Jian Hou, Yu Wang 0017, Fei Wang 0001
ICDM4
2021 Enhancing Blackslist-Based Packet Filtration Using Blockchain in Wireless Sensor Networks
Wenjuan Li 0001, Weizhi Meng 0001, Yu Wang 0017, Jin Li 0002
WASA (2)3
2021 Mobile network traffic pattern classification with incomplete a priori information
Zhiping Jin, Zhibiao Liang, Yu Wang 0017, Weizhi Meng 0001
Comput. Commun.3
2021 DRL-R: Deep reinforcement learning approach for intelligent routing in software-defined data-center networks
Waixi Liu 0001, Jun Cai 0002, Qing Chun Chen, Yu Wang 0017
J. Netw. Comput. Appl.4
2021 Recent Advances in Blockchain and Artificial Intelligence Integration: Feasibility Analysis, Research Issues, Applications, Challenges, and Future Work
abstract
Blockchain constructs a distributed point-to-point system, which is a secure and verifiable mechanism for decentralized transaction validation and is widely used in financial economy, Internet of Things, large data, cloud computing, and edge computing. On the other hand, artificial intelligence technology is gradually promoting the intelligent development of various industries. As two promising technologies today, there is a natural advantage in the convergence between blockchain and artificial intelligence technologies. Blockchain makes artificial intelligence more autonomous and credible, and artificial intelligence can prompt blockchain toward intelligence. In this paper, we analyze the combination of blockchain and artificial intelligence from a more comprehensive and three-dimensional point of view. We first introduce the background of artificial intelligence and the concept, characteristics, and key technologies of blockchain and subsequently analyze the feasibility of combining blockchain with artificial intelligence. Next, we summarize the research work on the convergence of blockchain and artificial intelligence in home and overseas within this category. After that, we list some related application scenarios about the convergence of both technologies and also point out existing problems and challenges. Finally, we discuss the future work.
Xifei Song, Lei Liu 0031, Yu Wang 0017, Dapeng Lan
Secur. Commun. Networks5
2020 A Framework of Blockchain-Based Collaborative Intrusion Detection in Software Defined Networking
Wenjuan Li 0001, Jiao Tan, Yu Wang 0017
NSS3
2020 TruNeo: an integrated pipeline improves personalized true tumor neoantigen identification
abstract
BACKGROUND: Neoantigen-based personal vaccines and adoptive T cell immunotherapy have shown high efficacy as a cancer treatment in clinical trials. Algorithms for the accurate prediction of neoantigens have played a pivotal role in such studies. Some existing bioinformatics methods, such as MHCflurry and NetMHCpan, identify neoantigens mainly through the prediction of peptide-MHC binding affinity. However, the predictive accuracy of immunogenicity of these methods has been shown to be low. Thus, a ranking algorithm to select highly immunogenic neoantigens of patients is needed urgently in research and clinical practice. RESULTS: We develop TruNeo, an integrated computational pipeline to identify and select highly immunogenic neoantigens based on multiple biological processes. The performance of TruNeo and other algorithms were compared based on data from published literature as well as raw data from a lung cancer patient. Recall rate of immunogenic ones among the top 10-ranked neoantigens were compared based on the published combined data set. Recall rate of TruNeo was 52.63%, which was 2.5 times higher than that predicted by MHCflurry (21.05%), and 2 times higher than NetMHCpan 4 (26.32%). Furthermore, the positive rate of top 10-ranked neoantigens for the lung cancer patient were compared, showing a 50% positive rate identified by TruNeo, which was 2.5 times higher than that predicted by MHCflurry (20%). CONCLUSIONS: TruNeo, which considers multiple biological processes rather than peptide-MHC binding affinity prediction only, provides prioritization of candidate neoantigens with high immunogenicity for neoantigen-targeting personalized immunotherapies.
Yunxia Tang, Yu Wang 0017, Linmin Peng, Guochao Wei, Jin Li 0002, Zhibo Gao
BMC Bioinform.2
2020 AAMcon: an adaptively distributed SDN controller in data center networks
Waixi Liu 0001, Yu Wang 0017, Hongjian Liao, Zhong-Wei Liang, Xiaochu Liu
Frontiers Comput. Sci.2
2020 Detecting insider attacks in medical cyber-physical networks based on behavioral profiling
Weizhi Meng 0001, Wenjuan Li 0001, Yu Wang 0017, Man Ho Au
Future Gener. Comput. Syst.3
2020 Toward supervised shape-based behavioral authentication on smartphones
Wenjuan Li 0001, Yu Wang 0017, Jin Li 0002, Yang Xiang 0001
J. Inf. Secur. Appl.2
2020 A swipe-based unlocking mechanism with supervised learning on smartphones: Design and evaluation
Wenjuan Li 0001, Jiao Tan, Weizhi Meng 0001, Yu Wang 0017
J. Netw. Comput. Appl.4
2020 Fine-grained flow classification using deep learning for software defined data center networks
Waixi Liu 0001, Jun Cai 0002, Yu Wang 0017, Qing Chun Chen, Jia-Qi Zeng
J. Netw. Comput. Appl.3
2019 A Comparative Study on Network Traffic Clustering
Hanxiao Xue, Guocheng Wei, Lisu Wu, Yu Wang 0017
NSS5
2019 Adaptive machine learning-based alarm reduction via edge computing for distributed intrusion detection systems
abstract
Summary To protect assets and resources from being hacked, intrusion detection systems are widely implemented in organizations around the world. However, false alarms are one challenging issue for such systems, which would significantly degrade the effectiveness of detection and greatly increase the burden of analysis. To solve this problem, building an intelligent false alarm filter using machine learning classifiers is considered as one promising solution, where an appropriate algorithm can be selected in an adaptive way in order to maintain the filtration accuracy. By means of cloud computing, the task of adaptive algorithm selection can be offloaded to the cloud, whereas it could cause communication delay and increase additional burden. In this work, motivated by the advent of edge computing, we propose a framework to improve the intelligent false alarm reduction for DIDS based on edge computing devices. Our framework can provide energy efficiency as the data can be processed at the edge for shorter response time. The evaluation results demonstrate that our framework can help reduce the workload for the central server and the delay as compared to the similar studies.
Yu Wang 0017, Weizhi Meng 0001, Wenjuan Li 0001, Zhe Liu 0001, Hanxiao Xue
Concurr. Comput. Pract. Exp.1
2019 Foreword to the special issue on security, privacy, and social networks
abstract
Social computing and cloud computing are the major trends of technology development in recent years. With the unparalleled popularity, social networks and cloud platforms have become part of our daily lives. Users have produced big data that are beyond the ability of commonly used computer software and hardware tools to capture, manage, and process within a tolerable elapsed time. It has been widely recognized that security and privacy are the key challenges for social network and cloud services due to their scale, complexity, and heterogeneity. The goal of this special issue is to promote research on security, privacy, and social networks. Eleven papers were carefully selected from open submissions and invited from the best original presentations at the 13th International Conference on Information Security Practice and Experience (ISPEC 2017) and the Third International Symposium on Security and Privacy in Social Networks and Big Data (SocialSec 2017). These research papers address the state-of-the-art technologies related to security, privacy, and social networks. The papers are organized under the following topics: security and privacy in social network and web, big data and information security, hardware and software security, and network security. Social media has greater and greater influence on society in recent years. Hu et al1 present an interesting study on the influence of negative opinions spreading in social media during election period. Unlike existing approaches that rely on sentiment analysis and emotional words, the authors take advantage of nouns with emotional context to determine the election preference of each user more accurately. Protecting social networks from security threats and preserving user privacy are the key challenges in social network security. To protect social network users against cross-site scripting worms, Gupta et al2 propose a client-server JavaScript code rewriting-based framework. A Java-based prototype has been developed by the authors, and the authors test its malicious script alleviation capability on several web applications. Yang et al3 propose a novel privacy-preserving authentication protocol for anonymous web browsing, which help users avoid being monitored by the web server on the basis of the identity. In particular, the proposed protocol makes use of a pseudoidentity mechanism and an identity-based elliptic-curve cryptography algorithm. In the area of big data and cloud computing, secure nearest neighbor query over encrypted data is an important issue. Zhu et al4 put forward an efficient attack against the CloudBI-II scheme that is designed for resisting the collusion of cloud server and query users. Accordingly, the authors present an enhanced scheme that can resist the collusion attack. Outsourcing heavy computational tasks to cloud service providers has become popular in the cloud era. As commercial cloud service providers are not trusted, preserving the integrity of computational results becomes an important challenge. Yang et al5 propose a verifiable computation scheme that can protect the output privacy. Ciphertext-policy attribute-based encryption (CP-ABE) is widely used for data access control in cloud storage, which gives data owners direct and flexible control on access policies. Zhang et al6 present a multiauthority attribute-based encryption scheme with constant-size ciphertexts and user revocation for threshold access policy, which addresses the practical challenges in CP-ABE. The security and reliability of information transmission in vehicular ad hoc networks have attracted a lot of research efforts in recent years. Wang et al7 propose a neighborhood trustworthiness-based vehicle-to-vehicle authentication scheme, which uses cloud computing to evaluate the trustworthiness of vehicles for emergent information delivery. The security of ARM embedded devices is important as they are becoming increasingly ubiquitous. Chang et al8 propose a hardware-assisted memory isolation protection mechanism using the B method, and present an implementation of the proposed system on an ARM-based platform. Function-call graph matching is useful in binary code analysis for software security purposes. Huang et al9 propose a function-call graph matching method based on the Hungarian algorithm. The proposed method solves the maximum weight matching problem in polynomial time, which allows matching between graphs of large scale. This special issue would not be complete without covering network security. Yang et al10 use software-defined network techniques to build a moving target defense model that maps physical network elements to a considerably large address space and creates different times of validity randomly to generate mapping addresses. The proposed model helps make it more difficult for attackers to find the targets in a network. Shan et al11 propose a node importance scheme to a community-based caching scheme with network coding for information centric networking. Experimental results indicate that the proposed scheme can improve network performance including average download time, cache hit rate, and instantaneous hop reduction rate. The articles presented in this special issue report recent advances in some areas of security, privacy, and social networks, including security and privacy in social network and web, big data and information security, hardware and software security, and network security. We hope the readers can benefit from the insights of these works and make further contributions to these important and rapidly growing fields. We are grateful to the authors who submitted papers to this special issue. We would also like to thank the reviewers for their hard work and their valuable feedback to the authors. Finally, we would like to express our sincere gratitude to Professor Geoffrey Fox, the Editor in Chief, for providing the opportunity and assistance to edit this special issue in the international journal of Concurrency and Computation: Practice and Experience.
Yang Xiang 0001, Md. Zakirul Alam Bhuiyan, Aniello Castiglione, Yu Wang 0017
Concurr. Comput. Pract. Exp.4
2019 Protecting VNF services with smart online behavior anomaly detection method
Yuxia Cheng, Huijuan Yao, Yu Wang 0017, Yang Xiang 0001, Hongpei Li
Future Gener. Comput. Syst.3
2019 Designing collaborative blockchained signature-based intrusion detection in IoT environments
Wenjuan Li 0001, Steven Tug, Weizhi Meng 0001, Yu Wang 0017
Future Gener. Comput. Syst.4
2019 IoT-FBAC: Function-based access control scheme using identity-based encryption in IoT
Hongyang Yan, Yu Wang 0017, Chunfu Jia, Jin Li 0002, Yang Xiang 0001, Witold Pedrycz
Future Gener. Comput. Syst.2
2018 Enhancing Intelligent Alarm Reduction for Distributed Intrusion Detection Systems via Edge Computing
Weizhi Meng 0001, Yu Wang 0017, Wenjuan Li 0001, Zhe Liu 0001, Jin Li 0002, Christian W. Probst
ACISP2
2018 A Data-driven Attack against Support Vectors of SVM
abstract
Machine learning (ML) is commonly used in multiple disciplines and real-world applications, such as information retrieval, financial systems, health, biometrics and online social networks. However, their security profiles against deliberate attacks have not often been considered. Sophisticated adversaries can exploit specific vulnerabilities exposed by classical ML algorithms to deceive intelligent systems. It is emerging to perform a thorough security evaluation as well as potential attacks against the machine learning techniques before developing novel methods to guarantee that machine learning can be securely applied in adversarial setting. In this paper, an effective attack strategy for crafting foreign support vectors in order to attack a classic ML algorithm, the Support Vector Machine (SVM) has been proposed with mathematical proof. The new attack can minimize the margin around the decision boundary and maximize the hinge loss simultaneously. We evaluate the new attack in different real-world applications including social spam detection, Internet traffic classification and image recognition. Experimental results highlight that the security of classifiers can be worsened by poisoning a small group of support vectors.
Shigang Liu, Jun Zhang 0010, Yu Wang 0017, Wanlei Zhou 0001, Yang Xiang 0001, Olivier Y. de Vel
AsiaCCS3
2018 Analyzing Use of High Privileges on Android: An Empirical Case Study of Screenshot and Screen Recording Applications
Mark Huasong Meng, Guangdong Bai, Joseph K. Liu, Xiapu Luo, Yu Wang 0017
Inscrypt5
2018 Evaluating the Impact of Intrusion Sensitivity on Securing Collaborative Intrusion Detection Networks Against SOOA
David Madsen, Wenjuan Li 0001, Weizhi Meng 0001, Yu Wang 0017
ICA3PP (4)4
2018 Towards Securing Challenge-Based Collaborative Intrusion Detection Networks via Message Verification
Wenjuan Li 0001, Weizhi Meng 0001, Yu Wang 0017, Jinguang Han, Jin Li 0002
ISPEC3
2018 PrivacySearch: An End-User and Query Generalization Tool for Privacy Enhancement in Web Search
Francisco-Javier Rodrigo-Ginés, Javier Parra-Arnau, Weizhi Meng 0001, Yu Wang 0017
NSS4
2018 JFCGuard: Detecting juice filming charging attack via processor usage analysis on smartphones
Weizhi Meng 0001, Lijun Jiang, Yu Wang 0017, Jin Li 0002, Jun Zhang 0010, Yang Xiang 0001
Comput. Secur.3
2018 Special issue on social network security and privacy
abstract
Abstract This special issue contains 28 full papers selected from the Computer Animation
Miroslaw Kutylowski, Yu Wang 0017, Shouhuai Xu, Laurence T. Yang
Concurr. Comput. Pract. Exp.2
2018 Enhancing network capacity by weakening community structure in scale-free network
Jun Cai 0002, Yu Wang 0017, Yan Liu 0042, Jian-Zhen Luo, Wenguo Wei, Xiaoping Xu
Future Gener. Comput. Syst.2
2018 TouchWB: Touch behavioral user authentication based on web browsing on smartphones
Weizhi Meng 0001, Yu Wang 0017, Duncan S. Wong, Sheng Wen, Yang Xiang 0001
J. Netw. Comput. Appl.2
2018 A fog-based privacy-preserving approach for distributed signature-based intrusion detection
Yu Wang 0017, Weizhi Meng 0001, Wenjuan Li 0001, Jin Li 0002, Waixi Liu 0001, Yang Xiang 0001
J. Parallel Distributed Comput.1
2017 Home Location Protection in Mobile Social Networks: A Community Based Method (Short Paper)
Bo Liu 0001, Wanlei Zhou 0001, Shui Yu 0001, Kun Wang 0005, Yu Wang 0017, Yong Xiang 0001, Jin Li 0002
ISPEC5
2017 Exploring Energy Consumption of Juice Filming Charging Attack on Smartphones: A Pilot Study
Lijun Jiang, Weizhi Meng 0001, Yu Wang 0017, Chunhua Su, Jin Li 0002
NSS3
2017 Addressing the class imbalance problem in Twitter spam detection using ensemble learning
Shigang Liu, Yu Wang 0017, Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001
Comput. Secur.2
2017 Statistical Features-Based Real-Time Detection of Drifted Twitter Spam
abstract
Twitter spam has become a critical problem nowadays. Recent works focus on applying machine learning techniques for Twitter spam detection, which make use of the statistical features of tweets. In our labeled tweets data set, however, we observe that the statistical properties of spam tweets vary over time, and thus, the performance of existing machine learning-based classifiers decreases. This issue is referred to as “Twitter Spam Drift”. In order to tackle this problem, we first carry out a deep analysis on the statistical features of one million spam tweets and one million non-spam tweets, and then propose a novel Lfun scheme. The proposed scheme can discover “changed” spam tweets from unlabeled tweets and incorporate them into classifier's training process. A number of experiments are performed to evaluate the proposed scheme. The results show that our proposed Lfun scheme can significantly improve the spam detection accuracy in real-world scenarios.
Chao Chen 0015, Yu Wang 0017, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Geyong Min
IEEE Trans. Inf. Forensics Secur.2
2016 Fuzzy-Based Feature and Instance Recovery
Shigang Liu, Jun Zhang 0010, Yu Wang 0017, Yang Xiang 0001
ACIIDS (1)3
2016 An Ensemble Learning Approach for Addressing the Class Imbalance Problem in Twitter Spam Detection
Shigang Liu, Yu Wang 0017, Chao Chen 0015, Yang Xiang 0001
ACISP (1)2
2016 Gatekeeping Behavior Analysis for Information Credibility Assessment on Weibo
Bailin Xie, Yu Wang 0017, Chao Chen 0015, Yang Xiang 0001
NSS2
2016 Detection and classification of anomaly intrusion using hierarchy clustering and SVM
abstract
Anomaly detection as a kind of intrusion detection is good at detecting the unknown attacks or new attacks, and it has attracted much attention during recent years. In this paper, a new hierarchy anomaly intrusion detection model that combines the fuzzy c-means (FCM) based on genetic algorithm and SVM is proposed. During the process of detecting intrusion, the membership function and the fuzzy interval are applied to it, and the process is extended to soft classification from the previous hard classification. Then a fuzzy error correction sub interval is introduced, so when the detection result of a data instance belongs to this range, the data will be re-detected in order to improve the effectiveness of intrusion detection. Experimental results show that the proposed model can effectively detect the vast majority of network attack types, which provides a feasible solution for solving the problems of false alarm rate and detection rate in anomaly intrusion detection model. Copyright © 2016 John Wiley & Sons, Ltd.
Chenghua Tang, Yang Xiang 0001, Yu Wang 0017, Junyan Qian, Baohua Qiang
Secur. Commun. Networks3
2016 A General Collaborative Framework for Modeling and Perceiving Distributed Network Behavior
abstract
Collaborative Anomaly Detection CAD is an emerging field of network security in both academia and industry. It has attracted a lot of attention, due to the limitations of traditional fortress-style defense modes. Even though a number of pioneer studies have been conducted in this area, few of them concern about the universality issue. This work focuses on two aspects of it. First, a unified collaborative detection framework is developed based on network virtualization technology. Its purpose is to provide a generic approach that can be applied to designing specific schemes for various application scenarios and objectives. Second, a general behavior perception model is proposed for the unified framework based on hidden Markov random field. Spatial Markovianity is introduced to model the spatial context of distributed network behavior and stochastic interaction among interconnected nodes. Algorithms are derived for parameter estimation, forward prediction, backward smooth, and the normality evaluation of both global network situation and local behavior. Numerical experiments using extensive simulations and several real datasets are presented to validate the proposed solution. Performance-related issues and comparison with related works are discussed.
Yi Xie 0002, Yu Wang 0017, Haitao He, Yang Xiang 0001, Shunzheng Yu, Xincheng Liu
IEEE/ACM Trans. Netw.2
2015 Unknown pattern extraction for statistical network protocol identification
abstract
The past decade has seen a lot of research on statistics-based network protocol identification using machine learning techniques. Prior studies have shown promising results in terms of high accuracy and fast classification speed. However, most works have embodied an implicit assumption that all protocols are known in advance and presented in the training data, which is unrealistic since real-world networks constantly witness emerging traffic patterns as well as unknown protocols in the wild. In this paper, we revisit the problem by proposing a learning scheme with unknown pattern extraction for statistical protocol identification. The scheme is designed with a more realistic setting, where the training dataset contains labeled samples from a limited number of protocols, and the goal is to tell these known protocols apart from each other and from potential unknown ones. Preliminary results derived from real-world traffic are presented to show the effectiveness of the scheme.
Yu Wang 0017, Chao Chen 0015, Yang Xiang 0001
LCN1
2014 Windows Event Forensic Process
Quang Do, Ben Martini, Jonathan Looi, Yu Wang 0017, Kim-Kwang Raymond Choo
IFIP Int. Conf. Digital Forensics4
2014 Internet traffic clustering with side information
Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Wanlei Zhou 0001, Bailin Xie
J. Comput. Syst. Sci.1
2014 Internet Traffic Classification Using Constrained Clustering
abstract
Statistics-based Internet traffic classification using machine learning techniques has attracted extensive research interest lately, because of the increasing ineffectiveness of traditional port-based and payload-based approaches. In particular, unsupervised learning, that is, traffic clustering, is very important in real-life applications, where labeled training data are difficult to obtain and new patterns keep emerging. Although previous studies have applied some classic clustering algorithms such as K-Means and EM for the task, the quality of resultant traffic clusters was far from satisfactory. In order to improve the accuracy of traffic clustering, we propose a constrained clustering scheme that makes decisions with consideration of some background information in addition to the observed traffic statistics. Specifically, we make use of equivalence set constraints indicating that particular sets of flows are using the same application layer protocols, which can be efficiently inferred from packet headers according to the background knowledge of TCP/IP networking. We model the observed data and constraints using Gaussian mixture density and adapt an approximate algorithm for the maximum likelihood estimation of model parameters. Moreover, we study the effects of unsupervised feature discretization on traffic clustering by using a fundamental binning method. A number of real-world Internet traffic traces have been used in our evaluation, and the results show that the proposed approach not only improves the quality of traffic clusters in terms of overall accuracy and per-class metrics, but also speeds up the convergence.
Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Wanlei Zhou 0001, Guiyi Wei, Laurence T. Yang
IEEE Trans. Parallel Distributed Syst.1
2013 Network traffic clustering using Random Forest proximities
abstract
The recent years have seen extensive work on statistics-based network traffic classification using machine learning (ML) techniques. In the particular scenario of learning from unlabeled traffic data, some classic unsupervised clustering algorithms (e.g. K-Means and EM) have been applied but the reported results are unsatisfactory in terms of low accuracy. This paper presents a novel approach for the task, which performs clustering based on Random Forest (RF) proximities instead of Euclidean distances. The approach consists of two steps. In the first step, we derive a proximity measure for each pair of data points by performing a RF classification on the original data and a set of synthetic data. In the next step, we perform a K-Medoids clustering to partition the data points into K groups based on the proximity matrix. Evaluations have been conducted on real-world Internet traffic traces and the experimental results indicate that the proposed approach is more accurate than the previous methods.
Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010
ICC1
2013 Unsupervised traffic classification using flow statistical properties and IP packet payload
Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Yu Wang 0017
J. Comput. Syst. Sci.4
2013 Modeling Oscillation Behavior of Network Traffic by Nested Hidden Markov Model with Variable State-Duration
abstract
Network traffic modeling is a fundamental problem in communication. A traffic model should be able to capture and reproduce various properties of a real trace. Despite the widespread success of most numerical models in various applications, few actually focus on the oscillation behavior proven to be one of the basic properties in network traffic. In this paper, a new mathematical method is proposed to model and synthesize stationary and nonstationary oscillatory processes of network traffic. The proposed model is based on the structure of the hierarchical hidden Markov model, which includes two nested hidden Markov chains and one observable process. The first-layer hidden Markov chain with variable state-duration controls the time-varying oscillatory process. Conditional on the first-layer Markov chain, the local fluctuation process is modeled by the second-layer hidden Markov chain. Algorithms are derived for inference of model parameters and traffic synthesis. The proposed approach is compared with four classical models for performance evaluation. The selected performance criterion includes time structure, statistical properties, self-similarity, queuing behavior and multiscale properties. The flexibility and accuracy of the proposed model results in a close fit to the real traces.
Yi Xie 0002, Jiankun Hu, Yang Xiang 0001, Shui Yu 0001, Shensheng Tang, Yu Wang 0017
IEEE Trans. Parallel Distributed Syst.6
2013 Network Traffic Classification Using Correlation Information
abstract
Traffic classification has wide applications in network management, from security monitoring to quality of service measurements. Recent research tends to apply machine learning techniques to flow statistical feature based classification methods. The nearest neighbor (NN)-based method has exhibited superior classification performance. It also has several important advantages, such as no requirements of training procedure, no risk of overfitting of parameters, and naturally being able to handle a huge number of classes. However, the performance of NN classifier can be severely affected if the size of training data is small. In this paper, we propose a novel nonparametric approach for traffic classification, which can improve the classification performance effectively by incorporating correlated information into the classification process. We analyze the new classification approach and its performance benefit from both theoretical and empirical perspectives. A large number of experiments are carried out on two real-world traffic data sets to validate the proposed approach. The results show the traffic classification performance can be improved significantly even under the extreme difficult circumstance of very few training samples.
Jun Zhang 0010, Yang Xiang 0001, Yu Wang 0017, Wanlei Zhou 0001, Yong Xiang 0001
IEEE Trans. Parallel Distributed Syst.3
2012 Internet traffic clustering with constraints
abstract
Due to the limitations of the traditional port-based and payload-based traffic classification approaches, the past decade has seen extensive work on utilizing machine learning techniques to classify network traffic based on packet and flow level features. In particular, previous studies have shown that the unsupervised clustering approach is both accurate and capable of discovering previously unknown application classes. In this paper, we explore the utility of side information in the process of traffic clustering. Specifically, we focus on the flow correlation information that can be efficiently extracted from packet headers and expressed as instance-level constraints, which indicate that particular sets of flows are using the same application and thus should be put into the same cluster. To incorporate the constraints, we propose a modified constrained K-Means algorithm. A variety of real-world traffic traces are used to show that the constraints are widely available. The experimental results indicate that the constrained approach not only improves the quality of the resulted clusters, but also speeds up the convergence of the clustering process.
Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Shunzheng Yu
IWCMC1
2012 Generating regular expression signatures for network traffic classification in trusted network management
Yu Wang 0017, Yang Xiang 0001, Wanlei Zhou 0001, Shunzheng Yu
J. Netw. Comput. Appl.1
2011 A novel semi-supervised approach for network traffic clustering
abstract
Network traffic classification is an essential component for network management and security systems. To address the limitations of traditional port-based and payload-based methods, recent studies have been focusing on alternative approaches. One promising direction is applying machine learning techniques to classify traffic flows based on packet and flow level statistics. In particular, previous papers have illustrated that clustering can achieve high accuracy and discover unknown application classes. In this work, we present a novel semi-supervised learning method using constrained clustering algorithms. The motivation is that in network domain a lot of background information is available in addition to the data instances themselves. For example, we might know that flow f1and f2are using the same application protocol because they are visiting the same host address at the same port simultaneously. In this case, f1and f2shall be grouped into the same cluster ideally. Therefore, we describe these correlations in the form of pair-wise must-link constraints and incorporate them in the process of clustering. We have applied three constrained variants of the K-Means algorithm, which perform hard or soft constraint satisfaction and metric learning from constraints. A number of real-world traffic traces have been used to show the availability of constraints and to test the proposed approach. The experimental results indicate that by incorporating constraints in the course of clustering, the overall accuracy and cluster purity can be significantly improved.
Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Shunzheng Yu
NSS1
2010 Automatic Application Signature Construction from Unknown Traffic
abstract
Identifying applications and classifying network traffic flows according to their source applications are critical for a broad range of network activities. Such classifications can be based on information derived from packet header fields and payload content, or statistical characteristics of flows and communication patterns of hosts. However, most of present methods rely on some forms of priori knowledge. In this paper, an application signature based traffic classification system with a novel approach to fully automate the process of deriving signatures from unknown traffic is proposed. The key idea is to combine traffic clustering based on statistical flow properties in order to generate clusters dominated by a single application on the one hand, and application signature construction solely based on payload content from each cluster on the other hand. Evaluation using real-world traffic traces indicate that the proposed approach is highly effective.
Yu Wang 0017, Yang Xiang 0001, Shunzheng Yu
AINA1
2010 An automatic application signature construction system for unknown traffic
abstract
Abstract Identifying applications and classifying network traffic flows according to their source applications are critical for a broad range of network activities. Such a decision can be based on packet header fields, packet payload content, statistical characteristics of traffic and communication patterns of network hosts. However, most present techniques rely on some sort ofa prioriknowledge, which means they require labor‐intensive preprocessing before running and cannot deal with previously unknown applications. In this paper, we propose a traffic classification system based on application signatures, with a novel approach to fully automate the process of deriving signatures from unidentified traffic. The key idea is to integrate statistics‐based flow clustering with payload‐based signature matching method, so as to eliminate the requirement of pre‐labeled training data sets. We evaluate the efficiency of our approach using real‐world traffic trace, and the results indicate that signature classifiers built from clustered data and pre‐labeled data are able to achieve similar high accuracy better than 99%. Copyright © 2010 John Wiley & Sons, Ltd.
Yu Wang 0017, Yang Xiang 0001, Shunzheng Yu
Concurr. Comput. Pract. Exp.1
2009 Automatic Network Protocol Automaton Extraction
abstract
Protocol reverse engineering, the process of (re)constructing the protocol context of communication sessions by an implementation, which involves translating a sequence of packets into protocol messages, grouping them into sessions, and modeling state transitions in the protocol state machine, is well-known to be invaluable for many network security applications, including intrusion prevention and detection, traffic normalization, and penetration testing, etc. However, current practice in deriving protocol specifications is either mostly manual or focusing on automatic reverse engineering the message format only and leaving the protocol state machine inverse undone. Although regular expressions offer superior expressive ability and flexibility, application protocols are described by regular expression manually based on sufficiently understanding protocol itself. At present there is not an effect method to realize classification, recognition and control automatically for the known applications and the unknown applications in future. In this paper a novel approach is presented to model network application specification. In this work, the whole automatic protocol reverse engineering is realized through accomplishing the protocol state machine, and then the FSMs are translated to corresponding regular expressions to enrich and update the pattern database. This approach uses grammatical inference and is motivated by the observation that an implementation of the protocol is inherently a state transition process, the state machine model the essence exactly. The important significance is to describe various state protocols with a common method through modeling the protocol state transition, including known and unknown ones. This approach had been implemented in the system and evaluated using real-world implementations of three different protocols: HTTP, SMTP, FTP, and compared the extracted protocol to the corresponding other newly system, such as 17-filter.
Ming-Ming Xiao, Shunzheng Yu, Yu Wang 0017
NSS3