Lizhi Peng

dblp:61/2166 · DBLP profile ↗
← Back
57ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-6009-522XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-author · 6 since 2021Computer networks · 12 · 2 first-author · 6 since 2021Systems, architecture and hardware · 10 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PEDTA: Parallel Analysis for High-Precision Website Fingerprinting in Real-World Networks
Lizhi Peng, Liang Jiao, Yutong Yan, Ze Kang
ICIC (2)2
2026 HF-Transformer: A Non-Pretrained Encrypted Network Traffic Classification Model Based on Packet Header Fields
Lizhi Peng, Peiqiang Liu, Yingshuo Bao, Bo Yang 0001
INFOCOM2
2026 NT-Transformer: A Non-Pretrained Encrypted Network Traffic Classification Model
abstract
Network traffic classification plays an indispensable role in network management, Quality of Service (QoS), and cybersecurity. With the widespread encryption techniques applied to network traffic, it has become increasingly challenging to classify network traffic into different management groups accurately. In recent years, pre-training Transformer-based models have been successfully applied to Natural Language Processing (NLP), and researchers have also introduced such models into encrypted network traffic analysis. However, besides the similarities of words in NLP and byte codes in network traffic, there exist essential differences between them, which may cause inefficacy of the pretrained model when being applied to new traffic data. In this paper, we propose a non-pretrained encrypted network traffic classification model based on Transformer called NT-Transformer, which can directly learn labeled network traffic features at two levels of granularity, namely, byte level (uni-gram or bi-gram) and flow level (packet size and packet inter-arrival time), without the relatively expensive pre-training procedure of unlabeled data. This method is validated on three public datasets and three sets of recently collected network traffic data. Experimental results indicate that in some scenarios, pretrained models offer limited performance gains when applied to new encrypted network traffic data not encountered during pretraining, and NT-Transformer with uni-gram byte representation outperforms the state-of-the-art models in terms of pushing the F1 score up by 0.25% - 2.24%.
Lizhi Peng, Peiqiang Liu, Yingshuo Bao, Bo Yang 0001
IEEE Trans. Netw. Serv. Manag.2
2025 Semi-GD: Graph-Level Consistency Regularization and Dynamic Threshold Boosted Semi-Supervised Learning
abstract
The rapid development in pseudo-labeling and consistency regularization has significantly advanced semi-supervised learning in computer vision. However, these methods have notable limitations. Pseudo-labeling approaches suffer from confirmation bias, as the model trained on self-generated labels tends to perpetuate or even amplify inaccuracies. Existing consistency regularization techniques focus on the consistency of individual instances, but overlook the alignment of relationships between samples. We propose a new semi-supervised learning method called Semi-GD, to effectively address these challenges. Our method implements two types of matching: prediction matching, which enforces instance consistency, and graph matching, which enforces the consistency of affinity relationships among batch samples. These two matching processes mutually reinforce each other to achieve a robust model. Furthermore, a simple yet effective thresholding strategy is introduced in prediction matching process to select high-confidence samples, alleviating confirmation bias. The target graph is constructed under dual constraints from cluster labels and pseudo-labels, enabling more accurate graph learning. In summary, Semi-GD utilizes data at both the individual sample level (through prediction matching) and the inter-sample relationship level (through graph matching) to improve model performance. Extensive experiments demonstrate the effectiveness of our method. Notably, on CIFAR10 and CIFAR100 with only four labels per class, Semi-GD reduces the error rate by 8.86% and 11.43% compared to FixMatch, respectively.
Ruixue Wang, Lizhi Peng, ShengBao Li, Huawei Yang
IJCNN2
2025 SCANet: Multi-Channel Representation with Synthesis Convolution and Composite Attention for Malicious Network Traffic Detection
abstract
Malicious network traffic detection is crucial for safeguarding cybersecurity. However, most existing approaches heavily rely on statistical features manually crafted by domain experts from network flows. This process is cumbersome and lacks real-time adaptability to address the latest cybersecurity threats. Furthermore, the quality of the feature set directly impacts the model’s performance. To address these challenges, we propose a deep learning-based model for malicious traffic detection that focuses on two critical aspects: input data representation and model architecture. For input data representation, we introduce a novel method that characterizes packets as multichannel color data, significantly enhancing data effectiveness compared to single-channel grayscale representations. Regarding model architecture, we have designed a convolutional neural network that integrates a synthesis convolution pattern with a composite attention mechanism. The synthesis convolution operates through the combined action of 2D convolution and 1D convolution, effectively adapting to the characteristics of the input network traffic data and efficiently extracting deep features from raw data. The composite attention mechanism, named the bi-dimensional fusion attention mechanism, further improves the model’s ability to identify key features while reducing redundant information. Experimental results on three datasets demonstrate that the proposed data representation method significantly improves the performance of all compared traffic analysis models. Additionally, the proposed model achieves the best detection results across all datasets, especially in multi-class classification tasks, highlighting the superiority of its architecture.
Qingsheng Yang, Lizhi Peng, Jianjun Dong
IJCNN2
2025 Imbalanced ensemble learning leveraging a novel data-level diversity metric
Ying Pang, Lizhi Peng, Haibo Zhang 0001, Bo Yang 0001
Pattern Recognit.2
2024 Burst Sequence Based Graph Neural Network for Video Traffic Identification
Yingshuo Bao, ShengBao Li, Lizhi Peng
ICDF2C (2)5
2024 Collaborative Learning With Heterogeneous Local Models: A Rule-Based Knowledge Fusion Approach
abstract
Federated Learning (FL) has emerged as a promising collaborative learning paradigm that enables to train machine learning models across decentralized devices, while keeping the training data localized to preserve user privacy. However, the heterogeneity in both decentralized training data and distributed computing resources has posed significant challenges to the design of effective and efficient FL schemes. Most existing solutions either focus on tackling a single type of heterogeneity, or are unable to fully support model heterogeneity with low communication overhead, fast convergence, and good interpretability. In this paper, we present CloREF, a novel rule-based collaborative learning framework that allows devices in FL to use completely different local learning models to cater to both data and resource heterogeneity. In CloREF, each rule is represented as a linear model, which provides good interpretability. Each participating device chooses a local model and trains it using its local data. The decision boundary of each trained local model is then approximated using a set of rules, which effectively bridges the gap arising from model heterogeneity. All participating devices collaborate to select the optimal set of rules as the global model, employing evolutionary optimization to effectively fuse the knowledge acquired from all local models. Experimental results on both synthesized and real-world datasets demonstrate that the rules generated by our proposed method can mimic the behaviors of various learning models with high fidelity ($\gt $0.95 in most tests), and CloREF gives competitive performance in accuracy, AUC, and communication overhead, compared with both the best-performing model trained centrally and several state-of-the-art model-heterogeneous federated learning schemes.
Ying Pang, Haibo Zhang 0001, Jeremiah D. Deng, Lizhi Peng, Fei Teng 0001
IEEE Trans. Knowl. Data Eng.4
2023 An Early Stage Identification of Cryptomining Behavior with DNS Requests
Yihang Hao, Mengda Lyu, Xiaojie Yu, Bo Yang 0001, Lizhi Peng
ADMA (5)6
2023 A Novel Adaptive Distribution Distance-Based Feature Selection Method for Video Traffic Identification
Shuaili Liu, Qingsheng Yang, Zhongfeng Qu, Lizhi Peng
ADMA (4)5
2023 Adversarial Attack with Genetic Algorithm against IoT Malware Detectors
abstract
The exponential growth and sophistication of Internet of Things (IoT) malware behavior have resulted in new detection technologies capable of defending IoT devices against some threats. However, their success has stimulated the interest of attackers attempting to circumvent current IoT malware detectors. Among detection technologies, the detectors trained based on Uniform Resource Locator (URL) requests have become popular. To draw attention to the safety of the detectors, we propose a grey-box method to attack detectors based on URL requests without breaking malicious functions of URL requests. The key idea is to add perturbations to the tail of URLs. Specifically, this method is based on a Genetic Algorithm (GA) to find suitable perturbations and optimizes the process of adversarial attacks through a dynamic number of evolution directions and a maximum generation limit. The effectiveness of our adversarial attack is demonstrated by experimental results based on a widely used public dataset CSIC2010 and several representative detectors. As far as we know, this is the first time an adversarial attack against IoT detectors based on URL requests has been done. The method has an attack success rate of more than 92 %. Furthermore, experiment results show that the method can reduce query numbers while maintaining the attack success rate.
Shanshan Wang 0003, Wenyue Wang, Daokuan Bai, Lizhi Peng
ICC6
2023 Devils in the Clouds: An Evolutionary Study of Telnet Bot Loaders
abstract
One of the innovations brought by Mirai and its derived malware is the adoption of self-contained loaders for infecting IoT devices and recruiting them in botnets. Functionally decoupled from other botnet components and not embedded in the payload, loaders cannot be analysed using conventional approaches that rely on honeypots for capturing samples. Different approaches are necessary for studying the loaders evolution and defining a genealogy. To address the insufficient knowledge about loaders' lineage in existing studies, in this paper, we propose a semantic-aware method to measure, categorize, and compare different loader servers, with the goal of highlighting their evolution, independent from the payload evolution. Leveraging behavior-based metrics, we cluster the discovered loaders and define eight families to determine the genealogy and draw a homology map. Our study shows that the source code of Mirai is evolving and spawning new botnets with new capabilities, both on the client side and the server side. In turn, shedding light on the infection loaders can help the cybersecurity community to improve detection and prevention tools.
Yuhui Zhu, Qiben Yan 0001, Shanshan Wang 0003, Alberto Giaretta 0001, Enlong Li, Lizhi Peng, Mauro Conti
ICC7
2023 D-CoA: Probability Density-based Confidence Assignment Semi-Supervised Clustering Ensemble
abstract
In recent years, semi-supervised clustering ensembles have gained significant attention in the field of clustering. These ensembles leverage limited prior knowledge to guide the clustering process. Nonetheless, throughout the integration process, owing to the employment of diverse fundamental clustering algorithms, disparate algorithms generate incongruent results. Consequently, ascertain a more reliable conclusion becomes a formidable challenge. In this paper, we introduce D-CoA (Density based Confidence of Assignment), a novel semi-supervised clustering ensemble algorithm, designed to navigate the intricacies of this issue. D-CoA incorporates sub-clusters using the Minimum Spanning Tree algorithm and evaluates the confidence of each amalgamated cluster to establish prior confidence. Subsequently, a Mixture Density Network is leveraged to calculate the probability density distribution for each merged cluster, yielding posterior confidence scores for resolving conflicts in sample assignments. By adopting this methodology, we amplify the reliability of conflicting sample assignments. Empirical results demonstrate the versatility of this approach, as it effectively accommodates probability distributions of varying shapes, thus demonstrating its feasibility and scalability for diverse data types.
Lizhi Peng, Yihang Hao, Bo Yang 0001
ICPADS2
2023 Sketch-based Real-time Intrusion Detection Framework for Industrial Internet of Things
abstract
Intrusion detection technology is of great significance to enhance the network security protection of industrial Internet of Things (IIoT) and ensure the efficient implementation of the production process. Aiming at the problems of poor real-time performance and high false positive rate (FPR) of existing intrusion detection methods in IIoT, a sketch-based intrusion detection framework is proposed. The framework employs an improved sketch algorithm as a primary classification module, which is able to process raw traffic in real-time, and perform traffic splitting and filtering efficiently. To improve accuracy, the filtered traffic is fed into a machine learning (ML) based secondary classification module for further analysis. Compared to direct analysis, the sketch can filter out a portion of the benign traffic, thus improving the efficiency of the secondary module. Our framework can work as a pre-module for existing ML based intrusion detection methods, giving them the ability to process traffic in real-time and reduce FPRs. We validate the performance of the framework by testing it on the latest publicly available dataset, Modbus 2023.
Mengda Lyu, Lizhi Peng, Bo Yang 0001
ICPADS2
2023 Live-Stream Identification Based on Reasoning Network with Core Traffic Set
Yingshuo Bao, Shuaili Liu, Zhongfeng Qu, Lizhi Peng
KSEM (2)4
2023 ECM-EFS: An ensemble feature selection based on enhanced co-association matrix
abstract
Currently, feature selection faces a huge challenge that no single feature selection method can effectively deal with various data sets for all real cases. Ensemble learning is a potential promising solution to address this problem. We propose an ensemble feature selection method based on enhanced co-association matrix (ECM-EFS). Positive-co-association matrix (PCM), negative-co-association matrix (NCM), and relative-co-association matrix (RCM) are first introduced to discover the relationship among features by ensembling the results in multiple feature selection methods. To further produce a more stable feature selection result, “Feature Kernel” is also introduced and used as a starting point for feature selection. Comparative experiments with four state-of-the-art methods have confirmed that the ECM-EFS can provide more robust results. Moreover, compared with traditional ensemble feature selection methods, our method can compensate information loss and reduce computational cost significantly.
Yihang Hao, Bo Yang 0001, Lizhi Peng
Pattern Recognit.4
2022 Video traffic identification with a distribution distance-based feature selection
abstract
The boosting of video traffic in recent years has brought severe challenges to Internet management as video traffic shows wide variety, including harmful video traffic. For example, game video traffic should be strictly supervised for teenagers. Therefore, effectively identifying video traffic has become a fundamental issue in network management, and one of the popular and difficult subjects in existing research. To this end, we propose to extract a large-scale feature set for video traffic identification in this paper. Then, a new distribution distance-based feature selection (DDFS) approach is proposed to obtain an effective feature subset. To test the effectiveness of the extracted feature and the DDFS, we designed a video traffic collection architecture to collect different video traffic data. We carried out a set of comparative experiments on the collected dataset, the experimental data suggest that the proposed approach can achieve an accuracy of over 90% for video scene traffic identification and 98% for cloud gaming video traffic identification. Additionally, the comparison of DDFS with four feature selection methods shows that DDFS is a practical feature selection technique for video traffic identification.
Shuaili Liu, Peifa Sun, Yingshuo Bao, Lizhi Peng
IPCCC5
2022 Rule-Based Collaborative Learning with Heterogeneous Local Learning Models
Ying Pang, Haibo Zhang 0001, Jeremiah D. Deng, Lizhi Peng, Fei Teng 0001
PAKDD (1)4
2022 An early stage convolutional feature extracting method using for mining traffic detection
Peifa Sun, Mengda Lyu, Bo Yang 0001, Lizhi Peng
Comput. Commun.5
2022 Gaussian Distribution Based Oversampling for Imbalanced Data Classification
abstract
The imbalanced data classification problem widely exists in many real-world applications. Data resampling is a promising technique to deal with imbalanced data through either oversampling or undersampling. However, the traditional data resampling approaches simply take into account the local neighbor information to generate new instances in linear ways, leading to the generation of incorrect and unnecessary instances. In this study, we propose a new data resampling technique, namely, Gaussian Distribution based Oversampling (GDO), to handle the imbalanced data for classification. In GDO, anchor instances are selected from the minority class instances in a probabilistic way by taking into account the density and distance information carried by the minority instances. Then new minority instances are generated following a Gaussian distribution model. The proposed method is validated in experimental study by comparing with seven imbalanced learning approaches on 40 data sets from the KEEL repository and 10 large data sets from the UCI repository. Experimental results show that our method outperforms the other compared methods in terms of AUC, G-mean and memory usage with an increase in running time. We also apply GDO to deal with two real imbalanced data classification problems: Internet video traffic identification and metastasis detection of esophageal cancer. The experimental results once again validate the effectiveness of our approach.
Yuxi Xie, Haibo Zhang 0001, Lizhi Peng
IEEE Trans. Knowl. Data Eng.4
2021 IEdroid: Detecting Malicious Android Network Behavior Using Incremental Ensemble of Ensembles
abstract
Malware detection has attracted widespread attention due to the growing malware sophistication. Machine learning based methods have been proposed to find traces of malware by analyzing network traffic. However, network traffic exhibits a series of growing and changing states, which makes it challenging to design a detection model that can detect malicious traffic over a long period without the need for costly retraining. In this paper, we present, IEdroid, an Android malicious network behavior detection method that leverages incremental ensembles for model update. Specifically, we train multiple classifiers to form an interim ensemble in distributed cluster environment, and update the interim ensemble by removing and adding classifiers. The generated model is composed of multiple interim ensembles that can adapt to the network traffic. We evaluated the performance of IEdroid using a dataset consisting of 98,565 benign and 41,267 malicious flows. Results show that IEdroid can effectively detect malicious traffic compared with state-of-the-art detection models. The experiment trained IEdroid on datasets incrementally for 10 times without a significant loss on accuracy, precision, recall, and F-Measure, compared with re-training from scratch with full data.
Anli Yan, Haibo Zhang 0001, Qiben Yan 0001, Lizhi Peng
ICPADS6
2021 AndroCreme: Unseen Android Malware Detection Based on Inductive Conformal Learning
abstract
Android platform is facing serious malware threats due to its popularity, as evidenced by the drastic increase on the number of mobile malware families and variants in recent years. Detecting malware variants and zero-day malware is a critical challenge that must be addressed to protect mobile devices against malware attacks. In this study, we present AndroCreme, a novel network intrusion detection system (NIDS) that can identify unseen malware by analyzing the network behavior of Android malware. To address the temporal bias issue in NIDS, we propose a method for rapid iterative update of the model based on data selection and data size limitation. The selection of effective data is carried out by induction and conformal technology, and the data scale is controlled by the method of time window and data cycle selection. To further achieve fast training speed and high efficiency, we leverage a gradient boosting framework that uses a tree-based learning algorithm, namely, LightGBM, as the meta predictor. We evaluate the performance of AndroCreme over 400K real-world network flows, which are collected from over 30K Android benignware and 21K malware applications. The experimental results show that, compared with the retraining method using all data, AndroCreme requires only a small amount of datareduce more than 3x to obtain better detection performance, which effectively solves the temporal bias.
Lizhi Peng, Yuhui Zhu
TrustCom4
2021 Effective detection of mobile malware behavior based on explainable deep neural network
Anli Yan, Haibo Zhang 0001, Lizhi Peng, Qiben Yan 0001, Muhammad Umair Hassan, Bo Yang 0001
Neurocomputing4
2020 Network-based Malware Detection with a Two-tier Architecture for Online Incremental Update
abstract
As smartphones carry more and more private information, it has become the main target of malware attacks. Threats on mobile devices have become increasingly sophisticated, making it imperative to develop effective tools that are able to detect and counter such threats. Unfortunately, existing malware detection tools based on machine learning techniques struggle to keep up due to the difficulty in performing online incremental update on the detection models. In this paper, a Two-tier Architecture Malware Detection (TAMD) method is proposed, which can learn from the statistical features of network traffic to detect malware. The first layer of TAMD identifies uncertain samples in the training set through a preliminary classification, whereas the second layer builds an improved classifier by filtering out such samples. We enhance TAMD with an incremental leaning based technique (TAMD-IL), which allows to incrementally update the detection models without retraining it from scratch by removing and adding sub-models in TAMD. We experimentally demonstrate that TAMD outperforms the existing methods with up to 98.72% on precision and 96.57% on recall. We also evaluate TAMD-IL on four concept drift datasets and compare it with classical machine learning algorithms, two state-of-the-art malware detection technologies, and three incremental learning technologies. Experimental results show that TAMD-IL is efficient in terms of both update time and memory usage.
Anli Yan, Riccardo Spolaor, Shuaishuai Tan, Lizhi Peng, Bo Yang 0001
IWQoS6
2020 Gradient descent evolved imbalanced data gravitation classification with an application on Internet video traffic identification
Anqi Teng, Lizhi Peng, Yuxi Xie, Haibo Zhang 0001
Inf. Sci.2
2020 Deep and broad URL feature mining for android malware detection
Shanshan Wang 0003, Qiben Yan 0001, Ke Ji, Lizhi Peng, Bo Yang 0001, Mauro Conti
Inf. Sci.5
2019 A Prognosis Method for Esophageal Squamous Cell Carcinoma Based on CT Image and Three-Dimensional Convolutional Neural Networks
Kaipeng Fan, Jifeng Guo 0002, Bo Yang 0001, Lin Wang 0004, Lizhi Peng, Ajith Abraham
ISDA5
2019 Ranking-based biased learning swarm optimizer for large-scale optimization
Hanbo Deng, Lizhi Peng, Haibo Zhang 0001, Bo Yang 0001
Inf. Sci.2
2019 Imbalanced learning based on adaptive weighting and Gaussian function synthesizing with an application on Android malware detection
Ying Pang, Lizhi Peng, Bo Yang 0001, Hongli Zhang 0001
Inf. Sci.2
2019 A mobile malware detection method using behavior features in network traffic
Shanshan Wang 0003, Qiben Yan 0001, Bo Yang 0001, Lizhi Peng, Zhongtian Jia
J. Netw. Comput. Appl.5
2019 Accelerating data gravitation-based classification using GPU
Lizhi Peng, Haibo Zhang 0001, Houcine Hassan, Yuehui Chen, Bo Yang 0001
J. Supercomput.1
2018 A Fast and Effective Detection of Mobile Malware Behavior Using Network Traffic
Shanshan Wang 0003, Lizhi Peng, Yuliang Shi
ICA3PP (4)4
2018 Accurate Identification of Internet Video Traffic Using Byte Code Distribution Features
Yuxi Xie, Hanbo Deng, Lizhi Peng
ICA3PP (1)3
2018 Machine learning based mobile malware detection using highly imbalanced network traffic
abstract
In recent years, the number and variety of malicious mobile apps have increased drastically, especially on Android platform, which brings insurmountable challenges for malicious app detection. Researchers endeavor to discover the traces of malicious apps using network traffic analysis. In this study, we combine network traffic analysis with machine learning methods to identify malicious network behavior, and eventually to detect malicious apps. However, most network traffic generated by malicious apps is benign, while only a small portion of traffic is malicious, leading to an imbalanced data problem when the traffic model skews towards modeling the benign traffic. To address this problem, we introduce imbalanced classification methods, including the synthetic minority oversampling technique (SMOTE) + support vector machine (SVM), SVM cost-sensitive (SVMCS), and C4.5 cost-sensitive (C4.5CS) methods. However, when the imbalance rate reaches a certain threshold, the performance of common imbalanced classification algorithms degrades significantly. To avoid performance degradation, we propose to use the imbalanced data gravitation-based classification (IDGC) algorithm to classify imbalanced data. Moreover, we develop a simplex imbalanced data gravitation classification (S-IDGC) model to further reduce the time costs of IDGC without sacrificing the classification performance. In addition, we propose a machine learning based comparative benchmark prototype system, which provides users with substantial autonomy, such as multiple choices of the desired classifiers or traffic features. Using this prototype system, users can compare the detection performance of different classification algorithms on the same data set, as well as the performance of a specific classification algorithm on multiple data sets.
Qiben Yan 0001, Hongbo Han, Shanshan Wang 0003, Lizhi Peng, Lin Wang 0004, Bo Yang 0001
Inf. Sci.5
2017 Top-k Merit Weighting PBIL for Optimal Coalition Structure Generation of Smart Grids
Sean Hsin-Shyuan Lee, Jeremiah D. Deng, Lizhi Peng, Martin K. Purvis, Maryam Purvis
ICONIP (4)3
2017 Imbalanced traffic identification using an imbalanced data gravitation-based classification model
Lizhi Peng, Haibo Zhang 0001, Yuehui Chen, Bo Yang 0001
Comput. Commun.1
2017 A fast feature weighting algorithm of data gravitation classification
Lizhi Peng, Hongli Zhang 0001, Haibo Zhang 0001, Bo Yang 0001
Inf. Sci.1
2017 A novel semi-supervised learning method for Internet application identification
Zhusong Liu, Lizhi Peng, Lin Wang 0004, Lei Zhang 0085
Soft Comput.3
2017 Flexible neural trees based early stage identification for IP traffic
Lizhi Peng, Chong-zhi Gao, Bo Yang 0001, Yuehui Chen, Jin Li 0002
Soft Comput.2
2017 Adaptive Message Routing and Replication in Mobile Opportunistic Networks for Connected Communities
abstract
Mobile opportunistic networking is a promising technology that can supplement existing cellular and WiFi networks to provide desirable services for smart and connected communities. Message routing is the most compelling challenge in mobile opportunistic networks due to the lack of contemporaneous end-to-end paths and the resource constraints at mobile devices. To improve the probability of successful message delivery, most existing routing schemes use the past contact history to predict future contacts for message forwarding, and exploit message replication and redundancy for multicopy routing. However, most existing prediction-based routing schemes simply use the average pairwise contact probability as the routing metric and neglect the benefits of exploring fine-grained contact information such as pairwise repeated contact patterns to improve the accuracy of predicting future contacts. Moreover, there is no efficient mechanism that can adaptively control message replication in a decentralized manner to achieve both high probability of successful message delivery and low message overhead. To address these problems, we present FGAR, a routing protocol designed for mobile opportunistic networks by leveraging fine-grained contact characterization and adaptive message replication. In FGAR, contact history is characterized in a fine-grained manner with timing information using a sliding window mechanism, and future contacts are predicted based on the fine-grained contact information, thereby improving the accuracy of contact prediction. We further design an efficient message replication scheme in which message replication is controlled in a fully decentralized manner by taking into account the expected message delivery probability, the replication history, and the quality of the encountered device. A replica is generated only when it is necessary to fulfill the expected message delivery probability. We evaluate our scheme through trace-driven simulations, and the simulation results show that FGAR outperforms existing schemes. In comparison with PRoPHET, FGAR can achieve more than 20% improvement on average on successful message delivery, whereas the message overhead has been reduced by a factor up to 15.
Haibo Zhang 0001, Luming Wan, Yawen Chen 0001, Laurence T. Yang, Lizhi Peng
ACM Trans. Internet Techn.5
2016 SMOTE-DGC: An Imbalanced Learning Approach of Data Gravitation Based Classification
Lizhi Peng, Haibo Zhang 0001, Bo Yang 0001, Yuehui Chen, Xiaoqing Zhou
ICIC (2)1
2016 Extraction of Feature Points on 3D Meshes Through Data Gravitation
Chengwei Wang, Dan Kang, Xiuyang Zhao, Lizhi Peng, Caiming Zhang 0001
ICIC (2)4
2016 TrafficAV: An effective and explainable detection of mobile malware behavior using network traffic
abstract
Android has become the most popular mobile platform due to its openness and flexibility. Meanwhile, it has also become the main target of massive mobile malware. This phenomenon drives a pressing need for malware detection. In this paper, we propose TrafficAV, which is an effective and explainable detection of mobile malware behavior using network traffic. Network traffic generated by mobile app is mirrored from the wireless access point to the server for data analysis. All data analysis and malware detection are performed on the server side, which consumes minimum resources on mobile devices without affecting the user experience. Due to the difficulty in identifying disparate malicious behaviors of malware from the network traffic, TrafficAV performs a multi-level network traffic analysis, gathering as many features of network traffic as necessary. The proposed method combines network traffic analysis with machine learning algorithm (C4.5 decision tree) that is capable of identifying Android malware with high accuracy. In an evaluation with 8,312 benign apps and 5,560 malware samples, TCP flow detection model and HTTP detection model all perform well and achieve detection rates of 98.16% and 99.65%, respectively. In addition, for the benefit of user, TrafficAV not only displays the final detection results, but also analyzes the behind-the-curtain reason of malicious results. This allows users to further investigate each feature's contribution in the final result, and to grasp the insights behind the final decision.
Shanshan Wang 0003, Lei Zhang 0085, Qiben Yan 0001, Bo Yang 0001, Lizhi Peng, Zhongtian Jia
IWQoS6
2015 A Real-time Android Malware Detection System Based on Network Traffic Analysis
Hongbo Han, Qiben Yan 0001, Lizhi Peng, Lei Zhang 0085
ICA3PP (3)4
2015 Effective packet number for early stage internet traffic identification
Lizhi Peng, Bo Yang 0001, Yuehui Chen
Neurocomputing1
2014 Feature Selection Toward Optimizing Internet Traffic Behavior Identification
Lizhi Peng, Shupeng Zhao, Lei Zhang 0085, Shan Jing
ICA3PP (2)2
2014 Feature Evaluation for Early Stage Internet Traffic Identification
Lizhi Peng, Hongli Zhang 0001, Bo Yang 0001, Yuehui Chen
ICA3PP (1)1
2014 A new approach for imbalanced data classification based on data gravitation
Lizhi Peng, Hongli Zhang 0001, Bo Yang 0001, Yuehui Chen
Inf. Sci.1
2011 A parallel evolving algorithm for flexible neural tree
Lizhi Peng, Bo Yang 0001, Lei Zhang 0085, Yuehui Chen
Parallel Comput.1
2010 Traffic identification using flexible neural trees
abstract
Traditional traffic classification techniques like port-based and payload-based techniques are becoming ineffective owning to more and more Internet applications using dynamic port number and encryption techniques. Therefore, in the past few years, many researches have addressed machine learning-based techniques. Most researches of machine learning-based traffic identification use traffic samples collected on key nodes of networks for their learning. These samples do not have accurate application information i. e. the ground truth which is crucial for machine learning algorithms. In this paper, we first designed a distributed host based traffic collecting platform (DHTCP) to gather traffic samples with accurate application information on user hosts. Then we built a data set using DHTCP, and applied Flexible Neural Trees (FNT) - a special kind of artificial neural network which has been successfully applied in many areas, for traffic identification. Web and P2P traffics were studied in our work. Although the proposed technique is at an early stage of development, experimental results show that it is a promising solution of Internet traffic identification.
Lizhi Peng, Hongli Zhang 0001, Bo Yang 0001, Yuehui Chen, Mahmoud T. Qassrawi
IWQoS1
2009 Data gravitation based classification
Lizhi Peng, Bo Yang 0001, Yuehui Chen, Ajith Abraham
Inf. Sci.1
2007 Automatic Design of Hierarchical Takagi-Sugeno Type Fuzzy Systems Using Evolutionary Algorithms
abstract
This paper presents an automatic way of evolving hierarchical Takagi-Sugeno fuzzy systems (TS-FS). The hierarchical structure is evolved using probabilistic incremental program evolution (PIPE) with specific instructions. The fine tuning of the if - then rule's parameters encoded in the structure is accomplished using evolutionary programming (EP). The proposed method interleaves both PIPE and EP optimizations. Starting with random structures and rules' parameters, it first tries to improve the hierarchical structure and then as soon as an improved structure is found, it further fine tunes the rules' parameters. It then goes back to improve the structure and the rules' parameters. This loop continues until a satisfactory solution (hierarchical TS-FS model) is found or a time limit is reached. The proposed hierarchical TS-FS is evaluated using some well known benchmark applications namely identification of nonlinear systems, prediction of the Mackey-Glass chaotic time-series and some classification problems. When compared to other neural networks and fuzzy systems, the developed hierarchical TS-FS exhibits competing results with high accuracy and smaller size of hierarchical architecture.
Yuehui Chen, Bo Yang 0001, Ajith Abraham, Lizhi Peng
IEEE Trans. Fuzzy Syst.4
2006 A DGC-Based Data Classification Method Used for Abnormal Network Intrusion Detection
Bo Yang 0001, Lizhi Peng, Yuehui Chen, Hanxing Liu, Runzhang Yuan
ICONIP (3)2
2006 Gene Expression Profiling Using Flexible Neural Trees
Yuehui Chen, Lizhi Peng, Ajith Abraham
IDEAL2
2006 Hierarchical Radial Basis Function Neural Networks for Classification Problems
Yuehui Chen, Lizhi Peng, Ajith Abraham
ISNN (1)2
2006 Exchange Rate Forecasting Using Flexible Neural Trees
Yuehui Chen, Lizhi Peng, Ajith Abraham
ISNN (2)2
2006 Stock Index Modeling Using Hierarchical Radial Basis Function Networks
Yuehui Chen, Lizhi Peng, Ajith Abraham
KES (3)2