VLDB 2026 Research / reviewers in the wild / expert
Junfeng Wang 0003
dblp:15/885-3
· DBLP profile ↗
50ranked-venue papers
1as first author
33since 2021 · last 2026
0000-0003-1699-2270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 9 since 2021Security and privacy · 15 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RuleLLM: LLM-driven rule generation for anomaly network traffic identificationabstractAbstract The ongoing evolution of network traffic attacks poses a considerable challenge for intrusion detection systems. While intelligent models can effectively identify traffic patterns, detection rules are typically easier to comprehend than these complex models. Currently, the generation of detection rules relies heavily on expert knowledge and insights from the security community alongside traditional automated generation methods. However, these generated rules often prove incomplete and inaccurate, failing to adapt to the constantly changing network environment and the rapid advancements in attack technologies. To fill this gap, this paper presents a traffic detection rule generation method called RuleLLM based on a large language model (LLM). RuleLLM establishes a framework for rule generation that utilizes this model and incorporates domain knowledge into the pretrained model using techniques such as Low-Rank Adaptation fine-tuning, few-sample fast engineering, and a feedback-driven revision mechanism. The effectiveness of the rules generated by this model has been validated through a real detection engine. Experimental results indicate that the proposed traffic rule generation method, which leverages an LLM, can analyze original traffic data comprehensively and effectively. With fine-tuning based on a small number of samples, a high detection rate of 91.8% was achieved, showcasing its ability to generate effective rules and enhance network defense capabilities. Tongcan Lin, Junfeng Wang 0003 |
Comput. J. | 2 |
| 2026 | LiteJam: A Lightweight Deep Learning Architecture for Real-Time GNSS Interference Detection and Characterization in UAVsabstractGlobal Navigation Satellite System (GNSS) interference poses a serious threat to Unmanned Aerial Vehicles (UAVs), potentially leading to navigation failures, airspace violations, or even loss of flight control. Although deep learning methods have demonstrated strong performance in interference detection and characterization, most models remain too computationally expensive for onboard deployment due to high computational cost. To address this challenge, we propose LiteJam, a lightweight architecture that utilizes pre-correlation in-phase and quadrature (I/Q) data to construct pseudo-image representations without requiring additional hardware. Specifically, LiteJam adopts a multi-scale convolutional architecture to capture interference patterns, employs a dynamic sparse attention mechanism to adaptively emphasize spatio-spectral cues, and leverages a hierarchical multi-head module for interference detection and characterization. Experimental results show that LiteJam outperforms all baselines. The F1-score of interference classification is 95.74%, outperforming lightweight baselines by 4.37%–24.51%, and generalizes well across diverse scenarios, while maintaining high computational efficiency for real-time UAV applications. Our codes are available at https://github.com/CynthiaCYX/LiteJam. Yuxue Chen, Junfeng Wang 0003, Zhiyang Fang, Tianjie Ni, Jiaxuan Geng, Wenhan Ge |
IEEE Internet Things J. | 2 |
| 2026 | ThreatMAMBA: Achieving High-Robustness Cyber Threat Attribution During the Evolution of Attacks
Wenhan Ge, Junfeng Wang 0003, Zeyuan Cui, Zhiyang Fang, Weilu Zhan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | MT-DEGCL: Multi-Task Encrypted Traffic Classification With Dual Embedding and Graph Contrastive LearningabstractAlthough encryption offers strong anonymity, it also facilitates the concealment of malicious activities, allowing adversaries to evade detection, and posing a great challenge to cybersecurity surveillance. Many existing encrypted traffic classification methods struggle to integrate flow- and packet-level tasks effectively, as they are trained independently, which is redundancy. Additionally, packet header and payload are treated equally, leading to the rich information in raw bytes remains fully unexplored, particularly in the abundant payload data. Moreover, they neglect the semantic invariance and common features between data samples, which ultimately results in suboptimal performance. To address these challenges, we propose an effective Multi-Task model using Dual Embedding and Graph Contrastive Learning (MT-DEGCL). Based on the byte-packet-flow structure of network traffic, a parallel dual embedding embeds the header and payload separately, followed by a cross-gated feature fusion strategy to capture the strong local packet-level representation. Then, we construct the traffic interaction graph and further utilize graph contrastive learning to extract the robust global flow-level representation. Finally, a multi-task model is trained for joint flow- and packet-level classification, leveraging the complementary learning between tasks to enhance overall performance. The experimental results on four real datasets highlight the effectiveness of MT-DEGCL, demonstrating superior performance in both tasks. Specifically, on the ISCX-Tor dataset, MT-DEGCL achieves F1 scores of 98.63% for flow-level classification and 98.10% at the packet level, surpassing the state-of-the-art (i.e., DE-GNN) by 2.03% and 83.21%, respectively. Furthermore, MT-DEGCL maximizes the rich information in raw payload bytes, significantly reducing or even nearly eliminating classification loss when using only payload data. Xiaolan Zhu, Junfeng Wang 0003, Wenhan Ge, Xinbo Han |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | RTsFCM: a robust two-stage flow correlation method for traffic tracking in anonymous communicationabstractAbstract Anonymous communication serves as the preferred tool for cyber attackers to evade detection, posing a serious threat to cyberspace security. Accurately tracking the attackers in anonymous communication is crucial for defending against attacks. Flow correlation is an effective method that can link flows in the anonymous network. Existing flow correlation methods usually rely on a long observation, resulting in reduced correlation precision and limited generalization ability within anonymous communication. To address this issue, we propose a robust two-stage flow correlation method called RTsFCM via Siamese network and ensemble voting scheme. In the first stage, a Siamese network with shared weights is utilized to automatically extract the multilevel features from ingress flow and egress flow, respectively. Further, they are concatenated to generate a more expressive feature set to enhance true positive rate (TPR). In the second stage, flow pairs are firstly divided into a series of partially overlapping sub-flows(windows) in view of flow duration. Then, pairwise comparison for each window is conducted independently and the ensemble voting scheme is adopted across these windows to reduce the false positive rate (FPR) significantly. Experimental results show that RTsFCM is superior to the state of the art, achieving over a 4% increase in both TPR and F1_score. Simultaneously, it obtains an FPR as low as 0.68%, utilizing the packet timing characteristics within the initial portion of a flow. Xiaolan Zhu, Junfeng Wang 0003, Zihua Song, Peng Wu 0036 |
Comput. J. | 2 |
| 2025 | AdvOT: Oblivious transfer based on generative adversarial networks against multiple attackers
Zhentian Zhong, Junfeng Wang 0003 |
Neurocomputing | 5 |
| 2025 | RaxCS: Towards cross-language code summarization with contrastive pre-training and retrieval augmentation
Kaiyuan Yang 0004, Junfeng Wang 0003, Zihua Song |
Inf. Softw. Technol. | 2 |
| 2025 | Perturbation-Resilient for Temporal-Camouflaged IoT AttacksabstractThe growing adoption of Internet of Things (IoT) devices has introduced significant challenges to network security due to their heterogeneous nature and temporal pattern vulnerabilities. Among the emerging threats, adversarial attacks targeting IoT network traffic have gained attention for their ability to evade traditional and machine learning-based Network Intrusion Detection Systems (NIDS). While prior work has focused on static adversarial perturbations, these approaches fail to account for the temporal dynamics inherent in IoT traffic. IoT networks exhibit time-dependent patterns driven by device behavior, environmental factors and user interactions, creating an opportunity for more sophisticated adversarial strategies. This paper introduces a co-evolutionary adversarial framework termed dynamic Adversarial Temporal-Camouflaged Perturbation (ATCP) for IoT network attack traffic. ATCP dynamically segments network traffic into temporal intervals and applies targeted adversarial perturbations to each segment. By leveraging the temporal characteristics of IoT traffic, the proposed method generates subtle yet effective adversarial modifications that confuse NIDS by disrupting their ability to model time-dependent traffic patterns. Unlike static perturbation methods, ATCP provides valuable insights into adversarial attack methodologies, lays the foundation for developing more robust IoT security frameworks, and adapts to evolving traffic dynamics, making it a more effective and robust NIDS in real-world scenarios. Extensive experiments conducted on real-world IoT network datasets demonstrate that the proposed method achieves high evasion rates against Machine Learning (ML) based NIDS while preserving the functional integrity of IoT communications. Notably, among the four NIDS evaluated, KitNET experiences the most significant degradation, with its detection rate dropping from 93.07% to 18.55% after applying ATCP. Furthermore, ATCP exhibits strong adaptability across diverse IoT device types and network configurations, highlighting its generalizability. Zhentian Zhong, Linfeng Tan, Junfeng Wang 0003 |
IEEE Internet Things J. | 5 |
| 2025 | Toward Dynamic Topology Obfuscation for IoT Networks With Evolutionary DefenseabstractTraditional defenses against DDoS and LFAs in IoT networks, such as static network topology obfuscation and traffic engineering, suffer from critical limitations. These include high computational overhead that is incompatible with low-power devices, rigid policies that are vulnerable to evolving attacks, and isolated modules that create exploitable gaps. Accordingly, a new problem setting for IoT network defense is introduced, i.e., dynamic open-world network defense that poses two critical tasks: 1) Open defense, aiming not only to mitigate attacks from known patterns but also to detect and counteract unknown adversarial strategies through real-time traffic analysis and adaptive obfuscation. 2) Dynamic adaptation, aiming to continuously learn new attack types (e.g., novel probing techniques or topology inference methods) without retraining the entire model while preserving robustness against previously encountered threats and avoiding catastrophic forgetting of learned defense policies. To address these problems, we propose OptiPathNet, a unified framework integrating multi-objective evolutionary optimization with dynamic topology obfuscation and real-time traffic analysis. At its core, OptiPathNet leverages two innovations. First is a GCMOEA-based method that dynamically balances security, latency and energy efficiency through Pareto-optimal trade-offs, reducing attackers’ topology inference accuracy while ensuring control packet delays and consumption. Moreover, a synergistic detection-obfuscation framework combining lightweight traffic classification with adaptive obfuscation policies, enabling protocol-compliant noise injection to disrupt both reconnaissance for traceroute-based probing and forensic analysis for timing-based attack traces. Evaluations across IoT scenarios, i.e., smart cities and industrial systems, demonstrate OptiPathNet’s superiority over the state-of-the-art methods. Reduces the structural similarity between the attacker-inferred and real topologies by 43%, while the delay overhead was reduced by 17%, and achieves optimal defense in 95% of the scenarios tested. By transforming IoT’s inherent constraints, including resource limitations and dynamic topologies, into defensive assets, OptiPathNet establishes a scalable, adaptive security paradigm for next-generation networks, bridging the gap between theoretical innovation and practical deployment in resource-constrained environments. Zhentian Zhong, Junfeng Wang 0003, Zhiping Cai |
IEEE Internet Things J. | 5 |
| 2025 | Enhancing vulnerability repair through the extraction and matching of repair patterns
Xiansheng Cao, Junfeng Wang 0003, Peng Wu 0036 |
J. Syst. Softw. | 2 |
| 2025 | CorreFlow: A Covert Fingerprinting Modulation for Flow Correlation in Open Heterogeneous NetworksabstractThe constantly changing landscape of the Internet presents a significant challenge in the detection and tracking of covert attackers and their sophisticated methods. To address this issue, various techniques, such as network flow watermarking (NFW) and traffic correlation, embed attack labels in data streams to identify attack pathways or aid post-analysis. However, existing solutions are often tailored to specific scenarios, resulting in lacking robustness, adaptability, and anonymity under non-cooperative or incomplete information heterogeneous environments. To this end, this paper proposesCorreFlow, a Transfer Learning (TL) based invisible network flow correlation framework utilizing time channel graph fingerprinting modulation. It considers the fragmentation and reassembly of data packets during transmission. In simple network environments,CorreFlowutilizes TL for rapid correlation across flows, enabling efficient linkage of related traffic segments. In complex heterogeneous network, where traditional correlation methods may fail due to encryption and variability, it leverages Inter Packet Delay (IPD) for encrypted flow matching and accurately identifies the optimal watermark point. Multiple experiments conducted on real network traffic and public datasets have demonstrated thatCorreFlowachieves highly efficient traffic correlation with minimal false positive rate, improved adaptability, and steganography. Specifically, it has achieved over 97.31% accuracy in various network environments and promotes network traffic correlation in open heterogeneous network environments from low correlation to 95%. Junfeng Wang 0003, Wenhan Ge, Lingfeng Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | WF-TFC: An Open-World Few-Shot Anonymous Website Fingerprinting via Time-Frequency ConsistencyabstractWhile Tor provides strong anonymity, it also facilitates the concealment of malicious activities, which poses a significant challenge to cybersecurity surveillance. As an effective anti-anonymity technique, Website Fingerprinting(WF) enables the inference of which websites a user is visiting, thereby uncovering potential attacker activities. State-of-the-art(SOTA) methods have demonstrated remarkable effectiveness. However, a large number of labeled traffic is required to ensure effectiveness, and without timely updates, these models will encounter serious challenges of concept drift due to the dynamic nature of website content and network conditions. The core reasons lie in the independently and identically distributed assumption, while in challenging open-world scenarios, the long-term spatial and temporal dynamics complicates data consistency and effective knowledge transfer. To address these issues, this paper presents WF-TFC, an open-world few-shot anonymous WF model via self-supervised contrastive learning and time-frequency consistency. It aligns time- and frequency-based representations in the latent time-frequency space, enhancing the sustained effectiveness of inherent patterns across various websites. Consequently, it accommodates diverse few-shot target domains with varying dynamics, facilitating data consistency and knowledge transfer in unobserved long-term temporal and spatial environments. For instance, with only 5 traces per website, WF-TFC achieves 92.62% accuracy on traces collected six weeks after pre-training, exceeding the SOTA(i.e., NetCLR) by 2.12%. On similar but mutually exclusive traces, it attains an F1 score of 87.20%, surpassing the SOTA by 6.12%. Xiaolan Zhu, Junfeng Wang 0003, Wenhan Ge, Yizhao Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | SeqMask: Behavior Extraction Over Cyber Threat Intelligence Via Multi-Instance LearningabstractAbstract Identification and extraction of Tactics, Techniques and Procedures (TTPs) for Cyber Threat Intelligence (CTI) restore the full picture of cyber attacks and guide the analysts to assess the system risk. Existing frameworks can hardly provide uniform and complete processing mechanisms for TTPs information extraction without adequate knowledge background. A multi-instance learning approach named SeqMask is proposed in this paper as a solution. SeqMask extracts behavior keywords from CTI evaluated by the semantic impact, and predicts TTPs labels by conditional probabilities. Still, the framework has two mechanisms to determine the validity of keywords. One using expert experience verification. The other verifies the distortion of the classification effect by blocking existing keywords. In the experiments, SeqMask reached 86.07% and 73.99% in F1 scores for TTPs classifications. For the top 20% of keywords, the expert approval rating is 92.20%, where the average repetition of keywords whose scores between 100% and 90% is 60.02%. Particularly, when the top 65% of the keywords were blocked, the F1 decreased to about 50%; when removing the top 50%, the F1 was under 31%. Further, we also validate the possibility of extracting TTPs from full-size CTI and malware whose F1 are improved by 2.16% and 0.81%. Wenhan Ge, Junfeng Wang 0003 |
Comput. J. | 2 |
| 2024 | A survey of strategy-driven evasion methods for PE malware: Transformation, concealment, and attack
Jiaxuan Geng, Junfeng Wang 0003, Zhiyang Fang, Wenhan Ge |
Comput. Secur. | 2 |
| 2024 | VulMPFF: A Vulnerability Detection Method for Fusing Code Features in Multiple PerspectivesabstractSource code vulnerabilities are one of the significant threats to software security. Existing deep learning‐based detection methods have proven their effectiveness. However, most of them extract code information on a single intermediate representation of code (IRC), which often fails to extract multiple information hidden in the code fully, significantly limiting their performance. To address this problem, we propose VulMPFF, a vulnerability detection method that fuses code features under multiple perspectives. It extracts IRC from three perspectives: code sequence, lexical and syntactic relations, and graph structure to capture the vulnerability information in the code, which effectively realizes the complementary information of multiple IRCs and improves vulnerability detection performance. Specifically, VulMPFF extracts serialized abstract syntax tree as IRC from code sequence, lexical and syntactic relation perspective, and code property graph as IRC from graph structure perspective, and uses Bi‐LSTM model with attention mechanism and graph neural network with attention mechanism to learn the code features from multiple perspectives and fuse them to detect the vulnerabilities in the code, respectively. We design a dual‐attention mechanism to highlight critical code information for vulnerability triggering and better accomplish the vulnerability detection task. We evaluate our approach on three datasets. Experiments show that VulMPFF outperforms existing state‐of‐the‐art vulnerability detection methods (i.e., Rats, FlawFinder, VulDeePecker, SySeVR, Devign, and Reveal) in Acc and F1 score, with improvements ranging from 14.71% to 145.78% and 152.08% to 344.77%, respectively. Meanwhile, experiments in the open‐source project demonstrate that VulMPFF has the potential to detect vulnerabilities in real‐world environments. Xiansheng Cao, Junfeng Wang 0003, Peng Wu 0036, Zhiyang Fang |
IET Inf. Secur. | 2 |
| 2024 | GMADV: An android malware variant generation and classification adversarial training framework
Shuangcheng Li, Zhangguo Tang, Huanzhou Li, Jian Zhang 0056, Han Wang 0044, Junfeng Wang 0003 |
J. Inf. Secur. Appl. | 6 |
| 2024 | MetaCluster: A Universal Interpretable Classification Framework for CybersecurityabstractRising cyber threats have created an immediate demand for Deep Learning (DL) in cybersecurity. Nevertheless, the opaque nature of DL models poses challenges in deploying, collaborating, and assessing their effectiveness in less reliable cybersecurity environments. Despite eXplainable Artificial Intelligence (XAI) playing a role in enhancing cybersecurity analytics, the limited task scope, the propensity for data overfitting, and the stochastic explanations hinder its broader application. To fill the gap, this paper introduces a generic interpretable classification framework, named MetaCluster. MetaCluster generates semantic prototypes for features, patterns, and domains at varying granular levels by following three fundamental steps: embedding representations, acquiring prototypes, and aggregating semantics. These mechanisms guarantee that MetaCluster achieves critical information extraction and reliable classification at minimal cost. The experiments encompass cybersecurity classification tasks and assess the interpretability of the framework. These tasks encompass malware family classification, threat behavior analysis, and malicious traffic identification. In particular, when compared to other DL models, MetaCluster exhibits a significant reduction in parameter consumption by 79.52% to 91.78%, and boosts operational speed up to 71.37%, while its F1 scores remain stable or slightly increase. Additionally, MetaCluster possesses the ability to assess and visually represent the significance of image, text, and statistical features. This capability leads to a reduction of Mean Squared Error (MSE) between expected and actual predictions by 0.0101 to 0.1020. Wenhan Ge, Zeyuan Cui, Junfeng Wang 0003, Binhui Tang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | DOMR: Toward Deep Open-World Malware RecognitionabstractDeep learning has been widely used for Android malware family recognition, but current deep learning-based approaches make the closed-world assumption that malware families encountered during testing are available at training phase. Unfortunately, this assumption is often violated in practice due to the constant emergence of novel categories and the huge cost of collecting abundant training classes, causing serious failures to the existing approaches. Accordingly, a new problem setting for Android malware family recognition is introduced, i.e., deep open-world malware recognition that poses two critical tasks: 1) Open recognition, aiming to not only classify malware from known families (present in training) but detect malware from unknown families (absent in training); 2) Incremental update, aiming to learn about the detected unknown/new categories without retraining from scratch and catastrophically forgetting the previously learned known/old classes. This paper formalizes the problem and proposes a novel solution called DOMR to address the above two tasks in a unified framework. The core of DOMR is an episode-based representation learning scheme that mimics the open-world setting through episodic training to learn a generalizable representation. The key insight is that the training process following the open-world setting forces the representation to accumulate experience in open recognition, thereby facilitating both the classification of known family instances and the detection of unknown family instances at inference. Given this representation, multiple one-vs-rest classifiers are subsequently built to make the final recognition decision through an aggregative strategy. Comparative experiments show that DOMR outperforms start-of-the-art methods, with macro-averaged F1-scores obtained on two datasets reaching 80.88% and 56.17% in the open case, and 79.34% and 49.55% in the incremental case, respectively. Ablation studies further analyze the effectiveness of DOMR in achieving the open recognition and incremental update goals. Junfeng Wang 0003 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Tactics And Techniques Classification In Cyber Threat IntelligenceabstractAbstract Completing the classification of tactics and techniques in cyber threat intelligence (CTI) is an important way to obtain tactics, techniques and procedures (TTPs) and portray the behavior of cyber attacks. However, the high level of abstraction of tactics and techniques information and their presence in CTI, usually in the form of natural language text, make it difficult for traditional manual analysis methods and feature engineering-based machine learning methods to complete the classification of tactics and techniques effectively. Meanwhile, flat deep learning methods do not perform well in classifying more fine-grained techniques due to their inability to exploit the hierarchical relationship between tactics and techniques. Therefore, this paper regards the tactics and techniques of TTPs defined in Adversarial Tactics, Techniques and Common Knowledge knowledge base as labels and proposes a Convolutional Neural Network (CNN) model based on hierarchical knowledge migration and attention mechanism for classifying tactics and techniques in CTI, named HM-ACNN (CNN based on hierarchical knowledge migration and attention mechanism). HM-ACNN classifies tactics and techniques into two phases, and the underlying network model for both phases is the Attention-based CNN network. The first step in HM-ACNN is converting the CTI text into a two-dimensional image based on the word embedding model, and then start training the classification of tactics through the CNN structure based on the attention mechanism before the classification of techniques. Secondly, after the tactics classification training is completed, the tactic-to-technique knowledge migration is then completed by transforming the parameters of the CNN layer and the attention layer in the tactics classification process based on the special hierarchical relationship between tactics and techniques. Then, the classification of techniques is finished by fine-tuning. The experimental results show that HM-ACNN performs well in the tactics and techniques classification tasks, and the metric F1 values reach 93.66% and 86.29%, which are better than other models such as CNN, Recurrent Neural Network and CRNN (Recurrent Convolutional Neural Networks). Zhongkun Yu, Junfeng Wang 0003, Binhui Tang |
Comput. J. | 2 |
| 2023 | CI_GRU: An efficient DGA botnet classification model based on an attention recurrence plot
Han Wang 0044, Zhangguo Tang, Huanzhou Li, Jian Zhang 0056, Shuangcheng Li, Junfeng Wang 0003 |
Comput. Networks | 6 |
| 2023 | Explainable cyber threat behavior identification based on self-adversarial topic generation
Wenhan Ge, Junfeng Wang 0003, Tongcan Lin, Binhui Tang |
Comput. Secur. | 2 |
| 2023 | DroidRL: Feature selection for android malware detection with reinforcement learning
Yinwei Wu, Meijin Li, Junfeng Wang 0003, Zhiyang Fang, Luyu Cheng |
Comput. Secur. | 5 |
| 2023 | HGIVul: Detecting inter-procedural vulnerabilities based on hypergraph convolution
Zihua Song, Junfeng Wang 0003, Kaiyuan Yang 0004, Jigang Wang |
Inf. Softw. Technol. | 2 |
| 2023 | Learning a holistic and comprehensive code representation for code summarization
Kaiyuan Yang 0004, Junfeng Wang 0003, Zihua Song |
J. Syst. Softw. | 2 |
| 2022 | F2DC: Android malware classification based on raw traffic and neural networks
Junfeng Wang 0003 |
Comput. Networks | 2 |
| 2022 | Markov-GAN: Markov image enhancement method for malicious encrypted traffic classificationabstractAbstract The rapidly growing encrypted traffic hides a large number of malicious behaviours. The difficulty of collecting and labelling encrypted traffic makes the class distribution of dataset seriously imbalanced, which leads to the poor generalisation ability of the classification model. To solve this problem, a new representation learning method in encrypted traffic and its diversity enhancement model are proposed, which uses the diversity of images to represent the diversity of traffic samples. First, the encrypted traffic is transformed into Markov images. Then, a diversity maximisation Markov‐GAN based on the Simpson index is designed to generate new Markov images. Finally, the balanced Markov image set is sent to the CNN for classification. Experimental results show that the proposed method can predict the whole dataset space with only a few original samples. And the classification accuracies under different imbalance degrees are significantly improved, all of which are over 90%. The enhanced Markov image set can effectively alleviate performance generalisation deviation caused by different network depths. Even an ordinary CNN has almost the same classification effect as VGG13 and VGG16. Compared with other data enhancement methods, the Markov‐GAN only needs to balance the transform domain dataset, which is lightweight, easy to train and has stronger amplification ability. Zhangguo Tang, Junfeng Wang 0003, Baoguo Yuan, Huanzhou Li, Jian Zhang 0056, Han Wang 0044 |
IET Inf. Secur. | 2 |
| 2022 | Enhancing software modularization via semantic outliers filtration and label propagation
Kaiyuan Yang 0004, Junfeng Wang 0003, Zhiyang Fang, Peng Wu 0036, Zihua Song |
Inf. Softw. Technol. | 2 |
| 2022 | IoT Malware Classification Based on Lightweight Convolutional Neural NetworksabstractInternet of Things (IoT) is hard to deploy adequate security defenses due to the diversity of architectures as well as the limited computing and storage capabilities, which makes it more vulnerable to malware. With the massive deployment of IoT devices, how to accurately identify and classify the malware variants is crucial to IoT security. However, existing methods of IoT malware classification generally support specific platform or require complex models to achieve higher accuracies. To solve these problems, this article proposes an IoT malware classification method based on lightweight convolutional neural networks (LCNNs). First, the malware binaries are converted into multidimensional Markov images. Then, the LCNN is designed with two new operations, depthwise convolution and channel shuffle, for malware images classification. Compared with other deep learning-based methods such as VGG16, the designed LCNN can greatly reduce trainable parameters while maintaining accuracy. The generated model of LCNN is only about 1 MB, while that of VGG16 is 552.57 MB. The average accuracies of the proposed method are higher than that of gray images on multiple IoT malware data sets, all of which are over 95%. Compared with the state-of-the-art low-level features-based methods, the average accuracy of the proposed method is 99.356% on the Microsoft data set even if the model is tiny. The results show that the proposed method is not only suitable for IoT environments but also has high accuracy. Baoguo Yuan, Junfeng Wang 0003, Peng Wu 0036, Xianguo Qing |
IEEE Internet Things J. | 2 |
| 2022 | LSTM Based Phishing Detection for Big Email DataabstractIn recent years, cyber criminals have successfully invaded many important information systems by using phishing mail, causing huge losses. The detection of phishing mail from big email data has been paid public attention. However, the camouflage technology of phishing mail is becoming more and more complex, and the existing detection methods are unable to confront with the increasingly complex deception methods and the growing number of emails. In this article, we proposed an LSTM based phishing detection method for big email data. The new method includes two important stages, sample expansion stage and testing stage under sufficient samples. In the sample expansion stage, we combined KNN with K-Means to expand the training data set, so that the size of training samples can meet the needs of in-depth learning. In the testing stage, we first preprocess these samples, including generalization, word segmentation and word vector generation. Then, the preprocessed data is used to train a LSTM model. Finally, on the basis of the trained model, we classify the phishing emails. By experiment, we evaluate the performance of the proposed method, and experimental results show that the accuracy of our phishing detection method can reach 95 percent. Qi Li 0057, Mingyu Cheng, Junfeng Wang 0003 |
IEEE Trans. Big Data | 3 |
| 2021 | Software plagiarism detection in multiprogramming languages using machine learning approachabstractSummary The Software plagiarism, which arises the problem of software piracy is a growing major concern nowadays. It is a serious risk to the software industry that gives huge economic damages every year. The customers may develop a modified version of the original software in other types of programming languages. Furthermore, the plagiarism detection in different types of source codes is a challenging task because each source code may have specific syntax rules. In this paper, we proposed a methodology for software plagiarism detection in multiprogramming languages based on machine learning approaches. The Principal Component Analysis (PCA) is applied for features extraction from source codes without losing the actual information. It extracts features by factor analysis and converts the dataset into normalized linear principal components which are further useful for predictions analysis. Then, the multinomial logistic regression model (MLR) is applied to these components to classify the source codes documents based on predictions. It gives the generalization of logistic regression to handle multiclass problems. Further, the predictors' performance in MLR is evaluated by 2 tailed z test. To apply the experiment, the dataset is collected in five different and popular languages, ie, C, C++, Java, C#, and Python. Each programming language taken in two different case studies, ie, binary search and Stack. Farhan Ullah 0001, Junfeng Wang 0003, Masood Habib, Shehzad Khalid |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | CNN-Based Malware Variants Detection Method for Internet of ThingsabstractMalware has become one of the most serious security threats to the Internet of Things (IoT). Detection of malware variants can inhibit the spread of malicious code from the traditional network to the IoT, and can also inhibit the spread of malicious code within the IoT, which is of great significance to the security detection and defense of the IoT. Since the terminals and the operating systems of IoT are very different from the traditional network, when malicious code is transferred from the traditional network to the IoT platform, the characteristics of the variants may change significantly. As a result, malicious code variant detection methods for traditional platforms cannot be directly applied to the IoT. In this article, a malware variant detection method for the IoT is proposed. First, we propose a feature representation method based on RGB image for IoT to solve the problem of representation difficulty caused by platform difference, which pays more attention to the assembly code and developer information of the malware. The generated image has richer texture information, which can dig out the deep association between the IoT variants and the original malicious code. Moreover, this article improves the convolutional neural network model by combining the self-attention mechanism and spatial pyramid pooling to solve the problem of large differences in the size of IoT malware. Experimental results show that our method can be used in cross-platform to detect malware variants in the IoT effectively. Qi Li 0057, Jiaxin Mi, Junfeng Wang 0003, Mingyu Cheng |
IEEE Internet Things J. | 4 |
| 2021 | IRTS: An Intelligent and Reliable Transmission Scheme for Screen Updates Delivery in DaaSabstractDesktop-as-a-service (DaaS) has been recognized as an elastic and economical solution that enables users to access personal desktops from anywhere at any time. During the interaction process of DaaS, users rely on screen updates to perceive execution results remotely, and thus the reliability and timeliness of screen updates transmission have a great influence on users’ quality of experience (QoE). However, the efficient transmission of screen updates in DaaS is facing severe challenges: most transmission schemes applied in DaaS determine sending strategies in terms of pre-set rules, lacking the intelligence to utilize bandwidth rationally and fit new network scenarios. Meanwhile, they tend to focus on reliability or timeliness and perform unsatisfactorily in ensuring reliability and timeliness simultaneously, leading to lower transmission efficiency of screen updates and users’ QoE when network conditions turn unfavorable. In this article, an intelligent and reliable end-to-end transmission scheme (IRTS) is proposed to cope with the preceding issues. IRTS draws support from reinforcement learning by adopting SARSA, an online learning method based on the temporal difference update rule, to grasp the optimal mapping between network states and sending actions, which extricates IRTS from the reliance on pre-set rules and augments its adaptability to different network conditions. Moreover, IRTS guarantees reliability and timeliness via an adaptive loss recovery method, which intends to recover lost screen updates data automatically with fountain code while controlling the number of redundant packets generated. Extensive performance evaluations are conducted, and numerical results show that IRTS outperforms the reference schemes in display quality, end-to-end delay/delay jitter, and fairness when transferring screen updates under various network conditions, proving that IRTS can enhance the transmission efficiency of screen updates and users’ QoE in DaaS. Hongdi Zheng, Junfeng Wang 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Cognitive Covert Traffic Synthesis Method Based on Generative Adversarial NetworkabstractIn the intelligent era of human‐computer symbiosis, the use of machine learning method for covert communication confrontation has become a hot topic of network security. The existing covert communication technology focuses on the statistical abnormality of traffic behavior and does not consider the sensory abnormality of security censors, so it faces the core problem of lack of cognitive ability. In order to further improve the concealment of communication, a game method of “cognitive deception” is proposed, which is aimed at eliminating the anomaly of traffic in both behavioral and cognitive dimensions. Accordingly, a Wasserstein Generative Adversarial Network of Covert Channel (WCCGAN) model is established. The model uses the constraint sampling of cognitive priors to construct the constraint mechanism of “functional equivalence” and “cognitive equivalence” and is trained by a dynamic strategy updating learning algorithm. Among them, the generative module adopts joint expression learning which integrates network protocol knowledge to improve the expressiveness and discriminability of traffic cognitive features. The equivalent module guides the discriminant module to learn the pragmatic relevance features through the activity loss function of traffic and the application loss function of protocol for end‐to‐end training. The experimental results show that WCCGAN can directly synthesize traffic with comprehensive concealment ability, and its behavior concealment and cognitive deception are as high as 86.2% and 96.7%, respectively. Moreover, the model has good convergence and generalization ability and does not depend on specific assumptions and specific covert algorithms, which realizes a new paradigm of cognitive game in covert communication. Zhangguo Tang, Junfeng Wang 0003, Huanzhou Li, Jian Zhang 0056 |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | Byte-level malware classification based on markov images and deep learning
Baoguo Yuan, Junfeng Wang 0003, Peng Wu 0036, Xuhua Bao |
Comput. Secur. | 2 |
| 2020 | Detection of Fake IoT App Based on Multidimensional SimilarityabstractWith low cost and high profit, fake IoT apps are an increasing risk to the security of the IoT ecosystem. In this article, we propose a novel fake IoT app detection method, referred as MSimDroid, based on multidimensional similarity to mitigate the threat. MSimDroid focuses on the distribution channels of fake apps, that is, app markets, and it consists of whole app similarity, resource similarity, code similarity, and their joint strategy. For similarity calculation, we design a distinctive algorithm based on the feature of different fake patterns. For joint strategy, which is the scheduler of multiple algorithms, it balances the accuracy and time consuming of MSimDroid. Experiments demonstrate that the accuracy of MSimDroid is more than 99.31% on ground-truth data set and 97.43% in the wild. The IoT apps from multiple well-known app markets reveal that the average proportion of fake apps is about 14.66%, and that of mixed-mode apps (including IoT and nonIoT apps) is 10.78%. Besides, it finds that about 0.58% of IoT apps suffer from malice, while the average ratio of mixed-mode apps is 1.06%. Peng Wu 0036, Junfeng Wang 0003, Baoguo Yuan, Wenyuan Kuang |
IEEE Internet Things J. | 3 |
| 2020 | Plagiarism detection in students' programming assignments based on semantics: multimedia e-learning based smart assessment methodology
Farhan Ullah 0001, Junfeng Wang 0003, Sohail Jabbar, Zhiming Wu, Shehzad Khalid |
Multim. Tools Appl. | 2 |
| 2019 | Data-flow bending: On the effectiveness of data-flow integrity
Junfeng Wang 0003 |
Comput. Secur. | 2 |
| 2019 | AeroMRP: A Multipath Reliable Transport Protocol for Aeronautical Ad Hoc NetworksabstractIn order to support Internet of Things, an aeronautical ad hoc network (AANET) has been proposed to solve the problem of air information isolated island. AANET is a special type of heterogeneous ad hoc network between commercial aircraft to share data and in-flight Internet access. Today’s civil airliners are typically equipped with multiple advanced network interfaces, which can form multiple independent communication paths between the source and the destination nodes in AANETs. Thus, users expect to utilize multiple networking interfaces of airliners simultaneously to support data transfer services. In this paper, we present a multipath reliable transport protocol for AANETs, called AeroMRP, which exploits path diversity provided by heterogeneous aeronautical networks. AeroMRP takes advantage of raptor codes to restore from packet loss in the aeronautical network environments while reducing the impact of the head-of-line blocking issue in the networks with multiple paths. A redundancy rate selection algorithm is adopted to determine the redundancy on the fly by considering the characteristics of the target applications and the time-varying aeronautical networks. Moreover, AeroMRP implements feedback-based packet scheduling considering not only aeronautical network status but also the acknowledgment scheme in congestion control for better use of different paths. We evaluate the performance of AeroMRP by ns-3 under a variety of simulation scenarios. The simulations demonstrate that AeroMRP can provide efficient and reliable data transmission services in heterogeneous multipath AANETs. Junfeng Wang 0003 |
IEEE Internet Things J. | 2 |
| 2019 | A QoE-perceived screen updates transmission scheme in desktop virtualization environment
Hongdi Zheng, Junfeng Wang 0003 |
Multim. Tools Appl. | 3 |
| 2018 | A New Software Birthmark based on Weight Sequences of Dynamic Control Flow Graph for Plagiarism DetectionabstractWith the large-scale development of open source software, software plagiarism has become a serious threat to software industry and intellectual property. As the latest technique of plagiarism detection, dynamic software birthmark has attracted much attention in recent years. Most of the existing dynamic birthmarks focus on how to resist obfuscation techniques such as compiler optimizations and strong obfuscations implemented in tools. However, they pay little attention to packers, especially encryption packer which is commonly used in software protection as well as plagiarism. When used to encrypt software, the decryption code is added to the binary. It is hard to distinguish the original parts of software from the decryption parts using traditional dynamic birthmarks. In this paper, we propose a new dynamic software birthmark called weight sequences birthmark (WSB) which is based on weight sequences of dynamic control flow graph (DCFG). The weight sequences are used as characteristics, which make full use of the different patterns of dynamic basic block replications between the original code and the decryption code. Compared with the-state-of-art dynamic key instruction sequence birthmark (DKISB), the new birthmark can resist encryption packer effectively. Furthermore, WSB shows better credibility than DKISB when distinguishing independent programs. The comprehensive experiments illustrate that the value of extended F-measure can reach 96.8%, indicating that it is a high-quality birthmark which satisfies both the credibility and the resiliency. Baoguo Yuan, Junfeng Wang 0003, Zhiyang Fang |
Comput. J. | 2 |
| 2018 | Elastically Reliable Video Transport Protocol Over Lossy Satellite LinksabstractAs there is a growing demand for providing high-definition (HD) video and satellite TV services over satellite networks, interconnection and multimedia applications need new-type satellite services to ensure good quality-of-experience (QoE) for these Internet flows. The characteristics of satellite channels (notably the lossy and long-delay environment) can heavily influence transferring protocols exclusively designed for terrestrial network, and using a radically reliable method (TCP) or unreliable method (UDP) is not the best way for the transmission of multimedia applications over satellite networks, such as HD video. In order to cope with the contradiction between high quality and real-time transfer, this paper presents a retransmission-based partially reliable transfer protocol, called automatically partially reliable transfer protocol (APRT). For the purpose of tracing the network status, the hidden Markov model (HMM) is employed to depict reliable level of the transfer strategy. Through off-line training initialization and online prediction, current network status is then mapped to a reliable level, which can represent good video quality and minimum packet delay. Meanwhile, APRT can retransmit the lost packets as the network changing without losing key frames. Extensive simulation results show that for both single and coexisting flow scenarios, APRT can improve video quality of services (QoSs) against reliable and unreliable transfer strategies in terms of different metrics. For single video flow, PSNR of APRT outperforms than that of other strategies under lossy and long-delay networks up to 36.53% and achieves short packet delay. For coexisting scenario, APRT shows good performance in terms of latency and video quality. Junfeng Wang 0003 |
IEEE J. Sel. Areas Commun. | 2 |
| 2018 | FRUDP: A Reliable Data Transport Protocol for Aeronautical Ad Hoc NetworksabstractIn this paper, we cope with the reliable and efficient data transfer problem in aeronautical ad hoc networks (AANETs). AANETs are different from terrestrial ad hoc networks and these distinctions feature AANETs with large propagation delay, highly dynamic network topology, and high error probability, which pose many new challenges for data transfer in such environments. In this paper, we propose a protocol, called FRUDP, which combines reliable UDP and fountain code to achieve reliable and efficient data transfer in AANETs. It employs Raptor codes to recover from data losses to mitigate retransmission caused by high channel losses in AANETs. A code rate determining algorithm is developed to estimate the redundancy on the fly by taking into full consideration the time-varying aeronautical network conditions. Moreover, the improved selective negative acknowledgment scheme is utilized to solve the bandwidth asymmetry problem in AANETs and a congestion control scheme is deployed to decouple congestion decisions from packet losses in order to avoid the erroneous congestion decisions due to frequent link handoffs. The proposed FRUDP is fully implemented directly on top of UDP in Linux and examined over a variety of emulation network environments. The experimental results show that FRUDP is suitable for reliable and efficient data transfer service in AANETs. Junfeng Wang 0003 |
IEEE J. Sel. Areas Commun. | 2 |
| 2018 | Predicting Long-Term Trajectories of Connected Vehicles via the Prefix-Projection TechniqueabstractThe vehicle location prediction based on their spatial and temporal information is an important and difficult task in many applications. In the last few years, devices, such as connected vehicles, smart phones, GPS navigation systems, and smart home appliances, have amassed the large stores of geographic data. The task of leveraging this data by employing moving objects database techniques to predict spatio-temporal locations in an accurate and efficient fashion, comprising a complete trajectory remains an actively researched area. Existing methods for frequent sequential pattern mining tend to be limited to predicting short-term partial trajectories, at extremely high computational costs. In order to address these limitations, we designed a prefix-projection-based trajectory prediction algorithm called PrefixTP, which contains three essential phases. First, data collection, connected vehicles equipped with sensors comprise a vehicle grid and generate copious amounts of spatio-temporal data, in order to communicate and share traffic information. Second, model training, examining only the prefix subsequences, and projecting only their corresponding postfix subsequences into projected sets. Finally, trajectory matching, recursively finding postfix sequences meeting the requirement of minimum support count, and outputting the most frequent sequential pattern as the most probable trajectory. Fundamentally, PrefixTP supports three trajectory matching strategies which encompass all possibilities of prediction. Extensive experiments were conducted using real world GPS data sets, and the results show, when comparing predicted complete trajectories against partial short-term trajectories with a guarantee of real-time forecasting, that PrefixTP outperforms first-order, second-order Markov models, and Apriori-based trajectory prediction algorithm. Shaojie Qiao, Nan Han, Junfeng Wang 0003, Rong-Hua Li 0001, Louis Alberto Gutierrez, Xindong Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Online autogenerated congestion control for high-speed transfer over high BDP networksabstractWith the exponential growth of network bandwidth and application requirements, the existing transmission control protocols have led to the issues of efficiency and applicability. In order to solve the network utility issues and accelerate the data delivery legitimately, the swift and self‐learned transmission methods have attracted attention gradually due to its adaptability and predictability. In this research, we propose a high‐speed transport protocol, called Hita, based on network performance learning framework to cope with high‐speed transfer challenge over high bandwidth delay product networks. The key idea of Hita is to select preferable network performance metrics to reflect network property variability and build corresponding window control model. The proposed protocol, which adopts the online‐learning methods, establishes the relation model between the transmission performance metrics and the corresponding congestion window. Afterwards, by determining the direction of window adjustment, the optimal sending window size can be predicted and approached quickly. Numerical results show that Hita can use the limited bandwidth more adequately. For simulation experiments, Hita achieves higher throughput while maintaining a lower packet delay comparing with other protocols. Moreover, Hita shows good performance in terms of intra‐protocol fairness, friendliness and protocol stability. For real‐world network experiments, Hita achieves more throughput than the other high‐speed protocols as well. Junfeng Wang 0003 |
IET Commun. | 2 |
| 2017 | Multiple QoS Parameters-Based Routing for Civil Aeronautical Ad Hoc NetworksabstractAeronautical ad hoc network (AANET) can be applied as in-flight communication systems to allow aircraft to communicate with the ground, in complement to other existing communication systems to support Internet of Things. However, the unique features of civil AANETs present a great challenge to provide efficient and reliable data delivery in such environments. In this paper, we propose a multiple quality of service parameters-based routing protocol (MQSPR), to improve the overall network performance for communication between aircraft and the ground. The proposed MQSPR integrates path availability period, residual path load capacity and path latency in route selection and presents a broadcast optimization scheme to minimize flooding. The primary aim of MQSPR is to maintain long link durations, achieve path load balancing and reduce end-to-end delay to satisfy the requirements of civil aviation communication services. The simulation scenario and real-world scenario are set up, respectively, and the experimental results show that the proposed MQSPR can achieve high ground connectivity while effectively increase the path durations, improve the packet delivery ratio and perform the path load balancing. In addition, the flexibility of MQSPR is demonstrated by considering weighting factors of path selection parameters. Junfeng Wang 0003 |
IEEE Internet Things J. | 2 |
| 2016 | Adaptive low-priority transfer for high bandwidth and delay product networksabstractAs the exponential growth of the Internet, there is an increasing need to provide different types of services for numerous applications. Among these services, low-priority data transfer across wide area network has attracted much attention and has been used in a number of applications, such as data backup and system updating. Although the design of low-priority data transfer has been investigated adequately in low speed networks at transport layer, it becomes more challenging for the design of low-priority data transfer with the adaptation to high bandwidth delay product networks than the previous ones. This paper proposes an adaptive low-priority protocol to achieve high utilization and fair sharing of links in high bandwidth delay product networks, which is implemented at transport layer with an end-to-end approach. The designed protocol implements an adaptive congestion control mechanism to adjust the congestion window size by appropriate amount of spare bandwidth. The improved congestion mechanism is intent to make as much use of the available bandwidth without disturbing the regular transfer as possible. Experiments demonstrate that the adaptive low-priority protocol achieve efficient and fair bandwidth utilization, and remain non-intrusive to high priority traffic. Copyright © 2016 John Wiley & Sons, Ltd. Junfeng Wang 0003, Sunyoung Han |
Wirel. Commun. Mob. Comput. | 2 |
| 2015 | A static Android malicious code detection method based on multi-source fusionabstractThe rapid development of mobile malwares makes the traditional signature-based and single-feature based malware detection methods a challenging task. The surge of new malwares with more complex structures and dynamic characteristics leads to efficient fusion of multi-source malicious information more difficult in detection. In this paper, we propose a new multi-source based method to detect Android malwares by emphasizing on the traditional static features, control flow graph, and repacking characteristics. Each category of features is treated as an independent information source in feature extracting rules building and classification. Then, the Dempster-Shafer algorithm is used to fuse these information sources. This method can improve accuracy of malware detection without adding too many instability characteristics that are extracted from disassembled codes, and have better performance in the resistance to code obfuscation technologies. To verify our method, different categories of apps are collected to build the dataset in our experiment. Based on the dataset, our method can achieve 97% detection accuracy and 1.9% false positive rate. Copyright © 2015John Wiley & Sons, Ltd. Yao Du 0004, Junfeng Wang 0003 |
Secur. Commun. Networks | 3 |
| 2014 | Adaptive congestion control framework and a simple implementation on high bandwidth-delay product networks
Junfeng Wang 0003, Sunyoung Han |
Comput. Networks | 2 |
| 2011 | A Zone-Diffusion Based Routing Protocol for LEO Satellite Networks
Junfeng Wang 0003 |
WASA | 2 |
| 2008 | Segment-based adaptive hyper-Erlang model for long-tailed network traffic approximation
Junfeng Wang 0003, Chundong She |
J. Supercomput. | 1 |