Yong Fang 0002

dblp:41/3085-2 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 14 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 An effective and stealthy XSS adversarial sample generation method against deep learning detection models
Yong Fang 0002, Yaochang Xu, Yijia Xu
Neurocomputing2
2026 MPS-Fuzz: An Enhanced Fine-Grained Fuzzing Based on Units With Multiple Inputs and Outputs
abstract
Edge coverage-guided fuzzing has demonstrated remarkable achievements in vulnerability discovery. Some studies with fine-grained coverage metrics have been proposed to enhance the vulnerability mining capabilities of fuzzing by capturing more program paths. However, this refinement often results in a significant increase in seeds, which are highly homogeneous and may limit vulnerability detection. Additionally, finer granularity requires more bitmap hits, increasing the risk of hash collisions. To address these shortages, the paper proposes the structure of a basic block unit with multiple predecessors and successors (referred to as MPS). Then, a fine-grained coverage method called MPS-Fuzz is designed based on the MPS structure. In this approach, it is convenient to exclude basic blocks involving loop structures when determining MPS units, which helps reduce seed homogeneity. Additionally, we introduce an additional bitmap to record the coverage status of MPS units, ensuring that the collision rate of the edge bitmap does not increase. Moreover, these additional operations do not incur excessive time overhead. To demonstrate the properties of the MPS-Fuzz, we implement our approach on AFL and conduct experiments on 16 benchmarks from FuzzBench and Unifuzz. The result indicates that, after 24-hour fuzzing, MPS-Fuzz explores an average of 9.6% more edges and an average of 25.7% more bugs than AFL. Compared to other fine-grained coverage methods (N-gram and PathAFL), MPS-Fuzz also achieves better performance. Moreover, MPS-Fuzz has discovered a previously unknown bug on real-world program and got a CVE assigned.
Ximing Fan, Yong Fang 0002, Peng Jia 0005, Hongwei Li 0001, Yijia Xu, Qinying Wang, Shouling Ji
IEEE Trans. Dependable Secur. Comput.2
2026 Web Page Tampering Detection Based on Dynamic Temporal Graph Pre-Training
abstract
Web page tampering detection is crucial in web threat perception. Current methods rely on monitoring historical changes of web pages to identify anomalies. These approaches often struggle to effectively distinguish between tampering and benign changes, especially in the presence of numerous dynamic pages. Furthermore, the increasing complexity of website structures places more resource demands on tampering monitoring and makes some malicious alterations more covert and challenging to detect. We propose a web page tampering detection based on pretraining with dynamic temporal graphs. The core of the method involves constructing a website temporal graph model based on evolutionary information, and enhances the graph feature perturbations to expose concealed tampering behaviors. Specifically, the framework's autoencoder is composed of enhanced DySAT, enabling it to handle dynamic data. We introduce DySAT, bolstered with GATv2, to capture dynamic attention. Additionally, we design a temporal masking mechanism and prediction error to improve the effectiveness of generative self-supervised learning in temporal graph pretraining. Experimental results on Webpage Tampering Dataset (WPT-Dataset) demonstrate that our method outperforms other comparative approaches in terms of both detection efficacy and stability. Furthermore, the research findings on the anomaly detector and model performance provide direction for the practical application of our method.
Yijia Xu, Qiang Zhang 0057, Zhonglin Liu, Cheng Huang 0003, Yong Fang 0002
IEEE Trans. Dependable Secur. Comput.6
2026 One Trigger, Multiple Victims: Clean-Label Neighborhood Backdoor Attacks on Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have achieved remarkable success in modeling structured data. Recent studies, however, reveal that they are highly vulnerable to backdoor attacks, which can implant triggers into training data to mislead predictions on nodes injected with triggers while maintaining accuracy on clean inputs. Despite recent advances, existing graph backdoor attacks often rely on explicit training interventions and substantial trigger injection while focusing solely on single-node misclassification, which limits their practicality in real-world deployments. To address these limitations, we propose a clean-label graph backdoor attack that induces one-hop neighborhood misclassification under a minimal trigger injection budget. Without altering target nodes’ features or labels, our method attaches a single trigger node to a target node, thereby misclassifying both the target and its immediate neighbors as the target class. To maximize effectiveness while preserving stealthiness, we propose a poisoned node selection strategy guided by semantic consistency and structural activeness, and design a conditional diffusion-based trigger generator optimized with multiple auxiliary objectives. Extensive experiments on multiple real-world benchmarks and mainstream GNN architectures show that our approach achieves over 95% attack success rate on both target nodes and their neighbors in most settings, including under state-of-the-art defenses. These findings underscore the urgent need for more robust graph learning systems and reveal novel attack surfaces in graph security.
Huaxin Deng, Yong Fang 0002, Qiang Zhang 0057, Yang Liu 0003, Yijia Xu
IEEE Trans. Inf. Forensics Secur.2
2025 Agent Behavior: The Regulatory Object of the Agent-Centric Online Ecosystem in Digital Age
Qiang Zhang 0057, Pei Yan, Yijia Xu, Xinfeng Li, Hongyi Cai, Chuanpo Fu, Yong Fang 0002, Yang Liu 0003
ICECCS7
2025 Directed fuzzing based on path constraints and deviation path correction
Hongsheng Zuo, Yong Fang 0002, Peng Jia 0005, Ximing Fan, Yijia Xu
Inf. Softw. Technol.2
2024 Few-shot graph classification on cross-site scripting attacks detection
Hongyu Pan, Yong Fang 0002, Wenbo Guo 0011, Yijia Xu, Changhui Wang
Comput. Secur.2
2024 Multi-target label backdoor attacks on graph neural networks
abstract
Graph neural networks have been shown to have characteristics that make them susceptible to backdoor attacks, and many recent works have proposed feasible graph backdoor attack methods. However, existing graph backdoor attack methods only target one-to-one attack types and lack graph backdoor attack methods that can address one-to-many attack requirements. This paper is the first research work on one-to-many type graph backdoor attacks and proposes the backdoor attack method MLGB, which can achieve multi-target label attacks for GNN node classification tasks. We designed encoding mechanisms to allow MLGB to customize triggers for different target labels and ensure differentiation between triggers for different target labels through loss functions. Additionally, we designed an innovative poisoned node selection method to improve the efficiency of MLGB’s attacks further. Extensive experiments were conducted to validate MLGB’s effectiveness across multiple datasets and model architectures, demonstrating its robustness against graph backdoor attack defense mechanisms. Furthermore, ablation experiments and explainability analyses were conducted to provide deeper insights into MLGB. Our work reveals that graph neural networks are also vulnerable to one-to-many type backdoor attacks, which is important for practitioners to understand model risks comprehensively.
Huaxin Deng, Yijia Xu, Zhonglin Liu, Yong Fang 0002
Pattern Recognit.5
2024 Graph Mining for Cybersecurity: A Survey
abstract
The explosive growth of cyber attacks today, such as malware, spam, and intrusions, has caused severe consequences on society. Securing cyberspace has become a great concern for organizations and governments. Traditional machine learning based methods are extensively used in detecting cyber threats, but they hardly model the correlations between real-world cyber entities. In recent years, with the proliferation of graph mining techniques, many researchers have investigated these techniques for capturing correlations between cyber entities and achieving high performance. It is imperative to summarize existing graph-based cybersecurity solutions to provide a guide for future studies. Therefore, as a key contribution of this work, we provide a comprehensive review of graph mining for cybersecurity, including an overview of cybersecurity tasks, the typical graph mining techniques, and the general process of applying them to cybersecurity, as well as various solutions for different cybersecurity tasks. For each task, we probe into relevant methods and highlight the graph types, graph approaches, and task levels in their modeling. Furthermore, we collect open datasets and toolkits for graph-based cybersecurity. Finally, we present an outlook on the potential directions of this field for future research.
Bo Yan 0005, Cheng Yang 0002, Chuan Shi 0001, Yong Fang 0002, Qi Li 0057, Yanfang Ye 0001, Junping Du 0001
ACM Trans. Knowl. Discov. Data4
2023 An Empirical Study of Malicious Code In PyPI Ecosystem
abstract
PyPI provides a convenient and accessible package management platform to developers, enabling them to quickly implement specific functions and improve work efficiency. However, the rapid development of the PyPI ecosystem has led to a severe problem of malicious package propagation. Malicious developers disguise malicious packages as normal, posing a significant security risk to end-users. To this end, we conducted an empirical study to understand the characteristics and current state of the malicious code lifecycle in the PyPI ecosystem. We first built an automated data collection framework and collated a multi-source malicious code dataset containing 4,669 malicious package files. We preliminarily classified these malicious code into five categories based on malicious behaviour characteristics. Our research found that over 50 % of malicious code exhibits multiple malicious behaviours, with information stealing and command execution being particularly prevalent. In addition, we observed several novel attack vectors and anti-detection techniques. Our analysis revealed that 74.81 % of all malicious packages successfully entered end-user projects through source code installation, thereby increasing security risks. A real-world investigation showed that many reported malicious packages persist in PyPI mirror servers globally, with over 72 % remaining for an extended period after being discovered. Finally, we sketched a portrait of the malicious code lifecycle in the PyPI ecosystem, effectively reflecting the characteristics of malicious code at different stages. We also present some suggested mitigations to improve the security of the Python open-source ecosystem.
Wenbo Guo 0011, Zhengzi Xu, Cheng Huang 0003, Yong Fang 0002, Yang Liu 0003
ASE5
2023 MFXSS: An effective XSS vulnerability detection method in JavaScript based on multi-feature model
Zhonglin Liu, Yong Fang 0002, Cheng Huang 0003, Yijia Xu
Comput. Secur.2
2023 PWAGAT: Potential Web attacker detection based on graph attention network
Yijia Xu, Yong Fang 0002, Zhonglin Liu, Qiang Zhang 0057
Neurocomputing2
2022 Web Attack Payload Identification and Interpretability Analysis Based on Graph Convolutional Network
abstract
Web attack payload identification is a significant part of the Web defense system. The current Web attack payload identification usually combines natural language processing and deep learning to automatically build a detection model to intercept malicious payloads. However, these detection methods ignore the bidirectional association between fields and is prone to the payload dilution problem for long strings. In addition, the weak interpretability of deep learning models makes it difficult for researchers to solve the problem of model pollution and adjust the model according to the prediction logic. Therefore, this paper proposes a new Web attack payload identification method based on Graph Convolutional Network (GCN), which can effectively extract Web payload features and help model interpretability analysis. The core of this method is to transform the text feature problem into a graph feature extraction problem and to understand the structure and content of the Web payload from the graph perspective. The method performs node embedding on the Web payload graph through GCN, then converts the embedding vector into a graph feature vector through a feature fusion method. The node ablation method is used to analyze malicious payloads' interpretability and calculate the predicted impact rate of nodes inside the graph structure. The experiments on the CSIC 2010 v2 HTTP dataset show that the method proposed in this paper has high accuracy for identifying Web attack payloads, and the node embedding of the Relational Graph Convolutional Network (RGCN) method is more suitable for identifying Web attack payloads than other GCN methods. The research results of the paper show that the model interpretability analysis based on the Web payload graph is reasonable and can effectively assist researchers in adjusting the model and preventing the problem of model pollution.
Yijia Xu, Yong Fang 0002, Zhonglin Liu
MSN2
2022 Viopolicy-Detector: An Automated Approach to Detecting GDPR Suspected Compliance Violations in Websites
abstract
To provide users with personalized services, the website collects and tracks user’s activity data. At the same time, each website uses a privacy policy to ensure the legality of these actions. The purpose of the implementation of the General Data Protection Regulation (GDPR) is to protect the privacy of user data. Because GDPR is a programmatic regulation, there is no specific guidance on what a privacy policy should contain. Therefore, there may still be potential violations on the website, thus cause a risk of leak users’ private data. In this paper, we define a violating behavior that data collected by the website without a declaration in the privacy policy is illegal. To complete the violating behavior detection, we first interpret the GDPR and analyze 1000 website privacy policies to present a personal data classification including eight categories. Based on this, we propose a privacy policy annotation scheme including these eight categories and collect 145 related Web APIs. Then we propose an automated method to detect GDPR suspected compliance violations in websites. On the one hand we use the multi-label text classification model to extract data collection stated in the privacy policy, with a precision of 0.9817. For another, we dynamically monitor the JavaScript calls of the website related to personal data collection during user visits. Finally, we compare the two results to determine whether violating behaviors appeared. We use this method to detect the European top 500 websites (actually 451 websites). A total of 159 (35.3%) websites appear in violation of the GDPR. We analyze the detection results from different perspectives, including statistics on the types of data declared in the privacy policy, statistics on data collected by the website, and which data collection is likely to cause violations. Then we classify the violating websites and find that websites in the Social category present the most violations. Finally, we count the rankings of the offending websites. Surprisingly, top-ranking sites are even more prone to breaches. There are even some globally well-known websites with violations, such as BBC, Nokia, Ebay, Google etc.
Haoran Ou, Yong Fang 0002, Yongyan Guo, Wenbo Guo 0011, Cheng Huang 0003
RAID2
2022 JStrong: Malicious JavaScript detection based on code semantic representation and graph neural network
Yong Fang 0002, Chaoyi Huang, Minchuan Zeng, Zhiying Zhao, Cheng Huang 0003
Comput. Secur.1
2022 HyVulDect: A hybrid semantic vulnerability mining system based on graph neural network
Wenbo Guo 0011, Yong Fang 0002, Cheng Huang 0003, Haoran Ou, Chun Lin, Yongyan Guo
Comput. Secur.2
2022 GraphXSS: An efficient XSS payload detection approach based on graph convolutional network
Zhonglin Liu, Yong Fang 0002, Cheng Huang 0003, Jiaxuan Han
Comput. Secur.2
2022 LMTracker: Lateral movement path detection based on heterogeneous graph embedding
Yong Fang 0002, Congshuang Wang, Zhiyang Fang, Cheng Huang 0003
Neurocomputing1
2022 HGHAN: Hacker group identification based on heterogeneous graph attention network
Yijia Xu, Yong Fang 0002, Cheng Huang 0003, Zhonglin Liu
Inf. Sci.2
2022 Embedding vector generation based on function call graph for effective malware detection and classification
Xiao-Wang Wu, Yong Fang 0002, Peng Jia 0005
Neural Comput. Appl.3
2021 No Pie in the Sky: The Digital Currency Fraud Website Detection
Haoran Ou, Yongyan Guo, Chaoyi Huang, Zhiying Zhao, Wenbo Guo 0011, Yong Fang 0002, Cheng Huang 0003
ICDF2C6
2021 CyberEyes: Cybersecurity Entity Recognition Model Based on Graph Convolutional Network
abstract
Abstract Cybersecurity has gradually become the public focus between common people and countries with the high development of Internet technology in daily life. The cybersecurity knowledge analysis methods have achieved high evolution with the help of knowledge graph technology, especially a lot of threat intelligence information could be extracted with fine granularity. But named entity recognition (NER) is the primary task for constructing security knowledge graph. Traditional NER models are difficult to determine entities that have a complex structure in the field of cybersecurity, and it is difficult to capture non-local and non-sequential dependencies. In this paper, we propose a cybersecurity entity recognition model CyberEyes that uses non-local dependencies extracted by graph convolutional neural networks. The model can capture both local context and graph-level non-local dependencies. In the evaluation experiments, our model reached an F1 score of 90.28% on the cybersecurity corpus under the gold evaluation standard for NER, which performed better than the 86.49% obtained by the classic CNN-BiLSTM-CRF model.
Yong Fang 0002, Yuchi Zhang, Cheng Huang 0003
Comput. J.1
2021 A communication-channel-based method for detecting deeply camouflaged malicious traffic
Yong Fang 0002, Rongfeng Zheng, Shan Liao
Comput. Networks1
2021 Effective method for detecting malicious PowerShell scripts based on hybrid features☆
Yong Fang 0002, Cheng Huang 0003
Neurocomputing1
2021 PBDT: Python Backdoor Detection Model Based on Combined Features
abstract
Application security is essential in today’s highly development period. Backdoor is a means by which attackers can invade the system to achieve illegal purposes and damage users’ rights. It has posed a serious threat to network security. Thus, it is urgent to take adequate measures to defend such attacks. Previous research work was mainly focused on numerous PHP webshells, with less research on Python backdoor files. Language differences make the method not entirely applicable. This paper proposes a Python backdoor detection model named PBDT based on combined features. The model summarizes the common functional modules and functions in the backdoor files and extracts the number of calls in the text to form sample features. What is more, we consider the text’s statistical characteristics, including the information entropy, the longest string, etc., to identify the obfuscated Python code. Besides, the opcode sequence is used to represent code characteristics, such as TF-IDF vector and FastText classifier, to eliminate the influence of interference items. Finally, we introduce the Random Forest algorithm to build a classifier. Covering most types of backdoors, some samples are obfuscated, the model achieves an accuracy of 97.70%, and the TNR index is as high as 98.66%, showing a good classification performance in Python backdoor detection.
Yong Fang 0002, Mingyu Xie, Cheng Huang 0003
Secur. Commun. Networks1
2020 EmailDetective: An Email Authorship Identification And Verification Model
abstract
Abstract Emails are often used to illegal cybercrime today, so it is important to verify the identity of the email author. This paper proposes a general model for solving the problem of anonymous email author attribution, which can be used in email authorship identification and email authorship verification. The first situation is to find the author of an anonymous email among the many suspected targets. Another situation is to verify if an email was written by the sender. This paper extracts features from the email header and email body and analyzes the writing style and other behaviors of email authors. The behaviors of email authors are extracted through a statistical algorithm from email headers. Moreover, the author’s writing style in the email body is extracted by a sequence-to-sequence bidirectional long short-term memory (BiLSTM) algorithm. This model combines multiple factors to solve the problem of anonymous email author attribution. The experiments proved that the accuracy and other indicators of proposed model are better than other methods. In email authorship verification experiment, our average accuracy, average recall and average F1-score reached 89.9%. In email authorship identification experiment, our model’s accuracy rate is 98.9% for 10 authors, 92.9% for 25 authors and 89.5% for 50 authors.
Yong Fang 0002, Cheng Huang 0003
Comput. J.1
2020 Detecting malicious JavaScript code based on semantic analysis
Yong Fang 0002, Cheng Huang 0003, Yaoyao Qiu
Comput. Secur.1
2017 Gossip: Automatically Identifying Malicious Domains from Mailing List Discussions
abstract
Domain names play a critical role in cybercrime, because they identify hosts that serve malicious content (such as malware, Trojan binaries, or malicious scripts), operate as command-and-control servers, or carry out some other role in the malicious network infrastructure. To defend against Internet attacks and scams, operators widely use blacklisting to detect and block malicious domain names and IP addresses. Existing blacklists are typically generated by crawling suspicious domains, manually or automatically analyzing malware, and collecting information from honeypots and intrusion detection systems. Unfortunately, such blacklists are difficult to maintain and are often slow to respond to new attacks. Security experts set up and join mailing lists to discuss and share intelligence information, which provides a better chance to identify emerging malicious activities. In this paper, we design Gossip, a novel approach to automatically detect malicious domains based on the analysis of discussions in technical mailing lists (particularly on security-related topics) by using natural language processing and machine learning techniques. We identify a set of effective features extracted from email threads, users participating in the discussions, and content keywords, to infer malicious domains from mailing lists, without the need to actually crawl the suspect websites. Our result shows that Gossip achieves high detection accuracy. Moreover, the detection from our system is often days or weeks earlier than existing public blacklists.
Cheng Huang 0003, Shuang Hao 0001, Luca Invernizzi, Yong Fang 0002, Christopher Krügel, Giovanni Vigna
AsiaCCS5
2016 A study on Web security incidents in China by analyzing vulnerability disclosure platforms
Cheng Huang 0003, Yong Fang 0002, Zheng Zuo
Comput. Secur.3