Peian Yang

dblp:260/3151 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
19since 2021 · last 2026
0009-0004-7226-8177ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MGDA: A provenance graph-based framework for threat detection and attack scenario reconstruction
abstract
Advanced persistent threat (APT) attacks are sophisticated, stealthy, and persistent, posing significant challenges to timely detection and investigation in modern network environments. Provenance graph analysis has become an important method for APT detection due to its ability to capture detailed causal relationships among system entities. However, existing methods suffer from several limitations: (1) lack of labeled attack data, (2) lack of high-level semantics in attack scenario reconstruction, and (3) high computational overhead limiting practical deployment. In this paper, we propose MGDA, a self-supervised method for effective and accurate threat detection as well as interpretable attack scenario reconstruction. MGDA introduces a multi-view masked graph autoencoder that jointly captures deep semantic features and structural patterns, enabling accurate detection of stealthy and unknown attacks. In the reconstruction phase, MGDA combines contextual analysis with rule-based attack pattern matching to produce attack scenario graphs that incorporate high-level semantics. We evaluate MGDA on three widely used datasets, including both real-world and simulated network attacks. The results demonstrate that MGDA achieves an average precision of 97.58% and F1-score of 98.03% in threat detection, outperforming state-of-the-art approaches. In addition, the automatically reconstructed scenario graphs help identify potential multi-step attacks and their stages, aiding analysts in conducting efficient network attack investigations.
Mengjiao Cui, Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng
Comput. Networks6
2025 AGLHunter: Automated Threat Hunting Using In-Context Learning-Enhanced LLM
abstract
Advanced Persistent Threats (APTs) are characterized by their persistence, sophistication, and stealth, posing significant challenges to network detection. Existing research on attack detection leveraging Provenance Graphs (PGs) has proven effective in correlating system entities and capturing persistence. However, the exponential growth of audit logs makes large-scale data storage and processing difficult. In addition, current threat hunting methods rely heavily on manually crafted attack query graphs, which are limited by expert knowledge and lack automated solutions. In this paper, we propose AGLHunter, an automated threat hunting system designed to enhance automation and efficiency while maintaining high detection accuracy. Our system leverages the In-Context Learning (ICL) capability of the Large Language Model (LLM) to automatically construct query graphs from Cyber Threat Intelligence (CTI) reports. Next, we extract suspicious subgraphs from PGs and employ graph representation learning to match these sub graphs with the query graphs, enabling efficient and accurate threat hunting. We use DARPA TC and OpTC datasets to evaluate AGLHunter's performance. The results show that AGLHunter not only achieves higher automation but also shows superior performance with reduced memory usage. AGLHunter, leveraging ICL-enhanced LLM, improved the F1 score for query graph construction by 13.6%, reduced the overall hunting time by more than 170 seconds, and maintained high detection accuracy.
Mengjiao Cui, Zhengwei Jiang, Yepeng Yao, Qiying He, Peian Yang, Huamin Feng
CSCWD6
2025 From Threat Report to ATT&CK: Automated Extraction and Reasoning of TTPs Using Large Language Models
abstract
The escalating frequency and increasing complexity of cyber attacks underscore the importance of Cyber Threat Intelligence (CTI). Tactics, Techniques, and Procedures (TTPs), as advanced CTI capable of characterizing adversarial behaviors and intentions, have garnered increased attention. However, TTPs are predominantly found embedded within unstructured natural language texts of threat reports. The accurate extraction and standardization of TTPs pose significant challenges. Existing methods exhibit limitations in terms of accuracy, generalizability, and interpretability. This paper presents a pipeline for automatically extracting TTPs from threat reports and providing rationales using large language models. To support this approach, we have developed three datasets using advanced commercial LLMs for data synthesis. These datasets are made publicly available to facilitate further research. Experimental results demonstrate the superior performance of our proposed approach, achieving an F1-score of 97.15% and accuracies of 79.22% and 92.97% in the respective tasks. These results surpass state-of-the-art methods by 15.39%, 12.87%, and 27.57%, respectively. To the best of our knowledge, this paper is the first to simultaneously extract TTPs while providing the underlying rationales for the extraction. This novel approach significantly improves the usability of the results by providing a richer context for threats.
Fangming Dong, Zhengwei Jiang, Qiying He, Peian Yang, Yepeng Yao
CSCWD5
2025 TIMFuser: A multi-granular fusion framework for cyber threat intelligence
Zhengwei Jiang, Kai Zhang 0035, Zhiting Ling, Yizhe You, Peian Yang, Huamin Feng
Comput. Secur.7
2024 TiGNet: Joint entity and relation triplets extraction for APT campaign threat intelligence
abstract
Contemporary cybersecurity faces escalating challenges from sophisticated threats, notably Advanced Persistent Threats (APTs). Addressing these challenges necessitates a collaborative, multidisciplinary approach that transcends traditional boundaries. Gathering cyber threat intelligence (CTI) on APT campaigns and constructing a comprehensive knowledge graph empowers defenders to track the latest trends in these campaigns, update defense strategies, and attain crucial advantages in defense measures. Previous works used relation extraction techniques to obtain entity-relation triplets for constructing threat intelligence knowledge graphs. However, these works either rely on pipeline workflow, are susceptible to exposure errors and error propagation, or use sequence annotation method, which combines entity and relation labels but lacks the ability to extract single entity overlaps (SEO) or subject-object overlaps (SOO) triplets. This paper introduces TiGNet, a novel method that transforms the entity-relation triplet’s extraction task into multiple token-span recognition tasks utilizing token-pair matrices. Additionally, we integrate GlobalPointer to incorporate token position information into the token-pair matrix, significantly enhancing extraction performance. To facilitate method evaluation, we annotated a Chinese entity-relation triplets dataset about APT campaigns, named APT-Triplets, comprising 9711 triplets encompassing seven triplet types. Our evaluation demonstrates that TiGNet improves the F1 score of 3.79-5.59 compared to previous joint extraction methods. Furthermore, it outperforms methods based on large language models (LLMs) in terms of both extraction performance and inference time. These results underscore TiGNet’s capacity to accurately and swiftly extract threat intelligence, facilitating the construction of the APT campaign knowledge graphs, empowering defenders to track evolving trends and fortify defense strategies collaboratively.
Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Huamin Feng, Peian Yang
CSCWD6
2024 CTIMiner: Cyber Threat Intelligence Mining Using Adaptive Multi-task Adversarial Active Learning
Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng
ICDF2C (1)5
2024 APTChaser: Cyber Threat Attribution via Attack Technique Modeling
Peian Yang, Zhengwei Jiang, Mengjiao Cui, Yizhe You
ICDF2C (1)2
2024 CTIFuser: Cyber Threat Intelligence Fusion via Unsupervised Learning Model
abstract
Cyber attack campaigns are becoming increasingly complex and severe, causing significant impacts on institutions and individuals. Cyber Threat Intelligence (CTI) provides important evidential knowledge about attackers and is critical to the shift from reactive to proactive defense against cyber attacks. Attack detection based on Indicators of Compromise (IOCs), a type of CTI, is vulnerable to the limitation of insufficient context of attack scenarios. In contrast, attack behavior intelligence is associated with information on attackers’ techniques, targets, and intentions, providing a solid foundation for security practitioners to conduct attack investigations or other applications. Many current CTI mining systems are limited to extracting CTI from a single source, leading to challenges such as fragmented attack behavior view and low-value density. To address these issues, we propose an unsupervised fusion framework named CTIFuser, which includes a comprehensive pipeline of four subtasks aimed at mining and fusing multi-source attack behaviors at the attack technique level. In our evaluation of 739 real-world CTI reports from 542 sources, experimental results demonstrate that CTIFuser can obtain a complete view of the attack behaviors at the attack technique level.
Zhengwei Jiang, Peian Yang, Mengjiao Cui, Fangming Dong, Huamin Feng
ISPA4
2024 P-TIMA: a framework of T witter threat intelligence mining and analysis based on a prompt-learning NER model
abstract
Abstract Open-source information platforms such as Twitter continuously provide the latest threat intelligence, including new vulnerabilities and in-the-wild exploitations of advanced persistent threat (APT) groups. Automated extraction of threat intelligence from Twitter has become crucial for defenders to access up-to-date threat knowledge. However, existing studies mainly rely on supervised learning methods to extract threat intelligence knowledge, such as entities, which require a large amount of annotated data. This paper presents Threat Intelligence Mining and Analysis based on Prompt Learning (P-TIMA), a framework specifically crafted for extracting and analyzing threat intelligence from Twitter. P-TIMA employs our innovative few-shot entity recognition method, SecEntPrompt (SEP), built on prompt learning, to extract vulnerability intelligence from Twitter. Additionally, P-TIMA analyzes and profiles the overarching vulnerability intelligence obtained from Twitter, along with in-the-wild exploitation intelligence of APT groups. The SEP improves the average entity recognition F1 score by 3.62-4.40 compared with the best-performing comparison model and outperforms the method based on the large language model on recognition performance and inference time. To validate our framework, we apply P-TIMA to extract vulnerability-related threat intelligence from real Twitter data. Through case studies, we then analyze trends in vulnerability threats and the exploitation capabilities of APT groups. In conclusion, our framework provides a more efficient and accurate method for extracting threat intelligence from Twitter, enabling defenders to stay up-to-date with the latest threat trends and helping them improve their defense strategies against cyber attacks.
Yizhe You, Zhengwei Jiang, Peian Yang, Kai Zhang 0035, Xuren Wang, Chenpeng Tu, Huamin Feng
Comput. J.3
2024 A survey of large language models for cyber threat detection
Mengjiao Cui, Yiyang Cao, Peian Yang, Bo Jiang 0013, Zhigang Lu 0002, Baoxu Liu
Comput. Secur.5
2023 Phishsifter: An Enhanced Phishing Pages Detection Method Based on the Relevance of Content and Domain
abstract
A Phishing website is used to steal users’ private information. The accelerated development of phishing kits has made it convenient to create such websites, which has become a persistent security threat. In this article, we propose a novel method to detect phishing webpages based on the relevance of the webpage content and domain. For phishing webpages whose domain is relevant to the content, we use the target identification method to identify the target brand. We use two components, the website logo and domain, to identify phishing sites, which increases the accuracy of identification. For irrelevant websites, we use a feature-based approach to distinguish phishing webpages. The experiment shows that the accuracy of target identification is 97.21%, while the false positive rate is 1.47%. The accuracy of the feature-based method is 98.32%. The proposed scheme can meet the needs of practical applications and provide an interpretation of the classification results.
Zhengwei Jiang, Zhiting Ling, Peian Yang
CSCWD6
2023 Simplified-Xception: A New Way to Speed Up Malicious Code Classification
abstract
Traditional malicious code detection methods require a lot of manpower and resources, which makes the research of malicious code very difficult. The selection of malicious code features mainly relies on the subjective analysis and selection of experts, which has a large impact on the detection effect of the model. In this paper, malicious codes are converted into greyscale images as model inputs, and features are automatically extracted using a deep-learning model. An improved convolutional neural network model based on Xception (Simplified Xception) is proposed for malicious code family classification. The model reduces the number of modules in the original model and adds a depth-separable convolutional layer with a step size of 2 to enhance the generated grey-scale images. The model is compared with CNN models, ResNet50, and improved models related to Inception. The experimental results show that the accuracy of SimplifiedXception is 98%, which is better than other related models. Compared to the Xception model, the accuracy of the Simplified-Xception model was improved by 1.3% and the number of parameters was reduced by half.
Xinshuai Zhu, Songheng He, Xuren Wang, Peian Yang, Yuxia Fu
CSCWD6
2023 FineCTI: A Framework for Mining Fine-grained Cyber Threat Information from Twitter Using NER Model
abstract
To timely respond to cyber threats related to a specific IT infrastructure called fine-grained (e.g., Windows or Linux), security analysts need to require timely and comprehensive threat information. Twitter, as a vital source of real-time threat information, provides abundant but overwhelming information due to the increased data sources. Automatically mining and summarizing fine-grained threat information from Twitter can help security analysts maintain the infrastructure’s security. Most existing studies focus on classification, which carries less threat information. Some works use clustering based on text similarity relying on the embedding of text obtained from pre-trained models, which cannot be applied to short text, resulting in noisy clusters. Several works build topic models. However, the incoherent topic keywords are difficult to understand and analyze. To overcome these challenges, we design a FineCTI framework to mine the threat information related to the specific infrastructure on Twitter and generate a detailed threat information summary that is machine-readable and human-readable, efficiently reducing information overload. FineCTI optimizes the feature extraction part based on the named entity recognition model and performs clustering based on features extracted, thus effectively reducing the influence of sparsity of tweets on the clustering result and with the V-measure score improved by 7%. The cluster analysis results show that we can mine the fine-grained threats up to 15 days before the official disclosure date.
Kai Zhang 0035, Zhengwei Jiang, Peian Yang, Xuren Wang, Huamin Feng
TrustCom5
2022 Cyber Threat Intelligence Entity Extraction Based on Deep Learning and Field Knowledge Engineering
abstract
The typical domain characteristics of Cyber Threat Intelligence (CTI), such as fuzzy entity boundary, polysemy or a single word corresponding to multiple word expressions and so on, makes the entity recognition result be worse than we expected. In addition, there are many challenges in directly migrating entity recognition models from the general field to CTI field. Therefore, we propose a deep learning entity recognition model with supplementing the domain knowledge engineering, which takes the open-source Cyber Threat Intelligence entity recognition as the research object, covering natural language processing, deep learning and cyber threat intelligence fields. Firstly, we use BERT model to obtain the dynamic word vector, then encode the word sequence by using BiLSTM-CRF, and finally improve the recognition result by using knowledge engineering of the Cyber Threat Intelligence to help increase the accuracy of entity recognition. Besides, we verify the effectiveness of the proposed model by experiments.
Xuren Wang, Runshi Liu, Zhiting Ling, Peian Yang
CSCWD6
2022 TriCTI: an actionable cyber threat intelligence discovery system via trigger-enhanced neural network
abstract
Abstract The cybersecurity report provides unstructured actionable cyber threat intelligence (CTI) with detailed threat attack procedures and indicators of compromise (IOCs), e.g., malware hash or URL (uniform resource locator) of command and control server. The actionable CTI, integrated into intrusion detection systems, can not only prioritize the most urgent threats based on the campaign stages of attack vectors (i.e., IOCs) but also take appropriate mitigation measures based on contextual information of the alerts. However, the dramatic growth in the number of cybersecurity reports makes it nearly impossible for security professionals to find an efficient way to use these massive amounts of threat intelligence. In this paper, we propose a trigger-enhanced actionable CTI discovery system (TriCTI) to portray a relationship between IOCs and campaign stages and generate actionable CTI from cybersecurity reports through natural language processing (NLP) technology. Specifically, we introduce the “campaign trigger” for an effective explanation of the campaign stages to improve the performance of the classification model. The campaign trigger phrases are the keywords in the sentence that imply the campaign stage. The trained final trigger vectors have similar space representations with the keywords in the unseen sentence and will help correct classification by increasing the weight of the keywords. We also meticulously devise a data augmentation specifically for cybersecurity training sets to cope with the challenge of the scarcity of annotation data sets. Compared with state-of-the-art text classification models, such as BERT, the trigger-enhanced classification model has better performance with accuracy (86.99%) and F1 score (87.02%). We run TriCTI on more than 29k cybersecurity reports, from which we automatically and efficiently collect 113,543 actionable CTI. In particular, we verify the actionability of discovered CTI by using large-scale field data from VirusTotal (VT). The results demonstrate that the threat intelligence provided by VT lacks a part of the threat context for IOCs, such as the Actions on Objectives campaign stage. As a comparison, our proposed method can completely identify the actionable CTI in all campaign stages. Accordingly, cyber threats can be identified and resisted at any campaign stage with the discovered actionable CTI.
Jian Liu 0008, Yitong He, Xuren Wang, Zhengwei Jiang, Peian Yang
Cybersecur.7
2022 TIM: threat context-enhanced TTP intelligence mining on unstructured threat data
abstract
Abstract TTPs (Tactics, Techniques, and Procedures), which represent an attacker’s goals and methods, are the long period and essential feature of the attacker. Defenders can use TTP intelligence to perform the penetration test and compensate for defense deficiency. However, most TTP intelligence is described in unstructured threat data, such as APT analysis reports. Manually converting natural language TTPs descriptions to standard TTP names, such as ATT&CK TTP names and IDs, is time-consuming and requires deep expertise. In this paper, we define the TTP classification task as a sentence classification task. We annotate a new sentence-level TTP dataset with 6 categories and 6061 TTP descriptions from 10761 security analysis reports. We construct a threat context-enhanced TTP intelligence mining (TIM) framework to mine TTP intelligence from unstructured threat data. The TIM framework uses TCENet (Threat Context Enhanced Network) to find and classify TTP descriptions, which we define as three continuous sentences, from textual data. Meanwhile, we use the element features of TTP in the descriptions to enhance the TTPs classification accuracy of TCENet. The evaluation result shows that the average classification accuracy of our proposed method on the 6 TTP categories reaches 0.941. The evaluation results also show that adding TTP element features can improve our classification accuracy compared to using only text features. TCENet also achieved the best results compared to the previous document-level TTP classification works and other popular text classification methods, even in the case of few-shot training samples. Finally, the TIM framework organizes TTP descriptions and TTP elements into STIX 2.1 format as final TTP intelligence for sharing the long-period and essential attack behavior characteristics of attackers. In addition, we transform TTP intelligence into sigma detection rules for attack behavior detection. Such TTP intelligence and rules can help defenders deploy long-term effective threat detection and perform more realistic attack simulations to strengthen defense.
Yizhe You, Zhengwei Jiang, Peian Yang, Baoxu Liu, Huamin Feng, Xuren Wang
Cybersecur.4
2021 Extracting Threat Intelligence Relations Using Distant Supervision and Neural Networks
Yali Luo, Shengqin Ao, Changxin Su, Peian Yang, Zhengwei Jiang
IFIP Int. Conf. Digital Forensics5
2021 FSSRE: Fusing Semantic Feature and Syntactic Dependencies Feature for threat intelligence Relation Extraction
abstract
Threat intelligence relation extraction plays an important role in threat intelligence text analysis and processing.To extract the relation between two threat entities in a sentence, we develop a novel framework called FSSRE which fuses sematic feature and syntactic dependencies feature for threat intelligence relation extraction.We utilize graph convolutional networks (GCN) to extract syntactic dependencies features, and utilize Sentence-BERT to extract contextual semantic features.To keep vital information with irrelevant content removed to the most extent, we further apply a novel pruning strategy, SDP-VP, to the input trees.With retaining the shortest path and nodes that are 𝑲 hops away from nodes on the shortest path, we give the edge connected to the verb nodes a weight of 𝒘 times.We create an advanced persistent threat (APT) intelligence entities and intra-sentence relations dataset, APTER-SENT, for that there is no public dataset can be used for relation extraction research in the threat intelligence field.Experimental results on APTER-SENT demonstrate improved performance over competitive baselines.At the same time, we also conducted experiments on the SemEval-2010 dataset.The results of the experiment indicate that our method is still effective on this dataset.
Xuren Wang, Mengbo Xiong, Famei He, Peian Yang, Binghua Song, Zhengwei Jiang, Zihan Xiong
SEKE4
2021 CAN: Complementary Attention Network for Aspect Level Sentiment Classification in Social E-Commerce
Yali Luo, Zhengwei Jiang, Peian Yang, Xuren Wang
WCNC4
2020 Towards Comprehensive Detection of DNS Tunnels
abstract
The Domain Name System (DNS) is a fundamental service of the Internet, and the DNS tunnel is one of the most threatening abuses of DNS, posing a huge threat to user privacy and Internet security. Attackers conceal the information into DNS packets to evade firewalls and intrusion detection systems. Recently, newly developed DNS tunnels used by Advanced Persist Threat groups tend to use A and AAAA resource records (RRs) for transmission, making them more invisible and more threatening. Previous DNS tunnel detection approaches mainly focus on subdomains and TXT RRs, but less attention has been paid to newly developed DNS tunnels based on A and AAAA RRs. In this paper, we present a novel DNS tunnel detection method that can detect newly developed A and AAAA RR based DNS tunnels. Since DNS tunnels will transmit a large amount of encrypted or encoded data in the DNS queries and responses, we extracted novel features from domains and 4 types of RRs (A, AAAA, TXT and CNAME RRs) that are most commonly used for tunneling to measure the amount and content of information exchanged between the authoritative nameservers and the clients. We also analyze the detection capabilities when different features were used. The anomaly detection algorithm is employed on domains related features and 4 types of RRs related features, respectively. The overlaps of outliers will be marked as DNS tunnels. Our approach has been evaluated on real-world network traffic. The experimental results show that our approach can detect all DNS tunnels in the dataset with a extremely low false positive rate.
Meng Luo 0006, Qiuyun Wang, Yepeng Yao, Xuren Wang, Peian Yang, Zhengwei Jiang
ISCC5