Kai Zhang 0035

dblp:55/957-35 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-1424-7883ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MGDA: A provenance graph-based framework for threat detection and attack scenario reconstruction
abstract
Advanced persistent threat (APT) attacks are sophisticated, stealthy, and persistent, posing significant challenges to timely detection and investigation in modern network environments. Provenance graph analysis has become an important method for APT detection due to its ability to capture detailed causal relationships among system entities. However, existing methods suffer from several limitations: (1) lack of labeled attack data, (2) lack of high-level semantics in attack scenario reconstruction, and (3) high computational overhead limiting practical deployment. In this paper, we propose MGDA, a self-supervised method for effective and accurate threat detection as well as interpretable attack scenario reconstruction. MGDA introduces a multi-view masked graph autoencoder that jointly captures deep semantic features and structural patterns, enabling accurate detection of stealthy and unknown attacks. In the reconstruction phase, MGDA combines contextual analysis with rule-based attack pattern matching to produce attack scenario graphs that incorporate high-level semantics. We evaluate MGDA on three widely used datasets, including both real-world and simulated network attacks. The results demonstrate that MGDA achieves an average precision of 97.58% and F1-score of 98.03% in threat detection, outperforming state-of-the-art approaches. In addition, the automatically reconstructed scenario graphs help identify potential multi-step attacks and their stages, aiding analysts in conducting efficient network attack investigations.
Mengjiao Cui, Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng
Comput. Networks5
2025 TIMFuser: A multi-granular fusion framework for cyber threat intelligence
Zhengwei Jiang, Kai Zhang 0035, Zhiting Ling, Yizhe You, Peian Yang, Huamin Feng
Comput. Secur.3
2024 TiGNet: Joint entity and relation triplets extraction for APT campaign threat intelligence
abstract
Contemporary cybersecurity faces escalating challenges from sophisticated threats, notably Advanced Persistent Threats (APTs). Addressing these challenges necessitates a collaborative, multidisciplinary approach that transcends traditional boundaries. Gathering cyber threat intelligence (CTI) on APT campaigns and constructing a comprehensive knowledge graph empowers defenders to track the latest trends in these campaigns, update defense strategies, and attain crucial advantages in defense measures. Previous works used relation extraction techniques to obtain entity-relation triplets for constructing threat intelligence knowledge graphs. However, these works either rely on pipeline workflow, are susceptible to exposure errors and error propagation, or use sequence annotation method, which combines entity and relation labels but lacks the ability to extract single entity overlaps (SEO) or subject-object overlaps (SOO) triplets. This paper introduces TiGNet, a novel method that transforms the entity-relation triplet’s extraction task into multiple token-span recognition tasks utilizing token-pair matrices. Additionally, we integrate GlobalPointer to incorporate token position information into the token-pair matrix, significantly enhancing extraction performance. To facilitate method evaluation, we annotated a Chinese entity-relation triplets dataset about APT campaigns, named APT-Triplets, comprising 9711 triplets encompassing seven triplet types. Our evaluation demonstrates that TiGNet improves the F1 score of 3.79-5.59 compared to previous joint extraction methods. Furthermore, it outperforms methods based on large language models (LLMs) in terms of both extraction performance and inference time. These results underscore TiGNet’s capacity to accurately and swiftly extract threat intelligence, facilitating the construction of the APT campaign knowledge graphs, empowering defenders to track evolving trends and fortify defense strategies collaboratively.
Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Huamin Feng, Peian Yang
CSCWD3
2024 CTIMiner: Cyber Threat Intelligence Mining Using Adaptive Multi-task Adversarial Active Learning
Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng
ICDF2C (1)3
2024 P-TIMA: a framework of T witter threat intelligence mining and analysis based on a prompt-learning NER model
abstract
Abstract Open-source information platforms such as Twitter continuously provide the latest threat intelligence, including new vulnerabilities and in-the-wild exploitations of advanced persistent threat (APT) groups. Automated extraction of threat intelligence from Twitter has become crucial for defenders to access up-to-date threat knowledge. However, existing studies mainly rely on supervised learning methods to extract threat intelligence knowledge, such as entities, which require a large amount of annotated data. This paper presents Threat Intelligence Mining and Analysis based on Prompt Learning (P-TIMA), a framework specifically crafted for extracting and analyzing threat intelligence from Twitter. P-TIMA employs our innovative few-shot entity recognition method, SecEntPrompt (SEP), built on prompt learning, to extract vulnerability intelligence from Twitter. Additionally, P-TIMA analyzes and profiles the overarching vulnerability intelligence obtained from Twitter, along with in-the-wild exploitation intelligence of APT groups. The SEP improves the average entity recognition F1 score by 3.62-4.40 compared with the best-performing comparison model and outperforms the method based on the large language model on recognition performance and inference time. To validate our framework, we apply P-TIMA to extract vulnerability-related threat intelligence from real Twitter data. Through case studies, we then analyze trends in vulnerability threats and the exploitation capabilities of APT groups. In conclusion, our framework provides a more efficient and accurate method for extracting threat intelligence from Twitter, enabling defenders to stay up-to-date with the latest threat trends and helping them improve their defense strategies against cyber attacks.
Yizhe You, Zhengwei Jiang, Peian Yang, Kai Zhang 0035, Xuren Wang, Chenpeng Tu, Huamin Feng
Comput. J.5
2023 FineCTI: A Framework for Mining Fine-grained Cyber Threat Information from Twitter Using NER Model
abstract
To timely respond to cyber threats related to a specific IT infrastructure called fine-grained (e.g., Windows or Linux), security analysts need to require timely and comprehensive threat information. Twitter, as a vital source of real-time threat information, provides abundant but overwhelming information due to the increased data sources. Automatically mining and summarizing fine-grained threat information from Twitter can help security analysts maintain the infrastructure’s security. Most existing studies focus on classification, which carries less threat information. Some works use clustering based on text similarity relying on the embedding of text obtained from pre-trained models, which cannot be applied to short text, resulting in noisy clusters. Several works build topic models. However, the incoherent topic keywords are difficult to understand and analyze. To overcome these challenges, we design a FineCTI framework to mine the threat information related to the specific infrastructure on Twitter and generate a detailed threat information summary that is machine-readable and human-readable, efficiently reducing information overload. FineCTI optimizes the feature extraction part based on the named entity recognition model and performs clustering based on features extracted, thus effectively reducing the influence of sparsity of tweets on the clustering result and with the V-measure score improved by 7%. The cluster analysis results show that we can mine the fine-grained threats up to 15 days before the official disclosure date.
Kai Zhang 0035, Zhengwei Jiang, Peian Yang, Xuren Wang, Huamin Feng
TrustCom3
2022 TI-Prompt: Towards a Prompt Tuning Method for Few-shot Threat Intelligence Twitter Classification*
abstract
Obtaining the latest Threat Intelligence (TI) via Twitter has become one of the most important methods for defenders to catch up with emerging cyber threats. Existing TI Twitter classification works mainly based on supervised learning methods. Such approaches require large amounts of annotated data and are difficult to be transferred to other TI Twitter classification tasks. This paper proposes a prompt-based method for classifying TI on Twitter, named TI-Prompt. TI-Prompt lever-ages the prompt-tuning method with two templates in different TI Twitter classification tasks. TI-Prompt also uses a semantic similarity-based approach to automatically enrich the prompt verbalizer without expert knowledge and a verbalizer refinement method to calibrate the verbalizer based on the training data. We evaluate TI-Prompt with binary and multi-classification tasks on two Twitter Threat Intelligence datasets. Evaluation results show that the proposed TI-Prompt improves 5-10% over the best performance of previous supervised learning methods under the few-shot settings. Compared to the general prompt-tuning methods, the proposed prompt-tuning templates can also improve the classification performance by 2–5%. Meanwhile, the proposed verbalizer enrichment method and refinement method improve classification accuracy by 1–4% compared with the general single-word verbalizer prompt method. Therefore, TI-Prompt can be extended to other Threat Intelligence classification tasks without requiring large amounts of training data, significantly reducing the annotation cost.
Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Xuren Wang, Shirui Wang, Huamin Feng
COMPSAC3
2021 Representativeness-Based Instance Selection for Intrusion Detection
abstract
With the continuous development of network technology, an intrusion detection system needs to face detection efficiency and storage requirement when dealing with large data. A reasonable way of alleviating this problem is instance selection, which can reduce the storage space and improve intrusion detection efficiency by selecting representative instances. An instance is representative not only in its class but also in different classes. This representativeness reflects the importance of an instance. Since the existing instance selection algorithm does not take into account the above situations, some selected instances are redundant and some important instances are removed, increasing storage space and reducing efficiency. Therefore, a new representativeness of instance is proposed and considers not only the influence of all instances of the same class on the selected instance but also the influence of instances of different classes on the selected instance. Moreover, it considers the influence of instances of different classes as an advantageous factor. Based on this representativeness, two instance selection algorithms are proposed to handle balanced and imbalanced data problems for intrusion detection. One is a representative-based instance selection for balanced data, which is named RBIS and selects the same proportion of instances from each class. The other is a representative-based instance selection for imbalanced data, which is named RBIS-IM and selects important majority instances according to the number of instances of the minority class. Compared with other algorithms on the benchmark data sets of intrusion detection, experimental results verify the effectiveness of the proposed RBIS and RBIS-IM algorithms and demonstrate that the proposed algorithms can achieve a better balance between accuracy and reduction rate or between balanced accuracy and reduction rate.
Fei Zhao 0004, Yang Xin 0001, Kai Zhang 0035, Xinxin Niu
Secur. Commun. Networks3