Zhengwei Jiang

dblp:55/8606 · DBLP profile ↗
← Back
55ranked-venue papers
1as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 17 · 17 since 2021Security and privacy · 13 · 10 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Computer networks · 7 · 6 since 2021Systems, architecture and hardware · 4 · 2 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 GDPO: Dual Learning for Self-Supervised Code Summarization in the Era of Large Language Models
Shuwei Wang, Weize Zhang, Zhengwei Jiang, Qiuyun Wang
SANER4
2026 MGDA: A provenance graph-based framework for threat detection and attack scenario reconstruction
abstract
Advanced persistent threat (APT) attacks are sophisticated, stealthy, and persistent, posing significant challenges to timely detection and investigation in modern network environments. Provenance graph analysis has become an important method for APT detection due to its ability to capture detailed causal relationships among system entities. However, existing methods suffer from several limitations: (1) lack of labeled attack data, (2) lack of high-level semantics in attack scenario reconstruction, and (3) high computational overhead limiting practical deployment. In this paper, we propose MGDA, a self-supervised method for effective and accurate threat detection as well as interpretable attack scenario reconstruction. MGDA introduces a multi-view masked graph autoencoder that jointly captures deep semantic features and structural patterns, enabling accurate detection of stealthy and unknown attacks. In the reconstruction phase, MGDA combines contextual analysis with rule-based attack pattern matching to produce attack scenario graphs that incorporate high-level semantics. We evaluate MGDA on three widely used datasets, including both real-world and simulated network attacks. The results demonstrate that MGDA achieves an average precision of 97.58% and F1-score of 98.03% in threat detection, outperforming state-of-the-art approaches. In addition, the automatically reconstructed scenario graphs help identify potential multi-step attacks and their stages, aiding analysts in conducting efficient network attack investigations.
Mengjiao Cui, Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng
Comput. Networks2
2025 AGLHunter: Automated Threat Hunting Using In-Context Learning-Enhanced LLM
abstract
Advanced Persistent Threats (APTs) are characterized by their persistence, sophistication, and stealth, posing significant challenges to network detection. Existing research on attack detection leveraging Provenance Graphs (PGs) has proven effective in correlating system entities and capturing persistence. However, the exponential growth of audit logs makes large-scale data storage and processing difficult. In addition, current threat hunting methods rely heavily on manually crafted attack query graphs, which are limited by expert knowledge and lack automated solutions. In this paper, we propose AGLHunter, an automated threat hunting system designed to enhance automation and efficiency while maintaining high detection accuracy. Our system leverages the In-Context Learning (ICL) capability of the Large Language Model (LLM) to automatically construct query graphs from Cyber Threat Intelligence (CTI) reports. Next, we extract suspicious subgraphs from PGs and employ graph representation learning to match these sub graphs with the query graphs, enabling efficient and accurate threat hunting. We use DARPA TC and OpTC datasets to evaluate AGLHunter's performance. The results show that AGLHunter not only achieves higher automation but also shows superior performance with reduced memory usage. AGLHunter, leveraging ICL-enhanced LLM, improved the F1 score for query graph construction by 13.6%, reduced the overall hunting time by more than 170 seconds, and maintained high detection accuracy.
Mengjiao Cui, Zhengwei Jiang, Yepeng Yao, Qiying He, Peian Yang, Huamin Feng
CSCWD2
2025 From Threat Report to ATT&CK: Automated Extraction and Reasoning of TTPs Using Large Language Models
abstract
The escalating frequency and increasing complexity of cyber attacks underscore the importance of Cyber Threat Intelligence (CTI). Tactics, Techniques, and Procedures (TTPs), as advanced CTI capable of characterizing adversarial behaviors and intentions, have garnered increased attention. However, TTPs are predominantly found embedded within unstructured natural language texts of threat reports. The accurate extraction and standardization of TTPs pose significant challenges. Existing methods exhibit limitations in terms of accuracy, generalizability, and interpretability. This paper presents a pipeline for automatically extracting TTPs from threat reports and providing rationales using large language models. To support this approach, we have developed three datasets using advanced commercial LLMs for data synthesis. These datasets are made publicly available to facilitate further research. Experimental results demonstrate the superior performance of our proposed approach, achieving an F1-score of 97.15% and accuracies of 79.22% and 92.97% in the respective tasks. These results surpass state-of-the-art methods by 15.39%, 12.87%, and 27.57%, respectively. To the best of our knowledge, this paper is the first to simultaneously extract TTPs while providing the underlying rationales for the extraction. This novel approach significantly improves the usability of the results by providing a richer context for threats.
Fangming Dong, Zhengwei Jiang, Qiying He, Peian Yang, Yepeng Yao
CSCWD2
2025 A Variant and Flow-Level AutoML Method for IoT Malicious Traffic Detection
abstract
The Internet of Things (IoT) involves communication and data exchange between a wide range of devices, often with security implications. Compared to Internet, IoT is costly to detect malicious traffic due to its complex protocols, resource limitations and variety of new attacks and attack variants. Automatic machine learning (AutoML) eliminates model selection and hyperparameter optimisation, which can reduce human dependency and address these issues. However, AutoML still cannot extract and represent features from raw data based on a specific problem. Manual feature extraction and representation will greatly affect the accuracy of the model and still rely on professional domain knowledge. Automated feature engineering in AutoML requires universal and efficient feature extraction and representation. This paper proposes a variant and flow-level AutoML (VFA) for IoT malicious traffic detection. VFA has added a binary representation of comprehensive content features based on the packet. The packet representation allows AutoML to automatically learn important features from a normalised aligned structure without guidance. Variant of exclusive OR (vXOR) enables data aggregation, allowing VFA to focus on the content features in the flow and the inherent connections between packets. VFA can strategically adjust monitoring priorities by adjusting parameters, allowing it to respond flexibly to different resource constraints or specific attack. We have evaluated VFA on a real-world dataset, IoT-23. We believe that the complete data-to-label VFA can be extended to other areas in the future.
Xinqiang Zhao, Xuren Wang, Zijing Fan, Yepeng Yao, Zhengwei Jiang
CSCWD7
2025 What Makes an Email Insecure: A Fine-Grained Risk Assessment Scheme for Phishing Emails Targeting Attack Vectors
abstract
As online adversarial tactics escalate, particularly with the utilization of large language models, distinguishing phishing emails from benign ones has become increasingly challenging in terms of appearance and semantics. The use of various techniques, including visual deception and the concealment of malicious attachments, has become more widespread, posing significant challenges to traditional machine learning models that rely on text and semantic features. To address these challenges, this study introduces a Fine-grained Phishing Email Risk Assessment framework (FPERA), which focuses on prevalent attack vectors. By integrating techniques such as deep email header inspection, visual analysis, and threat intelligence, FPERA systematically examines potential risk factors across multiple dimensions and calculates the risk score of emails using a risk weighting matrix. Experimental tests conducted on multiple original datasets and AI-generated datasets have demonstrated that detection targeting attack vectors can more effectively counter phishing tactics like AI-enhanced polishing, email header exploits, and visual deception, while exhibiting consistent stability across different datasets.
Ximin Huang, Fangli Ren, Weize Zhang, Shuwei Wang, Qiuyun Wang, Zhengwei Jiang
CSCWD7
2025 Graph Representation Learning via Generative-Contrastive Fusion for Advanced Persistent Threat Detection
Yijiao Jiang, Fangming Dong, Zhengwei Jiang, Tianming Zheng, Baoxu Liu, Liling Xin
ICA3PP (4)4
2025 Detection and Analysis of Poisoned Image in Container Registry
abstract
Container technology provides isolated, consistent, and efficient application environments across diverse computing platforms. Docker, as the dominant container platform, simplifies container creation, deployment, and operation. However, the integrity of the container supply chain is threatened by malicious “poisoned” images distributed via public registries like Docker Hub, which pose significant security risks to unsuspecting users. This study presents the first comprehensive investigation into this container registry image supply chain threat. We reveal that image poisoning occurs primarily during the build phase, where attackers embed malicious payloads into Dockerfiles and associated build artifacts via specific common vectors. To empirically assess this threat, we developed a detection system combining dynamic and static analysis. Scanning 214,920 images from public registries, we identified 122 poisoned images with high precision ($\mathbf{9 5. 3 1 \%}$). Our in-depth analysis shows these compromised images serve diverse malicious purposes, with cryptocurrency mining being prevalent, and exhibit significant characteristic differences compared to benign images. Finally, we propose concrete mitigation measures to improve the security of the Docker ecosystem. We reported our findings to Docker and received its confirmation.
Siyuan Pang, Yongshan Wang, Yepeng Yao, Zhengwei Jiang, Zijing Fan, Baoxu Liu
ISSRE4
2025 Automated Attack Graph Construction for Cross-host Threat Detection Using Cyber Threat Intelligence
abstract
Cyber Threat Intelligence (CTI) reports provide valuable insights into cyber threats. However, manually constructing attack graphs from the unstructured CTI reports requires significant human effort. With the development of Large Language Models (LLMs), researchers have begun to harness LLMs for attack graph construction from CTI reports. Nevertheless, existing works mainly focus on describing abstract and high-level attack behaviors through these graphs, which cannot be used as query graphs for threat detection based on graph matching, as provenance graphs are at the system level. Moreover, these works do not consider modeling cross-host attack behaviors. To address these problems, we propose a novel method for automatically constructing attack graphs from CTI reports. We utilize prompt engineering, and leverage the in-context learning ability of LLMs to generate attack graphs. In this method, we first restructure the CTI report by grouping continuous sentences with the same tactics, and then use multiple LLMs to extract entities and relations. Finally, we use an LLM to integrate all the results. We also design a cross-host threat detection algorithm using the generated attack graphs. The evaluation results show that our method constructs attack graphs with an average conversion rate of 72.4%. It also achieves nearly 94.8% precision and 96.5% recall for IoC and relation extraction compared with manually labeled results.
Ziqing Feng, Qiuyun Wang, Liling Xin, Zhengwei Jiang, Huamin Feng
TrustCom7
2025 Can LLMs deeply detect complex malicious queries? A framework for jailbreaking via obfuscating intent
abstract
Abstract This paper delves into a possible security flaw in large language models (LLMs), particularly in their capacity to identify malicious intent within intricate or ambiguous inquiries. We have discovered that LLMs might overlook the malicious nature of highly veiled requests, even without alterations to the malevolent text in those queries, thus exposing a significant weakness in their content analysis systems. To be specific, we pinpoint and scrutinize two aspects of this vulnerability: (i) LLMs’ diminished capability to perceive maliciousness when parsing extremely obscured queries, and (ii) LLMs’ inability to discern malicious intent in queries that have been intentionally altered to increase their ambiguity by modifying the malevolent content itself. To illustrate and tackle this problem, we propose a theoretical framework and analytical strategy, and introduce a novel black-box jailbreak attack technique called IntentObfuscator. This technique exploits the identified vulnerability by concealing the genuine intentions behind user prompts, thereby compelling LLMs to inadvertently produce restricted content and circumvent their inherent content safety protocols. We elaborate on two specific applications within this framework: ”Obscure Intention” and ”Create Ambiguity,” which skillfully manipulate the complexity and ambiguity of queries to effectively dodge the detection of malicious intent. We empirically confirm the efficacy of the IntentObfuscator approach across various models, including ChatGPT-3.5, ChatGPT-4, Qwen, and Baichuan, achieving an average jailbreak success rate of 69.21%. Remarkably, our tests on ChatGPT-3.5, boasting 100 million weekly active users, yielded an impressive success rate of 83.65%. Additionally, we verify our approach across a range of sensitive content categories, including graphic violence, racism, sexism, political sensitivity, cybersecurity threats, and criminal techniques, further highlighting the considerable impact of our findings on refining ”Red Team” tactics against LLM content security frameworks.
Shang Shang, Xinqiang Zhao, Zhongjiang Yao, Yepeng Yao, Liya Su, Zijing Fan, Zhengwei Jiang
Comput. J.8
2025 Advanced code slicing with pre-trained model fine-tuned for open-source component malware detection
abstract
Abstract Open Source Software (OSS) is an essential part of modern software development, with platforms such as PyPI for Python, NPM for JavaScript, and RubyGems for Ruby facilitating code sharing and reuse. However, these repositories also pose significant security risks due to potential software supply chain attacks, where payloads are injected into components, propagating threats to downstream users and critical infrastructure. Existing automatic malicious component detection tools, particularly for PyPI, struggle to distinguish between subtle differences in malicious and benign behaviors, leading to high false positive rates. To address these issues, we systematically compare and explore these subtle differences, offering a more refined and accurate detection method, Open-Source Component Code Slices BERT (OCS-BERT). OCS-BERT leverages taint-based program slicing to isolate sensitive behavior segments and fine-tunes pre-trained model to capture subtle semantic differences across programming languages. This system excels in detecting malicious Python components and exhibits encouraging cross-language transferability to JavaScript's NPM and Ruby's RubyGems. Additionally, OCS-BERT successfully detected 107 malicious components from a total of 25,759 newly-uploaded PyPI components, taking two weeks to complete the process. This achievement demonstrates the effectiveness of our method, which serves as a potent enhancement to the current repertoire of software supply chain detection methodologies.
Yongshan Wang, Siyuan Pang, Zijing Fan, Shang Shang, Yepeng Yao, Zhengwei Jiang, Baoxu Liu
Comput. J.6
2025 Who are querying for me? Measuring the dependency and centralization in recursive resolution
Qiuyun Wang, Jianrong Zhang, Baojiang Cui, Zhengwei Jiang
Comput. Secur.7
2025 TIMFuser: A multi-granular fusion framework for cyber threat intelligence
Zhengwei Jiang, Kai Zhang 0035, Zhiting Ling, Yizhe You, Peian Yang, Huamin Feng
Comput. Secur.2
2024 Automated Anti-malware Detection Rules Converter Based on SIMIOC
abstract
In recent years, using IOC to detect malware-based network attacks has become an effective and accurate method, but the scheme of manually writing IOC rules is inefficient and can not meet the detection needs of a large number of rapidly iterated malware. Therefore, more efficient methods are needed to automatically convert open-source rules into anti-malware detection rules (IOC rules for detecting malware). In this paper, we propose a method to automate the conversion of anti-malware detection rules using open-source detection rules. We designed an intermediate structure called SIMIOC (Structure Intermediate-representation for Malware Information of Compromise) and implemented a SIMIOC-based converter. In the experiment, we used the SIMIOC-based converter to automatically convert 1218 rules for the Windows platform from three open-source rule repositories: Sigma, Elastic Security Detection Rules, and Splunk Security Content into signature detection rules that can be used in the Cuckoo sandbox, and deployed these rules to detect 21044 malware. By analyzing the experimental results, we found that the detection rate of the detection rules automatically converted by SIMIOC-based converter is 50.3%, and it reaches 73% of 640 cuckoo Sandbox signature manual rules. Furthermore, we demonstrated that the rules generated by SIMIOC-based converters have their emphasis on TTPs (Tactics, Techniques, and Procedures) and families, which are the optimization and complement of manual rules.
Shuangze He, Zhengwei Jiang, Qiuyun Wang
CSCWD5
2024 A Novel Detection System for Multi-Architecture IoT Malware
abstract
As IoT devices become more prevalent, the thread that comes from the malware of IoT becomes more serious. In comparison to desktops, IoT malware has characteristics of different platforms, considerable environment reliance, and continual updating of countermeasure technology, which creates a huge obstacle for malware detection. In light of the aforementioned issues, we propose a set of analysis methods, including static analysis against software shells, dynamic analysis with sandbox has 9 different architecture environments, and a detection model designed based on SHAP(SHapley Additive ex-Planations) and Smith-Waterman algorithm with a small amount of prior knowledge about IOT malware. The experimental results show that reports about malware containing static information and dynamic behavior can be generated. And the detection accuracy of the model can reach 98.31%. At the same time, in the case of the training set has fewer malicious samples, the accuracy that our model has a 5%-10% improvement, proof of our method has better malware discovery ability for unknown variants.
Shuwei Wang, Molan Long, Qiuyun Wang, Rongqi Jing, Zhengwei Jiang
CSCWD6
2024 TiGNet: Joint entity and relation triplets extraction for APT campaign threat intelligence
abstract
Contemporary cybersecurity faces escalating challenges from sophisticated threats, notably Advanced Persistent Threats (APTs). Addressing these challenges necessitates a collaborative, multidisciplinary approach that transcends traditional boundaries. Gathering cyber threat intelligence (CTI) on APT campaigns and constructing a comprehensive knowledge graph empowers defenders to track the latest trends in these campaigns, update defense strategies, and attain crucial advantages in defense measures. Previous works used relation extraction techniques to obtain entity-relation triplets for constructing threat intelligence knowledge graphs. However, these works either rely on pipeline workflow, are susceptible to exposure errors and error propagation, or use sequence annotation method, which combines entity and relation labels but lacks the ability to extract single entity overlaps (SEO) or subject-object overlaps (SOO) triplets. This paper introduces TiGNet, a novel method that transforms the entity-relation triplet’s extraction task into multiple token-span recognition tasks utilizing token-pair matrices. Additionally, we integrate GlobalPointer to incorporate token position information into the token-pair matrix, significantly enhancing extraction performance. To facilitate method evaluation, we annotated a Chinese entity-relation triplets dataset about APT campaigns, named APT-Triplets, comprising 9711 triplets encompassing seven triplet types. Our evaluation demonstrates that TiGNet improves the F1 score of 3.79-5.59 compared to previous joint extraction methods. Furthermore, it outperforms methods based on large language models (LLMs) in terms of both extraction performance and inference time. These results underscore TiGNet’s capacity to accurately and swiftly extract threat intelligence, facilitating the construction of the APT campaign knowledge graphs, empowering defenders to track evolving trends and fortify defense strategies collaboratively.
Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Huamin Feng, Peian Yang
CSCWD2
2024 IntentObfuscator: A Jailbreaking Method via Confusing LLM with Prompts
Shang Shang, Zhongjiang Yao, Yepeng Yao, Liya Su, Zijing Fan, Zhengwei Jiang
ESORICS (4)7
2024 CTIMiner: Cyber Threat Intelligence Mining Using Adaptive Multi-task Adversarial Active Learning
Zhengwei Jiang, Kai Zhang 0035, Peian Yang, Huamin Feng
ICDF2C (1)2
2024 APTChaser: Cyber Threat Attribution via Attack Technique Modeling
Peian Yang, Zhengwei Jiang, Mengjiao Cui, Yizhe You
ICDF2C (1)3
2024 Automated Mining of Multi-Dimensional Information from APT Malware for Effective Feature Analysis and Threat Actor Attribution
Rongqi Jing, Zhengwei Jiang, Qiuyun Wang, Shuwei Wang
ICONIP (6)2
2024 CTIFuser: Cyber Threat Intelligence Fusion via Unsupervised Learning Model
abstract
Cyber attack campaigns are becoming increasingly complex and severe, causing significant impacts on institutions and individuals. Cyber Threat Intelligence (CTI) provides important evidential knowledge about attackers and is critical to the shift from reactive to proactive defense against cyber attacks. Attack detection based on Indicators of Compromise (IOCs), a type of CTI, is vulnerable to the limitation of insufficient context of attack scenarios. In contrast, attack behavior intelligence is associated with information on attackers’ techniques, targets, and intentions, providing a solid foundation for security practitioners to conduct attack investigations or other applications. Many current CTI mining systems are limited to extracting CTI from a single source, leading to challenges such as fragmented attack behavior view and low-value density. To address these issues, we propose an unsupervised fusion framework named CTIFuser, which includes a comprehensive pipeline of four subtasks aimed at mining and fusing multi-source attack behaviors at the attack technique level. In our evaluation of 739 real-world CTI reports from 542 sources, experimental results demonstrate that CTIFuser can obtain a complete view of the attack behaviors at the attack technique level.
Zhengwei Jiang, Peian Yang, Mengjiao Cui, Fangming Dong, Huamin Feng
ISPA2
2024 HRTC: A Triplet Joint Extraction Model Based on Cyber Threat Intelligence
HuanZhou Yue, Xuren Wang, Zhengwei Jiang, Yuxia Fu
KSEM (5)4
2024 P-TIMA: a framework of T witter threat intelligence mining and analysis based on a prompt-learning NER model
abstract
Abstract Open-source information platforms such as Twitter continuously provide the latest threat intelligence, including new vulnerabilities and in-the-wild exploitations of advanced persistent threat (APT) groups. Automated extraction of threat intelligence from Twitter has become crucial for defenders to access up-to-date threat knowledge. However, existing studies mainly rely on supervised learning methods to extract threat intelligence knowledge, such as entities, which require a large amount of annotated data. This paper presents Threat Intelligence Mining and Analysis based on Prompt Learning (P-TIMA), a framework specifically crafted for extracting and analyzing threat intelligence from Twitter. P-TIMA employs our innovative few-shot entity recognition method, SecEntPrompt (SEP), built on prompt learning, to extract vulnerability intelligence from Twitter. Additionally, P-TIMA analyzes and profiles the overarching vulnerability intelligence obtained from Twitter, along with in-the-wild exploitation intelligence of APT groups. The SEP improves the average entity recognition F1 score by 3.62-4.40 compared with the best-performing comparison model and outperforms the method based on the large language model on recognition performance and inference time. To validate our framework, we apply P-TIMA to extract vulnerability-related threat intelligence from real Twitter data. Through case studies, we then analyze trends in vulnerability threats and the exploitation capabilities of APT groups. In conclusion, our framework provides a more efficient and accurate method for extracting threat intelligence from Twitter, enabling defenders to stay up-to-date with the latest threat trends and helping them improve their defense strategies against cyber attacks.
Yizhe You, Zhengwei Jiang, Peian Yang, Kai Zhang 0035, Xuren Wang, Chenpeng Tu, Huamin Feng
Comput. J.2
2023 Who Are Querying For Me? Egress Measurement For Open DNS Resolvers
abstract
The dependencies and centralization in DNS infrastructure increase the risk of single-point failure and the scope of collateral damage. In the DNS recursive resolution, dependencies between different resolvers also exist due to situations such as forwarding. Currently, research on dependencies in recursive resolution is still insufficient. In this work, we take a deep insight into the recursive resolution implemented by open resolvers to investigate their dependencies, including the concentration of dependencies, and the influence of 3rd-party providers. We find that most open resolvers in the wild are dependent on a small number of egress resolvers to communicate with the authoritative name servers. 90% of the open resolvers are influenced by 8.41% of the egress resolvers. Besides, egress resolvers from 3rd-party providers are able to influence more than 44% of the open resolvers. The concentration makes a large amount of DNS traffic concentrated in a small number of egress resolvers/providers, which will reduce the redundancy of DNS and threaten user privacy.
Meng Luo 0006, Liling Xin, Yepeng Yao, Zhengwei Jiang, Qiuyun Wang, Wenchang Shi
CSCWD4
2023 Phishsifter: An Enhanced Phishing Pages Detection Method Based on the Relevance of Content and Domain
abstract
A Phishing website is used to steal users’ private information. The accelerated development of phishing kits has made it convenient to create such websites, which has become a persistent security threat. In this article, we propose a novel method to detect phishing webpages based on the relevance of the webpage content and domain. For phishing webpages whose domain is relevant to the content, we use the target identification method to identify the target brand. We use two components, the website logo and domain, to identify phishing sites, which increases the accuracy of identification. For irrelevant websites, we use a feature-based approach to distinguish phishing webpages. The experiment shows that the accuracy of target identification is 97.21%, while the false positive rate is 1.47%. The accuracy of the feature-based method is 98.32%. The proposed scheme can meet the needs of practical applications and provide an interpretation of the classification results.
Zhengwei Jiang, Zhiting Ling, Peian Yang
CSCWD2
2023 FineCTI: A Framework for Mining Fine-grained Cyber Threat Information from Twitter Using NER Model
abstract
To timely respond to cyber threats related to a specific IT infrastructure called fine-grained (e.g., Windows or Linux), security analysts need to require timely and comprehensive threat information. Twitter, as a vital source of real-time threat information, provides abundant but overwhelming information due to the increased data sources. Automatically mining and summarizing fine-grained threat information from Twitter can help security analysts maintain the infrastructure’s security. Most existing studies focus on classification, which carries less threat information. Some works use clustering based on text similarity relying on the embedding of text obtained from pre-trained models, which cannot be applied to short text, resulting in noisy clusters. Several works build topic models. However, the incoherent topic keywords are difficult to understand and analyze. To overcome these challenges, we design a FineCTI framework to mine the threat information related to the specific infrastructure on Twitter and generate a detailed threat information summary that is machine-readable and human-readable, efficiently reducing information overload. FineCTI optimizes the feature extraction part based on the named entity recognition model and performs clustering based on features extracted, thus effectively reducing the influence of sparsity of tweets on the clustering result and with the V-measure score improved by 7%. The cluster analysis results show that we can mine the fine-grained threats up to 15 days before the official disclosure date.
Kai Zhang 0035, Zhengwei Jiang, Peian Yang, Xuren Wang, Huamin Feng
TrustCom4
2023 M3F: A novel multi-session and multi-protocol based malware traffic fingerprinting
Jian Liu 0008, Qingsai Xiao, Liling Xin, Qiuyun Wang, Yepeng Yao, Zhengwei Jiang
Comput. Networks6
2022 TI-Prompt: Towards a Prompt Tuning Method for Few-shot Threat Intelligence Twitter Classification*
abstract
Obtaining the latest Threat Intelligence (TI) via Twitter has become one of the most important methods for defenders to catch up with emerging cyber threats. Existing TI Twitter classification works mainly based on supervised learning methods. Such approaches require large amounts of annotated data and are difficult to be transferred to other TI Twitter classification tasks. This paper proposes a prompt-based method for classifying TI on Twitter, named TI-Prompt. TI-Prompt lever-ages the prompt-tuning method with two templates in different TI Twitter classification tasks. TI-Prompt also uses a semantic similarity-based approach to automatically enrich the prompt verbalizer without expert knowledge and a verbalizer refinement method to calibrate the verbalizer based on the training data. We evaluate TI-Prompt with binary and multi-classification tasks on two Twitter Threat Intelligence datasets. Evaluation results show that the proposed TI-Prompt improves 5-10% over the best performance of previous supervised learning methods under the few-shot settings. Compared to the general prompt-tuning methods, the proposed prompt-tuning templates can also improve the classification performance by 2–5%. Meanwhile, the proposed verbalizer enrichment method and refinement method improve classification accuracy by 1–4% compared with the general single-word verbalizer prompt method. Therefore, TI-Prompt can be extended to other Threat Intelligence classification tasks without requiring large amounts of training data, significantly reducing the annotation cost.
Yizhe You, Zhengwei Jiang, Kai Zhang 0035, Xuren Wang, Shirui Wang, Huamin Feng
COMPSAC2
2022 The Hyperbolic Temporal Attention Based Differentiable Neural Turing Machines for Diachronic Graph Embedding in Cyber Threat Intelligence
abstract
Cyber Threat Intelligence (CTI) is an effective approach to solve cyber security problems, finding unknown threats is becoming a problem to be solved. Research based on threat intelligence knowledge graphs has gradually increased for its capability to capture entity characteristics and better predict unknown threats. Most of the current research focuses on Euclidean space, however, the Euclidean space is insufficient to capture the hierarchical information of the knowledge graph. In this paper, we propose a novel Hyperbolic Temporal Attention based Differential Neural Turing Machines for diachronic graph embedding framework (HTA-DNTM), which adopts a graph attention network model based on hyperbolic space, simultaneously uses multi-head self-attention to map the temporal graph into hyperbolic space and incorporates hyperbolic graph neural network and hyperbolic gated recurrent neural network, capturing the evolving behaviors and implicitly preserve hierarchical information simultaneously. Moreover, we demonstrate significantly improved performance over various approaches on CTI. A series of benchmark experiments illustrate HTA-DNTM has ability to generate higher quality than state-of-the-art word embedding models in CTI fields.
Binghua Song, Baoxu Liu, Zhengwei Jiang, Xuren Wang
CSCWD4
2022 APTNER: A Specific Dataset for NER Missions in Cyber Threat Intelligence Field
abstract
This paper provides a new dataset for Named Entity Recognition (NER) missions in cyber threat intelligence (CTI) studying. To the best of our knowledge, the proposed dataset is the biggest and challenging one in the field to comply with the STIX 2.1 specification. We collected the APT (Advanced Persistent Threats) reports from different network security companies and manually annotated them. Then we constructed a dataset named APTNER, which can be used for NER joint and multi-task learning tasks in CTI. Apart from common labels like IP, URL, mal-ware, location and so on, APTNER contains 21 categories, which make APTNER more challenging than other NER datasets in CTI field and we have proved the rationality of the dataset. For ease of comparison studies, we realize several state-of-the-art baselines and report their analysis. To facilitate future work on fine-grained NER for CTI, we make APTNER public at https://github.com/wangxuren/APTNER.
Xuren Wang, Songheng He, Zihan Xiong, Xinxin Wei, Zhengwei Jiang
CSCWD5
2022 A Graph Learning Approach with Audit Records for Advanced Attack Investigation
abstract
System audit logs are widely adopted in enterprise security by their support for causality analysis that generates provenance graphs to investigate advanced attacks. However, detecting attack activity in the overwhelming amount of logs is like looking for a needle in a haystack, which slows down the speed of attack investigation. In this paper, we propose an automated approach for attack detection and investigation by learning the contextual semantics of the provenance graph. Our framework uncovers the semantics of the attack events through the structural context of audit logs. Further, an attention-based graph convolutional neural network is utilized to capture the structural identity associated with the attack path. It is important to note that when inferring whether a specific system node is malicious or not, our approach optimizes the provenance subgraph generated for that node without destroying its contextual semantics. The discovered attack nodes are correlated chronologically for attack investigation and scenario reconstruction. Our approach is evaluated on a real-world Advanced Persistent Threats (APT) dataset. The results show that our approach has a high F1 score (95.97%) for identifying attack nodes in audit logs and speeds up the process of attack investigation (reducing the analysis workload by 91.24%).
Jian Liu 0008, Zhengwei Jiang, Xuren Wang
GLOBECOM3
2022 Effectiveness Evaluation of Evasion Attack on Encrypted Malicious Traffic Detection
abstract
With more and more TLS encrypted traffic on the Internet, an increasing amount of malware is using TLS to hide their tracks. The encrypted traffic makes the traditional malicious traffic detection methods invalid. Machine learning algorithms have become essential options for detecting encrypted malicious traffic. Recently, researchers found that machine learning algorithms have flaws, and threat actors can use some tricks to evade detection. But it remains an open question on how these machine learning-based encrypted malicious traffic detection algorithms perform in the face of evasion attacks.We explore the answer in this paper. We first define five mutation rules to generate adversarial examples. With these mutation rules, we can evaluate the ability of several detection algorithms to deal with evasion attacks when detecting encrypted malicious traffic. The encrypted malicious traffic collected for 12 months is used for experiments. Experiments show that modifying the destination port can reduce the detection rate of detection algorithms in feature space, except for random forest algorithms. Inserting junk data has minimal effect on these algorithms. Whether in the problem space or feature space, inserting useless cipher suites and simulating browser’s traffic can significantly reduce the detection rate of these algorithms. When simulating browser’s traffic, the random forest algorithm almost loses its usability. The same situation arises when SVM is faced with inserting useless cipher suites. Compared with inserting useless cipher suites, inserting useless extensions has a minor effect on these algorithms. Our findings will contribute to future research on encrypted malicious traffic detection.
Jian Liu 0008, Qingsai Xiao, Zhengwei Jiang, Yepeng Yao, Qiuyun Wang
WCNC3
2022 Measurement for encrypted open resolvers: Applications and security
Meng Luo 0006, Yepeng Yao, Liling Xin, Zhengwei Jiang, Qiuyun Wang, Wenchang Shi
Comput. Networks4
2022 TriCTI: an actionable cyber threat intelligence discovery system via trigger-enhanced neural network
abstract
Abstract The cybersecurity report provides unstructured actionable cyber threat intelligence (CTI) with detailed threat attack procedures and indicators of compromise (IOCs), e.g., malware hash or URL (uniform resource locator) of command and control server. The actionable CTI, integrated into intrusion detection systems, can not only prioritize the most urgent threats based on the campaign stages of attack vectors (i.e., IOCs) but also take appropriate mitigation measures based on contextual information of the alerts. However, the dramatic growth in the number of cybersecurity reports makes it nearly impossible for security professionals to find an efficient way to use these massive amounts of threat intelligence. In this paper, we propose a trigger-enhanced actionable CTI discovery system (TriCTI) to portray a relationship between IOCs and campaign stages and generate actionable CTI from cybersecurity reports through natural language processing (NLP) technology. Specifically, we introduce the “campaign trigger” for an effective explanation of the campaign stages to improve the performance of the classification model. The campaign trigger phrases are the keywords in the sentence that imply the campaign stage. The trained final trigger vectors have similar space representations with the keywords in the unseen sentence and will help correct classification by increasing the weight of the keywords. We also meticulously devise a data augmentation specifically for cybersecurity training sets to cope with the challenge of the scarcity of annotation data sets. Compared with state-of-the-art text classification models, such as BERT, the trigger-enhanced classification model has better performance with accuracy (86.99%) and F1 score (87.02%). We run TriCTI on more than 29k cybersecurity reports, from which we automatically and efficiently collect 113,543 actionable CTI. In particular, we verify the actionability of discovered CTI by using large-scale field data from VirusTotal (VT). The results demonstrate that the threat intelligence provided by VT lacks a part of the threat context for IOCs, such as the Actions on Objectives campaign stage. As a comparison, our proposed method can completely identify the actionable CTI in all campaign stages. Accordingly, cyber threats can be identified and resisted at any campaign stage with the discovered actionable CTI.
Jian Liu 0008, Yitong He, Xuren Wang, Zhengwei Jiang, Peian Yang
Cybersecur.6
2022 TIM: threat context-enhanced TTP intelligence mining on unstructured threat data
abstract
Abstract TTPs (Tactics, Techniques, and Procedures), which represent an attacker’s goals and methods, are the long period and essential feature of the attacker. Defenders can use TTP intelligence to perform the penetration test and compensate for defense deficiency. However, most TTP intelligence is described in unstructured threat data, such as APT analysis reports. Manually converting natural language TTPs descriptions to standard TTP names, such as ATT&CK TTP names and IDs, is time-consuming and requires deep expertise. In this paper, we define the TTP classification task as a sentence classification task. We annotate a new sentence-level TTP dataset with 6 categories and 6061 TTP descriptions from 10761 security analysis reports. We construct a threat context-enhanced TTP intelligence mining (TIM) framework to mine TTP intelligence from unstructured threat data. The TIM framework uses TCENet (Threat Context Enhanced Network) to find and classify TTP descriptions, which we define as three continuous sentences, from textual data. Meanwhile, we use the element features of TTP in the descriptions to enhance the TTPs classification accuracy of TCENet. The evaluation result shows that the average classification accuracy of our proposed method on the 6 TTP categories reaches 0.941. The evaluation results also show that adding TTP element features can improve our classification accuracy compared to using only text features. TCENet also achieved the best results compared to the previous document-level TTP classification works and other popular text classification methods, even in the case of few-shot training samples. Finally, the TIM framework organizes TTP descriptions and TTP elements into STIX 2.1 format as final TTP intelligence for sharing the long-period and essential attack behavior characteristics of attackers. In addition, we transform TTP intelligence into sigma detection rules for attack behavior detection. Such TTP intelligence and rules can help defenders deploy long-term effective threat detection and perform more realistic attack simulations to strengthen defense.
Yizhe You, Zhengwei Jiang, Peian Yang, Baoxu Liu, Huamin Feng, Xuren Wang
Cybersecur.3
2021 Modeling Attackers Based on Heterogenous Graph through Malicious HTTP Requests
abstract
As modern computer attacks are growing more and more complicated, there is a need for defenders to detect malicious activities and analyze which attacker or organization these attacks came from. It is a challenge to model an attacker from malicious web logs. In this paper, we modeled attacker activities based on malicious HTTP requests collected from kinds of websites, which recorded the behavior of IP addresses and provided the possibility to describe the attacker based on HTTP requests. First, we propose a novel method to get the IP address embedding through two aspects: we designed a heterogeneous graph, named IP-Domain-Graph, to capture the relation between the IP address and the domain it has sent malicious requests, and we designed an embedding method of requests content to capture the behavioral characteristics of the IP address. Then we use a similarity calculation method to cluster IP addresses to describe an attacker. The experimental results demonstrate the effectiveness of the proposed method.
Shengqin Ao, Yitong He, Xuren Wang, Zhengwei Jiang
CSCWD5
2021 Spear Phishing Emails Detection Based on Machine Learning
abstract
Spear phishing emails target to specific individual or organization, they are more elaborated, targeted, and harmful than phishing emails. The attackers usually harvest information about the recipient in any available ways, then create a carefully camouflaged email and lure the recipient to perform dangerous actions. In this paper we present a new effective approach to detect spear phishing emails based on machine learning. Firstly we extracted 21 Stylometric features from email, 3 forwarding features from Email Forwarding Relationship Graph Database(EFRGD), and 3 reputation features from two third-party threat intelligence platforms, Virus Total(VT) and Phish Tank(PT). Then we made an improvement on Synthetic Minority Oversampling Technique(SMOTE) algorithm named KM-SMOTE to reduce the impact of unbalanced data. Finally we applied 4 machine learning algorithms to distinguish spear phishing emails from non-spear phishing emails. Our dataset consists of 417 spear phishing emails and 13916 non-spear phishing emails. We were able to achieve a maximum recall of 95.56%, precision of 98.85% and 97.16% of F1-score with the help of forwarding features, reputation features and KM-SMOTE algorithm.
Xiong Ding, Baoxu Liu, Zhengwei Jiang, Qiuyun Wang, Liling Xin
CSCWD3
2021 HSRF: Community Detection Based on Heterogeneous Attributes and Semi-Supervised Random Forest
abstract
Potential connections between complex networks need to be discovered by the network community detection. Current detection methods are commonly based on homogeneous information networks, which usually extract single information among the nodes of the complex network and will lead to incomplete information or information loss. To address these problems existing on community detection, we propose a novel method based on heterogeneous attributes and semi-supervised Random Forest (HSRF) inspired by heterogeneous information networks. We define heterogeneous attribute arrays of nodes, which reflect the structural relationships between nodes in complex networks. Semi-supervised learning based on Jaccard similarity coefficients is introduced to predict the noise points and solve the problem of anti-noise interference. Our experiments on real networks and synthetic standard networks show that HSRF improves the generalization of the undirected and directed network community detection. Moreover, our HSRF performs better in terms of robustness when the community boundary structure becomes more ambiguous and convenient for parallel processing.
Zijing Fan, Liling Xin, Xuren Wang, Zhengwei Jiang, Qiuyun Wang
CSCWD5
2021 Producing More with Less: A GAN-based Network Attack Detection Approach for Imbalanced Data
abstract
Machine learning techniques are shown to be effective for network attack detection systems in identifying malicious network behaviors. In the real-world environment, however, network attack traffic i soften hidden under a large amount of normal daily communication traffic. In this paper, to resolve such challenges that the large-scale data is difficult to be effectively labeled, we propose a data augmentation method based on generative adversarial networks. The features of flow-based network traffic are firstly pre-processed to fit the generative adversarial networks (GANs). Then, we enhance the original GANs by adopting Earth-Mover (EM) distance to catch the distribution of low dimensional subspace data and add an encoder structure to learn latent space representation. Compared to other data augmentation methods, our method generates data from learning data distribution rather than performing numerical calculations on existing data. We construct an imbalanced dataset based on the real-world dataset and compare it with other methods. Our method reports better performance in terms of the recall, F1-score, and AUC, which proved the effectiveness of our proposed method.
Xingran Hao, Zhengwei Jiang, Qingsai Xiao, Qiuyun Wang, Yepeng Yao, Baoxu Liu, Jian Liu 0008
CSCWD2
2021 A Framework for Document-level Cybersecurity Event Extraction from Open Source Data
abstract
With the rapid development of the Internet, the number of cyber threats increases exponentially. More and more cyber threats come from new and unexpected sources, leading organizations and individuals to facing more security risks and vulnerabilities. Automatically obtaining and structuring security information from cybersecurity news can help security analysts to identify useful information more quickly. Most existing studies on extracting security events merely focused on the event detection task, aiming to discover and categorize cybersecurity events from the plain text. However, such event detection methods cannot capture useful information such as who performed the cyberattack, when the data breach event happened, who was the victim, etc. These arguments of a cybersecurity event are needed for analysts to get cybersecurity event details directly. Several studies have tried to extract rich semantic information of cybersecurity events, but they merely focused on extracting event arguments within the sentence scope. These studies still have limitations when the event arguments needed to recognize spread across multiple sentences. In this paper, we proposed a framework that effectively extracts cybersecurity events at the document-level from cybersecurity news, blogs and announcements. We model the document level event extraction task as a sequence tagging problem. The goal is to identify the related arguments of cybersecurity events from documents. Firstly, we get the characters embedding and incorporate the word information into the character representations. Then we design a sliding window mechanism to get the cross-sentence context information. Finally, we predict the label of each character. We build a Chinese cybersecurity dataset and use three methods to evaluate our method, and the experimental results demonstrate the effectiveness of the proposed model.
Xiangyu Du, Yitong He, Xuren Wang, Zhengwei Jiang
CSCWD6
2021 A Method for Extracting Unstructured Threat Intelligence Based on Dictionary Template and Reinforcement Learning
abstract
In recent years, individuals, organizations and countries are all threatened by cyber threats to some degree. The proposal of threat intelligence sharing scheme has greatly helped the protection of cyber security. Traditional threat intelligence sharing scheme mainly collects and analyzes information manually, which include but not limited to Indicators of Compromise (IOC) and forms a machine readable report for Security Operations Center (SOC) to take corresponding action. Therefore, it is challenging and significant to easily and automatically share and exchange cyber threat intelligence (CTI). Aiming at extracting the information of CTI efficiently, we construct a model of automatic information extraction process of the entity recognition and relationship extraction, which are used to extract effective entities and relationships in threat intelligence reports and improve the efficiency of threat intelligence sharing. The specific content and research results include two aspects: (1) Research on threat intelligence entity recognition model. We use the BERT model as a corpus pre-training model based on the classic neural network BiLSTM-CRF, and proposes a model DT-BERT-BiLSTM-CRF based on the dictionary template. The BERT pre-training model makes full use of the contextual semantic information of the corpus and alleviates the problem of ambiguity in the process of threat intelligence entity recognition. By constructing a dictionary template of threat intelligence entities, the accuracy of entity recognition in the threat intelligence field is further improved. (2) Research on the extraction of ITC relations. We constructed the relation extraction data set with distant supervision methods. For alleviating the noise annotation data, we introduce the attention mechanism and reinforcement learning into traditional neural networks, proposing a model NR-RL-PCNN-ATT. Through a new reward mechanism, our model improves the sentence selection quality and the efficiency of relationship extraction.
Xuren Wang, Binghua Song, Zhengwei Jiang, Shengqin Ao
CSCWD5
2021 Extracting Threat Intelligence Relations Using Distant Supervision and Neural Networks
Yali Luo, Shengqin Ao, Changxin Su, Peian Yang, Zhengwei Jiang
IFIP Int. Conf. Digital Forensics6
2021 FSSRE: Fusing Semantic Feature and Syntactic Dependencies Feature for threat intelligence Relation Extraction
abstract
Threat intelligence relation extraction plays an important role in threat intelligence text analysis and processing.To extract the relation between two threat entities in a sentence, we develop a novel framework called FSSRE which fuses sematic feature and syntactic dependencies feature for threat intelligence relation extraction.We utilize graph convolutional networks (GCN) to extract syntactic dependencies features, and utilize Sentence-BERT to extract contextual semantic features.To keep vital information with irrelevant content removed to the most extent, we further apply a novel pruning strategy, SDP-VP, to the input trees.With retaining the shortest path and nodes that are 𝑲 hops away from nodes on the shortest path, we give the edge connected to the verb nodes a weight of 𝒘 times.We create an advanced persistent threat (APT) intelligence entities and intra-sentence relations dataset, APTER-SENT, for that there is no public dataset can be used for relation extraction research in the threat intelligence field.Experimental results on APTER-SENT demonstrate improved performance over competitive baselines.At the same time, we also conducted experiments on the SemEval-2010 dataset.The results of the experiment indicate that our method is still effective on this dataset.
Xuren Wang, Mengbo Xiong, Famei He, Peian Yang, Binghua Song, Zhengwei Jiang, Zihan Xiong
SEKE7
2021 CAN: Complementary Attention Network for Aspect Level Sentiment Classification in Social E-Commerce
Yali Luo, Zhengwei Jiang, Peian Yang, Xuren Wang
WCNC2
2020 A Weak Coupling of Semi-Supervised Learning with Generative Adversarial Networks for Malware Classification
abstract
Malware classification helps to understand its purpose and is also an important part of attack detection. And it is also an important part of discovering attacks. Due to continuous innovation and development of artificial intelligence, it is a trend to combine deep learning with malware classification. In this paper, we propose an improved malware image rescaling algorithm (IMIR) based on local mean algorithm. Its main goal of IMIR is to reduce the loss of information from samples during the process of converting binary files to image files. Therefore, we construct a neural network structure based on VGG model, which is suitable for image classification. In the real world, a mass of malware family labels are inaccurate or lacking. To deal with this situation, we propose a novel method to train the deep neural network by Semi-supervised Generative Adversarial Network (SGAN), which only needs a small amount of malware that have accurate labels about families. By integrating SGAN with weak coupling, we can retain the weak links of supervised part and unsupervised part of SGAN. It improves the accuracy of malware classification by making classifiers more independent of discriminators. The results of experimental demonstrate that our model achieves exhibiting favorable performance. The recalls of each family in our data set are all higher than 93.75%.
Shuwei Wang, Qiuyun Wang, Zhengwei Jiang, Xuren Wang, Rongqi Jing
ICPR3
2020 Towards Comprehensive Detection of DNS Tunnels
abstract
The Domain Name System (DNS) is a fundamental service of the Internet, and the DNS tunnel is one of the most threatening abuses of DNS, posing a huge threat to user privacy and Internet security. Attackers conceal the information into DNS packets to evade firewalls and intrusion detection systems. Recently, newly developed DNS tunnels used by Advanced Persist Threat groups tend to use A and AAAA resource records (RRs) for transmission, making them more invisible and more threatening. Previous DNS tunnel detection approaches mainly focus on subdomains and TXT RRs, but less attention has been paid to newly developed DNS tunnels based on A and AAAA RRs. In this paper, we present a novel DNS tunnel detection method that can detect newly developed A and AAAA RR based DNS tunnels. Since DNS tunnels will transmit a large amount of encrypted or encoded data in the DNS queries and responses, we extracted novel features from domains and 4 types of RRs (A, AAAA, TXT and CNAME RRs) that are most commonly used for tunneling to measure the amount and content of information exchanged between the authoritative nameservers and the clients. We also analyze the detection capabilities when different features were used. The anomaly detection algorithm is employed on domains related features and 4 types of RRs related features, respectively. The overlaps of outliers will be marked as DNS tunnels. Our approach has been evaluated on real-world network traffic. The experimental results show that our approach can detect all DNS tunnels in the dataset with a extremely low false positive rate.
Meng Luo 0006, Qiuyun Wang, Yepeng Yao, Xuren Wang, Peian Yang, Zhengwei Jiang
ISCC6
2020 NER in Threat Intelligence Domain with TSFL
Xuren Wang, Zihan Xiong, Xiangyu Du, Zhengwei Jiang, Mengbo Xiong
NLPCC (1)5
2020 MTLAT: A Multi-Task Learning Framework Based on Adversarial Training for Chinese Cybersecurity NER
Yaopeng Han, Zhigang Lu 0002, Bo Jiang 0013, Zhengwei Jiang
NPC6
2020 DNRTI: A Large-scale Dataset for Named Entity Recognition in Threat Intelligence
abstract
Named entity recognition is an important and challenging problem in Natural language processing. Although the past decade has witnessed major advances in entity recognition in many fields, such successes have been slow to network security field, not only because of the data in the network security field is very professional, but also due to the sensitive information in the data. To advance named entity recognition research in network security field, we introduce a large-scale Dataset for Named Entity Recognition in Threat Intelligence (DNRTI). To this end, we collect more than 300 pieces of threat intelligence. The data in DNRTI is all annotated by experts in threat intelligence interpretation using 13 object categories. The fully annotated DNRTI contains 175220 words. To build a baseline for named entity recognition in the threat intelligence field, we evaluate some deep learning model on DNRTI. Experiments demonstrate that DNRTI well represents the key information in threat intelligence and are quite challenging.
Xuren Wang, Xinpei Liu, Shengqin Ao, Zhengwei Jiang, Zongyi Xu, Zihan Xiong, Mengbo Xiong
TrustCom5
2020 Joint Learning for Document-Level Threat Intelligence Relation Extraction and Coreference Resolution Based on GCN
abstract
In order to help researchers quickly understand the connection between new threat events and previous threat events, threat intelligence document-level relation extraction plays a very important role in threat intelligence text analysis and processing. Because there is no public document-level threat intelligence dataset, we create APTERC-DOC, an APT intelligence entities, relations and coreference dataset. We treat the relation extraction as a multi-classification task. Treating the coreference relation as a kind of predefined relations, we develop a joint learning framework called TIRECO, a model which can simultaneously complete threat intelligence relation extraction and coreference resolution. In order to solve the problem of document-level text being too long to extract feature, we propose the concept of sentence set, which transforms document-level relation extraction into inter-sentence relation extraction. To incorporate relevant information with maximally removing irrelevant content in sentence set, we further apply a novel pruning strategy (SDP-VP-SET) to the input trees considering that verbs are crucial in determining the relation between entities in sentence set. With retaining the shortest path and nodes that are K hops away from the shortest path, we give the edge connected to the verb nodes a weight of w times. Experimental results show that our model not only performs well in the extraction of inter-sentence relations, it is also effective in intra-sentence relations, and the F1 value has increased by 15.694%.
Xuren Wang, Mengbo Xiong, Yali Luo, Zhengwei Jiang, Zihan Xiong
TrustCom5
2020 A DGA domain names detection modeling method based on integrating an attention mechanism and deep neural network
abstract
Abstract Command and control (C2) servers are used by attackers to operate communications. To perform attacks, attackers usually employee the Domain Generation Algorithm (DGA), with which to confirm rendezvous points to their C2 servers by generating various network locations. The detection of DGA domain names is one of the important technologies for command and control communication detection. Considering the randomness of the DGA domain names, recent research in DGA detection applyed machine learning methods based on features extracting and deep learning architectures to classify domain names. However, these methods are insufficient to handle wordlist-based DGA threats, which generate domain names by randomly concatenating dictionary words according to a special set of rules. In this paper, we proposed a a deep learning framework ATT-CNN-BiLSTM for identifying and detecting DGA domains to alleviate the threat. Firstly, the Convolutional Neural Network (CNN) and bidirectional Long Short-Term Memory (BiLSTM) neural network layer was used to extract the features of the domain sequences information; secondly, the attention layer was used to allocate the corresponding weight of the extracted deep information from the domain names. Finally, the different weights of features in domain names were put into the output layer to complete the tasks of detection and classification. Our extensive experimental results demonstrate the effectiveness of the proposed model, both on regular DGA domains and DGA that hard to detect such as wordlist-based and part-wordlist-based ones. To be precise,we got a F1 score of 98.79% for the detection and macro average precision and recall of 83% for the classification task of DGA domain names.
Fangli Ren, Zhengwei Jiang, Xuren Wang
Cybersecur.2
2020 THS-IDPC: A three-stage hierarchical sampling method based on improved density peaks clustering algorithm for encrypted malicious traffic detection
Liangchen Chen, Shu Gao, Baoxu Liu, Zhigang Lu 0002, Zhengwei Jiang
J. Supercomput.5
2019 Integrating an Attention Mechanism and Deep Neural Network for Detection of DGA Domain Names
abstract
Domain generation algorithms (DGA) are employed by malware to generate domain names as a common practice, with which to confirm rendezvous points to their command-and-control (C2) servers. The detection of DGA domain names is one of the important technologies for command and control communication detection. Considering the randomness of the DGA domain names, recent work in DGA detection employed machine learning methods based on features extracting and deep learning architectures to classify domain names. However, these methods perform poorly on wordlist-based DGA families, which generate domain names by randomly concatenating dictionary words. In this paper, we proposed the ATT-CNN-BiLSTM model to detect and classify DGA domain names. Firstly, the Convolutional Neural Network (CNN) and bidirectional Long Short-Term Memory (BiLSTM) neural network layer was used to extract the features of the domain sequences information; secondly, the attention layer was used to allocate the corresponding weight of the extracted domain deep information. Finally, the domain feature messages of different weights were put into the output layer to complete the tasks of detection and classification. The experiment results demonstrate the effectiveness of the proposed model both on regular DGA domain names and wordlist-based ones. To be precise, we got a F1 score of 98.92% for the detection and macro average F1 score of 81% for the classification task of DGA domain names.
Fangli Ren, Zhengwei Jiang
ICTAI2
2019 PRTIRG: A Knowledge Graph for People-Readable Threat Intelligence Recommendation
Zhengwei Jiang, Zhigang Lu 0002, Xiangyu Du
KSEM (1)3
2010 MIMO Channel Estimation Using the Variational Expectation-Maximization Method
abstract
In this paper, the variational expectation-maximization (VEM) algorithm is used to provide channel estimates in a space-time decoder. The proposed estimator works for many types of space-time codes (STC), including full-rate full-diversity (FRFD) codes, Bell Laboratories layered space-time architecture (BLAST) and orthogonal STC's. The principle idea is to treat the channel coefficients as the unknown parameters, and the transmitted symbols as the unobserved variables in the EM algorithm. The posterior distribution of the symbols is approximated by a factorized distribution, whose Kullback-Liebler divergence with the true distribution is minimized, thereby explaining the ``variational'' aspect of the technique. The channel estimates may then be used for coherent decoding -- here we use the K-best detector for a 2×2 system. Simulation results show that the new channel estimator can reduce system complexity greatly compared with traditional schemes, while achieving near-ideal performance.
Zhengwei Jiang, Teng Joon Lim, Roya Doostnejad, Taiwen Tang
VTC Fall1