EDBT 2026 Demo / reviewers in the wild / expert
Runzi Zhang
dblp:210/3288
· DBLP profile ↗
15ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-2929-2484ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Apmp: APT attack detection in few-shot scenarios based on entity potential relationsabstractAbstract With the rapid development of information technology, advanced persistent threats (APTs) have led to numerous serious data breaches and information system disruptions, causing immense losses to governments, businesses, and individuals. APT attack activities are usually carried out stealthily and often require analyzing large amounts of audited data, making it difficult to handle APT attacks promptly. Existing work attempts to improve the handling and detection efficiency of APT attacks based on limited audit data. However, these methods only increase the number of attack samples by finding suspicious entities through rules, ignoring the attack features contained in potential relations between entities. In this paper, we propose a potential relation prediction-based method (APMP) for APT attack detection in few-shot scenarios, which exploits potential relations to find ignored attack features. Specifically, APMP extracts the information between entities and relations in the attack sequence to train the prediction model. The prediction model can predict potential relations between entities and map them into the provenance graph. In this way, APMP complements the potential relations between entities in the provenance graph and captures the attack-related information between entities, improving the results of attack detection. We evaluate APMP using ten real-world public APT attack datasets. The average evaluation precision of APMP attack detection is 100%, with a recall rate of 93.18% and an F1-score of 96.30%. The results show that our proposal can effectively detect APT attacks in few-shot scenarios. Tong Li 0001, Runzi Zhang, Zilong Wan, Zhen Yang 0004 |
Cybersecur. | 3 |
| 2025 | Enterprise Threat Detection with Explainable AI for Cloud-Network ConvergenceabstractWith the rapid development of cloud-network convergence, alert data processing has become a critical task in intelligent cloud-network operations and maintenance. Enterprise threat detection systems serve as a vital defense against cyberattacks. However, due to well-known challenges such as massive unlabeled data, false alerts, and lack of interpretability, threat detection remains a formidable challenge. To address these issues, we propose the Enterprise Threat Detection Explainer (ETDE), an interpretable graph neural network (GNN)-based framework designed to automatically identify attackers from intrusion prevention system (IPS) alerts. ETDE extracts relevant subgraph structures and features from security attribute graphs to model attack patterns, while providing comprehensive explanations to help security teams rapidly pinpoint threats and accelerate incident response. Evaluations on real-world operational data demonstrate that ETDE significantly reduces the manual workload for security analysts, enhancing both efficiency and reliability in threat detection. Fudi Wu, Xingwang Huang, Qiefu Wuri, Dujuan Gu, Runzi Zhang, Shen Gao |
ICNP | 6 |
| 2025 | A Context-Aware Clustering Approach for Assisting Operators in Classifying Security AlertsabstractModern software has evolved from delivering software products to web services and applications, which need to be protected by security operation centers (SOC) against ubiquitous cyber attacks. Numerous security alerts are continuously generated every day, which have to be efficiently and correctly processed to identify potential threats. Many AIOps (artificial intelligence for IT operations) approaches have been proposed to (semi-)automate the inspection of alerts so as to reduce manual effort as much as possible. However, due to the ever-complicating attacks, a significant amount of manual work is still required in practice to ensure correct analysis results. In this paper, we propose a Context-Aware cLustering approach for cLassifying sEcurity alErts (CALLEE), which fully exploits the rich relationships among alerts in order to precisely identify similar alerts, significantly reducing the workload of SOC. Specifically, we first design a core conceptual model to capture connections among security alerts, based on which we establish corresponding heterogeneous information networks. Next, we systematically design a set of meta-paths to profile typical alert scenarios precisely, contributing to obtaining the representation of security alerts. We then cluster security alerts based on their contextual similarities, considering the tradeoff between the number of clusters and the homogeneity of each cluster. Finally, security operators only need to manually inspect a limited number of alerts within each cluster, pragmatically reducing their workload while ensuring the accuracy of alert classification. To evaluate the effectiveness of our approach, we collaborate with our industrial partner and pragmatically apply the approach to a real alert dataset. The results show that our approach can reduce the workload of SOC by 99.76%, outperforming baseline approaches. In addition, we further investigate the integration of our proposal with the real business scenario of our industrial partner. The feedback from practitioners shows that CALLEE is pragmatically applicable and helpful in industrial settings. Yu Liu 0090, Tong Li 0001, Runzi Zhang, Mingkai Tong, Wenmao Liu, Zhen Yang 0004 |
IEEE Trans. Software Eng. | 3 |
| 2024 | VCRLog: Variable Contents Relationship Perception for Log-based Anomaly DetectionabstractLog-based anomaly detection is crucial for software reliability assurance. System logs are semi-structured data containing constant and variable contents, both of which can provide valuable features for anomaly detection. Due to variables being heterogeneous and discrete, there is a lack of effective approaches that can comprehensively incorporate features of variables into log-based anomaly detection. In this paper, we propose VCRLog, an anomaly detection method that mines the relationships among the heterogeneous and discrete variables and extracts important features contributing to anomaly detection. Firstly, considering parsing methods cannot accurately extract variables from logs, we propose a variable extraction method based on domain knowledge. Secondly, to capture and extract the relationship feature among heterogeneous and discrete variables, we design a conceptual model based on system operation to construct variable attributed graph, which can mine important feature vectors by structural embeddings. Finally, considering constants directly express the meaning of logs, we combine relationship vectors with semantic vectors of constants to achieve transformer-based anomaly detection. Experimental results show that our proposed method can accurately detect anomalies and maintain high accuracy as the training data size decreases, outperforming existing methods. Our source code and experimental data are publicly available at https://github.com/Fridaywjy/VCRLog. Jin-Yuan Wang, Tong Li 0001, Runzi Zhang, Zifang Tang, Di Wu 0064, Zhen Yang 0004 |
ISSRE | 3 |
| 2024 | Detecting APT attacks using an attack intent-driven and sequence-based learning approach
Tong Li 0001, Di Wu 0064, Runzi Zhang, Zhen Yang 0004 |
Comput. Secur. | 4 |
| 2023 | APM: An Attack Path-based Method for APT Attack Detection on Few-Shot LearningabstractAdvanced persistent threat (APT) attack leverages various intelligence-gathering techniques to obtain sensitive and critical information, imposing increasing threats to modern software enterprises. However, due to the persistent presence of APT attacks, it is difficult to effectively analyze a large amount of audit data for detecting such attacks, especially for small and medium-sized enterprises (SMEs). This limitation hinders security operation centers (SOC) from promptly handling APT attacks. In this paper, we propose an attack path-based method (APM) for APT attack detection on few-shot learning. Specifically, APM first identifies candidate malicious entities from the provenance graph, contributing to the completion of the missing attack paths. Secondly, we propose a systematic method to exploit potential attack behaviors in the attack path based on the identified candidate malicious entities. We evaluate APM through five APT attacks in realistic environments. Compared to existing baselines, the precision, recall, and F1-score of APM for attack detection increased by 0.28%, 1.64%, and 1.13%, respectively. The results show that our proposal can outperform baseline approaches and effectively detect APT attacks based on few-shot learning. Tong Li 0001, Runzi Zhang, Di Wu 0064, Zhen Yang 0004 |
TrustCom | 3 |
| 2022 | LogTracer: Efficient Anomaly Tracing Combining System Log Detection and Provenance GraphabstractInformation systems have penetrated into all areas of social life, however, unknown threats represented by APT attacks pose serious challenges to their security. In recent years, approaches based on log analysis and provenance graph have been extensively used in the anomaly detection and tracing of malicious attacks. However, traditional method has low detection accuracy, high complexity and low efficiency. To address those shortcomings, we propose an efficient anomaly tracing approach (LogTracer), which combines system log detection and provenance graph together. The proposed LogTracer extracts the attack path from provenance graph, which is constructed with the anomaly degrees of the system logs anomaly detection results. Compar-ative experiments with OmegaLog, NoDoze and ALchemist are conducted on a simulated dataset with 16 attack types totaling 290 million logs. The experimental results show that our method approximately 5.4x, 0.2x and 7.2x faster than these three methods in processing efficiency, and its malicious node coverage rate reaches 98.1%. Weina Niu, Zhenqi Yu, Zimu Li, Beibei Li 0002, Runzi Zhang, Xiaosong Zhang 0001 |
GLOBECOM | 5 |
| 2022 | A Novel Network Alert Classification Model based on Behavior Semantic
Zhanshi Li, Tong Li 0001, Runzi Zhang, Di Wu 0064, Zhen Yang 0004 |
SEKE | 3 |
| 2022 | Trine: Syslog anomaly detection with three transformer encoders in one generative adversarial network
Zhenfei Zhao, Weina Niu, Xiaosong Zhang 0001, Runzi Zhang, Zhenqi Yu, Cheng Huang 0003 |
Appl. Intell. | 4 |
| 2022 | Context2Vector: Accelerating security event triage via context representation learning
Runzi Zhang, Wenmao Liu, Dujuan Gu, Mingkai Tong, Jianxin Xue, Huanran Wang |
Inf. Softw. Technol. | 2 |
| 2021 | Integrating Heterogeneous Security Knowledge Sources for Comprehensive Security AnalysisabstractWith the fast growth of system complexity, it is increasingly difficult to comprehensively analyze security of such large-scale systems, which is a knowledge-intensive task. Although there are various available security knowledge sources, they are not well-connected with each other due to their heterogeneity and unstructured descriptions. In this paper, we propose a systematic approach to construct a comprehensive and reusable knowledge graph in the field of information security. Specifically, we first investigate heterogeneous security knowledge sources and establish a detailed ontology of information security, integrating various security conceptual models. Then, we train a security entity identifier based on active learning to extract security knowledge from unstructured descriptions. Such extracted knowledge is then fused to establish a comprehensive and reusable security knowledge graph based on the unified ontology. Finally, we illustrate the utility of our established knowledge graph with a set of exemplary queries and reasoning rules in the context of a real security scenario. Guodi Wang, Tong Li 0001, Zhen Yang 0004, Runzi Zhang |
COMPSAC | 5 |
| 2021 | An Evolutionary Study of IoT MalwareabstractRecent years have witnessed lots of attacks targeted at the widespread Internet of Things (IoT) devices and malicious activities conducted by compromised IoT devices. After some notorious IoT malware released their source code, many new variants emerge, which are usually more powerful and stealthy. Although numerous existing studies have analyzed some exposed families, there is a lack of systematic study to make full use of them, which can be a fundamental step for provenance, triage, labeling, lineage analysis, and authorship attribution. The key challenge of conducting an IoT malware evolutionary study is how to collect sufficient and accurate information about malware and identify the relationships among them. In this article, we take the first step to investigate the IoT malware evolution by leveraging the information from two sources that complement each other. First, we crawl online articles about IoT malware and employ natural language processing techniques to extract the features of malware samples and their relationships with other malware family, which allow us to form the basic lineage graph. Second, we collect real malware samples through our widely deployed honeypots and design a new classifier to group them into families and identify lineage relationships among them. Such results are used to enhance the basic lineage graph. Eventually, we construct the final lineage graph for 72 IoT malware families by correlating the information from the aforementioned sources, which can help the research community better understand and fight IoT malware now and in the future. Our study has been incorporated into the threat awareness system of NSFOCUS company. Huanran Wang, Weizhe Zhang, Peng Liu 0005, Xiapu Luo, Yang Liu 0039, Yan Li 0075, Wenmao Liu, Runzi Zhang, Xing Lan |
IEEE Internet Things J. | 11 |
| 2021 | A GAN and Feature Selection-Based Oversampling Technique for Intrusion DetectionabstractIn recent years, there have been numerous cyber security issues that have caused considerable damage to the society. The development of efficient and reliable Intrusion Detection Systems (IDSs) is an effective countermeasure against the growing cyber threats. In modern high-bandwidth, large-scale network environments, traditional IDSs suffer from a high rate of missed and false alarms. Researchers have introduced machine learning techniques into intrusion detection with good results. However, due to the scarcity of attack data, such methods’ training sets are usually unbalanced, affecting the analysis performance. In this paper, we survey and analyze the design principles and shortcomings of existing oversampling methods. Based on the findings, we take the perspective of imbalance and high dimensionality of datasets in the field of intrusion detection and propose an oversampling technique based on Generative Adversarial Networks (GAN) and feature selection. Specifically, we model the complex high-dimensional distribution of attacks based on Gradient Penalty Wasserstein GAN (WGAN-GP) to generate additional attack samples. We then select a subset of features representing the entire dataset based on analysis of variance, ultimately generating a rebalanced low-dimensional dataset for machine learning training. To evaluate the effectiveness of our proposal, we conducted experiments based on the NSL-KDD, UNSW-NB15, and CICIDS-2017 datasets. The experimental results show that our method can effectively improve the detection performance of machine learning models and outperform the baselines. Xiaodong Liu 0010, Tong Li 0001, Runzi Zhang, Di Wu 0064, Yongheng Liu, Zhen Yang 0004 |
Secur. Commun. Networks | 3 |
| 2020 | Far from classification algorithm: dive into the preprocessing stage in DGA detectionabstractDomain-Flux technique has been widely used by attackers to maintain a botnet for many years and the core of it is the adoption of domain generation algorithm (DGA). To combat attackers, there are lots of works in DGA domain detection area recently. But they usually collect quite limited data and conduct experiments in a closed dataset, meaning that the DGA data and the benign data they collected can not well represent the real distribution between them. Moreover, they handle the domains roughly and use the origin data to train the classifier directly, which is also not adequate to classify these two types of domains with lots of false positives and false negatives happening during the real-world deployment. In this paper, we conduct the first large-scale DGA domain analysis in traffic level and argue that the preprocessing stage is also vital for the final classifier, which is usually ignored by the existing works. We collect the largest amount of DGA domain data than prior works and collect DNS log offered by a big company, whose DNS data covers most important industries in China. Based on this data, we analyze the distribution of DGA domains in traffic and give quantifiable results showing that NXDomain (domain not exist) is more suitable for DGA detection. Moreover, we give detailed preprocessing steps to handle the original domains. Our experiment shows that with the preprocessing stage mentioned above, classifier performs better in DGA detection task. Our research indicates that improving the classification algorithm is far from enough in DGA detection and the preprocessing stage is also the key component in bringing the DGA detection methods from lab to product. Mingkai Tong, Runzi Zhang, Jianxin Xue, Wenmao Liu, Jiahai Yang 0001 |
TrustCom | 3 |
| 2020 | CMIRGen: Automatic Signature Generation Algorithm for Malicious Network TrafficabstractAlthough machine learning (ML) based solutions are ever-evolving for the attack defending paradigm, signatures of malicious network traffic are vital resources for intrusion detection systems (IDSs) and network forensic procedure, covering the lack of interpretability and stability for ML models. However, signature extraction is still a time and labor consuming task nowadays, resulting in possible increase of the attackers' dwell time. Existing automatic solutions rely too much on sequence similarity based and heuristic based methods, encountering performance degradation in large scale and dynamic network environment. In this paper, we present a novel method, called Clustering and Model Inference-based Rule Generation (CMIRGen), automatically generating token-set based signature rules for malicious traffic payloads to be inspected. CMIRGen leverages both optimized sequence similarity based and black-box model inference based methods to extract patterns from homogeneous and heterogeneous payloads respectively. Experimental evaluations have been conducted on several datasets and show the CMIRGen framework can extract discriminative signatures, presenting high recall rate and low false positive rate at the same time for malicious content recognition. Runzi Zhang, Mingkai Tong, Jianxin Xue, Wenmao Liu |
TrustCom | 1 |