EDBT 2026 Demo / reviewers in the wild / expert
Zhenyuan Li
dblp:252/3799
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 11 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Sands to Mansions: Actionable, Customizable and Causality-Preserving Cyberattack Emulation with LLM-Powered Symbolic Planning
Lingzhi Wang 0002, Zhenyuan Li, Zhengkai Wang, Xiangmin Shen, Yan Chen 0004 |
ACNS (3) | 2 |
| 2026 | Breaking the Bulkhead: Demystifying Cross-Namespace Reference Vulnerabilities in Kubernetes Operators
Zhaoxuan Jin, Zhenyuan Li, Yan Chen 0004 |
NDSS | 4 |
| 2026 | Incorporating Gradients to Rules: Toward Online, Adaptive Provenance-Based Intrusion DetectionabstractAs cyber-attacks become increasingly sophisticated and stealthy, accurately distinguishing between benign behavior and malicious intrusions has become both more critical and more challenging. Provenance-based intrusion detection systems (PIDS) show strong potential for detecting malicious activities through fine-grained causality analysis, which has gained significant attention from both industry and academia. Among the various PIDS approaches, rule-based systems are particularly favored for their low overhead, real-time detection capability, and interpretability. However, these systems face challenges in reducing false positive rates, primarily due to the lack of fine-tuned rules and specific environments. In this paper, we introduce CAPTAIN+, a rule-based PIDS that autonomously adapts to diverse environments online. Specifically, we propose three adaptive parameters to adjust the detection configuration for nodes, edges, and alarm generation thresholds. Initially, we build a differentiable tag propagation framework and utilize the gradient descent algorithm to optimize these adaptive parameters based on the training data. In this extended version, we integrate an online learning module into the detection stage to dynamically optimize adaptive parameters based on real-time feedback from the detection process. We evaluate CAPTAIN+ based on data from DARPA TC, OpTC datasets, and PKU ASAL datasets. The results demonstrate that CAPTAIN+ offers superior detection accuracy, lower detection latency, reduced runtime overhead, long-term resilience against concept drift, and more interpretable detection results compared to state-of-the-art PIDS. Zhenyuan Li, Lingzhi Wang 0002, Zhengkai Wang, Xiangmin Shen, Haitao Xu 0002, Yan Chen 0004, Shouling Ji |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
Xiangmin Shen, Lingzhi Wang 0002, Zhenyuan Li, Yan Chen 0004, Wencheng Zhao, Jiashui Wang |
AsiaCCS | 3 |
| 2025 | The Case for Learned Provenance-based System Behavior BaselineabstractProvenance graphs describe data flows and causal dependencies of host activities, enabling to track the data propagation and manipulation throughout the systems, which provide a foundation for intrusion detection. However, these Provenance-based Intrusion Detection Systems (PIDSes) face significant challenges in storage, representation, and analysis, which impede the efficacy of machine learning models such as Graph Neural Networks (GNNs) in processing and learning from these graphs. This paper presents a novel learning-based anomaly detection method designed to efficiently embed and analyze large-scale provenance graphs. Our approach integrates dynamic graph processing with adaptive encoding, facilitating compact embeddings that effectively address out-of-vocabulary (OOV) elements and adapt to normality shifts in dynamic real-world environments. Subsequently, we incorporate this refined baseline into a tag-propagation framework for real-time detection. Our evaluation demonstrates the method’s accuracy and adaptability in anomaly path mining, significantly advancing the state-of-the-art in handling and analyzing provenance graphs for anomaly detection. Zhenyuan Li, Yangyang Wei, Shouling Ji |
ICML | 2 |
| 2025 | Understanding the Business of Online Affiliate Marketing: An Empirical StudyabstractAffiliate marketing is a revenue-sharing marketing scheme by which an affiliate, such as a blogger or YouTuber, garners commissions for promoting a merchant's goods or services, thereby aiming to foster a mutually beneficial relationship between affiliates and merchants. Despite being a multi-billion-dollar global industry, affiliate marketing remains inadequately explored, and the research community lacks a comprehensive understanding of its intricate ecosystem. In this paper, we present the first comprehensive empirical study of the affiliate marketing ecosystem. We conduct thorough measurements to assess the prevalence of affiliate marketing, estimate the market size, and elucidate the characteristics of affiliates, merchants, and intermediary affiliate networks. Over a continuous span of 13 months, we monitored four of the most prominent affiliate aggregation platforms, yielding a substantial dataset. We observed 467,219 unique offers - tasks to be undertaken by affiliates - involving 37,109 merchants and 556 affiliate networks across the four platforms. Notably, these offers would cost the merchants more than 19 million USD for the completion of all the actions pre-defined in these offers, such as signing up or making a transaction. Additionally, we compiled a large-scale dataset comprising 124,462 affiliate links, enabling us to conduct a comprehensive investigation. Finally, we propose machine learning models incorporating the characteristics of affiliate links to detect real-world affiliate marketing campaigns. Haitao Xu 0002, Kaleem Ullah Qasim, Shuai Hao 0001, Wenrui Ma, Zhenyuan Li, Fan Zhang 0010, Zhao Li 0007 |
INFOCOM | 6 |
| 2025 | Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection
Lingzhi Wang 0002, Xiangmin Shen, Weijian Li 0002, Zhenyuan Li, R. Sekar 0001, Han Liu 0001, Yan Chen 0004 |
NDSS | 4 |
| 2025 | AutoSeg: Automatic micro-segmentation policy generation via configuration analysis
Zhaoxuan Jin, Zhenyuan Li, Yan Chen 0004 |
Comput. Secur. | 3 |
| 2024 | Exploring Depths of WebAudio: Advancing Greybox Fuzzing for Vulnerability Detection in SafariabstractWebAudio is a widely used audio processing API in popular browsers, which provides rich audio support for the exclusive browser Safari on macOS. Given its widespread use, it is critical to thoroughly test WebAudio to ensure its reliability. Traditional fuzzing techniques typically lack awareness of the input structure and fail to accommodate the unique characteristics of audio file formats, and cannot generate effective fuzzing input, thus falling short of effectively detecting vulnerabilities within WebAudio. In this work, we introduce Proteus, an advanced greybox fuzzer designed to achieve structure awareness through the use of input templates. Moreover, Proteus is equipped with high-level mutation operators, diverging from traditional bit-level manipulations, and incorporates a post-processing stage that repairs format constraints disrupted during mutation. These enhancements enable Proteus to explore new input domains effectively while maintaining file validity, significantly improving the depth and efficiency of the fuzzing process. Our evaluation confirms the effectiveness of Proteus. In the experiment of fuzzing WebAudio using CAF files, our tool exposed significantly more vulnerabilities than the baseline Honggfuzz without compromising efficiency. Excitingly, we have identified a vulnerability that can be exploited to gain control of the browser. Generally, Proteus has discovered 36 zero-day vulnerabilities in WebAudio on macOS 10.15.3, with 11 of these assigned CVEs. Jiashui Wang, Jundong Xie, Zhenyuan Li, Yan Chen 0004 |
APSEC | 4 |
| 2024 | Decoding the MITRE Engenuity ATT&CK Enterprise Evaluation: An Analysis of EDR Performance in Real-World EnvironmentsabstractEndpoint detection and response (EDR) systems have emerged as a critical component of enterprise security solutions, effectively combating endpoint threats like APT attacks with extended lifecycles. In light of the growing significance of endpoint detection and response (EDR) systems, many cybersecurity providers have developed their own proprietary EDR solutions. It's crucial for users to assess the capabilities of these detection engines to make informed decisions about which products to choose. This is especially urgent given the market's size, which is expected to reach around 3.7 billion dollars by 2023 and is still expanding. MITRE is a leading organization in cyber threat analysis. In 2018, MITRE started to conduct annual APT emulations that cover major EDR vendors worldwide. Indicators include telemetry, detection and blocking capability, etc. Nevertheless, the evaluation results published by MITRE don't contain any further interpretations or suggestions. Xiangmin Shen, Zhenyuan Li, Graham Burleigh, Lingzhi Wang 0002, Yan Chen 0004 |
AsiaCCS | 2 |
| 2024 | Expensive Optimization Based on Evolutionary Multi-Tasking and Hybrid Restart StrategyabstractEvolutionary Algorithms (EAs) can not handle expensive optimization problems (EOPs) well due to the limited function evaluations in EOPs. To address this challenge, surrogate-assisted evolutionary algorithms (SAEAs) have been widely used and obtained good performance. With the problem dimension increases, SAEAs encounter some challenges in relatively high complexity on the training time and prediction time. To address this, this article proposes a novel expensive optimization algorithm with evolutionary multi-tasking and hybrid restart strategy (HRS-EMT). In the surrogate model construct, two radial basis function (RBF) models with different kernel functions are trained on all evaluated data to provide diversity and then are solved by a multi-tasking optimizer to a better optimization performance. In the surrogate model management, HRS-EMT combines multiple RBF models into an ensemble RBF (ERBF) model, strategically applied in the initial population pre-selection of surrogate model. Based on the prediction of ERBF surrogate model, HRS-EMT can obtain a better initial population in high dimensions. HRS-EMT is validated on twelve benchmark functions and compared with other state-of-the-art SAEAs. Experimental studies have shown the superior or comparable performance to other popular SAEAs in addressing EOPs. Zhenyuan Li, Xiaoliang Ma 0001, Zexuan Zhu 0001, Yueyue Li |
CEC | 1 |
| 2024 | An Automated Alert Cross-Verification System with Graph Neural Networks for IDS EventsabstractIntrusion Detection Systems (IDSs) are vital in detecting network attacks and ensuring the confidentiality and integrity of network resources. Currently, industry-standard IDSs primarily rely on rule-based or anomaly-detection techniques. However, existing detection techniques often generate false positives and negatives, also known as the alert fatigue problem. This influx of incorrect events diminishes the IDS’s efficiency by overburdening security analysts. In this paper, we present ACVS, an innovative automated alert cross-verification system that leverages Graph Neural Networks for identifying misclassifications in security events. Initially, ACVS generates event graphs using attributes like IP addresses and timestamps from sequences of security events and then employs correlation analysis on these events, utilizing alert information to verify misclassifications. Finally, the system uses Graph Neural Networks to classify and correct these security events automatically. We conduct evaluations for ACVS on a substantial real-world dataset comprising over 5 million security events, which are categorized into 5 distinct groups. The results reveal that ACVS markedly enhances the accuracy of intrusion detection systems and substantially reduces the need for manual analysis. Yuanhui He, Feiyang Huang, Ziming Zhao 0008, Zhuoxue Song, Zhenyuan Li, Fan Zhang 0010 |
CSCWD | 7 |
| 2024 | Formal Verification Techniques for Post-quantum Cryptography: A Systematic Review
Yuexi Xu, Zhenyuan Li, Naipeng Dong, Veronika Kuchta, Dongxi Liu |
ICECCS | 2 |
| 2024 | Toward Dynamic Resource Allocation and Client Scheduling in Hierarchical Federated Learning: A Two-Phase Deep Reinforcement Learning ApproachabstractFederated learning (FL) is a viable technique to train a shared machine learning model without sharing data. Hierarchical FL (HFL) system has yet to be studied regrading its multiple levels of energy, computation, communication, and client scheduling, especially when it comes to clients relying on energy harvesting to power their operations. This paper presents a new two-phase deep deterministic policy gradient (DDPG) framework, referred to as “TP-DDPG”, to balance online the learning delay and model accuracy of an FL process in an energy harvesting-powered HFL system. The key idea is that we divide optimization decisions into two groups, and employ DDPG to learn one group in the first phase, while interpreting the other group as part of the environment to provide rewards for training the DDPG in the second phase. Specifically, the DDPG learns the selection of participating clients, and their CPU configurations and the transmission powers. A new straggler-aware client association and bandwidth allocation (SCABA) algorithm efficiently optimizes the other decisions and evaluates the reward for the DDPG. Experiments demonstrate that with substantially reduced number of learnable parameters, the TP-DDPG can quickly converge to effective polices that can shorten the training time of HFL by 39.4% compared to its benchmarks, when the required test accuracy of HFL is 0.9. Xiaojing Chen 0001, Zhenyuan Li, Wei Ni 0001, Xin Wang 0003, Shunqing Zhang, Yanzan Sun, Shugong Xu, Qingqi Pei |
IEEE Trans. Commun. | 2 |
| 2022 | AttacKG: Constructing Technique Knowledge Graph from Cyber Threat Intelligence Reports
Zhenyuan Li, Jun Zeng 0006, Yan Chen 0004, Zhenkai Liang |
ESORICS (1) | 1 |
| 2022 | Generic, efficient, and effective deobfuscation and semantic-aware attack detection for PowerShell scriptsabstractIn recent years, PowerShell has increasingly been reported as appearing in a variety of cyber attacks. However, because the PowerShell language is dynamic by design and can construct script fragments at different levels, state-of-the-art static analysis based PowerShell attack detection approaches are inherently vulnerable to obfuscations. In this paper, we design the first generic, effective, and lightweight deobfuscation approach for PowerShell scripts. To precisely identify the obfuscated script fragments, we define obfuscation based on the differences in the impacts on the abstract syntax trees of PowerShell scripts and propose a novel emulation-based recovery technology. Furthermore, we design the first semantic-aware PowerShell attack detection system that leverages the classic objective-oriented association mining algorithm and newly identifies 31 semantic signatures. The experimental results on 2342 benign samples and 4141 malicious samples show that our deobfuscation method takes less than 0.5 s on average and increases the similarity between the obfuscated and original scripts from 0.5% to 93.2%. By deploying our deobfuscation method, the attack detection rates for Windows Defender and VirusTotal increase substantially from 0.33% and 2.65% to 78.9% and 94.0%, respectively. Moreover, our detection system outperforms both existing tools with a 96.7% true positive rate and a 0% false positive rate on average. Chun-lin Xiong, Zhenyuan Li, Yan Chen 0004, Tiantian Zhu 0001, Jian Wang 0007 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | RATScope: Recording and Reconstructing Missing RAT Semantic Behaviors for Forensic Analysis on WindowsabstractRemote Access Trojan (RAT) attacks have become an extensively prevailing and serious threat to enterprise security. A forensic system targeting RAT attacks is needed to record and reconstruct fine-grained semantic behaviors of RATs. However, existing forensic systems suffer from various issues such as intrusive instrumentation, nontrivial recording overhead, and RAT behavior blindness. In this article, we first conduct a large-scale study of a representative set of real-world RAT families active from 1999 to 2016. This is the first study to understand the landscape of RATs in the literature. Based on the study, we then proposeRATScope, an instrumentation-free RAT forensic system targeting Windows platform. Specifically,RATScopeoffers an audit logging module to efficiently record system logs by leveraging Event Tracing for Windows (ETW), and provides a novel program behavior modeling technique to reconstruct semantic behaviors of RATs accurately. We implement a prototype ofRATScopeand evaluate the recording overhead and the behavior identification accuracy. The results show that the audit logging module only incurs 3.7 percent runtime overhead on average. Our system can achieve around 90 percent true positive rate in the cross-family experiment, around 80 percent true positive rate in the two-year spanning temporal experiment, and nearzerofalse positive rate. Runqing Yang, Xutong Chen, Haitao Xu 0002, Yueqiang Cheng, Chun-lin Xiong, Linqi Ruan, Mohammad Kavousi, Zhenyuan Li, Liheng Xu, Yan Chen 0004 |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2021 | A Data Processing Method for Load Data of Electric Boiler with Heat Reservoir
Zhenyuan Li, Baoju Li, Tao Peng 0003 |
ICIC (2) | 2 |
| 2021 | Threat detection and investigation with system-level provenance graphs: A survey
Zhenyuan Li, Qi Alfred Chen, Runqing Yang, Yan Chen 0004 |
Comput. Secur. | 1 |
| 2019 | Effective and Light-Weight Deobfuscation and Semantic-Aware Attack Detection for PowerShell ScriptsabstractIn recent years, PowerShell is increasingly reported to appear in a variety of cyber attacks ranging from advanced persistent threat, ransomware, phishing emails, cryptojacking, financial threats, to fileless attacks. However, since the PowerShell language is dynamic by design and can construct script pieces at different levels, state-of-the-art static analysis based PowerShell attack detection approaches are inherently vulnerable to obfuscations. To overcome this challenge, in this paper we design the first effective and light-weight deobfuscation approach for PowerShell scripts. To address the challenge in precisely identifying the recoverable script pieces, we design a novel subtree-based deobfuscation method that performs obfuscation detection and emulation-based recovery at the level of subtrees in the abstract syntax tree of PowerShell scripts. Building upon the new deobfuscation method, we are able to further design the first semantic-aware PowerShell attack detection system. To enable semantic-based detection, we leverage the classic objective-oriented association mining algorithm and newly identify 31 semantic signatures for PowerShell attacks. We perform an evaluation on a collection of 2342 benign samples and 4141 malicious samples, and find that our deobfuscation method takes less than 0.5 seconds on average and meanwhile increases the similarity between the obfuscated and original scripts from only 0.5% to around 80%, which is thus both effective and light-weight. In addition, with our deobfuscation applied, the attack detection rates for Windows Defender and VirusTotal increase substantially from 0.3% and 2.65% to 75.0% and 90.0%, respectively. Furthermore, when our deobfuscation is applied, our semantic-aware attack detection system outperforms both Windows Defender and VirusTotal with a 92.3% true positive rate and a 0% false positive rate on average. Zhenyuan Li, Qi Alfred Chen, Chun-lin Xiong, Yan Chen 0004, Tiantian Zhu 0001 |
CCS | 1 |