VLDB 2026 Research / reviewers in the wild / expert
Zhen Li 0027
dblp:74/2397-27
· DBLP profile ↗
23ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0002-0001-2998ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 11 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoK: Automated Vulnerability Repair: Methods, Tools, and Assessments
Zhen Li 0027, Kedie Shu, Shenghua Guan, Deqing Zou, Shouhuai Xu, Bin Yuan 0002, Hai Jin 0001 |
USENIX Security Symposium | 2 |
| 2025 | MalPacDetector: An LLM-Based Malicious NPM Package DetectorabstractThe Node Package Manager (NPM) registry contains millions of JavaScript packages widely shared between worldwide developers. However, NPM has also been abused by attackers to spread malicious packages, highlighting the importance of detecting malicious NPM packages. Existing malicious NPM package detectors suffer from, among other things, high false positives and/or high false negatives. In this paper, we propose a novel Malicious NPM Package Detector (MalPacDetector), which leverages Large Language Model (LLM) to automatically and dynamically generate features (rather than asking experts to manually define them). To evaluate the effectiveness of Mal-PacDetector and existing detectors, we construct a new NPM package dataset, which overcomes the weaknesses of existing datasets (e.g., a small number of examples and a high repetition rate of malicious fragments). The experimental results show that MalPacDetector outperforms existing detectors by achieving a false positive rate of 1. 3% and a false negative rate of 7. 5%. In particular, MalPacDetector detects 39 previously unknown malicious packages, which are confirmed by the NPM security team. Zhen Li 0027, Jixiang Qu, Deqing Zou, Shouhuai Xu, Ziteng Xu, Hai Jin 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | On the Effectiveness of Function-Level Vulnerability Detectors for Inter-Procedural VulnerabilitiesabstractSoftware vulnerabilities are a major cyber threat and it is important to detect them. One important approach to detecting vulnerabilities is to use deep learning while treating a program function as a whole, known as function-level vulnerability detectors. However, the limitation of this approach is not understood. In this paper, we investigate its limitation in detecting one class of vulnerabilities known as inter-procedural vulnerabilities, where the to-be-patched statements and the vulnerability-triggering statements belong to different functions. For this purpose, we create the first Inter-Procedural Vulnerability Dataset (InterPVD) based on C/C++ open-source software, and we propose a tool dubbed VulTrigger for identifying vulnerability-triggering statements across functions. Experimental results show that VulTrigger can effectively identify vulnerability-triggering statements and inter-procedural vulnerabilities. Our findings include: (i) inter-procedural vulnerabilities are prevalent with an average of 2.8 inter-procedural layers; and (ii) function-level vulnerability detectors are much less effective in detecting to-be-patched functions of inter-procedural vulnerabilities than detecting their counterparts of intra-procedural vulnerabilities. Zhen Li 0027, Ning Wang 0098, Deqing Zou, Ruqian Zhang, Shouhuai Xu, Chao Zhang 0008, Hai Jin 0001 |
ICSE | 1 |
| 2024 | Leakage of Authorization-Data in IoT Device Sharing: New Attacks and CountermeasureabstractDevice sharing among users is a common functionality in today's IoT clouds. Supporting device sharing are the delegation methods proposed by different IoT clouds, which we find are heterogeneous and ad-hoc IoT clouds use various data (e.g., device ID, product ID, and access token) as authorization certificates. In this paper, we report the first systematic study on how the authorization-data are managed in IoT device sharing. Our study brought to light the security risks in today's IoT authorization-data management, identifying 6 authorization-data leakage flaws. To mitigate such flaws, we propose an approach to hide the authorization-data from the delegatee (a.k.a., the user authorized to access the devices) without disrupting the device sharing services. We propose SecHARE, an automated tool to patch the vulnerable IoT clouds. We applied SecHARE to 3 popular open-source IoT clouds. Results have shown the compatibility, effectiveness, and efficiency of SecHARE. We have made SecHARE publicly available Bin Yuan 0002, Maogen Yang, Qunjinming Chen, Zhanxiang Song, Zhen Li 0027, Deqing Zou, Hai Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Toward Automated Attack Discovery in SDN Controllers Through Formal VerificationabstractSoftware-defined Network (SDN), presented to be a novel architecture of network because of its separation of data plane and control plane, brings centralization and extensibility to network management as well as new attacks that exploit the flexibility of SDN. OpenFlow, which is the protocol that is applied by the majority of SDN, leads to the widely used definition of the communication between the controller and the switch resulting in similar implementations regardless of different vendors. In this paper, we focus on the mechanisms of packet processing and topology discovery and their fundamental weaknesses caused by general implementations or device limitations. Despite the common vulnerabilities, the universal standard mechanisms of basic function in SDN also enlighten us to present an automated attack discovery method based on the formal verification with a generic model of SDN system. We describe the abstraction of the SDN components, their key functions, and communications along with the malicious operations that could be executed by malicious hosts and malicious switches and translate them into a formal model of the SDN system. The formal verification carried on with the assertion representing the security properties derived from the common vulnerabilities of the SDN system reports the potential attack paths each of which shows an attack process. Our evaluation shows that our method can discover feasible attack paths efficiently and effectively, with 23 attacks being identified, among which 2 are new. We further demonstrate the practicality of the 2 new attacks. Bin Yuan 0002, Chi Zhang 0117, Jiajun Ren, Qunjinming Chen, Biang Xu, Qiankun Zhang 0001, Zhen Li 0027, Deqing Zou, Fan Zhang 0024, Hai Jin 0001 |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2024 | Reducing the Impact of Time Evolution on Source Code Authorship Attribution via Domain AdaptationabstractSource code authorship attribution is an important problem in practical applications such as plagiarism detection, software forensics, and copyright disputes. Recent studies show that existing methods for source code authorship attribution can be significantly affected by time evolution, leading to a decrease in attribution accuracy year by year. To alleviate the problem of Deep Learning (DL)-based source code authorship attribution degrading in accuracy due to time evolution, we propose a new framework called Time D omain A daptation (TimeDA) by adding new feature extractors to the original DL-based code attribution framework that enhances the learning ability of the original model on source domain features without requiring new or more source data. Moreover, we employ a centroid-based pseudo-labeling strategy using neighborhood clustering entropy for adaptive learning to improve the robustness of DL-based code authorship attribution. Experimental results show that TimeDA can significantly enhance the robustness of DL-based source code authorship attribution to time evolution, with an average improvement of 8.7% on the Java dataset and 5.2% on the C++ dataset. In addition, our TimeDA benefits from employing the centroid-based pseudo-labeling strategy, which significantly reduced the model training time by 87.3% compared to traditional unsupervised domain adaptive methods. Zhen Li 0027, Chen Chen 0001, Qian Chen 0019 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | Enhancing Deep Learning-based Vulnerability Detection by Building Behavior Graph ModelabstractSoftware vulnerabilities have posed huge threats to the cyberspace security, and there is an increasing demand for automated vulnerability detection (VD). In recent years, deep learning-based (DL-based) vulnerability detection systems have been proposed for the purpose of automatic feature extraction from source code. Although these methods can achieve ideal performance on synthetic datasets, the accuracy drops a lot when detecting real-world vulnerability datasets. Moreover, these approaches limit their scopes within a single function, being not able to leverage the information between functions. In this paper, we attempt to extract the function's abstract behaviors, figure out the relationships between functions, and use this global information to assist DL-based VD to achieve higher performance. To this end, we build a Behavior Graph Model and use it to design a novel framework, namely VulBG. To examine the ability of our constructed Behavior Graph Model, we choose several existing DL-based VD models (e.g., TextCNN, ASTGRU, CodeBERT, Devign, and VulCNN) as our baseline models and conduct evaluations on two real-world datasets: the balanced$\text{FFMpeg}+\text{Qemu}$dataset and the unbalanced$\text{Chrome} +\text{Debian}$dataset. Experimental results indicate that VulBG enables all baseline models to detect more real vulnerabilities, thus improving the overall detection performance. Bin Yuan 0002, Yilin Fang, Yueming Wu 0001, Deqing Zou, Zhen Li 0027, Zhi Li 0048, Hai Jin 0001 |
ICSE | 6 |
| 2023 | Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code ClassifiersabstractDeep learning has been widely used in source code classification tasks, such as code classification according to their functionalities, code authorship attribution, and vulnerability detection. Unfortunately, the black-box nature of deep learning makes it hard to interpret and understand why a classifier (i.e., classification model) makes a particular prediction on a given example. This lack of interpretability (or explainability) might have hindered their adoption by practitioners because it is not clear when they should or should not trust a classifier's prediction. The lack of interpretability has motivated a number of studies in recent years. However, existing methods are neither robust nor able to cope with out-of-distribution examples. In this paper, we propose a novel method to produce Robust interpreters for a given deep learning-based code classifier; the method is dubbed Robin. The key idea behind Robin is a novel hybrid structure combining an interpreter and two approximators, while leveraging the ideas of adversarial training and data augmentation. Experimental results show that on average the interpreter produced by Robin achieves a 6.11% higher fidelity (evaluated on the classifier), 67.22% higher fidelity (evaluated on the approximator), and 15.87x higher robustness than that of the three existing interpreters we evaluated. Moreover, the interpreter is 47.31% less affected by out-of-distribution examples than that of LEMNA. Zhen Li 0027, Ruqian Zhang, Deqing Zou, Ning Wang 0098, Shouhuai Xu, Chen Chen 0001, Hai Jin 0001 |
ASE | 1 |
| 2023 | Network intrusion detection based on the temporal convolutional model
Ivandro Ortet Lopes, Deqing Zou, Ihsan H. Abdulqadder, Saeed Akbar, Zhen Li 0027, Francis A. Ruambo, Wagner Pereira |
Comput. Secur. | 5 |
| 2023 | A comparative study of adversarial training methods for neural models of source code
Zhen Li 0027, Yangrui Li, Qian Chen 0019 |
Future Gener. Comput. Syst. | 1 |
| 2022 | Generating Adversarial Source Programs Using Important Tokens-based Structural TransformationsabstractDeep learning models have been widely used in source code processing tasks, such as code captioning, code summarization, code completion, and code classification. Recent studies have shown that deep learning-based source code processing models are vulnerable. Attackers can generate adversarial examples by adding perturbations to source programs. Existing attack methods perturb a source program by renaming one or multiple variables in the program. These attack methods do not take into account the perturbation of the equivalent structural transformations of the source code. We propose a set of program transformations involving identifier renaming and structural transformations, which can ensure that the perturbed program retains the original semantics but can fool the source code processing model to change the original prediction result. We propose a novel method of applying semantics-preserving structural transformations to attack the source program pro-cessing model in the white-box setting. This is the first time that semantics-preserving structural transformations are applied to generate adversarial examples of source code processing models. We first find the important tokens in the program by calculating the contribution values of each part of the program, then select the best transformation for each important token to generate semantic adversarial examples. The experimental results show that the attack success rate of our attack method can improve 8.29 % on average compared with the state-of-the-art attack method; adversarial training using the adversarial examples generated by our attack method can reduce the attack success rates of source code processing models by 21.79% on average. Penglong Chen, Zhen Li 0027 |
ICECCS | 2 |
| 2022 | RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationabstractSource code authorship attribution is an important problem often encountered in applications such as software forensics, bug fixing, and software quality analysis. Recent studies show that current source code authorship attribution methods can be compromised by attackers exploiting adversarial examples and coding style manipulation. This calls for robust solutions to the problem of code authorship attribution. In this paper, we initiate the study on making Deep Learning (DL)-based code authorship attribution robust. We propose an innovative framework called Robust coding style Patterns Generation (RoPGen), which essentially learns authors' unique coding style patterns that are hard for attackers to manipulate or imitate. The key idea is to combine data augmentation and gradient augmentation at the adversarial training phase. This effectively increases the diversity of training examples, generates meaningful perturbations to gradients of deep neural networks, and learns diversified representations of coding styles. We evaluate the effectiveness of RoPGen using four datasets of programs written in C, C++, and Java. Experimental results show that RoPGen can significantly improve the robustness of DL-based code authorship attribution, by respectively reducing 22.8% and 41.0% of the success rate of targeted and untargeted attacks on average. Zhen Li 0027, Qian Chen 0019, Chen Chen 0001, Yayi Zou, Shouhuai Xu |
ICSE | 1 |
| 2022 | Towards Improving Multiple Authorship Attribution of Source CodeabstractSource code authorship attribution addresses the problems of copyright infringement disputes and plagiarism detection. However, most software projects are collaborative development projects. It is necessary to study multiple authorship attribution. Existing methods are not reliable in the domain of multiple authorship attribution. The reasons are as follows: i) It is a challenge to divide the code boundaries of different authors in a sample; ii) code segments belonging to different authors in a sample are usually small or incomplete. This paper proposes a method to address these challenges. We first divide the code sample into multiple lines, then integrate the code lines with similar author styles into code segments using Siamese networks. Finally, we use a path-based code representation and machine learning to identify authors. Experimental results show the method achieves an accuracy of 87.35% on C/C++ dataset and 91.35% on Java dataset, which performs better than existing methods. PengNan Hao, Zhen Li 0027, Cui Liu, Fanming Liu |
QRS | 2 |
| 2022 | SySeVR: A Framework for Using Deep Learning to Detect Software VulnerabilitiesabstractThe detection of software vulnerabilities (or vulnerabilities for short) is an important problem that has yet to be tackled, as manifested by the many vulnerabilities reported on a daily basis. This calls for machine learning methods for vulnerability detection. Deep learning is attractive for this purpose because it alleviates the requirement to manually define features. Despite the tremendous success of deep learning in other application domains, its applicability to vulnerability detection is not systematically understood. In order to fill this void, we propose thefirstsystematic framework for using deep learning to detect vulnerabilities in C/C++ programs with source code. The framework, dubbedSyntax-based,Semantics-based, andVectorRepresentations(SySeVR), focuses on obtaining program representations that can accommodate syntax and semantic information pertinent to vulnerabilities. Our experiments with four software products demonstrate the usefulness of the framework: we detect 15 vulnerabilities that are not reported in the National Vulnerability Database. Among these 15 vulnerabilities, seven are unknown and have been reported to the vendors, and the other eight have been “silently” patched by the vendors when releasing newer versions of the pertinent software products. Zhen Li 0027, Deqing Zou, Shouhuai Xu, Hai Jin 0001, Yawei Zhu, Zhaoxuan Chen |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | VulDeeLocator: A Deep Learning-Based Fine-Grained Vulnerability DetectorabstractAutomatically detecting software vulnerabilities is an important problem that has attracted much attention from the academic research community. However, existing vulnerability detectors still cannot achieve the vulnerability detection capability and the locating precision that would warrant their adoption for real-world use. In this article, we present a vulnerability detector that can simultaneously achieve a high detection capability and a high locating precision, dubbedVulnerabilityDeep learning-basedLocator(VulDeeLocator). In the course of designing VulDeeLocator, we encounter difficulties including how to accommodate semantic relations between the definitions of types as well as macros and their uses across files, how to accommodate accurate control flows and variable define-use relations, and how to achieve high locating precision. We solve these difficulties by using two innovative ideas: (i) leveraging intermediate code to accommodate extra semantic information, and (ii) using the notion ofgranularity refinementto pin down locations of vulnerabilities. When applied to 200 files randomly selected from three real-world software products, VulDeeLocator detects 18 confirmed vulnerabilities (i.e., true-positives). Among them, 16 vulnerabilities correspond to known vulnerabilities; the other two are not reported in the National Vulnerability Database (NVD) but have been “silently” patched by the vendor of Libav when releasing newer versions. Zhen Li 0027, Deqing Zou, Shouhuai Xu, Zhaoxuan Chen, Yawei Zhu, Hai Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Generating Adversarial Examples of Source Code Classification Models via Q-Learning-Based Markov Decision ProcessabstractAdversarial robustness becomes an essential concern in Deep Learning (DL)-based source code processing, as DL models are vulnerable to the deception by attackers. To address a new challenge posed by the discrete and structural nature of source code to generate adversarial examples for DL models, and the insufficient focus of existing methods on code structural features, we propose a Q-Learning-based Markov decision process (QMDP) performing semantically equivalent transformations on the source code structure. Two key issues are mainly addressed: (i) how to perform attacks on source code structural information and (ii) what transformations to perform when and where in the source code. We demonstrate that effectively tackling these two issues is crucial for generating adversarial examples for source code. By evaluating C/C++ programs working on the source code classification task, we verified that QMDP can effectively generate adversarial examples and improve the robustness of DL models over 44%. Chenxin Wang, Zhen Li 0027 |
QRS | 3 |
| 2021 | $\mu$μVulDeePecker: A Deep Learning-Based System for Multiclass Vulnerability DetectionabstractFine-grained software vulnerability detection is an important and challenging problem. Ideally, a detection system (or detector) not only should be able to detect whether or not a program contains vulnerabilities, but also should be able to pinpoint the type of a vulnerability in question. Existing vulnerability detection methods based on deep learning can detect the presence of vulnerabilities (i.e., addressing the binary classification or detection problem), but cannot pinpoint types of vulnerabilities (i.e., incapable of addressing multiclass classification). In this paper, we propose the first deep learning-based system for multiclass vulnerability detection, dubbed μ VulDeePecker. The key insight underlying μ VulDeePecker is the concept of code attention, which can capture information that can help pinpoint types of vulnerabilities, even when the samples are small. For this purpose, we create a dataset from scratch and use it to evaluate the effectiveness of μ VulDeePecker. Experimental results show that μ VulDeePecker is effective for multiclass vulnerability detection and that accommodating control-dependence (other than data-dependence) can lead to higher detection capabilities. Deqing Zou, Sujuan Wang, Shouhuai Xu, Zhen Li 0027, Hai Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | Interpreting Deep Learning-based Vulnerability Detector Predictions Based on Heuristic SearchingabstractDetecting software vulnerabilities is an important problem and a recent development in tackling the problem is the use of deep learning models to detect software vulnerabilities. While effective, it is hard to explain why a deep learning model predicts a piece of code as vulnerable or not because of the black-box nature of deep learning models. Indeed, the interpretability of deep learning models is a daunting open problem. In this article, we make a significant step toward tackling the interpretability of deep learning model in vulnerability detection. Specifically, we introduce a high-fidelity explanation framework, which aims to identify a small number of tokens that make significant contributions to a detector’s prediction with respect to an example. Systematic experiments show that the framework indeed has a higher fidelity than existing methods, especially when features are not independent of each other (which often occurs in the real world). In particular, the framework can produce some vulnerability rules that can be understood by domain experts for accepting a detector’s outputs (i.e., true positives) or rejecting a detector’s outputs (i.e., false-positives and false-negatives). We also discuss limitations of the present study, which indicate interesting open problems for future research. Deqing Zou, Yawei Zhu, Shouhuai Xu, Zhen Li 0027, Hai Jin 0001, Hengkai Ye |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2019 | AutoCVSS: An Approach for Automatic Assessment of Vulnerability Severity Based on Attack Process
Deqing Zou, Ju Yang, Zhen Li 0027, Hai Jin 0001, Xiaojing Ma 0002 |
GPC | 3 |
| 2018 | Automatically Identifying Security Bug Reports via Multitype Features Analysis
Deqing Zou, Zhijun Deng, Zhen Li 0027, Hai Jin 0001 |
ACISP | 3 |
| 2018 | VulDeePecker: A Deep Learning-Based System for Vulnerability Detection
Zhen Li 0027, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin 0001, Sujuan Wang, Zhijun Deng, Yuyi Zhong |
NDSS | 1 |
| 2017 | SCVD: A New Semantics-Based Approach for Cloned Vulnerable Code Detection
Deqing Zou, Hanchao Qi, Zhen Li 0027, Song Wu 0001, Hai Jin 0001, Guozhong Sun, Sujuan Wang, Yuyi Zhong |
DIMVA | 3 |
| 2016 | VulPecker: an automated vulnerability detection system based on code similarity analysis
Zhen Li 0027, Deqing Zou, Shouhuai Xu, Hai Jin 0001, Hanchao Qi |
ACSAC | 1 |