Yanjing Yang

dblp:252/8132 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 One Size Does Not Fit All: Investigating Efficacy of Perplexity in Detecting LLM-Generated Code
abstract
Large Language Model-Generated Code (LLMgCode) has become increasingly common in software development. So far LLMgCode has more quality issues than Human-Authored Code (HaCode). It is common for LLMgCode to mix with HaCode in a code change, while the change is signed by only human developers, without being carefully examined. Many automated methods have been proposed to detect LLMgCode from HaCode, in which the perplexity-based method ( Perplexity for short) is the state-of-the-art method. However, the efficacy evaluation of Perplexity has focused on detection accuracy. Yet it is unclear whether Perplexity is good enough in a wider range of realistic evaluation settings. To this end, we carry out a family of experiments to compare Perplexity against feature- and pre-training-based methods from three perspectives: detection accuracy , detection speed , and generalization capability . The experimental results show that Perplexity has the best generalization capability while having limited detection accuracy and detection speed. Based on that, we discuss the strengths and limitations of Perplexity , e.g., Perplexity is unsuitable for high-level programming languages. Finally, we provide recommendations to improve Perplexity and apply it in practice. As the first large-scale investigation on detecting LLMgCode from HaCode, this article provides a wide range of findings for future improvement.
Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Zeru Cheng, Bohan Liu 0003, Xin Zhou 0016, Alberto Bacchelli, Yin Kia Chiam, Thiam Kian Chiew
ACM Trans. Softw. Eng. Methodol.3
2026 Automated Localization of Affected Libraries and Versions from Vulnerability Reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Jinghao Hu 0001, Lanxin Yang, Bohan Liu 0003
IEEE Trans. Software Eng.4
2025 Securing Self-Managed Third-Party Libraries
abstract
Modern software development reuses third-party libraries to cut costs but may introduce vulnerabilities. A critical practice is to verify the security of third-party libraries against public vulnerability reports. Many automated methods have been proposed to identify vulnerable libraries from vulnerability reports. Existing methods are designed for the generic identification of vulnerable libraries, considering the security of all software libraries. Generic identification is inherently challenging, resulting in limited accuracy. However, organizations only consider the security of libraries they trust and use, by self-managing a library whitelist. Therefore, we propose LibGuard, a framework to adapt existing methods to help organizations secure the libraries they use. LibGuard supplies a library whitelist for existing methods and filters the results according to a threshold, facilitating the discovery of risks overlooked by organizations while controlling false alarms. LibGuard is implemented in two ways. The first attaches the whitelist after existing methods. The second integrates the whitelist into existing methods. We evaluated LibGuard using 5,107 vulnerability reports and the library whitelist built from 79 Google projects and 29 Huawei projects. The results show that the two implementations of LibGuard increase the average F1 score by 10.25% and 11.77%, respectively. Moreover, LibGuard performs stably during the extension of whitelists. To our knowledge, this paper is the first study dedicated to securing self-managed third-party libraries, offering insights into adapting generic software security management to self-managed contexts.
Xin Zhou 0016, Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Bohan Liu 0003, Hongshan Tang
ASE4
2025 Automated detection of affected libraries from vulnerability reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Runfeng Mao, Lanxin Yang, Haifeng Shen
Autom. Softw. Eng.4
2025 DLAP: A Deep Learning Augmented Large Language Model Prompting framework for software vulnerability detection
Yanjing Yang, Xin Zhou 0016, Runfeng Mao, Jinwei Xu, Lanxin Yang, Haifeng Shen, He Zhang 0001
J. Syst. Softw.1
2021 MSPLD: Shilling Attack Detection Model Based on Meta Self-Paced Learning
abstract
With the dramatic rise of recommendation systems, more and more companies use them to improve users' experience. However, the openness of the recommender systems makes them vulnerable to shilling attacks, which causes a bad impact on user satisfaction. Existing shilling attack detection models usually have problems in solving the noise of samples and labels. To this end, this paper proposes a shilling attack detection model based on meta self-paced learning, named MSPLD. Meta self-paced learning can make the model select samples from easy to difficult in the learning process, which can alleviate the problem that the model parameters are difficult to optimize due to the outliers or noises in the samples. Specifically, MSPLD adopts some methods to get the extraction of potential feature embedding vectors first. Second, metadata is selected adaptively by a regression method. Then, it uses the embedding vectors of malicious users and metadata as input data. Third, it optimizes the age loss function and the loss function of the classifier itself with bilateral optimization. The relationship between age and sample loss of the classifier will determine the weight of sample selection. Finally, using the tendency of the age gradually to select training samples from easy to difficult can improve the generalization ability of the models. Compared with the state-of-the-art detection models, the experimental results on two public datasets show that MSPLD can achieve better detection performance. Besides, we illustrate the training process of MSPLD to analyze the reason for the superiority of the model.
Yanjing Yang, Min Gao 0001, Yuerang Li, Jia Wang 0055, Quanwu Zhao
IJCNN1