Jiayuan Li 0002

dblp:138/8345-2 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Modubin: A Binary Modularization Approach Based on the Locality of Homologous Functions
Wenyan Yu, Lei Cui 0003, Jiayuan Li 0002, Hong Li 0004, Hongsong Zhu
ICPC3
2026 PVDetector: Pretrained Vulnerability Detection on Vulnerability-enriched Code Semantic Graph
abstract
Automated vulnerability detection is a critical issue in software security. The advent of Deep Learning (DL) has led to numerous studies employing DL to detect vulnerabilities in software source code. However, existing approaches still perform poorly, particularly with real-world vulnerabilities, due to the difficulty in accurately capturing their properties. To this end, we introduce PVDetector, a DL-based approach that utilizes rich code semantics, incorporates vulnerability knowledge, and leverages pretrained code representations for precise vulnerability detection. At its core, PVDetector employs a new model called Vulnerability-enriched Code Semantic Graph (VCSG), which accurately characterizes functions by distinguishing the semantics of identical variables and more finely capturing control dependencies, data dependencies, and vulnerability relationships. Additionally, we introduce four pretraining tasks specifically designed to learn the semantics of control, data, vulnerability, and variables from the VCSG model. These pretraining tasks significantly enhance PVDetector’s capability to detect vulnerabilities in downstream tasks. Experimental results indicate that PVDetector outperforms SOTAs by 5.0–12.5% in precision, 0.2–9.7% in recall, and 3.0–15.1% in F1-score. Additionally, it supports six programming languages and demonstrates high efficiency (e.g., 10.6 \(\times\) faster than DeepDFA). When applied to seven software products, PVDetector discovered 55 vulnerabilities, including 10 silently patched flaws that had not been previously reported.
Jiayuan Li 0002, Lei Cui 0003, Jie Zhang 0121, Rongrong Xi, Hongsong Zhu
ACM Trans. Softw. Eng. Methodol.1
2025 Steering Large Language Models for Vulnerability Detection
abstract
Vulnerability detection remains a critical challenge in the field of security. Many existing approaches extract code representations for vulnerability detection. However, these methods often focus on the overall semantics of the code, neglecting to specifically target vulnerability-related semantics. To address this limitation, we propose a novel LLM steering method designed to steer LLMs to focus on vulnerability concepts, thereby enhancing their performance in vulnerability detection. Specifically, we introduce a vulnerability steering vector that represents the concept of vulnerability in the representation space. This vector is generated using a paired vulnerability-patch function dataset, effectively capturing the essence of vulnerabilities. Experimental results demonstrate that the proposed method significantly improves LLMs' performance and notably outperforms existing SOTA methods in vulnerability detection tasks. Furthermore, we validate the cross-language transferability of the steering vector and explore the explainability of vulnerability detection.
Jiayuan Li 0002, Lei Cui 0003, Jie Zhang 0121, Haiqiang Fei, Hongsong Zhu
ICASSP1
2025 VulnTeam: A Team Collaboration Framework for LLM-based Vulnerability Detection
abstract
Software vulnerability detection is a critical challenge in cyber security. With the rise of deep learning and large language models (LLMs), numerous studies have applied these technologies to vulnerability detection. Existing approaches directly employ prompt engineering, chain-of-thought reasoning, and fine-tuning methods on LLMs, but achieve suboptimal results. To effectively leverage LLMs’ powerful reasoning capabilities for vulnerability detection, we propose VulnTeam, a novel team collaboration framework for LLM vulnerability detection inspired by human expert team collaboration. Specifically, we introduce a dual-stage fine-tuning approach where expert models are first fine-tuned using low-rank adaptation to detect vulnerabilities related to different vulnerability syntactic features, followed by instruction fine-tuning of a leader model responsible for the final decision-making. Ultimately, team members (expert models) and the team leader (leader model) collaborate to detect vulnerabilities. Our experimental evaluation across three LLMs and two datasets demonstrates that VulnTeam significantly enhances LLMs’ vulnerability detection performance (average F1-score improvement of 12.51%). Moreover, VulnTeam-enhanced LLMs substantially outperform previous state-of-the-art (SOTA) vulnerability detection methods (average F1-score improvement of 7.78%). Additionally, we analyze computational costs to validate VulnTeam’s practical applicability.
Jiayuan Li 0002, Lei Cui 0003, Wenyan Yu, Haiqiang Fei, Hongsong Zhu
IJCNN1
2024 EasyDetector: Using Linear Probe to Detect the Provenance of Large Language Models
abstract
The rapid development of large language models (LLMs) has driven significant advancements in various applications. However, the intellectual property of these models often faces risks due to unauthorized reproduction or encapsulation by third parties. In this paper, we propose EasyDetector, a novel approach to detect the provenance of LLMs using linear probes. Our method aims to identify the original source model, even if it has been fine-tuned or encapsulated into another model. Specifically, EasyDetector performs classification on the intermediate layer representations of the new model using linear probes of the original model. Models from the same source exhibit high accuracy, while models from different sources yield low accuracy. Extensive experiments on diverse LLMs demonstrate the effectiveness of EasyDetector in detecting model provenance. The proposed method is lightweight and applicable to various model architectures, holding significant importance for protecting the intellectual property of LLMs.
Jie Zhang 0121, Jiayuan Li 0002, Haiqiang Fei, Hongsong Zhu
TrustCom2