Yuanjun Gong

dblp:325/0896 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-4661-904XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments
abstract
[Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language models (LLMs) seem ideal candidates to replace human annotators. [Problem:] Coding security-specific aspects (code identifiers mentioned, lines-of-code-mentioned, security keywords mentioned) may require deeper contextual understanding than sentiment classification. [Objective:] explore whether LLMs can act as automated annotators for technical security comments by human subjects. [Method:] We prompt four best LLMs on LiveBench to detect nine security-relevant codes in free-text comments by human subjects analyzing vulnerable code snippets. Outputs are compared to the human annotators along with Cohen’s Kappa (chance-corrected accuracy). We test different prompts mimicking annotation best practices: emerging codes, a detailed codebook with examples, and conflicting examples. [Negative Results:] We observed marked improvements only with the code descriptions, but they are not uniform across codes and not sufficient to reliably replace a human annotator. [Limitations:] Additional studies with more LLMs and annotation tasks are needed.
Maria Camporese, Fabio Massacci, Yuanjun Gong
ICPC3
2026 Zero-shot temporal knowledge graph completion based on generative adversarial network
Lin Zhu 0014, Yuanjun Gong, Luyi Bai
World Wide Web (WWW)2
2024 Raisin: Identifying Rare Sensitive Functions for Bug Detection
abstract
Mastering the knowledge about the bug-prone functions (i.e., sensitive functions) is important to detect bugs. Some automated techniques have been proposed to identify the sensitive functions in large software systems, based on machine learning or natural language processing. However, the existing statistics-based techniques are not directly applicable to a special kind of sensitive functions, i.e., the rare sensitive functions, which have very few invocations even in large systems. Unfortunately, the rare ones can also introduce bugs. Therefore, how to effectively identify such functions is a problem deserving attention.
Jianjun Huang 0001, Jianglei Nie, Yuanjun Gong, Wei You 0001, Bin Liang 0002, Pan Bian
ICSE3
2024 SICode: Embedding-Based Subgraph Isomorphism Identification for Bug Detection
abstract
Given a known buggy code snippet, searching for similar patterns in a target project to detect unknown bugs is a reasonable approach. In practice, a search unit, such as a function, may appear quite different from the buggy snippet but actually contains a similar buggy substructure. Utilizing subgraph isomorphism identification can effectively hunt potential bugs by checking whether an approximate copy of the buggy subgraph exists within the target code graphs. Regrettably, subgraph isomorphism identification is an NP-complete problem.
Yuanjun Gong, Jianglei Nie, Wei You 0001, Wenchang Shi, Jianjun Huang 0001, Bin Liang 0002, Jian Zhang 0001
ICPC1
2022 Hunting bugs with accelerated optimal graph vertex matching
abstract
Various techniques based on code similarity measurement have been proposed to detect bugs. Essentially, the code fragment can be regarded as a kind of graph. Performing code graph similarity comparison to identify the potential bugs is a natural choice. However, the logic of a bug often involves only a few statements in the code fragment, while others are bug-irrelevant. They can be considered as a kind of noise, and can heavily interfere with the code similarity measurement. In theory, performing optimal vertex matching can address the problem well, but the task is NP-complete and cannot be applied to a large-scale code base. In this paper, we propose a two-phase strategy to accelerate code graph vertex matching for detecting bugs. In the first phase, a vertex matching embedding model is trained and used to rapidly filter a limited number of candidate code graphs from the target code base, which are likely to have a high vertex matching degree with the seed, i.e., the known buggy code. As a result, the number of code graphs needed to be further analyzed is dramatically reduced. In the second phase, a high-order similarity embedding model based on graph convolutional neural network is built to efficiently get the approximately optimal vertex matching between the seed and candidates. On this basis, the code graph similarity is calculated to identify the potential buggy code. The proposed method is applied to five open source projects. In total, 31 unknown bugs were successfully detected and confirmed by developers. Comparative experiments demonstrate that our method can effectively mitigate the noise problem, and the detection efficiency can be improved dozens of times with the two-phase strategy.
Yuanjun Gong, Bin Liang 0002, Jianjun Huang 0001, Wei You 0001, Wenchang Shi, Jian Zhang 0001
ISSTA2