Xiangyue Liu 0002

dblp:151/9210-2 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 PIONEER: improving the robustness of student models when compressing pre-trained models of code
Xiangyue Liu 0002, Lili Bo, Xiaoxue Wu 0001, Yun Yang 0003, Xiaobing Sun 0001
Autom. Softw. Eng.1
2026 AdvGen-X: Transferability driven adversarial example generation for pre-trained models of code
Xiangyue Liu 0002, Xiaobing Sun 0001, Lili Bo, Bin Li 0006, Xiaoxue Wu 0001, Sicong Cao, Yufei Hu
Empir. Softw. Eng.1
2025 Evaluating the Test Adequacy of Benchmarks for LLMs on Code Generation
abstract
ABSTRACT Code generation for users' intent has become increasingly prevalent with the large language models (LLMs). To automatically evaluate the effectiveness of these models, multiple execution‐based benchmarks are proposed, including specially crafted tasks, accompanied by some test cases and a ground truth solution. LLMs are regarded as well‐performed in code generation tasks if they can pass the test cases corresponding to most tasks in these benchmarks. However, it is unknown whether the test cases have sufficient test adequacy and whether the test adequacy can affect the evaluation. In this paper, we conducted an empirical study to evaluate the test adequacy of the execution‐based benchmarks and to explore their effects during evaluation for LLMs. Based on the evaluation of the widely used benchmarks, HumanEval, MBPP, and two enhanced benchmarks HumanEval+ and MBPP+, we obtained the following results: (1) All the evaluated benchmarks have high statement coverage (above 99.16%), low branch coverage (74.39%) and low mutation score (87.69%). Especially for the tasks with higher cyclomatic complexities in the HumanEval and MBPP, the mutation score of test cases is lower. (2) No significant correlation exists between test adequacy (statement coverage, branch coverage and mutation score) of benchmarks and evaluating results on LLMs at the individual task level. (3) There is a significant positive correlation between mutation score‐based evaluation and another execution‐based evaluation metric () on LLMs at the individual task level. (4) The existing test case augmentation techniques have limited improvement in the coverage of test cases in the benchmark, while significantly improving the mutation score by approximately 34.60% and also can bring a more rigorous evaluation to LLMs on code generation. (5) The LLM‐based test case generation technique (EvalPlus) performs better than the traditional search‐based technique (Pynguin) in improving the benchmarks' test quality and evaluation ability of code generation.
Xiangyue Liu 0002, Xiaobing Sun 0001, Lili Bo, Yufei Hu, Zhenlei Ye
J. Softw. Evol. Process.1
2025 Misactivation-Aware Stealthy Backdoor Attacks on Neural Code Understanding Models
abstract
Neural code models (NCMs) play a crucial role in helping developers solve code understanding tasks. Recent studies have exposed that NCMs are vulnerable to several security threats, among which backdoor attack is one of the toughest. It is usually achieved through data poisoning. Specifically, backdoored NCMs work normally on the clean example but produce attacker-expected output on the example injected with backdoor triggers. However, existing backdoor attacks against NCMs face two significant drawbacks: 1) lack of stealthiness, that is trigger tokens are easily detected by defense techniques/humans when they appear in excessive numbers; 2) damage to the model’s normal performance, that is partial trigger tokens may frequently appear as benign features in the clean samples, resulting in clean samples containing them may falsely activate the backdoor. To address these drawbacks, we propose a misactivation-aware stealthy backdoor attack against NCMs through data poisoning called MISNCM. MISNCM features target-biased trigger generation, thus achieving stealthy backdoor attacks. Moreover, we utilize misactivation-aware data poisoning to create calibration samples with partial trigger tokens to reduce false activations and ensure the regular performance of the model. We conduct comprehensive experiments to evaluate the effectiveness of MISNCM in attacking NCMs used for three code understanding tasks: defect detection, clone detection, and authorship attribution. The experimental results demonstrate that the triggers generated by MISNCM achieve an average attack success rate increase of 12.67% over IR and 8.38% over AFRAIDOOR. Furthermore, MISNCM achieves a 3.64% improvement in F1 score on the code clone detection task, and an average of 5.91% improvement in accuracy on the defect detection and authorship attribution tasks, compared with the two baselines.
Xiaobing Sun 0001, Yiran Xiao, Lili Bo, Weisong Sun, Xiangyue Liu 0002, Bin Li 0006, Jiale Zhang 0001
IEEE Trans. Software Eng.5
2016 Exploring topic models in software engineering data analysis: A survey
abstract
Topic models are shown to be effective to mine unstructured software engineering (SE) data. In this paper, we give a simple survey of exploring topic models to support various SE tasks between 2003 and 2015. The survey results show that there is an increasing concern in this area. Among the SE tasks, source code comprehension and software history comprehension are the mostly studied, followed by software defects prediction. However, there is still only a few studies on other SE tasks, such as feature location and regression testing.
Xiaobing Sun 0001, Xiangyue Liu 0002, Bin Li 0006, Yucong Duan, Jiajun Hu
SNPD2
2016 Code Comment Quality Analysis and Improvement Recommendation: An Automated Approach
abstract
Program comprehension is one of the first and most frequently performed activities during software maintenance and evolution. In a program, there are not only source code, but also comments. Comments in a program is one of the main sources of information for program comprehension. If a program has good comments, it will be easier for developers to understand it. Unfortunately, for many software systems, due to developers’ poor coding style or hectic work schedule, it is often the case that a number of methods and classes are not written with good comments. This can make it difficult for developers to understand the methods and classes, when they are performing future software maintenance tasks. To deal with this problem, in this paper we propose an approach which assesses the quality of a code comment and generates suggestions to improve comment quality. A user study is conducted to assess the effectiveness of our approach and the results show that our comment quality assessments are similar to the assessments made by our user study participants, the suggestions provided by our approach are useful to improve comment quality, and our approach can improve the accuracy of the previous comment quality analysis approaches.
Xiaobing Sun 0001, David Lo 0001, Yucong Duan, Xiangyue Liu 0002, Bin Li 0006
Int. J. Softw. Eng. Knowl. Eng.5
2014 Automatic generation of package diagram to understand Java packages
abstract
Program comprehension is a prerequisite in most software maintenance and evolution tasks. Given an unfamiliar system, it is difficult for practitioners to determine which software artifacts are relevant to the current task. Generally, there are a variety of packages in a Java software system. These packages often have different intents and different relationships between each other. Different information of packages and the relationships between different stereotypes packages form a signature of the system. This paper proposes a novel approach to automatically generate the description of the packages and its diagram to show relationships between the packages. The generated description and diagram can allow developers to more easily understand the main intent and structure of the system.
Xiaobing Sun 0001, Yun Li 0010, Xiangyue Liu 0002
ICIS4
2014 Supporting program comprehension with program summarization
abstract
A large amount of software maintenance effort is spent on program comprehension. How to accurately and quickly get the functional features in a program becomes a hot issue in program comprehension. Some studies in this area are focused on extracting the topics by analyzing linguistic information in the source code based on the textual mining techniques. However, the extracted topics are usually composed of some standalone words and difficult to understand. In this paper, we attempt to solve this problem based on a novel program summarization technique. First, we propose to use latent semantic indexing and clustering to group source artifacts with similar vocabulary to analyze the composition of each package in the program. Then, some topics composed of a vector of independent words can be extracted based on latent semantic indexing. Finally, we employ Minipar, a nature language parser, to help generate the summaries. The summaries can effectively organize the words from the topics in the form of the predefined sentence based on some rules. With such form of summaries, developers can understand what the features the program has and their corresponding source artifacts.
Xiaobing Sun 0001, Xiangyue Liu 0002, Yun Li 0010
ICIS3
2014 PFN: A novel program feature network for program comprehension
abstract
Program comprehension is one of the most frequently performed activities during software maintenance and evolution. In order to facilitate program comprehension, a variety of graphical models have been proposed in software engineering community to construct relationships between program elements. These graphical models are mostly used for understanding the system based on structural syntax dependencies between program elements. However, these graphical models fail to extract the functional or semantic features of the system. Thus, developers still cannot effectively identify the functional part in source code fit for their needs. This paper tries to fill this gap, and proposes a novel representation, program feature network (PFN), to identify the semantic features of the program at class level. PFN is generated based on the relational topic model, a hierarchical probabilistic model of networks. Based on PFN, the semantic features and the links between pairs of two classes in the program can be clearly shown. In addition, PFN can predict the possible links between the newly change request in existing program feature network rather than reconstructing the representation from the start.
Xiangyue Liu 0002, Xiaobing Sun 0001, Bin Li 0006, Junwu Zhu
ICIS1