Yiheng Shen 0002

dblp:257/3189-2 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0009-6278-633XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Divide and conquer: Optimizing code Chain-of-Thought in Small Language Model
Yiheng Shen 0002, Guang Yang 0019
Eng. Appl. Artif. Intell.1
2024 Bash comment generation via data augmentation and semantic-aware CodeBERT
Yiheng Shen 0002, Xiaolin Ju, Xiang Chen 0005, Guang Yang 0019
Autom. Softw. Eng.1
2024 Automatic smart contract comment generation via large language models and in-context learning
Xiang Chen 0005, Guang Yang 0019, Yiheng Shen 0002
Inf. Softw. Technol.4
2023 APICom: Automatic API Completion via Prompt Learning and Adversarial Training-based Data Augmentation
abstract
Based on developer needs and usage scenarios, API (Application Programming Interface) recommendation is the process of assisting developers in finding the required API among numerous candidate APIs. Previous studies mainly modeled API recommendation as the recommendation task, which can recommend multiple candidate APIs for the given query, and developers may not yet be able to find what they need. Motivated by the neural machine translation research domain, we can model this problem as the generation task, which aims to directly generate the required API for the developer query. After our preliminary investigation, we find the performance of this intuitive approach is not promising. The reason is that there exists an error when generating the prefixes of the API. However, developers may know certain API prefix information during actual development in most cases. Therefore, we model this problem as the automatic completion task and propose a novel approach APICom based on prompt learning, which can generate API related to the query according to the prompts (i.e., API prefix information). Moreover, the effectiveness of APICom highly depends on the quality of the training dataset. In this study, we further design a novel gradient-based adversarial training method ATCom for data augmentation, which can improve the normalized stability when generating adversarial examples. To evaluate the effectiveness of APICom, we consider a corpus of 33k developer queries and corresponding APIs. Compared with the state-of-the-art baselines, our experimental results show that APICom can outperform all baselines by at least 40.02%, 13.20%, and 16.31% in terms of the performance measures EM@1, MRR, and MAP. Finally, our ablation studies confirm the effectiveness of our component setting (such as our designed adversarial training method, our used pre-trained model, and prompt learning) in APICom.
Yafeng Gu, Yiheng Shen 0002, Xiang Chen 0005, Shaoyu Yang 0002, Zhixiang Cao
Internetware2
2023 An Empirical Study of Adversarial Training in Code Comment Generation
abstract
The code comment generation task is designed for developers to understand programs more quickly during development and maintenance.However, the existing automatic code comment generation models can not generate valuable comments for developers.It is necessary to explore a technology that can optimize the performance of code comment generation models without changing the model.We consider adversarial training as the experimental object, which can improve the robustness and generalization of the model.We present a large-scale study to experimentally validate the performance of gradient-based adversarial training methods in the code comment generation task.The results show that adversarial training can improve the model performance by generating adversarial examples without changing the model.Our empirical study can provide a new perspective for researchers to improve the performance of code comment generation models.
Yiheng Shen 0002, Xiaolin Ju, Xiang Chen 0005, Guang Yang 0019
SEKE1
2021 AGFL: A Graph Convolutional Neural Network-Based Method for Fault Localization
abstract
Fault localization techniques have been developed for decades. Spectrum Based Fault Localization (SBFL) is a popular strategy in this research topic. However, SBFL is well known for low accuracy, mainly due to simply using a coverage matrix of program executions. In this paper, we propose a method based on graph neural network (AGFL), characterized by the adjacent matrix of the abstract syntax tree and the word vector of each program token. Referring to the Dstar, we calculate the suspiciousness of the statements and rank these statements. The experiment carried on Defects4J, a widely used benchmark, reveals that AGFL can locate 178 of the 262 studied bugs within Top-1, while state-of-the-art techniques at most locate 148 within Top-1. We also investigate the impacts of hyper-parameters (e.g., epoch and learning rate). The results show that AGFL has the best effect when the epoch is 100 and the learning rate is 0.0001. This value of epoch and learning rate increases by 66% compared to the worst on Top-1.
Xiaolin Ju, Xiang Chen 0005, Hao Shen 0011, Yiheng Shen 0002
QRS5