Chenjie Shen

dblp:256/7617 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0000-6212-9721ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection
abstract
With the increasing security issues in blockchain, smart contract vulnerability detection has become a research focus. Existing vulnerability detection methods have their limitations: 1) Static analysis methods struggle with complex scenarios. 2) Methods based on specialized pre-trained models perform well on specific datasets but have limited generalization capabilities. In contrast, general-purpose Large Language Models (LLMs) demonstrate impressive ability in adapting to new vulnerability patterns. However, they often underperform on specific vulnerability types compared to methods based on specialized pre-trained models. We also observe that explanations generated by generalpurpose LLMs can provide fine-grained code understanding information, contributing to improved detection performance. Inspired by these observations, we propose SAEL, a LLMbased framework for smart contract vulnerability detection. First, we design prompts targeting specific smart contract vulnerabilities to guide general-purpose LLMs in detecting vulnerabilities and providing explanations. The detection results generated by LLMs serve as prediction features. Then, we employ prompt-tuning on CodeT5 and T5 respectively to process contract code and explanations, enhancing model performance on specific tasks. To leverage the strengths of each component, we introduce Adaptive Mixture-of-Experts, a dynamic architecture for smart contract vulnerability detection. This mechanism dynamically adjusts feature weights through a Gating Network, which selects the most relevant features by applying TopK filtering and Softmax normalization, and a Multi-Head Self-Attention mechanism, which enhances cross-feature relationships by processing multiple attention heads in parallel. This design ensures that prediction results for LLMs, explanation features, and contract code features are effectively integrated through gradient optimization. The loss function focuses on the independent prediction performance of each feature and the overall performance of weighted predictions. Experimental results show that SAEL outperforms existing methods in detecting various vulnerabilities.
Shiqi Cheng, Zhirong Huang, Chenjie Shen, Li Yang 0015, Fengjun Zhang, Jiajia Ma
ICSME5
2025 AUVANA: An Efficient and Automatic Approach to Variable Rename Refactoring via Large Pre-trained Language Model
abstract
Rename refactoring is an essential practice in software maintenance, and Variable Rename Refactoring (VRR) is much more challenging than other types of identifiers. Meaningful variable names are critical for code readability and maintainability, as inconsistent variable names can hinder developers from comprehending code. Existing VRR research primarily focuses on Variable Name Consistency Checking (VCC) or variable name recommendation independently, but merely checking inconsistencies or recommending variable names is insufficient: a fully automated process must identify inconsistent names and then rectify them.In this paper, we propose AUVANA, a novel language model based framework to fully AUtomate VAriable reNAme refactoring that automates VRR by integrating inconsistency detection and meaningful variable name generation in Java. Unlike rule-based or semi-automatic approaches, AUVANA eliminates manual effort through two synergistic components: 1) a VCC model that identifies inconsistent variable names and 2) a Variable Name Refactoring (VNR) model that generates consistent replacements. To bridge the gap between pre-training and fine-tuning, we leverage prompt-tuning to improve model performance and tackle the challenge of multiple variable name occurrences. Hard negatives are introduced to address data scarcity.Experimental results demonstrate that AUVANA outperforms SoTA methods. On JavaRef and TL-CodeSum datasets, AUVANA achieves 57.8% and 56.1% Exact Match (EM) accuracy for VNR, exceeding prior baselines by 7.64% and 5.65%, respectively. For VCC, AUVANA attains 95.6% and 94.8% overall accuracy on JavaRef and TL-CodeSum, respectively, showcasing its ability to accurately detect inconsistent variable names. User study demonstrates that AUVANA VRR performance surpasses human in efficiency, precision and EM Accuracy. Artifacts are released to support future research.
Shiqi Cheng, Chenjie Shen, Li Yang 0015, Fengjun Zhang, Chun Zuo
ISSRE2
2024 Dependency-Aware Method Naming Framework with Generative Adversarial Sampling
abstract
Method naming plays a vital role in code readability and maintenance. Researchers have proposed various approaches to automate method name recommendation and consistency checking task. However, two issues still remain unsolved: 1) Current work mainly focuses on local implementation and class-enclosed contexts, while dependency information is not fully exploited for the method name recommendation (MNR) task. 2) As a binary classification task, the method name consistency checking (MCC) task lacks high-quality negative samples severely, posing a challenge to train model. In this paper, we propose DMNA, a method naming framework with dependencies and generative adversarial sampling, which could help alleviate the above-mentioned problems. First, we introduce dependency information with other method contexts into training, which helps improve the performance of MNR task. Second, we leverage the model tuned for MNR task to generate high-quality adversarial samples for MCC task. Finally, we utilize prompt tuning to align the downstream task objective with the pre-training task, which helps alleviate the discrepancy problem and exploit the potential of pre-trained models. We validate the effectiveness of our approach on five widely-adopted datasets. Experimental results show that DMNA scores 49.1%, 58.5%, 63.5%, 75.4% on exact match accuracy for four MNR datasets, outperforming the SoTA baseline by at least 6.2%. And DMNA improves the accuracy of MCC task from 80.8% to 81.8%.
Chenjie Shen, Li Yang 0015, Chun Zuo
IJCNN1
2024 Exploring the impact of code review factors on the code review comment generation
Zhangyi Li, Chenjie Shen, Li Yang 0015, Chun Zuo
Autom. Softw. Eng.3