Chun Zuo

dblp:73/2500 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
16since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 12 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
abstract
Abstract Code review is a vital but demanding aspect of software development, generating significant interest in automating review comments. Traditional evaluation methods for these comments, primarily based on text similarity, face two major challenges: inconsistent reliability of human-authored comments in open-source projects and the weak correlation of text similarity with objectives like enhancing code quality and detecting defects. This study empirically analyzes benchmark comments using a novel set of criteria informed by prior research and developer interviews. We then similarly revisit the evaluation of existing methodologies. Our evaluation framework, DeepCRCEval, integrates human evaluators and Large Language Models (LLMs) for a comprehensive reassessment of current techniques based on the criteria set. Besides, we also introduce an innovative and efficient baseline, LLM-Reviewer, leveraging the few-shot learning capabilities of LLMs for a target-oriented comparison. Our research highlights the limitations of text similarity metrics, finding that less than 10% of benchmark comments are high quality for automation. In contrast, DeepCRCEval effectively distinguishes between high and low-quality comments, proving to be a more reliable evaluation mechanism. Incorporating LLM evaluators into DeepCRCEval significantly boosts efficiency, reducing time and cost by 88.78% and 90.32%, respectively. Furthermore, LLM-Reviewer demonstrates significant potential of focusing task real targets in comment generation.
Xiaojia Li, Zihan Hua, Shiqi Cheng, Li Yang 0015, Fengjun Zhang, Chun Zuo
FASE8
2025 Towards Practical Defect-Focused Automated Code Review
abstract
The complexity of code reviews has driven efforts to automate review comments, but prior approaches oversimplify this task by treating it as snippet-level code-to-text generation and relying on text similarity metrics like BLEU for evaluation. These methods overlook repository context, real-world merge request evaluation, and defect detection, limiting their practicality. To address these issues, we explore the full automation pipeline within the online recommendation service of a company with nearly 400 million daily active users, analyzing industry-grade C++ codebases comprising hundreds of thousands of lines of code. We identify four key challenges: 1) capturing relevant context, 2) improving key bug inclusion (KBI), 3) reducing false alarm rates (FAR), and 4) integrating human workflows. To tackle these, we propose 1) code slicing algorithms for context extraction, 2) a multi-role LLM framework for KBI, 3) a filtering mechanism for FAR reduction, and 4) a novel prompt design for better human interaction. Our approach, validated on real-world merge requests from historical fault reports, achieves a 2× improvement over standard LLMs and a 10× gain over previous baselines. While the presented results focus on C++, the underlying framework design leverages language-agnostic principles (e.g., AST-based analysis), suggesting potential for broader applicability.
Xiaojia Li, Jianbing Fang, Fengjun Zhang, Li Yang 0015, Chun Zuo
ICML7
2025 AUVANA: An Efficient and Automatic Approach to Variable Rename Refactoring via Large Pre-trained Language Model
abstract
Rename refactoring is an essential practice in software maintenance, and Variable Rename Refactoring (VRR) is much more challenging than other types of identifiers. Meaningful variable names are critical for code readability and maintainability, as inconsistent variable names can hinder developers from comprehending code. Existing VRR research primarily focuses on Variable Name Consistency Checking (VCC) or variable name recommendation independently, but merely checking inconsistencies or recommending variable names is insufficient: a fully automated process must identify inconsistent names and then rectify them.In this paper, we propose AUVANA, a novel language model based framework to fully AUtomate VAriable reNAme refactoring that automates VRR by integrating inconsistency detection and meaningful variable name generation in Java. Unlike rule-based or semi-automatic approaches, AUVANA eliminates manual effort through two synergistic components: 1) a VCC model that identifies inconsistent variable names and 2) a Variable Name Refactoring (VNR) model that generates consistent replacements. To bridge the gap between pre-training and fine-tuning, we leverage prompt-tuning to improve model performance and tackle the challenge of multiple variable name occurrences. Hard negatives are introduced to address data scarcity.Experimental results demonstrate that AUVANA outperforms SoTA methods. On JavaRef and TL-CodeSum datasets, AUVANA achieves 57.8% and 56.1% Exact Match (EM) accuracy for VNR, exceeding prior baselines by 7.64% and 5.65%, respectively. For VCC, AUVANA attains 95.6% and 94.8% overall accuracy on JavaRef and TL-CodeSum, respectively, showcasing its ability to accurately detect inconsistent variable names. User study demonstrates that AUVANA VRR performance surpasses human in efficiency, precision and EM Accuracy. Artifacts are released to support future research.
Shiqi Cheng, Chenjie Shen, Li Yang 0015, Fengjun Zhang, Chun Zuo
ISSRE6
2025 Breaking Task Isolation: Enhancing Code Review Automation with Mixture-of-Experts Large Language Models
abstract
The automation of code review activities has emerged as a critical research focus for optimizing development efficiency while ensuring code quality. While recent advancements in Large Language Models (LLMs) have shown promise, existing approaches predominantly isolate the three core code review tasks—review necessity prediction, review comment generation, and code refinement, overlooking their valuable interdependencies. Empirical analysis reveals that isolated-trained comment-generation models often produce superficial comments (e.g., “Undefined ‘userInput’”) due to insufficient understanding of defect patterns, which is what necessity prediction tasks precisely target. Recent efforts to model interdependencies through knowledge distillation remain constrained by static framework designs.To address these challenges, we present MoE-Reviewer, which adopts the Mixture-of-Experts (MoE) framework on the LLaMA model to tackle the interdependence of code review tasks. MoE-Reviewer enables collaborative modeling for the three tasks mentioned above. By integrating dynamic coordination routing strategies and fine-grained expert mechanisms, MoE-Reviewer facilitates effective knowledge sharing across tasks while mitigating parameter interference. Evaluations conducted on the CodeReviewer dataset demonstrated that MoE-Reviewer outperforms existing methods, achieving state-of-the-art performance with an F1-score of 73.2% and improving the BLEU score for review comment generation by 5.32 to 11.62. Additionally, routing analysis further validates the effectiveness of our approach.
Jiayue Tang, Li Yang 0015, Zhirong Huang, Fengjun Zhang, Chun Zuo
ISSRE7
2025 Leveraging Mixture-of-Experts Framework for Smart Contract Vulnerability Repair with Large Language Model
abstract
Smart contracts are a core component of blockchain ecosystems, but their transparency and immutability make them vulnerable to attacks, leading to significant financial losses. Thus, repairing vulnerabilities in smart contracts is crucial for establishing a trustworthy blockchain environment. Existing smart contract vulnerability repair methods suffer from a critical "one-for-all" design limitation, where a single model is tasked with fixing diverse vulnerability types, leading to suboptimal performance due to insufficient specialization. To address this, we propose MoEFix, a novel framework leveraging a Mixture-of-Experts (MoE) architecture tailored for smart contract characteristics. MoEFix partitions vulnerabilities into subspaces, trains specialized experts for each type (e.g., reentrancy, integer overflow), and employs a vulnerability-aware router to dynamically allocate repairs. We further redesign the repair workflow to align with large language models, enabling end-to-end secure contract generation instead of partial patches, and to achieve this, we curated a dataset of 1,391 contracts covering five critical vulnerability types.To validate our approach, we extend the benchmark PVD test suite. Experiments demonstrate that MoEFix outperforms state-of-the-art methods by 21.64% in overall accuracy, achieving improvements of 26.19% (reentrancy) and 23.08% (delegatecall) for specific vulnerabilities.
Xizhi Hou, Li Yang 0015, Jiayue Tang, Jiadong Xu, Yifei Liu 0002, Fengjun Zhang, Chun Zuo
ASE9
2025 EXE-Reviewer: Towards EXplainable and Effective Review Comments Generation
abstract
Modern code review is essential for software quality, but the complexity of codebases and time demands of manual reviews drive interest in automation for greater efficiency and consistency.However, current automated methods often fail to generate meaningful review comments and lack explainability, limiting developers' understanding and trust.This paper presents EXE-Reviewer, aimed at generating more EXplainable and Effective review comments.To enhance effectiveness, we integrate focus information into an existing model to improve its ability to extract key insights, thereby elevating comment quality.To improve explainability, we connect explanatory information (justification behind solutions) to causality, utilizing causality extraction techniques and introducing an explanatory loss.Furthermore, we devise two metrics to assess the quantity and quality of explanatory content, enhancing insight into the model's explanations.We compare EXE-Reviewer to state-ofthe-art methods in terms of effectiveness and explainability of the generated review comments.Experimental results show that EXE-Reviewer achieves a BLEU-4 score of 7.36%, surpassing the state-of-the-art baseline of 18.52%.Meanwhile, both explainability metrics and empirical study demonstrate notable improvements in explainability of the review comments generated by EXE-Reviewer, highlighting the effectiveness of our approach in generating accurate and comprehensible review comments to developers.
Yifei Liu 0002, Li Yang 0015, Xiaoxiao Ma 0005, Jiajia Ma, Fengjun Zhang, Chun Zuo
SEKE9
2024 Dependency-Aware Method Naming Framework with Generative Adversarial Sampling
abstract
Method naming plays a vital role in code readability and maintenance. Researchers have proposed various approaches to automate method name recommendation and consistency checking task. However, two issues still remain unsolved: 1) Current work mainly focuses on local implementation and class-enclosed contexts, while dependency information is not fully exploited for the method name recommendation (MNR) task. 2) As a binary classification task, the method name consistency checking (MCC) task lacks high-quality negative samples severely, posing a challenge to train model. In this paper, we propose DMNA, a method naming framework with dependencies and generative adversarial sampling, which could help alleviate the above-mentioned problems. First, we introduce dependency information with other method contexts into training, which helps improve the performance of MNR task. Second, we leverage the model tuned for MNR task to generate high-quality adversarial samples for MCC task. Finally, we utilize prompt tuning to align the downstream task objective with the pre-training task, which helps alleviate the discrepancy problem and exploit the potential of pre-trained models. We validate the effectiveness of our approach on five widely-adopted datasets. Experimental results show that DMNA scores 49.1%, 58.5%, 63.5%, 75.4% on exact match accuracy for four MNR datasets, outperforming the SoTA baseline by at least 6.2%. And DMNA improves the accuracy of MCC task from 80.8% to 81.8%.
Chenjie Shen, Li Yang 0015, Chun Zuo
IJCNN5
2024 Exploring the impact of code review factors on the code review comment generation
Zhangyi Li, Chenjie Shen, Li Yang 0015, Chun Zuo
Autom. Softw. Eng.5
2023 Detecting Flash Loan Based Attacks in Ethereum
abstract
Decentralized Finance (DeFi) ecosystem has grown rapidly in the past few years. In the DeFi ecosystem, flash loan is a novel type of uncollateralized loan with nearly negligible lending costs. Malicious attackers can easily borrow a large number of crypto assets, and utilize them to disrupt the price of crypto assets to make a profit. Many flash loan based price manipulation attacks have been reported recently, and caused immense economic losses, e.g., 30 million USD in a single attack. In this paper, we conduct an empirical study on real-world flash loan based attacks in the past two years and present three attack patterns for price manipulation attacks. Then, we propose an approach, LeiShen, to automatically detect price manipulation attacks with asset transfers. We evaluate LeiShen on the first 14,500,000 blocks in Ethereum, and detect 180 attacks with a precision of 78.9%. Among our newly-found attacks, the severest attack has caused a total loss of more than 6.1 million USD.
Qing Xia 0007, Zhirong Huang, Wensheng Dou, Yafeng Zhang, Fengjun Zhang, Geng Liang, Chun Zuo
ICDCS7
2023 DccGraph: Detecting Criminal Communities with Augmented Criminal Network Construction and Graph Neural Network
abstract
A criminal community is an interior group where individuals commit criminal activities with high intention. Therefore, the detection is of great importance to prevent potential crimes early in the stage. Prior studies focused on methods in modularity or network analysis based on topology. These approaches, however, do not work well for detecting minority communities, which is also a key issue in criminal detection. The main reasons are: 1) modularity-based approach cannot identify the inside community structure due to the resolution limit, 2) topology-based network analysis cannot fully leverage personal feature information, such as the amount and frequency of criminal transactions. To address these problems, this paper proposes a novel framework named DccGraph (Detect criminal communities using a Graph neural network) to enhance the overall performance of detecting criminal communities, especially minority ones. First, we extract the feature information of criminals and balance the distribution of criminal community members to construct an Augmented Criminal Network(ACN), which alleviates representation collapse and distinguish the feature of criminals in minority communities. In that case, it is capable to go beyond the resolution limit and locate minority communities effectively. Then, we design a criminal-oriented siamese graph encoder to capture both structural and feature information of criminals in the ACN. Specifically, feature interference and connection disturbance of criminals are employed to enrich the feature representation. To the best of our knowledge, DccGraph is the first framework to apply a graph neural network on criminal community detection. Experiments on several real-life dataset and benchmark datasets show that: DccGraph successfully outperforms eight baselines by 29.25%, 48.04%, 35.37%, and 40.98% on ACC, NMI, ARI, and F1, respectively. The dataset and the code for this framework are publicly available.
Yuanzhe Yang, Li Yang 0015, Lingwei Li, Xiaoxiao Ma 0005, Chun Zuo
IJCNN6
2023 LLaMA-Reviewer: Advancing Code Review Automation with Large Language Models through Parameter-Efficient Fine-Tuning
abstract
The automation of code review activities, a long-standing pursuit in software engineering, has been primarily addressed by numerous domain-specific pre-trained models. Despite their success, these models frequently demand extensive resources for pre-training from scratch. In contrast, Large Language Models (LLMs) provide an intriguing alternative, given their remarkable capabilities when supplemented with domain-specific knowledge. However, their potential for automating code review tasks remains largely unexplored.In response to this research gap, we present LLaMA-Reviewer, an innovative framework that leverages the capabilities of LLaMA, a popular LLM, in the realm of code review. Mindful of resource constraints, this framework employs parameter-efficient fine-tuning (PEFT) methods, delivering high performance while using less than 1% of trainable parameters.An extensive evaluation of LLaMA-Reviewer is conducted on two diverse, publicly available datasets. Notably, even with the smallest LLaMA base model consisting of 6.7B parameters and a limited number of tuning epochs, LLaMA-Reviewer equals the performance of existing code-review-focused models.The ablation experiments provide insights into the influence of various fine-tuning process components, including input representation, instruction tuning, and different PEFT methods. To foster continuous progress in this field, the code and all PEFT-weight plugins have been made open-source.
Xiaojia Li, Li Yang 0015, Chun Zuo
ISSRE5
2023 Automating Method Naming with Context-Aware Prompt-Tuning
abstract
Method names are crucial to program comprehension and maintenance. Recently, many approaches have been proposed to automatically recommend method names and detect inconsistent names. Despite promising, their results are still suboptimal considering the three following drawbacks: 1) These models are mostly trained from scratch, learning two different objectives simultaneously. The misalignment between two objectives will negatively affect training efficiency and model performance. 2) The enclosing class context is not fully exploited, making it difficult to learn the abstract functionality of the method. 3) Current method name consistency checking methods follow a generate-then-compare process, which restricts the accuracy as they highly rely on the quality of generated names and face difficulty measuring the semantic consistency.In this paper, we propose an approach named AUMENA to AUtomate MEthod NAming tasks with context-aware prompt-tuning. Unlike existing deep learning based approaches, our model first learns the contextualized representation(i.e., class attributes) of programming language and natural language through the pre-training model, then fully exploits the capacity and knowledge of large language model with prompt-tuning to precisely detect inconsistent method names and recommend more accurate names. To better identify semantically consistent names, we model the method name consistency checking task as a two-class classification problem, avoiding the limitation of previous generate-then-compare consistency checking approaches. Experiment results reflect that AUMENA scores 68.6%, 72.0%, 73.6%, 84.7% on four datasets of method name recommendation, surpassing the state-of-the-art baseline by 8.5%, 18.4%, 11.0%, 12.0%, respectively. And our approach scores 80.8% accuracy on method name consistency checking, reaching an 5.5% outperformance. All data and trained models are publicly available.
Lingwei Li, Li Yang 0015, Xiaoxiao Ma 0005, Chun Zuo
ICPC5
2023 Mining the Relationship between Object-Relational Mapping Performance Anti-patterns and Code Clones
abstract
The use of Object-Relational Mapping (ORM) in software development has become increasingly popular due to its superiority to simplify database interactions.Despite the prosperous development of ORM code smells detection tools for general code smell problems related to coupling and cohesion, these tools do not capture issues that are specific to ORM code statement context.In this work, we fill the gap wit h the potential performance anti-patterns of repetitive ORM code by heuristic analysis and code clone analysis on 6 open source ORM systems in Java and Python (Saler, Wagtail, Zulip, Taiga, Protal, and Roller).For each occurrence of this code smell, we distinguished problematic instances that potentially require further fixes among justifiable ones.Through our research, we identified four antipatterns associated with repetitive ORM code and proposed fix strategy for each of these anti-patterns.Additionally, our study delves into the relationship between repetitive ORM code anti-patterns and code clone, which reveals that a substantial proportion of repetitive ORM statements can be found in cloned code.Experiments show that repetitive ORM code can lead to a waste of system performance.This research highlights the impact of ORM code context on the proper use of ORM frameworks and emphasizes that copying ORM code without context evaluation can be detrimental to system performance.
Zeshan Xu, Li Yang 0015, Chun Zuo
SEKE4
2022 STPChain: a Crowdsourced Software Engineering Method for Software Traceability and Fine-grained Privacy Based on Blockchain
abstract
Crowdsourced software engineering (CSE) has be-come an increasingly popular way of software development owing to its flexibility. There exist two types of CSE participants: the requester, who posts a task including a set of requirements, and workers, including developers and testers, completing the task. Due to the centralized architecture, traditional CSE systems have raised the concern of untrustfulness since participants may collude with centralized platforms and behave maliciously to grab illegitimate interests. Some researches have utilized blockchain to solve the problem. Still, they either lack software traceability, which is significant for ensuring software quality, or cannot guar-antee users' transaction privacy owing to blockchain's openness. We propose STPChain, a CSE method for §_oftware Traceability and Privacy based on blockChain. To maintain software traceability, we articulate the process flow of CSE and implement it with smart contracts adapted to software development. The smart contracts ensure that every submission will automatically leave tamper-proof records. Credible software traceability links are realized through the records. To alleviate the transaction privacy leakage, we propose FGCA, a fine-grained CA (Certificate Authority) updating mechanism in consortium blockchain. Based on the finding that a traceability link belongs within a task, FGCA refines the digital certificates from user-level to task-level to conceal the connection between tasks and users. While guaranteeing transaction privacy, FGCA also keeps the traceability links by storing the relations between users' identifier and certificates. Security analysis demonstrates that STPChain can prevent malicious misbehaviors of participants in CSE. Case study illustrates that our method can be utilized to maintain traceability and save more than 70% of time when positioning relevant workers in CSE without transaction privacy leakage related to data openness. Performance experiments show the applicability of STPChain under CSE scenarios.
Li Yang 0015, Qing Xia 0007, Mingzhe Fang, Geng Liang, Chun Zuo
COMPSAC6
2022 AUGER: automatically generating review comments with pre-training models
abstract
Code review is one of the best practices as a powerful safeguard for software quality. In practice, senior or highly skilled reviewers inspect source code and provide constructive comments, consider- ing what authors may ignore, for example, some special cases. The collaborative validation between contributors results in code being highly qualified and less chance of bugs. However, since personal knowledge is limited and varies, the efficiency and effectiveness of code review practice are worthy of further improvement. In fact, it still takes a colossal and time-consuming effort to deliver useful review comments. This paper explores a synergy of multiple practical review comments to enhance code review and proposes AUGER (AUtomatically GEnerating Review comments): a review comments generator with pre-training models. We first collect empirical review data from 11 notable Java projects and construct a dataset of 10,882 code changes. By leveraging Text-to-Text Transfer Transformer (T5) models, the framework synthesizes valuable knowledge in the training stage and effectively outperforms baselines by 37.38% in ROUGE-L. 29% of our automatic review comments are considered useful according to prior studies. The inference generates just in 20 seconds and is also open to training further. Moreover, the performance also gets improved when thoroughly analyzed in case study.
Lingwei Li, Li Yang 0015, Huaxi Jiang, Tiejian Luo, Zihan Hua, Geng Liang, Chun Zuo
ESEC/SIGSOFT FSE8
2021 DeepRelease: Language-agnostic Release Notes Generation from Pull Requests of Open-source Software
abstract
The release note is an essential software artifact of open-source software that documents crucial information about changes, such as new features and bug fixes. With the help of release notes, both developers and users could have a general understanding of the latest version without browsing the source code. However, it is a daunting and time-consuming job for developers to produce release notes. Although prior studies have provided some automatic approaches, they generate release notes mainly by extracting information from code changes. This will result in language-specific and not being general enough to be applicable. Therefore, helping developers produce release notes effectively remains an unsolved challenge. To address the problem, we first conduct a manual study on the release notes of 900 GitHub projects, which reveals that more than 54% of projects produce their release notes with pull requests. Based on the empirical finding, we propose a deep learning based approach named DeepRelease (Deep learning based Release notes generator) to generate release notes according to pull requests. The process of release notes generation in DeepRelease includes the change entries generation and the change category (i.e., new features or bug fixes) generation, which are formulated as a text summarization task and a multi-class classification problem, respectively. Since DeepRelease fully employs text information from pull requests to summarize changes and identify the change category, it is language-agnostic and can be used for projects in any language. We build a dataset with over 46K release notes and evaluate DeepRelease on the dataset. The experimental results indicate that DeepRelease outperforms four baselines and can generate release notes similar to those manually written ones in a fraction of the time.
Huaxi Jiang, Li Yang 0015, Geng Liang, Chun Zuo
APSEC5
2005 An Effective Two-Stage Neural Network Model and Its Application on Flood Loss Prediction
Li Yang 0015, Chun Zuo, Yuguo Wang
ISNN (3)2
2004 Decision support system of flood disaster for property insurance: theory and practice
abstract
In the paper, the status of flood disaster in China and the progress of disaster prevention and reduction in the field of property insurance were analyzed. The characteristics and application fields of 3S (GIS, RS and GPS) were also introduced. According to the current need and future development of property insurance company, which were based on the investigation to the work of disaster prevention and reduction in property insurance and casualty company (abbreviated as PICC) China, the authors decided to apply 3S technologies to the field of property insurance and used the successful methods in foreign property insurance companies for reference to develop a decision support system of flood disaster for property insurance. The system linked well with the operational system of insurance company and realized seamless integration between different data sources. The work of disaster prevention and reduction in property insurance company was better organized both in theory ways and in key technologies. As a result, the economic benefit of property insurance company was really improved. The system was applied in Shenzhen, China and the result was satisfying
Lin Wang 0011, Qiming Qin, Vasit Sagan, Chun Zuo
IGARSS5