EDBT 2026 Demo / reviewers in the wild / expert
Li Yang 0015
dblp:09/3925-15
· DBLP profile ↗
23ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0001-8364-6525ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 15 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Binary Message Passing for Generalizable Semi-Supervised Graph Anomaly DetectionabstractGraph Neural Networks (GNNs) have achieved impressive performance in semi-supervised graph anomaly detection (GAD). While many GNN variants have been developed for this task, they largely focus on advanced message aggregation schemes, leaving the message routing aspect underexplored. We argue that the commonly used broadcast-based routing can also hinder generalization, particularly in the presence of rare and structurally challenging (vertices with a high-degree) anomalies. To address this, we propose Binary Message Passing (BMP), a novel routing paradigm that models the message flow of each vertex as a binary tree (BMP tree), where vanilla graph convolution is decoupled by its left and right subtrees. Each vertex recursively gathers information from neighbors with higher anomaly probabilities within each subtree, thereby amplifying the propagation of anomaly information across the topology. The anomaly probabilities are estimated and updated by the model itself, enabling adaptive, self-supervised routing over iterations. Furthermore, combining multiple BMP trees into a BMP forest provides multi-scale structural context, enhancing the expressiveness of final vertex embeddings. Extensive experiments show that BMP improves detection performance under limited supervision while exhibiting better generalization across structurally diverse anomalies. Li Yang 0015, Fengjun Zhang |
AAAI | 4 |
| 2026 | SQL-Commenter: Aligning Large Language Models for SQL Comment Generation with Direct Preference OptimizationabstractSQL query comprehension is a significant challenge in database and data analysis environments due to complex syntax, diverse join types, and deep nesting. Despite its critical role in backend development and data science, many queries, particularly within legacy systems, often lack adequate comments, which severely hinders code readability, maintainability, and knowledge transfer. Existing approaches to automated SQL comment generation face two main challenges: limited training datasets that inadequately represent real-world analytical queries involving multi-table joins, window functions, and complex aggregations, and an insufficient understanding of SQL-specific logical semantics and schema-related context by Large Language Models (LLMs), even after standard training. Our empirical analysis shows that even after continual pre-training and supervised fine-tuning, LLMs struggle to precisely understand complex SQL semantics, leading to inaccurate or incomplete comments. To address these challenges, we propose SQL-Commenter, an advanced comment generation method based on LLaMA-3.1-8B. First, we construct a comprehensive dataset containing longer, more complex SQL queries with expert-verified, detailed comments. Second, we perform continual pre-training using a large-scale SQL corpus to enhance the LLM’s understanding of SQL syntax and semantics. Then, we conduct supervised fine-tuning with our high-quality dataset. Finally, we introduce Direct Preference Optimization (DPO), which leverages human feedback to significantly improve comment quality. SQL-Commenter utilizes a preference-based loss function that encourages the LLM to increase the probability of preferred outputs while decreasing the probability of non-preferred outputs, thereby enhancing both fine-grained semantic learning, such as distinguishing between different join types, and context-dependent quality assessment based on business logic. We evaluate SQL-Commenter on the authoritative Spider and Bird benchmarks, where it significantly outperforms state-of-the-art baselines. On average, across these datasets, our method surpasses the strongest baseline (Qwen3-14B) by 9.29, 4.99, and 13.23 percentage points on BLEU-4, METEOR, and ROUGE-L, respectively. Moreover, human evaluation demonstrates the superior quality of comments generated by SQL-Commenter in terms of correctness, completeness, and naturalness. Li Yang 0015, Changzhi Deng, Jiajia Ma, Fengjun Zhang |
ICPC | 6 |
| 2026 | Strunkmap: An Abstract Approach to Understand Spatiotemporal Density DistributionabstractVisual analysis of spatiotemporal density distributions is crucial for understanding spatiotemporal dynamics. However, existing methods suffer from visual occlusion and information loss when simultaneously displaying multiple density distributions. We present Strunkmap as an abstract approach to address these challenges. We introduce anisotropic kernel density estimation to enhance the accuracy of density generation. We extract the trunks of density distributions to identify the overall spatial patterns. Path scanning and trunk-outline matching strategies are employed to preserve local spatial structure. We design a stacked trunk plot that enables lossless density representation while conserving substantial screen space. Based on the visual design, Strunkmap integrates multiple heatmaps within a single map to effectively display temporal evolution of density distributions without visual occlusion. Ablation studies and comparative experiments validate the superiority of Strunkmap in accuracy and efficiency for hotspot identification and trend exploration. Theoretical analysis demonstrates Strunkmap's scalability, which we further verify through large-scale spatiotemporal data visualization. Color encoding schemes and scaling ratios are discussed to illustrate the flexibility. Our evaluations with user feedback demonstrate that Strunkmap is a viable solution with significant potential to real-world applications. Zhirong Huang, Jiajia Ma, Shiqi Cheng, Ruize Zhou, Xiaoxiao Ma 0005, Li Yang 0015, Fengjun Zhang |
IEEE Trans. Vis. Comput. Graph. | 11 |
| 2025 | DeepCRCEval: Revisiting the Evaluation of Code Review Comment GenerationabstractAbstract Code review is a vital but demanding aspect of software development, generating significant interest in automating review comments. Traditional evaluation methods for these comments, primarily based on text similarity, face two major challenges: inconsistent reliability of human-authored comments in open-source projects and the weak correlation of text similarity with objectives like enhancing code quality and detecting defects. This study empirically analyzes benchmark comments using a novel set of criteria informed by prior research and developer interviews. We then similarly revisit the evaluation of existing methodologies. Our evaluation framework, DeepCRCEval, integrates human evaluators and Large Language Models (LLMs) for a comprehensive reassessment of current techniques based on the criteria set. Besides, we also introduce an innovative and efficient baseline, LLM-Reviewer, leveraging the few-shot learning capabilities of LLMs for a target-oriented comparison. Our research highlights the limitations of text similarity metrics, finding that less than 10% of benchmark comments are high quality for automation. In contrast, DeepCRCEval effectively distinguishes between high and low-quality comments, proving to be a more reliable evaluation mechanism. Incorporating LLM evaluators into DeepCRCEval significantly boosts efficiency, reducing time and cost by 88.78% and 90.32%, respectively. Furthermore, LLM-Reviewer demonstrates significant potential of focusing task real targets in comment generation. Xiaojia Li, Zihan Hua, Shiqi Cheng, Li Yang 0015, Fengjun Zhang, Chun Zuo |
FASE | 6 |
| 2025 | Towards Practical Defect-Focused Automated Code ReviewabstractThe complexity of code reviews has driven efforts to automate review comments, but prior approaches oversimplify this task by treating it as snippet-level code-to-text generation and relying on text similarity metrics like BLEU for evaluation. These methods overlook repository context, real-world merge request evaluation, and defect detection, limiting their practicality. To address these issues, we explore the full automation pipeline within the online recommendation service of a company with nearly 400 million daily active users, analyzing industry-grade C++ codebases comprising hundreds of thousands of lines of code. We identify four key challenges: 1) capturing relevant context, 2) improving key bug inclusion (KBI), 3) reducing false alarm rates (FAR), and 4) integrating human workflows. To tackle these, we propose 1) code slicing algorithms for context extraction, 2) a multi-role LLM framework for KBI, 3) a filtering mechanism for FAR reduction, and 4) a novel prompt design for better human interaction. Our approach, validated on real-world merge requests from historical fault reports, achieves a 2× improvement over standard LLMs and a 10× gain over previous baselines. While the presented results focus on C++, the underlying framework design leverages language-agnostic principles (e.g., AST-based analysis), suggesting potential for broader applicability. Xiaojia Li, Jianbing Fang, Fengjun Zhang, Li Yang 0015, Chun Zuo |
ICML | 6 |
| 2025 | SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability DetectionabstractWith the increasing security issues in blockchain, smart contract vulnerability detection has become a research focus. Existing vulnerability detection methods have their limitations: 1) Static analysis methods struggle with complex scenarios. 2) Methods based on specialized pre-trained models perform well on specific datasets but have limited generalization capabilities. In contrast, general-purpose Large Language Models (LLMs) demonstrate impressive ability in adapting to new vulnerability patterns. However, they often underperform on specific vulnerability types compared to methods based on specialized pre-trained models. We also observe that explanations generated by generalpurpose LLMs can provide fine-grained code understanding information, contributing to improved detection performance. Inspired by these observations, we propose SAEL, a LLMbased framework for smart contract vulnerability detection. First, we design prompts targeting specific smart contract vulnerabilities to guide general-purpose LLMs in detecting vulnerabilities and providing explanations. The detection results generated by LLMs serve as prediction features. Then, we employ prompt-tuning on CodeT5 and T5 respectively to process contract code and explanations, enhancing model performance on specific tasks. To leverage the strengths of each component, we introduce Adaptive Mixture-of-Experts, a dynamic architecture for smart contract vulnerability detection. This mechanism dynamically adjusts feature weights through a Gating Network, which selects the most relevant features by applying TopK filtering and Softmax normalization, and a Multi-Head Self-Attention mechanism, which enhances cross-feature relationships by processing multiple attention heads in parallel. This design ensures that prediction results for LLMs, explanation features, and contract code features are effectively integrated through gradient optimization. The loss function focuses on the independent prediction performance of each feature and the overall performance of weighted predictions. Experimental results show that SAEL outperforms existing methods in detecting various vulnerabilities. Shiqi Cheng, Zhirong Huang, Chenjie Shen, Li Yang 0015, Fengjun Zhang, Jiajia Ma |
ICSME | 7 |
| 2025 | AUVANA: An Efficient and Automatic Approach to Variable Rename Refactoring via Large Pre-trained Language ModelabstractRename refactoring is an essential practice in software maintenance, and Variable Rename Refactoring (VRR) is much more challenging than other types of identifiers. Meaningful variable names are critical for code readability and maintainability, as inconsistent variable names can hinder developers from comprehending code. Existing VRR research primarily focuses on Variable Name Consistency Checking (VCC) or variable name recommendation independently, but merely checking inconsistencies or recommending variable names is insufficient: a fully automated process must identify inconsistent names and then rectify them.In this paper, we propose AUVANA, a novel language model based framework to fully AUtomate VAriable reNAme refactoring that automates VRR by integrating inconsistency detection and meaningful variable name generation in Java. Unlike rule-based or semi-automatic approaches, AUVANA eliminates manual effort through two synergistic components: 1) a VCC model that identifies inconsistent variable names and 2) a Variable Name Refactoring (VNR) model that generates consistent replacements. To bridge the gap between pre-training and fine-tuning, we leverage prompt-tuning to improve model performance and tackle the challenge of multiple variable name occurrences. Hard negatives are introduced to address data scarcity.Experimental results demonstrate that AUVANA outperforms SoTA methods. On JavaRef and TL-CodeSum datasets, AUVANA achieves 57.8% and 56.1% Exact Match (EM) accuracy for VNR, exceeding prior baselines by 7.64% and 5.65%, respectively. For VCC, AUVANA attains 95.6% and 94.8% overall accuracy on JavaRef and TL-CodeSum, respectively, showcasing its ability to accurately detect inconsistent variable names. User study demonstrates that AUVANA VRR performance surpasses human in efficiency, precision and EM Accuracy. Artifacts are released to support future research. Shiqi Cheng, Chenjie Shen, Li Yang 0015, Fengjun Zhang, Chun Zuo |
ISSRE | 3 |
| 2025 | Breaking Task Isolation: Enhancing Code Review Automation with Mixture-of-Experts Large Language ModelsabstractThe automation of code review activities has emerged as a critical research focus for optimizing development efficiency while ensuring code quality. While recent advancements in Large Language Models (LLMs) have shown promise, existing approaches predominantly isolate the three core code review tasks—review necessity prediction, review comment generation, and code refinement, overlooking their valuable interdependencies. Empirical analysis reveals that isolated-trained comment-generation models often produce superficial comments (e.g., “Undefined ‘userInput’”) due to insufficient understanding of defect patterns, which is what necessity prediction tasks precisely target. Recent efforts to model interdependencies through knowledge distillation remain constrained by static framework designs.To address these challenges, we present MoE-Reviewer, which adopts the Mixture-of-Experts (MoE) framework on the LLaMA model to tackle the interdependence of code review tasks. MoE-Reviewer enables collaborative modeling for the three tasks mentioned above. By integrating dynamic coordination routing strategies and fine-grained expert mechanisms, MoE-Reviewer facilitates effective knowledge sharing across tasks while mitigating parameter interference. Evaluations conducted on the CodeReviewer dataset demonstrated that MoE-Reviewer outperforms existing methods, achieving state-of-the-art performance with an F1-score of 73.2% and improving the BLEU score for review comment generation by 5.32 to 11.62. Additionally, routing analysis further validates the effectiveness of our approach. Jiayue Tang, Li Yang 0015, Zhirong Huang, Fengjun Zhang, Chun Zuo |
ISSRE | 2 |
| 2025 | Leveraging Mixture-of-Experts Framework for Smart Contract Vulnerability Repair with Large Language ModelabstractSmart contracts are a core component of blockchain ecosystems, but their transparency and immutability make them vulnerable to attacks, leading to significant financial losses. Thus, repairing vulnerabilities in smart contracts is crucial for establishing a trustworthy blockchain environment. Existing smart contract vulnerability repair methods suffer from a critical "one-for-all" design limitation, where a single model is tasked with fixing diverse vulnerability types, leading to suboptimal performance due to insufficient specialization. To address this, we propose MoEFix, a novel framework leveraging a Mixture-of-Experts (MoE) architecture tailored for smart contract characteristics. MoEFix partitions vulnerabilities into subspaces, trains specialized experts for each type (e.g., reentrancy, integer overflow), and employs a vulnerability-aware router to dynamically allocate repairs. We further redesign the repair workflow to align with large language models, enabling end-to-end secure contract generation instead of partial patches, and to achieve this, we curated a dataset of 1,391 contracts covering five critical vulnerability types.To validate our approach, we extend the benchmark PVD test suite. Experiments demonstrate that MoEFix outperforms state-of-the-art methods by 21.64% in overall accuracy, achieving improvements of 26.19% (reentrancy) and 23.08% (delegatecall) for specific vulnerabilities. Xizhi Hou, Li Yang 0015, Jiayue Tang, Jiadong Xu, Yifei Liu 0002, Fengjun Zhang, Chun Zuo |
ASE | 4 |
| 2025 | EXE-Reviewer: Towards EXplainable and Effective Review Comments GenerationabstractModern code review is essential for software quality, but the complexity of codebases and time demands of manual reviews drive interest in automation for greater efficiency and consistency.However, current automated methods often fail to generate meaningful review comments and lack explainability, limiting developers' understanding and trust.This paper presents EXE-Reviewer, aimed at generating more EXplainable and Effective review comments.To enhance effectiveness, we integrate focus information into an existing model to improve its ability to extract key insights, thereby elevating comment quality.To improve explainability, we connect explanatory information (justification behind solutions) to causality, utilizing causality extraction techniques and introducing an explanatory loss.Furthermore, we devise two metrics to assess the quantity and quality of explanatory content, enhancing insight into the model's explanations.We compare EXE-Reviewer to state-ofthe-art methods in terms of effectiveness and explainability of the generated review comments.Experimental results show that EXE-Reviewer achieves a BLEU-4 score of 7.36%, surpassing the state-of-the-art baseline of 18.52%.Meanwhile, both explainability metrics and empirical study demonstrate notable improvements in explainability of the review comments generated by EXE-Reviewer, highlighting the effectiveness of our approach in generating accurate and comprehensible review comments to developers. Yifei Liu 0002, Li Yang 0015, Xiaoxiao Ma 0005, Jiajia Ma, Fengjun Zhang, Chun Zuo |
SEKE | 3 |
| 2025 | Topology Augmented Multi-Band and Multi-Scale Filtering for Graph Anomaly DetectionabstractGraph Anomaly Detection (GAD) has gained significant attention in areas such as financial risk control and social network security, becoming a critical research problem. Vanilla Graph Neural Networks (GNNs), a popular method for graph modeling, are known to perform poorly in GAD due to the assumption of homophily preferences. This article argues that the issue lies in the insufficient feature extraction ability caused by their single filtering property (low-pass filtering) and revealing the effectiveness of multi-band filtering to deal with GAD. From this, we note two other overlooked issues: (1) How can multi-band band-pass filtering further fuse multi-scale neighborhood information? (2) Adaptation between raw attributes of nodes and graph filters (graph topology). The former bridges the respective advantages of spectral domain and spatial domain, and the latter is an important bottleneck for the encoding capacity of the filters. To address these, we propose a new GAD method, Graph Perturbed Networks (GraphPN). Each hidden layer of GraphPN is a band-pass filter, enabling multi-band and multi-scale filtering through simple stacking and skip connections. We analyze its spectral locality and spatial locality to provide theoretical support. Additionally, GraphPN is supplemented with a tailored feature activation module to complete the adaptation of the above two. This module readjusts node indices and decouples graph convolution, introducing rich topological information to node attributes. In addition to further enhancing detection performance, another possibly counter-intuitive effect is that the distinguishability of the two classes of nodes is improved even before filtering. The proposed method performs well in real-world datasets compared with the current state-of-the-art baselines, which fully demonstrates its superiority. Codes are available at https://github.com/Thankstaro/GraphPN . Zhirong Huang, Li Yang 0015, Fengjun Zhang |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | Dependency-Aware Method Naming Framework with Generative Adversarial SamplingabstractMethod naming plays a vital role in code readability and maintenance. Researchers have proposed various approaches to automate method name recommendation and consistency checking task. However, two issues still remain unsolved: 1) Current work mainly focuses on local implementation and class-enclosed contexts, while dependency information is not fully exploited for the method name recommendation (MNR) task. 2) As a binary classification task, the method name consistency checking (MCC) task lacks high-quality negative samples severely, posing a challenge to train model. In this paper, we propose DMNA, a method naming framework with dependencies and generative adversarial sampling, which could help alleviate the above-mentioned problems. First, we introduce dependency information with other method contexts into training, which helps improve the performance of MNR task. Second, we leverage the model tuned for MNR task to generate high-quality adversarial samples for MCC task. Finally, we utilize prompt tuning to align the downstream task objective with the pre-training task, which helps alleviate the discrepancy problem and exploit the potential of pre-trained models. We validate the effectiveness of our approach on five widely-adopted datasets. Experimental results show that DMNA scores 49.1%, 58.5%, 63.5%, 75.4% on exact match accuracy for four MNR datasets, outperforming the SoTA baseline by at least 6.2%. And DMNA improves the accuracy of MCC task from 80.8% to 81.8%. Chenjie Shen, Li Yang 0015, Chun Zuo |
IJCNN | 4 |
| 2024 | Exploring the impact of code review factors on the code review comment generation
Zhangyi Li, Chenjie Shen, Li Yang 0015, Chun Zuo |
Autom. Softw. Eng. | 4 |
| 2023 | DccGraph: Detecting Criminal Communities with Augmented Criminal Network Construction and Graph Neural NetworkabstractA criminal community is an interior group where individuals commit criminal activities with high intention. Therefore, the detection is of great importance to prevent potential crimes early in the stage. Prior studies focused on methods in modularity or network analysis based on topology. These approaches, however, do not work well for detecting minority communities, which is also a key issue in criminal detection. The main reasons are: 1) modularity-based approach cannot identify the inside community structure due to the resolution limit, 2) topology-based network analysis cannot fully leverage personal feature information, such as the amount and frequency of criminal transactions. To address these problems, this paper proposes a novel framework named DccGraph (Detect criminal communities using a Graph neural network) to enhance the overall performance of detecting criminal communities, especially minority ones. First, we extract the feature information of criminals and balance the distribution of criminal community members to construct an Augmented Criminal Network(ACN), which alleviates representation collapse and distinguish the feature of criminals in minority communities. In that case, it is capable to go beyond the resolution limit and locate minority communities effectively. Then, we design a criminal-oriented siamese graph encoder to capture both structural and feature information of criminals in the ACN. Specifically, feature interference and connection disturbance of criminals are employed to enrich the feature representation. To the best of our knowledge, DccGraph is the first framework to apply a graph neural network on criminal community detection. Experiments on several real-life dataset and benchmark datasets show that: DccGraph successfully outperforms eight baselines by 29.25%, 48.04%, 35.37%, and 40.98% on ACC, NMI, ARI, and F1, respectively. The dataset and the code for this framework are publicly available. Yuanzhe Yang, Li Yang 0015, Lingwei Li, Xiaoxiao Ma 0005, Chun Zuo |
IJCNN | 2 |
| 2023 | Who Are the Money Launderers? Money Laundering Detection on Blockchain via Mutual Learning-Based Graph Neural NetworkabstractWith the development of blockchain technology, security concerns have become increasingly prominent in recent years. Money laundering through blockchain has been found to generate a significant amount of money and has become a serious threat. Towards money laundering detection in Bitcoin, conventional methods heavily rely on fixed expert rules, leading to low accuracy and poor scalability. Graph convolutional network approaches have improved this issue, but they fail to distinguish the importance of surrounding transactions and the structural information of different transactions. To solve above problems, we propose an approach to detect money laundering on blockchain by mining its transaction records, named AEtransGAT. First, we use a novel approach called transGat as an encoder to determine the significance of surrounding transactions by considering the transaction amount values of transaction flows. The original features and the features after graph embedding are combined to address the issue of feature distortion. Second, we deploy the graph autoencoder as the decoder to learn the overall structural information of different transactions, and the concatenated embedding is used to output the classification results as the detector. Finally, we propose our model based on mutual learning in this task which takes the advantages of both transactions classification loss and structure reconstruction loss. We validate the performance of our model on the Elliptic dataset which is the only large open source dataset in Bitcoin anti-money laundering. The results show that our method outperforms current state-of-the-art methods and is linearly scalable. Fengjun Zhang, Jiajia Ma, Li Yang 0015, Yuanzhe Yang |
IJCNN | 4 |
| 2023 | LLaMA-Reviewer: Advancing Code Review Automation with Large Language Models through Parameter-Efficient Fine-TuningabstractThe automation of code review activities, a long-standing pursuit in software engineering, has been primarily addressed by numerous domain-specific pre-trained models. Despite their success, these models frequently demand extensive resources for pre-training from scratch. In contrast, Large Language Models (LLMs) provide an intriguing alternative, given their remarkable capabilities when supplemented with domain-specific knowledge. However, their potential for automating code review tasks remains largely unexplored.In response to this research gap, we present LLaMA-Reviewer, an innovative framework that leverages the capabilities of LLaMA, a popular LLM, in the realm of code review. Mindful of resource constraints, this framework employs parameter-efficient fine-tuning (PEFT) methods, delivering high performance while using less than 1% of trainable parameters.An extensive evaluation of LLaMA-Reviewer is conducted on two diverse, publicly available datasets. Notably, even with the smallest LLaMA base model consisting of 6.7B parameters and a limited number of tuning epochs, LLaMA-Reviewer equals the performance of existing code-review-focused models.The ablation experiments provide insights into the influence of various fine-tuning process components, including input representation, instruction tuning, and different PEFT methods. To foster continuous progress in this field, the code and all PEFT-weight plugins have been made open-source. Xiaojia Li, Li Yang 0015, Chun Zuo |
ISSRE | 4 |
| 2023 | PSCVFinder: A Prompt-Tuning Based Framework for Smart Contract Vulnerability DetectionabstractWith the increasing security issues in the blockchain, smart contract vulnerability detection has gradually become the focus of research. Recently, many approaches have been proposed to detect smart contract vulnerabilities. Despite promising results, these approaches still have three drawbacks: 1) Symbolic execution and static analysis methods are constrained by predefined rules, which limits their adaptability to different vulnerabilities. 2) Most smart contract code contains abundant irrelevant information which is useless for vulnerability detection. 3) Pre-trained models fail to bridge the gap between pre-training and detecting smart contract vulnerabilities.To solve these problems, we propose an approach named PSCVFinder for detecting reentrancy vulnerability and times-tamp dependency vulnerability, which are two severe vulnerabilities in smart contract. To better detect these vulnerabilities, we propose CSCV which is a smart contract slicing method to reduce the irrelevant code. Unlike existing approaches, our model first learns the representation of programming language through the pre-training model, then fully exploits the capacity of large language model with prompt-tuning to precisely detect smart contract vulnerability. We conduct experiments on real-world dataset and the results reflect that PSCVFinder scores 93.83% and 93.49% on two kinds of vulnerabilities in F1-score, surpassing the state-of-the-art baseline by 1.14% and 4.02%, respectively. Xianglong Liu 0005, Li Yang 0015, Fengjun Zhang, Jiajia Ma |
ISSRE | 4 |
| 2023 | Automating Method Naming with Context-Aware Prompt-TuningabstractMethod names are crucial to program comprehension and maintenance. Recently, many approaches have been proposed to automatically recommend method names and detect inconsistent names. Despite promising, their results are still suboptimal considering the three following drawbacks: 1) These models are mostly trained from scratch, learning two different objectives simultaneously. The misalignment between two objectives will negatively affect training efficiency and model performance. 2) The enclosing class context is not fully exploited, making it difficult to learn the abstract functionality of the method. 3) Current method name consistency checking methods follow a generate-then-compare process, which restricts the accuracy as they highly rely on the quality of generated names and face difficulty measuring the semantic consistency.In this paper, we propose an approach named AUMENA to AUtomate MEthod NAming tasks with context-aware prompt-tuning. Unlike existing deep learning based approaches, our model first learns the contextualized representation(i.e., class attributes) of programming language and natural language through the pre-training model, then fully exploits the capacity and knowledge of large language model with prompt-tuning to precisely detect inconsistent method names and recommend more accurate names. To better identify semantically consistent names, we model the method name consistency checking task as a two-class classification problem, avoiding the limitation of previous generate-then-compare consistency checking approaches. Experiment results reflect that AUMENA scores 68.6%, 72.0%, 73.6%, 84.7% on four datasets of method name recommendation, surpassing the state-of-the-art baseline by 8.5%, 18.4%, 11.0%, 12.0%, respectively. And our approach scores 80.8% accuracy on method name consistency checking, reaching an 5.5% outperformance. All data and trained models are publicly available. Lingwei Li, Li Yang 0015, Xiaoxiao Ma 0005, Chun Zuo |
ICPC | 3 |
| 2023 | Mining the Relationship between Object-Relational Mapping Performance Anti-patterns and Code ClonesabstractThe use of Object-Relational Mapping (ORM) in software development has become increasingly popular due to its superiority to simplify database interactions.Despite the prosperous development of ORM code smells detection tools for general code smell problems related to coupling and cohesion, these tools do not capture issues that are specific to ORM code statement context.In this work, we fill the gap wit h the potential performance anti-patterns of repetitive ORM code by heuristic analysis and code clone analysis on 6 open source ORM systems in Java and Python (Saler, Wagtail, Zulip, Taiga, Protal, and Roller).For each occurrence of this code smell, we distinguished problematic instances that potentially require further fixes among justifiable ones.Through our research, we identified four antipatterns associated with repetitive ORM code and proposed fix strategy for each of these anti-patterns.Additionally, our study delves into the relationship between repetitive ORM code anti-patterns and code clone, which reveals that a substantial proportion of repetitive ORM statements can be found in cloned code.Experiments show that repetitive ORM code can lead to a waste of system performance.This research highlights the impact of ORM code context on the proper use of ORM frameworks and emphasizes that copying ORM code without context evaluation can be detrimental to system performance. Zeshan Xu, Li Yang 0015, Chun Zuo |
SEKE | 3 |
| 2022 | STPChain: a Crowdsourced Software Engineering Method for Software Traceability and Fine-grained Privacy Based on BlockchainabstractCrowdsourced software engineering (CSE) has be-come an increasingly popular way of software development owing to its flexibility. There exist two types of CSE participants: the requester, who posts a task including a set of requirements, and workers, including developers and testers, completing the task. Due to the centralized architecture, traditional CSE systems have raised the concern of untrustfulness since participants may collude with centralized platforms and behave maliciously to grab illegitimate interests. Some researches have utilized blockchain to solve the problem. Still, they either lack software traceability, which is significant for ensuring software quality, or cannot guar-antee users' transaction privacy owing to blockchain's openness. We propose STPChain, a CSE method for §_oftware Traceability and Privacy based on blockChain. To maintain software traceability, we articulate the process flow of CSE and implement it with smart contracts adapted to software development. The smart contracts ensure that every submission will automatically leave tamper-proof records. Credible software traceability links are realized through the records. To alleviate the transaction privacy leakage, we propose FGCA, a fine-grained CA (Certificate Authority) updating mechanism in consortium blockchain. Based on the finding that a traceability link belongs within a task, FGCA refines the digital certificates from user-level to task-level to conceal the connection between tasks and users. While guaranteeing transaction privacy, FGCA also keeps the traceability links by storing the relations between users' identifier and certificates. Security analysis demonstrates that STPChain can prevent malicious misbehaviors of participants in CSE. Case study illustrates that our method can be utilized to maintain traceability and save more than 70% of time when positioning relevant workers in CSE without transaction privacy leakage related to data openness. Performance experiments show the applicability of STPChain under CSE scenarios. Li Yang 0015, Qing Xia 0007, Mingzhe Fang, Geng Liang, Chun Zuo |
COMPSAC | 2 |
| 2022 | AUGER: automatically generating review comments with pre-training modelsabstractCode review is one of the best practices as a powerful safeguard for software quality. In practice, senior or highly skilled reviewers inspect source code and provide constructive comments, consider- ing what authors may ignore, for example, some special cases. The collaborative validation between contributors results in code being highly qualified and less chance of bugs. However, since personal knowledge is limited and varies, the efficiency and effectiveness of code review practice are worthy of further improvement. In fact, it still takes a colossal and time-consuming effort to deliver useful review comments. This paper explores a synergy of multiple practical review comments to enhance code review and proposes AUGER (AUtomatically GEnerating Review comments): a review comments generator with pre-training models. We first collect empirical review data from 11 notable Java projects and construct a dataset of 10,882 code changes. By leveraging Text-to-Text Transfer Transformer (T5) models, the framework synthesizes valuable knowledge in the training stage and effectively outperforms baselines by 37.38% in ROUGE-L. 29% of our automatic review comments are considered useful according to prior studies. The inference generates just in 20 seconds and is also open to training further. Moreover, the performance also gets improved when thoroughly analyzed in case study. Lingwei Li, Li Yang 0015, Huaxi Jiang, Tiejian Luo, Zihan Hua, Geng Liang, Chun Zuo |
ESEC/SIGSOFT FSE | 2 |
| 2021 | DeepRelease: Language-agnostic Release Notes Generation from Pull Requests of Open-source SoftwareabstractThe release note is an essential software artifact of open-source software that documents crucial information about changes, such as new features and bug fixes. With the help of release notes, both developers and users could have a general understanding of the latest version without browsing the source code. However, it is a daunting and time-consuming job for developers to produce release notes. Although prior studies have provided some automatic approaches, they generate release notes mainly by extracting information from code changes. This will result in language-specific and not being general enough to be applicable. Therefore, helping developers produce release notes effectively remains an unsolved challenge. To address the problem, we first conduct a manual study on the release notes of 900 GitHub projects, which reveals that more than 54% of projects produce their release notes with pull requests. Based on the empirical finding, we propose a deep learning based approach named DeepRelease (Deep learning based Release notes generator) to generate release notes according to pull requests. The process of release notes generation in DeepRelease includes the change entries generation and the change category (i.e., new features or bug fixes) generation, which are formulated as a text summarization task and a multi-class classification problem, respectively. Since DeepRelease fully employs text information from pull requests to summarize changes and identify the change category, it is language-agnostic and can be used for projects in any language. We build a dataset with over 46K release notes and evaluate DeepRelease on the dataset. The experimental results indicate that DeepRelease outperforms four baselines and can generate release notes similar to those manually written ones in a fraction of the time. Huaxi Jiang, Li Yang 0015, Geng Liang, Chun Zuo |
APSEC | 3 |
| 2005 | An Effective Two-Stage Neural Network Model and Its Application on Flood Loss Prediction
Li Yang 0015, Chun Zuo, Yuguo Wang |
ISNN (3) | 1 |