Wei Ma 0014

dblp:32/32-14 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-0044-466XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 25 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Temporal Requirement Formalization Using Large Language Models
Wei Ma 0014, Qian Wang 0002, Lingxiao Jiang, Dongsheng Li 0001
SANER2
2026 Effective Fine-tuning for Low-resource Languages: A Case Study of Cangjie
Zhaofeng Liu, Mingyi Zhou, Zihan Huang, Wei Ma 0014, Li Li 0029
Empir. Softw. Eng.6
2026 OptRCA: A More Efficient and Accurate Approach for Automated Root Cause Analysis and Explanation
abstract
With the development of automated software testing technology, software developers can get a large number of crash test cases in a short period of time. However, analyzing these crash test cases and finding their root cause is a time-consuming and labor-intensive task. Techniques based on reverse execution and backward taint analysis are proposed to locate the root cause, but can’t provide context information or explanation of the underlying fault. To address these two limitations, researchers have proposed an automated root cause analysis technique called AURORA. Although this technique provides powerful root cause analysis capabilities, it also have two obvious shortcomings. First, the results of root cause analysis are not accurate enough. Second, the efficiency of root cause analysis is not high enough. In order to improve these two shortcomings, we propose OptRCA, a more efficient and accurate approach for root cause analysis and explanation. Like AURORA’s fuzzing strategy, OptRCA is also designed based on AFL’s crash mode. The difference between them is mainly reflected in three points. First of all, the goal pursued by OptRCA is different from that of normal fuzzing technology. OptRCA pursues maximum correlation to ensure that as many crash test cases as possible are related to the same root cause. This test case with maximum correlation can greatly improve the accuracy of root cause analysis. Second, OptRCA proposed a more efficient non-crash test case retention strategy, which we named “Hill-Climbing Retention.” Using the hill-climbing retention method, OptRCA can obtain sufficient root cause information while retaining only a few non-crash test cases. Since the number of test cases is greatly reduced, the efficiency of OptRCA’s subsequent root cause analysis process is also greatly improved. In addition, OptRCA also optimizes the analysis formula to obtain more accurate analysis results. In the evaluation experimental results, OptRCA is significantly better than AURORA in terms of accuracy and efficiency. Quantitative analysis shows that OptRCA is 65% more accurate and 61% more efficient than AURORA.
Jingquan Ge, Yaowen Zheng, Yuekang Li, Wei Ma 0014, Sheikh Mahbub Habib, Praveen Kakkolangara, Gabriel Byman, Yang Liu 0003
ACM Trans. Softw. Eng. Methodol.4
2026 LIMR: Intent-Aware Mashup API Recommendation via LLM-Augmented Multi-Scale Fusion
abstract
The increasing availability of Web APIs has amplified the complexity of mashup creation, where developers must identify compatible and functionally relevant APIs based on often ambiguous natural language descriptions. Traditional methods also fall short in capturing hierarchical semantic cues, modeling compatibility, and aligning with developer intent. Although large language models (LLMs) offer strong generalization capabilities, they remain unreliable in mashup recommendation due to hallucinated outputs, limited controllability, and token-length constraints when dealing with large-scale API repositories. To overcome these limitations, we introduceLIMR, an intent-aware mashup recommendation framework that combines LLM-augmented semantic reasoning with structured, multi-scale neural modeling.LIMRfirst prompts a LLM to extract high-level intent from user requirements, which serves as a global semantic signal. This intent is fused with low-level, multi-scale features extracted by a convolutional encoder, which are designed to capture fine-grained lexical/phrasal patterns at different granularities and provide precise semantic grounding for API matching. These heterogeneous representations are further contextually refined through a Transformer-based interaction module. To handle nonlinear semantic dependencies and compositional complexity,LIMRintegrates a Kolmogorov-Arnold Network (KAN) with learnable activation functions, enhancing the model's capacity to capture intricate feature interactions. The entire framework is optimized via LLM, incorporating auxiliary objectives such as mashup category prediction and API quality estimation to guide generalization and reduce overfitting. Comprehensive experiments on the ProgrammableWeb and APIBench datasets show thatLIMRsignificantly outperforms state-of-the-art baselines, which the ranking-oriented metrics, including NDCG and mAP, achieves improvements of 17.1%–34.2% over the strongest competitors. These results confirm the effectiveness ofLIMR's hybrid design in delivering precise, robust, and intent-aware mashup API recommendations, especially in scenarios where LLMs alone fail to meet accuracy and scalability demands.
Yao Zhang 0019, Yude Bai, Minhong Dong, Keqing Cen, Ji Zhang 0001, Wei Ma 0014, Yongqiang Lyu 0001, Xiaohong Li 0001, Junjie Wang 0007, Lingxiao Jiang, Yang Liu 0003
IEEE Trans. Serv. Comput.7
2026 Towards Secure Program Partitioning for Smart Contracts With LLM's In-Context Learning
abstract
Smart contracts are highly susceptible to manipulation attacks due to the leakage of sensitive information. Addressing manipulation vulnerabilities is particularly challenging because they stem from inherent data confidentiality issues rather than straightforward implementation bugs. To tackle this by preventing sensitive information leakage, we present PARTITIONGPT, the first LLM-driven approach that combines static analysis with the in-context learning capabilities of large language models (LLMs) to partition smart contracts into critical (privileged) and normal codebases, guided by a few annotated sensitive data variables. We evaluated PARTITIONGPT on 18 annotated smart contracts containing 99 sensitive functions. The results demonstrate that PARTITIONGPT successfully generatescompilable, andverifiedpartitions, achieving a precision of 80% while reducing more than 26% code compared to functionlevel partitioning approach. Furthermore, we evaluated PARTITIONGPT on nine real-world manipulation attacks that led to a total loss of 25 million dollars, PARTITIONGPT effectively prevents eight cases, highlighting its potential for broad applicability and the necessity for secure program partitioning during smart contract development to diminish manipulation vulnerabilities.
Ye Liu 0012, Yuqing Niu, Chengyan Ma 0001, Ruidong Han, Wei Ma 0014, Yi Li 0008, Debin Gao, David Lo 0001
IEEE Trans. Software Eng.5
2026 Evaluating Large Language Models for Line-Level Vulnerability Localization
abstract
Recently, Automated Vulnerability Localization (AVL) has attracted growing attention, aiming to facilitate diagnosis by pinpointing the specific lines of code responsible for vulnerabilities. Large Language Models (LLMs) have shown potential in various domains, yet their effectiveness in line-level vulnerability localization remains underexplored.In this work, we present the first comprehensive empirical evaluation of LLMs for AVL. Our study examines 19 leading LLMs suitable for code analysis, including ChatGPT and multiple open-source models, spanning encoder-only, encoder-decoder, and decoder-only architectures, with model sizes from 60M to 70B parameters. We evaluate three paradigms—few-shot prompting, discriminative fine-tuning, and generative fine-tuning—with and without Low-Rank Adaptation (LoRA), on both a BigVul-derived dataset for C/C++ and a smart contract vulnerability dataset.Our results show that discriminative fine-tuning achieves substantial performance gains over existing learning-based AVL methods when sufficient training data is available. In low-data settings, prompting advanced LLMs such as ChatGPT proves more effective. We also identify challenges related to input length and unidirectional context during fine-tuning, and propose two remedial strategies: a sliding window approach and right-forward embedding, both of which yield significant improvements. Moreover, we provide the first assessment of LLM generalizability in AVL, showing that certain models can transfer effectively across Common Weakness Enumerations (CWEs) and projects. However, performance degrades notably for newly discovered vulnerabilities containing unfamiliar lexical or structural patterns, underscoring the need for continual adaptation. These findings offer practical guidance for deploying LLM-based AVL systems in realistic software security workflows.
Jian Zhang 0087, Chong Wang 0013, Anran Li 0001, Weisong Sun, Cen Zhang, Wei Ma 0014, Yang Liu 0003
IEEE Trans. Software Eng.6
2025 An Analytical Perspective on Software Engineering for Large Language Models
Tianlin Li, Chong Wang 0013, Jian Zhang 0087, Wei Ma 0014, Aishan Liu, Jingyi Wang 0004, Yang Liu 0003
ICECCS5
2025 Combining Fine-Tuning and LLM-Based Agents for Intuitive Smart Contract Auditing with Justifications
abstract
Smart contracts are decentralized applications built atop blockchains like Ethereum. Recent research has shown that large language models (LLMs) have potential in auditing smart contracts, but the state-of-the-art indicates that even GPT-4 can achieve only 30% precision (when both decision and justification are correct). This is likely because off-the-shelf LLMs were primarily pre-trained on a general text/code corpus and not fine-tuned on the specific domain of Solidity smart contract auditing. In this paper, we propose iAudit, a general framework that combines fine-tuning and LLM-based agents for intuitive smart contract auditing with justifications. Specifically, iAudit is inspired by the observation that expert human auditors first perceive what could be wrong and then perform a detailed analysis of the code to identify the cause. As such, iAudit employs a two-stage fine-tuning approach: it first tunes a Detector model to make decisions and then tunes a Reasoner model to generate causes of vulnerabilities. However, fine-tuning alone faces challenges in accurately identifying the optimal cause of a vulnerability. Therefore, we introduce two LLM-based agents, the Ranker and Critic, to iteratively select and debate the most suitable cause of vulnerability based on the output of the fine-tuned Reasoner model. To evaluate iAudit, we collected a balanced dataset with 1,734 positive and 1,810 negative samples to fine-tune iAudit. We then compared it with traditional fine-tuned models (CodeBERT, GraphCodeBERT, CodeT5, and UnixCoder) as well as prompt learning-based LLMs (GPT4, GPT-3.5, and CodeLlama-13b/34b). On a dataset of 263 real smart contract vulnerabilities, iAudit achieves an F1 score of 91.21% and an accuracy of 91.11%. The causes generated by iAudit achieved a consistency of about 38% compared to the ground truth causes.
Wei Ma 0014, Daoyuan Wu, Yuqiang Sun 0001, Tianwen Wang, Shangqing Liu, Jian Zhang 0087, Yue Xue, Yang Liu 0003
ICSE1
2025 Detecting DeFi Fraud With a Graph-Transformer Language Model
Wei Ma 0014, Jiaxi Qiu, Cong Wu 0003, Jing Chen 0003, Lingxiao Jiang, Shangqing Liu, Yang Liu 0003, Yang Xiang 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Assessing the Robustness of Test Selection Methods for Deep Neural Networks
abstract
Regularly testing deep learning-powered systems on newly collected data is critical to ensure their reliability, robustness, and efficacy in real-world applications. This process is demanding due to the significant time and human effort required for labeling new data. While test selection methods alleviate manual labor by labeling and evaluating only a subset of data while meeting testing criteria, we observe that such methods with reported promising results are simply evaluated, e.g., testing on original test data. The question arises: are they always reliable? In this paper, we explore when and to what extent test selection methods fail. First, we identify potential pitfalls of 11 selection methods based on their construction. Second, we conduct a study to empirically confirm the existence of these pitfalls. Furthermore, we demonstrate how pitfalls can break the reliability of these methods. Concretely, methods for fault detection suffer from data that are: 1) correctly classified but uncertain, or 2) misclassified but confident. Remarkably, the test relative coverage achieved by such methods drops by up to 86.85%. Besides, methods for performance estimation are sensitive to the choice of intermediate-layer output. The effectiveness of such methods can be even worse than random selection when using an inappropriate layer.
Yuejun Guo 0001, Xiaofei Xie, Maxime Cordy, Wei Ma 0014, Mike Papadakis, Lei Ma 0003, Yves Le Traon
ACM Trans. Softw. Eng. Methodol.5
2025 Open Source AI-based SE Tools: Opportunities and Challenges of Collaborative Software Learning
abstract
Large language models (LLMs) have become instrumental in advancing software engineering (SE) tasks, showcasing their efficacy in code understanding and beyond. AI code models have demonstrated their value not only in code generation but also in defect detection, enhancing security measures and improving overall software quality. They are emerging as crucial tools for both software development and maintaining software quality. Like traditional SE tools, open source collaboration is key in realizing the excellent products. However, with AI models, the essential need is in data. The collaboration of these AI-based SE models hinges on maximizing the sources of high-quality data. However, data, especially of high quality, often hold commercial or sensitive value, making them less accessible for open source AI-based SE projects. This reality presents a significant barrier to the development and enhancement of AI-based SE tools within the SE community. Therefore, researchers need to find solutions for enabling open source AI-based SE models to tap into resources by different organizations. Addressing this challenge, our position article investigates one solution to facilitate access to diverse organizational resources for open source AI models, ensuring that privacy and commercial sensitivities are respected. We introduce a governance framework centered on federated learning (FL), designed to foster the joint development and maintenance of open source AI code models while safeguarding data privacy and security. Additionally, we present guidelines for developers on AI-based SE tool collaboration, covering data requirements, model architecture, updating strategies, and version control. Given the significant influence of data characteristics on FL, our research examines the effect of code data heterogeneity on FL performance. We consider six different scenarios of data distributions and include four code models. We also include four most common FL algorithms. Our experimental findings highlight the potential for employing FL in the collaborative development and maintenance of AI-based SE models. We also discuss the key issues to be addressed in the co-construction process and future research directions.
Wei Ma 0014, Tao Lin 0004, Yaowen Zheng, Jingquan Ge, Jun Wang 0020, Jacques Klein, Tegawendé F. Bissyandé, Yang Liu 0003, Li Li 0029
ACM Trans. Softw. Eng. Methodol.2
2025 A Comprehensive Study of Governance Issues in Decentralized Finance Applications
abstract
Decentralized Finance (DeFi) is a prominent application of smart contracts, representing a novel financial paradigm in contrast to centralized finance. While DeFi applications are rapidly emerging on mainstream blockchain platforms, their quality varies greatly, presenting numerous challenges, particularly in terms of their governance mechanisms. In this paper, we present a comprehensive study of governance issues in DeFi applications. Initially, we collected 3,165 academic papers and numerous industry reports. After thorough screening, we selected 44 academic papers and 11 industry reports for detailed analysis. Drawing upon insights from industry reports and academic research articles, we develop a taxonomy to categorize these governance issues. We collect and build a dataset of 4,446 audit reports from seventeen Web3 security companies, categorizing their governance issues according to our constructed taxonomy. We conducted a thorough analysis of governance issues and identified vulnerabilities in the governance design and implementation, e.g., voting sybil attack and proposal front-running. Our statistical analysis indicates that a significant portion (35.48%) of governance-related issues is classified as severe. Within these, ownership-related problems constitute the largest share (65.38%). Despite DeFi governance being essential for the long-term success of DeFi projects, our data shows that both auditors and development teams have not fully grasped its significance. Based on audit reports, we also analyzed common vulnerabilities and issues in the governance domain. Our research identifies two primary categories of DeFi governance issues: technology-centric and human-centric. Technology-centric issues can be addressed through technology updates and iterations, whereas human-centric issues are influenced not only by the development team's technical skills but also by their understanding of DeFi governance. Data analysis reveals that design and implementation issues are frequently overlooked; although not directly associated with vulnerabilities, these issues can impact the equitable distribution of project benefits. Furthermore, our analysis of 104 projects’ tokenomics configurations, including 15 collected from DeFi platforms, uncovered 27 inconsistent configurations, with only two projects exhibiting no issues. This suggests that such issues are relatively common. We therefore advise project teams to ensure consistency between their tokenomics design and the actual code. Our study culminates in providing several key practical implications for various DeFi stakeholders, including developers, users, researchers, and regulators, aiming to deepen the understanding of DeFi governance issues and contribute to the robust growth of DeFi systems.
Wei Ma 0014, Chenguang Zhu 0002, Ye Liu 0012, Xiaofei Xie, Yi Li 0008
ACM Trans. Softw. Eng. Methodol.1
2025 An Empirical Study of Exploring the Capabilities of Large Language Models in Code Learning
abstract
Since the advent of ChatGPT, large language models (LLMs) have attracted widespread attention from academia and industry. They have also brought significant changes to software engineering. However, until now, there has been a lack of comprehensive studies comparing LLMs with previous smaller code pre-trained models. To address this gap, we conduct a study in this paper to illustrate the performance of LLMs in different software engineering tasks. Specifically, we select three open-source large language models, CodeGen, LLaMA, and StarCoder, for the research targets, and our study is conducted from four aspects, including code syntax understanding, code semantic reasoning, encoding representation quality, and adaptation performance for different software engineering tasks to compare LLMs with previous code pre-trained models. Four aspects build on each other, forming important components of AI for Software Engineering.We conclude that: (1) Compared with previous smaller pre-trained models like CodeBERT, LLMs exhibit distinct trends in how they learn code syntax or semantics as the number of layers increases. Additionally, mastering code semantics proves to be more challenging, with semantic information usually learned in the final layers; (2) Causal decoder architecture with left-to-right attention masking does not perform well in zero-shot tasks; (3) For classification tasks, the mean vector representation generated by LLMs over a sequence tends to outperform the last token representation in the sequence; (4) Incorporating parameter-efficient fine-tuning techniques into LLMs for downstream tasks can help LLMs achieve better performance than previous code pre-trained models on code generation tasks but may not be optimal in some code understanding tasks; (5) LoRA emerges as a more effective PEFT technique for LLMs in downstream code-related tasks. We hope these findings will better guide future researchers in designing more powerful code models.
Shangqing Liu, Daya Guo, Jian Zhang 0087, Wei Ma 0014, Yanzhou Li, Yang Liu 0003
IEEE Trans. Software Eng.4
2024 An Empirical Study on Noisy Label Learning for Program Understanding
abstract
Recently, deep learning models have been widely applied in program understanding tasks, and these models achieve state-of-the-art results on many benchmark datasets. A major challenge of deep learning for program understanding is that the effectiveness of these approaches depends on the quality of their datasets, and these datasets often contain noisy data samples. A typical kind of noise in program understanding datasets is label noise, which means that the target outputs for some inputs are incorrect.
Wenhan Wang, Yanzhou Li, Anran Li 0001, Jian Zhang 0087, Wei Ma 0014, Yang Liu 0003
ICSE5
2024 How Effective Are They? Exploring Large Language Model Based Fuzz Driver Generation
abstract
Fuzz drivers are essential for library API fuzzing. However, automatically generating fuzz drivers is a complex task, as it demands the creation of high-quality, correct, and robust API usage code. An LLM-based (Large Language Model) approach for generating fuzz drivers is a promising area of research. Unlike traditional program analysis-based generators, this text-based approach is more generalized and capable of harnessing a variety of API usage information, resulting in code that is friendly for human readers. However, there is still a lack of understanding regarding the fundamental issues on this direction, such as its effectiveness and potential challenges. To bridge this gap, we conducted the first in-depth study targeting the important issues of using LLMs to generate effective fuzz drivers. Our study features a curated dataset with 86 fuzz driver generation questions from 30 widely-used C projects. Six prompting strategies are designed and tested across five state-of-the-art LLMs with five different temperature settings. In total, our study evaluated 736,430 generated fuzz drivers, with 0.85 billion token costs ($8,000+ charged tokens). Additionally, we compared the LLM-generated drivers against those utilized in industry, conducting extensive fuzzing experiments (3.75 CPU-year). Our study uncovered that: 1) While LLM-based fuzz driver generation is a promising direction, it still encounters several obstacles towards practical applications; 2) LLMs face difficulties in generating effective fuzz drivers for APIs with intricate specifics. Three featured design choices of prompt strategies can be beneficial: issuing repeat queries, querying with examples, and employing an iterative querying process; 3) While LLM-generated drivers can yield fuzzing outcomes that are on par with those used in the industry, there are substantial opportunities for enhancement, such as extending contained API usage, or integrating semantic oracles to facilitate logical bug detection. Our insights have been implemented to improve the OSS-Fuzz-Gen project, facilitating practical fuzz driver generation in industry.
Cen Zhang, Yaowen Zheng, Mingqiang Bai, Yeting Li, Wei Ma 0014, Xiaofei Xie, Yuekang Li, Limin Sun 0001, Yang Liu 0003
ISSTA5
2024 Enhancing Code Vulnerability Detection via Vulnerability-Preserving Data Augmentation
abstract
Source code vulnerability detection aims to identify inherent vulnerabilities to safeguard software systems from potential attacks. Many prior studies overlook diverse vulnerability characteristics, simplifying the problem into a binary (0-1) classification task for example determining whether it is vulnerable or not. This poses a challenge for a single deep-learning based model to effectively learn the wide array of vulnerability characteristics. Furthermore, due to the challenges associated with collecting large-scale vulnerability data, these detectors often overfit limited training datasets, resulting in lower model generalization performance. To address the aforementioned challenges, in this work, we introduce a fine-grained vulnerability detector namely FGVulDet. Unlike previous approaches, FGVulDet employs multiple classifiers to discern characteristics of various vulnerability types and combines their outputs to identify the specific type of vulnerability. Each classifier is designed to learn type-specific vulnerability semantics. Additionally, to address the scarcity of data for some vulnerability types and enhance data diversity for learning better vulnerability semantics, we propose a novel vulnerability-preserving data augmentation technique to augment the number of vulnerabilities. Taking inspiration from recent advancements in graph neural networks for learning program semantics, we incorporate a Gated Graph Neural Network (GGNN) and extend it to an edge-aware GGNN to capture edge-type information. FGVulDet is trained on a large-scale dataset from GitHub, encompassing five different types of vulnerabilities. Extensive experiments compared with static-analysis-based approaches and learning-based approaches have demonstrated the effectiveness of FGVulDet.
Shangqing Liu, Wei Ma 0014, Jian Wang 0067, Xiaofei Xie, Yang Liu 0003
LCTES2
2024 Towards Exploring the Limitations of Test Selection Techniques on Graph Neural Networks: An Empirical Study
abstract
Abstract Graph Neural Networks (GNNs) have gained prominence in various domains, such as social network analysis, recommendation systems, and drug discovery, due to their ability to model complex relationships in graph-structured data. GNNs can exhibit incorrect behavior, resulting in severe consequences. Therefore, testing is necessary and pivotal. However, labeling all test inputs for GNNs can be prohibitively costly and time-consuming, especially when dealing with large and complex graphs. In response to these challenges, test selection has emerged as a strategic approach to alleviate labeling expenses. The objective of test selection is to select a subset of tests from the complete test set. While various test selection techniques have been proposed for traditional deep neural networks (DNNs), their adaptation to GNNs presents unique challenges due to the distinctions between DNN and GNN test data. Specifically, DNN test inputs are independent of each other, whereas GNN test inputs (nodes) exhibit intricate interdependencies. Therefore, it remains unclear whether DNN test selection approaches can perform effectively on GNNs. To fill the gap, we conduct an empirical study that systematically evaluates the effectiveness of various test selection methods in the context of GNNs, focusing on three critical aspects: 1) Misclassification detection : selecting test inputs that are more likely to be misclassified; 2) Accuracy estimation : selecting a small set of tests to precisely estimate the accuracy of the whole testing set; 3) Performance enhancement : selecting retraining inputs to improve the GNN accuracy. Our empirical study encompasses 7 graph datasets and 8 GNN models, evaluating 22 test selection approaches. Our study includes not only node classification datasets but also graph classification datasets. Our findings reveal that: 1) In GNN misclassification detection, confidence-based test selection methods, which perform well in DNNs, do not demonstrate the same level of effectiveness; 2) In terms of GNN accuracy estimation, clustering-based methods, while consistently performing better than random selection, provide only slight improvements; 3) Regarding selecting inputs for GNN performance improvement, test selection methods, such as confidence-based and clustering-based test selection methods, demonstrate only slight effectiveness; 4) Concerning performance enhancement, node importance-based test selection methods are not suitable, and in many cases, they even perform worse than random selection.
Xueqi Dang, Wei Ma 0014, Yuejun Guo 0001, Mike Papadakis, Maxime Cordy, Yves Le Traon
Empir. Softw. Eng.3
2024 Automated Commit Intelligence by Pre-training
abstract
GitHub commits, which record the code changes with natural language messages for description, play a critical role in software developers’ comprehension of software evolution. Due to their importance in software development, several learning-based works are conducted for GitHub commits, such as commit message generation and security patch identification. However, most existing works focus on customizing specialized neural networks for different tasks. Inspired by the superiority of code pre-trained models, which has confirmed their effectiveness across different downstream tasks, to promote the development of open-source software community, we first collect a large-scale commit benchmark including over 7.99 million commits across 7 programming languages. Based on this benchmark, we present CommitBART, a pre-trained encoder-decoder Transformer model for GitHub commits. The model is pre-trained by three categories (i.e., denoising objectives, cross-modal generation, and contrastive learning) for six pre-training tasks to learn commit fragment representations. Our model is evaluated on one understanding task and three generation tasks for commits. The comprehensive experiments on these tasks demonstrate that CommitBART significantly outperforms previous pre-trained works for code. Further analysis also reveals that each pre-training task enhances the model performance.
Shangqing Liu, Yanzhou Li, Xiaofei Xie, Wei Ma 0014, Guozhu Meng, Yang Liu 0003
ACM Trans. Softw. Eng. Methodol.4
2024 Unveiling Code Pre-Trained Models: Investigating Syntax and Semantics Capacities
abstract
Code models have made significant advancements in code intelligence by encoding knowledge about programming languages. While previous studies have explored the capabilities of these models in learning code syntax, there has been limited investigation on their ability to understand code semantics. Additionally, existing analyses assume that the number of edges between nodes at the abstract syntax tree (AST) is related to syntax distance, and also often require transforming the high-dimensional space of deep learning models to a low-dimensional one, which may introduce inaccuracies. To study how code models represent code syntax and semantics, we conduct a comprehensive analysis of seven code models, including four representative code pre-trained models (CodeBERT, GraphCodeBERT, CodeT5, and UnixCoder) and three large language models (LLMs) (StarCoder, CodeLlama and CodeT5+). We design four probing tasks to assess the models’ capacities in learning both code syntax and semantics. These probing tasks reconstruct code syntax and semantics structures (AST, control dependence graph (CDG), data dependence graph (DDG), and control flow graph (CFG)) in the representation space. These structures are core concepts for code understanding. We also investigate the syntax token role in each token representation and the long dependency between the code tokens. Additionally, we analyze the distribution of attention weights related to code semantic structures. Through extensive analysis, our findings highlight the strengths and limitations of different code models in learning code syntax and semantics. The results demonstrate that these models excel in learning code syntax, successfully capturing the syntax relationships between tokens and the syntax roles of individual tokens. However, their performance in encoding code semantics varies. CodeT5 and CodeBERT demonstrate proficiency in capturing control and data dependencies, whereas UnixCoder shows weaker performance in this aspect. We do not observe LLMs generally performing much better than pre-trained models. The shallow layers of LLMs perform better than their deep layers. The investigation of attention weights reveals that different attention heads play distinct roles in encoding code semantics. Our research findings emphasize the need for further enhancements in code models to better learn code semantics. This study contributes to the understanding of code models’ abilities in syntax and semantics analysis. Our findings provide guidance for future improvements in code models, facilitating their effective application in various code-related tasks.
Wei Ma 0014, Shangqing Liu, Xiaofei Xie, Wenhan Wang, Jie Zhang 0050, Yang Liu 0003
ACM Trans. Softw. Eng. Methodol.1
2023 Towards Understanding Model Quantization for Reliable Deep Neural Network Deployment
abstract
Deep Neural Networks (DNNs) have gained considerable attention in the past decades due to their astounding performance in different applications, such as natural language modeling, self-driving assistance, and source code understanding. With rapid exploration, more and more complex DNN architectures have been proposed along with huge pre-trained model parameters. A common way to use such DNN models in user-friendly devices (e.g., mobile phones) is to perform model compression before deployment. However, recent research has demonstrated that model compression, e.g., model quantization, yields accuracy degradation as well as output disagreements when tested on unseen data. Since the unseen data always include distribution shifts and often appear in the wild, the quality and reliability of models after quantization are not ensured. In this paper, we conduct a comprehensive study to characterize and help users understand the behaviors of quantization models. Our study considers four datasets spanning from image to text, eight DNN architectures including both feed-forward neural networks and recurrent neural networks, and 42 shifted sets with both synthetic and natural distribution shifts. The results reveal that 1) data with distribution shifts lead to more disagreements than without. 2) Quantization-aware training can produce more stable models than standard, adversarial, and Mixup training. 3) Disagreements often have closer top-1 and top-2 output probabilities, and Margin is a better indicator than other uncertainty metrics to distinguish disagreements. 4) Retraining the model with disagreements has limited efficiency in removing disagreements. We release our code and models as a new benchmark for further study of model quantization.
Yuejun Guo 0001, Maxime Cordy, Xiaofei Xie, Wei Ma 0014, Mike Papadakis, Yves Le Traon
CAIN5
2022 GraphCode2Vec: Generic Code Embedding via Lexical and Program Dependence Analyses
abstract
peer reviewed
Wei Ma 0014, Ezekiel O. Soremekun, Jie Zhang 0050, Mike Papadakis, Maxime Cordy, Xiaofei Xie, Yves Le Traon
MSR1
2022 On the use of commit-relevant mutants
abstract
Abstract Applying mutation testing to test subtle program changes, such as program patches or other small-scale code modifications, requires using mutants that capture the delta of the altered behaviours. To address this issue, we introduce the concept of commit-relevant mutants, which are the mutants that interact with the behaviours of the system affected by a particular commit. Therefore, commit-aware mutation testing, is a test assessment metric tailored to a specific commit. By analysing 83 commits from 25 projects involving 2,253,610 mutants in both C and Java, we identify the commit-relevant mutants and explore their relationship with other categories of mutants. Our results show that commit-relevant mutants represent a small subset of all mutants, which differs from the other classes of mutants (subsuming and hard-to-kill), and that the commit-relevant mutation score is weakly correlated with the traditional mutation score (Kendall/Pearson 0.15-0.4). Moreover, commit-aware mutation analysis provides insights about the testing of a commit, which can be more efficient than the classical mutation analysis; in our experiments, by analysing the same number of mutants, commit-aware mutants have better fault-revelation potential (30% higher chances of revealing commit-introducing faults) than traditional mutants. We also illustrate a possible application of commit-aware mutation testing as a metric to evaluate test case prioritisation.
Milos Ojdanic, Wei Ma 0014, Thomas Laurent 0003, Thierry Titcheu Chekam, Anthony Ventresque, Mike Papadakis
Empir. Softw. Eng.2
2021 MuDelta: Delta-Oriented Mutation Testing at Commit Time
abstract
To effectively test program changes using mutation testing, one needs to use mutants that are relevant to the altered program behaviours. In view of this, we introduce MuDelta, an approach that identifies commit-relevant mutants; mutants that affect and are affected by the changed program behaviours. Our approach uses machine learning applied on a combined scheme of graph and vector-based representations of static code features. Our results, from 50 commits in 21 Coreutils programs, demonstrate a strong prediction ability of our approach; yielding 0.80 (ROC) and 0.50 (PR Curve) AUC values with 0.63 and 0.32 precision and recall values. These predictions are significantly higher than random guesses, 0.20 (PR-Curve) AUC, 0.21 and 0.21 precision and recall, and subsequently lead to strong relevant tests that kill 45%more relevant mutants than randomly sampled mutants (either sampled from those residing on the changed component(s) or from the changed lines). Our results also show that MuDelta selects mutants with 27% higher fault revealing ability in fault introducing commits. Taken together, our results corroborate the conclusion that commit-based mutation testing is suitable and promising for evolving software.
Wei Ma 0014, Thierry Titcheu Chekam, Mike Papadakis, Mark Harman
ICSE1
2021 Towards Exploring the Limitations of Active Learning: An Empirical Study
abstract
Deep neural networks (DNNs) are increasingly deployed as integral parts of software systems. However, due to the complex interconnections among hidden layers and massive hyperparameters, DNNs must be trained using a large number of labeled inputs, which calls for extensive human effort for collecting and labeling data. Spontaneously, to alleviate this growing demand, multiple state-of-the-art studies have developed different metrics to select a small yet informative dataset for the model training. These research works have demonstrated that DNN models can achieve competitive performance using a carefully selected small set of data. However, the literature lacks proper investigation of the limitations of data selection metrics, which is crucial to apply them in practice. In this paper, we fill this gap and conduct an extensive empirical study to explore the limits of data selection metrics. Our study involves 15 data selection metrics evaluated over 5 datasets (2 image classification tasks and 3 text classification tasks), 10 DNN architectures, and 20 labeling budgets (ratio of training data being labeled). Our findings reveal that, while data selection metrics are usually effective in producing accurate models, they may induce a loss of model robustness (against adversarial examples) and resilience to compression. Overall, we demonstrate the existence of a trade-off between labeling effort and different model qualities. This paves the way for future research in devising data selection metrics considering multiple quality criteria.
Yuejun Guo 0001, Maxime Cordy, Xiaofei Xie, Wei Ma 0014, Mike Papadakis, Yves Le Traon
ASE5
2021 Test Selection for Deep Learning Systems
abstract
Testing of deep learning models is challenging due to the excessive number and complexity of the computations involved. As a result, test data selection is performed manually and in an ad hoc way. This raises the question of how we can automatically select candidate data to test deep learning models. Recent research has focused on defining metrics to measure the thoroughness of a test suite and to rely on such metrics to guide the generation of new tests. However, the problem of selecting/prioritising test inputs (e.g., to be labelled manually by humans) remains open. In this article, we perform an in-depth empirical comparison of a set of test selection metrics based on the notion of model uncertainty (model confidence on specific inputs). Intuitively, the more uncertain we are about a candidate sample, the more likely it is that this sample triggers a misclassification. Similarly, we hypothesise that the samples for which we are the most uncertain are the most informative and should be used in priority to improve the model by retraining. We evaluate these metrics on five models and three widely used image classification problems involving real and artificial (adversarial) data produced by five generation algorithms. We show that uncertainty-based metrics have a strong ability to identify misclassified inputs, being three times stronger than surprise adequacy and outperforming coverage-related metrics. We also show that these metrics lead to faster improvement in classification accuracy during retraining: up to two times faster than random selection and other state-of-the-art metrics on all models we considered.
Wei Ma 0014, Mike Papadakis, Anestis Tsakmalis, Maxime Cordy, Yves Le Traon
ACM Trans. Softw. Eng. Methodol.1
2020 Commit-Aware Mutation Testing
abstract
In Continuous Integration, developers want to know how well they have tested their changes. Unfortunately, in these cases, the use of mutation testing is suboptimal since mutants affect the entire set of program behaviours and not the changed ones. Thus, the extent to which mutation testing can be used to test committed changes is questionable. To deal with this issue, we define commit-relevant mutants; a set of mutants that affect the changed program behaviours and represent the commit-relevant test requirements. We identify such mutants in a controlled way, and check their relationship with traditional mutation score (score based on the entire set of mutants or on the mutants located on the commits). We conduct experiments in both C and Java, using 83 commits, 2,253,610 mutants from 25 projects. Our findings reveal that there is a relatively weak correlation (Kendall/Pearson 0.15-0.4) between the sought (commit-relevant) and traditional mutation scores, indicating the need for a commit-aware test assessment metric. Our analysis also shows that traditional mutation is far from the envisioned case as it loses approximately 50%-60% of the commit-relevant mutants when analysing 5-25 mutants. More importantly, our results demonstrate that traditional mutation has approximately 30% lower chances of revealing commit-introducing faults than commit-aware mutation testing.
Wei Ma 0014, Thomas Laurent 0003, Milos Ojdanic, Thierry Titcheu Chekam, Anthony Ventresque, Mike Papadakis
ICSME1