VLDB 2026 Research / reviewers in the wild / expert
Weifeng Sun 0004
dblp:88/6657-4
· DBLP profile ↗
29ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0001-6013-1369ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 26 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intention Chain-of-Thought Prompting with Dynamic Routing for Code GenerationabstractLarge language models (LLMs) exhibit strong generative capabilities and have shown great potential in code generation. Existing chain-of-thought (CoT) prompting methods enhance model reasoning by eliciting intermediate steps, but suffer from two major limitations: First, their uniform application tends to induce overthinking on simple tasks. Second, they lack intention abstraction in code generation, such as explicitly modeling core algorithmic design and efficiency, leading models to focus on surface-level structures while neglecting the global problem objective. Inspired by the cognitive economy principle of engaging structured reasoning only when necessary to conserve cognitive resources, we propose RoutingGen, a novel difficulty-aware routing framework that dynamically adapts prompting strategies for code generation. For simple tasks, it adopts few-shot prompting; for more complex ones, it invokes a structured reasoning strategy, termed Intention Chain-of-Thought (ICoT), which we introduce to guide the model in capturing task intention, such as the core algorithmic logic and its time complexity. Experiments across three models and six standard code generation benchmarks show that RoutingGen achieves state-of-the-art performance in most settings, while reducing total token usage by 46.37% on average across settings. Furthermore, ICoT outperforms six existing prompting baselines on challenging benchmarks. Li Huang 0006, Shaoxiong Zhan, Weifeng Sun 0004, Zhongxin Liu 0002, Meng Yan 0001 |
AAAI | 4 |
| 2026 | Exploring and improving knowledge distillation for pre-trained code models
Weifeng Sun 0004, Ruifeng Wu, Meng Yan 0001 |
Empir. Softw. Eng. | 1 |
| 2026 | On-the-Fly Generation-Quality Enhancement of Deep Code Models via Model CollaborationabstractThe growing prominence of deep code models in automating software engineering tasks is undeniable. However, their deployment encounters significant challenges in on-the-fly performance enhancement , which refers to dynamically improving the performance of deep code models during real-time execution. Conventional techniques, such as retraining or fine-tuning, are effective in controlled pre-deployment scenarios but fall short when adapting to on-the-fly adjustments post-deployment. CodeDenoise, a notable on-the-fly performance enhancement technology, leverages uncertainty-based methods to identify misclassified inputs and applies an input modification strategy to rectify classification errors. While effective for classification tasks, this approach is inapplicable to generative tasks due to two key challenges: ❶ Uncertainty-based methods are unsuitable for identifying challenging inputs , especially in generative tasks with diverse and open-ended outputs. Challenging inputs refers to a class of inputs where, due to the inherent complexity of the task or insufficient context in the input samples, the model struggles to generate high-quality outputs. ❷ Input modification strategies cannot be applied to generative tasks, as modifying the input can unpredictably affect the entire sequence of generated outputs. These limitations highlight the need for novel techniques that can enhance the generation quality of deep code models in real-time. To bridge this gap, we propose CodEn , a framework designed to enhance the generation quality of deployed deep code models through model collaboration and real-time output repair. CodEn employs an ensemble learning approach, integrating multiple generic output quality assessment metrics to identify challenging inputs . By combining these diverse metrics, CodEn overcomes the limitations of uncertainty-based methods, making it effective across various generative tasks. Additionally, we introduce an elaborate on-the-fly repair method for the outputs of challenging inputs , leveraging a Large Language Model (LLM) and a novel dual-prompt strategy. This strategy utilizes both generation and selection-based prompts to provide potential fixes and employs an adaptive mechanism to select the optimal output. Our experiments, conducted on 12 deep code models across three pre-trained code models, three popular code-related generation tasks, and four datasets, demonstrate the effectiveness of CodEn . For example, in the assertion generation task, CodEn enhances the Semantic Accuracy Match (SAM) of baseline models with improvements ranging from 12.14% to 21.65%. In the bug fixing task, CodEn achieves exact match gains ranging from 17.51% to 30.64% on TFix dataset. For the code summarization task, CodEn significantly boosts performance across key metrics: BLEU scores improved by 5.72%–11.79%, ROUGE-L by 4.41%–7.70%, METEOR by 7.51%–12.29%, and CIDEr by 8.09%–15.80%. Besides, we conduct experiments of CodEn on different open source LLMs and demonstrate that CodEn can still achieve significant improvements. Weifeng Sun 0004, Naiqi Huang, Meng Yan 0001, Zhongxin Liu 0002, Yan Lei 0005, David Lo 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2026 | Steer Your Model: Secure Code Generation With Contrastive DecodingabstractLarge Language Models (LLMs) specialized in code have demonstrated impressive capabilities in various programming tasks such as code generation. However, these models often generate vulnerable code due to inherent flaws in training datasets derived from large-scale, unfiltered open-source repositories. Existing methods like SVEN (prefix tuning) and CoSec (supervised co-decoding) attempt to address these risks but face challenges with transferability or inflexible security constraints. To mitigate these issues, we propose SCoDE, a two-stage approach for secure and functionally correct code generation. After an initial functional tuning phase, we integrate a plug-andplay security steering matrix at the model’s output embedding layer. This matrix can be transferred across models without modifying their original weights. During inference, we introduce a novel contrastive decoding mechanism that adaptively balances the base model’s functional logits with positive and negative security steering signals. Extensive experiments on 60 security scenarios and two standard benchmarks (HumanEval, MBPP) using StarCoder, Qwen2.5-Coder, and CodeLlama demonstrate that SCoDE enhances security while maintaining functional correctness. On average, SCoDE improves security by 28.09% over the original models, 12.07% over CoSec, and 4.82% over SVEN. For functional correctness, it achieves average gains of 55.03% on HumanEval and 41.81% on MBPP over the original models. Li Huang 0006, Meng Yan 0001, Weifeng Sun 0004, Zhongxin Liu 0002, Hongyu Zhang 0002, David Lo 0001 |
IEEE Trans. Software Eng. | 4 |
| 2026 | Cost-Effective Adversarial Attacks Against Code LLM With Model AttentionabstractCode LLMs (CLLMs) are vulnerable to adversarial attacks, where semantically identical code mutations mislead models into incorrect predictions. To address this, adversarial training has been proposed, retraining models with adversarial examples generated by attack methods. Among various attack approaches, black-box methods have attracted increasing attention due to their flexibility and applicability. However, existing black-box attack methods face two key challenges: 1) vast mutation spaces limit attack efficiency and effectiveness, and 2) resource-intensive model queries constrain scalability. These challenges hinder the practicality of black-box attacks, especially under resource constraints, prompting the critical question:Can we enhance the efficiency of existing attack methods without compromising their effectiveness?To answer this, we conduct an empirical study using Explainable AI (XAI) techniques to investigate differences between adversarial and non-adversarial (failure) examples. After analyzing state-of-the-art attack methods against two CLLMs, we introduce the concept ofmodel attention deviation, which quantifies differences in the model’s focus between unmutated (original) and mutated code. Our findings reveal that adversarial examples exhibit significant attention deviations, with the direction of deviation critically affecting attack success. Building on these insights, we propose ADVSEL, an efficient adversarial attack framework comprising two proxy components: the Attention Proxy Model (APM), which quickly estimates attention deviations to filter unpromising mutations, and the Deviation Direction Proxy Model (DDPM), which assesses whether attention shifts lead toward incorrect predictions. By integrating these proxy models with existing attack methods, ADVSELeffectively prioritizes promising mutations, significantly improving attack efficiency. Experimental evaluations across five CLLMs, four downstream tasks, and three attack methods demonstrate that ADVSEL maintains comparable attack success rates (a slight ASR reduction of 0.62%–0.70%) while significantly reducing model queries (by 34.98%–42.91%) and runtime (by 20.84%–21.45%). Under resource constraints, ADVSEL consistently outperforms baselines, highlighting its practical advantage in cost-effective adversarial evaluation. Weifeng Sun 0004, Naiqi Huang, Meng Yan 0001, Li Huang 0006, Zhongxin Liu 0002, Xiao Liu 0004, David Lo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2025 | Iterative Generation of Adversarial Example for Deep Code ModelsabstractDeep code models are vulnerable to adversarial attacks, making it possible for semantically identical inputs to trigger different responses. Current black-box attack methods typically prioritize the impact of identifiers on the model based on custom importance scores or program context and incrementally replace identifiers to generate adversarial examples. However, these methods often fail to fully leverage feedback from failed attacks to guide subsequent attacks, resulting in problems such as local optima bias and efficiency dilemmas. In this paper, we introduce ITGen, a novel black-box adversarial example generation method that iteratively utilizes feedback from failed attacks to refine the generation process. It employs a bitvectorbased representation of code variants to mitigate local optima bias. By integrating these bit vectors with feedback from failed attacks, ITGen uses an enhanced Bayesian optimization framework to efficiently predict the most promising code variants, significantly reducing the search space and thus addressing the efficiency dilemma. We conducted experiments on a total of nine deep code models for both understanding and generation tasks, demonstrating ITGen's effectiveness and efficiency, as well as its ability to enhance model robustness through adversarial finetuning. For example, on average, ITGen improves the attack success rate by 47.98 % and 69.70 % over the state-of-the-art techniques (i.e., ALERT and BeamAttack), respectively. Li Huang 0006, Weifeng Sun 0004, Meng Yan 0001 |
ICSE | 2 |
| 2025 | Tab: template-aware bug report title generation via two-phase fine-tuned models
Xiao Liu 0004, Yinkang Xu, Weifeng Sun 0004, Naiqi Huang, Dan Yang 0001, Meng Yan 0001 |
Autom. Softw. Eng. | 3 |
| 2025 | Neuron Semantic-Guided Test Generation for Deep Neural Networks FuzzingabstractIn recent years, significant progress has been made in testing methods for deep neural networks (DNNs) to ensure their correctness and robustness. Coverage-guided criteria, such as neuron-wise, layer-wise, and path-/trace-wise, have been proposed for DNN fuzzing. However, existing coverage-based criteria encounter performance bottlenecks for several reasons: ❶ Testing Adequacy : Partial neural coverage criteria have been observed to achieve full coverage using only a small number of test inputs. In this case, increasing the number of test inputs does not consistently improve the quality of models. ❷ Interpretability : The current coverage criteria lack interpretability. Consequently, testers are unable to identify and understand which incorrect attributes or patterns of the model are triggered by the test inputs. This lack of interpretability hampers the subsequent debugging and fixing process. Therefore, there is an urgent need for a novel fuzzing criterion that offers improved testing adequacy, better interpretability, and more effective failure detection capabilities for DNNs. To alleviate these limitations, we propose NSGen, an approach for DNN fuzzing that utilizes neuron semantics as guidance during test generation. NSGen identifies critical neurons, translates their high-level semantic features into natural language descriptions, and then assembles them into human-readable DNN decision paths (representing the internal decision of the DNN). With these decision paths, we can generate more fault-revealing test inputs by quantifying the similarity between original test inputs and mutated test inputs for fuzzing. We evaluate NSGen on popular DNN models (VGG16_BN, ResNet50, and MobileNet_v2) using CIFAR10, CIFAR100, Oxford 102 Flower, and ImageNet datasets. Compared to 12 existing coverage-guided fuzzing criteria, NSGen outperforms all baselines, increasing the number of triggered faults by 21.4% to 61.2% compared to the state-of-the-art coverage-guided fuzzing criterion. This demonstrates NSGen's effectiveness in generating fault-revealing test inputs through guided input mutation, highlighting its potential to enhance DNN testing and interpretability. Li Huang 0006, Weifeng Sun 0004, Meng Yan 0001, Zhongxin Liu 0002, Yan Lei 0005, David Lo 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | Exploring Automated Assertion Generation via Large Language ModelsabstractUnit testing aims to validate the correctness of software system units and has become an essential practice in software development and maintenance. However, it is incredibly time-consuming and labor-intensive for testing experts to write unit test cases manually, including test inputs (i.e., prefixes) and test oracles (i.e., assertions). Very recently, some techniques have been proposed to apply Large Language Models (LLMs) to generate unit assertions and have proven the potential in reducing manual testing efforts. However, there has been no systematic comparison of the effectiveness of these LLMs, and their pros and cons remain unexplored. To bridge this gap, we perform the first extensive study on applying various LLMs to automated assertion generation. The experimental results on two independent datasets show that studied LLMs outperform six state-of-the-art techniques with a prediction accuracy of 51.82%–58.71% and 38.72%–48.19%. The improvements achieve 29.60% and 12.47% on average. Besides, as a representative LLM, CodeT5 consistently outperforms all studied LLMs and all baselines on both datasets, with an average improvement of 13.85% and 26.64%, respectively. We also explore the performance of generated assertions in detecting real-world bugs, and find LLMs are able to detect 32 bugs from Defects4J on average, with an improvement of 52.38% against the most recent approach EditAS . Inspired by the findings, we construct a simplistic retrieval-and-repair-enhanced LLM-based approach by transforming the assertion generation problem into a program repair task for retrieved similar assertions. Surprisingly, such a simplistic approach can further improve the prediction accuracy of LLMs by 9.40% on average, leading to new records on both datasets. Besides, we provide additional discussions from different aspects (e.g., the impact of assertion types and test lengths) to illustrate the capacity and limitations of LLM-based approaches. Finally, we further pinpoint various practical guidelines (e.g., the improvement of multiple candidate assertions) for advanced LLM-based assertion generation in the near future. Overall, our work underscores the promising future of adopting off-the-shelf LLMs to generate accurate and meaningful assertions in real-world test cases and reduce the manual efforts of unit testing experts in practical scenarios. Quanjun Zhang, Weifeng Sun 0004, Chunrong Fang, Meng Yan 0001, Zhenyu Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | DeepVec: State-Vector Aware Test Case Selection for Enhancing Recurrent Neural NetworkabstractDeep Neural Networks (DNN) have realized significant achievements across various application domains. There is no doubt that testing and enhancing a pre-trained DNN that has been deployed in an application scenario is crucial, because it can reduce the failures of the DNN. DNN-driven software testing and enhancement require large amounts of labeled data. The high cost and inefficiency caused by the large volume of data of manual labeling, and the time consumption of testing all cases in real scenarios are unacceptable. Therefore, test case selection technologies are proposed to reduce the time cost by selecting and only labeling representative test cases without compromising testing performance. Test case selection based on neuron coverage (NC) or uncertainty metrics has achieved significant success in Convolutional Neural Networks (CNN) testing. However, it is challenging to transfer these methods to Recurrent Neural Networks (RNN), which excel at text tasks, due to the mismatch in model output formats and the reliance on image-specific characteristics. What’s more, balancing the execution cost and performance of the algorithm is also indispensable.In this paper, we propose a state-vector aware test case selection method for RNN models, namely DeepVec, which reduces the cost of data labeling and saves computing resources and balances the execution cost and performance. DeepVec selects data using uncertainty metric based on the norm of the output vector at each time step (i.e., state-vector), and similarity metric based on the direction angle of the state-vector. Because test cases with smaller state-vector norms often possess greater information entropy and similar changes of state-vector direction angle indicate similar RNN internal states. These metrics can be calculated with just a single inference, which gives it strong bug detection and model improvement capabilities. We evaluate DeepVec on five popular datasets, containing images and texts as well as commonly used 3 RNN classification models, and compare it with NC-based, uncertainty-based, and other black-box methods. Experimental results demonstrate that DeepVec achieves an average relative improvement of 12.5%-118.22% over baseline methods in selecting fault-revealing test cases with time costs reduced to only 1% to 1‱. At the same time, we find that the absolute accuracy improvement after retraining outperforms baseline methods by 0.29%-24.01% when selecting 15% data to retrain. Zhonghao Jiang, Meng Yan 0001, Li Huang 0006, Weifeng Sun 0004, Chao Liu 0014, David Lo 0001 |
IEEE Trans. Software Eng. | 4 |
| 2025 | Retrieval-Augmented Fine-Tuning for Improving Retrieve-and-Edit Based Assertion GenerationabstractUnit Testing is crucial in software development and maintenance, aiming to verify that the implemented functionality is consistent with the expected functionality. A unit test is composed of two parts: a test prefix, which drives the unit under test to a particular state, and a test assertion, which determines what the expected behavior is under that state. To reduce the effort of conducting unit tests manually, Yu et al. proposed an integrated approach (integrationfor short), combining information retrieval with a deep learning-based approach to generate assertions for test prefixes, and obtained promising results. In our previous work, we found that the overall performance ofintegrationis mainly due to its success in retrieving assertions. Moreover,integrationis limited to specific types of edit operations and struggles to understand the semantic differences between the retrieved focal-test (focal-testincludes a test prefix and a unit under test) and the input focal-test. Based on these insights, we then proposed a retrieve-and-edit approach namedEditAS to learn the assertion edit patterns to improve the effectiveness of assertion generation in our prior study. Despite being promising, we find that the effectiveness ofEditAS can be further improved. Our analysis shows that: ① The editing ability ofEditAS still has ample room for improvement. Its performance degrades as the edit distance between the retrieval assertion and ground truth increases. Specifically, the average accuracy ofEditAS is 12.38% when the edit distance is greater than 5. ②EditAS lacks a fine-grained semantic understanding of both the retrieved focal-test and the input focal-test themselves, which leads to many inaccurate token modifications. In particular, an average of 25.57% of the incorrectly generated assertions that need to be modified are not modified, and an average of 6.45% of the assertions that match the ground truth are still modified. Thanks to pre-trained models employing pre-training paradigms on large-scale data, they tend to have good semantic comprehension and code generation abilities. In light of this, we proposeEditAS2, which improves retrieval-and-edit based assertion generation through retrieval-augmented fine-tuning. Specifically,EditAS2first retrieves a similar focal-test from a predefined corpus and treats its assertion as a prototype. Then,EditAS2uses a pre-trained model, CodeT5, to learn the semantics of the input and similar focal-tests as well as assertion editing patterns to automatically edit the prototype. We first evaluate theEditAS2for its inference performance on two large-scale datasets, and the experimental results show thatEditAS2outperforms state-of-the-art assertion generation methods and pre-trained models, with average performance improvements of 15.93%-129.19% and 11.01%-68.88% in accuracy and CodeBLEU, respectively. We also evaluate the performance ofEditAS2in detecting real-world bugs from Defects4J. The experimental results indicate thatEditAS2achieves the best bug detection performance among all the methods. Weifeng Sun 0004, Meng Yan 0001, Xiaohong Zhang 0002, Hongyu Zhang 0002 |
IEEE Trans. Software Eng. | 2 |
| 2024 | DFEPT: Data Flow Embedding for Enhancing Pre-Trained Model Based Vulnerability DetectionabstractSoftware vulnerabilities represent one of the most pressing threats to computing systems. Identifying vulnerabilities in source code is crucial for protecting user privacy and reducing economic losses. Traditional static analysis tools rely on experts with knowledge in security to manually build rules for operation, a process that requires substantial time and manpower costs and also faces challenges in adapting to new vulnerabilities. The emergence of pre-trained code language models has provided a new solution for automated vulnerability detection. However, code pre-training models are typically based on token-level large-scale pre-training, which hampers their ability to effectively capture the structural and dependency relationships among code segments. In the context of software vulnerabilities, certain types of vulnerabilities are related to the dependency relationships within the code. Consequently, identifying and analyzing these vulnerability samples presents a significant challenge for pre-trained models. Zhonghao Jiang, Weifeng Sun 0004, Tao Wen 0012, Haibo Hu 0002, Meng Yan 0001 |
Internetware | 2 |
| 2024 | Toward Cost-Effective Adaptive Random Testing: An Approximate Nearest Neighbor ApproachabstractAdaptive Random Testing(ART) enhances the testing effectiveness (including fault-detection capability) ofRandom Testing(RT) by increasing the diversity of the random test cases throughout the input domain. Many ART algorithms have been investigated such asFixed-Size-Candidate-Set ART(FSCS) andRestricted Random Testing(RRT), and have been widely used in many practical applications. Despite its popularity, ART suffers from the problem of high computational costs during test-case generation, especially as the number of test cases increases. Although several strategies have been proposed to enhance the ART testing efficiency, such as theforgetting strategyand thek-dimensional tree strategy, these algorithms still face some challenges, including: (1) Although these algorithms can reduce the computation time, their execution costs are still very high, especially when the number of test cases is large; and (2) To achieve low computational costs, they may sacrifice some fault-detection capability. In this paper, we propose an approach based onApproximate Nearest Neighbors(ANNs), calledLocality-Sensitive Hashing ART(LSH-ART). When calculating distances among different test inputs, LSH-ART identifies the approximate (not necessarily exact) nearest neighbors for candidates in an efficient way. LSH-ART attempts to balance ART testing effectiveness and efficiency. Rubing Huang, Chenhui Cui, Junlong Lian, Dave Towey, Weifeng Sun 0004, Haibo Chen 0005 |
IEEE Trans. Software Eng. | 5 |
| 2024 | Method-Level Test-to-Code Traceability Link Construction by Semantic Correlation LearningabstractTest-to-code traceability links (TCTLs) establish links between test artifacts and code artifacts. These links enable developers and testers to quickly identify the specific pieces of code tested by particular test cases, thus facilitating more efficient debugging, regression testing, and maintenance activities. Various approaches, based on distinct concepts, have been proposed to establish method-level TCTLs, specifically linking unit tests to corresponding focal methods. Static methods, such as naming-convention-based methods, use heuristic- and similarity-based strategies. However, such methods face the following challenges: ① Developers, driven by specific scenarios and development requirements, may deviate from naming conventions, leading to TCTL identification failures. ② Static methods often overlook the rich semantics embedded within tests, leading to erroneous associations between tests and semantically unrelated code fragments. Although dynamic methods achieve promising results, they require the project to be compilable and the tests to be executable, limiting their usability. This limitation is significant for downstream tasks requiring massive test-code pairs, as not all projects can meet these requirements. To tackle the abovementioned limitations, we propose a novel static method-level TCTL approach, namedTestLinker. For the first challenge of existing static approaches,TestLinkerintroduces a two-phase TCTL framework to accommodate different project types in a triage manner. As for the second challenge, we employ thesemantic correlation learning, which learns and establishes the semantic correlations between tests and focal methods based on Pre-trained Code Models (PCMs).TestLinkerfurther establishes mapping rules to accurately link the recommended function name to the concrete production function declaration. Empirical evaluation on a meticulously labeled dataset reveals thatTestLinkersignificantly outperforms traditional static techniques, showing average F1-score improvements ranging from 73.48% to 202.00%. Moreover, compared to state-of-the-art dynamic methods,TestLinker, which only leverages static information, demonstrates comparable or even better performance, with an average F1-score increase of 37.40%. Weifeng Sun 0004, Zhenting Guo, Meng Yan 0001, Zhongxin Liu 0002, Yan Lei 0005, Hongyu Zhang 0002 |
IEEE Trans. Software Eng. | 1 |
| 2023 | Revisiting and Improving Retrieval-Augmented Deep Assertion GenerationabstractUnit testing validates the correctness of the unit under test and has become an essential activity in software development process. A unit test consists of a test prefix that drives the unit under test into a particular state, and a test oracle (e.g., assertion), which specifies the behavior in that state. To reduce manual efforts in conducting unit testing, Yu et al. proposed an integrated approach (integration for short), combining information retrieval with a deep learning-based approach, to generate assertions for a unit test. Despite being promising, there is still a knowledge gap as to why or where integration works or does not work. In this paper, we describe an in-depth analysis of the effectiveness of integration. Our analysis shows that: ① The overall performance of integration is mainly due to its success in retrieving assertions. ② integration struggles to understand the semantic differences between the retrieved focal-test (focal-test includes a test prefix and a unit under test) and the input focal-test, resulting in many tokens being incorrectly modified; ③ integration is limited to specific types of edit operations (i.e., replacement) and cannot handle token addition or deletion. To improve the effectiveness of assertion generation, this paper proposes a novel retrieve-and-edit approach named EDITAS. Specifically, Editas first retrieves a similar focal-test from a pre-defined corpus and treats its assertion as a prototype. Then, Editas reuses the information in the prototype and edits the prototype automatically. Editas is more generalizable than integration because it can ❶ comprehensively understand the semantic differences between input and similar focal-tests; ❷ apply appropriate assertion edit patterns with greater flexibility; and ❸ generate more diverse edit actions than just replacement operations. We conduct experiments on two large-scale datasets and the experimental results demonstrate that Editas outperforms the state-of-the-art approaches, with an average improvement of 10.00%-87.48% and 3.30%-42.65% in accuracy and BLEU score, respectively. Weifeng Sun 0004, Meng Yan 0001, Yan Lei 0005, Hongyu Zhang 0002 |
ASE | 1 |
| 2023 | Just-In-Time Method Name Updating With Heuristics and Neural ModelabstractEnsuring the quality and conciseness of method names is pivotal for the readability and maintainability of source code. However, for developers, it often presents challenges, particularly during the course of code evolution. Throughout this process, developers sometimes may neglect to update the method name, resulting in inconsistency which could potentially mislead developers and introduce future bugs. In this paper, we propose the task of “Just-In-Time (JIT) Method Name Updating” which automatically performs method name updates to avoid inconsistent names and fix them before being introduced into code bases. Specifically, we propose an approach that combines heuristic rules and a neural model. The heuristic rule-based component mainly focuses on the single-token changes for our empirical findings that the proportion of single-token modifications is extensive, and often corresponds to code-indicative updates. The neural model-based component is a customized Seq2seq model considering code changes and the new method body’s token type. To evaluate our approach, we conduct extensive experiments on the collected dataset with over 108K method name-body co-change samples from popular Java projects. The results show that our method outperforms the three baselines on all metrics. In particular, our approach achieves a significant improvement in Accuracy and improves method generation baseline by 23.5%. Zhenting Guo, Meng Yan 0001, Zhezhe Chen, Weifeng Sun 0004 |
QRS | 5 |
| 2023 | Extended Abstract of Candidate Test Set Reduction for Adaptive Random Testing: An Overheads Reduction TechniqueabstractThis document1is an extended abstract of a Science of Computer Programming paper, "Candidate Test Set Reduction for Adaptive Random Testing: An Overheads Reduction Technique," presented as a J1C2 (Journal publication first, Conference presentation following) at the 30th IEEE International Conference on Software Analysis, Evolution and Reengineering (Saner 2023).The paper presents a candidate set reduction strategy to enhance the Fixed-Sized-Candidate-Set version of Adaptive Random Testing (FSCS-ART). The proposed method reduces the number of randomly-generated candidate test cases by retaining valuable, unused candidates from previous iterations. As the computational costs associated with a stored/retained candidate are less than the costs associated with a randomly-generating one, the overall computational overheads of FSCS-ART are reduced. The reported experimental studies show that the proposed method has a comparable failure-detection effectiveness to FSCS-ART, but less computational overheads. Rubing Huang, Haibo Chen 0005, Weifeng Sun 0004, Dave Towey |
SANER | 3 |
| 2023 | An Adaptive Partition-Based Approach for Adaptive Random Testing on Real ProgramsabstractAdaptive random testing (ART) is a family of algorithms to enhance random testing (RT) by generating test cases extensively and evenly. For this purpose, many ART algorithms have been proposed, the most well-known and the first approach is the Fixed-Size-Candidate-Set ART (FSCS-ART). In recent years, researchers have also proposed many ART methods to continuously improve the performance of FSCS-ART, but the focus has been more on reducing the time overhead of FSCSART while retaining its failure detection effectiveness as much as possible due to the boundary effect. To alleviate the boundary effect and improve the effectiveness of FSCS-ART, this paper proposes an algorithm AP-FSCS-ART, an Adaptive Partition-based method on top of FSCS-ART. First, AP-FSCS-ART divides the entire input domain into external and internal sub-domains. Then, two different algorithms are adaptively applied to the two sub-domains to find the next test case from the randomly generated candidate test cases. During the selecting process, APFSCS-ART takes into account not only the most recently executed test case of a candidate test case but also its position relative to the input domain. Experiments using the 12 most common real programs and comparisons with other algorithms in this paper show that the AP-FSCS-ART algorithm has significantly better failure detection capability, with improvements from 8.8% to 11.4% compared to three state-of-the-art ART algorithms, including the FSCS-ART, FSCS-ctsr, and NNDC-ART. Yisheng Xia, Weifeng Sun 0004, Meng Yan 0001, Dan Yang 0001 |
SANER | 2 |
| 2023 | A first look at bug report templates on GitHub
Meng Yan 0001, Weifeng Sun 0004, Xiao Liu 0004, Yunsong Wu |
J. Syst. Softw. | 3 |
| 2023 | Revisiting the Identification of the Co-evolution of Production and Test CodeabstractMany software processes advocate that the test code should co-evolve with the production code. Prior work usually studies such co-evolution based on production-test co-evolution samples mined from software repositories. A production-test co-evolution sample refers to a pair of a test code change and a production code change where the test code change triggers or is triggered by the production code change. The quality of the mined samples is critical to the reliability of research conclusions. Existing studies mined production-test co-evolution samples based on the following assumption: if a test class and its associated production class change together in one commit, or a test class changes immediately after the changes of the associated production class within a short time interval, this change pair should be a production-test co-evolution sample . However, the validity of this assumption has never been investigated. To fill this gap, we present an empirical study, investigating the reasons for test code updates occurring after the associated production code changes, and revealing the pervasive existence of noise in the production-test co-evolution samples identified based on the aforementioned assumption by existing works. We define a taxonomy of such noise, including six categories (i.e., adaptive maintenance, perfective maintenance, corrective maintenance, indirectly related production code update, indirectly related test code update, and other reasons). Guided by the empirical findings, we propose CHOSEN (an identifi C ation met H od O f production-te S t co- E volutio N ) based on a two-stage strategy. CHOSEN takes a test code change and its associated production code change as input, aiming to determine whether the production-test change pair is a production-test co-evolution sample. Such identified samples are the basis of or are useful for various downstream tasks. We conduct a series of experiments to evaluate our method. Results show that (1) CHOSEN achieves an AUC of 0.931 and an F1-score of 0.928, significantly outperforming existing identification methods, and (2) CHOSEN can help researchers and practitioners draw more accurate conclusions on studies related to the co-evolution of production and test code. For the task of Just-In-Time (JIT) obsolete test code detection, which can help detect whether a piece of test code should be updated when developers modify the production code, the test set constructed by CHOSEN can help measure the detection method’s performance more accurately, only leading to 0.76% of average error compared with ground truth. In addition, the dataset constructed by CHOSEN can be used to train a better obsolete test code detection model, of which the average improvements on accuracy, precision, recall, and F1-score are 12.00%, 17.35%, 8.75%, and 13.50% respectively. Weifeng Sun 0004, Meng Yan 0001, Zhongxin Liu 0002, Xin Xia 0001, Yan Lei 0005, David Lo 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | VPP-ART: An Efficient Implementation of Fixed-Size-Candidate-Set Adaptive Random Testing Using Vantage Point PartitioningabstractAdaptive random testing(ART) is an enhancement ofrandom testing(RT), and aims to improve the RT failure-detection effectiveness by distributing test cases more evenly in the input domain. Many ART algorithms have been proposed, withfixed-size-candidate-setART (FSCS-ART) being one of the most effective and popular. FSCS-ART ensures high failure-detection effectiveness by selecting as the next test case the candidate farthest from previously executed test cases. Although FSCS-ART has good failure-detection effectiveness, it also faces some challenges, including heavy computational overheads. In this article, we propose an enhanced version of FSCS-ART,vantage point partitioning ART(VPP-ART). VPP-ART addresses the FSCS-ART computational overhead problem using VPP, while maintaining the failure-detection effectiveness. VPP-ART partitions the input domain space using amodified vantage point tree(VP-tree) and finds the approximate nearest executed test cases of a candidate test case in the partitioned subdomains—thereby significantly reducing the time overheads compared with the searches required for FSCS-ART. To enable the FSCS-ART dynamic insertion process, we modify the traditional VP-tree to support dynamic data. The simulation results show that VPP-ART has a much lower time overhead compared to FSCS-ART, but also delivers similar (or better) failure-detection effectiveness, especially in the higher dimensional input domains. According to statistical analyses, VPP-ART can improve on the FSCS-ART failure-detection effectiveness by approximately 50–58%. VPP-ART also compares favorably with theKD-tree-enhanced fixed-size-candidate-set ART(KDFC-ART) algorithms (a series of enhanced ART algorithms based on the KD-tree). Our experiments also show that VPP-ART is more cost-effective than FSCS-ART and KDFC-ART. Rubing Huang, Chenhui Cui, Dave Towey, Weifeng Sun 0004, Junlong Lian |
IEEE Trans. Reliab. | 4 |
| 2023 | Robust Test Selection for Deep Neural NetworksabstractDeep Neural Networks (DNNs) have been widely used in various domains, such as computer vision and software engineering. Although many DNNs have been deployed to assist various tasks in the real world, similar to traditional software, they also suffer from defects that may lead to severe outcomes. DNN testing is one of the most widely used methods to ensure the quality of DNNs. Such method needs rich test inputs with oracle information (expected output) to reveal the incorrect behaviors of a DNN model. However, manually labeling all the collected test inputs is a labor-intensive task, which delays the quality assurance process.Test selectiontackles this problem by carefully selecting a small, more suspicious set of test inputs to label, enabling the failure detection of a DNN model with reduced effort. Researchers have proposed different test selection methods, including neuron-coverage-based and uncertainty-based methods, where the uncertainty-based method is arguably the most popular technique. Unfortunately, existing uncertainty-based selection methods meet the performance bottleneck due to one or several limitations: 1) they ignore noisy data in real scenarios; 2) they wrongly exclude manyfailure-revealing test inputsbut rather include manysuccessful test inputs(referring to those test inputs that are correctly predicted by the model); 3) they ignore the diversity of the selected test set. In this paper, we propose RTS, a Robust Test Selection method for deep neural networks to overcome the limitations mentioned above. First, RTS divides all unlabeled candidate test inputs into noise set, successful set, and suspicious set and assigns different selection prioritization to divided sets, which effectively alleviates the impact of noise and improves the ability to identify suspect test inputs. Subsequently, RTS leverages a probability-tier-matrix-based test metric for prioritizing the test inputs in each divided set (i.e., suspicious, successful, and noise set). As a result, RTS can select more suspicious test inputs within a limited selection size. We evaluate RTS by comparing it with 14 baseline methods under 5 widely-used DNN models and 6 widely-used datasets. The experimental results demonstrate that RTS can significantly outperform all test selection methods in failure detection capability and the test suites selected by RTS have the best model optimization capability. For example, when selecting 2.5% test input, RTS achieves an improvement of 9.37%-176.75% over baseline methods in terms of failure detection. Weifeng Sun 0004, Meng Yan 0001, Zhongxin Liu 0002, David Lo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2022 | Candidate test set reduction for adaptive random testing: An overheads reduction technique
Rubing Huang, Haibo Chen 0005, Weifeng Sun 0004, Dave Towey |
Sci. Comput. Program. | 3 |
| 2022 | A nearest-neighbor divide-and-conquer approach for adaptive random testing
Rubing Huang, Weifeng Sun 0004, Haibo Chen 0005, Chenhui Cui |
Sci. Comput. Program. | 2 |
| 2021 | A Survey on Adaptive Random TestingabstractRandom testing (RT) is a well-studied testing method that has been widely applied to the testing of many applications, including embedded software systems, SQL database systems, and Android applications. Adaptive random testing (ART) aims to enhance RT's failure-detection ability by more evenly spreading the test cases over the input domain. Since its introduction in 2001, there have been many contributions to the development of ART, including various approaches, implementations, assessment and evaluation methods, and applications. This paper provides a comprehensive survey on ART, classifying techniques, summarizing application areas, and analyzing experimental evaluations. This paper also addresses some misconceptions about ART, and identifies open research challenges to be further investigated in the future work. Rubing Huang, Weifeng Sun 0004, Yinyin Xu, Haibo Chen 0005, Dave Towey, Xin Xia 0001 |
IEEE Trans. Software Eng. | 2 |
| 2020 | Poster: Is Euclidean Distance the best Distance Measurement for Adaptive Random Testing?abstractAdaptive random testing (ART) aims at enhancing the testing effectiveness of random testing (RT) by more evenly spreading test cases over the input domain. Many ART methods have been proposed, based on various, different notions. For example, distance-based ART (DART) makes use of the concept of distance to implement ART, attempting to generate new test cases that are far away from previously executed ones. The Euclidean distance has been a popular choice of distance metric, used in DART to evaluate the differences between test cases. However, is the Euclidean distance the most suitable choice for DART? To answer this question, we conducted a series of simulations to investigate the impact that the Euclidean distance, and its many variations, has on the testing effectiveness of DART. The results show that when the dimensionality of the input domain is low, the Euclidean distance may indeed be a good choice. However, when the dimensionality is high, it appears to be less suitable. Rubing Huang, Chenhui Cui, Weifeng Sun 0004, Dave Towey |
ICST | 3 |
| 2020 | Regression test case prioritization by code combinations coverage
Rubing Huang, Quanjun Zhang, Dave Towey, Weifeng Sun 0004, Jinfu Chen 0001 |
J. Syst. Softw. | 4 |
| 2020 | Abstract Test Case Prioritization Using Repeated Small-Strength Level-Combination CoverageabstractAbstract test cases (ATCs) have been widely used in practice, including in combinatorial testing and in software product line testing. When constructing a set of ATCs, due to limited testing resources in practice (e.g., in regression testing), test case prioritization (TCP) has been proposed to improve the testing quality, aiming at ordering test cases to increase the speed with which faults are detected. One intuitive and extensively studied TCP technique for ATCs is λ-wise Level-combination Coverage based Prioritization (λLCP), a static, black-box prioritization technique that only uses the ATC information to guide the prioritization process. A challenge facing λLCP, however, is the necessity for the selection of the fixed prioritization strength λ before testing-testers need to choose an appropriate λ value before testing begins. Choosing higher λ values may improve the testing effectiveness of λLCP (e.g., by finding faults faster), but may reduce the testing efficiency (by incurring additional prioritization costs). Conversely, choosing lower λ values may improve the efficiency, but may also reduce the effectiveness. In this paper, we propose a new family of λLCP techniques, Repeated Small-strength Level-combination Coverage-based Prioritization (RSLCP), that repeatedly achieves the full combination coverage at lower strengths. RSLCP maintains λLCP's advantages of being static and black box, but avoids the challenge of prioritization strength selection. We have performed an empirical study involving five different versions of each of five C programs. Compared with λLCP, and Incremental-strength LCP (ILCP), our results show that RSLCP could provide a good tradeoff between testing effectiveness and efficiency. Our results also show that RSLCP is more effective and efficient than two popular techniques of Similarity-based Prioritization (SP). In addition, the results of empirical studies also show that RSLCP can remain robust over multiple system releases. Rubing Huang, Weifeng Sun 0004, Tsong Yueh Chen, Dave Towey, Jinfu Chen 0001, Weiwen Zong, Yunan Zhou |
IEEE Trans. Reliab. | 2 |
| 2018 | On the Selection of Strength for Fixed-Strength Interaction Coverage Based PrioritizationabstractAbstract test cases are derived by modeling the system under test, and have been widely applied in practice, such as for software product line testing and combinatorial testing. Abstract test case prioritization (ATCP) is used to prioritize abstract test cases and aims at achieving higher rates of fault detection. Many ATCP algorithms have been proposed, using different prioritization criteria and information. One ATCP approach makes use of fixed-strength level-combinations information covered by abstract test cases, and is called fixed-strength interaction coverage based prioritization (FICBP). Before using FICBP, the prioritization strength λ needs to be decided. Previous studies have generally focused on λ values ranging between 1 and 6. However, no study has investigated the appropriateness of such a range, nor how to assign the prioritization strength for FICBP. To answer these questions, this paper reports on an empirical study involving four real-life programs (each of which with six versions). The experimental results indicate that λ should be set approximately equal to a value corresponding to half of the number of parameters, when testing resources are sufficient. Our results also show that when testing resources are limited or insufficient, either small or large λ values are suggested for FICBP. Rubing Huang, Weiwen Zong, Tsong Yueh Chen, Dave Towey, Jinfu Chen 0001, Yunan Zhou, Weifeng Sun 0004 |
COMPSAC (1) | 7 |