Zhanqi Cui

dblp:97/2440 · DBLP profile ↗
← Back
58ranked-venue papers
6as first author
51since 2021 · last 2026
0000-0002-5537-9236ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 3 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Detecting API compatibility issues of android applications based on screen transition graphs
Gaoyi Lin, Zhanqi Cui, Xiang Chen 0005
Empir. Softw. Eng.2
2026 Carbon-Aware Dynamic Task Scheduling in Hierarchical Cloud-Edge Systems for IoT Devices
abstract
With the widespread application of the Internet of Things (IoT), computing tasks on the terminal side have surged. Traditional cloud computing models, constrained by high network latency and overloaded central servers, can no longer effectively meet the dual requirements of real-time responsiveness and energy efficiency. The cloud–edge–device collaborative architecture, by enabling distributed resource scheduling, offers a promising solution to reduce both latency and energy consumption. However, optimizing carbon emissions under dynamic operating conditions remains a pressing and unresolved challenge. This paper proposes a carbon-aware dynamic scheduling framework for cloud–edge–device systems, which accounts for the stochastic nature of task arrivals, heterogeneous computing capabilities, and varying carbon intensity across devices and locations. A multi-layer carbon emission model is developed, and the long-term carbon minimization objective is formulated as a stochastic optimization problem. Using the Lyapunov drift-plus-penalty method, the problem is transformed into a tractable deterministic optimization framework, upon which a Carbon-Efficient Computation Offloading (CECO) algorithm is designed. CECO jointly optimizes local computation frequency, data transmission rate, and edge resource allocation to dynamically balance task queue stability and carbon emission intensity. Theoretical analysis and simulation results validate that the proposed algorithm significantly reduces system-level carbon emissions while maintaining quality of service, demonstrating strong potential for enabling green computing in intelligent distributed environments.
Juncai Gao, Zhuoyue Chen, Zhanqi Cui, Ying Chen 0010, Jiwei Huang
IEEE Internet Things J.4
2026 What developers ask about openai APIs: An empirical study on stack overflow
Xiang Chen 0005, Chaoyang Gao, Xiaolin Ju, Zhanqi Cui
J. Syst. Softw.5
2026 Defect prediction guided greybox fuzz testing
Haochen Jin, Zhanqi Cui, Xiang Chen 0005, Rongcun Wang, Xiulei Liu
J. Syst. Softw.2
2026 Edge Service-Oriented Game-Theoretic Joint Optimization of UAV Deployment and Hybrid-NOMA Task Offloading in MEC Networks
abstract
In light of the frequent disaster events worldwide, the risk of communication network disruptions is increasingly prominent. Unmanned Aerial Vehicle (UAV) with flexibility and maneuverability is considered an effective solution for rapid service provisioning. Meanwhile, Non-Orthogonal Multiple Access (NOMA) enables simultaneous access for multiple User Equipments (UEs) on the same frequency band, enhancing spectral efficiency. This paper investigates a system that integrates NOMA with UAV-assisted Mobile Edge Computing (MEC) to provide efficient computational services. First, we construct a system model of UE task offloading and UAV deployment to reduce the overall service cost. Each UE and UAV strive to minimize their own cost. Then, the problem is modeled as the User Task Offloading and UAV Deployment Game (UTUD Game), and it is theoretically proven that at least one Nash equilibrium strategy exists for the offloading and deployment selection. Further, a decentralized algorithm based on game theory named Distributed Iterative Co-optimization for Offloading and Deployment (DICOD) algorithm is proposed to achieve this strategy. The performance of the algorithm is analyzed theoretically. Finally, we conduct experiments to analyze the upper bound of convergence time and evaluate the algorithm's performance in comparison with some other benchmarks.
Ying Chen 0010, Jinze Shu, Jie Zhao 0041, Zhanqi Cui, Jiwei Huang
IEEE Trans. Serv. Comput.5
2025 LLMMutation: Mutation Testing for Large Language Models
abstract
The performance of Large Language Models (LLMs) heavily depends on the quality of their training data, while their capabilities are primarily evaluated through taskspecific test datasets. This evaluation approach makes the quality of test datasets a critical factor in ensuring LLM reliability. Even high-accuracy LLMs may exhibit weak generalizability and insufficient robustness when they are evaluated using flawed test sets with insufficient sample sizes or limited diversity. In traditional software testing, mutation testing is a well-established technique for assessing test suite effectiveness by measuring the ability to detect artificially injected faults. However, fundamental differences exist between traditional software and Transformerbased LLMs, preventing the direct application of classical mutation testing techniques. To address this, we propose LLMMutation, a dedicated test dataset quality evaluation method for LLMs, which applies multi-level mutation operators to introduce perturbations into the model. Experimental results demonstrate that models perturbed by different mutation operators exhibit significant accuracy degradation, which validate LLMMutation's effectiveness in fault injection. In addition, different test datasets exhibit distinct mutation score distributions under LLMMutation perturbations, which is positively correlated with the quality of the datasets. This confirms that the mutation score can effectively distinguish the quality of datasets used for the same task.
Xinhong Duan, Zhanqi Cui
ICTAI3
2025 TD4ITG: A Test Data Generation Method for Issue Title Generation Models
abstract
In open-source software platforms, users utilize issues to report software bugs or request new features. To improve the quality of issues, researchers have proposed several methods for issue title generation. It is widely recognized that deep learning models often suffer from robustness limitations, as minor input perturbations can lead to incorrect or significantly altered outputs. In this paper, we investigate the robustness of issue title generation models and propose a corresponding test data generation method, TD4ITG. This method leverages large language models in combination with chain-of-thought prompting to automatically generate test data to evaluate robustness. Experimental results demonstrate that both iTAPE and iTiger, two issue title generation models, exhibit robustness problems. Specifically, the test data generated by TD4ITG leads to a performance degradation of 21.73% for iTAPE, reducing its score to 75.00%, and a degradation of 17.34% for iTiger, reducing its score to 45.98%. Compared to MATS, a recently proposed testing method for text summarization models, TD4ITG is more effective in revealing the robustness limitations of the models.
Qifan He, Zhanqi Cui
SMC4
2025 MR-OT: a Metamorphic Testing Method for Object Tracking Models
abstract
To ensure the reliability of DNN models, researchers have proposed various testing methods. However, most existing methods focus on static tasks such as image recognition, neglecting the challenges posed by temporal continuity of input data and environmental changes in dynamic scenarios like object tracking and behavior detection. In this paper, we propose MR-OT, a metamorphic testing method for object tracking models in dynamic scenarios. We design five metamorphic relations to evaluate the robustness of the model under three common scenarios, which include natural weather variations, speed changes, and environmental disturbances. Additionally, leveraging generative adversarial networks (GANs) and other techniques, we generate large-scale, realistic, and temporally consistent test data based on original test data. Experimental results demonstrate that MR-OT achieves greater metamorphic relation violation rate, outperforming DeepTest by 0.06 to 0.49 and MT4MOT by 0.04 to 0.57, validating its effectiveness in uncovering robustness issues in dynamic tracking tasks.
Zhanqi Cui, Qifan He
SMC1
2025 Empowering LLM-based Software Defect Prediction with Chain-of-Thought and In-context Learning
abstract
Software defect prediction aims to identify potential defects by analyzing historical project data, thereby reducing development costs and improving software quality. Existing approaches rely on manually designed features or deep learning models, which struggle to capture complex semantics and contextual dependencies, and often require large volumes of labeled data. To address these limitations, this paper proposes ELASTIC(Empowering LLM-based Software Defect Prediction with Chain-of-Thought and In-context Learning), a novel software defect prediction method. ELASTIC employs MinHash and Locality Sensitive Hashing algorithms to retrieve similar code snippets and constructs enhanced prompts that combine code change context with stepwise reasoning paths. This approach enables effective defect prediction without requiring fine-tuning LLMs. Experimental results on the FunctionSStuBs4J dataset, which contains 21,047 samples, demonstrate that ELASTIC achieves an F1 improvement of 0.9% and a 31.9% increase in Recall compared to the state-of-the-art methods, outperforming baseline models such as COMPDEFECT.
Xinhong Duan, Zhanqi Cui
SMC3
2025 SPRoC: Semantics-Preserving Mutations for Robustness Evaluation of Code Generation Large Language Models
abstract
With the widespread use of large language models (LLMs) in code generation, their capabilities continues enhanced. However, LLMs still exhibit instability when faced with minor input prompt variations, which presents challenges for practical deployment. Existing prompt mutation methods, such as random insertions, deletions, or replacements without understanding prompt semantics and structure, still have some limitations. These methods fail to capture the diverse ways of real users express, limiting their ability to assess code generation robustness of LLMs. To address this, we propose SPRoC (Semantics-Preserving Robustness of Code generation), a method for evaluating LLM robustness through prompt mutation. Using a BERT-based model, SPRoC generates mutated prompts that maintain semantic consistency but offer diverse expressions. These prompts create a new dataset to verify the functionality of LLM-generated code. SPRoC compares the functional correctness of code before and after mutation to assess robustness of LLMs to input variations. We conduct experiments on the HumanEval dataset with several mainstream LLMs, including ChatGPT, DeepSeek, Claude, ERNIE, and Qwen, to evaluate performance under SPRoC mutations. Results show that SPRoC reduces the models’ Pass@k scores with minimal semantic changes, outperforming the baseline Radamsa method. In addition, SPRoC achieves better performance in terms of similarity metrics like BLEU and BERTScore, improving by 12.96% and 1.83%, respectively.
Qiancheng Shi, Qihong Han, Zhanqi Cui
SMC3
2025 Testing Autonomous Driving System via Object-Level Replacement
abstract
With the rapid advancement of autonomous driving technology, ensuring the robustness and reliability of decision-making modules has become a critical challenge for the safety of autonomous driving systems (ADSs). In this paper, we propose a novel method, SOGR (Semantic-Guided Object Replacement), to evaluate the decision consistency of ADSs by constructing highly misleading test images that preserve the semantic integrity of the original driving scenes. SOGR identifies important objects using Grad-CAM and YOLOv8, and replaces them with semantically equivalent objects generated by Stable Diffusion. Experiments conducted on the BDD100K dataset demonstrate that SOGR outperforms the pixel-level perturbation baseline DeepIA, achieving higher misleading rates while maintaining lower perceptual similarity (LPIPS) scores. These results indicate that SOGR can effectively expose model vulnerabilities while maintaining high visual realism, offering a practical and semantically grounded approach for robustness testing in real-world autonomous driving scenarios.
Songcheng Xie, Qifan He, Zhanqi Cui
SMC3
2025 SCoVerLLM: Smart Contract Vulnerability Detection via LLM-Based In-Context and Chain-of-Thought Prompts
abstract
As a key application of the blockchain technology, smart contracts have been adopted in various domains such as finance and the Internet of Things. However, their potential vulnerabilities can lead to significant economic losses, efficient and accurate vulnerability detection methods are essential to guarantee their security. Existing methods mostly rely on predefined rules or classification models, which suffer from high maintenance costs and limited semantic understanding about the smart contracts. To address this issue, this paper proposes SCoVerLLM (Smart Contract Vulnerability Detection via LLM-Based In-Context and Chain-of-Thought Prompts), which is designed to enhance the performance of smart contract vulnerability detection by using LLMs. SCoVerLLM combines prediction information generated by deep learning models with similar contract examples, and leverages In-Context Learning prompts and structured Chain-of-Thought templates to guide LLMs in step-by-step analyzing the logic of contracts for vulnerability detection. Experimental results show that SCoVerLLM outperforms existing four methods, including MANDO and Mythril, in terms of multiple metrics, with improvements of 10.72% to 19.20% in Accuracy, 8.70% to 18.51% in Precision, and 10.09% to 25.08% in F1.
Xiguo Gu, Weili Xu, Zhanqi Cui, Liwei Zheng
SMC4
2025 Towards prompt tuning-based software vulnerability assessment with continual learning
Jiacheng Xue, Xiang Chen 0005, Zhanqi Cui
Comput. Secur.4
2025 An ensemble-based transfer testing method for Large Language Models
abstract
Large Language Models (LLMs) can pose serious risks in real-world applications due to their potential for erroneous behavior, necessitating comprehensive and effective testing of LLMs. To assess the robustness of LLMs, adversarial attacks are typically conducted by constructing adversarial examples. Previous methods often require extensive queries and access to the internal information of the victim model. However, the internal information of most black-box LLMs is not accessible, rendering these testing methods infeasible. In addition, excessive queries to commercial black-box LLMs may incur substantial costs. To address these issues, this paper proposes an E nsemble-based T ransfer T esting method for L arge L anguage M odels (ETTLLM). In contrast to previous adversarial testing methods for LLMs, ETTLLM queries white-box surrogates rather than the victim model, thereby significantly reducing testing costs. Moreover, it enhances the transferability and generalization of adversarial examples across diverse real-world classification tasks. Compared to baselines, ETTLLM significantly reduced the number of queries to the victim model, with an average of 1.6 queries, just 1.2% of the baselines. Furthermore, the textual similarity and modification rate of the adversarial examples generated by ETTLLM differ from the baselines by no more than 1.6%, while achieving 70% of the attack success rate compared to the baselines.
Yuanxin Qiao, Yong Liu 0030, Xiang Chen 0005, Zhanqi Cui
Eng. Appl. Artif. Intell.4
2025 SEOCD: Detecting obsolete code comments by fusing semantic features and expert features
Zhanqi Cui, Shifan Liu, Li Li 0114, Liwei Zheng
Expert Syst. Appl.1
2025 A Change-Level Defect Prediction Approach based on Teacher-Student Network
abstract
Change-level defect prediction, also known as just-in-time (JIT) defect prediction, concentrates on predicting if a specific commit is likely to introduce defects. It effectively alleviates the limitations of traditional file-level defect prediction techniques, such as coarse-grained, hard to trace and poor timeliness. Currently, most change-level defect prediction techniques construct defect prediction models by using either expert features or semantic features. Recent studies have shown that the defect prediction performance can be enhanced by integrating these two types of features. However, obtaining expert features is not an easy task, due to missing historical data in real projects. To address the aforementioned problem, this paper proposes TS-SDP (Teacher–Student based Software Defect Prediction) based on teacher–student network. First, the source code is analyzed to extract expert features and semantic features. Then, a teacher–student network framework is constructed. In this framework, both features are used as inputs to the teacher network and only semantic features are used as inputs to the student network. The student network is enabled to learn about the expert features from the teacher network through the loss function. Finally, the student network is used to differentiate commits that are defect-inducing and those that are not, in the presence of only semantic features. The results of the experiments carried out on a dataset containing 21 different projects show that, when only semantic features are available, the cross-network knowledge dissemination between the teacher and student network makes it possible to predict defects. When compared to the state-of-the-art change-level defect prediction method, JIT-Fine, TS-SDP is 0.130, 0.114, 0.123 and 0.016 greater in [Formula: see text], [Formula: see text], F-measure and [Formula: see text], respectively.
Xinhong Duan, Xiguo Gu, Jiale Zhang 0002, Zhanqi Cui
Int. J. Softw. Eng. Knowl. Eng.4
2025 SCATCom: Code Comment Generation by Fusing Multi-Information
abstract
Several code comment generation approaches based on sequence-to-sequence (Seq2Seq) models have been proposed. Such approaches often extract structure information from abstract syntax trees (ASTs) using a certain serialization method. However, some structural information is inevitably lost while serializing ASTs. Furthermore, existing serialization methods only consider the “type” attribute of the nodes, neglecting the “value” attribute of the nodes. To further improve the performance of code comment generation, we propose a code comment generation approach, called SCATCom, which integrates a more comprehensive set of information from source code and ASTs, encompassing semantic, sequential, syntactic, and hierarchical structure information for code comment generation. Meanwhile, an AST traversal method, called V-POT, is presented, which considers both the “type” and the “value” attributes of the nodes. Experiments were designed and conducted on two commonly used datasets to validate the performance of our approach and the impact of five different serialization ways of ASTs on two code comment generation methods. The BLEU, METEOR, and ROUGE scores for our approach reach 52.6, 34.16, and 63.26 with an improvement of [Formula: see text], [Formula: see text], and [Formula: see text] compared to the baselines. It is evident that V-POT, which retains both the “type” and the “value” attributes, is superior to other methods that use only the “type” attribute.
Rongcun Wang, Xiang Chen 0005, Zhanqi Cui, Shujuan Jiang
Int. J. Softw. Eng. Knowl. Eng.4
2025 XL-HQL: A HQL query generation method via XLNet and column attention
Rongcun Wang, Yiqian Hou, Yuan Tian 0008, Zhanqi Cui, Shujuan Jiang
Inf. Softw. Technol.4
2025 BaSFuzz: Fuzz testing based on difference analysis for seed bytes
Wenwei Lan, Li Li 0114, Zhanqi Cui
J. Syst. Softw.5
2025 Learning never stops: Improving software vulnerability type identification via incremental learning
Jiacheng Xue, Xiang Chen 0005, Zhanqi Cui, Yong Liu 0030
J. Syst. Softw.3
2025 IATT: Interpretation Analysis-based Transferable Test Generation for Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have been widely used in various fields. However, it is essential to perform sufficient testing to detect internal defects before deploying CNNs, especially in security-sensitive scenarios. Generating error-inducing inputs to trigger erroneous behavior is the primary way to detect CNN model defects. However, in practice, when the model under test is a black-box CNN model without accessible internal information, in some scenarios it is still necessary to generate high-quality test inputs within a limited testing budget. In such a new scenario, a potential approach is to generate transferable test inputs by analyzing the internal knowledge of other white-box CNN models similar to the model under test, and then use transferable test inputs to test the black-box CNN model. The main challenge in generating transferable test inputs is how to improve their error-inducing capability for different CNN models without changing the test oracle. We found that different CNN models make predictions based on features of similar important regions in images. Adding targeted perturbations to important regions will generate transferable test inputs with high realism. Therefore, we propose the Interpretable Analysis-based Transferable Test (IATT) Generation method for CNNs, which employs interpretation methods of CNN models to explain and localize important regions in test inputs, using backpropagation optimizer and perturbation mask process to add targeted perturbations to these important regions, thereby generating transferable test inputs. This process is repeated to iteratively optimize the transferability and realism of the test inputs. To verify the effectiveness of IATT, we perform experimental studies on nine deep learning models, including ResNet-50 and Vit-B/16, and commercial computer vision system Google Cloud Vision , and compared our method with four state-of-the-art baseline methods. Experimental results show that transferable test inputs generated by IATT can effectively cause black-box target models to output incorrect results. Compared to existing testing and adversarial attack methods, the average Error-inducing Success Rate (ESR) in different testing scenarios is 18.1%–52.7% greater than the baseline methods. Additionally, the test inputs generated by IATT achieve high ESR while maintaining high realism.
Ruilin Xie, Xiang Chen 0005, Qifan He, Bixin Li, Zhanqi Cui
ACM Trans. Softw. Eng. Methodol.5
2024 Detecting Smart Contract Vulnerabilities based on Fusing Semantic and Syntax Structure Information
abstract
Due to the widespread application and economic value of smart contracts, they have become targets for attackers, leading to significant economic losses from vulnerabilities. Therefore, it is crucial to detect potential vulnerabilities in smart contracts before they are deployed. However, existing machine learning approaches often overlook the type information of nodes and edges, while those based on heterogeneous graphs only utilize the semantic information of smart contracts, neglecting the syntax structure information. This oversight compromises the performance in detecting vulnerabilities. To address these issues, we propose a novel smart contract vulnerability detection approach named HG-Detector(Heterogeneous Graph Detector), which stands for Heterogeneous Graph Detector. This approach integrates semantic and syntax structure information by employing a heterogeneous graph neural network to analyze the source code of smart contracts. It extracts both semantic and syntax structure information and then uses a classifier to detect potential vulnerabilities. Experimental results on a dataset comprising 1269 smart contracts show that, compared to MANDO, HG-Detector has achieved an average increase of 10.06% in Precision, an average increase of 1.61% in Recall, an average increase of 2.29% in the F1, and an average increase of 4.78% in Accuracy across seven types of vulnerabilities
Xiguo Gu, Xinhong Duan, Senlin Ren, Jiale Zhang 0002, Zhanqi Cui
ISPA5
2024 SDA-FirmFuzz: Fuzz Testing IoT Device Firmwares Based on Seed Differential Analysis
abstract
In recent years, as the Internet of Things (IoT) devices have been widely used in many fields.There have been attackers taking advantage of the vulnerabilities that exist in the firmware to take control of the devices.As a result, it is important to ensure that the secure of the firmware.Fuzz testing techniques have been proposed for testing firmwares which significantly improved the efficiency of detecting vulnerabilities.Existing firmware fuzz testing techniques mainly focused on the static analysis of firmware before fuzzing, improving the generality of emulation tools, and increasing emulation throughput to improve efficiency of fuzzing.In general, the quality of test cases (seeds) significantly affects the result of fuzz testing; and high-quality seeds can cover more edges and trigger more crashes.Therefore, this paper proposes a fuzz testing method SDA-FirmFuzz (Fuzz Testing IoT Device Firmware Based on Seed Differential Analysis) for IoT device firmwares base on seed differentiation analysis.By analyzing the difference of the seeds before fuzz testing, SDA-FirmFuzz enables the seeds with higher degree of difference are prioritized to be executed.Firstly, a similarity matrix of seeds is constructed based on the cosine similarities.Secondly, a weight matrix is obtained based on the similarity matrix calculation to obtain the similarity scores of the seeds.Finally, the seeds are reordered based on the similarity scores, after which a new seed queue is used to fuzzing the IoT device firmware.Experiments are carried out on six IoT device firmwares, and the experimental results show that SDA-FirmFuzz is able to cover 1.26 times more edges and trigger 32 more unique crashes than Firm-AFL on average.
Zheng-Wu Wang, Wenwei Lan, Zhanqi Cui
SEKE3
2024 Vulnerability Detection by Sequential Learning of Program Semantics via Graph Attention Networks
abstract
Vulnerability detection is a crucial aspect of protecting software systems from cyber attacks. However, some types of vulnerabilities are difficult to detect and require analyzing the source code from multi-views. To address this, we propose a general and easily extensible framework, SGVD(Sequential Graph Attention Networks for Vulnerability Detection). SGVD consists of a sequential module that uses the GAT to learn the semantic representations of the code and a novel Fused-Prediction module that extracts useful features from the multi-view source code. We evaluated this framework on a dataset that includes two large-scale open-source C projects. The experiments showed that SGVD had a superior performance compared to the existing advanced graph learning vulnerability detection tools Devign and ReGVd,with an average increase of 12.25% in Accuracy, 13.65% in Precision, 12.04% in F1 score, and 9.14% in Recall.
Li Li 0114, Qihong Han, Zhanqi Cui
SMC3
2024 Perturbing and Backtracking Based Textual Adversarial Attack
abstract
In the field of Natural Language Processing (NLP), Language Models (LMs) are widely applied in tasks such as text classification, machine translation, and knowledge reasoning. However, the defects of LMs make them vulnerable to adversarial attacks, resulting in substantial economic losses. Adversarial examples can effectively expose vulnerabilities of LMs and be used for adversarial training to improve the robustness of the models. Existing methods mostly generate adversarial examples by first selecting important tokens and then adding perturbations to them. Such methods require a large number of queries to the victim model, which is not applicable in scenarios where the query budget is limited. To address the imperative demands for more query-efficient adversarial example generation, this paper presents CBAPB, a Classification Boundary Adjacent Perturbation and Back-track based textual adversarial attack method, which initially introduces coarse-grained perturbations at random positions while preserving the original semantics of input examples until they reach the similarity threshold. Subsequently, fine-grained perturbation backtracking is conducted on all successfully misclassified examples to minimize perturbation magnitudes. We conduct multiple experiments on the Yelp Reviews, AG News, and DBpedia datasets by employing BERT as the victim model. Comparative analysis against baselines reveals that CBAPB requires merely 3.2% of the average query times of these baselines, while increasing the attack success rate by 7.6%, with only a slight decrease of 1.5% in textual similarity. Experimental results demonstrate the effectiveness of CBAPB, which is not only a query-efficient method but also with greater attack success rates.
Yuanxin Qiao, Ruilin Xie, Songcheng Xie, Zhanqi Cui
SMC4
2024 Combining Deep Learning and Expert Rules for Smart Contract Vulnerability Detection
abstract
Smart contracts usually hold a large amount of digital assets, which can cause substantial losses if these contracts have vulnerabilities. Thus, it is essential to adequately detect possible vulnerabilities in smart contracts before deployment. There are many types of vulnerabilities in smart contracts, and different detection methods have their own unique advantages, some vulnerabilities may be more suitable for expert rule-based methods, while some vulnerabilities are more suitable for deep learning-based methods. A single detection method usually fails to fully use its ability to detect vulnerabilities. To address the above problems, we propose a composite approach named CDE-VD (Combining Deep Learning and Expert Rules for Smart Contract Vulnerability Detection) to improve the performance of vulnerability detection. The method divides smart contract samples into deep learning-prone sam-ples and expert rule-prone samples by classifying them before detection, and extracts expert rule features to train the smart contract detection method classifier to predict the category of the samples under analysis, then selects the suitable method for detection. The experimental results show that the vulnerability detection performance of CDE-VD outperforms that of single detection methods. Compared with the SOTA method MANDO, CDE-VD achieves average improvements of 3.22%, 2.32%, 9.25%, and 6.54% in terms of the Accuracy, Precision, Recall, and F1-score for five categories of vulnerabilities such as access control and time manipulation, respectively, which indicates that category prediction of the smart contract samples could improve vulnerability detection performance.
Senlin Ren, Xiguo Gu, Liwei Zheng, Zhanqi Cui
SMC5
2024 Issue Title Generation: How Far Can Large Language Models Go?
abstract
In open-source software and platforms, developers utilize issues to record software failures or propose new features. The title of an issue, which is a mandatory field, should accurately describe the core content in a concise way. However, developers often face challenges in crafting high-quality issue titles due to insufficient experience or limited proficiency. As a result, researchers have proposed several methods for automatically generating issue titles, but typically relying on constructing large datasets to train models. Recently, Large Language Models (LLMs) have exhibited exceptional performance across a variety of general tasks, suggesting significant potential for issue title generation. Initial experiments indicate that the direct application of LLMs fails to yield satisfactory results. Therefore, we propose a method named LBITG (LLMs-Based Issue Title Generation). LBITG enhances the effectiveness of LLMs by providing contextual information through four types of prompts, which include example prompt and label prompt. These prompts serve as guidance for LLMs, thereby further improving their performance. Experimental results demonstrate that LBITG can significantly enhance the quality of issue titles generated by LLMs without any training. In the within-project scenario, LBITG achieves a minimum improvement of 111.29% in ROUGE, 104.54% in BLEU, and 188.48% in METEOR compared to iTAPE, and achieves performance comparable to that of the SOTA method iTiger. Moreover, LBITG demonstrates superior performance in the cross-project scenario, which outperforms iTiger by 25.33%, 30.14%, and 27.29% in terms of ROUGE-1, BLEU-1, and METEOR, respectively.
Shifan Liu, Qifan He, Songcheng Xie, Zhanqi Cui
SMC5
2024 TS-FL: Software Fault Localization Based on Teacher-Student Network
abstract
Automated fault localization methods can expedite the process for developers to locate faulty code in complex software systems. Existing fault localization methods improve performance by combining the suspicious scores from different kinds of fault localization methods. Among these, the suspicious scores of mutation-based fault localization methods, commonly referred as mutation features, have been proven to effectively enhance fault localization performance. However, collecting mutation features requires generating a large number of mutants and executing test cases for each mutant, which demands sub-stantial computational resources and time. Additionally, certain code statements lack mutation features because no mutant can be generated for them, which affect the performance of fault localization. To address this, this paper proposes a Teacher and Student network-based Fault Lecalization (TS-FL) method. Firstly, a BiLSTM-based classifier is used to extract the deep semantic features of code statements, and the suspicious scores calculated by spectrum-based and mutation-based fault localization methods are used as the spectrum features and mutation features of the code statements, respectively. Then, a teacher-student network is constructed, and a mutual learning strategy is used to collaboratively train the teacher and student network, enabling the student network to learn the mutation feature information from the teacher network and thereby enhance its fault localization performance. The experimental results on Defects4J show that, without using mutation features, TS-FL can locate 36, 36, and 35 more faulty statements than spectrum-based fault localization methods Ochiai, Tarantula, and DStar, and can locate 8 more faulty statements than deep learning-based fault localization method TRANSFER-FL, in terms of Top-1.
Jiale Zhang 0002, Liwei Zheng, Zhanqi Cui
SMC4
2024 A Fast Crash Reproduction Method for Android Applications Based on Widget Hierarchy Graphs
abstract
To improve the efficiency of fixing bugs, mobile application developers must reproduce bugs reported by testers or users as quickly as possible. In some cases, automated testing tools can help developers reproduce crashes. However, these tools were not designed for reproducing bug reports. They are not efficient at reproducing crashes. To focus testing resources on suspicious widgets, we propose CrPDroid, a fast crash reproduction method for Android applications based on widget hierarchy graphs. First, it builds a widget hierarchy graph by using the project file of the application under test; then, it locates suspicious widgets by analyzing the bug report and the project file of the application under test and calculates the fitness of each widget according to the widget hierarchy graph; finally, it uses the fitness of widgets to guide automated testing to reproduce crashes quickly. To evaluate the effectiveness of CrPDroid, experiments are conducted on real Android application bug reports, and the crash reproduction tool ReCDroid, ReproBot and automated testing tools APE and PUMA are compared with CrPDroid. The experimental results show that CrPDroid successfully reproduces 15 bug reports that cause Android app crashes. In addition, compared with APE, PUMA, ReCDroid and ReproBot, the average time for CrPDroid to reproduce crashes decreased by 76.87%, 81.94%, 95.58% and 76.55%, and the total number of operations on suspicious widgets in the same period of testing time increased by 44.07%, 87.57%, 88.70% and 68.93% on average, respectively.
Zhanqi Cui, Gaoyi Lin, Liwei Zheng
IEEE Internet Things J.1
2024 DPFuzz: A fuzz testing tool based on the guidance of defect prediction
Zhanqi Cui, Haochen Jin, Xiang Chen 0005, Rongcun Wang, Xiulei Liu
Sci. Comput. Program.1
2024 CrossFuzz: Cross-contract fuzzing for smart contract vulnerability detection
abstract
Smart contracts are computer programs that run on a blockchain. As the functions implemented by smart contracts become increasingly complex, the number of cross-contract interactions within them also rises. Consequently, the combinatorial explosion of transaction sequences poses a significant challenge for smart contract security vulnerability detection. Existing static analysis-based methods for detecting cross-contract vulnerabilities suffer from high false-positive rates and cannot generate test cases, while fuzz testing-based methods exhibit low code coverage and may not accurately detect security vulnerabilities. The goal of this paper is to address the above limitations and efficiently detect cross-contract vulnerabilities. To achieve this goal, we present CrossFuzz, a fuzz testing-based method for detecting cross-contract vulnerabilities. First, CrossFuzz generates parameters of constructors by tracing data propagation paths. Then, it collects inter-contract data flow information. Finally, CrossFuzz optimizes mutation strategies for transaction sequences based on inter-contract data flow information to improve the performance of fuzz testing. We implemented CrossFuzz, which is an extension of ConFuzzius, and conducted experiments on a real-world dataset containing 396 smart contracts. The results show that CrossFuzz outperforms xFuzz, a fuzz testing-based tool optimized for cross-contract vulnerability detection, with a 10.58% increase in bytecode coverage. Furthermore, CrossFuzz detects 1.82 times more security vulnerabilities than ConFuzzius. Our method utilizes data flow information to optimize mutation strategies. It significantly improves the efficiency of fuzz testing for detecting cross-contract vulnerabilities.
Huiwen Yang, Xiguo Gu, Xiang Chen 0005, Liwei Zheng, Zhanqi Cui
Sci. Comput. Program.5
2023 Software Fault Localization Based on Combining Information Retrieval and Mutation Analysis
abstract
Information Retrieval-based Bug Localization (IRBL) and Mutation-based Fault Localization (MBFL) are two widely used static and dynamic fault localization techniques, respectively. IRBL takes less time and utilizes more static information of software, while MBFL achieves high accuracy and the results are not easily affected by coincidental correctness test cases. However, the granularity of IRBL is coarse and MBFL consumes a lot of time to generate and execute mutants. In this paper, we propose IRMBFL (Information Retrieval and Mutation Analysis Based Software Fault Localization), a software fault localization technique that combines information retrieval and mutation analysis. First, the suspiciousness of source code files is measured by calculating the text similarity between the bug report and the source code to extract the files which may contain bugs. Then, the extracted files are mutated and tested. Finally, the bug statements are located by analyzing the changes in the execution results of the test cases. The experiments are conducted on the Defects4J dataset and$E_{inspect}{@} n$and EXAM are used as evaluation metrics to evaluate the performance of IRMBFL. The experimental results show that IRMBFL locates 14 and 3 more bug statements than BugLocator and Metallaxis for$E_{inspect}{@}n$when$n=1$. IRMBFL outperforms BugLocator on all projects and outperforms Met-allaxis on 2 out of 6 projects in terms of EXAM. In addition, the average bug localization time overhead of IRMBFL is reduced from 73.87% to 99.78% than Metallaxis.
Liwei Zheng, Li Li 0114, Zhanqi Cui
ATS5
2023 TBCUP: A Transformer-based Code Comments Updating Approach
Shifan Liu, Zhanqi Cui, Xiang Chen 0005, Li Li 0114, Liwei Zheng
COMPSAC2
2023 DeepIA: An Interpretability Analysis based Test Data Generation Method for DNN
abstract
Recently, deep neural networks (DNN) have been widely applied in various fields, such as image classification, even replace humans to make decisions in some specific tasks. However, like traditional software, DNNs inevitably contain defects. If defective DNN models are applied in safety-critical fields, such as autonomous driving and medical diagnosis, it may cause disastrous consequences. Therefore, effective testing methods are urgently needed to improve the reliability of DNNs. The existing DNN testing methods typically generate test data by either globally modifying the original data or taking adversarial approaches. The generated test data typically struggle to simultaneously achieve good performance in both the degree of difference from the original data and the Error-inducing Success Rate (ESR) with respect to the target DNN model. Moreover, the perturbation-based methods are difficult to be understood by humans. To address the above issue, this paper proposes DeepIA, an interpretability analysis based test data generation method for DNN. DeepIA analyzes the interpretability of decision-making behaviors for DNN. According to the interpretability analysis results, the original training data is split into different regions to evaluate their influences on decision-making results of the DNN. After that, the most significant regions of the original test data are transformed to generate new test data. Experimental results show that the interpretability method effectively enhances the misleading ability of DeepIA for the DNN model under test. Compared with DeepTest and DeepSearch, DeepIA can generate test data with minor permutations and greater ESR.
Qifan He, Ruilin Xie, Li Li 0114, Zhanqi Cui
QRS4
2023 Smart Contract Vulnerability Detection Based on Clustering Opcode Instructions
abstract
Smart contracts are programs running on the blockchain.In recent years, due to the continuous occurrence of smart contract security accidents, how to effectively detect vulnerabilities in smart contracts has received extensive attention.Machine learning-based vulnerability detection techniques have the advantage of not requiring expert rules.However, existing approaches have limitations in identifying vulnerabilities caused by version updates of smart contract compilers.In this paper, we propose OC-Detector, a smart contract vulnerabilities detection approach based on opcode instruction clustering.OC-Detector learns the characteristics of opcode instructions to cluster them and replaces opcode instructions belonging to the same cluster with the cluster number.After that, the similarity is calculated against the contract in the vulnerability database to identify vulnerabilities.Experimental results demonstrate that OC-Detector improves the F 1 value of detecting vulnerabilities from 0.04 to 0.40 compared to DC-Hunter, Securify, SmartCheck, and Osiris.Additionally, compared to DC-Hunter, F 1 value is improved by 0.27 when detecting vulnerabilities in smart contracts compiled by different version compilers.
Xiguo Gu, Huiwen Yang, Shifan Liu, Zhanqi Cui
SEKE4
2023 AFL2oop: Loop Coverage Guided Greybox Fuzz Testing
abstract
Fuzz testing automatically generates and executes test cases, to detect more defects by covering more logical and state spaces of the program under test (PUT).However, it becomes more difficult to adequately test the PUT with increasing size and code complexity.Studies have shown that complex code is more likely to contain defects, and the loop is one of the main reasons for increased code complexity.Therefore, it is necessary to thoroughly test the loops, but existing fuzzers cannot focus on the loops of the PUT.To address this issue, we design a loop interval coverage metric to measure the testing adequacy of the loop.Additionally, we propose a greybox fuzz testing approach named AFL 2 oop (AFL for Loop), which uses loop coverage as guidance.First, we analyze the loops of the PUT and expand the bitmap.Then, fuzz testing is guided by loop interval coverage and branch coverage.A prototype tool is implemented based on the proposed method, and experiments are carried out on four real-world software programs, such as LibXml2, LibMing, etc.The results show that AFL 2 oop achieves higher coverage, triggers more crashes, and reproduces defects faster than AFL and FairFuzz.
Haochen Jin, Liwei Zheng, Zhanqi Cui
SEKE3
2023 Widget Hierarchy Graph Guided Crash Reproduction Method for Android Applications (S)
abstract
To improve the efficiency of fixing bugs, mobile application developers must reproduce bugs reported by testers or users as quickly as possible.Automated testing tools can help but are not designed for reproducing bug reports.To improve the efficiency of reproducing crashes, we propose a widget hierarchy graph guided crash reproduction method for Android apps.It builds a widget hierarchy graph, locates suspicious widgets using bug reports and project files, calculates widget fitness, and guides automated testing to reproduce crashes quickly.To evaluate the effectiveness of the proposed method, experiments are conducted on real Android application bug reports and compared with the automated testing tools APE and PUMA.Experimental results show that our method successfully reproduces six bug reports that cause Android app crashes.In addition, compared with APE and PUMA, the average time for our method to reproduce crashes decreased by 51.94% and 71.47%.Index
Gaoyi Lin, Zhanqi Cui
SEKE3
2023 Fuzz Testing Based on Seed Diversity Analysis
abstract
Fuzz testing is a widely used technique to detect software defects and vulnerabilities. Coverage-guided fuzzing aims to improve code coverage by generating offspring test cases through mutation, executing the program under test, and retaining interesting seeds for subsequent mutations using customized genetic algorithms. However, existing fuzzing tools rarely consider the similarity between seeds during mutation. Mutating similar seeds frequently generates similar offspring test cases, which results in similar coverage and reduces the efficiency of fuzz testing. To alleviate the impact of this problem on fuzz testing, this paper proposes a fuzz testing method based on seed diversity analysis, which focuses on the characteristics of seeds and uses byte sequences as a feature to measure the similarity between seeds. It collects seeds that can cover new edges and constructs a shorter seed queue with significant differences based on this feature, which replaces the original seed queue for mutation. Based on the proposed method, we implement the prototype tools AFL-Varied and Neuzz-Varied. Compared with AFL and Neuzz on six projects, the edge coverage and basic block coverage can be increased by 214.57% and 233.33 % at most, respectively.
Wenwei Lan, Zhanqi Cui, Jiaming Zhang 0008, Xiguo Gu
SMC2
2023 Statement-Level Software Bug Localization Based on Information Retrieval and Spectrum
abstract
According to whether the program under test is executed, software bug localization methods can be divided into static bug localization and dynamic bug localization. Among them, Information Retrieval-based Bug Localization (IRBL) and Spectrum-based Fault Localization (SFL) are widely used static and dynamic bug localization methods, respectively. But the localization granularity of IRBL is coarse and the localization accuracy of SFL is easily reduced by the information which is unrelated to the bug. In order to refine the localization granularity of IRBL and improve the localization accuracy of SFL, this paper proposes ISBL (Combine Information Retrieval and Spectrum for Bug Localization), a statement-level software bug localization method based on information retrieval and spectrum. Firstly, the suspicious files are filtered using information retrieval technique, and then the suspicious files are used to reduce spectrum information for statement-level bug localization. To evaluate the performance of ISBL, experiments were conducted on the Defects4J dataset, and MRR and TOP@N were used as metrics for evaluation. As the experimental results show, for MRR, ISBL increased 3.0% and 3.1% compared to Ochiai and DStar, respectively; for TOP@1, ISBL locates 4 more bug statements than Ochiai and DStar.
Wenwei Lan, Zhanqi Cui
SMC4
2023 MOBTAG: Multi-Objective Optimization Based Textual Adversarial Example Generation
abstract
Natural language processing (NLP) models are vulnerable to adversarial examples. Generating high-quality adversarial examples, which expose the vulnerability of NLP models and can be used to evaluate and improve their robustness, deserves further research. Existing techniques of generating adversarial examples in the NLP field are typically based on greedy synonym replacements, which may result in out-of-context and unnatural perturbations, and are easily identifiable by humans. In this paper, we present MOB-TAG, a Multi-objective Optimization based Textual Adversarial Example Generation method, which includes three types of perturbations, and utilizes pre-trained models such as BERT and RoBERTa to generate high-quality adversarial examples. MOBTAG generates fluent and grammatical output through a mask-then-infill procedure, with introducing multi-objective optimization and genetic algorithm to pursue a high attack success rate while maintaining a high level of similarity and readability. Experimental results show that compared with methods such as TextFooler, BERTAttack, and CLARE, MOBTAG improves the attack success rate and the textual similarity by at least 11.8% and 0.09 on average, respectively.
Yuanxin Qiao, Ruilin Xie, Li Li 0114, Qifan He, Zhanqi Cui
SMC5
2023 SICUP: A Comment Updating Approach based on Structural Information
abstract
High quality code comments are of great value for program maintenance. However, during the development process, developers often neglect to update corresponding comments when changing code. In such case, inconsistent comments are introduced which affect the maintainability of the software. In previous work, code changes are usually performed by treating the code as ordinary text and the structural information of the code are ignored. In this paper, we propose an approach named SICUP (Structural Information based Comment UPdater) to provide a new solution for comment updating tasks. SICUP uses the structural information of the code to help updating comments by constructing different sequences of ASTs. Experiments on a popular dataset demonstrates that SICUP outperforms CUP, which is an effective deep learning-based approach in terms of accuracy and recall.
Shifan Liu, Zhanqi Cui, Ruilin Xie
SANER2
2023 On efficient matching of spatiotemporal rules
Zhanqi Cui
Future Gener. Comput. Syst.3
2023 CIDFuzz: Fuzz testing for continuous integration
abstract
Abstract As agile software development and extreme programing have become increasingly popular, continuous integration (CI) has become a widely used collaborative work method. However, it is common to make changes frequently to a project during CI. If existing testing methods are applied to CI directly, it will be difficult to make testing resources focus on changes generated by CI, which results in insufficient testing for changes. To solve this problem, we propose a fuzz testing method for CI. First, differential analysis is performed to determine the change points generated during CI, change points are added to the taint source set, and static analysis is conducted to calculate the distances between each basic block and the taint sources. Then, the project under test is instrumented according to the distances. During fuzz testing, testing resources are allocated based on seed coverage to test the change points effectively. Using the proposed methods, we implement CIDFuzz as a prototype tool, and experiments are conducted on four open‐source projects that use CI. Experimental results show that, compared with AFL and AFLGo, CIDFuzz can reduce the time costs of covering change points up to 39.59% and 41.64%, respectively. Also, CIDFuzz can reduce the time costs of reproducing vulnerabilities up to 34.78% and 25.55%.
Jiaming Zhang 0008, Zhanqi Cui, Xiang Chen 0005, Huiwen Yang, Liwei Zheng, Jianbin Liu
IET Softw.2
2023 OC-Detector: Detecting Smart Contract Vulnerabilities Based on Clustering Opcode Instructions
abstract
Smart contracts are programs running on blockchain. In recent years, due to the persistent occurrence of security-related accidents in smart contracts, the effective detection of vulnerabilities in smart contracts has received extensive attention from researchers and engineers. Machine learning-based vulnerability detection techniques have the advantage that they do not need expert rules for determining vulnerabilities. However, existing approaches cannot identify vulnerabilities when the versions of smart contract compilers are updated. In this paper, we propose OC-Detector (Opcode Clustering Detector), a smart contract vulnerability detection approach based on clustering opcode instructions. OC-Detector learns the characteristics of opcode instructions to cluster them and replaces opcode instructions belonging to the same cluster with the ID of the cluster. After that, the similarity between the contract under analysis and contracts in the vulnerability database is calculated to identify vulnerabilities. The experimental results demonstrate that OC-Detector improves the F1 value of detecting vulnerabilities from 0.04 to 0.40 compared to DC-Hunter, Securify, SmartCheck and Osiris. Additionally, compared to DC-Hunter, the F1 value is improved by 0.27 when detecting vulnerabilities in smart contracts compiled by different versions of compilers.
Xiguo Gu, Liwei Zheng, Huiwen Yang, Shifan Liu, Zhanqi Cui
Int. J. Softw. Eng. Knowl. Eng.5
2023 TitleGen-FL: Quality prediction-based filter for automated issue title generation
Xiang Chen 0005, Xuejiao Chen, Zhanqi Cui, Yun Miao, Jianmin Wang 0015
J. Syst. Softw.4
2023 Two-step multi-view data classification based on dynamic Graph-ELM
Li Li 0114, Qihong Han, Zhanqi Cui
Pattern Recognit. Lett.4
2023 Automated Question Title Reformulation by Mining Modification Logs From Stack Overflow
abstract
In Stack Overflow, developers may not clarify and summarize the critical problems in the question titles due to a lack of domain knowledge or poor writing skills. Previous studies mainly focused on automatically generating the question titles by analyzing the posts’ problem descriptions and code snippets. In this study, we aim to improve title quality from the perspective of question title reformulation and propose a novel approachQETRAmotivated by the findings of our formative study. Specifically, by mining modification logs from Stack Overflow, we first extract title reformulation pairs containing the original title and the reformulated title. Then we resort to multi-task learning by formalizing title reformulation for each programming language as separate but related tasks. Later we adopt a pre-trained model T5 to automatically learn the title reformulation patterns. Automated evaluation and human study both show the competitiveness ofQETRAafter compared with six state-of-the-art baselines. Moreover, our ablation study results also confirm that our studied question title reformulation task is more practical than the direct question title generation task for generating high-quality titles. Finally, we develop a browser plugin based onQETRAto facilitate the developers to perform title reformulation. Our study provides a new perspective for studying the quality of post titles and can further generate high-quality titles.
Xiang Chen 0005, Chunyang Chen 0001, Xiaofei Xie, Zhanqi Cui
IEEE Trans. Software Eng.5
2022 DeltaFuzz: Historical Version Information Guided Fuzz Testing
Jiaming Zhang 0008, Zhanqi Cui, Xiang Chen 0005, Huanhuan Wu, Liwei Zheng, Jian-Bin Liu
J. Comput. Sci. Technol.2
2021 CBFL: Improving Software Fault Localization by Analyzing Statement Complexity
abstract
Software fault localization, which is an important software quality assurance technology, provides the location of the faults in software to improve the efficiency of debugging and repairing. In previous research, software fault localization techniques, such as spectrum-based, mutation-based, and program slicing, have been widely used and achieved good results. However, many statements could have same suspicious values by using these techniques, which will consume large amount of manual effort to confirm and affect the accuracy of fault localization. For example, using Ochiai or DStar to locate 395 faulty versions of 6 projects in Defects4J, nearly 70 % of the faulty versions have more than one suspicious statement are ranked as top tied 1. To address the above problem, this paper proposes a complexity-based fault localization (CBFL) technique to further improve the accuracy of fault localization. Firstly, a set of metrics for measuring the complexity of statements is proposed, and the metrics of each statement in projects are extracted to construct a classification model. Then, the classification model is used to predict the faulty probability of the statements which are ranked as top tied 1 by SBFL, MBFL or other techniques, and these statements are reranked according to the estimated faulty probability to improve the accuracy of fault localization. This paper implements a fault localization tool CDStar based on the CBFL, and conducts experiments on the Defects4J dataset. Comparing with the DStar, the results show that CBFL outperforms DStar in terms of Einspect @1 and EXAM.
Haoren Wang, Haochen Jin, Zhanqi Cui, Rongcun Wang
QRS3
2021 Empirical studies on the impact of filter-based ranking feature selection on security vulnerability prediction
abstract
Abstract Security vulnerability prediction (SVP) can construct models to identify potentially vulnerable program modules via machine learning. Two kinds of features from different points of view are used to measure the extracted modules in previous studies. One kind considers traditional software metrics as features, and the other kind uses text mining to extract term vectors as features. Therefore, gathered SVP data sets often have numerous features and result in the curse of dimensionality. In this article, we mainly investigate the impact of filter‐based ranking feature selection (FRFS) methods on SVP, since other types of feature selection methods have too much computational cost. In empirical studies, we first consider three real‐world large‐scale web applications. Then we consider seven methods from three FRFS categories for FRFS and use a random forest classifier to construct SVP models. Final results show that given the similar code inspection cost, using FRFS can improve the performance of SVP when compared with state‐of‐the‐art baselines. Moreover, we use McNemar's test to perform diversity analysis on identified vulnerable modules by using different FRFS methods, and we are surprised to find that almost all the FRFS methods can identify similar vulnerable modules via diversity analysis.
Xiang Chen 0005, Zhidan Yuan, Zhanqi Cui, Dun Zhang, Xiaolin Ju
IET Softw.3
2021 Revisiting heterogeneous defect prediction methods: How far are we?
Xiang Chen 0005, Yanzhou Mu, Zhanqi Cui, Chao Ni 0001
Inf. Softw. Technol.4
2020 Large-Scale Empirical Studies on Effort-Aware Security Vulnerability Prediction Methods
abstract
Security vulnerability prediction (SVP) can identify potential vulnerable modules in advance and then help developers to allocate most of the test resources to these modules. To evaluate the performance of different SVP methods, we should take the security audit and code inspection into account and then consider effort-aware performance measures (such as ACC and Popt). However, to the best of our knowledge, the effectiveness of different SVP methods has not been thoroughly investigated in terms of effort-aware performance measures. In this article, we consider 48 different SVP methods, of which 36 are supervised methods and 12 are unsupervised methods. For the supervised methods, we consider 34 software-metric-based methods and two text-mining-based methods. For the software-metric-based methods, in addition to a large number of classification methods, we also consider four state-of-the-art methods (i.e., EALR, OneWay, CBS, and MULTI) proposed in recent effort-aware just-in-time defect prediction studies. For text-mining-based methods, we consider the Bag-of-Word model and the term-frequency-inverse-document-frequency model. For the unsupervised methods, all the modules are ranked in the ascendent order based on a specific metric. Since 12 software metrics are considered when measuring extracted modules, there are 12 different unsupervised methods. To the best of our knowledge, over 40 SVP methods have not been considered in previous SVP studies. In our large-scale empirical studies, we use three real open-source web applications written in PHP as benchmark. These three web applications include 3466 modules and 223 vulnerabilities in total. We evaluate these SVP methods both in the within-project SVP scenario and the cross-project SVP scenario. Empirical results show that two unsupervised methods [i.e., lines of code (LOC) and Halstead's volume (HV)] and four recently proposed state-of-the-art supervised methods (i.e., MULTI, OneWay, CBS, and EALR) can achieve better performance than the other methods in terms of effort-aware performance measures. Then, we analyze the reasons why these six methods can achieve better performance. For example, when using 20% of the entire efforts, we find that these six methods always require more modules to be inspected, especially for unsupervised methods LOC and HV. Finally, from the view of practical vulnerability localization, we find that all the unsupervised methods and the OneWay method have high false alarms before finding the first vulnerable module. This may have an impact on developers' confidence and tolerance, and supervised methods (especially MULTI and text-mining-based methods) are preferred.
Xiang Chen 0005, Yingquan Zhao, Zhanqi Cui, Guozhu Meng, Yang Liu 0003
IEEE Trans. Reliab.3
2019 Software defect number prediction: Unsupervised vs supervised methods
Xiang Chen 0005, Dun Zhang, Yingquan Zhao, Zhanqi Cui, Chao Ni 0001
Inf. Softw. Technol.4
2019 DP-Share: Privacy-Preserving Software Defect Prediction Model Sharing Through Differential Privacy
Xiang Chen 0005, Dun Zhang, Zhanqi Cui, Qing Gu 0001, Xiaolin Ju
J. Comput. Sci. Technol.3
2017 Applying Feature Selection to Software Defect Prediction Using Multi-objective Optimization
abstract
Software defect prediction can identify potential defective modules in advance and then provide guidances for software testers to allocate more testing resources on these modules. During the gathering process for defect prediction datasets, if multiple metrics are used to measure the program modules, it will result in curse of dimensionality. Feature selection is one of effective methods to alleviate this problem. However, designing effective feature selection methods is a great challenge. Motivated by the idea of search based software engineering, we formalize this problem as a multi-objective optimization problem, and then propose novel method MOFES. To verify the effectiveness of our proposed method, we choose PROMISE dataset gathered from real projects, and compare MOFES with some classical baseline methods. Final results show that our method has the advantages of selecting less features and achieving better prediction performance in most projects while its computational cost is acceptable.
Xiang Chen 0005, Yuxiang Shen, Zhanqi Cui, Xiaolin Ju
COMPSAC (2)3
2013 Verifying Aspect-Oriented Models against Crosscutting Properties
abstract
Dealing with crosscutting concerns has been a critical problem in software development processes. To facilitate handling crosscutting concerns at design phases, we proposed an aspect-oriented modeling and integration approach with UML activity diagrams. The primary concerns are depicted with UML activity diagrams as primary models, whereas crosscutting concerns are described with aspectual extended activity diagrams as aspect models. Aspect models can be integrated into primary models automatically. The AOM approach can reduce the complexity of design models. However, potential faults that violate desired properties of the software system might still be introduced during the modeling or integration processes. The verification technique is well-known for its ability to assure the correctness of models and uncover design problems before implementation. We propose a framework to verify aspect-oriented UML activity diagrams based on Petri net verification techniques. For verification purpose, we transform the integrated activity diagrams into Petri nets and prove the consistency of the transformation. Then, crosscutting concerns in system requirements are refined to properties in the form of CTL formulas. Finally, the Petri nets are verified against the formalized properties to report whether the aspect-oriented design models satisfies the requirements. Furthermore, we implement a tool named Jasmine-AOV to support the verification process. Case studies are conducted to evaluate the effectiveness of the proposed approach.
Zhanqi Cui, Linzhang Wang, Lei Bu, Xuandong Li
Int. J. Softw. Eng. Knowl. Eng.1
2012 Verifying Aspect-Oriented Activity Diagrams Against Crosscutting Properties with Petri Net Analyzer
Zhanqi Cui, Linzhang Wang, Lei Bu, Xuandong Li
SEKE1
2009 A Case Study for Fault Tolerance Oriented Programming in Multi-core Architecture
abstract
The multi-core architecture brings more and more challenges and means to common software developers. Reliable software system design approaches can give a high confidence that long-running online software systems run correctly. But anyway these approaches will certainly cause the loss of the efficiency. We found that the multi-core architecture is a quite suitable platform to support reliable software system design and can make the cost acceptable because of its advantages of the parallel performance and prevalence. In this paper we make use of the multi-core architecture to support software fault tolerance. This approach will make the integration of software fault tolerance and the multi-core architecture as a common design choice. According to the idea of software fault tolerance, for some key software units in a system we can develop N separate versions of them with equivalent functionalities. Each version is developed independently by an isolated group to prevent identical faults among versions. All implemented versions run separately from same initial conditions and inputs. Outputs of all redundant versions are submitted to a decision module that determines a single result from multiple results as the correct output. In this paper, we give a case study to show that with the multi-core architecture, the redundant versions of a key software unit can run in parallel on different cores to improve the efficiency.
Zhanqi Cui, Xuandong Li
HPCC2