Ziyuan Wang 0001

dblp:21/3543-1 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0002-0494-5285ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Working Towards Toxic datasets for LLM Safeguarding
abstract
Large language models (LLMs) have made remarkable strides, expanding their applications from casual dialogue to a wide spectrum of AI tasks. However, concerns about their reliability persist, particularly regarding their tendency to produce toxic content. In response, researchers have developed various toxicity benchmarks to evaluate LLM safety. Nevertheless, as LLMs scale and attack strategies grow more sophisticated, existing benchmarks fall short in assessing the safety of modern models. Addressing increasingly complex and diverse harmful inputs has become a major challenge. To mitigate these risks and leverage existing high-quality datasets, we propose a novel two-stage framework consisting of augmenting red teaming attack templates and generating a human preference dataset. First, we analyze existing red teaming methods, identifying key strategies that effectively exploit LLM vulnerabilities. Followed by augmenting high-quality red teaming attack templates using in-context learning and task-oriented prompt engineering. Then, we leverage the generated attack templates and existing toxicity datasets to generate diverse, high-quality toxic prompts across five threat scenarios: toxicity, stereotype bias, machine ethics, truthfulness, and privacy, intended to elicit toxic responses. To defend against these attacks, we design a Chain-of-Thought (CoT)-based safety guardrail, structured as toxic prompt, toxic response and safety prompt, to generate human preference responses across scenarios. Consequently, we construct the toxic preference dataset containing 7, 502 instances, each consisting of a toxic prompt, a toxic response and a human preference response. We fine-tune the Llama3-8B model on this dataset to enhance its robustness against red teaming attacks. Extensive experiments on five representative LLMs, including our fine-tuned model, demonstrate the effectiveness of our framework in both attack augmentation and human-aligned response generation.
Liuye Guo, Ziyuan Wang 0001, Tieke He
IJCNN2
2025 A fine-grained evaluation of mutation operators to boost mutation testing for deep learning systems
Zhiyi Zhang 0004, Yongming Yao, Ziyuan Wang 0001
Empir. Softw. Eng.4
2024 Towards Understanding Bugs in Go Programming Language
abstract
Go programming language is a powerful tool for modern software development. Ensuring the correct functionality of language features is paramount for the seamless execution of Go programs. Despite meticulous design and development efforts, bugs are an unavoidable aspect of any software, and Go is no exception. This paper presents a comprehensive empirical study based on an extensive analysis of 51,020 issue reports from Go’s repositories on GitHub, providing a panoramic view of the bug landscape within the Go ecosystem and its core packages. Our investigation reveals that: (1) Bugs are more prevalent in components such as Documentation, compiler/runtime, and builder. Windows OS and WebAssembly (WASM) exhibit a higher number of bugs compared to other system architectures. (2) Compilation and Building, Gocommand are the most affected language features by bugs. (3) Behavioral, compiler and runtime errors, and documentation are the three most common types of bug symptoms. (4) Approximately $35 \%$ of bugs are fixed within a day, and about $80 \%$ of bugs are fixed within a year. Roughly $5 \%$ are labeled as soon or release-blockers, statistics indicates that there’s a moderate positive correlation between bug priority and resolution time. (5) Incorrect code logic, improper condition checks, and documentation errors are the three most prevalent root causes of bugs.
Yaping Feng, Ziyuan Wang 0001
QRS2
2024 An Empirical Study on Bugs in Rust Programming Language
abstract
Rust is a young systems programming language that is type and memory safe. Despite Rust’s design focus on safety and correctness, bugs are inevitable in any software system, and Rust is no exception. By analyzing 10097 bug reports and 9360 related revisions in the Rust language, we found that the bugs are extremely unevenly distributed in components and source files; most of the general language features involved in Rust language bugs are Data types, Expressions and Assignment Statements, Rust-specific language features are mainly related to Traits and Ownership Systems; the main symptom of bugs is Internal Compilation Error (ICE); The bug fix work is not complicated; there is a significant correlation between priority and fix time; Semantic bugs are the most common root cause of bugs. These findings reveal some basic patterns of bugs in the Rust language, which can provide some help to Rust developers and maintainers to improve the quality of the Rust language and provide a better programming experience for users of Rust language.
Sijie Yu, Ziyuan Wang 0001
QRS2
2024 COPS: An improved information retrieval-based bug localization technique using context-aware program simplification
Ziyuan Wang 0001, Zhenyu Chen 0001, Baowen Xu
J. Syst. Softw.2
2024 Assessing Effectiveness of Test Suites: What Do We Know and What Should We Do?
abstract
Background. Software testing is a critical activity for ensuring the quality and reliability of software systems. To evaluate the effectiveness of different test suites, researchers have developed a variety of metrics. Problem. However, comparing these metrics is challenging due to the lack of a standardized evaluation framework including comprehensive factors. As a result, researchers often focus on single factors (e.g., size), which finally leads to different or even contradictory conclusions. After comparing dozens of pieces of work in detail, we have found two main problems most troubling to our community: (1) researchers tend to oversimplify the description of the ground truth they use, and (2) data involving real defects is not suitable for analysis using traditional statistical indicators. Objective. We aim at scrutinizing the whole process of comparing test suites for our community. Method. To hit this aim, we propose a framework ASSENT (ev A luating te S t S uite E ffective N ess me T rics) to guide the follow-up research for evaluating a test suite effectiveness metric. ASSENT consists of three fundamental components: ground truth, benchmark test suites, and agreement indicator. Its functioning is as follows: first, users clarify the ground truth for determining the real order in effectiveness among test suites. Second, users generate a set of benchmark test suites and derive their ground truth order in effectiveness. Third, users use the metric to derive the order in effectiveness for the same test suites. Finally, users calculate the agreement indicator between the two orders derived by two metrics. Result. With ASSENT, we are able to compare the accuracy of different test suite effectiveness metrics. We apply ASSENT to evaluate representative test suite effectiveness metrics, including mutation score and code coverage metrics. Our results show that, based on the real faults, mutation score, and subsuming mutation score are the best metrics to quantify test suite effectiveness. Meanwhile, by using mutants instead of real faults, test effectiveness will be overestimated by more than 20% in values. Conclusion. We recommend that the standardized evaluation framework ASSENT should be used for evaluating and comparing test effectiveness metrics in the future work.
Peng Zhang 0083, Yang Wang 0165, Xutong Liu 0003, Yibiao Yang, Yanhui Li 0001, Lin Chen 0015, Ziyuan Wang 0001, Chang-Ai Sun, Xiao Yu 0008, Yuming Zhou
ACM Trans. Softw. Eng. Methodol.8
2023 Toward Understanding Bugs in Swift Programming Language
abstract
Swift programming language has been widely used in IOS application development and has formed a perfect Apple development ecosystem due to its syntactic simplicity and functionality. However, as a complex programming language, Swift inevitably has problems, which may cause the program to fail to run normally. In this paper, we empirically analyze the ones in the Swift language. We collected 7446 bugs and 2749 revisions and manually analyzed the root causes of 180 bugs. We found that defects in Swift are unevenly distributed in components and source files; the test cases are small in size, and the complexity and workload of defect fixing are not significant; the symptoms of defects manifest themselves in various forms, but mainly in the form of Crash; and the root causes of defects are mostly semantic bugs. The conclusions drawn based on the findings are helpful for the development, testing, maintenance, and application of the Swift programming language.
Qianyue Wu, Sijie Yu, Ziyuan Wang 0001, Yaping Feng
QRS3
2023 BTM: Black-Box Testing for DNN Based on Meta-Learning
abstract
Deep learning is widely used in security fields like autonomous driving, but testing deep learning models poses challenges due to low generation efficiency and limited error detection. Current white-box test case generation methods rely on neuron coverage, but black-box testing is crucial when model details cannot be accessed. Prior methods also needed more consideration for error diversity under lower time resources. In this paper, we propose a novel approach based on the intuition that two models trained on the same classification task learn similar rules at a coarse-grained level. We employ meta-learning to train universal meta perturbation on a surrogate model, which increases neuron coverage on the surrogate model while inducing misclassifications in the original model (thus enhancing the adequacy of the initial tests). Subsequently, we apply data initialization to discover a wider range of errors faster. To accelerate test case generation and reduce resource consumption, we introduce a set of acceleration techniques based on image prediction probabilities, minimizing the time spent exploring irrelevant regions in images. Finally, we evaluate our approach on a well-known dataset and several well-known models, demonstrating its effectiveness.
Zhuangyu Zhang, Zhiyi Zhang 0004, Ziyuan Wang 0001, Fang Chen 0007
QRS3
2023 An empirical study on bugs in JavaScript engines
Ziyuan Wang 0001, Dexin Bu, Sijie Yu, Shanyi Gou, Aiyue Sun
Inf. Softw. Technol.1
2022 Context-Aware Program Simplification to Improve Information Retrieval-Based Bug Localization
abstract
Information Retrieval-based Bug localization (IRBL) techniques have become a hot research topic in bug localization due to their few external dependencies and low execution cost. However, existing IRBL techniques have many challenges regarding localization granularity and applicability. First, existing IRBL techniques have not yet achieved statement-level bug localization. Second, almost all studies are limited to Java-based projects, and the effectiveness of these techniques for other widely used programming languages (e.g., Python) is still unknown. The reason for these deficiencies is that existing IRBL techniques mainly employ conventional NLP techniques to analyze the bug reports and have not yet fully exploited the stack trace attached to the bug reports. To improve IRBL techniques in terms of localization granularity and adaptability, we propose a context-aware program simplification technique—COPS—that is able to localize defective statements in suspicious files by analyzing the stack trace in bug reports, which enables statement-level bug localization for Python-based projects. Experiments using 948 bug reports show that our technique can localize the buggy statements with 102.6% higher Top@10, 56.2% higher MAP@10, and 95.6% higher MRR@10 than the baseline. Compared with the state-of-the-art techniques, COPS can improve 19.1% in MAP@10 and achieve 92% buggy statement coverage with a full scope search. Experimental results show that COPS has higher bug localization effectiveness than existing IRBL techniques; and that COPS achieves the same effectiveness with higher execution efficiency than state-of-the-art statement-level defect techniques.
Ziyuan Wang 0001, Zhenyu Chen 0001, Baowen Xu
QRS2
2022 An Empirical Study on Bugs in PHP
abstract
PHP (Hypertext Preprocessor) is a scripting language that has been widely used in web development. This paper conducts an empirical study on bugs in PHP. By analyzing 35,921 bug reports, 6524 revisions, and root causes of randomly selected 500 bugs, we find that: (1) Among all the 385 versions involved in these bugs, there are the most bugs in PHP 4.0.4, PHP 4.0.6, and PHP 4.0.3; Documentation bugs are mainly distributed in PHP 4.y.z and PHP 5.y.z; Security bugs are distributed primarily in the relatively later normal versions of PHP 5.y.z. (2) Documentation, Compile, and Scripting Engine packages are greatly affected by bugs; 73.71% of documentation bugs affect documentation; PHAR, EXIF, and GD are more affected by security bugs. (3) It may be not difficult to repair most bugs since the number of modified lines of code and files are limited; However, nearly 11% of bugs need more than one year to repair; Compared with documentation bugs, security bugs are more difficult to be repaired; The duration of bugs in PHP 8.y.z is shorter than in other versions. (4) Semantic bugs and documentation bugs are the more common root causes of bugs than others. Besides, among semantic bugs, the “Missing Features” bugs and “Processing” bugs are more than others. These results could indicate some potential problems during the detecting and repairing of PHP’s bugs. These findings reveal some laws of bugs in PHP. It could assist developers of PHP in improving their development quality, assist maintainers of PHP in detecting and repairing bugs more effectively, and suggest users of PHP evade potential risks.
Ziyuan Wang 0001, Dexin Bu, Xingpeng Xuan
Int. J. Softw. Eng. Knowl. Eng.1
2022 Test case recommendation based on balanced distance of test targets
Weisong Sun, Quanjun Zhang, Chunrong Fang, Xingya Wang, Ziyuan Wang 0001
Inf. Softw. Technol.6
2022 Random or heuristic? An empirical study on path search strategies for test generation in KLEE
Zhiyi Zhang 0004, Ziyuan Wang 0001, Jiahao Wei, Yuqian Zhou
J. Syst. Softw.2
2022 Mutant Reduction Evaluation: What is There and What is Missing?
abstract
Background. Mutation testing is a commonly used defect injection technique for evaluating the effectiveness of a test suite. However, it is usually computationally expensive. Therefore, many mutation reduction strategies, which aim to reduce the number of mutants, have been proposed. Problem. It is important to measure the ability of a mutation reduction strategy to maintain test suite effectiveness evaluation. However, existing evaluation indicators are unable to measure the “order-preserving ability”, i.e., to what extent the mutation score order among test suites is maintained before and after mutation reduction. As a result, misleading conclusions can be achieved when using existing indicators to evaluate the reduction effectiveness. Objective. We aim to propose evaluation indicators to measure the “order-preserving ability” of a mutation reduction strategy, which is important but missing in our community. Method. Given a test suite on a Software Under Test (SUT) with a set of original mutants, we leverage the test suite to generate a group of test suites that have a partial order relationship in defect detecting ability. When evaluating a reduction strategy, we first construct two partial order relationships among the generated test suites in terms of mutation score, one with the original mutants and another with the reduced mutants. Then, we measure the extent to which the partial order under the original mutants remains unchanged in the partial order under the reduced mutants. The more partial order is unchanged, the stronger the Order Preservation ( OP ) of the mutation reduction strategy is, and the more effective the reduction strategy is. Furthermore, we propose Effort-aware Relative Order Preservation ( EROP ) to measure how much gain a mutation reduction strategy can provide compared with a random reduction strategy. Result. The experimental results show that OP and EROP are able to efficiently measure the “order-preserving ability” of a mutation reduction strategy. As a result, they have a better ability to distinguish various mutation reduction strategies compared with the existing evaluation indicators. In addition, we find that Subsuming Mutant Selection (SMS) and Clustering Mutant Selection (CMS) are more effective than the other strategies under OP and EROP. Conclusion. We suggest, for the researchers, that OP and EROP should be used to measure the effectiveness of a mutant reduction strategy, and for the practitioners, that SMS and CMS should be given priority in practice.
Peng Zhang 0083, Yang Wang 0165, Xutong Liu 0003, Yanhui Li 0001, Yibiao Yang, Ziyuan Wang 0001, Lin Chen 0015, Yuming Zhou
ACM Trans. Softw. Eng. Methodol.6
2022 An Empirical Study on Bugs in Python Interpreters
abstract
Python is an interpreted programming language that has been widely used in many fields. The successful execution of a Python program depends on both the correctness of Python program and the correctness of Python interpreter. As an infrastructure software, there are many bugs in the Python interpreter. Exploring the bugs in Python interpreters can help developers and maintainers of Python interpreters detect and fix bugs and help users of Python avoid risks. In this article, we conduct an empirical study on the bugs in two mainstream Python interpreters: CPython and PyPy. By analyzing 25 958 fixed bugs, 18 824 revisions, 2 116 test cases, and root causes of randomly sampled 510 bugs, we have summarized the following findings.1)The distribution of bugs in the Python interpreter is so uneven that the vast majority of bugs are distributed in a few components and source files.2)The scales of the testing programs that reveal bugs are small.3)The fixing works seem to be not complicated since the number of modified source files and lines of code are limited; however, most bugs need a long time to be fixed; nearly 15% of the bugs need more than one year to fix.4)The priorities of bugs are independent of their locations, but they significantly correlate with duration of bugs.5)Semantic bugs are the most frequent root causes of bugs, and their proportion exceeds other types of root causes.These results could indicate some potential problems during the detecting and fixing of Python interpreter’s bugs, and provide some assistance to developers and maintainers of Python interpreters, users of Python, as well as researchers in related fields.
Ziyuan Wang 0001, Dexin Bu, Aiyue Sun, Shanyi Gou, Yong Wang 0008, Lin Chen 0015
IEEE Trans. Reliab.1
2021 DeepBackground: Metamorphic testing for Deep-Learning-driven image recognition systems accompanied by Background-Relevance
Zhiyi Zhang 0004, Hongjing Guo, Ziyuan Wang 0001, Yuqian Zhou
Inf. Softw. Technol.4
2020 Prediction Method of Code Review Time Based on Hidden Markov Model
Weifeng Zhang 0001, Zhen Pan, Ziyuan Wang 0001
WISA3
2020 A Revisit of Metrics for Test Case Prioritization Problems
abstract
For the test case prioritization problems, the average percent of faults detected (APFD) and its variant versions are widely used as metrics to evaluate prioritized test suite’s efficiency of fault detection. By a revisit of metrics for test case prioritization, we observe that APFD is only available for the scenarios where all test suites under evaluation contain the same number of test cases. Such a limitation is often overlooked, and lead to incorrect results when comparing fault detection efficiency of test suites with different sizes. Moreover, APFD cannot precisely illustrate the process of fault detection in the real world. Besides the APFD, most of its variants, including the NAPFD and the APFD[Formula: see text], have similar problems. This paper points out these limitations in detail by analyzing the physical explanation of APFD series metrics formally. In order to eliminate these limitations, we propose a series of improved metrics, including the relative average percent of faults detected (RAPFD) and the relative cost-cognizant weighted average percent of faults detected (RAPFD[Formula: see text]), to evaluate the efficiency of the test suite. Furthermore, for the scenario of parallel testing, a series of metrics including the relative average percent of faults detected in parallel testing ([Formula: see text]-RAPFD) and the relative cost-cognizant weighted average percent of faults detected in parallel testing ([Formula: see text]-RAPFD[Formula: see text]) are proposed too. All the proposed metrics refer to both the speed of fault detection and the constraint of the testing resource. A formal analysis and some examples show that all the proposed metrics provide much more precise illustrations of the fault detection process.
Ziyuan Wang 0001, Chunrong Fang, Lin Chen 0015, Zhiyi Zhang 0004
Int. J. Softw. Eng. Knowl. Eng.1
2019 From Data Quality to Model Quality: An Exploratory Study on Deep Learning
abstract
In the field of deep learning, people strive to construct high-quality deep neural networks (DNNs) to improve the accuracy of predicting. As well known, the quality of training data have great impacts on the quality of DNN models, since all the DNN models are obtained by training using these training data. However, there is not any reported systematic study on how the quality of training data affects the quality of DNN model. To study the relationships between data quality and model quality, we mainly consider four aspects of data quality including Skewed Classes, Sample Complexity, Label Quality, and Noisy Data in this paper. We design experiments on MNIST and Cifar-10, and attempt to find out the influences of four aspects on the quality of DNN models. Pearson correlation coefficient and Spearman correlation coefficient are utilized to evaluate such influences. Experimental results show that all the four aspects of data quality have significant impacts on the quality of DNN models. It means that the decrease of data quality in these four aspects will reduce the accuracy of the DNN models.
Tianxing He, Shengcheng Yu, Ziyuan Wang 0001, Jieqiong Li, Zhenyu Chen 0001
Internetware3
2019 Leveraging keyword-guided exploration to build test models for web applications
Xiaofang Qi, Yun-Long Hua, Peng Wang 0004, Ziyuan Wang 0001
Inf. Softw. Technol.4
2019 How to Effectively Reduce Tens of Millions of Tests: An Industrial Case Study on Adaptive Random Testing
abstract
Running and analyzing a large number of tests in an industrial scenario is labor intensive and time consuming. Hence, it is necessary to select a smaller number of tests for cost reduction as well as fault detection. For a type of nonnumeric systems, the linear-order algorithm for adaptive random testing (ART) (LART) technique is proposed by making tests evenly spread in nonnumeric input domains. To further enhance LART in the industrial scenarios where the number of input categories is too large, a new technique called category selection-based ART (CSBART), in which partial categories are selected to calculate tests' distances to guide LART, is proposed in this article. The fault-coverage effectiveness of CSBART is evaluated via an empirical study on two large scale billing systems with tens of millions of test cases, and the results demonstrate the promising performance of the proposed CSBART. We also find that, after category selection, CSBART can outperform a more complex and widespread n-per cluster sampling technique that uses K-means clustering to certain extents.
Zhiyi Zhang 0004, Ziyuan Wang 0001, Ju Qian
IEEE Trans. Reliab.3
2017 Automated Testing of Web Applications Using Combinatorial Strategies
Xiaofang Qi, Ziyuan Wang 0001, Jun-Qiang Mao, Peng Wang 0004
J. Comput. Sci. Technol.2
2016 An Efficient Algorithm to Identify Minimal Failure-Causing Schemas from Exhaustive Test Suite
abstract
Combinatorial testing is widely used to detect failures caused by interactions among parameters for its efficiency and effectiveness.Fault localization plays an important role in this testing technique.And minimal failure-causing schema is the root cause of failure.In this paper, an efficient algorithm, which identifies minimal failure-causing schemas from existing failed test cases and passed test cases, is proposed to replace the basic algorithm with worse time performance.Time complexity of basic and improved algorithms is calculated and compared.The result shows that the method that utilizes the differences between failed test cases and passed test cases is better than the method that only uses the sub-schemas of those test cases.
Yuanchao Qi, Chiya Xu, Tieke He, Ziyuan Wang 0001
SEKE5
2016 Empirical analysis of network measures for predicting high severity software faults
Lin Chen 0015, Wanwangying Ma, Yuming Zhou, Lei Xu 0003, Ziyuan Wang 0001, Zhifei Chen, Baowen Xu
Sci. China Inf. Sci.5
2013 Generating Partial Covering Array for Locating Faulty Interactions in Combinatorial Testing
Ziyuan Wang 0001, Wujie Zhou, Weifeng Zhang 0001, Baowen Xu
SEKE1
2011 Cost-Cognizant Combinatorial Test Case Prioritization
abstract
Combinatorial testing has been widely used in practice. People usually assume all test cases in combinatorial test suite will run completely. However, in many scenarios where combinatorial testing is needed, for example the regression testing, the entire combinatorial test suite is not run completely as a result of test resource constraints. To improve the efficiency of testing, combinatorial test case prioritization technique is required. For the scenario of regression testing, this paper proposes a new cost-cognizant combinatorial test case prioritization technique, which takes both combination weights and test costs into account. Here we propose a series of metrics with physical meaning, which assess the combinatorial coverage efficiency of test suite, to guide the prioritization of combinatorial test cases. And two heuristic test case prioritization algorithms, which are based on total and additional techniques respectively, are utilized in our technique. Simulation experimental results illustrate some properties and advantages of proposed technique.
Ziyuan Wang 0001, Lin Chen 0015, Baowen Xu
Int. J. Softw. Eng. Knowl. Eng.1
2010 Cost-Effective Combinatorial Test Case Prioritization for Varying Combination Weights
Ziyuan Wang 0001, Baowen Xu, Lin Chen 0015, Zhenyu Chen 0001
SEKE1
2009 A New Mutation Analysis Method for Testing Java Exception Handling
abstract
Java exception mechanism can effectively free a program from abnormal exits and help developers locate faults with the exception tracing stacks. It is necessary to verify whether the exception handling constructs are arranged appropriately. Some approaches have been developed to evaluate the test sets and improve the quality of them, so that they can raise more number of exceptions in programs. Mutation analysis is a practical method to evaluate the quality of test sets. This paper presents some new mutation operators for Java exception handling constructs. Moreover, equivalent mutants can be identified by our approach. A case study illustrates the effectiveness and characteristic features of these mutation operators.
Changbin Ji, Zhenyu Chen 0001, Baowen Xu, Ziyuan Wang 0001
COMPSAC (2)4
2006 A New Heuristic for Test Suite Generation for Pair-wise Testing
Changhai Nie, Baowen Xu, Ziyuan Wang 0001
SEKE4