VLDB 2026 Research / reviewers in the wild / expert
Zhiyi Zhang 0004
dblp:84/19-4 · also Zhi-Yi Zhang 0004
· DBLP profile ↗
16ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-6671-4378ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A fine-grained evaluation of mutation operators to boost mutation testing for deep learning systems
Zhiyi Zhang 0004, Yongming Yao, Ziyuan Wang 0001 |
Empir. Softw. Eng. | 1 |
| 2025 | Efficient adaptive test case selection for DNNs robustness enhancement
Zhiyi Zhang 0004, Huanze Meng, Yuchen Ding, Shuxian Chen, Yongming Yao |
J. Syst. Softw. | 1 |
| 2024 | EATS: Efficient Adaptive Test Case Selection for Deep Neural NetworksabstractAs deep neural network (DNN) has made significant advancements across various fields, systematically testing DNN has become increasingly crucial. To uncover potential faults within DNN, a vast number of test cases and their correct labels are required, but the process of labeling is time-consuming and labor-intensive. To alleviate the burden on developers, test case selection techniques for DNN models have been proposed, assisting in the selection of test cases from large datasets that are more likely to reveal model faults. In this study, we introduce an efficient adaptive test case selection method based on the principle of uniform distribution of test cases, named EATS. In addition, we propose a test case optimization method and an image similarity calculation method. The optimization method can save time during the test case selection process, while the image similarity calculation method computes the degree of difference between images based on model uncertainty. EATS leverages the model’s uncertainty to achieve a more uniform distribution of selected test cases, aiming to select test cases that can induce a diversity of model fault predictions, thereby optimizing model performance. We conducted comparative experiments of EATS and other test case selection strategies on four common datasets and their corresponding DNN models. The experimental results show that EATS outperforms other methods in terms of the uniformity of test case distribution, diversity of errors discovered, and model optimization. It also demonstrates excellent time efficiency. Huanze Meng, Zhiyi Zhang 0004, Yuchen Ding, Shuxian Chen, Yongming Yao |
QRS | 2 |
| 2023 | A Fine-Grained Evaluation of Mutation Operators for Deep Learning Systems: A Selective Mutation ApproachabstractThe widespread adoption of deep learning (DL) has made it critical to ensure its reliability. Mutation testing has been employed in DL testing to assess test data quality, but it can be costly of a large number of generated mutants. Cost reduction can be achieved by selecting a sufficient subset of mutation operators. However, it remains unclear to what extent the DL mutation operators contribute to test effectiveness, making it challenging to determine which are useful mutation operators in a selective mutation approach. Zhiyi Zhang 0004, Yongming Yao |
Internetware | 2 |
| 2023 | BDGSE: A Symbolic Execution Technique for High MC/DCabstractModified Condition/Decision Coverage (MC/DC) is a test coverage standard with excellent fault detection capability, which is widely used in testing safety-critical software. Symbolic execution generates test cases automatically to achieve high code coverage. However, since symbolic execution uses short-circuit evaluation to evaluate decisions, it fails to guarantee the consistency of other conditions in a decision required by the MC/DC criterion. As a result, it is incapable of generating MC/DC test cases to test safety-critical systems adequately. To solve this problem, we propose Branch Dependence Guided Symbolic Execution (BDGSE), a symbolic execution technique for high MC/DC. The approach utilizes static analysis to compute the branch dependencies, then guides symbolic execution to selectively explore paths and simplify test cases, and finally generates a small number of test cases to achieve high MC/DC coverage. Our experimental results show that BDGSE can generate high-quality test cases. Although BDGSE generates fewer test cases, it can achieve higher MC/DC coverage than SPF and random methods, and the fault detection capability of test cases generated by BDGSE is comparable to that of test cases generated by SPF. Huangli Cai, Zhiyi Zhang 0004, Yifan Jian, Dan Li 0018 |
QRS | 2 |
| 2023 | BTM: Black-Box Testing for DNN Based on Meta-LearningabstractDeep learning is widely used in security fields like autonomous driving, but testing deep learning models poses challenges due to low generation efficiency and limited error detection. Current white-box test case generation methods rely on neuron coverage, but black-box testing is crucial when model details cannot be accessed. Prior methods also needed more consideration for error diversity under lower time resources. In this paper, we propose a novel approach based on the intuition that two models trained on the same classification task learn similar rules at a coarse-grained level. We employ meta-learning to train universal meta perturbation on a surrogate model, which increases neuron coverage on the surrogate model while inducing misclassifications in the original model (thus enhancing the adequacy of the initial tests). Subsequently, we apply data initialization to discover a wider range of errors faster. To accelerate test case generation and reduce resource consumption, we introduce a set of acceleration techniques based on image prediction probabilities, minimizing the time spent exploring irrelevant regions in images. Finally, we evaluate our approach on a well-known dataset and several well-known models, demonstrating its effectiveness. Zhuangyu Zhang, Zhiyi Zhang 0004, Ziyuan Wang 0001, Fang Chen 0007 |
QRS | 2 |
| 2023 | DeepRank: Test Case Prioritization for Deep Neural NetworksabstractDeep neural networks (DNNs) have been widely used in safety-critical fields such as autonomous driving and medical diagnosis.However, DNNs are easily disturbed to make wrong decisions, which may lead to loss of life or property.Therefore, it is vital to test DNN adequately.In practice, to reveal the incorrect behavior of DNN and improve its robustness, testers usually need massive labeled data to test and optimize DNN.However, labeling test inputs to detect the correctness of DNN predictions is an expensive and time-consuming task that even affects the efficiency of DNN testing.To relieve the labeling-cost problem, we propose DeepRank, a test case prioritization technique based on cross-entropy loss.The key idea of DeepRank is that the higher the loss value of a test case relative to the DNN, the more likely it is to be mispredicted and the more conducive it is to improve the robustness of the DNN through retraining.Therefore, the cross-entropy loss value can be used for test case prioritization.We experimentally validate our approach on two datasets and three DNNs models.The experimental results demonstrate that DeepRank is significantly better than existing test case prioritization methods regarding fault-revealing capability and retraining effectiveness. Zhiyi Zhang 0004, Yifan Jian |
SEKE | 2 |
| 2022 | Random or heuristic? An empirical study on path search strategies for test generation in KLEE
Zhiyi Zhang 0004, Ziyuan Wang 0001, Jiahao Wei, Yuqian Zhou |
J. Syst. Softw. | 1 |
| 2021 | DeepBackground: Metamorphic testing for Deep-Learning-driven image recognition systems accompanied by Background-Relevance
Zhiyi Zhang 0004, Hongjing Guo, Ziyuan Wang 0001, Yuqian Zhou |
Inf. Softw. Technol. | 1 |
| 2020 | Towards Generating Realistic and High Coverage Test Data for Constraint-Based Fault InjectionabstractGenerating faulty data is a key issue in fault injection. The faulty data include not only the ones of extreme values or bad formats, but also the ones which are logically unreasonable. Constraint-based fault injection which negates interface constraints to solve faulty data is effective for logically unreasonable data generation. However, the existing constraint-based approaches just solve brand new data for testing. Such brand new data may easily violate some hidden environment constraints on the test inputs and hence be nonrealistic. Besides, there can be different strategies to negate a constraint in order to derive the constraint-unsatisfied faulty data. What are the possible negation strategies and which strategies are better for high coverage fault injection are still unclear. To these ends, this paper presents a new constraint-based fault injection approach. The approach introduces 10 different strategies for constraint negation and relaxes constraint variables to generate faulty data instead of solving brand new data for fault injection. It can produce faulty data which are closer to the original non-faulty ones and hence likely to be more realistic. We experimentally investigated the effectiveness and cost of the introduced constraint negation strategies. The results provide insights for the application of these strategies in fault injection. Ju Qian, Fusheng Lin, Changjian Li 0004, Zhiyi Zhang 0004 |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2020 | A Revisit of Metrics for Test Case Prioritization ProblemsabstractFor the test case prioritization problems, the average percent of faults detected (APFD) and its variant versions are widely used as metrics to evaluate prioritized test suite’s efficiency of fault detection. By a revisit of metrics for test case prioritization, we observe that APFD is only available for the scenarios where all test suites under evaluation contain the same number of test cases. Such a limitation is often overlooked, and lead to incorrect results when comparing fault detection efficiency of test suites with different sizes. Moreover, APFD cannot precisely illustrate the process of fault detection in the real world. Besides the APFD, most of its variants, including the NAPFD and the APFD[Formula: see text], have similar problems. This paper points out these limitations in detail by analyzing the physical explanation of APFD series metrics formally. In order to eliminate these limitations, we propose a series of improved metrics, including the relative average percent of faults detected (RAPFD) and the relative cost-cognizant weighted average percent of faults detected (RAPFD[Formula: see text]), to evaluate the efficiency of the test suite. Furthermore, for the scenario of parallel testing, a series of metrics including the relative average percent of faults detected in parallel testing ([Formula: see text]-RAPFD) and the relative cost-cognizant weighted average percent of faults detected in parallel testing ([Formula: see text]-RAPFD[Formula: see text]) are proposed too. All the proposed metrics refer to both the speed of fault detection and the constraint of the testing resource. A formal analysis and some examples show that all the proposed metrics provide much more precise illustrations of the fault detection process. Ziyuan Wang 0001, Chunrong Fang, Lin Chen 0015, Zhiyi Zhang 0004 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2019 | How to Effectively Reduce Tens of Millions of Tests: An Industrial Case Study on Adaptive Random TestingabstractRunning and analyzing a large number of tests in an industrial scenario is labor intensive and time consuming. Hence, it is necessary to select a smaller number of tests for cost reduction as well as fault detection. For a type of nonnumeric systems, the linear-order algorithm for adaptive random testing (ART) (LART) technique is proposed by making tests evenly spread in nonnumeric input domains. To further enhance LART in the industrial scenarios where the number of input categories is too large, a new technique called category selection-based ART (CSBART), in which partial categories are selected to calculate tests' distances to guide LART, is proposed in this article. The fault-coverage effectiveness of CSBART is evaluated via an empirical study on two large scale billing systems with tens of millions of test cases, and the results demonstrate the promising performance of the proposed CSBART. We also find that, after category selection, CSBART can outperform a more complex and widespread n-per cluster sampling technique that uses K-means clustering to certain extents. Zhiyi Zhang 0004, Ziyuan Wang 0001, Ju Qian |
IEEE Trans. Reliab. | 1 |
| 2018 | Generating Realistic Logically Unreasonable Faulty Data for Fault InjectionabstractIn fault injection, we can use a logical constraint as an interface description and negate the constraint to derive logically unreasonable faulty data in order to test the dependability of a system. However, the existing constraint-based approaches only use constraint solving to generate brand new data for testing. Because the given constraints are often incomplete, such brand new data may not satisfy all the hidden constraints and hence can be nonrealistic. Besides, there can be many different strategies to negate a constraint in order to derive constraint-unsatisfied faulty data. Which negation strategy is the best choice for high coverage fault injection is still unclear. To these ends, this paper presents a new constraint-based fault injection technique which relaxes the constraint variables instead of solving brand new data for fault injection. With such an approach, the generated data can be more close to the original non-faulty data and hence are likely to be more realistic. We also investigated the effectiveness of different negation strategies on a constraint formula for fault injection. The experimental results indicate that our constraint relaxing approach does produce faulty data closer to the original ones. The results also provide insights for the application of constraint negation strategies in fault injection. Ju Qian, Fusheng Lin, Changjian Li 0004, Zhiyi Zhang 0004, Zhe Chen 0011 |
COMPSAC (2) | 4 |
| 2017 | An empirical study on constraint optimization techniques for test generation
Zhiyi Zhang 0004, Zhenyu Chen 0001, Ruizhi Gao, W. Eric Wong, Baowen Xu |
Sci. China Inf. Sci. | 1 |
| 2015 | EFSM-Based Test Case Generation: Sequence, Data, and OracleabstractModel-based testing has been intensively and extensively studied in the past decades. Extended Finite State Machine (EFSM) is a widely used model of software testing in both academy and industry. This paper provides a survey on EFSM-based test case generation techniques in the last two decades. All techniques in EFSM-based test case generation are mainly classified into three parts: test sequence generation, test data generation, and test oracle construction. The key challenges, such as coverage criterion and feasibility analysis in EFSM-based test case generation are discussed. Finally, we summarize the research work and present several possible research areas in the future. Zhenyu Chen 0001, Zhiyi Zhang 0004, Baowen Xu |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2012 | A New Approach to Evaluate Path Feasibility and Coverage Ratio of EFSM Based on Multi-objective Optimization
Zhenyu Chen 0001, Baowen Xu, Zhiyi Zhang 0004, Wujie Zhou |
SEKE | 4 |