Tiancheng Li 0005

dblp:92/250-5 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0004-9695-9220ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Identifying the Failure-Revealing Test Cases in Metamorphic Testing: A Statistical Approach
abstract
Metamorphic testing, thanks to its high failure-detection effectiveness especially in the absence of test oracle, has been widely applied in both the traditional context of software testing and other relevant fields such as fault localization and program repair. Its core element is a set of metamorphic relations, which are the necessary properties of the target algorithm in the form of the relationships among multiple inputs and corresponding expected outputs. When a relation is violated by the outputs of a group of test cases, namely metamorphic group of test cases, that are constructed based on the relation, a failure is said to be revealed. Traditionally, the primary task of software testing is to reveal failures. Therefore, from the perspective of software testing, it may not need to know which test case(s) in the metamorphic group cause the violation and thus the failure. However, such information is definitely helpful for other software engineering activities, such as software debugging. The current literature of metamorphic testing lacks a systematic mechanism of identifying the actual failure-revealing test cases, which hinders its applicability and effectiveness in other relevant fields. In this article, we propose a new technique for the FAILure-revealing Test case Identification in Metamorphic testing, namely FAILTIM. The approach is based on a novel application of statistical methods. More specifically, we leverage and adapt the basic ideas of spectrum-based techniques, which are originally used in fault localization, and propose the utilization of a set of risk formulas to estimate the suspiciousness of each individual test case in metamorphic groups. Failure-revealing test cases are then suggested according to their suspiciousness. A series of experiments have been conducted to evaluate the effectiveness and efficiency of FAILTIM using 9 subject programs and 30 risk formulas. The experimental results showed that the new approach can achieve a high accuracy in identifying the actual failure-revealing test cases in metamorphic testing. Consequently, our study will help boost the applicability and performance of metamorphic testing beyond testing to other software engineering areas. The present work also unfolds a number of research directions for further advancing the theory of metamorphic testing and more broadly, software testing.
Zheng Zheng 0001, Dai-Xu Ren, Huai Liu, Tsong Yueh Chen, Tiancheng Li 0005
ACM Trans. Softw. Eng. Methodol.5
2024 Coverage-guided fuzzing for deep reinforcement learning systems
abstract
While the past decade has witnessed a growing demand for employing deep reinforcement learning (DRL) in various domains to solve real-world problems, the reliability of DRL systems has become more of a concern. In particular, DRL agents are often trained on data from a potentially biased distribution over environmental settings, causing the trained agents to fail in certain cases despite high average-case performance. Hence, it is necessary and urgent to adequately test DRL agents to ensure the reliability of practical DRL systems. However, due to the fundamental difference in the programming paradigm and the development process , traditional software testing methodology cannot be applied directly to DRL systems. Given that, we introduce a novel testing framework for DRL systems, aiming to generate diverse test cases that can drive a DRL system to fail. Specifically, we design, implement and evaluate DRLFuzz, which is a coverage-guided fuzzing (CGF) framework for systematically testing DRL systems. Experimental results demonstrate that DRLFuzz can efficiently discover diverse failures in different DRL systems for various benchmark tasks. Compared with a random search baseline, DRLFuzz can generate 60% more failed cases in general. Additionally, the diversity of failed cases generated by DRLFuzz is increased by 4 . 6 % ∼ 14 . 1 % in terms of mean pairwise distance (MPD). Furthermore, our experiments also indicate that the failed cases generated by DRLFuzz can be utilized to fine-tune the DRL agent to eliminate the failures resulting from inadequate exploration during training and thus improve the reliability of DRL systems.
Xiaohui Wan, Tiancheng Li 0005, Weibin Lin, Zheng Zheng 0001
J. Syst. Softw.2
2021 An Empirical Study on Test Case Prioritization Metrics for Deep Neural Networks
abstract
Deep Neural Networks (DNNs) have been widely applied in safety and security domains. DNN testing is necessary to detect the incorrect behaviors of DNNs and guarantee the reliability of DNNs. Labeling test cases is costly that causes DNN testing a serious efficiency problem, which can be alleviated by just labeling test cases with higher priority rather than labeling them in a messy order. Therefore, test case prioritization for DNNs is extensively studied. This paper studies 11 test case prioritization metrics from the ratio of fault detection, accuracy, and correlation perspectives. We classify them into four categories: surprise adequacy, confidence dispersion, mutation uncertainty, and mutation rate. We perform an empirical study of the metrics on two benchmark datasets and DNN models. Our experimental results demonstrate the metrics based on confidence dispersion outperform others regarding effectiveness and efficiency. Meanwhile, we investigate two impact factors of metrics, including test suite size and mutation.
Beibei Yin, Zheng Zheng 0001, Tiancheng Li 0005
QRS4