EDBT 2026 Demo / reviewers in the wild / expert
Huai Liu
dblp:70/4981
· DBLP profile ↗
66ranked-venue papers
9as first author
36since 2021 · last 2026
0000-0003-3125-4399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 46 · 6 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Metamorphic Non-Functional Testing for Web Services: A Statistical Approach and a Case Study on Requestlatency Relations
Abrar Hossain Chy Toha, Huai Liu, Mengjiao Guo |
COMPSAC | 2 |
| 2026 | Toward Standardized Evaluation of Metamorphic Relations: A Structured Rubric and Human-LLM Comparison
Yifan Zhang 0016, Dave Towey, Matthew Pike, Quang-Hung Luu, Huai Liu, Tsong Yueh Chen |
COMPSAC | 5 |
| 2026 | Analysis of the Relationship between Editing Behaviors and Questions in Stack OverflowabstractStack Overflow is a popular online platform for program on various professional levels to ask and answer questions on a wide range of topics in computer programming. It has attracted lots of research on the analysis of user behaviors, especially those related to the asked questions. In this paper, we propose an approach for analyzing the correlation between the editing behaviors and the questions asked in Stack Overflow. The approach is particularly established on the foundation of the analysis of a series of editing operations that may exist between the questions and answers. The approach is applied to analyze real-life Stack Overflow data, and the experimental results demonstrate that the more edited questions are likely to get more answers. It is also found that the editing of the body part and the editing of the tag part will affect the quality of the final question. Ruobing Li, Yusi Chen, Huai Liu |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2026 | Introduction to the special issue on metamorphic testing
Huai Liu, Aldeida Aleti, Aitor Arrieta |
Inf. Softw. Technol. | 1 |
| 2026 | Advancing LLM-Generated Code Reliability: A Hybrid Approach for Hallucination DetectionabstractThe increasing use of Large Language Models (LLMs) for writing code has raised important concerns about “code hallucinations.” These occur when the generated code looks correct in terms of its structure (syntax) but contains mistakes in its meaning or logic. Such errors can then spread through software, leading to problems and inefficiencies in the final applications. Current research on finding these code hallucinations in LLM output often struggles with inefficiency. It also lacks a good collection of test cases specifically designed to properly evaluate how well different detection methods work. To address these issues, we introduce a new approach that effectively combines static and dynamic analysis techniques for hallucination detection (SDHD). While standard methods often fail to spot code hallucinations, SDHD shows significant improvement in performance across various datasets. For example, when tested on the MBPP, CodeHaluEval, and HalluCode datasets, SDHD achieved an average precision of 0.771, an average recall of 0.783, and an average F1-score of 0.776. These results are not just slightly better, but substantially higher than those of existing methods, clearly demonstrating SDHD’s superior effectiveness in overcoming the limitations of current hallucination detection approaches. Jiayi Dang, Huai Liu, Zhi Jin 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | FailMapper: Automated Generation of Unit Tests Guided by Failure ScenariosabstractThe automation of unit test generation has become a critical task for improving the overall efficiency of software development and testing. Many existing techniques attempt to generate a sufficient number of test cases to achieve high code coverage. However, it has been shown that a high coverage does not necessarily guarantee effective bug discovery. A potential enhancement is to guide the unit test generation based on bug properties. However, this solution is challenged by the large number and diversity of bug types, making it difficult to comprehensively summarize bug properties.We observe that failures, presented as the results of bugs, manifest in a limited number of scenarios. Therefore, instead of bug properties, in this paper, we propose an innovative framework, named FailMapper, which uses failure scenarios to guide the generation of unit tests. We summarize nine failure scenarios and design the corresponding failure-triggering test strategies. This significantly improves the efficacy of generating test cases towards triggering bugs. To systematically explore possible failure scenarios, FailMapper employs the Monte Carlo Tree Search algorithm to search for the faults that may lead to a failure. Experiments demonstrate that, on 50 known bugs in the Defects4J benchmark, FailMapper can detect many more bugs than five typical unit testing approaches, including EvoSuite, Randoop, CoverUp, HITS, and SymPrompt (40 versus at most 12, out of all 50 bugs). Meanwhile, FailMapper detects 12 out of 20 bugs in the GitBug-Java and Bears-benchmark datasets. We reveal 36 potential issues from 2 Apache projects, and 14 of them have been confirmed as bugs, further demonstrating FailMapper’s effectiveness. The experimental results show that our new framework can significantly enhance the overall efficacy of unit testing. Ruiqi Dong, Zehang Deng, Xiaogang Zhu 0001, Xiaoning Du 0001, Huai Liu, Shaohua Wang 0002, Sheng Wen, Yang Xiang 0001 |
ASE | 5 |
| 2025 | DiffFix: Incrementally Fixing AST Diffs via Context and Type InformationabstractThe abstract syntax tree differencing (ASTDiff) technique aims to capture syntactic code changes through comparing the differences between a pair of ASTs of a program, which has been widely used in various program analysis or testing tasks, such as code review, clone detection, and regression testing. A key issue for ASTDiff lies in the accurate mappings between nodes of two ASTs. However, most existing approaches often fail to generate such perfect diffs due to the gap between diverse code changes and unsound node matching heuristics. Our in-depth investigation reveals that most inaccurate mappings are caused by the ignorance of context- and/or type-specific constraints. Accordingly, we propose an AST diff fixing approach DiffFix that leverages both the node’s context and type constraints to iteratively and incrementally fix imperfect diffs. Comprehensive experiments have been conducted to evaluate the effectiveness of DiffFix through its application to fix diffs generated by five state-of-the-art ASTDiff techniques. The experimental results demonstrate that DiffFix can improve the perfect diff rate of these baseline techniques by 5.25% to 51.12% with negligible time overhead. Guofeng Zeng, Chang-Ai Sun, Kai Gao 0008, Huai Liu |
ASE | 4 |
| 2025 | Can Large Language Models Discover Metamorphic Relations? A Large-Scale Empirical StudyabstractSoftware testing is a mainstream approach for software quality assurance. One fundamental challenge for testing is that in many practical situations, it is very difficult to verify the correctness of test results given inputs for Software Under Test (SUT), which is known as the oracle problem. Metamorphic Testing (MT) is a software testing technique that can effectively alleviate the oracle problem. The core component of MT is a set of Metamorphic Relations (MRs), which are basically the necessary properties of SUT, represented in the form of relationship among multiple inputs and their corresponding expected outputs. Different methods have been proposed to support the systematic MR identification. However, most of them still rely heavily on test engineers' understanding of the SUT and involve massive manual work. Although a few preliminary studies have shown LLMs' viability in generating MRs, there does not exist a thorough and in-depth investigation on their capability in MR identification. We are thus motivated to conduct a comprehensive and large-scale empirical study to systematically evaluate the performance of LLMs in identifying appropriate MRs for a wide variety of software systems. This study makes use of 37 SUTs collected from previous MT studies. Prompts are constructed for two LLMs, gpt-3.5-turbo-1106 and gpt-4-1106-preview, to perform the MR identification for each SUT. The empirical results demonstrate that both LLMs can generate a large amount of MR candidates (MRCs). Among them, 29.86% and 43.79% of all MRCs are identified as the MRs valid for the corresponding SUT, respectively. In addition, 24.59% and 38.63% of all MRCs are MRs that had never been identified in previous studies. Our study not only reinforces LLM-based MR identification as a promising research direction for MT, but also provides some practical guidelines for how to further improve LLMs' performance in generating good MRs. Chang-Ai Sun, Huai Liu, Sijin Dong |
SANER | 3 |
| 2025 | MMF: A Lightweight Approach of Multimodel Fusion for Malware DetectionabstractNowadays, the Android system is widely used in mobile devices. The existence of malware in the Android system has posed serious security risks. Therefore, detecting malware has become a main research focus for Android devices. The existing malware detection methods include those based on static analysis, dynamic analysis, and hybrid analysis. The dynamic analysis and hybrid analysis methods require the simulation of malware’s execution in a certain environment, which often incurs high costs. With the aid of contemporary deep learning technology, static method can provide comparably good results without running software. To address these challenges, we propose a novel and efficient multimodel fusion (MMF) malware detection method. MMF innovatively integrates various static features, including application programming interface (API) call characteristics, request permission (RP) features, and bytecode image features. This fusion approach allows MMF to achieve high detection performance without the need for dynamic execution of the software. Compared to existing methods, MMF exhibits a higher accuracy rate of 99.4% and demonstrates superiority over baseline techniques in various metrics. Our comprehensive analysis and experiments confirm MMF’s effectiveness and efficiency in detecting malware, making a significant contribution to the field of Android malware detection. Mengbo Li, Li Li 0029, Huai Liu |
IET Softw. | 4 |
| 2025 | A Reinforcement Learning Based Approach to Partition Testing
Chang-Ai Sun, Ming-Jun Xiao, Hepeng Dai, Huai Liu |
J. Comput. Sci. Technol. | 4 |
| 2025 | TRALSem: A Robust Model for Textual Sentiment AnalysisabstractSentiment analysis has gained widespread applications across various domains due to its versatility and practicality. With the increasing availability of data and advancements in machine learning technologies, its utilization is expected to continue expanding. Deep learning (DL)-based sentiment analysis has demonstrated high accuracy and efficiency in numerous application areas, such as marketing, customer service, politics, healthcare, and finance, thereby highlighting its high potential. Despite recent progress, DL-based sentiment analysis methods still face significant challenges, particularly concerning the robustness of sentiment classification and scoring. To address these issues, our study introduces TRALSem, a novel text-centered sentiment analysis framework designed to tackle the unique challenges of sentiment classification and scoring, with a particular focus on enhancing the overall robustness of the model. We conducted extensive experiments on multilingual script datasets, including Chinese and English scripts, as well as the IMDB and SST datasets. The experimental results show that TRALSem significantly outperforms existing state-of-the-art sentiment analysis methods. In terms of sentiment classification, it achieves remarkable improvements, far surpassing the previous benchmarks. More importantly, TRALSem substantially enhances the model’s robustness, enabling it to maintain stable performance even in the face of complex and noisy data. It also significantly reduces the model’s sensitivity to noise data, effectively filtering out interference and providing more reliable and accurate sentiment analysis results. Moreover, it offers better interpretability to users, making the sentiment analysis process and outcomes more understandable and actionable. Jiayi Dang, Huai Liu, Zhi Jin 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Metamorphic Relation Generation: State of the Art and Research DirectionsabstractMetamorphic testing has become one mainstream technique to address the notorious oracle problem in software testing, thanks to its great successes in revealing real-life bugs in a wide variety of software systems. Metamorphic relations, the core component of metamorphic testing, have continuously attracted research interests from both academia and industry. In the last decade, a rapidly increasing number of studies have been conducted to systematically generate metamorphic relations from various sources and for different application domains. In this article, based on the systematic review on the state of the art for metamorphic relations’ generation, we summarize and highlight visions for further advancing the theory and techniques for identifying and constructing metamorphic relations and discuss promising research directions in related areas. Rui Li 0013, Huai Liu, Pak-Lok Poon, Dave Towey, Chang-Ai Sun, Zheng Zheng 0001, Zhiquan Zhou 0001, Tsong Yueh Chen |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | Identifying the Failure-Revealing Test Cases in Metamorphic Testing: A Statistical ApproachabstractMetamorphic testing, thanks to its high failure-detection effectiveness especially in the absence of test oracle, has been widely applied in both the traditional context of software testing and other relevant fields such as fault localization and program repair. Its core element is a set of metamorphic relations, which are the necessary properties of the target algorithm in the form of the relationships among multiple inputs and corresponding expected outputs. When a relation is violated by the outputs of a group of test cases, namely metamorphic group of test cases, that are constructed based on the relation, a failure is said to be revealed. Traditionally, the primary task of software testing is to reveal failures. Therefore, from the perspective of software testing, it may not need to know which test case(s) in the metamorphic group cause the violation and thus the failure. However, such information is definitely helpful for other software engineering activities, such as software debugging. The current literature of metamorphic testing lacks a systematic mechanism of identifying the actual failure-revealing test cases, which hinders its applicability and effectiveness in other relevant fields. In this article, we propose a new technique for the FAILure-revealing Test case Identification in Metamorphic testing, namely FAILTIM. The approach is based on a novel application of statistical methods. More specifically, we leverage and adapt the basic ideas of spectrum-based techniques, which are originally used in fault localization, and propose the utilization of a set of risk formulas to estimate the suspiciousness of each individual test case in metamorphic groups. Failure-revealing test cases are then suggested according to their suspiciousness. A series of experiments have been conducted to evaluate the effectiveness and efficiency of FAILTIM using 9 subject programs and 30 risk formulas. The experimental results showed that the new approach can achieve a high accuracy in identifying the actual failure-revealing test cases in metamorphic testing. Consequently, our study will help boost the applicability and performance of metamorphic testing beyond testing to other software engineering areas. The present work also unfolds a number of research directions for further advancing the theory of metamorphic testing and more broadly, software testing. Zheng Zheng 0001, Dai-Xu Ren, Huai Liu, Tsong Yueh Chen, Tiancheng Li 0005 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | Semantic Structure Invariance-Based Metamorphic Testing for Machine Translation SystemsabstractIn recent years, deep neural networks have been applied in machine translation systems, resulting in the so-called neural machine translation (NMT) models that can improve translation quality significantly. However, due to the brittleness of deep neural network, machine translation systems could return erroneous translations that lead to misunderstandings or even cause serious losses. To detect translation errors, various testing techniques have been proposed. As a popularly used technique, metamorphic testing mainly relies on text or syntactic structure of translations while ignoring the meaning of sentences (i.e., semantic information). Compared with text and syntactic information, semantic information of sentences is more stable when dealing with languages that have rich vocabulary and flexible word order. Motivated by this observation, we propose semantic structure invariance-based metamorphic testing (SSIMT) for machine translation systems. The key insight is that contextually similar sentences should typically have translations of similar semantic structures. Experiments have been conducted to evaluate SSIMT on two widely used machine translation systems, Microsoft Bing Translator and Google Translate with 600 seed sentences crawled from well-known news websites covering six different corpus topics. The experimental results show that SSIMT is able to find thousands of erroneous translations in both translation systems with high accuracy (over 70%). Translation errors reported by SSIMT covers a wide variety of common error types. Chang-Ai Sun, Jian Mu, Mingjun Xiao, Huai Liu, Pinjia He |
IEEE Trans. Reliab. | 4 |
| 2025 | Metamorphic Testing for Smart Contracts: A User-Behavior-Sequence-Aware Approach and Automation Tool
Chang-Ai Sun, Yuanrui Ji, Xinhui Zheng, Huai Liu |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | MRTCNN: A Lightweight Approach for Predicting Metamorphic RelationsabstractMetamorphic testing is a software testing technique that provides both a test case generation strategy and a test result verification mechanism. Its foundation is a set of metamorphic relations, which are basically the necessary properties of the software under test, represented in the form of relationships among multiple inputs and corresponding expected outputs. As the core element of metamorphic testing, metamorphic relations have attracted lots of research interests from different perspectives, among which one major direction is to identify metamorphic relations suitable for certain types of programs. In order to reduce the manual work in the identification process, machine learning techniques have been leveraged to predict valid metamorphic relations for scientific software. In this paper, we present a new approach for predicting metamorphic relations based on the deep learning of the program documentation. In particular, we make use of the text convolutional neural networks in the prediction and validation of proper metamorphic relations. Empirical studies have also been conducted to evaluate the applicability and performance of our approach. The experimental results demonstrate its effectiveness in predicting appropriate metamorphic relations for the testing of various Java programs. Compared with the existing baseline techniques, our approach improves the precision and accuracy of the metamorphic relation prediction process. This study also reveals potential research opportunities for advancing the performance of metamorphic testing. Huai Liu |
APSEC | 3 |
| 2024 | MT4SC: A User-Behavior-Sequence-Aware Metamorphic Testing Approach for Smart ContractsabstractSmart contracts are essential applications for blockchains, which have been used in a wide variety of fields and are handling large amounts of valuable assets. Once deployed on the blockchain network, smart contracts cannot be altered, thus making the pre-deployment testing of them extremely critical. Nevertheless, the features of smart contracts, especially their transaction-driven nature, pose huge challenges to testing. In particular, since their inputs are not static data but dynamic sequences of user behavior, it is very difficult to obtain a feasible oracle for testing, which refers to the systematic mechanism to verify the correctness of test results given any test input. This is the notorious oracle problem in the context of software testing, for which the metamorphic testing technique has been widely recognized as a simple yet effective solution. In this paper, we develop a comprehensive framework, namely MT4SC, for implementing metamorphic testing on smart contracts. Specifically, we propose a systematic way to construct metamorphic relations, the core component of metamorphic testing, based on the user behavior sequences. A series of experiments have been conducted to evaluate the performance of MT4SC on eight different smart contract scenarios. The experimental results demonstrate MT4SC’s high effectiveness in detecting potential faults in smart contracts, even without the need for test oracles. This study bolsters the research on the testing of smart contracts, thereby improving their quality and ultimately advancing the reliability of blockchains. Yuan-rui Ji, Chang-Ai Sun, Xin-hui Zheng, Huai Liu |
ICWS | 4 |
| 2024 | Set evolution based test data generation for killing stubborn mutants
Changqing Wei, Xiangjuan Yao, Dun-Wei Gong, Huai Liu, Xiangying Dang |
J. Syst. Softw. | 4 |
| 2024 | An Interleaving Guided Metamorphic Testing Approach for Concurrent ProgramsabstractConcurrent programs are normally composed of multiple concurrent threads sharing memory space. These threads are often interleaved, which may lead to some non-determinism in execution results, even for the same program input. This poses huge challenges to the testing of concurrent programs, especially on the test result verification—that is, the prevalent existence of the oracle problem. In this article, we investigate the application of metamorphic testing (MT), a mainstream technique to address the oracle problem, into the testing of concurrent programs. Based on the unique features of interleaved executions in concurrent programming, we propose an extended notion of metamorphic relations, the core part of MT, which are particularly designed for the testing of concurrent programs. A comprehensive testing approach, namely ConMT , is thus developed and a tool is built to automate its implementation on concurrent programs written in Java. Empirical studies have been conducted to evaluate the performance of ConMT, and the experimental results show that in addition to addressing the oracle problem, ConMT outperforms the baseline traditional testing techniques with respect to a higher degree of automation, better bug detection capability, and shorter testing time. It is clear that ConMT can significantly improve the cost-effectiveness for the testing of concurrent programs and thus advances the state of the art in the field. The study also brings novelty into MT, hence promoting the fundamental research of software testing. Chang-Ai Sun, Hepeng Dai, Ning Geng, Huai Liu, Tsong Yueh Chen, Peng Wu 0002, Yan Cai 0001, Jinqiu Wang |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | DFuzzer: Diversity-Driven Seed Queue Construction of Fuzzing for Deep Learning ModelsabstractIn light of high-performance computer processing, massive datasets, and mighty algorithms, we are rapidly entering an age where the advanced deep learning (DL) capabilities are integrated into the contemporary software systems to fulfill critical tasks. Like “traditional” software, DL systems are not immune to faults, some of which may even cause catastrophic disasters. As a mainstream testing technique for DL systems, fuzzing attempts to generate a large amount of semirandom yet syntactically valid test cases, from which the so-called adversarial inputs can be found, indicating the detection of faults. Test cases in fuzzing are generated based on a seed queue, which is constructed by randomly selecting seeds from the existing test suite (that is, the set of test cases). In this article, we propose a diversity-driven approach, namely DFuzzer, for constructing seed queues in fuzzing. We particularly develop two algorithms, namely DFuzzer-IB and DFuzzer-FB, based on the information theory and deep features, respectively, to improve the diversity of seed queues. Experimental studies have been conducted to evaluate the proposed techniques based on five fuzzers, three datasets, and seven DL models. The experimental results show that both strategies can significantly improve the performance of the state-of-the-art fuzzers for DL, including DeepXplore, DLFuzz, Tensorfuzz, DeepHunter, and DeepSmartFuzzer, not only in terms of finding more adversarial inputs for triggering faults but also achieving higher coverage. Our article demonstrates that the improved diversity of seed queues and the resultant test cases can help achieve a high testing effectiveness of fuzzing. Hepeng Dai, Chang-Ai Sun, Huai Liu, Xiangyu Zhang 0001 |
IEEE Trans. Reliab. | 3 |
| 2024 | Test Data Generation for Mutation Testing Based on Markov Chain Usage Model and Estimation of Distribution AlgorithmabstractMutation testing, a mainstream fault-based software testing technique, can mimic a wide variety of software faults by seeding them into the target program and resulting in the so-called mutants. Test data generated in mutation testing should be able to kill as many mutants as possible, hence guaranteeing a high fault-detection effectiveness of testing. Nevertheless, the test data generation can be very expensive, because mutation testing normally involves an extremely large number of mutants and some mutants are hard to kill. It is thus a critical yet challenging job to find an efficient way to generate a small set of test data that are able to kill multiple mutants at the same time as well as reveal those hard-to-detect faults. In this paper, we propose a new approach for test data generation in mutation testing, through the novel applications of the Markov chain usage model and the estimation of distribution algorithm. We first utilize the Markov chain usage model to reduce the so-called mutant branches in weak mutation testing and generate a minimal set of extended paths. Then, we regard the problem of generating test data as the problem of covering extended paths and use an estimation of distribution algorithm based on probability model to solve the problem. Finally, we develop a framework, TAMMEA, to implement the new approach of generating test data for mutation testing. The empirical studies based on fifteen object programs show that TAMMEA can kill more mutants using fewer test data compared with baseline techniques. In addition, the computation overhead of TAMMEA is lower than that of the baseline technique based on the traditional genetic algorithm, and comparable to that of the random method. It is clear that the new approach improves both the effectiveness and efficiency of mutation testing, thus promoting its practicability. Changqing Wei, Xiangjuan Yao, Dun-Wei Gong, Huai Liu |
IEEE Trans. Software Eng. | 4 |
| 2023 | A Trace-Log-Clusterings-Based Fault Localization Approach to Microservice SystemsabstractMicroservice architecture has been widely used for the development of large-scale distributed applications. Microservice systems normally have high complexity and loose coupling nature, which make it challenging to localize faults in them. Automated fault localization is particularly difficult for microservice systems, due to their unique features, such as frequent updates, complex dependencies, and multiple microservice instances. In this paper, we propose a fault localization approach for microservice systems based on trace log clusterings, called TLCluster. TLCluster first derives trace logs by collecting and combining communication messages and logs of microservice systems, then clusters trace logs for different business process categories, calculates similarities between normal and abnormal trace logs, and finally evaluates and ranks the suspiciousness scores of microservice instances. We conducted a series of experiments to evaluate the effectiveness of TLCluster using a large-scale microservice system. Experimental results show that our approach is able to effectively localize faults of microservice systems and demonstrates a better fault localization accuracy and precision compared with state-of-the-art baseline techniques. Chang-Ai Sun, Wanqing Zuo, Huai Liu |
ICWS | 4 |
| 2023 | RFLSem: A Lightweight Model for Textual Sentiment Analysis
Jiayi Dang, Huai Liu |
KSEM (4) | 3 |
| 2023 | User story clustering in agile development: a framework and an empirical study
Xiuyin Ma, Haoran Guo, Huai Liu |
Frontiers Comput. Sci. | 5 |
| 2023 | Evaluation and assessment of machine learning based user story grouping: A framework and empirical studies
Haoran Guo, Huai Liu |
Sci. Comput. Program. | 3 |
| 2023 | Feedback-Directed Metamorphic TestingabstractOver the past decade, metamorphic testing has gained rapidly increasing attention from both academia and industry, particularly thanks to its high efficacy on revealing real-life software faults in a wide variety of application domains. On the basis of a set of metamorphic relations among multiple software inputs and their expected outputs, metamorphic testing not only provides a test case generation strategy by constructing new (or follow-up) test cases from some original (or source) test cases, but also a test result verification mechanism through checking the relationship between the outputs of source and follow-up test cases. Many efforts have been made to further improve the cost-effectiveness of metamorphic testing from different perspectives. Some studies attempted to identify “good” metamorphic relations, while other studies were focused on applying effective test case generation strategies especially for source test cases. In this article, we propose improving the cost-effectiveness of metamorphic testing by leveraging the feedback information obtained in the test execution process. Consequently, we develop a new approach, namely feedback-directed metamorphic testing, which makes use of test execution information to dynamically adjust the selection of metamorphic relations and selection of source test cases. We conduct an empirical study to evaluate the proposed approach based on four laboratory programs, one GNU program, and one industry program. The empirical results show that feedback-directed metamorphic testing can use fewer test cases and take less time than the traditional metamorphic testing for detecting the same number of faults. It is clearly demonstrated that the use of feedback information about test execution does help enhance the cost-effectiveness of metamorphic testing. Our work provides a new perspective to improve the efficacy and applicability of metamorphic testing as well as many other software testing techniques. Chang-Ai Sun, Hepeng Dai, Huai Liu, Tsong Yueh Chen |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2023 | A Declarative Metamorphic Testing Framework for Autonomous DrivingabstractAutonomous driving has gained much attention from both industry and academia. Currently, Deep Neural Networks (DNNs) are widely used for perception and control in autonomous driving. However, several fatal accidents caused by autonomous vehicles have raised serious safety concerns about autonomous driving models. Some recent studies have successfully used the metamorphic testing technique to detect thousands of potential issues in some popularly used autonomous driving models. However, prior study is limited to a small set of metamorphic relations, which do not reflect rich, real-world traffic scenarios and are also not customizable. This paper presents a novel declarative rule-based metamorphic testing framework calledRMT.RMTprovides a rule template with natural language syntax, allowing users to flexibly specify an enriched set of testing scenarios based on real-world traffic rules and domain knowledge.RMTautomatically parses human-written rules to metamorphic relations using an NLP-based rule parser referring to an ontology list and generates test cases with a variety of image transformation engines. We evaluatedRMTon three autonomous driving models. With an enriched set of metamorphic relations,RMTdetected a significant number of abnormal model predictions that were not detected by prior work. Through a large-scale human study on Amazon Mechanical Turk, we further confirmed the authenticity of test cases generated byRMTand the validity of detected abnormal model predictions. James Xi Zheng, Tianyi Zhang 0001, Huai Liu, Guannan Lou, Miryung Kim, Tsong Yueh Chen |
IEEE Trans. Software Eng. | 4 |
| 2022 | DeepController: Feedback-Directed Fuzzing for Deep Learning SystemsabstractDeep learning (DL) systems are increasingly adopted in various fields, while fatal failures are still inevitable in them.One mainstream testing approach for DL is fuzzing, which can generate a large amount of semi-random yet syntactically valid test cases.Previous studies on fuzzing are mainly focused on selecting "quality" seeds or using "good" mutation strategies.In this paper, we attempt to improve the performance of fuzzing from a different perspective.A new fuzzer, namely DeepController, is accordingly developed, which makes use of the feedback information obtained in the test execution process to dynamically select seeds and mutation strategies.DeepController is evaluated through empirical studies on three datasets and eight DL models.The experimental results show that, with the same number of seeds, DeepController can generate more adversarial inputs and achieve higher neuron coverage than the state-of-the-art testing techniques for DL systems. Hepeng Dai, Chang-Ai Sun, Huai Liu |
SEKE | 3 |
| 2022 | Path-directed source test case generation and prioritization in metamorphic testing
Chang-Ai Sun, Baoli Liu, An Fu, Yiqiang Liu, Huai Liu |
J. Syst. Softw. | 5 |
| 2022 | Precise Learning of Source Code Contextual Semantics via Hierarchical Dependence Structure and Graph Attention Networks
Zhehao Zhao, Ge Li 0001, Huai Liu, Zhi Jin 0001 |
J. Syst. Softw. | 4 |
| 2022 | Enhancement of Mutation Testing via Fuzzy Clustering and Multi-Population Genetic AlgorithmabstractMutation testing, a fundamental software testing technique, which is a typical way to evaluate the adequacy of a test suite. In mutation testing, a set of mutants are generated by seeding the different classes of faults into a program under test. Test data shall be generated in the way that as many mutants can be killed as possible. Thanks to numerous tools to implement mutation testing for different languages, a huge amount of mutants are normally generated even for small-sized programs. However, a large number of mutants not only leads to a high cost of mutation testing, but also make the corresponding test data generation a non-trivial task. In this paper, we make use of intelligent technologies to improve the effectiveness and efficiency of mutation testing from two perspectives. A machine learning technique, namely fuzzy clustering, is applied to categorize mutants into different clusters. Then, a multi-population genetic algorithm via individual sharing is employed to generate test data for killing the mutants in different clusters in parallel when the problem of test data generation as an optimization one. A comprehensive framework, termed as$\mathbf {FUZGENMUT}$, is thus developed to implement the proposed techniques. The experiments based on nine programs of various sizes show that fuzzy clustering can help to reduce the cost of mutation testing effectively, and that the multi-population genetic algorithm improves the efficiency of test data generation while delivering the high mutant-killing capability. The results clearly indicate that the huge potential of using intelligent technologies to enhance the efficacy and thus the practicality of mutation testing. Xiangying Dang, Dun-Wei Gong, Xiangjuan Yao, Tian Tian 0010, Huai Liu |
IEEE Trans. Software Eng. | 5 |
| 2021 | Performance Analysis of Open-Source Hypervisors for Automotive SystemsabstractNowadays, automotive products are intelligence intensive and thus inevitably handle multiple functionalities under the current high-speed networking environment. The embedded virtualization has high potentials in the automotive industry, thanks to its advantages in function integration, resource utilization, and security. The invention of ARM virtualization extensions has made it possible to run open-source hypervisors, such as Xen and KVM, for embedded applications. Nevertheless, there is little work to investigate the performance of these hypervisors on automotive platforms. This paper presents a detailed analysis of different types of open-source hypervisors that can be applied in the ARM platform. We carry out the virtualization performance experiment from the perspectives of CPU, memory, file I/O, and some OS operation performance on Xen and Jailhouse. A series of microbenchmark programs have been designed, specifically to evaluate the real-time performance of various hypervisors and the relevant overhead. Compared with Xen, Jailhouse has better latency performance, stable latency, and little interference jitter. The performance experiment results help us summarize the advantages and disadvantages of these hypervisors in automotive applications. Zhengjun Zhang, Yanqiang Liu, Jiangtao Chen, Zhengwei Qi, Huai Liu |
ICPADS | 6 |
| 2021 | Metamorphic Testing on Multi-module UAV SystemsabstractRecent years have seen a rapid development of machine learning based multi-module unmanned aerial vehicle (UAV) systems. To address the oracle problem in autonomous systems, numerous studies have been conducted to use metamorphic testing to automatically generate test scenes for various modules, e.g., those in self-driving cars. However, as most of the studies are based on unit testing including end-to-end model-based testing, a similar testing approach may not be equally effective for UAV systems where multiple modules are working closely together. Therefore, in this paper, instead of unit testing, we propose a novel metamorphic system testing framework for UAV, named MSTU, to detect the defects in multi-module UAV systems. A preliminary evaluation plan to apply MSTU on an emerging autonomous multi-module UAV system is also presented to demonstrate the feasibility of the proposed testing framework. Rui Li 0013, Huai Liu, Guannan Lou, James Xi Zheng, Xiao Liu 0004, Tsong Yueh Chen |
ASE | 2 |
| 2021 | Software debugging analysis based on developer behavior data
Huai Liu, Chao Liu 0002 |
Frontiers Comput. Sci. | 3 |
| 2021 | Spectral clustering based mutant reduction for mutation testing
Changqing Wei, Xiangjuan Yao, Dun-Wei Gong, Huai Liu |
Inf. Softw. Technol. | 4 |
| 2021 | METRIC$^{+}$+: A Metamorphic Relation Identification Technique Based on Input Plus Output DomainsabstractMetamorphic testing is well known for its ability to alleviate the oracle problem in software testing. The main idea ofmetamorphic testing is to test a software system by checking whether each identified metamorphic relation (MR) holds among severalexecutions. In this regard, identifying MRs is an essential task in metamorphic testing. In view of the importance of this identificationtask, METRIC (METamorphic Relation Identification based on Category-choice framework) was developed to help software testersidentify MRs from a given set of complete test frames. However, during MR identification, METRIC primarily focuses on the inputdomain without sufficient attention given to the output domain, thereby hindering the effectiveness of METRIC. Inspired by this problem,we have extended METRIC into METRIC+by incorporating the information derived from the output domain for MR identification. A toolimplementing METRIC+has also been developed. Two rounds of experiments, involving four real-life specifications, have beenconducted to evaluate the effectiveness and efficiency of METRIC+. The results have confirmed that METRIC+is highly effective andefficient in MR identification. Additional experiments have been performed to compare the fault detection capability of the MRsgenerated by METRIC+and those bymMT (another MR identification technique). The comparison results have confirmed that the MRsgenerated by METRIC+are highly effective in fault detection. Chang-Ai Sun, An Fu, Pak-Lok Poon, Xiaoyuan Xie, Huai Liu, Tsong Yueh Chen |
IEEE Trans. Software Eng. | 5 |
| 2020 | TDD4Fog: A Test-Driven Software Development Platform for Fog Computing SystemsabstractAs an ideal infrastructure for smart services, Fog Computing is becoming the next wave of IT investment harnessing the successful models of Cloud Computing and latest technologies such as 5G and Internet of Things (IoT). However, the development of Fog Computing systems is a big challenge due to its complex, heterogeneous and distributed nature. Currently, there are a few SDKs released by some public Cloud service providers to support the development of Fog services in a top-down fashion as the key motive is to leverage their business Cloud services. However, Fog Computing systems are usually designed in a bottom-up fashion as the major functionalities are centred around the Edge Nodes and the End Devices. Meanwhile, significant efforts are required to verify the conformance of software behaviours as the collaboration between the End Devices, Edge Nodes and Cloud Servers is vital to the success of a Fog Computing System. Therefore, a holistically designed software development platform is urgently required. In this paper, we propose TDD4Fog, a test-driven software development platform for Fog Computing systems. Following the Test-Driven Development (TDD) methodology and a bottom-up design fashion, TDD4Fog supports the microservice architecture and provides the Test-Driven utilities such as metamorphic testing, mutation testing and random testing for the whole software development lifecycle of Fog Computing systems. To demonstrate the feasibility of TDD4Fog, we have presented some preliminary results on the key components of TDD4Fog and discussed some important future research directions. Rui Li 0013, Xiao Liu 0004, James Xi Zheng, Chong Zhang 0007, Huai Liu |
CCGRID | 5 |
| 2020 | A Lightweight Fault Localization Approach based on XGBoostabstractSoftware fault localization is one of the key activities in software debugging. The program spectrum-based approach is widely used in fault localization. However, lots of program information, for example, the sequence of the execution statement and statement semantics, is missing when such an approach is utilized, which affects the performance. XGBoost is an effective learning algorithm, which can use the characteristics of the training data to build a classification tree during training. In addition, XGBoost can iteratively adjust the information value of the feature, so that the training process retains the importance information of the feature. This paper proposes applying XGBoost into fault localization utilizing information of program execution behaviors. A novel method called XGB-FL is developed, where the program spectrum information is converted into a coverage matrix to train the XGBoost model. We can get the characteristics of the data through the trained model and the importance of the program statement in the classification process. This is also the basis for judging whether the statement is likely to contain a fault. Nine representative data sets have been chosen to evaluate the performance of XGB-FL. The experimental results show that XGB-FL can generally deliver a higher performance in fault localization than those baseline techniques, in terms of precision and efficiency. Huai Liu, Yixin Chen 0002 |
QRS | 3 |
| 2019 | Verification of Microservices Using Metamorphic Testing
James Xi Zheng, Huai Liu, Rongbin Xu, Dinesh Nagumothu, Ranjith Janapareddi, Er Zhuang, Xiao Liu 0004 |
ICA3PP (1) | 3 |
| 2019 | Genetic algorithm based test data generation for MPI parallel programs with blocking communication
Tian Tian 0010, Dun-Wei Gong, Fei-Ching Kuo, Huai Liu |
J. Syst. Softw. | 4 |
| 2019 | Adaptive Partition TestingabstractRandom testing and partition testing are two major families of software testing techniques. They have been compared both theoretically and empirically in numerous studies for decades, and it has been widely acknowledged that they have their own advantages and disadvantages and that their innate characteristics are fairly complementary to each other. Some work has been conducted to develop advanced testing techniques through the integration of random testing and partition testing, attempting to preserve the advantages of both while minimizing their disadvantages. In this paper, we propose a new testing approach, adaptive partition testing, where test cases are randomly selected from some partition whose probability of being selected is adaptively adjusted along the testing process. We particularly develop two algorithms, Markov-chain based adaptive partition testing and reward-punishment based adaptive partition testing, to implement the proposed approach. The former algorithm makes use of Markov matrix to dynamically adjust the probability of a partition to be selected for conducting tests; while the latter is based on a reward and punishment mechanism. We conduct empirical studies to evaluate the performance of the proposed algorithms using ten faulty versions of three large-scale open source programs. Our experimental results show that, compared with two baseline techniques, namely random partition testing (RPT) and dynamic random testing (DRT), our algorithms deliver higher fault-detection effectiveness with lower test case selection overhead. It is demonstrated that the proposed adaptive partition testing is an effective testing approach, taking advantages of both random testing and partition testing. Chang-Ai Sun, Hepeng Dai, Huai Liu, Tsong Yueh Chen, Kai-Yuan Cai |
IEEE Trans. Computers | 3 |
| 2018 | A Lightweight Program Dependence Based Approach to Concurrent Mutation AnalysisabstractMutation analysis is a classical software testing approach which attempts to imitate faults using a set of mutants. It has been advocated to be an appropriate technique for evaluating the quality of test suites as well as the effectiveness of a testing method. However, the applicability of mutation analysis, especially in many practical situations, has been hindered due to the high computation cost and the long execution time, which are mainly caused by the large number of mutants. Numerous studies, particularly those based on parallel computing, have been conducted to reduce the overhead of mutation analysis. In this paper, we aim to improve the efficiency of mutation analysis from a different perspective. We make use of lightweight program analysis techniques to identify a group of mutants that share the common execution traces before the mutation location, and then merge them into a synthesized program with the concurrent mechanism, on which mutation analysis can be efficiently executed without the duplicate execution of common traces. Our empirical study demonstrates that our approach can significantly decrease the computation overhead as well as shorten the execution time of mutation analysis, without jeopardizing its effectiveness. The in-depth analysis further shows that the effectiveness of our approach is positively correlated with the number of branches in the program under test. Our approach makes it possible to efficiently execute mutation analysis even without the need of advanced computer architectures. Chang-Ai Sun, Jingting Jia, Huai Liu, Xiangyu Zhang 0001 |
COMPSAC (1) | 3 |
| 2018 | Fault localisation for WS-BPEL programs based on predicate switching and program slicing
Chang-Ai Sun, Yufeng Ran, Caiyun Zheng, Huai Liu, Dave Towey, Xiangyu Zhang 0001 |
J. Syst. Softw. | 4 |
| 2018 | Automated Testing of WS-BPEL Service Compositions: A Scenario-Oriented ApproachabstractNowadays, service oriented architecture (SOA) has become one mainstream paradigm for developing distributed applications. As the basic unit in SOA, web services can be composed to construct complex applications. The quality of web services and their compositions is critical to the success of SOA applications. Testing, as a major quality assurance technique, is confronted with new challenges in the context of service compositions. In this paper, we propose a scenario-oriented testing approach that can automatically generate test cases for service compositions. Our approach is particularly focused on the service compositions specified by Business Process Execution Language for web services (WS-BPEL), a widely recognized executable service composition language. In the approach, a WS-BPEL service composition is first abstracted into a graph model; test scenarios are then derived from the model; finally, test cases are generated according to different scenarios. We also developed a prototype tool implementing the proposed approach, and an empirical study was conducted to demonstrate the applicability and effectiveness of our approach. The experimental results show that the automatic scenario-oriented testing approach is effective in detecting many types of faults seeded in the service compositions. Chang-Ai Sun, Huai Liu, Tsong Yueh Chen |
IEEE Trans. Serv. Comput. | 4 |
| 2017 | An Empirical Study on Mutation Testing of WS-BPEL ProgramsabstractNowadays, applications are increasingly deployed as Web services in the globally distributed cloud computing environment. Multiple services are normally composed to fulfill complex functionalities. Business Process Execution Language for Web Services (WS-BPEL) is an XML-based service composition language that is used to define a complex business process by orchestrating multiple services. Compared with traditional applications, WS-BPEL programs pose many new challenges to the quality assurance, especially testing, of service compositions. A number of techniques have been proposed for testing WS-BPEL programs, but only a few studies have been conducted to systematically evaluate the effectiveness of these techniques. Mutation testing has been widely acknowledged as not only a testing method in its own right but also a popular technique for measuring the fault-detection effectiveness of other testing methods. Several previous studies have proposed a family of mutation operators for generating mutants by seeding various faults into WS-BPEL programs. In this study, we conduct a series of empirical studies to evaluate the applicability and effectiveness of various mutation operators for WS-BPEL programs. The experimental results provide insightful and comprehensive guidance for mutation testing of WS-BPEL programs in practice. In particular, our work is the systematic study in the selection of effective mutation operators specifically for WS-BPEL programs. Chang-Ai Sun, Qiaoling Wang, Huai Liu, Xiangyu Zhang 0001 |
Comput. J. | 4 |
| 2017 | A path-aware approach to mutant reduction in mutation testing
Chang-Ai Sun, Feifei Xue, Huai Liu, Xiangyu Zhang 0001 |
Inf. Softw. Technol. | 3 |
| 2016 | A Cost-Effective Random Testing Method for Programs with Non-Numeric InputsabstractRandom testing (RT) has been widely used in the testing of various software and hardware systems. Adaptive random testing (ART) is a family of random testing techniques that aim to enhance the failure-detection effectiveness of RT by spreading random test cases evenly throughout the input domain. ART has been empirically shown to be effective on software with numeric inputs. However, there are two aspects of ART that need to be addressed to render its adoption more widespread-applicability to programs with nonnumeric inputs, and the high computation overhead of many ART algorithms. We present a linear-order ART algorithm for software with non-numeric inputs. The key requirement for using ART with non-numeric inputs is an appropriate “distance” measure. We use the concepts of categories and choices from category-partition testing to formulate such a measure. We investigate the failure-detection effectiveness of our technique by performing an empirical study on 14 object programs, using two standard metrics-F-measure and P-measure. Our ART algorithm statistically significantly outperforms RT on 10 of the 14 programs studied, and exhibits performance similar to RT on three of the four remaining programs. The selection overhead of our ART algorithm is close to that of RT. A. C. Barus, Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, Robert G. Merkel, Gregg Rothermel |
IEEE Trans. Computers | 4 |
| 2016 | Randomized Quasi-Random TestingabstractRandom testing is a fundamental testing technique that can be used to generate test cases for both hardware and software systems. Quasi-random testing was proposed as an enhancement to the cost-effectiveness of random testing: In addition to having similar computation overheads to random testing, it makes use of quasi-random sequences to generate low-discrepancy and low-dispersion test cases that help deliver high failure-detection effectiveness. Currently, few algorithms exist to generate quasi-random sequences, and these are mostly deterministic, rather than random. A previous study of quasi-random testing has examined two methods for randomizing quasi-random sequences to improve their applicability in testing. However, these randomization methods still have shortcomings-one method does not introduce much randomness to the test cases, while the other does not support incremental test case generation. In this paper, we present an innovative approach to incrementally randomizing quasi-random sequences. The test cases generated by this new approach show a high degree of randomness and evenness in distribution. We also conduct simulations and empirical studies to demonstrate the applicability and effectiveness of our approach in software testing. Huai Liu, Tsong Yueh Chen |
IEEE Trans. Computers | 1 |
| 2015 | Poster: Enhancing Partition Testing through Output VariationabstractA major test case generation approach is to divide the input domain into disjoint partitions, from which test cases can be selected. However, we observe that in some traditional approaches to partition testing, the same partition may be associated with different output scenarios. Such an observation implies that the partitioning of the input domain may not be precise enough for effective software fault detection. To solve this problem, partition testing should be fine-tuned to additionally use the information of output scenarios in test case generation, such that these test cases are more fine-grained not only with respect to the input partitions but also from the perspective of output scenarios. Huai Liu, Pak-Lok Poon, Tsong Yueh Chen |
ICSE (2) | 1 |
| 2015 | Evaluating and Comparing Fault-Based Testing Strategies for General Boolean Specifications: A Series of ExperimentsabstractA great amount of fault-based testing strategies have been proposed to generate test cases for detecting certain types of faults in Boolean specifications. However, most of the previous studies on these strategies were focused on the Boolean expressions in the disjunctive normal form (DNF), even the irredundant DNF (IDNF)—little work has been conducted to comprehensively investigate their performance on general Boolean specifications. In this study, we conducted a series of experiments to evaluate and compare 18 fault-based testing strategies using over 4000 randomly generated fault-seeded Boolean expressions. In the experiments, a testing strategy is regarded as effective and efficient if it can detect most of the seeded faults using a small number of test cases. Our experimental results show that if a testing strategy is highly effective and efficient when testing the Boolean expressions in the IDNF, it also shows high effectiveness and efficiency on general Boolean expressions. It is found that one family of fault-based testing strategies, namely MUMCUT, normally deliver the best performance among all the 18 strategies. Our study provides an in-depth understanding and insight of fault-based testing for general Boolean expressions. Chang-Ai Sun, Yimeng Zai, Huai Liu |
Comput. J. | 3 |
| 2015 | Enhancing mirror adaptive random testing through dynamic partitioning
Rubing Huang, Huai Liu, Jinfu Chen 0001 |
Inf. Softw. Technol. | 2 |
| 2014 | An Application of Adaptive Random Sequence in Test Case Prioritization
Tsong Yueh Chen, Huai Liu |
SEKE | 3 |
| 2014 | How Effectively Does Metamorphic Testing Alleviate the Oracle Problem?abstractIn software testing, something which can verify the correctness of test case execution results is called an oracle. The oracle problem occurs when either an oracle does not exist, or exists but is too expensive to be used. Metamorphic testing is a testing approach which uses metamorphic relations, properties of the software under test represented in the form of relations among inputs and outputs of multiple executions, to help verify the correctness of a program. This paper presents new empirical evidence to support this approach, which has been used to alleviate the oracle problem in various applications and to enhance several software analysis and testing techniques. It has been observed that identification of a sufficient number of appropriate metamorphic relations for testing, even by inexperienced testers, was possible with a very small amount of training. Furthermore, the cost-effectiveness of the approach could be enhanced through the use of more diverse metamorphic relations. The empirical studies presented in this paper clearly show that a small number of diverse metamorphic relations, even those identified in an ad hoc manner, had a similar fault-detection capability to a test oracle, and could thus effectively help alleviate the oracle problem. Huai Liu, Fei-Ching Kuo, Dave Towey, Tsong Yueh Chen |
IEEE Trans. Software Eng. | 1 |
| 2013 | Code Coverage of Adaptive Random TestingabstractRandom testing is a basic software testing technique that can be used to assess the software reliability as well as to detect software failures. Adaptive random testing has been proposed to enhance the failure-detection capability of random testing. Previous studies have shown that adaptive random testing can use fewer test cases than random testing to detect the first software failure. In this paper, we evaluate and compare the performance of adaptive random testing and random testing from another perspective, that of code coverage. As shown in various investigations, a higher code coverage not only brings a higher failure-detection capability, but also improves the effectiveness of software reliability estimation. We conduct a series of experiments based on two categories of code coverage criteria: structure-based coverage, and fault-based coverage. Adaptive random testing can achieve higher code coverage than random testing with the same number of test cases. Our experimental results imply that, in addition to having a better failure-detection capability than random testing, adaptive random testing also delivers a higher effectiveness in assessing software reliability, and a higher confidence in the reliability of the software under test even when no failure is detected. Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, W. Eric Wong |
IEEE Trans. Reliab. | 3 |
| 2012 | Comparison of adaptive random testing and random testing under various testing and debugging scenariosabstractSUMMARY Adaptive random testing is an enhancement of random testing. Previous studies on adaptive random testing assumed that once a failure is detected, testing is terminated and debugging is conducted immediately. It has been shown that adaptive random testing normally uses fewer test cases than random testing for detecting the first software failure. However, under many practical situations, testing should not be withheld after the detection of a failure. Thus, it is important to investigate the effectiveness with respect to the detection of multiple failures. In this paper, we compare adaptive random testing and random testing under various scenarios and examine whether adaptive random testing is still able to use fewer test cases than random testing to detect multiple software failures. Our study delivers some interesting results and highlights a number of promising research projects. Copyright © 2011 John Wiley & Sons, Ltd. Huai Liu, Fei-Ching Kuo, Tsong Yueh Chen |
Softw. Pract. Exp. | 1 |
| 2011 | Metamorphic Testing for Web Services: Framework and a Case StudyabstractService Oriented Architecture (SOA) has become a major application development paradigm. As a basic unit of SOA applications, Web services significantly affect the quality of the applications constructed from them. Since the development and consumption of Web services are completely separated under SOA environment, the consumers are normally provided with limited knowledge of the services and thus have little information about test oracles. The lack of source code and the restricted control of Web services limit the testability of Web services. To address the prominent oracle problem when testing Web services, we propose a metamorphic testing framework for Web services taking into account the unique features of SOA. We conduct a case study where the new metamorphic testing framework is employed to test a Web service that implements the electronic payment. The results of case study show the feasibility of the framework for web services, and also the efficiency of metamorphic testing. The work presented in the paper alleviates the test oracle problem when testing Web services under SOA. Chang-Ai Sun, Baohong Mu, Huai Liu, ZhaoShun Wang, Tsong Yueh Chen |
ICWS | 4 |
| 2011 | Adaptive random testing through test profilesabstractSUMMARY Random testing (RT), which simply selects test cases at random from the whole input domain, has been widely applied to test software and assess the software reliability. However, it is controversial whether RT is an effective method to detect software failures. Adaptive random testing (ART) is an enhancement of RT in terms of failure‐detection effectiveness. Its basic intuition is to evenly spread random test cases all over the input domain. There are various notions to achieve the goal of even spread, and each notion can be implemented by different algorithms. For example, ‘by exclusion’ and ‘by partitioning’ are two different notions to evenly spread test cases. Restricted random testing (RRT) is a typical algorithm for the notion of ‘by exclusion’, whereas the notion of ‘by partitioning’ can be implemented by either the technique of bisection (ART‐B) or the technique of random partitioning (ART‐RP). In this paper, we propose a generic approach that can be used to implement different notions. In the new approach, test cases are simply selected based on test profiles that are in turn designed according to certain notions. In this study, we design several test profiles for the notions of ‘by exclusion’ and ‘by partitioning’, and then use these profiles to illustrate our new approach. Our experimental results show that compared with the original RRT, ART‐B, and ART‐RP algorithms, our new approach normally brings at least a higher failure‐detection capability or a lower computational overhead. Copyright © 2011 John Wiley & Sons, Ltd. Huai Liu, Yansheng Lu, Tsong Yueh Chen |
Softw. Pract. Exp. | 1 |
| 2010 | Teaching an End-User Testing MethodologyabstractOne important focus of software engineering is how to develop quality software. Software testing is the main approach to the software quality assurance. Nowadays, more and more end-users write the program on their own but lack formal trainings on how to test their programs, and hence cannot guarantee the quality of their own software. Metamorphic testing is a simple, automatable, and cost-effective testing methodology. It is particularly suitable for end-users to test their own programs, because it does not demand the user to have great knowledge of software testing but knowledge of the program under development. In this paper, we report our experience in teaching metamorphic testing to various groups of students at Swinburne University of Technology, Melbourne, Australia. Our work not only enhances the teaching of software testing, but also fosters the training of end-user programmers. Huai Liu, Fei-Ching Kuo, Tsong Yueh Chen |
CSEE&T | 1 |
| 2009 | On the integration of metamorphic testing and model checking
Huai Liu, Daoming Wang, Huimin Lin, Tsong Yueh Chen |
IADIS AC (2) | 1 |
| 2009 | Dynamic Test Profiles in Adaptive Random Testing: A Case Study
Huai Liu, Fei-Ching Kuo, Tsong Yueh Chen |
SEKE | 1 |
| 2009 | An innovative approach for testing bioinformatics programs using metamorphic testingabstractBACKGROUND: Recent advances in experimental and computational technologies have fueled the development of many sophisticated bioinformatics programs. The correctness of such programs is crucial as incorrectly computed results may lead to wrong biological conclusion or misguided downstream experimentation. Common software testing procedures involve executing the target program with a set of test inputs and then verifying the correctness of the test outputs. However, due to the complexity of many bioinformatics programs, it is often difficult to verify the correctness of the test outputs. Therefore our ability to perform systematic software testing is greatly hindered. RESULTS: We propose to use a novel software testing technique, metamorphic testing (MT), to test a range of bioinformatics programs. Instead of requiring a mechanism to verify whether an individual test output is correct, the MT technique verifies whether a pair of test outputs conform to a set of domain specific properties, called metamorphic relations (MRs), thus greatly increases the number and variety of test cases that can be applied. To demonstrate how MT is used in practice, we applied MT to test two open-source bioinformatics programs, namely GNLab and SeqMap. In particular we show that MT is simple to implement, and is effective in detecting faults in a real-life program and some artificially fault-seeded programs. Further, we discuss how MT can be applied to test programs from various domains of bioinformatics. CONCLUSION: This paper describes the application of a simple, effective and automated technique to systematically test a range of bioinformatics programs. We show how MT can be implemented in practice through two real-life case studies. Since many bioinformatics programs, particularly those for large scale simulation and data analysis, are hard to test systematically, their developers may benefit from using MT as part of the testing strategy. Therefore our work represents a significant step towards software reliability in bioinformatics. Tsong Yueh Chen, Joshua W. K. Ho, Huai Liu, Xiaoyuan Xie |
BMC Bioinform. | 3 |
| 2009 | Adaptive random testing based on distribution metrics
Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu |
J. Syst. Softw. | 3 |
| 2009 | Application of a Failure Driven Test Profile in Random TestingabstractRandom testing techniques have been extensively used in reliability assessment, as well as in debug testing. When used to assess software reliability, random testing selects test cases based on an operational profile; while in the context of debug testing, random testing often uses a uniform distribution. However, generally neither an operational profile nor a uniform distribution is chosen from the perspective of maximizing the effectiveness of failure detection. Adaptive random testing has been proposed to enhance the failure detection capability of random testing by evenly spreading test cases over the whole input domain. In this paper, we propose a new test profile, which is different from both the uniform distribution, and operational profiles. The aim of the new test profile is to maximize the effectiveness of failure detection. We integrate this new test profile with some existing adaptive random testing algorithms, and develop a family of new random testing algorithms. These new algorithms not only distribute test cases more evenly, but also have better failure detection capabilities than the corresponding original adaptive random testing algorithms. As a consequence, they perform better than the pure random testing. Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu |
IEEE Trans. Reliab. | 3 |
| 2008 | Distributing test cases more evenly in adaptive random testing
Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu |
J. Syst. Softw. | 3 |
| 2008 | Enhancing adaptive random testing for programs with high dimensional input domains or failure-unrelated parameters
Fei-Ching Kuo, Tsong Yueh Chen, Huai Liu, Wing Kwong Chan |
Softw. Qual. J. | 3 |
| 2007 | On Test Case Distributions of Adaptive Random Testing
Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu |
SEKE | 3 |