VLDB 2026 Research / reviewers in the wild / expert
Huayao Wu
dblp:127/0799
· DBLP profile ↗
23ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-1383-5421ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Quality Assurance for Human-Machine-Thing Integrated Intelligent Software: A Software Cybernetics Perspective
Yulei Chen, Yuge Nie, Huayao Wu, Changhai Nie, William C. Chu |
COMPSAC | 3 |
| 2026 | Uncertainty and Geometric Dispersion-Driven Metamorphic Testing for DNN-Based SystemsabstractMetamorphic testing (MT) has emerged as a widely adopted technique for validating deep learning (DL) models in the absence of explicit test oracles. A key challenge in MT is the efficient selection of metamorphic groups (MGs), i.e. the source and follow-up test inputs, that are more likely to expose faults. To address this challenge, we propose UGD, a novel MT approach that integrates two complementary criteria, input uncertainty and geometric dispersion. In UGD, source inputs with high uncertainty are prioritized for testing, as these inputs are more likely to lie near decision boundaries and thereby reveal erroneous behaviors. For each selected source, a convex hull-based strategy is applied to choose follow-up inputs that are both distant from the source input and well-dispersed from each other. This design ensures that the generated MGs are diverse and fault-revealing. Extensive experiments demonstrate that UGD consistently outperforms existing baseline methods in terms of the number of violated MGs and unique faults detected, particularly on complex datasets such as ImageNet under limited test budgets. The results confirm that uncertainty is a reliable indicator for selecting fault-prone source inputs. Furthermore, geometric dispersion, guided by convex hulls, enhances fault detection by ensuring that follow-up inputs are sufficiently different from the source and diverse among themselves. Shengyou Hu, Wenyang Lyu, Huayao Wu, Xintao Niu, Changhai Nie |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2026 | How Composite Metamorphic Relations Enhance Test Effectiveness of DNN Testing: An Empirical Study
Huayao Wu, Peng Wang 0125, Shengyou Hu, Xintao Niu, Changhai Nie, Tsong Yueh Chen |
IEEE Trans. Software Eng. | 1 |
| 2025 | Cluster-Based Multi-Objective Metamorphic Test Case Pair Selection for Deep Neural NetworksabstractDue to the rapid development of deep neural networks (DNNs), ensuring their quality has become increasingly important.However, the test oracle problem poses an obstacle to DNN testing because of the massive unlabeled data.Metamorphic Testing (MT) has proven effective in alleviating the test oracle problem, and many efforts have been made to improve the cost-effectiveness of MT for DNNs.Some approaches focus on selecting good metamorphic relations (MRs), while others target the selection of suspicious source test cases.Since follow-up test cases are generated by combining source test cases with MRs, selecting effective pairs of source test cases and MRs is also quite essential and beneficial for MT.In this paper, we propose CMPS, a multi-objective black-box approach for metamorphic test case pair selection.Considering both uncertainty and diversity, CMPS aims to select pairs that can detect more unique faults in the model.It evaluates uncertainty based on model outputs and assesses diversity through clustering source test cases.Furthermore, CMPS can adaptively optimize the selection process based on feedback from the execution results of the selected pairs.We conduct extensive experiments on three datasets and five DNN models to evaluate CMPS's performance.The experimental results demonstrate that CMPS significantly outperforms baseline approaches in both failure triggering and fault detection. Jingling Wang, Shuwei Qiu, Peng Wang 0125, Jiyuan Song, Huayao Wu, Xintao Niu, Changhai Nie |
Internetware | 5 |
| 2025 | Boosting the Cost-Effectiveness of Metamorphic Test Case Pair Selection for DNN Testing with Surrogate ModelabstractWith its ability to alleviate the test oracle problem, Metamorphic Testing (MT) has been widely used to test Deep Neuron Networks (DNN). To improve failure detection ability of MT, recently, researchers have proposed uncertainty based methods to select Metamorphic test case Pairs (MPs) that are more likely to violate metamorphic relations. However, in these methods, the DNN under test needs to be frequently invoked to obtain the output probabilities of test cases for uncertainty calculation, potentially limiting their adoptions in resource-constrained test scenarios where the number of DNN calls should be minimized. To further boost the costeffectiveness of MT, in this paper, we propose MPSS, a black-box method that relies on a surrogate model to select failure-revealing MPs. In particular, MPSS aims to train and iteratively optimize a support vector machine to approximate the DNN classification boundaries in the latent space. Then, by analyzing the relative positions of both source and followup test cases of each MP to such boundaries, MPSS can effectively estimate whether the execution of this MP will lead to a metamorphic relation violation without actually calling the DNN model. Experimental results show that MPSS can increase the cost-effectiveness of MP selection by maximizing detected failures while minimizing DNN calling times under given test budgets in various situations. Jialin Fan, Jingling Wang, Shengyou Hu, Huayao Wu, Changhai Nie |
QRS | 4 |
| 2025 | Top-down: A better strategy for incremental covering array generation
Xintao Niu, Huayao Wu, Changhai Nie, Xiaoyin Wang, Jiaxi Xu |
Inf. Softw. Technol. | 3 |
| 2025 | A Systematic Literature Review on Fault Injection Testing of Microservice SystemsabstractThis paper presents the first comprehensive review of techniques that pertain to Fault Injection Testing (FIT) of Microservice systems. FIT is a popular resilience engineering technique for examining the correctness and robustness of fault-tolerance mechanisms in software systems. Despite its wide adoption in building Microservice systems of high reliability, the techniques and tools that underpin effective fault injection have not yet been systematically reviewed. To this end, a general FIT framework that consists of five key components is first summarized, with each component indicating a key design decision that should be carefully determined. Then, a systematic literature review (SLR) is performed to investigate the current practices that address the challenges associated with each of these components. Finally, the potential limitations and future research directions of FIT for Microservice systems are discussed. Senyao Yu, Huayao Wu, Xintao Niu, Changhai Nie |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | A Combinatorial Interaction Testing Method for Multi-Label Image ClassifierabstractMulti-label image classification is a critical task in computer vision, in which the correlations between labels are typically exploited by modern classifiers for an effective classification. In this study, we propose LV-CIT, a black-box testing method that applies Combinatorial Interaction Testing (CIT) to systematically test the ability of classifiers to handle such correlations. Specifically, LV-CIT views each label of the label space as an input-parameter taking binary values (indicating whether an object appears in an image), and manages to generate a label value covering array as the set of test cases to cover certain combinations of label values. Then, for each test case, LV-CIT relies on an object library to generate composite test images that perfectly match the specified labels, and reports classification errors if such labels cannot be correctly recognised. The experimental results on two popular datasets with six state-of-the-art image classifiers show that LV-CIT is more efficient than the existing CIT tools in generating label value covering arrays. LV-CIT is also effective in errors revelation, as it can find 111% more errors by using 20% fewer test images than the existing methods for testing multi-label image classifiers. Peng Wang 0125, Shengyou Hu, Huayao Wu, Xintao Niu, Changhai Nie |
ISSRE | 3 |
| 2024 | A method of multidimensional software aging prediction based on ensemble learning: A case of Android OS
Yuge Nie, Yulei Chen, Yujia Jiang, Huayao Wu, Beibei Yin, Kai-Yuan Cai |
Inf. Softw. Technol. | 4 |
| 2023 | ATOM: Automated Black-Box Testing of Multi-Label Image Classification SystemsabstractMulti-label Image Classification Systems (MICSs) developed based on Deep Neural Networks (DNNs) are extensively used in people's daily life. Currently, although there are a variety of approaches to test DNN-based systems, they typically rely on the internals of DNNs to design test cases, and do not take the core specification of MICS (i.e., correctly recognizing multiple objects in a given image) into account. In this paper, we propose ATOM, an automated and systematic black-box testing framework for testing MICS. Specifically, ATOM exploits the label combination as the testing adequacy criteria, hoping to systematically examine the impact of correlations between a fixed number of labels on the classification ability of MICS. Then, ATOM leverages image search engine and natural language processing to find test images that are not only common to the real-world, but also relevant to target label combinations. Finally, ATOM combines metamorphic testing and label information to realize test oracle identification, based on which the ability of MICS in classifying different label combinations is evaluated. To evaluate the effectiveness of ATOM, we have performed experiments on two popular datasets of MICS, VOC and COCO (each with five state-of-the-art DNN models), and one real-world photo tagging application from our industrial partner. The experimental results reveal that the performance of current DNN-based MICSs remains less satisfactory even in recognizing correlations between only two labels, as ATOM triggers a total number of 6,049 such label combination related errors for all MICSs studied. In particular, ATOM reports 587 error-revealing images for the industrial MICS, in which 92% of them are confirmed by the developers. Shengyou Hu, Huayao Wu, Peng Wang 0125, Yongjun Tu, Xiu Jiang, Xintao Niu, Changhai Nie |
ASE | 2 |
| 2023 | An Empirical Study to Identify Software Aging Indicators for Android OSabstractAndroid mobile devices have been suffering from performance degradation and increased failure rates during long-term operation, known as software aging. With the major changes in performance optimization and resource management in Android, it is spotted that the aging behavior of Android devices in the 2020s differs significantly from previous studies in resource utilization and performance metrics, which makes some classic metrics difficult to measure aging well, and new metrics are required to better describe the new phenomenon. Thus, we propose thread- and interface-level metrics to portray aging at a finer granularity and conduct an empirical study to reidentify classic and new software aging metrics in Android. Analysis confirms that software aging in Android is less reflected in global resources metrics but in more fine-grained ones, so thread- and interface-level metrics combined with specific classic resource and process-level metrics are helpful as indicators of software aging. These metrics have been confirmed and deployed for aging monitoring by our mobile phone manufacturer collaborators. A new experimental method customized for metric studies has also been adopted in this paper, significantly reducing data costs and interference in measurements. Yulei Chen, Yuge Nie, Beibei Yin, Zheng Zheng 0001, Huayao Wu |
QRS | 5 |
| 2023 | Enhancing Fault Injection Testing of Service Systems via Fault-Tolerance BottleneckabstractModern large-scale service systems are usually deployed with redundant components to ensure high dependability in distributed and volatile environments. Fault Injection Testing (FIT) is a popular technique for testing such systems, while the application of FIT to validating the correctness of redundant components remains a challenging task, especially when the system's structural information is unavailable when testing starts. In this study, we refer to a minimum set of faults that, when injected, will cut off all execution paths in a service system as afault-tolerance bottleneck, and we propose a novel Fault-tolerance Bottleneck driven Fault Injection (FBFI) approach to the exploration and validation of redundant components without prior knowledge of the system's business structure. The core idea of FBFI is to iteratively infer and inject bottlenecks of the business structure constructed so far. In this way, FBFI is able to discover and test redundant components by repeatedly triggering new system behaviors. The effectiveness and efficiency of FBFI is evaluated using two microservice benchmark systems with different deployment scales. The results reveal that FBFI is more practical and cost-effective than random and lineage-driven FIT approaches in testing service systems of high redundancy levels. Huayao Wu, Senyao Yu, Xintao Niu, Changhai Nie, Yu Pei 0001, Qiang He 0001, Yun Yang 0001 |
IEEE Trans. Software Eng. | 1 |
| 2022 | Combinatorial Testing of RESTful APIsabstractThis paper presents RestCT, a systematic and fully automatic approach that adopts Combinatorial Testing (CT) to test RESTful APIs. RestCT is systematic in that it covers and tests not only the interactions of a certain number of operations in RESTful APIs, but also the interactions of particular input-parameters in every single operation. This is realised by a novel two-phase test case generation approach, which first generates a constrained sequence covering array to determine the execution orders of available operations, and then applies an adaptive strategy to generate and refine several constrained covering arrays to concretise input-parameters of each operation. RestCT is also automatic in that its application relies on only a given Swagger specification of RESTful APIs. The creation of CT test models (especially, the inferring of dependency relationships in both operations and input-parameters), and the generation and execution of test cases are performed without any human intervention. Experimental results on 11 real-world RESTful APIs demonstrate the effectiveness and efficiency of RestCT. In particular, RestCT can find eight new bugs, where only one of them can be triggered by the state-of-the-art testing tool of RESTful APIs. Huayao Wu, Xintao Niu, Changhai Nie |
ICSE | 1 |
| 2022 | An Adaptive Penalty based Parallel Tabu Search for Constrained Covering Array Generation
Huayao Wu, Xintao Niu, Changhai Nie, Jiaxi Xu |
Inf. Softw. Technol. | 2 |
| 2022 | Enhance Combinatorial Testing With Metamorphic RelationsabstractDue to the effectiveness and efficiency in detecting defects caused by interactions of multiple factors, Combinatorial Testing (CT) has received considerable scholarly attention in the last decades. Despite numerous practical test case generation techniques being developed, there remains a paucity of studies addressing the automated oracle generation problem, which holds back the overall automation of CT. As a consequence, much human intervention is inevitable, which is time-consuming and error-prone. This costly manual task also restricts the application of higher testing strength, inhibiting the full exploitation of CT in the industrial practice. To bridge the gap between test designs and fully automated test flows, and to extend the applicability of CT, this paper presents a novel CT methodology, named COMER, to enhance the traditional CT by accounting for Metamorphic Relations (MRs). COMER puts a high priority on generating pairs of test cases which match the input rules of MRs, i.e., the Metamorphic Group (MG), such that the correctness can be automatically determined by verifying whether the outputs of these test cases violate their MRs. As a result, COMER can not only satisfy the t-way coverage as what CT does, but also automatically check test oracle as many violations as possible. Several empirical studies conducted on 31 real-world software projects have shown that COMER increased the number of metamorphic groups by an average factor of 75.9 and also increased the failure detection rate by an average factor of 11.3, when compared with CT, while the overall number of test cases generated by COMER barely increased. Xintao Niu, Yanjie Sun, Huayao Wu, Changhai Nie, Yu Lei 0001, Xiaoyin Wang |
IEEE Trans. Software Eng. | 3 |
| 2022 | A Theory of Pending Schemas in Combinatorial TestingabstractCombinatorial Testing (CT) is an effective testing technique for detecting failures which are triggered by the interactions of various factors that influence the behaviour of a system. Although many studies in CT have designed elaborate test suites (called covering arrays) to systemically check each possible factor interaction, they provide weak support to locate the concrete failure-inducing interactions, i.e., the Minimal Failure-causing Schemas (MFS). To this end, a variety of MFS identification approaches have been proposed. However, as this study reveals, these approaches suffer from various issues such as cannot identify multiple overlapping MFSs, cannot handle MFSs with high degrees, cannot be applied to systems with large number of parameters, etc. These issues are essentially caused by the exponential computing complexity of checking every interaction in the test cases. Therefore, they can only focus on a subset of all the possible interactions, resulting in many interactions unnoticed. Ignoring these unnoticed interactions could potentially cause failures that have never been systematically checked. Hence, it is beneficial for MFS identification approaches to identify these interactions. In order to account for these unnoticed interactions in CT, this study introduces the notion of pending schema, based on which a theoretical framework of CT schemas is established. In particular, we formally define the determinability of a schema in CT with respect to given information; as such, the yet-to-be determined schemas are exactly the pending schemas. The relationships between the different schemas (faulty, healthy, and pending) and test cases are also theoretically analyzed. Based on which, we further propose three formulas, along with three corresponding algorithms, for the identification of the pending schemas in failing test cases, and formally prove their correctness. As a result, we reduce the complexity of obtaining pending schemas with respect to the number of factors that may have influences on the software. Xintao Niu, Huayao Wu, Changhai Nie, Yu Lei 0001, Xiaoyin Wang |
IEEE Trans. Software Eng. | 2 |
| 2021 | Identifying Key Features from App User ReviewsabstractDue to the rapid growth and strong competition of mobile application (app) market, app developers should not only offer users with attractive new features, but also carefully maintain and improve existing features based on users' feedbacks. User reviews indicate a rich source of information to plan such feature maintenance activities, and it could be of great benefit for developers to evaluate and magnify the contribution of specific features to the overall success of their apps. In this study, we refer to the features that are highly correlated to app ratings as key features, and we present KEFE, a novel approach that leverages app description and user reviews to identify key features of a given app. The application of KEFE especially relies on natural language processing, deep machine learning classifier, and regression analysis technique, which involves three main steps: 1) extracting feature-describing phrases from app description; 2) matching each app feature with its relevant user reviews; and 3) building a regression model to identify features that have significant relationships with app ratings. To train and evaluate KEFE, we collect 200 app descriptions and 1,108,148 user reviews from Chinese Apple App Store. Experimental results demonstrate the effectiveness of KEFE in feature extraction, where an average F-measure of 78.13% is achieved. The key features identified are also likely to provide hints for successful app releases, as for the releases that receive higher app ratings, 70% of features improvements are related to key features. Huayao Wu, Wenjun Deng, Xintao Niu, Changhai Nie |
ICSE | 1 |
| 2021 | Comparative Analysis of Constraint Handling Techniques for Constrained Combinatorial TestingabstractConstraints depict the dependency relationships between parameters in a software system under test. Because almost all systems are constrained in some way, techniques that adequately cater for constraints have become a crucial factor for adoption, deployment and exploitation of Combinatorial Testing (CT). Currently, despite a variety of different constraint handling techniques available, the relationship between these techniques and the generation algorithms that use them remains unknown, yielding an important gap and pressing concern in the literature of constrained combination testing. In this article, we present a comparative empirical study to investigate the impact of four common constraint handling techniques on the efficiency of six representative (greedy and search-based) test suite generation algorithms. The results reveal that theVerifytechnique implemented with the Minimal Forbidden Tuple (MFT) approach is the fastest, while theReplacetechnique is promising for producing the smallest constrained covering arrays, especially for algorithms that construct test cases one-at-a-time. The results also show that there is an interplay between efficiency of the constraint handler and the test suite generation algorithm into which it is developed. Huayao Wu, Changhai Nie, Justyna Petke, Yue Jia 0001, Mark Harman |
IEEE Trans. Software Eng. | 1 |
| 2020 | An Empirical Comparison of Combinatorial Testing, Random Testing and Adaptive Random TestingabstractWe present an empirical comparison of three test generation techniques, namely, Combinatorial Testing (CT), Random Testing (RT) and Adaptive Random Testing (ART), under different test scenarios. This is the first study in the literature to account for the (more realistic) testing setting in which the tester may not have complete information about the parameters and constraints that pertain to the system, and to account for the challenge posed by faults (in terms of failure rate). Our study was conducted on nine real-world programs under a total of 1683 test scenarios (combinations of available parameter and constraint information and failure rate). The results show significant differences in the techniques' fault detection ability when faults are hard to detect (failure rates are relatively low). CT performs best overall; no worse than any other in 98 percent of scenarios studied. ART enhances RT, and is comparable to CT in 96 percent of scenarios, but its computational cost can be up to 3.5 times higher than CT when the program is highly constrained. Additionally, when constraint information is unavailable for a highly-constrained program, a large random test suite is as effective as CT or ART, yet its computational cost of test generation is significantly lower than that of other techniques. Huayao Wu, Changhai Nie, Justyna Petke, Yue Jia 0001, Mark Harman |
IEEE Trans. Software Eng. | 1 |
| 2016 | The optimal testing order in the presence of switching cost
Huayao Wu, Changhai Nie, Fei-Ching Kuo |
Inf. Softw. Technol. | 1 |
| 2015 | Combinatorial testing, random testing, and adaptive random testing for detecting interaction triggered failures
Changhai Nie, Huayao Wu, Xintao Niu, Fei-Ching Kuo, Hareton K. N. Leung, Charles J. Colbourn |
Inf. Softw. Technol. | 2 |
| 2015 | A Discrete Particle Swarm Optimization for Covering Array GenerationabstractSoftware behavior depends on many factors. Combinatorial testing (CT) aims to generate small sets of test cases to uncover defects caused by those factors and their interactions. Covering array generation, a discrete optimization problem, is the most popular research area in the field of CT. Particle swarm optimization (PSO), an evolutionary search-based heuristic technique, has succeeded in generating covering arrays that are competitive in size. However, current PSO methods for covering array generation simply round the particle's position to an integer to handle the discrete search space. Moreover, no guidelines are available to effectively set PSOs parameters for this problem. In this paper, we extend the set-based PSO, an existing discrete PSO (DPSO) method, to covering array generation. Two auxiliary strategies (particle reinitialization and additional evaluation of gbest) are proposed to improve performance, and thus a novel DPSO for covering array generation is developed. Guidelines for parameter settings both for conventional PSO (CPSO) and for DPSO are developed systematically here. Discrete extensions of four existing PSO variants are developed, in order to further investigate the effectiveness of DPSO for covering array generation. Experiments show that CPSO can produce better results using the guidelines for parameter settings, and that DPSO can generate smaller covering arrays than CPSO and other existing evolutionary algorithms. DPSO is a promising improvement on PSO for covering array generation. Huayao Wu, Changhai Nie, Fei-Ching Kuo, Hareton K. N. Leung, Charles J. Colbourn |
IEEE Trans. Evol. Comput. | 1 |
| 2012 | Search Based Combinatorial TestingabstractSearch techniques can dramatically change our ability to solve a host of problems in applied science and engineering, many search techniques have been developed and applied successfully in many fields, including search based software engineering (SBSE). As a key problem of combinatorial testing, covering array generation has been widely studied and many search techniques have been applied which can be named as search based combinatorial testing (SBCT). SBCT is a branch of search based software testing (SBST) within SBSE. In this paper, to explore the applicability and effectiveness of SBCT, we design six variants from existing search algorithms: Genetic Algorithm, Particle Swarm Optimization and Ant Colony Algorithm by reversing and randomizing their mechanisms. We study their effectiveness in terms of generating a covering array and compare their performance. Experiments show that these search techniques can work well with distinct performance in covering array generation. We believe that these search techniques can be further improved by fine-tuning their configuration and used in broad ranges of area. Changhai Nie, Huayao Wu, Yalan Liang, Hareton K. N. Leung, Fei-Ching Kuo, Zheng Li 0002 |
APSEC | 2 |