EDBT 2026 Demo / reviewers in the wild / expert
Xintao Niu
dblp:134/8989
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-5786-0894ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 4 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty and Geometric Dispersion-Driven Metamorphic Testing for DNN-Based SystemsabstractMetamorphic testing (MT) has emerged as a widely adopted technique for validating deep learning (DL) models in the absence of explicit test oracles. A key challenge in MT is the efficient selection of metamorphic groups (MGs), i.e. the source and follow-up test inputs, that are more likely to expose faults. To address this challenge, we propose UGD, a novel MT approach that integrates two complementary criteria, input uncertainty and geometric dispersion. In UGD, source inputs with high uncertainty are prioritized for testing, as these inputs are more likely to lie near decision boundaries and thereby reveal erroneous behaviors. For each selected source, a convex hull-based strategy is applied to choose follow-up inputs that are both distant from the source input and well-dispersed from each other. This design ensures that the generated MGs are diverse and fault-revealing. Extensive experiments demonstrate that UGD consistently outperforms existing baseline methods in terms of the number of violated MGs and unique faults detected, particularly on complex datasets such as ImageNet under limited test budgets. The results confirm that uncertainty is a reliable indicator for selecting fault-prone source inputs. Furthermore, geometric dispersion, guided by convex hulls, enhances fault detection by ensuring that follow-up inputs are sufficiently different from the source and diverse among themselves. Shengyou Hu, Wenyang Lyu, Huayao Wu, Xintao Niu, Changhai Nie |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2026 | How Composite Metamorphic Relations Enhance Test Effectiveness of DNN Testing: An Empirical Study
Huayao Wu, Peng Wang 0125, Shengyou Hu, Xintao Niu, Changhai Nie, Tsong Yueh Chen |
IEEE Trans. Software Eng. | 4 |
| 2025 | Cluster-Based Multi-Objective Metamorphic Test Case Pair Selection for Deep Neural NetworksabstractDue to the rapid development of deep neural networks (DNNs), ensuring their quality has become increasingly important.However, the test oracle problem poses an obstacle to DNN testing because of the massive unlabeled data.Metamorphic Testing (MT) has proven effective in alleviating the test oracle problem, and many efforts have been made to improve the cost-effectiveness of MT for DNNs.Some approaches focus on selecting good metamorphic relations (MRs), while others target the selection of suspicious source test cases.Since follow-up test cases are generated by combining source test cases with MRs, selecting effective pairs of source test cases and MRs is also quite essential and beneficial for MT.In this paper, we propose CMPS, a multi-objective black-box approach for metamorphic test case pair selection.Considering both uncertainty and diversity, CMPS aims to select pairs that can detect more unique faults in the model.It evaluates uncertainty based on model outputs and assesses diversity through clustering source test cases.Furthermore, CMPS can adaptively optimize the selection process based on feedback from the execution results of the selected pairs.We conduct extensive experiments on three datasets and five DNN models to evaluate CMPS's performance.The experimental results demonstrate that CMPS significantly outperforms baseline approaches in both failure triggering and fault detection. Jingling Wang, Shuwei Qiu, Peng Wang 0125, Jiyuan Song, Huayao Wu, Xintao Niu, Changhai Nie |
Internetware | 6 |
| 2025 | Top-down: A better strategy for incremental covering array generation
Xintao Niu, Huayao Wu, Changhai Nie, Xiaoyin Wang, Jiaxi Xu |
Inf. Softw. Technol. | 2 |
| 2025 | A Systematic Literature Review on Fault Injection Testing of Microservice SystemsabstractThis paper presents the first comprehensive review of techniques that pertain to Fault Injection Testing (FIT) of Microservice systems. FIT is a popular resilience engineering technique for examining the correctness and robustness of fault-tolerance mechanisms in software systems. Despite its wide adoption in building Microservice systems of high reliability, the techniques and tools that underpin effective fault injection have not yet been systematically reviewed. To this end, a general FIT framework that consists of five key components is first summarized, with each component indicating a key design decision that should be carefully determined. Then, a systematic literature review (SLR) is performed to investigate the current practices that address the challenges associated with each of these components. Finally, the potential limitations and future research directions of FIT for Microservice systems are discussed. Senyao Yu, Huayao Wu, Xintao Niu, Changhai Nie |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | A Combinatorial Interaction Testing Method for Multi-Label Image ClassifierabstractMulti-label image classification is a critical task in computer vision, in which the correlations between labels are typically exploited by modern classifiers for an effective classification. In this study, we propose LV-CIT, a black-box testing method that applies Combinatorial Interaction Testing (CIT) to systematically test the ability of classifiers to handle such correlations. Specifically, LV-CIT views each label of the label space as an input-parameter taking binary values (indicating whether an object appears in an image), and manages to generate a label value covering array as the set of test cases to cover certain combinations of label values. Then, for each test case, LV-CIT relies on an object library to generate composite test images that perfectly match the specified labels, and reports classification errors if such labels cannot be correctly recognised. The experimental results on two popular datasets with six state-of-the-art image classifiers show that LV-CIT is more efficient than the existing CIT tools in generating label value covering arrays. LV-CIT is also effective in errors revelation, as it can find 111% more errors by using 20% fewer test images than the existing methods for testing multi-label image classifiers. Peng Wang 0125, Shengyou Hu, Huayao Wu, Xintao Niu, Changhai Nie |
ISSRE | 4 |
| 2023 | ATOM: Automated Black-Box Testing of Multi-Label Image Classification SystemsabstractMulti-label Image Classification Systems (MICSs) developed based on Deep Neural Networks (DNNs) are extensively used in people's daily life. Currently, although there are a variety of approaches to test DNN-based systems, they typically rely on the internals of DNNs to design test cases, and do not take the core specification of MICS (i.e., correctly recognizing multiple objects in a given image) into account. In this paper, we propose ATOM, an automated and systematic black-box testing framework for testing MICS. Specifically, ATOM exploits the label combination as the testing adequacy criteria, hoping to systematically examine the impact of correlations between a fixed number of labels on the classification ability of MICS. Then, ATOM leverages image search engine and natural language processing to find test images that are not only common to the real-world, but also relevant to target label combinations. Finally, ATOM combines metamorphic testing and label information to realize test oracle identification, based on which the ability of MICS in classifying different label combinations is evaluated. To evaluate the effectiveness of ATOM, we have performed experiments on two popular datasets of MICS, VOC and COCO (each with five state-of-the-art DNN models), and one real-world photo tagging application from our industrial partner. The experimental results reveal that the performance of current DNN-based MICSs remains less satisfactory even in recognizing correlations between only two labels, as ATOM triggers a total number of 6,049 such label combination related errors for all MICSs studied. In particular, ATOM reports 587 error-revealing images for the industrial MICS, in which 92% of them are confirmed by the developers. Shengyou Hu, Huayao Wu, Peng Wang 0125, Yongjun Tu, Xiu Jiang, Xintao Niu, Changhai Nie |
ASE | 7 |
| 2023 | Toward More Efficient Statistical Debugging with Abstraction RefinementabstractDebugging is known to be a notoriously painstaking and time-consuming task. As one major family of automated debugging, statistical debugging approaches have been well investigated over the past decade, which collect failing and passing executions and apply statistical techniques to identify discriminative elements as potential bug causes. Most of the existing approaches instrument the entire program to produce execution profiles for debugging, thus incurring hefty instrumentation and analysis cost. However, as in fact a major part of the program code is error-free, full-scale program instrumentation is wasteful and unnecessary. This article presents a systematic abstraction refinement-based pruning technique for statistical debugging. Our technique only needs to instrument and analyze the code partially. While guided by a mathematically rigorous analysis, our technique is guaranteed to produce the same debugging results as an exhaustive analysis in deterministic settings. With the help of the effective and safe pruning, our technique greatly saves the cost of failure diagnosis without sacrificing any debugging capability. We apply this technique to two different statistical debugging scenarios: in-house and production-run statistical debugging. The comprehensive evaluations validate that our technique can significantly improve the efficiency of statistical debugging in both scenarios, while without jeopardizing the debugging capability. Zhiqiang Zuo 0002, Xintao Niu, Siyi Zhang 0011, Lu Fang 0003, Siau-Cheng Khoo, Shan Lu 0001, Chengnian Sun, Guoqing Harry Xu |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Enhancing Fault Injection Testing of Service Systems via Fault-Tolerance BottleneckabstractModern large-scale service systems are usually deployed with redundant components to ensure high dependability in distributed and volatile environments. Fault Injection Testing (FIT) is a popular technique for testing such systems, while the application of FIT to validating the correctness of redundant components remains a challenging task, especially when the system's structural information is unavailable when testing starts. In this study, we refer to a minimum set of faults that, when injected, will cut off all execution paths in a service system as afault-tolerance bottleneck, and we propose a novel Fault-tolerance Bottleneck driven Fault Injection (FBFI) approach to the exploration and validation of redundant components without prior knowledge of the system's business structure. The core idea of FBFI is to iteratively infer and inject bottlenecks of the business structure constructed so far. In this way, FBFI is able to discover and test redundant components by repeatedly triggering new system behaviors. The effectiveness and efficiency of FBFI is evaluated using two microservice benchmark systems with different deployment scales. The results reveal that FBFI is more practical and cost-effective than random and lineage-driven FIT approaches in testing service systems of high redundancy levels. Huayao Wu, Senyao Yu, Xintao Niu, Changhai Nie, Yu Pei 0001, Qiang He 0001, Yun Yang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2022 | Combinatorial Testing of RESTful APIsabstractThis paper presents RestCT, a systematic and fully automatic approach that adopts Combinatorial Testing (CT) to test RESTful APIs. RestCT is systematic in that it covers and tests not only the interactions of a certain number of operations in RESTful APIs, but also the interactions of particular input-parameters in every single operation. This is realised by a novel two-phase test case generation approach, which first generates a constrained sequence covering array to determine the execution orders of available operations, and then applies an adaptive strategy to generate and refine several constrained covering arrays to concretise input-parameters of each operation. RestCT is also automatic in that its application relies on only a given Swagger specification of RESTful APIs. The creation of CT test models (especially, the inferring of dependency relationships in both operations and input-parameters), and the generation and execution of test cases are performed without any human intervention. Experimental results on 11 real-world RESTful APIs demonstrate the effectiveness and efficiency of RestCT. In particular, RestCT can find eight new bugs, where only one of them can be triggered by the state-of-the-art testing tool of RESTful APIs. Huayao Wu, Xintao Niu, Changhai Nie |
ICSE | 3 |
| 2022 | An Adaptive Penalty based Parallel Tabu Search for Constrained Covering Array Generation
Huayao Wu, Xintao Niu, Changhai Nie, Jiaxi Xu |
Inf. Softw. Technol. | 3 |
| 2022 | Enhance Combinatorial Testing With Metamorphic RelationsabstractDue to the effectiveness and efficiency in detecting defects caused by interactions of multiple factors, Combinatorial Testing (CT) has received considerable scholarly attention in the last decades. Despite numerous practical test case generation techniques being developed, there remains a paucity of studies addressing the automated oracle generation problem, which holds back the overall automation of CT. As a consequence, much human intervention is inevitable, which is time-consuming and error-prone. This costly manual task also restricts the application of higher testing strength, inhibiting the full exploitation of CT in the industrial practice. To bridge the gap between test designs and fully automated test flows, and to extend the applicability of CT, this paper presents a novel CT methodology, named COMER, to enhance the traditional CT by accounting for Metamorphic Relations (MRs). COMER puts a high priority on generating pairs of test cases which match the input rules of MRs, i.e., the Metamorphic Group (MG), such that the correctness can be automatically determined by verifying whether the outputs of these test cases violate their MRs. As a result, COMER can not only satisfy the t-way coverage as what CT does, but also automatically check test oracle as many violations as possible. Several empirical studies conducted on 31 real-world software projects have shown that COMER increased the number of metamorphic groups by an average factor of 75.9 and also increased the failure detection rate by an average factor of 11.3, when compared with CT, while the overall number of test cases generated by COMER barely increased. Xintao Niu, Yanjie Sun, Huayao Wu, Changhai Nie, Yu Lei 0001, Xiaoyin Wang |
IEEE Trans. Software Eng. | 1 |
| 2022 | A Theory of Pending Schemas in Combinatorial TestingabstractCombinatorial Testing (CT) is an effective testing technique for detecting failures which are triggered by the interactions of various factors that influence the behaviour of a system. Although many studies in CT have designed elaborate test suites (called covering arrays) to systemically check each possible factor interaction, they provide weak support to locate the concrete failure-inducing interactions, i.e., the Minimal Failure-causing Schemas (MFS). To this end, a variety of MFS identification approaches have been proposed. However, as this study reveals, these approaches suffer from various issues such as cannot identify multiple overlapping MFSs, cannot handle MFSs with high degrees, cannot be applied to systems with large number of parameters, etc. These issues are essentially caused by the exponential computing complexity of checking every interaction in the test cases. Therefore, they can only focus on a subset of all the possible interactions, resulting in many interactions unnoticed. Ignoring these unnoticed interactions could potentially cause failures that have never been systematically checked. Hence, it is beneficial for MFS identification approaches to identify these interactions. In order to account for these unnoticed interactions in CT, this study introduces the notion of pending schema, based on which a theoretical framework of CT schemas is established. In particular, we formally define the determinability of a schema in CT with respect to given information; as such, the yet-to-be determined schemas are exactly the pending schemas. The relationships between the different schemas (faulty, healthy, and pending) and test cases are also theoretically analyzed. Based on which, we further propose three formulas, along with three corresponding algorithms, for the identification of the pending schemas in failing test cases, and formally prove their correctness. As a result, we reduce the complexity of obtaining pending schemas with respect to the number of factors that may have influences on the software. Xintao Niu, Huayao Wu, Changhai Nie, Yu Lei 0001, Xiaoyin Wang |
IEEE Trans. Software Eng. | 1 |
| 2021 | Identifying Key Features from App User ReviewsabstractDue to the rapid growth and strong competition of mobile application (app) market, app developers should not only offer users with attractive new features, but also carefully maintain and improve existing features based on users' feedbacks. User reviews indicate a rich source of information to plan such feature maintenance activities, and it could be of great benefit for developers to evaluate and magnify the contribution of specific features to the overall success of their apps. In this study, we refer to the features that are highly correlated to app ratings as key features, and we present KEFE, a novel approach that leverages app description and user reviews to identify key features of a given app. The application of KEFE especially relies on natural language processing, deep machine learning classifier, and regression analysis technique, which involves three main steps: 1) extracting feature-describing phrases from app description; 2) matching each app feature with its relevant user reviews; and 3) building a regression model to identify features that have significant relationships with app ratings. To train and evaluate KEFE, we collect 200 app descriptions and 1,108,148 user reviews from Chinese Apple App Store. Experimental results demonstrate the effectiveness of KEFE in feature extraction, where an average F-measure of 78.13% is achieved. The key features identified are also likely to provide hints for successful app releases, as for the releases that receive higher app ratings, 70% of features improvements are related to key features. Huayao Wu, Wenjun Deng, Xintao Niu, Changhai Nie |
ICSE | 3 |
| 2020 | Identifying Failure-Causing Schemas in the Presence of Multiple FaultsabstractCombinatorial testing (CT) has been proven effective in revealing the failures caused by the interaction of factors that affect the behavior of a system. The theory of Minimal Failure-Causing Schema (MFS) has been proposed to isolate the cause of a failure after CT. Most algorithms that aim to identify MFS focus on handling a single fault in the System Under Test (SUT). However, we argue that multiple faults are more common in practice, under which masking effects may be triggered so that some failures cannot be observed. The traditional MFS theory lacks a mechanism to handle such effects; hence, they may incorrectly isolate the MFS. To address this problem, we propose a new MFS model that takes into account multiple faults. We first formally analyze the impact of the multiple faults on existing MFS identifying algorithms, especially in situations where masking effects are triggered by multiple faults. We then develop an approach that can assist traditional algorithms to better handle multiple faults. Empirical studies were conducted using several kinds of open-source software, which showed that multiple faults with masking effects do negatively affect traditional MFS identifying approaches and that our approach can help to alleviate these effects. Xintao Niu, Changhai Nie, Yu Lei 0001, Hareton K. N. Leung, Xiaoyin Wang |
IEEE Trans. Software Eng. | 1 |
| 2020 | An Interleaving Approach to Combinatorial Testing and Failure-Inducing Interaction IdentificationabstractCombinatorial testing (CT) seeks to detect potential faults caused by various interactions of factors that can influence the software systems. When applying CT, it is a common practice to first generate a set of test cases to cover each possible interaction and then to identify the failure-inducing interaction after a failure is detected. Although this conventional procedure is simple and forthright, we conjecture that it is not the ideal choice in practice. This is because 1) testers desire to identify the root cause of failures before all the needed test cases are generated and executed 2) the early identified failure-inducing interactions can guide the remaining test case generation so that many unnecessary and invalid test cases can be avoided. For these reasons, we propose a novel CT framework that allows both generation and identification process to interact with each other. As a result, both generation and identification stages will be done more effectively and efficiently. We conducted a series of empirical studies on several open-source software, the results of which show that our framework can identify the failure-inducing interactions more quickly than traditional approaches while requiring fewer test cases. Xintao Niu, Changhai Nie, Hareton K. N. Leung, Yu Lei 0001, Xiaoyin Wang, Jiaxi Xu |
IEEE Trans. Software Eng. | 1 |
| 2015 | Combinatorial testing, random testing, and adaptive random testing for detecting interaction triggered failures
Changhai Nie, Huayao Wu, Xintao Niu, Fei-Ching Kuo, Hareton K. N. Leung, Charles J. Colbourn |
Inf. Softw. Technol. | 3 |