Shengyou Hu

dblp:360/7230 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0007-9995-1496ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Uncertainty and Geometric Dispersion-Driven Metamorphic Testing for DNN-Based Systems
abstract
Metamorphic testing (MT) has emerged as a widely adopted technique for validating deep learning (DL) models in the absence of explicit test oracles. A key challenge in MT is the efficient selection of metamorphic groups (MGs), i.e. the source and follow-up test inputs, that are more likely to expose faults. To address this challenge, we propose UGD, a novel MT approach that integrates two complementary criteria, input uncertainty and geometric dispersion. In UGD, source inputs with high uncertainty are prioritized for testing, as these inputs are more likely to lie near decision boundaries and thereby reveal erroneous behaviors. For each selected source, a convex hull-based strategy is applied to choose follow-up inputs that are both distant from the source input and well-dispersed from each other. This design ensures that the generated MGs are diverse and fault-revealing. Extensive experiments demonstrate that UGD consistently outperforms existing baseline methods in terms of the number of violated MGs and unique faults detected, particularly on complex datasets such as ImageNet under limited test budgets. The results confirm that uncertainty is a reliable indicator for selecting fault-prone source inputs. Furthermore, geometric dispersion, guided by convex hulls, enhances fault detection by ensuring that follow-up inputs are sufficiently different from the source and diverse among themselves.
Shengyou Hu, Wenyang Lyu, Huayao Wu, Xintao Niu, Changhai Nie
Int. J. Softw. Eng. Knowl. Eng.1
2026 How Composite Metamorphic Relations Enhance Test Effectiveness of DNN Testing: An Empirical Study
Huayao Wu, Peng Wang 0125, Shengyou Hu, Xintao Niu, Changhai Nie, Tsong Yueh Chen
IEEE Trans. Software Eng.3
2025 Boosting the Cost-Effectiveness of Metamorphic Test Case Pair Selection for DNN Testing with Surrogate Model
abstract
With its ability to alleviate the test oracle problem, Metamorphic Testing (MT) has been widely used to test Deep Neuron Networks (DNN). To improve failure detection ability of MT, recently, researchers have proposed uncertainty based methods to select Metamorphic test case Pairs (MPs) that are more likely to violate metamorphic relations. However, in these methods, the DNN under test needs to be frequently invoked to obtain the output probabilities of test cases for uncertainty calculation, potentially limiting their adoptions in resource-constrained test scenarios where the number of DNN calls should be minimized. To further boost the costeffectiveness of MT, in this paper, we propose MPSS, a black-box method that relies on a surrogate model to select failure-revealing MPs. In particular, MPSS aims to train and iteratively optimize a support vector machine to approximate the DNN classification boundaries in the latent space. Then, by analyzing the relative positions of both source and followup test cases of each MP to such boundaries, MPSS can effectively estimate whether the execution of this MP will lead to a metamorphic relation violation without actually calling the DNN model. Experimental results show that MPSS can increase the cost-effectiveness of MP selection by maximizing detected failures while minimizing DNN calling times under given test budgets in various situations.
Jialin Fan, Jingling Wang, Shengyou Hu, Huayao Wu, Changhai Nie
QRS3
2024 A Combinatorial Interaction Testing Method for Multi-Label Image Classifier
abstract
Multi-label image classification is a critical task in computer vision, in which the correlations between labels are typically exploited by modern classifiers for an effective classification. In this study, we propose LV-CIT, a black-box testing method that applies Combinatorial Interaction Testing (CIT) to systematically test the ability of classifiers to handle such correlations. Specifically, LV-CIT views each label of the label space as an input-parameter taking binary values (indicating whether an object appears in an image), and manages to generate a label value covering array as the set of test cases to cover certain combinations of label values. Then, for each test case, LV-CIT relies on an object library to generate composite test images that perfectly match the specified labels, and reports classification errors if such labels cannot be correctly recognised. The experimental results on two popular datasets with six state-of-the-art image classifiers show that LV-CIT is more efficient than the existing CIT tools in generating label value covering arrays. LV-CIT is also effective in errors revelation, as it can find 111% more errors by using 20% fewer test images than the existing methods for testing multi-label image classifiers.
Peng Wang 0125, Shengyou Hu, Huayao Wu, Xintao Niu, Changhai Nie
ISSRE2
2023 ATOM: Automated Black-Box Testing of Multi-Label Image Classification Systems
abstract
Multi-label Image Classification Systems (MICSs) developed based on Deep Neural Networks (DNNs) are extensively used in people's daily life. Currently, although there are a variety of approaches to test DNN-based systems, they typically rely on the internals of DNNs to design test cases, and do not take the core specification of MICS (i.e., correctly recognizing multiple objects in a given image) into account. In this paper, we propose ATOM, an automated and systematic black-box testing framework for testing MICS. Specifically, ATOM exploits the label combination as the testing adequacy criteria, hoping to systematically examine the impact of correlations between a fixed number of labels on the classification ability of MICS. Then, ATOM leverages image search engine and natural language processing to find test images that are not only common to the real-world, but also relevant to target label combinations. Finally, ATOM combines metamorphic testing and label information to realize test oracle identification, based on which the ability of MICS in classifying different label combinations is evaluated. To evaluate the effectiveness of ATOM, we have performed experiments on two popular datasets of MICS, VOC and COCO (each with five state-of-the-art DNN models), and one real-world photo tagging application from our industrial partner. The experimental results reveal that the performance of current DNN-based MICSs remains less satisfactory even in recognizing correlations between only two labels, as ATOM triggers a total number of 6,049 such label combination related errors for all MICSs studied. In particular, ATOM reports 587 error-revealing images for the industrial MICS, in which 92% of them are confirmed by the developers.
Shengyou Hu, Huayao Wu, Peng Wang 0125, Yongjun Tu, Xiu Jiang, Xintao Niu, Changhai Nie
ASE1