Zohreh Aghababaeyan

dblp:309/7615 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-9375-4095ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 DiffGAN: A Test Generation Approach for Differential Testing of Deep Neural Networks for Image Analysis
abstract
Deep Neural Networks (DNNs) are increasingly deployed across a wide range of applications, from image classification to autonomous driving. However, ensuring their reliability remains a challenge, and in many situations, alternative models with similar functionality and accuracy levels are available. Traditional accuracy-based evaluations often fail to capture behavioral differences between such models, particularly when testing datasets are limited, making it challenging to select or optimally combine models. Differential testing addresses this limitation by generating test inputs that expose discrepancies in the behavior of DNN models. However, existing differential testing approaches face significant limitations: many rely on access to model internals or are constrained by the availability of seed inputs, limiting their generalizability and effectiveness. In response to these challenges, we proposeDiffGAN, a black-box test generation approach for differential testing of DNN models. Our approach, though adaptable to other domains, is specific to DNN models for image classification tasks, a highly prevalent application area. Our method relies on a Generative Adversarial Network (GAN) and the Non-dominated Sorting Genetic Algorithm II (NSGA-II) to generate diverse and valid triggering inputs that effectively reveal behavioral discrepancies between models. Our method employs two custom fitness functions, one focused on diversity and the other on divergence, to guide the exploration of the GAN input space and identify discrepancies between the models’ outputs. By strategically searching the GAN input space, we show thatDiffGANcan effectively generate inputs with specific features that trigger differences in behavior for the models under test. Unlike traditional white-box methods,DiffGANdoes not require access to the internal structure of the models, which makes it applicable to a wider range of situations. We evaluateDiffGANon a benchmark comprising eight pairs of DNN models trained on two widely used image classification datasets. Our results demonstrate thatDiffGANsignificantly outperforms a state-of-the-art (SOTA) baseline, generating four times more triggering inputs, with higher diversity and validity, within the same testing budget. Furthermore, we show that the generated input can be used to improve the accuracy of a machine learning-based model selection mechanism, which dynamically selects the best-performing model based on input characteristics and can thus be used as a smart model output voting mechanism when using alternative models together.
Zohreh Aghababaeyan, Manel Abdellatif, Lionel C. Briand, S. Ramesh 0002
IEEE Trans. Software Eng.1
2024 DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
abstract
Deep neural networks (DNNs) are widely used in various application domains such as image processing, speech recognition, and natural language processing. However, testing DNN models may be challenging due to the complexity and size of their input domain. In particular, testing DNN models often requires generating or exploring large unlabeled datasets. In practice, DNN test oracles, which identify the correct outputs for inputs, often require expensive manual effort to label test data, possibly involving multiple experts to ensure labeling correctness. In this article, we propose DeepGD , a black-box multi-objective test selection approach for DNN models. It reduces the cost of labeling by prioritizing the selection of test inputs with high fault-revealing power from large unlabeled datasets. DeepGD not only selects test inputs with high uncertainty scores to trigger as many mispredicted inputs as possible but also maximizes the probability of revealing distinct faults in the DNN model by selecting diverse mispredicted inputs. The experimental results conducted on four widely used datasets and five DNN models show that in terms of fault-revealing ability, (1) white-box, coverage-based approaches fare poorly, (2) DeepGD outperforms existing black-box test selection approaches in terms of fault detection, and (3) DeepGD also leads to better guidance for DNN model retraining when using selected inputs to augment the training set.
Zohreh Aghababaeyan, Manel Abdellatif, Mahboubeh Dadkhah, Lionel C. Briand
ACM Trans. Softw. Eng. Methodol.1
2023 Black-Box Testing of Deep Neural Networks through Test Case Diversity
abstract
Deep Neural Networks (DNNs) have been extensively used in many areas including image processing, medical diagnostics and autonomous driving. However, DNNs can exhibit erroneous behaviours that may lead to critical errors, especially when used in safety-critical systems. Inspired by testing techniques for traditional software systems, researchers have proposed neuron coverage criteria, as an analogy to source code coverage, to guide the testing of DNNs. Despite very active research on DNN coverage, several recent studies have questioned the usefulness of such criteria in guiding DNN testing. Further, from a practical standpoint, these criteria are white-box as they require access to the internals or training data of DNNs, which is often not feasible or convenient. Measuring such coverage requires executing DNNs with candidate inputs to guide testing, which is not an option in many practical contexts. In this paper, we investigate diversity metrics as an alternative to white-box coverage criteria. For the previously mentioned reasons, we require such metrics to be black-box and not rely on the execution and outputs of DNNs under test. To this end, we first select and adapt three diversity metrics and study, in a controlled manner, their capacity to measure actual diversity in input sets. We then analyze their statistical association with fault detection using four datasets and five DNNs. We further compare diversity with state-of-the-art white-box coverage criteria. As a mechanism to enable such analysis, we also propose a novel way to estimate fault detection in DNNs. Our experiments show that relying on the diversity of image features embedded in test input sets is a more reliable indicator than coverage criteria to effectively guide DNN testing. Indeed, we found that one of our selected black-box diversity metrics far outperforms existing coverage criteria in terms of fault-revealing capability and computational time. Results also confirm the suspicions that state-of-the-art coverage criteria are not adequate to guide the construction of test input sets to detect as many faults as possible using natural inputs.
Zohreh Aghababaeyan, Manel Abdellatif, Lionel C. Briand, S. Ramesh 0002, Mojtaba Bagherzadeh
IEEE Trans. Software Eng.1