Mahboubeh Dadkhah

dblp:183/4650 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-0436-8369ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 MetaSel: A Test Selection Approach for Fine-Tuned DNN Models
abstract
Deep Neural Networks (DNNs) face challenges during deployment due to covariate shift, i.e., data distribution shifts between development and deployment contexts. Fine-tuning adapts pre-trained models to new contexts requiring smaller labeled sets. However, testing fine-tuned models under constrained labeling budgets remains a critical challenge. This paper introduces MetaSel, a new approach tailored for DNN models that have been fine-tuned to address covariate shift, to select tests from unlabeled inputs. MetaSel assumes that fine-tuned and pre-trained models share related data distributions and exhibit similar behaviors for many inputs. However, their behaviors diverge within the input subspace where fine-tuning alters decision boundaries, making those inputs more prone to misclassification. Unlike general approaches that rely solely on the DNN model and its input set, MetaSel leverages information from both the fine-tuned and pre-trained models and their behavioral differences to estimate misclassification probability for unlabeled test inputs, enabling more effective test selection. Our extensive empirical evaluation, comparing MetaSel against 11 state-of-the-art approaches and involving 68 fine-tuned models across weak, medium, and strong distribution shifts, demonstrates that MetaSel consistently delivers significant improvements in Test Relative Coverage (TRC) over existing baselines, particularly under highly constrained labeling budgets. MetaSel shows average TRC improvements of 28.46% to 56.18% over the most frequent second-best baselines while maintaining a high TRC median and low variability. Our results confirm MetaSel’s practicality, robustness, and cost-effectiveness for test selection in the context of fine-tuned models.
Amin Abbasishahkoo, Mahboubeh Dadkhah, Lionel C. Briand, Dayi Lin
IEEE Trans. Software Eng.2
2024 DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
abstract
Deep neural networks (DNNs) are widely used in various application domains such as image processing, speech recognition, and natural language processing. However, testing DNN models may be challenging due to the complexity and size of their input domain. In particular, testing DNN models often requires generating or exploring large unlabeled datasets. In practice, DNN test oracles, which identify the correct outputs for inputs, often require expensive manual effort to label test data, possibly involving multiple experts to ensure labeling correctness. In this article, we propose DeepGD , a black-box multi-objective test selection approach for DNN models. It reduces the cost of labeling by prioritizing the selection of test inputs with high fault-revealing power from large unlabeled datasets. DeepGD not only selects test inputs with high uncertainty scores to trigger as many mispredicted inputs as possible but also maximizes the probability of revealing distinct faults in the DNN model by selecting diverse mispredicted inputs. The experimental results conducted on four widely used datasets and five DNN models show that in terms of fault-revealing ability, (1) white-box, coverage-based approaches fare poorly, (2) DeepGD outperforms existing black-box test selection approaches in terms of fault detection, and (3) DeepGD also leads to better guidance for DNN model retraining when using selected inputs to augment the training set.
Zohreh Aghababaeyan, Manel Abdellatif, Mahboubeh Dadkhah, Lionel C. Briand
ACM Trans. Softw. Eng. Methodol.3
2024 TEASMA: A Practical Methodology for Test Adequacy Assessment of Deep Neural Networks
abstract
Successful deployment of Deep Neural Networks (DNNs), particularly in safety-critical systems, requires their validation with an adequate test set to ensure a sufficient degree of confidence in test outcomes. Although well-established test adequacy assessment techniques from traditional software, such as mutation analysis and coverage criteria, have been adapted to DNNs in recent years, we still need to investigate their application within a comprehensive methodology for accurately predicting the fault detection ability of test sets and thus assessing their adequacy. In this paper, we propose and evaluateTEASMA, a comprehensive and practical methodology designed to accurately assess the adequacy of test sets for DNNs. In practice,TEASMAallows engineers to decide whether they can trust high-accuracy test results and thus validate the DNN before its deployment. Based on a DNN model's training set,TEASMAprovides a procedure to build accurate DNN-specific prediction models of the Fault Detection Rate (FDR) of a test set using an existing adequacy metric, thus enabling its assessment. We evaluatedTEASMAwith four state-of-the-art test adequacy metrics: Distance-based Surprise Coverage (DSC), Likelihood-based Surprise Coverage (LSC), Input Distribution Coverage (IDC), and Mutation Score (MS). We calculated MS based on mutation operators that directly modify the trained DNN model (i.e., post-training operators) due to their significant computational advantage compared to the operators that modify the DNN's training set or program (i.e., pre-training operators). Our extensive empirical evaluation, conducted across multiple DNN models and input sets, including large input sets such as ImageNet, reveals a strong linear correlation between the predicted and actual FDR values derived from MS, DSC, and IDC, with minimum$R^{2}$values of 0.94 for MS and 0.90 for DSC and IDC. Furthermore, a low average Root Mean Square Error (RMSE) of 9% between actual and predicted FDR values across all subjects, when relying on regression analysis and MS, demonstrates the latter's superior accuracy when compared to DSC and IDC, with RMSE values of 0.17 and 0.18, respectively. Overall, these results suggest thatTEASMAprovides a reliable basis for confidently deciding whether to trust test results for DNN models.
Amin Abbasishahkoo, Mahboubeh Dadkhah, Lionel C. Briand, Dayi Lin
IEEE Trans. Software Eng.2
2020 A systematic literature review on semantic web enabled software testing
Mahboubeh Dadkhah, Saeed Araban, Samad Paydar
J. Syst. Softw.1
2016 Semantic-Based Test Case Generation
abstract
Software testing is a major V&V activity that revolves around quality test cases. Generating quality test cases is inherently knowledge intensive, tedious and expensive task that is traditionally done by humans. Therefore, a great deal of research has been done to facilitate and automate test case generation as much as possible. Given the knowledge intensity of software test case generation, various knowledge management techniques, such as semantic-based techniques, are applicable. The main focus of our research is automatic generation of quality test cases using available knowledge from the early stages of software development, i.e. the RE. Our goal is to develop a semantic web enabled framework for integrating knowledge from various requirement models to generate effective and efficient test cases automatically. We are going to apply semantic technology to facilitate the test case generation process by means of ontologies and Web of Data.
Mahboubeh Dadkhah
ICST1