VLDB 2026 Research / reviewers in the wild / expert
Tahereh Zohdinasab
dblp:297/2168
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-0191-1151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ICST Tool Competition 2025 - UAV Testing TrackabstractSimulation-based testing plays a crucial role in ensuring the safety of autonomous Unmanned Aerial Vehicles (UAVs); however, this area remains underexplored. The UAV Testing Competition aims to engage the software testing community by highlighting UAVs as an emerging and vital domain. This initiative offers a straightforward software platform and representative case studies to ease participants' entry into UAV testing, enabling them to develop their initial test generation tools for UAVs. In this second iteration of the competition, three tools were submitted, assessed, and thoroughly compared against each other, as well as the baseline approach. Our benchmarking framework analyzed their test generation capabilities across three distinct case studies. The resulting test suites were evaluated and ranked based on their failure detection and diversity. This paper provides an overview of the competition, detailing its context, platform, participating tools, evaluation methodology, and key findings. Sajad Khatiri, Tahereh Zohdinasab, Prasun Saurabh, Dmytro Humeniuk, Sebastiano Panichella |
ICST | 2 |
| 2024 | Focused Test Generation for Autonomous Driving SystemsabstractTesting Autonomous Driving Systems (ADSs) is crucial to ensure their reliability when navigating complex environments. ADSs may exhibit unexpected behaviours when presented, during operation, with driving scenarios containing features inadequately represented in the training dataset. To address this shift from development to operation, developers must acquire new data with the newly observed features. This data can be then utilised to fine tune the ADS, so as to reach the desired level of reliability in performing driving tasks. However, the resource-intensive nature of testing ADSs requires efficient methodologies for generating targeted and diverse tests. In this work, we introduce a novel approach, DeepAtash-LR , that incorporates a surrogate model into the focused test generation process. This integration significantly improves focused testing effectiveness and applicability in resource-intensive scenarios. Experimental results show that the integration of the surrogate model is fundamental to the success of DeepAtash-LR . Our approach was able to generate an average of up to 60× more targeted, failure-inducing inputs compared to the baseline approach. Moreover, the inputs generated by DeepAtash-LR were useful to significantly improve the quality of the original ADS through fine tuning. Tahereh Zohdinasab, Vincenzo Riccio, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | An Empirical Study on Low- and High-Level Explanations of Deep Learning MisbehavioursabstractBackground: Most quality assessment approaches for Deep Learning (DL) focus on finding misbehaviour-inducing inputs. However, it is difficult to clearly understand the causes of misbehaviours, due to the DL software opaqueness. Recent research proposed different techniques to explain DL misbehaviours, producing input explanations either at a “low level” (raw input elements) or at a “high level” (input features). Aims: We aim to compare the similarity between different explanations and assess to what extent they are understandable. Method: We have conducted an empirical study involving 3 state-of-the-art techniques for DL explanation in 13 configurations, applied to 2 different DL tasks. We have also collected answers from 48 questionnaires submitted to SE experts. Results: Low- and high-level techniques provide dissimilar explanations for the same inputs. However, experts deemed none of the explanations as useful in 28% of the cases. Conclusion: Despite the complementarity of existing explanations, further research is needed to produce better explanations. Tahereh Zohdinasab, Vincenzo Riccio, Paolo Tonella |
ESEM | 1 |
| 2023 | DeepAtash: Focused Test Generation for Deep Learning SystemsabstractWhen deployed in the operation environment, Deep Learning (DL) systems often experience the so-called development to operation (dev2op) data shift, which causes a lower prediction accuracy on field data as compared to the one measured on the test set during development. To address the dev2op shift, developers must obtain new data with the newly observed features, as these are under-represented in the train/test set, and must use them to fine tune the DL model, so as to reach the desired accuracy level. In this paper, we address the issue of acquiring new data with the specific features observed in operation, which caused a dev2op shift, by proposing DeepAtash, a novel search-based focused testing approach for DL systems. DeepAtash targets a cell in the feature space, defined as a combination of feature ranges, to generate misbehaviour-inducing inputs with predefined features. Experimental results show that DeepAtash was able to generate up to 29X more targeted, failure-inducing inputs than the baseline approach. The inputs generated by DeepAtash were useful to significantly improve the quality of the original DL systems through fine tuning not only on data with the targeted features, but quite surprisingly also on inputs drawn from the original distribution. Tahereh Zohdinasab, Vincenzo Riccio, Paolo Tonella |
ISSTA | 1 |
| 2023 | Efficient and Effective Feature Space Exploration for Testing Deep Learning SystemsabstractAssessing the quality of Deep Learning (DL) systems is crucial, as they are increasingly adopted in safety-critical domains. Researchers have proposed several input generation techniques for DL systems. While such techniques can expose failures, they do not explain which features of the test inputs influenced the system’s (mis-) behaviour. DeepHyperion was the first test generator to overcome this limitation by exploring the DL systems’ feature space at large. In this article, we propose DeepHyperion-CS , a test generator for DL systems that enhances DeepHyperion by promoting the inputs that contributed more to feature space exploration during the previous search iterations. We performed an empirical study involving two different test subjects (i.e., a digit classifier and a lane-keeping system for self-driving cars). Our results proved that the contribution-based guidance implemented within DeepHyperion-CS outperforms state-of-the-art tools and significantly improves the efficiency and the effectiveness of DeepHyperion . DeepHyperion-CS exposed significantly more misbehaviours for five out of six feature combinations and was up to 65% more efficient than DeepHyperion in finding misbehaviour-inducing inputs and exploring the feature space. DeepHyperion-CS was useful for expanding the datasets used to train the DL systems, populating up to 200% more feature map cells than the original training set. Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2021 | DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchabstractDeep Learning (DL) has been successfully applied to a wide range of application domains, including safety-critical ones. Several DL testing approaches have been recently proposed in the literature but none of them aims to assess how different interpretable features of the generated inputs affect the system's behaviour. Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo Tonella |
ISSTA | 1 |