VLDB 2026 Research / reviewers in the wild / expert
Sajad Khatiri
dblp:298/1109 · also Sajad Mazraeh Khatiri
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-0354-9747ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 4 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ICST Tool Competition 2026 - UAV Testing Track
Erdem Uysal, Gregory Loubet-Bonino, Prakash Aryan, Aren A. Babikian, Dmytro Humeniuk, Sajad Khatiri, Sebastiano Panichella |
ICST | 7 |
| 2025 | NN-SDCTest at the ICST 2025 Tool Competition - Self-Driving Car Testing TrackabstractTesting self-driving cars (SDCs) requires extensive simulation-based testing, making efficient test case selection important. This paper presents two approaches for test case selection in SDC testing: a curvature-based selector that analyzes road geometry and a graph neural network (GNN) based selector that learns failure patterns. The curvature-based approach uses road geometry analysis, turn detection, and group-based selection strategies, while the GNN approach has a four-layer neural architecture with feature engineering for predicting test failures. Both approaches are implemented and evaluated as part of the ICST 2025 Tool Competition for SDC testing. Our experimental results show that the GNN selector achieves superior computational efficiency (initialization: 1.45s vs 15.27s, selection: 0.30s vs 4.06s) and a better time-to-fault ratio (209.78 vs 239.48), while the curvature-based selector demonstrates stronger fault detection capabilities with a higher fault-to-selection ratio (0.214 vs 0.177). Both approaches maintain comparable diversity scores (0.037 and 0.039 respectively), demonstrating their effectiveness in achieving comprehensive test coverage. The comparative analysis provides insights into the strengths of geometric analysis and machine learning approaches in SDC test selection. Prakash Aryan, Sajad Khatiri |
ICST | 2 |
| 2025 | ICST Tool Competition 2025 - UAV Testing TrackabstractSimulation-based testing plays a crucial role in ensuring the safety of autonomous Unmanned Aerial Vehicles (UAVs); however, this area remains underexplored. The UAV Testing Competition aims to engage the software testing community by highlighting UAVs as an emerging and vital domain. This initiative offers a straightforward software platform and representative case studies to ease participants' entry into UAV testing, enabling them to develop their initial test generation tools for UAVs. In this second iteration of the competition, three tools were submitted, assessed, and thoroughly compared against each other, as well as the baseline approach. Our benchmarking framework analyzed their test generation capabilities across three distinct case studies. The resulting test suites were evaluated and ranked based on their failure detection and diversity. This paper provides an overview of the competition, detailing its context, platform, participating tools, evaluation methodology, and key findings. Sajad Khatiri, Tahereh Zohdinasab, Prasun Saurabh, Dmytro Humeniuk, Sebastiano Panichella |
ICST | 1 |
| 2025 | CertiFail at the ICST 2025 Tool Competition - Self-Driving Car Testing TrackabstractIn the context of Cyber Physical Systems like Self Driving Cars the transition from field operational testing to simulation based testing offer the advantage of lower cost, higher efficiency and the possibility of recording and repeating exact failing conditions. Yet traditional simulation based testing require execution of a long list of test cases, which means high computational costs and longer computational time. To combat this drawback, one can use different techniques to select and run test cases that are more likely to fail. In this work, we propose CertiFail a test case selection model that is an ensemble of several Machine learning models to more accurately predict if a test case is likely to fail. With CertiFail we achieve an accuracy of 73.7 percent in selection of test cases that fail. In addition CertiFail surpasses baseline models in terms of accurately predicting test cases that fail. Fasih Munir Malik, Sajad Khatiri |
ICST | 2 |
| 2025 | Bridging Research and Practice in Simulation-based Testing of Industrial Robot Navigation SystemsabstractEnsuring robust robotic navigation in dynamic environments is a key challenge, as traditional testing methods often struggle to cover the full spectrum of operational requirements. This paper presents the industrial adoption of Surrealist, a simulation-based test generation framework originally for UAVs, now applied to the ANYmal quadrupedal robot for industrial inspection. Our method uses a search-based algorithm to automatically generate challenging obstacle avoidance scenarios, uncovering failures often missed by manual testing. In a pilot phase, generated test suites revealed critical weaknesses in one experimental algorithm (40.3% success rate) and served as an effective benchmark to prove the superior robustness of another (71.2% success rate). The framework was then integrated into the ANYbotics workflow for a six-month industrial evaluation, where it was used to test five proprietary algorithms. A formal survey confirmed its value, showing it enhances the development process, uncovers critical failures, provides objective benchmarks, and strengthens the overall verification pipeline. Sajad Khatiri, Francisco E. Vina, Maximilian Wulf, Paolo Tonella, Sebastiano Panichella |
ASE | 1 |
| 2025 | When uncertainty leads to unsafety: Empirical insights into the role of uncertainty in unmanned aerial vehicle safetyabstractAbstract Despite the recent developments in obstacle avoidance and other safety features, autonomous Unmanned Aerial Vehicles (UAVs) continue to face safety challenges. No previous work investigated the relationship between the behavioral uncertainty of a UAV, characterized in this work by inconsistent or erratic control signal patterns, and the unsafety of its flight. By quantifying uncertainty, it is possible to develop a predictor for unsafety, which acts as a flight supervisor. We conducted a large-scale empirical investigation of safety violations using PX4-Autopilot, an open-source UAV software platform. Our dataset of over 5,000 simulated flights, created to challenge obstacle avoidance, allowed us to explore the relation between uncertain UAV decisions and safety violations: up to 89% of unsafe UAV states exhibit significant decision uncertainty, and up to 74% of uncertain decisions lead to unsafe states. Based on these findings, we implemented Superialist (Supervising Autonomous Aerial Vehicles), a runtime uncertainty detector based on autoencoders, the state-of-the-art technology for anomaly detection. Superialist achieved high performance in detecting uncertain behaviors with up to 96% precision and 93% recall. Despite the observed performance degradation when using the same approach for predicting unsafety (up to 74% precision and 87% recall), Superialist enabled early prediction of unsafe states up to 50 seconds in advance. Sajad Khatiri, Fatemeh Mohammadi Amin, Sebastiano Panichella, Paolo Tonella |
Empir. Softw. Eng. | 1 |
| 2025 | A Roadmap for Simulation-Based Testing of Autonomous Cyber-Physical Systems: Challenges and Future DirectionabstractAs the era of autonomous cyber-physical systems (ACPSs), such as unmanned aerial vehicles and self-driving cars, unfolds, the demand for robust testing methodologies is key to realizing the adoption of such systems in real-world scenarios. However, traditional software testing paradigms face unprecedented challenges in ensuring the safety and reliability of these systems. In response, this article pioneers a strategic roadmap for simulation-based system-level testing of ACPSs, specifically focusing on autonomous systems. Our article discusses the relevant challenges and obstacles of ACPSs, focusing on test automation and quality assurance, hence advocating for tailored solutions to address the unique demands of autonomous systems. While providing concrete definitions of test cases within simulation environments, we also accentuate the need to create new benchmark assets and the development of automated tools tailored explicitly for autonomous systems in the software engineering community. This article not only highlights the relevant, pressing issues the software engineering community should focus on (in terms of practices, expected automation, and paradigms), but it also outlines ways to tackle them. By outlining the various domains and challenges of simulation-based testing/development for ACPSs, we provide directions for future research efforts. Christian Birchler, Sajad Khatiri, Pooja Rani 0001, Timo Kehrer, Sebastiano Panichella |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Simulation-based Test Case Generation for Unmanned Aerial Vehicles in the Neighborhood of Real FlightsabstractUnmanned aerial vehicles (UAVs), also known as drones, are acquiring increasing autonomy. With their commercial adoption, the problem of testing their functional and non-functional, and in particular their safety requirements has become a critical concern. Simulation-based testing represents a fundamental practice, but the testing scenarios considered in software-in-the-loop testing may not be representative of the actual scenarios experienced in the field.In this paper, we propose SURREALIST (teSting UAVs in the neighboRhood of REAl flIghtS), a novel search-based approach that analyses the logs from real UAV flights and automatically generates simulation-based test cases in the neighborhood of such real flights, thereby improving the realism and representativeness of the simulation-based tests. This is done in two steps: first, SURREALIST faithfully replicates the given UAV flight in the simulation environment, generating a simulation-based test that mirrors a pre-logged real-world behavior. Then, it smoothly manipulates the replicated flight conditions to discover slightly modified test cases that are challenging or trigger misbehaviors of the UAV under test in simulation. In our experiments, we were able to replicate a real flight accurately in the simulation environment and to expose unstable and potentially unsafe behavior in the neighborhood of a replicated flight, which even led to crashes. Sajad Khatiri, Sebastiano Panichella, Paolo Tonella |
ICST | 1 |
| 2023 | Machine learning-based test selection for simulation-based testing of self-driving cars softwareabstractAbstract Simulation platforms facilitate the development of emerging Cyber-Physical Systems (CPS) like self-driving cars (SDC) because they are more efficient and less dangerous than field operational test cases. Despite this, thoroughly testing SDCs in simulated environments remains challenging because SDCs must be tested in a sheer amount of long-running test cases. Past results on software testing optimization have shown that not all the test cases contribute equally to establishing confidence in test subjects’ quality and reliability, and the execution of “safe and uninformative” test cases can be skipped to reduce testing effort. However, this problem is only partially addressed in the context of SDC simulation platforms. In this paper, we investigate test selection strategies to increase the cost-effectiveness of simulation-based testing in the context of SDCs. We propose an approach called SDC-Scissor (SDC coS t-effeC tI ve teS t S electOR) that leverages Machine Learning (ML) strategies to identify and skip test cases that are unlikely to detect faults in SDCs before executing them. Our evaluation shows that SDC-Scissor outperforms the baselines. With the Logistic model, we achieve an accuracy of 70%, a precision of 65%, and a recall of 80% in selecting tests leading to a fault and improved testing cost-effectiveness. Specifically, SDC-Scissor avoided the execution of 50% of unnecessary tests as well as outperformed two baseline strategies. Complementary to existing work, we also integrated SDC-Scissor into the context of an industrial organization in the automotive domain to demonstrate how it can be used in industrial settings. Christian Birchler, Sajad Khatiri, Bill Bosshard, Alessio Gambi, Sebastiano Panichella |
Empir. Softw. Eng. | 2 |
| 2023 | Cost-effective simulation-based test selection in self-driving cars softwareabstractSimulation environments are essential for the continuous development of complex cyber-physical systems such as self-driving cars (SDCs). Previous results on simulation-based testing for SDCs have shown that many automatically generated tests do not strongly contribute to the identification of SDC faults, hence do not contribute towards increasing the quality of SDCs. Because running such “uninformative” tests generally leads to a waste of computational resources and a drastic increase in the testing cost of SDCs, testers should avoid them. However, identifying “uninformative” tests before running them remains an open challenge. Hence, this paper proposes SDC-Scissor, a framework that leverages Machine Learning (ML) to identify SDC tests that are unlikely to detect faults in the SDC software under test, thus enabling testers to skip their execution and drastically increase the cost-effectiveness of simulation-based testing of SDCs software. Our evaluation concerning the usage of six ML models on two large datasets characterized by 22'652 tests showed that SDC-Scissor achieved a classification F1-score up to 96%. Moreover, our results show that SDC-Scissor outperformed a randomized baseline in identifying more failing tests per time unit. Webpage & Video: https://github.com/ChristianBirchler/sdc-scissor Christian Birchler, Nicolas Erni, Sajad Khatiri, Alessio Gambi, Sebastiano Panichella |
Sci. Comput. Program. | 3 |
| 2023 | Single and Multi-objective Test Cases Prioritization for Self-driving Cars in Virtual EnvironmentsabstractTesting with simulation environments helps to identify critical failing scenarios for self-driving cars (SDCs). Simulation-based tests are safer than in-field operational tests and allow detecting software defects before deployment. However, these tests are very expensive and are too many to be run frequently within limited time constraints. In this article, we investigate test case prioritization techniques to increase the ability to detect SDC regression faults with virtual tests earlier. Our approach, called SDC-Prioritizer , prioritizes virtual tests for SDCs according to static features of the roads we designed to be used within the driving scenarios. These features can be collected without running the tests, which means that they do not require past execution results. We introduce two evolutionary approaches to prioritize the test cases using diversity metrics (black-box heuristics) computed on these static features. These two approaches, called SO-SDC-Prioritizer and MO-SDC-Prioritizer , use single-objective and multi-objective genetic algorithms ( GA ), respectively, to find trade-offs between executing the less expensive tests and the most diverse test cases earlier. Our empirical study conducted in the SDC domain shows that MO-SDC-Prioritizer significantly ( P - value <=0.1 e -10) improves the ability to detect safety-critical failures at the same level of execution time compared to baselines: random and greedy-based test case orderings. Besides, our study indicates that multi-objective meta-heuristics outperform single-objective approaches when prioritizing simulation-based tests for SDCs. MO-SDC-Prioritizer prioritizes test cases with a large improvement in fault detection while its overhead (up to 0.45% of the test execution cost) is negligible. Christian Birchler, Sajad Khatiri, Pouria Derakhshanfar, Sebastiano Panichella, Annibale Panichella |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2022 | Cost-effective Simulation-based Test Selection in Self-driving Cars Software with SDC-ScissorabstractSimulation platforms facilitate the continuous development of complex systems such as self-driving cars (SDCs). However, previous results on testing SDCs using simulations have shown that most of the automatically generated tests do not strongly contribute to establishing confidence in the quality and reliability of the SDC. Therefore, those tests can be characterized as “uninformative”, and running them generally means wasting precious computational resources. We address this issue with SDC-Scissor, a framework that leverages Machine Learning to identify simulation-based tests that are unlikely to detect faults in the SDC software under test and skip them before their execution. Consequently, by filtering out those tests, SDC-Scissor reduces the number of long-running simulations to execute and drastically increases the cost-effectiveness of simulation-based testing of SDCs software. Our evaluation concerning two large datasets and around 12'000 tests showed that SDC-Scissor achieved a higher classification F1-score (between 47% and 90%) than a randomized baseline in identifying tests that lead to a fault and reduced the time spent running uninformative tests (speedup between 107% and 170%). Webpage & Video: https://github.com/ChristianBirchler/sdc-scissor Christian Birchler, Nicolas Erni, Sajad Khatiri, Alessio Gambi, Sebastiano Panichella |
SANER | 3 |