Hassan Sartaj

dblp:245/8614 · DBLP profile ↗
← Back
17ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0001-5212-9787ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 12 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 REST API Testing in DevOps: A Study on an Evolving Healthcare IoT Application
abstract
Healthcare Internet of Things (IoT) applications often integrate various third-party healthcare applications and medical devices through REST APIs, resulting in complex and interdependent networks of REST APIs. Oslo City’s healthcare department collaborates with various industry partners to develop these applications, enriched with diverse REST APIs that evolve during the DevOps process to accommodate evolving needs such as new features, services, and devices. Oslo City’s primary goal is to utilize automated solutions for continuous testing of REST APIs at each evolution stage to ensure dependability. Although the literature offers various automated REST API testing tools, their effectiveness in regression testing of the evolving REST APIs of healthcare IoT applications within a DevOps context remains undetermined. This article evaluates state-of-the-art and well-established REST API testing tools—specifically, RESTest, EvoMaster, Schemathesis, RESTler, and RestTestGen—for the regression testing of a real-world healthcare IoT application, considering failures, faults, coverage, regressions, and cost. We conducted experiments using all accessible REST APIs (17 APIs with 120 endpoints), and 14 releases evolved during DevOps. Overall, all tools generated tests leading to several failures, 18 potential faults, up to 84% coverage, and 23 regressions. Over 70% of tests generated by all tools fail to detect failures, resulting in significant overhead.
Hassan Sartaj, Shaukat Ali 0001, Julie Marie Gjøby
ACM Trans. Softw. Eng. Methodol.1
2025 LLMs in the Heart of Differential Testing: A Case Study on a Medical Rule Engine
abstract
The Cancer Registry of Norway (CRN) uses an automated cancer registration support system (CaReSS) to support core cancer registry activities, i.e., data capture, data curation, and producing data products and statistics for various stakeholders. GURI is a core component of CaReSS, which is responsible for validating incoming data with medical rules. Such medical rules are manually implemented by medical experts based on medical standards, regulations, and research. Since large language models (LLMs) have been trained on a large amount of public information, including these documents, they can be employed to generate tests for GURI. Thus, we propose an LLM-based test generation and differential testing approach (LLMeDiff) to test GURI. We experimented with four different LLMs, two medical rule engine implementations, and 58 real medical rules to investigate the hallucination, success, time efficiency, and robustness of the LLMs to generate tests, and these tests' ability to find potential issues in GURI. Our results showed that GPT-3.5 hallucinates the least, is the most successful, and is generally the most robust; however, it has the worst time efficiency. Our differential testing revealed 22 medical rules where implementation inconsistencies were discovered (e.g., regarding handling rule versions). Finally, we provide insights for practitioners and researchers based on the results.
Erblin Isaku, Christoph Laaber, Hassan Sartaj, Shaukat Ali 0001, Thomas Schwitalla, Jan Nygård
ICST3
2025 Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
abstract
Self-adaptive robots (SARs) in complex, uncertain environments must proactively detect and address abnormal behaviors, including out-of-distribution (OOD) cases. To this end, digital twins offer a valuable solution for OOD detection. Thus, we present a digital twin-based approach for OOD detection (ODiSAR) in SARs. ODiSAR uses a Transformer-based digital twin to forecast SAR states and employs reconstruction error and Monte Carlo dropout for uncertainty quantification. By combining reconstruction error with predictive variance, the digital twin effectively detects OOD behaviors, even in previously unseen conditions. The digital twin also includes an explainability layer that links potential OOD to specific SAR states, offering insights for self-adaptation. We evaluated ODiSAR by creating digital twins of two industrial robots: one navigating an office environment, and another performing maritime ship navigation. In both cases, ODiSAR forecasts SAR behaviors (i.e., robot trajectories and vessel motion) and proactively detects OOD events. Our results showed that ODiSAR achieved high detection performance—up to 98% AUROC, 96% TNR@TPR95, and 95% F1-score—while providing interpretable insights to support self-adaptation.
Erblin Isaku, Hassan Sartaj, Shaukat Ali 0001, Beatriz Sanguino, Guoyuan Li, Houxiang Zhang, Thomas Peyrucain
ASE2
2025 Uncertainty-aware environment simulation of medical devices digital twins
Hassan Sartaj, Shaukat Ali 0001, Julie Marie Gjøby
Softw. Syst. Model.1
2025 Search-Based MC/DC Test Data Generation With OCL Constraints
abstract
ABSTRACT System‐level testing of avionics software systems requires compliance with different international safety standards such as DO‐178C. An important consideration of the avionics industry is automated test data generation according to the criteria suggested by safety standards. One of the recommended criteria by DO‐178C is the modified condition/decision coverage (MC/DC) criterion. Current model‐based test data generation approaches use constraints written in Object Constraint Language (OCL) and apply search techniques to generate test data. These approaches either do not support MC/DC criterion or suffer from performance issues while generating test data for large‐scale avionics systems. In this paper, we propose an effective way to automate MC/DC test data generation during model‐based testing. We develop a strategy that utilizes case‐based reasoning (CBR) and range reduction heuristics designed to solve MC/DC‐tailored OCL constraints. We performed an empirical study to compare our proposed strategy for MC/DC test data generation using CBR, range reduction, both CBR and range reduction, with an original search algorithm, and random search. We also empirically compared our strategy with existing constraint‐solving approaches. The results show that both CBR and range reduction for MC/DC test data generation outperform the baseline approach. Moreover, the combination of both CBR and range reduction for MC/DC test data generation is an effective approach compared to existing constraint solvers.
Hassan Sartaj, Muhammad Zohaib Z. Iqbal, Atif A. A. Jilani, Muhammad Uzair Khan
Softw. Test. Verification Reliab.1
2025 MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
abstract
Testing healthcare Internet of Things (IoT) applications at system and integration levels necessitates integrating numerous medical devices. Challenges of incorporating medical devices are: (i) their continuous evolution, making it infeasible to include all device variants and (ii) rigorous testing at scale requires multiple devices and their variants, which is time-intensive, costly, and impractical. Our collaborator, Oslo City’s health department, faced these challenges in developing automated test infrastructure, which our research aims to address. In this context, we propose a meta-learning-based approach ( MeDeT ) to generate digital twins (DTs) of medical devices and adapt DTs to evolving devices. We evaluate MeDeT in Oslo City’s context using five widely used medical devices integrated with a real-world healthcare IoT application. Our evaluation assesses MeDeT ’s ability to generate and adapt DTs across various devices and versions using different few-shot methods, the fidelity of these DTs, the scalability of operating 1,000 DTs concurrently, and the associated time costs. Results show that MeDeT can generate DTs with over 96% fidelity, adapt DTs to different devices and newer versions with reduced time cost (around one minute), and operate 1,000 DTs in a scalable manner while maintaining the fidelity level, thus serving in place of physical devices for testing.
Hassan Sartaj, Shaukat Ali 0001, Julie Marie Gjøby
ACM Trans. Softw. Eng. Methodol.1
2024 Digital Twins Environment Simulation for Testing Healthcare IoT Applications
abstract
Healthcare applications using the Internet of Things (IoT) architecture are connected with various medical devices designed to serve patients. Rigorous system testing of health care IoT applications requires integrating multiple medical devices to ensure the dependability of these applications. The integration of numerous physical medical devices with varying versions is a costly and time-consuming process. In this regard, our previous work introduced the concept of employing digital twins (DTs) as substitutes for physical devices for testing purposes. Specifically, we presented a model-based approach to generate DTs of medicine dispensers. The evaluation of our approach with a Karie medicine dispenser demonstrated 92% fidelity of Karie DTs. From our experiences, we observed that the real operating environment of medical devices involves several non-deterministic factors, essential for DTs to reflect devices' behavior precisely. Therefore, we plan to devise a methodology to model and simulate the environment of medical devices DTs, taking into account environmental uncertainties. We intend to empirically evaluate our methodology in the real-world context to analyze the simulation of behavioral models of the environment and uncertain events generated for DTs.
Hassan Sartaj, Shaukat Ali 0001, Julie Marie Gjøby
COMPSAC1
2024 Automated system-level testing of unmanned aerial systems
Hassan Sartaj, Asmar Muqeet, Muhammad Zohaib Z. Iqbal, Muhammad Uzair Khan
Autom. Softw. Eng.1
2024 Model-based digital twins of medicine dispensers for healthcare IoT applications
abstract
Summary Healthcare applications with the Internet of Things (IoT) are often safety‐critical, thus, require extensive testing. Such applications are often connected to smart medical devices from various vendors. System‐level testing of such applications requires test infrastructures physically integrating medical devices, which is time and monetary‐wise expensive. Moreover, applications continuously evolve, for example, introducing new devices and users and updating software. Nevertheless, a test infrastructure enabling testing with a few devices is insufficient for testing healthcare IoT systems, hence compromising their dependability. In this paper, we propose a model‐based approach for the creation and operation of digital twins (DTs) of medicine dispensers as a replacement for physical devices to support the automated testing of IoT applications at scale. We evaluate our approach with an industrial IoT system with medicine dispensers in the context of Oslo City and its industrial partners, providing healthcare services to its residents. We study the fidelity of DTs in terms of their functional similarities with their physical counterparts: medicine dispensers. Results show that the DTs behave more than 92% similar to the physical medicine dispensers, providing a faithful replacement for the dispenser.
Hassan Sartaj, Shaukat Ali 0001, Tao Yue 0002, Kjetil Moberg
Softw. Pract. Exp.1
2023 Cost Reduction on Testing Evolving Cancer Registry System
abstract
The Cancer Registration Support System (CaReSS), built by the Cancer Registry of Norway (CRN), is a complex real-world socio-technical software system that undergoes continuous evolution in its implementation. Consequently, continuous testing of CaReSS with automated testing tools is needed such that its dependability is always ensured. Towards automated testing of a key software subsystem of CaReSS, i.e., GURI, we present a real-world application of an extension to the open-source tool EvoMaster, which automatically generates test cases with evolutionary algorithms. We named the extension EvoClass, which enhances EvoMaster with a machine learning classifier to reduce the overall testing cost. This is imperative since testing with EvoMaster involves sending many requests to GURI deployed in different environments, including the production environment, whose performance and functionality could potentially be affected by many requests. The machine learning classifier of EvoClass can predict whether a request generated by EvoMaster will be executed successfully or not; if not, the classifier filters out such requests, consequently reducing the number of requests to be executed on GURI. We evaluated EvoClass on ten GURI versions over four years in three environments: development, testing, and production. Results showed that EvoClass can significantly reduce the testing cost of evolving GURI without reducing testing effectiveness (measured as rule coverage) across all three environments, as compared to the default EvoMaster. Overall, EvoClass achieved ≈31% of overall cost reduction. Finally, we report our experiences and lessons learned that are equally valuable for researchers and practitioners.
Erblin Isaku, Hassan Sartaj, Christoph Laaber, Tao Yue 0002, Shaukat Ali 0001, Thomas Schwitalla, Jan Nygård
ICSME2
2023 Testing Real-World Healthcare IoT Application: Experiences and Lessons Learned
abstract
Healthcare Internet of Things (IoT) applications require rigorous testing to ensure their dependability. Such applications are typically integrated with various third-party healthcare applications and medical devices through REST APIs. This integrated network of healthcare IoT applications leads to REST APIs with complicated and interdependent structures, thus creating a major challenge for automated system-level testing. We report an industrial evaluation of a state-of-the-art REST APIs testing approach (RESTest) on a real-world healthcare IoT application. We analyze the effectiveness of RESTest’s testing strategies regarding REST APIs failures, faults in the application, and REST API coverage, by experimenting with six REST APIs of 41 API endpoints of the healthcare IoT application. Results show that several failures are discovered in different REST APIs with ≈56% coverage using RESTest. Moreover, nine potential faults are identified. Using the evidence collected from the experiments, we provide our experiences and lessons learned.
Hassan Sartaj, Shaukat Ali 0001, Tao Yue 0002, Kjetil Moberg
ESEC/SIGSOFT FSE1
2021 Automated Approach for System-level Testing of Unmanned Aerial Systems
abstract
Unmanned aerial systems (UAS) have a large number of applications in civil and military domains. UAS rely on various avionics systems that are safety-critical and mission-critical. A major requirement of international safety standards is to perform rigorous system-level testing of avionics systems, including software systems. The current industrial practice is to manually create test scenarios, manually or automatically execute these scenarios using simulators, and manual evaluation of the outcomes. A fundamental part of system-level testing of such systems is the simulation of environmental context. The test scenarios typically consist of setting certain environment conditions and testing the system under test in these settings. The state-of-the-art approaches available for this purpose also require manual test scenario development and manual test evaluation. In this research work, we propose an approach to automate the system-level testing of the UAS. The proposed approach (AITester) utilizes model-based testing and artificial intelligence (AI) techniques to automatically generate, execute, and evaluate various test scenarios. The test scenarios are generated on the fly, i.e., during test execution based on the environmental context at runtime. We develop a toolset to support automation. We perform a pilot experiment using a widely-used open-source autopilot, ArduPilot. The preliminary results show that the AITester is effective and efficient in violating autopilot expected behavior.
Hassan Sartaj
ASE1
2021 Testing cockpit display systems of aircraft using a model-based approach
Hassan Sartaj, Muhammad Zohaib Z. Iqbal, Muhammad Uzair Khan
Softw. Syst. Model.1
2020 CDST: A Toolkit for Testing Cockpit Display Systems
abstract
Avionics are highly critical systems that require extensive testing governed by international safety standards. Cockpit Display Systems (CDS) are an essential component of modern aircraft cockpits and display information from the user application using various widgets. A significant step in the testing of avionics is to evaluate whether these CDS are displaying the correct information. A common industrial practice is to manually test the information on these CDS by taking the aircraft into different scenarios during the simulation. Given the large number of scenarios to test, manual testing of such behavior is a laborious activity. In the previous work, we proposed a model-based testing approach for the CDS. In this paper, we present a CDST (CDS Testing) toolkit that automates the process of testing CDS. We discuss the workflow and architecture of the tool and also demonstrates the tool on an industrial case study. The results show that the tool is able to generate, execute, and evaluate the test cases and identify 3 bugs in the case study.
Hassan Sartaj, Muhammad Zohaib Z. Iqbal, Muhammad Uzair Khan
ICST1
2019 A Model-Based Testing Approach for Cockpit Display Systems of Avionics
abstract
Avionics are highly critical systems that require extensive testing governed by international safety standards. Cockpit Display Systems (CDS) are an essential component of modern aircraft cockpits and display information from the user application (UA) using various widgets. A significant step in the testing of avionics is to evaluate whether these CDS are displaying the correct information. A common industrial practice is to manually test the information on these CDS by taking the aircraft into different scenarios during the simulation. Such testing is required very frequently and at various changes in the avionics. Given the large number of scenarios to test, manual testing of such behavior is a laborious activity. In this paper, we propose a model-based strategy for automated testing of the information displayed on CDS. Our testing approach focuses on evaluating that the information from the user applications is being displayed correctly on the CDS. For this purpose, we develop a profile for capturing the details of different widgets of the display screens using models. The profile is based on the ARINC 661 standard for Cockpit Display Systems. The expected behavior of the CDS visible on the screens of the aircraft is captured using constraints written in Object Constraint Language. We apply our approach on an industrial case study of a Primary Flight Display (PFD) developed for an aircraft. Our results showed that the proposed approach is able to automatically identify faults in the simulation of PFD. Based on the results, it is concluded that the proposed approach is useful in finding display faults on avionics CDS.
Muhammad Zohaib Z. Iqbal, Hassan Sartaj, Muhammad Uzair Khan, Fitash Ul Haq, Ifrah Qaisar
MoDELS2
2019 A Search-Based Approach to Generate MC/DC Test Data for OCL Constraints
Hassan Sartaj, Muhammad Zohaib Z. Iqbal, Atif A. A. Jilani, Muhammad Uzair Khan
SSBSE1
2019 AspectOCL: using aspects to ease maintenance of evolving constraint specification
Muhammad Uzair Khan, Hassan Sartaj, Muhammad Zohaib Z. Iqbal, Muhammad Usman 0013, Numra Arshad
Empir. Softw. Eng.2