Jan Nygård

dblp:154/4319 · also Jan F. Nygård, Jan Franz Nygård · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-9655-7003ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 5 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Quantum Neural Network Classifier for Cancer Registry System Testing: A Feasibility Study
abstract
With the rapid advancement of quantum computing, research on quantum machine learning (QML) algorithms has grown significantly. Among these, the Quantum Neural Network (QNN) stands out as one of the promising algorithms that integrates the principles of quantum computing with artificial neural networks to process data. Inspired by applications of QNN across fields, we investigate their use in software testing for the Cancer Registry of Norway (CRN), part of the Norwegian Institute of Public Health (NIPH), responsible for cancer statistics among the Norwegian population. CRN develops a complex socio-technical software system, Cancer Registration Support System ( \(\mathtt{CaReSS}\) ), interacting with many entities (e.g., hospitals, medical laboratories, and other patient registries) to achieve its task. For cost-effective testing of \(\mathtt{CaReSS}\) , CRN has employed \(\mathtt{EvoMaster}\) , an AI-based REST API testing tool combined with an integrated classical machine learning model \(\mathtt{EvoClass}\) . Within this context, we propose \(\mathtt{EvoQlass}\) to investigate the feasibility of using, inside \(\mathtt{EvoMaster}\) , a QNN classifier, instead of the existing classical machine learning model. Results indicate that \(\mathtt{EvoQlass}\) can achieve performance comparable to that of \(\mathtt{EvoClass}\) . We further explore the effects of various QNN configurations on performance and offer recommendations for optimal QNN settings for future QNN developers.
Xinyi Wang 0004, Shaukat Ali 0001, Paolo Arcaini, Narasimha Raghavan, Jan Nygård
ACM Trans. Softw. Eng. Methodol.5
2025 LLMs in the Heart of Differential Testing: A Case Study on a Medical Rule Engine
abstract
The Cancer Registry of Norway (CRN) uses an automated cancer registration support system (CaReSS) to support core cancer registry activities, i.e., data capture, data curation, and producing data products and statistics for various stakeholders. GURI is a core component of CaReSS, which is responsible for validating incoming data with medical rules. Such medical rules are manually implemented by medical experts based on medical standards, regulations, and research. Since large language models (LLMs) have been trained on a large amount of public information, including these documents, they can be employed to generate tests for GURI. Thus, we propose an LLM-based test generation and differential testing approach (LLMeDiff) to test GURI. We experimented with four different LLMs, two medical rule engine implementations, and 58 real medical rules to investigate the hallucination, success, time efficiency, and robustness of the LLMs to generate tests, and these tests' ability to find potential issues in GURI. Our results showed that GPT-3.5 hallucinates the least, is the most successful, and is generally the most robust; however, it has the worst time efficiency. Our differential testing revealed 22 medical rules where implementation inconsistencies were discovered (e.g., regarding handling rule versions). Finally, we provide insights for practitioners and researchers based on the results.
Erblin Isaku, Christoph Laaber, Hassan Sartaj, Shaukat Ali 0001, Thomas Schwitalla, Jan Nygård
ICST6
2023 Securing Federated GANs: Enabling Synthetic Data Generation for Health Registry Consortiums
abstract
In this work, we review the architecture design of existing federated General Adversarial Networks (GAN) solutions and highlight the security and trust-related weaknesses in the existing designs. We then describe how these weaknesses make existing designs unsuitable for the requirements needed for a consortium of health registries working towards generating synthetic data sets for research purposes. Moreover, we propose how these weaknesses can be addressed with our novel architecture solution. Our architecture solution combines several building blocks to generate synthetic data in a decentralised setting. Federated GANs, Consortium blockchains, and Shamir Secret Sharing algorithm are the core building blocks of our proposed architecture solution. Finally, we discuss our proposed solution’s advantages, disadvantages and future research directions.
Narasimha Raghavan, Jan Nygård
ARES2
2023 Cost Reduction on Testing Evolving Cancer Registry System
abstract
The Cancer Registration Support System (CaReSS), built by the Cancer Registry of Norway (CRN), is a complex real-world socio-technical software system that undergoes continuous evolution in its implementation. Consequently, continuous testing of CaReSS with automated testing tools is needed such that its dependability is always ensured. Towards automated testing of a key software subsystem of CaReSS, i.e., GURI, we present a real-world application of an extension to the open-source tool EvoMaster, which automatically generates test cases with evolutionary algorithms. We named the extension EvoClass, which enhances EvoMaster with a machine learning classifier to reduce the overall testing cost. This is imperative since testing with EvoMaster involves sending many requests to GURI deployed in different environments, including the production environment, whose performance and functionality could potentially be affected by many requests. The machine learning classifier of EvoClass can predict whether a request generated by EvoMaster will be executed successfully or not; if not, the classifier filters out such requests, consequently reducing the number of requests to be executed on GURI. We evaluated EvoClass on ten GURI versions over four years in three environments: development, testing, and production. Results showed that EvoClass can significantly reduce the testing cost of evolving GURI without reducing testing effectiveness (measured as rule coverage) across all three environments, as compared to the default EvoMaster. Overall, EvoClass achieved ≈31% of overall cost reduction. Finally, we report our experiences and lessons learned that are equally valuable for researchers and practitioners.
Erblin Isaku, Hassan Sartaj, Christoph Laaber, Tao Yue 0002, Shaukat Ali 0001, Thomas Schwitalla, Jan Nygård
ICSME7
2023 Automated Test Generation for Medical Rules Web Services: A Case Study at the Cancer Registry of Norway
abstract
The Cancer Registry of Norway (CRN) collects, curates, and manages data related to cancer patients in Norway, supported by an interactive, human-in-the-loop, socio-technical decision support software system. Automated software testing of this software system is inevitable; however, currently, it is limited in CRN’s practice. To this end, we present an industrial case study to evaluate an AI-based system-level testing tool, i.e., EvoMaster, in terms of its effectiveness in testing CRN’s software system. In particular, we focus on GURI, CRN’s medical rule engine, which is a key component at the CRN. We test GURI with EvoMaster’s black-box and white-box tools and study their test effectiveness regarding code coverage, errors found, and domain-specific rule coverage. The results show that all EvoMaster tools achieve a similar code coverage; i.e., around 19% line, 13% branch, and 20% method; and find a similar number of errors; i.e., 1 in GURI’s code. Concerning domain-specific coverage, EvoMaster’s black-box tool is the most effective in generating tests that lead to applied rules; i.e., 100% of the aggregation rules and between 12.86% and 25.81% of the validation rules; and to diverse rule execution results; i.e., 86.84% to 89.95% of the aggregation rules and 0.93% to 1.72% of the validation rules pass, and 1.70% to 3.12% of the aggregation rules and 1.58% to 3.74% of the validation rules fail. We further observe that the results are consistent across 10 versions of the rules. Based on these results, we recommend using EvoMaster’s black-box tool to test GURI since it provides good results and advances the current state of practice at the CRN. Nonetheless, EvoMaster needs to be extended to employ domain-specific optimization objectives to improve test effectiveness further. Finally, we conclude with lessons learned and potential research directions, which we believe are applicable in a general context.
Christoph Laaber, Tao Yue 0002, Shaukat Ali 0001, Thomas Schwitalla, Jan Nygård
ESEC/SIGSOFT FSE5
2023 EvoCLINICAL: Evolving Cyber-Cyber Digital Twin with Active Transfer Learning for Automated Cancer Registry System
abstract
The Cancer Registry of Norway (CRN) collects information on cancer patients by receiving cancer messages from different medical entities (e.g., medical labs, hospitals) in Norway. Such messages are validated by an automated cancer registry system: GURI. Its correct operation is crucial since it lays the foundation for cancer research and provides critical cancer-related statistics to its stakeholders. Constructing a cyber-cyber digital twin (CCDT) for GURI can facilitate various experiments and advanced analyses of the operational state of GURI without requiring intensive interactions with the real system. However, GURI constantly evolves due to novel medical diagnostics and treatment, technological advances, etc. Accordingly, CCDT should evolve as well to synchronize with GURI. A key challenge of achieving such synchronization is that evolving CCDT needs abundant data labelled by the new GURI. To tackle this challenge, we propose EvoCLINICAL, which considers the CCDT developed for the previous version of GURI as the pretrained model and fine-tunes it with the dataset labelled by querying a new GURI version. EvoCLINICAL employs a genetic algorithm to select an optimal subset of cancer messages from a candidate dataset and query GURI with it. We evaluate EvoCLINICAL on three evolution processes. The precision, recall, and F1 score are all greater than 91%, demonstrating the effectiveness of EvoCLINICAL. Furthermore, we replace the active learning part of EvoCLINICAL with random selection to study the contribution of transfer learning to the overall performance of EvoCLINICAL. Results show that employing active learning in EvoCLINICAL increases its performances consistently.
Chengjie Lu, Qinghua Xu, Tao Yue 0002, Shaukat Ali 0001, Thomas Schwitalla, Jan Nygård
ESEC/SIGSOFT FSE6
2022 Matrix factorization for the reconstruction of cervical cancer screening histories and prediction of future screening results
abstract
BACKGROUND: Mass screening programs for cervical cancer prevention in the Nordic countries have strongly reduced cancer incidence and mortality at the population level. An alternative to the current mass screening is a more personalised screening strategy adapting the recommendations to each individual. However, this necessitates reliable risk prediction models accounting for disease dynamics and individual data. Herein we propose a novel matrix factorisation framework to classify females by the time-varying risk of being diagnosed with cervical cancer. We cast the problem as a time-series prediction model where the data from females in the Norwegian screening population are represented as sparse vectors in time and then combined into a single matrix. Using novel temporal regularisation and discrepancy terms for the cervical cancer screening context, we reconstruct complete screening profiles from this scarce matrix and use these to predict the next exam results indicating the risk of cervical cancer. The algorithm is validated on both synthetic and registry screening data by measuring the probability of agreement (PoA) between Kaplan-Meier estimates. RESULTS: In numerical experiments on synthetic data, we demonstrate that the novel regularisation and discrepancy term can improve the data reconstruction ability as well as prediction performance over varying data scarcity. Using a hold-out set of screening data, we compare several numerical models and find that the proposed framework attains the strongest PoA. We observe strong correlations between the empirical survival curves from our method and the hold-out data, and evaluate the ability of our framework to predict the females' next results for up to five years ahead in time using only their current screening histories as input. CONCLUSIONS: We have proposed a matrix factorization model for predicting future screening results and evaluated its performance in a female cohort to demonstrate the potential for developing prediction models for more personalized cervical cancer screening.
Geir Severin R. E. Langberg, Mikal Stapnes, Jan Nygård, Mari Nygård, Markus Grasmair, Valeriya Naumova
BMC Bioinform.3
2021 DeCanSec: A Decentralized Architecture for Secure Statistical Computations on Distributed Health Registry Data
abstract
The architectures presented in the literature, and current practices and solutions for computing statistics on data from health registries distributed across the world are manual and suffers from security and privacy problems. In this paper, we suggest a solution design with a infrastructure architecture providing improved security, automation and privacy guarantees compared to the related works. Our solution builds on top of the key research accomplishments from several areas such as distributed computing, blockchain, cryptography, and medical informatics rather than completely re-inventing the wheel from scratch for the healthcare domain. The proposed architecture is currently being prototyped in the Cancer Registry of Norway.
Narasimha Raghavan, Jan Nygård
ARES2
2019 Automated Refactoring of OCL Constraints with Search
abstract
Object Constraint Language (OCL) constraints are typically used to provide precise semantics to models developed with the Unified Modeling Language (UML). When OCL constraints evolve regularly, it is essential that they are easy to understand and maintain. For instance, in cancer registries, to ensure the quality of cancer data, more than one thousand medical rules are defined and evolve regularly. Such rules can be specified with OCL. It is, therefore, important to ensure the understandability and maintainability of medical rules specified with OCL. To tackle such a challenge, we propose an automated search-based OCL constraint refactoring approach (SBORA) by defining and applying four semantics-preserving refactoring operators (i.e., Context Change, Swap, Split and Merge) and three OCL quality metrics (Complexity, Coupling, and Cohesion) to measure the understandability and maintainability of OCL constraints. We evaluate SBORA along with six commonly used multi-objective search algorithms (e.g., Indicator-Based Evolutionary Algorithm (IBEA)) by employing four case studies from different domains: healthcare (i.e., cancer registry system from Cancer Registry of Norway (CRN)), Oil&Gas (i.e., subsea production systems), warehouse (i.e., handling systems), and an open source case study named SEPA. Results show: 1) IBEA achieves the best performance among all the search algorithms and 2) the refactoring approach along with IBEA can manage to reduce on average 29.25 percent Complexity and 39 percent Coupling and improve 47.75 percent Cohesion, as compared to the original OCL constraint set from CRN. To further test the performance of SBORA, we also applied it to refactor an OCL constraint set specified on the UML 2.3 metamodel and we obtained positive results. Furthermore, we conducted a controlled experiment with 96 subjects and results show that the understandability and maintainability of the original constraint set can be improved significantly from the perspectives of the 96 participants of the controlled experiment.
Hong Lu 0005, Shuai Wang 0001, Tao Yue 0002, Shaukat Ali 0001, Jan Nygård
IEEE Trans. Software Eng.5
2018 Stratifying Cervical Cancer Risk with Registry Data
abstract
The cervical cancer screening programmes in Sweden and Norway have successfully reduced the frequency of cervical cancer incidence but have not implemented any form of evaluation for screening needs. This means that the screening frequency for individuals can be suboptimal, increasing either the cost of the programme or the risk of missing an early stage cancer development. We developed a framework for assessing an individual's risk of cervical cancer based on their available screening history and computing a primary risk factor called CRS from a data-driven separation model together with multiple derived attributes. The results show that this approach is highly practical, validates against multiple established trends, and can be effective in personalizing the screening needs for individuals.
Nicholas Baltzer, Mari Nygård, Karin Sundstrom, Joakim Dillner, Jan Nygård, Jan Komorowski
eScience5
2018 Automated refactoring of OCL constraints with search
abstract
Object Constraint Language (OCL) constraints are typically used for providing precise semantics to models developed with the Unified Modeling Language (UML). When OCL constraints evolve in a regular basis, it is essential that they are easy to understand and maintain. For instance, in cancer registries, to ensure the quality of cancer data, more than one thousand medical rules are defined and evolve regularly. Such rules can be specified with OCL. It is, therefore, important to ensure the understandability and maintainability of medical rules specified with OCL.
Hong Lu 0005, Shuai Wang 0001, Tao Yue 0002, Shaukat Ali 0001, Jan Nygård
ICSE5
2017 RCIA: Automated Change Impact Analysis to Facilitate a Practical Cancer Registry System
abstract
The Cancer Registry of Norway (CRN) employs a cancer registry system to collect cancer patient data (e.g., diagnosis and treatments) from various medical entities (e.g., clinic hospitals). The collected data are then checked for validity (i.e., validation) and assembled as cancer cases (i.e., aggregation) based on more than 1000 cancer coding rules in the system. However, it is frequent in practice that the collected cancer data changes due to various reasons (e.g., different treatments) and the cancer coding rules can also change/evolve due to new medical knowledge. Thus, such a cancer registry system requires an efficient means to automatically analyze these changes and provide consequent impacts to medical experts for further actions. This paper proposes an automated Rule-based Change Impact Analysis (CIA) approach named RCIA that includes: 1) a change classification to capture the potential changes that can occur at CRN; 2) in total 80 change impact analysis rules including 50 dependency rules and 30 impact rules; and 3) an efficient algorithm to analyze changes and produce consequent impacts. We evaluate RCIA via a case study with 12 real change sets from CRN and a conducted interview. The results showed that RCIA managed to produce 100% actual change impacts and the medical expert at CRN is quite positive to apply RCIA to facilitate their cancer registry system. We also shared a set of lessons learned based on the collaboration with CRN.
Shuai Wang 0001, Thomas Schwitalla, Tao Yue 0002, Shaukat Ali 0001, Jan Nygård
ICSME5
2017 IOCL: An interactive tool for specifying, validating and evaluating OCL constraints
Hammad Muhammad, Tao Yue 0002, Shuai Wang 0001, Shaukat Ali 0001, Jan Nygård
Sci. Comput. Program.5
2016 MBF4CR: A Model-Based Framework for Supporting an Automated Cancer Registry System
Shuai Wang 0001, Hong Lu 0005, Tao Yue 0002, Shaukat Ali 0001, Jan Nygård
ECMFA5