VLDB 2026 Research / reviewers in the wild / expert
Upulee Kanewala
dblp:139/3849
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-5954-1616ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Robustness of Large Language Models Used in Healthcare Through Prompt Engineering
Jonathan O'Berry, Indika Kahanda, Upulee Kanewala |
ICMLA | 3 |
| 2025 | Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPTabstractLarge Language Models (LLMs) have made significant strides in Natural Language Processing but remain vulnerable to fairness-related issues, often reflecting biases inherent in their training data. These biases pose risks, particularly when LLMs are deployed in sensitive areas such as healthcare, finance, and law. This paper introduces a metamorphic testing approach to systematically identify fairness bugs in LLMs. We define and apply a set of fairness-oriented metamorphic relations (MRs) to assess the LLaMA and GPT model, a state-of-the-art LLM, across diverse demographic inputs. Our methodology includes generating source and follow-up test cases for each MR and analyzing model responses for fairness violations. The results demonstrate the effectiveness of MT in exposing bias patterns, especially in relation to tone and sentiment, and highlight specific intersections of sensitive attributes that frequently reveal fairness faults. This research improves fairness testing in LLMs, providing a structured approach to detect and mitigate biases and improve model robustness in fairness-sensitive applications. Harishwar Reddy, Madhusudan Srinivasan, Upulee Kanewala |
SERA | 3 |
| 2025 | Testing research software: an in-depth survey of practices, methods, and tools
Nasir U. Eisty, Upulee Kanewala, Jeffrey C. Carver |
Empir. Softw. Eng. | 2 |
| 2024 | Poster: Towards Understanding Root Causes of Real Failures in Healthcare Machine Learning ApplicationsabstractMachine learning (ML) is widely used in healthcare applications to diagnose diseases, forecast disease progression, develop personalized treatment plans, and aid in drug discovery and development [1]. The development of ML applications is inherently different from other applications. Instead of explicitly coding the program's logic, ML applications learn this logic using a machine learning algorithm and provided data. Thus, faults in ML applications, as opposed to in others, can manifest in all these components, such as the application itself, incorrect use of the machine learning algorithms or libraries, and issues with data used for training. Thus, understanding these various root causes of real faults would help to develop effective testing techniques for these applications. Therefore, we analyzed 50 real-life faults from four ML healthcare applications to better understand the faults presented in this domain. Guna Sekaran Jaganathan, Nazmul Kazi, Indika Kahanda, Upulee Kanewala |
ICST | 4 |
| 2024 | MLHCBugs: A Framework to Reproduce Real Faults in Healthcare Machine Learning ApplicationsabstractMachine Learning (ML) is the field of study that allows computers to learn from experiences without being explicitly programmed [1]. ML models are currently used in many safety-critical applications in healthcare [2]–[4] and survival analyses [5]. Thus, faults in this software can directly impact the quality of human life. In an ML application, the program logic is typically derived by a ML algorithm using the currently available data (i.e., training data) rather than explicitly being programmed [6]. Therefore, the program's behavior would evolve as it is exposed to new data. Further, healthcare ML applications are inherently complex and typically constructed by the interconnection of several components, such as data that is used to derive the logic, the ML framework that contains the algorithms used by the program, and the program itself that is written by the programmer for a specific task involved with healthcare [7]. Faults in any of these components may produce an observable incorrect output or the statistical nature of these programs may mask the incorrect output altogether, making it more challenging to understand the root causes of these failures. Guna Sekaran Jaganathan, Nazmul Kazi, Indika Kahanda, Upulee Kanewala |
ICST | 4 |
| 2023 | A Test Suite Minimization Technique for Testing Numerical ProgramsabstractMetamorphic testing is a technique that uses metamorphic relations (i.e., necessary properties of the software under test), to construct new test cases (i.e., follow-up test cases), from existing test cases (i.e., source test cases). Metamorphic testing allows for the verification of testing results without the need of test oracles (a mechanism to detect the correctness of the outcomes of a program), and it has been widely used in many application domains to detect real-world faults. Numerous investigations have been conducted to further improve the effectiveness of metamorphic testing. Recent studies have emerged suggesting a new research direction on the generation and selection of source test cases that are effective in fault detection. Herein, we present two important findings: i) a mutant reduction strategy that is applied to increase the testing efficiency of source test cases, and ii) a test suite minimization technique to help reduce the testing costs without trading off fault-finding effectiveness. To validate our results, an empirical study was conducted to demonstrate the increase in efficiency and fault-finding effectiveness of source test cases. The results from the experiment provide evidence to support our claims. Prashanta Saha, Clemente Izurieta, Upulee Kanewala |
SERA | 3 |
| 2022 | Using Metamorphic Relations to Improve The Effectiveness of Automatically Generated Test CasesabstractAutomated test case generation has helped to reduce the cost of testing. However, developing effective test oracles for these automatically generated test cases still remains a challenge. Metamorphic testing (MT) has become a well-known software testing approach over the years. This testing technique can effectively alleviate the oracle problem faced when testing using metamorphic relations (MRs) to determine whether a test case is passed or failed. In this work, we conduct an empirical study on an open source linear algebra library to evaluate whether MRs can be utilized to improve the fault detection effectiveness of automatically generated test cases. Our experiment suggests that MRs can help to improve the fault detection effectiveness of automatically generated test cases. Prashanta Saha, Upulee Kanewala |
SERA | 2 |
| 2022 | Metamorphic relation prioritization for effective regression testingabstractSummary Metamorphic testing (MT) is widely used for testing programs that face the oracle problem. It uses a set of metamorphic relations (MRs), which are relations among multiple inputs and their corresponding outputs to determine whether the program under test is faulty. Typically, MRs vary in their ability to detect faults in the program under test, and some MRs tend to detect the same set of faults. In this paper, we propose approaches to prioritize MRs to improve the efficiency and effectiveness of MT for regression testing. We present two MR prioritization approaches: (i) fault‐based and (ii) coverage‐based. To evaluate these MR prioritization approaches, we conduct experiments on three complex open‐source software systems. Our results show that the MR prioritization approaches developed by us significantly outperform the current practice of executing the source and follow‐up test cases of the MRs in an ad‐hoc manner in terms of fault detection effectiveness. Further, fault‐based MR prioritization leads to reducing the number of source and follow‐up test cases that needs to be executed as well as reducing the average time taken to detect a fault, which would result in saving time and cost during the testing process. Madhusudan Srinivasan, Upulee Kanewala |
Softw. Test. Verification Reliab. | 2 |
| 2021 | Contextual Understanding and Improvement of Metamorphic Testing in Scientific Software DevelopmentabstractBackground: Metamorphic testing emerges as a simple and effective approach for testing scientific software; yet, its adoption in actual scientific software projects is less studied. Zedong Peng, Upulee Kanewala, Nan Niu |
ESEM | 2 |
| 2016 | Predicting metamorphic relations for testing scientific software: a machine learning approach using graph kernelsabstractSummary Comprehensive, automated software testing requires an oracle to check whether the output produced by a test case matches the expected behaviour of the programme. But the challenges in creating suitable oracles limit the ability to perform automated testing in some programmes, and especially in scientific software. Metamorphic testing is a method for automating the testing process for programmes without test oracles. This technique operates by checking whether the programme behaves according to properties called metamorphic relations. A metamorphic relation describes the change in output when the input is changed in a prescribed way. Unfortunately, finding the metamorphic relations satisfied by a programme or function remains a labour‐intensive task, which is generally performed by a domain expert or a programmer. In this work, we propose a machine learning approach for predicting metamorphic relations that uses a graph‐based representation of a programme to represent control flow and data dependency information. In earlier work, we found that simple features derived from such graphs provide good performance. An analysis of the features used in this earlier work led us to explore the effectiveness of several representations of those graphs using the machine learning framework of graph kernels, which provide various ways of measuring similarity between graphs. Our results show that a graph kernel that evaluates the contribution of all paths in the graph has the best accuracy and that control flow information is more useful than data dependency information. The data used in this study are available for download at http://www.cs.colostate.edu/saxs/MRpred/functions.tar.gz to help researchers in further development of metamorphic relation prediction methods. Copyright © 2015 John Wiley & Sons, Ltd. Upulee Kanewala, James M. Bieman, Asa Ben-Hur |
Softw. Test. Verification Reliab. | 1 |
| 2014 | Testing scientific software: A systematic literature review
Upulee Kanewala, James M. Bieman |
Inf. Softw. Technol. | 1 |
| 2013 | Using machine learning techniques to detect metamorphic relations for programs without test oraclesabstractMuch software lacks test oracles, which limits automated testing. Metamorphic testing is one proposed method for automating the testing process for programs without test oracles. Unfortunately, finding the appropriate metamorphic relations required for use in metamorphic testing remains a labor intensive task, which is generally performed by a domain expert or a programmer. In this work we present a novel approach for automatically predicting metamorphic relations using machine learning techniques. Our approach uses a set of features developed using the control flow graph of a function for predicting likely metamorphic relations. We show the effectiveness of our method using a set of real world functions often used in scientific applications. Upulee Kanewala, James M. Bieman |
ISSRE | 1 |