VLDB 2026 Research / reviewers in the wild / expert
Indika Kahanda
dblp:49/8344
· DBLP profile ↗
17ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0002-4536-6917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Robustness of Large Language Models Used in Healthcare Through Prompt Engineering
Jonathan O'Berry, Indika Kahanda, Upulee Kanewala |
ICMLA | 2 |
| 2025 | Genotype-to-Phenotype Associations in Yeast with Frequented Region Variants and Deep LearningabstractPhenotypes are the observable characteristics of an individual organism. Predicting quantitative phenotypes from genomic variation remains challenging when causal signals span both local motifs and distal regulatory context. Building on Frequented Regions (FRs), which represent subsequences conserved across genomes, extracted from a pangenome graph, we compare six modeling strategies on five yeast growth phenotypes: Random Forest (RF) on FR counts, RF on FR sequences, 1D convolutional neural networks on FR sequences, Long Short-Term Memory (LSTM) on FR sequences, a Genome-wide Association Study (GWAS) baseline, and a sequence-based Enformer model trained on raw FR nucleotide windows. Across the five phenotypes, all sequence-based baselines improve upon RF (FR counts) and GWAS, confirming the value of sequence context. Enformer consistently outperforms CNN/LSTM on all five phenotypes and surpasses RF (FR-sequences) on three of five, while remaining competitive on the others. These results indicate that when long-range dependencies contribute to trait variation, transformer-based modeling of raw sequence windows may yield tangible gains over k-mer and local-pattern learners; conversely, for phenotypes dominated by short-range signals, lightweight baselines may remain competitive. These findings suggest that while short-range motif statistics can suffice for certain phenotypes, deep learning architectures that integrate positional context and distal interactions can yield additional gains, particularly when phenotypic variation is linked to dispersed regulatory signals. Tejaswi Vemuri, Trung Dinh, Thiruvarangan Ramaraj, Joann Mudge, Brendan Mumey, Indika Kahanda |
ICMLA | 7 |
| 2024 | Poster: Towards Understanding Root Causes of Real Failures in Healthcare Machine Learning ApplicationsabstractMachine learning (ML) is widely used in healthcare applications to diagnose diseases, forecast disease progression, develop personalized treatment plans, and aid in drug discovery and development [1]. The development of ML applications is inherently different from other applications. Instead of explicitly coding the program's logic, ML applications learn this logic using a machine learning algorithm and provided data. Thus, faults in ML applications, as opposed to in others, can manifest in all these components, such as the application itself, incorrect use of the machine learning algorithms or libraries, and issues with data used for training. Thus, understanding these various root causes of real faults would help to develop effective testing techniques for these applications. Therefore, we analyzed 50 real-life faults from four ML healthcare applications to better understand the faults presented in this domain. Guna Sekaran Jaganathan, Nazmul Kazi, Indika Kahanda, Upulee Kanewala |
ICST | 3 |
| 2024 | MLHCBugs: A Framework to Reproduce Real Faults in Healthcare Machine Learning ApplicationsabstractMachine Learning (ML) is the field of study that allows computers to learn from experiences without being explicitly programmed [1]. ML models are currently used in many safety-critical applications in healthcare [2]–[4] and survival analyses [5]. Thus, faults in this software can directly impact the quality of human life. In an ML application, the program logic is typically derived by a ML algorithm using the currently available data (i.e., training data) rather than explicitly being programmed [6]. Therefore, the program's behavior would evolve as it is exposed to new data. Further, healthcare ML applications are inherently complex and typically constructed by the interconnection of several components, such as data that is used to derive the logic, the ML framework that contains the algorithms used by the program, and the program itself that is written by the programmer for a specific task involved with healthcare [7]. Faults in any of these components may produce an observable incorrect output or the statistical nature of these programs may mask the incorrect output altogether, making it more challenging to understand the root causes of these failures. Guna Sekaran Jaganathan, Nazmul Kazi, Indika Kahanda, Upulee Kanewala |
ICST | 3 |
| 2023 | Genotype-to-Phenotype Associations with Frequented Region VariantsabstractA pangenome represents the entire sequence content and variation of a population. As collections of complete reference quality genomes become more common, so does the prevalence of pangenomes, necessitating the need for scalable computational methods for their analysis. Previously, we developed FindFRs for identifying Frequented Regions in pangenome graphs, where a Frequented Region is a subgraph that is frequently traversed by multiple sequences. In this work, we propose FindFRs3, which is an updated version of FindFRs capable of identifying Frequented Regions with improved runtime and memory efficiency, enabling the analysis of much larger pangenome graphs. In addition, FindFRs3 identifies Frequented Region Variants (the unique subpaths through each region). We demonstrate the utility of these variants by using them as input features for machine learning models that can predict genotype-to-phenotype associations in a large yeast pangenome. Biological insights gained from these variants show that this novel technique allows for a more nuanced and detailed analysis of larger pangenomes. Indika Kahanda, Buwani Manuweera, Brendan Mumey, Thiruvarangan Ramaraj, Alan M. Cleary, Joann Mudge |
BIBM | 1 |
| 2023 | UNF-IDT: Automated Irony Detection in English TweetsabstractThe usage of social media platforms has increased tremendously in the past decade. Twitter is one of the most popular platforms for sharing opinions and feelings on various topics, organizations rely on Twitter data to analyze and gather insights for their businesses using Natural Language Processing (NLP) techniques. However, there are various challenges in analyzing such a huge volume of tweets that are in text format. One such challenge is to identify irony in tweets, which has a significant impact on analyzing sentiments. To overcome this challenge, we propose a solution that automatically recognizes the presence of irony in text and classifies the type of irony. The solution consists of two tasks: Task A performs binary classification to annotate whether irony is expressed or not in each tweet, while Task B performs multi-class classification to classify the type of irony expressed in the tweets. We used the dataset from the “SemEval-2018 Task 3: Irony detection in English tweets” challenge to train and test the proposed solution. The dataset contains 3,817 English tweets for training and 784 English tweets for testing both Tasks A and B. We developed different machine learning models by leveraging traditional classifiers, neural networks, and large language models. Our UNF-IDT (Irony Detector in Text) model developed using BERT (Bidirectional Encoder Representations from Transformers) was able to achieve an F1 score of 0.757 for Task A (Binary Classification) and an F1 score of 0.449 for Task B (Multi-class Classification), retrospectively obtaining 4th and 9th ranks, in the two tasks respectively. Guna Sekaran Jaganathan, Grentina Kilungeja, Indika Kahanda |
ICMLA | 3 |
| 2023 | Enhancing Transfer Learning of LLMs through Fine- Tuning on Task - Related Corpora for Automated Short-Answer GradingabstractAutomated short-answer grading (ASAG) is a cru-cial element of any intelligent tutoring platform. Machine Learning (ML) has shown great promise for ASAG. However, this task remains challenging even for Deep Learning (DL) approaches and Large Language Models (LLMs), requiring semantic inference and textual entailment recognition. The SemEval-2013 Task 7, The Joint Student Response Analysis and 8th Recognizing Textual Entailment Challenge, is a benchmark widely used for research on ASAG. The SciEntsBank data included in this collection contains nearly 11,000 answers to 197 assessment questions in 15 different science domains. Despite the popularity, only a few researchers have explored the potential of DL or LLMs for this task. In this project, we explore the effectiveness of the RoBERTa Large model, an LLM trained on an extensive text corpus for language comprehension. By fine-tuning the model on the Multi-Genre Natural Language Inference (MNLI) corpus for semantic inference and subsequently on the SciEntsBank dataset, with a focus on the 3-way labels of correct, incorrect, and contradictory, we achieved a weighted Fl-score of 0.77, 0.72, and 0.72 on unseen answers, questions, and domains, respectively. Notably, our model significantly benefits from fine-tuning on the MNLI corpus, particularly in enhancing its performance on the contradictory class (which constitutes only 10% of the dataset) through transfer learning leading to significant improvements on the more challenging test sets: unseen questions and unseen domains. Nazmul Kazi, Indika Kahanda |
ICMLA | 2 |
| 2023 | Zero-Shot Information Extraction with Community-Fine-Tuned Large Language Models From Open-Ended Interview TranscriptsabstractMachine learning holds significant promise for automating and optimizing text data analysis. However, resource-intensive tasks like data annotation, model training, and parameter tuning often limit its practicality for one-time data extraction, medium-sized datasets, or short-term projects. There are many community-fine-tuned large language models (CLLMs) that are fine-tuned on task-specific datasets and can demonstrate impressive performance on unseen data without further fine-tuning. Adopting a hybrid approach of leveraging CLLMs for rapid text data extraction and subsequently hand-curating the inaccurate outputs can yield high-quality results, workload balance, and improved efficiency. This project applies CLLMs to three tasks involving the analysis of open-ended survey responses: semantic text matching, exact answer extraction, and sentiment analysis. We present our overall process and discuss several seemingly simple yet effective techniques that we employ to improve model performance without fine-tuning the CLLMs on our own data. Our results demonstrate high precision in semantic text matching (0.92) and exact answer extraction (0.90), while the sentiment analysis model shows room for improvement (precision: 0.65, recall: 0.94, F1: 0.77). This study showcases the potential of CLLMs in open-ended survey text data analysis, particularly in scenarios with limited resources and scarce labeled data. Nazmul Kazi, Indika Kahanda, S. Indu Rupassara, John W. Kindt |
ICMLA | 2 |
| 2023 | Enhancing Health Information Retrieval with Large Language Models: A Study on MedQuAD DatasetabstractGiven the enormous volume of textual data generated in the healthcare sector, effective and accurate retrieval systems are essential. A major challenge is presented by the explosive growth of scientific publications and medical information. In this study, an advanced pipeline was developed to enhance information retrieval in the healthcare domain. The pipeline has two components: information retrieval and evaluation. The information retrieval component is composed of the retriever, which uses BM25, and the reader, which is powered by a pre-trained Large Language Model. The evaluation component, which uses standardized dataset formats such as SQuAD, provides a framework for evaluating system performance and comparing different parameters. Based on the Cancer category in the MedQuAD dataset, the retriever component showed a strong recall of 0.881 and a Mean Reciprocal Rank score of 0.804, demonstrating its effectiveness in retrieving relevant information and accurate ranking. A Semantic Answer Similarity score of 0.677 for the reader component indicates room for improvement. This work has implications for healthcare providers and the text-mining community working in health information retrieval. Prajwol Lamichhane, Indika Kahanda |
ICMLA | 2 |
| 2023 | Patronizing and Condescending Language DetectionabstractPatronizing and Condescending Language (PCL) is the language used by an individual that denotes a superior attitude toward other individuals, especially those that are members of minority or marginalized groups. Due to the prevalence of patronizing and condescending language in today's society, there has been a focus on attempting to automate its detection and classification. In this study, we develop classifiers for this problem using the data from the previously concluded “SemEval 2022 Task 4: Patronizing and Condescending Language Detection” competition. We explore the implementation of three traditional machine learning algorithms and three transformer-based algorithms for both binary and multi-label PCL classification. Our results are in line with that of the original SemEval findings and help demonstrate the need for additional work. The development of a larger, more-balanced dataset would ensure more consistent and transferable results. Conrad Testagrose, Athlene V. Jones, Indika Kahanda |
ICMLA | 3 |
| 2022 | Impact of Concatenation of Digital Craniocaudal Mammography Images on a Deep-Learning Breast-Density Classifier Using Inception-V3 and ViTabstractBreast density is an indicator of a patient’s predisposed risk of breast cancer. Although not fully understood, increased breast density increases the likelihood of developing breast cancer. Accurate assessment of breast density from mammogram images is a challenging task for the radiologist. A patient’s breast density is assigned to one of four categories outlined by Breast Imaging and Reporting Data Systems (BIRADS). There have been efforts to identify automated approaches to assist radiologists in the classification of a patient’s breast density. The interest in using deep learning to fulfill this need for an automated approach has seen a significant increase in recent years. The preprocessing techniques used to develop these deep learning approaches often have a profound impact on the model’s accuracy and clinical viability. In this paper, we outline a novel image preprocessing technique where we concatenate individual mammogram images and compare the results using this technique between Inception-v3 and a vision transformer (ViT). The results are compared using the area under (AUC) the receiver operator characteristics (ROC) curves and traditional accuracy metrics. Conrad Testagrose, Vikash Gupta, Barbaros S. Erdal, Robert W. Maxwell, Xudong Liu 0003, Indika Kahanda, Sherif Elfayoumy, William Klostermeyer, Mutlu Demirer |
BIBM | 7 |
| 2021 | DeepPPPred: Deep Ensemble Learning with Transformers, Recurrent and Convolutional Neural Networks for Human Protein-Phenotype Co-mention ClassificationabstractThe extensive collection of biomedical literature is arguably the best source of knowledge and information on the latest scientific findings and fundamental problems for the biological and clinical communities. However, these articles contain unstructured text; therefore, this valuable knowledge may remain hidden without manual curation, which is tedious and time-consuming due to the rapid growth of publication. The relationships and associations between human proteins and phenotypic abnormalities associated with human disease are one such area of valuable knowledge. This situation calls for the development of accurate computational tools capable of automatically inferring these associations from text data, assisting human curators in expediting their triage and information extraction tasks. This work develops DeepPPPred, a deep ensemble learning model for protein-phenotype co-mention classification at the sentence level. In particular, DeepPPPred combines Support Vector Machines, Transformer models, Recurrent Neural Networks, and Convolutional Neural Networks via stacking. Our experimental results obtained using a manually curated gold-standard dataset demonstrate that DeepPPPred can provide state-of-the-art performance while outperforming all its competitors. This is the first study that develops deep learning models for the problem of classifying human protein-phenotype co-mentions. Our findings have implications for the biological and clinical communities and text mining and natural language processing developers working on biomedical relation extraction. Morteza Pourreza Shahri, Katrina Lyon, Julia Schearer, Indika Kahanda |
BIBM | 4 |
| 2021 | BioSGAN: Protein-Phenotype Co-mention Classification Using Semi-Supervised Generative Adversarial NetworksabstractValuable and relevant information that relates human proteins with their phenotypes in biomedical literature stays hidden from biomedical scientists due to the rapid rise in biomedical publications. Previous studies that developed computational methods to extract this knowledge mostly rely on rule-based linguistic patterns and supervised machine learning approaches. In this work, we propose the use of generative adversarial networks to develop a novel method called BioSGAN for the protein-phenotype co-mention classification task. We demonstrate the potential associated with combining a small labeled dataset with vast unlabelled biomedical text data extracted from Medline abstracts and PubMed Central open Access full-text in a semi-supervised machine learning framework. Our method achieves state-of-the-art performance for classifying the validity of a given sentence-level co-mention of a human protein and phenotype by convincingly outperforming a traditional machine learning-based counterpart. These findings have implications for biocurators, researchers, and the text mining community involved with biomedical relation extraction. Francis Anokye, Indika Kahanda |
CBMS | 2 |
| 2021 | Deep semi-supervised learning ensemble framework for classifying co-mentions of human proteins and phenotypesabstractBACKGROUND: Identifying human protein-phenotype relationships has attracted researchers in bioinformatics and biomedical natural language processing due to its importance in uncovering rare and complex diseases. Since experimental validation of protein-phenotype associations is prohibitive, automated tools capable of accurately extracting these associations from the biomedical text are in high demand. However, while the manual annotation of protein-phenotype co-mentions required for training such models is highly resource-consuming, extracting millions of unlabeled co-mentions is straightforward. RESULTS: In this study, we propose a novel deep semi-supervised ensemble framework that combines deep neural networks, semi-supervised, and ensemble learning for classifying human protein-phenotype co-mentions with the help of unlabeled data. This framework allows the ability to incorporate an extensive collection of unlabeled sentence-level co-mentions of human proteins and phenotypes with a small labeled dataset to enhance overall performance. We develop PPPredSS, a prototype of our proposed semi-supervised framework that combines sophisticated language models, convolutional networks, and recurrent networks. Our experimental results demonstrate that the proposed approach provides a new state-of-the-art performance in classifying human protein-phenotype co-mentions by outperforming other supervised and semi-supervised counterparts. Furthermore, we highlight the utility of PPPredSS in powering a curation assistant system through case studies involving a group of biologists. CONCLUSIONS: This article presents a novel approach for human protein-phenotype co-mention classification based on deep, semi-supervised, and ensemble learning. The insights and findings from this work have implications for biomedical researchers, biocurators, and the text mining community working on biomedical relationship extraction. Morteza Pourreza Shahri, Indika Kahanda |
BMC Bioinform. | 2 |
| 2020 | miRNAFinder: A pre-microRNA classifier for plants and analysis of feature impactabstractMicroRNAs (miRNAs) are endogenous small noncoding RNAs that play an important role in post-transcriptional gene regulation. Several machine learning-based studies have been conducted for miRNA identification with the use of miRNA features. It is difficult to classify real and pseudo-pre-miRNAs in plant species than that in animals since plant pre-miRNAs are more diverse than the animal pre-miRNAs. Therefore, this study is focused on classifying real and pseudo precursor miRNAs (pre-miRNAs) in plants. We have introduced a machine learning model based on a 280 feature set including compositional, sequence-based, and thermodynamic features. Classification performance is tested and compared, considering different feature sets and four different classifiers. Random forest classifier results in the best classification performance with all 280 features with a 97% accuracy for the testing dataset. Puwasuru Ihalagedara, Sandali Lokuge, Shyaman Jayasundara, Damayanthi Herath, Indika Kahanda |
CIBCB | 5 |
| 2019 | Exploring Frequented Regions in Pan-Genomic GraphsabstractWe consider the problem of identifying regions within a pan-genome De Bruijn graph that are traversed by many sequence paths. We define such regions and the subpaths that traverse them as frequented regions (FRs). In this work, we formalize the FR problem and describe an efficient algorithm for finding FRs. Subsequently, we propose some applications of FRs based on machine-learning and pan-genome graph simplification. We demonstrate the effectiveness of these applications using data sets for the organisms Staphylococcus aureus (bacterium) and Saccharomyces cerevisiae (yeast). We corroborate the biological relevance of FRs such as identifying introgressions in yeast that aid in alcohol tolerance, and show that FRs are useful for classification of yeast strains by industrial use and visualizing pan-genomic space. Alan M. Cleary, Thiruvarangan Ramaraj, Indika Kahanda, Joann Mudge, Brendan Mumey |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2009 | Using Transactional Information to Predict Link Strength in Online Social Networks
Indika Kahanda, Jennifer Neville |
ICWSM | 1 |