Kalpana Raja

dblp:144/0666 · also Raja Kalpana · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-3156-4197ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 SemNovel - A new approach to detecting semantic novelty of biomedical publications using embeddings of large language models
Xueqing Peng, Yutong Xie 0007, Brian D. Ondov, Kalpana Raja, Qijia Liu, Qiaozhu Mei, Hua Xu 0001
J. Biomed. Informatics5
2024 Advancing entity recognition in biomedicine via instruction tuning of large language models
abstract
MOTIVATION: Large Language Models (LLMs) have the potential to revolutionize the field of Natural Language Processing, excelling not only in text generation and reasoning tasks but also in their ability for zero/few-shot learning, swiftly adapting to new tasks with minimal fine-tuning. LLMs have also demonstrated great promise in biomedical and healthcare applications. However, when it comes to Named Entity Recognition (NER), particularly within the biomedical domain, LLMs fall short of the effectiveness exhibited by fine-tuned domain-specific models. One key reason is that NER is typically conceptualized as a sequence labeling task, whereas LLMs are optimized for text generation and reasoning tasks. RESULTS: We developed an instruction-based learning paradigm that transforms biomedical NER from a sequence labeling task into a generation task. This paradigm is end-to-end and streamlines the training and evaluation process by automatically repurposing pre-existing biomedical NER datasets. We further developed BioNER-LLaMA using the proposed paradigm with LLaMA-7B as the foundational LLM. We conducted extensive testing on BioNER-LLaMA across three widely recognized biomedical NER datasets, consisting of entities related to diseases, chemicals, and genes. The results revealed that BioNER-LLaMA consistently achieved higher F1-scores ranging from 5% to 30% compared to the few-shot learning capabilities of GPT-4 on datasets with different biomedical entities. We show that a general-domain LLM can match the performance of rigorously fine-tuned PubMedBERT models and PMC-LLaMA, biomedical-specific language model. Our findings underscore the potential of our proposed paradigm in developing general-domain LLMs that can rival SOTA performances in multi-task, multi-domain scenarios in biomedical and health applications. AVAILABILITY AND IMPLEMENTATION: Datasets and other resources are available at https://github.com/BIDS-Xu-Lab/BioNER-LLaMA.
Vipina Kuttichi Keloth, Qianqian Xie, Xueqing Peng, Yan Wang 0015, Andrew Zheng, Melih Selek, Kalpana Raja, Chih-Hsuan Wei, Qiao Jin 0001, Zhiyong Lu, Qingyu Chen 0001, Hua Xu 0001
Bioinform.8
2023 Towards precise PICO extraction from abstracts of randomized controlled trials using a section-specific learning approach
abstract
MOTIVATION: Automated extraction of participants, intervention, comparison/control, and outcome (PICO) from the randomized controlled trial (RCT) abstracts is important for evidence synthesis. Previous studies have demonstrated the feasibility of applying natural language processing (NLP) for PICO extraction. However, the performance is not optimal due to the complexity of PICO information in RCT abstracts and the challenges involved in their annotation. RESULTS: We propose a two-step NLP pipeline to extract PICO elements from RCT abstracts: (i) sentence classification using a prompt-based learning model and (ii) PICO extraction using a named entity recognition (NER) model. First, the sentences in abstracts were categorized into four sections namely background, methods, results, and conclusions. Next, the NER model was applied to extract the PICO elements from the sentences within the title and methods sections that include >96% of PICO information. We evaluated our proposed NLP pipeline on three datasets, the EBM-NLPmoddataset, a randomly selected and reannotated dataset of 500 RCT abstracts from the EBM-NLP corpus, a dataset of 150 COVID-19 RCT abstracts, and a dataset of 150 Alzheimer's disease (AD) RCT abstracts. The end-to-end evaluation reveals that our proposed approach achieved an overall micro F1 score of 0.833 on the EBM-NLPmod dataset, 0.928 on the COVID-19 dataset, and 0.899 on the AD dataset when measured at the token-level and an overall micro F1 score of 0.712 on EBM-NLPmod dataset, 0.850 on the COVID-19 dataset, and 0.805 on the AD dataset when measured at the entity-level. AVAILABILITY: Our codes and datasets are publicly available at https://github.com/BIDS-Xu-Lab/section_specific_annotation_of_PICO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vipina Kuttichi Keloth, Kalpana Raja, Yong Chen 0016, Hua Xu 0001
Bioinform.3
2023 Serial KinderMiner (SKiM) discovers and annotates biomedical knowledge using co-occurrence and transformer models
abstract
BACKGROUND: The PubMed archive contains more than 34 million articles; consequently, it is becoming increasingly difficult for a biomedical researcher to keep up-to-date with different knowledge domains. Computationally efficient and interpretable tools are needed to help researchers find and understand associations between biomedical concepts. The goal of literature-based discovery (LBD) is to connect concepts in isolated literature domains that would normally go undiscovered. This usually takes the form of an A-B-C relationship, where A and C terms are linked through a B term intermediate. Here we describe Serial KinderMiner (SKiM), an LBD algorithm for finding statistically significant links between an A term and one or more C terms through some B term intermediate(s). The development of SKiM is motivated by the observation that there are only a few LBD tools that provide a functional web interface, and that the available tools are limited in one or more of the following ways: (1) they identify a relationship but not the type of relationship, (2) they do not allow the user to provide their own lists of B or C terms, hindering flexibility, (3) they do not allow for querying thousands of C terms (which is crucial if, for instance, the user wants to query connections between a disease and the thousands of available drugs), or (4) they are specific for a particular biomedical domain (such as cancer). We provide an open-source tool and web interface that improves on all of these issues. RESULTS: We demonstrate SKiM's ability to discover useful A-B-C linkages in three control experiments: classic LBD discoveries, drug repurposing, and finding associations related to cancer. Furthermore, we supplement SKiM with a knowledge graph built with transformer machine-learning models to aid in interpreting the relationships between terms found by SKiM. Finally, we provide a simple and intuitive open-source web interface ( https://skim.morgridge.org ) with comprehensive lists of drugs, diseases, phenotypes, and symptoms so that anyone can easily perform SKiM searches. CONCLUSIONS: SKiM is a simple algorithm that can perform LBD searches to discover relationships between arbitrary user-defined concepts. SKiM is generalized for any domain, can perform searches with many thousands of C term concepts, and moves beyond the simple identification of an existence of a relationship; many relationships are given relationship type labels from our knowledge graph.
Robert J. Millikin, Kalpana Raja, John W. Steill, Cannon Lock, Xuancheng Tu, Lam C. Tsoi, Finn Kuusisto, Zijian Ni, Miron Livny, Brian Bockelman, James A. Thomson, Ron M. Stewart
BMC Bioinform.2
2023 Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001
J. Biomed. Informatics12
2021 Advancement in predicting interactions between drugs used to treat psoriasis and its comorbidities by integrating molecular and clinical resources
abstract
OBJECTIVE: Drug-drug interactions (DDIs) can result in adverse and potentially life-threatening health consequences; however, it is challenging to predict potential DDIs in advance. We introduce a new computational approach to comprehensively assess the drug pairs which may be involved in specific DDI types by combining information from large-scale gene expression (984 transcriptomic datasets), molecular structure (2159 drugs), and medical claims (150 million patients). MATERIALS AND METHODS: Features were integrated using ensemble machine learning techniques, and we evaluated the DDIs predicted with a large hospital-based medical records dataset. Our pipeline integrates information from >30 different resources, including >10 000 drugs and >1.7 million drug-gene pairs. We applied our technique to predict interactions between 37 611 drug pairs used to treat psoriasis and its comorbidities. RESULTS: Our approach achieves >0.9 area under the receiver operator curve (AUROC) for differentiating 11 861 known DDIs from 25 750 non-DDI drug pairs. Significantly, we demonstrate that the novel DDIs we predict can be confirmed through independent data sources and supported using clinical medical records. CONCLUSIONS: By applying machine learning and taking advantage of molecular, genomic, and health record data, we are able to accurately predict potential new DDIs that can have an impact on public health.
Matthew T. Patrick, Redina Bardhi, Kalpana Raja, Kevin He, Lam C. Tsoi
J. Am. Medical Informatics Assoc.3
2016 Classification of clinically useful sentences in clinical evidence resources
Mohammad Amin Morid, Marcelo Fiszman, Kalpana Raja, Siddhartha Jonnalagadda, Guilherme Del Fiol
J. Biomed. Informatics3
2015 Classification of Clinically Useful Sentences in MEDLINE
Mohammad Amin Morid, Siddhartha Jonnalagadda, Marcelo Fiszman, Kalpana Raja, Guilherme Del Fiol
AMIA4
2015 Automated Citation Retrieval System for Clinical Knowledge Management
Kalpana Raja, Andrew J. Sauer, Melanie R. Klerer, Siddhartha Jonnalagadda
AMIA1
2015 Agile text mining for the 2014 i2b2/UTHealth Cardiac risk factors challenge
abstract
This paper describes the use of an agile text mining platform (Linguamatics' Interactive Information Extraction Platform, I2E) to extract document-level cardiac risk factors in patient records as defined in the i2b2/UTHealth 2014 challenge. The approach uses a data-driven rule-based methodology with the addition of a simple supervised classifier. We demonstrate that agile text mining allows for rapid optimization of extraction strategies, while post-processing can leverage annotation guidelines, corpus statistics and logic inferred from the gold standard data. We also show how data imbalance in a training set affects performance. Evaluation of this approach on the test data gave an F-Score of 91.7%, one percent behind the top performing system.
James Cormack, Chinmoy Nath, David Milward, Kalpana Raja, Siddhartha Jonnalagadda
J. Biomed. Informatics4
2015 HPIminer: A text mining system for building and visualizing human protein interaction networks and pathways
abstract
The knowledge on protein-protein interactions (PPI) and their related pathways are equally important to understand the biological functions of the living cell. Such information on human proteins is highly desirable to understand the mechanism of several diseases such as cancer, diabetes, and Alzheimer's disease. Because much of that information is buried in biomedical literature, an automated text mining system for visualizing human PPI and pathways is highly desirable. In this paper, we present HPIminer, a text mining system for visualizing human protein interactions and pathways from biomedical literature. HPIminer extracts human PPI information and PPI pairs from biomedical literature, and visualize their associated interactions, networks and pathways using two curated databases HPRD and KEGG. To our knowledge, HPIminer is the first system to build interaction networks from literature as well as curated databases. Further, the new interactions mined only from literature and not reported earlier in databases are highlighted as new. A comparative study with other similar tools shows that the resultant network is more informative and provides additional information on interacting proteins and their associated networks.
Suresh Subramani, Kalpana Raja, Pankaj Moses Monickaraj, Jeyakumar Natarajan
J. Biomed. Informatics2
2014 ProNormz - An integrated approach for human proteins and protein kinases normalization
Suresh Subramani, Kalpana Raja, Jeyakumar Natarajan
J. Biomed. Informatics2