Vipina Kuttichi Keloth

dblp:234/3249 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-6919-1122ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 Information extraction from clinical notes: are we ready to switch to large language models?
abstract
OBJECTIVES: To assess the performance, generalizability, and computational efficiency of instruction-tuned Large Language Model Meta AI (LLaMA)-2 and LLaMA-3 models compared to bidirectional encoder representations from transformers (BERT) for clinical information extraction (IE) tasks, specifically named entity recognition (NER) and relation extraction (RE). MATERIALS AND METHODS: We developed a comprehensive annotated corpus of 1588 clinical notes from 4 data sources-UT Physicians (UTP) (1342 notes), Transcribed Medical Transcription Sample Reports and Examples (MTSamples) (146), Medical Information Mart for Intensive Care (MIMIC)-III (50), and Informatics for Integrating Biology and the Bedside (i2b2) (50), capturing 4 clinical entities (problems, tests, medications, other treatments) and 16 modifiers (eg, negation, certainty). Large Language Model Meta AI-2 and LLaMA-3 were instruction-tuned for clinical NER and RE, and their performance was benchmarked against BERT. RESULTS: Large Language Model Meta AI models consistently outperformed BERT across datasets. In data-rich settings (eg, UTP), LLaMA achieved marginal gains (approximately 1% improvement for NER and 1.5%-3.7% for RE). Under limited data conditions (eg, MTSamples, MIMIC-III) and on the unseen i2b2 dataset, LLaMA-3-70B improved F1 scores by over 7% for NER and 4% for RE. However, performance gains came with increased computational costs, with LLaMA models requiring more memory and Graphics Processing Unit (GPU) hours and running up to 28 times slower than BERT. DISCUSSION: While LLaMA models offer enhanced performance, their higher computational demands and slower throughput highlight the need to balance performance with practical resource constraints. Application-specific considerations are essential when choosing between LLMs and BERT for clinical IE. CONCLUSION: Instruction-tuned LLaMA models show promise for clinical NER and RE tasks. However, the tradeoff between improved performance and increased computational cost must be carefully evaluated. We release our Kiwi package (https://kiwi.clinicalnlp.org/) to facilitate the application of both LLaMA and BERT models in clinical IE applications.
Xu Zuo, Yujia Zhou 0003, Xueqing Peng, Jimin Huang, Vipina Kuttichi Keloth, Vincent J. Zhang, Ruey-Ling Weng, Cathy Shyr, Qingyu Chen 0001, Xiaoqian Jiang, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.6
2025 Developing and sustaining inclusive language in biomedical informatics communications: an AMIA Board of Directors endorsed paper on the Inclusive Language and Context Style Guidelines
abstract
OBJECTIVES: In 2023, AMIA's Inclusive Language and Context Style Guidelines (the "Guidelines") were approved by the Board of Directors and made a publicly available resource. This work began in 2021 through AMIA's DEI Task Force and subsequent DEI Committee; many members provided input, feedback, and time to create the Guidelines. In this paper, the authors provide a transparent account of the origin, development, contents, and dissemination of the Guidelines and share plans for their future development and use. MATERIALS AND METHODS: Our approach to drafting, refining, and distributing the Guidelines included consulting existing language guides, AMIA member reviews, external expert reviews, webinars, and workshops. Through an iterative approach to drafting and refining the Guidelines, the authors consulted relevant language guidelines and many experts throughout and beyond the AMIA community. RESULTS: The Inclusive Language Context Guidelines were formally approved by the AMIA Board of Directors on February 15, 2023. The Guidelines included four principles to be considered in scientific communications: Plurality, Precision, Transparency, and Destigmatization. DISCUSSION: A moment of vulnerability where an AMIA member raised concerns about the use of harmful language during a presentation resulted in the creation of a principled approach to support inclusive language within biomedical and health informatics communications. We envision that the Guidelines will support health equity by challenging dominant public narratives around health, fostering stronger interdisciplinary collaboration and critical thinking about the impact of language, and creating a more welcoming environment for the broader AMIA community. This work could not have been completed without the support of many AMIA members and other researchers in biomedical and health informatics. The Guidelines are a living document that will continue to be updated with input and feedback from the AMIA community into the future.
Oliver J. Bear Don't Walk IV, Shefali Haldar, Duo Helen Wei, Hu Huang 0004, Rebecca L. Rivera, Jungwei Fan 0001, Vipina Kuttichi Keloth, Tiffany I. Leung, Pooja M. Desai, Diane M. Korngiebel, Lisa Grossman Liu, Adrienne Pichon, Vignesh Subbian, Tony Solomonides, Laura K. Wiley, Omolola Ogunyemi, Gretchen Purcell Jackson, Irene Dankwa-Mullan, Lisa Dirks, Avery Rose Everhart, Andrea G. Parker, Bradley E. Iott, Clair A. Kronk, Randi E. Foraker, Krista G. Martin, Tara Anand, Salvatore G. Volpe, Nathan Yung, Rubina F. Rizvi, Robert James Lucero, Tiffani J. Bright
J. Am. Medical Informatics Assoc.7
2025 Ontology enrichment using a large language model: Applying lexical, semantic, and knowledge network-based similarity for concept placement
abstract
OBJECTIVE: Ontologies are essential for representing the knowledge of a domain. To make ontologies useful, they must encompass a comprehensive domain view. To achieve ontology enrichment, there is a need to discover new concepts to be added, either because they were missed in the first place, or the state-of-the-art has advanced to develop new real-world concepts. Our goal is to develop an automatic enrichment pipeline using a seed ontology, a Large Language Model (LLM), and source of text. The pipeline is applied to the domain of Social Determinants of Health (SDoH), using PubMed as a source of concepts. In this work, the applicability and effectiveness of the enrichment pipeline is demonstrated by extending the SDoH Ontology called SOHOv1, however our methodology could be used in other domains as well. METHODS: We first retrieved PubMed abstracts of candidate articles with existing SOHOv1 concepts as search terms. Next, we used GPT-4-1201 to extract semantic triples from the abstracts. We identified concepts from these triples utilizing lexical, semantic, and knowledge network-based filtering. We also compared the granularity of semantic triples extracted with our method to the triples in the SemMedDB (Semantic MEDLINE Database). The results were evaluated by human experts and standard ontology tools for checking consistency and semantic correctness. RESULTS: We expanded SOHOv1, which contained 173 concepts and 585 axioms, including 207 logical axioms to SOHOv2, which contains 572 concepts, 1,542 axioms, including 725 logical axioms. Our methods identified more concepts than those extracted from SemMedDB for the same task. While we have shown the feasibility of our approach for an SDoH ontology, the methodology is generalizable to other ontologies with an existing seed ontology and text corpus. CONCLUSIONS: The contributions of this work are: Extracting semantic triples from PubMed abstracts using GPT-4-1201 utilizing prompt chaining; showing the superiority of triples from GPT-4-1201 over triples from SemMedDB for SDoH; using lexical and semantic similarity search techniques with knowledge network-based search to identify the concepts to be added to the ontology; confirming the quality of the new concepts with human experts.
Navya Martin Kollapally, James Geller, Vipina Kuttichi Keloth, Zhe He 0001, Julia Xu
J. Biomed. Informatics3
2024 Using clinical entity recognition for curating an interface terminology to aid fast skimming of EHRs
abstract
Highlighting of Electronic Health Records (EHRs) involves marking essential content of EHR notes, corresponding to concepts of a clinical terminology. However, employing the best clinical terminology (SNOMED CT) for highlighting EHRs, captures only a portion of their crucial content. In this paper, we describe the curation of a Cardiology Interface Terminology (CIT) dedicated to the application of highlighting EHRs of cardiology patients. We utilize a Clinical-Named Entity Recognition (Clinical NER) approach for extracting phrases, of higher granularity than SNOMED CT concepts, from EHRs, for enriching CIT. For this purpose, we train a neural network model with BIOE-tagged (Beginning, Inside, End, and Outside) cardiology entities. Transfer Learning can be used to facilitate the curation of an interface terminology for highlighting EHRs for other specialties e.g. Nephrology. Large-scale highlighting enables overworked physicians and other healthcare providers to fast skim the dense volume of EHRs they regularly read. Secondary research and EHRs interoperability are other applications that can be supported by highlighting.
Navya Martin Kollapally, Mahshad Koohi Habibi Dehkordi, Yehoshua Perl, James Geller, Fadi P. Deek, Hao Liu 0025, Vipina Kuttichi Keloth, Gai Elhanan, Andrew J. Einstein, Shuxin Zhou
BIBM7
2024 Advancing entity recognition in biomedicine via instruction tuning of large language models
abstract
MOTIVATION: Large Language Models (LLMs) have the potential to revolutionize the field of Natural Language Processing, excelling not only in text generation and reasoning tasks but also in their ability for zero/few-shot learning, swiftly adapting to new tasks with minimal fine-tuning. LLMs have also demonstrated great promise in biomedical and healthcare applications. However, when it comes to Named Entity Recognition (NER), particularly within the biomedical domain, LLMs fall short of the effectiveness exhibited by fine-tuned domain-specific models. One key reason is that NER is typically conceptualized as a sequence labeling task, whereas LLMs are optimized for text generation and reasoning tasks. RESULTS: We developed an instruction-based learning paradigm that transforms biomedical NER from a sequence labeling task into a generation task. This paradigm is end-to-end and streamlines the training and evaluation process by automatically repurposing pre-existing biomedical NER datasets. We further developed BioNER-LLaMA using the proposed paradigm with LLaMA-7B as the foundational LLM. We conducted extensive testing on BioNER-LLaMA across three widely recognized biomedical NER datasets, consisting of entities related to diseases, chemicals, and genes. The results revealed that BioNER-LLaMA consistently achieved higher F1-scores ranging from 5% to 30% compared to the few-shot learning capabilities of GPT-4 on datasets with different biomedical entities. We show that a general-domain LLM can match the performance of rigorously fine-tuned PubMedBERT models and PMC-LLaMA, biomedical-specific language model. Our findings underscore the potential of our proposed paradigm in developing general-domain LLMs that can rival SOTA performances in multi-task, multi-domain scenarios in biomedical and health applications. AVAILABILITY AND IMPLEMENTATION: Datasets and other resources are available at https://github.com/BIDS-Xu-Lab/BioNER-LLaMA.
Vipina Kuttichi Keloth, Qianqian Xie, Xueqing Peng, Yan Wang 0015, Andrew Zheng, Melih Selek, Kalpana Raja, Chih-Hsuan Wei, Qiao Jin 0001, Zhiyong Lu, Qingyu Chen 0001, Hua Xu 0001
Bioinform.1
2024 Improving large language models for clinical named entity recognition via prompt engineering
abstract
IMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications.
Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2024 Ensemble pretrained language models to extract biomedical knowledge from literature
abstract
OBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research.
Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001
J. Am. Medical Informatics Assoc.9
2024 FedFSA: Hybrid and federated framework for functional status ascertainment across institutions
Sunyang Fu, Heling Jia, Maria Vassilaki, Vipina Kuttichi Keloth, Yifang Dang, Yujia Zhou 0003, Muskan Garg, Ronald C. Petersen, Jennifer L. St. Sauver, Sungrim Moon, Liwei Wang 0010, Andrew Wen, Fang Li 0011, Hua Xu 0001, Cui Tao, Jungwei Fan 0001, Sunghwan Sohn
J. Biomed. Informatics4
2023 Using annotation for computerized support for fast skimming of cardiology electronic health record notes
abstract
Under the circumstances prevalent in current healthcare, medical professionals such as physicians and nurses need to read large numbers of Electronic Health Record (EHR) notes. This demand fosters a situation in which providers typically do not read a whole note but quickly skim it, to capture its essential content. Consequently, by fast skimming, one may miss a critical medical fact. Annotation highlights important content in EHR notes, and enables healthcare professionals to perform fast skimming, thereby minimizing the risk of missing critical information, which is detrimental to patient care. We designed the Cardiology Interface Terminology (CIT) for the purpose of annotation of cardiology EHRs. We emphasize that by annotation we refer to highlighting important information in the EHR notes for enabling fast skimming rather than just recognizing names of diseases, drugs, etc. from Reference Terminologies as is usually done by existing Named Entity Recognition (NER) systems. The CIT design starts with the cardiology components of SNOMED CT. It is enhanced by mining phrases from cardiology EHRs, as potential CIT concepts, which are of higher granularity than SNOMED concepts. Machine learning (ML), the state of the art technique for mining concepts from EHRs, requires training data. However, there is no training data for designing CIT. In the first stage, we introduce an innovative semi-automatic method for mining concepts from EHRs, to replace costly manual mining. The only manual portion is the review of the automatically mined phrases, before their insertion as CIT concepts. The effectiveness of annotation of cardiology EHRs with CIT was evaluated utilizing proper metrics, and compared to annotation with SNOMED CT. In a future second stage, ML mining techniques will be used for enhancing CIT with extra concepts from EHRs, utilizing the concepts added in the first stage as training data. This work focuses on a novel semi-automated method to design the Cardiology Interface Terminology (CIT) for annotation of cardiology EHRs to support fast skimming of EHR notes. Similar interface terminologies for other medical specialties could be obtained from CIT using Transfer Learning.
Mahshad Koohi Habibi Dehkordi, Andrew J. Einstein, Shuxin Zhou, Gai Elhanan, Yehoshua Perl, Vipina Kuttichi Keloth, James Geller, Hao Liu 0025
BIBM6
2023 Towards precise PICO extraction from abstracts of randomized controlled trials using a section-specific learning approach
abstract
MOTIVATION: Automated extraction of participants, intervention, comparison/control, and outcome (PICO) from the randomized controlled trial (RCT) abstracts is important for evidence synthesis. Previous studies have demonstrated the feasibility of applying natural language processing (NLP) for PICO extraction. However, the performance is not optimal due to the complexity of PICO information in RCT abstracts and the challenges involved in their annotation. RESULTS: We propose a two-step NLP pipeline to extract PICO elements from RCT abstracts: (i) sentence classification using a prompt-based learning model and (ii) PICO extraction using a named entity recognition (NER) model. First, the sentences in abstracts were categorized into four sections namely background, methods, results, and conclusions. Next, the NER model was applied to extract the PICO elements from the sentences within the title and methods sections that include >96% of PICO information. We evaluated our proposed NLP pipeline on three datasets, the EBM-NLPmoddataset, a randomly selected and reannotated dataset of 500 RCT abstracts from the EBM-NLP corpus, a dataset of 150 COVID-19 RCT abstracts, and a dataset of 150 Alzheimer's disease (AD) RCT abstracts. The end-to-end evaluation reveals that our proposed approach achieved an overall micro F1 score of 0.833 on the EBM-NLPmod dataset, 0.928 on the COVID-19 dataset, and 0.899 on the AD dataset when measured at the token-level and an overall micro F1 score of 0.712 on EBM-NLPmod dataset, 0.850 on the COVID-19 dataset, and 0.805 on the AD dataset when measured at the entity-level. AVAILABILITY: Our codes and datasets are publicly available at https://github.com/BIDS-Xu-Lab/section_specific_annotation_of_PICO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vipina Kuttichi Keloth, Kalpana Raja, Yong Chen 0016, Hua Xu 0001
Bioinform.2
2023 Systematic design and data-driven evaluation of social determinants of health ontology (SDoHO)
abstract
OBJECTIVE: Social determinants of health (SDoH) play critical roles in health outcomes and well-being. Understanding the interplay of SDoH and health outcomes is critical to reducing healthcare inequalities and transforming a "sick care" system into a "health-promoting" system. To address the SDOH terminology gap and better embed relevant elements in advanced biomedical informatics, we propose an SDoH ontology (SDoHO), which represents fundamental SDoH factors and their relationships in a standardized and measurable way. MATERIAL AND METHODS: Drawing on the content of existing ontologies relevant to certain aspects of SDoH, we used a top-down approach to formally model classes, relationships, and constraints based on multiple SDoH-related resources. Expert review and coverage evaluation, using a bottom-up approach employing clinical notes data and a national survey, were performed. RESULTS: We constructed the SDoHO with 708 classes, 106 object properties, and 20 data properties, with 1,561 logical axioms and 976 declaration axioms in the current version. Three experts achieved 0.967 agreement in the semantic evaluation of the ontology. A comparison between the coverage of the ontology and SDOH concepts in 2 sets of clinical notes and a national survey instrument also showed satisfactory results. DISCUSSION: SDoHO could potentially play an essential role in providing a foundation for a comprehensive understanding of the associations between SDoH and health outcomes and paving the way for health equity across populations. CONCLUSION: SDoHO has well-designed hierarchies, practical objective properties, and versatile functionalities, and the comprehensive semantic and coverage evaluation achieved promising performance compared to the existing ontologies relevant to SDoH.
Yifang Dang, Fang Li 0011, Xinyue Hu 0002, Vipina Kuttichi Keloth, Sunyang Fu, Muhammad Amith, J. Wilfred Fan, Jingcheng Du, Evan Yu, Xiaoqian Jiang, Hua Xu 0001, Cui Tao
J. Am. Medical Informatics Assoc.4
2023 Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001
J. Biomed. Informatics1
2022 Section-specific Annotation of PICO for Medical Evidence Extraction: Applications to Articles of Randomized Controlled Trials for Alzheimer's Disease and COVID-19
Vipina Kuttichi Keloth, Hua Xu 0001
AMIA2
2021 Visual comprehension and orientation into the COVID-19 CIDO ontology
Yehoshua Perl, Yongqun He, Christopher Ochs, James Geller, Hao Liu 0025, Vipina Kuttichi Keloth
J. Biomed. Informatics7
2020 Generating Training Data for Concept-Mining for an 'Interface Terminology' Annotating Cardiology EHRs
abstract
Clinical data stored in EHRs could provide valuable knowledge for research if it were annotated properly. However, almost no EHR notes are currently annotated as the performance of off the shelf annotation tools is unsatisfactory. Concentrating on the cardiology specialty, we propose to design a Cardiology Interface Terminology dedicated to the annotation of EHR notes in cardiology. This interface terminology will be developed by the addition of high granularity concepts, mined from cardiology EHR notes, to an initial version reusing SNOMED CT cardiology subhierarchies. Using text mining NLP tools with machine learning for extending this interface terminology requires proper training data. In this paper, we discuss concept-mining of EHR notes, using concatenation and anchoring operations iteratively to create such training data. This approach can be applied to other medical specialties.
Vipina Kuttichi Keloth, Shuxin Zhou, Andrew J. Einstein, Gai Elhanan, Yan Chen 0009, James Geller, Yehoshua Perl
BIBM1
2020 Mining Concepts for a COVID Interface Terminology for Annotation of EHRs
abstract
The COVID-19 pandemic has overwhelmed the healthcare services of many countries with increased number of patients and also with a deluge of medical data. Furthermore, the emergence and global spread of new infectious diseases are highly likely to continue in the future. Incomplete data about presentations, signs, and symptoms of COVID-19 has had adverse effects on healthcare delivery. The EHRs of US hospitals have ingested huge volumes of relevant, up-to-date data about patients, but the lack of a proper system to annotate this data has greatly reduced its usefulness. We propose to design a COVID interface terminology for the annotation of EHR notes of COVID-19 patients. The initial version of this interface terminology was created by integrating COVID concepts from existing ontologies. Further enrichment of the interface terminology is performed by mining high granularity concepts from EHRs, because such concepts are usually not present in the existing reference terminologies. We use the techniques of concatenation and anchoring iteratively to extract high granularity phrases from the clinical text. In addition to increasing the conceptual base of the COVID interface terminology, this will also help in generating training data for large scale concept mining using machine learning techniques. Having the annotated clinical notes of COVID-19 patients available will help in speeding up research in this field.
Vipina Kuttichi Keloth, Shuxin Zhou, Luke Lindemann, Gai Elhanan, Andrew J. Einstein, James Geller, Yehoshua Perl
IEEE BigData1
2020 A review of auditing techniques for the Unified Medical Language System
abstract
OBJECTIVE: The study sought to describe the literature related to the development of methods for auditing the Unified Medical Language System (UMLS), with particular attention to identifying errors and inconsistencies of attributes of the concepts in the UMLS Metathesaurus. MATERIALS AND METHODS: We applied the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) approach by searching the MEDLINE database and Google Scholar for studies referencing the UMLS and any of several terms related to auditing, error detection, and quality assurance. A qualitative analysis and summarization of articles that met inclusion criteria were performed. RESULTS: Eighty-three studies were reviewed in detail. We first categorized techniques based on various aspects including concepts, concept names, and synonymy (n = 37), semantic type assignments (n = 36), hierarchical relationships (n = 24), lateral relationships (n = 12), ontology enrichment (n = 8), and ontology alignment (n = 18). We also categorized the methods according to their level of automation (ie, automated systematic, automated heuristic, or manual) and the type of knowledge used (ie, intrinsic or extrinsic knowledge). CONCLUSIONS: This study is a comprehensive review of the published methods for auditing the various conceptual aspects of the UMLS. Categorizing the auditing techniques according to the various aspects will enable the curators of the UMLS as well as researchers comprehensive easy access to this wealth of knowledge (eg, for auditing lateral relationships in the UMLS). We also reviewed ontology enrichment and alignment techniques due to their critical use of and impact on the UMLS.
Zhe He 0001, Duo Helen Wei, Vipina Kuttichi Keloth, Jungwei Fan 0001, Luke Lindemann, James J. Cimino, Yehoshua Perl
J. Am. Medical Informatics Assoc.4
2019 Measuring and Avoiding Information Loss During Concept Import from a Source to a Target Ontology
abstract
Comparing pairs of ontologies in the same biomedical content domain often uncovers surprising differences. In many cases these differences can be characterized as “density differences,” where one ontology describes the content domain with more concepts in a more detailed manner. Using the Unified Medical Language System across pairs of ontologies contained in it, these differences can be precisely observed and used as the basis for importing concepts from the ontology of higher density into the ontology of lower density. However, such an import can lead to an intuitive loss of information that is hard to formalize. This paper proposes an approach based on information theory that mathematically distinguishes between different methods of concept import and measures the associated avoidance of information loss.
James Geller, Shmuel Tomi Klein, Vipina Kuttichi Keloth
KEOD3
2019 Alternative classification of identical concepts in different terminologies: Different ways to view the world
Vipina Kuttichi Keloth, Zhe He 0001, Gai Elhanan, James Geller
J. Biomed. Informatics1
2018 How Sustainable are Biomedical Ontologies?
James Geller, Vipina Kuttichi Keloth, Mark A. Musen
AMIA2
2018 Leveraging Horizontal Density Differences between Ontologies to Identify Missing Child Concepts: A Proof of Concept
Vipina Kuttichi Keloth, Zhe He 0001, Yan Chen 0009, James Geller
AMIA1
2018 Extended Analysis of Topological-Pattern-Based Ontology Enrichment
Zhe He 0001, Vipina Kuttichi Keloth, Yan Chen 0009, James Geller
BIBM2