Navya Martin Kollapally

dblp:306/6731 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-4004-6508ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Benchmarking Automatic Speech Recognition for Aphasia: A Clinical Evaluation Framework
abstract
AI tools for speech therapy represent more than innovation; they signify necessity. With 89% of clinicians facing overwhelming caseloads and therapy wait times averaging 3-6 months, the demand for scalable rehabilitation support has never been higher. Although AI-driven communication platforms increasingly provide real-time feedback and sustained engagement, current Automatic Speech Recognition (ASR) systems still face significant limitations in accurately processing disordered speech such as aphasia. Aphasia is an acquired language disorder, the speech of individuals with aphasia is characterized by unintelligible words, jargon, or non-words. To add to this the speaker with aphasia may not recognize their errors and mostly have difficulty in comprehension. These limitations hinder equitable access to AI-driven rehabilitation tools. To bridge this gap, the first contribution of this study is evaluating four state-of-the-art ASR models, such as Whisper, NeMo-Conformer, Wav2Vec 2.0, and SpeechBrain, through the lens of Speech-Language Pathologist (SLP). The second contribution of the paper is utilizing a comprehensive benchmarking framework to assess how effectively these models capture clinically relevant aspects of aphasic speech, including lexical, syntactic, and fluency-related features. For evaluating the transcribed text, a combination of quantitative, qualitative (human-expert based), and linguistically grounded evaluation metrics is used, such as verb error rate, noun error rate, mean dependency length etc. of transcribed text.
Navya Martin Kollapally, Christa Akes
BIBM1
2025 Ontology enrichment using a large language model: Applying lexical, semantic, and knowledge network-based similarity for concept placement
abstract
OBJECTIVE: Ontologies are essential for representing the knowledge of a domain. To make ontologies useful, they must encompass a comprehensive domain view. To achieve ontology enrichment, there is a need to discover new concepts to be added, either because they were missed in the first place, or the state-of-the-art has advanced to develop new real-world concepts. Our goal is to develop an automatic enrichment pipeline using a seed ontology, a Large Language Model (LLM), and source of text. The pipeline is applied to the domain of Social Determinants of Health (SDoH), using PubMed as a source of concepts. In this work, the applicability and effectiveness of the enrichment pipeline is demonstrated by extending the SDoH Ontology called SOHOv1, however our methodology could be used in other domains as well. METHODS: We first retrieved PubMed abstracts of candidate articles with existing SOHOv1 concepts as search terms. Next, we used GPT-4-1201 to extract semantic triples from the abstracts. We identified concepts from these triples utilizing lexical, semantic, and knowledge network-based filtering. We also compared the granularity of semantic triples extracted with our method to the triples in the SemMedDB (Semantic MEDLINE Database). The results were evaluated by human experts and standard ontology tools for checking consistency and semantic correctness. RESULTS: We expanded SOHOv1, which contained 173 concepts and 585 axioms, including 207 logical axioms to SOHOv2, which contains 572 concepts, 1,542 axioms, including 725 logical axioms. Our methods identified more concepts than those extracted from SemMedDB for the same task. While we have shown the feasibility of our approach for an SDoH ontology, the methodology is generalizable to other ontologies with an existing seed ontology and text corpus. CONCLUSIONS: The contributions of this work are: Extracting semantic triples from PubMed abstracts using GPT-4-1201 utilizing prompt chaining; showing the superiority of triples from GPT-4-1201 over triples from SemMedDB for SDoH; using lexical and semantic similarity search techniques with knowledge network-based search to identify the concepts to be added to the ontology; confirming the quality of the new concepts with human experts.
Navya Martin Kollapally, James Geller, Vipina Kuttichi Keloth, Zhe He 0001, Julia Xu
J. Biomed. Informatics1
2024 Using clinical entity recognition for curating an interface terminology to aid fast skimming of EHRs
abstract
Highlighting of Electronic Health Records (EHRs) involves marking essential content of EHR notes, corresponding to concepts of a clinical terminology. However, employing the best clinical terminology (SNOMED CT) for highlighting EHRs, captures only a portion of their crucial content. In this paper, we describe the curation of a Cardiology Interface Terminology (CIT) dedicated to the application of highlighting EHRs of cardiology patients. We utilize a Clinical-Named Entity Recognition (Clinical NER) approach for extracting phrases, of higher granularity than SNOMED CT concepts, from EHRs, for enriching CIT. For this purpose, we train a neural network model with BIOE-tagged (Beginning, Inside, End, and Outside) cardiology entities. Transfer Learning can be used to facilitate the curation of an interface terminology for highlighting EHRs for other specialties e.g. Nephrology. Large-scale highlighting enables overworked physicians and other healthcare providers to fast skim the dense volume of EHRs they regularly read. Secondary research and EHRs interoperability are other applications that can be supported by highlighting.
Navya Martin Kollapally, Mahshad Koohi Habibi Dehkordi, Yehoshua Perl, James Geller, Fadi P. Deek, Hao Liu 0025, Vipina Kuttichi Keloth, Gai Elhanan, Andrew J. Einstein, Shuxin Zhou
BIBM1
2022 An Ontology for the Social Determinants of Health Domain
abstract
Social Determinants of Health (SDOH) are societal factors, such as where a person was born, grew up, works, lives, etc., along with socio-economic and community factors that affect an individual’s health. SDOH are correlated with many clinical outcomes, hence it is desirable to record SDOH data in Electronic Health Records (EHRs). Besides storing images, text, etc., EHRs rely on coded terms available in standard ontologies and terminologies to record observations and analyses. There is a substantial amount of research on understanding the clinical impact of SDOH, ranging from screening tools to practice-based interventions. However, there is no comprehensive collection of terms for recording SDOH observations in EHRs. Our research goal is to develop an ontology that covers the terms describing SDOH. We present a prototype ontology called Social Determinant of Health Ontology (SOHO) that covers relevant concepts and IS-A relationships describing impacts and associations of social determinants. We describe the evaluation techniques that we applied to SOHO, including human experts’ review and algorithmic evaluation.
Navya Martin Kollapally, Yan Chen 0009, Julia Xu, James Geller
BIBM1
2021 Detecting, Reporting And Alleviating Racial Biases In Standardized Medical Terminologies And Ontologies
abstract
Recently, the issue has been raised that personal and systemic biases in organizations, such as some police departments, have also been detected in healthcare organizations. Furthermore, victims of bias incidents often end up in the healthcare system for treatment. Providers use standardized terminologies to record the status of patients in EHRs. To accurately record patient data, these terminologies must contain all the terms that a healthcare provider needs, including terms that might be race-, ethnicity-, or gender-specific. Following reports about gaps in terminologies, we investigated the coverage with respect to such terms in major terminologies such as SNOMED CT, ICD-10, CPT, NCIt and MedDRA. To identify potentially missing terms, we drew on public databases and news articles describing incidents that resulted in minority members requiring medical attention after police interventions. We posit those terms should be added into medical terminologies to improve the ability to record incidents happening inside and outside of the healthcare system.
James Geller, Navya Martin Kollapally
BIBM2
2021 Health Ontology for Minority Equity (HOME)
Navya Martin Kollapally, Yan Chen 0009, James Geller
KEOD1