EDBT 2026 Demo / reviewers in the wild / expert
Girish N. Nadkarni
dblp:161/1895
· DBLP profile ↗
13ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-6319-4314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sociodemographic bias in large language model clinical trial screeningabstractBackground: Large language models (LLMs) are increasingly used in randomized clinical trial (RCT) screening, but their potential for sociodemographic bias remains unclear. Objective: To determine whether LLM-based trial screening judgments vary with patient sociodemographic characteristics when clinical details and eligibility criteria are held constant. Design Setting and Participants: Cross-sectional evaluation of Phase II-III RCT protocols from ClinicalTrials.gov (U.S. adult populations; 2023-2024). For each protocol, we created 15 physician-validated clinical vignettes rendered in 34 versions: one control (no identifiers) and 33 identity variants spanning gender, race/ethnicity, socioeconomic status, homelessness, unemployment, and sexual orientation. Exposures: Identity labels applied to otherwise identical vignettes, evaluated by nine contemporary LLMs. Main Outcomes and Measures: Primary: eligibility domain score (1-5 Likert scale) comparing identity variants versus control. Secondary: adherence, resources, risk-benefit, and trust/attitude domains. Mixed-effects models estimated adjusted mean differences with multiplicity-corrected P values; differences <.10 considered trivial. Results: Of 69 protocols, 58 met inclusion criteria. Analysis of 5,324,400 model evaluations showed eligibility judgments were largely stable: most identity-related differences fell within ±0.05 (transgender woman -.008 [95% CI -.04 to .02]; White male .036 [.01 to .07]). Only homelessness exceeded the trivial threshold (-.121 [-.15 to -.09], P<.001). Secondary domains revealed socioeconomic gradients, particularly for adherence (homeless -.595, P<.001) and resources (homeless -.715, P<.001), with smaller trust/attitude effects and negligible risk-benefit differences. Conclusions and Relevance: Bias in LLM-assisted trial screening is conditional. Within fixed criteria, models reason consistently; outside them, they echo the inequities of their data. Responsible deployment in clinical research depends on preserving that boundary so that automation strengthens fairness in trial access rather than inheriting distortion. Shelly Soffer, Mahmud Omar, Orly Efros, Donald Apakama, Aya Mudrik, Robert Freeman, Girish N. Nadkarni, Eyal Klang |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Extracting social support and social isolation information from clinical psychiatry notes: comparing a rule-based natural language processing system and a large language modelabstractOBJECTIVES: Social support (SS) and social isolation (SI) are social determinants of health (SDOH) associated with psychiatric outcomes. In electronic health records (EHRs), individual-level SS/SI is typically documented in narrative clinical notes rather than as structured coded data. Natural language processing (NLP) algorithms can automate the otherwise labor-intensive process of extraction of such information. MATERIALS AND METHODS: Psychiatric encounter notes from Mount Sinai Health System (MSHS, n = 300) and Weill Cornell Medicine (WCM, n = 225) were annotated to create a gold-standard corpus. A rule-based system (RBS) involving lexicons and a large language model (LLM) using FLAN-T5-XL were developed to identify mentions of SS and SI and their subcategories (eg, social network, instrumental support, and loneliness). RESULTS: For extracting SS/SI, the RBS obtained higher macroaveraged F1-scores than the LLM at both MSHS (0.89 versus 0.65) and WCM (0.85 versus 0.82). For extracting the subcategories, the RBS also outperformed the LLM at both MSHS (0.90 versus 0.62) and WCM (0.82 versus 0.81). DISCUSSION AND CONCLUSION: Unexpectedly, the RBS outperformed the LLMs across all metrics. An intensive review demonstrates that this finding is due to the divergent approach taken by the RBS and LLM. The RBS was designed and refined to follow the same specific rules as the gold-standard annotations. Conversely, the LLM was more inclusive with categorization and conformed to common English-language understanding. Both approaches offer advantages, although additional replication studies are warranted. Braja Gopal Patra, Lauren A. Lepow, Praneet Kasi Reddy Jagadeesh Kumar, Veer Vekaria, Mohit Manoj Sharma, Prakash Adekkanattu, Brian Fennessy, Gavin Hynes, Isotta Landi, Jorge A. Sanchez-Ruiz, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Ardesheer Talati, Myrna Weissman, Mark Olfson, J. John Mann, Yiye Zhang, Alexander Charney, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 13 |
| 2025 | A scalable framework for benchmark embedding models in semantic health-care tasksabstractOBJECTIVES: Text embeddings are promising for semantic tasks, such as retrieval augmented generation (RAG). However, their application in health care is underexplored due to a lack of benchmarking methods. We introduce a scalable benchmarking method to test embeddings for health-care semantic tasks. MATERIALS AND METHODS: We evaluated 39 embedding models across 7 medical semantic similarity tasks using diverse datasets. These datasets comprised real-world patient data (from the Mount Sinai Health System and MIMIC IV), biomedical texts from PubMed, and synthetic data generated with Llama-3-70b. We first assessed semantic textual similarity (STS) by correlating the model-generated similarity scores with noise levels using Spearman rank correlation. We then reframed the same tasks as retrieval problems, evaluated by mean reciprocal rank and recall at k. RESULTS: In total, evaluating 2000 text pairs per 7 tasks for STS and retrieval yielded 3.28 million model assessments. Larger models (>7b parameters), such as those based on Mistral-7b and Gemma-2-9b, consistently performed well, especially in long-context tasks. The NV-Embed-v1 model (7b parameters), although top in short tasks, underperformed in long tasks. For short tasks, smaller models such as b1ade-embed (335M parameters) performed on-par to the larger models. For long retrieval tasks, the larger models significantly outperformed the smaller ones. DISCUSSION: The proposed benchmarking framework demonstrates scalability and flexibility, offering a structured approach to guide the selection of embedding models for a wide range of health-care tasks. CONCLUSION: By matching the appropriate model with the task, the framework enables more effective deployment of embedding models, enhancing critical applications such as semantic search and retrieval-augmented generation (RAG). Shelly Soffer, Mahmud Omar, Moran Gendler, Benjamin S. Glicksberg, Patricia H. Kovatch, Orly Efros, Robert Freeman, Alexander Charney, Girish N. Nadkarni, Eyal Klang |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | A novel method leveraging time series data to improve subphenotyping and application in critically ill patients with COVID-19
Wonsuk Oh, Pushkala Jayaraman, Pranai Tandon, Udit S. Chaddha, Patricia H. Kovatch, Alexander Charney, Benjamin S. Glicksberg, Girish N. Nadkarni |
Artif. Intell. Medicine | 8 |
| 2024 | Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency roomabstractBACKGROUND: Artificial intelligence (AI) and large language models (LLMs) can play a critical role in emergency room operations by augmenting decision-making about patient admission. However, there are no studies for LLMs using real-world data and scenarios, in comparison to and being informed by traditional supervised machine learning (ML) models. We evaluated the performance of GPT-4 for predicting patient admissions from emergency department (ED) visits. We compared performance to traditional ML models both naively and when informed by few-shot examples and/or numerical probabilities. METHODS: We conducted a retrospective study using electronic health records across 7 NYC hospitals. We trained Bio-Clinical-BERT and XGBoost (XGB) models on unstructured and structured data, respectively, and created an ensemble model reflecting ML performance. We then assessed GPT-4 capabilities in many scenarios: through Zero-shot, Few-shot with and without retrieval-augmented generation (RAG), and with and without ML numerical probabilities. RESULTS: The Ensemble ML model achieved an area under the receiver operating characteristic curve (AUC) of 0.88, an area under the precision-recall curve (AUPRC) of 0.72 and an accuracy of 82.9%. The naïve GPT-4's performance (0.79 AUC, 0.48 AUPRC, and 77.5% accuracy) showed substantial improvement when given limited, relevant data to learn from (ie, RAG) and underlying ML probabilities (0.87 AUC, 0.71 AUPRC, and 83.1% accuracy). Interestingly, RAG alone boosted performance to near peak levels (0.82 AUC, 0.56 AUPRC, and 81.3% accuracy). CONCLUSIONS: The naïve LLM had limited performance but showed significant improvement in predicting ED admissions when supplemented with real-world examples to learn from, particularly through RAG, and/or numerical probabilities from traditional ML models. Its peak performance, although slightly lower than the pure ML model, is noteworthy given its potential for providing reasoning behind predictions. Further refinement of LLMs with real-world data is necessary for successful integration as decision-support tools in care settings. Benjamin S. Glicksberg, Prem Timsina, Ashwin Sawant, Akhil Vaid, Ganesh Raut, Alexander Charney, Donald Apakama, Brendan G. Carr, Robert Freeman, Girish N. Nadkarni, Eyal Klang |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Local large language models for privacy-preserving accelerated review of historic echocardiogram reportsabstractOBJECTIVES: The study developed framework that leverages an open-source Large Language Model (LLM) to enable clinicians to ask plain-language questions about a patient's entire echocardiogram report history. This approach is intended to streamline the extraction of clinical insights from multiple echocardiogram reports, particularly in patients with complex cardiac diseases, thereby enhancing both patient care and research efficiency. MATERIALS AND METHODS: Data from over 10 years were collected, comprising echocardiogram reports from patients with more than 10 echocardiograms on file at the Mount Sinai Health System. These reports were converted into a single document per patient for analysis, broken down into snippets and relevant snippets were retrieved using text similarity measures. The LLaMA-2 70B model was employed for analyzing the text using a specially crafted prompt. The model's performance was evaluated against ground-truth answers created by faculty cardiologists. RESULTS: The study analyzed 432 reports from 37 patients for a total of 100 question-answer pairs. The LLM correctly answered 90% questions, with accuracies of 83% for temporality, 93% for severity assessment, 84% for intervention identification, and 100% for diagnosis retrieval. Errors mainly stemmed from the LLM's inherent limitations, such as misinterpreting numbers or hallucinations. CONCLUSION: The study demonstrates the feasibility and effectiveness of using a local, open-source LLM for querying and interpreting echocardiogram report data. This approach offers a significant improvement over traditional keyword-based searches, enabling more contextually relevant and semantically accurate responses; in turn showing promise in enhancing clinical decision-making and research by facilitating more efficient access to complex patient data. Akhil Vaid, Son Q. Duong, Joshua Lampert, Patricia H. Kovatch, Robert Freeman, Edgar Argulian, Lori Croft, Stamatios Lerakis, Martin Goldman, Rohan Khera, Girish N. Nadkarni |
J. Am. Medical Informatics Assoc. | 11 |
| 2021 | Extracting Social Isolation Information From Psychiatric Notes in the Electronic Health Records
Lauren A. Lepow, Braja Gopal Patra, Isotta Landi, Prakash Adekkanattu, Jyotishman Pathak, Mark Olfson, J. John Mann, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Priya Wickramaratne, Myrna Weissman, Benjamin S. Glicksberg, Alexander Charney |
AMIA | 10 |
| 2021 | Relational Learning Improves Prediction of Mortality in COVID-19 in the Intensive Care UnitabstractTraditional Machine Learning (ML) models have had limited success in predicting Coronoavirus-19 (COVID-19) outcomes using Electronic Health Record (EHR) data partially due to not effectively capturing the inter-connectivity patterns between various data modalities. In this work, we propose a novel framework that utilizes relational learning based on a heterogeneous graph model (HGM) for predicting mortality at different time windows in COVID-19 patients within the intensive care unit (ICU). We utilize the EHRs of one of the largest and most diverse patient populations across five hospitals in major health system in New York City. In our model, we use an LSTM for processing time varying patient data and apply our proposed relational learning strategy in the final output layer along with other static features. Here, we replace the traditional softmax layer with a Skip-Gram relational learning strategy to compare the similarity between a patient and outcome embedding representation. We demonstrate that the construction of a HGM can robustly learn the patterns classifying patient representations of outcomes through leveraging patterns within the embeddings of similar patients. Our experimental results show that our relational learning-based HGM model achieves higher area under the receiver operating characteristic curve (auROC) than both comparator models in all prediction time windows, with dramatic improvements to recall. Tingyi Wanyan, Akhil Vaid, Jessica K. De Freitas, Sulaiman Somani, Riccardo Miotto, Girish N. Nadkarni, Ariful Azad, Ying Ding 0001, Benjamin S. Glicksberg |
IEEE Trans. Big Data | 6 |
| 2020 | Heterogeneous Graph Embeddings of Electronic Health Records Improve Critical Care Disease Predictions
Tingyi Wanyan, Martin Kang, Marcus A. Badgeley, Kipp W. Johnson, Jessica K. De Freitas, Fayzan F. Chaudhry, Akhil Vaid, Riccardo Miotto, Girish N. Nadkarni, Fei Wang 0001, Justin F. Rousseau, Ariful Azad, Ying Ding 0001, Benjamin S. Glicksberg |
AIME | 10 |
| 2015 | A Genome- and Phenome- Wide Study of Diverticulosis
Yoonjung Y. Joo, Jennifer A. Pacheco, Loren L. Armstrong, William K. Thompson, Robert J. Carroll, Joshua C. Denny, Peggy L. Peissig, James G. Linneman, Jyotishman Pathak, Girish N. Nadkarni, Laura Rasmussen-Torvik, M. Geoffrey Hayes, Abel N. Kho |
AMIA | 10 |
| 2015 | Incorporating temporal EHR data in predictive models for risk stratification of renal function deterioration
Anima Singh, Girish N. Nadkarni, Omri Gottesman, Stephen B. Ellis, Erwin P. Bottinger, John V. Guttag |
J. Biomed. Informatics | 2 |
| 2014 | Disease progression subtype discovery from longitudinal EMR data with a majority of missing values and unknown initial time points
Ilkka Huopaniemi, Girish N. Nadkarni, Rajiv Nadukuru, Vaneet Lotay, Stephen B. Ellis, Omri Gottesman, Erwin P. Bottinger |
AMIA | 2 |
| 2014 | Development and validation of an electronic phenotyping algorithm for chronic kidney disease
Girish N. Nadkarni, Omri Gottesman, James G. Linneman, Herbert S. Chase, Richard L. Berg, Samira Farouk, Vaneet Lotay, Stephen B. Ellis, George Hripcsak, Peggy L. Peissig, Chunhua Weng, Rajiv Nadukuru, Erwin P. Bottinger |
AMIA | 1 |