VLDB 2026 Research / reviewers in the wild / expert
Meliha Yetisgen
dblp:87/3886 · also Meliha Yetisgen-Yildiz
· DBLP profile ↗
75ranked-venue papers
10as first author
34since 2021 · last 2026
0000-0001-9919-9811ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 57 · 8 first-author · 25 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandra Baicoianu, Ana Neacsu, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Lea Reinartz, Benjamin Lecouteux, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Corneliu Florea, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Hendrik Damm, Henning Schäfer, Ivan Koychev, Josiane Mothe, Liviu-Daniel Stefan, Maja J. Hjuler, Mehmet Kurt, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Ivanovici, Ming Shan Hee, Mohammad El Sakka, Momina Ahsan, Obioma Pelka, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Bahadir Eryilmaz, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Yuri Prokopchuk, Zhuohan Xie |
ECIR (4) | 29 |
| 2026 | Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM ApproachesabstractLarge language models (LLMs) have shown considerable promise in clinical natural language processing, yet few domain-specific datasets exist to rigorously evaluate their performance on radiology tasks. In this work, we introduce an annotated corpus of 6,393 radiology reports from 586 patients, each labeled for follow-up imaging status, to support the development and benchmarking of follow-up adherence detection systems. Using this corpus, we systematically compared traditional machine-learning classifiers, including logistic regression (LR), support vector machines (SVM), Longformer, and a fully fine-tuned Llama3-8B-Instruct, with recent generative LLMs. To evaluate generative LLMs, we tested GPT-4o and the open-source GPT-OSS-20B under two configurations: a baseline (Base) and a task-optimized (Advanced) setting that focused inputs on metadata, recommendation sentences, and their surrounding context. A refined prompt for GPT-OSS-20B further improved reasoning accuracy. Performance was assessed using precision, recall, and F1 scores with 95% confidence intervals estimated via non-parametric bootstrapping. Inter-annotator agreement was high (F1 = 0.846). GPT-4o (Advanced) achieved the best performance (F1 = 0.832), followed closely by GPT-OSS-20B (Advanced; F1 = 0.828). LR and SVM also performed strongly (F1 = 0.776 and 0.775), underscoring that while LLMs approach human-level agreement through prompt optimization, interpretable and resource-efficient models remain valuable baselines. Namu Park, Giridhar Kaushik Ramachandran, Kevin Lybarger, Fei Xia 0004, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
LREC | 7 |
| 2026 | MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question AnsweringabstractEvaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research. Wen-Wai Yim, Asma Ben Abacha, Zixuan Yu, Robert Doerning, Fei Xia 0004, Meliha Yetisgen |
LREC | 6 |
| 2026 | RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings
Sitong Zhou, Meliha Yetisgen, Mari Ostendorf |
LREC | 2 |
| 2026 | Automated identification of incidentalomas requiring follow-up: A multi-anatomy evaluation of LLM-based and supervised approaches
Namu Park, Farzad Ahmed, Zhaoyi Sun, Kevin Lybarger, Ethan Breinhorst, Julie Hu, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
J. Biomed. Informatics | 9 |
| 2025 | Tailoring task arithmetic to address bias in models trained on multi-institutional datasets
Xiruo Ding, Zhecheng Sheng, Brian Hur, Justin Tauscher, Dror Ben-Zeev, Meliha Yetisgen, Serguei V. S. Pakhomov, Trevor Cohen |
J. Biomed. Informatics | 6 |
| 2025 | A scoping review of natural language processing in addressing medically inaccurate information: Errors, misinformation, and hallucination
Zhaoyi Sun, Wen-Wai Yim, Özlem Uzuner, Fei Xia 0004, Meliha Yetisgen |
J. Biomed. Informatics | 5 |
| 2025 | WoundcareVQA: A multilingual visual question answering benchmark dataset for wound careabstractOBJECTIVE: Introduce the task of wound care multimodal multilingual visual question answering, provide baseline performances, and identify areas of future study. METHODS: A dataset of wound care multimodal multilingual visual question answering (VQA) was created using consumer health questions asked online. Practicing US medical doctors were tasked with providing metadata and expert responses labels. Several instruct-enabled, multilingual visual question answering models (GPT-4o, Gemini-1.5-Pro, and Qwen-VL) were tested to benchmark performances. Finally, automatic evaluations were tested against domain expert response ratings. RESULTS: A multilingual dataset of 477 wound care cases, 768 responses, 748 images, 3k structured data labels, 1362 translation instances, and 10k judgments was constructed (https://osf.io/xsj5u/). Metadata scores ranged from 0.32-0.78 accuracy depending on classification type; response generation performances 0.06 BLEU, 0.66 BERTScore, 0.45 ROUGE-L in English and 0.12 BLEU, 0.69 BERTScore, and 0.50 ROUGE-L in Chinese. CONCLUSION: We construct and explore the tasks of multimodal, multilingual VQA. We hope the work here can inspire further research in wound care metadata classification, VQA response generation, and open response automatic evaluation. Wen-Wai Yim, Asma Ben Abacha, Robert Doerning, Chia-Yu Chen, Jiaying Xu, Anita Subbarao, Zixuan Yu, Fei Xia 0004, M. Kennedy Hall, Meliha Yetisgen |
J. Biomed. Informatics | 10 |
| 2024 | Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and MethodsabstractSocial determinants of health (SDoH) play a critical role in shaping health outcomes, particularly in pediatric populations where interventions can have long-term implications. SDoH are frequently studied in the Electronic Health Record (EHR), which provides a rich repository for diverse patient data. In this work, we present a novel annotated corpus, the Pediatric Social History Annotation Corpus (PedSHAC), and evaluate the automatic extraction of detailed SDoH representations using fine-tuned and in-context learning methods with Large Language Models (LLMs). PedSHAC comprises annotated social history sections from 1,260 clinical notes obtained from pediatric patients within the University of Washington (UW) hospital system. Employing an event-based annotation scheme, PedSHAC captures ten distinct health determinants to encompass living and economic stability, prior trauma, education access, substance use history, and mental health with an overall annotator agreement of 81.9 F1. Our proposed fine-tuning LLM-based extractors achieve high performance at 78.4 F1 for event arguments. In-context learning approaches with GPT-4 demonstrate promise for reliable SDoH extraction with limited annotated examples, with extraction performance at 82.3 F1 for event triggers. Yujuan Fu, Giridhar Kaushik Ramachandran, Nicholas J. Dobbins, Namu Park, Michael Leu, Abby R. Rosenberg, Kevin Lybarger, Fei Xia 0004, Özlem Uzuner, Meliha Yetisgen |
LREC/COLING | 10 |
| 2024 | A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language ModelsabstractMedical imaging is critical to the diagnosis, surveillance, and treatment of many health conditions, including oncological, neurological, cardiovascular, and musculoskeletal disorders, among others. Radiologists interpret these complex, unstructured images and articulate their assessments through narrative reports that remain largely unstructured. This unstructured narrative must be converted into a structured semantic representation to facilitate secondary applications such as retrospective analyses or clinical decision support. Here, we introduce the Corpus of Annotated Medical Imaging Reports (CAMIR), which includes 609 annotated radiology reports from three imaging modality types: Computed Tomography, Magnetic Resonance Imaging, and Positron Emission Tomography-Computed Tomography. Reports were annotated using an event-based schema that captures clinical indications, lesions, and medical problems. Each event consists of a trigger and multiple arguments, and a majority of the argument types, including anatomy, normalize the spans to pre-defined concepts to facilitate secondary use. CAMIR uniquely combines a granular event structure and concept normalization. To extract CAMIR events, we explored two BERT (Bi-directional Encoder Representation from Transformers)-based architectures, including an existing architecture (mSpERT) that jointly extracts all event information and a multi-step approach (PL-Marker++) that we augmented for the CAMIR schema. Namu Park, Kevin Lybarger, Giridhar Kaushik Ramachandran, Spencer Lewis, Aashka Damani, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
LREC/COLING | 8 |
| 2024 | To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and ReliabilityabstractUnpredictability, especially unpredictability with unknown error characteristics, is a highly undesirable trait, particularly in medical patient care applications. Although large pre-trained language models (LLM) have been applied to a variety of unseen tasks with highly competitive and successful results, their sensitivity to language inputs and resulting performance variability is not well-studied. In this work, we test state-of-the-art pre-trained language models from a variety of families to characterize their error generation and reliability in medical assessment ability. Particularly, we experiment with general medical assessment multiple choice tests, as well as their open-ended and true-false alternatives. We also profile model consistency, error agreements with each other and to humans; and finally, quantify their ability to recover and explain errors. The findings in this work can be used to give further information about medical models so that modelers can make better-informed decisions rather than relying on standalone performance metrics alone. Wen-Wai Yim, Yujuan Fu, Asma Ben Abacha, Meliha Yetisgen |
LREC/COLING | 4 |
| 2024 | Advancing Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications with ImageCLEF 2024
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Didier Schwab, Emmanuelle Esperança-Rodier, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ioan Coman, Johanna Schöler, Johannes Kiesel, Johannes Rückert, Louise Bloch, Martin Potthast, Maximilian Heinrich, Meliha Yetisgen, Michael Riegler 0001, Neal Snider, Pål Halvorsen, Raphael Brüngel, Steven Alexander Hicks, Vajira Thambawita, Vassili Kovalev, Yuri Prokopchuk, Wen-Wai Yim |
ECIR (6) | 29 |
| 2024 | DermaVQA: A Multilingual Visual Question Answering Dataset for Dermatology
Wen-Wai Yim, Yujuan Fu, Zhaoyi Sun, Asma Ben Abacha, Meliha Yetisgen, Fei Xia 0004 |
MICCAI (5) | 5 |
| 2024 | CACER: Clinical concept Annotations for Cancer Events and RelationsabstractOBJECTIVE: Clinical notes contain unstructured representations of patient histories, including the relationships between medical problems and prescription drugs. To investigate the relationship between cancer drugs and their associated symptom burden, we extract structured, semantic representations of medical problem and drug information from the clinical narratives of oncology notes. MATERIALS AND METHODS: We present Clinical concept Annotations for Cancer Events and Relations (CACER), a novel corpus with fine-grained annotations for over 48 000 medical problems and drug events and 10 000 drug-problem and problem-problem relations. Leveraging CACER, we develop and evaluate transformer-based information extraction models such as Bidirectional Encoder Representations from Transformers (BERT), Fine-tuned Language Net Text-To-Text Transfer Transformer (Flan-T5), Large Language Model Meta AI (Llama3), and Generative Pre-trained Transformers-4 (GPT-4) using fine-tuning and in-context learning (ICL). RESULTS: In event extraction, the fine-tuned BERT and Llama3 models achieved the highest performance at 88.2-88.0 F1, which is comparable to the inter-annotator agreement (IAA) of 88.4 F1. In relation extraction, the fine-tuned BERT, Flan-T5, and Llama3 achieved the highest performance at 61.8-65.3 F1. GPT-4 with ICL achieved the worst performance across both tasks. DISCUSSION: The fine-tuned models significantly outperformed GPT-4 in ICL, highlighting the importance of annotated training data and model optimization. Furthermore, the BERT models performed similarly to Llama3. For our task, large language models offer no performance advantage over the smaller BERT models. CONCLUSIONS: We introduce CACER, a novel corpus with fine-grained annotations for medical problems, drugs, and their relationships in clinical narratives of oncology notes. State-of-the-art transformer models achieved performance comparable to IAA for several extraction tasks. Yujuan Fu, Giridhar Kaushik Ramachandran, Ahmad Halwani, Bridget T. McInnes, Fei Xia 0004, Kevin Lybarger, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Large language models for biomedicine: foundations, opportunities, challenges, and best practicesabstractOBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications. Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Clinical natural language processing for secondary uses
Yanjun Gao, Diwakar Mahajan, Özlem Uzuner, Meliha Yetisgen |
J. Biomed. Informatics | 4 |
| 2023 | ImageCLEF 2023 Highlight: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Adrian Popescu 0001, Ahmad Idrissi-Yaghir, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Christoph M. Friedrich, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ihar Filipovich, Ioan Coman, Jérôme Deshayes-Chossart, Johanna Schöler, Johannes Rückert, Liviu-Daniel Stefan, Louise Bloch, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Gabriel Constantin, Neal Snider, Nikolaos Papachrysos, Pål Halvorsen, Raphael Brüngel, Serge Kozlovski, Steven Alexander Hicks, Thomas de Lange, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim |
ECIR (3) | 23 |
| 2023 | LeafAI: query generator for clinical cohort discovery rivaling a human programmerabstractOBJECTIVE: Identifying study-eligible patients within clinical databases is a critical step in clinical research. However, accurate query design typically requires extensive technical and biomedical expertise. We sought to create a system capable of generating data model-agnostic queries while also providing novel logical reasoning capabilities for complex clinical trial eligibility criteria. MATERIALS AND METHODS: The task of query creation from eligibility criteria requires solving several text-processing problems, including named entity recognition and relation extraction, sequence-to-sequence transformation, normalization, and reasoning. We incorporated hybrid deep learning and rule-based modules for these, as well as a knowledge base of the Unified Medical Language System (UMLS) and linked ontologies. To enable data-model agnostic query creation, we introduce a novel method for tagging database schema elements using UMLS concepts. To evaluate our system, called LeafAI, we compared the capability of LeafAI to a human database programmer to identify patients who had been enrolled in 8 clinical trials conducted at our institution. We measured performance by the number of actual enrolled patients matched by generated queries. RESULTS: LeafAI matched a mean 43% of enrolled patients with 27 225 eligible across 8 clinical trials, compared to 27% matched and 14 587 eligible in queries by a human database programmer. The human programmer spent 26 total hours crafting queries compared to several minutes by LeafAI. CONCLUSIONS: Our work contributes a state-of-the-art data model-agnostic query generation system capable of conditional reasoning using a knowledge base. We demonstrate that LeafAI can rival an experienced human programmer in finding patients eligible for clinical trials. Nicholas J. Dobbins, Weipeng Zhou, Kristine Lan, H. Nina Kim, Robert D. Harrington, Özlem Uzuner, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | Leveraging natural language processing to augment structured social determinants of health data in the electronic health recordabstractOBJECTIVE: Social determinants of health (SDOH) impact health outcomes and are documented in the electronic health record (EHR) through structured data and unstructured clinical notes. However, clinical notes often contain more comprehensive SDOH information, detailing aspects such as status, severity, and temporality. This work has two primary objectives: (1) develop a natural language processing information extraction model to capture detailed SDOH information and (2) evaluate the information gain achieved by applying the SDOH extractor to clinical narratives and combining the extracted representations with existing structured data. MATERIALS AND METHODS: We developed a novel SDOH extractor using a deep learning entity and relation extraction architecture to characterize SDOH across various dimensions. In an EHR case study, we applied the SDOH extractor to a large clinical data set with 225 089 patients and 430 406 notes with social history sections and compared the extracted SDOH information with existing structured data. RESULTS: The SDOH extractor achieved 0.86 F1 on a withheld test set. In the EHR case study, we found extracted SDOH information complements existing structured data with 32% of homeless patients, 19% of current tobacco users, and 10% of drug users only having these health risk factors documented in the clinical narrative. CONCLUSIONS: Utilizing EHR data to identify SDOH health risk factors and social needs may improve patient care and outcomes. Semantic representations of text-encoded SDOH information can augment existing structured data, and this more comprehensive SDOH representation can assist health systems in identifying and addressing these social needs. Kevin Lybarger, Nicholas J. Dobbins, Ritche Long, Angad P. Singh, Patrick Wedgeworth, Özlem Uzuner, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Advancements in extracting social determinants of health information from narrative textabstractSocial determinants of health (SDoH) are the conditions in which people are born, live, work, and age that affect personal well-being, health outcomes, and life expectancy.1 SDoH include a range of nonmedical factors, including substance use, quality of domestic life, marital status, employment status, education, race, geography, and other factors that impact health. Understanding patient SDoH can inform patient health care and has the potential to improve health outcomes and reduce health disparities.2,3 Patient SDoH information is documented in the electronic health record (EHR) and other health-related databases through structured data and free-text (natural language) documents, including patient notes. For many SDoH, the free-text descriptions capture social and behavioral factors with higher prevalence and more detail than is available through structured data. Utilizing free-text SDoH information in large-scale studies, clinical decision-support systems, and other secondary use applications, requires the automatic extraction of key aspects of the SDoH using natural language processing (NLP). NLP-based information extraction maps the unstructured, free-text descriptions of SDoH to structured semantic representations that can be combined with available structured data to create more complete patient profiles. Kevin Lybarger, Oliver J. Bear Don't Walk IV, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | The 2022 n2c2/UW shared task on extracting social determinants of healthabstractOBJECTIVE: The n2c2/UW SDOH Challenge explores the extraction of social determinant of health (SDOH) information from clinical notes. The objectives include the advancement of natural language processing (NLP) information extraction techniques for SDOH and clinical information more broadly. This article presents the shared task, data, participating teams, performance results, and considerations for future work. MATERIALS AND METHODS: The task used the Social History Annotated Corpus (SHAC), which consists of clinical text with detailed event-based annotations for SDOH events, such as alcohol, drug, tobacco, employment, and living situation. Each SDOH event is characterized through attributes related to status, extent, and temporality. The task includes 3 subtasks related to information extraction (Subtask A), generalizability (Subtask B), and learning transfer (Subtask C). In addressing this task, participants utilized a range of techniques, including rules, knowledge bases, n-grams, word embeddings, and pretrained language models (LM). RESULTS: A total of 15 teams participated, and the top teams utilized pretrained deep learning LM. The top team across all subtasks used a sequence-to-sequence approach achieving 0.901 F1 for Subtask A, 0.774 F1 Subtask B, and 0.889 F1 for Subtask C. CONCLUSIONS: Similar to many NLP tasks and domains, pretrained LM yielded the best performance, including generalizability and learning transfer. An error analysis indicates extraction performance varies by SDOH, with lower performance achieved for conditions, like substance use and homelessness, which increase health risks (risk factors) and higher performance achieved for conditions, like substance abstinence and living with family, which reduce health risks (protective factors). Kevin Lybarger, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Improving model transferability for clinical note section classification models using continued pretrainingabstractOBJECTIVE: The classification of clinical note sections is a critical step before doing more fine-grained natural language processing tasks such as social determinants of health extraction and temporal information extraction. Often, clinical note section classification models that achieve high accuracy for 1 institution experience a large drop of accuracy when transferred to another institution. The objective of this study is to develop methods that classify clinical note sections under the SOAP ("Subjective," "Object," "Assessment," and "Plan") framework with improved transferability. MATERIALS AND METHODS: We trained the baseline models by fine-tuning BERT-based models, and enhanced their transferability with continued pretraining, including domain-adaptive pretraining and task-adaptive pretraining. We added in-domain annotated samples during fine-tuning and observed model performance over a varying number of annotated sample size. Finally, we quantified the impact of continued pretraining in equivalence of the number of in-domain annotated samples added. RESULTS: We found continued pretraining improved models only when combined with in-domain annotated samples, improving the F1 score from 0.756 to 0.808, averaged across 3 datasets. This improvement was equivalent to adding 35 in-domain annotated samples. DISCUSSION: Although considered a straightforward task when performing in-domain, section classification is still a considerably difficult task when performing cross-domain, even using highly sophisticated neural network-based methods. CONCLUSION: Continued pretraining improved model transferability for cross-domain clinical note section classification in the presence of a small amount of in-domain labeled samples. Weipeng Zhou, Meliha Yetisgen, Majid Afshar, Yanjun Gao, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Extracting medication changes in clinical narratives using pre-trained language models
Giridhar Kaushik Ramachandran, Kevin Lybarger, Yaya Liu, Diwakar Mahajan, Jennifer J. Liang, Ching-Huei Tsou, Meliha Yetisgen, Özlem Uzuner |
J. Biomed. Informatics | 7 |
| 2022 | A Corpus of Radiology Reports From Multiple Imaging Modalities With Fine-grained Event-based Annotations
Kevin Lybarger, Namu Park, Sitong Zhou, Aashka Damani, Alison Brennan, Jagjeet Gill, Nianiella Dorvall, Vy Huynh, Spencer Lewis, Martin L. Gunn, Özlem Uzuner, Meliha Yetisgen |
AMIA | 12 |
| 2022 | Call for papers: Special issue on clinical natural language processing for secondary use applications
Meliha Yetisgen, Özlem Uzuner, Yanjun Gao, Diwakar Mahajan |
J. Biomed. Informatics | 1 |
| 2021 | Automatic Detection of Surgical Site Infections Using EHR Data
Arjun Chakraborty, Kevin Lybarger, Dustin Long, Vikas O. Shah, Meliha Yetisgen |
AMIA | 5 |
| 2021 | Automatic Assignment of Radiology Examination Protocols Using Pre-trained Language Models with Knowledge Distillation
Wilson Lau, Laura Aaltonen, Martin L. Gunn, Meliha Yetisgen |
AMIA | 4 |
| 2021 | A New Corpus for Clinical Findings in Radiology Reports
Wilson Lau, David Wayne, Spencer Lewis, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
AMIA | 6 |
| 2021 | Identifying ARDS using the Hierarchical Attention Network with Sentence Objectives Framework
Kevin Lybarger, Linzee Mabrey, Matthew Thau, Pavan K. Bhatraju, Mark M. Wurfel, Meliha Yetisgen |
AMIA | 6 |
| 2021 | An exploration of information extraction models on transcribed patient visits
Kevin Lybarger, Erica Qiao, Meliha Yetisgen |
AMIA | 3 |
| 2021 | Comparison of Three Phrase Chunking Approaches on Medical Concept Coverage in Clinical Text
Grace K. Turner, Meliha Yetisgen |
AMIA | 2 |
| 2021 | Transferability of neural network clinical deidentification systemsabstractOBJECTIVE: Neural network deidentification studies have focused on individual datasets. These studies assume the availability of a sufficient amount of human-annotated data to train models that can generalize to corresponding test data. In real-world situations, however, researchers often have limited or no in-house training data. Existing systems and external data can help jump-start deidentification on in-house data; however, the most efficient way of utilizing existing systems and external data is unclear. This article investigates the transferability of a state-of-the-art neural clinical deidentification system, NeuroNER, across a variety of datasets, when it is modified architecturally for domain generalization and when it is trained strategically for domain transfer. MATERIALS AND METHODS: We conducted a comparative study of the transferability of NeuroNER using 4 clinical note corpora with multiple note types from 2 institutions. We modified NeuroNER architecturally to integrate 2 types of domain generalization approaches. We evaluated each architecture using 3 training strategies. We measured transferability from external sources; transferability across note types; the contribution of external source data when in-domain training data are available; and transferability across institutions. RESULTS AND CONCLUSIONS: Transferability from a single external source gave inconsistent results. Using additional external sources consistently yielded an F1-score of approximately 80%. Fine-tuning emerged as a dominant transfer strategy, with or without domain generalization. We also found that external sources were useful even in cases where in-domain training data were available. Transferability across institutions differed by note type and annotation label but resulted in improved performance. Kahyun Lee, Nicholas J. Dobbins, Bridget T. McInnes, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Extracting COVID-19 diagnoses and symptoms from clinical text: A new annotated corpus and neural event extraction framework
Kevin Lybarger, Mari Ostendorf, Meliha Yetisgen |
J. Biomed. Informatics | 4 |
| 2021 | Annotating social determinants of health using active learning, and characterizing determinants using neural event extraction
Kevin Lybarger, Mari Ostendorf, Meliha Yetisgen |
J. Biomed. Informatics | 3 |
| 2020 | Jointly Learning Clinical Entities and Relations with Contextual Language Models and Explicit Context
Paul Barry, Sam Henry 0001, Meliha Yetisgen, Bridget T. McInnes, Özlem Uzuner |
AMIA | 3 |
| 2020 | A Novel Corpus With Detailed Annotations of Social Determinants of Health
Kevin Lybarger, Kylie Kerker, Jolie Shen, Erica Qiao, Özlem Uzuner, Mari Ostendorf, Meliha Yetisgen |
AMIA | 8 |
| 2020 | Alignment Annotation for Clinic Visit Dialogue to Clinical Note Sentence Language GenerationabstractFor every patient’s visit to a clinician, a clinical note is generated documenting their medical conversation, including complaints discussed, treatments, and medical plans. Despite advances in natural language processing, automating clinical note generation from a clinic visit conversation is a largely unexplored area of research. Due to the idiosyncrasies of the task, traditional methods of corpus creation are not effective enough approaches for this problem. In this paper, we present an annotation methodology that is content- and technique- agnostic while associating note sentences to sets of dialogue sentences. The sets can further be grouped with higher order tags to mark sets with related information. This direct linkage from input to output decouples the annotation from specific language understanding or generation strategies. Here we provide data statistics and qualitative analysis describing the unique annotation challenges. Given enough annotated data, such a resource would support multiple modeling methods including information extraction with template language generation, information retrieval type language generation, or sequence to sequence modeling. Wen-Wai Yim, Meliha Yetisgen, Jenny Huang, Micah Grossman |
LREC | 2 |
| 2020 | Predicting severe clinical events by learning about life-saving actions and outcomes using distant supervision
Dae Hyun Lee, Meliha Yetisgen, Lucy Vanderwende, Eric Horvitz |
J. Biomed. Informatics | 2 |
| 2019 | Automatic Identification of Social Determinants of Health from Clinical Records
Kevin Lybarger, Mari Ostendorf, Özlem Uzuner, Meliha Yetisgen |
AMIA | 4 |
| 2018 | Learning about Life-Saving Interventions to Predict the Risk of Acute Organ Failures
Dae Hyun Lee, Meliha Yetisgen, Eric Horvitz |
AMIA | 2 |
| 2018 | Using Neural Multi-task Learning to Extract Substance Abuse Information from Clinical Notes
Kevin Lybarger, Meliha Yetisgen, Mari Ostendorf |
AMIA | 2 |
| 2018 | Hierarchical Attention-Based Prediction Model for Discovering the Persistence of Chronic Opioid Therapy from a large Clinical Dataset
Ramón Maldonado, Mark Sullivan, Meliha Yetisgen, Sanda M. Harabagiu |
AMIA | 3 |
| 2017 | Automatic Identification of Substance Abuse from Social History in Clinical Text
Meliha Yetisgen, Lucy Vanderwende |
AIME | 1 |
| 2017 | Clustering Vital Sign Observations Using Unsupervised Random Forest
Dae Hyun Lee, Meliha Yetisgen |
AMIA | 2 |
| 2017 | Automatically Detecting Likely Edits in Clinical Notes Created Using Automatic Speech Recognition
Kevin Lybarger, Mari Ostendorf, Meliha Yetisgen |
AMIA | 3 |
| 2017 | Improving Electronic Inpatient Progress Notes Using Voice: Results from the VGEENS Project
Thomas H. Payne, Andrew Markiel, William D. Alonso, Ross J. Lordon, Kevin Lybarger, Meliha Yetisgen, Jennifer M. Zech, Andrew A. White |
AMIA | 6 |
| 2017 | Classification of hepatocellular carcinoma stages from free-text clinical and radiology reports
Wen-Wai Yim, Sharon W. Kwan, Guy Johnson, Meliha Yetisgen |
AMIA | 4 |
| 2017 | Classifying tumor event attributes in radiology reportsabstractRadiology reports contain vital diagnostic information that characterizes patient disease progression. However, information from reports is represented in free text, which is difficult to query against for secondary use. Automatic extraction of important information, such as tumor events using natural language processing, offers possibilities in improved clinical decision support, cohort identification, and retrospective evidence‐based research for cancer patients. The goal of this work was to classify tumor event attributes: negation, temporality, and malignancy, using biomedical ontology and linguistically enriched features. We report our results on an annotated corpus of 101 hepatocellular carcinoma patient radiology reports, and show that the improved classification improves overall template structuring. Classification performances for negation identification, past temporality classification, and malignancy classification were at 0.94, 0.62, and 0.77 F1, respectively. Incorporating the attributes into full templates led to an improvement of 0.72 F1 for tumor‐related events over a baseline of 0.65 F1. Improvement of negation, malignancy, and temporality classifications led to significant improvements in template extraction for the majority of categories. We present our machine‐learning approach to identifying these several tumor event attributes from radiology reports, as well as highlight challenges and areas for improvement. Wen-Wai Yim, Sharon W. Kwan, Meliha Yetisgen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Annotation of Tumor Reference Resolution and Tumor Characteristics for Cancer Liver Stage Prediction
Wen-Wai Yim, Tyler Denman, Sharon W. Kwan, Meliha Yetisgen |
AMIA | 4 |
| 2016 | Annotating and Detecting Medical Events in Clinical Notes
Prescott Klassen, Fei Xia 0004, Meliha Yetisgen |
LREC | 3 |
| 2016 | Tumor reference resolution and characteristic extraction in radiology reports for liver cancer stage prediction
Wen-Wai Yim, Sharon W. Kwan, Meliha Yetisgen |
J. Biomed. Informatics | 3 |
| 2015 | Annotating Recommendation Sentences in Radiology Reports
Meliha Yetisgen, Prescott Klassen, Lucas McCarthy, Thomas H. Payne, Martin L. Gunn |
AMIA | 1 |
| 2015 | Annotation of Disease Characteristics for Cancer Liver Stage Prediction
Wen-Wai Yim, Sharon W. Kwan, Guy Johnson, Meliha Yetisgen |
AMIA | 4 |
| 2014 | A New Corpus for Structured Microbiology Results
Wen-Wai Yim, Xavier Engle, Heather L. Evans, Meliha Yetisgen |
AMIA | 4 |
| 2014 | Annotating Clinical Events in Text Snippets for Phenotype Detection
Prescott Klassen, Fei Xia 0004, Lucy Vanderwende, Meliha Yetisgen |
LREC | 4 |
| 2013 | On-time clinical phenotype prediction based on narrative reports
Cosmin Adrian Bejan, Lucy Vanderwende, Heather L. Evans, Mark M. Wurfel, Meliha Yetisgen |
AMIA | 5 |
| 2013 | Assertion modeling and its role in clinical phenotype identification
Cosmin Adrian Bejan, Lucy Vanderwende, Fei Xia 0004, Meliha Yetisgen |
J. Biomed. Informatics | 4 |
| 2013 | Text classification for assisting moderators in online health communities
Jina Huh, Meliha Yetisgen, Wanda Pratt |
J. Biomed. Informatics | 2 |
| 2013 | A text processing pipeline to extract recommendations from radiology reports
Meliha Yetisgen, Martin L. Gunn, Fei Xia 0004, Thomas H. Payne |
J. Biomed. Informatics | 1 |
| 2012 | Assessing Pneumonia Identification from Time-Ordered Narrative Reports
Cosmin Adrian Bejan, Lucy Vanderwende, Mark M. Wurfel, Meliha Yetisgen |
AMIA | 4 |
| 2012 | Text Classification to Weave Medical Advice with Patient Experiences
Jina Huh, Meliha Yetisgen, Andrea L. Hartzler, David W. McDonald, Albert Park, Wanda Pratt |
AMIA | 2 |
| 2012 | Smoking Status Detection Across Domains
Michael Tepper, Fei Xia 0004, Meliha Yetisgen |
AMIA | 3 |
| 2012 | Statistical Section Segmentation in Free-Text Clinical Records
Michael Tepper, Daniel Capurro, Fei Xia 0004, Lucy Vanderwende, Meliha Yetisgen |
LREC | 5 |
| 2012 | Pneumonia identification using statistical feature selectionabstractOBJECTIVE: This paper describes a natural language processing system for the task of pneumonia identification. Based on the information extracted from the narrative reports associated with a patient, the task is to identify whether or not the patient is positive for pneumonia. DESIGN: A binary classifier was employed to identify pneumonia from a dataset of multiple types of clinical notes created for 426 patients during their stay in the intensive care unit. For this purpose, three types of features were considered: (1) word n-grams, (2) Unified Medical Language System (UMLS) concepts, and (3) assertion values associated with pneumonia expressions. System performance was greatly increased by a feature selection approach which uses statistical significance testing to rank features based on their association with the two categories of pneumonia identification. RESULTS: Besides testing our system on the entire cohort of 426 patients (unrestricted dataset), we also used a smaller subset of 236 patients (restricted dataset). The performance of the system was compared with the results of a baseline previously proposed for these two datasets. The best results achieved by the system (85.71 and 81.67 F1-measure) are significantly better than the baseline results (50.70 and 49.10 F1-measure) on the restricted and unrestricted datasets, respectively. CONCLUSION: Using a statistical feature selection approach that allows the feature extractor to consider only the most informative features from the feature space significantly improves the performance over a baseline that uses all the features from the same feature space. Extracting the assertion value for pneumonia expressions further improves the system performance. Cosmin Adrian Bejan, Fei Xia 0004, Lucy Vanderwende, Mark M. Wurfel, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 5 |
| 2011 | Dynamic categorization of clinical research eligibility criteria by hierarchical clustering
Jake Luo, Meliha Yetisgen, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2009 | A new evaluation methodology for literature-based discovery systems
Meliha Yetisgen, Wanda Pratt |
J. Biomed. Informatics | 1 |
| 2008 | Finding the Meaning of Medical Concept Correlations
Meliha Yetisgen, Wanda Pratt |
AMIA | 1 |
| 2007 | Extracting the meaning of medical concept correlationsabstractIn this paper, we propose a new method to extract the meaning of medical concept correlations from MEDLINE abstract sentences. Our method incorporates a medical knowledge base, natural language processing approaches, and text classification methods. We describe how we automatically created the training sets and report the results of our initial experiments. Meliha Yetisgen, Wanda Pratt |
K-CAP | 1 |
| 2007 | Response to ''Validating discovery in literature-based discovery"
Wanda Pratt, Meliha Yetisgen |
J. Biomed. Informatics | 2 |
| 2006 | Using statistical and knowledge-based approaches for literature-based discovery
Meliha Yetisgen, Wanda Pratt |
J. Biomed. Informatics | 1 |
| 2005 | The Effect of Feature Representation on MEDLINE Document Classification
Meliha Yetisgen, Wanda Pratt |
AMIA | 1 |
| 2005 | Data mining in deductive databases using query flocks
Ismail Hakki Toroslu, Meliha Yetisgen |
Expert Syst. Appl. | 2 |
| 2003 | A Study of Biomedical Concept Identification: MetaMap vs. People
Wanda Pratt, Meliha Yetisgen |
AMIA | 2 |
| 2003 | LitLinker: capturing connections across the biomedical literatureabstractThe explosive growth in the biomedical literature has made it difficult for researchers to keep up with advancements, even in their own narrow specializations. In addition, this current volume of information has created barriers that prevent researchers from exploring connections to their own work from other parts of the literature. Although potentially useful connections might permeate the literature, they will remain buried without new kinds of tools to help researchers capture new knowledge that bridges gaps across distinct sections of the literature. In this paper, we present LitLinker, a system that incorporates knowledge-based methodologies, natural-language processing techniques, and a data-mining algorithm to mine the biomedical literature for new, potential causal links between biomedical terms. Our results from a well-known text-mining example show that LitLinker can capture these novel, interesting connections in an open-ended fashion, with less manual intervention than in previous systems. Wanda Pratt, Meliha Yetisgen |
K-CAP | 2 |
| 2000 | Data Mining Using Query Flocks with Views
Meliha Yetisgen, Ismail Hakki Toroslu |
DEXA | 1 |