EDBT 2026 Demo / reviewers in the wild / expert
Joseph M. Plasek
dblp:119/3120
· DBLP profile ↗
21ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-9686-3876ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 5 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interactive active learning for literature screening: finetuning GPT with DeepSeek reasoning for cross-domain generalizationabstractOBJECTIVE: Automated literature screening in biomedical research is often hindered by domain shifts and scarcity of labeled data, which limit model accuracy and generalizability. While large language models (LLMs) perform well in zero-shot settings, they often fail to capture complex, domain-specific reasoning patterns. To address this limitation, this study investigates whether an interactive, weakly supervised learning framework combining GPT (generative pre-trained transformer)'s fine-tuning adaptability with DeepSeek's reasoning capabilities can improve literature screening performance across biomedical domains. MATERIALS AND METHODS: We developed an active learning framework that leverages model disagreement between GPT-4o and DeepSeek to improve literature screening performance. This process began with a labeled corpus of 6331 articles on large language models, from which a model disagreement analysis was performed to identify cases where GPT-4o misclassified and DeepSeek produced correct predictions. Three GPT variants-GPT-4o, GPT-4o-mini, and GPT-4.1-nano, were fine-tuned under standard supervised learning settings using these disagreement-based samples. Fine-tuning prompts incorporated classification labels and, when available, rationale traces generated by DeepSeek to provide reasoning-augmented weak supervision. Model performance was evaluated on an independent benchmark set of 291 annotated articles across 10 topic queries in cancer immunotherapy and LLMs in medicine, using standard evaluation metrics, with recall as the primary measure. RESULTS: Fine-tuning GPT models using disagreement-based examples significantly improved performance. GPT-4o-mini achieved the best overall results after fine-tuning, especially with the highest F1 score (0.93, P < .001) and recall (0.95, P < .001). Across the biomedical topics, fine-tuned models consistently outperformed their zero-shot counterparts without increasing reviewer workload. DISCUSSION: These findings demonstrate the effectiveness of disagreement-driven active learning in enhancing GPT-based biomedical literature screening. Lightweight models like GPT-4o-mini benefit most from targeted, reasoning-enriched training, highlighting their suitability for scalable deployment. CONCLUSION: This study introduces an interactive active learning framework that leverages fine-tuned LLMs with reasoning capabilities to enhance literature screening. The approach offers a scalable solution to more efficient and reliable information retrieval in systematic reviews. Joseph M. Plasek, Xinsong Du, Yifei Wang 0002, Zhengyang Zhou, John Lian, Ya-Wen Chuang, Pengyu Hong, Peter C. Hou, Li Zhou 0007 |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies
Richard Wyss, Jie Yang 0039, Sebastian Schneeweiss, Joseph M. Plasek, Li Zhou 0007, Thomas DeRamus, Janick Weberpals, Kerry Ngan, Theodore N. Tsacogianis, Kueiyu Joshua Lin |
J. Biomed. Informatics | 4 |
| 2024 | Large language models for biomedicine: foundations, opportunities, challenges, and best practicesabstractOBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications. Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Characterizing terminology applied by authors and database producers to informatics literature on consumer engagement with wearable devicesabstractOBJECTIVE: Identifying consumer health informatics (CHI) literature is challenging. To recommend strategies to improve discoverability, we aimed to characterize controlled vocabulary and author terminology applied to a subset of CHI literature on wearable technologies. MATERIALS AND METHODS: To retrieve articles from PubMed that addressed patient/consumer engagement with wearables, we developed a search strategy of textwords and Medical Subject Headings (MeSH). To refine our methodology, we used a random sample of 200 articles from 2016 to 2018. A descriptive analysis of articles (N = 2522) from 2019 identified 308 (12.2%) CHI-related articles, for which we characterized their assigned terminology. We visualized the 100 most frequent terms assigned to the articles from MeSH, author keywords, CINAHL, and Engineering Databases (Compendex and Inspec together). We assessed the overlap of CHI terms among sources and evaluated terms related to consumer engagement. RESULTS: The 308 articles were published in 181 journals, more in health journals (82%) than informatics (11%). Only 44% were indexed with the MeSH term "wearable electronic devices." Author keywords were common (91%) but rarely represented consumer engagement with device data, eg, self-monitoring (n = 12, 0.7%) or self-management (n = 9, 0.5%). Only 10 articles (3%) had terminology from all sources (authors, PubMed, CINAHL, Compendex, and Inspec). DISCUSSION: Our main finding was that consumer engagement was not well represented in health and engineering database thesauri. CONCLUSIONS: Authors of CHI studies should indicate consumer/patient engagement and the specific technology investigated in titles, abstracts, and author keywords to facilitate discovery by readers and expand vocabularies and indexing. Kristine M. Alpi, Christie L. Martin, Joseph M. Plasek, Scott M. Sittig, Catherine Arnott Smith, Elizabeth Weinfurter, Jennifer K. Wells, Rachel Wong, Robin Austin |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Learning from undercoded clinical records for automated International Classification of Diseases (ICD) codingabstractOBJECTIVES: To develop an unbiased objective for learning automatic coding algorithms from clinical records annotated with only partial relevant International Classification of Diseases codes, as annotation noise in undercoded clinical records used as training data can mislead the learning process of deep neural networks. MATERIALS AND METHODS: We use Medical Information Mart for Intensive Care III as our dataset. We employ positive-unlabeled learning to achieve unbiased loss estimation, which is free of misleading training signal. We then utilize reweighting mechanism to compensate for the imbalance between positive and negative samples. To further close the performance gap caused by poor quality annotation, we integrate the supervision provided by the automatic annotation tool Medical Concept Annotation Toolkit which can ease the heavy burden of manual validation. RESULTS: Our benchmarking results show that positive-unlabeled learning with reweighting outperforms competitive baseline methods over a range of missing label ratios. Integrating supervision provided by annotation tool further boosted the performance. DISCUSSION: Considering the annotation noise and severe imbalance, unbiased loss estimation and reweighting mechanism are both important for learning from undercoded clinical records. Unbiased loss requires the estimation of false negative ratios and estimation through trained models is practical and competitive. CONCLUSIONS: The combination of positive-unlabeled learning with reweighting and supervision provided by the annotation tool is a promising solution to learn from undercoded clinical records. Yun Xiong, Dan Shi 0006, Yifei Lin, Lifang He 0001, Yao Zhang 0009, Joseph M. Plasek, Li Zhou 0007, David W. Bates, Chunlei Tang |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | Using Twitter Data to Understand Public Perceptions of Approved versus Off-label Use for COVID-19-related Medications
Yining Hua, Jie Yang 0039, Shixu Lin, Joseph M. Plasek, David W. Bates, Li Zhou 0007 |
AMIA | 5 |
| 2022 | Sampling Adverse Drug Events in Outpatient Clinical Notes for Natural Language Processing Tasks
Joseph M. Plasek, Abigail Salem, Stuart R. Lipsitz, Mary G. Amato, Dinah Foer, Heba Edrees, Suzanne V. Blackley, Brett R. South, Amol Rajmane, Mario Lorenzo, Paul Felt, Brendan Bull, Gretchen Purcell Jackson, Henry Feldman, David W. Bates, Li Zhou 0007 |
AMIA | 1 |
| 2022 | Using Twitter data to understand public perceptions of approved versus off-label use for COVID-19-related medicationsabstractOBJECTIVE: Understanding public discourse on emergency use of unproven therapeutics is essential to monitor safe use and combat misinformation. We developed a natural language processing-based pipeline to understand public perceptions of and stances on coronavirus disease 2019 (COVID-19)-related drugs on Twitter across time. METHODS: This retrospective study included 609 189 US-based tweets between January 29, 2020 and November 30, 2021 on 4 drugs that gained wide public attention during the COVID-19 pandemic: (1) Hydroxychloroquine and Ivermectin, drug therapies with anecdotal evidence; and (2) Molnupiravir and Remdesivir, FDA-approved treatment options for eligible patients. Time-trend analysis was used to understand the popularity and related events. Content and demographic analyses were conducted to explore potential rationales of people's stances on each drug. RESULTS: Time-trend analysis revealed that Hydroxychloroquine and Ivermectin received much more discussion than Molnupiravir and Remdesivir, particularly during COVID-19 surges. Hydroxychloroquine and Ivermectin were highly politicized, related to conspiracy theories, hearsay, celebrity effects, etc. The distribution of stance between the 2 major US political parties was significantly different (P < .001); Republicans were much more likely to support Hydroxychloroquine (+55%) and Ivermectin (+30%) than Democrats. People with healthcare backgrounds tended to oppose Hydroxychloroquine (+7%) more than the general population; in contrast, the general population was more likely to support Ivermectin (+14%). CONCLUSION: Our study found that social media users with have different perceptions and stances on off-label versus FDA-authorized drug use across different stages of COVID-19, indicating that health systems, regulatory agencies, and policymakers should design tailored strategies to monitor and reduce misinformation for promoting safe drug use. Our analysis pipeline and stance detection models are made public at https://github.com/ningkko/COVID-drug. Yining Hua, Shixu Lin, Jie Yang 0039, Joseph M. Plasek, David W. Bates, Li Zhou 0007 |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Addressing Disparities in Diabetes Using Temporal Fairness Models
Joseph M. Plasek, Chunlei Tang, Yun Xiong, Yangyong Zhu, Yanming He, Patricia C. Dykes, David W. Bates, Li Zhou 0007 |
AMIA | 1 |
| 2021 | Estimating Time to Progression of Chronic Obstructive Pulmonary Disease With ToleranceabstractWe defined tolerance range as the distance of observing similar disease conditions or functional status from the upper to the lower boundaries of a specified time interval. A tolerance range was identified for linear regression and support vector machines to optimize the improvement rate (defined as IR) on accuracy in predicting mortality risk in patients with chronic obstructive pulmonary disease using clinical notes. The corpus includes pulmonary, cardiology, and radiology reports of 15,500 patients who died between 2011 and 2017. Their performance was compared against a long short-term memory recurrent neural network. The results demonstrate an overall improvement by those basic machine learning approaches after considering an optimal tolerance range: the average IR of linear regression was 90.1% and the maximum IR of support vector machines was 66.2%. There was a similitude between the time segments produced by our tolerance algorithms and those produced by the long short-term memory. Chunlei Tang, Joseph M. Plasek, Meihan Wan, Min-Jeoung Kang, Sevan M. Dulgarian, Yun Xiong, David W. Bates, Li Zhou 0007 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Following data as it crosses borders during the COVID-19 pandemicabstractData change the game in terms of how we respond to pandemics. Global data on disease trajectories and the effectiveness and economic impact of different social distancing measures are essential to facilitate effective local responses to pandemics. COVID-19 data flowing across geographic borders are extremely useful to public health professionals for many purposes such as accelerating the pharmaceutical development pipeline, and for making vital decisions about intensive care unit rooms, where to build temporary hospitals, or where to boost supplies of personal protection equipment, ventilators, or diagnostic tests. Sharing data enables quicker dissemination and validation of pharmaceutical innovations, as well as improved knowledge of what prevention and mitigation measures work. Even if physical borders around the globe are closed, it is crucial that data continues to transparently flow across borders to enable a data economy to thrive, which will promote global public health through global cooperation and solidarity. Joseph M. Plasek, Chunlei Tang, Yangyong Zhu, Yajun Huang, David W. Bates |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Harmonizing Oncology Data from Multiple Institutions with the Observational Medical Outcomes Partnership's Common Data Model
Joseph M. Plasek, John Weissert, Kourosh Ravvaz |
AMIA | 1 |
| 2019 | Data Reconstruction Based on Temporal Expressions in Clinical NotesabstractLearning representations of clinical notes poses challenges in handling complex content that necessitates preprocessing steps to make the data more suitable for data mining. An important issue, addressed here, is that of temporal expressions, where cues indicate the time when clinical events occur. We present a three-step data reconstruction algorithm for transforming similar clinical entities (e.g., symptoms, complications) into sequential data through unsupervised annotation of temporal expressions. First, the data reconstruction algorithm detects if an expression has temporal intent. Second, it decomposes and rewrites the expression into non-temporal sub-expression and temporal constraints. Finally, it clusters similar non-temporal sub-expressions by using unsupervised sentence embedding under the modified K-medoids paradigm. We experimented with our proposed algorithm on clinical notes associated with chronic obstructive pulmonary disease (COPD). Visualizing reconstruction results of cardiology reports for a longitudinal cohort of patients with COPD demonstrated that this algorithm is feasible. Chunlei Tang, Joseph M. Plasek, Yun Xiong, Min-Jeoung Kang, Patricia C. Dykes, David W. Bates, Li Zhou 0007 |
BIBM | 3 |
| 2018 | Comparing Machine Learning Algorithms to Predict Falls of Community Dwelling Older Adults
Rumei Yang, Joseph M. Plasek, Mollie R. Cummins, Katherine A. Sward |
AMIA | 2 |
| 2018 | RegionAl: an Optimized Regional Classifier to Predict Mortality in Chronic Obstructive Pulmonary Disease Patients
Chunlei Tang, Joseph M. Plasek, Yun Xiong, Li Zhou 0007, David W. Bates |
AMIA | 3 |
| 2018 | A Deep Learning Approach to Handling Temporal Variation in Chronic Obstructive Pulmonary Disease Progression
Chunlei Tang, Joseph M. Plasek, Yun Xiong, David W. Bates, Li Zhou 0007 |
BIBM | 2 |
| 2018 | A value set for documenting adverse reactions in electronic health recordsabstractObjective: To develop a comprehensive value set for documenting and encoding adverse reactions in the allergy module of an electronic health record. Materials and Methods: We analyzed 2 471 004 adverse reactions stored in Partners Healthcare's Enterprise-wide Allergy Repository (PEAR) of 2.7 million patients. Using the Medical Text Extraction, Reasoning, and Mapping System, we processed both structured and free-text reaction entries and mapped them to Systematized Nomenclature of Medicine - Clinical Terms. We calculated the frequencies of reaction concepts, including rare, severe, and hypersensitivity reactions. We compared PEAR concepts to a Federal Health Information Modeling and Standards value set and University of Nebraska Medical Center data, and then created an integrated value set. Results: We identified 787 reaction concepts in PEAR. Frequently reported reactions included: rash (14.0%), hives (8.2%), gastrointestinal irritation (5.5%), itching (3.2%), and anaphylaxis (2.5%). We identified an additional 320 concepts from Federal Health Information Modeling and Standards and the University of Nebraska Medical Center to resolve gaps due to missing and partial matches when comparing these external resources to PEAR. This yielded 1106 concepts in our final integrated value set. The presence of rare, severe, and hypersensitivity reactions was limited in both external datasets. Hypersensitivity reactions represented roughly 20% of the reactions within our data. Discussion: We developed a value set for encoding adverse reactions using a large dataset from one health system, enriched by reactions from 2 large external resources. This integrated value set includes clinically important severe and hypersensitivity reactions. Conclusion: This work contributes a value set, harmonized with existing data, to improve the consistency and accuracy of reaction documentation in electronic health records, providing the necessary building blocks for more intelligent clinical decision support for allergies and adverse reactions. Foster R. Goss, Kenneth H. Lai, Maxim Topaz, Warren W. Acker, Leigh Kowalski, Joseph M. Plasek, Kimberly G. Blumenthal, Diane L. Seger, Sarah P. Slight, Kin Wah Fung, Frank Y. Chang, David W. Bates, Li Zhou 0007 |
J. Am. Medical Informatics Assoc. | 6 |
| 2016 | Food entries in a large allergy data repositoryabstractOBJECTIVE: Accurate food adverse sensitivity documentation in electronic health records (EHRs) is crucial to patient safety. This study examined, encoded, and grouped foods that caused any adverse sensitivity in a large allergy repository using natural language processing and standard terminologies. METHODS: Using the Medical Text Extraction, Reasoning, and Mapping System (MTERMS), we processed both structured and free-text entries stored in an enterprise-wide allergy repository (Partners' Enterprise-wide Allergy Repository), normalized diverse food allergen terms into concepts, and encoded these concepts using the Systematized Nomenclature of Medicine - Clinical Terms (SNOMED-CT) and Unique Ingredient Identifiers (UNII) terminologies. Concept coverage also was assessed for these two terminologies. We further categorized allergen concepts into groups and calculated the frequencies of these concepts by group. Finally, we conducted an external validation of MTERMS's performance when identifying food allergen terms, using a randomized sample from a different institution. RESULTS: We identified 158 552 food allergen records (2140 unique terms) in the Partners repository, corresponding to 672 food allergen concepts. High-frequency groups included shellfish (19.3%), fruits or vegetables (18.4%), dairy (9.0%), peanuts (8.5%), tree nuts (8.5%), eggs (6.0%), grains (5.1%), and additives (4.7%). Ambiguous, generic concepts such as "nuts" and "seafood" accounted for 8.8% of the records. SNOMED-CT covered more concepts than UNII in terms of exact (81.7% vs 68.0%) and partial (14.3% vs 9.7%) matches. DISCUSSION: Adverse sensitivities to food are diverse, and existing standard terminologies have gaps in their coverage of the breadth of allergy concepts. CONCLUSION: New strategies are needed to represent and standardize food adverse sensitivity concepts, to improve documentation in EHRs. Joseph M. Plasek, Foster R. Goss, Kenneth H. Lai, Jason J. Lau, Diane L. Seger, Kimberly G. Blumenthal, Paige G. Wickner, Sarah P. Slight, Frank Y. Chang, Maxim Topaz, David W. Bates, Li Zhou 0007 |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | An Evaluation of a Natural Language Processing Tool for Identifying and Encoding Allergy Information in Emergency Department Clinical Notes
Foster R. Goss, Joseph M. Plasek, Jason J. Lau, Diane L. Seger, Frank Y. Chang, Li Zhou 0007 |
AMIA | 2 |
| 2013 | Evaluating standard terminologies for encoding allergy informationabstractOBJECTIVE: Allergy documentation and exchange are vital to ensuring patient safety. This study aims to analyze and compare various existing standard terminologies for representing allergy information. METHODS: Five terminologies were identified, including the Systemized Nomenclature of Medical Clinical Terms (SNOMED CT), National Drug File-Reference Terminology (NDF-RT), Medication Dictionary for Regulatory Activities (MedDRA), Unique Ingredient Identifier (UNII), and RxNorm. A qualitative analysis was conducted to compare desirable characteristics of each terminology, including content coverage, concept orientation, formal definitions, multiple granularities, vocabulary structure, subset capability, and maintainability. A quantitative analysis was also performed to compare the content coverage of each terminology for (1) common food, drug, and environmental allergens and (2) descriptive concepts for common drug allergies, adverse reactions (AR), and no known allergies. RESULTS: Our qualitative results show that SNOMED CT fulfilled the greatest number of desirable characteristics, followed by NDF-RT, RxNorm, UNII, and MedDRA. Our quantitative results demonstrate that RxNorm had the highest concept coverage for representing drug allergens, followed by UNII, SNOMED CT, NDF-RT, and MedDRA. For food and environmental allergens, UNII demonstrated the highest concept coverage, followed by SNOMED CT. For representing descriptive allergy concepts and adverse reactions, SNOMED CT and NDF-RT showed the highest coverage. Only SNOMED CT was capable of representing unique concepts for encoding no known allergies. CONCLUSIONS: The proper terminology for encoding a patient's allergy is complex, as multiple elements need to be captured to form a fully structured clinical finding. Our results suggest that while gaps still exist, a combination of SNOMED CT and RxNorm can satisfy most criteria for encoding common allergies and provide sufficient content coverage. Foster R. Goss, Li Zhou 0007, Joseph M. Plasek, Carol A. Broverman, George A. Robinson, Blackford Middleton, Roberto A. Rocha |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Mapping Partners Master Drug Dictionary to RxNorm using an NLP-based approach
Li Zhou 0007, Joseph M. Plasek, Lisa M. Mahoney, Frank Y. Chang, Dana DiMaggio, Roberto A. Rocha |
J. Biomed. Informatics | 2 |