EDBT 2026 Demo / reviewers in the wild / expert
Yonghui Wu 0001
dblp:26/2189-1
· DBLP profile ↗
54ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-6780-6135ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 52 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysisabstractOBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases. Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor |
J. Am. Medical Informatics Assoc. | 93 |
| 2026 | Natural language generation in healthcare: A review of methods and applications
Mengxian Lyu, Jinqian Pan, Cheng Peng 0009, Sankalp Talankar, Yonghui Wu 0001 |
J. Biomed. Informatics | 7 |
| 2026 | A study of large language models for patient information extraction: Model architecture, fine-tuning strategy, and multi-task instruction tuning
Cheng Peng 0009, Mengxian Lyu, Daniel Paredes, Yaoyun Zhang, Yonghui Wu 0001 |
J. Biomed. Informatics | 6 |
| 2025 | Comparative Evaluation of Clinical Large Language Models and Machine Learning to Predict Antimicrobial Resistance in Hospital-Onset Sepsis
Scott A. Cohen, Jiang Bian 0001, Christina Boucher 0001, Yonghui Wu 0001, Mattia Prosperi |
AIME (1) | 5 |
| 2025 | MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic ApproximationabstractModel merging has emerged as an effective approach to combining multiple single-task models into a multitask model. This process typically involves computing a weighted average of the model parameters without additional training. Existing model-merging methods focus on improving average task accuracy. However, interference and conflicts between the objectives of different tasks can lead to trade-offs during the merging process. In real-world applications, a set of solutions with various trade-offs can be more informative, helping practitioners make decisions based on diverse preferences. In this paper, we introduce a novel and low-compute algorithm, Model Merging with Amortized Pareto Front (MAP). MAP efficiently identifies a Pareto set of scaling coefficients for merging multiple models, reflecting the trade-offs involved. It amortizes the substantial computational cost of evaluations needed to estimate the Pareto front by using quadratic approximation surrogate models derived from a preselected set of scaling coefficients. Experimental results on vision and natural language processing tasks demonstrate that MAP can accurately identify the Pareto front, providing practitioners with flexible solutions to balance competing task objectives. We also introduce Bayesian MAP for scenarios with a relatively low number of tasks and Nested MAP for situations with a high number of tasks, further reducing the computational cost of evaluation. Zhiqi Bu, Suyuchen Wang, Jie Fu 0001, Yonghui Wu 0001, Jiang Bian 0002, Yong Chen 0016, Yoshua Bengio |
ICLR | 7 |
| 2025 | Leveraging undecided cases in chart-reviewed phenotypes to enhance EHR-based association studies
Xinyao Jian, Dazheng Zhang, Zehao Yu 0001, Hua Xu 0001, Jiang Bian 0001, Yonghui Wu 0001, Jiayi Tong, Yong Chen 0016 |
J. Biomed. Informatics | 6 |
| 2025 | Scaling up biomedical vision-language models: Fine-tuning, instruction tuning, and multi-modal learning
Cheng Peng 0009, Kai Zhang 0039, Mengxian Lyu, Lichao Sun 0001, Yonghui Wu 0001 |
J. Biomed. Informatics | 6 |
| 2024 | Comprehensive Study on German Language Models for Clinical and Biomedical Text UnderstandingabstractRecent advances in natural language processing (NLP) can be largely attributed to the advent of pre-trained language models such as BERT and RoBERTa. While these models demonstrate remarkable performance on general datasets, they can struggle in specialized domains such as medicine, where unique domain-specific terminologies, domain-specific abbreviations, and varying document structures are common. This paper explores strategies for adapting these models to domain-specific requirements, primarily through continuous pre-training on domain-specific data. We pre-trained several German medical language models on 2.4B tokens derived from translated public English medical data and 3B tokens of German clinical data. The resulting models were evaluated on various German downstream tasks, including named entity recognition (NER), multi-label classification, and extractive question answering. Our results suggest that models augmented by clinical and translation-based pre-training typically outperform general domain models in medical contexts. We conclude that continuous pre-training has demonstrated the ability to match or even exceed the performance of clinical models trained from scratch. Furthermore, pre-training on clinical data or leveraging translated texts have proven to be reliable methods for domain adaptation in medical NLP tasks. Ahmad Idrissi-Yaghir, Amin Dada, Henning Schäfer, Kamyar Arzideh, Giulia Baldini 0001, Jan Trienes, Max Hasin, Jeanette Bewersdorff, Cynthia Sabrina Schmidt, Marie Bauer, Kaleb E. Smith, Jiang Bian 0001, Yonghui Wu 0001, Jörg Schlötterer, Torsten Zesch, Peter A. Horn, Christin Seifert, Felix Nensa, Jens Kleesiek, Christoph M. Friedrich |
LREC/COLING | 13 |
| 2024 | Generative large language models are all-purpose text analytics engines: text-to-text learning is all your needabstractOBJECTIVE: To solve major clinical natural language processing (NLP) tasks using a unified text-to-text learning architecture based on a generative large language model (LLM) via prompt tuning. METHODS: We formulated 7 key clinical NLP tasks as text-to-text learning and solved them using one unified generative clinical LLM, GatorTronGPT, developed using GPT-3 architecture and trained with up to 20 billion parameters. We adopted soft prompts (ie, trainable vectors) with frozen LLM, where the LLM parameters were not updated (ie, frozen) and only the vectors of soft prompts were updated, known as prompt tuning. We added additional soft prompts as a prefix to the input layer, which were optimized during the prompt tuning. We evaluated the proposed method using 7 clinical NLP tasks and compared them with previous task-specific solutions based on Transformer models. RESULTS AND CONCLUSION: The proposed approach achieved state-of-the-art performance for 5 out of 7 major clinical NLP tasks using one unified generative LLM. Our approach outperformed previous task-specific transformer models by ∼3% for concept extraction and 7% for relation extraction applied to social determinants of health, 3.4% for clinical concept normalization, 3.4%-10% for clinical abbreviation disambiguation, and 5.5%-9% for natural language inference. Our approach also outperformed a previously developed prompt-based machine reading comprehension (MRC) model, GatorTron-MRC, for clinical concept and relation extraction. The proposed approach can deliver the "one model for all" promise from training to deployment using a unified generative LLM. Cheng Peng 0009, Xi Yang 0015, Aokun Chen, Zehao Yu 0001, Kaleb E. Smith, Anthony B. Costa, Mona Flores, Jiang Bian 0001, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Model tuning or prompt Tuning? a study of large language models for clinical concept and relation extraction
Cheng Peng 0009, Xi Yang 0015, Kaleb E. Smith, Zehao Yu 0001, Aokun Chen, Jiang Bian 0001, Yonghui Wu 0001 |
J. Biomed. Informatics | 7 |
| 2024 | Identifying social determinants of health from clinical narratives: A study of performance, documentation ratio, and potential bias
Zehao Yu 0001, Cheng Peng 0009, Xi Yang 0015, Chong Dang, Prakash Adekkanattu, Braja Gopal Patra, Yifan Peng 0002, Jyotishman Pathak, Debbie L. Wilson, Ching-Yuan Chang, Wei-Hsuan Lo-Ciganic, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
J. Biomed. Informatics | 16 |
| 2023 | The role of health system penetration rate in estimating the prevalence of type 1 diabetes in children and adolescents using electronic health recordsabstractOBJECTIVE: Having sufficient population coverage from the electronic health records (EHRs)-connected health system is essential for building a comprehensive EHR-based diabetes surveillance system. This study aimed to establish an EHR-based type 1 diabetes (T1D) surveillance system for children and adolescents across racial and ethnic groups by identifying the minimum population coverage from EHR-connected health systems to accurately estimate T1D prevalence. MATERIALS AND METHODS: We conducted a retrospective, cross-sectional analysis involving children and adolescents <20 years old identified from the OneFlorida+ Clinical Research Network (2018-2020). T1D cases were identified using a previously validated computable phenotyping algorithm. The T1D prevalence for each ZIP Code Tabulation Area (ZCTA, 5 digits), defined as the number of T1D cases divided by the total number of residents in the corresponding ZCTA, was calculated. Population coverage for each ZCTA was measured using observed health system penetration rates (HSPR), which was calculated as the ratio of residents in the corresponding ZTCA and captured by OneFlorida+ to the overall population in the same ZCTA reported by the Census. We used a recursive partitioning algorithm to identify the minimum required observed HSPR to estimate T1D prevalence and compare our estimate with the reported T1D prevalence from the SEARCH study. RESULTS: Observed HSPRs of 55%, 55%, and 60% were identified as the minimum thresholds for the non-Hispanic White, non-Hispanic Black, and Hispanic populations. The estimated T1D prevalence for non-Hispanic White and non-Hispanic Black were 2.87 and 2.29 per 1000 youth, which are comparable to the reference study's estimation. The estimated prevalence of T1D for Hispanics (2.76 per 1000 youth) was higher than the reference study's estimation (1.48-1.64 per 1000 youth). The standardized T1D prevalence in the overall Florida population was 2.81 per 1000 youth in 2019. CONCLUSION: Our study provides a method to estimate T1D prevalence in children and adolescents using EHRs and reports the estimated HSPRs and prevalence of T1D for different race and ethnicity groups to facilitate EHR-based diabetes surveillance. Piaopiao Li, Tianchen Lyu, Khalid Alkhuzam, Eliot Spector, William T. Donahoo, Sarah Bost, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Desmond A. Schatz, Mark A. Atkinson, Michael J. Haller, Elizabeth Shenkman, Yi Guo 0005, Jiang Bian 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Clinical concept and relation extraction using prompt-based machine reading comprehensionabstractOBJECTIVE: To develop a natural language processing system that solves both clinical concept extraction and relation extraction in a unified prompt-based machine reading comprehension (MRC) architecture with good generalizability for cross-institution applications. METHODS: We formulate both clinical concept extraction and relation extraction using a unified prompt-based MRC architecture and explore state-of-the-art transformer models. We compare our MRC models with existing deep learning models for concept extraction and end-to-end relation extraction using 2 benchmark datasets developed by the 2018 National NLP Clinical Challenges (n2c2) challenge (medications and adverse drug events) and the 2022 n2c2 challenge (relations of social determinants of health [SDoH]). We also evaluate the transfer learning ability of the proposed MRC models in a cross-institution setting. We perform error analyses and examine how different prompting strategies affect the performance of MRC models. RESULTS AND CONCLUSION: The proposed MRC models achieve state-of-the-art performance for clinical concept and relation extraction on the 2 benchmark datasets, outperforming previous non-MRC transformer models. GatorTron-MRC achieves the best strict and lenient F1-scores for concept extraction, outperforming previous deep learning models on the 2 datasets by 1%-3% and 0.7%-1.3%, respectively. For end-to-end relation extraction, GatorTron-MRC and BERT-MIMIC-MRC achieve the best F1-scores, outperforming previous deep learning models by 0.9%-2.4% and 10%-11%, respectively. For cross-institution evaluation, GatorTron-MRC outperforms traditional GatorTron by 6.4% and 16% for the 2 datasets, respectively. The proposed method is better at handling nested/overlapped concepts, extracting relations, and has good portability for cross-institute applications. Our clinical MRC package is publicly available at https://github.com/uf-hobi-informatics-lab/ClinicalTransformerMRC. Cheng Peng 0009, Xi Yang 0015, Zehao Yu 0001, Jiang Bian 0001, William R. Hogan, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Contextualized medication information extraction using Transformer-based deep learning architectures
Aokun Chen, Zehao Yu 0001, Xi Yang 0015, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
J. Biomed. Informatics | 6 |
| 2021 | A Study of Social and Behavioral Determinants of Health in Lung Cancer Patients Using Transformers-based Natural Language Processing Models
Zehao Yu 0001, Xi Yang 0015, Chong Dang, Songzi Wu, Prakash Adekkanattu, Jyotishman Pathak, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
AMIA | 11 |
| 2021 | Developing an Ontology for Social and Behavioral Determinants of Health
Hansi Zhang, Xi Yang 0015, Thomas J. George, William R. Hogan, Jiang Bian 0001, Yonghui Wu 0001 |
AMIA | 6 |
| 2021 | Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets
Pengfei Yin, Yongqiu Li, Xing He 0003, Jingcheng Du, Cui Tao, Yi Guo 0005, Mattia Prosperi, Pierangelo Veltri, Xi Yang 0015, Yonghui Wu 0001, Jiang Bian 0001 |
AMIA | 11 |
| 2021 | Extracting social determinants of health from electronic health records using natural language processing: a systematic reviewabstractOBJECTIVE: Social determinants of health (SDoH) are nonclinical dispositions that impact patient health risks and clinical outcomes. Leveraging SDoH in clinical decision-making can potentially improve diagnosis, treatment planning, and patient outcomes. Despite increased interest in capturing SDoH in electronic health records (EHRs), such information is typically locked in unstructured clinical notes. Natural language processing (NLP) is the key technology to extract SDoH information from clinical text and expand its utility in patient care and research. This article presents a systematic review of the state-of-the-art NLP approaches and tools that focus on identifying and extracting SDoH data from unstructured clinical text in EHRs. MATERIALS AND METHODS: A broad literature search was conducted in February 2021 using 3 scholarly databases (ACL Anthology, PubMed, and Scopus) following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 6402 publications were initially identified, and after applying the study inclusion criteria, 82 publications were selected for the final review. RESULTS: Smoking status (n = 27), substance use (n = 21), homelessness (n = 20), and alcohol use (n = 15) are the most frequently studied SDoH categories. Homelessness (n = 7) and other less-studied SDoH (eg, education, financial problems, social isolation and support, family problems) are mostly identified using rule-based approaches. In contrast, machine learning approaches are popular for identifying smoking status (n = 13), substance use (n = 9), and alcohol use (n = 9). CONCLUSION: NLP offers significant potential to extract SDoH data from narrative clinical notes, which in turn can aid in the development of screening tools, risk prediction models, and clinical decision support systems. Braja Gopal Patra, Mohit Manoj Sharma, Veer Vekaria, Prakash Adekkanattu, Olga V. Patterson, Benjamin S. Glicksberg, Lauren A. Lepow, Euijung Ryu, Joanna M. Biernacka, Al'ona Furmanchuk, Thomas J. George, William R. Hogan, Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Myrna Weissman, Priya Wickramaratne, J. John Mann, Mark Olfson, Thomas R. Campion Jr., Mark G. Weiner, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 13 |
| 2020 | Developing and Validating a Computable Phenotype for the Identification of Transgender and Gender Nonconforming Individuals and Subgroups
Yi Guo 0005, Xing He 0003, Tianchen Lyu, Hansi Zhang, Yonghui Wu 0001, Xi Yang 0015, Zhaoyi Chen, Merry J. Markham, François Modave, Mengjun Xie, William R. Hogan, Christopher A. Harle, Elizabeth Shenkman, Jiang Bian 0001 |
AMIA | 5 |
| 2020 | Assessing the practice of data quality evaluation in a national clinical data research network through a systematic scoping review in the era of real-world dataabstractOBJECTIVE: To synthesize data quality (DQ) dimensions and assessment methods of real-world data, especially electronic health records, through a systematic scoping review and to assess the practice of DQ assessment in the national Patient-centered Clinical Research Network (PCORnet). MATERIALS AND METHODS: We started with 3 widely cited DQ literature-2 reviews from Chan et al (2010) and Weiskopf et al (2013a) and 1 DQ framework from Kahn et al (2016)-and expanded our review systematically to cover relevant articles published up to February 2020. We extracted DQ dimensions and assessment methods from these studies, mapped their relationships, and organized a synthesized summarization of existing DQ dimensions and assessment methods. We reviewed the data checks employed by the PCORnet and mapped them to the synthesized DQ dimensions and methods. RESULTS: We analyzed a total of 3 reviews, 20 DQ frameworks, and 226 DQ studies and extracted 14 DQ dimensions and 10 assessment methods. We found that completeness, concordance, and correctness/accuracy were commonly assessed. Element presence, validity check, and conformance were commonly used DQ assessment methods and were the main focuses of the PCORnet data checks. DISCUSSION: Definitions of DQ dimensions and methods were not consistent in the literature, and the DQ assessment practice was not evenly distributed (eg, usability and ease-of-use were rarely discussed). Challenges in DQ assessments, given the complex and heterogeneous nature of real-world data, exist. CONCLUSION: The practice of DQ assessment is still limited in scope. Future work is warranted to generate understandable, executable, and reusable DQ measures. Jiang Bian 0001, Tianchen Lyu, Alexander T. Loiacono, Tonatiuh Mendoza Viramontes, Gloria P. Lipori, Yi Guo 0005, Yonghui Wu 0001, Mattia Prosperi, Thomas J. George, Christopher A. Harle, Elizabeth Shenkman, William R. Hogan |
J. Am. Medical Informatics Assoc. | 7 |
| 2020 | Identifying relations of medications with adverse drug events using recurrent convolutional neural networks and gradient boostingabstractOBJECTIVE: To develop a natural language processing system that identifies relations of medications with adverse drug events from clinical narratives. This project is part of the 2018 n2c2 challenge. MATERIALS AND METHODS: We developed a novel clinical named entity recognition method based on an recurrent convolutional neural network and compared it to a recurrent neural network implemented using the long-short term memory architecture, explored methods to integrate medical knowledge as embedding layers in neural networks, and investigated 3 machine learning models, including support vector machines, random forests and gradient boosting for relation classification. The performance of our system was evaluated using annotated data and scripts provided by the 2018 n2c2 organizers. RESULTS: Our system was among the top ranked. Our best model submitted during this challenge (based on recurrent neural networks and support vector machines) achieved lenient F1 scores of 0.9287 for concept extraction (ranked third), 0.9459 for relation classification (ranked fourth), and 0.8778 for the end-to-end relation extraction (ranked second). We developed a novel named entity recognition model based on a recurrent convolutional neural network and further investigated gradient boosting for relation classification. The new methods improved the lenient F1 scores of the 3 subtasks to 0.9292, 0.9633, and 0.8880, respectively, which are comparable to the best performance reported in this challenge. CONCLUSION: This study demonstrated the feasibility of using machine learning methods to extract the relations of medications with adverse drug events from clinical narratives. Xi Yang 0015, Jiang Bian 0001, Ruogu Fang, Ragnhildur I. Bjarnadottir, William R. Hogan, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2020 | Clinical concept extraction using transformersabstractOBJECTIVE: The goal of this study is to explore transformer-based models (eg, Bidirectional Encoder Representations from Transformers [BERT]) for clinical concept extraction and develop an open-source package with pretrained clinical models to facilitate concept extraction and other downstream natural language processing (NLP) tasks in the medical domain. METHODS: We systematically explored 4 widely used transformer-based architectures, including BERT, RoBERTa, ALBERT, and ELECTRA, for extracting various types of clinical concepts using 3 public datasets from the 2010 and 2012 i2b2 challenges and the 2018 n2c2 challenge. We examined general transformer models pretrained using general English corpora as well as clinical transformer models pretrained using a clinical corpus and compared them with a long short-term memory conditional random fields (LSTM-CRFs) mode as a baseline. Furthermore, we integrated the 4 clinical transformer-based models into an open-source package. RESULTS AND CONCLUSION: The RoBERTa-MIMIC model achieved state-of-the-art performance on 3 public clinical concept extraction datasets with F1-scores of 0.8994, 0.8053, and 0.8907, respectively. Compared to the baseline LSTM-CRFs model, RoBERTa-MIMIC remarkably improved the F1-score by approximately 4% and 6% on the 2010 and 2012 i2b2 datasets. This study demonstrated the efficiency of transformer-based models for clinical concept extraction. Our methods and systems can be applied to other clinical tasks. The clinical transformer package with 4 pretrained clinical models is publicly available at https://github.com/uf-hobi-informatics-lab/ClinicalTransformerNER. We believe this package will improve current practice on clinical concept extraction and other tasks in the medical domain. Xi Yang 0015, Jiang Bian 0001, William R. Hogan, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Identifying Cancer Patients at Risk for Heart Failure Using Machine Learning Methods
Xi Yang 0015, Nida Waheed, Keith March, Jiang Bian 0001, William R. Hogan, Yonghui Wu 0001 |
AMIA | 7 |
| 2018 | Combine Factual Medical Knowledge and Distributed Word Representation to Improve Clinical Named Entity Recognition
Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Yi Guo 0005, Hua Xu 0001, William R. Hogan |
AMIA | 1 |
| 2018 | Computable Eligibility Criteria through Ontology-driven Data Access: A Case Study of Hepatitis C Virus Trials
Hansi Zhang, Zhe He 0001, Xing He 0003, Yi Guo 0005, David R. Nelson, François Modave, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Jiang Bian 0001 |
AMIA | 7 |
| 2018 | PIE: A prior knowledge guided integrated likelihood estimation method for bias reduction in association studies using electronic health records dataabstractOBJECTIVES: This study proposes a novel Prior knowledge guided Integrated likelihood Estimation (PIE) method to correct bias in estimations of associations due to misclassification of electronic health record (EHR)-derived binary phenotypes, and evaluates the performance of the proposed method by comparing it to 2 methods in common practice. METHODS: We conducted simulation studies and data analysis of real EHR-derived data on diabetes from Kaiser Permanente Washington to compare the estimation bias of associations using the proposed method, the method ignoring phenotyping errors, the maximum likelihood method with misspecified sensitivity and specificity, and the maximum likelihood method with correctly specified sensitivity and specificity (gold standard). The proposed method effectively leverages available information on phenotyping accuracy to construct a prior distribution for sensitivity and specificity, and incorporates this prior information through the integrated likelihood for bias reduction. RESULTS: Our simulation studies and real data application demonstrated that the proposed method effectively reduces the estimation bias compared to the 2 current methods. It performed almost as well as the gold standard method when the prior had highest density around true sensitivity and specificity. The analysis of EHR data from Kaiser Permanente Washington showed that the estimated associations from PIE were very close to the estimates from the gold standard method and reduced bias by 60%-100% compared to the 2 commonly used methods in current practice for EHR data. CONCLUSIONS: This study demonstrates that the proposed method can effectively reduce estimation bias caused by imperfect phenotyping in EHR-derived data by incorporating prior information through integrated likelihood. Jing Huang 0021, Rui Duan 0004, Rebecca A. Hubbard, Yonghui Wu 0001, Jason H. Moore, Hua Xu 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | CLAMP - a toolkit for efficiently building customized clinical natural language processing pipelinesabstractExisting general clinical natural language processing (NLP) systems such as MetaMap and Clinical Text Analysis and Knowledge Extraction System have been successfully applied to information extraction from clinical text. However, end users often have to customize existing systems for their individual tasks, which can require substantial NLP skills. Here we present CLAMP (Clinical Language Annotation, Modeling, and Processing), a newly developed clinical NLP toolkit that provides not only state-of-the-art NLP components, but also a user-friendly graphic user interface that can help users quickly build customized NLP pipelines for their individual applications. Our evaluation shows that the CLAMP default pipeline achieved good performance on named entity recognition and concept encoding. We also demonstrate the efficiency of the CLAMP graphic user interface in building customized, high-performance NLP pipelines with 2 use cases, extracting smoking status and lab test values. CLAMP is publicly available for research use, and we believe it is a unique asset for the clinical NLP community. Ergin Soysal, Min Jiang 0007, Yonghui Wu 0001, Serguei V. S. Pakhomov, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | A study of generalizability of recurrent neural network-based predictive models for heart failure onset risk using a large and heterogeneous EHR data set
Laila Rasmy, Yonghui Wu 0001, Ningtao Wang, W. Jim Zheng, Fei Wang 0001, Hulin Wu, Hua Xu 0001, Degui Zhi |
J. Biomed. Informatics | 2 |
| 2017 | Clinical Named Entity Recognition Using Deep Learning Models
Yonghui Wu 0001, Min Jiang 0007, Jun Xu 0007, Degui Zhi, Hua Xu 0001 |
AMIA | 1 |
| 2017 | CLAMP - A User-Centric Clinical Natural Language Processing Toolkit
Hua Xu 0001, Ergin Soysal, Min Jiang 0007, Yonghui Wu 0001 |
AMIA | 5 |
| 2017 | Detecting Contradictory and Consistent Citations in Biomedical Literature
Jun Xu 0007, Yonghui Wu 0001, Yaoyun Zhang, Qiang Wei 0002, Hua Xu 0001 |
AMIA | 2 |
| 2017 | Detecting Body Location Modifiers of Disorders in Clinical Texts via Sequence Labeling
Jun Xu 0007, Yonghui Wu 0001, Yaoyun Zhang, Hua Xu 0001 |
AMIA | 2 |
| 2017 | Evaluating Word Embeddings from Multiple Domains for Symptom Recognition in Psychiatric Notes
Yaoyun Zhang, Hee-Jin Lee, Yonghui Wu 0001, Hua Xu 0001 |
AMIA | 3 |
| 2017 | A long journey to short abbreviations: developing an open-source framework for clinical abbreviation recognition and disambiguation (CARD)abstractOBJECTIVE: The goal of this study was to develop a practical framework for recognizing and disambiguating clinical abbreviations, thereby improving current clinical natural language processing (NLP) systems' capability to handle abbreviations in clinical narratives. METHODS: We developed an open-source framework for clinical abbreviation recognition and disambiguation (CARD) that leverages our previously developed methods, including: (1) machine learning based approaches to recognize abbreviations from a clinical corpus, (2) clustering-based semiautomated methods to generate possible senses of abbreviations, and (3) profile-based word sense disambiguation methods for clinical abbreviations. We applied CARD to clinical corpora from Vanderbilt University Medical Center (VUMC) and generated 2 comprehensive sense inventories for abbreviations in discharge summaries and clinic visit notes. Furthermore, we developed a wrapper that integrates CARD with MetaMap, a widely used general clinical NLP system. RESULTS AND CONCLUSION: CARD detected 27 317 and 107 303 distinct abbreviations from discharge summaries and clinic visit notes, respectively. Two sense inventories were constructed for the 1000 most frequent abbreviations in these 2 corpora. Using the sense inventories created from discharge summaries, CARD achieved an F1 score of 0.755 for identifying and disambiguating all abbreviations in a corpus from the VUMC discharge summaries, which is superior to MetaMap and Apache's clinical Text Analysis Knowledge Extraction System (cTAKES). Using additional external corpora, we also demonstrated that the MetaMap-CARD wrapper improved MetaMap's performance in recognizing disorder entities in clinical notes. The CARD framework, 2 sense inventories, and the wrapper for MetaMap are publicly available at https://sbmi.uth.edu/ccb/resources/abbreviation.htm . We believe the CARD framework can be a valuable resource for improving abbreviation identification in clinical NLP systems. Yonghui Wu 0001, Joshua C. Denny, S. Trent Rosenbloom, Randolph A. Miller, Dario A. Giuse, Carmelo Blanquicett, Ergin Soysal, Jun Xu 0007, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | An Empirical Study for Impacts of Measurement Errors on EHR based Association Studies
Rui Duan 0004, Ming Cao 0005, Yonghui Wu 0001, Jing Huang 0021, Joshua C. Denny, Hua Xu 0001, Yong Chen 0016 |
AMIA | 3 |
| 2016 | What Can Neural Networks Learn from Unlabeled Clinical Narratives?
Yonghui Wu 0001, Jun Xu 0007, Yaoyun Zhang, Hua Xu 0001 |
AMIA | 1 |
| 2016 | Extracting genetic alteration information for personalized cancer therapy from ClinicalTrials.govabstractOBJECTIVE: Clinical trials investigating drugs that target specific genetic alterations in tumors are important for promoting personalized cancer therapy. The goal of this project is to create a knowledge base of cancer treatment trials with annotations about genetic alterations from ClinicalTrials.gov. METHODS: We developed a semi-automatic framework that combines advanced text-processing techniques with manual review to curate genetic alteration information in cancer trials. The framework consists of a document classification system to identify cancer treatment trials from ClinicalTrials.gov and an information extraction system to extract gene and alteration pairs from the Title and Eligibility Criteria sections of clinical trials. By applying the framework to trials at ClinicalTrials.gov, we created a knowledge base of cancer treatment trials with genetic alteration annotations. We then evaluated each component of the framework against manually reviewed sets of clinical trials and generated descriptive statistics of the knowledge base. RESULTS AND DISCUSSION: The automated cancer treatment trial identification system achieved a high precision of 0.9944. Together with the manual review process, it identified 20 193 cancer treatment trials from ClinicalTrials.gov. The automated gene-alteration extraction system achieved a precision of 0.8300 and a recall of 0.6803. After validation by manual review, we generated a knowledge base of 2024 cancer trials that are labeled with specific genetic alteration information. Analysis of the knowledge base revealed the trend of increased use of targeted therapy for cancer, as well as top frequent gene-alteration pairs of interest. We expect this knowledge base to be a valuable resource for physicians and patients who are seeking information about personalized cancer therapy. Jun Xu 0007, Hee-Jin Lee, Yonghui Wu 0001, Yaoyun Zhang, Liang-Chin Huang, Amber M. Johnson, Vijaykumar Holla, Ann M. Bailey, Trevor Cohen, Funda Meric-Bernstam, Elmer V. Bernstam, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Citation Sentiment Analysis in Clinical Trial Papers
Jun Xu 0007, Yaoyun Zhang, Yonghui Wu 0001, Hua Xu 0001 |
AMIA | 3 |
| 2015 | Clinical Language Annotation, Modeling, and Processing Toolkit (CLAMP) - a user-centric NLP system
Ergin Soysal, Min Jiang 0007, Yonghui Wu 0001, Hua Xu 0001 |
AMIA | 4 |
| 2015 | Recognizing Disjoint Clinical Concepts in Clinical Text Using Machine Learning-based Methods
Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001, Yaoyun Zhang, Hua Xu 0001 |
AMIA | 4 |
| 2015 | A Study of Neural Word Embeddings for Named Entity Recognition in Clinical Text
Yonghui Wu 0001, Jun Xu 0007, Min Jiang 0007, Yaoyun Zhang, Hua Xu 0001 |
AMIA | 1 |
| 2015 | Deciphering Signaling Pathway Networks to Understand the Molecular Mechanisms of Metformin ActionabstractA drug exerts its effects typically through a signal transduction cascade, which is non-linear and involves intertwined networks of multiple signaling pathways. Construction of such a signaling pathway network (SPNetwork) can enable identification of novel drug targets and deep understanding of drug action. However, it is challenging to synopsize critical components of these interwoven pathways into one network. To tackle this issue, we developed a novel computational framework, the Drug-specific Signaling Pathway Network (DSPathNet). The DSPathNet amalgamates the prior drug knowledge and drug-induced gene expression via random walk algorithms. Using the drug metformin, we illustrated this framework and obtained one metformin-specific SPNetwork containing 477 nodes and 1,366 edges. To evaluate this network, we performed the gene set enrichment analysis using the disease genes of type 2 diabetes (T2D) and cancer, one T2D genome-wide association study (GWAS) dataset, three cancer GWAS datasets, and one GWAS dataset of cancer patients with T2D on metformin. The results showed that the metformin network was significantly enriched with disease genes for both T2D and cancer, and that the network also included genes that may be associated with metformin-associated cancer survival. Furthermore, from the metformin SPNetwork and common genes to T2D and cancer, we generated a subnetwork to highlight the molecule crosstalk between T2D and cancer. The follow-up network analyses and literature mining revealed that seven genes (CDKN1A, ESR1, MAX, MYC, PPARGC1A, SP1, and STK11) and one novel MYC-centered pathway with CDKN1A, SP1, and STK11 might play important roles in metformin's antidiabetic and anticancer effects. Some results are supported by previous studies. In summary, our study 1) develops a novel framework to construct drug-specific signal transduction networks; 2) provides insights into the molecular mode of metformin; 3) serves a model for exploring signaling pathways to facilitate understanding of drug action, disease pathogenesis, and identification of drug targets. Jingchun Sun, Min Zhao 0006, Peilin Jia, Lily Wang 0001, Yonghui Wu 0001, Carissa Iverson, Yubo Zhou, Erica A. Bowton, Dan M. Roden, Joshua C. Denny, Melinda Aldrich, Hua Xu 0001, Zhongming Zhao |
PLoS Comput. Biol. | 5 |
| 2014 | Development of a Unified Computable Problem-Medication Knowledge base
Yonghui Wu 0001, Adam Wright, Hua Xu 0001, Allison B. McCoy, Dean F. Sittig |
AMIA | 1 |
| 2014 | Domain Adaptation for Semantic Role Labeling of Clinical Text
Yaoyun Zhang, Buzhou Tang, Min Jiang 0007, Yonghui Wu 0001, Hua Xu 0001 |
AMIA | 5 |
| 2013 | Building a Large Clinical Abbreviation Sense Inventory from Discharge Summaries
Yonghui Wu 0001, S. Trent Rosenbloom, Joshua C. Denny, Randolph A. Miller, Dario A. Giuse, Hua Xu 0001 |
AMIA | 1 |
| 2013 | A hybrid system for temporal information extraction from clinical textabstractOBJECTIVE: To develop a comprehensive temporal information extraction system that can identify events, temporal expressions, and their temporal relations in clinical text. This project was part of the 2012 i2b2 clinical natural language processing (NLP) challenge on temporal information extraction. MATERIALS AND METHODS: The 2012 i2b2 NLP challenge organizers manually annotated 310 clinic notes according to a defined annotation guideline: a training set of 190 notes and a test set of 120 notes. All participating systems were developed on the training set and evaluated on the test set. Our system consists of three modules: event extraction, temporal expression extraction, and temporal relation (also called Temporal Link, or 'TLink') extraction. The TLink extraction module contains three individual classifiers for TLinks: (1) between events and section times, (2) within a sentence, and (3) across different sentences. The performance of our system was evaluated using scripts provided by the i2b2 organizers. Primary measures were micro-averaged Precision, Recall, and F-measure. RESULTS: Our system was among the top ranked. It achieved F-measures of 0.8659 for temporal expression extraction (ranked fourth), 0.6278 for end-to-end TLink track (ranked first), and 0.6932 for TLink-only track (ranked first) in the challenge. We subsequently investigated different strategies for TLink extraction, and were able to marginally improve performance with an F-measure of 0.6943 for TLink-only track. Buzhou Tang, Yonghui Wu 0001, Min Jiang 0007, Yukun Chen 0001, Joshua C. Denny, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | MedEx-UIMA - An Open-Source System for Medication Information Extraction from Clinical Text
Anushi Shah, Min Jiang 0007, Yonghui Wu 0001, Joshua C. Denny, Hua Xu 0001 |
AMIA | 3 |
| 2012 | Clinical Entity Recognition Using Structural Support Vector Machines
Buzhou Tang, Yonghui Wu 0001, Min Jiang 0007, Hua Xu 0001 |
AMIA | 2 |
| 2012 | A comparative study of current clinical natural language processing systems on handling abbreviations in discharge summaries
Yonghui Wu 0001, Joshua C. Denny, S. Trent Rosenbloom, Randolph A. Miller, Dario A. Giuse, Hua Xu 0001 |
AMIA | 1 |
| 2012 | DTome: a web-based tool for drug-target interactome constructionabstractBACKGROUND: Understanding drug bioactivities is crucial for early-stage drug discovery, toxicology studies and clinical trials. Network pharmacology is a promising approach to better understand the molecular mechanisms of drug bioactivities. With a dramatic increase of rich data sources that document drugs' structural, chemical, and biological activities, it is necessary to develop an automated tool to construct a drug-target network for candidate drugs, thus facilitating the drug discovery process. RESULTS: We designed a computational workflow to construct drug-target networks from different knowledge bases including DrugBank, PharmGKB, and the PINA database. To automatically implement the workflow, we created a web-based tool called DTome (Drug-Target interactome tool), which is comprised of a database schema and a user-friendly web interface. The DTome tool utilizes web-based queries to search candidate drugs and then construct a DTome network by extracting and integrating four types of interactions. The four types are adverse drug interactions, drug-target interactions, drug-gene associations, and target-/gene-protein interactions. Additionally, we provided a detailed network analysis and visualization process to illustrate how to analyze and interpret the DTome network. The DTome tool is publicly available at http://bioinfo.mc.vanderbilt.edu/DTome. CONCLUSIONS: As demonstrated with the antipsychotic drug clozapine, the DTome tool was effective and promising for the investigation of relationships among drugs, adverse interaction drugs, drug primary targets, drug-associated genes, and proteins directly interacting with targets or genes. The resultant DTome network provides researchers with direct insights into their interest drug(s), such as the molecular mechanisms of drug actions. We believe such a tool can facilitate identification of drug targets and drug adverse interactions. Jingchun Sun, Yonghui Wu 0001, Hua Xu 0001, Zhongming Zhao |
BMC Bioinform. | 2 |
| 2012 | Large-scale prediction of adverse drug reactions using chemical, biological, and phenotypic properties of drugsabstractOBJECTIVE: Adverse drug reaction (ADR) is one of the major causes of failure in drug development. Severe ADRs that go undetected until the post-marketing phase of a drug often lead to patient morbidity. Accurate prediction of potential ADRs is required in the entire life cycle of a drug, including early stages of drug design, different phases of clinical trials, and post-marketing surveillance. METHODS: Many studies have utilized either chemical structures or molecular pathways of the drugs to predict ADRs. Here, the authors propose a machine-learning-based approach for ADR prediction by integrating the phenotypic characteristics of a drug, including indications and other known ADRs, with the drug's chemical structures and biological properties, including protein targets and pathway information. A large-scale study was conducted to predict 1385 known ADRs of 832 approved drugs, and five machine-learning algorithms for this task were compared. RESULTS: This evaluation, based on a fivefold cross-validation, showed that the support vector machine algorithm outperformed the others. Of the three types of information, phenotypic data were the most informative for ADR prediction. When biological and phenotypic features were added to the baseline chemical information, the ADR prediction model achieved significant improvements in area under the curve (from 0.9054 to 0.9524), precision (from 43.37% to 66.17%), and recall (from 49.25% to 63.06%). Most importantly, the proposed model successfully predicted the ADRs associated with withdrawal of rofecoxib and cerivastatin. CONCLUSION: The results suggest that phenotypic information on drugs is valuable for ADR prediction. Moreover, they demonstrate that different models that combine chemical, biological, or phenotypic information can be built from approved drugs, and they have the potential to detect clinically important ADRs in both preclinical and post-marketing phases. Yonghui Wu 0001, Yukun Chen 0001, Jingchun Sun, Zhongming Zhao, Xue-wen Chen 0001, Michael E. Matheny, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | A new clustering method for detecting rare senses of abbreviations in clinical notes
Hua Xu 0001, Yonghui Wu 0001, Noémie Elhadad, Peter D. Stetson, Carol Friedman |
J. Biomed. Informatics | 2 |
| 2009 | STRank: A SiteRank Algorithm Using Semantic Relevance and Time FrequencyabstractMost of the researches on Web information processing are concentrated on the Web pages and the hyperlinks among them. One of the important facts that a Web page is just one building block of the whole Website had been ignored. But the situation is gradually changed in recent years for the needs of Website reputation calculation, the high level Website structure mining etc. It causes the Website ranking become one of the hot research topics and various site ranking algorithms, such as SiteRank, AggregateRank etc., had been proposed. But most of existing Website ranking algorithm just take use of Website link graphs and the content of Websites are usually not put into consideration. It is obviously not enough for a reliable ranking of Websites. To address this issue, this paper introduces two content based features, i.e., semantic relevance and time frequency and proposes a new STRank algorithm based on these two features. We firstly conduct a series of experiments to verify the feasibility of these two factors in site ranking task. Then the semantic relevance is applied in the calculation of transition probability, and the updating frequency of sites is combined into the ranking task. Since traditional Kendall's ¿ distance and Spearman's footrule distance is not appropriate for the evaluation of site ranking, we make some modifications accordingly to evaluate Website ranking algorithms. Finally, our experiments show that the STRank algorithm outperforms existing approaches on both effectiveness and efficiency. Hongzhi Guo 0007, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001 |
SMC | 5 |
| 2008 | Genre identification of Chinese finance text using machine learning methodabstractDocument genre information is one of the most distinguishing features in information retrieval, which brings order to the search results. What the genre classification concerned is not the topic but the genre of document. In this paper, we examine the effectiveness of using machine learning techniques to solve genre classification of Chinese text with the same topic, viz. finance. Based on the likelihood ratio test, we present a new method for selecting feature terms, which can improve the performance clearly and perform better than others with up to 80% terms removal. In empirical results with SVMs classifier on the real world corpora, we find that this method can gain a better selecting effect and likelihood ratio is a reliable measure for selecting informative features. Jun Xu 0007, Xiaolong Wang 0001, Yonghui Wu 0001 |
SMC | 4 |