Xu Zuo

dblp:47/6320 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Information extraction from clinical notes: are we ready to switch to large language models?
abstract
OBJECTIVES: To assess the performance, generalizability, and computational efficiency of instruction-tuned Large Language Model Meta AI (LLaMA)-2 and LLaMA-3 models compared to bidirectional encoder representations from transformers (BERT) for clinical information extraction (IE) tasks, specifically named entity recognition (NER) and relation extraction (RE). MATERIALS AND METHODS: We developed a comprehensive annotated corpus of 1588 clinical notes from 4 data sources-UT Physicians (UTP) (1342 notes), Transcribed Medical Transcription Sample Reports and Examples (MTSamples) (146), Medical Information Mart for Intensive Care (MIMIC)-III (50), and Informatics for Integrating Biology and the Bedside (i2b2) (50), capturing 4 clinical entities (problems, tests, medications, other treatments) and 16 modifiers (eg, negation, certainty). Large Language Model Meta AI-2 and LLaMA-3 were instruction-tuned for clinical NER and RE, and their performance was benchmarked against BERT. RESULTS: Large Language Model Meta AI models consistently outperformed BERT across datasets. In data-rich settings (eg, UTP), LLaMA achieved marginal gains (approximately 1% improvement for NER and 1.5%-3.7% for RE). Under limited data conditions (eg, MTSamples, MIMIC-III) and on the unseen i2b2 dataset, LLaMA-3-70B improved F1 scores by over 7% for NER and 4% for RE. However, performance gains came with increased computational costs, with LLaMA models requiring more memory and Graphics Processing Unit (GPU) hours and running up to 28 times slower than BERT. DISCUSSION: While LLaMA models offer enhanced performance, their higher computational demands and slower throughput highlight the need to balance performance with practical resource constraints. Application-specific considerations are essential when choosing between LLMs and BERT for clinical IE. CONCLUSION: Instruction-tuned LLaMA models show promise for clinical NER and RE tasks. However, the tradeoff between improved performance and increased computational cost must be carefully evaluated. We release our Kiwi package (https://kiwi.clinicalnlp.org/) to facilitate the application of both LLaMA and BERT models in clinical IE applications.
Xu Zuo, Yujia Zhou 0003, Xueqing Peng, Jimin Huang, Vipina Kuttichi Keloth, Vincent J. Zhang, Ruey-Ling Weng, Cathy Shyr, Qingyu Chen 0001, Xiaoqian Jiang, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.2
2025 Incorporating preprints in systematic reviews: a preliminary study of a novel method for rapid evidence synthesis
abstract
OBJECTIVES: By October 1, 2024, over 450,000 COVID-19 manuscripts were published, with 10% posted as unreviewed preprints. While they accelerate knowledge sharing, their inconsistent quality complicates systematic studies. MATERIALS AND METHODS: We propose a 2-stage method to include preprints in meta-analyses. In Stage A, preprints are integrated through restriction or imputation and weighted by a confidence score reflecting their publication likelihood. In Stage B, we assess and adjust for potential publication or reporting biases. RESULTS: This preliminary study employed a 2-stage procedure validated with 2 COVID-19 treatment case studies. For hydroxychloroquine, the relative risk (RR) was 1.06 [95% CI: 0.62, 1.80], suggesting no mortality benefit over placebo. For corticosteroids, the RR was 0.88 [95% CI: 0.62, 1.27], which, while not statistically significant, aligns with evidence supporting a mortality benefit. DISCUSSION: Our research aims to bridge a significant methodological gap by providing a solution for timely evidence synthesis, particularly in the face of the overwhelming number of publications surrounding COVID-19. CONCLUSION: This preliminary study presents a method to efficiently synthesize COVID-19 research, including non-peer-reviewed preprints, to support clinical and policy decisions amidst the information surge.
Jiayi Tong, Yifei Sun 0007, Rebecca A. Hubbard, M. Elle Saine, Hua Xu 0001, Xu Zuo, Chunhua Weng, Christopher H. Schmid, Stephen E. Kimmel, Craig A. Umscheid, Adam Cuker, Yong Chen 0016
J. Am. Medical Informatics Assoc.6
2025 Improving entity recognition using ensembles of deep learning and fine-tuned large language models: A case study on adverse event extraction from VAERS and social media
Deepthi Viswaroopan, William He, Jianfu Li, Xu Zuo, Hua Xu 0001, Cui Tao
J. Biomed. Informatics5
2024 Improving large language models for clinical named entity recognition via prompt engineering
abstract
IMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications.
Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.6
2024 Relation extraction using large language models: a case study on acupuncture point locations
abstract
OBJECTIVE: In acupuncture therapy, the accurate location of acupoints is essential for its effectiveness. The advanced language understanding capabilities of large language models (LLMs) like Generative Pre-trained Transformers (GPTs) and Llama present a significant opportunity for extracting relations related to acupoint locations from textual knowledge sources. This study aims to explore the performance of LLMs in extracting acupoint-related location relations and assess the impact of fine-tuning on GPT's performance. MATERIALS AND METHODS: We utilized the World Health Organization Standard Acupuncture Point Locations in the Western Pacific Region (WHO Standard) as our corpus, which consists of descriptions of 361 acupoints. Five types of relations ("direction_of", "distance_of", "part_of", "near_acupoint", and "located_near") (n = 3174) between acupoints were annotated. Four models were compared: pre-trained GPT-3.5, fine-tuned GPT-3.5, pre-trained GPT-4, as well as pretrained Llama 3. Performance metrics included micro-average exact match precision, recall, and F1 scores. RESULTS: Our results demonstrate that fine-tuned GPT-3.5 consistently outperformed other models in F1 scores across all relation types. Overall, it achieved the highest micro-average F1 score of 0.92. DISCUSSION: The superior performance of the fine-tuned GPT-3.5 model, as shown by its F1 scores, underscores the importance of domain-specific fine-tuning in enhancing relation extraction capabilities for acupuncture-related tasks. In light of the findings from this study, it offers valuable insights into leveraging LLMs for developing clinical decision support and creating educational modules in acupuncture. CONCLUSION: This study underscores the effectiveness of LLMs like GPT and Llama in extracting relations related to acupoint locations, with implications for accurately modeling acupuncture knowledge and promoting standard implementation in acupuncture training and practice. The findings also contribute to advancing informatics applications in traditional and complementary medicine, showcasing the potential of LLMs in natural language processing.
Xueqing Peng, Jianfu Li, Xu Zuo, Su-Yuan Peng, Donghong Pei, Cui Tao, Hua Xu 0001, Na Hong
J. Am. Medical Informatics Assoc.4
2024 Confidence score: a data-driven measure for inclusive systematic reviews considering unpublished preprints
abstract
OBJECTIVES: COVID-19, since its emergence in December 2019, has globally impacted research. Over 360 000 COVID-19-related manuscripts have been published on PubMed and preprint servers like medRxiv and bioRxiv, with preprints comprising about 15% of all manuscripts. Yet, the role and impact of preprints on COVID-19 research and evidence synthesis remain uncertain. MATERIALS AND METHODS: We propose a novel data-driven method for assigning weights to individual preprints in systematic reviews and meta-analyses. This weight termed the "confidence score" is obtained using the survival cure model, also known as the survival mixture model, which takes into account the time elapsed between posting and publication of a preprint, as well as metadata such as the number of first 2-week citations, sample size, and study type. RESULTS: Using 146 preprints on COVID-19 therapeutics posted from the beginning of the pandemic through April 30, 2021, we validated the confidence scores, showing an area under the curve of 0.95 (95% CI, 0.92-0.98). Through a use case on the effectiveness of hydroxychloroquine, we demonstrated how these scores can be incorporated practically into meta-analyses to properly weigh preprints. DISCUSSION: It is important to note that our method does not aim to replace existing measures of study quality but rather serves as a supplementary measure that overcomes some limitations of current approaches. CONCLUSION: Our proposed confidence score has the potential to improve systematic reviews of evidence related to COVID-19 and other clinical conditions by providing a data-driven approach to including unpublished manuscripts.
Jiayi Tong, Chongliang Luo, Yifei Sun 0007, Rui Duan 0004, M. Elle Saine, Yifan Peng 0002, Anchita Batra, Anni Pan, Olivia Wang, Ruowang Li, Arielle Marks-Anglin, Xu Zuo, Yulun Liu 0004, Jiang Bian 0001, Stephen E. Kimmel, Keith Hamilton, Adam Cuker, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016
J. Am. Medical Informatics Assoc.15
2022 Extracting Cancer Chemotherapy and Response Information from Clinical Notes following the RECIST Definition
Xu Zuo, Natalie Gregoriou, Jianfu Li, Jeremy Warner, Yang Ping
AMIA1
2022 ClinicalLayoutLM: A Pre-trained Multi-modal Model for Understanding Scanned Document in Electronic Health Records
abstract
Scanned documents (e.g., faxes) are still widely used in clinical practice and are prevalent in Electronic Health Records (EHR). Unlocking information in scanned documents in EHRs is critical for clinical operation and research. However, it is challenging as it requires converting images to texts before applying information extraction technologies. Here we propose a multi-modal approach (ClinicalLayoutLM) that jointly models text extracted from Optical Character Recognition (OCR) and layout/image information to classify scanned clinical documents into different categories (e.g., lab reports and CT scans). Using a clinical corpus of 348, 311 scanned documents, we continually pretrained ClinicalLayoutLM based on LayoutLMv3, a multi-modal model from the open domain. For the task to classify the scanned clinical documents into 16 categories, ClinicalLayoutLM achieved an F1 score of 0.9051, which outperformed the baseline model (0.8840) that was based on text from OCR only. ClinicalLayoutLM is the first of its kind of multi-modal models for clinical documents and we believe it could benefit other clinical natural language processing (NLP) tasks such as layout analysis, information extraction and so on. The code is available at https://github.com/UTHealth-CCB/ClinicalLayoutLM and the pre-trained model is available upon request.
Qiang Wei 0002, Xu Zuo, Omer Anjum, Ryan Denlinger, Elmer V. Bernstam, Martin J. Citardi, Hua Xu 0001
IEEE Big Data2
2021 How do we share data in COVID-19 research? A systematic review of COVID-19 datasets in PubMed Central Articles
abstract
OBJECTIVE: This study aims at reviewing novel coronavirus disease (COVID-19) datasets extracted from PubMed Central articles, thus providing quantitative analysis to answer questions related to dataset contents, accessibility and citations. METHODS: We downloaded COVID-19-related full-text articles published until 31 May 2020 from PubMed Central. Dataset URL links mentioned in full-text articles were extracted, and each dataset was manually reviewed to provide information on 10 variables: (1) type of the dataset, (2) geographic region where the data were collected, (3) whether the dataset was immediately downloadable, (4) format of the dataset files, (5) where the dataset was hosted, (6) whether the dataset was updated regularly, (7) the type of license used, (8) whether the metadata were explicitly provided, (9) whether there was a PubMed Central paper describing the dataset and (10) the number of times the dataset was cited by PubMed Central articles. Descriptive statistics about these seven variables were reported for all extracted datasets. RESULTS: We found that 28.5% of 12 324 COVID-19 full-text articles in PubMed Central provided at least one dataset link. In total, 128 unique dataset links were mentioned in 12 324 COVID-19 full text articles in PubMed Central. Further analysis showed that epidemiological datasets accounted for the largest portion (53.9%) in the dataset collection, and most datasets (84.4%) were available for immediate download. GitHub was the most popular repository for hosting COVID-19 datasets. CSV, XLSX and JSON were the most popular data formats. Additionally, citation patterns of COVID-19 datasets varied depending on specific datasets. CONCLUSION: PubMed Central articles are an important source of COVID-19 datasets, but there is significant heterogeneity in the way these datasets are mentioned, shared, updated and cited.
Xu Zuo, Yong Chen 0016, Lucila Ohno-Machado, Hua Xu 0001
Briefings Bioinform.1
2020 Normalizing Clinical Document Titles to LOINC Document Ontology: an Initial Study
Xu Zuo, Jianfu Li, Bo Zhao 0001, Yujia Zhou 0003, Jon D. Duke, Karthik Natarajan, George Hripcsak, Nigam H. Shah, Juan M. Banda, Ruth M. Reeves, Hua Xu 0001
AMIA1
2008 HTS filter subsystem for future mobile communication system
Lan Fang, Xinjie Zhao 0003, Xu Zuo, TieGe Zhou, ShaoLin Yan
Sci. China Ser. F Inf. Sci.4