Hua Xu 0001

dblp:31/4114-1 · DBLP profile ↗
← Back
206ranked-venue papers
18as first author
54since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 197 · 18 first-author · 52 since 2021Artificial intelligence and machine learning · 9 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Information extraction from clinical notes: are we ready to switch to large language models?
abstract
OBJECTIVES: To assess the performance, generalizability, and computational efficiency of instruction-tuned Large Language Model Meta AI (LLaMA)-2 and LLaMA-3 models compared to bidirectional encoder representations from transformers (BERT) for clinical information extraction (IE) tasks, specifically named entity recognition (NER) and relation extraction (RE). MATERIALS AND METHODS: We developed a comprehensive annotated corpus of 1588 clinical notes from 4 data sources-UT Physicians (UTP) (1342 notes), Transcribed Medical Transcription Sample Reports and Examples (MTSamples) (146), Medical Information Mart for Intensive Care (MIMIC)-III (50), and Informatics for Integrating Biology and the Bedside (i2b2) (50), capturing 4 clinical entities (problems, tests, medications, other treatments) and 16 modifiers (eg, negation, certainty). Large Language Model Meta AI-2 and LLaMA-3 were instruction-tuned for clinical NER and RE, and their performance was benchmarked against BERT. RESULTS: Large Language Model Meta AI models consistently outperformed BERT across datasets. In data-rich settings (eg, UTP), LLaMA achieved marginal gains (approximately 1% improvement for NER and 1.5%-3.7% for RE). Under limited data conditions (eg, MTSamples, MIMIC-III) and on the unseen i2b2 dataset, LLaMA-3-70B improved F1 scores by over 7% for NER and 4% for RE. However, performance gains came with increased computational costs, with LLaMA models requiring more memory and Graphics Processing Unit (GPU) hours and running up to 28 times slower than BERT. DISCUSSION: While LLaMA models offer enhanced performance, their higher computational demands and slower throughput highlight the need to balance performance with practical resource constraints. Application-specific considerations are essential when choosing between LLMs and BERT for clinical IE. CONCLUSION: Instruction-tuned LLaMA models show promise for clinical NER and RE tasks. However, the tradeoff between improved performance and increased computational cost must be carefully evaluated. We release our Kiwi package (https://kiwi.clinicalnlp.org/) to facilitate the application of both LLaMA and BERT models in clinical IE applications.
Xu Zuo, Yujia Zhou 0003, Xueqing Peng, Jimin Huang, Vipina Kuttichi Keloth, Vincent J. Zhang, Ruey-Ling Weng, Cathy Shyr, Qingyu Chen 0001, Xiaoqian Jiang, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.13
2026 AcuKG: a comprehensive knowledge graph for medical acupuncture
abstract
BACKGROUND: Acupuncture, a key modality in traditional Chinese medicine, is gaining global recognition as a complementary therapy and a subject of increasing scientific interest. However, fragmented and unstructured acupuncture knowledge spread across diverse sources poses challenges for semantic retrieval, reasoning, and in-depth analysis. To address this gap, we developed AcuKG, a comprehensive knowledge graph that systematically organizes acupuncture-related knowledge to support sharing, discovery, and artificial intelligence-driven innovation in the field. METHODS: AcuKG integrates data from multiple sources, including online resources, guidelines, PubMed literature, ClinicalTrials.gov, and multiple ontologies (SNOMED CT, UBERON, and MeSH). We employed entity recognition, relation extraction, and ontology mapping to establish AcuKG, with human-in-the-loop to ensure data quality. Two cases evaluated AcuKG's usability: (1) how AcuKG advances acupuncture research for obesity and (2) how AcuKG enhances large language model (LLM) application on acupuncture question-answering. RESULTS: AcuKG comprises 1839 entities and 11 527 relations, mapped to 1836 standard concepts in 3 ontologies. Two use cases demonstrated AcuKG's effectiveness and potential in advancing acupuncture research and supporting LLM applications. In the obesity use case, AcuKG identified highly relevant acupoints (eg, ST25, ST36) and uncovered novel research insights based on evidence from clinical trials and literature. When applied to LLMs in answering acupuncture-related questions, integrating AcuKG with GPT-4o and LLaMA 3 significantly improved accuracy (GPT-4o: 46% → 54%, P = .03; LLaMA 3: 17% → 28%, P = .01). CONCLUSION: AcuKG is an open dataset that provides a structured and computational framework for acupuncture applications, bridging traditional practices with acupuncture research and cutting-edge LLM technologies.
Xueqing Peng, Su-Yuan Peng, Jianfu Li, Donghong Pei, Fang Li 0011, Yongqun He, Cui Tao, Hua Xu 0001, Na Hong
J. Am. Medical Informatics Assoc.13
2025 MapExplorer: New Content Generation from Low-Dimensional Visualizations
abstract
Low-dimensional visualizations, or "projection maps" are widely used in scientific research and creative industries to interpret largescale and complex datasets.These visualizations not only support the understanding of existing knowledge spaces but are often used implicitly to guide exploration into unknown areas.While such visualizations can be created through various methods such as TSNE or UMAP, there is no systematic way to leverage them for
Xingjian Zhang 0002, Ziyang Xiong, Yutong Xie 0007, Tolga Ergen, Dongsub Shim, Hua Xu 0001, Honglak Lee, Qiaozhu Mei
KDD (2)7
2025 LitFM: A Retrieval Augmented Structure-aware Foundation Model For Citation Graphs
abstract
With the advent of large language models (LLMs), managing scientific literature via LLMs has become a promising direction of research. However, existing approaches often overlook the rich structural and semantic relevance among scientific literature, limiting their ability to discern the relationships between pieces of scientific knowledge, and suffer from various types of hallucinations. These methods also focus narrowly on individual downstream tasks, limiting their applicability across use cases. We propose LitFM, the first literature foundation model designed for a wide variety of practical downstream tasks on domain-specific literature, with a focus on citation information. At its core, LitFM contains a novel graph retriever that can provide accurate and diverse recommendations for LLM to integrate graph structure information and relevant literature. LitFM also leverages a knowledge-infused LLM, fine-tuned through a well-developed instruction paradigm. It enables LitFM to extract domain-specific knowledge from literature and reason relationships among them. By integrating citation graphs during both training and inference, LitFM can generalize to unseen papers and accurately assess their relevance within existing literature. Additionally, we introduce new large-scale literature citation benchmark datasets on three academic fields, featuring sentence-level citation information and local context. Extensive experiments validate the superiority of LitFM, achieving 28.1% improvement on retrieval task in precision, and an average improvement of 7.52% over state-of-the-art across six downstream literature-related tasks.
Ali Maatouk, Ngoc Bui, Qianqian Xie, Leandros Tassiulas, Hua Xu 0001, Jie Shao 0001, Rex Ying
KDD (2)7
2025 Incorporating preprints in systematic reviews: a preliminary study of a novel method for rapid evidence synthesis
abstract
OBJECTIVES: By October 1, 2024, over 450,000 COVID-19 manuscripts were published, with 10% posted as unreviewed preprints. While they accelerate knowledge sharing, their inconsistent quality complicates systematic studies. MATERIALS AND METHODS: We propose a 2-stage method to include preprints in meta-analyses. In Stage A, preprints are integrated through restriction or imputation and weighted by a confidence score reflecting their publication likelihood. In Stage B, we assess and adjust for potential publication or reporting biases. RESULTS: This preliminary study employed a 2-stage procedure validated with 2 COVID-19 treatment case studies. For hydroxychloroquine, the relative risk (RR) was 1.06 [95% CI: 0.62, 1.80], suggesting no mortality benefit over placebo. For corticosteroids, the RR was 0.88 [95% CI: 0.62, 1.27], which, while not statistically significant, aligns with evidence supporting a mortality benefit. DISCUSSION: Our research aims to bridge a significant methodological gap by providing a solution for timely evidence synthesis, particularly in the face of the overwhelming number of publications surrounding COVID-19. CONCLUSION: This preliminary study presents a method to efficiently synthesize COVID-19 research, including non-peer-reviewed preprints, to support clinical and policy decisions amidst the information surge.
Jiayi Tong, Yifei Sun 0007, Rebecca A. Hubbard, M. Elle Saine, Hua Xu 0001, Xu Zuo, Chunhua Weng, Christopher H. Schmid, Stephen E. Kimmel, Craig A. Umscheid, Adam Cuker, Yong Chen 0016
J. Am. Medical Informatics Assoc.5
2025 CDEMapper: enhancing National Institutes of Health common data element use with large language models
abstract
OBJECTIVE: Common Data Elements (CDEs) standardize data collection and sharing across studies, enhancing data interoperability and improving research reproducibility. However, implementing CDEs presents challenges due to the broad range and variety of data elements. This study aims to develop a CDE mapping tool to bridge the gap between local data elements and National Institutes of Health (NIH) CDEs. METHODS: We propose CDEMapper, a large language model (LLM)-powered mapping tool designed to assist in mapping local data elements to NIH CDEs. CDEMapper has 3 core modules: (1) CDE indexing and embeddings. NIH CDEs were indexed and embedded to support semantic search; (2) CDE recommendations. The tool combines Elasticsearch (BM25 methods) with GPT services to recommend candidate CDEs and their permissible values; and (3) Human review. Users review and select the best match for their data elements and value sets. We evaluate the tool's recommendation accuracy and usability against manual annotations and testing. RESULTS: CDEMapper offers a publicly available, LLM-powered, and intuitive user interface that consolidates essential and advanced mapping services into a streamlined pipeline. The evaluation results demonstrated that the augmented BM25 with GPT embeddings and a GPT ranker achieved the overall best performance. The usability test also highlighted the effectiveness and efficiency of our tool. DISCUSSIONS AND CONCLUSIONS: This work opens up the potential of using LLMs to assist with CDE mapping when aligning local data elements with NIH CDEs. Additionally, this effort helps researchers better understand the gaps between their data elements and NIH CDEs while promoting CDE reusability.
Yan Wang 0015, Jimin Huang, Yujia Zhou 0003, Xubing Hao, Pritham Ram, Lingfei Qian, Qianqian Xie, Ruey-Ling Weng, Fongci Lin, Licong Cui, Xiaoqian Jiang, Hua Xu 0001, Na Hong
J. Am. Medical Informatics Assoc.15
2025 TopicForest: embedding-driven hierarchical clustering and labeling for biomedical literature
Chia-Hsuan Chang, Brian D. Ondov, Bin Choi, Xueqing Peng, Hua Xu 0001
J. Biomed. Informatics6
2025 Leveraging undecided cases in chart-reviewed phenotypes to enhance EHR-based association studies
Xinyao Jian, Dazheng Zhang, Zehao Yu 0001, Hua Xu 0001, Jiang Bian 0001, Yonghui Wu 0001, Jiayi Tong, Yong Chen 0016
J. Biomed. Informatics4
2025 Evaluating the Bias, type I error and statistical power of the prior Knowledge-Guided integrated likelihood estimation (PIE) for bias reduction in EHR based association studies
abstract
• Question: How does PIE perform in various types of real-world scenarios, in terms of estimation and hypothesis testing? • Findings: Under non-differential misclassification, PIE had a smaller bias in estimated associations compared to the naïve method, but it had similar type I error and power. • The bias reduction of PIE was superior when the prior distribution of sensitivity and specificity of the phenotyping algorithm is more accurate (i.e., close to the true operating characteristics of the phenotyping algorithm). The impact of prior is relatively small when the outcome has low prevalence and is larger when the outcome is common. • PIE can effectively reduce the bias due to phenotyping error under a wide spectrum of real-world settings. However, its main advantage is in the reduction of bias in estimation but not in hypothesis testing. Binary outcomes in electronic health records (EHR) derived using automated phenotype algorithms may suffer from phenotyping error, resulting in bias in association estimation. Huang et al. [1] proposed the Prior Knowledge-Guided Integrated Likelihood Estimation (PIE) method to mitigate the estimation bias, however, their investigation focused on point estimation without statistical inference, and the evaluation of PIE therein using simulation was a proof-of-concept with only a limited scope of scenarios. This study aims to comprehensively assess PIE’s performance including (1) how well PIE performs under a wide spectrum of operating characteristics of phenotyping algorithms under real-world scenarios (e. g., low prevalence, low sensitivity, high specificity); (2) beyond point estimation, how much variation of the PIE estimator was introduced by the prior distribution; and (3) from a hypothesis testing point of view, if PIE improves type I error and statistical power relative to the naïve method (i.e., ignoring the phenotyping error). Synthetic data and use-case analysis were utilized to evaluate PIE. The synthetic data were generated under diverse outcome prevalence, phenotyping algorithm sensitivity, and association effect sizes. Simulation studies compared PIE under different prior distributions with the naïve method, assessing bias, variance, type I error, and power. Use-case analysis compared the performance of PIE and the naïve method in estimating the association of multiple predictors with COVID-19 infection. PIE exhibited reduced bias compared to the naïve method across varied simulation settings, with comparable type I error and power. As the effect size became larger, the bias reduced by PIE was larger. PIE has superior performance when prior distributions aligned closely with true phenotyping algorithm characteristics. Impact of prior quality was minor for low-prevalence outcomes but large for common outcomes. In use-case analysis, PIE maintains a relatively accurate estimation across different scenarios, particularly outperforming the naïve approach under large effect sizes. PIE effectively mitigates estimation bias in a wide spectrum of real-world settings, particularly with accurate prior information. Its main benefit lies in bias reduction rather than hypothesis testing. The impact of the prior is small for low-prevalence outcomes.
Naimin Jing, Jiayi Tong, James Weaver, Patrick B. Ryan, Hua Xu 0001, Yong Chen 0016
J. Biomed. Informatics6
2025 BiomedRAG: A retrieval augmented large language model for biomedicine
abstract
Retrieval-augmented generation (RAG) involves a solution by retrieving knowledge from an established database to enhance the performance of large language models (LLM). , these models retrieve information at the sentence or paragraph level, potentially introducing noise and affecting the generation quality. To address these issues, we propose a novel BiomedRAG framework that directly feeds automatically retrieved chunk-based documents into the LLM. Our evaluation of BiomedRAG across four biomedical natural language processing tasks using eight datasets demonstrates that our proposed framework not only improves the performance by 9.95% on average, but also achieves state-of-the-art results, surpassing various baselines by 4.97%. BiomedRAG paves the way for more accurate and adaptable LLM applications in the biomedical domain.
Halil Kilicoglu, Hua Xu 0001, Rui Zhang 0028
J. Biomed. Informatics3
2025 A comparative study of recent large language models on generating hospital discharge summaries for lung cancer patients
Fang Li 0011, Na Hong, Manqi Li, Kirk Roberts, Licong Cui, Cui Tao, Hua Xu 0001
J. Biomed. Informatics8
2025 Improving entity recognition using ensembles of deep learning and fine-tuned large language models: A case study on adverse event extraction from VAERS and social media
Deepthi Viswaroopan, William He, Jianfu Li, Xu Zuo, Hua Xu 0001, Cui Tao
J. Biomed. Informatics6
2025 SemNovel - A new approach to detecting semantic novelty of biomedical publications using embeddings of large language models
Xueqing Peng, Yutong Xie 0007, Brian D. Ondov, Kalpana Raja, Qijia Liu, Qiaozhu Mei, Hua Xu 0001
J. Biomed. Informatics8
2025 Impacts of sample weighting on transferability of risk prediction models across EHR-Linked biobanks with different recruitment strategies
Maxwell Salvatore, Alison M. Mondul, Christopher R. Friese, David A. Hanauer, Hua Xu 0001, Celeste Leigh Pearce, Bhramar Mukherjee
J. Biomed. Informatics5
2024 Advancing entity recognition in biomedicine via instruction tuning of large language models
abstract
MOTIVATION: Large Language Models (LLMs) have the potential to revolutionize the field of Natural Language Processing, excelling not only in text generation and reasoning tasks but also in their ability for zero/few-shot learning, swiftly adapting to new tasks with minimal fine-tuning. LLMs have also demonstrated great promise in biomedical and healthcare applications. However, when it comes to Named Entity Recognition (NER), particularly within the biomedical domain, LLMs fall short of the effectiveness exhibited by fine-tuned domain-specific models. One key reason is that NER is typically conceptualized as a sequence labeling task, whereas LLMs are optimized for text generation and reasoning tasks. RESULTS: We developed an instruction-based learning paradigm that transforms biomedical NER from a sequence labeling task into a generation task. This paradigm is end-to-end and streamlines the training and evaluation process by automatically repurposing pre-existing biomedical NER datasets. We further developed BioNER-LLaMA using the proposed paradigm with LLaMA-7B as the foundational LLM. We conducted extensive testing on BioNER-LLaMA across three widely recognized biomedical NER datasets, consisting of entities related to diseases, chemicals, and genes. The results revealed that BioNER-LLaMA consistently achieved higher F1-scores ranging from 5% to 30% compared to the few-shot learning capabilities of GPT-4 on datasets with different biomedical entities. We show that a general-domain LLM can match the performance of rigorously fine-tuned PubMedBERT models and PMC-LLaMA, biomedical-specific language model. Our findings underscore the potential of our proposed paradigm in developing general-domain LLMs that can rival SOTA performances in multi-task, multi-domain scenarios in biomedical and health applications. AVAILABILITY AND IMPLEMENTATION: Datasets and other resources are available at https://github.com/BIDS-Xu-Lab/BioNER-LLaMA.
Vipina Kuttichi Keloth, Qianqian Xie, Xueqing Peng, Yan Wang 0015, Andrew Zheng, Melih Selek, Kalpana Raja, Chih-Hsuan Wei, Qiao Jin 0001, Zhiyong Lu, Qingyu Chen 0001, Hua Xu 0001
Bioinform.13
2024 AutoCriteria: a generalizable clinical trial eligibility criteria extraction system powered by large language models
abstract
OBJECTIVES: We aim to build a generalizable information extraction system leveraging large language models to extract granular eligibility criteria information for diverse diseases from free text clinical trial protocol documents. We investigate the model's capability to extract criteria entities along with contextual attributes including values, temporality, and modifiers and present the strengths and limitations of this system. MATERIALS AND METHODS: The clinical trial data were acquired from https://ClinicalTrials.gov/. We developed a system, AutoCriteria, which comprises the following modules: preprocessing, knowledge ingestion, prompt modeling based on GPT, postprocessing, and interim evaluation. The final system evaluation was performed, both quantitatively and qualitatively, on 180 manually annotated trials encompassing 9 diseases. RESULTS: AutoCriteria achieves an overall F1 score of 89.42 across all 9 diseases in extracting the criteria entities, with the highest being 95.44 for nonalcoholic steatohepatitis and the lowest of 84.10 for breast cancer. Its overall accuracy is 78.95% in identifying all contextual information across all diseases. Our thematic analysis indicated accurate logic interpretation of criteria as one of the strengths and overlooking/neglecting the main criteria as one of the weaknesses of AutoCriteria. DISCUSSION: AutoCriteria demonstrates strong potential to extract granular eligibility criteria information from trial documents without requiring manual annotations. The prompts developed for AutoCriteria generalize well across different disease areas. Our evaluation suggests that the system handles complex scenarios including multiple arm conditions and logics. CONCLUSION: AutoCriteria currently encompasses a diverse range of diseases and holds potential to extend to more in the future. This signifies a generalizable and scalable solution, poised to address the complexities of clinical trial application in real-world settings.
Surabhi Datta, Kyeryoung Lee, Hunki Paek, Frank J. Manion, Nneka Ofoegbu, Jingcheng Du, Liang-Chin Huang, Hua Xu 0001
J. Am. Medical Informatics Assoc.11
2024 Improving large language models for clinical named entity recognition via prompt engineering
abstract
IMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications.
Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.12
2024 Relation extraction using large language models: a case study on acupuncture point locations
abstract
OBJECTIVE: In acupuncture therapy, the accurate location of acupoints is essential for its effectiveness. The advanced language understanding capabilities of large language models (LLMs) like Generative Pre-trained Transformers (GPTs) and Llama present a significant opportunity for extracting relations related to acupoint locations from textual knowledge sources. This study aims to explore the performance of LLMs in extracting acupoint-related location relations and assess the impact of fine-tuning on GPT's performance. MATERIALS AND METHODS: We utilized the World Health Organization Standard Acupuncture Point Locations in the Western Pacific Region (WHO Standard) as our corpus, which consists of descriptions of 361 acupoints. Five types of relations ("direction_of", "distance_of", "part_of", "near_acupoint", and "located_near") (n = 3174) between acupoints were annotated. Four models were compared: pre-trained GPT-3.5, fine-tuned GPT-3.5, pre-trained GPT-4, as well as pretrained Llama 3. Performance metrics included micro-average exact match precision, recall, and F1 scores. RESULTS: Our results demonstrate that fine-tuned GPT-3.5 consistently outperformed other models in F1 scores across all relation types. Overall, it achieved the highest micro-average F1 score of 0.92. DISCUSSION: The superior performance of the fine-tuned GPT-3.5 model, as shown by its F1 scores, underscores the importance of domain-specific fine-tuning in enhancing relation extraction capabilities for acupuncture-related tasks. In light of the findings from this study, it offers valuable insights into leveraging LLMs for developing clinical decision support and creating educational modules in acupuncture. CONCLUSION: This study underscores the effectiveness of LLMs like GPT and Llama in extracting relations related to acupoint locations, with implications for accurately modeling acupuncture knowledge and promoting standard implementation in acupuncture training and practice. The findings also contribute to advancing informatics applications in traditional and complementary medicine, showcasing the potential of LLMs in natural language processing.
Xueqing Peng, Jianfu Li, Xu Zuo, Su-Yuan Peng, Donghong Pei, Cui Tao, Hua Xu 0001, Na Hong
J. Am. Medical Informatics Assoc.8
2024 Ensemble pretrained language models to extract biomedical knowledge from literature
abstract
OBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research.
Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001
J. Am. Medical Informatics Assoc.16
2024 Large language models for biomedicine: foundations, opportunities, challenges, and best practices
abstract
OBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications.
Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang
J. Am. Medical Informatics Assoc.3
2024 Confidence score: a data-driven measure for inclusive systematic reviews considering unpublished preprints
abstract
OBJECTIVES: COVID-19, since its emergence in December 2019, has globally impacted research. Over 360 000 COVID-19-related manuscripts have been published on PubMed and preprint servers like medRxiv and bioRxiv, with preprints comprising about 15% of all manuscripts. Yet, the role and impact of preprints on COVID-19 research and evidence synthesis remain uncertain. MATERIALS AND METHODS: We propose a novel data-driven method for assigning weights to individual preprints in systematic reviews and meta-analyses. This weight termed the "confidence score" is obtained using the survival cure model, also known as the survival mixture model, which takes into account the time elapsed between posting and publication of a preprint, as well as metadata such as the number of first 2-week citations, sample size, and study type. RESULTS: Using 146 preprints on COVID-19 therapeutics posted from the beginning of the pandemic through April 30, 2021, we validated the confidence scores, showing an area under the curve of 0.95 (95% CI, 0.92-0.98). Through a use case on the effectiveness of hydroxychloroquine, we demonstrated how these scores can be incorporated practically into meta-analyses to properly weigh preprints. DISCUSSION: It is important to note that our method does not aim to replace existing measures of study quality but rather serves as a supplementary measure that overcomes some limitations of current approaches. CONCLUSION: Our proposed confidence score has the potential to improve systematic reviews of evidence related to COVID-19 and other clinical conditions by providing a data-driven approach to including unpublished manuscripts.
Jiayi Tong, Chongliang Luo, Yifei Sun 0007, Rui Duan 0004, M. Elle Saine, Yifan Peng 0002, Anchita Batra, Anni Pan, Olivia Wang, Ruowang Li, Arielle Marks-Anglin, Xu Zuo, Yulun Liu 0004, Jiang Bian 0001, Stephen E. Kimmel, Keith Hamilton, Adam Cuker, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016
J. Am. Medical Informatics Assoc.22
2024 A span-based model for extracting overlapping PICO entities from randomized controlled trial publications
abstract
OBJECTIVES: Extracting PICO (Populations, Interventions, Comparison, and Outcomes) entities is fundamental to evidence retrieval. We present a novel method, PICOX, to extract overlapping PICO entities. MATERIALS AND METHODS: PICOX first identifies entities by assessing whether a word marks the beginning or conclusion of an entity. Then, it uses a multi-label classifier to assign one or more PICO labels to a span candidate. PICOX was evaluated using 1 of the best-performing baselines, EBM-NLP, and 3 more datasets, ie, PICO-Corpus and randomized controlled trial publications on Alzheimer's Disease (AD) or COVID-19, using entity-level precision, recall, and F1 scores. RESULTS: PICOX achieved superior precision, recall, and F1 scores across the board, with the micro F1 score improving from 45.05 to 50.87 (P ≪.01). On the PICO-Corpus, PICOX obtained higher recall and F1 scores than the baseline and improved the micro recall score from 56.66 to 67.33. On the COVID-19 dataset, PICOX also outperformed the baseline and improved the micro F1 score from 77.10 to 80.32. On the AD dataset, PICOX demonstrated comparable F1 scores with higher precision when compared to the baseline. CONCLUSION: PICOX excels in identifying overlapping entities and consistently surpasses a leading baseline across multiple datasets. Ablation studies reveal that its data augmentation strategy effectively minimizes false positives and improves precision.
Yiliang Zhou, Hua Xu 0001, Chunhua Weng, Yifan Peng 0002
J. Am. Medical Informatics Assoc.4
2024 Complementary and Integrative Health Information in the literature: its lexicon and named entity recognition
abstract
OBJECTIVE: To construct an exhaustive Complementary and Integrative Health (CIH) Lexicon (CIHLex) to help better represent the often underrepresented physical and psychological CIH approaches in standard terminologies, and to also apply state-of-the-art natural language processing (NLP) techniques to help recognize them in the biomedical literature. MATERIALS AND METHODS: We constructed the CIHLex by integrating various resources, compiling and integrating data from biomedical literature and relevant sources of knowledge. The Lexicon encompasses 724 unique concepts with 885 corresponding unique terms. We matched these concepts to the Unified Medical Language System (UMLS), and we developed and utilized BERT models comparing their efficiency in CIH named entity recognition to well-established models including MetaMap and CLAMP, as well as the large language model GPT3.5-turbo. RESULTS: Of the 724 unique concepts in CIHLex, 27.2% could be matched to at least one term in the UMLS. About 74.9% of the mapped UMLS Concept Unique Identifiers were categorized as "Therapeutic or Preventive Procedure." Among the models applied to CIH named entity recognition, BLUEBERT delivered the highest macro-average F1-score of 0.91, surpassing other models. CONCLUSION: Our CIHLex significantly augments representation of CIH approaches in biomedical literature. Demonstrating the utility of advanced NLP models, BERT notably excelled in CIH entity recognition. These results highlight promising strategies for enhancing standardization and recognition of CIH terminology in biomedical contexts.
Huixue Zhou, Robin Austin, Sheng-Chieh Lu, Greg M. Silverman, Halil Kilicoglu, Hua Xu 0001, Rui Zhang 0028
J. Am. Medical Informatics Assoc.7
2024 Call for papers: Special issue on biomedical multimodal large language models - novel approaches and applications
Jiang Bian 0001, Yifan Peng 0002, Eneida A. Mendonça, Imon Banerjee, Hua Xu 0001, Casey Overby Taylor, Anália Maria Garcia Lourenço, Alejandro Rodríguez González, Elena Tutubalina
J. Biomed. Informatics5
2024 FedFSA: Hybrid and federated framework for functional status ascertainment across institutions
Sunyang Fu, Heling Jia, Maria Vassilaki, Vipina Kuttichi Keloth, Yifang Dang, Yujia Zhou 0003, Muskan Garg, Ronald C. Petersen, Jennifer L. St. Sauver, Sungrim Moon, Liwei Wang 0010, Andrew Wen, Fang Li 0011, Hua Xu 0001, Cui Tao, Jungwei Fan 0001, Sunghwan Sohn
J. Biomed. Informatics14
2024 A scoping review of fair machine learning techniques when using real-world data
abstract
OBJECTIVE: The integration of artificial intelligence (AI) and machine learning (ML) in health care to aid clinical decisions is widespread. However, as AI and ML take important roles in health care, there are concerns about AI and ML associated fairness and bias. That is, an AI tool may have a disparate impact, with its benefits and drawbacks unevenly distributed across societal strata and subpopulations, potentially exacerbating existing health inequities. Thus, the objectives of this scoping review were to summarize existing literature and identify gaps in the topic of tackling algorithmic bias and optimizing fairness in AI/ML models using real-world data (RWD) in health care domains. METHODS: We conducted a thorough review of techniques for assessing and optimizing AI/ML model fairness in health care when using RWD in health care domains. The focus lies on appraising different quantification metrics for accessing fairness, publicly accessible datasets for ML fairness research, and bias mitigation approaches. RESULTS: We identified 11 papers that are focused on optimizing model fairness in health care applications. The current research on mitigating bias issues in RWD is limited, both in terms of disease variety and health care applications, as well as the accessibility of public datasets for ML fairness research. Existing studies often indicate positive outcomes when using pre-processing techniques to address algorithmic bias. There remain unresolved questions within the field that require further research, which includes pinpointing the root causes of bias in ML models, broadening fairness research in AI/ML with the use of RWD and exploring its implications in healthcare settings, and evaluating and addressing bias in multi-modal data. CONCLUSION: This paper provides useful reference material and insights to researchers regarding AI/ML fairness in real-world health care data and reveals the gaps in the field. Fair AI/ML in health care is a burgeoning field that requires a heightened research focus to cover diverse applications and different types of RWD.
Yu Huang 0018, Jingchuan Guo, Wei-Han William Chen, Hsin-Yueh Lin, Huilin Tang, Fei Wang 0001, Hua Xu 0001, Jiang Bian 0001
J. Biomed. Informatics7
2024 Balancing the efforts of chart review and gains in PRS prediction accuracy: An empirical study
Yuqing Lei, Adam Christian Naj, Hua Xu 0001, Ruowang Li, Yong Chen 0016
J. Biomed. Informatics3
2024 Developing deep learning-based strategies to predict the risk of hepatocellular carcinoma among patients with nonalcoholic fatty liver disease from electronic health records
abstract
OBJECTIVE: The accuracy of deep learning models for many disease prediction problems is affected by time-varying covariates, rare incidence, covariate imbalance and delayed diagnosis when using structured electronic health records data. The situation is further exasperated when predicting the risk of one disease on condition of another disease, such as the hepatocellular carcinoma risk among patients with nonalcoholic fatty liver disease due to slow, chronic progression, the scarce of data with both disease conditions and the sex bias of the diseases. The goal of this study is to investigate the extent to which the aforementioned issues influence deep learning performance, and then devised strategies to tackle these challenges. These strategies were applied to improve hepatocellular carcinoma risk prediction among patients with nonalcoholic fatty liver disease. METHODS: We evaluated two representative deep learning models in the task of predicting the occurrence of hepatocellular carcinoma in a cohort of patients with nonalcoholic fatty liver disease (n = 220,838) from a national EHR database. The disease prediction task was carefully formulated as a classification problem while taking censorship and the length of follow-up into consideration. RESULTS: We developed a novel backward masking scheme to deal with the issue of delayed diagnosis which is very common in EHR data analysis and evaluate how the length of longitudinal information after the index date affects disease prediction. We observed that modeling time-varying covariates improved the performance of the algorithms and transfer learning mitigated reduced performance caused by the lack of data. In addition, covariate imbalance, such as sex bias in data impaired performance. Deep learning models trained on one sex and evaluated in the other sex showed reduced performance, indicating the importance of assessing covariate imbalance while preparing data for model training. CONCLUSIONS: The strategies developed in this work can significantly improve the performance of hepatocellular carcinoma risk prediction among patients with nonalcoholic fatty liver disease. Furthermore, our novel strategies can be generalized to apply to other disease risk predictions using structured electronic health records, especially for disease risks on condition of another disease.
Yujia Zhou 0003, Ruoxing Li, Kenneth D. Chavin, Hua Xu 0001, Liang Li 0026, David J. H. Shih, W. Jim Zheng
J. Biomed. Informatics6
2024 Artificial intelligence-powered pharmacovigilance: A review of machine and deep learning in clinical text-based adverse drug event detection for benchmark datasets
Zehan Li, Zenan Sun, Fang Li 0011, Susan H. Fenton, Hua Xu 0001, Cui Tao
J. Biomed. Informatics7
2024 Improving tabular data extraction in scanned laboratory reports using deep learning models
Qiang Wei 0002, Xinghan Chen, Jianfu Li, Cui Tao, Hua Xu 0001
J. Biomed. Informatics6
2024 Leveraging error-prone algorithm-derived phenotypes: Enhancing association studies for risk factors in EHR data
Jiayi Tong, Jessica Chubak, Thomas Lumley, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016
J. Biomed. Informatics6
2024 Augmenting biomedical named entity recognition with general-domain resources
abstract
OBJECTIVE: Training a neural network-based biomedical named entity recognition (BioNER) model usually requires extensive and costly human annotations. While several studies have employed multi-task learning with multiple BioNER datasets to reduce human effort, this approach does not consistently yield performance improvements and may introduce label ambiguity in different biomedical corpora. We aim to tackle those challenges through transfer learning from easily accessible resources with fewer concept overlaps with biomedical datasets. METHODS: We proposed GERBERA, a simple-yet-effective method that utilized general-domain NER datasets for training. We performed multi-task learning to train a pre-trained biomedical language model with both the target BioNER dataset and the general-domain dataset. Subsequently, we fine-tuned the models specifically for the BioNER dataset. RESULTS: We systematically evaluated GERBERA on five datasets of eight entity types, collectively consisting of 81,410 instances. Despite using fewer biomedical resources, our models demonstrated superior performance compared to baseline models trained with additional BioNER datasets. Specifically, our models consistently outperformed the baseline models in six out of eight entity types, achieving an average improvement of 0.9% over the best baseline performance across eight entities. Our method was especially effective in amplifying performance on BioNER datasets characterized by limited data, with a 4.7% improvement in F1 scores on the JNLPBA-RNA dataset. CONCLUSION: This study introduces a new training method that leverages cost-effective general-domain NER datasets to augment BioNER models. This approach significantly improves BioNER model performance, making it a valuable asset for scenarios with scarce or costly biomedical datasets. We make data, codes, and models publicly available via https://github.com/qingyu-qc/bioner_gerbera.
Hyunjae Kim, Chih-Hsuan Wei, Jaewoo Kang, Zhiyong Lu, Hua Xu 0001, Qingyu Chen 0001
J. Biomed. Informatics7
2023 Towards precise PICO extraction from abstracts of randomized controlled trials using a section-specific learning approach
abstract
MOTIVATION: Automated extraction of participants, intervention, comparison/control, and outcome (PICO) from the randomized controlled trial (RCT) abstracts is important for evidence synthesis. Previous studies have demonstrated the feasibility of applying natural language processing (NLP) for PICO extraction. However, the performance is not optimal due to the complexity of PICO information in RCT abstracts and the challenges involved in their annotation. RESULTS: We propose a two-step NLP pipeline to extract PICO elements from RCT abstracts: (i) sentence classification using a prompt-based learning model and (ii) PICO extraction using a named entity recognition (NER) model. First, the sentences in abstracts were categorized into four sections namely background, methods, results, and conclusions. Next, the NER model was applied to extract the PICO elements from the sentences within the title and methods sections that include >96% of PICO information. We evaluated our proposed NLP pipeline on three datasets, the EBM-NLPmoddataset, a randomly selected and reannotated dataset of 500 RCT abstracts from the EBM-NLP corpus, a dataset of 150 COVID-19 RCT abstracts, and a dataset of 150 Alzheimer's disease (AD) RCT abstracts. The end-to-end evaluation reveals that our proposed approach achieved an overall micro F1 score of 0.833 on the EBM-NLPmod dataset, 0.928 on the COVID-19 dataset, and 0.899 on the AD dataset when measured at the token-level and an overall micro F1 score of 0.712 on EBM-NLPmod dataset, 0.850 on the COVID-19 dataset, and 0.805 on the AD dataset when measured at the entity-level. AVAILABILITY: Our codes and datasets are publicly available at https://github.com/BIDS-Xu-Lab/section_specific_annotation_of_PICO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vipina Kuttichi Keloth, Kalpana Raja, Yong Chen 0016, Hua Xu 0001
Bioinform.5
2023 Systematic design and data-driven evaluation of social determinants of health ontology (SDoHO)
abstract
OBJECTIVE: Social determinants of health (SDoH) play critical roles in health outcomes and well-being. Understanding the interplay of SDoH and health outcomes is critical to reducing healthcare inequalities and transforming a "sick care" system into a "health-promoting" system. To address the SDOH terminology gap and better embed relevant elements in advanced biomedical informatics, we propose an SDoH ontology (SDoHO), which represents fundamental SDoH factors and their relationships in a standardized and measurable way. MATERIAL AND METHODS: Drawing on the content of existing ontologies relevant to certain aspects of SDoH, we used a top-down approach to formally model classes, relationships, and constraints based on multiple SDoH-related resources. Expert review and coverage evaluation, using a bottom-up approach employing clinical notes data and a national survey, were performed. RESULTS: We constructed the SDoHO with 708 classes, 106 object properties, and 20 data properties, with 1,561 logical axioms and 976 declaration axioms in the current version. Three experts achieved 0.967 agreement in the semantic evaluation of the ontology. A comparison between the coverage of the ontology and SDOH concepts in 2 sets of clinical notes and a national survey instrument also showed satisfactory results. DISCUSSION: SDoHO could potentially play an essential role in providing a foundation for a comprehensive understanding of the associations between SDoH and health outcomes and paving the way for health equity across populations. CONCLUSION: SDoHO has well-designed hierarchies, practical objective properties, and versatile functionalities, and the comprehensive semantic and coverage evaluation achieved promising performance compared to the existing ontologies relevant to SDoH.
Yifang Dang, Fang Li 0011, Xinyue Hu 0002, Vipina Kuttichi Keloth, Sunyang Fu, Muhammad Amith, J. Wilfred Fan, Jingcheng Du, Evan Yu, Xiaoqian Jiang, Hua Xu 0001, Cui Tao
J. Am. Medical Informatics Assoc.13
2023 Blockchain-enabled immutable, distributed, and highly available clinical research activity logging system for federated COVID-19 data analysis from multiple institutions
abstract
OBJECTIVE: We aimed to develop a distributed, immutable, and highly available cross-cloud blockchain system to facilitate federated data analysis activities among multiple institutions. MATERIALS AND METHODS: We preprocessed 9166 COVID-19 Structured Query Language (SQL) code, summary statistics, and user activity logs, from the GitHub repository of the Reliable Response Data Discovery for COVID-19 (R2D2) Consortium. The repository collected local summary statistics from participating institutions and aggregated the global result to a COVID-19-related clinical query, previously posted by clinicians on a website. We developed both on-chain and off-chain components to store/query these activity logs and their associated queries/results on a blockchain for immutability, transparency, and high availability of research communication. We measured run-time efficiency of contract deployment, network transactions, and confirmed the accuracy of recorded logs compared to a centralized baseline solution. RESULTS: The smart contract deployment took 4.5 s on an average. The time to record an activity log on blockchain was slightly over 2 s, versus 5-9 s for baseline. For querying, each query took on an average less than 0.4 s on blockchain, versus around 2.1 s for baseline. DISCUSSION: The low deployment, recording, and querying times confirm the feasibility of our cross-cloud, blockchain-based federated data analysis system. We have yet to evaluate the system on a larger network with multiple nodes per cloud, to consider how to accommodate a surge in activities, and to investigate methods to lower querying time as the blockchain grows. CONCLUSION: Blockchain technology can be used to support federated data analysis among multiple institutions.
Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim 0001, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Christian Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo Koola, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nicholas R. Anderson 0001, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu 0001, Yong K. Choi, Kai Zheng 0002, Yujia Zhou 0003, Rachel A Zucker
J. Am. Medical Informatics Assoc.55
2023 An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)
abstract
Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts.
Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001
J. Am. Medical Informatics Assoc.27
2023 Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001
J. Biomed. Informatics27
2023 A hierarchical strategy to minimize privacy risk when linking "De-identified" data in biomedical research consortia
Lucila Ohno-Machado, Xiaoqian Jiang, Tsung-Ting Kuo, Shiqiang Tao, Pritham Ram, Guo-Qiang Zhang 0001, Hua Xu 0001
J. Biomed. Informatics8
2022 Section-specific Annotation of PICO for Medical Evidence Extraction: Applications to Articles of Randomized Controlled Trials for Alzheimer's Disease and COVID-19
Vipina Kuttichi Keloth, Hua Xu 0001
AMIA3
2022 Existing and emerging privacy challenges and solutions for federated data coordination
Tsung-Ting Kuo, Xiaoqian Jiang, Hua Xu 0001, Li Xiong 0001, Lucila Ohno-Machado
AMIA3
2022 Fast Textual Corpus Retrieval from Electronic Health Records using PDF Parser and TxtSplit Tools
Qiang Wei 0002, Yujia Zhou 0003, Hua Xu 0001
AMIA4
2022 Automated Identification of Missing IS-A Relations in the Human Phenotype Ontology
Maryamsadat Mohtashamian, Rashmie Abeysinghe, Xubing Hao, Hua Xu 0001, Licong Cui
AMIA5
2022 ClinicalLayoutLM: A Pre-trained Multi-modal Model for Understanding Scanned Document in Electronic Health Records
abstract
Scanned documents (e.g., faxes) are still widely used in clinical practice and are prevalent in Electronic Health Records (EHR). Unlocking information in scanned documents in EHRs is critical for clinical operation and research. However, it is challenging as it requires converting images to texts before applying information extraction technologies. Here we propose a multi-modal approach (ClinicalLayoutLM) that jointly models text extracted from Optical Character Recognition (OCR) and layout/image information to classify scanned clinical documents into different categories (e.g., lab reports and CT scans). Using a clinical corpus of 348, 311 scanned documents, we continually pretrained ClinicalLayoutLM based on LayoutLMv3, a multi-modal model from the open domain. For the task to classify the scanned clinical documents into 16 categories, ClinicalLayoutLM achieved an F1 score of 0.9051, which outperformed the baseline model (0.8840) that was based on text from OCR only. ClinicalLayoutLM is the first of its kind of multi-modal models for clinical documents and we believe it could benefit other clinical natural language processing (NLP) tasks such as layout analysis, information extraction and so on. The code is available at https://github.com/UTHealth-CCB/ClinicalLayoutLM and the pre-trained model is available upon request.
Qiang Wei 0002, Xu Zuo, Omer Anjum, Ryan Denlinger, Elmer V. Bernstam, Martin J. Citardi, Hua Xu 0001
IEEE Big Data8
2022 Combining human and machine intelligence for clinical trial eligibility querying
abstract
OBJECTIVE: To combine machine efficiency and human intelligence for converting complex clinical trial eligibility criteria text into cohort queries. MATERIALS AND METHODS: Criteria2Query (C2Q) 2.0 was developed to enable real-time user intervention for criteria selection and simplification, parsing error correction, and concept mapping. The accuracy, precision, recall, and F1 score of enhanced modules for negation scope detection, temporal and value normalization were evaluated using a previously curated gold standard, the annotated eligibility criteria of 1010 COVID-19 clinical trials. The usability and usefulness were evaluated by 10 research coordinators in a task-oriented usability evaluation using 5 Alzheimer's disease trials. Data were collected by user interaction logging, a demographic questionnaire, the Health Information Technology Usability Evaluation Scale (Health-ITUES), and a feature-specific questionnaire. RESULTS: The accuracies of negation scope detection, temporal and value normalization were 0.924, 0.916, and 0.966, respectively. C2Q 2.0 achieved a moderate usability score (3.84 out of 5) and a high learnability score (4.54 out of 5). On average, 9.9 modifications were made for a clinical study. Experienced researchers made more modifications than novice researchers. The most frequent modification was deletion (5.35 per study). Furthermore, the evaluators favored cohort queries resulting from modifications (score 4.1 out of 5) and the user engagement features (score 4.3 out of 5). DISCUSSION AND CONCLUSION: Features to engage domain experts and to overcome the limitations in automated machine output are shown to be useful and user-friendly. We concluded that human-computer collaboration is key to improving the adoption and user-friendliness of natural language processing.
Yilu Fang, Betina Ross S. Idnay, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Karen Marder, Hua Xu 0001, Rebecca Schnall, Chunhua Weng
J. Am. Medical Informatics Assoc.7
2022 Discovering novel drug-supplement interactions using SuppKG generated from the biomedical literature
abstract
OBJECTIVE: Develop a novel methodology to create a comprehensive knowledge graph (SuppKG) to represent a domain with limited coverage in the Unified Medical Language System (UMLS), specifically dietary supplement (DS) information for discovering drug-supplement interactions (DSI), by leveraging biomedical natural language processing (NLP) technologies and a DS domain terminology. MATERIALS AND METHODS: We created SemRepDS (an extension of an NLP tool, SemRep), capable of extracting semantic relations from abstracts by leveraging a DS-specific terminology (iDISK) containing 28,884 DS terms not found in the UMLS. PubMed abstracts were processed using SemRepDS to generate semantic relations, which were then filtered using a PubMedBERT model to remove incorrect relations before generating SuppKG. Two discovery pathways were applied to SuppKG to identify potential DSIs, which are then compared with an existing DSI database and also evaluated by medical professionals for mechanistic plausibility. RESULTS: SemRepDS returned 158.5% more DS entities and 206.9% more DS relations than SemRep. The fine-tuned PubMedBERT model (significantly outperformed other machine learning and BERT models) obtained an F1 score of 0.8605 and removed 43.86% of semantic relations, improving the precision of the relations by 26.4% over pre-filtering. SuppKG consists of 56,635 nodes and 595,222 directed edges with 2,928 DS-specific nodes and 164,738 edges. Manual review of findings identified 182 of 250 (72.8%) proposed DS-Gene-Drug and 77 of 100 (77%) proposed DS-Gene1-Function-Gene2-Drug pathways to be mechanistically plausible. DISCUSSION: With added DS terminology to the UMLS, SemRepDS has the capability to find more DS-specific semantic relationships from PubMed than SemRep. The utility of the resulting SuppKG was demonstrated using discovery patterns to find novel DSIs. CONCLUSION: For the domain with limited coverage in the traditional terminology (e.g., UMLS), we demonstrated an approach to leverage domain terminology and improve existing NLP tools to generate a more comprehensive knowledge graph for the downstream task. Even this study focuses on DSI, the method may be adapted to other domains.
Dalton Schutte, Jake Vasilakes, Anusha Bompelli, Marcelo Fiszman, Hua Xu 0001, Halil Kilicoglu, Jeffrey R. Bishop, Terrence Adam, Rui Zhang 0028
J. Biomed. Informatics6
2022 Novel informatics approaches to COVID-19 Research: From methods to applications
Hua Xu 0001, David L. Buckeridge, Fei Wang 0001, Peter Tarczy-Hornoch
J. Biomed. Informatics1
2021 Generalizable Gated Recurrent Neural Network based model to predict COVID-19 patient outcomes on admission
Laila Rasmy, Bijun S. Kannadath, Masayuki Nigo, Ziqian Xie, Bingyu Mao, Khush A. Patel, Yujia Zhou 0003, Hua Xu 0001, Degui Zhi
AMIA9
2021 CovRNN: predicting outcomes of COVID-19 patients on admission using their electronic health records with minimal data processing
Laila Rasmy, Masayuki Nigo, Bijun S. Kannadath, Ziqian Xie, Bingyu Mao, Khush A. Patel, Wanheng Zhang, Yujia Zhou 0003, Angela Ross, Hua Xu 0001, Degui Zhi
AMIA10
2021 How do we share data in COVID-19 research? A systematic review of COVID-19 datasets in PubMed Central Articles
abstract
OBJECTIVE: This study aims at reviewing novel coronavirus disease (COVID-19) datasets extracted from PubMed Central articles, thus providing quantitative analysis to answer questions related to dataset contents, accessibility and citations. METHODS: We downloaded COVID-19-related full-text articles published until 31 May 2020 from PubMed Central. Dataset URL links mentioned in full-text articles were extracted, and each dataset was manually reviewed to provide information on 10 variables: (1) type of the dataset, (2) geographic region where the data were collected, (3) whether the dataset was immediately downloadable, (4) format of the dataset files, (5) where the dataset was hosted, (6) whether the dataset was updated regularly, (7) the type of license used, (8) whether the metadata were explicitly provided, (9) whether there was a PubMed Central paper describing the dataset and (10) the number of times the dataset was cited by PubMed Central articles. Descriptive statistics about these seven variables were reported for all extracted datasets. RESULTS: We found that 28.5% of 12 324 COVID-19 full-text articles in PubMed Central provided at least one dataset link. In total, 128 unique dataset links were mentioned in 12 324 COVID-19 full text articles in PubMed Central. Further analysis showed that epidemiological datasets accounted for the largest portion (53.9%) in the dataset collection, and most datasets (84.4%) were available for immediate download. GitHub was the most popular repository for hosting COVID-19 datasets. CSV, XLSX and JSON were the most popular data formats. Additionally, citation patterns of COVID-19 datasets varied depending on specific datasets. CONCLUSION: PubMed Central articles are an important source of COVID-19 datasets, but there is significant heterogeneity in the way these datasets are mentioned, shared, updated and cited.
Xu Zuo, Yong Chen 0016, Lucila Ohno-Machado, Hua Xu 0001
Briefings Bioinform.4
2021 The application of artificial intelligence and data integration in COVID-19 studies: a scoping review
abstract
OBJECTIVE: To summarize how artificial intelligence (AI) is being applied in COVID-19 research and determine whether these AI applications integrated heterogenous data from different sources for modeling. MATERIALS AND METHODS: We searched 2 major COVID-19 literature databases, the National Institutes of Health's LitCovid and the World Health Organization's COVID-19 database on March 9, 2021. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guideline, 2 reviewers independently reviewed all the articles in 2 rounds of screening. RESULTS: In the 794 studies included in the final qualitative analysis, we identified 7 key COVID-19 research areas in which AI was applied, including disease forecasting, medical imaging-based diagnosis and prognosis, early detection and prognosis (non-imaging), drug repurposing and early drug discovery, social media data analysis, genomic, transcriptomic, and proteomic data analysis, and other COVID-19 research topics. We also found that there was a lack of heterogenous data integration in these AI applications. DISCUSSION: Risk factors relevant to COVID-19 outcomes exist in heterogeneous data sources, including electronic health records, surveillance systems, sociodemographic datasets, and many more. However, most AI applications in COVID-19 research adopted a single-sourced approach that could omit important risk factors and thus lead to biased algorithms. Integrating heterogeneous data for modeling will help realize the full potential of AI algorithms, improve precision, and reduce bias. CONCLUSION: There is a lack of data integration in the AI applications in COVID-19 research and a need for a multilevel AI framework that supports the analysis of heterogeneous data from different sources.
Yi Guo 0005, Yahan Zhang, Tianchen Lyu, Mattia Prosperi, Fei Wang 0001, Hua Xu 0001, Jiang Bian 0001
J. Am. Medical Informatics Assoc.6
2021 Extracting postmarketing adverse events from safety reports in the vaccine adverse event reporting system (VAERS) using deep learning
abstract
OBJECTIVE: Automated analysis of vaccine postmarketing surveillance narrative reports is important to understand the progression of rare but severe vaccine adverse events (AEs). This study implemented and evaluated state-of-the-art deep learning algorithms for named entity recognition to extract nervous system disorder-related events from vaccine safety reports. MATERIALS AND METHODS: We collected Guillain-Barré syndrome (GBS) related influenza vaccine safety reports from the Vaccine Adverse Event Reporting System (VAERS) from 1990 to 2016. VAERS reports were selected and manually annotated with major entities related to nervous system disorders, including, investigation, nervous_AE, other_AE, procedure, social_circumstance, and temporal_expression. A variety of conventional machine learning and deep learning algorithms were then evaluated for the extraction of the above entities. We further pretrained domain-specific BERT (Bidirectional Encoder Representations from Transformers) using VAERS reports (VAERS BERT) and compared its performance with existing models. RESULTS AND CONCLUSIONS: Ninety-one VAERS reports were annotated, resulting in 2512 entities. The corpus was made publicly available to promote community efforts on vaccine AEs identification. Deep learning-based methods (eg, bi-long short-term memory and BERT models) outperformed conventional machine learning-based methods (ie, conditional random fields with extensive features). The BioBERT large model achieved the highest exact match F-1 scores on nervous_AE, procedure, social_circumstance, and temporal_expression; while VAERS BERT large models achieved the highest exact match F-1 scores on investigation and other_AE. An ensemble of these 2 models achieved the highest exact match microaveraged F-1 score at 0.6802 and the second highest lenient match microaveraged F-1 score at 0.8078 among peer models.
Jingcheng Du, Yang Xiang 0003, Madhuri Sankaranarayanapillai, Yuqi Si, Huy Anh Pham, Hua Xu 0001, Yong Chen 0016, Cui Tao
J. Am. Medical Informatics Assoc.8
2021 Privacy-protecting, reliable response data discovery using COVID-19 patient observations
abstract
OBJECTIVE: To utilize, in an individual and institutional privacy-preserving manner, electronic health record (EHR) data from 202 hospitals by analyzing answers to COVID-19-related questions and posting these answers online. MATERIALS AND METHODS: We developed a distributed, federated network of 12 health systems that harmonized their EHRs and submitted aggregate answers to consortia questions posted at https://www.covid19questions.org. Our consortium developed processes and implemented distributed algorithms to produce answers to a variety of questions. We were able to generate counts, descriptive statistics, and build a multivariate, iterative regression model without centralizing individual-level data. RESULTS: Our public website contains answers to various clinical questions, a web form for users to ask questions in natural language, and a list of items that are currently pending responses. The results show, for example, that patients who were taking angiotensin-converting enzyme inhibitors and angiotensin II receptor blockers, within the year before admission, had lower unadjusted in-hospital mortality rates. We also showed that, when adjusted for, age, sex, and ethnicity were not significantly associated with mortality. We demonstrated that it is possible to answer questions about COVID-19 using EHR data from systems that have different policies and must follow various regulations, without moving data out of their health systems. DISCUSSION AND CONCLUSIONS: We present an alternative or a complement to centralized COVID-19 registries of EHR data. We can use multivariate distributed logistic regression on observations recorded in the process of care to generate results without transferring individual-level data outside the health systems.
Jihoon Kim 0001, Larissa Neumann, Paulina Paul, Michele E. Day, Michael Aratow, Douglas S. Bell, Jason N. Doctor, Christian Hinske, Xiaoqian Jiang, Katherine K. Kim, Michael E. Matheny, Daniella Meeker, Mark J. Pletcher, Lisa M. Schilling, Spencer L. SooHoo, Hua Xu 0001, Kai Zheng 0002, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.16
2021 Are synthetic clinical notes useful for real natural language processing tasks: A case study on clinical entity recognition
abstract
OBJECTIVE: : Developing clinical natural language processing systems often requires access to many clinical documents, which are not widely available to the public due to privacy and security concerns. To address this challenge, we propose to develop methods to generate synthetic clinical notes and evaluate their utility in real clinical natural language processing tasks. MATERIALS AND METHODS: : We implemented 4 state-of-the-art text generation models, namely CharRNN, SegGAN, GPT-2, and CTRL, to generate clinical text for the History and Present Illness section. We then manually annotated clinical entities for randomly selected 500 History and Present Illness notes generated from the best-performing algorithm. To compare the utility of natural and synthetic corpora, we trained named entity recognition (NER) models from all 3 corpora and evaluated their performance on 2 independent natural corpora. RESULTS: : Our evaluation shows GPT-2 achieved the best BLEU (bilingual evaluation understudy) score (with a BLEU-2 of 0.92). NER models trained on synthetic corpus generated by GPT-2 showed slightly better performance on 2 independent corpora: strict F1 scores of 0.709 and 0.748, respectively, when compared with the NER models trained on natural corpus (F1 scores of 0.706 and 0.737, respectively), indicating the good utility of synthetic corpora in clinical NER model development. In addition, we also demonstrated that an augmented method that combines both natural and synthetic corpora achieved better performance than that uses the natural corpus only. CONCLUSIONS: : Recent advances in text generation have made it possible to generate synthetic clinical notes that could be useful for training NER models for information extraction from natural clinical notes, thus lowering the privacy concern and increasing data availability. Further investigation is needed to apply this technology to practice.
Jianfu Li, Yujia Zhou 0003, Xiaoqian Jiang, Karthik Natarajan, Serguei V. S. Pakhomov, Hua Xu 0001
J. Am. Medical Informatics Assoc.7
2021 COVID-19 SignSym: a fast adaptation of a general clinical NLP tool to identify and normalize COVID-19 signs and symptoms to OMOP common data model
abstract
The COVID-19 pandemic swept across the world rapidly, infecting millions of people. An efficient tool that can accurately recognize important clinical concepts of COVID-19 from free text in electronic health records (EHRs) will be valuable to accelerate COVID-19 clinical research. To this end, this study aims at adapting the existing CLAMP natural language processing tool to quickly build COVID-19 SignSym, which can extract COVID-19 signs/symptoms and their 8 attributes (body location, severity, temporal expression, subject, condition, uncertainty, negation, and course) from clinical text. The extracted information is also mapped to standard concepts in the Observational Medical Outcomes Partnership common data model. A hybrid approach of combining deep learning-based models, curated lexicons, and pattern-based rules was applied to quickly build the COVID-19 SignSym from CLAMP, with optimized performance. Our extensive evaluation using 3 external sites with clinical notes of COVID-19 patients, as well as the online medical dialogues of COVID-19, shows COVID-19 SignSym can achieve high performance across data sources. The workflow used for this study can be generalized to other use cases, where existing clinical natural language processing tools need to be customized for specific information needs within a short time. COVID-19 SignSym is freely accessible to the research community as a downloadable package (https://clamp.uth.edu/covid/nlp.php) and has been used by 16 healthcare organizations to support clinical research of COVID-19.
Noor Abu-El-Rub, Josh Gray, Huy Anh Pham, Yujia Zhou 0003, Frank J. Manion, Xing Song, Hua Xu 0001, Masoud Rouhizadeh, Yaoyun Zhang
J. Am. Medical Informatics Assoc.9
2020 The Open Health Natural Language Processing Collaboratory
Xiaoqian Jiang, Serguei V. S. Pakhomov, Chunhua Weng, Hua Xu 0001
AMIA5
2020 LANN: An integrated online annotation tool for information extraction
Yaoyun Zhang, Hua Xu 0001
AMIA3
2020 Normalizing Clinical Document Titles to LOINC Document Ontology: an Initial Study
Xu Zuo, Jianfu Li, Bo Zhao 0001, Yujia Zhou 0003, Jon D. Duke, Karthik Natarajan, George Hripcsak, Nigam H. Shah, Juan M. Banda, Ruth M. Reeves, Hua Xu 0001
AMIA12
2020 Opioid2FHIR: A system for extracting FHIR-compatible opioid prescriptions from clinical text
abstract
Background: The opioid crisis is a national public health emergency in US. Especially, prescription opioids contributed significantly to drug overdose deaths. To improve the surveillance of prescription opioid overdose, it is critical to accurately collect prescription opioid information and calculate morphine milligram equivalents (MMEs). However, plenty of detailed information is only contained in the free text Sig component of electronic health record (EHR) prescriptions and need to be extracted first. Moreover, it is also indispensable to normalize opioid information extracted from multiple heath care facilities to clinical data standards such as Fast Healthcare Interoperability Resources (FHIR) for efficient clinical decision support. However, few efforts are spent in this direction at present. Methods: In this study, we designed and implemented a system that can automatically extract opioid information from free text in Sig and map them to FHIR. The system, named as Opioid2FHIR, applied multiple natural language processing (NLP) techniques for opioid information extraction and normalization. In order to reduce manual efforts, a general-purpose medication IE model was first leveraged. Based on 1, 000 opioid prescription records randomly selected from MIMICIII, post-processing rules were designed to adapt the IE model to opioid medications. Concept normalization models were also built to transform and map the extracted medication elements to fine-granular standard concepts in FHIR. The system was evaluated on another 1, 000 opioid prescription records in MIMICIII. Results: Opioid2FHIR obtained an F-measure of 0.963 for medication information extraction and an accuracy of 0.987 for medical concept normalization. Conclusions: A clinical NLP application to EHR opioid scripts would fill a current gap in available batch script processing tools and would greatly enhance individual prescription processing limitations of prescription drug monitoring programs and clinical MME calculators.
William Christopher Mathews, Huy Anh Pham, Hua Xu 0001, Yaoyun Zhang
BIBM4
2020 COVID-19 TestNorm: A tool to normalize COVID-19 testing names to LOINC codes
abstract
Large observational data networks that leverage routine clinical practice data in electronic health records (EHRs) are critical resources for research on coronavirus disease 2019 (COVID-19). Data normalization is a key challenge for the secondary use of EHRs for COVID-19 research across institutions. In this study, we addressed the challenge of automating the normalization of COVID-19 diagnostic tests, which are critical data elements, but for which controlled terminology terms were published after clinical implementation. We developed a simple but effective rule-based tool called COVID-19 TestNorm to automatically normalize local COVID-19 testing names to standard LOINC (Logical Observation Identifiers Names and Codes) codes. COVID-19 TestNorm was developed and evaluated using 568 test names collected from 8 healthcare systems. Our results show that it could achieve an accuracy of 97.4% on an independent test set. COVID-19 TestNorm is available as an open-source package for developers and as an online Web application for end users (https://clamp.uth.edu/covid/loinc.php). We believe that it will be a useful tool to support secondary use of EHRs for research on COVID-19.
Jianfu Li, Ekin Soysal, Jiang Bian 0001, Scott L. DuVall, Elizabeth Hanchrow, Kristine E. Lynch, Michael E. Matheny, Karthik Natarajan, Lucila Ohno-Machado, Serguei V. S. Pakhomov, Ruth M. Reeves, Amy M. Sitapati, Swapna Abhyankar, Theresa A. Cullen, Jami Deckard, Xiaoqian Jiang, Robert Murphy, Hua Xu 0001
J. Am. Medical Informatics Assoc.20
2020 Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithm
abstract
OBJECTIVES: We propose a one-shot, privacy-preserving distributed algorithm to perform logistic regression (ODAL) across multiple clinical sites. MATERIALS AND METHODS: ODAL effectively utilizes the information from the local site (where the patient-level data are accessible) and incorporates the first-order (ODAL1) and second-order (ODAL2) gradients of the likelihood function from other sites to construct an estimator without requiring iterative communication across sites or transferring patient-level data. We evaluated ODAL via extensive simulation studies and an application to a dataset from the University of Pennsylvania Health System. The estimation accuracy was evaluated by comparing it with the estimator based on the combined individual participant data or pooled data (ie, gold standard). RESULTS: Our simulation studies revealed that the relative estimation bias of ODAL1 compared with the pooled estimates was <3%, and the ratio of standard errors was <1.25 for all scenarios. ODAL2 achieved higher accuracy (with relative bias <0.1% and ratio of standard errors <1.05). In real data analysis, we investigated the associations of 100 medications with fetal loss during pregnancy. We found that ODAL1 provided estimates with relative bias <10% for 85% of medications, and ODAL2 has relative bias <10% for 99% of medications. For communication cost, ODAL1 requires transferring p numbers from each site to the local site and ODAL2 requires transferring (p×p+p) numbers from each site to the local site, where p is the number of parameters in the regression model. CONCLUSIONS: This study demonstrates that ODAL is privacy-preserving and communication-efficient with small bias and high statistical efficiency.
Rui Duan 0004, Mary Regina Boland, Howard H. Chang, Hua Xu 0001, Haitao Chu, Christopher H. Schmid, Christopher B. Forrest, John H. Holmes, Martijn J. Schuemie, Jesse A. Berlin, Jason H. Moore, Yong Chen 0016
J. Am. Medical Informatics Assoc.6
2020 Learning from local to global: An efficient distributed algorithm for modeling time-to-event data
abstract
OBJECTIVE: We developed and evaluated a privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) without sharing patient-level information across sites. MATERIALS AND METHODS: Using patient-level data from a single site combined with only aggregated information from other sites, we constructed a surrogate likelihood function, approximating the Cox partial likelihood function obtained using patient-level data from all sites. By maximizing the surrogate likelihood function, each site obtained a local estimate of the model parameter, and the ODAC estimator was constructed as a weighted average of all the local estimates. We evaluated the performance of ODAC with (1) a simulation study and (2) a real-world use case study using 4 datasets from the Observational Health Data Sciences and Informatics network. RESULTS: On the one hand, our simulation study showed that ODAC provided estimates nearly the same as the estimator obtained by analyzing, in a single dataset, the combined patient-level data from all sites (ie, the pooled estimator). The relative bias was <0.1% across all scenarios. The accuracy of ODAC remained high across different sample sizes and event rates. On the other hand, the meta-analysis estimator, which was obtained by the inverse variance weighted average of the site-specific estimates, had substantial bias when the event rate is <5%, with the relative bias reaching 20% when the event rate is 1%. In the Observational Health Data Sciences and Informatics network application, the ODAC estimates have a relative bias <5% for 15 out of 16 log hazard ratios, whereas the meta-analysis estimates had substantially higher bias than ODAC. CONCLUSIONS: ODAC is a privacy-preserving and noniterative method for implementing time-to-event analyses across multiple sites. It provides estimates on par with the pooled estimator and substantially outperforms the meta-analysis estimator when the event is uncommon, making it extremely suitable for studying rare events and diseases in a distributed manner.
Rui Duan 0004, Chongliang Luo, Martijn J. Schuemie, Jiayi Tong, C. Jason Liang, Howard H. Chang, Mary Regina Boland, Jiang Bian 0001, Hua Xu 0001, John H. Holmes, Christopher B. Forrest, Sally C. Morton, Jesse A. Berlin, Jason H. Moore, Kevin B. Mahoney, Yong Chen 0016
J. Am. Medical Informatics Assoc.9
2020 Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical events
abstract
OBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making.
Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao
J. Am. Medical Informatics Assoc.13
2020 SCOR: A secure international informatics infrastructure to investigate COVID-19
abstract
Global pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale.
Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux
J. Am. Medical Informatics Assoc.26
2020 Representation of EHR data for predictive modeling: a comparison between UMLS and other terminologies
abstract
OBJECTIVE: Predictive disease modeling using electronic health record data is a growing field. Although clinical data in their raw form can be used directly for predictive modeling, it is a common practice to map data to standard terminologies to facilitate data aggregation and reuse. There is, however, a lack of systematic investigation of how different representations could affect the performance of predictive models, especially in the context of machine learning and deep learning. MATERIALS AND METHODS: We projected the input diagnoses data in the Cerner HealthFacts database to Unified Medical Language System (UMLS) and 5 other terminologies, including CCS, CCSR, ICD-9, ICD-10, and PheWAS, and evaluated the prediction performances of these terminologies on 2 different tasks: the risk prediction of heart failure in diabetes patients and the risk prediction of pancreatic cancer. Two popular models were evaluated: logistic regression and a recurrent neural network. RESULTS: For logistic regression, using UMLS delivered the optimal area under the receiver operating characteristics (AUROC) results in both dengue hemorrhagic fever (81.15%) and pancreatic cancer (80.53%) tasks. For recurrent neural network, UMLS worked best for pancreatic cancer prediction (AUROC 82.24%), second only (AUROC 85.55%) to PheWAS (AUROC 85.87%) for dengue hemorrhagic fever prediction. DISCUSSION/CONCLUSION: In our experiments, terminologies with larger vocabularies and finer-grained representations were associated with better prediction performances. In particular, UMLS is consistently 1 of the best-performing ones. We believe that our work may help to inform better designs of predictive models, although further investigation is warranted.
Laila Rasmy, Firat Tiryaki, Yujia Zhou 0003, Yang Xiang 0003, Cui Tao, Hua Xu 0001, Degui Zhi
J. Am. Medical Informatics Assoc.6
2020 A study of deep learning approaches for medication and adverse drug event extraction from clinical text
abstract
OBJECTIVE: This article presents our approaches to extraction of medications and associated adverse drug events (ADEs) from clinical documents, which is the second track of the 2018 National NLP Clinical Challenges (n2c2) shared task. MATERIALS AND METHODS: The clinical corpus used in this study was from the MIMIC-III database and the organizers annotated 303 documents for training and 202 for testing. Our system consists of 2 components: a named entity recognition (NER) and a relation classification (RC) component. For each component, we implemented deep learning-based approaches (eg, BI-LSTM-CRF) and compared them with traditional machine learning approaches, namely, conditional random fields for NER and support vector machines for RC, respectively. In addition, we developed a deep learning-based joint model that recognizes ADEs and their relations to medications in 1 step using a sequence labeling approach. To further improve the performance, we also investigated different ensemble approaches to generating optimal performance by combining outputs from multiple approaches. RESULTS: Our best-performing systems achieved F1 scores of 93.45% for NER, 96.30% for RC, and 89.05% for end-to-end evaluation, which ranked #2, #1, and #1 among all participants, respectively. Additional evaluations show that the deep learning-based approaches did outperform traditional machine learning algorithms in both NER and RC. The joint model that simultaneously recognizes ADEs and their relations to medications also achieved the best performance on RC, indicating its promise for relation extraction. CONCLUSION: In this study, we developed deep learning approaches for extracting medications and their attributes such as ADEs, and demonstrated its superior performance compared with traditional machine learning algorithms, indicating its uses in broader NER and RC tasks in the medical domain.
Qiang Wei 0002, Zongcheng Ji, Zhiheng Li 0004, Jingcheng Du, Jun Xu 0007, Yang Xiang 0003, Firat Tiryaki, Stephen Wu 0004, Yaoyun Zhang, Cui Tao, Hua Xu 0001
J. Am. Medical Informatics Assoc.12
2020 Deep learning in clinical natural language processing: a methodical review
abstract
OBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research. MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers. RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) and word2vec embeddings (74.1%) were the most popular methods; the information extraction tasks of text classification, named entity recognition, and relation extraction were dominant (89.2%). However, there was a "long tail" of other methods and specific tasks. Most contributions were methodological variants or applications, but 20.8% were new methods of some kind. The earliest adopters were in the NLP community, but the medical informatics community was the most prolific. DISCUSSION: Our analysis shows growing acceptance of deep learning as a baseline for NLP research, and of DL-based NLP in the medical community. A number of common associations were substantiated (eg, the preference of recurrent neural networks for sequence-labeling named entity recognition), while others were surprisingly nuanced (eg, the scarcity of French language clinical NLP with deep learning). CONCLUSION: Deep learning has not yet fully penetrated clinical NLP and is growing rapidly. This review highlighted both the popular and unique trends in this active field.
Stephen Wu 0004, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiang Wei 0002, Yang Xiang 0003, Bo Zhao 0001, Hua Xu 0001
J. Am. Medical Informatics Assoc.12
2020 A study of entity-linking methods for normalizing Chinese diagnosis and procedure terms to ICD codes
Zongcheng Ji, Stephen Wu 0004, Weiyan Lin, Wenzhen Li, Guohong Xiao, Hua Xu 0001, Yi Zhou 0005
J. Biomed. Informatics10
2019 Achievability to Extract Specific Date Information for Cancer Research
Liwei Wang 0010, Jason A. Wampfler, Angela Dispenzieri, Hua Xu 0001
AMIA4
2019 CLAMP-ATR: A deep learning pipeline for attribute recognition of clinical concepts
Yaoyun Zhang, Hua Xu 0001
AMIA3
2019 Relation Extraction from Clinical Narratives Using Pre-trained Language Models
Qiang Wei 0002, Zongcheng Ji, Yuqi Si, Jingcheng Du, Firat Tiryaki, Stephen Wu 0004, Cui Tao, Kirk Roberts, Hua Xu 0001
AMIA10
2019 Ontology of Consumer Health Vocabulary: providing a formal and interoperable semantic resource for linking lay language and medical terminology
abstract
The Consumer Health Vocabulary has been an important contribution to the health informatics field since its introduction in 2006. Many studies have utilized the vocabulary for various scientific research to bridge the gap between consumers and health experts. Given the flat file format of the Consumer Health Vocabulary dataset, we developed a SKOS-based ontology of the dataset. As an ontology, this dataset can be semantically linked to other resources to provide consumer-level meaning. In addition with this artifact, we plan to further expand the terminology.
Muhammad Amith, Licong Cui, Kirk Roberts, Hua Xu 0001, Cui Tao
BIBM4
2019 Extracting entities with attributes in clinical text via joint deep learning
abstract
OBJECTIVE: Extracting clinical entities and their attributes is a fundamental task of natural language processing (NLP) in the medical domain. This task is typically recognized as 2 sequential subtasks in a pipeline, clinical entity or attribute recognition followed by entity-attribute relation extraction. One problem of pipeline methods is that errors from entity recognition are unavoidably passed to relation extraction. We propose a novel joint deep learning method to recognize clinical entities or attributes and extract entity-attribute relations simultaneously. MATERIALS AND METHODS: The proposed method integrates 2 state-of-the-art methods for named entity recognition and relation extraction, namely bidirectional long short-term memory with conditional random field and bidirectional long short-term memory, into a unified framework. In this method, relation constraints between clinical entities and attributes and weights of the 2 subtasks are also considered simultaneously. We compare the method with other related methods (ie, pipeline methods and other joint deep learning methods) on an existing English corpus from SemEval-2015 and a newly developed Chinese corpus. RESULTS: Our proposed method achieves the best F1 of 74.46% on entity recognition and the best F1 of 50.21% on relation extraction on the English corpus, and 89.32% and 88.13% on the Chinese corpora, respectively, which outperform the other methods on both tasks. CONCLUSIONS: The joint deep learning-based method could improve both entity recognition and relation extraction from clinical text in both English and Chinese, indicating that the approach is promising.
Xue Shi, Yingping Yi, Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Zongcheng Ji, Yaoyun Zhang, Hua Xu 0001
J. Am. Medical Informatics Assoc.9
2019 Enhancing clinical concept extraction with contextual embeddings
abstract
OBJECTIVE: Neural network-based representations ("embeddings") have dramatically advanced natural language processing (NLP) tasks, including clinical NLP tasks such as concept extraction. Recently, however, more advanced embedding methods and representations (eg, ELMo, BERT) have further pushed the state of the art in NLP, yet there are no common best practices for how to integrate these representations into clinical tasks. The purpose of this study, then, is to explore the space of possible options in utilizing these new models for clinical concept extraction, including comparing these to traditional word embedding methods (word2vec, GloVe, fastText). MATERIALS AND METHODS: Both off-the-shelf, open-domain embeddings and pretrained clinical embeddings from MIMIC-III (Medical Information Mart for Intensive Care III) are evaluated. We explore a battery of embedding methods consisting of traditional word embeddings and contextual embeddings and compare these on 4 concept extraction corpora: i2b2 2010, i2b2 2012, SemEval 2014, and SemEval 2015. We also analyze the impact of the pretraining time of a large language model like ELMo or BERT on the extraction performance. Last, we present an intuitive way to understand the semantic information encoded by contextual embeddings. RESULTS: Contextual embeddings pretrained on a large clinical corpus achieves new state-of-the-art performances across all concept extraction tasks. The best-performing model outperforms all state-of-the-art methods with respective F1-measures of 90.25, 93.18 (partial), 80.74, and 81.65. CONCLUSIONS: We demonstrate the potential of contextual embeddings through the state-of-the-art performance these methods achieve on clinical concept extraction. Additionally, we demonstrate that contextual embeddings encode valuable semantic information not accounted for in traditional word representations.
Yuqi Si, Hua Xu 0001, Kirk Roberts
J. Am. Medical Informatics Assoc.3
2019 Cost-aware active learning for named entity recognition in clinical text
abstract
OBJECTIVE: Active Learning (AL) attempts to reduce annotation cost (ie, time) by selecting the most informative examples for annotation. Most approaches tacitly (and unrealistically) assume that the cost for annotating each sample is identical. This study introduces a cost-aware AL method, which simultaneously models both the annotation cost and the informativeness of the samples and evaluates both via simulation and user studies. MATERIALS AND METHODS: We designed a novel, cost-aware AL algorithm (Cost-CAUSE) for annotating clinical named entities; we first utilized lexical and syntactic features to estimate annotation cost, then we incorporated this cost measure into an existing AL algorithm. Using the 2010 i2b2/VA data set, we then conducted a simulation study comparing Cost-CAUSE with noncost-aware AL methods, and a user study comparing Cost-CAUSE with passive learning. RESULTS: Our cost model fit empirical annotation data well, and Cost-CAUSE increased the simulation area under the learning curve (ALC) scores by up to 5.6% and 4.9%, compared with random sampling and alternate AL methods. Moreover, in a user annotation task, Cost-CAUSE outperformed passive learning on the ALC score and reduced annotation time by 20.5%-30.2%. DISCUSSION: Although AL has proven effective in simulations, our user study shows that a real-world environment is far more complex. Other factors have a noticeable effect on the AL method, such as the annotation accuracy of users, the tiredness of users, and even the physical and mental condition of users. CONCLUSION: Cost-CAUSE saves significant annotation cost compared to random sampling.
Qiang Wei 0002, Yukun Chen 0001, Mandana Salimi, Joshua C. Denny, Qiaozhu Mei, Thomas A. Lasko, Qingxia Chen, Stephen Wu 0004, Amy Franklin, Trevor Cohen, Hua Xu 0001
J. Am. Medical Informatics Assoc.11
2018 Development of HeartData, a Data Discovery Index Prototype for Cardiovascular data
Mandana Salimi, Anupama E. Gururaj, Cui Tao, Degui Zhi, Hua Xu 0001
AMIA6
2018 Methods and Tools to Enhance Rigor and Reproducibility of Biomedical Research
Halil Kilicoglu, Aurélie Névéol, Timothy Clark, Hua Xu 0001, Neil R. Smalheiser
AMIA4
2018 Building a Biomedical Bilingual Parallel Corpus between Chinese and English from PubMed
Lingyi Tang, Yaoyun Zhang, Hua Xu 0001
AMIA3
2018 Clinical text annotation - what factors are associated with the cost of time?
Qiang Wei 0002, Amy Franklin, Trevor Cohen, Hua Xu 0001
AMIA4
2018 Combine Factual Medical Knowledge and Distributed Word Representation to Improve Clinical Named Entity Recognition
Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Yi Guo 0005, Hua Xu 0001, William R. Hogan
AMIA5
2018 Detecting Anatomical Entities in Clinical Text
Jun Xu 0007, Yaoyun Zhang, Jingcheng Du, Firat Tiryaki, Hua Xu 0001
AMIA6
2018 CLAMP-PA: A machine learning based pre-annotation pipeline for corpus construction of clinical concepts
Yaoyun Zhang, Qiang Wei 0002, Hua Xu 0001
AMIA4
2018 Guideline-driven data quality assessment- A case study of warfarin management
Yaoyun Zhang, Hua Xu 0001, Chunhua Weng
AMIA3
2018 DataMed - an open source discovery index for finding biomedical datasets
abstract
OBJECTIVE: Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. MATERIALS AND METHODS: DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. RESULTS AND CONCLUSION: Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community.
Anupama E. Gururaj, Ibrahim Burak Özyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Yueling Li, Nansu Zong, Min Jiang 0007, Deevakar Rogith, Mandana Salimi, Hyeon-Eui Kim, Philippe Rocca-Serra, Alejandra N. González-Beltrán, Claudiu Farcas, Todd Johnson, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Ian Fore, Lucila Ohno-Machado, Jeffrey S. Grethe, Hua Xu 0001
J. Am. Medical Informatics Assoc.24
2018 User needs analysis and usability assessment of DataMed - a biomedical data discovery index
abstract
OBJECTIVE: To present user needs and usability evaluations of DataMed, a Data Discovery Index (DDI) that allows searching for biomedical data from multiple sources. MATERIALS AND METHODS: We conducted 2 phases of user studies. Phase 1 was a user needs analysis conducted before the development of DataMed, consisting of interviews with researchers. Phase 2 involved iterative usability evaluations of DataMed prototypes. We analyzed data qualitatively to document researchers' information and user interface needs. RESULTS: Biomedical researchers' information needs in data discovery are complex, multidimensional, and shaped by their context, domain knowledge, and technical experience. User needs analyses validate the need for a DDI, while usability evaluations of DataMed show that even though aggregating metadata into a common search engine and applying traditional information retrieval tools are promising first steps, there remain challenges for DataMed due to incomplete metadata and the complexity of data discovery. DISCUSSION: Biomedical data poses distinct problems for search when compared to websites or publications. Making data available is not enough to facilitate biomedical data discovery: new retrieval techniques and user interfaces are necessary for dataset exploration. Consistent, complete, and high-quality metadata are vital to enable this process. CONCLUSION: While available data and researchers' information needs are complex and heterogeneous, a successful DDI must meet those needs and fit into the processes of biomedical researchers. Research directions include formalizing researchers' information needs, standardizing overviews of data to facilitate relevance judgments, implementing user interfaces for concept-based searching, and developing evaluation methods for open-ended discovery systems such as DDIs.
Ram Dixit, Deevakar Rogith, Vidya Narayana, Mandana Salimi, Anupama E. Gururaj, Lucila Ohno-Machado, Hua Xu 0001, Todd R. Johnson
J. Am. Medical Informatics Assoc.7
2018 PIE: A prior knowledge guided integrated likelihood estimation method for bias reduction in association studies using electronic health records data
abstract
OBJECTIVES: This study proposes a novel Prior knowledge guided Integrated likelihood Estimation (PIE) method to correct bias in estimations of associations due to misclassification of electronic health record (EHR)-derived binary phenotypes, and evaluates the performance of the proposed method by comparing it to 2 methods in common practice. METHODS: We conducted simulation studies and data analysis of real EHR-derived data on diabetes from Kaiser Permanente Washington to compare the estimation bias of associations using the proposed method, the method ignoring phenotyping errors, the maximum likelihood method with misspecified sensitivity and specificity, and the maximum likelihood method with correctly specified sensitivity and specificity (gold standard). The proposed method effectively leverages available information on phenotyping accuracy to construct a prior distribution for sensitivity and specificity, and incorporates this prior information through the integrated likelihood for bias reduction. RESULTS: Our simulation studies and real data application demonstrated that the proposed method effectively reduces the estimation bias compared to the 2 current methods. It performed almost as well as the gold standard method when the prior had highest density around true sensitivity and specificity. The analysis of EHR data from Kaiser Permanente Washington showed that the estimated associations from PIE were very close to the estimates from the gold standard method and reduced bias by 60%-100% compared to the 2 commonly used methods in current practice for EHR data. CONCLUSIONS: This study demonstrates that the proposed method can effectively reduce estimation bias caused by imperfect phenotyping in EHR-derived data by incorporating prior information through integrated likelihood.
Jing Huang 0021, Rui Duan 0004, Rebecca A. Hubbard, Yonghui Wu 0001, Jason H. Moore, Hua Xu 0001, Yong Chen 0016
J. Am. Medical Informatics Assoc.6
2018 CLAMP - a toolkit for efficiently building customized clinical natural language processing pipelines
abstract
Existing general clinical natural language processing (NLP) systems such as MetaMap and Clinical Text Analysis and Knowledge Extraction System have been successfully applied to information extraction from clinical text. However, end users often have to customize existing systems for their individual tasks, which can require substantial NLP skills. Here we present CLAMP (Clinical Language Annotation, Modeling, and Processing), a newly developed clinical NLP toolkit that provides not only state-of-the-art NLP components, but also a user-friendly graphic user interface that can help users quickly build customized NLP pipelines for their individual applications. Our evaluation shows that the CLAMP default pipeline achieved good performance on named entity recognition and concept encoding. We also demonstrate the efficiency of the CLAMP graphic user interface in building customized, high-performance NLP pipelines with 2 use cases, extracting smoking status and lab test values. CLAMP is publicly available for research use, and we believe it is a unique asset for the clinical NLP community.
Ergin Soysal, Min Jiang 0007, Yonghui Wu 0001, Serguei V. S. Pakhomov, Hua Xu 0001
J. Am. Medical Informatics Assoc.7
2018 Toward a normalized clinical drug knowledge base in China - applying the RxNorm model to Chinese clinical drugs
abstract
Objective: In recent years, electronic health record systems have been widely implemented in China, making clinical data available electronically. However, little effort has been devoted to making drug information exchangeable among these systems. This study aimed to build a Normalized Chinese Clinical Drug (NCCD) knowledge base, by applying and extending the information model of RxNorm to Chinese clinical drugs. Methods: Chinese drugs were collected from 4 major resources-China Food and Drug Administration, China Health Insurance Systems, Hospital Pharmacy Systems, and China Pharmacopoeia-for integration and normalization in NCCD. Chemical drugs were normalized using the information model in RxNorm without much change. Chinese patent drugs (i.e., Chinese herbal extracts), however, were represented using an expanded RxNorm model to incorporate the unique characteristics of these drugs. A hybrid approach combining automated natural language processing technologies and manual review by domain experts was then applied to drug attribute extraction, normalization, and further generation of drug names at different specification levels. Lastly, we reported the statistics of NCCD, as well as the evaluation results using several sets of randomly selected Chinese drugs. Results: The current version of NCCD contains 16 976 chemical drugs and 2663 Chinese patent medicines, resulting in 19 639 clinical drugs, 250 267 unique concepts, and 2 602 760 relations. By manual review of 1700 chemical drugs and 250 Chinese patent drugs randomly selected from NCCD (about 10%), we showed that the hybrid approach could achieve an accuracy of 98.60% for drug name extraction and normalization. Using a collection of 500 chemical drugs and 500 Chinese patent drugs from other resources, we showed that NCCD achieved coverages of 97.0% and 90.0% for chemical drugs and Chinese patent drugs, respectively. Conclusion: Evaluation results demonstrated the potential to improve interoperability across various electronic drug systems in China.
Li Wang 0077, Yaoyun Zhang, Min Jiang 0007, Jiancheng Dong, Yun Liu 0020, Cui Tao, Guoqian Jiang, Yi Zhou 0005, Hua Xu 0001
J. Am. Medical Informatics Assoc.10
2018 Interactive medical word sense disambiguation through informed learning
abstract
Objective: Medical word sense disambiguation (WSD) is challenging and often requires significant training with data labeled by domain experts. This work aims to develop an interactive learning algorithm that makes efficient use of expert's domain knowledge in building high-quality medical WSD models with minimal human effort. Methods: We developed an interactive learning algorithm with expert labeling instances and features. An expert can provide supervision in 3 ways: labeling instances, specifying indicative words of a sense, and highlighting supporting evidence in a labeled instance. The algorithm learns from these labels and iteratively selects the most informative instances to ask for future labels. Our evaluation used 3 WSD corpora: 198 ambiguous terms from Medical Subject Headings (MSH) as MEDLINE indexing terms, 74 ambiguous abbreviations in clinical notes from the University of Minnesota (UMN), and 24 ambiguous abbreviations in clinical notes from Vanderbilt University Hospital (VUH). For each ambiguous term and each learning algorithm, a learning curve that plots the accuracy on the test set against the number of labeled instances was generated. The area under the learning curve was used as the primary evaluation metric. Results: Our interactive learning algorithm significantly outperformed active learning, the previous fastest learning algorithm for medical WSD. Compared to active learning, it achieved 90% accuracy for the MSH corpus with 42% less labeling effort, 35% less labeling effort for the UMN corpus, and 16% less labeling effort for the VUH corpus. Conclusions: High-quality WSD models can be efficiently trained with minimal supervision by inviting experts to label informative instances and provide domain knowledge through labeling/highlighting contextual features.
Yue Wang 0035, Kai Zheng 0002, Hua Xu 0001, Qiaozhu Mei
J. Am. Medical Informatics Assoc.3
2018 Predict effective drug combination by deep belief network and ontology fingerprints
Guocai Chen, Alex Tsoi, Hua Xu 0001, W. Jim Zheng
J. Biomed. Informatics3
2018 A study of generalizability of recurrent neural network-based predictive models for heart failure onset risk using a large and heterogeneous EHR data set
Laila Rasmy, Yonghui Wu 0001, Ningtao Wang, W. Jim Zheng, Fei Wang 0001, Hulin Wu, Hua Xu 0001, Degui Zhi
J. Biomed. Informatics8
2017 A Natural Language Processing System for Biomedical Dataset Retrieval
Jun Xu 0007, Anupama E. Gururaj, Lucila Ohno-Machado, Hua Xu 0001
AMIA6
2017 Meeting User Needs for a Data Discovery Index of Biomedical Big Data
Ram Dixit, Deevakar Rogith, Vidya Narayana, Mandana Salimi, Anupama E. Gururaj, Lucila Ohno-Machado, Hua Xu 0001, Todd R. Johnson
AMIA7
2017 AMICUS: A Metasystem for Interoperation and Combination of UIMA Systems
Gregory P. Finley, Benjamin Knoll, Reed McEwan, Genevieve B. Melton, Hua Xu 0001, Serguei V. S. Pakhomov
AMIA6
2017 Leveraging existing corpora for de-identification of psychiatric notes using domain adaptation
Hee-Jin Lee, Yaoyun Zhang, Kirk Roberts, Hua Xu 0001
AMIA4
2017 Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Noémie Elhadad, Sumithra Velupillai, Hua Xu 0001, Guergana K. Savova
AMIA4
2017 Information Retrieval for Biomedical Datasets: The 2016 bioCADDIE Challenge
Kirk Roberts, Anupama E. Gururaj, Saeid Pournejati, Trevor Cohen, William R. Hersh, Dina Demner-Fushman, Lucila Ohno-Machado, Hua Xu 0001
AMIA9
2017 Identifying Disease Hierarchical Relation with Concept Embeddings and Distant Supervision
Lingyi Tang, Yaoyun Zhang, Hua Xu 0001
AMIA3
2017 Interactive Medical Word Sense Disambiguation with Instance and Feature Labeling
Yue Wang 0035, Kai Zheng 0002, Hua Xu 0001, Qiaozhu Mei
AMIA3
2017 Matching Consumer Health Vocabulary with Professional Medical Terms Through Concept Embedding
Yue Wang 0035, Jian Tang 0005, V. G. Vinod Vydiswaran, Kai Zheng 0002, Hua Xu 0001, Qiaozhu Mei
AMIA5
2017 Clinical Named Entity Recognition Using Deep Learning Models
Yonghui Wu 0001, Min Jiang 0007, Jun Xu 0007, Degui Zhi, Hua Xu 0001
AMIA5
2017 CLAMP - A User-Centric Clinical Natural Language Processing Toolkit
Hua Xu 0001, Ergin Soysal, Min Jiang 0007, Yonghui Wu 0001
AMIA1
2017 Detecting Contradictory and Consistent Citations in Biomedical Literature
Jun Xu 0007, Yonghui Wu 0001, Yaoyun Zhang, Qiang Wei 0002, Hua Xu 0001
AMIA5
2017 Detecting Body Location Modifiers of Disorders in Clinical Texts via Sequence Labeling
Jun Xu 0007, Yonghui Wu 0001, Yaoyun Zhang, Hua Xu 0001
AMIA4
2017 Evaluating Word Embeddings from Multiple Domains for Symptom Recognition in Psychiatric Notes
Yaoyun Zhang, Hee-Jin Lee, Yonghui Wu 0001, Hua Xu 0001
AMIA4
2017 A pilot study of mining association between psychiatric stressors and symptoms in tweets
abstract
Suicide is a significant public health issue, causing huge impacts on individuals as well as their families. Psychiatric stressors are major suicide risk factors and can profoundly impact a person in many aspects. In order to facilitate the understanding of psychiatric stressors and the associated symptoms, we extracted stressors and symptoms terms from major online knowledge repositories and psychiatric clinical notes. The current vocabulary collection contains 1,292 psychiatric symptoms and 715 psychiatric stressors, which was leveraged to study the associations between stressors and symptoms in a corpus of suicide related tweets. Using Chi-Square test with Bonferroni correction, 3,500 symptom-stressor pairs were identified with significant association (p-value<;0.01).
Jingcheng Du, Yaoyun Zhang, Cui Tao, Hua Xu 0001
BIBM4
2017 Towards practical temporal relation extraction from clinical notes: An analysis of direct temporal relations
abstract
Following the conventions developed in general domain, most of the current work on clinical temporal relation identification aims to identify a comprehensive set of temporal relations from source documents. This includes both explicit relations that is described in the documents and implicit relations that are identifiable only through inference. Although such an approach may provide a complete view of temporal information provided in a document, some temporal relations may not be practically essential, depending on the clinical application at hand. In addition, the performances of current systems that identify both explicit and implicit relations are still low and how to enhance the performances to be enough for practical use is not clear yet. In this paper, we propose focus on a subset of temporal relations, in order to provide insights into how to develop practically useful temporal information extraction methods for clinical text. We focus on “direct” temporal relations, which are intra-sentential temporal relations between a time expression and an event mention with limited syntactic distance. A corpus of 120 discharge summaries is constructed, leveraging an existing corpus, the 2012 i2b2 corpus. We show that the direct temporal relations constitute a major category of temporal relations. In addition, we show that the performance of the state-of-the art temporal relation extraction system, which is developed for both implicit and explicit relations, on direct temporal relations is still low. This indicates the need for development of methods tailored to direct temporal relations.
Hee-Jin Lee, Yaoyun Zhang, Jun Xu 0007, Cui Tao, Hua Xu 0001, Min Jiang 0007
BIBM5
2017 Knowledge-Based Approach for Named Entity Recognition in Biomedical Literature: A Use Case in Biomedical Software Identification
Muhammad Amith, Yaoyun Zhang, Hua Xu 0001, Cui Tao
IEA/AIE (2)3
2017 Interweaving Domain Knowledge and Unsupervised Learning for Psychiatric Stressor Extraction from Clinical Notes
Olivia R. Zhang, Yaoyun Zhang, Jun Xu 0007, Kirk Roberts, Xiang Y. Zhang, Hua Xu 0001
IEA/AIE (2)6
2017 Identification of adverse drug-drug interactions through causal association rule discovery from spontaneous adverse event reports
Ruichu Cai, Yong Hu 0002, Brittany Melton, Michael E. Matheny, Hua Xu 0001, Lemuel R. Waitman
Artif. Intell. Medicine6
2017 CNN-based ranking for biomedical entity normalization
abstract
BACKGROUND: Most state-of-the-art biomedical entity normalization systems, such as rule-based systems, merely rely on morphological information of entity mentions, but rarely consider their semantic information. In this paper, we introduce a novel convolutional neural network (CNN) architecture that regards biomedical entity normalization as a ranking problem and benefits from semantic information of biomedical entities. RESULTS: The CNN-based ranking method first generates candidates using handcrafted rules, and then ranks the candidates according to their semantic information modeled by CNN as well as their morphological information. Experiments on two benchmark datasets for biomedical entity normalization show that our proposed CNN-based ranking method outperforms traditional rule-based method with state-of-the-art performance. CONCLUSIONS: We propose a CNN architecture that regards biomedical entity normalization as a ranking problem. Comparison results show that semantic information is beneficial to biomedical entity normalization and can be well combined with morphological information in our CNN architecture for further improvement.
Haodi Li, Qingcai Chen, Buzhou Tang, Xiaolong Wang 0001, Hua Xu 0001
BMC Bioinform.5
2017 Investigating MicroRNA and transcription factor co-regulatory networks in colorectal cancer
abstract
BACKGROUND: Colorectal cancer (CRC) is one of the most common malignancies worldwide with poor prognosis. Studies have showed that abnormal microRNA (miRNA) expression can affect CRC pathogenesis and development through targeting critical genes in cellular system. However, it is unclear about which miRNAs play central roles in CRC's pathogenesis and how they interact with transcription factors (TFs) to regulate the cancer-related genes. RESULTS: To address this issue, we systematically explored the major regulation motifs, namely feed-forward loops (FFLs), that consist of miRNAs, TFs and CRC-related genes through the construction of a miRNA-TF regulatory network in CRC. First, we compiled CRC-related miRNAs, CRC-related genes, and human TFs from multiple data sources. Second, we identified 13,123 3-node FFLs including 25 miRNA-FFLs, 13,005 TF-FFLs and 93 composite-FFLs, and merged the 3-node FFLs to construct a CRC-related regulatory network. The network consists of three types of regulatory subnetworks (SNWs): miRNA-SNW, TF-SNW, and composite-SNW. To enhance the accuracy of the network, the results were filtered by using The Cancer Genome Atlas (TCGA) expression data in CRC, whereby we generated a core regulatory network consisting of 58 significant FFLs. We then applied a hub identification strategy to the significant FFLs and found 5 significant components, including two miRNAs (hsa-miR-25 and hsa-miR-31), two genes (ADAMTSL3 and AXIN1) and one TF (BRCA1). The follow up prognosis analysis indicated all of the 5 significant components having good prediction of overall survival of CRC patients. CONCLUSIONS: In summary, we generated a CRC-specific miRNA-TF regulatory network, which is helpful to understand the complex CRC regulatory mechanisms and guide clinical treatment. The discovered 5 regulators might have critical roles in CRC pathogenesis and warrant future investigation.
Jiamao Luo, Huilin Niu, Jing Wang 0026, Qi Liu 0024, Zhongming Zhao, Hua Xu 0001, Yanqing Ding, Jingchun Sun, Qingling Zhang 0003
BMC Bioinform.8
2017 A long journey to short abbreviations: developing an open-source framework for clinical abbreviation recognition and disambiguation (CARD)
abstract
OBJECTIVE: The goal of this study was to develop a practical framework for recognizing and disambiguating clinical abbreviations, thereby improving current clinical natural language processing (NLP) systems' capability to handle abbreviations in clinical narratives. METHODS: We developed an open-source framework for clinical abbreviation recognition and disambiguation (CARD) that leverages our previously developed methods, including: (1) machine learning based approaches to recognize abbreviations from a clinical corpus, (2) clustering-based semiautomated methods to generate possible senses of abbreviations, and (3) profile-based word sense disambiguation methods for clinical abbreviations. We applied CARD to clinical corpora from Vanderbilt University Medical Center (VUMC) and generated 2 comprehensive sense inventories for abbreviations in discharge summaries and clinic visit notes. Furthermore, we developed a wrapper that integrates CARD with MetaMap, a widely used general clinical NLP system. RESULTS AND CONCLUSION: CARD detected 27 317 and 107 303 distinct abbreviations from discharge summaries and clinic visit notes, respectively. Two sense inventories were constructed for the 1000 most frequent abbreviations in these 2 corpora. Using the sense inventories created from discharge summaries, CARD achieved an F1 score of 0.755 for identifying and disambiguating all abbreviations in a corpus from the VUMC discharge summaries, which is superior to MetaMap and Apache's clinical Text Analysis Knowledge Extraction System (cTAKES). Using additional external corpora, we also demonstrated that the MetaMap-CARD wrapper improved MetaMap's performance in recognizing disorder entities in clinical notes. The CARD framework, 2 sense inventories, and the wrapper for MetaMap are publicly available at https://sbmi.uth.edu/ccb/resources/abbreviation.htm . We believe the CARD framework can be a valuable resource for improving abbreviation identification in clinical NLP systems.
Yonghui Wu 0001, Joshua C. Denny, S. Trent Rosenbloom, Randolph A. Miller, Dario A. Giuse, Carmelo Blanquicett, Ergin Soysal, Jun Xu 0007, Hua Xu 0001
J. Am. Medical Informatics Assoc.10
2016 GWAS Finder: Search Engine for GWAS Datasets in Biomedical Literature
Yaoyun Zhang, Hua Xu 0001
AMIA3
2016 An Empirical Study for Impacts of Measurement Errors on EHR based Association Studies
Rui Duan 0004, Ming Cao 0005, Yonghui Wu 0001, Jing Huang 0021, Joshua C. Denny, Hua Xu 0001, Yong Chen 0016
AMIA6
2016 A Scalable Dataset Indexing Infrastructure for the bioCADDIE Data Discovery System
Jeffrey S. Grethe, Ibrahim Burak Özyurt, Hua Xu 0001, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado
AMIA3
2016 Building Successful Natural Language Processing Applications in Clinical Research and Healthcare Operations
Yang Huang 0008, Hua Xu 0001, Joshua C. Denny
AMIA2
2016 Literature-Based Discovery of Confounding in Observational Clinical Data
Scott A. Malec, Hua Xu 0001, Elmer V. Bernstam, Sahiti Myneni, Trevor Cohen
AMIA3
2016 Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001
AMIA12
2016 Applying Active Learning to Clinical Abbreviation Disambiguation in Real Time
Sungrim Moon, Yukun Chen 0001, Joshua C. Denny, S. Trent Rosenbloom, Ky Nguyen, Tolulola Dawodu, Hua Xu 0001
AMIA8
2016 Semantic Relatedness and Similarity between Biomedical Concepts
Sungrim Moon, Trevor Cohen, Hua Xu 0001
AMIA3
2016 Clinical Word Sense Disambiguation with Interactive Search and Classification
Yue Wang 0035, Kai Zheng 0002, Hua Xu 0001, Qiaozhu Mei
AMIA3
2016 A Study of Active Learning for Document Selection in Clinical Named Entity Recognition
Qiang Wei 0002, Yukun Chen 0001, Sungrim Moon, Trevor Cohen, Hua Xu 0001
AMIA5
2016 What Can Neural Networks Learn from Unlabeled Clinical Narratives?
Yonghui Wu 0001, Jun Xu 0007, Yaoyun Zhang, Hua Xu 0001
AMIA4
2016 Development of DataMed, a Data Discovery Index Prototype by bioCADDIE: Laying the Groundwork for Biomedical Data Discovery
Hua Xu 0001, Jeffrey S. Grethe, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Yueling Li, Ibrahim Burak Özyurt, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado
AMIA1
2016 Semantic Role Labeling of Clinical Text: Comparing Syntactic Parsers and Features
Yaoyun Zhang, Min Jiang 0007, Hua Xu 0001
AMIA4
2016 Extracting genetic alteration information for personalized cancer therapy from ClinicalTrials.gov
abstract
OBJECTIVE: Clinical trials investigating drugs that target specific genetic alterations in tumors are important for promoting personalized cancer therapy. The goal of this project is to create a knowledge base of cancer treatment trials with annotations about genetic alterations from ClinicalTrials.gov. METHODS: We developed a semi-automatic framework that combines advanced text-processing techniques with manual review to curate genetic alteration information in cancer trials. The framework consists of a document classification system to identify cancer treatment trials from ClinicalTrials.gov and an information extraction system to extract gene and alteration pairs from the Title and Eligibility Criteria sections of clinical trials. By applying the framework to trials at ClinicalTrials.gov, we created a knowledge base of cancer treatment trials with genetic alteration annotations. We then evaluated each component of the framework against manually reviewed sets of clinical trials and generated descriptive statistics of the knowledge base. RESULTS AND DISCUSSION: The automated cancer treatment trial identification system achieved a high precision of 0.9944. Together with the manual review process, it identified 20 193 cancer treatment trials from ClinicalTrials.gov. The automated gene-alteration extraction system achieved a precision of 0.8300 and a recall of 0.6803. After validation by manual review, we generated a knowledge base of 2024 cancer trials that are labeled with specific genetic alteration information. Analysis of the knowledge base revealed the trend of increased use of targeted therapy for cancer, as well as top frequent gene-alteration pairs of interest. We expect this knowledge base to be a valuable resource for physicians and patients who are seeking information about personalized cancer therapy.
Jun Xu 0007, Hee-Jin Lee, Yonghui Wu 0001, Yaoyun Zhang, Liang-Chin Huang, Amber M. Johnson, Vijaykumar Holla, Ann M. Bailey, Trevor Cohen, Funda Meric-Bernstam, Elmer V. Bernstam, Hua Xu 0001
J. Am. Medical Informatics Assoc.13
2016 Automated identification of molecular effects of drugs (AIMED)
abstract
INTRODUCTION: Genomic profiling information is frequently available to oncologists, enabling targeted cancer therapy. Because clinically relevant information is rapidly emerging in the literature and elsewhere, there is a need for informatics technologies to support targeted therapies. To this end, we have developed a system for Automated Identification of Molecular Effects of Drugs, to help biomedical scientists curate this literature to facilitate decision support. OBJECTIVES: To create an automated system to identify assertions in the literature concerning drugs targeting genes with therapeutic implications and characterize the challenges inherent in automating this process in rapidly evolving domains. METHODS: We used subject-predicate-object triples (semantic predications) and co-occurrence relations generated by applying the SemRep Natural Language Processing system to MEDLINE abstracts and ClinicalTrials.gov descriptions. We applied customized semantic queries to find drugs targeting genes of interest. The results were manually reviewed by a team of experts. RESULTS: Compared to a manually curated set of relationships, recall, precision, and F2 were 0.39, 0.21, and 0.33, respectively, which represents a 3- to 4-fold improvement over a publically available set of predications (SemMedDB) alone. Upon review of ostensibly false positive results, 26% were considered relevant additions to the reference set, and an additional 61% were considered to be relevant for review. Adding co-occurrence data improved results for drugs in early development, but not their better-established counterparts. CONCLUSIONS: Precision medicine poses unique challenges for biomedical informatics systems that help domain experts find answers to their research questions. Further research is required to improve the performance of such systems, particularly for drugs in development.
Safa Fathiamini, Amber M. Johnson, Alejandro Araya, Vijaykumar Holla, Ann M. Bailey, Beate Litzenburger, Nora S. Sanchez, Yekaterina Khotskaya, Hua Xu 0001, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen
J. Am. Medical Informatics Assoc.10
2015 Citation Sentiment Analysis in Clinical Trial Papers
Jun Xu 0007, Yaoyun Zhang, Yonghui Wu 0001, Hua Xu 0001
AMIA6
2015 Recent Advances in Computational Drug Repositioning
Atul J. Butte, Nigam H. Shah, Nicholas P. Tatonetti, Hua Xu 0001
AMIA4
2015 Real Time Active Learning Study for Clinical Named Entity Recognition
Yukun Chen 0001, Sungrim Moon, Thomas A. Lasko, Qiaozhu Mei, Trevor Cohen, Qingxia Chen, Joshua C. Denny, Hua Xu 0001
AMIA9
2015 Clinical Language Annotation, Modeling, and Processing Toolkit (CLAMP) - a user-centric NLP system
Ergin Soysal, Min Jiang 0007, Yonghui Wu 0001, Hua Xu 0001
AMIA5
2015 Recognizing Disjoint Clinical Concepts in Clinical Text Using Machine Learning-based Methods
Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Yonghui Wu 0001, Yaoyun Zhang, Hua Xu 0001
AMIA7
2015 A Study of Neural Word Embeddings for Named Entity Recognition in Clinical Text
Yonghui Wu 0001, Jun Xu 0007, Min Jiang 0007, Yaoyun Zhang, Hua Xu 0001
AMIA5
2015 A comparative study of disease genes and drug targets in the human protein interactome
abstract
BACKGROUND: Disease genes cause or contribute genetically to the development of the most complex diseases. Drugs are the major approaches to treat the complex disease through interacting with their targets. Thus, drug targets are critical for treatment efficacy. However, the interrelationship between the disease genes and drug targets is not clear. RESULTS: In this study, we comprehensively compared the network properties of disease genes and drug targets for five major disease categories (cancer, cardiovascular disease, immune system disease, metabolic disease, and nervous system disease). We first collected disease genes from genome-wide association studies (GWAS) for five disease categories and collected their corresponding drugs based on drugs' Anatomical Therapeutic Chemical (ATC) classification. Then, we obtained the drug targets for these five different disease categories. We found that, though the intersections between disease genes and drug targets were small, disease genes were significantly enriched in targets compared to their enrichment in human protein-coding genes. We further compared network properties of the proteins encoded by disease genes and drug targets in human protein-protein interaction networks (interactome). The results showed that the drug targets tended to have higher degree, higher betweenness, and lower clustering coefficient in cancer Furthermore, we observed a clear fraction increase of disease proteins or drug targets in the near neighborhood compared with the randomized genes. CONCLUSIONS: The study presents the first comprehensive comparison of the disease genes and drug targets in the context of interactome. The results provide some foundational network characteristics for further designing computational strategies to predict novel drug targets and drug repurposing.
Jingchun Sun, Kevin W. Zhu, W. Jim Zheng, Hua Xu 0001
BMC Bioinform.4
2015 Validating drug repurposing signals using electronic health records: a case study of metformin associated with reduced cancer mortality
abstract
OBJECTIVES: Drug repurposing, which finds new indications for existing drugs, has received great attention recently. The goal of our work is to assess the feasibility of using electronic health records (EHRs) and automated informatics methods to efficiently validate a recent drug repurposing association of metformin with reduced cancer mortality. METHODS: By linking two large EHRs from Vanderbilt University Medical Center and Mayo Clinic to their tumor registries, we constructed a cohort including 32,415 adults with a cancer diagnosis at Vanderbilt and 79,258 cancer patients at Mayo from 1995 to 2010. Using automated informatics methods, we further identified type 2 diabetes patients within the cancer cohort and determined their drug exposure information, as well as other covariates such as smoking status. We then estimated HRs for all-cause mortality and their associated 95% CIs using stratified Cox proportional hazard models. HRs were estimated according to metformin exposure, adjusted for age at diagnosis, sex, race, body mass index, tobacco use, insulin use, cancer type, and non-cancer Charlson comorbidity index. RESULTS: Among all Vanderbilt cancer patients, metformin was associated with a 22% decrease in overall mortality compared to other oral hypoglycemic medications (HR 0.78; 95% CI 0.69 to 0.88) and with a 39% decrease compared to type 2 diabetes patients on insulin only (HR 0.61; 95% CI 0.50 to 0.73). Diabetic patients on metformin also had a 23% improved survival compared with non-diabetic patients (HR 0.77; 95% CI 0.71 to 0.85). These associations were replicated using the Mayo Clinic EHR data. Many site-specific cancers including breast, colorectal, lung, and prostate demonstrated reduced mortality with metformin use in at least one EHR. CONCLUSIONS: EHR data suggested that the use of metformin was associated with decreased mortality after a cancer diagnosis compared with diabetic and non-diabetic cancer patients not on metformin, indicating its potential as a chemotherapeutic regimen. This study serves as a model for robust and inexpensive validation studies for drug repurposing signals using EHR data.
Hua Xu 0001, Melinda Aldrich, Qingxia Chen, Neeraja B. Peterson, Mia A. Levy, Anushi Shah, Xiaoyang Ruan, Min Jiang 0007, Jamii St Julien, Jeremy L. Warner, Carol Friedman, Dan M. Roden, Joshua C. Denny
J. Am. Medical Informatics Assoc.1
2015 Domain adaptation for semantic role labeling of clinical text
abstract
OBJECTIVE: Semantic role labeling (SRL), which extracts a shallow semantic relation representation from different surface textual forms of free text sentences, is important for understanding natural language. Few studies in SRL have been conducted in the medical domain, primarily due to lack of annotated clinical SRL corpora, which are time-consuming and costly to build. The goal of this study is to investigate domain adaptation techniques for clinical SRL leveraging resources built from newswire and biomedical literature to improve performance and save annotation costs. MATERIALS AND METHODS: Multisource Integrated Platform for Answering Clinical Questions (MiPACQ), a manually annotated SRL clinical corpus, was used as the target domain dataset. PropBank and NomBank from newswire and BioProp from biomedical literature were used as source domain datasets. Three state-of-the-art domain adaptation algorithms were employed: instance pruning, transfer self-training, and feature augmentation. The SRL performance using different domain adaptation algorithms was evaluated by using 10-fold cross-validation on the MiPACQ corpus. Learning curves for the different methods were generated to assess the effect of sample size. RESULTS AND CONCLUSION: When all three source domain corpora were used, the feature augmentation algorithm achieved statistically significant higher F-measure (83.18%), compared to the baseline with MiPACQ dataset alone (F-measure, 81.53%), indicating that domain adaptation algorithms may improve SRL performance on clinical text. To achieve a comparable performance to the baseline method that used 90% of MiPACQ training samples, the feature augmentation algorithm required <50% of training samples in MiPACQ, demonstrating that annotation costs of clinical SRL can be reduced significantly by leveraging existing SRL resources from other domains.
Yaoyun Zhang, Buzhou Tang, Min Jiang 0007, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2015 A study of active learning methods for named entity recognition in clinical text
Yukun Chen 0001, Thomas A. Lasko, Qiaozhu Mei, Joshua C. Denny, Hua Xu 0001
J. Biomed. Informatics5
2015 Identifying risk factors for heart disease over time: Overview of 2014 i2b2/UTHealth shared task Track 2
abstract
The second track of the 2014 i2b2/UTHealth natural language processing shared task focused on identifying medical risk factors related to Coronary Artery Disease (CAD) in the narratives of longitudinal medical records of diabetic patients. The risk factors included hypertension, hyperlipidemia, obesity, smoking status, and family history, as well as diabetes and CAD, and indicators that suggest the presence of those diseases. In addition to identifying the risk factors, this track of the 2014 i2b2/UTHealth shared task studied the presence and progression of the risk factors in longitudinal medical records. Twenty teams participated in this track, and submitted 49 system runs for evaluation. Six of the top 10 teams achieved F1 scores over 0.90, and all 10 scored over 0.87. The most successful system used a combination of additional annotations, external lexicons, hand-written rules and Support Vector Machines. The results of this track indicate that identification of risk factors and their progression over time is well within the reach of automated systems.
Amber Stubbs, Christopher Kotfila, Hua Xu 0001, Özlem Uzuner
J. Biomed. Informatics3
2015 Ease of adoption of clinical natural language processing software: An evaluation of five systems
abstract
OBJECTIVE: In recognition of potential barriers that may inhibit the widespread adoption of biomedical software, the 2014 i2b2 Challenge introduced a special track, Track 3 - Software Usability Assessment, in order to develop a better understanding of the adoption issues that might be associated with the state-of-the-art clinical NLP systems. This paper reports the ease of adoption assessment methods we developed for this track, and the results of evaluating five clinical NLP system submissions. MATERIALS AND METHODS: A team of human evaluators performed a series of scripted adoptability test tasks with each of the participating systems. The evaluation team consisted of four "expert evaluators" with training in computer science, and eight "end user evaluators" with mixed backgrounds in medicine, nursing, pharmacy, and health informatics. We assessed how easy it is to adopt the submitted systems along the following three dimensions: communication effectiveness (i.e., how effective a system is in communicating its designed objectives to intended audience), effort required to install, and effort required to use. We used a formal software usability testing tool, TURF, to record the evaluators' interactions with the systems and 'think-aloud' data revealing their thought processes when installing and using the systems and when resolving unexpected issues. RESULTS: Overall, the ease of adoption ratings that the five systems received are unsatisfactory. Installation of some of the systems proved to be rather difficult, and some systems failed to adequately communicate their designed objectives to intended adopters. Further, the average ratings provided by the end user evaluators on ease of use and ease of interpreting output are -0.35 and -0.53, respectively, indicating that this group of users generally deemed the systems extremely difficult to work with. While the ratings provided by the expert evaluators are higher, 0.6 and 0.45, respectively, these ratings are still low indicating that they also experienced considerable struggles. DISCUSSION: The results of the Track 3 evaluation show that the adoptability of the five participating clinical NLP systems has a great margin for improvement. Remedy strategies suggested by the evaluators included (1) more detailed and operation system specific use instructions; (2) provision of more pertinent onscreen feedback for easier diagnosis of problems; (3) including screen walk-throughs in use instructions so users know what to expect and what might have gone wrong; (4) avoiding jargon and acronyms in materials intended for end users; and (5) packaging prerequisites required within software distributions so that prospective adopters of the software do not have to obtain each of the third-party components on their own.
Kai Zheng 0002, V. G. Vinod Vydiswaran, Yang Liu 0019, Yue Wang 0035, Amber Stubbs, Özlem Uzuner, Anupama E. Gururaj, Samuel Bayer, John S. Aberdeen, Anna Rumshisky, Serguei V. S. Pakhomov, Hua Xu 0001
J. Biomed. Informatics13
2015 Deciphering Signaling Pathway Networks to Understand the Molecular Mechanisms of Metformin Action
abstract
A drug exerts its effects typically through a signal transduction cascade, which is non-linear and involves intertwined networks of multiple signaling pathways. Construction of such a signaling pathway network (SPNetwork) can enable identification of novel drug targets and deep understanding of drug action. However, it is challenging to synopsize critical components of these interwoven pathways into one network. To tackle this issue, we developed a novel computational framework, the Drug-specific Signaling Pathway Network (DSPathNet). The DSPathNet amalgamates the prior drug knowledge and drug-induced gene expression via random walk algorithms. Using the drug metformin, we illustrated this framework and obtained one metformin-specific SPNetwork containing 477 nodes and 1,366 edges. To evaluate this network, we performed the gene set enrichment analysis using the disease genes of type 2 diabetes (T2D) and cancer, one T2D genome-wide association study (GWAS) dataset, three cancer GWAS datasets, and one GWAS dataset of cancer patients with T2D on metformin. The results showed that the metformin network was significantly enriched with disease genes for both T2D and cancer, and that the network also included genes that may be associated with metformin-associated cancer survival. Furthermore, from the metformin SPNetwork and common genes to T2D and cancer, we generated a subnetwork to highlight the molecule crosstalk between T2D and cancer. The follow-up network analyses and literature mining revealed that seven genes (CDKN1A, ESR1, MAX, MYC, PPARGC1A, SP1, and STK11) and one novel MYC-centered pathway with CDKN1A, SP1, and STK11 might play important roles in metformin's antidiabetic and anticancer effects. Some results are supported by previous studies. In summary, our study 1) develops a novel framework to construct drug-specific signal transduction networks; 2) provides insights into the molecular mode of metformin; 3) serves a model for exploring signaling pathways to facilitate understanding of drug action, disease pathogenesis, and identification of drug targets.
Jingchun Sun, Min Zhao 0006, Peilin Jia, Lily Wang 0001, Yonghui Wu 0001, Carissa Iverson, Yubo Zhou, Erica A. Bowton, Dan M. Roden, Joshua C. Denny, Melinda Aldrich, Hua Xu 0001, Zhongming Zhao
PLoS Comput. Biol.12
2014 Automated Assessment of Medical Students' Clinical Exposures according to AAMC Geriatric Competencies
Yukun Chen 0001, Jesse O. Wrenn, Hua Xu 0001, Anderson Spickard III, Ralf Habermann, James S. Powers, Joshua C. Denny
AMIA3
2014 A Preliminary Study of Coupling Transfer Learning with Active Learning for Clinical Named Entity Recognition between Two Institutions
Yukun Chen 0001, Yaoyun Zhang, Qiaozhu Mei, Dead Account, Joshua C. Denny, Hua Xu 0001
AMIA6
2014 Building a Treebank of hospital discharge summaries
Yang Huang 0008, Jungwei Fan 0001, Elly W. Yang, Hua Xu 0001
AMIA4
2014 Applying Active Learning to Word Sense Disambiguation in a Real-Time Setting
Sungrim Moon, Yukun Chen 0001, Joshua C. Denny, Hua Xu 0001
AMIA5
2014 A study of synonym extraction from clinical texts using semantic vector models
Sungrim Moon, Trevor Cohen, Hua Xu 0001
AMIA3
2014 Identifying Plausible Adverse Drug Reactions Using Knowledge Extracted from the Literature
Ning Shang 0004, Hua Xu 0001, Thomas C. Rindflesch, Trevor Cohen
AMIA2
2014 Identifying Metastases from Pathology Reports in Lung Cancer Patients
Ergin Soysal, Jeremy L. Warner, Joshua C. Denny, Hua Xu 0001
AMIA4
2014 An Integrative Framework for Drug Target Prediction and Repurposing
Jingchun Sun, Cui Tao, Kevin W. Zhu, W. Jim Zheng, Hua Xu 0001
AMIA6
2014 Development of a Unified Computable Problem-Medication Knowledge base
Yonghui Wu 0001, Adam Wright, Hua Xu 0001, Allison B. McCoy, Dean F. Sittig
AMIA3
2014 Mining electronic health record data to detect drug-repurposing signals for cancers
Hua Xu 0001, Qingxia Chen, Jeremy L. Warner, Min Jiang 0007, Anushi Shah, Melinda Aldrich, Joshua C. Denny
AMIA1
2014 Domain Adaptation for Semantic Role Labeling of Clinical Text
Yaoyun Zhang, Buzhou Tang, Min Jiang 0007, Yonghui Wu 0001, Hua Xu 0001
AMIA6
2014 DTMBIO 2014: International Workshop on Data and Text Mining in Biomedical Informatics
abstract
Held each year in conjunction with one of the largest data management conferences, CIKM, the Eighth ACM International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 14) is organized to bring together researchers interested in development and application of cutting-edge biomedical and healthcare technology. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 14 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research.
Luonan Chen, Doheon Lee, Hua Xu 0001, Min Song 0001
CIKM3
2014 PhenDisco: phenotype discovery system for the database of genotypes and phenotypes
abstract
The database of genotypes and phenotypes (dbGaP) developed by the National Center for Biotechnology Information (NCBI) is a resource that contains information on various genome-wide association studies (GWAS) and is currently available via NCBI's dbGaP Entrez interface. The database is an important resource, providing GWAS data that can be used for new exploratory research or cross-study validation by authorized users. However, finding studies relevant to a particular phenotype of interest is challenging, as phenotype information is presented in a non-standardized way. To address this issue, we developed PhenDisco (phenotype discoverer), a new information retrieval system for dbGaP. PhenDisco consists of two main components: (1) text processing tools that standardize phenotype variables and study metadata, and (2) information retrieval tools that support queries from users and return ranked results. In a preliminary comparison involving 18 search scenarios, PhenDisco showed promising performance for both unranked and ranked search comparisons with dbGaP's search engine Entrez. The system can be accessed at http://pfindr.net.
Son Doan, Ko-Wei Lin, Mike Conway, Lucila Ohno-Machado, Alexander Hsieh, Stephanie Feudjio Feupe, Asher Garland, Mindy K. Ross, Xiaoqian Jiang, Seena Farzaneh, Rebecca Walker, Neda Alipanah, Hua Xu 0001, Hyeon-Eui Kim
J. Am. Medical Informatics Assoc.14
2014 Research and applications: Assisted annotation of medical free text using RapTAT
abstract
OBJECTIVE: To determine whether assisted annotation using interactive training can reduce the time required to annotate a clinical document corpus without introducing bias. MATERIALS AND METHODS: A tool, RapTAT, was designed to assist annotation by iteratively pre-annotating probable phrases of interest within a document, presenting the annotations to a reviewer for correction, and then using the corrected annotations for further machine learning-based training before pre-annotating subsequent documents. Annotators reviewed 404 clinical notes either manually or using RapTAT assistance for concepts related to quality of care during heart failure treatment. Notes were divided into 20 batches of 19-21 documents for iterative annotation and training. RESULTS: The number of correct RapTAT pre-annotations increased significantly and annotation time per batch decreased by ~50% over the course of annotation. Annotation rate increased from batch to batch for assisted but not manual reviewers. Pre-annotation F-measure increased from 0.5 to 0.6 to >0.80 (relative to both assisted reviewer and reference annotations) over the first three batches and more slowly thereafter. Overall inter-annotator agreement was significantly higher between RapTAT-assisted reviewers (0.89) than between manual reviewers (0.85). DISCUSSION: The tool reduced workload by decreasing the number of annotations needing to be added and helping reviewers to annotate at an increased rate. Agreement between the pre-annotations and reference standard, and agreement between the pre-annotations and assisted annotations, were similar throughout the annotation process, which suggests that pre-annotation did not introduce bias. CONCLUSIONS: Pre-annotations generated by a tool capable of interactive training can reduce the time required to create an annotated document corpus by up to 50%.
Glenn T. Gobbel, Jennifer H. Garvin, Ruth M. Reeves, Robert M. Cronin, Julia Heavirland, Jenifer Williams, Allison Weaver, Shrimalini Jayaramaraja, Dario A. Giuse, Theodore Speroff, Steven H. Brown, Hua Xu 0001, Michael E. Matheny
J. Am. Medical Informatics Assoc.12
2014 Research and applications: A comprehensive study of named entity recognition in Chinese clinical text
abstract
OBJECTIVE: Named entity recognition (NER) is one of the fundamental tasks in natural language processing. In the medical domain, there have been a number of studies on NER in English clinical notes; however, very limited NER research has been carried out on clinical notes written in Chinese. The goal of this study was to systematically investigate features and machine learning algorithms for NER in Chinese clinical text. MATERIALS AND METHODS: We randomly selected 400 admission notes and 400 discharge summaries from Peking Union Medical College Hospital in China. For each note, four types of entity-clinical problems, procedures, laboratory test, and medications-were annotated according to a predefined guideline. Two-thirds of the 400 notes were used to train the NER systems and one-third for testing. We investigated the effects of different types of feature including bag-of-characters, word segmentation, part-of-speech, and section information, and different machine learning algorithms including conditional random fields (CRF), support vector machines (SVM), maximum entropy (ME), and structural SVM (SSVM) on the Chinese clinical NER task. All classifiers were trained on the training dataset and evaluated on the test set, and micro-averaged precision, recall, and F-measure were reported. RESULTS: Our evaluation on the independent test set showed that most types of feature were beneficial to Chinese NER systems, although the improvements were limited. The system achieved the highest performance by combining word segmentation and section information, indicating that these two types of feature complement each other. When the same types of optimized feature were used, CRF and SSVM outperformed SVM and ME. More specifically, SSVM achieved the highest performance of the four algorithms, with F-measures of 93.51% and 90.01% for admission notes and discharge summaries, respectively.
Jianbo Lei, Buzhou Tang, Xueqin Lu, Kaihua Gao, Min Jiang 0007, Hua Xu 0001
J. Am. Medical Informatics Assoc.6
2014 Determining molecular predictors of adverse drug reactions with causality analysis based on structure learning
abstract
OBJECTIVE: Adverse drug reaction (ADR) can have dire consequences. However, our current understanding of the causes of drug-induced toxicity is still limited. Hence it is of paramount importance to determine molecular factors of adverse drug responses so that safer therapies can be designed. METHODS: We propose a causality analysis model based on structure learning (CASTLE) for identifying factors that contribute significantly to ADRs from an integration of chemical and biological properties of drugs. This study aims to address two major limitations of the existing ADR prediction studies. First, ADR prediction is mostly performed by assessing the correlations between the input features and ADRs, and the identified associations may not indicate causal relations. Second, most predictive models lack biological interpretability. RESULTS: CASTLE was evaluated in terms of prediction accuracy on 12 organ-specific ADRs using 830 approved drugs. The prediction was carried out by first extracting causal features with structure learning and then applying them to a support vector machine (SVM) for classification. Through rigorous experimental analyses, we observed significant increases in both macro and micro F1 scores compared with the traditional SVM classifier, from 0.88 to 0.89 and 0.74 to 0.81, respectively. Most importantly, identified links between the biological factors and organ-specific drug toxicities were partially supported by evidence in Online Mendelian Inheritance in Man. CONCLUSIONS: The proposed CASTLE model not only performed better in prediction than the baseline SVM but also produced more interpretable results (ie, biological factors responsible for ADRs), which is critical to discovering molecular activators of ADRs.
Ruichu Cai, Yong Hu 0002, Michael E. Matheny, Jingchun Sun, Hua Xu 0001
J. Am. Medical Informatics Assoc.7
2014 Identifying plausible adverse drug reactions using knowledge extracted from the literature
Ning Shang 0004, Hua Xu 0001, Thomas C. Rindflesch, Trevor Cohen
J. Biomed. Informatics2
2013 Informatics Infrastructure for Routine Personalized Medicine
Elmer V. Bernstam, Ken Chen 0001, Hua Xu 0001, Funda Meric-Bernstam
AMIA3
2013 A Study of Active Learning Methods for Clinical Entities Recognition
Yukun Chen 0001, Thomas A. Lasko, Qiaozhu Mei, Joshua C. Denny, Hua Xu 0001
AMIA5
2013 Ontology-Based Entity Extraction of Quality Metrics from Narrative Texts
Sina Madani, Dean F. Sittig, Hua Xu 0001, Parsa Mirhaji, Kim Dunn, Reza Alemy
AMIA3
2013 Word Sense Disambiguation of Clinical Abbreviations with Hyperdimensional Computing
Sungrim Moon, Bjoern-Toby Berster, Hua Xu 0001, Trevor Cohen
AMIA3
2013 Building a Large Clinical Abbreviation Sense Inventory from Discharge Summaries
Yonghui Wu 0001, S. Trent Rosenbloom, Joshua C. Denny, Randolph A. Miller, Dario A. Giuse, Hua Xu 0001
AMIA6
2013 Understanding Consumers' Information Needs about Breast Cancer by Analyzing Online Questions
Yaoyun Zhang, Hua Xu 0001
AMIA2
2013 Correlating adverse drug reactions with biological pathways in humans
abstract
It has been well recognized that adverse drug reactions (ADRs) are a significant cause of morbidity and mortality. There is a growing interest in investigating biological pathways involved in cellular response to drugs. Based on examining the co-occurrence of drugs in pathway activity and ADR profiles, in this paper, we propose a new method to explore the relationship between biological pathways and ADRs at a large scale. Using sparse canonical correlation analysis of 495 drugs with two profiles for 173 pathways and 1385 ADRs, a total of 80 correlated sets of pathways and ADRs were extracted. To evaluate the performance of our method, extracted correlated components were used to retrieve known ADR profiles from drug pathway profiles using a 5-fold cross validation. A relatively high prediction performance (AUC: 0.881) was achieved. This work provides a foundation for future investigation of ADRs in the context of biological pathways under different conditions.
Huiru Zheng, Haiying Wang 0001, Hua Xu 0001, Zhongming Zhao, Francisco Azuaje
BIBM3
2013 Messaging to your doctors: understanding patient-provider communications via a portal system
abstract
The patient portal is a relatively new healthcare information technology that enables patients more convenient access to their healthcare information and allows them to send messages to their doctors. Our study examines the themes discussed in these messages and the different ways in which patients communicate with their providers via a portal employed in a large medical center. We also explore the differences between the patient portal and more traditional communication media, and investigated the advantages and potential problems of the portal system. Our findings show a wide variety of topics discussed in the communication messages (such as medication, appointments, laboratory tests, etc.) and how patients provide information, consult with their providers, and express psychosocial and emotional needs. We argue that the patient portal improves the accuracy of communication and could facilitate illness management for patients, especially over a longer term. However, messaging through the patient portal is not popular among patients and the simultaneous use of multiple communication media may create information gaps. More research is needed to better elucidate barriers to the use of patient portals and the optimal methods of communication and information integration given different contexts.
Si Sun, Xiaomu Zhou, Joshua C. Denny, S. Trent Rosenbloom, Hua Xu 0001
CHI5
2013 DTMBIO 2013: international workshop on data and text mining in biomedical informatics
abstract
The organizers of ACM Seventh International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 13) are pleased to announce that the seventh DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 13 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research.
Atul J. Butte, Doheon Lee, Hua Xu 0001, Min Song 0001
CIKM3
2013 Applying active learning to supervised word sense disambiguation in MEDLINE
abstract
OBJECTIVES: This study was to assess whether active learning strategies can be integrated with supervised word sense disambiguation (WSD) methods, thus reducing the number of annotated samples, while keeping or improving the quality of disambiguation models. METHODS: We developed support vector machine (SVM) classifiers to disambiguate 197 ambiguous terms and abbreviations in the MSH WSD collection. Three different uncertainty sampling-based active learning algorithms were implemented with the SVM classifiers and were compared with a passive learner (PL) based on random sampling. For each ambiguous term and each learning algorithm, a learning curve that plots the accuracy computed from the test set as a function of the number of annotated samples used in the model was generated. The area under the learning curve (ALC) was used as the primary metric for evaluation. RESULTS: Our experiments demonstrated that active learners (ALs) significantly outperformed the PL, showing better performance for 177 out of 197 (89.8%) WSD tasks. Further analysis showed that to achieve an average accuracy of 90%, the PL needed 38 annotated samples, while the ALs needed only 24, a 37% reduction in annotation effort. Moreover, we analyzed cases where active learning algorithms did not achieve superior performance and identified three causes: (1) poor models in the early learning stage; (2) easy WSD cases; and (3) difficult WSD cases, which provide useful insight for future improvements. CONCLUSIONS: This study demonstrated that integrating active learning strategies with supervised WSD methods could effectively reduce annotation cost and improve the disambiguation models.
Yukun Chen 0001, Hongxin Cao, Qiaozhu Mei, Kai Zheng 0002, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2013 Research and applications: Syntactic parsing of clinical text: guideline and corpus development with handling ill-formed sentences
abstract
OBJECTIVE: To develop, evaluate, and share: (1) syntactic parsing guidelines for clinical text, with a new approach to handling ill-formed sentences; and (2) a clinical Treebank annotated according to the guidelines. To document the process and findings for readers with similar interest. METHODS: Using random samples from a shared natural language processing challenge dataset, we developed a handbook of domain-customized syntactic parsing guidelines based on iterative annotation and adjudication between two institutions. Special considerations were incorporated into the guidelines for handling ill-formed sentences, which are common in clinical text. Intra- and inter-annotator agreement rates were used to evaluate consistency in following the guidelines. Quantitative and qualitative properties of the annotated Treebank, as well as its use to retrain a statistical parser, were reported. RESULTS: A supplement to the Penn Treebank II guidelines was developed for annotating clinical sentences. After three iterations of annotation and adjudication on 450 sentences, the annotators reached an F-measure agreement rate of 0.930 (while intra-annotator rate was 0.948) on a final independent set. A total of 1100 sentences from progress notes were annotated that demonstrated domain-specific linguistic features. A statistical parser retrained with combined general English (mainly news text) annotations and our annotations achieved an accuracy of 0.811 (higher than models trained purely with either general or clinical sentences alone). Both the guidelines and syntactic annotations are made available at https://sourceforge.net/projects/medicaltreebank. CONCLUSIONS: We developed guidelines for parsing clinical text and annotated a corpus accordingly. The high intra- and inter-annotator agreement rates showed decent consistency in following the guidelines. The corpus was shown to be useful in retraining a statistical parser that achieved moderate accuracy.
Jungwei Fan 0001, Elly W. Yang, Min Jiang 0007, Rashmi Prasad, Richard M. Loomis, Daniel Zisook, Joshua C. Denny, Hua Xu 0001, Yang Huang 0008
J. Am. Medical Informatics Assoc.8
2013 Comparative analysis of pharmacovigilance methods in the detection of adverse drug reactions using electronic medical records
abstract
OBJECTIVE: Medication safety requires that each drug be monitored throughout its market life as early detection of adverse drug reactions (ADRs) can lead to alerts that prevent patient harm. Recently, electronic medical records (EMRs) have emerged as a valuable resource for pharmacovigilance. This study examines the use of retrospective medication orders and inpatient laboratory results documented in the EMR to identify ADRs. METHODS: Using 12 years of EMR data from Vanderbilt University Medical Center (VUMC), we designed a study to correlate abnormal laboratory results with specific drug administrations by comparing the outcomes of a drug-exposed group and a matched unexposed group. We assessed the relative merits of six pharmacovigilance measures used in spontaneous reporting systems (SRSs): proportional reporting ratio (PRR), reporting OR (ROR), Yule's Q (YULE), the χ(2) test (CHI), Bayesian confidence propagation neural networks (BCPNN), and a gamma Poisson shrinker (GPS). RESULTS: We systematically evaluated the methods on two independently constructed reference standard datasets of drug-event pairs. The dataset of Yoon et al contained 470 drug-event pairs (10 drugs and 47 laboratory abnormalities). Using VUMC's EMR, we created another dataset of 378 drug-event pairs (nine drugs and 42 laboratory abnormalities). Evaluation on our reference standard showed that CHI, ROR, PRR, and YULE all had the same F score (62%). When the reference standard of Yoon et al was used, ROR had the best F score of 68%, with 77% precision and 61% recall. CONCLUSIONS: Results suggest that EMR-derived laboratory measurements and medication orders can help to validate previously reported ADRs, and detect new ADRs.
Eugenia R. McPeek Hinz, Michael E. Matheny, Joshua C. Denny, Jonathan S. Schildcrout, Randolph A. Miller, Hua Xu 0001
J. Am. Medical Informatics Assoc.7
2013 Research and applications: Machine learning for predicting the response of breast cancer to neoadjuvant chemotherapy
abstract
OBJECTIVE: To employ machine learning methods to predict the eventual therapeutic response of breast cancer patients after a single cycle of neoadjuvant chemotherapy (NAC). MATERIALS AND METHODS: Quantitative dynamic contrast-enhanced MRI and diffusion-weighted MRI data were acquired on 28 patients before and after one cycle of NAC. A total of 118 semiquantitative and quantitative parameters were derived from these data and combined with 11 clinical variables. We used Bayesian logistic regression in combination with feature selection using a machine learning framework for predictive model building. RESULTS: The best predictive models using feature selection obtained an area under the curve of 0.86 and an accuracy of 0.86, with a sensitivity of 0.88 and a specificity of 0.82. DISCUSSION: With the numerous options for NAC available, development of a method to predict response early in the course of therapy is needed. Unfortunately, by the time most patients are found not to be responding, their disease may no longer be surgically resectable, and this situation could be avoided by the development of techniques to assess response earlier in the treatment regimen. The method outlined here is one possible solution to this important clinical problem. CONCLUSIONS: Predictive modeling approaches based on machine learning using readily available clinical and quantitative MRI data show promise in distinguishing breast cancer responders from non-responders after the first cycle of NAC.
Subramani Mani, Yukun Chen 0001, Lori R. Arlinghaus, A. Bapsi Chakravarthy, Vandana G. Abramson, Sandeep R. Bhave, Mia A. Levy, Hua Xu 0001, Thomas E. Yankeelov
J. Am. Medical Informatics Assoc.9
2013 A hybrid system for temporal information extraction from clinical text
abstract
OBJECTIVE: To develop a comprehensive temporal information extraction system that can identify events, temporal expressions, and their temporal relations in clinical text. This project was part of the 2012 i2b2 clinical natural language processing (NLP) challenge on temporal information extraction. MATERIALS AND METHODS: The 2012 i2b2 NLP challenge organizers manually annotated 310 clinic notes according to a defined annotation guideline: a training set of 190 notes and a test set of 120 notes. All participating systems were developed on the training set and evaluated on the test set. Our system consists of three modules: event extraction, temporal expression extraction, and temporal relation (also called Temporal Link, or 'TLink') extraction. The TLink extraction module contains three individual classifiers for TLinks: (1) between events and section times, (2) within a sentence, and (3) across different sentences. The performance of our system was evaluated using scripts provided by the i2b2 organizers. Primary measures were micro-averaged Precision, Recall, and F-measure. RESULTS: Our system was among the top ranked. It achieved F-measures of 0.8659 for temporal expression extraction (ranked fourth), 0.6278 for end-to-end TLink track (ranked first), and 0.6932 for TLink-only track (ranked first) in the challenge. We subsequently investigated different strategies for TLink extraction, and were able to marginally improve performance with an F-measure of 0.6943 for TLink-only track.
Buzhou Tang, Yonghui Wu 0001, Min Jiang 0007, Yukun Chen 0001, Joshua C. Denny, Hua Xu 0001
J. Am. Medical Informatics Assoc.6
2013 Development and evaluation of an ensemble resource linking medications to their indications
abstract
OBJECTIVE: To create a computable MEDication Indication resource (MEDI) to support primary and secondary use of electronic medical records (EMRs). MATERIALS AND METHODS: We processed four public medication resources, RxNorm, Side Effect Resource (SIDER) 2, MedlinePlus, and Wikipedia, to create MEDI. We applied natural language processing and ontology relationships to extract indications for prescribable, single-ingredient medication concepts and all ingredient concepts as defined by RxNorm. Indications were coded as Unified Medical Language System (UMLS) concepts and International Classification of Diseases, 9th edition (ICD9) codes. A total of 689 extracted indications were randomly selected for manual review for accuracy using dual-physician review. We identified a subset of medication-indication pairs that optimizes recall while maintaining high precision. RESULTS: MEDI contains 3112 medications and 63 343 medication-indication pairs. Wikipedia was the largest resource, with 2608 medications and 34 911 pairs. For each resource, estimated precision and recall, respectively, were 94% and 20% for RxNorm, 75% and 33% for MedlinePlus, 67% and 31% for SIDER 2, and 56% and 51% for Wikipedia. The MEDI high-precision subset (MEDI-HPS) includes indications found within either RxNorm or at least two of the three other resources. MEDI-HPS contains 13 304 unique indication pairs regarding 2136 medications. The mean±SD number of indications for each medication in MEDI-HPS is 6.22 ± 6.09. The estimated precision of MEDI-HPS is 92%. CONCLUSIONS: MEDI is a publicly available, computable resource that links medications with their indications as represented by concepts and billing codes. MEDI may benefit clinical EMR applications and reuse of EMR data for research.
Wei-Qi Wei, Robert M. Cronin, Hua Xu 0001, Thomas A. Lasko, Lisa Bastarache, Joshua C. Denny
J. Am. Medical Informatics Assoc.3
2013 Research and applications: ICD-9 tobacco use codes are effective identifiers of smoking status
abstract
OBJECTIVE: To evaluate the validity of, characterize the usage of, and propose potential research applications for International Classification of Diseases, Ninth Revision (ICD-9) tobacco codes in clinical populations. MATERIALS AND METHODS: Using data on cancer cases and cancer-free controls from Vanderbilt's biorepository, BioVU, we evaluated the utility of ICD-9 tobacco use codes to identify ever-smokers in general and high smoking prevalence (lung cancer) clinic populations. We assessed potential biases in documentation, and performed temporal analysis relating transitions between smoking codes to smoking cessation attempts. We also examined the suitability of these codes for use in genetic association analyses. RESULTS: ICD-9 tobacco use codes can identify smokers in a general clinic population (specificity of 1, sensitivity of 0.32), and there is little evidence of documentation bias. Frequency of code transitions between 'current' and 'former' tobacco use was significantly correlated with initial success at smoking cessation (p<0.0001). Finally, code-based smoking status assignment is a comparable covariate to text-based smoking status for genetic association studies. DISCUSSION: Our results support the use of ICD-9 tobacco use codes for identifying smokers in a clinical population. Furthermore, with some limitations, these codes are suitable for adjustment of smoking status in genetic studies utilizing electronic health records. CONCLUSIONS: Researchers should not be deterred by the unavailability of full-text records to determine smoking status if they have ICD-9 code histories.
Laura K. Wiley, Anushi Shah, Hua Xu 0001, William S. Bush
J. Am. Medical Informatics Assoc.3
2012 Extracting Semantic Lexicons from Discharge Summaries using Machine Learning and the C-Value Method
Min Jiang 0007, Joshua C. Denny, Buzhou Tang, Hongxin Cao, Hua Xu 0001
AMIA5
2012 A Study of Transportability of an Existing Smoking Status Detection Module across Institutions
Anushi Shah, Min Jiang 0007, Neeraja B. Peterson, Melinda Aldrich, Qingxia Chen, Erica A. Bowton, Joshua C. Denny, Hua Xu 0001
AMIA11
2012 MedEx-UIMA - An Open-Source System for Medication Information Extraction from Clinical Text
Anushi Shah, Min Jiang 0007, Yonghui Wu 0001, Joshua C. Denny, Hua Xu 0001
AMIA5
2012 Understanding Patient-Provider Communication via a Patient Portal
Si Sun, Xiaomu Zhou, Hua Xu 0001, Joshua C. Denny
AMIA3
2012 Clinical Entity Recognition Using Structural Support Vector Machines
Buzhou Tang, Yonghui Wu 0001, Min Jiang 0007, Hua Xu 0001
AMIA4
2012 A comparative study of current clinical natural language processing systems on handling abbreviations in discharge summaries
Yonghui Wu 0001, Joshua C. Denny, S. Trent Rosenbloom, Randolph A. Miller, Dario A. Giuse, Hua Xu 0001
AMIA6
2012 Electronic health record data suggests metformin improves cancer survival: A new model for drug repurposing studies
Hua Xu 0001, Melinda Aldrich, Qingxia Chen, Neeraja B. Peterson, Mia A. Levy, Anushi Shah, Carol Friedman, Joshua C. Denny
AMIA1
2012 Combining Corpus-derived Sense Profiles with Estimated Frequency Information to Disambiguate Clinical Abbreviations
Hua Xu 0001, Peter D. Stetson, Carol Friedman
AMIA1
2012 DTMBIO 2012: international workshop on data and text mining in biomedical informatics
abstract
The organizers of ACM Sixth International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 12) are happy announce that the sixth DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 12 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research.
Min Song 0001, Doheon Lee, Hua Xu 0001, Sophia Ananiadou
CIKM3
2012 DTome: a web-based tool for drug-target interactome construction
abstract
BACKGROUND: Understanding drug bioactivities is crucial for early-stage drug discovery, toxicology studies and clinical trials. Network pharmacology is a promising approach to better understand the molecular mechanisms of drug bioactivities. With a dramatic increase of rich data sources that document drugs' structural, chemical, and biological activities, it is necessary to develop an automated tool to construct a drug-target network for candidate drugs, thus facilitating the drug discovery process. RESULTS: We designed a computational workflow to construct drug-target networks from different knowledge bases including DrugBank, PharmGKB, and the PINA database. To automatically implement the workflow, we created a web-based tool called DTome (Drug-Target interactome tool), which is comprised of a database schema and a user-friendly web interface. The DTome tool utilizes web-based queries to search candidate drugs and then construct a DTome network by extracting and integrating four types of interactions. The four types are adverse drug interactions, drug-target interactions, drug-gene associations, and target-/gene-protein interactions. Additionally, we provided a detailed network analysis and visualization process to illustrate how to analyze and interpret the DTome network. The DTome tool is publicly available at http://bioinfo.mc.vanderbilt.edu/DTome. CONCLUSIONS: As demonstrated with the antipsychotic drug clozapine, the DTome tool was effective and promising for the investigation of relationships among drugs, adverse interaction drugs, drug primary targets, drug-associated genes, and proteins directly interacting with targets or genes. The resultant DTome network provides researchers with direct insights into their interest drug(s), such as the molecular mechanisms of drug actions. We believe such a tool can facilitate identification of drug targets and drug adverse interactions.
Jingchun Sun, Yonghui Wu 0001, Hua Xu 0001, Zhongming Zhao
BMC Bioinform.3
2012 Portability of an algorithm to identify rheumatoid arthritis in electronic health records
abstract
OBJECTIVES: Electronic health records (EHR) can allow for the generation of large cohorts of individuals with given diseases for clinical and genomic research. A rate-limiting step is the development of electronic phenotype selection algorithms to find such cohorts. This study evaluated the portability of a published phenotype algorithm to identify rheumatoid arthritis (RA) patients from EHR records at three institutions with different EHR systems. MATERIALS AND METHODS: Physicians reviewed charts from three institutions to identify patients with RA. Each institution compiled attributes from various sources in the EHR, including codified data and clinical narratives, which were searched using one of two natural language processing (NLP) systems. The performance of the published model was compared with locally retrained models. RESULTS: Applying the previously published model from Partners Healthcare to datasets from Northwestern and Vanderbilt Universities, the area under the receiver operating characteristic curve was found to be 92% for Northwestern and 95% for Vanderbilt, compared with 97% at Partners. Retraining the model improved the average sensitivity at a specificity of 97% to 72% from the original 65%. Both the original logistic regression models and locally retrained models were superior to simple billing code count thresholds. DISCUSSION: These results show that a previously published algorithm for RA is portable to two external hospitals using different EHR systems, different NLP systems, and different target NLP vocabularies. Retraining the algorithm primarily increased the sensitivity at each site. CONCLUSION: Electronic phenotype algorithms allow rapid identification of case populations in multiple sites with little retraining.
Robert J. Carroll, William K. Thompson, Anne E. Eyler, Arthur M. Mandelin, Tianxi Cai, Raquel M. Zink, Jennifer A. Pacheco, Chad S. Boomershine, Thomas A. Lasko, Hua Xu 0001, Elizabeth W. Karlson, Raúl G. Pérez, Vivian S. Gainer, Shawn N. Murphy, Eric M. Ruderman, Richard M. Pope, Robert M. Plenge, Abel N. Kho, Katherine P. Liao, Joshua C. Denny
J. Am. Medical Informatics Assoc.10
2012 Large-scale prediction of adverse drug reactions using chemical, biological, and phenotypic properties of drugs
abstract
OBJECTIVE: Adverse drug reaction (ADR) is one of the major causes of failure in drug development. Severe ADRs that go undetected until the post-marketing phase of a drug often lead to patient morbidity. Accurate prediction of potential ADRs is required in the entire life cycle of a drug, including early stages of drug design, different phases of clinical trials, and post-marketing surveillance. METHODS: Many studies have utilized either chemical structures or molecular pathways of the drugs to predict ADRs. Here, the authors propose a machine-learning-based approach for ADR prediction by integrating the phenotypic characteristics of a drug, including indications and other known ADRs, with the drug's chemical structures and biological properties, including protein targets and pathway information. A large-scale study was conducted to predict 1385 known ADRs of 832 approved drugs, and five machine-learning algorithms for this task were compared. RESULTS: This evaluation, based on a fivefold cross-validation, showed that the support vector machine algorithm outperformed the others. Of the three types of information, phenotypic data were the most informative for ADR prediction. When biological and phenotypic features were added to the baseline chemical information, the ADR prediction model achieved significant improvements in area under the curve (from 0.9054 to 0.9524), precision (from 43.37% to 66.17%), and recall (from 49.25% to 63.06%). Most importantly, the proposed model successfully predicted the ADRs associated with withdrawal of rofecoxib and cerivastatin. CONCLUSION: The results suggest that phenotypic information on drugs is valuable for ADR prediction. Moreover, they demonstrate that different models that combine chemical, biological, or phenotypic information can be built from approved drugs, and they have the potential to detect clinically important ADRs in both preclinical and post-marketing phases.
Yonghui Wu 0001, Yukun Chen 0001, Jingchun Sun, Zhongming Zhao, Xue-wen Chen 0001, Michael E. Matheny, Hua Xu 0001
J. Am. Medical Informatics Assoc.8
2012 Applying active learning to assertion classification of concepts in clinical text
Yukun Chen 0001, Subramani Mani, Hua Xu 0001
J. Biomed. Informatics3
2012 A new clustering method for detecting rare senses of abbreviations in clinical notes
Hua Xu 0001, Yonghui Wu 0001, Noémie Elhadad, Peter D. Stetson, Carol Friedman
J. Biomed. Informatics1
2011 A study of machine-learning-based approaches to extract clinical entities and their assertions from discharge summaries
abstract
OBJECTIVE: The authors' goal was to develop and evaluate machine-learning-based approaches to extracting clinical entities-including medical problems, tests, and treatments, as well as their asserted status-from hospital discharge summaries written using natural language. This project was part of the 2010 Center of Informatics for Integrating Biology and the Bedside/Veterans Affairs (VA) natural-language-processing challenge. DESIGN: The authors implemented a machine-learning-based named entity recognition system for clinical text and systematically evaluated the contributions of different types of features and ML algorithms, using a training corpus of 349 annotated notes. Based on the results from training data, the authors developed a novel hybrid clinical entity extraction system, which integrated heuristic rule-based modules with the ML-base named entity recognition module. The authors applied the hybrid system to the concept extraction and assertion classification tasks in the challenge and evaluated its performance using a test data set with 477 annotated notes. MEASUREMENTS: Standard measures including precision, recall, and F-measure were calculated using the evaluation script provided by the Center of Informatics for Integrating Biology and the Bedside/VA challenge organizers. The overall performance for all three types of clinical entities and all six types of assertions across 477 annotated notes were considered as the primary metric in the challenge. RESULTS AND DISCUSSION: Systematic evaluation on the training set showed that Conditional Random Fields outperformed Support Vector Machines, and semantic information from existing natural-language-processing systems largely improved performance, although contributions from different types of features varied. The authors' hybrid entity extraction system achieved a maximum overall F-score of 0.8391 for concept extraction (ranked second) and 0.9313 for assertion classification (ranked fourth, but not statistically different than the first three systems) on the test data set in the challenge.
Min Jiang 0007, Yukun Chen 0001, S. Trent Rosenbloom, Subramani Mani, Joshua C. Denny, Hua Xu 0001
J. Am. Medical Informatics Assoc.7
2011 Data from clinical notes: a perspective on the tension between structure and flexible documentation
abstract
Clinical documentation is central to patient care. The success of electronic health record system adoption may depend on how well such systems support clinical documentation. A major goal of integrating clinical documentation into electronic heath record systems is to generate reusable data. As a result, there has been an emphasis on deploying computer-based documentation systems that prioritize direct structured documentation. Research has demonstrated that healthcare providers value different factors when writing clinical notes, such as narrative expressivity, amenability to the existing workflow, and usability. The authors explore the tension between expressivity and structured clinical documentation, review methods for obtaining reusable data from clinical notes, and recommend that healthcare providers be able to choose how to document patient care based on workflow and note content needs. When reusable data are needed from notes, providers can use structured documentation or rely on post-hoc text processing to produce structured data, as appropriate.
S. Trent Rosenbloom, Joshua C. Denny, Hua Xu 0001, Nancy M. Lorenzi, William W. Stead, Kevin B. Johnson
J. Am. Medical Informatics Assoc.3
2011 Facilitating pharmacogenetic studies using electronic health records and natural-language processing: a case study of warfarin
abstract
OBJECTIVE: DNA biobanks linked to comprehensive electronic health records systems are potentially powerful resources for pharmacogenetic studies. This study sought to develop natural-language-processing algorithms to extract drug-dose information from clinical text, and to assess the capabilities of such tools to automate the data-extraction process for pharmacogenetic studies. MATERIALS AND METHODS: A manually validated warfarin pharmacogenetic study identified a cohort of 1125 patients with a stable warfarin dose, in which 776 patients were managed by Coumadin Clinic physicians, and the remaining 349 patients were managed by their providers. The authors developed two algorithms to extract weekly warfarin doses from both data sets: a regular expression-based program for semistructured Coumadin Clinic notes; and an advanced weekly dose calculator based on an existing medication information extraction system (MedEx) for narrative providers' notes. The authors then conducted an association analysis between an automatically extracted stable weekly dose of warfarin and four genetic variants of VKORC1 and CYP2C9 genes. The performance of the weekly dose-extraction program was evaluated by comparing it with a gold standard containing manually curated weekly doses. Precision, recall, F-measure, and overall accuracy were reported. Associations between known variants in VKORC1 and CYP2C9 and warfarin stable weekly dose were performed with linear regression adjusted for age, gender, and body mass index. RESULTS: The authors' evaluation showed that the MedEx-based system could determine patients' warfarin weekly doses with 99.7% recall, 90.8% precision, and 93.8% accuracy. Using the automatically extracted weekly doses of warfarin, the authors successfully replicated the previous known associations between warfarin stable dose and genetic variants in VKORC1 and CYP2C9.
Hua Xu 0001, Min Jiang 0007, Matthew Oetjens, Erica A. Bowton, Andrea H. Ramirez, Janina M. Jeff, Melissa A. Basford, Jill M. Pulley, James D. Cowan, Marylyn D. Ritchie, Daniel R. Masys, Dan M. Roden, Dana C. Crawford, Joshua C. Denny
J. Am. Medical Informatics Assoc.1
2011 Applying semantic-based probabilistic context-free grammar to medical language processing - A preliminary study on parsing medication sentences
Hua Xu 0001, Samir AbdelRahman, Yanxin Lu, Joshua C. Denny, Son Doan
J. Biomed. Informatics1
2010 Identifying potential drugs that induce QT prolongation using electronic medical records
abstract
Table 1 Potential drugs that prolong QT interval with significance level of 0.001 Drug Chi-square Evidence Amiodarone 39.21 Known reaction Potassium supplements 24.78 Treatmenttypically given to people with long QT intervals to keep it normal Procainamide 22.11 Known reaction Sotalol 21.62 Known reaction Warfarin 18.42 No evidence found Meperidine 18.13 No evidence found Oxycodone 17.08 No evidence found Promethazine 12.90 No evidence found
Joshua C. Denny, Subramani Mani, Yukun Chen 0001, Yong Hu 0002, Hua Xu 0001
BMC Bioinform.6
2010 Extracting timing and status descriptors for colonoscopy testing from electronic medical records
abstract
Colorectal cancer (CRC) screening rates are low despite confirmed benefits. The authors investigated the use of natural language processing (NLP) to identify previous colonoscopy screening in electronic records from a random sample of 200 patients at least 50 years old. The authors developed algorithms to recognize temporal expressions and 'status indicators', such as 'patient refused', or 'test scheduled'. The new methods were added to the existing KnowledgeMap concept identifier system, and the resulting system was used to parse electronic medical records (EMR) to detect completed colonoscopies. Using as the 'gold standard' expert physicians' manual review of EMR notes, the system identified timing references with a recall of 0.91 and precision of 0.95, colonoscopy status indicators with a recall of 0.82 and precision of 0.95, and references to actually completed colonoscopies with recall of 0.93 and precision of 0.95. The system was superior to using colonoscopy billing codes alone. Health services researchers and clinicians may find NLP a useful adjunct to traditional methods to detect CRC screening status. Further investigations must validate extension of NLP approaches for other types of CRC screening applications.
Joshua C. Denny, Josh F. Peterson, Neesha N. Choma, Hua Xu 0001, Randolph A. Miller, Lisa Bastarache, Neeraja B. Peterson
J. Am. Medical Informatics Assoc.4
2010 Integrating existing natural language processing tools for medication extraction from discharge summaries
abstract
OBJECTIVE: To develop an automated system to extract medications and related information from discharge summaries as part of the 2009 i2b2 natural language processing (NLP) challenge. This task required accurate recognition of medication name, dosage, mode, frequency, duration, and reason for drug administration. DESIGN: We developed an integrated system using several existing NLP components developed at Vanderbilt University Medical Center, which included MedEx (to extract medication information), SecTag (a section identification system for clinical notes), a sentence splitter, and a spell checker for drug names. Our goal was to achieve good performance with minimal to no specific training for this document corpus; thus, evaluating the portability of those NLP tools beyond their home institution. The integrated system was developed using 17 notes that were annotated by the organizers and evaluated using 251 notes that were annotated by participating teams. MEASUREMENTS: The i2b2 challenge used standard measures, including precision, recall, and F-measure, to evaluate the performance of participating systems. There were two ways to determine whether an extracted textual finding is correct or not: exact matching or inexact matching. The overall performance for all six types of medication-related findings across 251 annotated notes was considered as the primary metric in the challenge. RESULTS: Our system achieved an overall F-measure of 0.821 for exact matching (0.839 precision; 0.803 recall) and 0.822 for inexact matching (0.866 precision; 0.782 recall). The system ranked second out of 20 participating teams on overall performance at extracting medications and related information. CONCLUSIONS: The results show that the existing MedEx system, together with other NLP components, can extract medication information in clinical text from institutions other than the site of algorithm development with reasonable performance.
Son Doan, Lisa Bastarache, Sergio Klimkowski, Joshua C. Denny, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2010 Application of information technology: MedEx: a medication information extraction system for clinical narratives
abstract
Medication information is one of the most important types of clinical data in electronic medical records. It is critical for healthcare safety and quality, as well as for clinical research that uses electronic medical record data. However, medication data are often recorded in clinical notes as free-text. As such, they are not accessible to other computerized applications that rely on coded data. We describe a new natural language processing system (MedEx), which extracts medication information from clinical notes. MedEx was initially developed using discharge summaries. An evaluation using a data set of 50 discharge summaries showed it performed well on identifying not only drug names (F-measure 93.2%), but also signature information, such as strength, route, and frequency, with F-measures of 94.5%, 93.9%, and 96.0% respectively. We then applied MedEx unchanged to outpatient clinic visit notes. It performed similarly with F-measures over 90% on a set of 25 clinic visit notes.
Hua Xu 0001, Shane P. Stenner, Son Doan, Kevin B. Johnson, Lemuel R. Waitman, Joshua C. Denny
J. Am. Medical Informatics Assoc.1
2009 Development of a Natural Language Processing System to Identify Timing and Status of Colonoscopy Testing in Electronic Medical Records
Joshua C. Denny, Josh F. Peterson, Neesha N. Choma, Hua Xu 0001, Randolph A. Miller, Lisa Bastarache, Neeraja B. Peterson
AMIA4
2009 Research Paper: Methods for Building Sense Inventories of Abbreviations in Clinical Notes
abstract
OBJECTIVE: To develop methods for building corpus-specific sense inventories of abbreviations occurring in clinical documents. DESIGN: A corpus of internal medicine admission notes was collected and instances of each clinical abbreviation in the corpus were clustered to different sense clusters. One instance from each cluster was manually annotated to generate a final list of senses. Two clustering-based methods (Expectation Maximization--EM and Farthest First--FF) and one random sampling method for sense detection were evaluated using a set of 12 clinical abbreviations. MEASUREMENTS: The clustering-based sense detection methods were evaluated using a set of clinical abbreviations that were manually sense annotated. "Sense Completeness" and "Annotation Cost" were used to measure the performance of different methods. Clustering error rates were also reported for different clustering algorithms. RESULTS: A clustering-based semi-automated method was developed to build corpus-specific sense inventories for abbreviations in hospital admission notes. Evaluation demonstrated that this method could largely reduce manual annotation cost and increase the completeness of sense inventories when compared with a manual annotation method using random samples. CONCLUSION: The authors developed an effective clustering-based method for building corpus-specific sense inventories for abbreviations in a clinical corpus. To the best of the authors knowledge, this is the first time clustering technologies have been used to help building sense inventories of abbreviations in clinical text. The results demonstrated that the clustering-based method performed better than the manual annotation method using random samples for the task of building sense inventories of clinical abbreviations.
Hua Xu 0001, Peter D. Stetson, Carol Friedman
J. Am. Medical Informatics Assoc.1
2008 Methods for Building Sense Inventories of Abbreviations in Clinical Notes
Hua Xu 0001, Peter D. Stetson, Carol Friedman
AMIA1
2008 Research Paper: Automated Acquisition of Disease-Drug Knowledge from Biomedical and Clinical Documents: An Initial Study
abstract
OBJECTIVE: Explore the automated acquisition of knowledge in biomedical and clinical documents using text mining and statistical techniques to identify disease-drug associations. DESIGN: Biomedical literature and clinical narratives from the patient record were mined to gather knowledge about disease-drug associations. Two NLP systems, BioMedLEE and MedLEE, were applied to Medline articles and discharge summaries, respectively. Disease and drug entities were identified using the NLP systems in addition to MeSH annotations for the Medline articles. Focusing on eight diseases, co-occurrence statistics were applied to compute and evaluate the strength of association between each disease and relevant drugs. RESULTS: Ranked lists of disease-drug pairs were generated and cutoffs calculated for identifying stronger associations among these pairs for further analysis. Differences and similarities between the text sources (i.e., biomedical literature and patient record) and annotations (i.e., MeSH and NLP-extracted UMLS concepts) with regards to disease-drug knowledge were observed. CONCLUSION: This paper presents a method for acquiring disease-specific knowledge and a feasibility study of the method. The method is based on applying a combination of NLP and statistical techniques to both biomedical and clinical documents. The approach enabled extraction of knowledge about the drugs clinicians are using for patients with specific diseases based on the patient record, while it is also acquired knowledge of drugs frequently involved in controlled trials for those same diseases. In comparing the disease-drug associations, we found the results to be appropriate: the two text sources contained consistent as well as complementary knowledge, and manual review of the top five disease-drug associations by a medical expert supported their correctness across the diseases.
Elizabeth S. Chen, George Hripcsak, Hua Xu 0001, Marianthi Markatou, Carol Friedman
J. Am. Medical Informatics Assoc.3
2007 A Study of Abbreviations in Clinical Notes
Hua Xu 0001, Peter D. Stetson, Carol Friedman
AMIA1
2007 Gene symbol disambiguation using knowledge-based profiles
abstract
MOTIVATION: The ambiguity of biomedical entities, particularly of gene symbols, is a big challenge for text-mining systems in the biomedical domain. Existing knowledge sources, such as Entrez Gene and the MEDLINE database, contain information concerning the characteristics of a particular gene that could be used to disambiguate gene symbols. RESULTS: For each gene, we create a profile with different types of information automatically extracted from related MEDLINE abstracts and readily available annotated knowledge sources. We apply the gene profiles to the disambiguation task via an information retrieval method, which ranks the similarity scores between the context where the ambiguous gene is mentioned, and candidate gene profiles. The gene profile with the highest similarity score is then chosen as the correct sense. We evaluated the method on three automatically generated testing sets of mouse, fly and yeast organisms, respectively. The method achieved the highest precision of 93.9% for the mouse, 77.8% for the fly and 89.5% for the yeast. AVAILABILITY: The testing data sets and disambiguation programs are available at http://www.dbmi.columbia.edu/~hux7002/gsd2006
Hua Xu 0001, Jungwei Fan 0001, George Hripcsak, Eneida A. Mendonça, Marianthi Markatou, Carol Friedman
Bioinform.1
2007 Using contextual and lexical features to restructure and validate the classification of biomedical concepts
abstract
BACKGROUND: Biomedical ontologies are critical for integration of data from diverse sources and for use by knowledge-based biomedical applications, especially natural language processing as well as associated mining and reasoning systems. The effectiveness of these systems is heavily dependent on the quality of the ontological terms and their classifications. To assist in developing and maintaining the ontologies objectively, we propose automatic approaches to classify and/or validate their semantic categories. In previous work, we developed an approach using contextual syntactic features obtained from a large domain corpus to reclassify and validate concepts of the Unified Medical Language System (UMLS), a comprehensive resource of biomedical terminology. In this paper, we introduce another classification approach based on words of the concept strings and compare it to the contextual syntactic approach. RESULTS: The string-based approach achieved an error rate of 0.143, with a mean reciprocal rank of 0.907. The context-based and string-based approaches were found to be complementary, and the error rate was reduced further by applying a linear combination of the two classifiers. The advantage of combining the two approaches was especially manifested on test data with sufficient contextual features, achieving the lowest error rate of 0.055 and a mean reciprocal rank of 0.969. CONCLUSION: The lexical features provide another semantic dimension in addition to syntactic contextual features that support the classification of ontological concepts. The classification errors of each dimension can be further reduced through appropriate combination of the complementary classifiers.
Jungwei Fan 0001, Hua Xu 0001, Carol Friedman
BMC Bioinform.2
2007 Natural language processing and visualization in the molecular imaging domain
P. Karina Tulipano, Ying Tao, William S. Millar, Pat Zanzonico, Katherine Kolbert, Hua Xu 0001, Hong Yu 0001, Lifeng Chen, Yves A. Lussier, Carol Friedman
J. Biomed. Informatics6
2006 A Natural Language Processing (NLP) Tool to Assist in the Curation Of the Laboratory Mouse Tumor Biology Database
Hua Xu 0001, Debra M. Krupke, Judith A. Blake, Carol Friedman
AMIA1
2006 Machine learning and word sense disambiguation in the biomedical domain: design and evaluation issues
abstract
BACKGROUND: Word sense disambiguation (WSD) is critical in the biomedical domain for improving the precision of natural language processing (NLP), text mining, and information retrieval systems because ambiguous words negatively impact accurate access to literature containing biomolecular entities, such as genes, proteins, cells, diseases, and other important entities. Automated techniques have been developed that address the WSD problem for a number of text processing situations, but the problem is still a challenging one. Supervised WSD machine learning (ML) methods have been applied in the biomedical domain and have shown promising results, but the results typically incorporate a number of confounding factors, and it is problematic to truly understand the effectiveness and generalizability of the methods because these factors interact with each other and affect the final results. Thus, there is a need to explicitly address the factors and to systematically quantify their effects on performance. RESULTS: Experiments were designed to measure the effect of "sample size" (i.e. size of the datasets), "sense distribution" (i.e. the distribution of the different meanings of the ambiguous word) and "degree of difficulty" (i.e. the measure of the distances between the meanings of the senses of an ambiguous word) on the performance of WSD classifiers. Support Vector Machine (SVM) classifiers were applied to an automatically generated data set containing four ambiguous biomedical abbreviations: BPD, BSA, PCA, and RSV, which were chosen because of varying degrees of differences in their respective senses. Results showed that: 1) increasing the sample size generally reduced the error rate, but this was limited mainly to well-separated senses (i.e. cases where the distances between the senses were large); in difficult cases an unusually large increase in sample size was needed to increase performance slightly, which was impractical, 2) the sense distribution did not have an effect on performance when the senses were separable, 3) when there was a majority sense of over 90%, the WSD classifier was not better than use of the simple majority sense, 4) error rates were proportional to the similarity of senses, and 5) there was no statistical difference between results when using a 5-fold or 10-fold cross-validation method. Other issues that impact performance are also enumerated. CONCLUSION: Several different independent aspects affect performance when using ML techniques for WSD. We found that combining them into one single result obscures understanding of the underlying methods. Although we studied only four abbreviations, we utilized a well-established statistical method that guarantees the results are likely to be generalizable for abbreviations with similar characteristics. The results of our experiments show that in order to understand the performance of these ML methods it is critical that papers report on the baseline performance, the distribution and sample size of the senses in the datasets, and the standard deviation or confidence intervals. In addition, papers should also characterize the difficulty of the WSD task, the WSD situations addressed and not addressed, as well as the ML methods and features used. This should lead to an improved understanding of the generalizablility and the limitations of the methodology.
Hua Xu 0001, Marianthi Markatou, Rositsa Dimova, Carol Friedman
BMC Bioinform.1
2003 Facilitating Research in Pathology using Natural Language Processing
Hua Xu 0001, Carol Friedman
AMIA1