EDBT 2026 Demo / reviewers in the wild / expert
Yanshan Wang
dblp:45/11295
· DBLP profile ↗
55ranked-venue papers
11as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 51 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDoH-GPT: using large language models to extract social determinants of healthabstractOBJECTIVE: Extracting social determinants of health (SDoHs) from medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. Here, we introduce SDoH-GPT, a novel framework leveraging few-shot learning large language models (LLMs) to automate the extraction of SDoH from unstructured text, aiming to improve both efficiency and generalizability. MATERIALS AND METHODS: SDoH-GPT is a framework including the few-shot learning LLM methods to extract the SDoH from medical notes and the XGBoost classifiers which continue to classify SDoH using the annotations generated by the few-shot learning LLM methods as training datasets. The unique combination of the few-shot learning LLM methods with XGBoost utilizes the strength of LLMs as great few shot learners and the efficiency of XGBoost when the training dataset is sufficient. Therefore, SDoH-GPT can extract SDoH without relying on extensive medical annotations or costly human intervention. RESULTS: Our approach achieved tenfold and twentyfold reductions in time and cost, respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of LLM and XGBoost can ensure high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. DISCUSSION: This study has verified SDoH-GPT on three datasets and highlights the potential of leveraging LLM and XGBoost to revolutionize medical note classification, demonstrating its capability to achieve highly accurate classifications with significantly reduced time and cost. CONCLUSION: The key contribution of this study is the integration of LLM with XGBoost, which enables cost-effective and high quality annotations of SDoH. This research sets the stage for SDoH can be more accessible, scalable, and impactful in driving future healthcare solutions. Bernardo Scapini Consoli, Xizhi Wu, Song Wang 0026, Yanshan Wang, Justin F. Rousseau, Thomas Hartvigsen, Li Shen 0001, Huanmei Wu, Yifan Peng 0002, Qi Long, Tianlong Chen 0001, Ying Ding 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2026 | CPGPrompt: translating clinical guidelines into large language model-executable decision supportabstractOBJECTIVE: Clinical practice guidelines (CPGs) provide evidence-based recommendations for patient care; however, integrating them into artificial intelligence (AI) remains challenging. Previous approaches, such as rule-based systems or black-box AI models, face significant limitations, including poor interpretability, inconsistent adherence to guidelines, and narrow domain applicability. To address this, we develop and validate CPGPrompt, an auto-prompting system that converts narrative clinical guidelines into large language models (LLMs). MATERIALS AND METHODS: Our framework translates CPGs into structured decision trees and utilizes an LLM to dynamically navigate them for patient case evaluation. Synthetic vignettes were generated across 3 domains-headache, lower back pain, and prostate cancer-and distributed into 4 categories to test different decision scenarios. System performance was assessed on both binary specialty referral decisions and fine-grained pathway classification tasks. RESULTS: The binary specialty referral classification achieved consistently strong performance across all domains (F1: 0.85-1.00), with high recall (1.00 ± 0.00). In contrast, multiclass pathway assignment showed reduced performance, with domain-specific variations: headache (F1: 0.47), lower back pain (F1: 0.72), and prostate cancer (F1: 0.77). DISCUSSION: Domain-specific performance differences reflected the structure of each guideline. The headache guideline highlighted challenges with negation handling. The lower back pain guideline required temporal reasoning. In contrast, prostate cancer pathways benefited from quantifiable laboratory tests, resulting in more reliable decision-making. CONCLUSION: CPGPrompt demonstrates generalizability across diverse clinical domains while maintaining high sensitivity for referral decisions. Its transparent, auditable framework enables the systematic identification of failure modes and provides advantages over black-box AI approaches. However, persistent challenges with subjective clinical assessments indicate a need for targeted improvements and greater clinical robustness. Ruiqi Deng, Geoffrey Martin, Tony Wang, Yi Liu 0059, Chunhua Weng, Yanshan Wang, Justin F. Rousseau, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Machine learning applications related to suicide in military and Veterans: A scoping literature review
Yishu Wei, Yanshan Wang, Yunyu Xiao, Ronald K. Poropatich, Gretchen L. Haas, Yiye Zhang, Chunhua Weng, Jinze Liu, Lisa A. Brenner, James M. Bjork, Yifan Peng 0002 |
J. Biomed. Informatics | 3 |
| 2024 | TrojFSP: Trojan Insertion in Few-shot Prompt TuningabstractMengxin Zheng, Jiaqi Xue, Xun Chen, Yanshan Wang, Qian Lou, Lei Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Mengxin Zheng, Yanshan Wang, Qian Lou |
NAACL-HLT | 4 |
| 2024 | A taxonomy for advancing systematic error analysis in multi-site electronic health record-based clinical concept extractionabstractBACKGROUND: Error analysis plays a crucial role in clinical concept extraction, a fundamental subtask within clinical natural language processing (NLP). The process typically involves a manual review of error types, such as contextual and linguistic factors contributing to their occurrence, and the identification of underlying causes to refine the NLP model and improve its performance. Conducting error analysis can be complex, requiring a combination of NLP expertise and domain-specific knowledge. Due to the high heterogeneity of electronic health record (EHR) settings across different institutions, challenges may arise when attempting to standardize and reproduce the error analysis process. OBJECTIVES: This study aims to facilitate a collaborative effort to establish common definitions and taxonomies for capturing diverse error types, fostering community consensus on error analysis for clinical concept extraction tasks. MATERIALS AND METHODS: We iteratively developed and evaluated an error taxonomy based on existing literature, standards, real-world data, multisite case evaluations, and community feedback. The finalized taxonomy was released in both .dtd and .owl formats at the Open Health Natural Language Processing Consortium. The taxonomy is compatible with several different open-source annotation tools, including MAE, Brat, and MedTator. RESULTS: The resulting error taxonomy comprises 43 distinct error classes, organized into 6 error dimensions and 4 properties, including model type (symbolic and statistical machine learning), evaluation subject (model and human), evaluation level (patient, document, sentence, and concept), and annotation examples. Internal and external evaluations revealed strong variations in error types across methodological approaches, tasks, and EHR settings. Key points emerged from community feedback, including the need to enhancing clarity, generalizability, and usability of the taxonomy, along with dissemination strategies. CONCLUSION: The proposed taxonomy can facilitate the acceleration and standardization of the error analysis process in multi-site settings, thus improving the provenance, interpretability, and portability of NLP models. Future researchers could explore the potential direction of developing automated or semi-automated methods to assist in the classification and standardization of error analysis. Sunyang Fu, Liwei Wang 0010, Andrew Wen, Nansu Zong, Anamika Kumari, Rui Zhang 0028, Yanshan Wang, Jennifer L. St. Sauver, Sunghwan Sohn |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Distilling large language models for matching patients to clinical trialsabstractOBJECTIVE: The objective of this study is to systematically examine the efficacy of both proprietary (GPT-3.5, GPT-4) and open-source large language models (LLMs) (LLAMA 7B, 13B, 70B) in the context of matching patients to clinical trials in healthcare. MATERIALS AND METHODS: The study employs a multifaceted evaluation framework, incorporating extensive automated and human-centric assessments along with a detailed error analysis for each model, and assesses LLMs' capabilities in analyzing patient eligibility against clinical trial's inclusion and exclusion criteria. To improve the adaptability of open-source LLMs, a specialized synthetic dataset was created using GPT-4, facilitating effective fine-tuning under constrained data conditions. RESULTS: The findings indicate that open-source LLMs, when fine-tuned on this limited and synthetic dataset, achieve performance parity with their proprietary counterparts, such as GPT-3.5. DISCUSSION: This study highlights the recent success of LLMs in the high-stakes domain of healthcare, specifically in patient-trial matching. The research demonstrates the potential of open-source models to match the performance of proprietary models when fine-tuned appropriately, addressing challenges like cost, privacy, and reproducibility concerns associated with closed-source proprietary LLMs. CONCLUSION: The study underscores the opportunity for open-source LLMs in patient-trial matching. To encourage further research and applications in this field, the annotated evaluation dataset and the fine-tuned LLM, Trial-LLAMA, are released for public use. Mauro Nievas, Aditya Basu, Yanshan Wang, Hrituraj Singh |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Large language models for biomedicine: foundations, opportunities, challenges, and best practicesabstractOBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications. Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Extraction of sleep information from clinical notes of Alzheimer's disease patients using natural language processingabstractOBJECTIVES: Alzheimer's disease (AD) is the most common form of dementia in the United States. Sleep is one of the lifestyle-related factors that has been shown critical for optimal cognitive function in old age. However, there is a lack of research studying the association between sleep and AD incidence. A major bottleneck for conducting such research is that the traditional way to acquire sleep information is time-consuming, inefficient, non-scalable, and limited to patients' subjective experience. We aim to automate the extraction of specific sleep-related patterns, such as snoring, napping, poor sleep quality, daytime sleepiness, night wakings, other sleep problems, and sleep duration, from clinical notes of AD patients. These sleep patterns are hypothesized to play a role in the incidence of AD, providing insight into the relationship between sleep and AD onset and progression. MATERIALS AND METHODS: A gold standard dataset is created from manual annotation of 570 randomly sampled clinical note documents from the adSLEEP, a corpus of 192 000 de-identified clinical notes of 7266 AD patients retrieved from the University of Pittsburgh Medical Center (UPMC). We developed a rule-based natural language processing (NLP) algorithm, machine learning models, and large language model (LLM)-based NLP algorithms to automate the extraction of sleep-related concepts, including snoring, napping, sleep problem, bad sleep quality, daytime sleepiness, night wakings, and sleep duration, from the gold standard dataset. RESULTS: The annotated dataset of 482 patients comprised a predominantly White (89.2%), older adult population with an average age of 84.7 years, where females represented 64.1%, and a vast majority were non-Hispanic or Latino (94.6%). Rule-based NLP algorithm achieved the best performance of F1 across all sleep-related concepts. In terms of positive predictive value (PPV), the rule-based NLP algorithm achieved the highest PPV scores for daytime sleepiness (1.00) and sleep duration (1.00), while the machine learning models had the highest PPV for napping (0.95) and bad sleep quality (0.86), and LLAMA2 with finetuning had the highest PPV for night wakings (0.93) and sleep problem (0.89). DISCUSSION: Although sleep information is infrequently documented in the clinical notes, the proposed rule-based NLP algorithm and LLM-based NLP algorithms still achieved promising results. In comparison, the machine learning-based approaches did not achieve good results, which is due to the small size of sleep information in the training data. CONCLUSION: The results show that the rule-based NLP algorithm consistently achieved the best performance for all sleep concepts. This study focused on the clinical notes of patients with AD but could be extended to general sleep information extraction for other diseases. Sonish Sivarajkumar, Thomas Yu Chow Tam, Haneef Ahamed Mohammad, Samuel Viggiano, David Oniani, Shyam Visweswaran, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness
Qiao Jin 0001, Denis Jered McInerney, Yong Chen 0016, Fei Wang 0001, Curtis L. Cole, Qian Yang 0004, Yanshan Wang, Bradley A. Malin, Mor Peleg, Byron C. Wallace, Zhiyong Lu, Chunhua Weng, Yifan Peng 0002 |
J. Biomed. Informatics | 8 |
| 2023 | Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001 |
J. Biomed. Informatics | 17 |
| 2023 | Fair patient model: Mitigating bias in the patient representation learned from the electronic health records
Sonish Sivarajkumar, Yufei Huang 0001, Yanshan Wang |
J. Biomed. Informatics | 3 |
| 2022 | A Pilot Study of Automated Approaches to Translate Health Illiterate Language
Renuk DeAlmeida, Dinuk DeAlmeida, Sreekanth Sreekumar, Yanshan Wang |
AMIA | 4 |
| 2022 | HealthPrompt: A zero-shot learning paradigm for Clinical Natural Language Processing
Sonish Sivarajkumar, Yanshan Wang |
AMIA | 2 |
| 2022 | A Novel Zero-shot Learning Framework for Clinical Text Classification
Sonish Sivarajkumar, Yanshan Wang |
AMIA | 2 |
| 2021 | Patient Asynchronous Response to Coronavirus Disease 2019 (COVID-19): A Retrospective Analysis of Patient Portal Messages
Ming Huang 0006, Aditya Khurana, George M. Mastorakos, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Christi A. Patten, Jungwei Fan 0001 |
AMIA | 8 |
| 2021 | COVID-19 Dashboard: Visual Exploration of the Regional Pandemic Trend
Liwei Wang 0010, Andrew Wen, Ming Huang 0006, Yanshan Wang |
AMIA | 5 |
| 2021 | Disparity analysis of patient portal messaging use for COVID-19 in urban versus rural locality
Ming Huang 0006, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Nansu Zong, Yue Yu 0012, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Chyke Doubeni, Jungwei Fan 0001, Christi A. Patten |
AMIA | 6 |
| 2021 | Detecting Major Depressive Disorder from Clinical Notes using Neural Language Models with Distant Supervision
Bhavani Singh Agnikula Kshatriya, Nicolas A. Nunez, Manuel Gardea-Resendez, Euijung Ryu, Brandon J. Coombes, Sunyang Fu, Mark A. Frye, Joanna M. Biernacka, Yanshan Wang |
AMIA | 9 |
| 2021 | Corrigendum to: The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task on clinical concept normalization for clinical recordsabstractObjective The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task track 3, focused on medical concept normalization (MCN) in clinical records. This track aimed to assess the state of the art in identifying and matching salient medical concepts to a controlled vocabulary. In this paper, we describe the task, describe the data set used, compare the participating systems, present results, identify the strengths and limitations of the current state of the art, and identify directions for future research. Materials and methods Participating teams were provided with narrative discharge summaries in which text spans corresponding to medical concepts were identified. This paper refers to these text spans as mentions. Teams were tasked with normalizing these mentions to concepts, represented by concept unique identifiers, within the Unified Medical Language System. Submitted systems represented 4 broad categories of approaches: cascading dictionary matching, cosine distance, deep learning, and retrieve-and-rank systems. Disambiguation modules were common across all approaches. Results A total of 33 teams participated in the MCN task. The best-performing team achieved an accuracy of 0.8526. The median and mean performances among all teams were 0.7733 and 0.7426, respectively. Conclusions Overall performance among the top 10 teams was high. However, several mention types were challenging for all teams. These included mentions requiring disambiguation of misspelled words, acronyms, abbreviations, and mentions with more than 1 possible semantic type. Also challenging were complex mentions of long, multi-word terms that may require new ways of extracting and representing mention meaning, the use of domain knowledge, parse trees, or hand-crafted rules. Sam Henry 0001, Yanshan Wang, Feichen Shen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | An aberration detection-based approach for sentinel syndromic surveillance of COVID-19 and other novel influenza-like illnesses
Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Sunyang Fu, Sunghwan Sohn, Jacob A. Kugel, Vinod Kaggal, Ming Huang 0006, Yanshan Wang, Feichen Shen, Jungwei Fan 0001 |
J. Biomed. Informatics | 10 |
| 2020 | Deep Learning Approach to Parse Eligibility Criteria in Dietary Supplements Clinical Trials Following OMOP Common Data Model
Anusha Bompelli, Jianfu Li, Yiqi Xu, Yanshan Wang, Terrence Adam, Zhe He 0001, Rui Zhang 0028 |
AMIA | 5 |
| 2020 | Annotating Chronic Pain Episodes in EHR Text: Guideline Development and Corpus Analysis
Luke A. Carlson, Molly M. Jeffery, Sunyang Fu, Rozalina G. McCoy, Yanshan Wang, W. M. Hooten, Jennifer L. St. Sauver, Jungwei Fan 0001 |
AMIA | 6 |
| 2020 | Big Impact from Small Data: Unsupervised Machine Learning Approaches for Chronic Pain Patient Subgrouping
Luke A. Carlson, Jennifer L. St. Sauver, Sunyang Fu, Ahmad P. Tafti, Jungwei Fan 0001, Molly M. Jeffery, Rozalina G. McCoy, Yanshan Wang |
AMIA | 9 |
| 2020 | The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task on clinical concept normalization for clinical recordsabstractOBJECTIVE: The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task track 3, focused on medical concept normalization (MCN) in clinical records. This track aimed to assess the state of the art in identifying and matching salient medical concepts to a controlled vocabulary. In this paper, we describe the task, describe the data set used, compare the participating systems, present results, identify the strengths and limitations of the current state of the art, and identify directions for future research. MATERIALS AND METHODS: Participating teams were provided with narrative discharge summaries in which text spans corresponding to medical concepts were identified. This paper refers to these text spans as mentions. Teams were tasked with normalizing these mentions to concepts, represented by concept unique identifiers, within the Unified Medical Language System. Submitted systems represented 4 broad categories of approaches: cascading dictionary matching, cosine distance, deep learning, and retrieve-and-rank systems. Disambiguation modules were common across all approaches. RESULTS: A total of 33 teams participated in the MCN task. The best-performing team achieved an accuracy of 0.8526. The median and mean performances among all teams were 0.7733 and 0.7426, respectively. CONCLUSIONS: Overall performance among the top 10 teams was high. However, several mention types were challenging for all teams. These included mentions requiring disambiguation of misspelled words, acronyms, abbreviations, and mentions with more than 1 possible semantic type. Also challenging were complex mentions of long, multi-word terms that may require new ways of extracting and representing mention meaning, the use of domain knowledge, parse trees, or hand-crafted rules. Sam Henry 0001, Yanshan Wang, Feichen Shen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Recommendations for patient similarity classes: results of the AMIA 2019 workshop on defining patient similarityabstractDefining patient-to-patient similarity is essential for the development of precision medicine in clinical care and research. Conceptually, the identification of similar patient cohorts appears straightforward; however, universally accepted definitions remain elusive. Simultaneously, an explosion of vendors and published algorithms have emerged and all provide varied levels of functionality in identifying patient similarity categories. To provide clarity and a common framework for patient similarity, a workshop at the American Medical Informatics Association 2019 Annual Meeting was convened. This workshop included invited discussants from academics, the biotechnology industry, the FDA, and private practice oncology groups. Drawing from a broad range of backgrounds, workshop participants were able to coalesce around 4 major patient similarity classes: (1) feature, (2) outcome, (3) exposure, and (4) mixed-class. This perspective expands into these 4 subtypes more critically and offers the medical informatics community a means of communicating their work on this important topic. Nathan D. Seligson, Jeremy L. Warner, William S. Dalton, Robert S. Miller, Debra Patt, Kenneth L. Kehl, Matvey Palchuk, Gil Alterovitz, Laura K. Wiley, Ming Huang 0006, Feichen Shen, Yanshan Wang, Khoa A. Nguyen, Anthony F. Wong, Funda Meric-Bernstam, Elmer V. Bernstam, James L. Chen |
J. Am. Medical Informatics Assoc. | 13 |
| 2020 | Clinical concept extraction: A methodology review
Sunyang Fu, David Chen 0003, Sijia Liu 0002, Sungrim Moon, Kevin J. Peterson, Feichen Shen, Liwei Wang 0010, Yanshan Wang, Andrew Wen, Sunghwan Sohn |
J. Biomed. Informatics | 9 |
| 2020 | Unsupervised machine learning for the discovery of latent disease clusters and patient subgroups using electronic health records
Yanshan Wang, Terry M. Therneau, Elizabeth J. Atkinson, Ahmad P. Tafti, Shreyasee Amin, Andrew H. Limper, Sundeep Khosla |
J. Biomed. Informatics | 1 |
| 2019 | Trajectories of Functional Status of Cognitively-impaired Patients in EHRs
Yanshan Wang, Sunghwan Sohn |
AMIA | 1 |
| 2019 | Clinical Use of an Information Retrieval Framework for Cohort Discovery from Electronic Health Records
Yanshan Wang, Andrew Wen, Sijia Liu 0002, Jennifer L. St. Sauver, Adil E. Bharucha, Chunhua Weng |
AMIA | 1 |
| 2019 | Summarization of Patient Education Material for New Media Platform
Yanshan Wang, Paul R. Kingsbury, Michael Panzer, Richard Van Ert |
AMIA | 2 |
| 2019 | Enhancing Clinical Information Retrieval through Context-Aware Queries and IndicesabstractThe big data revolution has created a hefty demand for searching large-scale electronic health records (EHRs) to support clinical practice, research, and administration. Despite the volume of data involved, fast and accurate identification of clinical narratives pertinent to a clinical case being seen by any given provider is crucial for decision-making at the point of care. In the general domain, this capability is accomplished through a combination of the inverted index data structure, horizontal scaling, and information retrieval (IR) scoring algorithms. These technologies are also being used in the clinical domain, but have met limited success, particularly as clinical cases become more complex. One barrier affecting clinical performance is that contextual information, such as negation, temporality, and the subject of clinical mentions, impact clinical relevance but is not considered in general IR methodologies. In this study, we implemented a solution by identifying and incorporating the aforementioned semantic contexts as part of IR indexing/scoring with Elasticsearch. Experiments were conducted in comparison to baseline approaches with respect to: 1) evaluation of the impact on the quality (relevance) of the returned results, and 2) evaluation of the impact on execution time and storage requirements. The results showed a 5.1-23.1% improvement in retrieval quality, along with achieving 35% faster query execution time. Cost-wise, the solution required 1.5-2 times larger space and about 3 times increase in indexing time. The higher relevance demonstrated the merit of incorporating contextual information into clinical IR, and the near-constant increase in time and space suggested promising scalability. Andrew Wen, Yanshan Wang, Vinod Kaggal, Sijia Liu 0002, Jungwei Fan 0001 |
IEEE BigData | 2 |
| 2019 | HPO2Vec+: Leveraging heterogeneous knowledge resources to enrich node embeddings for the Human Phenotype Ontology
Feichen Shen, Su-Yuan Peng, Yadan Fan, Andrew Wen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010 |
J. Biomed. Informatics | 6 |
| 2018 | Information Extraction for Data Population in Lung Cancer Clinical Research
Liwei Wang 0010, Yanshan Wang, Jason A. Wampfler |
AMIA | 3 |
| 2018 | EMIRS: An Electronic Medical Information Retrieval System by Leveraging both Structured and Unstructured Electronic Health Records
Yanshan Wang, Andrew Wen, Sijia Liu 0002 |
AMIA | 1 |
| 2018 | Annotating Cohort Data Elements with OHDSI Common Data Model to Promote Research Reproducibility
Yanshan Wang, Henry Wang, Benjamin Yan, Feichen Shen, Kevin J. Peterson, Walter A. Rocca, Jennifer L. St. Sauver |
BIBM | 2 |
| 2018 | Clinical documentation variations and NLP system portability: a case study in asthma birth cohorts across institutionsabstractOBJECTIVE: To assess clinical documentation variations across health care institutions using different electronic medical record systems and investigate how they affect natural language processing (NLP) system portability. MATERIALS AND METHODS: Birth cohorts from Mayo Clinic and Sanford Children's Hospital (SCH) were used in this study (n = 298 for each). Documentation variations regarding asthma between the 2 cohorts were examined in various aspects: (1) overall corpus at the word level (ie, lexical variation), (2) topics and asthma-related concepts (ie, semantic variation), and (3) clinical note types (ie, process variation). We compared those statistics and explored NLP system portability for asthma ascertainment in 2 stages: prototype and refinement. RESULTS: There exist notable lexical variations (word-level similarity = 0.669) and process variations (differences in major note types containing asthma-related concepts). However, semantic-level corpora were relatively homogeneous (topic similarity = 0.944, asthma-related concept similarity = 0.971). The NLP system for asthma ascertainment had an F-score of 0.937 at Mayo, and produced 0.813 (prototype) and 0.908 (refinement) when applied at SCH. DISCUSSION: The criteria for asthma ascertainment are largely dependent on asthma-related concepts. Therefore, we believe that semantic similarity is important to estimate NLP system portability. As the Mayo Clinic and SCH corpora were relatively homogeneous at a semantic level, the NLP system, developed at Mayo Clinic, was imported to SCH successfully with proper adjustments to deal with the intrinsic corpus heterogeneity. Sunghwan Sohn, Yanshan Wang, Chung-Il Wi, Elizabeth A. Krusemark, Euijung Ryu, Mir H. Ali, Young J. Juhn |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | A comparison of word embeddings for the biomedical natural language processingabstractBACKGROUND: Word embeddings have been prevalently used in biomedical Natural Language Processing (NLP) applications due to the ability of the vector representations being able to capture useful semantic properties and linguistic relationships between words. Different textual resources (e.g., Wikipedia and biomedical literature corpus) have been utilized in biomedical NLP to train word embeddings and these word embeddings have been commonly leveraged as feature input to downstream machine learning models. However, there has been little work on evaluating the word embeddings trained from different textual resources. METHODS: In this study, we empirically evaluated word embeddings trained from four different corpora, namely clinical notes, biomedical publications, Wikipedia, and news. For the former two resources, we trained word embeddings using unstructured electronic health record (EHR) data available at Mayo Clinic and articles (MedLit) from PubMed Central, respectively. For the latter two resources, we used publicly available pre-trained word embeddings, GloVe and Google News. The evaluation was done qualitatively and quantitatively. For the qualitative evaluation, we randomly selected medical terms from three categories (i.e., disorder, symptom, and drug), and manually inspected the five most similar words computed by embeddings for each term. We also analyzed the word embeddings through a 2-dimensional visualization plot of 377 medical terms. For the quantitative evaluation, we conducted both intrinsic and extrinsic evaluation. For the intrinsic evaluation, we evaluated the word embeddings' ability to capture medical semantics by measruing the semantic similarity between medical terms using four published datasets: Pedersen's dataset, Hliaoutakis's dataset, MayoSRS, and UMNSRS. For the extrinsic evaluation, we applied word embeddings to multiple downstream biomedical NLP applications, including clinical information extraction (IE), biomedical information retrieval (IR), and relation extraction (RE), with data from shared tasks. RESULTS: The qualitative evaluation shows that the word embeddings trained from EHR and MedLit can find more similar medical terms than those trained from GloVe and Google News. The intrinsic quantitative evaluation verifies that the semantic similarity captured by the word embeddings trained from EHR is closer to human experts' judgments on all four tested datasets. The extrinsic quantitative evaluation shows that the word embeddings trained on EHR achieved the best F1 score of 0.900 for the clinical IE task; no word embeddings improved the performance for the biomedical IR task; and the word embeddings trained on Google News had the best overall F1 score of 0.790 for the RE task. CONCLUSION: Based on the evaluation results, we can draw the following conclusions. First, the word embeddings trained from EHR and MedLit can capture the semantics of medical terms better, and find semantically relevant medical terms closer to human experts' judgments than those trained from GloVe and Google News. Second, there does not exist a consistent global ranking of word embeddings for all downstream biomedical NLP applications. However, adding word embeddings as extra features will improve results on most downstream tasks. Finally, the word embeddings trained from the biomedical domain corpora do not necessarily have better performance than those trained from the general domain corpora for any downstream biomedical NLP task. Yanshan Wang, Sijia Liu 0002, Naveed Afzal, Majid Rastegar-Mojarad, Liwei Wang 0010, Feichen Shen, Paul R. Kingsbury |
J. Biomed. Informatics | 1 |
| 2018 | Clinical information extraction applications: A literature reviewabstractBACKGROUND: With the rapid adoption of electronic health records (EHRs), it is desirable to harvest information and knowledge from EHRs to support automated systems at the point of care and to enable secondary use of EHRs for clinical and translational research. One critical component used to facilitate the secondary use of EHR data is the information extraction (IE) task, which automatically extracts and encodes clinical information from text. OBJECTIVES: In this literature review, we present a review of recent published research on clinical information extraction (IE) applications. METHODS: A literature search was conducted for articles published from January 2009 to September 2016 based on Ovid MEDLINE In-Process & Other Non-Indexed Citations, Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science, and ACM Digital Library. RESULTS: A total of 1917 publications were identified for title and abstract screening. Of these publications, 263 articles were selected and discussed in this review in terms of publication venues and data sources, clinical IE tools, methods, and applications in the areas of disease- and drug-related studies, and clinical workflow optimizations. CONCLUSIONS: Clinical IE has been used for a wide range of applications, however, there is a considerable gap between clinical studies using EHR data and studies using clinical IE. This study enabled us to gain a more concrete understanding of the gap and to provide potential solutions to bridge this gap. Yanshan Wang, Liwei Wang 0010, Majid Rastegar-Mojarad, Sungrim Moon, Feichen Shen, Naveed Afzal, Sijia Liu 0002, Yuqun Zeng, Saeed Mehrabi 0003, Sunghwan Sohn |
J. Biomed. Informatics | 1 |
| 2017 | A Resource for Clinical Semantic Textual Similarity
Naveed Afzal, Yanshan Wang, Feichen Shen, Liwei Wang 0010 |
AMIA | 2 |
| 2017 | Leveraging Collaborative Filtering to Accelerate Rare Disease Diagnosis
Feichen Shen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010, Naveed Afzal |
AMIA | 3 |
| 2017 | Accelerating Rare Disease Diagnosis with Collaborative Filtering
Feichen Shen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010, Naveed Afzal |
AMIA | 3 |
| 2017 | Semantic Representations of Medical Terms: A Comparison Study of Word Embeddings Trained on Clinical Notes and PubMed Articles
Yanshan Wang, Naveed Afzal, Liwei Wang 0010, Feichen Shen |
AMIA | 1 |
| 2017 | Recommending education materials for diabetic questions using information retrieval approaches
Yuqun Zeng, Yanshan Wang, Feichen Shen, Sijia Liu 0002, Liwei Wang 0010, Majid Rastegar-Mojarad, Xu-Sheng Liu |
AMIA | 2 |
| 2017 | Medical concept intersection between outside medical records and consultant notes: A case study in transferred cardiovascular patientsabstractOne of the promises of “meaningful use” of Electronic Health Records (EHRs) is to facilitate digital information exchange between healthcare providers through continuity of care documents. Despite such promise, outside medical records (OMRs) of referral patients including clinical notes, lab test results or diagnostic test reports are frequently provided through fax or print out. Moreover, it is not clear how much information in those OMRs is utilized when providing care at the early stage. In this study, we collected clinical concepts automatically from OMRs through optical character recognition (OCR) technology and then performed a quantitative analysis of concepts presented in OMRs and concepts captured in clinical notes at Mayo Clinic. We also investigated information from OMRs not captured in initial consultant notes but presented over subsequent consultant notes. We identified 12.93% of concepts from OMRs were identified in clinical documents within three months. Among those overlapping concepts, 26.74% of them were not captured in initial consultant notes. Our study presents that clinical information from OMRs is important for patient care. Also, the delayed presence of information in clinical notes may indicate important information from OMRs is not fully utilized earlier in the care. Sungrim Moon, Sijia Liu 0002, Paul R. Kingsbury, David Chen 0003, Yanshan Wang, Feichen Shen, Rajeev Chaudhry |
BIBM | 5 |
| 2017 | Intrainstitutional EHR collections for patient-level information retrievalabstractResearch in clinical information retrieval has long been stymied by the lack of open resources. However, both clinical information retrieval research innovation and legitimate privacy concerns can be served by the creation of intrainstitutional, fully protected resources. In this article, we provide some principles and tools for information retrieval resource‐building in the unique problem setting of patient‐level information retrieval, following the tradition of the Cranfield paradigm. We further include an analysis of parallel information retrieval resources at Oregon Health & Science University and Mayo Clinic that were built on these principles. Stephen T. Wu, Sijia Liu 0002, Yanshan Wang, Tamara Timmons, Harsha Uppili, Steven Bedrick, William R. Hersh |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | A Topic-modeling Based Framework for Drug-drug Interaction Classification from Biomedical Text
Dingcheng Li, Sijia Liu 0002, Majid Rastegar-Mojarad, Yanshan Wang, Vipin Chaudhary, Terry M. Therneau |
AMIA | 4 |
| 2016 | Drug-drug Interaction Detection with A Topic-modeling Based Framework Augmented with Distant-supervision
Dingcheng Li, Sijia Liu 0002, Majid Rastegar-Mojarad, Yanshan Wang, Vipin Chaudhary, Terry M. Therneau |
AMIA | 4 |
| 2016 | Asthma Ascertainment NLP System Portability across Institutions
Sunghwan Sohn, Yanshan Wang, Chung-Il Wi, Elizabeth A. Krusemark, Euijung Ryu, Mir H. Ali, Young J. Juhn |
AMIA | 2 |
| 2016 | A Part-Of-Speech Weighting Scheme for Clinical Information Retrieval
Yanshan Wang, Stephen T. Wu, Dingcheng Li |
AMIA | 1 |
| 2016 | Probabilistic Population-level Modeling of Disease Event Timelines
Stephen T. Wu, Yanshan Wang, Sunghwan Sohn, Chung-Il Wi, Elizabeth A. Krusemark, Young J. Juhn |
AMIA | 2 |
| 2016 | Answering patients' questions using expert-vetted online resources: A case study of diabetes
Yuqun Zeng, Liwei Wang 0010, Yanshan Wang, Dingcheng Li |
AMIA | 3 |
| 2016 | Answering diabetic patients' questions using expert-vetted online resources: A case studyabstractAs an increasing number of patients seek medical information online, it is crucial to bring expert-vetted information to patients. Multiple expert-vetted online resources exist for their accuracy and authority of medical knowledge. However, it is not clear which resource is better in meeting the information needs of the patients. Utilizing a collection of questions raised by patients with diabetes retrieved from an online forum, we manually evaluated three widely used expert-vetted online resources (WebMD, MedlinePlus, and UpToDate) in terms of content coverage and time spent to find answers. The results indicated that WebMD had slightly better content coverage with less time spent in finding answers. Leveraging the Natural Language Processing (NLP) techniques and a clinical terminology resource, the Unified Medical Language System (UMLS), we demonstrated that WebMD had a higher mapping rate of clinical concepts comparing to other resources. Yuqun Zeng, Xu-Sheng Liu, Liwei Wang 0010, Yanshan Wang |
BIBM | 5 |
| 2016 | Indexing by Latent Dirichlet Allocation and an Ensemble ModelabstractThe contribution of this article is twofold. First, we present Indexing by latent Dirichlet allocation (LDI), an automatic document indexing method. Many ad hoc applications, or their variants with smoothing techniques suggested in LDA‐based language modeling, can result in unsatisfactory performance as the document representations do not accurately reflect concept space. To improve document retrieval performance, we introduce a new definition of document probability vectors in the context of LDA and present a novel scheme for automatic document indexing based on LDA. Second, we propose an Ensemble Model (EnM) for document retrieval. EnM combines basic indexing models by assigning different weights and attempts to uncover the optimal weights to maximize the mean average precision. To solve the optimization problem, we propose an algorithm, which is derived based on the boosting method. The results of our computational experiments on benchmark data sets indicate that both the proposed approaches are viable options for document retrieval. Yanshan Wang, In-Chan Choi |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | A Part-Of-Speech term weighting scheme for biomedical information retrieval
Yanshan Wang, Stephen T. Wu, Dingcheng Li, Saeed Mehrabi 0003 |
J. Biomed. Informatics | 1 |
| 2012 | A Text Classification Method based on Latent Topics
Yanshan Wang, In-Chan Choi |
ICORES | 1 |