EDBT 2026 Demo / reviewers in the wild / expert
Liwei Wang 0010
dblp:47/1798-10
· DBLP profile ↗
38ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0001-9970-8604ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 37 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering signature disease trajectories in pancreatic cancer and soft-tissue sarcoma from longitudinal patient records
Liwei Wang 0010, Andrew Wen, Qiuhao Lu, Jinlian Wang, Xiaoyang Ruan, Adriana Gamboa, Neha Malik, Christina L. Roland, Matthew H. G. Katz, Heather Lyu |
J. Biomed. Informatics | 1 |
| 2024 | A taxonomy for advancing systematic error analysis in multi-site electronic health record-based clinical concept extractionabstractBACKGROUND: Error analysis plays a crucial role in clinical concept extraction, a fundamental subtask within clinical natural language processing (NLP). The process typically involves a manual review of error types, such as contextual and linguistic factors contributing to their occurrence, and the identification of underlying causes to refine the NLP model and improve its performance. Conducting error analysis can be complex, requiring a combination of NLP expertise and domain-specific knowledge. Due to the high heterogeneity of electronic health record (EHR) settings across different institutions, challenges may arise when attempting to standardize and reproduce the error analysis process. OBJECTIVES: This study aims to facilitate a collaborative effort to establish common definitions and taxonomies for capturing diverse error types, fostering community consensus on error analysis for clinical concept extraction tasks. MATERIALS AND METHODS: We iteratively developed and evaluated an error taxonomy based on existing literature, standards, real-world data, multisite case evaluations, and community feedback. The finalized taxonomy was released in both .dtd and .owl formats at the Open Health Natural Language Processing Consortium. The taxonomy is compatible with several different open-source annotation tools, including MAE, Brat, and MedTator. RESULTS: The resulting error taxonomy comprises 43 distinct error classes, organized into 6 error dimensions and 4 properties, including model type (symbolic and statistical machine learning), evaluation subject (model and human), evaluation level (patient, document, sentence, and concept), and annotation examples. Internal and external evaluations revealed strong variations in error types across methodological approaches, tasks, and EHR settings. Key points emerged from community feedback, including the need to enhancing clarity, generalizability, and usability of the taxonomy, along with dissemination strategies. CONCLUSION: The proposed taxonomy can facilitate the acceleration and standardization of the error analysis process in multi-site settings, thus improving the provenance, interpretability, and portability of NLP models. Future researchers could explore the potential direction of developing automated or semi-automated methods to assist in the classification and standardization of error analysis. Sunyang Fu, Liwei Wang 0010, Andrew Wen, Nansu Zong, Anamika Kumari, Rui Zhang 0028, Yanshan Wang, Jennifer L. St. Sauver, Sunghwan Sohn |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | FedFSA: Hybrid and federated framework for functional status ascertainment across institutions
Sunyang Fu, Heling Jia, Maria Vassilaki, Vipina Kuttichi Keloth, Yifang Dang, Yujia Zhou 0003, Muskan Garg, Ronald C. Petersen, Jennifer L. St. Sauver, Sungrim Moon, Liwei Wang 0010, Andrew Wen, Fang Li 0011, Hua Xu 0001, Cui Tao, Jungwei Fan 0001, Sunghwan Sohn |
J. Biomed. Informatics | 11 |
| 2023 | GRU-D-Weibull: A novel real-time individualized endpoint predictionabstractIn the era of healthcare digital transformation, using electronic health record (EHR) data to generate various endpoint estimates for active monitoring is highly desirable in chronic disease management. However, traditional predictive modeling strategies leveraging well-curated data sets can have limited real-world implementation potential due to various data quality issues in EHR data. We propose a novel predictive modeling approach, GRU-D-Weibull, which models Weibull distribution leveraging gated recurrent units with decay (GRU-D), for real-time individualized endpoint prediction and population level risk management using EHR data. We systematically evaluated the performance and showcased the real-world implementability of the proposed approach through individual level endpoint prediction using a cohort of patients with chronic kidney disease stage 4 (CKD4). A total of 536 features including ICD/CPT codes, medications, lab tests, vital measurements, and demographics were retrieved for 6879 CKD4 patients. The performance metrics including C-index, L1-loss, Parkes' error, and predicted survival probability at time of event were compared between GRU-D-Weibull and other alternative approaches including accelerated failure time model (AFT), XGBoost based AFT (XGB(AFT)), random survival forest (RSF), and Nnet-survival. Both in-process and post-process calibrations were experimented on GRU-D-Weibull generated survival probabilities. GRU-D-Weibull demonstrated C-index of ~0.7 at index date, which increased to ~0.77 at 4.3 years of follow-up, comparable to that of RSF. GRU-D-Weibull achieved absolute L1-loss of ~1.1 years (sd≈0.95) at CKD4 index date, and a minimum of ~0.45 year (sd≈0.3) at 4 years of follow-up, comparing to second-ranked RSF of ~1.4 years (sd≈1.1) at index date and ~0.64 years (sd≈0.26) at 4 years. Both significantly outperform competing approaches. GRU-D-Weibull constrained predicted survival probability at time of event to smaller and more fixed range than competing models throughout follow-up. Significant correlations were observed between prediction error and missing proportions of all major categories of input features at index date (Corr ~0.1 to ~0.3), which faded away within 1 year after index date as more data became available. Through post training recalibration, we achieved a close alignment between the predicted and observed survival probabilities across multiple prediction horizons at different time points during follow-up. GRU-D-Weibull shows advantages over competing methods in handling missingness commonly encountered in EHR data and providing both probability and point estimates for diverse prediction horizons during follow-up. The experiment highlights the potential of GRU-D-Weibull as a suitable candidate for individualized endpoint risk management, utilizing real-time clinical data to generate various endpoint estimates for monitoring. Additional research is warranted to evaluate the influence of different data quality aspects on prediction performance. Furthermore, collaboration with clinicians is essential to explore the integration of this approach into clinical workflows and evaluate its effects on decision-making processes and patient outcomes. Accurate prediction models for individual-level endpoints and time-to-endpoints are crucial in clinical practice. In this study, we propose a novel approach, GRU-D-Weibull, which combines gated recurrent units with decay (GRU-D) to model the Weibull distribution. Our method enables real-time individualized endpoint prediction and population-level risk management. Using a cohort of 6879 patients with stage 4 chronic kidney disease (CKD4), we evaluated the performance of GRU-D-Weibull in endpoint prediction. The C-index of GRU-D-Weibull was ~0.7 at the index date and increased to ~0.77 after 4.3 years of follow-up, similar to random survival forest. Our approach achieved an absolute L1-loss of ~1.1 years (SD≈0.95) at the CKD4 index date and a minimum of ~0.45 years (SD≈0.3) at 4 years of follow-up, outperforming competing methods significantly. GRU-D-Weibull consistently constrained the predicted survival probability at the time of an event within a smaller and more fixed range compared to other models throughout the follow-up period. We observed significant correlations between the error in point estimates and missing proportions of input features at the index date (correlations from ~0.1 to ~0.3), which diminished within 1 year as more data became available. By post-training recalibration, we aligned the predicted and observed survival probabilities across multiple prediction horizons at different time points during follow-up. The findings underscore the promise of the GRU-D-Weibull architecture as a suitable candidate for real-time individualized CKD4 endpoint risk management, calling for additional refinements and clinical confirmation. Xiaoyang Ruan, Liwei Wang 0010, Charat Thongprayoon, Wisit Cheungpasitporn |
Artif. Intell. Medicine | 2 |
| 2023 | An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)abstractDespite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts. Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Quality Assessment of Functional Status Documentation in EHR Across Institutions
Sunyang Fu, Maria Vassilaki, Omar A. Ibrahim, Ronald C. Petersen, Jennifer L. St. Sauver, Liwei Wang 0010, Jungwei Fan 0001, Sunghwan Sohn |
AMIA | 6 |
| 2022 | Towards User-centered Corpus Development: Lessons Learnt from Designing and Developing MedTator
Sunyang Fu, Liwei Wang 0010, Andrew Wen, Sijia Liu 0002, Sungrim Moon, Kurt Miller |
AMIA | 3 |
| 2022 | MedTator: a serverless annotation tool for corpus developmentabstractSUMMARY: Building a high-quality annotation corpus requires expenditure of considerable time and expertise, particularly for biomedical and clinical research applications. Most existing annotation tools provide many advanced features to cover a variety of needs where the installation, integration and difficulty of use present a significant burden for actual annotation tasks. Here, we present MedTator, a serverless annotation tool, aiming to provide an intuitive and interactive user interface that focuses on the core steps related to corpus annotation, such as document annotation, corpus summarization, annotation export and annotation adjudication. AVAILABILITY AND IMPLEMENTATION: MedTator and its tutorial are freely available from https://ohnlp.github.io/MedTator. MedTator source code is available under the Apache 2.0 license: https://github.com/OHNLP/MedTator. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sunyang Fu, Liwei Wang 0010, Sijia Liu 0002, Andrew Wen |
Bioinform. | 3 |
| 2022 | CancerBERT: a cancer domain-specific language model for extracting breast cancer phenotypes from electronic health recordsabstractOBJECTIVE: Accurate extraction of breast cancer patients' phenotypes is important for clinical decision support and clinical research. This study developed and evaluated cancer domain pretrained CancerBERT models for extracting breast cancer phenotypes from clinical texts. We also investigated the effect of customized cancer-related vocabulary on the performance of CancerBERT models. MATERIALS AND METHODS: A cancer-related corpus of breast cancer patients was extracted from the electronic health records of a local hospital. We annotated named entities in 200 pathology reports and 50 clinical notes for 8 cancer phenotypes for fine-tuning and evaluation. We kept pretraining the BlueBERT model on the cancer corpus with expanded vocabularies (using both term frequency-based and manually reviewed methods) to obtain CancerBERT models. The CancerBERT models were evaluated and compared with other baseline models on the cancer phenotype extraction task. RESULTS: All CancerBERT models outperformed all other models on the cancer phenotyping NER task. Both CancerBERT models with customized vocabularies outperformed the CancerBERT with the original BERT vocabulary. The CancerBERT model with manually reviewed customized vocabulary achieved the best performance with macro F1 scores equal to 0.876 (95% CI, 0.873-0.879) and 0.904 (95% CI, 0.902-0.906) for exact match and lenient match, respectively. CONCLUSIONS: The CancerBERT models were developed to extract the cancer phenotypes in clinical notes and pathology reports. The results validated that using customized vocabulary may further improve the performances of domain specific BERT models in clinical NLP tasks. The CancerBERT models developed in the study would further help clinical decision support. Liwei Wang 0010, Rui Zhang 0028 |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Patient Asynchronous Response to Coronavirus Disease 2019 (COVID-19): A Retrospective Analysis of Patient Portal Messages
Ming Huang 0006, Aditya Khurana, George M. Mastorakos, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Christi A. Patten, Jungwei Fan 0001 |
AMIA | 6 |
| 2021 | COVID-19 Dashboard: Visual Exploration of the Regional Pandemic Trend
Liwei Wang 0010, Andrew Wen, Ming Huang 0006, Yanshan Wang |
AMIA | 2 |
| 2021 | Disparity analysis of patient portal messaging use for COVID-19 in urban versus rural locality
Ming Huang 0006, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Nansu Zong, Yue Yu 0012, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Chyke Doubeni, Jungwei Fan 0001, Christi A. Patten |
AMIA | 4 |
| 2021 | Deep Learning Approaches for Breast Cancer Characteristics Extraction from Electronic Health Records
Liwei Wang 0010, Sunyang Fu, Chetan Shenoy, Anne H. Blaes, Rui Zhang 0028 |
AMIA | 2 |
| 2021 | An aberration detection-based approach for sentinel syndromic surveillance of COVID-19 and other novel influenza-like illnesses
Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Sunyang Fu, Sunghwan Sohn, Jacob A. Kugel, Vinod Kaggal, Ming Huang 0006, Yanshan Wang, Feichen Shen, Jungwei Fan 0001 |
J. Biomed. Informatics | 2 |
| 2020 | Investigating the impact of disease and health record duration on the eMERGE algorithm for rheumatoid arthritisabstractOBJECTIVE: The study sought to determine the dependence of the Electronic Medical Records and Genomics (eMERGE) rheumatoid arthritis (RA) algorithm on both RA and electronic health record (EHR) duration. MATERIALS AND METHODS: Using a population-based cohort from the Mayo Clinic Biobank, we identified 497 patients with at least 1 RA diagnosis code. RA case status was manually determined using validated criteria for RA. RA duration was defined as time from first RA code to the index date of biobank enrollment. To simulate EHR duration, various years of EHR lookback were applied, starting at the index date and going backward. Model performance was determined by sensitivity, specificity, positive predictive value, negative predictive value, and area under the curve (AUC). RESULTS: The eMERGE algorithm performed well in this cohort, with overall sensitivity 53%, specificity 99%, positive predictive value 97%, negative predictive value 74%, and AUC 76%. Among patients with RA duration <2 years, sensitivity and AUC were only 9% and 54%, respectively, but increased to 71% and 85% among patients with RA duration >10 years. Longer EHR lookback also improved model performance up to a threshold of 10 years, in which sensitivity reached 52% and AUC 75%. However, optimal EHR lookback varied by RA duration; an EHR lookback of 3 years was best able to identify recently diagnosed RA cases. CONCLUSIONS: eMERGE algorithm performance improves with longer RA duration as well as EHR duration up to 10 years, though shorter EHR lookback can improve identification of recently diagnosed RA cases. Vanessa L. Kronzer, Liwei Wang 0010, John M. Davis III, Jeffrey A. Sparks, Cynthia S. Crowson |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical eventsabstractOBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making. Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 11 |
| 2020 | Clinical concept extraction: A methodology review
Sunyang Fu, David Chen 0003, Sijia Liu 0002, Sungrim Moon, Kevin J. Peterson, Feichen Shen, Liwei Wang 0010, Yanshan Wang, Andrew Wen, Sunghwan Sohn |
J. Biomed. Informatics | 8 |
| 2019 | Association between Cardiotoxic Chemotherapy and Myocardial Infarction
Liwei Wang 0010, Suzette J. Bielinski, Paul A. Decker, Jill M. Killian, Nicholas B. Larson, Rui Zhang 0028 |
AMIA | 1 |
| 2019 | Achievability to Extract Specific Date Information for Cancer Research
Liwei Wang 0010, Jason A. Wampfler, Angela Dispenzieri, Hua Xu 0001 |
AMIA | 1 |
| 2019 | Ensembles of natural language processing systems for portable phenotyping solutions
Cong Liu 0020, Casey N. Ta, James R. Rogers, Ziran Li, Alex M. Butler, Ning Shang 0004, Fabricio Sampaio Peres Kury, Liwei Wang 0010, Feichen Shen, Lyudmila Ena, Carol Friedman, Chunhua Weng |
J. Biomed. Informatics | 9 |
| 2019 | HPO2Vec+: Leveraging heterogeneous knowledge resources to enrich node embeddings for the Human Phenotype Ontology
Feichen Shen, Su-Yuan Peng, Yadan Fan, Andrew Wen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010 |
J. Biomed. Informatics | 7 |
| 2018 | ARETA: A Corpus for Asthma Related Event Temporal Association
Sijia Liu 0002, Liwei Wang 0010, Sunghwan Sohn, Liping Xia |
AMIA | 2 |
| 2018 | Information Extraction for Data Population in Lung Cancer Clinical Research
Liwei Wang 0010, Yanshan Wang, Jason A. Wampfler |
AMIA | 1 |
| 2018 | Leveraging Association Rule Mining to Detect Pathophysiological Mechanisms of Chronic Kidney Disease Complicated by Metabolic Syndrome
Su-Yuan Peng, Yadan Fan, Liwei Wang 0010, Andrew Wen, Xu-Sheng Liu, Feichen Shen |
BIBM | 3 |
| 2018 | A comparison of word embeddings for the biomedical natural language processingabstractBACKGROUND: Word embeddings have been prevalently used in biomedical Natural Language Processing (NLP) applications due to the ability of the vector representations being able to capture useful semantic properties and linguistic relationships between words. Different textual resources (e.g., Wikipedia and biomedical literature corpus) have been utilized in biomedical NLP to train word embeddings and these word embeddings have been commonly leveraged as feature input to downstream machine learning models. However, there has been little work on evaluating the word embeddings trained from different textual resources. METHODS: In this study, we empirically evaluated word embeddings trained from four different corpora, namely clinical notes, biomedical publications, Wikipedia, and news. For the former two resources, we trained word embeddings using unstructured electronic health record (EHR) data available at Mayo Clinic and articles (MedLit) from PubMed Central, respectively. For the latter two resources, we used publicly available pre-trained word embeddings, GloVe and Google News. The evaluation was done qualitatively and quantitatively. For the qualitative evaluation, we randomly selected medical terms from three categories (i.e., disorder, symptom, and drug), and manually inspected the five most similar words computed by embeddings for each term. We also analyzed the word embeddings through a 2-dimensional visualization plot of 377 medical terms. For the quantitative evaluation, we conducted both intrinsic and extrinsic evaluation. For the intrinsic evaluation, we evaluated the word embeddings' ability to capture medical semantics by measruing the semantic similarity between medical terms using four published datasets: Pedersen's dataset, Hliaoutakis's dataset, MayoSRS, and UMNSRS. For the extrinsic evaluation, we applied word embeddings to multiple downstream biomedical NLP applications, including clinical information extraction (IE), biomedical information retrieval (IR), and relation extraction (RE), with data from shared tasks. RESULTS: The qualitative evaluation shows that the word embeddings trained from EHR and MedLit can find more similar medical terms than those trained from GloVe and Google News. The intrinsic quantitative evaluation verifies that the semantic similarity captured by the word embeddings trained from EHR is closer to human experts' judgments on all four tested datasets. The extrinsic quantitative evaluation shows that the word embeddings trained on EHR achieved the best F1 score of 0.900 for the clinical IE task; no word embeddings improved the performance for the biomedical IR task; and the word embeddings trained on Google News had the best overall F1 score of 0.790 for the RE task. CONCLUSION: Based on the evaluation results, we can draw the following conclusions. First, the word embeddings trained from EHR and MedLit can capture the semantics of medical terms better, and find semantically relevant medical terms closer to human experts' judgments than those trained from GloVe and Google News. Second, there does not exist a consistent global ranking of word embeddings for all downstream biomedical NLP applications. However, adding word embeddings as extra features will improve results on most downstream tasks. Finally, the word embeddings trained from the biomedical domain corpora do not necessarily have better performance than those trained from the general domain corpora for any downstream biomedical NLP task. Yanshan Wang, Sijia Liu 0002, Naveed Afzal, Majid Rastegar-Mojarad, Liwei Wang 0010, Feichen Shen, Paul R. Kingsbury |
J. Biomed. Informatics | 5 |
| 2018 | Clinical information extraction applications: A literature reviewabstractBACKGROUND: With the rapid adoption of electronic health records (EHRs), it is desirable to harvest information and knowledge from EHRs to support automated systems at the point of care and to enable secondary use of EHRs for clinical and translational research. One critical component used to facilitate the secondary use of EHR data is the information extraction (IE) task, which automatically extracts and encodes clinical information from text. OBJECTIVES: In this literature review, we present a review of recent published research on clinical information extraction (IE) applications. METHODS: A literature search was conducted for articles published from January 2009 to September 2016 based on Ovid MEDLINE In-Process & Other Non-Indexed Citations, Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science, and ACM Digital Library. RESULTS: A total of 1917 publications were identified for title and abstract screening. Of these publications, 263 articles were selected and discussed in this review in terms of publication venues and data sources, clinical IE tools, methods, and applications in the areas of disease- and drug-related studies, and clinical workflow optimizations. CONCLUSIONS: Clinical IE has been used for a wide range of applications, however, there is a considerable gap between clinical studies using EHR data and studies using clinical IE. This study enabled us to gain a more concrete understanding of the gap and to provide potential solutions to bridge this gap. Yanshan Wang, Liwei Wang 0010, Majid Rastegar-Mojarad, Sungrim Moon, Feichen Shen, Naveed Afzal, Sijia Liu 0002, Yuqun Zeng, Saeed Mehrabi 0003, Sunghwan Sohn |
J. Biomed. Informatics | 2 |
| 2017 | A Resource for Clinical Semantic Textual Similarity
Naveed Afzal, Yanshan Wang, Feichen Shen, Liwei Wang 0010 |
AMIA | 4 |
| 2017 | Leveraging Collaborative Filtering to Accelerate Rare Disease Diagnosis
Feichen Shen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010, Naveed Afzal |
AMIA | 4 |
| 2017 | Accelerating Rare Disease Diagnosis with Collaborative Filtering
Feichen Shen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010, Naveed Afzal |
AMIA | 4 |
| 2017 | Semantic Representations of Medical Terms: A Comparison Study of Word Embeddings Trained on Clinical Notes and PubMed Articles
Yanshan Wang, Naveed Afzal, Liwei Wang 0010, Feichen Shen |
AMIA | 3 |
| 2017 | Recommending education materials for diabetic questions using information retrieval approaches
Yuqun Zeng, Yanshan Wang, Feichen Shen, Sijia Liu 0002, Liwei Wang 0010, Majid Rastegar-Mojarad, Xu-Sheng Liu |
AMIA | 5 |
| 2016 | Ontology-Based Analysis of Adverse Drug Reactions Associated with Anti-infection Drugs in China
Yuying Cao, Liwei Wang 0010, Yongqun He |
AMIA | 3 |
| 2016 | Discovering Associations between Problem List and Practice Setting
Liwei Wang 0010, Dingcheng Li, Majid Rastegar-Mojarad |
AMIA | 1 |
| 2016 | Answering patients' questions using expert-vetted online resources: A case study of diabetes
Yuqun Zeng, Liwei Wang 0010, Yanshan Wang, Dingcheng Li |
AMIA | 2 |
| 2016 | Answering diabetic patients' questions using expert-vetted online resources: A case studyabstractAs an increasing number of patients seek medical information online, it is crucial to bring expert-vetted information to patients. Multiple expert-vetted online resources exist for their accuracy and authority of medical knowledge. However, it is not clear which resource is better in meeting the information needs of the patients. Utilizing a collection of questions raised by patients with diabetes retrieved from an online forum, we manually evaluated three widely used expert-vetted online resources (WebMD, MedlinePlus, and UpToDate) in terms of content coverage and time spent to find answers. The results indicated that WebMD had slightly better content coverage with less time spent in finding answers. Leveraging the Natural Language Processing (NLP) techniques and a clinical terminology resource, the Unified Medical Language System (UMLS), we demonstrated that WebMD had a higher mapping rate of clinical concepts comparing to other resources. Yuqun Zeng, Xu-Sheng Liu, Liwei Wang 0010, Yanshan Wang |
BIBM | 3 |
| 2013 | Coreference Resolution from Medical Corpus with Topic Modeling
Dingcheng Li, Liwei Wang 0010, Cui Tao, Christopher G. Chute |
AMIA | 2 |
| 2013 | Profiling adverse drug events of cancer drug ingredients using normalized AERS dataabstractTo facilitate the utilization of FDA's adverse event reporting system (AERS) for data mining, we previously normalized AERS and aggregated related data into a data set (AERS-DM). In this paper, we aim to demonstrate the data mining potential of AERS-DM by profiling cancer drug ingredients. Findings suggest that the co-relationship may exist between adverse drug events (ADEs) and mechanism of action of cancer drug ingredients, between ADEs and physiologic effect, and between ADEs and treatment intention. We speculate that such co-relationship may provide a new direction to explore the etiology of ADEs. In addition, age and sex differences in ADEs for those ingredients are revealed, among them what haven't been discovered before may be used as hypotheses for detecting drug safety signal for further investigations. In conclusion, the discoveries in this study show the potential of AERS-DM in data mining for profiling ADEs of cancer drug ingredients. Liwei Wang 0010 |
BIBM | 1 |
| 2012 | A Preliminary Study on Normalizing AERS Drug Names Using RxNorm
Liwei Wang 0010, Guoqian Jiang |
AMIA | 1 |