Sijia Liu 0002

dblp:128/6972-2 · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
9since 2021 · last 2023
0000-0001-9763-1164ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2023 An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)
abstract
Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts.
Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001
J. Am. Medical Informatics Assoc.1
2022 Towards User-centered Corpus Development: Lessons Learnt from Designing and Developing MedTator
Sunyang Fu, Liwei Wang 0010, Andrew Wen, Sijia Liu 0002, Sungrim Moon, Kurt Miller
AMIA5
2022 MedTator: a serverless annotation tool for corpus development
abstract
SUMMARY: Building a high-quality annotation corpus requires expenditure of considerable time and expertise, particularly for biomedical and clinical research applications. Most existing annotation tools provide many advanced features to cover a variety of needs where the installation, integration and difficulty of use present a significant burden for actual annotation tasks. Here, we present MedTator, a serverless annotation tool, aiming to provide an intuitive and interactive user interface that focuses on the core steps related to corpus annotation, such as document annotation, corpus summarization, annotation export and annotation adjudication. AVAILABILITY AND IMPLEMENTATION: MedTator and its tutorial are freely available from https://ohnlp.github.io/MedTator. MedTator source code is available under the Apache 2.0 license: https://github.com/OHNLP/MedTator. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sunyang Fu, Liwei Wang 0010, Sijia Liu 0002, Andrew Wen
Bioinform.4
2022 Developing an ETL tool for converting the PCORnet CDM into the OMOP CDM to facilitate the COVID-19 data integration
Yue Yu 0012, Nansu Zong, Andrew Wen, Sijia Liu 0002, Daniel J. Stone, David Knaack, Alanna M. Chamberlain, Emily R. Pfaff, Davera Gabriel, Christopher G. Chute, Nilay Shah, Guoqian Jiang
J. Biomed. Informatics4
2021 Patient Asynchronous Response to Coronavirus Disease 2019 (COVID-19): A Retrospective Analysis of Patient Portal Messages
Ming Huang 0006, Aditya Khurana, George M. Mastorakos, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Christi A. Patten, Jungwei Fan 0001
AMIA7
2021 Disparity analysis of patient portal messaging use for COVID-19 in urban versus rural locality
Ming Huang 0006, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Nansu Zong, Yue Yu 0012, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Chyke Doubeni, Jungwei Fan 0001, Christi A. Patten
AMIA5
2021 FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang
AMIA2
2021 On Constraints and Considerations for Extending Support for Natural Language Processing-Based FHIR Resource Generation
Andrew Wen, Luke V. Rasmussen, Daniel J. Stone, Sijia Liu 0002, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Yuan Luo 0001, Fei Wang 0001, Jyotishman Pathak, Guoqian Jiang
AMIA4
2021 An aberration detection-based approach for sentinel syndromic surveillance of COVID-19 and other novel influenza-like illnesses
Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Sunyang Fu, Sunghwan Sohn, Jacob A. Kugel, Vinod Kaggal, Ming Huang 0006, Yanshan Wang, Feichen Shen, Jungwei Fan 0001
J. Biomed. Informatics4
2020 Predicting Section Location of Clinical Sentences using BERT Encoder - A Pilot Study
Sijia Liu 0002, Sunyang Fu, Sungrim Moon, Andrew Wen
AMIA1
2020 Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical events
abstract
OBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making.
Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao
J. Am. Medical Informatics Assoc.10
2020 Clinical concept extraction: A methodology review
Sunyang Fu, David Chen 0003, Sijia Liu 0002, Sungrim Moon, Kevin J. Peterson, Feichen Shen, Liwei Wang 0010, Yanshan Wang, Andrew Wen, Sunghwan Sohn
J. Biomed. Informatics4
2019 Clinical Use of an Information Retrieval Framework for Cohort Discovery from Electronic Health Records
Yanshan Wang, Andrew Wen, Sijia Liu 0002, Jennifer L. St. Sauver, Adil E. Bharucha, Chunhua Weng
AMIA3
2019 Enhancing Clinical Information Retrieval through Context-Aware Queries and Indices
abstract
The big data revolution has created a hefty demand for searching large-scale electronic health records (EHRs) to support clinical practice, research, and administration. Despite the volume of data involved, fast and accurate identification of clinical narratives pertinent to a clinical case being seen by any given provider is crucial for decision-making at the point of care. In the general domain, this capability is accomplished through a combination of the inverted index data structure, horizontal scaling, and information retrieval (IR) scoring algorithms. These technologies are also being used in the clinical domain, but have met limited success, particularly as clinical cases become more complex. One barrier affecting clinical performance is that contextual information, such as negation, temporality, and the subject of clinical mentions, impact clinical relevance but is not considered in general IR methodologies. In this study, we implemented a solution by identifying and incorporating the aforementioned semantic contexts as part of IR indexing/scoring with Elasticsearch. Experiments were conducted in comparison to baseline approaches with respect to: 1) evaluation of the impact on the quality (relevance) of the returned results, and 2) evaluation of the impact on execution time and storage requirements. The results showed a 5.1-23.1% improvement in retrieval quality, along with achieving 35% faster query execution time. Cost-wise, the solution required 1.5-2 times larger space and about 3 times increase in indexing time. The higher relevance demonstrated the merit of incorporating contextual information into clinical IR, and the near-constant increase in time and space suggested promising scalability.
Andrew Wen, Yanshan Wang, Vinod Kaggal, Sijia Liu 0002, Jungwei Fan 0001
IEEE BigData4
2019 HPO2Vec+: Leveraging heterogeneous knowledge resources to enrich node embeddings for the Human Phenotype Ontology
Feichen Shen, Su-Yuan Peng, Yadan Fan, Andrew Wen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010
J. Biomed. Informatics5
2018 ARETA: A Corpus for Asthma Related Event Temporal Association
Sijia Liu 0002, Liwei Wang 0010, Sunghwan Sohn, Liping Xia
AMIA1
2018 Leveraging Electronic Health Records for Identification of Features Associated with Sudden Death for Hypertrophic Cardiomyopathy Patients
Sungrim Moon, Sijia Liu 0002, Sujith Samudrala, Jane L. Shellum, Jeffrey B. Geske, Peter A. Noseworthy, Steve Ommen, Rajeev Chaudhry, Rick Nishimura, Adelaide M. Arruda-Olson
AMIA2
2018 EMIRS: An Electronic Medical Information Retrieval System by Leveraging both Structured and Unstructured Electronic Health Records
Yanshan Wang, Andrew Wen, Sijia Liu 0002
AMIA3
2018 A comparison of word embeddings for the biomedical natural language processing
abstract
BACKGROUND: Word embeddings have been prevalently used in biomedical Natural Language Processing (NLP) applications due to the ability of the vector representations being able to capture useful semantic properties and linguistic relationships between words. Different textual resources (e.g., Wikipedia and biomedical literature corpus) have been utilized in biomedical NLP to train word embeddings and these word embeddings have been commonly leveraged as feature input to downstream machine learning models. However, there has been little work on evaluating the word embeddings trained from different textual resources. METHODS: In this study, we empirically evaluated word embeddings trained from four different corpora, namely clinical notes, biomedical publications, Wikipedia, and news. For the former two resources, we trained word embeddings using unstructured electronic health record (EHR) data available at Mayo Clinic and articles (MedLit) from PubMed Central, respectively. For the latter two resources, we used publicly available pre-trained word embeddings, GloVe and Google News. The evaluation was done qualitatively and quantitatively. For the qualitative evaluation, we randomly selected medical terms from three categories (i.e., disorder, symptom, and drug), and manually inspected the five most similar words computed by embeddings for each term. We also analyzed the word embeddings through a 2-dimensional visualization plot of 377 medical terms. For the quantitative evaluation, we conducted both intrinsic and extrinsic evaluation. For the intrinsic evaluation, we evaluated the word embeddings' ability to capture medical semantics by measruing the semantic similarity between medical terms using four published datasets: Pedersen's dataset, Hliaoutakis's dataset, MayoSRS, and UMNSRS. For the extrinsic evaluation, we applied word embeddings to multiple downstream biomedical NLP applications, including clinical information extraction (IE), biomedical information retrieval (IR), and relation extraction (RE), with data from shared tasks. RESULTS: The qualitative evaluation shows that the word embeddings trained from EHR and MedLit can find more similar medical terms than those trained from GloVe and Google News. The intrinsic quantitative evaluation verifies that the semantic similarity captured by the word embeddings trained from EHR is closer to human experts' judgments on all four tested datasets. The extrinsic quantitative evaluation shows that the word embeddings trained on EHR achieved the best F1 score of 0.900 for the clinical IE task; no word embeddings improved the performance for the biomedical IR task; and the word embeddings trained on Google News had the best overall F1 score of 0.790 for the RE task. CONCLUSION: Based on the evaluation results, we can draw the following conclusions. First, the word embeddings trained from EHR and MedLit can capture the semantics of medical terms better, and find semantically relevant medical terms closer to human experts' judgments than those trained from GloVe and Google News. Second, there does not exist a consistent global ranking of word embeddings for all downstream biomedical NLP applications. However, adding word embeddings as extra features will improve results on most downstream tasks. Finally, the word embeddings trained from the biomedical domain corpora do not necessarily have better performance than those trained from the general domain corpora for any downstream biomedical NLP task.
Yanshan Wang, Sijia Liu 0002, Naveed Afzal, Majid Rastegar-Mojarad, Liwei Wang 0010, Feichen Shen, Paul R. Kingsbury
J. Biomed. Informatics2
2018 Clinical information extraction applications: A literature review
abstract
BACKGROUND: With the rapid adoption of electronic health records (EHRs), it is desirable to harvest information and knowledge from EHRs to support automated systems at the point of care and to enable secondary use of EHRs for clinical and translational research. One critical component used to facilitate the secondary use of EHR data is the information extraction (IE) task, which automatically extracts and encodes clinical information from text. OBJECTIVES: In this literature review, we present a review of recent published research on clinical information extraction (IE) applications. METHODS: A literature search was conducted for articles published from January 2009 to September 2016 based on Ovid MEDLINE In-Process & Other Non-Indexed Citations, Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science, and ACM Digital Library. RESULTS: A total of 1917 publications were identified for title and abstract screening. Of these publications, 263 articles were selected and discussed in this review in terms of publication venues and data sources, clinical IE tools, methods, and applications in the areas of disease- and drug-related studies, and clinical workflow optimizations. CONCLUSIONS: Clinical IE has been used for a wide range of applications, however, there is a considerable gap between clinical studies using EHR data and studies using clinical IE. This study enabled us to gain a more concrete understanding of the gap and to provide potential solutions to bridge this gap.
Yanshan Wang, Liwei Wang 0010, Majid Rastegar-Mojarad, Sungrim Moon, Feichen Shen, Naveed Afzal, Sijia Liu 0002, Yuqun Zeng, Saeed Mehrabi 0003, Sunghwan Sohn
J. Biomed. Informatics7
2018 Modeling asynchronous event sequences with RNNs
Stephen T. Wu, Sijia Liu 0002, Sunghwan Sohn, Sungrim Moon, Chung-Il Wi, Young J. Juhn
J. Biomed. Informatics2
2017 Leveraging Collaborative Filtering to Accelerate Rare Disease Diagnosis
Feichen Shen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010, Naveed Afzal
AMIA2
2017 Accelerating Rare Disease Diagnosis with Collaborative Filtering
Feichen Shen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010, Naveed Afzal
AMIA2
2017 Recommending education materials for diabetic questions using information retrieval approaches
Yuqun Zeng, Yanshan Wang, Feichen Shen, Sijia Liu 0002, Liwei Wang 0010, Majid Rastegar-Mojarad, Xu-Sheng Liu
AMIA4
2017 Medical concept intersection between outside medical records and consultant notes: A case study in transferred cardiovascular patients
abstract
One of the promises of “meaningful use” of Electronic Health Records (EHRs) is to facilitate digital information exchange between healthcare providers through continuity of care documents. Despite such promise, outside medical records (OMRs) of referral patients including clinical notes, lab test results or diagnostic test reports are frequently provided through fax or print out. Moreover, it is not clear how much information in those OMRs is utilized when providing care at the early stage. In this study, we collected clinical concepts automatically from OMRs through optical character recognition (OCR) technology and then performed a quantitative analysis of concepts presented in OMRs and concepts captured in clinical notes at Mayo Clinic. We also investigated information from OMRs not captured in initial consultant notes but presented over subsequent consultant notes. We identified 12.93% of concepts from OMRs were identified in clinical documents within three months. Among those overlapping concepts, 26.74% of them were not captured in initial consultant notes. Our study presents that clinical information from OMRs is important for patient care. Also, the delayed presence of information in clinical notes may indicate important information from OMRs is not fully utilized earlier in the care.
Sungrim Moon, Sijia Liu 0002, Paul R. Kingsbury, David Chen 0003, Yanshan Wang, Feichen Shen, Rajeev Chaudhry
BIBM2
2017 Intrainstitutional EHR collections for patient-level information retrieval
abstract
Research in clinical information retrieval has long been stymied by the lack of open resources. However, both clinical information retrieval research innovation and legitimate privacy concerns can be served by the creation of intrainstitutional, fully protected resources. In this article, we provide some principles and tools for information retrieval resource‐building in the unique problem setting of patient‐level information retrieval, following the tradition of the Cranfield paradigm. We further include an analysis of parallel information retrieval resources at Oregon Health & Science University and Mayo Clinic that were built on these principles.
Stephen T. Wu, Sijia Liu 0002, Yanshan Wang, Tamara Timmons, Harsha Uppili, Steven Bedrick, William R. Hersh
J. Assoc. Inf. Sci. Technol.2
2016 A Topic-modeling Based Framework for Drug-drug Interaction Classification from Biomedical Text
Dingcheng Li, Sijia Liu 0002, Majid Rastegar-Mojarad, Yanshan Wang, Vipin Chaudhary, Terry M. Therneau
AMIA2
2016 Drug-drug Interaction Detection with A Topic-modeling Based Framework Augmented with Distant-supervision
Dingcheng Li, Sijia Liu 0002, Majid Rastegar-Mojarad, Yanshan Wang, Vipin Chaudhary, Terry M. Therneau
AMIA2