Hao Liu 0054

dblp:09/3214-54 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-1975-1272ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 11 since 2021
YearPublicationVenuePosition
2025 Advancing Drug-Drug Interaction Prediction using Multi-Modal Feature Integration with Graph Neural Networks
abstract
Pharmaceutical treatments are essential for managing medical conditions, but drug-drug interactions (DDIs) pose significant risks. This research integrates Knowledge Graphs and Graph Neural Networks to predict DDIs by exploring drug relationships. We construct a knowledge graph using DrugBank data (1,000 drugs, 155,774 interactions) enriched with PubChem features, then enhance our approach by integrating transformer-based embeddings (ChemBERTa, SPECTER, and SBERT) to create 1152-dimensional feature vectors. Formulating DDI prediction as link prediction, we compare three GNN architectures: Graph Convolutional Network (GCN), GraphSAGE, and Graph Attention Network (GAT). With basic molecular features, GCN achieved 75.65% accuracy (80.17% F1). After multimodal integration, performance improved across all models, with GAT showing the greatest enhancement (80.61% accuracy, 82.57% F1). These results highlight the value of integrating diverse data modalities for DDI prediction and the potential for enhancing medication safety in polypharmacy scenarios.
Ernest C. Chianumba, Hao Liu 0054, Aparna S. Varde
BIBM2
2023 The suitability of UMLS and SNOMED-CT for encoding outcome concepts
abstract
OBJECTIVE: Outcomes are important clinical study information. Despite progress in automated extraction of PICO (Population, Intervention, Comparison, and Outcome) entities from PubMed, rarely are these entities encoded by standard terminology to achieve semantic interoperability. This study aims to evaluate the suitability of the Unified Medical Language System (UMLS) and SNOMED-CT in encoding outcome concepts in randomized controlled trial (RCT) abstracts. MATERIALS AND METHODS: We iteratively developed and validated an outcome annotation guideline and manually annotated clinically significant outcome entities in the Results and Conclusions sections of 500 randomly selected RCT abstracts on PubMed. The extracted outcomes were fully, partially, or not mapped to the UMLS via MetaMap based on established heuristics. Manual UMLS browser search was performed for select unmapped outcome entities to further differentiate between UMLS and MetaMap errors. RESULTS: Only 44% of 2617 outcome concepts were fully covered in the UMLS, among which 67% were complex concepts that required the combination of 2 or more UMLS concepts to represent them. SNOMED-CT was present as a source in 61% of the fully mapped outcomes. DISCUSSION: Domains such as Metabolism and Nutrition, and Infections and Infectious Diseases need expanded outcome concept coverage in the UMLS and MetaMap. Future work is warranted to similarly assess the terminology coverage for P, I, C entities. CONCLUSION: Computational representation of clinical outcomes is important for clinical evidence extraction and appraisal and yet faces challenges from the inherent complexity and lack of coverage of these concepts in UMLS and SNOMED-CT, as demonstrated in this study.
Abigail M. Newbury, Hao Liu 0054, Betina Ross S. Idnay, Chunhua Weng
J. Am. Medical Informatics Assoc.2
2023 A data-driven approach to optimizing clinical study eligibility criteria
Yilu Fang, Hao Liu 0054, Betina Ross S. Idnay, Casey N. Ta, Karen Marder, Chunhua Weng
J. Biomed. Informatics2
2022 Criteria2Query 2.0: Combining Machine Efficiency and Human Intelligence to Define a More Accurate and Feasible Cohort for Clinical Trial Recruitment
Betina Ross S. Idnay, Yilu Fang, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Rebecca Schnall, Chunhua Weng
AMIA4
2022 Combining human and machine intelligence for clinical trial eligibility querying
abstract
OBJECTIVE: To combine machine efficiency and human intelligence for converting complex clinical trial eligibility criteria text into cohort queries. MATERIALS AND METHODS: Criteria2Query (C2Q) 2.0 was developed to enable real-time user intervention for criteria selection and simplification, parsing error correction, and concept mapping. The accuracy, precision, recall, and F1 score of enhanced modules for negation scope detection, temporal and value normalization were evaluated using a previously curated gold standard, the annotated eligibility criteria of 1010 COVID-19 clinical trials. The usability and usefulness were evaluated by 10 research coordinators in a task-oriented usability evaluation using 5 Alzheimer's disease trials. Data were collected by user interaction logging, a demographic questionnaire, the Health Information Technology Usability Evaluation Scale (Health-ITUES), and a feature-specific questionnaire. RESULTS: The accuracies of negation scope detection, temporal and value normalization were 0.924, 0.916, and 0.966, respectively. C2Q 2.0 achieved a moderate usability score (3.84 out of 5) and a high learnability score (4.54 out of 5). On average, 9.9 modifications were made for a clinical study. Experienced researchers made more modifications than novice researchers. The most frequent modification was deletion (5.35 per study). Furthermore, the evaluators favored cohort queries resulting from modifications (score 4.1 out of 5) and the user engagement features (score 4.3 out of 5). DISCUSSION AND CONCLUSION: Features to engage domain experts and to overcome the limitations in automated machine output are shown to be useful and user-friendly. We concluded that human-computer collaboration is key to improving the adoption and user-friendliness of natural language processing.
Yilu Fang, Betina Ross S. Idnay, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Karen Marder, Hua Xu 0001, Rebecca Schnall, Chunhua Weng
J. Am. Medical Informatics Assoc.4
2022 Ontology-based categorization of clinical studies by their conditions
Hao Liu 0054, Simona Carini, Zhehuan Chen, Spencer Phillips Hey, Ida Sim, Chunhua Weng
J. Biomed. Informatics1
2021 Misalignment between COVID-19 hotspots and clinical trial sites
abstract
Hundreds of interventional clinical trials have been launched in the United States to identify effective treatment strategies for combating the coronavirus disease 2019 (COVID-19) pandemic. However, to date, only a small fraction of these trials have completed enrollment, delaying the scientific investigation of COVID-19 and its treatment options. This study presents novel metrics to examine the geographic alignment between COVID-19 hotspots and interventional clinical trial sites and evaluate trial access over time during the evolving pandemic. Using temporal COVID-19 case data from USAFacts.org and trial data from ClinicalTrials.gov, U.S. counties were categorized based on their numbers of cases and trials. Our analysis suggests that alignment and access have worsened as the pandemic shifted over time. We recommend strategies and metrics to evaluate the alignment between cases and trials. Future studies are warranted to investigate the impact of the misalignment of cases and clinical trial sites on clinical trial recruitment.
Lauren Franks, Hao Liu 0054, Mitchell S. V. Elkind, Muredach P. Reilly, Chunhua Weng, Shing M. Lee
J. Am. Medical Informatics Assoc.2
2021 Towards clinical data-driven eligibility criteria optimization for interventional COVID-19 clinical trials
abstract
OBJECTIVE: This research aims to evaluate the impact of eligibility criteria on recruitment and observable clinical outcomes of COVID-19 clinical trials using electronic health record (EHR) data. MATERIALS AND METHODS: On June 18, 2020, we identified frequently used eligibility criteria from all the interventional COVID-19 trials in ClinicalTrials.gov (n = 288), including age, pregnancy, oxygen saturation, alanine/aspartate aminotransferase, platelets, and estimated glomerular filtration rate. We applied the frequently used criteria to the EHR data of COVID-19 patients in Columbia University Irving Medical Center (CUIMC) (March 2020-June 2020) and evaluated their impact on patient accrual and the occurrence of a composite endpoint of mechanical ventilation, tracheostomy, and in-hospital death. RESULTS: There were 3251 patients diagnosed with COVID-19 from the CUIMC EHR included in the analysis. The median follow-up period was 10 days (interquartile range 4-28 days). The composite events occurred in 18.1% (n = 587) of the COVID-19 cohort during the follow-up. In a hypothetical trial with common eligibility criteria, 33.6% (690/2051) were eligible among patients with evaluable data and 22.2% (153/690) had the composite event. DISCUSSION: By adjusting the thresholds of common eligibility criteria based on the characteristics of COVID-19 patients, we could observe more composite events from fewer patients. CONCLUSIONS: This research demonstrated the potential of using the EHR data of COVID-19 patients to inform the selection of eligibility criteria and their thresholds, supporting data-driven optimization of participant selection towards improved statistical power of COVID-19 trials.
Jae Hyun Kim, Casey N. Ta, Cong Liu 0020, Cynthia Sung 0002, Alex M. Butler, Latoya A. Stewart, Lyudmila Ena, James R. Rogers, Anna Ostropolets, Patrick B. Ryan, Hao Liu 0054, Shing M. Lee, Mitchell S. V. Elkind, Chunhua Weng
J. Am. Medical Informatics Assoc.12
2021 The COVID-19 Trial Finder
abstract
Clinical trials are the gold standard for generating reliable medical evidence. The biggest bottleneck in clinical trials is recruitment. To facilitate recruitment, tools for patient search of relevant clinical trials have been developed, but users often suffer from information overload. With nearly 700 coronavirus disease 2019 (COVID-19) trials conducted in the United States as of August 2020, it is imperative to enable rapid recruitment to these studies. The COVID-19 Trial Finder was designed to facilitate patient-centered search of COVID-19 trials, first by location and radius distance from trial sites, and then by brief, dynamically generated medical questions to allow users to prescreen their eligibility for nearby COVID-19 trials with minimum human computer interaction. A simulation study using 20 publicly available patient case reports demonstrates its precision and effectiveness.
Yingcheng Sun, Alex M. Butler, Fengyang Lin, Hao Liu 0054, Latoya A. Stewart, Jae Hyun Kim, Betina Ross S. Idnay, Qingyin Ge, Xinyi Wei, Cong Liu 0020, Chi Yuan, Chunhua Weng
J. Am. Medical Informatics Assoc.4
2021 A knowledge base of clinical trial eligibility criteria
Hao Liu 0054, Chi Yuan, Alex M. Butler, Yingcheng Sun, Chunhua Weng
J. Biomed. Informatics1
2021 Building an OMOP common data model-compliant annotated corpus for COVID-19 clinical trials
Yingcheng Sun, Alex M. Butler, Latoya A. Stewart, Hao Liu 0054, Chi Yuan, Christopher T. Southard, Jae Hyun Kim, Chunhua Weng
J. Biomed. Informatics4
2020 Potential Role of Clinical Trial Eligibility Criteria in Electronic Phenotyping
Hao Liu 0054, Chi Yuan, Alex M. Butler, Yingcheng Sun, Chunhua Weng
AMIA1