EDBT 2026 Demo / reviewers in the wild / expert
Cong Liu 0020
dblp:95/6404-20
· DBLP profile ↗
15ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0001-6024-3037ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fine-tuning large language models for rare disease concept normalizationabstractOBJECTIVE: We aim to develop a novel method for rare disease concept normalization by fine-tuning Llama 2, an open-source large language model (LLM), using a domain-specific corpus sourced from the Human Phenotype Ontology (HPO). METHODS: We developed an in-house template-based script to generate two corpora for fine-tuning. The first (NAME) contains standardized HPO names, sourced from the HPO vocabularies, along with their corresponding identifiers. The second (NAME+SYN) includes HPO names and half of the concept's synonyms as well as identifiers. Subsequently, we fine-tuned Llama 2 (Llama2-7B) for each sentence set and conducted an evaluation using a range of sentence prompts and various phenotype terms. RESULTS: When the phenotype terms for normalization were included in the fine-tuning corpora, both models demonstrated nearly perfect performance, averaging over 99% accuracy. In comparison, ChatGPT-3.5 has only ∼20% accuracy in identifying HPO IDs for phenotype terms. When single-character typos were introduced in the phenotype terms, the accuracy of NAME and NAME+SYN is 10.2% and 36.1%, respectively, but increases to 61.8% (NAME+SYN) with additional typo-specific fine-tuning. For terms sourced from HPO vocabularies as unseen synonyms, the NAME model achieved 11.2% accuracy, while the NAME+SYN model achieved 92.7% accuracy. CONCLUSION: Our fine-tuned models demonstrate ability to normalize phenotype terms unseen in the fine-tuning corpus, including misspellings, synonyms, terms from other ontologies, and laymen's terms. Our approach provides a solution for the use of LLMs to identify named medical entities from clinical narratives, while successfully normalizing them to standard concepts in a controlled vocabulary. Cong Liu 0020, Jingye Yang, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Converting OMOP CDM to phenopackets: A model alignment and patient data representation evaluation
Kayla Schiffer, Cong Liu 0020, Tiffany Callahan, Casey N. Ta, Jordan G. Nestor, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2023 | Characterizing variability of electronic health record-driven phenotype definitionsabstractOBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic. Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | Deep learning for rare disease: A scoping review
Cong Liu 0020, Zhehuan Chen, Yingcheng Sun, James R. Rogers, Wendy K. Chung, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2022 | Leveraging electronic health record data for clinical trial planning by assessing eligibility criteria's impact on patient count and safety
James R. Rogers, Jovana Pavisic, Casey N. Ta, Cong Liu 0020, Ali Soroush, Ying Kuen Cheung, George Hripcsak, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2021 | Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei |
AMIA | 11 |
| 2021 | Towards clinical data-driven eligibility criteria optimization for interventional COVID-19 clinical trialsabstractOBJECTIVE: This research aims to evaluate the impact of eligibility criteria on recruitment and observable clinical outcomes of COVID-19 clinical trials using electronic health record (EHR) data. MATERIALS AND METHODS: On June 18, 2020, we identified frequently used eligibility criteria from all the interventional COVID-19 trials in ClinicalTrials.gov (n = 288), including age, pregnancy, oxygen saturation, alanine/aspartate aminotransferase, platelets, and estimated glomerular filtration rate. We applied the frequently used criteria to the EHR data of COVID-19 patients in Columbia University Irving Medical Center (CUIMC) (March 2020-June 2020) and evaluated their impact on patient accrual and the occurrence of a composite endpoint of mechanical ventilation, tracheostomy, and in-hospital death. RESULTS: There were 3251 patients diagnosed with COVID-19 from the CUIMC EHR included in the analysis. The median follow-up period was 10 days (interquartile range 4-28 days). The composite events occurred in 18.1% (n = 587) of the COVID-19 cohort during the follow-up. In a hypothetical trial with common eligibility criteria, 33.6% (690/2051) were eligible among patients with evaluable data and 22.2% (153/690) had the composite event. DISCUSSION: By adjusting the thresholds of common eligibility criteria based on the characteristics of COVID-19 patients, we could observe more composite events from fewer patients. CONCLUSIONS: This research demonstrated the potential of using the EHR data of COVID-19 patients to inform the selection of eligibility criteria and their thresholds, supporting data-driven optimization of participant selection towards improved statistical power of COVID-19 trials. Jae Hyun Kim, Casey N. Ta, Cong Liu 0020, Cynthia Sung 0002, Alex M. Butler, Latoya A. Stewart, Lyudmila Ena, James R. Rogers, Anna Ostropolets, Patrick B. Ryan, Hao Liu 0054, Shing M. Lee, Mitchell S. V. Elkind, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | The COVID-19 Trial FinderabstractClinical trials are the gold standard for generating reliable medical evidence. The biggest bottleneck in clinical trials is recruitment. To facilitate recruitment, tools for patient search of relevant clinical trials have been developed, but users often suffer from information overload. With nearly 700 coronavirus disease 2019 (COVID-19) trials conducted in the United States as of August 2020, it is imperative to enable rapid recruitment to these studies. The COVID-19 Trial Finder was designed to facilitate patient-centered search of COVID-19 trials, first by location and radius distance from trial sites, and then by brief, dynamically generated medical questions to allow users to prescreen their eligibility for nearby COVID-19 trials with minimum human computer interaction. A simulation study using 20 publicly available patient case reports demonstrates its precision and effectiveness. Yingcheng Sun, Alex M. Butler, Fengyang Lin, Hao Liu 0054, Latoya A. Stewart, Jae Hyun Kim, Betina Ross S. Idnay, Qingyin Ge, Xinyi Wei, Cong Liu 0020, Chi Yuan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 10 |
| 2019 | DQueST: dynamic questionnaire for search of clinical trialsabstractOBJECTIVE: Information overload remains a challenge for patients seeking clinical trials. We present a novel system (DQueST) that reduces information overload for trial seekers using dynamic questionnaires. MATERIALS AND METHODS: DQueST first performs information extraction and criteria library curation. DQueST transforms criteria narratives in the ClinicalTrials.gov repository into a structured format, normalizes clinical entities using standard concepts, clusters related criteria, and stores the resulting curated library. DQueST then implements a real-time dynamic question generation algorithm. During user interaction, the initial search is similar to a standard search engine, and then DQueST performs real-time dynamic question generation to select criteria from the library 1 at a time by maximizing its relevance score that reflects its ability to rule out ineligible trials. DQueST dynamically updates the remaining trial set by removing ineligible trials based on user responses to corresponding questions. The process iterates until users decide to stop and begin manually reviewing the remaining trials. RESULTS: In simulation experiments initiated by 10 diseases, DQueST reduced information overload by filtering out 60%-80% of initial trials after 50 questions. Reviewing the generated questions against previous answers, on average, 79.7% of the questions were relevant to the queried conditions. By examining the eligibility of random samples of trials ruled out by DQueST, we estimate the accuracy of the filtering procedure is 63.7%. In a study using 5 mock patient profiles, DQueST on average retrieved trials with a 1.465 times higher density of eligible trials than an existing search engine. In a patient-centered usability evaluation, patients found DQueST useful, easy to use, and returning relevant results. CONCLUSION: DQueST contributes a novel framework for transforming free-text eligibility criteria to questions and filtering out clinical trials based on user answers to questions dynamically. It promises to augment keyword-based methods to improve clinical trial search. Cong Liu 0020, Chi Yuan, Alex M. Butler, Richard D. Carvajal, Ziran Ryan Li, Casey N. Ta, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Facilitating phenotype transfer using a common data model
George Hripcsak, Ning Shang 0004, Peggy L. Peissig, Luke V. Rasmussen, Cong Liu 0020, Barbara Benoit, Robert J. Carroll, David Carrell, Joshua C. Denny, Ozan Dikilitas, Vivian S. Gainer, Kayla Marie Howell, Jeffrey G. Klann, Iftikhar J. Kullo, Todd Lingren, Frank D. Mentch, Shawn N. Murphy, Karthik Natarajan, Chunhua Weng |
J. Biomed. Informatics | 5 |
| 2019 | Ensembles of natural language processing systems for portable phenotyping solutions
Cong Liu 0020, Casey N. Ta, James R. Rogers, Ziran Li, Alex M. Butler, Ning Shang 0004, Fabricio Sampaio Peres Kury, Liwei Wang 0010, Feichen Shen, Lyudmila Ena, Carol Friedman, Chunhua Weng |
J. Biomed. Informatics | 1 |
| 2019 | Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2018 | Dynamic Questionnaire Generation for Efficient Patient Search of Trials
Cong Liu 0020, Chi Yuan, Eric Pua, Chunhua Weng |
AMIA | 1 |
| 2017 | Prediction of human QT prolongation liability based on pre-clinical RNA expression profilesabstractMarked drug-induced prolongation of the QT interval on the electrocardiogram is associated with Torsades de Pointes (TdP), a potentially life-threatening cardiac arrhythmia. Assessment of QT prolongation liability in the drug development process is required but is time and resource intensive. Current pre-clinical safety assessments use patch clamp analysis of the Human Ether-a-Go-Go (hERG) channel, but analyses have broadened to include patch clamp analysis of other ion channels and the use of in silico models. This investigation describes a method for predicting drug-induced QT prolongation liability in humans based on an association with RNA microarray expression profiles from rat liver data, and machine learning implemented in open-source software. Recently reported hERG patch clamp sensitivities and specificities range from between 64-82% and 75-88% respectively. Classification in this study was done using drugs known to prolong the QT interval vs. those that do not, regardless of the drugs' respective indication(s), and then further sub-classified by indication which resulted in 76 sub-groups. Classifier results in this project using 10-fold cross validation had average sensitivities of 85% and specificities of 90% using all available datasets as input, and a mean sensitivity and specificity of 92% and 94%, respectively across 76 drug sub-classifications. While an association between rat liver RNA expression profiles and QT prolongation in human heart tissue does not imply that a specific genetic expression profile is responsible for the QT prolongation, these results suggest that machine learning of gene expression profiles to predict QT liability may be used as a surrogate biomarker as part of the pre-clinical cardiac safety assessment of drugs. Dennis M. Bergau, Cong Liu 0020, Hui Lu 0004 |
BIBM | 2 |
| 2016 | A novel scoring estimator to screening for oncogenic chimeric transcripts in cancer transcriptome sequencingabstractBased on various genomic information of chimeric transcript, recent studies used machine-learning methods to predict the oncogenic potentials for chimeric transcripts, however these works ignored transcriptional signature of those chimeric transcripts. Based on clonal evolution theory, we hypothesized that a chimeric transcript is more likely to be an oncogenic `driver' mutation, if the neoplastic cells harboring this chimeric mutation has larger clonal size than other neoplastic cells in a particular tumor. Here we proposed a novel method, called iFCR (internal Fusion Clone Ratio), to estimate the ratio of subclone carrying chimeric transcripts to the rest of neoplastic cells in transcriptome sequencing data. To evaluate our hypothesis, we applied iFCR method on two public cancer transcriptome sequencing datasets, one for breast cancer cell line and the other for prostate tumors with adjacent normal tissues. Our results demonstrated that the chimeric transcripts in tumor samples appear to have higher iFCR value than normal tissues, the most frequent prostate cancer fusion mutation, TMPRSS2- ERG, has remarkably higher iFCR value in all three independent patients. Our work providing a novel point of view for screening oncogenesis chimeric transcripts in cancer research. Jianlei Gu, Shi-Yi Liu, Cong Liu 0020, Hui Lu 0004 |
BIBM | 4 |