VLDB 2026 Research / reviewers in the wild / expert
Chunhua Weng
dblp:60/4820
· DBLP profile ↗
162ranked-venue papers
16as first author
54since 2021 · last 2026
0000-0002-9624-0214ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 160 · 15 first-author · 54 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CPGPrompt: translating clinical guidelines into large language model-executable decision supportabstractOBJECTIVE: Clinical practice guidelines (CPGs) provide evidence-based recommendations for patient care; however, integrating them into artificial intelligence (AI) remains challenging. Previous approaches, such as rule-based systems or black-box AI models, face significant limitations, including poor interpretability, inconsistent adherence to guidelines, and narrow domain applicability. To address this, we develop and validate CPGPrompt, an auto-prompting system that converts narrative clinical guidelines into large language models (LLMs). MATERIALS AND METHODS: Our framework translates CPGs into structured decision trees and utilizes an LLM to dynamically navigate them for patient case evaluation. Synthetic vignettes were generated across 3 domains-headache, lower back pain, and prostate cancer-and distributed into 4 categories to test different decision scenarios. System performance was assessed on both binary specialty referral decisions and fine-grained pathway classification tasks. RESULTS: The binary specialty referral classification achieved consistently strong performance across all domains (F1: 0.85-1.00), with high recall (1.00 ± 0.00). In contrast, multiclass pathway assignment showed reduced performance, with domain-specific variations: headache (F1: 0.47), lower back pain (F1: 0.72), and prostate cancer (F1: 0.77). DISCUSSION: Domain-specific performance differences reflected the structure of each guideline. The headache guideline highlighted challenges with negation handling. The lower back pain guideline required temporal reasoning. In contrast, prostate cancer pathways benefited from quantifiable laboratory tests, resulting in more reliable decision-making. CONCLUSION: CPGPrompt demonstrates generalizability across diverse clinical domains while maintaining high sensitivity for referral decisions. Its transparent, auditable framework enables the systematic identification of failure modes and provides advantages over black-box AI approaches. However, persistent challenges with subjective clinical assessments indicate a need for targeted improvements and greater clinical robustness. Ruiqi Deng, Geoffrey Martin, Tony Wang, Yi Liu 0059, Chunhua Weng, Yanshan Wang, Justin F. Rousseau, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 6 |
| 2026 | A critical evaluation of generative query expansion on biomedical literature retrievalabstractOBJECTIVE: To evaluate the effectiveness of generative query expansion for biomedical literature retrieval. MATERIALS AND METHODS: We thoroughly examined eight generative query expansion methods using three large language models across five datasets for biomedical literature retrieval. We further performed a quantitative analysis, including performance comparisons, rank transition analysis, and article-type effect analysis. We also conducted a qualitative examination of representative cases, from which we derived an error taxonomy. RESULTS: On BioASQ-Y/N, GPT-4o-based query expansion shifts Recall@10 to 0.417-0.512 and nDCG@10 to 0.358-0.479, relative to a baseline of 0.491 and 0.456. For PubMedQA, Precision@1 ranges from 0.764 to 0.876 and nDCG@10 from 0.847 to 0.931, compared with baseline values of 0.893 and 0.935. For 2019-Trec-PM, query expansion yields Recall@100 of 0.217-0.256 and nDCG@100 of 0.272-0.312, versus a baseline of 0.227 and 0.274. Similarly, for 2018-TREC-PM, Recall@100 spans 0.169-0.227 and nDCG@100 spans 0.195-0.250, relative to baseline scores of 0.164 and 0.191. For 2017-TREC-PM, Recall@100 and nDCG@100 fall within 0.111-0.139 and 0.154-0.191 under query expansion, compared with baseline metrics of 0.102 and 0.147. Both general-purpose and domain-specific Llama-based models demonstrate similar performance to GPT-4o. DISCUSSION AND CONCLUSION: The impact of query expansion varies significantly by the expansion methods and type of evidence, but is relatively agnostic to backbone model choice. Notably, query expansion primarily affects article ranking but has a limited impact on the screening stage. Our findings underscore the unique challenges of biomedical literature retrieval and highlight the need to develop domain-specific information retrieval techniques. Yilu Fang, Fangyi Chen, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 5 |
| 2026 | A data-driven method for research trend analysis in a scientific discipline: Application to the journal of biomedical informaticsabstractOBJECTIVE: Accurately characterizing research trends is critical for identifying cutting-edge scientific breakthroughs in their infancy and informing strategic priorities. This research contributes a pipeline that utilizes generative AI technologies to develop research topic taxonomies from publication keywords and analyze keyword evolution within topics, methodological and domain trends, and topic co-occurrences. We demonstrated the pipeline by conducting a retrospective analysis of biomedical informatics research trends in the Journal of Biomedical Informatics (JBI). METHODS: We identified the JBI publications with keywords available on PubMed, spanning 2011-2025. We downloaded all the keywords and categorized them into methodological innovations and health domains, identified topics, assigned topic names, and constructed their hierarchies, all using large-language models (LLMs). We introduced an automated method for evaluating topics, leveraging MeSH terminology as the underlying knowledge base. RESULTS: Using 6,930 unique keywords from 2,427 publications, we derived 1,028 distinct topics related to methodological innovations, with each topic associated with medians of four keywords (Q1: 2, Q3: 13) and six publications (Q1: 2, Q3: 19). We identified 904 topics related to health domains, with each topic associated with three keywords (Q1: 1, Q3: 11) and four publications (Q1: 1, Q3: 15). Based on the topics, we analyzed the prominent research areas, trends in publication volume, evolution of keyword distributions within each topic, and patterns of co-occurring topics. Among the 2,379 eligible publications, 2,009 (84.4%) exhibited overlap between the keyword-derived MeSH terms and the MeSH terms assigned to the publication by the National Library of Medicine. CONCLUSION: This study presents a method that leverages modern generative AI technologies for retrospective analysis of a scientific field to identify emerging topics and to detect shifts in scholarly focus. Illustrated by data for JBI and correlated with historical background events and policy changes, our findings demonstrate the effectiveness and utility of the methods while providing a powerful lens to understand the evolution of biomedical informatics research priorities in JBI. Yilu Fang, Samir Sanchez Tejada, Fangyi Chen, Edward H. Shortliffe, Vimla L. Patel, Mor Peleg, Chunhua Weng |
J. Biomed. Informatics | 8 |
| 2025 | Semi-supervised learning from small annotated data and large unlabeled data for fine-grained Participants, Intervention, Comparison, and Outcomes entity recognitionabstractOBJECTIVE: Extracting PICO elements-Participants, Intervention, Comparison, and Outcomes-from clinical trial literature is essential for clinical evidence retrieval, appraisal, and synthesis. Existing approaches do not distinguish the attributes of PICO entities. This study aims to develop a named entity recognition (NER) model to extract PICO entities with fine granularities. MATERIALS AND METHODS: Using a corpus of 2511 abstracts with PICO mentions from 4 public datasets, we developed a semi-supervised method to facilitate the training of a NER model, FinePICO, by combining limited annotated data of PICO entities and abundant unlabeled data. For evaluation, we divided the entire dataset into 2 subsets: a smaller group with annotations and a larger group without annotations. We then established the theoretical lower and upper performance bounds based on the performance of supervised learning models trained solely on the small, annotated subset and on the entire set with complete annotations, respectively. Finally, we evaluated FinePICO on both the smaller annotated subset and the larger, initially unannotated subset. We measured the performance of FinePICO using precision, recall, and F1. RESULTS: Our method achieved precision/recall/F1 of 0.567/0.636/0.60, respectively, using a small set of annotated samples, outperforming the baseline model (F1: 0.437) by more than 16%. The model demonstrates generalizability to a different PICO framework and to another corpus, which consistently outperforms the benchmark in diverse experimental settings (P-value < .001). DISCUSSION: We developed FinePICO to recognize fine-grained PICO entities from text and validated its performance across diverse experimental settings, highlighting the feasibility of using semi-supervised learning (SSL) techniques to enhance PICO entities extraction. Future work can focus on optimizing SSL algorithms to improve efficiency and reduce computational costs. CONCLUSION: This study contributes a generalizable and effective semi-supervised approach leveraging large unlabeled data together with small, annotated data for fine-grained PICO extraction. Fangyi Chen, Yilu Fang, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Mini-mental status examination phenotyping for Alzheimer's disease patients using both structured and narrative electronic health record featuresabstractOBJECTIVE: This study aims to automate the prediction of Mini-Mental State Examination (MMSE) scores, a widely adopted standard for cognitive assessment in patients with Alzheimer's disease, using natural language processing (NLP) and machine learning (ML) on structured and unstructured EHR data. MATERIALS AND METHODS: We extracted demographic data, diagnoses, medications, and unstructured clinical visit notes from the EHRs. We used Latent Dirichlet Allocation (LDA) for topic modeling and Term-Frequency Inverse Document Frequency (TF-IDF) for n-grams. In addition, we extracted meta-features such as age, ethnicity, and race. Model training and evaluation employed eXtreme Gradient Boosting (XGBoost), Stochastic Gradient Descent Regressor (SGDRegressor), and Multi-Layer Perceptron (MLP). RESULTS: We analyzed 1654 clinical visit notes collected between September 2019 and June 2023 for 1000 Alzheimer's disease patients. The average MMSE score was 20, with patients averaging 76.4 years old, 54.7% female, and 54.7% identifying as White. The best-performing model (ie, lowest root mean squared error (RMSE)) is MLP, which achieved an RMSE of 5.53 on the validation set using n-grams, indicating superior prediction performance over other models and feature sets. The RMSE on the test set was 5.85. DISCUSSION: This study developed a ML method to predict MMSE scores from unstructured clinical notes, demonstrating the feasibility of utilizing NLP to support cognitive assessment. Future work should focus on refining the model and evaluating its clinical relevance across diverse settings. CONCLUSION: We contributed a model for automating MMSE estimation using EHR features, potentially transforming cognitive assessment for Alzheimer's patients and paving the way for more informed clinical decisions and cohort identification. Betina Ross S. Idnay, Fangyi Chen, Casey N. Ta, Matthew W. Schelke, Karen Marder, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Incorporating preprints in systematic reviews: a preliminary study of a novel method for rapid evidence synthesisabstractOBJECTIVES: By October 1, 2024, over 450,000 COVID-19 manuscripts were published, with 10% posted as unreviewed preprints. While they accelerate knowledge sharing, their inconsistent quality complicates systematic studies. MATERIALS AND METHODS: We propose a 2-stage method to include preprints in meta-analyses. In Stage A, preprints are integrated through restriction or imputation and weighted by a confidence score reflecting their publication likelihood. In Stage B, we assess and adjust for potential publication or reporting biases. RESULTS: This preliminary study employed a 2-stage procedure validated with 2 COVID-19 treatment case studies. For hydroxychloroquine, the relative risk (RR) was 1.06 [95% CI: 0.62, 1.80], suggesting no mortality benefit over placebo. For corticosteroids, the RR was 0.88 [95% CI: 0.62, 1.27], which, while not statistically significant, aligns with evidence supporting a mortality benefit. DISCUSSION: Our research aims to bridge a significant methodological gap by providing a solution for timely evidence synthesis, particularly in the face of the overwhelming number of publications surrounding COVID-19. CONCLUSION: This preliminary study presents a method to efficiently synthesize COVID-19 research, including non-peer-reviewed preprints, to support clinical and policy decisions amidst the information surge. Jiayi Tong, Yifei Sun 0007, Rebecca A. Hubbard, M. Elle Saine, Hua Xu 0001, Xu Zuo, Chunhua Weng, Christopher H. Schmid, Stephen E. Kimmel, Craig A. Umscheid, Adam Cuker, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 8 |
| 2025 | A method for characterizing disease progression from acute kidney injury to chronic kidney disease
Yilu Fang, Jordan G. Nestor, Casey N. Ta, Jerard Kneifati-Hayek, Chunhua Weng |
J. Biomed. Informatics | 5 |
| 2025 | CLEAR: A vision to support clinical evidence lifecycle with continuous learningabstractHuman knowledge of diseases, treatments, and prevention techniques is constantly evolving. The generation of clinical evidence using randomized controlled trials on human subjects occurs notably slowly and inefficiently. The Learning Health System (LHS) has been proposed to facilitate the continuous improvement of individual and population health through a cycle of knowledge, practice, and data. However, the gap between the demand for high-quality evidence to support clinical decisions and the available evidence continues to enlarge. While the current LHS vision articulates the integration of Real-World Data (RWD), the rapid generation of RWD often outpaces the rate of effective evidence synthesis and implementation. Considering this, we propose a new framework that more effectively leverages RWD to support the entire clinical evidence lifecycle through a continuous learning mechanism. This framework, powered by modern data science and informatics, offers enhanced scalability and efficiency. In this vision, specifically, RWD is integrated into the clinical evidence lifecycle via four closed feedback loops: 1) guiding research prioritization and study design, 2) facilitating clinical guideline development, 3) assisting guideline evaluation, and 4) supporting shared decision-making. Our framework enables rapid responsiveness to emerging health data and evolving healthcare needs, timely development of clinical guidelines to optimize clinical recommendations, and sustained improvements in clinical practice and patient outcomes. This vision calls for informatics support for an efficient, scalable, and stakeholder-aware clinical evidence lifecycle. Yilu Fang, Fangyi Chen, George Hripcsak, Yifan Peng 0002, Patrick B. Ryan, Chunhua Weng |
J. Biomed. Informatics | 7 |
| 2025 | Scalable scientific interest profiling using large language models
Yilun Liang, Edward Sun, Betina Ross S. Idnay, Yilu Fang, Fangyi Chen, Casey N. Ta, Yifan Peng 0002, Chunhua Weng |
J. Biomed. Informatics | 9 |
| 2025 | Machine learning applications related to suicide in military and Veterans: A scoping literature review
Yishu Wei, Yanshan Wang, Yunyu Xiao, Ronald K. Poropatich, Gretchen L. Haas, Yiye Zhang, Chunhua Weng, Jinze Liu, Lisa A. Brenner, James M. Bjork, Yifan Peng 0002 |
J. Biomed. Informatics | 8 |
| 2024 | Participant-guided development of bilingual genomic educational infographics for Electronic Medical Records and Genomics Phase IV studyabstractOBJECTIVE: Developing targeted, culturally competent educational materials is critical for participant understanding of engagement in a large genomic study that uses computational pipelines to produce genome-informed risk assessments. MATERIALS AND METHODS: Guided by the Smerecnik framework that theorizes understanding of multifactorial genetic disease through 3 knowledge types, we developed English and Spanish infographics for individuals enrolled in the Electronic Medical Records and Genomics Network. Infographics were developed to explain concepts in lay language and visualizations. We conducted iterative sessions using a modified "think-aloud" process with 10 participants (6 English, 4 Spanish-speaking) to explore comprehension of and attitudes towards the infographics. RESULTS: We found that all but one participant had "awareness knowledge" of genetic disease risk factors upon viewing the infographics. Many participants had difficulty with "how-to" knowledge of applying genetic risk factors to specific monogenic and polygenic risks. Participant attitudes towards the iteratively-refined infographics indicated that design saturation was reached. DISCUSSION: There were several elements that contributed to the participants' comprehension (or misunderstanding) of the infographics. Visualization and iconography techniques best resonated with those who could draw on prior experiences or knowledge and were absent in those without. Limited graphicacy interfered with the understanding of absolute and relative risks when presented in graph format. Notably, narrative and storytelling theory that informed the creation of a vignette infographic was most accessible to all participants. CONCLUSION: Engagement with the intended audience who can identify strengths and points for improvement of the intervention is necessary to the development of effective infographics. Aimiel Casillan, Michelle E. Florido, Jamie Galarza-Cornejo, Suzanne Bakken, John A. Lynch, Wendy K. Chung, Kathleen F. Mittendorf, Eta S. Berner, John J. Connolly, Chunhua Weng, Ingrid A. Holm, Atlas Khan, Krzysztof Kiryluk, Nita A. Limdi, Lynn Petukhova, Maya Sabatello, Julia Wynn |
J. Am. Medical Informatics Assoc. | 10 |
| 2024 | Knowledge-guided generative artificial intelligence for automated taxonomy learning from drug labelsabstractOBJECTIVES: To automatically construct a drug indication taxonomy from drug labels using generative Artificial Intelligence (AI) represented by the Large Language Model (LLM) GPT-4 and real-world evidence (RWE). MATERIALS AND METHODS: We extracted indication terms from 46 421 free-text drug labels using GPT-4, iteratively and recursively generated indication concepts and inferred indication concept-to-concept and concept-to-term subsumption relations by integrating GPT-4 with RWE, and created a drug indication taxonomy. Quantitative and qualitative evaluations involving domain experts were performed for cardiovascular (CVD), Endocrine, and Genitourinary system diseases. RESULTS: 2909 drug indication terms were extracted and assigned into 24 high-level indication categories (ie, initially generated concepts), each of which was expanded into a sub-taxonomy. For example, the CVD sub-taxonomy contains 242 concepts, spanning a depth of 11, with 170 being leaf nodes. It collectively covers a total of 234 indication terms associated with 189 distinct drugs. The accuracies of GPT-4 on determining the drug indication hierarchy exceeded 0.7 with "good to very good" inter-rater reliability. However, the accuracies of the concept-to-term subsumption relation checking varied greatly, with "fair to moderate" reliability. DISCUSSION AND CONCLUSION: We successfully used generative AI and RWE to create a taxonomy, with drug indications adequately consistent with domain expert expectations. We show that LLMs are good at deriving their own concept hierarchies but still fall short in determining the subsumption relations between concepts and terms in unregulated language from free-text drug labels, which is the same hard task for human experts. Yilu Fang, Patrick Ryan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Sociotechnical feasibility of natural language processing-driven tools in clinical trial eligibility prescreening for Alzheimer's disease and related dementiasabstractBACKGROUND: Alzheimer's disease and related dementias (ADRD) affect over 55 million globally. Current clinical trials suffer from low recruitment rates, a challenge potentially addressable via natural language processing (NLP) technologies for researchers to effectively identify eligible clinical trial participants. OBJECTIVE: This study investigates the sociotechnical feasibility of NLP-driven tools for ADRD research prescreening and analyzes the tools' cognitive complexity's effect on usability to identify cognitive support strategies. METHODS: A randomized experiment was conducted with 60 clinical research staff using three prescreening tools (Criteria2Query, Informatics for Integrating Biology and the Bedside [i2b2], and Leaf). Cognitive task analysis was employed to analyze the usability of each tool using the Health Information Technology Usability Evaluation Scale. Data analysis involved calculating descriptive statistics, interrater agreement via intraclass correlation coefficient, cognitive complexity, and Generalized Estimating Equations models. RESULTS: Leaf scored highest for usability followed by Criteria2Query and i2b2. Cognitive complexity was found to be affected by age, computer literacy, and number of criteria, but was not significantly associated with usability. DISCUSSION: Adopting NLP for ADRD prescreening demands careful task delegation, comprehensive training, precise translation of eligibility criteria, and increased research accessibility. The study highlights the relevance of these factors in enhancing NLP-driven tools' usability and efficacy in clinical research prescreening. CONCLUSION: User-modifiable NLP-driven prescreening tools were favorably received, with system type, evaluation sequence, and user's computer literacy influencing usability more than cognitive complexity. The study emphasizes NLP's potential in improving recruitment for clinical trials, endorsing a mixed-methods approach for future system evaluation and enhancements. Betina Ross S. Idnay, Jianfang Liu, Yilu Fang, Alex Hernandez, Shivani Kaw, Alicia Etwaru, Janeth Juarez Padilla, Sergio Ozoria Ramirez, Karen Marder, Chunhua Weng, Rebecca Schnall |
J. Am. Medical Informatics Assoc. | 10 |
| 2024 | Large language models in biomedicine and health: current research landscape and future directionsabstractLarge language models in biomedicine and health: current research landscape and future directionsLarge language models (LLMs) are a specialized type of generative artificial intelligence (AI) focused on generating natural language text.These models are developed through extensive training on massive amounts of text data and use deep learning algorithms to generate new text that closely resembles human-generated text.Generative AI methods, including LLMs, are rapidly transforming various domains, including biomedicine and healthcare.[1][2][3][4][5][6] They have already demonstrated remarkable potential as a means to process and analyze large amounts of text, interpret natural language, and generate new content in these domains.For example, Nori et al reported that GPT-4 is able to correctly answer the majority of questions from medical practice licensing exams, comfortably obtaining a passing grade.7 Similarly, Stribling et al found that this model exceeded the average performance of students in the graduate medical sciences on the majority of examinations, including strong performance on short answer and essay questions.8 Even though passing the exam is not the same as applying the knowledge in a real-world setting, these results demonstrate that LLMs can generate appropriate multiple-choice and narrative responses to questions framed in natural language.ChatGPT, first released in November 2022, has garnered phenomenal attention from both the scientific community and a broader society.A keyword search of "large language models" OR "ChatGPT" in PubMed returned over 4500 articles that discuss the technology and its implications for various topics, including medical informatics, by the end of June 2024.In addition, LLM-based technologies have already been deployed in several healthcare systems and are offered as integrated products for use in the clinic within vendor electronic health record systems (for thoughts on initial evaluations of an early product, see Garcia et al 9 and Tai-Seale et al 10 ).This rapid adoption of LLMs like ChatGPT brings an unprecedented opportunity to use this novel AI technology to transform healthcare and medicine.Despite their potential benefits, LLMs can sometimes produce invalid and unsubstantiated responses, a phenomenon known as the "hallucination and confabulation issue" in the literature, or biased responses, due to the biases inherent in their training data.[11][12][13][14][15][16][17] With this great potential also comes the need for trustworthy and responsible development and use of technology.As we continue to explore the capabilities of ChatGPT and other LLMs, it is critical to address related ethical, legal, and social issues to ensure that the technology is used in ways that are safe, fair, trustworthy, and beneficial for all.In the context of biomedicine and healthcare, it is particularly important to engage stakeholders, such as AI researchers, developers of data-driven clinical decision support, care providers, and system implementers from both academic medical centers and industry, to ensure responsible use of LLMs for good.To accelerate research and development in this area, we issued a call for submissions in Summer 2023, specifically focusing on the intersection of biomedicine/health and LLMs, and invited contributions on all related aspects.We invited submissions that report on innovative informatics methods development and evaluation, as well as studies that demonstrate the effectiveness/limitations of LLMs methodologies in healthcare.We particularly encouraged submissions that address the challenges and opportunities of this intersection and offer new insights into how these fields can work together to advance healthcare.This editorial provides an overview of the papers accepted in this Focus Issue.We highlight major themes and unique aspects of the research papers in medical LLMs, discuss ongoing challenges, and recommend future research directions.Box 1 lists the relevant large language model terms and abbreviations used in this editorial. Overall statistics of the Focus IssueThis JAMIA Focus Issue on LLMs in biomedicine and health has drawn enthusiasm from many researchers across different research disciplines.In total, we received over 150 submissions from authors in 25 countries and regions across 6 continents worldwide.The rigorous JAMIA peer review process was applied to all submissions, 41 of which were ultimately accepted for publication in the Focus Issue (Table 1).The majority of the accepted papers were authored by authors in North America, followed by those in Asia and Europe (Figure 1A).The Focus Issue highlights the nature of multi-disciplinary collaboration in medical informatics research across the broad JAMIA community.The number of authors per paper varies from 1 to 23, with an average of 7.3.Many papers feature authors with diverse expertise from different departments and organizations.The authors' expertise spans a wide range of fields, including computer science, data science, informatics, statistics, medicine, nursing, clinical services, public health policies, and more.Several papers also demonstrate scientific collaborations across different sectors, including academia, government labs, research institutes, hospitals, and industry.Additionally, a few papers showcase international collaborations among authors. Zhiyong Lu, Yifan Peng 0002, Trevor Cohen, Marzyeh Ghassemi, Chunhua Weng, Shubo Tian |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | Fine-tuning large language models for rare disease concept normalizationabstractOBJECTIVE: We aim to develop a novel method for rare disease concept normalization by fine-tuning Llama 2, an open-source large language model (LLM), using a domain-specific corpus sourced from the Human Phenotype Ontology (HPO). METHODS: We developed an in-house template-based script to generate two corpora for fine-tuning. The first (NAME) contains standardized HPO names, sourced from the HPO vocabularies, along with their corresponding identifiers. The second (NAME+SYN) includes HPO names and half of the concept's synonyms as well as identifiers. Subsequently, we fine-tuned Llama 2 (Llama2-7B) for each sentence set and conducted an evaluation using a range of sentence prompts and various phenotype terms. RESULTS: When the phenotype terms for normalization were included in the fine-tuning corpora, both models demonstrated nearly perfect performance, averaging over 99% accuracy. In comparison, ChatGPT-3.5 has only ∼20% accuracy in identifying HPO IDs for phenotype terms. When single-character typos were introduced in the phenotype terms, the accuracy of NAME and NAME+SYN is 10.2% and 36.1%, respectively, but increases to 61.8% (NAME+SYN) with additional typo-specific fine-tuning. For terms sourced from HPO vocabularies as unseen synonyms, the NAME model achieved 11.2% accuracy, while the NAME+SYN model achieved 92.7% accuracy. CONCLUSION: Our fine-tuned models demonstrate ability to normalize phenotype terms unseen in the fine-tuning corpus, including misspellings, synonyms, terms from other ontologies, and laymen's terms. Our approach provides a solution for the use of LLMs to identify named medical entities from clinical narratives, while successfully normalizing them to standard concepts in a controlled vocabulary. Cong Liu 0020, Jingye Yang, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | A span-based model for extracting overlapping PICO entities from randomized controlled trial publicationsabstractOBJECTIVES: Extracting PICO (Populations, Interventions, Comparison, and Outcomes) entities is fundamental to evidence retrieval. We present a novel method, PICOX, to extract overlapping PICO entities. MATERIALS AND METHODS: PICOX first identifies entities by assessing whether a word marks the beginning or conclusion of an entity. Then, it uses a multi-label classifier to assign one or more PICO labels to a span candidate. PICOX was evaluated using 1 of the best-performing baselines, EBM-NLP, and 3 more datasets, ie, PICO-Corpus and randomized controlled trial publications on Alzheimer's Disease (AD) or COVID-19, using entity-level precision, recall, and F1 scores. RESULTS: PICOX achieved superior precision, recall, and F1 scores across the board, with the micro F1 score improving from 45.05 to 50.87 (P ≪.01). On the PICO-Corpus, PICOX obtained higher recall and F1 scores than the baseline and improved the micro recall score from 56.66 to 67.33. On the COVID-19 dataset, PICOX also outperformed the baseline and improved the micro F1 score from 77.10 to 80.32. On the AD dataset, PICOX demonstrated comparable F1 scores with higher precision when compared to the baseline. CONCLUSION: PICOX excels in identifying overlapping entities and consistently surpasses a leading baseline across multiple datasets. Ablation studies reveal that its data augmentation strategy effectively minimizes false positives and improves precision. Yiliang Zhou, Hua Xu 0001, Chunhua Weng, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | Promoting equity in clinical research: The role of social determinants of health
Betina Ross S. Idnay, Yilu Fang, Edward Stanley, Brenda Ruotolo, Wendy K. Chung, Karen Marder, Chunhua Weng |
J. Biomed. Informatics | 7 |
| 2024 | Criteria2Query 3.0: Leveraging generative large language models for clinical trial eligibility query generation
Jimyung Park, Yilu Fang, Casey N. Ta, Betina Ross S. Idnay, Fangyi Chen, Rebecca Shyu, Emily R. Gordon, Matthew E. Spotnitz, Chunhua Weng |
J. Biomed. Informatics | 11 |
| 2024 | Converting OMOP CDM to phenopackets: A model alignment and patient data representation evaluation
Kayla Schiffer, Cong Liu 0020, Tiffany Callahan, Casey N. Ta, Jordan G. Nestor, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2024 | Rare disease diagnosis using knowledge guided retrieval augmentation for ChatGPT
Charlotte Zelin, Wendy K. Chung, Mederic Jeanne, Chunhua Weng |
J. Biomed. Informatics | 5 |
| 2024 | Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness
Qiao Jin 0001, Denis Jered McInerney, Yong Chen 0016, Fei Wang 0001, Curtis L. Cole, Qian Yang 0004, Yanshan Wang, Bradley A. Malin, Mor Peleg, Byron C. Wallace, Zhiyong Lu, Chunhua Weng, Yifan Peng 0002 |
J. Biomed. Informatics | 13 |
| 2023 | Characterizing variability of electronic health record-driven phenotype definitionsabstractOBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic. Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 10 |
| 2023 | EvidenceMap: a three-level knowledge representation for medical evidence computation and comprehensionabstractOBJECTIVE: To develop a computable representation for medical evidence and to contribute a gold standard dataset of annotated randomized controlled trial (RCT) abstracts, along with a natural language processing (NLP) pipeline for transforming free-text RCT evidence in PubMed into the structured representation. MATERIALS AND METHODS: Our representation, EvidenceMap, consists of 3 levels of abstraction: Medical Evidence Entity, Proposition and Map, to represent the hierarchical structure of medical evidence composition. Randomly selected RCT abstracts were annotated following EvidenceMap based on the consensus of 2 independent annotators to train an NLP pipeline. Via a user study, we measured how the EvidenceMap improved evidence comprehension and analyzed its representative capacity by comparing the evidence annotation with EvidenceMap representation and without following any specific guidelines. RESULTS: Two corpora including 229 disease-agnostic and 80 COVID-19 RCT abstracts were annotated, yielding 12 725 entities and 1602 propositions. EvidenceMap saves users 51.9% of the time compared to reading raw-text abstracts. Most evidence elements identified during the freeform annotation were successfully represented by EvidenceMap, and users gave the enrollment, study design, and study Results sections mean 5-scale Likert ratings of 4.85, 4.70, and 4.20, respectively. The end-to-end evaluations of the pipeline show that the evidence proposition formulation achieves F1 scores of 0.84 and 0.86 in the adjusted random index score. CONCLUSIONS: EvidenceMap extends the participant, intervention, comparator, and outcome framework into 3 levels of abstraction for transforming free-text evidence from the clinical literature into a computable structure. It can be used as an interoperable format for better evidence retrieval and synthesis and an interpretable representation to efficiently comprehend RCT findings. Tian Kang, Yingcheng Sun, Jae Hyun Kim, Casey N. Ta, Adler J. Perotte, Kayla Schiffer, Mutong Wu, Nour Fahmy, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | The suitability of UMLS and SNOMED-CT for encoding outcome conceptsabstractOBJECTIVE: Outcomes are important clinical study information. Despite progress in automated extraction of PICO (Population, Intervention, Comparison, and Outcome) entities from PubMed, rarely are these entities encoded by standard terminology to achieve semantic interoperability. This study aims to evaluate the suitability of the Unified Medical Language System (UMLS) and SNOMED-CT in encoding outcome concepts in randomized controlled trial (RCT) abstracts. MATERIALS AND METHODS: We iteratively developed and validated an outcome annotation guideline and manually annotated clinically significant outcome entities in the Results and Conclusions sections of 500 randomly selected RCT abstracts on PubMed. The extracted outcomes were fully, partially, or not mapped to the UMLS via MetaMap based on established heuristics. Manual UMLS browser search was performed for select unmapped outcome entities to further differentiate between UMLS and MetaMap errors. RESULTS: Only 44% of 2617 outcome concepts were fully covered in the UMLS, among which 67% were complex concepts that required the combination of 2 or more UMLS concepts to represent them. SNOMED-CT was present as a source in 61% of the fully mapped outcomes. DISCUSSION: Domains such as Metabolism and Nutrition, and Infections and Infectious Diseases need expanded outcome concept coverage in the UMLS and MetaMap. Future work is warranted to similarly assess the terminology coverage for P, I, C entities. CONCLUSION: Computational representation of clinical outcomes is important for clinical evidence extraction and appraisal and yet faces challenges from the inherent complexity and lack of coverage of these concepts in UMLS and SNOMED-CT, as demonstrated in this study. Abigail M. Newbury, Hao Liu 0054, Betina Ross S. Idnay, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Clinical and temporal characterization of COVID-19 subgroups using patient vector embeddings of electronic health recordsabstractOBJECTIVE: To identify and characterize clinical subgroups of hospitalized Coronavirus Disease 2019 (COVID-19) patients. MATERIALS AND METHODS: Electronic health records of hospitalized COVID-19 patients at NewYork-Presbyterian/Columbia University Irving Medical Center were temporally sequenced and transformed into patient vector representations using Paragraph Vector models. K-means clustering was performed to identify subgroups. RESULTS: A diverse cohort of 11 313 patients with COVID-19 and hospitalizations between March 2, 2020 and December 1, 2021 were identified; median [IQR] age: 61.2 [40.3-74.3]; 51.5% female. Twenty subgroups of hospitalized COVID-19 patients, labeled by increasing severity, were characterized by their demographics, conditions, outcomes, and severity (mild-moderate/severe/critical). Subgroup temporal patterns were characterized by the durations in each subgroup, transitions between subgroups, and the complete paths throughout the course of hospitalization. DISCUSSION: Several subgroups had mild-moderate severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections but were hospitalized for underlying conditions (pregnancy, cardiovascular disease [CVD], etc.). Subgroup 7 included solid organ transplant recipients who mostly developed mild-moderate or severe disease. Subgroup 9 had a history of type-2 diabetes, kidney and CVD, and suffered the highest rates of heart failure (45.2%) and end-stage renal disease (80.6%). Subgroup 13 was the oldest (median: 82.7 years) and had mixed severity but high mortality (33.3%). Subgroup 17 had critical disease and the highest mortality (64.6%), with age (median: 68.1 years) being the only notable risk factor. Subgroups 18-20 had critical disease with high complication rates and long hospitalizations (median: 40+ days). All subgroups are detailed in the full text. A chord diagram depicts the most common transitions, and paths with the highest prevalence, longest hospitalizations, lowest and highest mortalities are presented. Understanding these subgroups and their pathways may aid clinicians in their decisions for better management and earlier intervention for patients. Casey N. Ta, Jason Zucker 0001, Po-Hsiang Chiu, Yilu Fang, Karthik Natarajan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | A data-driven approach to optimizing clinical study eligibility criteria
Yilu Fang, Hao Liu 0054, Betina Ross S. Idnay, Casey N. Ta, Karen Marder, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2022 | Optimizing Clinical Research Eligibility Prescreening: An Iterative Usability Evaluation of an NLP-driven Cohort Identification Tool
Betina Ross S. Idnay, Yilu Fang, Caitlin N. Dreisbach, Karen Marder, Chunhua Weng, Rebecca Schnall |
AMIA | 5 |
| 2022 | Criteria2Query 2.0: Combining Machine Efficiency and Human Intelligence to Define a More Accurate and Feasible Cohort for Clinical Trial Recruitment
Betina Ross S. Idnay, Yilu Fang, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Rebecca Schnall, Chunhua Weng |
AMIA | 7 |
| 2022 | Leveraging Hierarchical Concept Relations to Improve Automatic Detection of Off-Label Drug Use in Electronic Health Records Data
Kayla Schiffer, Chunhua Weng, Yoolim A. Choi |
AMIA | 2 |
| 2022 | A Clinical Data Summarization Pipeline for Patients at Risk for Esophageal Adenocarcinoma Undergoing Endoscopic Surveillance
Ali Soroush, Courtney J. Diamond, Haley Zylberberg, Julian Abrams, Chunhua Weng |
AMIA | 5 |
| 2022 | An interactive fitness-for-use data completeness tool to assess activity tracker dataabstractOBJECTIVE: To design and evaluate an interactive data quality (DQ) characterization tool focused on fitness-for-use completeness measures to support researchers' assessment of a dataset. MATERIALS AND METHODS: Design requirements were identified through a conceptual framework on DQ, literature review, and interviews. The prototype of the tool was developed based on the requirements gathered and was further refined by domain experts. The Fitness-for-Use Tool was evaluated through a within-subjects controlled experiment comparing it with a baseline tool that provides information on missing data based on intrinsic DQ measures. The tools were evaluated on task performance and perceived usability. RESULTS: The Fitness-for-Use Tool allows users to define data completeness by customizing the measures and its thresholds to fit their research task and provides a data summary based on the customized definition. Using the Fitness-for-Use Tool, study participants were able to accurately complete fitness-for-use assessment in less time than when using the Intrinsic DQ Tool. The study participants perceived that the Fitness-for-Use Tool was more useful in determining the fitness-for-use of a dataset than the Intrinsic DQ Tool. DISCUSSION: Incorporating fitness-for-use measures in a DQ characterization tool could provide data summary that meets researchers needs. The design features identified in this study has potential to be applied to other biomedical data types. CONCLUSION: A tool that summarizes a dataset in terms of fitness-for-use dimensions and measures specific to a research question supports dataset assessment better than a tool that only presents information on intrinsic DQ measures. Sylvia Cho, Ipek Ensari, Noémie Elhadad, Chunhua Weng, Jennifer M. Radin, Brinnae Bent, Pooja M. Desai, Karthik Natarajan |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Combining human and machine intelligence for clinical trial eligibility queryingabstractOBJECTIVE: To combine machine efficiency and human intelligence for converting complex clinical trial eligibility criteria text into cohort queries. MATERIALS AND METHODS: Criteria2Query (C2Q) 2.0 was developed to enable real-time user intervention for criteria selection and simplification, parsing error correction, and concept mapping. The accuracy, precision, recall, and F1 score of enhanced modules for negation scope detection, temporal and value normalization were evaluated using a previously curated gold standard, the annotated eligibility criteria of 1010 COVID-19 clinical trials. The usability and usefulness were evaluated by 10 research coordinators in a task-oriented usability evaluation using 5 Alzheimer's disease trials. Data were collected by user interaction logging, a demographic questionnaire, the Health Information Technology Usability Evaluation Scale (Health-ITUES), and a feature-specific questionnaire. RESULTS: The accuracies of negation scope detection, temporal and value normalization were 0.924, 0.916, and 0.966, respectively. C2Q 2.0 achieved a moderate usability score (3.84 out of 5) and a high learnability score (4.54 out of 5). On average, 9.9 modifications were made for a clinical study. Experienced researchers made more modifications than novice researchers. The most frequent modification was deletion (5.35 per study). Furthermore, the evaluators favored cohort queries resulting from modifications (score 4.1 out of 5) and the user engagement features (score 4.3 out of 5). DISCUSSION AND CONCLUSION: Features to engage domain experts and to overcome the limitations in automated machine output are shown to be useful and user-friendly. We concluded that human-computer collaboration is key to improving the adoption and user-friendliness of natural language processing. Yilu Fang, Betina Ross S. Idnay, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Karen Marder, Hua Xu 0001, Rebecca Schnall, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 9 |
| 2022 | Assessing socioeconomic bias in machine learning algorithms in health care: a case study of the HOUSES indexabstractOBJECTIVE: Artificial intelligence (AI) models may propagate harmful biases in performance and hence negatively affect the underserved. We aimed to assess the degree to which data quality of electronic health records (EHRs) affected by inequities related to low socioeconomic status (SES), results in differential performance of AI models across SES. MATERIALS AND METHODS: This study utilized existing machine learning models for predicting asthma exacerbation in children with asthma. We compared balanced error rate (BER) against different SES levels measured by HOUsing-based SocioEconomic Status measure (HOUSES) index. As a possible mechanism for differential performance, we also compared incompleteness of EHR information relevant to asthma care by SES. RESULTS: Asthmatic children with lower SES had larger BER than those with higher SES (eg, ratio = 1.35 for HOUSES Q1 vs Q2-Q4) and had a higher proportion of missing information relevant to asthma care (eg, 41% vs 24% for missing asthma severity and 12% vs 9.8% for undiagnosed asthma despite meeting asthma criteria). DISCUSSION: Our study suggests that lower SES is associated with worse predictive model performance. It also highlights the potential role of incomplete EHR data in this differential performance and suggests a way to mitigate this bias. CONCLUSION: The HOUSES index allows AI researchers to assess bias in predictive model performance by SES. Although our case study was based on a small sample size and a single-site study, the study results highlight a potential strategy for identifying bias by using an innovative SES measure. Young J. Juhn, Euijung Ryu, Chung-Il Wi, Katherine S. King, Momin M. Malik, Santiago Romero-Brufau, Chunhua Weng, Sunghwan Sohn, Richard R. Sharp, John D. Halamka |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | EHR-based cohort assessment for multicenter RCTs: a fast and flexible model for identifying potential study sitesabstractOBJECTIVE: The Recruitment Innovation Center (RIC), partnering with the Trial Innovation Network and institutions in the National Institutes of Health-sponsored Clinical and Translational Science Awards (CTSA) Program, aimed to develop a service line to retrieve study population estimates from electronic health record (EHR) systems for use in selecting enrollment sites for multicenter clinical trials. Our goal was to create and field-test a low burden, low tech, and high-yield method. MATERIALS AND METHODS: In building this service line, the RIC strove to complement, rather than replace, CTSA hubs' existing cohort assessment tools. For each new EHR cohort request, we work with the investigator to develop a computable phenotype algorithm that targets the desired population. CTSA hubs run the phenotype query and return results using a standardized survey. We provide a comprehensive report to the investigator to assist in study site selection. RESULTS: From 2017 to 2020, the RIC developed and socialized 36 phenotype-dependent cohort requests on behalf of investigators. The average response rate to these requests was 73%. DISCUSSION: Achieving enrollment goals in a multicenter clinical trial requires that researchers identify study sites that will provide sufficient enrollment. The fast and flexible method the RIC has developed, with CTSA feedback, allows hubs to query their EHR using a generalizable, vetted phenotype algorithm to produce reliable counts of potentially eligible study participants. CONCLUSION: The RIC's EHR cohort assessment process for evaluating sites for multicenter trials has been shown to be efficient and helpful. The model may be replicated for use by other programs. Sarah J. Nelson, Bethany Drury, Daniel Hood, Jeremy Harper, Tiffany Bernard, Chunhua Weng, Nan Kennedy, Bernard LaSalle, Ramkiran Gouripeddi, Consuelo H. Wilkins, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | A research agenda to support the development and implementation of genomics-based clinical informatics tools and resourcesabstractOBJECTIVE: The Genomic Medicine Working Group of the National Advisory Council for Human Genome Research virtually hosted its 13th genomic medicine meeting titled "Developing a Clinical Genomic Informatics Research Agenda". The meeting's goal was to articulate a research strategy to develop Genomics-based Clinical Informatics Tools and Resources (GCIT) to improve the detection, treatment, and reporting of genetic disorders in clinical settings. MATERIALS AND METHODS: Experts from government agencies, the private sector, and academia in genomic medicine and clinical informatics were invited to address the meeting's goals. Invitees were also asked to complete a survey to assess important considerations needed to develop a genomic-based clinical informatics research strategy. RESULTS: Outcomes from the meeting included identifying short-term research needs, such as designing and implementing standards-based interfaces between laboratory information systems and electronic health records, as well as long-term projects, such as identifying and addressing barriers related to the establishment and implementation of genomic data exchange systems that, in turn, the research community could help address. DISCUSSION: Discussions centered on identifying gaps and barriers that impede the use of GCIT in genomic medicine. Emergent themes from the meeting included developing an implementation science framework, defining a value proposition for all stakeholders, fostering engagement with patients and partners to develop applications under patient control, promoting the use of relevant clinical workflows in research, and lowering related barriers to regulatory processes. Another key theme was recognizing pervasive biases in data and information systems, algorithms, access, value, and knowledge repositories and identifying ways to resolve them. Ken Wiley, Laura Findley, Madison Goldrich, Teji Rakhra-Burris, Ana Stevens, Pamela Williams, Carol J. Bult, Rex L. Chisholm, Patricia Deverka, Geoffrey S. Ginsburg, Eric D. Green, Gail P. Jarvik, George A. Mensah, Erin Ramos, Mary Relling, Dan M. Roden, Robb Rowley, Gil Alterovitz, Samuel J. Aronson, Lisa Bastarache, James J. Cimino, Erin L. Crowgey, Guilherme Del Fiol, Robert R. Freimuth, Mark A. Hoffman, Janina M. Jeff, Kevin B. Johnson, Kensaku Kawamoto, Subha Madhavan, Eneida A. Mendonça, Lucila Ohno-Machado, Siddharth Pratap, Casey Overby Taylor, Marylyn D. Ritchie, Nephi Walton, Chunhua Weng, Teresa Zayas-Cabán, Teri A. Manolio, Marc S. Williams |
J. Am. Medical Informatics Assoc. | 36 |
| 2022 | Deep learning for rare disease: A scoping review
Cong Liu 0020, Zhehuan Chen, Yingcheng Sun, James R. Rogers, Wendy K. Chung, Chunhua Weng |
J. Biomed. Informatics | 8 |
| 2022 | Ontology-based categorization of clinical studies by their conditions
Hao Liu 0054, Simona Carini, Zhehuan Chen, Spencer Phillips Hey, Ida Sim, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2022 | Leveraging electronic health record data for clinical trial planning by assessing eligibility criteria's impact on patient count and safety
James R. Rogers, Jovana Pavisic, Casey N. Ta, Cong Liu 0020, Ali Soroush, Ying Kuen Cheung, George Hripcsak, Chunhua Weng |
J. Biomed. Informatics | 8 |
| 2021 | Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei |
AMIA | 24 |
| 2021 | Cognitive Function Characterization Using Electronic Health Records Notes
Adrienne Pichon, Betina Ross S. Idnay, Rebecca Schnall, Karen Marder, Chunhua Weng |
AMIA | 5 |
| 2021 | Harmonization of Measurement Codes for Concept-Oriented Lab Data Retrieval
Matthew E. Spotnitz, Jason Patterson, Vojtech Huser, Chunhua Weng, Karthik Natarajan |
AMIA | 4 |
| 2021 | Misalignment between COVID-19 hotspots and clinical trial sitesabstractHundreds of interventional clinical trials have been launched in the United States to identify effective treatment strategies for combating the coronavirus disease 2019 (COVID-19) pandemic. However, to date, only a small fraction of these trials have completed enrollment, delaying the scientific investigation of COVID-19 and its treatment options. This study presents novel metrics to examine the geographic alignment between COVID-19 hotspots and interventional clinical trial sites and evaluate trial access over time during the evolving pandemic. Using temporal COVID-19 case data from USAFacts.org and trial data from ClinicalTrials.gov, U.S. counties were categorized based on their numbers of cases and trials. Our analysis suggests that alignment and access have worsened as the pandemic shifted over time. We recommend strategies and metrics to evaluate the alignment between cases and trials. Future studies are warranted to investigate the impact of the misalignment of cases and clinical trial sites on clinical trial recruitment. Lauren Franks, Hao Liu 0054, Mitchell S. V. Elkind, Muredach P. Reilly, Chunhua Weng, Shing M. Lee |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | A systematic review on natural language processing systems for eligibility prescreening in clinical researchabstractOBJECTIVE: We conducted a systematic review to assess the effect of natural language processing (NLP) systems in improving the accuracy and efficiency of eligibility prescreening during the clinical research recruitment process. MATERIALS AND METHODS: Guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) standards of quality for reporting systematic reviews, a protocol for study eligibility was developed a priori and registered in the PROSPERO database. Using predetermined inclusion criteria, studies published from database inception through February 2021 were identified from 5 databases. The Joanna Briggs Institute Critical Appraisal Checklist for Quasi-experimental Studies was adapted to determine the study quality and the risk of bias of the included articles. RESULTS: Eleven studies representing 8 unique NLP systems met the inclusion criteria. These studies demonstrated moderate study quality and exhibited heterogeneity in the study design, setting, and intervention type. All 11 studies evaluated the NLP system's performance for identifying eligible participants; 7 studies evaluated the system's impact on time efficiency; 4 studies evaluated the system's impact on workload; and 2 studies evaluated the system's impact on recruitment. DISCUSSION: NLP systems in clinical research eligibility prescreening are an understudied but promising field that requires further research to assess its impact on real-world adoption. Future studies should be centered on continuing to develop and evaluate relevant NLP systems to improve enrollment into clinical studies. CONCLUSION: Understanding the role of NLP systems in improving eligibility prescreening is critical to the advancement of clinical research recruitment. Betina Ross S. Idnay, Caitlin N. Dreisbach, Chunhua Weng, Rebecca Schnall |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | UMLS-based data augmentation for natural language processing of clinical research literatureabstractOBJECTIVE: The study sought to develop and evaluate a knowledge-based data augmentation method to improve the performance of deep learning models for biomedical natural language processing by overcoming training data scarcity. MATERIALS AND METHODS: We extended the easy data augmentation (EDA) method for biomedical named entity recognition (NER) by incorporating the Unified Medical Language System (UMLS) knowledge and called this method UMLS-EDA. We designed experiments to systematically evaluate the effect of UMLS-EDA on popular deep learning architectures for both NER and classification. We also compared UMLS-EDA to BERT. RESULTS: UMLS-EDA enables substantial improvement for NER tasks from the original long short-term memory conditional random fields (LSTM-CRF) model (micro-F1 score: +5%, + 17%, and +15%), helps the LSTM-CRF model (micro-F1 score: 0.66) outperform LSTM-CRF with transfer learning by BERT (0.63), and improves the performance of the state-of-the-art sentence classification model. The largest gain on micro-F1 score is 9%, from 0.75 to 0.84, better than classifiers with BERT pretraining (0.82). CONCLUSIONS: This study presents a UMLS-based data augmentation method, UMLS-EDA. It is effective at improving deep learning models for both NER and sentence classification, and contributes original insights for designing new, superior deep learning approaches for low-resource biomedical domains. Tian Kang, Adler J. Perotte, Youlan Tang, Casey N. Ta, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | A neuro-symbolic method for understanding free-text medical evidenceabstractOBJECTIVE: We introduce Medical evidence Dependency (MD)-informed attention, a novel neuro-symbolic model for understanding free-text clinical trial publications with generalizability and interpretability. MATERIALS AND METHODS: We trained one head in the multi-head self-attention model to attend to the Medical evidence Ddependency (MD) and to pass linguistic and domain knowledge on to later layers (MD informed). This MD-informed attention model was integrated into BioBERT and tested on 2 public machine reading comprehension benchmarks for clinical trial publications: Evidence Inference 2.0 and PubMedQA. We also curated a small set of recently published articles reporting randomized controlled trials on COVID-19 (coronavirus disease 2019) following the Evidence Inference 2.0 guidelines to evaluate the model's robustness to unseen data. RESULTS: The integration of MD-informed attention head improves BioBERT substantially in both benchmark tasks-as large as an increase of +30% in the F1 score-and achieves the new state-of-the-art performance on the Evidence Inference 2.0. It achieves 84% and 82% in overall accuracy and F1 score, respectively, on the unseen COVID-19 data. CONCLUSIONS: MD-informed attention empowers neural reading comprehension models with interpretability and generalizability via reusable domain knowledge. Its compositionality can benefit any transformer-based architecture for machine reading comprehension of free-text medical evidence. Tian Kang, Ali Turfah, Adler J. Perotte, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Towards clinical data-driven eligibility criteria optimization for interventional COVID-19 clinical trialsabstractOBJECTIVE: This research aims to evaluate the impact of eligibility criteria on recruitment and observable clinical outcomes of COVID-19 clinical trials using electronic health record (EHR) data. MATERIALS AND METHODS: On June 18, 2020, we identified frequently used eligibility criteria from all the interventional COVID-19 trials in ClinicalTrials.gov (n = 288), including age, pregnancy, oxygen saturation, alanine/aspartate aminotransferase, platelets, and estimated glomerular filtration rate. We applied the frequently used criteria to the EHR data of COVID-19 patients in Columbia University Irving Medical Center (CUIMC) (March 2020-June 2020) and evaluated their impact on patient accrual and the occurrence of a composite endpoint of mechanical ventilation, tracheostomy, and in-hospital death. RESULTS: There were 3251 patients diagnosed with COVID-19 from the CUIMC EHR included in the analysis. The median follow-up period was 10 days (interquartile range 4-28 days). The composite events occurred in 18.1% (n = 587) of the COVID-19 cohort during the follow-up. In a hypothetical trial with common eligibility criteria, 33.6% (690/2051) were eligible among patients with evaluable data and 22.2% (153/690) had the composite event. DISCUSSION: By adjusting the thresholds of common eligibility criteria based on the characteristics of COVID-19 patients, we could observe more composite events from fewer patients. CONCLUSIONS: This research demonstrated the potential of using the EHR data of COVID-19 patients to inform the selection of eligibility criteria and their thresholds, supporting data-driven optimization of participant selection towards improved statistical power of COVID-19 trials. Jae Hyun Kim, Casey N. Ta, Cong Liu 0020, Cynthia Sung 0002, Alex M. Butler, Latoya A. Stewart, Lyudmila Ena, James R. Rogers, Anna Ostropolets, Patrick B. Ryan, Hao Liu 0054, Shing M. Lee, Mitchell S. V. Elkind, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 15 |
| 2021 | Contemporary use of real-world data for clinical trial conduct in the United States: a scoping reviewabstractOBJECTIVE: Real-world data (RWD), defined as routinely collected healthcare data, can be a potential catalyst for addressing challenges faced in clinical trials. We performed a scoping review of database-specific RWD applications within clinical trial contexts, synthesizing prominent uses and themes. MATERIALS AND METHODS: Querying 3 biomedical literature databases, research articles using electronic health records, administrative claims databases, or clinical registries either within a clinical trial or in tandem with methodology related to clinical trials were included. Articles were required to use at least 1 US RWD source. All abstract screening, full-text screening, and data extraction was performed by 1 reviewer. Two reviewers independently verified all decisions. RESULTS: Of 2020 screened articles, 89 qualified: 59 articles used electronic health records, 29 used administrative claims, and 26 used registries. Our synthesis was driven by the general life cycle of a clinical trial, culminating into 3 major themes: trial process tasks (51 articles); dissemination strategies (6); and generalizability assessments (34). Despite a diverse set of diseases studied, <10% of trials using RWD for trial process tasks evaluated medications or procedures (5/51). All articles highlighted data-related challenges, such as missing values. DISCUSSION: Database-specific RWD have been occasionally leveraged for various clinical trial tasks. We observed underuse of RWD within conducted medication or procedure trials, though it is subject to the confounder of implicit report of RWD use. CONCLUSION: Enhanced incorporation of RWD should be further explored for medication or procedure trials, including better understanding of how to handle related data quality issues to facilitate RWD use. James R. Rogers, Ying Kuen Cheung, George Hripcsak, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 6 |
| 2021 | The COVID-19 Trial FinderabstractClinical trials are the gold standard for generating reliable medical evidence. The biggest bottleneck in clinical trials is recruitment. To facilitate recruitment, tools for patient search of relevant clinical trials have been developed, but users often suffer from information overload. With nearly 700 coronavirus disease 2019 (COVID-19) trials conducted in the United States as of August 2020, it is imperative to enable rapid recruitment to these studies. The COVID-19 Trial Finder was designed to facilitate patient-centered search of COVID-19 trials, first by location and radius distance from trial sites, and then by brief, dynamically generated medical questions to allow users to prescreen their eligibility for nearby COVID-19 trials with minimum human computer interaction. A simulation study using 20 publicly available patient case reports demonstrates its precision and effectiveness. Yingcheng Sun, Alex M. Butler, Fengyang Lin, Hao Liu 0054, Latoya A. Stewart, Jae Hyun Kim, Betina Ross S. Idnay, Qingyin Ge, Xinyi Wei, Cong Liu 0020, Chi Yuan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 12 |
| 2021 | From clinical trials to clinical practice: How long are drugs tested and then used by patients?abstractOBJECTIVE: Evidence is scarce regarding the safety of long-term drug use, especially for drugs treating chronic diseases. To bridge this knowledge gap, this research investigated the differences in drug exposure between clinical trials and clinical practice. MATERIALS AND METHODS: We extracted drug follow-up times from clinical trials in ClinicalTrials.gov and compared the difference between clinical trials and real-world usage data for 914 drugs taken by 96 645 927 patients. RESULTS: A total of 17.5% of drugs had longer median exposure in practice than in trials, 6% of patients had extended exposure to at least 1 drug, and drugs treating nervous system disorders and cardiovascular diseases were the most common among drugs with high rates of extended exposure. CONCLUSIONS: For most of patients, the drug use length is shorter than the tested length in clinical trials. Still, a remarkable number of patients experienced extended drug exposure, particularly for drugs treating nervous system disorders or cardiovascular disorders. Chi Yuan, Patrick B. Ryan, Casey N. Ta, Jae Hyun Kim, Ziran Li, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 6 |
| 2021 | A conceptual framework for external validity
Amelia J. Averitt, Patrick B. Ryan, Chunhua Weng, Adler J. Perotte |
J. Biomed. Informatics | 3 |
| 2021 | Similarity-based health risk prediction using Domain Fusion and electronic health records data
Chi Yuan, Ning Shang 0004, Natalie A. Bello, Krzysztof Kiryluk, Chunhua Weng |
J. Biomed. Informatics | 7 |
| 2021 | A knowledge base of clinical trial eligibility criteria
Hao Liu 0054, Chi Yuan, Alex M. Butler, Yingcheng Sun, Chunhua Weng |
J. Biomed. Informatics | 5 |
| 2021 | Clinical comparison between trial participants and potentially eligible patients using electronic health record data: A generalizability assessment method
James R. Rogers, George Hripcsak, Ying Kuen Cheung, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2021 | Building an OMOP common data model-compliant annotated corpus for COVID-19 clinical trials
Yingcheng Sun, Alex M. Butler, Latoya A. Stewart, Hao Liu 0054, Chi Yuan, Christopher T. Southard, Jae Hyun Kim, Chunhua Weng |
J. Biomed. Informatics | 8 |
| 2020 | Impact of IMPACT: Longitudinal Analysis of an Integrated Participant Scheduling System in a Clinical Research Setting
Alex M. Butler, Yat S. So, Linda Busacca, Karen Marder, Henry N. Ginsberg, Dianne Frederick, Ismael Castaneda, Elizabeth Guerrido, Chunhua Weng |
AMIA | 10 |
| 2020 | The Open Health Natural Language Processing Collaboratory
Xiaoqian Jiang, Serguei V. S. Pakhomov, Chunhua Weng, Hua Xu 0001 |
AMIA | 4 |
| 2020 | Potential Role of Clinical Trial Eligibility Criteria in Electronic Phenotyping
Hao Liu 0054, Chi Yuan, Alex M. Butler, Yingcheng Sun, Chunhua Weng |
AMIA | 5 |
| 2020 | Characterizing database granularity using SNOMED-CT hierarchy
Anna Ostropolets, Christian G. Reich, Patrick B. Ryan, Chunhua Weng, Anthony Molinaro, Frank J. DeFalco, Jitendra Jonnagaddala, Siaw-Teng Liaw, Hokyun Jeon, Rae Woong Park, Matthew E. Spotnitz, Karthik Natarajan, Kristin Kostka, George Argyriou, Robert T. Miller, Andrew E. Williams, Evan P. Minty, José D. Posada, George Hripcsak |
AMIA | 4 |
| 2020 | Contemporary Use of Real World Data for Clinical Trial Conduct
James R. Rogers, Patrick B. Ryan, George Hripcsak, Chunhua Weng |
AMIA | 6 |
| 2020 | Understanding the nature and scope of clinical research commentaries in PubMedabstractScientific commentaries are expected to play an important role in evidence appraisal, but it is unknown whether this expectation has been fulfilled. This study aims to better understand the role of scientific commentary in evidence appraisal. We queried PubMed for all clinical research articles with accompanying comments and extracted corresponding metadata. Five percent of clinical research studies (N = 130 629) received postpublication comments (N = 171 556), resulting in 178 882 comment-article pairings, with 90% published in the same journal. We obtained 5197 full-text comments for topic modeling and exploratory sentiment analysis. Topics were generally disease specific with only a few topics relevant to the appraisal of studies, which were highly prevalent in letters. Of a random sample of 518 full-text comments, 67% had a supportive tone. Based on our results, published commentary, with the exception of letters, most often highlight or endorse previous publications rather than serve as a prominent mechanism for critical appraisal. James R. Rogers, Hollis Mills, Lisa Grossman Liu, Andrew Goldstein, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | A graph-based method for reconstructing entities from coordination ellipsis in medical textabstractOBJECTIVE: Coordination ellipsis is a linguistic phenomenon abound in medical text and is challenging for concept normalization because of difficulty in recognizing elliptical expressions referencing 2 or more entities accurately. To resolve this bottleneck, we aim to contribute a generalizable method to reconstruct concepts from medical coordinated elliptical expressions in a variety of biomedical corpora. MATERIALS AND METHODS: We proposed a graph-based representation model and built a pipeline to reconstruct concepts from coordinated elliptical expressions in medical text (RECEEM). There are 4 modules: (1) identify all possible candidate conjunct pairs from original coordinated elliptical expressions, (2) calculate coefficients for candidate conjuncts using the embedding model, (3) select the most appropriate decompositions by global optimization, and (4) rebuild concepts based on a pathfinding algorithm. We evaluated the pipeline's performance on 2658 coordinated elliptical expressions from 3 different medical corpora (ie, biomedical literature, clinical narratives, and eligibility criteria from clinical trials). Precision, recall, and F1 score were calculated. RESULTS: The F1 scores for biomedical publications, clinical narratives, and research eligibility criteria were 0.862, 0.721, and 0.870, respectively. RECEEM outperformed 2 previously released methods. By incorporating RECEEM into 2 existing NLP tools, the F1 scores increased from 0.248 to 0.460 and from 0.287 to 0.630 on concept mapping of 1125 coordination ellipses. CONCLUSIONS: RECEEM improves concept normalization for medical coordinated elliptical expressions in a variety of biomedical corpora. It outperformed existing methods and significantly enhanced the performance of 2 notable NLP systems for mapping coordination ellipses in the evaluation. The algorithm is open sourced online (https://github.com/chiyuan1126/RECEEM). Chi Yuan, Ning Shang 0004, Ziran Li, Ruxin Zhao, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 6 |
| 2020 | Adapting electronic health records-derived phenotypes to claims data: Lessons learned in using limited clinical data for phenotyping
Anna Ostropolets, Christian G. Reich, Patrick B. Ryan, Ning Shang 0004, George Hripcsak, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2020 | Deep phenotyping: Embracing complexity and temporality - Towards scalability, portability, and interoperability
Chunhua Weng, Nigam H. Shah, George Hripcsak |
J. Biomed. Informatics | 1 |
| 2019 | Completeness and Concordance in Electronic Health Record (EHR) Documentation of Chemotherapy-Induced Nausea and Vomiting (CINV) in Pediatric Cancer
Melissa Beauchemin, Rebecca Schnall, Chunhua Weng |
AMIA | 3 |
| 2019 | Hidden Gaps in Using Common Data Models to Achieve Interoperability Between Electronic Phenotypes and Clinical Data
Fabricio Sampaio Peres Kury, Li-heng Fu, Chi Yuan, Ida Sim, Simona Carini, Chunhua Weng |
AMIA | 6 |
| 2019 | Clinical Use of an Information Retrieval Framework for Cohort Discovery from Electronic Health Records
Yanshan Wang, Andrew Wen, Sijia Liu 0002, Jennifer L. St. Sauver, Adil E. Bharucha, Chunhua Weng |
AMIA | 6 |
| 2019 | User engagement with web-based genomics education videos and implications for designing scalable patient education materials
Julia Wynn, Yat So, Suzanne Bakken, Chunhua Weng, Wendy K. Chung |
AMIA | 6 |
| 2019 | DQueST: dynamic questionnaire for search of clinical trialsabstractOBJECTIVE: Information overload remains a challenge for patients seeking clinical trials. We present a novel system (DQueST) that reduces information overload for trial seekers using dynamic questionnaires. MATERIALS AND METHODS: DQueST first performs information extraction and criteria library curation. DQueST transforms criteria narratives in the ClinicalTrials.gov repository into a structured format, normalizes clinical entities using standard concepts, clusters related criteria, and stores the resulting curated library. DQueST then implements a real-time dynamic question generation algorithm. During user interaction, the initial search is similar to a standard search engine, and then DQueST performs real-time dynamic question generation to select criteria from the library 1 at a time by maximizing its relevance score that reflects its ability to rule out ineligible trials. DQueST dynamically updates the remaining trial set by removing ineligible trials based on user responses to corresponding questions. The process iterates until users decide to stop and begin manually reviewing the remaining trials. RESULTS: In simulation experiments initiated by 10 diseases, DQueST reduced information overload by filtering out 60%-80% of initial trials after 50 questions. Reviewing the generated questions against previous answers, on average, 79.7% of the questions were relevant to the queried conditions. By examining the eligibility of random samples of trials ruled out by DQueST, we estimate the accuracy of the filtering procedure is 63.7%. In a study using 5 mock patient profiles, DQueST on average retrieved trials with a 1.465 times higher density of eligible trials than an existing search engine. In a patient-centered usability evaluation, patients found DQueST useful, easy to use, and returning relevant results. CONCLUSION: DQueST contributes a novel framework for transforming free-text eligibility criteria to questions and filtering out clinical trials based on user answers to questions dynamically. It promises to augment keyword-based methods to improve clinical trial search. Cong Liu 0020, Chi Yuan, Alex M. Butler, Richard D. Carvajal, Ziran Ryan Li, Casey N. Ta, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 7 |
| 2019 | Criteria2Query: a natural language interface to clinical databases for cohort definitionabstractOBJECTIVE: Cohort definition is a bottleneck for conducting clinical research and depends on subjective decisions by domain experts. Data-driven cohort definition is appealing but requires substantial knowledge of terminologies and clinical data models. Criteria2Query is a natural language interface that facilitates human-computer collaboration for cohort definition and execution using clinical databases. MATERIALS AND METHODS: Criteria2Query uses a hybrid information extraction pipeline combining machine learning and rule-based methods to systematically parse eligibility criteria text, transforms it first into a structured criteria representation and next into sharable and executable clinical data queries represented as SQL queries conforming to the OMOP Common Data Model. Users can interactively review, refine, and execute queries in the ATLAS web application. To test effectiveness, we evaluated 125 criteria across different disease domains from ClinicalTrials.gov and 52 user-entered criteria. We evaluated F1 score and accuracy against 2 domain experts and calculated the average computation time for fully automated query formulation. We conducted an anonymous survey evaluating usability. RESULTS: Criteria2Query achieved 0.795 and 0.805 F1 score for entity recognition and relation extraction, respectively. Accuracies for negation detection, logic detection, entity normalization, and attribute normalization were 0.984, 0.864, 0.514 and 0.793, respectively. Fully automatic query formulation took 1.22 seconds/criterion. More than 80% (11+ of 13) of users would use Criteria2Query in their future cohort definition tasks. CONCLUSIONS: We contribute a novel natural language interface to clinical databases. It is open source and supports fully automated and interactive modes for autonomous data-driven cohort definition by researchers with minimal human effort. We demonstrate its promising user friendliness and usability. Chi Yuan, Patrick B. Ryan, Casey N. Ta, Yixuan Guo, Ziran Li, Jill Hardin, Rupa Makadia, Ning Shang 0004, Tian Kang, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 11 |
| 2019 | Semi-supervised learning to improve generalizability of risk prediction models
Shengqiang Chi, Yu Tian 0002, Jun Li 0099, Xiang-Xing Kong, Kefeng Ding, Chunhua Weng, Jingsong Li 0001 |
J. Biomed. Informatics | 7 |
| 2019 | Facilitating phenotype transfer using a common data model
George Hripcsak, Ning Shang 0004, Peggy L. Peissig, Luke V. Rasmussen, Cong Liu 0020, Barbara Benoit, Robert J. Carroll, David Carrell, Joshua C. Denny, Ozan Dikilitas, Vivian S. Gainer, Kayla Marie Howell, Jeffrey G. Klann, Iftikhar J. Kullo, Todd Lingren, Frank D. Mentch, Shawn N. Murphy, Karthik Natarajan, Chunhua Weng |
J. Biomed. Informatics | 19 |
| 2019 | Ensembles of natural language processing systems for portable phenotyping solutions
Cong Liu 0020, Casey N. Ta, James R. Rogers, Ziran Li, Alex M. Butler, Ning Shang 0004, Fabricio Sampaio Peres Kury, Liwei Wang 0010, Feichen Shen, Lyudmila Ena, Carol Friedman, Chunhua Weng |
J. Biomed. Informatics | 14 |
| 2019 | Pathway analysis of genomic pathology tests for prognostic cancer subtyping
Olga Lyudovyk, Yufeng Shen, Nicholas P. Tatonetti, Susan J. Hsiao, Mahesh M. Mansukhani, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2019 | Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng |
J. Biomed. Informatics | 19 |
| 2018 | Sequencing EHR for Disease Subtyping
Po-Hsiang Chiu, Ning Shang 0004, Chunhua Weng |
AMIA | 3 |
| 2018 | An Exploration of the Terminology of Clinical Cognition and Reasoning
James J. Cimino, Ziran Li, Chunhua Weng |
AMIA | 3 |
| 2018 | Clinical Concept Value Sets and Interoperability in Health Data Analytics
Sigfried Gold, Andrea Batch, Robert C. McClure, Guoqian Jiang, Hadi Kharrazi, Rishi Saripalle, Vojtech Huser, Chunhua Weng, Nancy K. Roderer, Ana Szarfman, Niklas Elmqvist, David Gotz |
AMIA | 8 |
| 2018 | Dynamic Questionnaire Generation for Efficient Patient Search of Trials
Cong Liu 0020, Chi Yuan, Eric Pua, Chunhua Weng |
AMIA | 5 |
| 2018 | Guideline-driven data quality assessment- A case study of warfarin management
Yaoyun Zhang, Hua Xu 0001, Chunhua Weng |
AMIA | 4 |
| 2018 | Empowering genomic medicine by establishing critical sequencing result data flows: the eMERGE exampleabstractThe eMERGE Network is establishing methods for electronic transmittal of patient genetic test results from laboratories to healthcare providers across organizational boundaries. We surveyed the capabilities and needs of different network participants, established a common transfer format, and implemented transfer mechanisms based on this format. The interfaces we created are examples of the connectivity that must be instantiated before electronic genetic and genomic clinical decision support can be effectively built at the point of care. This work serves as a case example for both standards bodies and other organizations working to build the infrastructure required to provide better electronic clinical decision support for clinicians. Samuel J. Aronson, Lawrence J. Babb, Darren C. Ames, Richard A. Gibbs, Eric Venner, John J. Connelly, Keith Marsolo, Chunhua Weng, Marc S. Williams, Andrea L. Hartzler, Wayne H. Liang, James D. Ralston, Emily Beth Devine, Shawn N. Murphy, Christopher G. Chute, Pedro J. Caraballo, Iftikhar J. Kullo, Robert R. Freimuth, Luke V. Rasmussen, Firas H. Wehbe, Josh F. Peterson, Jamie R. Robinson, Ken Wiley, Casey Overby Taylor |
J. Am. Medical Informatics Assoc. | 8 |
| 2018 | The representativeness of eligible patients in type 2 diabetes trials: a case study using GIST 2.0abstractOBJECTIVE: The population representativeness of a clinical study is influenced by how real-world patients qualify for the study. We analyze the representativeness of eligible patients for multiple type 2 diabetes trials and the relationship between representativeness and other trial characteristics. METHODS: Sixty-nine study traits available in the electronic health record data for 2034 patients with type 2 diabetes were used to profile the target patients for type 2 diabetes trials. A set of 1691 type 2 diabetes trials was identified from ClinicalTrials.gov, and their population representativeness was calculated using the published Generalizability Index of Study Traits 2.0 metric. The relationships between population representativeness and number of traits and between trial duration and trial metadata were statistically analyzed. A focused analysis with only phase 2 and 3 interventional trials was also conducted. RESULTS: A total of 869 of 1691 trials (51.4%) and 412 of 776 phase 2 and 3 interventional trials (53.1%) had a population representativeness of <5%. The overall representativeness was significantly correlated with the representativeness of the Hba1c criterion. The greater the number of criteria or the shorter the trial, the less the representativeness. Among the trial metadata, phase, recruitment status, and start year were found to have a statistically significant effect on population representativeness. For phase 2 and 3 interventional trials, only start year was significantly associated with representativeness. CONCLUSIONS: Our study quantified the representativeness of multiple type 2 diabetes trials. The common low representativeness of type 2 diabetes trials could be attributed to specific study design requirements of trials or safety concerns. Rather than criticizing the low representativeness, we contribute a method for increasing the transparency of the representativeness of clinical trials. Anando Sen, Andrew Goldstein, Shreya Chakrabarti, Ning Shang 0004, Tian Kang, Anil Yaman, Patrick B. Ryan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 8 |
| 2018 | A conceptual framework for evaluating data suitability for observational studiesabstractOBJECTIVE: To contribute a conceptual framework for evaluating data suitability to satisfy the research needs of observational studies. MATERIALS AND METHODS: Suitability considerations were derived from a systematic literature review on researchers' common data needs in observational studies and a scoping review on frequent clinical database design considerations, and were harmonized to construct a suitability conceptual framework using a bottom-up approach. The relationships among the suitability categories are explored from the perspective of 4 facets of data: intrinsic, contextual, representational, and accessible. A web-based national survey of domain experts was conducted to validate the framework. RESULTS: Data suitability for observational studies hinges on the following key categories: Explicitness of Policy and Data Governance, Relevance, Availability of Descriptive Metadata and Provenance Documentation, Usability, and Quality. We describe 16 measures and 33 sub-measures. The survey uncovered the relevance of all categories, with a 5-point Likert importance score of 3.9 ± 1.0 for Explicitness of Policy and Data Governance, 4.1 ± 1.0 for Relevance, 3.9 ± 0.9 for Availability of Descriptive Metadata and Provenance Documentation, 4.2 ± 1.0 for Usability, and 4.0 ± 0.9 for Quality. CONCLUSIONS: The suitability framework evaluates a clinical data source's fitness for research use. Its construction reflects both researchers' points of view and data custodians' design features. The feedback from domain experts rated Usability, Relevance, and Quality categories as the most important considerations. Ning Shang 0004, Chunhua Weng, George Hripcsak |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | A method for harmonization of clinical abbreviation and acronym sense inventories
Lisa Grossman Liu, Elliot G. Mitchell, George Hripcsak, Chunhua Weng, David K. Vawdrey |
J. Biomed. Informatics | 4 |
| 2018 | The ranking of scientists
Chunhua Weng, Andrew Goldstein, Chi Yuan, Zhiping Zhou |
J. Biomed. Informatics | 1 |
| 2018 | Call for papers: Deep phenotyping for Precision Medicine
Chunhua Weng, Nigam H. Shah, George Hripcsak |
J. Biomed. Informatics | 1 |
| 2017 | Clinical Trial Eligibility Criteria Fail to Meet Burden of Generalizability
Amelia J. Averitt, Chunhua Weng, Adler J. Perotte |
AMIA | 2 |
| 2017 | Discordances between Patient-Reported Family History and Family Histories in EHR
Jung Hoon Son, Chunhua Weng |
AMIA | 2 |
| 2017 | Criteria2Query: Automatically Transforming Clinical Research Eligibility Criteria Text to OMOP Common Data Model (CDM)-based Cohort Queries
Chi Yuan, Patrick B. Ryan, Yixuan Guo, Tian Kang, Chunhua Weng |
AMIA | 6 |
| 2017 | Evidence appraisal: a scoping review, conceptual framework, and research agendaabstractOBJECTIVE: Critical appraisal of clinical evidence promises to help prevent, detect, and address flaws related to study importance, ethics, validity, applicability, and reporting. These research issues are of growing concern. The purpose of this scoping review is to survey the current literature on evidence appraisal to develop a conceptual framework and an informatics research agenda. METHODS: We conducted an iterative literature search of Medline for discussion or research on the critical appraisal of clinical evidence. After title and abstract review, 121 articles were included in the analysis. We performed qualitative thematic analysis to describe the evidence appraisal architecture and its issues and opportunities. From this analysis, we derived a conceptual framework and an informatics research agenda. RESULTS: We identified 68 themes in 10 categories. This analysis revealed that the practice of evidence appraisal is quite common but is rarely subjected to documentation, organization, validation, integration, or uptake. This is related to underdeveloped tools, scant incentives, and insufficient acquisition of appraisal data and transformation of the data into usable knowledge. DISCUSSION: The gaps in acquiring appraisal data, transforming the data into actionable information and knowledge, and ensuring its dissemination and adoption can be addressed with proven informatics approaches. CONCLUSIONS: Evidence appraisal faces several challenges, but implementing an informatics research agenda would likely help realize the potential of evidence appraisal for improving the rigor and value of clinical evidence. Andrew Goldstein, Eric Venker, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 3 |
| 2017 | EliIE: An open-source information extraction system for clinical trial eligibility criteriaabstractOBJECTIVE: To develop an open-source information extraction system called Eligibility Criteria Information Extraction (EliIE) for parsing and formalizing free-text clinical research eligibility criteria (EC) following Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) version 5.0. MATERIALS AND METHODS: EliIE parses EC in 4 steps: (1) clinical entity and attribute recognition, (2) negation detection, (3) relation extraction, and (4) concept normalization and output structuring. Informaticians and domain experts were recruited to design an annotation guideline and generate a training corpus of annotated EC for 230 Alzheimer's clinical trials, which were represented as queries against the OMOP CDM and included 8008 entities, 3550 attributes, and 3529 relations. A sequence labeling-based method was developed for automatic entity and attribute recognition. Negation detection was supported by NegEx and a set of predefined rules. Relation extraction was achieved by a support vector machine classifier. We further performed terminology-based concept normalization and output structuring. RESULTS: In task-specific evaluations, the best F1 score for entity recognition was 0.79, and for relation extraction was 0.89. The accuracy of negation detection was 0.94. The overall accuracy for query formalization was 0.71 in an end-to-end evaluation. CONCLUSIONS: This study presents EliIE, an OMOP CDM-based information extraction system for automatic structuring and formalization of free-text EC. According to our evaluation, machine learning-based EliIE outperforms existing systems and shows promise to improve. Tian Kang, Shaodian Zhang, Youlan Tang, Gregory William Hruby, Alex Rusanov, Noémie Elhadad, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 7 |
| 2016 | SMASH: A Data-driven Informatics Method to Assist Experts in Characterizing Semantic Heterogeneity among Data Elements
William Brown III 0001, Chunhua Weng, David K. Vawdrey, Alex Carballo-Dieguez, Suzanne Bakken |
AMIA | 2 |
| 2016 | Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001 |
AMIA | 8 |
| 2016 | A Method for Enhancing the Portability of Electronic Phenotyping Algorithms: An eMERGE Pilot Study
Ning Shang 0004, Chunhua Weng, George Hripcsak |
AMIA | 2 |
| 2016 | Multivariate analysis of the population representativeness of related clinical studies
Zhe He 0001, Patrick B. Ryan, Julia Hoxha, Simona Carini, Ida Sim, Chunhua Weng |
J. Biomed. Informatics | 7 |
| 2016 | DREAM: Classification scheme for dialog acts in clinical research query mediation
Julia Hoxha, Praveen Chandar Ravichandran, Zhe He 0001, James J. Cimino, David A. Hanauer, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2016 | Automated learning of domain taxonomies from text using background knowledge
Julia Hoxha, Guoqian Jiang, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2016 | Leveraging dialog systems research to assist biomedical researchers' interrogation of Big Clinical Data
Julia Hoxha, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2016 | Facilitating biomedical researchers' interrogation of electronic health record data: Ideas from outside of biomedical informatics
Gregory William Hruby, Konstantina Matsoukas, James J. Cimino, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2016 | Prediction of black box warning by mining patterns of Convergent Focus Shift in clinical trial study populations using linked public data
Handong Mao, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2016 | GIST 2.0: A scalable multi-trait metric for quantifying population representativeness of individual clinical studies
Anando Sen, Shreya Chakrabarti, Andrew Goldstein, Patrick B. Ryan, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2015 | Simulation-based Evaluation of the Generalizability Index for Study Traits
Zhe He 0001, Praveen Chandar Ravichandran, Patrick B. Ryan, Chunhua Weng |
AMIA | 4 |
| 2015 | Developing an Ontology from HIV-associated Elements in Research
William Brown III 0001, Chunhua Weng, David K. Vawdrey, Alex Carballo-Dieguez, Suzanne Bakken |
AMIA | 2 |
| 2015 | What Are Frequent Data Requests from Researchers? A Conceptual Model of Researchers' EHR Data Needs for Comparative Effectiveness Research
Gregory William Hruby, Praveen Chandar Ravichandran, Julia Hoxha, Eneida A. Mendonça, David A. Hanauer, Chunhua Weng |
AMIA | 6 |
| 2015 | ClinicalTrials.gov: Adding Value through Informatics
Vojtech Huser, Alexa T. McCray, Neil R. Smalheiser, Asba Tasneem, Chunhua Weng |
AMIA | 5 |
| 2015 | ArticlesAboutMe.org: Disseminating Clinical Trials Results to Patients
Vojtech Huser, Anil Yaman, Chunhua Weng, James J. Cimino |
AMIA | 3 |
| 2015 | Quality Assurance of Cancer Study Common Data Elements Using A Post-Coordination Approach
Guoqian Jiang, Harold R. Solbrig, Eric Prud'hommeaux, Cui Tao, Chunhua Weng, Christopher G. Chute |
AMIA | 5 |
| 2015 | Initial Readability Assessment of Clinical Trial Eligibility Criteria
Tian Kang, Noémie Elhadad, Chunhua Weng |
AMIA | 3 |
| 2015 | Exploration of Temporal ICD Coding Bias Related to Acute Diabetic Conditions
Mollie McKillop, Fernanda Polubriaginof, Chunhua Weng |
AMIA | 3 |
| 2015 | Desiderata for Major Eligibility Criteria in Breast Cancer Clinical Trials
Matthew L. Paulson, Chunhua Weng |
AMIA | 2 |
| 2015 | Similarity-Based Recommendation of New Concepts to a Terminology
Praveen Chandar Ravichandran, Anil Yaman, Julia Hoxha, Zhe He 0001, Chunhua Weng |
AMIA | 5 |
| 2015 | Understanding Challenges and Opportunities in Precision Medicine and Interoperability Using Informatics Approaches
Jeremiah Geronimo Ronquillo, Chunhua Weng, William T. Lester |
AMIA | 2 |
| 2015 | Representation of Genetic Variants in Genomic Sequencing Reports
Evelyn Rustia, Chunhua Weng, Carol Friedman |
AMIA | 2 |
| 2015 | A Framework for Assessing Clinical Data Suitability for Observational Study
Ning Shang 0004, Chunhua Weng, George Hripcsak |
AMIA | 2 |
| 2015 | A Guideline for Assessing EHR Data Quality for Secondary Use
Nicole Gray Weiskopf, Chunhua Weng |
AMIA | 2 |
| 2015 | Case-based reasoning using electronic health records efficiently identifies eligible patients for clinical trialsabstractOBJECTIVE: To develop a cost-effective, case-based reasoning framework for clinical research eligibility screening by only reusing the electronic health records (EHRs) of minimal enrolled participants to represent the target patient for each trial under consideration. MATERIALS AND METHODS: The EHR data--specifically diagnosis, medications, laboratory results, and clinical notes--of known clinical trial participants were aggregated to profile the "target patient" for a trial, which was used to discover new eligible patients for that trial. The EHR data of unseen patients were matched to this "target patient" to determine their relevance to the trial; the higher the relevance, the more likely the patient was eligible. Relevance scores were a weighted linear combination of cosine similarities computed over individual EHR data types. For evaluation, we identified 262 participants of 13 diversified clinical trials conducted at Columbia University as our gold standard. We ran a 2-fold cross validation with half of the participants used for training and the other half used for testing along with other 30 000 patients selected at random from our clinical database. We performed binary classification and ranking experiments. RESULTS: The overall area under the ROC curve for classification was 0.95, enabling the highlight of eligible patients with good precision. Ranking showed satisfactory results especially at the top of the recommended list, with each trial having at least one eligible patient in the top five positions. CONCLUSIONS: This relevance-based method can potentially be used to identify eligible patients for clinical trials by processing patient EHR data alone without parsing free-text eligibility criteria, and shows promise of efficient "case-based reasoning" modeled only on minimal trial participants. Riccardo Miotto, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Visual aggregate analysis of eligibility features of clinical trials
Zhe He 0001, Simona Carini, Ida Sim, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2015 | Combining expert knowledge and knowledge automatically acquired from electronic data sources for continued ontology evaluation and improvement
Claire L. Gordon, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2014 | A Method for Analyzing Commonalities in Clinical Trial Target Populations
Zhe He 0001, Simona Carini, Tianyong Hao, Ida Sim, Chunhua Weng |
AMIA | 5 |
| 2014 | Integrating Diverse HIV-associated Datasets via Semantic Harmonization
William Brown III 0001, Chunhua Weng, David K. Vawdrey, Suzanne Bakken |
AMIA | 2 |
| 2014 | Could Patient Self-reported Health Data Complement EHR for Phenotyping?
Daniel Fort, Adam B. Wilcox, Chunhua Weng |
AMIA | 3 |
| 2014 | What Is Asked in Clinical Data Request Forms? A Multi-site Thematic Analysis of Forms Towards Better Data Access Support
David A. Hanauer, Gregory William Hruby, Daniel Fort, Luke V. Rasmussen, Eneida A. Mendonça, Chunhua Weng |
AMIA | 6 |
| 2014 | Design and Evaluation of a Glomerular Disease Ontology
Jamie S. Hirsch, Chunhua Weng |
AMIA | 2 |
| 2014 | Patients Screening for Clinical Trials Using EHR Representation Similarities
Riccardo Miotto, Chunhua Weng |
AMIA | 2 |
| 2014 | Development and validation of an electronic phenotyping algorithm for chronic kidney disease
Girish N. Nadkarni, Omri Gottesman, James G. Linneman, Herbert S. Chase, Richard L. Berg, Samira Farouk, Vaneet Lotay, Stephen B. Ellis, George Hripcsak, Peggy L. Peissig, Chunhua Weng, Rajiv Nadukuru, Erwin P. Bottinger |
AMIA | 11 |
| 2014 | Network Analysis of Common Eligibility Criteria in Cancer Clinical Trials
Chin S. Park, Jacqueline Merrill, Chunhua Weng |
AMIA | 3 |
| 2014 | Unsupervised Time-Series Clustering for Identifying Uncontrolled Type-2 Diabetic Patients
Patric V. Prado, Chunhua Weng |
AMIA | 2 |
| 2014 | Associating co-authorship patterns with publications in high-impact journalsabstractOBJECTIVES: To develop a method for investigating co-authorship patterns and author team characteristics associated with the publications in high-impact journals through the integration of public MEDLINE data and institutional scientific profile data. METHODS: For all current researchers at Columbia University Medical Center, we extracted their publications from MEDLINE authored between years 2007 and 2011 and associated journal impact factors, along with author academic ranks and departmental affiliations obtained from Columbia University Scientific Profiles (CUSP). Chi-square tests were performed on co-authorship patterns, with Bonferroni correction for multiple comparisons, to identify team composition characteristics associated with publication impact factors. We also developed co-authorship networks for the 25 most prolific departments between years 2002 and 2011 and counted the internal and external authors, inter-connectivity, and centrality of each department. RESULTS: Papers with at least one author from a basic science department are significantly more likely to appear in high-impact journals than papers authored by those from clinical departments alone. Inclusion of at least one professor on the author list is strongly associated with publication in high-impact journals, as is inclusion of at least one research scientist. Departmental and disciplinary differences in the ratios of within- to outside-department collaboration and overall network cohesion are also observed. CONCLUSIONS: Enrichment of co-authorship patterns with author scientific profiles helps uncover associations between author team characteristics and appearance in high-impact journals. These results may offer implications for mentoring junior biomedical researchers to publish on high-impact journals, as well as for evaluating academic progress across disciplines in modern academic medical centers. Michael E. Bales, Daniel Dine, Jacqueline Merrill, Stephen B. Johnson, Suzanne Bakken, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2014 | From expert-derived user needs to user-perceived ease of use and usefulness: A two-phase mixed-methods evaluation framework
Mary Regina Boland, Alex Rusanov, Yat So, Carlos Lopez-Jimenez, Linda Busacca, Richard C. Steinman, Suzanne Bakken, J. Thomas Bigger, Chunhua Weng |
J. Biomed. Informatics | 9 |
| 2014 | Clustering clinical trials with similar eligibility criteria features
Tianyong Hao, Alex Rusanov, Mary Regina Boland, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2014 | Automatic generation of investigator bibliographies for institutional research networking systemsabstractOBJECTIVE: Publications are a key data source for investigator profiles and research networking systems. We developed ReCiter, an algorithm that automatically extracts bibliographies from PubMed using institutional information about the target investigators. METHODS: ReCiter executes a broad query against PubMed, groups the results into clusters that appear to constitute distinct author identities and selects the cluster that best matches the target investigator. Using information about investigators from one of our institutions, we compared ReCiter results to queries based on author name and institution and to citations extracted manually from the Scopus database. Five judges created a gold standard using citations of a random sample of 200 investigators. RESULTS: About half of the 10,471 potential investigators had no matching citations in PubMed, and about 45% had fewer than 70 citations. Interrater agreement (Fleiss' kappa) for the gold standard was 0.81. Scopus achieved the best recall (sensitivity) of 0.81, while name-based queries had 0.78 and ReCiter had 0.69. ReCiter attained the best precision (positive predictive value) of 0.93 while Scopus had 0.85 and name-based queries had 0.31. DISCUSSION: ReCiter accesses the most current citation data, uses limited computational resources and minimizes manual entry by investigators. Generation of bibliographies using named-based queries will not yield high accuracy. Proprietary databases can perform well but requite manual effort. Automated generation with higher recall is possible but requires additional knowledge about investigators. Stephen B. Johnson, Michael E. Bales, Daniel Dine, Suzanne Bakken, Paul J. Albert, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2013 | Design and Evaluation of a Bacterial Clinical Infectious Diseases Ontology
Claire L. Gordon, Stephanie Pouch, Lindsay G. Cowell, Mary Regina Boland, Heather L. Platt, Albert Goldfain, Chunhua Weng |
AMIA | 7 |
| 2013 | eTACTS: an Eligibility Tag Cloud-based Clinical Trial Search Engine
Riccardo Miotto, Chunhua Weng |
AMIA | 2 |
| 2013 | Sick Patients Have More Data: The Non-Random Completeness of Electronic Health Records
Nicole Gray Weiskopf, Alex Rusanov, Chunhua Weng |
AMIA | 3 |
| 2013 | Towards Computational Reuse of Clinical Research Eligibility Criteria with Collaboration across Academia, Industry, and Standardization Organizations
Chunhua Weng, Michael N. Cantor, Adel Taweel, Theodoros N. Arvanitis, Rebecca Daniels Kush |
AMIA | 1 |
| 2013 | Using Electronic Health Records To Assess Generalizability of Clinical Trials
Chunhua Weng, George Hripcsak, Yuan Zhang 0004, J. Thomas Bigger |
AMIA | 1 |
| 2013 | A centralized research data repository enhances retrospective outcomes research capacity: a case reportabstractThis paper describes our considerations and methods for implementing an open-source centralized research data repository (CRDR) and reports its impact on retrospective outcomes research capacity in the urology department at Columbia University. We performed retrospective pretest and post-test analyses of user acceptance, workflow efficiency, and publication quantity and quality (measured by journal impact factor) before and after the implementation. The CRDR transformed the research workflow and enabled a new research model. During the pre- and post-test periods, the department's average annual retrospective study publication rate was 11.5 and 25.6, respectively; the average publication impact score was 1.7 and 3.1, respectively. The new model was adopted by 62.5% (5/8) of the clinical scientists within the department. Additionally, four basic science researchers outside the department took advantage of the implemented model. The average proximate time required to complete a retrospective study decreased from 12 months before the implementation to <6 months after the implementation. Implementing a CRDR appears to be effective in enhancing the outcomes research capacity for one academic department. Gregory William Hruby, James McKiernan, Suzanne Bakken, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical researchabstractOBJECTIVE: To review the methods and dimensions of data quality assessment in the context of electronic health record (EHR) data reuse for research. MATERIALS AND METHODS: A review of the clinical research literature discussing data quality assessment methodology for EHR data was performed. Using an iterative process, the aspects of data quality being measured were abstracted and categorized, as well as the methods of assessment used. RESULTS: Five dimensions of data quality were identified, which are completeness, correctness, concordance, plausibility, and currency, and seven broad categories of data quality assessment methods: comparison with gold standards, data element agreement, data source agreement, distribution comparison, validity checks, log review, and element presence. DISCUSSION: Examination of the methods by which clinical researchers have investigated the quality and suitability of EHR data for research shows that there are fundamental features of data quality, which may be difficult to measure, as well as proxy dimensions. Researchers interested in the reuse of EHR data for clinical research are recommended to consider the adoption of a consistent taxonomy of EHR data quality, to remain aware of the task-dependence of data quality, to integrate work on data quality assessment from other fields, and to adopt systematic, empirically driven, statistically based methods of data quality assessment. CONCLUSION: There is currently little consistency or potential generalizability in the methods used to assess EHR data quality. If the reuse of EHR data for clinical research is to become accepted, researchers should adopt validated, systematic methods of EHR data quality assessment. Nicole Gray Weiskopf, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | A human-computer collaborative approach to identifying common data elements in clinical trial eligibility criteria
Jake Luo, Riccardo Miotto, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2013 | eTACTS: A method for dynamically filtering clinical trial search resultsabstractOBJECTIVE: Information overload is a significant problem facing online clinical trial searchers. We present eTACTS, a novel interactive retrieval framework using common eligibility tags to dynamically filter clinical trial search results. MATERIALS AND METHODS: eTACTS mines frequent eligibility tags from free-text clinical trial eligibility criteria and uses these tags for trial indexing. After an initial search, eTACTS presents to the user a tag cloud representing the current results. When the user selects a tag, eTACTS retains only those trials containing that tag in their eligibility criteria and generates a new cloud based on tag frequency and co-occurrences in the remaining trials. The user can then select a new tag or unselect a previous tag. The process iterates until a manageable number of trials is returned. We evaluated eTACTS in terms of filtering efficiency, diversity of the search results, and user eligibility to the filtered trials using both qualitative and quantitative methods. RESULTS: eTACTS (1) rapidly reduced search results from over a thousand trials to ten; (2) highlighted trials that are generally not top-ranked by conventional search engines; and (3) retrieved a greater number of suitable trials than existing search engines. DISCUSSION: eTACTS enables intuitive clinical trial searches by indexing eligibility criteria with effective tags. User evaluation was limited to one case study and a small group of evaluators due to the long duration of the experiment. Although a larger-scale evaluation could be conducted, this feasibility study demonstrated significant advantages of eTACTS over existing clinical trial search engines. CONCLUSION: A dynamic eligibility tag cloud can potentially enhance state-of-the-art clinical trial search engines by allowing intuitive and efficient filtering of the search result space. Riccardo Miotto, Silis Y. Jiang, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2013 | Unsupervised mining of frequent tags for clinical eligibility text indexing
Riccardo Miotto, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2013 | Defining and measuring completeness of electronic health records for secondary useabstractWe demonstrate the importance of explicit definitions of electronic health record (EHR) data completeness and how different conceptualizations of completeness may impact findings from EHR-derived datasets. This study has important repercussions for researchers and clinicians engaged in the secondary use of EHR data. We describe four prototypical definitions of EHR completeness: documentation, breadth, density, and predictive completeness. Each definition dictates a different approach to the measurement of completeness. These measures were applied to representative data from NewYork-Presbyterian Hospital's clinical data warehouse. We found that according to any definition, the number of complete records in our clinical database is far lower than the nominal total. The proportion that meets criteria for completeness is heavily dependent on the definition of completeness used, and the different definitions generate different subsets of records. We conclude that the concept of completeness in EHR is contextual. We urge data consumers to be explicit in how they define a complete record and transparent about the limitations of their data. Nicole Gray Weiskopf, George Hripcsak, Sushmita Swaminathan, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2013 | An Integrated Model for Patient Care and Clinical Trials (IMPACT) to support clinical research visit scheduling workflow for future learning health systemsabstractWe describe a clinical research visit scheduling system that can potentially coordinate clinical research visits with patient care visits and increase efficiency at clinical sites where clinical and research activities occur simultaneously. Participatory Design methods were applied to support requirements engineering and to create this software called Integrated Model for Patient Care and Clinical Trials (IMPACT). Using a multi-user constraint satisfaction and resource optimization algorithm, IMPACT automatically synthesizes temporal availability of various research resources and recommends the optimal dates and times for pending research visits. We conducted scenario-based evaluations with 10 clinical research coordinators (CRCs) from diverse clinical research settings to assess the usefulness, feasibility, and user acceptance of IMPACT. We obtained qualitative feedback using semi-structured interviews with the CRCs. Most CRCs acknowledged the usefulness of IMPACT features. Support for collaboration within research teams and interoperability with electronic health records and clinical trial management systems were highly requested features. Overall, IMPACT received satisfactory user acceptance and proves to be potentially useful for a variety of clinical research settings. Our future work includes comparing the effectiveness of IMPACT with that of existing scheduling solutions on the market and conducting field tests to formally assess user adoption. Chunhua Weng, Solomon Berhe, Mary Regina Boland, Junfeng Gao, Gregory William Hruby, Richard C. Steinman, Carlos Lopez-Jimenez, Linda Busacca, George Hripcsak, Suzanne Bakken, J. Thomas Bigger |
J. Biomed. Informatics | 1 |
| 2012 | Establishing Data Governance for Comparative Effectiveness Research: Experience of the Washington Heights Inwood Informatics Infrastructure for Comparative Effectiveness Research (WICER)
Suzanne Bakken, J. Thomas Bigger, Bernadette Boden-Albala, Penny Feldman, Kathleen Gallagher, Peter D. Stetson, Chunhua Weng, Adam B. Wilcox |
AMIA | 7 |
| 2012 | Analysis of Query Negotiation between a Researcher and a Query Expert
Gregory William Hruby, Adam B. Wilcox, Chunhua Weng |
AMIA | 3 |
| 2012 | An ontology-based Clinical Phenotyping Framework For eMERGE Network Algorithms
Nanfang Xu, Chunhua Weng |
AMIA | 2 |
| 2012 | Clinical research informatics: a conceptual perspectiveabstractClinical research informatics is the rapidly evolving sub-discipline within biomedical informatics that focuses on developing new informatics theories, tools, and solutions to accelerate the full translational continuum: basic research to clinical trials (T1), clinical trials to academic health center practice (T2), diffusion and implementation to community practice (T3), and 'real world' outcomes (T4). We present a conceptual model based on an informatics-enabled clinical research workflow, integration across heterogeneous data sources, and core informatics tools and platforms. We use this conceptual model to highlight 18 new articles in the JAMIA special issue on clinical research informatics. Michael G. Kahn, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Using EHRs to integrate research with patient care: promises and challengesabstractClinical research is the foundation for advancing the practice of medicine. However, the lack of seamless integration between clinical research and patient care workflow impedes recruitment efficiency, escalates research costs, and hence threatens the entire clinical research enterprise. Increased use of electronic health records (EHRs) holds promise for facilitating this integration but must surmount regulatory obstacles. Among the unintended consequences of current research oversight are barriers to accessing patient information for prescreening and recruitment, coordinating scheduling of clinical and research visits, and reconciling information about clinical and research drugs. We conclude that the EHR alone cannot overcome barriers in conducting clinical trials and comparative effectiveness research. Patient privacy and human subject protection policies should be clarified at the local level to exploit optimally the full potential of EHRs, while continuing to ensure participant safety. Increased alignment of policies that regulate the clinical and research use of EHRs could help fulfill the vision of more efficiently obtaining clinical research evidence to improve human health. Chunhua Weng, Paul Appelbaum, George Hripcsak, Ian M. Kronish, Linda Busacca, Karina W. Davidson, J. Thomas Bigger |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | EliXR: an approach to eligibility criteria extraction and representationabstractOBJECTIVE: To develop a semantic representation for clinical research eligibility criteria to automate semistructured information extraction from eligibility criteria text. MATERIALS AND METHODS: An analysis pipeline called eligibility criteria extraction and representation (EliXR) was developed that integrates syntactic parsing and tree pattern mining to discover common semantic patterns in 1000 eligibility criteria randomly selected from http://ClinicalTrials.gov. The semantic patterns were aggregated and enriched with unified medical language systems semantic knowledge to form a semantic representation for clinical research eligibility criteria. RESULTS: The authors arrived at 175 semantic patterns, which form 12 semantic role labels connected by their frequent semantic relations in a semantic network. EVALUATION: Three raters independently annotated all the sentence segments (N=396) for 79 test eligibility criteria using the 12 top-level semantic role labels. Eight-six per cent (339) of the sentence segments were unanimously labelled correctly and 13.8% (55) were correctly labelled by two raters. The Fleiss' κ was 0.88, indicating a nearly perfect interrater agreement. CONCLUSION: This study present a semi-automated data-driven approach to developing a semantic network that aligns well with the top-level information structure in clinical research eligibility criteria text and demonstrates the feasibility of using the resulting semantic role labels to generate semistructured eligibility criteria with nearly perfect interrater reliability. Chunhua Weng, Xiaoying Wu 0001, Jake Luo, Mary Regina Boland, Dimitri Theodoratos, Stephen B. Johnson |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Dynamic categorization of clinical research eligibility criteria by hierarchical clustering
Jake Luo, Meliha Yetisgen, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2011 | Combining PubMed knowledge and EHR data to develop a weighted bayesian network for pancreatic cancer prediction
Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2010 | Formal representation of eligibility criteria: A literature review
Chunhua Weng, Samson W. Tu, Ida Sim, Rachel L. Richesson |
J. Biomed. Informatics | 1 |
| 2009 | Case Report: Electronic Screening Improves Efficiency in Clinical Trial RecruitmentabstractThis study evaluated the performance of an electronic screening (E-screening) method and used it to recruit patients for the NIH sponsored ACCORD trial. Out of the 193 E-screened patients, 125 met the age criterion ("age>or=40"). For all of these 125 patients, the performance of E-screening was compared with investigator review. E-screening achieved a negative predictive accuracy of 100% (95% CI: 98-100%), a positive predictive accuracy of 13% (95% CI: 6-13%), a sensitivity of 100% (95% CI: 45-100%), and a specificity of 84% (95% CI: 82-84%). The method maximized the use of a patient database query (i.e., excluded ineligible patients with a 100% accuracy and automatically assembled patient information to facilitate manual review of only patients who were classified as "potentially eligible" by E-screening) and significantly reduced the screening burden associated with the ACCORD trial. Samir R. Thadani, Chunhua Weng, J. Thomas Bigger, John F. Ennever, David Wajngurt |
J. Am. Medical Informatics Assoc. | 2 |
| 2009 | A review of auditing methods applied to the content of controlled biomedical terminologies
Jungwei Fan 0001, David M. Baorto, Chunhua Weng, James J. Cimino |
J. Biomed. Informatics | 4 |
| 2008 | Comparing ICD9-Encoded Diagnoses and NLP-Processed Discharge Summaries for Clinical Trials Pre-Screening: A Case Study
Herbert S. Chase, Chintan Patel, Carol Friedman, Chunhua Weng |
AMIA | 5 |
| 2008 | Understanding Interdisciplinary Health Sciences Collaborations: A Campus-Wide Survey of Obesity Experts
Chunhua Weng, Dympna Gallagher, Michael E. Bales, Suzanne Bakken, Henry N. Ginsberg |
AMIA | 1 |
| 2007 | User-centered semantic harmonization: A case study
Chunhua Weng, John H. Gennari, Douglas B. Fridsma |
J. Biomed. Informatics | 1 |
| 2006 | A Call for Collaborative Semantic Harmonization
Chunhua Weng, Douglas B. Fridsma |
AMIA | 1 |
| 2005 | Why It Is Hard to Support Group Work in Distributed Healthcare Organizations: Empirical Knowledge of the Social-Technical Gap
Chunhua Weng |
AMIA | 1 |
| 2004 | The multiple views of inter-organizational authoringabstractCollaborative authoring is a common workplace task. Yet, despite improvements in word processors, communication software, and file sharing, many problems continue to plague co-authors. We conducted a qualitative study in a setting where participants are loosely connected, physically separated, and work together over a period of 4-9 months to author a complex technical document-a clinical trial protocol. Our study differs from most prior work in that the collaboration is longer-lived, and that the collaborators do not share equivalent status, background, nor domains of expertise. Our data demonstrates that the participants do not share the same view or representation of the authoring process, even though it has a long organizational history. Nonetheless, the participants can still coordinate their activity while maintaining only partially consistent representations of what they are doing. We contend that partial consistency in the participants' concept of the collaborative process is a feature for their asynchronous collaboration at a distance. Based on our findings we suggest a number of improvements for both tools and tool usage that have direct impact on support for collaborative authoring. David W. McDonald, Chunhua Weng, John H. Gennari |
CSCW | 2 |
| 2004 | Asynchronous collaborative writing through annotationsabstractAnnotation is central to iterative reviewing and revising activities in asynchronous collaborative writing. Currently most digital annotation models and systems assume static context information and provide far less functionality than physical annotations. We extend prior annotation research by Marshall and Cadiz and design an activity-oriented annotation model to mimic the rich functionality of physical annotations for an enhanced collaborative writing process. In this model, we define an annotation life cycle and support annotation version control. We implement a collaborative writing system that supports improved in-situ communication and cross-role feedback based on our annotation model. Chunhua Weng, John H. Gennari |
CSCW | 1 |
| 2003 | Scenario-based Participatory Design of A Collaborative Clinical Trial Protocol Authoring System
Chunhua Weng, John H. Gennari, David W. McDonald |
AMIA | 1 |
| 2002 | Temporal knowledge representation for scheduling tasks in clinical trial protocols
Chunhua Weng, Michael G. Kahn, John H. Gennari |
AMIA | 1 |