EDBT 2026 Demo / reviewers in the wild / expert
Casey N. Ta
dblp:212/2584
· DBLP profile ↗
18ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0002-4679-805XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mini-mental status examination phenotyping for Alzheimer's disease patients using both structured and narrative electronic health record featuresabstractOBJECTIVE: This study aims to automate the prediction of Mini-Mental State Examination (MMSE) scores, a widely adopted standard for cognitive assessment in patients with Alzheimer's disease, using natural language processing (NLP) and machine learning (ML) on structured and unstructured EHR data. MATERIALS AND METHODS: We extracted demographic data, diagnoses, medications, and unstructured clinical visit notes from the EHRs. We used Latent Dirichlet Allocation (LDA) for topic modeling and Term-Frequency Inverse Document Frequency (TF-IDF) for n-grams. In addition, we extracted meta-features such as age, ethnicity, and race. Model training and evaluation employed eXtreme Gradient Boosting (XGBoost), Stochastic Gradient Descent Regressor (SGDRegressor), and Multi-Layer Perceptron (MLP). RESULTS: We analyzed 1654 clinical visit notes collected between September 2019 and June 2023 for 1000 Alzheimer's disease patients. The average MMSE score was 20, with patients averaging 76.4 years old, 54.7% female, and 54.7% identifying as White. The best-performing model (ie, lowest root mean squared error (RMSE)) is MLP, which achieved an RMSE of 5.53 on the validation set using n-grams, indicating superior prediction performance over other models and feature sets. The RMSE on the test set was 5.85. DISCUSSION: This study developed a ML method to predict MMSE scores from unstructured clinical notes, demonstrating the feasibility of utilizing NLP to support cognitive assessment. Future work should focus on refining the model and evaluating its clinical relevance across diverse settings. CONCLUSION: We contributed a model for automating MMSE estimation using EHR features, potentially transforming cognitive assessment for Alzheimer's patients and paving the way for more informed clinical decisions and cohort identification. Betina Ross S. Idnay, Fangyi Chen, Casey N. Ta, Matthew W. Schelke, Karen Marder, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | A method for characterizing disease progression from acute kidney injury to chronic kidney disease
Yilu Fang, Jordan G. Nestor, Casey N. Ta, Jerard Kneifati-Hayek, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2025 | Scalable scientific interest profiling using large language models
Yilun Liang, Edward Sun, Betina Ross S. Idnay, Yilu Fang, Fangyi Chen, Casey N. Ta, Yifan Peng 0002, Chunhua Weng |
J. Biomed. Informatics | 7 |
| 2024 | Criteria2Query 3.0: Leveraging generative large language models for clinical trial eligibility query generation
Jimyung Park, Yilu Fang, Casey N. Ta, Betina Ross S. Idnay, Fangyi Chen, Rebecca Shyu, Emily R. Gordon, Matthew E. Spotnitz, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2024 | Converting OMOP CDM to phenopackets: A model alignment and patient data representation evaluation
Kayla Schiffer, Cong Liu 0020, Tiffany Callahan, Casey N. Ta, Jordan G. Nestor, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2023 | EvidenceMap: a three-level knowledge representation for medical evidence computation and comprehensionabstractOBJECTIVE: To develop a computable representation for medical evidence and to contribute a gold standard dataset of annotated randomized controlled trial (RCT) abstracts, along with a natural language processing (NLP) pipeline for transforming free-text RCT evidence in PubMed into the structured representation. MATERIALS AND METHODS: Our representation, EvidenceMap, consists of 3 levels of abstraction: Medical Evidence Entity, Proposition and Map, to represent the hierarchical structure of medical evidence composition. Randomly selected RCT abstracts were annotated following EvidenceMap based on the consensus of 2 independent annotators to train an NLP pipeline. Via a user study, we measured how the EvidenceMap improved evidence comprehension and analyzed its representative capacity by comparing the evidence annotation with EvidenceMap representation and without following any specific guidelines. RESULTS: Two corpora including 229 disease-agnostic and 80 COVID-19 RCT abstracts were annotated, yielding 12 725 entities and 1602 propositions. EvidenceMap saves users 51.9% of the time compared to reading raw-text abstracts. Most evidence elements identified during the freeform annotation were successfully represented by EvidenceMap, and users gave the enrollment, study design, and study Results sections mean 5-scale Likert ratings of 4.85, 4.70, and 4.20, respectively. The end-to-end evaluations of the pipeline show that the evidence proposition formulation achieves F1 scores of 0.84 and 0.86 in the adjusted random index score. CONCLUSIONS: EvidenceMap extends the participant, intervention, comparator, and outcome framework into 3 levels of abstraction for transforming free-text evidence from the clinical literature into a computable structure. It can be used as an interoperable format for better evidence retrieval and synthesis and an interpretable representation to efficiently comprehend RCT findings. Tian Kang, Yingcheng Sun, Jae Hyun Kim, Casey N. Ta, Adler J. Perotte, Kayla Schiffer, Mutong Wu, Nour Fahmy, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Clinical and temporal characterization of COVID-19 subgroups using patient vector embeddings of electronic health recordsabstractOBJECTIVE: To identify and characterize clinical subgroups of hospitalized Coronavirus Disease 2019 (COVID-19) patients. MATERIALS AND METHODS: Electronic health records of hospitalized COVID-19 patients at NewYork-Presbyterian/Columbia University Irving Medical Center were temporally sequenced and transformed into patient vector representations using Paragraph Vector models. K-means clustering was performed to identify subgroups. RESULTS: A diverse cohort of 11 313 patients with COVID-19 and hospitalizations between March 2, 2020 and December 1, 2021 were identified; median [IQR] age: 61.2 [40.3-74.3]; 51.5% female. Twenty subgroups of hospitalized COVID-19 patients, labeled by increasing severity, were characterized by their demographics, conditions, outcomes, and severity (mild-moderate/severe/critical). Subgroup temporal patterns were characterized by the durations in each subgroup, transitions between subgroups, and the complete paths throughout the course of hospitalization. DISCUSSION: Several subgroups had mild-moderate severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections but were hospitalized for underlying conditions (pregnancy, cardiovascular disease [CVD], etc.). Subgroup 7 included solid organ transplant recipients who mostly developed mild-moderate or severe disease. Subgroup 9 had a history of type-2 diabetes, kidney and CVD, and suffered the highest rates of heart failure (45.2%) and end-stage renal disease (80.6%). Subgroup 13 was the oldest (median: 82.7 years) and had mixed severity but high mortality (33.3%). Subgroup 17 had critical disease and the highest mortality (64.6%), with age (median: 68.1 years) being the only notable risk factor. Subgroups 18-20 had critical disease with high complication rates and long hospitalizations (median: 40+ days). All subgroups are detailed in the full text. A chord diagram depicts the most common transitions, and paths with the highest prevalence, longest hospitalizations, lowest and highest mortalities are presented. Understanding these subgroups and their pathways may aid clinicians in their decisions for better management and earlier intervention for patients. Casey N. Ta, Jason Zucker 0001, Po-Hsiang Chiu, Yilu Fang, Karthik Natarajan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | A data-driven approach to optimizing clinical study eligibility criteria
Yilu Fang, Hao Liu 0054, Betina Ross S. Idnay, Casey N. Ta, Karen Marder, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2022 | Leveraging electronic health record data for clinical trial planning by assessing eligibility criteria's impact on patient count and safety
James R. Rogers, Jovana Pavisic, Casey N. Ta, Cong Liu 0020, Ali Soroush, Ying Kuen Cheung, George Hripcsak, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2021 | UMLS-based data augmentation for natural language processing of clinical research literatureabstractOBJECTIVE: The study sought to develop and evaluate a knowledge-based data augmentation method to improve the performance of deep learning models for biomedical natural language processing by overcoming training data scarcity. MATERIALS AND METHODS: We extended the easy data augmentation (EDA) method for biomedical named entity recognition (NER) by incorporating the Unified Medical Language System (UMLS) knowledge and called this method UMLS-EDA. We designed experiments to systematically evaluate the effect of UMLS-EDA on popular deep learning architectures for both NER and classification. We also compared UMLS-EDA to BERT. RESULTS: UMLS-EDA enables substantial improvement for NER tasks from the original long short-term memory conditional random fields (LSTM-CRF) model (micro-F1 score: +5%, + 17%, and +15%), helps the LSTM-CRF model (micro-F1 score: 0.66) outperform LSTM-CRF with transfer learning by BERT (0.63), and improves the performance of the state-of-the-art sentence classification model. The largest gain on micro-F1 score is 9%, from 0.75 to 0.84, better than classifiers with BERT pretraining (0.82). CONCLUSIONS: This study presents a UMLS-based data augmentation method, UMLS-EDA. It is effective at improving deep learning models for both NER and sentence classification, and contributes original insights for designing new, superior deep learning approaches for low-resource biomedical domains. Tian Kang, Adler J. Perotte, Youlan Tang, Casey N. Ta, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Towards clinical data-driven eligibility criteria optimization for interventional COVID-19 clinical trialsabstractOBJECTIVE: This research aims to evaluate the impact of eligibility criteria on recruitment and observable clinical outcomes of COVID-19 clinical trials using electronic health record (EHR) data. MATERIALS AND METHODS: On June 18, 2020, we identified frequently used eligibility criteria from all the interventional COVID-19 trials in ClinicalTrials.gov (n = 288), including age, pregnancy, oxygen saturation, alanine/aspartate aminotransferase, platelets, and estimated glomerular filtration rate. We applied the frequently used criteria to the EHR data of COVID-19 patients in Columbia University Irving Medical Center (CUIMC) (March 2020-June 2020) and evaluated their impact on patient accrual and the occurrence of a composite endpoint of mechanical ventilation, tracheostomy, and in-hospital death. RESULTS: There were 3251 patients diagnosed with COVID-19 from the CUIMC EHR included in the analysis. The median follow-up period was 10 days (interquartile range 4-28 days). The composite events occurred in 18.1% (n = 587) of the COVID-19 cohort during the follow-up. In a hypothetical trial with common eligibility criteria, 33.6% (690/2051) were eligible among patients with evaluable data and 22.2% (153/690) had the composite event. DISCUSSION: By adjusting the thresholds of common eligibility criteria based on the characteristics of COVID-19 patients, we could observe more composite events from fewer patients. CONCLUSIONS: This research demonstrated the potential of using the EHR data of COVID-19 patients to inform the selection of eligibility criteria and their thresholds, supporting data-driven optimization of participant selection towards improved statistical power of COVID-19 trials. Jae Hyun Kim, Casey N. Ta, Cong Liu 0020, Cynthia Sung 0002, Alex M. Butler, Latoya A. Stewart, Lyudmila Ena, James R. Rogers, Anna Ostropolets, Patrick B. Ryan, Hao Liu 0054, Shing M. Lee, Mitchell S. V. Elkind, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | From clinical trials to clinical practice: How long are drugs tested and then used by patients?abstractOBJECTIVE: Evidence is scarce regarding the safety of long-term drug use, especially for drugs treating chronic diseases. To bridge this knowledge gap, this research investigated the differences in drug exposure between clinical trials and clinical practice. MATERIALS AND METHODS: We extracted drug follow-up times from clinical trials in ClinicalTrials.gov and compared the difference between clinical trials and real-world usage data for 914 drugs taken by 96 645 927 patients. RESULTS: A total of 17.5% of drugs had longer median exposure in practice than in trials, 6% of patients had extended exposure to at least 1 drug, and drugs treating nervous system disorders and cardiovascular diseases were the most common among drugs with high rates of extended exposure. CONCLUSIONS: For most of patients, the drug use length is shorter than the tested length in clinical trials. Still, a remarkable number of patients experienced extended drug exposure, particularly for drugs treating nervous system disorders or cardiovascular disorders. Chi Yuan, Patrick B. Ryan, Casey N. Ta, Jae Hyun Kim, Ziran Li, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | DQueST: dynamic questionnaire for search of clinical trialsabstractOBJECTIVE: Information overload remains a challenge for patients seeking clinical trials. We present a novel system (DQueST) that reduces information overload for trial seekers using dynamic questionnaires. MATERIALS AND METHODS: DQueST first performs information extraction and criteria library curation. DQueST transforms criteria narratives in the ClinicalTrials.gov repository into a structured format, normalizes clinical entities using standard concepts, clusters related criteria, and stores the resulting curated library. DQueST then implements a real-time dynamic question generation algorithm. During user interaction, the initial search is similar to a standard search engine, and then DQueST performs real-time dynamic question generation to select criteria from the library 1 at a time by maximizing its relevance score that reflects its ability to rule out ineligible trials. DQueST dynamically updates the remaining trial set by removing ineligible trials based on user responses to corresponding questions. The process iterates until users decide to stop and begin manually reviewing the remaining trials. RESULTS: In simulation experiments initiated by 10 diseases, DQueST reduced information overload by filtering out 60%-80% of initial trials after 50 questions. Reviewing the generated questions against previous answers, on average, 79.7% of the questions were relevant to the queried conditions. By examining the eligibility of random samples of trials ruled out by DQueST, we estimate the accuracy of the filtering procedure is 63.7%. In a study using 5 mock patient profiles, DQueST on average retrieved trials with a 1.465 times higher density of eligible trials than an existing search engine. In a patient-centered usability evaluation, patients found DQueST useful, easy to use, and returning relevant results. CONCLUSION: DQueST contributes a novel framework for transforming free-text eligibility criteria to questions and filtering out clinical trials based on user answers to questions dynamically. It promises to augment keyword-based methods to improve clinical trial search. Cong Liu 0020, Chi Yuan, Alex M. Butler, Richard D. Carvajal, Ziran Ryan Li, Casey N. Ta, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Criteria2Query: a natural language interface to clinical databases for cohort definitionabstractOBJECTIVE: Cohort definition is a bottleneck for conducting clinical research and depends on subjective decisions by domain experts. Data-driven cohort definition is appealing but requires substantial knowledge of terminologies and clinical data models. Criteria2Query is a natural language interface that facilitates human-computer collaboration for cohort definition and execution using clinical databases. MATERIALS AND METHODS: Criteria2Query uses a hybrid information extraction pipeline combining machine learning and rule-based methods to systematically parse eligibility criteria text, transforms it first into a structured criteria representation and next into sharable and executable clinical data queries represented as SQL queries conforming to the OMOP Common Data Model. Users can interactively review, refine, and execute queries in the ATLAS web application. To test effectiveness, we evaluated 125 criteria across different disease domains from ClinicalTrials.gov and 52 user-entered criteria. We evaluated F1 score and accuracy against 2 domain experts and calculated the average computation time for fully automated query formulation. We conducted an anonymous survey evaluating usability. RESULTS: Criteria2Query achieved 0.795 and 0.805 F1 score for entity recognition and relation extraction, respectively. Accuracies for negation detection, logic detection, entity normalization, and attribute normalization were 0.984, 0.864, 0.514 and 0.793, respectively. Fully automatic query formulation took 1.22 seconds/criterion. More than 80% (11+ of 13) of users would use Criteria2Query in their future cohort definition tasks. CONCLUSIONS: We contribute a novel natural language interface to clinical databases. It is open source and supports fully automated and interactive modes for autonomous data-driven cohort definition by researchers with minimal human effort. We demonstrate its promising user friendliness and usability. Chi Yuan, Patrick B. Ryan, Casey N. Ta, Yixuan Guo, Ziran Li, Jill Hardin, Rupa Makadia, Ning Shang 0004, Tian Kang, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Sex, obesity, diabetes, and exposure to particulate matter among patients with severe asthma: Scientific insights from a comparative analysis of open clinical data sources during a five-day hackathonabstractThis special communication describes activities, products, and lessons learned from a recent hackathon that was funded by the National Center for Advancing Translational Sciences via the Biomedical Data Translator program ('Translator'). Specifically, Translator team members self-organized and worked together to conceptualize and execute, over a five-day period, a multi-institutional clinical research study that aimed to examine, using open clinical data sources, relationships between sex, obesity, diabetes, and exposure to airborne fine particulate matter among patients with severe asthma. The goal was to develop a proof of concept that this new model of collaboration and data sharing could effectively produce meaningful scientific results and generate new scientific hypotheses. Three Translator Clinical Knowledge Sources, each of which provides open access (via Application Programming Interfaces) to data derived from the electronic health record systems of major academic institutions, served as the source of study data. Jupyter Python notebooks, shared in GitHub repositories, were used to call the knowledge sources and analyze and integrate the results. The results replicated established or suspected relationships between sex, obesity, diabetes, exposure to airborne fine particulate matter, and severe asthma. In addition, the results demonstrated specific differences across the three Translator Clinical Knowledge Sources, suggesting cohort- and/or environment-specific factors related to the services themselves or the catchment area from which each service derives patient data. Collectively, this special communication demonstrates the power and utility of intense, team-oriented hackathons and offers general technical, organizational, and scientific lessons learned. Karamarie Fecho, Stanley C. Ahalt, Saravanan Arunachalam, James Champion, Christopher G. Chute, Sarah Davis, Kenneth Gersing, Gwênlyn Glusman, Jennifer Hadlock, Jewel Lee, Emily R. Pfaff, Max Robinson, Eric Sid, Casey N. Ta, Hao Xu 0006, Richard L. Zhu, Qian Zhu 0003, David B. Peden |
J. Biomed. Informatics | 14 |
| 2019 | Ensembles of natural language processing systems for portable phenotyping solutions
Cong Liu 0020, Casey N. Ta, James R. Rogers, Ziran Li, Alex M. Butler, Ning Shang 0004, Fabricio Sampaio Peres Kury, Liwei Wang 0010, Feichen Shen, Lyudmila Ena, Carol Friedman, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2019 | Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng |
J. Biomed. Informatics | 4 |
| 2018 | Extended Lifetime In Vivo Pulse Stimulated Ultrasound ImagingabstractAn on-demand long-lived ultrasound contrast agent that can be activated with single pulse stimulated imaging (SPSI) has been developed using hard shell liquid perfluoropentane filled silica 500-nm nanoparticles for tumor ultrasound imaging. SPSI was tested on LnCAP prostate tumor models in mice; tumor localization was observed after intravenous (IV) injection of the contrast agent. Consistent with enhanced permeability and retention, the silica nanoparticles displayed an extended imaging lifetime of 3.3±1 days (mean±standard deviation). With added tumor specific folate functionalization, the useful lifetime was extended to 12 ± 2 days; in contrast to ligand-based tumor targeting, the effect of the ligands in this application is enhanced nanoparticle retention by the tumor. This paper demonstrates for the first time that IV injected functionalized silica contrast agents can be imaged with an in vivo lifetime ~500 times longer than current microbubble-based contrast agents. Such functionalized long-lived contrast agents may lead to new applications in tumor monitoring and therapy. James Wang, Christopher V. Barback, Casey N. Ta, Joi Weeks, Natalie Gude, Robert F. Mattrey, Sarah L. Blair, William C. Trogler, Hotaik Lee, Andrew C. Kummel |
IEEE Trans. Medical Imaging | 3 |