Betina Ross S. Idnay

dblp:288/9570 · also Betina Idnay · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-4318-5987ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 13 since 2021
YearPublicationVenuePosition
2025 Mini-mental status examination phenotyping for Alzheimer's disease patients using both structured and narrative electronic health record features
abstract
OBJECTIVE: This study aims to automate the prediction of Mini-Mental State Examination (MMSE) scores, a widely adopted standard for cognitive assessment in patients with Alzheimer's disease, using natural language processing (NLP) and machine learning (ML) on structured and unstructured EHR data. MATERIALS AND METHODS: We extracted demographic data, diagnoses, medications, and unstructured clinical visit notes from the EHRs. We used Latent Dirichlet Allocation (LDA) for topic modeling and Term-Frequency Inverse Document Frequency (TF-IDF) for n-grams. In addition, we extracted meta-features such as age, ethnicity, and race. Model training and evaluation employed eXtreme Gradient Boosting (XGBoost), Stochastic Gradient Descent Regressor (SGDRegressor), and Multi-Layer Perceptron (MLP). RESULTS: We analyzed 1654 clinical visit notes collected between September 2019 and June 2023 for 1000 Alzheimer's disease patients. The average MMSE score was 20, with patients averaging 76.4 years old, 54.7% female, and 54.7% identifying as White. The best-performing model (ie, lowest root mean squared error (RMSE)) is MLP, which achieved an RMSE of 5.53 on the validation set using n-grams, indicating superior prediction performance over other models and feature sets. The RMSE on the test set was 5.85. DISCUSSION: This study developed a ML method to predict MMSE scores from unstructured clinical notes, demonstrating the feasibility of utilizing NLP to support cognitive assessment. Future work should focus on refining the model and evaluating its clinical relevance across diverse settings. CONCLUSION: We contributed a model for automating MMSE estimation using EHR features, potentially transforming cognitive assessment for Alzheimer's patients and paving the way for more informed clinical decisions and cohort identification.
Betina Ross S. Idnay, Fangyi Chen, Casey N. Ta, Matthew W. Schelke, Karen Marder, Chunhua Weng
J. Am. Medical Informatics Assoc.1
2025 Scalable scientific interest profiling using large language models
Yilun Liang, Edward Sun, Betina Ross S. Idnay, Yilu Fang, Fangyi Chen, Casey N. Ta, Yifan Peng 0002, Chunhua Weng
J. Biomed. Informatics4
2024 Sociotechnical feasibility of natural language processing-driven tools in clinical trial eligibility prescreening for Alzheimer's disease and related dementias
abstract
BACKGROUND: Alzheimer's disease and related dementias (ADRD) affect over 55 million globally. Current clinical trials suffer from low recruitment rates, a challenge potentially addressable via natural language processing (NLP) technologies for researchers to effectively identify eligible clinical trial participants. OBJECTIVE: This study investigates the sociotechnical feasibility of NLP-driven tools for ADRD research prescreening and analyzes the tools' cognitive complexity's effect on usability to identify cognitive support strategies. METHODS: A randomized experiment was conducted with 60 clinical research staff using three prescreening tools (Criteria2Query, Informatics for Integrating Biology and the Bedside [i2b2], and Leaf). Cognitive task analysis was employed to analyze the usability of each tool using the Health Information Technology Usability Evaluation Scale. Data analysis involved calculating descriptive statistics, interrater agreement via intraclass correlation coefficient, cognitive complexity, and Generalized Estimating Equations models. RESULTS: Leaf scored highest for usability followed by Criteria2Query and i2b2. Cognitive complexity was found to be affected by age, computer literacy, and number of criteria, but was not significantly associated with usability. DISCUSSION: Adopting NLP for ADRD prescreening demands careful task delegation, comprehensive training, precise translation of eligibility criteria, and increased research accessibility. The study highlights the relevance of these factors in enhancing NLP-driven tools' usability and efficacy in clinical research prescreening. CONCLUSION: User-modifiable NLP-driven prescreening tools were favorably received, with system type, evaluation sequence, and user's computer literacy influencing usability more than cognitive complexity. The study emphasizes NLP's potential in improving recruitment for clinical trials, endorsing a mixed-methods approach for future system evaluation and enhancements.
Betina Ross S. Idnay, Jianfang Liu, Yilu Fang, Alex Hernandez, Shivani Kaw, Alicia Etwaru, Janeth Juarez Padilla, Sergio Ozoria Ramirez, Karen Marder, Chunhua Weng, Rebecca Schnall
J. Am. Medical Informatics Assoc.1
2024 Promoting equity in clinical research: The role of social determinants of health
Betina Ross S. Idnay, Yilu Fang, Edward Stanley, Brenda Ruotolo, Wendy K. Chung, Karen Marder, Chunhua Weng
J. Biomed. Informatics1
2024 Criteria2Query 3.0: Leveraging generative large language models for clinical trial eligibility query generation
Jimyung Park, Yilu Fang, Casey N. Ta, Betina Ross S. Idnay, Fangyi Chen, Rebecca Shyu, Emily R. Gordon, Matthew E. Spotnitz, Chunhua Weng
J. Biomed. Informatics5
2023 The suitability of UMLS and SNOMED-CT for encoding outcome concepts
abstract
OBJECTIVE: Outcomes are important clinical study information. Despite progress in automated extraction of PICO (Population, Intervention, Comparison, and Outcome) entities from PubMed, rarely are these entities encoded by standard terminology to achieve semantic interoperability. This study aims to evaluate the suitability of the Unified Medical Language System (UMLS) and SNOMED-CT in encoding outcome concepts in randomized controlled trial (RCT) abstracts. MATERIALS AND METHODS: We iteratively developed and validated an outcome annotation guideline and manually annotated clinically significant outcome entities in the Results and Conclusions sections of 500 randomly selected RCT abstracts on PubMed. The extracted outcomes were fully, partially, or not mapped to the UMLS via MetaMap based on established heuristics. Manual UMLS browser search was performed for select unmapped outcome entities to further differentiate between UMLS and MetaMap errors. RESULTS: Only 44% of 2617 outcome concepts were fully covered in the UMLS, among which 67% were complex concepts that required the combination of 2 or more UMLS concepts to represent them. SNOMED-CT was present as a source in 61% of the fully mapped outcomes. DISCUSSION: Domains such as Metabolism and Nutrition, and Infections and Infectious Diseases need expanded outcome concept coverage in the UMLS and MetaMap. Future work is warranted to similarly assess the terminology coverage for P, I, C entities. CONCLUSION: Computational representation of clinical outcomes is important for clinical evidence extraction and appraisal and yet faces challenges from the inherent complexity and lack of coverage of these concepts in UMLS and SNOMED-CT, as demonstrated in this study.
Abigail M. Newbury, Hao Liu 0054, Betina Ross S. Idnay, Chunhua Weng
J. Am. Medical Informatics Assoc.3
2023 A data-driven approach to optimizing clinical study eligibility criteria
Yilu Fang, Hao Liu 0054, Betina Ross S. Idnay, Casey N. Ta, Karen Marder, Chunhua Weng
J. Biomed. Informatics3
2022 Optimizing Clinical Research Eligibility Prescreening: An Iterative Usability Evaluation of an NLP-driven Cohort Identification Tool
Betina Ross S. Idnay, Yilu Fang, Caitlin N. Dreisbach, Karen Marder, Chunhua Weng, Rebecca Schnall
AMIA1
2022 Criteria2Query 2.0: Combining Machine Efficiency and Human Intelligence to Define a More Accurate and Feasible Cohort for Clinical Trial Recruitment
Betina Ross S. Idnay, Yilu Fang, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Rebecca Schnall, Chunhua Weng
AMIA1
2022 Combining human and machine intelligence for clinical trial eligibility querying
abstract
OBJECTIVE: To combine machine efficiency and human intelligence for converting complex clinical trial eligibility criteria text into cohort queries. MATERIALS AND METHODS: Criteria2Query (C2Q) 2.0 was developed to enable real-time user intervention for criteria selection and simplification, parsing error correction, and concept mapping. The accuracy, precision, recall, and F1 score of enhanced modules for negation scope detection, temporal and value normalization were evaluated using a previously curated gold standard, the annotated eligibility criteria of 1010 COVID-19 clinical trials. The usability and usefulness were evaluated by 10 research coordinators in a task-oriented usability evaluation using 5 Alzheimer's disease trials. Data were collected by user interaction logging, a demographic questionnaire, the Health Information Technology Usability Evaluation Scale (Health-ITUES), and a feature-specific questionnaire. RESULTS: The accuracies of negation scope detection, temporal and value normalization were 0.924, 0.916, and 0.966, respectively. C2Q 2.0 achieved a moderate usability score (3.84 out of 5) and a high learnability score (4.54 out of 5). On average, 9.9 modifications were made for a clinical study. Experienced researchers made more modifications than novice researchers. The most frequent modification was deletion (5.35 per study). Furthermore, the evaluators favored cohort queries resulting from modifications (score 4.1 out of 5) and the user engagement features (score 4.3 out of 5). DISCUSSION AND CONCLUSION: Features to engage domain experts and to overcome the limitations in automated machine output are shown to be useful and user-friendly. We concluded that human-computer collaboration is key to improving the adoption and user-friendliness of natural language processing.
Yilu Fang, Betina Ross S. Idnay, Yingcheng Sun, Hao Liu 0054, Zhehuan Chen, Karen Marder, Hua Xu 0001, Rebecca Schnall, Chunhua Weng
J. Am. Medical Informatics Assoc.2
2021 Cognitive Function Characterization Using Electronic Health Records Notes
Adrienne Pichon, Betina Ross S. Idnay, Rebecca Schnall, Karen Marder, Chunhua Weng
AMIA2
2021 A systematic review on natural language processing systems for eligibility prescreening in clinical research
abstract
OBJECTIVE: We conducted a systematic review to assess the effect of natural language processing (NLP) systems in improving the accuracy and efficiency of eligibility prescreening during the clinical research recruitment process. MATERIALS AND METHODS: Guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) standards of quality for reporting systematic reviews, a protocol for study eligibility was developed a priori and registered in the PROSPERO database. Using predetermined inclusion criteria, studies published from database inception through February 2021 were identified from 5 databases. The Joanna Briggs Institute Critical Appraisal Checklist for Quasi-experimental Studies was adapted to determine the study quality and the risk of bias of the included articles. RESULTS: Eleven studies representing 8 unique NLP systems met the inclusion criteria. These studies demonstrated moderate study quality and exhibited heterogeneity in the study design, setting, and intervention type. All 11 studies evaluated the NLP system's performance for identifying eligible participants; 7 studies evaluated the system's impact on time efficiency; 4 studies evaluated the system's impact on workload; and 2 studies evaluated the system's impact on recruitment. DISCUSSION: NLP systems in clinical research eligibility prescreening are an understudied but promising field that requires further research to assess its impact on real-world adoption. Future studies should be centered on continuing to develop and evaluate relevant NLP systems to improve enrollment into clinical studies. CONCLUSION: Understanding the role of NLP systems in improving eligibility prescreening is critical to the advancement of clinical research recruitment.
Betina Ross S. Idnay, Caitlin N. Dreisbach, Chunhua Weng, Rebecca Schnall
J. Am. Medical Informatics Assoc.1
2021 The COVID-19 Trial Finder
abstract
Clinical trials are the gold standard for generating reliable medical evidence. The biggest bottleneck in clinical trials is recruitment. To facilitate recruitment, tools for patient search of relevant clinical trials have been developed, but users often suffer from information overload. With nearly 700 coronavirus disease 2019 (COVID-19) trials conducted in the United States as of August 2020, it is imperative to enable rapid recruitment to these studies. The COVID-19 Trial Finder was designed to facilitate patient-centered search of COVID-19 trials, first by location and radius distance from trial sites, and then by brief, dynamically generated medical questions to allow users to prescreen their eligibility for nearby COVID-19 trials with minimum human computer interaction. A simulation study using 20 publicly available patient case reports demonstrates its precision and effectiveness.
Yingcheng Sun, Alex M. Butler, Fengyang Lin, Hao Liu 0054, Latoya A. Stewart, Jae Hyun Kim, Betina Ross S. Idnay, Qingyin Ge, Xinyi Wei, Cong Liu 0020, Chi Yuan, Chunhua Weng
J. Am. Medical Informatics Assoc.7