VLDB 2026 Research / reviewers in the wild / expert
Cathy Shyr
dblp:324/2819
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-7466-0034ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information extraction from clinical notes: are we ready to switch to large language models?abstractOBJECTIVES: To assess the performance, generalizability, and computational efficiency of instruction-tuned Large Language Model Meta AI (LLaMA)-2 and LLaMA-3 models compared to bidirectional encoder representations from transformers (BERT) for clinical information extraction (IE) tasks, specifically named entity recognition (NER) and relation extraction (RE). MATERIALS AND METHODS: We developed a comprehensive annotated corpus of 1588 clinical notes from 4 data sources-UT Physicians (UTP) (1342 notes), Transcribed Medical Transcription Sample Reports and Examples (MTSamples) (146), Medical Information Mart for Intensive Care (MIMIC)-III (50), and Informatics for Integrating Biology and the Bedside (i2b2) (50), capturing 4 clinical entities (problems, tests, medications, other treatments) and 16 modifiers (eg, negation, certainty). Large Language Model Meta AI-2 and LLaMA-3 were instruction-tuned for clinical NER and RE, and their performance was benchmarked against BERT. RESULTS: Large Language Model Meta AI models consistently outperformed BERT across datasets. In data-rich settings (eg, UTP), LLaMA achieved marginal gains (approximately 1% improvement for NER and 1.5%-3.7% for RE). Under limited data conditions (eg, MTSamples, MIMIC-III) and on the unseen i2b2 dataset, LLaMA-3-70B improved F1 scores by over 7% for NER and 4% for RE. However, performance gains came with increased computational costs, with LLaMA models requiring more memory and Graphics Processing Unit (GPU) hours and running up to 28 times slower than BERT. DISCUSSION: While LLaMA models offer enhanced performance, their higher computational demands and slower throughput highlight the need to balance performance with practical resource constraints. Application-specific considerations are essential when choosing between LLMs and BERT for clinical IE. CONCLUSION: Instruction-tuned LLaMA models show promise for clinical NER and RE tasks. However, the tradeoff between improved performance and increased computational cost must be carefully evaluated. We release our Kiwi package (https://kiwi.clinicalnlp.org/) to facilitate the application of both LLaMA and BERT models in clinical IE applications. Xu Zuo, Yujia Zhou 0003, Xueqing Peng, Jimin Huang, Vipina Kuttichi Keloth, Vincent J. Zhang, Ruey-Ling Weng, Cathy Shyr, Qingyu Chen 0001, Xiaoqian Jiang, Kirk Roberts, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2025 | Reply to Layne et al.'s Letter to the EditorabstractWe appreciate Layne et al.’s comments regarding our recent study on leveraging AI to generate lay summaries of scientific abstracts.1 We agree with the authors that readability is a key aspect of writing effective lay summaries. The authors noted that we “prompted ChatGPT-4 to craft a lay summary ‘in lay language at a 6th grade reading level’ with no other focus in the prompt on readability or other suggestions by the American Medical Association (AMA) recommendations for lay summaries.”2 In response to this, we would like to respectfully point out that our prompt also emphasizes succinctness (“under 100 words”) and clear focus on the key components of a scientific abstract (“highlight the study purpose, methods, key findings, and practical importance of these findings”). These elements are critical to readability and align with the AMA’s checklist for creating written materials for a lay audience.2 We acknowledge that readability formulas (eg, Flesch–Kincaid readability score, SMOG Index) can serve as useful tools for assessing the difficulty of the vocabulary and sentences in lay summaries. However, it is well known that these formulas overlook important factors that influence ease of reading, including content and the reader’s prior knowledge; as a result, their assessments can be inconsistent and often inaccurate.3,4 In the 2020 US Department of Health and Human Services’ guidance document on using readability formulas, it is cautioned that “relying on a grade level score can mislead you into thinking that your materials are clear and effective when they are not.”5 These limitations underscore the need for more comprehensive methods to evaluate the effectiveness of AI-generated content for lay audiences. As AI’s vision and language capabilities continue to evolve, there is potential to leverage these advancements to generate multimodal lay summaries, including AI-generated illustrations to supplement written information and translations into multiple languages. Consequently, evaluation methods would need to evolve and adapt to rigorously assess multimodal content designed for diverse audiences. Key considerations include assessing the accuracy, clarity, potential harm, and cultural relevance to the target audience to better understand the real-world impact of AI-generated materials on public comprehension and engagement with scientific results. Cathy Shyr, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Empowering the biomedical research community: Innovative SAS deployment on the All of Us Researcher WorkbenchabstractOBJECTIVES: The All of Us Research Program is a precision medicine initiative aimed at establishing a vast, diverse biomedical database accessible through a cloud-based data analysis platform, the Researcher Workbench (RW). Our goal was to empower the research community by co-designing the implementation of SAS in the RW alongside researchers to enable broader use of All of Us data. MATERIALS AND METHODS: Researchers from various fields and with different SAS experience levels participated in co-designing the SAS implementation through user experience interviews. RESULTS: Feedback and lessons learned from user testing informed the final design of the SAS application. DISCUSSION: The co-design approach is critical for reducing technical barriers, broadening All of Us data use, and enhancing the user experience for data analysis on the RW. CONCLUSION: Our co-design approach successfully tailored the implementation of the SAS application to researchers' needs. This approach may inform future software implementations on the RW. Izabelle P. Humes, Cathy Shyr, Moira Dillon, Zhongjie Liu, Jennifer Peterson, Chris De St. Jeor, Jacqueline Malkes, Hiral Master, Brandy Mapes, Romuladus Azuine, Nakia Mack, Bassent Abdelbary, Joyonna Gamble-George, Emily Goldmann, Stephanie Cook, Fatemeh Choupani, Rubin Baskir, Sydney J. McMaster, Chris Lunt, Karriem Watson, Minnkyong Lee, Sophie Schwartz, Ruchi Munshi, David Glazer, Eric Banks, Anthony Philippakis, Melissa A. Basford, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Leveraging artificial intelligence to summarize abstracts in lay language for increasing research accessibility and transparencyabstractOBJECTIVE: Returning aggregate study results is an important ethical responsibility to promote trust and inform decision making, but the practice of providing results to a lay audience is not widely adopted. Barriers include significant cost and time required to develop lay summaries and scarce infrastructure necessary for returning them to the public. Our study aims to generate, evaluate, and implement ChatGPT 4 lay summaries of scientific abstracts on a national clinical study recruitment platform, ResearchMatch, to facilitate timely and cost-effective return of study results at scale. MATERIALS AND METHODS: We engineered prompts to summarize abstracts at a literacy level accessible to the public, prioritizing succinctness, clarity, and practical relevance. Researchers and volunteers assessed ChatGPT-generated lay summaries across five dimensions: accuracy, relevance, accessibility, transparency, and harmfulness. We used precision analysis and adaptive random sampling to determine the optimal number of summaries for evaluation, ensuring high statistical precision. RESULTS: ChatGPT achieved 95.9% (95% CI, 92.1-97.9) accuracy and 96.2% (92.4-98.1) relevance across 192 summary sentences from 33 abstracts based on researcher review. 85.3% (69.9-93.6) of 34 volunteers perceived ChatGPT-generated summaries as more accessible and 73.5% (56.9-85.4) more transparent than the original abstract. None of the summaries were deemed harmful. We expanded ResearchMatch's technical infrastructure to automatically generate and display lay summaries for over 750 published studies that resulted from the platform's recruitment mechanism. DISCUSSION AND CONCLUSION: Implementing AI-generated lay summaries on ResearchMatch demonstrates the potential of a scalable framework generalizable to broader platforms for enhancing research accessibility and transparency. Cathy Shyr, Randall W. Grout, Nan Kennedy, Yasemin Akdas, Maeve Tischbein, Joshua Milford, Jason Tan, Kaysi Quarles, Terri L. Edwards, Laurie L. Novak, Jules White, Consuelo H. Wilkins, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Illuminating the landscape of high-level clinical trial opportunities in the All of Us Research ProgramabstractOBJECTIVE: With its size and diversity, the All of Us Research Program has the potential to power and improve representation in clinical trials through ancillary studies like Nutrition for Precision Health. We sought to characterize high-level trial opportunities for the diverse participants and sponsors of future trial investment. MATERIALS AND METHODS: We matched All of Us participants with available trials on ClinicalTrials.gov based on medical conditions, age, sex, and geographic location. Based on the number of matched trials, we (1) developed the Trial Opportunities Compass (TOC) to help sponsors assess trial investment portfolios, (2) characterized the landscape of trial opportunities in a phenome-wide association study (PheWAS), and (3) assessed the relationship between trial opportunities and social determinants of health (SDoH) to identify potential barriers to trial participation. RESULTS: Our study included 181 529 All of Us participants and 18 634 trials. The TOC identified opportunities for portfolio investment and gaps in currently available trials across federal, industrial, and academic sponsors. PheWAS results revealed an emphasis on mental disorder-related trials, with anxiety disorder having the highest adjusted increase in the number of matched trials (59% [95% CI, 57-62]; P < 1e-300). Participants from certain communities underrepresented in biomedical research, including self-reported racial and ethnic minorities, had more matched trials after adjusting for other factors. Living in a nonmetropolitan area was associated with up to 13.1 times fewer matched trials. DISCUSSION AND CONCLUSION: All of Us data are a valuable resource for identifying trial opportunities to inform trial portfolio planning. Characterizing these opportunities with consideration for SDoH can provide guidance on prioritizing the most pressing barriers to trial participation. Cathy Shyr, Lina M. Sulieman, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 1 |