Euijung Ryu

dblp:166/8209 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-6281-8738ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Extracting social support and social isolation information from clinical psychiatry notes: comparing a rule-based natural language processing system and a large language model
abstract
OBJECTIVES: Social support (SS) and social isolation (SI) are social determinants of health (SDOH) associated with psychiatric outcomes. In electronic health records (EHRs), individual-level SS/SI is typically documented in narrative clinical notes rather than as structured coded data. Natural language processing (NLP) algorithms can automate the otherwise labor-intensive process of extraction of such information. MATERIALS AND METHODS: Psychiatric encounter notes from Mount Sinai Health System (MSHS, n = 300) and Weill Cornell Medicine (WCM, n = 225) were annotated to create a gold-standard corpus. A rule-based system (RBS) involving lexicons and a large language model (LLM) using FLAN-T5-XL were developed to identify mentions of SS and SI and their subcategories (eg, social network, instrumental support, and loneliness). RESULTS: For extracting SS/SI, the RBS obtained higher macroaveraged F1-scores than the LLM at both MSHS (0.89 versus 0.65) and WCM (0.85 versus 0.82). For extracting the subcategories, the RBS also outperformed the LLM at both MSHS (0.90 versus 0.62) and WCM (0.82 versus 0.81). DISCUSSION AND CONCLUSION: Unexpectedly, the RBS outperformed the LLMs across all metrics. An intensive review demonstrates that this finding is due to the divergent approach taken by the RBS and LLM. The RBS was designed and refined to follow the same specific rules as the gold-standard annotations. Conversely, the LLM was more inclusive with categorization and conformed to common English-language understanding. Both approaches offer advantages, although additional replication studies are warranted.
Braja Gopal Patra, Lauren A. Lepow, Praneet Kasi Reddy Jagadeesh Kumar, Veer Vekaria, Mohit Manoj Sharma, Prakash Adekkanattu, Brian Fennessy, Gavin Hynes, Isotta Landi, Jorge A. Sanchez-Ruiz, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Ardesheer Talati, Myrna Weissman, Mark Olfson, J. John Mann, Yiye Zhang, Alexander Charney, Jyotishman Pathak
J. Am. Medical Informatics Assoc.11
2022 Comprehensive Evaluation of Health Impact of an ML-Based CDS Solution Integrated with Remote Device: Protocol for RCT in Children with Asthma
Lynnea Myers, Shauna M. Overgaard, Tracey Brereton, Jason Greenwood, Joshua Ohde, Matthew Spiten, Kathy Ihrke, Kevin Peterson, Ashwani Khurana, Euijung Ryu, Madison Roy, Chung-Il Wi, Bjorn Nordlund, Young J. Juhn
AMIA11
2022 The role of individual-level socioeconomic status on bias of machine learning algorithm
Euijung Ryu, Katherine S. King, Sunghwan Sohn, Chung-Il Wi, Momin M. Malik, Richard R. Sharp, John D. Halamka, Young J. Juhn
AMIA1
2022 Assessing socioeconomic bias in machine learning algorithms in health care: a case study of the HOUSES index
abstract
OBJECTIVE: Artificial intelligence (AI) models may propagate harmful biases in performance and hence negatively affect the underserved. We aimed to assess the degree to which data quality of electronic health records (EHRs) affected by inequities related to low socioeconomic status (SES), results in differential performance of AI models across SES. MATERIALS AND METHODS: This study utilized existing machine learning models for predicting asthma exacerbation in children with asthma. We compared balanced error rate (BER) against different SES levels measured by HOUsing-based SocioEconomic Status measure (HOUSES) index. As a possible mechanism for differential performance, we also compared incompleteness of EHR information relevant to asthma care by SES. RESULTS: Asthmatic children with lower SES had larger BER than those with higher SES (eg, ratio = 1.35 for HOUSES Q1 vs Q2-Q4) and had a higher proportion of missing information relevant to asthma care (eg, 41% vs 24% for missing asthma severity and 12% vs 9.8% for undiagnosed asthma despite meeting asthma criteria). DISCUSSION: Our study suggests that lower SES is associated with worse predictive model performance. It also highlights the potential role of incomplete EHR data in this differential performance and suggests a way to mitigate this bias. CONCLUSION: The HOUSES index allows AI researchers to assess bias in predictive model performance by SES. Although our case study was based on a small sample size and a single-site study, the study results highlight a potential strategy for identifying bias by using an innovative SES measure.
Young J. Juhn, Euijung Ryu, Chung-Il Wi, Katherine S. King, Momin M. Malik, Santiago Romero-Brufau, Chunhua Weng, Sunghwan Sohn, Richard R. Sharp, John D. Halamka
J. Am. Medical Informatics Assoc.2
2021 Detecting Major Depressive Disorder from Clinical Notes using Neural Language Models with Distant Supervision
Bhavani Singh Agnikula Kshatriya, Nicolas A. Nunez, Manuel Gardea-Resendez, Euijung Ryu, Brandon J. Coombes, Sunyang Fu, Mark A. Frye, Joanna M. Biernacka, Yanshan Wang
AMIA4
2021 Extracting Social Isolation Information From Psychiatric Notes in the Electronic Health Records
Lauren A. Lepow, Braja Gopal Patra, Isotta Landi, Prakash Adekkanattu, Jyotishman Pathak, Mark Olfson, J. John Mann, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Priya Wickramaratne, Myrna Weissman, Benjamin S. Glicksberg, Alexander Charney
AMIA8
2021 Extracting social determinants of health from electronic health records using natural language processing: a systematic review
abstract
OBJECTIVE: Social determinants of health (SDoH) are nonclinical dispositions that impact patient health risks and clinical outcomes. Leveraging SDoH in clinical decision-making can potentially improve diagnosis, treatment planning, and patient outcomes. Despite increased interest in capturing SDoH in electronic health records (EHRs), such information is typically locked in unstructured clinical notes. Natural language processing (NLP) is the key technology to extract SDoH information from clinical text and expand its utility in patient care and research. This article presents a systematic review of the state-of-the-art NLP approaches and tools that focus on identifying and extracting SDoH data from unstructured clinical text in EHRs. MATERIALS AND METHODS: A broad literature search was conducted in February 2021 using 3 scholarly databases (ACL Anthology, PubMed, and Scopus) following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 6402 publications were initially identified, and after applying the study inclusion criteria, 82 publications were selected for the final review. RESULTS: Smoking status (n = 27), substance use (n = 21), homelessness (n = 20), and alcohol use (n = 15) are the most frequently studied SDoH categories. Homelessness (n = 7) and other less-studied SDoH (eg, education, financial problems, social isolation and support, family problems) are mostly identified using rule-based approaches. In contrast, machine learning approaches are popular for identifying smoking status (n = 13), substance use (n = 9), and alcohol use (n = 9). CONCLUSION: NLP offers significant potential to extract SDoH data from narrative clinical notes, which in turn can aid in the development of screening tools, risk prediction models, and clinical decision support systems.
Braja Gopal Patra, Mohit Manoj Sharma, Veer Vekaria, Prakash Adekkanattu, Olga V. Patterson, Benjamin S. Glicksberg, Lauren A. Lepow, Euijung Ryu, Joanna M. Biernacka, Al'ona Furmanchuk, Thomas J. George, William R. Hogan, Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Myrna Weissman, Priya Wickramaratne, J. John Mann, Mark Olfson, Thomas R. Campion Jr., Mark G. Weiner, Jyotishman Pathak
J. Am. Medical Informatics Assoc.8
2020 Assessing Clinician's Adherence to Asthma Guidelines for Asthma Control Status using Natural Language Processing in a Primary Care Setting
Elham Sagheb, Chung-Il Wi, Pragya Shrestha, Euijung Ryu, Miguel Park, Barbara P. Yawn, Young J. Juhn, Sunghwan Sohn
AMIA5
2018 Clinical documentation variations and NLP system portability: a case study in asthma birth cohorts across institutions
abstract
OBJECTIVE: To assess clinical documentation variations across health care institutions using different electronic medical record systems and investigate how they affect natural language processing (NLP) system portability. MATERIALS AND METHODS: Birth cohorts from Mayo Clinic and Sanford Children's Hospital (SCH) were used in this study (n = 298 for each). Documentation variations regarding asthma between the 2 cohorts were examined in various aspects: (1) overall corpus at the word level (ie, lexical variation), (2) topics and asthma-related concepts (ie, semantic variation), and (3) clinical note types (ie, process variation). We compared those statistics and explored NLP system portability for asthma ascertainment in 2 stages: prototype and refinement. RESULTS: There exist notable lexical variations (word-level similarity = 0.669) and process variations (differences in major note types containing asthma-related concepts). However, semantic-level corpora were relatively homogeneous (topic similarity = 0.944, asthma-related concept similarity = 0.971). The NLP system for asthma ascertainment had an F-score of 0.937 at Mayo, and produced 0.813 (prototype) and 0.908 (refinement) when applied at SCH. DISCUSSION: The criteria for asthma ascertainment are largely dependent on asthma-related concepts. Therefore, we believe that semantic similarity is important to estimate NLP system portability. As the Mayo Clinic and SCH corpora were relatively homogeneous at a semantic level, the NLP system, developed at Mayo Clinic, was imported to SCH successfully with proper adjustments to deal with the intrinsic corpus heterogeneity.
Sunghwan Sohn, Yanshan Wang, Chung-Il Wi, Elizabeth A. Krusemark, Euijung Ryu, Mir H. Ali, Young J. Juhn
J. Am. Medical Informatics Assoc.5
2017 Bayesian Prediction of Asthma Exacerbation in Children
Sunghwan Sohn, Young J. Juhn, Sungrim Moon, Chung-Il Wi, Katherine S. King, Euijung Ryu
AMIA6
2016 Asthma Ascertainment NLP System Portability across Institutions
Sunghwan Sohn, Yanshan Wang, Chung-Il Wi, Elizabeth A. Krusemark, Euijung Ryu, Mir H. Ali, Young J. Juhn
AMIA5