Lillian Sung

dblp:234/6547 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0003-0951-3091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Systematic review of foundation models for structured electronic health records
abstract
PURPOSE: Foundation models pretrained on structured electronic health record (EHR) data promise improved predictive performance, sample efficiency and resilience to distribution shifts. However, model design, scale and use remain unclear. Objectives were to characterize foundation models pretrained on structured EHR data; examine temporal trends in model application and scale, architecture and design; and assess the extent to which publications omitted methodological details. METHODS: We searched MEDLINE and Embase (2018-October 2025) for foundation models pretrained on structured EHR data using self-supervised learning and applied to clinical prediction tasks. Study selection and data abstraction were performed in duplicate. Characteristics were summarized and stratified by median publication year. RESULTS: Fifty-three studies were included; publications increased over time. Most datasets (79%) originated from the United States. None pretrained exclusively on pediatric cohorts. Model architecture shifted towards transformers (P = .013) with longer context windows (P = .028), while application shifted from exclusively embedding-based toward generative or mixed use (P < .001). Choices regarding feature inclusion, temporal representation, self-supervised objective and downstream adaptation remained heterogeneous. Only 26% of studies evaluated transfer to external datasets, and none described clinical deployment. Key indicators of scale and compute were frequently unreported. CONCLUSIONS: EHR foundation models are proliferating and increasingly transformer-based and generative. Yet methodological choices and reporting remain fragmented, indicating design trade-offs and best practices for EHR foundation models have not yet been established. None describe clinical deployment. Future work should clarify which design choices improve performance, robustness and transferability, increase reporting transparency and identify if they can be implemented to improve patient-important outcomes.
Lin Lawrence Guo, Santiago Eduardo Arciniegas, Adam Paul Yan, Jason Alan Fries, George Tomlinson, Lillian Sung
J. Am. Medical Informatics Assoc.6
2026 Leveraging clinical epidemiology concepts to strengthen machine learning fairness evaluations
abstract
OBJECTIVES: The increasing use of machine learning (ML) in clinical care makes fairness a central issue. Fairness, defined as the absence of disparities across individuals or subgroups, shares several parallels with concepts in clinical epidemiology. The objective was to apply clinical epidemiology frameworks to the evaluation of ML fairness, both to enhance understanding of these concepts and to strengthen fairness assessments. METHODS: This manuscript addresses: (1) the connection between clinical epidemiology bias terms and ML fairness concepts; (2) the relationship between diagnostic testing metrics and fairness criteria; (3) issues arising from multiple testing; and (4) strategies for fairness considerations leveraging clinical epidemiology principles. RESULTS: Unfairness can arise at different stages: before model development, during model development and post-deployment. The root of unfairness can be conceptualized as a result of clinical epidemiology bias concepts, such as selection, measurement, model misspecification, cognitive and implementation bias. Common approaches to evaluating fairness involve comparing model performance metrics across 1 or more subgroups. Four widely used fairness criteria are independence, separation, sufficiency and predictive parity. They can be assessed using diagnostic testing metrics. Fairness evaluations are vulnerable to multiple testing issues, with subgroup analyses posing particular risks for spurious findings. Solutions can leverage established clinical epidemiology principles such as pre-specifying the analytic strategy. CONCLUSIONS: Many parallels exist between ML fairness and clinical epidemiology, including the conceptualization of the root causes of unfairness, the articulation of fairness criteria, and considerations related to multiple testing. Methodologically sound fairness approaches can leverage well-established principles from clinical epidemiology.
Lin Lawrence Guo, Santiago Eduardo Arciniegas, Adam Paul Yan, George Tomlinson, Melissa Beauchemin, Stephen Pfohl, Lillian Sung
J. Am. Medical Informatics Assoc.7
2025 Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
abstract
Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal performance across subgroups is an unreliable measure of fairness when data are representative of the relevant populations but reflective of real-world disparities. Furthermore, when data are not representative due to selection bias, both disaggregated evaluation and alternative approaches based on conditional independence testing may be invalid without explicit assumptions regarding the bias mechanism. We use causal graphical models to characterize fairness properties and metric stability across subgroups under different data generating processes. Our framework suggests complementing disaggregated evaluations with explicit causal assumptions and analysis to control for confounding and distribution shift, including conditional independence testing and weighted performance estimation. These findings have broad implications for how practitioners design and interpret model assessments given the ubiquity of disaggregated evaluation.
Stephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras, Vishwali Mhasawade, Olawale Salaudeen, Awa Dieng, Shannon Sequeira, Santiago Eduardo Arciniegas, Lillian Sung, Nnamdi Ezeanochie, Heather Cole-Lewis, Katherine A. Heller, Oluwasanmi Koyejo, Alexander D'Amour
NeurIPS10
2023 Self-supervised machine learning using adult inpatient data produces effective models for pediatric clinical prediction tasks
abstract
OBJECTIVE: Development of electronic health records (EHR)-based machine learning models for pediatric inpatients is challenged by limited training data. Self-supervised learning using adult data may be a promising approach to creating robust pediatric prediction models. The primary objective was to determine whether a self-supervised model trained in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients, for pediatric inpatient clinical prediction tasks. MATERIALS AND METHODS: This retrospective cohort study used EHR data and included patients with at least one admission to an inpatient unit. One admission per patient was randomly selected. Adult inpatients were 18 years or older while pediatric inpatients were more than 28 days and less than 18 years. Admissions were temporally split into training (January 1, 2008 to December 31, 2019), validation (January 1, 2020 to December 31, 2020), and test (January 1, 2021 to August 1, 2022) sets. Primary comparison was a self-supervised model trained in adult inpatients versus count-based logistic regression models trained in pediatric inpatients. Primary outcome was mean area-under-the-receiver-operating-characteristic-curve (AUROC) for 11 distinct clinical outcomes. Models were evaluated in pediatric inpatients. RESULTS: When evaluated in pediatric inpatients, mean AUROC of self-supervised model trained in adult inpatients (0.902) was noninferior to count-based logistic regression models trained in pediatric inpatients (0.868) (mean difference = 0.034, 95% CI=0.014-0.057; P < .001 for noninferiority and P = .006 for superiority). CONCLUSIONS: Self-supervised learning in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients. This finding suggests transferability of self-supervised models trained in adult patients to pediatric patients, without requiring costly model retraining.
Joshua Lemmon, Lin Lawrence Guo, Ethan Steinberg, Keith E. Morse, Scott L. Fleming, Catherine Aftandilian, Stephen Pfohl, José D. Posada, Nigam H. Shah, Jason Alan Fries, Lillian Sung
J. Am. Medical Informatics Assoc.11
2022 Real-World Data at the Point-of-Care: Exploring Education Needs Across All Stakeholders
Julian Z. Genkins, Dev Dash, Lillian Sung, Saurabh Gombar, Alison Callahan
AMIA3
2021 A Retrospective Analysis of Machine Learning Driven Antibiotic Selection in the Emergency Department
Conor K. Corbin, Arhana Chattopadhyah, Lillian Sung, Amy Chang, Stan Deresinski, Jonathan H. Chen
AMIA3