Stephen Pfohl

dblp:225/6445 · also Stephen R. Pfohl, Stephen Robert Pfohl · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-0551-9664ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Leveraging clinical epidemiology concepts to strengthen machine learning fairness evaluations
abstract
OBJECTIVES: The increasing use of machine learning (ML) in clinical care makes fairness a central issue. Fairness, defined as the absence of disparities across individuals or subgroups, shares several parallels with concepts in clinical epidemiology. The objective was to apply clinical epidemiology frameworks to the evaluation of ML fairness, both to enhance understanding of these concepts and to strengthen fairness assessments. METHODS: This manuscript addresses: (1) the connection between clinical epidemiology bias terms and ML fairness concepts; (2) the relationship between diagnostic testing metrics and fairness criteria; (3) issues arising from multiple testing; and (4) strategies for fairness considerations leveraging clinical epidemiology principles. RESULTS: Unfairness can arise at different stages: before model development, during model development and post-deployment. The root of unfairness can be conceptualized as a result of clinical epidemiology bias concepts, such as selection, measurement, model misspecification, cognitive and implementation bias. Common approaches to evaluating fairness involve comparing model performance metrics across 1 or more subgroups. Four widely used fairness criteria are independence, separation, sufficiency and predictive parity. They can be assessed using diagnostic testing metrics. Fairness evaluations are vulnerable to multiple testing issues, with subgroup analyses posing particular risks for spurious findings. Solutions can leverage established clinical epidemiology principles such as pre-specifying the analytic strategy. CONCLUSIONS: Many parallels exist between ML fairness and clinical epidemiology, including the conceptualization of the root causes of unfairness, the articulation of fairness criteria, and considerations related to multiple testing. Methodologically sound fairness approaches can leverage well-established principles from clinical epidemiology.
Lin Lawrence Guo, Santiago Eduardo Arciniegas, Adam Paul Yan, George Tomlinson, Melissa Beauchemin, Stephen Pfohl, Lillian Sung
J. Am. Medical Informatics Assoc.6
2025 Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
abstract
Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal performance across subgroups is an unreliable measure of fairness when data are representative of the relevant populations but reflective of real-world disparities. Furthermore, when data are not representative due to selection bias, both disaggregated evaluation and alternative approaches based on conditional independence testing may be invalid without explicit assumptions regarding the bias mechanism. We use causal graphical models to characterize fairness properties and metric stability across subgroups under different data generating processes. Our framework suggests complementing disaggregated evaluations with explicit causal assumptions and analysis to control for confounding and distribution shift, including conditional independence testing and weighted performance estimation. These findings have broad implications for how practitioners design and interpret model assessments given the ubiquity of disaggregated evaluation.
Stephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras, Vishwali Mhasawade, Olawale Salaudeen, Awa Dieng, Shannon Sequeira, Santiago Eduardo Arciniegas, Lillian Sung, Nnamdi Ezeanochie, Heather Cole-Lewis, Katherine A. Heller, Oluwasanmi Koyejo, Alexander D'Amour
NeurIPS1
2024 Proxy Methods for Domain Adaptation
abstract
We study the problem of domain adaptation under distribution shift, where the shift is due to a change in the distribution of an unobserved, latent variable that confounds both the covariates and the labels. In this setting, neither the covariate shift nor the label shift assumptions apply. Our approach to adaptation employs proximal causal learning, a technique for estimating causal effects in settings where proxies of unobserved confounders are available. We demonstrate that proxy variables allow for adaptation to distribution shift without explicitly recovering or modeling latent variables. We consider two settings, (i) Concept Bottleneck: an additional “concept” variable is observed that mediates the relationship between the covariates and labels; (ii) Multi-domain: training data from multiple source domains is available, where each source domain exhibits a different distribution over the latent confounder. We develop a two-stage kernel estimation approach to adapt to complex distribution shifts in both settings. In our experiments, we show that our approach outperforms other methods, notably those which explicitly recover the latent confounder.
Katherine Tsai, Stephen Pfohl, Olawale Salaudeen, Nicole Chiou, Matt J. Kusner, Alexander D'Amour, Oluwasanmi Koyejo, Arthur Gretton
AISTATS2
2023 Adapting to Latent Subgroup Shifts via Concepts and Proxies
abstract
We address the problem of unsupervised domain adaptation when the source domain differs from the target domain because of a shift in the distribution of a latent subgroup. When this subgroup confounds all observed data, neither covariate shift nor label shift assumptions apply. We show that the optimal target predictor can be non-parametrically identified with the help of concept and proxy variables available only in the source domain, and unlabeled data from the target. The identification results are constructive, immediately suggesting an algorithm for estimating the optimal predictor in the target. For continuous observations, when this algorithm becomes impractical, we propose a latent variable model specific to the data generation process at hand. We show how the approach degrades as the size of the shift changes, and verify that it outperforms both covariate and label shift adjustment.
Ibrahim Alabdulmohsin, Nicole Chiou, Alexander D'Amour, Arthur Gretton, Oluwasanmi Koyejo, Matt J. Kusner, Stephen Pfohl, Olawale Salaudeen, Jessica Schrouff, Katherine Tsai
AISTATS7
2023 Self-supervised machine learning using adult inpatient data produces effective models for pediatric clinical prediction tasks
abstract
OBJECTIVE: Development of electronic health records (EHR)-based machine learning models for pediatric inpatients is challenged by limited training data. Self-supervised learning using adult data may be a promising approach to creating robust pediatric prediction models. The primary objective was to determine whether a self-supervised model trained in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients, for pediatric inpatient clinical prediction tasks. MATERIALS AND METHODS: This retrospective cohort study used EHR data and included patients with at least one admission to an inpatient unit. One admission per patient was randomly selected. Adult inpatients were 18 years or older while pediatric inpatients were more than 28 days and less than 18 years. Admissions were temporally split into training (January 1, 2008 to December 31, 2019), validation (January 1, 2020 to December 31, 2020), and test (January 1, 2021 to August 1, 2022) sets. Primary comparison was a self-supervised model trained in adult inpatients versus count-based logistic regression models trained in pediatric inpatients. Primary outcome was mean area-under-the-receiver-operating-characteristic-curve (AUROC) for 11 distinct clinical outcomes. Models were evaluated in pediatric inpatients. RESULTS: When evaluated in pediatric inpatients, mean AUROC of self-supervised model trained in adult inpatients (0.902) was noninferior to count-based logistic regression models trained in pediatric inpatients (0.868) (mean difference = 0.034, 95% CI=0.014-0.057; P < .001 for noninferiority and P = .006 for superiority). CONCLUSIONS: Self-supervised learning in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients. This finding suggests transferability of self-supervised models trained in adult patients to pediatric patients, without requiring costly model retraining.
Joshua Lemmon, Lin Lawrence Guo, Ethan Steinberg, Keith E. Morse, Scott L. Fleming, Catherine Aftandilian, Stephen Pfohl, José D. Posada, Nigam H. Shah, Jason Alan Fries, Lillian Sung
J. Am. Medical Informatics Assoc.7
2021 Learning decision thresholds for risk stratification models from aggregate clinician behavior
abstract
Using a risk stratification model to guide clinical practice often requires the choice of a cutoff-called the decision threshold-on the model's output to trigger a subsequent action such as an electronic alert. Choosing this cutoff is not always straightforward. We propose a flexible approach that leverages the collective information in treatment decisions made in real life to learn reference decision thresholds from physician practice. Using the example of prescribing a statin for primary prevention of cardiovascular disease based on 10-year risk calculated by the 2013 pooled cohort equations, we demonstrate the feasibility of using real-world data to learn the implicit decision threshold that reflects existing physician behavior. Learning a decision threshold in this manner allows for evaluation of a proposed operating point against the threshold reflective of the community standard of care. Furthermore, this approach can be used to monitor and audit model-guided clinical decision making following model deployment.
Birju Patel, Ethan Steinberg, Stephen Pfohl, Nigam H. Shah
J. Am. Medical Informatics Assoc.3
2021 An empirical characterization of fair machine learning for clinical risk prediction
Stephen Pfohl, Agata Foryciarz, Nigam H. Shah
J. Biomed. Informatics1
2021 Language models are an effective representation learning technique for electronic health record data
Ethan Steinberg, Kenneth Jung, Jason Alan Fries, Conor K. Corbin, Stephen Pfohl, Nigam H. Shah
J. Biomed. Informatics5
2019 Creating Fair Models of Atherosclerotic Cardiovascular Disease Risk
abstract
Guidelines for the management of atherosclerotic cardiovascular disease (ASCVD) recommend the use of risk stratification models to identify patients most likely to benefit from cholesterol-lowering and other therapies. These models have differential performance across race and gender groups with inconsistent behavior across studies, potentially resulting in an inequitable distribution of beneficial therapy. In this work, we leverage adversarial learning and a large observational cohort extracted from electronic health records (EHRs) to develop a "fair" ASCVD risk prediction model with reduced variability in error rates across groups. We empirically demonstrate that our approach is capable of aligning the distribution of risk predictions conditioned on the outcome across several groups simultaneously for models built from high-dimensional EHR data. We also discuss the relevance of these results in the context of the empirical trade-off between fairness and model performance.
Stephen Pfohl, Ben J. Marafino, Adrien Coulet, Fátima Rodriguez, Latha Palaniappan, Nigam H. Shah
AIES1
2018 Transfer learning to adapt predictive models for pediatric patients in the EHR
Stephen Pfohl, Nigam H. Shah
AMIA1