EDBT 2026 Demo / reviewers in the wild / expert
Marzyeh Ghassemi
dblp:145/6563
· DBLP profile ↗
48ranked-venue papers
2as first author
41since 2021 · last 2026
0000-0001-6349-7251ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 2 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SFTMix: Elevating Language Model Instruction Tuning with Mixup RecipeabstractTo acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-response pairs using next-token prediction (NTP). Efforts to improve instruction tuning often focus on higher-quality supervised fine-tuning (SFT) datasets, typically requiring data filtering with proprietary LLMs or human annotation. In this paper, we take a different approach by proposing SFTMix, a novel Mixup-based recipe that elevates LLM instruction tuning without relying on well-curated datasets. We observe that LLMs exhibit uneven confidence across the semantic representation space. We argue that examples with different confidence levels should play distinct roles in instruction tuning: Confident data is prone to overfitting, while unconfident data is harder to generalize. Based on this insight, SFTMix leverages training dynamics to identify examples with varying confidence levels. We then interpolate them to bridge the confidence gap and apply a Mixup-based regularization to support learning on these additional, interpolated examples. We demonstrate the effectiveness of SFTMix in both instruction-following and healthcare-specific SFT tasks, with consistent improvements across LLM families and SFT datasets of varying sizes and qualities. Extensive analyses across six directions highlight SFTMix's compatibility with data selection, adaptability to compute-constrained scenarios, and scalability to broader applications. Yuxin Xiao, Shujian Zhang, Marzyeh Ghassemi, Wenxuan Zhou 0005 |
ACL (1) | 3 |
| 2025 | Vision-Language Models Do Not Understand NegationabstractMany practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain objects but not others. Despite advancements in vision-language models (VLMs) through large-scale training, their ability to comprehend negation remains underexplored. This study addresses the question: how well do current VLMs understand negation? We introduce NegBench, a new benchmark designed to evaluate negation understanding across 18 task variations and 79k examples spanning image, video, and medical datasets. The benchmark consists of two core tasks designed to evaluate negation understanding in diverse multimodal settings: Retrieval with Negation and Multiple Choice Questions with Negated Captions. Our evaluation reveals that modern VLMs struggle significantly with negation, often performing at chance level. To address these shortcomings, we explore a data-centric approach wherein we finetune CLIP models on large-scale synthetic datasets containing millions of negated captions. We show that this approach can result in a 10% increase in recall on negated queries and a 28% boost in accuracy on multiple-choice questions with negated captions. Kumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li 0013, Philip Torr 0001, Marzyeh Ghassemi |
CVPR | 7 |
| 2025 | MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent SimulationsabstractWe present a novel, open-source social network simulation framework MOSAIC where generative language agents predict user behaviors such as liking, sharing, and flagging content.This simulation combines LLM agents with a directed social graph to analyze emergent deception behaviors and gain a better understanding of how users determine the veracity of online social content.By constructing user representations from diverse fine-grained actual user personas, our system enables multi-agent simulations that model content dissemination and engagement dynamics at scale.Within this framework, we evaluate three different content moderation strategies with simulated misinformation dissemination, and we find that they not only mitigate the spread of non-factual content but also increase user engagement.In addition, we analyze the trajectories of popular content in our simulations, and explore whether simulation agents' articulated reasoning for their social interactions truly aligns with their collective engagement patterns.We opensource our simulation software to encourage further research within AI and social sciences: Genglin Liu, Vivian T. Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, Saadia Gabriel |
EMNLP | 5 |
| 2025 | LEMoN: Label Error Detection using Multimodal NeighborsabstractLarge repositories of image-caption pairs are essential for the development of vision-language models. However, these datasets are often extracted from noisy data scraped from the web, and contain many mislabeled instances. In order to improve the reliability of downstream models, it is important to identify and filter images with incorrect captions. However, beyond filtering based on image-caption embedding similarity, no prior works have proposed other methods to filter noisy multimodal data, or concretely assessed the impact of noisy captioning data on downstream training. In this work, we propose, theoretically justify, and empirically validate LEMoN, a method to identify label errors in image-caption datasets. Our method leverages the multimodal neighborhood of image-caption pairs in the latent space of contrastively pretrained multimodal models to automatically identify label errors. Through empirical evaluations across eight datasets and twelve baselines, we find that LEMoN outperforms the baselines by over 3% in label error detection, and that training on datasets filtered using our method improves downstream captioning performance by more than 2 BLEU points over noisy training. Haoran Zhang 0003, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong, Marzyeh Ghassemi |
ICML | 7 |
| 2025 | Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple InteractionsabstractDespite extensive safety alignment efforts, large language models (LLMs) remain vulnerable to jailbreak attacks that elicit harmful behavior. While existing studies predominantly focus on attack methods that require technical expertise, two critical questions remain underexplored: (1) Are jailbroken responses truly useful in enabling average users to carry out harmful actions? (2) Do safety vulnerabilities exist in more common, simple human-LLM interactions? In this paper, we demonstrate that LLM responses most effectively facilitate harmful actions when they are both *actionable* and *informative*---two attributes easily elicited in multi-step, multilingual interactions. Using this insight, we propose HarmScore, a jailbreak metric that measures how effectively an LLM response enables harmful actions, and Speak Easy, a simple multi-step, multilingual attack framework. Notably, by incorporating Speak Easy into direct request and jailbreak baselines, we see an average absolute increase of $0.319$ in Attack Success Rate and $0.426$ in HarmScore in both open-source and proprietary LLMs across four safety benchmarks. Our work reveals a critical yet often overlooked vulnerability: Malicious users can easily exploit common interaction patterns for harmful intentions. Yik Siu Chan, Narutatsu Ri, Yuxin Xiao, Marzyeh Ghassemi |
ICML | 4 |
| 2025 | Aggregation Hides Out-of-Distribution Generalization Failures from Spurious CorrelationsabstractBenchmarks for out-of-distribution (OOD) generalization often reveal a strong positive correlation between in-distribution (ID) and OOD accuracy across models, a phenomenon known as “accuracy-on-the-line.” This pattern is commonly interpreted as evidence that spurious correlations—relationships that improve ID but harm OOD performance—are rare in practice. We show that this positive correlation can be an artifact of aggregating heterogeneous OOD examples. Using a simple gradient-based method, OODSelect, we identify semantically coherent OOD subsets where accuracy-on-the-line breaks down. Across widely used distribution-shift benchmarks, OODSelect uncovers subsets—sometimes comprising more than half of the standard OOD set—where higher ID accuracy predicts lower OOD accuracy. These results suggest that aggregate metrics can mask critical failure modes in OOD robustness. We release code and the identified subsets to support further research. Olawale Salaudeen, Haoran Zhang 0003, Kumail Alhamoud, Sara Beery, Marzyeh Ghassemi |
NeurIPS | 5 |
| 2025 | Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language ModelsabstractFor an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can also convey implicit information. Recent work shows that \textit{syntactic templates}---frequent sequences of Part-of-Speech (PoS) tags---are prevalent in training data and often appear in model outputs. In this work we characterize syntactic templates, domain, and semantics in task-instruction pairs. We identify cases of spurious correlations between syntax and domain, where models learn to associate a domain with syntax during training; this can sometimes override prompt semantics. Using a synthetic training dataset, we find that the syntactic-domain correlation can lower performance (mean 0.51 +/- 0.06) on entity knowledge tasks in OLMo-2 models (1B-13B). We introduce an evaluation framework to detect this phenomenon in trained models, and show that it occurs on a subset of the FlanV2 dataset in open (OLMo-2-7B; Llama-4-Maverick), and closed (GPT-4o) models. Finally, we present a case study on the implications for LLM security, showing that unintended syntactic-domain correlations can be used to bypass refusals in OLMo-2-7B Instruct and GPT-4o. Our findings highlight two needs: (1) to explicitly test for syntactic-domain correlations, and (2) to ensure \textit{syntactic} diversity in training data, specifically within domains, to prevent such spurious correlations. Chantal Shaib, Vinith M. Suriyakumar, Byron C. Wallace, Marzyeh Ghassemi |
NeurIPS | 4 |
| 2025 | An Investigation of Memorization Risk in Healthcare Foundation ModelsabstractFoundation models trained on large-scale de-identified electronic health records (EHRs) hold promise for clinical applications. However, their capacity to memorize patient information raises important privacy concerns. In this work, we introduce a suite of black-box evaluation tests to assess privacy-related memorization risks in foundation models trained on structured EHR data. Our framework includes methods for probing memorization at both the embedding and generative levels, and aims to distinguish between model generalization and harmful memorization in clinically relevant settings. We contextualize memorization in terms of its potential to compromise patient privacy, particularly for vulnerable subgroups. We validate our approach on a publicly available EHR foundation model and release an open-source toolkit to facilitate reproducible and collaborative privacy assessments in healthcare AI. Sana Tonekaboni, Lena Stempfle, Adibvafa Fallahpour, Walter Gerych, Marzyeh Ghassemi |
NeurIPS | 5 |
| 2025 | KScope: A Framework for Characterizing the Knowledge Status of Language ModelsabstractCharacterizing a large language model's (LLM's) knowledge of a given question is challenging.
As a result, prior work has primarily examined LLM behavior under knowledge conflicts, where the model's internal parametric memory contradicts information in the external context.
However, this does not fully reflect how well the model knows the answer to the question.
In this paper, we first introduce a taxonomy of five knowledge statuses based on the consistency and correctness of LLM knowledge modes.
We then propose KScope, a hierarchical framework of statistical tests that progressively refines hypotheses about knowledge modes and characterizes LLM knowledge into one of these five statuses.
We apply KScope to nine LLMs across four datasets and systematically establish:
(1) Supporting context narrows knowledge gaps across models.
(2) Context features related to difficulty, relevance, and familiarity drive successful knowledge updates.
(3) LLMs exhibit similar feature preferences when partially correct or conflicted, but diverge sharply when consistently wrong.
(4) Context summarization constrained by our feature analysis, together with enhanced credibility, further improves update effectiveness and generalizes across LLMs. Yuxin Xiao, Shan Chen 0004, Jack Gallifant, Danielle S. Bitterman, Thomas Hartvigsen, Marzyeh Ghassemi |
NeurIPS | 6 |
| 2025 | On Group Sufficiency Under Label BiasabstractReal-world classification datasets often contain label bias, where observed labels differ systematically from the true labels at different rates for different demographic groups. Machine learning models trained on such datasets may then exhibit disparities in predictive performance across these groups. In this work, we characterize the problem of learning fair classification models with respect to the underlying ground truth labels when given only label biased data. We focus on the particular fairness definition of group sufficiency, i.e. equal calibration of risk scores across protected groups. We theoretically show that enforcing fairness with respect to label biased data necessarily results in group miscalibration with respect to the true labels. We then propose a regularizer which minimizes an upper bound on the sufficiency gap by penalizing a conditional mutual information term. Across experiments on eight tabular, image, and text datasets with both synthetic and real label noise, we find that our method reduces the sufficiency gap by up to 7.2% with no significant decrease in overall accuracy. Haoran Zhang 0003, Olawale Salaudeen, Marzyeh Ghassemi |
NeurIPS | 3 |
| 2025 | What's in a Query: Polarity-Aware Distribution-Based Fair RankingabstractMachine learning-driven rankings, where individuals (or items) are ranked in response to a query, mediate search exposure or attention in a variety of safety-critical settings. Thus, it is important to ensure that such rankings are fair. Under the goal of equal opportunity, attention allocated to an individual on a ranking interface should be proportional to their relevance across search queries. In this work, we examine amortized fair ranking -- where relevance and attention are cumulated over a sequence of user queries to make fair ranking more feasible in practice. Unlike prior methods that operate on expected amortized attention for each individual, we define new divergence-based measures for attention distribution-based fairness in ranking (DistFaiR), characterizing unfairness as the divergence between the distribution of attention and relevance corresponding to an individual over time. This allows us to propose new definitions of unfairness, which are more reliable at test time. Second, we prove that group fairness is upper-bounded by individual fairness under this definition for a useful class of divergence measures, and experimentally show that maximizing individual fairness through an integer linear programming-based optimization is often beneficial to group fairness. Lastly, we find that prior research in amortized fair ranking ignores critical information about queries, potentially leading to a fairwashing risk in practice by making rankings appear more fair than they actually are. Aparna Balagopalan, Kai Wang 0040, Olawale Salaudeen, Asia J. Biega, Marzyeh Ghassemi |
WWW | 5 |
| 2024 | Identifying Implicit Social Biases in Vision-Language ModelsabstractVision-language models, like CLIP (Contrastive Language Image Pretraining), are becoming increasingly popular for a wide range of multimodal retrieval tasks. However, prior work has shown that large language and deep vision models can learn historical biases contained in their training sets, leading to perpetuation of stereotypes and potential downstream harm. In this work, we conduct a systematic analysis of the social biases that are present in CLIP, with a focus on the interaction between image and text modalities. We first propose a taxonomy of social biases called So-B-It, which contains 374 words categorized across ten types of bias. Each type can lead to societal harm if associated with a particular demographic group. Using this taxonomy, we examine images retrieved by CLIP from a facial image dataset using each word as part of a prompt. We find that CLIP frequently displays undesirable associations between harmful words and specific demographic groups, such as retrieving mostly pictures of Middle Eastern men when asked to retrieve images of a "terrorist". Finally, we conduct an analysis of the source of such biases, by showing that the same harmful stereotypes are also present in a large image-text dataset used to train CLIP models for examples of biases that we find. Our findings highlight the importance of evaluating and addressing bias in vision-language models, and suggest the need for transparency and fairness-aware curation of large pre-training datasets. Kimia Hamidieh, Haoran Zhang 0003, Walter Gerych, Thomas Hartvigsen, Marzyeh Ghassemi |
AIES (1) | 5 |
| 2024 | Time2Stop: Adaptive and Explainable Human-AI Loop for Smartphone Overuse InterventionabstractDespite a rich history of investigating smartphone overuse intervention techniques, AI-based just-in-time adaptive intervention (JITAI) methods for overuse reduction are lacking. We develop Time2Stop, an intelligent, adaptive, and explainable JITAI system that leverages machine learning to identify optimal intervention timings, introduces interventions with transparent AI explanations, and collects user feedback to establish a human-AI loop and adapt the intervention model over time. We conducted an 8-week field experiment (N=71) to evaluate the effectiveness of both the adaptation and explanation aspects of Time2Stop. Our results indicate that our adaptive models significantly outperform the baseline methods on intervention accuracy (>32.8% relatively) and receptivity (>8.0%). In addition, incorporating explanations further enhances the effectiveness by 53.8% and 11.4% on accuracy and receptivity, respectively. Moreover, Time2Stop significantly reduces overuse, decreasing app visit frequency by 7.0 ∼ 8.9%. Our subjective data also echoed these quantitative measures. Participants preferred the adaptive interventions and rated the system highly on intervention time accuracy, effectiveness, and level of trust. We envision our work can inspire future research on JITAI systems with a human-AI loop to evolve with users. Adiba Orzikulova, Zhipeng Li 0001, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Marzyeh Ghassemi, Sung-Ju Lee 0001, Anind K. Dey, Xuhai Xu |
CHI | 7 |
| 2024 | MisinfoEval: Generative AI in the Era of "Alternative Facts"abstractThe spread of misinformation on social media platforms threatens democratic processes, contributes to massive economic losses, and endangers public health.Many efforts to address misinformation focus on a knowledge deficit model and propose interventions for improving users' critical thinking through access to facts.Such efforts are often hampered by challenges with scalability, and by platform users' personal biases.The emergence of generative AI presents promising opportunities for countering misinformation at scale across ideological barriers.In this paper, we introduce a framework (Mis-infoEval) for generating and comprehensively evaluating large language model (LLM) based misinformation interventions.We present (1) an experiment with a simulated social media environment to measure effectiveness of misinformation interventions, and (2) a second experiment with personalized explanations tailored to the demographics and beliefs of users with the goal of countering misinformation by appealing to their pre-existing values.Our findings confirm that LLM-based interventions are highly effective at correcting user behavior (improving overall user accuracy at reliability labeling by up to 41.72%).Furthermore, we find that users favor more personalized interventions when making decisions about news reliability and users shown personalized interventions have significantly higher accuracy at identifying misinformation. Saadia Gabriel, Liang Lyu 0001, James Siderius, Marzyeh Ghassemi, Jacob Andreas, Asuman E. Ozdaglar |
EMNLP | 4 |
| 2024 | Views Can Be Deceiving: Improved SSL Through Feature Space AugmentationabstractSupervised learning methods have been found to exhibit inductive biases favoring simpler features. When such features are spuriously correlated with the label, this can result in suboptimal performance on minority subgroups. Despite the growing popularity of methods which learn from unlabeled data, the extent to which these representations rely on spurious features for prediction is unclear. In this work, we explore the impact of spurious features on Self-Supervised Learning (SSL) for visual representation learning. We first empirically show that commonly used augmentations in SSL can cause undesired invariances in the image space, and illustrate this with a simple example. We further show that classical approaches in combating spurious correlations, such as dataset re-sampling during SSL, do not consistently lead to invariant representations. Motivated by these findings, we propose LateTVG to remove spurious information from these representations during pre-training, by regularizing later layers of the encoder via pruning. We find that our method produces representations which outperform the baselines on several benchmarks, without the need for group or label information during SSL. Kimia Hamidieh, Haoran Zhang 0003, Swami Sankaranarayanan, Marzyeh Ghassemi |
ICLR | 4 |
| 2024 | Measuring Stochastic Data Complexity with Boltzmann Influence FunctionsabstractEstimating the uncertainty of a model’s prediction on a test point is a crucial part of ensuring reliability and calibration under distribution shifts.A minimum description length approach to this problem uses the predictive normalized maximum likelihood (pNML) distribution, which considers every possible label for a data point, and decreases confidence in a prediction if other labels are also consistent with the model and training data. In this work we propose IF-COMP, a scalable and efficient approximation of the pNML distribution that linearizes the model with a temperature-scaled Boltzmann influence function. IF-COMP can be used to produce well-calibrated predictions on test points as well as measure complexity in both labelled and unlabelled settings. We experimentally validate IF-COMP on uncertainty calibration, mislabel detection, and OOD detection tasks, where it consistently matches or beats strong baseline methods. Nathan Ng 0001, Roger B. Grosse, Marzyeh Ghassemi |
ICML | 3 |
| 2024 | Position: Application-Driven Innovation in Machine LearningabstractIn this position paper, we argue that application-driven research has been systemically under-valued in the machine learning community. As applications of machine learning proliferate, innovative algorithms inspired by specific real-world challenges have become increasingly important. Such work offers the potential for significant impact not merely in domains of application but also in machine learning itself. In this paper, we describe the paradigm of application-driven research in machine learning, contrasting it with the more standard paradigm of methods-driven research. We illustrate the benefits of application-driven machine learning and how this approach can productively synergize with methods-driven work. Despite these benefits, we find that reviewing, hiring, and teaching practices in machine learning often hold back application-driven innovation. We outline how these processes may be improved. David Rolnick, Alán Aspuru-Guzik, Sara Beery, Bistra Dilkina, Priya L. Donti, Marzyeh Ghassemi, Hannah Kerner, Claire Monteleoni, Esther Rolf, Milind Tambe |
ICML | 6 |
| 2024 | Asymmetry in Low-Rank Adapters of Foundation ModelsabstractParameter-efficient fine-tuning optimizes large, pre-trained foundation models by updating a subset of parameters; in this class, Low-Rank Adaptation (LoRA) is particularly effective. Inspired by an effort to investigate the different roles of LoRA matrices during fine-tuning, this paper characterizes and leverages unexpected asymmetry in the importance of low-rank adapter matrices. Specifically, when updating the parameter matrices of a neural network by adding a product $BA$, we observe that the $B$ and $A$ matrices have distinct functions: $A$ extracts features from the input, while $B$ uses these features to create the desired output. Based on this observation, we demonstrate that fine-tuning $B$ is inherently more effective than fine-tuning $A$, and that a random untrained $A$ should perform nearly as well as a fine-tuned one. Using an information-theoretic lens, we also bound the generalization of low-rank adapters, showing that the parameter savings of exclusively training $B$ improves the bound. We support our conclusions with experiments on RoBERTa, BART-Large, LLaMA-2, and ViTs. The code and data is available at https://github.com/Jiacheng-Zhu-AIML/AsymmetryLoRA Kristjan Greenewald, Kimia Nadjahi, Haitz Sáez de Ocáriz Borde, Rickard Brüel Gabrielsson, Leshem Choshen, Marzyeh Ghassemi, Mikhail Yurochkin, Justin Solomon 0001 |
ICML | 7 |
| 2024 | FedMedICL: Towards Holistic Evaluation of Distribution Shifts in Federated Medical Imaging
Kumail Alhamoud, Yasir Ghunaim, Motasem Alfarra, Thomas Hartvigsen, Philip Torr 0001, Bernard Ghanem, Adel Bibi, Marzyeh Ghassemi |
MICCAI (10) | 8 |
| 2024 | MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingabstractFoundation models are becoming valuable tools in medicine. Yet despite their promise, the best way to leverage Large Language Models (LLMs) in complex medical tasks remains an open question. We introduce a novel multi-agent framework, named **M**edical **D**ecision-making **Agents** (**MDAgents**) that helps to address this gap by automatically assigning a collaboration structure to a team of LLMs. The assigned solo or group collaboration structure is tailored to the medical task at hand, a simple emulation inspired by the way real-world medical decision-making processes are adapted to tasks of different complexities. We evaluate our framework and baseline methods using state-of-the-art LLMs across a suite of real-world medical knowledge and clinical diagnosis benchmarks, including a comparison of
LLMs’ medical complexity classification against human physicians. MDAgents achieved the **best performance in seven out of ten** benchmarks on tasks requiring an understanding of medical knowledge and multi-modal reasoning, showing a significant **improvement of up to 4.2\%** ($p$ < 0.05) compared to previous methods' best performances. Ablation studies reveal that MDAgents effectively determines medical complexity to optimize for efficiency and accuracy across diverse medical tasks. Notably, the combination of moderator review and external medical knowledge in group collaboration resulted in an average accuracy **improvement of 11.8\%**. Our code can be found at https://github.com/mitmedialab/MDAgents. Yubin Kim 0002, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, Hae Won Park 0001 |
NeurIPS | 8 |
| 2024 | BendVLM: Test-Time Debiasing of Vision-Language EmbeddingsabstractVision-language (VL) embedding models have been shown to encode biases present in their training data, such as societal biases that prescribe negative characteristics to members of various racial and gender identities. Due to their wide-spread adoption for various tasks ranging from few-shot classification to text-guided image generation, debiasing VL models is crucial. Debiasing approaches that fine-tune the VL model often suffer from catastrophic forgetting. On the other hand, fine-tuning-free methods typically utilize a ``one-size-fits-all" approach that assumes that correlation with the spurious attribute can be explained using a single linear direction across all possible inputs. In this work, we propose a nonlinear, fine-tuning-free approach for VL embedding model debiasing that tailors the debiasing operation to each unique input. This allows for a more flexible debiasing approach. Additionally, we do not require knowledge of the set of inputs a priori to inference time, making our method more appropriate for online tasks such as retrieval and text guided image generation. Walter Gerych, Haoran Zhang 0003, Kimia Hamidieh, Eileen Pan, Maanas Kumar Sharma, Thomas Hartvigsen, Marzyeh Ghassemi |
NeurIPS | 7 |
| 2024 | Improving Subgroup Robustness via Data SelectionabstractMachine learning models can often fail on subgroups that are underrepresented
during training. While dataset balancing can improve performance on
underperforming groups, it requires access to training group annotations and can
end up removing large portions of the dataset. In this paper, we introduce
Data Debiasing with Datamodels (D3M), a debiasing approach
which isolates and removes specific training examples that drive the model's
failures on minority groups. Our approach enables us to efficiently train
debiased classifiers while removing only a small number of examples, and does
not require training group annotations or additional hyperparameter tuning. Saachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas, Marzyeh Ghassemi, Aleksander Madry |
NeurIPS | 5 |
| 2024 | Large language models in biomedicine and health: current research landscape and future directionsabstractLarge language models in biomedicine and health: current research landscape and future directionsLarge language models (LLMs) are a specialized type of generative artificial intelligence (AI) focused on generating natural language text.These models are developed through extensive training on massive amounts of text data and use deep learning algorithms to generate new text that closely resembles human-generated text.Generative AI methods, including LLMs, are rapidly transforming various domains, including biomedicine and healthcare.[1][2][3][4][5][6] They have already demonstrated remarkable potential as a means to process and analyze large amounts of text, interpret natural language, and generate new content in these domains.For example, Nori et al reported that GPT-4 is able to correctly answer the majority of questions from medical practice licensing exams, comfortably obtaining a passing grade.7 Similarly, Stribling et al found that this model exceeded the average performance of students in the graduate medical sciences on the majority of examinations, including strong performance on short answer and essay questions.8 Even though passing the exam is not the same as applying the knowledge in a real-world setting, these results demonstrate that LLMs can generate appropriate multiple-choice and narrative responses to questions framed in natural language.ChatGPT, first released in November 2022, has garnered phenomenal attention from both the scientific community and a broader society.A keyword search of "large language models" OR "ChatGPT" in PubMed returned over 4500 articles that discuss the technology and its implications for various topics, including medical informatics, by the end of June 2024.In addition, LLM-based technologies have already been deployed in several healthcare systems and are offered as integrated products for use in the clinic within vendor electronic health record systems (for thoughts on initial evaluations of an early product, see Garcia et al 9 and Tai-Seale et al 10 ).This rapid adoption of LLMs like ChatGPT brings an unprecedented opportunity to use this novel AI technology to transform healthcare and medicine.Despite their potential benefits, LLMs can sometimes produce invalid and unsubstantiated responses, a phenomenon known as the "hallucination and confabulation issue" in the literature, or biased responses, due to the biases inherent in their training data.[11][12][13][14][15][16][17] With this great potential also comes the need for trustworthy and responsible development and use of technology.As we continue to explore the capabilities of ChatGPT and other LLMs, it is critical to address related ethical, legal, and social issues to ensure that the technology is used in ways that are safe, fair, trustworthy, and beneficial for all.In the context of biomedicine and healthcare, it is particularly important to engage stakeholders, such as AI researchers, developers of data-driven clinical decision support, care providers, and system implementers from both academic medical centers and industry, to ensure responsible use of LLMs for good.To accelerate research and development in this area, we issued a call for submissions in Summer 2023, specifically focusing on the intersection of biomedicine/health and LLMs, and invited contributions on all related aspects.We invited submissions that report on innovative informatics methods development and evaluation, as well as studies that demonstrate the effectiveness/limitations of LLMs methodologies in healthcare.We particularly encouraged submissions that address the challenges and opportunities of this intersection and offer new insights into how these fields can work together to advance healthcare.This editorial provides an overview of the papers accepted in this Focus Issue.We highlight major themes and unique aspects of the research papers in medical LLMs, discuss ongoing challenges, and recommend future research directions.Box 1 lists the relevant large language model terms and abbreviations used in this editorial. Overall statistics of the Focus IssueThis JAMIA Focus Issue on LLMs in biomedicine and health has drawn enthusiasm from many researchers across different research disciplines.In total, we received over 150 submissions from authors in 25 countries and regions across 6 continents worldwide.The rigorous JAMIA peer review process was applied to all submissions, 41 of which were ultimately accepted for publication in the Focus Issue (Table 1).The majority of the accepted papers were authored by authors in North America, followed by those in Asia and Europe (Figure 1A).The Focus Issue highlights the nature of multi-disciplinary collaboration in medical informatics research across the broad JAMIA community.The number of authors per paper varies from 1 to 23, with an average of 7.3.Many papers feature authors with diverse expertise from different departments and organizations.The authors' expertise spans a wide range of fields, including computer science, data science, informatics, statistics, medicine, nursing, clinical services, public health policies, and more.Several papers also demonstrate scientific collaborations across different sectors, including academia, government labs, research institutes, hospitals, and industry.Additionally, a few papers showcase international collaborations among authors. Zhiyong Lu, Yifan Peng 0002, Trevor Cohen, Marzyeh Ghassemi, Chunhua Weng, Shubo Tian |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Evaluating the Impact of Social Determinants on Health Prediction in the Intensive Care UnitabstractSocial determinants of health (SDOH) – the conditions in which people live, grow, and age – play a crucial role in a person’s health and well-being. There is a large, compelling body of evidence in population health studies showing that a wide range of SDOH is strongly correlated with health outcomes. Yet, a majority of the risk prediction models based on electronic health records (EHR) do not incorporate a comprehensive set of SDOH features as they are often noisy or simply unavailable. Our work links a publicly available EHR database, MIMIC-IV, to well-documented SDOH features. We investigate the impact of such features on common EHR prediction tasks across different patient populations. We find that community-level SDOH features do not improve model performance for a general patient population, but can improve data-limited model fairness for specific subpopulations. We also demonstrate that SDOH features are vital for conducting thorough audits of algorithmic biases beyond protective attributes. We hope the new integrated EHR-SDOH database will enable studies on the relationship between community health and individual outcomes and provide new benchmarks to study algorithmic biases beyond race, gender, and age. Ming-Ying Yang, Gloria Hyun-Jung Kwak, Tom J. Pollard, Leo A. Celi, Marzyeh Ghassemi |
AIES | 5 |
| 2023 | When Personalization Harms Performance: Reconsidering the Use of Group Attributes in PredictionabstractMachine learning models are often personalized with categorical attributes that define groups. In this work, we show that personalization with *group attributes* can inadvertently reduce performance at a *group level* -- i.e., groups may receive unnecessarily inaccurate predictions by sharing their personal characteristics. We present formal conditions to ensure the *fair use* of group attributes in a prediction task, and describe how they can be checked by training one additional model. We characterize how fair use conditions be violated due to standard practices in model development, and study the prevalence of fair use violations in clinical prediction tasks. Our results show that personalization often fails to produce a tailored performance gain for every group who reports personal data, and underscore the need to evaluate fair use when personalizing models with characteristics that are protected, sensitive, self-reported, or costly to acquire. Vinith M. Suriyakumar, Marzyeh Ghassemi, Berk Ustun |
ICML | 2 |
| 2023 | Change is Hard: A Closer Look at Subpopulation ShiftabstractMachine learning models often perform poorly on subgroups that are underrepresented in the training data. Yet, little is understood on the variation in mechanisms that cause subpopulation shifts, and how algorithms generalize across such diverse shifts at scale. In this work, we provide a fine-grained analysis of subpopulation shift. We first propose a unified framework that dissects and explains common shifts in subgroups. We then establish a comprehensive benchmark of 20 state-of-the-art algorithms evaluated on 12 real-world datasets in vision, language, and healthcare domains. With results obtained from training over 10,000 models, we reveal intriguing observations for future progress in this space. First, existing algorithms only improve subgroup robustness over certain types of shifts but not others. Moreover, while current algorithms rely on group-annotated validation data for model selection, we find that a simple selection criterion based on worst-class accuracy is surprisingly effective even without any group information. Finally, unlike existing works that solely aim to improve worst-group accuracy (WGA), we demonstrate the fundamental tradeoff between WGA and other important metrics, highlighting the need to carefully choose testing metrics. Code and data are available at: https://github.com/YyzHarry/SubpopBench. Yuzhe Yang 0003, Haoran Zhang 0003, Dina Katabi, Marzyeh Ghassemi |
ICML | 4 |
| 2023 | "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsabstractMachine learning models frequently experience performance drops under distribution shifts. The underlying cause of such shifts may be multiple simultaneous factors such as changes in data quality, differences in specific covariate distributions, or changes in the relationship between label and features. When a model does fail during deployment, attributing performance change to these factors is critical for the model developer to identify the root cause and take mitigating actions. In this work, we introduce the problem of attributing performance differences between environments to distribution shifts in the underlying data generating mechanisms. We formulate the problem as a cooperative game where the players are distributions. We define the value of a set of distributions to be the change in model performance when only this set of distributions has changed between environments, and derive an importance weighting method for computing the value of an arbitrary set of distributions. The contribution of each distribution to the total performance change is then quantified as its Shapley value. We demonstrate the correctness and utility of our method on synthetic, semi-synthetic, and real-world case studies, showing its effectiveness in attributing performance changes to a wide range of distribution shifts. Haoran Zhang 0003, Harvineet Singh, Marzyeh Ghassemi, Shalmali Joshi |
ICML | 3 |
| 2023 | Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsabstractDeployed language models decay over time due to shifting inputs, changing user needs, or emergent world-knowledge gaps. When such problems are identified, we want to make targeted edits while avoiding expensive retraining. However, current model editors, which modify such behaviors of pre-trained models, degrade model performance quickly across multiple, sequential edits. We propose GRACE, a \textit{lifelong} model editing method, which implements spot-fixes on streaming errors of a deployed model, ensuring minimal impact on unrelated inputs. GRACE writes new mappings into a pre-trained model's latent space, creating a discrete, local codebook of edits without altering model weights. This is the first method enabling thousands of sequential edits using only streaming errors. Our experiments on T5, BERT, and GPT models show GRACE's state-of-the-art performance in making and retaining edits, while generalizing to unseen inputs. Our code is available at [github.com/thartvigsen/grace](https://www.github.com/thartvigsen/grace}). Thomas Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Marzyeh Ghassemi |
NeurIPS | 5 |
| 2023 | VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual PerceptionabstractAI alignment refers to models acting towards human-intended goals, preferences, or ethical principles. Analyzing the similarity between models and humans can be a proxy measure for ensuring AI safety. In this paper, we focus on the models' visual perception alignment with humans, further referred to as AI-human visual alignment. Specifically, we propose a new dataset for measuring AI-human visual alignment in terms of image classification. In order to evaluate AI-human visual alignment, a dataset should encompass samples with various scenarios and have gold human perception labels. Our dataset consists of three groups of samples, namely Must-Act (i.e., Must-Classify), Must-Abstain, and Uncertain, based on the quantity and clarity of visual information in an image and further divided into eight categories. All samples have a gold human perception label; even Uncertain (e.g., severely blurry) sample labels were obtained via crowd-sourcing. The validity of our dataset is verified by sampling theory, statistical theories related to survey design, and experts in the related fields. Using our dataset, we analyze the visual alignment and reliability of five popular visual perception models and seven abstention methods. Our code and data is available at https://github.com/jiyounglee-0523/VisAlign. Seungho Kim, Seunghyun Won, Joonseok Lee, Marzyeh Ghassemi, James Thorne, Jaeseok Choi, O.-Kil Kwon, Edward Choi 0003 |
NeurIPS | 5 |
| 2022 | Write It Like You See It: Detectable Differences in Clinical Notes by Race Lead to Differential Model RecommendationsabstractClinical notes are becoming an increasingly important data source for machine learning (ML) applications in healthcare. Prior research has shown that deploying ML models can perpetuate existing biases against racial minorities, as bias can be implicitly embedded in data. In this study, we investigate the level of implicit race information available to ML models and human experts and the implications of model-detectable differences in clinical notes. Our work makes three key contributions. First, we find that models can identify patient self-reported race from clinical notes even when the notes are stripped of explicit indicators of race. Second, we determine that human experts are not able to accurately predict patient race from the same redacted clinical notes. Finally, we demonstrate the potential harm of this implicit information in a simulation study, and show that models trained on these race-redacted clinical notes can still perpetuate existing biases in clinical treatment decisions. Hammaad Adam, Ming-Ying Yang, Kenrick Cato, Ioana Baldini, Charles Senteio, Leo A. Celi, Jiaming Zeng, Moninder Singh, Marzyeh Ghassemi |
AIES | 9 |
| 2022 | Get To The Point! Problem-Based Curated Data Views To Augment Care For Critically Ill PatientsabstractElectronic health records in critical care medicine offer unprecedented opportunities for clinical reasoning and decision making. Paradoxically, these data-rich environments have also resulted in clinical decision support systems (CDSSs) that fit poorly into clinical contexts, and increase health workers cognitive load. In this paper, we introduce a novel approach to designing CDSSs that are embedded in clinical workflows, by presenting problem-based curated data views tailored for problem-driven discovery, team communication, and situational awareness. We describe the design and evaluation of one such CDSS, In-Sight, that embodies our approach and addresses the clinical problem of monitoring critically ill pediatric patients. Our work is the result of a co-design process, further informed by empirical data collected through formal usability testing, focus groups, and a simulation study with domain experts. We discuss the potential and limitations of our approach, and share lessons learned in our iterative co-design process. Minfan Zhang, Daniel Ehrmann, Mjaye Mazwi, Danny Eytan, Marzyeh Ghassemi, Fanny Chevalier |
CHI | 5 |
| 2022 | Understanding the Variance Collapse of SVGD in High Dimensions
Jimmy Ba, Murat A. Erdogdu, Marzyeh Ghassemi, Shengyang Sun, Taiji Suzuki, Denny Wu, Tianzong Zhang |
ICLR | 3 |
| 2022 | Improving Mutual Information Estimation with Annealed and Energy-Based Bounds
Rob Brekelmans, Sicong Huang 0001, Marzyeh Ghassemi, Greg Ver Steeg, Roger B. Grosse, Alireza Makhzani |
ICLR | 3 |
| 2022 | Is Fairness Only Metric Deep? Evaluating and Addressing Subgroup Gaps in Deep Metric Learning
Natalie Dullerud, Karsten Roth, Kimia Hamidieh, Nicolas Papernot, Marzyeh Ghassemi |
ICLR | 5 |
| 2022 | If Influence Functions are the Answer, Then What is the Question?abstractInfluence functions efficiently estimate the effect of removing a single training data point on a model's learned parameters. While influence estimates align well with leave-one-out retraining for linear models, recent works have shown this alignment is often poor in neural networks. In this work, we investigate the specific factors that cause this discrepancy by decomposing it into five separate terms. We study the contributions of each term on a variety of architectures and datasets and how they vary with factors such as network width and training time. While practical influence function estimates may be a poor match to leave-one-out retraining for nonlinear networks, we show that they are often a good approximation to a different object we term the proximal Bregman response function (PBRF). Since the PBRF can still be used to answer many of the questions motivating influence functions, such as identifying influential or mislabeled examples, our results suggest that current algorithms for influence function estimation give more informative results than previous error analyses would suggest. Juhan Bae, Nathan Ng 0001, Alston Lo, Marzyeh Ghassemi, Roger B. Grosse |
NeurIPS | 4 |
| 2021 | Making Health AI Work in the Real World: Strategies, innovations, and best practices for using AI to improve care delivery
Suchi Saria, Marzyeh Ghassemi, Ziad Obermeyer, Karandeep Singh, Pei-Yun S. Hsueh, Eric J. Topol |
AMIA | 2 |
| 2021 | Pulling Up by the Causal Bootstraps: Causal Data Augmentation for Pre-training DebiasingabstractMachine learning models achieve state-of-the-art performance on many supervised learning tasks. However, prior evidence suggests that these models may learn to rely on "shortcut" biases or spurious correlations (intuitively, correlations that do not hold in the test as they hold in train) for good predictive performance. Such models cannot be trusted in deployment environments to provide accurate predictions. While viewing the problem from a causal lens is known to be useful, the seamless integration of causation techniques into machine learning pipelines remains cumbersome and expensive. In this work, we study and extend a causal pre-training debiasing technique called causal bootstrapping (CB) under five practical confounded-data generation-acquisition scenarios (with known and unknown confounding). Under these settings, we systematically investigate the effect of confounding bias on deep learning model performance, demonstrating their propensity to rely on shortcut biases when these biases are not properly accounted for. We demonstrate that such a causal pre-training technique can significantly outperform existing base practices to mitigate confounding bias on real-world domain generalization benchmarking tasks. This systematic investigation underlines the importance of accounting for the underlying data-generating mechanisms and fortifying data-preprocessing pipelines with a causal framework to develop methods robust to confounding biases. Sindhu C. M. Gowda, Shalmali Joshi, Haoran Zhang 0003, Marzyeh Ghassemi |
CIKM | 4 |
| 2021 | Simultaneous Similarity-based Self-Distillation for Deep Metric LearningabstractDeep Metric Learning (DML) provides a crucial tool for visual similarity and zero-shot retrieval applications by learning generalizing embedding spaces, although recent work in DML has shown strong performance saturation across training objectives. However, generalization capacity is known to scale with the embedding space dimensionality. Unfortunately, high dimensional embeddings also create higher retrieval cost for downstream applications. To remedy this, we propose S2SD - Simultaneous Similarity-based Self-distillation. S2SD extends DML with knowledge distillation from auxiliary, high-dimensional embedding and feature spaces to leverage complementary context during training while retaining test-time cost and with negligible changes to the training time. Experiments and ablations across different objectives and standard benchmarks show S2SD offering highly significant improvements of up to 7% in Recall@1, while also setting a new state-of-the-art. Karsten Roth, Timo Milbich, Björn Ommer, Joseph Paul Cohen, Marzyeh Ghassemi |
ICML | 5 |
| 2021 | Medical Dead-ends and Learning to Identify High-Risk States and TreatmentsabstractMachine learning has successfully framed many sequential decision making problems as either supervised prediction, or optimal decision-making policy identification via reinforcement learning. In data-constrained offline settings, both approaches may fail as they assume fully optimal behavior or rely on exploring alternatives that may not exist. We introduce an inherently different approach that identifies "dead-ends" of a state space. We focus on patient condition in the intensive care unit, where a "medical dead-end" indicates that a patient will expire, regardless of all potential future treatment sequences. We postulate "treatment security" as avoiding treatments with probability proportional to their chance of leading to dead-ends, present a formal proof, and frame discovery as an RL problem. We then train three independent deep neural models for automated state construction, dead-end discovery and confirmation. Our empirical results discover that dead-ends exist in real clinical data among septic patients, and further reveal gaps between secure treatments and those administered. Mehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, Marzyeh Ghassemi |
NeurIPS | 4 |
| 2021 | Characterizing Generalization under Out-Of-Distribution Shifts in Deep Metric LearningabstractDeep Metric Learning (DML) aims to find representations suitable for zero-shot transfer to a priori unknown test distributions. However, common evaluation protocols only test a single, fixed data split in which train and test classes are assigned randomly. More realistic evaluations should consider a broad spectrum of distribution shifts with potentially varying degree and difficulty.In this work, we systematically construct train-test splits of increasing difficulty and present the ooDML benchmark to characterize generalization under out-of-distribution shifts in DML. ooDML is designed to probe the generalization performance on much more challenging, diverse train-to-test distribution shifts. Based on our new benchmark, we conduct a thorough empirical analysis of state-of-the-art DML methods. We find that while generalization tends to consistently degrade with difficulty, some methods are better at retaining performance as the distribution shift increases. Finally, we propose few-shot DML as an efficient way to consistently improve generalization in response to unknown test shifts presented in ooDML. Timo Milbich, Karsten Roth, Samarth Sinha, Ludwig Schmidt, Marzyeh Ghassemi, Björn Ommer |
NeurIPS | 5 |
| 2021 | Learning Optimal Predictive ChecklistsabstractChecklists are simple decision aids that are often used to promote safety and reliability in clinical applications. In this paper, we present a method to learn checklists for clinical decision support. We represent predictive checklists as discrete linear classifiers with binary features and unit weights. We then learn globally optimal predictive checklists from data by solving an integer programming problem. Our method allows users to customize checklists to obey complex constraints, including constraints to enforce group fairness and to binarize real-valued features at training time. In addition, it pairs models with an optimality gap that can inform model development and determine the feasibility of learning sufficiently accurate checklists on a given dataset. We pair our method with specialized techniques that speed up its ability to train a predictive checklist that performs well and has a small optimality gap. We benchmark the performance of our method on seven clinical classification problems, and demonstrate its practical benefits by training a short-form checklist for PTSD screening. Our results show that our method can fit simple predictive checklists that perform well and that can easily be customized to obey a rich class of custom constraints. Haoran Zhang 0003, Quaid Morris, Berk Ustun, Marzyeh Ghassemi |
NeurIPS | 4 |
| 2020 | SSMBA: Self-Supervised Manifold Based Data Augmentation for Improving Out-of-Domain RobustnessabstractModels that perform well on a training domain often fail to generalize to out-of-domain (OOD) examples.Data augmentation is a common method used to prevent overfitting and improve OOD generalization.However, in natural language, it is difficult to generate new examples that stay on the underlying data manifold.We introduce SSMBA, a data augmentation method for generating synthetic training examples by using a pair of corruption and reconstruction functions to move randomly on a data manifold.We investigate the use of SSMBA in the natural language domain, leveraging the manifold assumption to reconstruct corrupted text with masked language models.In experiments on robustness benchmarks across 3 tasks and 9 datasets, SSMBA consistently outperforms existing data augmentation methods and baseline models on both in-domain and OOD data, achieving gains of 0.8% accuracy on OOD Amazon reviews, 1.8% accuracy on OOD MNLI, and 1.4 BLEU on in-domain IWSLT14 German-English. 1 Nathan Ng 0001, Kyunghyun Cho, Marzyeh Ghassemi |
EMNLP (1) | 3 |
| 2019 | The Cells Out of Sample (COOS) dataset and benchmarks for measuring out-of-sample generalization of image classifiersabstractUnderstanding if classifiers generalize to out-of-sample datasets is a central problem in machine learning. Microscopy images provide a standardized way to measure the generalization capacity of image classifiers, as we can image the same classes of objects under increasingly divergent, but controlled factors of variation. We created a public dataset of 132,209 images of mouse cells, COOS-7 (Cells Out Of Sample 7-Class). COOS-7 provides a classification setting where four test datasets have increasing degrees of covariate shift: some images are random subsets of the training data, while others are from experiments reproduced months later and imaged by different instruments. We benchmarked a range of classification models using different representations, including transferred neural network features, end-to-end classification with a supervised deep CNN, and features from a self-supervised CNN. While most classifiers perform well on test datasets similar to the training dataset, all classifiers failed to generalize their performance to datasets with greater covariate shifts. These baselines highlight the challenges of covariate shifts in image data, and establish metrics for improving the generalization capacity of image classifiers. Alex Lu 0002, Amy X. Lu, Wiebke Schormann, Marzyeh Ghassemi, David W. Andrews, Alan M. Moses |
NeurIPS | 4 |
| 2018 | Semi-Supervised Biomedical Translation With Cycle Wasserstein Regression GANsabstractThe biomedical field offers many learning tasks that share unique challenges: large amounts of unpaired data, and a high cost to generate labels. In this work, we develop a method to address these issues with semi-supervised learning in regression tasks (e.g., translation from source to target). Our model uses adversarial signals to learn from unpaired datapoints, and imposes a cycle-loss reconstruction error penalty to regularize mappings in either direction against one another. We first evaluate our method on synthetic experiments, demonstrating two primary advantages of the system: 1) distribution matching via the adversarial loss and 2) regularization towards invertible mappings via the cycle loss. We then show a regularization effect and improved performance when paired data is supplemented by additional unpaired data on two real biomedical regression tasks: estimating the physiological effect of medical treatments, and extrapolating gene expression (transcriptomics) signals. Our proposed technique is a promising initial step towards more robust use of adversarial signals in semi-supervised regression, and could be useful for other tasks (e.g., causal inference or modality translation) in the biomedical field. Matthew B. A. McDermott, Tom Yan, Tristan Naumann, Nathan Hunt, Harini Suresh, Peter Szolovits, Marzyeh Ghassemi |
AAAI | 7 |
| 2017 | Understanding vasopressor intervention and weaning: risk prediction in a public heterogeneous clinical time series databaseabstractBACKGROUND: The widespread adoption of electronic health records allows us to ask evidence-based questions about the need for and benefits of specific clinical interventions in critical-care settings across large populations. OBJECTIVE: We investigated the prediction of vasopressor administration and weaning in the intensive care unit. Vasopressors are commonly used to control hypotension, and changes in timing and dosage can have a large impact on patient outcomes. MATERIALS AND METHODS: We considered a cohort of 15 695 intensive care unit patients without orders for reduced care who were alive 30 days post-discharge. A switching-state autoregressive model (SSAM) was trained to predict the multidimensional physiological time series of patients before, during, and after vasopressor administration. The latent states from the SSAM were used as predictors of vasopressor administration and weaning. RESULTS: The unsupervised SSAM features were able to predict patient vasopressor administration and successful patient weaning. Features derived from the SSAM achieved areas under the receiver operating curve of 0.92, 0.88, and 0.71 for predicting ungapped vasopressor administration, gapped vasopressor administration, and vasopressor weaning, respectively. We also demonstrated many cases where our model predicted weaning well in advance of a successful wean. CONCLUSION: Models that used SSAM features increased performance on both predictive tasks. These improvements may reflect an underlying, and ultimately predictive, latent state detectable from the physiological time series. Mike Wu, Marzyeh Ghassemi, Mengling Feng, Leo A. Celi, Peter Szolovits, Finale Doshi-Velez |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | A Multivariate Timeseries Modeling Approach to Severity of Illness Assessment and Forecasting in ICU with Sparse, Heterogeneous Clinical DataabstractThe ability to determine patient acuity (or severity of illness) has immediate practical use for clinicians. We evaluate the use of multivariate timeseries modeling with the multi-task Gaussian process (GP) models using noisy, incomplete, sparse, heterogeneous and unevenly-sampled clinical data, including both physiological signals and clinical notes. The learned multi-task GP (MTGP) hyperparameters are then used to assess and forecast patient acuity. Experiments were conducted with two real clinical data sets acquired from ICU patients: firstly, estimating cerebrovascular pressure reactivity, an important indicator of secondary damage for traumatic brain injury patients, by learning the interactions between intracranial pressure and mean arterial blood pressure signals, and secondly, mortality prediction using clinical progress notes. In both cases, MTGPs provided improved results: an MTGP model provided better results than single-task GP models for signal interpolation and forecasting (0.91 vs 0.69 RMSE), and the use of MTGP hyperparameters obtained improved results when used as additional classification features (0.812 vs 0.788 AUC). Marzyeh Ghassemi, Marco A. F. Pimentel, Tristan Naumann, Thomas Brennan, David A. Clifton, Peter Szolovits, Mengling Feng |
AAAI | 1 |
| 2014 | Unfolding physiological state: mortality modelling in intensive care unitsabstractAccurate knowledge of a patient's disease state and trajectory is critical in a clinical setting. Modern electronic healthcare records contain an increasingly large amount of data, and the ability to automatically identify the factors that influence patient outcomes stand to greatly improve the efficiency and quality of care. We examined the use of latent variable models (viz. Latent Dirichlet Allocation) to decompose free-text hospital notes into meaningful features, and the predictive power of these features for patient mortality. We considered three prediction regimes: (1) baseline prediction, (2) dynamic (time-varying) outcome prediction, and (3) retrospective outcome prediction. In each, our prediction task differs from the familiar time-varying situation whereby data accumulates; since fewer patients have long ICU stays, as we move forward in time fewer patients are available and the prediction task becomes increasingly difficult. We found that latent topic-derived features were effective in determining patient mortality under three timelines: inhospital, 30 day post-discharge, and 1 year post-discharge mortality. Our results demonstrated that the latent topic features important in predicting hospital mortality are very different from those that are important in post-discharge mortality. In general, latent topic features were more predictive than structured features, and a combination of the two performed best. The time-varying models that combined latent topic features and baseline features had AUCs that reached 0.85, 0.80, and 0.77 for in-hospital, 30 day post-discharge and 1 year post-discharge mortality respectively. Our results agreed with other work suggesting that the first 24 hours of patient information are often the most predictive of hospital mortality. Retrospective models that used a combination of latent topic features and structured features achieved AUCs of 0.96, 0.82, and 0.81 for in-hospital, 30 day, and 1-year mortality prediction. Our work focuses on the dynamic (time-varying) setting because models from this regime could facilitate an on-going severity stratification system that helps direct care-staff resources and inform treatment strategies. Marzyeh Ghassemi, Tristan Naumann, Finale Doshi-Velez, Nicole Brimmer, Rohit Joshi, Anna Rumshisky, Peter Szolovits |
KDD | 1 |
| 2013 | Probabilistically Populated Medical Record Templates: Reducing Clinical Documentation Time Using Patient Cooperation
Tristan Naumann, Marzyeh Ghassemi, Andreea Bodnari, Rohit Joshi |
AMIA | 2 |