EDBT 2026 Demo / reviewers in the wild / expert
Matthew M. Churpek
dblp:190/9704
· DBLP profile ↗
21ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0002-4030-5250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 17 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable multimodal deep learning models for variable-length sequences in critically ill patients
Jennifer Martin, Majid Afshar, Askar Safipour Afshar, John R. Caskey, Dmitriy Dligach, Yanjun Gao, Jifan Gao, Guanhua Chen 0002, Anoop M. Mayampurath, Matthew M. Churpek |
J. Biomed. Informatics | 10 |
| 2025 | Toward digital twins in the intensive care unit: a medication management case studyabstractOBJECTIVE: To evaluate the efficacy of digital twins developed using a large language model (LLaMA-3), fine-tuned with Low-Rank Adapters (LoRA) on intensive care units (ICU) physician notes, and to determine whether specialty-specific training enhances treatment recommendation accuracy compared to other ICU specialties or zero-shot baselines. MATERIALS AND METHODS: Digital twins were created using LLaMA-3 fine-tuned on discharge summaries from the Medical Information Mart for Intensive Care III dataset, where medications were masked to construct training and testing datasets. The medical ICU dataset (1000 notes) was used for evaluation, and performance was assessed using Bidirectional Encoder Representations from Transformers Score (BERTScore) and ROUGE-L. A zero-shot baseline model, relying solely on contextual instructions without training, was also evaluated. While our approach moves toward digital twin capabilities, it does not incorporate real-time, patient-specific electronic health records data and can be viewed as an ICU specialty-level language model adaptation. RESULTS: Models fine-tuned on medical ICU notes achieved the highest BERTScore (0.842), outperforming models trained on other specialties or mixed datasets. Zero-shot models showed the lowest performance, highlighting the importance of training. DISCUSSION: The findings demonstrate that specialty-specific training significantly improves treatment recommendation accuracy in digital twins compared to generalized or zero-shot approaches. Tailoring models to specific ICU domains strengthens their clinical decision-support capabilities. CONCLUSION: Context-specific fine-tuning of LLMs is crucial for developing effective digital twins, offering foundational insights for personalized clinical decision support. Behnaz Eslami, Majid Afshar, Samie Tootooni, Timothy A. Miller, Matthew M. Churpek, Yanjun Gao, Dmitriy Dligach |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Lessons learned on information retrieval in electronic health records: a comparison of embedding models and pooling strategiesabstractOBJECTIVES: Applying large language models (LLMs) to the clinical domain is challenging due to the context-heavy nature of processing medical records. Retrieval-augmented generation (RAG) offers a solution by facilitating reasoning over large text sources. However, there are many parameters to optimize in just the retrieval system alone. This paper presents an ablation study exploring how different embedding models and pooling methods affect information retrieval for the clinical domain. MATERIALS AND METHODS: Evaluating on 3 retrieval tasks on 2 electronic health record (EHR) data sources, we compared 7 models, including medical- and general-domain models, specialized encoder embedding models, and off-the-shelf decoder LLMs. We also examine the choice of embedding pooling strategy for each model, independently on the query and the text to retrieve. RESULTS: We found that the choice of embedding model significantly impacts retrieval performance, with BGE, a comparatively small general-domain model, consistently outperforming all others, including medical-specific models. However, our findings also revealed substantial variability across datasets and query text phrasings. We also determined the best pooling methods for each of these models to guide future design of retrieval systems. DISCUSSION: The choice of embedding model, pooling strategy, and query formulation can significantly impact retrieval performance and the performance of these models on other public benchmarks does not necessarily transfer to new domains. The high variability in performance across different query phrasings suggests that the choice of query may need to be tuned and validated for each task, or even for each institution's EHR. CONCLUSION: This study provides empirical evidence to guide the selection of models and pooling strategies for RAG frameworks in healthcare applications. Further studies such as this one are vital for guiding empirically-grounded development of retrieval frameworks, such as in the context of RAG, for the clinical domain. Skatje Myers, Timothy A. Miller, Yanjun Gao, Matthew M. Churpek, Anoop M. Mayampurath, Dmitriy Dligach, Majid Afshar |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | Explaining alerts from a pediatric risk prediction model using clinical textabstractOBJECTIVE: Risk prediction models are used in hospitals to identify pediatric patients at risk of clinical deterioration, enabling timely interventions and rescue. The objective of this study was to develop a new explainer algorithm that uses a patient's clinical notes to generate text-based explanations for risk prediction alerts. MATERIALS AND METHODS: We conducted a retrospective study of 39 406 patient admissions to the American Family Children's Hospital at the University of Wisconsin-Madison (2009-2020). The pediatric Calculated Assessment of Risk and Triage (pCART) validated risk prediction model was used to identify children at risk for deterioration. A transformer model was trained to use clinical notes from the 12-hour period preceding each pCART score to predict whether a patient was flagged as at risk. Then, label-aware attention highlighted text phrases most important to an at-risk alert. The study cohort was randomly split into derivation (60%) and validation (20%) data, and a separate test (20%) was used to evaluate the explainer's performance. RESULTS: Our pCART Explainer algorithm performed well in discriminating at-risk pCART alert vs no alert (c-statistic 0.805). Sample explanations from pCART Explainer revealed clinically important phrases such as "rapid breathing," "fall risk," "distension," and "grunting," thereby demonstrating excellent face validity. DISCUSSION: The pCART Explainer could quickly orient clinicians to the patient's condition by drawing attention to key phrases in notes, potentially enhancing situational awareness and guiding decision-making. CONCLUSION: We developed pCART Explainer, a novel algorithm that highlights text within clinical notes to provide medically relevant context about deterioration alerts, thereby improving the explainability of the pCART model. Samuel Nycklemoe, Sriharsha Devarapu, Yanjun Gao, Kyle A. Carey, Nicholas Kuehnel, Neil Munjal, Priti Jani, Matthew M. Churpek, Dmitriy Dligach, Majid Afshar, Anoop M. Mayampurath |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | Development and external validation of deep learning clinical prediction models using variable-length time series dataabstractOBJECTIVES: To compare and externally validate popular deep learning model architectures and data transformation methods for variable-length time series data in 3 clinical tasks (clinical deterioration, severe acute kidney injury [AKI], and suspected infection). MATERIALS AND METHODS: This multicenter retrospective study included admissions at 2 medical centers that spanned 2007-2022. Distinct datasets were created for each clinical task, with 1 site used for training and the other for testing. Three feature engineering methods (normalization, standardization, and piece-wise linear encoding with decision trees [PLE-DTs]) and 3 architectures (long short-term memory/gated recurrent unit [LSTM/GRU], temporal convolutional network, and time-distributed wrapper with convolutional neural network [TDW-CNN]) were compared in each clinical task. Model discrimination was evaluated using the area under the precision-recall curve (AUPRC) and the area under the receiver operating characteristic curve (AUROC). RESULTS: The study comprised 373 825 admissions for training and 256 128 admissions for testing. LSTM/GRU models tied with TDW-CNN models with both obtaining the highest mean AUPRC in 2 tasks, and LSTM/GRU had the highest mean AUROC across all tasks (deterioration: 0.81, AKI: 0.92, infection: 0.87). PLE-DT with LSTM/GRU achieved the highest AUPRC in all tasks. DISCUSSION: When externally validated in 3 clinical tasks, the LSTM/GRU model architecture with PLE-DT transformed data demonstrated the highest AUPRC in all tasks. Multiple models achieved similar performance when evaluated using AUROC. CONCLUSION: The LSTM architecture performs as well or better than some newer architectures, and PLE-DT may enhance the AUPRC in variable-length time series data for predicting clinical outcomes during external validation. Fereshteh S. Bashiri, Kyle A. Carey, Jennie Martin, Jay L. Koyner, Dana P. Edelson, Emily R. Gilbert, Anoop M. Mayampurath, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning modelsabstractOBJECTIVE: The timely stratification of trauma injury severity can enhance the quality of trauma care but it requires intense manual annotation from certified trauma coders. The objective of this study is to develop machine learning models for the stratification of trauma injury severity across various body regions using clinical text and structured electronic health records (EHRs) data. MATERIALS AND METHODS: Our study utilized clinical documents and structured EHR variables linked with the trauma registry data to create 2 machine learning models with different approaches to representing text. The first one fuses concept unique identifiers (CUIs) extracted from free text with structured EHR variables, while the second one integrates free text with structured EHR variables. Temporal validation was undertaken to ensure the models' temporal generalizability. Additionally, analyses to assess the variable importance were conducted. RESULTS: Both models demonstrated impressive performance in categorizing leg injuries, achieving high accuracy with macro-F1 scores of over 0.8. Additionally, they showed considerable accuracy, with macro-F1 scores exceeding or near 0.7, in assessing injuries in the areas of the chest and head. We showed in our variable importance analysis that the most important features in the model have strong face validity in determining clinically relevant trauma injuries. DISCUSSION: The CUI-based model achieves comparable performance, if not higher, compared to the free-text-based model, with reduced complexity. Furthermore, integrating structured EHR data improves performance, particularly when the text modalities are insufficiently indicative. CONCLUSIONS: Our multi-modal, multiclass models can provide accurate stratification of trauma injury severity and clinically relevant interpretations. Jifan Gao, Guanhua Chen 0002, Ann P. O'Rourke, John R. Caskey, Kyle A. Carey, Madeline Oguss, Anne Stey, Dmitriy Dligach, Timothy A. Miller, Anoop M. Mayampurath, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | Comparison of time series clustering methods for identifying novel subphenotypes of patients with infectionabstractOBJECTIVE: Severe infection can lead to organ dysfunction and sepsis. Identifying subphenotypes of infected patients is essential for personalized management. It is unknown how different time series clustering algorithms compare in identifying these subphenotypes. MATERIALS AND METHODS: Patients with suspected infection admitted between 2014 and 2019 to 4 hospitals in Emory healthcare were included, split into separate training and validation cohorts. Dynamic time warping (DTW) was applied to vital signs from the first 8 h of hospitalization, and hierarchical clustering (DTW-HC) and partition around medoids (DTW-PAM) were used to cluster patients into subphenotypes. DTW-HC, DTW-PAM, and a previously published group-based trajectory model (GBTM) were evaluated for agreement in subphenotype clusters, trajectory patterns, and subphenotype associations with clinical outcomes and treatment responses. RESULTS: There were 12 473 patients in training and 8256 patients in validation cohorts. DTW-HC, DTW-PAM, and GBTM models resulted in 4 consistent vitals trajectory patterns with significant agreement in clustering (71-80% agreement, P < .001): group A was hyperthermic, tachycardic, tachypneic, and hypotensive. Group B was hyperthermic, tachycardic, tachypneic, and hypertensive. Groups C and D had lower temperatures, heart rates, and respiratory rates, with group C normotensive and group D hypotensive. Group A had higher odds ratio of 30-day inpatient mortality (P < .01) and group D had significant mortality benefit from balanced crystalloids compared to saline (P < .01) in all 3 models. DISCUSSION: DTW- and GBTM-based clustering algorithms applied to vital signs in infected patients identified consistent subphenotypes with distinct clinical outcomes and treatment responses. CONCLUSION: Time series clustering with distinct computational approaches demonstrate similar performance and significant agreement in the resulting subphenotypes. Sivasubramanium Bhavani, Li Xiong 0001, Abish Pius, Matthew W. Semler, Edward T. Qian, Philip A. Verhoef, Chad Robichaux, Craig M. Coopersmith, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | DR.BENCH: Diagnostic Reasoning Benchmark for Clinical Natural Language Processing
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, John R. Caskey, Brihat Sharma, Matthew M. Churpek, Majid Afshar |
J. Biomed. Informatics | 6 |
| 2023 | Progress Note Understanding - Assessment and Plan Reasoning: Overview of the 2022 N2C2 Track 3 shared task
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Matthew M. Churpek, Özlem Uzuner, Majid Afshar |
J. Biomed. Informatics | 4 |
| 2022 | Explaining Alerts from a Pediatric Deterioration Prediction Model Using Clinical Text
Anoop M. Mayampurath, Kyle A. Carey, Priti Jani, Majid Afshar, Matthew M. Churpek, Dmitriy Dligach |
AMIA | 5 |
| 2022 | Summarizing Patients' Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence ModelsabstractAutomatically summarizing patients’ main problems from daily progress notes using natural language processing methods helps to battle against information and cognitive overload in hospital settings and potentially assists providers with computerized diagnostic decision support. Problem list summarization requires a model to understand, abstract, and generate clinical documentation. In this work, we propose a new NLP task that aims to generate a list of problems in a patient’s daily care plan using input from the provider’s progress notes during hospitalization. We investigate the performance of T5 and BART, two state-of-the-art seq2seq transformer architectures, in solving this problem. We provide a corpus built on top of progress notes from publicly available electronic health record progress notes in the Medical Information Mart for Intensive Care (MIMIC)-III. T5 and BART are trained on general domain text, and we experiment with a data augmentation method and a domain adaptation pre-training method to increase exposure to medical vocabulary and knowledge. Evaluation methods include ROUGE, BERTScore, cosine similarity on sentence embedding, and F-score on medical concepts. Results show that T5 with domain adaptive pre-training achieves significant performance gains compared to a rule-based system and general domain pre-trained language models, indicating a promising direction for tackling the problem summarization task. Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Dongfang Xu, Matthew M. Churpek, Majid Afshar |
COLING | 5 |
| 2022 | Hierarchical Annotation for Building A Suite of Clinical Natural Language Processing Tasks: Progress Note UnderstandingabstractApplying methods in natural language processing on electronic health records (EHR) data has attracted rising interests. Existing corpus and annotation focus on modeling textual features and relation prediction. However, there are a paucity of annotated corpus built to model clinical diagnostic thinking, a processing involving text understanding, domain knowledge abstraction and reasoning. In this work, we introduce a hierarchical annotation schema with three stages to address clinical text understanding, clinical reasoning and summarization. We create an annotated corpus based on a large collection of publicly available daily progress notes, a type of EHR that is time-sensitive, problem-oriented, and well-documented by the format of Subjective, Objective, Assessment and Plan (SOAP). We also define a new suite of tasks, Progress Note Understanding, with three tasks utilizing the three annotation stages. This new suite aims at training and evaluating future NLP models for clinical text understanding, clinical knowledge representation, inference and summarization. Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Samuel Tesch, Ryan Laffin, Matthew M. Churpek, Majid Afshar |
LREC | 6 |
| 2022 | Identifying infected patients using semi-supervised and transfer learningabstractOBJECTIVES: Early identification of infection improves outcomes, but developing models for early identification requires determining infection status with manual chart review, limiting sample size. Therefore, we aimed to compare semi-supervised and transfer learning algorithms with algorithms based solely on manual chart review for identifying infection in hospitalized patients. MATERIALS AND METHODS: This multicenter retrospective study of admissions to 6 hospitals included "gold-standard" labels of infection from manual chart review and "silver-standard" labels from nonchart-reviewed patients using the Sepsis-3 infection criteria based on antibiotic and culture orders. "Gold-standard" labeled admissions were randomly allocated to training (70%) and testing (30%) datasets. Using patient characteristics, vital signs, and laboratory data from the first 24 hours of admission, we derived deep learning and non-deep learning models using transfer learning and semi-supervised methods. Performance was compared in the gold-standard test set using discrimination and calibration metrics. RESULTS: The study comprised 432 965 admissions, of which 2724 underwent chart review. In the test set, deep learning and non-deep learning approaches had similar discrimination (area under the receiver operating characteristic curve of 0.82). Semi-supervised and transfer learning approaches did not improve discrimination over models fit using only silver- or gold-standard data. Transfer learning had the best calibration (unreliability index P value: .997, Brier score: 0.173), followed by self-learning gradient boosted machine (P value: .67, Brier score: 0.170). DISCUSSION: Deep learning and non-deep learning models performed similarly for identifying infection, as did models developed using Sepsis-3 and manual chart review labels. CONCLUSION: In a multicenter study of almost 3000 chart-reviewed patients, semi-supervised and transfer learning models showed similar performance for model discrimination as baseline XGBoost, while transfer learning improved calibration. Fereshteh S. Bashiri, John R. Caskey, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
J. Am. Medical Informatics Assoc. | 13 |
| 2022 | A scoping review of publicly available language tasks in clinical natural language processingabstractOBJECTIVE: To provide a scoping review of papers on clinical natural language processing (NLP) shared tasks that use publicly available electronic health record data from a cohort of patients. MATERIALS AND METHODS: We searched 6 databases, including biomedical research and computer science literature databases. A round of title/abstract screening and full-text screening were conducted by 2 reviewers. Our method followed the PRISMA-ScR guidelines. RESULTS: A total of 35 papers with 48 clinical NLP tasks met inclusion criteria between 2007 and 2021. We categorized the tasks by the type of NLP problems, including named entity recognition, summarization, and other NLP tasks. Some tasks were introduced as potential clinical decision support applications, such as substance abuse detection, and phenotyping. We summarized the tasks by publication venue and dataset type. DISCUSSION: The breadth of clinical NLP tasks continues to grow as the field of NLP evolves with advancements in language systems. However, gaps exist with divergent interests between the general domain NLP community and the clinical informatics community for task motivation and design, and in generalizability of the data sources. We also identified issues in data preparation. CONCLUSION: The existing clinical NLP tasks cover a wide range of topics and the field is expected to grow and attract more attention from both general domain NLP and clinical informatics community. We encourage future work to incorporate multidisciplinary collaboration, reporting transparency, and standardization in data preparation. We provide a listing of all the shared task papers and datasets from this review in a GitLab repository. Yanjun Gao, Dmitriy Dligach, Leslie Christensen, Samuel Tesch, Ryan Laffin, Dongfang Xu, Timothy A. Miller, Özlem Uzuner, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 9 |
| 2021 | Bias Assessment and Correction in Machine Learning Algorithms: A Use-Case in a Natural Language Processing Algorithm to Identify Hospitalized Patients with Unhealthy Alcohol Use
Marissa Borgese, Cara Joyce, Emily E. Anderson, Matthew M. Churpek, Majid Afshar |
AMIA | 4 |
| 2021 | Sepsis Prediction Using Semi-Supervised and Transfer Learning
John R. Caskey, Fereshteh S. Bashiri, Anoop M. Mayampurath, Nicole Dussault, Jay Dumanian, Sivasubramanium Bhavani, Kyle A. Carey, Emily R. Gilbert, Christopher J. Winslow, Nirav Shah 0004, Dana P. Edelson, Majid Afshar, Matthew M. Churpek |
AMIA | 13 |
| 2021 | A multi-label classifier to screen different types of substance misuse in hospitalized patients
Brihat Sharma, Dmitriy Dligach, Hale Thomson, Matthew M. Churpek, Niranjan S. Karnik, Majid Afshar |
AMIA | 4 |
| 2021 | The Addition of United States Census-Tract Data Does Not Improve the Prediction of Substance Misuse
Daniel To, Cara Joyce, Sujay Kulshrestha, Brihat Sharma, Dmitriy Dligach, Matthew M. Churpek, Majid Afshar |
AMIA | 6 |
| 2021 | Bias and fairness assessment of a natural language processing opioid misuse classifier: detection and mitigation of electronic health record data disadvantages across racial subgroupsabstractOBJECTIVES: To assess fairness and bias of a previously validated machine learning opioid misuse classifier. MATERIALS & METHODS: Two experiments were conducted with the classifier's original (n = 1000) and external validation (n = 53 974) datasets from 2 health systems. Bias was assessed via testing for differences in type II error rates across racial/ethnic subgroups (Black, Hispanic/Latinx, White, Other) using bootstrapped 95% confidence intervals. A local surrogate model was estimated to interpret the classifier's predictions by race and averaged globally from the datasets. Subgroup analyses and post-hoc recalibrations were conducted to attempt to mitigate biased metrics. RESULTS: We identified bias in the false negative rate (FNR = 0.32) of the Black subgroup compared to the FNR (0.17) of the White subgroup. Top features included "heroin" and "substance abuse" across subgroups. Post-hoc recalibrations eliminated bias in FNR with minimal changes in other subgroup error metrics. The Black FNR subgroup had higher risk scores for readmission and mortality than the White FNR subgroup, and a higher mortality risk score than the Black true positive subgroup (P < .05). DISCUSSION: The Black FNR subgroup had the greatest severity of disease and risk for poor outcomes. Similar features were present between subgroups for predicting opioid misuse, but inequities were present. Post-hoc mitigation techniques mitigated bias in type II error rate without creating substantial type I error rates. From model design through deployment, bias and data disadvantages should be systematically addressed. CONCLUSION: Standardized, transparent bias assessments are needed to improve trustworthiness in clinical machine learning models. Hale M. Thompson, Brihat Sharma, Sameer Bhalla, Randy Boley, Connor McCluskey, Dmitriy Dligach, Matthew M. Churpek, Niranjan S. Karnik, Majid Afshar |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | A Computable Phenotype for Acute Respiratory Distress Syndrome Using Natural Language Processing and Machine Learning
Majid Afshar, Cara Joyce, Anthony Oakey, Perry Formanek, Philip Yang, Matthew M. Churpek, Richard S. Cooper, Ron Price, Susan J. Zelisko, Dmitriy Dligach |
AMIA | 6 |
| 2016 | Development and validation of an electronic medical record-based alert score for detection of inpatient deterioration outside the ICU
Patricia Kipnis, Benjamin J. Turk, David A. Wulf, Juan Carlos LaGuardia, Vincent X. Liu, Matthew M. Churpek, Santiago Romero-Brufau, Gabriel J. Escobar |
J. Biomed. Informatics | 6 |